* feat(reports): log behandlingsregler changes and program versions (BFNAR 2013:2 p. 9.16) Part 3 of the behandlingshistorik series (#1787 report, #1790 PDF). BFNAR 2013:2 punkt 9.16 second paragraph requires the behandlingshistorik to record "forandringar i bokforingssystemet som paverkar bokforingsposternas behandling samt nar dessa forandringar infordes", and BFN's commentary names behandlingsregler (automatkonteringar, fasta procentsatser) and new program versions as the examples. Until now both changed without a trace. Audit triggers on the behandlingsregler tables and the import logs: mapping_rules, booking_template_library, categorization_templates, salary_payroll_config, sie_imports, bank_file_imports. categorization_templates learns on every booking (occurrence_count, confidence, last_seen_date), so those telemetry-only updates are excluded by a WHEN clause the same way the api_keys request counters are (20260721115701): only real rule changes are logged. Measured against prod that is roughly 3 800 new audit rows a month against an audit_log already taking 371 688, so about +1 %. app_releases is an append-only log of program versions seen in production, written by the runtime the first time a build answers a request. Vercel exposes no build hook we can trust to write the row, so /api/version records it inside after(): the handler returns synchronously and a floating promise could be frozen before the insert lands, which is how a version log ends up silently empty. The service client is constructed lazily so the constantly polled public probe pays nothing once the module guard is set. Program versions are rolled up per Swedish calendar day in the report. main takes ~570 merges a month, so one event per version would be on the order of 7 000 a fiscal year: enough to trip the PDF's own 4 000-event guard and bury the ~400 events a real company's year contains. The statutory unit is the date, and the same sentence qualifies the requirement to changes that affect processing, which a deploy list cannot distinguish anyway. app_releases keeps the per-version truth for anyone who needs to go deeper. AuditLogEntry.user_id becomes string | null. The column is nullable and write_audit_log() falls back to auth.uid(), which is NULL for a service-role or global write; the company-less salary_payroll_config rows are the first that routinely hit it, and the read model already coded for it. Also restores the point citations the 2026-07-27 pass removed while the chapter was unverified: it is kapitel 9, not kapitel 8 (which is arkivering), verified against BFN's consolidated text. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY * test(pg): fix two fixture bugs in the behandlingshistorik trigger tests pg-real caught both, and neither is in the migration: the inserts fail before the trigger is reached. mapping_rules.rule_type is constrained to mcc_code / merchant_name / description_pattern / amount_threshold / combined; the test used 'merchant'. booking_template_library's btl_insert policy requires current_user_can_write() and company_id = current_active_company_id(), so the authenticated insert needs a company_members row and a user_preferences.active_company_id, the same setup booking-template-hidden.pg.test.ts uses. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY * test(pg): assert the booking-template audit row inside the user transaction withUserContext always rolls back, so the audit row the trigger writes is gone before an outside connection can see it. The trigger fires in the same transaction as the write, so the assertion belongs there too. The other cases in this file write on the pool (autocommit) and are unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY * fix(reports): name every build id in the per-day program-version entry Raised by the compliance review on #2097: the roll-up listed five ids and a count, which leaves an auditor unable to reconstruct which versions ran that day. app_releases keeps the full record, but the report is the surface anyone actually reads. A day is bounded by the deploy rate (~19), so the full list stays one readable cell. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
287 lines
11 KiB
TypeScript
287 lines
11 KiB
TypeScript
/**
|
|
* Turning a mailbox hit into an underlag the user can approve.
|
|
*
|
|
* Lives in core rather than in the mail extension because it writes documents
|
|
* and inbox items, and an extension may never import another extension. The
|
|
* mail extension only ever hands over bytes.
|
|
*
|
|
* The hunt does NOT book anything and does not link anything by itself: it
|
|
* stores the receipt, records where it came from, and stages the pairing. The
|
|
* document becomes räkenskapsinformation only when a human approves.
|
|
*/
|
|
import type { SupabaseClient } from '@supabase/supabase-js'
|
|
import { uploadDocument } from '@/lib/core/documents/document-service'
|
|
import { appendProcessingHistory } from '@/lib/processing-history/append'
|
|
import { getMailSearchService, type MailCandidate } from '@/lib/mail-search/service'
|
|
import { createLogger } from '@/lib/logger'
|
|
|
|
const log = createLogger('receipt-hunt-ingest')
|
|
|
|
/**
|
|
* What the bytes actually are, rather than what the mail claims.
|
|
*
|
|
* A mail's declared content type is untrusted metadata. Forwarded receipts
|
|
* routinely arrive as `application/octet-stream` whatever they really are, and
|
|
* uploadDocument validates the content against the type it is given, so
|
|
* trusting the mail means every such receipt is rejected at the door. Measured
|
|
* on a real mailbox: the first live fetch, an Elgiganten PDF, failed exactly
|
|
* this way.
|
|
*
|
|
* Magic bytes first, then the filename, then whatever the mail said.
|
|
*/
|
|
export function sniffMimeType(bytes: Buffer, declared: string, filename: string): string {
|
|
const head = bytes.subarray(0, 12)
|
|
if (head.subarray(0, 4).toString('latin1') === '%PDF') return 'application/pdf'
|
|
if (head[0] === 0xff && head[1] === 0xd8 && head[2] === 0xff) return 'image/jpeg'
|
|
if (head.subarray(0, 8).toString('latin1') === '\x89PNG\r\n\x1a\n') return 'image/png'
|
|
if (head.subarray(0, 4).toString('latin1') === 'GIF8') return 'image/gif'
|
|
if (
|
|
head.subarray(0, 4).toString('latin1') === 'RIFF' &&
|
|
bytes.subarray(8, 12).toString('latin1') === 'WEBP'
|
|
) {
|
|
return 'image/webp'
|
|
}
|
|
|
|
const ext = filename.toLowerCase().match(/\.([a-z0-9]+)$/)?.[1]
|
|
const byExt: Record<string, string> = {
|
|
pdf: 'application/pdf',
|
|
jpg: 'image/jpeg',
|
|
jpeg: 'image/jpeg',
|
|
png: 'image/png',
|
|
gif: 'image/gif',
|
|
webp: 'image/webp',
|
|
}
|
|
if (ext && byExt[ext]) return byExt[ext]
|
|
|
|
return declared
|
|
}
|
|
|
|
/** Largest attachment worth pulling. Receipts are small; anything larger is a report. */
|
|
const MAX_ATTACHMENT_BYTES = 10 * 1024 * 1024
|
|
|
|
export interface IngestedReceipt {
|
|
documentId: string
|
|
inboxItemId: string
|
|
fileName: string
|
|
mailbox: string
|
|
}
|
|
|
|
/**
|
|
* Provenance written onto the inbox item.
|
|
*
|
|
* Deliberately in `channel_context` and not in `extracted_data`: retrying
|
|
* extraction overwrites extracted_data wholesale, and the record of which
|
|
* mailbox a receipt came from must survive that. Same rule the WhatsApp intake
|
|
* follows.
|
|
*/
|
|
function buildChannelContext(candidate: MailCandidate, attachmentId: string) {
|
|
return {
|
|
channel: 'mail_hunt',
|
|
mail_message_id: candidate.messageId,
|
|
mail_attachment_id: attachmentId,
|
|
// Message + attachment, because one forward can carry receipts for several
|
|
// different purchases and each must be able to land separately. Taken from
|
|
// the attachment being stored, not from index 0: filing a later attachment
|
|
// under its sibling's key would block the sibling from ever landing.
|
|
mail_file_key: `${candidate.messageId}::${attachmentId}`,
|
|
mail_mailbox: candidate.mailbox,
|
|
mail_provider: candidate.provider,
|
|
mail_subject: candidate.subject,
|
|
mail_from: candidate.from,
|
|
mail_received_at: candidate.receivedAt,
|
|
fetched_at: new Date().toISOString(),
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Fetch the first usable attachment on a candidate and file it as an inbox item.
|
|
*
|
|
* Returns null when there is nothing to store (body-only receipt, oversized
|
|
* attachment, or a duplicate we have already ingested). Never throws for one
|
|
* bad message: a single unreadable attachment must not abort a night's hunt.
|
|
*/
|
|
export async function ingestMailCandidate(
|
|
supabase: SupabaseClient,
|
|
companyId: string,
|
|
userId: string,
|
|
candidate: MailCandidate,
|
|
/**
|
|
* What this document is paperwork for, as the reading model saw it.
|
|
*
|
|
* Stored so a later run compares like with like. Deriving it again from the
|
|
* extraction would compare the model's "Norwegian" against the PDF's
|
|
* "Norwegian Air Shuttle AOC AS" and conclude they are two suppliers, which
|
|
* is how an already-held receipt got fetched a second time.
|
|
*/
|
|
receiptIdentity?: string,
|
|
): Promise<IngestedReceipt | null> {
|
|
if (candidate.attachmentIds.length === 0) return null
|
|
|
|
const service = getMailSearchService()
|
|
|
|
for (const [index, attachmentId] of candidate.attachmentIds.entries()) {
|
|
// Per attachment, not per message: the check has to name the file it is
|
|
// about, and it sits inside the loop so trying a second attachment is not
|
|
// suppressed by the first one already being filed.
|
|
const fileKey = `${candidate.messageId}::${attachmentId}`
|
|
const { data: existing } = await supabase
|
|
.from('invoice_inbox_items')
|
|
.select('id')
|
|
.eq('company_id', companyId)
|
|
.eq('source', 'mail_hunt')
|
|
.eq('channel_context->>mail_file_key', fileKey)
|
|
.maybeSingle()
|
|
if (existing) continue
|
|
|
|
let fetched
|
|
try {
|
|
fetched = await service.fetchAttachment(candidate.connectionId, candidate.messageId, attachmentId)
|
|
} catch (error) {
|
|
log.warn('could not fetch attachment', {
|
|
messageId: candidate.messageId,
|
|
error: error instanceof Error ? error.message : String(error),
|
|
})
|
|
continue
|
|
}
|
|
if (!fetched) continue
|
|
if (fetched.bytes.byteLength > MAX_ATTACHMENT_BYTES) continue
|
|
|
|
try {
|
|
// The name the search already reported beats the one the provider
|
|
// re-derives on fetch: a second lookup can come back empty and fall back
|
|
// to a generic "underlag.pdf", throwing away "2332687551.pdf".
|
|
const knownName = candidate.attachmentNames?.[index]
|
|
const fileName = knownName && knownName.length > 0 ? knownName : fetched.filename
|
|
|
|
const document = await uploadDocument(
|
|
supabase,
|
|
userId,
|
|
companyId,
|
|
{
|
|
name: fileName,
|
|
buffer: fetched.bytes.buffer.slice(
|
|
fetched.bytes.byteOffset,
|
|
fetched.bytes.byteOffset + fetched.bytes.byteLength,
|
|
) as ArrayBuffer,
|
|
type: sniffMimeType(fetched.bytes, fetched.mimeType, fileName),
|
|
},
|
|
{ upload_source: 'mail_hunt', dedupeByContent: true },
|
|
)
|
|
|
|
// Content already archived for this company: the provenance key above
|
|
// only catches the SAME message re-hunted, while this catches the same
|
|
// receipt arriving through another inbox or channel (forwards are
|
|
// common). Skip ONLY when an inbox item already carries the document:
|
|
// then the receipt is in the Underlag flow (or handled). When the match
|
|
// is a document that never passed the inbox (a manually attached copy,
|
|
// an archival file), fall through and file an item for the EXISTING
|
|
// document: silently dropping the receipt could leave a real
|
|
// affärshändelse without underlag routing (BFL 5 kap).
|
|
if (document.deduplicated) {
|
|
const { data: dupItem, error: dupErr } = await supabase
|
|
.from('invoice_inbox_items')
|
|
.select('id')
|
|
.eq('company_id', companyId)
|
|
.eq('document_id', document.id)
|
|
.limit(1)
|
|
.maybeSingle()
|
|
// Fail closed: a broken lookup must not file a second item.
|
|
if (dupErr) throw new Error(dupErr.message)
|
|
if (dupItem) {
|
|
// Behandlingshistorik, not just an app log: the dedupe decision is
|
|
// part of the auditable trail (BFNAR 2013:2 p. 9.16).
|
|
try {
|
|
await appendProcessingHistory({
|
|
companyId,
|
|
correlationId: crypto.randomUUID(),
|
|
aggregateType: 'Document',
|
|
aggregateId: document.id,
|
|
eventType: 'DocumentDuplicateSkipped',
|
|
// No mailbox address here: the processing-history payload
|
|
// contract is pseudonymous IDs only (never emails). Which
|
|
// mailbox first delivered the receipt is on the existing
|
|
// item's channel_context.
|
|
payload: {
|
|
channel: 'mail_hunt',
|
|
document_id: document.id,
|
|
inbox_item_id: (dupItem as { id: string }).id,
|
|
mail_message_id: candidate.messageId,
|
|
reason: 'duplicate_content',
|
|
},
|
|
actor: { type: 'system', id: 'receipt-hunt' },
|
|
occurredAt: new Date(),
|
|
})
|
|
} catch (histErr) {
|
|
log.warn('could not append DocumentDuplicateSkipped', {
|
|
error: histErr instanceof Error ? histErr.message : String(histErr),
|
|
})
|
|
}
|
|
log.info('skipped duplicate attachment content', {
|
|
messageId: candidate.messageId,
|
|
documentId: document.id,
|
|
})
|
|
continue
|
|
}
|
|
}
|
|
|
|
// uploadDocument emits document.uploaded and awaits its handlers, so the
|
|
// extraction extension has already read the amount, date and vendor out
|
|
// of this file by the time we get here. Copying it onto the inbox item is
|
|
// what lets the deterministic matcher pair the receipt on its amount:
|
|
// the pool is read from invoice_inbox_items, and a row with no
|
|
// extracted_data can never match anything.
|
|
const { data: extractedRow } = await supabase
|
|
.from('document_attachments')
|
|
.select('extracted_data')
|
|
.eq('id', document.id)
|
|
.maybeSingle()
|
|
const extracted = (extractedRow as { extracted_data?: Record<string, unknown> } | null)
|
|
?.extracted_data
|
|
|
|
const { data: item, error } = await supabase
|
|
.from('invoice_inbox_items')
|
|
.insert({
|
|
company_id: companyId,
|
|
user_id: userId,
|
|
document_id: document.id,
|
|
source: 'mail_hunt',
|
|
status: 'received',
|
|
email_from: candidate.from,
|
|
email_subject: candidate.subject,
|
|
email_received_at: candidate.receivedAt,
|
|
extracted_data: extracted ?? null,
|
|
channel_context: {
|
|
...buildChannelContext(candidate, attachmentId),
|
|
...(receiptIdentity ? { receipt_identity: receiptIdentity } : {}),
|
|
},
|
|
})
|
|
.select('id')
|
|
.single()
|
|
|
|
if (error) {
|
|
// 23505 is the partial unique index doing its job: another run got
|
|
// there first, which is a success from the caller's point of view.
|
|
if (error.code === '23505') return null
|
|
throw new Error(error.message)
|
|
}
|
|
|
|
return {
|
|
documentId: document.id,
|
|
inboxItemId: (item as { id: string }).id,
|
|
fileName,
|
|
mailbox: candidate.mailbox,
|
|
}
|
|
} catch (error) {
|
|
log.warn('could not store hunted receipt', {
|
|
messageId: candidate.messageId,
|
|
error: error instanceof Error ? error.message : String(error),
|
|
})
|
|
// Magic-byte rejection and the like: try the next attachment rather than
|
|
// failing the whole candidate.
|
|
continue
|
|
}
|
|
}
|
|
|
|
return null
|
|
}
|