Files
accounted/lib/invoices/duplicate-payment-candidates.ts
T
a08bf51ced feat(reports): log behandlingsregler changes and program versions (BFNAR 2013:2 p. 9.16) (#2097)
* feat(reports): log behandlingsregler changes and program versions (BFNAR 2013:2 p. 9.16)

Part 3 of the behandlingshistorik series (#1787 report, #1790 PDF). BFNAR
2013:2 punkt 9.16 second paragraph requires the behandlingshistorik to record
"forandringar i bokforingssystemet som paverkar bokforingsposternas behandling
samt nar dessa forandringar infordes", and BFN's commentary names
behandlingsregler (automatkonteringar, fasta procentsatser) and new program
versions as the examples. Until now both changed without a trace.

Audit triggers on the behandlingsregler tables and the import logs:
mapping_rules, booking_template_library, categorization_templates,
salary_payroll_config, sie_imports, bank_file_imports. categorization_templates
learns on every booking (occurrence_count, confidence, last_seen_date), so
those telemetry-only updates are excluded by a WHEN clause the same way the
api_keys request counters are (20260721115701): only real rule changes are
logged. Measured against prod that is roughly 3 800 new audit rows a month
against an audit_log already taking 371 688, so about +1 %.

app_releases is an append-only log of program versions seen in production,
written by the runtime the first time a build answers a request. Vercel exposes
no build hook we can trust to write the row, so /api/version records it inside
after(): the handler returns synchronously and a floating promise could be
frozen before the insert lands, which is how a version log ends up silently
empty. The service client is constructed lazily so the constantly polled public
probe pays nothing once the module guard is set.

Program versions are rolled up per Swedish calendar day in the report. main
takes ~570 merges a month, so one event per version would be on the order of
7 000 a fiscal year: enough to trip the PDF's own 4 000-event guard and bury the
~400 events a real company's year contains. The statutory unit is the date, and
the same sentence qualifies the requirement to changes that affect processing,
which a deploy list cannot distinguish anyway. app_releases keeps the
per-version truth for anyone who needs to go deeper.

AuditLogEntry.user_id becomes string | null. The column is nullable and
write_audit_log() falls back to auth.uid(), which is NULL for a service-role or
global write; the company-less salary_payroll_config rows are the first that
routinely hit it, and the read model already coded for it.

Also restores the point citations the 2026-07-27 pass removed while the chapter
was unverified: it is kapitel 9, not kapitel 8 (which is arkivering), verified
against BFN's consolidated text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY

* test(pg): fix two fixture bugs in the behandlingshistorik trigger tests

pg-real caught both, and neither is in the migration: the inserts fail
before the trigger is reached.

mapping_rules.rule_type is constrained to mcc_code / merchant_name /
description_pattern / amount_threshold / combined; the test used
'merchant'.

booking_template_library's btl_insert policy requires
current_user_can_write() and company_id = current_active_company_id(),
so the authenticated insert needs a company_members row and a
user_preferences.active_company_id, the same setup
booking-template-hidden.pg.test.ts uses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY

* test(pg): assert the booking-template audit row inside the user transaction

withUserContext always rolls back, so the audit row the trigger writes
is gone before an outside connection can see it. The trigger fires in
the same transaction as the write, so the assertion belongs there too.
The other cases in this file write on the pool (autocommit) and are
unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY

* fix(reports): name every build id in the per-day program-version entry

Raised by the compliance review on #2097: the roll-up listed five ids
and a count, which leaves an auditor unable to reconstruct which
versions ran that day. app_releases keeps the full record, but the
report is the surface anyone actually reads. A day is bounded by the
deploy rate (~19), so the full list stays one readable cell.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3P2hr19PhQuCoTSGoegcY

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 20:31:10 +02:00

271 lines
9.1 KiB
TypeScript

import type { SupabaseClient } from '@supabase/supabase-js'
import {
DUPLICATE_AMOUNT_TOLERANCE_PCT,
DUPLICATE_DATE_WINDOW_DAYS,
escapeLikePattern,
normalizeOcrReference,
} from './duplicate-payment-guard'
import {
invoiceAmountSek,
magnitudesWithinTolerance,
normalizeCurrencyCode,
planAmountSweeps,
type ComparableAmount,
} from './duplicate-guard-currency'
import { resolveTransactionAmountSek } from '@/lib/transactions/booking-duplicate-detection'
import { createLogger } from '@/lib/logger'
const log = createLogger('invoices/duplicate-payment-candidates')
export type DuplicatePaymentMatchReason =
| 'ocr_exact'
| 'name_amount_fuzzy'
| 'amount_only'
export interface DuplicatePaymentCandidate {
id: string
date: string
amount: number
description: string | null
merchant_name: string | null
reference: string | null
match_reason: DuplicatePaymentMatchReason
match_confidence: number
}
const MATCH_REASON_RANK: Record<DuplicatePaymentMatchReason, number> = {
ocr_exact: 0,
name_amount_fuzzy: 1,
amount_only: 2,
}
const MATCH_REASON_CONFIDENCE: Record<DuplicatePaymentMatchReason, number> = {
ocr_exact: 0.99,
name_amount_fuzzy: 0.7,
amount_only: 0.5,
}
interface CustomerInvoice {
invoice_number: string | null
customer_name: string | null | undefined
/**
* `invoices.currency`; null means SEK (the column default). REQUIRED rather
* than optional on purpose: an optional field silently reads as SEK for any
* caller that forgets it, which is exactly how a 1 000 EUR payment came to be
* banded against a kronor column. TypeScript now refuses the call instead.
*/
currency: string | null
/** `invoices.total`, in `currency`. Pro-rates `total_sek` down to the payment. */
total: number | null
/** `invoices.total_sek`: the stored SEK view of `total`. */
total_sek: number | null
/** `invoices.exchange_rate`: SEK per unit of `currency`. */
exchange_rate: number | null
}
type Row = {
id: string
date: string
amount: number
description: string | null
merchant_name: string | null
reference: string | null
currency: string | null
amount_sek: number | null
exchange_rate: number | null
}
/**
* Scan unlinked positive (inbound) business bank transactions that could be
* the payment for this customer invoice. Used by the mark-paid duplicate
* guard: callers route the user to "link existing" instead of double-booking.
*
* Customer-side adaptations vs the supplier guard:
* - amount > 0 (inbound) instead of < 0
* - matches BOTH `merchant_name` AND `description` (banks often describe an
* inbound payment by payer name without populating merchant_name)
* - per-candidate scoring with OCR (invoice_number normalized) as the
* strongest signal
*
* Units: `paymentAmount` is denominated in the INVOICE's currency (that is what
* `invoices.remaining_amount` and `total` are stored in), while
* `transactions.amount` is denominated in the bank row's own currency. The
* plus-minus tolerance band is therefore planned per currency by
* `planAmountSweeps` and re-checked per row by `magnitudesWithinTolerance`:
* band and column always share a unit, and a candidate that cannot be brought
* into a shared unit is excluded rather than compared as a raw number. A SEK
* invoice produces exactly one sweep with the same band as before.
*
* The merchant_name and description searches are issued as two separate
* parameterised `.ilike()` queries and deduplicated by id. We deliberately
* avoid `.or('merchant_name.ilike.%X%,description.ilike.%X%')` because that
* interpolates the customer name into PostgREST's filter-DSL string, where
* `escapeLikePattern` only neutralises the LIKE wildcards (`%_\\`) and not
* the DSL chars (`,`, `.`, `(`, `)`). A name like `Acme,fake.eq.true` would
* otherwise inject a synthetic filter clause.
*/
export async function findDuplicatePaymentCandidatesForInvoice(
supabase: SupabaseClient,
params: {
companyId: string
invoice: CustomerInvoice
/** The payment being booked, in `invoice.currency`. */
paymentAmount: number
paymentDate: string
},
): Promise<DuplicatePaymentCandidate[]> {
const { companyId, invoice, paymentAmount, paymentDate } = params
const customerName = invoice.customer_name
if (!customerName) return []
const paymentCurrency = normalizeCurrencyCode(invoice.currency)
const reference: ComparableAmount = {
amount: paymentAmount,
currency: paymentCurrency,
sek: invoiceAmountSek({
amount: paymentAmount,
currency: paymentCurrency,
total: invoice.total,
totalSek: invoice.total_sek,
exchangeRate: invoice.exchange_rate,
}),
}
const { sweeps, crossCurrencyUnverifiable } = planAmountSweeps(
reference,
DUPLICATE_AMOUNT_TOLERANCE_PCT,
)
if (sweeps.length === 0) return []
if (crossCurrencyUnverifiable) {
// A foreign invoice with neither a usable total_sek nor an exchange_rate
// cannot be stated in kronor, so kronor bank rows are excluded rather than
// compared raw (a raw compare reads 1 000 kr as 1 000 EUR). Same-currency
// rows are still swept. Logged for the same reason the supplier-side twin
// logs it: an unevaluated candidate set is not a clean "no duplicate", and
// the gap must be visible in behandlingshistorik (BFNAR 2013:2 p. 9.16)
// rather than pass silently.
log.warn('duplicate-payment guard: cross-currency candidates not evaluated', {
reason: 'invoice_missing_sek_value',
companyId,
currency: paymentCurrency,
invoiceNumber: invoice.invoice_number,
})
}
const dateMs = new Date(paymentDate).getTime()
const dateLow = new Date(dateMs - DUPLICATE_DATE_WINDOW_DAYS * 24 * 3600 * 1000)
.toISOString()
.split('T')[0]
const dateHigh = new Date(dateMs + DUPLICATE_DATE_WINDOW_DAYS * 24 * 3600 * 1000)
.toISOString()
.split('T')[0]
const pattern = `%${escapeLikePattern(customerName)}%`
const base = (sweepIndex: number) => {
const sweep = sweeps[sweepIndex]
return supabase
.from('transactions')
.select(
'id, date, amount, description, merchant_name, reference, currency, amount_sek, exchange_rate',
)
.eq('company_id', companyId)
.eq('is_business', true)
.is('invoice_id', null)
.is('supplier_invoice_id', null)
.gt('amount', 0)
.or(sweep.currencyFilter)
.gte('amount', sweep.low)
.lte('amount', sweep.high)
.gte('date', dateLow)
.lte('date', dateHigh)
}
const responses = await Promise.all(
sweeps.flatMap((_sweep, i) => [
base(i).ilike('merchant_name', pattern).order('date', { ascending: false }).limit(5),
base(i).ilike('description', pattern).order('date', { ascending: false }).limit(5),
]),
)
const merged = new Map<string, Row>()
for (const res of responses) {
for (const row of (res.data ?? []) as Row[]) {
if (!merged.has(row.id)) merged.set(row.id, row)
}
}
const data = Array.from(merged.values())
.filter((row) =>
magnitudesWithinTolerance(reference, rowAmount(row), DUPLICATE_AMOUNT_TOLERANCE_PCT),
)
.sort((a, b) => (a.date < b.date ? 1 : a.date > b.date ? -1 : 0))
.slice(0, 5)
if (data.length === 0) return []
const invoiceOcr = normalizeOcrReference(invoice.invoice_number)
const searchTerms = customerName
.toLowerCase()
.split(/\s+/)
.filter((term) => term.length > 2)
const candidates: DuplicatePaymentCandidate[] = data.map((row) => {
const reason = scoreCandidate({
row,
invoiceOcr,
searchTerms,
})
return {
id: row.id,
date: row.date,
amount: row.amount,
description: row.description,
merchant_name: row.merchant_name,
reference: row.reference,
match_reason: reason,
match_confidence: MATCH_REASON_CONFIDENCE[reason],
}
})
candidates.sort((a, b) => MATCH_REASON_RANK[a.match_reason] - MATCH_REASON_RANK[b.match_reason])
return candidates
}
/**
* A bank row as a comparable amount. `resolveTransactionAmountSek` is the one
* definition of "this bank line in kronor" (shared with the booking-time
* duplicate guard) and returns null rather than falling back to the raw foreign
* number.
*/
function rowAmount(row: Row): ComparableAmount {
return {
amount: Number(row.amount),
currency: normalizeCurrencyCode(row.currency),
sek: resolveTransactionAmountSek({
amount: row.amount,
currency: row.currency,
amount_sek: row.amount_sek,
exchange_rate: row.exchange_rate,
}),
}
}
function scoreCandidate(args: {
row: { reference: string | null; description: string | null; merchant_name: string | null }
invoiceOcr: string
searchTerms: string[]
}): DuplicatePaymentMatchReason {
const { row, invoiceOcr, searchTerms } = args
if (invoiceOcr && row.reference) {
if (normalizeOcrReference(row.reference) === invoiceOcr) {
return 'ocr_exact'
}
}
if (searchTerms.length > 0) {
const haystack = `${row.description ?? ''} ${row.merchant_name ?? ''}`.toLowerCase()
if (searchTerms.some((term) => haystack.includes(term))) {
return 'name_amount_fuzzy'
}
}
return 'amount_only'
}