feat(inbox): match non-invoice documents via prominent amounts (#2048)

* feat(inbox): match non-invoice documents via prominent amounts

Bankintyg, bank agreements and other documentKind "other" PDFs carry no
invoice-style total, so extraction correctly left totals.total null and the
document became structurally unmatchable: findUnderlagCandidates hard-drops
items without a comparable amount and the picker lost the 40% amount signal.

- extraction: new prominentAmounts[] field (amount + document's own label),
  populated only when totals.total is null; account/org/phone/reference
  numbers and zero amounts excluded. totals.total semantics untouched.
- matching: bestProminentAmountVariance() tries each printed amount and
  feeds calculateMatchConfidence at reduced weight (0.3 vs 0.4) in both the
  agent candidate scorer and TransactionMatchPicker.
- UI: inbox rail shows the detected amounts read-only for such documents,
  list falls back to a single distinct prominent amount, and extraction no
  longer reads as "found nothing".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

* fix(inbox): discount prominent-amount fallback instead of reweighting it

Skeptic pass refutations on the first commit: normalized weighting made a
reduced amount weight self-defeating. Date + exact fallback amount with no
merchant scored (0.25+0.3)/0.55 = 1.0 ("100% sakerhet" on a wrong same-day
transaction), and a DISAGREEING fallback amount scored above a disagreeing
invoice total (0.67 vs 0.60) because shrinking the weight also shrank the
penalty.

- score fallbacks at full amount weight, then multiply by a flat
  FALLBACK_CONFIDENCE_FACTOR (0.85): agreement caps below certainty,
  disagreement stays at least as damning as for a real total.
- agent candidate surface additionally requires the document date within
  DATE_TOLERANCE_DAYS, so an avtal listing 349 kr no longer matches every
  future 349 kr charge from the same counterparty.
- bestProminentAmountVariance returns which amount matched + its document
  label, and the match reason names it ("Exakt belopp i dokumentet: 2 500
  SEK (Engangspris)"): no more bare "Exakt belopp" reaching the agent while
  total_amount is null.
- prompt: prominentAmounts restricted to non-invoice documentKinds, and
  never a parking spot for an unreadable invoice total.
- fix the stale "deliberately the same list" comment on
  EXTRACTED_FIELD_ACCESSORS.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

* fix(receipt-hunt): never propose on the prominent-amounts fallback

Second skeptic pass: the nightly hunt is a third consumer of
scoreUnderlagCandidates and inherited the fallback unaware. A bankintyg
whose printed "Insatt belopp" equals a same-day outflow scores 0.85, which
clears CERTAIN_CONFIDENCE (0.8) and skips LLM adjudication, on a pairing
wrong by construction (the hunt scans outflows only; "Insatt belopp"
labels an inflow), with document_amount null in the approval preview.

UnderlagCandidate now carries amountSource ('total' | 'prominent') and
selectProposals drops fallback-scored candidates. Non-invoice documents
stay reachable through the manual picker and the agent candidate surface,
both of which have a human reading the amounts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

* fix(inbox): round fallback confidence via roundOre, not the naive pattern

The two confidence discounts (and their test) tripped the naive-ore-round
antipattern ratchet (625 vs baseline 622); use the sanctioned helper.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Mattsson
2026-08-30 23:52:29 +02:00
committed by GitHub
co-authored by Claude Fable 5
parent 84ba8323b4
commit 516e8b62ff
14 changed files with 493 additions and 19 deletions
@@ -15,6 +15,7 @@ function extraction(partial: {
total?: number | null
vat?: number | null
currency?: string
prominentAmounts?: { amount: number; label: string | null }[]
}): InvoiceExtractionResult {
return {
supplier: {
@@ -28,6 +29,7 @@ function extraction(partial: {
lineItems: [],
totals: { subtotal: null, vatAmount: partial.vat ?? null, total: partial.total ?? null },
vatBreakdown: [],
prominentAmounts: partial.prominentAmounts,
confidence: 0.9,
} as InvoiceExtractionResult
}
@@ -71,6 +73,7 @@ describe('scoreUnderlagCandidates', () => {
expect(out).toHaveLength(1)
expect(out[0].inbox_item_id).toBe('item-1')
expect(out[0].confidence).toBeGreaterThanOrEqual(CANDIDATE_MIN_CONFIDENCE)
expect(out[0].amountSource).toBe('total')
// and it brings the captured answers along with it
expect(out[0].channelContext?.representation?.purpose).toBe('kundmöte')
})
@@ -139,6 +142,110 @@ describe('scoreUnderlagCandidates', () => {
expect(out[0].inbox_item_id).toBe('strong')
})
it('proposes a non-invoice document via its prominent amounts', () => {
// The Robotministeriet case: an SEB account agreement (documentKind
// "other") has no "Att betala" total, only "Anslutnings-/Engångspris
// 2 500". The bank charges AVGIFT -2500 the same day. Before the
// prominentAmounts fallback this document was structurally unmatchable.
const avgiftTx = {
...tx,
description: 'AVGIFT',
merchant_name: null,
amount: -2500,
date: '2026-08-26',
}
const out = scoreUnderlagCandidates(avgiftTx, [
{
id: 'item-avtal',
document_id: 'doc-avtal',
extracted_data: extraction({
supplier: 'SEB',
date: '2026-08-26',
total: null,
prominentAmounts: [
{ amount: 2500, label: 'Anslutnings-/Engångspris' },
],
}),
channel_context: null,
},
])
expect(out).toHaveLength(1)
expect(out[0].inbox_item_id).toBe('item-avtal')
expect(out[0].confidence).toBeGreaterThanOrEqual(CANDIDATE_MIN_CONFIDENCE)
// ...but never as certainty: a printed figure is not an invoice total.
expect(out[0].confidence).toBeLessThan(1)
// Tagged so low-scrutiny consumers (the nightly hunt) can exclude it.
expect(out[0].amountSource).toBe('prominent')
// The reason names WHICH figure matched, since total_amount stays null.
expect(out[0].total_amount).toBeNull()
// toLocaleString('sv-SE') groups with a non-breaking space (U+00A0).
expect(out[0].matchReasons.join(' ')).toContain(`2${' '}500`)
expect(out[0].matchReasons.join(' ')).toContain('Anslutnings-/Engångspris')
})
it('does not let a prominent amount alone carry a dateless document over the floor', () => {
// Amount agreement without a date is weaker than a total + date pair; the
// candidate surface trades recall for precision, so this stays in the
// manual picker only.
const out = scoreUnderlagCandidates({ ...tx, amount: -25000 }, [
{
id: 'item-intyg',
document_id: 'doc-intyg',
extracted_data: extraction({
supplier: 'SEB',
date: null,
total: null,
prominentAmounts: [{ amount: 25000, label: 'Insatt belopp' }],
}),
channel_context: null,
},
])
expect(out).toEqual([])
})
it('does not match an avtal to a later charge on amount + merchant alone', () => {
// A Telia avtal listing 349 kr must not surface for every future 349 kr
// Telia charge: the fallback requires the document date to agree within
// the normal tolerance.
const out = scoreUnderlagCandidates(
{ ...tx, description: 'TELIA SVERIGE AB', merchant_name: 'Telia Sverige AB', amount: -349 },
[
{
id: 'item-telia',
document_id: 'doc-telia',
extracted_data: extraction({
supplier: 'Telia Sverige AB',
date: '2026-01-15',
total: null,
prominentAmounts: [{ amount: 349, label: 'Månadspris' }],
}),
channel_context: null,
},
],
)
expect(out).toEqual([])
})
it('rejects a non-invoice document whose prominent amounts all disagree', () => {
// Same-day, same merchant, but the printed amounts match nothing: the
// discount keeps this under the floor, where a real disagreeing invoice
// total would sit exactly at it.
const out = scoreUnderlagCandidates(tx, [
{
id: 'item-wrong',
document_id: 'doc-wrong',
extracted_data: extraction({
supplier: 'Espresso House',
date: '2026-05-12',
total: null,
prominentAmounts: [{ amount: 9999, label: 'Pris' }],
}),
channel_context: null,
},
])
expect(out).toEqual([])
})
it('returns nothing for a transaction with no date or amount', () => {
expect(
scoreUnderlagCandidates({ ...tx, date: null }, [
+61 -9
View File
@@ -24,11 +24,15 @@
import type { SupabaseClient } from '@supabase/supabase-js'
import {
CONVERTED_AMOUNT_TOLERANCE_PERCENT,
DATE_TOLERANCE_DAYS,
FALLBACK_CONFIDENCE_FACTOR,
amountVarianceForMatch,
bestProminentAmountVariance,
calculateMatchConfidence,
calculateMerchantSimilarity,
} from '@/lib/documents/core-receipt-matcher'
import { resolveSekAmount } from '@/lib/bookkeeping/currency-utils'
import { roundOre } from '@/lib/money'
import type { InboxChannelContext, InvoiceExtractionResult } from '@/types'
/**
@@ -62,6 +66,13 @@ export interface UnderlagCandidate {
currency: string | null
/** 0-1 from the shared receipt matcher. */
confidence: number
/**
* Where the amount signal came from: an invoice-style total, or the
* prominent-amounts fallback for non-invoice documents (bankintyg, avtal).
* Consumers that act with less human scrutiny (the nightly receipt hunt)
* must treat 'prominent' as weaker evidence or exclude it.
*/
amountSource: 'total' | 'prominent'
/** Swedish reasons the match scored, for display. */
matchReasons: string[]
/** Answers already captured for this item, so they travel with it. */
@@ -103,6 +114,11 @@ function extractionSignals(extracted: InvoiceExtractionResult | null | undefined
total: extracted?.totals?.total ?? null,
vat: extracted?.totals?.vatAmount ?? null,
currency: (extracted?.invoice?.currency || 'SEK').toUpperCase(),
// Non-invoice documents (bankintyg, avtal) carry no total but often show
// the money amount anyway; the extractor lists those here.
prominentAmounts: (extracted?.prominentAmounts ?? []).filter(
(a) => Number.isFinite(a.amount) && a.amount !== 0,
),
}
}
@@ -128,11 +144,11 @@ export function scoreUnderlagCandidates(
for (const item of items) {
const sig = extractionSignals(item.extracted_data)
// An extraction with neither a date nor a total carries no signal the
// An extraction with neither a date nor any amount carries no signal the
// matcher can use; scoring it returns noise dressed as confidence.
if (!sig.date && sig.total == null) continue
if (!sig.date && sig.total == null && sig.prominentAmounts.length === 0) continue
const amountVariance = amountVarianceForMatch(
let amountVariance = amountVarianceForMatch(
sig.total,
sig.currency,
// A SEK value only when someone resolved a rate for this receipt.
@@ -144,6 +160,29 @@ export function scoreUnderlagCandidates(
txSek,
)
const dateVariance = sig.date
? Math.abs((new Date(sig.date).getTime() - txDateMs) / (1000 * 60 * 60 * 24))
: Number.POSITIVE_INFINITY
// A document with no invoice-style total (bankintyg, avtal: documentKind
// "other") but visible amounts falls back to the closest prominent
// amount. Two guards keep this precision-first: the document's date must
// agree within the normal tolerance (an avtal listing 349 kr must not
// match every future 349 kr charge from the same counterparty on amount +
// merchant alone), and the confidence is discounted below so a fallback
// can never present as certainty.
const fallbackMatch =
sig.total == null && amountVariance == null && dateVariance <= DATE_TOLERANCE_DAYS
? bestProminentAmountVariance(
sig.prominentAmounts,
sig.currency,
tx.amount,
txCurrency,
txSek,
)
: null
if (fallbackMatch) amountVariance = fallbackMatch.variance
// No comparable amount means no candidate. calculateMatchConfidence drops
// the amount signal when it cannot normalise the currencies, which leaves
// date + merchant carrying the whole normalised score: a same-day receipt
@@ -154,21 +193,33 @@ export function scoreUnderlagCandidates(
// through the picker; they are just not proposed.
if (amountVariance == null) continue
const dateVariance = sig.date
? Math.abs((new Date(sig.date).getTime() - txDateMs) / (1000 * 60 * 60 * 24))
: Number.POSITIVE_INFINITY
const similarity = sig.supplier ? calculateMerchantSimilarity(sig.supplier, txMerchant) : 0
// A converted total is judged against the wider bar, because the rate
// spread is a known error rather than a disagreement about the sum.
const converted = sig.currency !== txCurrency && item.sek_total != null
const { confidence, matchReasons } = calculateMatchConfidence(
const scoredMatch = calculateMatchConfidence(
dateVariance,
amountVariance,
similarity,
undefined,
converted ? CONVERTED_AMOUNT_TOLERANCE_PERCENT : undefined,
sig.currency !== txCurrency && item.sek_total != null
? CONVERTED_AMOUNT_TOLERANCE_PERCENT
: undefined,
)
let confidence = scoredMatch.confidence
let matchReasons = scoredMatch.matchReasons
if (fallbackMatch) {
confidence = roundOre(confidence * FALLBACK_CONFIDENCE_FACTOR)
// Name the figure that matched. A bare "Exakt belopp" would reach the
// agent while total_amount stays null: certainty without a number the
// agent or the user could check against the document.
const label = fallbackMatch.label ? ` (${fallbackMatch.label})` : ''
matchReasons = matchReasons.map((reason) =>
reason.startsWith('Exakt belopp') || reason.startsWith('Belopp ±')
? `${reason} i dokumentet: ${fallbackMatch.amount.toLocaleString('sv-SE')} ${sig.currency}${label}`
: reason,
)
}
if (confidence < CANDIDATE_MIN_CONFIDENCE) continue
scored.push({
@@ -180,6 +231,7 @@ export function scoreUnderlagCandidates(
vat_amount: sig.vat,
currency: sig.currency,
confidence,
amountSource: fallbackMatch ? 'prominent' : 'total',
matchReasons,
channelContext: item.channel_context ?? null,
})
@@ -1,12 +1,15 @@
import { describe, it, expect } from 'vitest'
import {
FALLBACK_CONFIDENCE_FACTOR,
levenshteinDistance,
normalizeMerchantName,
normalizeForMatch,
calculateMerchantSimilarity,
calculateMatchConfidence,
amountVarianceForMatch,
bestProminentAmountVariance,
} from '../core-receipt-matcher'
import { roundOre } from '@/lib/money'
describe('levenshteinDistance', () => {
it('returns 0 for identical strings', () => {
@@ -163,6 +166,56 @@ describe('amountVarianceForMatch', () => {
})
})
describe('bestProminentAmountVariance', () => {
it('picks the closest of several printed amounts and names it', () => {
// An agreement listing both a monthly price and a one-off price: the
// one-off 2500 matches the -2500 AVGIFT charge exactly.
const best = bestProminentAmountVariance(
[
{ amount: 49, label: 'Månadspris' },
{ amount: 2500, label: 'Engångspris' },
],
'SEK',
-2500,
'SEK',
-2500,
)
expect(best).toEqual({ variance: 0, amount: 2500, label: 'Engångspris' })
})
it('returns null when nothing is comparable', () => {
expect(bestProminentAmountVariance([], 'SEK', -2500, 'SEK', -2500)).toBeNull()
// Cross-currency without a rate stays incomparable, like a total would.
expect(
bestProminentAmountVariance([{ amount: 2500, label: null }], 'EUR', -2500, 'SEK', -2500),
).toBeNull()
// Zero amounts carry no signal (amountVarianceForMatch drops them).
expect(
bestProminentAmountVariance([{ amount: 0, label: null }], 'SEK', -2500, 'SEK', -2500),
).toBeNull()
})
it('the discount factor keeps fallback agreement below certainty', () => {
// Exact date + exact amount + no merchant normalises to 1.0; a fallback
// match must not present that as certainty (this exact geometry scored
// "100% säkerhet" on a wrong same-day transaction before the factor).
const { confidence } = calculateMatchConfidence(0, 0, 0)
const discounted = roundOre(confidence * FALLBACK_CONFIDENCE_FACTOR)
expect(confidence).toBe(1)
expect(discounted).toBeLessThan(1)
})
it('a disagreeing fallback amount scores no better than a disagreeing total', () => {
// Renormalized weights made a wrong fallback amount OUTSCORE a wrong
// invoice total (0.67 vs 0.60 with exact date + merchant); the factor
// approach scores both at full weight and then discounts the fallback.
const asTotal = calculateMatchConfidence(0, 1.4, 0.9).confidence
const asFallback = roundOre(asTotal * FALLBACK_CONFIDENCE_FACTOR)
expect(asFallback).toBeLessThanOrEqual(asTotal)
expect(asFallback).toBeLessThan(0.6)
})
})
describe('normalizeForMatch', () => {
it('leaves the frozen key normalizer alone', () => {
// normalizeMerchantName feeds a PERSISTED unique key with a SQL mirror.
+67
View File
@@ -37,6 +37,24 @@ export const AMOUNT_TOLERANCE_PERCENT = 0.05
export const CONVERTED_AMOUNT_TOLERANCE_PERCENT = 0.09
export const MIN_MATCH_CONFIDENCE = 0.4
/**
* Flat discount on a confidence scored from a prominentAmounts fallback
* (bankintyg, avtal, contracts: no invoice-style total). Such an amount is one
* of possibly several figures printed on the document rather than "what the
* buyer pays", so an agreement is real evidence but must stay weaker than a
* total agreeing.
*
* A discount FACTOR, deliberately not a reduced amount weight inside
* calculateMatchConfidence: the confidence is normalised over the included
* weights, so shrinking the amount weight both let a date+amount-only fallback
* reach 1.0 ((0.25+0.3)/0.55) and, when the amount DISAGREED, shrank the
* penalty so a wrong fallback amount outscored a wrong invoice total
* (0.67 vs 0.60). Scoring at full weight and discounting the result keeps
* agreement capped below certainty and disagreement at least as damning as it
* is for a real total.
*/
export const FALLBACK_CONFIDENCE_FACTOR = 0.85
/**
* Normalize a merchant name for comparison.
* Removes special characters, Swedish company suffixes, and extra whitespace.
@@ -263,6 +281,55 @@ export function amountVarianceForMatch(
return null
}
export interface ProminentAmountMatch {
variance: number
/** The printed amount that produced the variance. */
amount: number
/** The document's own label for it ("Insatt belopp", "Engångspris"). */
label: string | null
}
/**
* Fallback amount variance for documents with no invoice-style total but one
* or more prominent amounts (bankintyg "Insatt belopp", an agreement's
* "Engångspris", ...). Tries each amount against the transaction and returns
* the smallest variance, or null when none is comparable.
*
* Same-currency only by construction: prominent amounts never carry a resolved
* SEK value, so a cross-currency pair stays incomparable (receiptSek = null in
* amountVarianceForMatch) exactly like a cross-currency total without a rate.
* Returns the closest amount with its variance and the document's own label.
*
* Callers must multiply the resulting confidence by
* FALLBACK_CONFIDENCE_FACTOR (see its comment for why a factor, not a weight),
* and should surface WHICH amount matched: a bare "Exakt belopp" with no
* number attached is certainty the reader cannot check.
*/
export function bestProminentAmountVariance(
amounts: readonly { amount: number; label: string | null }[],
receiptCurrency: string,
txAmount: number,
txCurrency: string,
txSek: number,
): ProminentAmountMatch | null {
let best: ProminentAmountMatch | null = null
for (const candidate of amounts) {
if (!Number.isFinite(candidate.amount)) continue
const variance = amountVarianceForMatch(
candidate.amount,
receiptCurrency,
null,
txAmount,
txCurrency,
txSek,
)
if (variance != null && (best == null || variance < best.variance)) {
best = { variance, amount: candidate.amount, label: candidate.label }
}
}
return best
}
/**
* Calculate a weighted match confidence score from date, amount, and merchant signals.
* Weights: amount 40%, merchant 35%, date 25%.
+19
View File
@@ -68,6 +68,25 @@ describe('selectProposals', () => {
expect(selectProposals([tx()], [], noSuppression)).toEqual([])
})
it('never proposes on the prominent-amounts fallback', () => {
// A bankintyg (documentKind "other", no invoice-style total) whose printed
// "Insatt belopp" happens to equal a same-day outflow scores 0.85 on the
// shared matcher, which would clear CERTAIN_CONFIDENCE and skip
// adjudication, on a pairing that is wrong by construction: the hunt scans
// outflows only, and "Insatt belopp" labels an inflow. Fallback-scored
// candidates are for the picker and the agent, never the nightly hunt.
const bankintyg = item(
{ id: 'item-intyg', document_id: 'doc-intyg' },
{
supplier: { name: null },
totals: { total: null, vatAmount: null },
documentKind: 'other',
prominentAmounts: [{ amount: 438.75, label: 'Insatt belopp' }],
},
)
expect(selectProposals([tx()], [bankintyg], noSuppression)).toEqual([])
})
it('skips a transaction that already has a live proposal', () => {
const result = selectProposals([tx()], [item()], {
claimedTransactionIds: new Set(['tx-1']),
+9
View File
@@ -172,6 +172,15 @@ export function selectProposals(
const scored = scoreUnderlagCandidates(tx, pool as never[]).filter(
(candidate) =>
candidate.document_id != null &&
// Never propose on the prominent-amounts fallback (non-invoice
// documents: bankintyg, avtal). The hunt's thresholds were calibrated
// for invoice-style totals; a fallback pair can reach 0.85 on date +
// printed-figure alone, which clears CERTAIN_CONFIDENCE and skips
// adjudication, and its preview would show document_amount null. The
// hunt also scans outflows only, so an inflow-labeled figure
// ("Insatt belopp") is guaranteed to pair with the wrong row. Those
// documents stay reachable through the picker and agent candidates.
candidate.amountSource !== 'prominent' &&
!spentDocumentIds.has(candidate.document_id) &&
!suppression.claimedDocumentIds.has(candidate.document_id) &&
!suppression.rejectedPairs.has(pairKey(tx.id, candidate.document_id)),