feat(inbox): match non-invoice documents via prominent amounts (#2048)

* feat(inbox): match non-invoice documents via prominent amounts

Bankintyg, bank agreements and other documentKind "other" PDFs carry no
invoice-style total, so extraction correctly left totals.total null and the
document became structurally unmatchable: findUnderlagCandidates hard-drops
items without a comparable amount and the picker lost the 40% amount signal.

- extraction: new prominentAmounts[] field (amount + document's own label),
  populated only when totals.total is null; account/org/phone/reference
  numbers and zero amounts excluded. totals.total semantics untouched.
- matching: bestProminentAmountVariance() tries each printed amount and
  feeds calculateMatchConfidence at reduced weight (0.3 vs 0.4) in both the
  agent candidate scorer and TransactionMatchPicker.
- UI: inbox rail shows the detected amounts read-only for such documents,
  list falls back to a single distinct prominent amount, and extraction no
  longer reads as "found nothing".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

* fix(inbox): discount prominent-amount fallback instead of reweighting it

Skeptic pass refutations on the first commit: normalized weighting made a
reduced amount weight self-defeating. Date + exact fallback amount with no
merchant scored (0.25+0.3)/0.55 = 1.0 ("100% sakerhet" on a wrong same-day
transaction), and a DISAGREEING fallback amount scored above a disagreeing
invoice total (0.67 vs 0.60) because shrinking the weight also shrank the
penalty.

- score fallbacks at full amount weight, then multiply by a flat
  FALLBACK_CONFIDENCE_FACTOR (0.85): agreement caps below certainty,
  disagreement stays at least as damning as for a real total.
- agent candidate surface additionally requires the document date within
  DATE_TOLERANCE_DAYS, so an avtal listing 349 kr no longer matches every
  future 349 kr charge from the same counterparty.
- bestProminentAmountVariance returns which amount matched + its document
  label, and the match reason names it ("Exakt belopp i dokumentet: 2 500
  SEK (Engangspris)"): no more bare "Exakt belopp" reaching the agent while
  total_amount is null.
- prompt: prominentAmounts restricted to non-invoice documentKinds, and
  never a parking spot for an unreadable invoice total.
- fix the stale "deliberately the same list" comment on
  EXTRACTED_FIELD_ACCESSORS.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

* fix(receipt-hunt): never propose on the prominent-amounts fallback

Second skeptic pass: the nightly hunt is a third consumer of
scoreUnderlagCandidates and inherited the fallback unaware. A bankintyg
whose printed "Insatt belopp" equals a same-day outflow scores 0.85, which
clears CERTAIN_CONFIDENCE (0.8) and skips LLM adjudication, on a pairing
wrong by construction (the hunt scans outflows only; "Insatt belopp"
labels an inflow), with document_amount null in the approval preview.

UnderlagCandidate now carries amountSource ('total' | 'prominent') and
selectProposals drops fallback-scored candidates. Non-invoice documents
stay reachable through the manual picker and the agent candidate surface,
both of which have a human reading the amounts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

* fix(inbox): round fallback confidence via roundOre, not the naive pattern

The two confidence discounts (and their test) tripped the naive-ore-round
antipattern ratchet (625 vs baseline 622); use the sanctioned helper.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hqm9QgdyNAFaWiz6Ww7pgb

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Mattsson
2026-08-30 23:52:29 +02:00
committed by GitHub
co-authored by Claude Fable 5
parent 84ba8323b4
commit 516e8b62ff
14 changed files with 493 additions and 19 deletions
@@ -242,7 +242,19 @@ function timeAgo(iso: string): string {
}
function pickAmount(item: InboxItem): number | null {
return item.extracted_data?.totals?.total ?? null
const total = item.extracted_data?.totals?.total
if (total != null) return total
// Non-invoice documents (bankintyg, avtal) have no total; when exactly one
// distinct amount was read off the document, that is the amount to show.
// Two or more stay ambiguous and render as no amount.
const distinct = [
...new Set(
(item.extracted_data?.prominentAmounts ?? [])
.map((a) => a.amount)
.filter((a) => Number.isFinite(a) && a !== 0),
),
]
return distinct.length === 1 ? distinct[0] : null
}
function pickCurrency(item: InboxItem): string {
@@ -271,14 +283,18 @@ function hasAnyExtractedField(data: InvoiceExtractionResult | null): boolean {
s?.name || s?.orgNumber || s?.vatNumber || s?.bankgiro || s?.plusgiro ||
inv?.invoiceNumber || inv?.invoiceDate || inv?.dueDate || inv?.paymentReference ||
t?.subtotal != null || t?.vatAmount != null || t?.total != null ||
(data.lineItems?.length ?? 0) > 0 || (data.vatBreakdown?.length ?? 0) > 0
(data.lineItems?.length ?? 0) > 0 || (data.vatBreakdown?.length ?? 0) > 0 ||
(data.prominentAmounts?.length ?? 0) > 0
)
}
/**
* The fields the extraction is scored against, in the order a person reads
* them. Deliberately the same list hasAnyExtractedField checks, so the summary
* cannot claim a field the "is anything here at all" test does not count.
* them. A subset of what hasAnyExtractedField checks (that test also counts
* lineItems, vatBreakdown and prominentAmounts), so the "fält ifyllda" counter
* never claims a field the "is anything here at all" test does not count: the
* reverse can differ, e.g. a bankintyg with only prominentAmounts has fields
* but counts 0 here.
*/
const EXTRACTED_FIELD_ACCESSORS: ((d: InvoiceExtractionResult) => unknown)[] = [
(d) => d.supplier?.name,
@@ -2893,7 +2909,10 @@ function FieldsRail({
{/* AI classification: what kind of document this is and how it was
paid. Read-only context above the editable fields; absent for
extractions from before the fields existed. */}
{(data?.documentKind || data?.payment?.method || data?.pages) && (
{(data?.documentKind ||
data?.payment?.method ||
data?.pages ||
(data?.totals?.total == null && (data?.prominentAmounts?.length ?? 0) > 0)) && (
<div className="border-b px-4 py-3 text-xs space-y-1">
{data?.documentKind && (
<div className="flex gap-2">
@@ -2911,6 +2930,23 @@ function FieldsRail({
</span>
</div>
)}
{/* Amounts read off a non-invoice document (bankintyg, avtal):
without this row the empty "Totalt" field reads as if extraction
missed them. */}
{data?.totals?.total == null && (data?.prominentAmounts?.length ?? 0) > 0 && (
<div className="flex gap-2">
<span className="text-muted-foreground w-14 shrink-0">{t('prominent_amounts_label')}</span>
<span className="tabular-nums">
{(data?.prominentAmounts ?? [])
.map((a) =>
a.label
? `${a.label}: ${formatCurrency(a.amount, data?.invoice?.currency ?? 'SEK')}`
: formatCurrency(a.amount, data?.invoice?.currency ?? 'SEK'),
)
.join(' · ')}
</span>
</div>
)}
{data?.pages && (
<div className="text-muted-foreground">
{t('pages_partial_note', {
+42 -4
View File
@@ -17,11 +17,14 @@ import { cn, formatCurrency, formatDate } from '@/lib/utils'
import { Loader2, Search } from 'lucide-react'
import { Input } from '@/components/ui/input'
import {
FALLBACK_CONFIDENCE_FACTOR,
amountVarianceForMatch,
bestProminentAmountVariance,
calculateMatchConfidence,
calculateMerchantSimilarity,
} from '@/lib/documents/core-receipt-matcher'
import { resolveSekAmount } from '@/lib/bookkeeping/currency-utils'
import { roundOre } from '@/lib/money'
import type { InvoiceExtractionResult } from '@/types'
import { getErrorMessage as getUserErrorMessage } from '@/lib/errors/get-error-message'
@@ -132,6 +135,17 @@ export default function TransactionMatchPicker({
const total = extractedData?.totals?.total ?? null
const receiptCurrency = (extractedData?.invoice?.currency ?? 'SEK').toUpperCase()
const supplier = extractedData?.supplier?.name ?? null
// Non-invoice documents (bankintyg, avtal) have no total but often show the
// money amount anyway; those feed a discounted amount fallback below.
const prominentAmounts = useMemo(
() =>
total == null
? (extractedData?.prominentAmounts ?? []).filter(
(a) => Number.isFinite(a.amount) && a.amount !== 0,
)
: [],
[extractedData, total],
)
// SEK value of the underlag total. For a SEK underlag that's the total
// itself; for a foreign one it needs the fetched FX rate.
@@ -254,7 +268,7 @@ export default function TransactionMatchPicker({
// Currency-aware variance: null when uncomparable, which makes the
// matcher drop the amount signal instead of matching 750 EUR to 750 SEK.
const amountVariance = amountVarianceForMatch(
let amountVariance = amountVarianceForMatch(
total,
receiptCurrency,
receiptSek,
@@ -263,17 +277,36 @@ export default function TransactionMatchPicker({
txSek,
)
// No total (bankintyg, avtal): fall back to the closest prominent
// amount, discounted below so a printed figure never presents as the
// certainty a real total gives.
const fallbackMatch =
total == null && prominentAmounts.length > 0
? bestProminentAmountVariance(
prominentAmounts,
receiptCurrency,
tx.amount,
txCurrency,
txSek,
)
: null
if (fallbackMatch) amountVariance = fallbackMatch.variance
const dateVariance = Math.abs(
(new Date(tx.date).getTime() - invoiceDate.getTime()) / (1000 * 60 * 60 * 24),
)
const merchant = tx.merchant_name || tx.description || ''
const similarity = supplier ? calculateMerchantSimilarity(supplier, merchant) : 0
const { confidence, matchReasons } = calculateMatchConfidence(
const scoredMatch = calculateMatchConfidence(
dateVariance,
amountVariance,
similarity,
MATCH_DATE_TOLERANCE_DAYS,
)
const matchReasons = scoredMatch.matchReasons
const confidence = fallbackMatch
? roundOre(scoredMatch.confidence * FALLBACK_CONFIDENCE_FACTOR)
: scoredMatch.confidence
return {
id: tx.id,
@@ -289,7 +322,7 @@ export default function TransactionMatchPicker({
})
scored.sort((a, b) => b.confidence - a.confidence)
return scored
}, [rawRows, invoiceDate, total, receiptCurrency, receiptSek, supplier])
}, [rawRows, invoiceDate, total, prominentAmounts, receiptCurrency, receiptSek, supplier])
// Instant client-side narrowing while the debounced server query catches up.
const filtered = useMemo(() => {
@@ -344,7 +377,7 @@ export default function TransactionMatchPicker({
{/* Underlag reference: what we're matching against, so a currency or
amount mismatch with a candidate is obvious at a glance. */}
{(total != null || supplier) && (
{(total != null || prominentAmounts.length > 0 || supplier) && (
<div className="rounded-lg border bg-muted/30 px-3 py-2 text-xs flex items-center gap-x-3 gap-y-1 flex-wrap">
<span className="text-muted-foreground shrink-0">Underlag</span>
{supplier && <span className="font-medium truncate">{supplier}</span>}
@@ -359,6 +392,11 @@ export default function TransactionMatchPicker({
)}
</span>
)}
{total == null && prominentAmounts.length > 0 && (
<span className="tabular-nums font-medium shrink-0">
{prominentAmounts.map((a) => formatCurrency(a.amount, receiptCurrency)).join(' · ')}
</span>
)}
{hasInvoiceDate && rawInvoiceDate && (
<span className="text-muted-foreground tabular-nums shrink-0">
{formatDate(rawInvoiceDate)}