* feat(parties): selection-step evaluation against document-anchored truth Scores which similar keys are the same party without human labels: every key in the set carries an org number OCR-read from a linked invoice, so two keys are the same party exactly when the org numbers agree. Rules and the Bedrock model both land at 0.91 pair precision; the residual false merges are different legal entities sharing a trade name, which text cannot and should not separate. Results and caveats in the README. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(parties): make the model selector opt-in in the selection eval Sending voucher key text to the AI provider now requires --llm; the default run scores the rules selector only and makes no network call. Documents that the opt-in path uses the same configured provider the production categorizer already sends the same text to. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Parties, phase 0: measure before building
Phase 0 of the Kontakter plan (design doc: the "Vem du gör affärer med" artifact, 2026-09). Nothing in this phase changes product behaviour. It makes the resolver measurable so the model steps in later phases can be judged against labelled data instead of opinions.
What ships in this phase
- Migration
20260902120000: enablespg_trgm(trigram blocking of counterparty keys) and drops the two context-graph tables whose feature code was never merged and which prod no longer has. Test:tests/pg/parties-phase0.pg.test.ts. draw-golden-set.sql: the reproducible draw of the labelling sample and the payee-change base rate. Read-only.
The golden set (never in git)
The draw returns customer voucher text, which can contain person names. The
repository is public, so the drawn rows and the labels live in
dev_docs/parties/golden/ (gitignored). Only aggregates and this vocabulary
are versioned.
Label vocabulary for the pre-classifier, one per key:
| label | meaning | example |
|---|---|---|
party |
a real counterpart the company pays or invoices | Levfakt Beijer Byggmaterial AB (2089), Loopia |
category |
an expense description with no counterpart in it | Inköp av varor, Banktjänster, Fika |
payroll |
salary, benefits, expense claims to a person | Löneutbetalning: 2024.M.01 - anställd: 5 |
adjustment |
periodisering, kostnadsföring, omföring, lagerförändring | Periodisering av leverantörsfaktura 1311 |
authority |
Skatteverket, Bolagsverket, Transportstyrelsen, kommun | Bolagsverket ändra bolagsordning |
bank |
bank fees and bank products | Baspaket Bank, Avgift Amex |
intermediary |
payment rail or marketplace, not the counterpart | Klarna, Stripe, Zettle, AmazonMktplc |
unsure |
cannot tell from the text alone |
A key whose text names a vendor but is booked as a category still gets
party: the resolver decides identity, the account comes from the ledger.
Edge rules settled with the founder on 2026-09-02, after the first labelling round:
- State fees and taxes where the counterpart is fixed (Bolagsverket
registration fees, Skatteverket, Transportstyrelsen vehicle tax) are
authority. A kommun or a state company that invoices for a service is apartylike any supplier. - A marketplace or payment processor that the company actually pays (Amazon
Marketplace purchases, Stripe or Amex fees) is a
party.intermediaryis reserved for a rail that only carries someone else's money: Klarna, Swish and Zettle payouts, "försäljning via" lines. - A card-platform line that carries an employee's name (Pleo, Google Ads with
the cardholder appended) is a
party; the name is the cardholder, not the counterpart. - Insurance and pension premiums are
party: the insurer is the counterpart, whatever the 74xx account says.
Measured on prod, 2026-09-02
- 113,032 imported expense vouchers across 385 real companies; 99.5% yield a key; 71% belong to a key that repeats within the company.
- Payee-identity base rate: 365 suppliers have an established bankgiro on two or more documents; 13 of them (3.6%) carry two established bankgiro numbers, with an average span of 29 days between them, which reads as parallel accounts or OCR variance rather than changes over time. Any distinct value at all: 8.4%. Consequence for the signal: a new bankgiro is rare enough to warrant a "verify" sentence, never a block, and a value must be seen on two documents before it counts as known.
Pre-classifier shadow evaluation, 2026-09-02
eval-preclassifier.ts scores three routers on the same 180 held-out keys
(20 rows, chosen by md5 of the key, serve as few-shot examples and are never
scored). Rows the founder marked unsure always receive a label, so the
honest number is the one that excludes them.
| Router | Strict agreement | Excluding founder-unsure | Party TPR | Party TNR |
|---|---|---|---|---|
| Rules v0, lexicon from BAS account names | 0.917 | 0.965 | 0.992 | 0.930 |
| Rules v0, lexicon incl. BAS descriptions | 0.911 | 0.959 | 0.984 | 0.930 |
| Sonnet 5 on Bedrock EU, zero-shot | 0.861 | 0.906 | 0.906 | 0.977 |
| Sonnet 5 on Bedrock EU, 20 founder examples | 0.878 | 0.924 | 0.930 | 0.977 |
Read with two caveats. The rules were written after the labels were seen,
so their score is optimistic; the model scores are honest. And most of the
remaining disagreements are definitions, not errors: whether a Bolagsverket
fee is authority or intermediary, whether a marketplace or payment
processor (Amazon Marketplace, Stripe, Amex fees) is party or
intermediary, whether a card-platform line with a person's name (Pleo,
Google Ads) is party or payroll, and whether a pension insurance premium
is party or payroll. Settle those in the vocabulary above, relabel the
handful of rows, and re-run.
BAS description tokens in the lexicon make Google and Facebook look generic because the descriptions name them as examples; the account-name lexicon is the default for that reason. The model's remaining party misses are text with a person's name or a card descriptor, exactly where a hard key from a document would decide instead.
Selection-step evaluation against document-anchored truth, 2026-09-02
eval-selection.ts scores the selection step (which similar keys in the same
company are the same party) without human labels: every key in the set has an
org number read by OCR from an invoice linked to its voucher, so two keys are
the same party exactly when the org numbers agree. Keys mapping to several
org numbers are excluded. 225 anchors, 1,038 candidate pairs (747 same, 291
different), 44 anchors with no true match. Draw: draw-selection-gold.sql.
The model selector is opt-in (--llm): it sends the same voucher text the
production categorizer already sends to the configured AI provider (Bedrock EU
on hosted), so run it only with an env file whose AI settings you have checked.
| Selector | Pair precision | Pair recall | Non-match recall | Anchors fully right | "None" precision / recall |
|---|---|---|---|---|---|
| Rules (same core name) | 0.909 | 0.776 | 0.801 | 0.547 | 0.40 / 0.86 |
| Sonnet 5 on Bedrock EU, zero-shot | 0.910 | 0.904 | 0.770 | 0.756 | 0.68 / 0.77 |
What the false merges are: both selectors merge "Levfakt Fortnox (325)" with "Levfakt Fortnox (154)", "Vattenfall Kundservice" with "Vattenfall Eldistribution", "Jämtkraft" with "Jämtkraft" under another org number. Those are different legal entities sharing a trade name, or OCR variance in the org number. Text cannot separate them and should not try: the hard key from the document is the referee, which is exactly the resolver's precedence order. What the model catches that the rules miss: "lån themax" versus "lån från themax", and paraphrases with a changed word order.
Read with one caveat: sales-side vouchers ("kundbet ...") contribute noisy truth because the org number on the linked document can be the customer's; the next draw restricts to expense-side vouchers.
Cost: about 1,300 input tokens per anchor, 298k tokens for the run.
Still to confirm in this phase
- SCB access: the free Företagsregistret API carries F-skattstatus, Momsstatus and Arbetsgivarstatus (verified on scb.se 2026-09-02); access requested by email, credentials pending.
- The pre-classifier and selection agreement bars (0.85 on the labelled set) are enforced by the shadow harness in phase 1, not here.