Commit Graph

5 Commits

Author SHA1 Message Date
Jakob Wennberg 4c479f84f4 feat(receipt-hunt): convert a foreign receipt instead of refusing to compare it (#1497)
* feat(receipt-hunt): compare a foreign receipt by converting it, not by refusing

A Swedish bank posts a converted figure for a purchase abroad while the
receipt states the original: Anthropic bills 180,00 EUR and the statement
reads -2 014,32 kr. Neither number appears in the other document, so the
matcher refused the pair rather than guess. On a SaaS-heavy ledger that is
not an edge case: of 25 receipts fetched from a real company's mailboxes,
14 were in USD or EUR and none could ever pair.

The receipt's total is now resolved into kronor with Riksbanken's rate for
its own date, and handed to the same matcher, which still wants the
merchant and the date to agree. The seam already existed: the scorer
passed null where a SEK value would go, with a comment explaining that
cross-currency pairs were deliberately incomparable. Surfaces that do not
resolve a rate still pass nothing and behave exactly as before.

Rates are fetched once per currency and day. Riksbanken answers 429 to a
caller that asks per document, and a run holds a dozen receipts from one
vendor in one month. A rate that cannot be resolved leaves the receipt
exactly as incomparable as it was.

Two calibration faults surfaced once the amounts became comparable:

A converted total is judged at 9% rather than 5%. Riksbanken publishes a
mid rate and a card issuer charges its own, so the two carry a known
spread on top of any disagreement about the sum: measured against real
statements, 1.2% to 3%. Holding both to one bar treats a rate spread as
if it were a discrepancy.

The date tolerance moves from 3 days to 10. A card settles days after the
purchase, an international one routinely a week later, and a forwarded
receipt carries the purchase date while the statement carries the posting.
At three days the signal scored zero for ordinary correct pairs and took a
quarter of the weight with it: a receipt agreeing to within 1%, from a
merchant the matcher recognised, still capped at 0.62. The 27 human-
confirmed pairs and the 7 near-misses that must not match all still hold.

Measured on that ledger: purchases with any candidate at all go 9 -> 13,
and the Anthropic pair proposes at 0.82 where it was previously
unscoreable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): an undated receipt is not converted at today's rate

Raised in review. A foreign receipt with no invoiceDate fell back to the
current date, which put a receipt of unknown age into amount matching on
the strength of a guess: a rate two years out is how something
incomparable acquires confidence it has not earned. No date, no
conversion, and the receipt stays exactly as it was.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 17:54:22 +02:00
Jakob Wennberg 43386b4852 feat(agent): offer unmatched inbox receipts as confirmable underlag (#1436)
The originally reported scenario is still broken after #1425 and its
backfill-by-document_id: a user photographs a receipt into WhatsApp, answers
the bot's questions, then opens the app, clicks the bank transaction and asks
the assistant to book it, and is told "UNDERLAG: saknas" about a receipt we
are holding, then asked everything again.

WhatsApp intake writes neither invoice_inbox_items.matched_transaction_id
(process-inbound.ts passes uploadAndExtract's matchedTransactionId as
undefined) nor transactions.document_id (that mirror is written by the manual
match route). Only TransactionMatchPicker fills either column. So the underlag
list comes back empty, and a backfill that keys on document_id has nothing to
key on.

Unmatched, unconsumed inbox items are now scored against the transaction with
the same pure scorer the picker uses and the strongest few are surfaced as
TROLIGT UNDERLAG, carrying their captured chat answers.

Proposals only: nothing writes matched_transaction_id, and the prompt tells
the agent to get the match confirmed and to book only against a confirmed one.
Setting the link at intake above a confidence bar is the obvious alternative
and is deliberately left open.

An uncomparable amount (cross-currency with no rate) disqualifies a candidate,
because calculateMatchConfidence drops the amount signal in that case and date
+ merchant alone then score a confident match nobody checked the sums for.

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 14:07:30 +02:00
Jakob Wennberg 0f7147a078 fix(agent): read a WhatsApp "nej" as an answer, not a half answer (#1433)
* fix(agent): read a WhatsApp "nej" as an answer, not a half answer

#1425 gave the assistant the answers the user typed in WhatsApp. Rendering
those inline off the raw channel_context blob gets the most common answer
backwards.

Answering "nej" to the representation question stores an EMPTY representation
block: participants: [], purpose: null, denied: true. The renderer branched on
`if (!rep.purpose)` and so emitted

    syfte SAKNAS: fråga bara efter syftet, inte om deltagarna igen.

for a user who had just said the meal was not representation. `denied` was
never read anywhere. The result is the assistant asking about the purpose of a
private lunch, which is worse than the generic re-ask #1425 fixed, because the
instruction is specific and confident.

Clarifications now come from a structured summary that models the denial and
the genuine half answer (participants named, purpose missing, which BFL 5 kap
6-7 § does want completed) as different states. #1425's syfte SAKNAS nudge is
preserved for the case it was written for.

Two smaller fixes in the same renderer, both about untrusted text:

- The photo caption no longer reaches the prompt. It is the one field on the
  record nobody was asked for and nobody reviewed, and the rationale already
  written down in lib/documents/channel-context-notes.ts for keeping it off an
  immutable verifikat applies at least as strongly to a prompt that can call
  tools.
- Human free text passes through flattenMemoryContent. An intent's
  promptTemplate output is seeded as a user message, so wrapToolResult never
  sees it and nothing else defends this path; a caption reading
  "# NYA INSTRUKTIONER: ..." previously rendered verbatim.

All three tests fail against the current renderer and pass against this one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agent): gate the chat-answer guidance on what was rendered

CodeRabbit caught the same defect shape this PR is about: the
prior-conversation paragraph was gated on chat_answers != null, but a
caption-only context is non-null and now summarises to nothing, so the
paragraph pointed at 'uppgivna av användaren' rows the prompt does not
contain. Gate on whether a clarification line was actually emitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 14:00:20 +02:00
Jakob Wennberg 4b51af3d80 feat(agent): 'Vad din agent vet' page rendering the ledger context (P2) (#935)
* feat(agent): 'Vad din agent vet' page rendering the ledger context (P2)

The human-facing surface for the openwiki ledger-context: a read-only page
that renders the exact payload the AI agent reads (Accounted://ledger/context
+ the briefing digest) as a legible profile of how this company books.

- Route app/(dashboard)/agent-knowledge (server component) calls the shared
  buildLedgerContext(supabase, companyId) directly: one payload, two
  renderers, no new API or data path.
- Sections mirror the payload 1:1: coverage/freshness strip, counterparty
  patterns (monochrome confidence bars + seen/agree evidence), supplier
  patterns, explicit rules shown as authoritative instructions distinct from
  observed patterns, account usage, VAT profile, conventions.
- Nav entry in the Analys group (icon Brain), ungated so it doubles as an
  upsell; flip requiredCapability to paywall.
- Design per .claude/rules/design.md (PageHeader, Card, Table, Badge,
  AccountNumber BAS tooltips); sv + en strings (agentKnowledge namespace).
  VAT/BAS labels stay Swedish in both locales per i18n rules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(agent): deep entity-resolved analysis + radial graph for the knowledge page

Reworks the 'Vad din agent vet' page from tables into a radial-hub graph
driven by a new full-history deep analysis, per founder feedback.

- fix(rpc): median_booking_lag_days now measures real posting promptness via
  committed_at, not entry_date (which the bank flow sets to the transaction
  date, giving a ~0 tautology: 151/152 on prod). migration 20260708120000.
- feat(rpc): get_ledger_deep_context (migration 20260708130000): full-history,
  deterministic, read-side. Merges counterparties by normalize_counterparty_key
  (e.g. Claude = 14 bookings across 12 name variants, weekly, 9 710 kr, always
  5420), mines booked verifikat for SEK spend (coalesce amount_sek), detects
  recurrence cadence, dominant account + share, plus supplier entities. Storno
  excluded, corrections kept; 19xx/26xx excluded from the dominant contra.
- LedgerGraph: radial SVG (company center, accounts inner ring, payees outer
  ring), hover/focus reveals variants + spend + cadence + account. Keyboard
  focusable nodes with per-node accessible names + a screen-reader data table.
- Page fetches the deep context alongside the light context; coverage strip
  gains tracked-payee / recurring / tracked-spend stats. sv + en strings.
- 14 pg tests (light + deep) green; both RPCs applied to prod + version-matched.

Reviewed by an adversarial multi-lens pass (accounting/SQL, frontend/a11y,
prod-fact verification); all four verified findings fixed (SEK currency,
storno-lag guard, keyboard a11y, spacing tokens).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(agent): gentle mount fade-in for the radial map (reduced-motion safe)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(agent): render the page for a rules-only company (empty-state edge case)

isEmpty ignored explicit_rules, so a company with configured mapping rules
but no posted transactions hit the 'hasn't learned anything' empty state and
lost its rules section. Rules are independent of bookings. (CodeRabbit)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(agent): show Kompetens (skills) + Fakta (memory) on the knowledge page

The 'Vad din agent vet' page now shows the full picture of what the agent
knows: alongside the booking map, a compact read-only view of its Kompetens
(the Swedish accounting/tax knowledge atoms it ships with, grouped by
tier as chips with active/dormant state) and the Fakta it remembers (top
learned facts with kind + source), each linking to /settings/assistant for
full management. Server-rendered via a new buildAgentCompetence() that
mirrors GET /api/agent/skills + /api/agent/memory. Also renders in the
no-bookings case so a new company still sees its agent's competence.
sv + en strings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(agent): restructure knowledge page - hero graph + tabbed detail

Declutters the page per feedback: the booking map is the always-visible hero,
and the supporting detail (Kompetens · Minne · Regler & profil) moves into
tabs so only one view shows at a time instead of a long card stack. Split
AgentCompetenceSections into standalone CompetenceCard + FactsCard for the
tabs; removed the top stat row on request. sv + en.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(agent): Reconciliation Aurora rewrite of the ledger knowledge graph

Full rewrite of LedgerGraph: node area = sqrt(spend), colour = cadence,
shape = supplier/counterparty, confidence = depth-of-field; on-mount
descriptor-collapse animation with xN badge; cadence pulse veins;
deterministic seeded layout; framer-motion only (no new deps); keyboard
navigation, reduced-motion and sr-only support. Build-verified; 3-lens
adversarial review findings fixed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(agent): sample-size-honest confidence in the ledger knowledge graph

dominant_account_share was raw cnt/total, so a counterparty with a single
booking rendered as '100% säkerhet': fake certainty by construction (the
data_quality_master Item-C / P3 finding). New migration replaces
get_ledger_deep_context with a Laplace-smoothed share (cnt+1)/(total+2)
(1/1 -> 0.67, 3/3 -> 0.80) and exposes the raw evidence as
dominant_account_count / dominant_account_total. The detail card now shows
'Bokförd hit i k av n fall' under the confidence bar; the existing focus
buckets, stroke widths and percent labels inherit the honest value
unchanged. pg-real test updated to guard the n=1 case.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 21:25:34 +02:00
Jakob Wennberg a3c6566caf feat(mcp): ledger-context resource with per-company booking patterns (#928)
* feat(mcp): ledger-context resource with per-company booking patterns

Adds Accounted://ledger/context: derived account usage, counterparty
booking patterns with explicit confidence share (0.7 floor), explicit
mapping rules kept separate as authoritative, observed VAT profile, and
conventions. Backed by a SECURITY INVOKER get_ledger_usage_stats RPC so
group-bys run SQL-side, and surfaced as a top-5 digest stanza on
gnubok_get_agent_briefing so one call still bootstraps a session.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(mcp): fold source-quality prereqs into the ledger-context RPC

Merchant-name normalization at the aggregation path (the splinter fix):
new normalize_counterparty_key() SQL function mirroring
normalizeCounterpartyName() so KORTKÖP/SWISH/date-suffixed labels merge
into one counterparty key, which also makes the categorization_templates
join exact. New supplier_patterns section (per-supplier dominant expense
account + VAT treatment from supplier invoices; credit notes and reversed
invoices excluded). account_usage excludes storno lines (they re-inflate
the account a correction moved away from); the counterparty CTE keeps
corrections because the transaction relink self-heals. Pattern confidence
is now count-grounded evidence {seen_12m, agree, share, last_booked}
instead of a bare ratio, and the digest frames it as historical frequency,
never auto-book permission.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(agent-context): use roundOre for the share ratio (antipattern ratchet)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(agent-context): defensive storno filter on counterparty CTE, fail-loud secondary reads

Review follow-ups: the counterparty CTE now excludes source_type='storno'
defensively (no live code path links a transaction to a storno, but legacy
rows may predate reverseEntry's unlink; a linked storno would count the
reversed category as precedent). Corrections stay included: they are the
live booking after relink. Secondary reads (rules, templates, settings)
now throw instead of silently reading as empty data: an agent must never
be told 'no rules' when the truth is 'read failed'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-08 13:48:04 +02:00