Commit Graph

340 Commits

Author SHA1 Message Date
Jakob Wennberg 1d635b0d25 feat(receipt-hunt): look for receipts on request, from the mailbox settings page (#1496)
* feat(receipt-hunt): a button that looks for receipts on request

The nightly cron exists but still does not search mailboxes, and for a
good reason: a sweep of one real 172-message mailbox took over 600s,
against a scheduled function's 300. Pressing a button is the honest shape
for work that big. A bounded pass reports what it found and how much is
left, and the person decides whether to press again; a nightly run could
only truncate silently.

POST /api/receipt-hunt/run searches the mailboxes for eight purchases and
fetches at most ten receipts per press. Gated on the AI tier, because
reading the amount out of a PDF is what makes a fetched attachment
matchable at all: without it the hunt would file documents that can never
pair, which is worse than not running. Writes no journal entries; every
pairing is still a proposal waiting for approval.

The button lives on the mailbox settings page, which already ships, and
says what happened in words rather than a spinner that stops: "3 underlag
hämtade. 12 köp kvar att söka igenom."

huntCompany gains maxReceipts so a manual pass can carry a different
budget from a nightly one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): stop a manual press spending its budget on the wrong purchases

The first real press searched eight purchases, read forty mails and found
nothing, which looked like a broken model. It was the ordering.

Purchases are searched largest first, which is right for a nightly sweep
that eventually covers everything and wrong for a button pressed a few
times. On a real ledger the largest rows are the least likely to have a
findable receipt: rent already invoiced, bare payment references, direct
debits. Those filled the forty-mail cap, so the productive purchases
further down the list, the ones whose receipts are actually sitting in the
mailbox, were never read at all.

The cap was the binding constraint, not the time: eight purchases and
forty mails took 43s of the 300 available. A press now searches 25
purchases and reads 100 mails, measured at 85s and finding 7 underlag on
the same ledger that returned 0 before.

huntCompany gains maxMails alongside maxReceipts, so a manual pass can
carry a different budget from a nightly one rather than sharing an
environment default with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): one underlag per purchase, and a press that fits its budget

Your second press exposed three things, none of which a dry run reaches.

It took 5.8 minutes. The 85s I measured was a dry run, which never
fetches, uploads or extracts; each fetched receipt costs about another
37s because it is downloaded, stored, and then read by a model that opens
the PDF. Seven of them ran past the 300s a serverless function gets, so
in production that press would have been killed. Four receipts per press
keeps a full pass inside the budget.

It fetched seven receipts and proposed nothing. A single mail carries the
invoice AND the receipt for one purchase under different names
("Invoice-E19DBF63-0021.pdf" beside "Receipt-2066-0204-8388.pdf"), and
the same receipt reaches a second mailbox on a different message. Each
was fetched separately, so the pool filled with identical candidates and
the matcher refused to propose any of them rather than flip a coin. The
per-run key is now the vendor and the total, which is what identifies a
purchase; the filename only decides when no amount was read. Nine
duplicates already in the pool were removed.

And with the duplicates gone it still proposed nothing, for a separate
reason: "Utlägg Norwegian" scored 0.18 against "Norwegian Air Shuttle
AOC AS". Utlägg is Swedish for an expense reimbursement, bank vocabulary
rather than a company, and leaving it in broke the token-subset match, so
an exact 1 998 kr pair leaned entirely on a date eight days out and fell
under the floor. Stripped, along with överföring, via internet, bg-bet
and autogiro, in the comparison path only.

normalizeMerchantName is untouched: it is the persisted konteringskarta
key with a SQL mirror, and its 22 string pins and the 27-pair golden set
still pass.

Measured after: the Norwegian pair proposes at 0.72.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* perf(mail): read each message once per press, not once per query

The third press proposed a pairing, which the second had not, but still
ran 5.2 minutes against a function's 300s. Cutting receipts from seven to
four had only saved 36s, which said the receipts were never the cost.

Every search fetched a full message for every hit, and a press searches
many purchases across every connected mailbox. One receipt mail answers
several of those queries, so 25 purchases against 2 mailboxes could ask
Gmail for well over a thousand messages to end up with a hundred distinct
ones. Deduplication happened in the caller, too late to save the work.

A mail's content never changes, so it is now read once per mailbox and
kept, bounded at a thousand entries and evicting oldest first. Measured
on the same ledger: 55 purchases and 100 mails now take 102s, where 25
purchases alone previously cost around 264s before a single receipt was
fetched.

clearMessageCache exists because tests reuse message ids and production
does not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): do not refetch a receipt the company already holds

The fourth press fetched four documents the company already had:
Bolagsverket, Supabase twice, Uber. They came back because I had deleted
them as duplicates, and the cross-run check is the message and attachment
id, which lives in the rows I removed.

That was my mistake, but it exposed a real gap. The vendor-and-total key
only deduplicates inside a single pass. Across passes the same purchase
still arrives as an invoice in one mail and a receipt in another, with
different file keys, and both were fetched: the pool fills with identical
candidates and the matcher then refuses to choose between them, which is
how a press can fetch four documents and propose nothing.

The pass now starts from what the company already holds, so its budget
goes on documents that are actually missing.

Receipts per press drops to three. Measured on this ledger, a fetched
receipt costs about 50s from download to a stored amount, and that is the
model reading the PDF rather than the network: seven took 5.8 minutes and
four took 5.1, both past the 300s a function gets. Three fits, but it is a
stopgap. Doing the fetch inside the request is the wrong shape for work
this slow, and the fix is to move it off the request rather than keep
shaving this number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mail): a mail body must not outlive the run that read it

Two findings from the review, both correct.

The message cache held whole MailCandidate values, and one of those fields
is the mail body. The contract says a body is read once to extract fields
and discarded, and a process-global cache quietly broke that: bodies of
one company's mail sat in memory across requests until eviction or a
restart. The MailSearchService contract now has releaseCache, the Gmail
adapter clears its messages, and the hunt calls it in a finally so a
failed run releases them too.

The duplicate key accepted an empty vendor, so two unrelated documents
that happened to cost the same collapsed into one candidate. Those now
fall back to the file they came from: without a vendor there is nothing
to anchor an amount to.

The same finding caught something worse that I had introduced one commit
earlier. The persistent check derived its key from the stored extraction
while the fetch derived one from the reading model, so a document filed
as "Norwegian Air Shuttle AOC AS" did not recognise an incoming
"Norwegian" and was fetched again. Rather than guess at aliases, which
would fold "Google Cloud" into "Google Workspace", the identity is now
written onto the row when the receipt is filed and read back verbatim.
Rows filed before that fall back to the extraction.

receiptIdentity is one exported helper with its own tests, used by both
sides, instead of the same expression written twice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): a monthly subscription is not a duplicate of last month

The review caught that my duplicate key was worse than the problem it
solved. Anthropic bills the same amount every month, and keying on vendor
and total alone made July look like a duplicate of June: every later
receipt from any recurring supplier would have been suppressed forever,
silently. Duplicates block one proposal; that would have lost a receipt
per month per subscription.

The identity now carries the document date. Two documents for one
purchase share a date; June and July do not.

Two smaller faults in the same key. The amount was serialised as a raw
float, so 0.1 + 0.2 read as a different total from 0.3; it is rounded to
öre like every other money comparison in this codebase. And a document
with no vendor was identified by its filename alone, which collapses two
unrelated papers whenever a billing system attaches "invoice.pdf": those
now carry the message they came from.

The key is versioned so a future change to its shape cannot be mistaken
for a match against rows written under the old one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 16:08:06 +02:00
Jakob Wennberg 38f5d9812e feat(receipt-hunt): find receipts in connected mailboxes and pair them on the amount (#1492)
* feat(receipt-hunt): nightly matcher pairing unbooked purchases with held receipts

Stages an attach_document_to_transaction proposal for every unbooked card
purchase whose receipt the company already holds, so the underlag is attached
before the transaction is booked and the gap never forms. When the user later
books it, categorize-core.ts propagates the document onto the new verifikat
through the matched_transaction_id link the executor writes.

Deliberately scoped to UNBOOKED transactions. The posted-verifikat backlog is
96% imported history whose originals live in the previous system, so it stays a
pull (the verifikat_missing_document worklist) rather than a nightly push.

Ranking reuses scoreUnderlagCandidates; the pool is loaded once per company
instead of per transaction, which removes both the N+1 and the newest-50
truncation a per-transaction lookup imposes on a deep backlog.

Five guards, each mutation-tested: a confidence floor above the shared
candidate floor, an ambiguity margin so two equally-good receipts are left to
the picker rather than coin-flipped, one-receipt-one-purchase, one live
proposal per purchase, and permanent suppression of pairs a human rejected.
Suppression is derived from pending_operations history rather than a new table:
terminal rows are immutable and a rejection is already the durable "no".

Runs 05:30 UTC, after the 05:00 bank sync. Gated on RECEIPT_HUNT_COMPANY_IDS,
which hunts nobody when unset so enabling it stays a deliberate act. No
migration, no journal writes, no UI: proposals land in the existing Granskning
queue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(receipt-hunt): dry-run mode for provkörning against a real ledger

Returns the pairings a run would stage without writing any of them, so a
company can see tonight's proposals before they reach the granskningskö and so
the matcher can be validated against production data without staging an
operation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(matching): fold Swedish bank descriptors so receipts reach their purchases

calculateMerchantSimilarity compared raw bank descriptors, so a receipt from
"Alviks kött och fisk" scored 0.125 against the bank's own row for it,
"Alviks koett och fisk K3667 Kortköp/uttag" — an öre-exact pair no threshold
could reach. Adds normalizeForMatch, used for similarity only, which folds what
the card rails add and never changes identity: the K#### token, Kortköp/uttag
verbs, a leading "Kortköp YYMMDD", trailing /YY-MM-DD dates, reference numbers
glued to the name, domain wrappers, legal forms, and the three ways banks mangle
Swedish letters (ö, transliterated "oe", and ?? mojibake). Processor markers
become spaces because the merchant sits before the star in GOOGLE*PLAY and after
it in K*IKEA GALLE. Token-subset containment is scored level with substring
containment so a receipt's legal name matches the bank's trading name.

normalizeMerchantName is left byte-identical and now documents why: it is a
transitive input to categorization_templates.counterparty_name, a persisted
UNIQUE key with a hand-written SQL mirror the ledger-context RPC recomputes at
query time. Changing it would make stored keys stop equalling computed ones, so
the konteringskarta join misses and insertOrUpdateTemplate inserts a second row
per merchant instead of migrating the occurrence counts.

Aggressive folding is safe because it is applied to both sides of every
comparison, so an over-eager fold still matches; the risk is collision between
different merchants, which the new tests guard.

Measured on 27 receipt/transaction pairs humans actually confirmed in
production: recall 27/27, and 0/7 false positives on deliberately similar but
distinct merchants. Full unit suite unchanged (13,004 passing), including the 22
string pins on the frozen key path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(mail): read-only Gmail connector so receipts are found without forwarding

Forwarding was the only way a receipt reached Accounted, and it is both
unpopular (97% of companies with the problem have never used their inbox
address) and fragile: Arcim's own forward has been off for weeks and nobody
noticed. This lets the hunt look in the mailbox instead.

Scope is gmail.readonly and nothing else. It can search and download attachment
bytes, and it structurally cannot send, modify or delete: the promise the
consent screen makes is enforced by the grant, not by our code being careful.
The consequence is deliberate: the agent can prepare a forward for a portal-link
receipt but can never send one itself.

Query-then-classify, never sync. For each unexplained purchase we run a
provider-side search in a -3/+10 day window, pull metadata for a handful of
hits, and keep nothing. No mailbox is mirrored and no message body is stored,
which is what keeps this inside Google's Limited Use terms and GDPR data
minimisation. Mail is searched only for purchases Underlag could not already
explain, so a receipt we already hold never costs a mailbox read.

The query ORs merchant against amount rather than requiring both: demanding both
misses every rebrand and reseller (Anthropic bills as Claude), while the amount
alone is a strong filter inside two weeks.

mail_connections is service-role only with RLS enabled and zero policies,
because the row holds a live refresh token and RLS cannot hide a column.
Uniqueness is (company, provider, address) so a second mailbox is additive and a
reconnect updates in place. Tokens are AES-256-GCM under their own key by
preference, since a mail grant reads correspondence rather than backups.

Core reaches the extension through a registered service, mirroring
lib/email/service.ts, so lib/receipt-hunt never imports from @/extensions and a
zero-extension build still compiles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(mail): connect UI and ingest, making the hunt reach into the mailbox

Two halves that together make the connector usable.

Ingest (lib/receipt-hunt/ingest.ts, core): fetches the attachment, files it as a
document and an inbox item with source 'mail_hunt', then stages the pairing.
It lives in core because it writes documents and inbox items, and an extension
may never import another extension; the mail extension only ever hands over
bytes.

No re-matching for a hunted receipt: it was fetched WHILE SEARCHING for a
specific purchase, so the pairing is known by construction. The search is a
deliberately broad OR query, which is exactly why the proposal still goes to a
human with the mailbox, sender and subject written on it rather than being
linked automatically.

Provenance goes in channel_context, never extracted_data, because retrying
extraction overwrites extracted_data wholesale and the record of which mailbox
a receipt came from has to survive that. A partial unique index on
(company_id, channel_context->>'mail_message_id') makes re-runs and the same
receipt arriving in two mailboxes idempotent, and a 23505 is treated as success
rather than an error.

Guards, both mutation-tested: a duplicate message costs no provider call, and an
oversized attachment is skipped rather than stored. One unreadable attachment
falls through to the next and never aborts a night's hunt.

UI: /settings/mail lists connected mailboxes with their health, connects a new
one through a user-gesture tab (opened before the await, so popup blockers do
not eat it), and disconnects behind a ConfirmDialog that states the outcome up
front, including that already-approved receipts stay because they belong to the
bookkeeping now. Strings in sv and en; the read-only promise is spelled out on
the page rather than buried in a consent screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mail): renumber migrations to clear a version collision on main

20260806150000 was already taken by preserve_preset_committed_at, and
woocommerce_connections plus enforce_balance_on_posted_insert landed after this
branch was cut. Two files sharing a version breaks every fresh database, which
only shows up on a clean setup rather than on an already-migrated one.

Applied to prod under the new versions (20260807090000 / 20260807090100), so
schema_migrations matches these filenames exactly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): make the mailbox search actually able to find an underlag

A provkörning against a real ledger returned the same seven unrelated
messages for every purchase, all reporting no attachments. Three separate
causes, each fixed and pinned:

1. `getMessageSummary` asked Gmail for `format=metadata`, which returns
   headers and omits `payload.parts` entirely. Every message therefore
   looked attachment-free, `bodyIsReceipt` was always true, and the
   `found.find(c => c.attachmentIds.length > 0)` guard in the hunt could
   never select anything: the feature could not file a single receipt.
   Gmail has no format that returns MIME structure without the body, so
   the body now comes down the wire; it is read for nothing and stored
   nowhere.

2. The bank's description is not a merchant name. "Lön Juli Jakob
   Överföring via internet" searched for "Juli" and matched most of the
   mailbox. Month names and payment-rail boilerplate are now stopwords.

3. Salary and tax runs are a company's largest outgoing rows, so they
   consumed the whole search budget hunting receipts that cannot exist.
   `canHaveEmailReceipt` skips them for the mail leg only. Deliberately
   narrow: a supplier invoice paid over bankgiro does arrive by mail, and
   an "Utlägg" reimbursement has a real receipt behind it.

Measured on the same ledger: 22 hits, 0 with attachments, 0 ingestable
-> 4 hits, all with attachments, 3 of 4 correct (Elgiganten, Sting,
Anthropic). The fourth matched a Stockholm billing address, which is why
every proposal still waits for a human.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(receipt-hunt): let a model resolve merchants and pick the receipt

The keyword hunt was failing for reasons regex tuning cannot reach, all
measured against a real mailbox rather than assumed:

- `from:anthropic.com` returns 0. Receipts arrive here by being
  forwarded, so the sender is the user, not the vendor.
- The exact charged amount returns 0. The bank posts a converted SEK
  figure that appears nowhere in a USD receipt.
- A date window around the purchase returns 0, while the same merchant
  search without one returns 10+. A forward is stamped when it was
  forwarded, sometimes months later.

So the query now searches merchant names across the whole mailbox, and
precision is restored by judgement rather than by syntax. Two model calls
per run, both through forced tool use so the reply is a shape and not
prose to be parsed:

1. `planMerchantGroups` resolves bank descriptors to merchants and merges
   repeats. Six Anthropic subscriptions become one search and one
   decision instead of six of each.
2. `assignReceipts` decides which mail, and which attachment on it, is
   the receipt for which charge, and says why in a sentence the reviewer
   reads.

The attachment, not the message, is the unit of an underlag: a single
forward routinely carries receipts for several purchases ("Fwd: Kvitton
februari" has five). Migration 20260807103000 moves the dedupe key from
message to message+attachment, with a backfill, because the old index
would have silently blocked every receipt after the first in a forward.

The model may not produce any number that reaches the ledger. It returns
ids, a confidence and a reason; amounts, dates and the write stay in
deterministic code. Its answer is validated, not trusted: an unknown
message id, an invented filename or a low confidence drops the pairing,
and any failed call proposes nothing at all. Every result still waits
for a human.

Measured on the same ledger: 0 receipts that could ever be filed -> 3
correct pairings (Elgiganten, Sting office invoice, Anthropic), each
with a stated reason. The five remaining Anthropic charges are dated
after 2026-06-15, when forwarding to the connected mailbox stopped; the
model declined them correctly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(receipt-hunt): amount first, and drop the confidence scoring

Three findings from how others build this, applied.

Production email search (Superhuman, Haystack 2026) reports that recall
comes from loosening retrieval and letting the model filter downstream,
not from tightening the query. Retrieval depth per merchant 12 -> 25, and
purchases the planner cannot name a merchant for are now searched by
amount alone instead of skipped: a line like "1260525758758
Europabetalning" identifies no merchant but is a real supplier payment
whose invoice may carry exactly that total.

Reconciliation engines weight amount far above date (Midday: 35% vs 5%)
because banks post late while amounts do not drift. The Gmail query now
leads with the amount and ORs the merchant, rather than dropping the
amount whenever a merchant alias exists. Still an OR: a receipt billed in
USD never contains the SEK figure the bank charged.

The confidence score is gone entirely. Research on verbalised confidence
finds it badly calibrated, clustered on round-number anchors and barely
better than chance at separating a model's own right answers from its
wrong ones. That matched what this ran into: the model anchored on 0.6 /
0.7 / 0.75 / 0.9, and the 0.7 threshold discarded two correct pairings.
It is replaced by an observation rather than a self-assessment, whether
the charged amount is actually visible in the mail, which is what a
reviewer checks first and what sorts the queue.

Also fixes a real defect the run exposed: the one-file-one-purchase guard
only held within a merchant group, so when the planner split one landlord
into "Sting" and "Kontorsplatser" both 15 000 kr charges were assigned the
same invoice. A file is now claimed once per run, which is the duplicate
underlag BFL forbids.

Measured on the same ledger: 3 -> 5 pairings, no duplicate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(receipt-hunt): harvest receipts, then pair them on the amount

Splits the mailbox leg in two along the line of what each side can
actually know.

The model was being asked which purchase a mail belonged to. Deciding
that needs the amount; the amount lives inside the PDF; a Gmail preview
essentially never shows it. Measured over a real mailbox, every single
pairing came back "belopp ej synligt": it was answering without the
deciding evidence, which is why it declined five of six repeat
subscriptions and why two correct pairings sat just under a threshold.

Now it answers only what a subject, a sender and a preview line support:
is this mail an underlag, and which attachment is it. Then the receipt is
fetched, the extraction that already runs on document.uploaded reads its
amount, date and vendor, and the pairing is the same deterministic
amount-and-merchant match every other underlag goes through. Amount
becomes decisive for real rather than as an instruction the model could
not act on.

The load-bearing fix is small: ingest now copies the extraction result
onto the inbox item. The pool is read from invoice_inbox_items, so a
hunted receipt with no extracted_data could never have matched anything,
and the whole mail leg was quietly incapable of producing a pairing on
amount.

Consequences, all deliberate:
- Harvesting runs BEFORE the pool is read, so a receipt found tonight is
  paired tonight rather than a night later.
- One staging path instead of two. Mail-sourced proposals carry the same
  preview and confidence as every other, plus where they came from.
- Deduped on the attachment filename, not on the message: the same
  invoice arrives as an original, a reminder and two forwards, and the
  old key filed "Invoice_13041840.pdf" four times over.
- Capped at 8 receipts per merchant per run.

Measured on the same ledger: 5 pairings attempted from thin evidence ->
16 real documents identified, each waiting on an amount it can be checked
against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(receipt-hunt): the model reads mail, arithmetic does the matching

Collapses the mailbox leg to one model call that extracts fields, and
hands every judgement back to deterministic code.

Gone: resolving bank descriptors to merchant names, deciding which mail
belongs to which charge, and the confidence score gating the result.
Three prompts and two model calls become one, and mail-intelligence.ts
drops from 450 lines to 250.

What made this possible was measuring what a mail actually contains. The
body was being downloaded and thrown away in favour of a 200-character
snippet, and the body is where a forwarded receipt quotes its original
sender and its original date. That is the purchase date, the thing whose
absence forced the date window off entirely and made the old design miss
five of six repeat subscriptions. It was there all along.

So the model now answers only what text can support: is this an underlag,
from whom, when, and for how much if the mail says so. Fields, not
judgements. Everything after is arithmetic:

- Retrieval is deterministic. No model decides what to search for.
- Fetching is gated by worthFetching(): a stated amount is enough on its
  own, a vendor needs a plausible date, and a mail found by a purchase's
  own search is evidence in itself. That last rule is what handles a
  supplier the bank and the invoice name differently ("Kontorsplatser j
  BG" against "Stockholm Innovation & Growth AB"), which is what the
  deleted merchant-resolution call used to buy.
- The pairing is the existing scorer, reached the same way as every other
  underlag: fetch, let the extraction that already runs on upload read
  the PDF, match on the amount. Amount is decisive in fact rather than as
  an instruction the model could not act on.

Also adds the Swedish thousands-space amount formats to the query.
Measured: the Sting invoice is findable as "15 000,00" and "15 000" and
by no ungrouped form at all, so every amount search was missing them.

Measured on the same ledger: 5 thin pairings -> 8 real documents, each
with a vendor and a true purchase date, waiting on the amount in its own
PDF. Currency is never converted to make a number agree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): trust the bytes, not the mail, when filing an attachment

Found by the first live run, which fetched nothing and reported success.
Three defects, each invisible to a dry run because a dry run never
downloads anything.

1. Gmail declares a forwarded PDF as application/octet-stream, and
   uploadDocument validates content against the declared type, so the
   upload was rejected: "Filinnehållet matchar inte den angivna
   filtypen". Every forwarded receipt with a generic MIME type would
   have failed this way, silently, since ingest swallows one bad
   attachment to protect the rest of the run. The type is now sniffed
   from the magic bytes, then the filename, and only then from what the
   mail claimed.

2. The filename was re-derived by a second full message fetch inside
   fetchAttachment, which came back empty and fell back to a generic
   "underlag.pdf", discarding the real "2332687551.pdf" the search had
   already reported. The known name now wins.

3. The provkörning script imported lib/init instead of calling
   ensureInitialized(), so document.uploaded reached no handler and
   nothing was ever extracted. It also used static imports, which are
   hoisted and ran before .env.local was read, leaving the extraction
   extension unable to build a Supabase client. Both are script defects,
   not product defects: the cron route calls ensureInitialized() at
   module level as the architecture requires. The script now loads the
   environment first and imports dynamically.

Also makes the per-run fetch cap tunable (RECEIPT_HUNT_MAX_RECEIPTS) so a
pilot can be held to a couple of documents, and adds --live to the
script, which is the only way it writes anything.

Verified end to end against a real ledger, every link exercised for the
first time: two attachments fetched from Gmail, stored with their real
names and types, extraction run on both, the amount copied onto the inbox
item, and the deterministic matcher pairing Elgiganten 21 639,00 kr from
the PDF against the -21 639 kr card purchase at 0.85, staged into
Granskning as attach_document_to_transaction. The second document, a
Bolagsverket filing receipt, carries no total and correctly paired with
nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(receipt-hunt): sweep a whole mailbox, and stop lending one receipt twice

A backfill on a real ledger, 22 documents fetched from 172 messages.

Batches the extraction (25 mails per call) so a first run on an existing
company can read the whole mailbox instead of the 40 mails one call can
carry, and makes the per-run caps tunable
(RECEIPT_HUNT_MAX_MAILS, RECEIPT_HUNT_MAX_RECEIPTS) so a pilot can be
bounded. The nightly caps stay where they are: they pace the review
queue, and a backlog is a different job from a nightly tick.

Two defects the backfill exposed, neither reachable from a dry run:

The one-receipt-one-purchase rule only held inside a single run.
`spentDocumentIds` is per-invocation, so an H&M receipt was proposed
against a -358 kr purchase on one pass and a -354 kr purchase on the
next, and approving both would have put the same underlag on two
verifikat. A live proposal now claims its document across runs, the same
way it already claimed its transaction.

A document reported with no filename, on a message carrying five
attachments, was not an answer but a shrug: the caller fetched
attachment number one and hoped. Those are dropped now. A body-only
receipt, where there is nothing to choose between, still passes.

Measured after the sweep: 21 of 22 documents read correctly, and the
binding constraint on this ledger is no longer retrieval but currency.
Ten receipts are in SEK and five of those pair on the amount; twelve are
in USD or EUR, where the bank charged a converted figure that appears
nowhere in the receipt, so no comparison is possible and none is
attempted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(mail): show the provider's own mark on the mailbox settings page

Someone connecting a mailbox is picking an account at a provider, and the
provider's mark is how they recognise which one. A generic envelope
glyph said "mail" when the question is "whose".

The Google "G" already existed, drawn inline inside GoogleAuthButton for
the sign-in flow. It moves to components/ui/provider-marks so there is
one definition rather than two, and a Microsoft square joins it for the
Graph connector. Both stay inline: no external host is contacted for an
icon before anyone has agreed to anything.

These are the only coloured glyphs in an achromatic interface, which is
deliberate rather than an oversight. A brand mark is identity, not
chrome, and Google's terms require its mark unaltered rather than tinted
to match a palette.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(archive): drop the duplicate mail_connections exclusion left by the rebase

Main added the table to ARCHIVE_EXCLUDED_TABLES while this branch was
open, so rebasing produced the key twice and the zero-extension build
failed to type check. Main's entry stays, in its alphabetical place, and
keeps the sentence that answers the retention question: the grants are
not räkenskapsinformation, but the receipts they find are archived as
documents.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mail): record who disconnected a mailbox, without keeping the token

Raised by the compliance review: disconnect() hard-deleted the row with
no trace, and which mailboxes feed underlag into the books is a control
over how räkenskapsinformation is produced (BFNAR 2013:2 kap 8), so
switching one off should be reconstructable years later.

Written by hand rather than by the write_audit_log trigger the accounting
tables use. That trigger copies the whole row into audit_log, which here
would mean copying an encrypted refresh token into a second table and
keeping it after the entire point of the delete was to destroy it. The
sibling credential table shopify_connections omits the trigger for the
same reason. Only the address and provider are recorded, pinned by a test
that fails if a credential ever reaches the audit entry.

The review's two other flags were checked rather than assumed: nothing
purges mail_hunt documents, and categorize-core.ts:403 does carry the
attached document onto the verifikat when the transaction is booked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(mail): bound every outbound call, and stop the token widening itself

Four findings from the review, each checked against the code first.

Neither the Gmail API nor Google's token endpoint had a deadline. Both
are awaited inside Promise.all across mailboxes, so one stalled request
held the whole company's hunt open until the platform killed the run.
Both now carry a 15s AbortSignal, which turns a stall into one mailbox
missing from tonight's sweep.

`include_granted_scopes: 'true'` let Google fold scopes this app was
granted elsewhere into the token issued for a mailbox, so a grant could
carry more authority than the consent screen showed. Removed, and pinned
by a test asserting the parameter is absent.

disconnect() ignored both statement results: a failed delete still wrote
an audit entry claiming the mailbox was disconnected while the credential
was live, and a failed audit insert passed silently. The delete now
throws, so the entry is never written for a delete that did not happen.
The audit failure is logged rather than rolled back: the two can now only
diverge one way, credential gone and note missing, and recreating a
credential to keep them in step would be worse than a missing note.

The fifth finding is real and stays open by choice, recorded in
DECISIONS.md: the cron still passes searchMail=false. A sweep of one
172-message mailbox took over 600s against a maxDuration of 300, so
enabling the mailbox leg nightly would time out mid-run. That flag and
RECEIPT_HUNT_COMPANY_IDS get flipped together once the per-company budget
is measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(receipt-hunt): file each attachment under its own identity

Four more findings from the review. The first is a real defect.

ingestMailCandidate loops over candidate.attachmentIds, but the dedupe
key, the mail_attachment_id provenance and the filename were all read
from index 0. Storing the second attachment therefore recorded the
first one's key and name, which mislabels the row and, because the key
is unique, permanently blocks the first attachment from ever landing.
Masked today only because the hunt narrows to a single attachment before
calling in, so nothing in the current path exercises it. All three now
come from the attachment actually being stored, and the duplicate
pre-check moved inside the loop so trying a second attachment is not
suppressed by the first already being filed. Mutation-tested.

The per-run fetch key was the bare filename, which is not an identity:
"invoice.pdf" is what half the world's billing systems attach, so a
second supplier's invoice would be dropped as a duplicate of the first.
Scoped by vendor as well, keeping the behaviour it was written for, one
fetch for an invoice that arrives as an original, a reminder and two
forwards.

Adds tests/pg/mail-hunt-file-dedupe.pg.test.ts for the new unique index:
five attachments from one forward all land, the same attachment is
refused twice, two companies hold the same file independently, other
inbox sources are untouched by the partial predicate, and the
message-scoped predecessor is gone. Written against CI's Postgres; there
is no local DATABASE_URL here, so CI is what exercises it.

--live now refuses unless RECEIPT_HUNT_CONFIRM names the same company.
The script writes to whatever .env.local points at, which for this repo
is production, and a recalled command should not be able to fire it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(test): cast the jsonb parameter so Postgres can type it

pg-real could not determine the type of $3 inside jsonb_build_object.
An explicit ::text is what the other pg tests do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:23:42 +02:00
Jakob Wennberg 51bf1b77bf feat(import): mount the knowledge-graph theater on the Arcim migrating step (#1486)
* feat(import): stream real Arcim migration progress as NDJSON (#1485)

* feat(import): stream real Arcim migration progress as NDJSON

The /migrate route ran the orchestrator to completion and answered with one
JSON blob, so the wizard faked its progress bar: a hardcoded 55% anchor and
a static step label for a phase that can take minutes. The orchestrator has
had a real onProgress channel (eight emit points with Swedish step labels
and anchors) since it was written; the route just never passed it.

Now a request with Accept: application/x-ndjson gets a streamed response:
one line per orchestrator progress event, then a terminal done line with
the results or an error line carrying the same structured envelope the
JSON path returns (the 200 status is already committed once the stream
opens). Callers without the header keep the original single-JSON contract,
so pre-deploy tabs and the existing error-mapping tests are untouched.

The wizard opts in, drives MigratingStep from the real labels and anchors
(mapped onto the 55-100 slice of the wizard bar), and treats a dropped
connection as unconfirmed rather than failed, since the migration keeps
running server-side and a blind retry could double-import.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: record the opt-in NDJSON streaming decision for /migrate

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* feat(import): mount the knowledge-graph theater on the Arcim migrating step

The migration wizard already holds the fully parsed SIE client-side
(SIEData.parsed from /sie-data), so the same TheaterCanvas that carries
the /import flow can build the company's knowledge graph while the
migration runs: no server change, and the plain progress card stays as
the fallback whenever no parsed SIE exists (e.g. providers without SIE).

Unlike /import's fixed narration script, the wizard knows exactly what
the server is doing: phase 1 posts one SIE file at a time and phase 2
streams the orchestrator's real progress events. ArcimMigrationTheater
therefore narrates by printing those real step labels once each as they
arrive, and keys the canvas to the same milestones: the GL skeleton
(rings, buckets, accounts) builds during the journal writes, counterparty
waves attach while customers and suppliers import, and reconciliation
pulses. Real progress bar and elapsed counter stay visible throughout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(import): review fixes on the migration theater and stream close

CodeRabbit on #1486: (1) the aria-live region wrapped the per-second
elapsed counter, so a screen reader re-announced the timer every second
and drowned out the real step labels; the live region now covers only the
narration list and the timer row is aria-hidden (the Progress bar exposes
its own ARIA value). (2) controller.close() in the stream's finally block
throws if the reader already cancelled, escaping start() as an unhandled
rejection; now guarded like send().

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 09:43:32 +02:00
Jakob Wennberg 1d01010928 feat(import): stream real Arcim migration progress as NDJSON (#1485)
* feat(import): stream real Arcim migration progress as NDJSON

The /migrate route ran the orchestrator to completion and answered with one
JSON blob, so the wizard faked its progress bar: a hardcoded 55% anchor and
a static step label for a phase that can take minutes. The orchestrator has
had a real onProgress channel (eight emit points with Swedish step labels
and anchors) since it was written; the route just never passed it.

Now a request with Accept: application/x-ndjson gets a streamed response:
one line per orchestrator progress event, then a terminal done line with
the results or an error line carrying the same structured envelope the
JSON path returns (the 200 status is already committed once the stream
opens). Callers without the header keep the original single-JSON contract,
so pre-deploy tabs and the existing error-mapping tests are untouched.

The wizard opts in, drives MigratingStep from the real labels and anchors
(mapped onto the 55-100 slice of the wizard bar), and treats a dropped
connection as unconfirmed rather than failed, since the migration keeps
running server-side and a blind retry could double-import.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: record the opt-in NDJSON streaming decision for /migrate

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 09:26:27 +02:00
Jakob Wennberg ce1e0b7642 feat(banking): the bank-connect payoff (#1477)
* feat(banking): the bank-connect payoff: point the first sync onward

Fifth activation slice. Every successful connect passes through the
sync-progress dialog's done state, which used to end on a bare Klar
that stranded the user on the settings panel. Now it is the payoff:
the imported count stays, the work now waiting gets named (N att
bokföra · M matchar fakturor, from the same worklist counts endpoint
the dashboard pane refetches), and the primary action becomes
Visa N att bokföra -> /transactions, where realtime rows, the match
pills and the Att bokföra badge already deliver the rest. Zero-import
syncs keep the plain Klar.

Also mounts the fully-built-but-orphaned BankSyncSinceLastVisit pill
on the transactions footer, so returning users get the same payoff
line for the nightly cron (N nya transaktioner sedan sist).

Strings stay hardcoded Swedish inside the extension, matching every
neighboring string in the dialog.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(banking): review triage: no stale counts, no marker advance on error

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 20:52:25 +02:00
Mattsson c187fabf92 feat(shopify): Shopify order/refund feed into the transactions inbox (#1474)
* feat(shopify): Shopify order/refund feed into the transactions inbox

New extensions/general/shopify feed extension, modeled on the WooCommerce
feed: connect a Shopify store with Dev Dashboard custom-app client
credentials (client credentials grant, ~24h tokens, never stored), then a
nightly cron + manual sync imports paid orders and refunds via the GraphQL
Admin API (pinned 2026-07) into the transactions inbox on clearing account
1584. Feed-only: nothing auto-books. Zero PII fields are queried, keeping
the app outside Shopify's protected customer data program.

- shopify_connections migration (RLS, revoke-never-delete, encrypted
  client id/secret) + shopify_sync capability and bank_sync-mirrored
  backfill
- frozen external_id scheme shopify_{shop_domain}_order|refund_{id},
  scoped on the shop domain so reconnects never re-import
- cursor sync on updated_at windows with 24h overlap, lock-date drop at
  map time, ingest-failure cursor floor, deadline stop-and-resume,
  revoked-credential flip
- /import card + settings panel, sv/en i18n, cron 03:15 in vercel.json +
  regenerated Docker crontabs, logo, events, panel registry
- 65 unit tests + pg-real RLS test; extensions.schema.json enum also
  gains the missing stripe entry (pre-existing drift)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(shopify): review findings from PR 1474

- token exchange: a 429 that survives every retry is throttling, not a
  credential failure; stop remapping retryable 4xx to 401 so sustained
  throttling can no longer flip the connection to revoked and delete the
  stored credentials (CodeRabbit critical)
- order sync: advance a scanned-through watermark (run start, capped by
  the failure floor) after a fully-listed window, so empty first runs and
  quiet stores rotate to the back of the cron's oldest-first selection
  instead of permanently occupying the 50-connection batch (CodeRabbit
  major, starvation)
- add handler-level tests for the orders cron route (auth 401, disabled
  503, unconfigured no-op, query failure, capability skip, happy path,
  per-connection failure isolation, revoked marking)
- add 401 tests for /sync, /transaction-sync and /disconnect; pin the
  cursor floor rule with a two-order page; stub the encryption key via
  vi.stubEnv
- note in the panel description (sv/en) that orders can mix VAT rates and
  must be split at booking (Swedish review advisory)
- DECISIONS.md: wrap underscore identifiers in backticks (MD037)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 12:44:08 +02:00
Mattsson 39f4ecdad4 fix(providers): surface migration step errors; INK2 SRU 7104; non-modal invoice dialog (#1465)
* feat(mileage): körjournal with milersättning booking, MCP tools and CSV export

New mileage_trips table (RLS, booked-delete trigger per BFL retention),
lib/mileage service reusing the payroll schablon rates, /api/mileage routes
(trips CRUD, period booking to 7331, salary-run push, körjournal CSV),
Körjournal dashboard page + nav, and three staged MCP tools (search-only
catalog). Trips book as one verifikat per period via the engine; salary
path inserts mileage_taxfree line items. mileage_trips classified in the
full-archive export.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(mileage): use shared roundOre helper per tightened ratchet baseline

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): pending_operations op-type migration + Swedish review findings

- New migration pair adds log_mileage_trip/book_mileage_period to the
  pending_operations operation_type CHECK (pg-real audit).
- bookMileagePeriod refuses a period spanning several employees and names
  the employee in the verifikationstext when scoped (BFL motpart).
- vehicle_registration required for förmånsbil trips (schema, service,
  MCP staging, UI surfaces the field).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): claim-first booking, CSV injection guard and driver column

- bookMileagePeriod claims trips (draft to booked CAS) before creating the
  verifikat, so a concurrent second booking loses the race instead of
  double-booking; claim reverts if verifikat creation fails.
- Körjournal CSV neutralizes formula-injection triggers (OWASP) and adds a
  Förare column naming the employee per trip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): resolve CodeRabbit + Swedish review round: race, drift and hardening

- Copying a round trip no longer re-doubles the stored distance.
- pushMileageToSalaryRun claims trips before inserting line items (retry can
  no longer double-pay); CLAIM_LOST replaces misleading NO_TRIPS on lost races.
- Booked trips are DB-immutable via a BEFORE UPDATE trigger (new migration
  20260807113215): only claim/link/revert transitions and notes edits pass.
- Cross-year periods rejected (schablon rates are per calendar year); payroll
  config year read from the date string, not TZ-dependent getFullYear().
- MCP staged bookings freeze the previewed trip set (trip_ids in params) and
  the commit fails on drift; validation errors return 400, not 500.
- PATCH enforces the förmånsbil regnr rule on the effective row; export
  validates dates before they reach the Content-Disposition header; employee_id
  is verified company-scoped on trip creation; stale orphaned claims released.
- UI: fetch flags reset in finally; ICU plural for draft summary; distance
  stored at the column's 1-decimal precision.
- Tests: [id] route suite, pushMileageToSalaryRun suite, claim-race, drift,
  cross-year and update-trigger pg cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): revert-to-draft must clear salary_run_id at the trigger level

New migration 20260807114924 replaces the booked-immutability function: a
booked -> draft revert now rejects rows keeping salary_run_id, closing the
DB-level double-pay path CodeRabbit flagged. pg test pins both directions;
the CLAIM_LOST unit test now asserts the revert.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): company-scope employee_id on PATCH (Superagent P2)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(mileage): valid v4 uuid in cross-company employee PATCH test

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(providers): surface migration step errors instead of silent empty syncs

A Visma company without the API module activated (403 ErrorCode 4002,
"No access to module: api_standard") failed every provider call during
migration, yet the wizard reported success with zero rows and mapped the
403 to "reconnect", which loops forever since OAuth succeeds against
Visma's shared identity server. A real user burned time re-syncing and
reconnecting, then filed the config issue as a bug.

- New PROVIDER_API_MODULE_INACTIVE code; classifyProviderError reads the
  error body and recognizes the module error before the 403 to
  AUTH_EXPIRED mapping. Registry entry carries the remediation in
  Swedish and English (activate the API under Appar och tillagg, paid
  add-on on smaller plans, clear standardforetag, SIE fallback).
- Orchestrator: connection-level failures (auth expired, license
  missing, module inactive) rethrow and abort the doomed run so /migrate
  answers with the typed code; other step failures stay non-fatal but
  land on results.stepErrors instead of only in server logs.
- /preview fails fast on the two subscription codes so the user reads
  the remediation at connect time, before any sync.
- Wizard: preview treats the new code like the Fortnox license case
  (CTA + SIE fallback); the result step renders error cards per cause
  and says "Migrering delvis genomford" instead of "Allt ar uppdaterat";
  the completion toast is honest on partial failure.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ink2): SRU field 1.1 is 7104, not 7113 (Skatteverket rejects 7113)

The INK2 huvudblankett code for 1.1 Overskott av naringsverksamhet is
7104 per Skatteverket's official 2025P4 faltkoder (INK2_SKV2002-33-01-24-04).
We emitted 7113, which does not exist on INK2, so filoverforing rejected
every profitable company's BLANKETTER.SRU with 'UPPGIFT 7113 ar inte ett
giltigt postnamn' (reported by a user for FY 2024-10-07..2025-12-31).
Underskott (7114) was already correct.

The wrong code originated in the swedish-sru-filing skill reference;
fixed there too and regenerated the atom seed. All other emitted
INK2/INK2R/INK2S codes verified against the official 2025P4 lists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoices): keep the AI chat usable over the new-invoice dialog

The new-invoice dialog was a modal Radix dialog: modal mode sets body
pointer-events: none, aria-hidden on body siblings, and a focus trap, so
the agent sheet (z-60, painted above the dialog) was visible but dead:
clicks swallowed, input unfocusable, and all three dismiss paths
preventDefaulted, leaving no way out except the header X.

Now non-modal: page modality is restored by hand instead. A new
DialogVeil primitive supplies the backdrop (Radix renders no overlay in
non-modal mode) at z-40, under dialog content (z-50) and the agent sheet
(z-60), and inert on #dash-shell blocks pointer, keyboard, and AT access
to the page behind while the sheet (a body-level sibling) stays live.
The lazy-load fallback dialog on /invoices gets the same treatment so a
hung or 404'd chunk cannot dead-lock the route.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 16:04:58 +02:00
Mattsson 799fa1246a fix(vat): downgrade per-voucher RC basis gaps only under per-rate evidence (#1464)
* fix(vat): downgrade per-voucher RC basis gaps only under per-rate evidence

Per-voucher RC basis gap findings (findRcBasisGaps) blocked "Skicka till
Skatteverket" as ERROR even when the flagged vouchers were legitimate
moms-only rattelseverifikat whose basbelopp lives in another (often
reversed) verifikat. In that state no arrangement of vouchers satisfies
both the per-voucher scan and the aggregate basis/moms identity, so the
block was unfixable: every correction voucher joined the blocklist it
was meant to clear (Orto Engineering 3DJake support case, 2026-08).

The gap finding now downgrades to a non-blocking WARNING only when ALL
of the following hold, otherwise the blocking ERROR stays exactly as
before:

- the 44xx/45xx RC basis accounts, grouped per momssats
  (RC_BASIS_ACCOUNTS_BY_RATE), match ruta 30/31/32 two-sided within a
  0.5 kr ore epsilon per rate;
- no moms box (ruta 30/31/32) is negative;
- the aggregate RC_OUTPUT_MISSING check has not fired;
- the caller supplied the evidence at all (older wire payloads and
  totals-less contexts keep the blocking behavior).

A first cross-rate-sum predicate was refuted by adversarial review: a
wrong-rate fiktiv moms voucher (12% moms "covered" by a 25% basis)
reached parity and unblocked a 7 800 kr under-declaration, and a
net-negative rate box made the summed comparison vacuous (textbook
FK004 state filing). Rutor 20-24 are partitioned by purchase type, not
rate, so the certificate must come from account totals; both
counterexamples plus the tolerance-hole case (shortfall inside the
aggregate 0.5% tolerance still blocks) are locked in as regression
tests.

The evidence travels as rcBasisByRate on the declaration payload
(rcBasisTotalsByRate projection), consumed by the web view and the MCP
completeness checks; rc-basis-gaps.ts derives its flat account set from
the same rate-grouped single source so scan and evidence cannot drift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(vat): refuse gap downgrade on non-finite evidence; pin the ore epsilon

Review findings, one pass:

- CodeRabbit (major): rcBasisByRate arrives as unvalidated JSON in the
  web view; a missing or non-numeric field made every per-rate
  comparison evaluate against NaN, which compares false and PASSED the
  predicate, relaxing the filing gate in the unsafe direction. The
  predicate now refuses the downgrade outright on any non-finite basis
  or moms figure, covering both the web and MCP callers.
- CodeRabbit (nit): added a 0.51 kr drift case so a future widening of
  the 0.5 kr epsilon fails a test instead of slipping through green.

Declined with reasons (recorded in the PR summary): requiring textual
voucher-to-voucher references before downgrading (belongs to the
rattelse documentation flow, and would reintroduce the unfixable block
this PR removes); epsilon stacking across rates (max 1.5 kr, immaterial
at whole-krona filing and below the aggregate tolerance); explicit
negative-basis guard (all negative-basis paths already block via the
two-sided mismatch or the negative-moms guard, now plus the finite
guard).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 14:45:22 +02:00
Mattsson 4b0a185876 fix(invoice-inbox): parse AI extraction output wrapped in markdown fences (#1460)
* fix(invoice-inbox): parse AI extraction output wrapped in markdown fences

Since the Sonnet 5 switch (2d543ac99, 2026-07-27) the model intermittently
wraps its JSON answer in ```json fences or adds a short preamble despite
the JSON-only system-prompt rule. JSON.parse(rawText) then threw, the
catch swallowed the error into emptyResult(), and the user got a blank
extraction form: 10-20% of prod receipt extractions since July 28 landed
empty (confidence 0) while the Bedrock call was still paid for.

Slice the raw response from the first '{' to the last '}' before parsing.
Fenced, prefixed, and suffixed outputs now parse; brace-less prose
refusals fall through to the existing empty-result path unchanged. Covers
both pipelines (invoice-inbox upload/email/whatsapp and the
document-extraction extension) since they share extractInvoiceFields.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoice-inbox): depth-aware JSON extraction instead of naive brace slice

Review findings (CodeRabbit, PR Agent, compliance swarm) converged on the
same edge case: first-'{'/last-'}' slicing picks a wrong span when the
model's surrounding prose itself contains braces. Replace it with a
string- and escape-aware balanced scan that returns the first candidate
JSON.parse accepts; prose-only responses still fall through unchanged to
the empty-result path. Two regression tests: braces in surrounding prose,
braces inside JSON string values.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoice-inbox): bound the JSON candidate scan against pathological input

Compliance swarm round 2 (A.8.28/A.8.29, non-blocking): the balanced-brace
scan restarted from every '{' with no bound, worst-case quadratic on
adversarially brace-laden text. Cap input length at 256 KB and candidate
attempts at 50; real model output is capped by MAX_TOKENS at roughly 33 KB
so genuine responses never come near either bound. Exhausted or oversized
input falls through unchanged to the existing empty-result path. Two tests:
100k-brace pathological input completes fast and lands empty, oversized
input skips scanning entirely.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:41:27 +02:00
Jakob Wennberg 70845edf69 feat(transactions): structured transaction_method instead of channel-in-the-name (#1459)
* feat(transactions): structured transaction_method instead of channel-in-the-name

Swedish bank feeds embed the payment channel in the description string
("Vercel Jul Överföring via internet", "ANTHROPIC* ... Kortköp/uttag"):
the PSD2 remittance array is joined into one string and the ISO 20022
type codes were dropped at insert. This promotes the channel to data:

- transactions.transaction_method (text + CHECK closed vocabulary: card,
  transfer, bankgiro, plusgiro, swish, autogiro, e_invoice, international,
  deposit, withdrawal, salary, fee, interest, adjustment) plus verbatim
  bank_transaction_code / proprietary_bank_transaction_code evidence
  columns (data_quality_master Appendix B "Layer-A capture").
- classifyTransactionMethod() in lib/transactions/transaction-method.ts:
  explicit source method (Stripe txn.type) > trailing Swedish channel
  phrase > ISO 20022 family/subfamily > proprietary-code keywords > MCC.
  It also splits the clean display title off the description.
- Ingest stores the clean title as description and the full bank string
  as original_description; dedup is untouched (external_id is date+öre,
  the content bridge reads original_description and is prefix-based, and
  a trailing strip leaves a prefix). Enable Banking passes the codes
  through; the Stripe feed sets methods from its balance-txn types.
- Backfill migration classifies existing rows from the description text
  (+ MCC and Stripe prefixes) and strips unedited titles; user-edited
  titles are never rewritten.
- mapping-engine also matches original_description so user rules written
  against the full bank text keep firing.
- UI: the inbox row shows the clean name; clicking it now folds out
  "Betalsätt: Kortköp" etc. (sv/en), making every classified row
  expandable.

A card purchase implies a physical receipt, a Bankgiro/e-invoice payment
implies a supplier invoice: downstream automations can now branch on the
rail instead of regexing display strings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(transactions): anchor counterparty-template identity on original_description

Audit follow-up to the phrase-strip change: counterparty template lookup
AND learning derived their key from merchant_name || description. With
the working title now stripped ("SPOTIFY AB Kortköp" -> "SPOTIFY AB"),
templates learned from the full bank string would only re-match via the
occurrence-gated single-token tier, and single-token counterparties with
fewer than 3 bookings would silently stop matching.

Both sides now read merchant_name || original_description || description:
the immutable bank original is identical across eras (and across user
renames), so every stored key and alias keeps matching exactly. Same
anchoring rationale as buildMerchantHistory in category-suggestions.

Existing tests that relied on the fixture's default original_description
now state it explicitly; two new regression tests pin the era stability
(lookup via alias on the full string, learning key derivation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(transactions): review follow-ups on method classification

- methodFromCodes: two-pass subfamily-then-family scan so a SALA/XBCT
  refinement on the proprietary code beats a bare family match on the
  ISO code, matching the documented precedence; pinned by a test.
- mapping-engine: regression tests for merchant/description patterns
  that only match original_description, including the invalid-regex
  substring fallback and the no-match default.
- Stripe: regression test for the SDK-unmodeled 'tax' balance-txn type
  mapping to 'fee'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(transactions): scope method classification to feed rows + adjective guard

Multi-bank risk hardening before the backfill ships:

- Feed-row scope: classification and title stripping now require a real
  import feed (import_source present, not manual/mcp), both at the
  ingest boundary (USER_CREATED_IMPORT_SOURCES, now exported) and in
  every backfill statement. User-authored titles like "Egen insättning"
  on manual/MCP rows are never classified and never rewritten.
- Adjective guard (TS + SQL): a strip that would leave the title ending
  in a possessive/scope adjective (egen/eget/privat/intern/extern ...)
  is skipped, so "Egen insättning" stays whole even on bank-feed rows;
  the method column still classifies (deposit).
- Unknown bank phrasings remain untouched by construction: an unmatched
  phrase means no method and no rewrite, so the worst case for any bank
  whose vocabulary we have not seen is the status quo.

Pinned by new unit + pg-real cases (user-created exclusion for
NULL/manual/mcp, adjective guard, feed defaults in the pg fixture).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): re-timestamp transaction_method migrations after rebase

Main gained migrations dated 20260729-20260730 (already applied to prod)
while this branch carried 20260728 versions, which would have applied
out-of-order on merge. The files have never reached prod, so renaming to
current timestamps is safe and removes any dependence on the integration's
out-of-order handling. All code/doc references updated; the pg test reads
the backfill by its new filename.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): bump transaction_method versions past prod's max

Main's newest applied migration is 20260730090000 (future-leaning
timestamp), so the previous 202607300731xx rename still sorted before
prod's tail and risked a silent skip on merge-time apply. Versions are
now 20260730100000/20260730100100, strictly after everything applied to
prod. References updated; full migration stream replays clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(transactions): final review round: keyboard guard + bank_connection_id feed marker

- TransactionInboxCard: row-level Enter/Space handling now ignores events
  bubbling from nested controls, so keyboard activation of Bokför / the
  overflow menu is no longer cancelled by the (now much more common)
  expandable row.
- Feed predicate parity with isImportedTransaction(): a live
  bank_connection_id marks a feed row even when import_source is unset
  (the oldest PSD2 rows predate that column), in both the ingest
  classifier and every backfill statement: those legacy rows now get
  classified instead of being skipped as user-created.
- pg fixture typing uses the TransactionMethod union.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): re-timestamp transaction_method migrations past prod's 20260807 tail

Prod max applied is 20260807170000 (verified by name via list_migrations);
the 20260730-stamped pair would sort before it. References in code,
tests, and DECISIONS.md updated to the new versions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migrations): enforce, not assume, original_description preservation in the title strip

The strip UPDATE now fills a NULL original_description from the
pre-strip description in the same statement. Prod has zero such rows
(0/25,566 feed-scope rows, verified read-only), and 20260605120000's
backfill plus ingest make the NULL case unreachable on any DB that
replayed history, but the migration should not depend on that history
to avoid losing the only copy of a bank string.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: record the compliance-review triage of the backfill's booked-row title strip

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
2026-08-08 11:58:51 +02:00
Jakob Wennberg a49d75db77 fix(migration): Visma pagination + chunk-insert resilience (the '300 misslyckades' case) (#1455)
* fix(providers): paginate Visma eAccounting with $page/$pagesize

eAccounting silently ignores OData $top/$skip, so every request returned
page 1 and getPaginated appended the first page TotalNumberOfPages times:
customers were imported in triplicate and invoice chunks hit unique
violations. Also stop on an empty page so a stale Meta can never loop or
duplicate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migration): survive bad rows in entity imports instead of failing whole chunks

One PostgREST insert per 500-row chunk is all-or-nothing, so a single
duplicate reported every row as failed ('300 misslyckades') with no cause
shown. Now: dedupe repeats within the fetched data (paging faults, source
duplicates), fall back to per-row inserts when a chunk is rejected, store
empty invoice numbers as NULL instead of colliding '', surface the first
DB error in the result UI, and mark all-failed steps with an error icon.
Sales invoices also carry remaining_amount so open invoices no longer
land as settled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migration): never per-row retry after a successful bulk insert with short read-back

A succeeded statement whose .select() returns fewer rows than sent means
the rows ARE in the table; retrying them one by one would duplicate every
unreturned row. Pair what came back and report the tail instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migration): count stub-insert casualties as failed and sample enrichment errors

Review follow-ups: invoices dropped because their customer/supplier stub
insert errored are DB failures, not matching misses; classifying them as
noMatch rendered a green result row with the database error hidden.
Enrichment failures now also feed errorSample.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 10:45:02 +02:00
Mattsson 7411a0171b feat(mileage): körjournal with milersättning booking, MCP tools and CSV export (#1448)
* feat(mileage): körjournal with milersättning booking, MCP tools and CSV export

New mileage_trips table (RLS, booked-delete trigger per BFL retention),
lib/mileage service reusing the payroll schablon rates, /api/mileage routes
(trips CRUD, period booking to 7331, salary-run push, körjournal CSV),
Körjournal dashboard page + nav, and three staged MCP tools (search-only
catalog). Trips book as one verifikat per period via the engine; salary
path inserts mileage_taxfree line items. mileage_trips classified in the
full-archive export.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(mileage): use shared roundOre helper per tightened ratchet baseline

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): pending_operations op-type migration + Swedish review findings

- New migration pair adds log_mileage_trip/book_mileage_period to the
  pending_operations operation_type CHECK (pg-real audit).
- bookMileagePeriod refuses a period spanning several employees and names
  the employee in the verifikationstext when scoped (BFL motpart).
- vehicle_registration required for förmånsbil trips (schema, service,
  MCP staging, UI surfaces the field).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): claim-first booking, CSV injection guard and driver column

- bookMileagePeriod claims trips (draft to booked CAS) before creating the
  verifikat, so a concurrent second booking loses the race instead of
  double-booking; claim reverts if verifikat creation fails.
- Körjournal CSV neutralizes formula-injection triggers (OWASP) and adds a
  Förare column naming the employee per trip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): resolve CodeRabbit + Swedish review round: race, drift and hardening

- Copying a round trip no longer re-doubles the stored distance.
- pushMileageToSalaryRun claims trips before inserting line items (retry can
  no longer double-pay); CLAIM_LOST replaces misleading NO_TRIPS on lost races.
- Booked trips are DB-immutable via a BEFORE UPDATE trigger (new migration
  20260807113215): only claim/link/revert transitions and notes edits pass.
- Cross-year periods rejected (schablon rates are per calendar year); payroll
  config year read from the date string, not TZ-dependent getFullYear().
- MCP staged bookings freeze the previewed trip set (trip_ids in params) and
  the commit fails on drift; validation errors return 400, not 500.
- PATCH enforces the förmånsbil regnr rule on the effective row; export
  validates dates before they reach the Content-Disposition header; employee_id
  is verified company-scoped on trip creation; stale orphaned claims released.
- UI: fetch flags reset in finally; ICU plural for draft summary; distance
  stored at the column's 1-decimal precision.
- Tests: [id] route suite, pushMileageToSalaryRun suite, claim-race, drift,
  cross-year and update-trigger pg cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): revert-to-draft must clear salary_run_id at the trigger level

New migration 20260807114924 replaces the booked-immutability function: a
booked -> draft revert now rejects rows keeping salary_run_id, closing the
DB-level double-pay path CodeRabbit flagged. pg test pins both directions;
the CLAIM_LOST unit test now asserts the revert.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mileage): company-scope employee_id on PATCH (Superagent P2)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(mileage): valid v4 uuid in cross-company employee PATCH test

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 14:17:21 +02:00
Mattsson 93f81f03e8 feat(providers): WINT migration provider behind WINT_MIGRATION_ENABLED (#1446)
* feat(providers): WINT migration provider behind WINT_MIGRATION_ENABLED

Adds WINT (wint.se) as a sixth migration provider, built against the
OpenAPI specs WINT's own API host serves publicly. Tier A scope: only the
partner-facing v1 endpoints are used; the general ledger is fetched as
vouchers/accounts and rendered as SIE 4E by our own sie-builder, with
opening balances for earlier years derived backward from the current-year
Ib anchor. Auth is the user's WINT login exchanged once for a JWT pair;
the password is never stored.

Ships dark: the wizard shows a disabled "Kommer snart" card, and the
server-side /connect gate rejects WINT until WINT_MIGRATION_ENABLED=true.
Live verification against a real WINT account is still outstanding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(providers): harden WINT provider per PR #1446 review findings

Addresses CodeRabbit and Swedish accounting review feedback in one pass:

- Ib anchor selection now uses WINT's unfiltered fiscal-year list, so an
  active year outside the allowed import window can never silently anchor
  the wrong year; the voucher chain is extended through the anchor and a
  per-year fetch failure fails that year loudly instead of sinking the
  whole migration.
- Auth token exchange is strict: only LoginState Success with a complete
  access+refresh pair mints a consent (a pair without a refresh token is
  unrefreshable and would break days later).
- WintApiError no longer retains full response bodies (bounded 300-char
  diagnostic; bodies can carry customer data and errors get logged).
- sie-builder refuses to render structurally invalid vouchers (missing
  account number or booking date) and documents deleted-voucher gaps in a
  #PROSA record per BFL 5 kap 6-7 §.
- Account classification: 20xx is equity, 83xx is financial income.
- SIE validator accepts EUBAS97 as BAS-based (standard kontoplanstyp; it
  previously produced a false non-BAS warning on every WINT/Bollbok file).
- New tests: resolveConsent WINT refresh flow, credential upsert payload
  (no mail/password persisted), WINT fetch failure path, EUBAS97 warning
  regression, builder invalid-data rejection, vi.clearAllMocks hygiene.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(import): pin EUBAS97 acceptance to the exact SIE spec value

Review follow-up on PR #1446: match EUBAS97 exactly instead of any
EUBAS* prefix, so the non-BAS kontoplan warning stays pinned to the four
kontoplanstyp values the SIE 4B spec enumerates (BAS95, BAS96, EUBAS97,
NE2007) rather than silently accepting unknown future variants.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 11:07:14 +02:00
Mattsson 707d597b2e feat(woocommerce): store order/refund feed extension (#1442)
* feat(woocommerce): store order/refund feed extension

Connect a WooCommerce store via the wc-auth key handshake (manual key
fallback) with per-store consumer key/secret AES-256-GCM encrypted at rest,
and import paid orders and refunds into the transactions inbox as a
bank-style feed on the 1680 cash account. Feed-only: nothing auto-books,
gateway fees/payouts are out of scope (core wc/v3 does not expose them).

Sync is cursor-paginated on modified_after (offset pages only inside
same-second date_modified ties), terminates on an empty page, holds the
cursor below failed refund fetches / ingest errors / deadline-skipped work,
checks the time budget between refund fetches, and drops rows dated on or
before bookkeeping_locked_through on every run. Nightly cron gated on the
extension registry + new paid capability woocommerce_sync (backfilled to
existing bank_sync grant holders).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migrations): move woocommerce migrations past main's 20260806090000

origin/main gained 20260806090000_recurring_schedule_interval_months while
this branch was in flight; identical version timestamps abort the Supabase
apply, so the two new migrations move to 20260806170000/20260806170100.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(woocommerce): resolve CodeRabbit review findings

- callback 503s early when WOOCOMMERCE_CREDENTIALS_ENCRYPTION_KEY is
  unset: encryptCredential would otherwise throw after the probe and
  strand the pending row without error_message
- disconnect and upstream-revoke clear the encrypted consumer key/secret:
  nothing reads them after revoke and keeping decryptable dead
  credentials is unnecessary retention
- manual sync gets a 240s time budget and the panel reports a truncated
  run as 'partial, sync again' instead of a normal completion
- listOrderRefunds terminates on an empty batch (hosts may cap per_page),
  dedupes by id against hosts that ignore page, and caps total pages
- unparseable money strings count as errors and log instead of being
  silently identical to a zero total
- pg test uses per-run unique store URLs so committed rows cannot hit
  the store_url partial unique index across pg-real runs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(woocommerce): resolve CodeRabbit cycle-2 findings

- listOrderRefunds throws when the page cap is exhausted with data still
  flowing, instead of returning a silently partial list the sync cursor
  would advance past; the error routes into the existing held-cursor
  refund-retry path
- partial sync results keep the row-error count, and the partial toast
  string surfaces it (ICU plural, hidden at zero) in both locales

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: retrigger CI after dropped push event

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 23:30:00 +02:00
Mattsson a0ca692fed feat(invoices): quarterly, half-yearly and yearly recurring invoice schedules (#1438)
* fix(mcp): offer the link tool in the uncategorized-transactions VAT blocker

The gnubok_vat_close_check blocker hint only named categorize/auto-match,
both of which create new bookkeeping. For a transaction whose
affarshandelse is already booked on an existing verifikat, following the
hint would double-book, so agents dead-ended the case into "contact
support" (2026-08-06 support mail from Orto Engineering). The hint now
also names gnubok_link_transaction_to_journal_entry, is extracted as an
exported constant pinned by a test, and the tool joins the
categorize_month recommended loadout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(invoices): quarterly, half-yearly and yearly recurring schedules

User request: recurring invoice schedules only supported monthly cadence.
Adds interval_months (SMALLINT 1-12, default 1) to
recurring_invoice_schedules; the UI offers manadsvis/kvartalsvis/
halvarsvis/arsvis presets while API and MCP accept any 1-12.

The cron advances next_run_date by whole intervals from the due date, and
the new rollNextRunDateForward() helper rolls missed or edited interval
schedules on their own month grid so a quarterly Jan/Apr/Jul/Oct schedule
missed in an outage rolls Jan 15 to Apr 15, never Feb 15. Monthly
(interval 1) keeps its existing today-anchored recompute semantics
unchanged. Changing the interval alone never touches next_run_date: the
new cadence applies from the next run, so an edit can never pull a send
earlier.

Existing rows default to 1 and behave byte-identically. The MCP slice of
this feature (interval_months on the three recurring-schedule tools in
server.ts) was committed in d2600907f alongside the VAT-blocker hint fix
by a parallel session sharing this worktree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoices): address PR #1438 review findings

CodeRabbit round 1, all three findings:
- MCP descriptions now state the full accepted interval range (any integer
  1-12) instead of enumerating only the 1/3/6/12 presets, and qualify that
  changing ONLY interval_months leaves next_run_date untouched.
- assertValidCadence rejects fractional day_of_month.
- rollNextRunDateForward rejects calendar-invalid anchors that pass the
  shape regex (2026-13-05, 2026-02-31), with regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 15:43:54 +02:00
Jakob Wennberg d41ef2a909 feat(sandbox): seed payroll, articles and a year of ledger history; calm the connect CTAs (#1437)
* feat(sandbox): seed payroll, articles and a year of ledger history; calm the connect CTAs

The sandbox showed neither Löner nor a usable set of reports, and the
"connect X" surfaces were oversized boxed cards.

Sandbox seed:
- pays_salaries + employer_registered, so Löner and Anställda appear at all
  (an enskild firma is not an employer by default). Both seeded employees are
  employment_type 'employee': an EF may employ staff, just not its own owner.
- Two employees, one booked and one open lönekörning, and the three verifikat
  the booked run must have posted (7210/2710/1930, 7510/2731, 7290+7519/
  2920+2940). Skatteavdrag comes from the real Skatteverket 2026 tables.
- Year-to-date ledger history, January through last month, with the quarterly
  momsredovisning cleared to 2650 and paid on the SFL deadline. Without the
  settlement the demo collected VAT all year and never remitted it, which left
  an implausible bank balance and 155 813 kr of moms "att betala".
- The history is exempted through journal_entry_no_doc_required, the same way
  the SIE-import opt-in treats imported books: its kvitton live in the previous
  system, and unflagged it put 39 "verifikat utan underlag" on the home screen.
- Artikelregister, and the BAS accounts the K1 chart omits for an enskild firma.
- History is numbered before the invoice and payroll vouchers so the series runs
  forwards through the year, and its writes are batched.

Connect CTAs:
- Bank picker: a two-column grid of 95px bordered logo cards becomes flat
  hairline rows, Lucide icons, and a quiet inline connecting state.
- Cloud backup: each provider collapses to one row; the BFL note is shown once
  for the section and names only configured destinations.
- Hem first-run: only the active step argues its case, but every not-done step
  keeps a reachable action. The Skatteverket nudge becomes one quiet sentence.

Mobile assistant FAB: a fresh open is desktop-only, since the bottom nav already
has an Assistent tab. A collapsed session keeps its handle everywhere except
/chat, which is itself the way back to the conversation.

Also closes a real hole: /api/salary/runs/[id]/payslips/send had no sandbox
guard, and a seeded booked run put "Skicka lönebesked" one click from an
anonymous visitor with live Resend behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sandbox): check the two unchecked Supabase errors and tighten review nits

CodeRabbit review on #1437.

Major: two calls discarded their error and continued with null data. A failed
chart_of_accounts re-select would have written account_id: null onto every
ledger-history and salary voucher line, and a failed next_voucher_number would
have inserted a posted verifikat with no number, which is a hole in the
verifikationsserie (BFNAR 2013:2). Both now throw, and a null voucher number is
rejected explicitly.

Minor: the A-004 note claimed a 10 % markup on numbers that are 11.1 %; the
salary breakdown test's name said the opposite of its assertions after the
switch to the real tax table; the ledger-history doc still said 4 to 6 verifikat
per month before the quarterly momsredovisning added a seventh in March, May and
June.

Bank picker: the spinner is aria-hidden, so loading and connecting had no text
equivalent and a failed bank fetch was never announced. Added role="status" with
an sr-only label, and role="alert" on the error line.

Declined: confirm-before-disconnect on the cloud-backup row. Disconnect was
unconfirmed before this PR too, so adding a dialog is a behaviour change beyond
the redesign rather than a fix to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 15:30:04 +02:00
Jakob Wennberg 5b0ca3d874 fix(copy): make K3 and year-end claims match what the code actually does (#1431)
* fix(copy): make K3, leasing and year-end claims match what the code does

Follow-up to the batch that removed the uppskjuten-skatt posting on
obeskattade reserver (K3 29.37 gross in juridisk person) and added the K2
asset-account gate. Six user-facing strings still described the old
behaviour or made claims the code cannot support.

1. Arsredovisning page: the K3 explainer promised an uppskjuten skatt-not
   and a materiella anlaggningstillgangar-not in every K3 document. Both
   are conditional (a 2240/8940 balance, assets in the register) and the
   first is now absent in the normal case. The kassaflodesanalys is
   dropped with a warning when it cannot be generated, so it is named
   only when the document actually carries one.

2. Regelverk settings: kassaflodesanalys was presented as following from
   K3. It follows from being ett storre foretag
   (swedish-year-end-closing/references/reporting-and-filing.md:10,
   legal-framework.md:42); the copy now says the product includes one and
   states the storre-foretag rule separately. Komponentavskrivning was
   presented as optional under K3; it is mandatory where component useful
   lives differ materially (k2-vs-k3.md:5, asset-accounting
   references/depreciation.md:33).

3. Note 1 and the Uppskjutna skatter-not no longer claim the 2240 balance
   is hanforlig till obeskattade reserver. deriveLatentTaxMovement reads
   the 2240/8940 balances only, and under K3 that account carries deferred
   tax on all temporary differences (k2-vs-k3.md:11-13).

4. The deferredTax 'unknown' branch emitted the gross-reserve statement,
   which is the denial phrased positively: the same affirmative claim
   about books that could not be read. It now emits no deferred-tax
   paragraph at all; build-data already warns on that path.

5. Capitalized-lease detection looked at 1260/1269 only. On the shipped
   BAS 2026 chart 1260 is a free inventarier account and 1269 is ack.
   avskrivningar pa datorer, so owned computers were reported as leased,
   while 1217/1227 (finansiellt leasade) were missed. Detection now reads
   the company's own account names in kontogrupp 12, which is where BAS
   keeps capitalized leases (leasing-and-disposal.md:28) and which owned
   inventarier on 1220 never matches. 1720 forutbetalda leasingavgifter
   stays out: that is the operational treatment.

6a. gnubok_year_end_readiness listed FX revaluation as a blocker (it is a
   warning) and omitted UNBOOKED_TRANSACTIONS, the common one. The
   description now names every actionable blocker kind, within the
   280-char budget, and a test pins it against YEAR_END_BLOCKER_KIND.

6b. companies.accounting_framework defaults to 'k2', so every enskild
   firma hit the K2 asset gate and was handed a BFNAR 2016:10 punkt 10.4
   citation plus a K3 remedy it cannot take: a sole trader prepares ett
   forenklat arsbokslut, not an arsredovisning (legal-framework.md:29,
   :48). entity_type now rides along on the companies read the routes
   already do, and non-AB entities get wording with no citation and no
   K3, keeping the 1090 remedy. The K1 counterpart of punkt 10.4 is not
   sourced in the repo skills, so nothing was invented in its place.

* fix(copy): close the review findings on the copy-truth sweep

Three follow-ups from the source and code reviews. (1) The K2/K3 help text had upgraded a vague sentence into a definite boundary claim ('gransen gar vid <trosklar>'), which excludes the other routes into mandatory K3 that are live right now for this control's audience: noterade vardepapper, and from fiscal years starting after 2025-12-31 also utlandsk filial, kryptotillgangar, aktierelaterade ersattningar and fastighetsbolag. An AB in one of those categories would have read the sentence and stayed on a regelverk it may no longer use. (2) hasCapitalizedLeaseAsset compared per-side cumulative totals, so a lease acquired earlier and disposed this year still claimed the balance sheet carries a leased asset; it now compares the net balance. (3) The K3 warning enumerated a kassaflodesanalys the document may not contain, contradicting the newly conditional page copy on the same screen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 11:54:57 +02:00
Jakob Wennberg 78a37f6396 fix(year-end): typed preflight blocker codes so remediation links render (#1420)
* fix(year-end): typed preflight blocker codes so remediation links render

validateYearEndReadiness emits Swedish blocker strings but the wizard's
BlockerRow matched English phrases, so no remediation link ever rendered,
and the voucher-gap branch pointed at /bookkeeping/voucher-gaps which only
exists as an API route. Blockers now carry stable machine codes end to end
(YearEndBlockerCode on YearEndValidation.blockers, mirrored additively as
blockerItems on BokslutReadinessReport); errors stays the plain string
mirror so the v1 compliance check and MCP tool keep their exact shapes.
BlockerRow matches on code and links only to pages that exist; the
voucher-gap and dead-link branches are removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(year-end): code the unbooked-transaction blockers #1414 added

#1414 landed two new blockers in validateYearEndReadiness using the old
errors.push style, which this branch had already renamed to a typed
blockers array. Merging main left them referencing a variable that no
longer exists.

Converted both to the typed scheme: UNBOOKED_TRANSACTIONS (the safety
guard that stops executeYearEndClosing from aborting at the step 7 lock
AFTER the closing entry posted at step 4) and UNBOOKED_CHECK_FAILED (the
fail-closed variant). Neither behaviour changes; both keep their Swedish
wording verbatim.

The MCP year_end_readiness classifier now routes on the stable
YearEndBlockerCode instead of regexing the Swedish message, with the
wording heuristic kept as a fallback for an unmapped or legacy English
message. The public `kind` values are unchanged, so MCP consumers see the
same output; both new codes map to 'unbooked_transactions' as before,
since an agent reacts to "we could not tell" the same way it reacts to a
real count.

UNBOOKED_TRANSACTIONS gets a /transactions remediation link in the
preflight step: that page is where a transaction is booked or marked
private, the two remedies the message names. UNBOOKED_CHECK_FAILED gets
none: the remedy is to re-run the check.

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 10:04:42 +02:00
Jakob Wennberg 576ed4d290 fix(agent): let the assistant see answers the user gave in WhatsApp (#1425)
Two field findings from the first live receipts.

1. The assistant re-asked for information the user had already given.
   The user answered the representation question in WhatsApp ("Elias
   Karlsson från Canguro Media, Jakob Wennberg från Arcim"), the answer
   was stored correctly on invoice_inbox_items.channel_context, and then
   the in-app assistant said it could see no participant names and asked
   for them again. The intent's inbox query selected only document_id and
   extracted_data, and nothing under lib/agent/ read channel_context at
   all. It is now selected, threaded onto each underlag as chat_answers,
   and rendered into the prompt as "uppgivna av användaren" with an
   explicit instruction that human answers outrank anything read off the
   image and must never be re-asked. Also backfilled by document_id: a
   receipt can reach the intent through the document paths without its
   inbox row being matched to the transaction.

2. The representation question accepted half an answer in silence.
   Naming participants but no purpose stored purpose=null and replied
   "Tack!", leaving the deduction undocumented while looking complete.
   Skatteverket wants both (BFL 5 kap 6-7 §). It now asks once, for the
   missing half only, and keeps the question open so the reply routes
   back to the same receipt. Anti-loop: the follow-up fires only when no
   representation block exists yet, so a second incomplete answer is
   taken as-is rather than nagging.

Tests cover the prompt half and the query half separately: the earlier
prompt tests injected chat_answers directly and would have stayed green
with the column still missing from the select, which is precisely how the
bug shipped. Both mutation-checked.

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 20:08:15 +02:00
Jakob Wennberg c84f951a5c fix(whatsapp-inbox): name both ways to route receipts for multi-company senders (#1424)
Field feedback after the first live receipt: the link confirmation told a
multi-company sender only to set a default company in the panel, so the
per-receipt path looked unsupported even though it is the one that
actually runs when no default is set. Now it names both: set a default,
or send one receipt at a time and answer the company question after each.

The guidance also moves to its own paragraph. Run together with the AI
disclosure it read as one sentence, which is how a real user came away
believing the bot had called itself a "mänsklig AI-assistent".

New copy test asserts the promise (both options present, disclosure kept
in its own paragraph, single-company message unchanged) rather than the
exact wording, so a rewrite stays free but a dropped option fails.
Mutation-checked against the previous copy.

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 19:43:37 +02:00
Jakob Wennberg c0825e9bd2 fix(bokslut): surface unbooked transactions and AR/AP tie-outs in year-end preflight (#1414)
* fix(bokslut): surface unbooked transactions and AR/AP tie-outs in year-end preflight

Two gaps in the year-end readiness layer:

1. Unbooked bank transactions were enforced only by lockPeriod, which runs
   at step 7 of executeYearEndClosing, AFTER the closing entry has posted at
   step 4. A period with unbooked transactions reported ready: true from
   gnubok_year_end_readiness and the wizard, then aborted mid-flow, leaving
   a posted closing entry on an unlocked, unclosed period. The readiness
   check now runs the same counter as the lock guard
   (countUnbookedInPeriod, so the number reconciles with the "att bokföra"
   badge) as a blocking error, failing closed if the check cannot run. The
   lockPeriod guard stays as defense in depth. The MCP classifier tags the
   new blocker as kind unbooked_transactions.

2. The Phase-1 avstamningar (kundreskontra vs 1510, leverantörsreskontra vs
   2440) existed as reports (lib/reports/ar-reconciliation.ts,
   supplier-reconciliation.ts) but were wired only to the ledger report
   routes, never to the bokslut preflight. The readiness aggregator now runs
   both tie-outs and surfaces mismatches as warning-severity reminders with
   deep links, mirroring the bank-reconciliation reminder. Warnings only,
   never blockers: a difference can be legitimate (FX-settled partials).
   Skipped entirely for kontantmetod companies, where open invoices are
   deliberately not on 1510/2440 until the year-end conversion exists and
   the tie-out is permanently unreconciled by construction. Unconvertible-FX
   rows produce a "could not reconcile" message instead of a phantom
   difference.

YearEndValidation gains an optional unbookedTransactionCount field; the v1
compliance endpoint and MCP readiness tool pick the new blocker up
automatically since they share the same engine.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): classify the next-period-IB readiness blocker instead of kind other

The blocker "Nästa räkenskapsperiod har redan ingående balanser bokförda"
was the only validateYearEndReadiness error with no classifier regex, so it
always surfaced as kind: 'other'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(bokslut): log swallowed AR/AP tie-out failures in the readiness aggregator

Compliance-review finding: a rejected tie-out produced no reminder and no
log entry, making a failed avstämning control indistinguishable from a
reconciled one. Still degrades to no reminder (advisory check), but the
rejection reason is now traceable, mirroring the unbooked-transaction
check's logging.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 18:05:38 +02:00
Jakob Wennberg 88760ae6f6 fix(whatsapp-inbox): harden against adversarial review findings (#1342)
* fix(whatsapp-inbox): erase the WhatsApp channel on account deletion

whatsapp_phone_links relied on the auth.users ON DELETE CASCADE, but
Accounted never deletes auth.users: account deletion is
anonymize_user_account plus a ~100-year ban that keeps the auth row as a
tombstone, so the cascade never fires and nothing revokes the link. After
erasure the link stayed active with a decryptable phone_enc,
lookupActiveLink kept resolving the number, and every further inbound
message was persisted with body_text and the verbatim raw_payload while
the bot kept replying: GDPR Art 17 plus continued collection with no
lawful basis.

The RPC is re-created verbatim from 20260724150000 with one added block
that revokes and crypto-shreds the link, resets its conversation, nulls
body_text/raw_payload on that link's messages and deletes outstanding
link codes, plus a guarded repair pass for tombstones anonymized before
this migration. Covered by a pg-real test that fails against the previous
definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): pepper the link-code hash and bound code minting

hashLinkCode stored a bare sha256 over CODE_ALPHABET^6 = 30^6 values
behind a fixed 'AC-' prefix. The module cited the invite-token pattern,
but invite tokens are 256-bit random; this space enumerates offline in
about a second, so hashing at rest protected nothing. The sibling
phone-crypto.ts already states the team's own threat model for a LARGER
space ("a plain sha256 would be brute-forceable ... hence the pepper"),
so link codes now hash through the same env-mandated pepper.

/link/start was also an authenticated unbounded INSERT that left every
earlier code valid. Minting now burns the caller's unused codes (the code
the panel shows is the only one that works) and is capped per TTL window,
with the route answering 429 instead of throwing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): harden the conversation layer against the review findings

Pre-merge hardening of the unshipped chat layer. Every change below has a
test that fails without it.

Lifecycle and races:
- conversation writes go through updateConversation(), an optimistic
  compare-and-set on updated_at (the trigger makes it a revision counter).
  The ack winner, the answer worker, the pin refresh and the sweep hold
  different claims, so blind whole-jsonb writes resurrected answered
  questions, wiped pending_question and dropped queue entries.
- terminal markStatus writes are guarded on processing_status='processing'
  so a losing worker cannot overwrite the winner's 'done' and null its
  inbox_item_id.
- the message -> inbox item path is idempotent: a pre-check plus a 23505
  fallback adopt the item a concurrent worker created, instead of throwing
  after the WORM document is already committed.
- PROCESSING_STUCK_MS 90s -> 5 min. The enforced step budget of one media
  row already exceeds 90s, so the sweep was re-claiming live workers.
- sweep 2b re-arms only when the conversation itself has been quiet, not
  just the rows: pending_ack=false plus unacked rows is also the state of a
  live finalize, which produced a duplicate combined ack.
- pin expiry re-checks against fresh state instead of writing back a stale
  whole context, which reverted company choices applied mid-pass.
- askNextQueuedQuestion claims the pop before sending, so two answer
  workers cannot ask the same question twice.

Company question:
- the state is rolled back when the M6 send fails, so the next receipt
  re-asks instead of parking receipts behind a question nobody received.
- applyCompanyChoice claims the open question (company_options) rather
  than the state: a double tap confirms once, a transient membership-query
  error is no longer read as "not a member", and a LATE answer still lands.
- at the 48h TTL the parked receipts are kept, not discarded: options and
  staged rows survive so a late digit or tap still files them, and only
  rows past Meta's ~30-day media window get the terminal marker.
- an out-of-range digit or a typed company name now gets the options
  repeated instead of silence or the "I cannot answer questions" reply.

Inline dispositions:
- stop/start/byt/company answers run their side effect BEFORE the terminal
  wamid row, with a SELECT pre-check for dedupe. Writing the row 'done'
  first made them at-most-once: a crash in between lost the action forever.

Copy and answers:
- acks state the extracted currency instead of labelling every total 'kr'.
- M17 stops promising "about 10 minutes" when the daily quota tripped.
- M18 is sent once per message tracked by the outbound row, so a file
  whose first attempt died still reaches the sender, including from the
  max-attempts path.
- M11 no longer claims the number is disconnected: 'stopp' pauses, and
  muted senders now persist no chat content at all.
- 'byt' is recognized in every state but awaiting_company (m6-confirm
  teaches the word, and it was being stored as answer data instead).
- text sent while a re-send question is open is kept as a note on THAT
  receipt with the question left open, instead of binding to another
  receipt's question.
- a quoted reply wins over the pending question and is appended when the
  quoted question is already answered, so corrections stop landing on the
  wrong receipt.
- context answers keep raw_answer + answered_at like representation does.
- finalizeBurst checks the send result: on failure it rolls the question
  back and leaves the rows unacked for the sweep.

PII:
- the sender's plaintext number is stripped from raw_payload before it is
  persisted; replies decrypt the link's phone_enc instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(whatsapp-inbox): record the erasure path and the hardening decisions

RoPA gains the account-deletion row (immediate, not via the cron: the
auth.users cascade never fires because the row is tombstoned) plus the
two new security measures, and its "never in the clear" phone claim is
now true of the stored payload. DECISIONS.md records the non-obvious
calls: revoke-not-delete on erasure, commit-then-roll-back for the
company question, keeping expired company choices answerable, the
compare-and-set conversation write, effect-before-terminal-row for inline
dispositions, honest M11 copy, and the raw_payload redaction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): stop the answer re-claim from following a confirm with M16

A worker that died after applying an answer and sending its confirmation
leaves the row 'processing'. The sweep re-runs it, resolveAnswerTarget
finds the question already answered, and the user got "I did not
understand" immediately after the confirmation they had just received.
The fallback is now first-attempt only.

The catch comment claiming the sweep retries these rows is corrected
too: 'error' is terminal for the sweep, and nothing on the answer path
throws anyway (interpretChatAnswer degrades, sends never throw,
supabase-js returns errors), so the catch is a programming-error net.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): drop the amount floor on the representation question

The Swedish compliance review on #1340 caught a real error in the trigger
rules: the representation question only fired above 150 kr, but the duty to
document deltagare and syfte is what makes the expense deductible at all
(BFL 5 kap 6-7 §) and it is not conditioned on any amount. The 300 kr per
person figure I had in mind is the VAT-deduction base cap, a different rule.
A 120 kr business lunch would have been booked with no participant trail,
which is exactly the deduction Skatteverket denies later.

Noise stays bounded by the triggers that were already there: the question
fires only for receipt-shaped documents from restaurant, cafe or hotel
merchants, at most once per receipt, twice per burst and six times per
sender per day, and a single "nej" dismisses it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
2026-08-05 15:47:23 +02:00
Jakob Wennberg bf5ca2c615 feat(whatsapp-inbox): GDPR retention cron and RoPA entry (#1341)
PR5b, the final code piece of the WhatsApp intake track. A daily cron
(04:15) enforces the channel's retention table; the receipt itself stays
7-year WORM under BFL and is never touched.

Retention actions (lib/retention.ts, each isolated and idempotent):
- whatsapp_messages transcripts past 90 days: body_text + raw_payload
  cleared in id batches under a wall-clock budget; the row skeleton
  (wamid, direction, timestamps, status, inbox_item_id) survives for
  audit. Only rows still carrying content match.
- Rows with phone_link_id IS NULL (unknown senders, orphans) past
  30 days: deleted.
- Link codes expired more than 24h ago: deleted, used or not.
- Sender rate counters idle 2+ days: deleted (minute/day window keys
  are dead weight after that).
- Links revoked 90+ days ago: phone_enc crypto-shredded to '' (column
  is NOT NULL), one-shot via neq guard; phone_hash and phone_masked
  kept for uniqueness history and audit display.

Route mirrors the sweep cron exactly: withCronContext + registry gate
(503 EXTENSION_DISABLED when the extension is off). vercel.json gets
the schedule and both Docker crontabs are regenerated.

Compliance: new whatsapp.receipt_intake activity in .compliance/ropa.yaml
covering purpose, Art 6(1)(b)/(c)/(f) bases with the Art 14(5)(b) note
for third-party attendee names, Meta Platforms Ireland as processor
(Cloud API, EU SCC addendum, Local Storage region DE), the differentiated
retention table, and security measures.

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 15:35:41 +02:00
Jakob Wennberg 629069e281 feat(whatsapp-inbox): conversation layer with clarifying questions (#1340)
PR4 of the WhatsApp intake track: turns the per-message PR3 pipeline into a
conversation. Media replies are burst-debounced into ONE combined ack (M4
single / M5 numbered list) sent by the single winner of the atomic
pending_ack claim; losers stay silent. Multi-company senders get the company
question (reply buttons <=3, list 4-10, numbered text >10) with an 8h
sliding pin ('byt' clears it); their receipts park as staged message rows
until the answer and then run through the normal intake path.

Clarifying questions are evaluated per receipt after extraction, max one per
receipt, priority unreadable > representation > partial, keyed on the
Phase-0 classification (legibility/documentKind/merchantCategory) with
heuristic fallbacks (compressed-chat-photo signal, extended meal regex).
Budgets: <=2 content questions per burst, <=6 per sender per Stockholm day;
over budget acks only and flags the item moved_to_app. Questions expire
after 48h (sweep, silent hand-off) and are asked exactly once.

Free-text answers route through the ONE new LLM call
(lib/interpret-answer.ts): Sonnet via Bedrock, max_tokens 600, no thinking,
forced tool call validated by Zod with hard caps, gated by
checkAgentRateLimit, reply framed as untrusted data. Any failure degrades to
storing the raw text as a note; exact 'nej' short-circuits without the LLM.
Answers land in invoice_inbox_items.channel_context
(representation/user_note/quality) with ChannelQuestionAsked/Answered
processing-history events. Late answers match by quoted wamid or the most
recent open question within 7 days.

New per-minute sweep cron (registry-gated physical route, 503
EXTENSION_DISABLED when off) re-claims stuck rows (max 3 attempts), rescues
crashed burst acks, expires questions and pins. One new migration
(20260802210000) adds whatsapp_messages.acked_at, the relational burst-
membership marker, with pg-real coverage for the single-winner claim.

Verified: full vitest suite (12270), pg-real against a migrated
supabase/postgres 15 (977), lint 0 errors, tsc at the 405 baseline,
check:guards green, crontabs regenerated. Mutation-checked the debounce
claim and the daily budget gate.

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 15:27:00 +02:00
Jakob Wennberg d1c411ad6f feat(invoice-inbox): surface WhatsApp chat context in booking flows (#1339)
* feat(invoice-inbox): surface WhatsApp chat context in booking flows

The WhatsApp intake bot writes verified human answers (photo caption,
representation deltagare + syfte, sender note, open-question state) to
invoice_inbox_items.channel_context. This makes the in-app booking flows
READ it:

- New core renderer lib/documents/channel-context-notes.ts: deterministic
  compact Swedish line ("Representation: Anna Berg (Volvo), Jakob W ·
  Syfte: uppföljning av avtal"), capped at 220 chars by dropping whole
  participant names ("… och N till"), never mid-name. Representation
  first, then user_note; caption only when nothing else exists.
- FieldsRail "Från WhatsApp" block in InvoiceInboxWorkspace: caption,
  deltagare, syfte, anteckning rows plus an ochre AttnLine when a chat
  question expired unanswered (pending_question.status = moved_to_app).
- Notes threading: book-direct and convert default their notes to the
  rendered string server-side when the request carries none (a supplied
  value always wins); BookDirectlyDialog prefills its notes input with
  the same string so the user can edit it before it lands. Bulk-book
  (categorize-core) joins the shared batch note with the per-item
  rendered context so the representation trail survives batch booking.
- Inbox list: whatsapp rows get a chat icon and a quiet "Fråga obesvarad"
  badge for moved_to_app items. No worklist count change: unresolved
  whatsapp items are already counted by countInboxDocuments.
- sv/en strings for every new key; renderer + route + bulk tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoice-inbox): honor cleared notes and keep unreviewed captions out of verifikat

Two adversarial-review findings on the WhatsApp surfacing flows, both about
text that lands on an immutable verifikat.

1. A cleared note was silently re-applied. BookDirectlyDialog prefills the
   rendered chat context, and the dialog sent `notes.trim() || undefined`
   while book-direct and convert defaulted from channel_context on any falsy
   value. A user who read the prefill, disagreed and deleted it therefore got
   it written back onto a posted entry, removable only through a formal
   rättelse. Both code comments claimed "an edited value always wins", which
   was false for exactly that edit. Now PRESENCE of the field decides: the
   dialog always submits `notes` (empty string included) and the routes only
   default when the field is absent from the request (MCP, older clients).
   The Zod `.optional()` carrying that distinction is documented at the
   schema so it is not "tidied" into a `.default('')` later.

2. The photo caption was auto-burned into verifikat text with no review.
   renderChannelContextNotes fell back to the raw caption, and bulk-book
   appended the result per item with no per-item notes field at all (the MCP
   approval preview deliberately shows no per-item PII either), so unreviewed
   chat text reached a WORM record nobody had seen. The renderer now takes
   { includeCaption } and leaves the caption out by DEFAULT: representation
   answers and user_note are replies to a question the bot asked, the caption
   is not. Only the Bokför direkt prefill opts in, where the user reads the
   string in an editable field before booking.

Tests: cleared-notes and whitespace-cleared on book-direct, cleared-notes on
convert, caption-never-defaulted on both routes, caption-not-threaded in
bulk-book, and the renderer's opt-in. The book-direct cleared-notes tests
were mutation-checked (restoring the truthiness fallback fails them).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(reports): export the chat answers behind a verifikat in the full archive

The verifikat line caps the representation trail at 220 chars and drops whole
participant names ("… och N till"); the complete list (deltagare, syfte,
raw_answer) exists only in invoice_inbox_items.channel_context. That table was
in ARCHIVE_EXCLUDED_TABLES with a rationale predating channel_context ("inbox
workflow state"), so a company leaving Accounted and keeping the full-archive
export as its BFL 7-year record kept an incomplete deltagare documentation for
its representation deductions.

Dumped as a column PROJECTION, not the whole row: the new
MasterDataTableSpec.columns narrows the select to the underlag provenance
(document, matched transaction, created verifikat / leverantörsfaktura) plus
channel_context, so the answers are tied to what was booked from them while
the inbox workflow state (email bodies, OCR output, error messages) stays out
of the archive. The documents themselves remain in dokument/.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 14:54:30 +02:00
Jakob Wennberg 398c734b93 feat(whatsapp-inbox): intake extension with webhook, phone linking and receipt ack (#1338)
Webhook lifecycle: GET hub.challenge handshake (constant-time verify-token
compare); POST verifies X-Hub-Signature-256 over the RAW body before any
parse, Zod-parses the envelope, persists inbound rows (partial-unique wamid
= dedupe against Meta's up-to-7-day redelivery), acks 200 fast and defers
media processing via the after() idiom. Rejected and rate-limited content
always acks 200 and lands as skipped/error rows, never a retryable status.

Linking: the settings panel (Installningar -> WhatsApp) mints AC- one-time
codes (sha256 stored, 10 min TTL, single use, ambiguity-free alphabet); the
webhook consumes the code, binds phone to user (HMAC-peppered hash + AES-256-
GCM at rest) and confirms with M3. Keyword commands stopp/start/hjalp;
unknown senders get one throttled M1 greeting (1/h, 3/day) behind the
sender-quota RPC, with no media download and no content persistence.

Intake worker: atomic claim on the message row (the durable job record),
company resolution (default -> sole membership -> M6 fallback, no item),
per-company inbox quota (ack-and-drop, M17 once per 10 min per sender),
MIME allowlist, 10 MB stream-checked media download, exact sha256 duplicate
check, then the shared uploadAndExtract funnel (source 'whatsapp',
channel_context caption, whatsapp_message_id) and the M4 ack with extracted
merchant/total/date. Failures wrap to 'error' + error_message + one M18.

uploadAndExtract widened: source 'whatsapp', optional channelMeta + actorId;
email/upload paths behaviorally unchanged.

Deferred to PR4: burst debounce + combined ack (M5), in-chat company choice
(M6 buttons + 8h pin), clarifying questions M7-M10, interpret-answer LLM
call, sweep cron, retention cron.

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 14:47:08 +02:00
Jakob Wennberg 02e48efe11 refactor(invoice-inbox): export uploadAndExtract from lib (#1336)
Pure move of uploadAndExtract and its private helpers (sanitisers,
page-count/slice, sandbox check, MIME/size consts) from index.ts into
lib/upload-and-extract.ts so a future channel extension can import the
shared funnel the way document-extraction already imports
extract-invoice-fields. No behavior change; existing tests prove it.

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 14:28:03 +02:00
Jakob Wennberg 3b3adf96c7 feat(invoice-inbox): receipt-aware extraction + OCR pipeline fixes (#1331)
Receipts and invoices were extracted through one invoice-shaped prompt
with no document classification. The extractor now also returns
documentKind, payment method (+ card last4), purchaseTime,
merchantCategory and legibility, validated with .catch(null) so a
hallucinated label degrades to unknown instead of sinking the parse.
The FieldsRail shows type and payment method above the editable fields.

Pipeline fixes, all verified against real failure paths:
- PDFs >3 pages: extract from a pdf-lib slice of the first 3 pages
  instead of skipping entirely (issue #553 gate); truncation recorded
  in extracted_data.pages and shown in the UI.
- Oversized images (>4 MB, over Bedrock's 5 MB cap): downscale to
  <=2000px JPEG via sharp before base64, instead of erroring to an
  empty result.
- HEIC/HEIF: attempt sharp transcode to JPEG; when libvips lacks HEIF
  (prebuilt binaries), fall through to today's behavior but show an
  explicit hint instead of silently blank fields.
- Oresavrundning: prompt rule + totals.roundingAmount so receipt totals
  reconcile with subtotal+VAT for exact-amount transaction matching.
- retry-extraction overwrites extracted_data wholesale: a confirm now
  guards against silently destroying manual field edits.

Deliberately NOT added: retry-on-transient-Bedrock-error; the SDK
already retries twice by default (maxRetries=2).

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 14:20:40 +02:00
Jakob Wennberg 9f5a43310b fix(salary): recompute entitled_days on existing ledger rows and record pre-cutover taken days (#1403)
The vacation ledger sync carried entitled_days verbatim on existing open
rows while re-deriving accrued and taken, so a stale entitled value (for
example the flat 25 stored before Semesterlagen 7 § pro-rating existed)
survived every sync. The recompute loop now re-derives entitled the same
way the lazy-seed path does, with the opening-balance cutover still
outranking recomputation for the year containing cutover_date.

Opening balances could also not record paid vacation days already taken
in the cutover year under the previous payroll system. New additive
column employee_opening_balances.vacation_days_taken_this_year (NUMERIC
NOT NULL DEFAULT 0, CHECK 0..40) threaded through the shared service,
the Zod schema, the MCP staging tool (schema + mergeable fields), the
staged-operation executor, the v1 REST routes, and the employee editor
form. Ledger semantics for the cutover year, on both seed and recompute
paths: entitled = remaining + taken_this_year, taken = booked-run taken
+ taken_this_year, so remaining keeps meaning remaining and the seeded
value survives every subsequent sync.

Fixes #1347

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 19:35:02 +02:00
Mattsson 00ae3540db feat(customers): carry contact person and invoice copy recipients through migration (#1392)
* feat(customers): carry contact person and invoice copy recipients through migration

Extends the arcim-migration entity mapper, Fortnox provider mapper, canonical
DTOs, customer APIs (web + v1) and invoice send flows so contact person and
customer-level invoice CC/BCC addresses survive provider migrations. NULL
means unconfigured and empty means an explicit clear, so re-syncs enrich
legacy gaps without resurrecting deliberately removed values. Fortnox fixed
assets are split into a dedicated follow-up issue.

Fixes #1345

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(db): bump customer metadata migration past pack-slug version

Main already contains 20260803230000; keep new versions strictly newest so
Supabase branching applies them in order.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(customers): complete Customer type consumers and make enrichment payload resolvable

The preview-pdf mock customer and the makeCustomer fixture now carry the
three new metadata fields, fixing the type-check failure in Build (zero
extensions) and Vercel.

The enrichment update in the migration orchestrator now spells its payload
as an object literal typed CustomerMetadataEnrichment (absent keys drop at
serialization), so the phantom-column guard resolves the columns instead of
counting another unresolvable dynamic payload past its ceiling. The cc/bcc
guards also verify element types instead of casting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 10:00:03 +02:00
Jakob Wennberg 267ed6c1bb feat(mcp): expose the konteringspaket catalogue as a resource (#1395)
An agent proposing a booking currently invents the account numbers, and a
plausible guess (6071 instead of 6072, or the full cost instead of the 80%
deductible share) produces a verifikat that posts and is wrong. This turns that
into naming a reviewed template.

Read from packs/ rather than the database on purpose: legal_note has no column
in booking_template_library, so the database copy cannot answer "when does this
template apply", and that note is the one field an agent cannot derive from
account numbers.

Carries the amount maths explicitly (vat lines from vat_rate, everything else
from ratio) with an instruction not to fudge an amount to force a balance, plus
notes that entity_type is binding, that a company may hold templates beyond this
list, and that posting goes through the journal-entry tools so period locks and
the balance check still apply.

A load failure returns an explicit error rather than an empty list: an agent
shown 12 of 26 templates concludes the other 14 do not exist and hand-rolls
accounts for them.

Verified the packs are actually reachable at runtime before relying on this: a
production build traces all 26 YAML files into the serverless function, so the
loader is not a local-only convenience.

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 22:17:27 +02:00
Mattsson 281cb3989d fix(build): break Turbopack chunk-name hash collision from #1385's import edge (#1394)
Since the #1385 squash-merge every production build failed with 'Two or
more assets with different content were emitted to the same output path'
on [root-of-the-server]__1ge0sz5._.js: two distinct server chunk groups
(an AWS smithy helper chunk and the withRouteContext auth chunk) hash to
the same chunk name. The graph change that tipped the chunk layout into
the colliding state was entity-mapper.ts (arcim-migration extension
entry graph) importing lib/vat/supplier-invoice-line-checks, which
drags lib/money into the extension root.

Move normalizeVatRateToFraction into an import-free leaf module
(lib/vat/vat-rate-unit.ts), re-export it from
supplier-invoice-line-checks for all existing callers, and point
entity-mapper at the leaf. Behavior is unchanged (511 targeted tests
pass); the server chunk graph returns to the pre-#1385 shape that
builds cleanly. Verified: npm run build fails on ff864ad3d and passes
with this change.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 19:28:15 +02:00
Mattsson 8443062b1f fix(vat): enforce fraction unit for supplier invoice vat_rate writes (#1385)
Closes the remaining #310 write paths: credit-note item copies (web, v1,
pending-operations) and arcim-migration supplier imports now normalize
vat_rate to the decimal-fraction unit before storage, and a NOT VALID
CHECK constraint guards every new supplier_invoice_items row. Customer
invoice items deliberately stay percent; legacy supplier rows are left
untouched so posted-entry reversals reuse the exact original values.

Fixes #310

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 18:49:46 +02:00
Mattsson 8299ee9fb4 fix(bookkeeping): resolve settlement account in all categorization flows and ship mis-booking audit (#1383)
Completes the #985/#986/#987 caller sweep: categorize-core, v1
batch-categorize, pending-operation edits and the MCP categorize path now
resolve the settlement leg from the transaction's cash account instead of
inheriting a hardcoded or stale account. Extends the correct_entry preview
with currency, tax and dimension line metadata so staged corrections
preserve full line fidelity. Adds a read-only audit query and a runbook for
reviewing and correcting historical mis-bookings via staged storno with
explicit approval; no automated bulk mutation.

Fixes #1001

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 18:45:38 +02:00
Mattsson 5d7952a01e feat(mcp): model-free document upload via signed URL (#1378)
* feat(mcp): model-free document upload via signed URL (#748)

Adds gnubok_create_document_upload + gnubok_complete_document_upload so
document bytes reach storage through a short-lived signed PUT URL and
never pass through the model context. Fixes silent base64 corruption on
real-size PDFs and the context blowup on batch uploads.

- pending/ staage keys with TTL cleanup; completion validates magic
  bytes + SHA-256, moves bytes to the WORM key and adopts the reserved
  UUID as document id, making retries and concurrent completions
  idempotent
- legacy gnubok_upload_document kept for clients without file access,
  description now points to the signed-URL pair; shared mime resolution
  and inbox-item creation extracted
- both new tools mapped in TOOL_SCOPE_MAP (transactions:write) and
  MCP_TOOL_CAPABILITY_MAP (ai) so the paywall and scope gates hold
- payload guard ceiling 58.5K to 59K after trimming the create tool's
  outputSchema to upload_id/upload_url/expires_at

Fixes #748

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): satisfy capability-map lock and phantom-column scanner

The exact-entries lock in capability-maps.test.ts now includes the
signed-URL pair as dispatch-only AI tools, and the inbox insert uses a
literal payload (explicit UUID instead of a conditional spread) so the
no-phantom-columns scanner can resolve every column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 18:41:05 +02:00
Mattsson cd7d7f52b9 feat(invoices): per-recipient email delivery outcomes (#1384)
* feat(invoices): per-recipient email delivery outcomes

Resend delivery webhooks identify affected addresses in data.to, so one
message with CC recipients can carry independent To/CC outcomes instead
of masking the failing address into the aggregate reason text.

- new apply_invoice_delivery_provider_event RPC merges each reported
  recipient onto its immutable To/CC position with the same rank and
  timestamp ordering as the aggregate status (retry and out-of-order safe)
- recipient map is PII-free: keyed to:N / cc:N, BCC and unmatched
  recipients are never represented, and the map is cleared on PII redaction
- delivery summaries, API route and MCP tool expose the sanitized map;
  the route re-sanitizes as defense in depth
- UI shows a per-recipient status list under the aggregate outcome

The prod ops check in issue #1350 (webhook registered in Resend and
RESEND_DELIVERY_WEBHOOK_SECRET set in Vercel) cannot be verified from the
repo and remains a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(invoices): commit provider event before cross-context read

The BCC-leak test applied the event inside the rollback-scoped service
role helper and then asserted through a separate member context, so the
applied status was rolled back before the read. Use the committing
runAsServiceRole helper for the apply, matching how the summary read is
performed in its own context.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 17:56:37 +02:00
Mattsson 8a2498a987 fix(mcp): add true continuation to invoice list tools (#1327)
* fix(mcp): add true list continuation

Signed-off-by: Emil <emilmattsson14@gmail.com>

* test(mcp): cover pagination edge cases

Signed-off-by: Emil <emilmattsson14@gmail.com>

* fix(mcp): validate pagination offsets

Signed-off-by: Emil <emilmattsson14@gmail.com>

* fix(mcp): preserve continuation without counts

Signed-off-by: Emil <emilmattsson14@gmail.com>

---------

Signed-off-by: Emil <emilmattsson14@gmail.com>
2026-08-01 19:03:51 +02:00
Mattsson bfbd926950 fix: paginate MCP inbox items (#1329)
* fix: paginate MCP inbox items

* fix: validate MCP inbox cursors

* test: reset MCP inbox pagination state
2026-08-01 18:50:42 +02:00
Mattsson bcf919d995 fix: expose inbox filenames via MCP (#1328)
* fix: expose inbox filenames via MCP

* fix: handle joined inbox attachments
2026-08-01 18:30:27 +02:00
Mattsson fd376eff94 fix: audit Cloud Backup OAuth redirects (#1324)
* fix(cloud-backup): pin OAuth callback origin

* fix: reject non-web cloud backup origins
2026-08-01 16:08:37 +02:00
Jakob Wennberg 0f1c7c9365 fix(import): refuse a Bokio connection that opens a different company (#1315)
A Bokio integration token is scoped to one Bokio company and the company id is typed in by hand, so credentials for the user's other company imported that company's customers, suppliers and invoices with no error at all. Probe /companies/{id} before storing, mirroring the Bjorn Lunden /details probe, and refuse on a confident org-number mismatch.

Also surface the inbox mail body when nothing was attached: it was captured in email_body_text and never read back, which made Gmail's forwarding-confirmation mail unreadable and the forward impossible to complete.
2026-07-30 19:07:40 +02:00
Jakob Wennberg 4a38fa30ed fix(recon): share the cash-account scope and cover the no-1930 case (#1309)
PR #1295 (144cc514) fixed computeVatCloseCheck on main while this branch was
fixing it a second, different way. This rebuilds the branch on top of that
merge instead of re-landing the duplicate: main's local
getScopedReconciliationStatus stays as the MCP entry point, its lookup body
moves into lib/reconciliation/cash-account-scope.ts, and the pieces main does
not have are added on top.

Why the lookup has to live in lib/: the bokslut readiness aggregator has the
same defect and is core code, which must never import from @/extensions/. It
called getReconciliationStatus with 4 positional args, so cashAccountId stayed
undefined and scopeTransactionsToAccount fell through to its currency-only
filter: the bank side summed every SEK cash account while the GL side stayed on
1930, and the wizard reported "Bankavstamningen visar en differens" with zero
unmatched transactions and zero unmatched GL lines to point at. Same shape as
the MCP blocker in #1290, different surface.

Decisions taken deliberately, not by taking 'ours':

1. Lookup errors fail CLOSED, everywhere. resolveCashAccountScope throws
   "Kunde inte hamta kassakonto <n>" instead of returning the unscoped
   fallback. The earlier version on this branch logged and fell back, which
   turned a transient DB error or an RLS denial straight back into the #1290
   pooling path. Main's contract wins; its merged test asserting exactly that
   still passes untouched.

2. The default resolution no longer hard-codes 1930. With no account_number
   argument the resolver tries 1930 and, only if the company has no such row,
   falls back to its primary cash account. Measured read-only on prod
   2026-07-30: 2 companies have no 1930 cash_accounts row while running two SEK
   cash accounts each and zero journal_entry_lines on 1930, so the check
   compared their entire SEK bank volume against an empty GL side, i.e. a
   high-severity bank_unreconciled blocker with count 0 that no user action
   could clear. Roughly 20x the difference the issue reported. A caller that
   NAMES an account gets no fallback, so gnubok_get_reconciliation_status still
   rejects "Okant kassakonto 9999" rather than silently answering about a
   different account. 1367 companies have a 1930 row and are unaffected,
   including the 5 whose 1930 row is not the primary one.

3. The blocker message names the resolved account instead of a literal 1930:
   pointing a user at 1930 when the reconciliation ran on 1935 sends them to an
   account with no lines on it.

4. warnIfUnscopedAcrossCashAccounts logs a warning when a run left
   cashAccountId undefined AND the rows it fetched really do span more than one
   cash account. Kept on the write path too: an unscoped runReconciliation can
   persist a wrong journal_entry_id, which does not clear itself later.

Duplicate regression suite collapsed: the branch's
vat-close-check-bank-scope.test.ts overlapped main's
vat-close-check-reconciliation-scope.test.ts case for case, so only the cases
main lacked were merged into main's file (the primary-account fallback, the
message naming the resolved account, the blocker still firing on a genuine
scoped difference, and the tool handler's own scope resolution).

Residuals are now tracked issues, not code comments:

- #1298: the post-sync runReconciliation sweeps in
  app/api/extensions/enable-banking/sync/cron/route.ts and
  extensions/general/enable-banking/index.ts still run unscoped. They are write
  paths.
- #1299: booked transactions with a NULL cash_account_id whose verifikat has no
  line on the primary account still inflate the bank total. Measured all-time
  on prod: 294 booked NULL rows, 22 of them on verifikat with no 1930 line,
  4 companies, net -4170.31 kr with monthly swings from -18055.82 kr to
  +38086.00 kr. Fixing it means finishing the cash_account_id backfill.

Also: the file-global logger mock in bank-reconciliation.test.ts now wraps the
real module and swaps only warn, instead of substituting a four-method stub
whose child() returned undefined for the entire module graph of that suite.

Verified: npx vitest run over lib/reconciliation, lib/bokslut,
extensions/general/mcp-server, app/api/reconciliation, app/api/extensions,
app/api/v1, app/api/bookkeeping, app/api/transactions, lib/pending-operations,
lib/bookkeeping, lib/invoices, lib/transactions, lib/reports: all green.
eslint on the 8 changed files: 0 errors (18 pre-existing unused-import
warnings in server.ts). tsc --noEmit: 405 errors, byte-identical to the
origin/main baseline. check:guards passes. No migration, so no pg-real test.

Fixes #1290

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 18:50:19 +02:00
Mattsson 144cc51458 fix(mcp): scope VAT close reconciliation account (#1295)
Resolve the VAT close reconciliation scope using the selected cash account, currency, and unassigned-transaction behavior. Add regression coverage for cross-account leakage and fail closed on lookup errors.\n\nCloses #1290
2026-07-30 11:29:50 +02:00
Mattsson 6318501b71 fix(vat): recover ruta 05 for null-rate custom accounts (#1296) 2026-07-30 11:28:50 +02:00
Mattsson 17a7a62ceb fix(reports): stop the resultatavslut zeroing declarations, and make the mistake uninventable (#1293)
* fix(settings): explain why account deletion is blocked

The delete-account button was disabled while the user still owned
companies, but the reason only lived behind the "?" on the blocker row,
so the greyed-out button read as broken. Surface it as one visible attn
sentence directly under the button, and point aria-describedby at it
whenever the button is disabled, not only on a load error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(enable-banking): share one PSD2 consent across a user's companies

Connecting the same bank for a second company required a second BankID, and
at SEB that new authorization silently revoked the first one. A user with four
companies at one bank therefore signed four times a quarter and ended up with
three dead feeds, each still rendering as "Aktiv" with a stale last_synced_at
until someone pressed Synka.

Prod says this is not one customer: every SEB customer holding connections in
more than one company has had an earlier company stop syncing at the moment
the next was authorized, most of them while the consent was still formally
valid for weeks. The same measurement over other banks is far quieter, so the
one-active-session-per-PSU limit is real and ASPSP-side.

Enable Banking already supports the shape we want. POST /auth carries no
account restriction, so a session covers every account the user ticked at the
bank, and GET /accounts/{uid}/transactions takes no session id, so a second
company can sync its own accounts from an existing session. bank_connections
has no unique constraint on session_id, so this needs no migration.

Adds lib/session-sharing.ts plus GET /reusable-sessions and POST /attach. When
a live session in another of the user's companies still exposes accounts no
company syncs, the settings panel offers to reuse it: the new row shares
session_id and consent_expires, carries only the unclaimed accounts, and lands
in pending_selection so the existing IBAN-aware account picker does the ledger
mapping. Only the consent is shared; accounts, cash_accounts and transactions
stay strictly per-company.

Sharing a session changes three lifecycle paths, all handled here:

- Disconnect and reconnect now refcount before revoking. A blind revoke would
  take down a sibling company's feed, which is the exact failure this removes.
  The count runs on a service-role client because RLS hides a sibling in a
  company the user has since left, and it fails closed: an uncertain count is
  treated as shared, since a lingering consent lapses on its own in 90 days
  while a wrongly revoked one kills a working feed.
- A renewed consent fans out to every company sharing the old session, and
  re-points their account uids by IBAN. Several ASPSPs reissue uids on
  re-authorization, so carrying the session id alone would have left siblings
  calling retired uids and re-broken them every quarter. This is also why the
  superseded session_id is no longer nulled at /connect: the callback needs it.
- The nightly probe runs once per distinct session and applies the verdict to
  every row holding it, and expiry mails are keyed per (user, session), so one
  dead consent is one probe and one mail rather than four of each.

Only enabled cash_accounts rows count as claiming an IBAN. The callback mirrors
every account in a consent, deselected ones included, so counting any row as a
claim would leave nothing offerable once the first company connects.

An account handed to a company also stops being offered while that company's
picker is still open, closing the window where two companies could book the
same physical account.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ink2): read the resultaträkning from the pre-closing books

INK2R summed journal entries raw, so it included the resultatavslut that
zeroes every P&L account into 2099 at year-end. Nettoomsättning, kostnader,
periodiseringsfond and skatt all came out as 0, which cascaded into INK2S
7650/7651 and the taxable result. INK2 is always filed after bokslut, so
this was every real declaration, and nothing warned: with the P&L at zero
the balance sheet still tied out.

INK2R now reads two views of the same period. The balance sheet comes from
the closed books so 7302 keeps arets resultat via 2099; the income statement
comes from the pre-closing books via excludeFinalClosingEntry, which drops
only fiscal_periods.closing_entry_id so skatt and bokslutsdispositioner stay
on the form (7525, 7528). The equity adjustment is now conditional on a
posted closing entry having moved the result into 2099.

Second, independent bug: accounts were mapped by BAS number with no regard
for the sign of the balance, so konto 1630 with a credit was reported as a
negative fordran instead of a skatteskuld and konto 2641 with a debit was
netted off the liabilities. The three sign-reclassification rules the K2
iXBRL mapper already had are extracted to lib/reports/sign-reclassification
.ts and applied to INK2R too, so both statutory reports present the same
balance sheet. Only the rule table is shared: k2-mapper keeps its sumOre
arithmetic because the iXBRL path is ore-exact while INK2R truncates per
SFL 22:1.

NE-bilaga had the same empty-resultatrakning bug and gets the same fix.

Adds the closed-period coverage that was missing: the old tests only
exercised the mapping table against an open period, the one state in which
the engine happened to work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(reports): make the year-end closing decision explicit at every call site

generateTrialBalance took two optional booleans, so a caller that never
thought about the resultatavslut silently got 'include'. That is the wrong
default for anything summing class 3-8: the closing verifikat posts the
mirror image of every P&L account into 2099 inside the same period, so the
report reads ZERO across the board while the balance sheet still ties out
and nothing warns.

The booleans are replaced by a required
closingEntry: 'include' | 'exclude-final' | 'exclude-all-year-end'
with no default, so the build fails until each call site decides. All 40
were audited individually; every one keeps its current behaviour except
the two that were provably broken:

  - Resultatrapport read zero on every line for a closed year, in JSON,
    PDF and XLSX, and its prior-year comparison column read zero for
    anyone whose previous year was closed.
  - Resultat per projekt (dimension-pnl) had the same defect and must
    stay in lockstep with Resultatrapport to keep reconciling.

Both now pass 'exclude-all-year-end', which keeps them agreeing with the
formal Resultaträkning rather than pre-empting Stage 2 of #1051
(DECISIONS.md:632).

Deliberately unchanged and recorded in DECISIONS.md: the KPI expense
composition, which is blank for a closed year but cannot be fixed without
a migration and a displayed-figure change, and getBookedBolagsskatt, whose
contract is an open period and whose call chain already caused a
too-high-tax customer bug once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(vat): keep the resultatavslut out of the momsdeklaration

The closing verifikat posts the mirror image of every P&L account into
2099 inside the same fiscal period. Revenue accounts drive rutor 05, 39
and 40, so any VAT period containing the fiscal-year end reported NEGATED
turnover once the year was closed. get_vat_declaration_totals already
excluded vat_settlement and opening_balance entries, but not this one.

Reproduced read-only against production: for December of a closed year
the December declaration reported ruta 39 = -794 734 kr. After the fix
that period reports 0 and the January period carrying the real sale is
unchanged at 794 734 kr.

Keyed on fiscal_periods.closing_entry_id, not source_type = 'year_end':
avskrivningar, periodiseringsfond and skatt share that source_type and
must keep whatever VAT effect they carry. A reversed closing entry is
retained together with its storno so the pair still nets to zero, the
same predicate trial-balance.ts uses for closingEntry: 'exclude-final'.

Migration applied to the staging branch only; prod gets it via merge.
The pg test is written but has NOT been executed locally (no DATABASE_URL
configured and no local Postgres), so CI is its first real run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(kpi): keep the resultatavslut off the monthly chart

The monthly income/expense chart summed every posted entry in the fiscal
period. The closing verifikat posts the mirror image of every P&L account,
so once a year was closed the fiscal-year-end month charted the whole
year's revenue as negative income.

Measured read-only on production: 28 companies across 34 month-rows. The
worst case charted December income as -10 347 459,81 kr where the real
figure is +12,88 kr. Other examples: -1 868 731 -> +128 730,
-1 850 501 -> +431 709.

Both paths are fixed together so they keep agreeing: the RPC's monthly
section now joins the tb_ex_ye_entries CTE it already computes for
tb_ex_year_end, and monthly-breakdown.ts (the dimension-filtered fallback
and the MCP path) gains the matching source_type filter plus the
storno/correction chain of REVERSED year-end entries, so an undone bokslut
does not leave half a pair behind.

Migration 20260723180000 had recorded the omission as deliberate, on the
grounds that it mirrored the JS scan. It did, but the JS scan was wrong.

Migration applied to the staging branch (function body identical; three
comment lines differ from the committed file). Prod gets the file via merge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(reports): pin every statement generator against a closed fiscal year

The per-generator suites all exercised an OPEN fiscal period, which is the
one state in which a generator that forgets the resultatavslut happens to
work. Declarations are filed AFTER bokslut, so the untested state was the
only state that occurs in production. That is why the same defect could
ship three times.

Two new suites over one shared fixture (closed-year-fixture.ts, a synthetic
closed AB with a resultatavslut, a credit 1630 and a debit 2641):

  closed-year-statements.test.ts enumerates the generators and asserts each
  reports the year's revenue rather than zero, plus its own bottom line. The
  table IS the checklist: a new report either appears in it or nothing stops
  it shipping with this bug. Verified by regressing income-statement back to
  closingEntry 'include', which fails 2 of its assertions.

  cross-surface-agreement.test.ts asserts the surfaces agree with each
  other, which is what every customer complaint actually was. INK2R and the
  K2 årsredovisning must produce the same årets resultat, the same fritt
  eget kapital, the same sign reclassifications and the same balance total.
  The operational family (Resultaträkning, Resultatrapport) must agree
  internally, and the gap BETWEEN the families is asserted explicitly as
  bokslutsdispositioner + skatt, so when Stage 2 of #1051 lands the test
  names the expectation to change instead of failing vaguely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(guards): ratchet against new reports that scan the ledger directly

A statement generator that aggregates journal_entry_lines itself has to
remember, on its own, that the resultatavslut posts the mirror image of
every P&L account into 2099 inside the same fiscal period. Three forgot,
and each read ZERO revenue for a closed year while the balance sheet still
tied out, so nothing warned.

generateTrialBalance now requires an explicit closingEntry mode, which makes
that decision a compile error. This guard is what keeps NEW reports on that
path: any generator under lib/reports or lib/bokslut that reads
journal_entry_lines and is not in the baseline set fails CI. Verified by
adding a throwaway report, which the guard rejects by name.

Voucher and line listings (general-ledger, journal-register, SIE export,
reconciliation, diagnostics) are sanctioned: they show the ledger as posted
and have no closingEntry decision to make.

Four existing lib/bokslut files are grandfathered rather than migrated. One
of them is a genuine open follow-up recorded in DECISIONS.md:
sarskild-loneskatt-calculator sums 7410-7419 with no year-end exclusion, so
its basis reads ~0 if it runs against an already-closed period. Left alone
deliberately: it is a tax figure whose call chain has caused a customer bug
before and deserves its own verified change.

Also ratchets naive-ore-round down 646 -> 641.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(reports): pin where sign reclassification applies, in both directions

No behaviour change. The sweep asked whether the 1630/2641 sign
reclassification should be extended to the remaining balance-sheet
surfaces; the answer is that there are none left.

Both STATUTORY presentations already have it: the K2 iXBRL årsredovisning
since 2026-07-23 and INK2R since 2026-07-29. The other two balance-sheet
surfaces must NOT have it: /rapporter Balansräkning and Balansrapport are
organised by account number under BAS-prefix headings, and balansrapport
documents an invariant that depends on every row staying debit-positive
where it was booked. Moving konto 1630 into a liability section would break
the add-the-rows-to-verify-the-balance property and hide the account from
anyone looking it up by number.

Asserting both halves is the point. The first half stops the
reclassification silently disappearing from one statutory surface again,
which is how a customer ended up comparing two of our own reports against
each other. The second half stops a future sweep "fixing" the operational
reports into disagreeing with their own documented contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(reports): detect statement disagreement instead of waiting for a customer

Every year-end problem reported so far was a DISAGREEMENT between two of
our own screens, not a single wrong screen. The årsredovisning said one
figure, INK2 said another, and the customer did the reconciliation for us.
Nothing in the product noticed, because each screen tied out on its own.

Two additions:

  INK2R self-checks. On a closed year it compares the årets resultat it is
  about to declare against the booked konto 2099, and warns in Swedish when
  they disagree. This is the alarm that was missing: when INK2R reported
  0 kr against a booked 469 542 kr, the balance sheet still balanced, so no
  warning fired. Mirrors the equivalent check k2-mapper has had since
  2026-07-23, so both statutory reports now catch the same fault.

  reconcileStatements + GET /api/reports/statement-reconciliation return
  årets resultat from every surface side by side, grouped into families.
  ledger + statutory must agree and a mismatch is named; operational
  legitimately differs by bokslutsdispositioner + skatt until Stage 2 of
  #1051 lands, so that gap is explained rather than flagged.

The visual panel is deliberately not built here: it needs a
/frontend-design pass against the locked concept conventions plus sv/en
strings, and the warning above already puts the alarm where the user looks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(reports): address review findings from PR #1293

pg-real (7 failures, one signature): the new fixture called
insertFiscalPeriod({ isClosed: true }) and then inserted journal entries
into it, so enforce_period_lock (migration 017, legally required) refused
the write. Not worked around: the RPC's predicate keys on
fiscal_periods.closing_entry_id and never reads is_closed, so the fixture
now links the closing entry and leaves the period open, which exercises the
path that actually matters.

CodeRabbit, closed-year-fixture: EX_YEAR_END_ROWS dropped only the P&L legs
of the year_end entries (8811, 8910) and left their balance-sheet legs
(2125, 2512) at pre-closing values, so the 'exclude-all-year-end' view sat
160 000 kr out of balance and misrepresented what generateTrialBalance
returns. Latent, because today's consumers read class 3-8 only, but a shared
fixture that does not balance is a trap for the next consumer. Both legs now
go, and a new test asserts all three views sum to zero.

CodeRabbit, INK2 totals: renamed totals.resultAfterFinancial to
aretsResultat. It holds the result after bokslutsdispositioner AND skatt,
which is årets resultat, not resultat efter finansiella poster, and
build-data.ts uses the old name correctly for the different subtotal. The UI
already labelled the value "Årets resultat", so the name was simply wrong.

CodeRabbit, statement-reconciliation: the statutory branch called a
generator and caught any throw as "wrong entity type", mapping genuine
failures to a null figure that the comparison then skipped, so a real bug in
a declaration generator made the function report isReconciled: true. That is
the opposite of its purpose. It now dispatches on entity_type and surfaces a
generation failure as a named disagreement.

CodeRabbit, enable-banking (Emil's call to include): fetchClaimedIbans
returned an empty Set on a cash_accounts read failure, which is
indistinguishable from "nothing is claimed" and made every IBAN in the
session offerable, including accounts another company already books to. Its
own comment said it failed closed and its log said "offering nothing"; it
failed open. Returns null now, and findReusableSessions offers nothing when
the claimed set is unavailable. The test that pinned the fail-open asserted
toHaveLength(1) under the name "offers nothing"; it now asserts []. Also
removed an em dash per CLAUDE.md.

The remaining enable-banking finding (consent-expiry cooldown stamped only
on the selected connection, so it leaks one duplicate mail per sibling
company) is deliberately left to Emil: it changes email-sending behaviour in
his feature rather than fixing a stated contract.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(reports): resolve second-round review findings on PR #1293

pg-real, two NEW signatures (the closed-period one from cycle 1 is gone):

kpi-report-aggregates-rpc.pg.test.ts asserted the exact contract migration
20260730090000 deliberately changes. Its comment read "year_end entries are
NOT excluded from monthly" and expected December expenses 1250. That fixture's
December holds only year-end-chain entries, so with the fix the month drops
out of the chart entirely, which is the correct operational view: a month
whose only activity is bokslut has no operating result. Assertion and file
docstring updated to the new contract rather than the test being removed.

vat-totals-closing-entry.pg.test.ts passed the wrong account arrays. p_net_
accounts is VAT_SETTLEMENT_NET_ACCOUNTS (2650/1650, the momsredovisning
settlement pair), not the output-VAT accounts. Putting 2611 there made the
extra year_end entry match the settlement-SHAPE detector, so an ordinary
sale-with-VAT was classified a momsredovisning and dropped, and the test read
0 instead of 10 000. The RPC was right; the fixture was not.

CodeRabbit, statement-reconciliation: resolveEntityType checked neither
query's error, so a genuine DB failure (RLS, permissions, connectivity)
returned null indistinguishably from "no entity type set", fell into the
unsupported-form branch and reported isReconciled: true. That is the same
silent-false-reconciled bug the cycle-1 refactor closed, one level down. The
companies error now throws; a missing company_settings ROW stays tolerated,
because .single() errors on zero rows and many companies have none. Mirrors
the pattern the INK2 and NE engines already use.

Still open by Emil's explicit choice: the consent-expiry cooldown is stamped
only on the connection it was handed, so it leaks one duplicate mail per
sibling company on the shared session. That changes email-sending behaviour
in his feature rather than fixing a stated contract, so it stays his.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:03:05 +02:00
Jakob Wennberg 16f34fb214 fix(arcim): resolve OAuth redirect_uri identically in authorize and exchange (#1287)
* fix(arcim): resolve OAuth redirect_uri identically in authorize and exchange

The authorize leg honored the FORTNOX_REDIRECT_URI / VISMA_REDIRECT_URI
override while the token-exchange leg hardcoded the NEXT_PUBLIC_APP_URL
fallback. After the app-domain cutover (2026-07-21) the Fortnox env var
still pointed at app.gnubok.se while NEXT_PUBLIC_APP_URL moved to
app.accounted.se, so the two redirect_uri values differed and Fortnox
rejected every code exchange (RFC 6749 4.1.3). The failure was invisible:
the error popup posted its message from the old-domain origin, the
wizard's event.origin check dropped it, and the popup closed itself.

Both legs now resolve through one resolveArcimCallbackUrl() helper, and
the error popup stays open with the reason on screen so a dropped
postMessage can never again turn into "nothing happens".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: log OAuth popup and redirect-uri rollout decisions

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:47:07 +02:00
Jakob Wennberg ef25a87d75 feat(mcp): Tasks extension (io.modelcontextprotocol/tasks) (#1283)
* feat(mcp): speak spec revision 2026-07-28 (stateless core)

Adopt the 2026-07-28 MCP spec revision on the connector endpoint while
keeping every handshake-era client (2025-06-18 and earlier) byte-identical:

- Accept per-request _meta protocol negotiation
  (io.modelcontextprotocol/protocolVersion); unsupported versions return
  UnsupportedProtocolVersionError (-32022) with the supported list.
- Implement server/discover (spec MUST): supported revisions, capabilities
  including the extensions field, identity, instructions, freshness hints.
- Decorate results for stateless clients: required resultType, serverInfo
  in _meta, and CacheableResult ttlMs/cacheScope on tools/list,
  prompts/list, resources/list, resources/read.
- Validate the standard Mcp-Method/Mcp-Name request headers when present
  (HeaderMismatchError -32020); absence stays accepted.
- Declare the ratified MCP Apps extension (io.modelcontextprotocol/ui) in
  capabilities; the widgets already use the ratified mime type and
  _meta.ui.resourceUri shape, so no widget changes are needed.
- OAuth: include the RFC 9207 iss parameter on every authorization
  response (success and error) and advertise
  authorization_response_iss_parameter_supported in RFC 8414 metadata.

Resource-not-found already used -32602 and tools/list ordering was already
deterministic; both are covered by the new test file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(mcp): Tasks extension (io.modelcontextprotocol/tasks)

Durable handles for long-running MCP tool calls, per the official Tasks
extension. A client that declares the extension in its per-request
capabilities gets a CreateTaskResult (resultType: "task") immediately;
the work completes after the response via after() and lands in the new
mcp_tasks table for tasks/get polling. Clients that did not declare the
extension are never handed a task (spec MUST).

- New mcp_tasks table (migration 20260729094000): company-scoped SELECT
  RLS, service-role-only writes (mirrors pending_operations), 1-hour
  expiry, status lifecycle CHECK. pg-real coverage included; triaged as
  excluded in the full-archive backup contract (transient state).
- tasks/get (creator-scoped), tasks/cancel (cooperative, working-only
  flip), tasks/update (ack no-op: no input_required flows yet).
- Tool opt-in via shouldRunAsTask predicate; first producer is
  gnubok_audit_package, the one genuinely long-running blocking call
  (multi-minute ZIP generation). estimate_only stays synchronous.
- Tool failures complete the task with the standard isError envelope,
  exactly what the synchronous call would have returned; the failed
  status stays reserved for infrastructure errors.
- server/discover and initialize now advertise the tasks extension.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): creator-only task RLS, enforced expiry sweep, RoPA entry

Compliance-swarm follow-ups on the mcp_tasks migration (editing the
migration is safe: it has not shipped beyond the ephemeral PR preview):

- SELECT RLS tightened from company-wide to auth.uid() = user_id so the
  DB grant matches the creator-scoped tasks/get contract; task results
  carry raw tool output (Art. 5(1)(c)). pg test now proves a same-company
  colleague cannot read the row.
- The 1-hour retention is now enforced, not aspirational: createMcpTask
  opportunistically deletes expired rows on every creation
  (idx_mcp_tasks_expires), best-effort (Art. 5(1)(e)).
- RoPA entry mcp.async_task_handles added to .compliance/ropa.yaml
  (Art. 30).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): literal terminal-update payload for the phantom-column guard

The conditional spreads in resolveMcpTask made the payload unresolvable
for the no-phantom-columns guard (362 > 360 ceiling). A literal payload
writing null for absent terminal fields is equivalent here: the terminal
transition sets the complete terminal state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:15:06 +02:00
Jakob Wennberg 4501f118c2 feat(mcp): approval-queue MCP Apps widget for staged operations (#1278)
* feat(mcp): approval-queue MCP Apps widget for staged operations

gnubok_list_pending_operations(render_ui=true) now renders an interactive
approval queue (claude.ai / Claude Desktop) where the user approves or
rejects each staged operation with a click. High-risk operations arm the
approve button and the second click sends confirmed=true, so the BFL
5 kap 5 acknowledgment is a first-party human action instead of the
agent asserting confirmed=true on the user's behalf (the audit weakness
flagged in dev_docs/erpclaw_analysis.md).

- New widget ui://pending-operations/app.html following the established
  self-contained postMessage/JSON-RPC pattern (no fetch, theme-aware,
  Swedish labels, expandable preview_data per row).
- Result-level _meta.ui hint gated on render_ui=true, mirroring the VAT
  report wiring; the tool stays data-only by default.
- Widget tool references project per namespace (accounted_* clients see
  accounted_ names inside the HTML).
- tools/list payload ceiling 58K -> 58.5K per the in-test convention:
  prose trimmed to the floor first, remainder is wire contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): time out the widget RPC bridge so a silent host cannot strand a row

Review follow-up: sendRequest never settled if the host dropped a
response, leaving op._working=true forever with the approve/reject
buttons gone. A 30s timeout rejects the promise; the existing catch
paths restore the row with an error message so the user can retry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:02:17 +02:00
Jakob Wennberg dc5aea4a35 feat(mcp): speak spec revision 2026-07-28 (stateless core) (#1277)
* feat(mcp): speak spec revision 2026-07-28 (stateless core)

Adopt the 2026-07-28 MCP spec revision on the connector endpoint while
keeping every handshake-era client (2025-06-18 and earlier) byte-identical:

- Accept per-request _meta protocol negotiation
  (io.modelcontextprotocol/protocolVersion); unsupported versions return
  UnsupportedProtocolVersionError (-32022) with the supported list.
- Implement server/discover (spec MUST): supported revisions, capabilities
  including the extensions field, identity, instructions, freshness hints.
- Decorate results for stateless clients: required resultType, serverInfo
  in _meta, and CacheableResult ttlMs/cacheScope on tools/list,
  prompts/list, resources/list, resources/read.
- Validate the standard Mcp-Method/Mcp-Name request headers when present
  (HeaderMismatchError -32020); absence stays accepted.
- Declare the ratified MCP Apps extension (io.modelcontextprotocol/ui) in
  capabilities; the widgets already use the ratified mime type and
  _meta.ui.resourceUri shape, so no widget changes are needed.
- OAuth: include the RFC 9207 iss parameter on every authorization
  response (success and error) and advertise
  authorization_response_iss_parameter_supported in RFC 8414 metadata.

Resource-not-found already used -32602 and tools/list ordering was already
deterministic; both are covered by the new test file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): Mcp-Name covers params.uri, base64 sentinel, version-header consistency

Review follow-ups against the transport spec text: Mcp-Name mirrors
params.name OR params.uri (resources/read), values arrive base64-wrapped
in the =?base64?...?= sentinel and must be decoded before comparison, and
an MCP-Protocol-Version header that disagrees with the _meta protocol
version is a HeaderMismatch. Absence of any header stays accepted since
this server supports handshake-era clients (spec-sanctioned leniency).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:02:05 +02:00
Jakob Wennberg 80a14ddfd2 feat(mcp): dimension parity for the write/read tool edges (#1274)
Closes the MCP dimension gaps found in the 2026-07-28 audit:

- gnubok_bulk_book_inbox_items accepts a shared dimensions bag through
  all three layers (tool schema + BulkBookInboxSchema + categorize-core
  BulkBookInboxInput), resolve-don't-select with echoed resolutions; the
  web inbox bulk-book route and the pending-op executor inherit it via
  the shared schema.
- gnubok_create_employee / gnubok_update_employee accept
  default_dimensions (names resolve to codes; {} clears on update).
  The command layer already persisted the field: only the MCP boundary
  blocked it, leaving payroll tagging dashboard-only.
- gnubok_query_journal: dimensions bag filter (jsonb containment via
  the GIN index, covers custom dims the legacy project/cost_center
  filters cannot) + include_dimensions to return each line's bag.
  The wide full-match fetch stays dims-free unless something needs it.
- gnubok_list_invoices / gnubok_list_supplier_invoices return
  default_dimensions (agents could set invoice bags but never read
  them back).
- Discoverability: create_voucher, categorize_transaction,
  correct_entry, update_invoice descriptions now name dimensions;
  categorize_month and invoice_run loadouts include
  gnubok_list_dimensions. Trimmed new schema prose to stay under the
  tools/list payload budget.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 09:39:53 +02:00