diff --git a/DECISIONS.md b/DECISIONS.md index 062337a9..e40902f3 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -1457,6 +1457,15 @@ One line per decision: `[YYYY-MM-DD] : `. Appended by agents and [2026-09-01] F2 bank-data staleness: ship freshness reads only (last_synced_at/consent_expires/error_message on gnubok_connect_bank + new GET /api/v1/.../bank-connections, scope companies:read mirroring the MCP mapping): the daily cron already syncs server-side, so visibility is what the incident lacked; an agent-triggerable sync is a product bet (EB call cost, runaway agents) and was deferred by Emil. [2026-09-01] Verifikationsserie in the Ny verifikation modal is a closed dropdown instead of a one-letter free-text field: a typo there silently opens a brand-new series with its own number sequence, and the letters only mean anything if everyone uses the same ones. The letters are NOT prescribed by law (BFL 5 kap. 7 § requires only unbroken systematic numbering within each series), and the incumbents disagree: Björn Lundén uses A Huvudserie, F Kundfakturor, I Inbetalningar, L Leverantörsfakturor, N Löner, U Utbetalningar, J Bokslut. We ship FORTNOX's table verbatim (A Redovisning, B Kundfakturor, C Inbetalningar från kunder, D Leverantörsfakturor, E Utbetalningar till leverantörer, F Kassa, G Avskrivning, H Periodisering, I Bokslut, J Revisor, K Lön, L Kontantfaktura, M Momsrapport), from their own Systemdokumentation, because Fortnox is the system most companies migrate here from and an imported ledger should keep its meaning. REJECTED an earlier draft that labelled A as Kundfakturor: A is the general series manual entries land in (the one point Fortnox and BL agree on, and Fortnox allows manuell kontering ONLY in A), and migration 20260526120700 ships every source_type defaulting to 'A', so every existing company's A series already holds everything. Calling it Kundfakturor would mislabel their entire history and the modal's own default. The list is closed but any letter the company already configured, or that a draft was saved with, is appended so no existing value can fall out of the picker. Also: tabbing or clicking into an untouched amount field now proposes the outstanding difference (pre-selected, so typing replaces it) when the row already has an account and the difference belongs on that side. This deliberately reverses part of the note in updateLine that said a balancing amount must never auto-fill: that note was about filling on ACCOUNT selection, which stole the amount before the user had a chance to split it. Filling on focus keeps the split case intact because the proposal is selected text, and it fixes the common moms case where the last line is just the remainder. [2026-09-01] Settings PUT cross-field VAT validations scoped to touched field groups (vat-completeness, 40m-monthly, periodisk sammanstallning), not fixed at onboarding: partial saves from surfaces without VAT fields (invoice bank-details dialog) were hard-blocked by pre-existing vat_registered-without-number state (Marketio Lab case). The invariant still holds on every save that touches its group; explicit null now counts as a clear instead of falling back to the stored value during validation. Onboarding-side VAT number collection left as follow-up. +[2026-08-24] Declined to enable Peppol access for Low-Stack Technologies (enskild firma, personnummer-based org nr): both send validation and receive registration deliberately refuse personnummer identifiers (GDPR; DIGG recommends 0088 GLN) and no GLN support exists, so enabling would only surface errors; request row left as 'requested' pending founder call. +[2026-08-26] invoice@arcim.io email unblock (Jakob's own test account, EXECUTED on prod, never re-run): the in-app delete had already anonymized + banned tombstone 7086c4d1 at 06:48 UTC, but the tombstone keeps auth.users.email by design (2026-07-24), which blocks re-signup. Replicated the 2026-07-30 Orback pattern: email scrambled to anonymized+@tombstone.invalid, the stale auth.sessions row deleted, AND the Google auth.identities row (sub 103263015636623485631) deleted rather than scrubbed: an OAuth identity matches on (provider, provider_id), not email, so a scrubbed email alone would have re-bound the next Google sign-in to the banned tombstone. Hard delete of auth.users was never an option: companies.created_by cascades and the 4 154 posted CashLeads Media AB entries are BFL 7 kap. 2 § retained (delete_user_account was dropped in 20260706100000 for exactly this). The 3 archived companies stay as an inert tenant with 0 members. +[2026-08-31] Demonstrationsinspelning (record-your-work) prototyped INSIDE the app rather than as ambient screen capture: an in-app recorder captures semantic events (chosen konto, momskod, dimension per invoice) instead of pixels, which removes the OCR/masking problem, keeps client personal data out of the recording entirely, and gives a deterministic compile target. Research the same day found continuous employee screen recording is the most legally hostile design in EU workplace privacy (IMY: regular real-time monitoring "som regel inte tillåtet"; consent invalid in employment; CNIL Amazon France 32 MEUR; byrå screens also carry client data no client DPA covers), while the horizontal market (Skan AI, 63 MUSD Aug 2026) is on-prem and works-council gated. Ambient capture parked behind two founder gates: a named design-partner byrå and acceptance of AI Act Annex III provider status (deferred to 2027-12-02 by the July 2026 Digital Omnibus). Founder has said high-risk status is acceptable. +[2026-08-31] Recording compiles to a RUTIN (process level), not only a flow (task level): a routine groups the recorded steps by surface in first-visit order and carries a per-step honest verdict (körs automatiskt / förbereds åt dig / stannar hos dig), with the field-level kontering rules nested inside the relevant step as a linked flow. Chosen over making the recording produce ever-larger node graphs because a byrå's unit of work is a work session, not a trigger, and because the canvas only fits about four columns; a routine is also where the trust ladder belongs (the month report is prepared but never auto-sent, and a final "Ditt godkännande" step is always appended). Rutin orchestrates, flow executes, agent supplies knowledge. +[2026-08-31] Core premise of the record-and-compile product VALIDATED against prod data before building anything (read-only aggregates, no PII): across 302 companies with 20+ posted 2026 entries (50 499 entries), 86.7 % of entries fall into an account-shape that recurs 3+ times inside the same company, and each company uses only ~30 distinct shapes for ~167 entries. Sharper test on the exact claim the demo makes: of 1 785 repeat counterparties (10 852 entries), 88.8 % of entries match the dominant shape for that counterparty and 69.5 % of counterparties are booked identically every single time. So a rule learned from one demonstration would be right ~9 times in 10, and the residue is exactly the "skickas till dig i stället för att gissas" case already modelled in the prototype. Redundancy check: categorization_templates is live for 222 of 509 active companies (proves appetite for the rule layer) but the PROCESS layer is empty (recurring_invoice_schedules used by 2 companies), so the routine altitude is the unoccupied one. +[2026-08-31] CORRECTION + sharper reframing of the record-and-compile bet, from prod data: 81.9 % of posted 2026 entries are source_type='import' (SIE history migrated from other systems), so the earlier 86.7 %/88.8 % repetition figures are dominated by imported books. Re-measured on NATIVE work created in Accounted (imports, storno, correction and opening balances excluded): 143 companies, 6 998 entries, 72.1 % of entries sit in a shape recurring 3+ times, and 427 repeat counterparties covering 2 291 entries are 91.6 % predictable from the counterparty alone. Cross-company: for counterparties seen at 3+ companies, 69.9 % of entries land on the same shape, so a brand-new company can be seeded from the anonymised crowd at ~70 % and converge to ~92 % on its own history. Strategic consequence: the recorder learns something the ledger mostly already knows, so the predictor (derive counterparty -> treatment from each company's own ledger, including the imported SIE) should ship BEFORE the recorder; the recorder is reserved for the process/sequence layer the ledger genuinely lacks, and the migration wedge should read the old system's data rather than watch a human use its screens. +[2026-08-31] Podcast (Neil Movva, Sail Research, Invest Like The Best) re-read properly: the thesis is long-horizon BACKGROUND agents, throughput over latency, and driving token cost toward zero by scavenging chips/power nobody else wants. Consequences adopted for Accounted: (1) accounting is an unusually VERIFIABLE domain and the DB triggers are already the grader, so the long-deferred agent-eval gym is the asset this argues is scarce, not a nice-to-have; (2) accounting work is the ideal background workload (nobody waits, batchable, retryable, P99-tolerant) so non-interactive AI belongs in a nightly batch lane behind the existing getAiService() seam in lib/ai/provider.ts, using the ~24 crons that already run 03:00-05:30; (3) "own your intelligence" is satisfied by an exportable per-company CONTEXT artifact (the derived konteringskarta), not fine-tuned weights, which matches Movva's in-context-learning-over-fine-tuning view and our own 91.6 % counterparty predictability; (4) cheap background tokens make consensus-as-confidence affordable (sample several models per proposal; agreement + historical-rule match earns the autonomy ceiling). REVERSES the earlier on-prem-hardware leaning: if background inference trends to near-free from aggregated scavenged supply, buying a box for a 20-person byrå is the wrong bet except where sovereignty is contractually required. +[2026-08-31] Counterparty logos: server-side resolve-and-cache (logo.dev or similar into the existing public `logos` bucket) chosen over shipping a curated static brand pack. The pack looked cheaper because it needs no API and no table, but it needs TWO hand-built artifacts (150 optimized assets AND the alias table mapping "ICA SUPERMARKET KUNGSHOLMEN 4711" to ica.se), never covers the long tail, and means redistributing 150 companies' trademarks from an AGPL repo. What survives from the pack idea is only the initials fallback plus a small manual override list. Direct browser-to-logo.dev was rejected outright: it is a cold third-party origin per user, a new underbiträde row on the privacy page, ad-blockable, and it contradicts the deliberate PostHog `/rl` same-origin proxy pattern in next.config.ts. The Supabase `logos` bucket wins on latency too, since the browser already holds a warm HTTP/2 connection to that origin for the data queries. +[2026-08-31] BrandMark tile is ACHROMATIC (bg-secondary + muted text), not a deterministic per-merchant colour hue: the palette is an achromatic foundation where semantic colour is data-only, and components/ui/provider-marks.tsx records that its Google/Microsoft marks are "the only coloured glyphs in an otherwise achromatic interface". Colour stays confined to the resolved-logo path so it is one toggle at visual sign-off. brandInitials() lives in a dependency-free lib/brand/initials.ts rather than reusing normalizeCounterpartyName(), because that module drags the Supabase client, VAT entry generation and the dimension resolver into the client bundle of the transactions page (PAGE_SIZE = 200 rows). Canonical normalization stays the server-side resolver's identity so learning and lookup keep one key. [2026-09-01] mcp.tool_called gets errorCause = errorCauseTag(err) on the two execution catch paths only (#2051): SQLSTATE or coded-error code, else the error class name, capped at 64 chars; a plain Error deliberately tags null because the class name 'Error' is noise, and pre-execution denials pass nothing since their errorCode already IS the vocabulary. Raw driver messages stay out of event_log on purpose: a constraint-violation message can quote row values. [2026-09-01] counterparty_aliases joins the categorization_templates audit-trigger strip list (20260901200000) instead of staying logged: prod falsified the original exclusion list within 30 minutes of 20260901103000 going live (15 of the first 16 UPDATE audit rows were alias+learning noise, ~800/day projected vs ~50/day of real rule changes), because the learning path merges aliases in the same write that bumps occurrence_count. Explicit trade-off: a human editing ONLY aliases is no longer logged; accepted since alias growth is overwhelmingly automatic and any change also touching accounts/VAT/pattern/active still logs (first real one, 19:02:17Z same day, captured correctly). Pre-fix noise rows stay in audit_log (append-only) and the read model stops labelling the column so they render as no-ops. [2026-09-01] MCP catalog budget attacked at the duplicated staged envelope rather than by demoting more reads: measuring the payload by segment showed outputSchema is 38 % of the whole catalog (23 290 tokens) and STAGED_OPERATION_SCHEMA alone 14 736 of it, the same envelope transmitted 58 times, while descriptions (what the three previous rounds trimmed) are only 10 %. period_status now carries its shape in one sentence instead of declared JSON Schema, matching actor/approve/preview which were always bare objects; 2 552 tokens reclaimed with no tool demoted and no field removed. Every edit is in the LOOSER direction because the server emits structuredContent for every tool and the documented failure mode is a declaration too tight making a strict client reject a successful call. next kept additionalProperties: false: staging.test.ts pins it closed and a guard whose reason is not in front of you is not one to loosen for 420 tokens. Ceiling ratcheted to 60 000 rather than the usual ~300 margin, leaving ~1 070 deliberate working margin: server.ts took 70 commits in 14 days and the previous 116-token margin is what starts the ratchet-block-bump-demote cycle visible in the bench log.