Files
accounted/supabase/migrations/20260828120000_company_settings_data_analysis_opt_in.sql
T
ad8566f1ae feat(settings): per-company data-analysis opt-in gating the calibration corpus (#1346) (#2007)
* feat(settings): per-company opt-in for data analysis of bookkeeping outcomes (#1346)

Adds company_settings.data_analysis_opt_in (default false, no grandfathering)
and gates every path that reads bookkeeping outcomes across companies on it:
POST /api/agent/categorize/outcome stops writing calibration samples for
companies that have not opted in, and the backtest / calibration-fit scripts
filter to opted-in company ids. One helper (lib/company/data-analysis.ts)
is the single gate for future analysis paths. A toggle on Inställningar >
Företag states plainly what is analysed (proposed vs booked account, amount,
confidence; no free text, no personal data) in sv and en. The flag is UI-only
by design: consent is a human action, so it is absent from the v1 REST / MCP
settings pick lists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

* fix(settings): make data-analysis consent copy true for the backtest path (#1346)

Addresses adversarial review findings on PR #2007:

- Findings 1-3 (consent narrower than the gated processing): the flag also
  gates scripts/backtest-categorize.ts, which re-runs transaction
  descriptions, merchant names and matched underlag through the model. The
  sv/en toggle help and disclosure now state that explicitly as "evaluation
  runs" and no longer claim that free text or underlag are excluded. The
  migration header and COMMENT, the lib/company/data-analysis.ts docstring,
  the backtest script header and the DECISIONS line say the same. Kept the
  gate (un-gating would put the script back to reading every company with
  no consent at all). A test pins that both locales name those inputs and
  contain no "no free text / no underlag" denial.
- Finding 4 (member sees an active switch that RLS rejects): the toggle is
  now enabled only for owner/admin, matching the company_settings update
  policy; the disclosure says only administrators can change the choice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

* fix(scripts): address round-2 review findings (#1346)

1. [minor] Opted-in company filter was an unbounded PostgREST `in` list in
   the URL (scripts/fit-categorize-calibration.ts, scripts/backtest-categorize.ts).
   Both scripts now read the opted-in ids through a shared, paginated helper
   (listDataAnalysisOptedInCompanyIds, fetchAllRows so the pre-fetch no longer
   caps at 1000) and query per chunk of 100 ids (chunkCompanyIds). The fit
   script pages each chunk on the id PK; the backtest merges per-chunk
   results and re-cuts to the N most recent overall. Early exit on zero
   opt-ins is kept. Pinned with tests in lib/company/__tests__/data-analysis.test.ts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

* fix(scripts): coerce a null transaction description in the backtest (#1346)

The typed row from the chunked consent query made description nullable,
which TransactionForSelect does not accept; fall back to the original
description or an empty string, as the untyped row did implicitly before.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 17:38:36 +02:00

35 lines
2.2 KiB
SQL

-- Per-company opt-in for analysis of bookkeeping data (#1346).
--
-- Adds the consent flag that gates every path where a company's bookkeeping
-- data is read across companies for Accounted's own analysis. Two scopes,
-- both stated in the consent copy (messages/*.json data_analysis.*):
-- 1. Booking outcomes (proposed vs booked account, amount, confidence):
-- the auto-booking calibration corpus (categorize_calibration_samples,
-- written by POST /api/agent/categorize/outcome) and the calibration-fit
-- script.
-- 2. Evaluation runs (scripts/backtest-categorize.ts, founder-run): the
-- company's transaction descriptions, counterparty names and matched
-- underlag (receipts, invoices) are re-run through the same AI model
-- and processors as regular booking. That text can contain personal
-- data (e.g. names in Swish payments), so the disclosure says so.
-- Enforced server-side by lib/company/data-analysis.ts and by the scripts'
-- own opted-in filter, never by the client.
--
-- Default false for everyone, no grandfathering: a company contributes
-- nothing until a company owner or admin flips the toggle in
-- Inställningar > Företag (the company_settings RLS update policy is admin
-- only, so the switch is admin only too). Turning it off stops new collection immediately.
-- The flag does not cover product-usage analytics or MCP reliability
-- telemetry (no bookkeeping content, documented under legitimate interest in
-- .compliance/ropa.yaml). Mirrors mileage_enabled (20260812193500).
--
-- pg-test: skip (plain column addition, no trigger/RPC/RLS)
ALTER TABLE public.company_settings
ADD COLUMN data_analysis_opt_in boolean NOT NULL DEFAULT false;
COMMENT ON COLUMN public.company_settings.data_analysis_opt_in IS
'Consent flag: when true, this company''s booking outcomes (proposed vs booked account, confidence, amount) may be read across companies for the auto-booking calibration corpus, and in evaluation runs (scripts/backtest-categorize.ts) its transaction descriptions, counterparty names and matched underlag may be re-run through the AI model. Default false; enforced server-side (lib/company/data-analysis.ts).';
NOTIFY pgrst, 'reload schema';