feat(settings): per-company data-analysis opt-in gating the calibration corpus (#1346) (#2007)

* feat(settings): per-company opt-in for data analysis of bookkeeping outcomes (#1346)

Adds company_settings.data_analysis_opt_in (default false, no grandfathering)
and gates every path that reads bookkeeping outcomes across companies on it:
POST /api/agent/categorize/outcome stops writing calibration samples for
companies that have not opted in, and the backtest / calibration-fit scripts
filter to opted-in company ids. One helper (lib/company/data-analysis.ts)
is the single gate for future analysis paths. A toggle on Inställningar >
Företag states plainly what is analysed (proposed vs booked account, amount,
confidence; no free text, no personal data) in sv and en. The flag is UI-only
by design: consent is a human action, so it is absent from the v1 REST / MCP
settings pick lists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

* fix(settings): make data-analysis consent copy true for the backtest path (#1346)

Addresses adversarial review findings on PR #2007:

- Findings 1-3 (consent narrower than the gated processing): the flag also
  gates scripts/backtest-categorize.ts, which re-runs transaction
  descriptions, merchant names and matched underlag through the model. The
  sv/en toggle help and disclosure now state that explicitly as "evaluation
  runs" and no longer claim that free text or underlag are excluded. The
  migration header and COMMENT, the lib/company/data-analysis.ts docstring,
  the backtest script header and the DECISIONS line say the same. Kept the
  gate (un-gating would put the script back to reading every company with
  no consent at all). A test pins that both locales name those inputs and
  contain no "no free text / no underlag" denial.
- Finding 4 (member sees an active switch that RLS rejects): the toggle is
  now enabled only for owner/admin, matching the company_settings update
  policy; the disclosure says only administrators can change the choice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

* fix(scripts): address round-2 review findings (#1346)

1. [minor] Opted-in company filter was an unbounded PostgREST `in` list in
   the URL (scripts/fit-categorize-calibration.ts, scripts/backtest-categorize.ts).
   Both scripts now read the opted-in ids through a shared, paginated helper
   (listDataAnalysisOptedInCompanyIds, fetchAllRows so the pre-fetch no longer
   caps at 1000) and query per chunk of 100 ids (chunkCompanyIds). The fit
   script pages each chunk on the id PK; the backtest merges per-chunk
   results and re-cuts to the N most recent overall. Early exit on zero
   opt-ins is kept. Pinned with tests in lib/company/__tests__/data-analysis.test.ts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

* fix(scripts): coerce a null transaction description in the backtest (#1346)

The typed row from the chunked consent query made description nullable,
which TransactionForSelect does not accept; fall back to the original
description or an empty string, as the untyped row did implicitly before.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Jakob Wennberg
2026-08-28 17:38:36 +02:00
committed by GitHub
co-authored by Claude Fable 5 Jakob Wennberg
parent 33a58bec51
commit ad8566f1ae
16 changed files with 500 additions and 28 deletions
+52 -13
View File
@@ -13,6 +13,15 @@
* npx tsx scripts/backtest-categorize.ts [N]
* rm .env.local
*
* Consent: only companies with company_settings.data_analysis_opt_in = true
* are read (#1346). This script goes beyond booking outcomes: it reads each
* transaction's description, merchant name and matched underlag (via
* gatherUnderlag) and sends them to the model again, so the consent copy in
* messages/*.json (data_analysis.settings_toggle_help) explicitly names
* "evaluation runs" with exactly those inputs. Do not add inputs here that
* the copy does not name. Nobody is opted in by default, so an empty run is
* the expected state until an admin flips the toggle in Inställningar > Företag.
*
* Leakage caveat: a known vendor's counterparty template may already reflect
* the very booking under test, inflating the "deterministic nailed it" segment.
* The "model had to decide" segment below is the leakage-free measure.
@@ -29,23 +38,53 @@ async function main() {
const { gatherCandidates } = await import('../lib/agent/categorize/candidates')
const { gatherUnderlag } = await import('../lib/agent/categorize/underlag')
const { selectAccount } = await import('../lib/agent/categorize/select-account')
const { chunkCompanyIds, listDataAnalysisOptedInCompanyIds } = await import('../lib/company/data-analysis')
const url = process.env.NEXT_PUBLIC_SUPABASE_URL!
const key = process.env.SUPABASE_SERVICE_ROLE_KEY!
const supabase = createClient(url, key)
// Recent booked expense transactions with a counterparty.
const { data: txs, error } = await supabase
.from('transactions')
.select('id, company_id, merchant_name, description, original_description, amount, date, currency, document_id, journal_entry_id')
.not('journal_entry_id', 'is', null)
.lt('amount', 0)
.eq('is_business', true)
.not('merchant_name', 'is', null)
.order('created_at', { ascending: false })
.limit(N)
if (error) throw error
const rows = txs ?? []
// Consent gate (#1346): only companies that opted in to data analysis.
const optedInIds = await listDataAnalysisOptedInCompanyIds(supabase)
if (optedInIds.length === 0) {
console.log('\nNo company has opted in to data analysis (company_settings.data_analysis_opt_in). Nothing to backtest.')
return
}
// Recent booked expense transactions with a counterparty. Queried per chunk
// of company ids (`.in()` lives in the GET query string), then merged and
// re-cut to the N most recent overall.
type Tx = {
id: string
company_id: string
merchant_name: string | null
description: string | null
original_description: string | null
amount: number
date: string
currency: string | null
document_id: string | null
journal_entry_id: string | null
created_at: string
}
const candidatesByChunk: Tx[] = []
for (const chunk of chunkCompanyIds(optedInIds)) {
const { data: txs, error } = await supabase
.from('transactions')
.select('id, company_id, merchant_name, description, original_description, amount, date, currency, document_id, journal_entry_id, created_at')
.in('company_id', chunk)
.not('journal_entry_id', 'is', null)
.lt('amount', 0)
.eq('is_business', true)
.not('merchant_name', 'is', null)
.order('created_at', { ascending: false })
.limit(N)
if (error) throw error
candidatesByChunk.push(...((txs ?? []) as Tx[]))
}
const rows = candidatesByChunk
.sort((a, b) => (a.created_at < b.created_at ? 1 : a.created_at > b.created_at ? -1 : 0))
.slice(0, N)
console.log(`\nBacktesting ${rows.length} booked transactions on ${process.env.BEDROCK_MODEL_ID ?? process.env.AI_MODEL ?? 'the configured model'}…\n`)
// Ground-truth debit account per journal entry (expense line, not cash/VAT).
@@ -101,7 +140,7 @@ async function main() {
const sel = await selectAccount({
transaction: {
merchantName: r.merchant_name,
description: r.description,
description: r.description ?? r.original_description ?? '',
amount: r.amount,
date: r.date,
currency: r.currency,
+28 -10
View File
@@ -11,6 +11,11 @@
*
* Note: .env.local points at production; this only SELECTs, so it is safe, but
* it is still the prod corpus you are reading.
*
* Consent: samples are only written for, and only read from, companies with
* company_settings.data_analysis_opt_in = true (#1346). The write side is
* gated in POST /api/agent/categorize/outcome; the read side filters again
* here so a company that opted out after contributing drops out of the fit.
*/
import { createClient } from '@supabase/supabase-js'
import {
@@ -21,6 +26,7 @@ import {
bandFor,
type Sample,
} from '@/lib/agent/categorize/calibration'
import { chunkCompanyIds, listDataAnalysisOptedInCompanyIds } from '@/lib/company/data-analysis'
const url = process.env.NEXT_PUBLIC_SUPABASE_URL
const key = process.env.SUPABASE_SERVICE_ROLE_KEY
@@ -31,18 +37,30 @@ if (!url || !key) {
const supabase = createClient(url, key)
async function main() {
// Consent gate (#1346): only companies that opted in to data analysis.
const optedInIds = await listDataAnalysisOptedInCompanyIds(supabase)
if (optedInIds.length === 0) {
console.log('\nNo company has opted in to data analysis (company_settings.data_analysis_opt_in). Nothing to fit.')
return
}
// Query per chunk of company ids: `.in()` goes into the GET query string, so
// one request per few hundred opted-in companies would hit URL limits.
const rows: { confidence: number; was_correct: boolean }[] = []
const PAGE = 1000
for (let from = 0; ; from += PAGE) {
const { data, error } = await supabase
.from('categorize_calibration_samples')
.select('confidence, was_correct')
.order('created_at', { ascending: false })
.range(from, from + PAGE - 1)
if (error) throw error
if (!data || data.length === 0) break
rows.push(...(data as { confidence: number; was_correct: boolean }[]))
if (data.length < PAGE) break
for (const chunk of chunkCompanyIds(optedInIds)) {
for (let from = 0; ; from += PAGE) {
const { data, error } = await supabase
.from('categorize_calibration_samples')
.select('confidence, was_correct')
.in('company_id', chunk)
.order('id', { ascending: true })
.range(from, from + PAGE - 1)
if (error) throw error
if (!data || data.length === 0) break
rows.push(...(data as { confidence: number; was_correct: boolean }[]))
if (data.length < PAGE) break
}
}
const samples: Sample[] = rows.map((r) => ({ confidence: Number(r.confidence), correct: r.was_correct }))