* feat(settings): per-company opt-in for data analysis of bookkeeping outcomes (#1346) Adds company_settings.data_analysis_opt_in (default false, no grandfathering) and gates every path that reads bookkeeping outcomes across companies on it: POST /api/agent/categorize/outcome stops writing calibration samples for companies that have not opted in, and the backtest / calibration-fit scripts filter to opted-in company ids. One helper (lib/company/data-analysis.ts) is the single gate for future analysis paths. A toggle on Inställningar > Företag states plainly what is analysed (proposed vs booked account, amount, confidence; no free text, no personal data) in sv and en. The flag is UI-only by design: consent is a human action, so it is absent from the v1 REST / MCP settings pick lists. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna * fix(settings): make data-analysis consent copy true for the backtest path (#1346) Addresses adversarial review findings on PR #2007: - Findings 1-3 (consent narrower than the gated processing): the flag also gates scripts/backtest-categorize.ts, which re-runs transaction descriptions, merchant names and matched underlag through the model. The sv/en toggle help and disclosure now state that explicitly as "evaluation runs" and no longer claim that free text or underlag are excluded. The migration header and COMMENT, the lib/company/data-analysis.ts docstring, the backtest script header and the DECISIONS line say the same. Kept the gate (un-gating would put the script back to reading every company with no consent at all). A test pins that both locales name those inputs and contain no "no free text / no underlag" denial. - Finding 4 (member sees an active switch that RLS rejects): the toggle is now enabled only for owner/admin, matching the company_settings update policy; the disclosure says only administrators can change the choice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna * fix(scripts): address round-2 review findings (#1346) 1. [minor] Opted-in company filter was an unbounded PostgREST `in` list in the URL (scripts/fit-categorize-calibration.ts, scripts/backtest-categorize.ts). Both scripts now read the opted-in ids through a shared, paginated helper (listDataAnalysisOptedInCompanyIds, fetchAllRows so the pre-fetch no longer caps at 1000) and query per chunk of 100 ids (chunkCompanyIds). The fit script pages each chunk on the id PK; the backtest merges per-chunk results and re-cuts to the N most recent overall. Early exit on zero opt-ins is kept. Pinned with tests in lib/company/__tests__/data-analysis.test.ts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna * fix(scripts): coerce a null transaction description in the backtest (#1346) The typed row from the chunked consent query made description nullable, which TransactionForSelect does not accept; fall back to the original description or an empty string, as the untyped row did implicitly before. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015nAd8XJ2RPCmG2eKoLBdna --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
Jakob Wennberg
parent
33a58bec51
commit
ad8566f1ae
@@ -13,6 +13,15 @@
|
||||
* npx tsx scripts/backtest-categorize.ts [N]
|
||||
* rm .env.local
|
||||
*
|
||||
* Consent: only companies with company_settings.data_analysis_opt_in = true
|
||||
* are read (#1346). This script goes beyond booking outcomes: it reads each
|
||||
* transaction's description, merchant name and matched underlag (via
|
||||
* gatherUnderlag) and sends them to the model again, so the consent copy in
|
||||
* messages/*.json (data_analysis.settings_toggle_help) explicitly names
|
||||
* "evaluation runs" with exactly those inputs. Do not add inputs here that
|
||||
* the copy does not name. Nobody is opted in by default, so an empty run is
|
||||
* the expected state until an admin flips the toggle in Inställningar > Företag.
|
||||
*
|
||||
* Leakage caveat: a known vendor's counterparty template may already reflect
|
||||
* the very booking under test, inflating the "deterministic nailed it" segment.
|
||||
* The "model had to decide" segment below is the leakage-free measure.
|
||||
@@ -29,23 +38,53 @@ async function main() {
|
||||
const { gatherCandidates } = await import('../lib/agent/categorize/candidates')
|
||||
const { gatherUnderlag } = await import('../lib/agent/categorize/underlag')
|
||||
const { selectAccount } = await import('../lib/agent/categorize/select-account')
|
||||
const { chunkCompanyIds, listDataAnalysisOptedInCompanyIds } = await import('../lib/company/data-analysis')
|
||||
|
||||
const url = process.env.NEXT_PUBLIC_SUPABASE_URL!
|
||||
const key = process.env.SUPABASE_SERVICE_ROLE_KEY!
|
||||
const supabase = createClient(url, key)
|
||||
|
||||
// Recent booked expense transactions with a counterparty.
|
||||
const { data: txs, error } = await supabase
|
||||
.from('transactions')
|
||||
.select('id, company_id, merchant_name, description, original_description, amount, date, currency, document_id, journal_entry_id')
|
||||
.not('journal_entry_id', 'is', null)
|
||||
.lt('amount', 0)
|
||||
.eq('is_business', true)
|
||||
.not('merchant_name', 'is', null)
|
||||
.order('created_at', { ascending: false })
|
||||
.limit(N)
|
||||
if (error) throw error
|
||||
const rows = txs ?? []
|
||||
// Consent gate (#1346): only companies that opted in to data analysis.
|
||||
const optedInIds = await listDataAnalysisOptedInCompanyIds(supabase)
|
||||
if (optedInIds.length === 0) {
|
||||
console.log('\nNo company has opted in to data analysis (company_settings.data_analysis_opt_in). Nothing to backtest.')
|
||||
return
|
||||
}
|
||||
|
||||
// Recent booked expense transactions with a counterparty. Queried per chunk
|
||||
// of company ids (`.in()` lives in the GET query string), then merged and
|
||||
// re-cut to the N most recent overall.
|
||||
type Tx = {
|
||||
id: string
|
||||
company_id: string
|
||||
merchant_name: string | null
|
||||
description: string | null
|
||||
original_description: string | null
|
||||
amount: number
|
||||
date: string
|
||||
currency: string | null
|
||||
document_id: string | null
|
||||
journal_entry_id: string | null
|
||||
created_at: string
|
||||
}
|
||||
const candidatesByChunk: Tx[] = []
|
||||
for (const chunk of chunkCompanyIds(optedInIds)) {
|
||||
const { data: txs, error } = await supabase
|
||||
.from('transactions')
|
||||
.select('id, company_id, merchant_name, description, original_description, amount, date, currency, document_id, journal_entry_id, created_at')
|
||||
.in('company_id', chunk)
|
||||
.not('journal_entry_id', 'is', null)
|
||||
.lt('amount', 0)
|
||||
.eq('is_business', true)
|
||||
.not('merchant_name', 'is', null)
|
||||
.order('created_at', { ascending: false })
|
||||
.limit(N)
|
||||
if (error) throw error
|
||||
candidatesByChunk.push(...((txs ?? []) as Tx[]))
|
||||
}
|
||||
const rows = candidatesByChunk
|
||||
.sort((a, b) => (a.created_at < b.created_at ? 1 : a.created_at > b.created_at ? -1 : 0))
|
||||
.slice(0, N)
|
||||
console.log(`\nBacktesting ${rows.length} booked transactions on ${process.env.BEDROCK_MODEL_ID ?? process.env.AI_MODEL ?? 'the configured model'}…\n`)
|
||||
|
||||
// Ground-truth debit account per journal entry (expense line, not cash/VAT).
|
||||
@@ -101,7 +140,7 @@ async function main() {
|
||||
const sel = await selectAccount({
|
||||
transaction: {
|
||||
merchantName: r.merchant_name,
|
||||
description: r.description,
|
||||
description: r.description ?? r.original_description ?? '',
|
||||
amount: r.amount,
|
||||
date: r.date,
|
||||
currency: r.currency,
|
||||
|
||||
@@ -11,6 +11,11 @@
|
||||
*
|
||||
* Note: .env.local points at production; this only SELECTs, so it is safe, but
|
||||
* it is still the prod corpus you are reading.
|
||||
*
|
||||
* Consent: samples are only written for, and only read from, companies with
|
||||
* company_settings.data_analysis_opt_in = true (#1346). The write side is
|
||||
* gated in POST /api/agent/categorize/outcome; the read side filters again
|
||||
* here so a company that opted out after contributing drops out of the fit.
|
||||
*/
|
||||
import { createClient } from '@supabase/supabase-js'
|
||||
import {
|
||||
@@ -21,6 +26,7 @@ import {
|
||||
bandFor,
|
||||
type Sample,
|
||||
} from '@/lib/agent/categorize/calibration'
|
||||
import { chunkCompanyIds, listDataAnalysisOptedInCompanyIds } from '@/lib/company/data-analysis'
|
||||
|
||||
const url = process.env.NEXT_PUBLIC_SUPABASE_URL
|
||||
const key = process.env.SUPABASE_SERVICE_ROLE_KEY
|
||||
@@ -31,18 +37,30 @@ if (!url || !key) {
|
||||
const supabase = createClient(url, key)
|
||||
|
||||
async function main() {
|
||||
// Consent gate (#1346): only companies that opted in to data analysis.
|
||||
const optedInIds = await listDataAnalysisOptedInCompanyIds(supabase)
|
||||
if (optedInIds.length === 0) {
|
||||
console.log('\nNo company has opted in to data analysis (company_settings.data_analysis_opt_in). Nothing to fit.')
|
||||
return
|
||||
}
|
||||
|
||||
// Query per chunk of company ids: `.in()` goes into the GET query string, so
|
||||
// one request per few hundred opted-in companies would hit URL limits.
|
||||
const rows: { confidence: number; was_correct: boolean }[] = []
|
||||
const PAGE = 1000
|
||||
for (let from = 0; ; from += PAGE) {
|
||||
const { data, error } = await supabase
|
||||
.from('categorize_calibration_samples')
|
||||
.select('confidence, was_correct')
|
||||
.order('created_at', { ascending: false })
|
||||
.range(from, from + PAGE - 1)
|
||||
if (error) throw error
|
||||
if (!data || data.length === 0) break
|
||||
rows.push(...(data as { confidence: number; was_correct: boolean }[]))
|
||||
if (data.length < PAGE) break
|
||||
for (const chunk of chunkCompanyIds(optedInIds)) {
|
||||
for (let from = 0; ; from += PAGE) {
|
||||
const { data, error } = await supabase
|
||||
.from('categorize_calibration_samples')
|
||||
.select('confidence, was_correct')
|
||||
.in('company_id', chunk)
|
||||
.order('id', { ascending: true })
|
||||
.range(from, from + PAGE - 1)
|
||||
if (error) throw error
|
||||
if (!data || data.length === 0) break
|
||||
rows.push(...(data as { confidence: number; was_correct: boolean }[]))
|
||||
if (data.length < PAGE) break
|
||||
}
|
||||
}
|
||||
|
||||
const samples: Sample[] = rows.map((r) => ({ confidence: Number(r.confidence), correct: r.was_correct }))
|
||||
|
||||
Reference in New Issue
Block a user