feat(agent): telemetry + CI-gate quick wins from the "AI systems that ship" audit (#677)

* feat(agent): telemetry completeness + durability, CI gates, commit_method provenance

Quick wins from the "Building AI systems that ship" audit:

- mcp.tool_called gains errorMessage (message_sv, truncated 500 chars) on
  all failure exits; new mcp.skill_loaded event on every gnubok_load_skill
  (all tiers) so atom usage is finally measurable
- event_log: (event_type, created_at) index; cleanup cron keeps
  mcp.*/agent.* telemetry 180 days (delivery events stay 30)
- CI: lint ratchet (npm run check:lint — 60 legacy errors baselined,
  fails only on NEW errors) and a pg-real coverage gate (migrations
  touching trigger/RPC/RLS/DEFERRABLE require a *.pg.test.ts change;
  escape hatch: -- pg-test: covered-by/skip)
- journal_entries.commit_method CHECK widened with 'api_key'/'agent';
  the MCP approve path records 'api_key' truthfully instead of
  'user_accept' (agent_first_vision §8 P0-1). 'agent' is reserved — ALL
  MCP traffic (incl. claude.ai OAuth, whose access_token is a minted
  API key) authenticates as api_key today

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(import): derive opening balances from prior-year #UB when SIE lacks #IB (#675)

SIE files exported without #IB 0 rows (only #UB -1) previously imported
with zero opening balances. getEffectiveOpeningBalances() now derives IB
from prior-year UB for balance-sheet accounts when explicit #IB is
absent, surfaces the derivation as an info issue in the import preview,
and excludes share-capital vouchers from opening-balance detection.
Detection regexes are shared between parser and importer so the two
checks cannot drift. 507 lib/import tests pass.

(Authored in a parallel session in this checkout; included per request.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): address PR #677 bot findings — RoPA entry, execFileSync, gate scope note

Triage of the compliance-swarm + Greptile findings:

Applied:
- .compliance/ropa.yaml: new mcp.telemetry processing activity declaring
  the 180-day mcp.*/agent.* retention, lawful basis, data categories, and
  the no-args/no-results minimisation (ISO A.8.10, GDPR Art.5(1)(c) —
  the retention split is now formally documented, referenced from the cron)
- check-pg-test-coverage.mjs: execFileSync with argv array — no shell, so
  a hostile base-ref can't inject (ASVS V13.2.1); verified an injection
  attempt exits 2 without executing
- check-pg-test-coverage.mjs: documented the PR-level (not per-migration)
  scope of the gate so reviewers know to check coverage per migration when
  a PR carries several risky migrations (Greptile P2)

Acknowledged, no change:
- errorMessage PII risk: messages are domain-mapped strings; event_log
  already persists far richer delivery payloads under the same RLS; now
  declared in ropa.yaml
- cron error envelope: errorResponse maps to the canonical safe envelope
  and the endpoint is CRON_SECRET-gated
- two-pass delete "partial state": TTL deletes are idempotent — the next
  daily run sweeps whatever a failed pass left behind
- skill_loaded actorLabel/sessionId: mirrors the pre-existing
  mcp.tool_called payload; sessionId is the join key the analytics exist for

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jakob Wennberg
2026-06-05 15:47:13 +02:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 076bb169f8
commit bc61862e76
24 changed files with 1733 additions and 62 deletions
@@ -0,0 +1,108 @@
/**
* Tests for the event_log cleanup cron's differentiated retention:
* delivery events at 30 days, agent telemetry (mcp.*, agent.*) at 180 days.
*/
import { describe, it, expect, vi, beforeEach } from 'vitest'
vi.mock('@/lib/auth/cron', () => ({
verifyCronSecret: vi.fn(() => null),
}))
interface FilterCall {
method: string
args: unknown[]
}
interface DeleteCapture {
filters: FilterCall[]
}
const deleteCalls: DeleteCapture[] = []
let deleteResults: Array<{ error: unknown; count: number | null }> = []
vi.mock('@/lib/supabase/server', () => ({
createServiceClient: vi.fn(() => ({
from: vi.fn(() => {
const capture: DeleteCapture = { filters: [] }
deleteCalls.push(capture)
const result = deleteResults.shift() ?? { error: null, count: 0 }
const chain: Record<string, unknown> = {}
chain.delete = vi.fn(() => chain)
chain.lt = vi.fn((...args: unknown[]) => {
capture.filters.push({ method: 'lt', args })
return chain
})
chain.not = vi.fn((...args: unknown[]) => {
capture.filters.push({ method: 'not', args })
return chain
})
// Thenable — awaiting the builder resolves the queued result.
chain.then = (resolve: (v: unknown) => unknown) => Promise.resolve(result).then(resolve)
return chain
}),
})),
}))
import { GET } from '../route'
function cronRequest(): Request {
return new Request('http://localhost:3000/api/events/cleanup/cron')
}
function daysAgo(iso: string): number {
return (Date.now() - new Date(iso).getTime()) / 86_400_000
}
beforeEach(() => {
vi.clearAllMocks()
deleteCalls.length = 0
deleteResults = []
})
describe('GET /api/events/cleanup/cron', () => {
it('runs two delete passes: delivery at 30 days (telemetry excluded), everything at 180', async () => {
deleteResults = [
{ error: null, count: 12 },
{ error: null, count: 3 },
]
const response = await GET(cronRequest())
const json = await response.json()
expect(json).toEqual({
success: true,
deleted: 15,
deletedDelivery: 12,
deletedTelemetry: 3,
})
expect(deleteCalls).toHaveLength(2)
// Pass 1: 30-day cutoff + telemetry exclusion filters.
const pass1 = deleteCalls[0]
const lt1 = pass1.filters.find((f) => f.method === 'lt')!
expect(lt1.args[0]).toBe('created_at')
expect(daysAgo(lt1.args[1] as string)).toBeCloseTo(30, 0)
const notFilters = pass1.filters.filter((f) => f.method === 'not')
expect(notFilters.map((f) => f.args)).toEqual([
['event_type', 'like', 'mcp.%'],
['event_type', 'like', 'agent.%'],
])
// Pass 2: 180-day cutoff, no exclusions — sweeps the telemetry rows.
const pass2 = deleteCalls[1]
const lt2 = pass2.filters.find((f) => f.method === 'lt')!
expect(daysAgo(lt2.args[1] as string)).toBeCloseTo(180, 0)
expect(pass2.filters.filter((f) => f.method === 'not')).toHaveLength(0)
})
it('short-circuits with an error envelope when the delivery pass fails', async () => {
deleteResults = [{ error: { message: 'boom', code: 'XX000' }, count: null }]
const response = await GET(cronRequest())
expect(response.status).toBeGreaterThanOrEqual(500)
// The 180-day pass never ran.
expect(deleteCalls).toHaveLength(1)
})
})
+48 -11
View File
@@ -5,26 +5,63 @@ import { errorResponse } from '@/lib/errors/get-structured-error'
/**
* GET /api/events/cleanup/cron — daily 02:00 UTC.
* Removes event_log rows older than 30 days.
*
* Differentiated retention:
* - Delivery events (invoice.created, transaction.synced, …): 30 days. They
* exist for external automation polling (n8n/Make/Zapier) and go stale fast.
* - Agent telemetry (mcp.*, agent.*): 180 days. Error-rate trends and
* skill-load correlation need more than one month of signal — a 30-day
* window made it impossible to tell whether a tool or skill change actually
* moved failure rates.
*
* Retention is declared in .compliance/ropa.yaml (id: mcp.telemetry).
*/
const DELIVERY_RETENTION_DAYS = 30
const TELEMETRY_RETENTION_DAYS = 180
export const GET = withCronContext('cron.events_cleanup', async (_request, ctx) => {
const supabase = createServiceClient()
const cutoff = new Date()
cutoff.setDate(cutoff.getDate() - 30)
const deliveryCutoff = new Date()
deliveryCutoff.setDate(deliveryCutoff.getDate() - DELIVERY_RETENTION_DAYS)
const telemetryCutoff = new Date()
telemetryCutoff.setDate(telemetryCutoff.getDate() - TELEMETRY_RETENTION_DAYS)
const { error, count } = await supabase
// Pass 1: delivery events past 30 days. Telemetry (mcp.*, agent.*) is
// excluded here and swept by the 180-day pass below.
const { error: deliveryError, count: deliveryCount } = await supabase
.from('event_log')
.delete({ count: 'exact' })
.lt('created_at', cutoff.toISOString())
.lt('created_at', deliveryCutoff.toISOString())
.not('event_type', 'like', 'mcp.%')
.not('event_type', 'like', 'agent.%')
if (error) {
ctx.log.error('event log cleanup failed', error)
return errorResponse(error, ctx.log, { requestId: ctx.requestId })
if (deliveryError) {
ctx.log.error('event log delivery cleanup failed', deliveryError)
return errorResponse(deliveryError, ctx.log, { requestId: ctx.requestId })
}
const deleted = count ?? 0
ctx.log.info('event log cleanup summary', { deleted, cutoff: cutoff.toISOString() })
// Pass 2: everything past 180 days — catches the telemetry rows pass 1 skipped.
const { error: telemetryError, count: telemetryCount } = await supabase
.from('event_log')
.delete({ count: 'exact' })
.lt('created_at', telemetryCutoff.toISOString())
return NextResponse.json({ success: true, deleted })
if (telemetryError) {
ctx.log.error('event log telemetry cleanup failed', telemetryError)
return errorResponse(telemetryError, ctx.log, { requestId: ctx.requestId })
}
const deletedDelivery = deliveryCount ?? 0
const deletedTelemetry = telemetryCount ?? 0
const deleted = deletedDelivery + deletedTelemetry
ctx.log.info('event log cleanup summary', {
deleted,
deletedDelivery,
deletedTelemetry,
deliveryCutoff: deliveryCutoff.toISOString(),
telemetryCutoff: telemetryCutoff.toISOString(),
})
return NextResponse.json({ success: true, deleted, deletedDelivery, deletedTelemetry })
})