Files
accounted/app/api/pending-operations/expire/cron/route.ts
T
Jakob WennbergandClaude Fable 5 25a7261eda fix(pending-ops): recovery sweep for operations stuck in committing (#843) (#1108)
The commit dispatcher claims an op with an atomic pending -> committing
CAS; if the process dies after side-effects post but before the terminal
committed write (or that write fails, the PR #841 log line), the row sat
in status='committing' forever: the expire cron only sweeps 'pending'.

Add lib/pending-operations/recover-stuck-committing.ts, invoked from the
existing daily expire cron (no new vercel.json entry):

- Only rows whose updated_at (the claim timestamp: the CAS bumps it via
  the update_updated_at_column trigger) is older than 15 minutes, well
  past the 300s Vercel function ceiling, so in-flight executors are
  never raced.
- Positive evidence that side-effects posted finalizes the row to
  committed with result_data.recovered=true. Evidence exists only where
  params identify a target with an unambiguous posted state:
  categorize_transaction (is_transaction_booked RPC, skipped for
  allow_duplicate), link_transaction_journal_entry (exact tx+entry
  link), match_transaction_invoice (invoice_payments pair row).
- No evidence: terminal rejected with an explanatory result_data,
  never back to pending (re-execution could duplicate side-effects
  that posted without a trace). Reason 'stuck_committing' is distinct
  from 'expired' so the UI badge never claims these rows.
- Every terminal write is CAS-guarded on status='committing'; probe
  errors skip the row for the next run.
- One structured 'pending_op_recovery' warn per row (count by outcome);
  runbook comment added next to the #841 finalize-failure log line.

Tests: unit coverage for the decision logic and cron wiring (401, sweep
invoked, failure isolation), plus a pg-real test proving row selection,
the trustworthy updated_at anchor, committing -> terminal transitions
through the real immutability/input-frozen triggers, and the
is_transaction_booked evidence substrate.

Fixes #843

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 18:30:33 +02:00

84 lines
3.5 KiB
TypeScript

import { createServiceClient } from '@/lib/supabase/server'
import { NextResponse } from 'next/server'
import { withCronContext } from '@/lib/api/with-cron-context'
import { errorResponse } from '@/lib/errors/get-structured-error'
import { recoverStuckCommittingOperations } from '@/lib/pending-operations/recover-stuck-committing'
/**
* GET /api/pending-operations/expire/cron, daily 02:30 UTC.
*
* Two sweeps per run:
*
* 1. Expiry: auto-rejects staged operations that have sat at status='pending'
* for more than 30 days.
* 2. Recovery (#843): drives rows stuck in status='committing' beyond the
* safe threshold to a terminal status; see
* lib/pending-operations/recover-stuck-committing.ts for the semantics.
*
* AI agents stage operations for human review; when the chat
* session is abandoned the proposal would otherwise linger in the worklist
* forever, asking the user to Godkänn/Avvisa something whose context they no
* longer remember. A 30-day-old proposal has lost its context regardless of
* risk level, so the sweep applies uniformly.
*
* Rows are flipped to 'rejected' (never deleted: the table is the audit
* trail) with the same result_data shape the commit dispatcher uses for its
* own auto-rejects (lib/pending-operations/commit.ts). The strict
* reason: 'expired' marker is what the /pending UI keys its
* "Utgick automatiskt" badge on. rejection_category/rejection_reason stay
* NULL: those carry user feedback semantics, and an expiry is not feedback.
*
* If you change EXPIRY_DAYS, update the user-facing copy that states the
* window: pending.auto_expiry_note + pending.auto_expired_detail in
* messages/{sv,en}.json and the static note in components/agent/ApprovalCard.tsx.
*/
const EXPIRY_DAYS = 30
export const GET = withCronContext('cron.pending_operations_expire', async (_request, ctx) => {
const supabase = createServiceClient()
const cutoff = new Date()
cutoff.setDate(cutoff.getDate() - EXPIRY_DAYS)
// CAS on status='pending': rows a concurrent commit has claimed (status
// 'committing') or already resolved are skipped; the status-immutability
// trigger never fires because OLD.status is always 'pending' here.
// result_data is NULL on pending rows, so plain assignment is the merge.
const { data, error } = await supabase
.from('pending_operations')
.update({
status: 'rejected',
resolved_at: new Date().toISOString(),
result_data: { auto_rejected: true, reason: 'expired' },
})
.eq('status', 'pending')
.lt('created_at', cutoff.toISOString())
.select('id, company_id')
if (error) {
ctx.log.error('pending operations expiry failed', error)
return errorResponse(error, ctx.log, { requestId: ctx.requestId })
}
const expired = data?.length ?? 0
ctx.log.info('pending operations expiry summary', {
expired,
cutoff: cutoff.toISOString(),
})
// Recovery sweep for rows stuck in 'committing' (#843). Isolated so a
// recovery failure never masks a successful expiry pass: the expiry update
// above has already been applied at this point.
let recovery: Record<string, unknown>
try {
recovery = { ...(await recoverStuckCommittingOperations(supabase, { log: ctx.log })) }
} catch (err) {
// Detail goes to the structured log only; the response carries a static
// marker so an operator sees the run partially failed.
ctx.log.error('pending_op_recovery sweep failed', err as Error)
recovery = { error: 'recovery_sweep_failed' }
}
return NextResponse.json({ success: true, expired, cutoff: cutoff.toISOString(), recovery })
})