Files
accounted/extensions/general/invoice-inbox/lib/sweep.ts
T
MattssonandClaude Fable 5 a1cafe495f fix(invoice-inbox): read whole PDFs (last-page slice + truncation retry) (#2014)
* fix(invoice-inbox): read whole PDFs (last-page slice + truncation retry)

PDF extraction read only part of well-structured PDFs, two confirmed
mechanisms (21-day prod window: 49 sliced docs, 29 silent empties):

- The auto-extract page budget was 3 (Bedrock-latency legacy, issue #553)
  and the slice kept only the first pages, so multi-page invoices lost the
  final page where totals, OCR and 'Att betala' sit. The budget is now 8 on
  pdf-native backends (Claude reads PDFs directly); the slice always keeps
  the last page. Rasterizing self-host backends keep the old budget of 3.
- A max_tokens-truncated model answer was parsed as-is, failed, and became
  an all-null extraction with no trace. extractFromDocument now reports
  stop_reason max_tokens / finish_reason length as truncated; the extractor
  retries once at double AI_EXTRACTION_MAX_TOKENS and logs
  ai_extraction_truncated either way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xosyW53HUa9JoFiDayhSk

* fix(invoice-inbox): sweep cutoff covers the slower two-call extraction

Skeptic finding on #2014: the crash-recovery sweep flipped 'processing'
rows to an empty skeleton after 2 minutes, but a deferred extraction can
now legitimately run 3-5 minutes (8 native pages plus one truncation
retry at a doubled token cap), so the sweep stole the row and the CAS
discarded the worker's real result. Cutoff raised to 10 minutes.

Also: pages_partial_note made period-agnostic (old rows were extracted
from first-pages-only slices, so naming the last page was retroactively
wrong for them), and two stale first-pages-only comments updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xosyW53HUa9JoFiDayhSk

* fix(invoice-inbox): keep the first extraction response when the retry throws

CodeRabbit finding on #2014: a throttled/failed retry call bubbled to the
outer catch before rawText was assigned, discarding a first response whose
text may parse fine despite the truncation flag. The retry is now caught
locally (logged as ai_extraction_retry_failed) and the first result flows on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012xosyW53HUa9JoFiDayhSk

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 18:57:02 +02:00

84 lines
3.3 KiB
TypeScript

/**
* Crash recovery for the staged upload.
*
* The web upload route inserts the inbox row as 'processing' and defers
* Bedrock extraction to an after() worker that can die with the serverless
* instance. A row stuck in 'processing' is durable state with no live owner:
* this sweep flips it to 'received' with the empty extraction skeleton so it
* becomes a normal manually-editable item. Deliberately NO re-extraction
* here: the UI retry button covers that, and a cron that silently re-spends
* Bedrock tokens on every crash would hide the crashes.
*
* Overlap with a slow live worker is safe: every mutation is a guarded claim
* on status='processing' (and extracted_data still NULL), so the sweep and
* the worker never both win one row.
*/
import type { SupabaseClient } from '@supabase/supabase-js'
import { createLogger } from '@/lib/logger'
import { emptyResult } from './extract-invoice-fields'
const log = createLogger('invoice-inbox/sweep')
/**
* A deferred extraction is now up to TWO Bedrock calls on up to 8 native PDF
* pages: the base call can spend its full output budget (~90-150s at 8192
* tokens) before the truncation retry doubles the cap and runs again. A live
* worker can therefore be legitimately silent for several minutes; a cutoff
* shorter than its worst case makes the sweep steal the row and discard the
* worker's real result via the status CAS. Ten minutes clears the two-call
* worst case with margin while still flipping genuinely crashed rows well
* before anyone files a support mail about a stuck item.
*/
export const PROCESSING_STUCK_MS = 10 * 60 * 1000
const BATCH = 50
export interface InboxSweepSummary {
/** Stale 'processing' rows flipped to 'received' with the empty skeleton. */
flipped: number
}
/** Run one sweep pass. Never throws. */
export async function runInboxSweep(supabase: SupabaseClient): Promise<InboxSweepSummary> {
const cutoff = new Date(Date.now() - PROCESSING_STUCK_MS).toISOString()
const { data: stale, error: selectError } = await supabase
.from('invoice_inbox_items')
.select('id')
.eq('status', 'processing')
.lt('created_at', cutoff)
.limit(BATCH)
if (selectError) {
log.error('stale-processing select failed', { error: selectError.message })
return { flipped: 0 }
}
const ids = ((stale ?? []) as Array<{ id: string }>).map((r) => r.id)
if (ids.length === 0) return { flipped: 0 }
// CAS: the status guard keeps a just-finished worker's real result, and
// the extracted_data-still-NULL guard keeps any fields a caller PUT onto
// the row in the meantime; a row that fails either guard is someone
// else's win, not ours.
const { data: claimed, error: updateError } = await supabase
.from('invoice_inbox_items')
.update({
status: 'received',
extracted_data: emptyResult() as unknown as Record<string, unknown>,
extraction_skipped: false,
})
.in('id', ids)
.eq('status', 'processing')
.is('extracted_data', null)
.select('id')
if (updateError) {
log.error('stale-processing flip failed', { error: updateError.message })
return { flipped: 0 }
}
const flipped = Array.isArray(claimed) ? claimed.length : 0
if (flipped > 0) {
log.info('flipped stale processing rows to received', { flipped })
}
return { flipped }
}