fix(invoice-inbox): parse AI extraction output wrapped in markdown fences (#1460)

* fix(invoice-inbox): parse AI extraction output wrapped in markdown fences

Since the Sonnet 5 switch (2d543ac99, 2026-07-27) the model intermittently
wraps its JSON answer in ```json fences or adds a short preamble despite
the JSON-only system-prompt rule. JSON.parse(rawText) then threw, the
catch swallowed the error into emptyResult(), and the user got a blank
extraction form: 10-20% of prod receipt extractions since July 28 landed
empty (confidence 0) while the Bedrock call was still paid for.

Slice the raw response from the first '{' to the last '}' before parsing.
Fenced, prefixed, and suffixed outputs now parse; brace-less prose
refusals fall through to the existing empty-result path unchanged. Covers
both pipelines (invoice-inbox upload/email/whatsapp and the
document-extraction extension) since they share extractInvoiceFields.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoice-inbox): depth-aware JSON extraction instead of naive brace slice

Review findings (CodeRabbit, PR Agent, compliance swarm) converged on the
same edge case: first-'{'/last-'}' slicing picks a wrong span when the
model's surrounding prose itself contains braces. Replace it with a
string- and escape-aware balanced scan that returns the first candidate
JSON.parse accepts; prose-only responses still fall through unchanged to
the empty-result path. Two regression tests: braces in surrounding prose,
braces inside JSON string values.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(invoice-inbox): bound the JSON candidate scan against pathological input

Compliance swarm round 2 (A.8.28/A.8.29, non-blocking): the balanced-brace
scan restarted from every '{' with no bound, worst-case quadratic on
adversarially brace-laden text. Cap input length at 256 KB and candidate
attempts at 50; real model output is capped by MAX_TOKENS at roughly 33 KB
so genuine responses never come near either bound. Exhausted or oversized
input falls through unchanged to the existing empty-result path. Two tests:
100k-brace pathological input completes fast and lands empty, oversized
input skips scanning entirely.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Mattsson
2026-08-08 13:41:27 +02:00
committed by GitHub
co-authored by Claude Fable 5
parent 70845edf69
commit 4b0a185876
3 changed files with 175 additions and 2 deletions
+2
View File
@@ -831,3 +831,5 @@ One line per decision: `[YYYY-MM-DD] <decision>: <why>`. Appended by agents and
[2026-07-28] Transaction method (structured payment rail): the trailing channel phrase ("Överföring via internet", "Kortköp/uttag") is stripped from transactions.description at INGEST and by a one-shot BACKFILL, not merely hidden at render: description is the mutable working title, original_description keeps the full bank string, and every dedup surface (external_id: date+öre only; content bridge: prefix-containment over original_description ?? description, and a trailing strip leaves a prefix) is provably unaffected. transaction_method is text + CHECK (repo convention, no PG enums) beside verbatim bank_transaction_code / proprietary_bank_transaction_code evidence columns per data_quality_master Appendix B Layer-A; the dead `enrichment` jsonb was NOT reused (the Gokind lesson: opaque blobs with no readers die). mapping-engine now also matches original_description so user rules written against the full bank text keep firing on stripped rows.
[2026-07-29] Transaction-method backfill scope: classification and title-stripping are FEED-ROW concepts (import_source present, not manual/mcp), enforced identically at ingest and in the 20260808090100 backfill, plus an adjective guard so "Egen insättning"/"Eget uttag"/"Intern överföring" keep their full titles even on feed rows (the phrase IS the meaning after a possessive/scope adjective). Chosen over vocabulary tweaks because the failure mode for unknown bank phrasings must be "row unchanged", and user-authored titles must never be rewritten by a channel vocabulary. A read-only prod dry-run script exists for coverage measurement but prod reads were left to the founder (permission-gated).
[2026-08-08] Compliance-bot finding on the transaction_method backfill (booked rows' titles rewritten without a rattelse trail) triaged as satisfied-by-design, not a blocker: BFL 5 kap 5 attaches to bokforingsposter, and the backfill touches no journal table; the verifikat description is snapshotted into journal_entries at commit and SIE #VER export reads journal_entries only (both verified in code, no report reads transactions.description lazily); the bank original is preserved byte-identical in original_description by the same UPDATE (enforced since 80ef1ee0, and prod has 0/25,566 feed rows lacking it). The stricter TRANSACTION_TITLE_LOCKED gate on booked rows blocks arbitrary user free-text renames, a different mutation class from a deterministic trailing-vocabulary strip that skips user-edited titles and keeps the original adjacent. Period-lock triggers sit on the journal tables and fiscal periods, not on transactions; the pg-upgrade CI run applied the backfill against seeded booked rows with all enforcement triggers active.
[2026-08-08] Fenced-JSON fix uses brace-slice, not fence-regex: also rescues preamble/postamble prose around the object, and degrades to the existing empty-result path when no braces exist.
[2026-08-08] extractJsonObject upgraded from brace-slice to depth-aware balanced scan after PR 1460 review: prose containing braces around the JSON no longer poisons the slice; first parseable candidate wins.
@@ -1,5 +1,8 @@
import { describe, it, expect, vi, beforeEach, afterAll } from 'vitest'
import { extractInvoiceFields } from '@/extensions/general/invoice-inbox/lib/extract-invoice-fields'
import {
extractInvoiceFields,
extractJsonObject,
} from '@/extensions/general/invoice-inbox/lib/extract-invoice-fields'
// Mock the Bedrock SDK so tests drive the JSON parser without
// network/credential needs.
@@ -157,6 +160,119 @@ describe('extractInvoiceFields', () => {
expect(data.supplier.name).toBeNull()
})
// ── Fenced / prefixed model output (Sonnet 5 regression, 2026-08) ──
// Sonnet 5 intermittently wraps the JSON in markdown fences despite the
// JSON-only instruction; a fifth of prod extractions came back empty
// because JSON.parse saw the backticks.
it('parses a response wrapped in ```json fences', async () => {
mockCreate.mockReturnValueOnce(
aiResponse('```json\n' + JSON.stringify(VALID_RESULT) + '\n```')
)
const { data } = await extractInvoiceFields({
buffer: Buffer.from('%PDF'),
mimeType: 'application/pdf',
fileName: 'kvitto.pdf',
})
expect(data.supplier.name).toBe('Anthropic, PBC')
expect(data.totals.total).toBe(6.25)
expect(data.confidence).toBe(1)
})
it('parses a response wrapped in bare ``` fences', async () => {
mockCreate.mockReturnValueOnce(
aiResponse('```\n' + JSON.stringify(VALID_RESULT) + '\n```')
)
const { data } = await extractInvoiceFields({
buffer: Buffer.from('%PDF'),
mimeType: 'application/pdf',
fileName: 'kvitto.pdf',
})
expect(data.totals.total).toBe(6.25)
expect(data.confidence).toBe(1)
})
it('parses a response with prose before and after the JSON object', async () => {
mockCreate.mockReturnValueOnce(
aiResponse(
'Here is the extracted data:\n```json\n' +
JSON.stringify(VALID_RESULT) +
'\n```\nLet me know if you need anything else.'
)
)
const { data } = await extractInvoiceFields({
buffer: Buffer.from('%PDF'),
mimeType: 'application/pdf',
fileName: 'kvitto.pdf',
})
expect(data.supplier.name).toBe('Anthropic, PBC')
expect(data.confidence).toBe(1)
})
it('parses JSON when the surrounding prose itself contains braces', async () => {
mockCreate.mockReturnValueOnce(
aiResponse(
'Note: fields use the shape {field: value}.\n```json\n' +
JSON.stringify(VALID_RESULT) +
'\n```\nAnything unclear {just ask}.'
)
)
const { data } = await extractInvoiceFields({
buffer: Buffer.from('%PDF'),
mimeType: 'application/pdf',
fileName: 'kvitto.pdf',
})
expect(data.supplier.name).toBe('Anthropic, PBC')
expect(data.totals.total).toBe(6.25)
expect(data.confidence).toBe(1)
})
it('handles braces inside JSON string values without ending the object early', async () => {
mockCreate.mockReturnValueOnce(
aiResponse(
'```json\n' +
JSON.stringify({
...VALID_RESULT,
supplier: { ...VALID_RESULT.supplier, address: 'Suite {B}, "Main" St 1' },
}) +
'\n```'
)
)
const { data } = await extractInvoiceFields({
buffer: Buffer.from('%PDF'),
mimeType: 'application/pdf',
fileName: 'kvitto.pdf',
})
expect(data.supplier.address).toBe('Suite {B}, "Main" St 1')
expect(data.confidence).toBe(1)
})
it('stays bounded on pathological brace-laden input and falls through unchanged', async () => {
// 100k unclosed braces: without the attempt cap this would scan
// quadratically; with it the helper bails fast and returns the input,
// which then lands in the existing empty-result path.
const pathological = '{'.repeat(100_000)
const startedAt = performance.now()
expect(extractJsonObject(pathological)).toBe(pathological)
expect(performance.now() - startedAt).toBeLessThan(1_000)
mockCreate.mockReturnValueOnce(aiResponse(pathological))
const { data } = await extractInvoiceFields({
buffer: Buffer.from('%PDF'),
mimeType: 'application/pdf',
fileName: 'f.pdf',
})
expect(data.totals.total).toBeNull()
expect(data.confidence).toBe(0)
})
it('skips scanning entirely for oversized input', async () => {
// Above the 256 KB cap the helper must not scan at all; the raw text
// passes through unchanged even though it contains valid JSON.
const oversized = 'x'.repeat(300 * 1024) + JSON.stringify(VALID_RESULT)
expect(extractJsonObject(oversized)).toBe(oversized)
})
it('returns empty result when AI response fails schema validation', async () => {
mockCreate.mockReturnValueOnce(
aiResponse({ supplier: { name: 'X' } /* missing required keys */ })
@@ -240,6 +240,61 @@ Rules:
- lineItems: include every line. Empty array is fine if the document has no itemised lines.
- vatBreakdown: include one entry per distinct VAT rate. Empty array is fine.`
// Sonnet 5 intermittently wraps its answer in markdown fences (```json ... ```)
// or adds prose around it, despite the JSON-only instruction in the system
// prompt. Scan for balanced top-level '{'..'}' candidates (string- and
// escape-aware, so braces inside JSON string values don't end a candidate
// early) and return the first one JSON.parse accepts; prose braces around the
// object form unparseable candidates and are skipped. Returns the input
// unchanged when no candidate parses, so the existing parse-failure path
// handles prose-only refusals. Zod validation downstream still rejects
// well-formed-but-wrong JSON.
// Bounds for the candidate scan below. Real model output is already capped
// by MAX_TOKENS (roughly 33 KB of text at 8192 tokens), so genuine responses
// never come near these; they exist so pathological or adversarially
// brace-laden text cannot make the scan quadratic (compliance review
// A.8.28). Oversized or exhausted inputs fall through to the raw text and
// land in the existing empty-result path.
const MAX_SCAN_INPUT_LENGTH = 256 * 1024
const MAX_CANDIDATE_ATTEMPTS = 50
export function extractJsonObject(raw: string): string {
if (raw.length > MAX_SCAN_INPUT_LENGTH) return raw
let attempts = 0
let start = raw.indexOf('{')
while (start !== -1 && attempts < MAX_CANDIDATE_ATTEMPTS) {
attempts++
let depth = 0
let inString = false
let escaped = false
for (let i = start; i < raw.length; i++) {
const ch = raw[i]
if (inString) {
if (escaped) escaped = false
else if (ch === '\\') escaped = true
else if (ch === '"') inString = false
} else if (ch === '"') {
inString = true
} else if (ch === '{') {
depth++
} else if (ch === '}') {
depth--
if (depth === 0) {
const candidate = raw.slice(start, i + 1)
try {
JSON.parse(candidate)
return candidate
} catch {
break
}
}
}
}
start = raw.indexOf('{', start + 1)
}
return raw
}
export function emptyResult(): InvoiceExtractionResult {
return {
documentKind: null,
@@ -422,7 +477,7 @@ export async function extractInvoiceFields(
})
}
const parsed = JSON.parse(rawText)
const parsed = JSON.parse(extractJsonObject(rawText))
const validated = ExtractionSchema.parse(parsed)
return {