* feat(ai): job-shaped AI service with OpenAI-compatible backend, extraction-first; stop extracting every inbox document twice Sovereign plan WS1 PR1 (#1406 Tier 2, extraction-first, aligned with the AI surface audit). lib/ai grows a job-shaped service (generateText / generateStructured / extractFromDocument; no streaming members yet, see plan rule R3): - services/anthropic-family delegates to the existing createAiClient() and sends the exact request literals the inbox extractor sent before (request-shape tests deep-equal them), so hosted Bedrock stays byte-identical. - services/openai-compatible talks to any chat-completions endpoint (BYO Swedish provider) via Vercel AI SDK 6.x, exact-pinned and guarded: images as parts, PDFs rasterized with poppler (AI_PDF_MODE) or sent natively, AI_VISION / AI_STRICT_JSON declared, honest skips (ai_no_vision, pdf_rasterizer_missing) instead of fake failures. - config.ts: AI_PROVIDER/AI_BASE_URL/AI_API_KEY/AI_MODEL and per-tier AI_*_MODEL with the legacy BEDROCK_* names kept as the same overrides; getAiStatus() is the single source of truth for "is AI wired up". - provider.ts: openai-compatible in the auto-detect chain (after Bedrock and the direct API); createAiClient() refuses it loudly. Document extraction moves onto the service and gets the audit's fixes: - Inbox documents were extracted TWICE (pipeline A ran inside uploadDocument() before the inbox row existed, so its dedupe branch never fired; 3 707 + 1 666 calls / 30 d). The inbox now declares extractionOwner on the upload, the extension yields, and the inbox mirrors its single outcome onto document_attachments from every writer (sync, deferred, attach, retry, MCP). - Every "no extraction will ever happen" outcome is stamped (skipped:no_ai_entitlement / ai_unconfigured / system_generated / ...); the status route maps the quiet ones to 'disabled' on the first poll instead of a 30 s client timeout. Prod showed 309 of the 327 never-extracted uploads were the paywall working silently. - Self-generated documents (our own invoice PDFs, payout files) are no longer OCR'd. - Agent invoke answers 503 ai_unconfigured when the deployment has no assistant backend, distinct from the paywall. Guard: new direct-ai-client antipattern check (shrink-only allowlist of the pre-abstraction SDK callers) plus exact pins for @anthropic-ai/sdk, ai and @ai-sdk/openai-compatible. Verified: 15 958 unit tests green, guards, lint ratchet, typecheck, and a live smoke against hosted Bedrock through the new service (ping, streamed tool turn, thinking+cache, PDF extraction). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ai): make AI_API_KEY optional for OpenAI-compatible endpoints (keyless local model servers) A local model server (llama.cpp's server, Ollama /v1, LM Studio, vLLM) usually has no auth. Before, the OpenAI-compatible backend required both AI_BASE_URL and AI_API_KEY to count as configured, so running Accounted on a local model meant setting a meaningless placeholder key. - resolveAiProvider / hasAiCredentials: a base URL alone is now enough. - services/openai-compatible: only send Authorization: Bearer when AI_API_KEY is set, so a keyless server is never handed an empty bearer; a hosted provider that needs a key still sets it. - Docs (SELF-HOSTING Option 3: local-model example, key marked optional), DECISIONS. Verified: with no AI_API_KEY, just AI_BASE_URL + AI_MODEL, getAiStatus() reports configured=true / provider=openai-compatible (live). lib/ai suite 71 green; tsc, guards, lint clean. Bedrock/Anthropic logic unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
207 lines
7.8 KiB
TypeScript
207 lines
7.8 KiB
TypeScript
import type { Extension } from '@/lib/extensions/types'
|
|
import type { SupabaseClient } from '@supabase/supabase-js'
|
|
import { extractInvoiceFields } from '@/extensions/general/invoice-inbox/lib/extract-invoice-fields'
|
|
import { getAiStatus } from '@/lib/ai'
|
|
import { hasCapability } from '@/lib/entitlements/has-capability'
|
|
import { CAPABILITY } from '@/lib/entitlements/keys'
|
|
import { createLogger } from '@/lib/logger'
|
|
import { createServiceClient } from '@/lib/supabase/server'
|
|
import type { DocumentAttachment } from '@/types'
|
|
|
|
const log = createLogger('document-extraction')
|
|
|
|
// Mime types the extraction can read (a vision model reads PDFs natively on
|
|
// Claude; OpenAI-compatible backends rasterize them). Anything else (HEIC,
|
|
// ZIP, TXT, bank JSON, …) is skipped: extracted_at still gets stamped so the
|
|
// row is marked as "attempted, not eligible".
|
|
const SUPPORTED_MIME_TYPES = new Set([
|
|
'application/pdf',
|
|
'image/jpeg',
|
|
'image/png',
|
|
'image/webp',
|
|
'image/gif',
|
|
])
|
|
|
|
// AI-extraction extension: paid AI tier only.
|
|
//
|
|
// Subscribes to the existing document.uploaded event bus topic and runs the
|
|
// configured extraction model (reusing invoice-inbox's extractInvoiceFields)
|
|
// on every uploaded receipt or invoice. Writes the result to
|
|
// document_attachments.extracted_data so the agent intent capture can use
|
|
// it without re-asking the user.
|
|
//
|
|
// Idempotency: skips when extracted_at is already set on the row.
|
|
//
|
|
// Ownership: documents that enter through the invoice inbox (web upload,
|
|
// email, WhatsApp, MCP upload) are extracted by the inbox itself, which
|
|
// mirrors its result onto document_attachments when done. The inbox marks
|
|
// the upload with `extractionOwner: 'invoice-inbox'` on the event payload
|
|
// and this handler yields. It used to race instead: the inbox row did not
|
|
// exist yet when this handler ran inside uploadDocument(), so every inbox
|
|
// document was extracted twice (two paid model calls per receipt).
|
|
//
|
|
// Every outcome that means "no extraction will ever happen here" is stamped
|
|
// as `skipped:<reason>` so the polling UI stops immediately instead of
|
|
// timing out: paywall (`no_ai_entitlement`), no AI configured on this
|
|
// deployment (`ai_unconfigured`), self-generated documents (our own invoice
|
|
// PDFs, payout files: `system_generated`), unsupported types.
|
|
//
|
|
// Free tier: disable this extension in extensions.config.json. Uploads
|
|
// still work; the agent intent will see null extracted_data and either
|
|
// ask the user or call gnubok_get_document_content at chat-time.
|
|
export const documentExtractionExtension: Extension = {
|
|
id: 'document-extraction',
|
|
name: 'AI document extraction',
|
|
version: '1.1.0',
|
|
|
|
eventHandlers: [
|
|
{
|
|
eventType: 'document.uploaded',
|
|
handler: async (payload) => {
|
|
const { document, companyId, extractionOwner } = payload as {
|
|
document: DocumentAttachment
|
|
userId: string
|
|
companyId: string
|
|
extractionOwner?: 'invoice-inbox'
|
|
}
|
|
if (extractionOwner === 'invoice-inbox') return
|
|
await extractAndPersist(document, companyId)
|
|
},
|
|
},
|
|
],
|
|
}
|
|
|
|
async function stamp(
|
|
supabase: SupabaseClient,
|
|
documentId: string,
|
|
extractionModel: string
|
|
): Promise<void> {
|
|
const { error } = await supabase
|
|
.from('document_attachments')
|
|
.update({ extracted_at: new Date().toISOString(), extraction_model: extractionModel })
|
|
.eq('id', documentId)
|
|
if (error) log.warn('stamp failed', { doc: documentId, extractionModel, err: error.message })
|
|
}
|
|
|
|
async function extractAndPersist(
|
|
document: DocumentAttachment,
|
|
companyId: string,
|
|
): Promise<void> {
|
|
// Service-role client: the handler runs out-of-band of the request that
|
|
// emitted the event, so we don't have user cookies. RLS doesn't fit:
|
|
// events have no user context.
|
|
const supabase: SupabaseClient = createServiceClient()
|
|
|
|
// Cheap gates first, straight from the event payload: no DB round trip for
|
|
// the thousands of bank-statement JSON documents a sync produces.
|
|
const mimeType = (document.mime_type as string | null) ?? null
|
|
if (!mimeType || !SUPPORTED_MIME_TYPES.has(mimeType)) {
|
|
await stamp(supabase, document.id, 'skipped:unsupported_mime')
|
|
return
|
|
}
|
|
// Documents the system produced itself (our own invoice PDFs, payment
|
|
// files, filings) carry nothing to extract; reading them back with a paid
|
|
// model is waste, on hosted and doubly so on a BYO-key self-host.
|
|
if (document.upload_source === 'system') {
|
|
await stamp(supabase, document.id, 'skipped:system_generated')
|
|
return
|
|
}
|
|
|
|
// Idempotency guard: never re-extract a row that already has extracted_at.
|
|
const { data: existing, error: existingErr } = await supabase
|
|
.from('document_attachments')
|
|
.select('id, mime_type, storage_path, extracted_at')
|
|
.eq('id', document.id)
|
|
.single()
|
|
if (existingErr || !existing) {
|
|
log.warn('document not found, skipping extraction', {
|
|
doc: document.id,
|
|
err: existingErr?.message,
|
|
})
|
|
return
|
|
}
|
|
if (existing.extracted_at) {
|
|
return
|
|
}
|
|
|
|
// No AI configured on this deployment: stamp and stop. This is the
|
|
// self-host "key not set yet" state; the status route reports it as
|
|
// disabled on the first poll instead of after a 30 s timeout.
|
|
if (!getAiStatus().configured) {
|
|
await stamp(supabase, document.id, 'skipped:ai_unconfigured')
|
|
return
|
|
}
|
|
|
|
// Paywall: the free/manual tier never triggers paid extraction. Stamp it so
|
|
// the row reads "attempted, not entitled" rather than NULL forever (which
|
|
// the polling UI could only interpret by timing out).
|
|
if (!(await hasCapability(supabase, companyId, CAPABILITY.ai))) {
|
|
log.info('extraction skipped, ai capability not entitled', { doc: document.id, companyId })
|
|
await stamp(supabase, document.id, 'skipped:no_ai_entitlement')
|
|
return
|
|
}
|
|
|
|
// Download the file from Supabase Storage. The bucket is private: the
|
|
// service-role client can read any path.
|
|
const storagePath = existing.storage_path as string | null
|
|
if (!storagePath) {
|
|
log.warn('document has no storage_path, skipping', { doc: document.id })
|
|
await stamp(supabase, document.id, 'failed:no_storage_path')
|
|
return
|
|
}
|
|
|
|
const { data: blob, error: dlErr } = await supabase.storage
|
|
.from('documents')
|
|
.download(storagePath)
|
|
if (dlErr || !blob) {
|
|
log.warn('storage download failed', { doc: document.id, err: dlErr?.message })
|
|
await stamp(supabase, document.id, 'failed:storage_download')
|
|
return
|
|
}
|
|
const buffer = Buffer.from(await blob.arrayBuffer())
|
|
|
|
let extractedData: Record<string, unknown> | null = null
|
|
let model: string
|
|
try {
|
|
const { data, rawText, model: usedModel, skipped } = await extractInvoiceFields({
|
|
buffer,
|
|
mimeType,
|
|
fileName: (document.file_name as string) || 'document',
|
|
})
|
|
// extractInvoiceFields returns an "empty" result on failure rather than
|
|
// throwing. `skipped` means no model call was made (and why); a null
|
|
// rawText with no skip means the call failed or the JSON parse did.
|
|
if (skipped) {
|
|
await stamp(supabase, document.id, `skipped:${skipped}`)
|
|
return
|
|
}
|
|
if (!rawText) {
|
|
await stamp(supabase, document.id, 'failed:no_raw_text')
|
|
return
|
|
}
|
|
extractedData = data as unknown as Record<string, unknown>
|
|
model = usedModel ?? 'unknown'
|
|
} catch (err) {
|
|
log.warn('extraction threw', {
|
|
doc: document.id,
|
|
err: err instanceof Error ? err.message : String(err),
|
|
})
|
|
await stamp(supabase, document.id, 'failed:exception')
|
|
return
|
|
}
|
|
|
|
const { error: updateErr } = await supabase
|
|
.from('document_attachments')
|
|
.update({
|
|
extracted_data: extractedData,
|
|
extracted_at: new Date().toISOString(),
|
|
extraction_model: model,
|
|
})
|
|
.eq('id', document.id)
|
|
if (updateErr) {
|
|
log.warn('persist failed', { doc: document.id, err: updateErr.message, companyId })
|
|
return
|
|
}
|
|
log.info('extraction persisted', { doc: document.id, model, companyId })
|
|
}
|