* feat(ai): job-shaped AI service with OpenAI-compatible backend, extraction-first; stop extracting every inbox document twice Sovereign plan WS1 PR1 (#1406 Tier 2, extraction-first, aligned with the AI surface audit). lib/ai grows a job-shaped service (generateText / generateStructured / extractFromDocument; no streaming members yet, see plan rule R3): - services/anthropic-family delegates to the existing createAiClient() and sends the exact request literals the inbox extractor sent before (request-shape tests deep-equal them), so hosted Bedrock stays byte-identical. - services/openai-compatible talks to any chat-completions endpoint (BYO Swedish provider) via Vercel AI SDK 6.x, exact-pinned and guarded: images as parts, PDFs rasterized with poppler (AI_PDF_MODE) or sent natively, AI_VISION / AI_STRICT_JSON declared, honest skips (ai_no_vision, pdf_rasterizer_missing) instead of fake failures. - config.ts: AI_PROVIDER/AI_BASE_URL/AI_API_KEY/AI_MODEL and per-tier AI_*_MODEL with the legacy BEDROCK_* names kept as the same overrides; getAiStatus() is the single source of truth for "is AI wired up". - provider.ts: openai-compatible in the auto-detect chain (after Bedrock and the direct API); createAiClient() refuses it loudly. Document extraction moves onto the service and gets the audit's fixes: - Inbox documents were extracted TWICE (pipeline A ran inside uploadDocument() before the inbox row existed, so its dedupe branch never fired; 3 707 + 1 666 calls / 30 d). The inbox now declares extractionOwner on the upload, the extension yields, and the inbox mirrors its single outcome onto document_attachments from every writer (sync, deferred, attach, retry, MCP). - Every "no extraction will ever happen" outcome is stamped (skipped:no_ai_entitlement / ai_unconfigured / system_generated / ...); the status route maps the quiet ones to 'disabled' on the first poll instead of a 30 s client timeout. Prod showed 309 of the 327 never-extracted uploads were the paywall working silently. - Self-generated documents (our own invoice PDFs, payout files) are no longer OCR'd. - Agent invoke answers 503 ai_unconfigured when the deployment has no assistant backend, distinct from the paywall. Guard: new direct-ai-client antipattern check (shrink-only allowlist of the pre-abstraction SDK callers) plus exact pins for @anthropic-ai/sdk, ai and @ai-sdk/openai-compatible. Verified: 15 958 unit tests green, guards, lint ratchet, typecheck, and a live smoke against hosted Bedrock through the new service (ping, streamed tool turn, thinking+cache, PDF extraction). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ai): make AI_API_KEY optional for OpenAI-compatible endpoints (keyless local model servers) A local model server (llama.cpp's server, Ollama /v1, LM Studio, vLLM) usually has no auth. Before, the OpenAI-compatible backend required both AI_BASE_URL and AI_API_KEY to count as configured, so running Accounted on a local model meant setting a meaningless placeholder key. - resolveAiProvider / hasAiCredentials: a base URL alone is now enough. - services/openai-compatible: only send Authorization: Bearer when AI_API_KEY is set, so a keyless server is never handed an empty bearer; a hosted provider that needs a key still sets it. - Docs (SELF-HOSTING Option 3: local-model example, key marked optional), DECISIONS. Verified: with no AI_API_KEY, just AI_BASE_URL + AI_MODEL, getAiStatus() reports configured=true / provider=openai-compatible (live). lib/ai suite 71 green; tsc, guards, lint clean. Bedrock/Anthropic logic unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
136 lines
4.7 KiB
TypeScript
136 lines
4.7 KiB
TypeScript
// Job-shaped AI service interface.
|
|
//
|
|
// Call sites describe WHAT they need (a text answer, a schema-shaped object,
|
|
// the fields read out of a document) rather than HOW a particular backend is
|
|
// spoken to. The Anthropic-family service (Bedrock and the direct API) keeps
|
|
// hosted byte-identical by delegating to the existing client factory in
|
|
// lib/ai/provider.ts; the OpenAI-compatible service talks to any endpoint
|
|
// that implements the chat-completions API (the Swedish inference providers a
|
|
// sovereign self-host points at) through the Vercel AI SDK.
|
|
//
|
|
// Streaming members (the chat loop) are deliberately absent until the chat
|
|
// runtime decision is taken: see the Sovereign plan, alignment rule R3.
|
|
|
|
export type AiProviderKind = 'bedrock' | 'anthropic' | 'openai-compatible'
|
|
|
|
/**
|
|
* Model tiers. `assistant` is the conversational/default tier, `heavy` the
|
|
* deep-reasoning tier (supplier-invoice review, VAT review, bokslut), and
|
|
* `extraction` the document-reading tier (a vision model on OpenAI-compatible
|
|
* endpoints; Claude reads PDFs natively).
|
|
*/
|
|
export type AiTier = 'assistant' | 'heavy' | 'extraction'
|
|
|
|
export interface AiCapabilities {
|
|
/** PDF bytes can be sent as a native document part, no rasterization. */
|
|
pdfNative: boolean
|
|
/** Images (and therefore scanned receipts) can be read at all. */
|
|
imageInput: boolean
|
|
toolUse: boolean
|
|
/** The backend can be forced to answer with one named tool. */
|
|
forcedToolChoice: boolean
|
|
/** The backend enforces a JSON schema on the output (response_format). */
|
|
strictJsonSchema: boolean
|
|
}
|
|
|
|
export type AiImageMediaType = 'image/jpeg' | 'image/png' | 'image/webp' | 'image/gif'
|
|
|
|
export type AiDocumentInput =
|
|
| { kind: 'pdf'; data: Buffer; fileName?: string }
|
|
| { kind: 'image'; data: Buffer; mediaType: AiImageMediaType }
|
|
/** Plain text already extracted by the caller (HTML mail invoices). Works on every model, vision or not. */
|
|
| { kind: 'text'; text: string }
|
|
|
|
export interface AiUsage {
|
|
inputTokens: number | null
|
|
outputTokens: number | null
|
|
cacheCreationInputTokens: number | null
|
|
cacheReadInputTokens: number | null
|
|
}
|
|
|
|
export interface GenerateTextRequest {
|
|
tier: AiTier
|
|
system?: string
|
|
prompt: string
|
|
maxTokens: number
|
|
}
|
|
|
|
export interface GenerateTextResult {
|
|
text: string
|
|
model: string
|
|
usage: AiUsage
|
|
}
|
|
|
|
export interface GenerateStructuredRequest {
|
|
tier: AiTier
|
|
system?: string
|
|
prompt: string
|
|
maxTokens: number
|
|
schema: {
|
|
name: string
|
|
description?: string
|
|
/** JSON Schema (draft-07 subset) for the expected object. Hand-maintained by the caller. */
|
|
jsonSchema: Record<string, unknown>
|
|
}
|
|
}
|
|
|
|
export interface GenerateStructuredResult {
|
|
/** The model's object, NOT validated: callers run their own Zod parse. */
|
|
value: unknown
|
|
model: string
|
|
usage: AiUsage
|
|
}
|
|
|
|
export interface ExtractFromDocumentRequest {
|
|
document: AiDocumentInput
|
|
/** Byte-stable system prompt. The Anthropic-family service marks it as a prompt-cache breakpoint. */
|
|
system: string
|
|
/** Trailing user instruction placed after the document part(s). */
|
|
instruction: string
|
|
maxTokens: number
|
|
/**
|
|
* Optional JSON schema for the answer. Used only when strict JSON mode is
|
|
* on AND the backend supports it; otherwise the model answers in prose and
|
|
* the caller's JSON extraction + Zod parse do the work (works everywhere).
|
|
*/
|
|
jsonSchema?: Record<string, unknown>
|
|
}
|
|
|
|
export type ExtractionSkipReason =
|
|
| 'ai_unconfigured'
|
|
| 'ai_no_vision'
|
|
| 'pdf_rasterizer_missing'
|
|
| 'pdf_rasterize_failed'
|
|
|
|
export type ExtractFromDocumentResult =
|
|
| { ok: true; text: string; model: string; usage: AiUsage; pagesRasterized?: number }
|
|
| { ok: false; skipped: ExtractionSkipReason }
|
|
|
|
export interface AiService {
|
|
readonly provider: AiProviderKind
|
|
readonly capabilities: AiCapabilities
|
|
/** Provider-form model id for a tier (Bedrock inference-profile prefix applied, etc.). */
|
|
modelFor(tier: AiTier): string
|
|
generateText(req: GenerateTextRequest): Promise<GenerateTextResult>
|
|
generateStructured(req: GenerateStructuredRequest): Promise<GenerateStructuredResult>
|
|
extractFromDocument(req: ExtractFromDocumentRequest): Promise<ExtractFromDocumentResult>
|
|
}
|
|
|
|
export type AiPdfMode = 'native' | 'rasterize'
|
|
|
|
export interface AiStatus {
|
|
provider: AiProviderKind
|
|
/** Credentials AND (for OpenAI-compatible) a model id are present. */
|
|
configured: boolean
|
|
reason: 'ok' | 'no_credentials' | 'no_model'
|
|
capabilities: AiCapabilities
|
|
models: Record<AiTier, string | null>
|
|
pdfMode: AiPdfMode
|
|
/**
|
|
* Whether the in-app assistant (chat loop) can run. The loop still speaks
|
|
* the Anthropic messages surface directly, so it needs the Anthropic family
|
|
* until its streaming port lands; extraction and single-call jobs do not.
|
|
*/
|
|
assistantAvailable: boolean
|
|
}
|