feat(ai): job-shaped AI service with OpenAI-compatible backend, extraction-first; stop extracting every inbox document twice (#1740)
* feat(ai): job-shaped AI service with OpenAI-compatible backend, extraction-first; stop extracting every inbox document twice Sovereign plan WS1 PR1 (#1406 Tier 2, extraction-first, aligned with the AI surface audit). lib/ai grows a job-shaped service (generateText / generateStructured / extractFromDocument; no streaming members yet, see plan rule R3): - services/anthropic-family delegates to the existing createAiClient() and sends the exact request literals the inbox extractor sent before (request-shape tests deep-equal them), so hosted Bedrock stays byte-identical. - services/openai-compatible talks to any chat-completions endpoint (BYO Swedish provider) via Vercel AI SDK 6.x, exact-pinned and guarded: images as parts, PDFs rasterized with poppler (AI_PDF_MODE) or sent natively, AI_VISION / AI_STRICT_JSON declared, honest skips (ai_no_vision, pdf_rasterizer_missing) instead of fake failures. - config.ts: AI_PROVIDER/AI_BASE_URL/AI_API_KEY/AI_MODEL and per-tier AI_*_MODEL with the legacy BEDROCK_* names kept as the same overrides; getAiStatus() is the single source of truth for "is AI wired up". - provider.ts: openai-compatible in the auto-detect chain (after Bedrock and the direct API); createAiClient() refuses it loudly. Document extraction moves onto the service and gets the audit's fixes: - Inbox documents were extracted TWICE (pipeline A ran inside uploadDocument() before the inbox row existed, so its dedupe branch never fired; 3 707 + 1 666 calls / 30 d). The inbox now declares extractionOwner on the upload, the extension yields, and the inbox mirrors its single outcome onto document_attachments from every writer (sync, deferred, attach, retry, MCP). - Every "no extraction will ever happen" outcome is stamped (skipped:no_ai_entitlement / ai_unconfigured / system_generated / ...); the status route maps the quiet ones to 'disabled' on the first poll instead of a 30 s client timeout. Prod showed 309 of the 327 never-extracted uploads were the paywall working silently. - Self-generated documents (our own invoice PDFs, payout files) are no longer OCR'd. - Agent invoke answers 503 ai_unconfigured when the deployment has no assistant backend, distinct from the paywall. Guard: new direct-ai-client antipattern check (shrink-only allowlist of the pre-abstraction SDK callers) plus exact pins for @anthropic-ai/sdk, ai and @ai-sdk/openai-compatible. Verified: 15 958 unit tests green, guards, lint ratchet, typecheck, and a live smoke against hosted Bedrock through the new service (ping, streamed tool turn, thinking+cache, PDF extraction). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ai): make AI_API_KEY optional for OpenAI-compatible endpoints (keyless local model servers) A local model server (llama.cpp's server, Ollama /v1, LM Studio, vLLM) usually has no auth. Before, the OpenAI-compatible backend required both AI_BASE_URL and AI_API_KEY to count as configured, so running Accounted on a local model meant setting a meaningless placeholder key. - resolveAiProvider / hasAiCredentials: a base URL alone is now enough. - services/openai-compatible: only send Authorization: Bearer when AI_API_KEY is set, so a keyless server is never handed an empty bearer; a hosted provider that needs a key still sets it. - Docs (SELF-HOSTING Option 3: local-model example, key marked optional), DECISIONS. Verified: with no AI_API_KEY, just AI_BASE_URL + AI_MODEL, getAiStatus() reports configured=true / provider=openai-compatible (live). lib/ai suite 71 green; tsc, guards, lint clean. Bedrock/Anthropic logic unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
Jakob Wennberg
parent
fbf47649f2
commit
c7a75d069d
+23
-7
@@ -68,21 +68,37 @@ RECEIPT_HUNT_COMPANY_IDS=
|
||||
# NEXT_PUBLIC_GOOGLE_AUTH_ENABLED=true
|
||||
|
||||
# ── Optional: extension features (core runs without these) ─
|
||||
# AI features (document extraction + AI assistant). Two ways to provide a key;
|
||||
# set one of them. AI_PROVIDER (bedrock|anthropic) forces the choice if both
|
||||
# are present, which otherwise resolves to Bedrock.
|
||||
# AI features (document extraction + AI assistant). Three ways to provide a
|
||||
# backend; set one of them. AI_PROVIDER (bedrock|anthropic|openai-compatible)
|
||||
# forces the choice if several are present; otherwise Bedrock wins, then the
|
||||
# direct Anthropic API, then an OpenAI-compatible endpoint.
|
||||
#
|
||||
# 1. Claude via AWS Bedrock. Needs an AWS account with Bedrock model access to
|
||||
# Claude. Keeps inference in eu-north-1, which is what hosted runs.
|
||||
# AWS_ACCESS_KEY_ID=
|
||||
# AWS_SECRET_ACCESS_KEY=
|
||||
# AWS_REGION=eu-north-1
|
||||
# BEDROCK_MODEL_ID=
|
||||
#
|
||||
# 2. Claude via the direct Anthropic API. No AWS account needed, so this is
|
||||
# usually the self-hosted option. Note that it has no EU-residency
|
||||
# guarantee: use Bedrock if you need one.
|
||||
# 2. Claude via the direct Anthropic API. No AWS account needed. Note that it
|
||||
# has no EU-residency guarantee: use Bedrock if you need one.
|
||||
# ANTHROPIC_API_KEY=
|
||||
#
|
||||
# 3. Any OpenAI-compatible endpoint (chat-completions API), e.g. a Swedish
|
||||
# inference provider for a sovereign self-host. Document extraction and
|
||||
# single-call AI jobs run here; the in-app chat assistant does not yet.
|
||||
# A model id is required (no default exists for an arbitrary endpoint).
|
||||
# AI_BASE_URL=https://api.example.se/v1
|
||||
# AI_API_KEY=
|
||||
# AI_MODEL= # default model for every tier
|
||||
# AI_EXTRACTION_MODEL= # per-tier overrides (also AI_ASSISTANT_MODEL, AI_HEAVY_MODEL);
|
||||
# # legacy BEDROCK_MODEL_ID / BEDROCK_SONNET_MODEL_ID /
|
||||
# # BEDROCK_OPUS_MODEL_ID keep working as the same overrides
|
||||
# AI_VISION=true # openai-compatible only: set false for a text-only model
|
||||
# # (images/PDFs are then skipped honestly; HTML mail still extracts)
|
||||
# AI_PDF_MODE=auto # auto (Claude: native, others: rasterize with poppler) | native | rasterize
|
||||
# AI_PDF_MAX_PAGES=4 # pages rasterized per PDF
|
||||
# AI_STRICT_JSON=false # openai-compatible only: response_format json_schema when the provider enforces it
|
||||
# AI_EXTRACTION_MAX_TOKENS=8192 # output cap for document extraction (legacy BEDROCK_MAX_TOKENS)
|
||||
# AI_PROVIDER=
|
||||
# Bank connections (Enable Banking)
|
||||
# ENABLE_BANKING_APP_ID=
|
||||
|
||||
Reference in New Issue
Block a user