Files
accounted/lib/ai/index.ts
T
c7a75d069d feat(ai): job-shaped AI service with OpenAI-compatible backend, extraction-first; stop extracting every inbox document twice (#1740)
* feat(ai): job-shaped AI service with OpenAI-compatible backend, extraction-first; stop extracting every inbox document twice

Sovereign plan WS1 PR1 (#1406 Tier 2, extraction-first, aligned with the
AI surface audit).

lib/ai grows a job-shaped service (generateText / generateStructured /
extractFromDocument; no streaming members yet, see plan rule R3):
- services/anthropic-family delegates to the existing createAiClient()
  and sends the exact request literals the inbox extractor sent before
  (request-shape tests deep-equal them), so hosted Bedrock stays
  byte-identical.
- services/openai-compatible talks to any chat-completions endpoint
  (BYO Swedish provider) via Vercel AI SDK 6.x, exact-pinned and
  guarded: images as parts, PDFs rasterized with poppler (AI_PDF_MODE)
  or sent natively, AI_VISION / AI_STRICT_JSON declared, honest skips
  (ai_no_vision, pdf_rasterizer_missing) instead of fake failures.
- config.ts: AI_PROVIDER/AI_BASE_URL/AI_API_KEY/AI_MODEL and per-tier
  AI_*_MODEL with the legacy BEDROCK_* names kept as the same overrides;
  getAiStatus() is the single source of truth for "is AI wired up".
- provider.ts: openai-compatible in the auto-detect chain (after Bedrock
  and the direct API); createAiClient() refuses it loudly.

Document extraction moves onto the service and gets the audit's fixes:
- Inbox documents were extracted TWICE (pipeline A ran inside
  uploadDocument() before the inbox row existed, so its dedupe branch
  never fired; 3 707 + 1 666 calls / 30 d). The inbox now declares
  extractionOwner on the upload, the extension yields, and the inbox
  mirrors its single outcome onto document_attachments from every
  writer (sync, deferred, attach, retry, MCP).
- Every "no extraction will ever happen" outcome is stamped
  (skipped:no_ai_entitlement / ai_unconfigured / system_generated /
  ...); the status route maps the quiet ones to 'disabled' on the first
  poll instead of a 30 s client timeout. Prod showed 309 of the 327
  never-extracted uploads were the paywall working silently.
- Self-generated documents (our own invoice PDFs, payout files) are no
  longer OCR'd.
- Agent invoke answers 503 ai_unconfigured when the deployment has no
  assistant backend, distinct from the paywall.

Guard: new direct-ai-client antipattern check (shrink-only allowlist of
the pre-abstraction SDK callers) plus exact pins for @anthropic-ai/sdk,
ai and @ai-sdk/openai-compatible.

Verified: 15 958 unit tests green, guards, lint ratchet, typecheck, and a
live smoke against hosted Bedrock through the new service (ping, streamed
tool turn, thinking+cache, PDF extraction).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ai): make AI_API_KEY optional for OpenAI-compatible endpoints (keyless local model servers)

A local model server (llama.cpp's server, Ollama /v1, LM Studio, vLLM)
usually has no auth. Before, the OpenAI-compatible backend required both
AI_BASE_URL and AI_API_KEY to count as configured, so running Accounted on a
local model meant setting a meaningless placeholder key.

- resolveAiProvider / hasAiCredentials: a base URL alone is now enough.
- services/openai-compatible: only send Authorization: Bearer when AI_API_KEY
  is set, so a keyless server is never handed an empty bearer; a hosted
  provider that needs a key still sets it.
- Docs (SELF-HOSTING Option 3: local-model example, key marked optional),
  DECISIONS.

Verified: with no AI_API_KEY, just AI_BASE_URL + AI_MODEL, getAiStatus()
reports configured=true / provider=openai-compatible (live). lib/ai suite
71 green; tsc, guards, lint clean. Bedrock/Anthropic logic unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 19:39:08 +02:00

67 lines
2.0 KiB
TypeScript

import { getAiStatus, readAiConfig, type ResolvedAiConfig } from './config'
import { createAnthropicFamilyService } from './services/anthropic-family'
import { createOpenAICompatibleService } from './services/openai-compatible'
import type { AiService } from './types'
export type {
AiCapabilities,
AiDocumentInput,
AiImageMediaType,
AiPdfMode,
AiProviderKind,
AiService,
AiStatus,
AiTier,
AiUsage,
ExtractFromDocumentRequest,
ExtractFromDocumentResult,
ExtractionSkipReason,
GenerateStructuredRequest,
GenerateStructuredResult,
GenerateTextRequest,
GenerateTextResult,
} from './types'
export { getAiStatus, readAiConfig, resolveTierModel } from './config'
export { extractJsonObject } from './json'
// One service per resolved configuration. Keyed on the non-secret parts of
// the config plus credential presence, so a changed environment (tests, the
// smoke script) gets a fresh service while a long-lived process reuses one.
let cached: { key: string; service: AiService } | null = null
function cacheKey(cfg: ResolvedAiConfig): string {
return JSON.stringify({
provider: cfg.provider,
configured: cfg.configured,
baseUrl: cfg.baseUrl,
hasKey: !!cfg.apiKey,
models: cfg.models,
vision: cfg.vision,
strictJson: cfg.strictJson,
pdfMode: cfg.pdfMode,
pdfMaxPages: cfg.pdfMaxPages,
})
}
/**
* The AI service for this deployment. Never throws on construction: an
* unconfigured deployment still gets a service whose extractFromDocument
* answers `skipped: ai_unconfigured`, so upload paths degrade quietly.
*/
export function getAiService(): AiService {
const cfg = readAiConfig()
const key = cacheKey(cfg)
if (cached && cached.key === key) return cached.service
const service =
cfg.provider === 'openai-compatible'
? createOpenAICompatibleService(cfg)
: createAnthropicFamilyService(cfg)
cached = { key, service }
return service
}
/** Tests only: drop the cached service so the next call re-reads the environment. */
export function resetAiServiceForTests(): void {
cached = null
}