* feat(ai): resolve the Claude backend from the environment Tier 1 of #1406: a self-hosted deployment can now run every AI feature on a plain ANTHROPIC_API_KEY, with no AWS account. Hosted behaviour is unchanged. lib/ai/provider.ts resolves the backend once, from the environment: AI_PROVIDER explicit override, bedrock|anthropic AWS static key pair Bedrock ANTHROPIC_API_KEY the direct Anthropic API nothing set Bedrock, so the AWS credential provider chain (instance profile, IRSA) still resolves Bedrock deliberately wins when both credential sets are present. EU residency in eu-north-1 is a BFL/GDPR posture rather than a default, so adding an Anthropic key for an experiment must not silently move production inference out of the region. AI_PROVIDER is the way to say you meant it. Model ids are written bare in code and prefixed to eu.anthropic.* only for Bedrock, which needs the cross-region inference profile for on-demand throughput. An operator override that already carries a prefix passes through untouched, so BEDROCK_MODEL_ID and friends keep working as written. Converted call sites: the agent composer, invoice-inbox extraction, the document-extraction model label, and both receipt-hunt clients. The last two are not named in the issue, which predates receipt-hunt landing in main. @anthropic-ai/sdk is declared at 0.95.0, the version @anthropic-ai/bedrock-sdk 0.29.1 already pulled in transitively, so the lockfile dedupes to one copy with no new download. scripts/smoke-bedrock.ts becomes scripts/smoke-ai.ts and grows two steps. Unit tests can only prove which provider and model id get resolved; they cannot prove the resulting request is one the backend accepts. The script now sends real traffic over all three shapes the app uses: a plain create, a streamed turn carrying adaptive thinking, an effort level, an hour-long cache breakpoint and a tool, and document extraction end to end when given a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> * docs(self-hosting): document the AI smoke test The script added alongside the provider split is what closes the #1406 acceptance criterion ("document extraction and the assistant both work"), so a self-hoster needs to know it exists. Covers both invocations and states that it exits non-zero, which is what makes it usable as a post-deploy check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> * test(ai): split the smoke test's thinking probe from its tool probe The combined probe could not falsify what it claimed to. It asked a question that needs a tool call, so the tool was used and adaptive thinking correctly declined to reason about it: the zero thinking-block count that came back was uninformative rather than a signal. 2a keeps the tool and drops thinking. 2b asks a question with several dependent steps (reverse charge, then a partial deduction, then the affected boxes) so that a model honouring the parameter must reason, and reports the thinking text length as well as the block count, since display:"summarized" can yield blocks with empty text. The cached system prompt is also padded past the 1024-token minimum cacheable prefix. Below that the API caches nothing and reports no error, so the old probe's cache counters read zero whether or not caching worked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> * fix(document-extraction): stop requiring AWS_REGION in the manifest The extension now needs one of two credential sets, AWS static keys or ANTHROPIC_API_KEY, and the manifest schema cannot express "one of". Since requiredEnvVars only drives a build-time warning and never gates anything, listing AWS_REGION told every self-hoster running the direct API to set a variable that has no effect for them. The description was also still promising Sonnet 4.6 via Bedrock specifically, which is no longer what the extension does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> * fix(ai): read documentKind defensively in the smoke test The field arrived with the receipt-aware extraction work, so referencing it directly stops the script compiling against any checkout from before that landed. tsconfig includes **/*.ts and next.config does not disable type checking, so on such a checkout this failed the production build rather than just the script: caught while preparing a test branch for a self-hosted instance that had not synced yet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> * fix(deps): restore the nested @swc/helpers entry in the lockfile Declaring @anthropic-ai/sdk with `npm install --package-lock-only` also pruned node_modules/next-intl/node_modules/@swc/helpers@0.5.23, an optional peer entry the local npm 11 considers redundant and the image's npm 10.9.8 does not. The result passed every local check and failed `npm ci` inside the Docker build, which is the only place the lockfile is actually enforced. The lockfile is now the previous one plus the single root dependency line, verified with `npm ci --dry-run`. @anthropic-ai/sdk needed nothing else: it was already in the tree as a transitive dependency of @anthropic-ai/bedrock-sdk. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> * Update DECISIONS.md Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> * Update Docker documentation for AI provider credentials Clarify the role of credentials in AI provider selection and document extraction requirements. * Update SELF-HOSTING.md with smoke-ai script details Clarify usage of smoke-ai script for credential checks and document extraction. * Improve error handling and logging in smoke-ai script * fix(ai): complete plain-key self-hosting path Signed-off-by: Emil <emilmattsson14@gmail.com> --------- Signed-off-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> Signed-off-by: Emil <emilmattsson14@gmail.com> Co-authored-by: Bjorn Bergenheim <29535152+bjornbergenheim@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
211 lines
7.1 KiB
TypeScript
211 lines
7.1 KiB
TypeScript
/**
|
|
* The pairs the arithmetic cannot settle.
|
|
*
|
|
* A weighted formula is the right instrument for the clear cases: it is free,
|
|
* instant, reproducible years later for an audit, and it cannot invent a
|
|
* merchant. It is a poor instrument for the middle. Its weights are chosen by
|
|
* hand, and a pair agreeing to within 1% from a recognised merchant can still
|
|
* land at 0.62 because a date drifted, which says more about the constants than
|
|
* about the receipt.
|
|
*
|
|
* So the formula keeps what it is good at and hands over what it is not. Only
|
|
* the uncertain band is sent here, which on a real ledger was eight pairs
|
|
* against five it had already settled.
|
|
*
|
|
* The question asked is deliberately a yes or no with a reason, never a score.
|
|
* An earlier design in this feature asked the model to rate its own certainty
|
|
* and it anchored on round numbers, which the calibration literature predicts:
|
|
* verbalised confidence is badly calibrated and barely separates a model's
|
|
* right answers from its wrong ones. Judging concrete evidence and explaining
|
|
* the judgement is a different task, and one it is good at.
|
|
*
|
|
* Nothing here books anything. An accepted pair becomes the same proposal a
|
|
* human approves, carrying the reason so the approval is checking an argument
|
|
* rather than trusting a verdict.
|
|
*/
|
|
import { z } from 'zod'
|
|
import { createAiClient, toProviderModelId, type AiClient } from '@/lib/ai/provider'
|
|
import { createLogger } from '@/lib/logger'
|
|
|
|
const log = createLogger('receipt-hunt-adjudicate')
|
|
|
|
const MODEL = toProviderModelId(
|
|
process.env.RECEIPT_HUNT_MODEL_ID ||
|
|
process.env.BEDROCK_MODEL_ID ||
|
|
'claude-sonnet-5'
|
|
)
|
|
|
|
export interface UncertainPair {
|
|
/** Stable handle for this pair, opaque to the model beyond matching it back. */
|
|
key: string
|
|
purchase: {
|
|
description: string
|
|
amount: number
|
|
currency: string
|
|
date: string
|
|
}
|
|
receipt: {
|
|
vendor: string | null
|
|
total: number | null
|
|
currency: string | null
|
|
/** The total in kronor when a rate was resolved, so both sides compare. */
|
|
sekTotal?: number | null
|
|
date: string | null
|
|
fileName: string | null
|
|
}
|
|
/** What the formula made of it, as context rather than as an instruction. */
|
|
confidence: number
|
|
matchReasons: string[]
|
|
}
|
|
|
|
export interface Verdict {
|
|
key: string
|
|
accept: boolean
|
|
reason: string
|
|
}
|
|
|
|
const VerdictSchema = z.object({
|
|
verdicts: z.preprocess(
|
|
(v) => {
|
|
if (typeof v !== 'string') return v
|
|
try {
|
|
return JSON.parse(v)
|
|
} catch {
|
|
return v
|
|
}
|
|
},
|
|
z
|
|
.array(
|
|
z.object({
|
|
key: z.string().min(1),
|
|
accept: z.coerce.boolean(),
|
|
reason: z.string().min(1).max(300),
|
|
}),
|
|
)
|
|
.default([]),
|
|
),
|
|
})
|
|
|
|
const TOOL = {
|
|
type: 'object',
|
|
properties: {
|
|
verdicts: {
|
|
type: 'array',
|
|
items: {
|
|
type: 'object',
|
|
properties: {
|
|
key: { type: 'string', description: 'The pair key exactly as given.' },
|
|
accept: {
|
|
type: 'boolean',
|
|
description: 'True only if this document is the underlag for this purchase.',
|
|
},
|
|
reason: { type: 'string', description: 'One short sentence, in Swedish.' },
|
|
},
|
|
required: ['key', 'accept', 'reason'],
|
|
},
|
|
},
|
|
},
|
|
required: ['verdicts'],
|
|
}
|
|
|
|
const SYSTEM = `Du avgör om en handling hör till ett visst köp.
|
|
|
|
Du får par som en beräkning inte kunde avgöra själv. För varje par: är den här
|
|
handlingen underlaget för det här köpet? Svara ja eller nej och säg varför.
|
|
|
|
Det här gör paren svåra, och inget av det är i sig skäl att säga nej:
|
|
|
|
- Datumen glider. Ett kortköp bokförs hos banken dagar efter att det gjordes,
|
|
utrikes gärna en vecka, och ett vidarebefordrat kvitto bär köpets datum medan
|
|
kontoutdraget bär bokföringsdagen.
|
|
- Beloppen kan skilja någon procent när kvittot är i annan valuta. Banken drog
|
|
ett omräknat belopp till sin egen kurs; vi har räknat om till Riksbankens.
|
|
Ett par procents skillnad är växelkursen, inte olika belopp.
|
|
- Bankens text är inte ett handlarnamn. Den är avhuggen och innehåller
|
|
betalvägar: "ANTHROPIC* CLAUDE SUB" och "Anthropic, PBC" är samma leverantör.
|
|
|
|
Säg nej när något faktiskt talar emot: fel storleksordning på beloppet, en
|
|
handling som avser en annan period, eller en handlare som inte rimligen är
|
|
samma. Säg nej också när du helt enkelt inte kan avgöra det: en människa läser
|
|
ditt skäl och ett vagt ja kostar mer än ett ärligt nej.
|
|
|
|
Beräkningens poäng och skäl finns med som bakgrund. Den har redan vägt in
|
|
belopp, handlare och datum, så håll dig inte till den: du ser saker den inte
|
|
kan väga.
|
|
|
|
reason: en kort mening på svenska om varför paret hör ihop eller inte.`
|
|
|
|
function client(): AiClient {
|
|
return createAiClient()
|
|
}
|
|
|
|
/**
|
|
* Settle the pairs the formula could not.
|
|
*
|
|
* One call for the batch: the pairs are independent, but a run holds a handful
|
|
* of them and a call each would be latency for nothing.
|
|
*
|
|
* Returns only the pairs it was given, and only accepted ones. A failed call
|
|
* accepts nothing, which leaves the run exactly where the arithmetic left it.
|
|
*/
|
|
export async function adjudicate(pairs: readonly UncertainPair[]): Promise<Verdict[]> {
|
|
if (pairs.length === 0) return []
|
|
|
|
const known = new Set(pairs.map((p) => p.key))
|
|
const payload = {
|
|
pairs: pairs.map((p) => ({
|
|
key: p.key,
|
|
kop: {
|
|
text: p.purchase.description,
|
|
belopp: p.purchase.amount,
|
|
valuta: p.purchase.currency,
|
|
datum: p.purchase.date,
|
|
},
|
|
handling: {
|
|
leverantor: p.receipt.vendor,
|
|
belopp: p.receipt.total,
|
|
valuta: p.receipt.currency,
|
|
belopp_i_kronor: p.receipt.sekTotal ?? null,
|
|
datum: p.receipt.date,
|
|
fil: p.receipt.fileName,
|
|
},
|
|
berakningen_sa: { poang: p.confidence, skal: p.matchReasons },
|
|
})),
|
|
}
|
|
|
|
try {
|
|
const response = await client().messages.create({
|
|
model: MODEL,
|
|
max_tokens: 4096,
|
|
system: SYSTEM,
|
|
tools: [{ name: 'verdicts', description: 'Return the result in this exact shape.', input_schema: TOOL as never }],
|
|
tool_choice: { type: 'tool', name: 'verdicts' },
|
|
messages: [{ role: 'user', content: JSON.stringify(payload, null, 1) }],
|
|
})
|
|
|
|
const block = response.content.find((c) => c.type === 'tool_use')
|
|
if (!block || block.type !== 'tool_use') throw new Error('model did not use the tool')
|
|
const parsed = VerdictSchema.parse(block.input)
|
|
|
|
const seen = new Set<string>()
|
|
const out: Verdict[] = []
|
|
for (const v of parsed.verdicts) {
|
|
// Only pairs we asked about, and each answered once: a key we never sent
|
|
// would attach a document to a purchase nobody weighed.
|
|
if (!known.has(v.key) || seen.has(v.key)) continue
|
|
seen.add(v.key)
|
|
if (!v.accept) continue
|
|
out.push({ key: v.key, accept: true, reason: v.reason })
|
|
}
|
|
|
|
log.info('adjudicated uncertain pairs', { asked: pairs.length, accepted: out.length })
|
|
return out
|
|
} catch (error) {
|
|
log.warn('adjudication failed, proposing none of the uncertain pairs', {
|
|
pairs: pairs.length,
|
|
error: error instanceof Error ? error.message : String(error),
|
|
})
|
|
return []
|
|
}
|
|
}
|