* fix(mcp): make agent failures diagnosable from the log, not just from the source Mining event_log for mcp.tool_called: VALIDATION_ERROR ran at about one a day until 2026-08-25, then jumped to a hundred a day and stayed there. One integration's gnubok_get_kpi_report has been refused 604 times over seven days and is still failing. The cause is the unknown-parameter guard from #1856, which is correct and stays. The caller is even told precisely what is wrong: getStructuredError puts "Unknown parameter "x" for <tool>. Valid parameters: ..." into message_en, which the agent receives. What was broken is what we recorded about it. Two things, both cheap: - Telemetry logged only message_sv, which for VALIDATION_ERROR is the registry default "Förfrågan innehåller ogiltiga uppgifter." and names neither the parameter nor the tool. A seven-day outage was indistinguishable from a typo, and the cause was only findable by reading the dispatcher source. errorDetail now carries message_en, stored only when it differs from errorMessage, so the many domain failures whose message_sv is already the specific text cost nothing extra. - errorKind said 'company_access_denied' for all 604, because the arg guard throws inside the company-routing try. That is an active misdirection: it sends triage looking for a tenancy bug that does not exist. Exactly two things in that block raise VALIDATION_ERROR, the unknown-parameter guard and a malformed company_id, and neither is a permissions failure, so they now log as 'invalid_arguments'. Tests reproduce the production call shape and were verified to fail when either fix is reverted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): derive the telemetry payload type from the event contract Two review findings, both valid. The local ToolCalledPayload interface was a hand-maintained duplicate that omitted sessionId and half the errorKind union. That is not cosmetic: it is why errorDetail could be added to the emitter and to lib/events/types.ts while the test file type-checked against a stale shape. Deriving it from EventPayload<'mcp.tool_called'> removes the drift. The null-detail assertion was also conditional, so it would have passed on a wrong-but-present value. Replaced with two determinate cases: a scope denial carries both languages without duplicating either, and the unknown-tool exit, which supplies no diagnostic, stores null. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): pin the scope name in the diagnostic assertion 'A different string' passes on any placeholder. The reason errorDetail is worth storing is that it names what the caller lacks, so assert that. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Accounted MCP server
JSON-RPC 2.0 server exposing the Accounted bookkeeping engine to MCP clients (Claude Desktop, Claude Code, etc.). Endpoint: /api/extensions/ext/mcp-server/mcp. Add ?tool_namespace=accounted for the Accounted tool names. Requests without it retain the legacy Gnubok namespace. OAuth and stdio bridges live alongside the API surface: see app/api/mcp-oauth/, packages/accounted-mcp/, and the compatibility package in packages/gnubok-mcp/.
Tool authoring contract
Enforced by tests in __tests__/: these are not style preferences, they're guard rails.
additionalProperties: falseon everyinputSchema. Guarded bystrict-schemas.test.ts. Forces clear rejections on hallucinated fields instead of silent ignores.- Descriptions ≤ 280 chars. Guarded by
output-schema.test.ts. NoArgs:/Returns:/Examples:prose: those belong in JSON Schema. Use agent-native hints ("Use to…", "Call X first", "HIGH risk"). - Staged-operation envelope for write tools:
outputSchema: STAGED_OPERATION_SCHEMA(server.ts). Fields:staged, risk_level, actor, message, preview, period_status?, next?. Thestaged: trueboolean is the explicit completion signal; agents must not infer completion from prose. Do NOT introduce a parallel{ success, shouldContinue, output }envelope. period_statusthreading: any tool that ties to a fiscal-period-bound date (categorize, mark paid, create voucher, correct/reverse entry, approve supplier invoice) passesdateForPeriodChecktostagePendingOperation. Response then includesperiod_status: { period_id, status: open|locked|closed, lock_date }so widgets and agents disable writes without round-trips.- Scope mapping: every new tool needs an entry in
lib/auth/api-keys.tsTOOL_SCOPE_MAP. Missing entries default to deny. - Tests for new write tools: add staging-gate coverage to
__tests__/voucher-tools.test.ts(or a sibling) plus executor coverage tolib/pending-operations/__tests__/voucher-executors.test.tsif the tool stages a newoperation_type.
Determinism / cache stability
Tool definitions (name, description, inputSchema, outputSchema, annotations) are declared as static object literals at module load: no timestamps, no UUIDs, no Date/Math.random in the definition layer. This makes the tools/list JSON payload byte-stable across requests, which lets agent-side prompt caches stay warm. Do not introduce per-request non-determinism into the definitions block. Anything time-bound or random belongs inside execute().
For internal Anthropic API usage (the SDK is called from lib/ai/provider.ts and lib/ai/services/anthropic-family.ts; features such as extensions/general/invoice-inbox/lib/extract-invoice-fields.ts go through getAiService() in lib/ai rather than the SDK directly): annotate stable prefixes with cache_control: { type: 'ephemeral' } and log usage.cache_read_input_tokens for hit-ratio observability. The 1h TTL from the agent-native API plan (item 10) requires the direct Anthropic API; Accounted's Bedrock path defaults to a shorter TTL.
Payload-size watchdog
payload-size.bench.test.ts enforces a tools/list JSON payload ceiling. If the test fires, the right answer is rarely "raise the ceiling". Instead, trim descriptions or set specialized wide tools to catalogVisibility: 'search'. Those tools remain discoverable with full schemas through gnubok_search_tools and callable through tools/call on the wire without bloating the default catalog. Claude.ai only calls tools present in tools/list, so a tool that a user or a skill must call directly stays in the default catalog.
Where things live
server.ts: the tools array + JSON-RPC dispatchertool-result.ts:withNext(),toToolError()response helpersresources/: read-onlyAccounted://URIs, registered inresources/index.ts:company/current,period/active,recent-activity,capabilities,attention,chart-of-accounts,settings/vat-treatments,booking-templates,ledger/context,reconciliation/summarywidgets/: inline HTML widgets (receipt-matcher, vat-review, pending-operations)prompts/: slash-command-style promptsskills/: domain-knowledge skill bodies served viagnubok_load_skillpublic-tools.ts: lazy authentication (issue #1814).ANONYMOUS_METHODS(initialize, ping, tools/prompts/resources listing) and the threePUBLIC_TOOLS(gnubok_search_tools,gnubok_list_skills,gnubok_load_skill) answer without credentials, rate-limited per truncated IP; every othertools/callgets a transport-level 401 +WWW-AuthenticatefromhandleMcpRequestinserver.ts, which is the challenge clients turn into their Connect prompttasks.ts: MCP Tasks extension (io.modelcontextprotocol/tasks): durable handles for long-running tool calls, rows inmcp_tasks(service-role writes only)origin-guard.ts: Origin-header validation on the Streamable HTTP endpoint (DNS-rebinding defence required by the MCP spec)staging-pii-guard.ts: refuses a plaintext personnummer in stagedpending_operationsparams/preview, so every staging tool inherits the encrypt-at-staging ruletool-namespace.ts:?tool_namespace=accountedhandling (accounted_*aliases for the canonicalgnubok_*ids)__tests__/: strictness guards + per-tool coverage