Files
accounted/extensions/general/mcp-server
cce0de5704 feat(mcp): already-explained voucher guard at stage and commit for match_batch_allocate (#2294) (#2346)
* feat(mcp): already-explained voucher guard at stage and commit for match_batch_allocate

The dashboard match-batch route refused BATCH_TX_POSSIBLE_DUPLICATE when
posted, unlinked vouchers already summed to the bank row (PR #2300), but the
MCP door (gnubok_match_batch_allocate staging + commitMatchBatchAllocate)
called the RPC with no guard, so an agent could book a Bankgirot aggregate a
second time. The detector existed once; the guard lived in one door.

One shared decision helper, lib/invoices/already-explained-guard.ts, now
sits on top of the existing detectors (no fork) and is called by the
dashboard route, the MCP staging tools and the commit executors:

- gnubok_match_batch_allocate refuses to stage, coded
  BATCH_TX_POSSIBLE_DUPLICATE, naming the vouchers, the reconcile_match /
  link_transaction_to_journal_entry call that resolves the row, and the
  exact force + expected_journal_entry_ids binding.
- commitMatchBatchAllocate runs the same guard before the RPC and
  re-validates a staged force binding against the set detected at commit,
  so a stale approval cannot book a duplicate; 409 auto-rejects with the
  vouchers in result_data.
- force + expected_journal_entry_ids on the tool mirror MatchBatchSchema;
  an honoured override stages with a compliance_warning and, after the
  booking succeeds, writes BankTransactionDuplicateDismissed to
  behandlingshistorik (dashboard route included; it only logged before).
- gnubok_match_transaction_to_invoice and commitMatchTransactionInvoice get
  the dashboard's 1:1 soft-duplicate guard (MATCH_INVOICE_POSSIBLE_DUPLICATE
  / MATCH_INVOICE_FORCE_CANDIDATE_MISMATCH) with force +
  expected_journal_entry_id; at commit it runs before the storno.
- Registry: both duplicate codes gain retryable: false and a remediation.

Catalog payload held under the 60K ceiling by trimming the two tools' own
descriptions (59 988 measured, ledger entry in payload-size.bench.test.ts).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SaJfqNi4VmsG8FMKq99G6

* docs(decisions): record the 2026-09-06 ten-issue batch's first-principles choices

Carries the DECISIONS.md lines for PRs #2337 #2339 #2340 #2341 #2342 #2343 #2344 #2345 #2346 #2347 in one place so the ten branches do not conflict on this file.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SaJfqNi4VmsG8FMKq99G6

* fix(mcp): refuse an unverifiable forced override, surface a failed duplicate check, validate the binding (#2294 review)

Review round on PR #2346 (CodeRabbit + compliance):

- guardAlreadyExplained returned 'clear' when the detector threw even with
  force=true, so a forced 1:N override could book without re-validating
  expected_journal_entry_ids and left no behandlingshistorik record. It now
  returns a distinct 'unverifiable' outcome under force (mirrors
  guardDuplicatePaymentVoucher); the dashboard route, the MCP staging tool
  and the commit executor all refuse it with the new registry code
  BATCH_TX_EXPLAINED_CHECK_FAILED (409, retryable, remediation). Regression
  tests on every caller.
- A detector failure without force still fails open at stage time, but no
  longer silently: the tools track onDetectError and stage a
  complianceNote, so preview_data.compliance_warning is set on both
  match_batch_allocate (GenericPreview renders it) and
  match_transaction_invoice (MatchTransactionInvoicePreview now renders
  data.compliance_warning through AttnLine).
- expected_journal_entry_ids / expected_journal_entry_id are validated at
  the MCP boundary (array of 1 to 10 non-empty strings / non-empty string)
  and refused with VALIDATION_ERROR instead of being silently filtered.
  No schema description text added: catalog payload unchanged.
- RoPA: .compliance/ropa.yaml gains bookkeeping.duplicate_dismissal_history
  for the BankTransactionDuplicateDismissed record (Art. 6(1)(c), BFNAR
  2013:2 p. 9.16, retention per BFL 7 kap, stored in processing_history).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:56:36 +02:00
..
2026-07-24 15:03:50 +02:00
2026-07-24 15:03:50 +02:00

Accounted MCP server

JSON-RPC 2.0 server exposing the Accounted bookkeeping engine to MCP clients (Claude Desktop, Claude Code, etc.). Endpoint: /api/extensions/ext/mcp-server/mcp. Add ?tool_namespace=accounted for the Accounted tool names. Requests without it retain the legacy Gnubok namespace. OAuth and stdio bridges live alongside the API surface: see app/api/mcp-oauth/, packages/accounted-mcp/, and the compatibility package in packages/gnubok-mcp/.

Tool authoring contract

Enforced by tests in __tests__/: these are not style preferences, they're guard rails.

  1. additionalProperties: false on every inputSchema. Guarded by strict-schemas.test.ts. Forces clear rejections on hallucinated fields instead of silent ignores.
  2. Descriptions ≤ 280 chars. Guarded by output-schema.test.ts. No Args: / Returns: / Examples: prose: those belong in JSON Schema. Use agent-native hints ("Use to…", "Call X first", "HIGH risk").
  3. Staged-operation envelope for write tools: outputSchema: STAGED_OPERATION_SCHEMA (server.ts). Fields: staged, risk_level, actor, message, preview, period_status?, next?. The staged: true boolean is the explicit completion signal; agents must not infer completion from prose. Do NOT introduce a parallel { success, shouldContinue, output } envelope.
  4. period_status threading: any tool that ties to a fiscal-period-bound date (categorize, mark paid, create voucher, correct/reverse entry, approve supplier invoice) passes dateForPeriodCheck to stagePendingOperation. Response then includes period_status: { period_id, status: open|locked|closed, lock_date } so widgets and agents disable writes without round-trips.
  5. Scope mapping: every new tool needs an entry in lib/auth/api-keys.ts TOOL_SCOPE_MAP. Missing entries default to deny.
  6. Tests for new write tools: add staging-gate coverage to __tests__/voucher-tools.test.ts (or a sibling) plus executor coverage to lib/pending-operations/__tests__/voucher-executors.test.ts if the tool stages a new operation_type.

Determinism / cache stability

Tool definitions (name, description, inputSchema, outputSchema, annotations) are declared as static object literals at module load: no timestamps, no UUIDs, no Date/Math.random in the definition layer. This makes the tools/list JSON payload byte-stable across requests, which lets agent-side prompt caches stay warm. Do not introduce per-request non-determinism into the definitions block. Anything time-bound or random belongs inside execute().

For internal Anthropic API usage (the SDK is called from lib/ai/provider.ts and lib/ai/services/anthropic-family.ts; features such as extensions/general/invoice-inbox/lib/extract-invoice-fields.ts go through getAiService() in lib/ai rather than the SDK directly): annotate stable prefixes with cache_control: { type: 'ephemeral' } and log usage.cache_read_input_tokens for hit-ratio observability. The 1h TTL from the agent-native API plan (item 10) requires the direct Anthropic API; Accounted's Bedrock path defaults to a shorter TTL.

Payload-size watchdog

payload-size.bench.test.ts enforces a tools/list JSON payload ceiling. If the test fires, the right answer is rarely "raise the ceiling". Instead, trim descriptions or set specialized wide tools to catalogVisibility: 'search'. Those tools remain discoverable with full schemas through gnubok_search_tools and callable through tools/call on the wire without bloating the default catalog. Claude.ai only calls tools present in tools/list, so a tool that a user or a skill must call directly stays in the default catalog.

Where things live

  • server.ts: the tools array + JSON-RPC dispatcher
  • tool-result.ts: withNext(), toToolError() response helpers
  • resources/: read-only Accounted:// URIs, registered in resources/index.ts: company/current, period/active, recent-activity, capabilities, attention, chart-of-accounts, settings/vat-treatments, booking-templates, ledger/context, reconciliation/summary
  • widgets/: inline HTML widgets (receipt-matcher, vat-review, pending-operations)
  • prompts/: slash-command-style prompts
  • skills/: domain-knowledge skill bodies served via gnubok_load_skill
  • public-tools.ts: lazy authentication (issue #1814). ANONYMOUS_METHODS (initialize, ping, tools/prompts/resources listing) and the three PUBLIC_TOOLS (gnubok_search_tools, gnubok_list_skills, gnubok_load_skill) answer without credentials, rate-limited per truncated IP; every other tools/call gets a transport-level 401 + WWW-Authenticate from handleMcpRequest in server.ts, which is the challenge clients turn into their Connect prompt
  • tasks.ts: MCP Tasks extension (io.modelcontextprotocol/tasks): durable handles for long-running tool calls, rows in mcp_tasks (service-role writes only)
  • origin-guard.ts: Origin-header validation on the Streamable HTTP endpoint (DNS-rebinding defence required by the MCP spec)
  • staging-pii-guard.ts: refuses a plaintext personnummer in staged pending_operations params/preview, so every staging tool inherits the encrypt-at-staging rule
  • tool-namespace.ts: ?tool_namespace=accounted handling (accounted_* aliases for the canonical gnubok_* ids)
  • __tests__/: strictness guards + per-tool coverage