Files
accounted/extensions/general/mcp-server
272d19b287 fix(supplier-invoices): duplicate-payment guard matches abbreviated bank text and shares one detector with the customer side (#2299) (#2345)
* fix(supplier-invoices): duplicate-payment guard matches abbreviated bank text and shares one detector with the customer side

The mark-paid guard probed merchant_name for the FULL supplier name, so the
row that paid Hi3G Access AB (bank text "HI3G", merchant_name empty) never
matched and the payment was booked twice (#2299).

- counterpartyNeedle(): first distinctive token of the name (alnum, legal
  forms dropped, >= 2 chars so initialisms like SJ and 3M survive), probed on
  merchant_name OR description in one .or() per currency sweep; the alnum
  shape is what makes the DSL interpolation safe.
- findDuplicatePaymentCandidatesForSupplierInvoice() beside the customer
  detector; both share the sweep and the scorer. The dashboard route's inline
  copy is deleted; the v1 supplier mark-paid door gets the guard it lacked.
- New match_reason already_booked (row already carries a verifikat, booked
  straight from the bank side): ranked first, carries journal_entry_id, and
  the dialogs, MCP path and pending-operation commit word the remedy as a
  rattelse rather than "link it".
- Customer side gets the same token prefilter and classification.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SaJfqNi4VmsG8FMKq99G6

* test(invoices): align customer mark-paid queued mocks with the one-probe duplicate guard

The customer detector now issues one .or() counterparty probe per currency
sweep instead of two ILIKE queries, so every queued answer after the guard
was consumed one step early: the aggregate-sweep [] became company_settings,
the settings row hit the entry builder, and two tests saw 500 / the wrong
voucher id. Each guard block now enqueues one probe plus the aggregate sweep;
the 409 tests drop the second-probe entry that is no longer read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SaJfqNi4VmsG8FMKq99G6

* fix(invoices): one logic expression per duplicate-payment sweep, never two or= params

The sweep chain carried two .or() calls (currency clause, then name probe).
postgrest-js appends a query parameter per call, so the client sent or=
twice, and whether PostgREST ANDs a repeated key was never proven in this
repo; had it kept one, the currency predicate would be gone and foreign rows
banded against a kronor figure.

counterpartySweepLogic() now nests both groups under one and() inside a
single top-level or(): and(or(<currency>),or(merchant_name.ilike.*x*,
description.ilike.*x*)). The sweep issues exactly one .or() per currency.

Proof at three levels: unit tests pin the helper's string; a fake-fetch test
runs the real postgrest-js builder and asserts exactly one or= search param
per request; a tool-pg test seeds right-currency+hit, wrong-currency+hit
(with an amount_sek that would pass every JS check) and right-currency+miss
rows against a real PostgREST and asserts, for both detectors and both
sweeps, that only the first comes back, from PostgREST's own response.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019SaJfqNi4VmsG8FMKq99G6

* fix(invoices): name storno as the already_booked remedy, never "makulera"

A posted verifikat is never deleted; it is corrected by a storno entry
(BFL 5 kap 5 §). The already_booked remedy text in the error catalogue, the
MCP and pending-operation messages and both UI descriptions now say so:
"vänd en av verifikationerna med storno och koppla underlaget till den som
blir kvar" / "reverse one of the two vouchers with a storno entry and attach
the underlag to the remaining one".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:53:34 +02:00
..
2026-07-24 15:03:50 +02:00
2026-07-24 15:03:50 +02:00

Accounted MCP server

JSON-RPC 2.0 server exposing the Accounted bookkeeping engine to MCP clients (Claude Desktop, Claude Code, etc.). Endpoint: /api/extensions/ext/mcp-server/mcp. Add ?tool_namespace=accounted for the Accounted tool names. Requests without it retain the legacy Gnubok namespace. OAuth and stdio bridges live alongside the API surface: see app/api/mcp-oauth/, packages/accounted-mcp/, and the compatibility package in packages/gnubok-mcp/.

Tool authoring contract

Enforced by tests in __tests__/: these are not style preferences, they're guard rails.

  1. additionalProperties: false on every inputSchema. Guarded by strict-schemas.test.ts. Forces clear rejections on hallucinated fields instead of silent ignores.
  2. Descriptions ≤ 280 chars. Guarded by output-schema.test.ts. No Args: / Returns: / Examples: prose: those belong in JSON Schema. Use agent-native hints ("Use to…", "Call X first", "HIGH risk").
  3. Staged-operation envelope for write tools: outputSchema: STAGED_OPERATION_SCHEMA (server.ts). Fields: staged, risk_level, actor, message, preview, period_status?, next?. The staged: true boolean is the explicit completion signal; agents must not infer completion from prose. Do NOT introduce a parallel { success, shouldContinue, output } envelope.
  4. period_status threading: any tool that ties to a fiscal-period-bound date (categorize, mark paid, create voucher, correct/reverse entry, approve supplier invoice) passes dateForPeriodCheck to stagePendingOperation. Response then includes period_status: { period_id, status: open|locked|closed, lock_date } so widgets and agents disable writes without round-trips.
  5. Scope mapping: every new tool needs an entry in lib/auth/api-keys.ts TOOL_SCOPE_MAP. Missing entries default to deny.
  6. Tests for new write tools: add staging-gate coverage to __tests__/voucher-tools.test.ts (or a sibling) plus executor coverage to lib/pending-operations/__tests__/voucher-executors.test.ts if the tool stages a new operation_type.

Determinism / cache stability

Tool definitions (name, description, inputSchema, outputSchema, annotations) are declared as static object literals at module load: no timestamps, no UUIDs, no Date/Math.random in the definition layer. This makes the tools/list JSON payload byte-stable across requests, which lets agent-side prompt caches stay warm. Do not introduce per-request non-determinism into the definitions block. Anything time-bound or random belongs inside execute().

For internal Anthropic API usage (the SDK is called from lib/ai/provider.ts and lib/ai/services/anthropic-family.ts; features such as extensions/general/invoice-inbox/lib/extract-invoice-fields.ts go through getAiService() in lib/ai rather than the SDK directly): annotate stable prefixes with cache_control: { type: 'ephemeral' } and log usage.cache_read_input_tokens for hit-ratio observability. The 1h TTL from the agent-native API plan (item 10) requires the direct Anthropic API; Accounted's Bedrock path defaults to a shorter TTL.

Payload-size watchdog

payload-size.bench.test.ts enforces a tools/list JSON payload ceiling. If the test fires, the right answer is rarely "raise the ceiling". Instead, trim descriptions or set specialized wide tools to catalogVisibility: 'search'. Those tools remain discoverable with full schemas through gnubok_search_tools and callable through tools/call on the wire without bloating the default catalog. Claude.ai only calls tools present in tools/list, so a tool that a user or a skill must call directly stays in the default catalog.

Where things live

  • server.ts: the tools array + JSON-RPC dispatcher
  • tool-result.ts: withNext(), toToolError() response helpers
  • resources/: read-only Accounted:// URIs, registered in resources/index.ts: company/current, period/active, recent-activity, capabilities, attention, chart-of-accounts, settings/vat-treatments, booking-templates, ledger/context, reconciliation/summary
  • widgets/: inline HTML widgets (receipt-matcher, vat-review, pending-operations)
  • prompts/: slash-command-style prompts
  • skills/: domain-knowledge skill bodies served via gnubok_load_skill
  • public-tools.ts: lazy authentication (issue #1814). ANONYMOUS_METHODS (initialize, ping, tools/prompts/resources listing) and the three PUBLIC_TOOLS (gnubok_search_tools, gnubok_list_skills, gnubok_load_skill) answer without credentials, rate-limited per truncated IP; every other tools/call gets a transport-level 401 + WWW-Authenticate from handleMcpRequest in server.ts, which is the challenge clients turn into their Connect prompt
  • tasks.ts: MCP Tasks extension (io.modelcontextprotocol/tasks): durable handles for long-running tool calls, rows in mcp_tasks (service-role writes only)
  • origin-guard.ts: Origin-header validation on the Streamable HTTP endpoint (DNS-rebinding defence required by the MCP spec)
  • staging-pii-guard.ts: refuses a plaintext personnummer in staged pending_operations params/preview, so every staging tool inherits the encrypt-at-staging rule
  • tool-namespace.ts: ?tool_namespace=accounted handling (accounted_* aliases for the canonical gnubok_* ids)
  • __tests__/: strictness guards + per-tool coverage