Files
accounted/extensions/general/mcp-server
Jakob WennbergandClaude Fable 5 4501f118c2 feat(mcp): approval-queue MCP Apps widget for staged operations (#1278)
* feat(mcp): approval-queue MCP Apps widget for staged operations

gnubok_list_pending_operations(render_ui=true) now renders an interactive
approval queue (claude.ai / Claude Desktop) where the user approves or
rejects each staged operation with a click. High-risk operations arm the
approve button and the second click sends confirmed=true, so the BFL
5 kap 5 acknowledgment is a first-party human action instead of the
agent asserting confirmed=true on the user's behalf (the audit weakness
flagged in dev_docs/erpclaw_analysis.md).

- New widget ui://pending-operations/app.html following the established
  self-contained postMessage/JSON-RPC pattern (no fetch, theme-aware,
  Swedish labels, expandable preview_data per row).
- Result-level _meta.ui hint gated on render_ui=true, mirroring the VAT
  report wiring; the tool stays data-only by default.
- Widget tool references project per namespace (accounted_* clients see
  accounted_ names inside the HTML).
- tools/list payload ceiling 58K -> 58.5K per the in-test convention:
  prose trimmed to the floor first, remainder is wire contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(mcp): time out the widget RPC bridge so a silent host cannot strand a row

Review follow-up: sendRequest never settled if the host dropped a
response, leaving op._working=true forever with the approve/reject
buttons gone. A 30s timeout rejects the promise; the existing catch
paths restore the row with an error message so the user can retry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:02:17 +02:00
..
2026-07-15 15:53:15 +02:00
2026-07-24 15:03:50 +02:00
2026-07-24 15:03:50 +02:00
2026-07-24 15:03:50 +02:00

Accounted MCP server

JSON-RPC 2.0 server exposing the Accounted bookkeeping engine to MCP clients (Claude Desktop, Claude Code, etc.). Endpoint: /api/extensions/ext/mcp-server/mcp. Add ?tool_namespace=accounted for the Accounted tool names. Requests without it retain the legacy Gnubok namespace. OAuth and stdio bridges live alongside the API surface: see app/api/mcp-oauth/, packages/accounted-mcp/, and the compatibility package in packages/gnubok-mcp/.

Tool authoring contract

Enforced by tests in __tests__/: these are not style preferences, they're guard rails.

  1. additionalProperties: false on every inputSchema. Guarded by strict-schemas.test.ts. Forces clear rejections on hallucinated fields instead of silent ignores.
  2. Descriptions ≤ 280 chars. Guarded by output-schema.test.ts. No Args: / Returns: / Examples: prose: those belong in JSON Schema. Use agent-native hints ("Use to…", "Call X first", "HIGH risk").
  3. Staged-operation envelope for write tools: outputSchema: STAGED_OPERATION_SCHEMA (server.ts). Fields: staged, risk_level, actor, message, preview, period_status?, next?. The staged: true boolean is the explicit completion signal; agents must not infer completion from prose. Do NOT introduce a parallel { success, shouldContinue, output } envelope.
  4. period_status threading: any tool that ties to a fiscal-period-bound date (categorize, mark paid, create voucher, correct/reverse entry, approve supplier invoice) passes dateForPeriodCheck to stagePendingOperation. Response then includes period_status: { period_id, status: open|locked|closed, lock_date } so widgets and agents disable writes without round-trips.
  5. Scope mapping: every new tool needs an entry in lib/auth/api-keys.ts TOOL_SCOPE_MAP. Missing entries default to deny.
  6. Tests for new write tools: add staging-gate coverage to __tests__/voucher-tools.test.ts (or a sibling) plus executor coverage to lib/pending-operations/__tests__/voucher-executors.test.ts if the tool stages a new operation_type.

Determinism / cache stability

Tool definitions (name, description, inputSchema, outputSchema, annotations) are declared as static object literals at module load: no timestamps, no UUIDs, no Date/Math.random in the definition layer. This makes the tools/list JSON payload byte-stable across requests, which lets agent-side prompt caches stay warm. Do not introduce per-request non-determinism into the definitions block. Anything time-bound or random belongs inside execute().

For internal Anthropic API usage (today only extensions/general/invoice-inbox/lib/extract-invoice-fields.ts): annotate stable prefixes with cache_control: { type: 'ephemeral' } and log usage.cache_read_input_tokens for hit-ratio observability. The 1h TTL from the agent-native API plan (item 10) requires the direct Anthropic API; Accounted's Bedrock path defaults to a shorter TTL.

Payload-size watchdog

payload-size.bench.test.ts enforces a tools/list JSON payload ceiling. If the test fires, the right answer is rarely "raise the ceiling". Instead, trim descriptions or set specialized wide tools to catalogVisibility: 'search'. Those tools remain discoverable with full schemas through gnubok_search_tools and callable through tools/call without bloating the default catalog.

Where things live

  • server.ts: the tools array + JSON-RPC dispatcher
  • tool-result.ts: withNext(), toToolError() response helpers
  • resources/: read-only Accounted:// URIs (active company, period, recent activity, capabilities, attention items, voucher gaps, chart of accounts, VAT treatments)
  • widgets/: inline HTML widgets (receipt-matcher, vat-review)
  • prompts/: slash-command-style prompts
  • skills/: domain-knowledge skill bodies served via gnubok_load_skill
  • __tests__/: strictness guards + per-tool coverage