* feat(api): surface the registry's worked examples in the OpenAPI spec and generated skill EndpointDefinition.example is required and every one of the 125 v1 endpoints populates example.response, but generateOpenApiSpec() never emitted it. The examples reached only the docs markdown builder, so /api/v1/openapi.json carried none and the generated skills/accounted-api had zero json blocks in all 12 reference files: every agent reading the spec or installing the skill got schemas with no concrete body. Emit example on the application/json media types (request body and 200 response) and teach the portable renderOperationMd to print it as a fenced json block. 178 worked examples now reach the skill. SKILL.md is unchanged: the examples land in the on-demand reference files, not the entry file. Attached to JSON media types only, so a multipart body and a binary application/pdf response do not advertise an example they cannot send. Adds the one missing example.request (currency-revaluation) so the new exhaustive coverage assertions hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(api): emit Retry-After on a v1 429 so the documented contract is real The published accounted-api skill has told agents to honor Retry-After on a 429 since it shipped, but no /api/v1 route ever sent one: the wrapper's auth failure path early-returns through v1ErrorResponseFromCode, whose finalize() set only X-Request-Id and Gnubok-Version. Unattended clients had nothing to pace against and had to back off blindly. 60 seconds is an exact upper bound rather than a guess: the rate limiter is a fixed one-minute tumbling window per key row and the limited branch does not slide it. The value moves into an exported constant next to that limiter, so the MCP server's hardcoded '60' now reads from the same place. Also corrects the withApiV1 doc comment, which claimed step 8 stamps X-RateLimit-Limit. It never did. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(mcp): guard the tools/list payload for the namespace new installs get The payload ratchet only ever serialized the gnubok_* projection. The accounted_* projection is inherently larger (every tool reference gains 3 chars, ~209 tokens across the default catalog) and CLAUDE.md points new MCP installs at exactly that namespace, so the payload a new user's client receives was never measured. It had already drifted ~90 tokens past the 63.4K ceiling while the guarded number sat comfortably under it. Measure both and assert on the larger. The ceiling moves to 63.6K to cover the real worst case; this buys no new catalog surface. A second test pins the direction of the delta so Math.max cannot silently stop describing reality. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(mcp): make search-only read tools reachable, and put the payload ceiling into reverse DECISIONS.md records on 2026-08-26 that gnubok_reconcile_match had to be promoted back into the default catalog because "a search-only tool is uncallable on Claude.ai". That is a client-side limit, not a server one: the tools/call dispatcher has always resolved names against the whole tools array, and isDefaultCatalogTool gates only what tools/list shows. So catalogVisibility: 'search' was unusable as a payload lever for reads, and the ceiling could only ever go up. gnubok_call_tool gives such a client one visible name to forward through. It is a rewrite in the dispatcher rather than a forwarding wrapper: {tool, arguments} is rebound to the inner tool BEFORE resolution, so the scope check, unknown-argument guard, company routing, test-key write block, staging _meta and telemetry all apply to the real target instead of being bypassed. Reads only; a write must be named directly so its approval contract stays visible. Alongside it, gnubok_get_agent_briefing's outputSchema drops 7743 to 4565 chars. Four sub-schemas whose interiors were documentation rather than contract are condensed to a permissive object plus a fuller description; agent-briefing.test.ts already pins their runtime shape, so nothing is left unguarded. Net on the guarded (accounted) projection: 63 491 to 62 942 tokens, with the new tool included. The ceiling moves 63.6K DOWN to 63.1K, the first tightening in that ledger, and the note now says to demote a read before proposing a bump. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
252 lines
8.9 KiB
TypeScript
252 lines
8.9 KiB
TypeScript
/**
|
|
* Tests for the gnubok_call_tool bridge in the MCP dispatcher.
|
|
*
|
|
* Server-side, every tool has always been callable: `tools/call` resolves the
|
|
* name against the whole `tools` array, and `isDefaultCatalogTool` gates only
|
|
* what tools/list SHOWS. The failure was purely client-side, and DECISIONS.md
|
|
* records the consequence on 2026-08-26: `gnubok_reconcile_match` had to be
|
|
* promoted back into the default catalog because "a search-only tool is
|
|
* uncallable on Claude.ai".
|
|
*
|
|
* The bridge gives such a client one visible name to forward through. It is
|
|
* implemented as a REWRITE ahead of tool resolution rather than as a wrapper
|
|
* that calls the inner tool's execute(), because everything between resolution
|
|
* and execute (scope check, unknown-argument guard, company routing, the
|
|
* test-key write block, staging _meta, telemetry) must apply to the real
|
|
* target. These tests exist to prove it does.
|
|
*/
|
|
import { describe, it, expect, vi, beforeEach } from 'vitest'
|
|
import { eventBus } from '@/lib/events/bus'
|
|
|
|
vi.mock('@/lib/supabase/server', () => ({
|
|
createClient: vi.fn(),
|
|
createServiceClient: vi.fn(),
|
|
}))
|
|
|
|
vi.mock('@/lib/auth/api-keys', async (importOriginal) => {
|
|
const actual = await importOriginal<typeof import('@/lib/auth/api-keys')>()
|
|
const chain: unknown = new Proxy(
|
|
{},
|
|
{
|
|
get(_t, prop) {
|
|
if (prop === 'then') {
|
|
return (resolve: (v: unknown) => void) => resolve({ data: null, error: null })
|
|
}
|
|
return () => chain
|
|
},
|
|
},
|
|
)
|
|
const membershipChain: unknown = new Proxy(
|
|
{},
|
|
{
|
|
get(_t, prop) {
|
|
if (prop === 'then') {
|
|
return (resolve: (v: unknown) => void) =>
|
|
resolve({
|
|
data: { company_id: '11111111-1111-4111-8111-111111111111', role: 'owner' },
|
|
error: null,
|
|
})
|
|
}
|
|
return () => membershipChain
|
|
},
|
|
},
|
|
)
|
|
return {
|
|
...actual,
|
|
extractBearerToken: vi.fn().mockReturnValue('test-token'),
|
|
validateApiKey: vi.fn().mockResolvedValue({
|
|
userId: 'user-1',
|
|
companyId: '11111111-1111-4111-8111-111111111111',
|
|
scopes: ['transactions:read', 'reports:read', 'pending_operations:approve'],
|
|
apiKeyId: 'key-1',
|
|
apiKeyName: 'Live Key',
|
|
mode: 'live',
|
|
}),
|
|
createServiceClientNoCookies: vi.fn(() => ({
|
|
from: (table: string) => (table === 'company_members' ? membershipChain : chain),
|
|
rpc: () => chain,
|
|
})),
|
|
}
|
|
})
|
|
|
|
vi.mock('@/lib/entitlements/has-capability', async (importOriginal) => {
|
|
const actual = await importOriginal<typeof import('@/lib/entitlements/has-capability')>()
|
|
return { ...actual, hasCapability: vi.fn().mockResolvedValue(true) }
|
|
})
|
|
|
|
import { handleMcpRequest, tools, isDefaultCatalogTool } from '../server'
|
|
import { validateApiKey, extractBearerToken } from '@/lib/auth/api-keys'
|
|
|
|
function mcpToolCall(name: string, args: Record<string, unknown> = {}): Request {
|
|
return new Request('http://localhost:3000/api/extensions/ext/mcp-server/mcp', {
|
|
method: 'POST',
|
|
headers: { 'Content-Type': 'application/json', Authorization: 'Bearer test-token' },
|
|
body: JSON.stringify({
|
|
jsonrpc: '2.0',
|
|
id: 1,
|
|
method: 'tools/call',
|
|
params: { name, arguments: args },
|
|
}),
|
|
})
|
|
}
|
|
|
|
interface ToolCalledEvent {
|
|
tool: string
|
|
success: boolean
|
|
isError: boolean
|
|
errorKind: string | null
|
|
latencyMs: number
|
|
}
|
|
|
|
function captureNextToolCalled(): Promise<ToolCalledEvent> {
|
|
return new Promise((resolve) => {
|
|
const off = eventBus.on('mcp.tool_called', (payload) => {
|
|
off()
|
|
resolve(payload as unknown as ToolCalledEvent)
|
|
})
|
|
})
|
|
}
|
|
|
|
async function parsedToolResult(
|
|
response: Response,
|
|
): Promise<{ isError: boolean; payload: Record<string, unknown> }> {
|
|
const json = await response.json()
|
|
const result = json.result as { isError?: boolean; content: { text: string }[] }
|
|
return { isError: result.isError === true, payload: JSON.parse(result.content[0].text) }
|
|
}
|
|
|
|
const bridgeTool = tools.find((t) => t.name === 'gnubok_call_tool')!
|
|
|
|
describe('gnubok_call_tool registration', () => {
|
|
it('is in the default catalog and read-only', () => {
|
|
expect(bridgeTool).toBeDefined()
|
|
expect(isDefaultCatalogTool(bridgeTool)).toBe(true)
|
|
expect(bridgeTool.annotations.readOnlyHint).toBe(true)
|
|
})
|
|
|
|
it('has no direct implementation: the dispatcher rewrite is load-bearing', async () => {
|
|
// If this ever resolves instead of throwing, the rewrite was removed and
|
|
// every bridged call would have skipped the read-only check above it.
|
|
await expect(
|
|
bridgeTool.execute({}, 'company-id', 'user-id', {} as never, { type: 'api_key' }),
|
|
).rejects.toThrow(/no direct implementation/i)
|
|
})
|
|
})
|
|
|
|
describe('gnubok_call_tool bridge', () => {
|
|
beforeEach(() => {
|
|
vi.clearAllMocks()
|
|
eventBus.clear()
|
|
})
|
|
|
|
it('forwards to the inner tool and attributes telemetry to it, not to the wrapper', async () => {
|
|
const eventPromise = captureNextToolCalled()
|
|
|
|
await handleMcpRequest(mcpToolCall('gnubok_call_tool', { tool: 'gnubok_list_skills' }))
|
|
|
|
const event = await eventPromise
|
|
expect(event.tool).toBe('gnubok_list_skills')
|
|
expect(event.errorKind).not.toBe('bridge_refused')
|
|
})
|
|
|
|
it('reaches a search-only read tool, which is the whole point', async () => {
|
|
const searchOnlyRead = tools.find(
|
|
(t) => !isDefaultCatalogTool(t) && t.annotations.readOnlyHint === true,
|
|
)!
|
|
expect(searchOnlyRead).toBeDefined()
|
|
const eventPromise = captureNextToolCalled()
|
|
|
|
await handleMcpRequest(mcpToolCall('gnubok_call_tool', { tool: searchOnlyRead.name }))
|
|
|
|
const event = await eventPromise
|
|
expect(event.tool).toBe(searchOnlyRead.name)
|
|
expect(event.errorKind).not.toBe('bridge_refused')
|
|
})
|
|
|
|
it('refuses a write target so the staging and approval contract stays visible', async () => {
|
|
const eventPromise = captureNextToolCalled()
|
|
|
|
const response = await handleMcpRequest(
|
|
mcpToolCall('gnubok_call_tool', {
|
|
tool: 'gnubok_approve_pending_operation',
|
|
arguments: { operation_id: 'op-1' },
|
|
}),
|
|
)
|
|
const { isError, payload } = await parsedToolResult(response)
|
|
|
|
expect(isError).toBe(true)
|
|
expect(JSON.stringify(payload)).toContain('gnubok_approve_pending_operation')
|
|
const event = await eventPromise
|
|
expect(event.errorKind).toBe('bridge_refused')
|
|
// Refused before execute(): nothing is staged, nothing is approved.
|
|
expect(event.latencyMs).toBe(0)
|
|
})
|
|
|
|
it('refuses a call with no tool name', async () => {
|
|
const eventPromise = captureNextToolCalled()
|
|
|
|
const response = await handleMcpRequest(mcpToolCall('gnubok_call_tool', {}))
|
|
const { isError } = await parsedToolResult(response)
|
|
|
|
expect(isError).toBe(true)
|
|
const event = await eventPromise
|
|
expect(event.errorKind).toBe('bridge_refused')
|
|
})
|
|
|
|
it('enforces the INNER tool scope, not the wrapper (which has none)', async () => {
|
|
vi.mocked(validateApiKey).mockResolvedValueOnce({
|
|
userId: 'user-1',
|
|
companyId: '11111111-1111-4111-8111-111111111111',
|
|
// Deliberately omits transactions:read, which the inner tool requires.
|
|
scopes: ['reports:read'],
|
|
apiKeyId: 'key-1',
|
|
apiKeyName: 'Narrow Key',
|
|
mode: 'live',
|
|
} as Awaited<ReturnType<typeof validateApiKey>>)
|
|
const eventPromise = captureNextToolCalled()
|
|
|
|
const response = await handleMcpRequest(
|
|
mcpToolCall('gnubok_call_tool', { tool: 'gnubok_list_cash_accounts' }),
|
|
)
|
|
const { isError } = await parsedToolResult(response)
|
|
|
|
expect(isError).toBe(true)
|
|
const event = await eventPromise
|
|
expect(event.errorKind).toBe('scope_denied')
|
|
expect(event.tool).toBe('gnubok_list_cash_accounts')
|
|
})
|
|
|
|
it('applies the unknown-argument guard to the inner tool', async () => {
|
|
const response = await handleMcpRequest(
|
|
mcpToolCall('gnubok_call_tool', {
|
|
tool: 'gnubok_list_skills',
|
|
arguments: { nonexistent_parameter: 1 },
|
|
}),
|
|
)
|
|
const { isError, payload } = await parsedToolResult(response)
|
|
|
|
expect(isError).toBe(true)
|
|
expect(JSON.stringify(payload)).toContain('nonexistent_parameter')
|
|
})
|
|
|
|
it('is closed to anonymous callers: the pre-auth gate keys on the outer name', async () => {
|
|
// gnubok_call_tool is deliberately absent from PUBLIC_TOOLS, so an
|
|
// unauthenticated client cannot use it as a lever at all. Nothing is lost:
|
|
// all three public tools are in the default catalog already.
|
|
vi.mocked(extractBearerToken).mockReturnValueOnce(null)
|
|
|
|
const response = await handleMcpRequest(
|
|
mcpToolCall('gnubok_call_tool', { tool: 'gnubok_list_skills' }),
|
|
)
|
|
expect(response.status).toBe(401)
|
|
})
|
|
|
|
it('reports an unknown inner tool through the normal unknown-tool path', async () => {
|
|
const response = await handleMcpRequest(
|
|
mcpToolCall('gnubok_call_tool', { tool: 'gnubok_not_a_real_tool' }),
|
|
)
|
|
const json = (await response.json()) as { error?: { message?: string } }
|
|
expect(json.error?.message).toContain('gnubok_not_a_real_tool')
|
|
})
|
|
})
|