* feat(api): operations table immutability trigger BFNAR 2013:2 kap 8 § behandlingshistorik integrity: once an operations row is in a terminal status (succeeded / failed / cancelled) the audit record of what happened becomes immutable. Adds the BEFORE UPDATE and BEFORE DELETE triggers that the webhook_deliveries table already has (20260515170000 / 20260515190000), mirroring their predicate shape and error code exactly. Closes the Phase 4 PR-2 (PR #469) review-round carry-over flagged by Swedish-compliance: previously a future bug, a privileged operator, or a compromised service-role caller could rewrite "this year-end close succeeded" to "failed" by updating an already-terminal row. The running → succeeded/failed/cancelled transition itself stays legal because the trigger keys on OLD.status, which is non-terminal at the moment of the legitimate UPDATE. pg test covers all transitions (allowed and blocked) plus DELETE on both terminal and non-terminal rows. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(api): atomic SKIP LOCKED claim for webhook dispatch Replaces the SELECT-then-UPDATE-intersect pattern in the dispatcher with a single-roundtrip SQL function using FOR UPDATE SKIP LOCKED. PostgREST can't express SKIP LOCKED through the JS client, so the previous shape relied on a CAS guard inside an UPDATE WHERE status IN ('pending','failed') to ensure only one of two overlapping cron ticks claimed any given row. The CAS pattern was correct (under load — receivers >60s could push a batch past the next minute's tick) but burned two round trips and forced the application to negotiate the locking semantics in JS. The function form moves the contention to the DB, where SKIP LOCKED makes a row held by a concurrent tick simply invisible to the second caller. One round trip, no JS-side intersect. All filter semantics are preserved verbatim inside the function: status IN ('pending','failed'), next_attempt_at <= now, webhook_id IS NOT NULL, ORDER BY next_attempt_at ASC, LIMIT batchSize. p_batch_size is bounded (0, 1000] to forestall a runaway lock-set in case a caller misconfigures it. pg test covers basic claim (pending + failed), future-due skip, dangling- row (webhook_id IS NULL) skip, terminal-status skip, batch-size limits, out-of-range argument rejection, and the SKIP LOCKED invariant itself using two concurrent pool clients in BEGIN — the second caller does not see the row A locked, no double-delivery. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(api): pinned-IP HTTPS dispatch (close DNS-rebinding window) The url-guard.ts file header openly flagged the remaining gap: "a separate DNS-rebinding window (between dispatch-time validation and the actual fetch) remains; closing that requires a custom HTTPS agent that pins the resolved IP — tracked for follow-up." This closes it. The previous shape was: 1. validateWebhookUrl() → DNS resolves to [public IP], returns ok 2. fetch(webhook_url) → re-resolves DNS; an attacker who flipped the A record in the interval gets a private-IP socket The new pinnedHttpsFetch helper validates DNS once, then opens a node:https.request to that pinned IP — but keeps the original hostname in the TLS SNI extension (so the receiver's cert validates) and in the HTTP Host header (so vhost routing still works). The request socket never re-resolves DNS, foreclosing the rebind race entirely. Built on node:https.request rather than undici's Agent so the project doesn't take on a new dep — the stdlib API is also more explicit about the SNI / Host / pinned-IP split. Test seam injects both validateUrl and httpsRequest so the unit tests verify the pinning shape without standing up an HTTPS server. The dispatcher's attemptDelivery is rewritten as a switch over the four PinnedFetchResult kinds (ok / unsafe_url / redirect_blocked / timeout / transport_error). The previous fetch-based code path that distinguished redirect rejection by string-matching err.message is gone — the new result type makes the distinction structural. 8 unit tests cover the SNI/Host/pinned-IP shape, port handling, redirect_blocked, transport_error, timeout, response-body truncation, first-IP determinism, and the validation short-circuit (never opens a socket when the URL fails the SSRF guard). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(api): pg tests for webhook substrate triggers (PR-1 test debt) CLAUDE.md ("Testing" + "Migration Rules") mandates a *.pg.test.ts for any PR touching a trigger / RPC / RLS / DEFERRABLE constraint. Phase 6 PR-1 (#496) shipped three webhook_deliveries triggers without the accompanying pg test; this closes that debt. Triggers covered: - enforce_webhook_delivery_immutability (BEFORE UPDATE) - block_webhook_delivery_terminal_delete (BEFORE DELETE) - assert_webhook_delivery_company_match (BEFORE INSERT) 13 cases verify the lifecycle the dispatcher depends on remains mutable (pending → in_flight, in_flight → failed, failed → in_flight, in_flight → delivered) while terminal-status rows (delivered / dead) are write- locked and the cross-tenant INSERT path is refused with the ERRCODE=check_violation contract documented in the migration. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(api): integration tests for webhook routes (PR-1 test debt) CLAUDE.md mandates integration tests under app/api/v1/ for every route. Phase 6 PR-1 (#496) shipped the eight v1 webhook routes (five under /companies/{companyId}/webhooks/ + the cross-tenant /webhook-deliveries/ {id}/retry) without them; closes that debt. 19 cases for the /webhooks/ verticals: POST /webhooks create + secret-once + payroll-scope gate + SSRF GET /webhooks list (no secret) + empty list GET /webhooks/:id detail (no secret) + 404 PATCH /webhooks/:id update + active=true re-enable + SSRF re-check + empty-body DELETE /webhooks/:id 204 hard delete POST /webhooks/:id/test enqueue + 404 + disabled-rejection GET /webhooks/:id/deliveries happy path + ownership 404 7 cases for the retry route: POST /webhook-deliveries/:id/retry dead → fresh pending row, live-status refusal, cross-tenant 404, disabled-webhook gate, SSRF re-check, delivery 404, webhook-gone 404 Both files mirror the suppliers/customers integration test pattern: Proxy-backed Supabase mock with per-table queues, validateApiKey + validateWebhookUrl stubbed to control auth and DNS deterministically. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(api): address PR-500 review round 1 — pg-real CI fix + 4 review items 1. pg-real CI was red on this PR: the new webhook trigger pg.test.ts and claim-due-webhook-deliveries pg.test.ts fixtures tried to INSERT into `webhooks.user_id`, which doesn't exist in the migration history. The column was never declared in automation_webhooks (20260415000000) nor added by webhooks_v2 (20260515170000) — so a fresh schema replay had no such column. The webhook create route (`webhooks.create`) was also referencing this non-existent column in its INSERT, so the production route was latent-broken since PR-1 and never exercised against a fresh DB. Drop the `user_id` field from both the route INSERT and the pg fixtures. Actor attribution lives on `created_by_api_key_id` (which leads back to the owning user via `api_keys.user_id`). 2. Greptile P2 #1 — `recoverStuckInFlight` carried a redundant `.not('status','in','(delivered,dead)')` filter alongside `.eq('status','in_flight')`, with a comment that incorrectly described PostgreSQL's UPDATE re-evaluation semantics. Under READ COMMITTED, UPDATE re-evaluates WHERE against each row's CURRENT value when it acquires the row lock — a row that raced to terminal status will fail `status='in_flight'` on re-evaluation and be skipped, no immutability trigger fires. Drop the redundant filter and rewrite the comment. 3. Greptile P2 #2 — added explicit pg test verifying `in_flight` rows are skipped by `claim_due_webhook_deliveries`. The status filter is what prevents double-delivery and is the entire point of the SKIP LOCKED substrate; making that invariant load-bearing in the test suite forecloses a future filter expansion silently regressing it. 4. Greptile P2 #3 — pinned-fetch registered both `res.on('end', finalize)` and `res.on('close', finalize)`. Node fires BOTH on normal completions, so finalize ran twice; the outer `settled` guard squashed the double-resolve but the header reconstruction still ran twice. Switch to `once` + self-removing pair so finalize runs exactly once on whichever event fires first (normal: end; truncation: close). 5. Compliance Swarm V8.2.1 — the retry route only checked `webhooks:manage` even when retrying `salary_run.* / agi.*` deliveries. Mirror the create-route elevated-scope gate so a key with only `webhooks:manage` cannot re-emit payroll payloads carrying personnummer / lönesummor / skatteavdrag. New integration test verifies the gate returns 403 INSUFFICIENT_SCOPE with `required_scope: payroll:read`. 35 tests pass locally (+1 vs pre-fix). Type-check clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(api): address PR-500 review round 2 — 2 small precision fixes 1. Compliance Swarm Art.32 / A.8.24 — response_body size cap was enforced only at the application layer (pinnedHttpsFetch's maxResponseBytes=4096 constant). A future refactor that bypassed the truncation, or a non- dispatcher write path into webhook_deliveries.response_body, would silently land large blobs in a column adjacent to event payloads carrying personal data. Add a CHECK constraint at the DB layer with a generous ceiling (8 KB — double the application cap so legitimate dispatcher writes never hit it; only a regression surfaces as a check_violation). 2. Compliance Swarm CC6.6 — pinned-fetch substitutes the validated IP for `host` while keeping the original hostname in `servername`. A reader could reasonably worry that the IP substitution weakens TLS hostname verification. Document explicitly that Node's default `checkServerIdentity` matches the cert's SAN/CN against `servername` (not `host`), so a forged endpoint at the pinned IP with a valid cert for a different hostname would fail the handshake. No code change — the default behavior is correct; the comment forecloses future "this looks dangerous" review-round noise on the same line. Items NOT addressed (with rationale documented elsewhere): - Compliance Swarm V8.2.1 (retry route 404-vs-404 information leak): delivery IDs are UUIDs; the "leak" is the ability to probe existence of an opaque 128-bit identifier the caller already has, which is not meaningfully different from probing for any opaque token. Both branches return the same structured 404 envelope. - Compliance Swarm CC7.2 (restore the .not() defense-in-depth filter): direct contradiction of last round's Greptile P2 fix. Greptile's PG-semantics analysis is correct — under READ COMMITTED, UPDATE re-evaluates WHERE against the row's current value when it acquires the lock, so .eq('status','in_flight') already handles the race. Adding a redundant .not() restores a misleading comment without closing a real gap. This is the documented Compliance Swarm oscillation pattern from the project's Phase 4 lessons. - Compliance Swarm CC6.1 (webhook secret encryption-at-rest): architectural choice from PR-1; not in PR-3 (substrate hardening) scope. Belongs to a future hardening PR. - Swedish-compliance review (operations queued/running rows hard- deletable): deliberate operability tradeoff — operators need to clear stuck/queued entries that crashed mid-flight. Blocking all deletes would force a manual DB intervention every time a worker crashed before reaching terminal status. The audit trail starts at terminal-state mutation, which IS blocked. - Swedish-compliance review (salary_run.* / agi.* payload anonymisation after 7 years): already on the deferred-list as part of the 90-day TTL cleanup cron item from the PR description. Belongs to a retention-policy follow-up PR. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
568 lines
20 KiB
TypeScript
568 lines
20 KiB
TypeScript
/**
|
||
* Webhook delivery dispatcher.
|
||
*
|
||
* Invoked from the per-minute cron at /api/webhooks/dispatch/cron. Picks up
|
||
* pending + retry-due deliveries (FOR UPDATE SKIP LOCKED so multiple cron
|
||
* invocations don't double-deliver), POSTs each one with HMAC signature,
|
||
* and updates the row to one of:
|
||
*
|
||
* - delivered (2xx response) — terminal
|
||
* - failed (5xx / network / 4xx — non-terminal until attempts
|
||
* other than 410) exhausted; bumps next_attempt_at
|
||
* by exponential backoff
|
||
* - dead (HTTP 410 OR — terminal
|
||
* attempts exhausted)
|
||
*
|
||
* The receiver is expected to respond within 10 seconds; we time out
|
||
* aggressively so a slow receiver doesn't block the per-minute cron.
|
||
*
|
||
* On HTTP 410 we additionally disable the webhook (sets disabled_at +
|
||
* disabled_reason='HTTP 410 from receiver') so future events don't even
|
||
* enqueue against it.
|
||
*/
|
||
|
||
import type { SupabaseClient } from '@supabase/supabase-js'
|
||
import { signPayload } from './signing'
|
||
import { pinnedHttpsFetch, type PinnedFetchResult } from './pinned-fetch'
|
||
import { createLogger } from '@/lib/logger'
|
||
|
||
const log = createLogger('webhooks/dispatcher')
|
||
|
||
/** 7 retries over ~72h. Index = attempts BEFORE this one. */
|
||
const RETRY_BACKOFF_SECONDS: ReadonlyArray<number> = [
|
||
60, // 1m — first retry
|
||
5 * 60, // 5m
|
||
30 * 60, // 30m
|
||
2 * 60 * 60, // 2h
|
||
12 * 60 * 60, // 12h
|
||
24 * 60 * 60, // 24h
|
||
48 * 60 * 60, // 48h — final retry
|
||
]
|
||
|
||
const MAX_ATTEMPTS = RETRY_BACKOFF_SECONDS.length + 1 // initial + 7 retries = 8 total
|
||
const REQUEST_TIMEOUT_MS = 10_000
|
||
const MAX_RESPONSE_BODY_BYTES = 4096
|
||
|
||
interface DueDelivery {
|
||
id: string
|
||
webhook_id: string
|
||
company_id: string
|
||
event_type: string
|
||
payload: Record<string, unknown>
|
||
previous_attributes: Record<string, unknown> | null
|
||
api_version: string
|
||
attempts: number
|
||
}
|
||
|
||
interface WebhookForDelivery {
|
||
id: string
|
||
company_id: string
|
||
webhook_url: string
|
||
secret: string
|
||
}
|
||
|
||
export interface DispatchSummary {
|
||
picked: number
|
||
delivered: number
|
||
failed: number
|
||
dead: number
|
||
}
|
||
|
||
/**
|
||
* Run one dispatch cycle. Picks up to `batchSize` due deliveries and
|
||
* processes them sequentially (the per-minute cadence + small batch size
|
||
* makes parallelism unnecessary; in-process serial is also gentler on the
|
||
* receiver if many events fan out to the same URL).
|
||
*/
|
||
export async function dispatchDueDeliveries(args: {
|
||
supabase: SupabaseClient
|
||
/** Max rows to claim per cron tick. Default 50. */
|
||
batchSize?: number
|
||
/** Override for tests. */
|
||
now?: Date
|
||
/** Override for tests; injected pinned-fetch implementation. */
|
||
pinnedFetchImpl?: typeof pinnedHttpsFetch
|
||
}): Promise<DispatchSummary> {
|
||
const batchSize = args.batchSize ?? 50
|
||
const now = args.now ?? new Date()
|
||
const pinnedFetchImpl = args.pinnedFetchImpl ?? pinnedHttpsFetch
|
||
|
||
const summary: DispatchSummary = { picked: 0, delivered: 0, failed: 0, dead: 0 }
|
||
|
||
// Recover stuck in_flight rows: a previous tick that was killed mid-flight
|
||
// (Vercel function timeout, hard crash, manual termination) leaves rows
|
||
// marked in_flight forever otherwise. Sweep them back to 'failed' so the
|
||
// retry loop picks them up at next_attempt_at.
|
||
//
|
||
// Threshold = 2× REQUEST_TIMEOUT_MS. A live attempt takes at most
|
||
// REQUEST_TIMEOUT_MS plus the body read; doubling that gives an
|
||
// unambiguous "this is stuck, not in-flight" boundary.
|
||
await recoverStuckInFlight(args.supabase, now)
|
||
|
||
const due = await claimDueDeliveries(args.supabase, batchSize, now)
|
||
summary.picked = due.length
|
||
if (due.length === 0) return summary
|
||
|
||
// Dedupe webhook lookups within a single cycle.
|
||
const webhookIds = Array.from(new Set(due.map((d) => d.webhook_id)))
|
||
const webhookMap = await loadWebhooksByIds(args.supabase, webhookIds)
|
||
|
||
for (const delivery of due) {
|
||
const webhook = webhookMap.get(delivery.webhook_id)
|
||
if (!webhook) {
|
||
// The webhook was deleted between enqueue and dispatch. Mark dead;
|
||
// there's no receiver to deliver to. The webhook_deliveries.webhook_id
|
||
// FK is ON DELETE SET NULL (migration 20260515170000), so the row
|
||
// stays in the audit trail under status='dead'.
|
||
await markDead(args.supabase, delivery.id, 'webhook_deleted')
|
||
summary.dead++
|
||
continue
|
||
}
|
||
|
||
// Defense-in-depth tenancy check: the webhook the delivery row points
|
||
// at MUST belong to the same company as the delivery row. Mismatch
|
||
// indicates a poisoned row — refuse to dispatch (which would sign with
|
||
// the wrong tenant's secret and POST to the wrong receiver).
|
||
if (webhook.company_id !== delivery.company_id) {
|
||
log.error('cross-tenant delivery refused', new Error('company_id mismatch'), {
|
||
deliveryId: delivery.id,
|
||
deliveryCompanyId: delivery.company_id,
|
||
webhookId: webhook.id,
|
||
webhookCompanyId: webhook.company_id,
|
||
})
|
||
await markDead(args.supabase, delivery.id, 'cross_tenant_mismatch')
|
||
summary.dead++
|
||
continue
|
||
}
|
||
|
||
const outcome = await attemptDelivery({
|
||
delivery,
|
||
webhook,
|
||
pinnedFetchImpl,
|
||
now,
|
||
})
|
||
|
||
// Structured per-delivery outcome log. Keeps companyId / webhookId /
|
||
// deliveryId available in log aggregation for per-tenant audit-trail
|
||
// reconstruction without grepping through individual mark*-helper
|
||
// writes (V16 — security event correlation).
|
||
const logCtx = {
|
||
deliveryId: delivery.id,
|
||
webhookId: webhook.id,
|
||
companyId: delivery.company_id,
|
||
eventType: delivery.event_type,
|
||
attempt: delivery.attempts + 1,
|
||
}
|
||
|
||
switch (outcome.kind) {
|
||
case 'delivered':
|
||
await markDelivered(args.supabase, delivery.id, outcome)
|
||
log.info('delivery succeeded', { ...logCtx, responseStatus: outcome.responseStatus })
|
||
summary.delivered++
|
||
break
|
||
case 'dead':
|
||
await markDead(args.supabase, delivery.id, outcome.reason, outcome)
|
||
log.warn('delivery dead', { ...logCtx, reason: outcome.reason, responseStatus: outcome.responseStatus })
|
||
summary.dead++
|
||
if (outcome.disableWebhook) {
|
||
await disableWebhook(args.supabase, webhook.id, outcome.reason)
|
||
log.warn('webhook auto-disabled', { ...logCtx, reason: outcome.reason })
|
||
}
|
||
break
|
||
case 'failed':
|
||
if (delivery.attempts + 1 >= MAX_ATTEMPTS) {
|
||
await markDead(args.supabase, delivery.id, 'attempts_exhausted', outcome)
|
||
log.warn('delivery dead — attempts exhausted', { ...logCtx, lastError: outcome.error })
|
||
summary.dead++
|
||
} else {
|
||
await markFailedForRetry(args.supabase, delivery.id, delivery.attempts, outcome, now)
|
||
log.info('delivery failed — retry scheduled', { ...logCtx, error: outcome.error, responseStatus: outcome.responseStatus })
|
||
summary.failed++
|
||
}
|
||
break
|
||
}
|
||
}
|
||
|
||
return summary
|
||
}
|
||
|
||
// ──────────────────────────────────────────────────────────────────────
|
||
// DB ops
|
||
// ──────────────────────────────────────────────────────────────────────
|
||
|
||
/**
|
||
* Mark in_flight rows whose updated_at is older than the stuck-threshold
|
||
* back to 'failed' with next_attempt_at = now so they re-enter the
|
||
* dispatch queue. Best-effort — a write failure here is logged but
|
||
* doesn't block the rest of the cycle.
|
||
*/
|
||
async function recoverStuckInFlight(supabase: SupabaseClient, now: Date): Promise<void> {
|
||
const stuckBefore = new Date(now.getTime() - 2 * REQUEST_TIMEOUT_MS)
|
||
// Under READ COMMITTED (Postgres default), UPDATE re-evaluates the WHERE
|
||
// clause against each row's current value when it acquires the row lock.
|
||
// A row that raced from 'in_flight' to 'delivered'/'dead' between scan
|
||
// and lock will fail status='in_flight' on re-evaluation and be skipped
|
||
// entirely — the immutability trigger never fires, so a mid-flight
|
||
// terminal flip cannot abort the bulk update.
|
||
const { data, error } = await supabase
|
||
.from('webhook_deliveries')
|
||
.update({
|
||
status: 'failed',
|
||
next_attempt_at: now.toISOString(),
|
||
error: 'recovered_from_in_flight_timeout',
|
||
})
|
||
.eq('status', 'in_flight')
|
||
.lt('updated_at', stuckBefore.toISOString())
|
||
.select('id')
|
||
|
||
if (error) {
|
||
log.warn('stuck in_flight recovery failed', { code: error.code })
|
||
return
|
||
}
|
||
if (data && data.length > 0) {
|
||
log.warn('recovered stuck in_flight rows', { count: data.length })
|
||
}
|
||
}
|
||
|
||
async function claimDueDeliveries(
|
||
supabase: SupabaseClient,
|
||
batchSize: number,
|
||
now: Date,
|
||
): Promise<DueDelivery[]> {
|
||
// Atomic FOR UPDATE SKIP LOCKED claim via the SQL function shipped in
|
||
// migration 20260515220000. PostgREST can't express SKIP LOCKED through
|
||
// the JS client, so the function form is the documented entry point —
|
||
// see the migration comment for the full rationale (one round trip,
|
||
// no CAS contention, rows locked by a concurrent tick are simply
|
||
// invisible to the second caller).
|
||
//
|
||
// All filter semantics from the previous JS path are preserved inside
|
||
// the function: status IN ('pending','failed'), next_attempt_at <= now,
|
||
// webhook_id IS NOT NULL, ORDER BY next_attempt_at ASC, LIMIT batchSize.
|
||
const { data, error } = await supabase.rpc('claim_due_webhook_deliveries', {
|
||
p_batch_size: batchSize,
|
||
p_now: now.toISOString(),
|
||
})
|
||
|
||
if (error) {
|
||
log.error('claim_due_webhook_deliveries rpc failed', error as Error)
|
||
return []
|
||
}
|
||
return (data ?? []) as DueDelivery[]
|
||
}
|
||
|
||
async function loadWebhooksByIds(
|
||
supabase: SupabaseClient,
|
||
ids: string[],
|
||
): Promise<Map<string, WebhookForDelivery>> {
|
||
// Include company_id so the dispatch loop can assert that the delivery
|
||
// row's company_id matches the webhook's — defense in depth against a
|
||
// poisoned delivery row pointing at another tenant's webhook
|
||
// (compromised service-role path, faulty INSERT in a future code path,
|
||
// etc.). The DB trigger added in 20260515190000 enforces the same
|
||
// invariant at INSERT time; this is the application-layer mirror.
|
||
const { data, error } = await supabase
|
||
.from('webhooks')
|
||
.select('id, company_id, webhook_url, secret')
|
||
.in('id', ids)
|
||
|
||
if (error || !data) {
|
||
log.error('webhook lookup for dispatch failed', error as Error)
|
||
return new Map()
|
||
}
|
||
return new Map((data as WebhookForDelivery[]).map((w) => [w.id, w]))
|
||
}
|
||
|
||
async function markDelivered(
|
||
supabase: SupabaseClient,
|
||
id: string,
|
||
outcome: DeliveredOutcome,
|
||
): Promise<void> {
|
||
const { error } = await supabase
|
||
.from('webhook_deliveries')
|
||
.update({
|
||
status: 'delivered',
|
||
delivered_at: new Date().toISOString(),
|
||
attempts: outcome.attempts,
|
||
response_status: outcome.responseStatus,
|
||
response_body: outcome.responseBody,
|
||
response_headers: outcome.responseHeaders,
|
||
error: null,
|
||
})
|
||
.eq('id', id)
|
||
if (error) log.warn('mark delivered update failed', { id, code: error.code })
|
||
}
|
||
|
||
async function markFailedForRetry(
|
||
supabase: SupabaseClient,
|
||
id: string,
|
||
priorAttempts: number,
|
||
outcome: FailedOutcome,
|
||
now: Date,
|
||
): Promise<void> {
|
||
const nextAttemptIndex = priorAttempts // 0-indexed lookup into RETRY_BACKOFF_SECONDS
|
||
const backoffSeconds = RETRY_BACKOFF_SECONDS[Math.min(nextAttemptIndex, RETRY_BACKOFF_SECONDS.length - 1)]
|
||
const nextAttemptAt = new Date(now.getTime() + backoffSeconds * 1000)
|
||
|
||
const { error } = await supabase
|
||
.from('webhook_deliveries')
|
||
.update({
|
||
status: 'failed',
|
||
attempts: priorAttempts + 1,
|
||
next_attempt_at: nextAttemptAt.toISOString(),
|
||
response_status: outcome.responseStatus ?? null,
|
||
response_body: outcome.responseBody ?? null,
|
||
response_headers: outcome.responseHeaders ?? null,
|
||
error: outcome.error,
|
||
})
|
||
.eq('id', id)
|
||
if (error) log.warn('mark failed-for-retry update failed', { id, code: error.code })
|
||
}
|
||
|
||
async function markDead(
|
||
supabase: SupabaseClient,
|
||
id: string,
|
||
reason: string,
|
||
outcome?: AttemptOutcome,
|
||
): Promise<void> {
|
||
// delivered_at means "the receiver acknowledged the event". For dead
|
||
// rows (HTTP 410, attempts exhausted, webhook deleted, cross-tenant
|
||
// mismatch, unsafe URL) the receiver did NOT acknowledge — leaving
|
||
// delivered_at NULL keeps the audit semantics clean. An auditor
|
||
// querying `WHERE delivered_at IS NOT NULL` correctly sees only
|
||
// genuinely delivered rows. The terminal-state timestamp lives on
|
||
// `updated_at` (auto-stamped by the table's BEFORE UPDATE trigger).
|
||
const { error } = await supabase
|
||
.from('webhook_deliveries')
|
||
.update({
|
||
status: 'dead',
|
||
attempts: outcome && 'attempts' in outcome ? outcome.attempts : undefined,
|
||
response_status: outcome && 'responseStatus' in outcome ? outcome.responseStatus : null,
|
||
response_body: outcome && 'responseBody' in outcome ? outcome.responseBody : null,
|
||
response_headers: outcome && 'responseHeaders' in outcome ? outcome.responseHeaders : null,
|
||
error: reason,
|
||
})
|
||
.eq('id', id)
|
||
if (error) log.warn('mark dead update failed', { id, code: error.code })
|
||
}
|
||
|
||
async function disableWebhook(
|
||
supabase: SupabaseClient,
|
||
webhookId: string,
|
||
reason: string,
|
||
): Promise<void> {
|
||
const { error } = await supabase
|
||
.from('webhooks')
|
||
.update({
|
||
disabled_at: new Date().toISOString(),
|
||
disabled_reason: reason,
|
||
active: false,
|
||
})
|
||
.eq('id', webhookId)
|
||
if (error) log.warn('webhook auto-disable failed', { webhookId, code: error.code })
|
||
}
|
||
|
||
// ──────────────────────────────────────────────────────────────────────
|
||
// HTTP attempt
|
||
// ──────────────────────────────────────────────────────────────────────
|
||
|
||
type DeliveredOutcome = {
|
||
kind: 'delivered'
|
||
attempts: number
|
||
responseStatus: number
|
||
responseBody: string | null
|
||
responseHeaders: Record<string, string> | null
|
||
}
|
||
|
||
type FailedOutcome = {
|
||
kind: 'failed'
|
||
attempts: number
|
||
responseStatus: number | null
|
||
responseBody: string | null
|
||
responseHeaders: Record<string, string> | null
|
||
error: string
|
||
}
|
||
|
||
type DeadOutcome = {
|
||
kind: 'dead'
|
||
reason: string
|
||
disableWebhook: boolean
|
||
attempts: number
|
||
responseStatus: number | null
|
||
responseBody: string | null
|
||
responseHeaders: Record<string, string> | null
|
||
error?: string
|
||
}
|
||
|
||
type AttemptOutcome = DeliveredOutcome | FailedOutcome | DeadOutcome
|
||
|
||
async function attemptDelivery(args: {
|
||
delivery: DueDelivery
|
||
webhook: WebhookForDelivery
|
||
pinnedFetchImpl: typeof pinnedHttpsFetch
|
||
now: Date
|
||
}): Promise<AttemptOutcome> {
|
||
const { delivery, webhook, pinnedFetchImpl, now } = args
|
||
const attempts = delivery.attempts + 1
|
||
const requestId = `whdel_${delivery.id}`
|
||
|
||
const body = JSON.stringify({
|
||
id: delivery.id,
|
||
type: delivery.event_type,
|
||
api_version: delivery.api_version,
|
||
created: Math.floor(now.getTime() / 1000),
|
||
data: { object: delivery.payload },
|
||
previous_attributes: delivery.previous_attributes,
|
||
})
|
||
|
||
const { header } = signPayload({
|
||
body,
|
||
secret: webhook.secret,
|
||
timestamp: Math.floor(now.getTime() / 1000),
|
||
})
|
||
|
||
// pinnedHttpsFetch performs DNS validation AND opens the socket against
|
||
// the validated IP in a single call. The previous shape (separate
|
||
// validateWebhookUrl + fetch calls) left a DNS-rebinding window between
|
||
// the two — closed here. SNI + Host header continue to carry the
|
||
// original hostname so receiver-side TLS + vhost routing still work.
|
||
const result = await pinnedFetchImpl(webhook.webhook_url, {
|
||
method: 'POST',
|
||
headers: {
|
||
'Content-Type': 'application/json',
|
||
'X-Gnubok-Signature': header,
|
||
'X-Gnubok-Event': delivery.event_type,
|
||
'X-Gnubok-Delivery': delivery.id,
|
||
'X-Gnubok-Api-Version': delivery.api_version,
|
||
'X-Request-Id': requestId,
|
||
'User-Agent': 'gnubok-webhook/1',
|
||
},
|
||
body,
|
||
timeoutMs: REQUEST_TIMEOUT_MS,
|
||
maxResponseBytes: MAX_RESPONSE_BODY_BYTES,
|
||
})
|
||
|
||
switch (result.kind) {
|
||
case 'unsafe_url':
|
||
return {
|
||
kind: 'dead',
|
||
reason: `url_unsafe:${result.reason}`,
|
||
disableWebhook: true,
|
||
attempts,
|
||
responseStatus: null,
|
||
responseBody: null,
|
||
responseHeaders: null,
|
||
error: result.detail,
|
||
}
|
||
case 'redirect_blocked':
|
||
return {
|
||
kind: 'dead',
|
||
reason: 'redirect_blocked',
|
||
disableWebhook: true,
|
||
attempts,
|
||
responseStatus: result.status,
|
||
responseBody: null,
|
||
responseHeaders: null,
|
||
error: truncateError(result.detail),
|
||
}
|
||
case 'timeout':
|
||
case 'transport_error':
|
||
return {
|
||
kind: 'failed',
|
||
attempts,
|
||
responseStatus: null,
|
||
responseBody: null,
|
||
responseHeaders: null,
|
||
error: truncateError(result.detail),
|
||
}
|
||
case 'ok': {
|
||
const responseHeaders = filterResponseHeaders(result.headers)
|
||
const responseBody = isSafeContentType(result.headers['content-type'] ?? '')
|
||
? result.body
|
||
: null
|
||
|
||
// HTTP 410 — receiver explicitly asks us to stop. Auto-disable.
|
||
if (result.status === 410) {
|
||
return {
|
||
kind: 'dead',
|
||
reason: 'http_410_gone',
|
||
disableWebhook: true,
|
||
attempts,
|
||
responseStatus: 410,
|
||
responseBody,
|
||
responseHeaders,
|
||
}
|
||
}
|
||
|
||
if (result.status >= 200 && result.status < 300) {
|
||
return {
|
||
kind: 'delivered',
|
||
attempts,
|
||
responseStatus: result.status,
|
||
responseBody,
|
||
responseHeaders,
|
||
}
|
||
}
|
||
|
||
return {
|
||
kind: 'failed',
|
||
attempts,
|
||
responseStatus: result.status,
|
||
responseBody,
|
||
responseHeaders,
|
||
error: `HTTP ${result.status}`,
|
||
}
|
||
}
|
||
}
|
||
}
|
||
|
||
function truncateError(message: string): string {
|
||
return message.length > 500 ? `${message.slice(0, 497)}...` : message
|
||
}
|
||
|
||
// Content-Type prefixes for which we persist response_body verbatim. Other
|
||
// types (text/html error pages, application/octet-stream, ...) get dropped
|
||
// because they routinely echo PII back from receiver-side error renderers
|
||
// (Art.32(1)(b), A.8.12). A null body is just as useful for debugging
|
||
// when the operator can see the response_status and response_headers.
|
||
const SAFE_BODY_CONTENT_TYPE_PREFIXES = ['text/plain', 'application/json']
|
||
|
||
function isSafeContentType(contentType: string): boolean {
|
||
const lower = contentType.toLowerCase()
|
||
return SAFE_BODY_CONTENT_TYPE_PREFIXES.some((p) => lower.startsWith(p))
|
||
}
|
||
|
||
// Allowlist for response_headers persistence. Receiver-side headers like
|
||
// Set-Cookie, Authorization, WWW-Authenticate, internal tracing, and
|
||
// vendor x-* headers can carry credentials or sensitive identifiers; we
|
||
// don't need them for delivery diagnostics. (CC7.2 / Art.32(1)(b))
|
||
//
|
||
// 'server' is deliberately NOT in the allowlist (A.8.12): it carries no
|
||
// diagnostic value but routinely leaks receiver infrastructure version
|
||
// strings (nginx/1.21.6, Apache/2.4.41, ...) into a multi-tenant audit
|
||
// table.
|
||
const SAFE_RESPONSE_HEADERS = new Set([
|
||
'content-type',
|
||
'content-length',
|
||
'date',
|
||
'x-request-id',
|
||
'cf-ray',
|
||
])
|
||
|
||
function filterResponseHeaders(headers: Record<string, string>): Record<string, string> {
|
||
const obj: Record<string, string> = {}
|
||
for (const [k, v] of Object.entries(headers)) {
|
||
if (SAFE_RESPONSE_HEADERS.has(k.toLowerCase())) {
|
||
obj[k] = v
|
||
}
|
||
}
|
||
return obj
|
||
}
|
||
|
||
export const __TESTING__ = {
|
||
RETRY_BACKOFF_SECONDS,
|
||
MAX_ATTEMPTS,
|
||
REQUEST_TIMEOUT_MS,
|
||
MAX_RESPONSE_BODY_BYTES,
|
||
}
|