Files
accounted/extensions/general/whatsapp-inbox/lib/sweep.ts
T
88760ae6f6 fix(whatsapp-inbox): harden against adversarial review findings (#1342)
* fix(whatsapp-inbox): erase the WhatsApp channel on account deletion

whatsapp_phone_links relied on the auth.users ON DELETE CASCADE, but
Accounted never deletes auth.users: account deletion is
anonymize_user_account plus a ~100-year ban that keeps the auth row as a
tombstone, so the cascade never fires and nothing revokes the link. After
erasure the link stayed active with a decryptable phone_enc,
lookupActiveLink kept resolving the number, and every further inbound
message was persisted with body_text and the verbatim raw_payload while
the bot kept replying: GDPR Art 17 plus continued collection with no
lawful basis.

The RPC is re-created verbatim from 20260724150000 with one added block
that revokes and crypto-shreds the link, resets its conversation, nulls
body_text/raw_payload on that link's messages and deletes outstanding
link codes, plus a guarded repair pass for tombstones anonymized before
this migration. Covered by a pg-real test that fails against the previous
definition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): pepper the link-code hash and bound code minting

hashLinkCode stored a bare sha256 over CODE_ALPHABET^6 = 30^6 values
behind a fixed 'AC-' prefix. The module cited the invite-token pattern,
but invite tokens are 256-bit random; this space enumerates offline in
about a second, so hashing at rest protected nothing. The sibling
phone-crypto.ts already states the team's own threat model for a LARGER
space ("a plain sha256 would be brute-forceable ... hence the pepper"),
so link codes now hash through the same env-mandated pepper.

/link/start was also an authenticated unbounded INSERT that left every
earlier code valid. Minting now burns the caller's unused codes (the code
the panel shows is the only one that works) and is capped per TTL window,
with the route answering 429 instead of throwing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): harden the conversation layer against the review findings

Pre-merge hardening of the unshipped chat layer. Every change below has a
test that fails without it.

Lifecycle and races:
- conversation writes go through updateConversation(), an optimistic
  compare-and-set on updated_at (the trigger makes it a revision counter).
  The ack winner, the answer worker, the pin refresh and the sweep hold
  different claims, so blind whole-jsonb writes resurrected answered
  questions, wiped pending_question and dropped queue entries.
- terminal markStatus writes are guarded on processing_status='processing'
  so a losing worker cannot overwrite the winner's 'done' and null its
  inbox_item_id.
- the message -> inbox item path is idempotent: a pre-check plus a 23505
  fallback adopt the item a concurrent worker created, instead of throwing
  after the WORM document is already committed.
- PROCESSING_STUCK_MS 90s -> 5 min. The enforced step budget of one media
  row already exceeds 90s, so the sweep was re-claiming live workers.
- sweep 2b re-arms only when the conversation itself has been quiet, not
  just the rows: pending_ack=false plus unacked rows is also the state of a
  live finalize, which produced a duplicate combined ack.
- pin expiry re-checks against fresh state instead of writing back a stale
  whole context, which reverted company choices applied mid-pass.
- askNextQueuedQuestion claims the pop before sending, so two answer
  workers cannot ask the same question twice.

Company question:
- the state is rolled back when the M6 send fails, so the next receipt
  re-asks instead of parking receipts behind a question nobody received.
- applyCompanyChoice claims the open question (company_options) rather
  than the state: a double tap confirms once, a transient membership-query
  error is no longer read as "not a member", and a LATE answer still lands.
- at the 48h TTL the parked receipts are kept, not discarded: options and
  staged rows survive so a late digit or tap still files them, and only
  rows past Meta's ~30-day media window get the terminal marker.
- an out-of-range digit or a typed company name now gets the options
  repeated instead of silence or the "I cannot answer questions" reply.

Inline dispositions:
- stop/start/byt/company answers run their side effect BEFORE the terminal
  wamid row, with a SELECT pre-check for dedupe. Writing the row 'done'
  first made them at-most-once: a crash in between lost the action forever.

Copy and answers:
- acks state the extracted currency instead of labelling every total 'kr'.
- M17 stops promising "about 10 minutes" when the daily quota tripped.
- M18 is sent once per message tracked by the outbound row, so a file
  whose first attempt died still reaches the sender, including from the
  max-attempts path.
- M11 no longer claims the number is disconnected: 'stopp' pauses, and
  muted senders now persist no chat content at all.
- 'byt' is recognized in every state but awaiting_company (m6-confirm
  teaches the word, and it was being stored as answer data instead).
- text sent while a re-send question is open is kept as a note on THAT
  receipt with the question left open, instead of binding to another
  receipt's question.
- a quoted reply wins over the pending question and is appended when the
  quoted question is already answered, so corrections stop landing on the
  wrong receipt.
- context answers keep raw_answer + answered_at like representation does.
- finalizeBurst checks the send result: on failure it rolls the question
  back and leaves the rows unacked for the sweep.

PII:
- the sender's plaintext number is stripped from raw_payload before it is
  persisted; replies decrypt the link's phone_enc instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(whatsapp-inbox): record the erasure path and the hardening decisions

RoPA gains the account-deletion row (immediate, not via the cron: the
auth.users cascade never fires because the row is tombstoned) plus the
two new security measures, and its "never in the clear" phone claim is
now true of the stored payload. DECISIONS.md records the non-obvious
calls: revoke-not-delete on erasure, commit-then-roll-back for the
company question, keeping expired company choices answerable, the
compare-and-set conversation write, effect-before-terminal-row for inline
dispositions, honest M11 copy, and the raw_payload redaction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): stop the answer re-claim from following a confirm with M16

A worker that died after applying an answer and sending its confirmation
leaves the row 'processing'. The sweep re-runs it, resolveAnswerTarget
finds the question already answered, and the user got "I did not
understand" immediately after the confirmation they had just received.
The fallback is now first-attempt only.

The catch comment claiming the sweep retries these rows is corrected
too: 'error' is terminal for the sweep, and nothing on the answer path
throws anyway (interpretChatAnswer degrades, sends never throw,
supabase-js returns errors), so the catch is a programming-error net.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(whatsapp-inbox): drop the amount floor on the representation question

The Swedish compliance review on #1340 caught a real error in the trigger
rules: the representation question only fired above 150 kr, but the duty to
document deltagare and syfte is what makes the expense deductible at all
(BFL 5 kap 6-7 §) and it is not conditioned on any amount. The 300 kr per
person figure I had in mind is the VAT-deduction base cap, a different rule.
A 120 kr business lunch would have been booked with no participant trail,
which is exactly the deduction Skatteverket denies later.

Noise stays bounded by the triggers that were already there: the question
fires only for receipt-shaped documents from restaurant, cafe or hotel
merchants, at most once per receipt, twice per burst and six times per
sender per day, and a single "nej" dismisses it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
2026-08-05 15:47:23 +02:00

385 lines
15 KiB
TypeScript

/**
* Per-minute crash-recovery sweep for the WhatsApp channel.
*
* The webhook 200s fast and defers all real work to after() invocations that
* can die with the serverless instance. Everything here is a re-derivation
* from durable state, so a lost invocation is a latency regression, never a
* lost message:
*
* 1. Re-claim whatsapp_messages stuck in 'received' (>60s) or 'processing'
* (>5 min, safely above the worst-case live worker); after MAX_ATTEMPTS
* they land in 'error' and the sender gets one M18.
* 2. Claim stale pending_ack conversations (debounce crash) and send the
* combined ack; re-arm conversations whose winner died after claiming
* but before sending (done rows left unacked).
* 3. Expire questions past the 48h TTL: conversation back to idle, the
* item's pending_question -> moved_to_app. NEVER sends anything: the 24h
* service window is long gone, and v1 sends no templates. Company
* questions keep their options and their parked receipts, so a late
* answer still files them (see the pass itself).
* 4. Clear expired 8h company pins.
*/
import type { SupabaseClient } from '@supabase/supabase-js'
import { createLogger } from '@/lib/logger'
import type { WhatsAppConversation, WhatsAppMessage } from '@/types'
import {
COMPANY_CHOICE_EXPIRED,
QUESTION_TTL_MS,
STAGED_AWAITING_COMPANY,
getContext,
resolveRecipient,
updateConversation,
type ConversationContext,
} from './conversation'
import { finalizeBurst, processInboundMessage, sendErrorNoticeOnce } from './process-inbound'
import { appendQuestionHistory, updateItemContext } from './item-context'
const log = createLogger('whatsapp-inbox/sweep')
const RECEIVED_STUCK_MS = 60 * 1000
/**
* A 'processing' row is only stuck if no live worker can still be on it.
* The enforced step budget of one media row is markRead (10s) + media lookup
* (10s) + download (30s) + Bedrock extraction (the cron route budgets 10-60s,
* with no short SDK timeout), under a maxDuration of 300s, and there is no
* heartbeat between the claim and the terminal write. 90s therefore re-claimed
* live workers on ordinary large PDFs and ran two of them on the same message.
* A crashed row waiting five minutes is a latency regression; two concurrent
* workers are a correctness problem.
*/
const PROCESSING_STUCK_MS = 5 * 60 * 1000
const ACK_STALE_MS = 60 * 1000
const UNACKED_REARM_MS = 120 * 1000
const MAX_ATTEMPTS = 3
const BATCH = 25
/** Staged receipts stay answerable while Meta still serves their media
* (~30 days). Past that the marker is honest: nothing can recover them. */
const STAGED_MEDIA_MAX_AGE_MS = 30 * 24 * 60 * 60 * 1000
export interface SweepSummary {
reclaimedReceived: number
reclaimedProcessing: number
erroredMaxAttempts: number
finalizedAcks: number
expiredQuestions: number
clearedPins: number
}
interface StuckRow {
id: string
attempts: number
conversation_id: string | null
direction: string
message_type: string
sender_phone_hash: string | null
phone_link_id: string | null
correlation_id: string | null
raw_payload: Record<string, unknown> | null
}
/**
* Park a row that ran out of attempts, and tell the sender once. Without the
* notice a file whose FIRST attempt died with the instance ends terminally
* with no ack and no error: the burst ack only lists ingested rows, so that
* receipt simply vanishes from the conversation.
*/
async function markMaxAttempts(
supabase: SupabaseClient,
row: StuckRow,
fromStatus: 'received' | 'processing',
): Promise<void> {
const { data: parked } = await supabase
.from('whatsapp_messages')
.update({ processing_status: 'error', error_message: 'Max attempts exceeded' })
.eq('id', row.id)
.eq('processing_status', fromStatus)
.select('id')
if (Array.isArray(parked) && parked.length === 0) return
if (row.message_type === 'text') return // M18 is about files
const link = row.phone_link_id
? await loadPhoneLink(supabase, row.phone_link_id)
: null
const to = resolveRecipient(row as unknown as WhatsAppMessage, link)
if (!to) return
await sendErrorNoticeOnce(supabase, {
to,
senderPhoneHash: row.sender_phone_hash,
phoneLinkId: row.phone_link_id,
conversationId: row.conversation_id,
correlationId: row.correlation_id,
})
}
async function loadPhoneLink(
supabase: SupabaseClient,
phoneLinkId: string,
): Promise<{ phone_enc: string | null } | null> {
const { data } = await supabase
.from('whatsapp_phone_links')
.select('phone_enc')
.eq('id', phoneLinkId)
.maybeSingle()
return (data as { phone_enc: string | null } | null) ?? null
}
/** Run one sweep pass. Never throws. */
export async function runSweep(supabase: SupabaseClient): Promise<SweepSummary> {
const summary: SweepSummary = {
reclaimedReceived: 0,
reclaimedProcessing: 0,
erroredMaxAttempts: 0,
finalizedAcks: 0,
expiredQuestions: 0,
clearedPins: 0,
}
const finalizeConversations = new Set<string>()
const now = Date.now()
// ── 1a. Stuck 'received' rows ──────────────────────────────
try {
const cutoff = new Date(now - RECEIVED_STUCK_MS).toISOString()
const { data } = await supabase
.from('whatsapp_messages')
.select(
'id, attempts, conversation_id, direction, message_type, sender_phone_hash, phone_link_id, correlation_id, raw_payload',
)
.eq('processing_status', 'received')
.lt('created_at', cutoff)
.order('created_at', { ascending: true })
.limit(BATCH)
for (const row of ((data ?? []) as StuckRow[])) {
if (row.attempts >= MAX_ATTEMPTS) {
await markMaxAttempts(supabase, row, 'received')
summary.erroredMaxAttempts++
continue
}
const outcome = await processInboundMessage(supabase, row.id)
summary.reclaimedReceived++
if (outcome.kind === 'media_processed' && outcome.conversationId) {
finalizeConversations.add(outcome.conversationId)
}
}
} catch (err) {
log.error('sweep: received re-claim failed', err)
}
// ── 1b. Stuck 'processing' rows (claimed, then the worker died) ──
try {
const cutoff = new Date(now - PROCESSING_STUCK_MS).toISOString()
const { data } = await supabase
.from('whatsapp_messages')
.select(
'id, attempts, conversation_id, direction, message_type, sender_phone_hash, phone_link_id, correlation_id, raw_payload',
)
.eq('processing_status', 'processing')
.lt('updated_at', cutoff)
.order('updated_at', { ascending: true })
.limit(BATCH)
for (const row of ((data ?? []) as StuckRow[])) {
if (row.attempts >= MAX_ATTEMPTS) {
await markMaxAttempts(supabase, row, 'processing')
summary.erroredMaxAttempts++
continue
}
// Guarded reset back to 'received'; processInboundMessage re-claims.
const { data: reset } = await supabase
.from('whatsapp_messages')
.update({ processing_status: 'received' })
.eq('id', row.id)
.eq('processing_status', 'processing')
.select('id')
.maybeSingle()
if (!reset) continue
const outcome = await processInboundMessage(supabase, row.id)
summary.reclaimedProcessing++
if (outcome.kind === 'media_processed' && outcome.conversationId) {
finalizeConversations.add(outcome.conversationId)
}
}
} catch (err) {
log.error('sweep: processing re-claim failed', err)
}
// ── 2a. Stale pending_ack (the debounce worker died pre-claim) ──
try {
const cutoff = new Date(now - ACK_STALE_MS).toISOString()
const { data } = await supabase
.from('whatsapp_conversations')
.select('id')
.eq('pending_ack', true)
.lt('debounce_until', cutoff)
.limit(BATCH)
for (const row of ((data ?? []) as { id: string }[])) {
finalizeConversations.add(row.id)
}
} catch (err) {
log.error('sweep: stale pending_ack scan failed', err)
}
// ── 2b. Unacked ingested rows whose winner died post-claim ──
try {
const cutoff = new Date(now - UNACKED_REARM_MS).toISOString()
const { data } = await supabase
.from('whatsapp_messages')
.select('conversation_id')
.eq('direction', 'inbound')
.eq('processing_status', 'done')
.is('acked_at', null)
.not('inbox_item_id', 'is', null)
.not('conversation_id', 'is', null)
.lt('updated_at', cutoff)
.limit(BATCH * 2)
const conversationIds = [
...new Set(((data ?? []) as { conversation_id: string }[]).map((r) => r.conversation_id)),
]
for (const conversationId of conversationIds) {
// pending_ack=false plus unacked rows is ALSO the state of a live
// claimant between claimAck and its acked_at stamp, and the 120s cutoff
// above measures the ROWS' done-stamp, not when the ack was claimed. So
// the conversation's own updated_at (which claimAck bumps) is the second
// condition: without it the sweep re-armed under a working finalize and
// a second combined ack went out.
await supabase
.from('whatsapp_conversations')
.update({ pending_ack: true, debounce_until: new Date().toISOString() })
.eq('id', conversationId)
.eq('pending_ack', false)
.lt('updated_at', cutoff)
finalizeConversations.add(conversationId)
}
} catch (err) {
log.error('sweep: unacked re-arm failed', err)
}
for (const conversationId of finalizeConversations) {
await finalizeBurst(supabase, conversationId)
summary.finalizedAcks++
}
// ── 3. Question TTL (48h) ──────────────────────────────────
try {
const { data } = await supabase
.from('whatsapp_conversations')
.select('*')
.neq('state', 'idle')
.limit(BATCH * 2)
for (const conversation of ((data ?? []) as WhatsAppConversation[])) {
const context = getContext(conversation)
const askedAt = context.pending_question?.asked_at
const expired =
askedAt == null || now - new Date(askedAt).getTime() > QUESTION_TTL_MS
if (!expired) continue
// Current question -> moved_to_app on the item (company questions have
// no item; their parked rows get the expired marker instead).
const pending = context.pending_question
if (pending?.inbox_item_id) {
await updateItemContext(supabase, pending.inbox_item_id, (itemContext) => ({
...itemContext,
pending_question:
itemContext.pending_question && itemContext.pending_question.status === 'open'
? { ...itemContext.pending_question, status: 'moved_to_app' }
: itemContext.pending_question,
}))
await appendQuestionHistory(supabase, {
inboxItemId: pending.inbox_item_id,
eventType: 'ChannelQuestionExpired',
questionType: pending.type,
})
}
// Company questions are the one kind whose expiry used to DESTROY work:
// the parked receipts were stamped company_choice_expired, a marker no
// code reads, so they never became Underlag rows and nothing ever told
// the user. The 24h service window is long gone at 48h and v1 sends no
// templates, so the honest recovery is to keep accepting a LATE answer:
// the rows stay staged and company_options stay in the context, which
// classify() treats as an open choice even in idle. Only when Meta has
// stopped serving the media (~30 days) does the marker become true.
let keepCompanyOptions = false
if (conversation.state === 'awaiting_company') {
const staleCutoff = new Date(now - STAGED_MEDIA_MAX_AGE_MS).toISOString()
await supabase
.from('whatsapp_messages')
.update({ error_message: COMPANY_CHOICE_EXPIRED })
.eq('conversation_id', conversation.id)
.eq('processing_status', 'skipped')
.eq('error_message', STAGED_AWAITING_COMPANY)
.lt('created_at', staleCutoff)
const { count: stillStaged } = await supabase
.from('whatsapp_messages')
.select('id', { count: 'exact', head: true })
.eq('conversation_id', conversation.id)
.eq('processing_status', 'skipped')
.eq('error_message', STAGED_AWAITING_COMPANY)
keepCompanyOptions = (stillStaged ?? 0) > 0
}
// Queued questions expire with the episode.
for (const queued of context.question_queue ?? []) {
await updateItemContext(supabase, queued.inbox_item_id, (itemContext) => ({
...itemContext,
pending_question: itemContext.pending_question ?? {
type: queued.type,
asked_at: new Date().toISOString(),
status: 'moved_to_app',
},
}))
}
const nextContext: ConversationContext = {
...context,
recent_questions: (context.recent_questions ?? []).map((q) =>
q.status === 'open' && q.inbox_item_id === pending?.inbox_item_id
? { ...q, status: 'moved_to_app' }
: q,
),
}
delete nextContext.pending_question
if (!keepCompanyOptions) delete nextContext.company_options
delete nextContext.question_queue
await supabase
.from('whatsapp_conversations')
.update({ state: 'idle', context: nextContext as Record<string, unknown> })
.eq('id', conversation.id)
.eq('state', conversation.state)
summary.expiredQuestions++
}
} catch (err) {
log.error('sweep: question TTL pass failed', err)
}
// ── 4. Expired company pins (8h sliding) ───────────────────
try {
const { data } = await supabase
.from('whatsapp_conversations')
.select('*')
.not('company_id', 'is', null)
.limit(BATCH * 2)
for (const conversation of ((data ?? []) as WhatsAppConversation[])) {
const context = getContext(conversation)
const expiresAt = context.pin_expires_at
if (expiresAt != null && new Date(expiresAt).getTime() > now) continue
// Guarded, and re-checked against fresh state: this loop awaits a
// network round trip per row, so a company choice applied in between
// used to be reverted (company_id nulled, the pre-choice context
// restored) and the question re-asked seconds after the user answered.
const cleared = await updateConversation(supabase, conversation, (_current, currentContext) => {
const stillExpired =
currentContext.pin_expires_at == null ||
new Date(currentContext.pin_expires_at).getTime() <= Date.now()
if (!stillExpired) return null
const nextContext: ConversationContext = { ...currentContext }
delete nextContext.pin_expires_at
delete nextContext.pin_source
return { company_id: null, context: nextContext }
})
if (cleared) summary.clearedPins++
}
} catch (err) {
log.error('sweep: pin expiry pass failed', err)
}
return summary
}