* fix(whatsapp-inbox): erase the WhatsApp channel on account deletion whatsapp_phone_links relied on the auth.users ON DELETE CASCADE, but Accounted never deletes auth.users: account deletion is anonymize_user_account plus a ~100-year ban that keeps the auth row as a tombstone, so the cascade never fires and nothing revokes the link. After erasure the link stayed active with a decryptable phone_enc, lookupActiveLink kept resolving the number, and every further inbound message was persisted with body_text and the verbatim raw_payload while the bot kept replying: GDPR Art 17 plus continued collection with no lawful basis. The RPC is re-created verbatim from 20260724150000 with one added block that revokes and crypto-shreds the link, resets its conversation, nulls body_text/raw_payload on that link's messages and deletes outstanding link codes, plus a guarded repair pass for tombstones anonymized before this migration. Covered by a pg-real test that fails against the previous definition. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(whatsapp-inbox): pepper the link-code hash and bound code minting hashLinkCode stored a bare sha256 over CODE_ALPHABET^6 = 30^6 values behind a fixed 'AC-' prefix. The module cited the invite-token pattern, but invite tokens are 256-bit random; this space enumerates offline in about a second, so hashing at rest protected nothing. The sibling phone-crypto.ts already states the team's own threat model for a LARGER space ("a plain sha256 would be brute-forceable ... hence the pepper"), so link codes now hash through the same env-mandated pepper. /link/start was also an authenticated unbounded INSERT that left every earlier code valid. Minting now burns the caller's unused codes (the code the panel shows is the only one that works) and is capped per TTL window, with the route answering 429 instead of throwing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(whatsapp-inbox): harden the conversation layer against the review findings Pre-merge hardening of the unshipped chat layer. Every change below has a test that fails without it. Lifecycle and races: - conversation writes go through updateConversation(), an optimistic compare-and-set on updated_at (the trigger makes it a revision counter). The ack winner, the answer worker, the pin refresh and the sweep hold different claims, so blind whole-jsonb writes resurrected answered questions, wiped pending_question and dropped queue entries. - terminal markStatus writes are guarded on processing_status='processing' so a losing worker cannot overwrite the winner's 'done' and null its inbox_item_id. - the message -> inbox item path is idempotent: a pre-check plus a 23505 fallback adopt the item a concurrent worker created, instead of throwing after the WORM document is already committed. - PROCESSING_STUCK_MS 90s -> 5 min. The enforced step budget of one media row already exceeds 90s, so the sweep was re-claiming live workers. - sweep 2b re-arms only when the conversation itself has been quiet, not just the rows: pending_ack=false plus unacked rows is also the state of a live finalize, which produced a duplicate combined ack. - pin expiry re-checks against fresh state instead of writing back a stale whole context, which reverted company choices applied mid-pass. - askNextQueuedQuestion claims the pop before sending, so two answer workers cannot ask the same question twice. Company question: - the state is rolled back when the M6 send fails, so the next receipt re-asks instead of parking receipts behind a question nobody received. - applyCompanyChoice claims the open question (company_options) rather than the state: a double tap confirms once, a transient membership-query error is no longer read as "not a member", and a LATE answer still lands. - at the 48h TTL the parked receipts are kept, not discarded: options and staged rows survive so a late digit or tap still files them, and only rows past Meta's ~30-day media window get the terminal marker. - an out-of-range digit or a typed company name now gets the options repeated instead of silence or the "I cannot answer questions" reply. Inline dispositions: - stop/start/byt/company answers run their side effect BEFORE the terminal wamid row, with a SELECT pre-check for dedupe. Writing the row 'done' first made them at-most-once: a crash in between lost the action forever. Copy and answers: - acks state the extracted currency instead of labelling every total 'kr'. - M17 stops promising "about 10 minutes" when the daily quota tripped. - M18 is sent once per message tracked by the outbound row, so a file whose first attempt died still reaches the sender, including from the max-attempts path. - M11 no longer claims the number is disconnected: 'stopp' pauses, and muted senders now persist no chat content at all. - 'byt' is recognized in every state but awaiting_company (m6-confirm teaches the word, and it was being stored as answer data instead). - text sent while a re-send question is open is kept as a note on THAT receipt with the question left open, instead of binding to another receipt's question. - a quoted reply wins over the pending question and is appended when the quoted question is already answered, so corrections stop landing on the wrong receipt. - context answers keep raw_answer + answered_at like representation does. - finalizeBurst checks the send result: on failure it rolls the question back and leaves the rows unacked for the sweep. PII: - the sender's plaintext number is stripped from raw_payload before it is persisted; replies decrypt the link's phone_enc instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(whatsapp-inbox): record the erasure path and the hardening decisions RoPA gains the account-deletion row (immediate, not via the cron: the auth.users cascade never fires because the row is tombstoned) plus the two new security measures, and its "never in the clear" phone claim is now true of the stored payload. DECISIONS.md records the non-obvious calls: revoke-not-delete on erasure, commit-then-roll-back for the company question, keeping expired company choices answerable, the compare-and-set conversation write, effect-before-terminal-row for inline dispositions, honest M11 copy, and the raw_payload redaction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(whatsapp-inbox): stop the answer re-claim from following a confirm with M16 A worker that died after applying an answer and sending its confirmation leaves the row 'processing'. The sweep re-runs it, resolveAnswerTarget finds the question already answered, and the user got "I did not understand" immediately after the confirmation they had just received. The fallback is now first-attempt only. The catch comment claiming the sweep retries these rows is corrected too: 'error' is terminal for the sweep, and nothing on the answer path throws anyway (interpretChatAnswer degrades, sends never throw, supabase-js returns errors), so the catch is a programming-error net. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(whatsapp-inbox): drop the amount floor on the representation question The Swedish compliance review on #1340 caught a real error in the trigger rules: the representation question only fired above 150 kr, but the duty to document deltagare and syfte is what makes the expense deductible at all (BFL 5 kap 6-7 §) and it is not conditioned on any amount. The 300 kr per person figure I had in mind is the VAT-deduction base cap, a different rule. A 120 kr business lunch would have been booked with no participant trail, which is exactly the deduction Skatteverket denies later. Noise stays bounded by the triggers that were already there: the question fires only for receipt-shaped documents from restaurant, cafe or hotel merchants, at most once per receipt, twice per burst and six times per sender per day, and a single "nej" dismisses it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <jakob.wennberg@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com>
385 lines
15 KiB
TypeScript
385 lines
15 KiB
TypeScript
/**
|
|
* Per-minute crash-recovery sweep for the WhatsApp channel.
|
|
*
|
|
* The webhook 200s fast and defers all real work to after() invocations that
|
|
* can die with the serverless instance. Everything here is a re-derivation
|
|
* from durable state, so a lost invocation is a latency regression, never a
|
|
* lost message:
|
|
*
|
|
* 1. Re-claim whatsapp_messages stuck in 'received' (>60s) or 'processing'
|
|
* (>5 min, safely above the worst-case live worker); after MAX_ATTEMPTS
|
|
* they land in 'error' and the sender gets one M18.
|
|
* 2. Claim stale pending_ack conversations (debounce crash) and send the
|
|
* combined ack; re-arm conversations whose winner died after claiming
|
|
* but before sending (done rows left unacked).
|
|
* 3. Expire questions past the 48h TTL: conversation back to idle, the
|
|
* item's pending_question -> moved_to_app. NEVER sends anything: the 24h
|
|
* service window is long gone, and v1 sends no templates. Company
|
|
* questions keep their options and their parked receipts, so a late
|
|
* answer still files them (see the pass itself).
|
|
* 4. Clear expired 8h company pins.
|
|
*/
|
|
|
|
import type { SupabaseClient } from '@supabase/supabase-js'
|
|
import { createLogger } from '@/lib/logger'
|
|
import type { WhatsAppConversation, WhatsAppMessage } from '@/types'
|
|
import {
|
|
COMPANY_CHOICE_EXPIRED,
|
|
QUESTION_TTL_MS,
|
|
STAGED_AWAITING_COMPANY,
|
|
getContext,
|
|
resolveRecipient,
|
|
updateConversation,
|
|
type ConversationContext,
|
|
} from './conversation'
|
|
import { finalizeBurst, processInboundMessage, sendErrorNoticeOnce } from './process-inbound'
|
|
import { appendQuestionHistory, updateItemContext } from './item-context'
|
|
|
|
const log = createLogger('whatsapp-inbox/sweep')
|
|
|
|
const RECEIVED_STUCK_MS = 60 * 1000
|
|
/**
|
|
* A 'processing' row is only stuck if no live worker can still be on it.
|
|
* The enforced step budget of one media row is markRead (10s) + media lookup
|
|
* (10s) + download (30s) + Bedrock extraction (the cron route budgets 10-60s,
|
|
* with no short SDK timeout), under a maxDuration of 300s, and there is no
|
|
* heartbeat between the claim and the terminal write. 90s therefore re-claimed
|
|
* live workers on ordinary large PDFs and ran two of them on the same message.
|
|
* A crashed row waiting five minutes is a latency regression; two concurrent
|
|
* workers are a correctness problem.
|
|
*/
|
|
const PROCESSING_STUCK_MS = 5 * 60 * 1000
|
|
const ACK_STALE_MS = 60 * 1000
|
|
const UNACKED_REARM_MS = 120 * 1000
|
|
const MAX_ATTEMPTS = 3
|
|
const BATCH = 25
|
|
/** Staged receipts stay answerable while Meta still serves their media
|
|
* (~30 days). Past that the marker is honest: nothing can recover them. */
|
|
const STAGED_MEDIA_MAX_AGE_MS = 30 * 24 * 60 * 60 * 1000
|
|
|
|
export interface SweepSummary {
|
|
reclaimedReceived: number
|
|
reclaimedProcessing: number
|
|
erroredMaxAttempts: number
|
|
finalizedAcks: number
|
|
expiredQuestions: number
|
|
clearedPins: number
|
|
}
|
|
|
|
interface StuckRow {
|
|
id: string
|
|
attempts: number
|
|
conversation_id: string | null
|
|
direction: string
|
|
message_type: string
|
|
sender_phone_hash: string | null
|
|
phone_link_id: string | null
|
|
correlation_id: string | null
|
|
raw_payload: Record<string, unknown> | null
|
|
}
|
|
|
|
/**
|
|
* Park a row that ran out of attempts, and tell the sender once. Without the
|
|
* notice a file whose FIRST attempt died with the instance ends terminally
|
|
* with no ack and no error: the burst ack only lists ingested rows, so that
|
|
* receipt simply vanishes from the conversation.
|
|
*/
|
|
async function markMaxAttempts(
|
|
supabase: SupabaseClient,
|
|
row: StuckRow,
|
|
fromStatus: 'received' | 'processing',
|
|
): Promise<void> {
|
|
const { data: parked } = await supabase
|
|
.from('whatsapp_messages')
|
|
.update({ processing_status: 'error', error_message: 'Max attempts exceeded' })
|
|
.eq('id', row.id)
|
|
.eq('processing_status', fromStatus)
|
|
.select('id')
|
|
if (Array.isArray(parked) && parked.length === 0) return
|
|
if (row.message_type === 'text') return // M18 is about files
|
|
|
|
const link = row.phone_link_id
|
|
? await loadPhoneLink(supabase, row.phone_link_id)
|
|
: null
|
|
const to = resolveRecipient(row as unknown as WhatsAppMessage, link)
|
|
if (!to) return
|
|
await sendErrorNoticeOnce(supabase, {
|
|
to,
|
|
senderPhoneHash: row.sender_phone_hash,
|
|
phoneLinkId: row.phone_link_id,
|
|
conversationId: row.conversation_id,
|
|
correlationId: row.correlation_id,
|
|
})
|
|
}
|
|
|
|
async function loadPhoneLink(
|
|
supabase: SupabaseClient,
|
|
phoneLinkId: string,
|
|
): Promise<{ phone_enc: string | null } | null> {
|
|
const { data } = await supabase
|
|
.from('whatsapp_phone_links')
|
|
.select('phone_enc')
|
|
.eq('id', phoneLinkId)
|
|
.maybeSingle()
|
|
return (data as { phone_enc: string | null } | null) ?? null
|
|
}
|
|
|
|
/** Run one sweep pass. Never throws. */
|
|
export async function runSweep(supabase: SupabaseClient): Promise<SweepSummary> {
|
|
const summary: SweepSummary = {
|
|
reclaimedReceived: 0,
|
|
reclaimedProcessing: 0,
|
|
erroredMaxAttempts: 0,
|
|
finalizedAcks: 0,
|
|
expiredQuestions: 0,
|
|
clearedPins: 0,
|
|
}
|
|
const finalizeConversations = new Set<string>()
|
|
const now = Date.now()
|
|
|
|
// ── 1a. Stuck 'received' rows ──────────────────────────────
|
|
try {
|
|
const cutoff = new Date(now - RECEIVED_STUCK_MS).toISOString()
|
|
const { data } = await supabase
|
|
.from('whatsapp_messages')
|
|
.select(
|
|
'id, attempts, conversation_id, direction, message_type, sender_phone_hash, phone_link_id, correlation_id, raw_payload',
|
|
)
|
|
.eq('processing_status', 'received')
|
|
.lt('created_at', cutoff)
|
|
.order('created_at', { ascending: true })
|
|
.limit(BATCH)
|
|
for (const row of ((data ?? []) as StuckRow[])) {
|
|
if (row.attempts >= MAX_ATTEMPTS) {
|
|
await markMaxAttempts(supabase, row, 'received')
|
|
summary.erroredMaxAttempts++
|
|
continue
|
|
}
|
|
const outcome = await processInboundMessage(supabase, row.id)
|
|
summary.reclaimedReceived++
|
|
if (outcome.kind === 'media_processed' && outcome.conversationId) {
|
|
finalizeConversations.add(outcome.conversationId)
|
|
}
|
|
}
|
|
} catch (err) {
|
|
log.error('sweep: received re-claim failed', err)
|
|
}
|
|
|
|
// ── 1b. Stuck 'processing' rows (claimed, then the worker died) ──
|
|
try {
|
|
const cutoff = new Date(now - PROCESSING_STUCK_MS).toISOString()
|
|
const { data } = await supabase
|
|
.from('whatsapp_messages')
|
|
.select(
|
|
'id, attempts, conversation_id, direction, message_type, sender_phone_hash, phone_link_id, correlation_id, raw_payload',
|
|
)
|
|
.eq('processing_status', 'processing')
|
|
.lt('updated_at', cutoff)
|
|
.order('updated_at', { ascending: true })
|
|
.limit(BATCH)
|
|
for (const row of ((data ?? []) as StuckRow[])) {
|
|
if (row.attempts >= MAX_ATTEMPTS) {
|
|
await markMaxAttempts(supabase, row, 'processing')
|
|
summary.erroredMaxAttempts++
|
|
continue
|
|
}
|
|
// Guarded reset back to 'received'; processInboundMessage re-claims.
|
|
const { data: reset } = await supabase
|
|
.from('whatsapp_messages')
|
|
.update({ processing_status: 'received' })
|
|
.eq('id', row.id)
|
|
.eq('processing_status', 'processing')
|
|
.select('id')
|
|
.maybeSingle()
|
|
if (!reset) continue
|
|
const outcome = await processInboundMessage(supabase, row.id)
|
|
summary.reclaimedProcessing++
|
|
if (outcome.kind === 'media_processed' && outcome.conversationId) {
|
|
finalizeConversations.add(outcome.conversationId)
|
|
}
|
|
}
|
|
} catch (err) {
|
|
log.error('sweep: processing re-claim failed', err)
|
|
}
|
|
|
|
// ── 2a. Stale pending_ack (the debounce worker died pre-claim) ──
|
|
try {
|
|
const cutoff = new Date(now - ACK_STALE_MS).toISOString()
|
|
const { data } = await supabase
|
|
.from('whatsapp_conversations')
|
|
.select('id')
|
|
.eq('pending_ack', true)
|
|
.lt('debounce_until', cutoff)
|
|
.limit(BATCH)
|
|
for (const row of ((data ?? []) as { id: string }[])) {
|
|
finalizeConversations.add(row.id)
|
|
}
|
|
} catch (err) {
|
|
log.error('sweep: stale pending_ack scan failed', err)
|
|
}
|
|
|
|
// ── 2b. Unacked ingested rows whose winner died post-claim ──
|
|
try {
|
|
const cutoff = new Date(now - UNACKED_REARM_MS).toISOString()
|
|
const { data } = await supabase
|
|
.from('whatsapp_messages')
|
|
.select('conversation_id')
|
|
.eq('direction', 'inbound')
|
|
.eq('processing_status', 'done')
|
|
.is('acked_at', null)
|
|
.not('inbox_item_id', 'is', null)
|
|
.not('conversation_id', 'is', null)
|
|
.lt('updated_at', cutoff)
|
|
.limit(BATCH * 2)
|
|
const conversationIds = [
|
|
...new Set(((data ?? []) as { conversation_id: string }[]).map((r) => r.conversation_id)),
|
|
]
|
|
for (const conversationId of conversationIds) {
|
|
// pending_ack=false plus unacked rows is ALSO the state of a live
|
|
// claimant between claimAck and its acked_at stamp, and the 120s cutoff
|
|
// above measures the ROWS' done-stamp, not when the ack was claimed. So
|
|
// the conversation's own updated_at (which claimAck bumps) is the second
|
|
// condition: without it the sweep re-armed under a working finalize and
|
|
// a second combined ack went out.
|
|
await supabase
|
|
.from('whatsapp_conversations')
|
|
.update({ pending_ack: true, debounce_until: new Date().toISOString() })
|
|
.eq('id', conversationId)
|
|
.eq('pending_ack', false)
|
|
.lt('updated_at', cutoff)
|
|
finalizeConversations.add(conversationId)
|
|
}
|
|
} catch (err) {
|
|
log.error('sweep: unacked re-arm failed', err)
|
|
}
|
|
|
|
for (const conversationId of finalizeConversations) {
|
|
await finalizeBurst(supabase, conversationId)
|
|
summary.finalizedAcks++
|
|
}
|
|
|
|
// ── 3. Question TTL (48h) ──────────────────────────────────
|
|
try {
|
|
const { data } = await supabase
|
|
.from('whatsapp_conversations')
|
|
.select('*')
|
|
.neq('state', 'idle')
|
|
.limit(BATCH * 2)
|
|
for (const conversation of ((data ?? []) as WhatsAppConversation[])) {
|
|
const context = getContext(conversation)
|
|
const askedAt = context.pending_question?.asked_at
|
|
const expired =
|
|
askedAt == null || now - new Date(askedAt).getTime() > QUESTION_TTL_MS
|
|
if (!expired) continue
|
|
|
|
// Current question -> moved_to_app on the item (company questions have
|
|
// no item; their parked rows get the expired marker instead).
|
|
const pending = context.pending_question
|
|
if (pending?.inbox_item_id) {
|
|
await updateItemContext(supabase, pending.inbox_item_id, (itemContext) => ({
|
|
...itemContext,
|
|
pending_question:
|
|
itemContext.pending_question && itemContext.pending_question.status === 'open'
|
|
? { ...itemContext.pending_question, status: 'moved_to_app' }
|
|
: itemContext.pending_question,
|
|
}))
|
|
await appendQuestionHistory(supabase, {
|
|
inboxItemId: pending.inbox_item_id,
|
|
eventType: 'ChannelQuestionExpired',
|
|
questionType: pending.type,
|
|
})
|
|
}
|
|
// Company questions are the one kind whose expiry used to DESTROY work:
|
|
// the parked receipts were stamped company_choice_expired, a marker no
|
|
// code reads, so they never became Underlag rows and nothing ever told
|
|
// the user. The 24h service window is long gone at 48h and v1 sends no
|
|
// templates, so the honest recovery is to keep accepting a LATE answer:
|
|
// the rows stay staged and company_options stay in the context, which
|
|
// classify() treats as an open choice even in idle. Only when Meta has
|
|
// stopped serving the media (~30 days) does the marker become true.
|
|
let keepCompanyOptions = false
|
|
if (conversation.state === 'awaiting_company') {
|
|
const staleCutoff = new Date(now - STAGED_MEDIA_MAX_AGE_MS).toISOString()
|
|
await supabase
|
|
.from('whatsapp_messages')
|
|
.update({ error_message: COMPANY_CHOICE_EXPIRED })
|
|
.eq('conversation_id', conversation.id)
|
|
.eq('processing_status', 'skipped')
|
|
.eq('error_message', STAGED_AWAITING_COMPANY)
|
|
.lt('created_at', staleCutoff)
|
|
const { count: stillStaged } = await supabase
|
|
.from('whatsapp_messages')
|
|
.select('id', { count: 'exact', head: true })
|
|
.eq('conversation_id', conversation.id)
|
|
.eq('processing_status', 'skipped')
|
|
.eq('error_message', STAGED_AWAITING_COMPANY)
|
|
keepCompanyOptions = (stillStaged ?? 0) > 0
|
|
}
|
|
// Queued questions expire with the episode.
|
|
for (const queued of context.question_queue ?? []) {
|
|
await updateItemContext(supabase, queued.inbox_item_id, (itemContext) => ({
|
|
...itemContext,
|
|
pending_question: itemContext.pending_question ?? {
|
|
type: queued.type,
|
|
asked_at: new Date().toISOString(),
|
|
status: 'moved_to_app',
|
|
},
|
|
}))
|
|
}
|
|
|
|
const nextContext: ConversationContext = {
|
|
...context,
|
|
recent_questions: (context.recent_questions ?? []).map((q) =>
|
|
q.status === 'open' && q.inbox_item_id === pending?.inbox_item_id
|
|
? { ...q, status: 'moved_to_app' }
|
|
: q,
|
|
),
|
|
}
|
|
delete nextContext.pending_question
|
|
if (!keepCompanyOptions) delete nextContext.company_options
|
|
delete nextContext.question_queue
|
|
await supabase
|
|
.from('whatsapp_conversations')
|
|
.update({ state: 'idle', context: nextContext as Record<string, unknown> })
|
|
.eq('id', conversation.id)
|
|
.eq('state', conversation.state)
|
|
summary.expiredQuestions++
|
|
}
|
|
} catch (err) {
|
|
log.error('sweep: question TTL pass failed', err)
|
|
}
|
|
|
|
// ── 4. Expired company pins (8h sliding) ───────────────────
|
|
try {
|
|
const { data } = await supabase
|
|
.from('whatsapp_conversations')
|
|
.select('*')
|
|
.not('company_id', 'is', null)
|
|
.limit(BATCH * 2)
|
|
for (const conversation of ((data ?? []) as WhatsAppConversation[])) {
|
|
const context = getContext(conversation)
|
|
const expiresAt = context.pin_expires_at
|
|
if (expiresAt != null && new Date(expiresAt).getTime() > now) continue
|
|
// Guarded, and re-checked against fresh state: this loop awaits a
|
|
// network round trip per row, so a company choice applied in between
|
|
// used to be reverted (company_id nulled, the pre-choice context
|
|
// restored) and the question re-asked seconds after the user answered.
|
|
const cleared = await updateConversation(supabase, conversation, (_current, currentContext) => {
|
|
const stillExpired =
|
|
currentContext.pin_expires_at == null ||
|
|
new Date(currentContext.pin_expires_at).getTime() <= Date.now()
|
|
if (!stillExpired) return null
|
|
const nextContext: ConversationContext = { ...currentContext }
|
|
delete nextContext.pin_expires_at
|
|
delete nextContext.pin_source
|
|
return { company_id: null, context: nextContext }
|
|
})
|
|
if (cleared) summary.clearedPins++
|
|
}
|
|
} catch (err) {
|
|
log.error('sweep: pin expiry pass failed', err)
|
|
}
|
|
|
|
return summary
|
|
}
|