Files
accounted/app/api/sandbox/cleanup/cron/route.ts
T
bjornbergenheimandClaude Opus 5 43a71aec3c fix(supabase): stop server clients leaking a 30s refresh ticker per request (#1612)
* fix(supabase): stop server clients leaking a 30s refresh ticker per request

`autoRefreshToken` defaults to true in supabase-js, and off-browser
@supabase/auth-js starts the refresh ticker unconditionally:

    // in non-browser environments the refresh token ticker runs always
    this.startAutoRefresh()

That is a setInterval firing every 30 s. It calls unref(), so the process
still exits, tests pass, and Vercel never notices because the process is
torn down long before the tickers accumulate. But unref() does not make a
timer collectable: it stays registered in the event loop and remains a GC
root for its callback, which closes over the GoTrueClient, the
SupabaseClient, and the whole request scope around it.

A long-running self-hosted instance therefore leaks one timer plus one
entire request graph (socket, IncomingMessage, ServerResponse, headers,
route context: ~100 kB) per client constructed. One died of "JavaScript
heap out of memory" after 42 h, the last 24 of them completely idle. The
heap snapshot showed 445 retained request graphs and ~1050 Timeouts in
the 30 000 ms bucket, retained via `autoRefreshTicker`, and the rate
matched the traffic exactly: the Docker healthcheck polls /api/health
every 30 s and the webhook dispatch cron runs every minute, so
3 clients/min x 148 min = 444.

- new lib/supabase/service-client.ts: createServiceRoleClient() applies
  SERVER_AUTH_OPTIONS, spread LAST so a caller passing its own auth block
  cannot re-enable the ticker
- 22 call sites migrated; only booking-templates/sync/cron had ever
  passed the options itself
- guard 9 in no-new-antipatterns.mjs fails CI on any new value import of
  supabase-js's createClient outside the wrapper; type-only imports are
  fine. Verified to fail on a deliberate regression and pass once fixed
- browser clients untouched: a signed-in tab genuinely needs the refresh,
  and lib/supabase/client.ts is built on createBrowserClient anyway

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(checks): catch namespace imports in the leaky-supabase-client guard

The guard only matched named imports, so

    import * as sb from '@supabase/supabase-js'
    sb.createClient(url, key)

reached createClient through member access without ever naming it, and
passed. Verified against the real script before and after: the shape is
flagged now, and `import type * as sb` still passes.

Namespace value imports are treated as leaky outright rather than tracking
member access, which keeps the check a regex over source text with no new
dependency.

Review also suggested excluding *.test.tsx alongside *.test.ts. Skipped: the
repo has no .test.tsx files, and all four sibling checks in this file use
`.test.ts`. Diverging in one of them would read as an accident; if such files
appear, all four should change together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 14:30:17 +02:00

89 lines
3.4 KiB
TypeScript

import { createServiceRoleClient } from '@/lib/supabase/service-client'
import { NextResponse } from 'next/server'
import { withCronContext } from '@/lib/api/with-cron-context'
import { errorResponse, errorResponseFromCode } from '@/lib/errors/get-structured-error'
/**
* GET /api/sandbox/cleanup/cron: daily 04:00 UTC.
* Removes expired sandbox users (>24h old).
*
* Every PostgREST statement runs under authenticator's statement_timeout of
* 8s, and a function-level SET statement_timeout does NOT lift it (the timer
* arms when the top-level statement starts; verified empirically on prod
* 2026-08-07, same finding as the SIE import RPCs). One teardown costs
* ~220ms with the account_id index, so the run loops SMALL batches: each
* rpc() call is its own statement with its own 8s window, and the loop
* stops when a batch makes no progress (nothing left, or only failing
* users remain) or the route's time budget nears. Capacity per night is
* MAX_BATCHES * BATCH_LIMIT users; the nightly intake is a small fraction
* of that.
*/
export const maxDuration = 300
const BATCH_LIMIT = 10
const MAX_BATCHES = 25
const TIME_BUDGET_MS = 240_000
export const GET = withCronContext('cron.sandbox_cleanup', async (_request, ctx) => {
const supabaseUrl = process.env.NEXT_PUBLIC_SUPABASE_URL
const supabaseServiceKey = process.env.SUPABASE_SERVICE_ROLE_KEY
if (!supabaseUrl || !supabaseServiceKey) {
return errorResponseFromCode('INTERNAL_ERROR', ctx.log, {
requestId: ctx.requestId,
details: { reason: 'Missing Supabase configuration' },
})
}
const supabase = createServiceRoleClient(supabaseUrl, supabaseServiceKey)
const started = Date.now()
const totals = { cleaned: 0, failed: 0, orphans_removed: 0, batches: 0 }
for (let i = 0; i < MAX_BATCHES; i++) {
if (Date.now() - started > TIME_BUDGET_MS) break
const { data, error } = await supabase.rpc('cleanup_expired_sandbox_users', {
p_max_age_hours: 24,
p_limit: BATCH_LIMIT,
})
if (error) {
ctx.log.error('sandbox cleanup rpc failed', { error, ...totals })
return errorResponse(error, ctx.log, { requestId: ctx.requestId })
}
// Migration 20260807130000 changed the RPC's return from a bare integer
// to a {cleaned, failed, orphans_removed} summary; accept both shapes so
// deploy/migration ordering cannot break the cron.
const batch =
typeof data === 'number'
? { cleaned: data, failed: 0, orphans_removed: 0 }
: {
cleaned: Number(data?.cleaned ?? 0),
failed: Number(data?.failed ?? 0),
orphans_removed: Number(data?.orphans_removed ?? 0),
}
totals.cleaned += batch.cleaned
totals.failed += batch.failed
totals.orphans_removed += batch.orphans_removed
totals.batches += 1
// No progress means only permanently-failing users (retried nightly and
// reported below) or an empty backlog: looping further would spin on the
// same rows.
if (batch.cleaned + batch.orphans_removed === 0) break
}
// Per-user failures used to be swallowed as Postgres WARNINGs, which is how
// the cleanup sat broken for months; surface them at error level instead.
if (totals.failed > 0) {
ctx.log.error('sandbox cleanup completed with failures', totals)
} else {
ctx.log.info('sandbox cleanup summary', totals)
}
return NextResponse.json({ success: true, ...totals })
})