Files
accounted/app/api/documents/[id]/integrity/route.ts
T
Jakob WennbergandClaude Fable 5 968161b42b fix(documents): read attachments with service client so colleague uploads open (#1207)
The documents bucket SELECT policy only covers the uploader's own folder
(documents/{uid}/...), but document_attachments rows are company-scoped.
Every surface that touched storage with the user-bound client therefore
failed for attachments uploaded by another member of the same company
(colleague uploads, email-inbox ingest attributed to the company creator):

- GET /api/documents/:id 500ed with "Failed to create download URL", so
  viewing a bilaga on a verifikat or supplier invoice was broken for
  every member except the uploader (support case: Odin Aero, where all
  40 documents live in the owner's folder and the second member could
  open none of them).
- GET /api/documents/:id/integrity 500ed the same way.
- POST /api/documents/:id/verify failed the storage download.
- invoice-inbox retry-extraction could not download the attachment.
- cloud-backup user-triggered syncs silently dropped colleague-uploaded
  documents from the Drive archive (manifest rows flipped to 'error').

Fix: authorize on the user client (RLS + explicit company filter, plus
the membership check where present), then do the storage read with the
service-role client. This is the pattern the inline proxy route and the
v1 download route already use; these five call sites were left behind.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-26 12:34:11 +02:00

114 lines
4.3 KiB
TypeScript

import { NextResponse } from 'next/server'
import { z } from 'zod'
import { requireAuth } from '@/lib/auth/require-auth'
import { createServiceClient } from '@/lib/supabase/server'
import { validateDocumentMagicBytes } from '@/lib/core/documents/document-service'
import { createLogger } from '@/lib/logger'
const log = createLogger('documents.integrity')
const ParamsSchema = z.object({ id: z.string().uuid() })
/**
* GET /api/documents/:id/integrity
*
* Probes the actual stored bytes against the declared MIME type. Used by the
* Bilagor modal to surface a clear "this file is corrupt: please re-upload"
* warning instead of relying on the browser's PDF viewer error UI, which
* only fires after the user has already tried to view the file.
*
* Some legacy MCP uploads landed with non-PDF bytes under
* `mime_type = 'application/pdf'` because magic-byte validation was added
* after those rows were written. This endpoint lets the UI detect and steer
* the user toward replacing them.
*
* Response shape is intentionally minimal: { valid: boolean } only. The
* reason for an invalid result is logged server-side rather than returned
* to the client to avoid information disclosure (V1.2.5 / GDPR Art 25(2))
* and to keep this from being a probe surface for storage internals.
*/
export async function GET(
_request: Request,
{ params }: { params: Promise<{ id: string }> }
) {
const { user, supabase, error } = await requireAuth()
if (error) return error
const rawParams = await params
const parsed = ParamsSchema.safeParse(rawParams)
if (!parsed.success) {
return NextResponse.json({ error: 'Invalid document id' }, { status: 400 })
}
const { id } = parsed.data
// Filter to the current version. The integrity check is meaningful only
// on the live file; superseded versions are archived bytes and should
// not be re-probed (they're already preserved in the version chain
// exactly as uploaded).
const { data: doc, error: docError } = await supabase
.from('document_attachments')
.select('id, company_id, mime_type, storage_path')
.eq('id', id)
.eq('is_current_version', true)
.single()
if (docError || !doc) {
return NextResponse.json({ error: 'Document not found' }, { status: 404 })
}
// Tenant membership: even with the user-scoped supabase client below,
// we want a clear 404 rather than relying on a storage-layer RLS deny
// (which can present as a generic error). RLS on document_attachments
// is the primary control; this is defense in depth.
const { data: membership } = await supabase
.from('company_members')
.select('company_id')
.eq('company_id', doc.company_id)
.eq('user_id', user.id)
.maybeSingle()
if (!membership) {
return NextResponse.json({ error: 'Document not found' }, { status: 404 })
}
if (!doc.mime_type) {
return NextResponse.json({ data: { valid: true } })
}
// Download via the service-role client: the storage SELECT policy only
// covers the uploader's own folder (documents/{uid}/...), so the
// user-scoped client cannot read colleague-uploaded files even within
// the same company. The document_attachments RLS fetch plus the explicit
// membership check above are the authorization (same model as the
// inline proxy route).
const serviceClient = createServiceClient()
const { data: blob, error: downloadError } = await serviceClient.storage
.from('documents')
.download(doc.storage_path)
if (downloadError || !blob) {
log.error('storage download failed for integrity check', downloadError as Error, {
documentId: id,
companyId: doc.company_id,
})
return NextResponse.json({ error: 'Integrity check unavailable' }, { status: 500 })
}
// Only the first 16 bytes are needed for magic-byte detection (PDF/PNG
// use ≤8, WebP needs 12). Trimming here doesn't change bandwidth (the
// full blob is already downloaded) but it makes the intent explicit and
// keeps memory churn off the hot path for large PDFs.
const headerBuffer = await blob.slice(0, 16).arrayBuffer()
const magicError = validateDocumentMagicBytes(headerBuffer, doc.mime_type)
if (magicError) {
log.warn('document failed magic-byte integrity check', {
documentId: id,
companyId: doc.company_id,
reason: magicError,
})
}
return NextResponse.json({ data: { valid: magicError === null } })
}