Running it against the live service exposed two gaps. It announced "App
restarted" the moment the container existed, so the very next request got a 502
from an app that wasn't listening yet; it now polls /healthz until the app is
actually serving. And it only checked that the archive contained the right
files — it now opens the archived database and runs SQLite's integrity_check,
because an archive that exists but won't restore is the worst kind of backup.
End to end on the server, with a throwaway part and photo inserted to exercise
the verification path: 2.6 seconds of downtime, archive verified, service
answering immediately on return, throwaway data removed afterwards.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four audit findings on the photo feature.
The upload route was async, so its blocking SQLite work ran on the event loop:
under contention it stalled every other request for SQLite's busy timeout, not
just its own. It also read the photo count without the write lock, so
overlapping uploads all observed the same total and stored past the ceiling
together. It is a synchronous endpoint now, running in the threadpool, taking
BEGIN IMMEDIATE before re-checking the part, the ceiling and the position, and
committing before it returns. Reverting either half makes the new test die with
the same TimeoutError the audit reported.
The 8MB cap protected nothing: Starlette parses and spools an entire multipart
body before a route's dependencies run — before the login check — so the bytes
were already on disk by the time anything rejected them, and an anonymous
caller could make us write them. A plain ASGI middleware outside routing now
refuses an over-large body first, and Caddy enforces the same ceiling at the
edge.
The documented backup captured the database and the photos at two different
moments while the app stayed writable, so a photo deleted in between left the
saved database pointing at a file the archive did not contain. tools/backup.sh
stops the app for the few seconds the copy takes and verifies afterwards that
every referenced photo is in the archive.
Cleanup could destroy data rather than merely litter: prune-images could delete
a file between an upload writing it and inserting its row, and deletions
unlinked before their transaction committed. Pruning now ignores anything under
an hour old unless forced, and deletes commit before unlinking — an orphaned
file is recoverable, a row without its photo is not. check-images reports drift
in both directions and fails only on the direction that loses data.
Checks go from 271 to 291.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>