Fix upload concurrency, bound the request body, make backups consistent
Four audit findings on the photo feature. The upload route was async, so its blocking SQLite work ran on the event loop: under contention it stalled every other request for SQLite's busy timeout, not just its own. It also read the photo count without the write lock, so overlapping uploads all observed the same total and stored past the ceiling together. It is a synchronous endpoint now, running in the threadpool, taking BEGIN IMMEDIATE before re-checking the part, the ceiling and the position, and committing before it returns. Reverting either half makes the new test die with the same TimeoutError the audit reported. The 8MB cap protected nothing: Starlette parses and spools an entire multipart body before a route's dependencies run — before the login check — so the bytes were already on disk by the time anything rejected them, and an anonymous caller could make us write them. A plain ASGI middleware outside routing now refuses an over-large body first, and Caddy enforces the same ceiling at the edge. The documented backup captured the database and the photos at two different moments while the app stayed writable, so a photo deleted in between left the saved database pointing at a file the archive did not contain. tools/backup.sh stops the app for the few seconds the copy takes and verifies afterwards that every referenced photo is in the archive. Cleanup could destroy data rather than merely litter: prune-images could delete a file between an upload writing it and inserting its row, and deletions unlinked before their transaction committed. Pruning now ignores anything under an hour old unless forced, and deletes commit before unlinking — an orphaned file is recoverable, a row without its photo is not. check-images reports drift in both directions and fails only on the direction that loses data. Checks go from 271 to 291. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+58
-9
@@ -5,6 +5,7 @@ cannot help with — a forgotten password — and for the initial handover.
|
||||
|
||||
docker exec -it parts python -m app.admin set-password
|
||||
docker exec parts python -m app.admin show-status
|
||||
docker exec parts python -m app.admin check-images
|
||||
docker exec parts python -m app.admin prune-images
|
||||
"""
|
||||
|
||||
@@ -41,25 +42,72 @@ def clear_password(_argv):
|
||||
return 0
|
||||
|
||||
|
||||
def prune_images(_argv):
|
||||
# An upload writes its file before inserting the row, so for a moment a live
|
||||
# file legitimately has no row. Pruning anything younger than this would delete
|
||||
# a photo out from under a request that is still in flight.
|
||||
PRUNE_MIN_AGE_SECONDS = 3600
|
||||
|
||||
|
||||
def prune_images(argv):
|
||||
"""Delete image files with no row pointing at them.
|
||||
|
||||
Uploads write the file before the row, so an ill-timed crash can strand
|
||||
one. Nothing else creates orphans; this is a sweeper, not a routine step.
|
||||
Only files older than an hour, unless --all is given — and --all is only
|
||||
safe with the app stopped, because a file younger than its row is exactly
|
||||
what an in-flight upload looks like.
|
||||
"""
|
||||
import os
|
||||
import time
|
||||
|
||||
ignore_age = "--all" in argv
|
||||
with db.session() as conn:
|
||||
known = {
|
||||
os.path.basename(db.image_path(r["token"], r["mime"]))
|
||||
for r in conn.execute("SELECT token, mime FROM part_images").fetchall()
|
||||
}
|
||||
removed = 0
|
||||
removed = skipped = 0
|
||||
now = time.time()
|
||||
for name in os.listdir(db.IMAGE_DIR):
|
||||
if name not in known:
|
||||
os.remove(os.path.join(db.IMAGE_DIR, name))
|
||||
removed += 1
|
||||
print(f"{removed} orphaned image file(s) removed; {len(known)} kept.")
|
||||
if name in known:
|
||||
continue
|
||||
path = os.path.join(db.IMAGE_DIR, name)
|
||||
if not ignore_age and now - os.path.getmtime(path) < PRUNE_MIN_AGE_SECONDS:
|
||||
skipped += 1
|
||||
continue
|
||||
os.remove(path)
|
||||
removed += 1
|
||||
print(f"{removed} orphaned image file(s) removed; {len(known)} referenced file(s) kept.")
|
||||
if skipped:
|
||||
print(f"{skipped} skipped as too recent to be certain they are orphans "
|
||||
f"(stop the app and re-run with --all to include them).")
|
||||
return 0
|
||||
|
||||
|
||||
def check_images(_argv):
|
||||
"""Report drift in both directions between the database and the files."""
|
||||
import os
|
||||
|
||||
with db.session() as conn:
|
||||
rows = conn.execute("SELECT token, mime, part_id FROM part_images").fetchall()
|
||||
expected = {os.path.basename(db.image_path(r["token"], r["mime"])): r for r in rows}
|
||||
present = set(os.listdir(db.IMAGE_DIR))
|
||||
missing = sorted(set(expected) - present)
|
||||
orphans = sorted(present - set(expected))
|
||||
print(f"{len(expected)} referenced photo(s), {len(present)} file(s) on disk")
|
||||
for name in missing:
|
||||
print(f" MISSING FILE part {expected[name]['part_id']} {name}")
|
||||
for name in orphans:
|
||||
print(f" orphan file {name}")
|
||||
# A referenced photo with no file is data loss; a stray file is only clutter.
|
||||
return 1 if missing else 0
|
||||
|
||||
|
||||
def list_image_files(_argv):
|
||||
"""Every filename the database expects to exist. Used by the backup script."""
|
||||
import os
|
||||
|
||||
with db.session() as conn:
|
||||
for r in conn.execute("SELECT token, mime FROM part_images ORDER BY id").fetchall():
|
||||
print(os.path.basename(db.image_path(r["token"], r["mime"])))
|
||||
return 0
|
||||
|
||||
|
||||
@@ -79,7 +127,8 @@ def show_status(_argv):
|
||||
|
||||
|
||||
COMMANDS = {"set-password": set_password, "clear-password": clear_password,
|
||||
"show-status": show_status, "prune-images": prune_images}
|
||||
"show-status": show_status, "prune-images": prune_images,
|
||||
"check-images": check_images, "list-image-files": list_image_files}
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
|
||||
Reference in New Issue
Block a user