Resolve the backup volume from the container, and verify photo sizes

Three audit items on the backup path.

The script matched volume names by pattern and took the first hit, so a stale
or restored volume could be backed up instead of the live one — and every
verification step would then faithfully confirm the wrong database. It now asks
the container what is mounted at /data and refuses ambiguity. Tested against a
decoy volume that the old pattern would have matched first.

check-images treated a photo as healthy if a file with the right name existed,
so a truncated or partially restored file passed. It compares each file against
the byte count its row records now; a one-byte stand-in for a 123KB photo is
reported as WRONG SIZE and exits non-zero. The backup runs the same check
against the stopped volume and exits 2 when the source was already damaged —
still writing the archive, because a faithful copy of imperfect data is worth
having, but saying so.

The restart trap was installed after the app had already been stopped, so an
interrupt in between could leave the service down with nothing to bring it
back. The trap goes in first now, covers INT and TERM as well as EXIT, and
records whether the container was running beforehand so a backup of an
already-stopped app leaves it stopped.

Verified on the live host: healthy source exits 0, damaged source exits 2 with
the archive still written and verified, decoy volume correctly ignored, service
answering immediately afterwards, and the test rows removed.

Checks go from 290 to 294.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Jay
2026-08-25 10:12:45 -04:00
parent ccb3da51ad
commit 8d87f1c13d
4 changed files with 143 additions and 45 deletions
+21
View File
@@ -524,6 +524,27 @@ with TestClient(app) as client:
client.get(f"/api/parts/{sweep_part}/images/{kept}").status_code == 200)
check("check-images is clean when nothing has drifted",
_admin_img.main(["check-images"]) == 0)
# Presence is not health. A truncated file has the right name and the wrong
# contents, which a filename-only check waves through.
kept_path = os.path.join(_imgdb.IMAGE_DIR,
os.path.basename(_imgdb.image_path(kept, "image/png")))
full = open(kept_path, "rb").read()
with open(kept_path, "wb") as fh:
fh.write(full[:1])
check("check-images catches a truncated photo",
_admin_img.main(["check-images"]) == 1)
with open(kept_path, "wb") as fh:
fh.write(full)
check("check-images is clean again once it is restored",
_admin_img.main(["check-images"]) == 0)
with open(kept_path, "wb") as fh:
fh.write(full + b"trailing junk")
check("check-images catches a photo that grew",
_admin_img.main(["check-images"]) == 1)
with open(kept_path, "wb") as fh:
fh.write(full)
# A referenced photo whose file vanished is the direction that matters.
os.remove(os.path.join(_imgdb.IMAGE_DIR, os.listdir(_imgdb.IMAGE_DIR)[0]))
check("check-images reports a missing file as a failure",