fix(storage): a torn WAL never locks the workspace; bound failed recovery copies
A WAL sidecar shorter than its 32-byte header (a crash before the header was fully written) made every command fail on the new startup index recovery: 'Refusing schema preflight because the 20-byte WAL header is truncated'. doctor --repair refused on the same preflight, and each failed command copied the whole database family into a new .br_recovery run. 0.6.0 handled the same file normally. SQLite reads nothing from a WAL shorter than its header, so it holds no committed frame and the main database is the complete state: - Observational schema preflight treats a torn WAL like an empty one, byte-neutrally. Index recovery still refuses it. - The missing-index probes (startup and read-only snapshot) no longer treat a torn WAL as an index to rebuild. - Startup moves a torn WAL and its shared index into .br_recovery under the held family authority and only as the verified sole opener (with peers it leaves the family alone), and reports torn_wal_quarantined on stderr. - Automatic index recovery keeps one retained pre-state per incident: a byte-identical family whose automatic recovery already failed reports the retained copy instead of copying the family again. The explicit 'br doctor migrate-schema recover' still retries. Tests: e2e_torn_wal_startup covers 1- and 31-byte torn WALs with and without SHM (reads, writes, retained bytes, no new artifacts afterwards), doctor --repair on a torn WAL, empty and complete 32-byte WALs left alone, and repeated refused recoveries not growing .br_recovery. A unit test pins the preflight semantics for 1/20/31/32-byte WALs. The restored e2e_doctor_detects_and_quarantines_anomalous_wal_sidecar now passes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
J
Jeff Emanuel committed
b900260fb0b9f3d556ff0817e9fc0d900268a49f
Parent: adebe3f