SIGN IN SIGN UP

fix(storage): a torn WAL never locks the workspace; bound failed recovery copies

A WAL sidecar shorter than its 32-byte header (a crash before the header
was fully written) made every command fail on the new startup index
recovery: 'Refusing schema preflight because the 20-byte WAL header is
truncated'. doctor --repair refused on the same preflight, and each failed
command copied the whole database family into a new .br_recovery run.
0.6.0 handled the same file normally.

SQLite reads nothing from a WAL shorter than its header, so it holds no
committed frame and the main database is the complete state:
- Observational schema preflight treats a torn WAL like an empty one,
  byte-neutrally. Index recovery still refuses it.
- The missing-index probes (startup and read-only snapshot) no longer treat
  a torn WAL as an index to rebuild.
- Startup moves a torn WAL and its shared index into .br_recovery under the
  held family authority and only as the verified sole opener (with peers it
  leaves the family alone), and reports torn_wal_quarantined on stderr.
- Automatic index recovery keeps one retained pre-state per incident: a
  byte-identical family whose automatic recovery already failed reports the
  retained copy instead of copying the family again. The explicit
  'br doctor migrate-schema recover' still retries.

Tests: e2e_torn_wal_startup covers 1- and 31-byte torn WALs with and without
SHM (reads, writes, retained bytes, no new artifacts afterwards), doctor
--repair on a torn WAL, empty and complete 32-byte WALs left alone, and
repeated refused recoveries not growing .br_recovery. A unit test pins the
preflight semantics for 1/20/31/32-byte WALs. The restored
e2e_doctor_detects_and_quarantines_anomalous_wal_sidecar now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
J
Jeff Emanuel committed
b900260fb0b9f3d556ff0817e9fc0d900268a49f
Parent: adebe3f