perf(streams,tasks,chat): cut storage row writes to legacy parity on the streaming hot path (#2191)
* perf(streams,tasks,chat): cut storage row writes to legacy parity on the streaming hot path Streams' append fence becomes a read: a Durable Object executes one synchronous block at a time, so state-check + chunk-tail read + INSERT is exactly as atomic as the old guarded count-bump UPDATE — and one row write per append instead of two. The stream row is written only at open and settle; settlement stamps the final cursor, and live cursors/liveness derive from the chunk log through one #tail helper. Reader loops use narrow state reads and terminate on an empty poll observed terminal in the same synchronous block. The chat adapter's retention sweep decides abandonment in two phases (coarse row cutoff, then one indexed chunk-tail read per candidate), the legacy migration imports rows complete (final count and last activity up front; chunk imports are bare INSERTs — 1+N writes instead of 1+2N), destroy() no longer flushes chunks it deletes in the same call, the cleanup alarm scans the table once, and the write-only _segmentIndex field is gone. Tasks amortizes claim refreshes to one write per half claim-slack of wall time instead of one per step, journals already-elapsed sleeps born-completed in a single INSERT, skips startup job-queue upserts that already match the run's deadline, settles a parked-run cancel in one row write, dedupes identical status writes, and only re-syncs the wake mirror when its settle write actually landed. The in-suite benchmark now pins exact adapter/legacy write parity (240 vs 240 rows for 20 turns x 100 chunks; previously 440) and models the two-phase sweep (239 -> 42 rows read). * perf(streams,tasks,lifecycle): bill one row per hot-path write via WITHOUT ROWID tables; bound replay memory Cloudflare bills rowsWritten, which counts index maintenance — and an ordinary rowid table's PRIMARY KEY is a hidden UNIQUE index, so every chunk append billed 2 rows while total_changes() (the parity benchmark's metric) reported 1. Measured empirically in workerd and pinned by a new write-accounting test: rowid composite-PK insert = 2, WITHOUT ROWID = 1, each touched explicit index +1, untouched indexes free. All five capability tables (streams, stream chunks, task runs, task steps, jobs — none released) go WITHOUT ROWID; the aperture's rowid ordering tiebreaks become stream_id. The task runs table drops its (state, next_at) index — a billed tax on every claim, refresh, and settle, paid only to speed the startup reconcile's one scan of a retention-bounded table. A 100-chunk chat turn now bills 13 rows vs 33 for the legacy schema (and 34 for the capability shape this PR started from). Replay memory is bounded too: the chat adapter's chunk replay becomes a generator over paged reads (readChunks replaces the aperture's readAll), so a reconnecting client holds one page of segments instead of the whole stored turn. * style: format rfc-streams.md * fix(tests): update schema DDL snapshot for WITHOUT ROWID tables The snapshot pins the verbatim sqlite_master text; the WITHOUT ROWID change altered three tables' stored DDL. The templates now end exactly at the ROWID keyword so the stored text stays clean of trailing whitespace. * fix(streams): keep the stream metadata table a rowid table for deterministic newest-first ordering Review flagged that WITHOUT ROWID on cf_agents_streams dropped the rowid insertion-order tiebreak: same-tag rows sharing a created_at millisecond (a retried turn's shape) would order by random nanoid, so recovery's latest-row lookups could pick an older turn. rowid IS the right tiebreak, and its hidden-index cost lands once per stream open — per turn, never per chunk — so the metadata table stays a rowid table while the chunk log keeps the WITHOUT ROWID billing win. The aperture queries get their rowid tiebreak back, a new test pins three same-tag same-created_at rows to insertion order, and the write-accounting test now documents the deliberate 3-billed-row stream open.
M
Matt committed
b40bc5bbbfb96151baf7c524834cae27bcae0b6a
Parent: 91e6f15
Committed by GitHub <noreply@github.com>
on 9/1/2026, 10:27:47 AM