fix(core): drain step stream writes before completion (#3941)
## Summary & Motivation ### Situation - Workflow stream writers are expected to call `releaseLock()` when a step finishes writing so another step can acquire the stream. - The runtime observes that release and drains the server sink, but `step_completed` currently races the overall stream operation against 500ms. - A slow PUT can therefore continue under `waitUntil` while the next step starts and reads a stale tail. - Release can also happen while native `writer.write()` promises remain unsettled, leaving frames upstream of the server sink when a naive drain runs. - Writers intentionally kept locked must remain non-blocking so producer and consumer steps can overlap. ### Fix - Treat a writer released before step return as an implicit durable handoff boundary. - At step end, acquire the unlocked stream with a temporary writer and enqueue an internal checkpoint behind all writes queued by the released writer. - Once the checkpoint crosses serialization, wait for those frames to reach the server sink and drain the group-commit PUT before `step_completed`. - If the writer remains locked, do not wait for durability; preserve the existing 500ms inline-loop heuristic and background `waitUntil` lifecycle. - Drain failures or the 30-second safety timeout fail/retry the step. Client disconnect errors remain non-fatal. ## Test Plan - Covers released and held locks, release with unsettled writes, delayed first writer acquisition, forwarded writable arguments, drain timeout/failure, and multiple streams. - `pnpm --filter @workflow/core build` - `pnpm --filter @workflow/core typecheck` - `pnpm --filter @workflow/core test` — 2,405 passed, 3 expected failures, 1 skipped
A
Alex Langenfeld committed
c09c1bb6ea743d6c8e24574aa8dd5516fb287548
Parent: 17bd649
Committed by GitHub <noreply@github.com>
on 9/11/2026, 2:21:38 PM