SIGN IN SIGN UP

feat(examples): add Worker-native Codex harness (#2208)

* feat(examples): add Worker-native Codex harness

* chore(codex-harness): drop smoke and test harness, fix restart recovery

The example is the demo: remove smoke.mjs, the node:test adapter suite and
its tsx tsconfig, the /health, results and abort routes, and the
"static-wasm" mode/metadata leftovers. Remove the nx `build` script so CI
does not need a Rust toolchain; `start` and `deploy` still build the kernel.

Fix recovery after a Durable Object restart: a stored terminal action now
settles the operation instead of failing it, a replayed terminal operation
closes its stream, and an identical resubmit of a still-queued operation
re-syncs its Tasks wake. Join text blocks before trimming so boundary
whitespace survives, and stop gating UI completion on the demo file.

Trim deployment URLs and smoke timing tables from the RFCs.

Claude-Session: https://claude.ai/code/session_01KEFnjoMnZBMewsuD9qxGrL

* refactor(codex-harness): journal effects through Tasks step.do

Drop the harness's own cf_codex_effects table and run each model or
Workspace effect as a named Tasks step keyed by its effect ID. Tasks replays
settled results on later attempts and hands the step's AbortSignal to the
model call, so cancellation reaches in-flight work. Document the interrupted
step policy in the RFC: tools re-run idempotently, model rounds re-issue and
accept a rare duplicate call.

Claude-Session: https://claude.ai/code/session_01KEFnjoMnZBMewsuD9qxGrL

* feat(codex-harness): serve the UI over the WebSockets capability

Replace the HTTP polling routes with a WebSocket protocol on the
WebSockets capability: a session snapshot on connect, subscribe to
replay-then-tail an operation's Streams log, and submit and restart
commands. Operation state changes broadcast to every connection.

The client connects with useAgent from agents/react and a small
useCodexSession hook layers the protocol on that socket, so the whole
transcript and every operation's events reload from durable state on
reconnect. The worker routes with routeAgentRequest.

Claude-Session: https://claude.ai/code/session_01KEFnjoMnZBMewsuD9qxGrL

* chore(codex-harness): rename the worker and pin a dev inspector port

Deploy as codex-harness-example, give the Vite dev server its own inspector
port so it can run beside the other harness examples, and document setting
CLOUDFLARE_ACCOUNT_ID for logins with several accounts.

* fix(codex-harness): dedupe React so the client renders

agents/react resolved a second React copy from the agents package, which
broke every hook call in the browser.

* chore(codex-harness): shorten the empty-state copy

* fix(codex-harness): stop the snapshot and subscribe loop in the client

A closed stream ends as soon as it is replayed, and every stream end
asked for a snapshot that resubscribed to every operation. Subscribe to
each operation once per connection and refresh only after a live tail
ends, and scroll to the bottom only when content is added.

* feat(codex-harness): stress rig, payload caps, and transcript-size fixes

Add a synthetic-model stress worker (src/stress) with a driver and a CDP
heap profiler (scripts/) so the kernel, Tasks, Streams, and SQLite paths
can be pushed without Workers AI.

Fix what the runs broke:
- the 16-transition cap failed any turn with more than six tool calls;
  cap model rounds at 24 and transitions at 256 instead
- the stored model action repeated the checkpoint's input, halving the
  transcript a turn could hold before SQLITE_TOOBIG; store it without
  the input and rehydrate on read
- a 1 MB tool argument could not be journaled as a Tasks step result;
  bound prompts, tool arguments, and tool outputs at 256 KB, replacing
  oversized arguments with an error the model sees
- session listings carried every checkpoint; omit kernel state from
  listings and let the UI read one checkpoint on demand

Record the results and limits in the README and RFC.

* feat(codex-harness): keep the transcript in Sessions and lift the size limits

The kernel checkpoint no longer carries the transcript. It is a cursor over
one turn: phase, round, and pending tool calls, a few hundred bytes however
long the conversation is. Everything with size moves to the SDK's durable
primitives:

- prompts, assistant messages, and tool outputs are Sessions messages, so
  large ones chunk across rows; each round hydrates a byte-budgeted window
  and Sessions compacts the branch past a token threshold
- tool calls reach the kernel as a pointer to the stored assistant message
  and tool outputs are stored as tool messages, so Tasks journals effects
  by id instead of by value
- events append to the operation's Streams log in the same transaction as
  the checkpoint; the journal table is gone
- workspace_read takes offset and max_bytes, and the Workspace spills large
  files to an R2 bucket
- the prompt shows a marker for any tool input or output over 64 KB, with
  a ranged read to page it back, so one large write cannot evict the rest
  of the turn from the context window

Remove the 256 KB payload caps and the transition cap; the round cap is a
configurable option. The stress suite now completes 8 MB prompts, 8 MB tool
payloads, 60-round turns, 200 calls in a round, 200 turns on one object,
and 32 concurrent objects with a 0.5 KB checkpoint throughout.

* fix(codex-harness): retry a model round the provider fails

A transient Workers AI capacity error failed the whole turn. A round that
returns no usable response now throws inside its Tasks step, which retries
with backoff before the turn fails. Nothing is stored for a failed round,
so a retry starts clean.

* chore(codex-harness): hydrate 32 MiB of history per round, matching Think

* chore(codex-harness): drop the run details panel from the UI

* feat(codex-harness): tools and workspace sidebar

* chore(codex-harness): default ranged reads to offset 0 and drop the unused transport readFile hook
M
Matt committed
2847a28cde578e1a507e2b0afa237977c285d8d9
Parent: ec93caf
Committed by GitHub <noreply@github.com> on 9/8/2026, 3:56:17 PM