SIGN IN SIGN UP

Let Bonk judge whether a PR improves the eval runs (#608)

* Let Bonk judge whether a PR improves the eval runs

Bonk's verdict followed the pass-rate test alone, so a PR that only made
the agent's work better came back as "no regression". #600 cut the runs
that read files in a new, empty Gadget from 20 of 40 to 2 of 40, and
#601 raised cache hits on every task; both got a ⚪.

Bonk now decides from both sides' trajectories as well as
comparison.json: pass rates, wasted steps, tool errors, cache hit rate,
time and what the agent built. A change, better or worse, counts only
when the diff explains it, both sides show it, and it is larger than
run-to-run variation. The Tool errors and Prompt cache sections now show
both sides, and the cache section reports rises as well as falls.

* Call a mixed eval outcome inconclusive, not a regression

The regression rule came first and fired on any change for the worse, so a PR that also counted a change for the better could never reach the inconclusive case meant for it.

* Read the Bonk model from the BONK_MODEL repository variable

The three Bonk workflows each named their model twice, and the eval review had fallen behind on gpt-5.6-sol. BONK_MODEL is set to openai/gpt-6.1-sol.
A
Ashish Kumar Singh committed
2b81aef553a09a4f03d3f1394d5caee682fd5038
Parent: ec213ab
Committed by GitHub <noreply@github.com> on 9/30/2026, 5:50:55 PM