Let Bonk judge whether a PR improves the eval runs (#608)
* Let Bonk judge whether a PR improves the eval runs Bonk's verdict followed the pass-rate test alone, so a PR that only made the agent's work better came back as "no regression". #600 cut the runs that read files in a new, empty Gadget from 20 of 40 to 2 of 40, and #601 raised cache hits on every task; both got a ⚪. Bonk now decides from both sides' trajectories as well as comparison.json: pass rates, wasted steps, tool errors, cache hit rate, time and what the agent built. A change, better or worse, counts only when the diff explains it, both sides show it, and it is larger than run-to-run variation. The Tool errors and Prompt cache sections now show both sides, and the cache section reports rises as well as falls. * Call a mixed eval outcome inconclusive, not a regression The regression rule came first and fired on any change for the worse, so a PR that also counted a change for the better could never reach the inconclusive case meant for it. * Read the Bonk model from the BONK_MODEL repository variable The three Bonk workflows each named their model twice, and the eval review had fallen behind on gpt-5.6-sol. BONK_MODEL is set to openai/gpt-6.1-sol.
A
Ashish Kumar Singh committed
2b81aef553a09a4f03d3f1394d5caee682fd5038
Parent: ec213ab
Committed by GitHub <noreply@github.com>
on 9/30/2026, 5:50:55 PM