feat(evals): add mastra bench harness (integrations/mastra-sdk) (#2807)
## What Adds `--harness mastra` to evals, backed by a new private `@browserbasehq/stagehand-integrations-mastra-sdk` package. - `packages/integrations/mastra-sdk` — `runMastraSession`: in-process `@mastra/core` Agent + `@mastra/mcp` MCPClient (per-session UUID id), streamed tool-call/result/usage events, abort forwarding, status/stopReason normalization, redaction. - `packages/evals/framework/mastraToolAdapter.ts` — `via:"mcp"` mounts (stagehand_facade, playwright_mcp, chrome_devtools_mcp) with idempotent cleanup; step observations recorded. - `packages/evals/framework/harnesses/mastraAdapter.ts` — events → `NormalizedToolCall[]`. - Registered via `defineExternalHarness` on the shared external-runner skeleton from the base PR. ## Testing - Full unit gate green: evals 488, mastra-sdk 12, all other suites unchanged. - Connected smoke (`evals run b:webvoyager --harness mastra --tool stagehand_facade -l 1 -t 1 -e browserbase`): harness drove **7 Stagehand tool calls** through the facade and produced a final answer. Stacked on the wave-core PR. Implemented with Codex (gpt-5.6) under supervision; Claude + `codex exec review` findings (MCPClient id collision, env leakage, redaction) fixed in-branch. <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Adds a selectable `mastra` eval harness backed by the private `@browserbasehq/stagehand-integrations-mastra-sdk` package. Evals can now run tasks through `@mastra/core` and `@mastra/mcp`, sharing the external-runner prompt, parsing, grading, and metrics flow used by the other agent harnesses. - Wraps `Agent.stream` and `MCPClient` with per-session client IDs, abort forwarding, bounded discovery and cleanup, normalized usage, and sanitized output. - Mounts `stagehand_facade`, `playwright_mcp`, and `chrome_devtools_mcp` over stdio MCP, while exposing code surfaces through the existing in-process bridge. - Converts Mastra stream events into sanitized trajectories with paired tool calls and results, image extraction, and correctly concatenated reasoning and text fragments. - Treats an SDK error as a failed run even when the model emitted valid-looking success JSON first. - Registers `mastra` with `stagehand_facade` as the default tool surface, `openai/gpt-5.4-mini` as the default model, and `EVAL_MASTRA_MODELS` as the override. - Adds workspace dependencies, lockfile and Turbo targets, CI build caching, and unit coverage for the SDK, adapters, runner, registry, and planner. - Mastra requires `@mastra/core` and `@mastra/mcp`; MCP definitions must use stdio servers with a string `command`. - Supports step, MCP, evidence-capture, and teardown limits through `EVAL_MASTRA_MAX_STEPS`, `EVAL_MASTRA_MCP_TIMEOUT_MS`, `EVAL_CAPTURE_EVIDENCE_TIMEOUT_MS`, and `EVAL_AGENT_MOUNT_CLEANUP_TIMEOUT_MS`. <sup>Written for commit b9180990d29d12840ba61f6d876f9270fb5e75f5. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/browserbase/stagehand/pull/2807?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> --------- Co-authored-by: miguel <miguel@browserbase.com>
M
Miguel committed
7adfd33d3c8b7330ae6285cbedc08543b5e1449e
Parent: ad33f63
Committed by GitHub <noreply@github.com>
on 9/2/2026, 12:07:01 AM