100,000s
“100,000s of public eval dashboard views”
As published on braintrust.dev. Checked by usedby on Oct 9, 2026.
What happened
Browserbase, which provides infrastructure for AI agents to browse the web, uses Braintrust to add model observability on top of browser session replay and to power public benchmarks and evals for its open source Stagehand SDK. This helps it and its customers measure browser agent reliability and choose the best model.
Summary written by usedby from the source page, in English. The figures are those of Braintrust and Browserbase, not ours.
Browserbase maintains Stagehand, an open source SDK that provides a unified tool interface for agents to control browsers in TypeScript, Python, Go, and other languages. Alongside Stagehand, the team publishes benchmarks and evals powered by Braintrust that show which models perform best at browser tasks.
This is an important value for our customers because they don't know which model to choose. When we run all of these benchmarks and evals against Stagehand, it makes it really easy for our customers to know which model's going to work best.
Braintrust shows us the model's input and output and what actions it took.
What we're often missing is, what was the model thinking? That's where Braintrust comes in.
What the story claims, and what we checked
What we compared with the page.
- The figure: 100,000sCheckedPrinted word for word on the page, near the name of Browserbase.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Browserbase.
- Browserbase uses BraintrustCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.






