80%+
Score threshold for internal release
As published on braintrust.dev. Captured by usedby on Oct 6, 2026.
What happened
Box's product managers use Braintrust as the system of record for the datasets, experiments and metrics behind evals of the Box Agent. They run programmatic evals nightly, cluster failures with tags, grade outputs with LLM and code-based graders, and compare models against a fixed baseline dataset.
Summary written by usedby from the source page, in English. The figures are those of Braintrust and Box, not ours.
Box's path to launch was determined by datasets. Before opening the agent to internal dogfooding, the team built a set of roughly eighty questions covering core functionality and decided there would be no internal release until the agent scored above 80%.
What the story claims, and what we checked
We compared the story with its live page on Oct 6, 2026.
- The figure: 80%+CheckedThe figure is on the page; its label is our wording.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Box.
- Box uses BraintrustCheckedConfirmed line. Latest check across sources: Oct 6, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




