Skip to main content
usedby

80%+

Score threshold for internal release

As published on braintrust.dev. Captured by usedby on Oct 6, 2026.

What happened

Box's product managers use Braintrust as the system of record for the datasets, experiments and metrics behind evals of the Box Agent. They run programmatic evals nightly, cluster failures with tags, grade outputs with LLM and code-based graders, and compare models against a fixed baseline dataset.

Summary written by usedby from the source page, in English. The figures are those of Braintrust and Box, not ours.

Box's path to launch was determined by datasets. Before opening the agent to internal dogfooding, the team built a set of roughly eighty questions covering core functionality and decided there would be no internal release until the agent scored above 80%.

From the page. braintrust.dev, captured Oct 6, 2026

What the story claims, and what we checked

We compared the story with its live page on Oct 6, 2026.

  • The figure: 80%+CheckedThe figure is on the page; its label is our wording.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Box.
  • Box uses BraintrustCheckedConfirmed line. Latest check across sources: Oct 6, 2026.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

>0.9Eval system macro F1 scoreNavan uses Braintrust. Another customer of Braintrust~10 → 0 minsEngineering time to debug an issuePylon uses Braintrust. Another customer of Braintrust30-min“Continuous signal, 30-min fix cycles”Rex uses Braintrust. Another customer of Braintrust