# How Box uses Braintrust

**80%+** Score threshold for internal release

As published on [braintrust.dev](https://www.braintrust.dev/customers/box). Captured by usedby on 2026-10-06.

- Company: [Box](https://www.usedby.ai/companies/box.md)
- Tool: [Braintrust](https://www.usedby.ai/tools/braintrust.md)
- Teams: Product Management, Engineering

## What the story says

Box's product managers use Braintrust as the system of record for the datasets, experiments and metrics behind evals of the Box Agent. They run programmatic evals nightly, cluster failures with tags, grade outputs with LLM and code-based graders, and compare models against a fixed baseline dataset.

Summary written by usedby from the source page, in English. The figures are those of Braintrust and Box, not ours.

> Box's path to launch was determined by datasets. Before opening the agent to internal dogfooding, the team built a set of roughly eighty questions covering core functionality and decided there would be no internal release until the agent scored above 80%.

## What usedby checked

We compared the story with its live page on 2026-10-06.

- Checked: the figure 80%+ is on the page; its label is our wording.
- Checked: the passage quoted above is copied word for word from the page, near the name of Box.
- Checked: Box uses Braintrust. Confirmed line. Latest check across sources: 2026-10-06.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/box-braintrust · How we check: https://www.usedby.ai/methodology
