>0.9
Eval system macro F1 score
As published on braintrust.dev. Captured by usedby on Oct 6, 2026.
What happened
Navan built Miles, an AI voice agent that calls hotels to confirm bookings and provide payment details, and uses Braintrust to log every call and run automated evaluations that classify call outcomes. Calls needing human follow-up are surfaced in a filtered Braintrust dashboard for the payment operations team.
Summary written by usedby from the source page, in English. The figures are those of Braintrust and Navan, not ours.
- 0.56 to 0.89“improving from 0.56 to 0.89 through iterative refinement”
- 200 calls“Couldn't scale beyond 200 calls”
The team achieved over 0.9 macro F1 score across all control groups for their evaluation system, ensuring that their automated quality checks match human judgment with high accuracy.
Eval-driven development is the new test-driven development. Any projects that we take up, the first step is identifying the eval set.
What the story claims, and what we checked
We compared the story with its live page on Oct 6, 2026.
- The figure: >0.9CheckedThe figure is on the page; its label is our wording.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Navan.
- Navan uses BraintrustCheckedConfirmed line. Latest check across sources: Oct 6, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




