Skip to main content
usedby

>0.9

Eval system macro F1 score

As published on braintrust.dev. Captured by usedby on Oct 6, 2026.

What happened

Navan built Miles, an AI voice agent that calls hotels to confirm bookings and provide payment details, and uses Braintrust to log every call and run automated evaluations that classify call outcomes. Calls needing human follow-up are surfaced in a filtered Braintrust dashboard for the payment operations team.

Summary written by usedby from the source page, in English. The figures are those of Braintrust and Navan, not ours.

  • 0.56 to 0.89“improving from 0.56 to 0.89 through iterative refinement”
  • 200 calls“Couldn't scale beyond 200 calls”

The team achieved over 0.9 macro F1 score across all control groups for their evaluation system, ensuring that their automated quality checks match human judgment with high accuracy.

From the page. braintrust.dev, captured Oct 6, 2026

Eval-driven development is the new test-driven development. Any projects that we take up, the first step is identifying the eval set.

Sarav Bhatia, Senior Director of Software Engineering, Navan. Source, captured Oct 6, 2026

What the story claims, and what we checked

We compared the story with its live page on Oct 6, 2026.

  • The figure: >0.9CheckedThe figure is on the page; its label is our wording.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Navan.
  • Navan uses BraintrustCheckedConfirmed line. Latest check across sources: Oct 6, 2026.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

80%+Score threshold for internal releaseBox uses Braintrust. Another customer of Braintrust~10 → 0 minsEngineering time to debug an issuePylon uses Braintrust. Another customer of Braintrust30-min“Continuous signal, 30-min fix cycles”Rex uses Braintrust. Another customer of Braintrust