# How Navan uses Braintrust

**\>0.9** Eval system macro F1 score

As published on [braintrust.dev](https://www.braintrust.dev/customers/navan). Captured by usedby on 2026-10-06.

- Company: [Navan](https://www.usedby.ai/companies/navan.md)
- Tool: [Braintrust](https://www.usedby.ai/tools/braintrust.md)
- Industry: [Travel & Hospitality](https://www.usedby.ai/companies/industry/travel-hospitality.md)
- Teams: Engineering, Payment operations

## What the story says

Navan built Miles, an AI voice agent that calls hotels to confirm bookings and provide payment details, and uses Braintrust to log every call and run automated evaluations that classify call outcomes. Calls needing human follow-up are surfaced in a filtered Braintrust dashboard for the payment operations team.

Summary written by usedby from the source page, in English. The figures are those of Braintrust and Navan, not ours.

> The team achieved over 0.9 macro F1 score across all control groups for their evaluation system, ensuring that their automated quality checks match human judgment with high accuracy.

- **0.56 to 0.89**: “improving from 0.56 to 0.89 through iterative refinement”
- **200 calls**: “Couldn't scale beyond 200 calls”

> Eval-driven development is the new test-driven development. Any projects that we take up, the first step is identifying the eval set.
>
> Sarav Bhatia, Senior Director of Software Engineering

## What usedby checked

We compared the story with its live page on 2026-10-06.

- Checked: the figure \>0.9 is on the page; its label is our wording.
- Checked: the passage quoted above is copied word for word from the page, near the name of Navan.
- Checked: Navan uses Braintrust. Confirmed line. Latest check across sources: 2026-10-06.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/navan-braintrust · How we check: https://www.usedby.ai/methodology
