>0.9
Eval system macro F1 score
Tal como se publicó en braintrust.dev. Capturado por usedby el 6 oct 2026.
Qué pasó
Navan built Miles, an AI voice agent that calls hotels to confirm bookings and provide payment details, and uses Braintrust to log every call and run automated evaluations that classify call outcomes. Calls needing human follow-up are surfaced in a filtered Braintrust dashboard for the payment operations team.
Resumen escrito por usedby a partir de la página de origen, en inglés. Las cifras son de Braintrust y de Navan, no nuestras.
- 0.56 to 0.89“improving from 0.56 to 0.89 through iterative refinement”
- 200 calls“Couldn't scale beyond 200 calls”
The team achieved over 0.9 macro F1 score across all control groups for their evaluation system, ensuring that their automated quality checks match human judgment with high accuracy.
Eval-driven development is the new test-driven development. Any projects that we take up, the first step is identifying the eval set.
Lo que dice la historia, y lo que verificamos
Comparamos la historia con su página en línea el 6 oct 2026.
- La cifra: >0.9VerificadoLa cifra está en la página; su etiqueta es redacción nuestra.
- El pasaje citado arribaVerificadoCopiado palabra por palabra de la página, cerca del nombre de Navan.
- Navan usa BraintrustVerificadoLínea de nivel Confirmado. Última verificación entre todas las fuentes: 6 oct 2026.
- El resultado en síNo verificadoLo citamos; no lo medimos.




