# How Cloudflare uses Braintrust

As published on [braintrust.dev](https://www.braintrust.dev/customers/cloudflare). Captured by usedby on 2026-10-06.

- Company: [Cloudflare](https://www.usedby.ai/companies/cloudflare.md)
- Tool: [Braintrust](https://www.usedby.ai/tools/braintrust.md)
- Industry: [Cybersecurity](https://www.usedby.ai/companies/industry/cybersecurity.md)
- Teams: Agent Experience

## What the story says

Cloudflare's Agent Experience team uses Braintrust to evaluate its dashboard agent, running an LLM-as-a-judge over whole conversations to score resolution, gating skill, prompt and recipe changes in CI/CD, and benchmarking sub-agents against models of varying capacity.

Summary written by usedby from the source page, in English. The figures are those of Braintrust and Cloudflare, not ours.

> The dashboard agent's behavior lives in skills, recipes, and prompts, and every change runs through evals in CI/CD. Cloudflare built a shared component, now used across platform internals, that measures skill performance before and after a change.

> Instead of shipping changes, measuring them in production, and going back to the drawing board, we get to iterate directly in development, see how it performs, and then ship it to production.
>
> Kylie Czajkowski, Agent Experience Manager

## What usedby checked

We compared the story with its live page on 2026-10-06.

- Checked: the passage quoted above is copied word for word from the page, near the name of Cloudflare.
- Checked: Cloudflare uses Braintrust. Confirmed line. Latest check across sources: 2026-10-06.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/cloudflare-braintrust · How we check: https://www.usedby.ai/methodology
