# How Decagon uses Together AI

**6×**: “Decagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini.”

As published on [together.ai](https://www.together.ai/customers/decagon). Captured by usedby on 2026-10-03.

- Company: [Decagon](https://www.usedby.ai/companies/decagon.md)
- Tool: [Together AI](https://www.usedby.ai/tools/together-ai.md)
- Industry: [Artificial Intelligence](https://www.usedby.ai/companies/industry/artificial-intelligence.md)
- Teams: Research

## What the story says

Decagon runs production inference for its multi-model voice AI agents on Together AI, using hosted open-source and fine-tuned models on NVIDIA B200 GPUs. Together also helped train custom speculative decoding draft models and tune deployments to meet strict latency budgets.

Summary written by usedby from the source page, in English. The figures are those of Together AI and Decagon, not ours.

> Decagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini. This economic improvement made 24/7 voice deployments viable at scale.

- **\<400ms**: “Decagon reduced p95 model latency per turn from seconds to \<400ms on inputs up to tens of thousands of tokens.”

> Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up.
>
> Max Lu, Head of Research

## What usedby checked

We compared the story with its live page on 2026-10-03.

- Checked: the figure 6× is printed word for word on the page, near the name of Decagon.
- Checked: the passage quoted above is copied word for word from the page, near the name of Decagon.
- Checked: each number in our summary is printed on the page.
- Checked: Decagon uses Together AI. Confirmed line. Latest check across sources: 2026-10-03.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/decagon-together-ai · How we check: https://www.usedby.ai/methodology
