6×
“Decagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini.”
As published on together.ai. Captured by usedby on Oct 3, 2026.
What happened
Decagon runs production inference for its multi-model voice AI agents on Together AI, using hosted open-source and fine-tuned models on NVIDIA B200 GPUs. Together also helped train custom speculative decoding draft models and tune deployments to meet strict latency budgets.
Summary written by usedby from the source page, in English. The figures are those of Together AI and Decagon, not ours.
- <400ms“Decagon reduced p95 model latency per turn from seconds to <400ms on inputs up to tens of thousands of tokens.”
Decagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini. This economic improvement made 24/7 voice deployments viable at scale.
Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up.
What the story claims, and what we checked
We compared the story with its live page on Oct 3, 2026.
- The figure: 6×CheckedPrinted word for word on the page, near the name of Decagon.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Decagon.
- The numbers in our summaryCheckedEach one is printed on the page.
- Decagon uses Together AICheckedConfirmed line. Latest check across sources: Oct 3, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.





