Skip to main content
usedby

6×

“Decagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini.”

As published on together.ai. Captured by usedby on Oct 3, 2026.

What happened

Decagon runs production inference for its multi-model voice AI agents on Together AI, using hosted open-source and fine-tuned models on NVIDIA B200 GPUs. Together also helped train custom speculative decoding draft models and tune deployments to meet strict latency budgets.

Summary written by usedby from the source page, in English. The figures are those of Together AI and Decagon, not ours.

  • <400ms“Decagon reduced p95 model latency per turn from seconds to <400ms on inputs up to tens of thousands of tokens.”

Decagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini. This economic improvement made 24/7 voice deployments viable at scale.

From the page. together.ai, captured Oct 3, 2026

Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up.

Max Lu, Head of Research, Decagon. Source, captured Oct 3, 2026

What the story claims, and what we checked

We compared the story with its live page on Oct 3, 2026.

  • The figure: 6×CheckedPrinted word for word on the page, near the name of Decagon.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Decagon.
  • The numbers in our summaryCheckedEach one is printed on the page.
  • Decagon uses Together AICheckedConfirmed line. Latest check across sources: Oct 3, 2026.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

under two weeks“Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks”Deep Cogito uses Together AI. Another customer of Together AICursor uses Together AIAnother customer of Together AI2.5x“train models 2.5x faster than previously experienced on NVIDIA H200s”Mistral AI uses CoreWeave. Same industry: Artificial Intelligence