95%
“Time to First Token (TTFT): Improved by up to 95%”
As published on together.ai. Checked by usedby on Oct 9, 2026.
What happened
Arcee AI moved the inference for its specialized small language models from AWS-managed Kubernetes (EKS) to Together Dedicated Endpoints, a fully managed GPU deployment. The models power its Arcee Conductor routing and Arcee Orchestra workflow products.
Summary written by usedby from the source page, in English. The figures are those of Together AI and Arcee AI, not ours.
- 41+“Throughput (QPS): Achieved 41+ queries per second at 32 concurrent requests, surpassing Arcee AI's requirements.”
Time to First Token (TTFT): Improved by up to 95%, reducing latency of some models from 485ms on AWS to just 29ms on Together Dedicated Endpoints, exceeding expectations.
The migration was a very simple process. We put our models in a private Hugging Face repository, the Together AI team pulled them down and handed us the API, and we just plugged that into our app. That’s what we wanted–a fully managed experience. Everything has been great: full availability, no downtime.
What the story claims, and what we checked
What we compared with the page.
- The figure: 95%CheckedPrinted word for word on the page, near the name of Arcee AI.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Arcee AI.
- Arcee AI uses Together AICheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.





