40%
“40% lower overall latency.”
As published on baseten.co, December 2025. Checked by usedby on Oct 9, 2026.
What happened
Scaled Cognition runs its agentic models and agent workflows on Baseten's inference platform, using Custom Servers and a hybrid setup that combines its existing AWS environment with Baseten Cloud spillover capacity. It uses the platform for low-latency serving and for autoscaling many research experiments on its APT-1 model.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Scaled Cognition, not ours.
- <120 ms“<120 ms time to first token.”
The Baseten team was able to quickly implement a tailored solution in time for Scaled Cognition’s launch and active POCs, leading to:
Scaled Cognition has always been known for the quality of its agents, but performance is just as important. We partner with Baseten to ensure the lowest possible latency for our models and agentic workflows. It’s been a major differentiator for us in the market.
What the story claims, and what we checked
What we compared with the page.
- The figure: 40%CheckedPrinted word for word on the page, near the name of Scaled Cognition.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Scaled Cognition.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Scaled Cognition uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




