44%
“Speechify's cost per million characters dropped by 44%”
As published on baseten.co, May 2026. Checked by usedby on Oct 9, 2026.
What happened
Speechify runs its SIMBA text-to-speech models, plus text normalization, voice conversion, sound FX and page parsing models, on Baseten for real-time, high-volume inference. Researchers deploy models themselves with Truss, and traffic-based autoscaling replaced a self-managed GPU stack.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Speechify, not ours.
- 30–50%“Across the SIMBA TTS family, p99 inference latency dropped 30–50% post-migration”
- 4.5x“Replica startup became 4.5x faster”
- 50-70%“latency spikes dropped by 50-70%”
- 76 ms“Its new SIMBA 3.0 vLLM models run on Baseten with 76 ms p50 TTFB.”
Because of Baseten's efficient autoscaling, model performance and infrastructure optimizations, Speechify's cost per million characters dropped by 44%, as traffic on its platform grew 7% during the same period.
A researcher shipped SIMBA 3.0 vLLM into production by themselves. That would have taken days of work from our entire AI Platform team with our old infrastructure stack.
What the story claims, and what we checked
What we compared with the page.
- The figure: 44%CheckedPrinted word for word on the page, near the name of Speechify.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Speechify.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Speechify uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




