Skip to main content
usedby

44%

“Speechify's cost per million characters dropped by 44%”

As published on baseten.co, May 2026. Checked by usedby on Oct 9, 2026.

What happened

Speechify runs its SIMBA text-to-speech models, plus text normalization, voice conversion, sound FX and page parsing models, on Baseten for real-time, high-volume inference. Researchers deploy models themselves with Truss, and traffic-based autoscaling replaced a self-managed GPU stack.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Speechify, not ours.

  • 30–50%“Across the SIMBA TTS family, p99 inference latency dropped 30–50% post-migration”
  • 4.5x“Replica startup became 4.5x faster”
  • 50-70%“latency spikes dropped by 50-70%”
  • 76 ms“Its new SIMBA 3.0 vLLM models run on Baseten with 76 ms p50 TTFB.”

Because of Baseten's efficient autoscaling, model performance and infrastructure optimizations, Speechify's cost per million characters dropped by 44%, as traffic on its platform grew 7% during the same period.

From the page. baseten.co

A researcher shipped SIMBA 3.0 vLLM into production by themselves. That would have taken days of work from our entire AI Platform team with our old infrastructure stack.

Kai Krause, VP of Engineering and AI, Speechify. Source

What the story claims, and what we checked

What we compared with the page.

  • The figure: 44%CheckedPrinted word for word on the page, near the name of Speechify.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Speechify.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Speechify uses BasetenCheckedConfirmed line.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

under a week“Time-to-hire for key roles dropped from months to under a week”Speechify uses Juicebox. Another tool at Speechify30%“30% lower cost per call”you.com uses Baseten. Another customer of Baseten3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems uses Baseten. Another customer of Baseten