<300
milliseconds p99 latency
As published on baseten.co, October 2024. Checked by usedby on Oct 9, 2026.
What happened
Rime trains its own text-to-speech models and serves them through an enterprise API on Baseten's multi-region inference infrastructure. This lets it meet real-time latency and uptime requirements and scale GPU capacity as customer demand grows.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Rime, not ours.
- 200“Rime.ai offers a speech synthesis API with unmatched speed, lifelike accuracy, and support for over 200 distinct voices.”
Rime has consistently matched or beat their 300-millisecond p99 latency SLA for enterprise customers with optimized network serving and multi-region model deployments.
It would be hard or impossible to get the GPUs we need right when we need them on the same terms that Baseten gives through any cloud service provider – this flexibility is essential as we scale.
What the story claims, and what we checked
What we compared with the page.
- The figure: <300CheckedThe figure is on the page; its label is our wording.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Rime.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Rime uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




