# How Rime uses Baseten

**\<300** milliseconds p99 latency

As published on [baseten.co](https://www.baseten.co/resources/customers/rime/), October 2024. Captured by usedby on 2026-10-09.

- Company: [Rime](https://www.usedby.ai/companies/rime.md)
- Tool: [Baseten](https://www.usedby.ai/tools/baseten.md)

## What the story says

Rime trains its own text-to-speech models and serves them through an enterprise API on Baseten's multi-region inference infrastructure. This lets it meet real-time latency and uptime requirements and scale GPU capacity as customer demand grows.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Rime, not ours.

> Rime has consistently matched or beat their 300-millisecond p99 latency SLA for enterprise customers with optimized network serving and multi-region model deployments.

- **200**: “Rime.ai offers a speech synthesis API with unmatched speed, lifelike accuracy, and support for over 200 distinct voices.”

> It would be hard or impossible to get the GPUs we need right when we need them on the same terms that Baseten gives through any cloud service provider – this flexibility is essential as we scale.
>
> Lily Clifford, Co-Founder and CEO

Source: [baseten.co](https://www.baseten.co/resources/customers/rime/), captured 2026-10-09.

## What usedby checked

We compared the story with its live page on 2026-10-09.

- Checked: the figure \<300 is on the page; its label is our wording.
- Checked: the passage quoted above is copied word for word from the page, near the name of Rime.
- Checked: the publication date is read from the page’s own metadata, never guessed.
- Checked: Rime uses Baseten. Confirmed line. Latest check across sources: 2026-10-09.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/rime-baseten · How we check: https://www.usedby.ai/methodology
