# How Superhuman uses Baseten

**80%**: “80% lower latencies. Across models and architectures, BEI delivered an average of 80% reduction in all-important P95 latency.”

As published on [baseten.co](https://www.baseten.co/resources/customers/superhuman/), September 2025. Captured by usedby on 2026-10-09.

- Company: [Superhuman](https://www.usedby.ai/companies/superhuman.md)
- Tool: [Baseten](https://www.usedby.ai/tools/baseten.md)
- Industry: [DevTools & Infrastructure](https://www.usedby.ai/companies/industry/devtools-infrastructure.md)
- Teams: AI engineering

## What the story says

Superhuman runs dozens of custom and fine-tuned embedding models on Baseten to power AI classification, search and retrieval in its email app. Baseten Embeddings Inference and autoscaling GPU capacity serve these models with low latency.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Superhuman, not ours.

> Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.

> Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.
>
> Loïc Houssier, CTO

Source: [baseten.co](https://www.baseten.co/resources/customers/superhuman/), captured 2026-10-09.

## What usedby checked

We compared the story with its live page on 2026-10-09.

- Checked: the figure 80% is printed word for word on the page, near the name of Superhuman.
- Checked: the passage quoted above is copied word for word from the page, near the name of Superhuman.
- Checked: the publication date is read from the page’s own metadata, never guessed.
- Checked: Superhuman uses Baseten. Confirmed line. Latest check across sources: 2026-10-09.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/superhuman-baseten · How we check: https://www.usedby.ai/methodology
