80%
“80% lower latencies. Across models and architectures, BEI delivered an average of 80% reduction in all-important P95 latency.”
As published on baseten.co, September 2025. Checked by usedby on Oct 9, 2026.
What happened
Superhuman runs dozens of custom and fine-tuned embedding models on Baseten to power AI classification, search and retrieval in its email app. Baseten Embeddings Inference and autoscaling GPU capacity serve these models with low latency.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Superhuman, not ours.
Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.
Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.
What the story claims, and what we checked
What we compared with the page.
- The figure: 80%CheckedPrinted word for word on the page, near the name of Superhuman.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Superhuman.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Superhuman uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




