Skip to main content
usedby

80%

“80% lower latencies. Across models and architectures, BEI delivered an average of 80% reduction in all-important P95 latency.”

As published on baseten.co, September 2025. Checked by usedby on Oct 9, 2026.

What happened

Superhuman runs dozens of custom and fine-tuned embedding models on Baseten to power AI classification, search and retrieval in its email app. Baseten Embeddings Inference and autoscaling GPU capacity serve these models with low latency.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Superhuman, not ours.

Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.

From the page. baseten.co

Baseten cut our P95 latency by 80% across the dozens of fine-tuned embedding models that power core features in Superhuman's AI-native email app.

Loïc Houssier, CTO, Superhuman. Source

What the story claims, and what we checked

What we compared with the page.

  • The figure: 80%CheckedPrinted word for word on the page, near the name of Superhuman.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Superhuman.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Superhuman uses BasetenCheckedConfirmed line.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

2.5x“AI-recommended contacts in Common Room have a 2.5x rate of meetings booked compared to cold outbound contacts”Superhuman uses Common Room. Another tool at Superhuman30%“30% lower cost per call”you.com uses Baseten. Another customer of Baseten3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems uses Baseten. Another customer of Baseten