Skip to main content
usedby

10x

“Reduced cost by over 10x by shifting to Baseten”

As published on baseten.co, April 2026. Captured by usedby on Oct 5, 2026.

What happened

Hebbia runs a dedicated deployment of an open-source LLM on Baseten to serve latency-sensitive chat traffic for its institutional finance customers. It moved from a closed-source model provider and uses Baseten's elastic infrastructure to pay only for the GPUs it uses.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Hebbia, not ours.

  • 2.5x“2.5x improvement in TPS and 4x improvement in TTFT using the Baseten Inference Stack”

By shifting from a closed-source model provider to Baseten, Hebbia reduced inference costs by over 10x without compromising on the reliability or latency their enterprise customers depend on.

From the page. baseten.co, captured Oct 5, 2026

Our customers really care about speed, reliability, and the ability to deploy custom fine-tuned models for financial services. Baseten met all three criteria so far and it’s been a boon to our business.

George Sivulka, CEO, Hebbia. Source, captured Oct 5, 2026

What the story claims, and what we checked

We compared the story with its live page on Oct 5, 2026.

  • The figure: 10xCheckedPrinted word for word on the page, near the name of Hebbia.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Hebbia.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Hebbia uses BasetenCheckedConfirmed line. Latest check across sources: Oct 5, 2026.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

3x“3x cost savings when moving from closed-source to open-source models”Parallel Web Systems uses Baseten. Another customer of Baseten60%“~60% reduction in inference costs”EliseAI uses Baseten. Another customer of Baseten30%-80%“30%-80% faster image generation per model”Gamma uses Baseten. Another customer of Baseten