# How Hebbia uses Baseten

**10x**: “Reduced cost by over 10x by shifting to Baseten”

As published on [baseten.co](https://www.baseten.co/resources/customers/hebbia/), April 2026. Captured by usedby on 2026-10-05.

- Company: [Hebbia](https://www.usedby.ai/companies/hebbia.md)
- Tool: [Baseten](https://www.usedby.ai/tools/baseten.md)

## What the story says

Hebbia runs a dedicated deployment of an open-source LLM on Baseten to serve latency-sensitive chat traffic for its institutional finance customers. It moved from a closed-source model provider and uses Baseten's elastic infrastructure to pay only for the GPUs it uses.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Hebbia, not ours.

> By shifting from a closed-source model provider to Baseten, Hebbia reduced inference costs by over 10x without compromising on the reliability or latency their enterprise customers depend on.

- **2.5x**: “2.5x improvement in TPS and 4x improvement in TTFT using the Baseten Inference Stack”

> Our customers really care about speed, reliability, and the ability to deploy custom fine-tuned models for financial services. Baseten met all three criteria so far and it’s been a boon to our business.
>
> George Sivulka, CEO

## What usedby checked

We compared the story with its live page on 2026-10-05.

- Checked: the figure 10x is printed word for word on the page, near the name of Hebbia.
- Checked: the passage quoted above is copied word for word from the page, near the name of Hebbia.
- Checked: the publication date is read from the page’s own metadata, never guessed.
- Checked: Hebbia uses Baseten. Confirmed line. Latest check across sources: 2026-10-05.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/hebbia-baseten · How we check: https://www.usedby.ai/methodology
