Skip to main content
usedby

90%

“Sully reports over 90% reduction in inference costs”

As published on baseten.co, February 2026. Checked by usedby on Oct 9, 2026.

What happened

Sully.ai moved the inference behind its healthcare AI agents (clinical notes, decision support, medical coding) from closed-source models to open-source models served on Baseten. The aim was lower cost and latency and more control over model quality.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Sully AI, not ours.

  • 65%“This 65% reduction in median latency resulted in faster, more predictable responses during clinical workflows, improving overall experience for physicians.”
  • 29 million“December 2025: Sully has added ~29 million minutes to their customers' workforce overall”
  • 2.4+ hours“2.4+ hours time saving per physician”

Sully reports over 90% reduction in inference costs by switching from closed-source platforms to open-source models accessed via Baseten. It enabled them to reinvest these savings in the business to improve agents that help provide more value to practitioners and their patients.

From the page. baseten.co

At our scale, inference efficiency matters as much as model quality. Baseten enabled us to cut costs by 90% while delivering significantly faster, more predictable performance. That efficiency is what allows us to add millions of productive minutes back into our customers.

Ahmed Omar, CB=NB/SDER AND CEO, Sully AI. Source

What the story claims, and what we checked

What we compared with the page.

  • The figure: 90%CheckedPrinted word for word on the page, near the name of Sully AI.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Sully AI.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Sully AI uses BasetenCheckedConfirmed line.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

30%“30% lower cost per call”you.com uses Baseten. Another customer of Baseten3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems uses Baseten. Another customer of Baseten33%“High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers.”OpenCode uses Baseten. Another customer of Baseten