90%
“Sully reports over 90% reduction in inference costs”
As published on baseten.co, February 2026. Checked by usedby on Oct 9, 2026.
What happened
Sully.ai moved the inference behind its healthcare AI agents (clinical notes, decision support, medical coding) from closed-source models to open-source models served on Baseten. The aim was lower cost and latency and more control over model quality.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Sully AI, not ours.
- 65%“This 65% reduction in median latency resulted in faster, more predictable responses during clinical workflows, improving overall experience for physicians.”
- 29 million“December 2025: Sully has added ~29 million minutes to their customers' workforce overall”
- 2.4+ hours“2.4+ hours time saving per physician”
Sully reports over 90% reduction in inference costs by switching from closed-source platforms to open-source models accessed via Baseten. It enabled them to reinvest these savings in the business to improve agents that help provide more value to practitioners and their patients.
At our scale, inference efficiency matters as much as model quality. Baseten enabled us to cut costs by 90% while delivering significantly faster, more predictable performance. That efficiency is what allows us to add millions of productive minutes back into our customers.
What the story claims, and what we checked
What we compared with the page.
- The figure: 90%CheckedPrinted word for word on the page, near the name of Sully AI.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Sully AI.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Sully AI uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




