60%
“~60% reduction in inference costs”
As published on baseten.co, May 2026. Captured by usedby on Oct 5, 2026.
What happened
EliseAI fine-tunes small open-source models on Baseten Training and serves them on Baseten's inference platform to extract structured data from housing and healthcare conversations, including a live voice agent for leasing. Baseten's research team also helped with post-training techniques.
Summary written by usedby from the source page, in English. The figures are those of Baseten and EliseAI, not ours.
- 250ms“250ms p90 latency, down from 2.2 seconds on closed-source APIs”
- 99%“99% accuracy on one of the most important features in the leasing platform”
- 25x“Frontier-matching accuracy on a 4B parameter model, matching or exceeding models 25x its size on production benchmarks”
When their closed-source APIs hit a ceiling on cost, speed, and controllability, EliseAI set out to build their own custom models. They partnered with the Baseten research team to get there.
You can't treat fine-tuning as a one-and-done check item. This space moves so fast that you have to continuously keep up.
What the story claims, and what we checked
We compared the story with its live page on Oct 5, 2026.
- The figure: 60%CheckedPrinted word for word on the page, near the name of EliseAI.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of EliseAI.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- EliseAI uses BasetenCheckedConfirmed line. Latest check across sources: Oct 5, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




