30%
“30% lower cost per call”
As published on baseten.co, August 2026. Checked by usedby on Oct 9, 2026.
What happened
You.com runs open-source LLMs on Baseten's dedicated, multi-region deployments to handle the synthesis step of its Answer API, turning retrieved passages into cited, verified answers instead of routing each query to a frontier model.
Summary written by usedby from the source page, in English. The figures are those of Baseten and you.com, not ours.
- 93.48%“93.48% accuracy on SimpleQA, 2.67-second p50 latency, and $5 per 1,000 calls.”
Compared with routing synthesis to a frontier model, You.com saw:
We think open-weight models plus web search is a powerful alternative to routing everything through frontier model providers, at a fraction of the price and comparable quality for our developers.
What the story claims, and what we checked
What we compared with the page.
- The figure: 30%CheckedPrinted word for word on the page, near the name of you.com.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of you.com.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- you.com uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




