Skip to main content
usedby

30%

“30% lower cost per call”

As published on baseten.co, August 2026. Checked by usedby on Oct 9, 2026.

What happened

You.com runs open-source LLMs on Baseten's dedicated, multi-region deployments to handle the synthesis step of its Answer API, turning retrieved passages into cited, verified answers instead of routing each query to a frontier model.

Summary written by usedby from the source page, in English. The figures are those of Baseten and you.com, not ours.

  • 93.48%“93.48% accuracy on SimpleQA, 2.67-second p50 latency, and $5 per 1,000 calls.”

Compared with routing synthesis to a frontier model, You.com saw:

From the page. baseten.co

We think open-weight models plus web search is a powerful alternative to routing everything through frontier model providers, at a fraction of the price and comparable quality for our developers.

Saahil Jain, CTO, You.com. Source

What the story claims, and what we checked

What we compared with the page.

  • The figure: 30%CheckedPrinted word for word on the page, near the name of you.com.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of you.com.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • you.com uses BasetenCheckedConfirmed line.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems uses Baseten. Another customer of Baseten33%“High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers.”OpenCode uses Baseten. Another customer of Baseten44%“Speechify's cost per million characters dropped by 44%”Speechify uses Baseten. Another customer of Baseten