33%
“High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers.”
As published on baseten.co, June 2026. Checked by usedby on Oct 9, 2026.
What happened
OpenCode, an open-source AI coding agent, uses Baseten's Model APIs to serve open-source models through its Zen gateway. It relies on Baseten for fast, reliable inference with KV cache aware routing, cache token pricing, and dependable tool calling.
Summary written by usedby from the source page, in English. The figures are those of Baseten and OpenCode, not ours.
- 5x“5x TPS and higher stability versus closed-source providers”
- 100+ TPS“100+ TPS consistently across open-source models”
- 10M MAUs“OpenCode has grown from roughly 40k MAUs at the start of the partnership to 10M MAUs, with Baseten serving as one of their core inference providers across multiple models amidst massive user growth.”
High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers. OpenCode passed those savings directly to their users with discounted usage for Zen users.
Ever since the launch of Zen, we've gotten a lot of crazy feedback — it's fast and very good, like 90% as good as SOTA closed-source models, but it's just so fast that it changes people's behavior when they code.
What the story claims, and what we checked
What we compared with the page.
- The figure: 33%CheckedPrinted word for word on the page, near the name of OpenCode.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of OpenCode.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- OpenCode uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




