3x
“3x cost savings when moving from closed-source to open-source models”
As published on baseten.co, July 2026. Captured by usedby on Oct 5, 2026.
What happened
Parallel runs the multi-step agentic pipelines behind its web search and research APIs on Baseten. It started on pay-per-token Model APIs, then moved to dedicated deployments of an open-source model and now also trains and deploys fine-tuned models on the platform.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Parallel Web Systems, not ours.
- 50%“Baseten’s proprietary runtime cut latency by 50% and increased throughput by 3x through KV cache-aware routing and speculative decoding.”
- ~100x“Within a few months of their initial deployment, Parallel scaled traffic with Baseten ~100x.”
Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model. Eliminating redundant processing made it economically viable to run the kind of long, multi-turn agent workflows that deliver meaningfully better results than a single-shot query.
Our workloads can be very spiky. We need to be able to manage the peaks, but we also don't want to pay for the valleys. Pay-per-token with Model APIs enabled us to easily make the transition to open-source until our scale merited dedicated deployments optimized for throughput and flexible autoscaling.
What the story claims, and what we checked
We compared the story with its live page on Oct 5, 2026.
- The figure: 3xCheckedPrinted word for word on the page, near the name of Parallel Web Systems.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Parallel Web Systems.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Parallel Web Systems uses BasetenCheckedConfirmed line. Latest check across sources: Oct 5, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




