Skip to main content
usedby

3x

“3x cost savings when moving from closed-source to open-source models”

As published on baseten.co, July 2026. Captured by usedby on Oct 5, 2026.

What happened

Parallel runs the multi-step agentic pipelines behind its web search and research APIs on Baseten. It started on pay-per-token Model APIs, then moved to dedicated deployments of an open-source model and now also trains and deploys fine-tuned models on the platform.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Parallel Web Systems, not ours.

  • 50%“Baseten’s proprietary runtime cut latency by 50% and increased throughput by 3x through KV cache-aware routing and speculative decoding.”
  • ~100x“Within a few months of their initial deployment, Parallel scaled traffic with Baseten ~100x.”

Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model. Eliminating redundant processing made it economically viable to run the kind of long, multi-turn agent workflows that deliver meaningfully better results than a single-shot query.

From the page. baseten.co, captured Oct 5, 2026

Our workloads can be very spiky. We need to be able to manage the peaks, but we also don't want to pay for the valleys. Pay-per-token with Model APIs enabled us to easily make the transition to open-source until our scale merited dedicated deployments optimized for throughput and flexible autoscaling.

Matt Lee, Engineering Lead, Parallel Web Systems. Source, captured Oct 5, 2026

What the story claims, and what we checked

We compared the story with its live page on Oct 5, 2026.

  • The figure: 3xCheckedPrinted word for word on the page, near the name of Parallel Web Systems.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Parallel Web Systems.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Parallel Web Systems uses BasetenCheckedConfirmed line. Latest check across sources: Oct 5, 2026.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

60%“~60% reduction in inference costs”EliseAI uses Baseten. Another customer of Baseten10x“Reduced cost by over 10x by shifting to Baseten”Hebbia uses Baseten. Another customer of Baseten30%-80%“30%-80% faster image generation per model”Gamma uses Baseten. Another customer of Baseten