Skip to main content
usedby

50%

“50% lower inference latency through a combination of using a fine-tuned model and other model optimizations”

As published on baseten.co, November 2025. Checked by usedby on Oct 9, 2026.

What happened

Oxen AI runs its customers' model training jobs and fine-tuned model inference on Baseten's infrastructure, driven through the Baseten CLI, so customers can go from datasets to deployed models without managing GPUs. Baseten handles GPU provisioning and deprovisioning, multi-GPU clusters, and both asynchronous and dedicated inference.

Summary written by usedby from the source page, in English. The figures are those of Baseten and Oxen AI, not ours.

  • <1 minute“<1 minute training scheduling with no infrastructure overhead, for faster experimentation and iterations.”
  • 84%“84% cost savings: The total savings for inference were from $46,800 with Nano Banana to $7,530 with a fine-tuned Qwen-Image-Edit.”

Results for Oxen meant that inference time dropped by 50% through a combination of using a fine-tuned model and other model optimizations such as lightning LoRA and model graph compilation.

From the page. baseten.co

Baseten was a delight to integrate with. We never have to worry about GPU capacity, and can give our customers reliable and fast fine-tuning.

Greg Schoeninger, CEO, Oxen AI. Source

What the story claims, and what we checked

What we compared with the page.

  • The figure: 50%CheckedPrinted word for word on the page, near the name of Oxen AI.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Oxen AI.
  • The publication dateCheckedRead from the page’s own metadata, never guessed.
  • Oxen AI uses BasetenCheckedConfirmed line.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

30%“30% lower cost per call”you.com uses Baseten. Another customer of Baseten3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems uses Baseten. Another customer of Baseten33%“High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers.”OpenCode uses Baseten. Another customer of Baseten