50%
“50% lower inference latency through a combination of using a fine-tuned model and other model optimizations”
As published on baseten.co, November 2025. Checked by usedby on Oct 9, 2026.
What happened
Oxen AI runs its customers' model training jobs and fine-tuned model inference on Baseten's infrastructure, driven through the Baseten CLI, so customers can go from datasets to deployed models without managing GPUs. Baseten handles GPU provisioning and deprovisioning, multi-GPU clusters, and both asynchronous and dedicated inference.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Oxen AI, not ours.
- <1 minute“<1 minute training scheduling with no infrastructure overhead, for faster experimentation and iterations.”
- 84%“84% cost savings: The total savings for inference were from $46,800 with Nano Banana to $7,530 with a fine-tuned Qwen-Image-Edit.”
Results for Oxen meant that inference time dropped by 50% through a combination of using a fine-tuned model and other model optimizations such as lightning LoRA and model graph compilation.
Baseten was a delight to integrate with. We never have to worry about GPU capacity, and can give our customers reliable and fast fine-tuning.
What the story claims, and what we checked
What we compared with the page.
- The figure: 50%CheckedPrinted word for word on the page, near the name of Oxen AI.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Oxen AI.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Oxen AI uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




