50%
“50% lower inference latency through a combination of using a fine-tuned model and other model optimizations”
Tel que publié sur baseten.co, novembre 2025. Vérifié par usedby le 9 oct. 2026.
Ce qui s’est passé
Oxen AI runs its customers' model training jobs and fine-tuned model inference on Baseten's infrastructure, driven through the Baseten CLI, so customers can go from datasets to deployed models without managing GPUs. Baseten handles GPU provisioning and deprovisioning, multi-GPU clusters, and both asynchronous and dedicated inference.
Résumé rédigé par usedby à partir de la page source, en anglais. Les chiffres sont ceux de Baseten et de Oxen AI, pas les nôtres.
- <1 minute“<1 minute training scheduling with no infrastructure overhead, for faster experimentation and iterations.”
- 84%“84% cost savings: The total savings for inference were from $46,800 with Nano Banana to $7,530 with a fine-tuned Qwen-Image-Edit.”
Results for Oxen meant that inference time dropped by 50% through a combination of using a fine-tuned model and other model optimizations such as lightning LoRA and model graph compilation.
Baseten was a delight to integrate with. We never have to worry about GPU capacity, and can give our customers reliable and fast fine-tuning.
Ce que dit le témoignage, et ce que nous avons vérifié
Ce que nous avons comparé à la page.
- Le chiffre : 50%VérifiéImprimé mot pour mot sur la page, près du nom de Oxen AI.
- L’extrait cité plus hautVérifiéCopié mot pour mot depuis la page, près du nom de Oxen AI.
- La date de publicationVérifiéLue dans les métadonnées de la page, jamais devinée.
- Oxen AI utilise BasetenVérifiéLigne de niveau Confirmé.
- Le résultat lui-mêmeNon vérifiéNous le citons ; nous ne l’avons pas mesuré.




