50%
“50% lower inference latency through a combination of using a fine-tuned model and other model optimizations”
Tal como se publicó en baseten.co, noviembre de 2025. Verificado por usedby el 9 oct 2026.
Qué pasó
Oxen AI runs its customers' model training jobs and fine-tuned model inference on Baseten's infrastructure, driven through the Baseten CLI, so customers can go from datasets to deployed models without managing GPUs. Baseten handles GPU provisioning and deprovisioning, multi-GPU clusters, and both asynchronous and dedicated inference.
Resumen escrito por usedby a partir de la página de origen, en inglés. Las cifras son de Baseten y de Oxen AI, no nuestras.
- <1 minute“<1 minute training scheduling with no infrastructure overhead, for faster experimentation and iterations.”
- 84%“84% cost savings: The total savings for inference were from $46,800 with Nano Banana to $7,530 with a fine-tuned Qwen-Image-Edit.”
Results for Oxen meant that inference time dropped by 50% through a combination of using a fine-tuned model and other model optimizations such as lightning LoRA and model graph compilation.
Baseten was a delight to integrate with. We never have to worry about GPU capacity, and can give our customers reliable and fast fine-tuning.
Lo que dice la historia, y lo que verificamos
Lo que comparamos con la página.
- La cifra: 50%VerificadoImpresa palabra por palabra en la página, cerca del nombre de Oxen AI.
- El pasaje citado arribaVerificadoCopiado palabra por palabra de la página, cerca del nombre de Oxen AI.
- La fecha de publicaciónVerificadoLeída en los metadatos de la propia página, nunca adivinada.
- Oxen AI usa BasetenVerificadoLínea de nivel Confirmado.
- El resultado en síNo verificadoLo citamos; no lo medimos.




