Ir al contenido principal
usedby

50%

“50% lower inference latency through a combination of using a fine-tuned model and other model optimizations”

Tal como se publicó en baseten.co, noviembre de 2025. Verificado por usedby el 9 oct 2026.

Qué pasó

Oxen AI runs its customers' model training jobs and fine-tuned model inference on Baseten's infrastructure, driven through the Baseten CLI, so customers can go from datasets to deployed models without managing GPUs. Baseten handles GPU provisioning and deprovisioning, multi-GPU clusters, and both asynchronous and dedicated inference.

Resumen escrito por usedby a partir de la página de origen, en inglés. Las cifras son de Baseten y de Oxen AI, no nuestras.

  • <1 minute“<1 minute training scheduling with no infrastructure overhead, for faster experimentation and iterations.”
  • 84%“84% cost savings: The total savings for inference were from $46,800 with Nano Banana to $7,530 with a fine-tuned Qwen-Image-Edit.”

Results for Oxen meant that inference time dropped by 50% through a combination of using a fine-tuned model and other model optimizations such as lightning LoRA and model graph compilation.

De la página. baseten.co

Baseten was a delight to integrate with. We never have to worry about GPU capacity, and can give our customers reliable and fast fine-tuning.

Greg Schoeninger, CEO, Oxen AI. Fuente

Lo que dice la historia, y lo que verificamos

Lo que comparamos con la página.

  • La cifra: 50%VerificadoImpresa palabra por palabra en la página, cerca del nombre de Oxen AI.
  • El pasaje citado arribaVerificadoCopiado palabra por palabra de la página, cerca del nombre de Oxen AI.
  • La fecha de publicaciónVerificadoLeída en los metadatos de la propia página, nunca adivinada.
  • Oxen AI usa BasetenVerificadoLínea de nivel Confirmado.
  • El resultado en síNo verificadoLo citamos; no lo medimos.

Misma empresa, misma herramienta o mismo sector.

30%“30% lower cost per call”you.com usa Baseten. Otro cliente de Baseten3x“Baseten’s throughput and latency optimizations delivered 3x savings versus the equivalent closed-source model”Parallel Web Systems usa Baseten. Otro cliente de Baseten33%“High cache hit rates and cache token pricing enabled OpenCode to achieve a 33% blended cost reduction versus their previous providers.”OpenCode usa Baseten. Otro cliente de Baseten