3x
“saved the company 3x in inference costs compared to using OpenAI API for all listings”
Tel que publié sur anyscale.com. Capturé par usedby le 7 oct. 2026.
Ce qui s’est passé
Mirakl uses Anyscale's managed Ray platform to run fine-tuned, smaller open-source LLMs with LoRA adapters for batch inference. This automates catalog onboarding for marketplace sellers, with GPU clusters that autoscale to match traffic.
Résumé rédigé par usedby à partir de la page source, en anglais. Les chiffres sont ceux de Anyscale et de Mirakl, pas les nôtres.
- 80k+ token/sec“Consistent peak performance, handling 80k+ token/sec generations without latency.”
- 10 million“automating onboarding for over 10 million products each month”
- 20 GPU nodes“they can scale up to 20 GPU nodes during peak hours, but scale down to near-zero GPUs overnight”
- 90%“By intelligently routing up to 90% of catalog traffic to LLaMA 3.1 8B”
Investing in fine-tuned open-source models and leveraging Anyscale to manage Ray orchestration saved the company 3x in inference costs compared to using OpenAI API for all listings.
Using GPT-4 would have cost us more than half a million dollars per month… with Anyscale and fine-tuned LLaMA models, we brought that down by 3× while scaling to millions of products.
Ce que dit le témoignage, et ce que nous avons vérifié
Nous avons comparé le témoignage à sa page en ligne le 7 oct. 2026.
- Le chiffre : 3xVérifiéImprimé mot pour mot sur la page, près du nom de Mirakl.
- L’extrait cité plus hautVérifiéCopié mot pour mot depuis la page, près du nom de Mirakl.
- Mirakl utilise AnyscaleVérifiéLigne de niveau Confirmé. Dernière vérification, toutes sources confondues : 7 oct. 2026.
- Le résultat lui-mêmeNon vérifiéNous le citons ; nous ne l’avons pas mesuré.




