3x
“saved the company 3x in inference costs compared to using OpenAI API for all listings”
As published on anyscale.com. Captured by usedby on Oct 7, 2026.
What happened
Mirakl uses Anyscale's managed Ray platform to run fine-tuned, smaller open-source LLMs with LoRA adapters for batch inference. This automates catalog onboarding for marketplace sellers, with GPU clusters that autoscale to match traffic.
Summary written by usedby from the source page, in English. The figures are those of Anyscale and Mirakl, not ours.
- 80k+ token/sec“Consistent peak performance, handling 80k+ token/sec generations without latency.”
- 10 million“automating onboarding for over 10 million products each month”
- 20 GPU nodes“they can scale up to 20 GPU nodes during peak hours, but scale down to near-zero GPUs overnight”
- 90%“By intelligently routing up to 90% of catalog traffic to LLaMA 3.1 8B”
Investing in fine-tuned open-source models and leveraging Anyscale to manage Ray orchestration saved the company 3x in inference costs compared to using OpenAI API for all listings.
Using GPT-4 would have cost us more than half a million dollars per month… with Anyscale and fine-tuned LLaMA models, we brought that down by 3× while scaling to millions of products.
What the story claims, and what we checked
We compared the story with its live page on Oct 7, 2026.
- The figure: 3xCheckedPrinted word for word on the page, near the name of Mirakl.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Mirakl.
- Mirakl uses AnyscaleCheckedConfirmed line. Latest check across sources: Oct 7, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




