75%
“an overall 75% reduction in GPU compute cost”
As published on anyscale.com. Captured by usedby on Oct 7, 2026.
What happened
Coactive AI runs its Multimodal AI Platform on Anyscale's Ray-based managed compute, deployed in its own Kubernetes clusters on AWS and Azure. It uses Ray Serve for model serving and large-scale processing of image and video data, with fractional GPU allocation and autoscaling.
Summary written by usedby from the source page, in English. The figures are those of Anyscale and Coactive, not ours.
- 4x“an overall 4x cheaper per-image processing”
- 1 day“1 day to deploy new multimodal model endpoints, down from 1+ week”
- 25%“service definitions shrank to roughly 25% of their previous code size”
Anyscale delivered major cost gains through fractional GPU allocations, allowing Coactive to pack multiple model replicas onto a single GPU and directly reduce the number of GPUs needed for the same workload.
One of our applied AI engineers said, ‘we should use this model,’ and the next day it was running in production. Before Anyscale, that would’ve taken a week or more.
What the story claims, and what we checked
We compared the story with its live page on Oct 7, 2026.
- The figure: 75%CheckedPrinted word for word on the page, near the name of Coactive.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Coactive.
- Coactive uses AnyscaleCheckedConfirmed line. Latest check across sources: Oct 7, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




