# How vLLM uses Nebius AI Cloud

As published on [nebius.com](https://nebius.com/customer-stories/vllm). Captured by usedby on 2026-10-07.

- Company: [vLLM](https://www.usedby.ai/companies/vllm.md)
- Tool: [Nebius AI Cloud](https://www.usedby.ai/tools/nebius.md)

## What the story says

vLLM, an open-source LLM inference framework, uses Nebius compute clusters and storage to test, benchmark and optimize inference for large models such as DeepSeek R1. This includes validating optimizations and RLHF workloads before releasing them to the community.

Summary written by usedby from the source page, in English. The figures are those of Nebius AI Cloud and vLLM, not ours.

> By utilizing compute clusters, vLLM successfully scaled up inference experiments, integrating cutting-edge optimizations like multi-latent attention and multi-token prediction from the DeepSeek research paper into vLLM.

## What usedby checked

We compared the story with its live page on 2026-10-07.

- Checked: the passage quoted above is copied word for word from the page, near the name of vLLM.
- Checked: each number in our summary is printed on the page.
- Checked: vLLM uses Nebius AI Cloud. Confirmed line. Latest check across sources: 2026-10-07.
- Not checked: the result itself. We quote it; we did not measure it.

---
Source: https://www.usedby.ai/case-studies/vllm-nebius · How we check: https://www.usedby.ai/methodology
