2×
“SGLang achieved a 2× boost in throughput and markedly lower latency on one node.”
As published on nebius.com. Captured by usedby on Oct 7, 2026.
What happened
SGLang, an open-source serving framework for large language models, ran and benchmarked DeepSeek R1 on Nebius AI Cloud infrastructure. It used on-demand compute clusters to test inference optimizations such as new attention algorithms, FP8 matrix multiplication and kernel fusion.
Summary written by usedby from the source page, in English. The figures are those of Nebius AI Cloud and SGLang, not ours.
SGLang achieved a 2× boost in throughput and markedly lower latency on one node. In practice, this means faster answers from R1, even on long prompts or with dozens of users at once.
What the story claims, and what we checked
We compared the story with its live page on Oct 7, 2026.
- The figure: 2×CheckedPrinted word for word on the page, near the name of SGLang.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of SGLang.
- The numbers in our summaryCheckedEach one is printed on the page.
- SGLang uses Nebius AI CloudCheckedConfirmed line. Latest check across sources: Oct 7, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




