under two weeks
“Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks”
As published on together.ai. Captured by usedby on Oct 3, 2026.
What happened
Deep Cogito, an AI research lab, uses Together AI's GPU clusters to train its open-weight Cogito reasoning models and its dedicated inference endpoints to serve them in production. Together also builds the serving pipeline for hybrid reasoning and produces quantized variants of each model release.
Summary written by usedby from the source page, in English. The figures are those of Together AI and Deep Cogito, not ours.
- sub-500ms“Together delivers sub-500ms time to first token at sustained throughput of over 1,000 requests per minute with 99.9% uptime.”
- 1 million downloads“These releases drove significant adoption, contributing to over 1 million downloads across Hugging Face and Ollama.”
After each training run, Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks, without meaningful quality degradation.
Before Together, we were struggling with adapting to the compute environment for training large open models. Once we found Together, we were able to offload most of it to them and focus on model training and benchmarking, which is our core secret sauce.
What the story claims, and what we checked
We compared the story with its live page on Oct 3, 2026.
- The figure: under two weeksCheckedPrinted word for word on the page, near the name of Deep Cogito.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Deep Cogito.
- Deep Cogito uses Together AICheckedConfirmed line. Latest check across sources: Oct 3, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




