Skip to main content
usedby

under two weeks

“Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks”

As published on together.ai. Captured by usedby on Oct 3, 2026.

What happened

Deep Cogito, an AI research lab, uses Together AI's GPU clusters to train its open-weight Cogito reasoning models and its dedicated inference endpoints to serve them in production. Together also builds the serving pipeline for hybrid reasoning and produces quantized variants of each model release.

Summary written by usedby from the source page, in English. The figures are those of Together AI and Deep Cogito, not ours.

  • sub-500ms“Together delivers sub-500ms time to first token at sustained throughput of over 1,000 requests per minute with 99.9% uptime.”
  • 1 million downloads“These releases drove significant adoption, contributing to over 1 million downloads across Hugging Face and Ollama.”

After each training run, Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks, without meaningful quality degradation.

From the page. together.ai, captured Oct 3, 2026

Before Together, we were struggling with adapting to the compute environment for training large open models. Once we found Together, we were able to offload most of it to them and focus on model training and benchmarking, which is our core secret sauce.

Dhruv Malrana, Co-founder, Deep Cogito, Deep Cogito. Source, captured Oct 3, 2026

What the story claims, and what we checked

We compared the story with its live page on Oct 3, 2026.

  • The figure: under two weeksCheckedPrinted word for word on the page, near the name of Deep Cogito.
  • The passage quoted aboveCheckedCopied word for word from the page, near the name of Deep Cogito.
  • Deep Cogito uses Together AICheckedConfirmed line. Latest check across sources: Oct 3, 2026.
  • The result itselfNot checkedWe quote it; we did not measure it.

Same company, same tool or same industry.

6דDecagon achieved nearly 6× cost reduction per turn compared to closed models like GPT-5 mini.”Decagon uses Together AI. Another customer of Together AI70%“70% cost reduction for customers switching to open-source models on Together GPUs”Runware uses Together AI. Another customer of Together AI3–4 months“3–4 months of cumulative research time recovered”Scaled Cognition uses Together AI. Another customer of Together AI