under two weeks
“Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks”
Tal como se publicó en together.ai. Capturado por usedby el 3 oct 2026.
Qué pasó
Deep Cogito, an AI research lab, uses Together AI's GPU clusters to train its open-weight Cogito reasoning models and its dedicated inference endpoints to serve them in production. Together also builds the serving pipeline for hybrid reasoning and produces quantized variants of each model release.
Resumen escrito por usedby a partir de la página de origen, en inglés. Las cifras son de Together AI y de Deep Cogito, no nuestras.
- sub-500ms“Together delivers sub-500ms time to first token at sustained throughput of over 1,000 requests per minute with 99.9% uptime.”
- 1 million downloads“These releases drove significant adoption, contributing to over 1 million downloads across Hugging Face and Ollama.”
After each training run, Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks, without meaningful quality degradation.
Before Together, we were struggling with adapting to the compute environment for training large open models. Once we found Together, we were able to offload most of it to them and focus on model training and benchmarking, which is our core secret sauce.
Lo que dice la historia, y lo que verificamos
Comparamos la historia con su página en línea el 3 oct 2026.
- La cifra: under two weeksVerificadoImpresa palabra por palabra en la página, cerca del nombre de Deep Cogito.
- El pasaje citado arribaVerificadoCopiado palabra por palabra de la página, cerca del nombre de Deep Cogito.
- Deep Cogito usa Together AIVerificadoLínea de nivel Confirmado. Última verificación entre todas las fuentes: 3 oct 2026.
- El resultado en síNo verificadoLo citamos; no lo medimos.




