under two weeks
“Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks”
Tel que publié sur together.ai. Capturé par usedby le 3 oct. 2026.
Ce qui s’est passé
Deep Cogito, an AI research lab, uses Together AI's GPU clusters to train its open-weight Cogito reasoning models and its dedicated inference endpoints to serve them in production. Together also builds the serving pipeline for hybrid reasoning and produces quantized variants of each model release.
Résumé rédigé par usedby à partir de la page source, en anglais. Les chiffres sont ceux de Together AI et de Deep Cogito, pas les nôtres.
- sub-500ms“Together delivers sub-500ms time to first token at sustained throughput of over 1,000 requests per minute with 99.9% uptime.”
- 1 million downloads“These releases drove significant adoption, contributing to over 1 million downloads across Hugging Face and Ollama.”
After each training run, Together’s team shipped FP8, FP4, and INT4 quantized variants of the Cogito models in under two weeks, without meaningful quality degradation.
Before Together, we were struggling with adapting to the compute environment for training large open models. Once we found Together, we were able to offload most of it to them and focus on model training and benchmarking, which is our core secret sauce.
Ce que dit le témoignage, et ce que nous avons vérifié
Nous avons comparé le témoignage à sa page en ligne le 3 oct. 2026.
- Le chiffre : under two weeksVérifiéImprimé mot pour mot sur la page, près du nom de Deep Cogito.
- L’extrait cité plus hautVérifiéCopié mot pour mot depuis la page, près du nom de Deep Cogito.
- Deep Cogito utilise Together AIVérifiéLigne de niveau Confirmé. Dernière vérification, toutes sources confondues : 3 oct. 2026.
- Le résultat lui-mêmeNon vérifiéNous le citons ; nous ne l’avons pas mesuré.




