605 ms
P50 TTFT for Laguna M.1
As published on baseten.co, May 2026. Checked by usedby on Oct 10, 2026.
What happened
Poolside uses Baseten to host its Laguna coding models behind a whitelabeled, OpenAI-compatible production API. Baseten also provides per-key usage limits, billing webhooks, traffic sampling for model training, and an automated checkpoint deployment pipeline.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Poolside, not ours.
- 605ms“P50 TTFT: 146ms for Laguna XS.2 and 605ms for Laguna M.1”
- 3.9s“P90 TTFT: 1.5s for Laguna XS.2 and 3.9s for Laguna M.1”
Within 48 hours of Poolside creating their account, the model was live and responding to queries.
When we set out to make our Laguna family of models available to the world, we knew the inference layer would make or break the developer experience at launch. Baseten exceeded our quality and performance bar. The speed from conversion to production-grade, white labeled API was unlike anything we had seen from an infrastructure partner.
What the story claims, and what we checked
What we compared with the page.
- The figure: 605 msCheckedThe figure is on the page; its label is our wording.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Poolside.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Poolside uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.



