90 ms
“Goodcall achieves a 90 ms on average for time-to-first-audio”
As published on cartesia.ai, October 2024. Captured by usedby on Oct 8, 2026.
What happened
Goodcall runs an AI phone agent for reception, sales and service, and uses Cartesia's Sonic text-to-speech to generate the voices for its agents. It moved all of its text-to-speech generation from ElevenLabs to Cartesia.
Summary written by usedby from the source page, in English. The figures are those of Cartesia and Goodcall, not ours.
- 97%“consistently hit 97% interaction rate”
- 2,217“Goodcall has switched 100% of their text-to-speech generation for 2,217 unique voice agents from Eleven Labs to Cartesia.”
Industry-Leading Latency: With Sonic, Goodcall achieves a 90 ms on average for time-to-first-audio, including both model and network latency, which is more than four times as fast as the performance experienced with Eleven Labs.
I became an early adopter of Cartesia the day they launched as soon as I saw how low their latency was.
What the story claims, and what we checked
We compared the story with its live page on Oct 8, 2026.
- The figure: 90 msCheckedPrinted word for word on the page, near the name of Goodcall.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Goodcall.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Goodcall uses CartesiaCheckedConfirmed line. Latest check across sources: Oct 8, 2026.
- The result itselfNot checkedWe quote it; we did not measure it.




