~300ms
“~300ms reduction in end-to-end agent response time after integrating Deepgram’s streaming STT”
As published on deepgram.com. Checked by usedby on Oct 9, 2026.
What happened
SigmaMind AI uses Deepgram's Nova-3 and Flux models as the default real-time speech-to-text engine in its no-code platform for building voice AI agents. Transcripts, including interim results, feed its orchestration engine so agents can start acting before a speaker finishes.
Summary written by usedby from the source page, in English. The figures are those of Deepgram and SigmaMind AI, not ours.
- 50%“50% increase in outbound call conversion for a call center customer that migrated to SigmaMind, going live in just two weeks”
- 1 million+“1 million+ calls per month processed through SigmaMind’s platform, with 200+ hours of speech transcribed daily”
- 150“150 peak concurrent voice sessions handled per customer deployment without degradation”
- Sub-1-second“Sub-1-second voice-to-voice latency including telephony overhead, enabling natural conversational pacing”
By integrating Deepgram’s Nova-3, and Flux speech-to-text models as the default real-time transcription engine, SigmaMind reduced end-to-end agent response latency by roughly 300 milliseconds and enabled a new class of voice workflows where agents act on speech before a sentence is even finished.
When we began acting on interim transcripts and combined that with word timestamps, the agent could trigger API calls and follow-ups mid-utterance,” said Pratik Mundra, co-founder of SigmaMind AI. “That shift unlocked much richer, multi-step voice workflows.
What the story claims, and what we checked
What we compared with the page.
- The figure: ~300msCheckedPrinted word for word on the page, near the name of SigmaMind AI.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of SigmaMind AI.
- SigmaMind AI uses DeepgramCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




