300ms
“These optimizations cut latency from above a second to 300ms on Praktika's transcription workload, vs 1,000-1,500ms before.”
As published on baseten.co, July 2025. Checked by usedby on Oct 9, 2026.
What happened
Praktika runs its speech-to-text (Whisper transcription) workload on Baseten to power AI avatar language tutors, with an optimized runtime and autoscaling to keep conversational latency low. They shifted their traffic from a previous cloud vendor's inference solution using OpenAI compatible APIs.
Summary written by usedby from the source page, in English. The figures are those of Baseten and Praktika, not ours.
- 50%“50% cost savings”
- 7“Praktika recently launched 7 new languages.”
These optimizations cut latency from above a second to 300ms on Praktika's transcription workload, vs 1,000-1,500ms before.
We are very satisfied with the latency we achieved. The user experience feels so much faster. It's so fast we actually don't need as many replicas which helps keep costs down.
What the story claims, and what we checked
What we compared with the page.
- The figure: 300msCheckedPrinted word for word on the page, near the name of Praktika.
- The passage quoted aboveCheckedCopied word for word from the page, near the name of Praktika.
- The publication dateCheckedRead from the page’s own metadata, never guessed.
- Praktika uses BasetenCheckedConfirmed line.
- The result itselfNot checkedWe quote it; we did not measure it.




