07.02

Speech pipelines: ASR and TTS

skillmaxingInference EngineeringSpeech pipelines: ASR and TTS
Visual lesson

Speech pipelines: ASR and TTS

Included with Pro · 4:53
Lessons · 1,172 words

Text serving has a forgiving user: if a chat answer streams a little slowly, people keep reading. Audio has no such slack. Speech arrives at a fixed rate and synthesized speech must leave at a fixed rate, so a stage that falls behind is heard as a gap, a stutter, or an awkward silence.

Two numbers decide everything: how fast each stage runs compared with real time, and how long the whole chain takes from the user's last word to the first sound of the reply. Kite's voice mode chains four stages: voice activity detection (VAD), speech recognition (ASR), the chat LLM, and text-to-speech (TTS).

Pro

Continue reading

Pro unlocks every video lesson, the full notes, worked sizing and cost calculations, and runnable configs across the Inference Engineering course, from GPUs and kernels to quantization, speculative decoding, disaggregation and production autoscaling.

Unlock full access to
Python with AIInference EngineeringAI researchComplete system design
30% off with LAUNCH30first payment
Buy now
View all plans ·Already Pro? Sign in