Voice Agents Drop Speech Transcription
An open voice-agent stack now sends audio directly to a multimodal language model instead of converting speech to text first. Removing transcription could preserve prosody, simplify latency-sensitive pipelines and change the economics of real-time voice products. This is one implementation, so the line remains on watch until independent stacks demonstrate lower latency, better task performance or production adoption without a separate STT stage.