Voice AI is testing direct speech-to-speech runtimes
GitHub and Hugging Face cohorts compare native audio-in/audio-out systems with transcription-based voice-agent pipelines.
Early evidence; the market is not yet confirmed
84 subject-filtered models and implementations show an early split in voice-agent architecture. Explicit native audio-in/audio-out appears in 3.6% of the cohort, transcription cascades in 19%, real-time streaming in 57.1%, and interruption control in 9.5%. Direct speech-to-speech is now measurable, but it has not yet displaced modular STT–LLM–TTS pipelines.
Market snapshot
Comparable measurements from independent market surfaces.
Implementations and packages
Deduplicated and subject-filtered primary cohort.
Median repository stars
Calculated across repositories with at least one star.
Native audio-in/audio-out
3 items in the subject-filtered cohort.
Transcription-based pipelines
16 items in the subject-filtered cohort.
Search-match dynamics
Bars combine monthly GitHub query matches with captured Hugging Face models. Feature analysis separates explicit native-audio systems from STT–LLM–TTS cascades.
What exists inside the category
One item may contain more than one feature.
Native audio-in/audio-out
3 · 3.6%
STT–LLM–TTS cascades
16 · 19%
Real-time streaming
48 · 57.1%
Interruption and turn-taking
8 · 9.5%
Prosody and emotion
2 · 2.4%
Local voice deployment
18 · 21.4%
Multilingual speech
6 · 7.1%
Representative projects
Cross-source validation
These publications are not part of the primary numeric cohort.