AILANTA
All research
GitHub · Hugging Face · Discovery evidence

Voice AI is testing direct speech-to-speech runtimes

GitHub and Hugging Face cohorts compare native audio-in/audio-out systems with transcription-based voice-agent pipelines.

2026-08-1284 items2 source types1 validation posts
Main finding

Early evidence; the market is not yet confirmed

84 subject-filtered models and implementations show an early split in voice-agent architecture. Explicit native audio-in/audio-out appears in 3.6% of the cohort, transcription cascades in 19%, real-time streaming in 57.1%, and interruption control in 9.5%. Direct speech-to-speech is now measurable, but it has not yet displaced modular STT–LLM–TTS pipelines.

01

Market snapshot

Comparable measurements from independent market surfaces.

84

Implementations and packages

Deduplicated and subject-filtered primary cohort.

4

Median repository stars

Calculated across repositories with at least one star.

3.6%

Native audio-in/audio-out

3 items in the subject-filtered cohort.

19%

Transcription-based pipelines

16 items in the subject-filtered cohort.

02

Search-match dynamics

Bars combine monthly GitHub query matches with captured Hugging Face models. Feature analysis separates explicit native-audio systems from STT–LLM–TTS cascades.

Jan 26556 loaded
Feb 26929 loaded
Mar 2610110 loaded
Apr 268312 loaded
May 2611612 loaded
Jun 2610714 loaded
Jul 269818 loaded
Aug 26123 loaded
03

What exists inside the category

One item may contain more than one feature.

Native audio-in/audio-out

3 · 3.6%

STT–LLM–TTS cascades

16 · 19%

Real-time streaming

48 · 57.1%

Interruption and turn-taking

8 · 9.5%

Prosody and emotion

2 · 2.4%

Local voice deployment

18 · 21.4%

Multilingual speech

6 · 7.1%

04

Representative projects

05

Cross-source validation

These publications are not part of the primary numeric cohort.