Local AI is becoming a deployment stack
First-half data shows local model formats expanding, desktop runtimes growing, and browser package usage rising sharply. This is no longer a narrow reaction to cloud cost.
The market is moving at two speeds
The local-oriented share of a broad GitHub AI cohort rose from 10.1% to 13.2%. New supply is concentrating around GGUF and llama.cpp, while ONNX and WebGPU produce fewer new repositories but their established browser packages are seeing much higher usage.
Local share within AI tooling
The same broad GitHub selection method in both quarters.
+3.1 pp
18% Q/Q
+56.1%
+35.3%
New GitHub runtimes by month
Targeted searches deduplicated by repository ID.
April was the peak month for new local runtime projects. June's decline is not proof of market contraction: newer repositories had less time to reach the five-star inclusion threshold.
Formats are diverging
A project may match more than one technical format.
Hugging Face model supply is accelerating
Top-by-downloads cohorts within each filter; categories overlap.
GGUF, MLX, and quantized models show the clearest supply expansion. WebGPU model count increased, but median downloads remain extremely low—early supply rather than validated mass adoption.
Browser inference is already being used
Average daily downloads for a fixed benchmark of npm packages.
What this means for the market
Local AI is no longer one category
Distinct layers are forming around models, runtimes, desktop shells, browser inference, and mobile deployment.
ONNX is shifting from creation to usage
New ONNX repositories declined, while onnxruntime-web average daily downloads rose 107.8%. This signals maturing infrastructure, not disappearing demand.
Desktop currently leads mobile
New supply is concentrating in desktop runtimes and execution cores. Mobile models are present, but the product layer is developing more slowly.
Signals tested by this research
A direct link between AILANTA's early observations and the study's measured findings.
GitHub combines a broad control cohort with targeted searches across ten runtime themes from January–June 2026; repositories are deduplicated. Hugging Face takes up to 1,000 current top-by-downloads models per filter and then restricts them to H1 2026 creation dates. npm compares average daily downloads for a fixed package benchmark.
Hugging Face downloads and likes are current metrics, not historical snapshots. The Hugging Face sample does not estimate the entire Hub. GitHub stars introduce age and survivorship bias. The npm benchmark is not a full market-share estimate. Model size is excluded because the list API did not return sufficient usedStorage data.