AI infrastructure is splitting between hyperscale and the edge
A six-month cross-signal study measures centralized inference, local runtimes, routing infrastructure, model economics, and the role of Chinese open-weight ecosystems.
The market is forming as a barbell, not a winner-takes-all deployment stack
The barbell hypothesis is supported, with an important qualification: the two ends are visible on different surfaces. From January to June, the normalized local-project index moved +43%, centralized-inference projects +38%, and routing search intensity +148%. Hugging Face distribution moved faster than GitHub supply: Q2 produced 392 GGUF and 557 quantized models versus 118 and 125 in Q1. Routing is already operational infrastructure: 57.1% of the measured cohort supports multiple providers and 73.7% exposes cost controls.
Market snapshot
Four independent measurements describe different parts of the emerging market.
Local share of broad AI cohort
Q1 to Q2 across the same broad GitHub sampling frame.
Routing repositories
57.1% support multiple providers.
Chinese share of tracked local models
Within the fixed eight-family China/control comparison, not the whole Hub.
Local runtime download rate
Combined daily npm downloads, Q2 versus Q1, across six tracked packages.
The two ends move differently
Local supply expands more persistently. Centralized open-source supply remains large but does not accelerate after its March peak. Routing grows faster than either end.
This is not symmetric growth. The hyperscale end is increasingly defined by capital, power, and proprietary infrastructure, while GitHub is better at observing the spread of local runtimes and routing tools.
Distribution formats make the edge practical
Models created in Q1 and Q2 inside the fixed top-by-downloads Hugging Face sample.
Routing becomes the middle of the market
597 relevant repositories show deployment-mode choice becoming an operational decision.
The median latest-month download count across tracked routing packages is 1,191,643. This is a usage proxy, not a customer count.
Chinese models move into local distribution faster
Derivative models matched to declared base_model metadata inside the fixed China/control cohort.
Economics creates choice, not one winner
Cloud pricing and memory tiers are not directly comparable as TCO, but they show the range of choices a routing layer must control.
Representative projects
High-traction examples from the local and centralized GitHub cohorts.
Signals tested by this research
A direct link between AILANTA's early observations and the study's measured findings.
AI Model Routing
This signal matches the published research scope: A six-month cross-signal study of centralized inference, local runtimes, routing, model economics, and Chinese open-weight distribution.
AI Compute Capacity
This signal matches the published research scope: A six-month cross-signal study of centralized inference, local runtimes, routing, model economics, and Chinese open-weight distribution.
Local AI Runtime
This signal matches the published research scope: A six-month cross-signal study of centralized inference, local runtimes, routing, model economics, and Chinese open-weight distribution.
China Full-Stack AI
This signal matches the published research scope: A six-month cross-signal study of centralized inference, local runtimes, routing, model economics, and Chinese open-weight distribution.
Commodity Voice AI
Voice models provide an early application-level example of infrastructure moving to the edge.