AILANTA
All research
GitHub · Hugging Face · npm · OpenRouter

AI infrastructure is splitting between hyperscale and the edge

A six-month cross-signal study measures centralized inference, local runtimes, routing infrastructure, model economics, and the role of Chinese open-weight ecosystems.

28.07.20266,610 records1,770 GitHub4,493 Hugging Facefrozen snapshot
Main finding

The market is forming as a barbell, not a winner-takes-all deployment stack

The barbell hypothesis is supported, with an important qualification: the two ends are visible on different surfaces. From January to June, the normalized local-project index moved +43%, centralized-inference projects +38%, and routing search intensity +148%. Hugging Face distribution moved faster than GitHub supply: Q2 produced 392 GGUF and 557 quantized models versus 118 and 125 in Q1. Routing is already operational infrastructure: 57.1% of the measured cohort supports multiple providers and 73.7% exposes cost controls.

01

Market snapshot

Four independent measurements describe different parts of the emerging market.

10.1% → 13.2%

Local share of broad AI cohort

Q1 to Q2 across the same broad GitHub sampling frame.

597

Routing repositories

57.1% support multiple providers.

59.3%

Chinese share of tracked local models

Within the fixed eight-family China/control comparison, not the whole Hub.

+65.7%

Local runtime download rate

Combined daily npm downloads, Q2 versus Q1, across six tracked packages.

02

The two ends move differently

Local supply expands more persistently. Centralized open-source supply remains large but does not accelerate after its March peak. Routing grows faster than either end.

Centralized inferenceRouting and procurementLocal and edge
Jan
Feb
Mar
Apr
May
Jun

Index of new GitHub projects and routing search matches, January = 100. Series show change within each surface, not their relative market size.

This is not symmetric growth. The hyperscale end is increasingly defined by capital, power, and proprietary infrastructure, while GitHub is better at observing the spread of local runtimes and routing tools.

03

Distribution formats make the edge practical

Models created in Q1 and Q2 inside the fixed top-by-downloads Hugging Face sample.

GGUF
118
392
+232.2%
MLX
159
466
+193.1%
ONNX
96
79
-17.7%
WebGPU
57
155
+171.9%
Quantized
125
557
+345.6%
Mobile
183
191
+4.4%
+65.7%combined daily download rate across six local runtime packages
576,442npm downloads per day in Q2
04

Routing becomes the middle of the market

597 relevant repositories show deployment-mode choice becoming an operational decision.

57.1%Multiple providers
73.7%Cost controls
67.8%Fallback and retry
78.6%Observability
62.8%Governance

The median latest-month download count across tracked routing packages is 1,191,643. This is a usage proxy, not a customer count.

05

Chinese models move into local distribution faster

Derivative models matched to declared base_model metadata inside the fixed China/control cohort.

Chinese families8.5 days

75.4% of local derivatives appear within 30 days

Control cohort24.9 days

57.1% appear within 30 days

China share of local layer59.3%

inside the eight tracked families only

06

Economics creates choice, not one winner

Cloud pricing and memory tiers are not directly comparable as TCO, but they show the range of choices a routing layer must control.

Median cloud input$0.5

per 1M tokens

Median cloud output$1.8

318 paid models

China median 4-bit weights9.4 GB

30 base models with parameter metadata

Control median 4-bit weights1.4 GB

10 base models with parameter metadata

07

Representative projects

High-traction examples from the local and centralized GitHub cohorts.