Autonomous research agents are moving into experimentation
Scientific output is now accompanied by rapid open-source implementation. But the market is split: academic systems increasingly verify results, while most GitHub projects automate the workflow without a comparable control layer.
Research agents are growing faster than their trust infrastructure
Relevant papers rose from 51 in Q1 to 89 in Q2, a 74.5% increase. New GitHub implementations were almost flat at 155 versus 154. Verification appears in 60% of papers but only 5.8% of repositories.
Research accelerates as implementation normalizes
Targeted, deduplicated cohorts from January through June 2026.
+74.5%
-0.6%
12,465 total
current snapshot
The primary gap is result verification
Share of records where the capability is explicit in the title, abstract, description, or topics.
Verification is the clearest unresolved product category: an academic standard is forming, but it has not yet become normal in open-source implementations.
The market is still concentrated in ML engineering
Domain is conservatively classified from paper and repository text.
Medicine and materials each have nine papers but almost no new public implementations. This is an early sign of vertical supply, not a formed market.
A public stack is already forming
A separate catalog of exact laboratory and research-organization repositories; it is not mixed into the monthly cohort.
From paper to implementation in 1–42 days
Lag is measured between the first records in the same concept family and does not prove a direct paper-to-repository relationship.
What this means for the market
A new control layer
Experiment evaluation, reproducibility, tracing, and hypothesis comparison are becoming a distinct infrastructure layer.
Verticals remain open
Medicine, biology, mathematics, and materials appear in research but remain weakly represented by usable open systems.
The most autonomous agent will not necessarily win
Market advantage is shifting toward systems that can explain why an experiment was chosen and whether its result can be reproduced.
Signals tested by this research
A direct link between AILANTA's early observations and the study's measured findings.
The study uses targeted arXiv and GitHub queries from January through June 2026. GitHub returns up to 100 visible repositories per concept and month; records are deduplicated. HF Daily Papers is used as a current community-filtered validation snapshot. Product Hunt is fully excluded.
This is not a census of all papers or repositories. Capability classification relies on explicit textual evidence and is therefore a lower bound. Paper open-code share counts only explicit links or statements. Concept-family lag does not establish causality.