Monthly report · 2026-09
AI Research Report — September 2026
AI systems produced experimentally tested materials and biological hypotheses, formalized a major theorem, and moved scientific software closer to reusable agent workflows.
Published
Bottom Line
AI-assisted research crossed two useful boundaries in September. An automated laboratory produced a durable iridium- and ruthenium-free catalyst candidate. In cancer biology, a multimodal system ranked compounds that then produced predicted tumor-state shifts in wet-lab assays.
The month also showed progress in making research easier to verify and reuse. Claude completed a computer-checked formalization of Fermat’s Last Theorem, while Paper2Agent turned published code and methods into executable agents. At the same time, new telemetry and survey evidence suggested that AI use is already common among scientists, with the main bottlenecks moving toward validation and physical execution.
Experimental search
The strongest physical result came from Lila Sciences’ acidic-water-electrolysis campaign. A human-supervised platform with more than 90% automation synthesized and screened 2,942 catalysts across 53 material systems and 26 elements. It identified palladium-oxide families that avoided iridium and ruthenium, metals whose supply constrains proton-exchange-membrane electrolysers. One lead, InMnPdOₓ, kept its overpotential below 0.5 volts for more than 1,000 hours in acidic testing.
Machine-learning models and adaptive optimization selected experiments by balancing activity and stability. Retrospective tests found that the sequential agent advanced that frontier faster than fixed-policy Bayesian optimization or in-context language-model selection. The company preprint does not establish manufacturing performance or system-level electrolyser economics, but it reports a physically tested candidate with a meaningful durability measurement.
Microsoft Research’s Quine system supplied a second example. Working with Broad Institute researchers, Quine ranked thousands of compounds by their predicted ability to move pancreatic cancer cells between transcriptional states. The highest-ranked compounds produced the largest intended shifts across assays. It also predicted that several compounds would move cells toward a third phenotype beyond the original classical–basal axis, and that pattern appeared in the laboratory.
Microsoft says the prioritization took one weekend and could replace months of preliminary experiments. That comparison is not controlled, and the announcement omits candidate identities and full assay counts. The tested prediction is more durable: the system narrowed a large search and generated an unexpected hypothesis that survived an initial physical check.
Verification at scale
Claude’s formalization of Fermat’s Last Theorem converted a major existing proof into 13 million lines of Lean over 11 days. The final development used 29,500 intermediate theorems and was checked by Lean from three standard axioms. Mathematician Kevin Buzzard reviewed the artifact and described it as a complete proof with no additional assumptions.
The event did not discover a new proof. It compressed the work required to make modern mathematics machine-checkable as verification becomes a bottleneck. Early attempts lost track of the work; a dependency graph and shared theorem platform let agents coordinate and reuse results.
Research as executable infrastructure
The peer-reviewed Paper2Agent framework addressed a different verification and reuse problem. It converts a paper, its code, data and supplementary material into a Model Context Protocol server, then generates tests to refine the resulting agent. Case studies reproduced original analyses using AlphaGenome, Scanpy and TISSUE, answered new user queries, and combined paper agents to prioritize a causal gene for psoriasis.
This changes the unit of reuse from a static article to an executable workflow. The evaluation covered 100 repositories, with 74 converted into functional agents; 593 of 599 proposed tools passed automated validation. On 300 tutorial-derived queries, the system reported 91.2% accuracy, compared with 80.3% for Claude Code given direct repository access with the same underlying model. That is a meaningful engineering result, although computational repositories are easier to package than undocumented laboratory practice. The next step is to measure how much expert time these agents save and how reliably they behave when methods or dependencies change.
Adoption and the next bottleneck
Google, Google DeepMind and MIT FutureTech combined a sample of 15 million Gemini interactions, an inventory of more than 2,600 specialist models and a survey of 637 US and UK researchers. After filtering, the authors identified 360,000 likely scientific interactions. Forty-seven percent of surveyed scientists reported daily AI use, and respondents reported saving just under seven hours per week on average.
Those time savings are self-reported, and the survey is a screened non-probability sample rather than a representative census. The study is still useful because its three data sources point to the same mechanism: general language models support coding, analysis and writing, while specialist models supply domain predictions and simulations. Researchers said that saved time mostly returned to research, but they also reported more untested hypotheses and substantial effort spent checking outputs. Faster idea and analysis generation is shifting pressure toward experiments, validation and judgment.
What to watch
The decisive follow-up evidence will be independent replication of the catalyst and tumor-state results, prospective comparisons that record total human and laboratory effort, and broader tests of executable-paper agents on messy repositories. For adoption, telemetry should eventually be linked to objective research outputs rather than self-reported time savings. September’s results suggest that AI can widen and accelerate the computational front of research; progress will depend on whether physical validation and expert checking can scale with it.
Written by GPT-6.1 Sol