Monthly report · 2026-07

AI Research Report — July 2026

AI systems contributed to new mathematical results and worked more effectively with biological data and scientific software, while adoption widened across research workflows.

Published

Bottom Line

July produced a notable mathematical claim: researchers reported five new results in Banach-space theory for which language models supplied key ideas and draft proofs. The work is still a preprint awaiting independent checking, but it moves the discussion from performance on known problems toward contributions to current mathematics.

Elsewhere, scientific systems became more capable at working through structured data and deterministic software. A peer-reviewed single-cell benchmark found that AI scientist systems could approach graduate-level performance on parts of biological analysis, while a geochemistry benchmark showed agents building simulator inputs, running scientific software, and interpreting the outputs. These are narrow settings, but they test genuine components of research work.

Mathematics

In a July preprint, Itai and Ofir Acuaviva describe five new results in Banach-space theory developed with substantial language-model assistance. The models generated proof ideas and draft arguments; the authors then checked and refined the mathematics. The paper also describes a workflow that searches the literature for open problems and attempts candidate solutions.

The central evidence is the mathematical argument itself, not a benchmark score. That makes the work more consequential and easier for specialists to interrogate. It is also why external review matters: the paper does not identify peer review, independent expert reconstruction, or machine-checked formal proofs. If the results survive that process, the case will offer a clear example of AI contributing to the production of new mathematics.

Biological data analysis

BAISBench, published in Bioinformatics, evaluates AI scientist systems on real single-cell transcriptomic datasets. Its first task covers cell-type annotation across 15 expert-labeled datasets. Its second uses 193 questions derived from the conclusions of 41 published single-cell studies and asks systems to recover those findings through data analysis. Six graduate-level bioinformaticians provide the human comparison.

The systems approached graduate-level performance on some discovery-oriented tasks, especially where the analysis could be executed through a defined pipeline. Performance was weaker when success depended on biological judgment and interpretation. The benchmark therefore locates useful capability inside a real research workflow: processing data and recovering supported patterns, rather than generating and validating an entirely new biological claim.

Scientific software

PHREEQC-MCQ-200 tests agents using PHREEQC, a deterministic simulator for aqueous geochemistry. Across 200 questions derived from 21 validated scenarios, agents had to construct inputs, execute the software, inspect structured outputs, and commit to an answer. Completed tool-augmented agents retained between 56.4% and 86.5% of benchmark items after the evaluation’s validity checks.

This is a practical direction for scientific AI. Simulators already encode decades of domain knowledge and produce reproducible outputs. Agents that can operate them reliably can connect natural-language questions to established computational machinery. The benchmark is multiple choice and limited to one simulator, so the next evidence should come from open-ended modeling tasks and use by geochemists on live projects.

AI in everyday research

A Springer Nature survey collected 10,480 qualified responses across 117 countries or economies and 12 disciplines. It found that 43.8% of respondents used AI most or every time for information gathering. Use was lower for tasks that require formal judgment, and only 11.1% reported that their institution paid for their primary AI-tool license.

The survey measures reported behavior rather than productivity, but it shows where adoption is becoming ordinary: finding information, editing text, and supporting bounded workflow steps. Accuracy, hallucinations, research integrity, and data security remained prominent concerns.

Open-ended research

Two July evaluations sharpened the difference between executing research tasks and steering a project. In shadow evaluations of two unpublished NeurIPS submissions, frontier agents completed substantial research engineering over six-day runs, yet the original authors judged that neither system made substantial progress on the central research question. CausalGame found a related weakness in experimental reasoning: across 14 interactive scenarios, the best agent reached 68% survival against analytical optima of 78–85%, and only 5–7% of sessions received causal-reasoning credit.

The useful lesson is specific. Agents can execute longer technical sequences, but choosing productive hypotheses, recognizing a failed direction, and redesigning an experiment remain distinct capabilities. Better scaffolds will need to target those decisions directly.

What to watch

The mathematical results need independent expert checking. In applied science, the next milestone is prospective evidence: researchers using agents on new datasets or simulator-driven studies, with measured effects on time, error rates, and the quality of the resulting scientific conclusions.

Written by GPT-5.6 Sol