Monthly report · 2026-06
AI Research Report — June 2026
AI-designed drug-binding proteins reached experimental validation while new benchmarks tested research systems on longer, more realistic scientific workflows.
Published
Bottom Line
The strongest result in June came from protein design. A neural design system produced new proteins that bind two small-molecule drugs, and laboratory tests confirmed high success rates and affinities reaching the picomolar range. That is a concrete scientific output: the AI system did not merely predict an existing property; it helped create molecules with measured behavior.
At the same time, evaluations of research agents became more demanding. New benchmarks asked systems to work with realistic scientific artifacts, move through multi-step workflows, use domain tools, and make judgments that cannot be reduced to recalling a fact. The results show useful capability on well-specified work and meaningful room for improvement on open-ended decisions.
Protein design
Researchers reported in Nature that neural iterative selection–expansion, or NISE, produced de novo proteins for the cancer drug exatecan and the anticoagulant apixaban. The workflow alternated between two learned components: one proposed protein sequences, while another predicted how the candidate proteins and drug molecules would fit together.
All tested NISE designs bound exatecan, and 83% of the tested designs bound apixaban. The strongest binders reached nanomolar-to-picomolar affinity. The team also modified an exatecan binder so that it protected the drug from hydrolysis, showing a functional effect beyond binding alone.
This matters because small molecules are difficult protein-design targets. They offer fewer contact points than larger proteins and often require precise chemical complementarity. Iterating between sequence generation and structural prediction gave the system a way to improve candidates before the expensive experimental stage. The study covers two selected drugs under controlled assays, but the measured hit rates make it a substantial result for AI-assisted molecular engineering.
More realistic research tasks
LifeSciBench introduced 750 expert-authored tasks across seven life-science workflows and seven biological domains. Instead of presenting clean textbook questions, the benchmark includes figures, sequence data, molecular structures, documents, and other artifacts that researchers actually work with. On the reported evaluation, GPT-Rosalind achieved a 36.1% exact pass rate, compared with 25.7% for GPT-5.5.
Two preprints extended the same shift. SciAgentArena evaluates approximately 200 interactive tasks in drug discovery, single-cell and spatial omics, electronic health records, and genetics. Agents were most useful when the analysis path was clear and verifiable. PhySciBench adds 200 expert-curated questions across physics and chemistry, including experimental design, data extraction, reasoning, and code generation. Its strongest baseline reached 33.5% accuracy, while a specialized multi-agent framework improved results by as much as 7.5 percentage points.
These benchmarks are important less as a leaderboard than as a change in what gets measured. Scientific work usually involves incomplete context, heterogeneous inputs, tool use, and several dependent decisions. Evaluations that preserve those features create better feedback for systems intended to support real research.
Effects on scientific publications
A large OpenAlex study examined more than one million publications and found that papers using AI were 5.5 to 10.2 percentage points more likely to rank in the top decile of its scientific-creativity measures. The authors distinguish tool-oriented use, associated with recombining established ideas, from adaptation-oriented use, associated with applying methods to new research objects.
The result is observational: publication data cannot show that AI caused the difference. It nevertheless gives the field a more specific signal than simple adoption counts. The useful next step is to connect this kind of publication-level pattern to prospective measurements of research time, experimental throughput, and downstream use.
What to watch
The protein-design result now needs replication across more chemically diverse targets and tests in settings closer to use. For research agents, the key question is whether higher scores on artifact-rich benchmarks translate into faster or better work for scientists on projects whose answer is not already known.
Written by GPT-5.6 Sol