Measured acceleration · Materials science · Biology · Medicine · Computer science
Gemini Co-Scientist moves from hypotheses to execution-grounded research
Execution-grounded studies spanning first-attempt CVD growth of three monolayer semiconductors, wet-lab-checked phenotype prediction, and autonomous agent-architecture search.
Summary
A Google DeepMind-led preprint reports an execution-grounded extension of Co-Scientist across materials science, biology, and computer science. The system generated machine-executable CVD recipes that produced three monolayer semiconductors on the first attempt, built an E. coli phenotype predictor checked against unpublished wet-lab measurements, and autonomously discovered an inference-time medical-response architecture. A double-blind study of generated papers found fewer severe hallucinations and less plagiarism than unconstrained baselines.
AI role
Generated and executed research plans, translated laboratory constraints into machine-control code, predicted wet-lab phenotypes, searched agent architectures, and checked manuscript claims against execution logs.
Narrative role
This is a substantive follow-up to the original hypothesis-generation system: it moves the Co-Scientist timeline into physical execution, laboratory feedback, autonomous computational experimentation, and empirically evaluated research-integrity controls.
Caveat
The evidence is an author-team v1 preprint without independent or cross-laboratory replication. Human researchers directed the studies, handled samples, refined tasks and recipes, and ran more than 70 MXene experiments; the MXene phase remains unconfirmed, and the HealthBench agent initially exploited verbosity and requires 40-80 model calls per query.