Google DeepMind researchers showed on August 28 that Co-Scientist, its multi-agent research system, has evolved beyond hypothesis generation into a closed-loop partner that plans experiments, operates laboratory equipment, and writes scientific papers. The work, led by Samuel Schmidgall and published in Nature, was highlighted by The Decoder and extends the multi-agent science systems the lab validated in 2025.
What Changed
Co-Scientist, built on Gemini models, now runs a three-stage pipeline covering ideation, experimentation, and manuscript writing. An evolutionary agent system generates and refines hypotheses using Bayesian-rated pairwise tournaments with Upper Confidence Bound selection and automated literature review. A second evolutionary framework writes, executes, and iterates on experimental code. A third synthesizes results into structured papers, with a new verification module that cross-checks numeric claims against actual execution logs to suppress fabricated output.
Lab Results
In materials science, the system was paired with a semi-automated high-temperature furnace and chemical vapor deposition reactor. It generated complete growth recipes for the two-dimensional material MXene and, using Gemini 3 Deep Think for direct equipment control, synthesized three semiconductor thin films (monolayer MoS2, MoSe2, and WS2) on the first attempt, cutting recipe development from days to minutes. In biology, Co-Scientist autonomously built an image-analysis pipeline that predicted emergent swarming patterns of engineered E. coli colonies, matching unpublished wet-lab results on three of four morphological features. In computer science, it designed Agent_H, a medical AI architecture that outperformed six frontier models, including GPT-5 and Claude Opus 5, on HealthBench benchmarks.
Reliability Gains
In a double-blind evaluation with 30 domain experts reviewing 150 autonomously generated papers, Co-Scientist fabricated key results in 4 percent of cases with reliability modules active, versus 46 percent without them and 90 percent for a comparison system. The integrated safety architecture rejected 98.7 percent of potentially harmful research directions. Researchers acknowledged remaining limits, including selective reporting, recipes that may not transfer between labs, and benchmark gains that do not always align with expert clinical judgment.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.