The hard part may be moving from an idea to a fact
AI agents are getting good at producing scientific hypotheses. Testing those hypotheses is not getting fast at the same rate.
That is the central argument in a new Google DeepMind policy essay called ‘Conjecture Machines’. The authors say science may be approaching a lopsided moment: ideas and candidate solutions become cheap and plentiful, while the work that turns one of them into reliable knowledge stays physical, expensive and slow.
The distinction matters. A plausible explanation is not a discovery. A proposed drug is not a treatment. A new material described by a model still has to be made, measured and compared with what already exists.
In software and mathematics, some checks can run inside a computer. A program either passes a test or it does not. A formal proof can be verified. Biology, chemistry and much of the physical world are less forgiving. Cells need time to grow. Reactions take time. Equipment and trained people are finite.
DeepMind calls the resulting gap a validation bottleneck. The phrase is useful because it puts attention on the part of AI-for-science that a model launch can easily hide.
There is early evidence, but it is still early
The essay is not built only on a future scenario. It points to results from Co-Scientist, Google's multi-agent system for generating and refining research hypotheses.
A peer-reviewed Nature paper published in May describes three biomedical applications that reached wet-lab checks. In one, the system proposed drug candidates for acute myeloid leukaemia. In another, it suggested treatment targets for liver fibrosis. It also reconstructed an unpublished explanation for how a mechanism of antibiotic resistance can spread across bacterial species.
Those results show that an agent can contribute ideas that survive an initial encounter with laboratory evidence. They do not show that it can run science alone.
The paper repeatedly describes scientists in the loop. Experts set the goal, added constraints, selected which proposals deserved scarce lab time and interpreted the result. The biomedical validations were deliberately limited, and the authors say larger experimental testing is difficult precisely because it takes so much time and money.
Some performance measures were also narrow. Expert preference tests covered 11 research questions. Other comparisons used the system's own tournament-style rating. The Nature paper says the small evaluations need further study.
So the confirmed result is modest and still important: several agent-generated hypotheses passed defined in-vitro checks under expert supervision. A general claim that AI now discovers better science than people would go well beyond the evidence.
A faster idea machine changes what becomes scarce
Scientists already live with too much to read and too little time. Agents can search papers, query databases, write analysis code and propose links across fields. They can also run many lines of reasoning in parallel.
That lowers the cost of making a candidate idea. It may help a small team explore directions that would otherwise stay untouched. It can also create a pile of polished, cited and plausible work that someone still has to inspect.
One hallucinated detail can invalidate a long analysis. A proposal can be logically neat and biologically wrong. Even a sound result may be too incremental to deserve the next available experiment.
This is where human judgement does not disappear. It moves. Researchers may spend less time producing a first list of possibilities and more time deciding which question matters, which evidence is missing and which result is strong enough to trust.
The same pressure reaches peer review and grant funding. If agents help people draft more papers and applications, writing quality becomes a weaker signal. Reviewers receive more material without more hours in the week. An agent that helps write a submission may eventually sit opposite another agent helping to audit it.
That could improve error detection. It could also produce an arms race in volume. The outcome depends on the rules institutions build around the tools.
DeepMind's four-part answer
The essay proposes four priorities for governments and science funders.
First, researchers need access to capable agents. DeepMind argues that this may require new funding models because token and compute costs do not fit neatly into a conventional lab budget.
Second, publicly funded data should be easier for agents to use, with documented interfaces, reliable metadata and stronger controls for sensitive material. A folder of scanned PDFs is technically public and still difficult to use well.
Third, experimental capacity needs investment. That includes ordinary shared facilities as well as automated labs, where robotics can run some tests at higher volume. DeepMind has built a wet lab inside the Francis Crick Institute and is funding independent researchers who can test agent-generated ideas. That makes the recommendation relevant, but it also gives the company a direct interest in the direction it describes.
Fourth, reviewers should be allowed to use agents carefully, especially for checking citations, methods and internal consistency. The authors also suggest better disclosure of how AI contributed to a result.
None of these steps makes verification free. That is the point. If hypothesis generation becomes abundant, the systems that choose and check hypotheses become core scientific infrastructure.
What remains open
‘Conjecture Machines’ is a policy essay written from inside Google DeepMind. Its interviews are mainly with DeepMind researchers, and many examples feature DeepMind systems. It is a useful argument, not an independent forecast of how science will change.
We still do not know how often agent-generated ideas fail before a promising example reaches publication, whether the gains transfer across disciplines or how much the full workflow costs. It is also unclear who will get access to the best agents and the facilities needed to test their output.
There is a risk of concentration. A lab with models, private datasets, compute and automated experiments could move much faster than a university group with none of them. Opening data and shared facilities may narrow that gap. It may also be difficult and expensive enough that only a few institutions can do it well.
The near-term lesson is simpler. An AI-generated hypothesis should enter the same chain of evidence as any other hypothesis. Its fluency does not shorten the chain.
Agents may help science ask more questions. The harder task is deciding which questions deserve the world's limited capacity to find out.
Sources
- Google DeepMind — Conjecture MachinesPrimary company policy essay published in July 2026. Source for the validation-bottleneck argument and policy recommendations; it is identified as a DeepMind perspective, not independent evidence.
- Nature — Accelerating scientific discovery with Co-ScientistPeer-reviewed primary research published 19 May 2026. Source for the system design, expert-in-the-loop evaluations, wet-lab validations and stated limitations.
- Google DeepMind — Co-Scientist research overviewOfficial technical overview and application examples, used to cross-check the company description and availability of Co-Scientist.



