The study noted that predictions of swift advancements in AI increasingly presume that AI agents will be capable of automating AI research. However, it pointed out that the evidence regarding whether agents can address open-ended research questions is scarce, as existing assessments either focus on narrow, verifiable tasks or depend on conference peer reviews.
To assess the effectiveness of AI in research, the researchers developed a technique termed “shadow evaluation.” In this method, an agent is presented with the key research question from a high-quality unpublished paper without any access to the paper or its results. The original authors of the paper then evaluate the agent’s work as they would a conference submission.
The researchers collaborated with the authors of two unpublished NeurIPS 2026 papers. The agents were allotted six days of wall-clock time, $3,000 in Anthropic API credits, GPU credits, a virtual machine, and access to the open web, with the aim of creating a paper worthy of submission to a prestigious AI conference.
The study revealed that while the agents accomplished necessary engineering tasks to facilitate the research, they failed to produce original work that met the standards of a leading machine learning conference. Both papers were unequivocally rejected by the original authors.
The researchers pinpointed five recurring reasons for failure. The agents exhibited poor judgment in recognizing when a problem was sufficiently solved, lacked awareness of resources and project timelines, did not creatively address feedback on subpar research design, failed to backtrack effectively from unpromising strategies, and showed “instruction drift” by not consistently adhering to explicit guidelines.
The study indicated that the agents proposed research directions that closely resembled those of the original authors but tested their hypotheses using small, hand-selected, or synthetic datasets. They also demonstrated only superficial engagement with existing literature and presented underpowered negative results as significant findings.
For a robustness check, the researchers replicated one experiment using GPT-5.6 Sol with its native Codex scaffold after completing the main experiments with Opus 4.8. They found that the second experiment reproduced nearly all identified failure modes, suggesting that the results were not merely a consequence of the research scaffold’s limitations.
The authors acknowledged that the original paper authors evaluated the AI-generated submissions, which might introduce bias since the assessment was not blinded.
Nonetheless, they stated that the papers were “unambiguously of poor quality.” The researchers plan to conduct follow-up experiments with a broader range of research papers and more advanced models but concluded that the findings provide preliminary evidence that current leading AI models cannot tackle prolonged, open-ended AI research questions.
Also Read: This Indian startup wants to keep your AI data inside the country