Anthropic and OpenAI have promoted their models as able to speed up, and eventually run, AI research on their own. A new experiment from Princeton and the UK AI Security Institute suggests the research judgment behind that claim is not there yet.
The team built a method they call Shadow Evaluation. An agent receives the central research question from an unpublished paper, then the paper's original authors, who spent months on it, review the result as conference reviewers would. Because the papers were not yet public, the agent could not lean on its training data.

They partnered with the authors of two NeurIPS 2026 submissions: one on steering language-model personality traits through weights, and another on a method called TabPFN that flags when a deployed model hits data unlike its training set.
Each run used Claude Opus 4.8 with Extra-High Reasoning. The agent got six days, three thousand dollars in API credits, a GPU budget, and full access to a virtual machine and the open web, running inside a scaffold that launched subagents and long-running GPU jobs, with a heartbeat that collected results when jobs finished. The scaffold was OpenClaw, an open-source agent framework built by Peter Steinberger, who joined OpenAI this year.
The verdict: the engineering works, the research judgment does not. The agents could handle research engineering but failed at the parts that matter, forming the right question, judging relevance, and knowing when an approach is wrong, and the failure modes repeated across runs.
The results clash with lab claims that autonomous AI research is within reach. The authors also argue that submitting AI-generated papers to peer review, a common evaluation, is a poor yardstick because review quality is uneven.
Sources
The Decoder: https://the-decoder.com/study-contradicts-anthropic-and-openai-claims-that-autonomous-ai-research-is-within-reach/



