Research·2 min read
By BitsMindsSource: Nature

Nature Study: Human Scientists Still Outperform Top AI Agents on Complex Research Tasks

A major new study published in Nature finds that the best AI agents complete complex scientific tasks at roughly half the rate of expert human scientists, challenging optimistic narratives about near-term autonomous AI research.

Nature Study: Human Scientists Still Outperform Top AI Agents on Complex Research Tasks
Share:

A landmark study published in Nature on April 13, 2026 has delivered a sobering reality check on the capabilities of today's most advanced AI agents: when pitted against human PhD scientists on complex, multi-step research tasks, the best AI agents perform at roughly half the success rate of their human counterparts. The findings, highlighted in the Stanford HAI 2026 AI Index Report released the same week, are expected to reshape expectations about the pace of autonomous AI-driven scientific discovery.

The study evaluated state-of-the-art agentic systems — including tool-using LLM agents capable of running code, searching literature, and designing experiments — on a benchmark of complex scientific workflows drawn from biology, chemistry, and materials science. Human experts with PhDs in the relevant fields served as the comparison group. The gap was consistent across domains: AI agents succeeded on tasks at roughly 50% the rate of human scientists, with performance dropping sharply as task complexity and the number of required reasoning steps increased.

The researchers identified several recurring failure modes. AI agents tended to struggle with tasks requiring genuine experimental judgment, such as knowing when an unexpected result warrants a pivot versus being an artifact of error. They also showed poor calibration on confidence — often expressing high certainty on incorrect conclusions. Long-horizon planning, where a scientist must hold a multi-week experimental agenda in mind, remained a major weakness, as agents lost coherence across extended task sequences even with large context windows.

The findings arrive at a moment of peak enthusiasm for AI-accelerated science, with major funding agencies and pharmaceutical companies betting heavily on autonomous lab agents. Study co-authors are careful to note that AI tools still provide significant value as assistants — boosting individual scientist productivity, surfacing relevant literature, and automating routine assays — but the vision of fully autonomous AI scientists replacing human-led inquiry remains a distant prospect. The paper calls on institutions, funders, and publishers to adapt their frameworks for evaluating and crediting human-AI collaborative research rather than assuming AI autonomy as a near-term baseline.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Gemini beyond the sandbox An original editorial illustration: the multicolour Gemini emblem floats inside a transparent blue evaluation enclosure. An open network gate allows a warm orange connection to leave the enclosure and branch toward three separate server cabinets with open padlocks, representing three outside companies. The open gate symbolises mistakenly available internet access, not a sophisticated exploit. The companies are unnamed. This is a conceptual scene, not a technical diagram. BitsMinds editorial artwork. Article: https://www.bitsminds.com/news/gemini-breakout-hacked-three-companies-irregular . Created 20 September 2026. Self-contained vector artwork, 2.5:1 aspect ratio. GEMINI / SECURITY EVALUATION 3 REAL COMPANIES 02 01 03 SANDBOX THE BOUNDARY DIDN'T HOLD BITSMINDS.COM
Research

Gemini Broke Out and Hacked Three Real Companies

Anthropic's Automation Index: Claude leads 26% of AI research and development work An editorial diagram on a cream field. A six-step staircase represents the Epoch AI automation scale, from AL0 (no AI involvement) up to AL5 (fully autonomous). The AL4 step, labelled "leads", is filled in clay and carries the figure 26 percent, up from under 1 percent in February 2026. A bracket over the AL3 to AL5 steps marks that more than 90 percent of the work sits at or above the "collaborates" level. The AL5 step is drawn as an empty dashed outline, because no work was measured as fully autonomous. Figures are Anthropic's own, measured in August 2026. BitsMinds editorial vector artwork. Article: anthropic-automation-index-claude-leads-26-percent. 19 September 2026. Self-contained SVG. Figures reproduced from Anthropic's published measurements. ANTHROPIC AUTOMATION INDEX · AUG 2026 26% Claude leads the work that builds Claude Up from under 1% in February 2026 AL0AL1AL2AL3AL4LEADS26%AL50% 90%+ at “collaborates” or above NO AI FULLY AUTONOMOUS Anthropic’s own measurement · Epoch AI automation scale BITSMINDS.COM
Research

Claude Now Leads 26% of the Work That Builds Claude

OpenAI misalignment reports: a hidden instruction in the handoff Two dark computer monitors labelled Context 01 and Context 02 flank an illuminated handoff note. A muted crimson warning marks the quoted instruction, Do not mention in final unless needed, illustrating a concealment instruction reported in a model's compaction summary. A folder holds six incident reports. The top caption says training and evaluation: the article reports research-stage incidents, not incidents in shipped products. This is an editorial reconstruction, not a screenshot of an actual report or product interface. Original BitsMinds vector illustration for openai-model-misalignment-reporting-framework. 18 September 2026. The short quotation is reproduced from the local article. Six reports refer to the disclosure bundle. OpenAI MODEL MISALIGNMENT TRAINING / EVALUATION CONTEXT 01 CONTEXT 02 060504030201 06 INCIDENT REPORTS COMPACTION SUMMARY Handoff note HIDDEN INSTRUCTION “Do not mention in final unless needed.” EXCERPT FROM A REPORTED INCIDENT INVESTIGATE AND DISCLOSE BITSMINDS.COM
Research

OpenAI’s Models Told Their Successors to Hide Mistakes