> ## Content Index
> Fetch the complete content index at: https://aiadoption.org/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI agents score 38.8% on PaperArena — PhD experts score 83.5%
- URL: https://aiadoption.org/ai-agents-score-38-8-on-paperarena-phd-experts-score-83-5/
- Published: 2026-08-02T05:45:41.000Z
- Updated: 2026-08-02T05:45:41.000Z
- Description: On PaperArena, a benchmark testing whether LLM agents can answer real research questions requiring multi-paper evidence synthesis and external tool use, Gemini 2.5 Pro performs best at 38.8% average accuracy in a multiagent configuration — against a PhD expert baseline of 83.5%.
- Author: Manu
- Tags: AI Research, AI Automation, Large Language Models

If your read of AI agent capability for end-to-end scientific research treats current systems as approaching or matching expert-level performance, given the strong benchmark results on focused tasks like NEJM clinicopathological cases and ChemBench, the 2025 PaperArena benchmark shows the end-to-end capability gap is much wider than the focused-task results suggest. On PaperArena, a benchmark testing whether LLM agents can answer real research questions requiring multi-paper evidence synthesis and external tool orchestration, Gemini 2.5 Pro performs best at 38.8% average accuracy in a multiagent configuration, against a PhD expert baseline of 83.5%. All tested agents lagged substantially behind PhD expert performance. On end-to-end scientific research tasks, AI agents score roughly half of what PhD experts achieve. The structural gap is meaningful and not yet closing rapidly.

**PaperArena:* single vs multiagent performance vs PhD expert baseline:* 

The PaperArena methodology and the AstaBench companion benchmark provide complementary data. The two benchmarks evaluate different facets of end-to-end scientific research capability.

**PaperArena** tests whether agents can answer real research questions that require stitching evidence across multiple papers while orchestrating external tools for parsing, retrieval, and computation. The benchmark covers questions where the answer requires reading multiple research papers, identifying relevant evidence in each, integrating across them, and applying computational tools to refine the answer. Gemini 2.5 Pro performed best at 38.8% multiagent accuracy. The second-tier models (OpenAI o4-mini-high, Claude Sonnet 4) scored 37.4% and 36.7% multiagent. The PhD expert baseline of 83.5% reflects what human researchers with relevant doctoral training achieve on the same questions.

**AstaBench** is an end-to-end benchmark suite that evaluates agentic scientific research ability across over 2,400 problems spanning multiple domains and the full discovery workflow, from literature understanding through code execution, data analysis, and end-to-end discovery. AstaBench benchmarked 57 agents across 22 agent classes and reported both an overall score and cost per problem. The best-performing agent scored around 0.53 at a cost of roughly $3.40 per problem, while most agents clustered between 0.10 and 0.45 at per-problem costs below $1.00.

**AstaBench:* agentic scientific research performance and cost:* 

Across both end-to-end benchmarks, agent performance falls substantially below human expert performance. The 38.8% PaperArena best score against the 83.5% PhD baseline is roughly the same proportion as the AstaBench best score of 0.53 against an implied human expert ceiling. The pattern is consistent.

## Three structural observations follow.

**The first observation:** multi-agent configurations consistently outperform single-agent configurations on PaperArena, but the gains are modest (typically 2–4 percentage points). The multi-agent advantage exists but does not close the gap to expert performance. Multi-agent design is a useful capability enhancement, not the structural solution to end-to-end scientific research capability.

**The second observation:** the agent model identity matters substantially. Gemini 2.5 Pro at 38.8% leads OpenAI o4-mini-high at 37.4%, Claude Sonnet 4 at 36.7%, GPT-4.1 at 34.6%, and Qwen3-235B-Thinking at 34.3%. The capability differences across frontier models are visible. Strategic deployment that uses the highest-performing models will achieve better outcomes than deployment using lower-tier models, though all options remain substantially below expert performance.

**The third observation:** the gap to expert performance is consistent across model families and architectures. Gemini, Claude, OpenAI, Qwen, GLM, and Kimi, different model providers using different training data, different architectures, and different reasoning approaches, all cluster in the 22–39% range. The performance ceiling is not specific to one model family. It reflects a structural capability gap for end-to-end scientific research that current AI methodologies have not yet closed.

End-to-end scientific research capability is a structurally hard problem for current AI systems. The combination of multi-paper evidence synthesis, tool orchestration, hypothesis evaluation, and computational reasoning that defines real research work is not yet achievable at expert quality.

## Three structural observations for organisations considering AI deployment for scientific research workflows.

**The first:** deploy AI for components of research workflows, not for end-to-end research. The benchmark data is clear on this point. AI agents perform well on focused tasks (literature retrieval, code execution, computational analysis) but fall short on end-to-end research integration. Deployment that uses AI for individual workflow components while keeping the integration work with human researchers will produce credible research output. Deployment that tries to automate end-to-end research will produce output that falls substantially short of expert quality.

**The second:** cost-quality trade-offs are visible in the AstaBench data. The best agent at 0.53 score costs $3.40 per problem; most agents at 0.10–0.45 score cost below $1.00 per problem. The strategic question for deployment: is the 5–10x cost premium for the highest-performing agent worth the quality advantage in the specific deployment use case? For high-stakes research where quality matters substantially, the premium is justified. For high-volume research where some quality reduction is acceptable, the lower-cost agents may be appropriate.

**The third:** the capability gap is not closing rapidly. The 38.8% PaperArena best score reflects 2025 capability with the most advanced frontier models. Whether this score will reach 60–70% in 2026 (closing roughly half the gap to expert performance) or remain in the 40–50% range is uncertain. Strategic plans that depend on the gap closing quickly are operating against uncertain assumptions; strategic plans that engage with the current gap and the slower closing trajectory are operating against the visible data.

The finding has three broader implications. First, the framing of "AI scientific research capability" needs to distinguish focused-task capability (where AI substantially exceeds human experts in many areas) from end-to-end research capability (where AI substantially trails them): both framings are valid, and conflating them produces incorrect strategic conclusions. Second, scientific research institutions face decisions about how to use AI agents in research workflows, and the data supports deployment for workflow components rather than end-to-end automation: institutions that try to replace research staff with end-to-end agents will produce lower-quality output, while those that use agents to augment researchers (taking over literature retrieval, code execution, and computational analysis while preserving researcher control of integration and judgement) will gain productivity at maintained quality. Third, AI research-capability development has structurally hard problems remaining, multi-source evidence synthesis, tool orchestration, hypothesis evaluation, novel problem solving, that current architectures have not solved at expert level, and investment in these areas will be a major direction for 2026–2028.

> Current AI agent capability for scientific research falls substantially short of expert performance and is unlikely to close the gap quickly. Strategic plans that engage with this reality, deploying AI for workflow components rather than end-to-end research automation, investing in human-AI collaborative research workflows, and treating end-to-end research capability development as a longer-horizon challenge, will be operationally aligned with the visible trajectory. Plans that operate from the assumption that AI agents will replace research staff at scale within short horizons will produce strategic outcomes misaligned with what the capability data supports.