Home / Benchmarks & Leaderboards

Benchmarks & Leaderboards

90 entries

Benchmarks 85

CORE-Bench

Leaderboard2024

Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Interactive Environments for Empowering LLM Agents in Machine Learning Engineering

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

13

SUPER

2024

Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

FEABench

NeurIPS Workshop2024

Evaluating Language Models on Multiphysics Reasoning Ability

REPRO-Bench

ACL Findings2025

Can Agentic AI Systems Assess the Reproducibility of Social Science Research?

Can Coding Agents Reproduce Findings in Computational Materials Science?

A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

ResearchQA

TACL2026

Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics

PeerRead

NAACL2018

A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications

FLAWS

2025

A Benchmark for Error Identification and Localization in Scientific Papers

Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents

MASSW

2024

A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows

A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds

15

Probing the limitations of multimodal language models for chemistry and materials research

BioLP-bench

bioRxiv2024

Measuring understanding of biological lab protocols by large language models

An Improved Benchmark for AI Systems Performing Biology Research

63

CritPt

2025

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

SciEval

AAAI2024

A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research

SciBench

ICML2024

Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

EAIRA

2025

Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

Leaderboards 5

Public third-party evaluations of frontier model autonomy and long-horizon task capability, with released task suites.

Nothing matches those filters.