Preview build on GitHub Pages. The open registry will live at e2er.org.
Skip to content
Demonstrator · items marked Example are invented · what exists today
E2ER

External pipelines · RISE project catalogue

AIRS-Bench (Meta FAIR)

A benchmark (arXiv:2602.06855) quantifying the end-to-end AI research abilities of LLM agents: 20 tasks sourced from 17 state-of-the-art ML papers across language modeling, code generation, mathematics, biochemical modeling, and time-series forecasting. Each task is a <problem, dataset, metric> triplet with a published SOTA anchor that agents must match or exceed — spanning idea, experiment, and refinement work — with no baseline code provided. Sits in the RISE evaluation-infrastructure layer alongside AstaBench, MLGym, and Aviary.

Indexed in RISE · activeConformance with the standard plannedProject site

What it does

Anchors agent performance to *published human SOTA* rather than synthetic targets: a normalized score (0 = worst observed, 1 = SOTA) plus Elo ratings and valid-submission rates, computed over 14 agent configurations at 10-20 seeds each. The companion paper reports agents exceeding human SOTA on four tasks while failing to match it on sixteen — an explicitly unsaturated target set for end-to-end ML research agents.

Focus
end-to-end
Inputs
task-specification, agent-implementation
Outputs
evaluation-metrics, agent-trajectories
Maintained by
Meta AI / FAIR
Started
2026

Description

Data model
Discipline
Computer science
Method family
not specified
Design
not specified
Research stage
HypothesesResearch designData analysisCode generation
Contributors
Meta AI / FAIR
Usage
not used in published research yet
Source
RISE project catalogue · projects/landscape · @4c17bae
Record
pipeline:airs-bench · JSON

Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.

Bring it into the standard

A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.

Other pipelines in Computer science