External pipelines · RISE project catalogue
AIRS-Bench (Meta FAIR)
A benchmark (arXiv:2602.06855) quantifying the end-to-end AI research abilities of LLM agents: 20 tasks sourced from 17 state-of-the-art ML papers across language modeling, code generation, mathematics, biochemical modeling, and time-series forecasting. Each task is a <problem, dataset, metric> triplet with a published SOTA anchor that agents must match or exceed — spanning idea, experiment, and refinement work — with no baseline code provided. Sits in the RISE evaluation-infrastructure layer alongside AstaBench, MLGym, and Aviary.
Where it sits
left: what it builds on · right: what builds on it · pale: exampleContributed by Meta AI / FAIR
How studies reach it
No published study reaches it yet.
Disciplines it reaches
No study reaches it yet.
Solid: published studies. Light: examples.
Computed from the records on this site: what each study, template and specialist names as used, which study extends which, and who contributed what. 0 studies in total.
What it does
Anchors agent performance to *published human SOTA* rather than synthetic targets: a normalized score (0 = worst observed, 1 = SOTA) plus Elo ratings and valid-submission rates, computed over 14 agent configurations at 10-20 seeds each. The companion paper reports agents exceeding human SOTA on four tasks while failing to match it on sixteen — an explicitly unsaturated target set for end-to-end ML research agents.
- Focus
- end-to-end
- Inputs
- task-specification, agent-implementation
- Outputs
- evaluation-metrics, agent-trajectories
- Maintained by
- Meta AI / FAIR
- Started
- 2026
Description
Data model- Discipline
- Computer science
- Method family
- not specified
- Design
- not specified
- Research stage
- HypothesesResearch designData analysisCode generation
- Contributors
- Meta AI / FAIR
- Usage
- not used in published research yet
- Source
- RISE project catalogue · projects/landscape · @4c17bae
- Record
- pipeline:airs-bench · JSON
Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.
Bring it into the standard
A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.