External pipelines · RISE project catalogue
AstaBench (AI2)
An evaluation framework from AI2 for measuring scientific-research abilities of AI agents. 2,400+ examples across 11 benchmarks covering literature search, code execution, data analysis, and end-to-end discovery. Built on the InspectAI framework. Sits in the *RISE evaluation infrastructure* layer alongside Aviary and MLGym.
Where it sits
left: what it builds on · right: what builds on it · pale: exampleContributed by Allen Institute for AI
How studies reach it
No published study reaches it yet.
Disciplines it reaches
No study reaches it yet.
Solid: published studies. Light: examples.
Computed from the records on this site: what each study, template and specialist names as used, which study extends which, and who contributed what. 0 studies in total.
What it does
The most-scoped benchmark suite for scholarly-research agent abilities specifically — not generic agent benchmarks, not domain-specific science tasks (cf. BixBench), but a curated spectrum of *research skills* with standardized tools and execution environments for fair efficiency-comparable runs.
- Focus
- end-to-end
- Inputs
- task-specification, agent-implementation
- Outputs
- agent-trajectories, leaderboard-submissions, efficiency-metrics
- Architecture
- tool-use, artifact-versioning
- Maintained by
- Allen Institute for AI
- Started
- 2025
Description
Data model- Discipline
- General
- Method family
- not specified
- Design
- not specified
- Research stage
- Literature discoveryData analysisCode generation
- Contributors
- Allen Institute for AI
- Usage
- not used in published research yet
- Source
- RISE project catalogue · projects/landscape · @4c17bae
- Record
- pipeline:asta-bench · JSON
Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.
Bring it into the standard
A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.