External pipelines · RISE project catalogue
AARRI-Bench
"Act As a Real Research Intern" (arXiv:2606.07462) — 82 containerized scenarios in standardized Harbor task format, each with an assertion-based verifier, probing whether LLM agents show the professionalism of human researchers in *granular* research situations (citation integrity, ablation-completeness audits, dead-end recognition, contradictory-advisor merging) rather than end-to-end execution. Inaugural stage of a planned three-stage AARR series (AARRI intern -> AARRA agent -> AARRS scientist); sits in the RISE evaluation-infrastructure layer alongside AstaBench, MLGym, and AIRS-Bench.
Where it sits
left: what it builds on · right: what builds on it · pale: exampleContributed by AARR-bench team
How studies reach it
No published study reaches it yet.
Disciplines it reaches
No study reaches it yet.
Solid: published studies. Light: examples.
Computed from the records on this site: what each study, template and specialist names as used, which study extends which, and who contributed what. 0 studies in total.
What it does
Isolates the micro-level judgment failures that end-to-end benchmarks average away — context sensitivity, independent judgment, knowing when to quit, collaboration under conflicting guidance — as individually verifiable scenarios. The 11-author paper reports the best configuration (Mini-SWE-Agent with Claude Opus 4.7) at a 68.3% success rate, quantifying a specific gap between frontier agent harnesses and human research interns.
- Focus
- end-to-end
- Inputs
- task-specification, agent-implementation
- Outputs
- pass-fail-verdicts, evaluation-metrics
- Maintained by
- AARR-bench team
- Started
- 2026
Description
Data model- Discipline
- Computer science
- Method family
- not specified
- Design
- not specified
- Research stage
- HypothesesLiterature synthesisResearch designData analysisCode generationRevision and editing
- Contributors
- AARR-bench team
- Usage
- not used in published research yet
- Source
- RISE project catalogue · projects/landscape · @4c17bae
- Record
- pipeline:aarri-bench · JSON
Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.
Bring it into the standard
A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.