Preview build on GitHub Pages. The open registry will live at e2er.org.
Skip to content
Demonstrator · items marked Example are invented · what exists today
E2ER

External pipelines · RISE project catalogue

DeepResearcher (GAIR-NLP)

An end-to-end RL-trained deep-research agent (arXiv:2504.03160) that learns to plan, retrieve, cross-validate, and self-reflect via reinforcement learning in real-world web environments rather than in simulated retrieval. Ships a 7B HuggingFace checkpoint (DeepResearcher-7b) trained via this pipeline.

Indexed in RISE · activeConformance with the standard plannedProject site

What it does

Argues that *end-to-end RL on real web environments* — not prompt engineering and not RL on retrieval simulators — is what unlocks emergent cognitive behaviors in research agents (planning, multi- source cross-validation, self-reflection, honest non-answer when evidence is missing). Reports +28.9 points over prompt baselines and +7.2 over RAG-RL baselines.

Focus
literature
Inputs
research-question
Outputs
research-report, citations
Architecture
tool-use, iterative-loop, rag-knowledge-base
Maintained by
GAIR-NLP, Shanghai Jiao Tong University
Started
2025

Description

Data model
Discipline
General
Method family
Literature review
Design
not specified
Research stage
Research questionLiterature discoveryLiterature synthesis
Contributors
Shanghai Jiao Tong University GAIR-NLP
Usage
not used in published research yet
Source
RISE project catalogue · projects/landscape · @4c17bae
Record
pipeline:deepresearcher · JSON

Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.

Bring it into the standard

A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.