External pipelines · RISE project catalogue
Asta AutoDiscovery
Ai2's autonomous data-driven discovery agent (formerly AutoDS; relaunched inside AstaLabs on 2026-02-12): pointed at a structured dataset, it generates natural-language hypotheses, proposes experiment plans, writes and executes Python analyses — up to 500 experiments in a session — and ranks the resulting findings by Bayesian surprise, the shift from the LLM's prior to posterior belief in each hypothesis. Sits at the hypothesis-generation → data-analysis → code-generation slice of the pipeline; no literature layer and no paper drafting.
Where it sits
left: what it builds on · right: what builds on it · pale: exampleContributed by Allen Institute for AI
How studies reach it
No published study reaches it yet.
Disciplines it reaches
No study reaches it yet.
Solid: published studies. Light: examples.
Computed from the records on this site: what each study, template and specialist names as used, which study extends which, and who contributed what. 0 studies in total.
What it does
The first production discovery agent to use Bayesian surprise as the objective: an MCTS search with progressive widening treats surprisal as reward, so the system hunts belief-shifting findings rather than confirmations. Early-access users have generated 46K+ hypotheses across oncology, neuroscience, climate science, and the social sciences, and several independently verified social-science findings were published in a peer-reviewed paper (arXiv:2511.12529).
- Focus
- ideation
- Inputs
- structured-dataset
- Outputs
- ranked-hypotheses, experiment-code, statistical-results
- Architecture
- tool-use, iterative-loop
- Maintained by
- Allen Institute for AI (Ai2)
- Started
- 2025
Description
Data model- Discipline
- General
- Method family
- not specified
- Design
- not specified
- Research stage
- HypothesesData analysisCode generation
- Contributors
- Allen Institute for AI
- Usage
- not used in published research yet
- Source
- RISE project catalogue · projects/landscape · @4c17bae
- Record
- pipeline:asta-autodiscovery · JSON
Solid tags are declared by the source or mapped from its terms; dashed tags are inferred by a published rule. Hover a tag for its provenance.
Bring it into the standard
A pipeline built outside E2ER can meet the standard by describing its steps as a template, attaching the floor of checks and publishing evaluation records. Its authors keep ownership and credit.