Evaluations · real record
A run with its checks switched to record only
On 5 August 2026 E2ER ran one research question on the empirical template with governance set to off. The checks ran and wrote their findings to files without stopping the run. The draft was rejected at the end. The figures below come from those files.
Numbers in the tables
The draft's tables held 110 values. 43 matched the result file they were taken from. 53 were paired with a value in a result file and differed from it; 20 of those differences were critical, 33 major. 14 values had no result file behind them at all.
Source: results/number_verification.json of the run. Recomputed with E2ER at 3b91f0e, which gives the same 110, 43, 53 and 14.
References
The draft cited 18 references. 7 were confirmed by title in Crossref or Semantic Scholar. The other 11 could not be confirmed automatically. They include Kyle (1985), three SEC orders and a court case, so most are real sources. The check failed on them because the bibliography entries still carried LaTeX formatting (\newblock) and the titles did not match closely enough. The citation check passed the run anyway, because it was not in strict mode.
Source: reviews/citation_integrity.json of the run.
Numbers in the text
The first check of numbers in the prose flagged 284 of 423 as not matching any result. Most flags were false: years, single-digit numbers and formatting values were counted as results. After the check was fixed, a re-run on the same draft flagged 2. Both are real errors: the draft gives the VIX as averaging 19.2 before the ETF approval and 18.1 after it, while the only VIX figure in the results is an overall mean of 16.82, and the results hold no values for the two periods.
Source: E2ER commit 35d5d7f (10 August 2026), which fixed the check and records the re-run.
Scope
One run of one question with one model. A comparison across governance settings is planned; 8 of the 9 pilot runs for it could not be measured.