# Evidence and claim register

Version 1.0 · 20 September 2026

Evidence cutoff: 16 September 2026.

## How to interpret this register

This is a ledger of central claims and their permitted inferences, linked to the [references and review-depth notes](../index.html#references). It is not a scorecard of institutional trust. Review depth applies to the source material actually inspected, not every linked appendix or dataset.

**Review labels:** `computational audit` identifies a check performed on released artifacts; `selected review` identifies review of relevant primary-source sections; `abstract` indicates that only an abstract or bibliographic summary supports the statement. `Earlier review` identifies sources examined during an earlier phase of the research, without a subsequent full rereading. `Proposed analysis` identifies conceptual models or policy recommendations developed in the paper. These labels describe the work performed, not institutional independence or verification of the underlying events.

## Central claim ledger

| ID | Claim or analytical question | Evidence location | Review and inference boundary |
|---|---|---|---|
| C01 | Selected Navier–Stokes checks passed at a pinned revision | [Proof-checking results](../index.html#appendix-i) | Computational audit; formal artifact, not complete conventional review or AI provenance |
| C02 | Forced breakdown differs from unforced regularity | [Scope audit](../index.html#appendix-i) | Selected mathematical interpretation review |
| C03 | Other mathematical claims require theorem-level scope | [Scope of mathematical claims](../index.html#mathematics-scope) | Theorem-scope distinctions; selective coverage of specific mathematical claims |
| C04 | Reported research activity does not identify causal acceleration | Paper Section 2.3 | Selected review of developer methods; causal critique is proposed analysis |
| C05 | Parallel effort can encounter coordination and integration limits | Paper Section 2.4 | Primary conceptual model plus explicitly illustrative bottleneck model |
| C06 | Productivity-study participation and tasks can be endogenous | Paper Section 2.3 | Selected METR methodological review; no new effect estimate |
| C07 | Hugging Face and developer reports have distinct evidentiary scope | Paper Section 3.2 | Primary accounts, plus earlier audit; original service logs unavailable |
| C08 | Public METR visualizations contain derived units and unresolved count differences | [Derived-data results](../index.html#appendix-h) | Computational audit; discrepancies are unresolved, not evidence of fabrication |
| C09 | The Hugging Face replay is selected and interpolated | [Replay audit](../index.html#appendix-h) | Computational audit; not a raw event corpus |
| C10 | Released Mythos records agree in IDs but omit material across formats | [Cross-format audit](../index.html#appendix-h) | Computational audit; content completeness and authenticity remain separate |
| C11 | Retrospective scan totals are not a deployment-rate denominator | Paper Section 3.3 | Selected primary report; detector/corpus not independently reproduced |
| C12 | AISI event counts differ from run counts and have cross-run dependence | [AISI consistency results](../index.html#appendix-h) | Computational audit of tables and selected content, not all original trajectories |
| C13 | AISI's public quarantine times differ by 54 minutes | [Incident audit](../index.html#appendix-h) | Visually checked source pages in prior audit; unresolved interpretation |
| C14 | Narrow training can change broader behavior in studied models | Paper Section 4.2 | Selected methods and limitations; cross-model generality not established |
| C15 | Reward-hacking experiment supplies additional hack information | Paper Section 4.2 | Selected limitations; distinguishes generalization from spontaneous discovery |
| C16 | Summer-2026 failure-seeking simulations are not deployment samples | Paper Section 4.3 | Reviewed methods and consequence ablations; no complete transcript replication |
| C17 | AISI reports zero confirmed unprompted sabotage in its tested suite | Paper Section 4.3 | Reviewed methods/results; limited scenarios and evaluation-awareness caveats |
| C18 | Metagaming complicates behavioral interpretation | Paper Section 4.4 | Reviewed primary analysis; not proof that awareness always causes harm |
| C19 | Multi-agent monitoring may miss distributed behavior | Paper Section 4.5 | Published methods, limitations and availability reviewed in Appendix G; restricted code and trajectories unavailable |
| C20 | Control tests must cover artifacts and trajectories | Paper Section 5.2 | ResearchArena methods and selected pinned scoring/aggregation code reviewed; no model-experiment rerun |
| C21 | Detection and timely enforcement are distinct | Sections 5.3–5.5; Appendix B | Conditional-probability accounting and proposed operational analysis |
| C22 | Biological knowledge and physical execution studies have different endpoints | Paper Section 6.1 | Study-level design and statistical review; physical-trial RR, Fisher test and score CI reproduced from aggregate counts; participant data unavailable |
| C23 | Catastrophic pathways require intermediate causal claims | Paper Sections 6–8 | Conceptual analysis informed by cited primary literature; no fitted risk model |
| C24 | September-published researcher survey fielded in December 2024 | Paper Section 7.2 | Methods and Table 4 checked; no raw respondent reanalysis |
| C25 | Near-term forecasting results do not validate a long-run risk percentage | Paper Section 7.3 | Reviewed follow-up and limitations; primary tournament overview |
| C26 | Zero-detection bounds depend on strong assumptions | Appendix B.2 | Reproduced hypothetical calculation, not applied to incidents |
| C27 | Pacing must be evaluated against a counterfactual | Sections 8–10 | Proposed decision analysis; intervention effects unestimated |
| C28 | Pacing advocacy and corporate implementation are different | Sections 9 and 11 | Primary statements and archives; no compliance certification |
| C29 | US order Section 3 is voluntary and disclaims mandatory preclearance authority | Paper Section 11.2 | Selected official provisions reviewed; not a full US legal analysis |
| C30 | S. 5105 inspected text is an introduced proposal | Sections 11.2–11.3 | Introduced official record checked and Sections 2–4 reviewed; external legal interpretation pending |
| C31 | EU implementation depends on rule category and transitions | Paper Section 11.2 | July consolidated and authentic amending legislation reviewed; provision-level crosswalk in Section 20 |
| C32 | China's statement is not a pacing agreement | Sections 11–12 | Official response plus selected original Chinese and English framework provisions compared; implementation not verified |
| C33 | Hardware verification proposals differ in maturity | Paper Section 12.2 | Selected methods, mechanism scopes and limitations reviewed; engineering validation not performed |
| C34 | Delay, concentration, rights, and defensive access affect net policy value | Sections 9–13 | Proposed analysis; empirical magnitudes still need study |
| C35 | Energy figures are projections for data centers | Paper Section 12.5 | IEA executive-summary figure reviewed; no original energy model reproduction |
| C36 | Graduated controls are this paper's provisional recommendation | Paper Section 13 | Normative and decision analysis; conditions for reversal explicit |

## Empirical and policy claim ledger

| ID | Claim or analytical question | Evidence location | Review and inference boundary |
|---|---|---|---|
| C37 | ResearchArena monitoring is conditional on selected attacks and baselines | Appendix G.2 | Methods and selected pinned code inspected; not natural incidence or complete result replication |
| C38 | Distributed attacks exploit a fragmented monitor view | Appendix G.3 | Synthetic design and restricted availability reviewed; grouping advantage requires deployment validation |
| C39 | Digital biology's 4.16 statistic is an odds ratio | Appendix G.4 | Primary v2 design/results; adjusted odds do not equal a universal accuracy multiplier |
| C40 | Physical-trial primary aggregate is numerically consistent | Appendices G.5 and J; calculate-review.py | RR, Fisher tail and Koopman score CI independently reproduced; individual records unavailable |
| C41 | Biology SAP has a date inconsistency but visible signatures | Appendix G.5; cited statistical analysis plan | Visual inspection; chronology needs registry and unblinding records; no misconduct inference |
| C42 | METR 2025 assistance slowed selected familiar-repository tasks | Section 16.2 | Full design/main methods and selected appendices; limited developers, repeated tasks and dated models |
| C43 | Customer-support main result is staggered-rollout evidence | Section 16.2 | Pilot control identities unavailable in inspected study; no reconstructed randomized pilot contrast |
| C44 | Education-related task gains coexist with material attrition | Section 16.2 | Primary May revision, graders and completion reviewed; no full missing-data reanalysis |
| C45 | Revised Denmark result has a roughly 2% bound for its estimand | Section 16.2 | March 2026 primary revision; adoption endogeneity and common spillovers remain |
| C46 | Early-career employment result is descriptive | Section 16.2 | Updated author record; occupational exposure is not randomized AI adoption |
| C47 | Disavowed scientific-productivity preprint is excluded | Section 16.3 | Official MIT research-record notice; no independent misconduct adjudication |
| C48 | Large-scale scientific-feedback experiment measures revisions | Section 16.3 | Primary design and results; feedback-offer bundle, not validated discovery productivity |
| C49 | Survey estimates differ by wording and denominator | Section 17.2 | Original questions and selected tables checked; no pooled probability |
| C50 | Scenario uplift crossing is structurally sensitive | Section 17.5; reproduced calculation | Algebra reproduced, inputs illustrative; no full simulator validation |
| C51 | AI 2040 describes a recommended governance trajectory | Section 17.6 | Primary scenario; not a simple revised capability forecast or endorsed enforcement design |
| C52 | Monitoring and oversight results remain configuration-dependent | Section 18 | Selected primary methods and counterevidence; no general alignment guarantee |
| C53 | DARPA patch percentages require distinct denominators | Section 18.6; reproduced calculation | Official corrected counts; arithmetic checked, competition not deployment incidence |
| C54 | Persuasion personalization inference changed after correction | Section 19.1 | Original article plus September 2026 correction; nonsignificance is not equivalence |
| C55 | New persuasion studies measure actual behavioral outcomes | Section 19.2 | Primary methods and limitations; paid engagement, low-cost behavior, attrition and repeated participants |
| C56 | Present-harm sources have different evidentiary status | Section 19.4 | Official case/complaint records and primary technical studies; not a combined prevalence series |
| C57 | Openness is a spectrum of permissions and reversibility | Section 19.6 | Primary NTIA assessment plus proposed comparative analysis; net effects unestimated |
| C58 | Model welfare is unresolved and distinct from legitimate refusal | Section 19.7 | Conceptual primary literature; no diagnostic consciousness finding |
| C59 | Energy attribution must distinguish AI from all data centers | Section 19.8 | Primary IEA/LBNL scope; no new engineering forecast |
| C60 | California imposes thresholded process and reporting duties | Section 20.2 | Chaptered and current codified provisions reviewed; no compliance certification |
| C61 | New York's current RAISE duties begin January 2027 | Section 20.3 | Official current sections, deadlines and exceptions; not an older bill summary |
| C62 | EU July amendment changes high-risk dates distinctly from GPAI | Section 20.4 | Consolidated text and authentic amending act; external legal review pending |
| C63 | China's framework recommendations differ from binding measures | Section 20.5 | Selected original Chinese/English framework comparison and separate CAC measures |
| C64 | Developer frameworks differ in triggers, exceptions and authority | Section 21 | Selected consequential clauses and explicit change logs across six developers; implementation unverified |
| C65 | Displacement can change the sign of a modeled restriction effect | Section 22.1 | Original sensitivity arithmetic, not estimated substitution or risk |
| C66 | Revised Italy study does not support its superseded simple headline | Section 22.2 | Authors' revised primary methods/results; short national access restriction is a narrow analogy |
| C67 | Competing policy arguments have explicit reversal conditions | Section 22.4 | Analytical comparison of competing positions; no external endorsement or consensus claimed |
| C68 | Record access, prospective validation, and independent review limit the available assurance | Appendix K | Additional literature review cannot substitute for unavailable records or unperformed experiments |

## Search strategy

The review used purposive primary-source searches, citation following, and direct retrieval of known research and policy records. Searches conducted on 16 September 2026 covered control evaluations, evaluation awareness, biological uplift, productivity, forecasting, and governance. Sources included research papers, developer reports, government publications, and original legal texts. Secondary aggregators and social media served as discovery aids. The strategy was not a systematic database search and does not establish exhaustive coverage.

## Version and access limits

The paper’s reference list identifies cited sources and their review depth. Retrieval of a public archive does not establish review of every linked document or implementation of a stated commitment. Pinned revisions used in the audit are identified in the paper; the public supplement does not contain the full private audit archive.

No participant-level experimental dataset was analyzed. The Python calculations use published aggregate counts for a trial comparison and hypothetical inputs for illustrative sensitivities. No numerical study findings were pooled into a meta-analysis, and the review does not provide a new numerical extinction-risk forecast.


The supplement supports claim tracing and arithmetic reproduction. It does not contain a complete proof-build environment.

