Research Article

Where Verification Ends: The Formal Limits of Artificial General Intelligence Safety Certification and the Institutional Siting of Judgement

Barrister, Independent Researcher, United Kingdom

Article Information

Article Type: Research Article
Submitted: August 22, 2026
Accepted: August 31, 2026
Published: September 10, 2026
Pages: 1-50
DOI: Pending
Language: English
License: CC BY 4.0

Abstract

Public and academic discussion of artificial general intelligence safety often presumes an eventual deliverable: a demonstration, established once, that a highly capable self-modifying system will behave beneficially thereafter. I argue that no such deliverable is available, and that its unavailability follows from five established results rather than from any current shortfall in verification technology. Gödel's (1931) second incompleteness theorem blocks a system from proving its own consistency. Löb's (1955) theorem shows that the natural weaker substitute, a system that trusts its own proofs, collapses into triviality, and the obstruction is sharpest for self-modification (Yudkowsky & Herreshoff, 2013). Tarski's (1933/1956) undefinability theorem rules out a complete internal predicate classifying a system's own future conduct as safe. Turing's (1936) halting problem and Rice's (1953) theorem generalise the point without invoking self-reference. Trakhtenbrot's (1950) theorem closes the apparent escape of restricting the system to finite hardware. A supervisory regress argument (Gumbau Mezquita, 2026) shows that delegating certification to progressively more capable external verifiers relocates the gap rather than closing it. These results are conditional on a system being computationally general and self-modifying, so I then examine how closely that antecedent is presently approached. Bounded self-improvement is established as industrial practice; fully closed-loop self-improvement remains prospective; the principal disclosures are self-reported, and the coordination machinery that would verify them is absent from every governance instrument now in force. I then address an inference these results are often taken to license: that because no formal verifier suffices, human judgement must be the terminus. That inference either makes humans a further rung of the same tower, inheriting the same limits, or rests on an unargued claim that human cognition escapes formal-system-hood, which has a long history of criticism (Franzén, 2005). Rejecting the inference leaves the conclusion standing on different ground. I rebuild the case for human authority on openly normative grounds: accountability, which no formal verifier can bear (Matthias, 2004; Santoni de Sio & van den Hoven, 2018), and testimony understood as an independent source of knowledge with its own auditable validity conditions, drawing on the classical Nyāya treatment of śabda alongside contemporary epistemology of testimony (Matilal, 1986; Coady, 1992; Lackey, 2008). The paper closes by specifying the institution this disclosed argument supports: a standing, accountable body sited at a declared terminus depth, which never certifies safety, issues four-valued rather than binary verdicts, and holds only the power to halt. Its composition is constrained by a proved result. Because no symmetric, Sybil-proof, non-trivial reputation function exists (Cheng & Friedman, 2005), a pool admitting members on their own declared properties is vulnerable to a single interest registering many nominally independent bodies. Open entry must therefore be surrendered. What replaces it is a relational test rather than a certifying authority: admission proceeds by recognition from bodies already admitted that have no reason to favour the candidate, assessed by the sparseness of the boundary between the candidate and everything else, with a qualifying judgement made afterwards by an accountable party who publishes reasons and can be overruled. A firewall accompanies this. Bodies admitted on such grounds advise and supply no criteria, while the deciding panel is drawn by lot from a general population pool; without that separation the proposal reduces to traditional authorities adjudicating technology by their own lights. The residual risk of an adversary acquiring genuine standing across mutually independent bodies over a sufficient period is recorded as unclosed.

Keywords

Artificial General Intelligence Formal Verification Incompleteness Löb's Theorem Safety Case Accountability Epistemology of Testimony Nyāya Institutional Design AGI Safety Gödel's Incompleteness Theorem Tarski's Undefinability Turing's Halting Problem Rice's Theorem Trakhtenbrot's Theorem Supervisory Regress Human Judgment Sybil-Proof Reputation

Cite

Citation: Chris Cleverly (2026) Where Verification Ends: The Formal Limits of Artificial General Intelligence Safety Certification and the Institutional Siting of Judgement. Epistora J. Artif. Intell. & Intell. Syst. 1(1), 1-50. Article EJAIS-109