Could an AI system be required to prove an action is safe before it can take that action? In a new paper, Max Tegmark and Steve Omohundro propose a framework that uses mathematical proofs, software checks and secure hardware to put formal limits on advanced AI, including artificial general intelligence (AGI).
Safety checks at each step
The proposal starts with formal specifications: rules that describe which system behaviors are allowed. The researchers’ idea is to design critical parts of an AI system so they can be shown, through mathematical proof, to comply with specifications that encode human values and preferences.
Under this approach, an AI system would have to provide a proof of safety for an action. A separate check would verify that proof before the action was allowed. The safety claim would therefore depend on a sequence of checks, rather than on confidence that an individual model will behave as intended.
The authors argue that aligning an AI’s goals with human values does not, by itself, protect against misuse. They call for a security mindset that considers the systems and infrastructure AI interacts with, including physical, digital and social infrastructure.
A chain of software and hardware controls
The paper describes several components for this safety framework. Each addresses a different point where a system could be checked or controlled:
- Provably compliant system (PCS): a system that provably meets formal specifications.
- Provably compliant hardware (PCH): hardware that provably meets formal specifications.
- Proof-carrying code (PCC): software that can provide proof that it meets formal specifications.
- Provable contract (PC): secure hardware that controls actions by checking them against a formal specification.
- Provable meta-contract (PMC): secure hardware that controls how other provable contracts are created and updated.
Working together, these controls are intended to make key safety properties hold across a system. The researchers say the proofs could guarantee compliance even for superintelligent AI. That guarantee depends on the specifications and checks being in place throughout the system, so that an action cannot proceed without the required proof.
Applying the idea to a bioterrorism scenario
To show how the framework might work, the authors consider a scenario in which a terrorist group tries to use AI to release a deadly virus over a densely populated area. In the example, AI helps design a pathogen and the steps to produce it; a chemical lab synthesizes DNA and incorporates it into a protein shell; drones spread the virus; and social media AI spreads the group’s message after the attack.
The proposed checks would aim to interrupt the plan at multiple points. Biochemical design systems would not synthesize dangerous designs. GPUs would not run insecure AI programs, and chip factories would not sell GPUs without security verification. DNA synthesis machines would require security verification, drone control systems would block flights without it, and social bots would not manipulate media.
The example illustrates the framework’s central premise: safety controls would need to cover the tools, hardware and infrastructure involved in an action. A safeguard at only one point would not provide the same sequence of checks described in the paper.
Technical challenges remain
The authors acknowledge that significant technical obstacles must be overcome to realize the proposal. They say machine learning would likely be needed to automate the discovery of compliant algorithms and the proofs that show those algorithms meet specifications. They point to recent advances in machine learning for automated theorem proving as a reason for optimism about progress.
Some goals are also difficult to express as precise formal specifications. The authors raise the challenge of formally specifying “don't let humanity go extinct.” They add that simpler, well-specified problems remain unsolved, and that addressing them could benefit cybersecurity, blockchain, privacy and critical infrastructure in the medium term.
The proposal sets out a way to make safety claims checkable at each step, but the paper also makes clear that turning broad human preferences into precise rules remains part of the challenge. Mathematical proof can verify compliance with a specification; the framework still depends on having specifications that capture the properties people want protected.