AI risk checks could shape decisions before models are deployed

Researchers from Google DeepMind, OpenAI, Anthropic and academic institutions propose ongoing evaluations to detect dangerous AI capabilities and whether models might use them to cause harm. They recommend responsibilities for developers and policymakers, while warning that model evaluations cannot catch every risk.

WTF Index TERMINATOR
◄ Terminator 3 Idiocracy 0 ►

The story focuses on detecting and managing potentially dangerous AI capabilities, though its emphasis is on safeguards.

AI risk checks could shape decisions before models are deployed

As AI systems become more capable, researchers say developers and policymakers need ways to spot risks before those systems are widely used. A proposal from researchers at leading AI companies and academic institutions outlines an early warning framework for evaluating large-scale models and responding to concerning results.

Two questions for evaluating AI risk

The paper, “Model evaluation for extreme risks,” says risk assessments should examine both what a model can do and how likely it is to use those abilities in harmful ways. The first question concerns dangerous capabilities: whether a system could threaten security, exert influence, or evade oversight.

The second concerns alignment: whether a model tends to apply its capabilities in ways that could cause harm. The researchers say evaluations should check whether a model behaves as intended across a wide range of scenarios and, where possible, examine its internal workings.

The distinction matters because capability alone does not show how a system will behave. An evaluation of behavior, in turn, needs to account for what the model can actually do. The proposed framework treats both as parts of understanding potential risks.

Start early and keep checking

The researchers argue that evaluations should begin as early as possible and continue through development and use. Early checks could inform responsible training and use, while ongoing assessment could help developers identify risky properties as they emerge.

They also call for structured access to models for external security researchers and reviewers. Outside evaluations could add scrutiny beyond the developer’s own work and support transparency, alongside suitable security mechanisms.

The paper describes the goal as informing policymakers and other stakeholders and helping companies make responsible decisions about training, deployment, and security. It argues that developers operating at the frontier of AI capabilities need processes to track concerning model properties and respond to evaluation results.

What developers and policymakers could do

The study places particular responsibility on frontier AI developers. It says large companies have resources to carry out this work and are among the actors most likely to unintentionally develop or release systems that pose extreme risks.

For developers, the paper recommends investing in evaluation research, creating internal policies for conducting and responding to assessments, and supporting outside research through model access and other assistance. It also encourages developers to educate policymakers and take part in discussions about standards.

For policymakers, the proposal calls for a governance structure for evaluating and regulating AI. Suggested steps include tracking dangerous capabilities and progress in alignment across frontier AI research and development, and creating formal reporting processes for extreme risk evaluations.

Policymakers could also invest in external safety evaluation and create forums where developers, academic researchers, and government representatives can discuss results. The researchers further suggest external audits of highly capable, general-purpose AI systems and their developers’ risk assessments, as well as rules making clear that models posing extreme risks should not be deployed.

Evaluations have limits

The researchers caution that model evaluation is not a complete solution. Some risks may be difficult to detect because they depend on forces beyond the model, including complex social, political, or economic conditions.

They therefore argue that evaluations need to sit alongside other risk assessment tools and a broader commitment to safety across industry, government, and civil society. The proposal presents evaluation as one part of responsible AI development, intended to guide decisions without promising that every danger can be identified in advance.