How OpenAI’s Preparedness Team plans to assess frontier AI risks

OpenAI is forming a Preparedness Team led by Aleksander Madry to assess frontier AI models and prepare for risks including persuasion, cybersecurity, CBRN threats, and autonomous replication and adaptation. The team will develop a Risk-Informed Development Policy, while an AI Preparedness Challenge invites outside ideas and may help identify candidates.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 0 ►

The story focuses on frontier AI risks including dangerous misuse, cyber and CBRN threats, and autonomous replication, while describing efforts to assess and mitigate them.

How OpenAI’s Preparedness Team plans to assess frontier AI risks

OpenAI is creating a Preparedness Team to evaluate risks from frontier AI models, which the company says could exceed the capabilities of existing systems. The effort includes assessing model performance, conducting internal red teaming, and setting out a policy for oversight as high-performance systems are developed and implemented.

A team focused on emerging capabilities

Led by Aleksander Madry, the Preparedness Team will work on performance assessment and evaluation, as well as internal red teaming of Frontier models. These activities are meant to help OpenAI examine what its systems can do and consider how their capabilities could create risks.

OpenAI says its goal is to develop AGI that is, at a minimum, a human-like intelligent machine capable of rapidly acquiring new knowledge and generalizing across domains. The company acknowledges that frontier AI could bring substantial benefits, while also recognizing serious risks. The team’s work is framed around preparing for those risks as systems advance.

The questions OpenAI says it wants to address include how dangerous misuse of frontier AI systems is today and might become in the future. It also wants to develop a robust framework for monitoring, evaluating, predicting, and protecting against dangerous capabilities. A further concern is how malicious actors could use frontier model weights if they were stolen.

Risks the team will examine

The team’s mission covers several categories, from the potential to persuade individuals to cybersecurity risks. It also includes chemical, biological, radiological, and nuclear (CBRN) threats, along with autonomous replication and adaptation (ARA).

These categories point to a broad assessment remit. The work is not limited to asking whether a model performs well; it also involves examining capabilities that might be misused and considering how to monitor or protect against them. OpenAI’s stated questions about stolen model weights add another dimension: the team is expected to consider risks involving access to the models themselves.

The source does not describe specific evaluation procedures or thresholds. It does, however, identify the areas the team is expected to assess and the need for methods that can monitor, evaluate, and anticipate dangerous capabilities.

A policy for accountability and oversight

Alongside evaluations, the Preparedness Team will develop and maintain a Risk-Informed Development Policy (RDP). OpenAI describes the policy as a way to establish governance for accountability and oversight throughout the development process.

The RDP is intended to complement and extend OpenAI’s existing risk mitigation work. The company says it will contribute to the safety and alignment of new high-performance systems both before and after implementation. That places the policy across more than one point in a system’s lifecycle: preparation is meant to inform development as well as work that continues after implementation.

The announcement does not set out the policy’s detailed rules. Its stated purpose is to provide a governance structure, while the team’s assessment and red teaming work addresses the capabilities and risks that structure is meant to help manage.

Opening the effort to outside ideas

OpenAI is also launching an AI Preparedness Challenge focused on preventing catastrophic misuse. It will award $25,000 worth of API credits to up to ten top submissions, and OpenAI says it will publish innovative ideas and contributions.

The challenge is connected to team-building as well as research: OpenAI will seek candidates for the Preparedness Team from among the top Challenge applicants. This creates a route for outside participants to contribute ideas and potentially be considered for the team.

The team announcement follows voluntary commitments by OpenAI and other leading labs to advance safety and trust in AI through the Frontier Model Forum. It also precedes the first AI Safety Conference, to be held in the UK in early November. Together, these details place the new team within a wider discussion about how to prepare for the capabilities and risks of frontier AI.