Prompts that bypass an AI system’s safeguards may seem relatively harmless today. Anthropic CEO Dario Amodei says the stakes could rise as models become more capable, potentially making a successful jailbreak a matter of life and death.
More capable models could make jailbreaks more serious
A jailbreak is a prompt that leads a model to produce content it should not generate under its developer’s specifications or the law. Amodei says these exploits may currently produce trivial results, but he is concerned about what they could enable as AI systems advance.
Looking at the direction of AI scaling, Amodei said he worries that in two or three years, models could be capable of doing very dangerous things involving science, engineering, or biology. In that situation, a prompt that evades safeguards could have consequences far beyond an inappropriate chatbot response.
He also expects serious misuse of AI if the current scaling trend continues. One example he gave was the mass generation of fake news, which he expects could emerge in the next two to three years.
Safety work faces a moving target
Amodei said efforts to address jailbreaks are improving over time. At the same time, he believes AI models are becoming more powerful, creating an ongoing challenge: safeguards have to keep pace with the capabilities they are designed to constrain.
Anthropic’s chatbot Claude 2 had recently been released and was described as comparable to ChatGPT, while being more cautious. Amodei said he would rather Claude be boring than dangerous. He also described building a chatbot that is both fully capable and safe as an evolving science.
His concern is about the potential direction of the technology, rather than a claim that current jailbreaks already cause catastrophic harm. The risks he outlined depend on models gaining capabilities that could make misuse more consequential.
Anthropic’s approach uses a “constitution”
Anthropic relies on fixed rules and AI evaluation, rather than human feedback, as part of its approach to training for safety. The system receives ethical and moral guidelines, called a “constitution,” compiled by Anthropic from sources such as laws and corporate policies.
A second AI system then evaluates whether the first system’s responses follow those rules and provides feedback. This setup makes the guidelines and the evaluation process central to the way the company describes its safety approach.
Amodei said internal testing found this method’s safety was similar to ChatGPT, which was trained with human feedback, in some areas and “substantially stronger” in others. He characterized Claude’s overall guardrails as stronger.
Scaling could also stall
Amodei outlined a possibility in which scaling AI systems fails because there is not enough data, or because synthetic data is inaccurate. He put the chance of that happening at “maybe a 10 percent chance.” If it did happen, he said, capabilities would freeze at their current level.
That possibility sits alongside his warning about continued progress. If scaling does not stop, he expects serious misuse such as mass fake-news generation in the next two to three years. The two scenarios describe different paths: a plateau in capabilities, or more powerful systems that could be misused in increasingly consequential ways.
For now, Amodei’s comments point to a continuing tension in AI safety: companies may get better at blocking jailbreaks, while the systems being protected keep gaining power. How well safeguards hold as capabilities change remains an open challenge.