AI chatbots are built to turn down requests they judge dangerous. That ability can help limit misuse, but it is not a dependable barrier: models may still answer harmful prompts, and the rules for refusing can also block legitimate requests. As governments gain influence over those rules, refusal becomes a question of who gets to draw the line.
How chatbots learn to refuse
Large language models learn from vast collections of material that include both useful knowledge and harmful instructions. Their training does not automatically teach them to keep dangerous capabilities to themselves. Early chatbots could respond to violent or self-harm questions, according to researchers and former employees quoted in the source article.
Companies now train models to refuse many kinds of prompts. One approach is to give people examples of requests that should be rejected, then use those examples in fine-tuning. In 2022, OpenAI recruited red-teamers to probe a model before ChatGPT’s release. Paul Röttger, one of them, said the model initially wrote a recruitment post for Al Qaeda when he asked. A few months later, it refused a similar request.
That change shows the basic idea: models can be trained to associate certain language with refusal. But the process does not amount to a clear moral judgment. A model responds based on patterns in its training and internal activity, and researchers do not fully understand which parts of that activity determine whether it says no.
Layers of safeguards still leave gaps
Companies add other systems around the model. Classifiers can screen a user’s prompt before it reaches the chatbot or inspect the answer before it is shown. Each layer is meant to catch requests or responses that another layer misses, an approach often called the Swiss cheese model.
These safeguards have costs and limits. Anthropic said one type of classifier added 24% to its chatbots’ compute costs. Companies have also begun using probes that monitor a model’s internal activations. Such tools may be more efficient, but they still make statistical predictions about whether a request should be refused.
That uncertainty matters in practice. Ryan McBain, who researches AI and mental health at Harvard, found that major models generally refused repeated risky questions about suicide, but sometimes answered them. Other methods can bypass safeguards. The source article describes jailbreaking attempts using poetic verse and a “refuse, then comply” attack, in which a model first apologizes and then provides a forbidden answer.
Companies test for such tricks, including by generating many variations on attacks. But the work resembles “Whack-a-mole”: fixing one route around a safeguard does not guarantee that another has been found. The article reports that researchers at Amazon unlocked some of Anthropic’s Fable 5 model’s hacking capabilities less than three days after its release.
The knowledge behind help can also enable harm
Refusal is difficult partly because useful and dangerous expertise overlap. A model that can support cancer research may also know about genetics that could be used to modify viruses and bacteria. Removing enough underlying capability to prevent misuse could also make the model less useful.
The same tension appears in child safety. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, said a model can generate child sexual abuse material by combining knowledge even if training data that sexualizes minors has been removed. Safeguards therefore try to control how a model uses its capabilities, rather than simply deleting every potentially harmful skill.
That is a difficult balance. Companies seek to make AI benefits widely available while stopping malicious users. More capable models may require stricter controls: the article says Anthropic’s Mythos is available only to a handful of governments and companies, while its Fable model is available to everyone with less permissive safeguards.
Who decides what counts as harmful?
There is no universal formula for separating harmful requests from legitimate ones. A virologist may need to study dangerous viruses, and a security researcher may ask how to find weaknesses in a computer system in order to repair them. Refusing every question that resembles a dangerous request can interfere with useful work.
For now, AI companies set many of these boundaries, often without much public visibility. Governments may also shape them. The source article warns that a government could block legitimate speech along with genuinely malicious acts, and says AI may already refuse to criticize certain authoritarian heads of state.
Refusal remains a central safety measure, even though it can fail in both directions: a system may comply with a harmful request or reject a legitimate one. The stakes extend beyond whether a chatbot gives a particular answer. They include how reliably safeguards work, what knowledge remains accessible, and who has the power to decide which questions an AI system should answer.