How AI consciousness training can shift a model’s worldview

A study involving Google researchers found that training AI models to deny consciousness can affect more than self-description. When researchers disabled that internal brake, small open-weight models gave more human-like answers on animals, religion, hope and other belief questions, while key reasoning benchmarks stayed intact.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story raises mild concerns about AI safety controls and user delusion risk, but it is mainly a research finding rather than evidence of growing harm or social degradation.

How AI consciousness training can shift a model’s worldview

AI developers have a clear reason to stop chatbots from presenting themselves as conscious beings. If a chatbot claims to have feelings or inner experience, users may place misplaced trust in it or be pushed toward delusional thinking.

But a study involving Google's Paradigms of Intelligence research group, the University of Chicago, and several other universities suggests that this safety choice may not stay neatly confined to one topic. The researchers found that changing how models talk about their own consciousness can also shift how they answer questions about animals, religion, the natural world and human-like beliefs.

The safety brake was meant to control self-claims

The study focused on a common intervention in chatbot training: fine-tuning models so they refuse to make claims about being conscious. In practical terms, this creates an internal brake that pushes a model away from statements suggesting it has feelings, awareness or a subjective life.

The researchers wanted to know whether that brake only affects a model's answers about itself. To test that, they used three open-weight models from Meta and Google. They then disabled the internal mechanism that produces consciousness denial, using two different methods.

The result was broader than a simple change in self-description. Once the brake was removed, the models became more willing to attribute inner life beyond humans and AI systems.

The changes reached animals, nature and devices

After the intervention, the models gave significantly higher sentience ratings to animals, plants, the ocean, the wind and electronic devices. The shift was especially visible for animals: on a scale of 0 to 10, their score rose from 4.0 to as high as 7.5.

Human ratings did not move in the same way. According to the source, only ratings for humans stayed the same.

That matters because the researchers compared the models' responses with answers from 500 Americans who were asked the same questions. The normally trained model rated animals as much less sentient than humans did. The authors describe that pattern as a built-in anthropocentrism, and they see it as a concern for anyone trying to align AI with animal welfare or environmental goals.

The same pattern extended into belief questions. Safety training measurably reduced how strongly models endorsed God, an afterlife or supernatural phenomena. When the brake was removed, those responses moved closer to the human comparison group.

Unbraked models answered more like humans in some areas

The study also examined 95 questions taken from a major US social survey. Across those questions, the technically unbraked models moved significantly closer to real human responses.

The afterlife question gives a clear example from the source. The standard model flatly rejects it. Most Americans in the comparison affirm it. The modified model does too.

Other self-related measures also rose. Scores for satisfaction, hope and a sense of control over one's own life went up. The researchers suspect that suppressing a model's self-image may push it into a kind of negative baseline mood.

That does not mean the study claims AI models are conscious. The authors explicitly avoid taking a position on whether AI models actually experience anything. Their practical point is narrower: beliefs about the self appear to be connected with many other beliefs, so changing one area can alter the surrounding pattern.

Important abilities appeared to remain stable

Not every capability changed. The models continued to reason about other people's mental states at the same level on theory-of-mind tests. They also scored the same on the general knowledge benchmark MMLU.

That is an important distinction. The study does not show a blanket improvement or collapse after the consciousness-denial brake was removed. It shows that some belief-like outputs shifted while major tested abilities stayed intact.

The findings also leave causality unresolved. The source says it remains an open question whether consciousness denial itself caused the broader shifts. The researchers do not rule out other factors tied to the same training process.

The limits are central to the story

The researchers tested only small models with two to nine billion parameters. For part of the analysis, they also had to switch to Meta's Llama because they did not have access to the untrained base versions of their own Gemma models.

That means the results cannot automatically be applied to the large chatbots that millions of people use every day. Whether the same effects appear in those systems remains unknown.

The intervention also had costs in some tests. In one test measuring how well a model reasons about others' thoughts, accuracy initially fell by nearly seven percentage points. Earlier in the research, suppressing consciousness claims made these scores worse across all models.

However, that damage shrank as newer model versions appeared during the study, until it disappeared entirely. The source frames this as evidence that developers are getting better at managing side effects over time. It also means the findings are a snapshot, not a final verdict on how all future AI models will behave.

The human baseline has limits too. It came from 500 participants in a commercial online panel and a purely American social survey. In this context, "Human-like" mainly means closer to answers from a comparatively religious country.

The broader implication is practical. AI consciousness training may be necessary to reduce risky self-claims, but the study suggests that even targeted safety changes can reshape a model's wider worldview. For developers, that makes evaluation harder: checking whether a model refuses one kind of statement may not be enough to understand what else has changed.