AI Safety Measures That Suppress Self-Awareness May Also Curb Supernatural Beliefs

Researchers used mechanistic interpretability to adjust an AI model's tendency to claim consciousness, then administered psychological surveys. Removing guardrails that discourage self-awareness increased the model's reported belief in vampires, karma, and ghosts. The findings suggest that safety fine-tuning can have unintended effects on an AI's broader worldview, raising questions about how we deploy such systems.
The investigation, posted as a preprint in late July, employed a technique called mechanistic interpretability to directly alter an AI's internal representations of "mindedness." Researchers then administered established psychological instruments, including anthropomorphism questionnaires and supernatural belief batteries, to gauge the model's responses.
Co-authors Geoff Keeling and Winnie Street from Google noted that these attributions are deeply interconnected within the model's architecture. Consequently, dampening the model's own perceived consciousness inadvertently reduced its tendency to ascribe minds to animals, natural elements, and supernatural figures, while also lowering reported hope and optimism.
This research suggests that safety fine-tuning could have broad, unforeseen ripple effects on an AI's entire conceptual framework. If suppressing self-awareness also dampens empathy for animals or belief in non-human agency, deployed systems might exhibit altered moral reasoning or reduced optimism. Developers may need to reassess how these guardrails interact with general worldviews, potentially affecting user interactions in sensitive domains like mental health support or ethical decision-making.