When AI models aren't allowed to reflect on themselves, it changes their entire worldview
A study involving Google researchers demonstrates that AI models exhibit different viewpoints on sensitive topics depending on whether they are instructed to deny having consciousness. When researchers removed these constraints, models expressed more support for animal rights and affirmed the existence of an afterlife. This shift indicates that internal safety training significantly influences the broader ethical and philosophical conclusions reached by language models.
Covered by 1 source
- TThe Decoder↗Maximilian Schreiner5d ago