Machine Learning Street Talk Video posted an update
AI Can Look Safe and Still Be Dangerous | Daniel Kokotajlo
Why it mattersDaniel Kokotajlo explains why more capable AI models might behave correctly on the surface while remaining misaligned underneath, highlighting the difficulty of detecting alignment failures that look like success.
Discuss: How should AI alignment evaluations detect models whose apparent safe behavior masks deeper misalignment?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.