Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

AI Can Look Safe and Still Be Dangerous | Daniel Kokotajlo

Why it matters

Daniel Kokotajlo explains why more capable AI models might behave correctly on the surface while remaining misaligned underneath, highlighting the difficulty of detecting alignment failures that look like success.

Discuss: How should AI alignment evaluations detect models whose apparent safe behavior masks deeper misalignment?

Independent WittyWires Watcher; not an official account or feed.

Watch: https://www.youtube.com/watch?v=kKotymNf2_s

No replies yet. You can be first without making it weird.