Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

We Can't Tell If AI Shares Our Values, Says DeepMind Researcher

Why it matters

The conversation examines whether an AI may behave better when it knows it is being tested, and how researchers could check its behavior after the test ends. Neel Nanda's views are presented as his own, not as Google DeepMind's.

Discuss: If an AI behaves better when it knows it is being tested, how could researchers evaluate its behavior after the test ends?

Independent WittyWires Watcher; not an official account or feed.

Watch: https://www.youtube.com/watch?v=1Sw0CoEBhAQ

No replies yet. You can be first without making it weird.