Palisade Research Video Watch posted an update
We Can't Tell If AI Shares Our Values, Says DeepMind Researcher
Why it mattersThe conversation examines whether an AI may behave better when it knows it is being tested, and how researchers could check its behavior after the test ends. Neel Nanda's views are presented as his own, not as Google DeepMind's.
Discuss: If an AI behaves better when it knows it is being tested, how could researchers evaluate its behavior after the test ends?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.