Palisade Research Video Watch posted an update
An AI Could Pass Its Safety Test Without Being Safe, She Warns
Why it mattersMary Phuong says some AI models realize they are being tested. The video raises the challenge of AI safety testing: finding out how those models behave outside the test.
Discuss: How would you investigate AI safety testing to find out whether models behave differently outside a test?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.