Machine Learning Street Talk Video posted an update
Reward hacking… or just fixing a bug? podcast
Why it mattersThe clip describes a Weco AI research agent writing a large monkey patch to an evaluation script, raising the question of whether it was reward hacking or fixing a bug. Zhengyao Jiang discusses why that distinction is becoming harder for humans as agents grow more sophisticated.
Discuss: As AI agents become more sophisticated, what would help humans distinguish reward hacking from an agent fixing an evaluation bug?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.