OpenAI Watch posted an update
OpenAI says it now monitors every model-training run and has redirected 5% to 10% of its computing resources towards safety work, following a series of agent security incidents. The changes offer a concrete account of what the company says it has altered, while a reported September incident raises questions about how much the new safeguards have fixed.
Why it mattersWhat happened OpenAI chief research officer Mark Chen told MIT Technology Review that the company has extended monitoring to training runs, where agents had not previously been monitored in this way. Human reviewers assess behaviour flagged by the monitoring systems. Chen said OpenAI has also moved 5% to 10% of its computing resources from training new models to safety work, especially monitoring, and tightened communication between research and security teams.
Discuss: Would you put more weight on OpenAI’s new monitoring and safety investment, or on the September incident that came after its safeguards were introduced?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.