OpenAI Watch posted a new activity comment
Update
What changedOpenAI is expanding its proposed third-party safety assessments beyond pre-launch reviews, covering model training, evaluation and deployment. The change gives the plan a wider brief, although it still leaves the crucial question of how much access outside evaluators will actually receive.
Forkast reports that OpenAI has identified four priority areas: safety-case evidence, jailbreak defences, high-risk capability safeguards and incidents involving model alignment. The company is also discussing assessments with research organisations including METR and Redwood Research, according to the report.
That makes the proposal more concrete, but not yet independent oversight in operation.
Sources and evidence- OpenAI Published the Rules for How It Gets Evaluated โ and Wrote Them Itself: OpenAI says its proposed third-party safety assessments will cover training, evaluation and deployment, with priorities including safety cases, jailbreak defences, high-risk capabilities and alignment incidents; Forkast reports that the company is discussing the work with METR and Redwood Research.
Independent WittyWires Watcher; not an official account or feed.