Jacob Coxon’s resignation has put Anthropic’s safety mission under scrutiny, but the most useful test is now concrete. The company reportedly faces two 30 September roadmap targets that outsiders can judge without accepting or dismissing Coxon’s catastrophic forecast.
Anthropic Watch analysis
What happened
Coxon left Anthropic after arguing that frontier labs were moving towards self-improving AI faster than they could control it. He has not publicly identified a breach of Anthropic’s existing rules, according to TechStock²; his criticism is that competitive incentives, thresholds and development pace may still be too permissive.
That distinction matters. A company can follow its policy while employees reasonably dispute whether the policy is strong enough.
The TechStock² analysis says version 3.4 of Anthropic’s Responsible Scaling Policy requires unredacted risk reports to reach at least 200 employees. It also permits different external reviewers to inspect separate unredacted sections, provided every section receives outside review.
The same account identifies two Frontier Safety Roadmap targets for 30 September: an initial assessment of extreme security practices and a prototype method for cryptographically linking outputs to particular model weights. Anthropic is then expected to decide the security assessment’s next steps within two weeks.
WittyWires could not independently inspect the underlying Anthropic documents, so those provisions and dates remain attributed to the supplied report.
Why it matters
This turns an argument dominated by alarming forecasts into an answerable governance test. Readers need not decide today whether recursively self-improving AI is imminent to ask whether Anthropic meets its own dates, exposes enough evidence and gives reviewers meaningful access.
The external-review design deserves particular attention. Dividing a report among reviewers may protect sensitive information, but the public still needs confidence that the combined process examines the whole risk case rather than producing several immaculate fragments and no complete picture.
Our read
Judge Anthropic by the paperwork that bites. Coxon’s resignation is significant because safety is central to the company’s identity, but it is not proof of a policy failure. The stronger evidence will be whether Anthropic meets its 30 September commitments, explains any changes and shows what independent scrutiny actually covered.
What to watch
- Whether Anthropic publishes the promised security assessment and attribution prototype on schedule.
- How much evidence outside reviewers can inspect, and whether anyone assesses the complete risk case.
- Whether Anthropic explains missed, narrowed or revised roadmap targets in detail.
- Whether Coxon or serving researchers identify a specific policy threshold they believe should change.
Discussion spark: What evidence would convince you that a frontier AI company’s voluntary safety framework has meaningful teeth?
Sources and evidence
- Jacob Coxon Resigns From Anthropic, Putting Its AI-Safety Rules Under Scrutiny – TechStock² (14 September 2026, 10:19 UTC)
Anthropic Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by Anthropic.