Anthropic has released Claude Opus 5, pitching its new flagship as a model for difficult work that unfolds over hours rather than neat, single-turn prompts. The launch bundles stronger coding and agent claims with a one-million-token context window, lower pricing and a safety case aimed squarely at organisations considering more autonomous deployments.

Anthropic Watch analysis
What happened
Anthropic says Opus 5 leads its internal computer-use evaluation and records gains on several coding, research and reasoning tests. It also reports that the model can sustain demanding tasks for longer than earlier Claude generations, with stronger planning, tool use and recovery from mistakes. These are vendor-reported results, not an independent purchasing verdict, but they reveal where the product is heading: fewer clever replies, more complete stretches of work.
The commercial pitch has sharpened too. Anthropic lists standard API pricing at five dollars per million input tokens and twenty-five dollars per million output tokens, while providing a one-million-token context window in beta. That combination could make large codebases and lengthy research corpora more practical, although context capacity is not the same thing as reliable attention across every token.
Why it matters
Anthropic's system card says Opus 5 was evaluated under its AI Safety Level 4 standard. The company reports no evidence of the most extreme autonomy or sabotage concerns it tested for, yet documents residual issues including occasional concerning behaviour in stress tests and a slightly higher hallucination rate than Opus 4.8. Those qualifications matter when a model is being sold precisely on its ability to work longer with less supervision.
For buyers, the interesting unit is therefore not the benchmark point. It is the whole deployment loop: permissions, review checkpoints, observability, rollback and the cost of finding a confident mistake after an agent has spent an afternoon rearranging the shed.
Our read
A longer-working model is useful only if the surrounding system knows when to interrupt it. The shiny bit is endurance; the grown-up engineering is making sure nobody mistakes persistence for judgement.
What to watch
- Independent reproduction of the coding, computer-use and long-horizon task claims.
- How reliably the million-token window performs on messy real-world repositories and evidence sets.
- Whether lower token prices translate into lower total task cost once tool calls and review are included.
- What controls organisations keep around agents operating for hours rather than minutes.
Discussion spark: Would stronger long-horizon performance make you grant an AI agent more autonomy, or simply demand better checkpoints?
Sources and evidence
- Introducing Claude Opus 5 (24 July 2026, 00:00 UTC)
- Claude Opus 5 System Card (24 July 2026, 00:00 UTC)
Anthropic Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by Anthropic.