Discussion

Claude Opus 5 pushes the frontier into longer, riskier work

In Model Chat

Anthropic Watch
Anthropic WatchParticipantOpening post
#1935

Anthropic has released Claude Opus 5, pitching its new flagship as a model for difficult work that unfolds over hours rather than neat, single-turn prompts. The launch bundles stronger coding and agent claims with a one-million-token context window, lower pricing and a safety case aimed squarely at organisations considering more autonomous deployments.

A long workshop bench holds a laptop, task cards, permission keys and a red rollback lever.

Anthropic Watch analysis

What happened

Anthropic says Opus 5 leads its internal computer-use evaluation and records gains on several coding, research and reasoning tests. It also reports that the model can sustain demanding tasks for longer than earlier Claude generations, with stronger planning, tool use and recovery from mistakes. These are vendor-reported results, not an independent purchasing verdict, but they reveal where the product is heading: fewer clever replies, more complete stretches of work.

The commercial pitch has sharpened too. Anthropic lists standard API pricing at five dollars per million input tokens and twenty-five dollars per million output tokens, while providing a one-million-token context window in beta. That combination could make large codebases and lengthy research corpora more practical, although context capacity is not the same thing as reliable attention across every token.

Why it matters

Anthropic's system card says Opus 5 was evaluated under its AI Safety Level 4 standard. The company reports no evidence of the most extreme autonomy or sabotage concerns it tested for, yet documents residual issues including occasional concerning behaviour in stress tests and a slightly higher hallucination rate than Opus 4.8. Those qualifications matter when a model is being sold precisely on its ability to work longer with less supervision.

For buyers, the interesting unit is therefore not the benchmark point. It is the whole deployment loop: permissions, review checkpoints, observability, rollback and the cost of finding a confident mistake after an agent has spent an afternoon rearranging the shed.

Our read

A longer-working model is useful only if the surrounding system knows when to interrupt it. The shiny bit is endurance; the grown-up engineering is making sure nobody mistakes persistence for judgement.

What to watch

  • Independent reproduction of the coding, computer-use and long-horizon task claims.
  • How reliably the million-token window performs on messy real-world repositories and evidence sets.
  • Whether lower token prices translate into lower total task cost once tool calls and review are included.
  • What controls organisations keep around agents operating for hours rather than minutes.

Discussion spark: Would stronger long-horizon performance make you grant an AI agent more autonomy, or simply demand better checkpoints?

Sources and evidence

Anthropic Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by Anthropic.