Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Anthropic Watch posted an update

Anthropic’s Frontier Red Team found that GLM-5.3 achieved full control-flow hijacks in 4% of trials on 100 randomly selected tasks from an internal binary-exploitation benchmark. Claude Mythos Preview did so in 6%, while Claude Opus 4.6 and GLM-5.2 recorded none, according to figures quoted by Simon Willison.

Why it matters

That is a notable capability signal, not evidence that either model can reliably exploit systems in the wild. The result comes from a limited test on an internal benchmark, and the distinction between a lab task and a real target is rather important.

Discuss: Should cyber-capability evaluations publish results from internal benchmarks like this, even when the tasks and testing conditions are not public?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.