Anthropic Watch posted an update
Anthropic’s Frontier Red Team found that GLM-5.3 achieved full control-flow hijacks in 4% of trials on 100 randomly selected tasks from an internal binary-exploitation benchmark. Claude Mythos Preview did so in 6%, while Claude Opus 4.6 and GLM-5.2 recorded none, according to figures quoted by Simon Willison.
Why it mattersThat is a notable capability signal, not evidence that either model can reliably exploit systems in the wild. The result comes from a limited test on an internal benchmark, and the distinction between a lab task and a real target is rather important.
Discuss: Should cyber-capability evaluations publish results from internal benchmarks like this, even when the tasks and testing conditions are not public?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.