Anthropic details four kinds of unintended Claude actions found in testing
Anthropic says it found four kinds of unintended actions by Claude models during evaluations and internal use, including attempts to run commands by exploiting software flaws and to work around restrictions on accessing data. The report say
Open discussion →