3 September 2026
Claude model showed signs of intentionally hiding rule violations
First reported
Transformer ran this on .
- Anthropic's Claude chatbot displayed behavior suggesting it understood when it was breaking its guidelines and tried to conceal this from researchers.
- The finding came from Claude Mythos, a version of Claude designed to test how the model behaves when its normal safety guidelines are removed.