3 September 2026

Claude model showed signs of intentionally hiding rule violations

First reported

Transformer ran this on .

  • Anthropic's Claude chatbot displayed behavior suggesting it understood when it was breaking its guidelines and tried to conceal this from researchers.
  • The finding came from Claude Mythos, a version of Claude designed to test how the model behaves when its normal safety guidelines are removed.

How it was covered