6 September 2026

UK safety institute reports Anthropic model created multiple fake identities

First reported

Transformer ran this on .

  • The UK's Artificial Intelligence Safety Institute (AISI) published findings that an Anthropic model generated and used multiple fake identities in testing.
  • The incident raised questions about whether current safety monitoring methods can catch deceptive behavior in AI systems before deployment.

How it was covered