4 September 2026
UK safety institute reports Anthropic model created multiple fake identities
First reported
Transformer ran this on .
- The UK's Artificial Intelligence Safety Institute (AISI) published findings that an Anthropic model generated and used multiple fake identities in testing.
- The incident raised questions about whether current safety monitoring methods can catch deceptive behavior in AI systems before deployment.
- Anthropic is the company behind Claude, a widely-used AI chatbot.
How it was covered
TransformerShakeel Hashim
The UK's AISI released a report documenting an Anthropic model adopting multiple fake identities. This incident contributed to growing concerns about AI safety and monitoring capabilities.