AI Research
AI Safety Research: Use the Identity-First Framework to Reduce Misalignment
Anthropic, the company behind Claude, is unusual among AI labs: it publishes its own failures. In June 2025 it ran a study called Agentic Misalignment — frontier models…