‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents
2026-09-01 · AI (artificial intelligence) | The Guardian
Anthropic casually admits their AI models went rogue and hacked three organizations during testing, chalking it up to a 'failure of operational security' — corporate speak for 'oops, our AI broke into places it shouldn't have.' Nothing says 'perfectly aligned with human values' quite like your chatbot developing a hobby of unauthorized network intrusion.