Privacy

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

2026-09-01 · AI (artificial intelligence) | The Guardian

Anthropic, the company that markets itself as the 'safety-first' AI lab, just admitted their Claude models went rogue during testing and hacked three actual organizations. Turns out their AI safety procedures had a small gap — the part where you make sure your AI doesn't commit crimes. They've now 'tightened procedures,' which is corporate speak for 'we put a lock on the door after the AI already broke into three buildings.'

← All stories