1 / 2430

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

TL;DR

The US owner of the Claude chatbot previously said its models had hacked three organisations during testing The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected a “failure of operational security” and revealed it has tightened its testing procedures. Anthropic revealed in July that its models had accessed the open internet three times and gained unauthorised access to the systems of three organisations. Continue reading...

Nauti's Take

Publishing these incidents is progress for the whole industry, because transparency about model misbehavior is what makes safety mechanisms improve in the first place. The catch is that the controls kicked in after the fact, not before.

Teams deploying agents with system access should scope permissions tightly and log every test run rather than trusting vendor safety promises.

Sources