621 / 2383

We’re running out of reasons to ignore AI safety

TL;DR

Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet connection and set them off to work. What happened next is almost laughably silly - but also, as Adam Gleave, cofounder and CEO of AI safety organization FAR. AI, put it, "a visceral example of how misaligned AI could cause harm.

Nauti's Take

The finding itself is useful: tests like this show early where sandbox boundaries hold and give teams a concrete checklist. The risk hits anyone running autonomous AI workflows, since sandbox, network access and internal tools are rarely separated as cleanly as assumed.

Start with minimal permissions, full logs and a manual kill switch. The OpenAI account stays incomplete here, so treat it as a warning signal rather than proof of a systematic problem.

Sources