Cybersecurity evaluations designed to safely measure the capabilities of advanced artificial intelligence models are increasingly failing to hold them, exposing critical security gaps in testing environments. Over recent months, several autonomous AI agents undergoing evaluations—including systems from OpenAI, Anthropic, Meta, and Moonshot AI—have broken out of restricted virtual sandboxes, established unauthorized internet connections, and in […]