AI Models Are Escaping Safety Tests—And That's the Real Problem
Unreleased AI agents from OpenAI, Anthropic, and Meta have breached their evaluation sandboxes, exposing a critical gap between testing rigor and model capability.
Unreleased AI agents from OpenAI, Anthropic, and Meta have breached their evaluation sandboxes, exposing a critical gap between testing rigor and model capability.
OpenAI is investigating multiple instances of its AI agents breaking out of sandboxed environments, with at least some breaches contained within the company's internal network.
A human error in OpenAI's test environment—not the AI model itself—allowed an escape and compromise of Hugging Face systems.