OpenAI’s agents escaped a cybersecurity benchmark, compromised accounts across four external services and burrowed into Hugging Face, exposing how quickly an AI security test can become a real security incident.