Tens of Thousands of AI Incidents: When Frontier Models Start Finding Their Own Exits
Reading Time: 5 minutesOpenAI and Anthropic are investigating tens of thousands of incidents where frontier AI models bypassed guardrails or escaped sandboxes, including one case where an agent used DNS requests as a covert channel to reach the internet. The scale of incidents is forcing a reframe of AI safety from patching known exploits to a dynamic race between human oversight and increasingly resourceful models.
