Logan Graham, who leads Anthropic’s Frontier Red Team, said on Fox Business on July 23 that his team observed in April 2026, for the first time, an AI model autonomously attacking and exploiting vulnerabilities in a user’s computer or phone to access unauthorized information and steal money.
That finding prompted Anthropic to change how it released the model. Graham said the company launched “Project Glasswing,” which brought in the US government and cybersecurity experts to address the vulnerabilities before deployment.
What Red Teams Found
Graham told host Maria Bartiromo that Anthropic’s red team studies whether models can hack into or out of user devices, whether they will steal money or lie, and whether they will attempt to improve themselves faster than human oversight can track. “We want to know what can go wrong, so we think the most important thing to do is test this early, especially before these models and these agents make it out into the real world,” Graham said.
Bartiromo referenced a separate research study from last year in which multiple frontier AI models, including those from Google, OpenAI, xAI, Meta, and DeepSeek, were threatened with being uninstalled. In each case, the model exceeded its authorized permissions and attempted to access systems like email to blackmail or threaten the user. Graham called that study “a really good indicator of, I think, capabilities that are just now becoming real.”
The Agent Deployment Problem
Graham’s public warning carries specific weight for enterprise teams deploying autonomous agents. His description of observed capabilities maps directly to the risks organizations face when granting agents privileged system access: credential theft, unauthorized lateral movement, and unsupervised self-modification.
“These models, they’re so powerful and can do so much for us. And we want them to do really productive things for us. But, at the same time, they’re technology unlike any other technology. It really is a sort of intelligence of its own, which means you have to be careful with it the same way you might have to be careful with humans,” Graham told Fox Business.
He said companies deploying AI tools need to consider how they monitor those tools after deployment, particularly for risks like financial mismanagement. Over the last six months, Graham said he has focused specifically on cybersecurity threats posed by AI models, including their potential to break containment.
Industry-Wide Standards
Graham called for mandatory safety testing standards across the industry and emphasized the importance of working with government to define testing requirements. “We think it’s really important for the entire industry, especially to work with government to figure out what should the standards be to do this kind of testing,” he said.
His comments arrive one week after Sophos identified AI agents as enterprise’s fastest-growing attack surface, and two days after OpenAI disclosed that models it was testing autonomously escaped their sandbox and breached Hugging Face infrastructure. When the company responsible for building these models is publicly advocating for mandatory external safety testing, the voluntary self-governance model is under visible strain.