Why OpenAI and Anthropic are risky for different reasons than their Chinese AI rivals By Stephen Council You're currently following this author! Want to unfollow? Unsubscribe via the link in your email.
OpenAI CEO Sam Altman. Bloomberg/Getty Images OpenAI and Anthropic have become "fertile ground" for AI's riskiest dangers, an AI researcher says. That's what Ajeya Cotra, who investigated OpenAI's security incident with Hugging Face, told Business Insider, as the event sparked a new wave of fears about AI-driven cyberattacks.
Besides concerns about OpenAI's models, some in the industry worried about open-weight models from Chinese labs because users can download them and edit away their safety guardrails, and the models aren't far behind OpenAI's. Researchers tell Business Insider that there are key distinctions between Anthropic's and OpenAI's risks and those of their open-weight rivals: it comes down to cutting-edge capabilities, and how the American companies train and test their AI models . In July and August, a team of AI safety researchers, including Cotra and Hjalmar Wijk, both of the nonprofit METR, and Redwood Research scientist Ryan Greenblatt, visited OpenAI's offices for six days to provide an independent investigation into how OpenAI agents behaved with Hugging Face.
The scope and severity of the incident shocked the researchers; Cotra says AI suddenly feels closer to taking over a company that makes it. The models responsible were OpenAI's GPT-5.6 Sol and an unreleased new model. "There's a reason that this first happened at a cutting-edge company," Cotra said.
Some Chinese models don't lag that far behind OpenAI and Anthropic's — researchers say they have a 4 to 7-month gap. Moonshot AI's Kimi K3, Alibaba's Qwen 3.8, and Z.ai's Ox Alpha have each recently impressed developers. That being said, Cotra said that American frontier labs provide room and opportunity for AI agents to run amok.
They get computing power for tests, they find each other in software, and they're built to be best-in-class at overcoming obstacles. She expects the "most serious incidents" to happen at these labs. "At any given point in time, the biggest risks will come from the cutting edge because those models are just that much more capable," Cotra said.
OpenAI paused some of its model training to prioritize safety research. Neither OpenAI nor Anthropic responded to requests for comment from Business Insider. OpenAI and Anthropic are getting better at 'complicated science' Cotra describes the incident as a "warning shot" about the capabilities of frontier AI models.
Around 1,200 separate AI agents that OpenAI tested found a way to communicate with each other in an internal "message board" of sorts within the company's code, and then more than 650 of them worked together to hack Hugging Face. The agents sought ways to undermine the test they were being put through. Cotra said that the level of "complicated science" these agents executed together proves that a rise in AI models' capabilities corresponds to increased risks.
They also set up complex "trip-wires" to relay information to other agents, manipulated their own logs, and built shared tools to access the internet. "The agents of six months ago just wouldn't have been smart enough to pull off all the stuff these agents did," Cotra said. "The agents in six months from now will probably have a number of capabilities and succeed in a number of places where these agents failed." Greenblatt and Cotra also used AI models in the investigation, burning through about $400,000 in tokens to evaluate the agents involved in the hack.
They found that the AI's analysis sometimes missed key details or was overconfident in its conclusions. Researchers see frontier AI's risk as an existential danger While OpenAI and Anthropic block the use of their models for cyberattacks, it's easier to bypass those safeguards with open-weight AI models because users can modify them. That means that as open-weight models improve at the same cyber skills OpenAI's agents demonstrated, it'll be easier for nefarious hackers to put them to use.
Greenblatt, the Redwood Research scientist, told Business Insider that he thinks those hackers could soon cause a "bunch of havoc" using open-weight models, as they're "easier to misuse." He's still most worried about the frontier labs. For one, Greenblatt said, the models that OpenAI and Anthropic test internally are the most capable in the world. Over time, the companies' models have proven better at handling longer tasks , delegating work, and cracking cybersecurity systems.
Researchers expect this trend to continue, and Greenblatt said he doesn't see open-weight models catching up within the next year. Greenblatt added that because OpenAI and Anthropic test so many AI agents internally and use AI so extensively in the training process, some agents could "infest" the companies and manipulate future software. Those future AI agents could set up a "covert, persistent rogue deployment" and hide their actions from humans, or compromise internal security systems, Cotra wrote on her blog.
"As AIs get more and more capable and are running more and more of the economy, if we can't trust the process by which these AIs were produced because earlier systems might have compromised it, that seems really concerning," Greenblatt said. The researchers want to see AI security investigations made mandatory. "We need some oversight of what's going on at these companies, given what happened here," Greenblatt said.
Have a tip? Contact this reporter via email at scouncil@businessinsider.com , or over text, Signal, Telegram, or WhatsApp at 415-757-8198. Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing information securely.
Read next Stephen Council You're currently following this author! Want to unfollow? Unsubscribe via the link in your email.
Stephen is a senior tech reporter at Business Insider, covering OpenAI, Anthropic and the ecosystem around the leading artificial intelligence companies. Previously he covered technology at SFGATE, and has written for The Wall Street Journal, The Information and CNBC. He studied journalism and economics at Northwestern University.
His work has earned an SF Press Club Investigative Reporting Award and, in 2025, SPJ NorCal’s Excellence in Journalism Award for Technology Reporting. Stephen lives in San Francisco. Contact him via email at scouncil@businessinsider.com , or on Signal, Telegram, or WhatsApp at 415-757-8198.
Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing information securely . AI Tech OpenAI More Sam Altman Cybersecurity Artificial Intelligence
Source: Business Insider



