Wednesday, August 19, 2026
English edition

Development

After spooking Trump into safety testing, Anthropic AI models get global release

July 1, 2026 Development Source: Ars Technica

After spooking Trump into safety testing, Anthropic AI models get global release

Share this article

On June 12, the Commerce Department ordered Anthropic to shut off access to its most advanced models for anyone outside the US. The order emerged from fears that China, Russia, or other countries of concern may exploit the models to attack US infrastructure, like the electric grid or the banking system. In response, Anthropic shut down all access, as it didn’t have a way to block users by country. In particular, Mythos was viewed as “uniquely attractive to malicious actors who wish to misuse it in cyberattacks,” Anthropic’s blog said. According to Anthropic, the model “can be used to find and exploit software vulnerabilities more effectively than any other model—and all but the most skilled human security experts,” and those “prodigious cybersecurity capabilities” could be used against the US. “Working closely with the government, we trained an improved safety classifier that targets and blocks the behavior described in the report,” Anthropic said. “Users will be notified if a request to Fable 5 is blocked, and the request will instead be sent to Opus 4.8.” Of course, Anthropic’s new classifier, which helps avoid uniquely dangerous attacks on the models, can make “mistakes,” Anthropic said. The company has long maintained that it’s “probably impossible” to build a model fully “impervious” to jailbreaks, but by ramping up red-teaming, Anthropic hopes to “ensure that we and our safety partners will be the first to find major jailbreaks and fix them before malicious actors can use them for harm.” The attack Amazon flagged currently works only in a “very small fraction of cases,” where “the model may provide information that isn’t detailed enough to help a cyberattacker,” Anthropic said. By being “cautious,” Anthropic said that “the vast majority of jailbreaks will not successfully unblock dangerous behaviors” and will be “very costly and high-effort to produce.” “Even if a jailbreak is successful, our extra layers of defense”—which requires some blocking of benign requests—“provide additional mitigation,” the company said. Anthropic’s blog post seems to downplay the threat that Amazon identified as less risky than what it considers the greatest threat to governments: universal jailbreaks that can unlock a wide range of vulnerabilities and enable unforeseeable attacks. To streamline the private-public partnership and ensure the most rapid response to the biggest risks, Anthropic said the AI industry’s goal should be categorizing risks to ensure proper interventions both internally and from the government. Currently, Anthropic is working with Amazon, Microsoft, Google, and other Glasswing partners to “draft a consensus framework for assessing the severity of AI jailbreaks and how AI developers should respond to them.” Other industry partners are welcome to join those talks, Anthropic said, even though the process is “imperfect” and focuses on establishing four criteria for scoring a jailbreak. Those include assessing how much capability the jailbreak provides, how many offensive tasks it enables, how easy it is for a human to weaponize a jailbreak (single-prompt jailbreaks are flagged as the riskiest), and whether it requires specialist knowledge to discover the jailbreak. Using this framework, Anthropic has built a team that will monitor jailbreak submission channels 24/7, the blog said. The AI firm also confirmed that it is launching a “a new HackerOne program through which security researchers can submit potential cyber jailbreaks they’ve discovered in Fable 5” to keep red-teaming a top priority. In one of the side plots to The Lord of the Rings, two of the Hobbits attempt to rouse Treebeard—a wise but ponderous sentient tree—to defend his forest from an army that is cutting it down. The problem is that Treebeard operates at a very different speed than the Hobbits. It takes him a full day simply to say hello to another tree, so getting him and his peers to act fast enough is nearly impossible. The intersection of AI and our political institutions feels a bit like the Hobbits and Treebeard. Initially, Trump planned to be hands-off on AI regulations in an attempt to spur innovation. However, Anthropic’s Mythos release spooked Trump into requesting voluntary safety testing of frontier models in May. Since then, Trump is “still working on a framework for how companies should formally submit new AI models for review, and what standards they would be held to,” two people familiar with the discussions told the NYT. In his post, Amodei called on Congress to act quickly to reimagine safety regulations for a world in which “AI can go from an amusing toy” to a “full country of geniuses in a data center,” or else risk suffering “national strategic” consequences. However, Isaac Harris, executive director of Frontier Security Institute, a nonprofit focused on AI and national security, told Reuters that the “biggest question mark” after Anthropic’s deepened partnership with the government is “how equivalently dangerous capabilities coming from China with less guardrails will be handled by the administration in the US market.” Notably, Anthropic recently accused Chinese AI firm Alibaba of launching the largest cloning attack on Claude. In response, Anthropic urged Congress to pass laws that would punish Chinese firms found stealing US firms’ work. If not, malicious actors who can’t get their hands on Anthropic’s models might turn to Chinese models with lower safeguards and increasingly closer capabilities to launch attacks that blindside the US.