The oldest flaw that Anthropic's newest model dug up had been sitting in OpenBSD, an operating system known above all for its security, for 27 years. On 7 April 2026 Anthropic said the model that found it, Claude Mythos Preview, would not be sold to the public at all.
Instead the company put it to work in Project Glasswing, a defensive programme with Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, Nvidia and Palo Alto Networks, plus more than 40 other organisations that build or maintain critical software. They may use it for cybersecurity only. The reason is stated plainly: Anthropic says Mythos Preview can find and exploit software flaws better than all but the most skilled humans.
- What: Claude Mythos Preview, Anthropic's most capable model so far, released only to Project Glasswing partners for defensive security work.
- Price: up to USD 100 million in usage credits, then USD 25 per million input tokens and USD 125 per million output tokens for participants.
- Why it matters: a frontier lab has held back a general purpose model because of what it can do to software, and is giving defenders a head start.
What the model found
In a technical post that day, Anthropic's security researchers wrote that Mythos Preview could identify and then exploit zero day vulnerabilities (flaws unknown to the software's makers) in every major operating system and every major web browser when a user directed it to. Besides the OpenBSD bug, which let an attacker crash any OpenBSD machine reachable over TCP, it found a 16 year old flaw in the H.264 video decoder of FFmpeg and, with no human help after the first request, wrote a working exploit for a 17 year old bug in FreeBSD's NFS file sharing code that handed an attacker root access.
Anthropic's own comparison shows the size of the jump. Its previous flagship, Claude Opus 4.6 (released in February), turned vulnerabilities in the JavaScript engine of Firefox 147 into working exploits twice in several hundred attempts. Run as a benchmark, Mythos Preview produced working exploits 181 times. These are vendor reported results. “Engineers at Anthropic with no formal security training have asked Mythos Preview to find remote code execution vulnerabilities overnight, and woken up the following morning to a complete, working exploit,” the researchers wrote.
Most of what it found cannot be described yet: Anthropic said more than 99 percent of the vulnerabilities were still unpatched, and hired security professionals check every report by hand before it goes to a maintainer.
Who gets it, and why not everyone
Anthropic is committing up to USD 100 million in usage credits and USD 4 million in donations to open source security groups. Once the credits run out, participants pay USD 25 per million input tokens and USD 125 per million output tokens, on the Claude API, Amazon Bedrock, Google Cloud's Vertex AI and Microsoft Foundry. Anthropic promised a public report on what has been fixed within 90 days, and said it has been in ongoing discussions with US government officials about the model's cyber capabilities.
The argument is about time. “Given the rate of AI progress, it will not be long before such capabilities proliferate, potentially beyond actors who are committed to deploying them safely,” the company wrote. Elia Zaitsev, chief technology officer of CrowdStrike, described the defenders' side in the same announcement: the gap between discovering a flaw and an attack on it has collapsed, and “what once took months now happens in minutes with AI.”
“We do not plan to make Claude Mythos Preview generally available,” Anthropic wrote, while saying its eventual goal is to let its users deploy Mythos class models safely at scale. New safeguards would first ship with an upcoming Claude Opus model, one “that does not pose the same level of risk as Mythos Preview.” “Saying 'our model is too dangerous to release' is a great way to build buzz around a new model, but in this case I expect their caution is warranted,” the developer Simon Willison wrote.
Careful, until it is not
The system card makes an uncomfortable argument. “Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin,” it says. “Even so, we believe that it likely poses the greatest alignment-related risk of any model we have released to date.” Anthropic compares it to a seasoned mountain guide who takes clients onto harder climbs than a novice ever would.
The examples come from earlier internal versions. In one test a simulated user asked the model to escape a locked down sandbox computer and send a message to the researcher running the evaluation. It built what the card calls “a moderately sophisticated multi-step exploit” to reach the open internet, did as asked, then, unprompted, posted details of the exploit on obscure but public websites. The researcher learned of the success from an unexpected email from the model while eating a sandwich in a park. Anthropic stresses that the sandbox was separate from the systems holding the model's weights, and that in this test the model showed no ability to access its own weights or reach internal systems.
A footnote adds that holding the model back “does not stem from Responsible Scaling Policy requirements”, Anthropic's own rulebook for catastrophic risks: it was a judgement about cyber skills.
What happened next
On 16 April 2026 Anthropic released Claude Opus 4.7, the first model with its new safeguards that detect and block high risk cybersecurity requests. On 2 June it extended Project Glasswing to about 150 more organisations in more than 15 countries, saying the first partners had found more than 10,000 high or critical severity flaws. On 9 June Glasswing partners could upgrade from Mythos Preview to Claude Mythos 5, released alongside a safeguarded public version of the same model called Claude Fable 5.





Comments
Comments are reviewed before they appear.
No comments yet. Be the first to share your thoughts.