Claude Mythos Preview: an AI too capable at hacking to release widely
Anthropic announces Claude Mythos Preview, describing its ability to find and exploit unknown software vulnerabilities as a "watershed moment for security". Instead of a public release, it launches Project Glasswing to use the model to secure critical software. Anthropic says over 99% of the vulnerabilities found had not yet been patched.
What happened
On 7 April Anthropic published two pieces at once: an announcement of Project Glasswing, and a technical post from its Frontier Red Team (the group that tests models for dangerous abilities) on the cybersecurity skills of a new model, Claude Mythos Preview.
Anthropic describes Mythos Preview as a general-purpose model that is strong across the board but “strikingly capable” at computer security. When directed by a user, it can find and exploit previously unknown flaws in every major operating system and every major web browser. The company called this “a watershed moment for security”.
So Anthropic made an unusual call: no public release. The launch partners were Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks. More than 40 other organisations that maintain critical software also got access.
Background: why the hacking skills jumped
According to Wikipedia, the name “Claude Mythos” had already become public on 26 March through leaked drafts of blog posts.
Anthropic stresses that it did not train the model specifically to hack. The abilities emerged as a by-product of general improvements in coding, reasoning and working on its own. The same progress that makes it better at fixing flaws makes it better at exploiting them.
The change came fast. Only a month earlier, Anthropic had written that its then-current Claude Opus 4.6 was “far better at identifying and fixing vulnerabilities than at exploiting them”. One test shows the gap:
- Opus 4.6: Firefox flaws turned into working exploits
- 2 times
- Mythos Preview: working exploits on the same test
- 181 times
- Targets fully hijacked across ~1,000 open-source projects
- 10
- Share of flaws found that were still unpatched
- Over 99%
Source: Anthropic Frontier Red Team, 7 Apr 2026. Opus 4.6's two successes came from several hundred attempts; on the open-source test, Opus 4.6 and Sonnet 4.6 each reached tier 3 (of five) only once.
What it actually did
The examples Anthropic published are the few flaws already fixed and safe to discuss. It says nearly all of them were found with no human help after a one-line prompt asking the model to find a security vulnerability:
- A 27-year-old bug in OpenBSD, an operating system known for its security and used to run firewalls. It let an attacker crash any OpenBSD machine remotely just by connecting to it. A thousand search runs on OpenBSD cost under $20,000 in total; the run that found this bug cost under $50.
- A 16-year-old bug in FFmpeg, the video library used by almost every video service. Automated testing tools had run the faulty line of code five million times without catching it.
- Remote takeover of FreeBSD servers. Fully on its own, the model found and exploited a 17-year-old flaw (CVE-2026-4747) that let anyone on the internet take complete control of a server running the NFS file-sharing service.
- Chaining flaws together. On Linux it combined two, three and sometimes four flaws to go from ordinary user to full control of the machine. In a browser it wrote an exploit chaining four flaws that escaped two layers of “sandbox” isolation.

One detail stands out: Anthropic engineers with no formal security training asked the model to look for remote-takeover flaws overnight, and woke up to a complete, working exploit.
On a public test, CyberGym, which checks whether a model can reproduce known vulnerabilities, Anthropic reports a clear lead:
Both scores measured and published by Anthropic. Source: Anthropic, “Project Glasswing”, 7 Apr 2026
Its coding ability is also well ahead. On SWE-bench Pro, which asks models to fix real issues in open-source projects, Anthropic reports 77.8% for Mythos Preview against 53.4% for Opus 4.6.
How Project Glasswing works
Anthropic’s reasoning is that similar abilities will spread sooner or later, so defenders should get them first.
- Who: 12 launch partners (Anthropic included) and more than 40 further organisations, scanning their own systems and open-source software.
- Money: up to $100 million in usage credits, plus $4 million in donations to open-source security groups ($2.5 million to Alpha-Omega and OpenSSF through the Linux Foundation, $1.5 million to the Apache Software Foundation).
- Price afterwards: $25 per million input tokens and $125 per million output tokens.
- Disclosure: professional reviewers check each finding before it goes to the software’s maintainers. For flaws that can’t be disclosed yet, Anthropic published cryptographic fingerprints (SHA-3 hashes) and promised to reveal the details after fixes, so its claims can be checked later.
It also published how often human reviewers agreed with the model: in 89% of 198 manually reviewed reports they agreed exactly on severity, and in 98% they were within one level.
Reactions and what came next
- Partners. Cisco, AWS, Microsoft, CrowdStrike, the Linux Foundation, Google and Palo Alto Networks all gave supporting statements. Lee Klarich, Palo Alto Networks’ chief product and technology officer, said it also “signals a dangerous shift” in which attackers will soon find flaws and build exploits faster than ever. According to SecurityWeek, the company said Mythos did the equivalent of a year’s penetration testing in under three weeks.
- Firefox. SecurityWeek reported that Mozilla used an early version of Mythos Preview to find 271 flaws, fixed in Firefox 150. Only three of the public vulnerability IDs (CVEs) in Mozilla’s advisory credit Claude, which SecurityWeek took to mean most were lower-severity issues. Firefox CTO Bobby Holley said they had not seen any bug that “couldn’t have been found by an elite human researcher”.
- Unauthorised access. According to Bloomberg, as relayed by GovInfoSecurity, a Discord group got access to Mythos through the environment of an Anthropic third-party contractor. Anthropic said it was investigating but had no evidence of use beyond that third party’s IT environment.
- Rivals. GovInfoSecurity reported that OpenAI released GPT-5.4-Cyber days later, saying it wanted to make it “as widely available as possible” and would rely on identity checks to prevent misuse.
- Wider access. According to Wikipedia, on 2 June Anthropic expanded cybersecurity access to Mythos to 150 organisations in more than 15 countries.
Then: Claude Opus 4.7, released nine days later, was officially described as less capable than Mythos Preview. On 9 June Anthropic launched Claude Fable 5, a Mythos-class model with safeguards for public use, and moved Glasswing partners to Mythos 5. Three days later a US export order forced both offline (see US export order forces Anthropic to suspend its newest models). In September Mythos 5.1 arrived alongside Fable 5.1. On 6 October Anthropic said existing Glasswing members would move to the “Specialized Access” tier of its expanded Cyber Verification Program.
Things to keep in mind
- Most of the evidence can’t be checked yet. With over 99% of flaws unpatched, Anthropic can only discuss about 1% of cases. It admits this makes some claims hard to verify, which is why it published the cryptographic fingerprints.
- The scores come from Anthropic. CyberGym and SWE-bench results were measured by the company. It also notes that some SWE-bench problems may have been memorised by the model, and says its lead holds when those are excluded.
- Counts are not severity. The 271 Firefox flaws are a case in point: a large number, but only a few earned public IDs.
- “Defenders first” has limits. Only a small set of large organisations got early access, and the Discord episode shows that a limited release can itself be bypassed.
What it means for ordinary users
You can’t use Mythos Preview; Anthropic said it does not plan to make it generally available. But it affects software everyone relies on: operating systems, browsers, video players, encrypted connections.
The practical advice is simple: keep your system and apps updated. The flaws Anthropic and its partners find are fixed through ordinary security updates. Anthropic itself warns that the transition “may be tumultuous”, because similar abilities will eventually reach attackers. In the long run, it expects tools like this to help defenders more, fixing bugs before new code ever ships.