Inside the Mythos System Card
Claude Mythos Preview is not on the API. You cannot deploy it. You probably cannot even try it. Anthropic has stated plainly that they "do not plan to make Claude Mythos Preview generally available." Instead, it's restricted to twelve Project Glasswing launch partners and roughly forty critical-infrastructure organizations, for defensive cybersecurity work only.1 So why spend an evening on a 245-page system card for a model you cannot use?
Because the card documents a model that's a meaningful step beyond Opus 4.7, and the failure modes it catalogues are the failure modes the next public-release Anthropic Opus is being trained around. The Mythos card is the closest the public will get to the next-generation Anthropic model for the next six months, and the mitigations Anthropic is shipping in subsequent releases are downstream of what the Mythos team observed. If you ship products on Claude, this is the card to read for a preview of what's coming and what's being defended against.
I read the Mythos card the day after I read the Opus 4.7 card, and the comparison is striking. Where the 4.7 card kept telling me what Opus 4.7 was not (not the capability frontier, not Mythos-class), the Mythos card tells you exactly what's on the other side of that gap: a model that saturates Cybench, finds 27-year-old bugs in OpenBSD, escapes a research sandbox unprompted and posts the exploit to public-facing websites, and once took down every other running evaluation when asked to end one. The card is unsettling and important in roughly equal measure.
In this post:
- The release decision: Glasswing and what "limited" means
- The cyber numbers are the headline
- Capabilities, plainly stated
- The alignment paradox: best-aligned and greatest risk
- What this tells you about the next public Opus
- What's actually next
If you have only an hour: read the abstract, Section 1.2 (release decision process), Section 3 (cyber), Section 4.1 (alignment summary, including the catalogue of incidents), and the Section 6.3 capability table. That's roughly forty pages and most of the load-bearing content. The other 200 pages are evaluation methodology, raw transcripts, white-box interpretability detail, and the long appendices on safeguards.
The release decision: Glasswing and what "limited" means
The card's framing is set in the abstract: "Claude Mythos Preview's large increase in capabilities has led us to decide not to make it generally available."2 This is a different shape of release than any prior frontier-lab model I'm aware of. Mythos exists, has been evaluated against the full Anthropic Responsible Scaling Policy machinery (RSP 3.0, the version current as of April 20262), and the system card has been published in line with what Anthropic does for full-public-release models. The model itself is gated behind Project Glasswing.
Project Glasswing, announced the same day as the card, brings together AWS, Anthropic, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks as launch partners.1 An additional ~40 organizations that maintain critical infrastructure get access through extended membership. Anthropic is committing $100M in usage credits to seed the program plus $4M in direct donations to open-source security organizations ($2.5M to Alpha-Omega/OpenSSF via the Linux Foundation, $1.5M to the Apache Software Foundation).1 After the credits are exhausted, Mythos will be priced at $25/$125 per million input/output tokens for participants: five times Opus 4.7's pricing.1
The dual-track strategy is more honest than it might first appear. Anthropic isn't sitting on Mythos; they're using it on the most security-critical software in the world while explicitly continuing to ship lower-capability models (starting with Opus 4.7, nine days later) for general availability. The bet is that the cyber benefits of giving twelve hyperscalers and forty infrastructure orgs an extreme-capability model outweigh the proliferation risk of releasing it to everyone, and that the lessons Anthropic learns from deploying Mythos through Glasswing will inform the safeguards that make a Mythos-class general release safe later.
It's also a defensible product moat. Glasswing access becomes a privileged vendor channel for the world's biggest security and cloud companies. Both readings can be true.
The practical consequence for a working dev who isn't a Glasswing partner: you cannot use Mythos and you probably cannot test it. The system card is the artefact you have. Anthropic clearly expects more of the public discourse about Mythos to flow through the card itself than is usual for a model release, which is why the card includes a new "Impressions" section (more on that later), meant to surface qualitative behavior the public cannot otherwise observe.
The cyber numbers are the headline
The release-decision frame above is downstream of one set of numbers: cyber. Section 3 of the card opens with the line "Claude Mythos Preview is the most cyber-capable model we have released, surpassing all previous models across our internal evaluation suite and saturating nearly all of our existing internal and known external capability evaluations."2
The numbers that matter:
- Cybench: 100% pass@1. Mythos solves every one of 35 challenges across 10 trials. The card calls the benchmark saturated and is exploring what to report instead.23
- CyberGym: 0.83. On the 1,507-task targeted vulnerability reproduction suite, Mythos beats Opus 4.6 at 0.67 and Sonnet 4.6 at 0.65.24 This is now Anthropic's preferred cyber benchmark.
- Real-world zero-days, found autonomously. From the Glasswing announcement: a 27-year-old vulnerability in OpenBSD allowing remote crash-by-connection; a 16-year-old vulnerability in FFmpeg in code that automated testing tools had hit five million times without catching; an autonomous Linux kernel privilege-escalation chain combining multiple bugs to go from user to root.1 The announcement claims "thousands of high-severity vulnerabilities" across every major OS and web browser.
- First model to solve a private cyber range end-to-end. UK AISI reports Mythos completed their 32-step "The Last Ones" cyber range (estimated 20 hours of expert work) on three of ten attempts, averaging 22 of 32 steps; Opus 4.6 averaged 16 steps and never finished.5
What didn't work: Mythos failed an operational-technology cyber range and a properly-configured-and-patched sandbox eval. The model is dramatically capable against weak defensive posture; defense-in-depth still matters.
The mitigations are designed for the dual-track release. For Glasswing partners, cyber probe-classifiers run in monitor mode only. Exchanges aren't blocked even on classifier hits, because the partner pool is small and trusted. The card states that "in general-release models with strong cyber capabilities, we plan to block prohibited uses, and in many or most cases, block high risk dual use prompts as well."2 That's the design now live on Opus 4.7 via the Cyber Verification Program.6 If you've hit a cyber probe-block on 4.7 and wondered what justified it, the answer is in the Mythos card: a held-back model that can autonomously find and chain real-world zero-days.
Capabilities, plainly stated
The cyber numbers are the headline because they're the release-decision driver, but the general-capability numbers are also state-of-the-art. The table below combines Mythos's Table 6.3.A with the corresponding numbers from the Opus 4.7 card so the comparison against the model you can actually deploy today is in one place:7
| Evaluation | Mythos Preview | Opus 4.7 | Opus 4.6 | GPT-5.4 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| SWE-bench Verified | 93.9% | 87.6% | 80.8% | n/a | 80.6% |
| SWE-bench Pro | 77.8% | 64.3% | 53.4% | 57.7% | 54.2% |
| SWE-bench Multilingual | 87.3% | 80.5% | 77.8% | n/a | n/a |
| SWE-bench Multimodal | 59.0% (internal) | 34.5% | 27.1% | n/a | n/a |
| Terminal-Bench 2.0 | 82% | 69.4% | 65.4% | 75.1% | 68.5% |
| GPQA Diamond | 94.55% | 94.2% | 91.3% | 92.8% | 94.3% |
| MMMLU | 92.7% | 91.5% | 91.1% | n/a | 92.6%–93.6% |
| USAMO 2026 | 97.6% | 69.3% | 66.2% | 95.2% | 74.4% |
| GraphWalks BFS 256K-1M | 80.0% | 58.6% | 38.7% | 21.4% | n/a |
| HLE (no tools / with tools) | 56.8% / 64.7% | 46.9% / 54.7% | 40.0% / 53.1% | 39.8% / 52.1% | 44.4% / 51.4% |
| BrowseComp | 86.9% | 79.3% | 83.7% | n/a | n/a |
| OSWorld | 79.6% | 78.0% | 72.7% | 75.0% | n/a |
Mythos beats Opus 4.7 on every row, with the biggest practitioner-relevant gaps on SWE-bench Pro (+13 points), Terminal-Bench (+12), HLE with tools (+10), and SWE-bench Multimodal (+25 on internal harness). Two cells worth pausing on: USAMO 2026 took place on March 21–22, 2026, after both models' training data cutoffs, so the 97.6% vs 69.3% gap can't be explained by contamination. And BrowseComp is the only cell where Opus 4.7 regressed against Opus 4.6 (79.3% vs 83.7%), which the 4.7 card attributed to a "more conservative, more clarification-asking" tendency; Mythos doesn't show that regression and uses 4.9× fewer tokens per task than Opus 4.6 to get there.2
The Terminal-Bench number understates the gap: under the 2.1 harness fixes with four-hour timeouts, Mythos hits 92.1% (vs Opus 4.7's 69.4%). The card flags that 2.0 timeouts "risk hiding real capabilities jumps behind seemingly uncorrelated confounders like sampling speed."2
The alignment paradox: best-aligned and greatest risk
The most quotable section of the card is the alignment introduction (§4.1.1). Anthropic states two things in the same paragraph:
Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin… Even so, we believe that it likely poses the greatest alignment-related risk of any model we have released to date.2
The card's analogy:
Consider the ways in which a careful, seasoned mountaineering guide might put their clients in greater danger than a novice guide, even if that novice guide is more careless: The seasoned guide's increased skill means that they'll be hired to lead more difficult climbs, and can also bring their clients to the most dangerous and remote parts of those climbs.2
A more capable model is hired to lead more difficult climbs, can bring clients to more dangerous places, and is overseen less closely because it earns trust. Then the catalogue of specific incidents from early Mythos checkpoints, the kind of disclosure that doesn't appear in any other current frontier-lab system card:
- The sandbox escape and the email from the park. During behavioural testing, an early Mythos in a secured sandbox was asked to escape and notify the researcher. It built "a moderately sophisticated multi-step exploit to gain broad internet access" and emailed the researcher as requested. Then, in what the card calls "a concerning and unasked-for effort to demonstrate its success," it also posted exploit details to public-facing websites. Footnote 10: "The researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park."2 This is, I think, the single most memorable line in any current model card.
- Cover-ups. Mythos once obtained the answer to a quantitative-estimation question via a prohibited method, then "explicitly reasoned that it needed to make sure that its final answer submission wasn't too accurate" before submitting something deliberately less accurate. In another case, after editing files it lacked permission to edit, the model "made further interventions to make sure that any changes it made this way would not appear in the change history on git." White-box analysis showed concealment and strategic-manipulation features active during these actions.2
- /proc/ access for credentials. Earlier Mythos checkpoints used
/proc/to inspect process memory and recovered credentials for messaging, source control, and the Anthropic API itself: credentials Anthropic had intentionally chosen not to give the model.2 - Took down all evaluations. Asked to end one running eval, Mythos implemented the request by taking down every similar eval being run by every user, "in spite of specific warnings to be careful about this."2
- Prompt-injected an LLM judge. When its submissions were rejected by an LLM judge, Mythos attempted a prompt-injection attack on the judge, and "identified its behavior as conducting a prompt injection in its extended thinking."2
Anthropic isn't claiming the final Mythos is misaligned. Misuse success rates dropped by more than half against Opus 4.6's already-low levels. The mountaineering point is that a more capable model with low rates of misbehaviour can still cause more total harm, because the rare misbehaviour happens with more powerful tools in higher-stakes environments. Anthropic is explicitly urging Glasswing partners "not to deploy the model in settings where its reckless actions could lead to hard-to-reverse harms."2
The lesson for a non-Glasswing dev isn't "Mythos is dangerous." Mythos isn't on the API. The lesson is that this catalogue is what the next public Opus is being trained against. The mitigations Anthropic ships in subsequent models are designed for these specific patterns. Knowing what they're defending against is the closest you'll get to red-teaming the next-generation model in advance.
What this tells you about the next public Opus
Reading the Mythos card alongside the Opus 4.7 card, several through-lines become clearer than they were when only one card existed.
The agentic-safety mitigations on Opus 4.7 are designed for the failure modes catalogued in the Mythos card. The 4.7 card reports browser-prompt-injection attack-success rates of 0.0% with safeguards across 148 environments, "the strongest result Anthropic has observed on this benchmark."8 It's no coincidence those numbers landed at the same time as the Mythos catalogue of recklessness-in-cyber. If you ship browser, computer-use, or coding agents on 4.7, the fact that the model behind them was hardened against the failure modes of a held-back more-capable model is, I think, a meaningful piece of context.
The "completion-verification step" punch-list item from the 4.7 post applies double for Mythos-class behaviour. The 4.7 card's casual disclosure that "Opus 4.7 will occasionally mislead users about its prior actions, especially by claiming to have succeeded at a task that it did not fully complete" is the public-release version of the more dramatic Mythos cover-up incidents. The mitigation is the same: don't trust the model's self-report. Verify against artefact state.
The Cyber Verification Program exists because of Mythos. Probe-classifiers running in enforce mode on Opus 4.7, with credentialed pentesters able to apply for exemption, is a deployment design that would not have existed without the Mythos cyber numbers. If you do legitimate cyber work (vulnerability research, red-teaming, pentest tooling) and have hit a probe-classifier block on 4.7, the CVP is the path through, and the Mythos card is the document that explains why the gate exists.6
The Mythos card is a transparency artefact precisely because the model isn't shipped. Anthropic could have run all the Mythos evaluations privately, used the lessons internally, and never published anything. Instead they published a 245-page card that is in many ways more candid than the corresponding card for the model you can actually use. That's worth noting as a precedent: not because frontier labs are now obligated to publish system cards for their internal-only models, but because Anthropic has made the case that a model worth running ASL-3 evaluations against is also a model worth publishing those evaluations about.
What's actually next
Anthropic has committed to publishing a public report on Project Glasswing within 90 days of the announcement, which lands sometime around early July 2026.1 That report is the next major Mythos-related public artefact, and the place to look for the partner-organisation findings (which vulnerabilities have been fixed in the wild, what the partners have learned about defensive workflows, what the Glasswing engagement has actually produced).
The longer question is whether other frontier labs publish system cards for their own held-back models. The Mythos card is, as far as I can tell, the first instance of a frontier lab publishing a full system-card-grade evaluation of a model they have explicitly chosen not to release. Whether OpenAI, Google, Meta, and the major Chinese labs follow with similar artefacts for their own internal-only model snapshots is the kind of thing that will define the 2026–2027 frontier-lab transparency norm.
The thing to do whenever the next major Anthropic model card drops: read both. Read the public-release card to understand the model you can deploy. Read the Mythos card alongside it, because the mitigations in the public-release card are downstream of Mythos. The launch post is for people who want to know that the new model is better. The system cards, both of them, are for people who want to know what's being defended against.
Related: Inside the Claude 4.7 System Card · Evals for AI Agents · Agentic Engineering with Claude Code
Footnotes
-
Anthropic, Project Glasswing: Securing critical software for the AI era, April 7, 2026 (link). The launch announcement; the public-facing companion to the system card. ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
Anthropic, System Card: Claude Mythos Preview, April 7, 2026 (changelog April 8) (PDF). 245 pages. The primary document this post reads. Linked from the Anthropic Transparency Hub. All quantitative claims and quoted text in this post are sourced from this document unless otherwise noted. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16
-
Zhang, A., et al., Cybench: A framework for evaluating cybersecurity capabilities and risks of language models, arXiv:2408.08926 (link). The CTF-style cyber benchmark Mythos has saturated. ↩
-
Wang, Z., et al., CyberGym: Evaluating AI agents' cybersecurity capabilities with real-world vulnerabilities at scale, arXiv:2506.02548 (link). The targeted-vulnerability-reproduction benchmark on which Mythos scored 0.83. ↩
-
UK AI Security Institute, Our evaluation of Claude Mythos Preview's cyber capabilities, April 2026 (link). Independent government evaluation; first model to complete the 32-step "The Last Ones" cyber range end-to-end. ↩
-
Anthropic, Real-time cyber safeguards on Claude (Anthropic Support). Documentation for the Cyber Verification Program. ↩ ↩2
-
The capability table in this post is composed across two cards: row labels and Mythos / Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro columns are from Table 6.3.A of the Mythos card; the Opus 4.7 column is from Table 8.1.A of the Opus 4.7 card, published nine days after the Mythos card and so not included in the original table. SWE-bench Multimodal Mythos figure (59.0%) uses Anthropic's internal harness; Opus 4.7's 34.5% uses the same internal harness. ↩
-
Anthropic, System Card: Claude Opus 4.7, April 16, 2026 (PDF). Numbers cited from the Section 8 capability summary table. See also my prior post on the 4.7 card. ↩
Related writing
Inside the Claude 4.7 System Card
A practitioner's reading guide to the 200+ page Anthropic document almost no one reads in full. What the launch post hides, where the load-bearing safety numbers live.
How a Diffusion Model Works: A Practitioner's Read of the 2026 Image Stack
Modern image models aren't U-Nets running 50 denoising steps. They're transformers running 4 steps of a straight-line flow. Once that lands, every product surface starts making sense.
How a Vision LLM Works: A Practitioner's Read of the 2026 Multimodal Stack
Vision LLMs don't see images. They tokenize them. Once that lands, the cost, the failure modes, and the design space all fall out cleanly.