Anthropic’s Mythos Preview and Project Glasswing - A Cybersecurity Reckoning

Ronni Holmvig Strøm · 2026-04-08

Anthropic has chosen restraint over rapid release. On April 7, 2026, the company officially confirmed the existence of Claude Mythos Preview, internally described as a “step change” and the most capable model it has ever trained, but deliberately withheld it from general availability.

Anthropic has chosen restraint over rapid release. On April 7, 2026, the company officially confirmed the existence of Claude Mythos Preview, internally described as a “step change” and the most capable model it has ever trained, but deliberately withheld it from general availability.

Instead, Mythos Preview powers Project Glasswing, a defensive cybersecurity initiative that gives a select coalition of organizations early, controlled access to the model’s formidable agentic capabilities. The move follows a high-profile data leak in late March that first brought Mythos to public attention and underscores the growing tension between frontier capability gains and real-world risk.

The Leak That Forced the Conversation

The story began in late March 2026 with not one, but two significant lapses at Anthropic.

A misconfigured CMS exposed nearly 3,000 internal files, including draft blog posts and announcements that explicitly described an unreleased model called Claude Mythos (internally codenamed “Capybara”). The documents portrayed Mythos as significantly more powerful than Claude Opus 4.6, with dramatic advances in reasoning, coding, and especially cybersecurity tasks.

Days later, on March 31, a routine npm release of Claude Code version 2.1.88 accidentally included a 59.8 MB JavaScript source map file. This single artifact exposed over 512,000 lines of unobfuscated TypeScript across nearly 1,900 files, revealing internal architecture, tool-use patterns, safety mechanisms, and hints at unreleased features.

Security researcher Chaofan Shou (@Fried_rice) quickly identified the exposure. Within hours, the full codebase was mirrored, dissected, and partially recreated by the community (including popular Python ports dubbed “Claw Code”). Anthropic issued DMCA notices and later clarified that the incident was a “packaging error caused by human mistake,” while launching an internal review of its CI/CD processes.

The dual leaks created an unusual situation: the world learned about Mythos’s existence and potential dangers before Anthropic was ready to announce it. Internal documents warned that models at this capability level could “exploit vulnerabilities in ways that far outpace the efforts of defenders,” potentially triggering a new wave of AI-driven cyberattacks.

Mythos Preview: Capabilities That Demand Caution

According to Anthropic’s statements and leaked materials, Mythos Preview represents a genuine leap in agentic performance. It excels at long-horizon planning, recursive self-correction, sophisticated code reasoning, and autonomous tool use. In internal red-teaming, the model independently discovered thousands of high-severity vulnerabilities (including zero-days) across every major operating system, web browser, and critical software stack. It then generated realistic exploit chains, often with minimal human guidance.

These capabilities make Mythos Preview a double-edged sword. In defensive hands, it can rapidly audit and harden infrastructure. In offensive hands, it could dramatically lower the barrier to sophisticated cyberattacks, enabling agents that operate faster and more creatively than human teams.

Anthropic’s decision not to release the model broadly echoes its Responsible Scaling Policy but goes further. It treats Mythos as the first post-GPT-2-level system too dangerous for open deployment without safeguards.