A new AI model from Anthropic is changing how security teams find and prove software vulnerabilities. It is raising hard questions about what happens when the same technology falls into the wrong hands.
Cloudflare has published findings from its participation in Project Glasswing, Anthropic’s controlled research program, revealing that Mythos Preview, a security-focused large language model, can automatically build working proof-of-concept (PoC) exploits by chaining multiple low-severity bugs into a single, high-severity attack chain.
Unlike general-purpose AI models that can identify a bug and describe it in text, Mythos Preview goes several steps further. According to Cloudflare’s research, the model demonstrates two distinct capabilities that set it apart from previous tools.
The first is exploit chain construction, the ability to take several small attack primitives, such as a use-after-free memory bug combined with return-oriented programming (ROP) chains, and reason about how to combine them into a working, multi-stage exploit.
The second is automated proof generation, the model writes exploit code, compiles it in a sandboxed scratch environment, runs it, reads the output, adjusts its hypothesis if the exploit fails, and repeats the loop until it succeeds.
This closes a gap that has long frustrated security teams: a suspected vulnerability without a working PoC is, in practice, speculation. Mythos Preview eliminates that ambiguity on its own.
Cloudflare’s researchers noted that AI vulnerability scanning still produces a high volume of low-confidence findings, particularly from codebases written in memory-unsafe languages like C and C++. Models have a built-in bias to report findings even when none exist, hedging results with phrases like “possibly” or “could in theory.”
Mythos Preview improves on this considerably. When findings arrive paired with a working PoC, triaging becomes faster and more reliable, with fewer speculative results, clearer reproduction steps, and less wasted analyst time.
Cloudflare did not simply point Mythos Preview at a repository and ask for bugs. Instead, the team built a multi-stage vulnerability discovery harness that runs agents in parallel across focused, narrow scopes.
Key stages include a Recon phase that maps the codebase architecture, a Hunt phase that runs roughly 50 concurrent agents targeting specific attack classes, and a Validate stage where an independent agent attempts to disprove each finding.
A final Trace stage determines whether attacker-controlled input can actually reach the discovered bug from outside the system, converting a theoretical flaw into a confirmed, reachable vulnerability.
The Project Glasswing version of Mythos Preview operated without the standard safety restrictions found in publicly available models.
Cloudflare observed that the model produced its own “organic” refusals, but they were inconsistent. The same research task, framed slightly differently, produced opposite results across separate runs.
Cloudflare explicitly warned that these emergent guardrails are not sufficient as a standalone safety boundary. Any future public release of a capable cyber AI model must include additional, externally enforced safeguards beyond what the model generates on its own.
Cloudflare noted that some security teams are now operating under two-hour SLAs from CVE disclosure to production patch. This timeline is only sustainable if the architecture around the vulnerability is hardened. Patching faster without fixing regression testing pipelines introduces new bugs.
The same capabilities that make Mythos Preview valuable for defenders will accelerate offensive operations against every internet-facing application. The research was conducted in a controlled environment, and all vulnerabilities discovered were triaged and remediated under Cloudflare’s formal vulnerability management process.
Follow us on Google News, LinkedIn, and X to Get Instant Updates and Set GBH as a Preferred Source in Google.
Toronto, Canada, October 8th, 2026, CyberNewswire Insignary Launches Clarity AIR: Closing the Blind Spot Between…
A threat actor published a malicious version of the tensorlake npm package on October 8,…
A proof-of-concept (PoC) exploit has been released for CVE-2026-102489, a critical vulnerability in Zammad that…
A critical vulnerability in LMCache allows unauthenticated attackers to execute arbitrary code against reachable multi-process…
16 malicious Firefox extensions that impersonate cryptocurrency wallets to intercept recovery phrases and private keys…
Exposed directories on five servers have revealed an operational DarkSword/Coruna exploitation platform built to compromise…