An AI Found the Bug in July. Attackers Only Showed Up in October. The 80 Days in Between Are the Story.
The version that made the headlines
Attackers started probing an open source file server the day after researchers published the way in.
That is the story that travelled on Friday. Anthropic's restricted bug-hunting model, Mythos, found a critical flaw in Rejetto HTTP File Server. Horizon3, one of the firms with access to the model through Anthropic's Project Glasswing program, published the exploit chain on September 30. By Thursday evening, according to VulnCheck, its monitoring network was watching an actor in China probe vulnerable hosts in the US and Japan. By Friday there were four more hits, from two addresses in the same US subnet that VulnCheck researcher Patrick Garrity said appear to be a proxy.
The patch shipped on July 13. It landed 80 days before the first detection, and 79 days before the write-up that led straight to it. Rejetto had already fixed the flaw, and credited the find to Horizon3's researcher working, in the release notes' phrasing, "in collaboration with Claude and Anthropic Research." Anyone running a current version was never the target.
The interesting question is not how fast a model can find a bug. It is what the gap between those two dates says about the industry's ability to act on one.
What the model actually did
The flaw, tracked as CVE-2026-61500, is not a memory-safety bug or a parser overflow. It is a cryptography problem, and the way it was found is the part that will travel.
Rejetto HFS signs its session cookies with a key that comes from JavaScript's Math.random(). When the COOKIE_SIGN_KEYS environment variable is unset, that key is a value the server builds at startup. The design assumes the generator behind it is unpredictable. In V8, the JavaScript engine under Node.js, Math.random() runs on xorshift128+, which is fast, non-cryptographic, and fully reversible. Feed it enough observations and you can walk the internal state backwards to the value the process produced when it booted.
The observations were free. During the first step of the login handshake, HFS stored a raw Math.random() value in the session cookie. Those cookies are signed but not encrypted, so an unauthenticated client can read the number straight out of its own Set-Cookie header.
Mythos found both halves and, more to the point, treated them as one chain. In the model's own output, quoted in Horizon3's write-up, it noted that "the client receives the exact 52-bit double in its own Set-Cookie," that consecutive doubles allow state recovery, and that the recovered state "can be stepped backwards to the outputs consumed by randomId(30) at process start."
Then it wrote the exploit. Horizon3 says the model implemented an SMT solver using Microsoft's publicly available Z3, assembled a working proof of concept, and demonstrated arbitrary command execution. The write-up lists 12 samples of the leaking endpoint as enough to recover the state; the model's own analysis put the requirement at three to five consecutive doubles.
"Publicly-tooled" is the model's own description of the attack, and that is worth sitting with. Horizon3's researchers wrote that they "do not recall ever seeing an SMT solver being used to attack a cryptographic flaw like this in a real application." Not because it was novel mathematics, but because it was the kind of work that used to be too slow and too specialised to justify. Hanley's point is economic, not technical: the model removed both of the reasons a human team would have walked away.
The severity is not in dispute. NVD carries VulnCheck's scores at CVSS 3.1 9.8 and CVSS 4.0 9.3, weakness CWE-338, weak pseudo-random number generator. Versions 3.0.0 through 3.2.0 are affected. The fix, committed on July 10 and released as 3.2.1 three days later, changed two lines of real code: randomId(30) became randomBytes(32).toString('base64url'), and the session id became randomUUID(). Seven lines across three files, one of the smallest security patches you will ever see attached to a 9.8.
The eighty-day head start nobody used
Lay the dates out and the "exploited within 24 hours" framing starts to wobble.
The fix commit landed July 10. Version 3.2.1 shipped July 13, the same day the VulnCheck advisory and the NVD record were published. Nothing happened for ten weeks. A third-party proof-of-concept repository appeared on GitHub on September 26, per that repository's creation date. Horizon3 published its write-up on September 30, and Rejetto shipped 3.3.4 the same day with further security fixes. VulnCheck's sensors picked up the first activity the following evening.
So what happened within a day was a day after the exploit details went public and five days after working code appeared in a public repository. No source says which of those, if either, the attacker used. The timing is a pattern, not proof. But the pattern points at publicity and available tooling, not at the moment a model identified the bug back in the spring.
Two more details keep this in proportion. VulnCheck's detections come from its own canary network. They are sightings of scanning and attempted exploitation against hosts configured to record it, not confirmed compromises of production servers, and no reporting reviewed for this piece describes one. While VulnCheck's advisory lists CVE-2026-61500 in its own known-exploited catalogue, CISA's Known Exploited Vulnerabilities feed, checked at catalog version 2026.10.02, still does not. CISA does list HFS, but for the older 2.x flaw, CVE-2024-23692, added in July 2024 and long since associated with ransomware crews. Two catalogues, two answers, and the difference matters if your patch queue is driven by one of them.
A tracker, 286 CVEs, and two exploitation events
Garrity has been counting. His tracker holds the CVEs credited to Anthropic's research team or to Project Glasswing partners, and cross-references them against exploitation data to get at what he calls the real "danger factor."
On September 21, the count was 225, and exactly one of them, a SQL injection in Ghost CMS tracked as CVE-2026-26980, had been exploited in the wild. By October 2 the count was 286. The number of exploited entries was two.
That is under one percent, against a historical baseline Garrity puts at "just under one percent to two percent of vulnerabilities that get weaponized and used in the wild." His read, given to The Register, is blunt: "There's a big difference between finding vulnerabilities and whether they're actually useful to and will be used by threat actors." And what Anthropic is disclosing, he said, "isn't resulting in different outcomes from a threat perspective than a random selection of other vulnerabilities would."
There is a serious counterargument, and Cloudflare made it in detail in May after pointing Mythos at more than fifty of its own repositories. Its finding was that the model does something previous frontier models did not: it takes a batch of low-severity primitives that would sit invisible in a backlog and stitches them into one working exploit, then writes and runs the proof itself, adjusting when the code does not behave. That is a real change in what a defensive team can see. Cloudflare also flagged the failure mode that comes with it: "ask a model to find bugs, and it will find them, whether the code has any or not." That is why its harness runs fifty or so narrow hunters in parallel and puts an independent adversarial reviewer between every finding and the triage queue.
The number that should worry a security budget sits downstream of all of that. 1Password's researchers took 6,080 patches generated by two frontier models and found that 26 percent fully resolved the vulnerability; roughly 54 percent either failed to fix it, introduced a new one, or did both. Veracode, across more than 100 models and 80 coding tasks, measured a 56 percent average security pass rate for AI-generated code. Discovery is cheap now. Verification is not.
Garrity's summary is the least exciting and most useful line in the whole story: "The bar for vulnerability discovery is much lower with AI, but the real gap lies downstream in coordination, triage, remediation, and patch deployment, which is still largely people-intensive work."
What actually protects you is not the patch you hear about first
The HFS case is a clean illustration, because the chain has a soft last step.
Every path to code execution runs through the admin interface. The configuration option admin_net is documented as a netmask for the addresses allowed to reach the admin panel, and the default permits any address. Set it, and forging an administrator cookie buys an attacker very little. Leave COOKIE_SIGN_KEYS unset and you are trusting a generator that was never designed for security. Skip the upgrade and you are running code that has been publicly documented as exploitable for five days by the time anyone came looking.
Cloudflare put the strategic version of this better than any vendor summary: "Patching faster does not change the shape of the pipeline that produces the patch. If regression testing takes a day, you cannot get to a two-hour SLA without skipping it." The teams now operating under a two-hour CVE-to-patch target are, in its assessment, about to spend a lot of money learning that speed alone does not help. The defensive work that holds is architectural: keeping the bug unreachable, keeping one flaw from becoming a foothold, and being able to push a fix everywhere it runs at once.
This is the same problem hardware teams have been living with for years, and it is why we keep writing about it. You cannot patch what you do not know you are running. A file server embedded in a device that shipped three years ago is a line item in a bill of materials that nobody has audited since the vendor selection was signed. The software you did not write, on the board you did design, running on the network you do not fully control. That is a supply-chain visibility problem before it is a security problem. At DMC, we work with hardware companies on exactly that class of question: sourcing strategy, component and vendor visibility, and production planning when the thing you are shipping depends on parts and code you did not build. If your product roadmap needs the firmware and component inventory to survive the same audit as the bill of materials, let's talk.