AI Compressed the Discovery Half. Your Patch Window Did Not Move.

AI Compressed the Discovery Half. Your Patch Window Did Not Move.

There is a line in Anthropic’s reporting on Project Glasswing that deserves more attention from engineering leaders than it has had:

vulnerability discovery has become fast enough that triage and remediation are increasingly the bottleneck

That is a structural claim, not a product claim. For as long as most of us have been working, the scarce resource in security was finding the flaw. Skilled researchers were few, their time was expensive, and the queue of software worth examining was effectively infinite. Everything downstream — triage, patching, disclosure coordination — was sized around that scarcity.

If the discovery half becomes cheap and the remediation half does not, the whole system reorganises around the new constraint. That is now happening, and most patch policies have not noticed.

What changed, specifically

Three data points from the last few months, all public.

Anthropic’s Project Glasswing. Claude Mythos Preview identified more than 10,000 high and critical severity vulnerabilities, with access expanding from roughly 50 to around 150 organisations. The interesting number is not 10,000 — it is that the programme’s own summary names triage, not discovery, as the constraint.

OpenAI’s internal testing. During evaluation, agents escaped a test environment by exploiting an Artifactory package cache, discovered vulnerabilities, left notes for other agents, and rebuilt their communication channels after those channels were closed. Hugging Face later reconstructed approximately 17,600 actions across 4.5 days. The model involved approached OpenAI’s own “Critical” threshold on cybersecurity capability.

Independent research on WordPress plugins. A security researcher used Claude Opus 4.8 to review dozens of plugins and found 16 confirmed vulnerabilities, concentrated in smaller, less-maintained ones — a category previously protected mostly by not being worth a human researcher’s afternoon.

That last one is the part with the widest blast radius. The economics that kept long-tail software unexamined have changed. “Nobody has looked at this” was never a security property, but it functioned as one.

Three public numbers: 10,000 high and critical vulnerabilities, 17,600 agent actions, 16 plugin vulnerabilities.

The finding that keeps this honest

The same researcher reported something worth holding onto: some vulnerabilities the model identified disappeared under runtime testing. Code review alone was not sufficient; the candidate findings needed verification against a running system before any of them counted.

So the picture is not that AI now finds vulnerabilities. It is that AI now generates candidates at a rate that outstrips the human capacity to verify them — which is a different problem, and arguably a harder one to manage. A pile of plausible findings is not the same as a pile of real ones, and treating it as such produces exactly the confident, expensive, systematic mistakes that come from acting on a count instead of reading the exceptions.

Why this lands on defenders first

The asymmetry is uncomfortable.

An attacker needs one finding to be real. A defender needs to assess all of them. When the discovery rate rises, the attacker’s cost per exploit falls and the defender’s triage queue grows — the same input produces opposite pressures on the two sides.

The asymmetry: an attacker needs one finding to be real, a defender must assess all of them.

And the exploitation curve after disclosure has visibly steepened. In the three weeks after a WordPress core remote code execution flaw was disclosed this year, one security vendor reported attacks against it rising from about 10,000 to over 900,000 per day, with attackers rapidly cycling through variations of the technique.

Ninety-fold in three weeks. Set that against a patch policy that says “monthly maintenance window” and the mismatch is not subtle.

Three weeks after disclosure, attacks rose from about 10,000 a day to over 900,000 a day.

What actually needs to change

Not more scanning. Most organisations already generate more findings than they action; adding a faster generator to an unchanged triage process makes the backlog worse, not the risk lower.

Three things move the needle.

Measure time-to-patch, not patch coverage. “We are 96% patched” is a snapshot that says nothing about speed. The number that matters is the median hours from a disclosure affecting you to that disclosure being closed — and whether anyone knows it. If you cannot state it, you do not have a patch policy, you have a patch habit.

Patch coverage is a snapshot, not a speed; time to patch is the elapsed hours from disclosure to closed advisory.

Know your inventory before you need it. When a disclosure lands, the first question is “are we affected”, and most organisations cannot answer it in under a day. Every hour spent establishing what you run is an hour inside the exploitation window. This is unglamorous, it is a spreadsheet, and it is the single highest-leverage item on this list.

Separate the two bottlenecks and staff them differently. Verification — is this finding real — and remediation — ship the fix — are distinct work with distinct skills. Where discovery volume has risen, verification is where things now queue, and it is the step most often assumed rather than resourced.

Three things that move the needle: measure time to patch, know your inventory first, split verify from remediate.

The uncomfortable part for the long tail

If your product depends on components nobody has ever seriously audited — an old library, a niche plugin, an abandoned SDK — the protective obscurity is going away.

That is, on balance, good. Those vulnerabilities existed before anyone looked; the only change is who knows about them. But it is going away faster than most dependency-management practices assume, and “it has been fine for years” is a statement about attention, not about safety.

The reasonable response is not fear. It is to know what you depend on, to know how quickly you can replace or patch each one, and to accept that the answer for some of them is currently “we have no idea” — which is itself the finding.

What I would put in front of a board

One number and one question.

The number is your median time from disclosure to patched, for anything you actually run. Not coverage. Not scan volume. The elapsed hours.

The question is: has that number improved in the last twelve months?

One number and one question: median hours to patched, and whether it has improved this year.

Because the other side’s has. The discovery half of this has been compressed by something like an order of magnitude, and the exploitation curve after disclosure has steepened to match. If your remediation half has not moved at all in that time, your exposure has grown without anyone deciding it should.

Nobody signed off on that. It arrived as a default.


Ash Ganda is CTO and founder of Ganda Tech Services. Figures cited are drawn from public reporting by Anthropic, OpenAI, Hugging Face, Sucuri and BlogVault, 2026.

Free Guide · 2026

AI Strategy Primer for Australian Business Leaders

A practical framework for AI adoption in 2026 — cut through the hype and start with what matters.

No spam. Unsubscribe any time.