GLOBAL THREAT GROUP, INSIGHTS, RESEARCH | October 1, 2026

LLM-Assisted Vulnerability Research: Finding Real Bugs with Code-Reasoning Models

Introduction

Point a code-reasoning model at a codebase and it reads through it far faster than I can, then hands back a long list of things that look like bugs. Generating that list is the easy part now, but most of it is noise. Almost every candidate falls apart the first time I try to trigger it against a real build, and the handful that survive still have to be proven before they count as findings. Getting from that list to a proven bug is the slow, deliberate part of the job.

In this post I walk through the workflow that I use to close that gap, and two real bugs it has produced. Both are now fixed in Angular and rated HIGH severity:

  • CVE-2026-68945: a server-side render cache that could hand one visitor’s response to the next visitor, skipping the backend’s authorization check entirely.
  • CVE-2026-69151: a translation-file path that let lower-trust text compile into a live JavaScript event handler, turning a localization bundle into a script-injection vector.

These two are the CVEs I can name today; more from the same process are still working through coordinated disclosure.

Neither came from asking a model to “find bugs.” Both came from a disciplined loop: work out the rules a project’s code must never break, from its patch history and its own design, then keep the model’s attention on one small piece of code at a time, drive it through four focused audit passes, and believe nothing until I have proven it myself in a real runtime. The model does the reading and the guessing. I own the threat model, the runtime, and the final call.

That division of labor, not the automation, is what holds the whole thing together. The rest of this post builds it up, starting with why the obvious shortcut (just asking a model to find bugs) doesn’t work.

Why “Find All the Vulnerabilities” Fails

Ask a large language model (LLM) to “find security bugs in this repository” and it will answer at length, confidently – and almost none of it will be worth acting on.

What comes back is a list of plausible-looking weak points with no sense of which ones an attacker can reach, textbook bug patterns matched onto code that never runs, and not one claim that the model has tried to disprove. Separating the real from the decorative lands right back on you, now scaled up to the length of the model’s output.

A bare request gives the model no notion of reachability, no trust boundary, and no idea which bug classes a maintainer would actually accept. That means it falls back on the only thing it can do: it matches shapes that look like vulnerabilities and reports them with even confidence, whether they sit on a hot path or in dead code. The model writes clean, persuasive analysis, but nothing in the request forces it to be true.

Constrain the request to a single testable claim, anchored in the target’s own history, and the same model behaves differently:

Recent accepted vulnerabilities in this component involved validation/use mismatches. A patch introduced cumulative length validation in one path. Identify sibling paths that still validate entries independently while dispatching batches to a backend that consumes cumulative state. Demonstrate or refute this in the product runtime, with a negative control and a fixed-version comparison.

Now the model has a reachability target, a boundary that defines impact, and a claim that can be caught when it’s wrong. This still generates candidates, but each one arrives with the terms of its own refutation attached.

The difference between the two prompts is not the model or its size, but the question. And a good question starts from a threat model.

Threat Modeling: Rules the Code Must Never Break

Before I send any code to a model, I build a compact threat model from the project’s own security history: its advisories, its commits, its patch diffs. The point isn’t only to re-hunt old bugs. It’s to learn how the software expects its data to behave, so I can then look for places where that assumption quietly fails.

The output of that work is a short list of invariants. An invariant is a one-sentence rule the code assumes is always true and that a bug would break, for example, “the byte count validated at the perimeter must equal the amount the backend actually consumes.” That is the whole idea. An old validation-patch advisory becomes exactly that kind of rule. An advisory about a use-after-free (a bug where code keeps using an object after it’s been freed) becomes “no object may keep using itself after calling a callback that can destroy it.”

While an invariant is the rule the code must hold, once a bug shows that rule being broken, a variant is any other location where the same rule is broken in the same way.

Prior reported bugs make the hunt richer, because each one hands you an invariant the code already got wrong and a family of variants to chase, but they are not required. Where there are none, you derive the invariants from the code’s own contracts, a cache key that must stay unique or a trust boundary that must hold. This is where a model and an expert outrun what either does alone: the model maps the sinks and their connections across the whole surface far faster than a person could, while the researcher decides which invariants matter and where they break. The threat model is the method, and the pairing is what makes it powerful.

This is also why “find a use-after-free” is a useless prompt, while “determine whether this callback can free the parent object before the next line reads a member of it” is a good one. The first is a bug category, the second a specific claim a model and I can both test.

Nothing about this is specific to open-source. The method is the same whether the target is a public repository, a client’s proprietary codebase, or a closed black box, and what changes is only how much of the history is handed to you. Open-source projects give it away in public commits and advisories, which is why the examples in this post are open-source. On a closed target, you assemble the same ground truth yourself from decompiled interfaces, architecture documents, past vendor disclosures, and dynamic traces, and that is the slower, harder version of the identical work. A model can’t invent a threat model from nothing, and it works only from the evidence you give it.

Figure 1 – From a patch to its forgotten sibling.

Context Discipline: The Token Budget

Compressing history into an invariant is wasted effort if the model loses that invariant halfway through the job, which is exactly what happens when you ignore how its context works. A model doesn’t have your repository open the way an editor does. On every turn it sees one block of text – your prompt, the conversation so far, the code you pasted, the tool output that came back – and nothing else. That block is the context window, measured in tokens (roughly ¾ of a word each), and it is finite. Almost every failure I hit traces back to how that budget gets spent. The most relevant reasons for context failure are these three:

Lossy compaction. When a conversation nears the window limit, the tooling summarizes older turns to make room. Summaries keep the story (“found a validation gap”) and throw away the specifics upon which a finding depends: the exact line number, the build flag, the crash offset. After a couple of rounds of this, a session can sound completely confident while no longer holding a single piece of usable evidence. That’s the most dangerous failure mode, because nothing looks broken.

Degraded recall in the middle. A model recalls what sits at the start and end of a long context far better than what is buried in the middle. The effect has a name, “lost in the middle,” and it has outlasted several generations of far larger context windows. A 2025 context rot study found every one of eighteen frontier models still degrading as their input grew. That is the problem here, because the invariant, the one thing that has to hold from the first turn to the last, is exactly what sinks into that middle while the model keeps reasoning as if it still had it.

Cache invalidation. Reasoning over a large context is only affordable because the tooling caches the processed prefix. Churn the conversation and that cache goes cold, and every turn starts paying the full cost of reading everything again.

The fix is not a bigger window. That just moves the limit without solving the recall decay. The fix is to keep the working slice small enough that the invariant, the code under test, and the oracle (the check that separates a vulnerable build from a fixed one) all fit in the high-recall region at once. Everything else goes to disk.

Before I pause a session, I write a short HANDOFF.md at the repository root – the pinned commit, the current artifact, the live hypotheses, the dead ends already ruled out – so the next session resumes without re-deriving state. If that note runs longer than a page, the slice was too big and I split it.

Figure 2 – Keeping the slice in high-recall context.

The Four Audit Passes

With a threat model in hand and a list of candidate code paths to check, I run the model through four audit passes. Each one has its own input and its own output.

  1. Patch-diff variant farming. A patch shows the exact line the maintainers changed, and the code around it shows the neighborhood they left alone. I ask the model to find sibling paths that still contain the old, pre-patch logic. This is my highest-yield pass by a wide margin.
  2. Parser and serializer differentials. When a system implements the same concept more than once – two URL parsers, two header processors, two cache-key generators – those implementations can disagree, and the disagreement is a bug. I list the pairs and ask for inputs where they produce different results, aiming at request smuggling, server-side request forgery (SSRF), or cache confusion.
  3. Symmetry and dual checks. For every input-validation check, I look for its output twin: the cache layer, serializer, or redirect handler that’s supposed to enforce the same rule on the way out. A missing twin is a strong candidate.
  4. Cross-framework transfer. Once a bug pattern works in one target, I test sibling frameworks for the same underlying assumption. If it replicates, that’s a class-level design flaw, not one project’s slip.

Only the first pass looks backward at patches. The other three go looking for attack surface nobody has catalogued yet.

Those passes hand back a list of candidate bugs. I cut each surviving candidate down to a slice of under a hundred lines and send it through a dual-model gate, where one session is prompted to prove the bug exploitable and another, handed the same slice with no prior context, is asked to prove it safe. This is not a vote, because two runs of the same model share the same blind spots. I treat it as a disagreement detector instead.

When both sessions call the bug real, it still owes me a runtime proof; when both call it safe, I drop it; and a split, which flags the non-obvious edge cases, is the most useful of the three. The gate only filters, though, and whatever comes out of it I verify by hand, reading the code and reproducing the behavior myself, because two models agreeing is not proof.

Figure 3 – Four audit passes into a dual-model gate.

What a Verified Finding Looks Like

Before any of this touches a runtime, each candidate gets a spec card, four lines that fix exactly what is being claimed and what would prove it. Here’s one, from a sanitized browser-graphics example – a validation/use mismatch, where the check validates each draw command on its own but the backend runs the whole batch and walks the write cursor past the allowed buffer:

Invariant:  ValidateBatch must validate the cumulative byte count
            consumed by SubmitBatch, not the largest single entry.
Attacker:   A web page controls the draw list and the bound range.
Oracle:     The second draw must fail if draw1 + draw2 exceeds the
            bound range. A vulnerable build writes past the range;
            a fixed build returns an API error.
Stop:       Do not claim High unless the write crosses an allocation,
            ownership, process, origin, or authority boundary.

The oracle is the line that does the work, the exact condition that tells a vulnerable build from a fixed one. Here the oracle is checked across three runs, and the contrast between them is the proof:

separate_draw_control:    second_draw_error = INVALID_OPERATION
                          bytes_after_bound_changed = 0

batched_draw_vulnerable:  batch_error = NO_ERROR
                          bytes_after_bound_changed = 64

batched_draw_fixed:       batch_error = INVALID_OPERATION
                          bytes_after_bound_changed = 0

The first run is the negative control, where single checks reject the bad bounds. The second shows the bypass under batched execution, and the third confirms the fix. One log line would prove nothing, but the contrast across all three is what a triager can trust.

This example shows one more thing: because the writes stayed inside memory the calling page already owned, this was a constrained primitive (a real but limited capability), not full memory corruption. That distinction is the difference between a Medium and a High, and I’ll come back to it. The productive next question is never “can I make this proof-of-concept louder?” It’s “where else does this invariant fail?”

Figure 4 – A validation/use mismatch: control vs affected.

Two Real Findings in Angular

Here’s the method producing real, accepted bugs, both reported, confirmed, and patched upstream. One came from a forward-looking differential, the other from a patch variant.

CVE-2026-68945: two requests, one cache key

Advisory: GHSA-jhpw-976m-542j. Rated HIGH (CVSS v4.0) by the assigning authority.

The assumption, from history. Angular’s server-side rendering (SSR) builds the page on the server before sending it to the browser, and it caches the HTTP responses it fetches so the browser can reuse them during hydration instead of re-fetching. Any cache like that rests on one rule.

The invariant. Two HTTP requests that could receive security-distinct responses must never share a cache key.

The hypothesis (Pass 2, serializer differential). If the code builds cache keys by turning request parameters into a string, is there more than one way to write “the same” string? There was. The key generator joined repeated parameters with commas, so two genuinely different requests collapsed onto one key:

GET /api?role=user&role=admin   →  key "role=user,admin"
GET /api?role=user,admin        →  key "role=user,admin"

Figure 5 – Two requests, one cache key (CVE-2026-68945).

The proof. The oracle was a backend hit-counter, and two distinct requests produced exactly one backend call. The second visitor was served the first visitor’s cached response straight from the transfer cache, never reaching the backend’s authorization check. Angular fixed the key serialization in 20.3.27, 21.2.19, and 22.0.2 (PR #68571).

CVE-2026-69151: a translation file that could inject a script

Advisory: GHSA-jj27-h5hq-8×99. Rated HIGH (CVSS v4.0) by the assigning authority.

The assumption, from history. Angular’s template compiler treats internationalization (i18n) files as lower-trust: translators are supposed to supply text, not code. That’s a trust boundary, and trust boundaries are exactly what threat modeling tells you to test.

The invariant. Lower-trust translation content must never compile into an event handler.

The hypothesis. An earlier fix had already closed one hole where translated content bypassed the normal attribute checks. So, what other attributes still slip through the same translation path? Event handlers did. The compiler accepted i18n-onerror and other i18n-on* attributes, which meant a tampered translation file could drop live JavaScript into a static event handler.

The proof. The demonstration is a compile-time differential rather than a memory trace. Bind onerror on an element directly and the compiler rejects it. Mark the same attribute i18n-onerror and it passes the i18n collection path and compiles into the localized build intact. A translation file supplying that attribute’s value then controls a live handler, which is exactly what the advisory documents, a lower-trust translation file replacing a benign onerror=”void 0″ with arbitrary JavaScript that runs in the application’s origin. Angular fixed it in 20.3.27, 21.2.19, and 22.0.1 (PRs #68821 and #69306).

Figure 6 – i18n bypass: onerror vs i18n-onerror (CVE-2026-69151).

The bug wasn’t a sanitization slip but a broken contract between the compiler and untrusted input, and that input is realistic: localization bundles routinely come from outside translators or a shared repository. As with the cache-key bug, the fix seeds the next cycle, because patching one attribute is an invitation to check every sibling attribute immediately.

Candidate Classification

The fastest way to sink a valid report is to overstate it, so I rate a finding by one thing only: the security boundary its runtime proof actually crosses.

That’s why the graphics example earlier was a Medium, not a High. The writes stayed inside memory the page already owned, so nothing crossed an allocation, process, origin, or authority boundary. Both Angular bugs cross a real one: CVE-2026-68945 serves a response across users, and CVE-2026-69151 runs script in the application’s origin. A scanner’s score is only an interpretation of it, but the breached boundary is a fact, so document the boundary and let the rating follow.

Conclusion

The cost of generating a hypothesis, writing a negative control, and checking a patch’s variants has collapsed. What’s scarce now is the discipline to throw the bad ones away quickly and prove the good ones are real.

In practice, that comes down to a few habits: threat-model before you prompt, so every question is anchored to one rule; never report a finding without a negative control and a fixed-version comparison; and run every candidate past a fresh, independent model session whose only job is to disprove it. That session is only a filter, though, and the final sign-off on any finding is always a person’s.

This doesn’t automatically favor defenders. The same tools are a download away for anyone, and patch-diff variant farming rewards whoever moves first after a commit ships. This means the takeaway for anyone shipping code is direct: run these passes on your own patches before you release, because someone outside will run them after you ship.

INSIGHTS | September 9, 2026

Secure Software Development Lifecycle Practices

IBM’s 2026 Cost of a Data Breach Report found the global average breach cost reached USD $4.99 million, a record high driven by AI-powered attacks up 56% year over year.¹ Supply chain attacks compound the exposure: ReversingLabs research confirmed software supply chain attacks grew 1,300% over three years.² The Log4j vulnerability alone generated over 10 million attack attempts per hour at peak exploitation.³ Organizations that treat security as a final-stage gate accumulate deferred risk with every release.

Secure software development lifecycle practices embed security controls at every phase of the software development process, reducing remediation costs and organizational risk.

“The most effective secure software development lifecycle practices address security before software is deployed, not after vulnerabilities are discovered.”
– IOActive Security Team

Secure SDLC Best Practices: Controls, Outputs, and Metrics

Best PracticeWhat It DoesKey OutputSuccess MetricsRisk If Skipped
Security Requirements MappingTranslates compliance and business requirements into testable acceptance criteria before development beginsDocumented security requirements and acceptance criteria% requirements with security acceptance criteria; compliance gap rateMissing controls, regulatory gaps, late-stage rework
Threat Modeling (STRIDE/DREAD)Identifies attack paths in planned architecture using structured frameworks before a line of code is writtenThreat register, attack surface map# threats identified per review; % mitigated before development startsArchitectural flaws reach production with costly remediation
Static Code Analysis (SAST)Scans source code for insecure patterns at commit or build timeClean code reports, flagged findings# high/critical findings per 1,000 lines of code; mean time to remediateInjection flaws, hardcoded secrets, insecure patterns shipped
Dynamic Application Security Testing (DAST)Tests running applications against real attack patterns pre-releaseValidated vulnerability findings# runtime vulnerabilities found; % of API endpoints testedRuntime and authentication flaws missed before release
Software Composition Analysis (SCA) + SBOMInventories open-source dependencies and tracks known CVEs; SBOM documents every component in the buildCVE inventory, Software Bill of Materials% dependencies on a supported version; # libraries no longer actively maintainedVulnerable third-party libraries deployed; supply chain exposure
Artifact Signing and Configuration HardeningSigns build artifacts cryptographically and audits deployment configurations to prevent tamperingSigned artifacts, hardened deployment configurations% of artifacts signed; # misconfigurations resolved at deploy gateSupply chain tampering, misconfigured production environments
Vulnerability Management and SLA EnforcementPrioritizes and remediates vulnerabilities by CVSS score, production reachability, and CISA KEV statusRemediated CVE backlog, SLA compliance metricsMean time to remediate (MTTR) by severity tier; % SLA complianceUnpatched known-exploited vulnerabilities accumulate and expand breach risk
Security Governance and Program MaturityEstablishes named ownership, enforceable security gates, and maturity benchmarks aligned to NIST SP 800-218 or OWASP SAMMMaturity score, named control owners, audit evidence# controls with named owners; % security gates enforced in CI/CD; maturity level progressionControls exist on paper only; programs consistently underperform without governance infrastructure

Gaps at any single phase propagate forward. A vulnerability introduced at requirements and missed through design, development, and testing arrives in production fully formed and significantly more expensive to remediate.

Run Threat Modeling Before Writing Code

The highest-leverage security activity happens before a single line of code is written. Threat modeling, using frameworks such as STRIDE or DREAD, gives teams a structured method for identifying attack paths in planned architecture and building in controls before architectural decisions become costly to reverse. A research-fueled Secure Development Lifecycle engagement brings genuine adversarial thinking to this process, applying real-world offensive experience to surface the systemic risks that automated tools routinely overlook.

Threat Modeling Method Comparison

MethodWhat It DoesWho Runs ItWhat It ProvidesEffort LevelBest Fit
STRIDECategorizes threats by class: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of PrivilegeSecurity architects, senior developersComprehensive attack class mapping across trust boundariesMedium — 1 to 2 days per system reviewApplication architecture and trust boundary analysis
DREADScores threats on Damage, Reproducibility, Exploitability, Affected users, and Discoverability to produce a ranked risk listSecurity engineers, risk teamsPrioritized threat list by exploitability and business impactLow to medium — hours per threat registerRisk-ranking identified threats for remediation sequencing
Attack TreesDecomposes the path to a specific high-value target into branches of possible attack methodsPenetration testers, security researchersTargeted attack path analysis for critical or high-value componentsHigh — days to weeks for complex targetsComponent-level analysis of specific high-value targets

Method selection depends on the question the team needs to answer. STRIDE maps attack classes broadly across a system; DREAD narrows findings to prioritized risk; Attack Trees decompose the path to a specific high-value target. Most mature programs apply more than one method, selecting based on the architecture under review and the risk questions at hand.

Integrate Code Analysis and Supply Chain Controls Into Every Pipeline Stage

Static analysis (SAST) catches insecure code patterns during development, while dynamic testing (DAST) validates a running application against real attacker behavior. Supply chain exposure requires its own layer of controls. Sonatype’s research found open-source supply chain attacks growing at 742% annually.⁴ A dedicated Supply Chain Integrity assessment applies attacker-mindset analysis from silicon through the application layer. A 2025 third-party assessment of Microsoft’s Signing Transparency service confirmed strong implementation security and identified defense-in-depth improvements.⁵ Core controls include SBOM management, SCA in CI/CD pipelines, dependency pinning, and artifact signing.

Code Analysis and Supply Chain Controls

ControlWhat It DoesWho Runs ItIntegration PointRisk Addressed
SAST in CI/CDScans source code for insecure patterns automatically at commit or mergeDeveloper pipelines, security championsGit pre-commit hook or CI pipelineInsecure patterns caught before code reaches the main branch
DAST pre-releaseSends attack traffic to a running application to surface runtime flawsQA teams, security engineersStaging environment or pre-release gateRuntime and authentication flaws not visible to static analysis
SCA in build pipelineScans open-source dependencies against CVE databases with each buildDevOps, security teamsBuild pipeline or dependency resolution stepKnown CVEs in third-party libraries before they reach production
SBOM + artifact signingGenerates a complete bill of materials and cryptographically signs build outputs for every releaseDevSecOps, release engineeringRelease pipeline or deployment gateSupply chain tampering and untracked third-party components

Layering all four controls closes distinct gaps at different pipeline stages. SAST and DAST address first-party code at different points in the lifecycle. SCA and SBOM management address third-party and supply chain risk at the build and release stages. Each layer reinforces the others: SAST catches what code review misses, DAST catches what SAST cannot see at runtime, and SCA flags risks that originate entirely outside the team’s codebase.

Prioritize Vulnerabilities With Contextualized SLA Frameworks

Datadog’s 2025 DevSecOps research found that the median software dependency is 215 days behind its latest major version, with one in two services running libraries no longer actively maintained.⁶ Effective vulnerability management requires contextualized prioritization based on CVSS score and production reachability, SLA-based ownership, and regression testing to confirm fixes hold. Listings on the CISA Known Exploited Vulnerabilities catalog indicate confirmed real-world exploitation and warrant accelerated remediation timelines regardless of the base CVSS score.

Vulnerability Remediation SLA Framework

Priority TierCriteriaTarget SLA
CriticalCVSS 9.0+, internet-exposed, CISA KEV listed24–72 hours
HighCVSS 7.0–8.9, production reachable7–14 days
MediumCVSS 4.0–6.9, limited exposure30–60 days
LowBelow 4.0 or non-production90 days

SLAs paired with named ownership and enforced timelines convert a priority framework into a functioning remediation program. Without that operational infrastructure, vulnerability backlogs expand faster than teams can close them.

Establish Named Ownership and Track Security Program Maturity

Governance gaps are the most common point of failure in secure SDLC programs. Controls require named ownership and enforceable gates; programs without both consistently underperform regardless of tooling investment. The NIST Secure Software Development Framework (SP 800-218) and OWASP SAMM provide structured maturity models for building and measuring program progress.⁷ Red Team and Purple Team exercises close the feedback loop, testing whether development-phase controls hold under real adversarial simulation and informing the next cycle of program improvement.

Assess Your Current Maturity Level and Close the Gaps Systematically

Maturity LevelCharacteristicsNext Steps to Advance
Ad HocReactive, no formal process, incident-driven patchingAssign one control owner per business unit; document current state versus target state; implement a vulnerability tracking system
DefinedBasic security gates, SAST in CI/CD, vulnerability trackingAdd threat modeling at the design phase; adopt SLA-based remediation; integrate SCA into the build pipeline
ManagedThreat modeling, SCA, SLA-driven remediation, metrics reportingAutomate artifact signing; run quarterly adversarial simulations; establish an SBOM management program
OptimizedAutomated supply chain controls, adversarial simulation, SBOM managementIntegrate continuous Red Team feedback into the SDLC; benchmark the program annually against NIST SP 800-218 or OWASP SAMM

Most organizations enter at Ad Hoc or Defined. Identifying the gaps between the current state and the next level, then addressing them incrementally, produces durable outcomes. Each level builds on the controls established in the previous one, making systematic progression the most reliable path to program improvement.

Build Security Into Your Development Process With IOActive

Secure software development lifecycle practices address vulnerabilities at the phase where remediation is most efficient: before the code ships. Organizations that build security into design, development, and deployment produce software that holds up under real-world adversarial pressure and satisfies the compliance obligations regulators and insurers continue to tighten. IOActive’s 25+ years of research-driven assessments, landmark vulnerability disclosures, and straight-talk approach give clients an honest picture of where their program creates risk.

Sources

  1. IBM Security. Cost of a Data Breach Report 2026. https://www.ibm.com/reports/data-breach
  2. ReversingLabs. The State of Software Supply Chain Security 2024. https://www.reversinglabs.com/sscs-report-2024
  3. IOActive. Supply Chain Integrity. https://www.ioactive.com/service/supply-chain-integrity/
  4. Sonatype. State of the Software Supply Chain, 10-Year Look. 2024. https://www.sonatype.com/state-of-the-software-supply-chain/2024/10-year-look
  5. IOActive. Security Assessment of Microsoft’s Signing Transparency. 2025. https://www.ioactive.com/wp-content/uploads/2025/10/Microsoft-Signing-Transparency-Service-Security-Assessment-IOActive-Public-Facing-Report.pdf
  6. Datadog. State of DevSecOps 2025. April 2025. https://www.datadoghq.com/state-of-devsecops-2025/
  7. NIST. Secure Software Development Framework (SSDF) SP 800-218. https://csrc.nist.gov/projects/ssdf
INSIGHTS | August 18, 2026

The Five Eyes AI Shift in Cyber Risk Statement: What Industry Leaders Need to Know Now

Key Takeaways

  • On 22 June 2026, the leaders of the Five Eyes cyber security agencies issued a joint statement, The AI Shift in Cyber Risk: Why Leaders Must Act Now, warning that frontier AI is transforming cyber risk on a timeline measured in months, not years.
  • The statement is signed by the heads of the National Cyber Security Centre (NCSC, UK), Cybersecurity and Infrastructure Security Agency (CISA, US), National Security Agency (NSA, US), Australian Signals Directorate (ASD, Australia), Communications Security Establishment (CSE, Canada), and Government Communications Security Bureau (GCSB, New Zealand) — an unusually unified articulation of urgency from the Five Eyes partnership.
  • Cyber risk is explicitly reframed as a core business risk and board-level responsibility, not a technical issue to be delegated downward.
  • Five practical actions are prescribed: reduce attack surface, accelerate patching, address legacy systems, strengthen identity and access controls, and prepare for incidents before they happen.
  • The statement lands days after the NCSC’s own CEO disclosed that 75% of attacks on UK critical infrastructure over the past year are linked to hostile states, and weeks after NCSC guidance warning of an incoming “vulnerability patch wave” driven by AI-accelerated exploitation.
  • Organisations that treat this as a compliance afterthought will be out of step with where their regulators, and their adversaries, are already heading.

A Statement for a Narrowing Window

Joint statements from the Five Eyes cyber agencies are not issued lightly, and this one is notable for its tone as much as its content. Published on 22 June 2026, The AI Shift in Cyber Risk [1] is signed jointly by Stephanie Crowe (ASD), Rajiv Gupta (CSE), Catriona Robinson (GCSB), Richard Horne (NCSC), David Imbordino (NSA), and Nick Andersen (CISA). The framing is unambiguous: frontier AI models are expected to exceed current industry expectations, and the timeline for that shift is not years, it is months.

This is not an isolated warning. Five days before the statement was published, NCSC CEO Dr Richard Horne told the Royal United Services Institute’s Annual Security Lecture that the NCSC had managed more than 200 cyber incidents affecting the UK’s critical national infrastructure in the year to May 2026, with around 75% believed linked to hostile state actors including Russia, China, and Iran [2]. Horne went further, arguing that cyber security should no longer be framed as a risk to be tolerated within appetite, but as an ongoing contest with capable adversaries. He also pointed to an NCSC assessment that by 2028, AI-enabled capabilities will likely be used to exploit known vulnerabilities in legacy technology at scale across UK critical infrastructure.

That assessment builds directly on guidance the NCSC published in May 2026, warning organisations to prepare for a “vulnerability patch wave”: a forced correction in which AI-accelerated exploitation surfaces decades of accumulated technical debt across commercial, open source, and proprietary software simultaneously [3]. The Five Eyes statement should be read as the international consolidation of that warning, not a standalone development.

What Does the Statement Set Out?

The statement is short and deliberately free of new technical detail. Its purpose is to compress urgency into a small number of leadership-level actions, structured around four overarching asks and five practical steps [1].

Leaders are urged to:

  • Understand and assess risk, readiness, and accountability. Boards need a clear, current picture of organisational exposure, not a static risk register reviewed annually.
  • Prioritise foundational cyber security practices and controls. Sophistication in tooling does not substitute for getting the fundamentals right.
  • Empower cyber leaders with authority and resources. Cyber leadership requires the mandate to act, not just the responsibility to report.
  • Stay actively engaged as threats and guidance evolve. Static governance models cannot keep pace with a threat landscape that is itself accelerating.

Two structural points distinguish this statement from prior Five Eyes guidance. First, it reframes AI not solely as an adversary capability but as a defensive obligation: organisations are explicitly told to use AI deliberately to strengthen defence, not merely to improve efficiency. Second, it sets an expectation that breaches are not preventable in absolute terms; preparedness is reframed as the capability to contain incidents quickly before they escalate into operational and financial crises [1].

Who is Affected by the Statement?

Boards and Executive Leadership

The statement is addressed primarily upward, not downward. It states plainly that cyber resilience is not an IT issue, it is central to operational continuity and market trust, and that it is not enough to have controls; leaders must be confident those controls will perform during a real incident [1]. This places direct accountability on boards and executives to verify resilience, not simply to fund it.

CISOs and Security Leadership

For security leaders, the statement is a mandate to escalate. The call to empower cyber leaders with authority and resources [1] gives CISOs a clear external reference point when seeking budget, headcount, or the organisational authority to challenge unsafe trade-offs that have previously been accepted in the name of operational convenience.

Operators of Critical National Infrastructure

For CNI operators, the statement reinforces direction already set domestically. The NCSC’s own intervention five days prior, disclosing that three-quarters of attacks on UK CNI are state-linked [2], makes clear that the threat described in the Five Eyes statement is not a future scenario for this sector. It is the current operating environment.

Vendors and Technology Providers

The statement explicitly calls on leaders across industry, including vendors, to act now [1]. Combined with the NCSC’s separate warning that a wave of vulnerability disclosures is approaching across commercial, open source, and proprietary software [3], vendors should expect both faster exploitation of existing flaws and intensifying customer expectations around patch velocity and secure-by-design practice.

What are the Key Challenges Organisations Will Face?

Compressed Exploitation Timelines

The central technical claim underpinning the statement is that AI is shrinking the window between vulnerability discovery and exploitation [1]. Patch cadences and change-management processes built around weeks or months of lead time were not designed for this. Organisations with manual patching processes, particularly across operational technology environments with long update cycles, face a widening gap between the speed of the threat and the speed of their own response.

Your Adversary Already Has an AI Upgrade

Even if your organization hasn’t touched AI, your attackers have. The patch window that once gave defenders breathing room is effectively gone — Mandiant’s time-to-exploit tracking shows exploits now landing on or before the day a CVE goes public. Exploitation itself has become a commodity: working proof-of-concept exploits can be generated in about 15 minutes, and autonomous vulnerability discovery campaigns can be run for roughly $50. This isn’t theoretical. Recovered logs from a June incident (via OALABS) showed a single operator driving over 1,000 AI-agent sessions across 14+ companies, with the models flagging a policy violation only about 10 times — because every request was simply framed as “authorized red-team” work. Social engineering has scaled right alongside it: Arup lost $25.6 million in 2024 after a video call where every “colleague” on the line was a deepfake, and the economics of that kind of attack have only gotten cheaper since. None of this requires exotic new attack vectors. It’s the same attack surface organizations have always had, now facing an adversary that doesn’t sleep, doesn’t hesitate, and operates at machine speed.

Legacy and Unsupported Systems

The statement is blunt that unsupported systems are not just technical debt, they are strategic liabilities [1]. This sits uncomfortably with sectors where legacy estate is structural rather than incidental, where replacement cycles are measured in years and safety-critical considerations constrain how quickly systems can be patched or retired.

Governance Gaps Between Boards and Technical Teams

Repeated emphasis on board-level accountability assumes a level of cyber literacy that many boards do not yet have. The gap between technical risk and the language boards use to govern it remains one of the most persistent obstacles to the kind of confident assurance the statement demands.

AI as a Dual-Use Capability

The statement’s insistence that organisations use AI deliberately to strengthen defence, not just improve efficiency [1], is a meaningfully higher bar than most organisations’ current AI security posture. Many security teams have adopted AI tooling for productivity gains. Far fewer have built the detection, monitoring, and response capability the statement envisages, while simultaneously managing the new attack surface that frontier AI systems themselves introduce.

Shipping Faster Means Shipping More Attack Surface

AI coding assistants have undeniably accelerated software delivery — but speed and security are moving in opposite directions. Veracode’s testing across more than 100 models found that roughly 45% of AI-generated code carries a security weakness. What’s more concerning is the trendline: functional correctness has raced past 95%, while the security rate has stayed flat at around 55% for two years running. Newer, smarter models are writing code that works — not code that’s safe. The practical impact is that attack surface is now growing at delivery speed, with code merging faster than any human review process was designed to handle, whether that code was sanctioned by the organization or introduced through shadow AI use, often with no clear provenance to trace it back. At the same time, client appetite for security testing is rising faster than budgets are — creating a widening gap between how much validation organizations want and how much they’re actually resourcing. Closing that appetite-budget gap is, in many ways, the central challenge facing security teams right now.

Zero-Day Proliferation

The statement warns directly that as AI systems evolve, new and previously unknown vulnerabilities will emerge, including zero-day vulnerabilities [1]. Defence-in-depth, rather than reliance on any single control or technology, is positioned as the only credible response to a vulnerability landscape that is itself becoming less predictable.

How IOActive Can Help

IOActive’s work across offensive security, operational technology and industrial control system (OT/ICS) assessment, and critical infrastructure advisory positions us to support organisations translating this statement’s leadership-level mandate into operationally credible defence. Testing, not assumption, is how organisations find out whether their controls will actually perform under pressure, which is precisely the bar the statement sets.

Red Team and Purple Team Services

The statement’s call to verify that controls will perform during a real incident, not merely exist on paper [1], is best answered through adversarial testing rather than compliance review. IOActive’s Red Team operations emulate the tactics of the threat actors most likely to target a given organisation’s sector and assets, while our Purple Team engagements translate offensive findings directly into measurable improvements in detection and response.

Full Stack Security Assessments

Reducing attack surface and accelerating patching, two of the statement’s five practical actions [1], both depend on first knowing where exposure actually sits. Our full stack assessments examine internet-facing systems, cloud environments, and on-premises infrastructure together, identifying the specific systems an attacker would prioritise rather than producing a generic vulnerability count.

Supply Chain Integrity

The statement’s call for action extends explicitly to vendors [1]. IOActive’s Supply Chain Integrity service assesses the security posture of technology providers and critical third parties, reviewing firmware, embedded systems, and procurement processes for inherited risk before it becomes either a compliance finding or an incident vector.

Preparedness and Resilience Testing

The statement’s framing of preparedness as a capability to be trusted, not assumed, aligns directly with our tabletop exercise and crisis simulation work. We help organisations stress-test detection, escalation, and recovery processes ahead of a real incident, building the operational muscle memory boards are now being asked to assure.

Threat Modeling and Advisory

Understanding and assessing risk, readiness, and accountability, the statement’s first call to action [1], requires a structured, evidence-based view of organisational exposure. Our threat modelling and advisory engagements give CISOs and boards a shared, prioritised picture of risk that can be acted on with confidence rather than debated indefinitely.

The statement is explicit that delay carries growing and avoidable risk [1]. We recommend the following immediate actions.

  1. Brief your board now on the statement’s core claim, that cyber risk assumptions can become outdated in months, and translate that into business, financial, and reputational terms.
  2. Map your patch and change-management cycle against the compressed exploitation timelines the statement describes, identifying where current processes cannot keep pace.
  3. Inventory legacy and unsupported systems with explicit reference to their exposure on external attack surfaces, not just their internal criticality.
  4. Review identity and access controls across critical systems, with particular attention to permissions that have accumulated without recent review.
  5. Test your incident response plan through a structured exercise that assumes a breach has already occurred, rather than one that tests whether it can be prevented.
  6. Evaluate how AI is currently used across your security function, distinguishing tools adopted for efficiency from capability genuinely built to strengthen detection and response.

Conclusion

The Five Eyes statement is short by design, but its brevity should not be mistaken for limited weight. Six of the world’s most authoritative cyber security agencies have chosen to speak with one voice, in plain language, to say that the basis on which most organisations currently assess cyber risk is already out of date.

The statement does not introduce new technical obligations. It does something arguably more consequential: it removes the option of treating AI-accelerated cyber risk as a future planning consideration. Combined with the NCSC’s own recent disclosures on the scale of state-linked attacks against UK critical infrastructure and the coming vulnerability patch wave, the message to leaders is consistent and increasingly difficult to defer.

Organisations that wait for a forcing event, whether a breach, a regulatory deadline, or a sector-specific mandate, will be acting from a position of weakness. Those that act now, testing their assumptions rather than reviewing them, will be the ones still standing when the window the statement describes finally closes.

If you would like to discuss how your organisation’s current posture measures up against the expectations set out in this statement, or how IOActive can support your cyber resilience programme, we welcome the conversation.

References

[1] Australian Signals Directorate, Communications Security Establishment, Government Communications Security Bureau, National Cyber Security Centre (UK), National Security Agency, Cybersecurity and Infrastructure Security Agency. The AI Shift in Cyber Risk: Why Leaders Must Act Now. 22 June 2026. https://www.ncsc.gov.uk/news/the-ai-shift-in-cyber-risk-why-leaders-must-act-now

[2] National Cyber Security Centre. NCSC CEO: Hostile States Linked to Three-Quarters of Cyber Attacks Affecting UK’s Critical Systems. 17 June 2026. https://www.ncsc.gov.uk/news/ncsc-ceo-hostile-states-linked-to-three-quarters-of-cyber-attacks

[3] National Cyber Security Centre. Preparing for a ‘Vulnerability Patch Wave’. 1 May 2026. https://www.ncsc.gov.uk/blogs/prepare-for-vulnerability-patch-wave

WHITEPAPER | June 13, 2023

Drone Security and Fault Injection Attacks | Gabriel Gonzalez

Gabriel Gonzalez, IOActive Director of Hardware Security presents full technical detail of his research into drone security and side-channel/fault injection attacks in this whitepaper.

The use of Unmanned Aerial Vehicles (UAVs), commonly referred to as drones, continues to grow. Drones implement varying levels of security, with more advanced modules being resistant to typical embedded device attacks. IOActive’s interest is in developing one or more viable Fault Injection attacks against hardened UAVs.

This paper covers IOActive’s work in setting up a platform for launching side-channel and fault injection attacks using a commercially available UAV. We describe how we developed a threat model, selected a preliminary target, and prepared the components for attack, as well as discussing what we hoped to achieve and the final result of the project.

WHITEPAPER | September 25, 2018

Commonalities in Vehicle Vulnerabilities

With the connected car becoming commonplace in the market, vehicle cybersecurity continues to grow more important every year. At the forefront of security research, IOActive has amassed real-world vulnerability data illustrating the general issues and potential solutions to the cybersecurity threats today’s vehicles face.