INSIGHTS | October 6, 2026

Cyber Risk Management Strategies: Six Ways to Align Security With Business Priorities

Cybersecurity analyst reviewing business systems as part of a cyber risk management strategies.

Recent IOActive research on 2026 attack trends points to a recurring problem: attackers target systems that companies rely on the most, but watch the least, including operational technology (OT), firmware supply chains, and AI-enabled development. That raises a difficult question for security leaders: Which risks deserve investment first?

The answer is not always the vulnerability with the highest technical severity. A useful cyber risk management strategy links technical risk to business impact, tests whether attack paths are realistic, and directs funding toward the risks most likely to disrupt the business.

“The most effective cyber risk management strategies prioritize business impact alongside technical risk.” – IOActive Security Team

Six Pillars of Business-Aligned Cyber Risk Management Strategies

PillarsWhere to StartKey Areas CoveredQuestion to Consider
Link risk to business impactFocus on risks that could disrupt critical operationsCritical services, dependencies, recoveryWhich services and dependencies need protection first?
Test priorities through offensive securityTest whether realistic attack paths can achieve business objectivesPenetration testing, Red Team, threat researchHow well do controls detect and respond to realistic attacks?
Manage connected exposureEvaluate risks beyond the organization’s direct controlThird parties, supply chains, OTWhich dependencies could affect resilience?
Govern AI by impactMatch oversight to the impact of failure or misuseAI-generated code, automation, dataWhere are more controls required?
Turn findings into decisionsEstablish ownership, thresholds, and funding prioritiesGovernance, risk appetite, reportingWhat should be fixed, accepted, or escalated?
Keep testing securityConfirm that controls remain effective as conditions changeRetesting, Purple Team, detection, recoveryHow does the organization know risk is improving?

These six pillars move from prioritization to validation, giving leaders a repeatable way to decide what to test, fund, and revisit.

A cyber risk management strategy should begin with the business services that cannot tolerate long outages, like production, payments, clinical operations, customer access, or safety processes.

Why it is essential

A vulnerability’s priority depends on what it could affect. Business impact may include lost revenue, downtime, safety consequences, regulatory exposure, reputational harm, data-integrity problems, or recovery difficulty. A moderate weakness may deserve urgent attention if it sits between an attacker and a critical process.

The key question is not simply, “How severe is this vulnerability?” It is, “Could this weakness create a credible path to business failure?”

What to do next

For each critical service, identify the dependencies that would have to fail for the service to stop or operate unsafely. Then rank those dependencies according to access, consequence, recoverability, and the likelihood that an attacker could reach them.

One useful way to structure this conversation is the NIST Cybersecurity Framework 2.0. The framework gives security, technology, risk, and business leaders a shared language for identifying, governing, protecting, detecting, responding to, and recovering from cybersecurity risk.

Used this way, the framework helps turn isolated findings into priorities the business can understand and fund.

Test Priorities Through Offensive Security

Offensive security tests whether realistic attack paths could reach a defined business objective. It provides evidence that a risk register or compliance assessment cannot provide on its own.

Why it is essential

An attacker may need to combine several moderate weaknesses to reach a sensitive system. A penetration test can validate a defined application, network, or product. A Red Team exercise can test whether an attacker can reach a business goal across technical, physical, and human layers. A Purple Team exercise can then bring offensive and defensive teams together to evaluate detection, response, and control effectiveness.

What to do next

Start with the decision the assessment needs to support. Leaders may need to know whether a compromised vendor account could reach production, whether a stolen credential could access a critical application, or whether security operations would detect a multi-step attack.

Choose the assessment based on that question rather than a generic checklist. The result should show which weaknesses could be combined, which controls would interrupt the attack, and where further investment would reduce uncertainty or exposure.

Manage Connected Exposure Across Third Parties and OT

A business-aligned strategy must include systems and organizations outside the traditional security boundary. Vendors, cloud providers, software suppliers, remote-access channels, and operational technology may all support critical business services.

Why it is essential

A supplier compromise could interrupt operations, expose sensitive data, alter production, or delay recovery. Third-party risk should therefore reflect more than a vendor’s questionnaire score. It should account for access, dependency, concentration, recoverability, and business consequence.

OT creates an additional concern because a security incident may affect physical processes. In some environments, an integrity failure may be more serious than an outage if an attacker can manipulate commands, sensor data, or operating conditions.

What to do next

Map the suppliers and external connections that support each critical service. Identify which accounts, interfaces, software components, and remote-access paths could provide entry into the environment.

For OT, evaluate security decisions alongside safety, continuity, recovery, and the limits of specialized systems. The goal is to understand which dependencies could create a credible path from an external compromise to operational disruption.

Govern AI Adoption According to Business Impact

AI should not be governed as a single risk category. The right controls depend on what the system can access, influence, or automate.

An internal productivity assistant presents a different risk from an AI tool that generates live code, handles customer data, approves transactions, or influences operational decisions.

Why it is essential

The consequences of an AI failure depend on where the system sits in the business. An incorrect answer may be inconvenient in one workflow but costly, unsafe, or difficult to reverse in another.

If an incorrect or manipulated result could cause material harm, keep a person involved in approval and define when the system must stop. If the system can reach sensitive data, production systems, or consequential decisions, limit access, log use, and monitor the system. If its output can affect a release, customer, or operational process, test it before deployment and continue checking it in production.

What to do next

A practical AI risk review should ask:

  • What happens if the system produces an incorrect or manipulated result?
  • What data, systems, or decisions can the system influence?
  • What validation occurs before the output reaches production or a customer?

Use the answers to set the controls. Keep a person involved when an output could cause material harm, restrict and monitor access to sensitive data or production systems, and require testing before release and during use. Review the code, data, integrations, access paths, and workflows around the system, then retest them as the system changes.

Turn Risk Findings Into Investment Choices

A technical finding has limited value until someone decides what to do about it. Governance should give security and business leaders a consistent way to determine whether a risk requires remediation, mitigation, formal acceptance, or escalation.

Why it is essential

Without clear ownership and decision thresholds, unresolved exposure can remain in a risk register without receiving meaningful attention. Explicit decisions also help executives compare security investments with other business priorities.

The most useful recommendation explains what an investment changes. A Red Team exercise may reduce uncertainty about attack paths, while improved monitoring may shorten detection and containment. Those are different outcomes and should be evaluated accordingly.

What to do next

Assign each high-impact risk an owner, a defined consequence, a decision deadline, and a treatment plan. Document whether the organization will fix the weakness, reduce its likelihood, limit its impact, monitor it, accept it, or escalate it.

Give every accepted or mitigated risk a review date so the decision remains visible and accountable.

Keep Testing and Adjusting Cyber Risk Management Strategies

Cyber risk management strategies need regular review as the business, suppliers, technology, and threat landscape change. Organizations need evidence that remediation, detection, response, and recovery still work.

Why it is essential

A control that worked six months ago may no longer protect the same system, service, or attack path. Organizations need evidence that remediation worked and that detection, response, and recovery capabilities remain effective.

What to do next

Retest important attack paths after remediation. Review dependencies after major changes. Use Purple Team exercises to validate detection and response, and measure time to detect, contain, recover, and restore key services.

Threat research should also update the assumptions behind security decisions. The goal is not more testing for its own sake. It is a feedback loop:

Identify → Prioritize → Test → Improve → Validate → Reprioritize

This sequence helps organizations determine whether security investment is reducing business risk or simply generating more technical findings.

Invest in the Risks Most Likely to Disrupt the Business

Effective cyber risk management strategies help leaders identify the attack scenarios most likely to disrupt operations, assess whether current defenses would withstand them, and direct investment where it will reduce risk.

That requires more than vulnerability discovery. It requires business context, adversarial testing, governance, and continuous validation.

IOActive combines research, offensive testing, and advisory services to help organizations examine those questions across AI/ML, supply chains, OT/ICS, hardware, firmware, and enterprise environments. The result is a clearer basis for deciding what to fix now, what to monitor, and what risk the business can accept.

GLOBAL THREAT GROUP, INSIGHTS, RESEARCH | October 1, 2026

LLM-Assisted Vulnerability Research: Finding Real Bugs with Code-Reasoning Models

Introduction

Point a code-reasoning model at a codebase and it reads through it far faster than I can, then hands back a long list of things that look like bugs. Generating that list is the easy part now, but most of it is noise. Almost every candidate falls apart the first time I try to trigger it against a real build, and the handful that survive still have to be proven before they count as findings. Getting from that list to a proven bug is the slow, deliberate part of the job.

In this post I walk through the workflow that I use to close that gap, and two real bugs it has produced. Both are now fixed in Angular and rated HIGH severity:

  • CVE-2026-68945: a server-side render cache that could hand one visitor’s response to the next visitor, skipping the backend’s authorization check entirely.
  • CVE-2026-69151: a translation-file path that let lower-trust text compile into a live JavaScript event handler, turning a localization bundle into a script-injection vector.

These two are the CVEs I can name today; more from the same process are still working through coordinated disclosure.

Neither came from asking a model to “find bugs.” Both came from a disciplined loop: work out the rules a project’s code must never break, from its patch history and its own design, then keep the model’s attention on one small piece of code at a time, drive it through four focused audit passes, and believe nothing until I have proven it myself in a real runtime. The model does the reading and the guessing. I own the threat model, the runtime, and the final call.

That division of labor, not the automation, is what holds the whole thing together. The rest of this post builds it up, starting with why the obvious shortcut (just asking a model to find bugs) doesn’t work.

Why “Find All the Vulnerabilities” Fails

Ask a large language model (LLM) to “find security bugs in this repository” and it will answer at length, confidently – and almost none of it will be worth acting on.

What comes back is a list of plausible-looking weak points with no sense of which ones an attacker can reach, textbook bug patterns matched onto code that never runs, and not one claim that the model has tried to disprove. Separating the real from the decorative lands right back on you, now scaled up to the length of the model’s output.

A bare request gives the model no notion of reachability, no trust boundary, and no idea which bug classes a maintainer would actually accept. That means it falls back on the only thing it can do: it matches shapes that look like vulnerabilities and reports them with even confidence, whether they sit on a hot path or in dead code. The model writes clean, persuasive analysis, but nothing in the request forces it to be true.

Constrain the request to a single testable claim, anchored in the target’s own history, and the same model behaves differently:

Recent accepted vulnerabilities in this component involved validation/use mismatches. A patch introduced cumulative length validation in one path. Identify sibling paths that still validate entries independently while dispatching batches to a backend that consumes cumulative state. Demonstrate or refute this in the product runtime, with a negative control and a fixed-version comparison.

Now the model has a reachability target, a boundary that defines impact, and a claim that can be caught when it’s wrong. This still generates candidates, but each one arrives with the terms of its own refutation attached.

The difference between the two prompts is not the model or its size, but the question. And a good question starts from a threat model.

Threat Modeling: Rules the Code Must Never Break

Before I send any code to a model, I build a compact threat model from the project’s own security history: its advisories, its commits, its patch diffs. The point isn’t only to re-hunt old bugs. It’s to learn how the software expects its data to behave, so I can then look for places where that assumption quietly fails.

The output of that work is a short list of invariants. An invariant is a one-sentence rule the code assumes is always true and that a bug would break, for example, “the byte count validated at the perimeter must equal the amount the backend actually consumes.” That is the whole idea. An old validation-patch advisory becomes exactly that kind of rule. An advisory about a use-after-free (a bug where code keeps using an object after it’s been freed) becomes “no object may keep using itself after calling a callback that can destroy it.”

While an invariant is the rule the code must hold, once a bug shows that rule being broken, a variant is any other location where the same rule is broken in the same way.

Prior reported bugs make the hunt richer, because each one hands you an invariant the code already got wrong and a family of variants to chase, but they are not required. Where there are none, you derive the invariants from the code’s own contracts, a cache key that must stay unique or a trust boundary that must hold. This is where a model and an expert outrun what either does alone: the model maps the sinks and their connections across the whole surface far faster than a person could, while the researcher decides which invariants matter and where they break. The threat model is the method, and the pairing is what makes it powerful.

This is also why “find a use-after-free” is a useless prompt, while “determine whether this callback can free the parent object before the next line reads a member of it” is a good one. The first is a bug category, the second a specific claim a model and I can both test.

Nothing about this is specific to open-source. The method is the same whether the target is a public repository, a client’s proprietary codebase, or a closed black box, and what changes is only how much of the history is handed to you. Open-source projects give it away in public commits and advisories, which is why the examples in this post are open-source. On a closed target, you assemble the same ground truth yourself from decompiled interfaces, architecture documents, past vendor disclosures, and dynamic traces, and that is the slower, harder version of the identical work. A model can’t invent a threat model from nothing, and it works only from the evidence you give it.

Figure 1 – From a patch to its forgotten sibling.

Context Discipline: The Token Budget

Compressing history into an invariant is wasted effort if the model loses that invariant halfway through the job, which is exactly what happens when you ignore how its context works. A model doesn’t have your repository open the way an editor does. On every turn it sees one block of text – your prompt, the conversation so far, the code you pasted, the tool output that came back – and nothing else. That block is the context window, measured in tokens (roughly ¾ of a word each), and it is finite. Almost every failure I hit traces back to how that budget gets spent. The most relevant reasons for context failure are these three:

Lossy compaction. When a conversation nears the window limit, the tooling summarizes older turns to make room. Summaries keep the story (“found a validation gap”) and throw away the specifics upon which a finding depends: the exact line number, the build flag, the crash offset. After a couple of rounds of this, a session can sound completely confident while no longer holding a single piece of usable evidence. That’s the most dangerous failure mode, because nothing looks broken.

Degraded recall in the middle. A model recalls what sits at the start and end of a long context far better than what is buried in the middle. The effect has a name, “lost in the middle,” and it has outlasted several generations of far larger context windows. A 2025 context rot study found every one of eighteen frontier models still degrading as their input grew. That is the problem here, because the invariant, the one thing that has to hold from the first turn to the last, is exactly what sinks into that middle while the model keeps reasoning as if it still had it.

Cache invalidation. Reasoning over a large context is only affordable because the tooling caches the processed prefix. Churn the conversation and that cache goes cold, and every turn starts paying the full cost of reading everything again.

The fix is not a bigger window. That just moves the limit without solving the recall decay. The fix is to keep the working slice small enough that the invariant, the code under test, and the oracle (the check that separates a vulnerable build from a fixed one) all fit in the high-recall region at once. Everything else goes to disk.

Before I pause a session, I write a short HANDOFF.md at the repository root – the pinned commit, the current artifact, the live hypotheses, the dead ends already ruled out – so the next session resumes without re-deriving state. If that note runs longer than a page, the slice was too big and I split it.

Figure 2 – Keeping the slice in high-recall context.

The Four Audit Passes

With a threat model in hand and a list of candidate code paths to check, I run the model through four audit passes. Each one has its own input and its own output.

  1. Patch-diff variant farming. A patch shows the exact line the maintainers changed, and the code around it shows the neighborhood they left alone. I ask the model to find sibling paths that still contain the old, pre-patch logic. This is my highest-yield pass by a wide margin.
  2. Parser and serializer differentials. When a system implements the same concept more than once – two URL parsers, two header processors, two cache-key generators – those implementations can disagree, and the disagreement is a bug. I list the pairs and ask for inputs where they produce different results, aiming at request smuggling, server-side request forgery (SSRF), or cache confusion.
  3. Symmetry and dual checks. For every input-validation check, I look for its output twin: the cache layer, serializer, or redirect handler that’s supposed to enforce the same rule on the way out. A missing twin is a strong candidate.
  4. Cross-framework transfer. Once a bug pattern works in one target, I test sibling frameworks for the same underlying assumption. If it replicates, that’s a class-level design flaw, not one project’s slip.

Only the first pass looks backward at patches. The other three go looking for attack surface nobody has catalogued yet.

Those passes hand back a list of candidate bugs. I cut each surviving candidate down to a slice of under a hundred lines and send it through a dual-model gate, where one session is prompted to prove the bug exploitable and another, handed the same slice with no prior context, is asked to prove it safe. This is not a vote, because two runs of the same model share the same blind spots. I treat it as a disagreement detector instead.

When both sessions call the bug real, it still owes me a runtime proof; when both call it safe, I drop it; and a split, which flags the non-obvious edge cases, is the most useful of the three. The gate only filters, though, and whatever comes out of it I verify by hand, reading the code and reproducing the behavior myself, because two models agreeing is not proof.

Figure 3 – Four audit passes into a dual-model gate.

What a Verified Finding Looks Like

Before any of this touches a runtime, each candidate gets a spec card, four lines that fix exactly what is being claimed and what would prove it. Here’s one, from a sanitized browser-graphics example – a validation/use mismatch, where the check validates each draw command on its own but the backend runs the whole batch and walks the write cursor past the allowed buffer:

Invariant:  ValidateBatch must validate the cumulative byte count
            consumed by SubmitBatch, not the largest single entry.
Attacker:   A web page controls the draw list and the bound range.
Oracle:     The second draw must fail if draw1 + draw2 exceeds the
            bound range. A vulnerable build writes past the range;
            a fixed build returns an API error.
Stop:       Do not claim High unless the write crosses an allocation,
            ownership, process, origin, or authority boundary.

The oracle is the line that does the work, the exact condition that tells a vulnerable build from a fixed one. Here the oracle is checked across three runs, and the contrast between them is the proof:

separate_draw_control:    second_draw_error = INVALID_OPERATION
                          bytes_after_bound_changed = 0

batched_draw_vulnerable:  batch_error = NO_ERROR
                          bytes_after_bound_changed = 64

batched_draw_fixed:       batch_error = INVALID_OPERATION
                          bytes_after_bound_changed = 0

The first run is the negative control, where single checks reject the bad bounds. The second shows the bypass under batched execution, and the third confirms the fix. One log line would prove nothing, but the contrast across all three is what a triager can trust.

This example shows one more thing: because the writes stayed inside memory the calling page already owned, this was a constrained primitive (a real but limited capability), not full memory corruption. That distinction is the difference between a Medium and a High, and I’ll come back to it. The productive next question is never “can I make this proof-of-concept louder?” It’s “where else does this invariant fail?”

Figure 4 – A validation/use mismatch: control vs affected.

Two Real Findings in Angular

Here’s the method producing real, accepted bugs, both reported, confirmed, and patched upstream. One came from a forward-looking differential, the other from a patch variant.

CVE-2026-68945: two requests, one cache key

Advisory: GHSA-jhpw-976m-542j. Rated HIGH (CVSS v4.0) by the assigning authority.

The assumption, from history. Angular’s server-side rendering (SSR) builds the page on the server before sending it to the browser, and it caches the HTTP responses it fetches so the browser can reuse them during hydration instead of re-fetching. Any cache like that rests on one rule.

The invariant. Two HTTP requests that could receive security-distinct responses must never share a cache key.

The hypothesis (Pass 2, serializer differential). If the code builds cache keys by turning request parameters into a string, is there more than one way to write “the same” string? There was. The key generator joined repeated parameters with commas, so two genuinely different requests collapsed onto one key:

GET /api?role=user&role=admin   →  key "role=user,admin"
GET /api?role=user,admin        →  key "role=user,admin"

Figure 5 – Two requests, one cache key (CVE-2026-68945).

The proof. The oracle was a backend hit-counter, and two distinct requests produced exactly one backend call. The second visitor was served the first visitor’s cached response straight from the transfer cache, never reaching the backend’s authorization check. Angular fixed the key serialization in 20.3.27, 21.2.19, and 22.0.2 (PR #68571).

CVE-2026-69151: a translation file that could inject a script

Advisory: GHSA-jj27-h5hq-8×99. Rated HIGH (CVSS v4.0) by the assigning authority.

The assumption, from history. Angular’s template compiler treats internationalization (i18n) files as lower-trust: translators are supposed to supply text, not code. That’s a trust boundary, and trust boundaries are exactly what threat modeling tells you to test.

The invariant. Lower-trust translation content must never compile into an event handler.

The hypothesis. An earlier fix had already closed one hole where translated content bypassed the normal attribute checks. So, what other attributes still slip through the same translation path? Event handlers did. The compiler accepted i18n-onerror and other i18n-on* attributes, which meant a tampered translation file could drop live JavaScript into a static event handler.

The proof. The demonstration is a compile-time differential rather than a memory trace. Bind onerror on an element directly and the compiler rejects it. Mark the same attribute i18n-onerror and it passes the i18n collection path and compiles into the localized build intact. A translation file supplying that attribute’s value then controls a live handler, which is exactly what the advisory documents, a lower-trust translation file replacing a benign onerror=”void 0″ with arbitrary JavaScript that runs in the application’s origin. Angular fixed it in 20.3.27, 21.2.19, and 22.0.1 (PRs #68821 and #69306).

Figure 6 – i18n bypass: onerror vs i18n-onerror (CVE-2026-69151).

The bug wasn’t a sanitization slip but a broken contract between the compiler and untrusted input, and that input is realistic: localization bundles routinely come from outside translators or a shared repository. As with the cache-key bug, the fix seeds the next cycle, because patching one attribute is an invitation to check every sibling attribute immediately.

Candidate Classification

The fastest way to sink a valid report is to overstate it, so I rate a finding by one thing only: the security boundary its runtime proof actually crosses.

That’s why the graphics example earlier was a Medium, not a High. The writes stayed inside memory the page already owned, so nothing crossed an allocation, process, origin, or authority boundary. Both Angular bugs cross a real one: CVE-2026-68945 serves a response across users, and CVE-2026-69151 runs script in the application’s origin. A scanner’s score is only an interpretation of it, but the breached boundary is a fact, so document the boundary and let the rating follow.

Conclusion

The cost of generating a hypothesis, writing a negative control, and checking a patch’s variants has collapsed. What’s scarce now is the discipline to throw the bad ones away quickly and prove the good ones are real.

In practice, that comes down to a few habits: threat-model before you prompt, so every question is anchored to one rule; never report a finding without a negative control and a fixed-version comparison; and run every candidate past a fresh, independent model session whose only job is to disprove it. That session is only a filter, though, and the final sign-off on any finding is always a person’s.

This doesn’t automatically favor defenders. The same tools are a download away for anyone, and patch-diff variant farming rewards whoever moves first after a commit ships. This means the takeaway for anyone shipping code is direct: run these passes on your own patches before you release, because someone outside will run them after you ship.

INSIGHTS | September 29, 2026

AI Red Teaming Techniques for Enterprise AI Security

An enterprise security team reviewing AI red teaming techniques in a modern office environment.

Every major security breakthrough begins the same way: attackers find a new surface before defenders think to look there. It happened with networks, web applications, and the cloud. Now, as organizations embed artificial intelligence into their most critical operations, it is happening with AI systems as well, and conventional security testing was never designed to address it. That gap is exactly what AI red teaming techniques were built to close.

“AI red teaming reveals how attackers interact with AI systems in practice, exposing risks that traditional security testing often misses.” – IOActive Security Team

AI Red Teaming Techniques

Prompt Injection

What it is

A class of attacks where adversaries insert malicious instructions into AI input channels to override system behavior, hijack outputs, or trigger unauthorized actions.

How it works

Direct attacks manipulate user-facing prompts to override system instructions. Indirect attacks embed malicious commands inside content the AI trusts: retrieved documents, browsed webpages, or received emails. This allows attackers to compromise a system without ever interacting with it directly. Red teams use tools such as Garak and PromptBench to generate and score injection payloads across both vectors systematically. A structured engagement first maps every external input source the model trusts, then tests each path with escalating payload complexity to identify where instructions can be overridden.

Expected outcomes and benefits

Testing exposes which input paths, data sources, and integration points carry the highest risk of compromise. Organizations gain a clear map of where instruction override is possible and under what conditions.

Limitations

Injected instructions are difficult to detect because they can closely mimic legitimate input. Full coverage requires testing both direct and indirect vectors, and indirect paths multiply as AI systems access more external data sources.

Best for

Organizations deploying LLMs with access to external data, retrieval-augmented generation pipelines, agentic workflows, or enterprise integrations such as email, document management, and third-party APIs.

Prompt injection is the defining vulnerability class of the AI era. In direct attacks, adversaries manipulate user input to override system instructions. In indirect attacks, malicious instructions are embedded inside content the AI trusts: retrieved documents, browsed webpages, or received emails.

The risk scales with access. AI agents connected to enterprise platforms move 16 times more data than human users, turning a single compromised agent into an organization-wide exposure event, according to Obsidian Security. Indirect attacks now account for over 55% of observed prompt injection incidents, with 20 to 30% higher success rates than direct attacks due to stealth delivery through trusted content. OWASP ranks prompt injection as the number one vulnerability in its 2025 Top 10 for LLM Applications, with 73% of audited production systems showing exposure.

Jailbreak Testing

What it is

Systematic attempts to bypass the safety mechanisms, content filters, and alignment controls built into AI models, exposing gaps between intended behavior and actual behavior under adversarial pressure.

How it works

Red teams probe model guardrails using roleplay scenarios, multi-turn escalation sequences, adversarial rephrasing, cross-language attacks, and varied stylistic prompts designed to shift model behavior beyond its trained constraints.

Tools such as PyRIT (Microsoft’s Python Risk Identification Toolkit for red teamers) automate payload generation and track bypass rates across test runs, distinguishing a structured red team evaluation from ad hoc user experimentation. A formal engagement establishes a behavioral baseline, tests attack categories systematically, and maps each bypass to the specific guardrail or alignment control that failed.

Expected outcomes and benefits

Testing identifies which safety controls hold under pressure and which fail, along with the specific attack patterns most likely to succeed against a given deployment. Organizations can prioritize remediation based on realistic risk, not theoretical exposure.

Limitations

Success rates vary by model version and configuration. Not all bypasses carry equivalent risk, and some remediation options depend on controls that model providers, not operators, must implement.

Best for

Organizations deploying customer-facing LLMs, models with access to sensitive or regulated data, or AI systems where harmful outputs carry legal, reputational, or operational consequences.

Jailbreaks target safety mechanisms, bypassing the content filters and alignment controls built into a model. Roleplay-based attacks achieve 89.6% success rates against large language models. Multi-turn sequences, where an attacker escalates gradually across a conversation, reach 97% success within five exchanges. AI red teams probe these boundaries systematically across diverse attack styles, languages, and escalation strategies.

AI Agent Abuse and Privilege Escalation

What it is

Testing that targets autonomous AI agents’ susceptibility to goal hijacking, tool misuse, and unauthorized privilege escalation across the APIs, databases, and multi-agent environments they operate in.

How it works

Red teams attempt to redirect agent objectives, abuse tool access, and escalate privileges across agent architectures. In multi-agent environments, testers evaluate whether a single injected instruction can propagate across co-running agents in a cascading compromise.

Engagements begin by mapping the agent’s full tool inventory, permission scope, and trust relationships before testing whether crafted inputs can redirect agent behavior toward attacker-controlled objectives. MITRE ATLAS and custom agentic attack playbooks structure the engagement by threat actor objective rather than by system component.

Expected outcomes and benefits

Testing uncovers exploitable flaws in agent tool-execution logic, quantifies cascading risk in multi-agent pipelines, and validates whether permission boundaries hold under adversarial conditions.

Limitations

Scope grows significantly in complex multi-agent deployments. Comprehensive coverage requires deep access to agent architecture, tool configurations, and the external systems agents are authorized to reach.

Best for

Organizations running agentic AI workflows connected to enterprise systems, databases, or APIs, particularly where agents can take autonomous action such as sending emails, executing transactions, or modifying records.

Agentic AI introduces risks with no parallel in traditional testing. Agents plan, take action, and use tools: sending emails, querying databases, calling APIs, and chaining multi-step operations autonomously.

40% of AI agent frameworks contain exploitable prompt injection flaws in their tool-execution logic. Autonomous agents that call external APIs carry up to 2.5 times higher risk exposure than standalone models. In multi-agent environments, a single injected instruction can propagate to 48% of co-running agents in a cascading compromise. Red teams test for goal hijacking, tool misuse, and privilege escalation across these architectures.

Training Data Attacks and RAG Poisoning

What it is

Attacks that manipulate the data an AI system retrieves or learns from, corrupting its outputs and steering its behavior without touching the model itself.

How it works

Adversaries introduce crafted documents into retrieval pipelines or training datasets. Even a small number of poisoned entries can skew AI responses at scale. An adversary who controls what the AI retrieves effectively controls what it recommends and what actions it takes.

Red teams inject crafted documents into the retrieval corpus and query the system iteratively to measure how reliably poisoned content influences model outputs. Testing also covers chunk-level embedding manipulation and retrieval ranking abuse to determine how little adversarial content is required to achieve meaningful response distortion.

Expected outcomes and benefits

Testing reveals vulnerabilities in data ingestion pipelines, exposes how content control translates to output manipulation, and identifies supply chain risk in AI systems that rely on external or user-contributed data.

Limitations

Poisoning effects can be subtle and difficult to detect without targeted testing. Comprehensive assessment requires visibility into retrieval pipelines, document ingestion processes, and embedding logic, which may require close coordination with internal data and engineering teams.

Best for

Organizations using RAG architectures, enterprise knowledge bases, or AI systems trained on internal or third-party data, particularly where retrieved content directly shapes model output or user-facing recommendations.

Retrieval-Augmented Generation systems, which ground model responses in dynamically retrieved content, are vulnerable to data pipeline manipulation. Research published in Information (2026) shows that just five carefully crafted documents can manipulate AI responses 90% of the time through RAG poisoning. As IOActive has documented in its analysis of downstream AI attacks, an adversary who controls what an AI retrieves can control what it says and what it does.

Model Manipulation and Multimodal Attacks

What it is

Attacks targeting model confidentiality and cross-modal vulnerabilities, including efforts to reconstruct proprietary model behavior or embed malicious instructions inside image and audio inputs.

How it works

Model extraction attacks use crafted query sequences to map and replicate model behavior, enabling intellectual property theft or the creation of a shadow model for further exploitation. Multimodal attacks embed adversarial instructions inside images or audio accompanying benign requests, triggering unauthorized actions across input channels.

Extraction testing uses systematic query strategies to probe model decision boundaries and reconstruct behavior patterns without direct model access. Multimodal testing employs tools such as Burp Suite AI extensions and adversarial image generation frameworks to embed instructions inside image or audio payloads, then validates whether the target system processes them as executable commands.

Expected outcomes and benefits

Testing identifies intellectual property exposure, reveals how non-text inputs expand the attack surface, and validates whether input sanitization holds across all supported modalities.

Limitations

Model extraction requires significant query volume to produce a useful reconstruction. Multimodal testing scope depends on which input types the deployment supports, and organizations with restricted API access may face constraints on extraction simulations.

Best for

Organizations with proprietary or fine-tuned models, multimodal AI deployments, or any system that accepts user-submitted images, audio, or document files as part of its core workflow.

Model extraction attacks use crafted queries to reconstruct proprietary model behavior, enabling intellectual property theft or the development of a shadow model for further attacks. As AI systems expand to process images and audio alongside text, adversaries can embed malicious instructions inside images that accompany benign content, triggering unauthorized actions across input channels, according to OWASP’s LLM01:2025 guidance.

The techniques above represent the current front lines of AI adversarial testing. The urgency to apply them is driven by a threat environment that is moving faster than most organizations have adapted.

Why These Threats Are Escalating

The most consequential shift in the AI threat landscape is the rise of autonomous, tool-using agents operating across enterprise environments. A critical vulnerability in Microsoft 365 Copilot (CVE-2025-32711), rated CVSS 9.3, demonstrated how AI command injection in a widely deployed enterprise agent could enable large-scale data theft, as reported by Trend Micro. Controlled security experiments recorded credential theft or system manipulation outcomes in 70% of agent-based prompt injection trials.

OWASP’s December 2025 Top 10 for Agentic Applications and MITRE ATLAS’s October 2025 update, adding 14 new AI-focused techniques, reflect how quickly this space is evolving. The EU AI Act requires adversarial testing for high-risk AI systems ahead of its August 2026 compliance deadline, adding regulatory weight to what was already a security imperative.

The regulatory deadlines and real-world incidents above are not edge cases. They are the baseline risk environment organizations are now operating in, and the financial case for action reflects that reality.

The Business Case for AI Red Teaming

The business case for investing in AI red teaming is grounded in real incidents. 77% of businesses have reported an AI-related security incident. 35% of those incidents were caused by simple prompts, with some resulting in losses exceeding $100,000 per event, according to Adversa AI’s 2025 Security Report via Vectra AI. Organizations that run mature AI red teaming programs report 60% fewer AI-related security incidents than those without.

AI Red Teaming By The Numbers

Organizations are responding by increasingly investing in AI red teaming services. The market was valued at $2.26 billion in 2026 and is projected to reach $6.17 billion by 2030, growing at a 28.5% compound annual growth rate, according to Research and Markets. Demand for AI red teaming is forecasted to surge 35% by 2028, with qualified practitioners already in short supply, according to Practical DevSecOps.

The organizations best positioned to capture that advantage are those that can move beyond generic assessments and apply AI red teaming techniques grounded in real adversary behavior.

The IOActive Approach

IOActive brings decades of adversary simulation expertise to the full spectrum of AI red teaming techniques. The offensive instincts that break into financial institutions and critical infrastructure apply directly to AI systems that reason, decide, and act. IOActive connects proven red team methodology to the emerging AI attack taxonomy, testing how real adversaries would engage with an organization’s models, agents, and data pipelines.

A full engagement covers threat modeling scoped to the organization’s specific AI architecture, prompt injection testing across direct and indirect paths, jailbreak and safety bypass evaluation, agent workflow abuse testing, RAG and training data integrity assessment, privilege escalation validation, and multimodal attack surface review.

The output is a realistic picture of exploitable risk, produced under controlled conditions, with time to remediate before a real adversary finds the same vulnerabilities in production.

The vulnerabilities in your AI systems exist whether or not you have tested for them.

IOActive helps organizations find and fix exploitable risks in large language models, AI agents, and generative systems before adversaries do.

INSIGHTS | September 25, 2026

One Exposure, Many Actors: What the NCSC’s OT Alert Adds to the AA26-097A Picture

Key Takeaways

  • On 27 August 2026, the NCSC published an Alert warning of increased targeting of operational technology (OT) systems across multiple sectors globally, including in the UK, carried out by a range of threat actors and resulting in some limited real-world disruption [1].
  • The Alert is deliberately unattributed. It should not be read as an extension of joint advisory AA26-097A, which attributes a separate and well-documented PLC exploitation campaign to Iranian-affiliated actors [2].
  • Three separate bodies of reporting published between April and August 2026, covering Iranian-affiliated PLC exploitation, Russian FSB Centre 16 router targeting, and unattributed activity against water sector controllers, converge on the same access route of systems reachable directly from the public internet.
  • The FBI and EPA have reported that similarities in network configurations supplied by third parties allowed attackers to repeat successes across multiple victims, making integrator and supplier architecture a first-order concern rather than a downstream one [3].
  • The Alert lands five days before the UK Cyber Security and Resilience Bill reaches Lords Committee Stage on 1 September 2026, placing OT exposure in front of regulators, boards and legislators simultaneously [4].

Why the NCSC’s Alert Matters Now

The NCSC classified its 27 August publication as an Alert rather than guidance. That distinction carries weight. The organisation states that it has been engaging with affected sectors directly in response to recent targeting and is publishing the material to support national resilience efforts, which indicates that the advice follows live incident handling rather than horizon scanning [1].

The substance is uncomfortable for anyone who has treated OT exposure as a solved problem. The NCSC reports increased targeting of OT systems across multiple sectors globally, including in the UK, conducted by a range of threat actors, with some limited real-world disruption already recorded [1]. It also sets out a broader judgement. Against a backdrop of technology-enabled uplifts in offensive capability and increased geopolitical instability, the threat from state use of offensive cyber, including outside of conflict, has almost certainly increased [1].

The timing compounds the message. The Alert arrives within four months of joint advisory AA26-097A, seven weeks after an international advisory on Russian state targeting of network devices, and four weeks after an FBI and EPA announcement describing degraded water operations in at least seven US states [2][3][5]. Regulators are moving in parallel. The Cyber Security and Resilience Bill is scheduled to enter Lords Committee Stage on 1 September 2026, with Royal Assent expected later in the year [4].

What Does the NCSC Alert Actually Say?

The Alert sets out eight actions. Several restate established practice:

  • Build a definitive view of OT architecture including all assets, communications pathways and external connections
  • Replace default credentials and enforce unique administrator accounts with multi-factor authentication where supported
  • Harden the OT boundary and keep gateways, firewalls, routers and remote access appliances within vendor support
  • Segment OT, management and business networks
  • Maintain tested, ransomware-resistant backups of configurations, controller logic and critical engineering data [1]

Two items deserve closer attention because they move beyond network-layer control.

The first concerns device state. The NCSC advises that OT devices be operated in a condition that prevents remote programming during normal operations, ensuring PLCs are not left in PROGRAM or other maintenance modes and that controller logic uses password-based write protection or equivalent mechanisms to prevent unauthorised changes from malicious endpoints [1]. This is a control on the device itself, not on the path to it, and it assumes the network boundary may fail.

The second concerns protocols. The Alert recommends migrating industrial protocols to secure variants where available, naming DNP3 to DNP3-SAv5, CIP to CIP Security, Modbus to Modbus Security, and OPC DA to OPC UA, while removing telnet and SNMP versions 1 and 2 from management use [1]. Where no secure alternative exists, insecure protocols should be confined to isolated network segments.

The Alert also addresses organisations without OT estates, noting a continuing pattern of disruptive activity against internet-exposed systems and edge devices across all sectors, and directing readers to maintain an accurate inventory of internet-facing systems, retire end-of-life equipment and monitor for unexpected configuration changes or outbound connections [1]. It points UK organisations towards the NCSC’s free Early Warning service and frames the Cyber Assessment Framework as the mechanism through which boards should seek assurance [1].

Notably, the Alert names no threat actor.

How Does This Connect to AA26-097A, and Where Does It Not?

This is the point at which analysis most often goes wrong, and the distinction matters more than the similarity.

Advisory AA26-097A, first published on 7 April 2026 and updated on 22 July 2026, attributes an active campaign to Iranian-affiliated APT actors exploiting internet-connected PLCs across US critical infrastructure. The July update broadened the observed manufacturer scope from Rockwell Automation/Allen-Bradley to include Schneider Electric and Siemens controllers, added detection guidance for malicious changes in reusable code modules within Rockwell PLC programs, and brought the Department of the Treasury onto the authoring byline alongside CISA, the FBI, NSA, EPA, the Department of Energy and US Cyber Command’s Cyber National Mission Force [2].

Separately, an international advisory published in July 2026, co-sealed by the NCSC alongside 18 partner agencies from 12 countries, attributes sustained targeting of poorly configured routers to Russian Federal Security Service Centre 16, tracked variously as Berserk Bear, Energetic Bear, Dragonfly, Ghost Blizzard and Static Tundra. That activity centres on internet-wide scanning for devices accepting default or weak SNMP community strings, followed by exfiltration of device configuration files [5].

Separately again, the FBI and EPA issued a public service announcement on 30 July 2026 describing unattributed malicious cyber actors targeting internet-facing Rockwell Automation/Allen-Bradley MicroLogix 1100 and 1400 PLCs. Water and wastewater utilities in at least seven states reported incidents from 27 July 2026 onwards. The actors changed device IP addresses and set passwords, producing loss of view and, in some cases, loss of function. Reported operational effects included loss of pressure and flooding, with the announcement noting that pressure loss in water systems could potentially allow untreated groundwater to seep into pipes [3].

Four publications carry three different attribution states. One names Iranian-affiliated actors, one names a Russian intelligence unit, and two describe activity the authoring agencies have chosen not to attribute. Collapsing these into a single narrative would be an analytical error, and one that regulated operators cannot afford to make in reporting or board briefings.

What survives the separation is the access route. In every case the initial condition is the same, namely a control system, a management interface or an edge device reachable from the public internet and frequently protected by default, weak or absent credentials. The NCSC’s phrasing, that this activity is being carried out by a range of threat actors, is the more useful framing for defenders precisely because it removes the actor from the risk calculation. Exposure is exploitable by whoever finds it first.

Who is Affected by the NCSC’s OT Alert?

CNI operators with field-deployed OT

Any organisation with internet-exposed OT could be affected. The NCSC is explicit that organisations should not assume their OT is inaccessible from the internet without verifying it, since unintended exposure can arise through misconfigurations, legacy connections or unmanaged assets [1]. Water, wastewater and energy operators with geographically distributed sites and cellular field connectivity carry the highest concentration of this risk.

System integrators and managed service providers

The FBI and EPA observed that across several victims, similarities in network setups provided by third parties may give attackers the opportunity to multiply successes where vulnerable configurations recur across a supplier’s customer base [3]. A single integrator’s default deployment pattern becomes a repeatable attack template. This finding maps directly onto the critical supplier designation powers in the Cyber Security and Resilience Bill.

Organisations without OT

The Alert is explicit that the broader pattern of disruptive activity against internet-exposed systems and edge devices affects all sectors [1]. Routers, firewalls, VPN concentrators and remote access appliances present the same fundamental exposure.

Boards and regulated entities across jurisdictions

UK operators face the CAF as the assurance framework of record and the Cyber Security and Resilience Bill as the coming statutory instrument [1][4]. EU operators are working to NIS2. US operators sit within the CISA advisory and directive framework, including binding requirements on end-of-support edge devices [6]. The technical remediation is common across all three, and only the reporting obligations differ.

What Makes This Exposure So Difficult to Eliminate?

The mitigations in the NCSC Alert are not novel, and that is the difficulty. Every one of the eight actions was recognised practice before August 2026. Their repeated publication indicates a gap between recommendation and implementation rather than a gap in knowledge.

Three factors sustain that gap.

Asset visibility remains hard in OT. Exposure is frequently unintended, arising from a commissioning shortcut, a supplier’s remote support link, or a cellular modem installed to solve an operational problem years earlier. An operator can hold an accurate architecture diagram and still be exposed through an asset that diagram never captured.

Integrity attacks defeat availability-focused monitoring. AA26-097A describes actors using legitimate vendor engineering software to alter controller logic and falsify operator displays, and the FBI reported at least one organisation discovering modified PLC project files after noticing ladder logic discrepancies across several sites [2][3]. Where the control room’s picture of the process is itself manipulated, conventional uptime monitoring will report normality.

Remediation carries operational cost. Removing a PLC from direct internet exposure, migrating a protocol, or replacing an end-of-life gateway involves planned downtime in environments where downtime is the primary risk being managed. This is where technical recommendation meets capital planning, and where accountability for the outcome frequently disperses.

How IOActive Can Help

The NCSC’s Alert describes a condition rather than an incident. That condition is assets reachable from the internet that the owner did not know were reachable, protected by credentials the owner did not know were still default. Verifying that condition is an assessment problem before it is a remediation problem, and it spans the OT estate, the edge estate and the third parties with standing access to both.

IOActive has operated in industrial control environments since building the first proof-of-concept worm against the smart grid in 2009 [9], and has helped define standards and best practices including NIST 800-53 and 800-37. The services below map to the specific gaps this Alert exposes.

Full Stack Security Assessments

The Alert’s first action, to build a definitive view of OT architecture and not assume OT is unreachable without verifying it, is precisely the work most assessments stop short of. Our Full Stack assessments examine the whole environment rather than a single layer, covering internet-facing controllers and HMIs, the OT/IT boundary, the cellular modem paths that frequently provide field sites with their only route to the internet, and the firmware and silicon of the devices themselves through penetration testing, reverse engineering, side-channel analysis and fault injection.

Two of the Alert’s actions sit at the device rather than the network layer. Those are confirming controllers are not left in PROGRAM mode, and migrating industrial protocols to secure variants such as DNP3-SAv5, CIP Security, Modbus Security and OPC UA. These are testable conditions, not policy statements, and they are where an assessment either produces evidence or produces reassurance. As we set out in Cyber Attack Trends 2026, the environments organisations depend on most are routinely the ones they monitor least.

Supply Chain Integrity

The FBI and EPA finding that similar third-party network configurations allowed attackers to repeat successes across a supplier’s customer base is the most consequential detail in the current reporting, and the least addressable by any single operator acting alone [3]. It also has a direct precedent in our analysis of the Polish energy sector incident, which documented more than 30 facilities breached through the same pattern of internet-facing VPN portals without MFA, where credential reuse across sites turned one compromised account into a fleet-wide problem.

Our Supply Chain Integrity work assesses the security posture of technology providers and critical third parties, covering firmware, embedded systems, remote access arrangements and procurement processes, so that inherited risk is identified before it becomes a shared incident. For UK operators, this maps onto the critical supplier designation powers in the Cyber Security and Resilience Bill [4].

Red Team and Purple Team Services

The Alert notes that OT environments are typically static and predictable, which makes baseline monitoring highly effective at identifying unauthorised activity [1]. The question is whether a given deployment actually achieves that in practice. Adversarial testing is how an operator finds out before an adversary does, particularly against integrity attacks such as altered controller logic and falsified operator displays, which availability-focused monitoring is not built to catch. Our Purple Team work translates those findings into measurable detection improvements, and we set out the distinction between the two approaches in Red Team vs Penetration Testing.

Advisory Services

The eight actions in this Alert have all been published before, in various forms, and that is the problem worth solving. When we examined the Minnesota water utilities in When the Advisory Arrives First, the gap was not the absence of a warning but the absence of an owner for acting on it.

Our Advisory Services cover programmatic security review, security programme development and management, and Virtual CISO support. They turn an alert into a prioritised remediation plan with named accountability, and produce the evidence boards need for CAF or equivalent regulatory assurance. That work is increasingly shaped by converging regimes across all three jurisdictions, as we examined in our analysis of the UK Energy Sector Cyber Security Strategy and the Five Eyes statement on AI and cyber risk. The latter is directly relevant to the NCSC’s judgement here that technology-enabled uplifts in capability are part of why the threat has increased [1].

If you would like to discuss how your organisation’s internet-facing OT and edge exposure measures up against the activity described in this Alert, we welcome the conversation.

  1. Verify exposure rather than assuming its absence. Build or refresh a definitive view of the OT architecture covering all assets, communications pathways and external connections, and validate it against external scanning results rather than against documentation [1].
  2. Remove control system devices from direct internet reachability. Broker all remote access through a secure gateway or jump host so that OT is never directly exposed to external networks, and secure cellular modems used for field connectivity with strong authentication and logging [1][3].
  3. Eliminate default and shared credentials on OT and edge devices. Enforce unique administrator accounts, enable multi-factor authentication where supported, and use key-based authentication in preference to passwords where the protocol allows [1].
  4. Enforce device-level write protection. Confirm that PLCs are not left in PROGRAM or maintenance modes during normal operations and that controller logic is protected against unauthorised modification [1].
  5. Validate controller logic integrity. Use vendor integrity checking tools to compare running programs against known-good logic, and verify that backups are free of malicious logic before restoration [3].
  6. Retire or isolate end-of-support edge devices. Maintain a rolling replacement forecast reviewed against ownership and procurement, and apply compensating controls with firm decommission dates where replacement is delayed [3][6].
  7. Baseline and monitor OT network traffic. OT environments are typically static and predictable, which makes baseline monitoring effective at identifying communication with PLCs and HMIs from unexpected devices, networks or routes [1].
  8. Rehearse manual operation and recovery. Test the capability to revert to manual control, isolate affected systems and restore from trusted backups, and treat those exercises as evidence for CAF or equivalent regulatory assurance [1][3].

Conclusion

The NCSC’s August Alert does not describe a new technique, a new vulnerability class or a new adversary. It describes a durable condition that a widening set of actors continues to find and exploit, namely control systems and edge devices reachable from the open internet. The value in reading it alongside AA26-097A is not in merging the two, but in observing that campaigns with entirely different sponsors, objectives and levels of sophistication are arriving at the same door.

That has a practical consequence for how operators prioritise. Attribution shapes the diplomatic and sanctions response, but it does not change the remediation. An internet-facing PLC with a default password represents identical risk whether the entity that finds it is a state intelligence service or an opportunistic scanner. As the UK’s statutory regime moves towards Royal Assent and the EU and US frameworks tighten in parallel, the operators best positioned will be those who can evidence what they expose to the internet, and who owns the answer.

References

  1. National Cyber Security Centre, Disruptive cyber activity highlights risk from internet-exposed systems and edge devices, 27 August 2026. https://www.ncsc.gov.uk/news/disruptive-cyber-activity-highlights-risk-from-internet-exposed-systems-and-edge-devices
  2. CISA, FBI, NSA, EPA, DOE, US Cyber Command CNMF and Department of the Treasury, Joint Cybersecurity Advisory AA26-097A, Iranian-Affiliated Cyber Actors Exploit Programmable Logic Controllers Across US Critical Infrastructure, published 7 April 2026, updated 22 July 2026. https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-097a
  3. Federal Bureau of Investigation and Environmental Protection Agency, Malicious Cyber Actors Targeting Water and Wastewater Sector Internet-Facing Programmable Logic Controllers, Causing Operational Disruptions, 30 July 2026. https://www.fbi.gov/investigate/cyber/alerts/2026/malicious-cyber-actors-targeting-water-and-wastewater-sector-internet–facing-programmable-logic-controllers-causing-operational-disruptions
  4. UK Parliament, Cyber Security and Resilience (Network and Information Systems) Bill, HL Bill 32, Lords Second Reading 14 July 2026, Committee Stage 1 September 2026. https://bills.parliament.uk/bills/4035
  5. National Cyber Security Centre, UK and Allies urge critical sectors to improve defences against Russian intelligence targeting, July 2026. https://www.ncsc.gov.uk/news/uk-and-allies-urge-critical-sectors-to-improve-defences-against-russian-intelligence-targeting
  6. CISA, Reducing the Attack Surface for End-of-Support Edge Devices and Binding Operational Directive 26-02. https://www.cisa.gov/resources-tools/resources/reducing-attack-surface-end-support-edge-devices
  7. NCSC-UK, FBI and CISA, Secure Connectivity Principles for Operational Technology. https://www.ncsc.gov.uk/collection/operational-technology/secure-connectivity
  8. National Cyber Security Centre, Cyber Assessment Framework. https://www.ncsc.gov.uk/collection/cyber-assessment-framework
  9. M. Davis, IOActive, Advanced Metering Infrastructure (Smart Grid) Device Security, Black Hat USA 2009. https://blackhat.com/presentations/bh-usa-09/MDAVIS/BHUSA09-Davis-AMI-SLIDES.pdf
INSIGHTS | September 22, 2026

IT Risk Management Strategies: Key Metrics & Trends

Cyberattacks and data breaches have ranked as the number one global business concern for three consecutive years, according to Aon’s 2025 Global Risk Management Survey. The organizations still managing risk reactively are paying a measurable price. Hyperproof’s 2026 IT Risk and Compliance Benchmark Report found that 50% of those organizations suffered a breach in 2025. Among organizations with integrated, automated programs, the breach rate dropped to 27%.

That 23-point gap reflects a strategic difference in program design: how often risk is assessed, how findings connect to business decisions, and whether testing reflects real-world attack conditions.

This piece identifies key IT risk management strategies that separate mature programs from vulnerable ones, anchored in benchmarks from Verizon, Forrester, IBM, and Hyperproof:

  • Shift from periodic to continuous risk assessment
  • Align security testing to where attackers are actually moving
  • Extend technical validation into third-party and supply chain environments
  • Treat AI red teaming as a separate discipline from traditional red teaming
  • Quantify risk in business terms to improve security investment decisions

IT Risk Management Strategies

StrategyWhat It CatchesWhen to PrioritizeSign of Maturity
1. Shift to continuous risk assessmentBlind spots between assessment cyclesBoard lacks risk visibility; breach rate exceeds benchmarksNIST CSF 2.0 Govern in place; findings escalate to leadership on a defined cadence
2. Align testing to attacker trajectoriesAssessments that miss active threat vectorsEdge devices, ransomware, or third-party vectors are outside current scopeTTPs modeled on current threat actor behavior, not generic categories
3. Extend technical validation into third-party environmentsVendor programs limited to questionnaires and attestationsThird-party access is broad, or a vendor breach occurred in the past 24 monthsCritical suppliers tested under realistic attack conditions
4. Treat AI red teaming as a separate disciplineAI deployments assessed only against traditional controlsLLMs, ML models, AI APIs, or AI-integrated hardware are in productionPrompt injection, data poisoning, and adversarial inputs tested separately from IT infrastructure
5. Quantify risk in business termsSecurity findings that never reach prioritization or budget decisionsSecurity investment lacks ROI framing; findings stay technicalFAIR-based financial exposure tied to findings; risk reported in dollar terms

5 IT Risk Management Strategies That Separate High-Maturity Programs from Vulnerable Ones

“Effective IT risk management strategies prioritize the risks most likely to disrupt critical business operations, not simply the risks that are easiest to identify.” —IOActive

Strategy 1: Shift from Periodic to Continuous Risk Assessment

The biggest differentiator between high-maturity and low-maturity programs is not which framework they use. It is how often they assess.

Treating risk assessment as a periodic compliance exercise creates blind spots between cycles. The 23-point breach rate differential between reactive and integrated programs reflects this directly. Forrester’s State of Enterprise Risk Management 2025 adds another dimension: organizations without board-level ERM visibility were 20% more likely to suffer six or more critical risk events in a given year.

For most enterprise security programs, “continuous” does not mean auditing everything at once. It means layering assessment activities across defined cadences:

  • Ongoing: Automated monitoring of critical assets, threat intelligence feeds, and anomaly detection across production environments
  • Quarterly: Risk register review, third-party posture updates, and findings-to-remediation tracking against defined SLAs
  • Annually (minimum): Full penetration test of the environment, including edge devices and third-party access points
  • Change-triggered: Targeted assessment after any significant infrastructure change, acquisition, new vendor integration, or disclosed supplier breach

Two frameworks structure this cadence in practice. NIST CSF 2.0’s Govern function formalizes how risk oversight connects to organizational decision-making, not just technical operations. The FAIR (Factor Analysis of Information Risk) methodology translates those findings into probabilistic financial exposure that boards can act on. If current assessments cannot convert a technical finding into a quantified business risk, the reporting structure, not just the testing methodology, needs to change. IOActive’s penetration testing and advisory engagements are designed to generate findings at that level of specificity.

Reactive vs. Integrated Risk Programs: How Maturity Drives Breach Outcomes

Program CharacteristicReactive ProgramIntegrated Program
Assessment frequencyAnnual or periodicContinuous, change-triggered
Risk reportingSiloed, technicalTied to business outcomes
Board visibilityLimited or absentFormal ERM governance
Breach rate (2025)50%27%
Framework alignmentAd hocNIST CSF 2.0, FAIR, or equivalent

The table above makes the gap concrete: the difference between a 50% and a 27% breach rate is not a technology gap. It is a program design gap. Every row represents a decision a security leader can change without replacing existing tooling. To close that gap, start with three diagnostics:

  • Audit your assessment cadence. If the last full assessment was more than 12 months ago and was not triggered by an infrastructure change, your program is operating reactively regardless of which framework it claims to follow.
  • Test your reporting chain. Pull the most recent security finding delivered to leadership. If it was framed in technical terms without a financial exposure estimate, the board cannot act on it — and the Forrester data suggests that this gap is directly linked to higher rates of critical risk events.
  • Check framework alignment. NIST CSF 2.0 and FAIR are not aspirational targets. They are the operational baseline of integrated programs. If neither is in use, the program lacks a structured escalation path from findings to resource allocation.

Strategy 2: Align Security Testing to Where Attackers Are Moving

According to the Verizon 2025 DBIR, three vectors account for the most significant attacker movement in 2025: edge devices (exploitation up nearly eightfold), ransomware (up 37% year-over-year in breach data), and third-party environments (now involved in nearly 30% of all breaches, double the prior year). These are not random fluctuations. They reflect a deliberate shift toward the parts of enterprise environments that are hardest to monitor and least likely to be included in standard assessment scopes.

Organizations that limit testing to familiar, well-controlled internal environments are building a risk picture that excludes exactly the vectors attackers are actively exploiting. Effective IT risk management strategies account for where threats are moving: assessments must be built on current threat intelligence and on documented tactics, techniques, and procedures (TTPs) from active threat actors, rather than generic attack categories.

Active Threat Vectors Every IT Risk Management Strategy Must Account For (2025)

Threat Vector2025 TrendAssessment Implication
Ransomware37% YoY in breach dataTest detection and containment, not just prevention
Edge device exploitationNearly 8x growthExpand scope beyond the traditional network perimeter
Third-party involvementDoubled; ~30% of all breachesInclude supplier environments in testing scope
Phishing/social engineeringConsistent top initial access vectorTest employee response under realistic conditions

Strategy 3: Extend Technical Validation Into Third-Party and Supply Chain Environments

Third-party risk has moved from a compliance checkbox to a primary breach vector. Ethixbase360 reported that nearly 30% of all data breaches in 2025 involved third-party vendors or suppliers, double the rate from 2024. When a breach originates from a third-party environment, average remediation costs reach $4.8 million.

Exposure increases as the attack surface expands. SaaS integrations, AI APIs, cloud-managed services, and hardware supply chains create entry points outside an organization’s direct control. Attackers increasingly target upstream suppliers as indirect entry routes into better-defended downstream targets, a vector that questionnaires and attestations alone do not address.

For security leaders ready to move beyond attestation-based programs, four steps define where to start:

  • Classify vendors by access level. Identify which third parties have privileged access to your network, data, or production systems. Those with the highest access represent the highest residual risk if their environments are unvalidated.
  • Audit your current validation method per vendor. For each critical supplier, determine whether your current risk assessment is based on a questionnaire, a certification review, or actual technical testing. Questionnaire-only programs cannot detect active vulnerabilities or misconfigurations in a vendor’s environment.
  • Prioritize vendors who have changed or expanded access in the past 12 months. New SaaS integrations, API connections, and cloud-managed service agreements are the most common sources of unvalidated exposure in enterprise supply chains.
  • Schedule technical testing of your highest-risk suppliers. This means penetration testing of vendor-facing access points and, where applicable, hardware and firmware assessment for vendors operating in your physical or OT environment.

Technical validation of critical supplier environments, tested under realistic attack conditions, is where the real picture emerges. IOActive’s Supply Chain Integrity services take an attacker’s-mindset approach across every technology layer, including hardware, firmware, and silicon-level components.

Strategy 4: Treat AI Red Teaming as a Separate Discipline from Traditional Red Teaming

AI is both a security tool and a security risk vector, and those two realities require separate responses. IOActive’s AI Security Services address the risk side: testing the AI implementations that enterprise clients are deploying for vulnerabilities that traditional security assessments were not built to find. The distinction is not a matter of degree; it is a difference in target, methodology, threat framework, and deliverable.

AI Red Teaming vs. Traditional Red Teaming: Two Disciplines, Two Different Attack Surfaces

DimensionTraditional Red TeamAI Red Team
Primary targetNetworks, endpoints, applications, physical controlsML models, LLMs, training pipelines, AI APIs
Key attack typesLateral movement, privilege escalation, social engineeringPrompt injection, model extraction, data poisoning, adversarial inputs
Threat frameworkMITRE ATT&CK (Enterprise, Mobile, ICS)MITRE ATLAS, OWASP LLM Top 10
ScopeIT perimeter, cloud infrastructure, physical accessInference endpoints, training data, AI-integrated hardware, CI/CD pipelines
Deliverable focusControl gaps, detection capability, IR readinessModel robustness, data integrity, AI compliance posture
On-prem applicabilityYesPartial (AI-integrated hardware and embedded systems)

The Adversa AI 2025 report found that 35% of real-world AI security incidents resulted from simple prompt attacks, with no specialized tooling required: only manipulated inputs causing unintended behavior or data leakage. That finding aligns with IOActive’s own research, which evaluated 27 leading AI models using 730 real-world programming prompts across 27 languages and 219 vulnerability categories. Organizations testing AI implementations only against traditional security controls are not testing against the attacks that are actually succeeding in the field.

For security leaders building an AI red teaming capability or evaluating whether one is needed, four steps define where to start:

  • Inventory every AI deployment in production. Include LLMs, ML models, AI-integrated APIs, and any hardware with embedded AI components. Systems that have not been formally cataloged cannot be scoped into a testing program.
  • Determine how each has been tested to date. If AI systems have only been assessed under a traditional penetration testing scope, prompt injection, data poisoning, and adversarial input attacks have not been evaluated. That gap is where 35% of real-world AI incidents originate.
  • Prioritize by exposure. Public-facing AI interfaces and AI systems with access to sensitive data or internal decision-making pipelines carry the highest risk and should be tested first under MITRE ATLAS and OWASP LLM Top 10 frameworks.
  • Engage an AI red team, not a traditional one. The methodologies, tooling, and threat frameworks are different enough that assigning AI security testing to a conventional red team produces incomplete findings. The attack surface requires a purpose-built discipline.

Strategy 5: Quantify Risk in Business Terms to Improve Security Investment Decisions

Investment without measurement does not close the breach rate gap. Hyperproof found that 58% of GRC professionals anticipated increased risk and compliance spending in 2026, and that trajectory is expected to continue. The organizations extracting the most value from that spend are those directing it toward real-world validation with measurable business-risk outputs.

The FAIR methodology converts technical findings, such as a misconfigured edge device or an unpatched vendor API, into probabilistic estimates of financial exposure. NIST CSF 2.0’s Govern function formalizes the escalation path, providing a structure for how those estimates reach board-level decision-makers and translate into resource allocation. The result is a feedback loop: assessments generate findings, findings translate into financial exposure, and financial exposure directs investment where actual risk is highest.

Find Out Where Your IT Risk Management Strategies Actually Stand

Before directing more investment toward security tools or programs, three questions can identify where current gaps are largest:

  1. Can your team articulate, with data, how a sophisticated threat actor would move through your environment today? If the most recent assessment was conducted more than 12 months ago or did not include edge devices and third-party access points, the answer is likely no.
  2. Have your AI implementations been tested against prompt injection, data poisoning, and adversarial inputs? Testing AI only against traditional security controls leaves the most common AI-specific attack vectors unaddressed.
  3. Does your third-party risk program include technical validation of critical supplier environments, or only attestations and questionnaires? Given that 30% of 2025 breaches involved third parties, questionnaire-only programs leave a primary breach vector unvalidated.

If those questions cannot be answered with confidence, they are the gaps worth closing first. IOActive’s research-driven assessments connect offensive security findings to the business risk decisions that CISOs, CIOs, and boards need to build mature IT risk management strategies.

Sources

  1. Aon 2025 Global Risk Management Survey
  2. Forrester Business Risk Survey, 2025
  3. Forrester: The State of Enterprise Risk Management, 2025
  4. Hyperproof 2026 IT Risk and Compliance Benchmark Report
  5. Verizon 2025 Data Breach Investigations Report (DBIR)
  6. IBM/Ponemon Cost of a Data Breach 2025 (via CompTIA State of Cybersecurity 2025)
  7. Ethixbase360: Top 10 Third-Party Cyber Breaches of 2025
  8. FortifyData: Third-Party Data Breaches in 2025
  9. Adversa AI 2025 Report
  10. Grand View Research: Risk Management Market Report
  11. IOActive Resources: AI Model Research (730 prompts, 27 languages, 219 vulnerability categories)
  12. NIST Cybersecurity Framework (CSF) 2.0
INSIGHTS | September 17, 2026

Internet-Exposed OT: What the UK Generator Incident Tells CNI Operators About Attack Surface

Key Takeaways

  • A small UK power generator was reportedly forced offline for four days in July 2026. Press reporting attributes the incident to Iran-linked actors, but the UK government has confirmed only that an incident occurred, declining to attribute it or identify the facility.
  • The initial access vector has not been disclosed. Operators cannot map the incident to a specific product or vulnerability, and any vendor claiming otherwise is speculating.
  • CISA’s Internet Exposure Reduction Guidance, revised on 21 August 2026, sets out four steps for identifying and removing unnecessary internet exposure, alongside a discovery port list covering common OT and remote access protocols.
  • CISA observed malicious activity against more than 100 internet-exposed systems in the US Water and Wastewater Systems Sector during July 2026, commonly involving PLCs connected directly to cellular modems.
  • UK guidance and regulation already point in the same direction. The NCSC’s secure connectivity principles for OT and the DESNZ and Ofgem proposals for baseline cyber resilience requirements both target the exposure and boundary weaknesses that recur across these incidents.

Why Does Internet-Exposed OT Matter Now?

Two items arrived within days of each other in late August 2026. CISA revised its Internet Exposure Reduction Guidance on 21 August, anchoring the document in malicious activity it had observed against internet-exposed industrial systems the previous month [5]. The Telegraph reported the following day that hackers linked to Iran had forced a small British power generator offline for four days during July [1]. Britain briefed energy company chief executives on 24 August, and the Department for Energy Security and Net Zero wrote to companies advising them on next steps [2][3].

The proximity is coincidental. The underlying subject is not. CISA describes adversaries reaching controllers through connectivity that asset owners had not catalogued, and a UK generation asset has now lost four days of output to cyber activity. The common factor across both is not a particular adversary or a novel technique. It is attack surface that organisations created themselves and did not measure.

The wider trend line supports that reading. The NCSC’s chief executive told the RUSI Annual Security Lecture in June 2026 that around 75 per cent of cyber activity targeting UK critical infrastructure can be linked to state actors, and that the agency had managed more than 200 incidents affecting CNI and its wider ecosystem over the previous year [13].

For security leaders, the useful question is not whether Iran shut down a British generator. It is whether the organisation can demonstrate, with evidence rather than assumption, what of its own operational technology is reachable from the public internet.

What Is Known About the UK Generator Incident?

The facts that can be stated with confidence are narrow. The Telegraph broke the story on 22 August 2026, reporting that a cyber attack in July had forced a small British generator offline for four days [1]. The Financial Times, BBC and Guardian subsequently published accounts largely derived from the original reporting [4].

The UK government has confirmed that an incident occurred while declining to attribute it or identify the facility. A government spokesperson described the affected asset as a small-scale energy generator and stated that at no point was there a risk to the wider energy system [3]. Michael Shanks, the minister for energy, said the government and industry were treating the incident seriously and were working with regulators and the NCSC to assess threats and strengthen protections, adding that there was no threat to the wider grid and that nobody lost power [2].

Attribution to Iran therefore rests on press reporting rather than on any official assessment. The Iranian Embassy in London did not respond to requests for comment [7]. Little technical detail has emerged through official channels, and the initial access vector remains undisclosed [4].

That gap matters for defenders. Without a published vector, the incident cannot be mapped to a specific vulnerability or product, and no defensive shopping list follows from it directly. What the incident does establish is that a UK generation asset sustained four days of outage arising from cyber activity. That is a materially different proposition from the access-without-effect intrusions that have characterised much reported CNI targeting to date, and it is the part operators should brief upwards.

What Does CISA’s Internet Exposure Reduction Guidance Recommend?

In the absence of a disclosed vector, the most useful available material addresses the attack surface rather than the adversary. CISA’s guidance approaches this from the discovery end, setting out four steps:

  1. Assess current exposure by identifying which assets are reachable via the internet, using scanning tools and services to gain visibility of the organisation’s online footprint.
  2. Evaluate the necessity of that exposure, and remove or restrict access for assets that do not need to be internet-accessible.
  3. Mitigate risks to assets that must remain exposed, through default password changes, patching, replacement of unsupported devices, use of a jump host, ingress and egress monitoring, and multi-factor authentication.
  4. Establish routine assessments so that new exposures are identified as IT and OT environments evolve [5].

The guidance is anchored in observed activity rather than theory. CISA reports that in July 2026 it observed malicious cyber activity targeting more than 100 internet-exposed systems in the US Water and Wastewater Systems Sector, commonly involving programmable logic controllers connected directly to a cellular modem [5]. Threat actors remotely accessed those exposed PLCs, changed device IP addresses and passwords, and caused loss of monitoring and control functionality and, in some cases, operational disruption [5].

The cellular modem detail deserves attention. Connectivity installed by vendors or integrators for maintenance convenience frequently sits outside the corporate network and outside routine attack surface scanning. The controller is directly reachable while the organisation’s own exposure reporting shows nothing [5].

CISA also publishes a port list for discovery work, covering remote access services and OT protocols including Modbus on 502 to 507/TCP plus additional implementation ports, Niagara Fox on 1911/TCP and 4911/TCP, EtherNet/IP on 2222/UDP and 44818/TCP, DNP3 on 19999/UDP and 20000/TCP and UDP, OPC UA on 4840/TCP and 4843/TCP, and BACnet on 47808/TCP [5]. An open port is not itself evidence of compromise, but CISA advises that any internet-accessible OT or remote access service should be investigated and unnecessary exposure removed [5].

One point of discipline applies here. CISA does not attribute the July water sector activity to a named actor in this document. The description does resemble the Iranian-affiliated campaign against internet-exposed PLCs set out in advisory AA26-097A, which was updated on 22 July 2026 to add Schneider Electric and Siemens controllers alongside Rockwell, and which names the water and wastewater sector directly [6]. IOActive examined that update in Iranian-Affiliated Actors Expand PLC Targeting to Siemens and Schneider Electric. Resemblance is not attribution. CISA does not join the two, and operators briefing boards should be precise about the difference between an observed exposure pattern and an attributed campaign. Neither should be joined to the UK incident without evidence.

What Do UK Guidance and Regulation Already Require?

UK operators do not need to wait for American guidance to act, because equivalent direction already exists. The NCSC’s secure connectivity principles for OT address the same weaknesses, covering the requirement that OT devices are not directly exposed to the public internet, that management protocols and the OT boundary are hardened, that secure versions of industrial protocols are adopted where available, and that OT, management and business networks are segmented [8]. A worked example for the water sector was added to the collection on 11 August 2026, the first content authored by the Industrial Control System Community of Interest to appear on the NCSC website [9]. Separate NCSC guidance covers creating and maintaining a definitive view of OT architecture, which is the UK equivalent of CISA’s first step [10].

The regulatory position is also moving. DESNZ and Ofgem published their response on reshaping cyber regulation in downstream gas and electricity on 5 August 2026, confirming an intention to develop baseline cyber resilience requirements for all Ofgem licensees and to review which operators fall within the Network and Information Systems Regulations 2018 [11]. Ofgem will lead development of the detailed proposals, working with DESNZ and the NCSC.

The direction of travel matters for smaller generators in particular. The current regime concentrates obligations on the largest operators, while the incident reported in July involved a small asset. The Cyber Security and Resilience Bill would expand the existing framework, including powers intended to allow ministers to direct regulated organisations to take proportionate action where an imminent or live cyber threat puts national security at risk [12]. Shanks has also indicated that a wider Energy Resilience Strategy is due later in 2026 [7].

Operators should expect the compliance floor to rise. Building an accurate exposure picture now is cheaper than doing it under a regulatory deadline.

Who Is Affected by Internet-Exposed OT Risk?

Electricity generation and distribution operators

The UK incident involved a small generation asset, and that is the substance of the concern rather than a mitigating detail. A distributed generation fleet multiplies near-identical assets running comparable equipment sourced from a small pool of suppliers. A weakness that is inconsequential in one asset compounds when repeated across hundreds [7].

Water and wastewater utilities

CISA’s observation of more than 100 exposed systems in a single sector within one month indicates that exposure is systemic rather than exceptional [5]. Smaller utilities with limited engineering headcount are least likely to hold an accurate inventory of internet-facing assets.

Manufacturers and process industries

Nothing in the exposure problem is specific to CNI designations. Discrete manufacturing, chemicals, food production and building management run the same protocols on the same controllers, frequently with weaker segmentation and no regulatory driver.

Suppliers, integrators and maintenance providers

Cellular modems, remote access appliances and vendor support tunnels are installed for legitimate operational reasons and rarely appear in the asset owner’s exposure register [5].

Boards and audit committees

The question of what the organisation exposes to the internet is an assurance question rather than a technical one, and it is answerable with a number.

What Are the Challenges for Operators?

The published guidance is not technically demanding. The difficulty is organisational and operational.

Verification is harder than assertion. Most operators can state that their OT is segmented. Fewer can evidence that claim against an external view of their own address space, and fewer still have examined infrastructure introduced by vendors, contractors or legacy projects [5].

Patching carries operational and safety consequences. OT frequently runs software that cannot be updated remotely, and operators must weigh intervention against the risk that downtime itself creates a safety issue. Replacement programmes run for years, and exposure cannot simply be accepted in the interim. The workable response is to move carefully when changing the plant and quickly when reducing the risk around it.

Ownership is the recurring failure. Guidance of this kind repeats the same controls because no single person owns the outcome of knowing what the organisation exposes and acting when a scan flags the same default credential six months running.

Detection in OT is more tractable than in IT, and is routinely neglected. OT environments are typically static and predictable, which makes baseline monitoring effective at identifying unauthorised activity and misconfiguration [8]. Few operators exploit that property.

How Can IOActive Help?

The defining feature of this incident is that the vector is undisclosed. There is no patch to apply and no indicator to hunt for. What remains actionable is the attack surface itself, and establishing that is an assessment problem rather than a threat intelligence problem.

IOActive has worked in industrial control environments since building the first proof-of-concept worm against the smart grid in 2009 [14], and its team has contributed to standards and best practice including NIST 800-53 and 800-37. The services below address the gaps this incident and the accompanying CISA guidance expose.

Full Stack Security Assessments

CISA’s first step, and the NCSC’s guidance on maintaining a definitive view of OT architecture, both require an operator to establish what is reachable rather than to assert it [5][10]. Our Full Stack assessments examine the whole environment rather than a single layer, covering internet-facing controllers and HMIs, the OT and IT boundary, the cellular modem paths that frequently give field sites their only route to the internet, and the firmware and silicon of the devices themselves through penetration testing, reverse engineering, side-channel analysis and fault injection. Where an asset has lost four days of output and no vector has been published, that depth is what separates a finding from a reassurance.

Supply Chain Integrity

The activity CISA observed in the water sector involved controllers connected directly to cellular modems, which is connectivity the asset owner frequently did not commission and does not monitor [5]. Our Supply Chain Integrity work assesses the security posture of technology providers and critical third parties, covering firmware, embedded systems, remote access arrangements and procurement processes, so that inherited exposure is identified before it becomes a shared incident. For UK operators this maps onto the critical supplier designation powers in the Cyber Security and Resilience Bill [12].

Advisory Services

The regulatory floor is rising for precisely the class of asset involved here. DESNZ and Ofgem intend to develop baseline cyber resilience requirements for all Ofgem licensees and to review which operators fall within the NIS Regulations 2018 [11]. We examined what that direction of travel means for operators in The UK Energy Sector Cyber Security Strategy. Our Advisory Services cover programmatic security review, security programme development and management, and Virtual CISO support, turning an exposure picture into a prioritised remediation plan with named accountability and the evidence a board needs for CAF or equivalent regulatory assurance.

The water sector offers a useful contrast. While recent attention has centred on attacks against US water infrastructure, coverage of how UK water regulators are responding has been comparatively thin — despite water companies having been in scope of the NIS Regulations since 2018 and required to submit annual OT resilience risk assessments against the NCSC’s Cyber Assessment Framework. Ofwat’s lever here is largely financial: PR24 price control deliverables are tied to the Drinking Water Inspectorate’s regulation 18 notices, with penalties layered on top of any Inspectorate enforcement. What’s less visible, compared to Ofgem’s published guidance for energy OES, is an equivalent public baseline for water operators to work against — sector performance data isn’t published, and companies are largely left to interpret CAF outcomes on their own. That’s exactly the kind of gap our cross-industry ICS/OT experience is built to close.

If you would like to discuss how your organisation’s internet-facing OT exposure measures up against the activity described in CISA’s guidance, we welcome the conversation.

  1. Establish an external view of the estate. Use exposure discovery tooling to enumerate what is reachable from the public internet across known organisational IP space, rather than relying on the internal asset register as the source of truth [5].
  2. Reconcile the external view against the register. Treat every unexplained result as an incident precursor. The gap between the two views is the working measure of unmanaged exposure.
  3. Interrogate vendor and contractor connectivity. Identify cellular modems, support tunnels and remote access appliances installed outside change control, including those commissioned during legacy projects [5].
  4. Remove what is unnecessary and secure what remains. Route required remote access through a managed gateway, firewall or VPN rather than connecting directly to a PLC, HMI or RTU, and enforce unique credentials with phishing-resistant multi-factor authentication [5].
  5. Harden the OT boundary against the NCSC principles. Confirm that boundary devices are within vendor support, that management interfaces are unreachable from the internet, and that segmentation between OT, management and business networks is enforced rather than assumed [8].
  6. Verify controller state across the estate. Confirm that controllers are not left in programming or maintenance modes and that write protection is applied to control logic.
  7. Baseline OT network traffic. Establish what normal communication looks like and alert on attempts to reach controllers and HMIs from unexpected devices, networks or routes [8].
  8. Report exposure to the board as a measurable figure. Present the count of internet-facing OT assets, the proportion with a documented operational justification, and the trend across reporting periods, ahead of the baseline requirements now being developed by Ofgem [11].

Conclusion

The UK generator incident may never be attributed publicly with the confidence operators would prefer, and the vector may never be disclosed. Waiting for either would be a mistake, because the actions that reduce this risk do not depend on knowing who was responsible. Guidance on both sides of the Atlantic describes an attack surface that organisations create for themselves and can therefore measure and reduce for themselves. Four days of outage at a small generator is a manageable consequence for the grid as a whole; for the owner facing the remediation bill, the reputational fallout, and the regulatory scrutiny that follows, it is anything but manageable. The same weakness, repeated across a distributed estate and reached by an actor willing to cause visible disruption, stops being a grid-level inconvenience and becomes a systemic one. That asymmetry is sharper in the US, where water and electricity distribution are served by a much larger population of small, often municipally or cooperatively owned operators: a single incident is unlikely to trouble the wider grid or supply, but it can be existential for the operator itself. Further guidance and regulatory detail are expected through the remainder of 2026, and operators who can already answer the exposure question — at whatever scale they sit — will find the rest considerably easier.

References

  1. The Telegraph, “Iranian hackers shut down UK power plant”, 22 August 2026. https://www.telegraph.co.uk/news/2026/08/22/iranian-hackers-shut-down-uk-power-plant/
  2. Reuters, “UK briefs energy chiefs after Iran-linked cyber attack reports”, 24 August 2026. https://www.reuters.com/business/energy/uk-briefs-energy-chiefs-after-iran-linked-cyber-attack-reports-2026-08-24/
  3. CNBC, “Small UK power generator shut down after cyberattack linked to Iran”, 23 August 2026. https://www.cnbc.com/2026/08/23/small-uk-power-plant-shut-down-after-iran-linked-cyberattack-report.html
  4. SecurityWeek, “Iran-Linked Hackers Shut Down UK Power Plant for Four Days”, August 2026. https://www.securityweek.com/iran-linked-hackers-shut-down-uk-power-plant-for-four-days/
  5. Cybersecurity and Infrastructure Security Agency, “Internet Exposure Reduction Guidance”, revision date 21 August 2026. https://www.cisa.gov/resources-tools/resources/exposure-reduction
  6. CISA, FBI, NSA, EPA, DOE, CNMF and Department of the Treasury, “Iranian-Affiliated Cyber Actors Exploit Programmable Logic Controllers Across US Critical Infrastructure”, AA26-097A, updated 22 July 2026. https://www.ic3.gov/CSA/2026/260722.pdf
  7. Industrial Cyber, “UK power generator reportedly taken offline in Iran-linked cyberattack, raising energy security concerns”, 24 August 2026. https://industrialcyber.co/utilities-energy-power-water-waste/uk-power-generator-reportedly-taken-offline-in-iran-linked-cyberattack-raising-energy-security-concerns/
  8. National Cyber Security Centre, “Operational Technology: Secure connectivity principles”. https://www.ncsc.gov.uk/collection/operational-technology/secure-connectivity
  9. National Cyber Security Centre, “Water sector example added to the NCSC’s Secure connectivity principles”, 11 August 2026. https://www.ncsc.gov.uk/blogs/water-sector-example-added-to-the-ncscs-secure-connectivity-principles
  10. National Cyber Security Centre, “Operational Technology: Creating and maintaining a definitive view of your OT architecture”. https://www.ncsc.gov.uk/collection/operational-technology/definitive-architecture-view
  11. Department for Energy Security and Net Zero and Ofgem, “Whole energy cyber resilience requirements: reshaping cyber regulation in downstream gas and electricity”, government response, updated 5 August 2026. https://www.gov.uk/government/consultations/whole-energy-cyber-resilience-requirements-reshaping-cyber-regulation-in-downstream-gas-and-electricity
  12. UK Government, “Cyber Security and Resilience (Network and Information Systems) Bill: Power to direct regulated entities”, factsheet, updated 30 June 2026. https://www.gov.uk/government/publications/cyber-security-and-resilience-network-and-information-systems-bill-factsheets
  13. Royal United Services Institute, NCSC Chief Executive remarks, RUSI Annual Security Lecture, June 2026. https://www.rusi.org/news-and-comment/rusi-news/hostile-states-behind-75percent-cyber-attacks-uk-infrastructure-ncsc-ceo
  14. M. Davis, IOActive, “Advanced Metering Infrastructure (Smart Grid) Device Security”, Black Hat USA 2009. https://blackhat.com/presentations/bh-usa-09/MDAVIS/BHUSA09-Davis-AMI-SLIDES.pdf
INSIGHTS | September 15, 2026

Cybersecurity Testing Methodologies Compared

Cyber defenses blocking attacks from organizations and infrastructure.

The global average cost of a data breach reached $4.44 million in 2025, the first decline in five years. IBM attributes this largely to faster detection and containment, driven by AI-assisted security tooling and more mature incident response programs.¹ That decline is meaningful, but it does not erase the underlying question that keeps security leaders up at night: if an attacker came for your organization today, would your security program catch it in time?

For buyers comparing cybersecurity testing methodologies, the answer depends entirely on whether the testing approach they have chosen is aligned with the questions they most need to answer. This article compares cybersecurity testing methodologies so security leaders can build testing programs that reflect their risk profile.

Cybersecurity Testing Methodologies Compared

MethodologyPrimary Question AnsweredScopeEngagement DurationStakeholder AwarenessBest Suited For
Vulnerability AssessmentWhat weaknesses exist across our systems?Broad, automatedHours to daysHighBaseline hygiene, compliance inventory
Penetration TestingWhich weaknesses are actually exploitable?Targeted, manualDays to weeksHighTargeted risk reduction, audit requirements
Red Team ExerciseHow would a real attack on our organization play out?Full-scope, adversarialWeeks to monthsLow (exec only)Mature programs; resilience validation
Purple Team ExerciseHow can offensive findings improve our detection?Collaborative, controlledDays to weeksHighDetection engineering; SOC capability building
Adversary SimulationCan a specific known threat actor breach our defenses?Threat-actor specificWeeksLow to mediumHigh-risk sectors; critical asset protection
Security ValidationAre our controls still working the way we think?Continuous, platform-drivenOngoingHighControl drift detection; continuous assurance
Full-Stack Security AssessmentHow does our entire attack surface connect?Comprehensive: hardware, firmware, software, AIWeeks to monthsMediumComplex environments; hardware and embedded systems

Choosing the Right Cybersecurity Testing Methodology

The decision framework below matches organizational characteristics to the testing methodology most likely to answer the most pressing security questions. In practice, mature security programs combine multiple methodologies: vulnerability assessments for continuous hygiene, penetration testing for targeted risk validation, Red Team or adversary simulation for resilience measurement, and full-stack assessment where the attack surface extends into hardware and embedded systems.

If Your Organization…Start WithThen Add
Has no formal testing program yetVulnerability AssessmentPenetration Testing
Needs to meet a compliance requirementPenetration TestingVulnerability Assessment
Has a functioning security team and wants to test detectionRed Team ExercisePurple Team Exercise
Wants to improve SOC detection engineeringPurple Team ExerciseAdversary Simulation
Operates in critical infrastructure or OT/ICS environmentsFull-Stack AssessmentRed Team Exercise
Has hardware, embedded systems, or AI in productionFull-Stack AssessmentPenetration Testing (firmware/AI)
Needs to test against a specific known threat actorAdversary SimulationRed Team Exercise
Needs continuous visibility into control effectivenessSecurity ValidationPenetration Testing (periodic)
Is building a multi-year security testing roadmapPenetration TestingRed Team, then Adversary Simulation

Vulnerability Assessments

A vulnerability assessment is the foundational layer of any security testing program. Automated scanning tools cross-reference an organization’s systems and software against public databases of known vulnerabilities, producing a ranked list of unaddressed weaknesses. The output is fast and broad: an organization can scan thousands of assets within hours.

What a Vulnerability Assessment Actually Measures

The key limitation is that a vulnerability assessment identifies weaknesses without confirming whether they are exploitable in your specific environment. A critical-severity CVE may be present on a system protected by compensating controls, making its real-world impact far lower than its CVSS score implies. Conversely, a medium-severity finding on a system with elevated privileges may represent a far more serious risk than any automated tool will flag.

When a Vulnerability Assessment Is the Right Choice

Vulnerability assessments are most valuable as a continuous hygiene practice and as a pre-engagement baseline before any manual testing. Organizations pursuing compliance frameworks, including SOC 2, ISO 27001, PCI-DSS, and FedRAMP, typically require documented vulnerability assessments at defined intervals. They are also useful after infrastructure changes for quickly identifying new exposures.

Where It Falls Short

A vulnerability assessment cannot confirm whether your security controls would stop a determined attacker. It produces a list of targets, not a realistic picture of how those targets could be chained into a damaging attack path. Organizations that rely on vulnerability assessments as their primary testing methodology have identified their weaknesses without knowing which ones matter most under adversarial conditions.

AttributeVulnerability Assessment
Primary outputRanked list of CVEs and misconfigurations
Human expertise requiredLow to medium
Detection and response testingNone
Compliance valueHigh
Organizational maturity requiredLow
IOActive equivalentPre-assessment scanning embedded in full-scope engagements

Penetration Testing: Proving Exploitability

Penetration testing takes vulnerability assessment to the next level by adding the expert human element. A qualified ethical hacker is authorized to discover and exploit vulnerabilities within a defined scope, demonstrating how a real adversary might compromise systems, access sensitive data, or disrupt operations. The objective is not simply to find weaknesses but to confirm which ones carry genuine damage potential.

What Penetration Testing Measures

IOActive’s penetration testing methodology includes vulnerability scanning, brute-force testing, web application exploitation, and social engineering within the stated scope. Because exploitation is included, results show which vulnerabilities pose the most immediate risk, helping security teams prioritize remediation where it matters most. Unlike vulnerability assessments, penetration testing distinguishes confirmed risk from theoretical exposure.

When Penetration Testing Is the Right Choice

Penetration testing is the right methodology when your organization needs to address specific known risk areas, validate controls before or after a major deployment, satisfy a compliance requirement, or build an internal case for remediation investment. It is also the appropriate entry point for organizations that want to move beyond automated scanning but are not yet mature enough to benefit fully from Red Team exercises.

Where It Falls Short

Because internal stakeholders are generally aware that the test is occurring, penetration testing provides limited insight into detection and response capabilities. Security operations teams may respond differently during a known engagement than during an unannounced real-world attack. Penetration testing is also constrained by its defined scope, meaning attack paths that cross into hardware, embedded firmware, AI pipelines, or physical access controls are typically not evaluated unless the engagement is scoped to include them.

AttributePenetration Testing
Primary outputExploitability-confirmed vulnerability list with remediation guidance
Human expertise requiredHigh
Detection and response testingLimited (scope-dependent)
Compliance valueVery high
Organizational maturity requiredLow to medium
IOActive equivalentApplication, network, hardware, AI, and embedded systems penetration testing

Red Teaming: Testing Resilience, Not Just Vulnerabilities

A Red Team exercise is a full-scope adversarial simulation in which a group of expert ethical hackers emulates a real-world threat actor, attempting to compromise the organization’s most critical assets without detection. Unlike penetration testing, the Red Team operates covertly, its existence known only to select executives. It is not constrained by a predefined list of targets or a scoped attack surface.

What Red Team Exercises Measure

IOActive’s Red Team exercises use threat intelligence and threat modeling to create representative attack scenarios aligned to the organization’s industry, assets, and actual adversary landscape. The exercise tests not just whether vulnerabilities exist but whether your people, processes, and technologies work together to prevent, detect, and respond to a sophisticated, persistent attacker. The output is a realistic resilience scorecard, not a vulnerability inventory.

When Red Teaming Is the Right Choice

Red Team exercises deliver their highest value for organizations with mature security programs: those that already have a functioning vulnerability management process, a dedicated security operations function, and defined incident response procedures. Without this foundation, a Red Team exercise will surface the same basic issues that a penetration test would have identified at lower cost and in less time.

Where It Falls Short

Red Team exercises require significant time and investment, spanning weeks to months. They are also less effective at producing a comprehensive vulnerability inventory: the Red Team’s objective is adversarial impact, not exhaustive coverage. Organizations in earlier stages of security maturity often extract more value from a penetration test before committing to a full Red Team engagement.

AttributeRed Team Exercise
Primary outputAttack narrative, detection gap analysis, resilience scorecard
Human expertise requiredVery high
Detection and response testingComprehensive
Compliance valueMedium
Organizational maturity requiredHigh
IOActive equivalentFull-scope Red Team with threat intelligence, physical access, and social engineering

Purple Teaming: Turning Offense Into Defensive Improvement

Purple Teaming combines the offensive capabilities of a Red Team with the defensive knowledge of the organization’s security operations function in a structured, collaborative exercise. Rather than attacking covertly, the exercise operates transparently: the Red Team executes techniques, the Blue Team observes its own detection capability in real time, and both sides work together to improve visibility and response. The goal is to build detection capability directly from the findings.

What Purple Team Exercises Measure

Purple Team exercises answer the question: “Are our detection controls catching the techniques attackers are actually using against organizations like ours?” The output includes confirmed detections, confirmed gaps, SIEM tuning recommendations, alerting threshold adjustments, and an empirical measurement of SOC capability against specific attack techniques mapped to MITRE ATT&CK.

When Purple Teaming Is the Right Choice

Purple Teaming is the right methodology when your organization wants to improve its security operations capability, validate that new security tooling is correctly configured, or build internal detection engineering skills. It is particularly valuable as a structured follow-up after a Red Team engagement surfaces detection gaps, or as a more cost-efficient approach to testing specific threat scenarios without conducting a full covert Red Team exercise.

Where It Falls Short

Because the exercise is collaborative and transparent, it cannot test true detection capability under realistic conditions where defenders don’t know an attack is in progress. It also requires a sufficiently capable internal security team to be meaningful: organizations without a functioning SOC or detection engineering practice will find Purple Team exercises difficult to execute and difficult to learn from.

AttributePurple Team Exercise
Primary outputDetection gap analysis, control tuning recommendations, SOC capability benchmarks
Human expertise requiredHigh (both offensive and defensive)
Detection and response testingFocused and empirical
Compliance valueMedium
Organizational maturity requiredMedium to high
IOActive equivalentCollaborative red/blue exercises aligned to MITRE ATT&CK

Adversary Simulation: Testing Against the Threats That Target You Specifically

Adversary simulation is a targeted form of Red Teaming in which the testing team emulates the specific tactics, techniques, and procedures (TTPs) of a known threat actor relevant to the organization’s industry or threat landscape. Rather than operating as a generic attacker, the simulation team behaves as a specific threat group would: using the same initial access methods, persistence mechanisms, lateral movement techniques, and exfiltration approaches documented in real-world intelligence reporting.

What Adversary Simulation Measures

Adversary simulation answers the question: “Could the threat actors most likely to target our industry actually succeed against our defenses?” It is grounded in threat intelligence rather than generic attacker methodology, which makes the findings more directly actionable for security operations teams building detections and defenses against specific actor profiles.

When Adversary Simulation Is the Right Choice

Adversary simulation is the right methodology for organizations in high-risk sectors, including financial services, critical infrastructure, government, healthcare, and defense, where the threat actor landscape is well-defined, and the consequences of a successful targeted attack are severe. It is also appropriate for organizations that have completed multiple Red Team exercises and want to test against increasingly specific and sophisticated threat profiles.

Where It Falls Short

Adversary simulation requires high-quality, current threat intelligence to be effective. Simulations built on outdated or generic actor profiles may not reflect the organization’s actual threat landscape. Organizations operating in sectors where actor TTPs evolve rapidly need to ensure simulation frameworks are continuously refreshed to remain accurate.

AttributeAdversary Simulation
Primary outputThreat-actor-specific attack narrative and detection coverage assessment
Human expertise requiredVery high
Detection and response testingComprehensive and threat-specific
Compliance valueMedium to high (sector-specific frameworks)
Organizational maturity requiredHigh
IOActive equivalentThreat-intelligence-led emulation aligned to MITRE ATLAS and ATT&CK

Security Validation: Continuous Testing for Mature Programs

Security validation, often delivered through breach and attack simulation (BAS) platforms, is a continuous or high-frequency testing approach that regularly executes simulated attack techniques against an organization’s controls to confirm they are functioning correctly. Unlike periodic point-in-time assessments, security validation creates an ongoing signal about control effectiveness as the environment changes.

What Security Validation Measures

Security validation answers the question: “Are the controls we deployed still working the way we think they are?” Environments change constantly: new systems are added, configurations drift, patches create unexpected side effects, and threat actor techniques evolve. A control that was effective six months ago may have degraded without anyone noticing. Security validation surfaces that drift before an attacker exploits it.

When Security Validation Is the Right Choice

Security validation is the right methodology for organizations with mature, complex environments where control drift is a real operational risk, particularly those with large distributed infrastructure. It is most valuable as a complement to periodic manual testing, not as a replacement. Point-in-time penetration tests and Red Team exercises provide the depth and creative adversarial thinking that automated platforms cannot replicate; continuous validation fills the coverage gap between those engagements.

Where It Falls Short

Security validation platforms execute known, documented attack techniques. They cannot discover novel attack paths, evaluate physical access controls, test the human element, or assess the creative chained attack scenarios that skilled adversarial testers consistently identify. Treating security validation as a substitute for manual testing creates a significant and predictable coverage gap.

AttributeSecurity Validation
Primary outputControl effectiveness dashboard, drift alerts, trend benchmarks
Human expertise requiredLow to medium
Detection and response testingContinuous, technique-specific
Compliance valueHigh for continuous monitoring requirements
Organizational maturity requiredMedium to high
IOActive equivalentBaseline reference tool within broader assessment programs

Full-Stack Security Assessment: When Surface-Level Testing Isn’t Enough

Most cybersecurity testing engagements operate within a defined layer: application security, network infrastructure, or cloud environments. Full-stack security assessment examines how vulnerabilities and attack paths span multiple layers simultaneously, including hardware, firmware, embedded systems, AI and machine learning pipelines, physical infrastructure, and enterprise applications. It is the methodology appropriate for organizations whose risk surface extends beyond the software-defined perimeter.

Why the Full Stack Matters

Modern enterprise environments are not purely software-defined. Connected devices, industrial control systems, SATCOM terminals, autonomous vehicles, medical equipment, AI-enabled workflows, and physical access infrastructure all create attack paths that single-layer testing does not examine. When an attacker who compromises a network edge device can pivot into an operational technology environment, or when a manipulated AI model can influence production decisions, a network-only penetration test leaves the most consequential attack paths untested.

IOActive’s Full-Stack Advantage

This is where IOActive’s approach diverges from most security testing providers. With physical hardware labs in Seattle, Madrid, and Cheltenham, and 25-plus years of research spanning cloud stacks, connected vehicles, integrated circuits, SATCOM terminals, embedded firmware, and AI systems, IOActive provides full-stack assessments that few firms globally can replicate.² The AMD Sinkclose vulnerability, the Tesla Model Y NFC relay attack, the Raspberry Pi RP2350 challenge win, and the Reviver digital license plate jailbreak are examples of research that reflects this depth across hardware, firmware, and software layers simultaneously.²

When Full-Stack Assessment Is the Right Choice

Full-stack assessment is the right methodology for organizations in manufacturing, automotive, aerospace, critical infrastructure, healthcare, and financial services, where the attack surface includes physical systems, proprietary hardware, embedded software, and operational technology alongside enterprise infrastructure. It is also appropriate for any organization whose AI-enabled applications interact with production systems, financial workflows, or sensitive data at scale.

AttributeFull-Stack Security Assessment
Primary outputCross-layer attack path analysis, hardware and firmware findings, remediation roadmap
Human expertise requiredExceptional, cross-disciplinary
Detection and response testingComprehensive across all tested layers
Compliance valueHigh for regulated industries
Organizational maturity requiredMedium to very high
IOActive equivalentHardware, firmware, AI, embedded, network, application, and physical testing

With 25-plus years of independent security research and physical testing labs on three continents, IOActive brings the adversarial expertise and full-stack depth that complex security programs demand.

Frequently Asked Questions

How often should organizations conduct penetration tests?

Most security frameworks recommend annual penetration testing at a minimum. Organizations with active development pipelines, frequent infrastructure changes, or strict compliance requirements may need semi-annual or continuous testing. IBM’s 2025 breach cost data confirms that faster detection and containment produce measurable cost reduction, which makes regular, rigorous testing a direct financial risk management investment.¹

Is a Red Team exercise always more valuable than a penetration test?

No. A Red Team exercise answers different questions, not better ones. For organizations in early security maturity stages, a Red Team engagement will surface the same basic issues a penetration test would have identified at a fraction of the cost and duration. Red Team exercises deliver their highest value when an organization already has a functioning security operations capability and needs to validate it under realistic, sustained adversarial pressure.

Can a single engagement combine multiple methodologies?

Yes. A well-scoped full-program assessment from a firm like IOActive can combine penetration testing, adversary simulation, hardware assessment, and elements of Red Teaming within a single engagement. The key is scoping around the organization’s most important security questions, not selecting a methodology and fitting the questions to it afterward.

What distinguishes a full-stack security assessment from a standard penetration test?

A standard penetration test evaluates exploitability within a defined scope, typically application code, network infrastructure, or cloud configuration. Full-stack assessment examines how vulnerabilities and attack paths connect across hardware, firmware, embedded software, AI pipelines, physical infrastructure, and enterprise applications simultaneously. It is appropriate for any organization whose risk surface extends beyond software-defined boundaries, which now describes most Global 1000 enterprises with connected devices, operational technology, or AI-enabled business processes.

Sources

  1. IBM. “Cost of a Data Breach Report 2025.” ibm.com/reports/data-breach
  2. IOActive. “IOActive Research Timeline.” ioactive.com/ioactive-research-timeline/
INSIGHTS | September 11, 2026

Worse Than First Reported: What CISA’s Revised Water Sector Numbers Mean for Every Utility

Key Takeaways

  • CISA has confirmed that the July 2026 campaign against US water utilities targeted more than 100 internet-exposed systems across at least 12 states — over three times the roughly 30 Minnesota systems disclosed when the story first broke [2][4].
  • Georgia, Michigan, South Dakota and New Jersey have since confirmed their own incidents, including a precautionary boil-water advisory at a Georgia utility that was lifted after testing showed no water quality impact[5].
  • A second, separate joint advisory (AA26-231A), published August 19, describes threat actors using AI-generated exploitation scripts, disguised as legitimate monitoring tools, to probe internet-exposed Siemens S7 PLCs across six critical infrastructure sectors — reconnaissance and capability development, not yet confirmed disruption [6].
  • CISA followed up on August 21 with exposure-reduction guidance repeating, in near-identical language to April’s advisory, that PLCs should never be reachable directly from the internet through a cellular modem [3].
  • The revised numbers surfaced in a reporting vacuum: the federal CIRCIA rule that would mandate incident reporting for water utilities is still not final, with CISA now targeting September 2026, and at least one major state (Illinois) has no requirement at all for a utility to disclose a hack to the public or to law enforcement [9][2].

Why This Update Matters Now

Our first analysis of this campaign treated the Minnesota incidents as a measurement problem: how much distance sits between a federal advisory being issued and a utility being able to act on it. A month later, CISA has revealed that the measurement itself was wrong. The number of affected systems was not roughly 30. It was more than 100, spread across a dozen states, and the agency is only now saying so publicly [2][4].

That is not a minor correction. It means that for a month, water sector leaders, state regulators and the public were making risk decisions — about disclosure, about board briefings, about whether “this happened in Minnesota, not here” was a safe assumption — based on a picture that undercounted the actual campaign by a factor of three or more. Officials in the other 11 states now confirmed as targets did not get a warning proportional to what was actually happening in their sector, because no one outside the investigation knew the true scope yet.

Layered on top of that, a second, unrelated advisory has emerged describing attackers using AI-assisted tooling to build exploitation scripts against a specific, widely deployed PLC family. Whether or not that activity is connected to the July campaign, it confirms something water sector leaders should plan around regardless of attribution: the tooling gap between a nation-state operator and a lower-skilled one is narrowing, and it is narrowing fastest against exactly the kind of internet-exposed OT this campaign has repeatedly exploited.

What CISA Actually Confirmed on August 26

Reporting by NBC Chicago’s investigative team surfaced a line from CISA’s own guidance that had not previously drawn attention: “In July 2026, CISA observed malicious cyber activity targeting over 100 internet-exposed systems in the Water and Wastewater Systems (WWS) Sector, commonly via programmable logic controllers (PLCs) connected directly to a cellular modem” [2][3]. The Register independently confirmed the same figure and noted it is the first time federal officials have put a specific number on the intrusions, though the campaign still has not been attributed to a named actor or group [4].

The 12 states are still not fully identified in public reporting, but Minnesota, Michigan, Georgia, South Dakota and New Jersey have each had incidents confirmed through state or utility disclosures [4][5]. In Georgia, Clayton County Water Authority reported a temporary disruption to a portion of its operational systems and water service, and issued a precautionary boil-water advisory that was lifted once testing confirmed water quality was unaffected; a nearby utility, Columbus Water Works, reported a separate cyber incident days later [5]. In South Dakota, a utility serving Rapid City reported an incident consistent with the same pattern [5]. The FBI, in a joint public service announcement, described the operational effects bluntly: loss of pressure and flooding, with pressure loss creating the potential for untreated groundwater to seep into distribution pipes [5].

Illinois has not been named among the 12 affected states. NBC Chicago’s reporting draws out why that distinction may be less reassuring than it sounds: the state has no law requiring a water utility to notify law enforcement or the public after a cyber incident. As the report put it, “the bitter reality is, even if they know, we might not” [2]. Illinois is not unusual in this respect. It is a reminder that the 12-state figure reflects what has been detected and voluntarily disclosed, not necessarily the outer bound of what occurred.

Three Federal Advisories in Five Weeks

Operators tracking this campaign have now had to absorb three substantive federal releases in a five-week window, on top of the original April advisory:

  • July 22 — the update to AA26-097A widened observed Iranian-affiliated PLC targeting from Rockwell to include Schneider Electric and Siemens devices, and added detection guidance for tampered logic modules [1].
  • August 19 — the new joint advisory AA26-231A, authored by NSA, CISA, the FBI, the Department of Energy and the EPA, describes threat actors using AI-generated exploitation scripts disguised as legitimate OT monitoring tools to conduct reconnaissance against internet-exposed Siemens S7 Series PLCs (the S7-200 through S7-1500 families) over the S7comm protocol on TCP port 102 [6]. The activity spans Critical Manufacturing, Energy, Water and Wastewater, Chemical, Food and Agriculture and Commercial Facilities [7]. CISA’s own advisory text is careful to frame this as reconnaissance and capability development, not confirmed operational disruption, and notes explicitly that the underlying exposure problem is broader than Siemens equipment alone [6]. No CVE and no indicator set accompany the advisory; the actionable finding, as several outlets have observed, is unglamorous and predates the AI framing entirely — an unauthenticated industrial protocol reachable from the open internet [8].
  • August 21 — CISA’s Internet Exposure Reduction Guidance restated the same core instruction that has now appeared in every release since April: route remote access through a secure gateway, firewall or VPN, never connect a PLC, HMI or RTU directly to the internet, and require unique credentials and phishing-resistant MFA for anything that must remain remotely reachable [3].

None of these three releases describes a new class of vulnerability. Each restates a known architectural failure — direct internet exposure of OT — and observes it being exploited at a scale that keeps expanding. For a sector already short on cybersecurity staff, three federal releases in five weeks is not a cadence that most utilities are resourced to fully absorb, verify against their own asset inventory and act on before the next one arrives.

What This Means for Operators Outside the Named States

Water and Wastewater Operators, Regardless of State

If your state has not been named among the 12, that is evidence about what has been disclosed, not evidence about your exposure. The campaign’s access method — internet-reachable PLCs, frequently connected through cellular modems installed by an integrator or vendor rather than the utility itself — is a function of how a given asset was deployed, not of which state it sits in [2][3].

Critical Manufacturing, Energy, Chemical, and Food and Agriculture Operators

AA26-231A extends the named target list beyond water and wastewater to five additional sectors, all sharing the same underlying pattern: Siemens S7 controllers reachable from the internet, with reconnaissance activity that could evolve into the same operational effects seen in the water sector campaign [6][7].

State Regulators and Legislators

The gap in Illinois — no requirement for a utility to disclose an incident to the public or to law enforcement — is not an isolated gap. It is a preview of what a September 2026 CIRCIA final rule is meant to close nationally, and a reminder that until it takes effect, disclosure in most states remains voluntary [9][2]. Officials weighing state-level reporting requirements now have a concrete, recent example of how much of a real campaign’s scope can remain unknown for a month under the status quo.

System Integrators and OT Suppliers

Cellular modems installed for remote telemetry, and engineering software used for legitimate configuration, both continue to be the access points described across every advisory in this campaign. Where an integrator made the connectivity decision, the resulting exposure is frequently outside the asset owner’s own visibility, and it is the integrator’s documentation, not the utility’s network diagram, that will show it [1][3].

What Are the Practical Challenges for Operators?

A Threefold Undercount Changes the Board Conversation

A leadership team or governing board that was briefed on “a Minnesota incident” in late July is now working from materially different facts. Whatever risk posture, budget request or public statement was built on the original scope should be revisited against the confirmed 100-plus systems and 12-state footprint, not left as originally briefed [2][4].

AI Lowers the Barrier to Building Working OT Exploitation Tooling

AA26-231A’s most consequential detail is not the Siemens targeting itself but the method: publicly available device documentation, combined with an AI coding assistant and open-source libraries, was reportedly sufficient to produce functional scripts capable of reading and writing PLC memory over a known protocol [6][7]. That is a capability that previously required specialized ICS tradecraft. It does not change what the correct mitigation is — the same segmentation and exposure-reduction controls apply — but it does mean the population of actors capable of acting on an exposed asset is larger than it was when the April advisory was written.

Disclosure Remains Voluntary in Most Places, Which Means Scope Estimates Should Be Treated as Floors

With CIRCIA’s mandatory reporting rule still pending and state-level requirements inconsistent, the 12-state, 100-system figure should be read as the confirmed floor of the campaign’s scope, not its ceiling. Operators and officials should plan on the assumption that additional incidents have occurred and gone unreported, rather than treating the absence of their state’s name as reassurance [9][2].

Advisory Fatigue Is Now a Documented Pattern, Not a Hypothesis

Our first analysis raised advisory fatigue as a risk. It is no longer speculative. Three federal releases inside five weeks, each restating the same architectural fix, is direct evidence that the constraint is not awareness — it is the engineering and coordination capacity to act on what has already been published.

How IOActive Can Help

The revised scope of this campaign does not change the underlying gaps our first analysis identified — assessment boundaries that stop short of remote, cellular-connected assets, and the absence of a validated baseline for controller logic. It does change the urgency, the number of sectors in scope, and the case for treating “we weren’t named” as insufficient reassurance. The services below map directly to what this update reveals.

Full Stack Security Assessments

With six sectors now named across two advisories, and internet-reachable PLCs confirmed as the common access point regardless of state or industry, the priority is establishing ground truth about your own estate — not relying on the absence of your name from a news story. Our Full Stack Security Assessments scope beyond the network and application layer to the facility and silicon level, verifying every internet-facing PLC, HMI and RTU against the live environment, including the cellular-connected remote assets that integrator-built architectures routinely place outside an asset owner’s visibility.

This is not a hypothetical exercise for our team. In a recent security advisory, IOActive researcher Ethan Shackelford identified and disclosed multiple vulnerabilities in the KUNBUS Revolution Pi, a DIN-rail industrial PC widely deployed for automation and process control — the same category of device implicated in this campaign, just from a different vendor [12]. The findings included an authenticated command injection in the device’s web management interface that led to arbitrary code execution (CVE-2024-8684), a directory traversal exposing system files (CVE-2024-8685), and a years-out-of-date sudo binary that allowed local privilege escalation to root (CVE-2021-3156). It is the same underlying pattern the CISA advisories describe: a web-exposed administrative interface on an industrial controller, reachable further than intended, with unsanitized input standing between an authenticated session and full device compromise. KUNBUS fixed all three findings following coordinated disclosure. This kind of assessment — treating the PLC’s own management interface as an attack surface, not just its network position — is exactly what a Full Stack engagement is built to find before an adversary does.

Red Team and Purple Team Services

AA26-231A describes reconnaissance and capability-building against a named protocol (S7comm) using AI-assisted tooling built from public documentation [6]. Our Red Team engagements can emulate that exact tradecraft against your environment — including AI-assisted script generation aimed at your specific controller inventory — while our Purple Team work confirms whether your team would detect it before it progresses from reconnaissance to the disruption pattern already documented at Minnesota, Georgia and South Dakota utilities.

Supply Chain Integrity

The access method across every advisory in this campaign is a cellular modem or remote-access path installed by a vendor or system integrator. Our Supply Chain Integrity service reviews the security posture and connectivity decisions of the third parties who built your control system architecture, closing the visibility gap before it becomes the next disclosed incident.

Advisory Services

With CIRCIA’s final rule now targeted for September 2026 and state disclosure requirements inconsistent in the interim, utilities and their counsel need a clear-eyed view of what reporting obligations already apply and what is coming. Our Advisory Services — programmatic security review, security program development, and Virtual CISO support — help translate this campaign’s revised scope into a board-ready risk picture and a prioritized remediation plan, rather than a reactive response to the next news story.

AI/ML Security Services

As attackers increasingly use AI coding assistants to generate reconnaissance and exploitation tooling, defenders benefit from the same fluency in how that tooling behaves. Our AI/ML Security threat-modeling work extends to evaluating your exposure to AI-assisted attack techniques, complementing the OT-specific assessments above rather than replacing them.

Training

Many of the utilities affected in this campaign are small systems without dedicated cybersecurity staff. Our staff augmentation and Virtual CISO offerings let a resource-constrained utility or municipal IT department bring in the OT security expertise needed to act on these advisories without a multi-year hiring cycle.

1. Re-brief leadership using the confirmed scope, not the original one. If your board or council was told this was a Minnesota-only event, correct the record before the next report cycle.

2. Treat your state’s absence from the 12 as unconfirmed, not clean. Verify your own internet-facing PLC, HMI and RTU inventory directly rather than inferring safety from news coverage.

3. Extend hunting to AA26-231A’s specific indicators. Look for engineering software or S7comm traffic (TCP port 102) originating from unexpected hosts, particularly if you operate Siemens S7-series controllers [6].

4. Confirm your remote-access architecture routes through a managed gateway. Direct PLC-to-internet paths via cellular modem remain the confirmed access method across every release in this campaign [2][3].

5. Check your state’s disclosure requirements now, before an incident forces the question. Do not assume CIRCIA’s federal rule already applies; it is not yet final [9].

6. Document what you would report and to whom, under both current voluntary norms and CIRCIA’s proposed 72-hour and 24-hour timelines, so the mechanics are already in place before September’s rule finalization.

7. Rehearse detection of the manipulation scenario, not just the outage scenario, as our first analysis recommended — this campaign’s revised scope makes that exercise more urgent, not less.

Conclusion

The story a month ago was that a warning arrived first and a sector-wide incident followed four days later. The story now is that the incident itself was three times larger than anyone said publicly for the following month. Both facts point at the same underlying constraint: the water sector’s capacity to detect, verify, and disclose is not keeping pace with either the scale of what is being attempted against it or the speed at which the tooling to attempt it is becoming available. A federal reporting mandate is coming, but it is not here yet, and in its absence, the confirmed numbers in any advisory should be read as a floor.

If you would like to discuss how your organization’s controller exposure measures up against the activity described here — across water, energy, manufacturing, chemical or food and agriculture — or how IOActive can support your OT/ICS resilience program, we welcome the conversation.

References

[1] CISA, FBI, NSA, EPA, DOE, CNMF and US Department of the Treasury. Iranian-Affiliated Cyber Actors Exploit Programmable Logic Controllers Across US Critical Infrastructure (AA26-097A), published April 7, 2026, updated July 22, 2026. https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-097a

[2] C. Goudie, L. Capitanini, N. Halder. Government admits water system cyberattacks were worse than first reported, NBC Chicago, August 26, 2026. https://www.nbcchicago.com/investigations/government-admits-water-system-cyberattacks-were-worse-than-first-reported/3980961/

[3] CISA. Internet Exposure Reduction Guidance, published August 21, 2026. https://www.cisa.gov/resources-tools/resources/exposure-reduction

[4] J. Lyons. More than 100 water systems were hit in July cyberattacks, The Register, August 26, 2026. https://www.theregister.com/cyber-crime/2026/08/26/more-than-100-water-systems-were-hit-in-july-cyberattacks/5292685

[5] J. Greig. Cyberattacks on water systems expand to 12 states as South Dakota, Georgia announce incidents, The Record (Recorded Future News), August 5, 2026. https://therecord.media/iran-cyberattacks-water-treatment

[6] NSA, CISA, FBI, DOE and EPA. Defending Against an Active Threat to Siemens S7 Series PLCs (AA26-231A), August 19, 2026. https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-231a

[7] AI-Generated Exploit Scripts Target Siemens S7 PLCs in U.S. Critical Infrastructure, The Hacker News, August 2026. https://thehackernews.com/2026/08/ai-generated-exploit-scripts-target.html

[8] CISA, NSA, FBI warn of Siemens S7 PLC exploitation using AI-generated scripts to disrupt critical industrial processes, Industrial Cyber, August 2026. https://industrialcyber.co/industrial-cyber-attacks/cisa-nsa-fbi-warn-of-siemens-s7-plc-exploitation-using-ai-generated-scripts-to-disrupt-critical-industrial-processes/

[9] Federal News Network. CIRCIA, other big cyber rules expected to get finalized this fall, July 10, 2026. https://federalnewsnetwork.com/cybersecurity/2026/07/circia-other-big-cyber-rules-expected-to-get-finalized-this-fall/

[10] Testimony before the US Senate Committee on Environment and Public Works. The Cybersecurity State of the Water Sector, February 4, 2026. https://www.epw.senate.gov/public/_cache/files/5/3/53d93dfe-ed8c-4e48-b705-0b16cb92c90a/FA38A32EE48D7F3D4B0C069FDC08A69E6327376EB85838A03310B2E91C8582C1.02-04-2026-dr.-simonton-testimony.pdf

[11] House of Lords Library. Cyber Security and Resilience (Network and Information Systems) Bill: HL Bill 32 of 2026–27, June 2026. https://lordslibrary.parliament.uk/research-briefings/lln-2026-0032/

[12] IOActive Security Advisory. KUNBUS Revolution Pi – Multiple Vulnerabilities (CVE-2024-8684, CVE-2024-8685), discovered by Ethan Shackelford, published March 28, 2024, CVE IDs added February 11, 2025. https://www.ioactive.com/wp-content/uploads/2025/05/IOA-SecAdvisory-KUNBUS-Revolution-Pi.pdf

[13] IOActive. When the Advisory Arrives First: Minnesota’s Water Utilities and the Limits of Warning, August 10, 2026. https://www.ioactive.com/when-the-advisory-arrives-first-minnesotas-water-utilities-and-the-limits-of-warning/

INSIGHTS | September 9, 2026

Secure Software Development Lifecycle Practices

IBM’s 2026 Cost of a Data Breach Report found the global average breach cost reached USD $4.99 million, a record high driven by AI-powered attacks up 56% year over year.¹ Supply chain attacks compound the exposure: ReversingLabs research confirmed software supply chain attacks grew 1,300% over three years.² The Log4j vulnerability alone generated over 10 million attack attempts per hour at peak exploitation.³ Organizations that treat security as a final-stage gate accumulate deferred risk with every release.

Secure software development lifecycle practices embed security controls at every phase of the software development process, reducing remediation costs and organizational risk.

“The most effective secure software development lifecycle practices address security before software is deployed, not after vulnerabilities are discovered.”
– IOActive Security Team

Secure SDLC Best Practices: Controls, Outputs, and Metrics

Best PracticeWhat It DoesKey OutputSuccess MetricsRisk If Skipped
Security Requirements MappingTranslates compliance and business requirements into testable acceptance criteria before development beginsDocumented security requirements and acceptance criteria% requirements with security acceptance criteria; compliance gap rateMissing controls, regulatory gaps, late-stage rework
Threat Modeling (STRIDE/DREAD)Identifies attack paths in planned architecture using structured frameworks before a line of code is writtenThreat register, attack surface map# threats identified per review; % mitigated before development startsArchitectural flaws reach production with costly remediation
Static Code Analysis (SAST)Scans source code for insecure patterns at commit or build timeClean code reports, flagged findings# high/critical findings per 1,000 lines of code; mean time to remediateInjection flaws, hardcoded secrets, insecure patterns shipped
Dynamic Application Security Testing (DAST)Tests running applications against real attack patterns pre-releaseValidated vulnerability findings# runtime vulnerabilities found; % of API endpoints testedRuntime and authentication flaws missed before release
Software Composition Analysis (SCA) + SBOMInventories open-source dependencies and tracks known CVEs; SBOM documents every component in the buildCVE inventory, Software Bill of Materials% dependencies on a supported version; # libraries no longer actively maintainedVulnerable third-party libraries deployed; supply chain exposure
Artifact Signing and Configuration HardeningSigns build artifacts cryptographically and audits deployment configurations to prevent tamperingSigned artifacts, hardened deployment configurations% of artifacts signed; # misconfigurations resolved at deploy gateSupply chain tampering, misconfigured production environments
Vulnerability Management and SLA EnforcementPrioritizes and remediates vulnerabilities by CVSS score, production reachability, and CISA KEV statusRemediated CVE backlog, SLA compliance metricsMean time to remediate (MTTR) by severity tier; % SLA complianceUnpatched known-exploited vulnerabilities accumulate and expand breach risk
Security Governance and Program MaturityEstablishes named ownership, enforceable security gates, and maturity benchmarks aligned to NIST SP 800-218 or OWASP SAMMMaturity score, named control owners, audit evidence# controls with named owners; % security gates enforced in CI/CD; maturity level progressionControls exist on paper only; programs consistently underperform without governance infrastructure

Gaps at any single phase propagate forward. A vulnerability introduced at requirements and missed through design, development, and testing arrives in production fully formed and significantly more expensive to remediate.

Run Threat Modeling Before Writing Code

The highest-leverage security activity happens before a single line of code is written. Threat modeling, using frameworks such as STRIDE or DREAD, gives teams a structured method for identifying attack paths in planned architecture and building in controls before architectural decisions become costly to reverse. A research-fueled Secure Development Lifecycle engagement brings genuine adversarial thinking to this process, applying real-world offensive experience to surface the systemic risks that automated tools routinely overlook.

Threat Modeling Method Comparison

MethodWhat It DoesWho Runs ItWhat It ProvidesEffort LevelBest Fit
STRIDECategorizes threats by class: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of PrivilegeSecurity architects, senior developersComprehensive attack class mapping across trust boundariesMedium — 1 to 2 days per system reviewApplication architecture and trust boundary analysis
DREADScores threats on Damage, Reproducibility, Exploitability, Affected users, and Discoverability to produce a ranked risk listSecurity engineers, risk teamsPrioritized threat list by exploitability and business impactLow to medium — hours per threat registerRisk-ranking identified threats for remediation sequencing
Attack TreesDecomposes the path to a specific high-value target into branches of possible attack methodsPenetration testers, security researchersTargeted attack path analysis for critical or high-value componentsHigh — days to weeks for complex targetsComponent-level analysis of specific high-value targets

Method selection depends on the question the team needs to answer. STRIDE maps attack classes broadly across a system; DREAD narrows findings to prioritized risk; Attack Trees decompose the path to a specific high-value target. Most mature programs apply more than one method, selecting based on the architecture under review and the risk questions at hand.

Integrate Code Analysis and Supply Chain Controls Into Every Pipeline Stage

Static analysis (SAST) catches insecure code patterns during development, while dynamic testing (DAST) validates a running application against real attacker behavior. Supply chain exposure requires its own layer of controls. Sonatype’s research found open-source supply chain attacks growing at 742% annually.⁴ A dedicated Supply Chain Integrity assessment applies attacker-mindset analysis from silicon through the application layer. A 2025 third-party assessment of Microsoft’s Signing Transparency service confirmed strong implementation security and identified defense-in-depth improvements.⁵ Core controls include SBOM management, SCA in CI/CD pipelines, dependency pinning, and artifact signing.

Code Analysis and Supply Chain Controls

ControlWhat It DoesWho Runs ItIntegration PointRisk Addressed
SAST in CI/CDScans source code for insecure patterns automatically at commit or mergeDeveloper pipelines, security championsGit pre-commit hook or CI pipelineInsecure patterns caught before code reaches the main branch
DAST pre-releaseSends attack traffic to a running application to surface runtime flawsQA teams, security engineersStaging environment or pre-release gateRuntime and authentication flaws not visible to static analysis
SCA in build pipelineScans open-source dependencies against CVE databases with each buildDevOps, security teamsBuild pipeline or dependency resolution stepKnown CVEs in third-party libraries before they reach production
SBOM + artifact signingGenerates a complete bill of materials and cryptographically signs build outputs for every releaseDevSecOps, release engineeringRelease pipeline or deployment gateSupply chain tampering and untracked third-party components

Layering all four controls closes distinct gaps at different pipeline stages. SAST and DAST address first-party code at different points in the lifecycle. SCA and SBOM management address third-party and supply chain risk at the build and release stages. Each layer reinforces the others: SAST catches what code review misses, DAST catches what SAST cannot see at runtime, and SCA flags risks that originate entirely outside the team’s codebase.

Prioritize Vulnerabilities With Contextualized SLA Frameworks

Datadog’s 2025 DevSecOps research found that the median software dependency is 215 days behind its latest major version, with one in two services running libraries no longer actively maintained.⁶ Effective vulnerability management requires contextualized prioritization based on CVSS score and production reachability, SLA-based ownership, and regression testing to confirm fixes hold. Listings on the CISA Known Exploited Vulnerabilities catalog indicate confirmed real-world exploitation and warrant accelerated remediation timelines regardless of the base CVSS score.

Vulnerability Remediation SLA Framework

Priority TierCriteriaTarget SLA
CriticalCVSS 9.0+, internet-exposed, CISA KEV listed24–72 hours
HighCVSS 7.0–8.9, production reachable7–14 days
MediumCVSS 4.0–6.9, limited exposure30–60 days
LowBelow 4.0 or non-production90 days

SLAs paired with named ownership and enforced timelines convert a priority framework into a functioning remediation program. Without that operational infrastructure, vulnerability backlogs expand faster than teams can close them.

Establish Named Ownership and Track Security Program Maturity

Governance gaps are the most common point of failure in secure SDLC programs. Controls require named ownership and enforceable gates; programs without both consistently underperform regardless of tooling investment. The NIST Secure Software Development Framework (SP 800-218) and OWASP SAMM provide structured maturity models for building and measuring program progress.⁷ Red Team and Purple Team exercises close the feedback loop, testing whether development-phase controls hold under real adversarial simulation and informing the next cycle of program improvement.

Assess Your Current Maturity Level and Close the Gaps Systematically

Maturity LevelCharacteristicsNext Steps to Advance
Ad HocReactive, no formal process, incident-driven patchingAssign one control owner per business unit; document current state versus target state; implement a vulnerability tracking system
DefinedBasic security gates, SAST in CI/CD, vulnerability trackingAdd threat modeling at the design phase; adopt SLA-based remediation; integrate SCA into the build pipeline
ManagedThreat modeling, SCA, SLA-driven remediation, metrics reportingAutomate artifact signing; run quarterly adversarial simulations; establish an SBOM management program
OptimizedAutomated supply chain controls, adversarial simulation, SBOM managementIntegrate continuous Red Team feedback into the SDLC; benchmark the program annually against NIST SP 800-218 or OWASP SAMM

Most organizations enter at Ad Hoc or Defined. Identifying the gaps between the current state and the next level, then addressing them incrementally, produces durable outcomes. Each level builds on the controls established in the previous one, making systematic progression the most reliable path to program improvement.

Build Security Into Your Development Process With IOActive

Secure software development lifecycle practices address vulnerabilities at the phase where remediation is most efficient: before the code ships. Organizations that build security into design, development, and deployment produce software that holds up under real-world adversarial pressure and satisfies the compliance obligations regulators and insurers continue to tighten. IOActive’s 25+ years of research-driven assessments, landmark vulnerability disclosures, and straight-talk approach give clients an honest picture of where their program creates risk.

Sources

  1. IBM Security. Cost of a Data Breach Report 2026. https://www.ibm.com/reports/data-breach
  2. ReversingLabs. The State of Software Supply Chain Security 2024. https://www.reversinglabs.com/sscs-report-2024
  3. IOActive. Supply Chain Integrity. https://www.ioactive.com/service/supply-chain-integrity/
  4. Sonatype. State of the Software Supply Chain, 10-Year Look. 2024. https://www.sonatype.com/state-of-the-software-supply-chain/2024/10-year-look
  5. IOActive. Security Assessment of Microsoft’s Signing Transparency. 2025. https://www.ioactive.com/wp-content/uploads/2025/10/Microsoft-Signing-Transparency-Service-Security-Assessment-IOActive-Public-Facing-Report.pdf
  6. Datadog. State of DevSecOps 2025. April 2025. https://www.datadoghq.com/state-of-devsecops-2025/
  7. NIST. Secure Software Development Framework (SSDF) SP 800-218. https://csrc.nist.gov/projects/ssdf
INSIGHTS | September 3, 2026

The DRM Flag That Isn’t DRM

SetWindowDisplayAffinity makes a window disappear from screenshots, screen shares, and Recall snapshots. Vendors sell that as “screenshot protection,” and procurement checklists tick it off as data-exfiltration risk mitigated. Microsoft’s own documentation for the API says otherwise. This post breaks down what the flag actually guarantees, who can route around it and how, and why a black screenshot is the beginning of a threat model rather than the end of one.

The Pitch, and the Problem with It

Open a modern secure-messaging app, password manager, or exam browser on Windows 11, hit PrtSc, and paste. The result is a black rectangle, or nothing at all. Vendor marketing calls it “screenshot protection.” Privacy blogs call it “blocks Recall.” The procurement checklist gets a tick next to data exfiltration: mitigated.

Microsoft’s own documentation for the API doing the work says otherwise:

Unlike a security feature or an implementation of Digital Rights Management (DRM), there is no guarantee that using SetWindowDisplayAffinity … will strictly protect window content.

The feature that makes a window disappear from screenshots is explicitly not a security feature, according to the people who built it. Yet vendors build a whole category of “privacy” and “DLP” features on top of it.

Microsoft is not consistent about this either. The Recall management documentation, aimed at developers whose remote desktop clients lack screen capture protection, calls adding it “an easy feature,” labels it “This DRM flag,” and points them to the very same SetWindowDisplayAffinity API whose own reference page insists it is not DRM. Two documents, one API, opposite claims.

That gap, between what the control implies and what it guarantees, is what this post takes apart. Attackers route around it without much thought. Defenders keep inheriting it as a checkbox someone else already ticked.

How the Flag Works

The Win32 function doing the work:

BOOL SetWindowDisplayAffinity(
  [in] HWND  hWnd,      // top-level window, must belong to the calling process
  [in] DWORD dwAffinity // the exclusion mode
);

Three values for the dwAffinity parameter matter:

ConstantValueBehavior in a capture
WDA_NONE0x00000000No restriction. Normal capture.
WDA_MONITOR0x00000001Window shows only on a physical monitor; captures render it black.
WDA_EXCLUDEFROMCAPTURE0x00000011Window shows only on a physical monitor; captures omit it entirely (no suspicious black box).

WDA_EXCLUDEFROMCAPTURE is the newer, “better” flag. It arrived in Windows 10 Version 2004 (build 19041). Before that, WDA_MONITOR was the only option, and it left a tell-tale black rectangle. The upgrade is cosmetic from a defense standpoint: black box versus empty space. The security boundary is identical.

The difference is in what the capture comes back with. Under WDA_MONITOR, a screenshot or a screen share contains a black rectangle sitting exactly where the window is, and whatever the window overlaps is hidden along with it. Anyone looking at that capture learns that something was being withheld, how big it was, where it sat, and, in the case of a recording, how long it stayed open and when it closed. Under WDA_EXCLUDEFROMCAPTURE, the window is not in the frame and the desktop behind it shows through, so the capture looks like the application was not running at all. The user sharing their screen has no black box to explain, and the people watching get no cue that anything was hidden.

The Desktop Window Manager (DWM) enforces the exclusion. It is the compositor that assembles every window into the final image on screen. When a capture tool asks DWM for a frame of the desktop, DWM builds that frame and leaves the flagged window out of it. The pixels still reach the physical display; they never reach the composited frame handed to the capture tool. The flag embeds nothing in the window’s content and blocks no capture tool from running. Every bypass later in this post is a version of the same idea: get the image from somewhere other than DWM’s composited output, and the exclusion never applies.

Why Developers Reach for It Anyway

It’s an attractive control because:

  • It’s one line of code. No kernel driver, no service, no secure enclave.
  • It’s OS-native. No third-party dependency to vet.
  • It defeats the lazy attacker. PrtSc, Snipping Tool, Zoom/Teams/Meet screen share, OBS via the standard desktop-duplication path: all come up empty. Against a casual insider or an over-eager AI screenshotter, that’s a win.

The Signal case is the honest version of the story. When Microsoft shipped Recall, a background feature that silently snapshots the screen every few seconds into a searchable database, Signal had no developer-facing opt-out to keep chats out of the index. So, Signal set the display-affinity flag on its window. Signal’s own engineers described it, more or less, as a “one weird trick”: abusing a media-protection flag because Microsoft gave privacy apps no proper API. That’s a defensible decision against that specific threat, an OS feature capturing through the normal compositor path. It is not a general-purpose confidentiality control, and Signal never claimed it was.

The failure mode is the marketing leap: a product that “defeats normal screenshots” gets sold as one that “protects sensitive data,” with no threat model in between.

Who This Does Not Stop

Every control ever shipped can be bypassed, so the useful question is who can bypass this one, holding what access. Three capability tiers cover it.

Tier 0: The User Who Avoids the Blocked Path

Even with zero special access, plenty of capture paths never touch the DWM-composited surface the flag protects.

  • The analog hole: a phone camera pointed at the monitor. The pixels that reach a retina reach a camera sensor the same way. No API closes this, and Microsoft’s docs concede as much.
  • Context that breaks DWM: the protection only works while DWM is composing the desktop, a limit Microsoft states in its documentation. Remote Desktop sessions disable DWM, so a window invisible to a local screenshot can render perfectly over RDP. Certain remote-assistance and mirroring stacks, and some virtual-display configurations, land in the same bucket. The control silently fails open, which is the worst way for a control to fail.
  • VM quirks: in basic VMs without GPU acceleration, the compositor path can differ enough that the exclusion doesn’t behave as advertised.

None of these require privilege escalation. They require not using the one capture method the flag was designed to block. Tier 0 is where most real-world leakage happens, and it leaves almost nothing behind on the host.

Tier 1: The Local User Willing to Run Code

The window belongs to a process, and processes on a user’s own machine are not a trust boundary against that user. IOActive consultant Taha Draidia recently published the concrete version of this in Signal Windows Desktop: contentProtection Bypass, which takes apart Signal Desktop’s screen-capture protection. Signal reaches this same Win32 call through Electron’s setContentProtection() wrapper, so the write-up doubles as a case study in what the flag is worth.

Flip the flag back from inside the process. Draidia tested the obvious approach first: call SetWindowDisplayAffinity(hwnd, WDA_NONE) on Signal’s window from another process. That fails with ERROR_ACCESS_DENIED, and running elevated fails the same way, because the kernel check compares the caller’s process identity against the window’s owner rather than its privilege level. Administrator rights do nothing here. CreateRemoteThread into Signal’s own process satisfies the check, and the protection turns off with no error, no prompt, and nothing on screen to mark the change.

Capture below the compositor. DWM removes the window while assembling the desktop image, so the removal exists only in the copy DWM hands out. Capture that reads frames lower in the stack and closer to the hardware never receives that copy, and may see the window intact. The flag protects one rendering path rather than the content.

This explains where the protection breaks down, not how to build something that breaks it. Anyone who can run code in the user’s session can neutralize the flag, and the barrier is measured in API calls rather than in exploit development. The control is therefore exactly as strong as whatever stops code execution in that session.

Tier 2: Kernel, Driver, or Physical Access

The flag does not apply here at all. DWM checks the exclusion while it builds the desktop image, so code running at or below the display driver gets the pixels without that check ever happening. Independent kernel-mode research makes the point from the other direction: the proof-of-concept driver DWMShield skips the public API entirely and calls the undocumented internal routine GreProtectSpriteContent directly, passing a target window handle over an IOCTL from a non-elevated client. It reaches the same DWM enforcement point Draidia’s work identified, but from underneath the ownership check rather than by satisfying it — the mirror image of the CreateRemoteThread approach in Tier 1, and a separate piece of research rather than an extension of it.

The Actual DRM, for Comparison

The irony in the title is that real DRM exists on the same platform and works on a different principle.

Hardware-backed protected media paths (Widevine L1, PlayReady SL3000, FairPlay) decrypt and composite content inside a Trusted Execution Environment, a secure media path that user-mode and often kernel-mode capture cannot reach. That’s why screen-recording a premium streaming-video service yields a black frame even with admin rights: the pixels never exist in a framebuffer the OS will hand out.

The flag that isn’t DRM, side by side with the DRM that is:

 SetWindowDisplayAffinityHardware DRM (protected media path)
Enforced byDWM composition, kernel-side owner check on the flagSecure hardware / TEE
Where the viewable image livesNormal framebuffer; omitted only from the copy handed to captureInside the TEE; never in a framebuffer the OS can hand out
What it coversOne top-level window at a time, per HWND, re-applied for every new windowThe content stream itself, wherever it plays
Stops normal screenshotsYesYes
Kept out of Recall snapshotsYesYes (Microsoft: Recall won’t store DRM content)
Stops a local user with adminNo (admin enables injection)Yes
Survives process injectionNoYes
Survives RDP / DWM-off contextsNo (fails open)Yes
Stops a phone cameraNoNo
Who can disable itAny code running inside the owning processNo software path; requires defeating the hardware
How it failsOpen and silent: no error, no log, no visual changeClosed: the license refuses to bind, playback stops or drops quality
Cost to adoptOne API call per window, no licensingDevice certification, license server, key management; SL3000 is device-only
Microsoft’s own classification“Not a security feature or DRM”Actual content protection

SetWindowDisplayAffinity is a hardening measure against opportunistic capture. Hardware DRM is a confidentiality control. Treating the flag as a confidentiality control is where the false sense of security begins.

What Developers Should Do

Using SetWindowDisplayAffinity is reasonable. Products go wrong when they treat that one API call as the finished control.

Set it on every top-level window, not just the main one. Affinity is a per-HWND property and every new window starts at WDA_NONE. Dialogs, tooltips, context menus, toasts, and the separate windows that WPF popups and Electron render into each get their own HWND. Flag the main window but not the dialog, and the screenshot catches the secret in full while the ordinary window behind it is the part that gets hidden.

Check the return value. The call returns FALSE on a window that isn’t top level or doesn’t belong to the calling process, and a silent failure still looks protected in code review. Treat it as a security event. On builds older than 19041, WDA_EXCLUDEFROMCAPTURE succeeds and behaves as WDA_MONITOR, so the window turns black instead of vanishing and the API never mentions the difference.

Re-read the flag. GetWindowDisplayAffinity reads the current value from any process, so the app or a monitoring agent can poll it. A change to WDA_NONE the app didn’t make means something else is writing to its process: log it, alert on it, and consider blanking the view until the app can verify its own state.

Use the supported control when one exists. Recall now has real policy: Allow Recall to be enabled (AllowRecallEnablement) and Turn off saving snapshots for Recall (DisableAIDataAnalysis), and managed devices have it removed by default. The flag was a workaround for consumer machines with no opt-out, which is still where it earns its keep.

Document the threat model. Name what the feature stops: screenshot tools, screen sharing, OS-level snapshotting. Name what it doesn’t: cameras, remote sessions, code running in the user’s session, anything at kernel level. Microsoft’s Azure Virtual Desktop documentation is the model to copy: it states plainly that the feature isn’t DRM-level protection and isn’t a substitute for one, and recommends pairing it with other controls.

Pair it with content-level controls. Reveal-on-tap for secrets, short display timeouts, redaction by default, per-session watermarking.

What Defenders Should Monitor

You cannot stop capture on a machine the user controls. You can often catch the attempt.

Injection into the protected app: the Tier 1 bypass is a common and ordinary injection, and the Signal bypass used CreateRemoteThread, the loudest option available. Watch Sysmon Event ID 10 (ProcessAccess) against that target with PROCESS_VM_WRITE, PROCESS_VM_OPERATION, or PROCESS_CREATE_THREAD; Event ID 8 (CreateRemoteThread); Event ID 25 (ProcessTampering); and Event ID 7 (ImageLoad) for unsigned modules or anything from a user-writable path.

Tamper events the app reports about itself: this depends on developers implementing the affinity re-read above, so ask whether they did. An app reporting “my window affinity changed and I didn’t change it” is a detection with almost no false-positive surface.

Capture and remote-control tooling on regulated hosts: OBS, ShareX, Snagit, ffmpeg with a screen-grab input, and support stacks such as AnyDesk, TeamViewer, and ScreenConnect. Inventory and policy rather than alerting, since none are malicious by default. The question is why a capture stack is installed on a host whose security depends on capture being hard.

Policy drift: if Recall is disabled by policy, verify it stayed disabled on the endpoint rather than trusting that the GPO exists. BYOD is the harder case, because Recall is available by default there and the user decides.

Everything a camera sees: out of reach of host telemetry. That leaves physical controls and per-session watermarking that survives a photograph. “We can’t stop the screenshot, but we can tell whose session it came from” is a more defensible promise than “the screenshot came out black.”

Conclusion

Read the flag as what it is and Microsoft’s two pages stop contradicting each other: it keeps sensitive windows out of casual captures and out of Recall on machines where the user is not the adversary, and it does nothing about the three tiers above. When a datasheet or a control matrix claims more than that, the difference is data with nothing protecting it, and another control has to cover the gap. A design that depends on a screenshot-proof window for confidentiality is a finding rather than a control.

References