4 min read

New Approach to Measuring Control Effectiveness

Breach and attack simulation and continuous automated red teaming are making it possible to generate empirical, falsifiable data on control effectiveness — moving past theoretical mappings and expert judgment alone.
New Approach to Measuring Control Effectiveness

Organisations invest enormous amounts in cybersecurity controls, but struggle to measure and distinguish their actual contributions to stopping threat activity. The fundamental issue is that at the level of an individual organisation, serious incidents occur relatively infrequently compared with an enormous number of ever-evolving threats, controls and attack paths. Thus, even when an incident occurs, while it is possible to determine the pathway that succeeded, we do not observe all the threat activity that was curtailed by controls, or the plausible attack paths that weren’t tested.

As a result, practitioners must rely on anecdotal evidence, best practices and their own expertise to predict which exposures are riskiest and which controls deserve investment. The absence of reliable data, methods and tooling presents a serious challenge to crucial cybersecurity practices such as cyber threat exposure management (CTEM) and cyber risk quantification (CRQ).

New tooling, however, such as breach and attack simulation (BAS) and continuous automated red teaming (CART), has drastically reduced the cost of generating high-quality and highly tailored simulation and experimental data. This data, properly modelled, can be the basis of a new empirical approach to measuring control effectiveness.

The current state of evaluating control effectiveness

Experts in the cybersecurity community have recognised the importance of understanding control effectiveness and have pooled their expertise to do so. For example, the CIS Community Defense Model matches up CIS controls to MITRE ATT&CK techniques, which have been extracted from threat reporting. That mapping is very useful, but it is a theoretical approach and highly generalised. It does not tell us how reliably a control stops a technique in a particular organisation, which has its own specific constellation of assets, controls, threat environments, and exploitation events.

In the field of cyber risk quantification (CRQ), the CRQ process requires a quantitative statement of the impacts of controls on reducing various cyber risks. Typically, experts are asked to provide calibrated predictions on the extent to which a control will curtail an attack path. For example, an expert might say that the 90% confidence interval that phishing-resistant MFA reduces successful account takeover attempts is between 40% and 80%.

The problem is that humans struggle with predictions about complex scenarios, and they can’t improve unless they get repeated feedback on the accuracy and precision of their predictions. This is what academic prediction literature has determined after many large-scale trials, culminating in popular books such as Philip Tetlock’s Superforecasting.

The most fundamental principle espoused by Tetlock is that all useful predictions must be falsifiable, and any prediction method that purports to be superior (typically through reducing cognitive biases or quantitative sophistication) should be tested rather than accepted at face value. Only falsifiable predictions can be reliably scored, and “superforecasters” are those that have the best scores. In the cyber realm, it isn’t common practice to make verifiable predictions about control effectiveness. That means that even the most logical, most rigorous methodologies (and some of them are very impressive from that perspective) may be no more accurate than one based intuition. We just don’t know.

For cyber threat exposure management (CTEM), the lack of empirical data on control effectiveness weakens our ability to do exposure prioritisation. We would love to know whether those 10,000 critical overdue vulnerabilities are actually exploitable given the control stack, but how do we get an accurate understanding at scale and speed? While there are new and interesting approaches on the Unified Vulnerability Management (UVM) market now, those tools generally take a theoretical approach to exploitability rather than generating empirical data based on actual attempts to exploit. In other words, they examine whether the firewalls, endpoint detection, and other controls should theoretically prevent exploitation. It’s a major step in the right direction for sure, but theory isn’t empirical evidence.

From generalised data to tailored data

Using data to determine control effectiveness has been an industry goal for many years, and we're starting to see some payoff. Cyber insurer Marsh McLennan publicised 2023 research comparing control categories against cyber incident data. According to them, automated hardening showed the strongest association: organisations with it in place were nearly six times less likely to experience a cyber incident.

Marsh expanded that work in 2025 using thousands of organisations’ control implementation data and breach-related claims. The research again found that implementation comprehensiveness mattered. For example, each additional 25 percentage points of EDR deployment across workstations was correlated with a further 10% decrease in breach likelihood.

This is helpful trending, but these industry and sector-level observations are not specific to individual firms, each with their own specific constellation of assets, controls, and threat environments.

Solution: generating tailored data

It’s now possible to generate control effectiveness data ourselves using in situ simulations and experimentation. That used to be a non-starter because human red teams and pentesters are expensive. But breach and attack (BAS) simulation tools continuously exercise security controls using simulated attacker techniques. Continuous automated red teaming (CART) tools can exploit exposures and test out attack paths. A newer generation of AI-assisted offensive-security platforms is emerging and supercharging the potential of these tools to operate quickly and at scale.

In effect, the cost of producing large numbers of observations about the interactions of threat and control activity is falling quickly. We can now generate the data about how individual controls and combinations of controls change attacker success. That data can be snapped up by data scientists to produce prediction models about control effectiveness. These predictions are based on experiments and simulations tailored to the organisation’s environment, which is as close to real-life incidents as any particular organisation wants to get. It represents a major upgrade to the existing basis of understanding control effectiveness, contributing falsifiable predictions and empirical data to a cybersecurity domain that has been seeking a way to measure effectiveness for a long time.