Risk matrices and the residual / target / appetite triple
A risk matrix places a risk on a grid: consequence on one axis, likelihood on the other, with each cell carrying a band and a colour. It is the most widely used artefact in risk management and the most argued about, and a good deal of the argument comes from conflating two different things.
How do you rate a risk on one?
The standard gives a specific order at B.10.3.2, and it is not the order most workshops use:
To rate a risk, the user first finds the consequence descriptor that best fits the situation then defines the likelihood with which it is believed that consequence will occur.
Consequence first, then the likelihood of that consequence. This matters more than it sounds. The standard spells out why: "Where a range of different consequence values are possible from one event, the likelihood of any particular consequence will differ from the likelihood of the event that produces that consequence. Generally the likelihood of the specified consequence is used."
A data breach might be near-certain to happen in some minor form and very unlikely to reach the catastrophic end. Rating it "almost certain" against "catastrophic" combines the likelihood of the event with the severity of its worst consequence, and produces a number that describes nothing. Pick the consequence, then ask how likely that is.
The standard adds one more constraint that is easy to breach across a large register: "The way that likelihood is interpreted and used should be consistent across all risks being compared."
Designing the scales
This is where most of the value and most of the failure sits, and the standard is unusually prescriptive about it.
- Customise them. Always. The standard prints only part examples of its own scales, with a note saying this is deliberate, "to stress that the scales should always be customized". A shipped default, including ours, is a starting point.
- Tie them to objectives. Scales "should be directly connected to the objectives of the organization, and should extend from the maximum credible consequence to the lowest consequence of interest". Both ends are a decision: what is the worst thing that could credibly happen, and below what point do you stop caring.
- Step by an order of magnitude. "Generally, to be consistent with data, each scale point on the two scales will need to be an order of magnitude greater than the one before." This is the rule most matrices in the wild fail. If your consequence bands run $100k, $250k, $500k, $1M, $2M, the scale is roughly linear and the top band will fill up with everything serious.
- Test the draft. "Draft matrices need to be tested to ensure that the actions suggested by the matrix match the organization's attitude to risk and that users correctly understand the application of the scales." Rate a dozen known risks and check whether the answers match what the organisation would actually do about them.
- Consequence scales can run both ways. "The consequence scale (or scales) can depict positive or negative consequences."
Scales "can have any number of points", with three, four and five-point scales the most common, and can be "qualitative, semi-quantitative or quantitative".
The three plotted ratings
A single point tells you where a risk sits today. Beau-Tie plots up to three, because the distance between them is where the decision content lives.
Residual (R), where you are
The rating with today's controls running. ISO 31073:2022 defines residual risk (3.3.38) as the "risk remaining after risk treatment", and notes it "can contain unidentified risk" — it is what is left after treating what you found, not a full account of what is left. Plotted as a teal dot.
Inherent risk, the rating before any controls, is not plotted. The controls already exist, so an inherent rating describes a world that is not the one you are in. It remains a common requirement in audit and regulated-sector reporting, so treat that as our opinion rather than a settled position.
Target (T), where you are going
Where the residual should land once planned controls are live. Not a term either standard defines; it is practice. A target equal to the residual means no treatment is planned, which is fine when you are already within appetite and worth a conversation when you are not. Plotted as a confidence-green dot.
Appetite (A), what you will tolerate
ISO 31073:2022 3.3.27 defines risk appetite as the "amount and type of risk that an organization is willing to pursue or retain", and ISO 31000:2018 clause 5.2 puts establishing it with top management. It is policy rather than a per-risk judgement, and it is not the analyst's to set. Optional in Beau-Tie, because plenty of organisations have no stated appetite and inventing one to fill the field would be worse than leaving it empty. Plotted as a muted grey dot.
Reading the three together
- R above A, T above A. Outside appetite today, and the plan still leaves you outside. The treatment is not ambitious enough, or the appetite is wrong. Escalate and find out which.
- R above A, T below A. The plan closes the gap. Pressure-test the timeline: the risk sits at residual until the controls actually land, and planned controls have a habit of staying planned.
- R within A, T below R. Already tolerable and improving further. Defensible, but ask whether that treatment effort is better spent on something outside appetite.
- R within A, T equal to R. Within appetite, nothing planned. A healthy posture provided the rating is honest, which is the whole question.
Strengths and limitations
The standard states both at B.10.3.5, and its limitations list is blunter than most commentary written by people who dislike matrices.
Strengths. It "is relatively easy to use", "provides a rapid ranking of risks into different significance levels", gives "a clear visual display", and "can be used to compare risks with different types of consequence". That last one is the real reason matrices persist: it is the only cheap way to put a safety risk and a financial risk on the same page.
Limitations, quoted:
- "It requires good expertise to design a valid matrix."
- "It can be difficult to define common scales that apply across a range of circumstances relevant to an organization."
- "It is difficult to define the scales unambiguously to enable users to weight consequence and likelihood consistently."
- "The validity of risk ratings depends on how well the scales were developed and calibrated."
- "It requires a single indicative value for consequence to be defined, whereas in many situations a range of consequence values are possible and the ranking for the risk depends on which is chosen."
- "A properly calibrated matrix will involve very low likelihood levels for many individual risks which are difficult to conceptualize."
- "Its use is very subjective and different people often allocate very different ratings to the same risk. This leaves it open to manipulation."
- Risks cannot be directly aggregated from a matrix.
The standard's own further reading on this is Elmonstri, Review of the strengths and weaknesses of risk matrices, and Baybutt, Calibration of risk matrices for process safety. The best-known academic critique is Cox, "What's Wrong with Risk Matrices?", Risk Analysis 28(2), 2008, which shows that matrices can rank two risks in the opposite order to a coherent quantitative scoring. We point at it rather than quote it: we have not read the original, only the result as it is universally reported.
The pragmatic position: a matrix is the right tool for triage, for policy hooks, and for comparing unlike risks quickly. Where one risk is big enough to drive a material financial decision, complement it with a quantitative model. That is what Monte Carlo and single-event quantification are for.
Beau-Tie's default matrix
A 5×5 grid: likelihood Rare to Almost Certain, consequence Insignificant to Catastrophic, with bands Low, Moderate, High and Extreme. Fully editable from the Matrix dialog, and you can set a per-user default if your organisation always uses its own.
The colours are deliberately desaturated against the usual traffic-light heatmap so the eye lands on the plotted ratings rather than the background. The bands are also softer than they once were: Extreme is confined to the top-right corner and the Rare row is almost entirely Low. That is a defensible default and it is still a default. Per the standard, calibrate it against your own consequence scales before trusting what it tells you.