Two kinds of metric, both needed
Lagging metrics measure what already happened: tolerance breaches, incident duration, regulatory findings. They tell the board whether resilience held. Leading metrics measure capability before the incident: plans exercised, restore tests passed, crisis roles trained, supplier concentration. They tell the board whether resilience will hold. A dashboard with only lagging metrics is a rear-view mirror; with only leading ones it is a promise.
The ten metrics
| # | Metric | Type | Green | Red |
|---|---|---|---|---|
| M-01 | Impact tolerance breaches per quarter | Lagging | 0 | 2 or more |
| M-02 | Prioritised activities with a plan exercised in 12 months | Leading | 90-100 % | below 70 % |
| M-03 | Critical services where tested recovery time exceeds RTO | Leading | 0 | 3 or more |
| M-04 | Critical systems with a successful restore test in 6 months | Leading | 100 % | below 90 % |
| M-05 | Critical services depending on a single supplier | Leading | 0 | 3 or more |
| M-06 | Corrective actions overdue | Leading | below 10 % | above 25 % |
| M-07 | Crisis roles with trained primary and deputy | Leading | 100 % | below 80 % |
| M-08 | Mean time to detect, minutes | Lagging | below 15 | above 60 |
| M-09 | Mean time to activate the plan, minutes | Lagging | below 30 | above 60 |
| M-10 | Open regulatory findings | Lagging | 0 | 4 or more |
Thresholds in the table are starting points for a mid-size organisation in the Gulf; the file keeps them in editable columns. The rule that matters more than the numbers: every red threshold has a named reaction, otherwise the metric is statistics.
Download the template
Download the template. Metrics sheet with ten indicators, definitions, data sources, frequency, owners and thresholds; a one-page board view that pulls current values and statuses. Xlsx.
Download xlsx (operational-resilience-metrics-template.xlsx)
Status for M-01 is calculated by formula; for percentage metrics enter the status or adapt the formula to your thresholds. No registration, no forms. Need it adapted to your organisation or a full programme: see how we work.
Where the data comes from
Nothing in the ten requires new data collection. Incident and exercise logs give M-01, M-02, M-08 and M-09; IT test reports give M-03 and M-04; the supplier register gives M-05; the corrective action log gives M-06; training records give M-07; the findings register gives M-10. If one of these logs does not exist, that absence is the first finding, and the internal audit checklist will record it.
How to present them
One page, ten rows, three colours, trend arrows against last quarter. The board does not need the definitions; it needs to see which rows are red and what is being done. The structure of the full report is on the board reporting page, with a template.
Frequently asked questions
What are operational resilience metrics?
Measures that show whether an organisation can keep delivering critical services through disruption: lagging ones such as tolerance breaches and time to detect, and leading ones such as plans exercised, restore tests passed and crisis roles trained. The ten on this page cover both kinds.
How many resilience metrics should a board see?
Ten or fewer on one page, each with a threshold and a trend. More becomes an operational dashboard that belongs to management, not the board.
What is an impact tolerance?
The maximum level of disruption of a critical service the organisation is prepared to accept, expressed in time or volume. Breaches of tolerance are the single most important lagging metric and the one regulators ask about first.
How often should the metrics be updated?
Monthly for operational owners, quarterly for the board. Leading metrics change slowly; lagging ones are reported as incidents occur.