Home · Library · Operational resilience testing
Operational resilience guide

Operational resilience testing: how to prove that resilience works

Testing is where operational resilience stops being a document. A test takes an important business service, breaks a dependency it relies on, and measures whether the service stayed inside its impact tolerance. Everything else — plans, mappings, policies — is a hypothesis until a test confirms it.

What a test has to prove

A resilience test answers one question in a form the board and the regulator both accept: would this service stay inside its tolerance if the disruption happened today. That framing rules out most of what organisations call testing. A backup restored successfully proves the backup works; it says nothing about whether payments kept flowing to customers.

Three elements make a test count. The scenario has to be severe but plausible, not comfortable. The measurement has to be against the impact tolerance of a service, not against internal recovery targets. And the result has to be written down with the gaps named, including the ones nobody wants in a document.

Five types of test, and what each one is worth

TypeEffortWhat it provesWhat it cannot prove
Walkthrough of the planHoursPeople know the plan exists and can find their partThat the plan works under pressure
Tabletop exerciseHalf a dayDecisions, escalation, who calls whom, where the plan is silentThat systems and suppliers actually respond
Component testDaysOne dependency recovers within its target: a data centre, a system, a siteThat the whole service holds end to end
Service-level simulationDays to weeksThe important business service stays inside its impact tolerance end to endRare combined failures
Live disruption testWeeks of preparationReality: the service is genuinely switched to its alternative and runs thereNothing meaningful — this is the strongest evidence available

Most programmes stop at tabletops because they are safe. A tabletop is a good rehearsal of decisions and a poor test of capability. Regulators have noticed: the CBUAE, the Bank of England and the European supervisors all expect evidence that services were tested, not that meetings were held.

How to build a severe but plausible scenario

Severe but plausible is a deliberately narrow band. Too mild and the test passes without telling you anything; too extreme and everyone dismisses the result as unrealistic. The workable rule: take a disruption the organisation has genuinely survived somewhere in the world in the last five years, and apply it to your most concentrated dependency.

  1. Start from the service, not from the threat. Pick one important business service and its impact tolerance. The scenario exists to challenge that boundary.
  2. Find the concentration. Where does the service depend on a single provider, a single site, a single team or a single piece of software? That is where the scenario should land.
  3. Remove the dependency, not the building. A scenario that destroys everything teaches nothing. Take away one thing and keep the rest working.
  4. Set the clock. The test is against time: how long until the service breaches its tolerance, and did it.
  5. Decide in advance what failure looks like. Agreeing the pass criteria before the test prevents the result from being renegotiated afterwards.

Testing in banking: what supervisors look for

For financial institutions in the Gulf the requirement is explicit. The CBUAE operational resilience regulation expects mapped important business services, agreed impact tolerances and evidence of testing against severe but plausible scenarios, with results reviewed by the board. The transition period ends on 16 September 2026, and the evidence supervisors ask for is the test file: scenario, measurement, gaps, owners, dates.

The common finding in reviews is not the absence of testing. It is that the testing measured recovery of systems while the regulation asks about continuity of services. The two produce different conclusions from the same incident.

Writing up the result so it holds

A test write-up that survives scrutiny fits on two pages: the service and its tolerance, the scenario, what actually happened against the clock, whether the tolerance held, the gaps found with an owner and a date for each, and the date of the next test. Anything longer is usually protecting someone.

The uncomfortable part is naming gaps honestly. A test that finds nothing was either too easy or written up too politely, and both leave the organisation exactly where it started.

Frequently asked questions

How often should operational resilience testing happen?

Each important business service should face at least one meaningful test a year, and any service whose dependencies changed materially should be retested sooner. Annual tabletops for everything and real tests for nothing is the pattern supervisors criticise.

What is the difference between resilience testing and disaster recovery testing?

Disaster recovery testing proves a system or a site can be restored. Resilience testing proves the service the customer depends on stayed within tolerable harm. A successful DR test with a failed service outcome is a common and instructive result.

What makes a scenario severe but plausible?

It must have happened somewhere to someone comparable, and it must attack a real concentration in your own dependencies. Scenarios invented to be survivable produce comfortable, useless evidence.

Who should run the test?

The service owner runs it, the continuity or risk function designs and observes it, and internal audit reviews the evidence. A test run entirely by the team being tested rarely finds the awkward gap.

Does a passed test mean the service is resilient?

It means the service held under that scenario on that day. Resilience is a claim about the future, so the value of a test lies as much in the gaps it exposes as in the pass.

Learn this properlyERGP — the Executive Certificate in Enterprise Resilience Governance

Six modules, 94 chapters, a capstone defended before the examiner and a certificate anyone can verify. The first resilience governance certification fully available in Arabic, also in English. The AE/SCNS/NCEMA 7000 module is inside.

Explore the ERGP certification →