A resilience test answers one question in a form the board and the regulator both accept: would this service stay inside its tolerance if the disruption happened today. That framing rules out most of what organisations call testing. A backup restored successfully proves the backup works; it says nothing about whether payments kept flowing to customers.
Three elements make a test count. The scenario has to be severe but plausible, not comfortable. The measurement has to be against the impact tolerance of a service, not against internal recovery targets. And the result has to be written down with the gaps named, including the ones nobody wants in a document.
| Type | Effort | What it proves | What it cannot prove |
|---|---|---|---|
| Walkthrough of the plan | Hours | People know the plan exists and can find their part | That the plan works under pressure |
| Tabletop exercise | Half a day | Decisions, escalation, who calls whom, where the plan is silent | That systems and suppliers actually respond |
| Component test | Days | One dependency recovers within its target: a data centre, a system, a site | That the whole service holds end to end |
| Service-level simulation | Days to weeks | The important business service stays inside its impact tolerance end to end | Rare combined failures |
| Live disruption test | Weeks of preparation | Reality: the service is genuinely switched to its alternative and runs there | Nothing meaningful — this is the strongest evidence available |
Most programmes stop at tabletops because they are safe. A tabletop is a good rehearsal of decisions and a poor test of capability. Regulators have noticed: the CBUAE, the Bank of England and the European supervisors all expect evidence that services were tested, not that meetings were held.
Severe but plausible is a deliberately narrow band. Too mild and the test passes without telling you anything; too extreme and everyone dismisses the result as unrealistic. The workable rule: take a disruption the organisation has genuinely survived somewhere in the world in the last five years, and apply it to your most concentrated dependency.
For financial institutions in the Gulf the requirement is explicit. The CBUAE operational resilience regulation expects mapped important business services, agreed impact tolerances and evidence of testing against severe but plausible scenarios, with results reviewed by the board. The transition period ends on 16 September 2026, and the evidence supervisors ask for is the test file: scenario, measurement, gaps, owners, dates.
The common finding in reviews is not the absence of testing. It is that the testing measured recovery of systems while the regulation asks about continuity of services. The two produce different conclusions from the same incident.
A test write-up that survives scrutiny fits on two pages: the service and its tolerance, the scenario, what actually happened against the clock, whether the tolerance held, the gaps found with an owner and a date for each, and the date of the next test. Anything longer is usually protecting someone.
The uncomfortable part is naming gaps honestly. A test that finds nothing was either too easy or written up too politely, and both leave the organisation exactly where it started.
Each important business service should face at least one meaningful test a year, and any service whose dependencies changed materially should be retested sooner. Annual tabletops for everything and real tests for nothing is the pattern supervisors criticise.
Disaster recovery testing proves a system or a site can be restored. Resilience testing proves the service the customer depends on stayed within tolerable harm. A successful DR test with a failed service outcome is a common and instructive result.
It must have happened somewhere to someone comparable, and it must attack a real concentration in your own dependencies. Scenarios invented to be survivable produce comfortable, useless evidence.
The service owner runs it, the continuity or risk function designs and observes it, and internal audit reviews the evidence. A test run entirely by the team being tested rarely finds the awkward gap.
It means the service held under that scenario on that day. Resilience is a claim about the future, so the value of a test lies as much in the gaps it exposes as in the pass.
Six modules, 94 chapters, a capstone defended before the examiner and a certificate anyone can verify. The first resilience governance certification fully available in Arabic, also in English. The AE/SCNS/NCEMA 7000 module is inside.
Explore the ERGP certification →