Major IT system outage
Extended unavailability of mission-critical IT systems.
The risk event
Extended unavailability of mission-critical IT systems
A failure of one or more critical systems materially impacting customers or internal operations.
What could cause it, and what stops it
The left-hand side. Each cause is a plausible pathway to the event; the controls beneath it are the barriers that reduce the chance of that pathway completing.
Infrastructure failure
Hardware or platform failure in critical infrastructure.
- Redundant infrastructurePreventive · Effective
Active-active redundancy for tier-1 services.
- Capacity monitoringDetective · Effective
Continuous capacity headroom monitoring.
- Failover testingDirective · Limited
Quarterly failover and DR exercises.
Software defect
A defective change is deployed to production.
- Change managementPreventive · Limited
Standard change approval process.
- Staged rolloutsPreventive · Effective
Canary and progressive deployment.
- Rollback proceduresCorrective · Effective
Tested rollback runbooks per service.
Cyberattack / ransomware
Malicious encryption or destructive attack on IT estate.
- Endpoint protectionPreventive · Effective
EDR with behavioural detection.
- Network segmentationPreventive · Limited
Lateral-movement controls between zones.
- Immutable backupsCorrective · Highly effective
Air-gapped backups with rapid restore.
Third-party SaaS outage
A critical SaaS provider experiences extended downtime.
- SLA monitoringDetective · Limited
Real-time monitoring against SaaS commitments.
- Multi-cloud failoverPreventive · Planned · 12 months or more
Active failover to alternate provider.
What happens if it occurs, and what limits it
The right-hand side. Each consequence is an outcome the event could produce; the controls beneath it are what contains or recovers from that outcome once the event has already happened.
Service unavailability for customers
Customers cannot use products or services.
- Status pageCorrective · Effective
Public status page with timely updates.
- Customer comms protocolCorrective · Effective
Cadence and ownership for customer updates.
- SLA creditsCorrective · Limited
Automated SLA credit issuance.
Lost revenue
Direct revenue impact during the outage.
- Business interruption insuranceCorrective · Effective
BI cover for major incidents.
- Demand recapture planCorrective · Limited
Promotional levers to recover lost demand.
Data loss
Loss of customer or operational data.
- Backup schedulePreventive · Effective
RPO-aligned backup frequency.
- Restore testingDetective · Limited
Monthly restore tests with success metrics.
- RPO/RTO targetsDirective · Effective
Documented RPO/RTO per service tier.
Regulatory reporting failures
Inability to meet statutory reporting deadlines.
- Incident reporting workflowCorrective · Effective
Documented regulator notification process.
- Regulator notification SLADirective · Effective
Defined timeline commitments per regulator.
Where this template starts you
Ratings are a starting position, not a finding. They describe a generic organisation with the controls above in place; yours will differ, and the point of opening the template is to make them yours.
Make it yours
Opening the template loads it into the editor with everything above already in place. Rename the event, cut the causes that do not apply, and re-rate against your own matrix. Exports to PNG, PDF, Excel and PowerPoint are built in.