A resilience case study and reusable incident-management model covering prioritised recovery, communication, BCP, post-incident learning and risk transfer.
Process map
DetectScope / severity
known facts
known facts
TriageImpact / priority
dependencies
dependencies
MobiliseIncident lead
workstreams
workstreams
ContainProtect critical
minimise spread
minimise spread
RecoverRestore in
business order
business order
CommunicateExecs / users
cadence
cadence
ValidateServices / data
residual risk
residual risk
ImproveLessons / BCP
control uplift
control uplift
Establish command quickly
- Name an incident lead and decision path.
- Separate technical recovery, business continuity and communications workstreams.
- Create one operational picture of affected users, services and dependencies.
- Prioritise critical business services rather than restoring technology in arbitrary order.
Communicate as part of the control model
- Executives need consequence, trajectory, decisions and next checkpoint.
- Users need clear actions and realistic expectations.
- Set a reporting cadence frequent enough to maintain trust without distracting responders.
Business continuity during recovery
- Use manual or alternate processes where critical services can operate safely without normal technology.
- Track where workarounds introduce privacy, control or reconciliation risk.
- Plan the transition back to normal operation so temporary workarounds do not become permanent.
Post-incident learning
- Review technology dependencies, roles, BCP/DR realism, supplier response and communication effectiveness.
- Assess cyber-insurance or other risk-transfer arrangements against the actual event, evidence requirements and loss thresholds.
- Change plans, controls or architecture where the incident justifies it.
Experience context
The experience underpinning this piece includes leading BCP/incident response and recovery during the 19 July 2024 CrowdStrike outage, affecting most of an approximately 120-person environment across endpoints, servers and specialist compute, with rapid recovery limiting material financial loss.
Key lessons
- Incident leadership is prioritisation under uncertainty.
- Recovery order should follow business criticality.
- Communication cadence is an operational control.
- BCP and DR plans become credible when tested against real events.