ArcanaIncident-response documentationBrowse the feedTemplates
Back to feed
Playbook
PB-019

Major Security Incident Management

1. Purpose & Scope

  • Purpose:

    Provide a command-and-coordination playbook for Sev 0 or major security incidents. Use this playbook to establish incident command, run the war room, coordinate multiple technical playbooks, manage executive/legal communications, prioritise recovery, and track decisions until the incident is downgraded or closed.

  • Scope:

    Applies to security incidents with major business, customer, regulatory, operational, or executive impact. This playbook does not replace technical playbooks; it coordinates them when the incident exceeds a single-team or single-scenario response.

2. Incident Identification & Criteria

Incident Type: Major Security Incident Management

Trigger Conditions:

Initiate this playbook when any of the following occur:

  • Incident is declared Sev 0 or equivalent major incident
  • Multiple technical playbooks must run concurrently under one command structure
  • Critical business services, customer-facing systems, identity infrastructure, cloud tenant, backups, or security tooling are materially impacted
  • Executive, legal, regulatory, customer, public, vendor, or law enforcement coordination is required
  • Response will span multiple teams, geographies, shifts, or recovery phases

Severity Levels:

SeverityDescription
Sev 3N/A
Sev 2Cross-team incident with material risk, limited business impact, or executive awareness required
Sev 1Major incident with confirmed business, customer, data, production, identity, cloud, or recovery impact
Sev 0Enterprise-wide, critical-service, regulated-data, public, extortion, destructive, or crisis-level impact

3. Roles & Responsibilities

  • Incident Commander: Owns command, severity, priorities, operating cadence, decision log, escalation, and downgrade/closure decisions.
  • Deputy Incident Commander: Maintains action tracker, follows up on blockers, and supports continuity during extended response.
  • Technical Lead: Coordinates technical workstreams and ensures relevant playbooks/runbooks are executed.
  • Communications Lead: Coordinates stakeholder updates, executive briefings, internal messaging, and approved external communications.
  • Scribe: Maintains incident timeline, action tracker, decision log, and evidence references.
  • Other Roles:
    • Legal / Compliance: Assesses regulatory, contractual, legal, insurance, and disclosure obligations.
    • Business / Service Owners: Confirm business impact, recovery priority, customer impact, and operational constraints.
    • Vendor Owner / Procurement: Coordinates third-party support where vendor action is required.
    • Customer Support / Success: Coordinates customer-impact handling and support readiness.
    • Recovery Lead: Coordinates service restoration priorities, dependencies, and validation.

4. Initial Actions

  • Immediate Steps:

    • Confirm the major incident declaration, severity, Incident Commander, Technical Lead, Communications Lead, Scribe, and executive sponsor.
    • Open the major incident war room, action tracker, decision log, and operating cadence.
    • Page required engineering, security, platform, application, identity, cloud, network, legal, communications, support, and vendor-owner teams.
    • Identify active technical playbooks and assign owners for each workstream.
    • Start executive/legal communication process if major impact, customer impact, regulated data, or external communication is possible.

    Decision Point:

    • If the incident does not require major incident coordination โ†’ continue under the relevant technical playbook only.
    • If major coordination is required โ†’ continue this playbook alongside all relevant technical playbooks.
    • If external communications may be required โ†’ engage Legal/Compliance and Communications Lead before any external statement.

5. Investigation & Analysis

6. Containment, Eradication & Recovery

  • Containment Actions:

    • Confirm each technical workstream has an owner, containment objective, current action, risk, and rollback plan.
    • Prioritise containment actions that reduce active attacker access, customer impact, data exposure, destructive risk, and business-critical service impact.
    • Record containment decisions, risk acceptances, and business tradeoffs in the decision log.
  • Eradication Steps:

    • Confirm each technical playbook owns eradication of incident-specific root cause, persistence, malicious access, and vulnerable exposure.
    • Track cross-workstream dependencies such as identity resets, network segmentation, rebuilds, vendor support, legal holds, and business approvals.
    • Ensure eradication is not declared complete until technical leads confirm root cause and persistence have been addressed.
  • Recovery Steps:

    • Establish recovery priority based on business criticality, dependency order, customer impact, safety, and residual risk.
    • Gate service restoration on containment status, integrity validation, monitoring readiness, and business owner approval.
    • Downgrade the major incident only after the Incident Commander confirms stable containment, recovery path, communications status, and remaining risk.

7. Communication & Escalation

8. Post-Incident Activities

  • Lessons Learned:

    • Conduct a PIR covering command structure, severity decisions, communication cadence, decision quality, handoffs, technical workstream coordination, recovery prioritisation, and stakeholder outcomes.
    • Review whether the major incident was declared early enough and whether it was downgraded at the right time.
    • Identify coordination gaps, unclear ownership, delayed decisions, missing evidence, or communication failures.
  • Documentation Updates:

    • Update technical playbooks, SOPs, escalation paths, contact lists, status templates, decision log templates, and recovery priority guidance based on PIR findings.
    • Track all follow-up actions with owners, due dates, and validation criteria.

9. References & Linked Resources

10. Appendices

  • Expected Outputs:

    • Major incident declaration and severity record
    • Active workstream list and owners
    • War room link, action tracker, and decision log
    • Consolidated incident timeline
    • Executive/legal/customer/vendor communication records
    • Recovery priority and dependency plan
    • Downgrade or closure decision record
    • PIR action items and owners
  • Common Failure Modes:

    • Declaring a major incident too late
    • Running multiple technical playbooks without a single command structure
    • Missing decision logs or action owners
    • Allowing recovery work to start before containment and integrity validation
    • Sending external communications before legal and communications approval
    • Losing context during long-running or multi-geo handoffs

Contributor

Vishal Thakur GitHub: https://github.com/malienist

Contributed to the Arcana Incident Response Documentation Framework.