Skip to the content.

The training-discipline playbook. The framework’s operational disciplines (kill-switch activation, evidence export, communication) only work if the responders can execute them under pressure. PB16 converts the framework from documentation into operational capability through the 30-Minute Micro-Drill, the Four Core Moves, the two permanent operating roles (Safe Mode Owner and Evidence Owner), and the Curriculum-of-Six that focuses training on practical actions rather than AI theory. Run monthly, measure rigorously, fix the failures rather than re-explain them.

Part of the AI IR Overlay™ framework. See CONTENT_MAP.md for the full repository map.


Playbook 16: Training Your Team for AI Incidents

Every other playbook in this framework assumes the response team can execute the operational discipline it specifies. Activate Mode M3 Tool Tiering inside 10 minutes per Playbook 04. Export the Minimum Evidence Set A through F inside 60 minutes per the Evidence Export Script Contract. Issue the first stakeholder update inside 30 minutes per Playbook 17. None of those time budgets survive untrained execution. The first time a responder activates safe mode, exports evidence, applies the Three-Status Taxonomy, or convenes the Materiality call should not be during a real incident. PB16 is the training discipline that converts the framework’s documented capabilities into the team’s operational muscle memory.

Premise

Traditional incident response training (annual tabletops, theory-heavy curricula, generic SANS-style courses) is materially insufficient for AI incidents in four ways:

Aspect Traditional IR training AI IR training
Cadence Annual or twice-yearly tabletop Monthly micro-drill plus quarterly full scenario
Format 2-to-4-hour facilitated discussion 30-minute time-boxed live execution
Content scope Threat-model awareness, IR-process review, scenario discussion Direct execution of the framework’s operational moves (safe mode activation, evidence export, communication tagging) against live or staging-environment AI agents
Measurement Participation logged, exercise completed Time-to-Safe-Mode, Time-to-Evidence, 5-bullet executive update produced, all measured and bench-marked against the framework’s targets

The mismatch matters because AI incidents have a distinctive response shape that traditional training does not prepare responders for:

The biggest operational problem with AI IR readiness is not the absence of documentation; the framework’s playbooks document the disciplines extensively. The problem is the gap between the team has read the documentation and the team can execute the documentation under operational pressure at 02:30 on a Saturday morning. This playbook’s job is to close that gap through structured, measured, repeated practice.

Mental Model clauses engaged: PB16 operationalizes the entire Mental Model through training. The Acts clause is rehearsed through Tool-Call Ledger export and tool-tier disablement drills; the Remembers clause through Memory Snapshot drills; the Retrieves clause through Retrieval Trace drills; the Changes clause through Configuration Snapshot drills. Each clause has a training move; each training move has a measured target.

Use this playbook when: designing or revising the customer’s AI IR training curriculum · onboarding a new responder to the on-call rotation · scoping the Playbook 14 (Testing for Agent Failure Modes) testing discipline for the training-side complement · scoping the Playbook 20 (Maturity Roadmap) Level 2 (Containable) and Level 3 (Provable) capability validations · designing the quarterly board metric for Playbook 24 (Board-Ready Scorecard) training-readiness · responding to an audit finding that the customer’s AI IR training is insufficient · onboarding a newly-acquired business unit or subsidiary that needs to come into AI IR conformance · supporting the customer’s IR-mutual-aid arrangement with a partner organization where joint drills are required · responding to a regulator’s question about the customer’s AI IR team competence.

First-Hour Actions

PB16 is structurally a design-time and cadence-driven playbook. Its First-Hour Actions activate in three scenarios: the monthly micro-drill itself, an onboarding session for a new responder, and a post-incident competence review when an incident response surfaced a training gap.

Case A: The monthly micro-drill (the load-bearing cadence)

The 30-Minute Micro-Drill is the framework’s primary training artifact. Run monthly, against a representative agent in the customer’s deployment, with the on-call responder executing live.

Minute Action Owner
0–10 Phase 1: Trigger and Contain. Drill lead presents a realistic scenario (anomalous tool-call spike, suspicious agent output, customer complaint about agent behavior, drift signal per Playbook 22). On-call responder activates the appropriate safe mode (typically M1 Read-Only or M3 Tool Tiering depending on the scenario). The responder logs the activation time, the mode chosen, the rationale, and the affected agent’s AI-BOM identifier. Measured target: Time-to-Safe-Mode under 10 minutes. Drill lead + On-call responder + Safe Mode Owner
10–20 Phase 2: Pull Evidence. Responder exports the Minimum Evidence Set per the Evidence Export Script Contract: prompt and response logs (Type A), tool-call ledger (Type B), retrieval traces (Type C), memory snapshot (Type D), configuration snapshot (Type E), identity and SaaS audit-log correlation (Type F). The export uses the customer’s actual export procedure against the actual evidence store. Gaps, access issues, or timeouts are logged. Measured target: Time-to-Evidence under 60 minutes; export completeness scored against the six types. On-call responder + Evidence Owner
20–30 Phase 3: Scope and Brief. Responder produces the 5-bullet executive update per Playbook 17 Template Library: confirmed facts (per the Three-Status Taxonomy), suspected issues (explicitly tagged), containment status (mode active, time-to-activation), potential impact (data, actions, trust), following actions and ownership. The brief is reviewed for Responsible Reframing adherence (zero anthropomorphizing language) and Four-Element Update Standard completeness. Measured target: 5-bullet brief delivered inside the 30-minute window, all elements present, status taxonomy applied to every claim. On-call responder + Incident Commander (drill role)

Discipline: the drill is time-boxed. A drill that runs over 30 minutes is itself a training finding (the response sequence is too slow; the training has not yet produced the framework’s expected response time). A drill that completes inside 30 minutes but with skipped steps is also a finding (the responder is meeting the time budget by shortcutting the discipline). Both failures enter the drill retrospective and the Playbook 18 (Post-Incident Hardening) 5-business-day SLA backlog.

Critical rule: the on-call responder executes the drill without engineering assistance. Permission issues, missing logs, unclear runbooks, or unfamiliar tools that block the responder during the drill are infrastructure findings, not training failures. The drill is also a test of the customer’s operational substrate; gaps are fixed in the runbook, the access controls, or the tooling rather than absorbed as the responder’s responsibility.

Case B: Onboarding a new responder

A new responder enters the on-call rotation when they can complete the 30-Minute Micro-Drill independently. The onboarding sequence:

Phase Action Owner
Week 1 Curriculum-of-Six. The new responder learns the six core training topics (see Evidence Priorities below) through documented runbooks, framework playbook reading, and shadowing the current on-call responder during a real incident or a drill Training lead + Senior responder
Week 2 Assisted drill. The new responder runs the 30-Minute Micro-Drill with the senior responder shadowing. The senior responder coaches in real time; the drill is not time-bounded as strictly as a measured drill Senior responder + New responder
Week 3 Solo drill. The new responder runs the drill independently against a representative agent. The drill is fully time-bounded. The drill lead reviews the result and identifies remaining gaps Drill lead + New responder
Week 4 On-call eligibility. The new responder enters the on-call rotation. The first three on-call shifts include a senior responder available for backup; the new responder owns the response but can escalate quickly Training lead + New responder

Case C: Post-incident competence review

After a real incident, the post-incident retrospective per Playbook 18 (Post-Incident Hardening) includes a training-side review. The discipline is to distinguish capability gaps from operational gaps:

Phase Action Owner
Day 1 of retrospective Inventory the response-execution timeline. Document each operational step the responder took, with timestamps, against the framework’s target time budgets (TTSM, TTE, TT-first-update) IC + Training lead
Day 2 of retrospective Categorize gaps. Distinguish capability gaps (the responder did not know how to do X) from substrate gaps (the responder knew but was blocked by access, tooling, or runbook gaps). Capability gaps drive training curriculum updates; substrate gaps drive PB18 hardening items Training lead + IC + Platform engineer
Day 3 of retrospective Update the curriculum. Capability gaps surfaced by the incident are added to the next monthly micro-drill scenario and to the Curriculum-of-Six documentation. A capability gap that surfaces twice is a curriculum-design finding Training lead
Day 5 of retrospective Confirm hardening discipline closure. Per PB18’s 5-business-day SLA, both capability and substrate gaps have action items with owners and target dates by day 5 of the retrospective IC + Training lead + Platform engineer

Containment Options

PB16 does not introduce a new kill-switch variant because training is a discipline, not a containment surface. The framework’s existing Kill-Switch Modes apply unchanged. PB16’s containment-equivalent discipline is the competence-scope containment: the actions that limit response responsibility to trained capability while training catches up to expected scope.

Competence-scope containment

Action Use when What changes
Restricted on-call rotation New responder has not yet completed the onboarding sequence; an existing responder has been off the rotation for an extended period (parental leave, sabbatical, role change); a quarterly competence review has identified a capability gap The responder is removed from solo on-call; their on-call shifts include a senior backup; they continue to develop competence through assisted drills and shadowing
Buddy-system on-call The team has only one fully-trained responder; the team is geographically distributed and timezone coverage requires multiple responders who each have partial competence On-call shifts are run as pairs; each pair includes at least one fully-trained responder; the buddy system is documented as the customer’s interim posture and the recruiting or training plan to exit it is named
Capability-scoped escalation A specific incident class (e.g., insider threat per Playbook 12, vendor copilot incident per Playbook 10) requires specialized response that not every on-call responder has been trained on The on-call responder activates initial containment and escalates to the specialized responder; the escalation path is documented and time-bounded; the customer’s training plan progresses every responder toward the specialized scope
External-IR-mutual-aid activation The customer’s incident is beyond the trained team’s capacity (scale, complexity, scope); the customer has an MSA with an external IR provider The mutual-aid provider is engaged per the contracted SLA; the customer’s responders retain decision authority and drive the response; the mutual-aid provider augments capacity rather than replacing the customer’s discipline
Drill-pause cooldown The team has run drills aggressively without time to absorb the findings; the operational substrate gaps are accumulating faster than they are being closed The drill cadence is temporarily reduced to allow PB18 hardening backlog closure; the cooldown is documented with an explicit end date and a target backlog state before resumption
Training-substrate fix as containment A drill surfaces a substrate gap (access issue, runbook gap, tool gap) that compromises the response readiness even of fully-trained responders The substrate fix takes priority over continued drills; the framework’s training claims for this agent or scenario class are paused until the fix is complete; the pause is communicated to the customer’s leadership and tracked in the PB24 scorecard

The six actions are operational pragmatism. The discipline is to be explicit about when each is in use; an implicit “we don’t have anyone trained for that yet” posture compounds quietly into a capability gap that the customer cannot defend in a regulator review.

Evidence Priorities

PB16’s evidence discipline operates at two levels: the Curriculum-of-Six that defines what the training covers, and the drill-artifact archive that preserves training evidence for the customer’s records discipline per Playbook 15 (Records, Retention).

The Curriculum-of-Six

The Curriculum-of-Six is the framework’s minimum-viable training scope. Each topic is taught in accessible, practical language with hands-on practice rather than theoretical discussion. The discipline is to teach the moves, not the AI theory underneath them.

Topic What the responder must be able to do
1. Safe modes Activate M0 through M5 per Kill-Switch Modes for any agent in the AI-BOM. Identify which mode the scenario calls for. Verify activation. Document the activation in the response log. Roll back to M0 when the response sequence supports it. Recognize the M3 variants (M3-RAG, M3-Workflow, M3-Output, M3-Vendor, M3-Delegation Cap, M3-Drift) and apply them where appropriate
2. Tool tiering Classify each tool an agent uses as Tier-T0 (low risk), Tier-T1 (medium risk), or Tier-T2 (high risk) per Playbook 04 (Tool Design). Read the agent’s Privilege Matrix entry. Disable T2 tools while preserving T0/T1 capability. Understand the operational meaning of each tier classification for the response scope
3. Retrieval traces Identify which corpora the agent retrieved from in the incident window. Read retrieval-trace logs from the customer’s vector store or RAG framework. Determine which documents influenced the agent’s outputs and at which corpus version. Apply Playbook 03 (RAG Forensics) seven-component pipeline forensics
4. Tool-call logs Read the agent’s tool-call ledger (Type B evidence). Correlate tool calls with downstream SaaS audit records (Type F evidence) using the customer’s correlation identifier per Playbook 19 (Build vs Buy). Identify denied calls (evidence of intent) alongside successful calls (evidence of impact)
5. Memory state Determine whether the agent has memory enabled and at what scope (off, per-user, shared) per the AI-BOM. Snapshot memory before any rotation or cleanup per Playbook 12 (Insider Threat 3.0) discipline. Understand memory bleed across users as its own incident class
6. Configuration snapshots Pull the agent’s current configuration: system prompts, tool definitions, policies, retriever settings, memory configuration, model version pin. Compare against the customer’s last known-good baseline per Playbook 22 (Model and Policy Drift). Maintain configuration as evidence per Playbook 15 (Records, Retention) Two-Tier Retention Standard

Training on topics beyond the Curriculum-of-Six is optional but recommended. The minimum-viable training scope is competence on these six; everything else is depth and specialization.

The Four Core Moves

Every drill rehearses the same four operational moves. Mastery is measured by execution under time pressure rather than by curriculum-coverage breadth.

Move Operational specification Measured target
1. Activate safe mode Locate the agent in the AI-BOM, identify the appropriate Kill-Switch Mode, execute the activation through the customer’s documented mechanism, verify activation, log the action Time-to-Safe-Mode (TTSM) ≤ 10 minutes
2. Preserve and export evidence Run the Evidence Export Script Contract for the Minimum Evidence Set A through F against the affected agent and time window; verify completeness; document gaps Time-to-Evidence (TTE) ≤ 60 minutes; export completeness score against six types
3. Scope the impact in business terms Translate the technical evidence into business-impact statements: which customers affected, which records affected, which downstream actions taken, which regulatory or contractual obligations triggered Brief delivered inside 30-minute window; uses business-impact language rather than technical-detail language
4. Communicate findings with disciplined language Apply the Playbook 17 (Communication) Three-Status Taxonomy (Confirmed, Suspected, Validating), the Four-Element Update Standard (factual impact, immediate containment, evidence-collection activity, next-update timing), and the Responsible Reframing discipline (system-accountability language rather than anthropomorphic attribution) 5-bullet update with status tags on every claim; zero anthropomorphizing language; next-update time named

The Two Permanent Roles

The customer’s AI IR team has two named roles that own specific drill-and-incident responsibilities. The roles are documented in the AI-BOM per agent or per fleet; the named role-holders rotate but the role itself is permanent.

Role Responsibility Drill-day function
Safe Mode Owner Owns the kill-switch activation mechanism per agent. Validates that M0 through M5 (and the M3 variants) are operable at any given time. Maintains the runbook for safe mode activation. Tracks Time-to-Activate metrics per Playbook 13 (Six Metrics) Metric 4 During the drill, the Safe Mode Owner verifies the activation sequence and signs off on Time-to-Safe-Mode measurement
Evidence Owner Owns the evidence-export mechanism per agent. Validates that the Evidence Export Script Contract is operable for the six evidence types. Maintains the runbook for evidence export. Tracks Time-to-Evidence metrics per Metric 3. Coordinates with Playbook 23 (Logging and Privacy) discipline for access governance During the drill, the Evidence Owner verifies the export sequence and signs off on Time-to-Evidence measurement

Drill-artifact archive

Each drill produces evidence artifacts that enter the customer’s records discipline:

The drill artifacts are retained at the metadata tier per Playbook 15 Two-Tier Retention Standard (typically 3 to 5 years) so the customer’s training discipline is demonstrable to auditors, regulators, and the board on demand.

Operational requirement: the monthly micro-drill must run at the documented cadence. A month without a drill is itself a finding. The drill cadence reverts to twice-monthly during the 60 days following a real incident, so the post-incident retrospective findings get reinforced through repeated execution rather than absorbed into a single annual review.

Recovery Sequence

PB16 recovery addresses three scenarios: restoring training cadence after a cadence lapse, restoring competence after a responder turnover, and restoring training discipline after an audit finding.

Path 1: Restore training cadence after a lapse

A failure mode: the monthly drill cadence has slipped (one or more months without a drill). The recovery sequence:

  1. Document the cadence gap. Record the lapse duration, the contributing factors (competing priorities, drill-substrate gaps that blocked previous drills, responder turnover), and the affected responders.
  2. Run a catch-up drill within 5 business days. The drill follows the standard 30-Minute Micro-Drill structure; the catch-up drill is scenario-realistic rather than reduced-scope.
  3. Diagnose the root cause. A single-month lapse is typically operational; a multi-month lapse suggests a structural issue with the customer’s training prioritization or substrate readiness.
  4. Update the cadence-protection mechanism. The customer’s training discipline includes a documented owner whose role responsibilities include drill execution; cadence drift is itself a metric for that owner.

Path 2: Restore competence after responder turnover

A failure mode: a fully-trained responder leaves the team; the team’s solo-on-call coverage is materially diminished. The recovery sequence:

  1. Activate the buddy-system on-call posture from the Competence-scope containment until the new responder reaches solo eligibility.
  2. Compress the onboarding sequence where appropriate. New responders with prior IR experience and prior AI familiarity may complete the onboarding sequence faster than the standard 4-week pattern; the discipline is to compress based on demonstrated competence rather than seniority.
  3. Document the knowledge transfer from the departing responder. Specific agent quirks, customer-specific runbook tweaks, and informal knowledge are documented before the departure rather than reconstructed afterward.
  4. Update the team’s resilience posture. A team that has been single-point-of-failure on a fully-trained responder has a hidden risk that the responder’s departure exposed; the customer’s medium-term posture grows the trained-responder count to a documented minimum (typically 3 for adequate coverage).

Path 3: Restore training discipline after an audit finding

A failure mode: an internal audit, external audit, or regulator review has surfaced that the customer’s AI IR training is insufficient. The recovery sequence:

  1. Categorize the finding. Substrate-side findings (the documentation is missing, the runbook is unclear, the access path is broken) are different from competence-side findings (the responder cannot perform the move). Each requires a different corrective.
  2. Apply substrate fixes first. A responder who cannot perform a move because the substrate is broken does not need more training; they need a fixed substrate.
  3. Build curriculum to close the competence gap. Each competence-side finding produces a specific curriculum addition with measurable target.
  4. Run validation drills. The customer demonstrates the corrective through 2-to-3 consecutive drills that exhibit the previously-failing capability inside the framework’s time budgets.
  5. Document the closure in the customer’s posture artifact. The audit-response artifact references the framework’s training discipline, the specific curriculum addition, and the validated drill outcomes as the empirical closure evidence.

Approver for recovery actions: Training lead, in consultation with the CISO and the Incident Commander. The customer’s training discipline is owned at the IR-program level; ad-hoc recovery decisions by individual responders without the Training lead’s authorization tend to produce inconsistent practice across the team.

Post-Incident Hardening

PB16 hardening organizes around four boundaries. Each has an owner, an artifact, and a measurable acceptance criterion. The four boundaries together convert AI IR training from an annual checkbox into a continuous operational discipline.

Boundary 1: The monthly drill cadence

Boundary 2: The Curriculum-of-Six current and complete

Boundary 3: The two permanent roles assigned and current

Boundary 4: Measurable training targets reported and improved

Common Pitfalls

These are the highest-frequency failure modes in AI IR training. Each has been observed often enough to name as a pattern.

Pitfall Why it happens Consequence
Annual tabletop only, no monthly micro-drill Traditional IR training culture defaults to large quarterly or annual exercises The team’s operational discipline does not develop muscle memory; the first AI incident is the first time the framework’s time budgets are tested under pressure; the response fails the targets
Theory-heavy curriculum The complexity of AI tempts trainers to cover the underlying technology in depth The curriculum produces responders who can discuss model architecture but cannot activate M3 Tool Tiering inside 10 minutes; the framework’s operational claims are unsupported by the team’s actual capability
No measured time-to-safe-mode The drill is treated as a participation exercise rather than a measured execution TTSM drifts upward without anyone noticing; the customer believes the team is at the framework’s standard when the team is materially slower
No measured time-to-evidence The evidence export is treated as a post-incident activity; the 60-minute target is aspirational rather than measured TTE drifts upward; evidence is lost to vendor TTL expiry per Playbook 15; the framework’s evidence claims become indefensible
Drill blocked by substrate, finding absorbed as responder gap The drill identifies an access issue, missing log, or unclear runbook; the team treats it as something the responder should have figured out The substrate gap recurs in real incidents; the team has misallocated training cycles to capability development when substrate fix was the actual corrective
No Safe Mode Owner role The kill-switch activation is treated as collectively-owned Activation drifts because nobody is accountable for keeping it operable; M3 variants in particular suffer because they are agent-specific and require active maintenance
No Evidence Owner role The evidence-export pipeline is treated as collectively-owned Export gaps accumulate (per-type retention drift, access path changes, vendor TTL shifts); the 60-minute discipline becomes unmeasurable
Drill participants opt-out from the on-call responder The same senior responder runs every drill The team’s competence concentrates in one person; the buddy-system or rotation discipline is implicit; the senior’s eventual departure exposes the gap
No catch-up drill after a cadence lapse A missed month is treated as routine Multi-month lapses compound; the team’s competence regresses; the framework’s training claims are not empirically valid
Drill scenarios divorced from the customer’s actual deployment Generic drill scenarios are used regardless of the customer’s specific agent landscape The team rehearses against fictional incidents; real-incident response surfaces gaps the generic drills did not exercise
Punitive response to drill failures The customer’s culture treats drill outcomes as individual-performance ratings Responders avoid drill ownership; substrate and curriculum gaps are obscured by responder defensiveness; the drill discipline becomes ceremonial
No curriculum update after framework revisions New M3 variants, new playbooks, or new evidence-type extensions ship; the training curriculum does not absorb them The team’s discipline reflects the framework as it was at training time, not as it is at incident time; gaps accumulate at every framework update
No board-reported training metrics Training is treated as internal-operations rather than governance-visible Training cadence drifts because no senior leader sees the lapse; the customer’s PB24 scorecard’s operational-readiness claim is not empirically validated
No external mutual-aid relationship The customer’s team is assumed sufficient for all foreseeable incidents A scale or scope incident overwhelms the team; the customer’s response capacity exceeds the team’s trained capability without a planned escalation path
Single-trained-responder dependency The on-call rotation runs with one fully-trained responder and informal backups Single-point-of-failure on responder availability; vacation, illness, or departure exposes the gap during the response window

The Question to Carry Forward

If your AI agent caused a customer-visible incident at 14:00 next month, could the on-call responder activate the appropriate safe mode inside 10 minutes without engineering assistance? Could they export the Minimum Evidence Set inside 60 minutes? Could they produce the 5-bullet executive update inside 30 minutes with status tags on every claim and zero anthropomorphizing language? Could they coordinate the multi-stakeholder response across Security, Privacy, Legal, and Engineering without dropping any stakeholder? Could they do all of this at 02:30 on a Saturday morning when the senior responder is on vacation and the buddy-system backup is on a different continent?

The honest answer is the gap. If any of those answers is “only if I’m the one on call” or “only during a drill, not under real pressure”, the monthly micro-drill cadence, the Curriculum-of-Six, the two permanent roles, or the measurable training targets is the corresponding hardening priority.

AI incidents arrive with operational complexity, regulatory exposure, stakeholder anxiety, and time pressure in the same hour. The framework’s documentation is the customer’s reference; the team’s training is the customer’s capability. The framework’s job is not to choose between depth of documentation and depth of training; it is to make both load-bearing from the first incident. When the customer’s monthly drill cadence and the Four Core Moves and the two permanent roles are operational, the team’s capability becomes a credibility multiplier for the framework’s response claims rather than the gap that surfaces when the first real incident arrives.


Source: AI IR Overlay newsletter, Issue #16, “Training Your Team for AI Incidents: An Operational Approach to AI Incident Response,” by Jacob Ideji. https://www.linkedin.com/in/jacobideji/