Skip to the content.

Changelog

All notable changes to the AI IR Overlay framework live here.

The format follows Keep a Changelog. Versioning follows SemVer.

During the v0.x series, each substantive content drop ships as its own MINOR release. v1.0.0 arrives once the framework core is stable, the remaining playbooks are live, and a Steering Committee is in place.

Unreleased

Planned

v1.1 scope backlog (acknowledged gaps from expert-domain audit)

The following gaps were identified by the v0.33.0 expert-domain audit (CISO, NIST AI RMF/CSF/SP 800-61 r3 contributor, OWASP Agentic working group, securities counsel, GDPR/HIPAA counsel, AI safety researcher, evidence/forensics, ML engineer, AI red-teamer lenses). Each is acknowledged here as a v1.1 candidate; in-place qualification notes are added to the most directly affected files at v0.33.0.

0.35.0 · 2026-07-09 · Governance milestone: Second external Steering Committee member seated

Changed

Why now

v0.35.0 is the framework’s second governance-maturation release. The Steering Committee founding cohort now has 2 of 4 target external members seated:

Both members are participating in their personal capacity. Nothing in this framework should be construed as an endorsement, statement, or position of their current or former employers, including Amazon Web Services and Capital One Financial Corporation.

Dr. Willy Takang brings a multidisciplinary perspective combining cybersecurity, infrastructure engineering, cloud architecture, and secure software development to the framework. His experience spans Linux systems, cloud infrastructure, Kubernetes, containerization, infrastructure automation, Python development, observability, and incident response. He holds a Doctorate in Computer Science, a Master of Science in Information Security, an MBA, and a Bachelor of Science degree. His industry certifications include Cisco Certified Network Associate (CCNA), Cisco Certified Design Associate (CCDA), Cisco Certified Design Professional (CCDP), and AWS Certified Solutions Architect – Associate. His focus areas as a contributor include practical guidance and resources for AI incident response practitioners. His full bio and personal-capacity disclosure are in CONTRIBUTORS.md.

At v0.35.0, the 3-member threshold (Lead Maintainer + 2 or more external Steering Committee members) that the framework’s founding-cohort governance discipline referenced is now reached. Under the current governance framework, the Steering Committee may hold structured advisory votes going forward; formal voting discipline binding the Lead Maintainer activates only at v1.0.0 per the Future governance section of GOVERNANCE.md.

Dr. Willy Takang joins the framework as its second external contributor, with contributions in preparation for upcoming v0.35.x / v0.36.x releases. The specific artifact scope, review discipline, and release cadence for his contributions will be documented as they land.

This release is content-neutral: no playbooks, crosswalks, schemas, or reference implementations changed. The framework’s 24 playbooks, 3 crosswalks, 5 schemas, 2 reference implementations, and all other content remain at v0.33.0 state. What changed is the governance posture: the founding cohort has passed the 3-member threshold and moved from “single external voice” to “collaboratively-formed founding cohort.”

Progress on v1.0.0 governance gates:

0.34.0 · 2026-07-08 · Governance milestone: First external Steering Committee member seated

Changed

Why now

v0.34.0 is the framework’s first governance-maturation release. Two of the four v1.0.0 governance-gate items are now on track:

The remaining v1.0.0 governance-gate items are (a) trademark registration (currently unregistered word marks) and (b) additional Steering Committee members through the founding-cohort formation.

This release is content-neutral: no playbooks, crosswalks, schemas, or reference implementations changed. The framework’s 24 playbooks, 3 crosswalks, 5 schemas, 2 reference implementations, and all other content remain at v0.33.0 state. What changed is the governance posture: the framework is now a collaboratively-governed project rather than a single-maintainer artifact. Adopters, auditors, and readers can continue relying on the v0.33.0 content without any changes at this release.

The naming of Chukwunenye Amadi as a Steering Committee member is a personal-capacity participation. Nothing in this framework should be construed as an endorsement, statement, or position of his current or former employers, including Amazon Web Services.

0.33.0 · 2026-06-29 · Framework Matrix introduced (MATRIX.md)

Added

Changed

Why now

The framework reached content-gate completeness at v0.24.0 (all 24 playbooks shipped), repo-meta polish completeness at v0.32.0 (CONTRIBUTORS.md introduced, reading order restructured, framework/04 reference list expanded, README current-release line added, crosswalk attribution gaps closed, acronyms section expanded to 25 entries). What remained was an orientation gap: a board prospect, an onboarding engineer, or an auditor arriving at the repo had no single document that surfaced the entire framework in tabular form. The README walks the conceptual order across 15 reading-order items; CONTENT_MAP narrates how the pieces fit; the per-playbook prose covers depth. None of these compress the framework to the matrix view that a reader can absorb in 90 seconds before deciding where to dive in.

v0.33.0 closes that orientation gap with a single self-contained matrix document calibrated to four audiences:

The matrix is structurally distinct from MITRE ATT&CK Insights pages: ATT&CK Insights are narrative summaries with selected technique highlights, while this is the matrix view itself with full enumeration. The structural analogue is the ATT&CK matrix poster, not the Insights page.

After v0.33.0, the framework’s content, calibration, polish, repository-meta discipline, attribution discipline, acronym discipline, and tabular-orientation discipline are all at their highest-ever state. The remaining v1.0.0 blockers continue to be governance maturation (Steering Committee announcement, external contributors, trademark registration, production case studies) rather than content or orientation gaps.

0.32.0 · 2026-06-29 · P2 polish sweep (README reading-order completeness, CONTRIBUTORS.md initial file)

Added

Changed

Why now

This release closes the final 3 P2 polish items from the v0.31.0 holistic Steering-Committee-readiness critique:

P2.10 (framework/04 orphaned from README reading order): the v0.25.0 P0.2 calibration established framework/04 as the canonical source for the Materiality and Disclosure convening trigger. The v0.29.0 materiality canonicalization ripple added six playbooks (PB01, PB06, PB09, PB10, PB21, PB22) to the canonical convening list; combined with PB05, PB15, and PB23 (which referenced framework/04 from their own ship dates at v0.24.0, v0.19.0, and v0.20.0), nine playbooks convene the call as part of First-Hour Actions, with PB18 and PB24 gating downstream discipline on the materiality determination. But the README’s reading order had not been updated to reflect framework/04’s elevated role; it was listed only as a triage sub-reference (item 4’s link to the Six Triage Questions). v0.32.0 promotes framework/04 to a top-level reading-order item (item 4), making it discoverable as a framework foundation document on first read rather than as a cross-reference discovered through playbook navigation.

P2.11 (examples/incident-walkthrough.md is buried): the worked example is the framework’s strongest single adoption artifact (the synthetic incident response showing the framework operating as a coherent system end-to-end). Prior to v0.32.0, it was mentioned only in the “New here?” preamble and in QUICKSTART link references. A reader who proceeded sequentially through the README’s reading order would not encounter the walkthrough until they hit it incidentally through QUICKSTART. v0.32.0 elevates it to item 8 of the reading order, between the conceptual core (items 1-7) and the working artifacts (items 9-14). This positioning is intentional: a reader who has internalized the MVO, Mental Model, Maturity Roadmap, Materiality framework, Triage Questions, Kill-Switch Modes, and Minimum Evidence Set is exactly the audience the walkthrough is calibrated for. The walkthrough lands as the synthesis that connects the conceptual core to the operational artifacts.

P2.12 (No CONTRIBUTORS.md file): the GOVERNANCE.md file referenced CONTRIBUTORS.md as a future artifact (“once external contributors exist”) but the file did not exist. The chicken-and-egg problem: a prospect evaluating the framework’s contribution readiness has no signal that the contribution discipline is operational because the recognition artifact is absent. v0.32.0 ships an initial CONTRIBUTORS.md that names Jacob Ideji as Founding Maintainer, lists External Contributors as “None yet” with explicit framing, documents the five highest-velocity contribution paths, and establishes the recognition discipline (“lightweight and inclusive; every accepted PR earns a place”). The file itself signals readiness for contributions even at zero-external-contributor state.

After v0.32.0, all P0, P1, and P2 items from both the v0.24.0 original critique and the v0.27.0/v0.31.0 holistic re-audits are closed. The framework’s content, calibration, polish, and repository-meta discipline are at their highest-ever state. The remaining v1.0.0 blockers are governance maturation: production case studies, external contributions, trademark registration, and the Steering Committee announcement itself (the v1.0.0-rc1 cut). None of these are content gaps; all are lifecycle-stage gaps that no further calibration release can close.

0.31.0 · 2026-06-29 · P2-NEW sweep (Three Realities in PB01 First-Hour, maturity_target default docs, PB08 multi-agent depth)

Changed

Why now

This release closes three P2 items remaining from the v0.24.0 original critique and the v0.27.0 re-audit:

P2-NEW.1 (Three Realities in PB01 First-Hour Actions): the v0.25.0 P0.3 calibration added the Three Realities review to PB18 Boundary 4 (post-incident retrospective). The v0.27.0 re-audit identified that the Three Realities were enforced post-incident but not surfaced as a response-phase checkpoint, creating the risk that response teams under operational pressure would not apply the foundational lens during the response itself. v0.31.0 adds the explicit reflex to PB01’s First-Hour Actions, closing the response-phase gap.

P2-NEW.2 (maturity_target default behavior): the v0.26.0 P1.4 calibration added the maturity_target field to the AI-BOM schema. The v0.27.0 re-audit identified that the default behavior when the field was omitted was undocumented. v0.31.0 clarifies the framework’s “no silent default” position: schema requires the field; omission fails validation; adopters must explicitly opt-in to a maturity level.

P2.11 (PB08 multi-agent depth): the v0.24.0 critique identified that PB08 was light on multi-agent depth relative to 2026 production reality (26 KB vs newer playbooks at 35-50 KB). v0.31.0 adds the Multi-Agent Threat Patterns section with five named patterns covering orchestrator compromise, peer-to-peer compromise, MCP/A2A supply-chain compromise, memory-mediated cross-agent compromise, and cascading failure through trust chains. Each pattern names attack mechanism, topology surface, indicators, and containment variant.

After v0.31.0, all P0 and P1 items from both the v0.24.0 original critique and the v0.27.0 re-audit are closed, plus three P2 items. The one remaining P2 item (P2.12 No formal threat-model artifact) is explicitly framed in v0.27.0 as a future enhancement, not a v1.0.0 blocker. The framework’s only remaining v1.0.0 blocker is the Steering Committee announcement (governance gate, not content gate).

0.30.0 · 2026-06-29 · P1-NEW.3: QUICKSTART-startup Week-0 Pre-Adoption Readiness Check

Changed

Why now

This release closes P1-NEW.3 from the v0.27.0 holistic re-audit. The v0.26.0 initial QUICKSTART-startup release identified the 3-playbook + 2-template minimum subset and the 4-week adoption path, but the v0.27.0 re-audit found that the 4-week timeline was structurally bounded by three pre-Week-1 preconditions that the QUICKSTART did not surface:

  1. Vendor-copilot bottleneck: a startup’s Week 3 tabletop discovers that M3 (Tool Tiering) for a vendor copilot requires vendor support with a 24-72 hour SLA, not a 10-minute TTA. The maturity claim lands as “Level 2 for customer-managed agents, Level 1 for vendor copilots.” False confidence about maturity level.

  2. Evidence export infrastructure does not exist: a startup begins Week 1 with zero understanding of vendor log retention, API export mechanisms, or legal holds on evidence. By Week 4, they have not validated whether evidence export is even possible. Months of wasted effort claiming Level 2 while being structurally unable to reach Level 3.

  3. Tool reversibility discovery creates mid-project redesign: Week 2 tool tiering surfaces that 30-40% of high-risk tools cannot be reversed. The startup discovers it cannot deploy M1 (Read-Only) or M3 (Tool Tiering) as designed; all containment defaults to M4 (Full Disable), which breaks business continuity. Week 2-3 derails into tool-design discussions; the 4-week timeline becomes 6-8 weeks.

v0.30.0 closes the discovery-friction gap by adding the explicit Week 0 readiness check. The pattern: surface the structural preconditions as Week 0 actions rather than discovering them mid-project. A startup that completes Week 0 honestly will either confirm the 4-week timeline is achievable OR document the specific extension needed (1-3 additional weeks for vendor coordination, 2-4 additional weeks for tool-reversibility redesign, etc.). Either outcome is materially better than the prior “discover-during-Week-3” pattern.

Three secondary improvements:

After v0.30.0, the QUICKSTART-startup path is structurally honest about the discovery work the original 4-week timeline elided. The remaining P2-NEW items (Three Realities in PB01 First-Hour Actions, maturity_target default behavior documented) are polish-grade and not blockers for v1.0.0-rc1.

0.29.0 · 2026-06-29 · P1-NEW.2: Materiality canonicalization ripple closure (PB10/21/22)

Changed

Why now

This release closes P1-NEW.2 from the v0.27.0 holistic re-audit. The v0.25.0 P0.2 calibration established framework/04-materiality-and-disclosure.md as the canonical source for the convening trigger and updated PB06 and PB09 to reference it. The v0.27.0 re-audit identified that three additional playbooks (PB10, PB21, PB22) still restated trigger conditions locally rather than referencing the canonical list, producing the same drift risk the v0.25.0 fix set out to prevent.

v0.29.0 closes the ripple gap. After this release, all 6 playbooks that convene the Materiality and Disclosure call (PB01, PB06, PB09, PB10, PB21, PB22) reference the canonical trigger from framework/04 rather than restate it. PB05 introduces the canonical framing; PB18 verifies the convening determination is documented; PB24 audits the convening discipline at the scorecard level.

The pattern preserved across these playbooks: reference the canonical first; then add the playbook-specific commentary about which triggers are most commonly applicable for that incident class (vendor copilots, shadow agents, drift events). This pattern keeps the canonical source authoritative while allowing each playbook to honestly describe its scenario’s most-common convening conditions.

After v0.29.0, the framework’s calibration is materially more complete: 8 of 24 playbooks operationalize CIA+T (v0.28.0); 6 playbooks reference canonical materiality (v0.29.0); reference implementations are contract-conformant (v0.28.0); repo hygiene is clean. The remaining P1-NEW/P2-NEW items (QUICKSTART-startup pre-Week-0 checklist, Three Realities in PB01 response phase, maturity_target default behavior) are not blockers for v1.0.0-rc1.

0.28.0 · 2026-06-29 · P0-NEW sweep (Evidence Exporter conformance, CIA+T to PB12/21/22/23, pycache cleanup)

Changed

Pre-release housekeeping

Why now

This release closes three of the most impactful new findings from the v0.27.0 second holistic critique:

P0-NEW.1 (pycache hygiene): repo cleanliness; the .gitignore catches future git-tracked commits but doesn’t filter web-UI folder-uploads. Documented in this release for manual cleanup.

P0-NEW.2 (Evidence Exporter contract deviations): the most impactful single fix in this calibration cycle. The v0.26.0 reference implementation was actively misleading: adopters who forked it would ship non-compliant evidence artifacts. The rewrite restores conformance and demonstrates the contract elements (JSONL discipline, manifest schema, exit-code semantics, validate-access pre-flight) that the spec specifies. The kill_switch_demo did not have equivalent deviations and is unchanged in this release.

P0-NEW.3 (CIA+T ripple gap): the v0.25.0 P0.1 fix propagated CIA+T into the four playbooks named in that release (PB05, PB09, PB17, PB24). The v0.27.0 re-audit identified that the same framing should apply to four additional sensitive-incident playbooks: PB12 (Insider Threat), PB21 (Shadow AI), PB22 (Drift), PB23 (Logging and Privacy). v0.28.0 closes the ripple gap with CIA+T sections calibrated to each playbook’s specific incident class. Notably, each section introduces a Trust-dimension insight specific to that scenario: insider-threat triad-vs-CIA+T complementarity (PB12); latent regulatory exposure (PB21); silent customer-trust erosion (PB22); separate disclosure tracks (PB23).

After v0.28.0, the framework’s calibration story is more complete. The remaining P1/P2 findings from the v0.27.0 re-audit (materiality canonicalization ripple in PB10/19/21/22; Three Realities in PB01 response phase; QUICKSTART-startup pre-Week-0 checklist) are not yet addressed and remain open for a future calibration release.

0.27.0 · 2026-06-29 · P2 polish (CHANGELOG comma-annotation, PB02 mental-model cross-reference)

Changed

Why now

This release closes the two remaining P2 polish items from the v0.24.0 holistic critique:

After v0.27.0, the framework is content-complete (v0.24.0) + consistency-calibrated (v0.25.0) + adoption-experience-calibrated (v0.26.0) + cosmetically polished (v0.27.0). The remaining v1.0.0 work is governance maturation only (Steering Committee announcement, public-interface freeze, v1.0.0-rc1 release candidate).

0.26.0 · 2026-06-29 · P1 adoption-friction fixes (maturity-target schema, validator staleness, reference implementations, startup QUICKSTART)

Added

Changed

Why now

The v0.24.0 holistic critique identified four material adoption-friction items (P1.4 through P1.7). v0.26.0 closes all four in a single calibration release. The release does not change the framework’s content gate (24 playbooks shipped per v0.24.0); it improves the adopter experience for customers who would otherwise stall on schema rigidity, validator under-enforcement, missing reference code, or framework-density overwhelm.

P1.4 (Schema rigidity for early adopters): The v0.14.0 schema required all four kill-switch modes to be implemented=true for any AI-BOM to validate. This blocked early adopters at Level 1 (Aware) who had only inventory in place and were still building M1-M4. v0.26.0 introduces the explicit kill_switches.maturity_target field with maturity-conditional enforcement: Level 1 customers may declare unimplemented modes during initial adoption; Level 2 and above require full M1-M4 implementation. The framework’s discipline of honest self-assessment is preserved by requiring explicit level_1_aware declaration (no blank fields).

P1.5 (Validator under-enforcement of staleness): The framework’s MVO conformance criteria specified 7-day last_reviewed and 90-day tested_at windows, but the JSON schema had no maximum-date constraint and the Python validator did no temporal logic. AI-BOM files could pass validate.py while violating the operational SLA. v0.26.0’s validator now enforces staleness with a --strict flag for CI enforcement; permissive mode reports staleness as WARNINGs to support iterative adoption.

P1.6 (No reference implementations of the API contracts): The framework’s two operational contracts (evidence-export.spec.md and kill-switch-api.md) were exceptionally detailed but lacked working code. Every adopter re-invented the implementation independently, losing the convergence benefit the specs were designed to achieve. v0.26.0 ships minimal Python implementations of both contracts with stub adapters that demonstrate the contract’s shape (manifest discipline, integrity hashes, telemetry events, probe-after-activate, separation-of-duties). Adopters fork and replace the stubs with vendor-specific implementations; the contract conformance is preserved through the adapter swap.

P1.7 (No startup-minimum subset): The standard QUICKSTART.md targets well-resourced security teams with full platform control. For a 5-person team or a vendor-copilot-dependent organization, the 30-day path slipped to 6-8 weeks and faced blockers (instrumentation gaps, identity scope misalignment, control-plane access). v0.26.0 ships QUICKSTART-startup.md as the explicit minimum-viable alternative: 3 playbooks + 2 templates + 1 triage card, 4-week path, Level 2 (Containable) target, with explicit acknowledgment of what is deliberately deferred.

After v0.26.0, the framework’s content is structurally complete (per v0.24.0) AND adopter-experience-calibrated. The remaining v1.0 work is governance (Steering Committee announcement, public-interface freeze, v1.0.0-rc1 release candidate).

0.25.0 · 2026-06-29 · P0 consistency calibration (CIA+T propagation, materiality trigger canonicalization, Three Realities retrospective)

Changed

Why now

This release is a P0 consistency calibration following the v0.24.0 holistic quality critique. Three material consistency issues were identified between PB05 (the newly-shipped executive decision-making playbook), the canonical materiality framework, and the operational playbooks that reference both. The release does not add new playbooks or new disciplines; it propagates existing disciplines into the playbooks where they were missing, closing daylight that a regulator review or hostile audit would surface.

P0.1 (CIA+T propagation): Playbook 05’s CIA+T framing (elevating Trust to a peer dimension alongside Confidentiality, Integrity, and Availability) was the framework’s most defensible single innovation for executive decision-making in non-classic-breach AI incidents. However, the framing was operationalized only in PB05 itself. PB09 (Output Leakage), the framework’s canonical CIA+T scenario, applied traditional CIA framing without Trust. PB17 (Communication) drafted templates without referencing the impact taxonomy. PB24 (Scorecard) did not include a Trust-dimension scorecard item. A response team using PB09 for an output-leakage incident would not produce a Trust-dimension assessment. v0.25.0 closes the propagation gap: PB09 has the operational Impact-assessment table; PB17 maps CIA+T to each stakeholder class; PB24 has C5 as the scorecard signal.

P0.2 (materiality convening trigger canonicalization): The trigger for convening the Materiality and Disclosure call was stated in four different framings across the framework. PB05 used condition-based framing. PB06 used mode-based (“M3 or higher OR external recipients OR regulated data”). PB09 used the most permissive (“regardless of mode”). PB24 used a process-only formulation (“documented for M3 or higher”). The drift was operationally low-risk (the framework favors over-convening; under-convening is unlikely) but auditably visible. v0.25.0 establishes framework/04 as the canonical source of the trigger definition and reframes the downstream playbooks to reference rather than restate. The customer-facing trust impact trigger is added to the canonical list per the PB05 CIA+T framing.

P0.3 (Three Realities application review in PB18): PB02’s Three Realities were enforced as a curriculum prerequisite in PB16 onboarding and as a drill-evaluation criterion in the monthly micro-drill cadence, but no playbook closed the loop on the post-incident retrospective. A response team could complete an incident, generate the PB18 5-business-day hardening backlog, and never evaluate whether the Three Realities were applied during the response. v0.25.0 adds the Reality-application review as an explicit Boundary 4 bullet in PB18, closing the loop between conceptual foundation (PB02), training cadence (PB16), and operational hardening (PB18).

After v0.25.0, the framework’s content is structurally complete (24 of 24 playbooks shipped per v0.24.0) AND consistently calibrated across the playbooks that reference shared disciplines. The next framework work is governance maturation (Steering Committee announcement, public-interface freeze, v1.0.0-rc1 release candidate).

0.24.0 · 2026-06-29 · Playbook 05: Executive Decision-Making With AI in the Loop (content gate complete)

Added

Changed

Why now

PB05 closes the executive-decision-making discipline that is the final piece of the framework’s content gate. Every prior playbook specifies what the response team does (technical playbooks) or what the response team says (PB17) or how the team is trained (PB16) or how evidence is captured, retained, and proved (PB02, PB15, PB23). None of these specify how the executive team makes the decisions that the regulator, the customer, the board, and the press will hold the customer to. Executive decisions during AI incidents are made under three converging pressures (unfolding technical situation, running disclosure window, stakeholder anxiety) that produce a distinctive failure pattern that traditional-IR executive briefings do not address.

PB05 addresses this with the Executive Decision Packet (AI Edition) (the five-section structured update that gives executives decision-support rather than status-narration), the CIA+T framing (elevating Trust to peer status with Confidentiality, Integrity, and Availability because AI incidents commonly produce trust impact that exceeds traditional-CIA impact even when no classic breach has occurred), the 4-hour cadence (an explicit time budget for executive-layer information density), the 4/24/72-hour planning horizons (the structured forward-look that prevents decisions from being made on immediate-only information), and the Approval Receipt discipline (the four required elements that prevent human approval workflows from degrading into rubber-stamping). The playbook makes the difference between a defensible incident response and a defensible response coupled with credible executive accountability an empirical question (does the customer’s IC produce the first Decision Packet inside 60 minutes? does the Approval Receipt discipline apply to every Tier-T2 action?) rather than an asserted claim.

The playbook completes the framework’s executive-layer trio with PB17 (Communication Techniques) and PB24 (Board-Ready Scorecard). PB05 is the decision-during-incident; PB17 is the communication-of-the-decision; PB24 is the periodic governance review. Together they convert the framework’s executive-readiness from a written commitment into a measurable, drillable, board-defensible discipline.

After v0.24.0, the framework’s content gate is complete. All 24 playbooks (PB01 through PB24) are shipped. The framework’s coverage of the operational arc (Foundation, Prevention, Closure, Governance, Measurement and Depth, Operations), the six preconditions (procurement, inventory, change-event, proof, privacy, communication), the four discipline pairs (concepts-and-operations, capture-retain-prove triad, testing-and-training, governance-and-communication), and the executive-layer trio is comprehensive. The remaining v1.0 work is governance: the Steering Committee announcement and the public-interface freeze that converts the framework from single-maintainer pre-1.0 into a sustainable multi-maintainer artifact.

PB05 also closes a long-standing standards-gap pattern: NIST CSF 2.0 mandates organizational oversight (GV.OV) and risk-management decision-making (GV.RM) but does not specify the AI-specific executive-decision discipline. NIST AI RMF mandates organizational accountability for AI risk (GOVERN 4.1) but does not specify the operational mechanism. OWASP’s ASI09 Human-Agent Trust Exploitation addresses the trust-exploitation risk but does not address the executive-decision-making discipline that operationalizes accountable decisions under uncertainty. PB05 fills each of those gaps with concrete operational specifications that customers can adopt, regulators can audit against, and adopters can extend.

0.23.0 · 2026-06-29 · Playbook 02: Evidence Lives in New Places (foundational concepts)

Added

Changed

Why now

PB02 closes the conceptual-foundation gap that the framework’s operational playbooks have been depending on without specifying as a separate artifact. Every prior playbook references the Minimum Evidence Set A-F taxonomy and the Capture Order discipline; both flow from the three foundational realities of AI evidence that newsletter Issue 2 introduced. The original maintainer decision in v0.1.0 was to absorb PB02 into the framework core (evidence/minimum-evidence-set.md) because the operational substance (the A-F taxonomy, the Capture Order, the pitfalls list) was fully captured there. The decision had three downstream effects that v0.23.0 resolves:

First, the Three Realities were not named as principles. Issue 2’s conceptual contributions (“the actor is a workflow, not a workstation”; “the payload can be language, not malware”; “evidence is fragile”) were referenced implicitly across the framework but never named explicitly. Responders could read the operational playbooks without internalizing the mental shifts that the operations are built on; PB02 makes the realities explicit and named.

Second, the “every newsletter issue maps to one playbook in the framework” provenance principle from the README’s Provenance section was violated. The v0.23.0 promotion restores the principle: PB02 now has its own playbook (the conceptual companion) alongside the framework-core operational specification.

Third, newcomers lacked a pedagogical entry point. The framework’s existing entry points (QUICKSTART for 30-day adoption, examples/incident-walkthrough.md for an end-to-end worked example) are action-oriented and scenario-oriented respectively. PB02 adds a concepts-oriented entry point: “first principles of AI evidence, before any specific scenario.” The three entry points (action, scenario, concept) together cover the major learning modalities for newcomers.

The playbook uses a modified canonical skeleton that preserves coherence with the framework’s standard 9-section structure while adapting content to the foundational nature: First-Hour Actions become first-hour reflexes that apply to every AI incident; Containment Options become state-preservation discipline; Evidence Priorities map the Three Realities to the A-F taxonomy; Recovery Sequence is conceptual rather than operational. The structural coherence supports framework-wide consistency; the content adaptation honors PB02’s foundational pedagogical role.

The v1.0 criteria are updated accordingly: the target is now 24 playbooks (PB01 through PB24, no absorption), 23 of which are shipped after v0.23.0. The remaining playbook is PB05 (Executive Decision-Making With AI in the Loop). After PB05 ships, the content gate is fully closed and the v1.0 cut turns entirely on the Steering Committee announcement (the governance gate).

0.22.0 · 2026-06-29 · Playbook 16: Training Your Team for AI Incidents

Added

Changed

Why now

PB16 closes the training-discipline precondition that the framework’s operational playbooks have been depending on without specifying. Every prior playbook assumes the response team can execute its discipline under operational pressure: activate M3 Tool Tiering inside 10 minutes, export the Minimum Evidence Set A through F inside 60 minutes, issue the first stakeholder update inside 30 minutes, coordinate across Security, Privacy, Legal, and Engineering without dropping a stakeholder. None of those time budgets survive untrained execution. Traditional annual-tabletop IR training does not produce the muscle memory the framework’s targets require; the first AI incident becomes the first time the team’s discipline is tested under pressure, and the framework’s claims surface as gaps at the moment they are most damaging.

PB16 addresses this with the 30-Minute Micro-Drill (the framework’s primary training artifact, run monthly against a representative agent), the Four Core Moves (activate safe mode, preserve and export evidence, scope impact, communicate with disciplined language: the four operational moves rehearsed in every drill), the two permanent operating roles (Safe Mode Owner who owns the kill-switch mechanism per agent and validates that M0-M5 and the M3 variants are operable; Evidence Owner who owns the evidence-export mechanism and validates the Type A-F export pipeline), the Curriculum-of-Six (safe modes, tool tiering, retrieval traces, tool-call logs, memory state, configuration snapshots: practical hands-on coverage rather than AI theory), the monthly drill cadence (twice-monthly during the 60 days following a real incident so retrospective findings get reinforced through repeated execution), and the measurable training targets that feed PB13 Six Metrics and PB24 board scorecard. The playbook makes the difference between a documented framework and an executable framework an empirical question (does the on-call responder hit the framework’s time budgets in the monthly drill against a representative agent?) rather than an asserted competency.

The playbook completes the framework’s testing-and-training pair with PB14 (Testing for Agent Failure Modes). PB14 is system-side testing: does the substrate support the kill-switch ladder? Can the Drift Canary catch drift? Can the Reconstructability Test pass at 30 days? PB16 is human-side training: can the responders execute the documented discipline under pressure? Together they convert the framework’s operational claims from written commitments into empirically-validated capabilities. After v0.22.0, the framework’s response posture is testable on both axes: a CISO can use PB14’s quarterly testing cadence to validate the substrate readiness and PB16’s monthly micro-drill cadence to validate the team readiness.

PB16 also explicitly closes a long-standing standards-gap pattern: NIST CSF 2.0 mandates personnel training (PR.AT-01) and specialized-role training (PR.AT-02) but does not specify the AI-specific training discipline. NIST AI RMF mandates documented and trained human-AI roles (GOVERN 3.2) but does not specify the operational mechanism. OWASP’s ASI09 Human-Agent Trust Exploitation addresses the trust-exploitation risk but does not address the operator-training discipline that makes the framework’s countermeasures executable. PB16 fills each of those gaps with concrete operational specifications that customers can adopt, regulators can audit against, and adopters can extend.

After v0.22.0, the framework is structurally complete on content: all six preconditions are closed (procurement, inventory, change-event, proof, privacy, communication), the MVO-3 capture-retain-prove triad is shipped, the testing-and-training pair is complete, and the governance-and-communication discipline pair is operational. The single remaining drafted playbook (PB05 Executive Decision-Making With AI in the Loop) addresses an important human-side dimension but is not a structural precondition closure. The v1.0.0 cut now turns on the Steering Committee announcement (the documented governance gate) rather than the content gate.

0.21.0 · 2026-06-29 · Playbook 17: Communication Techniques for AI-Involved IR

Added

Changed

Why now

PB17 closes the communication-discipline precondition that the framework’s technical playbooks have been depending on without specifying. Every prior playbook specifies what the response team does operationally; none specify what the response team says. The result is that even technically excellent responses can produce poor stakeholder outcomes: a premature root-cause attribution that later evidence contradicts erodes credibility for the rest of the response window; an anthropomorphic attribution (“the AI did it”) undermines the customer’s accountability posture and creates legal exposure; a one-size message drafted for “stakeholders” satisfies none of the actual stakeholder classes; silence in the first 30 minutes is interpreted as either unaware or hiding. None of these failure modes are addressable through better technical response alone.

PB17 addresses this with the 30-minute first-update SLA (the empirical baseline for the customer’s communication-track discipline), the Three-Status Taxonomy (the shared vocabulary that prevents premature commitment to claims that exceed the available evidence), the Four-Element Update Standard (the structural minimum that first updates must contain), the Stakeholder Communication Matrix (the per-audience calibration that prevents one-size messaging), the Template Library (the pre-positioned communication asset that makes accurate, accountable language fast enough for the time pressure), and the Responsible Reframing discipline (the language pattern that operationalizes the Mental Model’s accountability framing in every communication artifact). The playbook makes the difference between a credibility-preserving response and a credibility-eroding response an empirical question (does the customer hit the 30-minute first-update SLA in the quarterly Communication Drill?) rather than an asserted competency.

The playbook completes the framework’s governance and communication discipline pair with PB24 (Board-Ready Scorecard). PB24 specifies what the board sees on quarterly cadence; PB17 specifies how every incident’s stakeholder communications hold trust through the response window. Together they convert the framework’s governance posture from a documentation artifact into a continuous trust-preservation discipline. After v0.21.0, the framework’s stakeholder-trust posture is testable as well as documented: a CISO can use PB17’s quarterly Communication Drill to validate the customer’s communication discipline against the same time pressures that real incidents impose, with the drill findings flowing into PB18’s 5-business-day hardening SLA backlog and PB13’s Six Metrics.

PB17 also explicitly closes a long-standing standards-gap pattern: NIST CSF 2.0 mandates incident communication (RS.CO, RC.CO) and stakeholder context (GV.OC) but does not specify the AI-specific communication discipline. NIST AI RMF mandates incident communication to AI actors and affected communities (MANAGE 4.3) but does not specify the operational mechanism. OWASP’s ASI09 Human-Agent Trust Exploitation addresses the trust-exploitation risk but does not address the trust-rebuilding response. PB17 fills each of those gaps with concrete operational specifications that customers can adopt, regulators can audit against, and adopters can extend.

0.20.0 · 2026-06-29 · Playbook 23: AI Logging and Privacy in a Multi-Stakeholder World

Added

Changed

Why now

PB23 closes the privacy-discipline precondition that the framework’s evidence-side playbooks have been depending on without specifying. Every prior playbook assumes that the evidence captured at incident time can be retained, accessed, and exported in a way that supports the framework’s forensic claims. None of those assumptions survive routine modern privacy reality: AI logs are not metadata-only logs; they contain prompt bodies, response bodies, retrieved document content, tool-call parameters, and memory content that may include PII, PHI, regulated identifiers, business-confidential information, and even credentials. Treating these logs as if they were traditional IR telemetry produces either overcollection findings under data-minimization regulatory regimes (GDPR Article 5(1)(c), CCPA personal-information limitation, HIPAA minimum-necessary standard) or undercollection findings under the framework’s own Reconstructability Test, or both simultaneously.

PB23 addresses this with the Multi-Stakeholder Governance Matrix (Security, Privacy, Legal, Engineering each name a defensible interest, a load-bearing artifact, and an acceptance criterion; a policy any one cannot defend is not yet a policy), the Three-Layer Logging Model (Layer 1 metadata broadly retained and load-bearing for the typical-incident reconstruction; Layer 2 selective payload narrowly triggered by documented high-risk actions, sensitive-corpus access, and active-incident windows; Layer 3 escalation capture under legal hold), the Forensically Useful standard (the six core questions logs must answer at Layer 1 alone, calibrating Layer 2 as the closure layer for specific gaps rather than the default), and the redaction-and-tokenization discipline (structural preservation that retains evidence value while removing sensitive content). The playbook makes the choice between a privacy-defensible posture and a forensic-defensible posture an empirical question (does the Three-Layer Logging Model satisfy all four stakeholder acceptance criteria simultaneously?) rather than a stakeholder-tribal one.

The playbook completes the framework’s capture / retain / prove triad on top of the Minimum Evidence Set: evidence/minimum-evidence-set.md establishes the A-F taxonomy and the 60-minute export discipline; PB23 specifies how each evidence type is captured at Layer 1, Layer 2, or Layer 3 with the multi-stakeholder discipline; Playbook 15 (Records, Retention) specifies how each captured evidence artifact is retained with chain-of-custody integrity through the regulatory and legal review window. After v0.20.0, the framework’s evidence claims are testable on the privacy axis as well as the forensic axis: a Data Protection Officer can use PB23’s Multi-Stakeholder Governance Matrix to validate the customer’s data-minimization posture against regulator review; a CISO can use PB15’s Reconstructability Test to validate the customer’s forensic posture against the 30-, 60-, and 90-day horizons that regulator and legal review windows typically span.

PB23 also explicitly closes a long-standing standards-gap pattern: NIST CSF 2.0 mandates data-at-rest protection (PR.DS-01) and access control (PR.AA) but does not specify the AI-specific multi-stakeholder governance discipline for logs that simultaneously serve as evidence and as regulated-data stores. NIST AI RMF mandates privacy risk measurement (MEASURE 2.10) but does not specify the operational mechanism. OWASP’s LLM02 Sensitive Information Disclosure addresses the data exposure in AI outputs but does not address the exposure in the logs themselves. PB23 fills each of those gaps with concrete operational specifications that customers can adopt, regulators can audit against, and adopters can extend.

0.19.0 · 2026-06-29 · Playbook 15: Records, Retention, and Proving What Happened

Added

Changed

Why now

PB15 closes the proof-discipline precondition that the framework’s response-side playbooks have been depending on without specifying. Every prior playbook assumes the evidence captured at incident time will be available, defensible, and reconstructable when the regulatory, legal, or business-trust review asks for it weeks or months later. None of those assumptions survive routine evidence-retention defaults: vendor TTLs on prompt-and-response logs are 24 to 72 hours; telemetry pipelines truncate event payloads at default size caps; vector indices retain only the current version; storage-tier transitions move evidence from warm to cold to inaccessible on automated schedules; sensitive-data redaction policies applied without forensic awareness destroy payload-class evidence at the same time they protect the data. The first time the customer finds out about an evidence-retention failure is often during the audit, regulator review, or post-incident retrospective where the evidence is needed and discovered missing.

PB15 addresses this with the Two-Tier Retention Standard (metadata-tier and payload-tier windows calibrated per evidence class), the incident-triggered legal-hold mechanism (the hold-scope identification, hold-class retention duration, hold notification, and hold release sequence that extends default retention past its window for events that require it), the chain-of-custody discipline (every access to the evidence store from the moment of incident declaration is access-logged), the tamper-evidence anchor (cryptographic integrity hashes computed at capture time and verifiable on subsequent access), and the quarterly Reconstructability Test that empirically validates the framework’s evidence claims at 30, 60, and 90 days. The playbook makes the difference between a defensible incident and an unprovable one an empirical question (does the Reconstructability Test pass at 30 days for the target scope?) rather than an asserted commitment.

The playbook completes the framework’s MVO-3 Evidence taxonomy depth coverage: evidence/minimum-evidence-set.md establishes the A-F taxonomy and the 60-minute export discipline, Playbook 03 (RAG Forensics) is the Type C deep-dive, Playbook 09 (Output Leakage) is the Type F deep-dive, and PB15 is the lifecycle deep-dive across all six evidence types (capture-to-disposal, two-tier retention, chain of custody, tamper-evidence, and reconstructability). After v0.19.0, the framework’s evidence claims are testable as well as documented: a CISO can use PB15’s Reconstructability Test to validate the customer’s posture against the 30-, 60-, and 90-day horizons that regulator and legal review windows typically span, with the test result entering Playbook 13 (Six Metrics) Metric 3 (Time-to-Evidence) and Playbook 24 (Board-Ready Scorecard) Evidence-domain signals.

PB15 also explicitly closes a long-standing standards-gap pattern: NIST CSF 2.0 mandates that incident records and incident data be preserved with integrity and provenance (RS.AN-06, RS.AN-07) but does not specify the AI-specific retention lifecycle, the two-tier retention discipline, the legal-hold mechanism for AI evidence, the chain-of-custody discipline for AI evidence stores, or the empirical validation cadence. PB15 fills each of those gaps with a concrete operational specification that customers can adopt, regulators can audit against, and adopters can extend.

0.18.0 · 2026-06-29 · Playbook 22: Model and Policy Drift

Added

Changed

Why now

PB22 closes the change-event precondition that the framework’s response-side playbooks have been depending on without specifying. Every prior playbook assumes the AI system it addresses is operating in a steady state: the AI-BOM entries reflect the current configuration, the canary baselines hold, the detection rules are tuned against the current behavior envelope. None of those assumptions survive routine production AI operation, where models are upgraded by vendors on their own cadence, system prompts are edited weekly, policies and moderation layers are tuned in response to user feedback, retriever parameters shift as the index is rebuilt, and tool schemas change as downstream APIs evolve. The first time the security team finds out about a drift event is often when downstream business owners report that “the AI changed” without a corresponding security or operational event.

PB22 addresses this with the change-window forensics discipline (Post-Change Configuration Snapshot, change-pipeline event ledger, Drift Canary pack), the layered rollback sequence (tool policies → retriever parameters → system prompt → policy and moderation configuration → memory and context window → tool schemas → retrieval index and corpus version → model version pin, with canary replay between each layer), and the M3-Drift kill-switch variant that scopes containment to the specific recently-changed component while pre-change state is restored. The playbook makes the distinction between routine production tuning and drift event an empirical question (does the Drift Canary pack pass against the post-change state?) rather than a judgment call.

The playbook also explicitly identifies the misdiagnosis cost as the dominant operational risk: a drift incident investigated as an external attack burns response capacity, may trigger disclosure protocols inappropriately, and ultimately fails to identify the actual cause because the change-window evidence has expired. PB22’s two-snapshot pattern (Post-Change and Pre-Change Configuration Snapshots) and the change-pipeline event ledger make change-window evidence load-bearing rather than incidental.

PB22 completes the framework’s pre-production-testing / continuous-monitoring pair with PB14 (Testing for Agent Failure Modes). PB14 catches drift before deployment with the canary pack; PB22 catches drift after deployment with the layered rollback discipline. Together they convert the framework’s continuous-monitoring capability from a written commitment into a change-event-verified reality. After v0.18.0, the framework’s coverage of the precondition chain → response → measurement → continuous monitoring arc is complete: PB19 selects the platform; PB04 tiers the tools; PB07 disciplines the credentials; PB21 brings shadow agents into inventory; the response-side playbooks execute the incident response; PB13/PB14 measure and test; PB11 detects; PB22 closes the change-event feedback loop that keeps all of the above accurate as the AI system evolves.

0.17.0 · 2026-06-29 · Playbook 19: Build vs Buy for Agent Controls

Added

Changed

Why now

PB19 closes the procurement-time precondition that determines whether the framework’s response playbooks are executable at all. Every prior playbook assumes the underlying AI platform can do specific things: activate read-only mode in under 10 minutes, export prompt and tool-call logs within 60 minutes, correlate tool calls to SaaS audit records with traceable identifiers, configure data retention to match the customer’s regulatory window. If the platform cannot do these things, the response playbooks cannot do these things either. The procurement decision is therefore not a vendor-selection question; it is an incident-readiness question.

PB19 addresses this with the Proof of Readiness Test (the operational substitute for vendor demonstration), the eight critical procurement questions (the framework’s distillation of platform dependencies into a procurement checklist), the Build vs Buy Decision Matrix (which capabilities are typically appropriate to buy and which to build), and the post-procurement hardening discipline that converts the readiness test’s findings into either contractual commitments, customer-side build commitments, or documented risk-acceptance with PB24 C4 scorecard tracking.

The playbook completes the framework’s pre-incident discipline pair with PB04 (Tool Design Is Containment). PB19 selects the platform; PB04 tiers the tools the platform exposes. Together they convert the framework’s response capability from a written commitment into a deployment-time-verified reality. After v0.17.0, the framework’s coverage of the procurement → tool design → response pre-incident chain is complete: a CISO can use PB19 to decide whether to buy a platform, PB04 to tier the platform’s tools, and the response-side playbooks (PB01, PB03, PB06, PB07, PB08, PB09, PB10, PB11, PB12, PB18) to execute the incident response the platform was selected to support.

PB19 also closes a long-standing strategic positioning gap: prior framework releases positioned the response discipline but did not specify how to verify the platform supports it. The eight critical procurement questions and the Proof of Readiness Test give CISOs a defensible artifact to take into vendor evaluations and a measurable check against existing deployments. The artifact is the kind of thing standards-body engagement (NIST AI Safety Institute, OWASP GenAI Security Project, ISO/IEC JTC 1/SC 42) can reference as a concrete operational test.

0.16.0 · 2026-06-29 · Playbook 21: Shadow AI (From Shadow IT to Shadow Agents)

Added

Changed

Why now

PB21 closes the inventory-gap precondition the framework’s response-side playbooks have been depending on without specifying. Every prior playbook assumes the agent it addresses is in the AI-BOM, has a documented identity, and has tier-classified tools. None of those assumptions hold for shadow agents. In a typical 2026 enterprise, the AI agents the security team knows about are a fraction of the agents actually running; the rest sit in product teams, marketing operations, finance automation, customer success workflows, and individual employee tooling. The first time the security team finds out about a shadow agent is often during the incident the shadow agent caused.

PB21 addresses this with the discovery boundary, the 24-hour intake standard, and the migrate / redesign / retire decision path that converts discovery into governed inventory growth rather than discovery into churn. The playbook’s identity-level containment discipline addresses the operational case where the agent’s runtime is not customer-modifiable (vendor-hosted, personal-account-hosted, no-code platform without admin access), where traditional tool-level kill-switches do not apply.

The playbook explicitly takes a non-punitive posture toward shadow agent creators: shadow AI emerges from organizational momentum and innovation, not malice. The response framework’s job is to make the governed path faster than the shadow path, not to make the shadow path more painful. When the legitimate path takes a day and the shadow path takes a day, the discovery boundary becomes a governance boundary rather than an arms race. PB21 names the governed integration path as the fourth hardening boundary explicitly to prevent the discovery boundary from becoming a treadmill.

After v0.16.0, the framework’s coverage of the input → context → output → identity → inventory preconditions for every other playbook is complete: PB06 (input), PB03 (context), PB09 (output), PB07 (identity), PB21 (inventory). The remaining shipped playbooks (PB01 keystone, PB04 tool design, PB08 multi-agent, PB10 vendor copilots, PB11 monitoring, PB12 insider threat, PB13 metrics, PB14 testing, PB18 hardening, PB20 maturity, PB24 board scorecard) all operate on top of this foundational coverage.

0.15.0 · 2026-06-29 · Playbook 09: Leakage Without a Breach (AI Output Incidents)

Added

Changed

Why now

PB09 closes the largest remaining operational-arc gap in the framework’s response posture. Prior releases shipped playbooks for input-side incidents (PB06 workflow injection), context-side incidents (PB03 RAG forensics), identity-side incidents (PB07 secrets and tokens, PB12 insider threat), tool-side incidents (PB04 tool design), multi-agent cascade (PB08), vendor-managed incidents (PB10), and detection (PB11). The framework had no dedicated response playbook for the dominant 2026 data-incident class: output leakage, the confidentiality failures stemming from authorized AI outputs that reach destinations they should not have reached, without classic breach signals firing.

PB09’s defensive thesis (treat output channels as exfiltration paths; ship output-layer DLP in the tool wrapper, not the system prompt; classify destinations and tier outputs accordingly; map output distribution before any cleanup begins) operationalizes the response discipline that practitioners have been improvising on a per-incident basis. The newsletter issue this playbook ships from (Issue 9: “Leakage Without a Breach”) explicitly cites the dominant 2026 data-incident pattern: a support copilot pasting internal escalation notes into a customer ticket, a sales assistant CC’ing the wrong customer on a contract draft, a code copilot logging credentials in a public response. None of these incidents trigger traditional breach detection. All of them are real exposures.

The M3-Output containment variant introduced in this release is the architectural parallel to PB06’s M3-Workflow (content-channel containment) and PB10’s M3-Vendor (vendor-side containment). The framework’s M3 family now covers the three production-relevant containment surfaces above the canonical M3 Tool Tiering: M3-Workflow for content channels feeding the agent, M3-Output for output channels the agent writes to, and M3-Vendor for vendor-managed deployments. Together with M3-RAG (retrieval-layer containment from PB03) and M3-Delegation Cap (inter-agent delegation from PB08), the framework now has five M3 variants documented in kill-switches/overview.md, each anchored to a source playbook.

PB09 also addresses the OWASP Top 10 for LLM Applications 2025.1 LLM02 Sensitive Information Disclosure category that prior framework releases addressed only implicitly. After v0.15.0, the framework has substantive playbook coverage of both the OWASP Agentic Top 10 ASI01-ASI10 and the operationally-significant categories of the LLM Top 10 (LLM02 by PB09; LLM03/LLM04 supply-chain and poisoning by PB04 + PB06 + PB10).

0.14.3 · 2026-06-28 · Calibration Tuning + Cross-Doc Coherence

Added

Changed

Why now

The v0.14.2 ship surfaced a second-order pattern through six follow-on hostile-critic verification passes: every internal fix that touched cross-doc claims had a non-trivial chance of introducing a new inconsistency (the “ripple”). Fix 65 introduced 1 inconsistency (PB10 callout cited PB24 functionality that did not exist). Fix 66 introduced 2 (framework/01 still said “the minimum standard” after the README softened it; PB10 callout conflicted with PB10’s own defensive thesis). Fix 67 introduced 1 substantive (C4 row referenced “in-flight hardening” as overbroad scope). Fix 68 introduced 1 (PB13’s precision attempt at Governance-to-Metric-1 mapping was technically wrong). Fix 69 introduced 1 stylistic (PB24 Related parenthetical still said “four domains derive from” after PB13 clarified Governance is structurally different).

Each ripple was small. Cumulatively they represented a recurring failure mode that the RELEASE_CHECKLIST did not yet catch. v0.14.3 closes all six ripples and adds the two RELEASE_CHECKLIST discipline guards (Cross-reference existence check, Attribution correctness check) that should bound this failure mode going forward.

The release also addresses two carryover positioning observations from the holistic full-repo review: the “Insider Threat 3.0” framing in PB12 was presented as established taxonomy (now explicitly framed as the framework’s coined working label), and framework/01 did not preempt the “how does MVO-1 differ from NIST CSF ID.AM” question (now explicitly framed as an AI-agent-specific overlay layered on existing standards-body disciplines, not a replacement).

v0.14.3 is the calibration-arc completion release. After this release, the framework is in the most internally-consistent state internal work can produce; further improvements depend on outside-in signals (adoption case studies in Discussions, community PRs, standards-body engagement).

0.14.2 · 2026-06-28 · Release Hygiene + Operational Entry Points

Added

Changed

Why now

The v0.14.1 fresh hostile-critic review (post-release) surfaced three new defects introduced by the release itself: CITATION.cff went stale again, three CHANGELOG link references pointed to release tags that were never cut (v0.12.0, v0.13.0, v0.14.0), and the validator workflow had never run because its push trigger did not include its own path. The same review surfaced two carryover issues from earlier passes: the README’s absolutist “the minimum standard” opening, and the absence of an operational entry point (“framework-as-architecture-document, not framework-as-operational-artifact”). v0.14.2 closes the release-hygiene defects, opens the operational entry point, and softens the README opening to match the project’s actual scale.

The operational entry point (QUICKSTART.md + examples/incident-walkthrough.md) is the load-bearing addition of this release. Prior releases shipped templates, schemas, playbooks, and crosswalks but no integrated path from “I have an AI agent in production” to “I have a defensible maturity claim”. A CISO opening the repo on Monday morning had no answer to the question “what do I actually do first?” beyond reading 43 markdown files. QUICKSTART.md answers that question with a 30-day path that produces a Level 2 (Containable) or Level 3 (Provable) maturity claim per framework/03-maturity-roadmap.md. The synthetic worked example in examples/incident-walkthrough.md then demonstrates the framework’s controls operating as a coherent system across a fictional workflow-injection incident, anchoring the framework’s abstract claims (six triage questions, six mode variants, six evidence types, six metrics) in a concrete operational sequence.

The release-hygiene additions (RELEASE_CHECKLIST.md + workflow workflow_dispatch trigger + CITATION.cff bump) close the meta-defect that prior releases had a non-repeatable release process. The checklist makes the release cycle a documented routine rather than a sequence of remembered steps.

0.14.1 · 2026-06-28 · Calibration Pass + Reference Validator

Added

Changed

Why now

The v0.14.0 ship brought the framework to a coherent baseline (schemas + crosswalks + 14 playbooks). A hostile-critic review identified a calibration gap: the framework’s positioning (governance scaffolding for a multi-party project, “field-tested” claim, unverifiable reviewer attribution) exceeded the project’s actual scale (single-maintainer pre-1.0 synthesis with no documented production deployments). v0.14.1 closes that gap so the framework’s presentation matches its real state, and ships a working reference validator so the schemas can be cited as live CI artifacts rather than static specifications.

0.14.0 · 2026-06-28 · Schemas Directory: Machine-Readable Contracts

Added

Changed

Why now

The framework’s prior releases ship the playbooks (the narrative discipline) and the templates (the human-readable starting points). The schemas directory ships the machine-readable contracts that connect adopter CI pipelines to the framework’s normative rules. Before v0.14.0, an adopter editing their AI-BOM YAML or Privilege Matrix CSV had no automated way to verify the file still satisfied the framework’s CI rules. The risk-tier vocabulary could drift from T0/T1/T2 to “low/medium/high”. A T2 row could ship to production without approval_required=yes. A write tool could ship without declared write_targets. The schemas close that gap: the same CI rules that appear as English prose in the playbooks now appear as JSON Schema conditional constraints that any modern validator can enforce.

The two markdown specs (kill-switch-api.md and evidence-export.spec.md) close a parallel gap on the runtime side. Before v0.14.0, the Kill-Switch Mode M0 through M5 ladder was specified by behavior and TTA target, but not by API surface. Adopters building their own kill-switch automation had to derive the API from the playbooks. PB10’s vendor copilots showed why the API surface needed formal specification: vendor-provided granular containment must be testable against a common contract. The Kill-Switch API spec and the Evidence Export Script spec together formalize the runtime contracts that PB13 Metric 2 (Time-to-Safe-Mode) and PB13 Metric 3 (Time-to-Evidence) implicitly measure against.

The schemas are not a new framework chapter. They are the machine-readable encoding of contracts that already exist in the narrative framework. The shipping discipline is the same as the rest of the repository: the schemas trace back to the playbooks that specify their rules, the playbooks trace back to the newsletter issues, and the newsletter issues trace back to the operational practice. Adopters now have CI-ready artifacts they can drop into their existing validation pipelines without writing custom validators against framework prose.

0.13.0 · 2026-06-28 · Playbook 10: Vendor Copilots and Mutual Responsibility

Added

Changed

Why now

PB10 closes one of the two largest gaps the maintainer identified in the framework’s coverage of 2026 production deployment patterns. Vendor copilots (Microsoft 365 Copilot, Salesforce Einstein, ServiceNow Now Assist, Google Workspace Gemini, GitHub Copilot, and the embedded copilots increasingly built into CRM, ERP, and ticketing platforms) are a common 2026 deployment pattern for AI agents in regulated enterprises. The framework’s existing playbooks assumed customer-managed agents; PB10 ships the operational discipline for the case where the customer is responsible to the regulator but the vendor controls the agent’s operational levers.

PB10’s defensive thesis (deploy behind customer-controlled identity boundary, contract for testable evidence and containment SLAs, rehearse quarterly through the Vendor Evidence Drill) is the framework’s first formal supply-chain response playbook. It maps directly to OWASP Agentic Top 10 ASI04 (Agentic Supply Chain Compromise), NIST CSF 2.0 GV.SC (supply chain risk management), and the EU AI Act provider/deployer distinction. Customers acting as deployers of vendor-provided AI systems can now point at a specific operational playbook when responding to regulator inquiry or auditor question about vendor copilot incident readiness.

The Materiality and Disclosure integration (from v0.12.0) extends naturally to vendor-copilot incidents: vendor incidents nearly always cross the convening threshold because external recipients (the vendor and downstream customers) are touched. PB10 makes that convening call explicit in its First-Hour Actions section.

0.12.0 · 2026-06-28 · Playbook 06 + Materiality and Disclosure

Added

Changed

Why now

PB06 closes one of the largest gaps the maintainer identified in the framework’s previous releases. Indirect prompt injection in workflow form (harmful instructions hidden in tickets, emails, web pages, and ingested documents the agent reads as part of normal operation) is a distinct attack surface from the chat-UI form that earlier framework releases addressed in passing. PB06 ships the architectural-defense framing (untrusted content never directly triggers Tier-2 tools) as an alternative to the prompt-engineering posture common in published competing frameworks.

The Materiality and Disclosure block addresses a gap the maintainer’s review of SEC Item 1.05, EU AI Act Article 26/73, NY DFS Part 500, and HIPAA §164.408 surfaced: prior framework releases shipped containment and evidence machinery but did not specify the convening protocol that determines whether a regulatory disclosure clock has started. The four-file chain (framework/04 authority, PB01 trigger, PB18 verification, PB24 governance signal) operationalizes the discipline across the response arc.

The Scope declaration and the Measurement Scope qualifier are the framework’s first formal defensive language additions. They bound the framework’s claims to a specific regulatory role (deployer obligations) and a specific measurement context (drill-measured targets, not live-incident performance). Both are the load-bearing language a publicly-traded adopter needs in their 10-K to defensibly cite framework conformance.

The metadata and cross-document consistency work brings the framework to a coherent v0.12.0 baseline. Twelve releases of patches, ghost references, stale version anchors, and contradictory cross-doc citations are reconciled.

0.11.0 · 2026-06-28 · Playbook 12: Insider Threat 3.0 (AI-Driven Misuse)

Added

Changed

Why now

PB12 closes the last remaining forthcoming reference in the framework’s foundational chapters (the Mental Model’s Related section). After this release, a reader following the foundational arc (README → Mental Model → Maturity Roadmap → MVO) hits zero unfinished references. The framework’s foundational narrative reads as complete.

PB12 also completes the rogue-agent coverage arc. Playbook 11 covers detection of capability-family signals that suggest rogue behavior; PB12 covers the response and investigation when those signals fire. The two playbooks form a matched detection-response pair, the same upstream-downstream pattern PB07 → PB11 → PB08 established.

PB12’s “Insider Threat 3.0” framing positions the framework ahead of the analyst category formation. Insider Threat 1.0 (humans with credentials, DLP era) and Insider Threat 2.0 (humans with anomalous behavior, UEBA era) are addressed by mature programs. The 3.0 generation, where AI agents mediate or constitute the insider action, is the rising 2026 CISO concern. PB12 ships the operational playbook before the category solidifies.

0.10.0 · 2026-06-26 · Playbook 08: Multi-Agent Systems Multiply Blast Radius

Added

Changed

Why now

PB08 closes the last remaining OWASP Top 10 for Agentic Applications category. With this release, the framework has substantive playbook coverage of all 10 OWASP ASI categories AND all 6 NIST CSF 2.0 functions. Combined with the v0.9.0 CSF DETECT closure, v0.10.0 reaches the v1.0.0-ready standards posture.

PB08 also extends the upstream-downstream contract pattern established by PB07 → PB11. Multi-agent topologies inherit the credential-event log schema from PB07 (each agent is a separate identity, no permission inheritance) and consume PB11 detection rules for inter-agent traffic anomalies.

0.9.0 · 2026-06-25 · Playbook 11: Monitoring That Truly Detects Agent Incidents

Added

Changed

Why now

PB11 closes the last remaining NIST CSF 2.0 deferral. With this release, the framework has substantive playbook coverage of all six CSF 2.0 functions (GOVERN, IDENTIFY, PROTECT, DETECT, RESPOND, RECOVER). PB11 also consumes the credential-event log schema specified in Playbook 07, honoring the upstream-downstream contract PB07 established. The framework’s IR loop (identify, protect, detect, respond, recover, improve) is now operationally complete.

PB11 covers OWASP ASI06 (Memory & Context Poisoning detection), ASI08 (Cascading Agent Failures early detection), and ASI10 (Rogue Agent drift detection).

0.8.0 · 2026-06-24 · Playbook 07: Secrets and Tokens in an Agent World

Added

Why now

PB07 closes the largest implicit gap in the live framework. Every shipped playbook references token rotation (PB01 names it as “the single most common evidence-destruction failure in AI IR”), yet no playbook specified the agent-specific PAM cadence, rotation discipline, OAuth grant lifecycle, or break-glass procedure. PB07 fills the gap. It also closes the NIST CSF 2.0 PR.AA-05 standards gap that was explicitly documented as deferred in the CSF crosswalk Status section.

PB07 directly extends Playbook 01’s privileged-identity lens with the credential-management operational discipline that lens implies. PB11 (Monitoring, forthcoming) will consume the credential-event log this playbook specifies.

0.7.0 · 2026-06-24 · Playbook 20: AI IR Maturity Roadmap (Operating View)

Added

Why now

PB20 completes the measurement and discipline triad with Playbook 13 (Six Metrics) and Playbook 14 (Testing for Agent Failure Modes). Together these three carry the Level 4 (Resilient) maturity claim from aspiration to operating reality.

0.6.2 · 2026-06-23 · Structural Cleanup

Documentation accuracy and navigation polish. No framework substance changes.

Fixed

Changed

Added

0.6.1 · 2026-06-23 · Documentation Polish

Accuracy and OSS-convention round. No framework substance changes.

Fixed

Added

Changed

0.6.0 · 2026-06-23 · Measurement Release

Added

0.5.0 · 2026-06-20 · Playbook 24: Board-Ready Scorecard

Added

0.4.0 · 2026-06-19 · Playbook 18: Post-Incident Hardening

Added

0.3.0 · 2026-06-18 · Playbook 04: Tool Design Is Containment

Added

0.2.0 · 2026-06-18 · Playbook 01: The Agent Is a Privileged Identity

Added

0.1.5 · 2026-06-17 · Crosswalk Expansion

Added

Changed

0.1.4 · 2026-06-17 · Content Accuracy Polish

Changed

0.1.3 · 2026-06-17 · Cosmetic Polish

Changed

0.1.2 · 2026-06-17 · OSS Conventions

Added

0.1.1 · 2026-06-17 · License and Lint Fixes

Added

Fixed

0.1.0 · Foundation

The founding release. Establishes the thesis, the framework core, the triage discipline, the containment ladder, and the evidence taxonomy. No playbooks in this release. Playbook 01 ships next, in v0.2.0.

Added: Framework Core

Added: Triage

Added: Containment

Added: Evidence

Added: Templates