Implementation & Operations Runbook
Phased execution plan (P0–P9) with gates, evidence and rollback stances, plus the operational procedure library (RB-01…RB-23).
Companion to the canonical model in data/*.json and the roadmap view.
Honesty header: the source document supplies themes, not an organization. Every duration, threshold and owner in this runbook is a placeholder to be bound during Phase 0–1 (assumptions A-01, A-05, A-10, A-12). Nothing here claims compliance with any framework (A-11).
1 · Conventions, gates & evidence rules
Every phase ends at a gate: a go/no-go review against explicit exit criteria, chaired by the accountable governance body. A gate passes only on evidence — artifacts a skeptical reviewer can open (test reports, policy exports, drill timings, signed records), never on verbal assurance. Evidence is filed in the ITSM/document system referenced from the exit-criteria record (control C-GOV-03).
Three rules apply throughout:
| Rule | Meaning |
|---|---|
| Rollback before cutover | No production cutover proceeds until its rollback has been rehearsed or a documented point-of-no-return has been signed by the accountable owner. |
| Placeholder discipline | Values marked placeholder (RTO/RPO bands, thresholds, volumes) must be replaced with organization-confirmed numbers at the phase that produces them — carrying a placeholder past its resolution phase is a gate finding. |
| Exceptions expire | Any deviation from this runbook is an exception with an owner, a compensating control and an expiry date (RB-20, principle PR-12). |
2 · Operating teams & responsibility model
Team names are role placeholders (A-12); one person may hold several roles in smaller organizations, but the accountability split — especially platform vs. domain vs. security — should survive any sizing.
| Team | Owns (Accountable) | Key interfaces |
|---|---|---|
| Executive technology steering | Funding, wave go/no-go, enterprise risk acceptance | All governance bodies escalate here |
| Enterprise architecture | Canonical model, ADRs, standards, exceptions, this package | ARB secretariat; every phase gate |
| Business capability / domain product teams | Domain services, data products, their SLOs and pagers | Platform (consumes), data governance (products), ARB (designs) |
| Platform engineering | Golden paths, IDP, CI/CD, container platform SLOs | Every delivery team; cloud foundation |
| Cloud foundation | Landing zones, network, hybrid connectivity, guardrails | Security engineering, platform engineering |
| Security engineering & operations | Identity, controls catalog, SIEM/SOAR, incident response | All teams; security & risk review body |
| Data platform & data governance | Lakehouse, catalog, quality gates; policy via council | Domain data-product owners, AI platform |
| AI platform & AI governance | AI gateway, RAG, evals, agent controls; approvals via council | Security, data governance, domain teams |
| Integration platform | API/event/B2B platforms, schema registry, delivery semantics | Domain teams, partner operations |
| SRE / service operations | Observability, incident process, capacity, backup/DR | Every service owner; reliability review |
| FinOps practice | Allocation, budgets, anomaly response, optimization backlog | Steering, platform, every product owner |
| Risk, compliance, privacy, legal, procurement | Regulatory overlays, vendor terms, contracts, records policy | P0/P1 checkpoints; OQ-02/03/07/10 answers |
3 · Phases P0–P9
Phases map to roadmap waves: wave-0 = P0–P4, wave-1 = P5–P7 (first passes), wave-2 = P7–P8, wave-3 = P9 steady state. Phases overlap deliberately; gates, not calendars, sequence the work.
P0 · Mobilization & Governance
- Objective
- Stand up the decision fabric before any technical work: sponsorship, scope, decision rights, registers.
- Entry criteria
- Executive sponsor identified; this package reviewed as the baseline proposal.
- Confirm sponsorship, scope, objectives, constraints, funding envelope and success measures; record them against the six business objectives (obj-*).
- Charter the governance bodies (steering, ARB, security & risk, data, AI, platform, API, reliability, FinOps) with decision rights, quorum, risk-tiered review thresholds and review SLAs.
- Adopt the architecture repository: this package's model, ADR process, risk / assumption / dependency / exception registers; bind owner placeholders (A-12) to names.
- Ratify, amend or reject the 16 proposed ADRs and 12 principles — they are proposals until this step.
- Engage risk/compliance/privacy/procurement on the open questions that gate design choices (OQ-02, OQ-03, OQ-07, OQ-10).
- Set the RACI and communication cadence; publish the decision-rights matrix.
- Gate (exit)
- Charters signed · decision-rights matrix published · ADR dispositions recorded · registers live with owners · compliance-scope checkpoints scheduled.
- Evidence
- Charter documents; ADR log entries; register exports; kickoff minutes.
- Primary owners
- Executive steering (A), enterprise architecture (R).
P1 · Discovery & Current-State Assessment
- Objective
- Replace the illustrative baseline (A-01) with evidence; produce the BIA that turns placeholder tiers into commitments.
- Entry criteria
- P0 gate passed; inventory tooling and interview access arranged.
- Inventory business capabilities, value streams, products, applications, infrastructure, integrations, data assets, identities, controls, operational processes, vendors, contracts, licenses and run costs.
- Map dependencies and data flows; reconcile against this package's model — correct the model where reality differs (the model bends to evidence, never the reverse).
- Assess each application: business criticality, technical health, security posture, lifecycle, cloud readiness, operational maturity, cost; record a 7R/8R disposition per application (see ADR discussion of the source's inconsistent R-lists in traceability).
- Perform the business impact analysis: downtime and data-loss cost per business service; assign service tiers T1–T4 with named business sign-off; replace tier placeholders (A-05/A-07).
- Answer open questions OQ-01…OQ-12 where discovery can; convert answers into model updates and decision inputs (OD-01 especially).
- Baseline current run cost and delivery metrics (DORA baseline) so improvement is measurable (S48).
- Gate (exit)
- Inventory sign-off by domain owners · BIA v1 signed with tier assignments · disposition register complete · updated canonical model published (v1.x).
- Evidence
- Inventory dataset; BIA report; disposition register; model diff record.
- Primary owners
- Enterprise architecture (A), domain teams + SRE + FinOps (R).
P2 · Target Architecture & Roadmap Ratification
- Objective
- Turn this package's proposed target state into the organization's ratified target and a priced, sequenced roadmap.
- Entry criteria
- P1 gate passed (evidence-based current state exists).
- Confirm architecture principles and quality-attribute scenarios; write measurable NFRs from the BIA and volume data (replacing OQ-04 placeholders in capacity records).
- Revise the logical target architecture and transition states against discovery findings; re-run the model validator after every edit.
- Decide the open decisions that gate foundation build: OD-01 (region/DR posture) with priced options, OD-03 (workflow realization), OD-06 stance review.
- Evaluate modernization approach per application from the P1 dispositions (retain / retire / replace / repurchase / rehost / relocate / replatform / refactor).
- Prioritize waves by business value, risk retired, dependency order, readiness, cost and team capacity; set wave exit criteria (the roadmap view holds the proposal).
- Price wave 0–1 and secure funding through steering.
- Gate (exit)
- Target architecture ratified by ARB · OD-01/OD-03 decided · wave plan funded · NFR register versioned with real numbers where discovery produced them.
- Evidence
- ARB minutes; decision records; funded roadmap; NFR register diff.
- Primary owners
- ARB (A), enterprise architecture (R), steering (funding).
P3 · Foundation Build (landing zones, identity, guardrails)
- Objective
- Build the governed ground everything else lands on. No production workloads before the guardrail validation gate.
- Entry criteria
- P2 gate passed; cloud provider selection complete (ADR-02 bound to a vendor by the organization).
- Provision the landing-zone hierarchy (management / identity / connectivity / shared services / prod / non-prod / sandbox) via IaC only (RB-03).
- Implement org-policy guardrails as code: regions, mandatory tags, encryption defaults, public-exposure bans, logging-on (C-CLD-01); wire violation reporting.
- Stand up network foundation: hub with inspected egress, segmentation baselines, private endpoints, DNS/edge, hybrid connectivity with failover test (RB-04, C-NET-02/03).
- Harden identity: workforce IdP MFA/conditional access, CIAM tenant, PAM with JIT + break-glass, HRIS-driven JML via IGA (RB-01/RB-02, C-ID-01…05).
- Deploy central logging, SIEM onboarding for foundation sources, secrets/KMS services with rotation procedures (RB-06), backup platform baseline.
- Deploy FinOps baseline: billing export, allocation tagging enforcement, budgets and anomaly alerts (C-GOV-02).
- Run the guardrail validation suite: attempt-to-violate tests (public bucket, untagged deploy, unencrypted store, console change) must all be blocked and alerted.
- Gate (exit)
- Guardrail validation report all-green · MFA coverage at target · break-glass drill done · connectivity failover test done · tag enforcement proven.
- Evidence
- Validation suite output; policy pack export; drill records; coverage dashboards.
- Primary owners
- Cloud foundation (A/R), security engineering (R), FinOps (R).
P4 · Platform Engineering & DevSecOps
- Objective
- Make the compliant path the easy path: golden-path pipelines, GitOps delivery, supply-chain enforcement, developer portal.
- Entry criteria
- P3 gate passed.
- Establish source-control standards: protected branches, mandatory review, separation of duties on deploy approvals (C-SC-03).
- Build golden-path pipeline templates: build → test → SAST/SCA/secret/container/IaC scans → SBOM → sign → attest → staged deploy (C-SC-01/02).
- Stand up the trusted artifact registry; enforce provenance verification at cluster admission — test that unsigned images are rejected.
- Deploy container platform (multi-AZ, admission-controlled) and PaaS runtimes per ADR-04; GitOps controllers reconcile all environments (C-CLD-02).
- Launch the internal developer portal: service catalog with ownership/SLO/tier metadata, scaffolding, environment self-service (RB-05).
- Define environment, repository, promotion, release, exception and support models; publish platform SLOs and the platform roadmap (platform-as-product).
- Ship one reference service end-to-end through the golden path to production as the proving exercise.
- Gate (exit)
- Reference service in prod via golden path only · admission rejects unsigned/unattested images (tested) · platform SLOs published · DORA baseline captured.
- Evidence
- Pipeline runs; admission test logs; portal catalog export; SLO dashboard.
- Primary owners
- Platform engineering (A/R), security engineering (R).
P5 · Integration & API Enablement
- Objective
- Light up governed connectivity: API management, events/queues with delivery-semantics discipline, B2B/EDI/MFT, iPaaS.
- Entry criteria
- P4 gate passed; API/event standards drafted by gov-api.
- Deploy edge API gateway + APIM lifecycle + developer portal; onboard gateway logs to SIEM (C-OPS-03).
- Stand up schema/contract registry with backward-compatibility enforcement wired into CI (C-INT-01).
- Deploy event streaming backbone and work queues; publish the streams-vs-queues decision tree (ADR-07) and DLQ/replay conventions (C-INT-02/03).
- Deploy B2B gateway (EDI/AS2) and MFT with non-repudiation logging; write the partner onboarding runbook into service (RB-08).
- Deploy iPaaS for SaaS-sync flow classes with error stores and replay.
- Publish API/event standards: naming, versioning, deprecation policy, idempotency and reconciliation requirements; automate linting in CI.
- Build the legacy facade + ACL: read APIs and change events over the legacy core without modifying it (quick win; strangler enabler).
- First DLQ + replay drill executed by the first consumer team (RB-10) — a consumer may not go live before its drill.
- Gate (exit)
- First 5 API products live with contracts and plans · schema gate active in CI · DLQ/replay drill evidence · legacy facade serving production reads · first partner migrated via RB-08.
- Evidence
- Portal catalog; registry compatibility logs; drill records; facade traffic dashboards.
- Primary owners
- Integration platform (A/R), gov-api (standards), domain teams (consumers).
P6 · Data & AI Platform
- Objective
- Governed data estate (lakehouse, catalog, quality, MDM, first products) and the AI foundation (gateway, RAG, evals, cost controls) — with autonomy explicitly gated.
- Entry criteria
- P5 event backbone live; data governance council operating; classification policy ratified.
- Deploy lakehouse zones with promotion quality gates; CDC/batch ingestion for first sources via RB-09 (catalog-first: no source lands unclassified).
- Deploy catalog/lineage, data-quality monitoring and MDM foundations; wire classification tags to access and masking policy (C-DP-01/03).
- Publish the first two data products (Customer 360, Orders) with contracts, freshness SLOs and owner pagers (RB-11); begin legacy-DW parallel run with variance reporting.
- Deploy the AI gateway with model allowlists, virtual keys, quota/cost metering, redacted logging and guardrail filters (C-AI-01/02); scan for and eliminate direct provider egress.
- Stand up prompt registry, embedding pipeline (cataloged sources only), vector store with entitlement metadata, and the authorization-aware RAG service; run entitlement-bypass tests as a release gate (C-AI-03).
- Stand up the ML platform: registry, promotion gates, scheduled evals incl. hallucination/drift suites (RB-22); AI observability + token budgets feeding FinOps (C-AI-06).
- Ship the first internal RAG assistant (employee-facing) with HITL feedback loop.
- Hold the line: no autonomous production actions — agent capabilities remain in transition status until the AI governance council approves risk classes, permissions, rollback, evidence and operating ownership (ADR-14 gates; RB-12 exists first).
- Gate (exit)
- 2 certified data products with SLO dashboards · classification coverage 100% of onboarded sources · RAG bypass tests passed · zero direct provider egress (scan) · eval + cost dashboards reviewed by gov-ai.
- Evidence
- Certification records; scan reports; red-team/bypass results; gateway metering exports.
- Primary owners
- Data platform (A/R for estate), AI platform (A/R for AI), councils (approvals).
P7 · Application Modernization & Migration Waves
- Objective
- Execute strangler extractions and migrations with parallel-run evidence and rehearsed rollback — value and risk-retirement per wave, not big-bang.
- Entry criteria
- P5 facade live; P4 golden paths proven; per-wave cutover plans approved.
- Select pilot workloads by measurable business value and controlled risk (from the P1 disposition register); confirm each has an owner, SLOs and a rollback stance before build starts.
- Wave 1 extractions: customer profile and order management services behind the BFF; parallel-run against the monolith with reconciliation (counts + amounts) until variance is below the agreed threshold.
- Per-domain cutover: freeze window (RB-13) → traffic shift → hypercare → rollback window closes only on stability evidence; the monolith path stays warm through the window.
- Wave 2 extractions: fulfillment orchestration, billing handoff, notification/document services; drain remaining ESB flow classes to APIM/events/iPaaS with per-flow drain-back procedures.
- Rehost the irreducible VM remainder with forward dispositions recorded (transitional by policy, ADR-04).
- For each wave: data migration with checksum/row-count verification, integration change tests, security validation (pen-test scope per change), performance tests against P2 NFRs, DR test participation, operational-readiness review (P8 checklist).
- Gate (exit, per wave)
- Reconciliation variance below threshold · rollback rehearsed before each cutover · ESB flow count trending to zero (wave 2) · no orphaned integrations (validator clean against updated model).
- Evidence
- Parallel-run reports; cutover + rollback drill records; flow migration register; test reports.
- Primary owners
- Domain teams (A/R per service), integration platform (R), EA (model updates).
P8 · Operational Readiness & Cutover Discipline
- Objective
- Prove each service is operable before it carries production traffic, and prove recovery works with timed evidence.
- Entry criteria
- Service candidate feature-complete; monitoring/alerting wired.
- Operational-readiness review per service: dashboards, alerts with owners, SLOs, on-call rota, runbooks, capacity plan, backup/restore tested, security detections onboarded, support ownership and vendor escalation paths named.
- Go/no-go review against explicit acceptance criteria; record approvals, rollback point, communication plan, freeze window and hypercare model.
- Execute restore, failover and cyber-vault drills for tier-1 services (RB-17); capture timings against BIA objectives (or placeholder bands until then — flagged).
- Deliver the OD-01 decision package: priced multi-region posture options against BIA downtime costs.
- Close each cutover with a hypercare exit review and a lessons-learned record feeding the runbook library.
- Gate (exit)
- ORR checklist evidence per service · timed recovery drill records for T1 · OD-01 decided by steering · hypercare exits clean.
- Evidence
- ORR records; drill timing reports; go/no-go minutes; OD-01 decision record.
- Primary owners
- SRE (A/R), service owners (R), steering (OD-01).
P9 · Day-2 Operations & Continuous Modernization
- Objective
- Run the estate as a product portfolio with measured reliability, security, cost and delivery health — and keep modernizing on cadence instead of by crisis.
- Entry criteria
- Wave-2 exit; governance cadence operating.
- Operate the standing processes: incident, problem, change, release, configuration, patch, vulnerability, capacity, performance, backup, DR testing, access reviews, cost management, compliance-overlay checks, vendor management.
- Track the metric set: SLO attainment and error budgets, security posture (findings ageing, detection coverage), cost (unit economics, budget variance, anomaly MTTR), DORA delivery metrics, platform adoption and DevEx survey, data quality scores, AI eval scores and incident counts, business-outcome KPIs bound at P0.
- Run the governance cadence: quarterly — standards/technology-radar refresh, exception-register review (expiries enforced), technical-debt review, risk register re-scoring; monthly — FinOps review, reliability review; per-release — ADR updates.
- Complete legacy decommission with evidence: archive per retention, contract wind-down, decommission certificates for the monolith, ESB and legacy DW (the R-01 kill shot).
- Graduate AI autonomy per evidence: agent use cases move assist → recommend → bounded-auto only via AI-governance decisions with eval and audit evidence (never by default).
- Feed everything back into this package: the model, registers and runbooks are living artifacts — a stale architecture repository is itself a gate finding at the quarterly review.
- Gate (standing)
- Quarterly governance review passes: registers current, exceptions unexpired, model matches deployed reality (spot-audited).
- Evidence
- Review minutes; metric dashboards; decommission certificates; register exports.
- Primary owners
- SRE + platform + EA (R), governance bodies (A per domain).
4 · Procedure library (RB-01 … RB-23)
Every procedure follows one template: objective, scope, owner, approvers, prerequisites, inputs, tools, security considerations, steps, validation, evidence, failure conditions, rollback, escalation, acceptance criteria. Placeholders (thresholds, durations) are bound at P0–P2. The 21 mandatory runbooks from the brief are all present; RB-08 (partner onboarding) and the RB-13/14 split are additions for operability.
RB-01 · Identity onboarding & offboarding
- Objective
- Provision and deprovision workforce identity and access from HRIS lifecycle events with zero orphan accounts.
- Scope
- Employees/contractors; workforce IdP, IGA, downstream SCIM-connected apps. (Customer/partner identity: CIAM policies, not this runbook.)
- Owner
- Security engineering (IGA operator)
- Approvers
- Manager (access requests); application owner (app roles)
- Prerequisites
- HRIS→IGA feed live; role catalog defined; birthright-access matrix approved
- Inputs
- JML event (hire/move/termination) with effective date, role, department
- Tools
- IGA (sec-iga), IdP (sec-idp-workforce), ITSM record
- Security considerations
- Terminations execute at effective time, not business hours; movers lose old access on grant of new (no accumulation); privileged roles never birthright.
- IGA receives JML event; validates payload against schema.
- Joiner: create identity, assign birthright bundle, enroll phishing-resistant MFA before first credential issue.
- Mover: compute role delta; revoke-then-grant with manager confirmation on additions.
- Leaver: disable sessions + tokens immediately, disable account, schedule deletion per retention, transfer owned resources, revoke PAM entitlements.
- Propagate via SCIM; verify downstream completion states.
- Write completion record to ITSM with timings.
- Validation
- Post-run reconciliation: IdP accounts vs HRIS active roster; zero unmatched.
- Evidence
- JML latency report; reconciliation output; MFA-enrollment record
- Failure conditions
- SCIM target down → queue and alert; conflicting identities → manual merge procedure
- Rollback
- Joiner/mover grants reversible via IGA transaction log; leaver disable is not rolled back without HR instruction
- Escalation
- Security operations (orphan or failed-termination cases are Sev-2)
- Acceptance criteria
- Termination-to-disable within agreed minutes placeholder; zero orphan accounts in weekly scan
RB-02 · Privileged access & emergency (break-glass) access
- Objective
- Grant time-boxed privileged access with approval and recording; provide audited break-glass when identity systems fail.
- Scope
- Cloud/platform/data admin roles; on-prem legacy admin; break-glass accounts
- Owner
- Security engineering (PAM)
- Approvers
- Resource owner + security duty officer (break-glass: post-use review board)
- Prerequisites
- PAM live; standing-privilege scan clean; break-glass credentials sealed with dual control
- Inputs
- Elevation request: role, scope, duration, justification, change/incident reference
- Tools
- PAM (sec-pam), IdP conditional access, SIEM
- Security considerations
- All sessions recorded; elevation duration capped; break-glass use pages security operations in real time.
- Requester files elevation with justification and reference.
- Approver validates scope is least-privilege for the task; approves with expiry.
- PAM issues JIT credential/role; session recording starts.
- Work performed; PAM auto-revokes at expiry (early release encouraged).
- Break-glass: unseal under dual control → use → rotate credential immediately → mandatory post-use review within 48h.
- Validation
- Weekly: zero standing privileged assignments outside PAM; recordings retrievable.
- Evidence
- Elevation log with approvals; session recordings; break-glass review minutes
- Failure conditions
- PAM outage → break-glass path; approval bypass detected → security incident (RB-16)
- Rollback
- Revoke elevation; rotate any credential exposed during session
- Escalation
- CISO office for repeated break-glass or bypass findings
- Acceptance criteria
- 100% privileged sessions via PAM or reviewed break-glass; zero unreviewed break-glass uses
RB-03 · Cloud account / subscription provisioning
- Objective
- Create governed isolation units (accounts/subscriptions/projects) inside the landing-zone hierarchy — never hand-built.
- Scope
- All new isolation units, all environments
- Owner
- Cloud foundation
- Approvers
- Platform governance (standard patterns auto-approved; deviations to ARB)
- Prerequisites
- Landing-zone IaC modules versioned; tag taxonomy live; budget owner named
- Inputs
- Request: product, environment class, data classification ceiling, owner, cost center, network needs
- Tools
- IaC/GitOps (pe-iac), policy engine (pe-policy), FinOps tooling
- Security considerations
- Guardrails attach at creation (before any workload); logging + SIEM onboarding are part of the module, not a follow-up.
- Requester submits via IDP portal; request renders to an IaC pull request.
- Policy checks validate naming, tags, budget, classification ceiling.
- Approval per pattern tier; merge triggers provisioning.
- Module applies: baseline policies, log routing, SIEM source, budget alerts, network attachment, break-glass role wiring.
- Smoke validation suite runs (attempt-to-violate tests).
- Record unit in the model (architecture.json deployment scope) and CMDB.
- Validation
- Validation suite green; unit visible in FinOps allocation within 24h.
- Evidence
- PR + pipeline run; validation report; CMDB record
- Failure conditions
- Partial apply → destroy and re-apply from IaC (idempotent); manual edits detected → drift incident
- Rollback
- Destroy via IaC if no workloads; else exception process
- Escalation
- Cloud foundation lead; ARB for pattern deviations
- Acceptance criteria
- Unit provisioned entirely from code; all baseline controls verified active
RB-04 · Network & private connectivity provisioning
- Objective
- Attach workloads/zones to the hub network with deny-by-default segmentation and private endpoints; provision hybrid links.
- Scope
- Spoke networks, private endpoints, firewall rules, hybrid circuits
- Owner
- Cloud foundation (network)
- Approvers
- Security engineering for cross-zone rules; steering for new circuits (cost)
- Prerequisites
- Hub + inspection live; IP plan; flow-request schema
- Inputs
- Flow request: source, destination, port/protocol, classification, justification, expiry
- Tools
- IaC modules, firewall policy as code, DNS management
- Security considerations
- No any-any rules; every rule carries owner + expiry; data-zone egress is allowlist-only.
- Request via portal → IaC PR with rule metadata.
- Automated checks: overlap, shadowing, zone-crossing policy, classification compatibility.
- Security approval for crossings; merge applies via GitOps.
- Private endpoints preferred over public exposure; DNS records automated.
- For circuits: order, test failover (primary down → VPN carries), document capacity baseline.
- Quarterly rule recertification: expired/unused rules removed.
- Validation
- Connectivity test from both sides; negative test (undeclared flow blocked).
- Evidence
- Rule register with owners/expiry; failover test record
- Failure conditions
- Rule breaks existing flow → GitOps revert; circuit degradation → RB-21 vendor path
- Rollback
- Git revert of rule commit
- Escalation
- Network on-call; security ops for suspicious flow requests
- Acceptance criteria
- Flow works; nothing else changed (diff-verified); recertification current
RB-05 · Application onboarding to the platform
- Objective
- Bring a new or migrated service onto golden paths with ownership, SLOs, security and observability wired from day one.
- Scope
- Services deploying to container platform or PaaS
- Owner
- Owning domain team (platform engineering supports)
- Approvers
- Platform (namespace/quota); security (if handling restricted data)
- Prerequisites
- Golden-path templates live; team on-call rota exists
- Inputs
- Service metadata: owner, tier, classification, dependencies, SLO targets
- Tools
- IDP portal scaffolding, pipeline templates, catalog
- Security considerations
- Workload identity from first deploy; no static secrets; threat-model checklist for restricted-data services.
- Scaffold from golden path (service/consumer/data-product template) — pipeline, observability, security defaults pre-wired.
- Register in service catalog with ownership, tier, SLOs, runbook link.
- Request namespace/environment via portal (invokes RB-03/04 as needed).
- Wire alerts to the team's rota; add synthetic check if user-facing.
- Pass the operational-readiness checklist (P8) before production traffic.
- Update architecture model (node + relationships) — validator must pass.
- Validation
- ORR checklist; catalog completeness check (no unknown-owner services).
- Evidence
- Catalog entry; ORR record; first deployment pipeline run
- Failure conditions
- Template deviation → exception with expiry or remediation
- Rollback
- Service removal: traffic off, data disposition per retention, catalog + model cleanup
- Escalation
- Platform engineering lead
- Acceptance criteria
- Deployed via golden path; observable; owned; modeled
RB-06 · Secrets, keys & certificate rotation
- Objective
- Rotate credentials, encryption keys and certificates on schedule and on demand (compromise) without outages.
- Scope
- Vault secrets, KMS keys, internal PKI certificates, partner certs (with RB-08)
- Owner
- Security engineering; service teams execute their consumer side
- Approvers
- Key owner (per key register)
- Prerequisites
- Key register with owners + schedules; expiry monitoring alerts ≥30d ahead
- Inputs
- Rotation trigger: schedule, compromise event, or policy change
- Tools
- Secrets manager (sec-secrets), KMS/PKI (sec-kms), pipelines
- Security considerations
- Compromise rotations treat old material as hostile: revoke, don't just replace; dual control on root/HSM material.
- Identify consumers from the register (and vault access logs as cross-check).
- Issue new material alongside old (dual-validity window where protocol allows).
- Roll consumers via redeploy/reload; verify each cut over.
- Revoke old material; for certs, confirm CRL/OCSP propagation.
- Compromise path: execute immediately, page consumers, follow with RB-16.
- Update register with rotation record.
- Validation
- No consumer using old material (log scan); expiry dashboard green.
- Evidence
- Rotation log; revocation record; consumer verification list
- Failure conditions
- Consumer breaks on new material → dual-validity fallback while fixing; missed consumer → incident
- Rollback
- Scheduled rotations: extend dual-validity; compromise rotations: no rollback
- Escalation
- Security on-call; CISO for root/HSM events
- Acceptance criteria
- Zero expiry-caused outages; rotation SLAs met; register accurate
RB-07 · API publication & version retirement
- Objective
- Publish APIs as products with contracts and plans; retire versions without breaking consumers silently.
- Scope
- All APIs on the gateway/APIM (internal cross-domain + external)
- Owner
- Providing team
- Approvers
- API governance (standards conformance; breaking changes)
- Prerequisites
- Contract in registry; lint + compatibility checks in CI; portal listing drafted
- Inputs
- OpenAPI/AsyncAPI contract, product plan (quotas, auth), SLO statement
- Tools
- APIM (int-apim), schema registry (int-schema), developer portal
- Security considerations
- AuthN/Z policy per C-APP-01 attached before exposure; external products get abuse-rate alerting.
- Contract-first: publish to registry; CI compatibility gate green.
- Attach gateway policies (auth, rate limits, quotas) from standard policy sets.
- Publish product + docs + changelog to portal; subscription workflow live.
- New version: publish vN+1 alongside vN; announce deprecation schedule per policy.
- Retirement: monitor vN traffic → contact remaining consumers → brownouts (announced) → retire at zero/forced date with governance sign-off.
- Update model relationship contracts.
- Validation
- Contract tests green; portal docs render; auth negative-tests pass.
- Evidence
- Registry entries; deprecation notices; retirement sign-off
- Failure conditions
- Breaking change slips through → hotfix vN, incident review on the gate gap
- Rollback
- Version routing revert at gateway (vN kept warm through deprecation window)
- Escalation
- API governance chair
- Acceptance criteria
- No consumer breakage without notified schedule; zero rogue endpoints in scan
RB-08 · Trading-partner onboarding (B2B / EDI / MFT)
- Objective
- Onboard a partner to governed document exchange with agreed standards, security and acknowledgment tracking — repeatable, not a project.
- Scope
- EDI/AS2/SFTP/API partners on the B2B gateway and MFT
- Owner
- Partner integration team
- Approvers
- Partner operations (business terms); security (credential exchange)
- Prerequisites
- Trading-partner agreement signed; document standards + canonical mappings versioned
- Inputs
- Partner profile: documents, direction, volumes, windows, standards version, contacts
- Tools
- B2B gateway (int-b2b), MFT (int-mft), test harness
- Security considerations
- Certificate/key exchange out-of-band verified; per-partner mailboxes; least-privilege partner identities; non-repudiation logs retained per agreement.
- Create partner profile + identities; exchange and verify certificates.
- Configure document routes + canonical mappings; unit-test transforms with partner samples.
- Connectivity test (AS2 MDN / SFTP handshake) both directions.
- End-to-end test: partner sends test set → canonical events verified → acks (997/CONTRL) returned; reverse direction likewise.
- Parallel run window against any legacy exchange path; reconcile counts.
- Go live; monitoring on windows + ack timeouts; hypercare for first cycles.
- Validation
- Ack tracking green over first N cycles placeholder; reconciliation zero-variance.
- Evidence
- Onboarding checklist; test transcripts; reconciliation report
- Failure conditions
- Mapping defects → error store + replay after fix; partner connectivity flaps → RB-21 posture
- Rollback
- Route back to legacy exchange path during parallel window
- Escalation
- Partner operations → partner's technical contact chain
- Acceptance criteria
- Lead time within target placeholder; no manual re-keying anywhere in the flow
RB-09 · Data-source onboarding
- Objective
- Land a new source into the lakehouse raw zone catalog-first: classified, owned, quality-checked from the first byte.
- Scope
- Operational DBs (CDC), SaaS extracts, files via MFT, event topics
- Owner
- Data platform team + source-domain owner
- Approvers
- Data governance council delegate (classification); source system owner
- Prerequisites
- Catalog live; classification policy; ingestion patterns published
- Inputs
- Source profile: system, entities, classification, PII fields, refresh needs, retention
- Tools
- CDC/batch ingestion (data-cdc), catalog (data-catalog), quality monitors (data-quality)
- Security considerations
- Read-scoped credentials; PII tagged at entry driving masking (C-DP-03); residency placement per overlay (C-DP-05).
- Register source + datasets in catalog with owner and classification (gate: unclassified = no pipeline).
- Provision read-scoped identity; connectivity via private path.
- Configure connector (CDC position / extract schedule / file manifest); initial snapshot.
- Attach quality expectations (freshness, volume, schema); route alerts to source-domain owner.
- Verify lineage capture; document backfill procedure.
- Promote to validated zone only after first quality-green window.
- Validation
- Row-count/checksum parity vs source sample; lag within SLO placeholder.
- Evidence
- Catalog entry; parity report; quality dashboard
- Failure conditions
- Schema drift → connector alert + contract discussion with source owner; lag breach → capacity review (RB-18)
- Rollback
- Disable connector; raw-zone data retained or purged per decision
- Escalation
- Data platform on-call; council for classification disputes
- Acceptance criteria
- Source flowing, classified, owned, quality-monitored, lineage-visible
RB-10 · Event / topic onboarding (incl. DLQ & replay drill)
- Objective
- Add a producer or consumer to the backbone with schema governance and proven failure handling — the drill is the license to operate.
- Scope
- Streaming topics and work queues
- Owner
- Producing/consuming team
- Approvers
- API/integration governance (schema); integration platform (capacity)
- Prerequisites
- Schema in registry (backward-compatible); consumer group named; DLQ conventions read
- Inputs
- Topic/queue spec: keying, retention class, expected volume, ordering needs
- Tools
- Backbone (int-events)/queues (int-queue), registry (int-schema)
- Security considerations
- ACLs bound to workload identity; PII minimized in payloads and tagged in schema.
- Register schema; CI compatibility gate green.
- Provision topic/queue with ACLs, partitions/retention per spec; DLQ created alongside.
- Producers implement outbox; consumers implement id-dedup (C-INT-03).
- Lag/DLQ dashboards + alerts wired to the owning team.
- Drill (mandatory): inject poison message in staging → observe bounded retries → DLQ lands → alert fires → replay procedure executed → duplicate-safety verified.
- Production enablement after drill evidence filed.
- Validation
- Drill evidence; lag under threshold at expected volume test.
- Evidence
- Drill record; registry entry; dashboard links
- Failure conditions
- Schema break attempt → CI gate blocks; DLQ growth in prod → owner paged, replay per drill
- Rollback
- Consumers resume from committed offsets; topics are append-only (no destructive rollback)
- Escalation
- Integration platform on-call
- Acceptance criteria
- No consumer in production without a passed DLQ/replay drill
RB-11 · Data-product publication
- Objective
- Certify and publish a domain data product with contract, freshness SLO, quality gates and an owner pager.
- Scope
- Curated/serving-zone products (dp-*)
- Owner
- Domain data-product owner
- Approvers
- Data governance council (certification)
- Prerequisites
- Sources onboarded (RB-09); product contract drafted (schema + SLOs + semantics)
- Inputs
- Product spec: consumers, refresh, quality rules, access policy, PII handling
- Tools
- Lakehouse, catalog, quality platform, BI semantic layer
- Security considerations
- Access via classification-driven policy; restricted columns masked by default; consent filters where personal data (C-PRV-01).
- Publish contract to catalog (schema, freshness SLO, quality rules, owner, support channel).
- Implement transformations with tests; lineage verified end-to-end.
- Quality gates green for the certification window placeholder.
- Council certification review: contract completeness, policy conformance, duplication check (reuse-first).
- Publish to consumers; register BI datasets against the product (not raw tables).
- Pager wired: freshness/quality breaches page the domain owner.
- Validation
- Contract tests in CI; consumer smoke queries.
- Evidence
- Certification record; SLO dashboard; lineage graph
- Failure conditions
- Quality regression → product marked degraded in catalog (consumers see status), fix-forward
- Rollback
- Version-pinned consumers; previous product version retained per retention
- Escalation
- Data governance council
- Acceptance criteria
- Certified, contracted, monitored, owned — and discoverable in the catalog
RB-12 · Agent tool authorization & kill switch
- Objective
- Grant AI agents scoped tool access through governance, and disable agents/models/tools instantly when needed.
- Scope
- All agentic workloads (ai-agent-orch, ai-tools); gateway-level model routes
- Owner
- AI platform team
- Approvers
- AI governance council (grants + autonomy levels); security engineering (permission review)
- Prerequisites
- Tool registry live; sandbox execution verified; audit trail capturing every action
- Inputs
- Grant request: use case, tool(s), scopes, transaction limits, autonomy level, accountable human
- Tools
- Tool registry (ai-tools), AI gateway, HITL service, SIEM
- Security considerations
- Least-privilege per tool (never a broad service account); limits enforced in the sandbox, not the prompt; injection red-team results current (C-AI-02) before customer-facing grants.
- Requester files grant with use-case risk class and accountable owner.
- Security reviews effective permissions (diff against least privilege); council approves level (assist → recommend → bounded-auto).
- Register tool grant: scopes, transaction caps, HITL thresholds, expiry.
- Deploy; verify audit trail shows every tool call with inputs/outputs (redacted per policy).
- Kill switch: on trigger (incident, eval regression, cost runaway) — disable scope at registry/gateway (agent, tool, model, or all), confirm halt, notify owner, file incident.
- Quarterly grant recertification; expired grants auto-disable.
- Validation
- Kill-switch drill per scope quarterly; sandbox escape tests on registry changes.
- Evidence
- Grant register; drill records; audit-trail samples
- Failure conditions
- Unauthorized action detected → kill switch + RB-16 security incident + council review
- Rollback
- Grants are revocable instantly; agent-performed business actions roll back via their transactional/compensation paths (bounded by design)
- Escalation
- AI governance council chair; CISO for security-relevant events
- Acceptance criteria
- Every agent action attributable to a live, scoped, unexpired grant; kill switch proven per scope
RB-13 · CI/CD release & change freeze
- Objective
- Ship changes through golden-path gates with separation of duties; operate freeze windows without blocking emergencies.
- Scope
- All production deployments via pipelines/GitOps
- Owner
- Delivery team; platform engineering owns the paths
- Approvers
- Deploy approver ≠ author (C-SC-03); freeze exceptions per RB-14
- Prerequisites
- Pipeline gates green (tests, scans, SBOM, signing); rollback target identified
- Inputs
- Change record: scope, risk class, rollback plan, verification plan
- Tools
- CI/CD (pe-cicd), GitOps (pe-iac), registry (pe-artifacts)
- Security considerations
- Only signed, attested artifacts admit (C-SC-02); pipeline identities are OIDC-federated, no static cloud keys.
- Merge to release branch triggers pipeline; all gates must pass (no manual skip without RB-14).
- Staged rollout (canary/progressive where supported); health checks gate promotion.
- Verification plan executed; change record closed with links.
- Freeze windows (cutovers, peak periods): declared in advance with scope + dates; only RB-14 emergencies deploy inside.
- Validation
- Post-deploy SLO watch window; drift scan clean.
- Evidence
- Pipeline run + attestations; change record; approval trail
- Failure conditions
- Health gate fails → auto-halt; SLO burn post-deploy → rollback decision within defined window
- Rollback
- GitOps revert to last good state; data migrations require their own tested down-path or roll-forward plan declared in the change record
- Escalation
- Service owner → platform on-call
- Acceptance criteria
- Change deployed with full gate evidence; rollback path proven available
RB-14 · Emergency release
- Objective
- Ship urgent fixes (Sev-1/2, critical vulnerability) fast without abandoning control — speed with evidence, not cowboy deploys.
- Scope
- Production emergencies only; invoked from RB-16/RB-23/RB-15
- Owner
- Incident commander + delivery team
- Approvers
- Service owner + duty manager (2-person rule maintained)
- Prerequisites
- Active incident/vulnerability record justifying emergency path
- Inputs
- Fix scope, risk statement, verification + rollback plan
- Tools
- Same pipelines with emergency profile (reduced non-safety gates, security gates retained)
- Security considerations
- Signing/attestation NEVER skipped; scans run post-hoc if bypassed, with findings triaged within 24h.
- Declare emergency in the incident record; approvers ack.
- Run emergency pipeline profile (fast tests + sign + attest).
- Deploy with heightened watch; verify fix against incident symptoms.
- Backfill skipped checks within 24h; convert to normal release retro-record.
- Post-incident review includes: was the emergency path justified?
- Validation
- Symptom resolution; post-hoc gate results.
- Evidence
- Incident link; emergency approval trail; backfill results
- Failure conditions
- Fix worsens state → immediate GitOps revert (step 3 watch)
- Rollback
- Revert to last good; incident continues
- Escalation
- Major-incident chain (RB-23)
- Acceptance criteria
- Emergency path used only with incident linkage; zero unsigned artifacts even under pressure
RB-15 · Vulnerability remediation
- Objective
- Drive findings (CNAPP, scanners, pen-tests, disclosures) to closure within severity SLAs, with expiring exceptions only.
- Scope
- Code, dependencies, images, IaC, cloud posture, endpoints
- Owner
- Owning team per asset; security engineering runs the program
- Approvers
- Security & risk review for exceptions/risk acceptance
- Prerequisites
- Findings routed to owners automatically (asset → owner from catalog)
- Inputs
- Finding: severity, exploitability/reachability, affected assets
- Tools
- Posture platform (sec-posture), pipelines, ITSM
- Security considerations
- Internet-reachable + exploited-in-wild class triggers emergency path (RB-14) regardless of scheduled SLAs.
- Triage: validate, deduplicate, score with reachability context.
- Assign to owner with SLA clock per severity thresholds placeholder.
- Remediate via normal release (RB-13) or emergency (RB-14) per class.
- Verify closure by rescan, not assertion.
- Exceptions: compensating control + owner + expiry via security review; auto-reopen at expiry.
- Monthly program review: ageing, recurrence patterns, SLA attainment.
- Validation
- Rescan-verified closure; exception register hygiene.
- Evidence
- SLA reports; rescan results; review minutes
- Failure conditions
- SLA breach → escalation ladder; recurring same-class findings → root-cause item in engineering backlog
- Rollback
- n/a (remediation changes follow RB-13/14 rollback)
- Escalation
- Security & risk review → steering for systemic under-resourcing
- Acceptance criteria
- SLA attainment at target; zero unexpired exceptions past due
RB-16 · Security incident response
- Objective
- Detect, contain, eradicate and recover from security events with evidence preservation and honest communication.
- Scope
- Suspected/confirmed security events across the estate incl. AI components
- Owner
- Security operations (incident lead)
- Approvers
- CISO for destructive containment; legal/privacy for notification decisions
- Prerequisites
- SIEM detections mapped; SOAR playbooks tested; contact tree current
- Inputs
- Alert/report with initial indicators
- Tools
- SIEM (sec-siem), SOAR (sec-soar), EDR/XDR, PAM, forensics tooling
- Security considerations
- Evidence preservation before eradication; out-of-band comms channel if identity/collab compromise suspected.
- Triage + classify severity; open case with timeline log.
- Contain: SOAR playbooks (token revoke, host isolate, account disable) — destructive steps human-approved (C-OPS-04); kill switch for AI scopes (RB-12) if implicated.
- Preserve evidence (images, logs, chain of custody).
- Eradicate: rotate credentials (RB-06 compromise path), patch/close vector (RB-15/14).
- Recover: staged restoration with heightened monitoring; validate from clean sources (cyber vault if backups suspect, RB-17).
- Notify per legal/privacy determination — no premature certainty in comms.
- Post-incident review: timeline, root cause, control gaps → register updates and detection improvements.
- Validation
- Eradication verified by hunt; recovered systems clean-scanned.
- Evidence
- Case record + timeline; containment approvals; PIR with actions
- Failure conditions
- Containment fails → widen isolation scope, escalate to crisis management
- Rollback
- n/a — forward through recovery
- Escalation
- CISO → executive crisis team → external IR retainer placeholder
- Acceptance criteria
- MTTD/MTTR within targets placeholder; PIR actions tracked to closure
RB-17 · Backup, restore, failover & failback
- Objective
- Prove recovery, not backups: scheduled restore tests, tier-ordered failover rehearsals, clean failback — with timings recorded.
- Scope
- Databases, object stores, SaaS data (per OQ-06 scope), platform state; cyber-vault recovery path
- Owner
- SRE; data owners validate content
- Approvers
- Service owner (failover); steering (regional invocation per OD-01 posture)
- Prerequisites
- Backup policies active with immutability on T1 sets; recovery sequence documented per tier (dependency-ordered)
- Inputs
- Trigger: scheduled test, incident, or DR invocation
- Tools
- Backup platform (ops-backup), cyber vault (ops-cyber-vault), DR orchestration (ops-dr), IaC
- Security considerations
- Vault credentials separate from production IdP; restored systems rejoin only after clean-scan when incident-driven (RB-16).
- Restore test (scheduled per tier): select restore point → restore to isolated environment → integrity checks (row counts, checksums, app-level smoke) → record timing vs objective band.
- Failover rehearsal: announce window → execute tier sequence (stores → services → gateway → channels per path-recovery) → synthetic journeys verify → record RTO-observed.
- Failback: reconcile deltas accrued during failover (reconciliation jobs + event replay), verified cutback, post-check.
- Cyber-vault drill (annual+): assume primary backups hostile → recover T1 set from vault → measure and record.
- File all timings; gaps become actions in the reliability review.
- Validation
- Integrity checks green; synthetic journeys pass post-failover; delta reconciliation zero-variance.
- Evidence
- Timed drill reports; restore integrity outputs; reconciliation records
- Failure conditions
- Restore fails integrity → escalate as Sev-2, fix backup chain, re-test within window
- Rollback
- Failover rehearsals include the failback leg by design
- Escalation
- SRE lead → crisis management for real events
- Acceptance criteria
- Every T1/T2 store has a restore test within its cadence; observed timings vs objectives published honestly
RB-18 · Capacity scaling
- Objective
- Scale ahead of demand using telemetry, not incidents; keep autoscaling bounded and cost-aware.
- Scope
- Cluster/node pools, database tiers, broker partitions, gateway units, GPU pools (if ADR-13 trigger fired)
- Owner
- SRE + platform engineering; FinOps consulted on step changes
- Approvers
- Service owner (cost delta above threshold placeholder)
- Prerequisites
- Capacity dashboards per component; scaling triggers defined at onboarding
- Inputs
- Utilization/queue-depth trends, forecast events (campaigns, seasonality — OQ-04 data)
- Tools
- Observability, autoscalers, IaC
- Security considerations
- Scale changes via IaC (no console); quota guardrails prevent runaway autoscale spend.
- Review capacity dashboards on cadence; project against growth assumptions.
- For projected breach: plan step (scale up/out, partition, cache, or re-architecture flag).
- Apply via IaC; verify headroom restored; update capacity record.
- For load events: pre-scale plan + post-event scale-down (FinOps checks the down happened).
- Load-test significant changes against NFR profiles.
- Validation
- Headroom targets met; no capacity-caused SLO burn.
- Evidence
- Capacity review minutes; change records; load-test reports
- Failure conditions
- Emergency scaling during incident → allowed via RB-14 profile, reviewed after
- Rollback
- IaC revert to prior sizing (data-tier downscales checked for storage fit first)
- Escalation
- SRE lead; steering if demand shifts break the funding envelope
- Acceptance criteria
- Zero capacity Sev-1s; scale-downs executed after events (cost evidence)
RB-19 · Cost anomaly response
- Objective
- Detect and resolve spend anomalies fast — waste, misconfiguration, or compromise — with the same discipline as availability incidents.
- Scope
- Cloud spend, SaaS licensing spikes, AI token consumption
- Owner
- FinOps practice; resource owner executes
- Approvers
- n/a (response); steering for structural changes
- Prerequisites
- Budgets + anomaly detection live (C-GOV-02); allocation tags enforced; AI metering feeding (C-AI-06)
- Inputs
- Anomaly alert: scope, magnitude, trend
- Tools
- Cost tooling (fin-tooling), AI observability (ai-observe), provider console (read)
- Security considerations
- Sudden compute/egress anomalies are treated as possible compromise until excluded — loop in security ops early (crypto-mining pattern).
- Alert triage within response SLA placeholder: real vs billing artifact.
- Attribute via tags to owner; classify: waste / misconfig / legitimate growth / suspicious.
- Waste/misconfig: owner remediates (rightsize, delete idle, fix retention/egress path) via RB-13.
- Suspicious: hand to RB-16 immediately.
- Legitimate: re-baseline budget with owner sign-off.
- Record in anomaly log; recurring patterns become policy-as-code rules.
- Validation
- Spend returns to baseline or approved new baseline.
- Evidence
- Anomaly log with MTTR; remediation change records
- Failure conditions
- Attribution impossible (untagged) → tagging gap is itself a finding on the deploy path
- Rollback
- n/a
- Escalation
- FinOps → steering (budget), security (suspicious)
- Acceptance criteria
- Anomaly MTTR within target; recurrence class rate declining
RB-20 · Architecture exception & expiration
- Objective
- Allow justified deviation from standards without silent drift: every exception owned, compensated, expiring, reviewed.
- Scope
- Deviations from principles, standards, golden paths, control requirements
- Owner
- Requesting team owns the exception; EA runs the register
- Approvers
- Risk-tiered: ARB, plus security & risk review where controls affected
- Prerequisites
- Standard being excepted is identified precisely (not "the rules")
- Inputs
- Exception request: standard, reason, scope, compensating control, expiry, owner
- Tools
- Exception register (gov-ea-repo); policy engine exemptions where applicable
- Security considerations
- Policy-as-code exemptions are scoped and time-boxed in code — the register and the enforcement must never disagree.
- File request with compensating control and expiry (no expiry, no exception).
- Review at the appropriate tier; approve/reject/modify with rationale recorded.
- Implement scoped policy exemption where enforcement is automated.
- Register entry links standard, owner, expiry, compensation.
- At expiry: auto-flag → remediate, renew (re-review), or escalate; exemption code removed on closure.
- Quarterly hygiene review: ageing, clustering (clusters = standard may be wrong — feed standards lifecycle).
- Validation
- Register vs policy-exemption diff clean.
- Evidence
- Register export; review minutes; expiry action records
- Failure conditions
- Expired-but-active exception found → gate finding on the owning team
- Rollback
- Remove exemption; component must conform or stop deploying
- Escalation
- ARB chair → steering for risk-acceptance above ARB threshold
- Acceptance criteria
- Zero unexpired-past-due exceptions; cluster analysis feeding standards review
RB-21 · Vendor outage & third-party dependency response
- Objective
- Ride out SaaS/cloud/provider outages with pre-decided degradation stances instead of improvisation.
- Scope
- SaaS SoRs (ERP/CRM/HRIS/ITSM), payment provider, model providers, cloud service events
- Owner
- Service owner of the consuming capability; vendor manager engages the vendor
- Approvers
- Business owner for degraded-mode entry where customer-visible
- Prerequisites
- Per-dependency degradation matrix documented (what still works, what queues, what stops); vendor escalation contacts + SLAs on file
- Inputs
- Vendor status signal or synthetic/consumer failure detection (do not wait for the vendor's status page)
- Tools
- Observability, incident tooling, queues (absorption), AI gateway fallback routes
- Security considerations
- Degraded modes must not bypass authentication or approval controls "temporarily".
- Detect + confirm scope (their outage vs our path to them).
- Declare incident (RB-23); enter the pre-decided degradation stance: queue writes with replay (ERP-class), read-from-cache with staleness banners (CRM-class), provider failover at gateway (model-class), defer (analytics-class).
- Open vendor escalation per contract; track their incident id.
- Communicate honestly to affected users (status page templates).
- On restoration: drain queues with reconciliation (C-INT-05), verify data integrity, exit degraded mode.
- PIR includes: did the degradation matrix hold? Update it.
- Validation
- Reconciliation zero-variance after drain; degradation matrix confirmed accurate.
- Evidence
- Incident record; vendor comms log; reconciliation report
- Failure conditions
- Queue overflow before restoration → shed per matrix priority order (business-approved)
- Rollback
- n/a — exit through recovery steps
- Escalation
- Vendor manager → contract remedies; steering for chronic offenders (portfolio decision)
- Acceptance criteria
- No data loss (queued or reconciled); degraded-mode entry within target minutes placeholder
RB-22 · Model & prompt promotion
- Objective
- Promote models and prompt versions through eval gates with rollback — treating AI artifacts with release discipline.
- Scope
- Model versions (hosted route configs + self-managed deployments), prompt registry versions, embedding/config changes affecting RAG
- Owner
- AI platform team + owning product team
- Approvers
- AI governance council for risk-classed use cases; product owner otherwise
- Prerequisites
- Eval suites (golden sets incl. hallucination/injection tests) current; baseline scores recorded
- Inputs
- Candidate: model/prompt version, eval results vs baseline, cost delta projection
- Tools
- ML platform registry (ai-mlplat), prompt registry (ai-prompt), AI observability (ai-observe)
- Security considerations
- Injection red-team suite must pass for user-input-exposed prompts (C-AI-02); provider/data-terms re-check when the model provider changes.
- Register candidate with eval results: quality vs baseline, safety suite, latency, cost-per-call.
- Approval per risk class; record decision.
- Staged rollout via gateway routing (shadow → percentage → full), watching eval-proxy metrics and cost.
- Full promotion or halt; prior version retained hot for instant rollback.
- Post-promotion: scheduled evals continue; drift alerts route to owner (C-AI-06).
- Validation
- Live metrics within expected band vs offline evals.
- Evidence
- Registry promotion record; eval reports; rollout metrics
- Failure conditions
- Live regression → instant route rollback; eval-vs-live divergence → eval-suite improvement action
- Rollback
- Gateway route revert to prior version (kept hot through the watch window)
- Escalation
- AI platform lead; council for safety regressions
- Acceptance criteria
- No unevaluated artifact serving production; rollback exercised at least once in staging per quarter
RB-23 · Service incident & major incident response
- Objective
- Restore service fast with clear command, honest comms and learning that sticks.
- Scope
- Availability/performance incidents (security events: RB-16, invoked jointly when both)
- Owner
- On-call responder → incident commander at Sev-2+
- Approvers
- Comms: service owner; crisis escalation: duty executive
- Prerequisites
- Severity matrix published; on-call rotas staffed; paging tested (ops-incident)
- Inputs
- Alert (SLO burn, synthetic failure, user reports)
- Tools
- Observability (ops-observability), paging, ITSM, status comms
- Security considerations
- If compromise is plausible, engage RB-16 immediately — availability restoration must not destroy evidence.
- Acknowledge within paging SLA; classify severity per matrix.
- Sev-2+: name incident commander (command ≠ keyboard), open incident channel + timeline.
- Stabilize first (rollback RB-13, failover RB-17, degrade RB-21, scale RB-18) — root cause comes later.
- Communicate on cadence: internal stakeholders + user-facing status where visible.
- Verify restoration via synthetics + SLO recovery; close hypercare watch.
- Blameless post-incident review within days placeholder: timeline, contributing factors, action items with owners/dates; feed runbooks, alerts, model updates.
- Validation
- SLO recovery sustained; PIR actions tracked to closure.
- Evidence
- Incident record + timeline; comms log; PIR document
- Failure conditions
- Stabilization exhausted → crisis management + manual-workaround activation per continuity plan
- Rollback
- n/a — restoration is the goal
- Escalation
- Severity ladder to duty executive; vendor chain via RB-21
- Acceptance criteria
- MTTA/MTTR within targets placeholder; zero repeat incidents from unactioned PIR items