← Back to interactive architecture

Implementation & Operations Runbook

Phased execution plan (P0–P9) with gates, evidence and rollback stances, plus the operational procedure library (RB-01…RB-23). Companion to the canonical model in data/*.json and the roadmap view.

Honesty header: the source document supplies themes, not an organization. Every duration, threshold and owner in this runbook is a placeholder to be bound during Phase 0–1 (assumptions A-01, A-05, A-10, A-12). Nothing here claims compliance with any framework (A-11).

1 · Conventions, gates & evidence rules

Every phase ends at a gate: a go/no-go review against explicit exit criteria, chaired by the accountable governance body. A gate passes only on evidence — artifacts a skeptical reviewer can open (test reports, policy exports, drill timings, signed records), never on verbal assurance. Evidence is filed in the ITSM/document system referenced from the exit-criteria record (control C-GOV-03).

Three rules apply throughout:

RuleMeaning
Rollback before cutoverNo production cutover proceeds until its rollback has been rehearsed or a documented point-of-no-return has been signed by the accountable owner.
Placeholder disciplineValues marked placeholder (RTO/RPO bands, thresholds, volumes) must be replaced with organization-confirmed numbers at the phase that produces them — carrying a placeholder past its resolution phase is a gate finding.
Exceptions expireAny deviation from this runbook is an exception with an owner, a compensating control and an expiry date (RB-20, principle PR-12).

2 · Operating teams & responsibility model

Team names are role placeholders (A-12); one person may hold several roles in smaller organizations, but the accountability split — especially platform vs. domain vs. security — should survive any sizing.

TeamOwns (Accountable)Key interfaces
Executive technology steeringFunding, wave go/no-go, enterprise risk acceptanceAll governance bodies escalate here
Enterprise architectureCanonical model, ADRs, standards, exceptions, this packageARB secretariat; every phase gate
Business capability / domain product teamsDomain services, data products, their SLOs and pagersPlatform (consumes), data governance (products), ARB (designs)
Platform engineeringGolden paths, IDP, CI/CD, container platform SLOsEvery delivery team; cloud foundation
Cloud foundationLanding zones, network, hybrid connectivity, guardrailsSecurity engineering, platform engineering
Security engineering & operationsIdentity, controls catalog, SIEM/SOAR, incident responseAll teams; security & risk review body
Data platform & data governanceLakehouse, catalog, quality gates; policy via councilDomain data-product owners, AI platform
AI platform & AI governanceAI gateway, RAG, evals, agent controls; approvals via councilSecurity, data governance, domain teams
Integration platformAPI/event/B2B platforms, schema registry, delivery semanticsDomain teams, partner operations
SRE / service operationsObservability, incident process, capacity, backup/DREvery service owner; reliability review
FinOps practiceAllocation, budgets, anomaly response, optimization backlogSteering, platform, every product owner
Risk, compliance, privacy, legal, procurementRegulatory overlays, vendor terms, contracts, records policyP0/P1 checkpoints; OQ-02/03/07/10 answers

3 · Phases P0–P9

Phases map to roadmap waves: wave-0 = P0–P4, wave-1 = P5–P7 (first passes), wave-2 = P7–P8, wave-3 = P9 steady state. Phases overlap deliberately; gates, not calendars, sequence the work.

P0 · Mobilization & Governance

Objective
Stand up the decision fabric before any technical work: sponsorship, scope, decision rights, registers.
Entry criteria
Executive sponsor identified; this package reviewed as the baseline proposal.
  1. Confirm sponsorship, scope, objectives, constraints, funding envelope and success measures; record them against the six business objectives (obj-*).
  2. Charter the governance bodies (steering, ARB, security & risk, data, AI, platform, API, reliability, FinOps) with decision rights, quorum, risk-tiered review thresholds and review SLAs.
  3. Adopt the architecture repository: this package's model, ADR process, risk / assumption / dependency / exception registers; bind owner placeholders (A-12) to names.
  4. Ratify, amend or reject the 16 proposed ADRs and 12 principles — they are proposals until this step.
  5. Engage risk/compliance/privacy/procurement on the open questions that gate design choices (OQ-02, OQ-03, OQ-07, OQ-10).
  6. Set the RACI and communication cadence; publish the decision-rights matrix.
Gate (exit)
Charters signed · decision-rights matrix published · ADR dispositions recorded · registers live with owners · compliance-scope checkpoints scheduled.
Evidence
Charter documents; ADR log entries; register exports; kickoff minutes.
Primary owners
Executive steering (A), enterprise architecture (R).

P1 · Discovery & Current-State Assessment

Objective
Replace the illustrative baseline (A-01) with evidence; produce the BIA that turns placeholder tiers into commitments.
Entry criteria
P0 gate passed; inventory tooling and interview access arranged.
  1. Inventory business capabilities, value streams, products, applications, infrastructure, integrations, data assets, identities, controls, operational processes, vendors, contracts, licenses and run costs.
  2. Map dependencies and data flows; reconcile against this package's model — correct the model where reality differs (the model bends to evidence, never the reverse).
  3. Assess each application: business criticality, technical health, security posture, lifecycle, cloud readiness, operational maturity, cost; record a 7R/8R disposition per application (see ADR discussion of the source's inconsistent R-lists in traceability).
  4. Perform the business impact analysis: downtime and data-loss cost per business service; assign service tiers T1–T4 with named business sign-off; replace tier placeholders (A-05/A-07).
  5. Answer open questions OQ-01…OQ-12 where discovery can; convert answers into model updates and decision inputs (OD-01 especially).
  6. Baseline current run cost and delivery metrics (DORA baseline) so improvement is measurable (S48).
Gate (exit)
Inventory sign-off by domain owners · BIA v1 signed with tier assignments · disposition register complete · updated canonical model published (v1.x).
Evidence
Inventory dataset; BIA report; disposition register; model diff record.
Primary owners
Enterprise architecture (A), domain teams + SRE + FinOps (R).

P2 · Target Architecture & Roadmap Ratification

Objective
Turn this package's proposed target state into the organization's ratified target and a priced, sequenced roadmap.
Entry criteria
P1 gate passed (evidence-based current state exists).
  1. Confirm architecture principles and quality-attribute scenarios; write measurable NFRs from the BIA and volume data (replacing OQ-04 placeholders in capacity records).
  2. Revise the logical target architecture and transition states against discovery findings; re-run the model validator after every edit.
  3. Decide the open decisions that gate foundation build: OD-01 (region/DR posture) with priced options, OD-03 (workflow realization), OD-06 stance review.
  4. Evaluate modernization approach per application from the P1 dispositions (retain / retire / replace / repurchase / rehost / relocate / replatform / refactor).
  5. Prioritize waves by business value, risk retired, dependency order, readiness, cost and team capacity; set wave exit criteria (the roadmap view holds the proposal).
  6. Price wave 0–1 and secure funding through steering.
Gate (exit)
Target architecture ratified by ARB · OD-01/OD-03 decided · wave plan funded · NFR register versioned with real numbers where discovery produced them.
Evidence
ARB minutes; decision records; funded roadmap; NFR register diff.
Primary owners
ARB (A), enterprise architecture (R), steering (funding).

P3 · Foundation Build (landing zones, identity, guardrails)

Objective
Build the governed ground everything else lands on. No production workloads before the guardrail validation gate.
Entry criteria
P2 gate passed; cloud provider selection complete (ADR-02 bound to a vendor by the organization).
  1. Provision the landing-zone hierarchy (management / identity / connectivity / shared services / prod / non-prod / sandbox) via IaC only (RB-03).
  2. Implement org-policy guardrails as code: regions, mandatory tags, encryption defaults, public-exposure bans, logging-on (C-CLD-01); wire violation reporting.
  3. Stand up network foundation: hub with inspected egress, segmentation baselines, private endpoints, DNS/edge, hybrid connectivity with failover test (RB-04, C-NET-02/03).
  4. Harden identity: workforce IdP MFA/conditional access, CIAM tenant, PAM with JIT + break-glass, HRIS-driven JML via IGA (RB-01/RB-02, C-ID-01…05).
  5. Deploy central logging, SIEM onboarding for foundation sources, secrets/KMS services with rotation procedures (RB-06), backup platform baseline.
  6. Deploy FinOps baseline: billing export, allocation tagging enforcement, budgets and anomaly alerts (C-GOV-02).
  7. Run the guardrail validation suite: attempt-to-violate tests (public bucket, untagged deploy, unencrypted store, console change) must all be blocked and alerted.
Gate (exit)
Guardrail validation report all-green · MFA coverage at target · break-glass drill done · connectivity failover test done · tag enforcement proven.
Evidence
Validation suite output; policy pack export; drill records; coverage dashboards.
Primary owners
Cloud foundation (A/R), security engineering (R), FinOps (R).

P4 · Platform Engineering & DevSecOps

Objective
Make the compliant path the easy path: golden-path pipelines, GitOps delivery, supply-chain enforcement, developer portal.
Entry criteria
P3 gate passed.
  1. Establish source-control standards: protected branches, mandatory review, separation of duties on deploy approvals (C-SC-03).
  2. Build golden-path pipeline templates: build → test → SAST/SCA/secret/container/IaC scans → SBOM → sign → attest → staged deploy (C-SC-01/02).
  3. Stand up the trusted artifact registry; enforce provenance verification at cluster admission — test that unsigned images are rejected.
  4. Deploy container platform (multi-AZ, admission-controlled) and PaaS runtimes per ADR-04; GitOps controllers reconcile all environments (C-CLD-02).
  5. Launch the internal developer portal: service catalog with ownership/SLO/tier metadata, scaffolding, environment self-service (RB-05).
  6. Define environment, repository, promotion, release, exception and support models; publish platform SLOs and the platform roadmap (platform-as-product).
  7. Ship one reference service end-to-end through the golden path to production as the proving exercise.
Gate (exit)
Reference service in prod via golden path only · admission rejects unsigned/unattested images (tested) · platform SLOs published · DORA baseline captured.
Evidence
Pipeline runs; admission test logs; portal catalog export; SLO dashboard.
Primary owners
Platform engineering (A/R), security engineering (R).

P5 · Integration & API Enablement

Objective
Light up governed connectivity: API management, events/queues with delivery-semantics discipline, B2B/EDI/MFT, iPaaS.
Entry criteria
P4 gate passed; API/event standards drafted by gov-api.
  1. Deploy edge API gateway + APIM lifecycle + developer portal; onboard gateway logs to SIEM (C-OPS-03).
  2. Stand up schema/contract registry with backward-compatibility enforcement wired into CI (C-INT-01).
  3. Deploy event streaming backbone and work queues; publish the streams-vs-queues decision tree (ADR-07) and DLQ/replay conventions (C-INT-02/03).
  4. Deploy B2B gateway (EDI/AS2) and MFT with non-repudiation logging; write the partner onboarding runbook into service (RB-08).
  5. Deploy iPaaS for SaaS-sync flow classes with error stores and replay.
  6. Publish API/event standards: naming, versioning, deprecation policy, idempotency and reconciliation requirements; automate linting in CI.
  7. Build the legacy facade + ACL: read APIs and change events over the legacy core without modifying it (quick win; strangler enabler).
  8. First DLQ + replay drill executed by the first consumer team (RB-10) — a consumer may not go live before its drill.
Gate (exit)
First 5 API products live with contracts and plans · schema gate active in CI · DLQ/replay drill evidence · legacy facade serving production reads · first partner migrated via RB-08.
Evidence
Portal catalog; registry compatibility logs; drill records; facade traffic dashboards.
Primary owners
Integration platform (A/R), gov-api (standards), domain teams (consumers).

P6 · Data & AI Platform

Objective
Governed data estate (lakehouse, catalog, quality, MDM, first products) and the AI foundation (gateway, RAG, evals, cost controls) — with autonomy explicitly gated.
Entry criteria
P5 event backbone live; data governance council operating; classification policy ratified.
  1. Deploy lakehouse zones with promotion quality gates; CDC/batch ingestion for first sources via RB-09 (catalog-first: no source lands unclassified).
  2. Deploy catalog/lineage, data-quality monitoring and MDM foundations; wire classification tags to access and masking policy (C-DP-01/03).
  3. Publish the first two data products (Customer 360, Orders) with contracts, freshness SLOs and owner pagers (RB-11); begin legacy-DW parallel run with variance reporting.
  4. Deploy the AI gateway with model allowlists, virtual keys, quota/cost metering, redacted logging and guardrail filters (C-AI-01/02); scan for and eliminate direct provider egress.
  5. Stand up prompt registry, embedding pipeline (cataloged sources only), vector store with entitlement metadata, and the authorization-aware RAG service; run entitlement-bypass tests as a release gate (C-AI-03).
  6. Stand up the ML platform: registry, promotion gates, scheduled evals incl. hallucination/drift suites (RB-22); AI observability + token budgets feeding FinOps (C-AI-06).
  7. Ship the first internal RAG assistant (employee-facing) with HITL feedback loop.
  8. Hold the line: no autonomous production actions — agent capabilities remain in transition status until the AI governance council approves risk classes, permissions, rollback, evidence and operating ownership (ADR-14 gates; RB-12 exists first).
Gate (exit)
2 certified data products with SLO dashboards · classification coverage 100% of onboarded sources · RAG bypass tests passed · zero direct provider egress (scan) · eval + cost dashboards reviewed by gov-ai.
Evidence
Certification records; scan reports; red-team/bypass results; gateway metering exports.
Primary owners
Data platform (A/R for estate), AI platform (A/R for AI), councils (approvals).

P7 · Application Modernization & Migration Waves

Objective
Execute strangler extractions and migrations with parallel-run evidence and rehearsed rollback — value and risk-retirement per wave, not big-bang.
Entry criteria
P5 facade live; P4 golden paths proven; per-wave cutover plans approved.
  1. Select pilot workloads by measurable business value and controlled risk (from the P1 disposition register); confirm each has an owner, SLOs and a rollback stance before build starts.
  2. Wave 1 extractions: customer profile and order management services behind the BFF; parallel-run against the monolith with reconciliation (counts + amounts) until variance is below the agreed threshold.
  3. Per-domain cutover: freeze window (RB-13) → traffic shift → hypercare → rollback window closes only on stability evidence; the monolith path stays warm through the window.
  4. Wave 2 extractions: fulfillment orchestration, billing handoff, notification/document services; drain remaining ESB flow classes to APIM/events/iPaaS with per-flow drain-back procedures.
  5. Rehost the irreducible VM remainder with forward dispositions recorded (transitional by policy, ADR-04).
  6. For each wave: data migration with checksum/row-count verification, integration change tests, security validation (pen-test scope per change), performance tests against P2 NFRs, DR test participation, operational-readiness review (P8 checklist).
Gate (exit, per wave)
Reconciliation variance below threshold · rollback rehearsed before each cutover · ESB flow count trending to zero (wave 2) · no orphaned integrations (validator clean against updated model).
Evidence
Parallel-run reports; cutover + rollback drill records; flow migration register; test reports.
Primary owners
Domain teams (A/R per service), integration platform (R), EA (model updates).

P8 · Operational Readiness & Cutover Discipline

Objective
Prove each service is operable before it carries production traffic, and prove recovery works with timed evidence.
Entry criteria
Service candidate feature-complete; monitoring/alerting wired.
  1. Operational-readiness review per service: dashboards, alerts with owners, SLOs, on-call rota, runbooks, capacity plan, backup/restore tested, security detections onboarded, support ownership and vendor escalation paths named.
  2. Go/no-go review against explicit acceptance criteria; record approvals, rollback point, communication plan, freeze window and hypercare model.
  3. Execute restore, failover and cyber-vault drills for tier-1 services (RB-17); capture timings against BIA objectives (or placeholder bands until then — flagged).
  4. Deliver the OD-01 decision package: priced multi-region posture options against BIA downtime costs.
  5. Close each cutover with a hypercare exit review and a lessons-learned record feeding the runbook library.
Gate (exit)
ORR checklist evidence per service · timed recovery drill records for T1 · OD-01 decided by steering · hypercare exits clean.
Evidence
ORR records; drill timing reports; go/no-go minutes; OD-01 decision record.
Primary owners
SRE (A/R), service owners (R), steering (OD-01).

P9 · Day-2 Operations & Continuous Modernization

Objective
Run the estate as a product portfolio with measured reliability, security, cost and delivery health — and keep modernizing on cadence instead of by crisis.
Entry criteria
Wave-2 exit; governance cadence operating.
  1. Operate the standing processes: incident, problem, change, release, configuration, patch, vulnerability, capacity, performance, backup, DR testing, access reviews, cost management, compliance-overlay checks, vendor management.
  2. Track the metric set: SLO attainment and error budgets, security posture (findings ageing, detection coverage), cost (unit economics, budget variance, anomaly MTTR), DORA delivery metrics, platform adoption and DevEx survey, data quality scores, AI eval scores and incident counts, business-outcome KPIs bound at P0.
  3. Run the governance cadence: quarterly — standards/technology-radar refresh, exception-register review (expiries enforced), technical-debt review, risk register re-scoring; monthly — FinOps review, reliability review; per-release — ADR updates.
  4. Complete legacy decommission with evidence: archive per retention, contract wind-down, decommission certificates for the monolith, ESB and legacy DW (the R-01 kill shot).
  5. Graduate AI autonomy per evidence: agent use cases move assist → recommend → bounded-auto only via AI-governance decisions with eval and audit evidence (never by default).
  6. Feed everything back into this package: the model, registers and runbooks are living artifacts — a stale architecture repository is itself a gate finding at the quarterly review.
Gate (standing)
Quarterly governance review passes: registers current, exceptions unexpired, model matches deployed reality (spot-audited).
Evidence
Review minutes; metric dashboards; decommission certificates; register exports.
Primary owners
SRE + platform + EA (R), governance bodies (A per domain).

4 · Procedure library (RB-01 … RB-23)

Every procedure follows one template: objective, scope, owner, approvers, prerequisites, inputs, tools, security considerations, steps, validation, evidence, failure conditions, rollback, escalation, acceptance criteria. Placeholders (thresholds, durations) are bound at P0–P2. The 21 mandatory runbooks from the brief are all present; RB-08 (partner onboarding) and the RB-13/14 split are additions for operability.

RB-01 · Identity onboarding & offboarding

Objective
Provision and deprovision workforce identity and access from HRIS lifecycle events with zero orphan accounts.
Scope
Employees/contractors; workforce IdP, IGA, downstream SCIM-connected apps. (Customer/partner identity: CIAM policies, not this runbook.)
Owner
Security engineering (IGA operator)
Approvers
Manager (access requests); application owner (app roles)
Prerequisites
HRIS→IGA feed live; role catalog defined; birthright-access matrix approved
Inputs
JML event (hire/move/termination) with effective date, role, department
Tools
IGA (sec-iga), IdP (sec-idp-workforce), ITSM record
Security considerations
Terminations execute at effective time, not business hours; movers lose old access on grant of new (no accumulation); privileged roles never birthright.
Steps
  1. IGA receives JML event; validates payload against schema.
  2. Joiner: create identity, assign birthright bundle, enroll phishing-resistant MFA before first credential issue.
  3. Mover: compute role delta; revoke-then-grant with manager confirmation on additions.
  4. Leaver: disable sessions + tokens immediately, disable account, schedule deletion per retention, transfer owned resources, revoke PAM entitlements.
  5. Propagate via SCIM; verify downstream completion states.
  6. Write completion record to ITSM with timings.
Validation
Post-run reconciliation: IdP accounts vs HRIS active roster; zero unmatched.
Evidence
JML latency report; reconciliation output; MFA-enrollment record
Failure conditions
SCIM target down → queue and alert; conflicting identities → manual merge procedure
Rollback
Joiner/mover grants reversible via IGA transaction log; leaver disable is not rolled back without HR instruction
Escalation
Security operations (orphan or failed-termination cases are Sev-2)
Acceptance criteria
Termination-to-disable within agreed minutes placeholder; zero orphan accounts in weekly scan

RB-02 · Privileged access & emergency (break-glass) access

Objective
Grant time-boxed privileged access with approval and recording; provide audited break-glass when identity systems fail.
Scope
Cloud/platform/data admin roles; on-prem legacy admin; break-glass accounts
Owner
Security engineering (PAM)
Approvers
Resource owner + security duty officer (break-glass: post-use review board)
Prerequisites
PAM live; standing-privilege scan clean; break-glass credentials sealed with dual control
Inputs
Elevation request: role, scope, duration, justification, change/incident reference
Tools
PAM (sec-pam), IdP conditional access, SIEM
Security considerations
All sessions recorded; elevation duration capped; break-glass use pages security operations in real time.
Steps
  1. Requester files elevation with justification and reference.
  2. Approver validates scope is least-privilege for the task; approves with expiry.
  3. PAM issues JIT credential/role; session recording starts.
  4. Work performed; PAM auto-revokes at expiry (early release encouraged).
  5. Break-glass: unseal under dual control → use → rotate credential immediately → mandatory post-use review within 48h.
Validation
Weekly: zero standing privileged assignments outside PAM; recordings retrievable.
Evidence
Elevation log with approvals; session recordings; break-glass review minutes
Failure conditions
PAM outage → break-glass path; approval bypass detected → security incident (RB-16)
Rollback
Revoke elevation; rotate any credential exposed during session
Escalation
CISO office for repeated break-glass or bypass findings
Acceptance criteria
100% privileged sessions via PAM or reviewed break-glass; zero unreviewed break-glass uses

RB-03 · Cloud account / subscription provisioning

Objective
Create governed isolation units (accounts/subscriptions/projects) inside the landing-zone hierarchy — never hand-built.
Scope
All new isolation units, all environments
Owner
Cloud foundation
Approvers
Platform governance (standard patterns auto-approved; deviations to ARB)
Prerequisites
Landing-zone IaC modules versioned; tag taxonomy live; budget owner named
Inputs
Request: product, environment class, data classification ceiling, owner, cost center, network needs
Tools
IaC/GitOps (pe-iac), policy engine (pe-policy), FinOps tooling
Security considerations
Guardrails attach at creation (before any workload); logging + SIEM onboarding are part of the module, not a follow-up.
Steps
  1. Requester submits via IDP portal; request renders to an IaC pull request.
  2. Policy checks validate naming, tags, budget, classification ceiling.
  3. Approval per pattern tier; merge triggers provisioning.
  4. Module applies: baseline policies, log routing, SIEM source, budget alerts, network attachment, break-glass role wiring.
  5. Smoke validation suite runs (attempt-to-violate tests).
  6. Record unit in the model (architecture.json deployment scope) and CMDB.
Validation
Validation suite green; unit visible in FinOps allocation within 24h.
Evidence
PR + pipeline run; validation report; CMDB record
Failure conditions
Partial apply → destroy and re-apply from IaC (idempotent); manual edits detected → drift incident
Rollback
Destroy via IaC if no workloads; else exception process
Escalation
Cloud foundation lead; ARB for pattern deviations
Acceptance criteria
Unit provisioned entirely from code; all baseline controls verified active

RB-04 · Network & private connectivity provisioning

Objective
Attach workloads/zones to the hub network with deny-by-default segmentation and private endpoints; provision hybrid links.
Scope
Spoke networks, private endpoints, firewall rules, hybrid circuits
Owner
Cloud foundation (network)
Approvers
Security engineering for cross-zone rules; steering for new circuits (cost)
Prerequisites
Hub + inspection live; IP plan; flow-request schema
Inputs
Flow request: source, destination, port/protocol, classification, justification, expiry
Tools
IaC modules, firewall policy as code, DNS management
Security considerations
No any-any rules; every rule carries owner + expiry; data-zone egress is allowlist-only.
Steps
  1. Request via portal → IaC PR with rule metadata.
  2. Automated checks: overlap, shadowing, zone-crossing policy, classification compatibility.
  3. Security approval for crossings; merge applies via GitOps.
  4. Private endpoints preferred over public exposure; DNS records automated.
  5. For circuits: order, test failover (primary down → VPN carries), document capacity baseline.
  6. Quarterly rule recertification: expired/unused rules removed.
Validation
Connectivity test from both sides; negative test (undeclared flow blocked).
Evidence
Rule register with owners/expiry; failover test record
Failure conditions
Rule breaks existing flow → GitOps revert; circuit degradation → RB-21 vendor path
Rollback
Git revert of rule commit
Escalation
Network on-call; security ops for suspicious flow requests
Acceptance criteria
Flow works; nothing else changed (diff-verified); recertification current

RB-05 · Application onboarding to the platform

Objective
Bring a new or migrated service onto golden paths with ownership, SLOs, security and observability wired from day one.
Scope
Services deploying to container platform or PaaS
Owner
Owning domain team (platform engineering supports)
Approvers
Platform (namespace/quota); security (if handling restricted data)
Prerequisites
Golden-path templates live; team on-call rota exists
Inputs
Service metadata: owner, tier, classification, dependencies, SLO targets
Tools
IDP portal scaffolding, pipeline templates, catalog
Security considerations
Workload identity from first deploy; no static secrets; threat-model checklist for restricted-data services.
Steps
  1. Scaffold from golden path (service/consumer/data-product template) — pipeline, observability, security defaults pre-wired.
  2. Register in service catalog with ownership, tier, SLOs, runbook link.
  3. Request namespace/environment via portal (invokes RB-03/04 as needed).
  4. Wire alerts to the team's rota; add synthetic check if user-facing.
  5. Pass the operational-readiness checklist (P8) before production traffic.
  6. Update architecture model (node + relationships) — validator must pass.
Validation
ORR checklist; catalog completeness check (no unknown-owner services).
Evidence
Catalog entry; ORR record; first deployment pipeline run
Failure conditions
Template deviation → exception with expiry or remediation
Rollback
Service removal: traffic off, data disposition per retention, catalog + model cleanup
Escalation
Platform engineering lead
Acceptance criteria
Deployed via golden path; observable; owned; modeled

RB-06 · Secrets, keys & certificate rotation

Objective
Rotate credentials, encryption keys and certificates on schedule and on demand (compromise) without outages.
Scope
Vault secrets, KMS keys, internal PKI certificates, partner certs (with RB-08)
Owner
Security engineering; service teams execute their consumer side
Approvers
Key owner (per key register)
Prerequisites
Key register with owners + schedules; expiry monitoring alerts ≥30d ahead
Inputs
Rotation trigger: schedule, compromise event, or policy change
Tools
Secrets manager (sec-secrets), KMS/PKI (sec-kms), pipelines
Security considerations
Compromise rotations treat old material as hostile: revoke, don't just replace; dual control on root/HSM material.
Steps
  1. Identify consumers from the register (and vault access logs as cross-check).
  2. Issue new material alongside old (dual-validity window where protocol allows).
  3. Roll consumers via redeploy/reload; verify each cut over.
  4. Revoke old material; for certs, confirm CRL/OCSP propagation.
  5. Compromise path: execute immediately, page consumers, follow with RB-16.
  6. Update register with rotation record.
Validation
No consumer using old material (log scan); expiry dashboard green.
Evidence
Rotation log; revocation record; consumer verification list
Failure conditions
Consumer breaks on new material → dual-validity fallback while fixing; missed consumer → incident
Rollback
Scheduled rotations: extend dual-validity; compromise rotations: no rollback
Escalation
Security on-call; CISO for root/HSM events
Acceptance criteria
Zero expiry-caused outages; rotation SLAs met; register accurate

RB-07 · API publication & version retirement

Objective
Publish APIs as products with contracts and plans; retire versions without breaking consumers silently.
Scope
All APIs on the gateway/APIM (internal cross-domain + external)
Owner
Providing team
Approvers
API governance (standards conformance; breaking changes)
Prerequisites
Contract in registry; lint + compatibility checks in CI; portal listing drafted
Inputs
OpenAPI/AsyncAPI contract, product plan (quotas, auth), SLO statement
Tools
APIM (int-apim), schema registry (int-schema), developer portal
Security considerations
AuthN/Z policy per C-APP-01 attached before exposure; external products get abuse-rate alerting.
Steps
  1. Contract-first: publish to registry; CI compatibility gate green.
  2. Attach gateway policies (auth, rate limits, quotas) from standard policy sets.
  3. Publish product + docs + changelog to portal; subscription workflow live.
  4. New version: publish vN+1 alongside vN; announce deprecation schedule per policy.
  5. Retirement: monitor vN traffic → contact remaining consumers → brownouts (announced) → retire at zero/forced date with governance sign-off.
  6. Update model relationship contracts.
Validation
Contract tests green; portal docs render; auth negative-tests pass.
Evidence
Registry entries; deprecation notices; retirement sign-off
Failure conditions
Breaking change slips through → hotfix vN, incident review on the gate gap
Rollback
Version routing revert at gateway (vN kept warm through deprecation window)
Escalation
API governance chair
Acceptance criteria
No consumer breakage without notified schedule; zero rogue endpoints in scan

RB-08 · Trading-partner onboarding (B2B / EDI / MFT)

Objective
Onboard a partner to governed document exchange with agreed standards, security and acknowledgment tracking — repeatable, not a project.
Scope
EDI/AS2/SFTP/API partners on the B2B gateway and MFT
Owner
Partner integration team
Approvers
Partner operations (business terms); security (credential exchange)
Prerequisites
Trading-partner agreement signed; document standards + canonical mappings versioned
Inputs
Partner profile: documents, direction, volumes, windows, standards version, contacts
Tools
B2B gateway (int-b2b), MFT (int-mft), test harness
Security considerations
Certificate/key exchange out-of-band verified; per-partner mailboxes; least-privilege partner identities; non-repudiation logs retained per agreement.
Steps
  1. Create partner profile + identities; exchange and verify certificates.
  2. Configure document routes + canonical mappings; unit-test transforms with partner samples.
  3. Connectivity test (AS2 MDN / SFTP handshake) both directions.
  4. End-to-end test: partner sends test set → canonical events verified → acks (997/CONTRL) returned; reverse direction likewise.
  5. Parallel run window against any legacy exchange path; reconcile counts.
  6. Go live; monitoring on windows + ack timeouts; hypercare for first cycles.
Validation
Ack tracking green over first N cycles placeholder; reconciliation zero-variance.
Evidence
Onboarding checklist; test transcripts; reconciliation report
Failure conditions
Mapping defects → error store + replay after fix; partner connectivity flaps → RB-21 posture
Rollback
Route back to legacy exchange path during parallel window
Escalation
Partner operations → partner's technical contact chain
Acceptance criteria
Lead time within target placeholder; no manual re-keying anywhere in the flow

RB-09 · Data-source onboarding

Objective
Land a new source into the lakehouse raw zone catalog-first: classified, owned, quality-checked from the first byte.
Scope
Operational DBs (CDC), SaaS extracts, files via MFT, event topics
Owner
Data platform team + source-domain owner
Approvers
Data governance council delegate (classification); source system owner
Prerequisites
Catalog live; classification policy; ingestion patterns published
Inputs
Source profile: system, entities, classification, PII fields, refresh needs, retention
Tools
CDC/batch ingestion (data-cdc), catalog (data-catalog), quality monitors (data-quality)
Security considerations
Read-scoped credentials; PII tagged at entry driving masking (C-DP-03); residency placement per overlay (C-DP-05).
Steps
  1. Register source + datasets in catalog with owner and classification (gate: unclassified = no pipeline).
  2. Provision read-scoped identity; connectivity via private path.
  3. Configure connector (CDC position / extract schedule / file manifest); initial snapshot.
  4. Attach quality expectations (freshness, volume, schema); route alerts to source-domain owner.
  5. Verify lineage capture; document backfill procedure.
  6. Promote to validated zone only after first quality-green window.
Validation
Row-count/checksum parity vs source sample; lag within SLO placeholder.
Evidence
Catalog entry; parity report; quality dashboard
Failure conditions
Schema drift → connector alert + contract discussion with source owner; lag breach → capacity review (RB-18)
Rollback
Disable connector; raw-zone data retained or purged per decision
Escalation
Data platform on-call; council for classification disputes
Acceptance criteria
Source flowing, classified, owned, quality-monitored, lineage-visible

RB-10 · Event / topic onboarding (incl. DLQ & replay drill)

Objective
Add a producer or consumer to the backbone with schema governance and proven failure handling — the drill is the license to operate.
Scope
Streaming topics and work queues
Owner
Producing/consuming team
Approvers
API/integration governance (schema); integration platform (capacity)
Prerequisites
Schema in registry (backward-compatible); consumer group named; DLQ conventions read
Inputs
Topic/queue spec: keying, retention class, expected volume, ordering needs
Tools
Backbone (int-events)/queues (int-queue), registry (int-schema)
Security considerations
ACLs bound to workload identity; PII minimized in payloads and tagged in schema.
Steps
  1. Register schema; CI compatibility gate green.
  2. Provision topic/queue with ACLs, partitions/retention per spec; DLQ created alongside.
  3. Producers implement outbox; consumers implement id-dedup (C-INT-03).
  4. Lag/DLQ dashboards + alerts wired to the owning team.
  5. Drill (mandatory): inject poison message in staging → observe bounded retries → DLQ lands → alert fires → replay procedure executed → duplicate-safety verified.
  6. Production enablement after drill evidence filed.
Validation
Drill evidence; lag under threshold at expected volume test.
Evidence
Drill record; registry entry; dashboard links
Failure conditions
Schema break attempt → CI gate blocks; DLQ growth in prod → owner paged, replay per drill
Rollback
Consumers resume from committed offsets; topics are append-only (no destructive rollback)
Escalation
Integration platform on-call
Acceptance criteria
No consumer in production without a passed DLQ/replay drill

RB-11 · Data-product publication

Objective
Certify and publish a domain data product with contract, freshness SLO, quality gates and an owner pager.
Scope
Curated/serving-zone products (dp-*)
Owner
Domain data-product owner
Approvers
Data governance council (certification)
Prerequisites
Sources onboarded (RB-09); product contract drafted (schema + SLOs + semantics)
Inputs
Product spec: consumers, refresh, quality rules, access policy, PII handling
Tools
Lakehouse, catalog, quality platform, BI semantic layer
Security considerations
Access via classification-driven policy; restricted columns masked by default; consent filters where personal data (C-PRV-01).
Steps
  1. Publish contract to catalog (schema, freshness SLO, quality rules, owner, support channel).
  2. Implement transformations with tests; lineage verified end-to-end.
  3. Quality gates green for the certification window placeholder.
  4. Council certification review: contract completeness, policy conformance, duplication check (reuse-first).
  5. Publish to consumers; register BI datasets against the product (not raw tables).
  6. Pager wired: freshness/quality breaches page the domain owner.
Validation
Contract tests in CI; consumer smoke queries.
Evidence
Certification record; SLO dashboard; lineage graph
Failure conditions
Quality regression → product marked degraded in catalog (consumers see status), fix-forward
Rollback
Version-pinned consumers; previous product version retained per retention
Escalation
Data governance council
Acceptance criteria
Certified, contracted, monitored, owned — and discoverable in the catalog

RB-12 · Agent tool authorization & kill switch

Objective
Grant AI agents scoped tool access through governance, and disable agents/models/tools instantly when needed.
Scope
All agentic workloads (ai-agent-orch, ai-tools); gateway-level model routes
Owner
AI platform team
Approvers
AI governance council (grants + autonomy levels); security engineering (permission review)
Prerequisites
Tool registry live; sandbox execution verified; audit trail capturing every action
Inputs
Grant request: use case, tool(s), scopes, transaction limits, autonomy level, accountable human
Tools
Tool registry (ai-tools), AI gateway, HITL service, SIEM
Security considerations
Least-privilege per tool (never a broad service account); limits enforced in the sandbox, not the prompt; injection red-team results current (C-AI-02) before customer-facing grants.
Steps
  1. Requester files grant with use-case risk class and accountable owner.
  2. Security reviews effective permissions (diff against least privilege); council approves level (assist → recommend → bounded-auto).
  3. Register tool grant: scopes, transaction caps, HITL thresholds, expiry.
  4. Deploy; verify audit trail shows every tool call with inputs/outputs (redacted per policy).
  5. Kill switch: on trigger (incident, eval regression, cost runaway) — disable scope at registry/gateway (agent, tool, model, or all), confirm halt, notify owner, file incident.
  6. Quarterly grant recertification; expired grants auto-disable.
Validation
Kill-switch drill per scope quarterly; sandbox escape tests on registry changes.
Evidence
Grant register; drill records; audit-trail samples
Failure conditions
Unauthorized action detected → kill switch + RB-16 security incident + council review
Rollback
Grants are revocable instantly; agent-performed business actions roll back via their transactional/compensation paths (bounded by design)
Escalation
AI governance council chair; CISO for security-relevant events
Acceptance criteria
Every agent action attributable to a live, scoped, unexpired grant; kill switch proven per scope

RB-13 · CI/CD release & change freeze

Objective
Ship changes through golden-path gates with separation of duties; operate freeze windows without blocking emergencies.
Scope
All production deployments via pipelines/GitOps
Owner
Delivery team; platform engineering owns the paths
Approvers
Deploy approver ≠ author (C-SC-03); freeze exceptions per RB-14
Prerequisites
Pipeline gates green (tests, scans, SBOM, signing); rollback target identified
Inputs
Change record: scope, risk class, rollback plan, verification plan
Tools
CI/CD (pe-cicd), GitOps (pe-iac), registry (pe-artifacts)
Security considerations
Only signed, attested artifacts admit (C-SC-02); pipeline identities are OIDC-federated, no static cloud keys.
Steps
  1. Merge to release branch triggers pipeline; all gates must pass (no manual skip without RB-14).
  2. Staged rollout (canary/progressive where supported); health checks gate promotion.
  3. Verification plan executed; change record closed with links.
  4. Freeze windows (cutovers, peak periods): declared in advance with scope + dates; only RB-14 emergencies deploy inside.
Validation
Post-deploy SLO watch window; drift scan clean.
Evidence
Pipeline run + attestations; change record; approval trail
Failure conditions
Health gate fails → auto-halt; SLO burn post-deploy → rollback decision within defined window
Rollback
GitOps revert to last good state; data migrations require their own tested down-path or roll-forward plan declared in the change record
Escalation
Service owner → platform on-call
Acceptance criteria
Change deployed with full gate evidence; rollback path proven available

RB-14 · Emergency release

Objective
Ship urgent fixes (Sev-1/2, critical vulnerability) fast without abandoning control — speed with evidence, not cowboy deploys.
Scope
Production emergencies only; invoked from RB-16/RB-23/RB-15
Owner
Incident commander + delivery team
Approvers
Service owner + duty manager (2-person rule maintained)
Prerequisites
Active incident/vulnerability record justifying emergency path
Inputs
Fix scope, risk statement, verification + rollback plan
Tools
Same pipelines with emergency profile (reduced non-safety gates, security gates retained)
Security considerations
Signing/attestation NEVER skipped; scans run post-hoc if bypassed, with findings triaged within 24h.
Steps
  1. Declare emergency in the incident record; approvers ack.
  2. Run emergency pipeline profile (fast tests + sign + attest).
  3. Deploy with heightened watch; verify fix against incident symptoms.
  4. Backfill skipped checks within 24h; convert to normal release retro-record.
  5. Post-incident review includes: was the emergency path justified?
Validation
Symptom resolution; post-hoc gate results.
Evidence
Incident link; emergency approval trail; backfill results
Failure conditions
Fix worsens state → immediate GitOps revert (step 3 watch)
Rollback
Revert to last good; incident continues
Escalation
Major-incident chain (RB-23)
Acceptance criteria
Emergency path used only with incident linkage; zero unsigned artifacts even under pressure

RB-15 · Vulnerability remediation

Objective
Drive findings (CNAPP, scanners, pen-tests, disclosures) to closure within severity SLAs, with expiring exceptions only.
Scope
Code, dependencies, images, IaC, cloud posture, endpoints
Owner
Owning team per asset; security engineering runs the program
Approvers
Security & risk review for exceptions/risk acceptance
Prerequisites
Findings routed to owners automatically (asset → owner from catalog)
Inputs
Finding: severity, exploitability/reachability, affected assets
Tools
Posture platform (sec-posture), pipelines, ITSM
Security considerations
Internet-reachable + exploited-in-wild class triggers emergency path (RB-14) regardless of scheduled SLAs.
Steps
  1. Triage: validate, deduplicate, score with reachability context.
  2. Assign to owner with SLA clock per severity thresholds placeholder.
  3. Remediate via normal release (RB-13) or emergency (RB-14) per class.
  4. Verify closure by rescan, not assertion.
  5. Exceptions: compensating control + owner + expiry via security review; auto-reopen at expiry.
  6. Monthly program review: ageing, recurrence patterns, SLA attainment.
Validation
Rescan-verified closure; exception register hygiene.
Evidence
SLA reports; rescan results; review minutes
Failure conditions
SLA breach → escalation ladder; recurring same-class findings → root-cause item in engineering backlog
Rollback
n/a (remediation changes follow RB-13/14 rollback)
Escalation
Security & risk review → steering for systemic under-resourcing
Acceptance criteria
SLA attainment at target; zero unexpired exceptions past due

RB-16 · Security incident response

Objective
Detect, contain, eradicate and recover from security events with evidence preservation and honest communication.
Scope
Suspected/confirmed security events across the estate incl. AI components
Owner
Security operations (incident lead)
Approvers
CISO for destructive containment; legal/privacy for notification decisions
Prerequisites
SIEM detections mapped; SOAR playbooks tested; contact tree current
Inputs
Alert/report with initial indicators
Tools
SIEM (sec-siem), SOAR (sec-soar), EDR/XDR, PAM, forensics tooling
Security considerations
Evidence preservation before eradication; out-of-band comms channel if identity/collab compromise suspected.
Steps
  1. Triage + classify severity; open case with timeline log.
  2. Contain: SOAR playbooks (token revoke, host isolate, account disable) — destructive steps human-approved (C-OPS-04); kill switch for AI scopes (RB-12) if implicated.
  3. Preserve evidence (images, logs, chain of custody).
  4. Eradicate: rotate credentials (RB-06 compromise path), patch/close vector (RB-15/14).
  5. Recover: staged restoration with heightened monitoring; validate from clean sources (cyber vault if backups suspect, RB-17).
  6. Notify per legal/privacy determination — no premature certainty in comms.
  7. Post-incident review: timeline, root cause, control gaps → register updates and detection improvements.
Validation
Eradication verified by hunt; recovered systems clean-scanned.
Evidence
Case record + timeline; containment approvals; PIR with actions
Failure conditions
Containment fails → widen isolation scope, escalate to crisis management
Rollback
n/a — forward through recovery
Escalation
CISO → executive crisis team → external IR retainer placeholder
Acceptance criteria
MTTD/MTTR within targets placeholder; PIR actions tracked to closure

RB-17 · Backup, restore, failover & failback

Objective
Prove recovery, not backups: scheduled restore tests, tier-ordered failover rehearsals, clean failback — with timings recorded.
Scope
Databases, object stores, SaaS data (per OQ-06 scope), platform state; cyber-vault recovery path
Owner
SRE; data owners validate content
Approvers
Service owner (failover); steering (regional invocation per OD-01 posture)
Prerequisites
Backup policies active with immutability on T1 sets; recovery sequence documented per tier (dependency-ordered)
Inputs
Trigger: scheduled test, incident, or DR invocation
Tools
Backup platform (ops-backup), cyber vault (ops-cyber-vault), DR orchestration (ops-dr), IaC
Security considerations
Vault credentials separate from production IdP; restored systems rejoin only after clean-scan when incident-driven (RB-16).
Steps
  1. Restore test (scheduled per tier): select restore point → restore to isolated environment → integrity checks (row counts, checksums, app-level smoke) → record timing vs objective band.
  2. Failover rehearsal: announce window → execute tier sequence (stores → services → gateway → channels per path-recovery) → synthetic journeys verify → record RTO-observed.
  3. Failback: reconcile deltas accrued during failover (reconciliation jobs + event replay), verified cutback, post-check.
  4. Cyber-vault drill (annual+): assume primary backups hostile → recover T1 set from vault → measure and record.
  5. File all timings; gaps become actions in the reliability review.
Validation
Integrity checks green; synthetic journeys pass post-failover; delta reconciliation zero-variance.
Evidence
Timed drill reports; restore integrity outputs; reconciliation records
Failure conditions
Restore fails integrity → escalate as Sev-2, fix backup chain, re-test within window
Rollback
Failover rehearsals include the failback leg by design
Escalation
SRE lead → crisis management for real events
Acceptance criteria
Every T1/T2 store has a restore test within its cadence; observed timings vs objectives published honestly

RB-18 · Capacity scaling

Objective
Scale ahead of demand using telemetry, not incidents; keep autoscaling bounded and cost-aware.
Scope
Cluster/node pools, database tiers, broker partitions, gateway units, GPU pools (if ADR-13 trigger fired)
Owner
SRE + platform engineering; FinOps consulted on step changes
Approvers
Service owner (cost delta above threshold placeholder)
Prerequisites
Capacity dashboards per component; scaling triggers defined at onboarding
Inputs
Utilization/queue-depth trends, forecast events (campaigns, seasonality — OQ-04 data)
Tools
Observability, autoscalers, IaC
Security considerations
Scale changes via IaC (no console); quota guardrails prevent runaway autoscale spend.
Steps
  1. Review capacity dashboards on cadence; project against growth assumptions.
  2. For projected breach: plan step (scale up/out, partition, cache, or re-architecture flag).
  3. Apply via IaC; verify headroom restored; update capacity record.
  4. For load events: pre-scale plan + post-event scale-down (FinOps checks the down happened).
  5. Load-test significant changes against NFR profiles.
Validation
Headroom targets met; no capacity-caused SLO burn.
Evidence
Capacity review minutes; change records; load-test reports
Failure conditions
Emergency scaling during incident → allowed via RB-14 profile, reviewed after
Rollback
IaC revert to prior sizing (data-tier downscales checked for storage fit first)
Escalation
SRE lead; steering if demand shifts break the funding envelope
Acceptance criteria
Zero capacity Sev-1s; scale-downs executed after events (cost evidence)

RB-19 · Cost anomaly response

Objective
Detect and resolve spend anomalies fast — waste, misconfiguration, or compromise — with the same discipline as availability incidents.
Scope
Cloud spend, SaaS licensing spikes, AI token consumption
Owner
FinOps practice; resource owner executes
Approvers
n/a (response); steering for structural changes
Prerequisites
Budgets + anomaly detection live (C-GOV-02); allocation tags enforced; AI metering feeding (C-AI-06)
Inputs
Anomaly alert: scope, magnitude, trend
Tools
Cost tooling (fin-tooling), AI observability (ai-observe), provider console (read)
Security considerations
Sudden compute/egress anomalies are treated as possible compromise until excluded — loop in security ops early (crypto-mining pattern).
Steps
  1. Alert triage within response SLA placeholder: real vs billing artifact.
  2. Attribute via tags to owner; classify: waste / misconfig / legitimate growth / suspicious.
  3. Waste/misconfig: owner remediates (rightsize, delete idle, fix retention/egress path) via RB-13.
  4. Suspicious: hand to RB-16 immediately.
  5. Legitimate: re-baseline budget with owner sign-off.
  6. Record in anomaly log; recurring patterns become policy-as-code rules.
Validation
Spend returns to baseline or approved new baseline.
Evidence
Anomaly log with MTTR; remediation change records
Failure conditions
Attribution impossible (untagged) → tagging gap is itself a finding on the deploy path
Rollback
n/a
Escalation
FinOps → steering (budget), security (suspicious)
Acceptance criteria
Anomaly MTTR within target; recurrence class rate declining

RB-20 · Architecture exception & expiration

Objective
Allow justified deviation from standards without silent drift: every exception owned, compensated, expiring, reviewed.
Scope
Deviations from principles, standards, golden paths, control requirements
Owner
Requesting team owns the exception; EA runs the register
Approvers
Risk-tiered: ARB, plus security & risk review where controls affected
Prerequisites
Standard being excepted is identified precisely (not "the rules")
Inputs
Exception request: standard, reason, scope, compensating control, expiry, owner
Tools
Exception register (gov-ea-repo); policy engine exemptions where applicable
Security considerations
Policy-as-code exemptions are scoped and time-boxed in code — the register and the enforcement must never disagree.
Steps
  1. File request with compensating control and expiry (no expiry, no exception).
  2. Review at the appropriate tier; approve/reject/modify with rationale recorded.
  3. Implement scoped policy exemption where enforcement is automated.
  4. Register entry links standard, owner, expiry, compensation.
  5. At expiry: auto-flag → remediate, renew (re-review), or escalate; exemption code removed on closure.
  6. Quarterly hygiene review: ageing, clustering (clusters = standard may be wrong — feed standards lifecycle).
Validation
Register vs policy-exemption diff clean.
Evidence
Register export; review minutes; expiry action records
Failure conditions
Expired-but-active exception found → gate finding on the owning team
Rollback
Remove exemption; component must conform or stop deploying
Escalation
ARB chair → steering for risk-acceptance above ARB threshold
Acceptance criteria
Zero unexpired-past-due exceptions; cluster analysis feeding standards review

RB-21 · Vendor outage & third-party dependency response

Objective
Ride out SaaS/cloud/provider outages with pre-decided degradation stances instead of improvisation.
Scope
SaaS SoRs (ERP/CRM/HRIS/ITSM), payment provider, model providers, cloud service events
Owner
Service owner of the consuming capability; vendor manager engages the vendor
Approvers
Business owner for degraded-mode entry where customer-visible
Prerequisites
Per-dependency degradation matrix documented (what still works, what queues, what stops); vendor escalation contacts + SLAs on file
Inputs
Vendor status signal or synthetic/consumer failure detection (do not wait for the vendor's status page)
Tools
Observability, incident tooling, queues (absorption), AI gateway fallback routes
Security considerations
Degraded modes must not bypass authentication or approval controls "temporarily".
Steps
  1. Detect + confirm scope (their outage vs our path to them).
  2. Declare incident (RB-23); enter the pre-decided degradation stance: queue writes with replay (ERP-class), read-from-cache with staleness banners (CRM-class), provider failover at gateway (model-class), defer (analytics-class).
  3. Open vendor escalation per contract; track their incident id.
  4. Communicate honestly to affected users (status page templates).
  5. On restoration: drain queues with reconciliation (C-INT-05), verify data integrity, exit degraded mode.
  6. PIR includes: did the degradation matrix hold? Update it.
Validation
Reconciliation zero-variance after drain; degradation matrix confirmed accurate.
Evidence
Incident record; vendor comms log; reconciliation report
Failure conditions
Queue overflow before restoration → shed per matrix priority order (business-approved)
Rollback
n/a — exit through recovery steps
Escalation
Vendor manager → contract remedies; steering for chronic offenders (portfolio decision)
Acceptance criteria
No data loss (queued or reconciled); degraded-mode entry within target minutes placeholder

RB-22 · Model & prompt promotion

Objective
Promote models and prompt versions through eval gates with rollback — treating AI artifacts with release discipline.
Scope
Model versions (hosted route configs + self-managed deployments), prompt registry versions, embedding/config changes affecting RAG
Owner
AI platform team + owning product team
Approvers
AI governance council for risk-classed use cases; product owner otherwise
Prerequisites
Eval suites (golden sets incl. hallucination/injection tests) current; baseline scores recorded
Inputs
Candidate: model/prompt version, eval results vs baseline, cost delta projection
Tools
ML platform registry (ai-mlplat), prompt registry (ai-prompt), AI observability (ai-observe)
Security considerations
Injection red-team suite must pass for user-input-exposed prompts (C-AI-02); provider/data-terms re-check when the model provider changes.
Steps
  1. Register candidate with eval results: quality vs baseline, safety suite, latency, cost-per-call.
  2. Approval per risk class; record decision.
  3. Staged rollout via gateway routing (shadow → percentage → full), watching eval-proxy metrics and cost.
  4. Full promotion or halt; prior version retained hot for instant rollback.
  5. Post-promotion: scheduled evals continue; drift alerts route to owner (C-AI-06).
Validation
Live metrics within expected band vs offline evals.
Evidence
Registry promotion record; eval reports; rollout metrics
Failure conditions
Live regression → instant route rollback; eval-vs-live divergence → eval-suite improvement action
Rollback
Gateway route revert to prior version (kept hot through the watch window)
Escalation
AI platform lead; council for safety regressions
Acceptance criteria
No unevaluated artifact serving production; rollback exercised at least once in staging per quarter

RB-23 · Service incident & major incident response

Objective
Restore service fast with clear command, honest comms and learning that sticks.
Scope
Availability/performance incidents (security events: RB-16, invoked jointly when both)
Owner
On-call responder → incident commander at Sev-2+
Approvers
Comms: service owner; crisis escalation: duty executive
Prerequisites
Severity matrix published; on-call rotas staffed; paging tested (ops-incident)
Inputs
Alert (SLO burn, synthetic failure, user reports)
Tools
Observability (ops-observability), paging, ITSM, status comms
Security considerations
If compromise is plausible, engage RB-16 immediately — availability restoration must not destroy evidence.
Steps
  1. Acknowledge within paging SLA; classify severity per matrix.
  2. Sev-2+: name incident commander (command ≠ keyboard), open incident channel + timeline.
  3. Stabilize first (rollback RB-13, failover RB-17, degrade RB-21, scale RB-18) — root cause comes later.
  4. Communicate on cadence: internal stakeholders + user-facing status where visible.
  5. Verify restoration via synthetics + SLO recovery; close hypercare watch.
  6. Blameless post-incident review within days placeholder: timeline, contributing factors, action items with owners/dates; feed runbooks, alerts, model updates.
Validation
SLO recovery sustained; PIR actions tracked to closure.
Evidence
Incident record + timeline; comms log; PIR document
Failure conditions
Stabilization exhausted → crisis management + manual-workaround activation per continuity plan
Rollback
n/a — restoration is the goal
Escalation
Severity ladder to duty executive; vendor chain via RB-21
Acceptance criteria
MTTA/MTTR within targets placeholder; zero repeat incidents from unactioned PIR items