Phased execution plan (P0–P9) with gates, evidence and rollback stances, plus the operational procedure library (RB-01…RB-23).
Companion to the canonical model in data/*.json and the roadmap view.
Honesty header: the source document supplies themes, not an organization. Every duration, threshold and owner in this
runbook is a placeholder to be bound during Phase 0–1 (assumptions A-01, A-05, A-10, A-12). Nothing here claims compliance with any framework (A-11).
1 · Conventions, gates & evidence rules
Every phase ends at a gate: a go/no-go review against explicit exit criteria, chaired by the accountable governance body.
A gate passes only on evidence — artifacts a skeptical reviewer can open (test reports, policy exports, drill timings, signed records),
never on verbal assurance. Evidence is filed in the ITSM/document system referenced from the exit-criteria record (control C-GOV-03).
Three rules apply throughout:
Rule
Meaning
Rollback before cutover
No production cutover proceeds until its rollback has been rehearsed or a documented point-of-no-return has been signed by the accountable owner.
Placeholder discipline
Values marked placeholder (RTO/RPO bands, thresholds, volumes) must be replaced with organization-confirmed numbers at the phase that produces them — carrying a placeholder past its resolution phase is a gate finding.
Exceptions expire
Any deviation from this runbook is an exception with an owner, a compensating control and an expiry date (RB-20, principle PR-12).
2 · Operating teams & responsibility model
Team names are role placeholders (A-12); one person may hold several roles in smaller organizations, but the accountability split —
especially platform vs. domain vs. security — should survive any sizing.
Replace the illustrative baseline (A-01) with evidence; produce the BIA that turns placeholder tiers into commitments.
Entry criteria
P0 gate passed; inventory tooling and interview access arranged.
Inventory business capabilities, value streams, products, applications, infrastructure, integrations, data assets, identities, controls, operational processes, vendors, contracts, licenses and run costs.
Map dependencies and data flows; reconcile against this package's model — correct the model where reality differs (the model bends to evidence, never the reverse).
Assess each application: business criticality, technical health, security posture, lifecycle, cloud readiness, operational maturity, cost; record a 7R/8R disposition per application (see ADR discussion of the source's inconsistent R-lists in traceability).
Perform the business impact analysis: downtime and data-loss cost per business service; assign service tiers T1–T4 with named business sign-off; replace tier placeholders (A-05/A-07).
Answer open questions OQ-01…OQ-12 where discovery can; convert answers into model updates and decision inputs (OD-01 especially).
Baseline current run cost and delivery metrics (DORA baseline) so improvement is measurable (S48).
Gate (exit)
Inventory sign-off by domain owners · BIA v1 signed with tier assignments · disposition register complete · updated canonical model published (v1.x).
Evidence
Inventory dataset; BIA report; disposition register; model diff record.
Primary owners
Enterprise architecture (A), domain teams + SRE + FinOps (R).
P2 · Target Architecture & Roadmap Ratification
Objective
Turn this package's proposed target state into the organization's ratified target and a priced, sequenced roadmap.
Entry criteria
P1 gate passed (evidence-based current state exists).
Confirm architecture principles and quality-attribute scenarios; write measurable NFRs from the BIA and volume data (replacing OQ-04 placeholders in capacity records).
Revise the logical target architecture and transition states against discovery findings; re-run the model validator after every edit.
Decide the open decisions that gate foundation build: OD-01 (region/DR posture) with priced options, OD-03 (workflow realization), OD-06 stance review.
Evaluate modernization approach per application from the P1 dispositions (retain / retire / replace / repurchase / rehost / relocate / replatform / refactor).
Prioritize waves by business value, risk retired, dependency order, readiness, cost and team capacity; set wave exit criteria (the roadmap view holds the proposal).
Price wave 0–1 and secure funding through steering.
Gate (exit)
Target architecture ratified by ARB · OD-01/OD-03 decided · wave plan funded · NFR register versioned with real numbers where discovery produced them.
Stand up network foundation: hub with inspected egress, segmentation baselines, private endpoints, DNS/edge, hybrid connectivity with failover test (RB-04, C-NET-02/03).
Harden identity: workforce IdP MFA/conditional access, CIAM tenant, PAM with JIT + break-glass, HRIS-driven JML via IGA (RB-01/RB-02, C-ID-01…05).
Deploy central logging, SIEM onboarding for foundation sources, secrets/KMS services with rotation procedures (RB-06), backup platform baseline.
Run the guardrail validation suite: attempt-to-violate tests (public bucket, untagged deploy, unencrypted store, console change) must all be blocked and alerted.
Gate (exit)
Guardrail validation report all-green · MFA coverage at target · break-glass drill done · connectivity failover test done · tag enforcement proven.
Evidence
Validation suite output; policy pack export; drill records; coverage dashboards.
Primary owners
Cloud foundation (A/R), security engineering (R), FinOps (R).
P4 · Platform Engineering & DevSecOps
Objective
Make the compliant path the easy path: golden-path pipelines, GitOps delivery, supply-chain enforcement, developer portal.
Entry criteria
P3 gate passed.
Establish source-control standards: protected branches, mandatory review, separation of duties on deploy approvals (C-SC-03).
Stand up the trusted artifact registry; enforce provenance verification at cluster admission — test that unsigned images are rejected.
Deploy container platform (multi-AZ, admission-controlled) and PaaS runtimes per ADR-04; GitOps controllers reconcile all environments (C-CLD-02).
Launch the internal developer portal: service catalog with ownership/SLO/tier metadata, scaffolding, environment self-service (RB-05).
Define environment, repository, promotion, release, exception and support models; publish platform SLOs and the platform roadmap (platform-as-product).
Ship one reference service end-to-end through the golden path to production as the proving exercise.
Gate (exit)
Reference service in prod via golden path only · admission rejects unsigned/unattested images (tested) · platform SLOs published · DORA baseline captured.
Evidence
Pipeline runs; admission test logs; portal catalog export; SLO dashboard.
Light up governed connectivity: API management, events/queues with delivery-semantics discipline, B2B/EDI/MFT, iPaaS.
Entry criteria
P4 gate passed; API/event standards drafted by gov-api.
Deploy edge API gateway + APIM lifecycle + developer portal; onboard gateway logs to SIEM (C-OPS-03).
Stand up schema/contract registry with backward-compatibility enforcement wired into CI (C-INT-01).
Deploy event streaming backbone and work queues; publish the streams-vs-queues decision tree (ADR-07) and DLQ/replay conventions (C-INT-02/03).
Deploy B2B gateway (EDI/AS2) and MFT with non-repudiation logging; write the partner onboarding runbook into service (RB-08).
Deploy iPaaS for SaaS-sync flow classes with error stores and replay.
Publish API/event standards: naming, versioning, deprecation policy, idempotency and reconciliation requirements; automate linting in CI.
Build the legacy facade + ACL: read APIs and change events over the legacy core without modifying it (quick win; strangler enabler).
First DLQ + replay drill executed by the first consumer team (RB-10) — a consumer may not go live before its drill.
Gate (exit)
First 5 API products live with contracts and plans · schema gate active in CI · DLQ/replay drill evidence · legacy facade serving production reads · first partner migrated via RB-08.
Integration platform (A/R), gov-api (standards), domain teams (consumers).
P6 · Data & AI Platform
Objective
Governed data estate (lakehouse, catalog, quality, MDM, first products) and the AI foundation (gateway, RAG, evals, cost controls) — with autonomy explicitly gated.
Entry criteria
P5 event backbone live; data governance council operating; classification policy ratified.
Deploy lakehouse zones with promotion quality gates; CDC/batch ingestion for first sources via RB-09 (catalog-first: no source lands unclassified).
Deploy catalog/lineage, data-quality monitoring and MDM foundations; wire classification tags to access and masking policy (C-DP-01/03).
Publish the first two data products (Customer 360, Orders) with contracts, freshness SLOs and owner pagers (RB-11); begin legacy-DW parallel run with variance reporting.
Deploy the AI gateway with model allowlists, virtual keys, quota/cost metering, redacted logging and guardrail filters (C-AI-01/02); scan for and eliminate direct provider egress.
Stand up prompt registry, embedding pipeline (cataloged sources only), vector store with entitlement metadata, and the authorization-aware RAG service; run entitlement-bypass tests as a release gate (C-AI-03).
Stand up the ML platform: registry, promotion gates, scheduled evals incl. hallucination/drift suites (RB-22); AI observability + token budgets feeding FinOps (C-AI-06).
Ship the first internal RAG assistant (employee-facing) with HITL feedback loop.
Hold the line: no autonomous production actions — agent capabilities remain in transition status until the AI governance council approves risk classes, permissions, rollback, evidence and operating ownership (ADR-14 gates; RB-12 exists first).
Gate (exit)
2 certified data products with SLO dashboards · classification coverage 100% of onboarded sources · RAG bypass tests passed · zero direct provider egress (scan) · eval + cost dashboards reviewed by gov-ai.
Select pilot workloads by measurable business value and controlled risk (from the P1 disposition register); confirm each has an owner, SLOs and a rollback stance before build starts.
Wave 1 extractions: customer profile and order management services behind the BFF; parallel-run against the monolith with reconciliation (counts + amounts) until variance is below the agreed threshold.
Per-domain cutover: freeze window (RB-13) → traffic shift → hypercare → rollback window closes only on stability evidence; the monolith path stays warm through the window.
Wave 2 extractions: fulfillment orchestration, billing handoff, notification/document services; drain remaining ESB flow classes to APIM/events/iPaaS with per-flow drain-back procedures.
Rehost the irreducible VM remainder with forward dispositions recorded (transitional by policy, ADR-04).
For each wave: data migration with checksum/row-count verification, integration change tests, security validation (pen-test scope per change), performance tests against P2 NFRs, DR test participation, operational-readiness review (P8 checklist).
Gate (exit, per wave)
Reconciliation variance below threshold · rollback rehearsed before each cutover · ESB flow count trending to zero (wave 2) · no orphaned integrations (validator clean against updated model).
Domain teams (A/R per service), integration platform (R), EA (model updates).
P8 · Operational Readiness & Cutover Discipline
Objective
Prove each service is operable before it carries production traffic, and prove recovery works with timed evidence.
Entry criteria
Service candidate feature-complete; monitoring/alerting wired.
Operational-readiness review per service: dashboards, alerts with owners, SLOs, on-call rota, runbooks, capacity plan, backup/restore tested, security detections onboarded, support ownership and vendor escalation paths named.
Go/no-go review against explicit acceptance criteria; record approvals, rollback point, communication plan, freeze window and hypercare model.
Execute restore, failover and cyber-vault drills for tier-1 services (RB-17); capture timings against BIA objectives (or placeholder bands until then — flagged).
Deliver the OD-01 decision package: priced multi-region posture options against BIA downtime costs.
Close each cutover with a hypercare exit review and a lessons-learned record feeding the runbook library.
Gate (exit)
ORR checklist evidence per service · timed recovery drill records for T1 · OD-01 decided by steering · hypercare exits clean.
Run the estate as a product portfolio with measured reliability, security, cost and delivery health — and keep modernizing on cadence instead of by crisis.
Complete legacy decommission with evidence: archive per retention, contract wind-down, decommission certificates for the monolith, ESB and legacy DW (the R-01 kill shot).
Graduate AI autonomy per evidence: agent use cases move assist → recommend → bounded-auto only via AI-governance decisions with eval and audit evidence (never by default).
Feed everything back into this package: the model, registers and runbooks are living artifacts — a stale architecture repository is itself a gate finding at the quarterly review.
SRE + platform + EA (R), governance bodies (A per domain).
4 · Procedure library (RB-01 … RB-23)
Every procedure follows one template: objective, scope, owner, approvers, prerequisites, inputs, tools, security considerations,
steps, validation, evidence, failure conditions, rollback, escalation, acceptance criteria. Placeholders (thresholds, durations)
are bound at P0–P2. The 21 mandatory runbooks from the brief are all present; RB-08 (partner onboarding) and the RB-13/14 split are additions for operability.
RB-01 · Identity onboarding & offboarding
Objective
Provision and deprovision workforce identity and access from HRIS lifecycle events with zero orphan accounts.
Scope
Employees/contractors; workforce IdP, IGA, downstream SCIM-connected apps. (Customer/partner identity: CIAM policies, not this runbook.)
HRIS→IGA feed live; role catalog defined; birthright-access matrix approved
Inputs
JML event (hire/move/termination) with effective date, role, department
Tools
IGA (sec-iga), IdP (sec-idp-workforce), ITSM record
Security considerations
Terminations execute at effective time, not business hours; movers lose old access on grant of new (no accumulation); privileged roles never birthright.
Steps
IGA receives JML event; validates payload against schema.
Joiner: create identity, assign birthright bundle, enroll phishing-resistant MFA before first credential issue.
Mover: compute role delta; revoke-then-grant with manager confirmation on additions.
Leaver: disable sessions + tokens immediately, disable account, schedule deletion per retention, transfer owned resources, revoke PAM entitlements.
Propagate via SCIM; verify downstream completion states.
Write completion record to ITSM with timings.
Validation
Post-run reconciliation: IdP accounts vs HRIS active roster; zero unmatched.
Evidence
JML latency report; reconciliation output; MFA-enrollment record
Failure conditions
SCIM target down → queue and alert; conflicting identities → manual merge procedure
Rollback
Joiner/mover grants reversible via IGA transaction log; leaver disable is not rolled back without HR instruction
Escalation
Security operations (orphan or failed-termination cases are Sev-2)
Acceptance criteria
Termination-to-disable within agreed minutes placeholder; zero orphan accounts in weekly scan
Grant request: use case, tool(s), scopes, transaction limits, autonomy level, accountable human
Tools
Tool registry (ai-tools), AI gateway, HITL service, SIEM
Security considerations
Least-privilege per tool (never a broad service account); limits enforced in the sandbox, not the prompt; injection red-team results current (C-AI-02) before customer-facing grants.
Steps
Requester files grant with use-case risk class and accountable owner.
Security reviews effective permissions (diff against least privilege); council approves level (assist → recommend → bounded-auto).
Service owner (failover); steering (regional invocation per OD-01 posture)
Prerequisites
Backup policies active with immutability on T1 sets; recovery sequence documented per tier (dependency-ordered)
Inputs
Trigger: scheduled test, incident, or DR invocation
Tools
Backup platform (ops-backup), cyber vault (ops-cyber-vault), DR orchestration (ops-dr), IaC
Security considerations
Vault credentials separate from production IdP; restored systems rejoin only after clean-scan when incident-driven (RB-16).
Steps
Restore test (scheduled per tier): select restore point → restore to isolated environment → integrity checks (row counts, checksums, app-level smoke) → record timing vs objective band.
Ride out SaaS/cloud/provider outages with pre-decided degradation stances instead of improvisation.
Scope
SaaS SoRs (ERP/CRM/HRIS/ITSM), payment provider, model providers, cloud service events
Owner
Service owner of the consuming capability; vendor manager engages the vendor
Approvers
Business owner for degraded-mode entry where customer-visible
Prerequisites
Per-dependency degradation matrix documented (what still works, what queues, what stops); vendor escalation contacts + SLAs on file
Inputs
Vendor status signal or synthetic/consumer failure detection (do not wait for the vendor's status page)
Tools
Observability, incident tooling, queues (absorption), AI gateway fallback routes
Security considerations
Degraded modes must not bypass authentication or approval controls "temporarily".
Steps
Detect + confirm scope (their outage vs our path to them).
Declare incident (RB-23); enter the pre-decided degradation stance: queue writes with replay (ERP-class), read-from-cache with staleness banners (CRM-class), provider failover at gateway (model-class), defer (analytics-class).
Open vendor escalation per contract; track their incident id.
Communicate honestly to affected users (status page templates).
On restoration: drain queues with reconciliation (C-INT-05), verify data integrity, exit degraded mode.
PIR includes: did the degradation matrix hold? Update it.
Validation
Reconciliation zero-variance after drain; degradation matrix confirmed accurate.
Observability (ops-observability), paging, ITSM, status comms
Security considerations
If compromise is plausible, engage RB-16 immediately — availability restoration must not destroy evidence.
Steps
Acknowledge within paging SLA; classify severity per matrix.
Sev-2+: name incident commander (command ≠ keyboard), open incident channel + timeline.
Stabilize first (rollback RB-13, failover RB-17, degrade RB-21, scale RB-18) — root cause comes later.
Communicate on cadence: internal stakeholders + user-facing status where visible.
Verify restoration via synthetics + SLO recovery; close hypercare watch.
Blameless post-incident review within days placeholder: timeline, contributing factors, action items with owners/dates; feed runbooks, alerts, model updates.
Validation
SLO recovery sustained; PIR actions tracked to closure.
Evidence
Incident record + timeline; comms log; PIR document
Failure conditions
Stabilization exhausted → crisis management + manual-workaround activation per continuity plan
Rollback
n/a — restoration is the goal
Escalation
Severity ladder to duty executive; vendor chain via RB-21
Acceptance criteria
MTTA/MTTR within targets placeholder; zero repeat incidents from unactioned PIR items