Designing a V2V Migration That Survives Contact With a 24/7 Clinical Workload
A field-tested strategy for moving a VMware estate to Azure without pretending lift-and-shift is easy: how to run the provider evaluation, design waves around downtime tolerance, protect data integrity, plan a rollback you can actually execute, and validate the result — in a regulated, always-on environment.
Every lift-and-shift proposal I have reviewed in the last decade opens with the same sentence: "This is a straightforward as-is migration — no application changes." That sentence is usually true about the virtual machines and almost never true about the project. The VMs are the easy part. Modern replication tooling will copy a running virtual disk into a cloud platform with a fidelity that would have looked like science fiction in 2012. What the tooling will not do is tell you which of your 140 machines can tolerate a two-hour outage, which application will quietly break when its database is 4 milliseconds further away, who signs off when a wave goes wrong at 3 a.m., or what your bill looks like in month seven once the migration discount stops flattering the numbers.
This post is a consolidated view of how I approach a virtual-to-virtual (V2V) migration off an on-premises or hosted VMware estate onto Azure. The framing is drawn from a composite of engagements in regulated, always-on environments — the recurring pattern I keep meeting, not any single customer. Picture a mid-sized US healthcare software provider: multi-tenant clinical applications reached around the clock by hospitals, nursing homes, provider groups and medical-device manufacturers; roughly 130 virtual machines on a hosted VMware platform; a contract renewal forcing a decision; and two stated objectives that are in permanent tension — improve performance and optimise consumption cost.
Nothing here is vendor-specific in a way that would not translate to AWS or Google Cloud. I am using Azure as the worked example because for this class of estate it is frequently, though not automatically, the right answer.
Start with the honest question: is the cloud even the answer?
A consulting engagement that begins with "we have chosen Azure, now plan the migration" has skipped the step that determines whether the project succeeds financially. The first deliverable should be a hosting evaluation, and it should genuinely consider co-location and dedicated hosting alongside the three hyperscalers. I have recommended against public cloud twice for workloads whose profile was flat, predictable, storage-heavy and fully depreciated. Both times the client was relieved someone had done the arithmetic rather than the marketing.
For the profile above, the evaluation usually comes down to four questions.
What does the estate actually look like? Not what the CMDB says. Two to four weeks of real telemetry — CPU, memory, disk IOPS and throughput, network flows, uptime patterns — collected from the hypervisor rather than from inside the guests. Discovery consistently finds 10–20% of machines that nobody can identify an owner for, plus a handful doing far more I/O than anyone expected.
Where does the licensing land? This is the single largest swing factor in a Microsoft-heavy estate and it is routinely omitted from comparisons. Existing Windows Server and SQL Server licences with Software Assurance carry across to Azure at a materially different effective rate than they do elsewhere, and end-of-support operating systems and database versions receive extended security updates on Azure that must be purchased in a co-location scenario. In several evaluations this one line item has been worth more than every compute-price difference combined.
What is the exit cost? Ask it about every option, including the incumbent. Data egress, contract termination, dual-running during transition, and the re-platforming effort you would face if you had to leave in three years. An evaluation that models entry cost but not exit cost is a sales document.
Who can operate it on Monday morning? A platform that your team cannot run at 2 a.m. without a partner on retainer is not cheaper, whatever the spreadsheet says. Where the client's requirement is to keep delivery in-house and avoid third-party processors — common in healthcare, where every additional vendor is another business-associate agreement and another breach-notification path — I weight native, first-party platform services heavily and treat third-party migration and backup tooling as a cost and compliance line, not a free convenience.
The comparison table that survives scrutiny
| Dimension | What to compare | The mistake to avoid |
|---|---|---|
| Compute | 3-year modelled run rate at right-sized capacity, not lifted-as-is capacity | Comparing list price to your current invoice |
| Storage | Tiered to measured IOPS and latency, not matched to current allocated size | Provisioning premium storage for every disk "to be safe" |
| Licensing | Existing entitlements, hybrid-use benefits, extended security updates | Treating licensing as a rounding error |
| Network | Egress, dedicated circuits, inter-region and cross-zone traffic | Forgetting that chatty applications generate internal charges |
| Resilience | Backup retention, DR target, and the cost of testing both | Costing the backup and not the restore |
| Commercial | Reserved capacity and savings-plan coverage at a realistic 60–75% | Modelling 100% reservation coverage on day one |
| Exit | Egress and re-platform cost if the decision is reversed in year three | Not modelling it at all |
Present the cost report as three numbers per option: day-one run rate (everything lifted as-is, no commitments), optimised run rate (right-sized after 30–60 days of production telemetry, with reservations applied), and three-year total cost including migration effort and dual-running. The gap between the first two numbers is typically 30–45%. Leadership needs to see both, because the first invoice will look like the project failed and someone must have already explained why.
Native IaaS or a VMware-compatible landing zone?
Once Azure is chosen, there is a second architectural fork that materially changes timeline, cost and risk: rehost onto native Azure IaaS, or move the vSphere estate onto a VMware-compatible private cloud inside Azure.
The VMware-compatible path preserves your existing hypervisor operating model. You keep vCenter, your existing networking constructs, your existing VM management practices, and you can move large numbers of machines with minimal per-VM engineering and near-zero guest-level change. It is the faster, lower-risk route when the estate is large, the deadline is contractual, or the workloads carry unusual OS and appliance dependencies. Its cost profile, however, is node-based rather than per-VM: you buy capacity in host increments, so the economics reward dense, well-utilised estates and punish small or shrinking ones.
Native IaaS is the better destination when the estate is a few hundred machines or fewer, when the long-term intent is to modernise (managed databases, platform services, autoscaling), and when you want per-VM cost granularity and the full native service catalogue without an intermediate layer. The trade-off is more per-machine work: driver and agent changes, sizing decisions, network redesign, and a genuine shift in the operating model your team must absorb.
For a ~130-VM estate with a modernisation ambition, native IaaS is usually the right destination — with one important caveat. If the deadline is driven by a contract expiry rather than by strategy, the VMware-compatible landing zone is a legitimate intermediate step: move to safety first, modernise on your own schedule. Just make sure the intermediate step is written down as an intermediate step, with a date and an owner, or it becomes permanent by default.
Design the waves around downtime tolerance, not around technology
The most consequential design decision in a V2V programme is how you group machines into migration waves. Get this wrong and every other control is compensating for it.
The instinct is to group by technology — all the web servers, then all the databases. This is almost always wrong. Group by application service and dependency boundary, so that everything which talks to everything else moves together within a single change window. Splitting a chatty application tier from its database across a wide-area link is the single most reliable way to generate an "everything is slow since the migration" escalation.
Then sequence the waves by risk and by downtime tolerance:
- Pilot wave — internal, non-clinical, fully owned by IT. Development, build agents, internal tooling. The purpose is to prove the pipeline, the runbook and the rollback, not to save money.
- Low-tolerance-for-nothing wave — batch, reporting, archive, back-office. Systems where a four-hour window on a Saturday is genuinely available.
- Business-critical waves — the clinical applications, split as finely as dependencies allow. Small waves. Multiple weekends. Each one rehearsed.
- The hard residue — the licence servers, the legacy appliance nobody has patched since 2019, the machine with a physical dongle, the integration that hard-codes an IP address. Identify these in week two, not in month five. They need individual plans and sometimes individual budgets.
For each wave, record three numbers agreed in writing with the application owner: maximum tolerable outage, maximum tolerable data loss, and the named person who can authorise both a cutover and a rollback at 3 a.m. without escalating. Migration plans fail at 3 a.m. because nobody knows who decides.
The seven problems that actually consume the project
1. Downtime in an environment that has no maintenance window
Hospitals do not stop. The mitigation is architectural, not heroic: use replication-based migration so that the bulk data movement happens while the source is live and serving traffic. Initial replication runs for as long as it takes — days, for large volumes — followed by periodic delta cycles that track only changed blocks. The actual outage is then reduced to a final delta sync, a controlled shutdown, and a start-up and validation in the target. For a well-behaved application tier that is 30–90 minutes, not eight hours.
Three practices make this reliable:
- Rehearse with test migrations into an isolated network. Every mainstream tool supports booting a replica in a sandbox without touching the source or the replication schedule. If a wave has not been booted and smoke-tested at least once before its real cutover, it is not ready.
- Watch the replication window, not the replication job. Delta cycles are scheduled relative to how long the previous cycle took; a machine with heavy write churn can drift into long cycles that quietly lengthen your cutover. Track per-VM cycle duration for a week before committing to a window.
- Respect source-side limits. Replication reads from the same storage the production workload is using. Aggressive parallelism during business hours is a self-inflicted performance incident on a platform you are about to leave — which nobody will remember when they describe how badly the migration went.
2. Data integrity
The rule I hold to is: the source remains authoritative until a declared point of no return. Before that point, the target is a copy under test. After it, the source is frozen — powered off, network-isolated, snapshotted, and explicitly retained on a defined retention schedule rather than deleted at the end of the weekend.
For file and OS data, block-level replication with change tracking is trustworthy provided you verify: boot, checksum a sample set, and confirm disk and volume counts. For transactional databases, I prefer database-native replication over generic block replication for the largest and most active instances — log shipping or native replication into the target, then a controlled role switch. It gives cleaner consistency semantics, a shorter cutover, and a far more credible rollback story. Application-consistent snapshots matter for the rest, and multi-VM consistency groups matter for anything with distributed state.
Validation is a checklist executed by the application owner, not by the migration engineer: row counts on key tables, most-recent-transaction timestamps, a reconciliation report against a pre-cutover baseline, and one real end-to-end business transaction. Sign-off is a name and a timestamp.
3. "As-is" does not mean "insecurely as-is"
A lift-and-shift copies the guest, but it does not copy the perimeter that was protecting it. This is where the majority of post-migration incidents originate, and it is entirely preventable.
Build the landing zone before the first machine moves: subscription and resource-group structure, network segmentation that reproduces (and usually improves on) the existing tiering, no public IP addresses on migrated servers by default, brokered administrative access rather than open management ports, centralised identity, platform-native threat detection, encryption at rest with clear key ownership, and diagnostic logging to an immutable, access-controlled destination from day one. Set guardrails as policy so drift is prevented rather than audited later.
Two things I insist on. First, the migration is the cheapest security remediation you will ever get — flat networks, legacy protocols and shared local administrator accounts can be fixed during a change window you already have. Take the opportunity where it does not add risk to the cutover itself, and log the ones you defer. Second, treat the migration tooling's own footprint as in-scope: it holds credentials to your entire estate and it replicates production data across a network boundary. It gets the same review as anything else that touches protected data.
4. Backup that has actually been restored
Every migration plan contains a backup requirement. Very few contain a restore test, which is the only part that matters.
Configure protection in the target before the wave goes live, not after: policy-driven backup with retention that matches your regulatory obligation, immutability and soft-delete enabled so that a compromised administrator account cannot destroy recovery points, and separation between the identity that runs the workload and the identity that can alter backup policy. Then, before the wave is declared complete, restore one machine and one database from the target's own backup and document the elapsed time. That number is your real recovery-time capability. It is often three times the figure in the DR plan, and it is much better to learn that on a Tuesday.
Keep the source platform's backups intact and accessible for a defined period after cutover — typically 30 to 90 days. Cancelling the old contract on the Monday after the final wave is a recurring, avoidable mistake.
5. Performance validation, and the trap inside it
Baseline before you move or you will argue forever. Capture, for each significant application: transaction response times at representative load, batch and report completion times, database wait statistics, and peak storage IOPS and latency. Without this, every post-migration complaint becomes an unfalsifiable claim, and you will spend months chasing ghosts.
Two failure modes account for most genuine performance regressions:
- Storage tiering by habit. On-premises estates often run on a single shared high-performance array, so every VM inherits good storage whether it needs it or not. In the cloud, storage performance is an explicit, priced choice. Match tiers to measured IOPS and latency requirements — under-provisioning the database causes an incident, over-provisioning the file servers causes a budget review.
- Latency amplification. An application that makes 400 sequential database calls per page is fine at 0.2 ms and unusable at 4 ms. This is why dependency-aware wave grouping matters, and why any period of hybrid operation with a split application tier must be short, deliberate, and monitored.
Also be precise about what "improve performance" means to the customer. Sometimes it means response time. Often it actually means predictability — no more 9 a.m. contention with a noisy neighbour on shared infrastructure. Those lead to different designs, and the difference should be settled before the architecture is drawn, not after.
6. Cost control after the applause stops
The first full month's invoice is the most dangerous moment in a cloud migration, because it arrives before any optimisation is possible and it is what people remember.
- Do not right-size during the migration. Lift like-for-like, prove stability, then optimise on 30–60 days of real production telemetry from the new platform. Changing capacity and location in the same change window makes every incident undiagnosable.
- Do not commit to reserved capacity on day one. Wait until the shape is stable, then cover the predictable base — typically 60–75% — with commitments and leave the remainder flexible.
- Turn off the lower environments. Development and test machines that ran continuously because the hypervisor was a sunk cost are now billed by the hour. Scheduled shutdown outside working hours is one of the fastest wins available and it is usually left undone for months.
- Tag from the first machine, not from the first invoice. Application, owner, environment, cost centre, data classification. Retro-fitting tags across a live estate is miserable work; enforcing them at deployment is a policy setting.
- Publish showback monthly to the people who can act on it. Cost visibility that stops at Finance changes nothing.
Expect and communicate a curve: day-one run rate high, month-three run rate materially lower after right-sizing and scheduling, month-six lower again after commitments. Set that expectation in the cost report so the trajectory reads as the plan working rather than the plan failing.
7. Compliance continuity
For protected health information the compliance requirements do not change because the hypervisor did. What changes is the evidence you need and who you need it from. Confirm the platform agreement covering protected data is executed and that every service you use is in scope. Keep encryption, retention, access control, audit logging and breach-notification processes continuous across the cutover — a gap in audit logging during a migration weekend is a finding, and it is trivially avoidable. Document data residency, and be deliberate about it: replication traffic, backup copies and DR targets all place data somewhere, and "somewhere" needs to be a decision rather than a default. Finally, run the migration itself through change control with per-wave records. Auditors do not object to migrations; they object to undocumented ones.
Rollback: the section most plans do not really have
Most migration plans contain a paragraph titled "Rollback" that says, in effect, we will restore from backup. That is not a rollback plan. It is a hope.
A usable rollback plan has four properties.
A defined point of no return, per wave. Before it, rollback means powering the source back on and reversing DNS — minutes, not hours, because the source was never destroyed. After it, rollback means restoring data written in the target back to the source, which is a different and much slower operation. Everyone in the change window must know which side of that line they are on at all times.
Named triggers, agreed in advance. Not "if things go badly". Write them: validation checklist fails; error rate above an agreed threshold sustained for 15 minutes; response time above an agreed multiple of baseline; any suspected data-integrity issue (immediate, no debate); cutover exceeding its window with no confirmed path to completion.
A time box. Every wave gets a hard decision time — typically 60–75% into the agreed window. At that moment the decision is made: continue or roll back. Teams roll back badly because they decided too late, having spent the remaining window troubleshooting.
A rehearsal. Roll back at least once during the pilot wave, deliberately, and time it. A rollback that has never been executed is an assumption.
The corollary is a discipline that must hold for the whole programme: do not decommission anything until the wave has been stable in production for an agreed period. Dual-running costs money for two to eight weeks. It is the cheapest insurance in the project, and the pressure to cancel it early should be resisted in writing.
Post-migration validation and hypercare
Cutover is not completion. Validation runs at four levels, and each has a different owner:
- Infrastructure (migration engineer): the machine boots, disks and volumes match, agents report, time sync and name resolution are correct, backup has run and been verified, monitoring is reporting.
- Application (application owner): services start clean, integrations connect, scheduled jobs run, certificates and licences are valid, authentication works for every identity type including service accounts.
- Business (a real user): a genuine end-to-end transaction, executed by someone who does the job daily and will notice something an engineer would not.
- Non-functional (architect): performance against the pre-migration baseline, security configuration against the landing-zone standard, cost against the model, and evidence captured for audit.
Then run a 30-day hypercare with elevated monitoring, a daily check-in for the first week, a standing bridge for the first weekend, and — critically — a decision review at day 30 that formally closes each wave and releases its rollback obligations. Hypercare is also when you collect the operational documentation, because the details are still fresh and the people who know them are still assigned.
The deliverables worth insisting on
If I were writing the statement of work, these are the artefacts I would hold the engagement to:
- Hosting and cost evaluation — all realistic options, three-year modelled TCO, day-one versus optimised run rate, exit cost, and a stated recommendation with its reasoning exposed.
- Migration plan — wave composition with dependency rationale, per-wave outage and data-loss tolerance, named decision-makers, rehearsal schedule, rollback triggers and time boxes.
- Target architecture — landing zone, network and segmentation design, identity and access model, backup and DR design with stated recovery objectives, and the operational model.
- Access and security configuration — administrative access model, guardrail policies, logging and monitoring design, and the compliance evidence map.
- Post-migration validation pack — the four-level checklist, baseline versus actual performance results, restore-test evidence, and formal wave sign-offs.
- Operations handover — runbooks, cost-management practice, and a documented list of everything deliberately deferred, so the next person inherits decisions rather than mysteries.
That last item is the one clients thank you for two years later. Every migration defers something. Writing down what you deferred, and why, is the difference between technical debt and technical mystery.
If you are planning a V2V programme off VMware — or trying to work out whether the cloud is even the right destination for your estate — I am happy to compare notes on the evaluation approach. The contact page is the best way to reach me.
- #Cloud Migration
- #Azure
- #VMware
- #Lift and Shift
- #V2V
- #Healthcare IT
- #HIPAA
- #FinOps
- #Disaster Recovery
- #Migration Strategy
Related articles
Practical Cloud Landing Zone Security Considerations
Landing zone security is mostly decided before the first workload arrives. The guardrails, identity boundaries, and logging defaults that matter — and the ones that just add friction.
3 min read