# 8. Cloud vs. On-Premises: What Goes Where?
Let's tackle the next one, which is a bit different in shape because it's not really a "vendor question" topic, it's an architectural decision framework. So I'm going to adapt the format slightly: questions you should ask yourself, your team, and your vendors to make smart cloud-vs-local decisions instead of being driven by vendor pitches or hype.
Context: For a decade, the answer to "should we move this to the cloud?" was reflexively yes. The pendulum is now swinging back. We're seeing major repatriation stories (37signals/Basecamp publicly leaving AWS, GEICO, Dropbox earlier, plus a quiet wave of companies pulling specific workloads back on-prem). At the same time, "go cloud" is still right for many things, and "go local" is right for others. The honest answer is it depends on workload, regulation, economics, and your team's skills, not on vendor narratives.
The trap most "regular" companies fall into: they make this decision based on the vendor in front of them at the time, not on a coherent framework. A CSP rep tells them everything should be in the cloud. A hardware vendor tells them on-prem is making a comeback. An MSP tells them hybrid is the answer (because that's their service offering). All of them are partly right and significantly self-interested.
The core question isn't "cloud or local?", it's "for this specific workload, what's the right home given cost, risk, latency, regulation, control, and skills, over a 5–10 year horizon?"
14 Questions to Ask Before Moving Workloads to the Cloud (or Bringing Them Back)
1. "What's the real total cost of ownership over 5 years, including egress, support, premium services, and the people to run it, vs. the local equivalent?"
Why: Cloud TCO is famously underestimated at the proposal stage. Egress fees, premium support, managed services markup, and talent costs add up. On-prem is overestimated for OpEx but underestimated for capital flexibility. Run the math honestly.
What to look for: A real spreadsheet with assumptions documented. Engage Finance, not just IT. Red flag: A vendor (cloud or hardware) presenting only their preferred scenario.
2. "How predictable is this workload's compute, storage, and bandwidth profile? Steady-state, spiky, or unpredictable?"
Why: Cloud's economic strength is elasticity. Steady-state predictable workloads (e.g., your production database running 24/7 at 60% utilization) are often cheaper on-prem at scale. Spiky/seasonal workloads (e.g., retail at Black Friday, batch ML training) thrive in the cloud.
What to look for: Honest workload characterization. Red flag: "Cloud handles all workload patterns equally well" (true, but at very different cost efficiencies).
3. "What data does this workload touch, where is it regulated, and what are the residency, sovereignty, and breach-notification implications?"
Why: GDPR, HIPAA, PCI, FedRAMP, ITAR, state-level laws, and emerging sovereignty rules (especially in the EU) all have data location implications. Some data simply cannot leave certain jurisdictions, certain providers, or certain control planes.
What to look for: Legal/compliance team consulted early, not as a last-mile checkbox. Red flag: "We'll figure out compliance after we move" (you won't, and you'll regret it).
4. "What's the latency requirement, and what's the cost of milliseconds for this workload?"
Why: Some workloads (manufacturing OT, real-time control systems, high-frequency trading, clinical imaging, retail POS) have hard latency requirements that the cloud, even with regional placement, can't meet reliably. Others are completely insensitive.
What to look for: Measured latency requirements with business consequences. Red flag: Hand-wavy "performance" requirements without numbers.
5. "If this provider has an outage, region failure, or terminates our contract, what's the business impact and recovery plan?"
Why: Cloud concentration risk is real. Multi-region, multi-cloud, and hybrid strategies are insurance, and they have costs. Some workloads warrant the insurance; others don't.
What to look for: Documented BCDR plans, tested failover. Red flag: "We trust [provider]" without contingency.
6. "Does this workload need integration with on-prem systems that aren't going anywhere, and what's the network, identity, and data sync complexity?"
Why: Hybrid is harder than pure cloud or pure on-prem. Data gravity is real, moving an app while its data dependencies stay home creates latency, cost, and reliability problems. The "lift and shift" trap.
What to look for: Honest assessment of dependencies. Red flag: "We'll just put a VPN in" (it's never just a VPN).
7. "What does the team need to operate this workload, and do we have those skills, can we hire them, or do we need a partner?"
Why: Cloud-native operations require different skills than traditional infrastructure. Misjudging this is the #1 reason cloud projects fail or get repatriated. A workload that "works in the cloud" but can't be effectively operated by your team is worse than one that stays local.
What to look for: Honest skills audit. Red flag: "The cloud is easier to manage" (it's different, sometimes easier, often not).
8. "What's the data egress story? How much data leaves this workload monthly, and what does that cost, including for backups, DR, and analytics?"
Why: Egress fees are the cloud's silent tax. Workloads that produce lots of data (video, logs, analytics, backups) become punitively expensive to move data out of once they're in. This is also the strongest lever cloud providers have to keep you locked in.
What to look for: Egress modeled in TCO, exit costs estimated. Red flag: "Egress is minimal" without measurement.
9. "How sensitive is this workload's intellectual property, and what's the trust model with the cloud provider's privileged access?"
Why: Cloud providers have privileged access to underlying infrastructure. For highly sensitive workloads (proprietary models, trade secrets, government data), confidential computing or on-prem may be the right answer. This is a risk-tolerance conversation, not a technical one.
What to look for: Threat model that includes the provider as an actor. Red flag: "We trust [hyperscaler] completely" without nuance.
10. "What's the lifecycle of this workload? Will it run for 10+ years, or is it experimental and likely to change?"
Why: Long-lived, stable workloads benefit from on-prem amortization. Short-lived, experimental, or rapidly evolving workloads benefit from cloud elasticity and managed services. Match the venue to the lifecycle.
What to look for: Honest lifecycle estimates. Red flag: "Everything is dynamic now" (untrue for most enterprise workloads).
11. "Can we lift and shift, or do we need to re-architect, and what's the realistic effort and cost of each?"
Why: Lift-and-shift cloud migrations often deliver the worst of both worlds: cloud pricing without cloud benefits. Re-architecting captures the elasticity and managed services value but takes 2–5x longer. Be honest about which you're really doing.
What to look for: A migration strategy that distinguishes the two, with realistic timelines. Red flag: "Lift and shift, then optimize later" (later rarely comes).
12. "What does our exit plan look like, both for leaving the cloud and for changing cloud providers?"
Why: Lock-in is a spectrum. Open formats (Kubernetes, Postgres, S3-compatible storage) are portable. Proprietary services (DynamoDB, BigQuery, Cosmos DB) are not. Architectural choices today determine flexibility a decade out.
What to look for: Conscious choices about which services to use, with portability in mind. Red flag: "We'll figure out exit if it ever happens" (the longer you wait, the more expensive it gets).
13. "What's our security model, and does it work consistently across cloud and on-prem, or are we doubling our control surface?"
Why: A hybrid environment with two security models (one for cloud, one for on-prem) is harder to defend than either alone. Identity, monitoring, network segmentation, and data protection should be consistent or you'll have gaps.
What to look for: A unified security architecture across venues. Red flag: Separate teams managing separate stacks with no shared visibility.
14. "What's our environmental and energy story, and does that factor into this decision?"
Why: This is becoming a real factor, not just for ESG reporting, but because hyperscalers are running into power constraints, and some on-prem deployments (especially with renewable contracts) can be more sustainable than commodity cloud regions. Increasingly relevant for regulated industries and public reporting.
What to look for: Sustainability data integrated with infrastructure decisions. Red flag: Treating this as a separate, downstream concern.
🎯 The Meta-Framework
Here's the simplified decision lens I'd give a CIO/CISO making these calls:
| Workload Trait | Lean Cloud | Lean On-Prem |
|---|---|---|
| Compute pattern | Spiky, unpredictable, seasonal | Steady-state, high-utilization |
| Data sensitivity | Standard / mainstream regulated | Highly classified / sovereign / IP-critical |
| Latency | Tolerant (>20ms ok) | Hard real-time (<5ms) |
| Lifecycle | Short, experimental, evolving | Long-lived, stable, mature |
| Skills | Cloud-native team | Traditional infrastructure team |
| Data gravity | Cloud-resident upstream/downstream | On-prem-resident dependencies |
| Scale | Variable, growing fast | Predictable, large steady-state |
| Cost predictability needs | OpEx-friendly, variable acceptable | CapEx-friendly, predictable preferred |
💡 Honest Observations
For "regular" companies, the patterns I see consistently:
- SaaS for commodity functions, almost always. Email, HR, CRM, collaboration, ITSM, buy these as SaaS. The economics and security are usually better than self-hosted, even on-prem. This is the "easy yes."
- Cloud for new development, especially elastic or experimental workloads. Greenfield apps, data analytics, ML, dev/test environments, these benefit most from cloud-native services.
- On-prem (or colo) for steady-state, predictable, data-heavy workloads at scale. If you're running a 24/7 database at consistent load with high I/O, the math often favors on-prem once you cross a certain scale (the "Basecamp threshold," roughly).
- Edge/local for latency-sensitive, OT, and air-gapped requirements. Manufacturing floors, hospitals, retail stores, remote sites, these often need local compute regardless of cloud strategy.
- The repatriation wave is real but selective. Companies aren't leaving the cloud entirely; they're moving specific workloads back where the economics or control demand it. This is mature, not regressive.
The biggest mistake I see: companies make this decision based on the current loudest vendor in the room rather than a workload-by-workload framework. A coherent strategy is "we have criteria, and each workload gets evaluated against them", not "we're a cloud-first company" or "we're keeping everything on-prem." Both extremes are usually wrong.
The second-biggest mistake: treating this as a one-time decision. The right answer changes as workloads evolve, costs shift, regulations change, and team skills develop. Revisit the decision every 2–3 years per significant workload.
The third-biggest mistake: ignoring the security model implications. Where workloads live changes how you defend them. CASB, SSE, ZTNA, SASE, these all exist because the perimeter dissolved when workloads spread out. Plan for the security architecture, not just the workload location.