# 15. Backup, Recovery, and Cyber Resilience
This is the topic nobody wants to spend money on, until they desperately need it. And then they wish they'd spent a lot more.
Context: Backup and recovery used to be an IT operations concern. Ransomware made it a security-critical, board-level conversation. The question is no longer "do we have backups?", almost everyone does. The question is "can we actually recover, how fast, and have we proven it?"
The threat has also shifted the requirements fundamentally. Traditional backup was designed around hardware failures, accidental deletions, and natural disasters. Ransomware adversaries specifically target backup infrastructure, because if they can corrupt or encrypt your backups, your only option is to pay. The best-resourced ransomware groups spend weeks inside environments identifying and destroying backup systems before triggering encryption. Your backup architecture needs to be built assuming a sophisticated adversary is looking for it.
The market spans a wide range:
- Traditional enterprise backup, Commvault, Veeam, IBM Spectrum Protect, Veritas. Mature, complex, feature-rich, often under-tested.
- Modern data protection platforms, Rubrik, Cohesity. Built with security-first thinking, immutability, and ransomware recovery as design principles rather than afterthoughts.
- Cloud-native backup, AWS Backup, Azure Backup, Google Cloud Backup. Convenient for cloud workloads, but recovery speed and cross-cloud portability can be limitations.
- Disaster Recovery as a Service (DRaaS), Zerto, VMware Site Recovery Manager, Veeam Cloud Connect. Focus on replication and rapid failover rather than traditional restore.
- Immutable storage, A capability across platforms, not a separate vendor category. Object lock, air-gapped vaults, write-once media. The non-negotiable requirement in the ransomware era.
The core question for any backup and recovery purchase: "When we're hit, not if, how fast can we be back, what will we lose, and can we prove it before we're in the middle of a crisis?"
14 Questions to Ask Backup & Recovery / Cyber Resilience Vendors
1. "What is your immutability architecture, specifically, how do backups become unmodifiable, and can an attacker with domain admin credentials delete or encrypt them?"
Why: This is the ransomware question. Most backup products added "immutability" as a feature, the maturity and completeness of that implementation varies enormously. Domain admin credentials, API access, and management console access are all potential paths to backup destruction. A vendor who can't explain the technical mechanism is selling a checkbox.
Good answer: Object lock with governance vs. compliance mode explained, air-gapped copy architecture, MFA-required deletion, management plane isolated from production environment. Red flag: "Our backups are immutable" without explaining the mechanism.
2. "Show me your 3-2-1-1 story, three copies, two media types, one offsite, one air-gapped or offline. Where does your product sit in that architecture?"
Why: The 3-2-1 rule has been extended to 3-2-1-1 in the ransomware era specifically because connected backups can be reached. An air-gapped or truly offline copy that can't be accessed from the network is the last line of defense. Not all vendors support this, and "cloud" is not the same as "air-gapped."
Good answer: Clear architecture supporting all four tiers, honest about what they cover vs. what you need to add. Red flag: "Cloud backup is your offsite and air-gap" (a cloud backup reachable via credentials isn't air-gapped).
3. "What are your tested RTO and RPO numbers, not aspirational, not theoretical, but from actual recovery tests in customer environments similar to mine?"
Why: RTO (Recovery Time Objective) and RPO (Recovery Point Objective) are almost universally optimistic until tested. Vendors will quote best-case numbers in clean environments. The realistic numbers in complex environments, partial recoveries, dependency ordering, database consistency, are often 5-10x worse.
Good answer: Tested numbers from reference customers in your size range, methodology for measurement, acknowledgment of variance across workload types. Red flag: SLA-level RTOs quoted without testing evidence.
4. "How do you handle recovery sequencing, databases with dependencies, applications that must restart in order, active directory, domain controllers?"
Why: Recovering individual files is straightforward. Recovering a production environment at scale requires orchestrated sequencing, if you bring up application servers before domain controllers, nothing works. This orchestration capability separates basic backup from genuine DR.
Good answer: Runbook automation, dependency mapping, tested recovery sequences, AD/DC-specific recovery capabilities. Red flag: "You define the recovery order in the runbook" (= you build it yourself with no tooling).
5. "How do you detect whether a backup set is already compromised, dormant ransomware, corrupted data, before you restore it?"
Why: Restoring from a backup that contains dormant malware or early-stage encryption resets you to the attack, not to safety. Some modern backup platforms scan backup sets for indicators of compromise before restoration. This is increasingly critical for environments where attackers dwell for weeks before triggering.
Good answer: Backup integrity scanning, malware detection in stored backup sets, anomaly detection on backup data changes, clean-room restore options. Red flag: "Our backups are encrypted and secure" (not the same as scanning for malware within them).
6. "Walk me through a ransomware recovery scenario end-to-end, from the moment we declare an incident to when production workloads are back online."
Why: Tabletop exercises are valuable; vendor-led demo walkthroughs of their own recovery process reveal gaps that polished documentation hides. Forces the vendor to be specific about what happens at each step, who does what, and what manual processes remain.
Good answer: Step-by-step walkthrough with honest identification of manual steps, time estimates, prerequisites. Red flag: High-level "our platform handles the recovery" with no procedural specifics.
7. "How is your backup management plane protected, credentials, access controls, network isolation, MFA? If an attacker compromises it, what can they do?"
Why: The backup management console is a high-value target. A sophisticated attacker who gains access to it can delete backup schedules, destroy retention policies, or trigger premature expiry of backup sets, silently, weeks before the ransomware detonates. The management plane is as critical to protect as the backups themselves.
Good answer: MFA on all console access, network isolation from production, role-based access, audit logging, alerting on destructive operations. Red flag: Management console accessible from the production network with standard AD credentials.
8. "How do you handle backup of cloud-native workloads, SaaS applications, containers, serverless, cloud databases, that traditional agents don't cover?"
Why: Most organizations have significant data in SaaS applications (M365, Salesforce, Google Workspace) that is not covered by traditional backup tools. Microsoft's shared responsibility model explicitly places Teams, SharePoint, and Exchange data protection in the customer's scope. Many organizations assume the cloud provider handles it. They don't.
Good answer: Native SaaS backup connectors (M365, Google Workspace, Salesforce), API-based backup for cloud databases and containers, honest about what falls outside scope. Red flag: "Your SaaS data is protected by the provider's replication" (true for availability, not for accidental deletion or ransomware).
9. "What's your cross-platform and cross-cloud story, can I restore an AWS workload to Azure, or vice versa, if I need to?"
Why: A DR scenario that requires you to stay on the same cloud provider that's having an outage or been compromised is a limited DR strategy. Cross-cloud recovery capability is increasingly relevant for cloud-heavy organizations.
Good answer: Documented cross-cloud recovery paths, tested cases, honest about performance and format limitations. Red flag: "We optimize for [single cloud] recovery."
10. "How often do you test recovery, and will you test our recovery, not just your own platform demo?"
Why: Untested backups are not backups, they are hopes. The industry standard is annual DR tests at minimum; best practice is quarterly tabletops and at least annual full recovery tests. Vendors who test their platform in isolation are not testing your specific environment, which is where the surprises live.
Good answer: Facilitated recovery testing, customer-specific test runbooks, DR test reporting, anomaly detection when backups haven't been validated. Red flag: "Recovery testing is your responsibility."
11. "What encryption and key management do you use, for backups at rest and in transit, and who holds the keys?"
Why: Backup sets contain a comprehensive picture of your environment. Encryption of backups matters, but so does key management. If the vendor holds your encryption keys and their key management system is compromised, or they go bankrupt, or they get acquired, your backups may become inaccessible.
Good answer: BYOK (Bring Your Own Key) option, hardware-backed key storage, clear key custody documentation. Red flag: "We encrypt everything" with no detail on key management or custody.
12. "What's your pricing model as data volume grows, and what happens to cost when recovery requires pulling large datasets from cloud storage?"
Why: Backup and recovery has two hidden cost traps: egress fees (the cost of pulling data out of cloud storage during recovery, which can be enormous during a large-scale incident) and capacity scaling costs (backup data grows significantly in ransomware scenarios due to encryption-in-progress creating many changed blocks).
Good answer: Transparent capacity and egress pricing, recovery cost modeling, flat-rate or capped options, experience with large-scale recovery billing. Red flag: "Pricing depends on usage" with no recovery cost scenario modeled.
13. "How long do you retain backup data, and what are the regulatory implications of recovery from older backup sets?"
Why: Compliance frameworks (HIPAA, PCI DSS, financial regulators) have data retention requirements that interact with backup strategy. Recovering from a 6-month-old backup to restore a database may re-introduce data you were legally required to delete. These intersections are real and often unaddressed.
Good answer: Configurable retention with legal hold capabilities, documented interaction with common regulatory frameworks, consultation on retention policy design. Red flag: "We retain backups per your configured policy" (true but doesn't address the regulatory intersection).
14. "If we experience a catastrophic event that makes your platform unavailable, vendor outage, licensing failure, acquisition, what's our recovery path without you?"
Why: Backup data in a proprietary format, accessible only through the vendor's platform, creates a dangerous dependency. If the vendor has an outage during your incident, or their licensing server can't be reached, or they've been acquired and the product is being sunset, can you still access your backups?
Good answer: Open or documented backup formats, offline recovery tools, local catalog copies, contractual data access rights. Red flag: "Our platform is always available" (that's not the question).
The Meta-Test
- Q1, Q2, and Q7 are the ransomware hardening test. Immutability, 3-2-1-1 architecture, and management plane protection are the three controls that specifically defend against sophisticated ransomware actors targeting backup infrastructure. Weak answers here mean your last line of defense has a door in it.
- Q3 and Q10 are the "have you actually tried this" test. RTOs that haven't been tested are fiction. Recovery tests that don't include your environment are dress rehearsals for the wrong play.
- Q5 is the underrated question. Restoring malware-contaminated backups is one of the most common ways ransomware incidents become repeat ransomware incidents.
Honest Observations
1. Most organizations have backup. Very few have recovery.
Having backup software installed and schedules running is table stakes. The test is whether you can actually recover, at the speed you need, with the data integrity you expect, in the environment you'll be working in during a crisis (likely degraded, understaffed, and under external pressure).
The only way to know this is to test it. Annual full recovery tests are the minimum; quarterly tabletops at minimum for IR planning. Organizations that haven't successfully completed a full recovery test in the last 12 months have backup, not recovery.
2. Ransomware has made the air-gap non-negotiable.
If all your backup copies are reachable from your network, even with credentials, a sophisticated attacker can reach them too. The 3-2-1-1 rule exists specifically for this: the fourth copy must be genuinely offline or isolated, not just "in the cloud."
This doesn't mean tape (though tape has made an ironic comeback precisely because it's offline by nature). It can mean an immutable cloud vault with object lock compliance mode, a physically isolated network segment for backup infrastructure, or a managed offline vault service. The mechanism matters less than the genuine isolation.
3. SaaS data is the backup blind spot in almost every organization.
M365, Google Workspace, Salesforce, ServiceNow, Workday, most organizations have significant business-critical data in SaaS applications with no independent backup. Microsoft's own documentation is explicit that Teams, SharePoint, and Exchange data protection is a customer responsibility. The provider's geo-redundant replication protects against hardware failures; it does not protect against accidental deletion, malicious deletion, or ransomware that encrypts cloud-synced files.
If you haven't independently verified what SaaS data you have and whether it's backed up by something other than the provider's native retention, this is worth doing before buying anything else.
4. The management console is a crown jewel, treat it like one.
Sophisticated ransomware groups have become skilled at identifying and targeting backup infrastructure specifically. The playbook: gain initial access, spend weeks identifying backup systems and schedules, delete or corrupt backup sets silently, then trigger encryption. By the time the encryption fires, the victims discover their last clean backup is weeks old.
The countermeasure: isolate the backup management plane from the production network, require MFA and separate credentials for backup administration, alert on any backup schedule modification or deletion, and maintain an out-of-band audit log the primary admin environment can't reach.
5. Recovery is a business function, not just a technology one.
The technology is necessary but not sufficient. Successful recovery from a major incident requires:
- A tested IR plan that covers who does what and in what order
- Pre-negotiated access to IR retainer resources who know your environment
- Executive decision-making clarity on when to restore vs. rebuild
- Communications plans for employees, customers, regulators, and press
- Regulatory notification timelines understood before they're needed
- Legal counsel pre-briefed on ransomware payment decisions and disclosure obligations
Organizations that treat backup as a storage problem and recovery as a technical problem find out the hard way that the hard parts are organizational. The technology just buys you the time and the options, your people and process determine whether you use them well.
6. For "regular" mid-market companies, the practical checklist before buying anything new:
- Audit what you actually have, backup schedules, retention policies, what's covered, what's not (SaaS, cloud workloads, OT).
- Test what you have, if you haven't done a full recovery test in 12 months, do that before buying new tools.
- Verify immutability, do your existing backups have a genuinely isolated, immutable copy?
- Assess the management plane, who has access to the backup console, from where, with what credentials?
- Then evaluate platforms, you'll buy much more intelligently once you know your actual gaps.