Ransomware backup strategy: immutable, tested, recoverable
Backups are the last reliable lifeline against ransomware. They must be immutable, tested regularly and segregated from the production network. Cyber insurers list 'segregated backups' among the five core controls; missing evidence in a claim eliminates cover. A ransomware recovery layer such as [Halcyon Anti-Ransomware](/en/services/halcyon-anti-ransomware) complements backups. It addresses the encryption attempt itself and shortens recovery time.
How this compares to neighbouring topics
This page covers the backup and recovery strategy against ransomware. The operational response during a crisis including forensics and regulatory notification is on Incident response within 72 hours. The framework for roles and notification chains sits on Create incident response plan. Technical prevention and recovery at endpoint level is on Halcyon Anti-Ransomware. Insurance warranties and claim logic are on SOC and cyber insurance. The operational picture is in the pillar SOC as a Service Switzerland.
Why backups so often fail in a real crisis
The most frequent causes of failed recovery are not technical defects but attack vectors against the backup layer itself. Ransomware groups deliberately target backup consoles, domain-joined backup servers and snapshot stores to delete or encrypt them before the main encryption event. A backup reachable with the same domain credentials as production is not a backup during a ransomware case.
- Backup accounts share the same Active Directory domain as production. Attackers with domain admin rights delete backups first.
- Snapshots and replicas live on the same storage system as the production volumes and are encrypted together.
- Restore tests only cover single files, never full systems or application stacks. The real crisis becomes the first full restore.
- The immutability window is shorter than the typical dwell time of attackers (currently 5 to 21 days). Attackers wait the window out.
- Recovery teams lack a documented recovery order. They improvise the sequence: database before application, identity before database, and network before identity.
Reference model: 3-2-1-1-0
The established pattern for ransomware-resilient backups is 3-2-1-1-0. It requires at least three data copies on two different media, including one off-site copy and one immutable or air-gapped copy. Restore tests must produce zero errors. The last two digits are the ransomware-critical part and appear explicitly on insurer proposal forms.
| Digit | Requirement | Ransomware relevance |
|---|---|---|
| 3 | Three copies of the data (production plus two backups). | Redundancy against loss of individual copies. |
| 2 | Two different media types. | Reduces the risk of systemic attacks on a single platform. |
| 1 | One copy off-site (cloud or second location). | Protection against physical destruction and ransomware at the primary site. |
| 1 | One copy immutable (object lock, WORM) or air-gapped. | Direct core protection against encryption and deletion by attackers. |
| 0 | Zero errors in restore tests (documented, at least quarterly). | Evidence of recoverability towards insurer and regulator. |
Getting immutability right
- Use object lock in compliance mode. Administrators can override governance mode, leaving backups vulnerable once attackers compromise backup admin rights.
- Retention window at least 30 days, better 60 to 90, given the dwell time of modern ransomware campaigns.
- Separate identity system for the backup console (separate tenant, separate MFA, no shared privilege inheritance with production).
- Monitoring on configuration changes to immutability policies. Any shortening of the window is an alert signal for the SOC.
- Air-gap option (tape, offline volumes) for the most critical datasets, where compliance requires it (FINMA, critical infrastructure).
Immutability is activated 'going forward' but does not apply to the existing backup sets. In an incident during the following days, exactly the old copies remain vulnerable.
Restore tests that insurers accept
- Restore tests at least quarterly, documented with date, system selection, duration, result, sign-off.
- Rotation principle: across twelve months every business-critical system is covered at least once in a restore test.
- One major recovery drill per year (full stack, several systems, isolated restore zone), integrated with the tabletop exercise.
- Restore zone: isolated network segment in which recovered systems are validated before returning to production. Attacker artefacts are detected here.
- Documented recovery order: identity and DNS first, then databases, then applications, then front ends. Teams practise this order in advance to avoid improvisation.
Backup is not recovery: think in two layers
A pure backup strategy addresses data recoverability while leaving downtime unanswered. Restoring entire environments typically takes days to weeks depending on volume, order and restore zone. A complementary recovery layer addresses encryption at endpoint level. Halcyon Anti-Ransomware detects the encryption attempt itself, aborts it and in many cases restores individual endpoints. The backup restore sequence then only needs to cover the systems lost. Halcyon complements backups by shortening the recovery window. It protects the production copy, which teams need first in a crisis.
Detect and stop the attack (Halcyon, EDR), preserve forensic evidence and isolate affected systems. Then restore from immutable backup and reintegrate systems through the restore zone. None of this works without the pre-documented IR plan.
What insurers check during a claim
- Immutability configuration on the day of the incident (screenshots, configuration state, retention window).
- Restore test history for the last twelve months with date, system, result and sign-off.
- Network segregation of the backup environment from the production domain (identity, network, privilege inheritance).
- MFA and access log for the backup console in the 30 days before the incident.
- Documented recovery order and its exercise history over the last twelve months.
If any of this evidence is missing, cover reduction or denial is likely, regardless of whether the control existed in reality. See SOC and cyber insurance for the warranty logic.
Frequently asked questions
Does cloud backup automatically count as 'off-site'?
Yes, provided the cloud tenant uses a separate identity, does not share domain trust with production and has object lock in compliance mode enabled. A cloud backup reachable with the same domain admin accounts does not meet the requirement.
How long does a realistic full restore take?
A full restore for a few hundred servers and several TB of data takes three to ten days with preparation. This requires a prepared recovery order, restore zone and network segments. Without preparation, recovery takes several times longer. Metrics such as MTTR are covered on [MTTD and MTTR in a SOC](/en/soc/soc-mttd-mttr).
Does Halcyon replace the backup strategy?
Halcyon complements the backup strategy. It addresses encryption itself on the endpoint through behavioural detection, interruption of encryption and recovery of key material for affected files. Backups remain the lifeline for systems that were outside the Halcyon scope or predated deployment, and for deleted or exfiltrated data. The two layers are complementary.
Must immutability cover the entire backup set?
Immutability must cover at least the copies needed to recover business-critical systems. Secondary datasets (long-term compliance archives) can follow their own regime but are usually not the time-critical rescue layer in a ransomware case.
How often should the recovery order be exercised?
Practise annually through a large full-stack restore drill, integrated with the tabletop exercise from the [incident response plan](/en/soc/create-incident-response-plan). Perform partial restores every quarter as well. Trying the order for the first time in a real crisis extends downtime by days.
Related terms
- Ransomware Ransomware is malicious software that encrypts data and demands a ransom for decryption.
- Incident Response Incident Response is the structured process of containing, eradicating, and recovering from a security incident.
- Malware Malware is the umbrella term for malicious software such as ransomware, Trojans or infostealers that damage systems or steal data.
- Playbook A playbook is a predefined procedure describing how a SOC responds to a specific type of security incident.