Backups and Ransomware Recovery

Almost every business that loses data to ransomware had backups. The backups existed, ran on schedule and reported success. They just could not be used to recover, for one of a small number of predictable reasons.

Why backups fail at the moment of use

The backup was encrypted too. If the backup destination is a network share or an always-connected drive that the server can write to, ransomware can write to it as well. A backup reachable from the compromised machine is not a backup.

Nobody ever tested a restore. The job reported success for three years. Nobody confirmed the files could actually be brought back, or how long it would take, or whether the person who set it up still works there.

The backup covered the wrong things. The file server was covered. The database, the mail archive, the line-of-business application configuration and the machine that runs the alarm reporting were not, because they were added after the backup was configured.

Recovery was slower than the business could survive. Restoring several terabytes over a domestic broadband connection from cloud storage is technically a recovery and commercially a disaster.

The 3-2-1 rule, and what it is actually for

Three copies of the data, on two different types of media, with one copy off site. It is old advice and it survives because each part addresses a distinct failure.

Three copies covers a single corruption. Two media types covers a technology-wide failure. One off site covers fire, flood and theft.

Ransomware added a fourth requirement that the original rule does not state: at least one copy must be immutable or offline - something the compromised environment cannot alter or delete even with administrator credentials. Immutable cloud storage, or media that is physically disconnected between runs. Without that, all three copies can be encrypted in the same event.

Decide two numbers first

Every sensible backup design starts from two questions the business has to answer, not IT.

How much data can you afford to lose? That sets backup frequency. Nightly backups mean up to a day of work gone.

How long can you be down? That sets the recovery method. An hour is a very different design from a week, and usually a much more expensive one.

Most organisations have never stated either number, which is why their backup design cannot be evaluated - there is nothing to evaluate it against.

Test the restore

This is the whole thing, and it is the step that is skipped.

Pick a date. Restore something real to a separate location. Time it. Open the restored files and confirm they are intact and current. Write down what went wrong, because something will. Do it again in six months, and make sure at least two people know how.

An untested backup is a belief, not a control.

Beyond the files

Recovery needs more than data. If your systems are down, do you have your suppliers’ contact details somewhere other than the system that is down? Your insurance policy number? The phone number for whoever supports the affected system? A short printed sheet solves a problem that is genuinely difficult at three in the morning.

And a decision worth making in advance rather than under pressure: your position on paying a ransom. Payment does not guarantee a working decryption key, may fund further offences, and can carry legal exposure. Deciding that calmly beforehand is better than deciding it in the first hour of an incident.

Related: phishing and business email compromise, which is how attackers most often get in to begin with.