A backup without a restore test is a hope
A backup without a restore test is a hope: it describes how things should have been, not what is actually there. This post is about why I stopped trusting green checkmarks and what one very concrete failure taught me.
For a long time I believed a backup was mainly a technical task: pick a target, set a schedule, see the green checkmark in the log the next morning. Done. These days I see it differently. Copying is only half the work. The other half is restoring what you copied, at a moment when nobody expects it, while keeping a clear head.
The Sunday when everything was in the open
It was a quiet Sunday evening. I only wanted to update a database for a small club website and had mistyped on the command line. Instead of running the dry run, the script ran against the live data. Within seconds a table that had been filled for years was empty. My first thought was relief: at least there are backups. My second thought came three hours later.
What a backup actually is
A backup is, to begin with, nothing more than a claim about the past. It says: at a certain point in time, something looked like this. Whether that claim is true, nobody knows until they have checked it. A log entry proves that a process ran. It does not prove that the data at the destination is complete, readable and consistent. That gap between running and working goes unnoticed as long as nothing happens.
- Only a restore tells you whether the chain from source to target really holds.
- A backup that has never been read is a promise, not evidence.
- The most dangerous state is not no backup but a backup that everyone trusts blindly.
The example: the file that was only half there
An acquaintance runs the till and the invoice archive for a small shop. Backups ran faithfully every night for years, to a remote target, encrypted, with logs. When the disk in the machine died, he was almost relaxed. After the restore, everything came back as well – except the invoices from the last quarter. The files existed and even had plausible sizes, but they opened empty. The cause was not a failed drive but a script that, during a migration, kept copying files for a while without writing their contents properly. Fifty-six nights in a row. Every log was green.
He had half a year of backups that looked like backups and were not. Nobody noticed until he actually needed them. That is the point: a flaw in a backup stays invisible until it becomes expensive.
Why the test is worth more than the job
A running job is reassuring; a passed restore test proves something. So I set myself a simple rule: if I cannot restore a backup within half an hour on a fresh system, I do not have a backup, I have an assumption. That applies to the database, to the configuration, to certificates and to everything people like to call a detail. The test is uncomfortable because it takes time and because it exposes mistakes. But a mistake I find now is an appointment in my calendar. A mistake I find later is an emergency.
How I do it today
My process is unspectacular, and that is deliberate. The host runs Proxmox VE, the data lives in a ZFS pool, encrypted and with snapshots. Next to that there is a second target, physically separated, and a directory with the configurations kept in Git, so I can trace every change. A small cron job checks only one thing each morning: is the latest backup younger than twenty-four hours, and can the archive be opened? If either question fails, I get a message before I need it.
- Monthly: pull one file or one database out of the backup and open it. Not just list it, actually open it.
- Quarterly: a full restore onto a test machine, by hand, with notes.
- The notes are the important part. If I have to guess next time, the test was incomplete.
- Always: keep at least one copy somewhere a fire cannot reach.
What matters is that the test machine is genuinely empty. Restoring onto the same system is not a test, it is self-confirmation. Only when I start from zero do I see what is actually missing – the passphrase nobody wrote down, the certificate that lived only on the old machine, the one setting no one remembers.
What I no longer believe
I no longer believe that a backup without a restore test is a backup. I also no longer believe that “it has been running for years” is an argument – the flaw in my example ran for fifty-six nights, and it would have kept running for six hundred. What I do believe: a system whose way back I have practised once a month forgives more mistakes than one that only knows a green checkmark. Calm in an emergency does not come from the tool, it comes from practice.
A small routine with a large effect
None of this needs heavy infrastructure. A reasonably clean schedule, a second target, an entry in the calendar and the willingness to restore something once a month and actually look at it. It takes less than an hour. It is the hour that makes the difference later, when everything else goes wrong and I have to decide soberly instead of hoping. In the end it is not about how many backups I have, but about whether I know where the way back is.