Every organisation we audit has backups. A green tick in a dashboard, a nightly job, a retention policy written down somewhere. Far fewer have ever restored one.
Those are different states. The first is a belief. Only the second is a capability.
How backups fail quietly
- The job has been erroring for months and the alert goes to an inbox nobody reads
- The database dumps fine but the uploaded files were never included
- Everything is backed up to the same server that fails
- The archive is encrypted and the key lives only on the laptop that died
- The restore works, but takes eleven hours — and you needed it in one
Every one of these passes a dashboard check. Every one is discovered at the worst possible moment.
The drill
Twice a year, block an afternoon and do this properly:
- Pick a real backup at random — not the newest one
- Restore it to a clean machine that is not your production server
- Start the application and sign in
- Open the five records that matter most and check them against what you expect
- Write down how long the whole thing took, from start to working system
The number that matters
That last figure is your real recovery time. Not the one in the contract — the one you measured. Compare it against how long the business can actually be down. If measured recovery is longer than tolerable downtime, you do not have a backup problem, you have an architecture problem, and no amount of extra backup frequency will fix it.
Restore drills are boring, which is why they get skipped. They are also the only thing separating a bad afternoon from a closed business.
- Infrastructure
- Security