Assess, fix, keep it that way

EdgeReadyEdgeResolveEdgeAssure

For investors

EdgeSignalAll services
How it works Find your path Case studies Security About What we take on Book a 20-minute triage call
Resilience

Backups that had never been restored, proven to work — and timed.

AWS Backup was switched on, and nobody had ever tried to restore from it. A backup you have not restored is a belief, not a control — the only way to know the real number is to run the recovery and read the clock.

← Some of our past work · anonymised prior delivery, no client named

Stack

AWS BackupAmazon RDSAmazon EBSAmazon S3CloudWatch alarms

Rough timeline

PhaseTypical duration
Workload criticality, RTO and RPO mapping1 week
Backup policy automation1–2 weeks
Controlled restore tests and runbooks1–2 weeks
The challenge

"We have backups" was the whole answer, and it was not a complete one.

AWS Backup being enabled answers whether data exists somewhere else. It says nothing about how long getting it back would take, or whether the restore actually works.

No stated recovery target

Without an RTO or RPO per workload, "backed up" and "recoverable in time" were being treated as the same claim.

A restore is disruptive to test carelessly

Which is exactly why it had never been tried — testing it safely needed its own plan.

Coverage had never been checked against need

A backup plan set up once, early on, and never revisited against what actually needed protecting since.

What we found

Coverage that varied by workload, and a recovery time nobody had measured.

The gaps were quiet ones — nothing was failing, because nothing had been tested to fail.

Backup coverage was inconsistent

Some workloads had a documented plan; others were covered only by a default nobody had revisited.

Retention did not match the recovery need

A short retention window on a workload where the real requirement was to go back further than that.

Nobody had timed anything

The honest answer to "how long would a restore take" was a guess, because it had never been measured.

What InfraEdge changed

Mapped, automated, restored for real, and the timing written down.

In that order deliberately — a stated target before a policy, and a real restore before a runbook claims it works.

ChangeWhat it did
Critical workloads mappedEach workload assigned a stated RTO and RPO, agreed with the team that owns it rather than set by default.
Backup policies automatedAWS Backup plans and vaults defined in Terraform, with retention and frequency set against the stated RPO for each workload.
Controlled restore tests runA restore performed for real, into an isolated environment, for each workload tier — not a dry run described in a document.
Recovery runbooks writtenStep by step, with the timing the restore actually took recorded alongside it, so the next person works from a measured number.
Alarming on backup healthCloudWatch alarms on failed or missed backup jobs, so a gap is caught the day it happens rather than the day someone needs the backup.
Outcome

From "we have backups" to a stated recovery time, proven once already.

The change that mattered was not the backup policy. It was running the restore for real and writing down what happened.

We have not published a number here because it belongs to that estate and that workload, not to a general claim — the method is the transferable part: map the workload, automate the policy against a stated target, restore it for real, measure it, then write it down.

Book a 20-minute triage call

Twenty minutes, no charge. We work out what would actually help — which is sometimes us and sometimes not. Nothing is priced on the call; if there is work worth doing, a written scope and a price reach you within 24 hours.