Recovery Readiness: A Four-Hour RTO That Was Really Thirty-One
A SaaS provider for clinical trial management (180 pharma and CRO customers) had a documented disaster recovery plan with a four-hour RTO — in their contracts, their SOC 2 report, and every security questionnaire. Nobody could remember a full restore ever being performed. Envion ran a full recovery rehearsal: production region declared lost, rebuild in the secondary region, restore all data. It took 31 hours — and exposed that ~40,000 uploaded trial documents existed in exactly one region and would have been permanently lost. Eleven months and four quarterly rehearsals later: full recovery in 3h 10m, zero unrecoverable data, manual steps down from 60+ to 9, and the contractual 4-hour RTO actually demonstrated.

The challenge
They had a documented disaster recovery plan with a four-hour recovery time objective and a one-hour recovery point objective. It was in their contracts, in their SOC 2 report, and in every security questionnaire they answered. Backups ran nightly and the monitoring said they succeeded.
Their new VP of Engineering had one question when she arrived: had anyone actually done it? Nobody could remember a full restore ever being performed. There had been a partial database restore in 2022 for a customer data issue. That was it.
She brought Envion in because she wanted an outside party to run the test and write down what happened, rather than an internal exercise that could be quietly abandoned if it went badly.
Decision path
Envion planned a full recovery exercise: production region declared lost, rebuild in the secondary region, restore all data, bring the platform to a state where customers could work. The team was asked not to prepare specially — whoever would be on call in a real event should run it, with the runbook they actually had.
It took 31 hours. Against a documented four.
Hours 0–3: nobody could start. The runbook's first step assumed access to a bastion host in the primary region — the region declared lost. The credentials to provision infrastructure in the secondary region were in a password manager that authenticated through an SSO provider whose configuration also lived in the primary region.
Hours 3–9: infrastructure drift. The secondary region was provisioned by Terraform that hadn't been applied in fourteen months, while production had been changed by hand during two incidents. The restored infrastructure came up and didn't match.
Hours 9–20: the backups. The database backup restored successfully — genuinely good news. But uploaded trial documents lived in object storage with cross-region replication configured for one of two buckets; roughly 40,000 files existed in exactly one region. Secrets were in a secrets manager never included in the backup scope at all. And the search index took six hours to rebuild.
Hours 20–28: dependency order. No documented startup sequence; services came up in the wrong order, failed health checks, retried, and hit rate limits on an external identity provider that locked the account for thirty minutes.
Hours 28–31: verification. No one had defined what "recovered" meant. The team eventually agreed on a smoke test they invented during the exercise.
Envion contribution
The RPO finding was worse than the RTO one. The 40,000 unreplicated documents would have been permanently lost in a real regional failure. In clinical trial management, that's not a service interruption — that's a regulatory event and potentially trial data integrity.
Envion did not recommend a rewrite of the plan. The plan wasn't wrong in its intent; it had never been tested, which is a different problem with a different fix.
Immediate, week one: enable replication on the second bucket, add the secrets manager to backup scope, and stand up a break-glass credential path with zero dependency on the primary region — printed and in a safe.
Weeks 2–8: eliminate infrastructure drift by making manual production changes technically impossible rather than discouraged; define and automate the service startup sequence; write a real verification checklist that defines "recovered" in terms of specific user journeys.
Ongoing, the actual recommendation: quarterly recovery exercises, rotating who leads them, with results reported to the board. An untested DR plan degrades continuously, and the only thing that reliably prevents that is someone senior expecting a number every quarter.
Envion also told them to change their contractual RTO until they could hit it — an uncomfortable conversation, and the right call: a four-hour commitment they couldn't meet was a larger liability than a twelve-hour one they could.
Delivery
The rehearsal program ran with the client's own on-call engineers leading, rotating leadership each quarter, with results reported to the board. Envion's role moved from running the first exercise to reviewing the later ones.
By the fourth rehearsal, in month 11, full recovery ran in 3 hours 10 minutes — inside the contractual RTO, which was then kept at four hours because it could finally be demonstrated.
Outcome and evidence
Eleven months on: full recovery time fell from 31 hours to 3h 10m, permanently unrecoverable data from ~40,000 documents to zero, manual recovery steps from 60+ to 9, circular dependencies on the primary region from 3 to 0, and infrastructure drift from fourteen months to enforced zero. Rehearsals went from zero in five years to quarterly.
They now include the most recent rehearsal result in security questionnaires. Their VP says it closes the DR section of enterprise reviews faster than the SOC 2 report does — a measured number beats an attestation.
| Metric | At first rehearsal | Rehearsal 4 |
|---|---|---|
| Full recovery time | 31 hours | 3h 10m |
| Data permanently unrecoverable | ~40,000 documents | 0 |
| Manual steps in recovery | 60+ | 9 |
| Circular dependencies on primary region | 3 | 0 |
| Infrastructure drift between regions | 14 months | 0 (enforced) |
| Rehearsals performed | 0 in 5 years | Quarterly |
| Contractual RTO | 4h (unachievable) | 4h (demonstrated) |
Client feedback
What the client says about this engagement

“We had a four-hour RTO in eleven contracts and it took us thirty-one hours in a controlled exercise where nobody was panicking and nothing was actually on fire. The forty thousand documents sitting in one region is the one I still think about — that's trial data, and it would have been gone.
Alex's framing stuck with me: we didn't have a disaster recovery plan, we had a disaster recovery document. The difference is whether anyone has ever run it.”
From the engagement lead
What I’d tell anyone considering this

“Run the exercise. That's the whole recommendation. Not a tabletop walkthrough where people describe what they'd do — an actual restore into actual infrastructure, led by whoever is genuinely on call, with a stopwatch running.
Two specific things to look for, because they're near-universal. Circular dependencies on the thing you've lost: credentials, runbooks, VPN, SSO, the wiki with the instructions. Trace every step of your recovery and ask what it depends on. If your recovery plan lives in a system hosted in the region you're recovering from, you don't have a recovery plan.
And things added after the backup scope was defined: new buckets, new databases, new secrets stores. Backup configuration is set once and the system keeps growing around it. Reconcile what exists against what's backed up — it's a half-day of work and it's where the permanent data loss hides. If you can't meet your contractual RTO, change the contract before someone else discovers you can't.”
Evidence gate. This page publishes only what Envion's project records and client disclosure permissions support. Outcomes are added once verified against a baseline, a measurement period, and an approved source.
FAQ
Questions about this case
Facing a similar challenge?
When did you last restore from backup, end to end, into empty infrastructure? If the answer is never — your RTO is a guess. Envion runs the rehearsal and writes down what happens.
Discuss a Similar ChallengeKeep exploring
Similar case studies
Executive Technology Leadership
Support for high-stakes product and AI decisions
Bring senior technology leadership into the business when the roadmap is unclear, delivery is at risk, an AI initiative needs stronger ownership, or the company needs an experienced technical voice before hiring a permanent CTO.
Discuss Interim CTO SupportCore responsibilities
- Align product and technology priorities with business goals and measurable outcomes.
- Review architecture, delivery risks, data foundations, security needs, and AI readiness.
- Lead internal teams and external partners through a practical execution plan.
- Clarify team structure, ownership, decision rights, and delivery cadence.
- Support investor, board, partner, and due-diligence conversations with credible technical judgment.
New experience
Prompt-to-Page — try it right here
Describe the landing page you want, in your own words. We turn it into a finished page and email you a private link in 5–10 minutes — no briefs, no calls, $0 to see the result.
- Describe what you want to create.
- We structure, write, and compose the page.
- You receive a private link when it is ready.
Start with a sentence — the interactive builder takes it from there.
Generate My PageSafe, respectful content only. No obligation.



