How Often Should You Test DR? Cadence, DORA, and Reality

How Often Should You Test DR? Cadence, DORA, and Reality

Ask ten IT directors how often they test disaster recovery and eight will say “annually” — because that is what the auditor asks for, not because anyone believes one test a year proves anything. In 2026 that answer is failing on two fronts at once. Regulators have put numbers on the requirement: DORA, now firmly in its enforcement phase for EU financial entities, mandates resilience testing of every system supporting a critical or important function at least yearly, with threat-led penetration testing every three years for significant firms. And operational reality keeps embarrassing the annual test — survey after survey finds that most organizations need six hours or more to restore critical workloads, against plans that promised one or two.

This brief gives you a defensible cadence by workload tier, explains why the annual compliance-style test produces exactly the six-hour surprise it is supposed to prevent, and walks through the non-disruptive testing tooling — Veeam, Zerto, Commvault, and DRaaS providers like Expedient — that has removed the last honest excuse for not testing quarterly.

What changed

Two things. First, regulation grew teeth. DORA’s testing programme (Article 24) requires EU financial entities to test all ICT systems supporting critical or important functions at least once a year — not the DR plan as a document, the systems themselves. Article 26 layers threat-led penetration testing on top, at least every three years, run against live production systems, with competent authorities empowered to shorten that interval based on risk profile. Supervisors began asking for testing evidence in 2025; in 2026 they are issuing findings. If you sell into or operate in EU financial services, “we have a DR plan” is no longer an answer to anything.

Second, the testing itself got cheap. A decade ago a full DR test meant a weekend, a change freeze, and a conference bridge full of tired people. Today Veeam boots backups in an isolated lab on a nightly schedule, Zerto runs a failover test against tier-1 workloads without touching production replication, and Commvault spins up an on-demand cleanroom in the cloud. The cost argument for annual-only testing died quietly, and most organizations haven’t noticed.

The takeaway for the meeting: regulators now mandate the floor, and tooling has collapsed the cost of exceeding it. Annual-only testing is a choice, and an increasingly hard one to defend.

The cadence answer, by tier

Cadence should follow workload tier, not the audit calendar. Here is the benchmark we hold clients to — the same tiering logic behind our RTO/RPO benchmarks for 2026.

Automated validationNon-disruptive failover testFull-scale exercise
Tier 1 (revenue-critical)Daily/nightlyQuarterlyAnnually
Tier 2 (important, hours of tolerance)WeeklyTwice a yearEvery 12–18 months
Tier 3 (deferrable)MonthlyAnnuallySampled in the annual exercise
DORA-regulated critical functionsContinuous where feasibleAt least annually (Art. 24 floor)Annually, plus TLPT every 3 years (Art. 26)

Diagnostics: if your tier-1 systems have never had a failover test outside the annual exercise, you are running on faith, not a recovery capability. If your automated backup validation is “the job completed successfully,” you have no validation at all — a completed job proves the backup wrote, not that it restores. And if the only person who has ever executed the runbook is the person who wrote it, your real single point of failure is a human.

Why compliance-style tests lie

The six-hour restore reality exists because organizations test to pass, not to fail. The annual exercise is announced months ahead. The scope is trimmed to what is known to work. The most experienced engineer drives. Dependencies that would complicate the result — DNS, authentication, the third-party API nobody owns — are declared out of scope. The test passes, the attestation is filed, and the actual recovery capability remains unmeasured.

Then a real event hits at 2 a.m. on a Saturday. The senior engineer is on a plane. Active Directory comes up after the applications that depend on it. The restore that took 90 minutes in the test takes six hours because the test restored one application and the incident requires forty, all contending for the same backup infrastructure. None of this is bad luck. It is the predictable output of a test designed to produce a green checkmark.

Failure-style testing inverts the design: unannounced windows, second-string operators, deliberately broken dependencies, restore-at-scale rather than restore-one-VM. It produces uglier reports and dramatically better recoveries. The verdict is simple: a DR test that cannot fail is not a test.

The tooling that removes the excuse

Veeam owns the automated-validation layer for most shops. SureBackup boots actual backups inside an isolated DataLab, confirms the OS starts, services initialize, and application checks pass — on a schedule, with no production impact. Paired with Recovery Orchestrator for documented, repeatable failover plans, it turns “did the backup work” into a nightly answered question. The honest limits: SureBackup proves individual machines recover, not that a forty-application business service fails over in order, and the deepest automation still assumes a largely virtualized estate. It fits organizations that want continuous restore assurance from the backup platform they already run.

Zerto (now under HPE) approaches from the replication side. Its journal-based continuous data protection captures every write, and its failover tests run against replica infrastructure without interrupting ongoing replication — a test measured in minutes, safe enough to run monthly or quarterly without a change board fight. Where it is strong: tier-1, low-RPO workloads where you need to prove seconds-of-data-loss recovery on demand. Where it is not: it is a replication product priced per protected workload, not a backup platform — most shops deploy it only for the top tier and cover the rest with something cheaper.

Commvault aimed its testing story at the cyber scenario. Cleanroom Recovery stands up an isolated, on-demand cloud environment where you rehearse recovering from a compromised state — validating that data is clean and applications actually function before anything touches production. That maps directly onto the post-ransomware question boards now ask, and it complements an isolated vault strategy — see our cyber recovery vault comparison for how the vault side stacks up. The trade-offs: rehearsals consume cloud compute you pay for, the experience is strongest in Azure, and it presumes Commvault is already your data protection standard.

Expedient represents the DRaaS answer: put recovery in a provider’s hands and make test failovers a contractual deliverable, with provider engineers assisting the exercise. For mid-market teams without a dedicated DR staff, a contracted, provider-assisted test is often the difference between quarterly testing happening and not. The limits are the DRaaS limits generally — a largely North American footprint (less directly useful to EU DORA entities), and you are testing the provider’s runbook as much as your own, which cuts both ways. Contract structure and test-failover inclusions vary widely across providers; our DRaaS pricing breakdown covers what to demand in the SLA.

What to do about it

  • Publish a 12-month test calendar with named owners: nightly automated validation, quarterly tier-1 failover tests, one annual full-scale exercise with an unannounced component.
  • Rotate operators. If the same engineer runs every test, you are testing the engineer, not the capability.
  • Score tests on time-to-service-restored against your stated RTO, not on “completed / not completed.” Track the trend quarterly.
  • If you are DORA-regulated, map every critical or important function to a test artifact with a date on it. Supervisors ask for evidence, not intentions.
  • Once a year, test at scale — dozens of workloads contending for restore bandwidth — because that is the shape of a real event.

Frequently asked questions

How often should disaster recovery be tested?

Tier 1 workloads: quarterly non-disruptive failover tests plus nightly automated backup validation, with one full-scale annual exercise. Tier 2: twice a year. Annual-only testing is a compliance floor, not a practice.

What does DORA require for resilience testing?

DORA’s Article 24 programme requires EU financial entities to test ICT systems supporting critical or important functions at least yearly. Article 26 adds threat-led penetration testing at least every three years for entities beyond the simplified regime, performed on live production systems.

What is a non-disruptive DR test?

A test that exercises recovery — booting backups in an isolated lab, failing over to replicas in a test bubble — without touching production. Veeam SureBackup, Zerto failover tests, and Commvault Cleanroom Recovery are the common implementations.

What should a DR test plan include?

Scope by tier, named operators and alternates, success criteria expressed as time-to-service-restored versus RTO, dependency order (identity and DNS first), an unannounced element, and a findings log with remediation owners and dates.

Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.