Skip to main content
← All posts·
Platform Engineering

Disaster Recovery for Kubernetes in Regulated Enterprises: RTO/RPO Planning Guide

Kubernetes disaster recovery guide for regulated enterprises. RTO/RPO planning, Velero backup strategies, multi-cluster failover, etcd backup and restore, stateful workload recovery, and DORA business continuity requirements for containerised infrastructure.

Luca Berton12 min read

DR for Kubernetes Is Different

Traditional disaster recovery focuses on VMs, databases, and storage. Kubernetes adds complexity: you need to recover not just data, but cluster state (etcd), application configuration (manifests, ConfigMaps, Secrets), persistent volumes, custom resources (CRDs), and the relationships between them. Getting this wrong means your recovery doesn't work when you need it most.

RTO/RPO Planning

Definitions

  • RTO (Recovery Time Objective): Maximum acceptable downtime. How quickly must the service be restored?
  • RPO (Recovery Point Objective): Maximum acceptable data loss. How much data can you afford to lose? (Measured in time: RPO of 1 hour = up to 1 hour of data loss)

DORA requirement: Article 11 requires documented RTOs and RPOs for all critical ICT services, reviewed annually and after significant changes.

TierRTORPOStrategy
Tier 1 (Critical)< 15 minutesNear-zeroActive-active multi-cluster
Tier 2 (Important)< 4 hours< 1 hourActive-passive with async replication
Tier 3 (Standard)< 24 hours< 24 hoursBackup and restore (Velero)
Tier 4 (Non-critical)< 72 hours< 72 hoursRebuild from Git (GitOps)

Kubernetes DR Strategies

Velero — Backup & Restore

  • Cluster state backup: Backs up all Kubernetes resources (Deployments, Services, ConfigMaps, Secrets, CRDs)
  • Persistent volume backup: Snapshots PVs via CSI snapshots or Restic/Kopia for file-level backup
  • Scheduled backups: Cron-based backup schedules with retention policies
  • Selective restore: Restore entire clusters, specific namespaces, or individual resources
  • Cross-cluster migration: Backup from one cluster, restore to another (DR or migration)

Multi-Cluster Failover

  • Active-active: Traffic distributed across clusters. If one fails, the other absorbs all traffic. Requires global load balancing and data replication.
  • Active-passive: Standby cluster kept in sync via GitOps (ArgoCD/Flux) and async data replication. Failover requires DNS switch and data promotion.
  • Pilot light: Minimal standby cluster (control plane + critical configs). Scale up workers on failover.
📘 Book

Kubernetes Recipes

A practical guide for container orchestration and deployment by Grzegorz Stencel & Luca Berton (Apress).

Watch on Skillshare →

What to Back Up

  1. etcd: The cluster's brain. Automated etcd snapshots to external storage (minimum every 30 minutes for critical clusters)
  2. Kubernetes resources: All manifests, ConfigMaps, Secrets, CRDs — Velero handles this
  3. Persistent volumes: Database data, file storage, stateful application data — CSI snapshots + cross-region replication
  4. GitOps repository: Your source of truth. Already backed up in Git, but ensure the Git hosting is also resilient.
  5. External dependencies: DNS records, certificates, secrets backend (Vault), container registry — these must also be recoverable
  6. Test your recovery — At least quarterly. A backup you haven't tested is not a backup. Document recovery time achieved vs RTO target.
disaster recovery
Kubernetes
RTO
RPO
Velero
business continuity
regulated enterprises

Need help applying this in your organization?

Get a free 30-minute assessment with actionable recommendations — whether we work together or not.

Book Your Free AI Platform Assessment

Or see AI readiness assessment scope & pricing

18+ years experience · Ex-Red Hat & Dell · Speaker at KubeCon EU 2026

Luca Berton

Written by

Luca Berton

CEO at Open Empower. 18+ years building enterprise infrastructure at JPMorgan Chase, Red Hat & Dell. Author of 9 technical books. Speaker at Red Hat Summit and KubeCon EU 2026. Instructor on Coursera, Pluralsight & Udemy.

Get more insights like this

Practical AI infrastructure and platform engineering guides — delivered to your inbox.

Subscribe to Newsletter →