SRE and DORA: A Natural Alignment
Site Reliability Engineering (SRE) and the Digital Operational Resilience Act (DORA) want the same thing: resilient, reliable, well-monitored systems with systematic incident management. The difference is that SRE is a practice, and DORA is a legal requirement. For regulated financial institutions, SRE practices provide the operational foundation that DORA demands.
SRE Practices Mapped to DORA
SLOs & Error Budgets → DORA Art. 5-16 (ICT Risk Management)
- Service Level Objectives (SLOs): Define measurable reliability targets for each critical service — 99.9% availability, P99 latency < 200ms, error rate < 0.1%
- Service Level Indicators (SLIs): The measurements that feed SLOs — request success rate, latency distribution, throughput
- Error budgets: The allowed amount of unreliability (0.1% for a 99.9% SLO = 43 minutes/month). When the budget is exhausted, prioritise reliability over features.
DORA alignment: SLOs quantify your risk appetite for ICT systems. Error budgets create a governance mechanism — when reliability drops below the target, engineering resources shift to resilience. This is exactly the risk-proportionate approach DORA requires.
Incident Management → DORA Art. 17-23
- On-call rotation: Named responders 24/7 for critical services, with escalation procedures
- Incident severity classification: P1-P4 aligned with DORA's materiality thresholds (clients affected, transactions impacted, duration)
- Incident commander role: Single point of coordination during major incidents
- Communication protocol: Internal stakeholders, management, and (for DORA) competent authorities notified within defined timelines
- Blameless postmortems: Root cause analysis focused on systemic improvements, not individual blame — with documented action items
Chaos Engineering → DORA Art. 24-27 (Resilience Testing)
- Game days: Planned exercises that simulate real failures (dependency outage, data centre loss, DNS failure)
- Chaos experiments: Controlled fault injection in production — pod killing, network latency, resource exhaustion
- Disaster recovery drills: Full DR failover and recovery exercises with measured RTO/RPO
- Load testing: Validate system behaviour under peak and stress conditions
DORA alignment: DORA Article 25 requires testing that includes scenario-based tests, performance testing, and penetration testing. SRE's chaos engineering practices directly satisfy these requirements.
SRE Team Structure for Regulated Enterprises
- Embedded SREs: SREs embedded in product teams for critical services — deep knowledge of the service they support
- Platform SREs: Manage shared infrastructure — Kubernetes clusters, databases, networking, observability
- Incident management team: Coordinate major incidents, maintain runbooks, drive postmortem process
- Ratio: Google's original guidance: 1 SRE per 5-10 developers. Adjust based on service criticality and regulatory requirements.
Kubernetes Recipes
Practical guide for container orchestration and deployment — hands-on patterns you can use today.
View on Amazon →Getting Started
- Define SLOs for critical services — Start with 3-5 most critical services. Define availability and latency SLOs.
- Implement SLO monitoring — Grafana + Prometheus with SLO dashboards and error budget burn-rate alerts
- Establish incident management process — Severity classification, on-call rotation, communication protocol, postmortem template
- Run your first game day — Simulate a realistic failure scenario with controlled conditions
- Map to DORA requirements — Document how each SRE practice satisfies specific DORA articles for audit readiness
Related Solution
Navigating AI adoption in a regulated environment? Our readiness assessment maps infrastructure, governance, and compliance gaps in 3-4 weeks.
Explore AI Readiness for Regulated Enterprises →
Luca Berton