Many multi-region cloud architectures that appear resilient on paper collapse during actual failures because they share hidden control-plane dependencies, according to a new technical report published on InfoQ by Alexey Golev. The report examines why systems designed for high availability frequently can't recover from real-world incidents, drawing on engineering case studies where redundant infrastructure failed to protect against outages. High availability addresses expected failures within design assumptions, while resilience involves recovering from conditions the system was never explicitly built to handle.

The report documents a production incident where a team upgraded public ingress load balancers to TLS 1.3 for compliance, which triggered a cascading failure that took roughly forty minutes to isolate. Route 53 HTTPS health checks require endpoints to support TLS 1.2, so when TLS 1.2 was disabled, the health checker couldn't complete handshakes and marked endpoints unhealthy. The CDN then stopped routing traffic to an entire region, redirecting users to a failover region thousands of kilometers away and causing latency spikes in affected areas. Inside the impacted region, services remained operational, load balancers showed healthy status, and dashboards displayed nothing unusual—the only visible symptom was that traffic had stopped arriving. The failure occurred in the control plane that decides where traffic flows, not in the application data plane, and was caught only by an engineer manually monitoring metrics during rollout.

According to the report, recovery ownership remains ambiguous in most organizations even when uptime ownership is clearly defined through on-call rotations and SLOs. Golev writes that "recovery paths decay because they are rarely exercised," and the first thirty minutes of major incidents are often lost to what he calls "dependency archaeology"—discovering that permissions have drifted, runbooks reference retired tooling, and standby environments are poorly understood. The report identifies a pattern where architecture diagrams suggest resilience exists, standby regions are provisioned, and runbooks are written, but actual recovery capability is assumed rather than demonstrated. Testing scenarios that genuinely matter requires dedicated engineering time, operational risk, and infrastructure maintained solely for events that may never occur, pushing organizations toward what the report terms "performative resilience."

The report explains that many ostensibly multi-region architectures are only multi-region in the data plane, while their recovery assumptions collapse onto a small number of shared control-plane dependencies that rarely appear on architecture diagrams until something breaks them. AWS STS is cited as a clear example—the global endpoint at sts.amazonaws.com is hosted in a single AWS region (US East N. Virginia) and doesn't provide automatic failover to endpoints in other regions, as AWS documentation explicitly states. When organizations build failover on top of the control plane through mechanisms like DNS-based failover via Route 53 health checks, they inherit those failure modes whether intended or not. The report argues that data planes are designed to keep working when control planes are impaired—an EC2 instance continues running even if the EC2 control plane degrades—but failover only works if the system that decides to fail over is itself functioning.

The report recommends explicit recovery ownership separate from on-call responsibilities, with recurring failover drills treated as an engineering commitment rather than an occasional exercise. Teams should start small with one service and one availability zone shift, as every exercise typically finds something broken—a revoked permission, a stale runbook, or a silent alarm. For organizations with sub-minute recovery objectives, the report suggests AWS Application Recovery Controller (ARC) for customer-facing critical paths like payment flows and authentication, where control-plane dependency is unacceptable, while DNS-based failover remains adequate for internal tools and low-traffic APIs. Golev concludes that in sufficiently complex systems, resilience is probabilistic rather than provable, so the realistic target is improving confidence by reducing unknowns, rehearsing coordination, and shortening the gap between failure and detection. The organizational challenge isn't just architectural—it's that recovery procedures never exercised are the ones that fail first, and that gap becomes visible only when the recovery path must actually be used. The strategic tension for engineering leaders is deciding how much to invest in proving recovery capability when the return materializes only during rare catastrophic events, balanced against the reputational and financial consequences of discovering broken failover paths in production.