In today’s digital-first world, cloud infrastructure is the backbone of business operations.
But as recent events have shown, relying on a single cloud provider, even ones as dominant as AWS or Azure, can expose organisations to significant risk.
The recent AWS and Azure outages are a stark reminder: even well-architected applications can stall when the control plane fails.
AWS and Azure Cloud Outages: A Wake-Up Call
On October 20th, 2025, AWS experienced a major disruption originating in its US-EAST-1 region. What began as a DNS resolution issue quickly cascaded into a broader failure, affecting services like DynamoDB and EC2 API endpoints. Despite being deployed across multiple Availability Zones (AZs), many applications became inaccessible, not because their data was lost, but because the control systems required to manage and scale them were down.
Just nine days later, on October 29th, Microsoft Azure suffered a global outage triggered by an inadvertent configuration change to its Azure Front Door (AFD) service. This incident disrupted access to critical services including Microsoft 365, Xbox Live, Minecraft, and even airline and retail websites such as Alaska Airlines, Heathrow Airport, and Costco. The outage lasted over eight hours and affected authentication, DNS resolution, and routing across multiple Azure regions. [engadget.com], [independent.co.uk]
These back-to-back incidents reveal a painful truth: multi-AZ or multi-region architecture alone is not enough. When the control plane fails, whether due to DNS issues, configuration errors, or routing failures, even geographically distributed workloads can become unusable. The resilience of cloud-native applications must extend beyond infrastructure redundancy to include robust failover strategies, decentralised control mechanisms, and real-time observability.
Why Hybrid Cloud Is the Strategic Answer
To effectively mitigate systemic risk, organisations must move beyond the limitations of a single cloud provider. A hybrid cloud architecture, whether it spans multiple hyperscalers or blends public cloud with private infrastructure, offers a more resilient and flexible foundation for modern digital operations.
This approach helps eliminate the vulnerability of relying on one vendor’s control plane, which can become a single point of failure during widespread outages. It also enables compliance with data sovereignty regulations by allowing sensitive data to remain within specific geographic boundaries. For many businesses, hybrid cloud provides a practical bridge for legacy systems that are too complex or costly to migrate entirely to the cloud.
But hybrid cloud is not merely a redundancy strategy. It represents a shift toward autonomy, giving organisations the freedom to architect systems that are designed to survive disruption, not just maintain uptime. By embracing this model, businesses gain the ability to adapt, recover, and continue operating even when their primary cloud provider falters.
Resilience Is a Design Principle, Not a Feature
True resilience isn’t something that can be purchased off the shelf, it must be intentionally built into the very foundation of your architecture. It begins with designing systems that prioritise portability and independence, ensuring that workloads aren’t tightly bound to any single cloud provider’s proprietary services. This autonomy allows organisations to adapt quickly when failure inevitably occurs.
Automation plays a critical role in this strategy. Recovery must be fast, consistent, and free from human error, which means relying on orchestrated processes that can restore services in seconds, not hours. Infrastructure-as-Code (IaC) tools like Terraform or Pulumi become essential here, enabling teams to replicate environments across clouds with precision and speed.
The goal isn’t to abandon hyperscalers altogether, but to design around their limitations. By embracing hybrid cloud principles, organisations gain control over their infrastructure, reduce systemic risk, and build the agility needed to survive in an unpredictable digital landscape.
Resilience as a Business Driver
Resilience in cloud architecture isn’t just a technical safeguard, it’s a strategic enabler of business performance. As organisations increasingly digitise their operations, the ability to withstand and recover from disruptions directly impacts financial outcomes, customer trust, and long-term competitiveness.
According to McKinsey, cloud platforms have the potential to unlock up to $3 trillion in EBITDA uplift globally, but much of this value hinges on the resilience of mission-critical applications. Yet, fewer than 10% of large enterprises have successfully migrated their tier-1 workloads to the cloud, largely due to concerns about availability and risk. This hesitation means many businesses are missing out on the full financial and operational benefits that resilient cloud infrastructure can offer. [mckinsey.com]
Hybrid and multi-cloud strategies are increasingly seen as essential to capturing this value. A recent whitepaper from Interactive found that 90% of enterprises had adopted hybrid or multi-cloud models by 2024, with resilience cited as a key driver. These architectures allow organisations to adapt quickly to changing conditions, optimise costs, and maintain continuity even during major outages. [interactive.com.au]
The financial stakes are high. Research from Frost & Sullivan shows that the average cost of downtime for large enterprises is $9,000 per minute, with some industries facing losses of up to $5 million per hour. Beyond the immediate financial impact, outages erode customer trust and brand reputation, intangibles that can take years to rebuild. [pages.awscloud.com]
Resilient cloud design also supports regulatory compliance and operational agility. In sectors like finance and healthcare, where data sovereignty and uptime are non-negotiable, hybrid cloud enables organisations to meet strict requirements while still leveraging the scalability of public cloud services.
Ultimately, resilience isn’t just about surviving failure, it’s about enabling growth. By investing in resilient cloud architecture, businesses position themselves to innovate faster, respond to market shifts more effectively, and deliver consistent value to customers, even in the face of disruption.
The Architect’s Role: Mapping the Blast Radius
One of the most complex challenges in building resilient cloud systems is understanding how services depend on one another. Hyperscale cloud providers often publish documentation that is intentionally vague, making it difficult to grasp the full extent of these interdependencies. As a result, predicting how a failure in one component might ripple through the system becomes a guessing game.
This is where independent cloud architects become indispensable. Their expertise allows them to go beyond vendor documentation, using external tools and real-world experience to uncover hidden links between services. They work to decouple mission-critical systems from proprietary APIs, ensuring that those systems can continue to function even if the cloud provider’s control plane becomes unavailable.
A key part of their role is validating disaster recovery plans that are truly independent. For instance, storing recovery scripts on the same infrastructure that might be compromised during an outage is a common oversight. A resilient architecture demands that such plans be accessible from outside the affected environment, ready to be executed without relying on the very systems that have failed.
Conclusion: Architect for Recovery, Not Perfection
But the recent, high-profile outages, within almost a week of each other, affecting both AWS and Azure the top two major hyperscalers have served as a sharp reminder: the promise of 99.999% uptime doesn’t mean zero risk. The future of resilience lies in architecting for recovery, not chasing unattainable uptime metrics.
Hybrid cloud is no longer a luxury for businesses, it’s a necessity. And the organisations that embrace it will be the ones that survive and thrive in an increasingly unpredictable digital landscape.
