Feature
The July AWS outage was short. Website recovery was not.
What failed in AWS us-west-2 on July 24, 2026, why downstream recovery took longer, and what website operators should check now.
Impetuous · · 4 Min Read

On July 24, 2026, AWS had a connectivity incident in us-west-2 (Oregon). It was not a global AWS shutdown, nor did every workload in the region stop running. The failure affected network routing between us-west-2 and the Seattle metro area. AWS said connectivity within the region was unaffected, even as some connections into the region timed out.
That creates a confusing failure mode for website operators: compute and storage can remain healthy while readers, external monitors, operators or third-party integrations cannot reach them.
For most affected connections, the initial interruption lasted about 20 minutes. Intermittent connectivity returned briefly during route reconvergence, and one affected AWS Direct Connect path was not restored until 12:12 UTC. Some downstream providers kept incidents open for hours afterward while clients reconnected and queued work drained.
The useful lesson is not that every affected website was down for 10 hours. It is that an upstream status page turning green does not establish that a website has recovered.
What happened on July 24
AWS’s final update gives this timeline:
| Time (UTC) | Event |
|---|---|
| 10:55 | Connectivity impact begins between us-west-2 and the Seattle metro area. |
| 11:01 | AWS engineers are automatically engaged. |
| 11:15 | Mitigation produces initial recovery. |
| 11:40 | AWS posts its first public update. |
| 11:47–11:59 | Route reconvergence causes intermittent connectivity for some customers. |
| 11:59 | All general routes are restored and service metrics return to prior levels. |
| 12:12 | The affected Direct Connect path through EqSe2 in Seattle is restored. |
| 13:01 | AWS publishes its closing update. |
AWS attributed the incident to networking devices responsible for routing from the region to the Seattle metro area. It listed ten affected services, including EC2, Elastic Load Balancing, API Gateway, Global Accelerator, VPC and Direct Connect. The AWS Health Dashboard contains the complete incident record.
This distinction matters when diagnosing a website. An internal health check can pass while an external reader receives a timeout because the two checks traverse different network paths.
Why the “10-hour outage” needs qualification
The AWS network fault and the downstream incidents did not share one duration.
IncidentHub identified nine incidents across seven providers whose updates connected their problems to AWS. The last provider in that set marked its incident resolved at 21:28 UTC—9 hours and 16 minutes after AWS restored the affected Direct Connect path. However, a status-page closure time does not prove that all customers were unable to use the service until that moment. IncidentHub’s reconstruction distinguishes AWS’s recovery from downstream status-page timelines.
The reported downstream behavior shows why recovery can take longer than the original fault. NinjaOne described clients reconnecting gradually through a randomized backoff path while its device-status processing absorbed the returning load. SparkPost reported that it had resumed delivery but still needed to drain mail queued during the interruption. Bex’s account explains the resulting recovery tail.
Common mechanisms include:
- many clients reconnecting when a route returns;
- requests, jobs or messages accumulated during the interruption;
- synchronized retries adding load to a recovering service; and
- health checks restoring traffic before every dependent component has stabilized.
AWS recommends exponential backoff, jitter and a maximum retry count. It also warns against retries at multiple application layers, which can compound into a retry storm, and against retrying non-idempotent operations without safeguards. The AWS Well-Architected reliability guidance provides the implementation detail.
What to check if your website was down
Build the incident timeline from your own measurements rather than inferring your impact from AWS’s timeline.
- Check from outside AWS. Use probes on at least two independent networks. Record DNS resolution, connection and TLS timing, HTTP status, time to first byte and a small content assertion.
- Compare edge and origin signals. A healthy origin alongside external timeouts points toward the network path, load balancer, CDN or DNS layer rather than an application-process failure.
- Inspect dependencies separately. Authentication, forms, email, consent systems, ad delivery, APIs and webhooks can fail while article pages still load.
- Measure recovery, not only reachability. Watch error rate, latency, queue depth, oldest-message age, retry volume and completion of critical actions. Keep the incident open until those measures return to an agreed baseline.
- Preserve evidence in UTC. Save monitor output, relevant logs, provider updates, configuration changes and manual interventions so the review can reconstruct one timeline.
Changes worth making
Start with a dependency register, not an automatic multi-region rebuild. For each infrastructure and SaaS dependency, record:
- its region or other known failure domain;
- the user-facing failure mode;
- timeout and retry ownership;
- the available failover path; and
- the metric that proves recovery.
Then test the controls that limit the recovery tail:
- explicit connection and request timeouts on remote calls;
- bounded retries for transient failures and safe operations, at one chosen layer;
- exponential backoff with jitter;
- queue limits, dead-letter handling and controlled drain rates;
- graceful degradation for nonessential integrations where the product permits it;
- a defined, tested failover trigger; and
- monitoring and incident communications outside the failure domain they observe.
For a mostly static publication, reducing the dynamic serving path may be more useful than duplicating every component. A private object-store origin behind a CDN, immutable releases and a tested rollback reduce the systems required to serve an article. The S3 static-site production setup covers that architecture, while the static website migration checklist provides a cutover and rollback structure.
Multi-Availability Zone deployment helps with zone-level failures, but it does not remove a shared regional network-path dependency. Multi-region delivery can reduce that exposure only if content replication, traffic routing, certificates, secrets, operational access and critical third-party dependencies also work during failover.
Use two recovery markers in the incident record: upstream restored and website recovered. Close the incident on the second marker, using measurements from the path readers actually use.