Ссылка
click to show
click to show
Amazon releases information about cause of major AWS us-east-1 outage
Summary
A bunch of websites and apps across the world were impacted on Monday when Amazon had a major outage in their us-east-1 region, with services degraded for over 14 hours. A concise chain of events that led to the outage was:
Amazon manages the DNS entries for a bunch of their services using automation, because there are far too many entries, and they change too frequently, to manage manually. Those DNS entries are what allow users to connect to those services.
That automated tooling contained a race condition, which had not been encountered before and which only happens when some services have fallen way behind
That race condition resulted in the main DNS entries for DynamoDB, a popular AWS-managed database (and one that is used internally by a bunch of AWS services), to be deleted in the us-east-1 region
That meant that nobody could connect to DynamoDB in us-east-1, which caused a bunch of websites etc to go down
A bunch of other AWS services and features, including the functionality that is used to provision new EC2 instances, depend on DynamoDB, so they also stopped working
The DNS issue was resolved in just under 3 hours, and DynamoDB was restored
Now, a bunch of the internal AWS services started working again, but had huge amounts of work to do, causing them to fail or have massive backlogs
This resulted in various functionality, including EC2 instance launching, being degraded for several more hours while AWS engineers worked to manage the backlog in the various services
Although this outage only impacted a single region of a single cloud provider, it was the most popular region of the most popular cloud provider, so a bunch of big websites were down or degraded as a result (a small sample includes eight sleep, snapchat, slack, and fortnite).
Quotes
Quote
Right before this event started, one DNS Enactor experienced unusually high delays needing to retry its update on several of the DNS endpoints. As it was slowly working through the endpoints, several other things were also happening. First, the DNS Planner continued to run and produced many newer generations of plans. Second, one of the other DNS Enactors then began applying one of the newer plans and rapidly progressed through all of the endpoints. The timing of these events triggered the latent race condition. When the second Enactor (applying the newest plan) completed its endpoint updates, it then invoked the plan clean-up process, which identifies plans that are significantly older than the one it just applied and deletes them. At the same time that this clean-up process was invoked, the first Enactor (which had been unusually delayed) applied its much older plan to the regional