PaaS + IaaS

This article is more than 1 year old

AWS power failure in US-EAST-1 region killed some hardware and instances

Some servers and EBS volumes won’t make it back

Wed 22 Dec 2021 // 23:58 UTC

A small group of sysadmins have a disaster recovery job on their hands, on top of Log4J fun, thanks to a power outage at Amazon Web Services’ USE1-AZ4 Availability Zone in the US-EAST-1 Region.

The lack of fun kicked off at 04:35AM Pacific Time (PST – aka 12:35 UTC) on December 22nd, when AWS noticed launch failures and networking issues for some instances in its Elastic Compute Cloud IaaS service.

26 minutes later the cloud colossus ‘fessed up to a power outage and recommended moving workloads to other parts of its cloud that were still receiving electricity.

Power was restored at 05:39AM PST and AWS reported slow recovery of services, however a 6:51AM update admitted that ongoing networking issues were hampering efforts at full restoration.

Services including Slack and Asana reported service difficulties as a result of AWS' mess.

At the time of writing, AWS has still not fully restored networking.

And restoration may not be possible for some customers: at the time of writing, the most recent update on AWS’ status page offers the following grim news:

As is often the case with a loss of power, there may be some hardware that is not recoverable, which will prevent us from fully recovering the affected EC2 instances and EBS volumes. We are not quite at that point yet in terms of recovery, but it is unlikely that we will recover all of the small number of remaining EC2 instances and EBS volumes.

That’s the digital equivalent of waking up to a lump of coal on Christmas Day.

The incident is AWS’s third US outage this month: on December 15th the operation’s US-WEST-1 and US-WEST-2 went missing for around 30 minutes. The US-EAST-1 region had a wobble on Dec 7. Technical issues were blamed. US-EAST-1 also browned out for eight hours in September 2021.

AWS advises customers not to rely on a single Availability Zone (AZ). The outfit’s architecture places two or more AZs within a single Region, and each Zone is physically distant from the others so that a single physical infrastructure incident can’t take out the whole Region. Using multiple regions therefore improves resilience – and cost.

Not every user follows AWS’ guidance about using multiple AZs, so when incidents like this strike their servers and data will become unavailable.

US-EAST-1 is AWS’ biggest and oldest region. Cloud economist Corey Quinn rates its importance as follows:

A multi-day full outage of us-east-1 will have an observable effect on the world economy. That is not an exaggeration.
— Corey Quinn (@QuinnyPig) December 8, 2021

AWS offers a service level agreement of 99.95 per cent uptime for compute instances – or just under 22 minutes a month of downtime. If AWS misses that mark, it offers a ten per cent service credit, a sum that grows to thirty per cent if uptime drops below 99 per cent. If uptime falls below 95 per cent, customers are given 100 per cent of their fees as credits.

AWS also automatically waives fees if EC2 Instance are unavailable for more than six minutes inside a single hour.

Good luck if you’re one of the AWS customers faced with the need for a sudden rebuild. ®

Topics

Special Features

Vendor Voice

Resources

PaaS + IaaS

AWS power failure in US-EAST-1 region killed some hardware and instances

Some servers and EBS volumes won’t make it back

More about

More about

Narrower topics

Broader topics

More about

More about

More about

Narrower topics

Broader topics

TIP US OFF

Other stories you might like

US-EAST-1 region is not the cloudy crock it's made out to be, claims AWS EC2 boss

AWS must pay $525M to cloud storage patent holder, says jury

Irish power crunch could be prompting AWS to ration compute resources

Getting on board with AI

911 goes MIA across multiple US states, cause unclear

Snowmobile, Amazon's truck-powered migration service, reaches the end of the road

Tencent Cloud to revisit design after circular dependencies slowed emergency API fix

Backblaze cloud storage buzzes with added Event Notifications

Misconfigured cloud server leaked clues of North Korean animation scam

Sacramento airport goes no-fly after AT&T internet cable snipped

Oracle scores big win with Fujitsu Japan for its Alloy partner cloud

Alleged cryptojacker accused of stealing $3.5M from cloud to mine under $1M in crypto

About Us

Our Websites

Your Privacy