Skip to main content
dario's.blog
Back to posts

How not to communicate a production incident

Last Friday I received an email that was straight out of my worst nightmares. It was an AWS notification about over-budget spending. My budget alert threshold was $9, while the actual amount spent was a whopping $518,615,489. That's almost 519 million dollars.

My first thought was that this was just a spam email and that it's probably nothing, especially because it was coming from my personal AWS account where I don't have many services. Still, I needed to check. After opening the Cost Explorer page in the AWS console, I saw a huge spike in the cost and usage graph:

For years, I have read horror stories about people leaving their AWS services unprotected, exposed for bad actors to take advantage of. Had I become one of those victims? How in the world did I manage to spend that much?

Fortunately for me and for many AWS customers, it turned out to be an issue with AWS's billing calculations, and it took almost a full day to fix. However, I want to do a retrospective on this issue as it was a very good example of how not to communicate with your customers in such a situation.

The misleading AI assistant and the bad joke

The first thing I did was check the actual S3 buckets and whether there was any trace of malicious traffic in them. Everything seemed correct, and there was nothing suspicious. Failing to find anything, I even asked Amazon Q, the AI-based assistant in the console. However, it only led me to falsely believe that someone generated a lot of temporary traffic on my S3 buckets:

Only after opening the support page in the AWS console did I find an almost-hidden notice about a "billing operational issue". Try to find it yourself in the screenshot below; it's not very obvious, is it?

Only then did I search X to see if anyone was having similar issues, and many people were already reporting it. This is when I confirmed it was definitely something on the AWS end.

The question is, why were the affected customers not informed proactively? If you've sent a miscalculated amount via email, then use the same channel to communicate that this number may be incorrect and to ask for patience while the issue is being resolved.

Instead, AWS posted the following notice on X more than 9 hours after the issue started affecting customers:

Plenty of people nearly had a heart attack over that alert, and the best AWS could do was joke about it? Not funny to the people on the receiving end. The brief apology came only on Sunday, far too little and far too late.

Conclusion

Mistakes happen. However, when they do, the best thing you can do for your customers is to communicate proactively, with empathy, and take responsibility for what happened. Especially if you operate at AWS's scale, hosting an estimated 10-15% of the internet's infrastructure. What AWS did on Friday and Saturday was a perfect example of how not to communicate with your customers.