From an email I just sent:
Hi everyone,
As a quick follow-up to my earlier message (which many of you probably didn't see), the global AWS outage has officially been resolved. In other words, the internet is no longer broken (I hope)...
I also wanted to send a quick postmortem on what happened, for those of you who missed the chaos.
This morning around 4:30 AM EST, I woke up to alarms and alerts firing off everywhere and some tickets (sleep is a luxury these days... anyone got tips?). At first, it looked like a widespread outage was hitting parts of the internet. I sent out a frantic alert around 5:30 AM EST, but ironically, that message didn’t get delivered for hours.
Why? Because AWS was the culprit. And as it turns out, when AWS has an issue, the entire internet breaks, including the SMTP relay we use to send emails.
From there, things got weird. I couldn’t update our status page. Our ticketing system started timing out. TrustedForm went down. Zapier connections stalled. Even my ring wouldn’t stop blaring because the Ring app was down, so I couldn’t even turn off my own alarm. I had to throw it into the washing machine so it wouldn't wake up my kids.
Despite the chaos, I stayed at my desk sending out individual messages to everyone I could think of, especially those running active ad campaigns. They probably didn't even get through. I knew how critical every lead was during those early hours.
It felt like the internet was collapsing around us. And to be honest, this was the first time in a long time I truly felt helpless. The last AWS outage this bad that I remember was back in 2016 or 2017... and back then I was just an employee enjoying a surprise day off. This time? Very different story.
What We Know Now
✅ The LeadCapture app, forms, and funnels stayed up during the entire outage, which was a huge win. It speaks to the infrastructure that we have in place now.
⚠️ The issues were mostly upstream and downstream, integrations, CRMs, OTP, and lead distribution systems like TrustedForm and Zapier were affected.
⚠️ Some data like TrustedForm certs weren’t captured. Some leads may have experienced processing delays.
⚠️ We saw elevated CPU alerts, but with AWS issues, it was hard to separate signal from noise.
If you were impacted by any of this, I’m truly sorry. The entire internet is at the mercy of this shared infrastructure like AWS, and today reminded us just how fragile it can be.
What Comes Next
Now that things are returning to normal, I’m thinking through how we can be better prepared next time. We’re going to:
* Build a proactive alert plan
* Create a downtime checklist
* Share best practices so you can react quickly on your end too to prevent ad spend loss.
Even when it’s “not our fault,” downtime, even in the integrations layer, can cause real issues. So we're committed to doing everything we can to prevent these major impact on our side when these issues happen.
Thanks for bearing with us. I appreciate all of you more than you know.
– John