At 04:09 UTC on July 19, 2024, airports stopped working. Hospitals stopped working. Banks stopped working. Retail terminals stopped working. Emergency services stopped taking calls in parts of the United States. Not all at once, but in a rolling wave that spread across every continent as businesses came online for their morning. The cause was a routine security update from a vendor called CrowdStrike, pushed to about 8.5 million Windows devices, that placed each machine in an unrecoverable boot loop.1
Vendor software update risk of that scale had never played out this publicly. Fixing it required physical access to every affected machine. Boot to safe mode. Delete a specific file. Reboot. Multiply by 8.5 million endpoints spread across data centers, offices, warehouses, kiosks, and hospitals, and you understand why parts of the world took a week to recover.
CrowdStrike’s own root cause analysis, published August 6, 2024, traced the outage to a mismatch inside a Rapid Response Content update: the sensor expected 20 input fields, the update delivered 21, and the Content Validator’s checks did not catch it. The bad file, Channel File 291, went to production, and the out-of-bounds memory read that resulted crashed every host that received it.2
Delta Air Lines, one of the most exposed businesses, canceled 7,000 flights over five days, disrupted 1.3 million passengers, and filed a $500 million lawsuit against CrowdStrike three months later. A Georgia judge has since allowed key claims in that suit to proceed.3
The most important thing a business owner can learn from the CrowdStrike outage is not the technical details. It is the vendor software update risk your business is running today, and what your MSP is doing about it.
What Actually Happened
CrowdStrike Falcon is an endpoint detection and response product installed on tens of millions of business Windows systems. It runs at the kernel level, which is how it can see and stop threats other tools miss. That kernel-level access is also why a bad content update from CrowdStrike could turn a working Windows machine into a doorstop before the operating system finished loading.
The update in question was not a full sensor update. It was a smaller, faster kind of update CrowdStrike calls a Rapid Response Content file, designed to push new detection rules to sensors quickly. The RCA describes what went wrong. A Template Type introduced back in February 2024 had been designed to accept 21 input fields. Every sensor deployed in the field expected only 20. The Content Validator, which is supposed to catch exactly this kind of mismatch before content ships, did not. When Channel File 291 hit the sensors on July 19, the sensor read past the end of the array it had allocated, triggered a fault, and Windows crashed.2
The blast radius was global because CrowdStrike deploys content updates to all customers on roughly the same schedule. There was no staged rollout that would have limited the damage to a subset of hosts. There was no automated rollback the moment crashes started. There was, at the customer level, no way to opt out of that specific update without also opting out of the security product itself.
What Vendor Software Update Risk Really Means
Every business runs on software from vendors that push updates. Windows itself. Microsoft 365. The line-of-business application every operator opens on Monday morning. The RMM your MSP uses. The EDR that protects your endpoints. The firewall. The backup product. The identity provider.
Every one of those vendors will push a bad update at some point. Most of the time it is small and gets patched quietly. Occasionally it is CrowdStrike-scale. The question is not whether it will happen. The question is what your business has in place to survive it when it does.
Vendor software update risk breaks down into four operational failures that businesses experienced during July 2024. Each one is a decision your MSP has made, whether they told you or not.
- No staged deployment. The update went to every host at once. A phased rollout to a small percentage of endpoints first would have caught the problem before it took down 8.5 million machines.
- No offline recovery plan. Boot-looped machines needed manual, physical intervention. Any organization with hundreds or thousands of endpoints scattered across sites had no realistic way to reach them all in a reasonable time.
- No documented dependency map. Businesses discovered which of their tools depended on CrowdStrike Falcon in the middle of the outage. That is not the time to find out.
- No vendor concentration awareness. Many businesses had the same vendor running on every workstation, every server, every domain controller. A single vendor’s bug became a single point of failure across the entire environment.
What Your MSP Should Be Doing Differently
Managing vendor software update risk is now on the table for every serious operator. The vendors are now, mostly, making changes. CrowdStrike has announced staged deployments and customer-side control over the timing of Rapid Response Content updates. Microsoft has pushed for changes to the way security tools access the Windows kernel. Those changes are good. They are also not your defense strategy.
Your MSP should be applying the same discipline to every vendor in your stack, not just to the one that made the news.
Staged patching for every critical tool. A pilot group of hosts receives updates first, on a delay, with a rollback plan. Your CFO’s laptop and the domain controller are not in the pilot group.
An out-of-band way to reach every endpoint. If the endpoint agent itself is broken, your MSP has to have another path in. That means physical access plans, secondary remote tools, and documented site contact information that is not stored in the same system that just went down.
A vendor dependency map you can see. Which of your systems depend on which vendors? Where is the single point of failure? Your MSP should be able to show you this on request, not build it during an incident.
An incident response plan that includes vendor failure. Not just ransomware. Not just outage of your own systems. The scenario where a trusted vendor breaks the world, and your business has to keep operating while the vendor works on a fix.
The Question Worth Asking
Ask your MSP one thing this week. If CrowdStrike had happened to your business, how would we have recovered, and how long would it have taken?
The good answer names the vendors in your environment, walks you through staged patching, describes the out-of-band recovery path, and quotes a realistic time to bring every endpoint back online. The bad answer is silence, or a shrug, or “we would have figured it out.” The vendors that broke the world last July are not the last ones that will. Vendor software update risk is a permanent condition. What separates the businesses that get through the next one from the ones that lose a week is whether their MSP already planned for it.
Where This Leaves You
CrowdStrike was not a security failure. It was an operations failure, inside the vendor, and inside every business that had no plan for what happens when a trusted update turns out to be poison. The next one may come from a different vendor and take a different shape. It will not care that yours is a small business. Your line-of-business application publishes patches too.
The businesses that survived July 19, 2024 with the least damage were the ones whose IT partner had already treated vendor updates as a risk to manage, not a background convenience. That is the standard your MSP should be operating at, whether you have asked them for it or not.
Sources
1 Cybersecurity and Infrastructure Security Agency (CISA), “Widespread IT Outage Due to CrowdStrike Update,” cisa.gov, July 19, 2024.
2 CrowdStrike, “Channel File 291 Incident RCA is Available,” crowdstrike.com, August 6, 2024.
3 CNBC, “Delta, CrowdStrike sue each other over widespread IT outage that caused thousands of cancellations,” cnbc.com, October 25, 2024.
About Brent Lacy: Brent Lacy has been in the IT industry since 1997. He moved into the managed services world around 2015 and was doing vCIO work before the title even existed. He writes about the operational discipline, trust-based relationships, and strategic thinking that separate MSPs built to last from those built to bill. He is the author of Rewired MSP: Mastery, Scalability and Performance, vCIO Rewired: Virtually Conquering IT Obstacles, and Near Miss: Preventable IT Failures Threatening Your Business Security.