A 3 GW data center load drop is your cloud's new risk
A transmission fault in Ashburn knocked over 3 GW of data center load off the PJM grid in seconds. What grid fragility in Data Center Alley means for your uptime.
On July 23, one transmission line faulted in Ashburn, Virginia, and more than three gigawatts of data center load disappeared from the PJM grid in a matter of seconds. Nobody's site went down. That's the part worth paying attention to — a data center load drop that large is a new category of correlated risk sitting underneath the region where a huge share of the internet actually runs.
What actually happened in Data Center Alley
Data Center Knowledge reported that the fault hit Ashburn — the densest concentration of data centers on earth — and hyperscale facilities responded exactly as designed: their own protection systems transferred them to backup generation. This was not utility-directed load shedding. Thousands of megawatts of demand simply stopped being demand, all at once, because a few hundred independent UPS controllers made the same decision inside the same second.
PJM confirmed the event produced a measurable frequency change but no reliability impact to the bulk power system, and Dominion Energy's operators stabilized things within minutes. NERC said it is reviewing the event through its normal monitoring process. TechCrunch's follow-up put finer numbers on it — roughly 3.1 GW off in about 30 seconds, a resulting supply surplus, and a recovery measured in double-digit minutes rather than the usual few seconds. The load that dropped was about 3% of PJM demand at the time.
This wasn't a surprise to regulators. NERC issued a rare Level 3 alert on May 5 over exactly this failure mode, per Utility Dive, with seven mandatory actions for transmission planners and operators due August 3, 2026 — modeling requirements, commissioning tests at full and no load, voltage variance testing, fault recording. The regulator saw the pattern before the July fault confirmed it.
Why a data center load drop matters for your business
Your cloud region is a physical place with a substation. us-east-1 is Northern Virginia. When you architect across three availability zones, you are buying independence from a rack failure, a switch failure, maybe a building. You are not buying independence from a regional grid disturbance that every facility in the metro reacts to simultaneously.
Nothing broke this time — the backup generation did its job. But the failure mode this exposes is correlation. The whole premise of AZ redundancy is that failures are independent, and a grid event is the clearest counterexample there is. Add the trajectory: data centers are heading from a few percent of PJM load toward a quarter of it over the next decade. The events get bigger from here, not smaller.
The move is not to panic-migrate. It's to know your actual blast radius. Which of your systems are single-region right now? When did you last fail one over on purpose, with real traffic, and time it? If your answer to that second question is "we haven't," your recovery time objective is a guess, not a number. Pick your two most revenue-critical services this quarter and prove the failover works. That's a one-week exercise that turns an assumption into a measurement.
Key takeaways
- A July 23 transmission fault in Ashburn caused more than 3 GW of data center load — about 3% of PJM demand — to disconnect within seconds as facilities self-transferred to backup power
- PJM reported a measurable frequency change but no reliability impact; NERC is reviewing the event
- NERC already issued a Level 3 alert on May 5 over sudden data center load loss, with mandatory actions due August 3, 2026
- Availability-zone redundancy buys you independence from local failures, not from a regional grid disturbance every facility reacts to at once
- Test a real failover on your two most revenue-critical services this quarter so your RTO is measured instead of assumed
Never actually tested your failover? We build systems with a documented recovery path and we run the drill, so your uptime story survives contact with a bad Tuesday. See how we build resilient systems or tell us what's single-region right now.
Sources: Data Center Knowledge, TechCrunch, Utility Dive.
- #data-centers
- #cloud-infrastructure
- #uptime
- #resilience
- #ai-infrastructure
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Qualcomm's price increase hits your hardware budget
Qualcomm told customers of a double-digit price increase on chips shipped after September 1. The AI buildout is now repricing hardware that has nothing to do with AI.
Read it24,000 exposed BMCs leak hashes: close your IPMI port
A 2004 protocol flaw with no patch is handing out password hashes from 24,000 internet-exposed server BMCs. The fix is network exposure, not a firmware update.
Read it