A facilities manager at a regional data center once told me his worst outage had nothing to do with a cyberattack or a software bug. A bad afternoon in July, an aging cooling unit, and a rack that hit shutdown temperature before anyone noticed the airflow problem. Data center resilience failed on him in the dullest way available.
Three hours down. Six figures gone. Root cause almost embarrassing in its simplicity: hot exhaust recirculating into equipment intakes, because nobody had separated it properly.
That story names something IT teams avoid. Data center resilience gets treated as a software discipline when half of it is physical, sitting quietly in a room most executives never enter. AI workloads have made that half considerably harder to ignore.
AI Workloads Broke the Thermal Assumptions
The numbers moved fast. Average rack density sat around 7 kW in 2021 and reached roughly 27 kW in 2026, according to AFCOM’s annual survey.
Averages understate it, though. Racks running H100 and H200 GPU clusters routinely pull 40 to 100 kW, against 5 to 10 kW for a conventional server build, and that shift took barely two years. Meanwhile, only about one in five operators report being ready for the 50 to 70 kW racks AI deployments now use routinely.
Past a certain point, air stops working at all. Air cooling generally gives out somewhere above 50 to 100 kW per rack, which is why liquid cooling stopped being a preference and became a requirement for current GPU hardware.
Most facilities hosting AI work were never built for it. Data center resilience assumptions written in 2018 do not survive a GPU cluster. They were built for mixed enterprise workloads, and they are now running equipment that generates several times the heat per square metre.
Physical Containment Is the Part Everyone Assumes Someone Else Handles
Every data center makes heat. Where that heat goes matters far more than people outside facilities management tend to realise.
Hot aisle containment solves one specific problem: keeping exhaust air off equipment intakes, so cooling systems stop fighting air they already cooled once.
Without it, the cooling plant works harder than it should, recycling warm air instead of drawing in properly conditioned supply. Energy costs climb, yes. Equipment failure is the larger risk. Servers running above their rated operating temperature fail sooner, and they pick their moment — peak load, not a quiet Tuesday when someone would notice.
Smaller sites often assume containment belongs to hyperscale operators with thousands of racks. Not really true. A modest server room benefits too, since heat problems refuse to scale down gracefully. Ten racks with no airflow separation run just as hot, proportionally, as a thousand. Add two GPU nodes to that room, and the proportion gets worse quickly.
Why Data Center Resilience Budgets Skip the Physical Layer
Here’s the imbalance worth naming. Most IT budgets pour money into disaster recovery, backup systems, failover architecture, and redundant cloud storage, while the physical environment stays an afterthought that facilities handles somewhere else.
That division collapses under any scrutiny. A backup is worthless if the hardware processing it overheats before recovery finishes.
Digital and physical resilience are two halves of one problem. Fund one, neglect the other, and you own expensive redundancy that fails precisely when it matters. The ownership question sits underneath this, which is why AI transformation so often turns out to be a governance problem before it becomes a technical one.
Cloud Backup Costs Creep Without Anyone Deciding
On the digital side, cloud backup solves a real problem while its pricing structure catches teams off guard constantly.
The cost of Azure Backup scales with storage volume and retention period. Both look manageable at setup. Both drift upward as retention policies extend and data accumulates month over month. Budget against year-one numbers and year three arrives substantially higher, purely because nobody revisited the retention settings.
AI workloads accelerate this too. Training data, checkpoints, and model artefacts consume storage at a rate traditional application backups never approached.
So build in a recurring cost review, quarterly at minimum, instead of setting a policy once and assuming the price holds. Storage is one of the few IT line items that reliably grows without anyone choosing to grow it.
Redundancy Without Testing Is Just an Assumption
Plenty of organisations build redundant systems, physical and digital, then never test whether the redundancy survives an actual failure.
Backup generators nobody has load-tested in a year. Failover configured correctly at setup, never verified after a software update quietly changed something underneath it.
Testing means pulling a plug or simulating a failure, not reading documentation. It is unglamorous work, and it gets deprioritised constantly. It is also the only way to learn whether the resilience you paid for exists. Automation helps here — infrastructure that detects and repairs its own failures shortens recovery — though automated remediation still assumes the physical layer holds.
The Gap Between Facilities and IT
Most data center resilience failures live between departments. Facilities assumes IT watches server health. IT assumes Facilities watches room temperature. Nobody owns the intersection.
That gap produced the July outage. It produces most physical failures, quietly, right up until it doesn’t. Predictive tooling closes part of it, since monitoring systems now flag degradation before an outage lands rather than paging someone afterwards.
Organisations that avoid this are not necessarily spending more. They stopped treating the server room and the disaster recovery plan as separate conversations, and started asking whether the systems protecting their data can survive the room those systems live in.
For anyone deploying AI hardware into existing space, that question stopped being theoretical.
Related: Why Did ChatGPT, Claude & Grok Fail at the Same Time?
