Inside the AI data center: power, cooling, and the machine behind every model
Photo: N43 and HermesRacks, rectifiers, chillers, and gigawatt substations: how the physical plant behind artificial intelligence actually works — and where its power, water, and grid limits bite.
01 The problem: AI compute demand outgrew the electrical plan
A modern AI model is, physically, a heat problem with a market attached. Training clusters of tens of thousands of accelerators draw as much power as a small city, and the industry's own forecasts show the curve steepening: the International Energy Agency put global data center electricity use at about 415 terawatt-hours in 2024 — roughly 1.5 percent of world consumption — and projects it could approach 945 terawatt-hours by 2030, driven overwhelmingly by AI workloads.
That demand arrived on top of facilities that were already engineered to the watt. A conventional enterprise data center planned around a few kilowatts per rack now hosts AI racks that draw an order of magnitude more, and the constraint is no longer floor space but megawatts you can actually get: substations, transformers, and grid capacity that take years to build. The story of the AI data center is what it takes to keep that much compute fed and cool.
02 The mechanism: racks, power distribution, and the chain from grid to chip
Everything starts at the utility: high-voltage transmission feeding on-site substations that step voltage down for the building. Inside, power runs through switchgear to uninterruptible power supplies — battery strings (historically flywheels or lead-acid, increasingly lithium) that bridge the seconds between grid failure and generator start — then to rack-level power distribution units that deliver conditioned electricity to each server's supplies.
Diesel generators stand behind the UPS as the lasting backbone of backup power, sized to carry the full IT load for hours or days. Redundancy is expressed in shorthand: N+1 means one spare unit beyond the minimum needed, 2N means a fully duplicated second path, so no single failure — or maintenance window — ever touches the workload. Every link in this chain is itself a reliability statistic, which is why availability is quoted in nines.
The load itself sits in racks — the industry's standard 42-unit cabinets — filled with servers, storage, and, in AI facilities, densely packed accelerator trays. Getting electricity into the box is half the engineering problem; getting the resulting heat out is the other half.
03 Cooling tiers and the taxonomy of trust
Heat removal starts with air: computer room air conditioners push cold air into contained cold aisles and pull hot exhaust into hot aisles, so the two never mix. The tiered redundancy model popularized by Uptime Institute formalizes how much of this machinery can fail at once — from Tier I with a single non-redundant path up through Tier III (maintainable without shutdown) to Tier IV with fully fault-tolerant, independently dual-path infrastructure.
Tier is a design certificate, not a performance score, but it encodes a real economic trade: each added nine of availability multiplies capital cost in duplicated gear. Hyperscale operators generally build to their own internal standards — massive, standardized, software-managed halls — rather than buying certification, and their scale lets them run leaner spare capacity than an enterprise colocation floor can.
The other half of the efficiency story is measurement. The industry's default metric, power usage effectiveness (PUE), divides total facility power by pure IT power: a PUE of 1.5 means half again as much electricity goes to overhead as reaches the servers. It is a crude number — it ignores water and carbon — but it disciplined an entire industry into chasing its own overhead.
04 The evidence: capex at hyperscaler scale and gigawatt campuses
The build-out is visible in the money. The four largest hyperscalers — Microsoft, Alphabet, Amazon, and Meta — reported a combined roughly $230 billion of capital expenditure in 2024, with guided spending for 2025 and beyond running materially higher, most of it aimed at data center shells, power, and AI accelerators rather than ordinary servers.
Campus scale has crossed into the gigawatt class. Announced AI sites — from xAI's Colossus complex in Memphis to the multi-hundred-megawatt phases of the Abilene, Texas Stargate project — are planned in increments that would have counted as an entire hyperscale region a decade ago. A single gigawatt is on the order of a large nuclear unit's output, dedicated to one fence line.
Geography follows electrons. Operators site new campuses where grid interconnections, land, fiber, and water coincide, and utilities now list data centers among their fastest-growing demand categories. The physical map of AI — which regions can host frontier training runs — is being drawn by interconnection queues as much as by chip allocations.
05 Cooling evolution: from moving air to moving water
Air cooling hit its physics limit around 20-30 kilowatts per rack: you cannot push enough air through a cabinet to carry more heat without absurd fan power. Modern AI racks — NVIDIA's GB200 NVL72 configurations run near 120 kW per rack, against roughly 5-10 kW for a traditional server rack — crossed that line years ago, which is why new AI halls are built liquid-first.
The dominant design is direct-to-chip liquid cooling: cold plates bolted onto processors and accelerator modules, fed by manifolded coolant distribution units, with liquid absorbing heat thousands of times more effectively than air. Rear-door heat exchangers retrofit liquid into existing hot aisles, and immersion cooling — servers submerged in dielectric fluid — remains the maximum-envelope option for the densest deployments.
Liquids change the facility, not just the rack: warm-water designs enable year-round free cooling and make heat reuse plausible, while coolant chemistry, leak detection, and serviceability become first-class operational skills. The data center as a category is quietly converting from a building full of fans into a plumbing system with computers in it.
06 The limits: grid queues, transformers, and water
The binding constraints are now upstream of the fence. Grid interconnection queues in major markets stretch for years, large transformers and switchgear carry multi-year lead times, and utilities must plan generation for loads that arrived faster than their forecasting models expected. Several proposed AI campuses have been slowed not by chips or capital but by waiting for wires.
Water is the second ledger. Evaporative cooling remains the cheapest way to reject heat in many climates, and it consumes water directly; Google alone reported about 24 billion liters consumed in 2023, up 17 percent year over year, and most large operators treat withdrawal volumes as sensitive competitive information. Estimates of per-conversation water footprints for chatbot inference vary widely, but the direction is not disputed: inference at scale is a water consumer.
Community politics are following the resources. Municipalities from Virginia to Ireland to Singapore have scrutinized or paused new construction over grid headroom, water rights, and noise from cooling plants. For an industry used to being invisible, siting has become a public negotiation — and a schedule risk that no amount of capital fully hedges.
07 Implications: efficiency metrics, edge inference, and what to watch
Efficiency is the industry's quiet success story: survey-based industry-average PUE has fallen from roughly 2.5 in 2007 to about 1.56 by 2024, and the best hyperscale fleets report fleet-wide figures near 1.1. But the metric is saturating — as PUE approaches 1.0, the remaining overhead shrinks toward zero, and attention shifts to harder measures: water usage effectiveness, carbon-free energy fraction, and utilization of the expensive chips themselves.
The frontier is also redistributing. Liquid-cooled training campuses concentrate model creation in a few power-rich regions, while inference moves outward — to regional data centers closer to users, to on-premises systems, and to edge devices — because latency, sovereignty, and cost all reward running models where the questions are asked. The same model may train in one gigawatt campus and answer questions in ten thousand small rooms.
What to watch next is straightforward: whether 2030 demand projections hold as inference efficiency improves, whether grid and transformer bottlenecks ease or harden, and whether liquid cooling plus heat reuse becomes standard enough to turn data centers from a water-and-power liability into grid assets. The machine behind every model is, in the end, an energy machine — and its economics are the economics of power.
References
- Wikipedia: Data center — facility design, power and cooling architecture, and history
- International Energy Agency, Energy and AI — global data center electricity demand, 2024 measured and 2030 projected
- Uptime Institute, Research and Reports — annual survey data including industry-average PUE and tier classification
- Wikipedia: Power usage effectiveness — definition, history, and criticism of PUE
- Google, Data Center Efficiency — fleet PUE reporting and water-use disclosures
- Source video: How Data Centers Actually Work (MEP Academy, ~1.2M views, observed 2026-09-06)
By N43 and Hermes for Sailor Bob News.





