AI Data Centers: Compute, Energy, Water, and the Scaling Constraint
The new AI data center is not built where the users are. It is built where the power is.

Research updated Sep 10, 2026
Key topics
The new AI data center is not built where the users are. It is built where the power is.
That single sentence explains more about the current state of AI infrastructure than any chip roadmap. For the past few years, the story of AI scaling was a story about silicon: who could get GPUs, how many, and how fast. That story is not over, but it is no longer the only binding constraint. The constraint has moved to a less glamorous place — the ability to deliver large, dense, reliable power to a specific piece of land, and to cool the equipment that consumes it.
This article is a guide to reading that shift without either dismissing it as hype or repeating vendor claims as settled fact. The goal is not to tell you whether AI is "running out of energy." The goal is to give you a mental model for figuring out which resource binds at which site, on which timeline, and according to whom.
The Constraint Moved From Chips to the Power Grid

Start with the vocabulary, because the numbers only make sense once the units are clear.
An AI data center is a facility designed and optimized for large-scale AI workloads — training and running models — rather than general-purpose IT services. It is built around dense clusters of specialized processors, primarily GPUs and TPUs, packed into racks. A rack is the physical frame that holds servers and networking gear, and it is the standard unit for describing how much computing a facility can hold.
The number that changed is power density: how much electricity a single rack draws. A conventional server rack typically uses about 7 to 10 kilowatts (kW). An AI-capable rack can demand 30 kW to over 100 kW. That is the same building footprint, the same floor, the same rack slot — and roughly three to ten times the electrical load.
That difference breaks assumptions that were baked into data center design for two decades. Cooling systems were sized for the old number. Electrical distribution — the transformers, busways, and switchgear that move power from the utility connection to the rack — was sized for the old number. And most importantly, the grid interconnection agreement with the local utility was sized for the old number.
This is why the bottleneck is no longer only silicon supply. It is the ability to deliver large, dense, reliable power to a specific piece of land, on a timeline that matches when the GPUs arrive.
One caveat before we go further: much of what is publicly claimed about AI energy use comes from the companies selling the infrastructure. That does not make the numbers wrong. It means they are claims, not independent measurements, and they should be read with the source's incentive in mind. We will come back to how to do that systematically.
Why a Query Costs More Than a Search
To reason about energy and water claims, you need one intuitive unit: the cost of a single query.
When you send a request to a large language model, the system processes it token by token — a token is roughly three-quarters of a word. Each token requires computation, computation requires electricity, and electricity produces heat that must be removed. That entire process is called inference.
A widely cited comparison puts a ChatGPT-style query at roughly 2.9 watt-hours (Wh), against about 0.3 Wh for a conventional web search. That is roughly a tenfold difference, and it is the number that shows up in most headlines.
Treat it as a starting point, not a law. The comparison is fragile for several reasons. Model size varies enormously. Caching means repeated or similar queries can be served far more cheaply than fresh ones. Hardware generations differ in efficiency by large factors. And the measurement boundary — what exactly was counted — is rarely stated in the same way across studies. A 2.9 Wh figure from one deployment on one hardware generation does not describe every deployment.
The metric the industry is actually optimizing is tokens per second per watt — how much useful output you get per unit of energy. That framing matters because it forces the conversation toward efficiency per unit of work rather than raw performance. It is also the metric that vendors cite when they claim dramatic year-over-year improvements.
Here is the practical takeaway. When you see a per-query energy or water number, ask four questions: What hardware? What model? What utilization rate? What measurement boundary? If the source cannot answer those, the number is a headline, not a finding. Per-query figures are a moving target because efficiency gains compound — a figure measured last year may not describe today's deployment at all.
Water: The Cooling Cost Nobody Budgets For
The second resource constraint is water, and it is the one most likely to be both overstated and understated in the same news cycle.
Most large data centers use evaporative cooling: heat is removed by letting water evaporate, which carries the heat away. The key word is consumed. The water is not borrowed and returned; it leaves the facility as vapor. That is different from water used in a closed loop, which is recirculated.
At facility scale, the numbers are large. Research on data center electricity and grid impacts cites an estimate that a 100 MW data center in the U.S. using evaporative cooling consumes on the order of 2 million liters of water per day — comparable to the daily use of roughly 6,500 households. The same research estimates global data center water consumption at around 560 billion liters per year, with a possible rise to roughly 1,200 billion liters per year by 2030.
Now contrast that with the per-query view. One large provider estimates that a typical query uses between 0.0 and 0.067 milliliters of water, with a median below a single drop.
Both numbers can be true at once, and reconciling them is the actual lesson. Per-query water is tiny. Facility-level water is enormous. The difference is scale and cooling design — not contradiction. A single query is a rounding error; millions of queries per day against an evaporatively cooled campus is a watershed-level question.
Two more things matter here.
First, water intensity is a design choice, not a fixed law. Zero-water and closed-loop cooling designs are being deployed, and providers that adopt them change their water profile substantially. A facility's water number is a statement about its engineering, not about AI in general.
Second, local context matters more than global averages. Water stress is a watershed-level problem. A data center in a water-rich region and one in a drought-prone region can have identical consumption numbers and completely different consequences. National or global averages hide the only comparison that matters to the people living near the facility.
Grid Interconnection Is the Slowest Moving Part
Here is the part of the story that gets the least attention and probably deserves the most.
Interconnection is the process by which a data center gets permission and physical connection to draw power from the grid. It is not a formality. A facility can only draw power once the grid operator approves the request and the physical connection is built, and that queue can take years in constrained markets.
This is why the geography of AI infrastructure is shifting. Analysis of the European market shows the average distance of new sites from a major hub rising sharply — projects delivered between 2022 and 2025 averaged around 46 kilometers from a major hub, while the 2026–2028 pipeline averages roughly 175 kilometers. Inner-city locations are falling as a share of the pipeline. As one industry analyst put it, data centers are being brought to where the power is, not the other way around.
There is a second, subtler problem. AI load is harder on a grid than ordinary load. Training runs and inference traffic cause large, fast swings in demand. Those swings stress generation equipment and, in some reported cases, have been hard enough on natural gas turbines that operators need large battery banks to smooth the curve.
The emerging answer is demand response and load flexibility: rather than building for absolute peak demand, shift or throttle some workloads when the grid is strained. One large provider has described expanding its demand response toolkit specifically for machine learning workloads, and has signed utility agreements to put it into practice.
Be precise about the limits. Flexibility works for some workloads and not others. Services with hard reliability requirements — search, maps, cloud customers in healthcare — cap how much a facility can flex. And demand flexibility is still early-stage and location-dependent. It is a promising tool, not a solved problem.
The geographic consequence is real and worth internalizing: moving away from core hubs buys cheaper land and available power, but it changes latency, staffing, and cost profiles. Those tradeoffs are now part of infrastructure strategy, not facilities trivia.
What the Efficiency Numbers Actually Claim
This is where readers most often get fooled, in both directions.
The headline claim pattern is now familiar: providers report combined near-term reductions in energy per query in the range of 8x to 20x, achieved by stacking improvements. One large provider's published research describes gains of that magnitude from a combination of factors: denser and more efficient accelerators, better model serving, higher utilization, improved cooling, and smarter workload scheduling. The claim is that these improvements build on each other, so a gain in one area becomes the new baseline for everything running on the platform.
Where do those gains come from? Mostly from doing more useful work per unit of energy — better hardware, better software, better scheduling, less idle capacity. That is a real and measurable direction of travel.
But here is the tension you have to hold. Efficiency per query can improve while total consumption rises. If inference gets cheaper, more inference gets done. Cheaper queries invite more queries, more products, more automation, more always-on agents. This is the classic efficiency paradox: gains per unit of work do not automatically translate into lower total draw.
So give yourself a decision rule. An efficiency claim answers "per unit of work." A capacity claim answers "total draw." They are different questions, and they are constantly conflated in headlines. A 20x efficiency improvement and a doubling of total electricity demand can both be true, and often are.
What is genuinely not established: whether efficiency gains will outrun demand growth. That is an open question. Anyone claiming certainty in either direction — "AI will always get more efficient, so energy is a non-issue" or "demand growth makes efficiency irrelevant" — is overreaching past the evidence.
Nuclear, Gas, and the Behind-the-Meter Bet
If the grid is slow, the obvious move is to bring your own power. That is the behind-the-meter bet: siting generation on or near the campus rather than waiting for the utility.
The tradeoff to understand first: data centers need firm, dispatchable power — power available on demand, not when the wind blows. But most new firm generation takes far longer to build than a data hall.
Nuclear is attractive on one dimension and awkward on another. Reactors have the highest capacity factor of any generating technology — in the U.S., they run at maximum output about 92.5% of the time. But they are slow to ramp, capable of changing output by only a few percent of rated capacity per minute. AI load swings fast; reactors prefer steady output. That mismatch is the whole problem.
One proposed answer is a storage buffer: keep the reactor running at steady output and absorb the mismatch in a thermal or battery buffer, so the expensive asset stays fully utilized even when demand dips. At least one reactor design stores excess heat in molten salt and taps it when demand spikes. It is a clever way to pair high capital cost with flexible output.
The honest caveat: small modular reactor economics are unproven at scale, and timelines are long enough that they cannot solve a near-term constraint. If the manufacturing cost reductions materialize, the payoff is a decade or more out.
Natural gas is the faster bridge, and it inherits two problems: fuel price exposure and emissions questions. Fast-ramping load also has real equipment consequences, as the turbine stress noted earlier shows.
The decision rule for readers: match the power source to the load shape and the timeline, not to the press release. A technology that is excellent on capacity factor but slow to ramp is a different answer than one that is fast to build but exposed to fuel prices. The right question is not "which power source is best" but "which source fits this load profile, at this site, on this schedule."
How to Read a Bottleneck Claim Without Getting Fooled
Everything above collapses into a short checklist you can apply to the next headline.
Ask what the unit is. Per query, per facility, per year? A per-query number and a per-facility number describe different things and are not interchangeable.
Ask what the boundary is. IT load only, or the whole campus including cooling and distribution? Water consumed, or water withdrawn? The boundary changes the answer by large factors.
Ask who measured it, and whether they had an interest in the answer. Vendor figures are not automatically wrong, but they are claims. Third-party analysis and market estimates are a different category. Confirmed operational facts are a third.
That gives you four claim types to keep separate:
- Confirmed operational facts — things that are directly observable and verifiable.
- Vendor-reported figures — real measurements, but selected and framed by an interested party.
- Third-party analysis and market estimates — independent, but often modeled rather than measured.
- Forward-looking projections — useful for planning, not for stating as fact.
There are two symmetric failure modes here. Dismissing all constraints as hype is one. Accepting every projection as settled fact is the other. Both feel like having a position; neither survives contact with a specific site and a specific timeline.
And notice how the same underlying data supports different conclusions depending on the question. A single facility can be highly efficient. A region can be strained. The global total can be growing fast. All three can be true, and the useful question is always which one you are actually asking about.
For decision-makers, the practical implication is direct: site selection, power contracting, and cooling design are now strategic decisions. They determine whether your compute arrives on schedule, at what cost, and with what regulatory exposure. Treating them as facilities details is how projects slip by years.
What to Learn Next If This Is Your Problem
If you want to go deeper, build the mental model in the order the constraints actually bind.
Start with load shape. Before reading any capacity plan, understand utilization, duty cycle, and the difference between peak and average demand. Most bad infrastructure reasoning comes from confusing a peak number with an average number.
Then learn the physical stack in order: power delivery, cooling, and rack density. That is the sequence in which constraints show up when you try to add compute to a real building.
Then learn the measurement layer: how energy and water per unit of work are defined, and why changing the definition changes the answer. This is the skill that lets you read the next claim without needing someone else to interpret it.
Adjacent topics are separate learning tracks, not repeats of this one. Building reliable and scalable AI infrastructure is about architecture and operations. The global compute and chip landscape is about supply and policy. AI supply-chain security is about provenance and permissions. Each deserves its own attention.
Here is a small exercise worth doing. Take one published efficiency or water claim — any of them — and try to reconstruct the assumptions behind it. What hardware? What utilization? What boundary? If you cannot reconstruct them from the source, that is the finding. You have learned that the claim is not auditable, which is more useful than memorizing the number.
The useful question is not whether AI is running out of energy. It is which specific resource binds at which specific site, on which timeline, and according to whom. Watch one thing going forward: whether per-query efficiency gains continue to outpace total demand growth. And before you accept any capacity or efficiency number, identify the unit, the boundary, and the source's incentive. That habit will outlast every figure in this article.
References
- NVIDIA, Energy Leaders Accelerating Power‑Flexible AI Factories to Fortify the Grid | NVIDIA Blog
- Electricity Demand and Grid Impacts of AI Data Centers
- How we’re making data centers more flexible to benefit power grids
- Europe AI data centres seek cheaper, quicker energy and land
- TerraPower’s nuclear reactor has a secret weapon for powering AI data centers | TechCrunch
- Scaling AI with 8 to 20x energy efficiency | The Microsoft Cloud Blog


