- Four different quantities travel under the same name. Direct on site consumption, direct withdrawal, indirect consumption from generating the electricity, and total withdrawal including the returned share.
- US data centers were estimated to consume about 449 million gallons a day directly in 2021, roughly 163.7 billion gallons for the year. The indirect figure for the electricity they used in 2023 is larger, around 211 billion gallons.
- Most public arguments are two people quoting different measures at each other. Once you ask which of the four a number refers to, the apparent contradictions mostly dissolve.
- The widely repeated projection of 4.2 to 6.6 billion cubic meters for global AI in 2027 is a withdrawal figure, not a consumption figure. The two are not interchangeable.
- Location decides whether any of it matters. The same gallon in a wet basin and a stressed one are not the same gallon.
A medium sized data center is reported to use as much water as a small town, or 110 million gallons a year, or nothing at all because it is air cooled, or five million gallons a day if it is one of the large ones. All four of those statements can be published in the same week about facilities in the same country, and all four can be accurate. They are answers to different questions.
This piece is not an argument about whether data centers use too much water. It is the thing you need before you can have that argument: a way to tell which quantity any given number describes. Every figure below is dated and attributed, and the aim is that you can come back in two years, read a new headline, and place it correctly.
What is the difference between withdrawal and consumption?
Withdrawal is all the water taken out of a river, a lake or an aquifer. Consumption is the part of that water which does not come back. The US Geological Survey, which has maintained these definitions for decades, describes withdrawal as "water removed from the ground or diverted from a surface-water source for use", and consumptive use as the part of the withdrawn water that is evaporated, transpired, put into products, drunk, or otherwise not available for immediate use.
The bridge between them is return flow, which USGS defines on the same page as water that reaches a groundwater or surface water source after release from the point of use and so becomes available again.
None of this is specific to computing. A power station withdraws enormous volumes and returns most of them warmer. An irrigated field withdraws less and returns almost none. The reason it matters here is that evaporative cooling, which is what most large facilities use, sits at the consumptive end: the water leaves as vapour rather than going back down the pipe.
The four quantities, and which numbers belong to each
Here is the table to keep. It maps the four measures against a real published figure for each, with the year and the source, so that a headline can be sorted in about ten seconds.
| Measure | What it counts | A published figure | Year and source |
|---|---|---|---|
| Direct on site consumption | Water evaporated at the facility itself | About 449 million gallons per day, roughly 163.7 billion gallons for the year, across US data centers | 2021, cited by EESI from a Nature paper |
| Direct on site withdrawal | Everything the facility takes in, including the share discharged again | Around 20 percent of the water withdrawn is discharged to municipal wastewater rather than evaporated | EESI, undated ratio |
| Indirect consumption | Water used to generate the electricity the facility consumes | About 211 billion gallons, at roughly 1.2 gallons per kilowatt hour against 176 terawatt hours of electricity | 2023, federal figures cited by EESI |
| Total withdrawal, direct plus indirect | Every drop taken from any source on the facility's behalf | 4.2 to 6.6 billion cubic meters projected for global AI | 2027 projection, Li and colleagues |
| Per facility figures | One building, not a national total | Up to about 110 million gallons a year for a medium facility, up to 5 million gallons a day for a large one | EESI |
Read across the first and third rows and the most common misunderstanding becomes visible. The indirect figure, 211 billion gallons for 2023, is larger than the direct figure of 163.7 billion for 2021. Different years, so not a clean comparison, but the ordering is the point: for a typical facility on a typical grid, the water spent making the electricity exceeds the water spent cooling the building. An operator who moves to a closed loop cooling system and announces a large reduction has reduced the smaller of the two numbers, and closed loop systems require more electricity, which raises the larger one.
Why does the same facility produce four different headlines?
Because each measure has a legitimate constituency. A municipal water utility cares about withdrawal, because that is what its permits and its pipe capacity are sized against. A watershed ecologist cares about consumption, because that is the volume that does not come back. A climate analyst cares about the indirect figure, because it tracks the grid. A resident at a planning meeting cares about all three and is usually handed one.
Add to that the ordinary reasons numbers diverge. Estimates from 2021 and estimates from 2023 describe an industry that doubled between 2018 and 2021 and doubled again after that. A national average hides a facility in Arizona and a facility in Ireland. And a good deal of the underlying data is simply not published: the Lincoln Institute reports operators arriving at negotiations with non disclosure agreements attached, leaving communities without figures for water, energy or emissions.
What about the bottle of water per prompt?
That figure is real, comes from a real research group, and is quoted more loosely than it deserves. The estimate originates with researchers at the University of California, Riverside. EESI states it as roughly one bottle, or 519 millilitres, for a 100 word prompt. The Lincoln Institute renders the same underlying work differently, as up to a bottle of freshwater for a chat session of about twenty queries.
Those are not the same claim, and the gap between them is a factor of twenty. Neither publication is careless. They are describing different model sizes, different data centers and different assumptions about which water is counted, which is exactly the problem this article is about.
Go back to the group's own paper and the vocabulary tightens. The abstract of Making AI Less Thirsty states that training GPT-3 in Microsoft's US data centers could directly evaporate 700,000 litres of clean freshwater, and separately projects global AI demand at 4.2 to 6.6 billion cubic meters of water withdrawal in 2027. Evaporate is a consumption verb. Withdrawal is the other measure entirely. The paper is precise about which is which; the retellings usually are not.
A quick test for any AI water figure you meet. Does it say evaporated, consumed or withdrawn? Does it cover the building only, or the electricity too? What year is the estimate from, and how much did the industry grow since? If a headline cannot answer those three questions, it is not wrong, it is unplaceable.
Where the water comes from matters more than the volume
A gallon evaporated beside a large river in a wet climate and a gallon evaporated from a stressed aquifer are the same number and not the same event. This is where the national totals become actively misleading, and where the local reporting is usually better than the global commentary.
The Lincoln Institute's account gives the concrete version. Google reported using over 5 billion gallons across its data centers in 2023, with 31 percent drawn from watersheds described as water scarce. A Meta facility in Newton County, Georgia is described as using 500,000 gallons a day, about a tenth of that county's consumption. Texas facilities are projected at 49 billion gallons in 2025 and as much as 399 billion by 2030. Roughly two thirds of the data centers built since 2022 have gone into water stressed regions, which is a siting choice rather than an accident: hot dry places have cheap land, cheap power and few clouds.
Concentration is the real story
Northern Virginia is the extreme case and worth understanding because everywhere else is on the same curve, further back. The Lincoln Institute describes it as the densest concentration of data centers anywhere in the world, roughly 300 facilities across a handful of counties with dozens more planned, in a corridor through which about two thirds of the world's internet traffic passes. Loudoun County alone holds 27 million square feet of capacity and expects property and real estate tax revenue from these facilities approaching 900 million dollars in fiscal year 2025, which is the reason local opposition rarely wins and also the reason the county board has begun discussing whether leaning that hard on one industry is wise.
The United Kingdom shows the same pattern under a wetter sky, which is the detail that undermines the assumption that rain solves this. Around 80 percent of British data centers sit in the Thames Water service area, with roughly another hundred proposed, in a city that receives under 25 inches of rain a year. London is drier than most people who live there believe. Ireland, meanwhile, is described as running facilities on polluting off grid generators, which converts a water and power problem into an air quality one.
None of those three places chose badly by their own lights. They chose for fibre, for land, for tax treatment and for proximity to customers, and water entered the calculation late or not at all. That sequencing is the thing to watch in your own region, because it is repeatable and it is early.
There is a second order effect that rarely appears in the totals. Evaporative cooling concentrates whatever was dissolved in the water, so the discharge that does return carries higher salt and contaminant loads, and water routed through a facility in one basin does not return to the basin it came from. Both are consumption in the USGS sense, and neither is captured by asking how many gallons went in.
Does using AI make you responsible for any of this?
Partly, and the honest accounting is smaller and stranger than the headlines imply. If you run a business and use hosted AI, your share is almost entirely the indirect figure: the water spent generating the electricity that ran the inference, in whichever region your provider's capacity happens to sit. You have no visibility into that and no control over it beyond choosing a provider.
What you can do is stop repeating figures you have not placed. The bottle per prompt claim in particular gets deployed both as an argument against using AI at all and as evidence that critics are innumerate, and both uses depend on not checking which measure it is. We are not neutral here, since MaShop runs on hosted open source inference rather than hardware we own, and our account of who builds MaShop and what we actually operate is the honest version of that footprint. Nobody serving a model from someone else's data center can tell you its water number with confidence, and you should be suspicious of anyone who claims otherwise.
How this connects to the rest of the AI infrastructure story
Water is one constraint among several that all bind in the same places at the same time. Power is the one that binds first: electric bills in the United States rose at about twice the rate of inflation over the past year, with costs spread across a service area while the tax benefits concentrate locally. Land is third, since the largest campuses cover hundreds of acres of impermeable surface, which changes local runoff before a single server is switched on.
The supply chain underneath has its own water bill, and it is not small. Chip fabrication needs about 1.5 gallons of tap water for every gallon of ultrapure water it uses, and a typical fabrication plant is described as consuming 10 million gallons of ultrapure water a day, comparable to 33,000 households. That is a manufacturing footprint, not a data center footprint, and it belongs in a fifth column that almost nobody adds. We looked at the capital side of that build out in our piece on the memory chip investment wave, and at who is building the silicon in the comparison of custom AI chips against the incumbent.
Whether the whole build out is sized to real demand is a separate question with its own evidence, which we went through in the five indicators worth watching on the AI bubble question. Water does not settle it either way. A facility that turns out to be unnecessary still evaporated the water while it ran.
How should a resident read a planning proposal?
Ask for four numbers rather than one, and ask for them in writing. What is the projected withdrawal, in gallons per day, at full build out. What is the projected consumption, meaning the share not returned. Which source, aquifer or municipal supply or reclaimed water, and what happens in a drought year. And what the discharge contains, since concentrated cooling water is a water quality question rather than a volume question.
Two further things are worth requesting because they are frequently withheld. The first is whether any figure supplied is covered by a non disclosure agreement, which the Lincoln Institute identifies as routine. The second is the cooling design, because a closed loop system genuinely reduces on site consumption while raising electricity use, and a proposal that mentions the first without the second is telling you half of a trade.
What would honest disclosure actually require?
Three fields, published annually, per facility. Withdrawal in gallons, consumption in gallons, and the name of the source basin. Everything else in this article is downstream of those three, and none of them is commercially sensitive in any serious sense: a competitor learns nothing from knowing how much a building evaporated.
The counter argument you will hear is that per facility figures invite misleading comparison, since a building running at half occupancy looks efficient and a building running flat out looks profligate. That is true and it is solved the way every other utility disclosure solves it, by publishing the denominator alongside the numerator. Report the water and report the compute or the megawatt hours it supported, and the ratio becomes comparable across sites. Operators already track water usage effectiveness internally for exactly this reason.
Until that happens, the practical advice is unglamorous. Prefer local reporting to national commentary, because the local reporter has the permit application and the commentator has an average. Prefer a company's own environmental report to a summary of it, since the summary is where the measure gets dropped. And treat any figure without a year attached as unusable, given an industry that doubled twice in five years.
What to remember
There is no single true number for data center water, and there was never going to be. There are four defensible quantities, each with a constituency and a legitimate use, and nearly every public disagreement about AI and water is two people holding different ones.
The figures that anchor the current debate are worth memorising with their labels attached. About 449 million gallons a day consumed directly across US facilities in 2021. About 211 billion gallons a year consumed indirectly through electricity in 2023. A 2027 projection of 4.2 to 6.6 billion cubic meters of withdrawal for global AI. Attach the measure to the number every time you quote it, and the conversation gets shorter and considerably more useful.