Making Tokens, Pt. 4: A City's Worth of Power
Part four of the Making Tokens series. By now we have accelerator chips, finished and packaged (Part 3). The next step is connecting them into a training system. This is where the power, cooling, and capital stop being abstract.
Pick a reference system
Use the H100 SXM as a reference, rather than calling it the universal frontier configuration in 2026. NVIDIA's specification lists 80 GB of HBM3, up to 700 W configurable thermal design power, and about 989 TFLOPs of dense FP16/BF16 tensor compute. The larger advertised 1,979 TFLOPs figure uses sparsity. Peak throughput is not sustained training performance.
Eight GPUs at 700 W total 5.6 kW for the GPUs alone. A whole server must also power CPUs, memory, fans, storage, and networking. NVIDIA rates the eight-GPU DGX H100 system at 10.2 kW maximum. That is a system rating, not a promise that every workload draws that much.
Scale the example
Suppose we deploy 100,000 of those GPUs. This is a calculation scenario, not a description of a named lab's training run:
- GPU power at the chosen rating: 100,000 × 700 W = 70 MW.
- IT load assumption: 110 MW, including the rest of the computing equipment.
- PUE assumption: 1.3. PUE is total facility power divided by IT power.
- Facility power: 110 MW × 1.3 = 143 MW, rounded to 150 MW for the examples below.
The electricity demand is substantial. For perspective, 150 MW is 15% of a 1,000 MW reactor's rated output. Capacity factor affects annual energy, not that nameplate comparison. It is not the output of 1.5 large nuclear reactors. See the Department of Energy's explanation of capacity factor.
Cooling does not always mean consuming water
Liquid cooling carries heat away from the chips. How the facility ultimately rejects that heat determines water consumption. Evaporative towers use water; dry heat rejection trades different power and temperature constraints. Water circulating in a closed chip-cooling loop is not automatically water consumed.
To keep the arithmetic inspectable, assume a site water usage effectiveness of 1 liter per kWh of IT energy. This is a chosen scenario, not an industry average. At 110 MW IT load:
110,000 kW × 24 hours × 1 L/kWh = 2.64 million liters per day, approximately 697,000 US gallons.
Change the cooling design, climate, or assumed WUE and the result changes. Electricity generation can also consume water upstream. That belongs in a separate boundary rather than being silently added to site consumption.
The energy bill
A facility drawing a constant 150 MW consumes:
- Annual energy: 150 MW × 8,760 hours = 1.314 TWh.
- At an assumed $0.05/kWh: $65.7 million per year.
- At an assumed $0.08/kWh: $105.12 million per year.
These are electricity scenarios, not quotes for a particular campus. Average load and local tariffs determine the actual bill.
For a 730-hour month, GPU energy at the chosen rating is 51.1 GWh. Facility energy at 150 MW is 109.5 GWh. At $0.05/kWh, the latter costs $5.475 million. Calling GPU-only energy the entire training electricity bill would leave out the rest of the system.
Capital is a separate ledger
Assume $30,000 per accelerator for a budgeting example: 100,000 × $30,000 = $3 billion. This is not a current market quote. Servers, networking, storage, land, buildings, power connections, and cooling add other costs.
If those accelerators are depreciated over four years, that assumption gives $750 million per year, or $62.5 million per month, before the other assets. Depreciation is an accounting allocation; it is not another electricity bill. A cloud renter instead faces rental pricing and utilization constraints.
There is no defensible universal "$500 million to $1 billion" training-run price here. A specific estimate needs the run duration, hardware, utilization, rental or ownership basis, and treatment of experiments and failed runs. Publicly announced campus budgets do not supply all of those inputs.
Where a cluster can go
Power delivery, interconnection timelines, cooling, local permits, and available land constrain siting. Cheap electricity is useful only if the site can get enough of it when needed. Inference also has latency constraints that training can sometimes avoid.
Do not collapse distinct projects into a regional list: Memphis is in Tennessee, not west Texas. And an announced site is not the same thing as commissioned capacity. The physical details determine when the compute actually becomes usable.
What's left
The result of training is a set of weights, together with the software and other artifacts needed to use them. Their size depends on parameter count and storage format. The upstream investment does not make a finished answer; it makes a model that still needs serving infrastructure.
In Part 5, we finally get to the cost of producing tokens. We will keep measured performance separate from a calculation assumption there too.