Where $100B in training compute actually goes
Frontier labs now talk about training budgets in nine and ten figures. This is a line-item tour of where the money actually lands — chips, buildings, power, talent, and data — and why the biggest constraint is no longer the silicon.
When Anthropic's CEO said frontier labs would soon spend close to a billion dollars on a single training run — and up to ten billion on runs in the next two years — it sounded like a boast. It was closer to a budget forecast. The numbers behind modern frontier training have quietly crossed from "expensive software project" into "infrastructure program" territory, and the money flows into surprisingly old-economy line items: silicon, steel, electricity, and people.
This article is a line-item tour of where the money goes. One caveat up front: the famous headline figures — "GPT-4 cost $40 million to train," "Gemini cost $191 million" — are not three witnesses describing the same event. They are three different measurement systems. Epoch AI, which models amortized hardware, put GPT-4's final training run at roughly $40 million and Gemini Ultra at about $30 million. The Stanford AI Index, which prices the same compute at cloud rental rates, landed on roughly $78 million for GPT-4 and $191 million for Gemini Ultra. Add failed runs, experiments, and R&D salaries and the numbers move again: Epoch AI estimates total development compute runs 1.2 to 4 times the final run. Nobody is lying; they are measuring different things.
With that settled, here is where a roughly $100 billion frontier training program — a plausible multi-year budget for a leading lab covering hardware, facilities, power, talent, and data — actually gets spent.
The chips: ~$45–50B, the biggest single line#
Roughly half the money buys silicon. A single NVIDIA H100 sold for around $28,000 while costing an estimated $3,320 to manufacture; its successor, the B200, sells for around $40,000 on an estimated $6,400 manufacturing cost, with high-bandwidth memory (HBM) making up 41–45% of the bill of materials. At roughly $40,000 a chip, $50 billion buys about 1.25 million top-tier GPUs.
That scale is no longer hypothetical. xAI's Colossus cluster in Memphis grew from 100,000 to 200,000 H100s, Anthropic's Project Rainier with AWS will bring online more than one million Trainium 2 chips, and OpenAI's Stargate project began its first phase in Abilene, Texas, with a publicly discussed vision of 5 gigawatts.
The chokepoint worth knowing about isn't the GPU die itself — it is the memory and the packaging. TSMC's advanced chip-on-wafer-on-substrate (CoWoS) lines, which bond GPU logic to HBM stacks, are described by analysts as fully booked with 52-to-78-week lead times, and NVIDIA alone is estimated to hold about 60% of that capacity. In other words, half the budget goes to chips, and a large fraction of the delay risk lives in the least glamorous steps of making them.
An important paradox: cost per unit of compute keeps falling — roughly 10x over the last decade and a half — while compute per frontier run rose about a billion-fold. Cheaper chips, costlier models: the B200 cuts cost per FLOP by an estimated 40–60% versus the H100, but absolute capital requirements rose 4.7x per rack.
The buildings: ~$20–25B, because GPUs don't sit in the sky#
Chips are useless without the industrial facility around them. Building an AI-optimized data center now costs $20 million or more per megawatt of IT load — roughly double to triple the $7–12 million of a conventional build — and fully built-out hyperscale campuses are modeled at $45–55 billion per gigawatt. The four largest hyperscalers planned a combined $630 billion in capital expenditure in 2026 alone, a 62% jump from the prior year.
What drives the premium? Three things: power delivery (substations, transformers, switchgear, backup generation, costing $500 million to $2 billion for a gigawatt-scale site), cooling (liquid cooling loops, chillers, cooling towers at $1.5–2.5 billion, since dense racks now hit 60–120 kW where a standard rack draws 5–15), and the networking fabric — InfiniBand switches, NICs, and cabling at $4–8 billion per gigawatt — that keeps hundreds of thousands of GPUs acting like one machine. None of this depreciates on software timelines; it's concrete, copper, and steel.
The power: ~$5–8B of lifetime electricity, and the real hard limit#
On a single training run, energy is a surprisingly small slice: Epoch AI's cost model puts it at 2–6% of total training cost. The reason it deserves its own line item is that power is increasingly the binding constraint on everything else. You can print more money and book more wafers; you cannot print megawatts.
The scale: AI data centers are forecast to consume 175 terawatt-hours globally in 2026, up 84% from the prior year, with US AI data centers alone drawing about 68 TWh. At typical commercial rates around $0.12–0.15 per kWh, a one-gigawatt campus burns on the order of $100 million-plus per year just in electricity — and that assumes the grid connection, the substation, and the community approval all exist. Communities near proposed sites are increasingly pushing back on power consumption, water usage, and minimal local employment, which is why site selection has become a strategic function, not an operations detail.
The talent: ~$10–15B, because the scarcest resource has a name#
In Epoch AI's cost model, staff costs were one of the two largest line items for GPT-4 and Gemini-class models — each in the tens of millions of dollars, on par with the chips themselves. That was for a $40–80 million training run. Scale the program to $100 billion and the people bill scales with it, because the labor market for frontier AI researchers is brutally thin: expert estimates put the number of people worldwide capable of pushing the frontier at only about 2,000.
The 2025 talent war made the pricing public. Meta reportedly offered compensation packages of up to $100 million — including nine-figure signing bonuses — to lure researchers from OpenAI, with CEO Mark Zuckerberg personally recruiting. Meta later paid Apple foundation-models lead Ruoming Pang a package reportedly worth more than $200 million. The arithmetic is cold: if one researcher can meaningfully change whether a $500 million training run succeeds, paying nine figures to hire them is rational.
The data: ~$5–8B, and rising fastest#
This is the quietest line item and possibly the fastest-growing. The public internet's readily available text has been largely exhausted by existing models, so frontier labs now buy, license, and manufacture their training data. In June 2025, Meta paid $14.3 billion for a 49% stake in Scale AI — not for a model or a chip, but for a position in human judgment: the annotation pipeline that turns raw data into training examples.
The economics are stark at both ends. Bulk annotation in low-wage markets pays as little as $1.32–$2 an hour, while US-based expert labelers on AI training platforms command $25–$50 an hour for generalists and $112–$150-plus for specialists. Gartner projected synthetic data would account for over 60% of AI model development data by 2027, which trades annotation cost for generation compute and curation. Either way, "the data is free because it's on the internet" was a pre-2024 argument.
The multiplier: failed runs, utilization, and the 1.2–4x nobody budgets for#
Add up the five line items above and you get the cost of the final, successful training run — the one that makes the press release. Epoch AI's estimate that total development compute is 1.2 to 4 times the final run is the line item nobody's spreadsheet wants to admit: the failed runs, the scaling-law experiments, the aborted checkpoints. At frontier scale, a cluster that costs tens of billions to build also has to be kept busy; idle GPUs depreciate just as fast as working ones. Cloud rental rates of $2 to $8.50 per GPU-hour (H100 to B200, on-demand) exist precisely because utilization is the whole game.
The takeaway#
If you had to summarize where $100 billion in frontier training compute goes, it looks roughly like this:
| Line item | Share of program | What it buys |
|---|---|---|
| Accelerators | ~45–50% | 1M+ GPUs; HBM and CoWoS are the chokepoints |
| Facilities + networking | ~20–25% | $20M+/MW builds; liquid cooling; InfiniBand fabric |
| Talent | ~10–15% | A few thousand researchers; nine-figure packages at the top |
| Data | ~5–8% | Annotation, licenses, synthetic generation |
| Power (lifetime) | ~5–8% | ~$100M+/yr per GW; the binding constraint |
| Failed runs + reserve | remainder | The 1.2–4x development multiplier |
Two conclusions matter for anyone watching the frontier. First, the cost structure of training a frontier model now looks like the cost structure of building a power plant or an airline fleet: capital-intensive, long-lead-time, and dominated by physical constraints. Second, the binding constraint has moved from silicon to electricity. Compute per dollar keeps falling; watts do not. The next era of the AI race will be won not by whoever designs the best chip, but by whoever secures the gigawatts — and the people — to run them.