The 9,396 records in this section are overwhelmingly reconstructions, not disclosures. This page says how each kind of number is arrived at and how much to trust it.
All of it comes from Epoch AI, an independent research institute that tracks the trajectory of AI. They publish their datasets under a CC BY 4.0 license, which permits reuse, redistribution and reproduction provided the source and authors are credited. That credit appears on every page of this section, next to the data it applies to.
Nothing here is scraped from their site. Each dataset is taken from the files they publish for download, normalized into a database, and re-published as individual pages. The research is theirs; the structure, the cross-linking and the prose are ours.
Training compute is almost never published by a lab. It is reconstructed either from the hardware used (chip count multiplied by throughput, training duration and a utilization factor) or from the architecture and dataset size. Each model record names which method produced its figure.
Data-center capacity and power come from construction permits, satellite imagery, power-purchase agreements, chip orders and reported spending. These carry roughly 1.4x to 1.6x of uncertainty depending on the metric — a site listed at 900 MW could plausibly be 600 or 1,400.
Chip supply and stock are modelled from earnings reports, shipment data and supply-chain capacity. They are published as a median with a 5th and 95th percentile band, and all three are stored rather than collapsed to the middle value.
Company financials come from press reporting, not filings — most of these companies are private. Different outlets report different figures for the same quarter, so each report is kept as its own dated point instead of being averaged into a single series. The headline figure on a company page is simply the most recent report, and it says which date it is from.
Epoch publishes a street address for each data center but no coordinates — their own map resolves them in the browser. The compute map therefore geocodes those addresses against OpenStreetMap, and records how far down each one resolved: to the building, the street, the district, the town, or only the county. A handful of sites list a construction description rather than a postal address; those are placed from the town named in the record and are never presented as more precise than that.
Each candidate match is checked against the country and, in the United States, the state parsed from the address, and rejected if it disagrees — without that check a matching street name in the wrong state will silently relocate a gigawatt site. GPU cluster coordinates come from Epoch’s own export and are used as published.
Attributions — who owns a data center, who uses its compute — carry an explicit confidence tag: confident, likely or speculative. Those tags are kept on the record and displayed, because “Microsoft probably owns this” and “Microsoft owns this” are different claims.
The same applies to model records, which carry a confidence rating on the whole entry, and to GPU clusters, which carry a certainty field and often a note explaining what is disputed. Where a figure has a percentile band, the band is stored; where a figure is missing, the page says nothing rather than guessing.
Epoch republishes often — some datasets change weekly. This site re-syncs from their published files and every record stores the date it was retrieved, which is what the source line at the bottom of each page reports. If a number here disagrees with theirs, the retrieval date is the first thing to check.