A building-level dataset of 2,696 United States data centers shows that more than one third lie within five kilometers of high wildfire hazard land
DOI:
https://doi.org/10.13021/jssr2026.5587Abstract
Artificial intelligence and cloud computing have concentrated computing capacity and electrical load into a small number of United States locations, which leaves critical digital infrastructure exposed to natural hazards and straining regional grids. Assessing that exposure requires knowing where each facility physically sits. Open sources are incomplete and unverified, and commercial directories are proprietary or aggregated above the building level, so no open dataset supports building-level hazard analysis. This study assembled a reproducible, public dataset of 2,696 data centers across the contiguous United States. A Python pipeline merged OpenStreetMap, PeeringDB, and Wikidata records, resolved duplicates, and graded every coordinate against an independent building footprint source and satellite imagery. Official hazard maps then supplied earthquake ground motion, lightning flash density, and wildfire hazard at each location. Virginia alone holds 409 facilities, which is 15 percent of the national total. Independent footprints place 1,546 coordinates on a building, individual review resolved a further 416, however 184 still remain unmatched. Sampling wildfire hazard at the building itself proved to not be very helpful, because 85 percent of facilities occupy land that the hazard map classifies as developed. When measuring across the surrounding area instead, 342 facilities fall within one kilometer of land rated High or Very High, and 953 within five kilometers. That exposure concentrates in Arizona, New Jersey, and New York rather than California, so wildfire risk to digital infrastructure may be underestimated outside of the states that are typically associated with it. Building-level location data of this type is what multi-hazard risk assessment of computing infrastructure has been missing.


