Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,476 characters · 29 sections · 37 citation commands
Certificates without Electrons? Theory and Evidence on Impacts from AI-Driven Power Demand
\let\WriteBookmarks\relax
\shorttitle{Certificates without Electrons}
\cormark[1] \ead{[email removed]} \credit{Conceptualization, Methodology, Software, Data Curation, Formal Analysis, Writing -- Original Draft, Writing -- Review & Editing, Visualization}
\ead{[email removed]} \credit{Conceptualization, Writing -- Review & Editing, Supervision}
\ead{[email removed]} \credit{Conceptualization, Writing -- Review & Editing, Supervision}
\cortext[1]{Corresponding author}
Over the past six years, data centers have expanded from consuming 1.9% to 4.4% of total U.S. electricity demand (Figure (ref), DOE). This rapid growth—driven largely by artificial intelligence—follows two decades of flat electricity consumption and coincides with nationwide efforts to electrify transport and heating. The combination poses a challenge: grid infrastructure was planned for stable demand, not simultaneous load growth from multiple sectors. Recent anecdotal evidence points to data centers reducing local power quality and raising prices Nicoletti_Malik_Tartar_2024, yet systematic empirical evidence on these grid-level impacts remains scarce.
Existing work does not provide causal estimates of how AI deployments affect power systems at the {\em grid level}. Techno-economic assessments from international organizations broadly characterize rising AI workloads in the data center market RN1, RN3. Micro-level studies examine energy performance of AI accelerators and identify workload management opportunities RN9, RN4. A large computer science literature quantifies the compute, energy, and environmental costs of model development and deployment strubell19, schwartz2020green, RN7, RN5, morrison2025holistically. These provide a foundation for understanding datacenter-level dynamics, but do not address the externalities experienced by other grid users—impacts on power quality, prices, and reliability that matter for energy security.
Quantifying and understanding these power grid level impacts is critical from the perspectives of energy reliability and security. In this paper, we ask seek to answer four specific questions: 1) What are the effects of AI model releases on local power quality and power frequency? 2) How do AI training and inference affect electricity consumption around data centers owned by model developers? 3) How do these impacts translate into price effects in the wholesale and retail markets? Additional demand will likely increase prices, but understanding the magnitude of these price increases is important to determining the severity of the issue. 4) How will these impacts be different based on counterfactual scenarios articulating different technical evolutions of AI computing at scale? These scenarios include a shift away from cloud-based inference to edge-based computing, a geographic realignment of data centers away from population clusters, an evolution towards primarily inference-based or primarily training-based workloads, and a significant improvement in overall energy efficiency.
One way to answer the above questions is to extrapolate from energy usage measurements (micro measures) within a datacenter to system-level (macro measures) impacts on the grid. This is however, is near infeasible because it requires a great deal of specific internal information that isolates what is run within the datacenter and when. While this data with level of detail is not available (for proprietary and other practical reasons), there are data sources that provide high-level observational data about various aspects of power consumption and quality at the grid-level over the time periods when models were trained and released. How can we answer specific questions about the impact of AI models with this high-level observational data alone?
We can draw upon econometric approaches for addressing this problem. In particular, we focus on the Difference-in-Differences (DiD) models which allows us to estimate the causal effects of AI model deployments on power systems, even when some contributing factors are unobservable. By comparing areas around data centers running AI models against control areas, and partitioning observations into pre- and post-release periods, we isolate AI-specific impacts while controlling for temporal and geographic factors.
There are multiple technical challenges in applying this standard DiD approach in our setting: (i) There is a lack of specific information about when and where models are trained and released. Further model releases can be staggered, affecting the validity of the standard model which only uses one timepoint. (ii) The DiD method at its core relies on the parallel trends assumption i.e. we would know what would happen if the AI model were not run; in other words the estimates are valid only if we have a fairly accurate estimate of the trend of the variable of interest over time. (iii) There could be other confounding factors contributing to change in the variable of interest (e.g. extreme weather events coinciding with model release times could also cause increased power demand). (iv) No one clean dataset exists including all data useful for analysis. Often data are at different scales, lack linking tables, or require significant processing prior to use for analysis. In some cases, this requires significant efforts for manual mapping, combination, and labeling.
Our solution methodology addresses the above challenges:\\ 1) Addressing Inexact Information and Staggered Releases We use a Staggered DiD approach with multiple horizons, which has been used to overcome similar challenges in other applications cengiz2019effect, deshpande2019screened. Staggered DiD with multiple horizons extends the standard two-period (pre/post), two-group setup to settings where treatment activates for different units at different times, leveraging this variation in rollout timing to assess different potential points of impact—analogous to how a staged feature rollout lets you compare early adopters against users still in the control group at each point in time. For the question on fossil fuel demand effects, we address confounds such as alternative fuel prices or generator efficiency issues using a two stage regression, where we first factor out the impact of these confounds, and explain the remaining demand using a second stage staggered DiD. Similarly, for estimating price effects, we need to isolate effects from other sources of power demand that can cause price changes. We handle confounds such as weather-driven and industry production driven shifts in demand (and thus price) via a two-stage regression.
\noindent2) Addressing validity of the DiD method and other confounds We conduct a wide-array of statistical robustness checks supported by integration of large amounts of data from a diverse array of sources. We run statistical tests to show that we cannot reject the null hypothesis of parallel trends. We also run placebo tests with random treatment dates and randomized locations to show that our treatment effects are not the result of randomness, and we vary treatment definitions geographically and temporally. We use time series testing for anomaly detection and structural breaks to show that the differences we find in our regressions show up in model free analysis. Together, these robustness checks provide evidence of an experiment that is valid across all relevant dimensions.
\noindent3) Addressing data disparity and relevance issues Training dates and facilities are generally unobserved, and for inference—which occurs continuously—these dates are not well-defined. We therefore rely primarily on model release dates. To strengthen identification, we supplement this with training locations, training dates, and API release dates gathered from press releases and news articles. We obtain verified training information for four models and confirmed API release dates for over half of the foundational models in our sample. We construct estimates using this most robust subset, then expand our model definitions to exploit exogenous variation for counterfactual analysis. Throughout, we use the most conservative feasible treated subset for each analysis. To construct suitable control groups, we employ propensity score matching as detailed in Appendix Subsection (ref), excluding observations without comparable pre-treatment trends.
We apply this methodology using publicly available data and proprietary on datacenter locations, power grid conditions, AI model release dates, and controls including weather and market prices. Our data creation process is described in detail in Section (ref). While all data pre-exist our analysis, the combination of the data creates a novel dataset that represents a significant contribution. Because the data are from government agencies, ISOs, or proprietary datasets, the data are highly trustworthy. We perform data validation to verify this. We estimate DiD models for outcomes measuring both power quality and electricity demand.
Our findings presented in Section (ref) speak to the large magnitude of the impacts of AI data centers and the increasing effects of training ever-larger models over time. We find that large model releases (including the likes of GPT-3, Claude 2.1, Llama 2 documented in (ref)) cause significant power quality deterioration in nearby areas -- well above half the standard deviation of U.S. power-quality distributions-- and increase fossil fuel consumption on the order of terawatt-hours--- equivalent to around 45,000 households annual consumption-- in both training and inference phases. We also find wholesale electricity prices increase in treated PJM zones with increases of 25% during treatment in the American Electric Power (AEP) PJM Zone. Using these estimates, we show in Subsection (ref) how altering the trajectory of AI along disparate dimensions impacts the grid. We are able to provide evidence for the separate impacts of training and inference, the effects of a shift to edge computing or more concentrated training loads, and how differential efficiency improvements impact AI's effects on power usage. We also show compelling evidence that significantly increasing on-site power generation provides a promising avenue for future reductions in grid impact.
This work makes three contributions. First, we develop a methodology for assessing macro-level grid impacts of AI models using publicly available data, even when fine-grained operational information is unavailable. Second, we provide the first causal estimates of these impacts for major frontier models, complementing the emerging literature on environmental and system-level effects of AI data centers murino2023sustainable, guidi2024environmental, thangam2024impact. Finally, we provide policy-relevant counterfactual estimates for the impacts of technical evolution of AI on the grid.
This paper contributes to four interrelated literatures. First, techno-economic assessments document the scale of data center electricity demand: Lawrence Berkeley estimates U.S. data center consumption could reach 6.7--12% of national electricity by 2028 shehabi2024, while utility five-year peak demand forecasts jumped from 38 GW to 128 GW between 2023 and 2024, with approximately 90 GW attributable to data centers gridstrategies2025. Recent simulation studies examine AI-driven load growth implications for grid planning lin2024exploding, chien2023adapting, but these assessments characterize aggregate trends rather than isolating causal mechanisms.
Second, computer science research quantifies AI's computational and environmental costs, from training emissions strubell19, schwartz2020green to debates over efficiency improvements RN2, RN8 and inference-time scaling kim2025cost, including data center carbon emissions datacenter-carbon. Related work has developed carbon intensity forecasting methods maji2022dacf, li2023gnn and examined power forecasting for grids optimizing-grid, carbon-intensity-forecasting, but not in the context of AI workload impacts on market outcomes. Recent empirical projections analyze the tension between AI energy growth and grid decarbonization maji2024crossroads, while others consider grid planning for AI demand ai-grid-planning, yet the causal effect of AI deployment on wholesale markets remains unstudied.
Third, emerging evidence points to grid-level externalities from data center concentration. Sensor data reveals spatial correlation between data center proximity and power quality degradation nicoletti2024, PJM capacity prices increased from \$30 to \$270/MW-day between 2023--2024 amid data center load growth cmu2025, and transmission costs are increasingly socialized across ratepayers jacobs2025. The systems literature has examined carbon-aware geographical load shifting using locational marginal prices lindberg2021guide and data center demand response participation chen2019datacenter, liu2014pricing, klingert2018mapping. These studies establish correlational patterns or propose optimization frameworks but lack causal identification.
Fourth, the econometric literature on electricity markets provides methodological foundations, including difference-in-differences for demand responses roth2024did and event studies for policy interventions pnnl2022, but has not applied quasi-experimental methods to AI deployment impacts.
We address this gap by treating AI model releases as plausibly exogenous demand shocks, using difference-in-differences to estimate causal effects on locational marginal prices and power quality. Unlike prior work building from computational requirements or facility audits, we analyze grid-level impacts using wholesale market data, capturing realized demand effects without requiring proprietary operational information.
Our goal is to investigate three primary effects of AI data centers on power markets. We also conduct counterfactual analyses that could inform policy: assessing different scenarios of parameter size, training versus inference proportions, edge compute preference, relocation to low-population areas etc.
First, we quantify the effect of frontier AI models on power quality. AI workloads affect power quality through three main mechanisms. First, data centers and AI computing facilities add substantial baseload demand that can strain grid capacity and compromise voltage regulation, especially during peak periods. Second, AI workloads fluctuate rapidly with computational needs, creating unpredictable load variations that challenge stability and frequency regulation. Third, the switching power supplies and electronics essential to AI hardware generate harmonic distortion, introducing waveform distortions that spread through distribution networks and degrade power quality. See Figure (ref) in Appendix for an illustration of this harmonic distortion.
Second, we analyze how training and inference workloads differentially affect local power demand given that training produces sharp, episodic spikes while inference generates steadier ongoing consumption.
Third, we estimate how wholesale electricity prices have responded to LLM development. AI workloads affect wholesale electricity prices through the demand channel, with potentially indeterminate long-run retail market impacts. Wholesale prices are determined hourly in day-ahead markets via sealed-bid, uniform-price auctions subject to network constraints. Locational price variation reflects transmission limitations within the network. Increased demand raises wholesale prices by necessitating dispatch of higher-cost marginal generators. The wholesale price increases are passed through to retail consumers with substantial lag. Understanding wholesale demand and price shifts thus provides useful intuition regarding the direction and magnitude of retail price changes.
We want to quantify the effects of running AI models on the power grid but we only have overall observational data about power demand, power quality, prices, exogenous controls, locations of data centers, and dates of model activity. We do not have access to data that specifically isolates the effects with respect to running the AI models because these require proprietary knowledge.
To address this challenge, we can turn to Difference-in-differences (DiD), a widely-used statistical estimation technique in econometrics for causal inference with observational data. This approach has been used to answer a wide range of empirical economics questions from observational data such as telecom demand from price and quantity variation, fowlie2012industrial evaluate environmental regulation in electricity using emissions and price data, and card1994minimum study minimum wage effects via cross-state employment variation.
The key idea behind this technique is to estimate treatment effects by comparing the change in outcomes over time for a treated group to the change for a control group, as illustrated in Figure (ref).
For power quality, as an example, this translates to estimating the change in utility-level power quality within the geographically-defined utility area before and after the release of a major AI model. The treated group in this case would be utility areas containing data centers engaged in model training or inference during the time period of activity, and the control group would be utilities without AI activity in the same time period. The outcome variable, power quality in our example, is modeled as a function of the treatment (i.e. the running of the AI model), and other contributing factors. Formally, we can specify the power quality effect as:
where $Y_{it}$ is power quality for utility $i$ at time $t$, $\text{Treatment}_i$ indicates proximity to a data center, and $\text{Post}_t$ captures periods after the release of a major AI model. The coefficient of interest, $\beta$, measures the causal impact of model releases on local power quality. $\gamma_i$ captures unobserved factors common across all utilities, $\delta_t$ captures unobserved factors that remain fixed across time, and $\varepsilon_{it}$ is an error term that captures random disturbances or shocks.
We can learn the co-efficients, given appropriate data about the outcome variables over time --- in this case with data about the power-quality values over time across a range of utilities containing AI data centers, and the knowledge of when and where the AI models of interest are run.
The validity of this basic DiD approach for our problem is affected by the following issues:
We address the first two challenges with a stacked modification of the DiD model ((ref)). We address the next two through appropriate statistical tests and robustness checks ((ref)) and the final issue through our data creation (section (ref)).
To address the inexact information about training windows and release times, we estimate dynamic treatment effects at multiple horizons relative to model release dates. This approach, referred to as an event study or stacked difference-in-differences in econometrics, offers three advantages over the standard pre/post comparison discussed earlier: (1) it does not require observing the true training start date, instead using release timing as the anchor; (2) it also provides a direct test of the parallel trends assumption through examination of pre-release coefficients; and (3) it allows the data to reveal the temporal profile of effects rather than imposing an arbitrary discontinuous treatment structure.
Formally, we define treatment exposure relative to each model's release date $r_m$ for both training and release: 1) Training windows: The period $[r_m - \bar{\tau}_{\text{train}}, r_m)$ preceding release, during which computational resources are devoted to model training. We remain agnostic about the exact training start date; our identifying assumption is that training activity is elevated in the months before release and that any pre-existing differential trends would manifest in the earlier leads. 2) Inference window: The period $[r_m, r_m + \bar{\tau}_{\text{infer}}]$ following release, during which the model is deployed for inference via API access or public availability.
Model releases occur at different calendar dates, creating a staggered adoption setting. Recent econometric literature has demonstrated that conventional two-way fixed effects (TWFE) estimators as is the case in the estimates from Equation (ref). can produce biased estimates under treatment effect heterogeneity goodman2021difference, sun2021treatment, callaway2021difference, borusyak2024revisiting. Treatment effect heterogeneity means the policy works differently for different places or periods, which can distort standard difference-in-differences estimates. The bias arises because TWFE implicitly uses already-treated units as controls for later-treated units, and differences in treatment effects across cohorts or over time contaminate the estimated average.
In our setting, consider regions as units and AI model releases as treatments inducing increased data center power demand. A "cohort" comprises regions whose data centers began significant AI workloads at the same time. Suppose Region A adopted AI workloads in 2022 while Region B adopted in 2023. When estimating treatment effects for Region B, TWFE implicitly uses Region A as a control—but Region A in 2023 is not untreated; its data centers have been running AI workloads for a year. If treatment effects evolve over time (e.g., power demand grows as AI infrastructure scales), comparing Region B's outcomes against Region A's already-treated outcomes introduces bias. Staggered design provides heterogeneity-robust estimators that solve for this exact issue.
To address this, we use a stacked DiD model cengiz2019effect, deshpande2019screened, where we construct a stacked dataset where each model release defines a separate “sub-experiment.” The idea here is to construct, for each treatment cohort, a separate sub-experiment using only units that are genuinely untreated at the time of that cohort's adoption. By stacking these sub-experiments—each with its own clean control group of not-yet-treated regions—we ensure that no already-treated unit ever serves as a control, thereby eliminating the bias that arises when treatment effects are heterogeneous across cohorts or time.
Formally, for each release event $m$, we create a dataset containing: (i) Treated units -- Regions linked to the releasing company's AI data centers. (ii) Control units -- Regions with no AI data center exposure during the event window, or regions whose treatment occurs sufficiently far from event $m$ that they serve as clean controls.
Following \, we then stack these sub-experiments and estimate:
where $\gamma_{u}$ are unit-by-event fixed effects, $\delta_{u}$ are time-by-event fixed effects, $\mathbf{X}_{ut}'\theta$ represents controls, and the coefficients $\{\beta_k\}$ trace out the dynamic treatment effect at each event-time horizon $k$. Standard errors are clustered at the unit level to account for correlation across events within the same region.
Our power quality analysis estimates the following stacked DiD specification at the utility-month level:
where:
We estimate an analogous specification for total harmonic distortion (THD):
Power demand analysis proceeds at the generator-month level, exploiting finer geographic resolution. We maintain the stacked DiD structure while incorporating instrumental variables to address the price endogeneity concern. Price endogeneity arises because electricity prices and demand are simultaneously determined—high demand pushes prices up, while high prices suppress demand—making it impossible to identify the causal effect of price on consumption from their correlation alone.
Two-stage least squares resolves this by isolating variation in price that is plausibly exogenous: in the first stage, we predict prices using an instrument that affects prices but has no direct effect on demand; in the second stage, we regress demand on these predicted prices rather than observed prices, recovering an unbiased estimate of the price elasticity.
\paragraph{First Stage.} We instrument for electricity prices using supply-side cost shifters:
where $\text{HeatRate}_{it}$ captures the thermal efficiency of generator $i$ and $\text{GasPrice}_t$ is the contemporaneous natural gas spot price. These instruments satisfy relevance (they are primary determinants of marginal generation costs) and exclusion (they affect demand only through their impact on prices, not through direct effects on AI workload scheduling or other consumption decisions).
\paragraph{Second Stage.} We estimate the stacked DiD specification using predicted prices:
where $D_{it}^k$ indicates event-time relative to model releases linked to AI data centers in generator $i$'s service area.
The coefficients $\{\beta_k^{FD}\}$ capture the dynamic effect of AI activity on fossil fuel demand. Negative values of $k$ (pre-release) correspond to the training phase; positive values correspond to inference. The price coefficient $\eta$ controls for supply-driven price movements and provides a scaling factor for welfare calculations.
Beyond quantity effects, we estimate how AI-driven demand increases affect electricity prices. The key economic insight is that price impacts depend on supply flexibility. In other words, when generators can easily expand output, increases in demand (e.g. AI-driven demand) can be easily absorbed with minimal increases in price. When the demand suddenly spikes (i.e. a demand shock), the prices will increase sharply as well. This is characterized using supply elasticity $\varepsilon^s$—the percentage change in quantity supplied per one percent price change—quantifies this responsiveness.
Our estimation proceeds in two steps. First, we use our IV estimates to quantify how AI activity shifts demand. Simply put, the average change in total fossil demand from AI demand is the DiD coefficient times the treatment indicator:
Second, we estimate a linear supply relationship using demand controls as instruments and supply shifters as controls. While power markets involve optimal dispatch, the fossil-fuel portion of PJM’s marginal cost curve is approximately linear over the operating range we study PJM_DataMiner2_2025. Let $p_i$ denote the market-clearing price in market $i$, and let $Q^s$ denote the supplied quantity. We model supply locally as linear in price, $Q^s = a + b \cdot p_i$, where the slope parameter $b$ captures the responsiveness of supplied quantity to price changes (i.e., how flexibly supply can adjust to incentives). Rearranging, the implied price response to a demand shift $\Delta AI_{it}$ is given by $\Delta p = \hat{\beta} \cdot \frac{\Delta AI_{it}}{b}$. Expressed in percentage terms, the price impact depends on the flexibility of supply with respect to price:
where $\varepsilon^s$ denotes the percentage change in supplied quantity per percentage change in price.
We estimate $\varepsilon^s$ via IV regression of quantity on price in log-log form, using demand-side instruments (weather-driven load variation and industrial production indices).
We focus on the PJM operator at the zonal level for this analysis. We restrict to a single ISO (operator) because differing market structures make prices incomparable across ISOs, and choose PJM because it is the largest (65 million people, 13 states), includes key data center states (Virginia, Pennsylvania), and provides 20 diverse pricing zones. Because observed training locations are outside PJM, this analysis uses the broader model set only.
We implement several diagnostic and robustness tests to assess the credibility of our estimates.
Our empirical strategy balances internal validity with external relevance. Our main sample comprises four models for which we observe exact training and inference dates as well as training location, and we employ propensity score matching (described in the appendix) to support the parallel trends assumption. We then report results for progressively larger samples of models and data centers, accepting weaker identification in exchange for broader coverage. This approach enables precise, well-identified estimates from our core sample while exploring heterogeneity and counterfactuals in extended samples, with appropriate caveats regarding the latter.
Our primary test of the identifying assumption examines whether pre-release coefficients $\{\beta_k : k < 0\}$ are jointly indistinguishable from zero. We report:
Significant pre-trends would undermine the causal interpretation of post-release estimates.
To verify that our estimates reflect genuine treatment effects rather than spurious correlations or specification artifacts, we conduct placebo tests with randomly assigned treatment:
Given uncertainty in linking models to specific training locations, we assess robustness to alternative treatment definitions in our less conservative analysis:
We examine whether treatment effects vary with observable model characteristics:
Heterogeneity analysis both informs mechanisms and serves as a specification check: if effects are concentrated among larger models or models from companies with substantial AI infrastructure, this lends credibility to the AI-driven interpretation.
We conduct three primary sets of counterfactual exercises to investigate how alternative evolutionary paths for AI compute might affect grid-level outcomes.
Model Scaling. First, we examine the grid implications of continued growth in AI model size. Using the Epoch database to characterize trends in model parameters, we provide descriptive analysis of how increasing model scale translates into grid-level impacts. We then relate our estimated power demand coefficients to power quality coefficients, enabling analysis of how changes in training and inference energy efficiency propagate to power quality outcomes.
Geographic Redistribution. Second, we leverage the geographic locations of data centers in conjunction with our econometric estimates to conduct counterfactuals focused on spatial reallocation of AI compute. We consider two scenarios. In the first, we simulate a shift from cloud-based to edge-based inference—reallocating inference demand from concentrated hyperscaler facilities to smaller data centers distributed closer to population centers. We implement this by redistributing our estimated inference demand impacts across data centers not currently classified as AI-focused. In the second scenario, we simulate relocating training workloads from population-dense regions on the East Coast to energy-abundant regions in the center of the country. For both geographic counterfactuals, we provide sensitivity analyses in the appendix to account for potential increases in communication costs and efficiency losses.
Colocated on-site Power Generation. Finally, we exploit a new feature in the Aterio data, namely a flag on data centers with on-site power generation. We re-estimate both our main models separately for generators with and generators without onsite power generation. We then report the impacts to power demand and power quality of either removing all onsite generation or having onsite generation for all data centers.
Together, these counterfactuals illuminate both how the grid might be affected by different trajectories for AI compute and which levers—geographic placement, workload composition, efficiency improvements—may be available to computer scientists and policymakers seeking to mitigate grid impacts.
Table (ref) summarizes the datasets we use to address each of our three research questions, distinguishing between variables of interest and the control variables described above. For variables of interest, we rely on Aterio, which provides comprehensive information on the location and history of data centers. The Aterio dataset documents when and where centers have been established, expanded, or retired, along with details on ownership and operational capacity (in MW). This allows us to identify data centers owned by specific hyperscalers and to quantify their scale over time. We then combine these data with external sources that provide the necessary controls. The data centers we use include ones that we could verifiably trace as having been the source of training of well-known large scale models such as GPT-3, PaLM, and Llama 2, among others.
We analyze four primary outcomes. The Consumer Power Quality Index (CPQI) from WhiskerLabs_THD_2024 provides a composite measure of consumer-facing reliability events (surges, sags, brownouts, interruptions), summarizing frequency and severity of power deviations at the household level. Total Harmonic Distortion (THD) data from the same source offer a technical measure of waveform distortion, quantifying voltage deviation from a clean 60-Hz sine wave; elevated THD indicates grid stress, reduces motor efficiency, and shortens equipment lifespan (Figure (ref)). Since THD data are more recent than CPQI, fewer model releases are available for THD analysis. We incorporate generator demand from EPA's CAMPD database, providing hourly plant-level demand spanning 2021–2023. Finally, we combine wholesale zonal price data from Independent System Operators (ISOs) to assess AI impacts on prices.
Our power quality regressions use a parsimonious specification with time and geographic fixed effects. Generator demand regressions employ a richer specification: we instrument for endogeneity using generator heat rates and natural gas prices (S&P Capital IQ Global) in a two-stage least squares framework, and control for weather conditions (Meteostat) to capture demand and renewable variability. EIA retail service territory files reconcile generator-, zonal-, and retail-level observations.
To integrate datasets, we harmonize spatial and temporal units across sources. Figure (ref) demonstrates the integration at each step for our modeling purposes. Data center locations from Aterio are mapped to ISO zones, EIA territories, and counties, enabling linkage with Whisker Labs reliability indexes (CPQI and THD) and CAMPD generator demand. We are the first study to integrate Aterio’s new field indicating which data centers are utilized in processing AI workloads. CPQI has 2,683 observations (mean 0.52, s.d. 0.42), while THD is more dispersed (mean 1.81, s.d. 6.77). CAMPD generator demand averages 218 MW (s.d. 158), and wholesale prices average \$51/MWh (s.d. 148). Weather data capture local conditions (mean temperature 17°C, precipitation 0.12 mm), and time series are aligned at hourly or monthly resolution. External controls—natural gas prices, weather, and ISO market data—are merged by geography and time, yielding a unified panel suitable for difference-in-differences and IV regressions.
Our demand model focuses on deregulated markets in the Eastern Interconnection and ERCOT, where data on zones and prices are readily available, while our power quality model draws on Whisker Labs data from 72 retail utilities nationwide (monthly, 2022–2025). The demand regressions use hourly data for several thousand generators (2021–2023). We concentrate on hyperscaler AI companies—including Meta, Microsoft, Amazon, and Google—with Anthropic linked to Google due to its cloud partnership. Appendix materials include maps, sample descriptions, and data tables. In order to determine which data centers should be considered AI data centers and should be included in our sample, we perform a linkage procedure described in appendix Subsection (ref).
We reports results using our stacked DiD estimator ((ref)) which is designed to handle staggered AI model releases. For all power quality and fossil demand effects, we first construct cohort-specific datasets following callaway2021difference, for each of the treatment events corresponding to major AI model releases through 2023 ((ref)). Table (ref) provides the full set of results from our main regressions including analysis of the impacts of on-site colocated generation for our counterfactuals.
We report the power quality effects of AI data centers in seventy one retail electric utility service territories estimated from power quality observations over 1136 utility-months. We use Consumer Power Quality Index (CPQI) and Total Harmonic Distortion (THD) as outcome variables in our stacked DiD models ( (ref)).
Panel A in Table (ref) reports results from the utility-level strict DiD specification. We use strict here to distinguish that due to the sample size, we prefer the traditional DiD as our main specification rather than the stacked alternative. The estimated CPQI coefficient speaks to the reliability of the power supply to the consumers. The CPQI coefficient of 0.3795 indicates a power quality deterioration level that corresponds to moving from a typical U.S. power area toward the bottom quartile of power quality—roughly the difference between $\sim$ 1 outage/year and $\sim$ 1.5–2 outages/year for the average customer. Other reliability events such as brownouts and surges are predicted to increase by 1-2 per year. Importantly, power surges are a known ignition source for electrical fires, contributing to the approximately 51,000 annual residential electrical fires in the U.S., with electrical arcing accounting for 74% of ignition sources usfa2019electrical.
The post-treatment coefficient of 0.0625 for THD indicates that utility territories containing AI data centers experienced a 0.0625 percentage point increase in readings above the 0.8 distortion threshold following AI model releases. While a 0.0625 percentage point increase may appear small in isolation, it is still problematic as sustained exposure to power distortions compounds to over the multi-year lifespans of household appliances, accelerating degradation of refrigerators, air conditioners, and other motor-driven equipment NEMA_MG1,Laughner2024THD.
Fossil fuel demand effects are estimated using demand data from a total of 35.6 million generator-hour observations\footnote{These data are from he Clean Air Markets Program (CAMPD) dataset and include hourly generation by all US utility-scale fossil generators EPA_CAPMD_PowerSector2022}, which includes 900,956 treated observations from generators located within 20 miles of AI-affiliated data centers and the other 34.7 million from other generators as control observations.
Average fossil fuel demand effects are shown in the IV 2SLS rows in Panel B of Table (ref).\footnote{Recall that we use two stage instrumental variables (IV) regression here, where the second stage isolates the AI model caused demand-side effects from supply shocks (i.e., generator heat rate and alternate fuel source prices) modeled in the first stage. Details on the specifics of the instrumental variables specification can be found in Appendix Subsection (ref) and results from the first stage regression are presented in Appendix.} Figure (ref) shows the coefficients spread two months before (training) and after (inference) model release.
The training phase fossil fuel power demand corresponds to the $t=-2$ pre-treatment coefficient of $6.33$ MWh, which we find to also be statistically significant result. Note that this is the hourly demand at a single generator. When the aggregate impact is calculated, they represent over 500 additional GWh added nearby to a data center over the inference period in some cases. The additional power is roughly equivalent to that of 47,000 household-years in the United States. For inference, the $t=+2$ coefficient shows a significant demand increase of 5.20 MWh, representing the demand shift at generators near data centers two months after AI model release, isolated from supply-side price variation through our instrumentation strategy.
We estimate wholesale electricity price effects by combining our demand estimates with supply curve parameters. This approach, described in Section (ref), recognizes that our IV strategy identifies demand shifts but does not directly estimate equilibrium price effects.\footnote{Our instrumental variables approach isolates exogenous variation in electricity demand driven by AI model releases. However, observed market prices reflect the intersection of supply and demand in equilibrium. To translate our demand estimates into price effects, we must incorporate information about supply-side responses, which requires additional assumptions about the shape of the supply curve.} The price impact of AI-induced demand depends on the slope of the supply curve\footnote{Recall, the supply curve describes the relationship between price and quantity supplied. Its slope reflects how much generators must be compensated to produce additional electricity—steeper slopes indicate that marginal generation costs rise quickly as output expands, typically because higher-cost units must be dispatched.}: inelastic supply, where quantity supplied responds weakly to price changes, implies large price effects, while elastic supply implies small effects. We use supply chain elasticities for PJM obtained from our supply-side instrumental variables regression (see Figure (ref) in Appendix (ref) for a map).
Figure (ref) showing price increase estimates by zone. On average AI training leads to substantial increases in wholesale prices in PJM with the highest increase surpassing 25% in the American Electric Power (AEP) Zone.
We also estimate the expected price rise for retail customers, using estimates from the literature as documented in Appendix Subsection (ref). While we are not able to provide definite evidence of the impacts on retail price pass through, it should be noted that retail electricity prices have increased overall at a rate surpassing inflation since 2020 Nicoletti_Malik_Tartar_2024,horsley2025electricbill,spp2024stateofthemarket. Also, while increased investment in electric generation through power purchase agreements by tech firms may in the long run lead to falling retail prices, it appears that at present data center power usage is increasing wholesale prices by raising total demand, which requires dispatching higher-cost generators and thus increases the market clearing price—a price paid by all retail consumers.
Using the most relevant examples for PJM of retail elasticity of between .53 and .56 defeuilley2023pennsylvania, the expected consumer price rise in AEP, a sub-zone in PJM, would be $\sim 13\%$. However, using recent long-run econometric estimates of full passthrough from mackay2024wholesale, the long-run expected rise for consumers may be closer to the full 25%.
We present two sets of analyses to address issues highlighted in Subsection (ref): formal testing to provide evidence for the parallel trends assumption, and placebo tests address potential issues related to external confounds or lack of specific knowledge about training. Additional tests are found in Appendix (ref).
\noindent1) Formal Testing for Parallel Trends: For power demand specifically, the formal joint test of pre-treatment coefficients using a Wald test of joint significance yields a chi-squared statistic of 0.99 (df = 1) with p-value = 0.3196. We cannot reject the null hypothesis of parallel pre-trends at conventional significance levels. This provides support for our assumption that treated and control generators followed similar demand trajectories prior to AI model releases i.e., our identification assumption is valid. Some of our samples demonstrate significant evidence against the the parallel trends test prior to propensity score matching (PSM) that we use for sample selection. However, across all matched samples for both power quality and power demand after PSM, we find no evidence of violations of the parallel trends assumption. This implies that the technique correctly handles the issue where it exists.
\noindent2) Placebo Tests We conduct two placebo exercises to demonstrate that the treatments are meaningful and not the result of random chance, we implement random date and random location placebo tests. Both yield a p-value of 1.00 for the coefficient in our fossil fuel demand, CPQI, and THD analyses, confirming that our results do not arise from spuriously chosen locations or dates.
We present three main counterfactual exercises to inform stakeholders and policymakers about how alternative evolutionary paths for AI compute might affect grid outcomes. These combine our econometric estimates with data on model characteristics and infrastructure configurations to explore three dimensions: model scaling, geographic redistribution, and workload composition.\\
\noindent1) Model Scaling. We combine data on model size (parameter counts) with our econometric estimates from our least conservative specification to extrapolate grid impacts at different scales.
(a) Power Quality. Fitting a trend line to power quality impacts for Llama-3.1 (405B), PaLM (540B), and Llama 4 Behemoth (2T), we find that scaling from a 2 trillion parameter model to a 4 trillion parameter model would increase power quality impact from 0.321 to 0.434---a deterioration of roughly 0.113 units, or just over 35% relative to the 2 trillion baseline. The relationship is exponential over log parameter counts, implying that larger models require proportionally greater parameter reductions to achieve the same percentage improvement in power quality. (b) Power Demand. Using an analogous extrapolation, we find that each 1% parameter increase leads to approximately 0.15% higher power demand within three months of model release. Scaling from 540 billion to one trillion parameters would increase total power demand by 10.5%. For DALL-E, halving parameters would reduce inference demand by approximately 300 GWh---equivalent to the annual consumption of 27,000--30,000 homes.\\
\noindent2) Training Compute. Training compute requirements also have measurable effects: a 1% increase in training FLOPs is associated with a 0.1% increase in power demand. For GPT-3.5, halving training compute would reduce energy use by enough to power 8,000--10,000 U.S.\ homes for a year.\\
\noindent3) Colocated On-site Generation. Data centers increasingly deploy colocated on-site generation to reduce grid dependence. Our inventory identifies 87 facilities with on-site capability and 12,401 without. We test whether on-site capacity moderates estimated effects by interacting treatment with on-site status in the utility-level strict DiD specification. Results indicate that on-site generation substantially mitigates measured grid impacts. The sign reversal for CPQI is notable: data centers with on-site generation are associated with improved power quality (negative coefficient), while those without are associated with deterioration (positive coefficient). This heterogeneity suggests that on-site generation absorbs demand spikes that would otherwise propagate to the grid, and that aggregate grid impacts may understate total AI energy consumption.
Figure (ref) extends this to policy-relevant scenarios. Remarkably, universal adoption of on-site power would shift impact of data centers from a net detriment to power quality to a net improvement.\\
\noindent4) Geographic Redistribution
(a) Distributed Inference (Edge Computing). In our sample where we know AI data centers but not specific training and inference sites, we simulate a shift from cloud-based to edge-based inference, redistributing inference demand from concentrated hyperscaler facilities to smaller data centers located closer to population centers. Under this scenario, total demand added in a particular zone falls from a baseline of 162 MWh per hour over the entire treatment period to 4 MWh per hour. Figure (ref) illustrates this effect of shifting the composition of inference.
(b) Concentration in Low-Population, High-Supply Regions. Alternatively, we consider relocating both training and inference to regions with abundant electricity supply and lower baseline load, such as areas in the central United States with substantial renewable or fossil capacity. We simulate this by shifting demand from the baseline case where training and inference are assumed distributed to a case where the only data center for training and inference is the largest data center owned by a hyperscaler. This leads to an extreme concentration of added demand with nearly 100 GWh added near data centers in training locations. Figure (ref) shows this scenario.
This analysis has the important caveat that shifting to one data center for all training activity is extreme, with many possible configurations between full decentralization and full centralization. Centralization may also yield other benefits, such as concentration near renewable generation and reduced energy consumption from lower communication costs.
Understanding how AI models impact U.S. power grids is critical for ensuring energy reliability and security. This work highlights the challenges in quantifying this impact using existing data sources. It introduces a rigorous econometric methodology to overcome the lack of fine-grained data about contributing factors. In particular, we use difference-in-differences regressions to provide econometric estimates of the amount of power being used by specific AI models as well as the power distortions being caused by AI and the impacts of AI on price. We find evidence that AI is making power quality significantly worse and increasing fossil fuel power demand. Using counterfactual analyses, we provide estimates of the potential impacts of increasing efficiency of AI on the power draw of data centers. We also show how changes in the geographic composition of AI, a potential shift to edge computing, and a shift in the ratio of training to inference conducted impact the power market.