EconBase
← Back to paper

A Machine Learning Approach to Measuring Climate Adaptation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

74,568 characters · 12 sections · 90 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Machine Learning Approach to Measuring Climate Adaptation

{0pt} {0pt} {0pt} {0pt} {0pt}

abstractI measure adaptation to climate change by comparing elasticities from short-run and long-run changes in damaging weather. I propose a debiased machine learning approach to flexibly measure these elasticities in panel settings. In a simulation exercise, I show that debiased machine learning has considerable benefits relative to standard machine learning or ordinary least squares, particularly in high-dimensional settings. I then measure adaptation to damaging heat exposure in United States corn and soy production. Using rich sets of temperature and precipitation variation, I find evidence that short-run impacts from damaging heat are significantly offset in the long run. I show that this is because the impacts of long-run changes in heat exposure do not follow the same functional form as short-run shocks to heat exposure. \paragraph{Keywords:} Climate change, directional derivative, machine learning, panel data \paragraph{JEL Classification:} C14, C33, C55, Q51, Q54

Introduction

Measuring adaptation to recent climate change is important for informing climate policy and projecting damages of future climate change. The extent of recent adaptation can signal when policy interventions are needed and give a sense of how much adaptation will be possible to projected changes. Researchers can study this by modeling how weather shocks impact economic outcomes, and examining how impacts change with more exposure to the weather shock (e.g. Burke2016) or over time (e.g. Barreca2016). These studies depend on accurate models of the relationship between weather and economic outcomes.

Machine learning (ML) can help model these relationships when researchers have rich data, but lack expert guidance to suggest a functional form. When domain experts provide a model of how weather shocks impact economic outcomes, researchers can use this to accurately measure and compare impacts. Otherwise the researcher must learn this relationship from data. Economists typically use classic statistical tools to model weather-economic relationships without imposing strong functional form assumptions Hsiang2016. Such tools work well with low-dimensional weather variation, but can lead to high variance or inconsistent estimates as the dimensionality of covariates increases. Even a model with only temperature and precipitation can become high-dimensional if a researcher flexibly models each variable and the interactions between them. ML is well suited for flexible modeling in such settings Mullainathan, and can help researchers take full advantage of high-dimensional weather variation in modern panel datasets.

I introduce an ML approach to study adaptation to damaging heat exposure in United States (U.S.) corn and soy production. High temperatures are generally damaging for crop growth, although adaptation may offset some of these damages in the long run. Schlenker2006 introduce a parsimonious model of how an annual shock of heat exposure impacts crop yield. They model crop yields as a piecewise-linear function of heat exposure below and above a crop-specific temperature threshold, where heat exposure below the threshold is beneficial and heat exposure above the threshold is damaging. By estimating the parameter on damaging heat using different sources of variation, Schlenker2009, Burke2016, and Lemoine2018 argue that the observed degree of adaptation to damaging heat exposure will not be sufficient to offset projected climate damages.

I measure the degree of adaptation after learning the crop yield-weather relationship from data. I compare impacts from short-run and long-run changes in damaging heat exposure. With a low-dimensional linear model, this is equivalent to the method from Burke2016. The approach allows me to incorporate more flexible models and high-dimensional weather variation, making the method suitable for applications with rich weather variation and little expert guidance on how that variation impacts economic outcomes. I use ML to model the crop yield-weather relationship, specifically Least Absolute Shrinkage and Selection Operator (Lasso) and a neural network (NNet). My ML models account for additive fixed effects while flexibly modeling temperature, precipitation, and interactions between them.

To estimate the degree of adaptation, I estimate the elasticity of crop yields with respect to damaging heat exposure. This elasticity summarizes how damaging heat exposure impacts crop yields, and can be computed for each model. There is no single parameter I can compare across models, as in Burke2016. Instead, I find the elasticity by fitting a regression function of log crop yields on temperature and precipitation and computing the average directional derivative in the direction of a marginal increase in damaging heat exposure. I estimate these regression functions via ordinary least squares (OLS) and ML.

I implement a debiasing procedure to address bias from standard ML models. ML approaches can induce bias from overfitting or regularization, but there are approaches to reduce this bias when estimating causal parameters or statistics based on regression functions Chernozhukov2018,chernozhukov2022automatic. I adapt the estimator from chernozhukov2022automatic. This approach uses double machine learning (DML), where standard ML estimates are corrected using a second ML algorithm. I adjust the second ML algorithm to suit the panel setting. I then apply this DML procedure to debias the elasticity estimates.

I compare elasticities to estimate the degree that short-run impacts from damaging heat exposure are offset in the longer term. I construct panel datasets of long-run and short-run variation in crop yield and weather variables from 1990-2019. I then estimate the elasticity of crop yield with respect to damaging heat exposure in each dataset. I compute this elasticity using OLS, DML, and ML without bias correction. Comparing these estimates, I find the extent that short-run impacts from damaging heat exposure are offset in the longer term.

Before taking this approach to the data, I conduct a simulation exercise to evaluate DML estimates of an elasticity relative to OLS and standard ML approaches. Each simulation trial uses the empirical distribution of temperature and precipitation from U.S. counties, but simulates outcome variables. I examine the performance of my DML procedure relative to OLS and ML without bias correction, comparing the bias and variance of recovering the true elasticity. I use three potential sets of weather variables: a simple, commonly used set of annual temperature and precipitation variables, a set that includes richer variation in annual temperature, and a set that includes monthly observations of precipitation and the rich variation in temperature.

This simulation exercise clearly highlights the benefits of using debiased machine learning for high-dimensional settings. Both NNet and Lasso have advantages over OLS in high-dimensional cases, although NNet has lower bias. With the rich set of annual temperature variation, DML estimates have significantly lower variance than OLS and lower bias than standard ML estimates. With the set of monthly temperature and precipitation, DML estimates have significantly less bias and variance than OLS, and lower bias than standard ML.

I then apply the approach to a dataset of U.S. crop yields, and find evidence that adaptation is offsetting impacts from damaging heat exposure. I implement Lasso, NNet, and OLS estimators on a county-level dataset of corn and soy yields from 1990 to 2019. I compare estimates of the elasticity using short-run and long-run variation. I take 500 bootstrap trials of the estimated elasticity, using the three sets of weather variation as in the simulation exercise. Using the simple annual set of weather variables as in Schlenker2009 and Burke2016, I find that there has been little to no significant adaptation to climate change in corn or soy production. This confirms the results from Burke2016.

Using more flexible sets of weather variation, I find that a large share of short-run impacts from damaging heat exposure are offset in the long run. With short-run variation, I find statistically and economically significant declines in yield from a marginal increase in damaging heat exposure. However, I do not find evidence of such declines when using long-run variation. These results hold for both corn and soy. The primary difference comes from using a richer set of weather variables. I make this same conclusion using OLS with the more flexible set of weather variation, although the DML approach results in smaller confidence intervals. This shows that a substantial degree of the short-run impacts from damaging heat exposure are offset in the long run, suggesting substantial adaptation to this heat exposure.

This result differs dramatically from the conclusions by Burke2016 and other analyses Schlenker2009,Lemoine2018. This is likely explained by model misspecification for the impact of a long-run shift in heat exposure on crop yields. I show that in the panel with long-run variation, the simple model from Schlenker2006 does not adequately summarize the flexible role of temperature variables. While damaging heat exposure is correlated with declines in crop yield, other temperature variation is better able to explain these declines. This suggests that there is limited adaptation to some damaging feature of climate change, but not to the specific feature of marginal increase in damaging heat exposure.

This paper is related to several literatures. First is a literature on estimating the degree of adaptation to climate change. Measuring adaptation to climate change requires understanding how weather influences economic outcomes. For a review of economics literature on measuring the economic impacts of the weather, see Dell2014. Hsiang2016 provides an overview of econometric approaches to measuring these impacts. Much of this literature focuses on agriculture, as this sector is directly exposed to weather and hence is particularly vulnerable to climate change shukla2019ipcc. The first approach to studying impacts of climate change used the Ricardian approach, where researchers compare the value of agricultural land in cross sections. Mendelsohn1994 forecast the impacts of climate change by regressing average temperature and agricultural property value in a cross section of U.S. counties. This approach is susceptible to omitted variable bias, and subsequent work has focused on addressing specific omitted variables such as endogenous changes in farmer technology kurukulasuriya2011adaptation or nonfarm income ortiz2020role. Other approaches to estimate the potential for adaptation in agriculture involve economy-wide simulations Costinot2016, production changes in historical migrations Sutch2011,Olmstead2011, natural experiments Hornbeck2012,hagerty2021adaptation, or panel approaches.

Panel approaches address omitted variable bias by identifying adaptation from annual or long-term variation within a panel dataset. Schlenker2009 uses panel data to estimate the elasticity of crop yields with respect to extreme heat exposure, and conclude that there is limited potential to adapt to climate change because these damages are similar in the southern and northern U.S. despite climatic differences. Barreca2016 use a flexible model of temperature exposure to document how the mortality consequences of extreme heat declined over the 20$^\text{th}$ century. Burke2016 show that panel variation can be used to estimate adaptation to recent climate change by using separate sources of variation to identify the impacts of weather shocks and shifts in average temperature. Lemoine2018 provides an alternate approach that partially identifies the degree of possible adaptation by considering the role of ex-ante and ex-post adaptation to heat exposure shocks.

My paper is most closely related to Burke2016. Like their paper, I estimate the degree that damages to corn and soy yields from short-run changes in weather are offset over longer exposures. I also use crop and weather data from U.S. agriculture. My approach differs because I consider richer sets of weather variables, and use DML to model learn the relationship between these data and crop yields. I conclude that there has been a higher degree of adaptation to damaging heat exposure.

Second is a growing literature on applying ML methods in economics. Kleinberg2015 discuss applications of predictive machine learning in economics, and varian2014big and Mullainathan provide a practical guide to algorithms. Several recent papers have used ML to measure important outcomes in environmental economics. Crane-Droesch2018 proposes a semi-parametric NNet and uses it to study the impact of climate change on corn yields. Deryugina2019 uses a ML approach to measure the costs of air pollution. burlig2020machine use ML to refine estimates of energy efficiency improvements. stetter2022using use a DML approach to measure effectiveness of an agricultural intervention, and klosin2022 introduce a DML approach to measure elasticities in a panel setting. There are also numerous applications within agriculture; for a review, Liakos2018.

My paper is most related to Crane-Droesch2018 and klosin2022. Crane-Droesch2018 estimates a NNet that accounts for unobservable county-level fixed effects, and uses the model to predict yield under counterfactual climate change scenarios. I use the same technique to address county-level fixed effects, although I modify the algorithm in order to recover derivatives from the network and to ensure that the network is differentiable. Crane-Droesch2018 studies the impacts of future climate change, while I use this tool to understand adaptation to recent climate change. Like klosin2022, I estimate the elasticity of crop yield with respect to an increase in damaging heat exposure. I consider more flexible representations of temperature variation, and I apply the estimator to measure adaptation to recent climate change.

Within the literature on machine learning, this paper applies results from the emerging field of DML. Chernozhukov2018 prove that sample splitting and constructing Neyman-orthogonal moment conditions can yield approximately debiased machine learning estimates in certain settings. Semenova2021 extend the Neyman-orthogonal moment condition approach to several other statistical targets, including structural derivatives. In an alternate approach, chernozhukov2022automatic,chernozhukov2022debiased give an approximately debiased estimator for a more general class of linear functionals based on the Riesz representation theorem. klosin2022 adopt this approach in panel settings and prove asymptotic normality of the estimator for the average derivative. Like klosin2022, my approach applies the result from chernozhukov2022automatic in panel settings; I use a different approach to address fixed effects and consider a more flexible representation of temperature variation.

The rest of this paper proceeds as follows. In (ref), I describe the data used for this project. I illustrate the degree of climate change and describe the transformations required to generate the growing degree days. In (ref), I describe the methods used; this includes details on how to compute average derivatives and an explanation of the debiased machine learning estimation approach. In (ref), I give details and results of the simulation exercise. In (ref), I present and discuss results from using the estimation procedure to measure the degree of adaptation to climate change. (ref) concludes.

\FloatBarrier

Data

For the empirical application, I use weather and crop data from U.S. corn and soy production from 1990-2019. I consider counties east of the 100\textdegree West meridian, which defines an agricultural region of significant corn and soy cultivation. From 1990-2019, this region produced over 93% of the nation's corn and over 99% of the nation's soy. Crop data are annual yield (bushels per acre) of corn and soy, from the U.S. Department of Agriculture's Survey of Agriculture. These data also include the area planted (acres) in each county.

Weather data are generously shared by Schlenker2009, who provide a gridded dataset of daily temperature and precipitation from a network of consistently reporting weather stations. As in Schlenker2009 and Ortiz-Bobea2013, I consider weather during the March-August growing season. I aggregate the gridded dataset to a county-level dataset of daily maximum temperature, minimum temperature, and precipitation. I then transform daily temperature exposure into growing degree days (GDD) at a monthly level and for the growing season. (ref) illustrates the transformation from daily temperature observations to GDD.

figure[figure omitted — 2,104 chars of source]

Throughout this paper, I consider three sets of weather variables per county: total growing season heat exposure above and below 29\textdegree C plus total growing season precipitation (Yearly Linear), total growing season heat exposure in each 1\textdegree C temperature bin plus total growing season precipitation (Yearly Flexible), and monthly heat exposure in each 1\textdegree C bin and monthly precipitation for each month of the growing season (Monthly Flexible). (ref) illustrates the transformations of temperature variables. The Yearly Linear transformation, shown in (ref), is widely used in economic analysis. Schlenker2009 demonstrate that, in a panel setting, regression based on heat exposure above and below a crop-specific damaging threshold explains as much variation in yields as more flexible models. Burke2016 identify 29\textdegree C as the threshold for damaging heat exposure in the long run, for both corn and soy. I use the 29\textdegree C threshold throughout the analysis, and refer to heat exposure above 29\textdegree C as damaging heat exposure.

figure[figure omitted — 1,351 chars of source]

Variation in the degree of climate change between counties is used to estimate the long-run elasticities in our sample. (ref) shows illustrates the difference in GDD exposure above 29\textdegree C per US county, from the period 1990-1999 to the period from 2010-2019. Heat exposure increases in 80% of counties in the sample, but there is substantial variation in the degree of warming even within states. There are many pairs of neighboring counties where one experienced an increase in damaging heat exposure and the other experienced a decrease. This variation is plausibly uncorrelated with unobservable factors such as soil quality or per-state fixed effects. See Burke2016 for a detailed argument that this type of variation can identify the effects of climate change.

While many counties experienced damaging heat exposure increases and decreases, only a handful had average yields decline between these time periods. (ref) and (ref) show changes in average yields of corn and soy over this time period. Only 4.03% of counties saw corn yields decline and only 3.96% saw soy yields decline, among counties that grew corn or soy in both 1990-1999 and 2010-2019. This suggests that improved farming technology increased yields throughout the sample, and that these improvements exceeded damages from increased damaging heat exposure.

\FloatBarrier

Methods

In this section, I discuss the methods used to estimate the degree of adaptation. I first define adaptation in terms of average directional derivatives, and describe how to measure these average directional derivatives. I then introduce the ordinary least squares (OLS), Lasso and neural network (NNet) procedures, including details on how to account for fixed effect terms in each model and the cross-folds training procedures. I also describe the double machine learning (DML) approach used to adjust for bias in the standard machine learning (ML) estimates.

Adaptation

I define adaptation as the amount of short-term impact to yield from damaging heat exposure that is offset in the longer term. This definition encompasses all adaptation behaviors an agent makes to their production technology, management practices, or variety choice within each crop as they are exposed to climate change. It does not capture some other important margins of adaptation such as crop switching or exit from agriculture. Burke2016 provide evidence that these margins of adaptation are limited. This definition also does not capture changes that alter the impact of damaging heat independently from an individual's exposure to climate change, such as the decline in heat-related mortality as studied by Barreca2016. First I define the estimation target, as introduced by Burke2016. I generalize this definition in terms of elasticities, and discuss how to estimate those elasticities using flexible functional forms and/or richer sets of weather variation.

Following Burke2016, I measure adaptation by comparing the elasticity of crop yields with respect to extreme heat using short-run and long-run variation. Short-run variation comes from year-to-year changes in an annual panel of weather observations and log crop yields. To capture long-run variation, I first average weather and crop yield data over a long period (I use 10-year periods) and then construct a two-period panel. The elasticity computed using short run variation captures the extent that a weather shock of extreme heat in a single year impacts yields, while the elasticity computed using long run variation captures the extent that exposure to a long period of increased heat exposure will impact average crop yields. Estimates using long-run variation arguably capture the extent of damages from climate change, because there is a change in average heat exposure over a relatively long period where farmers have time to adjust to those changes. (ref) illustrates long-run variation in the sample, comparing the differences in long-run average crop yields and damaging temperature exposure over the decades I study. Comparing impacts using these sources of variation, I can conclude whether the short-term damages are offset in the long term.

It is straightforward to recover these elasticities in a model where log yield is linear in total growing season heat exposure. Burke2016 use such a model, which I adapt below:

equation[equation omitted — 129 chars of source]

where $higher_{it}$ ($lower_{it}$) is total growing season heat exposure above (below) the damaging temperature threshold, $g$ is some function of precipitation, $a_i$ is an additive per-county fixed effect term, and $\varepsilon_{it}$ is an additive error term. Burke2016 estimate this equation with OLS, after using the within transformation to remove the fixed effect term\footnote{Burke2016 advocate using first differences to remove the fixed effect in the long-run variation panel; I use a within transformation approach as this is numerically equivalent to taking first differences in a two-period panel. }. The key parameter here is $\beta_2$, which captures the extent that log crop yield changes with marginal increase in damaging heat exposure. This is equivalent to the elasticity of crop yield with respect to damaging heat exposure. Note that $\beta_2$ is expected to be negative in the short run, as temperature shocks in this range are damaging to crop growth Schlenker2006,Schlenker2009. The estimate of the share of short-run damages that are offset in the longer term is therefore $(\hat{\beta}_2^{SR} - \hat{\beta}_2^{LR})/\hat{\beta}_2^{SR} = 1 - \hat{\beta}_2^{LR}/\hat{\beta}_2^{SR}$, where $\hat{\beta}_2^{SR}$ ($\hat{\beta}_2^{LR}$) is the estimate using the short-run (long-run) variation. Burke2016 take bootstrap samples of this ratio and fail to reject the null hypothesis that the ratio is different from 0.

To recreate this ratio with a more flexible functional form, I replace the parameter estimates with average directional derivatives. Consider a more general functional form:

equation[equation omitted — 64 chars of source]

Here, $X_{it}$ is a collection of weather variables, $\gamma$ is a general function, $a_i$ is an additive fixed effect term, and $\varepsilon_{it}$ is an additive error term. I use three different collections of weather variables for $X_{it}$, as described in (ref). As $y_{it}$ is the log of crop yields, the elasticity of crop yield with respect to some variable is equivalent to the average directional derivative of $\gamma$ with respect to that variable.

The analogue to $\beta_2$ from (ref) is therefore the average directional derivative of $\gamma$ with respect to $higher_{it}$. Let $\theta^{SR}$ ($\theta^{LR}$) be this average directional derivative using short-run (long-run) variation. I then take bootstrap samples of the ratio $1 - \hat{\theta}_2^{LR}/\hat{\theta}_2^{SR}$ to test the hypothesis that damages from short-run heat exposure are offset in the longer run.

This average directional derivative is equivalent to an average partial derivative when using the Yearly Linear set of weather variables. Specifically, take $\hat{\theta} = \mathbb{E}[\partial \hat{\gamma}(X_{it}) / \partial higher_{it}]$. When using a linear specification, this average derivative is equivalent to the parameter estimate from OLS. When using NNet, this can be recovered after training the network (see (ref)). When the specification involves basis functions (such as OLS with polynomial basis functions or Lasso), I recover this derivative by taking the average of the dot product of the gradient of the basis functions and the estimated coefficients. I describe this procedure in (ref).

When I use another set of weather variation, it is also necessary to account for the extent that each weather variable contributes to the total growing season heat exposure. I find this using the chain rule:

equation[equation omitted — 163 chars of source]

Where $X_{it}^{higher}$ is the set of weather variables that are summed to reach total growing season heat exposure above the temperature threshold. I set $\partial X_{it}/\partial higher_{it} = X_{it} / higher_{it}$; this captures the assumption that additional marginal heat will be distributed proportionally to the empirical heat exposure distribution.

This allows me to measure adaptation to climate change from recent panel variation while using a flexible model of high-dimensional weather variation. In the following sections, I describe how I estimate the regression function and average directional derivative using various estimation techniques.

Ordinary Least Squares methods

Ordinary Least Squares (OLS) is a classical statistics approach to estimating this elasticity. I use methods with a linear functional form (OLS Linear), and after applying a basis function transformation of polynomial functions and interactions (OLS Poly). In both cases, I use the within transformation to remove a county-level fixed effect term, and then include yearly fixed effect terms via dummy variables.

I use basis functions that include interactions and flexible functional forms of the original data. I specify polynomial basis functions of all terms, as well as interactions between these polynomial expansions. To produce a tractable model, I limit the space of potential interactions to interactions between heat exposure and precipitation variables within the same time period. For example, in the Monthly Flexible specification, I consider cumulative GDD in July between 28 and 29 C interacted with precipitation in July, as well as squared values of each term and the interactions between those polynomial expansions, but do not consider interactions of that variable with cumulative GDD in any other temperature bin, or precipitation in any other month. I use 3rd order polynomials for the Yearly Linear and Yearly Flexible variable sets, and 2nd order polynomials for the Monthly Flexible variable set. I then scale each flexible basis function so that it has mean zero and variance 1.

For OLS Linear, I use the identity set of basis functions; that is, $b(X_{it}) = X_{it}$. For Yearly Linear, $X_{it}$ has 3 covariates; for Yearly Flexible $X_{it}$ has 41; and for Monthly Flexible $X_{it}$ has 246. For OLS Poly, I use the set of basis functions described above. For Yearly Linear, $b(X_{it})$ has 30 covariates; for Yearly Flexible $b(X_{it})$ has 486; and for Monthly Flexible $b(X_{it})$ has 1464.

OLS assumes that the following is a true model of the relationship:

equation[equation omitted — 87 chars of source]

As is common in economics, I assume that we have relatively short panels where it is not possible to consistently estimate $a_i$ by including dummy variables. I therefore use the within transformation to remove county-level fixed effect terms:

equation[equation omitted — 95 chars of source]

where the double dot denotes the within transformation, i.e. $\ddot{y}_{it} := y_{it} - \text{mean}(\{y_{it} \forall t\})$ and $\ddot{b}(X_{it}) := \text{mean}(\{b(X_{it} \forall t\})$. Note that to construct $\ddot{b}(X_{it})$, the mean of all observations in that panel unit is subtracted after applying the basis function transformation. This ensures that $\beta_0$ is the same parameter vector between models. I use OLS to estimate $\hat{\beta}$.

I then compute the average directional derivative by projecting my estimate of $\beta_0$ on partial derivatives of the basis functions. Define the gradient of the basis function as $b_{higher}$; this is a 1 by $p$ dictionary of the derivative of each basis function with respect to $higher_{it}$. The true average directional derivative is then $\theta_0 = \mathbb{E}[b_{higher}(X_{it})\beta_0]$ and its estimate is $\hat{\theta} = \mathbb{E}[b_{higher}(X_{it})\hat{\beta}]$.

In specifications in my main analysis, I include per-year fixed effects as dummy variables. I assume that there are enough observations per time period to consistently estimate these variables separately. The derivative of each per-year fixed effect is zero, so including these terms does not change how I estimate the average directional derivative.

Machine Learning Methods

Machine Learning (ML) methods include a range of estimation techniques that allow consistent function approximation in high-dimensional settings, when the number of covariates is large relative to the number of observations. Such settings poses challenges for classical statistical methods such as OLS, binning, or kernel regression. I use Lasso and Neural Networks (NNets), although the same procedure could be used for another algorithm such as random forests, support vector machines, or other methods. I focus on these two machine learning algorithms because each enables the researcher to incorporate linear fixed effects and to evaluate derivatives without numerical differentiation. Standard ML methods can induce bias in regression analysis; I overcome this bias by using a procedure from chernozhukov2022automatic.

The average derivative is computed from a ML regression of the output variable on weather inputs. Write $\gamma(\cdot; \lambda)$ to denote the flexible machine learning function, emphasizing the dependence on the hyperparameter $\lambda$. The hyperparameter is a researcher-specified value that influences the behavior of the model, such as the regularization penalty in Lasso or the network width in NNet. The machine learner is estimated as a regression function: $\mathbb{E}[\ddot{y}_{it} | \ddot{X}_{it}] =\hat{\ddot{\gamma}}(X_{it}; \lambda)$. (ref) and (ref) give details on each estimation procedure. Let $m(\gamma, X_{it}; \lambda)$ denote the directional derivative of $\gamma(\cdot; \lambda)$ evaluated on observation $X_{it}$.

I use the double machine learning (DML) procedure from chernozhukov2022automatic to find an approximately debiased estimate of the average directional derivative. I discuss this procedure in more detail in (ref). Briefly, I estimate a second machine learner and use this to construct an estimate of the average directional derivative that is robust to errors in estimating either the first or second machine learner. Let $\alpha(X; \kappa)$ denote this second machine learner, emphasizing the dependence on hyperparameter $\kappa$. chernozhukov2022automatic show that the expression $\mathbb{E}[m(\hat{\gamma}, X_{it}; \lambda) + \hat{\alpha}(X_{it}; \kappa)\hat{\varepsilon}_{it}]$ is an approximately unbiased estimate of the true average directional derivative, where $\hat{\varepsilon}_{it}$ is the residual from estimating $y_{it}$. I modify the form of the doubly robust estimator in chernozhukov2022automatic to account for the panel structure of the data; details are in (ref).

I use a data-driven process to determine the value of the hyperparameters. First split each panel unit (a U.S. county) into one of the $k$ folds for cross validation. Grouping observations from each panel unit into the same fold reduces correlation between training and test data, as counties share unobservable characteristics that likely influence the distribution of weather and crop yields. Let $\mathcal{I}_\ell$ denote the set of indices in fold $\ell$, for $\ell \in \{1, 2, \dots, k\}$. For each of the $k$ folds, I train the ML and DML algorithm on data not in fold $\ell$, and evaluate the algorithm only on indices in fold $\ell$. Let $\hat{\gamma}_{\ell}$ and $\hat{\alpha}_{\ell}$ denote the ML and DML estimators trained on the set of indices not in fold $\ell$. Then select a hyperparameter value by searching over a grid of potential values, and selecting the value that minimizes a loss function. Let $\mathcal{L}_\gamma(\gamma_\ell, \mathcal{I}_{\ell}; \lambda)$ be the mean squared error of the function $\gamma_\ell$ with the hyperparameter $\lambda$ on the data in $\mathcal{I}_{\ell}$. Let $\mathcal{L}_\alpha(\alpha_\ell, \mathcal{I}_{\ell}; \kappa)$ be the loss function of the function $\alpha_\ell$ with the hyperparameter $\kappa$ on the data in $\mathcal{I}_{\ell}$; I describe this loss function in (ref).

Once the hyperparameter is selected, I evaluate the (debiased) score on the test sets using the same folds defined above. The full estimation procedure for the debiased score is below. To use this procedure without the debiasing correction, omit the bias correction term $\hat{\alpha}_\ell(X_{it}; \hat{\kappa})(\ddot{y}_{it} - \hat{\ddot{\gamma}}_\ell(X_{it}; \hat{\lambda}))$ from each step.

enumerate• Select hyperparameter $\hat{\lambda}$ that minimize test-set mean squared error of the regression: \[ \hat{\lambda} = \operatorname*{arg\,min}_{\lambda} \sum_{\ell = 1}^k \mathcal{L}_\gamma(\hat{\gamma}_\ell, \mathcal{I}_{\ell}; \lambda) \] • Select hyperparameter $\hat{\kappa}$ that minimize test-set loss of the double machine learner: \[ \hat{\kappa} = \operatorname*{arg\,min}_{\kappa} \sum_{\ell = 1}^k \mathcal{L}_\alpha(\hat{\alpha}_{\ell}, \mathcal{I}_{\ell}; \kappa) \] • Evaluate the debiased score using these hyperparameters: \[ \hat{\theta} = \frac{1}{N} \sum_{\ell = 1}^k \sum_{it \in \mathcal{I}_\ell} m(\hat{\gamma}_\ell, X_{it}; \hat{\lambda}) + \hat{\alpha}_\ell(X_{it}; \hat{\kappa})(\ddot{y}_{it} - \hat{\ddot{\gamma}}_\ell(X_{it}; \hat{\lambda})) \] • Find the asymptotic variance of the estimator, after adjusting for within-panel-unit correlations. Let $\hat{\theta}_{\ell; it} := m(\hat{\gamma}_\ell, X_{it}; \hat{\lambda}) + \hat{\alpha}_\ell(X_{it}; \hat{\kappa})(\ddot{y}_{it} - \hat{\ddot{\gamma}}_\ell(X_{it}; \hat{\lambda}))$ and $ \hat{\theta}_{\ell; i} := 1 / T \sum_{t = 1}^T \hat{\theta}_{\ell; it}$. Then the asymptotic variance is: \[ \hat{V} = \frac{1}{N} \sum_{\ell=1}^{k} \sum_{i \in \mathcal{I}_\ell} \left\{ \sum_{t = 1}^{T} (\hat{\theta}_{\ell; it} - \hat{\theta})^2 + 2 \sum_{t = 1}^{T - 1} \sum_{t' = t + 1}^{T} (\hat{\theta}_{\ell; it} - \hat{\theta}_{\ell; i})(\hat{\theta}_{\ell; it'} - \hat{\theta}_{\ell; i}) \right\} \]

In the following subsections, I describe how to train and evaluate the Lasso and NNet estimators, and introduce the DML procedure.

Lasso

Least absolute shrinkage and selection operator (Lasso) is a regression procedure that selects a sparse linear model from a researcher-specified set of basis functions. The procedure finds this sparse linear combination by minimizing squared error of estimation, while penalizing more complex models via regularization. The procedure is similar to OLS Poly, but estimates can differ greatly because of this penalization. The hyperparameter $\lambda$ is the magnitude of the regularization term, which I determine via the cross fitting procedure outlined above.

As with OLS, I assume that there is a true linear model of the form (ref) and take within transformations to remove county-level fixed effects to result in (ref). I use the same flexible set of basis functions used in OLS Poly, as defined in (ref). Unlike in OLS Poly, I assume that the true parameter vector is sparse and find an estimate by solving a regularized optimization problem.

For each cross-validation fold, I find the estimate $\hat{\beta}_\ell$ via the following minimization problem:

equation[equation omitted — 176 chars of source]

where $|\beta|_1$ is $\ell1$ penalty or the sum of the absolute value of each component of $\beta$. To incorporate yearly fixed effects, I include dummy variables and do not apply the regularization penalty to the coefficients on those dummy variables. This reflects the assumption that while coefficients in $\gamma$ are sparse, coefficients in the fixed effects terms are not belloni2016inference. Additional details on the Lasso procedure are included in (ref).

Once I have estimated $\hat{\beta}_\ell$, I compute the directional derivative by projecting my estimate of $\beta_0$ on partial derivatives of the basis functions. This is the same procedure to compute the derivative using OLS, as described in (ref).

Neural Network

A neural network (NNet) learns a relationship between inputs and outputs through an iterative process, where the best fitting model is selected from a large space of flexible transformations of all potential interactions of input features. NNets have several advantages for my application. First, it is straightforward to compute a gradient of the output of the entire network with respect to each input feature. Second, NNets can incorporate fixed effects in panel models, as demonstrated by Crane-Droesch2018. Third, estimating NNets does not require the researcher to specify basis functions.

I compute derivatives of the NNet by extracting gradients computed during training the network. A NNet is a weighted composition of activation functions (user-specified transformations) applied to an input vector. After a random initialization of internal parameters, the model goes through many iterations of predicting the output variable, calculating a loss from training data, and updating the parameter values. I use the mean squared error for this loss function. While training or evaluating the network, the algorithm computes the derivative of the output with respect to each input. This is commonly used to adjust parameter values during the training procedure, in a process known as back propagation. I also use these automatic derivatives to compute the partial derivative of the prediction with respect to each weather input.

I follow Crane-Droesch2018 to estimate the model with nonparametric treatment of weather inputs and additive linear fixed effects. Training this network involves an iterative procedure that can be interpreted as selecting basis functions from a large space of candidate functions. After this training procedure is completed, I use standard econometric techniques to account for linear terms, treating the nonlinear transformation of inputs as a feature in a linear model. Specifically, I use the within transformation to remove individual fixed effects and use OLS to find other linear terms such as yearly fixed effects. The iterative procedure is performed on the set of training data, and the OLS step is performed on the test set. This requires an architecture with a top layer that is linear in all fixed effect components and the output of the nonlinear network transformation. See Crane-Droesch2018 for more details on such networks and their performance relative to linear models or fully nonparametric models in estimating agricultural yields. Additional details of the NNet architecture and training procedure are included in (ref).

Double Machine Learning

Double machine learning (DML) is an approach to remove bias from standard machine learning algorithms. I use an approach introduced by chernozhukov2022automatic, relying on the statistical theory of the Riesz representer.

There are two related, yet orthogonal, problems in estimating the average derivative. First is the regression function, and second is the derivative operator on a regression function. chernozhukov2022automatic show how to conduct this second step without estimating a regression function, and that these two methods can be combined to construct an approximately debiased estimate of the original target derivative.

This second function can be estimated from data using the Riesz representation theorem. I denote this second function $\alpha_0$, and its estimate $\hat{\alpha}$. The Riesz representation implies that $\mathbb{E}[\alpha_0(X_{i,t}) \ddot{\gamma}_0(X_{i,t})] = m(\ddot{\gamma}_0, X_{i,t})$. chernozhukov2022automatic show how to use this fact to estimate $\alpha_0$ from data. In a panel setting, I use the within-transformed function $\ddot{\gamma}$; I use $m(\ddot{\gamma}_0, X_{i,t}) = m(\gamma_0, X_{i,t})$ because the functional $m$ is a directional derivative and the derivative of the mean value is 0. I assume that $\alpha_0$ is linear in the set of basis functions $b$. The true Riesz representer is then $\alpha_0(X_{it}) := \ddot{b}(X_{it}) \rho_0$ and its estimate is $\hat{\alpha}(X_{it}) := \ddot{b}(X_{it}) \hat{\rho}$. I estimate $\hat{\rho}$ using an optimization package, based on the moment conditions identified by chernozhukov2022automatic. klosin2022 demonstrate that this procedure is effective in panel settings. (ref) includes more details on the motivation of this Riesz representer and the estimation procedure.

This estimate is then used to construct the following doubly robust score: \[ \hat{\theta} = \frac{1}{N} \sum_{\ell = 1}^k \sum_{it \in \mathcal{I}_\ell} m(\hat{\gamma}_\ell, X_{it}; \hat{\lambda}) + (\ddot{y}_{it} - \hat{\ddot{\gamma}}_{\ell}(X_{it})) \hat{\alpha}_\ell(X_{it}; \hat{\kappa}) \]

\FloatBarrier

Simulation Exercise

I conduct a simulation exercise to compare the performance of the estimation procedure to OLS while varying the set of weather variables used. To match the setting as closely as possible, I use the empirical distribution of weather covariates in all trials. I then apply a simulated production function of the the piecewise-linear functional form from Schlenker2009.

The simulation exercise focuses on comparing ML, DML, and OLS in a regression setting with high-dimensional variation. OLS is correctly specified in all trials. In cases where OLS is not correctly specified, ML and DML would likely perform better because they are able to represent a richer set of flexible functional forms. Simulation exercises by Chernozhukov2018 demonstrate this property.

figure[figure omitted — 1,504 chars of source]

I conduct 1,000 Monte Carlo simulation trials of this estimation procedure. In each trial, I first randomly draw a 1,000 county sample, and randomly select a 2-year panel from these counties. I use the weather observations from this sample, guaranteeing that there are realistic correlations between weather variables. Then, I generate a $y$ variable according to the following function: $y_{it} = a_i + \beta_1 lower_{it} + \beta_2 higher_{it} + \beta_3 prec_{it} + \varepsilon_{it}$, where $lower$ ($higher$) is the total growing season accumulated GDD below (above) 29\textdegree C and $prec$ is the total growing season precipitation. This functional form closely matches the parsimonious form suggested by Schlenker2009. Note that in this functional form, OLS Linear is correctly specified for all sets of weather variables I consider. I set $\beta_1 = 0.02; \beta_2 = -0.05; \beta_3 = 0.001$, and take $a_i \sim N(1,1)$ and $\varepsilon_{it} \sim N(0,1)$.

table[table omitted — 522 chars of source]

The results of this simulation exercise are summarized visually in (ref) and numerically in (ref). I estimate $\beta_2$, the parameter of interest, using OLS, ML, and DML for weather variables in the Yearly Linear, Yearly Flexible, and Monthly Flexible sets (as illustrated in (ref)). For OLS, I OLS Linear and OLS Poly as described in (ref). For ML, I consider Lasso and NNet estimators as described in (ref) and (ref). For DML, I adjust each ML result with the DML procedure in (ref). I use Lasso DML and NNet DML to describe the Lasso and NNet estimates with the DML correction. I use the cross-folds and sample splitting procedure described in (ref).

With Yearly Linear weather inputs, OLS performs best among all models. OLS Poly and OLS Linear have the lowest bias, and OLS Linear has the lowest variance. The machine learning models performed reasonably well, especially with the DML correction, although they have greater variance and bias than the OLS results. The central estimate from both DML approaches lie within 7% of the true elasticity, and a 95% confidence interval contains the true elasticity.

With Yearly Flexible weather variables, ML estimates have considerably lower variance than OLS. The estimates using OLS Linear have a standard deviation approximately three times as large as the DML estimates, while estimates from OLS Poly have a standard deviation approximately ten times as large. There is substantial bias from the Lasso estimates, although the DML correction greatly reduces this bias. The NNet and NNet DML estimates have very little bias; central estimates from both are within 0.4% of the true value and have less bias than OLS Linear. The 95% confidence interval from both DML procedures contains the true elasticity, although the 95% confidence interval from Lasso without DML does not.

With Monthly Flexible variables, the benefits of using DML are even more dramatic. OLS Linear and OLS Poly have extremely high variance, and both have greater bias than the DML methods. The high variance is especially limiting, as confidence intervals using these approaches are uninformative. As (ref) shows, the interquartile range of the bootstrap estimates is not visible on a plot whose range is 8 times the magnitude of the true target elasticity. OLS Linear estimates a mean value of 2.359, relative to the true value of -0.05. OLS Poly has a bias of .018, which is greater than bias from Lasso DML (0.012) or NNet DML (0.0087).

Note that while mean squared error (MSE) can help suggest a preferred model, a straightforward comparison of MSE does not select the model with lowest bias. MSE appears higher among ML models than OLS because the MSE reported by ML is the MSE of the model evaluated on a test set. When MSE from an ML model is low, this indicates that ML is providing a better fit because the model extrapolates well to unseen data. When MSE from an OLS approach is low, this could indicate overfitting. OLS Poly has the lowest MSE in all trials, but is overfitting the data in the Yearly Flexible and Monthly Flexible cases. In the data generating process (DGP), the additive error term has variance 1. The MSE from OLS Poly is much lower than the true squared error from the DGP, indicating that OLS Poly is fitting the noise instead of the desired pattern in the data. The MSE from OLS Linear with Monthly Flexible weather terms is significantly lower than the true squared error, indicating that OLS Linear may be overfitting the data in that setting. This suggests that a comparison of MSE alone should not be used to select the preferred model, but the researcher can consider test-set MSE and performance in simulation trials.

This simulation exercise shows that when a researcher wishes to measure a relationship with a high-dimensional set of weather variables, DML can provide results with lower bias and less variance than OLS. With Yearly Flexible or Monthly Flexible weather, NNet DML has the lowest bias among all models, and the DML methods have significantly lower variance. OLS performs best in the Yearly Linear case, which is expected because OLS is correctly specified and the weather variation is low-dimensional.

\FloatBarrier

Results

I implement the above estimation procedure to study adaptation to damaging heat exposure in U.S. corn and soy production. When using the Yearly Linear set of weather variables (as in Burke2016), I find little to no evidence of adaptation. However, when using a richer set of weather variation, I find evidence that a considerable share of the short-run damages from extreme heat are offset with greater exposure to those temperatures. To help explain this discrepancy, I visualize OLS results with Yearly Flexible temperature variation to show that the simple linear model may not accurately describe the long-run role of extreme heat.

I run 500 bootstrap trials of the procedure described in (ref) to measure the elasticity of crop yields with respect to extreme heat. As described in (ref), I study corn and soy production from 1990-2019 in counties east of the 100 \textdegree W meridian. I use a panel of the full data to capture short-run variation, and capture long-run variation by comparing average yield and weather from 1990-1999 to 2010-2019. In each bootstrap trial, I draw a random subsamples of 80% of counties. This resampling scheme addresses the intertemporal correlation as suggested by kapetanios2008bootstrap, while avoiding the risk of own-sample bias from machine learning methods by including the same observations in train and test data.

I estimate the elasticity using the three sets of weather variation described in (ref) and all estimation approaches described in (ref). Each method uses the within transformation to remove individual fixed effects, and includes additive yearly fixed effects. For estimation methods that use a set of basis functions, I use polynomial expansions of all temperature and precipitation variables, and interactions of the polynomial expansions of each precipitation variable with the polynomial expansions of each temperature variable. I use third-order polynomial expansions for Yearly Linear and Yearly Flexible weather variables, and second-order polynomial expansions for Monthly Flexible. The machine learning models use a 5-fold cross validation procedure.

I only report results from OLS Linear, Lasso DML, and NNet DML in this section. The simulation results showed that the debiasing procedures are effective at reducing bias from naive machine learning methods, and that estimates using OLS Poly can have extremely high variance. Results from the machine learning models without the debiasing procedure and from OLS with polynomial basis functions are included in (ref). Generally, estimates using machine learning without bias correction are similar to their debiased counterparts, and estimates using OLS Poly have unacceptably high variance when using richer sets of weather variation.

figure[figure omitted — 2,383 chars of source]
table[table omitted — 588 chars of source]

Reassuringly, the results using Yearly Linear weather variation are similar to those of Burke2016. (ref) summarizes the results from the bootstrap trials. Long-run and short-run datasets both find that additional extreme heat decreases crop yields, with comparable magnitudes for both corn and soy cultivation. The magnitude of these elasticities are economically significant -- an estimate of -0.005 implies that crop yields decline by 0.5% for each additional day crops are exposed to temperatures above 29\textdegree C. The findings are similar to the results reported by Burke2016. They find that the elasticity of corn yield ranges from $-0.0037$ to $-0.0062$.

As in Burke2016, with Yearly Linear weather variation I find little to no evidence that damages from the short run are offset in the long run. (ref) shows the results of taking bootstrap samples of the ratio $1 - \hat{\theta}^{LR}/\hat{\theta}^{SR}$. As described in (ref), $\hat{\theta}^{SR} \ (\hat{\theta}^{LR}) $ is the estimated elasticity with short-run (long-run) variation. For corn cultivation, I fail to reject the null hypothesis that this ratio is different from 0 for any estimation method. I test at the $p = 0.05$ level, with a Bonferroni correction to account for taking three hypothesis tests. For soy cultivation, I find mixed results: for both OLS Linear and NNet DML I reject the null hypothesis at this level, while with Lasso DML I fail to reject the null hypothesis. For these soy estimates, the 95% confidence interval is $[0.1278, 0.3488]$ using OLS Linear and $[0.07153, 0.3649]$ using NNet DML. Confidence intervals are computed using the sample mean and standard deviation among bootstrap trials.

Using Yearly Flexible weather variables, I find significant evidence of adaptation. The panel with short-run variation finds statistically and economically significant damages from a marginal increase in extreme heat exposure. Central estimates range from $-0.01057$ to $-0.01248$ for corn, and $-0.006983$ to $-0.00788$ for soy. However, I fail to reject the null hypothesis that the long-run elasticity is different from zero, for both crops using all specifications. (ref) confirms this finding. For both crops and using all estimation methods, I reject the null hypothesis that the degree of short-run damages that are offset in the long run is equal to zero.

Using this set of weather variables, my preferred estimator is NNet DML. Simulation results showed the this estimator performed best with Yearly Flexible weather variables. The MSE also suggests that the NNet Double is performing well in this case. The MSE of NNet DML decreases for long-run regressions for both corn and soy. As the MSE of NNet DML is the test-set MSE, an improvement relative to the MSE from the Yearly Linear weather variables reflects am improvement in modeling the true functional form instead of overfitting. Using this preferred estimate, I find a 95% confidence interval that the ratio is within $[0.3023, 1.222]$ for corn and $[0.6331, 2.001]$ for soy.

The results from using the Monthly Flexible set of weather variables are similar to those using Yearly Flexible weather variables, although with greater variance. This greater variance is not surprising, as the simulation results demonstrated that DML and OLS have higher bias and variance using this set of weather variables. With short-run variation using this set of weather variables, there is an economically and statistically significant decline in yields associated with a marginal increase in heat exposure. With long-run variation, there is no evidence of significant declines in yields from a marginal increase in heat exposure. Reassuringly, this evidence is consistent with the result using Yearly Flexible weather variation. Given the greater variance of these estimates, I fail to reject the null hypothesis that the ratio of these elasticities is different from zero except using DML methods. Using OLS Linear, I find that a marginal increase in heat exposure significantly increases crop yields.

These results show that, when the temperature-crop yield relationship is modeled flexibly, there are not significant average declines in yields associated with a marginal change in long-run damaging heat exposure. I do not estimate a significant decline from long-run exposure to temperatures above 29\textdegree C because other variation is better able to explain differences in average long-run yields. This is not the case when using short-run variation, where flexible modeling confirms that a marginal increase in exposure to temperatures above 29\textdegree C is significantly associated with a decline in crop yields. This finding is novel, and suggests significant adaptation to damaging heat exposure.

figure[figure omitted — 1,821 chars of source]

To better explain this result, I examine the estimates of temperature coefficients from each weather bin of the Yearly Flexible weather variables. I use my NNet DML procedure to estimate these plots. (ref) includes results using alternate estimation procedures. Due to the computational cost of estimating each average derivative using the DML procedure, I do not replicate the DML procedure for each of these temperature coefficients; the results are similar. (ref) visualizes the results of these regressions. (ref) and (ref) show the results of this estimate using short-run variation. These results are similar to the results from Schlenker2009, who noted that the piecewise linear fit (the dotted line in these figures) explains most of the variation from the flexible modeling approach (the bar chart in these figures). Here, the piecewise linear model seems to be a good fit: before around 29\textdegree C, additional heat exposure is generally associated with an increase in yields. After this point, additional heat exposure is associated with a decline in yields.

The results using long-run variation do not share this pattern. (ref) and (ref) show the results of this estimate using long-run variation. The coefficients do not demonstrate a clear piecewise linear pattern. This suggests that the piecewise linear model relating temperature exposure and crop yield may not be appropriate in long-run models. In (ref), I replicate this figure for different time periods, starting from 1950-1979. These figures confirm that the piecewise linear model is a better fit when using short-run variation than when using long run variation. While the results from the Yearly Linear analysis show that long-run damaging heat exposure is correlated with a decline in crop yields, these results show that other temperature variation is better able to explain the long-run changes in crop yields. This shows the importance of flexible modeling to understanding the role of a marginal increase in damaging heat exposure.

Discussion

I introduced a DML procedure that can estimate average directional derivatives in high-dimensional settings, and demonstrated the benefits of this procedure over OLS estimation in a simulation exercise. The results show that this procedure can estimate average directional derivatives with lower bias and variance than OLS. The simulation results also show that my method is less reliable than having access to a true, parsimonious model of the underlying function. In settings where the researcher has high-dimensional weather variation and does not have a strong prior about the true functional form, this DML approach can be used to estimate elasticities and the degree of adaptation to a weather feature.

Applying this estimator to panel of U.S. corn and soy yields with a rich set of temperature variation, I conclude that there has been significant adaptation to damaging heat exposure. Using this flexible method, a panel of short-run damages finds evidence that a marginal increase in heat exposure above 29\textdegree C is damaging for both corn and soy yields. However, by constructing a dataset of average changes to capture long-run variation, I cannot reject the null hypothesis that a marginal increase in long-run damaging heat exposure is unrelated to long-run crop yields. OLS estimates using a Yearly Flexible set of weather variables support this finding. I demonstrate that long-run exposure to temperatures above 29\textdegree C is not clearly associated with yield declines, as is the case with short-run exposures. This implies that there has been significant adaptation to climate change in this setting, as the short-run damages are significantly offset in the long run.

This approach does not offer insight into what form this adaptation may take. Farmers have many possible adaptation mechanisms, such as investing in irrigation, purchasing improved seeds, or adjusting planting times. It is important to learn which mechanisms may be responsible for offsetting damages in the long run, in order to study the long-term effectiveness of these mechanisms or the ability to use them in other agricultural contexts. More research is needed to understand how these damages are offset.

The apparent contradiction between this result and prior literature suggests that while there is adaptation to a marginal increase in damaging heat exposure, there is limited adaptation to some other damaging feature of climate change. Prior literature, such as Burke2016, did not find evidence of substantial adaptation to damaging heat exposure. My analysis shows that while a long-run increase in damaging heat exposure is correlated with a decline in crop yields, variation in heat exposure at other temperature levels is better able to explain this pattern. This is evidence of omitted variable bias in long-run estimates using only beneficial and damaging heat exposure. While prior results indicate limited adaptation to some change in the temperature distribution, my detailed analysis finds that there has been substantial adaptation to damaging heat exposure.

\FloatBarrier \singlespacing \printbibliography

\doublespacing