Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,088 characters · 32 sections · 113 citation commands
Dynamic Biases of Static Panel Data Estimators
\doublespacing
\paragraph{Keywords:} treatment effects, fixed effects panel model, dynamic panel model, climate economics, environmental dynamics
JEL classification: C33, Q51
Treatment effects are often estimated with fixed effects panel models currie2020technology. These models accounted for 19% of the empirical articles published in the American Economic Review from 2010 to 2012 de2020two.\footnote{Fixed effects panel models are especially important in economic contexts where randomizing treatment is not feasible. For example, in environmental economics, it is impossible to randomize exposure to floods or temperature shocks to estimate their effects. Instead, researchers use observational data and typically assume that, conditional on location, the treatment is random.} It is common for these fixed-effects models to be static, meaning the models do not control for past outcomes and, therefore, do not account for dynamics. Static models are frequently used even when economic theory suggests a dynamic relationship between past and current outcomes. For example, past agricultural yields impact future yields through soil health and market demand griliches1963sources, human capital formation is dynamic cunha2007technology, past labor market states impacts future states blanchard1988beyond, past GDP impacts current GDP solow1956contribution, and migration flows are functions of historical migration chains massey1993theories. Despite this theoretical expectation, empirical papers studying these outcomes often run fixed effects analysis without controlling for past outcomes.\footnote{Examples include annan2015federal, burke2015global, cho2017effects, jessoe2018climate, drabo2015natural, mahajan2020taken, missirian2017asylum, graff2018temperature garg2020temperature.}
One reason why researchers often use static models is because they are concerned that adding past outcomes as controls can cause estimation problems. This concern is rooted in the understanding that panel models with both fixed effects and past outcomes are subject to Nickell bias nickell1981biases. Nickell bias arises from the estimation of fixed effects, in particular the failure of strict exogeneity conditions when dynamics are present, leading to biased estimates. Another reason why researchers use static models is because they think it is unnecessary to control for past outcomes when treatment is random. For example, if treatment is an exogenous temperature shock, it is commonly assumed that treatment is random conditional on fixed effects. Consequently, researchers believe they can obtain unbiased treatment effect estimates without controlling for past outcomes. Random treatment assignment ensures that excluding past outcomes does not lead to omitted variable bias. However, this paper shows that a different bias arises due to the omission of these dynamics and the fixed effect estimation.
The main contribution of this paper is two-fold. First, I identify “dynamic bias”, a bias that arises when static fixed effects panel models are used in settings with dynamics in outcomes. Dynamic bias occurs because past treatment is related to past outcomes through the outcome equation. Therefore, de-meaning or differencing to account for fixed effects generates confounding. This generated confounding leads to biased treatment effect estimates even if the treatment is random. Treatment effects are even more biased if treatment is related to past outcomes and therefore endogenous. A contribution of this paper is to explicitly characterize the resulting dynamic bias. Using analytical derivations, simulations, and applied examples, I demonstrate that dynamic bias is often substantial. The dynamic bias, caused by omitting past outcomes from the model, is often much larger than the Nickell bias caused by including past outcomes in the model.\footnote{In simulations, no matter how correlated past outcomes are with current outcomes, dynamic bias is larger than Nickell bias.} Therefore, when researchers avoid controlling for past outcomes because they are worried about introducing bias, they may in fact be making the bias on their treatment effect estimates larger.
The second contribution of the paper is to develop a new estimator that corrects for both dynamic bias and Nickell bias, which works even when the number of time periods $T$ is fixed. I create a novel bias correction, “dynamic biases correction” (DBC), by deriving a formula for the asymptotic\footnote{Under asymptotics where the number units $N \rightarrow \infty$ but number of time periods $T$ is fixed.} bias term following kiviet1995bias. The explicit formula can be derived because, given the model, we know exactly how the Nickell bias is generated: the errors become correlated with regressors due to the demeaning required for fixed effects estimation. Therefore, I analytically derive an expression for the demeaned errors and use this expression to calculate how the demeaned errors correlate with the demeaned regressors, leading to an explicit formula for the bias. The bias correction is achieved by subtracting an estimate of the bias from the original estimated coefficients.
The DBC correction for treatment effects is the first analytic bias correction that accommodates endogenous treatment related to past outcomes. Allowing for endogenous treatment is important in many economic settings because selection into treatment is often a function of past outcomes.\footnote{marx2022parallel and ghanem2022selection show how many economic models lead to treatment selection that depends on past outcomes. For example, selection into environmental policy treatment is often a function of past environmental conditions. Policies that target air pollution in particular areas are implemented because of past pollution rates chay2005does. As another example, regional deforestation protection in Brazil is based on past deforestation in the region harding2021commodity, assunccao2023optimal.} My correction also works when the treatment is exogenous, e.g. randomly assigned. Additionally, my correction is the first analytical correction not to impose homogeneity in treatment effects, even when treatment is endogenous. A large econometrics literature highlights problems that arise when homogeneity in treatment effects is assumed incorrectly -- and how important it is to allow for treatment effects to vary depending on the treatment group in panel data.\footnote{Discussed in sun2020estimating, callaway2021difference, goodman2021difference, de2024difference .}
The exact bias correction approach has some appealing characteristics in comparison to alternative corrections, which are based on instrumental variables holtz1988estimating, arellano1991some. The instrumental variable methods are based on using further outcome lags as instruments for outcome lags. The correct choice of instrument is often unclear and can lead to problems caused by weak instruments.\footnote{Problems caused by weak instruments are discussed by andrews2019weak, mikusheva2021many, mikusheva2024weak.} In simulations, I find that my analytical solution keeps standard errors as small as the original linear regressions and maintains proper coverage, as compared to instrumental variable methods which lead to larger standard errors.
The biases and proposed estimator are illustrated through Monte Carlo simulations and empirical replications. I generate simulation data where treatment is a function of past outomes, as well as data where treatment is random conditional on fixed effects. Even in the simulation with random treatment, treatment effect estimates from models omitting past outcomes are biased significantly more than models that control for past outcomes. In Monte Carlo simulations, the DBC estimator is unbiased and has smaller standard errors as compared to the Arellano-Bond-based alternative. To validate my results with real data, I use data from dell2012temperature. This paper studies the “contemporary causal effect of temperature on the development process” by using a yearly panel of countries with GDP and temperature information dell2012temperature. The treatment variable of interest is temperature, which is taken to be random, conditional on the country. Controlling for past outcomes significantly changes the results both when the outcome is GDP growth (10% change) and GDP levels (120% change).\footnote{The p-value for the GDP growth result is .06, so it is significant at the 10% level while the GDP level result is significant at the 5% level.}\footnote {Both GDP growth and levels are used as outcomes in the literature that studies the effect of temperature on economic outcomes newell2021gdp, nath2024much.}
The rest of the paper proceeds as follows. Section (ref) discusses related work in more detail. I provide empirical motivation and an overview of the biases discussed in this paper in Section (ref). Section (ref) presents theoretical results. Section (ref) introduces the DBC estimator of treatment effects. Section (ref) conducts a simulation study to illustrate the asymptotic properties of the estimators I propose. Section (ref) provides an empirical example illustrating how correcting for dynamic bias impacts treatment effect estimation. Section (ref) concludes.
This paper contributes to the literature on dynamic panel estimation with fixed effects. This literature has a long history, beginning with griliches1967distributed\footnote{griliches1967distributed discussed how time series regression parameters are estimated with bias when intercepts are included. This phenomenon was also studied by nerlove1971further. } and other researchers, who investigated how dynamics in outcomes lead to violations of the strict exogeneity of errors assumption, which is necessary for unbiased ordinary least square error (OLS) coefficient estimation including fixed effect estimation. nickell1981biases derived an explicit formula for the bias of OLS parameters when past outcomes and unit fixed effects are included in the regression, assuming all other regressors are strictly exogenous. Nickell bias can be thought of as a specific type of incidental parameter bias neyman1948consistent.\footnote{Incidental parameter bias arises in fixed effects panel models with the number of time periods $T$ is small relative to the number of units $N$. This bias occurs because each unit fixed effect is estimated only using a few observations, leading to biased estimates of the fixed effects.} The primary focus of the literature has been on investigating how including past outcomes in the regression leads to bias, particularly with regard to the coefficient on the past outcomes. However, in applied work the statistical object of interest is often the treatment effect estimate, rather than the dynamic process itself, and do not include past outcomes as controls. This paper is therefore uniquely contributing to the literature by focusing on the coefficient on the treatment effect rather than the coefficient on the past outcome, which is treated as a nuisance parameter. It is also the first to focus on parameter estimation when the past outcome is not included in the OLS regression model. By studying treatment effects in models that exclude past outcomes, I am able to characterize dynamic bias. This characterization reveals that dynamic bias can be much larger than the previously studied Nickell bias. In all simulation specifications, dynamic bias is larger than Nickell bias, even when treatment is random.
To correct for dynamic bias, I provide a bias-corrected (DBC) estimator. Given that Nickell bias is often much smaller than dynamic bias, my bias correction procedure calls for first controlling for past outcomes, and then correcting the resulting Nickell bias. In settings with fixed $T$, there are two main approaches for dealing with Nickell bias. The first and most well-known approach is that of holtz1988estimating and arellano1991some, which is based on instrumental variables. The second approach is an analytical method that I build upon, following the works of kiviet1995bias, juodis2015iterative, and breitung2022bias.
The instrumental variables (IV) approach is based on using past outcomes as instruments for endogenous regressors. Its goal is to use further lags of the outcomes, or treatments, as instruments for current differenced outcomes and treatments. These instruments may be weak, and the more time periods available the more possible instruments, which can lead to problems associated with many weak instruments mikusheva2024weak.\footnote{The weak instrument issue is partially addressed by blundell1998initial, who proposed a system GMM solution by including both first differences and levels of past outcomes as instruments. However, this approach still suffers from the weak instrument problem when the variance of individual effects is greater than the variance of the errors (see bun2010weak).} In practice, estimates obtained using instrumental variables are quite sensitive to the choice of instruments, making instrument selection a daunting task for applied researchers.\footnote{See Section (ref) for application to dell2012temperature. Depending on the instruments used, point estimates for both treatment and past outcomes flip signs.}
Instead of using instrumental variables, I follow kiviet1995bias and analytically correct the bias. The analytical bias correction avoids problems associated with instruments while still providing a correction in the fixed-T setting. The downside of the analytical bias correction method is that it requires analytical work that is outcome and treatment model-specific, which IV methods do not. The past analytical bias literature provided corrections for a specific class of models. Both kiviet1995bias and breitung2022bias focus on corrections for models with endogeneity from past outcomes and do not accommodate endogenous treatment. breitung2022bias extends the work of kiviet1995bias by allowing for multiple lags of the outcome variable and for incorporating the bias correction into a GMM framework. I build off breitung2022bias and also use a Generalized method of moments (GMM) framework for the DBC. Instead of GMM, juodis2015iterative provides an iterative analytical bias correction for Vector Autoregressive (VAR) systems. VAR models also do not accommodate endogenous treatments; endogenous treatments in time period $t$ impact outcomes in time period $t$. Therefore to allow for endogenous treatments I extend the analytical bias corrections to structural VARs (SVARs). This extension requires that I make a structural assumption to avoid simultaneity problems: treatment in a time period $t$ impacts the outcome in time period $t$, but the outcome in time period $t$ doesn't impact treatment in time period $t$. I allow past outcomes, like those from time period $t-1$, to impact treatment in time period $t$. This paper is the first to provide an analytical correction for settings with endogenous treatments. This analytical correction is also the first to allow for interaction terms between endogenous variables and exogenous variables, allowing treatment effects to vary based on observable characteristics.
I work in the fixed-T setting to help applied researchers sidestep the issue of having to guess whether they have “enough” time periods for a correction that yields valid inference. Although asymptotic normality is only guaranteed in a fixed-T setting with instrumental or analytical bias approaches, there are other approaches for de-biasing Nickell bias as long as T is allowed to grow. One correction for Nickell bias in the large T asymptotic setting is the jackknife approach (e.g., dhaene2015split), which involves taking samples of the panel data with different lengths of T to calculate the bias. This hands-off approach is easy to implement but results in larger standard errors than the analytical method fernandez2016individual. Fixed-T analytical corrections work with flexible linear models but do not allow for non-linear non-separable models (e.g., logit models). Marginal effects are not identified in fixed-T settings in non-separable models chernozhukov2013average. Many applied papers in environmental economics seek to estimate treatment effects, a type of marginal effect, so I impose flexible linearity to allow for their estimation.
This paper demonstrates the importance of explicitly including past outcomes in the regression model for panel data settings where there is a relationship between past and current outcomes,\footnote{A method for testing whether past outcomes influence current outcomes, as opposed to merely exhibiting autocorrelation in the model's errors, is discussed in chamberlain1982multivariate.} Currently, applied researchers often employ various alternative methods to account for outcome history without incorporating past outcomes into their models. Common approaches include adding time trends to the models or employing factor model-based methods.\footnote{For examples of applied research using factor models to control for outcome history, see damm2024beyond; for examples using time trends, see annan2015federal.} However, these methods do not control for dynamics in outcomes—that is, the influence of past outcomes on current outcomes, meaning the dynamics bias persists in these estimates.
Time trends are often introduced into models by adding unit-specific trends, which are constructed by interacting unit-specific dummy variables with a polynomial of the time variable. While these trends capture a general time-specific pattern for each unit, they fail to account for the influence of past outcomes on current outcomes, such as in autoregressive processes. In other words, time trends alone do not capture the dynamic relationships in the data where past values directly affect present outcomes. This is also seen empirically in the applied example in Section (ref). Another prevalent method for controlling for outcome history is the use of factor models. Factor model-based approaches, such as synthetic controls and synthetic differences-in-differences, arkhangelsky2021synthetic, abadie2010synthetic rely on low-rank factor assumptions that do not accommodate unit-specific dynamics.\footnote{In the context of panel data, a low-rank factor model aims to capture cross-sectional correlations among different units (e.g., individuals, firms, or countries) by assuming that these correlations can be explained by a limited number of common factors. Importantly, such models focus primarily on capturing variation between units rather than time-series dynamics within each unit.} Consequently, to properly account for the dynamics where past outcomes influence current outcomes, researchers must directly include past outcomes in their models.\footnote{In bio-statistics, the g-formula is also used to control for outcome history, but it does not accommodate selection on unobserved fixed effects robins1986new, naimi2017introduction.}
In the applied literature, when the correlation between past and current outcomes is large, another approach for dealing with dynamics is transforming outcomes to reduce the correlation. For example, GDP levels are correlated highly over time, with a correlation parameter of .95, the correlation of the difference transformed GDP (GDP growth) is smaller and closer to .3. However, transforming the outcomes through this approach still creates bias in treatment effect estimates. This is discussed in greater detail in the next section.
Before presenting the theoretical results of the paper, I preview the main empirical results of this paper. I give brief intuition for the origins of dynamic bias. I also compare dynamic bias and Nickell bias in simulations both when treatment is random and when it is endogenous (related to past outcomes). I then discuss how commonly used transformations of the outcome variable interact with these biases.
My applied example is based on the literature that studies the effect of temperature on GDP. The units are countries and the time periods are years. Suppose that the true model is the simple model given in Equation (ref): temperature, $\text{Temp}_{i,t}$, in every time period is an independent and identically distributed (i.i.d.) random shock conditional on country.\footnote{$\text{Temp}_{i,t} = c_i + u_{i,t}$ where $c_i$ is a country fixed effect and $u_{i,t}$ is a random shock.} In a fixed effects model, treatment is therefore as good as random conditional on the fixed effects. Additionally, past GDP is included in the true model, as solow1956contribution explains that past GDP impacts current GDP.\footnote{One way that past GDP affects current GDP through its impact on capital accumulation. Higher GDP in the past implies higher savings and investment, leading to a larger capital stock in the current period. Since the capital stock is an input in the production function, a larger capital stock results in higher current GDP.}
In this section, I introduce two models that could be used to estimate the treatment effect $\tau_0$. Researchers often estimate these models using OLS. A necessary condition for OLS to be unbiased is that treatment in time period $t$, $\text{Temp}_{i,t}$, is uncorrelated with the model error $e_{i,t}$ in time period $t$.
Researchers often estimate these models because they see that treatment in time period $t$, $\text{Temp}_{i,t}$, is uncorrelated with model error in time period $t$. However,researchers still have to estimate fixed effects $a_i$ by either using dummies, the within transformation, or first differences.\footnote{Running a regression with within-transformed data leads to same coefficients as running a regression with unit dummies wooldridge2010econometric. Running first-differences also generates bias of a similar form.} These transformations create new variables that are functions of data in multiple time periods. For example, consider the within transformation,
This new treatment variable $\widetilde{\text{Temp}}_{i,t}$ is now a function of $\text{Temp}_{i,s}$ in all time periods. Because of this, unbiased OLS estimation requires not only that treatment in time period $t$, $\text{Temp}_{i,t}$, is uncorrelated with model error in time period $t$, but also that regressors in all time periods are uncorrelated with model error, which is known as strict exogenieity. The fixed effect estimation turns past outcomes into generated confounders if they are correlated with either the outcome or the treatment.\footnote{In classic cross sectional causal inference we think of variables as confounders if they are correlated with both the outcome and treatment. } Both models, the Static Model and Dynamic Model, lead to biased treatment effects.
\paragraph{Variation on dynamic bias:} Before showing simulation results, I introduce a variation on the Static Model, which I call the Delta Model, often implemented in applied work that also leads to dynamic bias.
Some researchers suspect that highly persistent outcomes ($\rho_{10}$ close to 1) can cause problems with their analysis, so instead they study transformations of their outcomes, such as differences or growth. When taking the difference on the left-hand side, this imposes a $\rho_{10}$ coefficient of 1, since $1 \times \text{GDP}_{i,t-1}$ is subtracted from the model. The difference between 1 and true $\rho_{10}$ remains in the error of the model and therefore $\eta_{i,t} := (1 - \rho_{10}) \text{GDP}_{i,t -1} + \epsilon_{i,t} $.\footnote{Note here that only the outcome ($GDP_{i,t}$) is being transformed through differences - the variables on the right-hand side are not being differenced - therefore this transformation is not equivalent to the the first-differences transformation.} A type of dynamic bias occurs in this model because like in the Static Model, part of the outcome remains in the error term. It is the case that $\text{Temp}_{i,t-1}$ is correlated with $\eta_{i,t}$, discussed in detail in Appendix (ref), also leading to bias. I call this type of dynamic bias transformation bias.
Although all three models (Dynamic Model, Static Model, and Delta Model) lead to biased treatment effect estimates, the magnitude of the bias varies greatly. I run a simulation and present the results in Figure (ref) to highlight some key takeaways from these biases. I generate datasets based on the true model given in Equation (ref) with random treatment. I set the true treatment effect $\tau_0 = .5$, and simulate a variety of datasets. Each of the three panels in the figure corresponds to a different value of the correlation between past outcomes and current outcomes $\rho_{10} \in (.2,.5,.9)$ used to generate the data. I simulate datasets with 1000 units and vary the number of time periods along the x-axis. I estimate the treatment effects using OLS to estimate the three models above, and plot the value of the treatment effect estimates on the y-axis. The number of time periods in the dataset and the magnitude of $\rho_{10}$ significantly impact the magnitude of the bias.
The Dynamic Model (Equation (ref)) leads to the smallest treatment effect estimate bias out of all three models regardless of the value of $\rho_{10}$. The intuition for this is that only the Dynamic Model explicitly controls for past outcome, and so only in this model does treatment remain strictly exogenous.
As for whether the Static Model or Delta Model has the largest bias, that depends on the value of $\rho_{10}$. When $\rho_{10}$ is close to 1 (the right most panel), the bias of the Delta Model is smaller than the bias of the Static Model. This is because in the Delta Model, the past outcome with a coefficient of 1 is subtracted from the current outcome to create the transformed outcome $\Delta \text{GDP}_{i,t}$. Therefore if $\rho_{10}$ is close to 1, this subtracting is sort of controlling for the past outcome. The opposite is true when $\rho_{10}$ is small (close to 0, the left-most panel) - then subtracting out a value of 1 leads to more bias.
One key lesson for applied work is that the fewer the number of time periods in the panel dataset, the worse the bias of all these methods. This bias is important to keep in mind when shortening panel datasets; panel datasets are often shortened in environmental research when studies differentiate between temperature and climate change by using “long differences”.\footnote{Papers that use long differences include nordhaus2006geography, deryugina2014does, burke2015global.} When researchers implement long differences, they normally reduce the length of the panel, which in turn increases the bias of the above models. On the other hand, the longer the difference, the smaller the correlation between the outcomes in the two time periods, which can reduce the bias as well.
Treatment effect estimates are even more biased in settings where treatment is not random and is instead related to past outcomes, making it endogenous. Treatment can be related to past outcomes due to selection into treatment. marx2022parallel, ghanem2022selection show that many economic models lead to treatment selection based on past outcomes.
For example, selection into environmental policy treatment is often a function of past environmental conditions. Policies that target air pollution in specific areas are implemented because of those areas' past pollution levels chay2005does. In Brazil, regional deforestation protection is based on previous years deforestation harding2021commodity, assunccao2023optimal. In public finance, we might expect that health care spending in the past impacts health care plan choice, which then impacts future health care spending.\footnote{A key finding in kowalski2023reconciling was that emergency room utilization before the Oregon Health Insurance Experiment began affected enrollment into health insurance within the experiment, which in turn affected emergency room utilization.} In development economics, the assignment of drought relief in southern India has been found to be a function of the region's past outcomes tarquinio2022politics. In industrial economics, past firm productivity, which is the outcome, impacts output choices in current time periods olley1992dynamics.
When treatment is endogenous, the bias does not disappear as the number of time periods increases for the Static and Delta model. Since the treatment is endogenous, omitted variable bias arises when past outcomes are not explicitly included in the model. This omitted variable bias does not decrease, regardless of how large the number of time periods or units is. Simulations in Appendix (ref) demonstrate that even with a small degree of endogeneity, the bias in treatment effect estimates can be substantial. Omitted variable bias can be removed by controlling for past outcomes and correcting for the resulting Nickell bias, which is smaller than the omitted variable bias caused by failing to control for past outcomes.
Dynamic bias arises when there are dynamics in outcomes but researchers do not control for past outcomes when estimating treatment effects. Dynamic bias is explained and characterized in Section (ref). Correcting dynamic bias requires controlling for past outcomes. Estimators that include past outcomes as controls still suffer from Nickell bias, which is discussed in Section (ref). Section (ref) provides the dynamic bias-corrected (DBC) estimator which eliminates Nickell bias.
I use $Y_{i,t}$ to denote the outcome, $Y_{i,t-h}$ to denote the $h^{th}$ lag of the outcome, $D_{i,t}$ to denote treatment, and $W_{i,t}$ to denote other covariates. I use the superscript $t$ to denote the vector of variables up to time period $t$, for example $D_{i}^t := (D_{i,1}, D_{i,2}, \cdots, D_{i,t})$. The concatenation of all possible regressors is given by $X_{i,t} := (D_{i}^t, W_{i}^t, Y_{i}^{t-1})$.
Note that the covariates $W_{i,t}$ can include time period dummies. That is for each time period $t$ I can include a dummy $Q_t$ with 1 in the $t^{th}$ position. This enables the model to include time effects.
I use $\perp \!\!\! \perp$ to denote independence, and $\perp$ for uncorrelated. I use bar notation for averages over time, for example $\Bar{Y}_{i,-1} := \frac{1}{T} \sum_{t=1}^T Y_{i,t-1}$ and $\Bar{Y}_{i} := \frac{1}{T} \sum_{t=1}^T Y_{i,t}$. I use tilde to denote the within transformation, for example,
I begin by introducing the setting under the potential outcomes framework (neyman1923, rubin1974estimating). I observe outcomes $Y_{i,t}$ for a sample of units $i = 1, \dots, N$ for time periods $t = 1, \dots, T$. The number of time periods $T$ is fixed, but the cross-sectional dimension $N$ grows with more observations. Therefore, the formal asymptotic results are for the large $N$ setting where $N \rightarrow \infty$ and $T$ is fixed. This is also sometimes known as the “short panel”.
For now, let us assume that the treatment $D_{i,t}$ is continuous. For each value $d$ of the treatment $D_{i,t}$, unit $i$ has a corresponding potential outcome $Y_{i,t}(D_{i,t} = d)$. In the example in dell2012temperature, $D_{i,t}$ is temperature and $Y_{i,t}$ is country GDP. The authors are interested in estimating the “contemporary causal effect of temperature on the development process”.\footnote{One may also be interested in the estimation of the long term effect of treatment, as studied in hausman1985taxes, as opposed to the contemporary. I focus on the contemporary effect in this paper as it is the most common effect of interest in a set of surveyed applied papers.} Formally, the causal object of interest is average effect of marginally increasing temperature in the current time period ($D_{i,t}$) on GDP, holding all else fixed. This is also referred to the contemporary average partial derivative (APD) and is written as
where the expectation is over all units and time periods.
I impose structure on the potential outcomes to allow estimation. In the rest of this section, I showcase that dynamic bias is a problem even in a simple (though common) stylized model. I use the formula for the bias in this example to provide intuition. However, when I provide a bias correction in Section (ref) I introduce a more general model (Section (ref)), to which the DBC correction also applies.
This simple example is based off the setting of dell2012temperature. I include $Y_{i,t-1}$ in the outcome model because economic theory tells us that past GDP ($Y_{i,t-1}$) impacts current GDP ($Y_{i,t}$) solow1956contribution. As is standard in panel data, I impose additive fixed effects $a_i$, often called individual fixed effects, which can be correlated with covariates. In the treatment model in Equation (ref),treatment $D_{i,t}$ is also a function of individual fixed effects $c_i$.\footnote{It is important to note dynamic bias remains even if treatment is not a function of fixed effects and $D_{i,t} = u_{i,t}$.} This structure mimics dell2012temperature, who write that treatment is correlated with country fixed effects $a_i$. The presence of fixed effects in the true structural model generates the need to control for fixed effects in treatment effect estimation.
However, in the data we don't observe the fixed effects. Further, they cannot be estimated consistently in short panels, which is known as the incidental parameter problem neyman1948consistent. I bypass the estimation of fixed effects by using within-estimation, which leads to the same estimates of $\theta = (\rho_{10},\tau_0)$ as would explicitly estimating fixed effects by including fixed effect dummies for each unit.
The weak exogeneity assumption says the that error has zero conditional expectation given current and past covariates. Because the structural outcome model allows for lagged outcomes to impact current outcomes, I cannot make the more familiar strict exogeneity assumption, which would require that the regressors are uncorrelated with the error term across all time periods, meaning that current, past, and future values of the regressors do not influence the current error term.
Given the model laid out above, I now introduce dynamic bias and show how it impacts treatment effect estimates. Dynamic bias arises when there is a relationship between past outcomes and current outcomes in the true model but past outcomes are not included in the estimation model. A crucial feature of dynamic bias is that it arises even when the treatment is random.
Given the simple stylized model presented in Assumption (ref), in particular the homogeneous treatment assumption in Equation (ref), the treatment object of interest is simply $\tau_0$. This follows using our definition of the APD from Equation (ref) along with potential outcome given in Equation (ref).
In the context of dell2012temperature this would represent the contemporary causal effect of temperature on the development process.
In settings where treatment assignment is random, applied researchers often rely on a static model that does not account for past outcomes to estimate $\tau_0$. I write the static model in Equation (ref).
\paragraph{Model 1: Static Model}
Given the true model presented in Assumption (ref), the error in the static model is given by $e_{i,t} := \rho_{10} \text{Y}_{i,t -1} + \epsilon_{i,t}$. Since applied researchers do not observe fixed effects $a_i$, they can not use OLS to estimate equation (ref) directly. Instead they have to estimate the fixed effects with dummy variables, or by using the within or first difference transformation. I remove the fixed effects by using the within-transformation, which recall was defined in Equation (ref).
\paragraph{Model 2: Within-Transformed Static Model}
\paragraph{OLS Estimator.} To estimate the treatment effect, OLS is used to estimate Equation (ref). I call the coefficient estimated this way $\hat{\tau}^{D}$. Unfortunately $\hat{\tau}^{D}$ is inconsistent—no matter how large the number of units grows, the estimator does not converge to the true parameter value.
The proof of Theorem (ref) is given in Appendix (ref). This bias is what I term dynamic bias. This bias is of order $1/T$, and it diminishes as the number of time period increases.
Dynamic bias arises due to the estimation of fixed effects, which are typically estimated by including dummies, applying within-transformation, or first-differencing the data. This estimation generates confounding if past outcomes are not controlled for. In classic cross-sectional causal inference, a covariate is a confounder only if it is correlated with both the treatment and the outcomes. This reasoning explains why many researchers opt for static panel regressions; in their contexts, treatment is random, so although past outcomes affect current outcomes, they are not correlated with the treatment and, therefore, are not classic confounders angrist2009mostly. Due to fixed effects estimation, a past outcome is a generated confounder and generates bias if it is correlated with either the treatment or the outcome.
Estimation of models with fixed effects requires strict exogeneity for unbiased estimation. When treatment is random, and the past outcome is controlled for in the model, then treatment is a strictly exogenous regressor. However, when the past outcome is not controlled for, the past outcome is part of the model error, and treatment is no longer strictly exogenous. This leads to biased treatment effects.
To see algebraic intuition for why strict exogeneity does not hold, consider the within-transformation approach for estimating fixed effects. After applying the within transformation, the newly transformed variables become functions of data from all time periods. Consequently, the errors $\Tilde{e}_{i,t}$ are functions of errors $e_{i,t}$ across all time periods, causing the within-transformed error to be correlated with the treatment, which in turn leads to endogeneity.
Therefore $\widetilde{e}_{i,t}$ is a function of ${D}_{i,t}$.
However, if past outcomes had been included in the model, the treatment would have been a strictly exogenous regressor. Dynamic bias occurs when the outcome lag is not considered, leading to potentially large biases in treatment effect estimates. This dynamic bias can be mitigated by including past outcomes in the regression model.
However, even when past outcomes are included, the well-known Nickell bias remains. In the next section, I will discuss Nickell bias and then demonstrate in Section (ref) that, in simulations, Nickell bias is much smaller than dynamic bias. Regardless of the values of $\rho_{10}$ or $\tau_{0}$, I show that in simulation that dynamic bias is consistently larger than Nickell bias.
The DBC estimator works both when treatment is random and when treatment is a function of past outcomes and is therefore endogenous.
To explain how the bias correction works, I assume the following simple data-generating processes for the remainder of the section. However, the results of the paper hold for a richer class of models.\footnote{The richer class of models are given in Section (ref) below.} Equation (ref) shows that the treatment is a function of the past outcome. Recall biases get worse if, in addition to past outcome not being controlled for, treatment is related to past outcomes in the true model (not random). This is because in addition to dynamic bias there is now omitted variable bias. Equation (ref) shows that outcomes are a function of the past outcome and treatment. This simple example highlights the causal parameter of interest, inferential goal, and key assumptions of this approach.
where as before $\epsilon_{i,t}$ and $u_{i,t}$ are i.i.d mean zero true errors with variance $\sigma^2_{i,\epsilon}$ and $\sigma^2_{i,u}$.
Nickell bias occurs when a regression includes both estimated fixed effects and past outcome variables as regressors. Since fixed effects are not observed, they have to be estimated with dummies, the within-transformation, or first differences. These transformations impact all parts of the model, including the error terms. Taking the within transform as an example, the transformed error term becomes a function of error terms in all time periods. Therefore, if a regressor is a function of past error terms (like lagged outcome variables), it becomes correlated with the new within-transformed errors.
\paragraph{Model 3: Within-Transformed Dynamic Model}
\paragraph{OLS Estimator.} One could estimate treatment effects with a dynamic model by using OLS to estimate Equation (ref). This OLS regression leads to estimates of $\hat{\tau}^{NB}$ and $\hat{\rho}_1^{NB}$ that have Nickell bias. Though less common in practice, researchers could also use OLS to estimate Equation (ref) to get the estimate $\hat{\rho}_2^{NB}$. Our vector of estimated OLS parameters is $\hat{\theta}^{NB} := (\hat{\rho}_1^{NB}, \hat{\tau}^{NB}, \hat{\rho}_2^{NB}) $. The estimated residuals for a given parameter vector $\hat{\theta}$ are given by $\Tilde{\epsilon}_{i,t}(\hat{\theta}) := \Tilde{Y}_{i,t} - \hat{\tau} \Tilde{D}_{i,t} - \hat{\rho}_{1} \Tilde{Y}_{i,t-1} $ and $\Tilde{u}_{i,t}(\hat{\theta}) := \Tilde{D}_{i,t} - \hat{\rho}_2 \Tilde{Y}_{i,t -1}$ . By definition we have that $\Tilde{\epsilon}_{i,t}(\theta_0) = \Tilde{\epsilon}_{i,t}$ and $\Tilde{u}_{i,t}(\theta_0) = \Tilde{u}_{i,t} $.
To see concretely how the Nickell bias arises, focus on the OLS estimation of Equation (ref). I use the familiar formula for our OLS coefficients \big($\hat{\theta} = (\mathbb{X}'\mathbb{X})^{-1}\mathbb{X}'\mathbb{Y}$\big) plugging in our stack of covariates, a $(N \cdot N_T) \times 2$ matrix $\mathbb{X}$ whose row $i,t$ is $\mathbb{X}_{i,t} = [\Tilde{Y}_{i, t-1} \quad \Tilde{D}_{i,t}]'$, and the outcomes given in vector $\mathbb{Y}$ where the $i,t$-th element is $\Tilde{Y}_{i, t}$.
If the errors were strictly exogenous the expectation of term (ref).3 would be zero, and the OLS estimator would be unbiased. This would imply that ${\rm I}\kern-0.18em{\rm E}[{\Tilde{Y}_{i, t-1}} {\Tilde{\epsilon}_{i, t}} ] = 0$ and ${\rm I}\kern-0.18em{\rm E}[{\Tilde{D}_{i,t}}{\Tilde{\epsilon}_{i, t}} ] = 0$. However, given the model in Assumption (ref) the expectation of term (ref).3 is not zero. This is because ${\rm I}\kern-0.18em{\rm E}[{\Tilde{Y}_{i, t-1}} {\Tilde{\epsilon}_{i, t}} ] \neq 0$ and ${\rm I}\kern-0.18em{\rm E}[{\Tilde{D}_{i,t}} {\Tilde{\epsilon}_{i, t}} ] \neq 0$ . As an example, let us focus on why ${\rm I}\kern-0.18em{\rm E}[{\Tilde{Y}_{i, t-1}} {\Tilde{\epsilon}_{i, t}} ] \neq 0$. It follows that ${\Tilde{Y}_{i, t-1}}$ and ${\Tilde{\epsilon}_{i, t}}$ are correlated by observing that they are both correlated with {$\epsilon_{i,t -1}$}. First, ${\Tilde{Y}_{i, t-1}}$ is a function of ${Y}_{i,t-1}$, which itself is a function of {$\epsilon_{i,t -1}$}. Second, ${\Tilde{\epsilon}_{i, t}} = {\epsilon}_{i,t} - \bar{\epsilon}_{i,t}$ and $\bar{\epsilon}_{i,t} = \frac{1}{T} (e_{i,0} + \cdots {\epsilon_{i,t -1}}$ $+ \cdots + \epsilon_{i,T} )$. Similar logic is used to understand why ${\rm I}\kern-0.18em{\rm E}[{\Tilde{D}_{i,t}} {\Tilde{\epsilon}_{i, t}} ] \neq 0$.
Though my de-biased DBC estimator uses a GMM framework, I first build an understanding for how the bias correction works using OLS. Focusing on the OLS estimates in Equation (ref) , I calculate an expression for the asymptotic OLS estimator bias (term (ref).2 + (ref).3). The bias correction is created by calculating an estimate of this bias term, and the de-biased estimator by re-centering the GMM moment condition.
Rewriting Equation (ref), I get an expression of the asymptotic bias of the OLS estimator.
The bias correction comes from solving for $[\ref{eq:within_hat}.3]$ and subtracting the estimated bias out. Therefore the bias-corrected estimate $\hat{\theta}^{DBC} = \hat{\theta}^{NB} - \hat{\text{Bias}}$. Since I am interested in $\theta_0 = (\rho_{10}, \tau_0, \rho_{20})$, not just $\rho_{10}$ and $\tau_0$, I need to estimate both Equations $\eqref{eq:ols_within}$ and $\eqref{eq:ols_within_treat}$. In order to estimate both equation simultaneously, I cast the problem as a GMM system.
The original moment conditions, which contain bias, for estimating the parameters in Equations $\eqref{eq:ols_within}$ and $\eqref{eq:ols_within_treat}$ are given in Equation (ref). I index the moment conditions by $iT$ to make it clear that they are functions of the data and that the number of time periods $T$ is fixed.
Because of Nickell bias, the expectation of these moment equations are not zero at the true parameter $\theta_0$. For each moment in $\mathbf{m}_{iT}(\theta)$, however, I can create a new de-biased moment by subtracting the mean of each original moment equation at $\theta_0$. By construction, this new moment that is mean zero at the true parameters. This term that I subtract out is labeled $\mathbf{b}_{iT}(\theta_0)$ and is written
The analytic bias correction comes from solving for $\mathbf{b}_{iT}(\theta_0)$.
\paragraph{Bias Correction Formula Intuition.} A detailed calculation $\mathbf{b}_{iT}(\theta_0)$ is given in Appendix (ref). However, here, I give intuition for how the calculation works. As an example, consider $b_{\rho_1,iT}(\theta) = {\rm I}\kern-0.18em{\rm E}[\frac{1}{T} \sum_{t=1}^T \Tilde{Y}_{i,t-1}\Tilde{\epsilon}_{i,t}(\theta_0)]$. To calculate this expectation, we need to derive how $\Tilde{Y}_{i,t-1}$ relates to $\Tilde{\epsilon}_{i,t}$. To do this, it is helpful to unpack $\Tilde{Y}_{i,t-1}$ by looking just at ${Y}_{i,t-1}$ and how it depends on $\epsilon_{i,t}$ terms in all time periods. I use the outcome model in Assumption (ref) to write $Y_{i,t-1}$ as a sum of past outcomes and errors by iteratively plugging in the definition of past outcomes and treatment.\footnote{This formula uses the fact that treatment is varying over time. If it is the case that treatment is “absorbing” (once a unit starts treatment they stay treated), a different formula applies, given in Appendix (ref).}:
Given equation (ref), it is clear how $Y_{i,t-1}$ relates to $\epsilon_{i,t}$ in all time periods. The dependence comes from the $\sum_{j=0}^{t-2} (\rho_1 + \tau \rho_2)^j \epsilon_{i,t-j-1}$ term. Using the properties of geometric sums and accounting for the fact that I am calculating the expectation of the within-transformed $\Tilde{Y}_{i,t-1}$ and $\Tilde{\epsilon}_{i,t}$, I calculate the following bias correction terms.
Here $\phi(\theta) = (\rho_{10} + \tau_{0} \rho_{20}) $. Proofs are given in Appendix (ref). The true variance parameters $\sigma_{\epsilon,i}^2$ and $\sigma_{u,i}^2$ are not observed and have to be estimated. The estimated bias correction terms are the following.
\paragraph{Estimated Bias Correction Terms.}
Here ${\rm I}\kern-0.18em{\rm E}[\hat{\sigma}_{\epsilon, iT}^2(\theta_0)] = {\sigma}_{\epsilon, iT}^2$ and ${\rm I}\kern-0.18em{\rm E}[\hat{\sigma}_{u, iT}^2(\theta_0)] = {\sigma}_{u, iT}^2$.
The bias-corrected moment equations are the original OLS moment equations minus the bias correction term:
Given the definition of the model, only at $\theta_0$ do the residuals in the original moment equation equal the true noise variables in the bias correction term. Consequently, the moment conditions for the the de-biased moment equations are satisfied:
\paragraph{Bias-Corrected Estimator.}
The DBC estimator $\hat{\theta}_{DBC}$ is obtained by solving the GMM objective function under sample moment conditions:
I now derive asymptotic properties. Since the proposed estimator is based on GMM, asymptotic normality is established following standard results from the GMM framework.
This assumption ensures that the process is stationary, meaning that the statistical properties (such as the mean and variance) of the outcome are stable over time.
The complete expressions for $\Omega$, $G$ and the proof are given in Appendix (ref).
The bias correction formula for the more following more general model is given in Appendix (ref).
Allowing the past outcomes $Y_{i,t-1}$ through $Y_{i, t-h}$ to impact current potential outcomes follow from breitung2022bias and general VAR structure follows from juodis2015iterative. In an in-progress extension to this paper, I have a machine learning estimator that still imposes additive fixed effects and errors, but allows for more flexible modeling of the regressors, such as the interaction terms in the model above. A discussion of this extension is given in Appendix (ref). Note that a general nonparmetric form with a non-separable fixed effect is not possible, especially with fixed $T$.\footnote{The current methods that control for time varying unobserved heterogeneity rely on large $T$ asymptotics so they can estimate latent factor structures, which is not the setting of this paper moon2017dynamic. }
In this section, I present a simulation study to compare the dynamic and Nickell bias along several dimensions. I provide additional details for the simulation that I introduced in Section (ref) as well as additional results comparing my DBC estimator to other Nickell bias correction procedures.
I start with a Monte Carlo simulation using the structural model presented in Assumption (ref), which I reproduce here:
In Equation (ref), the outcome is a function of the treatment and the past outcome. For the treatment model (Equation (ref)), I make treatment a function of the fixed effect, in line with the common practice in applied research of adding fixed effects to account for correlation between treatment and a time-invariant individual-level term.\footnote{It is important to note dynamic bias remains even if treatment is not a function of fixed effects and $D_{i,t} = u_{i,t}$.} This model is the simplest setting, where $D_{i,t}$ is not impacted by past $Y_{i,t-1}$. The errors $u_{i,t}$ and $\epsilon_{i,t}$ are i.i.d random noise.
The individual fixed effects are drawn from a normal distribution $a_i \sim N(0,5).$ The treatment is set to $\tau_0 = .5$. I vary the values of $\rho_{10}$. \footnote{Though I show simulations for this setting, Nickell bias is smaller than dynamic bias in all DGPs I have studied. Additional simulation results for a wide variety of parameters are given in Appendix(ref)}
\paragraph{Simulation estimators:}
I compare the OLS estimates of three different models.
I repeat the same data generating process for a range of different panel lengths. For each number of time periods $N_T:$ 3 - 50, I generate datasets and plot the OLS estimate of the Static, Dynamic, and Delta Models.
Figure (ref) (replicated from Section (ref)) plots the results. On the x-axis we have the number of time periods in the generated panel data set, and then the y-axis marks the treatment estimate. The true treatment effect value is marked with a black horizontal line.
The plots show that even when treatment is random, dynamic bias and transformation bias are larger than Nickell bias. The Nickell bias is small because treatment effect coefficient is biased only because the coefficient on the past outcome is estimated with bias, which then leads to bias in the other parameters in the model.
Now let us consider the case when treatment is not random. I modify the DGP to follow the structural model outlined in Assumption (ref):
I set $\rho_{20} = .1$, and generate results analogously to those in the previous section, but now using the endogenous treatment simulation design.
Figure (ref) present the results. The bias gets larger the larger the absolute value of $\rho_{20}$. Recall that when treatment is endogenous, the bias for the Static and Delta model does not disappear as the number of time periods increases omitted variable bias arises when past outcomes are not explicitly included in the model. Note in Figure (ref), for the case of $\rho_{10} = .9$, the bias under the Static Model is so large it is omitted from the plot.
In these simulations I showcase the performance of my bias-corrected estimator. I also show how my estimator performs in comparison to the Arellano Bond estimator arellano1991some.
I create a simulation data generating process with an endogenous treatment following Equations (ref) and (ref). Again, the individual fixed effects are drawn from a standard normal distribution $a_i \sim N(0,5).$ The initial values $D_{i0} \sim N(a_i, 1) $. The vector of the true parameters is given by $\theta_0 = (\rho_{10}, \tau_0, \rho_{20}) = (.2,.5,.3)$.
\paragraph{Simulation Estimators:}
I run 1000 Monte Carlo simulations with the number of units $N = 1000$ and the number of time periods $N_T = 5$.
Table (ref) compares three different estimation strategies for treatment effects. The table compares my bias-corrected (DBC) estimator and ordinary least squares (OLS) estimates and Arellano Bond (AB) estimates. The DBC estimates are closest to the true treatment parameter $\tau_0 = .5$. The table presents the mean and standard deviation (sd) for each estimate, showing that DBC estimates yield proper 95% coverage, while OLS does not. The Arellano Bond (AB) estimator uses several lagged variables as instruments. The data was generated with random i.i.d errors, and so by construction the instruments are valid. The results demonstrate that AB models also achieve proper 95% coverage. I find that proper coverage is achieved regardless of the instruments used. However, the instrument choice does have a substantial effect on the standard deviations of the estimates. Using fewer instruments leads to an even larger standard deviation for the AB estimates. The large standard deviations explain why in practice for a particular dataset (as opposed to our Monte Carlo setting here where we average over 1000 datasets) the choice of instruments can lead to very different treatment effect estimates. In the table (ref) present the AB results in which I selected the instruments that led to the smallest standard deviation for treatment ($\tau$), using all possible instruments. Yet, even in this case, AB estimation still leads to standard deviations that are larger in comparison to DBC.
Table (ref) compares the same three different estimation strategies for $\rho_{10} = .2$. While OLS led to biased estimates of treatment effects, the bias of the $\rho_{10}$ is much larger and even changes sign. Both my DBC method and AB are able to achieve proper coverage of the true $\rho_{10}$, but the standard deviation of method is again substantially smaller for my estimator.
This paper is highly relevant for any applied research where past outcomes influence current outcomes. As discussed in the introduction, this includes studies focusing on variables such as agricultural yields, human capital, labor market outcomes, and migration flows.
To illustrate the application of my method, I focus on the relationship between temperature (as the treatment) and GDP (as the outcome). This relationship has been extensively explored in prior literature, and for this analysis, I utilize data from the seminal study by dell2012temperature.
dell2012temperature explores the relationship between temperature ($D_{i,t}$) and GDP ($Y_{i,t}$), and examines how rising temperatures can impact economic growth. The authors' results suggest that warmer temperatures can lead to reduced agricultural yields, decreased labor productivity, and increased health risk, all of which can hinder economic outcomes. The main result of their paper focuses on how higher temperatures affect poorer countries.
I use this empirical example to highlight two main points. First, I use the example to show how dynamic bias, transformation bias, and Nickell bias compare in a real world setting. This is discussed in detail in Section (ref). Second, I highlight how my bias correction performs in practice and compare it to Arellano Bond. This is discussed in detail in Section (ref).
In Table (ref), I present estimates from regressions of GDP growth and levels on temperature using data from dell2012temperature. The main results of dell2012temperature focus on the impact of temperature on the GDP of poor countries, so my sample is restricted to poor countries. I do this for simplicity so that this analysis can focus on one treatment effect estimate. I exactly replicate the original analysis that uses all countries dell2012temperature in Appendix (ref). The replication there is consistent with the results presented here.
In Table (ref), I present the subset of results focused on poor countries. In the first two columns, I run baseline regressions where I regress GDP growth on temperature in the first column and then the second column shows how these estimates change once lagged GDP growth is included. In the third column and fourth columns I include all the original controls used in dell2012temperature, which are 371 region and time controls . The third column of my Table (ref) follows the same specification given in the second column of Table 2 in dell2012temperature - the outcome is GDP growth, and the past outcome is not controlled for. The treatment effect estimate without controlling for the lagged outcome is -1.421. Then in Column 4 I include lagged GDP growth, and the treatment effect estimate changes to -1.279. This is a 10% difference in treatment effects that is significant at the $p < .1$ level.\footnote{To test whether the treatment estimate in Column 3 $(\hat{\tau}_3)$ is statistically significantly smaller than the treatment estimate in Column 4, $(\hat{\tau}_4)$ I conduct Wald test and the resulting p-value is .06. The Wald test requires accounting for the correlation between $\hat{\tau}_3$ and $\hat{\tau}_4$. These two estimates are highly correlated, which is why even though the estimates have a large variance, the null hypothesis is still rejected at the p < .1 level. }
These changes in coefficient occur despite the fact that I include the large number of time trend controls used in the original analysis. This highlights the point discussed in the introduction: time trend controls do not control for dynamics. To see this analytically, I run additional specifications to see how treatment effect estimates are impacted by the controls. In Table (ref) Column 1 and 2 I run the same specifications in Column 3 and 4 respectively, just without controls. Adding controls increases the estimated coefficient, whereas controlling for lags decreases the estimate -- demonstrating that controls do not correct for dynamic relationships in outcomes. Adding the lagged outcome changes the treatment effect estimate by the same percentage regardless of whether or not controls are included.
It is important to note that GDP growth is calculated by transforming GDP levels. Looking at GDP growth rather then GDP levels may somewhat reduce dynamic bias, but it introduces transformation bias, as discussed in Section (ref) and Appendix (ref). The magnitudes of both biases depend on $\rho_{10}$ and the number of time periods. In this particular setting, the true $\rho_{10}$ is very close to 1 and the number of time periods is 30, so the transformation bias is not as large as the dynamic bias.
To understand the impact of temperature on the original untransformed variable, I use GDP levels as the outcome in Table (ref) Column 5 and Column 6. This is important because as discussed in nath2024much, the economic literature is divided over whether temperature affects GDP levels or growth rates. Therefore, it is also important to study GDP levels. Levels of outcomes are often studied in economics papers. \footnote{Many economics papers study the levels of outcome variables without controlling for past outcomes. somanathan2021impact study the impact of temperature on economic productivity without controlling for levels, annan2015federal study the temperature effects of crop yield levels.} The impact of including past outcomes in the regression models is even greater when looking at GDP levels. Table (ref) Column 5 shows that without controlling for past outcome the treatment estimate is unexpectedly positive (174.202), and only becomes negative (-19.330) as expected when controlling for past GDP level in Column 6.
In this section I compare the performance of my bias correction method to Arellano Bond using real data from dell2012temperature. In the case of dell2012temperature, the number of time periods is 30, and so the expectation is that Nickell bias is relatively small.
I replicate the analysis of Table (ref) Column 2 and present the resulting estimates of $\tau_0$ and $\rho_{10}$ are presented in the “OLS with Lag” columns in Table (ref) and Table (ref), respectively. I do the analysis on a balanced panel; therefore, the OLS estimates of $\tau_0$ and $\rho_{10}$ are slightly different from those in Table (ref) where I used the unbalanced panel of dell2012temperature.\footnote{The extension to unbalanced panels can be implemented but is left for future work.}
I then implement my bias correction method for both parameters; the results are given in the bias-corrected columns labeled “DBC”. The bias-corrected method does not change the estimate of treatment $\tau_0$ effect much, but does lead to higher estimate of ${\rho}_{10}$. This is expected as Nickell bias is small in the setting when the number of time periods is 30, and also Nickell bias impacts the estimation of ${\rho}_{10}$ more than the estimation of $\tau_0$.
The tables also include the Arellano Bond estimates. arellano1991some is one of the most cited papers in economics as it was one of the first papers to correct for Nickell bias. This approach requires picking which past outcomes are used as instruments, but point estimates of both $\tau_0$ and $\rho_{10}$ are quite sensitive to this choice. I run three different Arellano Bond specifications, changing which lags used as instruments - these results are reported in columns “AB 1”, “AB 2” and “AB 3”.\footnote{For specification AB 1, I use the lag 2-5 for both variables, for AB 2 I used lags from 2-15, for AB 3 I use lags 10-15.} Looking first at the treatment point estimates in Table (ref), depending on the choice of instrument the point estimates go from negative (-.460) to positive and become an order of magnitude smaller (.060). A similar pattern occurs with the point estimates of $\hat{\rho}_{10}$ given in Table (ref). Depending on the choice of instrument, the point estimates go from negative (-.016) to positive and change by an order of magnitude larger (.101). Avoiding the instability associated with the choice of instrument is a benefit of my analytical approach.
The standard errors for all methods were calculated with the panel bootstrap clustered at the country level. The standard errors of the bias-corrected method are smaller than the standard errors of Arellano Bond method, and are close to the original OLS standard errors, for both $\hat{\tau}$ and $\hat{\rho}_{1}$. Depending on the instruments, the standard errors of AB vary, and can be twice as large as the OLS and DBC standard errors.
This paper identifies a source of bias in fixed effects panel models, which I term dynamic bias, which occurs when past outcomes are left out of the model. Economic research often uses static models even in settings in which economics theory tells us that past outcomes directly influence current outcomesgriliches1963sources,cunha2007technology, blanchard1988beyond,massey1993theories. Empirical papers studying these outcomes often run fixed effects analysis without controlling for past outcomes.\footnote{Examples include annan2015federal, burke2015global, cho2017effects, jessoe2018climate, drabo2015natural, mahajan2020taken, missirian2017asylum, graff2018temperature garg2020temperature.}
Through simulations I show that dynamic bias is larger than the well-known Nickell bias. While Nickell bias comes from including past outcomes in the model, dynamic bias occurs when past outcomes are excluded. Through simulations and real-world data, I demonstrate that ignoring past outcomes can lead to significantly biased treatment effect estimates, even when treatments, such as temperature, are randomly assigned. This bias is especially problematic in contexts like environmental economics, where past environmental conditions affect current environmental outcomes.
To solve this challenge, I developed a new estimator (the DBC estimator) that corrects both dynamic bias and Nickell bias. The estimator works well even when the number of time periods is small (fixed-T panels). It provides more accurate estimates of treatment effects by properly accounting for the influence of past outcomes. The DBC estimator performs better than alternative methods, such as instrumental variable techniques, which can suffer from weak instruments and larger standard errors.
I applied the DBC estimator to study the impact of temperature shocks on GDP, a key question in environmental economics. The results show that properly accounting for dynamic biases significantly changes the estimated effects. For instance, correcting for dynamic biases reduces the estimated impact of temperature shocks on GDP growth by 10% and on GDP levels by 120%.
This research offers a practical solution for applied researchers dealing with dynamic settings. It highlights the importance of including past outcomes to avoid large biases in treatment effect estimates. Future work will focus on extending this estimator to more complex models by incorporating machine learning and adapting it for spatial data, where additional geographic dynamics play an important role.