Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
21,309 characters · 8 sections · 10 citation commands
On the Effect of Imputation on the 2SLS Variance
\onehalfspacing \baselineskip=19pt
Missing data are frequently encountered in empirical studies in economics and social sciences. A popular method to handle missing data is the complete case approach, which excludes incomplete observations from the analysis. Among others, an alternative approach is regression imputation, which utilizes the complete observations to fill in the missing values. The imputed data is then used as if it was observed in the first place.
While the imputation of an exogenous regressor is a well-researched topic little1992missingx, little is known about the imputation of an endogenous regressor. Two-stage least squares (2SLS) estimation is a way to deal with the endogeneity of a regressor. MCDONOUGH2017 discussed the bias of 2SLS with imputation. Their analysis was mainly based on NAGAR1959's seminal work about the finite-sample bias of 2SLS. However, they did not discuss variance estimation, which is challenging as we show in this study. 2SLS inference is affected by the imputation implying that the conventional variance estimator, which ignores the imputation, is only valid if we are interested to test whether the parameter of the endogenous regressors is zero. Standard errors and confidence intervals based on the conventional variance are invalid. We obtain the asymptotic distribution and derive a heteroskedasticity robust variance estimator for 2SLS with regression imputation, which allows to construct valid standard errors, confidence intervals and conduct tests.
We focus on settings where the endogenous regressor is missing at random, the number of instruments is fixed as the sample size grows and the imputation method is regression imputation. The missing at random setting is essentially selction based on observables wooldridge2007 and covers many relevant applications wooldridge2007,graham2012,chaudhuri2018. To simplify the comparison with the conventional variance estimation, we assume missing completely at random in parts of the study.
We illustrate our theoretical results using Monte Carlo simulations. In the simulation results we focus on the key parameters in the interplay between 2SLS estimation and regression imputation. These key parameters are the fraction of missings, which can be observed easily, and the direction of the OLS bias, which is unobserved but very often researchers have strong prior beliefs about it.
In the next section we describe the setup including model, missing data structure and a brief description of regression imputation. In Section (ref), we derive the asymptotic variance and propose a variance estimator which accounts for the imputation and is robust to general forms of heteroskedasticity. Section (ref) shows Monte Carlo simulation results and Section (ref) concludes.
Consider the standard simultaneous equation model
with dependent variable $(y_i)$, endogenous regressor $(x_i)$ and a set of $L$ instruments $(Z_{i.})$.\footnote{To simplify notation, we abstract from exogenous control variables.} The parameter of interest is $\beta$. The two-stage least squares (2SLS) estimator is a way to deal with the endogeneity of $x$:
where $P_Z=Z(Z'Z)^{-1}Z'$. Throughout the paper we assume that 2SLS(ref)-2SLS(ref) are fulfilled in model ((ref)), which assures consistency and valid inference of the 2SLS estimator.
The analysis is, however, complicated by the fact that data on the endogenous regressor, $x$, is missing for some observations causing them to be incomplete. Therefore, $\widehat\beta$ as defined above cannot be calculated with the data at hand. Let the subscripts indicate the missing status, that is, $y_0, x_0, Z_0$ indicate the complete observations and $y_1, x_1, Z_1$ the observations with missing value of $x$. Let $\widehat{p}$ be the probability of missings in the endogenous regressor and assume
which implies that not only the number of incomplete observations ($n_1$) but also the number of complete observations ($n_0$) increase as $n$ increases. Without loss of generality we assume that the first $n_0$ observations of the matrix $Z$ and the vectors $x$ and $y$ are complete observation and the remaining $n_1$ observations are incomplete.
The missing data literature distinguishes three types of missing structures: missing completely at random (MCAR), missing at random (MAR)\footnote{That is the missing structure only depends on observables.}, and not missing at random (NMAR). While MCAR is the easiest to deal with, real data is often NMAR or MAR. The missing structure is called ignorable if the data is either MCAR or MAR. Since dealing with NMAR is very different from the other two types, we focus on missing (completely) at random in this exposition, and use the following additional assumption:
Regression imputation (RI), which is a two-step procedure, can be applied as a tool to fill in the missing values. In the first step the complete observations are used to regress the endogenous variable $(x_0)$ on the imputation variables to obtain the imputation parameters. These parameters are equal to the first stage estimates from the complete case approach $\left(\widehat\pi_{CC}=(Z_0'Z_0)^{-1}Z_0'x_0\right)$ if the imputation method incorporates the instruments.\footnote{MCDONOUGH2017 show in Monte Carlo simulations that imputation methods, which incorporate the instruments, produce the smallest finite sample bias of 2SLS estimation. } In the second step these estimates are employed to impute the missing values in $x_1$ by multiplying the imputation variables for the incomplete observations with the parameters obtained from the first step. The imputed variable is then used in the 2SLS estimation as if it was observed in the first place.\footnote{Clearly, the observations with missing values cannot add any information to the estimation of the first stage parameters. Hence, the relevant $F$-statistic for the first stage parameters should be based on the complete case observations.}
The model in ((ref)) can be restated in terms of $\widetilde{x}$ by adding an imputation error:
The imputation error $e_1$ and the composite errors of the imputed model ($\widetilde{v}$ & $\widetilde{u}$) are defined as
While the 2SLS estimator with regression imputation remains consistent if 2SLS (ref) to (ref) and (ref) to (ref) are fulfilled, inference is affected by the imputation. The independence across observations is violated in the imputed model as the complete data has been used to impute incomplete observations. In the next section, we derive a variance estimator which takes this into account.
We consider
and derive its limiting distribution under heteroskedasticity in the next proposition.
The first term of $\Omega$ corresponds to the standard asymptotic variance of 2SLS without missing data. The remaining terms can be attributed to regression imputation. If no data is missing ($p=0$) or $\beta=0$, $\Omega$ collapses to the standard asymptotic variance of 2SLS. The latter simplification (i.e., $\beta=0$) is a well-known result in the literature about generated regressors murphy1985estimation. It allows valid inference based on the conventional variance estimator if and only if we are interested in tests of $\beta=0$, which often is of major interest in applications. However, it does not justify the construction of standard errors or confidence intervals.
The estimation of $\Omega$ is challenging as we cannot obtain reliable residual estimates for the imputed observations, $u_{1}$ and $v_{1}$ . Hence, the first, fourth and fifth term of $\Omega$ cannot be simply estimated by using the corresponding residuals. However, it is possible to circumvent this issue by using the error of the imputed model $(\widetilde{u}_{1})$ for which residuals $(\widehat{\widetilde{u}}_{1})$ can be obtained. The variance estimator given in the following proposition is consistent under general forms of heteroskedasticity.
In the following we compare the conventional variance estimator, which ignores the imputation, with the true asymptotic variance. To simplify the comparison, we assume homoskedasticity and MCAR. The asymptotic variance is then stated in the following corollary. \paragraph{Corollary 1.} Under the conditions of Proposition (ref), 2SLS (ref) and MCAR the asymptotic variance is given by
The conventional variance estimator calculated by standard statistical software is given by
and its limit under homoskedasticity and MCAR is,
Comparing the limit of the conventional estimators with the true asymptotic variance in Corollary 1, we can see that the conventional estimator does not reflect the true variance of $\widehat\beta_{RI}$. For instance, while the asymptotic variance always increases with the missing probability, the conventional estimator could even decrease in it (if $-2\sigma_{uv}\beta>\sigma_{v}^2\beta^2$). In this case the limit of the conventional estimator is smaller than the asymptotic variance without any missing data problem, which is clearly counterintuitive. Moreover, the degree of endogeneity ($\sigma_{uv}$) erroneously affects the conventional limit while it has indeed no effect on the true asymptotic variance.
We illustrate our theoretical findings in Monte Carlo simulations. To implement the missing data problem, we first mimic the model in ((ref)) using the data generating process described below, and then randomly delete the value $x_i$ with probability $p$. The missing structure is thus MCAR. For each of the $n$ observations $Z_{i.}$ and $v_i$ are drawn from $$Z_{i.}\sim{}N\left(0_{L},\frac{1}{L}I_L\right) \ ; \hspace{0.5cm} v_i\sim N(0,1) \, , $$ where $L$ denotes the number of instruments. We follow hausman2012instrumental and chao2014testing and define the structural error as, $$u_i=\sigma_{uv}v_i+\sqrt{\frac{1-\sigma_{uv}^2}{\phi+(0.86^2)}}(\phi\epsilon_{1i}+0.86\epsilon_{2i}),\;\;\epsilon_{1i}\sim{}N(0,Z_{i.}'Z_{i.}),\;\epsilon_{2i}\sim{}N(0,0.86^2),\;\phi=5 \, ,$$ where $\phi$ defines the strength of the heteroskedasticity. We set $\beta=0.5$, $L=3$, $n=1000$, number of Monte Carlo repetitions $R=5000$, and the step size, which we use to alter the probability of missings in the different simulations, to $\Delta{}p=0.005$. As derived in the previous section, the sign of the endogeneity is crucial. Hence, we show results for $\sigma_{uv}=0.3$ and $\sigma_{uv}=-0.3$. We set $\pi=\sqrt{\frac{F\,L}{n}}$. Note that we divide by $n$ and not by $n_0$. Therefore, $F$ defines the first-stage $F$-statistic in the entire sample containing both complete and incomplete observations, and the first-stage $F$-statistic in the complete case sample gradually decreases with the missing probability. We set $F=100$. That is, the complete case $F$-statistic is around $100$ at $p=1$ and around $20$ at $p=0.8$.
Figure (ref) compares our consistent and the conventional standard errors to the observed root mean squared error (RMSE) obtained from the Monte Carlo simulations. It shows that the conventional estimator cannot properly describe the standard error of the 2SLS estimator with regression imputation. The missing probability has a linear effect on the conventional estimator (see eq. (ref)) while it nonlinearly affects the true asymptotic variance as shown in Corollary 1. Moreover, as expected the conventional standard error can even decrease as the missing probability increases ($\vert\sigma_{uv}\beta\vert>\sigma^2_v \beta^2$ in the right graph of Figure (ref)). The reason why MCDONOUGH2017, who did not derive a consistent variance estimator for regression imputation, have acceptable results in their simulations is solely due to their chosen parameters ($\sigma_{uv}>0$, $\beta>0$, low missing probability) and cannot be generalized to the entire parameter space. Our variance estimator ($\widehat{V}_{\widehat\beta_{RI}}$) performs well in both settings. Additionally, Figure (ref) shows how often the null hypothesis ($H_0:\,\beta=0.5$) is rejected at the $5\%$ nominal level. Again, our proposed variance estimator performs better than the conventional one---particularly, if $\sigma_{uv}<0$.
We investigate how two issues, which are likely present in many empirical studies, affect the estimation of causal effects, namely the endogeneity of regressors and missing data. If researchers use an instrumental variables regression after single imputation of missing values for an endogenous regressor, they have to be aware that conventional methods to estimate the variance fail to account for the imputation. The asymptotic variance of 2SLS is affected by the imputation implying that conventional methods cannot be used to construct standard errors, confidence intervals and conduct tests. We derive a heteroskedastic variance estimator which takes the imputation into account and is consistent. Monte Carlo simulations show that our estimator performs well while the conventional variance estimator does not.