EconBase
← Back to paper

On the Effect of Imputation on the 2SLS Variance

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

21,309 characters · 8 sections · 10 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the Effect of Imputation on the 2SLS Variance

abstractEndogeneity and missing data are common issues in empirical research. We investigate how both jointly affect inference on causal parameters. Conventional methods to estimate the variance, which treat the imputed data as if it was observed in the first place, are not reliable. We derive the asymptotic variance and propose a heteroskedasticity robust variance estimator for two-stage least squares which accounts for the imputation. Monte Carlo simulations support our theoretical findings. Key Words: endogeneity, instrumental variables, imputation, variance estimation

\onehalfspacing \baselineskip=19pt

Introduction

Missing data are frequently encountered in empirical studies in economics and social sciences. A popular method to handle missing data is the complete case approach, which excludes incomplete observations from the analysis. Among others, an alternative approach is regression imputation, which utilizes the complete observations to fill in the missing values. The imputed data is then used as if it was observed in the first place.

While the imputation of an exogenous regressor is a well-researched topic little1992missingx, little is known about the imputation of an endogenous regressor. Two-stage least squares (2SLS) estimation is a way to deal with the endogeneity of a regressor. MCDONOUGH2017 discussed the bias of 2SLS with imputation. Their analysis was mainly based on NAGAR1959's seminal work about the finite-sample bias of 2SLS. However, they did not discuss variance estimation, which is challenging as we show in this study. 2SLS inference is affected by the imputation implying that the conventional variance estimator, which ignores the imputation, is only valid if we are interested to test whether the parameter of the endogenous regressors is zero. Standard errors and confidence intervals based on the conventional variance are invalid. We obtain the asymptotic distribution and derive a heteroskedasticity robust variance estimator for 2SLS with regression imputation, which allows to construct valid standard errors, confidence intervals and conduct tests.

We focus on settings where the endogenous regressor is missing at random, the number of instruments is fixed as the sample size grows and the imputation method is regression imputation. The missing at random setting is essentially selction based on observables wooldridge2007 and covers many relevant applications wooldridge2007,graham2012,chaudhuri2018. To simplify the comparison with the conventional variance estimation, we assume missing completely at random in parts of the study.

We illustrate our theoretical results using Monte Carlo simulations. In the simulation results we focus on the key parameters in the interplay between 2SLS estimation and regression imputation. These key parameters are the fraction of missings, which can be observed easily, and the direction of the OLS bias, which is unobserved but very often researchers have strong prior beliefs about it.

In the next section we describe the setup including model, missing data structure and a brief description of regression imputation. In Section (ref), we derive the asymptotic variance and propose a variance estimator which accounts for the imputation and is robust to general forms of heteroskedasticity. Section (ref) shows Monte Carlo simulation results and Section (ref) concludes.

Setup

Model

Consider the standard simultaneous equation model

align[align omitted — 106 chars of source]

with dependent variable $(y_i)$, endogenous regressor $(x_i)$ and a set of $L$ instruments $(Z_{i.})$.\footnote{To simplify notation, we abstract from exogenous control variables.} The parameter of interest is $\beta$. The two-stage least squares (2SLS) estimator is a way to deal with the endogeneity of $x$:

align[align omitted — 84 chars of source]

where $P_Z=Z(Z'Z)^{-1}Z'$. Throughout the paper we assume that 2SLS(ref)-2SLS(ref) are fulfilled in model ((ref)), which assures consistency and valid inference of the 2SLS estimator.

assumption$plim\left(\frac{Z'u}{n}\right)=E[Z_{i.}u_i]=0 \, $; $plim\left(\frac{Z'v}{n}\right)=E[Z_{i.}v_i]=0 \,$.
assumption$plim\left(\frac{Z'x}{n}\right)=E[Z_{i.}x_i]=Q_{Zx}\neq0 \, $.
assumption$plim\left(\frac{Z'Z}{n}\right)=E[Z_{i.}Z_{i.}']=Q_{ZZ}$, with $Q_{ZZ}$ a finite and full rank matrix.
assumptionObservations are identically and independently distributed.
assumptionErrors are homoskedastatic $E[uu'|Z]=\sigma_{u}^2I_n \,$; $E[vv'|Z]=\sigma_{v}^2I_n \,$; $E[vu'|Z]=\sigma_{uv}I_n \, $; $\sigma_{u}^2$, $\sigma_{v}^2$ and $\sigma_{uv}$ are finite.

Missing data

The analysis is, however, complicated by the fact that data on the endogenous regressor, $x$, is missing for some observations causing them to be incomplete. Therefore, $\widehat\beta$ as defined above cannot be calculated with the data at hand. Let the subscripts indicate the missing status, that is, $y_0, x_0, Z_0$ indicate the complete observations and $y_1, x_1, Z_1$ the observations with missing value of $x$. Let $\widehat{p}$ be the probability of missings in the endogenous regressor and assume

assumption$\widehat{p}=\frac{n_1}{n}\rightarrow{}p<1$ as $n\to\infty \, $,

which implies that not only the number of incomplete observations ($n_1$) but also the number of complete observations ($n_0$) increase as $n$ increases. Without loss of generality we assume that the first $n_0$ observations of the matrix $Z$ and the vectors $x$ and $y$ are complete observation and the remaining $n_1$ observations are incomplete.

The missing data literature distinguishes three types of missing structures: missing completely at random (MCAR), missing at random (MAR)\footnote{That is the missing structure only depends on observables.}, and not missing at random (NMAR). While MCAR is the easiest to deal with, real data is often NMAR or MAR. The missing structure is called ignorable if the data is either MCAR or MAR. Since dealing with NMAR is very different from the other two types, we focus on missing (completely) at random in this exposition, and use the following additional assumption:

assumptionThe missing structure is ignorable.

Regression imputation

Regression imputation (RI), which is a two-step procedure, can be applied as a tool to fill in the missing values. In the first step the complete observations are used to regress the endogenous variable $(x_0)$ on the imputation variables to obtain the imputation parameters. These parameters are equal to the first stage estimates from the complete case approach $\left(\widehat\pi_{CC}=(Z_0'Z_0)^{-1}Z_0'x_0\right)$ if the imputation method incorporates the instruments.\footnote{MCDONOUGH2017 show in Monte Carlo simulations that imputation methods, which incorporate the instruments, produce the smallest finite sample bias of 2SLS estimation. } In the second step these estimates are employed to impute the missing values in $x_1$ by multiplying the imputation variables for the incomplete observations with the parameters obtained from the first step. The imputed variable is then used in the 2SLS estimation as if it was observed in the first place.\footnote{Clearly, the observations with missing values cannot add any information to the estimation of the first stage parameters. Hence, the relevant $F$-statistic for the first stage parameters should be based on the complete case observations.}

The model in ((ref)) can be restated in terms of $\widetilde{x}$ by adding an imputation error:

align[align omitted — 338 chars of source]

The imputation error $e_1$ and the composite errors of the imputed model ($\widetilde{v}$ & $\widetilde{u}$) are defined as

align[align omitted — 492 chars of source]

While the 2SLS estimator with regression imputation remains consistent if 2SLS (ref) to (ref) and (ref) to (ref) are fulfilled, inference is affected by the imputation. The independence across observations is violated in the imputed model as the complete data has been used to impute incomplete observations. In the next section, we derive a variance estimator which takes this into account.

Estimation and Inference

We consider

align[align omitted — 94 chars of source]

and derive its limiting distribution under heteroskedasticity in the next proposition.

propositionUnder Assumptions 2SLS (ref)-(ref), 2SLS (ref)-(ref), $E\left[y_i^4\right]$, $E\left[\widetilde{x}_i^4\right]$, $E\left[\Vert Z_{i.}\Vert^4\right]$ being finite and $n\rightarrow\infty$, we have that \begin{align} \begin{split} \sqrt{n}(\widehat{\beta}_{RI}-\beta)\xrightarrow[]{d}N(0,V_{\widehat\beta_{RI}}) \end{split} \end{align} with the asymptotic variance given by, \begin{align*} V_{\widehat\beta_{RI}}=\left(Q_{xZ}Q_{ZZ}^{-1}Q_{Zx}\right)^{-1}Q_{xZ}Q_{ZZ}^{-1}\,\Omega\,Q_{ZZ}^{-1}Q_{Zx}\left(Q_{xZ}Q_{ZZ}^{-1}Q_{Zx}\right)^{-1} \end{align*} where, \begin{align*} \Omega=&E[u_{i}^2 Z_{i.}Z_{i.}']-2pE[u_{0i}v_{0i}Z_{0i.}Z_{0i.}']Q_{Z_0Z_0}^{-1}Q_{Z_1Z_1}\beta\\ &+\frac{p^2}{1-p}Q_{Z_1Z_1}Q_{Z_0Z_0}^{-1}E[v_{0i}^{2}Z_{0i.}Z_{0i.}']Q_{Z_0Z_0}^{-1}Q_{Z_1Z_1}\beta^2\\ &+2pE[u_{1i}v_{1i}Z_{1i.}Z_{1i.}']\beta+pE[v_{1i}^2 Z_{1i.}Z_{1i.}']\beta^2 \end{align*}
proofSee Appendix (ref).

The first term of $\Omega$ corresponds to the standard asymptotic variance of 2SLS without missing data. The remaining terms can be attributed to regression imputation. If no data is missing ($p=0$) or $\beta=0$, $\Omega$ collapses to the standard asymptotic variance of 2SLS. The latter simplification (i.e., $\beta=0$) is a well-known result in the literature about generated regressors murphy1985estimation. It allows valid inference based on the conventional variance estimator if and only if we are interested in tests of $\beta=0$, which often is of major interest in applications. However, it does not justify the construction of standard errors or confidence intervals.

The estimation of $\Omega$ is challenging as we cannot obtain reliable residual estimates for the imputed observations, $u_{1}$ and $v_{1}$ . Hence, the first, fourth and fifth term of $\Omega$ cannot be simply estimated by using the corresponding residuals. However, it is possible to circumvent this issue by using the error of the imputed model $(\widetilde{u}_{1})$ for which residuals $(\widehat{\widetilde{u}}_{1})$ can be obtained. The variance estimator given in the following proposition is consistent under general forms of heteroskedasticity.

propositionUnder Assumptions 2SLS (ref)-(ref), 2SLS (ref)-(ref) and $E\left[y_i^4\right]$, $E\left[\widetilde{x}_i^4\right]$, $E\left[\Vert Z_{i.}\Vert^4\right]$ being finite, we have that \begin{align*} \widehat{V}_{\widehat{\beta}_{RI}}=\left(\widetilde{x}'P_Z\widetilde{x}\right)^{-1}\widetilde{x}'Z(Z'Z)^{-1}\widehat{W}_{RI}(Z'Z)^{-1}Z'\widetilde{x}\left(\widetilde{x}'P_Z\widetilde{x}\right)^{-1} \end{align*} is a consistent estimator of the asymptotic variance ($n\widehat{V}_{\widehat\beta_{RI}}\overset{p}{\to}V_{\widehat\beta_{RI}}$) with \begin{align*} \widehat{W}_{RI}=&\left(\sum_{i=1}^{n}\widehat{\widetilde{u}}_{i}^2Z_{i.}Z_{i.}'\right)-2\left(\sum_{i=1}^{n_0}\widehat{\widetilde{u}}_{i}\widehat{\widetilde{v}}_{i}Z_{i.}Z_{i.}'\right)\left(\sum_{i=1}^{n_0}Z_{i.}Z_{i.}'\right)^{-1}\left(\sum_{i=n_0+1}^{n}Z_{i.}Z_{i.}'\right)\widehat{\beta}_{RI}\\ &+\Bigg[\left(\sum_{i=n_0+1}^{n}Z_{i.}Z_{i.}'\right)\left(\sum_{i=1}^{n_0}Z_{i.}Z_{i.}'\right)^{-1}\left(\sum_{i=1}^{n_0}\widehat{\widetilde{v}}_{i}^2Z_{i.}Z_{i.}'\right)\left(\sum_{i=1}^{n_0}Z_{i.}Z_{i.}'\right)^{-1}\left(\sum_{i=n_0+1}^{n}Z_{i.}Z_{i.}'\right)\\ &-\sum_{i=n_0+1}^{n}Z_{i.}Z_{i.}'\left(\sum_{i=1}^{n_0}Z_{i.}Z_{i.}'\right)^{-1} \left(\sum_{i=1}^{n_0}\widehat{\widetilde{v}}_{i}^2Z_{i.}Z_{i.}'\right)\left(\sum_{i=1}^{n_0}Z_{i.}Z_{i.}'\right)^{-1}Z_{i.}Z_{i.}'\Bigg]\widehat{\beta}_{RI}^2 \end{align*} where $\widehat{\widetilde{u}}_i=y_i-\widetilde{x}_i\widehat\beta_{RI}$ and $\widehat{\widetilde{v}}_i=\widetilde{x}_i-Z_{i.}'\widehat\pi_{CC}$
proofSee Appendix (ref).

In the following we compare the conventional variance estimator, which ignores the imputation, with the true asymptotic variance. To simplify the comparison, we assume homoskedasticity and MCAR. The asymptotic variance is then stated in the following corollary. \paragraph{Corollary 1.} Under the conditions of Proposition (ref), 2SLS (ref) and MCAR the asymptotic variance is given by

align[align omitted — 233 chars of source]
proofSee Appendix (ref).

The conventional variance estimator calculated by standard statistical software is given by

align[align omitted — 367 chars of source]

and its limit under homoskedasticity and MCAR is,

align[align omitted — 273 chars of source]

Comparing the limit of the conventional estimators with the true asymptotic variance in Corollary 1, we can see that the conventional estimator does not reflect the true variance of $\widehat\beta_{RI}$. For instance, while the asymptotic variance always increases with the missing probability, the conventional estimator could even decrease in it (if $-2\sigma_{uv}\beta>\sigma_{v}^2\beta^2$). In this case the limit of the conventional estimator is smaller than the asymptotic variance without any missing data problem, which is clearly counterintuitive. Moreover, the degree of endogeneity ($\sigma_{uv}$) erroneously affects the conventional limit while it has indeed no effect on the true asymptotic variance.

Monte Carlo Simulation

We illustrate our theoretical findings in Monte Carlo simulations. To implement the missing data problem, we first mimic the model in ((ref)) using the data generating process described below, and then randomly delete the value $x_i$ with probability $p$. The missing structure is thus MCAR. For each of the $n$ observations $Z_{i.}$ and $v_i$ are drawn from $$Z_{i.}\sim{}N\left(0_{L},\frac{1}{L}I_L\right) \ ; \hspace{0.5cm} v_i\sim N(0,1) \, , $$ where $L$ denotes the number of instruments. We follow hausman2012instrumental and chao2014testing and define the structural error as, $$u_i=\sigma_{uv}v_i+\sqrt{\frac{1-\sigma_{uv}^2}{\phi+(0.86^2)}}(\phi\epsilon_{1i}+0.86\epsilon_{2i}),\;\;\epsilon_{1i}\sim{}N(0,Z_{i.}'Z_{i.}),\;\epsilon_{2i}\sim{}N(0,0.86^2),\;\phi=5 \, ,$$ where $\phi$ defines the strength of the heteroskedasticity. We set $\beta=0.5$, $L=3$, $n=1000$, number of Monte Carlo repetitions $R=5000$, and the step size, which we use to alter the probability of missings in the different simulations, to $\Delta{}p=0.005$. As derived in the previous section, the sign of the endogeneity is crucial. Hence, we show results for $\sigma_{uv}=0.3$ and $\sigma_{uv}=-0.3$. We set $\pi=\sqrt{\frac{F\,L}{n}}$. Note that we divide by $n$ and not by $n_0$. Therefore, $F$ defines the first-stage $F$-statistic in the entire sample containing both complete and incomplete observations, and the first-stage $F$-statistic in the complete case sample gradually decreases with the missing probability. We set $F=100$. That is, the complete case $F$-statistic is around $100$ at $p=1$ and around $20$ at $p=0.8$.

figure[figure omitted — 207 chars of source]
figure[figure omitted — 220 chars of source]

Figure (ref) compares our consistent and the conventional standard errors to the observed root mean squared error (RMSE) obtained from the Monte Carlo simulations. It shows that the conventional estimator cannot properly describe the standard error of the 2SLS estimator with regression imputation. The missing probability has a linear effect on the conventional estimator (see eq. (ref)) while it nonlinearly affects the true asymptotic variance as shown in Corollary 1. Moreover, as expected the conventional standard error can even decrease as the missing probability increases ($\vert\sigma_{uv}\beta\vert>\sigma^2_v \beta^2$ in the right graph of Figure (ref)). The reason why MCDONOUGH2017, who did not derive a consistent variance estimator for regression imputation, have acceptable results in their simulations is solely due to their chosen parameters ($\sigma_{uv}>0$, $\beta>0$, low missing probability) and cannot be generalized to the entire parameter space. Our variance estimator ($\widehat{V}_{\widehat\beta_{RI}}$) performs well in both settings. Additionally, Figure (ref) shows how often the null hypothesis ($H_0:\,\beta=0.5$) is rejected at the $5\%$ nominal level. Again, our proposed variance estimator performs better than the conventional one---particularly, if $\sigma_{uv}<0$.

Conclusion

We investigate how two issues, which are likely present in many empirical studies, affect the estimation of causal effects, namely the endogeneity of regressors and missing data. If researchers use an instrumental variables regression after single imputation of missing values for an endogenous regressor, they have to be aware that conventional methods to estimate the variance fail to account for the imputation. The asymptotic variance of 2SLS is affected by the imputation implying that conventional methods cannot be used to construct standard errors, confidence intervals and conduct tests. We derive a heteroskedastic variance estimator which takes the imputation into account and is consistent. Monte Carlo simulations show that our estimator performs well while the conventional variance estimator does not.