Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,051 characters · 13 sections · 33 citation commands
15.819Identification of Average Responses with Endogenous Controls
\thispagestyle{empty}
\setcounter{page}{0} \thispagestyle{empty} \pagestyle{plain}
We study a critical and prevalent, yet often overlooked, issue in empirical research: control variables may be endogenous. In models with endogenous treatments, researchers often leverage a conditional independence assumption (CIA) or instrumental variables (IV) to identify treatment effects, while rather casually assuming the control variables are exogenous.
In applications, additional control variables are often included either because they are relevant for both the treatment variable and the outcome, or because the IV is more likely to be valid when controls are included. However, empirical researchers often end up with control variables that are subject to additional endogeneity concerns, while finding instruments for every endogenous control can be challenging or impossible.
In this paper, we demonstrate how endogeneity of controls affects the identification of the average response to the treatment and, when it does, how the issue can be addressed in certain scenarios. To fix the idea, suppose we are given a parametric model
where $h$ is known up to unknown parameters ${\Greekmath 0112}$. $W$ is the treatment or policy variable of interest and $X$ is a vector of control variables. We assume that $W$ is scalar for notational convenience, but the results extend to the vector case with additional notation. For now, suppose $Y$ and $W$ are continuous random variables, and $h(w,x,{\Greekmath 010F};{\Greekmath 0112})$ is differentiable with respect to $w$. Suppose a researcher is interested in the average response with respect to $W$, defined as
We consider the parametric model as correctly specified in the sense that there exists ${\Greekmath 0112}$ such that (ref) holds almost surely\footnote{${\Greekmath 0112}$ need not to be unique since we are only interested in identifying ${\Greekmath 010C}$.}. However, we remain agnostic about the dependence between the $X$ and ${\Greekmath 0122}$. The goal is to establish identification of ${\Greekmath 010C}$ under CIA
where $\widetilde{X}$ is a collection of control variables such that $X\subseteq\widetilde{X}$ and $\widetilde{X}$ may contain extra variables excludable from the outcome equation (ref).
As we illustrate in the examples below, CIA alone does not suffice for identification of ${\Greekmath 010C}$ without further restrictions, for example, when $\widetilde{X}$ is endogenous. Instead of imposing the exogeneity of $\widetilde{X}$, we establish identification through an intuitive rank condition associated with $W$ and $\widetilde{X}$, effectively allowing controls to be endogenous. \\
Notation. Throughout the paper, we let $F(A,B,C)$, $F(A,B)$, and $F(A)$ denote the distribution functions of $(A,B,C)$, $(A,B)$, and $A$; $F_{A|B,C}$ denotes the distribution functions of $A$ conditional on $B,C$. Whenever a (conditional) density is invoked, we assume it exists and is well-defined. We denote (conditional) density functions $f$ similarly. $A \perp \!\!\! \perp B \ |\ C$ denotes the independence of $A$ and $B$ conditional on $C$. The supports of the random elements $A$ are denoted by their calligraphic letters $\mathcal{A}$. We use $P$ to denote the probability measure and $P^*$ to denote the bootstrap probability measure conditional on the observed data.
To illustrate the problem, let's consider the following linear model:
In this model, the average response with respect to $W$ is given by ${\Greekmath 010C} = {\Greekmath 011C}$. While it is assumed that ${\Greekmath 0122}$ is mean independent of $W$ conditional on $X$, ${\Greekmath 0122}$ may not be mean independent of $X$. For example, if the functional form of $X$ is misspecified and some nonlinear effects of $X$ end up being part of ${\Greekmath 0122}$, then $E[{\Greekmath 0122}|X]\ne 0$. Generally, this is more than a specification issue, although, as we show below, in many cases this can be mitigated by a flexible functional form.
There are two consequences of this dependence. First, without considering $X$, it may cause omitted variable bias. Second, in the presence of $X$, the dependence between $X$ and ${\Greekmath 0122}$ can also pollute the identification of ${\Greekmath 011C}$ even if $W$ is conditionally independent of ${\Greekmath 0122}$. In either case, ${\Greekmath 011C}$ is not identified by the linear projection parameters, so the OLS estimator is inconsistent and biased. To see the bias in the second case, let $D = (W,X')'$ and ${\Greekmath 0112} = ({\Greekmath 011C},{\Greekmath 0118}')'$, the linear projection parameters are defined as follows:
Therefore, without further restriction on $E[{\Greekmath 0122}|X]$ such as $E[{\Greekmath 0122}|X]=0$, ${\Greekmath 0112}$, or ${\Greekmath 011C}$ in particular, are not identified and OLS would produce a biased and inconsistent estimator. In a worse scenario, $X$ could be an outcome of $W$, in which $X$ is referred to as “bad control" in angrist2009mostly. With the presence of bad control, CIA is not likely to hold (lechner2008note), and it is also noted that the bad control can cause problems for identification even if treatments are randomly assigned (wooldridge2005violating).
However, as shown below, ${\Greekmath 011C}$ can still be identified as long as $X$ is not solely a function of $W$. More formally, this extra condition is referred to as the measurable separability, as introduced in florens1990elements:
At its essence, this assumption ensures that we can vary the value of $W$ while holding $X$ at a particular value. Note that this still allows the distribution of $X$ to depend on $W$, and vice versa. Although it does not rule out the bad control directly\footnote{For example, $X = W+ e$ where $e\perp \!\!\! \perp W$, then we can vary $W$ while fixing $X$, but in terms of potential outcome $X(w)\ne X$.}, it rules out situations where CIA is not likely to hold.
The idea is as follows. Under the parametric specification assumption (ref), there exists a nonparametric and nonseparable nesting model $m(W,X,{\Greekmath 0122})$ such that
It follows that the average response defined from this nesting model is equivalent to the average response defined by $h$:
It is shown later that, under CI given $\widetilde{X}$ as well as the measurable separability between $W$ and $\widetilde{X}$,
In this linear case (ref), we can take $\widetilde{X} = X$, so
and it is straightforward to examine $E[\partial_W E(Y|W,X)]$ directly. Under the same CI and the measurable separability as above,
where the last equality holds due to the measurable separability between $W$ and $X$. To see why this condition is necessary, suppose the measurable separability does not hold, e.g., $X=f(W)$ almost surely and they are not constants, then conditioning on $W=w, X=x$ necessitates $W=w, X=f(w)$. In that case, the last equality does not hold anymore.
Note that the differentiability with respect to $W$ is not essential. Given the conditional mean independence and the measurable separability of $W$ and $X$, the identification result extends to nonparametric regression models such as
From the conditional mean function
$f(w)$ is identified due to the measurable separability, even if $h(x)$ is not identified (i.e. $E[{\Greekmath 0122}|X=x]\neq0$). This result does not require differentiability of $E[Y|W=w,X=x]$ or $f(w)$.
In the case of a binary $W$, measurable separability between $W$ and $X$ allows for conditioning on $W=1$ and $W=0$ at different values of $X$. Combining with CI, the identification of ${\Greekmath 011C}$ is achieved as follows:
When the instrument is genuinely exogenous and affects the outcome solely through the treatment, the answer is yes. However, in practice, control variables are often included to either increase precision or to make the IV assumptions more plausible. In these cases, the endogeneity of these controls can again jeopardize identification. To illustrate, let's consider the linear model with an excludable IV:
In this model, the average response with respect to $W$ is also ${\Greekmath 011C}$. However, differing from the previous example, $W$ is not conditionally independent of ${\Greekmath 0122}$ even after conditioning on $X$. Meanwhile, without controlling $X$, $Z$ may not be a valid IV if $Z$ affects $Y$ through $X$ too. Therefore, endogeneity in $X$ creates the same dilemma, and neither ${\Greekmath 011C}$ nor ${\Greekmath 0118}$ is identified by the usual IV or 2SLS projection.
Nevertheless, ${\Greekmath 011C}$ can be nonparametrically identified through a control function approach\footnote{ It is not the only way of identification. In this linear triangular model, ${\Greekmath 011C}$ is also identified by $E[\partial_Z E[Y|X,Z]] / E[\partial_Z E[W|X,Z]] = {\Greekmath 011C}$.}: If $Z$ is independent of $({\Greekmath 0122},{\Greekmath 0111})$ conditional on $X$, then $W$ is independent of ${\Greekmath 0122}$ conditional on $X$ and ${\Greekmath 0111}$. Now, by taking $\widetilde{X} = \{X,{\Greekmath 0111}\}$ and assuming $\widetilde{X}$ is measurably separated from $W$\footnote{Because $ W = Z{\Greekmath 0119}_Z + X{\Greekmath 0119}_X + {\Greekmath 0111}$ with ${\Greekmath 0119}_Z\ne 0$, the measurable separability between $Z$ and $X$ also implies the measurable separability between $(X,{\Greekmath 0111})$ and $W$ in this case.}, the general result (ref) implies
This can also be obtained by examining $E[\partial_W E(Y|W,X,{\Greekmath 0111})]$ directly:
As illustrated in the examples above, nonparametric methods appear to be more robust to endogenous controls. This is first argued in frolich2008parametric, but many questions remain: First, to what extent are the nonparametric methods immune to the endogenous control? For example, some extra regularity conditions may be needed, such as the aforementioned measurable separability. Furthermore, we ask how broad the class of models is in which nonparametric methods remain valid in the presence of endogenous controls. If we can define such an admissible class of models, then the next question is, should researchers always use nonparametric methods? It is well-known that nonparametric methods can be less efficient, as the cost of being robust to specification, and, in this case, to endogenous control. Ideally, we would like a testing procedure that detects potentially endogenous controls in a broad class of models.
In this paper, we study a large class of parametric models that are nested in a nonseparable and nonparametric model, with controls that are potentially endogenous. We show that under CIA and the measurable separability condition, the identification of the average response associated with the treatment or policy variable is available for the nesting model even with endogenous controls. Since the average response of the nested parametric models is equivalent to the average response of the nesting nonseparable and nonparametric model, researchers can instead resort to the latter one to avoid the bias caused by endogenous controls. Note that the CIA can be either achieved by conditioning on the observable control variables or through the control function approach when there exist excludable exogenous variables.
Based on these results, we further propose a test for endogenous controls. For linear models, our test reduces to a Hausman-style specification test. For more general setups, the test follows the same principle by checking agreement between a restricted estimate and a more robust estimate of the average response. Particularly, we focus on the average derivative estimator through the series approximation and inference based on bootstrap. We examine the size and the power of the test both theoretically and through simulation.
In the simulation study, we present some examples of data generating processes that feature endogenous controls. By contrasting our approach and methods that do not take into account the endogenous controls, we highlight the practical importance of this issue and document the finite-sample performance of our procedure. To further illustrate our methods and draw practical implications, we revisit a classic study of the import-competition impact on labor market outcomes by autor2013china. While adding control variables can alleviate the bias from the concern of IV validity, it may not address—and can even mask—bias arising from endogeneity in the controls themselves. Our approach and test together provide a simple robustness check for detecting this additional source of bias.
The issue of endogenous controls is prevalent in empirical research but is not well studied in the econometrics literature. One exception outside our setting is regarding the regression discontinuity (RD) design, where il2013regression finds that endogenous controls yield asymptotic bias in the RD estimator while the inclusion of these relevant controls may offset this bias and improve some higher-order properties of the estimator. diegert2022assessing assesses the omitted variable bias when the controls are potentially correlated with the omitted variables in a sensitivity analysis framework. In a recent paper by iv2025, the issue of endogenous controls is attributed to misspecification, and their method of “strong exclusion” amounts to projecting out from the instrument $Z$ a conditional mean function of $Z$ given $X$, which is conceptually and econometrically equivalent to including more flexible functional forms of $ X$ as the control function.
The rest of the paper is outlined as follows. The main identification results of the average response under CIA and the measurable separability are given in Section (ref). The endogenous control test is proposed in Section (ref). Section (ref) presents the simulation study that compares our methods and those not robust to endogenous controls, and examines the finite sample performance of endogenous control test. Section (ref) provides the empirical application to illustrate the practical implications of our methods. Section (ref) concludes the paper with empirical recommendations.
To investigate the impact of endogenous controls in a general setting, we consider a broad class of parametric models as (ref). Let $m(W,X,{\Greekmath 0122})$ be a nesting nonseparable and nonparametric model,
A similar nonparametric and nonseparable model has been studied by altonji2005cross except that they consider $X$ as the excluded instruments and $X$ do not enter the outcome equation.\footnote{Our main motivation for model (ref), which differs from altonji2005cross, is that empirical researchers often seek to include $X$ as control variables in the outcome equation.} For a continuous treatment $W$ and a continuous outcome $Y$, the average response ${\Greekmath 010C}$ associated with $W$ can be defined as
If $Y$ is a binary outcome, we can write
where $m^*$ is implicitly defined. Following altonji2005cross, we partition ${\Greekmath 0122}$ as $(u,v)$ and implicitly define $u^*= u^*(W,X,v) $ as a solution of $m^*(W,X,u^*,v) = 0$. Suppose that, for fixed $(w,x,v)$, $m^*(w,x,u,v)$ has at least one root in $u$ and that $m^*(w,x,u,v)$ is strictly monotonic in $u$, then $u^*(W,X,v)$ is uniquely defined. Additionally, suppose $m^*(w,x,u,v)$ is continuously differentiable in $w$ and $u$ and that $\partial_{u}m^*(w,x,u^*,v)\ne 0$, then by implicit function theorem $u^*(w,x,v)$ is differentiable in $w$. Then, we can define the average response ${\Greekmath 010C}^*$ associated with a binary outcome as
For binary $W$, we can define a set of parameters analogously by replacing the derivatives above with the differences. For our main results, we focus on the continuous treatment case to simplify the exposition.
Let $\widetilde{X}$ denote a collection of control variables such that $X\subseteq \widetilde{X}$. We consider the scenario where ${\Greekmath 0122}$ is conditionally independent of $W$ given $\widetilde{X}$, but the dependence between $X$ and ${\Greekmath 0122}$ is unrestricted.
Assumption (ref) is a key condition for identifying the average response. Importantly, it does not rule out $X\subseteq \widetilde{X}$, or $X$ in particular, being endogenous. We emphasize that Assumption (ref) itself may not be sufficient for identification due to the endogeneity of $X$: for example, if $W$ is some function of $X$ only, which is dependent on ${\Greekmath 0122}$, then the average response would not be identified by Assumption (ref), and, in which case, even Assumption (ref) itself is not likely to hold. Thus, an additional restriction is needed to rule out such extreme cases. The next theorem shows that the measurable separability condition between $W$ and $\widetilde{X}$ suffices for the identification of the average response.
A proof is provided in Appendix. Theorem (ref) shows that even if ${\Greekmath 0122}$ and $X$ are potentially dependent, the average response associated with the treatment can still be identified as long as CIA and the measurable separability condition hold. By definition, measurable separability simply requires that $W$ and $X$ are not exclusively determined by each other, which is a very mild rank restriction. For further discussion, readers are referred to florens2008identification, where they also provide primitive conditions on the data generating process under which measurable separability between two random variables is guaranteed.
Theorem (ref) is a positive result: since the measurable separability condition is a very mild rank condition that holds in most settings, it justifies the prevalent use of potentially endogenous controls. Meanwhile, we have seen in previous examples that parametric models are not immune to the endogeneity of control, so this result encourages the use of fully flexible models in these scenarios. \\
Example 3 continued. For the binary model $Y = 1\{ {\Greekmath 011C} W + X'{\Greekmath 0118} + {\Greekmath 0122}>0 \} $, suppose ${\Greekmath 0122}|W,X \sim N({\Greekmath 0116}(X),{\Greekmath 011B}(X)^2)$, i.e. CIA holds for this model with $\widetilde{X} = X$. Following the calculation as (ref), the true average response is given by
By Theorem (ref), ${\Greekmath 010C}^*$ is identified as ${\Greekmath 010C}^*=E[\partial_W P(Y=1|W,X)]$ as long as $W$ is measurably separated from $X$. Indeed, $E[\partial_W P(Y=1|W,X)]= E\left[\frac{{\Greekmath 011C}}{{\Greekmath 011B}(X)} {\Greekmath 011E}\left( \frac{{\Greekmath 011C} W + X'{\Greekmath 0118}- {\Greekmath 0116}(X)}{{\Greekmath 011B}(X)} \right)\right]$ under CIA. However, ignoring the endogeneity of the controls with standard probit approach would impose $P(Y=1|W,X) = \Phi({\Greekmath 011C} W + X'{\Greekmath 0118})$ in constructing the likelihood function, which would fail to identify the true average response.
When the CIA is not plausible given observable controls $X$, it is common to consider exogenous and excludable IVs for identification. In applications, additional control variables are often included to make the exogeneity condition of IVs more likely to hold. Although control variables are explicitly or implicitly assumed to be exogenous, we caution that they may be endogenous in practice, while finding IVs for all endogenous controls is not possible. In this section, by utilizing excludable exogenous variables, we construct an extra control variable $V$ following the approach by imbens2009identification, such that CIA is satisfied given $\widetilde{X} = \{X,V\}$. Thus, combining with the results in the previous section, the average response to the treatment can be identified using IVs while explicitly allowing for endogenous controls.
The researcher starts from the parametric model (ref), nested in (ref), and the average response parameter $ {\Greekmath 010C}$ is defined the same way, except now the CIA is not available. Instead, suppose there exist exogenous and excludable variables IVs $Z$ such that the reduced form equation for $W$ is given by
where $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{dim}({\Greekmath 0111}) = \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{dim}(W) = 1$.\footnote{As discussed in imbens2007nonadditive, the same dimension of the treatment and the endogenous error term is essential for recovering all endogeneity in ${\Greekmath 0111}$ through the inverting procedure.} The exogeneity of the IV is represented by the conditional independence in the following assumption.
Note that we don't impose the exogeneity condition of $X$ with respect to either ${\Greekmath 0122}$ or ${\Greekmath 0111}$, yet $X$ is allowed in both outcome and the reduced form equations in a nonseparable way. The requirement for IV is relaxed, and $Z$ may well be dependent on $X$, which motivates the inclusion of $X$ in the model.
From a control-function perspective, the goal is to find a control variable $V$ such that
and $\widetilde{X}$ is measurably separated from $W$. Then, we can identify the average response by Theorem (ref). In the next theorem, we show that under Assumption (ref), ${\Greekmath 0111}$ can serve such a purpose; and so is $F_{\Greekmath 0111}({\Greekmath 0111})$ given that the CDF of ${\Greekmath 0111}$ is continuous and strictly increasing. If $q(Z,X,{\Greekmath 0111})$ is strictly monotonic in ${\Greekmath 0111}$ almost surely, then we might recover $F_{{\Greekmath 0111}}({\Greekmath 0111})$ from $F_{W|Z,X}(W)$, but this is possible only when both $Z$ and $X$ are independent of ${\Greekmath 0111}$ (see e.g. Matzkin2003nonparametric,imbens2009identification), which is not the case here since $Z$ is only conditionally independent and $X$ is allowed to depend on ${\Greekmath 0111}$. Nevertheless, it turns out $V=F_{W|Z,X}(W)$ is still a valid choice, for the identification of the average response, as long as the sigma algebra generated by $(X,{\Greekmath 0111})$ is the same as that by $(X,V)$, which is indeed the case as we show in the next theorem.
A proof can be found in Appendix.
Assumption (ref) (i) ensures the invertibility of $q(Z,X,{\Greekmath 0111})$ in ${\Greekmath 0111}$, which helps recover $F_{{\Greekmath 0111}|X}({\Greekmath 0111})$ through $F_{W|Z,X}(W)$. Assumption (ref) (ii) further guarantees the recovered $F_{{\Greekmath 0111}|X}({\Greekmath 0111})$ captures all endogeneity in ${\Greekmath 0111}$, once conditional on $X$. By Theorem (ref), CIA is satisfied with $\widetilde{X} = \{X,V\}$. Therefore, the identification of the average response associated with the treatment can be obtained by Theorem (ref).\footnote{Although $V$ does not enter the outcome equation $m(W,X,{\Greekmath 0122})$, we can always rewrite $\tilde{m}(W,X,V,{\Greekmath 0122})= m(W,X,{\Greekmath 0122})$ and treat $\tilde{m}$ as the outcome equation in Theorem (ref).}
In this section, we consider a test for endogenous controls in a large class of models, as an implication of the results presented in Section (ref). Consider the parametric model $Y = h(W,X,{\Greekmath 0122};{\Greekmath 0112})$, nested by the nonparametric and nonseparable model $Y = m(W,X,{\Greekmath 0122})$. To focus on the main idea, we limit our attention to the case where $\widetilde{X}=X$ are observable, and we observe an i.i.d sample of size $n$.\footnote{The analysis can be extended to the case of Section (ref) where the conditioning variable $V$ needs to be estimated, but that would complicate the asymptotic analysis and deviate from the main idea of the test.}.
Regardless of the continuous or binary outcomes, the average response is universally defined for parametric models and is equivalent to the average response defined using the nesting nonparametric model. In the linear model (ref), ${\Greekmath 010C} = {\Greekmath 011C}$. In the binary model of Example (ref), ${\Greekmath 010C}^* = E\left[\frac{{\Greekmath 011C}}{{\Greekmath 011B}(W,X)} {\Greekmath 011E}\left( \frac{{\Greekmath 011C} W + X'{\Greekmath 0118}- {\Greekmath 0116}(W,X)}{{\Greekmath 011B}(W,X)} \right)\right]$. Under CIA and the measurable separability, both are identified by Theorem (ref). However, neither is parametrically identified, and so the corresponding parametric estimators assuming exogeneity of the controls would be inconsistent. Meanwhile, when these models are correctly specified without endogenous controls, the parametric estimators corresponding to these specifications can be efficient (attaining the Cramer-Rao lower bound) under certain conditions. This observation naturally leads to a Hausman-type test on the endogenous control.
Formally, we consider the null and the alternative hypotheses as follows:
The null and alternative hypotheses are specified loosely to accommodate a broad class of models. The meaning of the null and the alternative varies over the specification of $ h(W,X,{\Greekmath 0122};{\Greekmath 0112})$. For example, in the linear model, $E[X{\Greekmath 0122}] = 0$ is necessary and sufficient for the null while $X\perp \!\!\! \perp {\Greekmath 0122}$ is not necessary. However, in the endogenous probit model of Example (ref), $X\perp \!\!\! \perp {\Greekmath 0122}$ is necessary and sufficient for the null while $E[X{\Greekmath 0122}] = 0$ is not sufficient. We will make the hypotheses concrete in a moment when we introduce the assumptions for the estimators under consideration.
To represent both binary and continuous outcomes, we denote the average response as ${\Greekmath 010C}_0:=E[\partial_Wh(W,X,{\Greekmath 0122};{\Greekmath 0112})]$ for both cases for the remainder of this section. Let $\hat{{\Greekmath 010C}}_p$ denote the parametric estimator under model (ref), but treating $X$ as exogenous, i.e. imposing $H_0$. Due to the identification results from last section, we consider an average derivative estimator $\hat{{\Greekmath 010C}}_{np}$: let $g(W,X) := E[Y|W,X]$, and
where $\hat{g}$ is a nonparametric estimator for $g$.
We define the limit of these two estimators as ${\Greekmath 010C}_p$ and ${\Greekmath 010C}_{np}$, respectively. With an appropriate choice of $\hat{g}$ and sufficient regularity conditions, the average derivative estimator is consistent for ${\Greekmath 010C}_{np} = E[\partial_Wg(W,X)]$ (e.g., see Example 3 of newey1994asymptotic). Under the conditions of Theorem 1, ${\Greekmath 010C}_{np} = {\Greekmath 010C}_0$, regardless of $H_0$ or $H_1$. Meanwhile, under the same conditions, the limit ${\Greekmath 010C}_p$ may change across data generating processes because the parametric model is correctly specified up to the exogeneity restrictions on the controls.
If, under $H_0$, ${\Greekmath 010C}_{p} = {\Greekmath 010C}_0$, and both estimators, $\hat{{\Greekmath 010C}}_p$ and $\hat{{\Greekmath 010C}}_{np}$, are asymptotically linear with some influence functions ${\Greekmath 0127}$ and ${\Greekmath 011E}$ respectively, then we can define ${\Greekmath 0120} = {\Greekmath 0127} - {\Greekmath 011E}$, and write
which, under common regularity conditions, is asymptotically normal with the asymptotic variance $ V = \lim_{n\to\infty} Var\left(\frac{1}{\sqrt{n}}\sum_{i=1}^n {\Greekmath 0120}(W_i,X_i)\right)$. When the parametric estimator $\hat{\Greekmath 010C}_p$ is efficient under the null, we can appeal to the classic result of hausman1978specification to express the $V$ as the difference of the asymptotic variances of $\hat{{\Greekmath 010C}}_{np}$ and $\hat{{\Greekmath 010C}}_p$. In general, however, assuming efficiency of the parametric estimator is not desirable. Moreover, since ${\Greekmath 0127}(W,X)$ depends on specific parametric estimator, and ${\Greekmath 011E}(W,X)$ may depend on extra infinite-dimensional nuisance parameter estimation, the analytical expression for $V$ is not straightforward. Alternatively, we can resort to a bootstrap approach, which is a common and general approach for inference that involves semiparametric estimation; see, for example, chen2003estimation.
Therefore, with ${\Greekmath 0112}_0 = {\Greekmath 010C}_p - {\Greekmath 010C}_{np}$ and $\hat{\Greekmath 0112} = \hat{{\Greekmath 010C}}_{p} - \hat{\Greekmath 010C}_{np}$, we consider a bootstrap procedure based on the non-studentized statistic $ \sqrt{n}\left(\hat{{\Greekmath 0112}}-{{\Greekmath 0112}}_0\right)$. Let $J_n^*$ denote the empirical bootstrap CDF of $\sqrt{n}(\hat{\Greekmath 0112}^*-\hat {\Greekmath 0112})$ where $\hat{\Greekmath 0112}^*$ is the bootstrap estimator for $\hat{\Greekmath 0112}$. We consider the bootstrap confidence region $$ \mathcal{B}_n({\Greekmath 010B}) = \{{\Greekmath 0112}\in \Theta: \hat{\Greekmath 0112} - n^{-1/2} \inf \{t: J_n^*(t)\geq 1- {\Greekmath 010B}/2 \}\leq {\Greekmath 0112} \leq \hat{\Greekmath 0112} - n^{-1/2}\inf \{t: J_n^*(t)\geq {\Greekmath 010B}/2 \}\}.$$ A brief algorithm for implementation of the test is given as follows.
We introduce a set of high-level conditions in terms of $\hat{{\Greekmath 010C}}_p$, $\hat{{\Greekmath 010C}}_{np}$, as well as their limits. Let $P^*$ denote the bootstrap probability measure conditional on the sample $\{Y_i,W_i,X_i\}_{i=1}^n$.
Assumption (ref) (i) operationalizes the null and alternative hypotheses. For example, under the linear model (ref), ${\Greekmath 010C}_0 = {\Greekmath 011C}$ and ${\Greekmath 010C}_p$ is the first element of the linear projection parameter ${\Greekmath 010D}_{LP}$, thus $B_n$ is the first element of $E[DD']^{-1}E[DE[{\Greekmath 0122}|X]]$ where $D = (W,X')'$. The upper bound in Assumption (ref)(iv) is a common finite variance condition for applying the central limit theorem. The lower bound excludes the case where the parametric estimator and the average derivative estimator agrees.
Assumption (ref) (ii) and (iii) impose asymptotic linear representations for the estimators and their bootstrap counterparts. For parametric estimators, asymptotic linear representation is very common. For example, it holds for M-estimators with twice continuously differentiable objective functions. For some estimators with non-smooth objective functions, e.g. quantile estimators, there are also existing results established under extra regularity conditions. The asymptotic linear representation for the average derivative estimator follows from the general results for semiparametric estimators. newey1994asymptotic provides a general approach for deriving the influence function for semiparametric estimators, and Example 3 of the same paper exemplifies the average derivative estimator using series approximation for the conditional expectation. See the next section for a concrete implementation of this estimator for simulation.
The asymptotic linear representations for bootstrap estimators expect slightly stronger regularity conditions. To see how this representation can be justified, we consider the process of the recentered scores $S_n(g):= \frac{1}{\sqrt{n}}\sum_{i=1}^n \left( \partial_W g(W_i,X_i) - E[ \partial_Wg(W_i,X_i)] \right)$ and the bootstrap process $S_n^*(g) = \frac{1}{\sqrt{n}}\sum_{i=1}^n \left( \partial_W g(W_i^*,X_i^*) - \partial_Wg(W_i,X_i)\right)$. Also, define the pathwise derivative of $S_n(g)$ in the direction of $\tilde g - g$, evaluated at $g$, as $\Gamma(g)[\tilde g - g]$. In a semiparametric GMM framework, Theorem B of chen2003estimation shows that under a slightly strengthen set of sufficient conditions for asymptotic normality of the GMM estimator, (ref) can be obtained by
Condition (ref) is also assumed in chen2003estimation and holds for, for example, series estimators(NEWEY1997147,chen2015optimal) and kernel regression estimators(hall1991convergence). For Condition (ref), we further denote $M_n(g) = \frac{1}{\sqrt{n}}\sum_{i=1}^n {g}(W_i,X_i)$ and $M_n(g) ^* =\frac{1}{\sqrt{n}}\sum_{i=1}^n {g}(W_i^*,X_i^*)$. By adding and subtracting $M_n^*(\hat g^*) - M^*_n(\hat g)$,
Under regularity conditions for $ \Gamma(\hat g)[\hat g^* - \hat g]$ on $P^*$, we can apply results of newey1994asymptotic (Lemma 5.1) to obtain (ref).
The following result shows that this bootstrap confidence region is asymptotically valid.
A proof can be found in Appendix. Theorem (ref) formalizes the Hausman-type principle for the endogenous control test. Unlike a specification test, the rejection here can occur even when the parametric model is correctly specified.
In this section, we use Monte Carlo simulations to demonstrate the performance of our proposed estimators and tests in finite samples: (i) finite-sample bias from endogenous controls, (ii) the robustness of our approach to such endogeneity, and (iii) the coverage and power properties of the proposed test. We consider two data generating processes (DGPs) that correspond to the two scenarios covered in Sections (ref) and (ref).
For concreteness, we consider an implementation using series estimator to estimate conditional mean functions and their derivatives, based on our constructive identification results. Suppose $(W,\widetilde{X})$ is $d$-dimensional. For some given $K$, we define a power series $\{\tilde{g}_k\}_{1\le k\le K}$ to approximate $g(W,\widetilde{X})=E[Y|W,\widetilde{X}]$:
where $G = (\tilde{g}_1',...,\tilde{g}'_K)'$; for each $k$, $$\tilde{g}_k= \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{vec}\left\{\prod_{j=1}^d q_j^{{\Greekmath 0115}_j}: \sum_{j=1}^d{\Greekmath 0115}_j = k, (q_1,...,q_d) = \left(l_1(W),l_2(\widetilde{X}_1) ,...,l_d(\widetilde{X}_{d-1})\right)\right\},$$ and $l=(l_1,...,l_d)$ is some transformation function differentiable to all orders with bounded derivatives and has $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{det}(\partial_z l(z))$ bounded away from zero. $\Pi = ({\Greekmath 0119}_1',...,{\Greekmath 0119}_K)'$ are the corresponding linear projection parameters. The associated average derivative estimator is defined as $$\hat{{\Greekmath 010C}}_{np}:= \frac{1}{n}\sum_{i=1}^n \partial_W \left[G(W_i,\widetilde{X}_i)'\hat\Pi\right],$$ where $\hat\Pi$ are the (penalized) least square estimators of $\Pi$. For our implementation in this section, we fix $K=3$ and set $l(x)=x$.
The first DGP covers the scenario where both the treatment and the controls are endogenous, while the treatment is conditionally independent of the unobserved determinants:
where ${\Greekmath 011C}= 1$; $(a,b)$ and $(p,q)$ are independent of each other, and each are jointly normal with mean zero, variance one, and covariance ${\Greekmath 011A}$; $N(0,1)$ denotes a random draw from a standard normal distribution, independent of $(a,b)$ and $(p,q)$. We observe that (1) $X = (X_1,X_2)'$ are relevant for both $Y$ and $W$; (2) $W$ and $X$ are dependent on $U$; (3) Conditional on $X$, $W$ is independent of $U$; and (4) $W$ and $X$ are measurably separated. As a result, linear projection parameters do not identify ${\Greekmath 011C}$. Meanwhile, ${\Greekmath 011C}$ is also the average response to $W$, and it is nonparametrically identified by Theorem (ref).
The second DGP covers the scenario where the IV is only conditionally valid, and the control is endogenous:
where $({\Greekmath 0118},{\Greekmath 0110})$ are also jointly normal with mean zero, variance one, and covariance ${\Greekmath 011A}$. We observe that (i) $W$ is not conditionally independent given $X$; (ii) $Z$ is a valid $IV$ only when conditional on $X$, but $X$ is endogenous; (iii) $Z$ and $X$ are measurably separated. (iv) The measurable separability between $Z$ and $X$ here also implies the measurable separability between $W$ and $(X,{\Greekmath 0111})$. As a result, ${\Greekmath 011C}$ is not identified by the usual IV projection, while ${\Greekmath 011C}$ is again, as an average response, nonparametrically identified: constructive identification is given by (ref) or footnote (ref) when the outcome equation is linear, and Theorem (ref) gives a more general approach when the outcome equation is nonparametric and nonseparable.
For DGP (1), Table (ref) compares the estimates of ${\Greekmath 011C}$ using (i) OLS without control, (ii) OLS with control, and (iii) $\hat{{\Greekmath 010C}}_{np}$ through the third-order polynomial series regression with $\widetilde{X} = X$. We compare these three approaches across three endogeneity levels denoted by ${\Greekmath 011A}$, with ${\Greekmath 011A} = 0$ corresponding to the exogenous control case and ${\Greekmath 011A} >0$ for endogenous controls. The results are clear and align with theoretical predictions. Conventional methods that assume the exogeneity of controls are severely biased unless the controls are indeed exogenous (${\Greekmath 011A}=0$) and the size of the bias depends on the strength of the endogeneity (${\Greekmath 011A}\neq0$), while nonparametric methods are robust to the endogeneity of controls and perform much better in terms of bias.
Table (ref) reports simulation results under DGP (2). The comparison is among the IV estimator without control, the IV estimator with control, nonparametric approach due to footnote (ref) using the series approach, control function approach due to Theorem (ref) with $\widetilde{X} = {X,V}$, where the conditional CDF $V$ is estimated by the series approach, too, with 10 quantiles grids between 0.05 and 0.95. The results also support our theories: The IV-based methods that implicitly impose the exogeneity of the controls fail to produce consistent estimates of the average response unless the exogeneity is indeed true, and the proposed approaches are robust to endogenous controls with much less bias.
In this section, we will focus on DGP(1) and demonstrate the finite sample performance of the proposed test in terms of empirical coverages and statistical powers. The implementation of the test is based on Algorithm 1 from last section.
In the setting of DGP (1), the covariance parameter ${\Greekmath 011A}$ controls the degree of endogeneity in controls. As a result, the null and alternative hypotheses of the test can be translated as
Therefore, we can simulate the coverage probability under the null ${\Greekmath 011A} = 0$, and we can choose a grid for ${\Greekmath 011A}>0$ to simulate the power of the test.
Table (ref) reports the empirical coverage under different sample sizes and number of Bootstrap replications. We find that the coverage gets closer to the nominal rate as the sample size increases, and the test tends to be conservative when the sample size is small.
Figure 1 displays the empirical powers of the test along a sequence of alternatives characterized by ${\Greekmath 011A}$. Overall, the pattern is as expected by the theory and the finite sample performance in terms of power is as desired. For small values of ${\Greekmath 011A}$ and when the sample size is small, the test is slightly conservative.
In this section, we revisit an empirical study of the import competition effect on the US labor market. During the late 20th and early 21st century, the world has witnessed a drastic surge of imports from the Chinese market. Meanwhile, there was a downturn of import-competing manufacturing in certain regions of the US, and it came with a higher unemployment rate and wage inequality in these regions. Therefore, a natural question arises: was the disruption of the local labor market mainly caused by the import competition, or was it rather a result of an overall economic transition in the US?
In a classic paper by autor2013china, the authors answer this question by exploiting the import-competition exposure variation across regional markets in the US. The idea is that regions with more initial specialization in labor-intensive industries are more exposed to the Chinese import competition. They measure the share of those industries at the start of the observation period and multiply it by the growth in US imports from China (shift), which produces a measure of import competition varied by regions. By taking commuting zones as analysis units, they are able to estimate effects on various labor market outcomes with a reasonable sample size.
There are two potential sources of endogeneity in this setup. Firstly, the import growth from China could be driven by unobserved local market transitions, such as industrial reallocation that causes a supply shortage. In that case, the growth in US imports from China could be driven by unobserved shocks that also move local outcomes. Secondly, the start-of-period measure of share in labor-intensive industries may also be related to other regional characteristics, such as educational attainment, which also determines the local labor market outcomes. To deal with the first type of endogeneity, they employ a measure of import growth in other high-income countries as an exogenous shift and multiply it by the same share variable to obtain the IV. They further assume that the growth in imports from China is not driven by demand shocks that are shared by the US and other high-income countries, so as to ensure the validity of the IV. For the second type of endogeneity, they take two approaches: (1) They first-difference both the treatment and the outcomes so that they are in terms of per capita changes. In a way, this is similar to removing the fixed effects in a linear panel model. (2) They add other start-of-period measures of regional market conditions and demographics to make the exogeneity of the share variables more plausible.
The results in Table 3 of autor2013china show that adding more controls reduces the size of the estimated (negative) impact due to Chinese imports. This can be regarded as correcting the bias due to the second type of endogeneity. However, adding more controls may not fully capture the unobserved determinants that jeopardize the exogeneity of the share variables, and doing so may introduce extra endogeneity from the control variables. To make the exogeneity condition of the IV more likely to hold while avoiding the extra bias due to endogenous controls, we propose to implement the nonparametric control function approach in Section (ref). We also illustrate how the endogenous control test works and its practical implications.
Our approach is also based on the exogeneity of the IV shift variables\footnote{Suppose the growth in imports from China is partially driven by positive demand shocks too, then the estimates of import-competition impact would look smaller because the positive demand would also generate positive outcomes in the US markets.}. Thus, we do not tend to evaluate or improve upon the choice of IV in the original study; instead, we intend to make this approach more robust by allowing more controls to achieve conditional validity of IV while relaxing the exogeneity requirement of the control variables.
To focus on our main concern, readers are referred to autor2013china for a detailed construction of the data. The main outcome of interest, $\Delta L$, is the regional decade-change (% pts) in the share of manufacturing employment; the explanatory variable of focus, $\Delta I\_US$, is a measure of changes in Chinese import exposure per worker in each region; and the IV, $I\_O$, is the same measure of changes in exposure but with Chinese imports to the US replaced by those to other high-income countries. Two periods of decade-long first-difference cross-sectional data are stacked together as a panel. Five sets of time-invariant control variables, $X$, are augmented. The baseline model is given as follows:
where ${\Greekmath 010D}_t$ is a time fixed effect and $e_{it}$ is the stochastic error term. Under the IV conditions and regularity conditions, ${\Greekmath 011C}$ is exactly identified by the linear IV approach, given that $X_i$ is uncorrelated to $e_{it}$.
To implement our approach as in Section (ref), we use the third-order Hermite polynomial series for approximating the nonparametric functions. Specifically, we estimate $F_{\Delta I\_US|\Delta I\_O,X}$ at a grid of values, taken from the unconditional quantiles of $\Delta I\_US$, using the third-order Hermite polynomial basis functions of $(\Delta I\_O,X)$. Due to a large number of basis functions as the dimension of control increases, we add a ridge penalty to the least square estimation of the finite series regression, which is also a common practice in nonparametric series IV for better computation properties as suggested by newey2003instrumental. The time fixed effects and dummy variables in $X$ are included linearly and not penalized. Using the ridge estimates, we obtain the fitted conditional CDF $\widehat{V}$ for each observation using linear interpolation. The conditional expectation $E(Y|W,V)$ is also approximated by the third-order Hermite polynomials with plugged-in $\widehat{V}$, and then the average derivative estimator is obtained from the least square estimator of the series regression with ridge penalty. Both ridge penalty levels are chosen by cross-validation. For inference, we obtain confidence intervals through cluster bootstrap that resamples with replacement at the state level.
Table (ref) displays our results. The first IV section of the table replicates the results from Table 3 in autor2013china, except that the standard errors and confidence intervals are replaced by the bootstrap versions to match those used for other estimators. ADE denotes the average derivative estimator with plugged-in $\widehat{V}$. The last section IV -- ADE gives the difference of these two estimators and is used for the test of endogenous controls. Six specifications in terms of controls are as follows: (1) No control; (2) start-of-period employment; (3) Census-division specific fixed effects + (1); (4) start-of-period shares of college attainment, immigrants, female employment + (2); (5) Routine-intensive occupations, offshorability index + (2); (6) All controls mentioned above.
First, looking at the IV estimates, we find that as more controls are included, the estimated effects get smaller: the magnitude drops from 0.746 (no control), to 0.610 (one control), 0.538 (census division dummies), 0.508 (demographics), 0.562 (industrial types), and 0.596 (all controls), all with a negative sign. The estimates of our approach display a similar pattern: the absolute sizes of the estimates reduce as more controls are included. As we discussed above, while the inclusion of more controls may solve the first type of bias, it may further introduce the second type, and our results provide alternative estimates robust to endogenous controls. If we believe more control variables help guard against the first type of bias, then the results in (6) suggest that after correcting the bias from the second type, the causal impact of Chinese import competition on the US market is not as large as what's found in the baseline results.
The standard errors of the nonparametric method are reasonably small, and some are even smaller than the linear IV estimates, which is attributed to the ridge regularization. However, the ridge penalty also introduces shrinkage bias on the slope estimators of the series regression, which in turn may cause shrinkage bias in the average derivative estimator. Therefore, we further conduct a robustness check by only penalizing the conditional CDF estimation and leaving the second-step average derivative estimator unpenalized. The results are displayed in Table (ref). We find that the estimates without penalization are indeed larger than their counterparts in Table (ref), but the pattern remains: although the standard errors increase, we still find our approach generates smaller estimates of the average response except for the first case, where no controls are included.
Lastly, the endogenous control test is conducted through estimates $\hat{\Greekmath 010C}_{p} -\hat{\Greekmath 010C}_{np} $ and the bootstrap 95% confidence interval from the last sections of Tables (ref) and (ref). Since row (1) does not include a control variable, our test does not apply there. Examining rows (2) - (6) in Table (ref), we find rejections of no endogenous control in the last three rows, suggesting the ADEs are preferable in specifications (4) - (6), where the 95% confidence interval excludes zero. Due to larger standard errors, the test results in Table (ref) are not very informative, given that none of the tests are rejected.
We address a critical, prevalent, yet often overlooked problem in empirical research: the endogeneity of control variables. Building on the insightful observation and discussion in frolich2008parametric that nonparametric estimation can help with the endogenous control problem, we provide constructive identification results for marginal effects in a simple linear model with or without IVs, and extend the results to a general class of nonseparable models.
Our results not only provide solutions for identifying marginal effects of the treatment in the presence of endogenous controls but also have important implications. Because the additional measurable separability condition we introduce imposes minimal practical restrictions, endogenous controls are generally innocuous in a nonparametric model. Furthermore, based on the identification results, we propose a test for endogenous controls.
For empirical studies, our results invite researchers to conduct robustness checks using nonparametric methods and the proposed test. In general, nonparametric approaches are more robust, not only with respect to specification concerns but also in light of endogenous controls.