Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
121,609 characters · 18 sections · 86 citation commands
Proxy Controls and Panel Data
A sizable portion of the empirical economist's working life is dedicated to diagnosing and accounting for the presence of confounding factors. Ideally, a researcher would control for these factors, but some important confounders may not be measured in the data. We suppose that the data contain some observables that can act as proxies (broadly defined) for these latent factors.
We refer to observables that provide noisy and possibly biased signals for unobserved confounders as `proxy controls' because they can act as proxies for factors one would like to control for. Test scores may be informative proxies for academic ability and years of experience for human capital. With panel data, past observations can provide a wealth of information about the latent characteristics of an individual. For example, past consumption habits are likely informative about the individual's consumption preferences.
The problem of mismeasured controls has long been acknowledged in the labor economics literature, particularly in the context of returns to schooling (see Griliches1977 for an early empirical example). Classical analyses assume additive linear specifications for the potential outcomes and the measurement error. Linearity may be implausible in some settings and precludes the study of nonlinearities and heterogeneity in treatment effects. Following the seminal work of Miao2018, an emerging literature considers proxy controls in non-linear, semiparametric, and nonparametric settings.
We develop new nonparametric identification results in the context of proxy controls. We identify the conditional (on observed treatments and possibly other covariates) average potential outcomes. Causal objects can be identified using either of two conditional moment restrictions. Using these dual characterizations, we show that estimation of causal effects in this setting is `well-posed' under our identifying assumptions. Well-posedness is crucial for deriving simple and transparent convergence rates for estimation methods based on our identification results. We show that the problem of proxy controls can be adapted to causal analysis using panel data and develop new nonparametric identification results for panel models with a fixed number of time periods.
In panels, observations from other time periods can be informative proxies for latent confounding factors. By definition, confounders are associated with treatments and potential outcomes, and so, if the confounding factors are persistent, past treatments and past outcomes are informative about the confounding factors. We provide conditions on the serial dependence structure of the data and latent variables so that one can form proxy controls from past observations that satisfy our identifying assumptions.
We suggest estimation and inference procedures based on our identification results. Our methods are fully non-parametric, they allow for continuous treatments and identification of causal effects conditional on continuous observable covariates. The procedures can be applied in both cross-sectional and panel settings. The estimation method is based on series regression, and therefore it also suggests a flexible parametric method if the number of series terms is simply held fixed rather than allowed to grow with the sample size.
We establish consistency and a convergence rate for our estimator under our identifying assumptions and primitive conditions of the kind employed in the literature on standard nonparametric regression. We give conditions under which our estimator can be asymptotically approximated by a Gaussian process. We develop a method for constructing uniform confidence bands that is based on the multiplier bootstrap and a related specification test.
To demonstrate the usefulness of our methodology we apply it to two very different real-world data problems. We use data from the Panel Survey of Income Dynamics (PSID) to estimate a structural Engel curve for food. We revisit the empirical setting of Fruehwirth2016 who use data from the Early Childhood Longitudinal Study of Kindergartners (ECLS-K) to estimate the causal impact of grade retention on the performance of US students in cognitive tests.
This paper contributes to an expanding body of recent research on the use of proxy controls. Miao2018 first established nonparametric identification of average potential outcomes and the marginal distribution of potential outcomes in cross-sectional settings when controls are mismeasured.
Miao2018 follows earlier work that detects and sometimes adjusts for the presence of measurement error or confounding using `negative controls': variables which lack a direct causal relationship to either the treatments or to outcomes (for a review see Shi2020a). The conditions Miao2018 and others place on proxy controls means that they are in fact, negative controls. Some of this earlier work utilizes serial dependence in the data to select valid negative controls, similar to our use of serial dependence restrictions to motivate choices of proxy controls (for example, Flanders2009).
We extend the identification results of Miao2018 to identify conditional average potential outcomes and we provide a dual identification result based on a different conditional moment restriction. This dual result plays a crucial role in our analysis of estimation using proxy controls. Our identifying assumptions also differ from Miao2018 in that we do not require one vector of proxies to be statistically complete for another. Instead, we use a condition of the type employed by Shi2020 in the context of categorical confounding and proxies.
Recent work adapts the proximal approach to a broad range of causal inference problems. For example, mediation analysis Dukes, synthetic control Qiu2022, and complex longditudinal studies Ying2021. The latter of these uses proxy controls to estimate causal effects in dynamic panel models as in this work, but has somewhat different aims. We use the panel structure to motivate choices of valid proxies formed from past treatments and/or outcomes. By contrast Ying2021 uses time-varying proxies to identify rich dynamic causal effects.
A number of papers that consider estimation using proxy controls utilize parametric specifications (for example, Tchetgen and Miao2018a). Much of the work in this literature considers semiparametric estimation, for example Cui2020 who derive semiparametric efficiency bounds. A flurry of recent work apply a kernel minimax estimation approach to this semiparametric problem, including Ghassami2021, Kallus2021, and Bennett2022. One innovation in these recent works is the use of Neyman-orthogonal learning, which can reduce bias and may have other advantages (Chernozhukov). Some of these later works utilize a dual characterization of causal objects using a `treatment bridge' akin to our result in Theorem 1.1.i. Nagasawa2018 considers identification and estimation with proxy controls using a control function approach.
We consider the fully nonparametric estimation and inference problems. Other work on these problems includes Singh2020 and Mastouri2021, who employ RKHS-based approaches.
One important point of distinction between our analysis of estimation using proxy controls and that elsewhere in the literature, is that we avoid assuming smoothness of solutions to the identifying conditional moment restrictions. These functions, sometimes referred to in the literature as `bridge functions', lack a clear structural interpretation, and under our identifying assumptions they may not be unique. As such, it is difficult to motivate a priori conditions that these functions have say, bounded derivatives or are well approximated by polynomials. Fortunately, we are able to establish asymptotic guarantees for our estimates without the need for such restrictions on the bridge functions as we discuss in Section 4.
Our work contributes to the broader literature on nonparametric and non-separable measurement error models and their applications. Hu2008 provide identification results for nonparametric and non-separable models with measurement error and present a related estimator. The results of Hu2008 have been used for estimation in panel models using a factor analytic approach. Notably in Hu2012 and Arellano2016. Wilhelmb applies their results to a panel model with noisy measurements of the covariates of interest. The results of Hu2008 is adapted to causal inference with mismeasured confounders in the context of regression discontinuity by Rokkanen2015 and may be adapted to the problem of proxy controls more generally, as discussed in Deaner2022. As shown in Deaner2022, the exclusion restrictions of Hu2008 differ from those in this paper and other closely related works. Moreover, the corresponding characterization of causal objects differs markedly and motivates a distinct estimation strategy.
Our panel analysis follows a long line of work in which observations from other periods are used to account for unobserved heterogeneity. This approach is the basis of classic methods like those of Hausman1981b, Holtz-Eakin1988b, and Arellano1991 and some more recent work that allows for non-linearity like Freyberger2018 and Evdokimov2009.
An extensive literature examines the effects of grade retention on cognitive and social success. For a meta-analysis see Jimerson2001. We build on the work of Fruehwirth2016. We use the cleaned data available with their paper and we estimate some of the same causal effects. Recent work to estimate consumer demand counterfactuals (in particular, structural Engel curves) in nonparametric/semi-parametric models includes the instrumental variables approach of Blundell2007 and the panel approach of Chernozhukov2015a. For a short survey see Lewbel.
Consider an outcome of interest $Y$ and a column vector of treatments $X$. We are interested in the causal effect of $X$ on $Y$ which we define in terms of potential outcomes. We denote by $Y(x)$ the potential outcome from a counterfactual level $x$ of the treatment.
The identification of causal effects of $X$ on $Y$ is complicated by the presence of confounders: factors which jointly determine $X$ and $Y$. In order to adjust for the confounding, researchers may control for a set of pre-treatment covariates $D$. A problem arises when there are additional confounding factors that are not captured in $D$. Let $W$ be vector of latent factors that includes all the remaining confounders. If $W$ were observed then we could control for $W$ and $D$ in the usual way to recover the effect of $X$ on $Y$.
For example, suppose $X$ is some feature of a student's education and $Y$ measures the student's adult wages. We may have access to covariates $D$ that represent demographic characteristics of the student. Controlling for $D$ is likely insufficient to recover the causal effect of $X$ on $Y$. This is because $D$ does not fully account for a persistent component of academic aptitude $W$ which may confound $X$ and $Y$.
While $W$ is not observed, we suppose that the researcher has access to two vectors of covariates $V$ and $Z$ which provide some noisy information about $W$. For example, $V$ and $Z$ may include sets of test scores which are informative about the student's underlying academic ability $W$. We refer to $V$ and $Z$ as `proxy controls' because they act as stand-ins (or `proxies') for latent factors which we would like to control for.\footnote{In the econometrics literature one sometimes says that a variable $\tilde{X}$ is a `proxy' for $X$ if $X=\tilde{X}+\epsilon$ with $\epsilon$ independent of $\tilde{X}$. We use the term more loosely to simply refer to a variable that is informative about another.}
The conditions we place on the two vectors of proxies are not symmetric and in particular they differ in terms of the restrictions on their relationships with treatments and outcomes. We place relatively weak restrictions on the relationship between $V$ and the outcome $Y$ and so we refer to $V$ as `outcome-aligned' proxies. We place only weak restrictions on the association between $Z$ and $X$, and thus we refer to $Z$ as `treatment-aligned' proxies.
An informal summary of the key identifying conditions is provided below. This may act as a check-list for applied researchers interested in applying proxy control methods. We formalize and examine these conditions later in this section.
We suggest researchers check these conditions in turn, first attempting to write down a minimal set of latent factors $W$ so that the first condition is credible before proceeding to assess the next.
\theoremstyle{definition} \newtheorem*{GRS}{Example 1: Returns to Schooling}
Assumptions 1-3 below strengthen the identifying assumptions in linear case so that we can achieve identification in nonparametric, non-separable models.
\theoremstyle{definition} \newtheorem*{A1}{Assumption 1 (Independence)} \begin{A1} i. $Y(x) \perp\!\!\!\perp (X,Z)|(W,D)$, and ii. $V \perp\!\!\!\perp (X,Z)|(W,D)$ \end{A1}
\newtheorem*{A2}{Assumption 2 (Informative Proxies)} \begin{A2} For every function $\delta$ with the property that $0<E[\delta(W)^2|X=x,D=d]<\infty $: \[ \text{i.}\, E\big[E[\delta(W)|X,D,Z]^2|X=x,D=d\big]>0 \] \[ \text{ii.}\,E\big[E[\delta(W)|X,D,V]^2|X=x,D=d\big]>0 \] \end{A2}
\theoremstyle{definition} \newtheorem*{A3}{Assumption 3 (Full Support)} \begin{A3} i. $W$ has a conditional probability density function $f_{W|XD}$ (with respect to some dominating measure) that is strictly positive on the support of $W$. ii. $E[Y^2]<\infty$ and $E[Y(x)^2]<\infty$. \end{A3}
Assumption 1 effectively requires $W$ contains not only all the latent confounders, but any unobserved factor that directly impacts both a) $V$ or $Y$, and b) $Z$ or $X$. If there were additional factors that determined say, $Z$ and $Y$, then $Z$ and $Y(x)$ would likely be correlated even after controlling for $X$ and $W$.
In practice, we suggest that researchers list a minimal set of factors $W$ that satisfy the condition in the previous paragraph before proceeding to justify the remaining assumptions. For example, in the Griliches setting, one must consider whether univariate ability and family characteristics are the only variables that impact one of adult wages or early test scores, and which also affect either time is schooling or late test scores. If the researcher believes that test-taking ability may also be an important determinant of both early and late test scores then this should be included in $W$.
Assumption 1 rules out certain causal relationships between the proxies, treatments and outcomes. The condition proscribes any direct causal relationship between outcome-aligned proxies $V$ and treatments, as well as any direct causation between treatment-aligned proxies $Z$ and the outcome. These exclusion restrictions hold in the Griliches model, for example the early test scores $V$ do not impact time in schooling $X$.
Assumption 1 is implied by a nonparametric analogue of the Griliches model. The model on the right of Figure 1.a has the same exclusion restrictions as ((ref)) and these can be justified by the same arguments. The structural functions $g_Y$, $g_X$ etc. are possibly non-linear and the noise terms are taken to be jointly independent rather than uncorrelated. In this model $Y(x)$ is equal to $g_Y(x,W,U_Y)$. The directed graph on the left of Figure 1.a encodes the exclusion restrictions of the model.
The model in Figure 1.a is a nonparametric structural equations model (NSEM) of the type in Pearl2009. NSEMs are structural (i.e., causal) models that can act as intuitive primitive conditions for conditional independence restrictions. As in Figure 1.a, the restrictions of any NSEM can be neatly encoded in a directed graph. The NSEMs associated with the other graphs in Figure 1 also imply Assumption 1. This would remain true if we include $D$ as a cause of all other variables. These models are not exhaustive. The two graphs generalize the model in Figure 1.a. They allow $V$ to impact $Y$, include unobservables $Q$ which may affect both $Y$ and $V$, and allow direct causation between $X$ and $Z$.
Any NSEM that implies Assumption 1 must have the following properties. There must be no direct causation between treatment-aligned proxies $Z$ and the outcome $Y$, nor between outcome-aligned proxies $V$ and treatment $X$. Neither vector of proxies $Z$ nor $V$ must directly impact the other, and $Y$ must not impact $V$.
Assumption 2 ensures that the proxies $V$ and $Z$ are each sufficiently informative about the unobserved confounders. Similar conditions are used for identification in nonparametric factor models (see Hu2008). Assumption 2 is a nonparametric analogue of the rank conditions for identification in the linear model in Example 1. In fact, the two coincide when the variables are jointly normal (see Newey2003).
For this reason we suggest that researchers assess the credibility of Assumption 2 by considering whether the corresponding rank conditions are credible, noting that the two coincide in Gaussian models but are otherwise non-nested. At the end of our discussion of Example 1 we note that researchers might justify the rank conditions by arguing that, for each factor in $W$, there is a proxy in both $V$ and $Z$ that is particularly strongly associated with this factor compared to the other factors.
The rank conditions require that the dimensions of $V$ and $Z$ are each weakly greater than that of $W$. That is, we have at least as many proxies in each vector as there are confounding factors. Strictly speaking, this order condition is not necessary for Assumption 2. However, because Assumption 2 coincides with the rank condition in the Gaussian case, it seems prudent that researchers consider the order condition when determining whether Assumption 2 is credible.
Recall that the rank conditions in Example 1 ensure that the proxies $V$ and $Z$ are relevant instruments for $W$ in the sense of linear instrumental variables. Similarly, Assumption 2 ensures $V$ and $Z$ are relevant instruments for $W$ as required in Nonparametric Instrumental Variables (NPIV) models. To be more precise, Assumption 2.i is a statistical completeness condition which identifies an NPIV model of the kind in Newey2003 and Ai2003 where $Z$ is an instrument for $W$, and $X$ and $D$ are exogenous regressors. 2.ii is identical but with $V$ in place of $Z$. Newey2003 state that completeness is analogous to the rank condition in linear instrumental variables. Strictly speaking, Assumption 2 imposes $L_2$-completeness, which is slightly weaker than the standard completeness condition (see Andrews2017b).
Note that Assumption 2 differs from the completeness conditions in Miao2018. In particular, Miao2018 require that one vector of proxies is statistically complete for the other. In Gaussian models this rules out the possibility that both vectors of proxies are of higher dimension than $W$. By contrast, if Assumption 2 holds for some $V$ and $Z$ then it also hold for any ${V}^*$ and ${Z}^*$ of which $V$ and $Z$ are subvectors.\footnote{Strictly speaking neither Assumption 2 nor the completeness conditions of Miao2018 are stronger. However, as we note in the proof of Theorem 1.1, Assumption 2 can be weakened slightly and identification still holds. If Assumption 1 holds then the weakened version of Assumption 2 is in fact weaker than the conditions in Miao2018.} Conditions of essentially the same form to Assumption 2 are employed by Shi2020 in the context of categorical confounding and proxies.
The full support condition in Assumption 3.i is required for identification even in the case with $W$ observed. In the case of a binary treatment, the condition ensures that the propensity score (from covariates $W$) is non-zero.
We state our main identification result Theorem 1.1 below. In the condition $\mathcal{V}_{x,d}$ and $\mathcal{Z}_{x,d}$ respectively denote the supports of $V$ and $Z$ conditional on $X=x$ and $D=d$.
\theoremstyle{plain} \newtheorem*{Th11}{Theorem 1.1 (Identification)} \begin{Th11} Suppose Assumptions 1-3 hold for $x=x_1,x_2$. i. For any function $\varphi$ with $E[\varphi(Z)^{2}|X=x_{1},D=d]<\infty$ and:
the conditional average potential outcome satisfies:
ii. For any function $\gamma$ with $E[\gamma(V)^{2}|X=x_1,D=d]<\infty$ and:
the conditional average potential outcome satisfies:
\end{Th11}
So long as either ((ref)) or ((ref)) has a solution, Theorem 1.1 identifies the conditional average potential outcome. Note $\varphi$ that satisfies ((ref)) may depend on $x_1$, $x_2$, and $d$. Similarly $\gamma$ in ((ref)) may depend on $x_1$ and $d$. $\gamma$ is sometimes referred to elsewhere in the literature as a `confounding bridge function' or `outcome bridge' and $\varphi$ (or a closely related object) as a `treatment bridge'.
The existence of solutions to ((ref)) and ((ref)) follows from Assumption 2 and a set of regularity conditions. The regularity conditions, which apply Picard's Criterion to this setting, are somewhat technical and so we relegate them to Appendix A along with further discussion. We provide semiparametric examples in which ((ref)) and ((ref)) admit solutions below. Note that our conditions do not imply that the solutions to these equations are unique.
Theorem 1.ii is similar to the result of Miao2018 but identifies the conditional average potential outcome rather than just an unconditional average. Thus it allows for identification of say, the average effect of treatment on the treated and other conditional average treatment effects.
Theorem 1.i is fully original. It provides an alternative path to identification using a moment condition that does not depend on the outcome $Y$. This result plays a crucial role in our analysis of nonparametric estimation using proxy controls, as we discuss in Section 4.
Note the similarity between ((ref)) and the moment condition that identifies $\beta$ in the linear model of Example 1. Assuming no additional controls $D$, and making the dependence of $\gamma$ on $x_1$ explicit, ((ref)) can be written as: \[ E[Y-\gamma(V,x_1)|X=x_1,Z=z]=0 \] If the above holds for all $x_1$ and $z$ in the joint support of $X$ and $Z$, then we have: \[ E\big[\big(Y-\gamma(V,X)\big)(X',Z')\big]=0 \] If we restrict $\gamma$ to be linear then the above is precisely the moment condition that identifies $\beta$ in the linear model.
\theoremstyle{definition} \newtheorem*{EE1}{Semiparametric Example 1: Multivariate Normal} \begin{EE1} Suppose for simplicity that there are no additional confounders $D$ and that $W$ and $Z$ are jointly normal conditional on each value $x$ of $X$, with a strictly positive definite conditional variance-covariance matrix. \[
|X=x\sim\mathcal{N}\bigg(
,
\bigg) \]
Then Assumption 2.i holds if and only if $\Sigma_{WZ|x}$ has rank equal to the dimension of $W$. Note that this implies $Z$ has weakly greater dimension than $W$. We can replace $Z$ with $V$ to get an equivalent characterization of Assumption 2.ii.
Suppose Assumptions 1-3 hold and the conditional variance of $W$ does not vary with $X$ (i.e., $\Sigma_{WW|x_{2}}=\Sigma_{WW|x_{1}}$). This holds for example if $W$, $Z$, and $X$ are jointly Gaussian. Then ((ref)) holds with $\varphi$ of the form below, where $\alpha$ is a scalar and $M$ a vector:
More generally, if $\Sigma_{WW|x_{2}}\neq \Sigma_{WW|x_{1}}$ but the two matrices are sufficiently close, ((ref)) holds with:
Where $P$ is a matrix which may depend on $x_1$ and $x_2$. The precise condition on $\Sigma_{WW|x_{2}}$ and $\Sigma_{WW|x_{1}}$ is given in Appendix A along with a proof of this result.
\end{EE1}
\theoremstyle{definition} \newtheorem*{SE2}{Semiparametric Example 2: Partially Linear Model} \begin{SE2}
Suppose for simplicity that there are no additional confounders $D$, and that conditional on $X=x_1$, $V$ and $W$ are jointly normally distributed with strictly positive definite variance-covariance matrix. In addition suppose that $E[Y^2|X=x_{1}]$ is finite and the conditional mean of $Y$ is linear in $W$ (but not necessarily $X$): \[ E[Y|W=w,X=x_{1}]=A(x_{1})+B(x_{1})w \] If Assumption 2 holds there exist a scalar $a$ and vector $b$ (possibly dependent on $x_1$) so that ((ref)) is satisfied by: \[ \gamma(V)=a+b'V \]
\end{SE2}
In the previous section we assume the existence of two vectors of proxies $V$ and $Z$ for the unobserved confounding factors. In panel settings, past and future treatments and outcomes are a natural source of relevant proxies. By definition, unobserved confounders affect both outcomes and treatments. If there is confounding in every period, and if the confounders are persistent, then treatments and outcomes from other periods are likely to be correlated with today's confounders.
We leverage serial independence restrictions on the treatments and outcomes to ensure that proxies $V$ and $Z$, and controls $D$, each composed of treatments and outcomes, are valid in the sense that Assumption 1 holds.
Our interest is in panel data with a fixed number of time periods. Observations are available from periods $t=1,...,T$, where $T$ is fixed. When $T$ is large there are more potential proxies and thus we can allow for richer panel dynamics while still finding proxies that satisfy Assumption $1$. Moreover, if the proxies $V$ and $Z$ contain more observables, then the informativeness conditions in Assumption 2 are more plausible.
In the discussion below we let $X_t$ denote the treatment at period $t$, and $Y_t$ the outcome at $t$. $W_t$ denotes the unobserved confounding factors at time $t$. In addition, we subscript a variable with $[s:r]$ with $s\leq r$ to denote the observations of that variable from periods $s$ to $r$ inclusive, for example $X_{[s:r]}=\{X_s,X_{s+1},...,X_{r-1},X_r\}$. For convenience, if $s> r$ we take $X_{[s:r]}=\emptyset$.
We begin by considering models in which the unobserved confounders are constant over time. That is, $W_t=W$ for all $t$. These can be understood as non-parametric fixed effects models in which a set of time-invariant individual-specific factors (or `fixed effects') $W$, can determine both treatments and outcomes.
Crucial to our analysis is that some of the observables exhibit Markov dependence. By this we mean that given the unobserved individual-characteristics $W$, the association between observables in the past and future is explained by observables in the intervening periods.
To see why Markov dependence is useful, suppose that treatments (and not necessarily outcomes) are first-order Markov dependent. That is, for each $t=3,4,..,T$:
Let $Y_t(x_t)$ denote the potential value of $Y_t$ under a counterfactual intervention that sets $X_t$ equal to the (non-random) value $x_t$. In addition to the above, suppose that the fixed-effects $W$ explain all confounding between the outcome in some period and the treatments up to that point so that the following holds for each $t=1,2,...,T$:
The condition above is a nonparametric analogue of the standard `weak exogeneity' or `predetermination' condition in linear panel data and time-series models. This condition allows for feedback from past outcomes to future treatments. However, it rules out any effect of past treatments on today's outcome unless that effect is mediated by today's treatment.
The Markov dependence condition ((ref)) and weak exogeneity ((ref)) imply vectors of proxies $V$ and $Z$ formed from past treatments are valid in the sense of Assumption $1.1$. In particular, if $T>3$ then Assumption $1$ holds with $X=X_T$, $Y=Y_T$, $D=X_{t-1}$, $V=X_{[1:t-2]}$, and $Z=X_{[t:T-1]}$, where $t$ can be any number between $3$ and $T-1$ inclusive. This follows from the more general results later in this section.
The Markovian dependence in ((ref)) is crucial because it ensures that after conditioning on $W$ and on treatment in period $t-1$, the treatments prior to this period are independent of those after. This allows us to compose $V$ and $Z$ from the history of treatments without violating Assumption 1.ii. In fact, as we discuss below, we can choose valid $V$ and $Z$ using treatments and/or outcomes from the past and/or future under a range of different modeling assumptions.
To better understand the kinds of modeling assumptions that justify Markovian dynamics, we frame the remainder of this subsection in the dynamic, nonparametric structural (i.e., causal) model given below which holds for $t=1,...,T$.
In the model above $g_{Y,t}$ and $g_{X,t}$ are non-random structural functions. The time subscripts indicate that these functions may vary over time. $U_{Y,t}$ and $U_{X,t}$ are period-specific and individual-specific noise terms. The model allows the dynamics of treatments and outcomes to vary both between individuals and over time due to the presence of $W$ and the time-dependence of the structural functions. In the model all observables are directly impacted by the confounders $W$, suggesting they constitute informative proxies.
The constants $\kappa_{YY}$, $\kappa_{YX}$, $\kappa_{XY}$, and $\kappa_{XX}$ determine how far back in the past an observable can be and still directly impact treatments and outcomes today. For example, if $\kappa_{YY}=1$ then $Y_{t-1}$ may directly affect $Y_t$ but earlier outcomes only impact $Y_t$ indirectly.
Note that treatments or outcomes in earlier periods may be impacted by treatments and outcomes from periods prior to period $1$, the first period for which there is data. Because the model need only hold for $t=1,...,T$ we avoid any initial conditions assumption.
In a structural model like ((ref)) potential outcomes can be defined by replacing a variable with a fixed value in every equation in which it appears. We focus on the potential value of $Y_t$ under a counterfactual value $x_t$ of $X_t$ which is given by: \[Y_t(x_t)= g_{Y,t}(Y_{[t-\kappa_{YY}:t-1]},X_{[t-\kappa_{YX}:t-1]},x_t,W,U_t) \]
Figure 2.a depicts the exclusion restrictions of the model ((ref)) in the special case of $\kappa_{YX}=0$ and $\kappa_{YY}=\kappa_{XY}=\kappa_{XX}=1$. We omit $W$, which directly affects all variables, for the purpose of legibility.
The model ((ref)) implies conditional independence restrictions. These conditions imply Assumption 1 holds for appropriate choices of proxies $V$ and $Z$, and conditioning variables $D$. In particular, in the special case depicted in Figure 2.a, one can confirm that Assumption 1 holds for $V=\{X_1,Y_1\}$, $Z=\{X_3\}$, $D=\{X_2,Y_2,Y_3\}$, $X=X_4$, and $Y=Y_5$.
These choices are indicated on the graph by color-coding the nodes. The nodes for variables in $V$ are colored purple, those in $Z$ are red, $D$ is yellow, $X$ is blue, and $Y$ is green. Note that the choice of $D$ `blocks' (or `$d$-separates') all paths from variables in $V$ to variables in $Z$, ensuring that $V$ and $Z$ are independent conditional on $D$ and $W$ (see Section 1.2.3 of Pearl2009).
We first consider the special case of ((ref)) in which $\kappa_{XY}=0$. Under this restriction $Y_t$ has no direct affect on future treatments. This has important consequences for the choice of proxies because our assumptions generally preclude the possibility that $V$, $Z$, or $D$ are caused by the outcome of interest (directly or indirectly). Thus when $\kappa_{XY}\neq0$ we can only form proxies using past observables whereas when $\kappa_{XY}=0$ we may use future treatments or outcomes as proxies.
We begin by considering the special case of $\kappa_{XY}=0$. Under this restriction we may use future treatments as treatment-aligned proxies. If, in addition, $\kappa_{YY}=0$ then we can use future outcomes as treatment-aligned proxies as well. The outcome-aligned proxies consist of lagged outcomes.
\theoremstyle{plain} \newtheorem*{P333}{Proposition 2.1} \begin{P333} Suppose $\kappa_{XY}=0$. Then for any $t$ the model ((ref)) implies the following:
If we also have $\kappa_{YY}=0$ then letting $k_X=\max\{\kappa_{XX},\kappa_{YX}-1\}$ we have:
\end{P333}
Under the conditional independence results from Proposition 2.1 we can compose the treatment aligned proxies $Z$ from future treatments and outcomes, and the outcome-aligned proxies $V$ using past outcomes. Theorem 2.2 does not require that the model ((ref)) holds. Rather, it requires the conclusions of Propositions 2.1 which are implied by, and are thus weaker than, the model ((ref)) with the restriction $\kappa_{XY}=0$.
\theoremstyle{plain} \newtheorem*{T222}{Theorem 2.1} \begin{T222} Suppose that for each $1\leq t\leq T$, ((ref)) and ((ref)) hold. Then for each $1 < t < T$, Assumption $1$ holds with $Y=Y_{t}$, $X=X_{t}$, $V=Y_{[1:t-1]}$, $Z=X_{[t+1:T]}$, and $D=X_{[t-\kappa_{XX}:t-1]}$. If ((ref)) and ((ref)) hold and we instead set $D=X_{[t-k_X:t-1]}$, then we can also include $Y_{[t+1:T]}$ in $Z$. \end{T222}
Theorem 2.1 suggests choices of proxies $V$ and $Z$, and additional controls $D$, so that Assumption 1 holds. The theorem defines these as sets of random variables but we can simply stack the elements to form vectors. The treatments and outcome of interest are those from period $t$. Note that $1<t<T$ and so we require at least three periods of data and we cannot use Theorem 2.1 to identify causal effects in the first and final period.
The controls $D$ are composed of lagged treatments. Consider the choice $D=X_{[t-\kappa_{XX}:t-1]}$. For all components of $D$ to be included in the observed data, we require that $t> \kappa_{XX}$. To ensure this to holds for at least one choice of $t$, the number of periods for which data is available must weakly exceed $\kappa_{XX}+2$.
Note that Proposition 2.1 and Theorem 2.1 together allow one to form valid proxies without any restrictions on $\kappa_{YX}$. That is, we can allow treatments to directly impact outcomes indefinitely far in the future and still form proxies so that Assumption 1 holds.
In the case of $\kappa_{XY}=0$ only, $V$ can be composed ot $t-1$ lagged outcomes and $Z$ composed of $T-t-1$ future treatments. If many periods of data are available then this may amount to a large number of proxies in $V$ and $Z$. While a larger set of proxies increases the credibility of the completeness conditions (Assumptions 2.i and 2.ii), nonparametric estimation with many proxies may result in noisy estimates. Therefore it may be preferable to include only a subset of these proxies. This is not problematic for Assumption 1 because if the condition holds for some choice of $V$ and $Z$, then it also holds for subvectors of these variables.
We apply Theorem 2.1 to the model in Figure 2.b. We let $D=X_1$, $V=Y_1$, $Z=X_3$, $X=X_2$, and $Y=Y_2$. As in Figure 2.a, the nodes are color-coded to indicate inclusion of a variable in either $D$, $V$, $Z$, $X$, or $Y$ and we omit $W$ which is a cause of all other variables. Some of the arrows are colored to aid legibility.
We now consider the case in which lagged outcomes may impact future treatments (that is, $\kappa_{XY}\neq 0$). This generally precludes the use of future observables as proxies. However, lagged observables may still form valid proxies under sufficient restrictions on the model ((ref)).
\theoremstyle{plain} \newtheorem*{P331}{Proposition 2.2} \begin{P331} For any $t$, the model ((ref)) implies the following conditional independence restriction where we define $k_X=\max\{\kappa_{XX},\kappa_{YX}\}$ and $k_Y=\max\{\kappa_{XY},\kappa_{YY}\}$.
Moreover, we have:
\end{P331}
Given the independence conditions in Proposition 2.2, Theorem 2.2 below states that Assumption $1$ holds for proxies $V$ and $Z$, and conditioning variables $D$, formed from appropriate choices of lagged observables. We focus on the case in which the causal effect of interest is of $X_T$ on $Y_T$, however for earlier periods causal effects can be identified in the same way by ignoring data from subsequent periods.
\theoremstyle{plain} \newtheorem*{T221}{Theorem 2.2} \begin{T221} Suppose for every $1\leq t\leq T$, ((ref)) and ((ref)) hold. Then for every $t$ that is strictly greater than $\min\{k_X,k_Y\}+1$ and strictly less than $T-\max\{\kappa_{YX},\kappa_{YY}\}$, Assumption $1$ holds with $Y=Y_T$ and $X=X_T$ so that $Y(x)=Y_T(x_T)$, and with $V$, $Z$, and $D$ defined as follows.
$V=X_{[1:t-k_X-1]}\cup Y_{[1:t-k_Y-1]}$, $Z=X_{[t:T-\kappa_{YX}-1]}\cup Y_{[t:T-\kappa_{YY}-1]}$, and $D=D_1\cup D_2$ where $D_1=X_{[t-k_X:t-1]}\cup Y_{[t-k_Y:t-1]}$ and $D_2=X_{[T-\kappa_{YX}:T-1]}\cup Y_{[T-\kappa_{YY}:T-1]}$.
\end{T221}
Theorem 2.2 shows that, under the conditional independence restrictions from Proposition 2.2, we can form valid proxies using past observables. $V$ consists of observables from some periods prior to a period $t$, $Z$ contains observables from periods after $t$. For there to exist a $1\leq t\leq T$ that satisfies the conditions of the theorem, the number of periods $T$ must be sufficiently large. In particular, we need $\min\{k_X,k_Y\}+\max\{\kappa_{YX},\kappa_{YY}\}+2<T$.
Note that the vector of conditioning variables $D$ is composed of two sub-vectors. Conditioning on the first of these, $D_1$ (along with $W$), renders $V$ independent of $Z$ and $X$. Conditioning on $D_2$ ensures potential outcomes are jointly independent of $X$ and $Z$.
We apply Theorem 2.1 to the model in Figure 1.c. As in Figures 1.a and 1.b, we omit $W$ from the graph. Applying Theorem 2.2 with $t=3$, we have $V=X_1$, $Z=X_3$, $D=X_2\cup Y_{[1:3]}$. These choices are color-coded as in Figures 1.a and 1.b. Applying Theorem 2.1 to the setting in Figure 1.a yields the choices for $V$, $Z$, $D$, $X$, and $Y$ described earlier and indicated by the color-coding on that graph. The stronger restrictions in Figure 1.a allow us to use $Y_1$ as an outcome-aligned proxy, whereas in Figure 1.c it must be included as an additional control.
We now consider models in which the confounding factors $W_t$ may vary over time. We generalize ((ref)) to the following model:
In this model past values of the confounders $W_t$ from multiple times periods may directly affect outcomes and treatments in each period. Moreover past confounding may directly determine future confounders.
We show that the results of Proposition 2.1 and Theorem 2.1 generalize straight-forwardly to this case. As before we must restrict the model so that past outcomes have no direct affect on future treatments. Similarly, we assume that past values of the outcomes do not directly impact future confounders. That is, we assume both $\kappa_{XY}=0$ and $\kappa_{XW}=0$.
\theoremstyle{plain} \newtheorem*{P24}{Proposition 2.3} \begin{P24} Suppose $\kappa_{XY}=\kappa_{WY}=0$. Letting ${k}_X=\max\{\kappa_{XX},\kappa_{WX}-1\}$ and defining ${k}_W=\max\{\kappa_{XW},\kappa_{WW}-1\}$,Then for any $t$ the model ((ref)) implies the following:
If we also have $\kappa_{YY}=0$ then letting $\bar{k}_X=\max\{\kappa_{XX},\kappa_{WX}-1,\kappa_{YX}-1\}$ and defining $\bar{k}_W=\max\{\kappa_{XW},\kappa_{WW}-1,\kappa_{YW}-1\}$, we have:
\end{P24}
\theoremstyle{plain} \newtheorem*{T223}{Theorem 2.3} \begin{T223} Suppose ((ref)), and ((ref)) hold for each $1\leq t\leq T$. Then for any $1<t<T$, Assumption $1$ holds with $Y=Y_{t}$, $X=X_{t}$, $V=Y_{[1:t-1]}$, $Z=X_{[t+1:T]}$, $D=X_{[t-k_{X}:t-1]}$, and $W=W_{[t-k_{W}:t]}$. If ((ref)) and ((ref)) hold and we instead set $D=X_{[t-\bar{k}_X:t-1]}$ and $W=W_{[t-\bar{k}_W:t]}$, then we can also include $Y_{[t+1:T]}$ in $Z$. \end{T223}
We apply Theorem 2.3 to the model in Figure 3.a. The choices of $V$, $Z$, $D$, $X$, and $Y$ are the same as in Figure 2.b. The nodes are color coded as before with the addition that the node for $W_2$ is colored orange to indicate that this variable plays the role of $W$ in Assumption 1.
In the case of time-varying confounders, the informativeness of the proxies is less evident. $W$ in Theorem 2.3 is composed of confounders from various periods, and these may not all directly cause or be caused by all elements of $V$ and $Z$. Of particular concern is that the choices of $D$, $Z$, $V$, and $W$ in Theorem 2.3 may lead to a necessary violation of Assumption 2. Consider an element $W_s$ of $W$ and let $W_{-s}$ denote the other elements of $W$. Assumption 2.i cannot hold if $W_s\perp \!\!\! \perp Z|(X_t,D,W_{-s})$, and similarly for Assumption 2.ii.\footnote{Strictly speaking Assumption 2 can hold only in the trivial case in which $W_s$ is non-random given $X_t$, $D$, and $W_{-s}$.}
However, we can rule out this possibility so long as we assume that the confounders are sufficiently persistent, that is, $\kappa_{WW}$ is sufficiently large. Consider the case in Theorem 2.3 in which $Z$ is composed of future treatments. If $\kappa_{WW}>\kappa_{XW}$, then one can show there is a path from each component of $W$ in the directed graph corresponding to ((ref)) to each component of $Z$, that is `unblocked' by $D$, $X_t$, and $W_{-s}$. This means that generically, each element of $Z$ is statistically associated with each element of $W$ even controlling for $D$, $X_t$, and the other elements of $W$. The same holds for each element of $V$. The same is true when $Z$ includes future treatments if we also assume that $\kappa_{WW}>\kappa_{YW}$. Note that these conditions on are sufficient to rule out necessary violations of Assumption 2 but they may not be necessary.
In Figure 2.a, we see that $V=Y_1$ is generically associated with $W_2$, even controlling for $D$ and $X_2$, because both are directly impacted by $W_1$. $Z=X_3$ is generically associated with $W_2$ even after controlling for $D$ and $X_2$ because $W_2$ impacts $W_3$ which is a direct cause of $X_3$.
Unfortunately, the violation of Assumption 2 described above precludes an extension of the case in which proxies are composed only of lagged observables to the setting with time-varying confounding.
In this section we describe our estimation and inference procedures. The key step in estimation corresponds to penalized sieve minimum distance (PSMD) estimation (see Cheng and Chen2015). Inference is based on the multiplier bootstrap (see for example Belloni2015). Our methods can be applied in panel settings or to cross-sectional data. To emphasize this generality we return to the notation in Section 1 in which we suppress time subscripts.
Let $\{(Y_{i},X_{i},Z_{i},V_{i},D_{i})\}_{i=1}^{n}$ be a random iid sample of $n$ observations of the variables $Y$, $X$, $Z$, $V$, and $D$. In the panel case, $Y_i$ and $X_i$ should be understood to come from one fixed period $t$. Our estimation method is of the sieve-type, and so we need to specify some vectors of basis functions. For each $n$ let $\phi_{n}(V,X,D)$ be a vector of transformations of $V$, $X$, and $D$. The practitioner estimates conditional means $g_i$, $\pi_{n,i}$, and $\alpha_n$ which are defined by:
The objects above can be estimated by fitted values from nonparametric regression. Many methods are available, for example local-linear regression, or series least-squares. We describe a particular procedure later in this sub-section.
Note that $\pi_{n,i}$ is a vector of the same length as $\phi_{n}(v,x,d)$ and in general we must perform a separate regression in order to estimate each of its components. However, we can reduce the number of regressions if $\phi_n$ has a multiplicative structure. Suppose that $\phi_{n}(v,x,d)$ is of the form $\rho_{n}(v)\otimes\chi_n(x,d)$ where `$\otimes$' is the Kronecker product, then we have: \[ \pi_{n,i} =E[\rho_{n}(V_i)|Z_i,X_i,D_i]\otimes\chi_{n}(X_i,D_i) \] Thus we need only perform one regression for each element of $\rho_n(v)$ rather than for each element of $\phi_n(v,x,d)$. The same applies for estimation of ${\alpha}_{n}(x_1,x_2,d)$.
Having obtained estimates $\hat{g}_i$, $\hat{\pi}_{n,i}$, and $\hat{\alpha}_n$ of $g_i$, $\pi_{n,i}$, and $\alpha_n$, the researcher evaluates a vector of coefficients $\hat{\theta}$. These coefficients minimize the penalized least-squares objective below:
$\lambda_{0}$ is a positive scalar penalty parameter that may change with the same size and controls the degree of regularization in the second stage. In our empirical applications we set $\lambda_0=\sqrt{n}$. The ridge regression problem has the closed-form solution $(\hat{\Sigma}_{\hat{\pi}}+\lambda_0 I)^{-1}\frac{1}{n}\sum_{i=1}^n \hat{\pi}_{n,i} \hat{g}$, where $I$ is the identity matrix and $\hat{\Sigma}_{\hat{\pi}}$ is equal to $\frac{1}{n}\sum_{i=1}^n \hat{\pi}_{n,i}\hat{\pi}_{n,i}'$. An estimate of the conditional average potential outcome is then given by:
A researcher may be interested in average potential outcomes conditional only on a function of $X_i$ and $D_i$. In order to achieve this one can replace $\alpha_n(x_1,x_2,d)$ with the mean of $\phi_n(V_i,x_1,D_i)$ conditional on that function of $X_i$ and $D_i$. For example, in order to estimate $E[Y(x_1)|X=x_2]$ we can replace $\hat{\alpha}_{n}(x_1,x_2,d)$ with $\hat{\alpha}_{n}(x_1,x_2)$ which estimates the regression function below: \[{\alpha}_{n}(x_1,x_2)=E[\phi_n(V_i,x_1,D_i)|X_i=x_2]\]
Suppose we are interested in the unconditional mean of potential outcomes $E[Y(x)]$. In this case we can replace $\hat{\alpha}_n(x_1,x_2,d)$ with $\hat{\alpha}_{n}(x)$ an estimate of $\alpha_n(x)=E[\phi_n(V,x,D)]$. We can simply use the sample mean: \[ \hat{\alpha}_n(x)=\frac{1}{n}\sum_{i=1}^n \phi_n(V_i,x,D_i) \] Then $\hat{\alpha}_{n}(x)'\hat{\theta}$ is an estimate of the average potential outcome $E[Y(x)]$.
In our empirical applications we use series ridge regression to estimate $g_i$, $\pi_{n,i}$, and $\alpha_n$. Some of our asymptotic results pertain to this particular choice of regression method. We assume that $\phi_n$ has the multiplicative structure defined earlier in this section. We must specify additional basis functions. Let $\zeta_{n}(Z,X,D)$ be a vector of transformations of $Z$, $X$, and $D$ and likewise for $\psi_n(Z,X,D)$. We let $\zeta_{n,i}=\zeta_{n}(Z_i,X_i,D_i)$ and similarly for $\psi_{n,i}$, $\rho_{n,i}$ and $\chi_{n,i}$. In addition, take $\rho_{n,i,k}$ to be the $k$-th component of the vector $\rho_{n,i}$. Define regression estimates $\hat{\beta}_g$, $\hat{\beta}_{\pi,k}$, and $\hat{\beta}_{\alpha,k}$ as follows:
$\lambda_g$, $\lambda_{\pi}$, and $\lambda_{\alpha}$ are penalty parameters. In our empirical applications we set these to zero. Stack the vectors $\hat{\beta}_{\pi,k}$ into a matrix $\hat{\beta}_{\pi}$ and similarly for $\hat{\beta}_{\alpha,k}$. Our estimates of $g_i$, $\pi_{n,i}$, and $\alpha_n$ are then:
In order to perform inference we specify a multiplier bootstrap procedure. Let $Q_{b,i}$ for each $i=1,...,n$ and $b\in\{1,...,B\}$ be iid standard exponential random variables that are independent of the data.\footnote{We follow (Belloni2015) and use the standard exponential, but $Q_{b,i}$ may have any distribution with mean and variance both equal to $1$ and $\max_{1\leq i\leq n }|Q_{b,i}|\precsim_{p}ln(n)$.}
For each $b$ we re-estimate $g_i$, $\pi_{n,i}$, and $\alpha_i$ using the method in the previous subsection, but weighting the contribution of the $i^{th}$ observation by $Q_{b,i}$ in each summation.
For example, consider $\hat{g}_i$, our estimate of $g_i$ defined in the previous subsection. We define the $b^{th}$ bootstrap estimate of $g_{i}$ by $\hat{g}_{b,i}=\psi_{n,i}'\hat{\beta}_{g,b}$ where $\hat{\beta}_{g,b}$ is given below: \[ \hat{\beta}_{g,b}=\big(\frac{1}{n}\sum_{i=1}^n Q_{b,i}\psi_{n,i}\psi_{n,i}'+\lambda_g I\big)^{-1}\frac{1}{n}\sum_{i=1}^n Q_{b,i}\psi_{n,i}y_{i} \] We define the $b^{th}$ bootstrap estimates of $\pi_{n,i}$ and ${\alpha}_{n}$, denoted by $\hat{\pi}_{b,n,i}$ and $\hat{\alpha}_{b,n}$, similarly, again replacing averages in the formulas with weighted averages.
Let $\hat{\theta}_{b}$ solve the weighted, penalized, least squares objective below: \[ \hat{\theta}_b=\arg\min_{\theta}\frac{1}{n}\sum_{i=1}^{n}Q_{b,i}\big(\hat{g}_{b,i}-\hat{\pi}_{b,n,i}'\theta\big)^{2}+\lambda_{0}\|\theta\|^2 \] Thus we obtain a bootstrap sample $\{\hat{\alpha}_{b,n}(x_{1},x_{2},d)'\hat{\theta}_{b}\}_{b=1}^B$. We can form a bootstrap sample that corresponds to estimates of say, the average potential outcome $E[Y(x)]$, similarly. We obtain bootstrap analogues of the estimator of $E[\phi_n(V,x_1,D)]$, denoted $\hat{\alpha}_{n}(x)$ earlier in this section. Again we achieve this simply by replacing averages with weighted averages in the formula. We thus obtain a bootstrap sample $\{\hat{\alpha}_{b,n}(x)'\hat{\theta}_{b}\}_{b=1}^B$.
We can use the bootstrap sample to form uniform confidence bands. A uniform confidence band is a collection of intervals, in our case, one for each value $x_1$, $x_2$, and $d$ in a set $\mathcal{S}$. Such a band is asymptotically valid if every interval contains the value of the function at the corresponding $x_1$, $x_2$, and $d$ with probability approaching the desired level. We consider $1-a$-level intervals of the form: \[ \bigg[\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}-\frac{\hat{\sigma}(x_{1},x_{2},d)}{\sqrt{n}}\hat{c}_{1-a},\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}+\frac{\hat{\sigma}(x_{1},x_{2},d)}{\sqrt{n}}\hat{c}_{1-a}\bigg] \] $\hat{\sigma}(x_{1},x_{2},d)/\sqrt{n}$ is an estimate of the standard deviation of $\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}$. In practice we use the pointwise standard deviation of the bootstrap sample as our estimate of $\hat{\sigma}(x_{1},x_{2},d)/\sqrt{n}$. $\hat{c}_{1-a}$ is a uniform critical value that is equal to the smallest scalar $c>0$ that satisfies the inequality below: \[ \frac{1}{B}\sum_{b=1}^{B}1\big\{\sup_{(x_{1},x_{2},d)\in\mathcal{S}}\big|\frac{\hat{\alpha}_{b,n}(x_{1},x_{2},d)'\hat{\theta}_{b}-\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}}{\hat{\sigma}(x_{1},x_{2},d)/\sqrt{n}}\big|\leq c\big\}\geq1-a \]
Again, we can adapt the procedure above to perform inference on say, $E[Y(x)]$ or $E[Y(x_1)|X=x_2]$, simply by replacing $\hat{\alpha}_{b,n}(x_{1},x_{2},d)$ with a bootstrap estimate of $E[\phi_n(V,x,D)]$ or $E[\phi_n(V,x_1,D)|X=x_2]$.
The application of our approach to panel models, as detailed in Section 2, requires a researcher to take a stance on the degree of Markovian dependence. Generally speaking, stronger restrictions on the degree of Markovian dependence allow for more precise estimates because the researcher has more proxies to choose from and needs to condition on a smaller set of observables.
In order to inform these modeling choices we suggest a Hausman specification test (Hausman1978). Given two alternative sets of modeling assumptions, the researcher constructs appropriate estimates $\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}$ and $\tilde{\alpha}_{n}(x_{1},x_{2},d)'\tilde{\theta}$ as specified in the previous subsections. We suppose that the first of these estimates is consistent under weaker conditions than the second. Thus the goal is to test the null hypothesis that both sets of assumptions hold against the alternative that the stronger conditions fail. It suffices to test the null that the difference between the probability limits of the two estimators is zero.
In some cases the two specification to be tested may require two different choices of additional controls $D$. In this case researchers can instead compare corresponding estimates $\hat{\alpha}_{n}(x_{1},x_{2})'\hat{\theta}$ and $\tilde{\alpha}_{n}(x_{1},x_{2})'\tilde{\theta}$ of $E[Y(x_1)|X=x_2]$. In this case one proceeds precisely as we describe below but using these estimates and their bootstrap analogues in place of estimates of $E[Y(x_1)|X=x_2,D=d]$.
In order to carry out the test, a researcher constructs uniform confidence bands for the difference between the two estimates. We denote this difference by $\hat{\Delta}(x_1 ,x_2 ,d)$. That is: \[ \hat{\Delta}(x_1 ,x_2 ,d)=\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}-\tilde{\alpha}_{n}(x_{1},x_{2},d)'\tilde{\theta} \]
In order to construct confidence bands, the researcher obtains a bootstrap sample $\{\hat{\alpha}_{b,n}(x_{1},x_{2},d)'\hat{\theta}_{b}\}_{b=1}^B$ corresponding to the estimator $\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}$, as described earlier in this section. Then, using the same exponential weights, the researcher constructs a bootstrap sample $\{\tilde{\alpha}_{b,n}(x_{1},x_{2},d)'\tilde{\theta}_{b}\}_{b=1}^B$ that corresponds to the estimator $\tilde{\alpha}_{n}(x_{1},x_{2},d)'\tilde{\theta}$.
For a level $1-a$ test, the researcher calculates a critical value $\hat{c}_{1-a}$ is a uniform critical value that is equal to the smallest scalar $c>0$ that satisfies the inequality below: \[ \frac{1}{B}\sum_{b=1}^{B}1\big\{\sup_{(x_{1},x_{2},d)\in\mathcal{S}}\big|\frac{\hat{\alpha}_{b,n}(x_{1},x_{2},d)'\hat{\theta}_{b}-\tilde{\alpha}_{b,n}(x_{1},x_{2},d)'\tilde{\theta}_{b}}{\hat{\sigma}_{\Delta}(x_{1},x_{2},d)/\sqrt{n}}\big|\leq c\big\}\geq1-a \] $\hat{\sigma}_{\Delta}(x_{1},x_{2},d)/\sqrt{n}$ is an estimate of the pointwise variance of $\hat{\Delta}(x_1 ,x_2 ,d)$ and can itself be obtained by the pointwise variance of the bootstrap estimates of this quantity.
Uniform confidence bands are then given below. The test rejects if zero is not contained within the bands at some point. \[ \bigg[\hat{\Delta}(x_{1},x_{2},d)-\frac{\hat{\sigma}_{\Delta}(x_{1},x_{2},d)}{\sqrt{n}}\hat{c}_{1-a},\hat{\Delta}(x_{1},x_{2},d)+\frac{\hat{\sigma}_{\Delta}(x_{1},x_{2},d)}{\sqrt{n}}\hat{c}_{1-a}\bigg] \]
In this section we provide statistical guarantees for the empirical methods in Section 3. A notable departure from existing analyses of PSMD estimators is that we do not require that there exists a smooth solution to either of the conditional moment restrictions in Theorem 1 that motivate our estimation method. This is important in our setting because solutions $\gamma$ and $\varphi$ to the equations in Theorem 1 lack structural interpretations and need not be unique. This contrasts with analysis in NPIV settings in which case one solution to the relevant conditional moment restriction is the object of interest and has a clear structural interpretation.
A key ingredient in our asymptotic analysis is summarized in Theorem 4.1 below which establishes `well-posedness' of the estimation problem. Recall that Theorem 1.1 provides two alternative characterizations of the conditional average potential outcome based on the solutions to conditional moment restrictions. Theorem 4.1 shows that if the moment condition in 1.1.i has a solution, then the characterization of the object of interest in 1.1.ii is well-posed, and vice versa. For simplicity we omit additional confounders $D$ but the theorem extends straight-forwardly to incorporate such covariates (we prove this more general version in the appendix).
\newtheorem*{Th12}{Theorem 4.1 (Well-Posedness)} \begin{Th12} Suppose there are no additional covariates $D$ and Assumptions 1-3 hold for $x=x_1,x_2$, then:
i. If ((ref)) holds with $E[\varphi(Z)^{2}|X=x_1]^{1/2}\leq c<\infty$, then for any ${\gamma}$ with $E[\gamma(V)^{2}|X=x_1]<\infty$:
ii. If ((ref)) holds with $E[\gamma(V)^{2}|X=x_1]^{1/2}\leq c<\infty$, then for any ${\varphi}$ with $E[\varphi(Z)^{2}|X=x_1]<\infty$:
\end{Th12}
To understand the implications of this result, consider Theorem 4.1.i. Suppose we find a function $\gamma$ that solves an empirical analogue of the moment restriction ((ref)) and then estimate its conditional mean $E[\gamma(V)|X=x_2]$. $\gamma$ will not satisfy the population moment restriction exactly, rather there is some error which is measured by the expectation on the right hand side of the inequality in 4.1.i. Theorem 4.1.i states that the difference between $E[\gamma(V)|X=x_2]$ and the object of interest is no greater than a factor $c$ times this error. $c$ here is the norm of a solution $\varphi$ to the moment condition ((ref)) in Theorem 1.1.i. In sum, the error in our estimate of the object of interest cannot be much greater than the error in the population moment restriction.
In our case $\phi(v,x_1)'\hat{\theta}$ is our $\gamma(v)$ and $\hat{\alpha}_n(x_1,x_2)'\hat{\theta}$ estimates its conditional mean. Theorem 4.1 shows that for good estimation of the conditional average potential outcome, $\gamma$ need not be close to an exact solution to the moment condition $\gamma_0$ (recall there may not be a unique solution), instead we only need that $\gamma$ approximately satisfies the population moment restriction. We can achieve this under weaker conditions than would be required for consistent estimation of a solution $\gamma_0$.
In the mathematical literature `well-posedness' is usually understood, at least in part, to mean that the solutions to an integral equation is insensitive to perturbations in the function on the right-hand side. Consider Theorem 4.1.i and define a function $g(Z)$. Suppose $E[\gamma(V)|Z,X=x_1]=g(Z)$ almost surely. Note that this condition can be understood as an integral equation. Now consider another function $\tilde{g}$ and let $\tilde{\gamma}$ satisfy $E[\tilde{\gamma}(V)|Z,X=x_1]=\tilde{g}(Z)$. From the proof of Theorem 4.1.i we see that if $g$ and $\tilde{g}$ are close in the sense that $E[|g(Z)-\tilde{g}(Z)|^2|X=x_1]^{1/2}$ is small, then ${\gamma}$ and $\tilde{\gamma}$ are close in terms of the semi-norm $|E[\gamma(V)-\tilde{\gamma}(V)|X=x_1]|$. In other words, a small perturbation to the function $g$ has only a small effect on the solution $\gamma$.
Well-posedness in the sense summarized in Theorem 4.1 allows us to establish statistical guarantees without any restrictions on the `sieve measure of ill-posedness' (Chen2015). Estimation of the `structural function' in NPIV models is typically not well-posed. While a solution $\gamma_0$ to the restriction ((ref)) is analogous to a structural function in NPIV, our interest is not in a solution $\gamma_0$ itself (which may not be unique) but rather, a linear functional thereof. Severini2012 note that estimation of a linear functional of the structural function in NPIV may be well-posed, and indeed existence of a solution to ((ref)) implies that the relevant condition in Severini2012 holds in our setting.
We now state the assumptions under which we establish consistency of our estimator. It is helpful to introduce some additional notation. Recall that in Section 3 we use vectors of basis functions $\psi_{n,i}$, $\rho_{n,i}$, $\zeta_{n,i}$, and $\chi_{n,i}$. It is convenient to define the vector of product basis functions $\kappa_{n,i}=\zeta_{n,i}\otimes\chi_{n,i}$. We let $k_{\psi,n}$ be the length of $\psi_{n,i}$, $k_{\rho,n}$ the length of $\rho_{n,i}$, and similarly for $k_{\zeta,n}$, $k_{\chi,n}$, and $k_{\kappa,n}$. We define $\Sigma_{\psi,n}=E[\psi_{n,i}\psi_{n,i}']$, $\Sigma_{\rho,n}=E[\rho_{n,i}\rho_{n,i}']$, and similarly for the other basis functions. We define reduced-form residuals as follows:
Throughout we let `$\|\cdot\|$' denote the Euclidean norm of a vector and the Euclidean operator norm of a matrix. In addition, we define the semi-norms $\|\cdot\|_{L_2}$ and $\|\cdot\|_{n}$ as follows. If $\delta$ is a vector-valued function of $Z$, $X$, and $D$ then we let $\|\delta\|_{L_2}^2=E[\|\delta(Z_i,X_i,D_i)\|^2]$ and $\|\delta\|_{n}^2=\frac{1}{n}\sum_{i=1}^n\|\delta(Z_i,X_i,D_i)\|^2$. For a vector-valued function $\delta$ and vector $\beta$ of the same length, we let $\|\delta'\beta\|_{L_2}^2=E[\|\delta(Z_i,X_i,D_i)'\beta\|^2]$, and similarly for the norm $\|\cdot\|_n$.
In order to specify the approximation properties of our sieve basis functions we must introduce spaces of smooth functions. In particular, we let $\Lambda^k_s(c)$ denote the set of H{\"o}lder smoothness class $s$ functions on $\mathbb{R}^k$ with semi-norm at most $c$. For example, $\Lambda^1_1(c)$ is the set of Lipschitz continuous functions on the real line with Lipschitz constant at most $c$. A formal definition can be found in section B.3 of the appendix. We let $\dim(X,D)$ denote the sum of the lengths of vectors $X$ and $D$, and similarly for other collections of observables.
Finally, we let $\mu_n$ denote the smallest eigenvalue of $\Sigma_{\pi}=E[\pi_{n,i}\pi_{n,i}']$. In our setting the reciprocal of $\mu_n$ is analogous to the sieve measure of ill-posedness. As suggested above, we are able to guarantee consistency without placing any restrictions on this quantity, however it shows up in our inference results.
\theoremstyle{definition} \newtheorem*{A51}{Assumption 4.1 (First Stage)} \begin{A51} For sequences $r_{g,n}$, $r_{\pi,n}$, $\bar{r}_{\pi,n}$, and $\bar{r}_{\alpha,n}$ all $o(1)$, i. $\|\hat{g}-g\|_{n}=O_p(r_{g,n})$, ii. for any sequence nonrandom sequence $\beta_n$ with $\|\beta_n\|\leq 1$, $\|(\hat{\pi }_n-\pi_n)'\beta_n\|_{n}=O_p(r_{\pi,n})$ iii. $\|\hat{\pi}_n-\pi_n\|_{L_2}=O_p(\bar{r}_{\pi,n})$, and iv. $\sup_{(x_1,x_2, d)\in \mathcal{S}}\|\hat{\alpha}_n(x_1,x_2,d)-\alpha_n(x_1,x_2,d)\|=O_p(\bar{r}_{\alpha,n})$. \end{A51}
Assumption 4.1 assumes certain convergence rates for the first stage nonparametric regression estimates. Rates of convergence for nonparametric regression estimators can be found in the literature, for example in (Belloni2015). In Lemmas C.3, C.4, and C.5 in the appendix we derive convergence rates for the first-stage estimators detailed in Section 3 under primitive conditions.
\theoremstyle{definition} \newtheorem*{A52}{Assumption 4.2 (Bases)} \begin{A52} i. The eigenvalues of $\Sigma_{\psi,n}$, $\Sigma_{\zeta,n}$, $\Sigma_{\chi,n}$, $\Sigma_{\rho,n}$, $\Sigma_{\phi,n}$, and $\Sigma_{\kappa,n}$, are bounded above and below away from zero uniformly over $n$. ii. $\|\psi_{n,i}\|\leq \xi_{\psi,n}$, $\|\rho_{n,i}\|\leq \xi_{\rho,n}$, $\|\kappa_{n,i}\|\leq \xi_{\kappa,n}$, and $\|\chi_{n,i}\|\leq \xi_{\chi,n}$.
iii. For each $s>0$ there is a sequence $\ell_{\rho,n}(s)\to 0$ so that for any $\delta\in\Lambda_{s}^{\dim(V)}(c)$: \[ \inf_{\beta\in\mathbb{R}^{k(n)}}E\big[(\delta(V)-\rho_{n}(V)'\beta\big)^{2}\big]^{1/2}\leq c\ell_{\rho,n}(s) \]
iv. For each $s>0$ there is a sequence $\ell_{\psi,n}(s)\to 0$ so that for any $\delta\in\Lambda_{s}^{\dim(X,Z,D)}(c)$, letting $\beta$ minimize $E\big[(\delta(X,Z,D)-\psi_{n}(X,Z,D)'\beta\big)^{2}\big]$ we have for all $x$, $z$, and $d$: \[ \big|\delta(x,z,d)-\psi_{n}(x,z,d)'\beta\big|\leq c\ell_{\psi,n}(s) \] The same holds with $\zeta$ in place of $\psi$.
v. Either $X$ and $D$ have a finite discrete support and $\chi_{n}(X,D)$ is a vector of binary indicators (one for each possible value of X and D), or For each $s>0$ there is a sequence $\ell_{\chi,n}(s)\to 0$ so that for any $\delta\in\Lambda_{s}^{\dim(X,D)}(c)$, letting $\beta$ minimize $E\big[(\delta(X,D)-\chi_{n}(X,D)'\beta\big)^{2}\big]$ we have: \[ \big|\delta(x,d)-\chi_{n}(x,d)'\beta\big|\leq c\ell_{\chi,n}(s) \]
\end{A52} \theoremstyle{definition} \newtheorem*{A53}{Assumption 4.3 (Densities)} \begin{A53} i. $V$ admits a probability density conditional on each value of $Z$, $X$, and $D$, that is bounded above and below away from zero on the the support of $V$. ii. Either $X$ and $D$ have a finite discrete support or they have a joint density that is bounded above and below away from zero on their rectangular joint support. \end{A53} \theoremstyle{definition} \newtheorem*{A54}{Assumption 4.4 (Smoothness)} \begin{A54} With probability $1$, i. the function that maps $v$ to $ \frac{f_{V|XDZ}(v|X_i,D_i,Z_i)}{f_{V}(v)}$ is an element of $\Lambda_{s}^{\dim(V)}(c)$, ii.$\frac{f_{V|XDZ}(V_i|\cdot,\cdot,\cdot )}{f_{V}(V_i)}\in\Lambda_{s}^{\dim(Z,X,D)}(c)$, and iii. $g\in\Lambda_{s}^{\dim(Z,X,D)}(c)$.
\end{A54} \theoremstyle{definition} \newtheorem*{A55}{Assumption 4.5 (Sieve Growth and Penalty)} \begin{A55} i. $\frac{\xi_{\psi,n}^{2}ln(k_{\psi,n})}{n}\prec1$, ii. $\frac{\xi_{\chi,n}^{2}ln(k_{\chi,n})}{n}\prec1$, iii. $\frac{\xi_{\zeta,n}^{2}ln(k_{\zeta,n})}{n}\prec1$, iv. $\frac{\xi_{{\kappa},n}^{2}ln(k_{\kappa,n})}{n}\prec1$
\end{A55}
\newtheorem*{A46}{Assumption 4.6 (Existence)} \begin{A46} i. There is a finite constant $c$ so that for each $x$ and $d$ in the joint support of $X$ and $D$, ((ref)) has a solution $\gamma(\cdot,x,d)$ with $E[\gamma(Z,x,d)^{2}|X=x,D=d]\leq c$. ii. There is a finite constant $c$ so that for each $(x_1,x_2,d)\in\mathcal{S}$, ((ref)) has a solution $\varphi(\cdot,x_1,x_2,d)$ so that $E[\varphi(Z,x_1,x_2,d)^{2}|X=x_1,D=d]\leq c$.
\end{A46} Assumption 4.2.i and 4.2.ii are standard. 4.2.i can be established under primitive conditions (see e.g., Belloni2015). The rates in 4.2.ii are readily available in the literature for most basis functions used in practice. For many popular bases the supremum of the norm of the vector of basis functions grows at the same rate as the square-root of the number of functions, so for example $\xi_{\psi,n}\precsim\sqrt{k_{\psi,n}}$ (again, see Belloni2015). Assumptions 4.2.iii-4.2.v specify the rate at which the basis functions can approximate smooth functions. Precise bounds for particular basis functions can be found in the approximation literature (see, for example, DeVore1993). For many commonly used bases we would have $\ell_{\psi,n}(s)\precsim k_{\psi,n}^{-s/\dim(Z,X,D)}$ and similarly for $\ell_{\rho,n}$ and $\ell_{\chi,n}$ with the number of basis functions and dimensions adjusted accordingly.
Assumption 4.3 is a standard regularity condition on joint probability densities. Assumption 4.4 imposes that some reduced-form objects be smooth so that we can approximate them using sieve basis functions. Smoothness of $g$ in its $Z$ argument follows from 4.5.i and the existence of a $\gamma$ that satisfies the conditions of Theorem 1.1.ii. Assumption 4.5 restricts the rate at which the numbers of basis functions can grow. This assumption allows us to apply Rudelson's matrix law of large numbers (Rudelson1999).
Assumption 4.6 states that the conditional moment restrictions in Theorem 1.1 admit solutions for various choices of $x_1$, $x_2$, and $d$, and moreover, that the norms of these solutions are uniformly bounded. We provide primitive conditions and further discussion in Appendix A.
The assumptions do not impose smoothness on any solutions $\gamma$ or $\varphi$ to the moment conditions in Theorem 1.1. This is important because these functions are not unique and lack a clear structural interpretation. This presents a challenge because, in effect, we find an approximate solution $\gamma$ that is a linear combination of the basis functions $\phi_n$. However, for consistency we only need that $g_i$ is well approximated by a linear combination of the components of $\pi_{n,i}$. Lemma 4.1 below shows that such an approximation result is attainable without imposing smoothness on $\gamma$ or $\varphi$.
\theoremstyle{plain} \newtheorem*{L39}{Lemma 4.1} \begin{L39} Suppose that for each $x$ and $d$ in the joint support of $X$ and $D$, Assumptions 1-3 hold and Assumptions 4.2, 4.3, 4.4, and 4.6.i hold. Then there is a sequence $\{\theta_{n}\}_{n=1}^{\infty}$ with $\|\theta_{n}\|\precsim1$ and a finite constant $c$ so that: \[ |g(z,x,d)-\pi_{n}(z,x,d)'\theta_{n}|\leq cr_{\theta} \] Where $r_{\theta,n}=\ell_{\rho,n}(s)$ if $X$ and $D$ have finite discrete support and otherwise $r_{\theta,n}=\xi_{\chi,n}\ell_{\rho,n}(s)+\ell_{\chi,n}(1)^{-\bar{s}}$ with $\bar{s}=\frac{\min\{s,1\}}{\min\{s,1\}+1}$. \end{L39}
Theorem 4.2 establishes a rate of convergence for the estimator in terms of the first stage convergence rates. This rate does not depend on any sieve measure of ill-posedness and suggests the estimator is consistent under fairly weak restrictions on the penalty $\lambda_0$.
\theoremstyle{plain} \newtheorem*{Th31}{Theorem 4.2 (Convergence)} \begin{Th31} Suppose the conditions of Lemma 4.1 hold along with Assumptions 4.5, and 4.6.ii. Assume $k_{\chi,n}\ell_{\zeta}(s)^{2}\precsim\bar{r}_{\pi,n}^{2}$ and $\bar{r}_{\pi,n}^{2}/\lambda_{0}\to0$. Then for any first-stage estimators $\hat{g}_{i}$, and $\hat{\alpha}_{n,i}$, and any estimator $\hat{\pi}_{n,i}$ of the form $\hat{\pi}_{n,i}=\hat{\beta}_{\pi}'\kappa_{n,i}$ for some $\hat{\beta}_{\pi}$, we have:
Where $r_{\theta,n}$ is as defined in Lemma 4.1. \end{Th31}
The rate in Theorem 4.2 suggests consistency is achieved so long as $\lambda_0$ converges to zero but not too quickly compared to the convergence rates of the first-stage nonparametric regression estimates. The condition that $\bar{r}_{\pi,n}^{2}/\lambda_{0}\to0$ suggests that $\lambda_0$ must shrink strictly more slowly than $n^{-1}$. We can choose $\lambda_0$ to optimize the rate in Theorem 4.2 and this yields the rate below. \[ \bar{r}_{\pi,n}\bar{\xi}_{\kappa,n}^{1/2}+r_{g,n}+r_{\theta,n}+\bar{r}_{\alpha,n}+r_{\pi,n} \]
Under standard conditions, nonparametric regression estimates converge more slowly than $n^{-1/2}$. Thus the rate in Theorem 4.2 is strictly slower than the parametric rate, even in the case of discrete $X$ and $D$. Theorem 4.3 in the next subsection shows that for certain choices of first-stage estimators it is possible to refine the rate in Theorem 4.3. This requires that the first-stage estimates are `under-smoothed', that is, the first-stage sieve dimensions are chosen so that the bias decreases strictly more quickly than the variance. In addition, to refine the rates we utilize restrictions on $\mu_n$ which corresponds to a sieve measure of ill-posedness.
We now establish asymptotic normality of our estimator and validity of the bootstrap inference procedure in Section 3.2. The results in this subsection pertain to the version of our estimator that uses the first-stage series ridge regressions as specified in Section 3.1. In the case of continuous treatments, the relevant notion of asymptotic normality is approximation by a sequence of Gaussian processes.
In addition, we provide potentially tighter convergence rates than those in Theorem 4.2 when the first-stage is carried out using series ridge regression. By restricting our attention to this particular choice of first-stage regression method, we can carefully disentangle the first-stage estimation errors into zero-mean components and bias. With sufficient under-smoothing, the first-stage bias becomes second-order, which allows for faster convergence of the second-stage estimates.
We use Assumption 4.7 below to establish asymptotic linearity of our estimator. To be more precise, after appropriate re-scaling, the condition helps to ensure that the estimation error is approximately equal to a re-scaled sample average of zero-mean iid random variables. This requires stronger restrictions on the rates at which the penalty parameter $\lambda_{0}$ goes to zero and the rate at which the dimension of the sieve spaces grow.
Assumption 4.7 refers to quantities $\bar{l}_{\alpha,n}$, $\bar{l}_{g,n}$, and $\bar{l}_{\pi,n}$, which are the rates of the linearization errors for the first-stage estimates $\hat{\alpha}_{n}$, $\hat{g}_{i}$, and $\hat{\pi}_{n,i}$ respectively. For more details on what these quantities represent see Lemmas C.3, C.4, and C.5 in the appendix. In these lemmas we establish the rates below under Assumptions 4.2, 4.4, and 4.5:
Assumption 4.7 also depends on rates of convergence $\bar{r}_{\alpha,n}$ and $\bar{r}_{\pi,n}$ which satisfy Assumptions 4.1.iii and 4.1.iv. In Lemmas C.4 and C.5 in the appendix we provide formulas for $\bar{r}_{\alpha,n}$ and $\bar{r}_{\pi,n}$ in terms of model primitives under Assumptions 4.2, 4.4, and 4.5. These rates could perhaps be tightened under additional smoothness restrictions on the basis functions using similar analytical techniques to those in Belloni2015.
\newtheorem*{A56}{Assumption 4.7 (Linearization)} \begin{A56}
i. $E[\epsilon_{i}^{2}|X_{i},D_{i},Z_{i}]$ is almost surely bounded above and below away from zero and likewise for $\|E[\upsilon_{n,i}\upsilon_{n,i}'|Z_{i},X_{i},D_{i}]\|$ and $\|E[u_{n,i}u_{n,i}'|X_{i},D_{i}]\|$ uniformly over $n$. ii. $\sqrt{\frac{\xi_{\kappa,n}^{2}(k_{\zeta,n}+k_{\psi,n})ln(k_{\kappa,n})}{n}}\leq r_{s,n}$ and $\sqrt{\frac{k_{\rho,n}k_{\chi,n}(\xi_{\kappa,n}^{2}k_{\zeta,n}+\xi_{\psi,n}^{2}k_{\psi,n})}{n}}\leq r_{s,n}$ where $r_{s,n}\prec0$, iii. $r_{l,n}=\sqrt{ln(n)}(\bar{l}_{\alpha,n}+\bar{l}_{g,n}+\bar{l}_{\pi,n})\prec1$, iv. $\sqrt{ln(n)n}r_{\theta,n}\prec1$ v. $r_{\lambda,n}=\bar{\xi}_{\kappa,n}\sqrt{\frac{\lambda_{0}n}{1+\mu_{n}/\lambda_{0}^{2}}}\prec1$, and vi. $r_{\mu,n}=\sqrt{\frac{(\bar{r}_{\alpha,n}^{2}+\bar{\xi}_{\kappa,n}^{2}\bar{r}_{\pi,n}^{2})(k_{\psi,n}+k_{\zeta,n})/\lambda_{0}}{1+\mu_{n}/\lambda_{0}^{2}}}\prec1$. \end{A56}
Assumption 4.7.i is standard. Assumption 4.6.ii limits the rate at which the number of basis functions may grow with the sample size. This condition ensures that some sample second moment matrices converge sufficiently quickly to their population counterparts. 4.7.iii requires that the linearization errors in the first-stage estimates shrink sufficiently quickly. This condition necessarily requires that the bias in the first stage estimates converges faster than $n^{-1/2}$ which means the first-stage estimates are under-smoothed. The $\sqrt{ln(n)}$ can be dropped if we only require a Gaussian approximation for our estimator and not for its bootstrap analogue. Assumption 4.7.iv ensures the second-stage sieve approximation error is negligible. Again, the $\sqrt{ln(n)}$ is only required for Gaussian approximation of the bootstrap estimator. Assumption 4.7.v requires that $\lambda_{0}$ goes to zero sufficiently quickly with the sample size so that the asymptotic bias due to regularization in the second stage is negligible.
Assumption 4.7.vi is more complex. It involves $\lambda_{0}$ as well as rates $\bar{r}_{\alpha,n}$ and $\bar{r}_{\pi,n}$. To see why we require Assumption 4.7.vi, recall that $\hat{\theta}$ in the second stage of our procedure is defined by $\hat{\theta}=(\hat{\Sigma}_{\hat{\pi}}+\lambda_0 I)^{-1}\frac{1}{n}\sum_{i=1}^{n}\hat{\pi}_{n,i}\hat{g}_{i}$. In order to linearize our estimator we must approximate $(\hat{\Sigma}_{\hat{\pi}}+\lambda_0 I)^{-1}$ with its population counterpart $(\Sigma_{\pi,n}+\lambda_{0}I)^{-1}$. The resulting approximation error is small when $\bar{r}_{\pi,n}$ is small, i.e., when $\hat{\pi}_{n,i}$ is close to $\pi_{n,i}$ in an appropriate sense. This approximation error also depends on both $\mu_{n}$ and $\lambda_{0}$, because for small $\mu_{n}$ and $\lambda_{0}$ the matrix $\Sigma_{\pi,n}+\lambda_{0}I$ is close to singular.
For both 4.7.v and 4.7.vi to hold, $\mu_{n}$ cannot go to zero too rapidly. If $\mu_{n}$ goes sufficiently quickly to zero then there is no sequence of penalty parameters $\lambda_{0}$ under which both conditions hold. The following condition on $\mu_{n}$ is sufficient to ensure a sequence of penalty parameters exists so that 4.7.v and 4.7.vi are both satisfied: \[ \frac{\mu_{n}^{2}}{(\bar{r}_{\alpha,n}^{2}+\bar{r}_{\pi,n}^{2})^{3}(k_{\psi,n}+k_{\zeta,n})^{3}/n}\to\infty \]
The rate at which $\mu_{n}$ goes to zero depends on the number of basis functions in $\rho_{n}$. To ensure $\mu_{n}$ does not shrink too quickly to zero, the dimension of $\rho_{n}$ must not grow too rapidly. The rate at which $\mu_{n}$ decreases with $\rho_{n}$ is akin to a restriction on the sieve measure of ill-posedness. Thus there is a marked distinction between our asymptotic inference results and the consistency results in Theorem 4.2. For consistency we do not require any restriction on the rate at which $\mu_{n}$ goes to zero and thus we need not place any restrictions on a sieve measure of ill-posedness. The reason for this difference is that consistency only requires $n^{-1/2}r_{\lambda,n}\prec0$ and $n^{-1/2}r_{\mu,n}\prec0$, which can hold even if we set $\mu_{n}=0$ in the formulas for $r_{\lambda,n}$ and $r_{\mu,n}$. We also require 4.7.v and 4.7.vi to achieve an $n^{-1/2}$-rate of convergence for our estimator in the case of discrete $X$ and $D$.
Finally, the following assumptions allow us to apply Yuriskii's coupling (see e.g., Pollard2001 or Belloni2015) to achieve uniform approximations of both our estimator and its bootstrap analogue by Gaussian processes.
\newtheorem*{A57}{Assumption 4.8 (Normality)} \begin{A57}
i. $E[\epsilon_{i}^{3}|x_{i},d_{i},z_{i}]$ is bounded above, ii. \[\frac{(k_{\psi,n}+k_{\zeta,n}+k_{\chi,n})(k_{\psi,n}\xi_{\psi}+k_{\zeta,n}\xi_{\zeta,n}\xi_{\rho,n}+k_{\chi,n}\xi_{\chi,n}\xi_{\rho,n})}{n^{1/2}/ln(n)}\prec1\] iii. $\sqrt{ln(n)}\|\psi_{n,i}\|\leq\xi_{\psi,n}$, $\sqrt{ln(n)}\|\rho_{n,i}\|\leq\xi_{\rho,n}$, $\sqrt{ln(n)}\|\kappa_{n,i}\|\leq\xi_{\kappa,n}$, and $\sqrt{ln(n)}\|\chi_{n,i}\|\leq\xi_{\chi,n}$.
\end{A57}
Assumption 4.8.i imposes a bound on conditional third moments. This is a standard condition that allows us to apply an appropriate central limit theorem under growing dimensions. Assumption 4.8.ii restricts the rate at which the sieve spaces grow with the sample size. This condition ensures that the third moment of a vector of zero-mean random variables, multiplied by the length of the vector, grows slowly enough so that we may apply Yuriskii's coupling. 4.8.iii strengthens 4.2.ii so that our earlier linearization arguments extend to the bootstrap estimator. It ensures that, not only does $\xi_{\psi,n}$ upper-bound the essential supremum of $\|\psi_{n,i}\|$, but that $\ensuremath{\max}_{1\leq i\leq n}\sqrt{Q_{b,i}}\psi_{n,i}\precsim_{p}\xi_{\psi,n}$ which allows us to linearize the bootstrap analogue of the estimator. This follows similar arguments to those in Belloni2015.
To state the theorem below we define a vector-valued function $s_{n}$. First define a zero-mean random vector $\eta_{n,i}$ as follows:
\[ \eta_{n,i}=
\]
In the above, $\Sigma_{\pi,\lambda_{0}}=\Sigma_{\pi,n}+\lambda_0 I$. $\tilde{\theta}_{n}$ is a $k_{\rho,n}$-by-$k_{\chi,n}$ matrix obtained by partitioning the vector $\theta_{n}$ that satisfies Lemma 4.1 into $k_{\rho,n}$ contiguous length-$k_{\chi,n}$ subvectors and stacking the transposes into a matrix. Let $\Omega_{n}=E[\eta_{n,i}\eta_{n,i}']$ and define a vector-valued function $s_{n}$ by: \[ s_{n}(x_{1},x_{2},d)=\Omega_{n}^{1/2}
\]
Theorem 4.3 below shows that the distribution of the estimation error is approximately equal to that of a Gaussian process whose pointwise variance is $\|s_{n}(x_{1},x_{2},d)\|^{2}/n$. Moreover, the difference between the bootstrap estimator and the original estimator can be approximated by an identical Gaussian process that is independent of the data (but not the bootstrap weights).
\theoremstyle{plain} \newtheorem*{Th33}{Theorem 4.3 (Asymptotic Normality)} \begin{Th33} Suppose Assumptions 1-3 and 4.2-4.8.ii all hold and the eigenvalues of $\Omega_{n}$ are bounded below away from zero uniformly over $n$.
Then $\|s_{n}(x_{1},x_{2},d)\|\precsim1+\bar{\xi}_{\kappa,n}+\xi_{\chi,n}^{2}$ and there is a sequence of mean-zero Gaussian random vectors $\text{\ensuremath{\mathcal{N}_{n}}}$ with identity variance covariance matrices so that:
Moreover, if 4.8.iii also holds, there exists a sequence of zero-mean Gaussian random vectors $\mathcal{N}_{n}$ with identity variance-covariance matrix which are independent of the data and:
}
\end{Th33}
Theorem 4.3 shows that the distribution of the estimation error can be well-approximated by a sequence of Gaussian processes. The theorem also provides a convergence rate for $\|s_{n}(x_{1},x_{2},d)\|$ and thus a rate for the pointwise standard error of the approximating Gaussian process. In fact, this rate, combined with the Gaussian approximation, implies that the estimator converges to the truth at rate $(1+\bar{\xi}_{\kappa,n}+\xi_{\chi,n}^{2})/\sqrt{n}$. If $X$ and $D$ have finite support then $\chi_n$ can be chosen so that both $\bar{\xi}_{\kappa,n}$ and $\xi_{\chi,n}$ are bounded and thus the estimator converges at the parametric rate. As discussed above, this relies crucially on under-smoothing in the first-stage regressions.
In addition, the theorem shows that the distribution of the estimation error can be approximated by its bootstrap analogue, that is, by the distribution of $\hat{\alpha}_{n}(x_{1},x_{2},d)'\hat{\theta}_{b}-\alpha_{n}(x_{1},x_{2},d)'\hat{\theta}$ where the data (but not the bootstrap weights) are treated as fixed. Validity of the bootstrap confidence bands in Section 3.2 follows immediately if the set $\mathcal{S}$ is finite. If $\mathcal{S}$ is infinite then asymptotically correct uniform coverage requires additional anti-concentration results on the suprema of Gaussian processes. Such results can be found in Chernozhukov2014.
Theorem 4.3 extends straight-forwardly to linear combinations of different estimators that all satisfy the relevant conditions. The only condition that must be strengthened is 4.8.ii in order to reflect the fact that the relevant normal random vector has greater dimension. Thus the theorem also justifies the specification test described in Section 3.2.
We apply our methodology to real data. In order to emphasize the applicability of our approach to both cross-sectional and panel models we present two separate empirical settings. In the first application we exploit the panel structure of the data to form proxies as suggested in Section 2. In our second application we use cross-sectional variation to estimate causal effects
A household's Engel curve for a particular class of good captures the relationship between the share of the household's budget spent on that class and the total expenditure of the household. An Engel curve is `structural' if it captures the effect of an exogenous change in total expenditure. Imagine an ideal experiment in which the household's total expenditure is chosen by a researcher using a random number generator and the household then chooses how to allocate that total expenditure between different classes of goods. Then the resulting relationship between the total expenditure and budget share is a structural Engel curve.
Nonparametric regression of the budget share spent on food and the total expenditure on certain classes of goods is unlikely to represent the average structural Engel curve. This is because total expenditure is chosen by the household and thus depends upon the household's underlying financial assets and consumption preferences. Household finances and preferences partially determine the share of expenditure that the household allocates to food.
We estimate average structural Engel curves for food eaten at home using data from the Panel Study of Income Dynamics (PSID). The PSID study follows US households over a number years and record expenditure on various classes of goods. We use ten periods of data from the surveys carried out every two years between 1999 and 2017. We drop all households whose household heads are not married or cohabiting and drop all households for which we lack the full ten periods of data, leaving us with $840$ households. We take as the total expenditure the sums of expenditures on food (both at home and away from home), housing, utilities, transportation, education, childcare and health-care.
We apply the approach to identification with fixed-$T$ panels described in Section 2.2, allowing for time-varying preferences and unobserved assets. In this setting the condition that outcomes do not feed back to future confounders nor treatments seems reasonable. While total expenditure on non-durables directly impacts household assets, the proportion of this that is allocated to food does not.
Let $X_{t}$ denote the total expenditure on non-durables in period $t$ and $Y_t$ the share of this expenditure allocated to food. We have $10$ periods of data and aim to estimate causal effects at time $t=4$. In line with Section 2.2, we set $V=Y_{[1:3]}$ and $Z=X_{[5:10]}$. We take $D=X_{[1:3]}$ which allows for the possibility that total expenditure in some period can directly impact expenditure at most three periods ahead and preferences/financial assets up to four periods ahead.
We apply our method using the first-stage series ridge regression procedures specified in Section 3.3. The vectors of basis functions include all squares and interactions of the variables. First-stage penalty parameters are set to zero and for the second stage penalty $\lambda_0$ we simply use the square root of the sample size.
Figure 4.a plots our nonparametric estimate of the average structural Engel curve for food. The figure shows a downward-sloping Engel curve that (with a log scale for total expenditure) is subtly concave. The downward slope of the curve suggests that food is a normal good, at least in aggregate.
The choice of $D$ used for Figure 4.a may be overly conservative. The results in Section 2.3 suggest that under stronger restrictions on the dynamics of the data-generating process (namely $\min\{\kappa_XX,\kappa_{WX}-1\}\leq 1$), we may set $D=X_3$. To test these stronger restrictions we apply the Hausman specification method proposed in Section 3.2.3. The results are shown in Figure 4.b. The uniform confidence bands for the difference between the two alternative estimates contains zero at all values of total expenditure and so we fail to reject these restrictions at the $95\%$ level.
Fruehwirth2016 examine the causal effect of being made to repeat a particular grade level on the cognitive development of US students. They use data from the ECLS-K panel study which contains panel data on the early cognitive development of US children. We use our methods to examine the effect of grade retention on the cognitive outcomes of children in the 1998-1999 kindergarten school year using cleaned data available with their paper. Following Fruehwirth2016, we take our outcome variables to be the tests scores in reading and math when aged approximately eleven. Also in line with Fruehwirth2016 our treatments are indicators for retention in kindergarten, `early' (in first or second grade) and `late' (in third or fourth grade). The cleaned data from Fruehwirth2016 contains only students who are retained at most once in the sample period and no students who skip a grade.
The ECLS-K dataset contains scores that measure a student's behavioral and social skills and their scores on a range of cognitive tests at different ages. To account for the confounding effect of unmeasured ability, Fruehwirth2016 estimate a latent factor model with a particular structure. They assume that all confounding between grade retention and potential future cognitive test scores is due entirely to the presence of three latent factors representing different dimensions of ability. Fruehwirth2016 then use test scores to recover the distribution of the latent factors and their loadings. They assume a particular multiplicative structure between the factors (which are time-invariant) and time-specific factor loadings in both their outcome and selection equations.
Like Fruehwirth2016 we wish to use test-scores to adjust for some latent ability to perform well academically $W$. Our methods allow us to avoid any strong assumptions on the factor structure. We take the treatment to be a three-dimensional vector of binary indicators that signify whether student is held back in kindergarten, early elementary, or late elementary. Note that by defining $X$ in this way we estimate the effects of all three treatments simultaneously.
We let the set of proxies $V$ contain the student's scores tests in kindergarten and $Z$ contains the test scores from early in elementary school (first or second grade). Some children are held back a grade in kindergarten and therefore treatment (being held back a grade) may have a causal effect on the proxies in $Z$. This is compatible with Assumption 1 (see Figure 1.b).
In the ECLS-K dataset that we use for our empirical analysis the cognitive and behavioral scores are not shared with the students, parents nor teachers, and thus they should not determine the decision to retain a child.\footnote{This was confirmed by email with the ECLS study director.}, hence there is no causal effect of $V$ on treatments $X$ nor any causal effect of $V$ on $Y$.
Our sample contains of $1,951$ individuals. $V$ contain the student's scores on three cognitive and three behavioral tests in kindergarten and $Z$ contain the scores on these tests from early elementary. Thus we have six proxies in each group. We understand to $W$ capture both underlying cognitive and behavioral characteristics. Given the discussion of Assumption 2, for identification to be plausible we need to believe that there are at most six such characteristics in all.
In this setting the full support condition (Assumption 3) requires the children of all underlying latent cognitive skills and behavioral characteristics have some probability of being held back. To support this assumption consider that children may be held back a year due to some disruption in their family life, because they are among the youngest of their peers, due to a period of illness, and any other of many reasons that are not likely to be very strongly correlated with underlying cognitive and behavioral skills. Note that no student is held back twice we understand the support of $X$ to be equal to the set of three binary indicators that sum up to weakly less than one.
The ECLS-K data also contain a number of additional observables which we include as controls $D$. These include vectors that capture family characteristics, features of each student's school, student charateristics, and the student's age.
Table 2 compares estimates of the average effect of treatment on the treated under different approaches to estimation. In all cases we control for the additional observables $D$ described above. In the first column are linear least-squares estimates of the average treatment effects when we include $D$ and treatments as regressors but no other variables. The estimated effects are all strongly negative and very statistically significant. In the second column, the kindergarten cognitive and behavioral scores are included along with $D$ as regressors in an additive linear specification, note that in every case the estimated negative effects are at least halved in magnitude compared to the case in which the scores are not included.
In the third column we apply our method with linear specifications (i.e., $\rho_n$, $\psi_n$, $\zeta_n$, and $\chi_n$ return their arguments) and with the first-stage penalty parameters set to zero and the second stage penalty set to $\sqrt{n}$. In every case the the estimated ATTs less strongly negative than they are when we simply control for kindergarten scores. Finally, the last column contains the ATTs from the series ridge method specified in Section 3. The vectors of basis functions include squares and interactions of variables. Penalty parameters are set to zero for the first-stage estimates and the second stage penalty $\lambda_0$ is set to $\sqrt{n}$. Compared to the linear case the ATTs for reading are less strongly negative, the only exception being for math scores of children retained in kindergarten.
The results in Table 2 are consistent with the notion that unmeasured ability biases the estimated ATTs downwards. Including kindergarten test scores as controls mitigates some of this bias, and the proxy controls method mitigates the bias further still, resulting in mostly less negative estimated ATTs.