Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,522 characters · 12 sections · 87 citation commands
Nonparametric Instrumental Variables Estimation Under Misspecification
Instrumental validity can be difficult to defend. The assumption rules out not only confounding between outcomes and instruments, but also any direct causal effect of the instruments on outcomes. Thus instruments may not be valid even if they are assigned by an ideal randomized experiment. To justify the use of instrumental variables (IV) methods, applied researchers usually argue not that instrumental validity holds exactly, but that any deviation from validity is small.\footnote{See Conley2008 for further discussion.} `Small' does not imply `inconsequential', and so it is important to assess the sensitivity of IV methods to a modest deviation from full instrumental validity.
In this work we consider the sensitivity of Nonparametric Instrumental Variables (NPIV) analysis (Newey2003, Ai2003) to misspecification, particularly that which results from invalid instruments. NPIV generalizes the linear IV model to allow for a flexible, non-linear `structural function'. The structural function is the key object of interest and typically has a causal interpretation. Identification and estimation is based on a conditional moment restriction. If instruments are invalid then the moment condition generally does not hold.
We show that identification in NPIV models is fragile. If only weak a priori conditions are placed on the structural function, then the identified set can be very large under only a slight relaxation of the NPIV moment condition. Estimation in NPIV models can be highly non-robust to misspecification. For estimators that impose only weak restrictions on the structural function, an arbitrarily small deviation from instrumental validity can impart a large asymptotic bias. This sensitivity is a consequence of the `ill-posedness' of the NPIV moment restriction.
If researchers place sufficiently strong restrictions (typically smoothness conditions) on the structural function, then robust identification is possible. Estimation that is robust to misspecification can be achieved by constraining the estimates so that they satisfy such smoothness conditions. In both cases, the sensitivity to misspecification is still greater in a certain sense than for standard nonparametric regression or parametric IV. Imposing strong smoothness restrictions in estimation carries a cost. If the true structural function does not obey these restrictions then an estimator that imposes them is necessarily inconsistent.
Two NPIV estimation methods which impose sufficient smoothness conditions are the procedures of Newey2003 and Blundell2007. A number of prominent NPIV procedures do not constrain the estimates to be smooth. For example, the methods described in Chen2018, Darolles, Hall2005, and Horowitz2011.
Robust estimation is also possible if the researcher is interested in a continuous functional of the structural function rather than the function itself. We provide necessary and sufficient conditions under which such a functional can be estimated robustly without the need to impose smoothness.
Our results suggest that researchers face a trade-off: impose too little smoothness and NPIV estimates are non-robust to invalid instruments, impose too much smoothness and the estimates are inconsistent even if instruments are valid. In response, we supplement NPIV estimation with a method for empirical sensitivity analysis. Methods for sensitivity analysis are increasingly popular in applied economics and include local methods (e.g., Andrews2017a, Armstrong, Weidner), and global methods (e.g., Altonji2005a, Masten2018, Oster2019)). We develop a technique for global sensitivity analysis in which the researcher estimates the identified set under relaxations of the NPIV moment condition and various smoothness assumptions. Thus researchers can assess which of their findings are robust to some misspecification of the NPIV moment condition and examine the trade-off between robustness and smoothness. Estimation of the identified set under weakened assumptions is used elsewhere for sensitivity analysis, for example in Masten2018. The set estimation problem is well-posed, and we derive convergence rates under primitive conditions.
We apply our methods to the empirical setting shared by Blundell2007 and Horowitz2011 who use NPIV methods to estimate shape-invariant Engel curves using data from The British Family Expenditure Survey. We argue that in this setting instruments are unlikely to be fully exogenous, in which case the NPIV moment restriction will not hold exactly. We use our methodology to assess which features of the structural Engel curve for food can be inferred robustly.
For a random variable $W$, the corresponding calligraphic letter $\mathcal{W}$ denotes its support and $\mu_W$ its probability distribution. If $W$ and $V$ are random variables then statements of the form $W=V$ or $W\leq V$ should be understood to hold with probability $1$. If $W$ is a random variable and $h$ a function on $\mathcal{W}$ then $|h|_\infty$ is the essential supremum of $|h(W)|$, we refer to this as the `sup norm'. If $\mathcal{S}$ is a subset of a topological space then $\bar{\mathcal{S}}$ is its closure and $int(\mathcal{S})$ its interior. For an operator $\mathbb{T}:\,\mathcal{B}_1\to\mathcal{B}_2$, if $\mathcal{S}\subseteq\mathcal{B}_1$ then $\mathbb{T}[\mathcal{S}]$ is the image of $\mathcal{S}$ under $\mathbb{T}$ and if $\mathcal{S}\subseteq\mathcal{B}_2$ then $\mathbb{T}^{-1}[\mathcal{S}]$ is the pre-image of $\mathcal{S}$. If $\mathcal{B}$ is a normed space then we denote its norm by $\|\cdot\|_\mathcal{B}$. If $\mathcal{B}$ is a set of functions with domain $\mathcal{W}$ then $\mathcal{B}^{d}$ contains all length-$d$ vectors of functions $r$ such that $r(w)=(h_1(w),h_2(w),...,h_d(w))$ for all $w\in\mathcal{W}$, where $h_1,h_2,...,h_d\in\mathcal{B}$. For a real vector $v$ we let $\|v\|_p$ denote the $\ell_p$ norm of $v$.
NPIV estimation is a flexible, nonparametric alternative to linear IV. It has been studied extensively, for example in Newey2003, Ai2003, Cheng, Darolles, Hall2005, Horowitz2011, and Chen2018. NPIV methods have been applied in empirical work by Blundell2007, Chen2018, and Horowitz2011 among others.
NPIV models relax the linearity assumption of linear IV but retain additive separability. Let $Y$ be an outcome of interest, $X$ an endogenous variable, and $Z$ an instrument. NPIV analysis assumes $Y=h_0(X)+\varepsilon$ where $h_0$ is a non-random `structural function' and the structural residual $\varepsilon$ satisfies $E[\varepsilon|Z]=0$. NPIV is `nonparametric' because the structural function is not assumed to take any particular parametric form.
The structural function $h_0$ is the key object of interest and typically has a causal interpretation. The use of instrumental variables is motivated by the policy-relevance of $h_0$. If $\varepsilon$ is correlated with $X$ then standard nonparametric regression will not recover $h_0$ but rather a reduced-form parameter which may not be useful for analyzing the effects of a policy.
We can obtain an NPIV model by imposing an additive structure for potential outcomes and assuming instruments are valid in the sense of Angrist. In particular, instruments are valid if they have no direct causal effect on outcomes and they are randomly assigned. Let $Y_{x,z}$ be the potential outcome from a counterfactual treatment level $x$ and counterfactual value of the instrument $z$. An additive model for $Y_{x,z}$ is given below.
The random variable $\varepsilon$ captures all heterogeneity in potential outcomes and without loss of generality $E[\varepsilon]=0$. The instrument has no direct effect on outcomes and so $Y_{x,z}$ does not depend on $z$, that is, $Y_{x,z}=Y_{x}$. If the instrument is randomly assigned then $E[\varepsilon|Z]=0$. In this case the NPIV model holds with $h_0$ the average potential outcome from treatment level $x$, that is $h_0(x)=E[Y_{x}]$.\footnote{In fact, we can allow for non-additive heterogeneity in potential outcomes so long as this heterogeneity is fully exogenous. Formally, suppose $Y_{x,z}=l_0(x,V)+\varepsilon$ where $l_0$ is a non-random function, $V\perp \!\!\! \perp(X,Z)$, and as before $E[\varepsilon|Z]=0$. Then the NPIV model holds with $h_0(x)=E[Y_{x}]$.}
Assumption 1.1 below formally states the key NPIV identifying moment restriction. It is equivalent to the model $Y=h_0(X)+\varepsilon$ with $E[\varepsilon|Z]=0$.
\theoremstyle{definition} \newtheorem*{A1.1}{Assumption 1.1 (Correct Specification)} \begin{A1.1} $E[Y-h_{0}(X)|Z]=0$ \end{A1.1}
We sometimes wish to place a priori restrictions on the structural function. To achieve this we can assume $h_0$ belongs to a restrictive subset $\mathcal{H}$ of the underlying space of functions $\mathcal{B}_X$.
\theoremstyle{definition} \newtheorem*{A1.2}{Assumption 1.2 (Parameter Space)} \begin{A1.2} $h_0\in \mathcal{H} \subseteq \mathcal{B}_X$. \end{A1.2}
Assumptions 1.1 and 1.2 are structural assumptions in the sense that one cannot confirm that they hold by looking at the joint distribution of the observables. The non-structural Assumptions 1.3 and 1.4 below are imposed throughout the NPIV literature. They restrict the joint distribution of $X$ and $Z$.
\theoremstyle{definition} \newtheorem*{A1.3}{Assumption 1.3 (Completeness)} \begin{A1.3} For any $h\in\mathcal{B}_X$, $E[h(X)|Z]=0$ if and only if $h(X)=0$. \end{A1.3} \newtheorem*{A1.4}{Assumption 1.4 (Compactness)} \begin{A1.4} Define the linear operator $\mathbb{A}$ by the equation $\mathbb{A}[h](Z)=E[h(X)|Z]$. $\mathbb{A}$ is a compact infinite-dimensional linear operator from the Banach space $\mathcal{B}_{X}$ into a Banach space $\mathcal{B}_{Z}$. \end{A1.4}
Assumption 1.3 imposes a type of statistical completeness. Completeness plays a role analogous to the rank condition for identification in linear IV (Newey2003).\footnote{For some work on statistical completeness see Andrews2017b, Canay2013, Chena, DHaultfoeuille2011, Freyberger2017a, and Hu2018.} If Assumption 1.3 holds then Assumptions 1.1 point identifies $h_0$.
Assumption 1.4 is a technical condition that allows us to apply powerful results from functional analysis to the NPIV estimation problem. For common choices of ${\mathcal{B}_X}$ and ${\mathcal{B}_Z}$, Assumption 1.4 holds under weak primitive conditions on the joint distribution of $X$ and $Z$ (Florens2011, Horowitz2011). However, the assumption rules out the case of $X=Z$ which corresponds to standard non-parametric regression.\footnote{When $X=Z$ and $\mathcal{B}_X=\mathcal{B}_Z$, $\mathbb{A}$ is the identity and cannot be compact unless $\mathcal{B}_X$ is finite-dimensional. See Theorem 2.20 in Kress2014.}
As a leading example we consider $\mathcal{B}_X$ and $\mathcal{B}_Z$ to be spaces of continuous functions equipped with the sup norm $|\cdot|_\infty$. In this case a sufficient condition for Assumption 1.4 is that $X$ and $Z$ have compact supports and admit a strictly positive continuous joint probability density.\footnote{This implies the conditional density of $X$ given $Z$ is continuous and so we can apply Theorem 2.21 in Kress2014.}
While our results apply for quite general choices of the underlying function spaces $\mathcal{B}_X$ and $\mathcal{B}_Z$, Assumption 1.5 below places some mild restrictions on these spaces. The condition ensures that means of functions in these spaces are finite. We also assume that the outcome has a finite mean.
\newtheorem*{A1.5}{Assumption 1.5 (First Moments)} \begin{A1.5} $E[|Y|]<\infty$, $E[|h(X)|]<\infty,\forall h\in\mathcal{B}_X$, and $E[|g(Z)|]<\infty,\forall g\in\mathcal{B}_Z$. The mapping from an element $h\in\mathcal{B}_X$ to its mean $E[h(X)]$ is continuous. \end{A1.5}
Let us define a function $g_0\in\mathcal{B}_Z$ by $g_0(Z)=E[Y|Z]$. Using $\mathbb{A}$ defined in Assumption 1.4, we can rewrite the moment condition in Assumption 1.1 as an equation $\mathbb{A}[h_{0}]=g_{0}$. If Assumption 1.3 holds then this equation has a unique solution $h_{0}=\mathbb{A}^{-1}[g_{0}]$. The quantity on the right-hand side depends only on the distribution of the observables, and so $h_0$ is identified. If the NPIV moment condition is misspecified then $h_0$ need not equal $\mathbb{A}^{-1}[g_{0}]$, and so we refer to the latter as the `pseudo-solution'.
NPIV estimation is often said to be `ill-posed'. This refers to the fact that under Assumptions 1.3 and 1.4, the pseudo-solution varies discontinuously with $g_0$. In fact, if we perturb $g_{0}$ by some arbitrarily small amount, this can induce an arbitrarily large change in the pseudo-solution. This is problematic for estimation because $g_0$ and $\mathbb{A}$ must be empirically estimated, and thus are subject to error. If we substitute these estimates into the pseudo-solution then ill-posedness suggests the resulting estimate of $h_0$ can be highly inaccurate, even if $g_0$ and $\mathbb{A}$ are precisely estimated.
To tackle this problem researchers have proposed numerous `regularization' methods. Regularization involves replacing the discontinuous operator $\mathbb{A}^{-1}$ in the pseudo-solution with a continuous approximation. This generally imparts bias and so the degree of regularization is reduced as the sample size grows.\footnote{Available regularization methods including Tikhonov regularization, projection onto a finite-dimensional sieve space, and many others. See Darolles for discussion.} Reducing the strength of regularization increases sensitivity to error in the estimates of $g_0$ and $\mathbb{A}$, but this error typically decreases with the sample size. If the regularization is relaxed sufficiently slowly then the reduction in error more than offsets the increased sensitivity.
A key insight in this work is that misspecification, due to say invalid instruments, perturbs $g_0$ away from the value it would take under correct specification. The pseudo-solution is highly sensitive to perturbations due to misspecification just as it is to estimation error. NPIV estimators are typically designed to converge in probability to the pseudo-solution (Ai2007), estimates that converge to the pseudo-solution under weak conditions are thus highly sensitive to misspecification, at least asymptotically. A similar sensitivity result applies for the identified set under a slight relaxation of the NPIV moment condition. To the best of our knowledge this point is entirely absent from the rest of the NPIV literature. The precise consequences for identification and estimation depend crucially on whether or not strong a priori conditions are imposed on the structural function. Strong restrictions can shrink the identified set and limit the conditions under which an NPIV estimator converges to the pseudo-solution.
The NPIV model is misspecified if Assumption 1.1 does not hold. That is, ${u}_0(Z)$ defined below is non-zero with positive probability:
\theoremstyle{definition} \newtheorem*{E1}{Example 1: Direct Effect of the Instrument} \begin{E1}
Suppose we relax the model for potential outcomes ((ref)) so that the instrument may have a direct and additively separable effect on outcomes:
\[ Y_{x,z}=h_{0}(x)+{u}_{0}(z)+\varepsilon \]
Without loss of generality we suppose $E[\varepsilon]=0$ and $E[u_{0}(Z)]=0$ so that $h_{0}(x)=E[Y_{x}]$. We assume as before that the instrument is randomly assigned so that $E[\varepsilon|Z]=0$. This model implies ((ref)). \end{E1}
\theoremstyle{definition} \newtheorem*{E2}{Example 2: Confounded Instrument} \begin{E2} Suppose that ((ref)) holds with $E[\varepsilon]=0$, but the instrument is not randomly assigned. Instead, $Z$ and $\epsilon$ are statistically associated due to their shared dependence on a latent factor $V$ and assume $Z\perp \!\!\! \perp \epsilon|V$. Let $f_{\varepsilon|V}$ and $f_{V|Z}$ be the conditional probability densities of $\varepsilon$ given $V$, and of $V$ given $Z$ respectively. In this case ((ref)) holds with $h_{0}(x)=E[Y_{x}]$ and ${u}_{0}$ given below: \[ {u}_{0}(z)=\int\int e f_{\varepsilon|V}(e|v)def_{V|Z}(v|z)dv \] \end{E2}
\theoremstyle{definition} \newtheorem*{E3}{Example 3: Nonseparable IV Model} \begin{E3} Consider the fully nonseparable potential outcomes model $Y_{xz}=l_{0}(x,\varepsilon)$. As before, instruments have no direct effect on outcomes. Let $X_{z}=r_{0}(z,V)$ be the potential treatment from a counterfactual level of the instrument $z$. In this nonseparable model random assignment of the instrument can be formalized as $Z\perp \!\!\! \perp(\varepsilon,V)$. Suppose $\varepsilon$ and $V$ admit a joint density $f_{\varepsilon V}$ and marginal densities $f_{\varepsilon}$ and $f_{V}$. Let $h_{0}(x)=E[Y_{x}]$ then ((ref)) holds with ${u}_{0}$ given below: \[ {u}_{0}(z)=\int\int l_{0}(r_{0}(z,v),e)\big(f_{\varepsilon V}(e,v)-f_{\varepsilon}(e)f_{V}(v)\big)dedv \] \end{E3}
\theoremstyle{definition} \newtheorem*{A1.1alt}{Assumption $1.1^*$ (Nearly Correct Specification)} \begin{A1.1alt} $E[Y-h_{0}(X)|Z]={u}_0(Z)$ where ${u}_0\in\mathcal{U}\subseteq \mathcal{B}_Z$ and $\|{u}_0\|_{\mathcal{B}_Z}\leq b$. \end{A1.1alt}
The scalar $b$ in Assumption $1.1^*$ controls the degree of misspecification. If $b=0$ then Assumption 1.1 holds and so the NPIV moment condition is correctly specified. If $b$ is small but non-zero then the two sides of the moment condition in Assumption 1.1 need not be equal, but they must be close to equal, with the distance between them measured by $\|\cdot\|_{\mathcal{B}_Z}$. If $\mathcal{B}_Z$ is equipped with the sup norm then $\|{u}_0\|_{\mathcal{B}_Z}\leq b$ is equivalent to $|E[Y-h_{0}(X)|Z]|\leq b$.
Other choices of norms correspond to stronger or weaker notions of misspecification. For example if we take $\|\cdot\|_{\mathcal{B}_Z}$ to be the $L_2(\mu_Z)$ norm then we restrict the second moment of ${u}_0(Z)$ to be less than $b$, which is weaker than the corresponding restriction on the sup norm.
Bounds on the norm of $u_0$ can be derived from more primitive conditions. Consider Example 1 above and let $\mathcal{B}_Z$ be a set of continuous functions equipped the sup norm. The direct causal effect of changing the instrument from $z_1$ to $z_2$ is $u_0(z_2)-u_0(z_1)$. Suppose that for all $z_1,z_2$ the magnitude of this effect is at most $b$, given $E[u_0(Z)]=0$ it follows that $\|u_0\|_{\mathcal{B}_Z}\leq b$. Conversely, $u_0(z_2)-u_0(z_1)$ must be bounded above by $2\|u_0\|_{\mathcal{B}_Z}$. Thus the norm of $u_0$ is small if and only if the instrument has at most a small direct impact on outcomes.
In Example 2 one can show the sup norm of $u_0$ must be less than the essential supremum of $2|E[\varepsilon|V]-E[\varepsilon]|$. This quantity measures the strength of the dependence between the structural residual $\varepsilon$ and the unobserved confounder $V$. In Example 3 one can derive the following bound: \[\sup_{z\in\mathcal{Z}}|u_0(z)|\leq 2\inf_{h,q\in\mathcal{B}_{X}}\sup_{x\in\mathcal{X},e\in\mathcal{E}}|l_{0}(x,e)-h(x)-q(e)|\] The expression on the right-hand side measures how closely $l_0$ can be approximated by an additively separable function. Thus a bound the approximation error implies a bound on the norm of $u_0$.
The restriction that $u_0\in \mathcal{U}$ allows us to incorporate additional conditions on $u_0$. In Examples 1 and 2, but not necessarily Example 3, $E[u_0(Z)]=0$ by construction, and so we may take $\mathcal{U}$ to contain only functions with zero-mean. $\mathcal{U}$ can also incorporate smoothness restrictions. In Example 1 $u_0$ is smooth if, loosely speaking, small changes in $Z$ have small direct causal effects on the outcome.
In order to accommodate cases in which $\mathcal{U}$ contains only zero-mean functions it is helpful to define the subspaces $\tilde{\mathcal{B}}_X$ and $\tilde{\mathcal{B}}_Z$ so that $\tilde{\mathcal{B}}_X$ contains all $h\in{\mathcal{B}}_X$ with $E[h(X)]=0$ and $\tilde{\mathcal{B}}_Z$ contains all $g\in{\mathcal{B}}_Z$ with $E[g(Z)]=0$.
Our first result concerns the size of the identified set under the a priori assumptions $1.1^*$ and $1.2$. The identified set $\Theta_b$ is the set of functions that are consistent with these two assumptions:
Theorem 1.1 concerns the diameter of the identified set. The diameter, denoted $diam(\Theta_b)$ is defined as follows: \[ diam(\Theta_b)=\sup_{h_1, h_2 \in \Theta_b}\|h_1-h_2\|_{\mathcal{B}_X} \]
\theoremstyle{plain}
Theorem 1.1 examines the diameter of the identified under different degrees of misspecification. The behavior of the set for small values of $b>0$ depends crucially on the strength of the a priori restrictions on the structural function, as captured in $\mathcal{H}$. Part a. of the theorem requires that the pseudo-solution lies in the interior of the closure of $\mathcal{H}$. Some spaces of smooth functions are sufficiently restrictive that they are compact, in which case the interior is empty and part a. cannot apply.\footnote{A compact subset of an infinite-dimensional Banach space must be closed and have an empty interior.} Part b. provides results for these more restrictive parameter spaces.
Under the conditions of part a., the diameter of the identified set does not shrink to zero with $b$. Instead it remains bounded below by a constant that depends on $\mathcal{H}$ and $\mathcal{U}$. Consider the extreme case in which $\mathcal{H}$ is dense in $\mathcal{B}_X$ and $\mathcal{U}$ contains an open ball centered at zero. Then diameter of the identified set is infinite, regardless of how small we make the bound $b>0$.\footnote{If $\mathcal{U}$ contains an open $\tilde{\mathcal{B}}_Z$-ball at zero then for sufficiently small $b>0$ we can replace $\mathcal{U}$ with $\tilde{\mathcal{B}}_Z$ (so $r_u=\infty$) and the resulting identified set contains the original as a subset.} Under the conditions of part b., the identified set shrinks to a point as the degree of misspecification is reduced. However, the diameter of the identified set shrinks to zero strictly more slowly than the bound $b$.
Let us compare to standard non-parametric regression. This case corresponds to the NPIV moment condition with $X=Z$. Let $\mathcal{B}_X=\mathcal{B}_Z$. Recall that Assumption 1.4 cannot hold in this setting and so Theorem 1.1 does not apply. In this case the identified set has diameter at most $2b$, regardless of the choice of $\mathcal{H}$ and $\mathcal{U}$. The identified set shrinks to zero at exactly the same rate as $b$. This is strictly faster than the rate at which the identified set shrinks in the NPIV case, even under the strong restrictions on $\mathcal{H}$ imposed by part b. of Theorem 1.1.
We now specify spaces $\mathcal{H}$ and $\mathcal{U}$ compatible with parts a. and b. of the theorem. For concreteness let $\mathcal{B}_X$ be the set of continuous functions with the sup norm and assume $X$ has compact support.
First take $\mathcal{H}$ to be the set of Lipschitz continuous functions whose magnitude is bounded by a known constant $\bar{c}<\infty$:
By the Stone-Weirstrauss theorem the set of Lipschitz functions is dense $\mathcal{B}_X$. Thus this set satisfies the conditions in part a. of Theorem 1.1 with $r_h=\bar{c}$.
Now suppose we further restrict $\mathcal{H}$ so that it contains only Lipschitz continuous functions with Lipschitz constant less than $C<\infty$:
The closure of this set does not contain an open ball so we cannot apply part a. of Theorem 1.1. In fact, the set is compact (see Freyberger2017) and it is absolutely convex and infinite-dimensional. Thus this choice of $\mathcal{H}$ satisfies the restriction in part b. of Theorem 1.1. Part b. also requires that the pseudo-solution $\mathbb{A}^{-1}[g_0]$ be in $\mathcal{H}$ and that $\frac{1}{\alpha}\mathbb{A}^{-1}[g_0]\in\mathcal{H}$ for some $\alpha\in(0,1)$. Given the choice of $\mathcal{H}$ above, this holds if and only if $\mathbb{A}^{-1}[g_0]$ is Lipschitz with constant strictly less than $C$ and bounded in magnitude by a constant strictly less than $\bar{c}$.
Compact parameter spaces like ((ref)) are employed in many NPIV papers. For example they appear in Newey2003, Ai2003, Blundell2007, Freyberger2017a, and Santos2012.
The distinction between the sets ((ref)) and ((ref)) may seem subtle. However, ((ref)) encodes much stronger knowledge about the structural function $h_0$. Not only do we know $h_0$ is Lipschitz continuous, we know it is Lipschitz continuous with a Lipschitz constant smaller than a known scalar $C$. A similar distinction holds for other smoothness conditions. For example, in Section 2 we consider a space of functions whose second derivatives are bounded. In order to achieve an identified set that shrinks to zero with $b$ we must be willing to assume a specific bound on the second derivatives of $h_0$. It is not enough to simply assume that there exists some unknown finite bound on the second derivatives.\footnote{For more examples of infinite-dimensional sets of compact functions see Freyberger2017. All examples in that paper also satisfy the absolute convexity requirement.}
The condition on $\mathcal{U}$ is the same for both parts of the Theorem, it states that ${\mathbb{A}^{-1}[\mathcal{U}]}$ contains an open $\tilde{\mathcal{B}}_X$-ball centered at zero of radius $r_u$. This condition is relatively weak in that it can hold even if $\mathcal{U}$ contains only very smooth functions. Suppose we impose the same strong Lipschitz condition on $u_0$ as in ((ref)):
Although this set is restrictive, ${\mathbb{A}^{-1}[\mathcal{U}]}$ can contain an open ball at zero. The following is a sufficient (but not necessary) condition for such a ball to exist given the choice of $\mathcal{U}$ above. Let $f_{X|Z}$ be the conditional probability density of $X$ given $Z$. Suppose $f_{X|Z}$ is well-defined and Lipschitz continuous in its $Z$ argument: \[\sup_{x\in\mathcal{X}}\sup_{ z_1\neq z_2\in \mathcal{Z}}\frac{|f(x|z_1)-f(x|z_2)|}{\|z_1-z_2\|_2}<\infty \] Denote the quantity on the right-hand side above by $c$. Then $\mathbb{A}^{-1}[\mathcal{U}]$ contains an open $\tilde{\mathcal{B}}_X$-ball of radius $\min\{\bar{c},C/c\}$.
The sensitivity of a given NPIV estimator to misspecification may be substantially worse than is suggested by Theorem 1.1. The reason for this is that under misspecification the probability limit of a particular NPIV estimate may be outside of the identified set. We show below that the sensitivity of an NPIV estimator to misspecification depends crucially on the conditions under which it is consistent under correct specification. In particular, an NPIV estimator that achieves consistency under only weak a priori conditions on the structural function is necessarily very sensitive to misspecification. Thus in order to obtain estimates that are robust to misspecification, one must forgo consistent estimation whenever the true structural function lies outside a restrictive parameter space.
In order to capture this formally we introduce a notion of asymptotic bias. In particular, we consider the largest possible (un-scaled) asymptotic bias under Assumptions $1.1^*$, 1.2, 1.3, and 1.4. We fix parameters that are not directly related to misspecification. That is, we fix the structural function $h_{0}$ and $\mu_{XZ\eta}$ the joint probability distribution of $X$, $Z$, and $\eta\equiv Y-E[Y|Z]$. The worst-case asymptotic bias of an estimator $\hat{h}_{n}$ for a given $b$ is then: \[ bias_{\hat{h}_{n}}(b)=\sup_{u_{0}\in \mathcal{U}:\,\|u_{0}\|_{\mathcal{B}_{Z}}\leq b}\inf \big[\epsilon : P(\|\hat{h}_{n}-h_{0}\|_{\mathcal{B}_{X}}\leq \epsilon)\to 1 \big] \] The above is a version of the `maximum bias' discussed in Huber2011. The definition above does not require that $\hat{h}_n$ has a probability limit.\footnote{The existence of a probability limit for an NPIV estimator is difficult to establish under misspecification, sufficient conditions for the existence of a probability limit for misspecified sieve minimum distance estimators can be found in Ai2007.} If $\hat{h}_n$ converges in probability to a limit $\tilde{h}$ then the infimum in the definition above is equal to $\|\tilde{h}-h_{0}\|_{\mathcal{B}_{X}}$.
\theoremstyle{plain}
Theorem 1.2 shows that there is a trade-off between consistency under correct specification and robustness to misspecification. Suppose that under correct specification and some restrictions on the distribution of observables, an NPIV estimator is consistent whenever $h_0\in\mathcal{S}$. Ideally $\mathcal{S}$ would be large so that consistency does not require strong a priori assumptions on the structural function. However, if $\mathcal{S}$ is large (in the sense given in the theorem), then the estimator must be very sensitive to certain kinds of misspecification. This is true even if the structural function $h_0$ in fact lies in a set $\mathcal{H}$ that is much more restrictive than $\mathcal{S}$.
Theorem 1.2 can be used to derive the sensitivity properties of particular NPIV estimators. To apply the theorem we can use existing asymptotic results to establish consistency over a sufficiently large space $\mathcal{S}$. In Appendix A1 we verify the key conditions of the theorem for some highly-cited NPIV estimators.
Some NPIV estimators are constructed in such a way that Theorem 1.2 cannot apply. These estimators are constrained so that $\hat{h}_n$ is always an element of a restrictive parameter space like ((ref)) whose closure has an empty interior. For example, let $\hat{Q}_n$ be an empirical objective function that may change with the sample size (for example, to incorporate a penalty function with decreasing strength) and let $\mathcal{S}_n$ be a finite-dimensional subset of $\mathcal{S}$ that can grow with the sample size. Many NPIV estimators take the form below, including those in Newey2003 and Blundell2007. \[ \hat{h}_n=\arg\min_{h\in\mathcal{S}_n}\hat{Q}_n(h) \]
Because $\mathcal{S}_n$ is a subset of $\mathcal{S}$, the estimate $\hat{h}_n$ above must be an element of $\mathcal{S}$ for all $n$. Such an estimator is necessarily inconsistent if $h_0\notin \mathcal{S}$. If $\mathcal{S}$ is sufficiently restrictive, for example if it takes the form ((ref)), then the sensitivity result in Theorem 1.2 cannot apply. Thus such estimators forgo consistency outside of a restrictive parameter space but have the advantage that they potentially avoid the non-robustness to misspecification suggested by the theorem. If the researcher is certain that $h_0$ really does satisfy the restrictions imposed by inclusion in $\mathcal{S}$, then there is no cost to inconsistency outside of $\mathcal{S}$.
Our next result establishes conditions on an estimator that ensure it is robust in the sense that $\lim_{b\to 0} bias_{\hat{h}_{n}}(b)=0$. However, analogous to Theorem 1.1, the rate at which the bias goes to zero must be strictly slower than the rate at which the degree of misspecification shrinks to zero
The key condition is that $\hat{h}_n$ converges in probability to a function $\mathbb{A}^{-1}P_{Z}[g_{0}]$. $P_Z$ is a projection onto $\mathbb{A}[\mathcal{S}]$. That is, a possibly non-linear operator that maps functions into $\mathbb{A}[\mathcal{S}]$ and leaves those already in this set unchanged. This condition implies the estimator has a probability limit in $\mathcal{S}$. Thus if $h_0$ is not in $\mathcal{S}$ then the estimator is necessarily inconsistent, even under correct specification. Conversely, if Assumption 1.1 holds and $h_0\in\mathcal{S}$, then $h_0=\mathbb{A}^{-1}P_{Z}[g_{0}]$ and so $\hat{h}_n$ is consistent.
\theoremstyle{plain}
Let us compare again with the case of standard nonparametric regression (where $Z=X$). Suppose that under some conditions on $\mu_{XZ\eta}$ a nonparametric regression estimator converges to the condition mean $g_0(X)=E[Y|X]$. If Assumption 1.1 holds then $g_0(X)=h_0(X)$ and so the estimator is consistent under correct specification. Under misspecification, the difference between the conditional mean and the structural function is simply $u_0(X)$. If we use the same norm to measure the degree of misspecification (i.e, the magnitude of $u_0$) and the worst-case bias, then this estimator has worst-case bias $b$. Thus the estimator is robust and has worst-case bias that shrinks at the same rate as the degree of misspecification. This does not require projection onto some restricted parameter space $\mathcal{S}$.
The sensitivity results in Theorems 1.1-1.3 apply to the identified set for the structural function itself and estimators of the structural function. If the object of interest is instead a smooth linear functional of the structural function then the identified set may shrink to zero with the degree of misspecification, even without any strong a priori smoothness restrictions. Moreover, it may be possible to construct estimates that a) are consistent under correct specification of the NPIV moment condition without any a priori restrictions on the true structural function and also b) are robust to the failure of instrumental validity.
The estimation of functionals of the structural function in NPIV models is analyzed extensively in the literature (Ai2003, Severini2012, Ichimuraa, and others). Following Severini2012, we let $\mathcal{B}_{X}=L_{2}(\mu_{X})$ and similarly $\mathcal{B}_{Z}=L_{2}(\mu_{Z})$. These choices of function spaces have the advantage that any continuous linear functional $\mathbb{L}$ takes the form $\mathbb{L}[h]=E[w(X)h(X)]$ where $w\in L_{2}(\mu_{X})$. This characterization follows by the Reisz representation theorem.
The identified set for the linear functional $\mathbb{L}[h_0]$ is simply $\mathbb{L}[\Theta_b]$. Let $\hat{l}_{n}$ be an estimator of $\mathbb{L}[h_{0}]$ and fix $h_{0}$ and $\mu_{XZ\eta}$. The worst-case asymptotic bias of $\hat{l}_{n}$ is defined below. \[ bias_{\hat{l}_{n}}(b)=\sup_{u_{0}\in R(\mathbb{A}):\,\|u_{0}\|_{L_{2}(\mu_{Z})}\leq b}\inf \big[\epsilon : P(|\hat{l}_{n}-\mathbb{L}[h_{0}]|)\leq \epsilon)\to 1 \big] \]
Severini2012 show that $E[w(X)h(X)]$ is estimable at rate $\sqrt{n}$ only if there exists a function $\alpha\in L_{2}(\mu_{Z})$ so that $ w(X)=E[\alpha(Z)|X]$. Theorem 1.4 shows that this condition is necessary and sufficient for robust estimation of $\mathbb{L}[h_{0}]$ to be possible without restricting the parameter space.
Theorem 1.4.a suggests that robust estimation of a continuous linear functional is possible without imposing any strong conditions on the NPIV estimates. However, the condition that $w(X)=E[\alpha(Z)|X]$ for some $\alpha$ may be hard to verify. In fact, we conjecture that for a given function $w$, it is not possible to empirically verify this condition.\footnote{More precisely, we conjecture that given a function $w$ any test of the null hypothesis that there is no $\alpha$ with $w(X)= E[\alpha(Z)|X]$ has power no greater than size against any alternative.}
The results in Section 1 show that identification and estimation in NPIV models entails a trade-off. Strong a priori restrictions allow for identification and estimation that is somewhat robust to misspecification. However, if the true structural function violates these restrictions then it will not lie in the corresponding identified set and an estimator that imposes the conditions will be inconsistent, even if the NPIV moment condition holds.
To help researchers assess the robustness of their findings, and to better assess the trade-off between robustness and strong smoothness conditions, we present a method for empirical sensitivity analysis in NPIV models. Our procedure is based on estimation of the identified set under the weakened instrumental validity condition in Assumption $1.1^*$ and Assumption 1.2. We estimate the identified set for a range of bounds $b$ on the degree of misspecification and a range of smoothness restrictions. The approach is similar in spirit to Masten2018 who analyze the sensitivity of treatment effect estimates to a small failure of unconfoundedness by weakening the assumptions used for point identification. More precisely, we estimate the identified set for linear functionals of $h_0$. An important special case is the identified set for $h_0(x)$ at some $x$ in the support of $X$.
Note that our approach assesses the robustness of empirical findings to deviations from instrumental validity and weakened smoothness conditions, it does not assess the sensitivity of any particular NPIV estimator. In Appendix A3 we also provide a method to directly assess the finite-sample sensitivity of specific NPIV point-estimates to invalid instruments.
Let $\Theta_b$ be the identified set as defined in ((ref)). We only restrict the degree of misspecification and thus we take $\mathcal{U}$ to be the whole space $\mathcal{B}_Z$. Recall that the identified set for a linear functional $\mathbb{L}$ is $\mathbb{L}[\Theta_b]$. Unlike in the previous section we do not restrict ourselves to only continuous linear functionals. Of particular interest is the identified set for the value of $h_0$ at a point $x$, this is the set $[h(x): h\in\Theta_b]$ and corresponds to $\mathbb{L}:h\mapsto h(x)$.
For the choices of $\mathcal{B}_Z$ and $\mathcal{H}$ we consider, $\mathbb{L}[\Theta_b]$ is an interval.\footnote{See Proposition 2.1 in the Appendix B for a proof.} Denote the lower and upper endpoints for the interval by $\underline{\theta}_{\mathbb{L}}$ and $\bar{\theta}_{\mathbb{L}}$, for simplicity we suppress the subscript $b$. We wish to estimate these end-points.
Assumption $1.1^*$ depends on the norm of the space $\mathcal{B}_Z$ and Assumption 1.2 on the choice of $\mathcal{H}$. We assume that these spaces are chosen so that the two assumptions can be written in terms of inequality constraints. \newtheorem*{A2.1}{Assumption 2.1} \begin{A2.1} i. $\mathcal{U}=\mathcal{B}_Z$ and $\|\cdot\|_{\mathcal{B}_Z}$ is the sup norm. Thus Assumption $1.1^*$ states that if $h=h_0$, then with probability $1$:
ii. $\mathcal{H}$ is the set of continuous bounded functions $h$ so that:
Where $\mathbb{T}$ is a known linear functional from $\mathcal{B}_{X}$ to $\mathcal{B}_{X}^{d}$, and $C\in \mathcal{B}_{X}^{d}$ is a known vector of functions. \end{A2.1} Under Assumption 2.1 $\underline{\theta}_\mathbb{L}$ and $\bar{\theta}_\mathbb{L}$ are defined as follows:
Our statistical analysis applies for more general constraints than ((ref)) and ((ref)). We detail the more general framework in Appendix A2.
The bound $b$ in Assumption $1.1^*$ determines the size of the identified set. We suggest the researcher estimate the identified set for a range of choices for $b$ corresponding to mild, moderate, and severe misspecification. To determine whether $b$ corresponds to say, `mild' misspecification, it is helpful to compare $b$ to an estimable reference quantity. This approach is common in the empirical sensitivity literature. For example, Masten2018 compare the degree of selection on unobservables to selection on observables, which is estimable. See also Imbens2003, Altonji2005a, and Oster2019.
We suggest calibrating of $b$ against the residual variation in the outcome. Consider the following decomposition: \[Y=E[h_0(X)|Z]+u_0(Z)+\eta\] In the context of Example 1 in Section 1, the first two terms on the right-hand side capture the additive indirect and direct effects of the instrument. The residual $\eta$ absorbs the effect of all other factors that influence $Y$. The variation in $\eta$ is a useful reference quantity against which to calibrate the bound $b$ on the magnitude of $u_0(Z)$. We measure the variation using a difference in quantiles.
Let $q_\eta(\alpha)$ denote the $\alpha$ quantile of $\eta$. We say $b$ is small if it is equal to $q_\eta(0.5+\tau)-q_\eta(0.5-\tau)$ for a small $\tau$. Suppose $b=q_\eta(0.6)-q_\eta(0.4)$ which corresponds to $\tau=0.1$. In Example 1, this restricts the direct causal effect of the instrument to be less than twice the effect of shifting $\eta$ from its $0.4$ to its $0.6$ quantile. In our application we take $\tau=0.01$, $0.05$, and $0.1$ to correspond to mild, moderate, and severe misspecification respectively. Note that $\eta=Y-E[Y|Z]$ is a reduced-form residual and thus we can estimate $\eta$ and its quantiles in a first stage.
To estimate $\underline{\theta}_\mathbb{L}$ and $\bar{\theta}_\mathbb{L}$ we must replace the optimization problems that define these quantities with feasible ones. Instead of optimizing over $\mathcal{B}_X$, we optimize over a finite-dimensional sieve space. We replace the conditional moment in ((ref)) with a non-parametric estimate. Finally, we enforce the inequalities ((ref)) and ((ref)) only on finite grids.
Let $\Phi_{n}$ be a length-$K_{n}$ column vector of basis functions on $\mathcal{X}$. In a first stage, we nonparametrically regress $Y$ and $\Phi_n(X)$ on $Z$ to obtain estimates $\hat{g}_{n}$ and $\hat{\Pi}_{n}$ of $g_{0}(Z)=E[Y|Z]$ and $\text{\ensuremath{\Pi}}_{n}(Z)=E[\Phi_{n}(X)|Z]$. Let $\mathcal{X}_{n}$ and $\mathcal{Z}_{n}$ be finite grids in the supports of $X$ and $Z$. We replace the constraints ((ref)) and ((ref)) with:
Where $\mathbb{T}[\Phi_n'](x)$ be the $d$-by-$K_n$ matrix whose $k^{th}$ column is $\mathbb{T}[\Phi_{n,k}](x)$ where $\Phi_{n,k}$ is the $k^{th}$ component of $\Phi_n$. Let $\mathbb{L}[\Phi_n]$ be the column-vector whose $k^{th}$ entry is $\mathbb{L}[\Phi_{n,k}]$. The estimates of $\bar{\theta}_\mathbb{L}$ and $\underline{\theta}_\mathbb{L}$ are respectively:
These are linear programming problems with $K_{n}$ scalar parameters and $2|\mathcal{Z}_{n}|+d|\mathcal{X}_{n}|$ constraints.
A vector-valued function $f$ on a domain $\mathcal{V}$ is of H{\"o}lder smoothness class $s\in (0,1]$ with constant $\xi$ if and only if: \[ \sup_{v_{1},v_{2}\in\mathcal{V}: v_1 \neq v_2}\frac{\|f(v_{1})-f(v_{2})\|_2}{\|v_{1}-v_{2}\|_{2}^s}=\xi \] Let $\lfloor s \rfloor$ denote the largest integer less than $s$. $f$ is of H{\"o}lder smoothness class $s>1$ with constant $\xi$ if and only if all the derivatives of each entry of $f$ of order weakly less than $\lfloor s \rfloor$ are uniformly bounded by $\xi$ and all the derivatives of order exactly $\lfloor s \rfloor$ are of H{\"o}lder smoothness class $ s - \lfloor s \rfloor$ with constant $\xi$. A function is Lipschitz continuous with constant $\xi$ if it is of H{\"o}lder smoothness class $1$ with constant $\xi$.
Let $D_{1,n}=\underset{{x_{1}\in\mathcal{X}}}{\sup}\underset{x_{2}\in\mathcal{X}_{n}}{\min}\|x_{1}-x_{2}\|_{2}$ and $D_{2,n}=\underset{{z_{1}\in\mathcal{Z}}}{\sup}\underset{{z_{2}\in\mathcal{Z}_{n}}}{\min}\|z_{1}-z_{2}\|_{2}$. Let $C_{n}=1+\sup_{\beta\in\mathbb{R}^{K_{n}}}\frac{\|\beta\|_{2}}{|\Phi_{n}'\beta|_{\infty}} $. \theoremstyle{definition} \newtheorem*{A2.2}{Assumption 2.2} \newtheorem*{A2.3}{Assumption 2.3} \newtheorem*{A2.4}{Assumption 2.4} \newtheorem*{A2.5}{Assumption 2.5} \begin{A2.2} $\mathbb{T}[h](x)\leq C(x)$ implies $|h(x)|\leq\bar{c}$ for some $0<\bar{c}<\infty$, and for some $\underline{c}>0$, $C(x)\geq \underline{c}$ for all $x\in\mathcal{X}$. \end{A2.2} \begin{A2.3} There is a sequence of positive scalars $a_{n}\to0$ so that: \[ |\hat{g}_{n}-g_{0}|_{\infty}+\sup_{\beta\in\mathbb{R}^{K_{n}}:\,\Phi_{n}'\beta\in\mathcal{H}}|(\hat{\Pi}_{n}-\Pi_{n})'\beta|_{\infty}=O_{p}(a_{n}) \] \end{A2.3} \begin{A2.4} There is a sequence of positive scalars $\kappa_{n}\to0$ so that for any $h\in\mathcal{H}$ there exists $\beta_{n}\in\mathbb{R}^{K_{n}}$ with $|\Phi_{n}'\beta_{n}-h|_{\infty}\leq\kappa_{n}$ and:
\end{A2.4} \begin{A2.5} i. $\Phi_{n}$, $C$, and each row of $\mathbb{T}[\Phi_{n}']$ are Lipschitz continuous with constant at most $\xi_{n}$. ii. With probability approaching $1$ both $\hat{g}_{n}$ and $\hat{\Pi}_{n}$ are Lipschitz continuous with constant at most $G_{n}$. iii. $D_{1,n},D_{2,n}\to0$, $C_{n}\xi_{n}D_{1,n}\to0$ and $C_{n}G_{n}D_{2,n}\to0$. \end{A2.5}
Assumption 2.2 implies that $\mathcal{H}$ is bounded and convex. Assumption 2.3 quantifies the estimation error in $\hat{g}_n$ and $\hat{\Pi}_n$. 2.4 quantifies the error from the replacing $\mathcal{B}_{X}$ with a sieve space. Theorems 2.2 and 2.3 establish primitive conditions for Assumptions 2.3 and 2.4. Assumption 2.5 controls the error from the use of grids $\mathcal{X}_{n}$ and $\mathcal{Z}_{n}$. \theoremstyle{plain}
Theorem 2.1 demonstrates the well-posedness of the set estimation problem. The first-stage rate $a_{n}$ is not premultiplied by some growing factor like a `sieve measure of ill-posedness' (Blundell2007).
$D_{1,n}$ and $D_{2,n}$ are small when the grids $\mathcal{X}_{n}$ and $\mathcal{Z}_{n}$ are dense. $a_{n}$ and $\kappa_{n}$ are independent of the grids, so if the grids grow dense quickly enough, the rates in the Theorem reduce to $a_n+\kappa_n$. This suggests the grids should be made as dense as computationally feasible.
If the dimension of the sieve space $K_{n}$ grows quickly then $\kappa_{n}$ converges rapidly to zero. Theorem 2.2 establishes a rate $a_{n}$ that is independent of $K_{n}$. This suggests that under the conditions of Theorem 2.2, $K_{n}$ should be made as large as is computational constraints allow.
We now provide primitive conditions for Assumption 2.3. Define series estimators $\hat{g}_n$ and $\hat{\Pi}_n$. Let $\Psi_{n}$ be a length-$L_{n}$ column vector of basis functions on $\mathcal{Z}$ and define $\hat{Q}_{n}=\frac{1}{n}\sum_{i=1}^{n}\Psi_{n}(Z_{i})\Psi_{n}(Z_{i})'$:
In the Assumptions below $\mathcal{N}(\mathcal{H},|\cdot|_{\infty},\delta)$ is the smallest number of $|\cdot|_{\infty}$-balls of radius $\delta$ that can cover $\mathcal{H}$ and $\dim (Z)$ is the dimension of $Z$.
\theoremstyle{definition} \newtheorem*{A2.6}{Assumption 2.6} \newtheorem*{A2.7}{Assumption 2.7} \newtheorem*{A2.8}{Assumption 2.8} \begin{A2.6} i. The eigenvalues of $Q_{n}=E[\Psi_{n}(Z_{i})\Psi_{n}(Z_{i})']$ are bounded uniformly above and away from zero. ii. $\mathcal{Z}$ is bounded and the distribution of $X$ given $Z$ admits a conditional density $f_{X|Z}$ so that $\forall x\in\mathcal{X}$, $f_{X|Z}(x,\cdot)$ is of H{\"o}lder smoothness class $s>0$ with constant at most $\bar{\ell}$. \end{A2.6} \begin{A2.7} For any $s>0$ there is a sequence $R_n(s)\to 0$ so that for any $g\in\mathcal{B}_{Z}$ of H{\"o}lder smoothness class $s$ with constant $\xi$: \[ \sup_{z\in\mathcal{Z}}|g(z)-\Psi_{n}(z)'Q_{n}^{-1}E[\Psi_{n}(Z)g(z)]|\leq \xi R_n(s) \] ii. For all $z\in\mathcal{Z}$, $\|\Psi_{n}(z)\|_{2}\leq\bar{\xi}_{n}$. $\alpha_{n}(z)=\frac{\Psi_{n}(z)}{\|\Psi_{n}(z)\|_2}$ is Lipschitz continuous with constant $\ell_{n}$. iii. $\int_{0}^{1}\sqrt{log\mathcal{N}(\mathcal{H},|\cdot|_{\infty},u)}du<\infty$. iv. $\frac{\bar{\xi}_{n}^{2}log(L_{n})}{n}\to0$ \end{A2.7} \begin{A2.8} i. The function $g_{0}(Z)=E[Y|Z]$ is of H{\"o}lder smoothness class $s>0$. For $m>2$, $E\big[|Y-E[Y]|^{m}\big|Z\big]<\infty$, $\bar{\xi}_{n}^{2m/(m-2)}log(L_{n})/n=O(1)$, $L_{n}log(L_{n})/(n^{1-2/m})=O(1)$ and $L_{n}^{2-2/\dim(Z)}/n=O(1)$. ii. $log(\ell_{n})=O\big(log(L_{n})\big)$, $\bar{\xi}_{n}=O(\sqrt{L_{n}})$ and $R_{n}(s)=O(L_{n}^{-s_0 (s)/\dim(Z)})$ for some function $s_0:\mathbb{R}_{++}\to \mathbb{R}_{++}$ and $R_n(s)=O(L_n^{-1/2})$. \end{A2.8}
Assumption 2.6.i is standard. Smoothness of the conditional density in 2.6.ii ensures that for any $h\in\mathcal{B}_{X}$, $\mathbb{A}[h]$ is smooth. Assumptions 2.7.i and 2.7.ii. can be verified for common basis functions. 2.7.iii is a condition on the metric entropy of $\mathcal{H}$, such conditions are commonplace in the sieve estimation literature (see Chen2007a). Spaces of smooth functions like those used in our empirical application typically obey the condition (see Wainwright2019 Chapter 5). Assumption 2.7.iv allows us to apply Rudelson's matrix law of large numbers (Rudelson1999).
Assumption 2.8.i allows us to apply results from Belloni2015 to derive a convergence rate for $\hat{g}_{n}$. 2.8.ii can be verified for a given choice of basis functions, for example if $s\geq 1/2$, then the Assumption holds for the one-dimensional B-spline case in Section 3.
\theoremstyle{plain}
If the conditions of the theorem hold, then setting $L_{n}$ optimally Assumption 2.3 holds with $a_{n}=\big( \frac{log(n)}{n} \big) ^{s_0(s)/(2 s_0(s)+\dim (Z))}$.
Finally, we show Assumption 2.4 holds with $\kappa_{n}=O(K_{n}^{-\frac{1}{2}})$ for the setting in our empirical application.
To demonstrate the usefulness of our methods we apply them to setting in Section 5.1 of Horowitz2011. Horowitz2011 estimates a shape-invariant Engel curve for food using data from the British Family Expenditure Survey.\footnote{We made use of the data file that accompanies Horowitz2011 and adapted the accompanying code in order to evaluate Horowitz's estimator and B-spline bases for our own methods.} Horowitz's application is in turn based on Blundell2007 who also carry out NPIV estimation of shape-invariant Engel curves and use the same data.
From Horowitz2011: “The data are 1655 household-level observations from the British Family Expenditure Survey. The households consist of married couples with an employed head-of-household between the ages of 25 and 55 years.”
Blundell2007 and Horowitz2011 aim to estimate `structural' Engel curves. Suppose a researcher were to exogenously allocate to a household a particular budget for non-durable goods. A structural Engel curve $h_0$ measures the share of that budget the household would choose to allocate to a class of goods as a function of the budget size.
In observational settings, the share of household wealth allocated to nondurables is decided by the household. Therefore, the budget for nondurables depends on household preferences. These preferences also determine the allocation of the budget to classes of non-durable goods. Thus expenditure on non-durable goods is endogenous.
To tackle this problem Blundell2007 and Horowitz2011 use household income as an instrument for nondurables expenditure. There are many reasons to think that household income and preferences are correlated. Both tastes and income may depend on household size and socio-economic status, and those with expensive tastes may seek high-paying jobs. However, some shocks to household income may result from exogenous external factors like shocks to production costs in an employed householder's firm. If one controls for a sufficiently rich set of household covariates this may absorb the endogenous variation in income leaving only the exogenous variation. If the data do not contain rich enough household observables, then some endogeneity likely remains and the income instrument is unlikely to be fully valid.
Blundell2007 and Horowitz2011 can only control for some coarse demographic variables.\footnote{Both papers control for demographics by analyzing a homogeneous sub-sample. Blundell2007 incorporate additional dummy variable controls.} Therefore, in this setting instrumental validity may not hold exactly and it is of interest to assess what empirical findings are robust to some failure of instrumental validity.
In this setting $Y$ is the share of total expenditure on non-durables that a household spends on food. $X$ is the logarithm of the household's total expenditure and $Z$ is the logarithm of household income. We take $\mathcal{H}$ to be the set of functions that map from $\mathcal{X}$ to the unit interval and have second derivative bounded in magnitude by a constant $c$. We present results for $c=2$ and $c=5$, for comparison, the magnitude of the second derivatives of Horowitz's estimated structural function never exceed $0.5$.
Following Subsection 2.1 we set $b=\hat{q}_\eta(0.5+\tau)-\hat{q}_\eta(0.5-\tau)$ for different values of $\tau$. $\hat{q}_\eta$ is an estimate of $q_\eta$ and is calculated by taking empirical quantiles of $Y_i -\hat{g}_n(Z_i)$. In particular we consider $\tau=0.01,0.05,0.1$ to correspond to mild, moderate, and severe misspecification. These values of $\tau$ correspond to $b=0.0043,0.024,0.046$.
Following Horowitz2011 we let $\Phi_{n}$ be fourth-order (cubic) B-spline basis functions with evenly-spaced knot points.\footnote{See Boor2002 for a practical introduction to B-splines.} We carry out our first-stage estimates using series regression onto cubic B-splines. Motivated by the results in Theorem 2.2 we set $K_{n}$ to be large, specifically we let $K_{n}=10$. The grid $\mathcal{X}_{n}$ consists of 100 evenly spaced points between the smallest and largest observed values of $X$. The grid $\mathcal{Z}_{n}$ consists of 100 evenly spaced points between the $0.005$ and $0.995$ quantiles of $Z$.
Figure 4.1 contains the results of our procedure. The figure contains six sub-figures each corresponding to a different set of values for $\tau$ and $c$. The lower and upper dotted lines represent $\underline{\hat{\theta}}_{h\mapsto h(x)}$ and $\hat{\bar{\theta}}_{{h\mapsto h(x)}}$, the upper and lower end points of the identified set for $h_0(x)$. The thick black line represents the half-way point between the end points, which is a point estimator with the smallest possible worst-case asymptotic bias under Assumptions $1.1^*$ and 1.2. The thin black line is the estimator from Horowitz2011.
The results in Section 1 suggest that if the bound $c$ on the second derivatives is loose then the identified set for $h_0(x)$ will be large, even if $\tau$ (and hence $b$) is small. This is clear in Figure 4.1 which shows that for a given $\tau$, the intervals are wider when $c=5$ than when $c=2$.
Apart from in the severely misspecified case of $\tau=0.1$ (Sub-figures (c) and (f)), the figures suggest a general downward slope in the Engel curve for medium values of total expenditure. More precisely, the lower envelope at log-total expenditure of $5$ exceeds the upper envelope at $6$, which implies the Engel curve has decreased between these two values. None of the results in Figure 4.1 provide evidence in favor of an upward sloping Engel curve for low values of total expenditure as found by Horowitz2011 and so we conclude then that this finding is not robust to even a mild failure of instrumental validity.
Our results show that identification and estimation in NPIV can be highly sensitive to misspecification, and that the sensitivity depends crucially upon the strength of a priori restrictions on the structural function. We develop a method for empirical sensitivity analysis in NPIV that allows researchers to better assess the relationship between the smoothness restrictions they impose and the robustness of their findings to misspecification.
We conjecture that our sensitivity results extend to a broader class of conditional moment restriction models. The non-robustness of NPIV estimators is tied to the ill-posedness of NPIV and a range of other nonparametric conditional moment restriction models are likewise ill-posed. It may be possible to adapt our empirical methods to these settings, although non-linearity of the conditional moment restriction would likely complicate estimation of the identified set. We leave these extensions for future work.