EconBase
← Back to paper

Nonparametric Identification and Estimation with Independent, Discrete Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,336 characters · 13 sections · 35 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Identification and Estimation with Independent, Discrete Instruments

abstractIn a nonparametric instrumental regression model, we strengthen the conventional moment independence assumption towards full statistical independence between instrument and error term. This allows us to prove identification results and develop estimators for a structural function of interest when the instrument is discrete, and in particular binary. When the regressor of interest is also discrete with more mass points than the instrument, we state straightforward conditions under which the structural function is partially identified, and give modified assumptions which imply point identification. These stronger assumptions are shown to hold outside of a small set of conditional moments of the error term. Estimators for the identified set are given when the structural function is either partially or point identified. When the regressor is continuously distributed, we prove that if the instrument induces a sufficiently rich variation in the joint distribution of the regressor and error term then point identification of the structural function is still possible. This approach is relatively tractable, and under some standard conditions we demonstrate that our point identifying assumption holds on a topologically generic set of density functions for the joint distribution of regressor, error, and instrument. Our method also applies to a well-known nonparametric quantile regression framework, and we are able to state analogous point identification results in that context.

\address{Department of Economics, Northwestern University} \email{[email removed]}

Introduction

In this paper we consider the identification and estimation of a structural function $g$, which satisfies the relation

align*[align* omitted — 26 chars of source]

in the case where $Y$ and $X$ are observable random variables and $U$ is unobserved. We are concerned with the case where the regressor $X$ is possibly endogenous so that $\mathrm{E}\left[ U|X \right] \neq 0$. This complicates estimation of $g$, as one does not have the relation $\mathrm{E}\left[ Y|X \right] = g(X)$, whereby $g(X)$ is identified and may be estimated by a broad range of kernel estimators. The typical approach to this problem is to introduce an instrumental variable $W$ which satisfies:

align*[align* omitted — 47 chars of source]

When $W$ and $X$ are both continuous random variables this problem has been studied extensively and $g$ can be estimated as the solution to an ill-posed inverse problem, cf.\ HH2005 and NP2003 among others. However, instruments in the applied literature are often discrete with few mass points. For instance, AK1991 use season of birth to instrument in a linear model for the number of years of education. See FH2015 for several more instances in which researchers have used instruments $W$ with fewer mass points than $X$. Generally $g$ is not point identified if this is the case: see Proposition 1 of FH2015. A common solution is to assume that $g$ is a linear function of $X$ which allows it to be identified and estimated with e.g.\ a binary instrument. However, this assumption is rarely justified in practice.

We deal first with the case in which $X$ has $K$ mass points and $W$ has $2$ mass points, with $2 < K$. Most of the results that we display for such binary instruments may be readily generalized for instruments which take on more than two mass points, i.e.\ when $W$ has $L$ mass points with $L < K$. For instance, one can restrict attention to a subpopulation on which $W$ has two points of support. To more conveniently treat this binary instrument we write, without loss of generality, $W \in \{0,1\}$. Our approach to this problem is motivated by an observation from D2014 that “in specific econometric applications, the conditional mean assumption is typically established by arguing that the stronger independence assumption holds". In mathematical terms, one typically argues that $\mathrm{E}\left[ U|W \right] = 0$ by making the stronger claim that in fact $U \protect\mathpalette{\protect\independenT}{\perp} W$. If this is held to be true, then for any sequence of integrable functions $\{f_m\}$, one has $\mathrm{E}\left[ f_m(U) |W = 0 \right] = \mathrm{E}\left[ f_m(U) | W = 1 \right]$. This is similar to the observation made by P2017 that independence actually provides infinitely many moment restrictions which can be used to identify $g$. We show that when $X$ is discrete, only finitely many such moment conditions can be used to partially identify $g$. Later, when we consider a continuously distributed $X$, we use the full power of the independence assumption $U \protect\mathpalette{\protect\independenT}{\perp} W$.

When $X$ has $K$ mass points we take $f_m: u \mapsto u^m$ to be the function raising $U$ to the $m^\text{th}$ integer power. In Section (ref) we show that considering the moment restriction $\mathrm{E}\left[ U^m|W = 0 \right] = \mathrm{E}\left[ U^m|W = 1 \right]$ for $m = 1, \ldots , K$ partially identifies the function $g$, which can be considered as a vector in $\mathbb{R}^K$ over the support of $X$, under a light relevance condition for $X$. Partial identification in this case means that $g$ is identified up to a set of size at most $K!$. The first results of this section, Theorem (ref) and Corollary (ref), only require that the first $K$ moments of $U$ exist and be independent of $W$, which is a consequence of full independence $U \protect\mathpalette{\protect\independenT}{\perp} W$, provided that the moments of $U$ exist. In other words, we do not require full independence to prove these results. We also show in Theorem (ref) that, even under our relevance condition for $X$, it is possible that $g$ is not point identified even by the full independence assumption $U \protect\mathpalette{\protect\independenT}{\perp} W$. Indeed, we exhibit random variables $Y$ and $X$ and functions $g_1$ and $g_2$ such that $Y - g_1(X) \protect\mathpalette{\protect\independenT}{\perp} W$ and $Y - g_2(X) \protect\mathpalette{\protect\independenT}{\perp} W$. These counterexamples can be found under very stringent assumptions on the conditional distributions $X|W = 0$ and $X|W = 1$, suggesting that point identification is not possible unless restrictions are also placed upon the distribution of $U$. Section (ref) takes the additional step making assumptions on the joint distribution of the vector $(X,W,U)$. It describes a condition on the conditional moments of $U$ which implies point identification. Proposition (ref) demonstrates that this additional condition is fulfilled by the majority of joint distributions for $(X,W,U)$, in a sense to be described further on. Hence, imposing $\mathrm{E}\left[ U^{m}|W = 0 \right] = \mathrm{E}\left[ U^m | W = 1 \right]$ for $m = 1, \ldots, K+1$ is often enough to ensure that the $K$-vector $g$ is point identified.

In Section (ref) we use our moment independence relations to present an estimator $\widehat{g}$ of the function $g$ when $g$ is point identified. The estimator is shown to be almost surely consistent under light conditions, including an identification condition for $g$. We then show that a modified version of $\widehat{g}$, which we denote $\widetilde{g}$, is $\sqrt{n}$-consistent for $g$ under an invertibility condition for a $K \times K$ matrix, $V$, whose coefficients are polynomials in the conditional moments of $Y$ given $X$ and $W$. P2017 provided an efficient estimator for $g$ in the case that $X$ takes on finitely many values (as $g$ can be treated as a vector); however, the proposed estimator required the user to provide basis functions satisfying certain regularity and approximation properties for the distributions of $U$ and $W$. In contrast, our estimator $\widetilde{g}$ requires the user to solve a multivariate polynomial system of fixed dimension with coefficients that are directly calculated from the data (solving multivariate polynomial systems is a well-studied problem and most mathematics packages include toolkits for this application, which are reviewed in\ C2005) and then minimize an objective function over the solutions obtained, which usually constitute only a finite set in $\mathbb{R}^K$. Thus, it might be expected that our estimator comes with a reduced computational cost. Theorem (ref) establishes asymptotic normality for the quantity $\sqrt{n} (\widetilde{g} - g)$ with a limiting variance which can be estimated consistently from the data.

In Section (ref), we address the case where the structural function $g$ is perhaps not point identified but only partially identified, and $X$ is either discrete or continuously distributed. Estimation of the identified set requires some uniform convergence results from empirical process theory, and our assumptions in this section are mostly standard therein (see VW1996). Our main result in this section is Proposition (ref), which states some convergence properties of an estimator for the identified set of $g$ under some standard integrability assumptions on the covering numbers for the class to which $g$ belongs. It is shown that the estimator enjoys the property of (asymptotically) containing the identified set for $g$ and excluding any element not in the identified set.

Section (ref) extends our analysis of identification to the case where $X$ is continuously distributed. Existing results on identification in our setting, such as the work contained in D2014 and more recently C2019, have typically focused on local identification (that is, identification in some neighborhood of the true structural function $g$) because the operator which arises out of our independence assumption is nonlinear and difficult to characterize. The problem is similar to that faced when considering identification in quantile regression models like the one discussed in CH2005, and is typically addressed by making a nonlinear completeness assumption which is highly intractable. We address this issue by linearizing the operator offered to us by the independence assumption with one higher dimension (see Theorem (ref)), which allows for the application of the more traditional theory of linear maps. Our method allows us to give a sufficient condition for global identification which is comparable to a standard instrument completeness condition (Lemma (ref)), and to show that this condition holds on dense and topologically generic sets (Corollaries (ref) and (ref)) under certain commonplace assumptions. Our method also applies to identification in quantile regression models, and we prove a new result in \S (ref) which augments the work done in C2005. A discussion of our work and a comparison to existing results are included in \S (ref).

All proofs are located in our Appendix, Section (ref), which concludes.

Model, Discrete Case

Suppose that we have the additively separable model

align[align omitted — 41 chars of source]

and an instrument $W$ which is strongly exogenous in that either $\mathrm{E}\left[ U^m|W = 0 \right] = \mathrm{E}\left[ U^m | W = 1 \right]$ for certain values of $m$, to be specified our assumptions, or there is full statistical independence and $U \protect\mathpalette{\protect\independenT}{\perp} W$. Here we shall consider a subcase where $X$ and $W$ take on finitely many values; $X \in [K]$, where throughout we define $[K] \equiv \{1, \ldots, K\}$, and $W \in \{0,1\}$ so that we have a binary instrument. When the support of $W$ contains more than two points, one can always partition $\text{supp}\left( W \right)$ into two sets and define a new binary instrument to be an indicator of either partition element, adapting our assumptions to the constructed instrument. One helpful observation is full statistical independence provides us with a number of moment conditions which we can use to identify $g$:

align[align omitted — 126 chars of source]

for all $m \in \mathbb{N}$, provided that the moments exist. When the moments of $U$ determine its distribution, (ref) is in fact equivalent to independence of $U$ and $W$. As $X$ takes on finitely many values we may regard $g$ as a vector in $\mathbb{R}^K$. Thus, for the sake of brevity let $p_k(\ell) \equiv \mathrm{P}\left( X = k | W = \ell \right)$ and $g_k \equiv g(k)$. In this section, we will interchangeably use vector and function notation for functions $h$ over $\text{supp}\left( X \right) = [K]$ in this manner. Then, with the law of iterated expectations we may rewrite (ref) as

align[align omitted — 182 chars of source]

for all $m \in \mathbb{N}$ such that $\mathrm{E}\left[ U^m \right]$ exists. The system (ref) may be viewed as a degree $n$ multivariate polynomial in $g_1, \ldots , g_K$. Note for instance that \[\mathrm{E}\left[ (Y - g_k)^m|W = \ell, X = k \right] = \sum_{j = 0}^m \binom{m}{j} \mathrm{E}\left[ Y^j | W = \ell, X = k \right] (-g_k)^{m-j} \] Importantly, the coefficients of these polynomials (in particular, the conditional moments of $Y$ given the $W$ and $X$) may be estimated directly from the data.

Identification, Discrete Case

Partial Identification

In this section we demonstrate that one can use the system (ref) to obtain some conditions for the partial identification of $\{g_k\}$. The relatively straightforward assumptions we require as as follows:

asm\normalfont For any strict subset $J \subsetneq [K]$, $\mathrm{P}\left( X \in J | W = 0 \right) \neq \mathrm{P}\left( X \in J| W = 1 \right)$. In other words, $\sum_{k \in J} (p_k(0) - p_k(1))$ is nonvanishing in $J \subsetneq [K]$.

It is straightforward to see that identification of $g$ can fail when Assumption (ref) is not fulfilled, in the absence of any restrictions on the distribution of $U$, as we demonstrate in the following lemma (whose proof allows for a large degree of flexibility in the conditional distributions of $U$ given $X$ and $W$).

lemmaSuppose that $K \ge 2$ and for some subset $J \subsetneq K$ one has $\mathrm{P}\left( X \in J| W = 0 \right) = \mathrm{P}\left( X \in J | W = 1 \right)$. Then there exist conditional distributions for $U|_{X, W}$ such that the model (ref) and restrictions $U \protect\mathpalette{\protect\independenT}{\perp} W$, $\mathrm{E}\left[ U \right] = 0$ do not point identify $g$, and in fact the identified set for $g$ contains a continuum of elements.

Lemma (ref) illustrates that Assumption (ref) is fundamental to identification and partial identification of $g$. Note that the next assumption that we make is weaker than specifying full independence $U \protect\mathpalette{\protect\independenT}{\perp} W$ in the case that $|\mathrm{E}\left[ U^m \right]| < \infty$ for $m \in [K]$. It provides $K$ (nonlinear, polynomial) equations that aid in partially identifying the $K$-vector $g$:

asm\normalfont $\mathrm{E}\left[ U \right] = 0$ and $\mathrm{E}\left[ U^m| W \right] = \mathrm{E}\left[ U^m \right] < \infty$ for $m = 1, \ldots , K$.

With these we may prove the following:

theoremIf Assumptions (ref) and (ref) hold then the set of possible solutions to the system of equations formed by (ref) for $m = 1, \ldots , K$ and $\mathrm{E}\left[ U \right] = 0$ in $\mathbb{R}^K$ is finite. In particular, the identified set of vectors $\{g_k\}_{k=1}^K$ is finite.

A loose upper bound on the size of the identified set is then provided by the B\'{e}zout Bound of Algebraic Geometry.

corollary[B\'{e}zout's Theorem] If Assumptions (ref) and (ref) hold, then $\{g_k\}_{k=1}^K$ is identified up to a set of size at most $K!$. In particular, the number of possible solutions to (ref) for $n = 1, \ldots , K$ is bounded above by $K!$.

One immediately asks whether it is possible to extend Assumption (ref) on the conditional distributions of the regressor $X$ conditional on $W$ to obtain point identification of the vector $\{g_k\}_{k=1}^K$. It turns out that this is not possible if the probability vectors $\{p_k(0)\}$ and $\{p_k(1)\}$ are presumed to be strictly positive, as we show next.

theoremLet $K \ge 3$ and $p_k(0), p_k(1) > 0$ for all $k \in [K]$. Then for every pair of probability vectors $\{p_k(0)\}_{k=1}^K, \{p_k(1)\}_{k=1}^K$, there are distributions for $Y$ and functions $g_1 \neq g_2: [K] \rightarrow \mathbb{R}$ such that \begin{align*} &Y - g_1(X) \mathpalette{\independenT}{\perp} W \\ &Y - g_2(X) \mathpalette{\independenT}{\perp} W. \end{align*}

Conditions for Point Identification when $X$ is Discrete

Theorem (ref) implies that sometimes it is not possible to point identify the function $g: [K] \rightarrow \mathbb{R}$ with only assumptions on the joint distribution of $X$ and $W$, such as Assumptions (ref) and (ref). In this section we establish conditions which allow $g$ to be point identified and attempt to show that these conditions are fulfilled in a wide variety of settings. The general strategy is as follows: first, use the multivariate polynomial equations (with variables $\{g_k\}_{k \in [K]}$) stated in (ref), for $m \in [K]$ as well as the normalizing relation $\mathrm{E}\left[ U \right] = 0$ to characterize the identified set down to a finite number of points in $\mathbb{R}^K$. Then, use the polynomial equation supplied by (ref) with $m = K+1$ to show that generally only one of the points in the aforementioned finite set satisfies $\mathrm{E}\left[ U^{K+1}| W = 0 \right] = \mathrm{E}\left[ U^{K+1}|W = 1 \right]$. The motivation of this section is: independence of the first $K$ moments of $U$ from $W$ typically is enough to obtain partial identification of the vector $g$ (Theorem (ref)), whereas independence of the first $K+1$ moments of $U$ is typically enough to point identify $g$.

To begin, note that for a particular vector $h \in \mathbb{R}^K$ which satisfies (ref) for some $m \in \mathbb{N}$ we may combine equations (ref) and (ref) to write equivalently:

align[align omitted — 185 chars of source]

where we have denoted $\delta_k \equiv h_k - g_k$. Fixing a joint distribution for $X, W$ which satisfies Assumption (ref), Theorem (ref) and Corollary (ref) imply that the size of the set of possible solutions to (ref) for $m = 1, \ldots , K$ is bounded above by $K!$ for every fixed set of $U$-moment vectors of the form:

align[align omitted — 142 chars of source]

Note that dependence on $\mathrm{E}\left[ U^K| W = \ell, X = K \right]$ is suppressed because when (ref) is expanded with $m = K$, the terms involving $K^\text{th}$ moments of $U$ vanish by moment independence. A useful observation is that $S_{K-1}$, which may in this context be called a vector of lower moments of $U$, only depends on the condition moments of $U^m$ for $m \le K-1$ (not $K$).

Denote the finite set of possible $\delta$ vectors permitted under Theorem (ref) as a function of $S_{K-1}$ by $A(S_{K-1})$, i.e.\ define

align*[align* omitted — 170 chars of source]

where for notational ease we use the convention $\mathrm{E}\left[ h - g \right] = \mathrm{E}\left[ h(X) - g(X) \right]$. Fixing $S_{K-1}$ and thus $A(S_{K-1})$, we note that in order to satisfy (ref) for $m = K+1$ we must have the relation:

align[align omitted — 205 chars of source]

where \[ P \left( S_{K-1}, \delta_k\right) \equiv K^{-1}\sum_{k=1}^K \left[ \sum_{j = 0}^{K-1} \binom{K+1}{j} \delta_k^{K+1 - j} \left(p_k(0)\mathrm{E}\left[ U^j|W = 0, X = k \right] - p_k(1) \mathrm{E}\left[ U^j|W = 1, X = k \right]\right) \right] \] is the term determined by $S_{K-1}$ and $\delta \in A(S_{k-1})$ in the binomial expansion of (ref). Suppose that the system formed by (ref) for $m = 1 , \ldots , K$ and the normalization relation $\mathrm{E}\left[ U \right] = 0$ does not identify the function $g$, so that $A(S_{k-1}) \supsetneq \{0\}$. If in addition the equation (ref) for $m = K+1$ does not identify $g$ then there exists $\delta \neq 0 \in A(S_{K-1}) \subset \mathbb{R}^K$ such that (ref) holds: that is, the existence of nontrivial $\delta$ implies that a specific linear relation must hold among the $K^\text{th}$ moments $\mathrm{E}\left[ U^K|W = 1, X = k \right]$ whose coefficients are determined by the lower moments of $U$ which are contained in $S_{K-1}$. This leads to the following observation, that point identification is generically fulfilled when one considers the $K^\text{th}$ conditional moments of $U$ in a sense that we make precise below.

The main result of this section is Proposition (ref), which distills our remarks so far into a genericity result on conditions for point identification. Below we give a brief introduction to its result.

A set $T_0 \subset T$ is commonly said to be meagre and its complement $T\setminus T_0$ generic if (topologically) $T_0$ is the countable union of relatively closed nowhere dense sets, or (measure-theoretically) there exists some well-behaved measure $\mu$ on $T$ such that $\mu(T_0) = 0$ and $\mu(T) > 0$. Fix a joint distribution for $X,W$ satisfying Assumptions (ref) and (ref) and let $S_{K-1}$ denote the vector of lower moments in (ref), assuming that the moments $\mathrm{E}\left[ U^m \right]$ exist for $m \in [K+1]$. Consider the set $T$ which consists of all possible vectors of $K^\text{th}$ moments of $U$ that extend $S_{K-1}$ and the moment independence assumption $\mathrm{E}\left[ U^K|W = 0 \right] = \mathrm{E}\left[ U^K| W = 1 \right]$. In imposing the first requirement, we mean that there exist actual probability distributions with lower moments in $S_{K-1}$ and $K^\text{th}$ moments in $T$. Formally, $T$ is a set of $2K$-vectors that verifies

align[align omitted — 458 chars of source]

$T$ is a subset of a hyperplane of $\mathbb{R}^{2K}$ corresponding with the linear restriction in the first line of (ref). Furthermore, define $T_0$ to be the subset of moment vectors in $T$ under which the restrictions $\mathrm{E}\left[ U^m|W = 0 \right] = \mathrm{E}\left[ U^m| W = 1 \right]$ for $m \in [K+1]$ and $\mathrm{E}\left[ U \right] = 0$ do not point identify $g$, i.e.\ those for which there exist a nonzero $\delta \in A(S_{K-1})$ such that $\mathrm{E}\left[ (U + \delta(X))^m|W = 0 \right] = \mathrm{E}\left[ (U + \delta(X))^m|W = 1 \right]$ for $m \in [K+1]$ and $\mathrm{E}\left[ \delta(X) \right] = 0$. Note that such a $\delta$ satisfies (ref) for $m \in [K+1]$ and also $\mathrm{E}\left[ \delta(X) \right] = 0$. The first finding of Proposition (ref) is that $T$ is “large" in the $(2K-1)$-dimensional hyperplane given by the linear restriction in (ref) in that it has non-empty interior in this hyperplane and is in fact assigned infinite measure by the unique translation-invariant Haar measure (completely analogous to Lebesgue measure) on that plane. The second is that $T_0$ is “small" in $T$: it consists of at most the intersection of finitely many $(2K-2)$-dimensional hyperplanes with $T$, which are each assigned Haar measure $0$. Hence, in light of both topological and measure theoretical conditions, point identification of $g$ can be regarded as generic in terms of the conditional $K^\text{th}$ moments of $U$.

propositionSuppose $\mathrm{E}\left[ |U|^{m} \right] < \infty$ and $\mathrm{E}\left[ U^m|W = \ell \right] = \mathrm{E}\left[ U^m \right]$ for $m \le K+1$ and that for all $\ell, k$, $U|_{W = \ell, X = k}$ has at least $\left\lfloor (K-1)/2 \right\rfloor+1$ points in its support. Suppose Assumptions (ref) and (ref) hold and let $S_{K-1}$ be defined as in (ref). Let $T \subset \mathbb{R}^{2K}$ denote the set of possible vectors $\{ \mathrm{E}\left[ U^K|W = \ell, X = k \right]: \ell \in \{0,1\}, X \in [K]\}$ which satisfy Assumption (ref) and extend $S_{K-1}$. Let $T_0\subset T$ denote the subset of possible $K^\text{th}$ moment vectors for which $\mathrm{E}\left[ U^m| W = 0 \right] = \mathrm{E}\left[ U^m|W = 1 \right]$ for $m \in [K+1]$ and $\mathrm{E}\left[ U \right] = 0$ does not point identify $g$ (i.e.\ there exists $0 \neq \delta \in A(S_{K-1})$ which satisfies (ref) for $m \in [K+1]$, and $\mathrm{E}\left[ \delta(X) \right] = 0$). Then $T_0$ is contained in a finite intersection of translated subspaces of strictly lower dimension than $T$, and $T$ has nonempty interior in the hyperplane implied by the linear restriction (ref). In particular, the standard $2K$-dimensional Haar measure $\mu$ supported on the hyperplane corresponding with the linear restriction $\mathrm{E}\left[ U^K| W = 0 \right] = \mathrm{E}\left[ U^K|W = 1 \right]$ satisfies $\mu(T) =\infty $ but $\mu(T_0) = 0$.

\noindentRemark: The requirement that the conditional distributions of $U$ have a lower bounded number of points of support is implied if $U$ has continuous conditional distributions. This requirement is made to ensure that we can flexibly supply $U$ with higher order moments.

As the elements $\delta \in A(S_{K-1})$ are solutions to nonlinear polynomial equations and $T_0$ is furthermore a function of those elements, there is no closed form expression for $T_0$ in terms of the $S_{K-1}$, which complicates the question of whether the structural function $g$ is point identified in any particular instance. However, Proposition (ref) assures us that there are conditions which guarantee that the identified set is a singleton, and that these conditions are fulfilled quite often when, conditional on the vector of lower moments $S_{K-1}$, the $K^\text{th}$ conditional moment vectors $\{\mathrm{E}\left[ U^K|W = \ell, X = k \right]: \ell \in \{0,1\}, k \in [K] \}$ arise from a prior distribution which is continuous with respect to Lebesgue measure on the hyperplane in $\mathbb{R}^{2K}$ that is consistent with the fundamental linear restriction in (ref). In fact, Fubini's theorem and Proposition (ref) clearly imply that, if this is the case, $g$ is point identified with prior probability $1$. Further on, we provide complementary conditions for point identification of the function $g$ which involve the joint distributions of $X$ and $U$ given $W$, and show that these conditions are topologically generic (Proposition (ref)). These conditions are specialized to the case where $X$ is continuously distributed, but they may readily be adapted to the discrete case heretofore considered.

Estimation when $g$ is Point Identified

In this section we adapt our identification result to propose estimators of the function $g$ in (ref). We begin by making the observation that the polynomial system expressed in (ref), associated with the moment independence condition $\mathrm{E}\left[ U^m|W \right] = \mathrm{E}\left[ U^m \right]$ for $m = 1, \ldots , K+1$ and the mean independence assumption $\mathrm{E}\left[ U \right] = 0$, point identify the function $g$ under conditions which are enumerated in Proposition (ref) and shown therein to be commonplace. Moreover, these same polynomials can be estimated directly from the data.

We make the following standard and straightforward assumption to ensure that the necessary laws of large numbers may be invoked. All asymptotic statements (e.g.\ almost surely, eventually) are made as the sample size $n$ goes to infinity.

asm$(X_i, W_i, Y_i)_{i=1}^n$ is an iid sample of $X, W, Y$. For all $\ell \in \{0,1\}$ and $k \in [K]$, $\mathrm{E}\left[ |Y|^{K+1} |W = \ell, X =k \right] < \infty$. $\mathrm{P}\left( W = 0 \right), \mathrm{P}\left( W = 1 \right) > 0$.

Recall that our identification result Theorem (ref) centered on a system of polynomial equations (ref). The central insight of this result was that the solution to the polynomial system could be characterized completely by finitely many of the equations, rather than the infinitely many equations utilized by P2017. Note that GMM is typically asymptotically efficient only if the true, efficient score function belongs to the closure of linear manifold spanned by the moment conditions, which typically requires the use of an asymptotically diverging number of moments (see CF2014). Therefore, our approach, which only uses a bounded number of moment conditions, is not expected to be asymptotically efficient. The advantage of our proposed methods is that they are straightforward to implement, and involve mainly solutions of systems of polynomial equations, around which a significant theory has been developed.

Our estimation strategy is based on estimating an empirical analogue to the system (ref) and solving for the function $g$ on $[K] = \mathrm{supp}( X)$. As shorthand, we continue to variably denote $g$ as a function $g: [K] \rightarrow \mathbb{R}$ and as a vector in $\mathbb{R}^K$. Hence, we may now define $\Gamma : \mathbb{R}^K \rightarrow \mathbb{R}^{K+2}$ to be the following vector valued function:

align*[align* omitted — 104 chars of source]

where $P_0(h) \equiv \sum_{k=1}^K p_k(0) \mathrm{E}\left[ Y - h_k|W = 0, X = k \right]$ and for $m \ge 1$, $P_m(h)$ is given as the left side of (ref), which is to say \[ P_m(h) \equiv \mathrm{E}\left[ (Y - h(X))^m | W = 0 \right] - \mathrm{E}\left[ (Y - h(X))^m| W = 1 \right]. \] Note that the Binomial theorem implies that for $m \ge 1$ we have

align[align omitted — 458 chars of source]

where we make the denotations

align[align omitted — 479 chars of source]

that is, $P_m(h)$ is indeed an $m^\text{th}$ degree multivariate polynomial in the coordinates of $h \in \mathbb{R}^K$.

Recall that under Assumptions (ref) and (ref), Corollary (ref) implies that $\Gamma$ attains at most $K!$ zeros in $\mathbb{R}^K$, one of which corresponds to the true function $g$ represented in (ref). Consider estimation of $\Gamma$ with the empirical function $\widehat{\Gamma}$:

align*[align* omitted — 145 chars of source]

where for $j = 0, \ldots ,K+1$, $\widehat{P}_j$ is the plug-in estimator of $P_j$ with coefficients $Q_{j,\ell,k}$ for $j \le K+1$ replaced with the plug-in estimator:

align*[align* omitted — 272 chars of source]

(where $\widehat{C}_{j,\ell,k}$ is defined as indicated). Letting $\mathrm{E}_{n}\left[ \cdot \right]$ denote expectation with respect to the empirical measure $\frac{1}{n}\sum_{i=1}^n \delta_{X_i}$, it is straightforward to verify that \[ \widehat{P}_m(h) = \mathrm{E}_{n}\left[ (Y - h(X))^m | W = 0 \right] - \mathrm{E}_{n}\left[ (Y - h(X))^m|W = 1 \right]. \] The resulting estimator $\widehat{\Gamma}$ enjoys almost sure uniform convergence to $\Gamma$ as a result of the type of functions (polynomial) under consideration.

lemmaUnder Assumption (ref), $\widehat{\Gamma}(h)$ converges uniformly almost surely to $\Gamma(h)$ over any bounded subset of $\mathbb{R}^K$.

We now consider the problem of estimating $g$. We supplant Assumptions (ref) and (ref) with an identification assumption on $g$. Namely, we suppose that the system of polynomials $P_0, \ldots, P_{K+1}$ uniquely identify $g$. Conditions for this to be the case are discussed in Proposition (ref), and are shown to hold outside of a very small (Haar measure $0$) subset of cases. Some alternate conditions which are sufficient for point identification may be drawn from the results of Section (ref), and in particular Theorem (ref) and its corollaries.

asm[Identification] \normalfont The vector $g \in \mathbb{R}^K$ uniquely satisfies $\Gamma(g) = 0$. Moreover, $\left\lVert g\right\rVert < R$ for some known constant $R$ (where $\left\lVert \cdot\right\rVert$ denotes the Euclidean norm).

We now define a preliminary estimator of $g$ as (unsurprisingly)

align*[align* omitted — 169 chars of source]

By Lemma (ref) and a standard extremum estimation argument, the following is then true:

lemmaUnder Assumptions (ref) and (ref), one has $\widehat{g} \overset{\mathrm{a.s.}}{\rightarrow} g$.

Standard arguments which establish the asymptotic normality of extremum estimators cannot be applied in situ to the estimator $\widehat{g}$ because the components of $\Gamma$ are not population moments but linear combinations of conditional moments.

We now turn our attention to a polynomial-based estimator of $g$ which has an asymptotic normality property. We denote this estimator by $\widetilde{g}$. The estimator $\widetilde{g}$ is estimated in two stages: first, solve a multivariate polynomial system of equations $\widehat{\Lambda}(h)$ for $h$ within the parameter set, and then minimizing the objective function $\widehat{\Gamma}(h)$ over the obtained solution set. The idea of the proof of asymptotic normality is to apply the implicit function theorem and make a Delta-method argument on the estimator. To this end, define the vector valued function $\Lambda: \mathbb{R}^K \rightarrow \mathbb{R}^K$ by

align*[align* omitted — 207 chars of source]

and for some fixed $R < \infty$, define the zero set $\mathcal{Z}_R$ of any function $\gamma: \mathbb{R}^K \rightarrow \mathbb{R}^K$ to be:

align*[align* omitted — 331 chars of source]

For the $R$ used in Assumption (ref), we now define our estimator $\widetilde{g}$ to be

align*[align* omitted — 155 chars of source]
remark\normalfont By Bernstein's theorem (see\ C2005 \S5, Theorem 5.4), a polynomial system of the form $\Gamma(h): \mathbb{R}^K \rightarrow \mathbb{R}^K$ generically admits $K! $ solutions in $(\mathbb{C}\setminus\{0\})^K$.\footnote{This may be calculated using standard formulae for mixed volume in terms of normal $K$-dimensional volume, and the formula for the volume of a typical simplex in $\mathbb{R}^K$ formed by the $K$ elementary vectors} Generic in this case means that there is a nonzero polynomial in the coefficients of the polynomials $P_0, \ldots , P_{K-1}$ such that the property holds whenever the nonzero polynomial is nonvanishing for $P_0, \ldots , P_{K-1}$. For any fixed $d \in \mathbb{N}$, it can be shown via induction on $d$ that the set of solutions (variety) for a multivariate polynomial on $\mathbb{R}^d$ has zero Lebesgue measure on $\mathbb{R}^d$. Hence, under the assumption that the coefficients of $P_1, \ldots , P_K$ avoid this zero-measure subset of $\mathbb{R}^d$, $d$ indicating the number of coefficients in those polynomials, then $\mathcal{Z}_R(\widehat{\Lambda})$ is nonempty for $R$ large enough. In our case this suggests that typically, upon solving for $\Lambda$, the user should obtain at most $K!$ solutions.

Application of the implicit function theorem requires that we assume an invertibility condition on the Jacobian matrix of $\Lambda$.

asm\normalfont The $K \times K$ matrix $V$ whose $(m,k)^\text{th}$ coordinate is given by \begin{align*} V_{m,k} &= \frac{\partial P_{m-1}}{\partial h_k} \Big|_{h = g} \\ & = \left\{ \begin{array}{ll} \sum_{j=0}^{m-2} \binom{m}{j} Q_{j,k,0} (m-j ) (-g_k)^{m-j - 1} & for m = 1\\ \sum_{j=0}^{m-2} \binom{m}{j}\left( Q_{j,k,1} - Q_{j,k,0} \right) (m-j ) (-g_k)^{m-j - 1} & for m \ge 2 \end{array} \right. \end{align*} for $m, k \in [K]$ is invertible.

If we had instead put $V_{m,k} = \frac{\partial P_{m}}{\partial h_k}\Big|_{h = g}$ in Assumption (ref), and thus omitted $P_0$, then $V$ would not be invertible; this stems from the fact that if $P_m(g) = 0$ for $m \ge 1$, then also $P_m(g + c \mathbf{1}_K) = 0$ for all $c \in \mathbb{R}$ (where $\mathbf{1}_K$ is the $K$-dimensional vector $(1,\ldots , 1)'$).

Letting $V$ denote the matrix of partial derivatives indicated in Assumption (ref), define the functions $\Psi_m: \mathbb{R}^{K} \times \mathbb{R}^{2K^2 + 2}, \, m = 0, \ldots , K-1$ by

align*[align* omitted — 355 chars of source]

where we define the indexing bijection $\iota: \{0 , \ldots , K - 1\} \times \{0,1\} \times \{1, \ldots , K\} \rightarrow \{1, \ldots , 2K^2 \}$ by $\iota(j,\ell,k) = 2Kj + 2k + \ell - 1$. Although the definition of $\Psi_m$ introduces substantial notational difficulty, comparison with (ref) reveals that $\Psi_m(v,w)$ just becomes $P_m$ with the proper choice of coefficient vector $w$. The difference from (ref) is that the definition of $\Psi_m$ allows the coefficients of the latter to differ from $Q_{j,k,\ell}$, as should be expected when those coefficients are estimated from data.

To continue, let $w^* \in \mathbb{R}^{2K^2 + 2}$ denote the vector satisfying $w_{\iota(j,\ell,k)}^* = \mathrm{E}\left[ Y^j \mathbf{1}_{W = \ell, X = k } \right]$ for $j \in \{0, \ldots, K-1\}$, $k \in [K]$, and $\ell \in \{0,1\}$, and also $w^*_{2K^2 + 1} = \mathrm{P}\left( W = 0 \right), w^*_{2K^2 + 2} = \mathrm{P}\left( W = 1 \right)$. By definition, $w^*$ is precisely the “correct" choice of $w$ which equates $\Psi_m(\cdot, w)$ with $P_m(\cdot)$. Define $\Psi: \mathbb{R}^{K} \times \mathbb{R}^{2K^2 + 2}$ to be the vector valued function

align*[align* omitted — 96 chars of source]

and let $\Delta = D_w \Psi(h, w^*)|_{h = g}$ denote its Jacobian matrix (with respect to $w$) evaluated at $(h,w) = (g, w^*)$. Finally, let the $(2K^2 + 2) \times (2K^2 + 2)$ matrix $\Omega$ by

align*[align* omitted — 765 chars of source]

to be the covariance matrix of the random vector

align*[align* omitted — 279 chars of source]

With our invertibility assumption, we may now state the following central limit theorem for the estimator $\widetilde{g}$:

theoremUnder Assumptions (ref)---(ref), the convergence \[ \sqrt{n} (\widetilde{g} - g) \overset{\mathrm{d}}{\rightarrow} N(0, \left((\mathrm{D}_w \omega(w^*))\, \Omega \, (\mathrm{D}_w \omega(w^*)) '\right)^{-1/2} ) \] holds with $D_w \omega(w^*) = -V^{-1} \Delta$. Moreover, if $\mathrm{E}\left[ Y^{2K - 2} \mathbf{1}_{ W = \ell, X = k} \right] < \infty$ for all $\ell \in \{0,1\}, X \in [K]$, then there exist consistent estimators $(\widehat{\mathrm{D}_w \omega(w^*)})$ and $\widehat{\Omega}$, for $\mathrm{D}_w(\omega(w^*))$ and $\Omega$, respectively.

Estimation when $g$ is Partially Identified

In this section we suppose that $g$ is possibly only partially identified. In other words, we consider the case where $g \in \mathcal{H}$ for some space of continuous functions $\mathcal{H}$ over the support of $X$, $\mathcal{X} = \supp X$. We let $X$ be either discrete or continuously distributed in this section. Let $\left\lVert \cdot\right\rVert_\infty$ indicate the supremum norm over $\mathcal{H}$, i.e.\ $\left\lVert h\right\rVert_\infty = \sup_{x \in \mathcal{X}} |h(x)|$. Assume that

asm\normalfont The standing model $Y = g(X) + U$ (ref) holds with $\supp W = \{0,1\}$, and \begin{enumerate} • $U \protect\mathpalette{\protect\independenT}{\perp} W$ and $\mathrm{E}\left[ U \right] = 0$$\mathrm{P}\left( W = 0 \right) \mathrm{P}\left( W = 1 \right) \neq 0$$g \in \mathcal{H}$ and $\mathcal{H}$ is bounded in $\left\lVert \cdot\right\rVert_\infty$$\limsup_{m \rightarrow \infty} \left( \frac{\mathrm{E}\left[ U^m \right]}{m!}\right)^{1/m} < \infty$. \end{enumerate}

Parts (1) and (2) of the assumption is familiar, (2) restricts $g$ to lie in our class $\mathcal{H}$ (to be defined shortly) and (3) is a regularity condition on the moments of $U$ which is satisfied by probability distributions whose Fourier transform exists in a complex neighborhood of the origin, or alternately whose Laplace transform exists in a neighborhood of the origin. Under Assumption (ref), the partially identified set for $g$ is:

align*[align* omitted — 176 chars of source]

A common object of interest is not the entire function $g$ but the value of $T g$, where $T$ is some complex-valued functional defined on $\mathcal{H}$. Then the partially identified set for $Lg$ is $L \mathcal{H}_0 = \{Lh: h \in \mathcal{H}_0\}$. Any set-valued estimator for $\mathcal{H}_0$ may readily extended to an estimator for $T\mathcal{H}_0$ by considering its image under $L$.

Our first objective is to more tractably characterize the identified set $\mathcal{H}_0$, which we do in the following lemma:

lemmaUnder Assumption (ref), the characteristic function $\mathrm{E}\left[ e^{it(Y - h(X))} \right]$ is holomorphic in a neighborhood of the origin in $\mathbb{C}$ for any $h \in \mathcal{H}$, and \begin{align*} \mathcal{H}_0 &= \left\{ h \in \mathcal{H}: \, \mathrm{E}\left[ Y - h(X) \right] = 0 and \sup_{t \in [0,1]} \left| \mathrm{E}\left[ e^{it(Y - h(X))}| W = 0 \right] - \mathrm{E}\left[ e^{it (Y - h(X)}|W = 1 \right]\right|= 0 \right\}. \end{align*}

Lemma (ref) provides a convenient characterization of $\mathcal{H}_0$ which we exploit. Notably, we may restrict attention to only $t$ within a compact subset of $\mathbb{R}$. To employ some of the tools of empirical process theory, we now make the following assumption on the class $\mathcal{H}$.

asm\normalfont The covering numbers $N(\varepsilon, \mathcal{H}, \left\lVert \cdot\right\rVert_\infty)$ satisfy the integrability condition \begin{align*} \int_0^\infty \sqrt{ \log N(\varepsilon, \mathcal{H}, \left\lVert \cdot\right\rVert_\infty )} \, \mathrm{d} \varepsilon < \infty. \end{align*}

Assumption (ref) is a uniform bound on the entropy of $\mathcal{H}$. It is satisfied if, for example, $\mathcal{X}$ is a bounded and convex subset of $\mathbb{R}^d$ with nonempty interior, and

align[align omitted — 240 chars of source]

for some constants $M$ and $\alpha > d/2$, where $k = (k_1, \ldots , k_d)$ is a multi-index of $d$ integers, $|k| = \sum_{j = 1}^d k_j$, and $D^k \equiv \frac{\partial^k}{\partial x_1^{k_1} \cdots \partial x_d^{k_d}}$ (see VW1996, Theorem 2.7.1, which is more general and applies to H\"{o}lder-continuous derivatives as well as Lipschitz in the case where $\alpha$ is not an integer). In fact, (ref) implies that one has $\log N(\varepsilon, \mathcal{H}, \left\lVert \cdot\right\rVert_\infty) \le K \left( \frac{1}{\varepsilon} \right)^{d/\alpha}$ for some constant $K$. S2012 gives conditions, in particular restrictions on the weighted Sobolev norms of functions in $\mathcal{H}$, which guarantee that this is the case. It is clear that if $\mathcal{X}$ is finite then Assumption (ref) is immediately satisfied as long as $\mathcal{H}$ is bounded.

Define the classes of functions from the sample space $\mathbb{R} \times \mathcal{X} \times \{0,1\}$ to $\mathbb{C}$ which we will subsequently consider by

align*[align* omitted — 256 chars of source]

When the uniform entropy condition holds for $\mathcal{H}$ under $\left\lVert \cdot\right\rVert_\infty$, we can infer that the class $\mathcal{F}$ obeys a Donsker theorem, in the sense that

lemmaLet Assumptions (ref) and (ref) hold and let $P$ denote the distribution of $(Y,X,U,W)$. Then the classes $\mathcal{E}$ and $\mathcal{F}$ are Glivenko-Cantelli and $P$-Donsker.

As we are dealing with classes of potentially complex valued random functions, we designate an empirical process $\mathbb{G}_n = \sqrt{n} ( \mathbb{P}_n - \mathbb{P})$ indexed by a class, say $\mathcal{F}$, $P$-Donsker if it converges in the weak sense to a complex valued Gaussian process $\mathbb{G}$ (that is, a process whose marginals are all joint complex Gaussian distributions) which takes on values in $L^\infty(\mathcal{F})$. All of the standard results from empirical process theory apply to complex-valued functions, which can be readily seen by limiting consideration to the real and complex parts of functions in $\mathcal{F}$ individually, and then noting that the union of two Glivenko-Cantelli classes is clearly Glivenko-Cantelli, whilst the union of two $P$-Donsker classes is also $P$-Donsker (K2008, Corollary 9.31).

From Lemma (ref), conclude (continuing to use the convention that $0/0 = 0$ when calculating conditional expectations with respect to empirical measure $\mathrm{E}_{n}\left[ \cdot \right]$):

propositionLet $\mathcal{F}_0 \equiv [0,1] \times \mathcal{H}$. Then for all $t \in [0,1]$ and $h \in \mathcal{H}$, one has the convergence \begin{align} &\sqrt{n} \Bigg[\mathrm{E}_{n}\left[ e^{it(Y - h(X))}|W = 0 \right] - \mathrm{E}_{n}\left[ e^{it(Y - h(X))}|W = 1 \right] \\ & \qquad \qquad \qquad - \mathrm{E}\left[ e^{it(Y - h(X))}|W = 0 \right] + \mathrm{E}\left[ e^{it(Y - h(X))}|W = 1 \right] \Bigg] \rightsquigarrow \mathbb{D}(t,h), \nonumber \end{align} where $\mathbb{D}$ is a tight mean-zero Gaussian process in $\ell^\infty(\mathcal{F}_0)$.

As a consequence of Lemma (ref) and Proposition (ref), we are able to state an estimator for the identified set $\mathcal{H}_0$ by substituting for moment equalities with their finite sample analogues. Let $\eta_n \rightarrow 0$ be a positive sequence, and define

align*[align* omitted — 299 chars of source]

where $\mathrm{E}_{n}\left[ \cdot \right]$ denotes the sample mean. Then we have the following convergence result:

propositionUnder Assumptions (ref) and (ref), if $\eta_n \rightarrow 0$ and $\eta_n \sqrt{n} \rightarrow \infty$ then for all sequences $(\alpha_n)_{n \in \mathbb{N}}$ satisfying $\liminf_{n \rightarrow \infty} \frac{\alpha_n}{2 \eta_n } > 1$, \begin{align*} &\mathrm{P}\left( \mathcal{H}_0 \subset \widehat{\mathcal{H}}_n \right) \rightarrow 1 and \mathrm{P}\left( \mathcal{H}_{\alpha_n} \cap \widehat{\mathcal{H}}_n = \emptyset \right) \rightarrow 1, \end{align*} where \[\mathcal{H}_\alpha = \left\{ h \in \mathcal{H}: \, \mathrm{E}\left[ Y - h(X) \right] \ge \alpha \text{ or } \sup_{t \in [0,1]} \left| \mathrm{E}\left[ e^{it(Y - h(X))}| W = 0 \right] - \mathrm{E}\left[ e^{it (Y - h(X))}|W = 1 \right]\right| \ge \alpha \right\}. \] Moreover, if $\eta_n \propto n^{-\gamma}$ and $\log N(\varepsilon, \mathcal{F}, \left\lVert \cdot\right\rVert_\infty) \le K \left( \frac{1}{\varepsilon} \right)^\omega$ for some $\gamma \in (0,1/2)$ and $\omega \in (0,2)$ then \begin{align*} \mathbf{1}_{\mathcal{H}_0 \subset \widehat{\mathcal{H}}_n} \overset{\mathrm{a.s.}}{\rightarrow} 1 and \mathbf{1}_{\mathcal{H}_{\alpha_n} \cap \widehat{\mathcal{H}}_n = \emptyset} \overset{\mathrm{a.s.}}{\rightarrow} 1. \end{align*}

The set $\widehat{\mathcal{H}}_n$ is thus a consistent estimator for $\mathcal{H}_0$ under the stated assumptions. Given that the bound $K!$ stated on the identified set in Theorem (ref) is large, the reader may find it practical even in the discrete case to reduce consideration of $\mathcal{H}_0$ down to its image under a linear map $L$ as previously discussed, and then estimate it by $L \widehat{\mathcal{H}}_n$ or a perhaps an interval containing that set.

Identification Results when $X$ is a Continuous Random variable

We now turn to the case where $X$ is a continuous random variable, considering first a scalar $X$ and binary instrument $W$ satisfying our standing model (ref) with $U \protect\mathpalette{\protect\independenT}{\perp} W$. We state conditions which allow the function $g$ to be point identified by the joint distribution of observables $(Y,X, W)$. Our main strategy here is to linearize the nonlinear operator imposed by the independence restriction, which allows us to more tractably state some sufficient conditions for identification. This avoids some of the difficulties encountered when dealing with strictly nonlinear operators (see for instance C2019, \S 2, and CH2005). Our approach, which is encapsulated in Theorem (ref), allows us to construct examples in which $g$ is point identified and then to prove that the conditions which enable point identification hold on a topologically generic set (Proposition (ref)) of density functions. This result reinforces Proposition (ref) in suggesting that our standing model with independence assumption might typically be enough to point identify $g$. Interestingly, our method generalizes to quantile regression models, and we are able to give a new identification result in that setting which extends the result of CH2005 (\S (ref)). A discussion of our results is given in \S (ref).

The main assumption that we will make in this section (Assumption (ref)) is in the familiar form of a restriction on the kernel of a linear operator, which is determined by the joint distribution of observables. To begin, consider the following base assumptions, which are regularity conditions on the joint distribution of $X$ and $U$:

asm\normalfont $X$ is supported on a compact set $\mathcal{X} \subset \mathbb{R}^{d_X}$, $d_X \in \mathbb{N}$. Moreover, $g \in C(\mathcal{X})$ and $\left\lVert g\right\rVert_\infty =\sup_{x \in \mathcal{X}}|g(x)| \le B$ for some known constant $B \le \infty$.
asm\normalfont The joint distribution of $(X,U)|W = w$ is continuous with respect to (product) Lebesgue measure $\lambda$ on $\mathcal{X} \times \mathbb{R}$ with continuous densities $f_\ell(x,u)\equiv f_{W = \ell}(x,u)$, for $\ell \in \{0,1\}$.

Assumption (ref) is used mainly to provide notational simplicity and to ensure the existence and convergence of certain integrals. One could most expediently address the case of an $X$ with unbounded support by considering the pushforward of $X$ by a continuous and invertible mapping $\psi$, say the probability integral transform. If $\psi$ maps $\text{supp}\left( X \right)$ into a bounded subset of $\mathcal{X}\subset \mathbb{R}^{d_X}$ then $\psi(X)$ may be considered instead of $X$ and assumptions on $g$ may be presumed to hold for $g \circ \psi$. Our notation and assumptions are for $X$ continuously distributed with respect to Lebesgue measure, but this regularity condition could readily be dropped in favor of another dominating measure (in the discrete $X$ case, consider counting measure). Note that Assumption (ref) allows the choice $B = \infty$, which places no restrictions on $g$ beyond continuity. Assumption (ref) is standard.

The constant $B$, which is the provided bound on the sup-norm of $g$, is fundamental in what follows. For any $t \in \mathbb{R}$ denote by $f_w^t(x,u)$ the transformed function $f_w(x, t + u)$. Define now the operator $T$ on $L^2(\mathcal{X} \times \mathbb{R}) $ by

align*[align* omitted — 374 chars of source]

where the last line follows by substituting $v = t + u$ in the definition of $(Th)(t)$. The most important aspect of $T$ is that it takes the familiar form of a linear operator between Banach spaces.

To develop some insight on $T$, we characterize the range of our operator $T$ in the following lemma. All $L^p$ spaces are taken with respect to Lebesgue measure unless otherwise noted, and spaces of continuous functions $C(\cdot)$ are equipped naturally with the sup-norm, which makes them complete metric spaces.

lemmaSuppose Assumptions (ref) and (ref) hold. Then the linear mapping $T:L^\infty(\mathcal{X} \times \mathbb{R})\rightarrow C(\mathbb{R})$ is bounded. If $B < \infty$, then additionally $T: L^2(\mathcal{X} \times \mathbb{R})\rightarrow C(\mathbb{R})$ is bounded. Moreover, if $f_0 - f_1 \in L^2(\mathcal{X} \times \mathbb{R})$ and $B< \infty$ then $T$ is a bounded linear map from $L^2(\mathcal{X} \times \mathbb{R})$ to $L^2(\mathbb{R})$.

The main assumption that we make is that the joint distribution of $X,U$ is sufficiently rich, in that the linear operator $T$ has small enough kernel. It will turn out that the stipulation $U \protect\mathpalette{\protect\independenT}{\perp} W$ implies that $\ker T$ is always nontrivial. Now, to proceed, we define the subspace of $x$-invariant functions $\mathcal{V}$ on $\mathcal{X} \times \mathbb{R}$ by

align*[align* omitted — 263 chars of source]

Define also the set of functions $\mathcal{W}$ by

align*[align* omitted — 216 chars of source]

where we have used $C(\mathcal{X})_+$ to denote the set of positive continuous real-valued functions over $\mathcal{X}$. Note that $\mathcal{V}$ forms a closed linear subspace of $L^2(\mathcal{X} \times \mathbb{R})$, and the restriction of its elements to $\mathcal{X} \times [0,2B]$ is a closed linear subspace of $L^2(\mathcal{X} \times [0,2B])$. Therefore, projection onto $\mathcal{V}$ is a well-defined and bounded operator.

Our key assumption, which is given in two (nonequivalent) forms, is as follows.

asm\normalfont When $T$ is viewed as an operator mapping $L^2(\mathcal{X} \times [0,2B]) \rightarrow C(\mathbb{R})$, either: \begin{enumerate} • $\ker{T} \cap \mathcal{W} = \emptyset$, • or more specifically $\ker{T}\subset \mathcal{V}$. \end{enumerate}

Note that we impose the constraint that $\delta(x)$ is nonconstant in the definition of $\mathcal{W}$. Hence, it may be readily be seen that $\mathcal{V} \cap \mathcal{W} = \emptyset$ so that Assumption (ref)(ii) is stronger than Assumption (ref)(i). It will turn out that Assumption (ref)(i) is necessary and sufficient for our purposes of identification, but Assumption (ref)(ii) has a more meaningful interpretation that we now turn to.

It is a fact that our standing assumption that $U \protect\mathpalette{\protect\independenT}{\perp} W$ implies that $\mathcal{V} \subset \ker{T}$; indeed, for any $h \in \mathcal{V}$ there exists by definition some function $H$ such that:

align*[align* omitted — 178 chars of source]

So under independence Assumption (ref)(ii) amounts to the condition that $\ker{T}$ is precisely $\mathcal{V}$. Indeed, have the following result clarifying the relationship between Assumption (ref)(ii) and independence $U \protect\mathpalette{\protect\independenT}{\perp} W$.

lemmaUnder Assumptions (ref) and (ref), independence $U \protect\mathpalette{\protect\independenT}{\perp} W$ is equivalent to the inclusion $\mathcal{V} \subset \ker{T}$.

Now, we are able to clarify the conditions necessary and sufficient to obtain point identification of $g$ under our stated regularity assumptions. The following theorem is an application of Green's theorem whose proof makes clear why we have introduced the operator $T$:

theoremSuppose Assumptions (ref), (ref), and the restrictions $U \protect\mathpalette{\protect\independenT}{\perp} W$ and $\mathrm{E}\left[ U \right] = 0$ hold. Then Assumption (ref)(i) holds if and only if $g$ is point identified in the set $\mathcal{G} \equiv \left\{ h \in C(\mathcal{X}): \left\lVert h\right\rVert_\infty \le B \right\}$. In particular, Assumption (ref)(ii) implies point identification of $g$.

Assumption (ref) is clearly central and merits further investigation. To further understand it we can rephrase our requirement in terms more closely resembling typical completeness assumptions, which typically appear as:

align[align omitted — 133 chars of source]

where $W$ is some instrument for an endogenous regressor $X$. This particular assumption has been addressed in varying forms; see A2017 for a recent treatment.

To place our Assumption (ref)(ii) in terms of the more familiar condition (ref), let $V$ be a random variable distributed as uniform $\mathcal{U}[0,2B]$, independently of $(U,X)$, which we assume has a distribution with density $f(x,u)$ on $\mathcal{X} \times \mathbb{R}$ in accordance with Assumption (ref).

lemmaLet $\widetilde{U} \equiv U + V$ where $V\sim \mathcal{U}[0,2B]$ is independent of $(U,X,W)$. Then under Assumptions (ref) and (ref), Assumption (ref)(ii) is equivalent to the following assertion: \[ \mathrm{E}\left[ h(X,V) | \widetilde{U}, W = 0 \right] \overset{\mathrm{a.s.}}{ = } \mathrm{E}\left[ h(X,V )| \widetilde{U}, W = 1 \right] \implies h \in \mathcal{V} \] whenever $h \in L^2(\mathcal{X}\times [0,2B])$.

More on Assumption (ref)

Given its utility, we wish to explore conditions under which the stronger Assumption (ref) holds, in particular with respect to the conditional density functions $f_0(x,u)$ and $f_1(x,u)$. To this end, define $\Gamma$ as the set of functions $\gamma(x,u)$ over $\mathcal{X} \times \mathbb{R}$ which are of the form $f_0(x,u)-f_1(x,u)$, where $f_0$ and $f_1$ are proper density functions, i.e.\ \[ \Gamma \equiv \left\{ \gamma : \, \gamma = f_0 - f_1, \text{ } f_0, f_1 \text{ are density functions over }\mathcal{X} \times \mathbb{R} \text{ satisfying } U \protect\mathpalette{\protect\independenT}{\perp} W \right\}. \] We use $U \protect\mathpalette{\protect\independenT}{\perp} W$ to indicate that one should have the relation $\int_\mathcal{X} f_0(x,u) \, \mathrm{d}x = \int_\mathcal{X} f_1(x,u) \, \mathrm{d} x$, for a.e.\ $u$, i.e.\ equality of the conditional distributions of $U$ given $W$, almost everywhere. It can be seen that

align[align omitted — 216 chars of source]

An alternate formulation is to impose the moment condition that $f_0$ satisfies $\int_\mathbb{R} \int_\mathcal{X} u f_0(x,u) \, \mathrm{d}x \, \mathrm{d} u = 0$ which has been employed throughout the paper. It may readily be seen that:

align*[align* omitted — 616 chars of source]

where inclusion in one direction is clear and inclusion in the other direction follows from the fact that, given $\gamma$ satisfying $\left\lVert \gamma\right\rVert_{L^1(\mathcal{X} \times \mathbb{R})} \le 2$ and $\int_\mathcal{X} \gamma(x,u) \, \mathrm{d} x \, \mathrm{d} u = 0$ for a.e. $x$, one may set

align[align omitted — 292 chars of source]

where $\gamma = \gamma^+- \gamma^-$, $\gamma^+, \gamma^- \ge 0$, and $\rho$ is an arbitrary probability density function defined on $\mathcal{X} \times \mathbb{R}$ chosen to satisfy $\int_\mathbb{R} \int_\mathcal{X} uf_0(x,u)\, \mathrm{d} x \, \mathrm{d} u = 0$ (and if necessary to ensure integrability of the functions). Then $\gamma = f_0 - f_1$. Note that (ref) only slightly differs from (ref) and only in those elements $\gamma$ for which $\left\lVert \gamma\right\rVert_{L^1(\mathcal{X} \times \mathbb{R})} = 2$. As the results established in this section concern density and genericity in the $L^1(\mathcal{X} \times \mathbb{R})$ norm and are not affected by the immaterial change from (ref) to (ref), we work with the more convenient definition (ref). Note that (ref) implies that $\Gamma$ is a closed set in $L^1(\mathcal{X}\times \mathbb{R})$; if for instance $\gamma_n$ is a sequence occurring in $\Gamma$ such that $\gamma_n \rightarrow \gamma$ in $L^1(\mathcal{X} \times \mathbb{R})$ then the Lebesgue differentiation theorem implies

align*[align* omitted — 447 chars of source]

For any $\gamma \in \Gamma$, define the linear operator $T_\gamma: L^2(\mathcal{X} \times \mathbb{R} ) \rightarrow C(\mathbb{R})$ (see Lemma (ref)) by

align*[align* omitted — 119 chars of source]

Then let $\Gamma_0 = \left\{ \gamma \in \Gamma: \, \gamma \text{ is continuous and } \ker{T_\gamma} = \mathcal{V} \right\}$. Then we have the following density result:

propositionIn the preceding notation, $\Gamma_0$ is dense in $\Gamma$ in the $L^1(\mathcal{X} \times \mathbb{R} )$-norm.

As an immediate corollary, we obtain:

corollaryFor any continuous probability distributions $f_0$, $f_1$ on $\mathcal{X} \times \mathbb{R}$ satisfying $U \protect\mathpalette{\protect\independenT}{\perp} W$, and $\varepsilon > 0$, there exist probability densities $f_0^\varepsilon$ and $f_1^\varepsilon$ such that $\left\lVert f_\ell^\varepsilon - f_\ell\right\rVert_{L^1(\mathcal{X} \times \mathbb{R} )} < \varepsilon$ for $\ell =0,1$, $f_0^\varepsilon - f_1^\varepsilon \in \Gamma_0$, and $\int_{\mathbb{R}} \int_\mathcal{X} u f_0^\varepsilon(x,u) \, \mathrm{d}x \, \mathrm{d} u = 0$.

Topological genericity of Assumption (ref)(i) under a Lipschitz restriction

By making an additional assumption on the smoothness of the function $g$, one can comment on the topologically genericity of Assumption (ref)(i). Recall that in a given topological space $(\mathfrak{X}, \mathcal{T})$ a set is called residual or comeagre if it contains a countable intersection of open dense sets (a dense $G_\delta$ set). When $(\mathfrak{X}, \mathcal{T})$ is a complete metric space, the Baire Category theorem implies that any residual set contains a dense set (additionally, any residual set is not countable), and residuality is used to define a generic property on a given topological space. For instance, the irrational numbers comprise a residual set in $\mathbb{R}$ whereas the rationals do not. Recall that $\Gamma$ is a closed set in $L^1(\mathcal{X} \times \mathbb{R} )$, and thus it is a complete metric space with the topology induced by $L^1(\mathcal{X} \times \mathbb{R} )$.

In this section, we will confine $g$ to belong to the class of Lipschitz-continuous functions on $\mathcal{X}$. Then if $h$ is also in the identified set and Lipschitz continuous, the difference $\delta = g - h$ is Lipschitz continuous. By contrapositive, if there are no Lipschitz continuous functions $\delta$ such satisfy $\mathbf{1}_{u \in [0, \delta(x)]} \in \ker{T}$, then it follows straightforwardly by the method of Theorem (ref) that $g$ is point identified under the Lipschitz restriction. Thus, analogously to $\mathcal{W}$, define \[ \mathcal{W}_\mathrm{Lip} \equiv \left\{ \mathbf{1}_{u \in [0, \delta(x)]} \text{ for some Lipschitz-continuous, nonconstant } \delta \in C(\mathcal{X})_+\right\} \] $\mathcal{W}_\mathrm{Lip}$ is the restriction of $\mathcal{W}$ to the Lipschitz case. Equipped with these definitions, we have the following result.

propositionLet $\Gamma_1 \equiv \left\{ \gamma \in \Gamma: \, \ker{T_\gamma } \cap \mathcal{W}_{\mathrm{Lip}} = \emptyset \right\}$; then $\Gamma_1$ is a residual set in $\Gamma$ in the topology induced from $L^1(\mathcal{X} \times \mathbb{R} )$.

The genericity result is also relevant when we consider densities, and not functions which are the difference of densities. For let $\mathfrak{X}$ denote the set of probability density functions over $\mathcal{X} \times \mathbb{R}$ equipped with the $L^1(\mathcal{X} \times \mathbb{R} )$ norm, and $\mathfrak{F} \subset \mathfrak{X} \times \mathfrak{X}$ the set of pairs of densities $(f_0, f_1)$ such that $f_0 - f_1 \in \Gamma$, with the induced product topology. Note that $\mathfrak{F}$ is manifestly closed therein. Let $\mathfrak{F}_1$ denote the set of pairs $(f_0, f_1)$ such that $f_0 - f_1 \in \Gamma_1$. Then:

corollaryWhen $\mathfrak{F}$ is equipped with its induced product topology, $\mathfrak{F}_1$ is a residual set in $\mathfrak{F}$.

In Corollary (ref) and the definition of $\mathfrak{F}$ we do not impose the additional moment requirement that $\int_\mathbb{R} \int_\mathcal{X} uf_0(x,u) \, \mathrm{d} x \, \mathrm{d} u = 0$ considered in (ref) because the set of densities which satisfy this condition, and indeed the weaker condition of mere integrability of $uf_0(x,u)$, is not a closed subset of $L^1(\mathcal{X}\times \mathbb{R})$. Imposing this requirement would require us to consider a stronger topology on $\mathfrak{F}$; results in this direction could certainly be made along the lines of Proposition (ref), but the $L^1$ topology is arguably the most natural when discussing the $L^1$-closed set of probability density functions.

Identification when $U$ has Compact Support

Given that the examples produced in Proposition (ref) and Corollary (ref) of conditional density functions which point identified $g$ had unbounded support, it may come as a surprise to the reader that there exist examples of density functions which point identify $g$ and have bounded support. For suppose that $\text{supp}\left( U \right) \subset [-C_1, C_2]$ for fixed constants $C_1, C_2 > 0$. We show the following density result which is a corollary of Proposition (ref) and Corollary (ref):

corollaryIf $C_1 + C_2 > 2B$ then for any continuous probability densities $f_0, f_1$ on $\mathcal{X} \times [-C_1 , C_2]$ satisfying $U \protect\mathpalette{\protect\independenT}{\perp} W$ and $\mathrm{E}\left[ U \right] = 0$ and any $\varepsilon > 0$ there exist probability densities $f_0^\varepsilon, f_1^\varepsilon \in L^1(\mathcal{X} \times [-C_1, C_2])$ such that $\left\lVert f_\ell^\varepsilon - f_\ell\right\rVert_{L^1(\mathcal{X} \times \mathbb{R})} < \varepsilon$ for $\ell = 0 , 1$, $f_0^\varepsilon - f_1^\varepsilon \in \Gamma_0$, and $\int_\mathbb{R} \int_\mathcal{X} uf_0^\varepsilon (x,u) \, \mathrm{d} x \, \mathrm{d} u = 0$.

Retaining the notation of \S (ref), let $\mathfrak{F}^{C_1,C_2}$ denote the set of elements in $\mathfrak{F}$ with support in $\mathcal{X} \times [-C_1, C_2]$. Similarly, let $\mathfrak{F}_1^{C_1, C_2}$ denote the elements $(f_0,f_1) \in \mathfrak{F}^{C_1,C_2}$ such that $\ker{T_{f_0 - f_1}} \cap \mathcal{W}_{\mathrm{Lip}} \cap \mathcal{W} = \emptyset$ (note now the dependence on the bound $B$ via $\mathcal{W}$). Then exactly the same arguments which led to Proposition (ref) and Corollary (ref) imply in light of Corollary (ref) that

corollaryIf $C_1 + C_2 > 2B$ then $\mathfrak{F}_1^{C_1, C_2}$ is a residual set in $\mathfrak{F}^{C_1, C_2}$.

Connection with Identification in Nonparametric Instrumental Variables Quantile Regression

Interestingly, it is possible to extend the methods used in this section to identification in a quantile regression model as considered by CH2005 and later by HL2007. Consider the framework

align[align omitted — 139 chars of source]

adopted in HL2007, where $\gamma$ is some known function mapping $\text{supp}\left( W \right)$ into $[0,1]$ and variables retain their interpretation from our standing model (ref). The model displayed in (ref) nests the model considered in HL2007 (consider the constant function $\gamma(w) = q$, $q$ fixed), who show that their model subsumes the setup considered by CH2005. For a random variable $Z$, say that $W$ is boundedly complete for $(X,U)$ if for all bounded functions $h: \text{supp}\left( (X,U) \right) \rightarrow \mathbb{R}$ one has $\mathrm{E}\left[ h(X,U)|W \right] \overset{\mathrm{a.s.}}{=} 0$ if and only if $h \overset{\mathrm{a.s.}}{=} 0$. In the spirit of Assumption (ref), say that $W$ is boundedly $X$-complete for $(X,U)$ if for all bounded functions $h: \text{supp}\left( (X,U) \right) \rightarrow \mathbb{R}$ one has $\mathrm{E}\left[ h(X,U)|W \right] \overset{\mathrm{a.s.}}{=} 0$ only if $h(X,U) \overset{\mathrm{a.s.}}{=} H(U)$ for some function $H$, i.e.\ $h$ does not depend on $X$. It should be clear that if $Z$ is boundedly complete for $(X,U)$, then it is boundedly $X$-complete for $(X,U)$.

Equipped with these definitions, we derive the following identification result for the structural function $g$ completely along the lines of Theorem (ref):

propositionSuppose model (ref) holds. If $W$ is boundedly complete for $(X,U)$, then $g$ is point identified. If $0$ is in the interior of $\text{supp}\left( U \right)$ and $W$ is boundedly $X$-complete for $(X,U)$, then $g$ is also point identified.

One helpful aspect of Proposition (ref) is that it sidesteps some issues faced when considering identification of nonlinear operators, which is faced by CH2005. A number of sufficient conditions for bounded completeness have been developed by e.g.\ D2011, to which we refer the interested reader. Roughly speaking, if

align*[align* omitted — 47 chars of source]

for some random disturbance $\varepsilon$ which is independent of $W$, where $\mu$ and $\nu$ are possibly vector-valued functions, then there are light conditions which can be made (see Assumptions 1-3 and 4 of D2011) to ensure that $W$ is boundedly complete for $(X,U)$.

Discussion, Comparison with Local Identification

The condition obtained by D2014 (see their equation (16)) and more recently considered by C2019 (see their Assumption 2.1) for local identification in our model is the relation: for $h$ satisfying $\mathrm{E}\left[ h(X) \right] = 0$ and $\mathrm{E}\left[ |h(X)|^2 \right] < \infty$,

align[align omitted — 180 chars of source]

Comparison with Assumption (ref)(ii) shows that the stronger assumption we make in order to obtain point identification of $g$ is stronger than (ref), as should be expected. For suppose that (ref) does not hold for some square integrable $h$: then for all $t \in \mathbb{R}$

align*[align* omitted — 630 chars of source]

Hence, $h \in \ker{T} \setminus \mathcal{V}$ and Assumption (ref)(ii) is violated. The relation of (ref) with the necessary and sufficient condition Assumption (ref)(i) is more difficult to ascertain, which may suggest that (ref) is not a necessary condition for local identification.

A typical completeness condition puts $\dim X = \dim W$ and asks that, conditional on some restrictions on the function $h$, $\mathrm{E}\left[ h(X)|W \right] \overset{\mathrm{a.s.}}{=} 0$ if and only if $h(X) \overset{\mathrm{a.s.}}{=} 0$. One of the most studied examples where the completeness condition is fulfilled puts $X = \mu(\nu(W) + \varepsilon)$, as in D2011; in this case, it is somewhat essential that $\dim W \ge \dim X$ and that $W$ satisfies a large support condition. Interestingly, in both of the settings we have discussed, identification has been shown to arise when a form of completeness condition holds for an instrument which has possibly lower dimension than its regressor. For instance, Lemma (ref) shows that our Assumption (ref) is tantamount to the assertion that a random variable $\widetilde{U} = U + V$ is complete for the vector $(X,V)$ in a sense defined there, and within the class of functions $\mathcal{V}$. Of course, $\dim(X,V) > \dim \widetilde{U}$, which imposes some difficulties when attempting to view our identification assumptions through the typical lens of instrument completeness. Moreover, in Proposition (ref) we require $W$ to act as a complete instrument for $(X,U)$, so that in order to apply conventional examples of completeness one would have to have $\dim W \ge \dim X + \dim U$. Hence, while the conditions enumerated in Lemma (ref) and Proposition (ref) are not necessary for identification, they suggest that to state examples of identified models in our framework is also to make progress on finding sufficient conditions for the completeness condition when the instrument has strictly lower degree than the regressor (and vice-versa).