EconBase
← Back to paper

Matching Points: Supplementing Instruments with Covariates in Triangular Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,786 characters · 23 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Matching Points: Supplementing Instruments with Covariates in Triangular Models

\nolinenumbers

abstractModels with a discrete endogenous variable are typically underidentified when the instrument takes on too few values. This paper presents a new method that matches pairs of covariates and instruments to restore point identification in this scenario in a triangular model. The model consists of a structural function for a continuous outcome and a selection model for the discrete endogenous variable. The structural outcome function must be continuous and monotonic in a scalar disturbance, but it can be nonseparable. The selection model allows for unrestricted heterogeneity. Global identification is obtained under weak conditions. The paper also provides estimators of the structural outcome function. Two empirical examples of the return to education and selection into Head Start illustrate the value and limitations of the method. \\ \\ \noindentKeywords: Nonparametric identification, triangular model, instrumental variable, endogeneity, generalized propensity score. \\ \\

\setcounter{page}{0} \thispagestyle{empty}

Introduction

This paper considers identification and estimation of the structural outcome function $\boldsymbol{g}^{*}\equiv (g^{*}_{d})_{d}$ in a triangular model:

linenomath*\begin{align} Y&=\sum_{d\in S(D)}\mathbbm{1}(D=d)\cdot g^{*}_{d}(\boldsymbol{X},U_{d})\\ D&=h(\boldsymbol{X},Z,\boldsymbol{V}) \end{align}

where both the endogenous variable $D\in S(D)$ and the instrumental variable $Z\in S(Z)$ are discrete, $\boldsymbol{X}$ is a vector of covariates, and the disturbances $(U_{d})_{d}$ and $\boldsymbol{V}$ are correlated (see different versions of the model in newey1999nonparametric, chesher2003identification, imbens2009identification, etc.).

In many applications, the instrument takes on fewer values than the endogenous variable, i.e., the cardinality of their support sets satisfy $|S(Z)|<|S(D)|$. The outcome function $\boldsymbol{g}^{*}$ is then in general underidentified. For instance, under exogeneity of $Z$ and certain shape restrictions or separability of $\bm{g}^{*}(\boldsymbol{X},\cdot)$, the $|S(D)|$-vector of the unknown parameters $\bm{g}^{*}(\boldsymbol{x}_{0},u)$ for some $\bm{X}=\bm{x}_{0}$ and $u$ may satisfy moment equations conditional on $Z$ and $\bm{X}=\bm{x}_{0}$ (e.g. newey2003instrumental and chernozhukov2005iv). Conditioning on each value of $Z$ generates one moment equation. The number of the equations is thus $|S(Z)|$, smaller than the number of the unknowns ($|S(D)|$). The classical order condition fails and so does identification of $\bm{g}^{*}(\boldsymbol{x}_{0},u)$.

This paper develops a novel approach to obtain point identification of $\boldsymbol{g}^{*}$ in this scenario. For a value of interest $\bm{X}=\bm{x}_{0}$, identification is achieved by finding special values $\bm{X}=\bm{x}_{m}$, the matching points, such that the function $\boldsymbol{g}^{*}(\boldsymbol{x}_{m},\cdot)$ can be expressed as a known mapping of $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},\cdot)$. As a consequence, both variation in $Z$ and local variation in $\bm{X}$ across $\bm{x}_{0}$ and $\bm{x}_{m}$ have identification power for $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},\cdot)$. Specifically, by substituting the mapping into the moment equations conditional on $Z$ and $\bm{X}=\bm{x}_{m}$ for a chosen $u$, the unknowns in these equations become $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u)$. Together with the moment equations conditional on $Z$ and $\bm{X}=\bm{x}_{0}$, the total number of the equations for $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u)$ is increased, while the number of the unknowns is unchanged. Even though the covariates are not excluded from the structural function, the special value $\bm{x}_{m}$ facilitates identification in a way an additional instrument value does. The effective support set of the instrument is thus enlarged, making identification possible.

Two key restrictions are needed in addition to the exogeneity of the instrument for the mapping of $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},\cdot)$ to $\boldsymbol{g}^{*}(\boldsymbol{x}_{m},\cdot)$ to be traced out before themselves are. First, the matching point $\bm{x}_{m}$ and the value of interest $\bm{x}_{0}$ need to generate the same selection patterns when paired with appropriate instrument values. That is, for some $z,z'\in S(Z)$, $\bm{x}_{m}$ needs to satisfy $h(\bm{x}_{m},z',\cdot)=h(\bm{x}_{0},z,\cdot)$. Second, each outcome disturbance $U_{d}$ is a scalar and $g^{*}_{d}(\bm{X},\cdot)$ is continuous and strictly increasing for all $d\in S(D)$ almost surely.

A consequence of the first restriction is that the distributions of the outcome disturbance $U_{D}\equiv\sum_{d\in S(D)}U_{d}$ conditional on $D$ given $(\bm{X},Z)=(\bm{x}_{0},z)$ and $(\bm{x}_{m},z')$ are equal, under some other assumptions. Intuitively, $D$ has the same degree of endogeneity under $(\bm{X},Z)=(\bm{x}_{m},z')$ and $(\bm{x}_{0},z)$. Continuity and monotonicity of $g^{*}_{d}(\bm{X},\cdot)$ imposed by the second key restriction then transform equality between the conditional distributions of $U_{D}$ into equality between the conditional distributions of $Y$ evaluated at $g_{d}^{*}(\bm{x}_{0},\cdot)$ and $g^{*}_{d}(\bm{x}_{m},\cdot)$ for each $d$. The mapping from $\bm{g}^{*}(\bm{x}_{0},\cdot)$ to $\bm{g}^{*}(\bm{x}_{m},\cdot)$ can then be traced out by inverting these observable distributions.

Two approaches are available to find the matching points $\bm{x}_{m}$ satisfying the first restriction. If $h$ is known and identified, one can obtain the matching points by searching for an $\bm{x}_{m}$ that matches $h$ with $\bm{x}_{0},z,z'$ fixed. If $h$ is unknown, I show that a statistical implication of the first key restriction is that the generalized propensity scores conditional on $(\bm{X},Z)=(\bm{x}_{0},z)$ and $(\bm{x}_{m},z')$ are equal. Therefore, one can recover the matching points robustly by searching for an $\bm{x}_{m}$ to match the generalized propensity scores without knowing $h$. A sufficient condition for all such $\bm{x}_{m}$ to be matching points is that $(\bm{X},Z)$ enters $h$ only via the generalized propensity scores. Usually, this requires the selection model to have an index structure such as discrete choice models (see heckman2005structural, for example). For selection models not satisfying these conditions, one can still use the solutions to the generalized propensity score matching as candidates of matching points, and their validity is testable. Therefore, the underlying selection model can be very general with no restrictions on the dimensionality or separability in the selection heterogeneity $\bm{V}$.

After the order condition is fulfilled using the instrument and the matching points, continuity and monotonicity of $\bm{g}^{*}(\bm{X},\cdot)$ also simplify the sufficient conditions for global identification. I show that the outcome function $\boldsymbol{g}^{*}(\bm{x}_{0},\cdot)$ is globally identified among monotonic functions if $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u)$ is only locally identified for all $u$. Hence, this new result only relies on local invertibility conditions for nonlinear functions, which are much weaker than the conditions in global inverse theorems widely adopted for global identification. This result also applies to the standard nonparametric quantile IV approach when the instrument has large support, and may be of independent interest. Based on the identification strategy, I construct a sieve estimator and derive its asymptotic properties under simple low level conditions.

An important special case of a continuous and strictly increasing structural outcome function is that it is additively separable in the disturbance. I show that under separability, the outcome function at a given $\bm{X}=\bm{x}_{0}$ solves a system of linear equations, preserving a similar structure as in the nonparametric IV approach with rich instruments (e.g. newey2003instrumental and das2005instrumental). I construct a closed form estimator that is easy to implement in practice. I apply it to a return to education application using the same extract from the 1979 National Longitudinal Surveys (NLS) as in card1995using. Adopting the binary proximity-to-college instrument, the returns of three levels of education, high school, some college, and college and above, are nonparametrically underidentified by the existing approaches. To apply my approach, I use the average of parents' years of schooling as a covariate to generate matching points. Identification is restored using the matching points. I then estimate the returns and find that they are increasing in the level of education and heterogeneous in parents' years of schooling.

It is worth noting that my approach hinges on the covariates' ability to offset the effect of the instrument on selection. For applications where the instrument has a dominant effect, matching points may not exist. As an illustration, I consider another empirical example of the choice of preschool programs. I use the Head Start Impact Study dataset following kline2016evaluating. A randomly assigned lottery granting access to Head Start serves as the instrument, while the endogenous variable is the multivalued preschool program choice. I find that the instrument has a much larger effect on selection than covariates such as the baseline test scores and family income. No matching points exist.

I defer a detailed discussion of the relation of my approach to the literature until Section (ref). Here let me only briefly highlight some major differences. Methods that circumvent the problem of a small-support instrument include imposing homogeneity between adjacent levels of $D$ when $D$ is ordered, or specifying a parametric form for $\boldsymbol{g}^{*}$ and using interactions between $Z$ and $\boldsymbol{X}$ as extra instruments by assuming $\boldsymbol{X}$ is exogenous. These methods would fail in a fully nonparametric model, as studied in this paper. torgovitsky2015identification,torgovitsky2017minimum and d2015identification show that a binary instrument can identify nonseparable models with a continuous $D$. Different from my approach, continuity in $D$ is indispensable, and they require $\bm{V}$ to be a scalar and the selection function $h$ to be strictly increasing in it. caetano2016identifying use covariates to identify models when the instruments do not have enough variation. Their approach does not rely on a selection model, but they need the covariates used for identification purposes separable from the model. In contrast, the covariates in my approach can enter the model in an arbitrary way. The idea in my approach of using shifts in some observables to compensate for a shift in a target variable to facilitate identification can also be seen in ichimura2000direct, vytlacil2007dummy and chen2020identification. They focus on different parameters than this paper, and the shifting variables and the target variables are also different. vuong2017counterfactual and feng2019estimation study the individual treatment effect of a binary $D$ and develop a concept called the counterfactual mapping. It is also an identifiable mapping linking two outcome functions, but at different values of $D$ and the same value of $\bm{X}$.

The rest of the paper is organized as follows. In Section (ref), I introduce the matching points and show how to use them to generate new moment equations and restore the order condition. How to find the matching points is also discussed. Given the fulfilled order condition, Section (ref) provides sufficient conditions for the global identification of the nonseparable model. Section (ref) discusses identification of a separable model as a special case under relaxed conditions. Section (ref) sketches estimation of the matching points and the separable model. Section (ref) shows Monte Carlo simulation results to illustrate the estimator's finite sample performance. Section (ref) presents two empirical applications. Section (ref) discusses the relation of my approach to the literature. Section (ref) concludes. Some additional results and the proofs of identification are in Appendix. In the Supplemental Material, I provide an estimator of the general nonseparable model, proofs of its asymptotic properties, and additional simulation results.

Notation

A vector valued function is said to be (strictly) increasing or monotonic if every component in it is (strictly) increasing or monotonic. Random variables are denoted by upper-case Latin letters and their realizations by the corresponding lower-cases. Bold Latin letters denote vectors or matrices. For two generic random variables $A$ and $B$, denote the conditional expectation of $A$ given $B=b$ by $\mathbb{E}_{A|B}(b)$, with similar notation for conditional cumulative distribution functions $F_{A|B}(a|b)$, densities $f_{A|B}(a|b)$ and variances $\mathbb{V}_{A|B}(b)$. Denote the support set of $A$ by $S(A)$, and the support of $A$ given $B=b$ by $S(A|B=b)$, or simply $S(A|b)$ when it does not cause confusion. For a finite set $G$, $|G|$ denotes its cardinality. Throughout, I assume all the random variables are in a common probability space with the measure function $\mathbb{P}$. Almost surely (a.s.) and measurable are always with respect to $\mathbb{P}$.

Matching Points and the Order Condition

To highlight the key features of my approach, I focus on a simple case where a single endogenous $D$ takes on three values ($|S(D)|=3$) and the instrument $Z$ is binary $(|S(Z)|=2)$. The discrete $D$ can be either ordered or unordered. Without loss of generality, let $S(D)=\{1,2,3\}$ and $S(Z)=\{0,1\}$. The general cases of arbitrary $|S(D)|>|S(Z)|$ and of multiple endogenous variables will be discussed in Appendix (ref). Besides the endogenous variable and the instrument, a continuous outcome variable $Y$ and a vector of covariates $\bm{X}$ are observable. The outcome is determined by the structural equation (ref). The function $\bm{g}^{*}$ will be referred to as the outcome function subsequently.

For a fixed value $\bm{x}_{0}\in S(\bm{X})$, $\bm{g}^{*}(\bm{x}_{0},\cdot)$ is underidentified by the standard IV approach (for example chernozhukov2005iv). This is because for any $u$, there can be only two moment equations for $\bm{g}^{*}(\bm{x}_{0},u)$ by conditioning on $(\bm{X},Z)=(\bm{x}_{0},0)$ and $(\bm{x}_{0},1)$. But $\bm{g}^{*}(\bm{x}_{0},u)$ contains three elements. The classical order condition thus fails. The idea of this paper is to generate more moment equations by conditioning on some other special values of the covariates, the matching points, denoted by $\bm{x}_{m}$. Consider the moment equations conditional on $(\bm{X},Z)=(\bm{x}_{m},0)$ and $(\bm{x}_{m},1)$. Although the unknowns in these equations are $\bm{g}^{*}(\bm{x}_{m},u)$, if $\bm{g}^{*}(\bm{x}_{m},u)$ can be rewritten as an known function of $\bm{g}^{*}(\bm{x}_{0},u)$, then substituting the function into these moment equations provide new equations for $\bm{g}^{*}(\bm{x}_{0},u)$.

To achieve this goal, a selection model for $D$ is needed. By the discreteness of $D$, rewrite the selection model (ref) as follows:

linenomath*\begin{equation} D=d if and only if h_{d}(\boldsymbol{X},Z,\boldsymbol{V})=1 \end{equation}

where for all $d\in S(D)$, the function $h_{d}(\boldsymbol{X},Z,\boldsymbol{V})\in\{0,1\}$ and $\sum_{d\in S(D)} h_{d}(\boldsymbol{X},Z,\boldsymbol{V})=1$ a.s. The random vector $\boldsymbol{V}$ contains unobservables that are correlated with the outcome disturbances $\{U_{d}\}$. I assume that for every $(\boldsymbol{x},z)\in S(\boldsymbol{X},Z)$, the indicator function $h_{d}(\boldsymbol{x},z,\cdot)$ is measurable on $S(\boldsymbol{V})$. Note that the covariates in the selection model are the same as those in the outcome function for simplicity. Additional covariates in the outcome function are indeed allowed. In that case, all the analysis in this paper can be viewed as conditional on them. In the rest of this section, I introduce the matching points of an arbitrary $\bm{x}_{0}\in S(\bm{X})$ and related concepts, and show how the mappings between the outcome functions at these points are identified, and how to use them to generate moment equations.

Matching Points and the M-Connected Set

A matching point $\bm{x}_{m}$ of a given point $\bm{x}_{0}\in S(\bm{X})$ is defined as follows:

definition[Matching Point and Matching Pair] A point $\boldsymbol{x}_{m}\in S(\boldsymbol{X})$ is a matching point of $\boldsymbol{x}_{0}\in S(\boldsymbol{X})$ if there exist $z, z'\in S(Z)$ such that the following equations hold for all $d\in S(D)$, $u\in S(U)$ and $\bm{v}\in S(\bm{V})$: \begin{linenomath*}\begin{align} h_{d}(\boldsymbol{x}_{m},z',\bm{v})&=h_{d}(\boldsymbol{x}_{0},z,\bm{v}),\\ F_{U_{d}\boldsymbol{V}|\bm{X}Z}(u,\bm{v}|\boldsymbol{x}_{m},z')&=F_{U_{d}\boldsymbol{V}|\bm{X}Z}(u,\bm{v}|\boldsymbol{x}_{0},z) \end{align}\end{linenomath*} The covariates-instrument combinations $(\boldsymbol{x}_{0},z)$ and $(\boldsymbol{x}_{m},z')$ form a matching pair.

Note that equation (ref) is satisfied for any $(\bm{x}_{0},z)$ and $(\bm{x}_{m},z')$ if $(\bm{X},Z)$ are jointly independent of $(U_{d},\bm{V})$ for all $d\in S(D)$. The independence of $Z$ is needed in this paper and will be introduced in the following subsections. The independence of the covariates is unnecessary, though commonly assumed in practice. It is worth noting that the conditions only need to be satisfied by the covariates used to generate the matching points. Finally, all these restrictions are indirectly testable under overidentification of $\bm{g}^{*}(\bm{x}_{0},\cdot)$.

Denote the generalized propensity score $\mathbb{P}(D=d|\boldsymbol{X}=\boldsymbol{x},Z=z)$ at an arbitrary point $(\bm{x},z)$ by $p_{d}(\boldsymbol{x},z)$. The following lemma is a direct consequence of equations (ref) and (ref).

lemSuppose $(\bm{x}_{0},z)$ and $(\bm{x}_{m},z')$ are a matching pair. For all $d\in S(D)$ and $u\in S(U)$, \begin{linenomath*}\begin{align} F_{U_{d}|D\boldsymbol{X}Z}(u|d,\boldsymbol{x}_{m},z')&=F_{U_{d}|D\boldsymbol{X}Z}(u|d,\boldsymbol{x}_{0},z)\\ p_{d}(\boldsymbol{x}_{m},z')&=p_{d}(\boldsymbol{x}_{0},z) \end{align}\end{linenomath*}
proofSee Appendix (ref).

Lemma (ref) does not require the instrument or the covariates to be exogenous. Also, no structures for the outcome functions are imposed so far. So the lemma may have a broader application than studied in this paper. In Section (ref), I will show under what conditions the mapping from $\bm{g}^{*}(\bm{x}_{0},\cdot)$ to $\bm{g}^{*}(\bm{x}_{m},\cdot)$ can be traced out from equation (ref), while equation (ref) provides a model-free way to find the matching points, discussed in Section (ref).

For each $\bm{x}_{0}$, one may find different matching points by pairing $\bm{x}_{0}$ with different values of $Z$; for some $\bm{x}_{m1}\neq \bm{x}_{m2}$, points $(\bm{x}_{0},0)$ and $(\bm{x}_{m1},1)$ and $(\bm{x}_{0},1)$ and $(\bm{x}_{m2},0)$ may form two different matching pairs. Furthermore, a matching point of $\bm{x}_{0}$ may have its own matching points besides $\bm{x}_{0}$. See the following example for an illustration.

example[Ordered Choice] Suppose $D$ is ordered and there is only one covariate $X$. Let $h_{1}(X,Z,V)=\mathbbm{1}(V<\kappa_{1}+\beta X+\alpha Z)$, $h_{3}(X,Z,V)=\mathbbm{1}(V\geq \kappa_{2}+\beta X+\alpha Z)$, and $h_{2}=1-h_{1}-h_{3}$. Assume $\alpha\cdot\beta \neq 0$, $\kappa_{1}<\kappa_{2}$, and $(X,Z)\protect\mathpalette{\protect\independenT}{\perp} (U_{d},V)$ for all $d$ where $V$ is continuously distributed on $\mathbb{R}$. For any fixed $x_{0}\in S(X)$, it has the following two matching points by equation (ref) if both are in $S(X)$: \begin{linenomath*}\begin{align} (z=0,z'=1):\beta x_{m1}+\alpha\cdot 1=\beta x_{0}+\alpha\cdot 0\implies x_{m1}=x_{0}-\frac{\alpha}{\beta}\\ (z=1,z'=0):\beta x_{m2}+\alpha\cdot 0=\beta x_{0}+\alpha\cdot 1\implies x_{m2}=x_{0}+\frac{\alpha}{\beta} \end{align}\end{linenomath*} Similarly, each of $x_{m1}$ and $x_{m2}$ also has two matching points: One is $x_{0}$, and the other is $x_{0}-2\frac{\alpha}{\beta}$ and $x_{0}+2\frac{\alpha}{\beta}$ respectively. This process can be continued until the boundaries of $S(X)$ are reached, illustrated in the following figure: \begin{figure}[H] \caption{The Pyramid of Matching Points} \end{figure} The horizontal axis in Figure (ref) is the value of the single index $x\beta+z\alpha$. Starting from $(x_{0},0)$ and $(x_{0},1)$, $x_{m1}$ and $x_{m2}$ are obtained by solving the equations below the horizontal axis. Repeat this procedure to match $(x_{m1},0)$ with $(x_{m3},1)$ and $(x_{m2},1)$ with $(x_{m4},0)$. Continuing the process, one can expect to see that the dotted points on this axis extend to both directions, until they reach the boundaries of $S(X)$.

In Example (ref), $x_{0}$ has two matching points: $x_{m1}$ and $x_{m2}$. Each of them has one extra descendent matching point $x_{m3}$ and $x_{m4}$. These points, in turn, have their matching points. Although they are no longer matching points of $\boldsymbol{x}_{0}$, the conditional distributions of $U_{d}$ at these descendants and at $\boldsymbol{x}_{0}$ paired with appropriate values of the instrument are still linked by repeatedly applying Lemma (ref). Motivated by this observation, let me introduce a concept that is more general than the matching point.

definition[M-Connected Set] A set $\mathcal{X}_{MC}(\boldsymbol{x}_{0})\subseteq S(\bm{X})$ is called the m-connected set of $\boldsymbol{x}_{0}$ if $\boldsymbol{x}_{0}\in \mathcal{X}_{MC}(\boldsymbol{x}_{0})$ and for any $\boldsymbol{x}\in \mathcal{X}_{MC}(\boldsymbol{x}_{0})$, there exists $\boldsymbol{x}_{1},\boldsymbol{x}_{2},...,\boldsymbol{x}_{k(\boldsymbol{x})}\in \mathcal{X}_{MC}(\boldsymbol{x}_{0})$ such that $\boldsymbol{x}_{j}$ is a matching point of $\boldsymbol{x}_{j-1}$, $j=1,...,k(\boldsymbol{x})$, and $\boldsymbol{x}$ is a matching point of $\boldsymbol{x}_{k(\boldsymbol{x})}$. Any two points in the m-connected set are said to be m-connected.

In Example (ref), any two points in the form of $x_{0}+c\frac{\alpha}{\beta},c\in\mathbb{Z}$, are m-connected provided that both are in $S(X)$. The m-connected set of $x_{0}$ is $\mathcal{X}_{MC}(x_{0})=\{x_{0}+c\frac{\alpha}{\beta},c\in\mathbb{Z}\}\cap S(X)$.

By construction, the m-connected set is the largest subset of $S(\bm{X})$ such that the relationship between the conditional distributions $U_{d}$ at any two elements in it can be established by recursively applying Lemma (ref).

The Fulfillment of the Order Condition

Now it is ready to present how to use Lemma (ref) to recover the mapping of $\bm{g}^{*}(\bm{x}_{0},\cdot)$ to $\bm{g}^{*}(\bm{x},\cdot)$ for any $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$, and how in turn to use the m-connected points to supplement the instrument.

The Mapping of $\bm{g}^{*}(\bm{x}_{0},\cdot)$ to $\bm{g}^{*}(\bm{x},\cdot)$

Matching $\bm{g}^{*}(\bm{x},\cdot)$ with $\bm{g}^{*}(\bm{x}_{0},\cdot)$ by Lemma (ref) needs the following assumptions.

assumption[Continuity and Monotonicity] For all $\boldsymbol{x}\in S(\boldsymbol{X})$, $\bm{g}^{*}(\boldsymbol{x},\cdot)$ is continuous and strictly increasing.
assumption[Normalization, Exogeneity, and Full Support] For all $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$ and $d\in S(D)$, conditional on $\bm{X}=\bm{x}$ , i) $U_{d}\sim\textrm{Unif}[0,1]$, ii) $(U_{d},\boldsymbol{V})$ is continuously distributed with $(U_{d},\boldsymbol{V})\protect\mathpalette{\protect\independenT}{\perp} Z$, and iii) $S(U_{d}|\boldsymbol{V})=S(U_{d})$ .

Assumption (ref) regulates the behavior of $\boldsymbol{g}^{*}(\boldsymbol{x},\cdot)$. Continuity and strict monotonicity are two standard requirements in the literature of nonseparable models with a scalar unobservable (e.g. matzkin2003nonparametric,matzkin2007nonparametric, chernozhukov2005iv, etc.). In addition to constructing moment conditions, I will show that they also deliver nice results for the global uniqueness of the solution to systems of nonlinear equations in Section (ref).

Assumption (ref) i) normalizes the distribution of $U_{d}$ for each $d\in S(D)$ conditional on an arbitrary point $\bm{x}$ in $\mathcal{X}_{MC}(\bm{x}_{0})$. Together with Assumption (ref), the normalization gives $g^{*}_{d}(\bm{x},\cdot)$ an interpretation of the counterfactual quantile function.

In Assumption (ref) ii), the requirement of joint independence between $Z$ and $(U_{d},\bm{V})$ is common for triangular models (e.g. imbens2009identification), yet it is usually made only conditional on the value of interest $\bm{X}=\bm{x}_{0}$. The assumption here is stronger since joint independence needs to hold conditional on every point in $\mathcal{X}_{MC}(\bm{x}_{0})$. The need arises from using the moment conditions conditional on the m-connected points instead of only on $\bm{X}=\bm{x}_{0}$, as will be seen later in this section. It is noteworthy that when the moment conditions obtained from the m-connected points are more than needed, joint independence can be relaxed to hold only conditional on points in a subset of $\mathcal{X}_{MC}(\bm{x}_{0})$.

In Assumption (ref) iii), the full support requirement is useful to identify $\bm{g}^{*}(\bm{x}_{0},\cdot)$ on the entire domain $[0,1]$. It can also be found in related work with a similar focus, for instance d2015identification, torgovitsky2015identification and vuong2017counterfactual.

Assumptions (ref) and (ref) guarantee that $S(Y|d,\boldsymbol{x},z)=S(Y|d,\bm{x})$, that $S(Y|d,\bm{x})$ is compact, and that the range of $g^{*}_{d}(\boldsymbol{x},\cdot)$ on $[0,1]$ is equal to $S(Y|d,\boldsymbol{x})$ for all $d\in S(D)$, $z\in S(Z)$ and $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$.

Under Assumption (ref), $F_{Y|D\bm{X}Z}\left(g^{*}_{d}(\bm{x},u)|d,\bm{x},z\right)=F_{U_{d}|D\bm{X}Z}\left(u|d,\bm{x},z\right)$ for any $(u,d,\bm{x},z)\in S(U_{D},D,\bm{X},Z)$. Equation (ref) in Lemma (ref) thus implies the following equation for a matching pair $(\bm{x}_{0},z)$ and $(\bm{x}_{m},z')$ for all $d\in S(D)$ and $u\in S(U_{d}|d,\bm{x}_{0},z)$:

linenomath*\begin{equation} F_{Y|D\bm{X}Z}\left(g^{*}_{d}(\bm{x}_{m},u)|d,\bm{x}_{m},z'\right)=F_{Y|D\bm{X}Z}\left(g^{*}_{d}(\bm{x}_{0},u)|d,\bm{x}_{0},z\right), \end{equation}

Let $F^{-1}_{Y|DXZ}:[0,1]\mapsto S(Y|D,X,Z)$ be the inverse of the cumulative conditional distribution function of $Y$ whenever it is strictly increasing on the support. Let $q_{d,(\bm{x},z),(\bm{x'},z')}(y)=F^{-1}_{Y|D\boldsymbol{X}Z}\big(F_{Y|D\boldsymbol{X}Z}(y|d,\boldsymbol{x},z)\big|d,\boldsymbol{x}',z'\big)$ for $y\in \mathbb{R}$. Under exogeneity of $Z$ and the full support condition in Assumption (ref), $F_{Y|D\bm{X}Z}\left(g^{*}_{d}(\bm{x}_{m},u)|d,\bm{x}_{m},z'\right)$ is strictly increasing in $g^{*}_{d}(\bm{x}_{m},u)$ for all $u\in [0,1]$. Therefore, equation (ref) implies

linenomath*\begin{equation} g^{*}_{d}(\boldsymbol{x}_{m},u)=q_{d,(\bm{x}_{0},z),(\bm{x}_{m},z')}\left(g^{*}_{d}(\boldsymbol{x}_{0},u)\right) \end{equation}

for all $d\in S(D)$ and $u\in [0,1]$. The mapping $q_{d,(\bm{x}_{0},z),(\bm{x}_{m},z')}$ is directly identifiable from the population, continuous and increasing on $\mathbb{R}$ and strictly increasing on $S(Y|d,\bm{x}_{0})$.

More generally, let $\varphi_{d}(g^{*}_{d}(\boldsymbol{x}_{0},u);\bm{x})$ denote the mapping from $g^{*}_{d}(\boldsymbol{x}_{0},u)$ to $g^{*}_{d}(\boldsymbol{x},u)$ for an arbitrary $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$. This mapping is also identified. Specifically, suppose $\bm{x}$ is m-connected with $\bm{x}_{0}$ such that $(\bm{x}_{0},z_{0})$ and $(\bm{x}_{1},z_{1})$, $(\bm{x}_{1},z_{1}')$ and $(\bm{x}_{2},z_{2})$, ..., $(\bm{x}_{k},z_{k}')$ and $(\bm{x},z)$ are matching pairs, where $z_{0},z_{1},z_{1}',z_{2},...,z_{k}',z\in S(Z)$, then for all $d\in S(D)$ and $u\in [0,1]$,

linenomath*\begin{equation} \varphi_{d}(g^{*}_{d}(\boldsymbol{x}_{0},u);\bm{x})=q_{d,(\bm{x}_{k},z_{k}'),(\bm{x},z)}\circ\cdots\circ q_{d,(\bm{x}_{1},z_{1}'),(\bm{x}_{2},z_{2})}\circ q_{d,(\bm{x}_{0},z_{0}),(\bm{x}_{1},z_{1})}(g_{d}^{*}(\bm{x}_{0},u)) \end{equation}

where $\circ$ denotes function composition.

The mapping $\varphi_{d}(\cdot;\bm{x})$ has some useful properties. First, it is continuous and increasing on $\mathbb{R}$, and strictly increasing on $S(Y|d,\bm{x}_{0})$. Second, for a matching pair $(\bm{x}_{m},z')$ and $(\bm{x}_{0},z)$, $\varphi_{d}(g^{*}_{d}(\boldsymbol{x}_{0},u);\bm{x}_{m})=q_{d,(\bm{x}_{0},z),(\bm{x}_{m},z')}(g^{*}_{d}(\boldsymbol{x}_{0},u))$. Third, since $(\bm{x}_{0},z)$ and itself form a matching pair by definition, $\varphi_{d}(g^{*}_{d}(\bm{x}_{0},u);\bm{x}_{0})=q_{d,(\bm{x}_{0},z),(\bm{x}_{0},z)}(g^{*}_{d}(\boldsymbol{x}_{0},u))=g^{*}_{d}(\boldsymbol{x}_{0},u)$.

The Moment Conditions

With the traced out mapping from $g^{*}_{d}(\bm{x}_{0},\cdot)$ to $g^{*}_{d}(\bm{x},\cdot)$ for each $d\in S(D)$ and $\bm{x}\in \mathcal{X}_{MC}(\bm{x}_{0})$, moment conditions can be constructed.

assumption[Rank Similarity] Conditional on any $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$, $\{U_{d}\}$ are identically distributed conditional on $\bm{V}$.

Assumption (ref) is adopted from chernozhukov2005iv to handle $d$-dependent outcome disturbances. It is worth emphasizing that Assumption (ref) is not needed to identify the mapping $\varphi_{d}$.

prop[Moment Condition] Under Assumptions (ref) to (ref), the following equation holds for all $z\in\{0,1\}$, $\bm{x}\in \mathcal{X}_{MC}(\bm{x}_{0})$ and $u\in[0,1]$, \begin{linenomath*}\begin{equation} \sum_{d\in S(D)}p_{d}(\boldsymbol{x},z)\cdot F_{Y|D\boldsymbol{X}Z}(\varphi_{d}(g^{*}_{d}(\boldsymbol{x}_{0},u);\bm{x})|d,\boldsymbol{x},z)=u \end{equation}\end{linenomath*}
proofSee Appendix (ref).

Proposition (ref) generalizes Theorem 1 in chernozhukov2005iv. To identify $\bm{g}^{*}(\bm{x}_{0},u)$, they only condition on $\bm{X}=\bm{x}_{0}$. In equation (ref), this is the case when $\bm{x}=\bm{x}_{0}$ because $\varphi_{d}(g^{*}_{d}(\boldsymbol{x}_{0},u);\bm{x}_{0})=g^{*}_{d}(\boldsymbol{x}_{0},u)$, and the equation becomes

linenomath*\begin{equation*} \sum_{d\in S(D)}p_{d}(\boldsymbol{x}_{0},z)\cdot F_{Y|D\boldsymbol{X}Z}(g^{*}_{d}(\boldsymbol{x}_{0},u)|d,\boldsymbol{x}_{0},z)=u \end{equation*}

As the generalized propensity scores and the conditional cumulative distribution functions are directly identified, two moment conditions are available for each $u\in [0,1]$ by setting $z=0$ and $1$, but there are three unknowns.

Now with the matching points and more generally, the m-connected points, equation (ref) induces a larger system of equations for $\bm{g}^{*}(\bm{x}_{0},u)$. The number of the equations is determined by the sizes of $S(Z)$ and $\mathcal{X}_{MC}(\bm{x}_{0})$. As $(\bm{X},Z)=(\bm{x}_{0},0)$ and $(\bm{x}_{0},1)$ already provide two moment equations, the order condition is fulfilled as long as one nontrivial matching point of $\bm{x}_{0}$ exists. More discussion on the number of the matching points and the size difference of $S(D)$ and $S(Z)$ can be found in Appendix (ref). With more matching points and the m-connected points, we may have overidentification, and the rank condition introduced in Section (ref) for global identification is more likely to be satisfied.

It is worth noting that although each m-connected point (including $\bm{x}_{0}$ by definition) can induce two moment equations with $z=0,1$, in the four equations generated by two adjacently m-connected points (i.e., one is a matching point of the other), one equation is redundant. For example, suppose $(\bm{x}_{0},0)$ and $(\bm{x}_{m},1)$ are a matching pair. Equation (ref) holds for $(\bm{x},z)=(\bm{x}_{0},0),(\bm{x}_{0},1),(\bm{x}_{m},0),(\bm{x}_{m},1)$, but it can be checked that the equations at $(\bm{x}_{0},0)$ and $(\bm{x}_{m},1)$ are identical. So in Figure (ref) (Example (ref)), ten covariate-instrument combinations are available to build moment equations for $\bm{g}^{*}(\boldsymbol{x}_{0},u)$, but four of them are redundant (one in each pair of the arrow-connected points).

Finally, utilizing the matching points is useful even when $S(D)=S(Z)$; the outcome function is then overidentified so that the validity of the instrument is testable together with the validity of the matching points.

Finding the Matching Points

So far, I assumed that the matching points existed and were known. In this subsection, I discuss the existence of the matching points and how to find them.

First, equation (ref) in Lemma (ref) provides a statistical implication for a matching point. For $\bm{x}_{m}\in S(\bm{X})$ to be a matching point of $\bm{x}_{0}$, there must exist $z,z'\in S(Z)$ such that

linenomath*\begin{equation} \big(p_{1}(\boldsymbol{x}_{m},z')-p_{1}(\boldsymbol{x}_{0},z)\big)^{2}+\big(p_{2}(\boldsymbol{x}_{m},z')-p_{2}(\boldsymbol{x}_{0},z)\big)^{2}=0 \end{equation}

The generalized propensity scores at $d=3$ are not included because they are matched automatically under (ref). Hence, for a given $\bm{x}_{0}$, if there does not exist $\bm{x}_{m},z$ and $z'$ that satisfy equation (ref), no matching point exists.

The existence of a solution to equation (ref) depends on how much $\boldsymbol{X}$ can affect the generalized propensity scores at $z'$ and how large $S(\bm{X})$ is. For instance, if the generalized propensity scores $(p_{1}(\bm{X},z'),p_{2}(\bm{X},z'))$ have full support, i.e., $(p_{1}(\cdot,z'),p_{2}(\cdot,z')):S(\boldsymbol{X})\mapsto [0,1]\times [0,1]$ is surjective, then a solution always exists. As another example, in Example (ref), $x_{0}\pm \alpha/\beta$ are two matching points if they are in $S(X)$. So $S(X)$ needs to be sufficiently large and/or $\alpha/\beta$ is small. The latter implies that the effects of $X$ on the generalized propensity scores are large compared to $Z$.

Conversely, not all the solutions to equation (ref) are necessarily to be matching points. Sufficient conditions for both equations (ref) and (ref) in Definition (ref) to hold for all the solutions to (ref) are that i) $(\bm{X},Z)$ enters the selection model only via the generalized propensity scores, and ii) $(\bm{X},Z)\protect\mathpalette{\protect\independenT}{\perp} (U_{d},\bm{V})$ for all $d\in S(D)$.

Condition i) is often imposed in the literature on local instrumental variable (LIV) and marginal treatment effect (MTE) heckman1999local, heckman2001policy,heckman2005structural,heckman2006understanding. Typically only selection models that are separable in $\bm{V}$ satisfy the condition. Appendices (ref) and (ref) provide some examples.

Note that the purpose of condition i) in this paper is different from that in the MTE and LIV literature. Here it guarantees that the selection model can be matched via the generalized propensity scores. Yet in the LIV and MTE literature, the instruments are continuous and the condition is imposed to obtain index sufficiency, that is, for any $d\in S(D)$ and $y\in S(Y|d,\bm{X})$, $F_{Y|D\bm{X}Z}(y|d,\bm{X},Z)=F_{Y|D\bm{Xp}}\left(y|d,\bm{X},\bm{p}(\bm{X},Z)\right)$ a.s. (heckman2005structural, p.678). Index sufficiency holds trivially when $Z$ is binary as in this paper's setup, provided that $\bm{p}(\bm{X},z)\neq \bm{p}(\bm{X},z')$ a.s. This is because in this case, $\bm{p}(\bm{X},Z)$ is a one-to-one function of $Z$ given $\bm{X}$. Hence, index sufficiency itself does not have identification power here.

When condition i) or ii) does not hold, for instance $\bm{X}$ is correlated with $(U_{d},\bm{V})$ for some $d\in S(D)$ or the selection model is nonseparable in $\bm{V}$, matching points may still exist because both requirements in Definition (ref) are local, yet conditions i) and ii) impose global restrictions on the selection model and the dependence of $(U_{d},\bm{V})$ on $(\bm{X},Z)$. Since the matching points form a subset of the solutions to equation (ref), one can treat the solutions as candidates for the matching points and obtain candidates for the m-connected points similarly. As long as the number of the moment equations from the instrument and such candidates is greater than $|S(D)|$, whether the candidates are truly matching points or m-connected points is testable.

Identification

The fulfilled order condition makes identification of the outcome function $\bm{g}^{*}(\bm{x}_{0},\cdot)$ possible. This section provides a new result on the uniqueness of the solution to the system of nonlinear equations characterized by Propositions (ref). The result relies on weaker conditions than those commonly used in the Hadamard-type global inverse function theorems. Since the moment conditions obtained by the nonparametric quantile IV approach with rich instruments are special cases of the ones obtained in this paper, this new result also applies there.

Proposition (ref) shows that for each $u\in [0,1]$, $\bm{g}^{*}(\bm{x}_{0},u)$ solves a system of nonlinear equations. Unlike identification of nonseparable models with a continuous $D$ chernozhukov2007instrumental,chen2014local, here we do not face the ill-posed problem due to the discreteness of $D$. Nonetheless, establishing global identification of $\bm{g}^{*}(\bm{x}_{0},u)$ is still demanding. The Jacobian matrix of the nonlinear equation system being full rank at $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u)$ only implies local identification of $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u)$, and stronger high level conditions are required for global identification chernozhukov2005iv.

However, the above approach with $u$ fixed does not exploit the structures of $\bm{g}^{*}(\bm{x}_{0},\cdot)$ as a function as well as the structures of the moment equations. By construction, $\bm{g}^{*}(\bm{x}_{0},\cdot)$ and the function $F_{Y|D\bm{X}Z}(\varphi_{d}(\cdot;\bm{x})|d,\bm{x},z)$ are continuous and strictly increasing on $[0,1]$ and on $S(Y|d,\bm{x}_{0})$ respectively. In this section, I show that with these properties, local identification of $\bm{g}^{*}(\bm{x}_{0},u)$ at every $u\in [0,1]$ guarantees global identification of $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},\cdot)$ in the class of monotonic functions. This notion of identification is from a solution path perspective, common in differential equations and defined as follows.

definition[Solution Path] For an interval $\mathcal{U}\subseteq\mathbb{R}$ and a system of equations $\boldsymbol{M}(\boldsymbol{y},u)=\boldsymbol{0}$ where $\boldsymbol{y}$ is a real vector and $u\in\mathcal{U}$, a solution path $\boldsymbol{y}^{*}(\cdot)$ is a function on $\mathcal{U}$ such that $\boldsymbol{M}(\boldsymbol{y}^{*}(u),u)=\boldsymbol{0}$ for all $u\in\mathcal{U}$.
lemLet $\mathcal{Y}\subseteq \mathbb{R}^{K}$, $\mathcal{U}\subseteq\mathbb{R}$ be a compact interval, and $\bm{M}(\cdot,\cdot):\mathcal{Y}\times \mathcal{U}\mapsto \mathbb{R}^{L}$ be continuously differentiable. Suppose there exists a continuous and weakly increasing function $\bm{y}^{*}:\mathcal{U}\mapsto\mathcal{Y}$ such that $\bm{M}(\bm{y}^{*}(u),u)=\bm{0}$ for all $u\in\mathcal{U}$ and $\bm{y}^{*}(u^{*})=\bm{c}$ for some $u^{*}\in \mathcal{U}$ and $\bm{c}\in\mathcal{Y}$. If $\bm{M}(\cdot,u)$ is strictly increasing in each argument and its Jacobian matrix at $\bm{y}^{*}(u)$, $\nabla\bm{M}(\bm{y}^{*}(u),u)$, is full rank for all $u\in\mathcal{U}$, then $\bm{y}^{*}$ is the unique weakly increasing solution path that passes through $(u^{*},\bm{c})$.
proofSee Appendix (ref).
remThe lemma also holds for a system of equations with a decreasing solution path: If $\bm{y}$ is decreasing, let $\tilde{\bm{M}}((-\bm{y}),u)\equiv-\bm{M}(-(-\bm{y}),u)$, and thus $-\bm{y}$ is increasing and $\tilde{\bm{M}}(\cdot,u)$ as a function of $-\bm{y}$ is strictly increasing in every argument. Lemma (ref) thus applies. Similarly, $\bm{M}(\cdot,u)$ can be strictly decreasing as well.

The proof of Lemma (ref) is in Appendix (ref). Here let me provide some heuristics to highlight the key roles played by monotonicity and continuity. Suppose $u^{*}$ is in the interior of $\mathcal{U}$ and there exists another increasing solution path $\tilde{\bm{y}}$ with $\tilde{\bm{y}}(u^{*})=\bm{c}$ but $\tilde{\bm{y}}(u)\neq \bm{y}^{*}(u)$ for all $u$ in some interval right to $u^{*}$. Then $\tilde{\bm{y}}(\cdot)$ cannot be continuous at $u^{*}$, otherwise there must exist some $u>u^{*}$ such that $\tilde{\bm{y}}(u)$ is close enough to $\bm{y}^{*}(u)$ that violates the local uniqueness of the solution implied by the full rank and continuous Jacobian. Consequently, $\tilde{\bm{y}}$ must jump up at $u^{*}$ since it is increasing. However, this is again not possible because otherwise, by continuity and monotonicity of $\bm{M}(\cdot,u^{*})$, $\bm{M}$ would jump up as well and thus the equation cannot hold at $\tilde{\bm{y}}(u^{*})$.

Although the uniqueness only holds among functions passing through the same point, this condition can be trivially satisfied in some special cases. For instance, let $\underline{u}$ be the lower boundary of $\mathcal{U}$. Suppose $\mathcal{Y}$ is the product of compact intervals and each component in $\bm{y}^{*}(\underline{u})$ equals the lower boundary of the corresponding interval, then all possible increasing solution paths $\tilde{\bm{y}}$ must satisfy $\tilde{\bm{y}}(\underline{u})=\bm{y}^{*}(\underline{u})$ because $\tilde{\bm{y}}(\underline{u})$ must be no smaller than $\bm{y}^{*}(\underline{u})$ to be in $\mathcal{Y}$, while if its greater than $\bm{y}^{*}(\underline{u})$, the system of equations cannot hold at $(\tilde{\bm{y}}(\underline{u}),\underline{u})$ because $\bm{M}(\cdot,\underline{u})$ is strictly increasing. This is indeed the case in this paper, as will be seen later in this section.

Lemma (ref) shows that monotonicity and continuity simplify the sufficient conditions usually required for the global uniqueness of a solution at a fixed $u$ (see ambrosetti1995primer for variants of Hadamard's theorem). Here, the Jacobian matrix is just required to be full rank along the unique solution path, which only guarantees the local uniqueness of the solution at each fixed $u$. The lemma thus says that the local uniqueness of the solution pointwise in $u$ implies the global uniqueness of a monotonic solution path. Finally, the result holds among a class of functions where discontinuous functions are allowed. This is crucial to obtain the other results in this section.

Now let us turn to global identification of $\bm{g}^{*}(\bm{x}_{0},\cdot)$. Let $\mathcal{Z}(\boldsymbol{x}_{0})\equiv \mathcal{X}_{MC}(\boldsymbol{x}_{0})\times S(Z)$. For any three points $\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3}\in\mathcal{Z}(\bm{x}_{0})$, let $\Psi(\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u);\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3}))$ denote the $3\times 1$ vector by stacking the left hand side of equation (ref) evaluated at theses points respectively. For example, the $k$-th component in $\Psi$ is $\sum_{d=1}^{3}p_{d}(\tilde{\boldsymbol{z}}_{k})\cdot F_{Y|D\boldsymbol{X}Z}(\varphi_{d}(\cdot)|d,\tilde{\boldsymbol{z}}_{k})$. Denote the vector $(u,u,u)'$ by $\bm{u}$. Then $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},\cdot)$ is one solution path to $\boldsymbol{M}(\boldsymbol{y},u)\equiv \Psi(\boldsymbol{y};\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3})-\boldsymbol{u}=\bm{0}$ on $[0,1]$. Let $\mathcal{G}$ be the set of all increasing functions defined on $[0,1]$:

linenomath*\begin{equation} \mathcal{G}\equiv\{\boldsymbol{g}:[0,1]\mapsto \mathbb{R}^{3} and is weakly increasing\} \end{equation}

The following theorem provides sufficient conditions that guarantee global identification of $\boldsymbol{g}^{*}(\bm{x}_{0},\cdot)$ in $\mathcal{G}$.

thm[Global Identification of $\bm{g}^{*}(\bm{x}_{0},\cdot)$] Under Assumptions (ref) to (ref), if there exist $\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3}\in \mathcal{Z}(\boldsymbol{x}_{0})$ such that $\Psi(\cdot;\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3})$ is continuously differentiable on $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$, and that its Jacobian matrix at $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},u)$ is full rank for all $u\in[0,1]$, then $\boldsymbol{g}^{*}(\boldsymbol{x}_{0},\cdot)$ is the unique solution path (up to $u=0,1$) to $\Psi(\cdot;\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3})-\boldsymbol{u}=0$ in $\mathcal{G}$.
proofSee Appendix (ref).
remThe conditioning points $\tilde{\bm{z}}_{1},\tilde{\bm{z}}_{2}$ and $\tilde{\bm{z}}_{3}$ do not necessarily include $(\boldsymbol{x}_{0},z)$ and $(\boldsymbol{x}_{0},z')$. It is possible to use any points in $\mathcal{Z}(\bm{x}_{0})$ to achieve full rankness. Meanwhile, once $\bm{g}^{*}(\bm{x}_{0},\cdot)$ is identified, $\bm{g}^{*}(\bm{x},\cdot)$ is identified for all $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$ by $\varphi_{d}(g^{*}_{d}(\bm{x}_{0},u);\bm{x})$.

The proof of Theorem (ref) consists of two steps. In the first step, I invoke Lemma (ref) to show that the uniqueness holds in a smaller space $\mathcal{G}^{*}\equiv \{\boldsymbol{g}:[0,1]\mapsto \prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})\text{ and is weakly increasing}\}$, a subset of $\mathcal{G}$ defined in equation (ref). This step is straightforward by treating $[0,1]$ as $\mathcal{U}$ and $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$ as $\mathcal{Y}$ in Lemma (ref). Since $\bm{g}^{*}(\bm{x}_{0},\cdot)$ is continuous on the closed interval $[0,1]$, $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$ is compact by Assumption (ref) and $g^{*}_{d}(\bm{x}_{0},0)$ and $g_{d}^{*}(\bm{x}_{0},1)$ are equal to the lower and the upper boundaries of $S(Y|d,\bm{x}_{0})$ for all $d\in S(D)$. Since $F_{Y|DXZ}\circ\varphi_{d}$ in $\Psi$ is strictly increasing, all possible solutions in $\mathcal{G}^{*}$ at $u=0$ and $1$ must also equal these boundaries to make the system of equations hold at these $u$. All the conditions in Lemma (ref) are then satisfied.

In the second step, I show that the uniqueness indeed holds in the larger space $\mathcal{G}$ by exploiting the properties of the cumulative distribution functions in $\Psi$. Outside $S(Y|d,\boldsymbol{x}_{0})$, $F_{Y|DXZ}(\cdot|d,\bm{x}_{0},z)$ is either $0$ or $1$ and equal to the value at the corresponding boundary of $S(Y|d,\boldsymbol{x}_{0})$. Therefore, if there exists a second solution path $\bm{g}(u)$ which may take on values outside $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$, there must also exist a function only taking on values within it (including the boundaries) which yields the same $\Psi$. That function is then in $\mathcal{G}^{*}$, so it is necessarily equal to $\bm{g}^{*}(\bm{x}_{0},\cdot)$. Therefore, $\bm{g}(u)$ has to be equal to $\bm{g}^{*}(\bm{x}_{0},u)$ for all $u\in (0,1)$ and can only take on values outside $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$ at $u=0$ or $1$. See the proof in Appendix (ref) for more details.

Allowing the parameter space to contain functions taking on values outside the conditional support $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$ is useful in estimation because then one does not need to accurately estimate $\prod_{d=1}^{3}S(Y|d,\boldsymbol{x}_{0})$ to obtain a consistent estimator of $\bm{g}^{*}(\bm{x}_{0},\cdot)$. For instance, one can focus on the following parameter space

linenomath*\begin{equation} \mathcal{G}_{0}\equiv \{\boldsymbol{g}:[0,1]\mapsto \prod_{d=1}^{3}S(Y|d) and is weakly increasing\} \end{equation}

where $S(Y|d)$ is easier to estimate than $S(Y|d,\bm{x}_{0})$ with an estimator converging much faster due to the discreteness of $D$.

Define

linenomath*\begin{equation} Q_{NSP}(\boldsymbol{g},u)\equiv \left(\Psi\left(\boldsymbol{g}(u);\tilde{\bm{z}}_{1},\tilde{\bm{z}}_{2},\tilde{\bm{z}}_{3}\right)-\boldsymbol{u}\big)'\boldsymbol{W}_{NSP}(u)\big(\Psi\left(\boldsymbol{g}(u);\tilde{\bm{z}}_{1},\tilde{\bm{z}}_{2},\tilde{\bm{z}}_{3}\right)-\boldsymbol{u}\right) \end{equation}

where $\boldsymbol{W}_{NSP}(u)$ is positive definite uniformly in $u\in [0,1]$. Theorem (ref) implies that $\bm{g}^{*}(\bm{x}_{0},\cdot)$ is the unique increasing function such that $\int_{0}^{1}Q_{NSP}(\cdot,u)du=0$. Beyond that, it is important to know if $\int_{0}^{1}Q_{NSP}(\boldsymbol{g}(u),u)du$ is well separated from $0$ when $\boldsymbol{g}(\cdot)$ is well separated from $\boldsymbol{g}^{*}(\bm{x}_{0},\cdot)$, for instance, whether the following inequality holds for any $\delta>0$ and any closed interval $\mathcal{U}_{0}$ in the interior of $[0,1]$:

linenomath*\begin{equation} \inf_{\substack{\boldsymbol{g}\in\mathcal{G}_{0}\\\sup_{u\in \mathcal{U}_{0}}|\boldsymbol{g}(u)-\boldsymbol{g}^{*}(x_{0},u)|\geq \delta}}\int_{0}^{1}Q_{NSP}(\boldsymbol{g}(u),u)du>0 \end{equation}

It can be verified that the infinite dimensional space $\mathcal{G}_{0}$ is not compact under the sup-metric. In general, inequality (ref) does not necessarily hold when the parameter space is noncompact even under the global uniqueness of $\bm{g}^{*}(\bm{x}_{0},\cdot)$ chen2007large,chen2012estimation. However, the following corollary shows that this is not a concern here. Again, monotonicity and continuity of $\bm{g}^{*}(\bm{x}_{0},\cdot)$ play the central role in it.

corUnder the conditions in Theorem (ref), inequality (ref) is true.
proofSee Appendix (ref).

It is noteworthy that since $\varphi_{d}(g^{*}_{d}(\boldsymbol{x}_{0},\cdot);\boldsymbol{x}_{0})=g^{*}_{d}(\boldsymbol{x}_{0},\cdot)$, Theorem (ref) and Corollary (ref) also apply to the standard nonparametric quantile IV approach when $D$ is discrete with $|S(Z)|\geq |S(D)|$ (for example chernozhukov2005iv).

Before closing this section, let me emphasize that global identification in terms of the solution path does not rule out the possibility that at some $u$, the solution to $\Psi(\cdot)=\boldsymbol{u}$ is not unique. This is expected because the conditions required here are much weaker than the sufficient conditions for global invertibility of $\Psi(\cdot)$ on $\Pi_{d=1}^{3} S(Y|d,\boldsymbol{x}_{0})$. Under this weaker notion of identification, one cannot estimate $\bm{g}^{*}(\bm{x}_{0},u)$ for a fixed $u$. In Appendix (ref) in the Supplemental Material, I provide an estimator that minimizes the sample analogue of $Q_{NSP}(\cdot,u)$ jointly at multiple nodes of $u$ in $[0,1]$ under a monotonicity constraint. The number of the nodes needs to grow to infinity slowly with the sample size.

A Special Case: The Separable Model

In some applications, the outcome disturbance $U_{d}$ may be additively separable. In this special case, identification results can be obtained under weaker conditions. Formally, suppose for each $d\in S(D)$, $g^{*}_{d}(\bm{X},U_{d})=m^{*}_{d}(\bm{X})+U_{d}$. The outcome equation (ref) can then be rewritten as

linenomath*\begin{equation} Y=\sum_{d\in S(D)}\mathbbm{1}(D=d)\cdot\left(m_{d}^{*}(\boldsymbol{X})+U_{d}\right)\ \end{equation}

Under separability, some requirements for the matching points and exogeneity of $Z$ can be relaxed. First, for $\bm{x}_{m}$ to be a matching point of $\bm{x}_{0}$, equation (ref) in Definition (ref) can be weakened such that for all $d\in S(D)$,

linenomath*\begin{equation} \mathbb{E}_{U_{d}|\boldsymbol{V}\boldsymbol{X}Z}(\bm{v},\boldsymbol{x}_{m},z')=\mathbb{E}_{U_{d}|\boldsymbol{V}\boldsymbol{X}Z}(\bm{v},\boldsymbol{x}_{0},z) and F_{\boldsymbol{V}|\bm{X}Z}(\bm{v}|\boldsymbol{x}_{m},z')=F_{\boldsymbol{V}|\bm{X}Z}(\bm{v}|\boldsymbol{x}_{0},z) \end{equation}

Essentially, for $U_{d}$, only mean dependence of $U_{d}$ on $\bm{V}$ are required to be the same conditional on $(\bm{X},Z)=(\bm{x}_{m},z')$ and on $(\bm{x}_{0},z)$.

Then silmilar to Lemma (ref), the generalized propensity scores at $(\bm{x}_{m},z')$ and $(\bm{x}_{0},z)$ are equal, and the following equation holds for all $d\in S(D)$:

linenomath*\begin{equation} \mathbb{E}_{U_{d}|D\boldsymbol{X}Z}(d,\boldsymbol{x}_{m},z')=\mathbb{E}_{U_{d}|D\boldsymbol{X}Z}(d,\boldsymbol{x}_{0},z) \end{equation}

With equation (ref), one can trace out the mapping from $m_{d}^{*}(\bm{x}_{0})$ to $m_{d}^{*}(\bm{x}_{m})$ for all $d\in S(D)$: Take expectations on both sides of equation (ref) conditional on $(D,\bm{X},Z)=(d,\bm{x},z)$:

linenomath*\begin{equation*} m_{d}^{*}(\boldsymbol{x})+\mathbb{E}_{U_{d}|D\boldsymbol{X}Z}(d,\boldsymbol{x},z)=\mathbb{E}_{Y|D\boldsymbol{X}Z}(d,\boldsymbol{x},z) \end{equation*}

Evaluate the above equation at $(\bm{x}_{m},z')$ and $(\bm{x}_{0},z)$ and subtract one from the other. The conditional expectations of $U_{d}$ are canceled out by equation (ref). For any $\bm{x},\bm{x}'\in S(\bm{X})$, let $\delta_{d,(\bm{x},z),(\bm{x}',z')}=\mathbb{E}_{Y|D\boldsymbol{X}Z}(d,\boldsymbol{x}',z')-\mathbb{E}_{Y|D\boldsymbol{X}Z}(d,\boldsymbol{x},z)$. Then we have

linenomath*\begin{equation} m_{d}^{*}(\boldsymbol{x}_{m})=m_{d}^{*}(\boldsymbol{x}_{0})+\delta_{d,(\bm{x}_{0},z),(\bm{x}_{m},z')} \end{equation}

The term $\delta_{d,(\bm{x}_{0},z),(\bm{x}_{m},z')}$ is directly identified from the population.

More generally, suppose $\bm{x}$ is m-connected with $\bm{x}_{0}$ such that $(\bm{x}_{0},z_{0})$ and $(\bm{x}_{1},z_{1})$, $(\bm{x}_{1},z_{1}')$ and $(\bm{x}_{2},z_{2})$, ..., $(\bm{x}_{k},z_{k}')$ and $(\bm{x},z)$ are matching pairs, where $z_{0},z_{1},z_{1}',...,z_{k}',z\in S(Z)$. Let

linenomath*\begin{equation*} \Delta_{d}(\bm{x}_{0},\bm{x})=\delta_{d,(\bm{x}_{0},z_{0}),(\bm{x}_{1},z_{1})}+\delta_{d,(\bm{x}_{1},z_{1}'),(\bm{x}_{2},z_{2})}\cdots+\delta_{d,(\bm{x}_{k},z_{k}'),(\bm{x},z)}, \end{equation*}

then we have $m^{*}_{d}(\bm{x})=m^{*}_{d}(\bm{x}_{0})+\Delta_{d}(\bm{x}_{0},\bm{x})$ for all $d\in S(D)$ and $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$. By construction, $\Delta_{d}(\bm{x}_{0},\bm{x})$ is directly identified. Meanwhile, for a matching pair $(\bm{x}_{0},z)$ and $(\bm{x}_{m},z')$, $\Delta_{d}(\bm{x}_{0},\bm{x}_{m})=\delta_{d,(\bm{x}_{0},z),(\bm{x}_{m},z')}$. In particular, $\Delta_{d}(\bm{x}_{0},\bm{x}_{0})=0$.

Finally, the exogeneity assumption of $Z$ and rank similarity can be relaxed as follows due to separability: For all $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$ and all $d\in S(D)$,

assumption[Normalization and Exogeneity] $\mathbb{E}_{U_{d}|\boldsymbol{X}}(\boldsymbol{x})=0$, $\mathbb{E}_{U_{d}|\boldsymbol{V}\boldsymbol{X}Z}(\boldsymbol{V},\boldsymbol{x},Z)=\mathbb{E}_{U_{d}|\boldsymbol{V}\boldsymbol{X}}(\boldsymbol{V},\boldsymbol{x})$ a.s., and $Z\protect\mathpalette{\protect\independenT}{\perp}\boldsymbol{V}$ conditional on $\boldsymbol{X}=\boldsymbol{x}$.
assumption[Mean Similarity] Conditional on $\bm{X}=\bm{x}$, $\{U_{d}\}$ have the same expectation conditional on $\bm{V}$.

Assumption (ref) imposes a location normalization and joint (mean) independence of $Z$ and $(U_{d},\bm{V})$ for each $d\in S(D)$ conditional on any m-connected point. Joint (mean) independence is standard in the literature on triangular models with a separable outcome functions newey1999nonparametric.

Assumption (ref) relaxes rank similarity (Assumption (ref)) in the nonseparable case. Instead of identical conditional distributions of $\{U_{d}\}$, only the conditional expectations are required to be identical.

The following proposition characterizes the moment condition for $\bm{m}^{*}(\bm{x}_{0})\equiv (m^{*}_{d}(\bm{x}_{0}))_{d}$.

prop[Moment Condition] Under Assumptions (ref) and (ref), the following equation holds for all $z\in S(Z)$ and $\bm{x}\in\mathcal{X}_{MC}(\bm{x}_{0})$, \begin{linenomath*}\begin{equation} \sum_{d\in S(D)}p_{d}(\boldsymbol{x},z)\cdot m_{d}^{*}(\boldsymbol{x}_{0})=\sum_{d\in S(D)}p_{d}(\boldsymbol{x},z)\cdot \left(\mathbb{E}_{Y|D\boldsymbol{X}Z}(d,\boldsymbol{x},z)-\Delta_{d}(\bm{x}_{0},\bm{x})\right) \end{equation}\end{linenomath*}
proofSee Appendix (ref).

Similar to the nonseparable case, when $\bm{x}=\bm{x}_{0}$, $\Delta_{d}(\bm{x}_{0},\bm{x})=0$ for all $d\in S(D)$ and then equation (ref) is back to equation (2.2) in newey2003instrumental or equation (2.5) in das2005instrumental.

Global identification of $\bm{m}^{*}(\bm{x}_{0})$ is straightforward to establish due to linearity of equation (ref). Recall the augmented set of the conditioning points $\mathcal{Z}(\boldsymbol{x}_{0})\equiv \mathcal{X}_{MC}(\boldsymbol{x}_{0})\times S(Z)$. Evaluating equation (ref) at any three points $\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3}\in\mathcal{Z}(\boldsymbol{x}_{0})$ yields a system of linear equations of $\bm{m}^{*}(\bm{x}_{0})$ with the coefficient matrix equal to $\Pi_{SP}(\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3})\equiv (\bm{p}(\tilde{z}_{1}),\bm{p}(\tilde{z}_{2}),\bm{p}(\tilde{z}_{3}))'$, where $\bm{p}$ is the column vector of the three generalized propensity scores $(p_{1},p_{2},p_{3})'$. Then $\bm{m}^{*}(\bm{x}_{0})$ is globally identified on $\mathbb{R}^{3}$ if

linenomath*\begin{equation} \Pi_{SP}(\tilde{\boldsymbol{z}}_{1},\tilde{\boldsymbol{z}}_{2},\tilde{\boldsymbol{z}}_{3}) is full rank. \end{equation}

Finally, let us discuss the sufficient and necessary conditions for the full rankness of $\Pi_{SP}$. For concreteness, let $(\boldsymbol{x}_{m},z')$ and $(\boldsymbol{x}_{0},z)$ be a matching pair. Let $\tilde{\bm{z}}_{1}=(\bm{x}_{0},z),\tilde{\bm{z}}_{2}=(\bm{x}_{0},z')$ and $\tilde{\bm{z}}_{3}=(\bm{x}_{m},z)$. The moment equation at $(\bm{x}_{m},z')$ is not included as it is identical with the equation at $(\boldsymbol{x}_{0},z)$.

Since the sum of the three columns in $\Pi_{SP}$ is always equal to $\boldsymbol{1}_{3\times 1}$, it can be shown that $\Pi_{SP}$ is full rank if and only if

linenomath*\begin{align} &\left[p_{1}(\boldsymbol{x}_{m},z)-p_{1}(\boldsymbol{x}_{0},z)\right]\cdot\left[p_{3}(\boldsymbol{x}_{0},z)-p_{3}(\boldsymbol{x}_{0},z')\right]\notag\\ \neq& \left[p_{1}(\boldsymbol{x}_{0},z)-p_{1}(\boldsymbol{x}_{0},z')\right]\cdot \left[p_{3}(\boldsymbol{x}_{m},z)-p_{3}(\boldsymbol{x}_{0},z)\right] \end{align}

Inequality (ref) does not hold if both sides are simultaneously zero. This is the case when $Z$ has no effect on $\boldsymbol{p}$ at $\boldsymbol{X}=\boldsymbol{x}_{0}$ or $\boldsymbol{X}\in\{\bm{x}_{0},\bm{x}_{m}\}$ has no effect on $\boldsymbol{p}$ at $Z=z$. Both can be ruled out by a local relevance condition saying that $\boldsymbol{X}$ and $Z$ have nonzero effects on the propensity scores at $(\boldsymbol{x}_{0},z)$.

Now suppose neither side is $0$. By $\boldsymbol{p}(\boldsymbol{x}_{0},z)=\boldsymbol{p}(\boldsymbol{x}_{m},z')$, inequality (ref) can be rewritten as

linenomath*\begin{equation} \frac{p_{1}(\boldsymbol{x}_{m},z)-p_{1}(\boldsymbol{x}_{0},z)}{p_{3}(\boldsymbol{x}_{m},z)-p_{3}(\boldsymbol{x}_{0},z)}\neq \frac{p_{1}(\boldsymbol{x}_{m},z')-p_{1}(\boldsymbol{x}_{0},z')}{p_{3}(\boldsymbol{x}_{m},z')-p_{3}(\boldsymbol{x}_{0},z')} \end{equation}

Inequality (ref) generally holds unless the propensity score differences are locally uniform. For example, it can be verified that the inequality holds in the ordered choice model in Example (ref) for almost all $x_{0}\in S(X)$ and their matching points unless $V$ is (locally) uniformly distributed. In particular, it holds in widely used logit and probit models.

Estimation

In this section, I illustrate how to estimate the matching points and the separable model given an i.i.d. sample $(Y_{i},D_{i},\boldsymbol{X}_{i},Z_{i})_{i=1}^{n}$. The estimator for the nonseparable model and its properties are introduced and proved in Appendix (ref) in the Supplemental Material. For illustrative purposes, I focus on the following benchmark case to highlight the key features of the estimation procedure: i) $\boldsymbol{X}$ is one dimensional, denoted by $X$, and ii) all the solutions to the generalized propensity score matching equation (ref) are the matching points. Let $(x_{0},0)$ and $(x_{m1},1)$, and $(x_{0},1)$ and $(x_{m2},0)$ be two matching pairs. This benchmark case is the simplest scenario where both the matching points and the outcome functions are overidentified, allowing me to introduce the overidentification tests. Extending a scalar $X$ to the multivariate case is straightforward. At the same time, whether the generalized propensity score matching is successful and whether a solution to the matching is a matching point are testable by the overidentification tests provided in this section.

Estimating the Matching Points

Let $S_{0}(X)$ be a compact subset in the interior of $S(X)$. For $z\in S(Z)$, let $\hat{\boldsymbol{p}}(\cdot,z)$ be a consistent estimator of the $3\times 1$ vector of the generalized propensity scores $\boldsymbol{p}(\cdot,z)$ uniformly on $S_{0}(X)$. For concreteness, I consider the following Nadaraya-Watson estimator for each component in it. Other common nonparametric estimators of conditional probability would work as well.

linenomath*\begin{equation} \hat{p}_{d}(x,z)=\frac{\sum_{i=1}^{n}\mathbbm{1}(D_{i}=d)K(\frac{X_{i}-x}{h_{x}})\mathbbm{1}(Z_{i}=z)}{\sum_{i=1}^{n} K(\frac{X_{i}-x}{h_{x}})\mathbbm{1}(Z_{i}=z)} \end{equation}

where $K(\cdot)$ is a kernel function and $h_{x}$ is the bandwidth converging to $0$.

Assume both matching points are in $S_{0}(X)$. The matching points can be estimated by the sample analogue of equation (ref). Denote $\Delta \hat{\boldsymbol{p}}(x_{1},x_{2})\equiv \big(\hat{p}_{1}(x_{1},1)-\hat{p}_{1}(x_{0},0),\hat{p}_{2}(x_{1},1)-\hat{p}_{2}(x_{0},0),\hat{p}_{1}(x_{2},0)-\hat{p}_{1}(x_{0},1),\hat{p}_{2}(x_{2},0)-\hat{p}_{2}(x_{0},1)\big)'$ and its propability limit by $\Delta \boldsymbol{p}(x_{1},x_{2})$. For some weighting matrix $\boldsymbol{W}_{xn}$ with a positive definite probability limit, let $\hat{Q}_{x}(x_{1},x_{2})\equiv \Delta \hat{\boldsymbol{p}}(x_{1},x_{2})'\boldsymbol{W}_{xn}\Delta\hat{\boldsymbol{p}}(x_{1},x_{2})$. Define the estimator $(\hat{x}_{m1},\hat{x}_{m2})$ as any point in $S_{0}^{2}(X)$ such that for some $a_{n}=o(1)$,

linenomath*\begin{equation} \hat{Q}_{x}(\hat{x}_{m1},\hat{x}_{m2})\leq \inf_{S_{0}^{2}(X)} \hat{Q}_{x}(x_{1},x_{2})+a_{n}^{2} \end{equation}

This estimator is adapted from chernozhukov2007estimation for partially identified parameters. It is applicable here because the solution to the generalized propensity score matching may not be unique. For simplicity, in this section I focus on the case where $(x_{m1},x_{m2})$ is unique and let $a_{n}=0$. Consistency and asymptotic normality of $(\hat{x}_{m1},\hat{x}_{m2})$ then follow from the standard arguments for (local) GMM estimators under Assumption (ref). The proofs are omitted. The general case with multiple solutions to the matching and $a_{n}>0$ is discussed in Appendix (ref) in the Supplemental Material.

assumptionLet $f_{XZ}\equiv \mathbb{P}_{Z|X}\cdot f_{X}$ and $f_{DXZ}\equiv \mathbb{P}_{D|XZ}\cdot f_{XZ}$. For every $d\in S(D)$ and $z\in S(Z)$, $f_{DXZ}(d,\cdot,z)$ and $f_{XZ}(\cdot,z)$ are three times continuously differentiable on $S(X)$ with bounded derivatives, and are bounded away from zero on $S_{0}(X)$. The kernel $K(\cdot)$ is positive, symmetric at $0$, continuously differentiable on $\mathbb{R}$ with bounded derivative, and supported on $[-1,1]$.

Under Assumption (ref), it can be shown that $||(\hat{x}_{m1},\hat{x}_{m2})-(x_{m1},x_{m2})||=o_{p}(1)$. For the asymptotic distribution, denote the Jacobian matrix of $\Delta\boldsymbol{p}(x_{1},x_{2})$ evaluated at $x_{m1}$ and $x_{m2}$ by $\partial_{x'}\Delta\boldsymbol{p}(x_{m1},x_{m2})$. Let $\tilde{\boldsymbol{z}}_{1},...,\tilde{\boldsymbol{z}}_{4}$ be $(x_{0},0)$, $(x_{0},1)$, $(x_{m1},1)$ and $(x_{m2},0)$ respectively. Let $\kappa\equiv \int K(v)^{2}dv$ and $\Sigma_{x}=\kappa

pmatrix[pmatrix omitted — 77 chars of source]

$, where for $k=1,2$, \[\Sigma_{xk}=

pmatrix[pmatrix omitted — 946 chars of source]

.\] Using the standard two step GMM procedure, let $\hat{\Sigma}_{x}^{-1}$ be the feasible optimal weighting matrix obtained by a consistent first step estimator (for instance using the identity matrix as the weighting matrix). If $(x_{m1},x_{m2})$ lies in the interior of $S^{2}_{0}(X)$ and $\partial_{x'}\Delta\boldsymbol{p}(x_{m1},x_{m2})$ is nonsingular, $(\hat{x}_{m1},\hat{x}_{m2})$ has the following asymptotic distribution under undersmoothing $h_{x}^{2}\cdot \sqrt{nh_{x}}=o(1)$ and $\sqrt{nh_{x}^{3}}\to\infty$ (guaranteeing the derivatives of $\hat{\bm{p}}(z,\cdot)$ are also uniformly consistent):

linenomath*\begin{equation} \sqrt{nh_{x}}\begin{pmatrix} \hat{x}_{m1}-x_{m1}\\ \hat{x}_{m2}-x_{m2} \end{pmatrix}\overset{d}{\to} \mathcal{N}\left(0,\left[\partial_{x'}\Delta\boldsymbol{p}(x_{m1},x_{m2})\Sigma_{x}^{-1}\partial_{x}\Delta\boldsymbol{p}(x_{m1},x_{m2})\right]^{-1}\right) \end{equation}

Since only one covariate is in the model, each matching point is a scalar that matches two generalized propensity scores. So $(x_{m1},x_{m2})$ is overidentified. The null hypothesis $\mathbb{H}_{0}: \Delta\boldsymbol{p}(x_{m1},x_{m2})=\boldsymbol{0}$ can be tested by the J test

linenomath*\[\mathcal{J}_{x}=nh_{x} \Delta\hat{\boldsymbol{p}}(\hat{x}_{m1},\hat{x}_{m2})'\widehat{\Sigma}_{x}^{-1}\Delta\hat{\boldsymbol{p}}(\hat{x}_{m1},\hat{x}_{m2})\]

Under the null, $\mathcal{J}_{x}\overset{d}{\to} \chi_{2}^{2}$. In addition to jointly testing whether $(x_{m1},x_{m2})$ solves the propensity score matching equation, one can separately test either one of them when needed. By block-diagonality of the asymptotic variance, $\hat{x}_{m1}$ and $\hat{x}_{m2}$ are asymptotically independent, and thus it is equivalent to estimate the two matching points separately. In each separate problem, the matching point is still overidentified. Let $\mathcal{J}_{x1}=nh_{x}\Delta\hat{\boldsymbol{p}}(\hat{x}_{m1})'(\kappa\widehat{\Sigma}_{x1})^{-1}\Delta\hat{\boldsymbol{p}}(\hat{x}_{m1})$ and $\mathcal{J}_{x2}=nh_{x}\Delta\hat{\boldsymbol{p}}(\hat{x}_{m2})'(\kappa\widehat{\Sigma}_{x2})^{-1}\Delta\hat{\boldsymbol{p}}(\hat{x}_{m2})$, where $\Delta\hat{\boldsymbol{p}}(\hat{x}_{m1})$ and $\Delta\hat{\boldsymbol{p}}(\hat{x}_{m2})$ are subvectors of $\Delta\hat{\boldsymbol{p}}(\hat{x}_{m1},\hat{x}_{m2})$ containing its first and last two elements respectively. Each test statistic converges in distribution to $\chi_{1}^{2}$ under the null.

Estimating the Separable Model

For the separable model, linearity of the moment conditions (ref) yields a closed form estimator. Assume $\Pi_{SP}\equiv \left(\bm{p}(x_{0},0),\bm{p}(x_{0},1),\bm{p}(x_{m1},0),\bm{p}(x_{m2},1)\right)'$ is full rank. Let $\hat{\delta}_{(x_{0},0),(\hat{x}_{m1},1)}=\widehat{\mathbb{E}}_{Y|DXZ}(d,\hat{x}_{m1},1)-\widehat{\mathbb{E}}_{Y|DXZ}(d,x_{0},0)$ and $\hat{\delta}_{(x_{0},1),(\hat{x}_{m2},0)}$ be defined similarly. Then with a weighting matrix $\bm{W}_{mn}$ that has a positive definite probability limit, let

linenomath*\begin{equation} \hat{\boldsymbol{m}}(x_{0}) =(\widehat{\Pi}_{SP}'\boldsymbol{W}_{mn}\widehat{\Pi}_{SP})^{-1}\cdot \widehat{\Pi}_{SP}'\boldsymbol{W}_{mn}\widehat{\Phi} \end{equation}

where $\widehat{\Pi}_{SP}=\left(\hat{\bm{p}}(x_{0},0),\hat{\bm{p}}(x_{0},1),\hat{\bm{p}}(\hat{x}_{m1},0),\hat{\bm{p}}(\hat{x}_{m2},1)\right)'$ and \[ \widehat{\Phi}=

pmatrix[pmatrix omitted — 448 chars of source]

. \] The propensity score estimators are as equation (ref). The conditional expectations can be estimated by the Nadaraya-Watson estimator:

linenomath*\begin{equation} \widehat{\mathbb{E}}_{Y|DXZ}(d,x,z)=\frac{\sum_{i=1}^{n}Y_{i}\mathbbm{1}(D_{i}=d)K(\frac{X_{i}-x}{h_{m}})\mathbbm{1}(Z_{i}=z)}{\sum_{i=1}^{n}\mathbbm{1}(D_{i}=d) K(\frac{X_{i}-x}{h_{m}})\mathbbm{1}(Z_{i}=z)} \end{equation}
assumptionThe conditional variance of $Y$, $\mathbb{V}_{Y|DXZ}(d,\cdot,z)$, is finite and continuous on $S(X)$ for each $(d,z)\in S(D,Z)$. The conditional expectation $\mathbb{E}_{Y|DXZ}(d,\cdot,z)$ is three times continuously differentiable on $S(X)$ with bounded derivatives.

Under Assumptions (ref) and (ref), every component in the right hand side of equation (ref) is uniformly consistent on $S_{0}(X)$. Under consistency of $(\hat{x}_{m1},\hat{x}_{m2})$, $||\hat{\boldsymbol{m}}(x_{0})-\boldsymbol{m}^{*}(x_{0})||=o_{p}(1)$.

For the asymptotic distribution, I let $h_{m}/h_{x}\to 0$ so that the impacts of estimating $(x_{m1},x_{m2})$ and the generalized propensity scores are negligible. Let $\tilde{\boldsymbol{z}}_{1},...,\tilde{\boldsymbol{z}}_{6}$ be $(x_{0},0)$, $(x_{m1},0)$, $(x_{m1},1)$, $(x_{0},1)$, $(x_{m2},1)$ and $(x_{m2},0)$. Let $\Sigma_{SP}=\kappa(\Sigma_{SP,1}+\Sigma_{SP,2}+\Sigma_{SP,3})$ and $\Sigma_{SP,d}$ ($d=1,2,3$) equal

linenomath*\begin{equation*} \setlength\arraycolsep{0pt} \begin{pmatrix} \frac{p_{d}(\tilde{\boldsymbol{z}}_{1})^{2}\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{1})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{1})}&0& \frac{p_{d}(\tilde{\boldsymbol{z}}_{1})p_{d}(\tilde{\boldsymbol{z}}_{2})\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{1})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{1})}&0\\ 0 &\frac{p_{d}(\tilde{\boldsymbol{z}}_{4})^{2}\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{4})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{4})}&0&\frac{p_{d}(\tilde{\boldsymbol{z}}_{4})p_{d}(\tilde{\boldsymbol{z}}_{5})\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{4})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{4})}\\ \frac{p_{d}(\tilde{\boldsymbol{z}}_{1})p_{d}(\tilde{\boldsymbol{z}}_{2})\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{1})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{1})}&0 & \sum_{k=1}^{3}\frac{p_{d}(\tilde{\boldsymbol{z}}_{2})^{2}\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{k})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{k})}&0\\ 0& \frac{p_{d}(\tilde{\boldsymbol{z}}_{4})p_{d}(\tilde{\boldsymbol{z}}_{5})\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{4})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{4})}& 0 & \sum_{k=4}^{6}\frac{p_{d}(\tilde{\boldsymbol{z}}_{5})^{2}\mathbb{V}_{Y|DXZ}(d,\tilde{\boldsymbol{z}}_{k})}{f_{DXZ}(d,\tilde{\boldsymbol{z}}_{k})} \end{pmatrix}. \end{equation*}

Let $W_{mn}=\hat{\Sigma}_{SP}^{-1}$. Then if $h_{m}^{2}\cdot \sqrt{nh_{m}}\to 0$ and $\sqrt{nh_{m}^{3}}\to\infty$,

linenomath*\begin{equation} \sqrt{nh_{m}}\left(\hat{\boldsymbol{m}}(x_{0})-\boldsymbol{m}^{*}(x_{0})\right) \overset{d}{\to} \mathcal{N}\big(0,(\Pi_{SP}'\Sigma_{SP}^{-1}\Pi_{SP})^{-1}\big) \end{equation}

Since there are four moment equations, $\bm{m}^{*}(\bm{x}_{0})$ is overidentified. Let the overidentification test statistic be $\mathcal{J}_{SP}=nh_{m}\big(\widehat{\Pi}_{SP}\hat{\boldsymbol{m}}(x_{0})-\widehat{\Phi}_{SP}\big)'\widehat{\Sigma}_{SP}^{-1}\big(\widehat{\Pi}_{SP}\hat{\boldsymbol{m}}(x_{0})-\widehat{\Phi}_{SP}\big)$. The null hypothesis is that all of the four moment conditions hold. Validity of the moment conditions jointly depend on the exogeneity of the instrument and validity of the matching points obtained by generalized propensity score matching. Under the null, it can be verified that $\mathcal{J}_{SP}\overset{d}{\to} \chi_{1}^{2}$ following the standard argument in the GMM framework.

Monte Carlo Simulations

This section illustrates the finite sample performance of the estimator. The endogenous variable $D$ follows the ordered choice model in Example (ref). The instrument $Z$ is binary. Let the outcome variable $Y$ be determined by the following model:

linenomath*\begin{align*} Y=[\gamma_{1}\mathbbm{1}(D=1)+&\gamma_{2}\mathbbm{1}(D=2)+\gamma_{3}\mathbbm{1}(D=3)]\cdot (X+1)+U \end{align*}

where $X$ is drawn from $\text{Unif}[-3,3]$, $Z$ from a Bernoulli distribution with parameter $0.5$, $[U,V]$ from $\mathcal{N}\left(0,

pmatrix[pmatrix omitted — 29 chars of source]

\right)$, and $X\mathpalette{\independenT}{\perp} Z\mathpalette{\independenT}{\perp} (U,V)$.

For the parameters, I set $(\gamma_{1},\gamma_{2},\gamma_{3},\kappa_{1},\kappa_{2})=(1.5,3,3.5,-0.7,0.1)$. The parameters $(\alpha,\beta)$ govern the strength of the instrument and the covariate. In this section I present the results under $(\alpha,\beta)=(0.8,0.4)$. The parameter values are selected for two reasons: i) all the generalized propensity scores are away from 0 so that in the simulated sample, there are a sufficient number of observations to estimate the conditional expectation and the generalized propensity score for each $d$, and ii) $X$ and $Z$ have large effects on the generalized propensity scores. Finally, I set $\rho=0.5$ and $x_{0}=0$. Additional simulation results for small $(\alpha,\beta)$, different $\rho$ and different $x_{0}$ are provided in Appendix (ref) in the Supplemental Material.

table[table omitted — 2,309 chars of source]

Table (ref) contains the results for sample size $n=1000$, 2000 and 3000. The number of simulation replications is 500. In each replication, I estimate $(x_{m1},x_{m2})$ using grid search with $500$ grid nodes. The propensity scores and the conditional expectations are estimated as proposed in Section (ref) with the biweight kernel. A smaller bandwidth is chosen when estimating the outcome function than the one used to estimate the matching points. The actual coverage probabilities of the confidence intervals for $\boldsymbol{m}^{*}(x_{0})$ are computed using the asymptotic variance estimator. The coverage probabilities of the overidentification tests for $(x_{m1},x_{m2})$ and for $\boldsymbol{m}^{*}(x_{0})$ are also reported.

As is shown in Table (ref), the variance of the estimator dominates in mean squared error (MSE) due to undersmoothing. The actual coverage probabilities are close to the nominal values for both the outcome function and the overidentification tests.

Empirical Applications

In this section, I use two empirical examples to illustrate the value and limitations of my approach. Section (ref) presents an application of the return to education. Section (ref) uses an example of the preschool program choice to illustrate when a matching point does not exist.

The Return to Education

In this application, I use the same extract from the 1979 NLS dataset as in card1995using and adopt the same instrument. The instrument equals $1$ if an individual grew up near an accredited four year college, and equals $0$ otherwise. The outcome variable $Y$ is the log wage and is assumed to be determined by a separable model. I use the average of parents' years of schooling as the matching covariate $X$, and $x_{0}$ is set equal to $10,11$ and $12$. Finally, I drop the individuals who were still enrolled in a school at the time of the survey. The remaining sample size is 2000.

I construct $D$ by dividing individuals' (or, children's) years of schooling into either two or three categories. In the case of a binary $D$, both my approach and the standard nonparametric IV approach can identify and estimate the outcome function at each level of $D$. Comparing the results obtained by these two approaches, they yield similar point estimates, but my approach attains smaller variances. When $D$ takes on three values, no existing method would work without imposing additional structures on the outcome function. My estimates are coherent with the empirical literature.

A Binary $D$

In this subsection, I assume that the latent selection mechanism only yields two outcomes: $D=1$ if an individual's years of schooling is greater than 12 and $D=0$ otherwise. As $|S(D)|=|S(Z)|=2$, at a fixed $x_{0}$, the outcome function $\bm{m}^{*}(x_{0})\equiv ({m}_{0}^{*}(x_{0}),{m}_{1}^{*}(x_{0}))'$ is just-identified by the standard nonparametric IV approach, and overidentified by my approach if a matching point exists.

For each value of $x_{0}$, I try to find two matching points $x_{m1}$ and $x_{m2}$ such that $(x_{0},0)$ and $(x_{m1},1)$ are a matching pair, and $(x_{0},1)$ and $(x_{m2},0)$ are another. As $D$ only takes on two values, each matching point only needs to match one propensity score. The propensity score is estimated as proposed in Section (ref). The kernel and the bandwidth follow those used in Section (ref).

Figure (ref) illustrates the propensity score matching for $x_{0}=12$. The red solid curves in the left and the right panels are $\hat{p}_{0}(x,1)-\hat{p}_{0}(12,0)$ and $\hat{p}_{0}(x,0)-\hat{p}_{0}(12,1)$ respectively. The intersection points of these propensity score differences with zero are the estimated matching points. The patterns for $x_{0}=10$ and $11$ are similar to Figure (ref) thus omitted. Figure (ref) implies that individuals whose parents have more years of schooling are more likely to attain post-high school education. From the values of the matching points, living close to a four year college and parents' education are substitutes for an individual's educational attainment. At $X=12$, an increase of about half a year in parents' education compensates for living far from a college.

figure[figure omitted — 186 chars of source]

Now the outcome function can be estimated using two approaches. The results are shown in Table (ref). The second row Matching indicates whether the matching points are estimated and used. When not using the matching points, $\bm{m}^{*}(x_{0})$ is estimated by the nonparametric IV approach by solving the sample analogue of equation (ref) with $(x,z)=(x_{0},0)$ and $(x_{0},1)$, and $\Delta_{d}(x_{0},x_{0})=0$. The standard errors in parentheses are computed using the asymptotic variance estimators. The $p$-values of the overidentification test for the outcome function are reported in the last row when applicable.

table[table omitted — 801 chars of source]

From Table (ref), we can make three observations. First, the estimates using the two approaches are very close. It provides evidence that the additional moment conditions brought in by the matching points are valid. The insignificant overidentification tests also suggest that both the instrument and the matching points are valid. Note that the overidentification test is unavailable when using the nonparametric IV approach. Second, the variances are lower using the new approach. Variance reduction is due to the use of more moment conditions. Consequently, the estimated effects can be more significant. For instance, though not reported here, $\hat{m}_{1}(12)-\hat{m}_{0}(12)$ is significant at 10% level using the IV approach but is significant at 1% level using my approach. Third, the return is increasing in the level of own education and heterogeneous in parents' education.

A Three-Valued $D$

Now assume the selection model yields three outcomes. The baseline level of $D$ is still at most high school but recoded by $D=1$. Post-high school education is further divided into two groups: $D=2$ if $12<\text{years of schooling}\leq 15$ (some college), and $D=3$ if $\text{years of schooling}> 15$ (college and above). In this case, no existing method can identify and estimate $\boldsymbol{m}^{*}(x_{0})$ without imposing additional assumptions on it.

Figure (ref) illustrates the matching points for $x_{0}=12$. Again, the plots for the other values of $x_{0}$ are omitted as the patterns are similar. The solid red curves are $\hat{p}_{1}(x,1)-\hat{p}_{1}(x_{0},0)$ and $\hat{p}_{1}(x,0)-\hat{p}_{1}(x_{0},1)$, while the dashed blue curves are $\hat{p}_{3}(x,1)-\hat{p}_{3}(x_{0},0)$ and $\hat{p}_{3}(x,0)-\hat{p}_{3}(x_{0},1)$. In theory, matching is successful if the solid and the dashed curves intersect with the horizontal line of zero at the same point. From the figure, the intersection points are indeed very close in both panels. The overidentification tests also support successful matching; $\mathcal{J}_{x1}$ and $\mathcal{J}_{x2}$ reported on top of the plots are insignificant in both cases. Finally, since the baseline level here (years of schooling $\leq 12$) is defined in the same way as in the case of a binary $D$, its propensity score is also equal to the previous case. Since this propensity score has to be matched in both cases, the matching points in these cases should be identical. Here the estimates are 11.54 and 12.34, indeed very close to those when $D$ is binary (11.47 and 12.37).

figure[figure omitted — 200 chars of source]

Next, let us turn to $\hat{\bm{m}}(x_{0})$ shown in Table (ref). In this case both the outcome function and the matching points are overidentified. The $p$-value for each overidentification test is presented in the bottom panel. First, we can see that none of the overidentification tests for $\bm{m}^{*}(x_{0})$ is significant at any reasonable level, similar to Table (ref) for the binary case. Meanwhile, the joint overidentification tests for the matching points are also insignificant, confirming that the single covariate matches all the generalized propensity scores. Second, the return to education is monotonic in the level of own education and heterogeneous in parents' years of schooling, while the difference in returns across adjacent own education levels is decreasing.

table[table omitted — 641 chars of source]

When Does the Matching Fail?

As shown in Lemma (ref), a matching pair necessarily matches all the generalized propensity scores. Matching may fail if the instrument has dominant effects on the generalized propensity scores such that shifting the covariates cannot compensate for those effects. As my identification strategy treats local variation in the covariates like instruments, one would hope that both such variation and the original instrument have comparably strong effects on selection. When the former is too weak, it loses identification power and this paper's approach would fail.

For illustration, I consider an application on preschool program selection, following kline2016evaluating using the Head Start Impact Study (HSIS) dataset. The endogenous variable $D$ takes on three values: participating in Head Start ($h$), participating in another competing preschool program ($c$), and not participating in any preschool programs ($n$). The binary instrument $Z$ is a random lottery granting access to Head Start. Candidates for the covariate $X$ are family income, baseline test score and the centers' quality index.

figure[figure omitted — 189 chars of source]

Figure (ref) shows the estimated generalized propensity scores using the baseline test score as $X$ and $x_{0}$ is equal to the sample median. Findings under other values of this covariate or using other covariates are similar. We see that if an individual wins the lottery, the probability of attending Head Start is very high, and not much affected by the baseline test score. On the contrary, when not winning the lottery, the individual would most likely not participate in any program, and in particular, the probability of attending Head Start is lowest for almost any baseline test score. A matching point does not exist in this example because shifting $X$ never offsets the dominant effects of $Z$ on the generalized propensity scores.

Relation to the Existing Literature

Triangular Models

Techniques that achieve point identification of triangular models often require the endogenous variable $D$ to be continuous. Different approaches are developed depending on whether $Z$ is also continuous or discrete.

Continuous $D$ and $Z$. The widely used control function approach usually needs a continuous instrument (e.g. newey1999nonparametric, chesher2003identification, florens2008identification, imbens2009identification, etc.). This approach allows the outcome heterogeneity to be multidimensional, but the first stage heterogeneity needs to be a scalar with the selection function $h$ strictly increasing in it.

A continuous $D$ and a binary $Z$. d2015identification and torgovitsky2015identification show that identification of a nonseparable outcome function increasing in the scalar disturbance can be achieved with a binary $Z$. gunsilius2018point extends the model to allow for multidimensional heterogeneity. Compared with my approach, I need to use covariates as an additional source of variation, but I allow for a discrete $D$ and more flexible selection heterogeneity.

Multidimensional $D$ (with a continuous component) and a binary $Z$. huang2019identification consider a triangular model with a separable outcome function where there are two endogenous variables and a single instrument that can be binary. In this respect, their paper focuses on a similar problem to mine since in my setup, a discrete $D$ is equivalent to multiple dummy endogenous variables while there is only one binary instrument. However, like the other papers discussed, they also need one of the endogenous variables to be continuous at least on a subset of its support, with a first stage where the unobservable must be a scalar.

Single Equation Approaches

Single equation approaches refer to methods that achieve identification without explicitly relying on a selection model. The typical IV approach for nonparametric identification falls into this category (e.g. newey2003instrumental, das2005instrumental, chernozhukov2005iv, chernozhukov2007instrumental, chen2014local, ect.). This approach requires $Z$ to have large support.

caetano2016identifying develop a strategy that achieves identification using small-support $Z$ when $D$ is multivalued. Their method does not rely on selection models. Similar to this paper, they utilize the covariates for identification, but the covariates need to be structurally separable from the model. Taking the nonseparable model with a discrete $D$ as an example, they essentially impose a single index structure: $Y=g^{*}_{D}(U)$ and $U=\phi(\boldsymbol{X},U_{0})$, where both $U$ and $U_{0}$ are unobserved. The function $\phi$ is real valued and $g^{*}_{D}(\cdot)$ is strictly increasing. In contrast, I allow all the covariates to enter the model in arbitrary ways, but a selection model, though very general, is needed. These two approaches are complementary.

Concluding Remarks

In this paper, I develop a novel approach to use covariates to identify structural outcome functions in a triangular model when the discrete endogenous variable takes on more values than the instrument. This paper illustrates that information on endogenous selection has large identifying power. The generalized propensity scores can provide useful information on the degree of endogeneity indexed by covariates-instrument combination. Across such combinations at which the endogenous variable has the same degree of endogeneity, extrapolation can be made to supplement the insufficient information from the instrument and facilitate identification.

Moving forward, it would be of interest to apply the idea in this paper to other scenarios, such as models with limited dependent variables, extrapolation in regression discontinuity designs, etc. Another direction is to generalize the outcome function by allowing for multidimensional heterogeneity.