EconBase
← Back to paper

Local Identification in Instrumental Variable Multivariate Quantile Regression Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

42,028 characters · 8 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Local Identification in Instrumental Variable Multivariate Quantile Regression Models

abstractIn the instrumental variable quantile regression (IVQR) model of chernozhukov2005iv, a one-dimensional unobserved rank variable monotonically determines a single potential outcome. In practice, when researchers are interested in multiple outcomes, it is common to estimate separate IVQR models for each of them. This approach implicitly assumes that the rank variable in each regression affects only its associated outcome, without influencing others. In reality, however, outcomes are often jointly determined by multiple latent factors, inducing structural correlations across equations. To address this limitation, we propose a nonlinear instrumental variable model that accommodates multivariate unobserved heterogeneity, where each component of the latent vector acts as a rank variable corresponding to an observed outcome. When both the treatment and the instrument are discrete, we show that the structural function in our model is locally identified under a sufficiently strong positive correlation between the treatment and the instrument.

Introduction

The instrumental variable quantile regression (IVQR) model introduced by chernozhukov2005iv has become a widely used tool for estimating quantiles of potential outcomes in the presence of endogeneity. See, for example, chernozhukov2004effects and autor2017effect. While the IVQR framework assumes a one-dimensional outcome variable, researchers are often interested in settings with multiple outcomes. For instance, chernozhukov2004effects examine the effect of 401(k) participation on several wealth measures, including total wealth and financial assets. In such cases, it is common to estimate separate IVQR models for each outcome dimension. However, this practice raises several conceptual and empirical concerns.

First, running separate models ignores the correlation structure among outcome variables and is therefore silent about their joint distribution. In the 401(k) example, this approach cannot capture the share of total wealth accounted for by financial assets.

More importantly, estimating separate quantile models makes rank similarity, a crucial assumption of chernozhukov2005iv, unrealistic. To apply the IVQR model in the 401(k) example, one must assume that individuals' holdings of financial assets depend solely on their preferences for those assets. In reality, individuals allocate wealth across multiple assets simultaneously, and their preferences over financial assets are inherently linked to preferences over other asset categories.

To address these limitations, we propose a new nonlinear model that accommodates multidimensional outcomes that may be correlated with each other. Specifically, we consider a multivariate extension of the potential outcome framework in chernozhukov2005iv. For a treatment $D = d$, let $Y_d$ denote the corresponding $p$-dimensional potential outcome. Conditional on covariates, we assume it can be represented as $Y_d = q_d^\ast (U_d)$, where $q_d^\ast$ is a structural function and $U_d$ is a $p$-dimensional rank variable. The rank variable $U_d$ is a random vector that captures unobserved heterogeneity among observationally identical individuals, and can be interpreted as reflecting latent individual characteristics.

Existing studies of endogenous quantile models impose various monotonicity restrictions on the structural function for identification. Extending the assumption in chernozhukov2005iv that the structural function is the quantile function of the potential outcome, we require its derivative to be symmetric and positive definite. This restriction implies that an increase in the $i$th component of $U_d$ increases the $i$th component of $Y_d$, while also allowing cross-dimensional effects: changes in one component of $U_d$ can affect other components of $Y_d$ positively or negatively. Thus, our framework can capture substitutability and complementarity between outcome dimensions---features that the standard IVQR approach cannot accommodate. Section (ref) presents two examples of structural functions satisfying this symmetry and positive definiteness condition.

Vector-valued functions with symmetric and positive definite derivatives play a central role in optimal transport theory, where they characterize maps minimizing quadratic transportation costs. Recent developments in statistics have shown that such maps possess many desirable properties that justify viewing them as multivariate analogues of quantile functions (see, e.g., ekeland2012comonotonic, chernozhukov2017monge, hallin2021distribution, ghosal2022multivariate). This perspective motivates us to refer to our framework as an instrumental variable multivariate quantile regression model.

For discrete treatments and instruments, our identification result generalizes the key insight of chernozhukov2005iv. They show that to identify quantiles of one-dimensional potential outcomes under binary treatment, an instrument with binary or richer support satisfying the full-rank condition is required. For higher-dimensional potential outcomes, one might expect that instruments with more than two support points are necessary. Surprisingly, our main result, Theorem (ref), shows that an instrument with only binary support suffices for identification. However, this comes at the cost of a stronger relevance condition: the instrument must be sufficiently positively correlated with the treatment, and the strength required increases with the dimension of the outcome vector.

Related literature. Our model builds on the IVQR framework proposed and developed by chernozhukov2005iv, chernozhukov2006instrumental, chernozhukov2007instrumental, and chernozhukov2013quantile. Unlike these studies, we allow for multidimensional outcomes and multidimensional unobserved heterogeneity.

Several related papers consider endogeneity in models with a single outcome variable. abadie2002instrumental analyze quantile treatment effects using the local average treatment effect framework of imbens1994identification. Their approach identifies treatment effects for compliers but is limited to binary treatments. Extending chesher2003identification, imbens2009identification adopt a control function approach under the assumption that the selection equation is monotone and the instrumental variable is independent of the unobserved disturbance. Nonparametric identification in triangular models with discrete instruments is studied by torgovitsky2015identification, and further extended to multivariate settings by gunsilius2023condition. matzkin2008identification investigates nonparametric identification in simultaneous equation models, where endogenous variables are continuous and their dimension must coincide with that of unobserved heterogeneity. In contrast, our framework accommodates discrete treatments and allows for arbitrary randomness in the treatment assignment process. Within the same simultaneous equation framework, blundell2017individual propose a nonseparable model that exploits proxy variables for unobserved heterogeneity.

Our key restriction on the structural function requires it to be a multivariate quantile function, as developed by ekeland2012comonotonic, chernozhukov2017monge, hallin2021distribution, and ghosal2022multivariate. Multivariate quantile functions are grounded in optimal transport theory (villani2003topics, villani2009optimal), which provides a natural framework for extending quantiles to random vectors. galichon2018optimal offers a comprehensive overview of economic applications of optimal transport. In this paper, we draw on results from this literature to establish the identification of our model.

Organization: This paper is organized as follows. Section (ref) introduces the IV multivariate quantile regression model. Section (ref) formally states the local identification result for the model and discusses the relationship to existing results. Section (ref) concludes. All proofs and mathematical preliminaries are provided in the Appendix.

IV multivariate quantile regression model

Model

Let $Y_d$ be a $p$-dimensional potential outcome vector, $D \in \cD$ an endogenous variable (treatment), $X \in \cX$ an exogenous covariate vector, and $Z \in \cZ$ an IV. For now, the treatment and IV can be either discrete or continuous, although we will focus on the discrete cases for identification in the next section. Let $\cU \subset \bR^p$ be a compact convex set with a piecewisely $C^1$ boundary $\partial \cU.$ Let $\mu$ be a reference probability measure of which support is $\cU.$ Assume $\mu$ is absolutely continuous with respect to the $p$-dimensional Lebesgue measure.

Our model extends chernozhukov2005iv so that it accounts for multidimensional outcome vectors. The following is the primitives.

assumptionThe random variables $((Y_d)_{d \in \cD}, D, X, Z, (U_d)_{d \in \cD})$ satisfy the following conditions with probability one: \begin{enumerate} [label=\textcolor{blue}{(A\arabic*)}, ref=A\arabic*] • For each $d \in \cD,$ it holds that $U_d \mid X \sim \mu,$ and there exists a function $q_d^\ast : \cU \times \cX \to \bR^p$ such that $Y_d = q_d^\ast (U_d, X),$ $q_d^\ast$ is continuously differentiable in the first variable, and its derivative $D q_d^\ast$ is symmetric and positive definite on $\mathrm{Int} (\cU) \times \cX.$ • For each $d \in \cD,$ $U_d$ is independent of $Z$ conditional on $X.$$D = \delta(Z, X, \nu)$ for some unknown function $\delta$ and random element $\nu.$ • Conditional on $(X, Z, \nu),$ $(U_d)_{d \in \cD}$ are identically distributed. • The observed random variables consist of $Y \coloneqq Y_D, D, X$ and $Z.$ \end{enumerate}

We refer to this framework as an IV multivariate quantile regression model. The term “multivariate quantile” follows the literature that extends the classical notion of quantile functions to multivariate random variables. Further details are provided in Appendix (ref). See also, for example, carlier2016vector, chernozhukov2017monge, hallin2021distribution, and ghosal2022multivariate.

The difference from chernozhukov2005iv appears in ((ref)). It follows from Theorem (ref) that for any reference measure $\mu,$ there exists $q_d^\ast$ with a symmetric and positive semi-definite derivative such that $Y_d = q_d^\ast (U_d, X),$ but it is not necessarily regular as specified in ((ref)). Assumption ((ref)) assumes the smoothness of $q_d^\ast$ and the strict positive definiteness of its derivative. This assumption holds if both $\mu$ and the distribution of $Y_d$ conditional on $X$ are sufficiently regular, according to the regularity theory of optimal transport (see, for example, caffarelli1992regularity). Hence, for any other sufficiently regular reference measure $\tilde \mu,$ condition ((ref)) is satisfied with some $\tilde q_d^\ast$ and $\tilde U_d.$ This implies that any causal parameter defined as a functional of the distributions of the potential outcomes, such as average/quantile treatment effects, is independent of the choice of the reference measure. Also, the continuity of $q_d^\ast$ imposed in ((ref)) implies that the range of $q_d^\ast (\cdot, X)$ is compact and so is the support of $Y_d$ conditional on $X.$

Condition ((ref)) implies that the IV is independent of potential outcomes conditional on covariates. Condition ((ref)) is a weak restriction that allows a broad class of assignment rules of treatments. In particular, $\nu$ in ((ref)) can be correlated with potential outcomes.

Condition ((ref)) is called the rank similarity (chernozhukov2005iv). The simplest form of rank similarity is rank invariance that requires $U_d = U$ for all $d \in \cD.$ In the wealth accumulation example discussed in Section (ref) and (ref), the rank invariance assumes that an individual's preference does not change regardless of their participation to 401(k). On the other hand, the rank similarity assumes that the preference may change depending on the participation status, but the individual cannot predict the change before deciding whether to participate 401(k) or not. For more details of these assumptions, see chernozhukov2005iv and chernozhukov2013quantile.

One limitation of the IV multivariate quantile regression model is that Condition ((ref)) depends on the choice of the reference measure $\mu$. In particular, the fact that Condition ((ref)) holds for a given $\mu$ does not guarantee that it will hold for another measure $\tilde{\mu}$. This dependence implies that the distribution $\mu$ of unobserved heterogeneity must be specified a priori. Although strong, such an assumption is common in related literature. For example, chernozhukov2021identification impose a similar condition for identifying hedonic equilibrium models, and also, discrete-choice models typically rely on a fixed distribution for unobserved heterogeneity.

The unobserved random vector $U_d$ captures heterogeneity in outcomes among observationally identical individuals. It is interpreted as a rank variable in the sense that, conditional on covariates $X,$ the $i$th component of $Y_d$ is monotonically determined by the corresponding component of $U_d.$ In particular, under the positive definiteness of $D q_d^\ast$ imposed in Assumption ((ref)), we have $\partial q_d^{\ast i}(u_d, x)/\partial u_d^i > 0$, where $q_d^\ast = (q_d^{\ast 1}, \dots, q_d^{\ast p})^\prime$ and $u_d = (u_d^1, \dots, u_d^p)^\prime$. Moreover, unlike in the one-dimensional setting, the $i$th component of $U_d$ may also influence the $j$th component of $Y_d$ for $i \neq j$. Specifically, the cross-partial derivative $\partial q_d^{\ast j}(u_d, x)/\partial u_d^i$ may be either positive or negative, as long as $D q_d^\ast$ remains positive definite. This feature enables the model to capture flexible correlations and potential substitutability or complementarity across different dimensions of the outcome vector.

Technically, any absolutely continuous probability measure on a compact convex set $\cU$ can be chosen as the reference measure $\mu$. In the absence of a strong reason to prefer a particular specification, it is natural to take $\cU = [0, 1]^p$ and $\mu = U[0, 1]^p$, as in carlier2016vector. Alternatively, one may consider the $p$-dimensional unit ball as $\cU$ and the spherical uniform distribution as $\mu$, following chernozhukov2017monge and hallin2021distribution.

By condition ((ref)), the observed outcome is written as the structural form $Y = q_D^\ast (U, X)$ where $U \coloneqq U_D.$ The following representation is the main testable implication of the model.

theoremSuppose Assumption (ref) holds. Then, it holds with probability one that for each measurable $B \subset \cU,$ \begin{align} \bP ( Y \in q_D^\ast (B, X) \mid X, Z ) = \mu (B) \end{align} where $q_D^\ast (B, X) = \{q_D^\ast (u, X) \mid u \in B\}.$ In particular, $U \mid X, Z \sim \mu$ holds.

As the LHS of ((ref)) is determined by the joint distribution of the observable variables, the equation gives a conditional moment restriction. In particular, any candidate structural function $(q_d){d \in \cD}$ must satisfy ((ref)) with $q_d^\ast$ replaced by $q_d$. For identification, we explore conditions under which $(q_d^\ast)_{d \in \cD}$ is the only function that satisfies the equation in the subsequent section.

For notational simplicity, let $q_d^\ast(u) \coloneqq q_d^\ast(u, x)$, suppressing the dependence on covariates $x$. All subsequent analysis should be understood as conditional on $X = x$. This simplification does not affect the identification results, since the conditional distribution of the observables given $X$ is identifiable.

When outcomes are univariate and the reference distribution is the uniform distribution on the unit interval, as in chernozhukov2005iv, $q_d^\ast (\tau)$ is the $\tau$-quantile of $Y_d$ and satisfies

align*[align* omitted — 102 chars of source]

for all $z \in \cZ,$ which corresponds to setting $\mu = U [0, 1]$ and $B = [0, \tau]$ in ((ref)). If the treatment variable is supported on a finite set $\cD = \{0, \dots, m - 1\}$, this is a simultaneous equation with $m$ unknown variables $q_0^\ast (\tau), \dots, q_{m - 1}^\ast (\tau).$ Hence, if the instrument $Z$ takes on more than or equal to $m$ values and the equation system is non-degenerate, it is expected that the solution is unique, at least locally. Indeed, chernozhukov2005iv show that this observation is correct. Checking the condition for each quantile level $\tau$ establishes the nonparametric identification of the structural function $q_d.$

This approach works in one dimension because, for each $\tau$, the values $(q_d^\ast(\tau))_{d \in \cD}$ are determined by the conditional moment restriction ((ref)) when we take the set $B = [0, \tau]$. In higher dimensions, however, this argument no longer applies. To see why, consider the case $p = 2$ with $\mu = U[0,1]^2$ and fix $(\tau^1, \tau^2) \in [0,1]^2$. For any set $B \subset [0,1]^2$, equation ((ref)) does not isolate $(q_d^\ast(\tau^1, \tau^2))_{d \in \cD}$ alone, because the image $q_d^\ast(B)$ can be fully nonlinear. In particular, the image depends not only on the value of $q_d^\ast$ at $(\tau^1, \tau^2)$ but also on its behavior along the boundary $\partial B$. Since the set $q_d^\ast(\partial B)$ is an infinite-dimensional object, evaluating ((ref)) for a fixed $B$ does not yield identification when the instrument has finite support. To address this issue, we instead interpret ((ref)) as a measure-valued equation, allowing us to exploit the relationships between $q_d^\ast(B)$ and $q_d^\ast(B')$ for different sets $B$ and $B'$. Such relationships are unnecessary in the univariate case but are crucial for identifying multivariate structural functions.

Examples

Example 1. We first consider an extension of the saving model proposed by chernozhukov2004effects. The paper investigates the effects of 401(k) participation on wealth accumulation. As measures of wealth, they consider total wealth, net financial assets, and net non-401(k) financial assets. Since chernozhukov2004effects apply the IVQR model to these variables separately, their analysis does not capture the correlation among them. For example, while total wealth is defined as net financial assets plus other assets (e.g., housing equity and the value of business, property, and motor vehicles), their model has no implication for how people distribute total wealth into these sub categories. We shall see that our multivariate model can take both variables into consideration simultaneously and allows us to discuss their joint distribution.

Let $U_d = (U_d^F, U_d^O)^\prime \sim \mu$ represent the preference for net financial assets and other assets under participation status $d \in \{0, 1\}.$ Let $y_d = (y_d^F, y_d^O)^\prime$ be net financial assets and other assets, respectively. An individual with type $U_d$ whose portfolio is $y_d$ receives a quasi-linear utility $U_d^\prime y_d + \psi_d (y_d),$ where $\psi_d$ is a deterministic concave function.\footnote{Notice that covariates are suppressed for simplicity.} Assuming an interior solution, the individual chooses the optimal portfolio $Y_d = (- \nabla \psi_d)^{-1} (U_d),$ which follows by the first order condition. It can be seen that the function $q_d (u) \coloneqq (- \nabla \psi_d)^{-1} (u)$ has a symmetric and positive definite derivative by the concavity of $\psi_d.$

The individual determines whether to participate 401(k) based on the eligibility $Z$ of 401(k) and unobservable heterogeneity $\nu.$ This is represented by $D = \delta (Z, \nu).$ The eligibility variable $Z$ is considered to work as a valid IV. The joint distribution of $(Y \coloneqq Y_D, D, Z)$ fits into the multivariate IVQR model under $Z \perp \!\!\! \perp U_d$ and the rank invariance/similarity.

With $q_d$ identified, we can obtain the joint distribution of net financial assets and other assets. This allows us to discuss the effect of 401(k) participation on the portfolio selection, which is impossible in the standard IVQR framework.

Example 2. We now show how a simple discrete choice problem fits naturally into the IV multivariate quantile framework. This example is inspired by shi2018estimating and fosgerau2020discrete. Consider a retailer offering $p$ differentiated goods in two sets of stores, indexed by $d \in \{0, 1\}.$ The products' qualities are captured by a latent quality vector $U = (U^1, \dots, U^p)^\prime \sim \mu,$ where $\mu$ is assumed known (e.g., from in-store audits of freshness and display quality). Each consumer in district $d$ has an unobservable preference vector $\varepsilon = (\varepsilon^1, \dots, \varepsilon^p)^\prime \sim \nu_d.$ The preference distributions $\nu_0, \nu_1$ are unknown and may differ because district $1$ is exposed to a marketing campaign. By choosing product $i,$ a consumer receives utility $U^i + \varepsilon^i.$

The analyst never observes $U$ or $\varepsilon$. Instead, one observes the aggregate market share vector $s_d (U) \coloneqq (s_d^1 (U), \dots, s_d^p (U))$ if the store is in district $d,$ where $$ s_d^i(u) \coloneqq \int \bI \left\{ i = \mathop{\mathrm{argmax}}_{j=1,\dots,p} (u^j + \varepsilon^j) \right\} \,d\nu_d(\varepsilon). $$ Economically, $s_d (u)$ represents how a marginal improvement in the quality $u$ redistributes sales across all $p$ goods under the taste regime $\nu_d$.

The Williams-Daly-Zachary theorem implies that the surplus function defined as $$ W_d (u) \coloneqq \int \max_{i = 1, \dots p} \left( u^i + \varepsilon^i \right) d \nu_d (\varepsilon) $$ is related to the market share vector via the equation $s_d (u) = \nabla W_d (u)$ (see, e.g., mcfadden1981econometric). It is not hard to see that $W_d$ is a convex function. Also, it is differentiable under mild regularity on $\nu_d.$ Hence, $s_d$ has a symmetric and positive definite derivative.

Let $D$ denote the district, or equivalently the treatment status, ($0$ or $1$) and suppose there exists a valid instrument $Z$ (e.g. randomized ad‐exposure eligibility). Then we observe i.i.d. draws of $(Y, D, Z)$ with $Y = s_D(U),$ $U \sim \mu,$ and $Z \perp \!\!\! \perp U.$ This setup satisfies the rank invariance and exogeneity conditions of the IV multivariate quantile model.

Applying a separate IVQR mode of chernozhukov2005iv to each product discards the rich cross‐product relationships. For instance, one-dimensional approaches cannot impose the natural constraint $\sum_{i=1}^p s_d^i(u) \leq 1,$ nor capture correlations in $U$ arising from similar products facing related quality shocks. The multivariate model preserves these economically meaningful cross‐product patterns while identifying the two demand-quality mappings $s_0$ and $s_1.$

Comparison with chernozhukov2005iv

Even when multiple outcome variables are present, one might consider applying the one-dimensional quantile model of chernozhukov2005iv to each component separately. However, this approach typically leads to the violation of the rank similarity assumption of chernozhukov2005iv. To see this, we revisit the wealth accumulation example of Section (ref). For $p = 2,$ consider a binary treatment environment and assume that $(Y_0, Y_1, D, Z, U_0, U_1)$ satisfies Assumption (ref) with the rank invariance $U = U_0 = U_1$ and $D = \delta (Z, U, \eta),$ where $\eta$ is a random variable independent of all the other variables. Recall that $Y_d = (Y_d^F, Y_d^O)^\prime$ is the vector consisting of net financial assets and other assets, $D$ is the participation status of 401(k), $Z$ is its eligibility, and $U = (U^F, U^O)^\prime$ is the preference for the corresponding assets. The potential outcome for treatment $d$ is componentwisely represented as

align*[align* omitted — 185 chars of source]

Since $U$ and $Z$ are assumed independent, it is expected that the treatment effect of $D$ is identified if $Z$ is sufficiently informative.

Consider an empirical researcher who is interested only in net financial assets $Y_d^F.$ In this case, it is common to apply the standard quantile model to the data $(Y^F, D, Z),$ as chernozhukov2004effects do. To see that this practice leads to the violation of the rank similarity assumption, notice that the structural equation under the one-dimensional model is $Y_d^F = \tilde q_d (\tilde U_d),$ where $\tilde U_d \sim U [0, 1]$ and $\tilde q_d$ is the quantile function of $Y_d^F.$ Then, $\tilde U_d$ is measurable in $U$ since $\tilde U_d = \tilde q_d^{-1} (q_d^{\ast F} (U)).$ The rank similarity requires $\tilde U_0 \overset{d}{=} \tilde U_1 \mid Z, U, \eta,$ but this is satisfied only when $\tilde q_0^{-1} \circ q_0^{\ast F} = \tilde q_1^{-1} \circ q_1^{\ast F},$ which does not hold in general.\footnote{For a counterexample, consider $q_0^\ast (u) = u$ and $q_1^\ast (u) = ((u^1 + u^2) / 2, (u^1 + u^2) / 2)^\prime.$} In other words, the instrument $Z$ is invalid in the sense that it is correlated with the unobserved heterogeneity $\tilde U \coloneqq \tilde U_D$ via the treatment variable $D.$

The failure of rank similarity arises because $q_d^{\ast F}$ depends not only on $U^F$ but also on $U^O,$ whereas chernozhukov2005iv do not allow such dependence---an assumption that is often unrealistic. Specifically, applying the framework of chernozhukov2005iv to the wealth accumulation example would require assuming that individuals determine their holdings of financial assets solely based on their preference for financial assets. In reality, individuals also consider their preference for other assets, since it affects how much they can allocate to financial assets. Our multivariate quantile model accommodates such interdependence by allowing the potential outcomes to depend jointly on multiple unobservables $U = (U^F, U^O)^\prime,$ thereby resolving this limitation.

In summary, the IV multivariate quantile model enables to consider multiple outcome variables at the same time, since it captures the correlation between different outcomes. Moreover, even if only some of the outcome variables are of interest, our model can alleviate the endogeneity by considering other outcomes together.

Identification

In this section, we consider the identification problem of the structural functions. We assume that the treatment and IV are supported on the same finite set, i.e., $\cD = \cZ = \{0, 1, \dots, m - 1\}.$ We will further restrict our attention to binary treatments/IVs in the later part of this section. More general cases are discussed in Appendix (ref).

The representation ((ref)) of Theorem (ref) and the change-of-variables formula imply that the true structural function $q^\ast = (q_0^\ast, \dots, q_{m - 1}^\ast)^\prime$ solves the measure-valued equation

align[align omitted — 185 chars of source]

where

align*[align* omitted — 240 chars of source]

For a scalar function $a : \cU \to \bR,$ a vector-valued function $b : \cU \to \bR^p,$ and a matrix-valued function $M : \cU \to \bR^{p \times p},$ we define three types of supremum norms as follows:

align*[align* omitted — 418 chars of source]

For a compact set $K$ in a Euclidean space, let $C^k (K; \bR^s)$ be the set of $\bR^s$-valued functions on $K$ that are $k$-times continuously differentiable on $\mathrm{Int} (K)$ and the derivatives can be continuously extended to $K.$

Our identification result below shows that $q = q^\ast$ is the locally unique solution to ((ref)) in a certain function class. To state the result formally, consider the normed space $\cQ \coloneqq \left(C^2 (\cU; \bR^p)\right)^m$ equipped with $\left\|q\right\|_\cQ \coloneqq \max_{d \in \cD} \left\|q_d\right\|_\infty.$ Also, for fixed constants $\overline \lambda > \underline \lambda > 0,$ define a subset $\tilde \cQ$ of $\cQ$ as

align*[align* omitted — 365 chars of source]

where $\lambda_{\text{max}} (A)$ is the largest eigenvalue of $A,$ and define $\lambda_{\text{min}} (A)$ similarly. For a constant $K > 0,$ the space in which the parameter is identified is

align*[align* omitted — 257 chars of source]

Since we are interested in identification, we assume the correct specification.

assumption$q^\ast \in \cQ_0.$

Assumption (ref) is weak under Assumption (ref), as it just requires that $q^\ast$ be smooth and that it have moderate derivatives.

The conditions in the definition of $\cQ_0$ restrict the class of admissible deviations from the true structural function. Such restrictions are common in the literature on nonparametric IV identification. To see their role, note that identifying $q$ requires inverting the linearization of the system of equations ((ref)). Although this linearized system is invertible under the standard full-rank condition, its inverse is generally discontinuous because of the infinite-dimensional nature of the problem, rendering the identification ill-posed. This difficulty is resolved by restricting the identification domain to $\cQ_0,$ thereby ruling out pathological deviations from the truth. Similar regularity conditions are imposed in related work, including chen2014local and centorrino2024iterative. In particular, Section 2.3 of chen2014local presents an example where identification fails once an analogous restriction is removed.

We further impose two regularity conditions. Let $\cY_d \subset \bR^p$ be the support of the distribution of $Y_d.$

assumptionFor each $d \in \cD,$ the support $\cY_d$ is a convex compact set with a piecewisely $C^2$ boundary $\partial \cY_d.$
assumptionFor each $(d, z) \in \cD \times \cZ,$ $f_{d, z} \in C^1 (\cY_d; \bR).$

Assumption (ref) requires that the support of a potential outcome be regular. The convexity of the support is necessary for many purposes in the optimal transport theory (see, for example, villani2003topics and figalli2017monge). The smoothness of the boundary is rarely a problem in practice, as it allows for kinks at some points. Combined with Assumption (ref), Assumption (ref) implies that $f_{d, z}$ is bounded.

We consider the identification of structural functions in the following sense.

definitionSuppose $q^\ast \in \cQ_0.$ We say the structural function $q^\ast$ is identified in $\cQ_0$ if the following holds: if a set of random variable $((\tilde Y_d)_{d \in \cD}, \tilde D, \tilde Z, (\tilde U_d)_{d \in \cD})$ satisfies Assumption (ref) with some $q \in \cQ_0 \setminus \{q^\ast\},$ then the joint distribution of $(\tilde Y \coloneqq \tilde Y_{\tilde D}, \tilde D, \tilde Z)$ is different from that of $(Y, D, Z).$ Also, we say the structural function $q^\ast$ is locally identified in $\cQ_0$ if there exists $\varepsilon > 0$ such that $q^\ast$ is identified in $\{q \in \cQ_0 \mid \left\|q - q^\ast\right\|_\cQ < \varepsilon\}.$

In what follows in this section, we focus on the case of $m = 2,$ i.e., $\cD = \cZ = \{0, 1\},$ to develop the idea clearly. Identification results for nonbinary treatments are given in Appendix (ref).

The following is our main theorem.

theoremSuppose that Assumptions (ref), (ref), (ref), (ref) hold. Then, $q^\ast$ is locally identified in $\cQ_0$ if \begin{align} 4 f_{0, 0} (y_0) f_{1, 1} (y_1) > \left(\frac{\overline \lambda}{\underline \lambda}\right)^{p + 1} ( f_{0, 1} (y_0) + f_{1, 0} (y_1) )^2 \end{align} holds for almost all $(y_0, y_1) \in \cY_0 \times \cY_1$ with respect to the Lebesgue measure.

Condition ((ref)) is a novel assumption that, to the best of our knowledge, has not appeared in the existing literature. Roughly speaking, it requires that $Z = 0$ ($Z = 1$) be sufficiently positively correlated with $D = 0$ ($D = 1$). To see this, let $g (y \mid d, z)$ denote the density of $Y$ conditional on $D = d$ and $Z = z,$ and suppose it satisfies $m < g (y \mid d, z) < M$ for some constants $M > m > 0.$ Then, the condition

align*[align* omitted — 221 chars of source]

is sufficient for condition ((ref)) to hold. This inequality is satisfied when $\bP (D = 0 \mid Z = 0)$ and $\bP (D = 1 \mid Z = 1)$ are sufficiently large. Since condition ((ref)) requires a positive correlation between $D$ and $Z,$ the labeling of the instrumental variable can be reversed if necessary to meet the condition.

Condition ((ref)) is stronger than chernozhukov2005iv's monotone likelihood ratio condition: $$ \frac{f_{1, 1} (y_1)}{f_{0, 1} (y_0)} > \frac{f_{1, 0} (y_1)}{f_{0, 0} (y_0)} \quad \text{for all $(y_0, y_1) \in \cY_0 \times \cY_1$} . $$ This can be shown as follows. Since $(f_{0, 1} (y_0) + f_{1, 0} (y_1))^2 \geq 4 f_{0, 1} (y_0) f_{1, 0} (y_1),$ equation ((ref)) implies

align*[align* omitted — 210 chars of source]

When $p = 1,$ condition ((ref)) can be relaxed as follows.

propositionSuppose that Assumptions (ref), (ref), (ref), (ref) hold. Then, $q^\ast$ is locally identified in $\cQ_0$ if \begin{align} \begin{pmatrix} f_{0, 0} (y_0) & f_{0, 1} (y_0) \\ f_{1, 0} (y_1) & f_{1, 1} (y_1) \end{pmatrix} \end{align} is positive definite for almost all $(y_0, y_1) \in \cY_0 \times \cY_1$ with respect to the Lebesgue measure.

Proposition (ref) is consistent with chernozhukov2005iv, who show that identification holds when matrix ((ref)) is full rank, since the determinant of any full-rank $2 \times 2$ matrix can be made positive by switching its columns if necessary.

Identification for $p = 1$ holds under a milder assumption because the proof of Proposition (ref) relies on a property that is valid only in one dimension. Specifically, it uses the fact that the cofactor matrix\footnote{For an invertible matrix $C$, its cofactor matrix is defined as $\hspace{0.1em}\mathrm{cof}\hspace{0.1em} (C) \coloneqq \det (C) C^{-1}$.} of any matrix equals the identity if and only if $p = 1$. See Lemma (ref) for details. This observation implies that replacing the positive definiteness of matrix ((ref)) with condition ((ref)) can be interpreted as the additional cost of allowing for multidimensional potential outcomes.

As the dimension $p$ increases, condition ((ref)) becomes more demanding, which is natural given the need to identify higher-dimensional structural functions. However, the minimum support size of the instrumental variable $Z$ required for identification remains independent of the dimension $p$ of the potential outcomes, as long as the IV is sufficiently correlated with the treatment variable.

Outline of the proof of Theorem (ref)

Roughly speaking, the proof of Theorem (ref) proceeds in two steps. In the first step, we show that the support $\cY_d$ of the potential outcome $Y_d$ is identified under condition ((ref)). If a candidate structural function $q_d$ implies a different support for the potential outcome, then it cannot be consistent with the joint distribution of the observable variables $(Y, D, Z).$ Therefore, the identification analysis can be restricted to the class of structural functions that preserve the support.

In the second step, we show that $q^\ast$ is the locally unique solution to the system ((ref)) of equations within the class of structural functions. The argument follows the logic of the implicit function theorem, which is commonly used in the identification analysis of nonlinear models with finite-dimensional parameters. Specifically, we linearize the nonlinear system ((ref)) around the true structural function $q^\ast.$ The key insight is that the positive correlation condition ((ref)) ensures that the “slope” of the linearized system is bounded away from zero, thereby guaranteeing the local uniqueness of the solution.

The full proof appears in Appendix (ref). The first step is established in Lemma (ref). The second step follows from a more general identification result presented in Appendix (ref).

Concluding Remarks

In this paper, we proposed a new nonlinear IV model that extends chernozhukov2005iv to accommodate the correlation among multiple outcome variables. A key identifying restriction is that the structural functions are assumed to have a symmetric and positive definite derivative with respect to the rank vector. We showed that if the instrumental variable is sufficiently positively correlated with the treatment variable, the structural functions are locally identified. The cost of multidimensionality is that the positive correlation condition becomes more demanding as the dimension of the potential outcomes increases. Nevertheless, the minimum support size of the instrumental variable required for identification remains the same as in the univariate case. These results clarify how identification in nonlinear IV models can be extended to multidimensional settings while maintaining a similar structure to the scalar case.