EconBase
← Back to paper

A condition for the identification of multivariate models with binary instruments -- with Corrigendum and Addendum

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

59,613 characters · 13 sections · 84 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A condition for the identification of multivariate models with binary instruments with Corrigendum and Addendum

abstractThis article introduces an empirical condition for the nonparametric point-identification of multivariate instrumental variable models with continuous endogenous variables using binary instruments. Verifying this condition can confirm point-identification in settings in which traditional approaches are not applicable. In particular, it shows that nonlinear instrumental variable models with general heterogeneity can be point-identified with only a binary instrument. This generalizes existing identification results which either restrict the unobserved heterogeneity substantially or require the instrument to have a large support. The main assumption on the instrumental variable model is cyclic monotonicity of its first stage, a multivariate generalization of the classical rank-invariance assumption for univariate models. Asymptotic convergence results for the empirical observable distributions are derived that allow to check the condition in practice. The identification rests on a fixed-set convergence result of cyclically monotone maps between quasi-concave functions. \\ JEL subject classification: C01, C14\\ Keywords: Cyclic monotonicity; fixed set iteration; instrumental variable; nonseparable model; optimal transportation

Introduction

The first step in a statistical analysis of any economic model is to check whether it is identified, i.e. whether the important features of the model could in theory be recovered from the observations if we had access to the population lewbel2019identification. The arguably most widely used criterion for identification is point-identification, which requires that the parameters of the model can be recovered uniquely from the theoretical population. It has become relatively common, especially in applied economic analyses, to first show nonparametric point-identification of the model of interest before estimating a parametric analogue of the model. This requires a flexible nonparametric point-identification result that can be applied to general economic models with many variables. In many practical settings, for instance in production-function estimation, the endogenous variables or treatments are continuous while researchers only have access to a binary instrument. This article provides a criterion for point-identification of the model in this setting.

We introduce the point-identification result for a general version of instrumental variable models, the “nonseparable triangular model”, which takes the form

equation[equation omitted — 167 chars of source]

The novelty is that the observable variables $Y\in\mathbb{R}^d$ and $X\in\mathbb{R}^k$ can be multivariate with an arbitrary dependence structure while allowing for arbitrarily high but still finite dimensional unobserved heterogeneity $\varepsilon\in\mathbb{R}^l$; in most cases it will be higher dimensional than the outcome of interest $Y$. $X$ is a vector of endogenous regressors which are correlated with the unobserved error term $\varepsilon$. $Z$ is a (binary or discrete) instrumental variable which is required to have an influence on $X$ but itself fully independent of the model.\footnote{The result is straightforward to extend to the discrete setting: there the method has to be applicable for at least two realizations of the instrument.} The independence requirement is captured by the restriction $Z\protect\mathpalette{\protect\independenT}{\perp} (\varepsilon,U)$, i.e. that $Z$ be independent of $\varepsilon$ and $U$ jointly. A standard restriction hoderlein2016erratum is that the unobservable $U$ has to be of the same dimension as the endogenous variable $X$, i.e. $U\in\mathbb{R}^k$. Note that we do not restrict the distribution of $Y$, it can be continuous, discrete, or mixed; we only require that $P_{X|Z}$, the conditional distribution of $X$ given $Z$, is absolutely continuous with respect to Lebesgue measure.

The main contribution of this article is a condition for the observable data-generating process, provided in Assumption (ref) below, that guarantees point-identification of $m$ in the case where $Z$ is binary but $X$ is continuous. This condition is a requirement on the intersection of the multivariate conditional cumulative distribution functions (CDF) $F_{X|Z=z}$ for the two realizations $z,z'$ of the binary $Z$, a multivariate generalization of the intersection requirement in torgovitsky2015identification. It also requires $F_{X|Z=z}$ to be quasi-concave for $z$ and $z'$. This is satisfied by standard distribution functions, for instance all unimodal distributions. Both conditions are made on the observable distribution and hence can serve as a practical check for point-identification of the model (ref).

Assumptions and main result

The goal is to derive a point-identification result for the function $m$. For this, the class of functions $m$ needs to be restricted. Surprisingly, we can allow for a large class of functions $m$, even those where the unobservable $\varepsilon$ is of higher dimension than the observable $Y$. The identification result depends more heavily on the functional form of $h$. As a result, the assumptions on $h$ are significantly stronger than on $m$, in line with the existing literature.

The classical assumption in the literature is a univariate $X$ and $U$ with a function $h$ that is strictly increasing and continuous in $U$ for all $z$ matzkin2003nonparametric, hoderlein2016erratum, imbens2009identification. We extend this assumption to the multivariate setting in a natural way: by assuming that $X$ and $U$ are of the same dimension and $h$ is cyclically monotone. The following definition is taken from villani2003topics and is analogous to the one in shi2018estimating.

definition[Cyclic monotonicity] A map $T:\mathbb{R}^k\to \mathbb{R}^k$ is cyclically monotone if its graph $\Gamma\coloneqq \{(x,Tx): x\in\mathcal{X}\}$ is cyclically monotone in the sense that for all $m\geq 1$, and for all $(x_1,Tx_1),\ldots,(x_m,Tx_m)\in\Gamma$, \[\sum_{i=1}^m\|x_i-Tx_i\|_2^2\leq\sum_{i=1}^m\|x_i-Tx_{i-1}\|_2^2,\] with the convention $x_0=x_m$, where $\|\cdot\|_2$ is the standard Euclidean norm.

Cyclic monotonicity is a natural generalization of univariate monotonicity to vector-valued functions $T:\mathbb{R}^k\to\mathbb{R}^k$, based on a classical result in rockafellar1997convex, which states that the graph of the gradient $\nabla \varphi(x)\coloneqq\left(\frac{\partial}{\partial x_j} \varphi(x)\right)_{j=1,\ldots,k}$ of some convex function $\varphi:\mathbb{R}^k\to\mathbb{R}$ is a cyclically monotone set. This is a natural generalization of univariate monotonicity, because the antiderivative $F(x)\coloneqq \int_a^x f(u)du$ of a monotonically increasing function $f:\mathbb{R}\to\mathbb{R}$ is a convex function rockafellar1997convex. In other words, a monotonically increasing function is the derivative of a convex function in the univariate setting.\footnote{Note that cyclic monotonicity monotonicty is stronger than simple multivariate monotonicity, where $T:\mathbb{R}^k\to\mathbb{R}^k$ is monotone if for every $x_0, x_1\in\mathbb{R}^k$ $(x_1-x_0)'(Tx_1-Tx_0)\geq0$, where $A'$ denotes the transpose of a matrix $A$. In fact, this condition of monotonicity corresponds to cycles of length $2$ in the definition of cyclic monotonicity. This implies immediately that every cyclically monotone function is monotone in this sense. In the univariate setting cyclic monotonicity and monotonicity coincide rockafellar1997convex.}

Based on this definition we can state the the assumptions on the model (ref) and the data-generating process.

assumption[Instrument] $Z$ is binary, i.e. $\mathcal{Z}=\{z,z'\}$. It is a valid instrument for $X$, i.e. (i) it generates exogenous variation in $X$ such that $F_{X|Z=z}(x)\neq F_{X|Z=z'}(x)$ for at least one $x\in\mathcal{X}$ and (ii) is independent of $\varepsilon$ and $U$.
assumption[First stage] \begin{enumerate} • The distribution of the unobservable $U$ is absolutely continuous with respect to Lebesgue measure. • $h(z,U)$ is cyclically monotone between $U$ and $X$ for $z$. • $h(z',u)=u$ for all $u\in\mathcal{U}$ for $z'$. \end{enumerate}
assumption[Second stage] \begin{enumerate} • $m(x,\cdot)$ is identified for exogenous $X$, i.e. for fixed $P_{Y|X=x}$ and $P_\varepsilon$ there exists a unique measure preserving map $m(x,\cdot)$ transporting $P_\varepsilon$ to $P_{Y|X=x}$ for all $x\in\mathcal{X}$. • $m(\cdot,\varepsilon)$ is uniformly continuous in probability in $X$. • The law $P_{\varepsilon|X=x}$ is absolutely continuous with respect to Lebesgue measure for all $x$. • The support $\mathcal{E}$ is convex and independent of $X$ and $Z$, i.e. $\mathcal{E}_{x,z}$ coincides with $\mathcal{E}$ for all $(x,z)\in\mathcal{X}\times\mathcal{Z}$. • $m(\bar{x},e)=e$ for all $e\in\mathcal{E}$ and one known $\bar{x}\in\mathcal{X}$. \end{enumerate}

The following is the formal definition of uniform continuity in probability.

definitionA sequence $\{m_n(\cdot)\}_{n\in\mathbb{N}}$ is said to converge uniformly in probability over a set $\mathcal{E}$ to a function $m(\cdot)$ if newey1991uniform \[ \sup_{e\in\mathcal{E}} \|m_n(e)-m(e)\|=o_P(1),\] where $\|\cdot\|$ is some norm. If for every sequence $\{x_n\}_{n\in\mathbb{N}}\in\mathcal{X}$ which converges to some $x\in\mathcal{X}$ the corresponding sequence $m(x_n,e)$ converges uniformly in probability to $m(x,e)$, we say that $m$ is uniformly continuous in probability in $x$.

Uniform continuity in probability is significantly weaker than assuming uniform continuity of $m$ in $x$. It is important, because we cannot guarantee full continuity of $m(x,e)$ in $x$ in many applications, but continuity in probability is satisfied. For instance, if $m(x,\varepsilon)$ is defined via optimal transport theory, one can prove (uniform) continuity in probability straightforwardly (see for instance Corollary 5.23 in villani2008optimal villani2008optimal). In terms of our identification result, the difference between continuity in probability and full continuity of $m$ in $x$ is that we will only be able to identify $m$ for almost every $e\in\mathcal{E}$ in the case of continuity in probability instead of every $e\in\mathcal{E}$. But this difference is immaterial in practice.

These assumptions are generalizations of classical univariate assumptions in the literature matzkin2003nonparametric, hoderlein2016erratum. The following is the main condition on the observable data-generating process announced in the introduction.

assumption[The condition on the data-generating process] The distributions $F$ and $G$ are continuously differentiable everywhere and quasi-concave with convex coinciding supports $\mathcal{X}_F\equiv\mathcal{X}_G\equiv\mathcal{X}\subseteq\mathbb{R}^k$. Moreover, one of the following two settings holds: \begin{itemize} • $\mathcal{X}$ is bounded above or bounded below. • $\mathcal{X}$ is allowed to be unbounded, but $F$ and $G$ intersect in the sense that there exist some quantile-values $\alpha,\beta\in(0,1)$ such that $G(x)>F(x)$ for all $x\in\mathcal{X}$ with $\alpha\leq G(x)<1$ and $G(x')<F(x')$ for all $x'\in\mathcal{X}$ with $0< G(x')\leq\beta$. \end{itemize}

The following provides a formal definition of quasi-concave functions.

definitionA function $F:\mathbb{R}^k\to\mathbb{R}$ is quasi-concave if its support $\mathcal{X}$ is convex and if for every $x,x'\in\mathcal{X}$ \[F((1-t)x+tx')\geq\min\{F(x),F(x')\},\qquad t\in[0,1].\] We call $F$ strictly quasi-concave if \[F((1-t)x+tx')>\min\{F(x),F(x')\},\qquad t\in[0,1].\]

Quasi-concave functions have a long tradition in economic theory as they induce convex isoquants mas1995microeconomic. We require quasi-concavity of the cumulative distribution functions $F$ and $G$. Many important cumulative distribution functions are (strictly) quasi-concave like the (multivariate) Normal-, uniform-, or beta-distribution---in particular, (strictly) log-concave distribution functions are (strictly) quasi-concave. For a reference on quasi-concave distributions, we refer to prekopa2013stochastic.

These assumptions lead to the following identification result.

theoremLet Assumptions (ref)-- (ref) hold for model (ref). The identified set for $m$ is \[\mathbb{I}\coloneqq \{\text{$m$ satisfies Assumption \ref{mpiso2}}:(\varepsilon_m,U)\protect\mathpalette{\protect\independenT}{\perp} Z\thickspace\text{for all $\varepsilon_m\in m^{-1}(Y,X)$}\}.\] If Assumption (ref) holds for the observed data-generating process, then $\mathbb{I}$ contains an $(X,\varepsilon)$-almost everywhere unique element $m$.

Theorem (ref) is an identification result in the sense that it does not provide us with a way to estimate the function $m$. It shows that the identified set $\mathbb{I}$ contains a unique element, just as the univariate result torgovitsky2015identification and d2015identification. Also, the definition of the identified set is slightly more general than the definition in torgovitsky2015identification to account for the fact that $m$ may not be invertible in $\varepsilon$, for instance when $\varepsilon$ is of higher dimension than $Y$. If $m$ is invertible in $\varepsilon$ then \[\mathbb{I}\coloneqq \{\text{$m$ satisfies Assumption \ref{mpiso2}}:(m^{-1}(X,Y),U)\protect\mathpalette{\protect\independenT}{\perp} Z\},\] which coincides with the identified set in torgovitsky2015identification.

Discussion of the assumptions and applicability

Assumption (ref) is the standard assumption for a theoretical instrument. It does not prescribe the strength of the instrument. $Z$ can be lower-dimensional than $X$. Allowing for a lower-dimensional $Z$ is relevant in many applications. For instance one could view $z$ and $z'$ as different markets of a product.

Assumption (ref) is also similar to classical assumptions in the identification literature, except for the requirement of cyclic monotonicity and the normalization. The classical identification results in the univariate case starting with matzkin2003nonparametric assume univariate $X$ and $U$ and require $h(x,\cdot)$ to be strictly increasing and continuous in $U$. The assumption of cyclic monotonicity implies that $U$ is of the same dimension as $X$, which has been shown to be a necessary assumption for point-identification in nonseparable triangular models hoderlein2016erratum. Part 3 is a normalization analogous to matzkin2003nonparametric. It fixes $h$ to be the identity for one $\bar{z}\in\{z,z'\}$. This assumption is stronger than in the univariate case, which rests on the control function approach and constructs a control function $V$ that “mimics” the behavior of $U$. Such a control-function approach is unavailable in the multivariate setting since the full ordering of the real line is lost in higher dimensions.

Assumptions (ref) (2) - (5) are regularity assumptions on the model in analogy to the univariate identification result in torgovitsky2015identification. Assumption (ref) (1) is not tautological. It simply requires that without the endogeneity problem, i.e. if $X$ were exogenous, $m$ could be point-identified via uniqueness. There are many assumptions in the literature to guarantee this (in particular all possible variations in the literature of Lemma 1 in matzkin2003nonparametric) some of which we show below. In particular, $m(x,\cdot)$ does not have to be invertible in $\varepsilon$ for identification, which seems to provide a novel result for these general models. In this respect, our identification result can be seen as a “dissection” of the point-identification argument into identification of $m$ in the exogenous setting and the endogenous case: if $m$ is identifiable in the exogenous setting then our approach shows that it can be identifiable in the endogenous setting too. Below we provide some sufficient conditions when $Y$ is continuous, but other conditions can be derived for binary or discrete settings.

When Theorem (ref) is applicable

Model (ref) is a generalization of the classical univariate one torgovitsky2015identification, d2015identification: in the case where $X$ and $U$ are univariate, strict cyclic monotonicity of $h$ reduces to a strictly increasing and continuous function in $U$. Moreover, requiring $m$ to be strictly increasing and continuous in $\varepsilon$ makes it satisfy Assumption (ref) (1) under a normalization matzkin2003nonparametric. Therefore, the proposed result is applicable in any setting where the first or second stage is univariate under the standard monotonicity assumptions, but it is applicable in more general settings that are economically relevant.

\paragraph{Economic conditions for $m(X,\varepsilon)$.} Assumption (ref) (1) only requires uniqueness (potentially under a normalization) of $m$ if $X$ were exogenous. The classical assumption of a strictly increasing and continuous function is a nonparametric assumption that follows from the theory of optimal transportation: given probability measures $P_\varepsilon$ and $P_Y$ on domains $\mathcal{E}$ and $\mathcal{Y}$ and a surplus function $s: \mathcal{E}\times\mathcal{Y}\to \mathbb{R}$, the Monge problem consists of finding a deterministic map $m(\varepsilon)$ that maximizes \[\int_{\mathcal{E}} s(\varepsilon,m(\varepsilon)) \medspace dP_\varepsilon(\varepsilon)\] such that it maps $P_\varepsilon$ to $P_Y$. This problem has a solution if $P_\varepsilon$ possesses a density with respect to Lebesgue measure, which we assume. Then it provides an optimal deterministic matching $m(\varepsilon)$ under maximal average surplus as measured by $s(\varepsilon, y)$ and has found applications in a variety of economic settings galichon2021unreasonable. Different choices of surplus functions $s(\varepsilon,y)$ produce matchings with different properties: for instance, if $s(\varepsilon,y)$ is the negative Euclidean distance $-|\varepsilon - y|^2$ then the optimal match $m(\varepsilon)$ is cyclically monotone. In the univariate setting, it produces the strictly increasing and continuous function $m(\varepsilon)$ from the classical setting. This idea has been used recently in chernozhukov2016monge. These are just two examples. More general nonparametric functional forms are possible by choosing different surplus functions. The following are two examples.\\ (i) The natural generalization of the univariate setting is to assume $Y$ and $\varepsilon$ are multivariate with absolutely continuous distributions and $\mathcal{E}$ and $\mathcal{Y}$ are of the same dimension. Then under a regularity assumption on the surplus function $m(x,\cdot)$ will be the unique generalized cyclically monotone map solving the Monge problem villani2003topics, chernozhukov2014single. This uniqueness makes it identifiable under a normalization like in Assumption (ref) (5), see the analysis in chernozhukov2014single. \\ (ii) This idea has been extended to the setting where the unobservable $\varepsilon$ is higher-dimensional than the observable outcome $Y$ in chiappori2015multi, chiappori2019multi and in particular mccann2018optimal. Under regularity assumptions on the surplus function $s(\varepsilon,y)$, the optimal map $m(x,\varepsilon)$ solving the Monge problem is the unique transport map between $P_\varepsilon$ and $P_{Y|X=x}$ mccann2018optimal. This uniqueness makes $m(x,\varepsilon)$ identifiable under the normalization in Assumption (ref) (5) as required for Assumption (ref) (1).

\paragraph{Economic conditions for $h(Z,U)$.} One classical argument for nonseparability of the model and monotonicity of $h$ in $U$ is maximization of quasi-linear functions. Consider for instance the example in imbens2009identification from the univariate setting. $Y$ is some outcome like life-time earnings or firm revenue and $m(X,\varepsilon)$ is some production function. An agent chooses input $X$ to maximize the expected outcome minus the costs associated to the value of $X$, given her information set. Suppose the information set consists of a (potentially lower dimensional) noisy signal $U$ of the unobserved input $\varepsilon$.\footnote{imbens2009identification assume $U$ and $\varepsilon$ are both one dimensional.} Then $X$ is obtained as the optimization problem \[X = \operatorname*{\arg\!\max}_x \left[E\left[m(x,\varepsilon) | U,Z\right] - c(x,Z)\right],\] which is a nonseparable first stage $X = h(Z,U)$. If the conditional expectation is linear in both $x$ and $U$ for fixed $z$, then the optimal $X$ will be a cyclically monotone function of $U$ for fixed $z$ by a classical theorem in rochet1987necessary. The same argument holds in the univariate setting: if the expectation is linear in both $X$ and $U$ then $h$ will be strictly increasing and $P_U$-almost everywhere continuous in $U$.

This is not the only instance of cyclic monotonicity in economics. A recent example is from multinomial choice models shi2018estimating, where the authors show that the “social surplus function” mcfadden1981econometric, i.e. the expected utility from a simple additive multinomial choice problem, is cyclically convex. Let $U\coloneqq (U_1,\ldots U_M)$ be the latent utility of an individual choosing between $M$ options, trying to maximize their utility, and let $\varepsilon\coloneqq (\varepsilon_1,\ldots,\varepsilon_M)$ be some idiosyncratic error term. Then the gradient of \[E\left[\max_{1\leq m\leq M} U_m+\varepsilon_m \vert U= u\right]\] is cyclically monotone shi2018estimating for any realization $u$.

As another example, cyclic monotonicity is naturally linked to demand analysis, in particular the Generalized Axiom of Revealed Preferences (GARP) as introduced in varian1982nonparametric and cyclic consistency introduced in afriat1967construction, see hadjisavvas2006handbook for an in-depth analysis. One of the earliest approaches in econometrics exploiting cyclic monotonicity between consumption and prices is browning1989nonparametric who tests rational expectation hypotheses. More recently, cyclic monotonicity has found uses in the semi-parametric estimation of multinomial choice models shi2018estimating, the identification of single-market Hedonic models chernozhukov2014single, and the construction of nonlinear analogues of the classical principal component analysis gunsilius2019independent.

When Theorem (ref) is not applicable

The two fundamental assumptions for Theorem (ref) are (i) absolute continuity of $F_{X|Z}$ and (ii) cyclic monotonicity of $h(x,\cdot)$.

If the conditional distribution $F_{X|Z}$ is not absolutely continuous, then Theorem (ref) is not applicable. This follows from the multivariate fixed-set method we show below, which requires an intersection of two continuous CDFs. Recently, feng2019matching introduced a complementary approach based on a similar idea that is applicable in settings where $F_{X|Z}$ is discrete. A general method that can handle mixed data most likely does not exist: the reason is that the interplay between the distributions $F_{X|Z}$ as the first stage of the model $h$ is crucial for a fixed-set convergence result we show below.

The cyclic monotonicity of $h$ is a natural generalization of the strict monotonicity of $h$ in the univariate case. It is still a restrictive functional assumption. To see this, recall the expected payoff maximization example \[X = \operatorname*{\arg\!\max}_x \left[E\left[m(x,\varepsilon) | U,Z\right] - c(x,Z)\right]\] which is cyclically monotone if the conditional expectation is linear in $X$ and $U$ for the two realizations of $Z$. This is a strong functional form restriction. Allowing for more general and nonlinear functional forms of $E\left[m(x,\varepsilon) | U,Z\right]$ leads to the concept of generalized cyclic monotonicity villani2003topics, which has been used in chernozhukov2014single. Unfortunately, the proposed fixed-set method on which our identification result relies would need to be adjusted each time for any new form of “generalized cyclic monotonicity”. So far, there does not seem to be a general approach for a large class of generalized cyclically monotone functions. Our result relies on classical cyclic monotonicity because we show that cyclically monotone maps between quasi-concave distribution functions take the form of metric projections onto convex isoquants. The corresponding iterative procedure can be analyzed and this is the main technical contribution of the article which we now explore.

The underlying idea: multivariate fixed-set iterations

The idea in the univariate case

The intuition for identification of the univariate version of model (ref) follows the idea from torgovitsky2015identification. The univariate model is the one where all variables, observable and unobservable, are univariate and where both $m$ and $h$ satisfy the rank-invariance condition in their second argument, i.e. they are strictly increasing and continuous functions in $\varepsilon$ and $U$, respectively.

$m(X,\varepsilon)$ in model (ref) maps the distribution $F_\varepsilon$ of the unobservable to the conditional counterfactual distribution $F_{Y(X)}$ for exogenous $X$, i.e. for an $X$ with $X\protect\mathpalette{\protect\independenT}{\perp} \varepsilon$. The notation $F_{Y(X)}$ is Rubin's counterfactual notation rubin1974estimating: due to the endogeneity problem, i.e. the fact that $X$ and $\varepsilon$ are not independent, $F_{Y(X)}$ is unobservable and the observable distribution $F_{Y|X}$ does not coincide with $F_{Y(X)}$. In order to identify $m$, we would need to know $F_{Y(X)}$, i.e. we want to know how the model would behave for a change in $X$ that does not affect the distribution of $\varepsilon$.

The idea is to vary $Z$ in such a way that $X$ varies but $\varepsilon$ stays constant. The fact that this is possible if $Z$ is only binary was first shown in torgovitsky2015identification and d2015identification. The argument rests on the fact that using the binary instrument $Z$ with realizations $z$ and $z'$, there are two maps that do not change the distribution $F_\varepsilon$ of $\varepsilon$, but change values of $X$. This hence captures the exogenous effect of $X$ on $Y$. These two maps are depicted in Figure (ref).

figure[figure omitted — 1,324 chars of source]

The first map changes $z\mapsto z'$ for a fixed $x$, i.e. switches the distributions $F_{X|Z=z}(x)$ and $F_{X|Z=z'}(x)$ (the “vertical” map in Figure (ref)). The fact that $F_{\varepsilon|X,Z}$ is not affected by this follows from the classical control variable approach imbens2009identification and because $Z$ has no effect on $m$ due to the exclusion restriction of $Z$. A control variable $V$ is such that $X \protect\mathpalette{\protect\independenT}{\perp} \varepsilon| V$ and can be constructed via $V = F_{X|Z}(X)$ imbens2009identification. In particular, it contains the same information as the unobservable $U$. Hence, by conditioning on $V$ and the fact that $Z$ is independent of $(\varepsilon, U)$, the vertical shift changes $Z$ but does not change $X$, which therefore does not affect the function $m(x,\varepsilon)$ we want to identify torgovitsky2015identification.

The second map is the change of quantiles (the “horizontal” map in Figure (ref)), which follows by the fact that the change $(x,z')\to (Tx,z)$ is performed in such a way that $F_{X|Z=z'}(x) = F_{X|Z=z}(Tx)$. This is again achieved via the control variable approach by defining $V=F_{X|Z}(X)$ and conditioning on this. This implies that the horizontal map does not affect the distribution of $\varepsilon$ and hence the function $m(\cdot,\varepsilon)$ we want to identify. If we keep alternating between these two maps for a given starting value $x_0$, this sequence $x_0,Tx_0,T(Tx_0),T(T(Tx_0)),\ldots$ will converge to the point $x^*$ where $F_{X|Z=z}$ and $F_{X|Z=z'}$ intersect. For points $x\leq x^*$ we need to iterate $z\mapsto z'$ and for points $z\geq z^*$ we need to iterate $z'\mapsto z$. This allows us to identify the function $m(x,\varepsilon)$ by comparing different points $x$ in this iterative approach to the point $x^*$, because the “vertical” and “horizontal” map do not change the distribution of $\varepsilon$ and hence keep $m(\cdot,\varepsilon)$ fixed.

Measure preserving isomorphisms and multivariate fixed-set iterations

Extending this idea to the multivariate setting, i.e. to the setting where $F_{Y|X}$, and in particular $F_{X|Z=z}$ and $F_{X|Z=z'}$ are multivariate, is the main contribution of this article. This multivariate setting allows for arbitrary dependence between the variables, extending the element-wise generalization of torgovitsky2015identification and d2015identification, which only works with the marginal distributions of each individual element $X_j$ of the vector $X\coloneqq (X_1, X_2,\ldots, X_k)\in\mathbb{R}^k$. There are mainly two reasons for why the univariate reasoning does not work in a higher dimensional setting. First, the map $(x,z)\mapsto (F_{X|Z=z}(x),z)$ used for generating the control variable is only invertible in the one-dimensional case, and we need to define an analogue in our multivariate setting. Second, a general sequencing argument is more intricate in higher dimensions.

Generalizing the map $(x,z)\mapsto (F_{X|Z=z}(x),z)$.

We now provide a formal argument for how we solve the first challenge; a rigorous argument using disintegrations is given in the appendix. We need to find a natural generalization of the “horizontal map” in Torgovitsky's argument; we show below that this is achieved by any map $h(x,\cdot)$ that is a measure preserving isomorphism einsiedler2013ergodic. A map $T:A\to B$ transporting a probability measure $P_A$ onto another probability measure $P_B$ is measure-preserving if it is measurable\footnote{Measurability of $T$ means that $\mathscr{A}=T^{-1}\mathscr{B}$, where $\mathscr{A}$ and $\mathscr{B}$ are the $\sigma$-algebras corresponding to $A$ and $B$, respectively. $T^{-1}A$ denotes the set of points $x\in \mathcal{X}$ such that $Tx\in A$.} and

equation[equation omitted — 35 chars of source]

for every set $S$ in the $\sigma$-algebra $\mathscr{B}$ corresponding to $Y$. If $T$ is invertible and its inverse is also measure-preserving, it is called a measure-preserving isomorphism. In short, a measure preserving isomorphism is a map that preserves probabilities, is invertible, and whose inverse also preserves probabilities.

We now argue that this is the required assumption on $h$ to solve the first challenge. In fact, since $h$ is is a measure preserving isomorphism between $P_U$ and $P_{Y|X=x}$ for all $x$, its inverse $U=h^{-1}(X,Z)$ is measure preserving too, which gives

equation[equation omitted — 131 chars of source]

The second equality in (ref) follows from the fact that $Z$ is an instrument and that $\varepsilon$ is independent of $Z$ conditional on $U$. The first equality holds by the following reasoning: the map $\phi: (X,Z) \mapsto (h^{-1}(X,Z),Z)$ is an invertible map that preserves probabilities since $z\mapsto z$ is the identity and $x\mapsto h^{-1}(x,z)$ is an optimal transport map preserving probabilities for all $z$, so that for every rectangle $A_x\times A_z\equiv(-\infty,x]\times (-\infty,z]\in\mathscr{B}_{\mathbb{R}^{k+m}}$ \[P_{X,Z}(A_x\times A_z) = P_{U,Z}(\phi^{-1}(A_x\times A_z))\equiv P_{U,Z}(h^{-1}(A_x,z)\times A_z).\] The same thing holds for the map $(\varepsilon,x,z)\mapsto (\varepsilon,h^{-1}(x,z),z)$, so that for every rectangle $A_\varepsilon\times A_x\times A_z\equiv (-\infty,\varepsilon]\times (-\infty,x]\times (-\infty,z] \in\mathscr{B}_{\mathbb{R}^{d+k+m}}$ \[P_{\varepsilon,X,Z}(A_\varepsilon\times A_x\times A_z) = P_{\varepsilon,U,Z}(\phi^{-1}(A_\varepsilon\times A_x\times A_z))=P_{\varepsilon,U,Z}(A_\varepsilon\times h^{-1}(A_x,z)\times A_z).\] Thus \[P_{\varepsilon|X,Z}(A_\varepsilon)=\frac{P_{\varepsilon,X,Z}(A_\varepsilon\times A_x\times A_z)}{P_{X,Z}(A_x\times A_z)} =\frac{P_{\varepsilon,U,Z}(A_\varepsilon\times h^{-1}(A_x,z)\times A_z)}{P_{U,Z}(h^{-1}(A_x,z)\times A_z)}=P_{\varepsilon|U,Z}(A_\varepsilon).\] The last thing to notice is that conditioning on measure zero events does not cause issues, because $(X,Z) \mapsto (h^{-1}(X,Z),Z)$ is measurable with measurable inverse by definition of a measure preserving isomorphism $h(z,u)$, so that their $\sigma$-algebras coincide, i.e. $\sigma(U,Z) = \sigma(X,Z)$.\footnote{Note that this conditioning is different from the approach in kasy2014instrumental. Kasy used the mapping $\psi:(X,Z)\mapsto (X,h^{-1}(X,U))$, where he defined the inverse of $h$ is with respect to $X$. This makes the composite map not invertible, as it becomes a map from $\mathbb{R}^2$ to $\mathbb{R}\times \mathbb{U}$, where $\mathbb{U}$ is the (in Kasy's case possibly infinite dimensional) metric space containing $U$. Therefore, the respective $\sigma$-algebras $\sigma(X,Z)$ and $\sigma(X,h^{-1}(X,U))$ need not coincide. In our case, however, we use the measure preserving isomorphism $\phi(X,Z)= (h^{-1}(X,Z),Z)$, which is measurable with measurable inverse, so that $\sigma(X,Z)$ and $\sigma(h^{-1}(X,Z),Z)$ coincide. It is exactly here where we make use of the fact that $U$ and $X$ must be of the same dimension.}

This argument shows that a measure preserving isomorphism is the required generalization of the “horizontal map”. Note that we do not need to make any assumptions on the function of the second stage. In particular, we do not need to make any assumptions on the dimension of $\varepsilon$ or any functional form assumptions on $m$. Also note that we use the random variable $U$ instead of constructing a control variable $V$ as in the univariate setting. The reason is that a control variable can not be generated in multivariate settings because of a lack of a full ordering. Therefore, we invoke Assumption (ref) (2) which allows us to condition on $U$ instead.

The proof of Theorem (ref) shows that the “vertical map” from the univariate setting still holds in the multivariate setting under the exclusion restriction on $Z$ and Assumption (ref) (2). We also condition on $U$ instead of constructing a control variable $V$. But conditional on $U$, $Z$ is independent of $\varepsilon$, so that changing $z\mapsto z'$ for fixed $x$ does not affect the distribution of $\varepsilon$ and hence the function $m(\cdot,\varepsilon)$ which is what we need for identification.

Multivariate fixed-set iterations.

The second, and more complicated, challenge to address, is to find a generalization of the classical fixed-point iterations result from the univariate setting. We do this by assuming that $h$ is cyclically monotone in $U$, which in our setting with absolutely continuous distributions will be a measure preserving isomorphism. The idea is to generalize the fixed-point iteration argument from Figure (ref) which has been used in many settings beyond the immediate application in econometric identification, in particular in microeconomics for establishing the existence of equilibria (e.g. \citeauthor*{hopenhayn1992entry} hopenhayn1992entry, \citeauthor*{mas1995microeconomic} mas1995microeconomic). This result could be of independent interest, so we phrase it in a general way.

We consider functions $F(x)$ and $G(x)$ on some set $\mathcal{X}\subset\mathbb{R}^k$, $k\geq2$, which in our case will be cumulative distribution functions on a support $\mathcal{X}$. The main challenge for a multivariate analogue of the fixed-point sequence is to define an appropriate version of the quantile-preserving map $T:\mathbb{R}\to\mathbb{R}$ from Figure (ref). The reason is that the quantiles for multivariate cumulative distribution functions are not points but isoquants, also called level sets.

definition[Isoquant] For every point $x_0\in\mathcal{X}$, $I_F(x_0)$ denotes the isoquant or level set of $F$ at $x_0$, which is the set of all $x\in\mathcal{X}$ which have the same value $F(x)$ as $x_0$: \[I_F(x_0)\coloneqq\{x\in\mathcal{X}:F(x)=F(x_0)\}.\] The epigraph of the isoquant $I_F(x_0)$ is \[I^{\uparrow}_F(x_0)\coloneqq\{x\in\mathcal{X}:F(x)\geq F(x_0)\},\] i.e. all values at or above the isoquant.

Our idea is not to find a unique fixed point, but a lower-dimensional fixed set compared to the ambient set $\mathbb{R}^k$. By the fact that this set will be lower-dimensional than the dimension of $\mathcal{X}$, it will have zero probability under absolutely continuous distributions, similar to the fixed point in the univariate case. The natural analogue of the quantile-preserving map $T$ in the univariate case is a cyclically monotone map $T:\mathbb{R}^k\to\mathbb{R}^k$. If the distributions $P_U$ and $P_{X|Z=z}$ are absolutely continuous with respect to Lebesgue measure, then by Brenier's theorem (brenier1991polar brenier1991polar or villani2003topics villani2003topics, Theorem 2.12) $h$ is invertible and in particular a measure-preserving isomorphism. Recall that we require $h$ to me a measure preserving isomorphism between $P_U$ and $P_{X|Z=z}$ for all $z$ for our argument.

In addition to cyclic monotonicity, we also need to restrict the possible functional forms of $F$ and $G$ in order to be able to derive the dynamics of cyclically monotone maps between them. We do this by assuming they are quasi-concave. We require strict quasi-concavity due to the following characterization of cyclically monotone maps between quasi-concave distribution functions with identical support.

lemmaLet $T$ be the cyclically monotone map between quasi-concave distribution functions $G$ and $F$ supported on $\mathcal{X}$.\footnote{Note that we say “the” cyclically monotone map, as it is unique between absolutely continuous distribution by Brenier's theorem villani2003topics.} Then, for each $x\in\mathcal{X}$, $T$ is either the metric projection of $x$ onto $I^{\uparrow}_{F}(x)$ or its inverse, which is the projection onto $I^{\uparrow}_{G}(x)$.

In short, this lemma characterizes cyclically monotone maps between quasi-concave distribution functions with the same support as the metric projection from isoquants of one function onto isoquants of the same value of the other function. The metric projection $T$ of $x$ onto the closed convex set $I_F(x^*)$ maps $x$ onto the point $y\in I^\uparrow_F(x^*)$ which is closest to $x$ in the sense that \[y=\operatorname*{\arg\!\min}_{s\in I^\uparrow_F(x^*)}\|x-s\|_2^2.\] Quasi-concavity of $F$ and $G$ guarantee that this map exists and is unique since $I^\uparrow_F(x^*)$ is a closed and convex subset of $\mathcal{X}$ in this case aliprantis2006infinite. Figure (ref) depicts the idea of Lemma (ref) in the case where the isoquants between $G$ and $F$ intersect at some $x^*$.

figure[figure omitted — 1,479 chars of source]

In the first example the point $x\in I_G(x_*)$ is outside the convex set $I^{\uparrow}_F(x_*)$, which means that $G(x)>F(x)$ there. Therefore, the cyclically monotone map in this case is the metric projection of $x$ onto $I^{\uparrow}_F(x_*)$, denoted by $Tx$. In the second case, the point $x'\in I_G(x_*)$ lies inside $I^{\uparrow}_F(x_*)$, because $G(x')<F(x')$. In this case, notice that if we exchange the roles of $G$ and $F$, it holds that there is a point (which we conveniently label $Tx'$), for which the cyclically monotone map $T^{-1}$ from $F$ onto $G$ is the metric projection of $Tx'\in I_F(x_*)$ onto $I^{\uparrow}_G(x_*)$.

A cyclically monotone map between the distribution functions “preserves the quantiles” as in the univariate case. Furthermore, it does so in a very simple manner, via projections. This allows us to derive the dynamics, analogously to the univariate setting. All we need for this is that the distribution functions $F$ and $G$ intersect appropriately. This is where the intersection condition from Assumption (ref) comes in. It implies the existence of a fixed set to which the dynamics for every point in $\mathcal{X}$ converge. This is analogous to the univariate case: if there exist quantiles $\alpha$ and $\beta$ as above then $F(x)$ and $G(x)$ must intersect at least once between these two quantiles.

lemma[Fixed-set dynamics] Under Assumption (ref) there exists a manifold $\mathcal{I}(F,G)\coloneqq\{x\in\mathcal{X}: F(x)=G(x)\}$ which generically is of lower dimension than $k$ and is hence of (Lebesgue-) measure zero. Furthermore, for every $x\in\mathcal{X}$, the repeated application of the cyclically monotone map $T$ between $F$ and $G$ or its inverse $T^{-1}$ will converge to a point $x^*\in\mathcal{I}(F,G)$.

The manifold $\mathcal{I}(F,G)$ in Lemma (ref) is the fixed-set we require for our subsequent identification result. Note that we require strict quasi-concavity to guarantee that $\mathcal{I}(F,G)$ is generically lower-dimensional and hence of measure zero. We need this lower-dimensional manifold for our identification result to hold for almost every $x$. This is also analogous to the univariate case, as a point in the univariate case is a lower-dimensional subset.

The genericity condition in the statement follows from the generality of Assumption (ref). A property is generic if the set of all possible elements which satisfy this property are of second category in the sense of Baire aliprantis2006infinite. Intuitively, genericity is the topological analogue of an “almost sure” property: instead of introducing a probability structure on a space one works with the topological structure. In this sense, one can view sets of second category as the topological analogues of sets of probability one. In our case we need the genericity statement because the distribution functions in Assumption (ref) (ii) will intersect generically in a lower-dimensional manifold.

Compare this to the univariate setting from torgovitsky2015identification. There, the author directly requires that the two distribution functions intersect in only one point, and not in a connected set. The reasoning is exactly the same as in our case: in the univariate case a point is a submanifold of measure zero, and one can only identify the function $m$ up to this lower-dimensional manifold. Instead of making the stronger assumption that the two distribution functions intersect in a certain manifold, we make the much easier to check assumption on the existence of the quantiles $\alpha$ and $\beta$ and show that this generically leads to the required intersection. One could do the exact same thing in the univariate case in torgovitsky2015identification. In fact, it follows from the same reasoning that the set of intersections of two univariate strictly increasing and continuous distribution functions will generically be lower-dimensional, i.e. a collection of disconnected points.

To make Assumption (ref) and Lemma (ref) more tangible, consider Figure (ref) which depicts the contour maps of bivariate t- and Normal distributions as well as their intersections.

figure[figure omitted — 864 chars of source]

The intersection of the two distribution functions, the dark black line, is one connected manifold in the left panel. It is not difficult to construct examples where $\mathcal{I}(F,G)$ consists of several manifolds, as depicted in the right panel of Figure (ref). The existence of several manifolds is a generalization of the univariate case where there can exist more points $x^*$ of equilibrium, i.e. several intersections between the distribution functions. We only need to require that at least one of these sub-manifolds satisfies property $(ii)$ in the case where the support is unbounded. This is provided by the “upper” manifold in the right panel of Figure (ref). As long as all of these manifolds are lower-dimensional than the support (which they clearly are in this example as they are univariate curves), we will be able to identify the model up to the union of these manifolds, which will be of measure zero.

Comparison to existing approaches

We now provide a general scenario where our proposed multivariate criterion based on Assumption (ref) is applicable, but the multivariate generalizations of the existing approaches in torgovitsky2015identification and d2015identification are not. By “multivariate generalizations of the existing approaches”, we mean the straightforward generalization of the univariate approaches to the setting of a multivariate $X$ by applying the requirement of intersection of the univariate CDF $F_{X|Z=z}$ and $F_{X|Z=z'}$ element-wise. This has been done in the supplement torgovitsky2015supplement for instance. For a vector-valued random variable $X\in\mathbb{R}^k$, it requires that each pair of marginal distributions $F_{X_i|Z=z}$ and $F_{X_i|Z=z'}$, $i=1,\ldots, k$, intersect in at most finitely many points, but it does not restrict the joint distributions $F_{X|Z=z}$ and $F_{X|Z=z'}$ over all elements $X_i$ as we do here.

Information from joint distributions is important in economic settings compared to mere information from the marginal distributions,. An example of this is from the literature on poverty indices duclos2006robust. In that paper, the authors provide an example (their Figure 4) of two bivariate CDF $F$ and $G$ with bounded support with identical marginal distributions, but where one CDF dominates the other in the first order. Their example illustrates the point that there are important multivariate relations between joint CDF that cannot be captured by looking at only the marginal distributions. Transported to our setting, their examples satisfies part (i) of Assumption (ref), and model (ref) would be identified if these two distributions corresponded to $F_{X|Z=z}$ and $F_{X|Z=z'}$ under a few further assumptions on the model as specified below. On the other hand, the existing element-wise generalizations of torgovitsky2015identification and d2015identification will not find that the model is identified, because the marginal distributions in this example all coincide and therefore violate the requirement that they intersect in finitely many points.

This example is one special case of the more general setting where our approach can provide point-identification while existing approaches cannot: copulas, where all marginal distributions are equal, but the dependence structure changes. The above example by duclos2006robust is but one example, and there are many more examples where the copulas of $F_{X|Z=z}$ and $F_{X|Z=z'}$ are such that they satisfy part (i) or (ii) of Assumption (ref). These are straightforward to check in practice. As an extreme example, any copula will be first order-stochastically dominated by the copula taking on the lower Fr\'echet-Hoeffding bound, so that part (i) of Assumption (ref) will always be satisfied in this setting where we consider copulas with bounded support.

Examples for copulas that satisfy part (ii) of Assumption (ref) also abound. A simple example is illustrated in the left panel of Figure (ref). It depicts the contour plots of a Gumbel (i.e. extreme-value) copula with parameter $\theta=2$ and an AMH copula ali1978class with parameter $\theta=1$. Their manifold of intersection inside the support, $\mathcal{I}(F,G)$, is depicted in bold. Part (ii) of Assumption (ref) is easily satisfied for this example, which is the main criterion for point-identification of model (ref). On the other hand, all marginals coincide, so that the element-wise generalizations of the univariate approaches are not able to provide identification.

Finally, the right panel of Figure (ref) provides an example where part (i) of Assumption (ref) is satisfied without first-order stochastic dominance. The manifold of intersection $\mathcal{I}(F,G)$ here does not go from boundary to boundary, which means that there do not exist the required quantiles $\alpha$ and $\beta$ for part (ii). Since the support is compact and the set of intersection is a closed curve that is lower dimensional, part (i) of Assumption (ref) is satisfied and identification is guaranteed. However, if the support of the two distributions were not compact, then our method could not provide identification of the model.

figure[figure omitted — 569 chars of source]

Asymptotic properties

The condition proposed in Assumption (ref) for the identification of model (ref) is a restriction on the observable distributions $F_{X|Z=z}$ and $F_{X|Z=z'}$, which can be estimated in practice. Part (i) of Assumption (ref) is often impossible to check nonparametrically: the issue is to check whether the support is finite. Part (ii) of Assumption (ref) works for unbounded supports, and will therefore often be the main empirical criterion to check in completely nonparametric models, i.e. models where the applied researcher does not want to make assumptions on the distributions.

In this section, we therefore focus on the potential region of intersection of $F_{X|Z=z}$ and $F_{X|Z=z'}$ as prescribed by part (ii) of Assumption (ref). Estimating this in practice is straightforward, even in high-dimensions: all one has to do is estimate the function $D(x)\coloneqq F_{X|Z=z}(x)-F_{X|Z=z'}(x)$ by its empirical analogue $\hat{D}_n(x)=\hat{F}_{X|Z=z;n}(x)-\hat{F}_{X|Z=z';n}(x)$, where $\hat{F}_{X|Z=z;n}$ is the empirical analogue of the conditional CDF for $n$ observations, which can either be calculated via the standard conditional empirical CDF \[\hat{F}_{X|Z=z;n}(x)\coloneqq \frac{\frac{1}{n}\sum_{i=1}^n \prod_{j=1}^k\mathds{1}(X_{ij}\leq x_j)\mathds{1}(Z_i=z)}{\frac{1}{n}\sum_{i=1}^n\mathds{1}(Z_i=z)}\] or by a kernel smoothing approach. The classical empirical approach seems preferable in high dimensions, i.e. for large $k$, as there exist efficient and fast methods to compute the empirical distribution function in high dimensions (e.g. bentley1980multidimensional and langrene2020fast). In this section, we therefore focus on the classical case and not its smoothed alternative.

The following is the main result of the section, providing uniform large sample properties of $\hat{D}_n(x)$. This result can be used to provide uniform confidence intervals for $\hat{D}$, which is important when trying to check a multivariate intersection based on quantiles. In contrast, existing approaches estimating intersections of CDF as in duclos2006robust only provide pointwise confidence intervals, which in general do not provide confidence intervals which uniformly cover the intersection manifold $\mathcal{I}(F_{X|Z=z}, F_{X|Z=z'})$.

propositionIf $\{(X_i, Z_i)\}_{i=1,\ldots,n}\coloneqq \left\{(X_{i1},\ldots, X_{ik}, Z_i)\right\}_{i=1,\ldots,n}$ is a sequence of iid random variables, then the empirical process \[\sqrt{n}\left(\hat{D}_n(x)-D(x)\right)\coloneqq\sqrt{n}\left(\left(\hat{F}_{X|Z=z;n}-\hat{F}_{X|Z=z';n}\right)-\left(F_{X|Z=z}-F_{X|Z=z'}\right)\right)\] satisfies \[\sqrt{n}\left(\hat{D}_n(x)-D(x)\right)\rightsquigarrow \mathbb{G}(x)\] uniformly, where $\mathbb{G}(x)$ is a mean-zero Gaussian process with covariance kernel \begin{multline}CoV(\mathbb{G})(x,x')= \frac{F_{X|z=z}(\min\{x,x'\})-F_{X|Z=z}(x)F_{X|Z=z}(x')}{P(Z=z)}\\+\frac{F_{X|z=z'}(\min\{x,x'\})-F_{X|Z=z'}(x)F_{X|Z=z'}(x')}{P(Z=z')}\end{multline} for any $x,x'\in\mathcal{X}$.

Proposition (ref) is helpful in practice, as the confidence intervals can be straightforwardly obtained as the limit process is Gaussian. It is also an indication that the classical bootstrap method is valid in this setting. Based on this, one can check if the graph of $F_{X|Z=z}$ lies above the graph of $F_{X|Z=z'}$ for all $x$ which satisfy $F_{X|Z=z}(x)>\alpha$ and the reverse condition for all $x$ which satisfy $F_{X|Z=z}(x)<\beta$ for some values $\beta\leq\alpha\in(0,1)$.

Conclusion

In this article we have proposed a simple condition for the nonparametric point-identification of instrumental variable models with general unobserved heterogeneity and a multivariate first- and second stage. The main result is a direct generalization of the result from torgovitsky2015identification for point-identification of nonseparable triangular models with discrete instruments. Interestingly we can allow for (almost) arbitrary heterogeneity in the second stage. Instead of using a control variable approach imbens2009identification which in this form is not available in higher dimensions, we assume implicitly that the unobservable of the first stage is known. The key is a novel fixed-set iteration result for cyclically monotone maps between quasi-concave distribution functions. The asymptotic distribution of the relevant statistic for checking the condition is derived: the limit distribution is Gaussian which makes it particularly straightforward to check the condition in practice.