EconBase
← Back to paper

A Distance Covariance-based Estimator

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,028 characters · 16 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Distance Covariance-based Estimator

refsection\begin{abstract} This paper proposes an estimator that relaxes the conventional relevance condition in instrumental variable (IV) analyses. The method allows endogenous covariates to be weakly correlated, uncorrelated, or even mean-independent---though not independent---of the instruments, enabling the use of the maximal set of relevant instruments in a given application. Identification is attainable without exclusion restrictions and without finite-moment assumptions on the disturbance term. Under either of two non-nested exogeneity conditions, combined with mild regularity conditions, the parameter of interest is identified. The estimator is shown to be consistent and asymptotically normal, and the relaxed relevance condition required for identification is testable. Keywords: distance covariance, dependence, weak instrument, endogeneity, $ U $-statistics JEL classification: C13, C14, C26 \end{abstract}

Introduction

Empirical work in economics often relies on instrumental variable (IV) methods. However, when instruments are weakly correlated with endogenous covariates, conventional IV methods such as two-stage least squares (TSLS), the control function (CF) method, and the generalised method of moments (GMM) become unreliable, leading to biased estimates and hypothesis tests with significant size distortions. Furthermore, conventional IV methods are infeasible when excluded instruments are unavailable or uncorrelated with the endogenous variables. These conventional methods are also highly sensitive to outliers or non-existent moments of the disturbance term $U$. While much of the econometric literature on the weak instrument problem is focused on detection and weak-instrument-robust inference, theoretical progress on estimation is scant andrews2019weak. This paper introduces a new single-step estimator that minimises a scalar-valued measure of stochastic dependence between a parametrised disturbance $U(\theta)$ and a set of instruments $Z$ using the distance covariance measure (dCov) proposed by szekely2007measuring. The proposed Minimum Dependence estimator (MDep) substantially relaxes the instrument relevance requirement, allows for instruments $Z$ that are not independent of covariates $X$, and remains robust even when the disturbance term $U$ lacks finite moments.

The MDep has remarkable features that render it fundamentally different from existing IV methods. (1) The non-independence identifying variation means the MDep can exploit the maximum number of instruments available in any given empirical setting.\footnote{In a class of single-index models, for example, MDep relevance requires that no non-trivial linear combination of $X$ be independent of $Z$.} (2) In the absence of excluded instruments, identification in the MDep framework continues to hold as long as covariates $X$ are not independent of instruments $Z$. (3) Although the MDep does not estimate a quantile model, it shares the “robustness” property of quantile estimators—see, e.g., powell1991estimation,oberhofer2016asymptotic—in that its asymptotic properties do not rely on the existence of moments of $U$. By replacing $Z$ with a bounded one-to-one mapping such that $ Z $ and the mapping generate the same Euclidean Borel field, one obviates moment existence conditions on $Z$ as well.\footnote{An example of such a mapping is $ z \mapsto \mathrm{atan}(z) $.} This third feature is important as economic theory can go as far as justifying the exogeneity of instruments, but typically cannot go far enough to justify the existence of moments of $U$. This paper appears to be the first to introduce an IV estimator that exploits identifying variation from arbitrary stochastic dependence---of unknown and unspecified form---between \( X \) and \( Z \), in a broad class of models.

As the form of identifying variation needs to be neither known nor specified, the MDep framework effectively eliminates the sensitivity of estimates to first-stage model specification.\footnote{dieterle2016simple, for example, uncovers substantial sensitivity of conclusions to specification (linear versus quadratic) of the first stage.} Thus, often-imposed linearity or monotonicity restrictions on first-stage relationships, e.g., wooldridge2010econometric,dhaultfoeuille-fevrier-2015-identification,torgovitsky2017minimum, are unnecessary in the MDep framework. Although this property is also shared by integrated conditional moment estimators (ICM hereafter), e.g., dominguez2004consistent,escanciano2006consistent,antoine2014conditional,escanciano2018simple,tsyawo2023feasible, it is worth emphasising that the MDep relevance condition is more general. For example, $ \mathbb{P}\big(\mathbb{E}[X\mid Z] \neq \mathbb{E}[X]\big) > 0 $ neither implies nor is implied by $ \mathbb{P}\big(\mathbb{E}[Z\mid X]\neq \mathbb{E}[Z] \big) > 0 $. For identification, the MDep exploits both forms of dependence, while ICM estimators can only exploit the former. The MDep can achieve identification without excludability; this is more general than similar identification highlighted in tsyawo2023feasible,gao-wang-2023linIV for the ICM and IV classes of estimators, respectively.

The rest of the paper is organised as follows. (ref) discusses strands of related literature, while (ref) describes the dCov measure and presents the MDep estimator. (ref) derives theoretical results viz. identification, consistency, asymptotic normality, consistency of the covariance matrix estimator, and testability of the MDep relevance condition. (ref) examines the small sample performance of the MDep via simulations, and (ref) concludes. All proofs are relegated to the Appendix. Additional theoretical and simulation results are available in the Supplemental Appendix.

\paragraph*{Notation:} Define $ \mathbb{E}_n[\xi_i] := \frac{1}{n}\sum_{i=1}^{n}\xi_i $ and $ \mathbb{E}_n [\xi_{ij}] := \frac{1}{n(n-1)}\sum_{i=1}^{n}\sum_{j\neq i}^{n}\xi_{ij} $. For a random variable $\xi$, let $\xi^\dagger$ denote its independent and identically distributed ($i.i.d.$) copy, and define its symmetrised version as $\widetilde{\xi}:= \xi - \xi^\dagger $. Similarly for observations $ i \neq j$, define $\widetilde{\xi}_{ij} := \xi_i - \xi_j $. Independence between random variables is denoted by $\xi_1 \protect\mathpalette{\protect\independenT}{\perp} \xi_2 $. Let $p_\xi$ denote the dimension of $\xi$, and define $[p]:= \{1,\ldots,p\} $ for $ p\in \mathbb{N} $. The symbol $ || \cdot || $ denotes the usual Euclidean norm; $ a \vee b:= \max\{a,b\} $; and $ a \wedge b := \min\{a,b\} $. Finally, let $\widetilde{\sigma}\big(\xi\big)$ denote the sigma-algebra generated by $ [\xi,\xi^\dagger] $, and define the sign function as \( \operatorname*{sgn}( \xi ) := \big(1-2\mathbbm{1} \{ \xi \leq 0 \}\big) \).

Related Literature

The MDep minimises a scalar-valued criterion of stochastic dependence between a parametrised error $U(\theta)$ and a set of instruments $Z$. This approach builds on the tradition of Minimum Distance from Independence (MDI) estimators initiated by manski1983closest and further developed by brown-wegkamp-2002weighted,komunjer-santos-2010semi,gao-galvao-2014-minimum,dhaultfoeuille-fevrier-2015-identification,torgovitsky2017minimum,poirier-2017-efficient. Of the foregoing, only torgovitsky2017minimum explicitly considers identification cum estimation under endogeneity, as does this paper. torgovitsky2017minimum specifies and models a first-stage infinite-dimensional nuisance parameter (the conditional distribution $X\mid Z$). komunjer-santos-2010semi,dhaultfoeuille-fevrier-2015-identification,torgovitsky2017minimum require that covariates be continuously distributed---a substantive restriction, e.g., in settings with endogenous binary treatment. This paper imposes no support restrictions on \([X, Z]\), thereby accommodating a broader class of models, covariates and instruments, allowing for potentially non-monotonic first-stage relationships, and obviating continuity assumptions in the first stage. Moreover, the current paper appears to be the first to provide a tractable IV relevance condition in the class of MDI estimators.

The MDep estimator is related to ICM estimators, e.g., dominguez2004consistent,escanciano2006consistent,antoine2014conditional,escanciano2018simple,wang2018consistent,antoine2022partially,tsyawo2023feasible,song-jiang-ke-2024estimation. This class of estimators minimises the mean dependence of $U(\theta)$ on $Z$. Continuum Moment (CM) estimators---a related class of estimators---convert mean-independence restrictions into a continuum of unconditional moment conditions indexed by a nuisance parameter on an index set, and are typically estimated using IV methods such as Two-Stage Least Squares (TSLS) or the Generalized Method of Moments (GMM) (see, e.g., carrasco2000generalization,donald2003empirical,hsu2011estimation,carrasco2015regularized). Despite the advantages of both the ICM and Continuum Moment (CM) classes of estimators, two key differences set the MDep apart. First, endogenous covariates in the MDep framework can be mean-independent but stochastically dependent on instruments, e.g., at some quantile(s) that need not be known or determined. Thus, ICM/CM-relevant instruments are MDep-relevant by construction, whereas the converse does not hold. Second, unlike ICM/CM estimators, which require the existence of at least the first two moments of the disturbance for consistency and asymptotic inference, the MDep obviates the existence of any moment of the disturbance. Mean independence assumptions apply to the ICM, CM, and conventional IV classes and are often imposed as replacements for distributional exogeneity conditions. While mean independence is implied by distributional exogeneity, this holds under the implicit assumption that the mean exists.

Some existing works consider IV estimation without excludability by exploiting and modelling non-linear forms of dependence between endogenous and exogenous covariates, e.g., cragg1997using,dagenais1997higher,lewbel1997constructing,erickson2002two,rigobon2003identification,klein2010estimating,gao-wang-2023linIV. Unlike the foregoing, the MDep does not require the practitioner to construct moments or model first-stage relationships. It suffices that there be dependence between covariates and instruments that ought not to be known, modelled, or estimated. To enhance the practicality of this important feature, this paper demonstrates the testability of the MDep relevance condition.

The econometric literature on weak instruments largely focuses on detection and weak-instrument-robust inference (e.g., staiger1997instrumental,andrews2006optimal,kleibergen2006generalized,andrews2016condInference,sanderson2016weak,andrews2017unbiased)---see andrews2019weak for a review. Normal distributions of conventional IV estimates can be poor and hypothesis tests based on them can be unreliable when instruments are weak nelson1990distribution,nelson1990somefurther,bound1995problems. The MDep gives a new perspective to handling weak IVs in empirical practice; IV- or ICM/CM-irrelevant instruments can be MDep-strong, and this condition is testable.

By extracting non-linear identifying variation in instruments in order to boost instrument strength, some works employ flexible methods such as the non-parametric IV, e.g., donald2001choosing,newey2003instrumental,donald2003empirical,kitamura2004empirical,das2005instrumental, machine learning techniques, e.g., chen2020mostly, and regularisation or moment selection schemes, e.g., ng2009selecting,darolles2011nonparametric,belloni2012sparse,hansen2014instrumental,carrasco2015regularized. While it is conceivable to take transformations of instruments to extract more identifying variation, this approach may be limited, for example, when available instruments are non-monotone in endogenous covariates.\footnote{E.g., $ X = X^* + U $, $ Z = |X^*| $, $ U \protect\mathpalette{\protect\independenT}{\perp} Z $, and $ X^* $ is symmetrically distributed with mean zero. $ \mathrm{cov}[X,Z]=0 $, and no measurable (feasible) transformation of $ Z $ can induce correlation with $ X $.} Further, the aforementioned approach usually results in high dimensionality, unlike the MDep, which remains parsimonious in $Z$.

The dCov measure is primarily used in tests of independence that are consistent against all forms of dependence, including linear, non-linear, monotone, and non-monotone alternatives. This feature of the dCov measure accounts for the weak relevance condition in the MDep framework. Several applications of the dCov measure have emerged since the seminal work szekely2007measuring---see, e.g., sheng2013direction,szekely2014partial,shao2014martingale,park2015partial,su2017martingale,davis2018applications,xu2020martingale. The current paper departs from this literature by leveraging the dCov for estimation and inference under possible endogeneity.

The MDep Estimator

This section presents (1) motivating illustrative examples highlighting the MDep's key features, (2) the dCov measure, (3) an interesting class of applicable models, and (4) the MDep estimator.

Motivating examples

The MDep estimator has unique strengths relative to existing IV estimators. To explore these, consider the linear model

equation*[equation* omitted — 52 chars of source]

in the following examples. $Z$ is MDep-relevant as long as it is not independent of any non-trivial linear combination of $X$.

example[Non-monotone first stage] Suppose $X_1$ and $Z_1$ are such that \[ X_1 = Z^* + U \quad \text{ and } \quad Z_1 = \mathbbm{1} \{ |Z^*| < -\Phi^{-1}(0.25) \} , \] where $ [Z^*, U] \sim \mathcal{N}(0,\mathrm{I}_2)$ and $\Phi^{-1}(\cdot)$ is quantile function of the standard normal distribution. Clearly, $X_1$ and $Z_1$ are related through $Z^*$. However, there is no possible transformation of $Z_1$, without extra information, that induces mean dependence or correlation between $X_1$ and $Z_1$.
example[Identification without excludability I - non-monotone first-stage] Consider a slight modification of Example 3.1, where \[ X_2 = Z = 0.2Z^* + Z^{*2}. \] $Z$ is MDep-relevant without being excluded, as no non-trivial linear combination of $X_1$ and $X_2$ is independent of $Z$.
example[Identification without excludability II - skedastic function] Consider the setting where \[ X_1 = U\sqrt{1+Z^2}, \quad X_2 = Z, \quad \text{and} \quad \mathbb{E}[U] = 0. \] $Z$ is not IV-relevant. Moreover, $X_1$ is mean-independent of $Z$. However, relevance in the MDep framework holds as any non-trivial linear combination of $X_1$ and $X_2$ is dependent on $Z$.
example[Non-existent first moment of $U$] The disturbance, $U$, follows the Cauchy distribution $U \sim \mathcal{C}\big(0,\ 0.1 + |Z_1| \big)$ with conditional scale heterogeneity. Existing IV methods, such as Conventional IV, non-parametric IV, ICM, and CM estimators, are inconsistent when the first moment of $U$ does not exist. The MDep, in contrast, is consistent.

The MDep explores identifying variation in all the examples given above while conventional IV, non-parametric IV, ICM, and CM methods fail. The above scenarios serve to highlight the remarkable features of the MDep relative to existing conventional methods.

The dCov measure

It is instructive to briefly review the distance covariance (dCov) measure introduced by szekely2007measuring, which underpins the MDep objective function.

definitionThe square of the distance covariance between random variables $ \Upsilon $ and $ Z $ with finite first moments is defined by szekely2007measuring as \begin{equation} \begin{split} \mathcal{V}^2(\Upsilon,Z) &= \int\big| \varphi_{\Upsilon,Z}(t,s) - \varphi_{\Upsilon}(t)\varphi_{Z}(s) \big|^2 w(t,s)dtds\\ & = \int \Big| \mathbb{E}\big[\exp(\iota(t'\Upsilon+s'Z))\big] - \mathbb{E}\big[\exp(\iota t'\Upsilon)\big]\mathbb{E}\big[\exp(\iota s'Z)\big] \Big|^2w(t,s)dtds \end{split} \end{equation} where $ \varphi_{\xi}(.)$ denotes the characteristic function of $\xi$, $\iota = \sqrt{-1} $, and the integrating measure $ w(t,s) $ is an arbitrary positive function for which the integral exists. The modulus is defined as $ |\zeta|^2 = \zeta\bar{\zeta} $, where $ \bar{\zeta} $ is the complex conjugate of $\zeta$.

Using the integrating measure $ w(t,s) = (c_{p_\Upsilon}c_{p_Z}||t||^{1+p_\Upsilon}||s||^{1+p_Z})^{-1} $ where $ c_p = \frac{\pi^{(1+p)/2}}{\Gamma((1+p)/2)}, \ p \geq 1 $, and $ \Gamma(\cdot) $ is the complete gamma function, szekely2007measuring obtains a distance covariance measure, which is shown in (ref) to have the representation \[ \mathcal{V}^2(\Upsilon,Z) = \mathbb{E}[\mathcal{Z}||\Upsilon-\Upsilon^\dagger||] \] where $ \mathcal{Z} := h(Z,Z^\dagger) $ and $h(z_a,z_b) := ||z_a - z_b|| - \mathbb{E}\big[||z_a - Z||+||Z - z_b||\big] + \mathbb{E}\big[||Z - Z^\dagger||\big]$. From (ref), one observes that $ |\varphi_{U,Z}(t,s) - \varphi_{U}(t)\varphi_{Z}(s)|_w^2 \geq 0 $.

This paper follows szekely2014partial in using the following algebraically equivalent form of the unbiased estimator of the distance covariance measure

equation[equation omitted — 172 chars of source]

where $ \mathcal{Z}_{ij,n} := h_n(Z_i,Z_j)$ such that

equation[equation omitted — 296 chars of source]

szekely2007measuring's szekely2007measuring integrating measure $ w(t,s) = (c_{p_\Upsilon}c_{p_Z}||t||^{1+p_\Upsilon}||s||^{1+p_Z})^{-1} $, besides yielding a reliable measure of dependence, results in a computationally tractable measure, which does not require numerical integration, obviates the choice of smoothing parameters (e.g., bandwidth or number of approximating terms in non-parametric approaches), and admits multiple instruments. The simplified formulation (ref) offers two key advantages for the proposed estimator: the permutation symmetry of $ \mathcal{Z}_{ij,n} = \mathcal{Z}_{ji,n} $ facilitates the use of U-statistic theory in establishing asymptotic normality and reduces the computational burden in evaluating (ref).

For ease of reference, the properties of the dCov measure in szekely2007measuring,szekely2009brownian are stated below.

property*The following properties hold for the distance covariance measure under the condition $ \mathbb{E}\big[||\Upsilon||^2 + ||Z||^2\big]< \infty $: \begin{enumerate}[label=(\alph*)] • $ \mathcal{V}^2(\Upsilon,Z) \geq 0 $; • $ \mathcal{V}^2(\Upsilon,Z) = 0 \text{ if and only if } \Upsilon $ and $ Z $ are independent; • $ \mathcal{V}^2(\Upsilon,Z) = \mathbb{E}\big[\mathcal{Z}||\Upsilon - \Upsilon^\dagger||\big] $; and • $ \mathbb{E}[\mathcal{V}_n^2(\Upsilon,Z)] = \mathcal{V}^2(\Upsilon,Z) $ for $ n>3 $ and $i.i.d.$ samples $ \big\{ [\Upsilon_i,Z_i] \, : \, i \in [n] \big\} $. \end{enumerate}

The properties are proved in the following: Property (ref) in szekely2009brownian, Property (ref) in szekely2007measuring, Property (ref) in (ref) of this paper, and Property (ref) in szekely2014partial.

Model specification

For a tractable characterisation and statistical testing of the MDep relevance identification condition, consider regression models in which the outcome \( Y_i \) is generated as

equation[equation omitted — 88 chars of source]

where $ G(\cdot) $ is a known invertible function and $ g(\cdot) $ is a known differentiable function with unknown parameter vector $ \theta_o \in \mathbb{R}^{p_\theta} $. $ \theta_{o,c} $ is the location parameter of $U(\theta_o)$, where $ U(\theta) := G^{-1}(Y) - g(X\theta) $ denotes the parametrised disturbance function. $X_i$ contains a constant term. The dependence of $ U_i(\theta) $ on $ X_i $ is suppressed for notational ease.

The class of models under consideration includes interesting examples such as the linear model $ U_i(\theta) = Y_i - X_i\theta $ (where the location parameter coincides with the intercept), non-linear parametric models, e.g., $ U_i(\theta) = Y_i - \exp(X_i\theta) $, fractional response models, e.g., $ U_i(\theta) = \log( Y_i/(1-Y_i)) - X_i\theta $, and special cases of Box-Cox models e.g., $ U_i(\theta) = \log(Y_i) - X_i\theta $. See (ref) for a more general class of applicable models.

Estimation

Let \( \big\{W_i = [Y_i, X_i, Z_i]: i \in [n] \big\} \) be a random sample of $W:= [Y, X, Z] $ defined on a probability space \( (\mathcal{W}, \mathscr{W}, \mathbb{P}) \). The MDep estimator is the minimiser of \( \mathcal{V}_n^2\big(U(\theta),Z\big) \), namely

equation[equation omitted — 207 chars of source]

where $\mathcal{Z}_{ij,n} := h_n(Z_i,Z_j)$ as defined in (ref) and $\widetilde{U}_{ij}(\theta) = U_i(\theta) - U_j(\theta)$. It may be of interest to estimate a location parameter for $U$, e.g., the median: $ \displaystyle \widehat{\theta}_{n,c} = \operatorname*{arg\,min}_{t} \sum_{i=1}^n | U_i(\widehat{\theta}_n) - t | $.\footnote{The asymptotic properties of $\widehat{\theta}_{n,c}$ are omitted since they can be derived straightforwardly from those of $\widehat{\theta}_n$.}

Following huber1967behavior, the minimand in ((ref)) is normalised as

equation[equation omitted — 189 chars of source]

in order to avoid unnecessary moment conditions on $ U $---e.g., powell1991estimation,oberhofer2016asymptotic. This holds even though the dCov measure itself requires the existence of $\mathbb{E}[|U|]$---cf. szekely2007measuring.

Asymptotic Theory

It follows from (ref) that the asymptotic theory for the MDep estimator belongs to the broader class of estimators based on \( U \)-statistics--type objective functions (e.g., honore1994pairwise,honore-powell-2005-pairwise,jochmans2013pairwise), as well as those involving non-smooth objective functions such as quantile regression (QR) (e.g., koenker1978regression,powell1991estimation,oberhofer2016asymptotic), instrumental-variable QR methods (e.g., chernozhukov2006instrumental,chernozhukov2008instrumental), and the control-function approach to QR of lee2007endogeneity. Let the Jacobian and its symmetrised version be defined by \[ X^g(\theta) := -\frac{\partial U(\theta)}{\partial \theta'} \text{ and } \widetilde{X^g}(\theta) := X^g(\theta)-X^{g\dagger}(\theta), \]respectively. Also, define \( \displaystyle \widetilde{X^{gg}}(\theta):= \frac{\partial \widetilde{X^g}(\theta)}{\partial \theta}. \) The parameter vector $ \theta_o$ is the MDep estimand, and $U:= U(\theta_o) - \theta_{o,c} $.

Regularity conditions

Two sets of regularity conditions imposed in the paper guarantee the consistency of the MDep estimator $ \widehat{\theta}_n $. The first set outlined in the following comprises smoothing and dominance conditions, ensuring that the difference between the normalised minimand and its expectation converges to zero uniformly in $ \theta \in \Theta $.

assumption[Regularity] \quad \begin{enumerate}[label=(\alph*)] • $ U(\theta) $ is measurable in $ [U,X] $ for all $ \theta $ and is twice continuously differentiable in $ \theta $ for all $ [U,X] $ on the support of $ [U_i,X_i] $. $ X^g(\theta) = g'(X\theta)X $ is measurable in $ X $ and $ \mathbb{P}\big( g'(X\theta) = 0 \big) < 1 $ for all $ \theta \in \Theta $. • For some constant $ C \in (0,\infty) $, \quad $\displaystyle \mathbb{E}\Big[ \Big( \{|\mathcal{Z}|\vee 1\} \cdot \{ \sup_{\theta \in \Theta }||\widetilde{X^g}(\theta)|| \vee 1 \} \Big)^4\Big] \leq C $, \quad $ \displaystyle \mathbb{E}\Big[ \sup_{\theta \in \Theta }\Big\lVert \{|\mathcal{Z}|\vee 1\} \cdot \widetilde{X^{gg}}(\theta) \Big\rVert^2 \Big] \leq C $, \quad and \quad $\displaystyle \mathbb{E}\Big[ \big\lVert \widetilde{Z} \big\rVert^4\Big] <\infty. $$ \Theta $ is a compact parameter space. \end{enumerate}

(ref) and the differentiability requirement in (ref)(ref) characterise an interesting class of models considered in this paper, e.g., the linear model. $ \widetilde{U}(\theta) = \widetilde{U} - \widetilde{X^g}(\bar{\theta})(\theta-\theta_o) $ for some $ \bar{\theta} $ lying on the line segment between $ \theta $ and $ \theta_o $ is a useful expression for subsequent analyses thanks to (ref)(ref) and the Mean-Value Theorem (MVT). The technical requirement $ \mathbb{P}\big(g'(X\theta) = 0 \big) < 1 $ is important for identification as the expression $ \widetilde{U}(\theta) = \widetilde{U} - \widetilde{X^g}(\bar{\theta})(\theta-\theta_o) $ with $ \widetilde{X^g}(\theta) = \big(g'(X\theta)X - g'(X^\dagger\theta)X^\dagger\big) $ shows that $ \widetilde{U}(\theta) $ can equal $ \widetilde{U} $ almost surely ($a.s.$) for some $ \theta\neq \theta_o $ if it is violated.

(ref)(ref) is an MDep analogue of uniform moment bounds; it implies \( \displaystyle \mathbb{E}\big[\sup_{\theta \in \Theta }||\mathcal{Z}\widetilde{X^g}(\theta)||^4\big] \leq C \) and \( \displaystyle \mathbb{E}\big[\sup_{\theta \in \Theta }||\widetilde{X^g}(\theta)||^4\big] \leq C.\) (ref)(ref) can be further weakened by replacing $Z$ with bounded one-to-one mappings such that $ Z $ and the mapping generate the same Euclidean Borel field, e.g., $ \mathrm{atan}(Z) $ ---see bierens1982consistent and szekely2007measuring---thereby allowing $Z$ (in addition to $U$) to have no finite moments. In that case, $\mathcal{Z}$ can be dropped from (ref)(ref). (ref)(ref) is required since the objective function (ref) is non-convex.

Identification and consistency

The second set of regularity conditions for consistency ((ref), (ref), and (ref)) are identification conditions that ensure that \( Q(\theta) := \mathbb{E}\big[ \mathcal{Z}(|\widetilde{U}(\theta)| - |\widetilde{U}|) \big] \) is uniquely minimised at $\theta_o$. The first identification assumption concerns the relevance condition in the MDep framework.

assumption[Relevance] $ X\uptau \not\protect\mathpalette{\protect\independenT}{\perp} Z $ for all $ \uptau \neq 0$.

(ref) is the condition of non-independence between non-trivial linear combinations of $ X $ and $ Z $; it is the MDep analogue of the relevance condition in the IV setting, e.g., wooldridge2010econometric, and an MDep analogue of the linear completeness condition in ICM estimators, e.g., escanciano2018simple,tsyawo2023feasible. In the IV setting, the relevance condition requires that no non-zero linear combination of $ X $ be uncorrelated with $ Z $. The ICM relevance condition requires that no non-zero linear combination of $ X $ be mean-independent of $ Z $. (ref) requires that no non-zero linear combination $ X $ be independent of $ Z $. As independence implies mean independence, which in turn implies uncorrelatedness, it follows that the MDep relevance condition ((ref)) is the weakest possible. In a simple case with a univariate $ X $, (ref) allows $ X $ to be uncorrelated, or mean-independent of $ Z $ as long as $ X $ is not independent of $ Z $. All IV-strong or ICM-strong instruments are therefore MDep-strong by construction. The converse is, however, not true. Like in the case of ICM estimators, (ref) can hold even if there are fewer instruments than covariates, e.g., tsyawo2023feasible. This feature of the MDep can be explored to attain identification without excludability: (ref).

remarkThe MDep accommodates the broadest possible set of instruments in any empirical setting: it includes all IV- and ICM-relevant instruments, and even those that are IV- or ICM-irrelevant yet dependent on covariates in the sense of (ref).

The single-index structure of the models in (ref) offers the advantage of a tractable characterisation and statistical testing of the relevance condition ((ref)). General non-single-index structures and non-additively separable disturbance functions can be considered at the cost of a less intuitive relevance identification condition.

remarkA more general class of applicable models accommodates potentially non-single-index structures, non-additive disturbances, or both, taking the form $Y = G(X, U; \theta_o)$, where $G(X, U; \theta)$ is invertible in $U$, such that $U(\theta_o) := G^{-1}(Y, X; \theta_o)$, and $X$ may be endogenous. In this broader setting, the relevance condition becomes $X^g(\theta)\uptau \not\protect\mathpalette{\protect\independenT}{\perp} Z$ for all $\uptau \neq 0$ and $\theta \in \Theta$.\footnote{See the discussion around (ref).}

Two non-nested exogeneity conditions apply under the MDep framework. The first is a standard MDI exogeneity condition of independence between $Z$ and $U$.

assumption[Exogeneity I] \( U \protect\mathpalette{\protect\independenT}{\perp} Z \).

From a model specification perspective, (ref) is testable using the tests of sen2014testing,davis2018applications,xu2021omnibus. (ref) rules out conditional scale heterogeneity, e.g., heteroskedasticity. However, exploiting the absolute value in the objective function (ref), the following exogeneity condition can also be exploited for identification.

namedassumption{(ref)$^\prime$}[Exogeneity II] \( \mathrm{med}\big[ (U - U^\dagger) \mid \widetilde{\sigma}( [X,Z] )\big] = 0 \ a.s. \)

Exogeneity in the MDep framework only requires either (ref) or (ref) to hold. Moreover, both exogeneity conditions are non-nested. Consider two DGPs with $X=Z+V$: (a) $U=\rho V + \xi$, $ \rho \neq 0 $ with $ Z, \, V, \, \xi $ all independent and (b) $U=|X|\xi, \, \xi \sim \mathcal{N}(0,1) $. (a) satisfies (ref) but not (ref), whereas (b) satisfies (ref) but not (ref).

(ref) accommodates some form of conditional scale heterogeneity of $U$ in $ [X,Z] $, e.g., conditional heteroskedasticity.\footnote{Heteroskedasticity in the traditional sense does not apply to heavy-tailed distributions such as the Cauchy. However, it is conceivable that the scale parameter of \( (U - U^\dagger) \mid [X, X^\dagger, Z, Z^\dagger] \) is non-degenerate.} In the aforementioned example (b), $(U-U^\dagger)\mid [X,X^\dagger,Z,Z^\dagger] \ \sim \mathcal{N}(0,|X|+|X^\dagger|) $, whence $ \mathrm{med}\big[ (U - U^\dagger) \mid X,X^\dagger,Z,Z^\dagger\big] = 0 \ a.s.$\footnote{This type of characterisation applies to the entire family of symmetric $\alpha$-stable distributions.} Unlike the ICM and conventional IV estimators, the MDep is not robust to arbitrary forms of heteroskedasticity if $\mathbb{E}[U^2]<\infty$.\footnote{If the violation of (ref) arises solely from arbitrary scale heterogeneity in \( U \), a potential remedy---left unexplored in this paper---is to estimate the conditional scale function alongside $\theta_o$ and scale-standardise $U(\theta)$ \`a la, e.g., wooldridge2010econometric,romano2017resurrecting,alejo2024endogenous.} (ref) requires that the median of $ (U - U^\dagger) $ conditional on $\widetilde{\sigma}( [X,Z] )$ be zero almost surely, thereby unifying the median, the mean (if it exists), and the mode (if $ \widetilde{U} $ is unimodal) as a natural point on which to impose exogeneity, thanks to symmetrisation. Unlike (ref), which is imposed on pairwise differences in disturbances, similar exclusion restrictions on conditional quantiles are imposed on the levels of disturbances for quantile estimators under (possible) endogeneity, see e.g., chernozhukov2006instrumental, lee2007endogeneity, and powell1991estimation. (ref) can be expressed as $ \mathbb{E}\big[\mathbbm{1} \{\widetilde{U} \leq 0 \} - 0.5 \mid \widetilde{\sigma}( [X,Z] )\big] = 0 \ \text{a.s.} $; this condition is testable from a model specification perspective using a suitable extension of, for example, ICM tests---see bierens1982consistent,dominguez2015simple,su2017martingale,xu2020martingale,jiang-tsyawo-2022consistent.\footnote{This task, however, is left for future work due to considerations of scope and space. }

remarkNeither (ref) nor (ref) requires the existence of any moment of \( U \). (ref) is tied to the integrating measure of szekely2007measuring, which yields the absolute value function in (ref). As a result, the MDep behaves like a specially weighted least absolute deviations (LAD) estimator on pairwise differences in disturbances. In contrast, arbitrary integrating measures in (ref) do not deliver this extra property.

The MDep objective function (ref) is non-convex because $ \mathcal{Z}_{ij,n} $ is not non-negative. This renders typical QR identification proof techniques that draw on the convexity of the objective function, e.g., koenker1978regression,powell1991estimation,oberhofer2016asymptotic, inapplicable. In contrast, this paper leverages the non-negativity and “omnibus" properties of the dCov measure---namely Properties (ref) and (ref)---to establish identification.

theoremSuppose (ref) hold. If, in addition, either Assumption (ref) or (ref) is satisfied, then for every $\varepsilon > 0$, there exists a constant $\delta_\varepsilon > 0$ such that \[ \inf_{\{\theta \in \Theta : \|\theta - \theta_o\| \ge \varepsilon\}} Q(\theta) > \delta_\varepsilon. \]

(ref) shows that under the given assumptions, the minimand $Q(\theta)$ has a unique minimum.

For illustrative purposes, consider the setting where \( \theta_o = 0 \), \( X \sim \mathrm{Ber}(0.5) \), \( X=Z \protect\mathpalette{\protect\independenT}{\perp} U \), and \( Y = X\theta_o + U \), under three distributions for \( U \): (a) \( U \sim \mathcal{N}(0,0.5) \), (b) \( U \sim \mathcal{C}(0,0.5) \), and (c) \( U \sim \mathcal{U}[0,\sqrt{6}] \). The corresponding population objective functions \( Q(\theta) := \mathbb{E}[\mathcal{Z}(|\widetilde{U}(\theta)| - |\widetilde{U}|)] \) are plotted in (ref). The minima are well defined, \( X \) is discrete, and \( U \), in case (b), lacks a finite first moment.

figure[figure omitted — 549 chars of source]

With the identification result in hand, this subsection concludes with a proof of consistency of the MDep. The following standard sampling scheme is imposed.

assumption$ \big\{W_i: i \in [n] \big\} $ are independently and identically ($i.i.d.$) distributed random vectors.
theoremSuppose the conditions of (ref) hold, then in addition to (ref), the MDep $\widehat{\theta}_n$ converges almost surely to $\theta_o$ as $n \to \infty$, i.e., \( \widehat{\theta}_n \xrightarrow{a.s.} \theta_o. \)

Conditional functionals and parameters of interest

Whenever elements of $\theta_o$ are themselves of interest, e.g., in a structural economic model with an economically meaningful $\theta_o$, the interpretation is direct. However, when $\theta_o$ is not of direct interest per se, but the partial effects obtained therefrom are, it is essential first to determine the identified conditional functional.

Consider the simple linear model $ Y = X\theta_o + U $ where $ X=Z $ and $p_X=1$. Under (ref), $ Q_{Y|X}(\tau|x) = x\theta_o $ for all $\tau \in (0,1)$ where $ Q_{Y|X}(\tau|x) $ is the $\tau$'th quantile of $Y$ conditional on $X=x$. When $ \mathbb{E}[|U|]<\infty $, then $ \mathbb{E}[Y|X=x] = x\theta_o $ as well. Under (ref), $ \mathrm{med}[ (Y - Y^\dagger) \mid (X-X^\dagger) ] = (X-X^\dagger)\theta_o $. Hence, $\theta_o$ is the median partial effect of a unit increase in $X$ on the outcome $Y$, relative to an observationally equivalent agent.

Unlike the simple linear example above, the partial effect of $X$ is not constant for non-linear $G(\cdot)$. For example, consider the model $ \log(Y) = X\theta_o + U $ where $G(\cdot)=\exp(\cdot)$. Under (ref), $ \mathrm{med}[ \log(Y) - \log(Y^\dagger) \mid X,X^\dagger] = \log\big(\mathrm{med}[(Y/Y^\dagger) \mid X,X^\dagger]\big) = (X-X^\dagger)\theta_o $, i.e., $ \displaystyle \mathrm{med}\Big[ \frac{Y-Y^\dagger}{Y^\dagger} \Big| (X-X^\dagger) \Big] = \exp\big((X - X^\dagger)\theta_o\big)-1$, and the partial effects are interpretable as changes in fractions or percentages. As the resulting partial effect is a function of $[X,X^\dagger]$, interesting summaries of this heterogeneity can be reported, such as the average partial effect or the partial effect at the average.

Asymptotic normality

Define the score function $ \mathcal{S}_n(\theta):= \mathbb{E}_n[\psi(W_i,W_j;\theta)]$ where $ \psi(W_i,W_j;\theta) := \mathcal{Z}_{ij}\operatorname*{sgn}\big( \widetilde{U}_{ij}(\theta) \big) \widetilde{X^g}_{ij}(\theta)' $, with $ \psi(W_i,W_j) := \psi(W_i,W_j;\theta_o) $. $ h(Z_i,Z_j) =: \mathcal{Z}_{ij} = \mathcal{Z}_{ji} $ and $ \operatorname*{sgn}( \widetilde{U}_{ij} ) \widetilde{X^g}_{ij} = \operatorname*{sgn}( \widetilde{U}_{ji} ) \widetilde{X}_{ji}^g $ with $ \widetilde{X^g}_{ij} := \widetilde{X^g}_{ij}(\theta_o) $ hence $ \psi(\cdot,\cdot) $ is permutation symmetric. Denote the cumulative distribution function and the probability density functions of $ \widetilde{U} $ conditional on $\widetilde{\sigma}( [X,Z] )$ by $ F_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(\cdot) $ and $ f_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(\cdot) $, respectively. Further, define $ v_n(\theta):= \sqrt{n}\big( \mathcal{S}_n(\theta) - \mathcal{S}(\theta) \big) $ where $ \mathcal{S}(\theta) := \mathbb{E}[\mathcal{S}_n(\theta)] = \mathbb{E}\big[\psi(W,W^\dagger;\theta)\big] $, $ \psi^{(1)}(W_i) := \mathbb{E}\big[\psi(W_i,W_j)|W_i\big] $, and the Hessian \[ \mathcal{H} := 2\mathbb{E}\big[f_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(0)\mathcal{Z}\widetilde{X^g}'\widetilde{X^g}\big] + \mathbb{E}\big[\operatorname*{sgn}(\widetilde{U}) \mathcal{Z}\widetilde{X^{gg}}\big]. \] Finally, let $ \partial^{-}|\hat{q}| $ and $ \partial^{+}|\hat{q}| $, respectively, denote the left- and right-derivatives of $ |q| $ with respect to $q$ at $ q=\hat{q} $.

assumption[Asymptotic Linearity of $ \widehat{\theta}_n $] \quad \begin{enumerate}[label=(\alph*)] • $ \theta_o $ is an interior point of $ \Theta $; • $ F_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(\cdot) $ is continuously differentiable with density $ f_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(\cdot) $, and there exists a constant $ f_o \in (0,\infty) $ such that, for all $ \epsilon $ in a neighbourhood of zero, $ \displaystyle f_o^{-1} < f_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(\epsilon) \leq \sup_{e \in \mathbb{R}}f_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(e) \leq f_o^{1/4} \ a.s. $$ \mathcal{H} $ is non-singular. \end{enumerate}

(ref)(ref) is standard. Conditions similar to (ref)(ref) are standard in the quantile regression literature---cf. lee2007endogeneity, chernozhukov2006instrumental, chernozhukov2008instrumental, powell1991estimation, oberhofer2016asymptotic), and xu2021omnibus. It ensures the Hessian is well-defined. As $ \mathcal{Z} $ has both negative and positive values in its support, the Hessian $ \mathcal{H} = 2\mathbb{E}\big[f_{\widetilde{U} \mid \widetilde{\sigma}( [X,Z] )}(0)\mathcal{Z}\widetilde{X^g}'\widetilde{X^g} \big] + \mathbb{E}\big[\operatorname{sgn}(\widetilde{U}) \mathcal{Z}\widetilde{X^{gg}}\big] $ cannot be positive definite by construction; non-singularity ((ref)(ref)) is thus required---cf. honore1994pairwise. The second term in the Hessian disappears if (ref) holds (which ensures $ \mathbb{E}\big[\operatorname*{sgn}(\widetilde{U}) \mid \widetilde{\sigma}([X,Z]) \big] = 0 \, a.s. $ ) or the model is linear (which implies $ \widetilde{X^{gg}}(\theta) = 0 $ for all $ \theta \in \Theta $ ).

Define $ \Omega := 4\mathbb{E}[\psi^{(1)}(W)\psi^{(1)}(W)'] $. The following theorem states the asymptotic linearity and normality of the MDep estimator.

theoremSuppose that (ref) holds in addition to the conditions of (ref). Then the MDep estimator \( \widehat{\theta}_n \) satisfies: \begin{enumerate}[(a)] • asymptotic linearity: \begin{equation*} \sqrt{n}(\widehat{\theta}_n - \theta_o) = -\,\mathcal{H}^{-1} \cdot \frac{2}{\sqrt{n}} \sum_{i=1}^{n} \psi^{(1)}(W_i) + o_p(1) \quad and\, ; \end{equation*} • asymptotic normality: \[ \sqrt{n}(\widehat{\theta}_n - \theta_o) \xrightarrow{d} \mathcal{N}\big(0,\, \mathcal{H}^{-1} \Omega \mathcal{H}^{-1}\big). \] \end{enumerate}

(ref) establishes the asymptotic normality of the MDep estimator. However, like other MDI estimators, the MDep is not efficient poirier-2017-efficient. Although a two-step procedure (not implemented in this paper) for achieving efficiency---along the lines of dominguez2004consistent---can be applied, it introduces several complications. Specifically, such an approach would require: (1) smoothing the inherently non-smooth moment equations of the MDep, (2) non-parametrically estimating components of the efficient GMM objective function, (3) selecting tuning parameters for both smoothing and estimation steps, and (4) accepting the risk of identification failure or inconsistency if the error term \( U \) lacks finite moments.

Consistent covariance matrix estimation

The preceding subsection established the asymptotic normality of the MDep estimator. Building on that result, this subsection introduces a consistent estimator of the asymptotic covariance matrix and proves its consistency. This consistency is crucial for conducting valid statistical inference, including $t$-tests, Wald tests, and the construction of confidence intervals. Define $\displaystyle \widehat{\psi}^{(1)}(W_i) := \frac{1}{n-1} \sum_{j\neq i}^{n} \widehat{\psi}(W_i,W_j) $ where $ \widehat{\psi}(W_i,W_j) := \mathcal{Z}_{ij,n}\operatorname*{sgn}\big( \widetilde{U}_{ij}(\widehat{\theta}_n) \big) \widetilde{X^g}_{ij}(\widehat{\theta}_n)'$. The estimators of $ \Omega $ and $ \mathcal{H} $ are given by \( \displaystyle \widehat{\Omega}_n = 4\mathbb{E}_n[\widehat{\psi}^{(1)}(W_i)\widehat{\psi}^{(1)}(W_i)'] \) \quad and

equation*[equation* omitted — 451 chars of source]

respectively, where $ \hat{c}_n $, a possibly random bandwidth sequence, and the uniform kernel, as proposed by powell1991estimation, is used to estimate the conditional density in $ \mathcal{H} $.\footnote{The second term in $ \widehat{\mathcal{H}}_n $ is identically zero for the linear model.} The estimator of the covariance matrix is $ \widehat{\mathcal{H}}_n^{-1} \widehat{\Omega}_n \widehat{\mathcal{H}}_n^{-1} $. An additional condition is imposed on the bandwidth sequence \( \hat{c}_n \) to ensure the consistency of \( \widehat{\mathcal{H}}_n \).

assumptionFor some non-stochastic sequence $ c_n $ with $ c_n \rightarrow 0 $ and $ \sqrt{n}c_n \rightarrow \infty $, $\displaystyle \operatorname*{plim}_{n \rightarrow \infty} (\hat{c}_n/c_n) = 1 $.

(ref) corresponds to powell1991estimation and is used to establish the consistency of $ \widehat{\mathcal{H}}_n $. It requires that the bandwidth sequence $\hat{c}_n$ satisfy the rate conditions $\hat{c}_n = o_p(1)$ and $\hat{c}_n^{-1} = o_p(\sqrt{n})$. The following theorem states the consistency of the covariance matrix estimator.

theoremSuppose the conditions of (ref) hold. If, in addition, (ref) holds, then $\widehat{\mathcal{H}}_n^{-1} \widehat{\Omega}_n \widehat{\mathcal{H}}_n^{-1} \xrightarrow{p} \mathcal{H}^{-1}\Omega\mathcal{H}^{-1} $ as $n \rightarrow \infty$.

Estimating the asymptotic covariance matrix involves specifying the bandwidth sequence $ \hat{c}_n $. The bandwidth sequence used throughout this paper follows the approach in koenker2005quantile and is given by \( \displaystyle \hat{c}_n = \sqrt{2} k_n \min\Big\{\widehat{\sigma}_{\widehat{U}},\ \frac{\mathrm{IQR}(\widehat{U})}{1.34}\Big\} \) where $k_n: = n^{-1/3}\Big(\frac{3}{4\pi}\big(\Phi^{-1}(0.975)\big)^2\Big)^{1/3} $ is the hall1988distribution bandwidth sequence. The terms $\widehat{\sigma}_{\widehat{U}}$ and $\mathrm{IQR}(\widehat{U})$ denote the sample standard deviation and inter-quartile range, respectively, of the residuals $ \big\{\widehat{U}_i, \ i \in [n] \big\} $.

Testing the MDep relevance condition

The weak relevance condition ((ref)) makes the MDep a powerful tool in a practitioner's toolkit, especially in dealing with unavailable or weak instruments. The practical usefulness of the MDep thus lies crucially in its testability. This subsection demonstrates the testability of the MDep relevance condition ((ref)) within the class of single-index models.\footnote{MDep relevance in the more general class in (ref) is left for future work.} Partition $ X $ as $ X = [D, \, Z_{-D}]$ where $ D \in \mathbb{R}^{p_D} $ and $Z_{-D} \in \mathbb{R}^{p_X - p_D} $, respectively, collect endogenous and exogenous covariates. Define $ \mathcal{D}_l(\gamma) := D_l - [D_{-l},\ Z_{-D}]\gamma, \ l\in [p_D] $ where $D_l$ denotes the $l$'th element of $D$ and $D_{-l} \in \mathbb{R}^{p_D - 1} $ excludes $D_l$ from $D$. Let $\mathbb{S}^p$ denote a compact subset of $\mathbb{R}^p,\, p\geq 1$. The following theorem shows the testability of the MDep relevance condition.

theoremSuppose (ref)(ref) holds, then a test of MDep relevance ((ref)) can be formulated via the following hypotheses: \begin{align*} \mathbb{H}_o &: \mathcal{D}_{l^*}(\gamma^*) \mathpalette{\independenT}{\perp} Z for some \{\gamma^*, \ l^* \} \in \mathbb{S}^{p_X - 1} \times [p_D]; and \\ \mathbb{H}_a &: \mathcal{D}_l(\gamma) \not \mathpalette{\independenT}{\perp} Z for all \{\gamma, \ l \} \in \mathbb{S}^{p_X - 1} \times [p_D]. \end{align*}

Thanks to Properties (ref) and (ref) of the dCov measure, $\mathbb{H}_o$ and $\mathbb{H}_a$ can be equivalently cast as

align*[align* omitted — 333 chars of source]

It follows from the above reformulation that (ref) is testable using tests of independence between MDep regression disturbance terms $ \mathcal{D}_l(\gamma), \, l\in [p_D] $ and $Z$, e.g., sen2014testing,davis2018applications,xu2021omnibus.

Simulation Experiments

This section examines the finite sample performance of the MDep using simulations. $Y = [X_1, X_2]\theta_o + U$ is the data-generating process, where $\theta_o = [0.5, \ -0.5]'$. Auxiliary variables include $\dot{X} \sim \mathcal{N}(0,\mathrm{I}_2) $, $V = Ua + \dot{U}\sqrt{1-a^2}, \, a=-0.2$, $\dot{U} \sim \mathcal{U}[-\sqrt{3},\sqrt{3}] $, and $ U\protect\mathpalette{\protect\independenT}{\perp} \dot{U} $. $U\sim (\chi_1^2-1)/\sqrt{2} $ unless otherwise specified. The data-generating processes (DGPs) considered are the following.

description$ U \sim \mathcal{N}(0,1), \ Z=X=\dot{X}$; • $U\mid X \sim \mathcal{C}\big(0,\ 0.1+|X_1|\big) $, $Z=X=\dot{X}$; • $X_1 = \dot{X}_1 + V$, $X_2=\dot{X}_2$, $Z = \big[\mathbbm{1} \{ |\dot{X}_1| < -\Phi^{-1}(0.25) \}, \, X_2 \big]$; • $X_1 = \mathbbm{1} \{V < - |\dot{X}_1| - \Phi^{-1}(0.25) \} $, $X_2=\dot{X}_2$, $Z = \dot{X}$; • $U\mid X \sim \mathcal{N}\big(0,\, (0.1+|X_1|)^{-2}\, \big) $, $Z=\dot{X}$, $ X_1 = \dot{X}_1 + \dot{U} $, $X_2 = \dot{X}_2$; • $\dot{Z} \sim \mathcal{N}(0,1) $, $X_1 = \dot{Z} + V$, $ Z = \dot{Z}^2 - a\dot{Z} $, $X_2=Z$; • $\Ddot{X} = \dot{X}/||\dot{X}|| $, $X_1 = \Ddot{X}_1 - aU$, $ Z = X_2 = \Ddot{X}_2 $; • $Z\sim \mathcal{N}(0,1)$, $X_1 = \dot{U}Z^2 - aU$, $X_2=Z$.

$ X:=[X_1, X_2] $ is exogenous in DGPs LM--0A and LM--0B, while $X_1$ is endogenous in the remaining DGPs.\footnote{Specifically, it is scale-endogenous in DGP LM--1C.} DGPs LM--1A, LM--1B, LM--2A, LM--2B, and LM--3 have non-monotone forms of relevance (see (ref)). A transformation of $Z$ in DGPs LM--1A and LM--2B that induces correlation (or mean-dependence) between $X_1$ and $Z$ is impossible. The identifying variation in LM--2B is implicit; $Z$ and the exogenous variation in $X_1$, namely $\Ddot{X}_1$ are defined on the unit circle, and one can only determine the other up to sign. Instrument relevance in LM--3 is in the “first-stage" skedastic function (see (ref)). There is MDep identification without excludability in LM--2A through LM--3 (see (ref)). In LM--1A, the excluded instrument is discrete; in LM--1B, the endogenous covariate is discrete; and in both cases, the first-stage relationships are non-monotone (see (ref)). Conditional scale heterogeneity in LM--0B and LM--1C does not violate (ref), and the first moment of $U$ in LM--0B does not exist (see (ref)).

table[table omitted — 4,732 chars of source]

For each of the DGPs, (ref) reports the median $t$-statistic (M-$t$), the median absolute deviation (MAD), the root mean squared error (RMSE), and the 5% rejection rate of the $t$-test of the null hypothesis $\theta_1 = 0.5$ across 1000 random samples with sample sizes \( n \in \{ 50, 100,200 \} \). Simulation results with larger samples and non-linear models are available in (ref) of the Supplemental Appendix. Competing estimators include the proposed MDep, conventional instrumental variables (IV) estimators---namely, two-stage least squares (TSLS) and Ordinary Least Squares (OLS)---as well as ICM estimators MMD and ESC6 of tsyawo2023feasible and escanciano2006consistent, respectively. Across all DGPs, the MDep exhibits stable fine-sample performance and clear robustness to weak or non-monotone instrument relevance, heavy-tailed distributions, heteroskedastic disturbances, and scale endogeneity in $U$ subject to (ref) or (ref). All estimators perform well in the baseline scenario LM--0A without endogeneity. However, in LM--0B, where the first moment of $U$ does not exist and its scale is heterogeneous in $X_1$, only the MDep estimator remains reliable---its bias and RMSE shrink steadily with the sample size---while all competing estimators exhibit explosive RMSEs and unreliable inference, underscoring MDep's robustness to infinite-variance disturbances.

In the weak, non-monotone, and discontinuous-covariate or instrument designs (LM--1A and LM--1B), MDep continues to dominate: its median bias and RMSE are modest and improve with $n$, whereas MMD, ESC6, and especially TSLS display severe finite-sample distortions and oversized rejection rates. Under conditional heteroskedasticity with scale endogeneity (LM--1C), all estimators improve markedly, but MDep achieves the lowest RMSE overall and the most stable rejection rates across sample sizes. Under endogeneity without excludability, where there is only one instrument for the two covariates (LM--2A, LM--2B, and LM--3), MDep again outperforms: its RMSEs remain small and converge rapidly, while competing estimators become erratic---showing extremely large RMSEs and severe over- or under-rejection. Overall, the simulations confirm that MDep provides accurate, numerically stable, and size-correct inference even in models featuring weak, non-monotone, or endogeneity without excludability, whereas the alternative estimators display unreliable behaviour under those conditions.

Conclusion

This paper introduces the MDep estimator, which weakens the relevance condition of conventional IV, ICM, CM, and non-parametric IV methods to stochastic dependence between non-trivial linear combinations of $X$ and $Z$. Thus, under the MDep framework, one can exploit the maximum number of relevant instruments possible in any given empirical setting, subject to either of two non-nested exogeneity conditions.

The MDep framework offers a fundamentally distinct and practically valuable approach to addressing several challenges: (1) the absence of excluded instruments, (2) the weak instrument problem, and (3) the non-existence or contamination of the disturbance term due to outliers or random noise with potentially undefined moments. Moreover, the use of bounded one-to-one transformations of \( Z \) obviates moment bounds on \( Z \), further enhancing robustness. Consistent estimation and reliable inference are feasible without excludability, provided endogenous covariates are non-linearly dependent (in the distributional sense) on exogenous covariates. The MDep handles the weak IV problem by admitting instruments of which endogenous covariates may be uncorrelated or even mean-independent but not independent.

Identification, consistency, and asymptotic normality hold in the MDep framework under mild regularity conditions. Moreover, the MDep covariance matrix estimator is shown to be consistent. To ensure the practical usefulness of the MDep, this paper shows the testability of the weak relevance condition. Illustrative examples backed by simulations showcase the remarkable properties of the MDep estimator vis-\`{a}-vis existing conventional IV and ICM methods.