Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,093 characters · 17 sections · 132 citation commands
Feasible IV Regression without Excluded Instruments
Instrumental variable (IV) methods play a pivotal role in the identification and inference on structural parameters in empirical work. When excluded instruments are unavailable or very weak, the usefulness of conventional IV methods is limited. Thanks to the weak relevance condition of Integrated Conditional Moment estimators (ICM hereafter), identification, consistency, and asymptotic normality are attainable even if excluded instruments are unavailable or very weak. These remarkable properties notwithstanding, multiplicative-kernel ICM estimators suffer diminished identification strength with attendant drawbacks, viz. large biases and severe size distortions even for a moderate but fixed number of instruments. In fact, the multiplicative structure of a kernel induces an artificial weak instrument problem in the presence of multiple instruments, irrespective of the true strength of non-parametric identification.\footnote{The term “kernel" in this paper refers to a user-fixed function as in $ U $-statistics -- see, e.g., lee1990u.} Moreover, in any given empirical context, the practical usefulness of the aforementioned features of ICM estimators crucially lies in the testability of the ICM relevance condition.
This paper makes a three-fold contribution in light of the foregoing. This paper (1) proposes and investigates the properties of a computationally fast linear ICM estimator which preserves identification strength in the presence of multiple instruments, (2) highlights a remarkable feature of ICM estimators, namely, consistent IV estimation without excluded instruments provided the endogenous covariates are non-linearly mean-dependent on exogenous covariates in a way that ought not to be known or modelled, and (3) extends the test of the relevance condition in escanciano2018simple from a single-covariate to a multiple-covariate setting. Although the artificial weak instrument problem is not new in the literature -- e.g., escanciano2006consistent -- this paper's characterisation of the problem is, nonetheless, interesting as it provides insight into the case of a moderately large but fixed number of instruments.
The Minimum Mean Dependence estimator (MMD hereafter) proposed in this paper minimises the mean dependence of an error term on a set of instruments using the Martingale Difference Divergence measure (MDD hereafter) of shao2014martingale. Thanks to this general notion of dependence, the MMD, like other estimators in the ICM-class, e.g., dominguez2004consistent, escanciano2006consistent, shin2008semiparametric, antoine2014conditional, escanciano2018simple, wang2018consistent, choi2021generalized, and antoine2022partially, weakens the IV-relevance condition from correlatedness to mean dependence and thus expands the set of relevant IVs in at least two significant ways. First, consistent IV estimation is feasible when there are no excluded instruments, provided endogenous covariates are non-linearly mean-dependent on exogenous covariates in a way that ought not to be known or modelled. This is a valid and non-spurious source of identifying variation as the structural economic model is generally agnostic about the dependence between endogenous and exogenous covariates. Second, as observed in, e.g., antoine2014conditional, escanciano2018simple, and antoine2022partially, endogenous covariates may be weakly correlated or even uncorrelated with but mean-dependent on instruments.
The above features of the ICM-class are not only potentially useful; they are also present in empirical settings. A specification without excluded instruments in the empirical example of this paper yields reasonable estimates and smaller standard errors relative to the IV. The p-value of ICM-relevance, $ 0.08 $, suggests a rejection of the null hypothesis of lack of non-parametric identification at the $ 10\% $ level, although no excluded instrument is used. In Table 1 of escanciano2018simple, the standard first-stage $ F $-statistics corresponding to the UK and the US are $ 2.52 $ and $ 2.93 $, respectively, which are indicative of very weak IVs judging by the stock2005testing rule-of-thumb of $ 10 $. The respective ICM-relevance p-values, 0.07 and 0.00, are indicative of ICM-strong instruments.
The MMD is identified, $ \sqrt{n} $-consistent, and asymptotically normal. Extensive Monte Carlo simulations in this paper suggest a favourably competitive performance of the MMD within and without the ICM class. This paper shows that the relevance condition of ICM estimators can be cast as an ICM specification test of, e.g., bierens1982consistent, bierens1990consistent, stute1997nonparametric, escanciano2006consistent, dominguez2015simple, su2017martingale, and jiangTsyawo2022aconsistent. The relevance test in this paper allows a multiple-covariate setting with one endogenous covariate; it thus generalises the single-covariate case of escanciano2018simple -- see Section 4.1 of the paper. A set of simulations in this paper with up to 32 relevant instruments underscore the drawbacks of multiplicative ICM kernels. With sample sizes up to 1000, multiplicative-kernel ICM estimators produce empirical sizes that appear to increase with the sample size. In fact, three out of four multiplicative-kernel ICM estimators have empirical sizes equal to 1 for a 5% level $ t $-test at sample sizes 500 and 1000. In contrast, the MMD and the ICM estimator of escanciano2006consistent, whose kernels are non-multiplicative and less sensitive to the dimension of the instrument vector, control size meaningfully and improve in performance with the sample size. The computational cost of the escanciano2006consistent estimator relative to the MMD, however, increases exponentially with the sample size.
The MMD complements the class of ICM estimators. ICM estimators extend a literature that is largely focussed on specification testing to estimation. The minimum distance estimator of wang2018consistent is the closest to the MMD as both are based on the same kernel. The focus differs, however, as unlike wang2018consistent which is concerned about the inconsistency problem highlighted in dominguez2004consistent, the current paper specialises to the linear model, provides a closed-form expression of the estimator, and investigates its theoretical properties in detail. Moreover, this paper, unlike wang2018consistent, highlights the adverse effects of multiplicative kernels in ICM estimators in the presence of a moderately sized instrument set. The current paper diverges from the aforementioned in at least three respects. First, available ICM estimators, save wang2018consistent and escanciano2006consistent, are based on multiplicative kernels. Multiplicative kernels are almost constant at zero and vary negligibly across observations for even moderate dimensions of the instrument -- see escanciano2022debiased for a related discussion.\footnote{This paper considers estimation with a finite set of covariates and instruments. Allowing the number of covariates, instruments, or both to grow with the sample size is beyond the scope of the current paper -- see antoine2022partially, escanciano2022debiased, and sun2022estimation for a discussion.} The MMD and escanciano2006consistent kernels, in contrast, retain meaningful variation even for a moderately large but finite number of instruments, as shown in the Monte Carlo simulations. The escanciano2006consistent estimator is, however, a factor $ O(n) $ more computationally costly than the MMD. A drawback of the MMD is that it requires moment bounds on the set of instruments, unlike bounded-kernel ICM estimators, e.g., dominguez2004consistent, escanciano2006consistent, antoine2014conditional, escanciano2018simple, and choi2021generalized. Second, the possibility of IV regression without excludability using ICM estimators does not appear to have been noticed before in the literature.\footnote{This generic claim is completely agnostic about the “first stage" and naturally holds for multiple endogenous covariates without excludability, unlike the 2017 working paper version of choi2021generalized.} Third, this paper provides a generally useful test of the ICM non-parametric identification condition in a multiple-covariate setting.
The rest of the paper is organised as follows. (ref) presents the MMD estimator, and (ref) collects theoretical results viz. consistency, asymptotic normality, consistency of the covariance matrix estimator, and the ICM relevance test. (ref) examines the small sample performance of the estimator and its covariance matrix estimator via simulations. An empirical example in (ref) illustrates the proposed estimator's practical usefulness, and (ref) concludes. All proofs and additional simulation results are relegated to the appendix and the online supplement.
\setcounter{equation}{0}
Consider the econometric model of interest: $ Y = X\theta_o + U $ where $ Y $ is the outcome, elements in the vector of covariates $ X \in \mathbb{R}^{p_x} $ may be endogenous, $ Z \in \mathbb{R}^{p_z} $ collects instrumental variables, $ E[U|Z] = 0 \ a.s. $, and $ \theta_o $ is the parameter vector of interest. $ X $ contains a constant term which ensures the error term $ U $ is zero-mean. The conventional IV uses the orthogonality condition $ E[UZ] = 0 $ which, as observed by dominguez2004consistent, fails to fully exploit the stated conditional moment restriction while an ICM estimator directly solves the conditional moment restriction $ E[U|Z] = 0 \ a.s. $ ICM measures convert the possibly continuum of orthogonality conditions implied by $ E[U|Z] = 0 \ a.s. $ to a scalar-valued objective function which is zero if and only if $ E[U|Z] = 0 \ a.s. $ (bierens1982consistent). As the proposed MMD is based on the Martingale Difference Divergence (MDD) ICM measure of shao2014martingale, it is instructive to briefly present the measure.
For a zero-mean $ W \in \mathbb{R} $ and a multivariate random variable $ Z \in \mathbb{R}^{p_z} $, the MDD is defined as the square root of \[ \mathrm{MDD}^2(W|Z) \equiv \frac{1}{c_{p_z}} \int \frac{|E[W\exp(\iota Zs)]|^2}{||s||^{1+p_z}}ds \] where $ c_{p_z} = \frac{\pi^{(1 + p_z)/2}}{\Gamma((1 + p_z)/2)} $, $ \Gamma(\cdot) $ is the complete gamma function $ \Gamma(p) \equiv \int_{0}^{\infty} t^{p-1}\exp(-t)dt $, $ \iota = \sqrt{-1} $ is the imaginary unit, and $ ||\cdot|| $ denotes the Euclidean norm. The MDD uses the non-integrable weight function $ w(s) = 1/(c_{p_z}||s||^{1+p_z}) $ proposed by szekely2007measuring.
The MDD belongs to the general class of ICM measures of mean dependence with the form
where different suitable choices of $ G(Z,s) $ and $ w(s) $ give rise to different ICM measures. For example, escanciano2018simple uses the complex exponential $ G(Z,s) = \exp(\iota Zs) $ with the standard normal probability density function as weight $ w(s) $ -- see dominguez2004consistent, escanciano2006consistent, bierens1990consistent, and kim2020robust for other examples. This paper's choice of the pair $ G(Z,s) = \exp(\iota Zs) $ and $ w(s) = 1/(c_{p_z}||s||^{1+p_z}) $ is borne out of a natural double requirement: tractability (i.e., without requiring numerical integration) and informativeness of the resulting ICM measure even for moderately large $ p_z $ -- the latter is discussed in (ref).
For ease of reference, the following theoretical results on the MDD measure are collected below as properties; see shao2014martingale for details. \setcounter{bean}{0}
The analytical expression of the MDD measure in Property (c) is preferred in setting up the objective function as it avoids working with the integral or derivative of the norm of a complex number. The expected value of the objective function is thus
where $ [Y^\dagger,X^\dagger,Z^\dagger] $ is an $ iid $ copy of $ [Y,X,Z] $. For a random variable indexed by $ i $ or $ (i,j) $, define $ E_n[\xi_i] \equiv \frac{1}{n}\sum_{i=1}^{n}\xi_i $ and $ E_n[\xi_{ij}] \equiv \frac{1}{n^2}\sum_{i=1}^{n}\sum_{j=1}^{n}\xi_{ij} $. The sample version of $ Q_o(\theta) $ is given by \[ Q_n(\theta) = - E_n[||Z_i - Z_j||(Y_i - X_i\theta)(Y_j - X_j\theta)]. \] The following result shows that the MMD is a linear IV estimator.
$ \hat{\theta}_n $ in (ref) is a linear IV-estimator -- see, e.g., wooldridge2010econometric -- with the constructed instrument $ \{h_n(Z_i), 1 \leq i \leq n\} $. It remains an IV estimator irrespective of the dimension of $ Z $ since $ ||Z_i - Z_j|| $ remains scalar-valued in the “over-identified" ($ p_z > p_x $), “just-identified" ($ p_z=p_x $), or “under-identified" ($ p_z < p_x $) case. Estimation thus proceeds by constructing $ \{h_n(Z_i), 1 \leq i \leq n\} $ and running a standard IV regression. It is shown in (ref) that inference follows through exactly as the standard IV. Unlike the ICM estimators of antoine2014conditional, escanciano2018simple, and choi2021generalized, where scale invariance is induced by scaling $ Z $, the MMD is not only naturally scale invariant but also invariant to affine transformations after rotation.\footnote{This is referred to as rigid motion invariance in the Statistics literature -- see, e.g., szekely2012uniqueness and shao2014martingale.} For all $ c, q \in \mathbb{R} $, $q\not = 0 $, orthonormal $ Q \in \mathbb{R}^{p_z\times p_z} $, and affine transformed $ Z $ after a rotation, i.e., $ \tilde{Z} \equiv c+qZQ $, $ \mathrm{MDD}^2(U|\tilde{Z}) = - E[||qZQ - qZ^\dagger Q||UU^\dagger] = |q|\mathrm{MDD}^2(U|Z) $ since $ ||ZQ-Z^\dagger Q|| = \sqrt{(Z-Z^\dagger )QQ'(Z-Z^\dagger )'}=||Z-Z^\dagger|| $.
The MMD belongs to the class of linear ICM estimators. Estimators in this class can be cast in the form (ref); they differ by the kernel $ K(Z_i,Z_j) $ which replaces $ ||Z_i-Z_j|| $ in (ref). Kernel functions result from the specific combination of $ G(Z,s) $ and $ w(s) $ in (ref). To see this, note that the generalised MDD measure (ref) is \[ \int| E[WG(Z,s)]|^2w(s)ds = E[WW^\dagger \int G(Z,s)\overline{G(Z^\dagger,s)}w(s)ds] = E[WW^\dagger K(Z,Z^\dagger)] \] where $ \overline{\xi} $ denotes the complex conjugate of $ \xi $.\footnote{$ \overline{G(Z^\dagger,s)} = G(Z^\dagger,s) $ for real-valued $ G(Z^\dagger,s) $.} Examples of kernels used in ICM estimators include $ K(Z_i,Z_j) = \exp(-0.5(Z_i-Z_j)\hat{V}_z^{-1}(Z_i-Z_j)')$ in the Integrated Instrumental Variable (IIV) estimator of escanciano2018simple where $ \hat{V}_z $ is the sample covariance of $ \{Z_i, 1\leq i \leq n\} $, $ K_n(Z_i,Z_j) = \frac{1}{n}\sum_{l=1}^{n} \mathrm{I}(Z_i \leq Z_l )\mathrm{I}(Z_j \leq Z_l) $ in dominguez2004consistent, and \[ K_n(Z_i,Z_j) = \frac{\pi^{(p_z/2)-1}}{\Gamma((p_z/2)+1)}\sum_{l=1}^{n} A_{ijl}^{(0)}/n \text{ where } A_{ijl}^{(0)} = \Big| \pi - \mathrm{arcos}\Big(\frac{(Z_i-Z_l)\cdot(Z_j-Z_l)}{||Z_i-Z_l||\times||Z_j-Z_l||}\Big) \Big|, \] and $ Z\cdot Z^\dagger $ denotes the dot product of $ Z $ and $ Z^\dagger $ in escanciano2006consistent.\footnote{The subscript $ n $ on the dominguez2004consistent and escanciano2006consistent kernels is meant to emphasise dependence on the sample.} On the estimator of escanciano2006consistent, the following computational simplifications apply: $ A_{ijl}^{(0)} = \pi $ if $ Z_l\not=Z_i = Z_j $, $ Z_l=Z_i\not=Z_j $, or $ Z_l=Z_j\not=Z_i $, and $ A_{ijl}^{(0)} = 2\pi $ if $ Z_l = Z_i = Z_j $ -- escanciano2006consistent. Let $ Y_i^* \equiv [Y_i, X_i] $ and $ \tilde{K}(Z_i-Z_j) $ be a symmetric and bounded density if $ i\not =j $ and zero otherwise. $ K(Z_i,Z_j) $ equals $ \tilde{K}(Z_i-Z_j) $ if $ i\not=j $ and $ -\tilde{\lambda} $ otherwise in the Weighted Minimum Distance (WMD) of antoine2014conditional, where $ \tilde{\lambda} $ equals the smallest eigen-value of $ ( E_n[{Y_i^*}'Y_i^*])^{-1} E_n[\tilde{K}(Z_i-Z_j){Y_i^*}'Y_j^*] $. Although fundamentally ICM, the antoine2014conditional kernel has an induced jackknifed limited information maximum likelihood (LIML)-like structure. Although the kernels of dominguez2004consistent and escanciano2006consistent (DL and ESC6 respectively hereafter) involve quite natural choices of weight $ w(s) $ such as the uniform density on the unit sphere, the empirical distribution of the data, or both, they come at an $ O(n^3) $ computational cost relative to the MMD or the IIV's $ O(n^2) $ for example. The aforementioned kernels, save the MMD's $ K(Z_i,Z_j) = -||Z_i-Z_j|| $ and ESC6's, are bounded multiplicative kernels. It is easily verified that the ESC6 kernel is also invariant to affine transformations after rotation.
As instruments $ Z $ enter the ICM measure (ref) through the kernel only, the variation of the kernel across varying values of $ Z $ is crucial for the measure to be informative of mean dependence. A useful characterisation of the dependence of the kernel on the dimension of $ Z $ can be achieved through appropriate means $ \zeta $, e.g., quadratic, geometric, or arithmetic, of measurable functions of $ [Z,Z^\dagger] $ as presented in (ref).\footnote{See the online supplement for details.} The characterisation is useful as it highlights the sensitivity of the kernel to $ p_z $.
The antoine2014conditional and escanciano2018simple kernels are not separately presented in (ref) as they are modifications of the Gaussian kernel $ K(Z,Z^\dagger)=\exp(-0.5||Z-Z^\dagger||^2) $. The almost sure limit of the DL kernel suggests that its characterisation depends on the joint distribution of $ Z $.\footnote{$ K_n(Z,Z^\dagger) \xrightarrow{a.s.} 1 - F_Z(Z\vee Z^\dagger) $ for the DL kernel where $ F_Z(\cdot) $ denotes the joint cumulative distribution function of $ Z $ and $ Z\vee Z^\dagger $ denotes the element-wise maxima of the vectors $ Z $ and $ Z^\dagger $. For example, the DL kernel with a multivariate normal $ Z $ has the Gaussian characterisation in (ref) with element-wise maxima $ Z\vee Z^\dagger $ as $ n\rightarrow \infty $ almost surely (a.s.). That of the ESC6 is expressed using the angular distance in kim2020robust.} In general, kernels with the multiplicative structure have the characterisation $ K(Z,Z^\dagger) \equiv \prod_{k=1}^{p_z} \varphi(Z_k,Z_k^\dagger) = \zeta_{m}^{p_z}, $ where $ \zeta_{m} \equiv \Big(\prod_{k=1}^{p_z} \varphi(Z_k,Z_k^\dagger)\Big)^{1/p_z} $ is the geometric mean of $ \{ \varphi(Z_k,Z_k^\dagger), \ 1\leq k \leq p_z \} $ and $ \varphi(Z_k,Z_k^\dagger) $ is a bounded function, e.g., the normal probability density function which gives a Gaussian kernel. When $ \varphi(Z_k,Z_k^\dagger) \in [0,1] $ for each $ k \in \{1,\ldots, p_z\} $ in particular, $ \mathrm{var}(K(Z,Z^\dagger)) \leq E[\prod_{k=1}^{p_z} \varphi(Z_k,Z_k^\dagger)^2] = E[\zeta_{m}^{2p_z}] $ can be very small even for moderate $ p_z $, say, $ p_z=18 $. Multiplicative kernels $ K(Z,Z^\dagger) $ tend to have negligible variation and thus result in an almost constant kernel problem even for a moderate dimension of $ Z $ as they are very sensitive to the dimension of $ Z $. The structure of the MMD and ESC6 kernels, in contrast, allows for an informative measure as both kernels are less sensitive to $ p_z $. To shed further light on the problem, note that by the Cauchy-Schwartz inequality, an ICM measure $ \mathrm{GMDD}^2(W|Z) $ satisfies
Irrespective of the strength of mean dependence, $ \mathrm{GMDD}^2(W|Z) $ can be approximately zero even for a moderate $ p_z $ as the variation of $ K(Z,Z^\dagger) $ is approximately zero. As the GMDD is a covariance (see (ref)) and the scale of $ K(Z,Z^\dagger) $ varies by kernel and instruments $ Z $, a correlation version of the GMDD (see, e.g., shao2014martingale) is unit- and scale-free and ensures comparability across different measures' informativeness of mean dependence.
Consider an illustrative example with $ Z \sim \mathcal{N}(0,I_{p_z}) $. The plot of the standard deviations of the kernels in (ref) clearly shows that an ICM measure using the Gaussian or DL kernel can be very small, irrespective of the actual strength of mean dependence.\footnote{For $ E[W^2]=1 $ and $ p_z = 18 $, $ \mathrm{GMDD}^2(W|Z) \leq (\mathrm{var}[K(Z,Z^\dagger)])^{1/2} \approx 7.1\times 10^{-4} $ for the Gaussian kernel while $ \mathrm{MDD}^2(W|Z) \leq (\mathrm{var}[K(Z,Z^\dagger)])^{1/2} \approx 0.99 $. See the online supplement for details.} In contrast, a measure based on the MMD kernel enjoys meaningful variation even for moderate $ p_z $. To illustrate the informativeness of the measure of mean dependence, (ref) plots the Generalised Martingale Difference Correlation (GMDC) -- a correlation analogue of the GMDD -- to ensure comparability of measures' informativeness across different kernels; see Section S2 of the online supplement for details.\footnote{The mean dependence of $ W $ on $ Z $ is modelled as $ E[W|Z] = \frac{1}{\sqrt{p_z}}\sum_{k=1}^{p_z}Z_k $, and the GMDCs are simulated.} One observes that for the same level of mean dependence of $ W $ on $ Z $, GMDCs based on multiplicative kernels are less informative of the strength of mean dependence relative to the MDD and ESC6. Moreover, notice that GMDCs of both the MDD and ESC6 coincide. This is not surprising as the characterisation of both kernels in (ref) suggests that the ratio of their GMDCs is approximately 1.
\setcounter{theorem}{0} \setcounter{equation}{0}
Regularity conditions for the MMD are grouped into two categories. The first category (Assumptions (ref) and (ref)) comprises sampling and dominance conditions.
The $ iid $ setting in (ref) allows for conditional heteroskedasticity $ \sigma^2(Z) \equiv E[U^2|Z] $ of unknown form. The boundedness assumptions are sufficient for the existence of the MDD measure (see shao2014martingale) and useful in establishing the asymptotic properties of the MMD estimator.
An unbounded MMD kernel comes at the cost of imposing a bound on the fourth moment of $ Z $ in (ref)(a). The MMD may therefore be sensitive to outliers in $ Z $. This is, however, no limitation of the MMD as $ Z $ can simply be replaced by an element-wise bounded one-to-one mapping such that $ Z $ and its mapping generate the same Euclidean Borel field, e.g., $ \mathrm{atan}(Z) $ -- see bierens1982consistent. The moment bound on $ ||Z|| $ is not necessary for bounded-kernel ICM estimators as the instruments $ Z $ enter the estimator only through the bounded kernel -- see, e.g., antoine2014conditional or kim2020robust. This relative advantage of bounded-kernel ICM estimators only applies to excluded instruments in $ Z $ as included instruments in $ X $ ought to obey the moment bounds imposed in (ref)(a) -- cf. escanciano2018simple. In effect, this advantage no longer holds when $ Z $ contains no excluded instrument. The foregoing also suggests that the MMD is not sensitive to outliers in $ Y $ or $ X $ any more than bounded-kernel ICM estimators. The bound on $ Z $ in (ref)(a) is standard in the IV context -- cf. hansen_2021econometrics. (ref)(b) places a standard condition on the error term $ U $, e.g., hansen2014instrumental and carrasco2015regularized.
The second set of regularity conditions is needed for identification; they are conventional IV-type identification conditions which ensure that the expected value of the objective function (ref) is uniquely minimised at $ \theta_o $. Define $ m(z) \equiv E[X|Z=z] $.
(ref)(a) and (ref)(b) are, respectively, ICM analogues of IV exclusion and relevance identification conditions.
(ref)(a) is a standard exogeneity condition for ICM estimators, e.g., antoine2014conditional and escanciano2018simple. Although stronger than the IV exclusion restriction ($ E[Z'U] = 0 $) when $ X=Z $ or $ Z $ contains excluded instruments, (ref)(a) can be viewed as weak in the sense that $ Z $ may simply include strictly exogenous elements in $ X $ without any excluded instruments provided (ref)(b) holds -- (ref)(b) can hold even for $ p_z<p_x $. (ref)(a) is typically tested using ICM specification tests.
(ref)(b) is an ICM relevance condition -- see, e.g., escanciano2018simple. It can be equivalently expressed as a linear completeness condition (LC hereafter) $ E[X|Z]\tau = 0 \implies \tau = 0 $ a.s. as termed by escanciano2018simple. (ref)(b) appears as an identification rank condition within the context of the non-parametric IV (donald2001choosing), as a Local Identifiability condition in the context of a non-linear ICM estimator (antoine2014conditional), and as a relevance condition in the context of a partially linear ICM estimator (antoine2022partially) -- see newey2003instrumental for a related discussion in the context of the fully non-parametric IV. The LC formulation of (ref)(b) is perhaps a more interesting interpretation of ICM relevance; it says no non-zero linear combination of $ X $ is mean-independent of $ Z $.\footnote{In the simple case of a univariate $ X $, (ref)(b) simply says $ X $ is mean-dependent on $ Z $. This is the case covered by the test in escanciano2018simple.} The LC test introduced in this paper draws on this insight. Although the completeness condition in the fully non-parametric context is not testable (canay2013testability), the LC condition of ICM estimators is testable (escanciano2018simple). Both ICM identification conditions in (ref) are hence testable. An analogous interpretation for the standard IV full rank (relevance) condition on $ E[Z'X] $, see, e.g., wooldridge2010econometric, is that no non-zero linear combination of $ X $ is uncorrelated with $ Z $. Since mean independence implies uncorrelatedness, one intuitively sees why (ref)(c) below holds. One may wonder how (ref)(b) translates into the context of the MMD and its specific kernel. This is addressed in (ref)(b) -- cf. escanciano2018simple.
Define the function $ h(z) \equiv E[||z-Z||X] $. The matrix $ - E[||Z-Z^{\dagger}||X'X^{\dagger}] $ can be shown to equal $ - E[h(Z)'X] $, and it is the expected value of $ - E_n[h_n(Z_i)'X_i] = - E_n[||Z_i-Z_j||X_j'X_i] $ in the MMD estimator (ref). Moreover, it is real, symmetric, and positive semi-definite. To see why it is positive semi-definite, observe that for all $ \uptau \in \mathcal{S}_{p_x} $ where $ \mathcal{S}_p \equiv \{ \uptau \in \mathbb{R}^p: ||\uptau|| = 1 \} $ denotes the space of vectors $ \uptau \in \mathbb{R}^p $ with unit Euclidean norm,
From the proof of (ref)(a), $ Q_o(\theta) - Q_o(\theta_o) = - (\theta-\theta_o)' E[h(Z)'X](\theta-\theta_o) $, thus $ - E[h(Z)'X] = - E[||Z-Z^{\dagger}||X'X^{\dagger}] $ informs the strength of identification. Since $ E[X\uptau|Z] = 0 $ is equivalent to $ E[(X_{-1}- E[X_{-1}])\uptau_{-1}|Z] = 0 $ for $ \uptau \in \mathcal{S}_{p_x} $, $ \uptau_{-1} \in \mathcal{S}_{p_x-1} $, and $ X_{-1}\in \mathbb{R}^{p_x-1} $ which is $ X $ with the constant term excluded, the condition $ p_x-1 = \mathrm{rank}\big( E[||Z-Z^{\dagger}||(X_{-1}- E[X_{-1}])'(X_{-1}^{\dagger}- E[X_{-1}])]\big) $ is equivalent to (ref)(b) in view of (ref)(b). From the foregoing, \[ \uptau_{-1}' E[K(Z-Z^{\dagger})(X_{-1}- E[X_{-1}])'(X_{-1}^{\dagger}- E[X_{-1}])]\uptau_{-1} = \mathrm{GMDD}^2(X_{-1}\uptau_{-1}|Z) \geq 0 \] whence \[ 0\leq \uptau_{-1}' E[K(Z,Z^{\dagger})||(X_{-1}- E[X_{-1}])'(X_{-1}^{\dagger}- E[X_{-1}])]\uptau_{-1} \leq \mathrm{var}[X_{-1}\uptau_{-1}]\sqrt{\mathrm{var}[K(Z,Z^{\dagger})]} \] by (ref) for any estimator in the ICM-class with a kernel $ K(Z,Z^\dagger) $. The above suggests the following unit- and scale-free measure of ICM identification strength: $ \mathrm{GMDC}(X_{-1}\uptau_{-1}^*|Z) $ where $ \uptau_{-1}^* $ is the eigen-vector associated with the smallest eigen-value of $ E[K(Z,Z^{\dagger})||(X_{-1}- E[X_{-1}])'(X_{-1}^{\dagger}- E[X_{-1}])] $. Under a standard moment boundedness condition on $ X $, e.g., (ref)(a), all eigen-values of $ E[K(Z,Z^{\dagger}) X'X^{\dagger}] $ can be very close to zero if $ \sqrt{\mathrm{var}[K(Z,Z^{\dagger})]} $ is. The almost constant kernel problem thus hurts identification in the context of ICM estimators since in the presence of multiple instruments, it is possible that $ \mathrm{GMDC}(X\uptau|Z) \approx 0 $ but $ E[X|Z]\uptau \not = 0 $ a.s. for some $ \uptau \in \mathcal{S}_{p_x} $. This implies the equivalence relationship in (ref)(b) is weakened for multiplicative-kernel ICM estimators, relative to the MMD and ESC6 for example, even for a moderate $ p_z $. In practice, the almost constant kernel problem induces an artificial weak instrument problem in multiplicative-kernel ICM estimators, irrespective of the actual strength of non-parametric ICM identification.
(ref)(c) shows that the LC condition (ref)(b) is weaker than the standard IV relevance condition. The weakening can be viewed along two dimensions. First, the dimension of $ Z $ can be less than that of $ X $, i.e., there can be fewer instruments than covariates. This is under-identification in the conventional IV setting.\footnote{This feature can be gleaned from dominguez2004consistent with $ X = [Z,Z^2] $ and univariate $ Z $. The authors, however, did not consider the general case in this paper where $ X $ may be an unknown function of $ Z $ or where some covariates in $ X $ are endogenous. sun2022estimation also relies on this feature to estimate the parameters on two endogenous covariates using only one excluded instrument.} An important implication of this is that a researcher who faces an endogeneity problem but lacks excluded instruments can still consistently estimate $ \theta_o $ provided the mean of endogenous covariates in $ X $ have some non-linear dependence on the exogenous covariates in $ X $. Take $ X = [D,Z] $, $ D = Z + Z^2 + V $, $ E[V|Z]=0 $, and $ Z \sim \mathcal{N}(0,1) $ as an illustrative example. There is no non-zero linear combination such that $ E[X\uptau|Z] = (\uptau_1+\uptau_2)Z + \uptau_1Z^2$ is not a function of $ Z $. (ref)(b) is thus satisfied without an excluded instrument. Second, endogenous covariates in $ X $ can be weakly correlated or uncorrelated with but mean-dependent on $ Z $ as observed in, e.g., antoine2014conditional, escanciano2018simple, and antoine2022partially.
Define $ A \equiv - E[h(Z)'X] $, $ \hat{A}_n \equiv -E_n [h_n(Z_i)'X_i] $, and $ B \equiv \mathrm{var}\big[\sqrt{n} E_n[h(Z_i)'U_i] \big] $. It is shown in the online supplement that $ \hat{A}_n \xrightarrow{a.s.} A $. Combining the expression of the MMD estimator (ref), the data generating process ((ref)), and the continuous mapping theorem, $ \hat{\theta}_n $ satisfies the following expansion:
It is shown in the online supplement that $ E_n[(h_n(Z_i) - h(Z_i))'U_i] = O_p(n^{-1}) $ thus, $ \sqrt{n} E_n[h_n(Z_i)'U_i] $ satisfies the following Bahadur expansion:
From (ref) and (ref), the MMD is asymptotically linear thus rendering the multivariate Lindeberg-L\'evy Central Limit Theorem applicable. The covariance matrix estimator comprises the sample analogue of $ A $ and a consistent estimator of $B $. Denote $ \hat{B}_n \equiv E_n[\hat{U}_i^2h_n(Z_i)'h_n(Z_i)] $ where $ \hat{U}_i \equiv Y_i - X_i\hat{\theta}_n $ is the MMD residual for observation $ i \in \{1,\ldots,n\} $. The covariance matrix estimator is given by $ \hat{A}_n^{-1}\hat{B}_n\hat{A}_n^{-1} $. The following theorem collects the asymptotic properties of the MMD viz. consistency, asymptotic normality, and consistency of the covariance matrix estimator.
Computationally, (ref)(c) shows that the only departure of the MMD (and the linear ICM-class cast in the IV formulation (ref)) from the standard IV is the construction of the instrument $ \{h_n(Z_i), 1 \leq i \leq n\} $; estimation and inference are the same. Existing IV routines are, therefore, applicable without further modification.
The significantly weaker LC condition ((ref)(b)) is helpful in addressing the problems of unavailable excluded instruments or potentially IV-weak but ICM-strong instruments. These advantages are only practically useful if the LC condition (ref)(b) is testable. To this end, this paper generalises the escanciano2018simple LC test to a multiple-covariate setting with a single endogenous covariate. This setting is quite prevalent in empirical practice.\footnote{The 2017 working paper version of choi2021generalized proposes a test of ICM relevance. The arguments of the authors rest on the structural model of interest and do not seem easily generalisable.} \footnote{For example, 101 out of 230 specifications (about 44%) surveyed in andrews2019weak (see page 732 of the paper) have a single endogenous covariate and a single excluded instrument. This means at least 101 of the specifications apply to this case since the number of excluded instruments in this case of the LC test is not limited to one; it can be zero, one, or more.} Partition $ X = [D, \tilde{X}] $ where $ \tilde{X} \in \mathbb{R}^{p_x-1} $ is a vector of exogenous covariates including the constant term and the univariate $ D $ is the endogenous covariate. Define $ \mathcal{E}^D(\eta) \equiv D - \tilde{X}\eta $. The following result provides justification for casting the test of (ref)(b) as an ICM specification test.
(ref) in conjunction with Property (b) shows that the test of (ref)(b) in a single-endogenous-covariate setting simply involves a linear ICM regression of $ D $ on $ \tilde{X} $ with $ Z $ as instruments, and an ICM specification test. Rejecting $ \mathbb{H}_o $ at a nominal level provides evidence of non-parametric identification, i.e., (ref)(b) holds. While the proposed LC test is simple and intuitive, it does not generalise to a multiple-endogenous-covariate setting in an obvious way. This is because the reasoning underpinning (ref) only justifies setting $ D $ to at least one endogenous covariate without identifying any particular one. This pursuit is therefore left for future work.
Although ICM estimators have emerged since dominguez2004consistent, no known study extensively compares their small sample performance. For instance, there is no small sample comparison of different ICM estimators in escanciano2006consistent, antoine2014conditional, or escanciano2018simple. In this regard, this section conducts extensive Monte Carlo simulations that examine the small sample performance of the MMD vis-\`a-vis existing ICM estimators. Following antoine2014conditional, estimators of the K-class are also included. In addition, this section uses simulations to examine the small sample effect of the almost constant kernel problem on ICM estimators.\footnote{See the online supplement for further simulation results.} ICM estimators considered include the Weighted Minimum Distance (WMD) estimator and its modification \`a la fuller1977some (WMDF) of antoine2014conditional, the minimum distance estimator of dominguez2004consistent (DL), the minimum distance estimator of escanciano2006consistent (ESC6), and the Integrated Instrumental Variable (IIV) estimator of escanciano2018simple. This section also considers the following estimators of the K-class: Two-Stage Least Squares (TSLS), the Jackknife Instrumental Variables Estimator (JIVE) of angrist1999jackknife, the Limited Information Maximum Likelihood (LIML) estimator, the jackknife LIML (HLIM) of hausman2012instrumental, and its modification \`a la fuller1977some HFUL.
Four different data-generating processes are considered. To examine the performance of estimators under the two features of ICM estimators, namely, IV regression without excluded instruments and IV regression with IV-weak but ICM-strong instruments, $ DGP_{0A} $, $ DGP_{0B} $, $ DGP_{1A} $, and $ DGP_{1B} $ are introduced. Define the function $ f_1(Z) \equiv \frac{2}{\sqrt{p_z}} \sum_{k=1}^{p_z}I\big(|Z_k| < - \Phi^{-1}(1/4)\big) $.
$ DGP_{0A} $ considers IV regression without excluded instruments with a single endogenous covariate, while $ DGP_{0B}$ extends $ DGP_{0A} $ to a two-endogenous-covariate setting. Estimators of the K-class are infeasible under $ DGP_{0A} $ and $ DGP_{0B} $ as they are not identified when $ p_z < p_x $. $ U $ and $ V $ are each standard normally distributed. $ \delta $ and $ \rho \equiv \mathrm{cov}[U,V] $ tune the strength of non-parametric identification and the degree of endogeneity respectively. For instance, $ \delta $ in $ DGP_{0A} $ tunes the non-linear mean dependence of $ D $ on $ Z $. In $ DGP_{1B} $, non-parametric identification is IV-irrelevant but ICM-relevant, and it depends on $ \delta $ since $ E[D|Z] = \sqrt{\delta} \frac{\mathrm{sin}(Z_1)\mathrm{sin}(Z_2)}{(1-\exp(-2))/4} \not = 0 $ a.s. for $ \delta \not = 0 $ but $ \mathrm{cov}[D,Z] = 0 $ for any $ \delta $. In $ DGP_{1A} $ (at $ \delta = 0 $) and $ DGP_{1B} $ (at $ \delta \not = 0 $), $ D $ is uncorrelated with but mean-dependent on $ Z $. For $ \delta \not = 0 $ in $ DGP_{1A} $, the linear projection of $ D $ on $ Z $ is largely dominated by its orthogonal complement; $ Z $ is thus IV-weak but ICM-strong in $ DGP_{1A} $ with $ \delta \in (0,1] $.\footnote{The first-stage $ F $-statistics are approximately $ 2.0 $ and $ 5.5 $ at $ \delta = 0.25 $ and $ \delta=0.5 $ respectively.} $ \Phi(\cdot) $ denotes the cumulative distribution function of the standard normal distribution. In all DGPs, the set of instruments is generated as $ Z \sim \mathcal{N}(0,\Omega) $, where the $ (k,l) $'th element of $ \Omega $ is $ \exp(-|k-l|) $. $ p_z = 2 $ in $ DGP_{1A} $ and $ DGP_{1B} $, while it is set to $ 1 $ in $ DGP_{0A} $, and $ DGP_{0B}$. $ \tilde{X} = Z_2 $ in both $ DGP_{1A} $ and $ DGP_{1B} $. $ \alpha_o=\beta_o=\gamma_o=1 $, and $ \rho = 0.5 $ in $ DGP_{0A}$ through $ DGP_{1B} $. 1000 Monte Carlo simulations are run for each specification.
(ref) present simulation results on $ DGP_{0A} $ through $ DGP_{1B} $ with sample size $ n=250 $.\footnote{Results for $ n=500 $ and other DGPs are available in the online supplement.} In each table, $ \delta $ is varied in order to examine the performance of estimators at different degrees of non-parametric identification. For each level of $ \delta $, the mean bias (MB), median absolute deviation (MAD), the root mean square error (RMSE), and the empirical rejection rate (Rej.) for a 5% $ t $-test of the null hypothesis $ \mathbb{H}_o:\ \beta = \beta_o $ are reported.
The results for $ DGP_{0A} $ and $ DGP_{0B} $ are presented in (ref). In these DGPs, all estimators in the ICM-class perform well even though there is endogeneity without an excluded instrument. Notice particularly that the MMD compares favourably with other estimators within the ICM-class in terms of MAD and RMSE. For example, the MMD is not dominated in terms of RMSE at all levels of $ \delta $ and in terms of MAD at $ \delta \in \{0.5,1.0\} $ for $ DGP_{0A} $ and $ DGP_{0B} $. Size control is also generally good across all ICM estimators.
From (ref), one notices in $ DGP_{1A} $ $ (\delta=0.0) $ and in $ DGP_{1B} $ that only estimators of the ICM-class have low bias and good size control as expected since $ Z $ is IV-irrelevant but ICM-relevant in these cases. The ESC6 and WMDF, in particular, are not dominated in terms of MAD and RMSE in $ DGP_{1A} $ and $ DGP_{1B} $, respectively. The ESC6 is, however, a little size-distorted at $ \delta=0.1 $ for $ DGP_{1B} $. Observe that even with $ \delta > 0 $ in $ DGP_{1A} $ where there is some linear dependence of $ D $ on $ Z $, the linear projection of $ D $ on $ Z $ is dominated by its complement, whence the poor performance of the K-class of estimators. In spite of the poor performance of all estimators in the K-class, the JIVE and HLIM surprisingly show good size control at $ \delta=0.5 $ for $ DGP_{1A} $.
In this subsection, the almost constant kernel problem within the ICM-class of estimators comes into focus. To this end, $ DGP_{4} $ is specified below with $ \rho = \mathrm{cov}[U,V]=0.5 $ and $ p_z \in \{8, 18, 32\} $ at $ n \in \{250, 500, 1000\} $. Quite importantly, $ DGP_{4} $ allows for several relevant instruments where each instrument contributes a small fraction to the total strength of identification:
For all $ p_z $ considered, notice from (ref) that MAD and RMSE are non-increasing in $ n $ for all estimators. The MMD and ESC6, specifically, have MAD and RMSE that are strictly decreasing in sample size. The performance of the MMD and ESC6 in terms of MAD and RMSE are quite indistinguishable. Only the MMD and ESC6 show meaningful improvement with sample size at $ p_z=32 $. For example, the MAD of the WMDF, DL, and IIV at $ p_z = 32 $ is practically unchanged across sample sizes, and doubling the sample size from $ 500 $ to $ 1000 $ leaves the RMSE of the DL at $ p_z=32 $ practically unchanged. At each sample size, the MAD and RMSE of the MMD and ESC6 are very stable across $ p_z $. In fact, $ \sqrt{n}\times $MAD for the MMD and ESC6 are approximately $ 0.5 $ across all $ n $ and $ p_z $ which is indicative of $ \sqrt{n} $-consistency. This is, however, not the case for the WMD, WMDF, DL, and IIV as $ \sqrt{n}\times $MAD is sensitive to $ p_z $ at all $ n $ and sensitive to $ n $ at $ p_z\in \{18,32\} $. As a case in point, $ \sqrt{n}\times $MAD of the WMDF, and IIV are increasing in $ n $ at $ p_z\in \{18, 32\} $. It, however, ought to be emphasised that these are small sample issues as this paper considers $ p_z $ fixed with respect to $ n $.
At $ (n, p_z) =(250, 32) $, all estimators suffer substantial size distortion, although that of the MMD and ESC6 is much less severe. The size distortion of the MMD and ESC6 declines with the sample size, whereas that of the other estimators (which have multiplicative kernels) worsens with sample size. In fact, the empirical sizes of the WMDF and IIV at ($ n\in \{500,1000\}, \ p_z = 18 $) and the WMDF, DL, and IIV at ($ n\in \{500,1000\}, \ p_z=32 $) equal 1. Even at $ p_z=8 $, the size distortion of the WMDF and IIV remains severe across sample sizes. The size distortion of the WMD is not as pronounced as that of the other ICM estimators with multiplicative kernels. This is perhaps attributable to the jackknifed LIML-like structure of its kernel which is intended to mitigate dispersion -- see antoine2014conditional. This small simulation exercise shows the adverse effect of the almost constant kernel problem on the quality of inference in the ICM-class of estimators with multiplicative kernels. The almost constant kernel problem appears to increase finite sample bias and induce severe size distortions in ICM estimators with multiplicative kernels. In comparison to results in (ref), one observes that multiplicative-kernel ICM estimators, namely, the WMD, DL, and IIV, perform reasonably well when the dimension of $ Z $ is small, say one or two. Their relative performance deteriorates for moderate dimensions of $ Z $.
This section presents an empirical example from hornung2014immigration which illustrates the practical usefulness of the MMD estimator. Quite importantly, the endogenous covariate is non-linearly mean-dependent on exogenous instruments and hence lends itself to IV regression without an excluded instrument using an ICM estimator. The motivation for using the MMD in an empirical application where an excluded instrument is available is to verify the reliability of the MMD without an excluded instrument in a real-data setting. Also, the sample size is small $ (n=150) $, the instruments are not very IV-strong, and the instrument set is moderately sized $ (p_z = 11) $.
hornung2014immigration seeks to identify the long-term impact of Huguenot skilled-worker migration on the productivity of textile manufactories in Prussia. The outcome variable is the value of a firm's output in a given town, and the covariate of interest is the population share of Huguenots in a town (Percent Huguenots). Other covariates include the number of workers, the number of looms, the value of material input, and regional and town-specific characteristics that can impact productivity -- see hornung2014immigration for details. As Huguenots are highly specialised in the textile industry, one can expect the population share of Huguenots (the endogenous covariate) to, for example, depend on the number of looms. As one cannot, a priori, rule out non-linearities in the dependence structure of covariates, IV without excludability is possible. Excluded instruments proposed by hornung2014immigration for Percent Huguenots include population losses during the Thirty Years' War and its interpolated version -- see hornung2014immigration for details.
(ref) presents the empirical results. The upper panel presents coefficient estimates with standard errors, the second presents p-values of the proposed LC test, the $ F $-statistic of kleibergen2006generalized rank test, and the third presents the specification test of su2017martingale. There are four specifications in (ref). The first compares the MMD to the OLS, the second uses population losses during the Thirty Years' War as the excluded instrument, the third uses its interpolated version, and the fourth runs MMD without instrumenting for the endogenous covariate.\footnote{MMD*(4) thus differs from MMD(1) by the exclusion (without replacement) of the endogenous covariate from the set of instruments.}
One observes that MMD estimates and standard errors are stable across specifications, while OLS/IV estimates show substantial variation. In the presence of excluded instruments (specifications (2) and (3)), MMD estimates are more precise than the IV. Interestingly, in the last column, where no excluded instrument is used, the MMD estimate is reasonably close to MMD estimates with excluded instruments and appears to be more precisely estimated than the IV and the MMD itself with excluded instruments. This shows that the unavailability of excluded instruments in this empirical example hurts neither the plausibility of the estimate nor its precision in a meaningful way.
This paper introduces a linear IV estimator to the ICM-class of estimators that minimises the mean dependence of an error term on a finite set of instruments. The proposed estimator, like that of escanciano2006consistent, is more robust to the almost constant kernel problem relative to estimators within the ICM-class with multiplicative kernels. This paper highlights a remarkable feature of the ICM-class that can address the challenge of unavailable excluded instruments. Given that the practical usefulness of ICM estimators in tackling the aforementioned challenge crucially lies in the testability of the LC condition, this paper shows that the LC test for ICM estimators can be cast as a standard ICM specification test. The estimator's closed-form expression makes it computationally fast to implement using available linear IV routines.
The type of estimator proposed in this paper enables the researcher to exploit unknown forms of identifying variation that cannot be exploited using conventional IV methods. Although this approach is not automatically applicable whenever excluded instruments are unavailable, nor is it a panacea for all forms of the weak IV problem, the testability of the LC condition makes the approach practically useful as a researcher is able to ascertain applicability in a given empirical context. Firstly, the empirical example in this paper shows an empirically relevant scenario where an ICM estimator still yields reliable estimates when the researcher lacks excluded instruments but likely faces an endogeneity problem. Secondly, the empirical example helps to demonstrate the robustness and reliability of inference that the MMD and the ICM-class as a whole offers. The simulation exercise offers insights into how ICM estimators can rescue a project from the otherwise hard-to-solve problem of unavailable excluded instruments or very IV-weak (but ICM-strong) instruments. It shows a favourably competitive performance of the proposed MMD relative to other estimators of the ICM- and K-classes, and calls attention to the severity of the almost constant kernel problem in multiplicative-kernel ICM-estimators. Although not entirely new to the literature, the characterisation of the almost constant kernel problem in this paper is, nonetheless, interesting as it sheds light on the case of a moderately large but fixed dimension of $ Z $.
This paper leaves a number of interesting avenues for future work, e.g., extending the current framework to non-linear models, models with non-smooth objective functions such as quantile regression, and multivariate outcomes. It will also be interesting to extend the LC test to a multiple-endogenous-covariate setting. Owing to the frequency of autocorrelation and clustering in empirical work, it will be interesting to extend the current $ iid $ framework for autocorrelation- and cluster-robust inference.
An earlier draft of this paper circulated under the title IV Regression with Possibly Uncorrelated Instruments. The author acknowledges very useful feedback from the Editor and two anonymous referees. This paper also benefited from the invaluable comments of Al-mouksit Akim, Firmin Ayivodji, Brantly Callaway, Weige Huang, Feiyu Jiang, Gilles Koumou, Abdul-Nasah Soale, Guy Tchuente, and participants at the 2021 Latin American Meeting of the Econometric Society (LAMES), 2022 North American Summer Meeting of the Econometric Society (NASMES), and the 2021/2022 Université Mohammed VI Polytechnique AIRESS Seminar Series.\\