EconBase
← Back to paper

An Identification and Dimensionality Robust Test for Instrumental Variables Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

245,363 characters · 44 sections · 146 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

An Identification-and Dimensionality-Robust Test for Instrumental Variables Models

\thispagestyle{empty}

abstractAbstract. Using modifications of Lindeberg's interpolation technique, I propose a new identification-robust test for the structural parameter in a heteroskedastic instrumental variables model. While my analysis allows the number of instruments to be much larger than the sample size, it does not require many instruments, making my test applicable in settings that have not been well studied. Instead, the proposed test statistic has a limiting chi-squared distribution so long as an auxiliary parameter can be consistently estimated. This is possible using machine learning methods even when the number of instruments is much larger than the sample size. To improve power, a simple combination with the sup-score statistic of BCCH-2012 is proposed. I point out that first-stage F-statistics calculated on LASSO selected variables may be misleading indicators of identification strength and demonstrate favorable performance of my proposed methods in both empirical data and simulation study. Keywords: Instrumental Variables, Weak Identification, High-Dimensional JEL Codes: C12, C36, C55

Introduction

When instruments are suspected to be weak, researchers may want to test hypotheses about structural parameters using testing procedures that are robust to identification strength. These procedures all rely on some conditions on the rate of growth of the number of instruments, \(d_z\), in relation to the sample size, \(n\). The initial identification robust tests developed in StaigerStock-1997, Moreira_2003, and Kleibergen-2005 are shown by Andrews-Stock-2007 to control size under heteroskedasticity when the number of instruments cubes is small relative to the sample size, \(d_z^3/n\to 0\). Meanwhile, recent and interesting “many-instrument” tests (crudu_mellace_sandor_2021, Mikusheva_Sun_2022, Matsushita_Otsu_2022, lim2022conditional) allow the number of instruments to be proportional to the sample size, \(d_z/n\to \varrho \in [0,1)\), but require that the number of instruments itself be large, \(d_z \to\infty\).

In practice, these conditions can be difficult to interpret and the variety of tests available under alternate regimes may make it difficult for the researcher to know which test, if any, should be applied in her exact setting. As examples, consider settings such as that of derenoncourt-2022, where \(d_z = 9\) and \(n = 130\), Paravisini-2014, where \(d_z = 10\) and \(n = 5{,}995\), and gilchrist-glassberg-2016 where \(d_z = 52\) and \(n = 1{,}671\). In all three cases, the number of instruments cubed, \(d_z^3 = 729\), \(1{,}000\) and \(140{,}608\), respectively, cannot be treated as negligible relative to the sample size. Indeed, all of these papers use post-LASSO estimates of the first-stage, suggesting concern about the large number of instruments by the authors. However as the number of instruments is only moderately large in all of these settings, asymptotic approximations that rely on \(d_z \to \infty\) may not accurately resemble the finite sample distribution of the test statistic. This can make size control of many-instrument tests questionable.\footnote{Mikusheva_Sun_2022 point out that, when errors are homoskedastic, under fixed \(d_z\) their proposed test statistic weakly converges to a distribution whose 95th quantile is close to that of the standard normal. However, it is unclear what the behavior of the test would be when errors are heteroskedastic.}

By comparison, the test proposed in this paper can be applied in any of the settings listed above as well as when the instruments are high-dimensional, \(d_z \gg n\), a setting for which there has been little progress on identification robust testing to this point. The main problem in these settings has been that the exact limiting behaviors of the regularized first-stage estimators used when \(d_z \gg n\) are difficult to analyze and typically unknown. Existing analyses sidestep this issue by assuming strong identification and exploiting a certain orthogonality, termed “Neyman Orthogonality” by CCDDHNR-2018, of the structural parameter estimate to the first-stage estimation error BCCH-2012. This approach is explicitly not applicable under weak identification where the first-stage estimation error is on the same or lesser order as the signal from the instruments and thus relevant to asymptotic analysis. As such, the few existing proposals for identification robust testing that allow \(d_z \gg n\) (BCCH-2012, gautier2021highdimensional, Mikusheva-2023-ManyTimeSeries) either fail to incorporate first-stage information or rely on sample splitting, both of which may reduce power.

To construct the test statistic I borrow an idea from Kleibergen_2002,Kleibergen-2005 and leverage a conditional slope parameter, which can be simply estimated using regularized methods even when \(d_z \gg n\), to partial out the structural error from the endogenous variable. I then use this partialled out version of the endogenous variable to construct first-stage estimates. The key idea is that, if variables in the model followed a jointly Gaussian distribution, we could exploit the fact that uncorrelated jointly Gaussian random variables are independent to show that these first-stage estimates are independent of the structural error under the null. Thus, even under weak identification where their behavior becomes relevant to the distribution of the proposed test statistic, one could easily derive the limiting \(\chi^2\) distribution of the proposed test statistic by conditioning on these estimates. Deriving the limiting distribution of the test statistic in the general model then reduces to showing that one can treat all observations as if they followed a jointly Gaussian distribution.

When the number of instruments is small, this can be justified in large samples through central limit and continuous mapping theorems. However, when the number of instruments is large these standard tools cannot be applied. Instead, my asymptotic analysis proceeds using modifications of Lindeberg's interpolation argument Lindeberg1922, which roughly proceeds by showing to be negligible the total change in distribution from a one-by-one replacement of each observation in the expression of the test statistic with a Gaussian version. The modifications of Lindeberg's original interpolation method are non-trivial and required to deal with the fact that derivatives of the test statistic with respect to individual observations may be unbounded generally, an issue arising from the “divide-by-zero” problem of weak identification ASS-2019.\footnote{Existing interpolation results require these derivatives to be bounded while in this setting they may not even have finite moments. Given this, these modified approaches may be of independent interest to a growing literature on direct Gaussian approximation techniques (Chatterjee-2006, cck2013, Pouzo-Quadratic, CMY-2020).} This interpolation argument requires some conditions on the first-stage estimates, in particular I require that the first-stage estimates take on a “jackknife-linear” form and suggest using a jackknife ridge regression in practice to allow \(d_z \gg n\). Interestingly, though, these first-stage estimates are not required to converge to limiting values which allows the researcher some flexibility if she wanted to use a different approach.

Through the Gaussian approximation result, I examine the power properties of my proposed testing procedure in local neighborhoods of the null. These local neighborhoods are characterized by a bounded local power index. In the case of a single endogenous variable I show that, under an additional regularity condition, the local power index diverging implies that the test is consistent. Unfortunately, the process of partialling out the structural error can introduce bias into the first-stage estimate under the alternative hypothesis. Against certain alternatives, this bias can be particularly pronounced and erase the first-stage signal from the instruments, a problem pointed out by Andrews-2016 in the context of the Kleibergen-2005 K-statistic. To address this, I propose a simple combination with the sup-score test of BCCH-2012, which can also allow \(d_z \gg n\). As with the Anderson-Rubin test, while the sup-score statistic does not incorporate first-stage information it does not face a power decline against particular alternatives.

Identification-robust testing procedures may be of particular interest in high-dimensional settings due to a lack of clarity on how to pretest for weak identification. Current empirical practice when using post-LASSO estimates of the first-stage appears to be to use standard t-test inference if the first stage F-statistic on the LASSO selected variables is larger than 10 (Paravisini-2014, gilchrist-glassberg-2016, derenoncourt-2022). Using a simple numerical demonstration, I argue that first stage F-statistics on LASSO selected variables may not be reliable indicators of identification strength. Given uncertainty about the strength of identification I apply the newly proposed testing procedures to the data of gilchrist-glassberg-2016 and generate weak instrument-robust confidence intervals for the effect of social spillovers on movie consumption. The newly proposed confidence intervals are smaller than those obtained by inverting the sup-score and many instrument tests across all specifications. To verify these results, I also generate confidence intervals in the data of AK-1991, again finding that the newly proposed methods generate tighter confidence bands than the many instruments methods. In (ref), I argue these improvements in power may be explained by my proposed test statistic's use of higher quality first-stage estimates and individual scores that are uncorrelated across observations.

Finally, I examine the applicability of the theoretical results in this paper through a simulation study. While existing tests seem to face size distortions in alternate regimes, the test based on my proposed test statistic has nearly exact size in a variety of settings. While this new test may have diminished power against certain alternatives, this deficiency is ameliorated through combination with the sup-score test. After combining the new test with the sup-score statistic, I find the newly proposed testing is also have favorable power properties compared to the many instruments and sup-score test in my simulation setup.

The outline of this paper is as follows. (ref) formally defines the model considered and introduces the jackknife K-statistic and it's combination with the sup-score test. (ref) demonstrates the usefuleness of the proposed testing procedures in two empirical applications while (ref) provides evidence from simulation study. (ref) provides an overview of the Gaussian approximation approach and characterizes the limiting behavior of the test statistic. (ref) uses this characterization to examine the power properties of the test and establishes the validity of the combination test. Proofs of the main results as well as the presentation of some auxiliary results are deferred to the Appendix.

\paragraph{Notation.} For any \(n \in \SN\) let \([n]\) denote the set \(\{1,\dots,n\}\). I work with a sequence of probability measures \(P_n\) on the data \(\{(y_i,x_i,z_i): i \in [n]\}\). To accommodate independent but not identically distributed observations, let \(\E_n[f_i] = n^{-1}\sum_{i=1}^n f_i\) denote the empirical expectation and \(\bar\E[f] = \E_n[\E[f_i]]\) denote the average expectation operator.

Model and Setup

Though the analysis below allows for exogenous regressors, to simplify the exposition I follow Mikusheva_Sun_2022 and assume that they have already been partialed out of both the outcome, \(y_i\), and the endogenous regressors, \(x_i\). As the controls are assumed to be of fixed dimension, this is without loss of generality.\footnote{For discussion refer to (ref).} Along with the first stage, the IV model can then be written as a system of simultaneous equations:

equation[equation omitted — 144 chars of source]

The researcher observes the outcome \(y_i \in \SR\), the endogenous variable \(x_i \in \SR^{d_x}\), and the instruments \(z_i \in \SR^{d_z}\) but neither the structural error \(\varepsilon_i \in \SR\) nor the first-stage errors \(v_i \in \SR^{d_x}\). The structural error is assumed to be conditional-mean independent of the instruments, \(\E[\varepsilon_i|z_i] = 0\). I denote \(\E[x_i|z_i]\) as \(\Pi_i \coloneqq \E[x_i|z_i]\) and make no assumptions about the functional form of the conditional expectation so the instruments are allowed to affect the endogenous variable in a nonlinear fashion.

The random variables \(\{(z_i,\varepsilon_i, v_i)\}_{i=1}^n\) are assumed to be independent across observations. Observations need not be identically distributed but the errors are assumed to have a common covariance structure conditional on the instruments \(z_i\): \[ \Var((\varepsilon_i, v_i)'|z_i) \coloneqq \Omega(z_i) =

pmatrix[pmatrix omitted — 104 chars of source]

\in \SR^{(1 + d_x) \times (1 + d_x)} \] As \(\Omega(z_i)\) is otherwise left unrestricted, the errors are allowed to be heteroskedastic. All results in this paper hold conditionally on a realization of the instruments \(\vz := (z_1',\dots, z_n') \in \SR^{n\,\times\,d_z}\) so from this point forth they are treated as fixed and all expectations can be understood as conditional on the instruments.

Under this setup, the researcher wishes to test a two-sided restriction on the structural parameter: \[ H_0: \beta = \beta_0\vsbox H_1: \beta \neq \beta_0 \] I am interested in constructing powerful tests for this null-alternate pair that are asymptotically valid under arbitrarily weak identification and with minimal restrictions on the number of instruments \(d_z\). To this end, define the null errors \(\varepsilon_i(\beta_0) \coloneqq y_i - x_i'\beta_0\). Using these, I construct a variable, \(r_i\), that is a “partialed-out” version of the endogenous variable satisfying \(\Cov(r_i,\eps_i(\beta_0))= 0\):

align*[align* omitted — 290 chars of source]

Each element of the nuisance parameter \(\rho(z_i)\), \(\rho_\ell(z_i)\) for \(\ell = 1,\dots,d_x\), can be interpreted as the (conditional) slope coefficient from a simple linear regression of \(x_{\ell i}\) on \(\eps_i(\beta_0)\). Thus, if \(\rho_\ell(\cdot)\) falls in some function class \(\Phi\) it can be estimated directly under \(H_0\) by solving empirical analogs of:\footnote{Under \(H_1\), \(\rho_\ell(z_i)\) can be estimated directly by solving empirical analogs of \(\rho_\ell(z_i) = \arg\min_{\phi \in \Phi} \E[(x_{\ell i} - \eta_i(\beta_0)\phi(z_i))^2]\) where \(\eta_i(\beta_0) = \eps_i(\beta_0) - \E[\eps_i(\beta_0)|z_i]\). This requires an initial estimate of \(\E[\eps_i(\beta_0)|z_i]\), however.} \[ \rho_\ell(z_i) = \arg\min_{\varphi\in\Phi} \bar\E[(x_{\ell i} - \eps_i(\beta_0)\varphi(z_i))^2]. \] While other estimators of \(\rho(\cdot)\) are, in principle, possible (see (ref)), I will focus on \(\ell_1\)-penalized/LASSO estimators. These estimators are consistent under the assumption that \(\rho(z_i)\) has an approximately sparse representation in some basis \(b(z_i) \coloneqq (b_1(z_i),\dots,b_{d_b}(z_i))' \in \SR^{d_b}\), that is \(\rho_\ell(z_i) = b(z_i)'\phi_\ell + \xi_{\ell i}\) where \(\xi_{\ell i}\) represents an approximation error that tends to zero with the sample size and \(\phi_\ell\) is sparse in the sense that many of its coefficients are zero. This allows for nesting of the low-dimensional case, where the number of instruments is fixed, and the high dimensional case, where the number of instruments is potentially much larger than the sample size, under a unified estimation procedure. Under homoskedasticity, \(\rho_\ell(z_i)\) is a constant function and thus has a spare representation in any basis that contains a constant term. In general, the aproximate sparsity assumption can either be interpreted as an assumption that there are only a few instruments that are important for explaining variation in the covariance matrix \(\Omega(z_i)\) or as an assumption that the function \(\rho(z_i)\) can be accurately approximated using only a smaller set of basis terms in \(b(z_i)\).

The parameter \(\phi_\ell\) can be estimated via LASSO:

equation[equation omitted — 169 chars of source]

or via post-LASSO, refitting an unpenalized version of (ref) using only the basis terms associated with nonzero coeffecients in the inital LASSO regression. The estimating procedure in (ref) is a simple \(\ell_1\)-penalized regression of \(x_{\ell i}\) against \(\eps_i(\beta_0)b(z_i)\). It can be easily implemented using out-of-the-box software available on most platforms. Under standard conditions, this leads to a consistent estimate of \(\rho_\ell(z_i)\) as long as the sparsity condition \(s^2\log^M(d_bn)/n \to 0\) where \(s\) is the number of nonzero elements of \(\phi_\ell\) and \(M\) is a positive constant that depends on the moment bounds imposed. The estimation procedure is discussed in more detail in (ref). With \(\hat\rho(z_i) \coloneqq b(z_i)'\hat\phi_\ell\), I construct the estimated version of \(r_{\ell i}\), \(\hat r_{\ell i} \coloneqq x_i - \hat\rho(z_i)\eps_i(\beta_0)\) for each \(\ell \in [d_x]\).

Test Statistic

The test statistic is based on an arbitrary jackknife-linear estimate of the first stage, \[ \widehat\Pi_{\ell i} = \sum_{j\neq i} h_{ij}\hat r_{\ell j},\;\; \ell \in [d_x] \] for some “hat” matrix \(H = [h_{ij}] \in \SR^{n\times n}\). The phrase “hat matrix” is borrowed from ordinary least squares (OLS) where the projection matrix, \(\vz(\vz'\vz)^{-1}\vz'\), is sometimes referred to as the hat matrix in the sense that \(\hat x = \vz(\vz'\vz)^{-1}\vz' x\). In practice, the hat matrix, \(H\), can be any matrix that depends only on \(\vz\). It is important to note that while \(\widehat\Pi_{\ell i}\) does not depend on \(\hat r_{\ell i}\), it may depend on \(z_i\) through the hat matrix \(H\). This gives the test power against alternatives where \(\E[\eps_i(\beta_0)z_i] \neq 0\). For technical reasons, I will assume that \(h_{ii} = 0\) for each \(i \in [n]\) so that \(\widehat\Pi_{\ell i}\) can be written as \(\widehat\Pi_{\ell i} = \sum_{j=1}^n h_{ij} r_{\ell j}\).

Formally, the only structure I require on the hat matrix \(H\) is a balanced-design condition described in (ref). However, for reasons explained in (ref) it may be optimal to introduce some regularization in estimating the first-stage models \(\widehat\Pi_{\ell i}\) so I suggest using a jackknife ridge regression procedure setting:

equation[equation omitted — 101 chars of source]

where \(\hat\pi_{(-i)}(\lambda)\) is the coefficient from a ridge regression of \(\hat r\) on \(\vz\), leaving out observation \(i\) and with penalty parameter set equal to \(\lambda\). \footnote{A ridge regression coefficient estimate from a regression of \(\tilde\vy \in \SR^n\) on \(\tilde\vx\in \SR^{k\times n}\) with penalty parameter \(\lambda \in \SR_+\) solves \(\hat\pi \in \arg\min_{\pi\in\SR^K} \|\tilde\vy - \pi'\tilde\vx\|_2^2 + \lambda\|\pi\|_2^2\). nyquist1988applications shows how the fitted values from a jackknife ridge regression can be calculated without having to recompute \(\hat\pi_{(-i)}\) for each observation. Angrist_Krueger_JIVE provides a similar analysis for jackknife OLS.} Following recommendations in harrell_2015 and lecture_notes_ridge, the penalty parameter \(\lambda^\star\) is set so that the effective degrees of freedom is no more than a fraction of the sample size: \[ \lambda^\star = \inf \{\lambda \geq 0: \trace(\vz(\vz'\vz + \lambda I_{d_z})^{-1}\vz') \leq n/5\} \] The jackknife ridge first-stage estimate has the benefit of being well defined even when the number of instruments is larger than the sample size. I stress, though, that the \(\widehat\Pi_{\ell i}\) estimators are not required to be consistent and the researcher may use any other hat matrices that she believes will lead to plausible first-stage estimates. Other possible choices of hat matrix include the jackknife OLS hat matrix of Angrist_Krueger_JIVE, the deleted diagonal projection matrix introduced in Chao-2012-JIVE and successfully used in KSS-2018,crudu_mellace_sandor_2021,Mikusheva_Sun_2022, and Matsushita_Otsu_2022, or hat matrices based on selecting instruments via some preliminary unsupervised technique such as principal component analysis (PCA). (ref) below discusses how the balanced-design condition may be verified for arbitrary choices of hat matrices.

For each \(i = 1,\dots,n\), define \(\widehat\Pi_i = (\widehat\Pi_{1i},\dots,\widehat\Pi_{d_xi})\in \SR^{d_x}\) and \(\widehat\Pi_{\eps i} = \eps_i(\beta_0)\widehat\Pi_i\). Collect these in the matrices

equation[equation omitted — 430 chars of source]

The jackknife K-statistic can then be defined

align[align omitted — 232 chars of source]

I will show that, under appropriate moment bounds and conditions on the hat matrix, \(H\), the limiting distribution of \(\JK(\beta_0)\) under \(H_0\) is \(\chi^2_{d_x}\). For exposition, I will largely focus on the case where \(d_x = 1\), in which case the form of the test statistic simplifies to \(\JK(\beta_0) = \big(\sum_{i=1}^n \eps_i(\beta_0)\widehat\Pi_i\big)^2/\sum_{i=1}^n \eps_i^2(\beta_0)\widehat\Pi_i^2\). The extension to \(d_x > 1\) is not immediate but is possible under strengthened moment conditions.

remark[] While use of first-stage estimates that are uncorrelated with the structural error is inspired by Kleibergen_2002,Kleibergen-2005, the form of the jackknife K-statistic is distinct from that of the original K-statistics. One major difference is in how both test statistics account for heteroskedasticity. The K-statistic of Kleibergen-2005 accounts for heteroskedastic errors using a \(d_z \times d_z\) matrix, which cannot be consistently estimated when \(d_z\) is large. In contrast, the jackknife K-statistic uses the heteroskedasticity robust variance estimate \((\widehat\Pi_\eps'\widehat\Pi_\eps)^{-1} \in \SR^{d_x\,\times\,d_x}\). Showing that these variance estimates can be used to account for heteroskedasticity is a feature of the direct Gaussian approximation approach. Under weak identification the distribution of the variance estimate is relevant to the distribution of the test-statistic. However, even when \(d_z \ll n\), the distribution of this variance estimate would be difficult to analyze using traditional central limit theorems as it is not a continuous function of a sample mean or even of a quadratic form.

Combination with Sup-Score Test

As will be discussed further in (ref), the test based on the \(\JK(\beta_0)\) statistic can have deficient power against certain alternatives. This loss of power is similar to that faced by tests based on the K-statistics of Kleibergen_2002,Kleibergen-2005 and derives from the fact that the process of partialling out the null errors, \(\eps(\beta_0)\), from the endgenous variables introduces bias into the first stage estimates, \(\widehat\Pi\), under the alternative hypothesis. Against certain alternatives, this bias can “erase” the first stage signal from the instruments.

The power deficiency in tests based on the K-statistic is typically adressed by combining the K-statistic with the Anderson-Rubin statistic based on a constructed conditioning variable. Prominent examples of such combinations include the celebrated conditional likelihood test of Moreira_2003, GMM-M test of Kleibergen-2005, and minimax regret tests of Andrews-2016. I take a similar approach here in combining the newly proposed tests with tests based on the sup-score statistic of BCCH-2012,

equation[equation omitted — 193 chars of source]

which have correct asymptotic size even when the instruments is much larger than the sample size, \(d_z \gg n\). A level \(\alpha \in (0,1)\) test based on the sup-score statistic rejects whenever \(S(\beta_0) > c^S_{1-\alpha}\) where, for \(e_1,\dots,e_n\) iid standard normal and generated independently of the data, \(c^S_{1-\alpha}\) is the simulated multiplier bootstrap critical value:\footnote{This conditional quantile can also be approximated using an empirical bootstrap procedure as demonstrated by deng2020beyond.} \[ c^S_{1-\alpha} \coloneqq (1-\alpha)\text{ quantile of } \sup_{1 \leq \ell \leq d_z} \bigg|\frac{\sum_{i=1}^n e_i \eps_i(\beta_0)z_{\ell i}}{(\sum_{i=1}^n z_{\ell i}^2)^{1/2}}\bigg|\text{ conditional on }\{(y_i,x_i,z_i)\}_{i=1}^n. \] As with the Anderson-Rubin test, tests based on the sup-score statistic may have suboptimal power properties in overidentified models as it does not incorporate first-stage information. However, the sup-score statistic does retain the benefit of directing power evenly in all directions, avoiding pitfalls of tests based on \(\JK(\beta_0)\) against certain alternatives.

The combination test decides which test to run based on an attempt to detect whether the alternative \(\beta\) is such that \(\E[\widehat\Pi_{\ell,i}^I] = 0\) for all \(i \in [n]\) and some \(\ell \in [d_x]\). When this happens, tests based on the \(\JK(\beta_0)\) statistic have trivial power against deviations in the \(\ell\textsuperscript{th}\) coordinate of \(\beta\), so in local neighborhoods of these values of \(\beta\), tests based on the sup-score statistic may be preferable. Detection of whether \(\E[\widehat\Pi_{\ell,i}^I] = 0\) is based on the conditioning statistic:

equation[equation omitted — 227 chars of source]

Under the assumption that \(\E[\widehat\Pi_i^I] = 0\) for all \(i \in [n]\), quantiles of the conditioning statistic can be simulated analogously to the sup-score critical value. For a new set of \(\{(e_{\ell1},\dots,e_{\ell n}): \ell \in [d_x]\}\) iid standard normal and generated independently of the data, and for any \(\theta\in (0,1)\), define the conditional quantile

equation[equation omitted — 286 chars of source]

The thresholding test decides which test to run by comparing the conditioning statistic \(C\) to a treshhold value \(\tau\),

equation[equation omitted — 255 chars of source]

In principle, the thresholding statistic has correct size for any (preset) choice of parameter \(\tau\). In practice, however, I find that setting \(\tau = c^C_{0.75}\) leads to a reasonable balance of power between local and distant alternatives.

Empirical Application

I apply the testing procedures proposed in this paper to the data of gilchrist-glassberg-2016, who examine the effect of social spillovers in movie ticket sales, and to the data of AK-1991, who examine the returns to education. In both studies the number of instruments, \(d_z = 52\) and \(d_z = 180\), respectively, cannot be treated as negligible relative to the sample size (\(n = 1{,}671\) and \(n = 329{,}509\), respectively). To deal with the large number of instruments, gilchrist-glassberg-2016 employ a post-LASSO estimate of the first stage. This strategy is also shown to work well in the data of AK-1991 by angrist2022machine. Using a simple simulation study, I demonstrate that the first-stage F-statistics on LASSO selected variables typically reported by authors can be misleading indicators of identification strength. When revisting the initial analyses using identification robust testing procedures, the confidence intervals constructed by inverting the tests proposed in (ref) are consistently narrower than those constructed from inverting the sup-score and many-instrument testing procedures.

Application to Social Spillovers in Movie Consumption

The gilchrist-glassberg-2016 sample consists of 1,671 opening weekend days between January 1, 2002 and January 1, 2012. For each opening weekend, the authors observe gross ticket sales for movies wide released in theaters in the United States with a run in theaters of at least six weeks.\footnote{An opening weekend day is a Friday, Saturday, or Sunday of opening weekend and a wide released movie is any movie that ever shows on 600 or more screens.} The data are obtained through Box Office Mojo, a subsidiary of the Internet Movie Database (IMDb).

The outcome variables of interest are gross ticket sales of movies that opened in a given weekend in the second through sixth weeks of their run, while the endogenous variable is the gross ticket sales of a movie in its opening weekend. To control for seasonal periodicity in both the supply of and demand for movies, a vector of date controls are included. Formally, the authors are interested in the parameters \(\beta_w\), \(w = 2,\dots,7\) from the linear IV model(s):

equation[equation omitted — 111 chars of source]

where, for \(w = 1,\dots,6\), \(\text{Sales}_{wi}^\perp\) represents gross national ticket sales, after the partialing out of date controls and a constant, \(7(w-1)\) days after day \(i\), of movies that opened on the opening weekend of \(i\). The variable \(\text{Sales}_{7i}^\perp = \sum_{w = 1}^6 \text{Sales}_{wi}^\perp\) denotes the cumulative national ticket sales in the second through sixth running weekends of movies who opened in weekend \(i\), after the partialing out of date controls and a constant. The parameter \(\beta_w\) represents the social spillover effect of strong opening weekend sales on sales in later weeks.

To instrument for sales on opening weekend the authors employ a vector of nationally aggregated weather measures. These weather measures include the proportion of movie theaters experiencing maximum temperatures in \(5^\circ\) Fahrenheit bins on the interval \([10^\circ, 100^\circ]\), the proportion of movie theaters experiencing precipitation levels in 0.25 inch per hour increments on the interval \([0,1.5]\), and the proportions of theaters experiencing any type of snow and of theaters experiencing any type of rain. Since unusually poor weather may cause people to substitute away from outdoor activities and into watching a movie, these measures provide a source of exogenous variation in opening weekend sales that can be used to identify the effect of social spillovers.

Putting together the nationally aggregated weather measures leaves gilchrist-glassberg-2016 with 48 linearly independent instrumental variables. \footnote{There are 52 instruments in total, but four linearly dependent ones are ignored in the following.} To handle the large number of instruments, the authors employ a post-LASSO estimate of the first stage BCCH-2012; they set the first-stage penalty parameter so that the number of instrument selected is one, two, or three. The resulting first-stage F-statistics using the selected instrument(s), 38.80, 25.86, and 20.95, respectively, seem to indicate strong identification. However, the first-stage F-statistic on the full set of instrumental variables is only 3.80. Moreover, since the LASSO objective is an \(\ell_1\) penalized version of the OLS loss, using the variables selected by LASSO may mechanically lead to higher F-statistics even if the underlying relationship between the instruments and the endogenous variables is weak.

(ref) provides evidence from a simple simulation experiment to demonstrate this. For the simulation experiment I generate an iid sample of size \(n = 1000\). For each \(i \in [n]\), I generate 10 mutually independent instruments \(Z_{ki} \sim N(0,1)\) for \(1 \leq k \leq 10\). The endogenous variable is generated to only have a weak relationship with the instruments, \(X_i = \frac{2}{\sqrt{n}}\sum_{\ell = 1}^{10} Z_{\ell i}+ v_i\), and the outcome is generated \(Y_i = X_i + \eps_i\) where \((\eps_i,v_i)\) are independent standard normals. From the initial set of 10 instrumental variables I generate an additional 55 technical instruments by squaring and taking all interactions between variables in the initial set. These generated instruments are correlated with the initial instruments but do not directly enter the first stage.

I then set the LASSO penalty so that only a certain number of instruments are chosen and report the resulting average first stage F-statistics and 95% confidence interval coverage over one thousand simulations. As a comparasion I also report the average first-stage F-statistics and 95% confidence interval coverage from the oracle estimator, which only uses the relevant 10 initial instruments. Despite the fact that the first-stage F-statstic on selected instruments is more than double the first-stage F-statistic using the oracle first stage estimator, the coverage rate of 95% confidence intervals based on LASSO selected instruments is significantly degraded compared to both the nominal coverage probability and the coverage probability using the oracle first-stage estimator.

table[table omitted — 733 chars of source]

Given a lack of clarity on the strength of identification, I seek to validate the results of gilchrist-glassberg-2016 using the weak identification testing procedures proposed in this paper. The setting is particularly suitable for weak IV testing using the jackknife K-statistic. With 48 instruments and a sample size of 1671, \(d_z^3 = 110{,}592 \gg n \), making the tests of Moreira_2003,Moreira_2009,Kleibergen-2005, and Andrews-2016 inapplicable. On the other hand, it is unclear whether asymptotic approximations based on \(d_z \to \infty\) will accurately describe the finite-sample distribution of test statistics with 48 instruments. Moreover, since fluctuations in movie theater attendance seem to be largely driven by either particularly cold or particularly hot weather (see Figure 4 in gilchrist-glassberg-2016), the nuisance parameter \(\rho(z_i)\) is plausibly approximately sparse.

I compare the 95% confidence intervals based on the jackknife K-statistic to those based on the jackknife Lagrange-Multiplier (JLM) statistic of Matsushita_Otsu_2022 and the sup-score statistic of BCCH-2012. Confidence intervals based on the jackknife AR statistic of Mikusheva_Sun_2022 are not reported as they were empty for all specifications, a result that could indicate misspecification of the linear model. Similarly, confidence intervals based on the thresholding statistic, implemented as recommended in (ref), are also not reported as they always align with the those of jackknife K-statistic. The narrower confidence bands of the jackknife K-statistic in specifications where the sup-score confidence interval is non-empty indicates higher power from the jackknife K-test in this setting, so the combination test suggesting its use may be expected. The jackknife K-statistic is implemented using the jackknife OLS hat matrix of Angrist_Krueger_JIVE, that is by setting \(\lambda = 0\) in (ref).

Tables (ref) reports the \(95\%\) confidence intervals for \(\beta_1,\dots,\beta_7\) generated by weak-instrument robust confidence intervals for three sets of instruments: the first is the initial set of 48 instruments in gilchrist-glassberg-2016, the second set includes only the temperature instruments for \(d_z = 36\), and the final includes the initial instruments as well as all interactions between the temperature instruments and the remaining instruments for \(d_z = 524\). For reference, I also provide point estimates and standard errors for \(\beta_1,\dots,\beta_7\) from gilchrist-glassberg-2016, Table 2. To facilitate comparison, these point estimates and standard errors come from a specification that uses all the instruments in the first stage of a 2SLS procedure.

Qualitatively, the results from the weak-instrument robust confidence intervals are similar to that of the author's original analysis; indeed the gilchrist-glassberg-2016 point estimates are always in the identification robust confidence intervals when using either the initial instrument set (\(d_z = 48\)) or the reduced instrument set (\(d_z = 36\)). However, when using the larger instrument set of \(d_z = 524\) we obtain confidence bands using the jackknife K-test that rule out the author's initial point estimates for the parameters \(\beta_5, \beta_6\) and \(\beta_7\) and suggest somewhat smaller social spillover effects in movie consumption. Across all specifications and instrument sets, confidence intervals obtained from inverting the jackknife K-test are consistently smaller than those obtained from inverting both the sup-score and jackknife Lagrange Multiplier test. These reductions in confidence interval length are most noticeable when using the complete set of interactions, across all parameters the jackknife K-confidence intervals are nearly half the length of their jackknife Lagrange Multiplier counterparts. As with the jackknife Anderson-Rubin test, the sup-score confidence intervals are also often empty which again could suggest misspecification of the linear IV model.

table[table omitted — 5,287 chars of source]

Application to Returns to Education

For additional comparison, I revisit the setting of AK-1991, who study the effect of education on log-wages, instrumenting for education using various combinations of quarter-of-birth (QOB), year-of-birth (YOB), and place-of-birth indicators (POB). This dataset has the benefit of being used in the empirical studies of both Mikusheva_Sun_2022 and Matsushita_Otsu_2022, facilitating an easy comparison of the performance of the newly proposed testing procedure to these existing “many-instrument” methods. Moreover, angrist2022machine report improved power in this dataset when using post-LASSO estimates in the first stage, suggesting that high-dimensional or machine-learning techniques can be useful in this setting.

(ref) displays the resulting identification robust confidence intervals using two set of instruments. The first set, which contains all QOB \(\times\) YOB and all QOB \(\times\) POB interactions for \(d_z = 180\), corresponds to the specification in Table VII of AK-1991. The second set of instruments, \(d_z = 1{,}530\), contains all QOB \(\times\) YOB \(\times\) POB interactions. The confidence intervals for the jackknife Anderson-Rubin (JAR) and jackknife Lagrange-Multiplier (JLM) tests are taken directly from Mikusheva_Sun_2022 and Matsushita_Otsu_2022, respectively. To implement the jackknife K-test I follow a similar procedure to that of the prior empirical exercise, setting \(\lambda = 0\) when constructing the \(\widehat\Pi_i\) values. However, for computational reasons I opt to split the data into 11 pieces when estimating \(\widehat\Pi_i\) rather than using a true “leave-one-out” jackknife approach. Confidence intervals from the combination test are not reported as the sup-score confidence interval is empty in both specifications.

The results in (ref) are similar to those in (ref). In both specifications the confidence intervals obtained from inverting the jackknife K-test are substantially smaller than those obtained by inverting the “many-instrument” JAR and JLM tests. Indeed, as before, the confidence intervals obtained from inverting the jackknife K-statistic are nearly half the length of those obtained from inverting the JLM test and less than a quarter the length of those obtained from inverting the JAR test. It is notable that the range of plausible returns to education values implied by the jackknife K-confidence interval with \(d_z = 1{,}530\) are strictly below that implied by the jackknife-K confidence interval with \(d_z = 180\). This may be related to misspecification of the linear model as indicated by the empty sup-score confidence interval. “Downward bias” of the confidence intervals when using the larger instrument set is also seen for the JAR and JLM confidence intervals, though to a lesser extent. (ref), below, provides reasoning for why confidence intervals based on the \(\JK(\beta_0)\) statistic may be tighter than those based on the JAR and JLM statistics.

table[table omitted — 1,008 chars of source]

As always, these are just particular empirical examples, and my results should not be interpreted as a critique of BCCH-2012, Mikusheva_Sun_2022, or Matsushita_Otsu_2022, whose prior work I relied upon and was inspired by.

Simulation Study

In a simple simulation study, I examine the performance of tests based on the \(\JK(\beta_0)\) statistic and compare it with that of other tests that may be used in settings where the number of instruments is nonnegligible as a fraction of sample size. I consider a reduced-form data-generating process (DGP) similar to that of Matsushita_Otsu_2022. The outcome variable, \(y_i\), and endogenous variable, \(x_i\), are generated according to

equation[equation omitted — 104 chars of source]

where \(\Pi_i = \frac{1}{r_n}\sum_{k=1}^5 \frac{3}{4}\bar z_{ki} + \frac{1}{4}\bar z_{ki}^2 + \frac{1}{4}\bar z_{ki}^3\) is a (dense) transformation of an initial set of instruments \(\bar z_i \in \SR^{15}\) generated as described below. The value of \(r_n\) varies depending on the strength of identification considered; under weak identification, \(r_n = n^{-1/2}\) while for power curves I consider an intermediate identification strength, \(r_n = n^{-1/3}\).\footnote{Intermediate identification strength is considered to let the power curves come up to one at the boundaries of the considered range of \(\beta\) values. Power curves with \(r_n = n^{-1/2}\) look similar, but shrunk towards zero.} To model heteroskedasticity, the errors \((e_i,v_i)\) are generated \(\eps_i = (1 + \varrho_1(\bar z_{1i}^2 + \bar z_{2i}^2 + \bar z_{2i}\bar z_{3i}))e_{1i}\), and \(v_i = \varrho_2(1 + \bar z_{1i})\eps_i + (1 - \varrho_2)^2e_{2i}\) where \(e_{1i}\) and \(e_{2i}\) are generated independently of each other and other variables in the model according to a Laplace distribution with location parameter \(\mu = 0\) and scale parameter \(b = 1\). Since jackknife K-statistic has an exact \(\chi^2\) distribution when the errors are jointly Gaussian and \(\rho(z_i)\) is known, I purposefully avoid normally distributed errors to investigate the quality of asymptotic approximations. The parameters \(\varrho_1\) and \(\varrho_2\) control the degree of heteroskedasticity and endogeneity, respectively.

table[table omitted — 3,006 chars of source]

I examine the size of the test under three different instrument regimes. In all three regimes, I begin with an initial set of instruments \(\bar z_i = (\bar z_{1i},\dots,\bar z_{15i})'\) generated independently across indices according to a multivariate Gaussian distribution with Toeplitz covariance structure, \(\Cov(\bar z_{\ell i},\bar z_{ki}) = 2^{-|\ell-k|}\). In the first regime, the instruments only include the initial set, \(z_i = \bar z_i\) so that \(d_z = 15\). In the second regime, the full set of instruments \(z_i\) additionally includes all quadratic and cubic terms, \((z_{\ell i}^2,z_{\ell i}^3)\), \(\ell = 1,\dots,15\) so that in total \(d_z = 45\). In the third regime, the full set of instrument includes the initial set of instruments, \(\bar z_i\), all quadratic and cubic terms (30 additional terms) and interactions of the initial set of instruments (\(\binom{15}{2} = 105\) additional terms), so that in total \(d_z = 150\). Under each regime, the full set of instruments is passed to the test statistics with no indication about which instruments correspond to the initial set, and thus no indication about which instruments are relevant to the DGP.

figure[figure omitted — 463 chars of source]

In constructing the jackknife K-statistic, I opt to use the deleted diagonal ridge matrix, \(H = [h_{ij}]\) where \(h_{ij} = [\vz(\vz'\vz + \lambda I_{d_z})\vz']_{ij}\bm{1}\{i \neq j\}\), instead of a true jackknife ridge procedure in order to make the simulations compuationally tractable. Following reccomendations in lecture_notes_ridge, the penalty paramter \(\lambda\) is set so that effective degrees of freedom of the resulting hat matrix is no more than \(n/5\).\footnote{Precisely, the penalty parameter is set \(\lambda = \max{0, (n/5)^{-1}d_1^2(\vz)}(d_z - n/5)\), where \(d_1^2(\vz)\) is the square of the maximum singular value of the design matrix \(\vz\).} To estimate the parameter \(\rho(z_i)\), I implement the default cross-validated \(\ell_1\)-penalized procedure of (ref) via the glmnet package in R glmnet. I use the full vector of instruments as the basis to approximate \(\rho(z_i)\).

I compare the simulated size of the jackknife K test and to the performance of the sup-score test, \(S(\beta_0)\), of BCCH-2012, the thresholding test introduced in (ref), the standard Anderson-Rubin (A.Rbn.) test of Anderson-Rubin-1949 and StaigerStock-1997, the jackknife AR test (JAR) of crudu_mellace_sandor_2021 and Mikusheva_Sun_2022, and the jackknife LM test (JLM) of Matsushita_Otsu_2022. Critical values of the sup-score and conditioning statistic are simulated with the procedures described in (ref) with 1000 bootstrap replications. For the combination test cutoff, I report results using two different quantiles of the conditioning statistic under the assumption that \(\E[\widehat\Pi_i^I] = 0\) for all \(i \in [n]\); \(\tau_{0.3}\) corresponding to the 30\textsuperscript{th} quantile and \(\tau_{0.75}\) corresponding to the 75\textsuperscript{th} quantile.

(ref) reports the simulated size for all tests under weak and strong identification, respectively. One can see that the \(\JK(\beta_0)\) statistic has nearly exact size in almost all the setups considered. In contrast, both the jackknife AR and jackknife LM test all exhibit moderate size distortions in various regimes. The jackknife AR test in particular appears to overreject in nearly all setups considered. This property also observed in the simulation study of Matsushita_Otsu_2022 and so may be driven by the similarity of this simulation setup to theirs. The jackknife LM statistic appears to have good size properties when \(d_z = 10\) and \(d_z = 150\) but is consistently (though only moderately) undersized in the intermediate regime where \(d_z = 45\). This size distortion does not improve when moving from \(n = 200\) to \(n = 500\), suggesting that the requirement that \(d_z \to \infty\) is important for the quality of finite-sample approximation by its limiting distribution. Though the good performance of the jackknife LM statistic when \(d_z = 10\) is notable, it should also be remarked that this is the setup with the least amount of correlation between the instruments.

The sup-score and Anderson Rubin test seem to be undersized in all regimes, with the Anderson-Rubin test nearly never rejecting when \(d_z = 150\). However, and in line with the theory, both of their size properties appear to improve when increasing the sample size from \(n = 200\) to \(n = 500\). It is possible that the size properties of the sup-score test could be improved by using an empirical bootstrap based approach to simulate the critical value, as proposed by deng2020beyond, however I do not consider that approach here. The threshholding test appears to inherit the conservative nature of the sup-score test, though to a lesser degree due to the combination with the \(\JK(\beta_0)\) test.

figure[figure omitted — 457 chars of source]

In addition to examining the size properties of the given tests I also investigate the power properties of the tests in this setting. (ref) plot calibrated local power curves under an intermediate identification strength where the first stage is in a \(n^{-1/3}\) neighborhood of zero for \(n \in \{200, 500\}\), the number of instrument is plausible large, \(d_z = 150\), \(\varrho_1 \in \{0.2,0.5\}\) and \(\varrho_2 \in \{0.3,0.6\}\). The critical value of each test is set to simulated 95\textsuperscript{th} quantile of the distribution of the corresponding test-statistic under \(H_0\). I compare the calibrated local power curves of the \(\JK(\beta_0)\) test, the combination test with cutoff \(\tau_{0.75}\), the jackknife AR test, the Jackknife LM test, and the sup-score test.\footnote{For both the jackknife AR and jackknife LM tests, I use cross-fit estimates of test statistic variances proposed and shown to improve power by Mikusheva_Sun_2022.}

In all setups considered, the jackknife K and combination tests have substantially stronger power than the jackknife AR, jackknife LM, and sup-score tests in local neighborhoods of the null as well as for negative values of \((\beta_0 - \beta)\). For values of \((\beta_0 - \beta)\) larger than 1.5, tests based on the jackknife K-statistic suffer from a loss of power as described in (ref). This power decline is largerly ameliorated by combining the jackknife K-statistic with the sup-score statistic and the thresholding test has good power properties over all alternatives considered. However, tests based on the jackknife AR or jackknife LM statistic can still provide better power than the threshholding test for very positive values of \((\beta_0 - \beta)\) in some setups.

Limiting Behavior of the Test Statistic

The limiting behavior of the test statistic is analyzed via a direct Gaussian approximation technique. In this section, I describe the approach and characterize the limiting behavior of the test statistic under local alternatives to \(H_0\). This direct approach has the advantage of not relying on any particular central limit theorem, which allows a great deal of flexibility in the choice of hat matrix \(H\) used to construct the first stage estimates. When there is only a single endogenous variable, \(d_x = 1\), the approach can be considerably simplified. The extension to \(d_x > 1\) requires a more involved argument which relies on stronger moment conditions.

Formally, I show that quantiles of the jackknife K-statistic can be approximated by analogous quantiles of the Gaussian statistic:

equation[equation omitted — 169 chars of source]

where, for each \(i \in [n]\), \((\tilde\eps_i(\beta_0),\tilde r_i')'\) are generated independently across indices following a Gaussian distribution with the same mean and covariance matrix as \((\eps_i(\beta_0),r_i')'\). Further, for \(\tilde\Pi_{\ell i} = \sum_{j\neq i} h_{ij} \tilde r_{\ell j}\) define \(\tilde\Pi_i \coloneqq (\tilde\Pi_{1i},\dots,\tilde\Pi_{d_xi})'\in\SR^{d_x}\), \(\tilde\Pi_{\eps i} \coloneqq (\E[\eps_i^2(\beta_0)])^{1/2}\tilde\Pi_i\),

align*[align* omitted — 314 chars of source]

As uncorrelated jointly Gaussian random variables are independent, under \(H_0\) the vector \(\tilde\eps(\beta_0)\) is mean zero and independent of \((\tilde\Pi, \tilde\Pi_\eps)\). Conditional on any realization of \((\tilde\Pi,\tilde\Pi_\eps)\) the \(\JK_G(\beta_0)\) statistic then follows a \(\chi^2_{d_x}\) distribution, and thus, its unconditional distribution is also \(\chi^2_{d_x}\).

To outline the approach, consider functions \(\varphi_\gamma(\cdot) \in C_b^3(\SR)\) that approximate the indicators \(\bm{1}\{\cdot \leq a\}\), where \(a \in \SR\) is arbitrary and as \(\gamma \to 0\) the quality of the approximation improves but the derivatives of \(\varphi_\gamma\) become larger in magnitutde. A primary goal is to show, for a sequence \(\gamma_n\) tending to zero, that

equation[equation omitted — 130 chars of source]

for a version of the test statistic, \(\JK_I(\beta_0)\), that could be constructed if \(\rho(\cdot)\) was known to the researcher. In order to establish (ref), existing interpolation methods cannot be applied as they require the derivative of the test statistic with respect to individual obserations are bounded. In this setting, the derivative of the test statistics with respect to terms in the denominator matrix, \(\widehat\Pi_\eps'\widehat\Pi_\eps\), may be as large as the inverse of the minimum eigenvalue of the denominator matrix. When identification is sufficiently weak, the eigenvalues of the denominator matrix can be arbitrarily close to zero and thus inverse of its minimum eigenvalue may not even have finite moments.

To get around this, I modify the argument by considering a “data-dependent” choice of approximation parameter \(\gamma_n\). This choice of approximation parameter inversely scales with the determinant of the denominator matrix and thus, since the determinant is the product of the eigenvalues, inversely scales with the minimum eigenvalue.\footnote{The determinant has the benefit of being a smooth function of elements of the matrix. This makes it nicer to work with than the minimum eigenvalue itself, which loses differentiability when the dimension of its eigenspace is larger than one.} Geometrically, this approach can be thought of as “stretching out” the function \(\varphi_{\gamma_n}(\cdot)\) in directions where the minimum eigenvalue of the denominator matrix is close to zero. Through the chain rule, this allows for control of the overall derivative of \(\varphi_{\gamma_n}(\JK_I(\beta_0))\) with respective to an individual observation.

High Level Assumptions

I now detail the assumptions needed for the argument, starting with the assumptions that are commmon to the cases with \(d_x = 1\) and \(d_x > 1\). In what comes below \(c > 1\) can be considered an arbitrary constant that may be updated upon each use but that does not depend on sample size \(n\).

assumption[Balanced Design] (i) Let \(s_{\ell, n}^{-2} = \max_{1 \leq i \leq n}\E[(\widehat\Pi_{\ell i}^I)^2]\) for each \(\ell \in [d_x]\); then, the minimum eigenvalue of the following matrix is bounded away from zero: \[ c^{-1} \leq \lambda_{\min}\E\begin{pmatrix} \frac{s_{\ell,n}s_{k,n}}{n}\sum_{i=1}^n (\widehat\Pi_{\ell i}^I)(\widehat\Pi_{k i}^I) \end{pmatrix}_{\substack{1 \leq \ell \leq d_x \\ 1 \leq k \leq d_x}} \] (ii) \(\max_i s_n \sum_{j\neq i} h_{ji}^2 \leq c\); and (iii) the following ratio is bounded away from zero: \(\frac{\sum_{k=2}^n\lambda_k^2(HH')}{\sum_{k=1}^n\lambda_k^2(HH')} \geq c^{-1}\) where \(\lambda_k(HH')\) represents the \(k^{\text{th}}\) largest eigenvalue of the matrix \(HH'\).

(ref)(i) requires that the average second moment of the infeasible first-stage estimators be on the same order as the maximum first-stage estimator second moment. This is imposed mainly to rule out hat matrices that are all zeroes or nearly all zeros so that the effective number of observations used to test the null is growing with the sample size. (ref) provides further intuition for this assumption and below discusses how this assumption and (ref)(ii) may be verified in practice. (ref) compares this balanced design assumption to that in the many-instruments literature crudu_mellace_sandor_2021,Mikusheva_Sun_2022,Matsushita_Otsu_2022,lim2022conditional, noting that their balanced design neither implies nor is implied by the one in this paper.

(ref)(ii) requires that the maximum leverage of any observation be bounded. When \(H\) is symmetric, it is automatically satisfied.\footnote{To see this for \(d_x = 1\) notice that \(s_n^{-2} = \max_i\E[(\widehat\Pi_i^I)^2] \geq \max_i \Var(\widehat\Pi_i^I) = \max_i \sum_{j\neq i} h_{ij}^2 \Var(r_j)\), while \(\Var(r_j)\) will be assumed bounded from below by \(c^{-1}\). Inverting this chain of inequalities yields that \(s_n^2 \sum_{j\neq i} h_{ij}^2\) is bounded from above uniformly over all \(i \in [n]\).} (ref)(iii) can be viewed as a mild technical requirement that there be more than one “effective” instrument in the hat matrix.\footnote{In the case of a standard projection matrix (no deleted diagonal), (ref)(iii) would be satisfied whenever \(\rank(z(z'z)^{-1}z) > 1\), which occurs whenever there are at least two linearly independent instruments.} This condition can be easily verified in practice by examining the eigenvalues of \(HH'\).

Next, I make a high level assumption that the estimation error in \(\hat \rho(\cdot)\) can be treated as negligible in both the numerator and denominator. Later, I will verify this assumption for the particular choice of \(\hat \rho(\cdot)\) described in (ref). For each \(\ell \in [d_x]\) define \(\widehat\Pi^I_{\ell,i} \coloneqq \sum_{j\neq i} h_{ij}r_{\ell j}\), the version of the first stage estimates that could be constructed if \(\rho(\cdot)\) was known to the researcher. Using these, define the magnitude of estimation error in the numerator and denominator as

align*[align* omitted — 345 chars of source]
assumption[Estimation Error] Estimation error in both the numerator and denominator of the test statistic can be treated as negligible, \((\Delta_N, \Delta_D)\to_p 0\).

Showing that \((\Delta_N, \Delta_D)\to_p 0\) implies that estimation error can be treated as negligible for the test statistic \(\JK(\beta_0)\) requires some care. In a standard approach, this would straightforwardly follow from application of the continous mapping theorem. However, this approach requires that the scaled numerator and denominator each have well defined distributional limits, something that is not required by the direct Gaussian approximation. Instead, I establish and make use of anticoncentration bounds to show that the basic results of the continuous mapping theorem still apply even when the numerator and denominator do not have weak limits.

Finally, in addition to characterizing the limiting distribution of \(\JK(\beta_0)\) under \(H_0\), I also examine the behavior of \(\JK(\beta_0)\) in local neighborhoods of the null. These local neighborhoods are characterized by the local power index \(P\), defined below, as well as an additional regularity condition that restricts the size of \(\E[\eps_i(\beta_0)]\) relative to \(\E[r_{\ell i}]\).

assumption[Local Identification] (i) The local power index is bounded \(P \leq c\) for \[ P = \sum_{\ell = 1}^{d_x} \E\bigg[\bigg(\frac{s_{\ell,n}}{\sqrt{n}}\sum_{i=1}^n \widehat\Pi_{\ell i}^I\Pi_i'(\beta - \beta_0)\bigg)^2\bigg] \] (ii) \(\E[(s_{n,\ell}\sum_{j\neq i} h_{ji} \eps_j(\beta_0))^2] \leq c\) for all \(\ell = 1,\dots,d_x\).

Under \(H_0\), (ref) is trivially satisfied since \((\beta - \beta_0) = 0\) and \(\sum_{j\neq i} s_{\ell,n}^2h_{ji}^2 \leq c\) for each \(\ell \in [d_x]\). The local power index is the second moment of the scaled numerator of the test statistic is a measure of the association between the true first stage \(\Pi_i\) and the first-stage estimates \(\widehat\Pi_i\). In (ref), I discuss how the strength of this association is related to the power of the test under local alternatives. (ref)(ii) can be roughly interpreted as requiring the local neighborhoods of \(H_0\) considered to be those in which the means of \((\eps_1(\beta_0),\dots,\eps_n(\beta_0))\) are of the same or lesser order than the means of \((r_1,\dots,r_n)\).

Limiting Behavior of the Test Statistic

Having detailed the assumptions needed for the argument, I now present results showing that, in local neighborhoods of \(H_0\), the distribution of the test statistic can be uniformly approximated by the distribution of the Gaussian statistic described in (ref). As mentioned, the argument can be simplified to require lighter moment bounds when \(d_x = 1\). These moment bounds will be made on the model primitives, \(\eta_i \coloneqq (\beta - \beta_0)v_i + \eps_i = \eps_i(\beta_0) - \E[\eps_i(\beta_0)]\) and \(\zeta_{\ell i} \coloneqq v_{\ell i} - \rho_\ell(z_i)\eta_i = r_{\ell i} - \E[r_{\ell i}]\) for each \(\ell \in [d_x]\).

theorem[Single Endogenous Variable] Suppose that (ref) hold. In addition suppose that (i) \(\{|\Pi_i| + |\rho(z_i)| + |(\beta - \beta_0)|\} \leq c\) and (ii) for any \(r,s \in \SZ_+\) satisfying \(r + s \leq 6\) \(c^{-1} \leq \E[|\eta_i|^r|\zeta_i|^s] \leq c\). Then, for \(d_x = 1\), \[ \sup_{a\in\SR} \big|\Pr(\JK(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big|\to 0. \] In particular, under \(H_0\), \(\JK(\beta_0) \rightsquigarrow \chi^2_1\).

In the case of \(d_x = 1\), I additionally show that the test based on the \(\JK_I(\beta_0)\) statistic is consistent whenever the power index diverges, \(P\to\infty\), and (ref)(ii) holds.

theorem[Consistency] Suppose that Assumptions (ref), (ref), and (ref)(ii) hold along with the additional conditions of (ref). Then, if \(P \to \infty\) the test based on \(\JK(\beta_0)\) is consistent; i.e for any fixed \(a \in \SR\), \(\Pr(\JK(\beta_0) \leq a) \to 0\).

The dependence of the consistency result on (ref)(ii) is a nontrivial restriction because of the bias taken on in constructing \(r_i\). In particular, against certain alternatives it is possible that \(\E[\widehat\Pi_i^I] = 0\) for all \(i \in [n]\) even under strong identification. This is an extreme case, however. In general, bias in \(\E[r_i]\) does not imply a violation of (ref)(ii), which requires only that the size of \(\E[r_i]\) be of a weakly greater order than that of \(\E[\eps_i(\beta_0)]\). Moreover, as discussed in (ref), (ref) does not necessarily rule out consistency when \(P \to \infty\) but (ref)(ii) fails.

Regardless, bias taken on in constructing \(r_i\) has consequences for the power of the test in finite samples. The combination with the sup-score test described in (ref) is an attempt to rectify this. While this attempt is not perfect, it appears to work well both in the empirical application to the data of gilchrist-glassberg-2016 and in the simulation study of (ref).

The argument when \(d_x > 1\) is considerably more involved than the case where \(d_x = 1\) and requires strengthened moment condition on the variables \(\eta_i\) and \(\zeta_i\). Given a random variable \(X\) and \(\upsilon > 0\) the Orlicz (quasi-)norm is defined \[ \|X\|_{\psi_\upsilon} \coloneqq \inf\{t > 0: \E\exp(|X|^\upsilon/t^\upsilon) \leq 2\} \] Random variables with a finite Orlicz norm for some \(\upsilon \in (0,1] \cup\{2\}\) are termed \(\alpha\)-sub-exponential random variables Gotze-subexponential-polynomial,sambale2022notes. This class of encompasses a wide range of potential distributions including all bounded and sub-Gaussian random random variables (with \(\upsilon = 2\)), all sub-exponential random variables such as Poisson or noncentral \(\chi^2\) random variables (with \(\upsilon = 1\)), as well as random variables with “fatter” tails such as Weibull distributed random variables with shape parameter \(\upsilon \in (0,1]\). Thus, while imposing that the variables \(\eta_i\) and \(\zeta_i\) are \(\alpha\)-sub-exponential is notably stronger than the finite sixth moments required by (ref), it may still be plausible in a wide range of empirical settings.

theorem[Uniform Approximation] Suppose that (ref) hold. In addition suppose that (i) \(c^{-1} \leq \lambda_{\min}(\E[\eta_i\eta_i']) \leq \lambda_{\max}(\E[\eta_i\eta_i']) \leq c\) and (ii) for some \(\nu \in (0,1]\cup\{2\}\) both \(\|\eta_i\|_{\psi_\nu} \leq c\) and \(\|\zeta_{\ell i}\|_{\psi_\nu} \leq c\). Then, \[ \sup_{a\in\SR}|\Pr(\JK(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)|\to 0 \] In particular, under \(H_0\), \(\JK(\beta_0)\rightsquigarrow \chi^2_{d_x}\).

While \(\JK_G(\beta_0)\) does not have a fixed distribution, examining its behavior is still tractable and allows for insight into the power properties of the jackknife K-test. In (ref), I use this result to analyze the local power of the proposed test. To improve power against certain alternatives, I suggest a combination with the sup-score statistic of BCCH-2012.

Controlling Estimation Error

The final step is to show that estimation error in \(\hat \rho\) can be treated as negligible, that is verify (ref). Establishing that this holds even when identification is weak makes use of the fact that estimation error in \(\hat \rho(\cdot)\) enters the test statistic only through its interaction with the implied error \(\eps_i(\beta_0)\). When identification is weak, the implied error is nearly conditional mean independent of the instruments and thus the product, \(\eps_i(\beta_0)\hat\rho(\cdot)\), is insensitive to small pertubations in \(\hat\rho(\cdot)\).\footnote{In the langauge of CCDDHNR-2018, this is termed “Neyman Orthogonality.” Under strong identification, Neyman orthogonality allows for the use of many machine learning techniques to estimate the first stage. This (approximate) orthogonality also holds under structural parameter sequences that are local to the null.}

In theory, this orthogonality could be combined with cross-fitting to allow for the use of other machine learning methods to estimate \(\hat\rho(\cdot)\), as in CCDDHNR-2018. This possibility is explored briefly in (ref). The use of other machine learning methods may be useful if sparsity of \(\rho(\cdot)\) is not a plausible assumption in the researcher's empirical setting as estimators such as random forests and neural networks can be consistent in high dimensional settings under alternate assumptions.\footnote{chi2022asymptotic provide results for random forests under a “sufficient impurity decrease” assumption while farrellNeuralNetwork2021 and Schmidt-Hieber-2020-neural-networks provide results for neural networks under a generalized heirarchical model assumption.} However, I focus on the \(\ell_1\)-penalized procedure proposed in (ref) both for expositional simplicity and because the sparsity assumption required for the consistency of this procedure mirrors that needed for the popular post-Lasso first-stage estimator.

assumption[Estimation Error] For each \(\ell \in [d_x]\) (i) there is a fixed constant \(\upsilon \in (0,1]\cup\{2\}\) such that \(\|\eta_i\|_{\psi_\upsilon} \leq c\); (ii) the basis terms \(b(z_i)\) are bounded, \(\|b(z_i)\|_\infty \leq c\); (iii) the approximation error satisfies \((\E_n[\xi_{\ell i}^2])^{1/2} = o(n^{-1/2})\); (iv) the researcher has access to an estimator \(\widehat\phi\) of \(\phi\) satisfying \(\log(d_b n)^{2/(\upsilon\wedge 1)}\|\widehat\phi_\ell- \phi_\ell\|_1 \to_p 0\); (v) the following moment bounds hold \begin{enumerate}[(va)] • \(\max_{1 \leq l \leq d_b} \big|\E\big[\frac{s_n}{\sqrt{n}}\sum_{i=1}^n \sum_{j\neq i} h_{ij}\eps_i(\beta_0)b_l(z_j)\eps_j(\beta_0)\big]\big| \leq c \)\(\max_{\substack{1 \leq i \leq n \\ 1 \leq l \leq d_b}} |\E[s_n\sum_{j\neq i} h_{ij}b_l(z_j)\eps_j(\beta_0)]| \leq c.\) \end{enumerate}

(ref)(i) ensures that \(\eta_i\) has finite exponential moments (i.e has a well defined moment generating function), which is required to allow the number of basis terms used to approximate \(\rho(\cdot)\) to grow at a near exponential rate compared to the sample size (\(d_b \gg n\)). When using fewer basis terms this assumption may be relaxed. (ref)(ii) is a standard condition in \(\ell_1\)-penalized estimation. At the cost of extra notation, it can be relaxed and the sup-norm of the basis terms can be allowed to grow slowly with the sample size to accommodate bases such as normalized b-splines or wavelets. (ref)(iii) is a bound on the rate of decay of the approximation error, similar to the approximate sparsity condition of BCCH-2012.

(ref)(iv) is a high-level condition on the rate of consistency of the parameter estimate \(\hat\phi\) in the \(\ell_1\) norm. This can be verified under approximate sparsity for both the LASSO estimator in (ref) or post-LASSO procedures based on refitting an unpenalized version of (ref) only using the basis terms selected in a LASSO first stage. See BCCH-2012,vanDerGreer2016,Tan-2017, and CS-2021 for references under various choices of penalty parameter. This condition allows for the dimensionality of the basis terms, \(d_b\), to grow near exponentially as a function of the sample size. Following the analysis of Tan-2017 this may be satisfied as long as \(s^2\log^{2(\upsilon + 1)/\upsilon}(d_b n)/n\to 0\), where the sparsity index \(s\) denotes the number of nonzero elements of \(\phi\).

(ref)(v) is a strengthening of the definition of local neighborhoods and can be interpreted similarly to (ref)(ii). Since the moment conditions in (ref)(va,vb) hold with \(b_\ell(z_j)\eps_j(\beta_0)\) replaced with \(r_j\), (ref)(v) can be interpreted as requiring that \(|\E[\sum_{j\neq i} h_{ij}b_\ell(z_j)\eps_j(\beta_0)]|\) is on the same order as \(|\E[\sum_{j\neq i} h_{ij}r_j]|\) for all \(i =1,\dots,n\) and \(\ell = 1,\dots,d_b\). As with (ref)(ii), it is trivially satisfied under \(H_0\) or, using the fact that \(\max_i \sum_{j\neq i} s_n^2 h_{ij}^2 \leq c\), whenever \(\E[\eps_i(\beta_0)] = \Pi_i(\beta - \beta_0)\) is in a \(\sqrt{n}\)-neighborhood of zero.

theorem[Estimation Error] Suppose that (ref) hold. Then \((\Delta_N,\Delta_D) \to_p 0\).
remarkWhen \(d_x = 1\), a sufficient condition for (ref)(i) is that there is some fixed quantile \(q \in (0,100)\) such that \((cq)^{-1} \leq \frac{q^\text{\tiny th}\text{-quantile of }\E[(\widehat\Pi_i^I)^2]}{\max_i \E[(\widehat\Pi_i^I)^2]}\). In practice this can be verified by checking that there is some quantile \(q\) such that both \begin{equation} \frac{q^{\tiny th}-quantile of \sum_{j\neq i} h_{ij}^2}{\max_i \sum_{j\neq i} h_{ij}^2} \andbox \frac{q^{\tiny th}-quantile of (\sum_{j\neq i} h_{ij} \hat r_j)^2}{\max_i (\sum_{j\neq i} h_{ij} \hat r_j)^2 } \end{equation} are bounded away from zero. Similarly, (ref)(ii) can be verified by checking that \(\max_i \sum_{j\neq i} h_{ji}^2 / \max_i \sum_{j\neq i} h_{ij}^2\) is bounded from above. The scaling factor \(s_n\) captures both the “size” of the elements in the hat matrix \(H\) and the strength of identification. If elements of the hat matrix are on the same order as a constant, one would expect \(s_n = O(n^{-1})\) under strong identification (\(\Pi_i\propto 1\)) while \(s_n = O(n^{-1/2})\) under weak identification (\(\Pi_i\lesssim n^{-1/2}\)).
remark[] The balanced-design condition in (ref)(i) is neither weaker nor stronger than that in the many instruments literature crudu_mellace_sandor_2021,Mikusheva_Sun_2022,Matsushita_Otsu_2022,lim2022conditional. These papers require that the projection matrix \(P = \vz(\vz'\vz)^{-1}\vz'\) satisfies \([P]_{ii} \leq \delta \leq 1\) for some value \(\delta\) and all \(i \in [n]\). Since \(P\) is idempotent, \([P]_{ii} = 1\) for some \(i \in [n]\) implies that \([P]_{ij} = 0\) for \(j \neq i\).\footnote{Since \(P\) is idempotent, \([P]_{ii} = \sum_{i=1}^n [P]_{ij}^2 = [P]_{ii}^2 + \sum_{j\neq i} [P]_{ij}^2\).} This would not violate (ref) if one were to take \(H\) such that \(h_{ij} = [P]_{ij}\bm{1}\{i \neq j\}\); \(\E[(\widehat\Pi_i^I)^2] = 0\) is allowed for a constant share of \(i \in [n]\). Conversely, if the instruments are fixed or grow slowly, it is possible to construct a projection matrix \(P\) of rank \(d_z\) where \([P]_{ii}\) is bounded away from one for all \(i\in [n]\), but “most” of the rows are zero. I view this as a theoretical edge case, however, that seems unlikely to result from real data.
remarkThe modified Lindeberg interpolation method allows me to give a nearly uniform explicit bound on the Gaussian approximation error in the case where \(d_x = 1\). In particular, I show that for any fixed value \(\Delta > 0\); \begin{align*} \sup_{a \leq \Delta} \big| \Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big| \leq C n^{-2/13} \end{align*} where \(C\) is a constant that depends only on \((c,\Delta)\) and \(\JK_I(\beta_0)\) is the version of the test statistic that could be constructed if \(\rho(\cdot)\) was known to the researcher. While it does not account for estimation error in \(\hat\rho(\cdot)\), obtaining an explicit bound reflects an improvement over the original analyses of K-statistics in Kleibergen_2002,Kleibergen-2005. These original studies rely on continuous mapping theorems to obtain the limiting chi-squared distributions, making the rate of decay of the approximation error difficult to analyze.
remark[] The interpolation argument relies on the fact that the first and second moments of \((\tilde\eps_i(\beta_0),\tilde r_i)\) are the same as the first and second moments of \((\eps_i(\beta_0),r_i)\) to match the first and moments of one-step deviations with Gaussian analogs. Without the jackknife form of \(\widehat\Pi_i^I\), these one step deviations would additionally contain cross-terms such as \(h_{ii}r_i\eps_i(\beta_0)\), for \(i \in [n]\). While the first moment of this cross-term is matched by the first moment of the Gaussian analog, \(h_{ii}\tilde\eps_i(\beta_0)\tilde r_i\), the second moment is not matched. This is manageable, however, so long as the terms \(h_{ii}\) are “small.” An example of when the \(h_{ii}\) terms are small is when \(H\) is taken to be the OLS projection matrix, \(H = \vz(\vz'\vz)^{-1}\vz\), and the number of instruments satisfies \(d_z^3/n \to 0\). See (ref) for details.
remark[] (ref) does not necessarily rule out that a test based on \(\JK_I(\beta_0)\) is consistent when \(P\to\infty\) but (ref)(ii) fails to hold. There is reason to believe that this issue can be overcome, AMS-2004 show that the K-statistic of Kleibergen_2002 is consistent against fixed alternatives under strong identification. However, a full consistency result is not pursued here and left to future work.
remark[] Approximate sparsity of \(\rho(z_i)\) may be a particularly palatable assumption in cases where the instrument set is generated by functions of a smaller initial set of instruments, as in AK-1991,Paravisini-2014,gilchrist-glassberg-2016, and derenoncourt-2022. In these cases, the dimensionality of the basis, \(d_b\), may not need to be much larger than the dimensionality of the instruments, \(d_z\), to provide a good approximation of \(\rho(z_i)\). Interestingly, if taking \(b(z_i) = z_i\) provides a good approximation of \(\rho(z_i)\), then consistency of \(\hat\phi\) in \(\ell_1\)-norm is achievable under \(d_z^2/n \to 0\) even if \(\phi\) is fully dense. This requirement is substantially weaker than the \(d_z^3/n\to 0\) requirement of the standard K-statistic.

Local Power and Improvements

Using the characterization of the limiting behavior of the test statistic derived in (ref), I analyze the local power properties of the test. Unfortunately, against certain alternatives the test statistic may have trivial power, a deficiency shared with the K-statistics of Kleibergen_2002,Kleibergen-2005. Finally, I describe how the the combination with the sup-score statistic, described in (ref), attempts to remedy this and formally establish its validity.

Local Power Properties

For exposition, I focus on the case where \(d_x = 1\). In local neighborhoods of \(H_0\), as defined in (ref), (ref) implies that the limiting behavior of \(\JK(\beta_0)\) can be analyzed by examining the behavior of the Gaussian analog statistic, \(\JK_G(\beta_0)\). Conditional on the vector \(\tilde r = (\tilde r_1,\dots,\tilde r_n)\), the distribution of \(\JK_G(\beta_0)\) is nearly non-central \(\chi^2_1\) with noncentrality parameter \(\mu(\tilde r)\), \(\JK_G(\beta_0) |\tilde r \sim A^2(\tilde r) \cdot \chi^2_1(\mu(\tilde r))\):

align*[align* omitted — 321 chars of source]

Under local alternatives, the terms \(\Pi_i^2(\beta - \beta_0)^2 \to 0\) so that \(A(\tilde r) \to 1\) and \(|\mu^2(\tilde r) - \mu_\infty^2(\tilde r)|\to 0\), where

equation[equation omitted — 187 chars of source]

The numerator of \(\mu^2_\infty(\tilde r)\) suggests that power is maximized when the first-stage estimate \(\tilde \Pi_i\) is close to the true first stage value \(\Pi_i\). Indeed, when errors are homoskedastic \(\mu_\infty^2(\tilde r)\) is maximized by setting \(\tilde \Pi_i = \Pi_i\) reflecting the classical result of Chamberlain-1987. The denominator of \(\mu^2_\infty(\tilde r)\) suggests that having first-stage estimates \(\tilde \Pi_i\) with low second moments may increase power. This guides the recommendation for the use of \(\ell_2\)-regularization in constructing the hat matrix, \(H\).

Unfortunately, estimators of \(\Pi_i\) based on \(r_i = x_i - \rho(z_i)\eps_i(\beta_0)\) may not be close to \(\Pi_i\) under \(H_1\). This is because the mean of \(r_i\) will in general differ from \(\Pi_i\) \[ \E[r_i] = \Pi_i - \rho(z_i)\Pi_i(\beta - \beta_0) \] This deficiency is inherited from the similarity of the \(\JK(\beta_0)\) statistic to the K-statistic. As pointed out by Moreira-2001, this need not be an issue as long as there is a fixed constant \(C\neq 0\) such that \(\E[r_i] = C\Pi_i\) for all \(i \in [n]\). However, in general, this will introduce bias into the first-stage estimates \(\widehat\Pi_i\) under \(H_1\). The power implications of this bias are particularly pronounced when \(\rho(z_i)\) is a constant \((\beta - \beta_0) = 1/\rho(z_i)\). In this case, \(\E[r_i]\), and thus \(\E[\tilde\Pi_i]\), will equal zero for each \(i\in[n]\), and the \(\JK(\beta_0)\) statistic will select a direction completely at random to direct power into.\footnote{AMS-2006 and Andrews-2016 point out this deficiency in the context of the K-statistics of Kleibergen_2002,Kleibergen-2005.}

A Simple Combination Test

To combat this loss of power for tests based on the K-statistic, a common strategy is to combine the K-statistic with the Anderson-Rubin statistic based on a conditioning statistic. While the Anderson-Rubin statistic does not have optimal power on its own, it has the benefit of directing power equally in all directions avoiding the pitfalls of the K-statistic which lacks power in certain directions. Prominent examples of such tests are the conditional likelihood ratio test of Moreira_2003, the GMM-M test of Kleibergen-2005, and the minimax regret tests of Andrews-2016. These combinations make use of the fact that the Anderson-Rubin statistic is asymptotically independent of both the K-statistic and the conditioning statistic.

Unfortunately, the asymptotic validity of these tests under heteroskedasticity is based on the assumption that \(d_z^3/n\to 0\), which may not reasonably describe many settings discussed above. Instead, to improve the power of tests based on the jackknife K-statistic, I consider a simple combination with the sup-score statistic of BCCH-2012. The test based on the sup-score statistic (ref) is similar in spirit to the Anderson-Rubin test but controls size even when \(d_z\) grows near exponentially as a function of the sample size.

equation[equation omitted — 196 chars of source]

A size \(\theta \in (0,1)\) test based on the sup-score statistic rejects whenever \(S(\beta_0) > c^S_{1-\theta}\) where, for \(e_1,\dots,e_n\) iid standard normal and generated independently of the data, \(c^S_{1-\theta}\) is the simulated multiplier bootstrap critical value: \[ c^S_{1-\theta} \coloneqq (1-\theta)\text{ quantile of } \sup_{1 \leq \ell \leq d_z} \bigg|\frac{\sum_{i=1}^n e_i \eps_i(\beta_0)z_{\ell i}}{(\sum_{i=1}^n z_{\ell i}^2)^{1/2}}\bigg|\text{ conditional on }\{(y_i,x_i,z_i)\}_{i=1}^n. \] As with the Anderson-Rubin test, tests based on the sup-score statistic may have suboptimal power properties in overidentified models as it does not incorporate first-stage information. However, the sup-score statistic does retain the benefit of directing power evenly in all directions, avoiding pitfalls of tests based on \(\JK(\beta_0)\) against certain alternatives.

The combination test will be based on an attempt to detect whether the alternative \(\beta\) is such that \(\E[\widehat\Pi_{\ell,i}^I] = 0\) for all \(i = 1,\dots, n\) and some \(\ell \in [d_x]\). When this is the case, the researcher would prefer to test the null hypothesis using the sup-score statistic. As mentioned in (ref), detection of whether \(\E[\widehat\Pi_{\ell,i}^I] = 0\) for some \(i \in [n]\) is based on the conditioning statistic:

equation[equation omitted — 221 chars of source]

Under the assumption that \(\E[\widehat\Pi_i^I] = 0\) for all \(i \in [n]\), quantiles of the conditioning statistic can be simulated analogously to the sup-score critical value. For a new set of \(e_1,\dots,e_n\) iid standard normal and generated independently of the data, and for any \(\theta\in (0,1)\), define the conditional quantile

equation[equation omitted — 279 chars of source]

Depending on the value of the conditioning statistic, the thresholding test decides whether the test based on \(\JK(\beta_0)\) or one based on \(S(\beta_0)\) should be run.

equation[equation omitted — 315 chars of source]

for some cutoff \(\tau\), which I take in the simulation study and empirical exercise to be the 75\textsuperscript{th} quantile of the distribution of \(C\) under the assumption that \(\E[\widehat\Pi_i^I] = 0, \forall i \in [n]\).

To show that the thresholding test controls size, I compare the rejection probability to that of a Gaussian analog. In addition to \(\JK_G(\beta_0)\), defined in (ref), define the Gaussian analogs of \(\S(\beta_0)\) and the conditioning statistic \(C\):

align*[align* omitted — 321 chars of source]

where, as in (ref), \((\tilde\eps_i(\beta_0),\tilde r_i)'\) are generated independently of each other and the data following a Gaussian distribution with the same mean and covariance matrix as \((\eps_i(\beta_0),r_i)\). Since \(\Cov(\tilde\eps_i(\beta_0),\tilde r_i) = 0\) under \(H_0\), the statistics \(C_G\) and \(S_G(\beta_0)\) are independent under the null. Similarly, the null distribution of \(\JK_G(\beta_0)\) is the same conditional on any realization of \((\tilde r_1,\dots,r_n)\); it is also independent of \(C_G\) under the null. The Gaussian analog thresholding test decides whether the researcher should run a test based on \(S_G(\beta_0)\) or \(\JK_G(\beta_0)\) depending on the value of \(C_G\) as in (ref).

The test statistics \(\JK_G(\beta_0)\) and \(S_G(\beta_0)\) are only marginally independent of the conditioning statistic \(C_G\) under the null. This limits the ways in which the test statistics can be combined using the conditioning statistic while still controlling size. This marginal independence in the Gaussian limit is enough, however, for the asymptotic validity of the thresholding test, \(T(\beta_0;\tau)\). To establish that the behavior of the pairs \((C,\JK(\beta_0))\) and \((C,S(\beta_0))\) can be approximated by the behavior of \((C_G,\JK_G(\beta_0))\) and \((C_G,S_G(\beta_0))\), respectively, I rely on the following assumption:

assumption[Combination Conditions] Assume that for each \(\ell \in [d_x]\) (i) there is a \(\upsilon \in (0,1]\cup\{2\}\) such that \(\|\zeta_{\ell i}\|_{\psi_\upsilon} \leq c\); (ii) \(\max_{i,j} |\frac{h_{ij}}{(\E_n[h_{ij}^2])^{1/2}}| + \max_{l,i}| \frac{z_{li}}{(\E_n[z_{li}^2])^{1/2}} | \leq c\); and (iii) \(\log^{7 + 4/\upsilon}(d_zn)/n\to 0\).

(ref)(i) is a strengthening of the moment bound on \(r_i\) similar to that of (ref)(i). As discussed, while more restrictive than the condition in (ref), this still allows for a wide range of potential distributions for \(r_i\). (ref)(ii) requires that the number of observations used to test \(\E[\widehat\Pi_i] = 0\) via the conditioning statistic and the number of observations used to test the null hypothesis via the sup-score test are both growing with the sample size. It can be verified by looking at the hat matrix \(H\) and the instruments. Finally, (ref)(iii) is a light requirement on the number of instruments \(d_z\) needed for the validity of the sup-score test. It allows the number of instruments to grow near exponentially as a function of sample size.

theorem[] Suppose the conditions of (ref) and (ref) hold. Then, \begin{enumerate} • the test based on \(T(\beta_0;\tau)\) has asymptotic size \(\alpha\) for any choice of cutoff \(\tau\), and • if \(\E[\widehat\Pi_i^I] = 0\) for all \(i \in [n]\), there exist sequences \(\delta_n \searrow 0\) and \(\beta_n \searrow 0\) such that with probability at least \(1 - \delta_n\), \[ \sup_{\theta\in (0,1)}\big|\Pr_e(C \leq c^C_{1-\theta}) - (1-\theta)\big| \leq \beta_n, \] where \(\Pr_e(\cdot)\) denotes the probability with respect to only the variables \(e_1,\dots,e_n\). \end{enumerate}

The first part of (ref) establishes the asymptotic validity of the thresholding test \(T(\beta_0;\tau)\) for any choice of cutoff \(\tau\). While not explicitly stated in the statment of the theorem, this result is uniform in the choice of \(\tau\); for any sequence \(\{\tau_n\}\subset \SR_+\) the sequnce of testing procedures \(T(\beta_0;\tau_n)\) will also have asymptotic size \(\alpha\). The proof of this statement follows the logic outlined above. The second part of (ref) establishes the validity of the multiplier bootstrap procedure to approximate quantiles of the conditioning statistic. It follows directly from results in belloni2018highdimensional after verifying that the conditions needed for error taken on from estimation of \(\rho(z_i)\) can treated as negligible under (ref).

In the case of a single endogenous variable, \(d_x = 1\), (ref) could be established under the lighter conditions of (ref) along with (ref). However, for brevity, I do not seperate the two cases here.

remark[] It is useful to compare the \(\JK(\beta_0)\) statistic to the JLM statistic of Matsushita_Otsu_2022, which also converges to a limiting \(\chi^2\) distribution when \(d_z \to \infty\). In the case where \(d_x = 1\) the JLM statistic can be expressed \begin{equation} JLM(\beta_0) \coloneqq \bigg(\frac{\sum_{i=1}^n \eps_i(\beta_0) \sum_{j\neq i}P_{ij}x_j}{\sum_{i=1}^n\eps_i^2(\beta_0)\big(\sum_{j\neq i} P_{ij} x_j\big)^2 + \sum_{i=1}^n\sum_{j\neq i}P_{ij}^2\eps_i(\beta_0)\eps_j(\beta_0)x_ix_j}\bigg)^2 \end{equation} where \(P = \vz(\vz'\vz)^{-1}\vz'\) is the standard OLS projection matrix. This expression looks simillar to that of the \(\JK(\beta_0)\) statistic with first stage estimates \(\widehat\Pi_i = \sum_{j\neq i} P_{ij}x_j\). From this expression, one can posit two potential reasons for the increased power of tests based on the \(\JK(\beta_0)\) statistic seen in the empirical applications in (ref) and the simulation study of (ref). The first is that, when the bias of \(r_i\) is not too adverse, first stage estimates based on the “true” jackknife ridge or jackknife OLS described in (ref) may be closer to the true first stage, \(\Pi_i\) than those based on the deleted-diagonal projection matrix. As seen in (ref), higher quality first stage estimates can improve the power of the test by increasing the correlation between these estimates and \(\eps_i(\beta_0)\) under \(H_1\). This loss in quality of first-stage estimates based on the deleted-diagonal projection matrix may be negligible when the diagonal elements, \(P_{ii}\), are small in which case the deleted diagonal estimates closely resemble standard OLS estimates. However, when either the number of instruments is large relative to the sample size or the instruments are highly correlated the diagonal elements \(P_{ii}\) will be large in which case estimates of \(\Pi_i\) based on the deleted-diagonal projection matrix may not be accurate. This pattern can be seen in both empirical applications in (ref); in both the data of gilchrist-glassberg-2016 and AK-1991 the improvments in power from using the \(\JK(\beta_0)\) statistic become more pronounced as the number of instruments increases. A second potential reason for improved power is that the \(\JK(\beta_0)\) statistic uses individual scores, \(\eps_i(\beta_0)\widehat\Pi_i\), that are uncorrelated with each other under \(H_0\). That is, for \(j \neq i\), \(\E[\eps_i(\beta_0)\widehat\Pi_i\eps_j(\beta_0)\widehat\Pi_j] = 0\). Thus, the second term in the denominator of (ref), which accounts for the covariance between individual scores in the numerator of the JLM statistic, does not appear in the expression of \(\JK(\beta_0)\). Note that this second term has a positive expectation under both positive selection, \(\E[\eps_i(\beta_0)x_i] > 0\) for all \(i \in [n]\), and negative selection, \(\E[\eps_i(\beta_0)x_i] < 0\) for all \(i \in [n]\). If this term is large, it can substantially increase the denominator of the JLM statistic relative to that of the \(\JK(\beta_0)\) statistic, reducing its power. This again may be likely when the diagonal elements, \(P_{ii}\), are large due to idempotency of the projection matrix: \(P_{ii} = \sum_{j=1}^n P_{ij}^2\). Moreover, by construction \(\Var(r_i) \leq \Var(x_i)\), so a large second term in the denominator of the JLM statistic may not be offset by a smaller first term, at least in local regions of \(H_0\). In sum, when bias taken on in constructing \(r_i\) is not too adverse, the \(\JK(\beta_0)\) statistic may have a larger numerator than the JLM statistic due to the use of higher quality first-stage estimates and a smaller denominator due to the use of uncorrelated individual scores. Since both the \(\JK(\beta_0)\) and JLM statistics are compared to the same \(\chi^2\) quantile, both of these properties may lead to more likely rejection of tests based on the \(\JK(\beta_0)\) statistic under \(H_1\). Matsushita_Otsu_2022 note that the power properties of the JLM statistic are similar to those of the JAR statistic of Mikusheva_Sun_2022, suggesting that the improvments in power compared to the JAR test seen in (ref) and (ref) may be explained similarly.
remark[] As mentioned by Andrews-2016 in the context of the standard K-statistic, the attempt to rectify the power deficiency via this particular conditioning statistic is not perfect. In particular, under heteroskedasticity, the means of the partialed-out endogenous variables, \(\E[r_i]\), may not be scaled versions of the true first stages. However, as long as \(\E[r_i] \neq 0\), one can still expect \(\E[\widehat\Pi_i^I] = \sum_{j\neq i} h_{ij}\Pi_i + (\beta - \beta_0)\sum_{j\neq i} h_{ij}\rho(z_i)\Pi_i\) to be related to the true fist stage \(\Pi_i\) and for the test to have nontrivial power. Moreover, in light of the dependence of the consistency result in (ref) on (ref)(ii), in the case where \(\E[\widehat\Pi_i] = 0\) for all \(i\in[n]\) it may be particularly important to avoid using the jackknife K-statistic to test \(H_0\).

Conclusion

I propose a new test for the structural parameter in a linear instrumental variables model. This test is based on a jackknife version of the K-statistic and the limiting behavior of the test is analyzed via a novel direct Gaussian approximation argument. I show that, as long as an auxiliary parameter can be consistently estimated, the test is robust to both the strength of identification and the number of instruments; the limiting distribution of the test statistic does not depend on either of these factors. Consistency of the auxiliary parameter can be achieved under approximate sparsity using simple-to-implement \(\ell_1\)-penalized methods.

I characterize the behavior of the jackknife K-statistic in local neighborhoods of the null. To address a power deficiency that tests based on jackknife K-statistic inherit from their non-jackknife namesakes, I propose a testing procedure that decides whether the researcher should run a test via the jackknife K-statistic or one via the sup-score statistic based on the value of a conditioning statistic. While this combination may not fully address the power decline, I show that it works well in a simulation study and leave further refinements to future work.

\startappendix

Proof of (ref)

(ref) follows from the following two main technical lemmas, the proofs of which will comprise the majority of this appendix section. Let \(\JK_I(\beta_0)\) be the version of the test statistic that could be constructed if \(\rho(\cdot)\) was known to the researcher, defined in more detail shortly.

lemma[Infeasible Uniform Approximation] Suppose that (ref) hold as well as the moment bounds of (ref). Then, \[ \sup_{a \in \SR}\big|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big| \to 0 \]
lemma[Estimation Error] Suppose that (ref) and (ref) hold as well as the moment bounds of (ref). Then, if \((\Delta_N,\Delta_D)\to_p0\), \[ \sup_{a \in \SR} \big|\Pr(\JK(\beta_0) \leq a) - \Pr(\JK_I(\beta_0) \leq a)\big|\to_p 0 \]

Proof of (ref)

Before proceeding, we will introduce some notation. Let \(\tilde H = s_n H\) and \(\tilde h_{ij} = s_n h_{ij}\), where \(s_n\) is as in (ref). Recall that \(\tilde h_{ii} = 0\) and define

align*[align* omitted — 411 chars of source]

where \((\tilde\eps_i(\beta_0),\tilde r_i)\) are jointly Gaussian with the same mean and covariance matrix as \((\eps_i(\beta_0),r_i)\) and \(\kappa_i^2(\beta_0) = \E[\eps_i^2(\beta_0)]\). Under this notation we can write \(\JK_I(\beta_0) = \frac{N^2}{D}\bm{1}_{\{D > 0\}}\) and \(\JK_G(\beta_0) = \frac{\tilde N^2}{\tilde D}\). Dealing with these forms of the statistics is difficult for the interpolation argument, since the denominator is random. Instead, we will notice that since \(D = 0 \implies N = 0\) and \(\Pr(\tilde D > 0) = 1\), for any \(a \geq 0\) we can rewrite the events

equation[equation omitted — 232 chars of source]

With this in mind define

align*[align* omitted — 84 chars of source]

Showing (ref) is then equivalent to showing that \(\sup_a|\Pr(\JK^a \leq 0) - \Pr(\tilde \JK^a \leq 0)|\to 0\). The statement \(\sup_{a < 0}\big|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big| = 0\) is immediate since both \(\JK_I(\beta_0)\) and \(\JK_G(\beta_0)\) are always weakly positive. It thus suffices to show \[ \sup_{a \geq 0} \big|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big| \to 0 \] We do so in a few lemmas, the final result being shown in (ref) at the bottom of this subsection.

lemma[Lindeberg Interpolation] Suppose that (ref) hold along with the conditions of (ref). Let \(\varphi(\cdot):\SR\to\SR\) be such that \(\varphi(\cdot) \in C_b^3(\SR)\) with \(L_2(\varphi) = \sup_x|\varphi''(x)|\) and \(L_3(\varphi) = \sup_x|\varphi'''(x)|\). Then, there is a constant \(M\) that depends only on the constant \(c\) such that: \[ |\E[\varphi(\JK^a) - \varphi(\tilde \JK^a)]| \leq \frac{M(a^3 \vee 1)}{\sqrt{n}}(L_2(\varphi) + L_3(\varphi)) \]
proof[Proof of (ref)] Begin by defining the leave-one-out numerator, denominator, and decomposed statistics \begin{align*} N_{-i} &:= \frac{1}{\sqrt{n}}\sum_{j \neq i} \dot\eps_j(\beta_0)\sum_{\ell \neq i} \tilde h_{j\ell }\dot r_\ell & D_{-i} &:= \frac{1}{n}\sum_{j \neq i}\ddot\eps_{j}^2(\beta_0)\big(\sum_{\ell\neq i} \tilde h_{j\ell} \dot r_\ell\big)^2 \\ \JK_{-i} &:= N_{-i}^2 - aD_{-i} & \end{align*} where for each \(\ell \in [n]\), \(\dot\eps_\ell(\beta_0)\) is equal to \(\eps_\ell(\beta_0)\) if \(\ell > i\) and \(\tilde \eps_\ell(\beta_0)\) if \(\ell < i\), \(\dot r_\ell\) is equal to \(r_\ell\) if \(\ell > i\) and \(\tilde r_\ell\) if \(\ell < i\), and \(\ddot\eps_\ell^2(\beta_0)\) is equal to \(\kappa_\ell^2(\beta_0)\) if \(\ell < i\) and \(\eps_\ell^2(\beta_0)\) if \(\ell > i\). While the definitions of \(\dot\eps_\ell, \dot r_\ell,\) and \(\ddot \eps_\ell\) depend on \(i\) because we will be considering only one deviation at a time, we will supress the dependence of these variables on \(i\) to simplify notation. Next, define the one-step deviations \begin{equation} \begin{split} \Delta_{1i} &:= \eps_i(\beta_0)\sum_{j=1}^n \tilde h_{ij} \dot r_j + r_i \sum_{j=1}^n \tilde h_{ji} \dot\eps_j(\beta_0)\\ \tilde \Delta_{1i} &:= \tilde\eps_i(\beta_0)\sum_{j=1}^n \tilde h_{ij} \dot r_j + \tilde r_i \sum_{j=1}^n \tilde h_{ji} \dot\eps_j(\beta_0) \\ \Delta_{2i} &:= \underbrace{a\eps_i^2(\beta_0)(\sum_{j=1}^n \tilde h_{ij} \dot r_j)^2 + a r_i^2 \sum_{j=1}^n \tilde h_{ji}^2 \ddot \eps_j^2(\beta_0)}_{\Delta_{2i}^a} + \underbrace{2a r_i\sum_{j=1}^n \ddot \eps_j^2(\beta_0)\sum_{\ell \neq i} \tilde h_{j \ell}\tilde h_{ji}\dot r_\ell}_{\Delta_{2i}^b}\\ \tilde\Delta_{2i} &:= \underbrace{a\kappa_i^2(\beta_0)(\sum_{j=1}^n \tilde h_{ij} \dot r_j)^2 + a \tilde r_i^2 \sum_{j=1}^n \tilde h_{ji}^2 \ddot \eps_j^2(\beta_0)}_{\Delta_{2i}^a} + \underbrace{2a \tilde r_i\sum_{j=1}^n \ddot \eps_j^2(\beta_0)\sum_{\ell \neq i} \tilde h_{j\ell} \tilde h_{ji}\dot r_\ell}_{\tilde\Delta_{2i}^b} \end{split} \end{equation} These one-step deviations contain all the terms associated with observation \(i\) in the expression of the numerator and denominator of the test statistics. To demonstrate, note that these one-step deviations satisfy \(N_{-1} + n^{-1/2}\Delta_{11} = N\) and \(aD_{-1} + n^{-1}\Delta_{21} = aD\) as \begin{align*} N &= \frac{1}{\sqrt{n}}\sum_{i=1}^n \eps_i(\beta_0) \sum_{j=1}^n \tilde h_{ij} r_j \\ &= \frac{1}{\sqrt{n}}\sum_{j > 1} \eps_j(\beta_0) \sum_{\ell = 1}^n \tilde h_{j\ell} r_j + \eps_1(\beta_0) \frac{1}{\sqrt{n}}\sum_{j > 1} \tilde h_{1j}r_j \\ &= \frac{1}{\sqrt{n}}\sum_{j > 1}\eps_j(\beta_0)\Big\{\tilde h_{j1} r_1 + \sum_{\ell > 1} h_{j\ell} r_\ell \Big\} + \eps_1(\beta_0) \frac{1}{\sqrt{n}}\sum_{j > 1}\tilde h_{1j}r_j \\ &= \underbrace{\frac{1}{\sqrt{n}}\sum_{j > 1}\eps_j(\beta_0)\sum_{\ell > 1}h_{j\ell} r_\ell}_{N_{-1}} + \underbrace{\eps_1(\beta_0)\frac{1}{\sqrt{n}}\sum_{j > 1}\tilde h_{1j} r_j + r_1 \frac{1}{\sqrt{n}}\sum_{j > 1}\tilde h_{j1}\eps_j(\beta_0)}_{n^{-1/2}\Delta_{11}} \intertext{and} D &= \frac{1}{n} \sum_{i=1}^n \eps_i^2(\beta_0)\big(\sum_{j=1}^n \tilde h_{ij} r_j)^2 \\ &= \frac{1}{n}\sum_{j > 1}\eps_j^2(\beta_0)\big(\sum_{\ell = 1}^n \tilde h_{j\ell}r_\ell\big)^2 + \eps_1^2(\beta_0)\frac{1}{n}\big(\sum_{j > 1} \tilde h_{1j}r_j\big)^2\\ &= \frac{1}{n}\sum_{j > 1}\eps_j^2(\beta_0) \big(\tilde h_{j1}r_1 + \sum_{\ell \neq 1} \tilde h_{\ell j}r_\ell\big)^2 + \eps_1^2(\beta_0)\frac{1}{n}\big(\sum_{j > 1}\tilde h_{1j}r_j\big)^2\\ &= \underbrace{\frac{1}{n}\sum_{j > 1}\eps_j^2(\beta_0)\big(\sum_{\ell > 1}\tilde h_{\ell, j}r_\ell\big)^2}_{D_{-1}} \\ &\;\;\;\;\;\;+ \underbrace{\eps_1^2(\beta_0)\frac{1}{n}\big(\sum_{j > 1} \tilde h_{1j}r_j\big)^2 + r_1^2\frac{1}{n}\sum_{j > 1} \tilde h_{j1}^2 \eps_j^2(\beta_0) + 2r_1\frac{1}{n}\sum_{j > 1}\eps_j^2(\beta_0)\sum_{\ell > 1}\tilde h_{\ell j} r_\ell}_{(an)^{-1}\Delta_{21}} \end{align*} Using the one-step deviations, write the difference \(\E[\varphi(K^a) - \varphi(\tilde K^a)]\) as a telescoping sum, one by one replacing \((\Delta_{1i},\Delta_{2i})\) with \((\tilde\Delta_{1i},\tilde\Delta_{2i})\) in the expressions of \(\JK^a = N^2 - aD\) until we arrive at \(\tilde \JK^a = \tilde N^2 - a \tilde D\). \begin{equation} \begin{split} \E[\varphi(\JK^a) - \varphi(\tilde \JK^a)] &= \sum_{i=1}^n \E[\varphi(\JK_{-i} + n^{-1/2}N_{-i}\Delta_{1i} + n^{-1}\Delta_{1i}^2 - n^{-1}\Delta_{2i})] \\ &\hphantom{=\sum_{i=1}^n\;\;} - \E[\varphi(\JK_{-i} + n^{-1/2}N_{-i}\tilde\Delta_{1i} + n^{-1}\tilde\Delta_{1i}^2 - n^{-1}\tilde\Delta_{2i})] \end{split} \end{equation} Via a second-order Taylor expansion, we can write each term inside the summand \begin{align*} \E[Term_i] &= \E[\varphi'(\JK_{-i})\{2n^{-1/2}N_{-i}(\Delta_{1i}-\tilde\Delta_{1i}) + n^{-1}(\Delta_{1i}^2 - \Delta_{1i}^2) - n^{-1}(\Delta_{2i} - \tilde\Delta_{2i})\}] \\ &+ \E[\varphi”(\JK_{-i})\{4n^{-1}N_{-i}^2(\Delta_{1i}^2 - \tilde\Delta_{1i}^2) + n^{-2}(\Delta_{1i}^4 - \tilde\Delta_{1i}^4) - n^{-2}(\Delta_{2i}^2 - \Delta_{2i}^2)\}] \\ &+ \E[\varphi”(\JK_{-i})\{4n^{-3/2}N_{-i}(\Delta_{1i}^3 - \tilde\Delta_{1i}^3) + 4n^{-3/2}N_{-i}(\Delta_{1i}\Delta_{2i} - \tilde\Delta_{1i}\tilde\Delta_{2i})\}] \\ &+ \E[\varphi”(\JK_{-i})\{2n^{-2}(\Delta_{1i}^2 \Delta_{2i} - \tilde\Delta_{1i}^2 \tilde\Delta_{2i})\}] + R_i + \tilde R_i \end{align*} where \(R_i\) and \(\tilde R_i\) are remainder terms to be examined later. Let \(\calF_{-i}\) denote the sigma algebra generated by all random variables whose index is not equal to \(i\). Since (a) for each \(i \in [n]\) the mean and covariance matrix of \((\eps_i(\beta_0), r_i)\) is the same as the mean and covariance matrix of \((\tilde\eps_i(\beta_0), \tilde r_i)\), (b) \(\E[\eps_i^2(\beta_0)] = \kappa_i^2(\beta_0)\), and (c) random variables are independent across indices, we have that \begin{equation} \begin{split} \E[\Delta_{1i} - \tilde\Delta_{1i}|\calF_{-i}] &= \E[\Delta_{1i}^2 - \tilde\Delta_{1i}^2|\calF_{-i}] = \E[\Delta_{2i} - \tilde\Delta_{2i}|\calF_{-i}] \\ &= \E[\Delta_{2i}^b - \tilde\Delta_{2i}^b|\calF_{-i}] = \E[\Delta_{1i}\Delta_{2i}^b - \tilde\Delta_{1i}\tilde\Delta_{2i}^b|\calF_{-i}] = 0 \end{split} \end{equation} Using this we can simplify the prior display \begin{align*} \E[Term_i] &= \underbrace{n^{-2}\E[\varphi”(\JK_{-i})(\Delta_{1i}^4 - \Delta_{1i}^4)]}_{\vA_i} - \underbrace{n^{-2}\E[\varphi”(\JK_{-i})((\Delta_{2i}^a)^2 - (\tilde\Delta_{2i}^a)^2)]}_{\vB_i} \\ &- \underbrace{2n^{-2}\E[\varphi”(\JK_{-i})(\Delta_{2i}^a \Delta_{2i}^b - \tilde\Delta_{2i}^a \tilde\Delta_{2i}^b)]}_{\vC_i} + \underbrace{4n^{-3/2}\E[\varphi”(\JK_{-i})N_{-i}(\Delta_{1i}^3 - \tilde\Delta_{1i}^3)]}_{\vD_i} \\ &+ \underbrace{4n^{-3/2}\E[\varphi”(\JK_{-i})N_{-i}(\Delta_{1i}\Delta_{2i}^a - \tilde\Delta_{1i}\tilde\Delta_{2i}^a)]}_{\vE_i} + \underbrace{2 n^{-2}\E[\varphi”(\JK_{-i})(\Delta_{1i}^2 \Delta_{2i} - \tilde\Delta_{1i}^2 \tilde\Delta_{2i})}_{\vF_i} \\ &+ R_i + \tilde R_i \end{align*} where for some \(\bar \JK_{1i}\) and \(\bar \JK_{2i}\) we can write \begin{align*} R_i &= \E[\varphi”'(\bar \JK_{1i})\{n^{-1/2}N_{-i}\Delta_{1i} + n^{-1}\Delta_{1i}^2 + n^{-1}\Delta_{2i}\}^3]\\ \tilde R_i &= \E[\varphi”'(\bar \JK_{2i})\{n^{-1/2}N_{-i}\tilde\Delta_{1i} + n^{-1}\tilde\Delta_{1i}^2 + n^{-1}\tilde\Delta_{2i}\}^3] \end{align*} Applications of (ref), Cauchy-Schwarz, and the generalized Hölder inequality,\footnote{\(\E[|fgk|]^3 \leq \E[|f|^3]\E[|g|^3]\E[|k|^3]\)} will allow us to bound for a fixed constant \(M\) that depends only on \(c\), \begin{align*} |\vA_i| &\leq \frac{M}{n^2}L_2(\varphi) & |\vB_i| &\leq \frac{Ma^2}{n^2}L_2(\varphi) & |\vC_i| &\leq \frac{Ma^2}{n^{3/2}}L_2(\varphi) \\ |\vD_i| &\leq \frac{M}{n^{3/2}}L_2(\varphi) & |\vE_i| &\leq \frac{M (a\vee 1)}{n^{3/2}}L_2(\varphi) & |\vF_i| &\leq \frac{Ma^3}{n^{3/2}}L_2(\varphi) \end{align*} and \begin{align*} |R_i| + |\tilde R_i| &\leq \frac{M}{n^{3/2}}L_3(\varphi) + \frac{Ma^3}{n^3}L_3(\varphi) \end{align*} Combining these bounds and summing over \(n\) gives the result.
lemma[Gaussian Denominator Anti-Concentration] Suppose that the conditions of (ref) and (ref) hold. Then, for any sequence \(\delta_n \searrow 0\), \[ \Pr(\tilde D \leq \delta_n) \to 0 \]
proof[Proof of (ref)] Since \(\kappa_i^2(\beta_0) \in [c^{-1},c]\) for all \(i = 1,\dots,n\) we have that \(\tilde D \geq \frac{c^{-1}}{n}\sum_{i=1}^n (\sum_{j=1}^n \tilde h_{ij}r_j)^2\). Then \begin{align*} \Pr(\tilde D \leq \delta_n) &\leq \Pr\big(\frac{1}{cn}\sum_{i=1}^n\big(\sum_{j=1}^n \tilde h_{ij}\tilde r_j\big)^2 \leq \tilde \delta_n \big) \\ &= \Pr\big(\|\tilde r' \bar H^{1/2}\|^2 \leq \delta_n\big) \numberthis \end{align*} where \(\tilde r := (\tilde r_1,\dots,\tilde r_n)'\in \SR^n\) and \(\bar H := \frac{1}{cn}\tilde H \tilde H' \in \SR^{n\times n}\). \(\bar H\) is symmetric and positive semidefinite so we can take \(\bar H^{1/2}\) to be its symmetric square root, which will also be symmetric and positive semidefinite (and thus not necessarily equal to \(\sqrt{\frac{c}{n}}\tilde H\)). I provide two bounds on (ref), the first of which corresponds to the strong identification setting while the second corresponds to weak identification. First Bound. Since \(\delta_n \searrow 0\) we will eventually have that \(\delta_n < c^{-1}/2\). When this happens we can bound using Chebyshev's inequality and \(c^{-1} < \E[r'\bar H r] < c\): \begin{align*} \Pr(\tilde r'\bar H \tilde r \leq \delta_n) &= \Pr(\tilde r'\bar H \tilde r - \E[\tilde r' \bar H \tilde r] \leq \delta_n - \E[\tilde r' \bar H \tilde r]) \\ &\leq \Pr(\tilde r'\bar H \tilde r - \E[r'\bar H r] \geq \E[\tilde r'\bar H \tilde r] - \delta_n) \\ &\leq \Pr(|\tilde r'\bar H \tilde r - \E[r'\bar H r]| \geq \frac{1}{2c}) \\ &\leq 2c\Var(r'\bar H r)\numberthis \end{align*} Under strong identification we will expect \(\Var(r'\bar H r) \to 0\). Second Bound. For the second bound, we will directly use bounds on the density of Gaussian quadratic forms from anticoncentration-bernoulli-gotze-et-al. The vector \(r'\bar H^{1/2}\) is Gaussian with covariance matrix \(\Sigma_r = \bar H^{1/2}\mR \bar H^{1/2}\) where \(\mR = \diag(\Var(r_1),\dots,\Var(r_n))\). Let \(\Lambda_1 = \sum_{k=1}^n \lambda_k^2(\Sigma_r)\) and \(\Lambda_2 = \sum_{k=2}^n\lambda_k^2(\Sigma_r)\). By (ref) and (ref), \(\Lambda_2/\Lambda_1\) is bounded away from zero. Using (ref) we can then bound for some constant \(C > 0\) \begin{equation} \begin{split} \Pr(\|r'H\|^{1/2} \leq \delta_n) \leq C\delta_n\Lambda_1^{-1} \end{split} \end{equation} Combining Bounds. To combine the bounds in (ref) and (ref), first write \[ \Var(\tilde r'\bar H\tilde r) = 2\trace(\mR\bar H\mR \bar H) + 4\mu_r\bar H \mR \bar H \mu_r \] for \(\mu_r = \E[r]\). Using the fact that \(\bar H^{1/2}\mR \bar H^{1/2}\) is symmetric positive definite we can bound: \begin{align*} \mu_r'\bar H \mR\bar H \mu_r &= (\mu_r'\bar H^{1/2})'(\bar H^{1/2}\mR\bar H^{1/2})(\bar H^{1/2}\mu_r) \\ &\leq \lambda_1(\bar H^{1/2}\mR\bar H^{1/2})\|\mu_r'\bar H^{1/2}\|^2 \\ &= \sqrt{\lambda_1^2(\bar H^{1/2}\mR\bar H^{1/2})}\|\mu_r'\bar H^{1/2}\|^2\\ &= \sqrt{\lambda_1(\bar H^{1/2}\mR\bar H \mR \bar H^{1/2})}\|\mu_r'\bar H^{1/2}\|^2\\ &\leq \sqrt{\trace(\bar H^{1/2}\mR\bar H \mR \bar H^{1/2})}\|\mu_r'\bar H^{1/2}\|^2\\ &= \sqrt{\trace(\mR\bar H \mR \bar H)}\|\mu_r'\bar H\|^2 \leq c^2\Lambda_1^{1/2}\numberthis \end{align*} where the first equality uses the symmetric square root of \(\bar H\), the first inequality comes from Courant-Fischer minmax principle and the third equality uses the fact that the eigenvalues of \(A^2\) are the squares of the eigenvalues of \(A\), for any generic symmetric matrix \(A\). The second inequality comes from the fact that a matrix times its transpose is always positive semidefinite and that for \(M\) psd, \(\lambda_1(M) \leq \sqrt{\trace(M^2)}\) since the trace is the sum of the (weakly positive) eigenvalues. The final inequality uses \(\mu_r'\bar H \mu_r = \frac{c}{n}\sum_{i=1}^n(\E[\tilde\Pi_i])^2 \leq \frac{c}{n}\sum_{i=1}^n \E[(\tilde\Pi_i)^2] \leq c^2 \). Combining (ref), (ref), and (ref) gives us \begin{equation} \Pr(\tilde D \leq \delta_n) \leq C\min\left\{\Lambda_1 + \Lambda_1^{1/2}, \delta_n \Lambda_1^{-1} \right\} \end{equation} Regardless of the behavior of \(\Lambda_1\), this tends to zero as \(\delta_n\to 0\).
remark[Final Anticoncentration Bound] To give an explicit bound on (ref) in terms of \(\delta_n\) we note that, if \(x^\star\) solves \[ x^\star + \sqrt{x^\star} = \frac{c}{x^\star} \] then for any \(x \geq 0\), \(\min\{x + \sqrt{x}, c/x\} \leq x^\star + \sqrt{x^\star}\). Using this, notice that \((x^\star)^2 + (x^\star)^{3/2} = c\) so that \(x^\star \leq \sqrt{c}\). This allows us to bound (ref) \begin{align*} \Pr(\tilde D \leq \delta_n) \leq C\min\{\Lambda_1 + \Lambda_1^{1/2},\delta_n\Lambda_1^{-1}\} \leq C(\delta_n^{1/2} + \delta_n^{1/4}) \end{align*}
lemmaLet \(X_n\) and \(Y_n\) be two sequences of random variables and let \(W_n = X_n/Y_n\). Then for any \(c \in \SR\) and any \(\delta > 0\): \begin{align*} \Pr(0 \leq X_n - cY_n \leq \delta) &\leq \Pr(c \leq W_n \leq \delta^{1/2} + c) + \Pr(Y_n \leq \delta^{1/2}) \intertext{and} \Pr( - \delta \leq X_n - cY_n \leq 0) &\leq \Pr( c - \delta^{1/2} \leq W_n \leq c) + \Pr(Y_n \leq \delta^{1/2}) \end{align*}
proofDefine the event \(\Omega = \{Y_n \geq \delta^{1/2}\}\). We can bound \begin{align*} \Pr(0 \leq X_n - cY_n \leq \delta) &= \Pr(cY_n \leq X_n \leq \delta + cY_n) \\ &\leq \Pr(\{cY_n \leq X_n \leq \delta + cY_n\} \cap \Omega) + \Pr(\Omega^c) \\ &= \Pr(\{c \leq W_n \leq \delta/Y_n + c\}\cap\Omega) + \Pr(\Omega^c) \\ &\leq \Pr(c \leq W_n \leq \delta^{1/2} + c) + \Pr(\Omega^c) \end{align*} The second statement of the lemma follows symmetrically.
lemma[] Suppose that \(X_n\) and \(Y_n\) are sequences of (real-valued) random variables such that \(Y_n = O_p(1)\) and for any \(x \in \SR\) \[ |\Pr(X_n \leq x) - \Pr(Y_n \leq x)| \to 0 \] Then \(X_n = O_p(1)\).
proofPick any \(\eps > 0\), and let \(M_{\eps/2}\) be such that \(\Pr(Y_n > M_{\eps /2}) \leq \eps/2 \) for all \(n \geq N_\eps\). In addition, let \(\tilde N_\eps\) be such that \(|\Pr(X_n \leq M_{\eps/2}) - \Pr(Y_n \leq M_{\eps/2})| \leq \eps/2\) for all \(n \geq \tilde N_\eps\). Then for all \(n \geq N_\eps \vee \tilde N_{\eps/2}\), \begin{align*} \Pr(X_n > M_{\eps/2}) &\leq \Pr(Y_n > M_{\eps/2}) + |\Pr(X_n > M_{\eps/2}) - \Pr(Y_n > M_{\eps/2})| \\ &\leq \eps/2 + |\Pr(Y_n \leq M_{\eps/2}) - \Pr(X_n \leq M_{\eps/2})| \\ &\leq \eps/2 + \eps/2 = \eps \end{align*}
lemma[] Suppose that \(X_n\) and \(Y_n\) are sequences of (real-valued) random variables such that \(Y_n = O_p(1)\) and for any \(\Delta \in \SR\) \[ \sup_{x \leq \Delta}|\Pr(X_n \leq x) - \Pr(Y_n \leq x)| \to 0 \] Then \(\sup_{x\in\SR}|\Pr(X_n \leq x) - \Pr(Y_n \leq x)| \to 0\).
proofPick an \(\eps > 0\). By (ref), \(X_n = O_p(1)\). Pick a constant \(M_{\eps/3}\) such that \(\Pr(X_n > M_{\eps/3}) \leq \eps/3\) and \(\Pr(Y_n > M_{\eps/3}) \leq \eps/3\). Then for any \(x \in \SR\) we can bound \(|\Pr(X_n \leq x) - \Pr(Y_n \leq x)|\) by considering two cases: Case 1. If \(x \leq M_{\eps /3}\), then, \begin{equation} |\Pr(X_n \leq x) - \Pr(Y_n \leq x)| \leq \sup_{x \leq M_{\eps / 3}}|\Pr(X_n \leq x) - \Pr(Y_n \leq x)| \end{equation} by hypothesis, there is an \(N_{\eps}\) such that for \(n \geq N_{\eps}\) the RHS of (ref) is less than \(\eps\). Case 2. If \(x > M_{\eps/3}\) we can bound \begin{align*} |\Pr(X_n \leq x) - \Pr(Y_n \leq x)| &\leq |\Pr(X_n \leq M_{\eps/3}) - \Pr(Y_n \leq M_{\eps/3})| \\ &\;\;\;+ |\Pr(M_{\eps/3} < X_n \leq x) - \Pr(M_{\eps/3} < Y_n \leq x)| \\ &\leq |\Pr(X_n \leq M_{\eps/3}) - \Pr(Y_n \leq M_{\eps/3})| + \eps/3 + \eps/3 \numberthis \end{align*} By hypothesis, there is an \(N_{\eps/3}\) such that \(|\Pr(X_n \leq M_{\eps/3}) - \Pr(Y_n \leq N_{\eps/3})| \leq \eps/3\). WLOG \(N_{\eps/3} \geq N_\eps\). Combining the bounds in (ref) and (ref), for any \(n \geq N_{\eps/3}\) and any \(x \in \SR\), \[ |\Pr(X_n \leq x) - \Pr(Y_n \leq x)| \leq \eps \] Since this holds for all \(x\), this gives the result.
lemma[Approximate Distribution] Under (ref) and the conditions of (ref) \[ \sup_{a \in \SR}|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)|\to 0 \]
proof[Proof of (ref)] First, fix a \(\Delta \geq 0\) and consider any \(a \leq \Delta\). As in (ref), let \(\tilde \varphi(\cdot): \SR \to \SR\) be three times continuously differentiable with bounded derivatives up to the third order such that \(\tilde\varphi(x)\) is 1 if \(x \leq 0\), \(\tilde\varphi(x)\) is decreasing if \(x \in (0,1)\), and \(\tilde\varphi(x)\) is zero if \(x \geq 1\). Consider a sequence \(\gamma_n \searrow 0\) slowly enough such that \((\gamma_n^{-2} + \gamma_n^{-3})/\sqrt{n} \to 0\) and define \(\varphi_n(x) = \tilde\varphi(\frac{x}{\gamma_n})\). By (ref) we can write for some constant \(M\) that depends only on \(\Delta\): \begin{align*} \Pr(\JK_I(\beta_0) \leq a) = \Pr(\JK^a \leq 0) &\leq \E[\varphi_n(\JK^a)] \\ &\leq \E[\varphi_n(\tilde \JK^a)] + \frac{M}{\sqrt{n}}(\gamma_n^2 + \gamma_n^{-3}) \\ &\leq \Pr(\tilde\JK^a \leq 0) + \Pr(0 \leq \tilde N^2 - a \tilde D \leq \gamma_n) + \frac{M}{\sqrt{n}}(\gamma_n^2 + \gamma_n^{-3}) \intertext{Applying (ref) and \(\{\tilde\JK^a \leq 0\} = \{\JK_G(\beta_0) \leq a\}\) gives:} &\leq \Pr(\JK_G(\beta_0) \leq a) + \underbrace{\Pr(a \leq \tilde N^2/\tilde D \leq a + \gamma_n^{1/2})}_{\vA} \\ &\;\;\;+ \underbrace{\Pr(\tilde D \leq \gamma_n^{1/2})}_{\vB}+ \frac{M}{\sqrt{n}}(\gamma_n^{-2} + \gamma_n^{-3}) \end{align*} By (ref), we can bound \(\vA \leq M\gamma_n^{1/2}\) while by (ref) and (ref), \(\vB \leq M \gamma_n^{1/4}\). Since \(\gamma_n\) is chosen such that \(\frac{M}{\sqrt{n}}(\gamma_n^{-2} + \gamma_n^{-3}) \to 0\) we can conclude that \(\Pr(\JK_I(\beta_0) \leq a) \leq \Pr(\JK_G(\beta_0) \leq a) + o(1)\). A symmetric argument with \(\varphi_n(x) = \tilde\varphi(1 - \frac{x}{\gamma_n})\) gives a lower bound so that, in total \[ \Pr(\JK_G(\beta_0) \leq a) - \ve \leq \Pr(\JK_I(\beta_0) \leq a) \leq \Pr(\JK_G(\beta_0) \leq a) + \ve \] where \[ \ve = M\big(\frac{\gamma_n^{-2} + \gamma_n^{-3}}{\sqrt{n}} + \gamma_n^{1/2} + \gamma_n^{1/4}\big) = o(1) \] Since the constant M depends only on \(\Delta\), this gives us that for any fixed \(\Delta > 0\) \begin{equation} \sup_{a \leq \Delta}\big|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big| \leq C\big(\frac{\gamma_n^{-2} + \gamma_n^{-3}}{\sqrt{n}} + \gamma_n^{1/2} + \gamma_n^{1/4}\big) = o(1) \end{equation} where \(C\) is a constant that depends only on \(\Delta\). Noting that the numerator \(\JK_G(\beta_0)\) is \(O_p(1)\) under (ref) while the inverse of the denominator of \(\JK_G(\beta_0)\) is \(O_p(1)\) by (ref), we can apply (ref). This step shows that the result in (ref) implies that the approximation error tends to zero uniformly over the real line, which is the desired result. Optimizing over \(\gamma_n\) in the expression of (ref) yields the rate of decay in (ref).

Proof of (ref)

proof[Proof of (ref)] For \(N\) and \(D\) defined at the top of (ref) define \(\widehat N = N + \Delta_N\) and \(\widehat D = D + \Delta_D\). We can then write \(\JK(\beta_0) = \widehat N^2 / \widehat D\) and rewrite \[ \JK(\beta_0) - \JK_I(\beta_0) = \frac{2ND\Delta_N + D\Delta_N - N^2\Delta_D}{D^2 + D\Delta_D} \] Apply (ref) to see that \(N^2 = O_p(1)\) while under (ref), \(D = O_p(1)\). Thus, \(2ND\Delta_n + D\Delta_n - N^2\Delta_D = o_p(1)\). Meanwhile, by (ref), \(\Pr(D^2 \leq \delta_n) \to 0\) for any sequence \(\delta_n \to 0\). Apply (ref) to obtain that \(|\JK(\beta_0) - \JK_I(\beta_0)|\to_p 0\). Finally, apply (ref) with \(X_n = \JK(\beta_0)\), \(Y_n = \JK_I(\beta_0)\) and \(Z_n = \JK_G(\beta_0)\). The density of \(Z_n\) is uniformly bounded by (ref) to show that the distribution of \(\JK(\beta_0)\) may be uniformly approximated by the distribution of \(\JK_G(\beta_0)\).
lemma[] Let \(A_n, B_n\) and \(Y_n\) be sequences of random variables such that \(A_n = o_p(1)\) and \(B_n = o_p(1)\). If \(Y_n\) is such that for any sequence \(\delta_n \to 0\), \(\Pr(|Y_n| \leq \delta_n)\to 0\), then, \[ \bigg|\frac{A_n}{Y_n + B_n}\bigg| = o_p(1) \]
proofFix any \(\eps > 0\). We show that \[ \bigg|\frac{A_n}{Y_n + B_n}\bigg| \leq \eps \] on an intersection of events whose probability tends to one. By (ref) there is a sequence \(\eps_n \searrow 0\) such that \[ \Pr(|A_n| \leq \eps_n) \to 1 \andbox \Pr(\eps|B_n| \leq \eps_n)\to 1 \] Consider the intersection of events \(\Omega_1 \cap\Omega_2\cap\Omega_3\) where \[ \Omega_1 := \{\eps|Y_n| \geq 2\eps_n\},\;\;\Omega_2 := \{\eps|B_n| \leq \eps_n\},\;\;\Omega_3 := \{|A_n| \leq \eps_n\} \] By assumption, \(\Pr(\Omega_1\cap\Omega_2\cap\Omega_3) \to 1\). On this event \(|Y_n + B_n| \geq \eps_n/\eps > 0\) and \(|A_n| \leq \eps_n\) so that \(|A_n/(Y_n + B_n)| \leq |\eps_n /(\eps_n/\eps)| \leq \eps\).
lemma[Denominator Interpolation] Suppose that the moment bounds of (ref) and (ref) hold. Let \(\varphi(\cdot):\SR \to \SR\) be such that \(\varphi(\cdot) \in C_b^3(\SR)\) with \(L_2(\varphi) = \sup_x|\varphi''(x)|\) and \(L_3(\varphi) = \sup_x|\varphi'''(x)|\). Then there is a constant \(M\) that depends only on the constant \(c\) such that: \[ |\E[\varphi(D) - \varphi(\tilde D)]| \leq \frac{M}{\sqrt{n}}(L_2(\varphi) + L_3(\varphi)) \]
proof[Proof of (ref)] We inherit the definitions of \(D_{-i}\), \(\Delta_{2i}^a\), \(\Delta_{2i}^b\), \(\tilde\Delta_{2i}^a\), and \(\tilde\Delta_{2i}^b\) from the proof of (ref) with \(a = 1\). Then, as before we can write \begin{align*} \E[\varphi(D) - \varphi(\tilde D)] &= \sum_{i=1}^n \E[\varphi(D_{-i} + n^{-1}\Delta_{2i}^a + n^{-1}\Delta_{2i}^b)] \\ &\hphantom{=\sum_{i=1}^n }-\E[\varphi(D_{-i} + n^{-1}\tilde\Delta_{2i}^a + n^{-1}\tilde\Delta_{2i}^b)] \end{align*} We examine each term via a second-order Taylor expansion around \(D_{-i}\) \begin{align*} \E[Term_i] &= \frac{1}{n}\E[\varphi'(D_{-i})\{(\Delta_{2i}^a - \tilde\Delta_{2i}^a) + (\Delta_{2i}^b - \tilde\Delta_{2i}^b)\}] \\ &+ \frac{1}{2n^2}\E[\varphi”(D_{-i})\{((\Delta_{2i}^a)^2 - (\tilde\Delta_{2i}^a)^2) + 2(\Delta_{2i}^a \Delta_{2i}^b - \tilde\Delta_{2i}^a \tilde\Delta_{2i}^b) + ((\Delta_{2i}^b)^2 - (\Delta_{2i}^b)^2)\}] \\ &+ R_i + \tilde R_i \end{align*} where \(R_i\) and \(\tilde R_i\) are remainder terms to be analyzed later. Using the restrictions in (ref) we can simplify the above display: \begin{align*} \E[Term_i] &= \underbrace{0.5n^{-2}\E[\varphi”(D_{-i})((\Delta_{2i}^a)^2 - (\tilde\Delta_{2i}^a)^2)]}_{\dot\vA_i} + \underbrace{n^{-2}\E[\varphi”(K_{-i})(\Delta_{2i}^a \Delta_{2i}^b - \tilde\Delta_{2i}^a \tilde\Delta_{2i}^b)}_{\dot\vB_i} \\ &+ R_i + \tilde R_i \end{align*} Using (ref) we can bound \begin{align*} |\vA_i| &\leq \frac{M}{n^2}L_2(\varphi) & |\vB_i| &\leq \frac{M}{n^{3/2}}L_2(\varphi) \end{align*} For some \(\bar D_{1i}\) and \(\bar D_{2i}\) we can express \begin{align*} R_i &= \E[\varphi”'(\bar D_{1i})\{n^{-1}\Delta_{2i}^a + \Delta_{2i}^b\}^3] \leq \frac{M}{n^{3/2}}L_3(\varphi) + \frac{M}{n^3}L_3(\varphi)\\ R_i &= \E[\varphi”'(\bar D_{2i})\{n^{-1}\tilde\Delta_{2i}^a + \tilde\Delta_{2i}^b\}^3] \leq \frac{M}{n^{3/2}}L_3(\varphi) + \frac{M}{n^3}L_3(\varphi) \end{align*} where the inequalities again come from applications of (ref). Combining these bounds and summing over the \(n\) terms gives the result.
lemma[Denominator anti-concentration] Suppose that the moment bounds of (ref) and (ref) hold. Then, for any sequence \(\delta_n \searrow 0\), \[ \Pr(D \leq \delta_n) \to 0 \]
proof[Proof of (ref)] Let \(\tilde \varphi(\cdot): \SR \to \SR\) be three times continuously differentiable with bounded derivatives up to the third order such that \(\tilde\varphi(x)\) is 1 if \(x \leq 0\), \(\tilde\varphi(x)\) is decreasing if \(x \in (0,1)\), and \(\tilde\varphi(x)\) is zero if \(x \geq 1\). Consider a second sequence \(\gamma_n \searrow 0\) slowly enough such that \((\gamma_n^{-2} + \gamma_n^{-3})/\sqrt{n} \to 0\). Take \(\varphi_n(x) = \tilde\varphi(\frac{x- \delta_n}{\gamma_n})\). By (ref) and since \(\tilde\varphi(\cdot)\) has bounded derivatives up to the third order, there is a fixed constant \(M_1 > 0\) that depends only on \(c\) such that \begin{align*} \Pr(D \leq \delta_n) &\leq \Pr(\tilde D \leq \delta_n + \gamma_n) + \frac{M_1}{\sqrt{n}}(\gamma_n^{-2} + \gamma_n^{-3}) \end{align*} Let \(\gamma_n\) be a sequence tending to zero such that \((\gamma_n^{-2} + \gamma_n^{-3})/\sqrt{n} \to 0\) and conclude by applying (ref).
lemma[] Let \(X_n\), \(Y_n\), and \(Z_n\) be sequences of random variables such that \(|X_n - Y_n| \to_p 0\), the distribution of \(Z_n\) is absolutely continuous with respect to Lebesgue measure and the density functions of \(Z_n\) are uniformly bounded and \(\sup_{a\in\SR}|\Pr(Y_n \leq a) - \Pr(Z_n \leq a)|\to 0.\) Then \(\sup_{a \in \SR}|\Pr(X_n \leq a) - \Pr(Z_n \leq a)|\to 0 .\)
proofFor any \(a \in \SR\) and \(\eps > 0\) we have that \(\{X_n \leq a\} \subseteq \{Y_n \leq a + \eps\} \cup \{|X_n - Y_n| > \eps\}\); thus, by applying union bound and rearranging we obtain: \begin{align*} \Pr(X_n \leq a) &\leq \Pr(Y_n \leq a + \eps) + \Pr(|Y_n - X_n| > \eps) \\ &\leq \Pr(Z_n \leq a + \eps) + |\Pr(Y_n \leq a + \eps) - \Pr(Z_n \leq a + \eps)| \\ &\;\;\;\;\;\;\;\;\;+ \Pr(|Y_n - X_n| > \eps) \intertext{so that} \Pr(X_n \leq a) - \Pr(Z_n \leq a) &\leq \Pr(a < Z_n \leq a+\eps) + |\Pr(Y_n \leq a + \eps) - \Pr(Z_n \leq a + \eps)|\\ &\;\;\;\;\;\;\;\;\;+ \Pr(|Y_n - X_n| >\eps) \end{align*} Let \(\eps_n \to 0\) be a sequence tending to zero such that \(\Pr(|X_n - Y_n| > \eps_n) \to 0\) ((ref)). Applying a supremum to the above display yields \begin{align*} \sup_{a\in \SR} \Pr(X_n \leq a) - \Pr(Z_n \leq a) &\leq \sup_{a\in \SR} \Pr(a < Z_n \leq a + \eps_n) \\ &\;+ \sup_{a \in \SR} |\Pr(Y_n \leq a + \eps_n) - \Pr(Z_n \leq a + \eps_n)|\\ &\;\;+ \Pr(|Y_n - X_n| >\eps_n) \end{align*} The first term goes to zero as \(\eps_n \to 0\) since \(Z_n\) has a uniformly bounded density; the second term goes to zero by \(\sup_{a \in \SR} |\Pr(Y_n \leq a) - \Pr(Z_n \leq a)| \to 0\) and the third term goes to zero by definition of \(\eps_n\) and \(|Y_n - X_n| \to_p 0\). We can apply a symmetric argument to show that \(\sup_{a\in\SR}\Pr(Z_n \leq a) - \Pr(X_n \leq a) \leq o(1)\) which completes the claim of the lemma.

Proof of (ref)

proof[Proof of (ref)] As at the top of (ref), recall that \(\tilde h_{ii} = 0\), and define \begin{align*} N &= \frac{1}{\sqrt{n}}\sum_{i=1}^n \eps_i(\beta_0)\sum_{j=1}^n \tilde h_{ij} r_j & D &= \frac{1}{n}\sum_{i=1}^n \eps_i^2(\beta_0)(\sum_{j=1}^n \tilde h_{ij} r_j)^2 \end{align*} where \(\tilde h_{ij} = s_n h_{ij}\). A primary goal is to show that tests based on the infeasible statistic, \(\JK_I(\beta_0)\), are consistent. That is, \(\Pr(\JK_I(\beta_0) \leq a) \to 0\) for any fixed \(a \in \SR_+\). The event \(\{\JK_I(\beta_0) \leq a\}\) is equivalently expressed \(\{N^2 - a D \leq 0\}\) so that \(\Pr(\JK(\beta_0) \leq a) = \Pr(N^2 - aD \leq 0)\). Under the moment bounds of (ref) and (ref), \(aD = O_p(1)\) so by (ref) it suffices to show that \(\Pr(|N| \leq M) \to 0\) for any fixed \(M \geq 0\). By assumption \(P = \E[N^2] \to \infty\) so we move to show that \(\Var(N) = O(1)\) and then apply (ref) to conclude. To this end, recall the definition of \(\eta_i = \eps_i(\beta_0) - \E[\eps_i(\beta_0)]\), define \(\mu_i= \E[\eps_i(\beta_0)] = \Pi_i(\beta - \beta_0)\), and let \begin{align*} N_1 &\coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \eta_i \sum_{j=1}^n \tilde h_{ij} r_j & N_2 &\coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n \mu_i \sum_{j=1}^n \tilde h_{ij} r_j \end{align*} Notice that \(N = N_1 + N_2\). To show that \(\Var(N_1) = O(1)\), define \(\va_i = \eta_i \sum_{j=1}^n \tilde h_{ij} r_j\). Since \(\E[\eta_i r_i] = 0\), we have that \(\Cov(\va_i,\va_j) = 0\) for \(i \neq j\). Thus, \[ \Var(N_1) = \Var(\sum_{i=1}^n \va_i/\sqrt{n}) = n^{-1}\sum_{i=1}^n \Var(\va_i) = n^{-1}\sum_{i=1}^n \Var(\eta_i)\E[(\sum_{j=1}^n \tilde h_{ij} r_j)^2] \leq c^2 \] where the final inequality follows from the upper bound on \(\Var(\eta_i)\) and by definition of \(\tilde h_{ij} = s_n h_{ij}\) from (ref). To show that \(\Var(N_2) = O(1)\) let \(\vb_i = \sum_{j=1}^n \tilde h_{ji}\tilde\Pi_j(\beta - \beta_0)\) and rewrite \(N_2 = \frac{1}{\sqrt{n}}\sum_{i=1}^n r_i \vb_i\). Under (ref)(ii), \(|\vb_i| = |\E[\sum_{j=1}^n \tilde h_{ji}\eps_j(\beta_0)]| \leq c^{1/2}\), so we can bound \[ \Var(N_2) = \Var(\sum_{i=1}^n r_i \vb_i /\sqrt{n}) = n^{-1}\sum_{i=1}^n \vb_i^2 \Var(r_i) \leq c^2 \] Since \(\Var(N) \leq 2\Var(N_1) + 2\Var(N_2)\), we can conclude that tests based on \(\JK_I(\beta_0)\) are consistent. Finally, we want to show that this fact, along with \((\Delta_N,\Delta_D)\to_p0\) implies that tests based on \(\JK(\beta_0)\) are consistent. To do this, notice that we can write \[ \JK(\beta_0) = \frac{(N + \Delta_N)^2}{D + \Delta_D} \] and thus that \(\JK(\beta_0) \leq a\) if and only if \[ \widehat\JK_a \coloneqq (N + \Delta_N)^2 - a(D + \Delta_D) = N^2 - aD + 2N\Delta_N + \Delta_N^2 - a\Delta_D \leq 0. \] Define \(\JK_a = N^2 - aD\). Using that \(\{\widehat\JK_a \leq 0\} \subseteq \Big\{\frac{\widehat\JK_a}{\JK_a}\JK_a \leq 0\Big\} \cup \big\{\JK_a \leq 0\}\) we can write \begin{align*} \Pr(\widehat\JK_a \leq 0) &\leq \Pr\Big(\frac{\widehat\JK_a}{\JK_a} \JK_a\leq 0\Bigg) + \Pr(\JK_a \leq 0) \\ &\leq 2\Pr(\JK_a \leq 0) + \Pr\Big(\frac{\widehat\JK_a}{JK_a} \leq \frac{1}{2}\Big) \end{align*} By consistency of the test based on the infeasible \(\JK_I(\beta_0)\) statistic, we have that \(\Pr(\JK_a \leq 0)\to 0\). Thus, it only remains to show that \(\Pr(\widehat\JK_a /\JK_a \leq 1/2) \to 0\). This, in turn, follows if \[ \frac{\widehat\JK_a - \JK_a}{\JK_a} = \frac{2N\Delta_N + \Delta_N^2 - a\Delta_D}{N^2 - aD}\to_p 0. \] The above results can be used to show that \(\Pr(|N^2 - aD| \leq \delta_n) \to 0\) for any sequence \(\delta_n \searrow 0\) so that \(\frac{1}{\JK_a} = O_p(1)\). Combined with \((\Delta_N,\Delta_D)\to_p0\) this implies that \(\{\Delta_N^2 - a\Delta_D\}/\JK_a \to_p0\). What remains is to show that \(2N\Delta_N/(N^2 - aD) \to_p 0\). Write \[ \frac{2N\Delta_N }{N^2 - aD} = \frac{\frac{2N\Delta_N}{N^2}}{1 - a\frac{D}{N^2}}. \] Since \(D = O_p(1)\) while \(\Pr(N^2 \leq M) \to 0\) for any \(M\) we have that \(D/N^2 \to_p 0\). Moreover, \(\Pr(|N| \leq M) \to 0\) for any fixed M implies \(N/N^2 = O_p(1)\) so that \(2N\Delta_N/N^2 \to_p0\). We can apply continuous mapping theorem to conclude.
lemma[] Suppose that \(X_n\) is a sequence of random variables such that \(\E[X_n^2] \to \infty\) while \(\Var(X_n) = O(1)\). Then, for any \(M \geq 0\), \(\Pr(|X_n| \leq M) \to 0\).
proofFirst, note that \(\Var(|X_n|) \leq \Var(X_n)\) so \(\Var(|X_n|) = O(1)\). Moreover \(\Var(|X_n|) = \E[X_n^2] - (\E[|X_n|])^2\), so \(\E[X_n^2] \to \infty\) and \(\Var(|X_n|) = O(1)\) implies that \(\E[|X_n|] \to \infty\). Then, \begin{align*} \Pr(|X_n| \leq M) &= \Pr(|X_n| - \E[|X_n|] \leq M - \E[|X_n|]) \\ &= \Pr(\E[|X_n|] - |X_n| \geq \E[|X_n| - M) \\ &\leq \Pr(|\E[|X_n|] - |X_n|| \geq \E[|X_n|] - M) \\ &\leq \frac{\Var(|X_n|)}{\E[|X_n|] - M} \end{align*} Since \(\Var(|X_n|) = O(1)\) but \(\E[|X_n|] \to \infty\), this tends to zero.
lemma[] Suppose that \(X_n\) and \(Y_n\) are random variables such that \(Y_n = O_p(1)\) and, for any \(M \geq 0\), \(\Pr(|X_n|\leq M) \to 0\). Then, for any \(M_1 \geq 0\), \(\Pr(X_n^2 - Y_n \leq M_1) \to 0\).
proofPick any \(\eps > 0\). We want to show that, eventually, \(\Pr(X_n^2 - Y_n > M_1) \geq 1- \eps\). Since \(Y_n = O_p(1)\), there is a fixed constant \(M_Y\) such that \(\Pr(|Y_n| \leq M_Y) \geq 1-\eps/2\). Since \(\Pr(|X_n| \leq M)\to 0\) for any \(M \geq 0\), there exists an \(N_X\) such that, for \(n \geq N_X\), \(\Pr(X_n^2 \leq M_1 + M_Y) \leq \eps/2\). A union bound completes the argument (on the eventuality \(n \geq N_X\)): \begin{align*} \Pr(X_n^2 - Y_n > M) &\geq \Pr(X_n^2 > M_1 + M_Y, |Y_n| \leq M_Y)\\ &= 1 - \Pr(\{X_n^2 < M_1 + M_Y\} \cup \{|Y_n| > M_Y\}) \\ &\geq 1 - \eps/2 - \eps/2 = 1-\eps \end{align*}

Proof of (ref)

I provide the proof for the case where \(d_x = 1\) with the general case following from symmetric logic. For any \(j = 1,\dots, d_b\) define the matrix \(B_j = \diag(b_j(z_1),\dots,b_j(z_n))\) and collect observations \(\eps(\beta_0) = (\eps_1(\beta_0),\dots,\eps_n(\beta_0))'\in \SR^n\), \(r = (r_1,\dots,r_n)'\in\SR^n\), \(\hat r = (\hat r_1,\dots, \hat r_n)'\in\SR^n\), and \(\xi = (\xi_1,\dots,\xi_n)' \in \SR^n\). In addition, collect \(b_\eps = (b_{\eps1},\dots,b_{\eps n})\in \SR^{d_b\times n}\) where \(b_{\eps i} = \eps_i(\beta_0)b(z_i) \in \SR^{d_b}\). Finally, let \(\mH = \frac{s_n}{\sqrt{n}}H\), \(\tilde H = s_n H\) and \(\tilde h_{ij} = s_n h_{ij}\).

Step 1: \(\Delta_N\to_p 0\). To show that \(\Delta_N \to_p 0\) write

align*[align* omitted — 324 chars of source]

To bound \(\vA\) we move to apply (ref) to the quadratic form \(\eps(\beta_0)'(\mH B_j)\eps(\beta_0)\). First notice that, under (ref)(v), we have \[ \|\E[\mH b_j\eps(\beta_0)]\|_2 = \frac{1}{n}\sum_{i=1}^n (\E[s_n\sum_{j\neq i} h_{ij} b(z_j)\eps_j(\beta_0)])^2 \leq c^2 \] In the notation of (ref) this give us an upper bound on \(\|\E f^{(1)}(X)\|_{\text{HS}}\). Next, (ref) gives us that the Frobenius norm of \(\mH = \frac{s_n}{\sqrt{n}}\mH\) is bounded, since the rows of \(s_n H\) are square summable, \(\sum_{j\neq i} (s_n h_{ij})^2 \leq c\) for all \(i = 1,\dots, n\). In the notation of (ref) this gives us an upper bound on \(\|\E f^{(2)}(X)\|_{\text{HS}}\). Applying (ref) and a union bound then gives us that

equation[equation omitted — 189 chars of source]

Since \(\max_{1 \leq j \leq d_b}|\E[\eps(\beta_0)'\mH B_j\eps(\beta_0)]| \leq c\) under (ref)(v), (ref) gives that \[ \max_{1 \leq j \leq d_b} |\eps(\beta_0)'\mH B_j \eps(\beta_0)| = O_p(\log^{2/a}(d_b)) \] Since \(\log^{2/a}(d_b)\|\widehat\gamma - \gamma\|_1 \to_p 0\) by assumption, this yields that \(\vA \to_p 0\).

To bound \(\vB\) see that \(\|\eps(\beta_0)'\mH\|_2 = \frac{s_n^2}{n}\sum_{i=1}^n (\sum_{j\neq i} h_{ij}\eps_i(\beta_0))^2 = O_p(1)\) under (ref)(ii) while under (ref) \(\|\xi\|_2 = o(1)\).

Step 2: \(\Delta_D\to_p 0\). Notice that \(a^2 - b^2 = 2b(a - b) + (a - b)^2\) and bound:

align*[align* omitted — 336 chars of source]

Since both \(\vE = O_p(1)\) and \(\vF = O_p(1)\) under the moment bounds of (ref) and (ref), it suffices to show that \[ \max_i |\sum_{j\neq i} \tilde h_{ij} (\widehat r_j - r_j)| \to_p 0 \] To do so write

align*[align* omitted — 379 chars of source]

To bound \(\vA\), note that by (ref)(v) \(\max_{i,j}|\E[\sum_{j\neq i} \tilde h_{ij} b(z_j) \eps_j(\beta_0)| \leq c\). Under (ref)(ii), \(\max_{i,j} \sum_{j\neq i} \tilde h_{ij}^2 b^2(z_j) \leq c^2\) so we can apply (ref) and a union bound to obtain that \[ \max_{\substack{1 \leq i \leq n \\ 1 \leq j \leq d_b}} \big|\sum_{j\neq i} \tilde h_{ij} b(z_j)\eps_j(\beta_0)\big| = O_p(\log^{1/a}(d_bn)) \] Along with the implied rate on \(\|\hat\gamma - \gamma\|_1\) from (ref)(iv) this shows that \(\vA \to_p 0\).

To show that \(\vB \to 0\), use Cauchy-Schwarz, \(\sum_{j\neq i} \tilde h_{ij}^2b^2(z_j) \leq c\) for any \(i,j\) by (ref)(ii), and \(\sum_{i=1}^n \xi_i^2 = o(1)\) by (ref)(iii).

Proof of (ref)

Throughout this section, define the scaled elements of the infeasible and gaussian numerators and denominators

align*[align* omitted — 515 chars of source]

Collect these in \(N = (N_1,\dots N_{d_x})' \in \SR^{d_x}\), \(\tilde N = (\tilde N_1,\dots, \tilde N_{d_x})' \in \SR^{d_x}\), \(D = [D_{\ell k}]_{\ell, k \in [d_x]} \in \SR^{d_x \times d_x}\), and \(\tilde D = [\tilde D_{\ell k}]_{\ell, k \in [d_x]} \in \SR^{d_x \times d_x}\). After multiplying by scaling matrix \(\diag(s_{1,n},\dots,s_{d_x,n})\) and the inverse of the scaling matrix we rewrite the infeasible and gaussian test statistics

align*[align* omitted — 132 chars of source]

As with (ref), the result (ref) follows directly from combining the following lemmas. The first is the main technical lemmma, and shows that the distribution of the infeasible statistic \(\JK_I(\beta_0)\) can be uniformly approximated by that of \(\JK_G(\beta_0)\). The proof of this technical lemma is involved and deferred to (ref). The second lemma establishes that estimation error can be treated as negligible. As with (ref), the main difficulty hear is in dealing with the fact that neither the numerator vector nor denominator matrix of the \(\JK(\beta_0)\) statistic may have stable limiting distributions.

lemma[] Suppose that (ref) hold as well as the moment conditions of (ref). Then, \[ \sup_{a\in\SR}\big|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big| \to 0. \]
proof[Proof of (ref)] (ref) follows as a consequence of the joint gaussian approximation with the combination statistic established in (ref).
lemma[] Suppose that (ref) hold along with the moment conditions of (ref). Then, if \((\Delta_N, \Delta_D) \to_p0\), \(\big|\JK(\beta_0) -\JK_I(\beta_0)\big|\to_p0\).
proof[Proof of (ref)] Define the matrix \(\Delta_{D} = [(\Delta_{D})_{\ell k}]_{\ell,k\in [d_x]}\) and the vector \(\Delta_N = [(\Delta_N)_{\ell}]_{\ell \in [d_x]}\) where \begin{align*} (\Delta_D)_{\ell k} &\coloneqq \frac{s_{\ell,n}s_{k,n}}{n}\sum_{i=1}^n \eps_i^2(\beta_0)\big(\widehat\Pi_{\ell, i} \widehat\Pi_{k,i} - \widehat\Pi_{\ell,i}^I\widehat\Pi_{k,i}^I\big) \\ (\Delta_N)_\ell &\coloneqq \frac{s_{\ell, n}}{\sqrt{n}}\sum_{i=1}^n \eps_i(\beta_0)(\widehat\Pi_{\ell,i} - \widehat\Pi^I_{\ell, i}) \end{align*} By assumption we have that \(\|\Delta_D\| \to_p 0\) and \(\|\Delta_N\|\to_p 0\). Using this notation, we can write the infeasible version of the test statistic as \(\JK^I(\beta_0) = N'D^{-1}N\) while the feasible version is written \(\JK(\beta_0) = (N + \Delta_N)'(D + \Delta_D)^{-1}(N + \Delta_N)\). Add and subtract \(D^{-1}\) to get \begin{align*} \JK(\beta_0) &= \big(N + \Delta_N\big)'\big( (D + \Delta_D)^{-1} \pm D^{-1}\big)\big(N + \Delta_N\big) \\ &= \JK^I(\beta_0) + N'\big( (D + \Delta_D)^{-1} - D^{-1}\big)N + \Delta_N\big((D + \Delta_D)^{-1} - D^{-1})N\\ &\;\;+ \Delta_N'\big((D + \Delta_D)^{-1} - D^{-1}\big)\Delta_N + N'D^{-1}\Delta_N + \Delta_ND^{-1}N + \Delta_ND^{-1}\Delta_N \end{align*} Via (ref) we have that \(\|D^{-1}\| = (\lambda_{\min}(D))^{-1} = O_p(1)\) and by assumption we have that \(\Delta_N \to_p 0\). It therefore suffices to show that \begin{equation} \|(D + \Delta_D)^{-1} - D^{-1}\| \to_p 0 \end{equation} To do so, we can use the following equality from horn2012matrix, p. 381. \[ \|(D + \Delta_D)^{-1} - D^{-1}\| \leq \frac{\|D^{-1}\|^2\|\Delta_D\|}{1 - \|D^{-1}\Delta_D\|} \] Since \(\|D^{-1}\| = O_p(1)\) and \(\Delta_D \to_p 0\), this gives (ref).

Proofs of Results in (ref)

The statement of (ref) relies on showing

align*[align* omitted — 295 chars of source]

In particular, since \((\JK_G(\beta_0) \perp C_G)\) and \((\S_G(\beta_0)\perp C_G)\) under \(H_0\), showing the above will imply the test based on \(T(\beta_0;\tau)\) has asymptotic size \(\alpha\) for any choice of cutoff \(\tau\). The second line in the above display follows imediately from (ref) after verifying (ref), below.

The first line in the top display relies on a joint interpolation of the infeasible \(\JK_I(\beta_0)\) test statistic and the infeasible conditioning statistic \(C_I\), which could be constructed if \(\rho(z_i)\) was known to the researcher.

equation[equation omitted — 189 chars of source]

This joint interpolation argument is rather involved however, and deferred to (ref). The interpolation argument for the conditioning statistic very closely follows the results in cck2013. The results of (ref) rely on showing that the difference between \(C\) and \(C_I\) can be treated as negligible. This in turn reduces to verifying (ref), which is done in (ref), below.

lemma[] Suppose that (ref) holds. Then there are sequences \(\delta_n \searrow 0\), \(\beta_n \searrow 0\) such that \[ \Pr\big(\max_{i \in [n]} n^{-1}\sum_{j=1}^n \dot h_{ij}^2(\widehat r_j - r_j)^2 > \delta_n^2 /\log^2(n)\big) \leq \beta_n \] where \(\dot h_{ij} = h_{ij}/(n^{-1}\sum_{j=1}^n h_{ij}^2)^{1/2}\).
proofIn view of (ref) it suffices to show \begin{equation} \max_{1 \leq i \leq n} \frac{1}{n}\sum_{j=1}^n\dot h_{ij}^2 (\hat r_i - r_i)^2 = o_p(1/\log^2(n)) \end{equation} Notice that we can bound \begin{align*} \max_{1 \leq i \leq n} \frac{1}{n}\sum_{j=1}^n (\hat r_i - r_i)^2 &= \max_{1 \leq i \leq n}\big|(\widehat\gamma - \gamma)'n^{-1}\sum_{j=1}^n\eps_j^2(\beta_0)b(z_i)b(z_j)'(\widehat\gamma - \gamma)\big| \\ &\;\;\; + \max_{1 \leq i \leq n} | n^{-1}\sum_{j=1}^n \dot h_{ij}^2\xi_j^2 |\\ &\leq \max_{\substack{1 \leq i \leq n \\ 1 \leq j,k \leq d_b}}\underbrace{\big|n^{-1}\sum_{j=1}^n \eps_j^2(\beta_0)b_j(z_j)b_{k}(z_j)\big|}_{\vA_{ijk}}\|\widehat\gamma - \gamma\|_1^2 \\ &\;\; + n^{-1/2}\max_{1 \leq i \leq n}(n^{-1}\sum_{j=1}^n \dot h_{ij}^4)^{1/2}(\sum_{j=1}^n \xi_j^4)^{1/2} \end{align*} Under (ref)(i,ii) each \(\vA_{ijk}\) is \(\upsilon\)-sub-exponential by (ref) (that is \(\|\vA_{ijk}\|_{\psi_\upsilon}\) is bounded). An application of (ref) then yields that \(\max_{i,j,k}|\vA_{ijk}| = O_p(\log^{1/\nu}(d_bn))\). Along with (ref)(iv) this gives that \(\max_{i,j,k}|\vA_{ijk}|\|\widehat\gamma - \gamma\|_1 = O_p(\log^{-3/(v\wedge 1)}(d_bn)) = o_p(\log^{-2}(n))\). Meanwhile by definition of \(\dot h_{ij}\), \(\max_i (n^{-1}\sum_{j=1}^n \dot h_{ij}^4)^{1/2} = O(1)\) while by (ref)(iii) \((\sum_{j=1}^n \xi_j^4)^{1/2} = o(1)\). Since \(\log^2(n)/\sqrt{n} \to 0\) this shows (ref).

Proof of (ref)

The first result in (ref) with \(\JK(\beta_0)\) and \(C\) replaced with their infeasible analogs \(\JK_I(\beta_0)\) and \(C_I\) follows from the argument in (ref). After verifying that \(|\JK(\beta_0) - \JK(\beta_0)| \to_p 0\) via (ref) and that (ref) is satisfied via (ref) follow the same steps as in the proof of belloni2018highdimensional, {Theorem 2.1} to see that approximation result holds for the feasible \(\JK(\beta_0)\) and \(C\).

For the second statement, I show that the conditions of (ref) are satisfied. To see that (ref)(i,ii) is satisfied under the moment assumptions of (ref) use (i) the definition of \(\dot h_{ij} = h_{ij}\big/(n^{-1}\sum_{j=1}^n h_{ij}^2)^{1/2}\); (ii) that the variance of each \(r_{j}\) is bounded away from zero and (iii) that the fourth moments of \(r_j\) are bounded from above. (ref)(iii) is satisfied with \(B_n = \log^{1/\upsilon}(n)\) by (ref)(i,iii) and (ref). Finally (ref) is satisfied by applying (ref). Apply (ref) to conclude.

Joint Gaussian Approximation of \(\JK(\beta_0)\) and \(C\)

(ref) rely on a joint interpolation of the conditioning and testing statistics as well as a joint interpolation of the conditioning and testing statistics. The joint interpolation of \(\JK(\beta_0)\) and the conditioning statistic \(C\) is given in (ref) after introducing some notation in (ref). The joint gaussian approximation of \(S(\beta_0)\) and \(C\) follows immediately from results in \cites{belloni2018highdimensional,CCK-2018-HDCLT}. The result is presented below for the general form of the \(\JK(\beta_0)\) statistic under \(H_0\) however the proof strategy is very similar when using the decomposed form of \(\JK(\beta_0)\) when \(d_x = 1\). This proof is available on request.

Notation

\paragraph{Jackknife Statistic Definitions.} Define \(\tilde h_{\ell, ij} = s_{n,\ell} h_{ij}\) for each \(\ell = 1,\dots,d_x\) and the scaled leave-one-out quasi-numerator and denominators

align*[align* omitted — 453 chars of source]

where \(\dot\eps_j(\beta_0)\) is equal to \(\tilde\eps_j(\beta_0)\) if \(j < i\) and equal to \(\eps_j(\beta_0)\) if \(j > i\), \(\dot r_{\ell j}\) is equal to \(\tilde r_{\ell j}\) if \(j < i\) and equal to \(r_j\) if \(j > i\), and \(\ddot \eps_j(\beta_0)\) is equal to \(\E[\eps_j^2(\beta_0)]\) if \(j < i\) and equal to \(\eps_j(\beta_0)\) if \(j > i\). As in the proof of (ref) while the definitions of \(\dot\eps_j(\beta_0), \dot r_{\ell j}\), and \(\ddot\eps_j(\beta_0)\) depend on \(i\) this dependence is suppressed to conslidate notation and since we only consider one step deviations at a time.

Also define the one step deviations

align*[align* omitted — 968 chars of source]

where

align*[align* omitted — 1,180 chars of source]

Notice that in this notation we can write the test statistic and gaussian test statistics, after scaling by \(\diag(s_{n,1},\dots,s_{n,d_x})\), as

align*[align* omitted — 331 chars of source]

In this proof we will use these representations for the test statistics. Finally define

align*[align* omitted — 174 chars of source]

\paragraph{Conditioning Statistic Definitions.} Let \(h_{\ell, ii} = 0\) for any \(\ell = 1,\dots,d_x\) and \(i= 1,\dots,n\). Define \(\tilde h_{\ell,ij} = h_{\ell,ij}/\omega_{\ell i}\) for \(\omega_{\ell i} = n^{-1}\sum_{j=1}^n |h_{\ell,ij}|\). Also define the one-step deviations:

align*[align* omitted — 565 chars of source]

Notice that \(C = \max_{1 \leq \iota \leq 2nd_x} (C_{-1} + \frac{1}{\sqrt{n}}\Delta_{C1})_\iota\) while \(\tilde C = \max_{1 \leq \iota \leq 2nd_x}(C_{-n} + \Delta_{Cn})_\iota \).

\paragraph{Function Definitions.} As in cck2013 consider the “smooth max” function, \(F_\beta:\SR^p \to \SR\) defined \[ F_\beta(z) = \beta^{-1}\log\bigg(\sum_{i=1}^n \exp(\beta z_i)\bigg) \] which satisfies \[ 0 \leq F_\beta(z) - \max_{1 \leq i \leq n} z_i \leq \beta^{-1}\log p. \] (ref) notes some useful properties of the smooth max function which we will use in the joint interpolation argument. In addition let \(\varphi(\cdot) \in C_b^3(\SR)\) be such that \(\varphi(x) = 1\) if \(x \leq 0\), \(\varphi'(x) < 0\) for \(x \in (0,1)\), and \(\varphi(x) = 0\) for \(x \geq 1\). For any \(\gamma > 0\) and \(a = (a_1,a_2)'\in \SR^2\) define the function \(\tilde \varphi(\cdot,\cdot,\cdot): \SR^{d_x} \times \vec(\SR^{d_x \times d_x}) \times \SR^{2nd_x} \to \SR\) via

equation[equation omitted — 149 chars of source]

where

align*[align* omitted — 217 chars of source]

The function \(\tilde\varphi_{\gamma,a}(\cdot,\cdot,\cdot)\) is meant to approximate the indicator function \(\bm{1}\{K(\beta_0) \leq a_1\}\bm{1}\{C \leq a_2\}\) with \(\gamma\) governing the quality of approximation. Where it is obvious, we will supress the subscripts \(\gamma,a\) from our notation.

Main Argument

lemma[Joint Lindeberg Interpolation] Suppose that (ref) hold as well as the moment conditions of (ref). Then there are fixed constants \(M_1,M_2\) such that \begin{equation} \left|\E[\tilde\varphi_{\gamma,a}(U,\vec(D),C) - \tilde\varphi_{\gamma,a}(\tilde U, \vec(\tilde D), \tilde C)]\right| \leq \frac{M_1\log^{M_2}(n)}{\sqrt{n}}(\gamma^{-1} + \gamma^{-2} + \gamma^{-3}) \end{equation}
proof[Proof of (ref)] We can bound the difference on the left hand side of (ref) using the telescoping sum \begin{equation} \begin{split} \sum_{i=1}^n\big| \E[\tilde\varphi_{\gamma,a}(U_{-i} &+ \Delta_{Ui}/\sqrt{n},\vec(D_{-i} + \Delta_{Di}/n), C_{-i} + \Delta_{Ci}/\sqrt{n})]\\ &-\E[\tilde\varphi_{\gamma,a}(U_{-i} + \Delta_{Ui}/\sqrt{n},\vec(D_{-i} + \Delta_{Di}/n), C_{-i} + \Delta_{Ci}/\sqrt{n})]\big| \end{split} \end{equation} By second degree Taylor expansion, we break each of the summands in (ref) into first order, second order, and remainder terms; each of which are bounded below. We make use of the following moment conditions implied by (i) indpendence of observations across \(i = 1,\dots,n\) and (ii) the mean and covariance matrix of \((\eps_i(\beta_0),r_i)\) being equal to the mean and covariance matrix of \((\tilde\eps_i(\beta_0),r_i)\) \begin{equation} \begin{split} 0 &=\E[\Delta_{Ui} - \tilde\Delta_{Ui}|\calF_{-i}] = \E[\Delta_{Ui}\Delta_{Ui}' - \tilde\Delta_{Ui}\tilde\Delta_{Ui}'|\calF_{-i}] = \E[\vec(\Delta_{Di}) - \vec(\tilde\Delta_{Di})|\calF_{-i}] \\ &= \E[\Delta_{Ci} - \tilde\Delta_{Ci}|\calF_{-i}] = \E[\Delta_{Ui}\otimes\vec(\Delta_{Di}^b)' - \tilde\Delta_{Ui}\otimes\vec(\tilde\Delta_{Di}^b)'|\calF_{-i}]\\ &= \E[\Delta_{Ci}\otimes\Delta_{Ui} - \tilde\Delta_{Ci}\otimes\tilde\Delta_{Ui}|\calF_{-i}] = \E[\Delta_{Ci}\otimes\vec(\tilde\Delta_{Di}^b) - \tilde\Delta_{Ci}\otimes\vec(\tilde\Delta_{Di}^b)|\calF_{-i}]\\ &= \E[\vec(\Delta_{Di}^b)\vec(\Delta_{Di}^b)' - \vec(\tilde\Delta_{Di}^b)\vec(\tilde\Delta_{Di}^b)'|\calF_{-i}] \end{split} \end{equation} where \(\calF_{-i}\) denotes the sub-sigma algebra generated by all observations not equal to \(i\), \(\otimes\) denotes the Kronecker product, and I apologize for the abuse of the equal sign in the above display. First Order Terms. First order terms can be expressed \begin{align*} First Order_i &= \sum_{\ell = 1}^{d_x} \E\left[\frac{\partial }{\partial U_\ell}\tilde\varphi(U_{-i},\vec(D_{-i}),C_{-i})((\Delta_{Ui})_\ell - (\tilde\Delta_{Ui})_\ell)\right]/\sqrt{n} \\ &+ \sum_{\ell=1}^{d_x}\sum_{m = 1}^{d_x} \E\left[\frac{\partial }{\partial D_{\ell m}}\tilde\varphi(U_{-i},\vec(D_{-i}),C_{-i})((\Delta_{Di})_{\ell m} - (\tilde\Delta_{Di})_{\ell m})\right]/n \\ &+ \sum_{\ell = 1}^{2nd_x} \E\left[\frac{\partial }{\partial C_{\ell}}\tilde\varphi(U_{-i}, \vec(D_{-i}), C_{-i})((\Delta_{Ci})_\ell - (\tilde\Delta_{Ci})_\ell)\right] /\sqrt{n} \end{align*} These terms are all equal to zero after applying the matched moments in (ref). Second Order Terms. After canceling out terms using the matched moments in (ref) the second order terms that remain can be expressed \begin{align*} 2nd Order_i &= \frac{1}{n^{3/2}}\sum_{\ell = 1}^{d_x} \sum_{m = 1}^{d_x} \sum_{n = 1}^{d_x} \underbrace{\E\bigg[\frac{\partial^2 }{\partial U_{\ell} \partial D_{mn}}\tilde\varphi(U_{-i},\vec(D_{-i}),C_{-i})((\Delta_{Ui})_\ell (\Delta_{Di}^a)_{mn} - (\tilde\Delta_{Ui})_\ell(\tilde\Delta_{Di}^a)_{mn}) \bigg]}_{\vA_{\ell mn}} \\ &= \frac{1}{n^{2}}\sum_{\ell = 1}^{d_x} \sum_{m = 1}^{d_x} \sum_{n = 1}^{d_x}\sum_{o=1}^{d_x} \underbrace{\E\bigg[\frac{\partial^2 }{\partial U_{\ell} \partial D_{mn}}\tilde\varphi(U_{-i},\vec(D_{-i}),C_{-i})((\Delta_{Di}^a)_{\ell m} (\Delta_{Di}^a)_{no} - (\tilde\Delta_{Di}^a)_{\ell m}(\tilde\Delta_{Di}^a)_{no}) \bigg]}_{\vB_{\ell m n o}} \\ &= \frac{2}{n^{2}}\sum_{\ell = 1}^{d_x} \sum_{m = 1}^{d_x} \sum_{n = 1}^{d_x}\sum_{o=1}^{d_x} \underbrace{\E\bigg[\frac{\partial^2 }{\partial U_{\ell} \partial D_{mn}}\tilde\varphi(U_{-i},\vec(D_{-i}),C_{-i})((\Delta_{Di}^b)_{\ell m} (\Delta_{Di}^a)_{no} - (\tilde\Delta_{Di}^a)_{\ell m}(\tilde\Delta_{Di}^b)_{no}) \bigg]}_{\vC_{\ell m n o}} \\ &= \frac{1}{n^{3/2}}\sum_{\ell = 1}^{2nd_x} \sum_{m = 1}^{d_x} \sum_{n = 1}^{d_x} \underbrace{\E\bigg[\frac{\partial^2 }{\partial C_{\ell} \partial D_{mn}}\tilde\varphi(U_{-i},\vec(D_{-i}),C_{-i})((\Delta_{Ci})_{\ell} (\Delta_{Di}^a)_{mn} - (\tilde\Delta_{Ci})_{\ell}(\tilde\Delta_{Di}^a)_{mn}) \bigg]}_{\vD_{\ell m n}} \end{align*} To bound each \(\vA_{\ell m n}\), \(\vB_{\ell mno}\), and \(\vC_{\ell m n o}\) we use the fact that the second order derivatives of \(\tilde\varphi\) are bounded up to a log power of \(n\) via repeated application of (ref). Under the moment conditions of (ref) the absolute value of terms \((\Delta_{Ui})_{\ell}\),\(|\Delta_{Di}^a|_{mn}\), and \((\Delta_{Di}^b/\sqrt{n})_{no}\) can also be shown to have bounded third moments via the exact same steps as in the proof of (ref). Putting these together with generalized Holder's inequality will yield a finite constants \(M_1\) and \(M_2\) such that \(|\vA_{lmn}| \leq M_1\log^{M_2}(n)(\gamma^{-1} + \gamma^{-2})\), \(\vB_{\ell mno} \leq M_1\log^{M_2}(n)(\gamma^{-1} + \gamma^{-2})\), and \(|\vC_{\ell mno}| \leq M_1\log^{M_2}(n)n^{1/2}(\gamma^{-1} + \gamma^{-2})\). To bound \(\vD_{\ell m n}\) terms notice that \begin{align*} \sum_{\ell = 1}^{2nd_x}\vD_{\ell mn} &=\sum_{\ell = 1}^{2nd_x} \E\bigg[\frac{\partial }{\partial D_{mn}}\phi(U_{-i},\vec(D_{-i}))\frac{\partial }{\partial C_\ell}\tau(C_{-i})((\Delta_{C-i})_\ell(\Delta_{Di}^a)_{mn} - (\tilde\Delta_{Ci})_\ell(\tilde\Delta_{Di}^a)_{mn}))\bigg] \intertext{Apply (ref) to bound \(\Delta_{Di}^a\), and (ref) to bound the derivative of \(\phi(\cdot)\) and Cauchy-Schwarz to split up the \(\Delta_{Ci}\) and \(\Delta_{Di}\) terms} &\leq \sqrt{M_1\log^{M_2}(n)\gamma^{-2}}\E\bigg[\sum_{\ell =1}^{2nd_x}(\partial_\ell \tau(C_{-i}))^2((\Delta_{Ci})_\ell + (\tilde\Delta_{Ci})_\ell)^2\bigg]^{1/2}\\ &\leq \sqrt{M_1\log^{M_2}(n)\gamma^{-2}}\E\bigg[\max_{1 \leq \ell \leq n} ((\Delta_{Ci})_{2\ell} + (\tilde\Delta_{Ci})_{2\ell})^2 \sum_{\ell = 1}^{2nd_x} (\partial_\ell\tau(C_{-i}))^2\bigg]^{1/2} \intertext{By (ref) and chain rule we have that \(\sum_{\ell = 1}^{2nd_x}(\partial_\ell \tau(C_{-i}))^2 \leq \gamma^{-2}\). Moreover \((\Delta_{Ci})^{a/2}_\ell\) is sub-exponential so via (ref) the second moment of the maximum is bounded by a power of \(\log(n)\). After updating the constant \(M_1\) and \(M_2\) this yields} &\leq M_1\log^{M_2}(n)\gamma^{-2} \end{align*} Putting these all together and summing over the remaining indices gives \begin{equation} |Second Order_i| \leq \frac{M_1\log^{M_2}(n)}{n^{3/2}}(\gamma^{-1} + \gamma^{-2}) \end{equation} Remainder Terms. The first remainder term can be expressed \begin{align*} \text{Remainder}_i &= \frac{1}{n^{3/2}}\sum_{\ell =1}^{d_x}\sum_{m = 1}^{d_x}\sum_{n=1}^{d_x} \E\bigg[\frac{\partial^3 }{\partial U_\ell \partial U_m \partial U_n}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ui})_\ell(\Delta_{Ui})_m(\Delta_{Ui})_n\bigg] \\ &+ \frac{1}{n^3}\sum_{(\ell,m)}\sum_{(n,o)}\sum_{(q,p)} \E\bigg[\frac{\partial^3 }{\partial D_{\ell m}\partial D_{no}\partial D_{pq}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Di})_{\ell m}(\Delta_{Di})_{no}(\Delta_{Di})_{qp}\bigg] \\ &+ \frac{1}{n^{3/2}}\sum_{\ell =1}^{2nd_x}\sum_{m = 1}^{2nd_x}\sum_{n=1}^{2nd_x} \E\bigg[\frac{\partial^3 }{\partial C_\ell \partial C_m \partial C_n}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ci})_\ell(\Delta_{Ci})_m(\Delta_{Ci})_n\bigg] \\ &+ \frac{1}{n^{2}}\sum_{\ell =1}^{d_x}\sum_{m = 1}^{d_x}\sum_{(n,o)} \E\bigg[\frac{\partial^3 }{\partial U_\ell \partial U_m \partial D_{no}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ui})_\ell(\Delta_{Ui})_m(\Delta_{Di})_{no}\bigg] \\ &+ \frac{1}{n^{5/2}}\sum_{\ell =1}^{d_x}\sum_{(m,n)}\sum_{(o,p)} \E\bigg[\frac{\partial^3 }{\partial U_\ell \partial D_{mn} \partial D_{op}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ui})_\ell(\Delta_{Di})_{mn}(\Delta_{Di})_{op}\bigg] \\ &+ \frac{1}{n^{5/2}}\sum_{\ell =1}^{2nd_x}\sum_{(m,n)}\sum_{(o,p)} \E\bigg[\frac{\partial^3 }{\partial C_\ell \partial D_{mn} \partial D_{op}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ci})_\ell(\Delta_{Di})_{mn}(\Delta_{Di})_{op}\bigg] \\ &+ \frac{1}{n^{2}}\sum_{\ell =1}^{2nd_x}\sum_{m = 1}^{2nd_x}\sum_{(n,o)} \E\bigg[\frac{\partial^3 }{\partial C_\ell \partial C_m \partial D_{no}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ci})_\ell(\Delta_{Ci})_m(\Delta_{Di})_{no}\bigg] \\ &+ \frac{1}{n^{3/2}}\sum_{\ell =1}^{2nd_x}\sum_{m = 1}^{2nd_x}\sum_{n=1}^{d_x} \E\bigg[\frac{\partial^3 }{\partial C_\ell \partial C_m \partial U_{n}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ci})_\ell(\Delta_{Ci})_m(\Delta_{Ui})_{n}\bigg] \\ &+ \frac{1}{n^{2}}\sum_{\ell =1}^{2nd_x}\sum_{m = 1}^{2nd_x}\sum_{n=1}^{d_x} \E\bigg[\frac{\partial^3 }{\partial C_\ell \partial C_m \partial U_{n}}\tilde\varphi(\bar U, \vec(\bar D), \bar C)(\Delta_{Ci})_\ell(\Delta_{Ci})_m(\Delta_{Ui})_{n}\bigg] \\ \end{align*} where \(\bar U, \vec(\bar D)\), and \(\bar C\) vary term by term but are always in the hyper-rectangles \([U_{-i}, U + \Delta_{Ui}]\), \([\vec(D_{-i}), \vec(D_{-i} + \Delta_{Di})]\), and \([C_{-i}, C_{-i} + \Delta_{Ci}]\), respectively. As such, any moment conditions that apply to \(U, D, C\) also apply to \((\bar U,\bar D, \bar C)\). Repeated application of generalized Hölder inequality, (ref) to bound moments of \(\Delta_{Ui}\) and \((\Delta_{Di}/\sqrt{n})\), (ref) to bound moments of the second and third derivatives of \(\phi(\tilde U, \vec(\tilde D))\), (ref) to bound the sums of derivatives of \(\tau(\tilde C)\), and (ref) to bound moments of \(\max_{1 \leq \ell \leq n} (\Delta_{Ci})_\ell\) will yield that \begin{equation} |\text{Remainder}_i| \leq \frac{M_1\log^{M_2}(n)}{n^{3/2}}(\gamma^{-1} + \gamma^{-2} + \gamma^{-3}) \end{equation} Symmetric logic will bound the other remainder term. Summing (ref) and (ref) over indices gives the result.
lemma[Denominator Anticoncentration] Suppose that (ref) hold as well as the moment conditions of (ref). Then for any sequence \(\delta_n \to 0\) we have that \(\Pr(\lambda_{\min}(\tilde D) \leq \tilde\delta_{n}) \to 0\).
proofBy (ref) it suffices to show that for any fixed \(a\in\calS^{d_x-1}\) and any \(\delta_n \to 0\), \(\Pr(a'Da \leq \delta_{n})\to 0\). For any such \(a\) write: \begin{align*} a'\tilde Da &= \frac{1}{n}\sum_{i=1}^n \E[\eps_i^2(\beta_0)]\big(\sum_{\ell=1}^{d_x}\sum_{j=1}^n a_\ell \tilde h_{\ell,ij} r_{\ell,j}\big)^2 \\ &\geq \frac{1}{cn}\sum_{i=1}^n\big(\sum_{\ell=1}^{d_x}\sum_{j=1}^n a_\ell \tilde h_{\ell,ij} r_{\ell,j}\big)^2 \intertext{Define \(\dot s_{n,j} = \max_{\{\ell: a_\ell \neq 0\}}s_{n,\ell}\) and \(\dot h_{ij} =s_n h_{ij}\)} &= \frac{1}{cn}\sum_{i=1}^n \big(\sum_{j=1}^n \dot h_{ij} \sum_{\ell=1}^{d_x} \frac{a_\ell s_{n,\ell}}{s_n} r_{\ell, j}\big)^2 \\ \end{align*} By the moment conditions required by (ref) we have that \(\lambda_{\min}(\E[D]) \geq \underline c\) so that \(\E[\frac{1}{n}\sum_{i=1}^n \big(\sum_{\ell=1}^{d_x} \sum_{j=1}^n a_\ell \tilde h_{\ell, ij}r_{\ell,j}\big)^2] \geq c^{-1}\). Moreover, by assumption, \(\Var(\sum_{\ell = 1}^{d_x} \frac{a_\ell s_{n,\ell}}{s_n})\) is bounded from above and below. Define the matrix \(\tilde H = [\dot h_{ij}]_{ij}\) and follow the same steps as (ref) to conclude.
lemma[Gaussian Approximation] Suppose that (ref) hold as well as the moment conditions of (ref). Then, \[ \sup_{a \in \SR} \big|\Pr(\JK_I(\beta_0) \leq a) - \Pr(\JK_G(\beta_0) \leq a)\big|\to 0 \]
proofLet \(a = (a_1,a_2)\) and \(\tilde\phi_{\gamma,a}\) be as in (ref): \begin{align*} \Pr(N'D^{-1}N \leq a_1, C \leq a_2) &\leq \E[\tilde\phi_{\gamma,a}(U, \vec(D), C)] \\ &\leq \E[\tilde\phi_{\gamma,a}(\tilde U, \vec(\tilde D), \tilde C)] + \frac{M_1\log^M_2(n)}{\sqrt{n}}(\gamma^{-1} + \gamma^{-2}) \\ &\leq \Pr(\tilde N'\tilde D^{-1}\tilde N \leq a_1, \tilde C \leq a_2) + \Pr(a_1 \leq \tilde N'\tilde D^{-1}N \leq a_1 + \gamma\lambda_{\min}^5(D)) \\ &\;\;\; + \Pr(a_2 \leq C \leq a_2 + \gamma) + \frac{M_1\log_2^{M_2}(n)}{\sqrt{n}}(\gamma^{-1} + \gamma^{-2} + \gamma^{-3}) \\ &\leq \Pr(\tilde N'\tilde D^{-1}\tilde N \leq a_1, \tilde C \leq a_2) + \Pr(a_1 \leq \tilde N'\tilde D^{-1}N \leq a_1 + \gamma\lambda_{\min}^5(D)) \\ &\;\;\; + \Pr(a_2 \leq C \leq a_2 + \gamma) + \frac{M_1\log_2^{M_2}(n)}{\sqrt{n}}(\gamma^{-1} + \gamma^{-2} + \gamma^{-3}) \\ \end{align*} Let \(\gamma \to 0\) at a rate such that \(\frac{\log^{M_2}(n)}{\sqrt{n}}\gamma^{-3}\to 0\) and apply (ref) to conclude as in the proof of (ref). A symmetric argument shows that the lower bound tends to zero.
lemma[] Let \(\Sigma_n \in \SR^{d\times d}\) be a sequence of random positive-semidefinite matrices. Suppose that for any fixed \(a \in \calS^{d-1}\) and any \(\delta_n \to 0\) we have that \(\Pr(a'\Sigma_na \leq \delta_n) \to 0\) and \(\Pr(\lambda_{\max}^2(\Sigma_n) \geq \delta_n^{-1}) \to 0\). Then for any \(\delta_n \to 0\), \(\Pr(\lambda_{\min}^2(\Sigma_n) \leq \delta_n)\to 0\).
proofTake any preliminary sequence \(\delta_n \to 0\). It suffices to show that there is another sequence \(\tilde\delta_n\) weakly larger than \(\delta_n/2\) such that \(\Pr(\lambda_{\min}^2(\Sigma_n) \leq \tilde\delta_n) \to 0\). For any \(m \in \SN\) let \(\calA_m\) be a set of points in \(\calS^{d-1}\) such that \[ \max_{a\in\calS^{d-1}} \min_{\tilde a \in \calA_m} \|a - \tilde a\| \leq \delta_m^2 \] From here let \(\tilde n_j\) be defined \[ \tilde n_j = \inf\{n \geq j: \min_{\tilde a \in \calA_{n,j}}\Pr(\tilde a'\Sigma_na \leq 2\delta_{n_j}) < \delta_{n_j}\} \] Define a new sequence \(\tilde\delta_n \to 0\), weakly larger than \(\delta_n\), via \[ \tilde\delta_n = \begin{cases} 1 &\text{if }0 \leq 0 \leq n < \tilde n_1 \\ \delta_i &\text{if } \tilde n_i \leq n < \tilde n_{i+1} \end{cases} \] and notice that, by definition \(\Pr(\min_{a \in \calA_{\tilde n_j}} a'\Sigma_na \leq 2\tilde\delta_n) < \delta_{\tilde n_j}\). We wish to show that \(\lambda_{\min}^2(\Sigma_n) > \tilde\delta_n\) on an intersection of events whose probability tends to one. Since \(\Sigma_n\) is positive semi-definite, \(\|x\|^2_{\Sigma_n} = x'\Sigma_nx\) defines a seminorm. By triangle inequality \[ \lambda_{\min}^2(\Sigma_{n_j}) \geq \min_{\calA_{n_j}} a'\Sigma_{n_j} a - \lambda_{\max}^2(\Sigma_{n})\tilde\delta_{n_j}^2 \] Define the events \[ \Omega_1 = \{\min_{\calA_{\tilde n_j}} a'\Sigma_{n} a \geq 2\tilde\delta_n\} \andbox \Omega_2 = \{\lambda_{\max}(\Sigma_n) \leq \tilde\delta_n^{-1/2}\} \] On the intersection of these events, whose probabilities tend to one, we have \(\lambda_{\min}^2(\Sigma_{n}) \geq \tilde\delta_{n}\).

\newblankpage

Incorporating Exogenous Controls

In this section, I analyze the model with exogeneous controls. To this end, define the vector \(z_2 = (z_{21}',\dots,z_{2n}')' \in \SR^{n \times d_c}\). Let \(P_2 = z_2(z_2'z_2)^{-1}z_2' \in \SR^{n\times n}\) denote the projection onto the column space of \(z_2\) and \(M_2 = I_n - P_2\) denote the projection onto to orthocomplement of the column space. Focus will be on the case where \(d_x = 1\) to simplify notation, but the basic concepts apply generally to \(d_x > 1\).

For \(y \coloneqq (y_1,\dots,y_n)' \in \SR^n\) and \(x \coloneqq (x_1',\dots,x_n')'\in \SR^{n\times}\) define \(y^\perp \coloneqq M_2y\) and \(x^\perp \coloneqq M_2 x\) as the “partialled out” versions of \(y\) and \(x\), respectively. Let \(y^\perp_i\) be the \(i\)\textsuperscript{th} element of \(y^\perp\) and \(x^\perp_i\) be the \(i\)\textsuperscript{th} element of \(x^\perp\). From here we can define \(\eps(\beta_0) \coloneqq y - x\beta_0 \), \(\eps^\perp(\beta_0) = M_2\eps(\beta_0)\) and \(r^\perp \coloneqq M_2r\) where as in the main text \(r = (r_1,\dots,r_n)'\) is constructed \(r_i = x_i - \rho(z_i)\eps_i(\beta_0)\). The definition of \(\rho(z_i)\) does not change after partialling out \(z_2\) since all expectations are understood to be conditional on the instruments \(z\). Notice that \(\eps^\perp(\beta_0)\) is mean zero. Finally I assume that the controls have been partialled out of hat matrix so that the effective hat matrix is \(M_2H\) and the vector \(\widehat\Pi \in \SR^n\) is defined \(\widehat\Pi = (M_2 H)(M_2 r)\). This does not make a difference for the numerator of the \(\JK(\beta_0)\) statistic but does affect the denominator slightly. When this is not done, inference may be conservative.

Using matrix notation in the numerator to make things clear, we can write the version of the \(\JK(\beta_0)\) statistic with the partialled out vectors, \(\eps^\perp(\beta_0)\) and \(r^\perp\), in terms of the original vectors, \(\eps(\beta_0)\) and \(r\),

align*[align* omitted — 385 chars of source]

where \(\vh_{ij} = [M_2 \tilde HM_2]_{ij}\), \(\tilde H = s_n H\), and \(m_{ij} = [M_2]_{ij}\). I seek to characterize the limiting distribution of \(\JK(\beta_0)\) under \(H_0\). To do so, we show that quantiles \(\JK(\beta_0)\) can be approximated by quantiles of the gaussian analog statistic

align*[align* omitted — 200 chars of source]

where \((\tilde\eps_i, \tilde \eps_i(\beta_0), \tilde r_i)\) are generated gaussian independent of the data and with the same mean and covariance as \((\eps_i, \eps_i(\beta_0),r_i)\). Since \(\Var(\tilde\eps(\beta_0)) = \Var(\eps_i)\) under \(H_0\), \(\E[\tilde\eps(\beta_0)'M_2] = 0\), and \(\tilde r \perp \tilde\eps(\beta_0)\), this gaussian analog statistic has a \(\chi^2_1\) distribution conditional on any realization of \(\tilde r\) and thus its unconditional distribution is also \(\chi^2_1\).

Showing that quantiles of \(\JK(\beta_0)\) can be approximated by quantiles of \(\tilde\JK(\beta_0)\) proceeds in two steps. In the first step, we show that \(\JK(\beta_0)\) converges in probability to an intermediate statistic.

align*[align* omitted — 201 chars of source]

We will then show that quantiles of this intermediate statistic can be approximated by quantiles of \(\tilde\JK(\beta_0)\). In view of (ref), it suffices to show for the first step that \(\Delta_D \to_p 0\), where

align*[align* omitted — 108 chars of source]

To do this, notice that under \(H_0\) we can write \(\eps_i^\perp(\beta_0) = \eps_i + z_{2i}'(\widehat\Gamma - \Gamma)\) where \(\widehat\Gamma = (z_2'z_2)^{-1}z_2\eps(\beta_0)\) is a \(\sqrt{n}\)-consistent estimate of \(\Gamma\). Exploiting this fact we get

align*[align* omitted — 212 chars of source]

Both of these terms will tend to zero by the consistency \(\widehat\Gamma\) to \(\Gamma\), giving that \(\Delta_D \to_p 0\).

In our second step, we argue that quantiles of \(\JK^{\text{int}}(\beta_0)\) can be approximated by quantiles of \(\JK_G(\beta_0)\). To make this comparasion, we can follow almost exactly the same steps as in (ref). The only difference between analysis in this case and analysis in the original case is that the partialling out of controls leads the test statistic to not strictly have a jackknife form; the effective hat matrix \(M_2HM_2\) no longer has a deleted diagonal. However, as I will argue below, this will not make a difference in the interpolation argument since the diagonal terms of \([P_2]_{ii}\) are small in the sense that they sum to \(d_c\).

The (ref) analog one step deviations for the numerator are given

align*[align* omitted — 340 chars of source]

where as (ref), a dotted variable is equal to the gaussian analog if \(j > i\) but equal to the standard version otherwise. The first and second moments of the first two terms in \(\Delta_{1i}\) can be matched with their gaussian analog terms as in the proof of (ref). While we cannot match seconds moments of the third term in the one step deviation, this sum of all these third terms can be treated as negligible after scaling by \(1/\sqrt{n}\) as \(\sum_{i=1}^n |\vh_{ii}| \lesssim d_c\). This is because \(M_2\tilde H M_2 = \tilde H - P_2 \tilde H - \tilde H P_2 - P_2 \tilde H P_2\). The matrix \(\tilde H\) has zeros on it's diagonal. Meanwhile

align*[align* omitted — 188 chars of source]

where the final inequality comes because the matrix \(P_2\) is symmetric and idempotent and since \(\Big(\sum_{j\neq i} H_{ji}^2\Big) \lesssim 1\) by (ref)(ii). A similar argument can be used to show that \([P_2 \tilde H P_2]_{ii}^2 \lesssim [P_2]_{ii} \). Since \(P_2\) is a projection matrix we must have that \(\|P_2He_j\| \leq \|He_j\|\) for any basis vector \(e_j \in \SR^n\). Thus \(\sum_{j=1}^n [P_2H]_{ji}^2 \leq \sum_{j=1}^n [H]_{ji}^2\). Finally, we can use the fact that the trace of \(P_2\) is equal to its rank to show that \(\sum_{i=1}^n|\vh_{ii}| \lesssim d_c\)

The one step deviations in the denominator can be bounded using the same logic. These one step deviations are given

align*[align* omitted — 620 chars of source]

where \(\ddot\eps_j\) is equal to \(\Var(\eps_j)\) if \(j < i\) and equal to \(\eps_j\) if \(j > i\). The first three terms in this expansion are can be dealt with exactly as in the proof of (ref). The fourth term is new, however summing over the fourth terms and scaling by \(1/n\) will be negligible as \(\sum_{i=1}^n |\vh_{ii}| \lesssim d_c\). After showing the lindeberg interpolation step, the rest of the proof follows exactly as in (ref).

Alternative Construction of Test Statistic via Cross Fitting

To accomodate a general class of estimators for \(\hat\rho(z_i)\) I propose a cross-fit form of the jackknife K-statistic. In this section I detail the cross-fitting procedure and present high level conditions needed for estimation error to be treated as negligible. These conditions can be satisfied by a large class of machine learning estimators under alternate conditions.

Cross Fit Test Statistic

To construct the cross fit test statistic, evenly (and randomly) split the sample into two subsets, \(\calI_1\) and \(\calI_2\) such that \(\calI_1 \cap \calI_2 = \emptyset\) and \(\calI_1 \cup \calI_2 = [n]\). For each \(k = 1,2\) construct an estimator \(\hat\rho^{(k)}(\cdot)\) using only the observations in \(\calI_k\). For observations in \(i \in \calI_1\) form the first-stage estimates \[ \widehat\Pi_i = \sum_{j \in \calI_1\setminus\{i\}} h_{ij}^{(1)}(x_j - \eps_j(\beta_0)\hat\rho^{(2)}(z_j)) \] where the weights \(h_{ij}^{(1)}\) come from a hat matrix \(H^{(1)}\) that only depends on the observations in \(\calI_1\), for example the ridge regression hat matrix of (ref) using only the instruments \((z_i)_{i\in\calI_1}\). First stage estimates for observations in \(\calI_2\) symmetrically, using the estimator \(\hat\rho^{(1)}(\cdot)\). The test statistic is then constructed as before \[ \JK(\beta_0) = \frac{\big(\sum_{i=1}^n \eps_i(\beta_0)\widehat\Pi_i\big)^2}{\sum_{i=1}^n \eps_i^2(\beta_0)(\widehat\Pi_i)^2} \] Notice that while only half the sample is being used to construct each first stage estimate, the full sample is still being used to test the null hypothesis.

Controlling Estimation Error

After writing the overall hat matrix as \[ H =

pmatrix[pmatrix omitted — 51 chars of source]

, \] analysis of the infeasible statistic proceeds as before. Thus, to show that the crossfit test statistic has a limiting \(\chi^2_1\) distribution by (ref) it suffices to state high-level conditions under which \((\Delta_N, \Delta_D)' \to_p 0\).

assumption[Cross-fit Conditions] Suppose (i) that there is a constant \(\nu \in (0,1] \cup \{2\}\) such that for all \(i \in[n]\), \(\|\eps_i\|_{\Psi_\nu} \leq c\) and that (ii) for \(k = 1,2\) \[ \max_{i \in I_k} (\hat\rho^{(-k))}(z_j) - \rho(z_j))^2 = o_p(\log^{1/\nu}(n)). \] where \(\hat\rho^{(-k)}(\cdot)\) indicates the estimator of \(\rho(\cdot)\) computed using observations in \([n]\setminus \calI_k\).
lemma[Negligible Cross-fit error] Suppose that Assumptions (ref), and (ref) hold as well as the moment conditions of (ref). Then, under \(H_0\), \((\Delta_N, \Delta_D)'\to_p 0\).
proofWe consider the statements \(\Delta_N\to_p 0\) and \(\Delta_D \to_p 0\) seperately. \paragraph*{\(\Delta_N\to_p 0\):} It suffices to show that \(\Delta_{N,1} \to_p 0\) for \[ \Delta_{N,1} = \frac{1}{\sqrt{n}}\sum_{i\in \calI_1} \eps_i(\beta_0)\sum_{j \in \calI_1\setminus\{i\}} \eps_j(\beta_0)\tilde h_{ij} (\hat\rho^{(2)}(z_j) - \rho(z_j)). \] The corresponding statement for \(\Delta_{N,2}\) follows from the same logic and \(\Delta_N = \Delta_{N,1} + \Delta_{N,2}\). Define the event \[ \Omega(\eps) \coloneqq \{ \max_{i\in\calI_k} (\hat\rho^{(2)}(z_j) - \rho(z_j))^2 \leq \eps\} \] and consider a sequence \(\eps_n \to 0\) such that \(\Pr(\Omega(\eps_n)) \geq 1 - \eps_n\). Noting that \(\Omega(\eps_n) \perp (\eps_i(\beta_0))_{i \in \calI_k}\) we write \[ \Delta_{N,1} = \eps(\beta_0)\mH\eps(\beta_0) \] where \(\mH \in \SR^{|\calI_1|\times |\calI_1|} = \frac{1}{\sqrt{n}}\big(\tilde h_{ij}(\hat\rho^{(1)}(z_j) - \rho(z_j))\big)_{i,j \in \calI_1}\). Since \(\E[\Delta_{N,1}|\Omega(\eps_n)] = 0\), an application of the generalized Hanson-Wright inequality, (ref), gives us that there is a sequence \(\delta_n \to 0\) such that \[ \Pr(\Delta_{N,1} \geq \eps_n|\Omega(\eps_n)) \leq \delta_n \] which allows us to conclude. \paragraph*{\(\Delta_D \to_p 0\):} As before, it suffices to show that \(\Delta_{D,1} \to_p 0\) where \[ \Delta_{N,1} = \frac{1}{n}\sum_{i \in \calI_1} \eps_i^2(\beta_0)\big\{\big(\sum_{j \in \calI_1\setminus\{i\}}\tilde h_{ij}\hat r_j\big)^2 - \big(\sum_{j \in \calI_1\setminus\{i\}}\tilde h_{ij} r_j\big)^2\} \] From the proof of (ref), it suffices to show that \[ \max_{i \in \calI_1} \big|\sum_{j \in \calI_1\setminus\{i\}} \tilde h_{ij}\eps_j(\beta_0)(\hat\rho^{(2)}(z_j) -\rho(z_j) \big|\to_p 0 \] Consider a sequence \(\eps_n \to 0\) such that \(\log^{1/\nu}(n)\eps_n \to 0\) and \(\Pr(\Omega(\eps_n)) \geq 1 - \eps_n\) and apply (ref) to conclude that there is a sequence \(\delta_n \to 0\) such that \[ \Pr\big(\max_{i \in \calI_1} \big|\sum_{j \in \calI_1\setminus\{i\}} \tilde h_{ij}\eps_j(\beta_0)(\hat\rho^{(2)}(z_j) -\rho(z_j) \big| \geq \eps_n\mid \Omega(\eps_n)\big) \leq \delta_n \] Since \(\Pr(\Omega(\eps_n))\to1\) this gives the result.

Relevant Moment Bounds

Moment Bounds for (ref)

Here I provide some lemmas that are useful in the proof of \Crefrange{lemma:lindeberg-interpolation}{lemma:approximate-distribution}

lemma[] Let \(\Delta_{1i}, \tilde\Delta_{1i}, \Delta_{2i}^a, \tilde\Delta_{2i}^a, \Delta_{2i}^b\),\(\tilde\Delta_{2i}^b\) be as in (ref). Then under (ref) and the moment conditions of (ref) there is a constant \(M > 0\) such that for any \(k=1,\dots,6\): \begin{align*} \E[|\Delta_{1i}|^k] &\leq M & \E[|\tilde\Delta_{1i}|^k] &\leq M \intertext{and for any \(k=1,\dots,3\):} \E[|\Delta_{2i}^a|^k] &\leq M\alpha^k & \E[|\tilde\Delta_{2i}^k|] &\leq M\alpha^k \\ \E[|\Delta_{2i}^b/\sqrt{n}|^k] &\leq M\alpha^k & \E[|\tilde\Delta_{2i}^b/\sqrt{n}|^k] &\leq M\alpha^k \end{align*}
proofFirst, since \[ \sum_{j=1}^n h_{ij}^2\E[(r_j - \E[r_j])^2 \leq\E[(\sum_{i=1}^n \tilde h_{ij}r_j)^2] \leq 1 \] the constants are bounded, \(\sum_{i=1}^n \tilde h_{ij}^2 \leq c\). Applying (ref) with \(X_i = h_{ij}r_j\) and \(X_i = h_{ij}\eps_j(\beta_0)\) we see that there is a constant \(A\) such that for any \(k = 1,\dots,6\) \begin{equation} \E\big[\big|\sum_{i=1}^n \tilde h_{ij}r_j\big|^k\big] \leq A\andbox \E\big[\big|\sum_{i=1}^n \tilde h_{ij}\eps_j(\beta_0)\big|^k\big] \leq A \end{equation} The bounds on \(\E[|\Delta_{1i}^k|]\) and \(\E[|\tilde\Delta_{1i}^k|]\) immediately follow from this result and the bounds on moments of \(r_i\) and \(\eps_i(\beta_0)\). The bounds on \(\E[|\Delta_{2i}^a|^k]\) and \(\E[|\tilde\Delta_{2i}^a|^k]\) also follow from (ref) after noting that there is a finite constant \(B\) such that: \[ \E[(\sum_{i=1}^n \tilde h_{ij}^2\eps_i^2(\beta_0))^k] \leq B \] Finally to bound \(\E[|\Delta_{2i}^b/\sqrt{n}|^k]\) and \(\E[|\tilde\Delta_{2i}^b/\sqrt{n}|^k]\) apply (ref) with \(v_j = \eps_j^2(\beta_0)\sum_{k\neq i,j}\tilde h_{jk}r_k\), noting that \(\E[|v_j|^3]\) is bounded by (ref).
lemma[] Let \(N\) and \(N_{-i}\) be defined as in (ref). Under (ref) and the moment conditions in (ref), there is a fixed constant \(M\) such that for all \(i=1,\dots,n\) and any \(k = 1,\dots,6\), \[ \E[|N|^k] + \E[|N_{-i}|^k] \leq M \]
proofWe show the bound for \(\E[|N|^k]\) and note that the bound for \(N_{-i}\) follows from symmetric logic. Write \(\eps_i(\beta_0) = \eta_i + \gamma_i\) where \(\gamma_i = \Pi_i(\beta - \beta_0)\) and \(\eta_i\) is mean zero. Decompose \(N = N_1 + N_2 + N_3\): \[ N_1 = \frac{1}{\sqrt{n}}\sum_{i=1}^n \eta_i \sum_{j=1}^n \tilde h_{ij} \dot r_j,\; N_2 = \frac{1}{\sqrt{n}}\sum_{i=1}^n r_i \sum_{j=1}^n \tilde h_{ji}\gamma_j, \andbox N_3 = \frac{1}{\sqrt{n}}\sum_{i=1}^n \eta_i \sum_{j=1}^n \tilde h_{ij}\E[r_j] \] where \(\dot r_j = r_j - \E[r_j]\). Since via (ref), \(\sum_{i=1}^n h_{ji}^2 \leq c\), and \(|\gamma_j| \leq c\), we can bound, \[ (\sum_{j=1}^n h_{ji}\gamma_j /\sqrt{n})^4 \leq (\frac{c}{\sqrt{n}}\sum_{i=1}^n |h_{ji}|)^4 \leq c^8 \implies (\sum_{j=1}^n h_{ji}\gamma_j /\sqrt{n})^6 \leq c^8 (\sum_{j=1}^n h_{ji}\gamma_j /\sqrt{n})^2 \] Under (ref), \(\E[N_2^2] \leq c\) while (ref) implies that \((\sum_{i=1}^n h_{ij}\E[r_j])^2 \leq c\) so that \(\E[N_3^2] \leq c^2\). An absolute bound on the higher moments of \(N_2\) then follows from an application of (ref) with \(X_i = r_i \sum_{j=1}^n h_{ji}\gamma_j/\sqrt{n}\). An absolute bound on the higher moments of \(N_3\) follows from symmetric logic. To bound higher moments of \(N_1\) define \(v_i = \sum_{j < i} \{\eta_i h_{ij}r_j + \dot r_i h_{ji}\eta_j\}\) and write \(N_1 = \frac{1}{\sqrt{n}}\sum_{i=2}^n v_i\). The sequence \(v_2,\dots,v_n\) is a martingale difference array. Via the same procedure as the bounds on \(\E[|\Delta_{1i}|^k]\) as in (ref) one can verify that there is a fixed constant \(M\) such that \(\E[|v_i|^k] \leq M\) for all \(k= 1,\dots,6\). The bound on the higher moments of \(N\) then follows from (ref). The bounds for moments of \(N_{-i}\) follow symmetric logic.
lemma[] Let \(\tilde N\) and \(\tilde D\) be defined as in (ref). Let \(f(\cdot, \tilde r)\) be the density function of \(\frac{\tilde N}{\tilde D^{1/2}}|\tilde r\). Under (ref) and the moment bounds of (ref), there is a constant \(M > 0\) such that \(\sup_x |f(x, \tilde r)| \leq M\) for almost all \(\tilde r\).
proofRecall that \[ \tilde N = \frac{1}{\sqrt{n}}\sum_{i=1}^n \tilde\eps_i(\beta_0)\sum_{j=1}^n \tilde h_{ij} \tilde r_j \andbox \tilde D^{1/2} = \sqrt{\frac{1}{n}\sum_{i=1}^n \kappa_i^2(\beta_0) (\sum_{j=1}^n \tilde h_{ij}\tilde r_j)^2} \] The distribution of \(\tilde\eps_i(\beta_0)|\tilde r_i\) is \[ \tilde\eps_i(\beta_0)|\tilde r \sim N\left(\mu_i(r_i), (1 - \rho_i^2)\Var(\eps_i(\beta_0)) \right) \] where \(\mu_i(r_i) = \Pi_i(\beta - \beta_0) + \frac{\Cov(\eps_i(\beta_0),r_i)}{\Var(r_i)}(r_i - \E[r_i])\) and \(\rho_i = \corr(\eps_i(\beta_0), r_i)\). Define \(\bar\Pi_i := \sum_{j=1}^n \tilde h_{ij} \tilde r_j\). Then, conditional on \(\tilde r\), \begin{equation} \frac{\tilde N}{\tilde D^{1/2}} \sim N\bigg(\frac{\frac{1}{\sqrt{n}}\sum_{i=1}^n \mu_i(r_i)\bar\Pi_i}{\sqrt{\frac{1}{n}\sum_{i=1}^n\kappa_i^2(\beta_0)\bar\Pi_i^2}}, \frac{\frac{1}{n}\sum_{i=1}^n (1 - \rho_i^2)\Var(\eps_i(\beta_0))\bar\Pi_i^2}{\frac{1}{n}\sum_{i=1}^n \kappa_i^2(\beta_0)\bar\Pi_i^2}\bigg) \end{equation} The maximum of the normal density is proportional to the inverse of the standard deviation so it suffices to show that the variance in (ref) is bounded away from zero. To this end, notice that under the moment bounds of (ref) and (ref) \[ (1-\delta^2)c^{-2}\leq (1-\rho_i^2)\frac{\Var(\eps_i(\beta_0))}{\kappa_i^2(\beta_0)} \leq c^2 \] By (ref) to this gives that the conditional variance is also larger than \((1 -\delta^2)c^{-2} > 0\).
lemma[] Let \(X_1,\dots,X_n\) be random variables such that \(\E[X_i] = \mu_i\) and \(\E[(\sum_{i=1}^n X_i)^2] \leq C\). Suppose that for any \(i = 1,\dots,n\) there is a constant \(U\) such that \[ \E[(X_i - \mu_i)^3] \leq U\E[(X_i - \mu_i)^2] \andbox \E[(X_i - \mu_i)^6]^{1/3}\leq U \E[(X_i -\mu_i)^2] \] Then \(\E[(\sum_{i=1}^n X_i)^6] \leq 64U^3C^3 + 32C^3\).
proofFirst write \begin{align*} \E[(\sum_{i=1}^n X_i)^2] = \sum_{i=1}^n \E(X_i - \mu_i)^2 + (\sum_{i=1}^n \mu_i)^2 \leq C \end{align*} To bound \(\E[(\sum_{i=1}^n X_i)^6]\) expand out \begin{align*} \E[(\sum_{i=1}^n X_i)^6] &= \E[(\sum_{i=1}^n (X_i - \mu_i) + \sum_{i=1}^n \mu_i)^6] \\ &\lesssim\E[(\sum_{i=1}^n (X_i - \mu_i))^6] + (\sum_{i=1}^n \mu_i)^6 \\ &= \sum_{i=1}^n \E[(X_i - \mu_i)^6] + \sum_{i=1}^n \sum_{j=1}^n \E[(X_i - \mu_i)^3(X_j - \mu_j)^3] \\ &\;\;\;\;\;\;+ \sum_{i=1}^n \sum_{j=1}^n \E[(X_i - \mu_i)^4(X_j - \mu_j)^2] \\ &+\sum_{i=1}^n \sum_{j=1}^n \sum_{k\neq i,j} \E[(X_i - \mu_i)^2(X_j-\mu_i)^2(X_k - \mu_k)^2] + (\sum_{i=1}^n \mu_i)^6\\ &\leq \sum_{i=1}^n \E[(X_i - \mu_i)^6] + \sum_{i=1}^n \sum_{j=1}^n \E[(X_i - \mu_i)^3]\E[(X_j - \mu_j)^3] \\ &+ \sum_{i=1}^n \sum_{j=1}^n \E[(X_i - \mu_i)^6]^{4/6}\E[(X_j - \mu_j)^6]^{2/6} \\ &+ \sum_{i=1}^n \sum_{j=1}^n \sum_{k\neq i,j} \E[(X_i - \mu_i)^6]^{1/3}\E[(X_j-\mu_i)^6]^{1/3}\E[(X_k - \mu_k)^6]^{1/3} \\ &+ C^3 \\ &= \bigg(\sum_{i=1}^n (\E[(X_i - \mu_i)^6])^{1/3}\bigg)^3 + \sum_{i=1}^n \sum_{j=1}^n \E[(X_i - \mu_i)^3]\E[(X_j - \mu_j)^3]+ C^3\\ &\leq \bigg(\sum_{i=1}^n (\E[(X_i - \mu_i)^6])^{1/3}\bigg)^3 + \bigg(\sum_{i=1}^n \E[(X_i - \mu_i)^3]\bigg)^2 + C^3\\ &\leq 2U^3\bigg(\sum_{i=1}^n \E[(X_i - \mu_i)^2]\bigg)^3 + C^3\\ &\leq 2U^3C^3 + C^3 \\ \end{align*} where the implied constant in the second line is 32 by an application of (ref), the third line comes from expanding out the power, the first inequality by application of Hölder's inequality, and the penultimate inequality comes from applying bounds on the third and sixth central moments in terms of the second moments.
lemma[] Let \(h = (h_1,\dots,h_n)\in \SR^n\) be such that \(\sum_{i=1}^n h_i^2 \leq b\). Suppose that \(X_1,\dots,X_n\) are such that \(\E[|X_i|^k] \leq M\) for all \(k = 1,2,3\). Then \[ \E\big[\big|\sum_{i=1}^n h_i^2 X_i\big|^3\big] \leq b^3M^3 \]
proofWe can expand out \begin{align*} \E\big[\big|\sum_{i=1}^n h_i^2X_i\big|^3] &\leq \sum_{i=1}^n \sum_{j=1}^n\sum_{k=1}^n h_i^2h_j^2h_k^2 \E[|X_i||X_j||X_k|] \\ &\leq M^3 \sum_{i=1}^n h_i^2 \sum_{j=1}^n h_j^2 \sum_{k=1}^n h_k^2\\ &\leq M^3\big(\sum_{i=1}^n h_i^2)^3 \leq c^3M^3 \end{align*}
lemma[] Let \(v_1,\dots,v_n\) be random variables such that \(\E[|v_i|^3] \leq M\) for all \(i=1,\dots,n\). Let \(h = (h_1,\dots,h_n)\in \SR^n\) be a vector of weights such that \(\|h\|_2 \leq c\). Then \[ \E\big[\big|\frac{1}{\sqrt{n}}\sum_{i=1}^n h_iv_i\big|^3\big] \leq c^3M \]
proofWe can expand out \begin{align*} \E\big[\big|\frac{1}{\sqrt{n}}\sum_{i=1}^n h_iv_i\big|^3\big] &\leq \frac{1}{n^{3/2}}\sum_{i=1}^n \sum_{j=1}^n\sum_{k=1}^n |h_i||h_j||h_k|\E[|v_i||v_j||v_k|] \\ &\leq \frac{M}{n^{3/2}}\sum_{i=1}^n |h_i| \sum_{j=1}^n |h_j| \sum_{k=1}^n |h_k| \leq \frac{M}{n^{3/2}}\|h\|_1^3 \leq Mc^3 \end{align*} where the second inequality follows from generalized Hölder's inequality, \[ |\E[fgh]| \leq (\E[|f|^3]\E[|g|^3]\E[|h|^3])^{1/3} \] and the fourth inequality from \(\|h\|_1 \leq \sqrt{n}\|h\|_2\).
lemma[] Let \(v_1,\dots,v_n\) be a martingale difference array such that \(\E[|v_i|^l] \leq M\) for all \(l=1,\dots,k\). Then there is a fixed constant \(C_k\) that only depends on \(k\) such that \[ \E[(\frac{1}{\sqrt{n}}\sum_{i=1}^n v_i)^k] \leq C_k M \]
proofWe move to apply (ref) with \(X_t = \sum_{i=1}^t v_i/\sqrt{n}\). \begin{align*} \E[(\frac{1}{\sqrt{n}}\sum_{i=1}^n v_i)^k] &\leq \E[(\max_{s \leq n} \sum_{t=1}^s X_s)^k] \\ &\leq C_k \E\big[\big(\sum_{i=1}^n v_i^2/n\big)^{k/2}\big]\leq C_k\E\big[\frac{1}{n}\sum_{i=1}^n v_i^k\big] \leq C_kM \end{align*} where the second inequality comes from (ref) and the third comes from an application of Jensen's inequality to the sample mean.

Useful Properties of Smooth Max

lemma[cck2013, Lemma A.2] For every \(1 \leq j,k,l \leq p\), \begin{align*} \partial_j F_\beta(z) &= \pi_j(z), & \partial_j\partial_k F_\beta(z) &= \beta w_{jk}(z), & \partial_j\partial_k\partial_l F_\beta(z) &= \beta^2 q_{jkl}(z) \end{align*} where for \(\delta_{jk} := \bm{1}\{j=k\}\), \begin{align*} \pi_j(z) &:= e^{\beta z_j}\bigg/\sum_{i=1}^n e^{\beta z_i},\hbox\;\;\;\; w_{jk} := (\pi_j\delta_{jk} - \pi_j\pi_k)(z) \\ q_{jkl}(z) &:= (\pi_j \delta_{jl}\delta_{jk} - \pi_j\pi_l \delta_{jk} - \pi_j\pi_k(\delta_{jl} + \delta_{kl}) + 2\pi_j\pi_k\pi_l)(z) \end{align*} Moreover, \begin{align*} \pi_j(z) &\geq 0, & \sum_{j=1}^p \pi_i(z) &= 1, & \sum_{j,k=1}^p |w_{jk}(z)| &\leq 2, & \sum_{j,k,l=1}^p |q_{jkl}| &\leq 6 \end{align*}
lemma[cck2013, {Lemma A.3}] For every \(x,z \in \SR^p\), \[ |F_\beta(x) -F_\beta(z)| \leq \max_{1 \leq j \leq p}|x_j - z_j|. \]
lemma[cck2013, {Lemma A.4}] Let \(\varphi(\cdot):\SR\to\SR\) be such that \(\varphi \in C_b^3(\SR)\) and define \(m:\SR^p \to \SR\), \(z \mapsto \varphi(F_\beta(z))\). The derivatives (up to the third order) of \(m\) are given \begin{align*} \partial_j m(z) &= (\partial g(F(\beta))\pi_j)(z) \\ \partial_j\partial_k m(z) &= (\partial^2 g(F_\beta)\pi_j\pi_k + \partial g(F_\beta)\beta w_{jk})(z) \\ \partial_j\partial_k\partial_l m(z) &= (\partial^3 g(F_\beta)\pi_j\pi_k\pi_l +\partial^2 g(F_\beta)\beta(w_{jk}\pi_l + w_{jl}\pi_k + w_{kl}\pi_j) + \partial g(F_\beta)\beta^2 q_{jkl})(z) \end{align*} where \(\pi_j, w_{jk}, q_{jkl}\) are as described in (ref).
lemma[cck2013, {Lemma A.5}] Define \(L_1(\varphi) = \sup_x |\varphi'(x)|, L_2(\varphi) = \sup_x |\varphi''(x)|\), and \(L_3(\varphi) = \sup_x |\varphi'''(x)|\). For every \(1 \leq j,k,l \leq p\), \begin{align*} |\partial_j \partial_k m(z)| \leq U_{jk}(z) \andbox |\partial_j\partial_k\partial_l m(z)| \leq U_{jkl}(z) \end{align*} where for \(W_{jk}(z) := (\pi_j\delta_{jk} + \pi_j\pi_k)(z)\), \begin{align*} U_{jk}(z) &:= (L_2 \pi_j\pi_k + L_1\beta W_{jk}(z) \\ U_{jkl}(z) &:= (L_3 \pi_j\pi_k\pi_l + L_2\beta(W_{jk}\pi_l + W_{jl}\pi_k + W_{kl}\pi_j) + L_1\beta^2 Q_{jkl})(z) \\ Q_{jkl}(z) &:= (\pi_j\delta_{jl}\delta_{jk} + \pi_j\pi_k\delta_{jk} + \pi_j\pi_k(\delta_{jl} + \delta_{kl}) + 2\pi_j\pi_k\pi_l)(z). \end{align*} Moreover, \begin{align*} \sum_{j,k=1}^p U_{jk}(z) \leq (L_2 + 2L_1\beta) \andbox \sum_{j,k,l=1}^p U_{jkl}(z) \leq (L_3 + 6L_2\beta + 6L_1 \beta^2). \end{align*}

Moment Bounds for (ref)

lemma[] Suppose that the moment conditions of (ref) hold and let \(N\) and \(D\) be as defined at the top of (ref) Then under \(H_0\), for any \(k\) there is a fixed constant \(C_k\) such that for any \(\ell = 1,\dots,d_x\) \[ \E[|N_\ell|^k] \leq C_k \andbox \E[|D_{\ell\ell}|^k] \leq C_k\log^{2k/a}(n) \]
proofLet \(\eta_{\ell i} = r_i - \E[r_i]\) and write \[ N_\ell= \underbrace{\frac{1}{\sqrt{n}}\sum_{i=1}^n \eps_i(\beta_0)\sum_{j=1}^n \tilde h_{ij}\eta_{\ell j}}_{N_\ell^1} + \frac{1}{\sqrt{n}}\underbrace{\sum_{i=1}^n \eps_i(\beta_0)\sum_{j=1}^n \tilde h_{ij}\E[r_{\ell j}]}_{N_\ell^2} \] To bound moments of \(N_\ell^1\) use the fact that \(N_\ell^1\) is a quadratic form in mean-zero \(a\)-sub-exponential variables. By (ref), \(N_\ell^1\) is therefore also \(a\)-sub-exponential with parameter \(a/2\); thus \((N_\ell^1)^{a/2}\) is sub-exponential and (ref) provides the moment bound for arbitrary moments. To bound moments of \(N_\ell^2\) we use the fact that \(\max_i \big|\sum_{j=1}^n \tilde h_{ij} \E[r_{\ell j}]\big|\) is bounded by assumption and apply Burkholder-Davis-Gundy ((ref)) after adding and subtracting \(\E[\eps_i(\beta_0)]\). To bound moments of \(D_{\ell\ell}\) we decompose \[ |D| \leq \frac{1}{n}\sum_{i=1}^n \eps_i^2(\beta_0) \max_{1 \leq i \leq n}\big|\sum_{j=1}^n h_{ij}r_j\big|^2 \] Apply (ref) to see that \(\sum_{j=1}^n h_{ij}r_j\) is \(\alpha\)-sub-exponential and (ref) to bound the RHS by a log-power of \(n\).

Matrix Derivative Lemmas

The purpose of this section is largely to establish some matrix derivative expressions that will be useful for the Lindeberg interpolation in

lemma[] Let \(D \in \SR^{d\times d}\) be a symmetric, real matrix such that \(\det(D)\neq 0\). Let \(N \in \SR^d\) be a vector. The derivatives up to the derivatives of quadratic form \(N'D^{-1}N\) are given. First Order: \begin{align*} \frac{\partial }{\partial N_l} &= 2\sum_{j=1}^d (D^{-1})_{jl} N_j,\hbox\hbox\hbox\hbox \frac{\partial }{\partial D_{l m}} = -2\sum_{j=1}^d \sum_{k=1}^d (D^{-1})_{jl}(D^{-1})_{k m} N_jN_k, \intertext{Second Order:} \frac{\partial^2}{\partial N_l N_m} &= 2(D^{-1})_{l m}, \hbox\hbox\hbox\hbox\frac{\partial^2}{\partial N_l \partial D_{pq}} = -2\sum_{j=1}^d (D^{-1})_{jp}(D^{-1})_{ql}N_j, \\ \frac{\partial^2 }{\partial D_{l m}\partial D_{qj}} &= \sum_{j=1}^d\sum_{k=1}^d\left\{ (D^{-1})_{l p}(D^{-1})_{qj})(D^{-1})_{km} + (D^{-1})_{kp}(D^{-1})_{mq}(D^{-1})_{l j}\right\}N_jN_k \intertext{Third Order:} \frac{\partial^3 }{\partial N_l \partial N_m \partial N_p} &= 0,\hbox\hbox\hbox\hbox \frac{\partial^3 }{\partial N_l \partial N_m \partial D_{pq}} = -2(D^{-1})_{l p}(D^{-1})_{qm} \\ \frac{\partial^3 }{\partial D_{l m}\partial D_{pq}\partial N_r} &= 2 \sum_{j=1}^d \bigg\{(D^{-1})_{l p}(D^{-1})_{qj}(D^{-1})_{rm} + (D^{-1})_{rp}(D^{-1})_{mq}(D^{-1}){l j}\bigg\}N_j \\ \frac{\partial^3 }{\partial D_{l m}D_{pq}D_{rs}} &= 2\sum_{j=1}^d \sum_{j=1}^d \bigg\{(D^{-1})_{l r}(D^{-1})_{ps}(D^{-1})_{qj}(D^{-1})_{km} + (D^{-1})_{l p}(D^{-1})_{qr}(D^{-1})_{js}(D^{-1})_{km} \\ &\hphantom{2\sum\sum\sum} +(D^{-1})_{l p}(D^{-1})_{qj}(D^{-1})_{kr}(D^{-1})_{ms} + (D^{-1})_{kr}(D^{-1})_{ps}(D^{-1})_{mq}(D^{-1})_{l j} \\ &\hphantom{2\sum\sum\sum } +(D^{-1})_{kp}(D^{-1})_{mr}(D^{-1})_{qs}(D^{-1})_{l j} + (D^{-1})_{rp}(D^{-1})_{mq}(D^{-1})_{l r}(D^{-1}){js}\bigg\}N_jN_k \end{align*}
proofThe derivative of an element of the the inverse of a matrix \(\mX\) can be expressed matrix-cookbook \begin{equation} \frac{\partial(\mX^{-1})_{kl}}{\partial \mX_{ij}} = - (\mX^{-1})_{ki}(\mX^{-1})_{jl} \end{equation} repeated application of this identity as well as the expression of the quadratic form \[ N'D^{-1}N = \sum_{j=1}^d\sum_{k=1}^d (D^{-1})_{jk}N_jN_k \] leads to the result, bearing in mind that the inverse of a symmetric matrix is symmetric.
lemma[] Let D be a symmetric positive definite matrix. Then, for any \(p > 3\), the derivatives of \((\det(D))^p\) are given up to the third order by \begin{align*} \frac{\partial\,(\det(D))^p}{\partial D_{lm}} &= p(\det(D))^{p-1}(D^{-1})_{lm} \\ \frac{\partial^2\,(\det(D))^p }{\partial D_{lm}\partial D_{pq}} &= \frac{p!}{(p-2)!}(\det(D))^{p-2}(D^{-1})_{pq}(D^{-1})_{lm} \\ &\;\;+ p(\det(D))^{p-1}(D^{-1})_{lp}(D^{-1})_{mq} \\ \frac{\partial^3\,(\det(D))^p }{\partial D_{lm}\partial D_{pq}\partial D_{rs}} &= \frac{p!}{(p-3)!}(\det(D))^{p-3}(D^{-1})_{rs}(D^{-1})_{pq}(D^{-1})_{lm} \\ &\;\; + \frac{p!}{(p-2)!}(\det(D))^{p-2}\bigg\{(D^{-1})_{pq}(D^{-1})_{lr}(D^{-1})_{ps} + (D^{-1})_{pr}(D^{-1})_{qs}(D^{-1})_{l m} \\ &\hphantom{\;\;+\frac{p!}{(p-2)!}(\det(D))^{p-2}\bigg\{(D^{-1})_{pq}(D^{-1})_{lr}(D^{-1})_{ps}}\;\,+(D^{-1})_{rs}(D^{-1})_{lp}(D^{-1})_{mq}\bigg\}\\ &\;\; + p(\det(D))^{p-1}\bigg\{(D^{-1})_{lr}(D^{-1})_{qs}(D^{-1})_{mq} + (D^{-1})_{lp}(D^{-1})_{mr}(D^{-1})_{q s}\bigg\} \end{align*}
proofWe can express the derivative of the detrminant matrix-cookbook, \begin{equation} \begin{split} \frac{\partial,\det(\mX)}{\partial\, \mX_{ij}}= \det(\mX)(\mX^{-1})_{ij} \end{split} \end{equation} Repeated application of this and (ref) yields the result.
lemma[] For any \(p > 4\) define the function \(\gamma(N,\vec(D)):\SR^d \times \SR^{d^2}\) by \[ \gamma(N,\vec(D)) := \begin{cases} (\det(D))^p(N'D^{-1}N - c) &\text{if } \det(D) \neq 0 \\ 0 &\text{if }\det(D) = 0 \end{cases} \] This function is thrice continously differentiable. Futher the \(k^\text{th}\) moments of all partial derivatives of this function up to the third order are bounded \[ \E[(\partial^\alpha \gamma(N,\vec(D))^k] \leq C_k (\max_{\iota \leq d}\E[|D_{\iota\iota}|^{2pdk}] \vee \max_{\iota \leq d}\E[|N_{\iota\iota}|^{6k}) \] where \(C_k\) is a positive constant that only depends on \(k\) and \(d\).
proofThe first statement is clear by examination of the derivatives in (ref) as well as the inequality (ref) below. For the moment bounds, we may extensive use of following bounds on elements of \(D^{-1}\) for a positive-definite \(D^{-1}\): \begin{equation} \begin{split} |\det(D)(D^{-1})_{jk}| \leq \det(D)\trace(D^{-1}) &\leq d\lambda_{\max}(D^{-1})\big(\prod_{m=1}^d \lambda_m(D)\big) \\ &= d\prod_{m=2}^d \lambda_m(D) \\ &\leq d\big(\sum_{m=2}^d \lambda_m(D)\big)^{d-1} \\ &\leq d(\trace(D))^{d-1} \end{split} \end{equation} where the first inequality uses the fact that the largest element of a positive semidefinite matrix is on the diagonal and the fact that the diagonal elements of a positive semidefinite matrix are weakly positive, the second inequality uses the fact that the trace is the sum of the eigenvalues and the determinant is the product of the eigenvalues, the equality comes from \(\frac{1}{\lambda_{\min}(D)} = \lambda_{\max}(D^{-1})\), the third inequality uses the AM-GM inequality and the fourth again uses that the trace is the sum of the (weakly positive) eigenvalues. The moment bounds follow from (ref) and the expressions in (ref). We give an example of how this is done for the first order derivatives, higher order derivatives follow from similar logic. For the following let \(A\) be an arbitrary random variable. First Order. \begin{align*} \E\bigg|A\frac{\partial \gamma}{\partial N_l}\bigg|^k &\lesssim \sum_{j=1}^d\E|(\trace(D))^{kdp}N_j^kA^k| \\ &\lesssim \sum_{j=1}^d\sum_{\iota =1}^d\E[D_{\iota\iota}^{kdp}N_j^kA^k] \\ &\leq \sum_{j=1}^d\sum_{\iota =1}^d \gamma^{2kdp} \E[N_j^{2k}A^{2k}]\\ \E\bigg|A \frac{\partial \gamma}{\partial D_{lm}}\bigg|^k &= p\E\bigg|A \det(D)^{p-1}\sum_{j=1}^d\sum_{j' = 1}^d (D^{-1})_{lm}(D^{-1})_{jj'}N_jN_{j'} \bigg|^k\\ &\lesssim p\sum_{j=1}^d\sum_{j'=1}^d \E[|(\trace(D))^{2k(d-1) + (p-3)kd}A^kN_j^kN_{j'}^k| \\ &\leq \sum_{j=1}^d \sum_{j'=1}^d \gamma^{2kd(p-1)}\E[A^{2k}N_j^{2k}N_{j'}^{2k}] \end{align*}

Technical Lemmas

Probability Lemmas

lemma[] Let \(X_n\) be a sequence of random variables such that \(X_n = o_p(1)\), that is for any \(\delta > 0\), \(\Pr(|X_n| \geq \delta) \to 0\). Then, there is a sequence \(\delta_n \to 0\) such that \(\Pr(|X_n| \geq \delta_n) \to 0\).
proofTake a preliminary sequence \(\tilde\delta_n \to 0\) and define \[ \tilde n_j = \inf\{n : \Pr(|X_n| > \tilde\delta_j) < \tilde\delta_j\} \] Because \(\Pr(|X_n| >\delta) \to 0\) for any fixed \(\delta\), we know that \(n_j\) is finite. Define a new sequence \(\delta_n \to 0\) as below: \begin{equation} \begin{split} \delta_n = \begin{cases} 1 &if 0 \leq n < \tilde n_1 \\ \tilde\delta_i &if \tilde n_i \leq n < \tilde n_{i+1} \end{cases} \end{split} \end{equation} By construction, this sequence satisfies \(\Pr(X_n \geq \delta_n) \leq \delta_n\) whenever \(n \geq n_1\).
lemma[] Suppose that \(X_1,\dots,X_n\) are \(\alpha\)-subexponential such that \(\Pr(|X_i| \geq t) \leq 2\exp(-t^\alpha/K)\) for all \(t \geq 0\) and fixed constants \(K\). For any \(p \geq 1\) there is a constant \(C\) that depends only on \(p,K\) such that: \[ \E\bigg[\max_{i \leq n} \frac{|X_j|^p}{(1 + \log i)^{p/\alpha}}\bigg] \leq C \] As a consequence \[ \E\big[\max_{i \leq n} |X_i|^p\big] \leq C(\log n)^{p/\alpha} \]
proofArgument below is provided for \(\alpha = 1\). This can be extended to \(\alpha \neq 1\) by noting that if \(\Pr(|X_i| \geq t) \leq 2\exp(-t^\alpha /K)\) for some \(\alpha > 0\) then \(\Pr(|X_i|^{\alpha} \geq t) \leq 2\exp(-t/K)\). \begin{align*} \E\max_{i \leq n} \frac{|X_i|^p}{(1 + \log i)^p} &= \int_{0}^{\infty} \Pr\bigg(\max_i \frac{|X_i|^p}{(1 + \log i)^p} > t\bigg)\,dt \\ &= \int_{0}^{2^{p/\alpha}} \Pr\bigg(\max_i \frac{|X_i|^p}{(1 + \log i)^p} > t\bigg)\,dt + \int_{2^{p/\alpha}}^{\infty} \Pr\bigg(\max_i \frac{|X_i|^p}{(1 + \log i)^p} > t\bigg)\,dt \\ &\leq 2^p + \int_{2^{p/\alpha}}^{\infty} \sum_{i=1}^n \Pr\bigg(\frac{|X_i|}{1 + \log i}> t^{1/p}\bigg)\,dt \\ &\leq 2^p + \int_{2^p}^{\infty} \sum_{i=1}^n 2\exp\bigg(-\frac{t^{1/p}(1 + \log i)}{K}\bigg)\,dt \\ &= 2^p + 2\sum_{i=1}^n \int_{2^p}^{\infty} \exp\bigg(-\frac{t^{1/p}}{K}\bigg)i^{-t^{1/p}}\,dt \\ &\leq 2^p+ 2\sum_{i=1}^n \int_{2^p}^{\infty}\exp(-t^{-1/p}/K)i^{-2}\,dt \\ &\leq 2^p+ 2\bigg(\sum_{i=1}^n i^{-2}\bigg)\bigg( \int_{2^p}^{\infty}\exp(-t^{-1/p}/K)\,dt \bigg) \end{align*} Both the integral and the summation are bounded, which gives the result.

Matrix Lemmas

lemma[] Given a matrix \(M\) and a matrix \(P\) of full rank, the matrix \(M\) and the matrix \(P^{-1}MP\) have the same eigenvalues.
proofSuppose \(\lambda\) is a eigenvalue of \(P^{-1}MP\) with eigenvector \(p\). Then \[ P^{-1}MPv = \lambda v \implies M(Pv) = \lambda Pv \] Hence \(Pv\) is an eigenvector of \(M\) with eigenvalue \(\lambda\). Similarly, given an eigenvector \(v\) of \(M\), it can be shown that \(P^{-1}v\) is an eigenvector of \(P^{-1}MP\); \[ P^{-1}MP(P^{-1}v) = P^{-1}Mv = \lambda P^{-1}v \]
lemma[] Let \(A \in \SR^{n\times n}\) and \(B\in \SR^{n\times n}\) be real symmetric positive semidefinite matrices. For an arbitary square matrix \(M\) let \(\lambda_k(M)\) denote the \(k^\text{\tiny th}\) largest eigenvalue of \(M\). Then for any \(k = 1,\dots,n\): \[ \lambda_k(A)\lambda_n(B) \leq \lambda_k(AB) \leq \lambda_k(A)\lambda_1(B) \]
lemma[] Let \(D \in \SR^{n \times n}\) be a diagonal real matrix such that \(d_{ii} \in [u, U]\) for all \(i =1 ,\dots, n\). Let \(A \in \SR^{n\times n}\) be a symmetric real matrix. For an arbitrary square matrix \(M\), let \(\lambda_k(M)\) denote the \(k^\text{\tiny th}\) largest eigenvalue of \(M\). Then for any \(k = 1,\dots,n\): \[ u\lambda_k(A^2) \leq \lambda_k(ADA) \leq U \lambda_k(A^2) \]
proofConsider any vector \(a \in \SR^n\) and define \(\va = a'H\). Then \begin{align*} \alpha'HDH\alpha = \va'D\va = \sum_{i=1}^n d_{ii}(\va_i)^2 &\in \bigg[u \sum_{i=1}^n (\va_i)^2,\, U \sum_{i=1}^n (\va_i)^2 \bigg] \\ &= \bigg[u \times a'H^2a,\, U \times a'H^2a\bigg] \end{align*} The result then follows from an application of Courant-Fischer-Weyl min-max principle.
lemma[] Let \(X_1,\dots,X_n\) denote i.i.d standard normal random variables and \(a_1,\dots,a_n\) denote weakly positive constants. Then \[ \Pr \left(\sum_{i=1}^n a_i X_i^2 \leq \eps \sum_{i=1}^n a_i\right) \leq \sqrt{e\eps} \]

Miscellaneous Lemmas

lemma[] Let \(a_1,\dots,a_n\) and \(b_1,\dots,b_n\) be two sequences of real numbers. If \(a_i \leq Ub_i\) for some \(U > 0\), then \(\sum_i a_i/\sum_i b_i \leq U\). Conversely if \(a_i \geq L b_i\) for some \(L > 0\) then \(\sum_i a_i/\sum_i b_i \geq L\).
proofReplace \(a_i \leq Ub_i\) for the upper bound and \(a_i \geq Lb_i\) for the lower bound.

The following is a standard bound, but it is used a lot so it is restated here.

lemma[] Let \(a_1,\dots,a_m\) be constants and \(p > 1\). Then \[ |a_1 + \dots a_m|^p \leq m^{p-1}\sum_{i=1}^m |a_i|^p \]
proofApply Hölder's inequality with \(\frac{1}{p} + \frac{p-1}{p} = 1\) to the vectors \((a_1,\dots,a_m)\in \SR^m\) and \((1,\dots,1)\in \SR^m\)

Assorted Results from Literature

Concentration Inequalities and Tail Bounds

theorem[Gotze-subexponential-polynomial*{Theorem 1.2}] Let \(X_1,\dots,X_n\) be independent random variables satisfying \(\|X_i\|_{\Psi_a}\leq M\) for some \(a \in (0,1] \cup \{2\}\) and let \(f:\SR^n\to\SR\) be a polynomial of total degree \(D \in \SN\). Then for all \(t > 0\); \[ \Pr(|f(X) - \E[f(X)]| \geq t) \leq 2\exp\bigg(- \frac{1}{C_{D,a}}\min_{1 \leq d \leq D}\bigg(\frac{t}{M^d\|\E f^{(d)}(X)\|_{\text{HS}}}\bigg)^{a/d}\bigg) \] In particular, if \(\|\E f^{(d)}(X)\|_{\text{HS}} \leq 1\) for \(d = 1,\dots D\), then \[ \E\exp\bigg(\frac{C_{D,a}}{M^a}|f(X)|^{\frac{a}{D}}\bigg) \leq 2, \] or equivalently \[ \|f(X)\|_{\Psi_{\frac{a}{D}}} \leq C_{d,a}M^D \]
theorem[Hoeffding's Inequality] Let \(X_1,\dots,X_n\) be independent, mean-zero sub-gaussian random variables, and let \(a = (a_1,\dots,a_n)\in \SR^n\). Then, for every \(t \geq 0\), we have \[ \Pr\bigg\{\bigg|\sum_{i=1}^n a_iX_i\bigg|\geq t\bigg\} \leq 2\exp\bigg(-\frac{ct^2}{K^2\|a\|_2^2}\bigg) \] where \(K = \max_i \|X_i\|_{\psi_2}\).
theorem[Burkholder-Davis-Gurdy for Discrete Time Martingales] For any \(1 \leq k < \infty\) there exist positive constants \(c_k\) and \(C_k\) such that for all local martingales with \(X_0 = 0\) and stopping times \(\tau\) \[ c_k \E\big[\big(\sum_{t=1}^\tau (X_t - X_{t-1})^2\big)^{k/2}\big] \leq \E\big[(\sup_{t \leq \tau} X_t)^k\big] \leq C_k\E\big[\big(\sum_{t=1}^\tau (X_t - X_{t-1})^2\big)^{k/2}\big] \]

Anticoncentration Bounds

Let \(\xi \in \SR^n\) follow a normal distribution on \(\SR^n\) with mean zero and covariance matrix \(\Sigma_\xi\). Order the eigenvalues of \(\Sigma_\xi\) in non-increasing order \(\lambda_{1\xi} \geq \lambda_{2\xi} \geq ... \geq \lambda_{n\xi}\). Define the quantities \[ \Lambda_{k\xi}^2 = \sum_{j=k}^\infty \lambda_{j\xi}^2,\;\; k =1,2 \]

theorem[anticoncentration-bernoulli-gotze-et-al, {Theorem 2.6}] Let \(\xi\) be a gaussian element with zero mean and covariance \(\Sigma_\xi\). Then it holds for any \(\va \in \SR^n\) that \[ \sup_{x \geq 0} p_\xi(x,\va) \lesssim (\Lambda_{1\xi}\Lambda_{2\xi})^{-1/2} \] where \(p_\xi(x,a)\) denotes the p.df of \(\|\xi - \va\|^2\).

We use the following anticoncentration lemma from Nazarov2003 noted in CCK-2018-HDCLT.

lemmaLet \(Y = (Y_1,\dots,Y_p)'\) be a centered Gaussian random vector in \(\SR^p\) such that \(\E[Y_j^2] \geq b\) for all \(j = 1,\dots,p\) and some constant \(b > 0\). Then for every \(y \in \SR^p\) and \(a > 0\), \[ \Pr(Y \leq y + a) - \Pr(Y \leq y) \leq Ca\sqrt{\log(p)} \] where \(C\) is a constant only depending on \(b\).

Gaussian Comparasions and Approximations

We also use the following gaussian approximation results from \cites{belloni2018highdimensional,CCK-2018-HDCLT}. Let \(X_1,\dots,X_n \in \SR^p\) be independent, mean zero, random vectors and let \(Y_1,\dots,Y_n \in \SR^p\) be independent random vectors such that \(Y_i \sim N(0, \E[X_iX_i'])\). Suppose that the researcher does not directly observe \(X_1,\dots,X_n\) but instead observes noisy estimates \(\widehat X_1,\dots,\widehat X_n \in \SR^p\).

Define the sums

align*[align* omitted — 120 chars of source]

Let \(\calA^{\text{re}}\) be the class of all hyperrectangles in \(\SR^p\); that is, \(\calA^{\text{re}}\) consists of all sets \(A\) of the form

equation*[equation* omitted — 128 chars of source]

for some \(-\infty \leq a_j \leq b_j \leq \infty\), \(j = 1,\dots, p\). Define \[ \rho_n(\calA^{\text{re}}) \coloneqq \sup_{A \in \calA^{\text{re}}}\big|\Pr(S_n^X \in A) - \Pr(S_n^Y \in A)\big| \] Bounding \(\rho_n(\calA^{\text{re}})\) relies on the following moment conditions:

assumption[] Suppose there are constants \(B_n \geq 1\), \(b > 0,q > 0\) such that \begin{enumerate}[(i)] • \(n^{-1}\sum_{i=1}^n \E[X_{ij}^2] \geq b\) for all \(j = 1,\dots,p\)\(n^{-1}\sum_{i=1}^n \E[|X_{ij}|^{2+k}] \leq B_n^k\) for all \(j = 1,\dots,p\) and \(k = 1,2\). • \(\E[(\max_{1 \leq j \leq p}|X_{ij}|/B_n)^4] \leq 1\) for all \(i = 1,\dots, n\) and \(\left(\frac{B_n^4\ln^7(pn)}{n}\right)^{1/6} \leq \delta_n\). \end{enumerate}

as well as the following bounds on the estimation error

assumption[] The estimates \(\hat X_1,\dots,\hat X_n\) satisfy \[ \Pr\left(\max_{1 \leq j \leq p} \E_n[(\widehat X_{ij} - X_{ij})^2] > \delta_n^2/\log^2(pn)\right)\leq \beta_n \]
theorem[belloni2018highdimensional, {Theorem 2.1}] Suppose that (ref) hold. Then there is a constant \(C\) which depends only on \(b\) such that \[ \rho_n(\calA^{\text{re}}) \leq C\{\delta_n + \beta_n\} \]

Let \(e_1,\dots,e_n \overset{\text{iid}}{\sim}N(0,1)\) be generated independently of the data. A gaussian bootstrap draw is defined \[ S_n^{X,\star} \coloneqq \frac{1}{\sqrt{n}}\sum_{i=1}^n e_i \widehat X_i \]

theorem[belloni2018highdimensional, {Theorem 2.2}] Suppose that (ref) hold. Then there is a constant \(C\) which depends only on \(b\) such that \[ \sup_{A \in \calA^{\text{re}}} \big|\Pr_e(S_n^{X,\star} \in A) - \Pr(S_n^Y \in A) \big| \leq C\delta_n \] with probability at least \(1 - \beta_n - (\log n)^{-2}\) where \(\Pr_e(\cdot)\) denotes the probability measure only taken with respect to the variables \(e_1,\dots,e_n\) conditional on the data used to estimate \(\widehat X\).