EconBase
← Back to paper

Testing Many Restrictions Under Heteroskedasticity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,269 characters · 20 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Many Restrictions Under Heteroskedasticity

\if00 \fi

\if10 \fi

abstractWe propose a hypothesis test that allows for many tested restrictions in a heteroskedastic linear regression model. The test compares the conventional F statistic to a critical value that corrects for many restrictions and conditional heteroskedasticity. This correction uses leave-one-out estimation to correctly center the critical value and leave-three-out estimation to appropriately scale it. The large sample properties of the test are established in an asymptotic framework where the number of tested restrictions may be fixed or may grow with the sample size, and can even be proportional to the number of observations. We show that the test is asymptotically valid and has non-trivial asymptotic power against the same local alternatives as the exact F test when the latter is valid. Simulations corroborate these theoretical findings and suggest excellent size control in moderately small samples, even under strong heteroskedasticity. Keywords: linear regression, ordinary least squares, many regressors, leave-out estimation, hypothesis testing, high-dimensional models. JEL codes: C12, C13, C21

\thispagestyle{empty}

Introduction

One of the central tenets in modern economic research is to consider models that allow for flexible specifications of heterogeneity and to establish whether {meaningful heterogeneity is present or absent in a particular empirical setting}. For example, abowd1999high study whether there is firm-specific heterogeneity in a linear model for individual log-wages, card2016bargaining,card2016firms ask if this heterogeneity varies by the individual's {gender or} education, and lachowska2019firm investigate whether the firm-specific heterogeneity is constant over time. Other work relies on similarly flexible models to investigate the presence of heterogeneity in health economics finkelstein2016sources and to study neighborhood effects chetty2018impacts. In all these examples, the absence of a particular dimension of heterogeneity corresponds to a hypothesis that imposes hundreds or thousands of restrictions on the model of interest. The present paper provides a tool to conduct a test of such hypotheses. {In contemporary work, kline2022systemic apply our proposed test in a study of discrimination among U.S. employers.}

We develop a test for hypotheses that impose multiple restrictions and establish its asymptotic validity in a heteroskedastic linear regression model where the number of tested restrictions may be fixed or increasing with the sample size. In particular, we allow for the number of restrictions and the sample size to be proportional. The exact F test{, which compares the F statistic to a quantile of the F distribution,} fails to control size in this environment. Instead, our proposed test rejects the null hypothesis if the F statistic exceeds a critical value that corrects for many restrictions and conditional heteroskedasticity. This critical value is a recentered and rescaled quantile of what is naturally called the {\it F-bar distribution} as it describes the distribution of a chi-bar-squared random variable divided by an independent chi-squared random variable over its degrees of freedom.\footnote{{Chi-bar-squared is a standard name used to describe a mixture of chi-squared distributions. See, e.g., chibar2 who studies asymptotic properties of chi-bar-squared distributions.}} This family of distributions can approximate both the finite sample properties of the F statistic under homoskedastic normal errors and---after recentering and rescaling---the asymptotic distribution of the F statistic in the presence of conditional heteroskedasticity and few or many restrictions.

The large sample validity of our proposed test holds uniformly in the number of regressors and tested restrictions. In combination with the F-bar distribution, the key to this uniformity is our proposed location and variance estimators that are used to recenter and rescale the critical value. {The location estimator utilizes unbiased leave-one-out estimators for individual error variances, while the variance estimator utilizes unbiased leave-three-out estimators for products of these variances. While the product of leave-one-out estimators is biased for the product of variances because of mutual dependence, dropping three observations in successive fashion breaks the dependence between estimators of individual error variances, and hence provides unbiasedness of their products.} The use of leave-three-out estimation to implement this idea is novel in the literature. Because the essential elements of the test are built on leave-out machinery, we will at times and for brevity refer to the proposed test using the acronym LO.

The LO test has exact asymptotic size when the regression design has full rank after leaving any combination of three observations out of the sample. This condition is satisfied in models with many continuous regressors and only a few discrete ones. However, the condition can fail when many discretely valued regressors are included, as occurs for models with {fixed individual or group effects.} { With group effects, {in particular,} leave-three-out {may not exist} when group sizes are two or three.} To handle such cases, the proposed test uses estimators for the products of individual error variances that are intentionally biased upward when the unbiased leave-three-out estimators do not exist. This construction ensures {large sample validity} but can potentially {lead to a slightly conservative test} when {a large fraction} of the leave-three-out estimators do not exist.

Using both theoretical arguments and simulations, huber1973robust and berndt1977conflict have highlighted the importance of allowing the number of regressors and potentially the number of tested restrictions to increase with sample size when studying asymptotic properties of inference procedures. {The latter paper specifically documents cases where asymptotically equivalent classical tests yield opposite outcomes when the number of tested restrictions is somewhat large.} Despite these early cautionary tales, most inference procedures that allow for proportionality between the number of regressors, sample size, and potentially the number of restrictions, are of a more recent vintage. Here, we survey the ones most relevant to the current paper and refer to anatolyev2019survey for a more extensive review of the literature.\looseness=-1\footnote{In analysis of variance contexts, which are special cases of linear regression, akritas2004anova and zhou2017hidim propose heteroskedasticity robust tests for equality of means that are, however, specific to their models. An expanding literature considers (outlier) robust estimation of linear high-dimensional regressions elkaroui2013hidim but does not provide valid tests of many restrictions.}

In homoskedastic regression models, anatolyev2012inference and calhoun2011many propose various corrections to classical tests that restore asymptotic validity in the presence of many restrictions. In heteroskedastic regressions with one tested restriction and many regressors, cattaneo2017inference show that the use of conventional Eicker-White standard errors and their “almost-unbiased” variations mackinnon2012hetero does not yield asymptotic validity. This failure may be viewed as a manifestation of the incidental parameters problem. To overcome this problem, cattaneo2017inference and subsequently anatolyev2018almostunbiased propose new versions of the Eicker-White standard errors, which restore size control in large samples. However, these proposals rely on the inversion of $n$-by-$n$ matrices ($n$ denotes sample size) that may fail to be invertible in examples of practical interest horn1975estimating,verdier2016estimation. rao1970heterovar's unbiased estimator for individual error variances is closely related to cattaneo2017inference's proposal and suffers from the same existence issue.\footnote{{{A recent use of rao1970heterovar's MINQUE estimator is juhl2014hetero, where it is incorporated into a test of heterogeneity in short panels with fixed effects. Here, the MINQUE estimator is applied to each cross-sectional unit, and its use therefore imposes a growing number of invertibility requirements.}}}

In homoskedastic regression models, Anatolyev (2012) and Calhoun (2011) propose various corrections to classical tests that restore asymptotic validity in the presence of many restrictions. In heteroskedastic regressions with one tested restriction and many regressors, Cattaneo et al. (2018b) show that the use of conventional Eicker-White standard errors and their “almost-unbiased” variations does not yield asymptotic validity. This failure may be viewed as a manifestation of the incidental parameters problem. To overcome this problem, Cattaneo et al. (2018b) and subsequently Anatolyev (2018) propose new versions of the Eicker-White standard errors, which restore size control in large samples. However, these proposals rely on the inversion of n-by-n matrices (n denotes sample size) that may fail to be invertible in examples of practical interest. Rao (1970)’s unbiased estimator for individual error variances is closely related to Cattaneo et al. (2018b)’s proposal and suffers from the same existence issue.

kline2018leave propose instead a version of the Eicker-White standard errors that relies only on leave-one-out estimators of individual error variances and show that its use leads to asymptotic size control when testing a single restriction.\footnote{jochmans2018variance additionally uses simulations to investigate the behavior of this variance estimator.} While this conclusion extends to hypotheses that involve a fixed and small number of restrictions through the use of a heteroskedasticity-robust Wald test, it fails to hold in cases of many restrictions. When testing many coefficients equal to zero, kline2018leave note that those leave-one-out individual variance estimators can be used to center the conventional F statistic\footnote{The use of leave-one-out estimation has a long tradition in the literature on instrumental variables phillips1977bias, and our test shares an algebraic representation with the adjusted J test analyzed in chao2014testing kline2018leave. An attractive feature of relying on leave-one-out is that challenging estimation of higher order error moments can be avoided, which is in contrast to the tests of calhoun2011many and anatolyev2013manyexogenous.} and propose a rescaling of the statistic that relies on successive sample splitting {kline2018leave as a tool of breaking dependence among different estimates when error variances enter as pairwise products. However, first, sample splitting places restrictions on the data that will fail when the number of regressors is larger than half of the sample size.\footnote{This phenomenon is akin to the situation in cattaneo2017inference, in which the `Hadamard square' of the orthogonal projection matrix may be non-invertible. cattaneo2017inference rule out this possibility by imposing a sufficient condition such that the number of covariates is no larger than half of the sample size.} Second, sample splitting means that the error variances are estimated only from a part of the sample, which is clearly inefficient and undesirable for conditional objects that require the use of as many observations as possible. Third, sample splitting may be undesirable because different ways of splitting the sample can lead to opposite conclusions.

We instead take another route and utilize information for estimation of error variances in the whole sample. In order to remove the dependence among individual estimates, we appeal to leave-three-out estimation instead of sample splitting. Importantly, leave-three-out estimation places much fewer restrictions on the number of regressors than sample splitting, exploits available sample information more efficiently, and does not require a researcher to choose a way to split the sample. Additionally, the robustified version of the LO test using the F-bar distribution enables asymptotic size control uniformly in the number of restrictions.}

We provide a theoretical study of the power properties under local and global alternatives. Under local alternatives, the asymptotic power curve of the proposed LO test is parallel to that of the exact F test when the latter is valid, e.g., under homoskedastic normal errors. While the curves are parallel, the LO test tends to have power somewhat below the exact F test. This loss in power stems from the estimation of individual error variances and can be viewed as a cost of using a test that is robust to general heteroskedasticity. This cost is largely monotone in the number of tested restrictions and disappears when the number of restrictions is small relative to sample size.

We also conduct a simulation study that documents excellent performance of the LO test in small and moderately sized samples. We document that the LO test delivers nearly exact size control in samples as small as $100$ observations in both homoskedastic and heteroskedastic environments. On the other hand, conventional tools such as the Wald test and the exact F test can exhibit severe size distortions and reject a true null with near certainty for some configurations. These findings are documented using two simulation settings: one with continuous regressors only, and one with a mix of both continuous and discrete regressors. In the latter setting, roughly $7\%$ of observations cause a full rank failure when leaving up to three observations out, but the proposed test shows almost no conservatism even in this adverse environment. When both the LO and exact F tests are valid, the simulations document a power loss that varies between being negligible and up to roughly $15$ percentage points, depending on the type of deviation from the null and sample size. For many applications, this range of power losses is a small cost to incur for being robust to heteroskedasticity.

The paper is organized as follows. Section (ref) introduces the setup and the proposed critical value in samples where all the leave-three-out estimators exist, while Section (ref) analyzes the asymptotic size and power of the LO test for such samples. Section (ref) describes the critical value for use in samples where the design loses full rank after leaving certain triples of observations out. Section (ref) discusses the results of simulation experiments, and Section (ref) concludes. Proofs of theoretical results and some clarifying but technical details are collected in the online supplemental Appendix. An R package AScode that implements the proposed test is available online.

Leave-out test

Consider a linear regression model

align[align omitted — 148 chars of source]

where an intercept is included in the regression function $\boldsymbol x_i'{\boldsymbol \beta}$ and the $n$ observed random vectors $\{(y_i,\boldsymbol x_i')'\}_{i=1}^n$ are independent across $i$. The dimension of the regressors $\boldsymbol x_i \in \mathbb{R}^m$ may be large relative to sample size {with $m < n$}, and there is conditional heteroskedasticity in the unobserved errors: {

align[align omitted — 128 chars of source]

The conditional variances are assumed to exist with no restrictions placed on the functional form, as in kline2018leave.}

The hypothesis of interest involves $r \leq m$ linear restrictions

align[align omitted — 89 chars of source]

where the matrix $\boldsymbol R \in \mathbb{R}^{r\times m}$ has full row rank $r$, and $\boldsymbol q \in \mathbb{R}^r$. Both $\boldsymbol R$ and $\boldsymbol q$ are specified by the researcher. Specifically, they are assumed to be known and are allowed to depend on the observed regressors. The space of alternatives is $H_A : \ \boldsymbol R{\boldsymbol \beta} \neq \boldsymbol q$.

The attention of the paper is on settings where the design matrix $\boldsymbol S_{xx} = \sum_{i=1}^n \boldsymbol x_i \boldsymbol x_i'$ has full rank so that $\hat{\boldsymbol \beta}=\boldsymbol S_{xx}^{-1} \sum_{i=1}^n \boldsymbol x_i y_i$, the ordinary least squares (OLS) estimator of $\boldsymbol \beta$, is defined. For compact reference, we define the degrees-of-freedom adjusted residual variance

align[align omitted — 145 chars of source]
rem{ We maintain in this paper that observations are independent across $i$, as it facilitates simplicity when discussing some of our high-level conditions. {We conjecture that the results of the paper} continue to hold under the weaker assumption that the error terms are mean zero and independent across $i$ when conditioning on all the regressors $\{\boldsymbol{x}_i\}_{i=1}^n$. }

Test statistic

Our proposed test rejects $H_0$ for large values of Fisher's F statistic,

align[align omitted — 270 chars of source]

which is a monotone transformation of the likelihood ratio statistic when the regression errors are homoskedastic normal. Since we do not impose normality, $F$ may be viewed as a quasi likelihood ratio statistic. {The behavior of the test is governed by the numerator of $F$, which we denote by ${\cal F}$:

align[align omitted — 243 chars of source]

By taking this statistic as a point of departure, we are able to construct a critical value that ensures size control in the presence of heteroskedasticity and an arbitrary number of restrictions.\footnote{An alternative approach might have taken a heteroskedasticity-robust Wald statistic $\boldsymbol R \boldsymbol S_{xx}^{-1} (\sum_{i=1}^n \boldsymbol x_i \boldsymbol x_i' \hat\varepsilon_i^2) \boldsymbol S_{xx}^{-1} \boldsymbol R'$, where $\{\hat \varepsilon_{i}\}_{i=1}^n$ are OLS residuals, in a similar attempt to ensure validity when the number of restrictions is proportional to the sample size. However, in such environments, any heteroskedasticity-robust Wald statistic relies on the inverse of a high-dimensional covariance matrix estimator, a feature that presents substantial challenges when attempting to control size. Specifically, the randomness in the residuals induced into this $r \times r$-matrix persists in large samples and is therefore a threat to valid inference. Our conjecture is that some regularization of the covariance matrix may be helpful in mitigating the noise arising from this estimated covariance matrix. In addition, it is not clear if the weighting behind the heteroskedasticity-robust Wald statistic and hence a potential test will preserve optimality under asymptotics where the number of restrictions is proportional to the sample size. We leave investigation of these difficult but interesting questions to future research.} The proposed critical value yields} asymptotic validity under two asymptotic frameworks, one where the number of restrictions is fixed, and one where the number of restrictions may grow as fast as proportionally to the sample size. To achieve such uniformity with respect to the number of restrictions, we rely on an auxiliary distribution, the F-bar distribution, that helps unite these two frameworks.

F-bar distribution

Our test rejects $H_0$ if Fisher's F exceeds a linearly transformed quantile of a distribution, which we call the F-bar distribution. We define this family of distributions and discuss its role before we turn to a description of the linear transformation mentioned above.

definition[F-bar distribution] Let $\boldsymbol w=(w_1,\dots,w_r)$ be a collection of non-negative weights summing to one, and $df$ be a positive real number. The F-bar distribution with weights $\boldsymbol w$ and degrees of freedom $df$, denoted by $\bar F_{{\boldsymbol w},df}$, is a distribution of \begin{align} \frac{\sum_{\ell=1}^r w_\ell Z_\ell}{Z_{0}/df }, \end{align} where $Z_0,Z_1,\dots,Z_r$ are mutually independent random variables with $Z_0 \sim \chi^2_{df}$ and $Z_\ell \sim \chi^2_{1}$ for $ 1\le \ell \le r$. Here, $\chi^2_\kappa$ denotes a chi-squared distribution with $\kappa>0$ degrees of freedom.

The name attached to this family originates from its close relationship to both the chi-bar-squared distribution and to Snedecor's F distribution, which we denote as $\bar \chi^2_{\boldsymbol w}$ and $F_{r,df}$, respectively. {In the Appendix, we show the following three essential properties of this family to be used later. First, Snedecor's F is a special case when the entries of $\boldsymbol w$ are all equal. Second, the limiting case of $\bar F_{{\boldsymbol w},df}$ when $df \rightarrow \infty$ is $\bar \chi^2_{\boldsymbol w}$. Third, the standard normal distribution, whose CDF is denoted as $\Phi$,} is also a limiting case since, as $df \rightarrow \infty$ and $\max_{1\le\ell \le r} w_\ell \rightarrow 0$,

align[align omitted — 143 chars of source]

for $\tau \in (0,1),$ where $q_\tau(G)$ denotes the $\tau$-th quantile of the distribution $G$. The centering and rescaling in (ref) are done according to the limiting mean and variance of the underlying random variable from Definition (ref) following the $\bar F_{{\boldsymbol w},df}$ distribution, while asymptotic normality results from mixing over infinitely many independent chi-squared variables.

Our reliance on the F-bar distribution is tied to its three properties described in the previous paragraph and three closely related observations about the F statistic. These observations are: (i) the F statistic is distributed as $F_{r,n-m}$ if the errors are homoskedastic normal, (ii) the F statistic converges in distribution (after rescaling) to a chi-bar-squared if the number of restrictions $r$ is fixed, and (iii) the F statistic converges in distribution (after centering and rescaling) to a standard normal as $r$ grows. Therefore, the class of F-bar distributions serves as a roof designed both to match the finite sample distribution of the F statistic in an important special case and to approximate each of the possible limiting distributions after a suitable linear transformation.

Critical value

The proposed critical value {for the F statistic $F$} at a nominal size $\alpha \in (0,1)$ is a linear transformation of $q_{1-\alpha}\big( \bar F_{\hat {\boldsymbol w}, n-m} \big)$ given by

align[align omitted — 243 chars of source]

{The hat over $c _\alpha$ emphasizes that this value is data-dependent. The quantities $\hat{E}_{\cal F}$ and $\hat{V}_{\cal F}$ are related to $\cal F$ in (ref), the numerator of the F statistic. From this point forward, all means and variances are conditional on the regressors $\{\boldsymbol x_i\}_{i=1}^n$, and those with a subscript $0$ are calculated under $H_0$. The quantity} $\hat{E}_{\cal F}$ is an unbiased estimator of the conditional mean $\mathbb{E}_0[{\cal F}]$, while $\hat{V}_{\cal F}$ is either an unbiased or positively biased estimator of the conditional variance $\mathbb{V}_0[{\cal F}-\hat{E}_{\cal F}]$ as explained further below. The estimated weights $\hat{\boldsymbol w} = (\hat w_1,\dots,\hat w_r)$ are constructed to be consistent for weights $\boldsymbol w_{\cal F}$ in those cases where ${\cal F}/\mathbb{E}_0[{\cal F}]$ converges in distribution to $\bar \chi^2_{\boldsymbol w_{\cal F}}$.

The critical value $\hat c_\alpha$ ensures asymptotic size control irrespective of whether $r$ is viewed as fixed or growing with the sample size $n$. To explain why $\hat c_\alpha$ provides such uniformity, we consider first the case where $r$ grows. In this case, it is illuminating to rewrite the rejection rule as an equivalent event

align[align omitted — 175 chars of source]

Since $\hat V_{\cal F}^{-1/2}({\cal F} - \hat{E}_{\cal F})$ is asymptotically normal under the null, the validity in large samples follows from the relationship between the F-bar and standard normal distributions given in (ref).\looseness=-1

When instead $r$ is viewed as asymptotically fixed, it is more informative to express the rejection region through the inequality

align[align omitted — 297 chars of source]

Note that rejecting when ${\cal F }/{\hat{E}_{\cal F}}$ exceeds the quantile $q_{1-\alpha}(\bar F_{\hat {\boldsymbol w},n-m})$ suffices for validity; for the case of a single restriction such an approach corresponds to the standard practice of comparing squares of a heteroskedasticity robust t statistic and the $(1-\alpha)$-th quantile of Student's t distribution with $n-m$ degrees of freedom.\footnote{When testing a single restriction, $\hat {\boldsymbol w}$ must equal unity so that $\bar F_{\hat {\boldsymbol w},n-m}= F_{1,n-m} = t_{n-m}^2$, and in this case ${\cal F }/{\hat{E}_{\cal F}}$ is the square of the t statistic studied in kline2018leave.} The last term on the right hand side of (ref) can then be viewed as a finite sample correction that adjusts the critical value up or down depending on the relative size of the variance estimator for the ratio ${\cal F }/{\hat{E}_{\cal F}}$, which is $\hat V_{\cal F}/\hat{E}_{\cal F}^2$, and the variance of the approximating distribution $\bar F_{\hat {\boldsymbol w},n-m}$, which is roughly $2\sum_{\ell=1}^r \hat w_\ell^2 +2/(n-m)$. As the ratio of these variances converges to unity when the number of restrictions is fixed, this term does not affect first order asymptotic validity.

Finally, note that if one is willing to rest on the assumption that the restrictions are numerous and the few restriction framework is superfluous, one might use the following simplified critical value not robust to few restrictions:\footnote{Such settings occur, for example, if the null of interest involves thousands of restrictions, in which case the two critical values $\hat c_\alpha$ and $\check c_\alpha$ are essentially equivalent but $\check c_\alpha$ is computationally simpler to construct as it circumvents computation of $\hat{\boldsymbol w}$.}

align[align omitted — 202 chars of source]

To complete the description of the proposed critical value, definitions of the quantities $\hat{E}_{\cal F}$, $\hat{V}_{\cal F}$ and $\hat{\boldsymbol w}$ are needed. Section (ref) describes how we rely on leave-one-out OLS estimators to construct $\hat{E}_{\cal F}$ and $\hat{\boldsymbol w}$. For $\hat{V}_{\cal F}$, Section (ref) provides the corresponding definition when it is possible to rely on leave-three-out OLS estimators, while Section (ref) introduces the form of $\hat{V}_{\cal F}$ for settings where some of the leave-three-out estimators cease to exist. In the former case, it is possible to ensure that $\hat{V}_{\cal F}$ is unbiased, while the latter introduces a (small) positive bias. We initially consider the former case, a framework where the design matrix has full rank when any three observations are left out of the sample, and relax this condition in Section (ref).

assumption$\sum_{\ell \neq i,j,k} \boldsymbol x_\ell \boldsymbol x_\ell'$ is invertible for every $i,j,k \in \{1,\dots,n\}$.

When $\boldsymbol x_i$ is identically and continuously distributed with unconditional second moment $\mathbb{E}[\boldsymbol x_i \boldsymbol x_i']$ of full rank, Assumption (ref) holds with probability one whenever $n - m \ge 3$. The asymptotic framework considers a setting where $n-m$ diverges so that Assumption (ref) must hold in sufficiently large samples with continuous regressors. This conclusion also applies when $\boldsymbol x_i$ includes a few discrete regressors and, in particular, an intercept. In settings with many discrete regressors, Assumption (ref) may fail to hold, even in large samples. For that reason, Section (ref) introduces the version of $\hat{V}_{\cal F}$ for empirical settings where the full rank condition is satisfied when any one observation is left out, but not necessarily when leaving two or three observations out.

Leave-out algebra

Before describing $\hat{E}_{\cal F}$, $\hat{V}_{\cal F}$, and $\hat{\boldsymbol w}$ in detail, we will reformulate Assumption (ref) using leave-out algebra. That is, we will derive an equivalent way of expressing this assumption while introducing notation that is essential for the construction of the critical value and for stating the asymptotic regularity conditions.

When $\boldsymbol S_{xx}$ has full rank, a direct implication of the Sherman-Morrison-Woodbury identity sherman1950adjustment,woodbury1949stability is that the leave-one-out design matrix $\sum_{j \neq i}\boldsymbol x_j \boldsymbol x_j'$ is invertible if and only if the statistical leverage of the $i$-th observation $P_{ii} = \boldsymbol x_i' \boldsymbol S_{xx}^{-1} \boldsymbol x_i$ is less than one. Letting $M_{ij} = \mathbf{1}\{i=j\} - \boldsymbol x_i'\boldsymbol S_{xx}^{-1} \boldsymbol x_j$ be elements of the residual projection matrix $\boldsymbol M$ associated with the regressor matrix, this condition on the leverage is equivalently stated as $M_{ii}$ being greater than zero. When $M_{ii}>0$ holds, we can additionally use SMW to represent the inverse of $\sum_{j \neq i} \boldsymbol x_j \boldsymbol x_j'$ as

align[align omitted — 228 chars of source]

which highlights the role of a non-zero $M_{ii}$.

The representation in (ref) can also be used to understand when the leave-two-out design matrix $\sum_{k \neq i,j}\boldsymbol x_k \boldsymbol x_k'$ has full rank, since (ref) can be used to compute leverages in a sample that excludes $i$. After leaving observation $i$ out, the leverage of a different observation $j$ is $\boldsymbol x_j'\big(\sum_{k \neq i} \boldsymbol x_k \boldsymbol x_k'\big)^{-1} \boldsymbol x_j$. To see when this leverage is less than one, note that (ref) yields

align[align omitted — 162 chars of source]

so that a necessary and sufficient condition for a full rank of $\sum_{k \neq i,j} \boldsymbol x_k \boldsymbol x_k'$ is that $D_{ij} >0$, where

align[align omitted — 127 chars of source]

and $\abs{\cdot}$ denotes the determinant.

Extending the previous argument to the case of leaving three observations out, we find that the invertibility of $\sum_{\ell \neq i,j,k}\boldsymbol x_\ell \boldsymbol x_\ell'$ for $i$, $j$, and $k$, all of which are different, is equivalent to $D_{ijk}>0$, where

align[align omitted — 239 chars of source]

This discussion reveals that Assumption (ref) can equivalently be stated as requiring full rank of $\boldsymbol S_{xx}$ and

align[align omitted — 120 chars of source]

In addition to facilitating an algebraic description of Assumption (ref), the quantities $M_{ii}, \ D_{ij},$ and $D_{ijk}$ also play a role in the computation of the proposed critical value. Specifically, they can be used to avoid explicitly computing the OLS estimates after leaving one, two, or three observations out. Additionally, since construction of $\hat{E}_{\cal F}$, $\hat{V}_{\cal F}$, and $\hat{\boldsymbol w}$ relies on dividing by $M_{ii}, \ D_{ij},$ and $D_{ijk}$, the study of the asymptotic size of the proposed testing procedure imposes a slight strengthening of (ref), which bounds the smallest $D_{ijk}$ away from zero.

Location estimator

{ Recall that we recenter the numerator of the F statistic ($\cal F$ in (ref)) by using $\hat{E}_{\cal F}$, which is an unbiased estimator of the conditional mean $\mathbb{E}_0[{\cal F}]$ under $H_0$. This mean equals}

align[align omitted — 92 chars of source]

where the values $B_{ij} = \boldsymbol x_i'\boldsymbol S_{xx}^{-1} \boldsymbol R'\!\left( \boldsymbol R\boldsymbol S_{xx}^{-1} \boldsymbol R' \right)^{-1} \! \boldsymbol R \boldsymbol S_{xx}^{-1} \boldsymbol x_j$ are observed and satisfy $\sum_{i=1}^{n}B _{ii} =r$. Furthermore, the exact null distribution of ${\cal F}/\mathbb{E}_0[{\cal F} ]$, under the additional condition of normally distributed regression errors, is $\bar \chi^2_{\boldsymbol w_{\cal F}}$, with $\boldsymbol w_{\cal F}$ containing the eigenvalues of the matrix\footnote{ { Under error normality, the exact null distribution of $\boldsymbol R\hat {\boldsymbol \beta} - \boldsymbol q $ is ${\cal N}\big( 0, \mathbb{V}\big[ \boldsymbol R \hat{\boldsymbol \beta} \big] \big)$, and it follows from Lemma 3.2 of vuong1989 that ${\cal F}/\mathbb{E}_0[{\cal F} ]$ is distributed as a weighted sum of chi-squares with weights that are eigenvalues of $( \boldsymbol R\boldsymbol S_{xx}^{-1} \boldsymbol R' )^{-1}\mathbb{V}\big[ \boldsymbol R \hat{\boldsymbol \beta} \big] /\mathbb{E}_0[{\cal F} ]=\varOmega \big(\sigma_1^2,\dots,\sigma_n^2 \big).$ } Note that the eigenvalues of $\varOmega \big(\sigma_1^2,\dots,\sigma_n^2 \big)$ are all real and non-negative as they can be expressed as the eigenvalues of the symmetric and positive semidefinite matrix $( \boldsymbol R\boldsymbol S_{xx}^{-1} \boldsymbol R' )^{-1/2} \mathbb{V}\big[ \boldsymbol R \hat{\boldsymbol \beta} \big] ( \boldsymbol R\boldsymbol S_{xx}^{-1} \boldsymbol R' )^{-1/2} /\mathbb{E}_0[{\cal F} ]$. Furthermore, the entries of $\boldsymbol w_{\cal F}$ sum to one as $\mathbb{E}_0[{\cal F} ]$ is the trace of $( \boldsymbol R\boldsymbol S_{xx}^{-1} \boldsymbol R' )^{-1} \mathbb{V}\big[ \boldsymbol R \hat{\boldsymbol \beta} \big]$. }

align[align omitted — 335 chars of source]

Both $\mathbb{E}_0[{\cal F} ]$ and $\boldsymbol w_{\cal F}$ are thus functions of $\{\sigma_i^2\}_{i=1}^n$, and the relevance of the vector $\boldsymbol w_{\cal F}$ for asymptotic size control transcends the normality assumption on the errors that we used in order to introduce it.

As shown in kline2018leave, the individual specific error variances can be estimated without bias for any value of $\boldsymbol\beta$ using leave-one-out estimators. Let the leave-$i$-out OLS estimator of $\boldsymbol\beta$ be $\hat {\boldsymbol\beta}_{-i} = \big(\sum_{j \neq i}\boldsymbol x_j\boldsymbol x_{j}'\big)^{-1} \sum_{j \neq i} \boldsymbol x_j y_j,$ and construct

align[align omitted — 111 chars of source]

With these leave-one-out estimators, we can estimate the null mean of ${\cal F}$ using

align[align omitted — 76 chars of source]

which ensures that the first moment of ${\cal F} - \hat{E}_{\cal F} $ is zero under the null. Since $\hat \sigma_i^2$ is unbiased for any value of $\boldsymbol\beta$, this centered statistic still has its expectation minimized under $H_0$, so that large values of the statistic can be taken as evidence against the null. Following the same approach, we can estimate $\boldsymbol w_{\cal F}$ using the sample analog $\check {\boldsymbol w} = (\check w_1,\dots,{ \check w_r})'$, where $\check w_\ell$ is the $\ell$-th eigenvalue of $\varOmega \!\left(\hat \sigma_1^2,\dots,\hat \sigma_n^2 \right)$. However, $\check {\boldsymbol w}$ may not have non-negative entries summing to one, so we ensure that these conditions hold by letting $\hat {\boldsymbol w} = (\hat w_1,\dots,{ \hat w_r})'$, where

align[align omitted — 97 chars of source]

While our construction of $\hat{E}_{\cal F}$ implies that the first moment of ${\cal F} - \hat{E}_{\cal F} $ is known when $H_0$ holds, its second moment still depends heavily on unknown parameters. Under $H_0$,

align[align omitted — 262 chars of source]

where $U_{ij} = 2\!\left( B_{ij} - {M_{ij}} \!\left({B_{ii}}/{M_{ii}}+{B_{jj}}/{M_{jj}}\right)\!/2 \right)^2$ and $V_{ij}=M_{ij} \!\left( {B_{ii}}/{M_{ii}}-{B_{jj}}/{M_{jj}}\right)$ are known quantities. This representation of the null-variance stems from writing ${\cal F} - \hat{E}_{\cal F}$ as a second order $U$-statistic with squared kernel weights of $U_{ij}/2$ plus a linear term with weights $\sum\nolimits_{j\neq i} V_{ij}\boldsymbol x_j'\boldsymbol\beta$ (see the Appendix for details).

Variance estimator

This subsection describes the construction of an unbiased estimator of the conditional variance ${\mathbb{V}}_{0}\big[{\cal F} - \hat{E}_{\cal F} \big]$. As is evident from the representation in (ref), this variance depends on products of second moments such as the product $\sigma_i^2 \sigma_j^2$. While $\hat \sigma_i^2$ and $\hat \sigma_j^2$ are unbiased for $\sigma_i^2$ and $\sigma_j^2$, their product is not unbiased, as the estimation error is correlated across the two estimators. Some of this dependence can be removed by leaving both $i$ and $j$ out, but a bias remains as the remaining sample is used in estimating both $\sigma_i^2$ and $\sigma_j^2$. We therefore propose a leave-three-out estimator of the variance product $\sigma_i^2 \sigma_j^2$. The product $\boldsymbol x_j'{\boldsymbol\beta}\,\boldsymbol x_k'{\boldsymbol\beta}\, \sigma_i^2$ appearing in the second component of ${\mathbb{V}}_{0}\big[{\cal F} - \hat{E}_{\cal F} \big]$ can similarly be estimated without bias using leave-three-out estimators.

Towards this end, let $\hat {\boldsymbol\beta}_{-ijk} = \big(\sum_{\ell \neq i,j,k}\boldsymbol x_\ell\boldsymbol x_{\ell}'\big)^{-1} \sum_{\ell \neq i,j,k}\boldsymbol x_\ell y_{\ell}$ denote the OLS estimator of $\boldsymbol\beta$ applied to the sample that leaves observations $i$, $j$, and $k$ out. Then, define a leave-three-out estimator of $\sigma_i^2$ as

align[align omitted — 104 chars of source]

When $j$ and $k$ are identical, only two observations are left out, and we also write $\hat {\boldsymbol \beta}_{-ij}$ and $\hat \sigma_{i,-j}^2$. To construct an estimator of $\sigma_i^2 \sigma_j^2$, we first write the leave-two-out variance estimator $\hat \sigma_{i,-j}^2$ as a weighted sum (see Section (ref) for details)

align[align omitted — 188 chars of source]

Then we multiply each summand above by a leave-three-out variance estimator $ \hat \sigma_{j,-ik}^2$, which leads to an unbiased estimator of $\sigma_i^2 \sigma_j^2$:

align[align omitted — 136 chars of source]

While this construction appears to treat $i$ and $j$ in an asymmetric fashion, we show to the contrary that (ref) is invariant to a permutation of the indices; $\widehat{\sigma_i^2 \sigma_j^2}=\widehat{\sigma_j^2 \sigma_i^2}$.

To understand why this proposal is unbiased for $\sigma_i^2 \sigma_j^2$, it is useful to highlight that $\hat \sigma_{j,-ik}^2$ is conditionally independent of $(y_i,y_k)$ and unbiased for $\sigma_{j}^2$, which, when coupled with (ref), leads to unbiasedness immediately:

align[align omitted — 288 chars of source]

An unbiased estimator of the variance expression in (ref) that utilizes the variance product estimator in (ref) is

align[align omitted — 255 chars of source]

Note that the product of the $(j=k)$-th terms in the second component generate, for each $i$, a term not present in (ref) and whose non-zero expectation contains $V_{ij}^2\sigma_i^2 \sigma_j^2$; hence the use of $U_{ij} - V_{ij}^2$ instead of $U_{ij}$ in the first component.

remIn the process of establishing the asymptotic validity of the proposed test, we show that the variance estimator $\hat{V}_{\cal F}$ is close to the null variance $\mathbb{V}_0 \big[{\cal F} - \hat{E}_{\cal F} \big]$. In particular, this property implies that the variance estimator is positive with probability approaching one in large samples. However, negative values may still emerge in small samples. In such cases, we propose to replace the variance estimator with an upward biased alternative that uses squared outcomes as estimators of all the error variances. This replacement is guaranteed positive, as is detailed in the Appendix, and therefore ensures that the critical value is always defined. Relatedly, Section (ref) considers settings where the design matrix may turn rank deficient after leaving certain triples of observations out of the sample. There, we similarly propose to use squared outcomes as estimators of some error variances, namely those whose observations cause rank deficiency when left out of the sample.
remNote that in finite samples, the proposed critical value $\hat c_\alpha$ is not invariant to the value of $\boldsymbol \beta$. In practice, this means that finite sample size and power may be influenced by the size of the regression coefficients. In Section (ref) we analyze via simulations the impact of the average signal size on the finite sample size and power through the regression $R^2$.

{

remThe null restrictions $\boldsymbol{R\beta }=\mathbf{q}$ can be imposed during the estimation of auxiliary quantities (i.e., $\sigma _{i}^{2},\ \sigma _{i}^{2}\sigma _{j}^{2},\dots$), and the restricted estimates may be used in place of our proposals that do not impose those restrictions (i.e., $\widehat{\sigma }_{i}^{2},\ \widehat{\sigma _{i}^{2}\sigma _{j}^{2}}, \dots)$.\footnote{{We thank an anonymous referee for raising this point.}} In the Appendix, we show how this idea can be implemented, and leave further investigation to future research. Incorporating restricted estimates may make the the procedure more complex and slow down computations, but it may also improve the efficiency of the auxiliary estimates.

}

Computational remarks

While the previous subsections introduced the location estimator $\hat{E}_{\cal F}$, variance estimator $\hat{V}_{\cal F}$, and empirical weights $\hat{\boldsymbol w}$ using leave-out estimators of $\boldsymbol\beta$, we note here that direct computation of $\hat {\boldsymbol\beta}_{-i}$, $\hat {\boldsymbol\beta}_{-ij}$, and $\hat {\boldsymbol\beta}_{-ijk}$ can be avoided by using the Sherman-Morrison-Woodbury (SMW) identity. Specifically, (ref) implies that

align[align omitted — 127 chars of source]

so that computation of $\hat {\boldsymbol\beta}_{-i}$ can be avoided when constructing the leave-one-out variance estimator $\hat \sigma _{i}^{2} = {y_i(y_i -\boldsymbol x_i'\hat {\boldsymbol\beta}_{-i})}$. Similarly, it is possible to show that for $i$ and $j$ not equal,

align[align omitted — 189 chars of source]

which leads to (ref), and, for $i$, $j$, and $k$, all of which are different,

align[align omitted — 257 chars of source]

These relationships allow for recursive computation of the leave-out residuals and therefore for simple construction of the variance estimators $\hat \sigma _{i}^{2}$, $\hat \sigma _{i,-j}^{2}$, and $\hat \sigma _{i,-jk}^{2}$ needed to compute the components of the critical value $c_{\alpha}$. In particular, the location estimator $\hat{E}_{\cal F}$ and empirical weights $\hat{\boldsymbol w}$, which require only the leave-one-out residuals, can be computed without explicit loops, by relying instead on elementary matrix operations applied to the matrices containing $M_{ij}$ and $B_{ij}$ as well as the data matrices. Similarly, all doubly indexed objects entering the variance estimator $\hat{V}_{\cal F}$ can be computed by elementary matrix operations. Those objects are $D_{ij}$, $V_{ij}$, $U_{ij}$, and the leave-two-out residuals. The remaining objects entering $\hat{V}_{\cal F} $ can be computed by a single loop across $i$ with matrices containing $D_{ijk}$ and leave-three-out residuals renewed at each iteration.

Additionally, the above representations of leave-out residuals demonstrate how $M_{ii}^{-1}$, $D_{ij}^{-1}$ and $D_{ijk}^{-1}$ enter the critical value, and thus highlight the need for bounding $D_{ijk}$ away from zero when analyzing the large sample properties of the proposed test.

remThe quantile $q_{1-\alpha}(\bar F_{\hat{\boldsymbol{w}},n-m})$ can easily be constructed by simulating the distribution of the random variable in (ref) conditional on the realized value of $\hat{\boldsymbol{w}}$.

Asymptotic size and power

This section studies the asymptotic properties of the proposed test. Specifically, we provide a set of regularity conditions under which the test has a correct asymptotic size and non-trivial power against local alternatives. All limits are taken as the sample size $n$ approaches infinity. In studying asymptotic size, we allow for the number of restrictions $r$ and/or number of regressors $m$ to be fixed or diverging with $n$, and show that the asymptotic size is controlled uniformly over the two situations by the test, which is therefore robust to the type of asymptotics. When studying asymptotic power, we focus on the case of many restrictions, i.e., $r$ diverging with $n$. The ordering $r \le m < n-3$ is maintained throughout.

{

Assumptions

To establish asymptotic validity of the proposed test, we impose some regularity conditions. We begin by outlining the assumptions for the sampling scheme.

assumption$\{(y_i,\boldsymbol x_i')'\}_{i=1}^n$ are i.i.d., $\mathbb{E}[\varepsilon_i | \boldsymbol x_i] =0$, $\max_{1\le i\le n} \big(\mathbb{E}[\varepsilon_i^4 |\boldsymbol x_i] + \sigma_i^{-2}\big) = O_p(1)$.

Assumption (ref)} places restrictions on the error conditional moments: an upper bound on the conditional fourth moments and lower bound on the skedastic function. Such restrictions are typically required when heteroskedasticity is allowed chao2012asymptotic,cattaneo2017inference,kline2018leave.

{Next, we impose regularity conditions on the regressors to ensure convergence of the centered statistic ${\cal F} - \hat{E}_{\cal F}$ to the normal distribution when the number of restrictions grow large. These conditions restrict the weights on the regression errors in the bilinear form of ${\cal F} - \hat{E}_{\cal F}$ \footnote{{See, e.g., expansion (ref) below and its extended version (ref) in the Appendix.}}. These weights contain various functions of regressors (in particular, potentially unbounded regression function values $\boldsymbol x_i'{\boldsymbol\beta}$), and the purpose of Assumption (ref) is to restrict their asymptotic behavior.

assumptionThere exists a sequence $\epsilon_n \rightarrow 0$ such that $(i)$ $\epsilon_n^{1/3} \max_{1\le i\le n} (\boldsymbol x_i'{\boldsymbol\beta})^2 = O_p( 1 )$ and (ii) at least one of the following two conditions is satisfied: \begin{enumerate} • $\max_{1\le i\le n} B_{ii} = o_p( \epsilon_n )$, • $\max_{1\le i\le n}\,(\sum_{j \neq i} V_{ij}\boldsymbol x_j'{\boldsymbol\beta} )^2/r = o_p( 1 )$ and $\epsilon_n r \rightarrow \infty$. \end{enumerate}

}

{Part $(i)$ of Assumption (ref),} which places bounds on regression function values, is used to control the variance of the leave-out estimators $\hat \sigma_i^2$ and $\hat \sigma_{i,-jk}^2$. This condition places a rate bound on extreme outliers among the individual signals, and its role is primarily to control certain higher order terms. The condition $(i)$ is used to establish both size control and local power properties, so we stress that it pertains to the actual data generating process, not just the hypothesized value of $\boldsymbol\beta$. Also note that {Assumption (ref)$(i)$} differs from Assumption 1(iii) of kline2018leave in that we allow $\boldsymbol x_i'\boldsymbol\beta$ to have an unbounded support so the maximum over $i$ may be slowly diverging with $n$. This relaxation is important, as it allows regressors with unbounded support and associated non-zero regression coefficients.

{Part $(ii)$ of Assumption (ref) is an analogue of the Lindeberg condition in the central limit theorem for weighted sums of independent random variables, in that it also controls the collective asymptotic behavior of weights, but the weights in a bilinear form of independent regression errors.} When the number of tested restrictions is fixed, {this condition} implies that the estimator of the tested contrasts $\boldsymbol R \hat{\boldsymbol\beta} - \boldsymbol q$ is asymptotically normal. When the number of restrictions is growing, {this condition} is weaker and involves only a high-level transformation of the regressors $\sum_{j \neq i} V_{ij}\boldsymbol x_j'\boldsymbol\beta$, which enters ${\cal F} - \hat E_{\cal F}$ as a weight on the $i$-th error term $\varepsilon_i$. To ensure that the asymptotic distribution of ${\cal F} - \hat E_{\cal F}$ does not depend on the unknown distribution of any one error term, we therefore require that no squared $\sum_{j \neq i} V_{ij}\boldsymbol x_j'\boldsymbol\beta$ dominates the variance ${\mathbb{V}}_{0}\big[{\cal F} - \hat{E}_{\cal F} \big]$, which in turn is proportional to $r$. {Assumption (ref)$(ii)$} can be verified in particular applications of interest. For example, the Appendix shows that {part $(ii)(b)$ of Assumption (ref)} holds in models characterized by group specific regressors.

The next assumption imposes the previously discussed regularity condition that the determinant $D_{ijk}$ is bounded away from zero for any $i$, $j$, and $k$, all of which are different. This condition will be relaxed in Section (ref), where such a version of $\hat V_{\cal F}$ is introduced that exists even when leaving two or three observations out leads to rank deficiency of the design.

assumption$\max_{i \neq j \neq k \neq i} D_{ijk}^{-1} = O_p(1)$.

Asymptotic size

Under the regularity conditions in Assumptions (ref) and (ref), the following theorem establishes the asymptotic validity of the proposed testing procedure.

theoremIf Assumptions (ref), (ref), (ref), and (ref) hold, then, under $H_0$, \begin{align} \lim_{n \rightarrow \infty} \mathbb{P}\left( F> \hat c_\alpha \right) = \alpha. \end{align}

{A discussion of the structure of the decision event $F > \hat c_\alpha$ may aid in understanding why size control occurs in the critical case of many restrictions, $r\to\infty$, and where the challenges come from. The critical value is then asymptotically close to $\check c _\alpha$ defined in (ref), and the decision event is equivalent to $\hat V_{\cal F}^{-1/2}({\cal F} - \hat{E}_{\cal F}) > \big(q_{1-\alpha}(F_{r,n-m})-1\big)/\sqrt{2/r+2/(n-m)}$. The right side is asymptotically standard normal, so the size control rests on asymptotic standard normality of $\hat V_{\cal F}^{-1/2}({\cal F} - \hat{E}_{\cal F})$, which, in turn, is shown using a central limit theorem for ${\cal F} - \hat{E}_{\cal F}$ and its asymptotically correct standardization by $\hat V_{\cal F}^{1/2}$.

The demeaned statistic ${\cal F} - \hat{E}_{\cal F}$ has a representation of} a bilinear form in independent, not necessarily identically distributed, random variables (see equation (3) in the Appendix):

align[align omitted — 153 chars of source]

where, in our case, $u_{j}=v_{j}=w_{j}=\varepsilon _{j}$ are regression errors, and the coefficients $c_{\boldsymbol \beta j}$ in the second, linear component, explicitly depend on the unknown parameters $\boldsymbol \beta $. {Note that the quadratic term in (ref) has a jackknife structure and lacks terms of the type $C_{ii}\varepsilon _{i}^2$, whose presence would introduce higher-order moments into the variance.\footnote{{See also calhoun2011many for an example of a structure where higher-order moments do arise and need to be tediously estimated.}} This further highlights the importance of demeaning the original statistic ${\mathcal{F}}$. }

The econometric theory literature is populated by central limit theorems (CLTs) handling the asymptotics of (ref) under various assumptions, starting from an early CLT in kelejian2001spatial that was originated in a spatial regression environment. More recent examples are the CLT of chao2012asymptotic, formulated in the many-weak-instrument context, the CLT of kline2018leave designed for many-regressor models, and the CLT of anatmik2021factor targeting big factor models. cattaneo2016alternative presented a unifying framework leading to the description of asymptotical behavior of V-statistics that further generalize bilinear forms like (ref). {Finally, kuersteiner2020 provide a CLT for expressions like (ref), allowing for data-dependent $C_{ij}$ and $c_{\boldsymbol \beta j}$, heteroskedasticity, and some forms of dependence across $j$.} In all these setups, the asymptotic normality of (ref) eventually results in an asymptotically normal test statistic when (ref) is suitably standardized.

The next challenge is converting the asymptotic normality of (ref) into an asymptotically valid test.\footnote{Often, the bilinear form (ref) has a tighter structure, which simplifies the emergence of asymptotic normality and further pivotization, removing the challenges handled in the present paper. For example, simplification may come from a slow growth of incidental parameters' dimensionality hong1995nparmtest, absence of the quadratic component in (ref) breitung2016hetpanel, or assumption of conditional homoskedasticity anatolyev2012inference. Likewise, in the many weak instrument literature, a variety of tests are also based on asymptotic normality of bilinear forms, with simplifying deviations from the general form (ref). In J type tests for validity of many instruments and thus many restrictions anatgosp2011test,leeokui2012hausmantest,chao2014testing, the dependence on a parameter of asymptotically fixed dimensionality can be handled using its plug-in estimate. In Anderson-Rubin type tests for few parameter restrictions anatgosp2011test,crudu2021ar,miksun2021manyweak, the number of parameter restrictions is asymptotically fixed, and one can use restricted null values of the parameters. In our situation, in contrast, both the restriction numerosity and parameter dimensionality are asymptotically increasing, at least when $r$ is asymptotically growing.} To obtain an asymptotically pivotal statistic, (ref) needs to be standardized by a consistent estimate of its (rescaled) variance (ref). Constructing such an estimate is difficult because of explicit dependence of the coefficients $c_{\boldsymbol \beta j}$ and implicit dependence of the regression errors $\varepsilon _{i}$ on $\boldsymbol \beta ,$ a high-dimensional parameter when regressors are many. {The estimator (ref) based on leave-three-out error variance estimates accomplishes this goal.}

As our treatment covers two asymptotically different frameworks under one roof, we use two CLTs in this paper. One, formulated as Lemma B.1 in kline2018leave, which builds on Lemmas A2.1 and A2.2 in solv2020robmany, pertains to the case of asymptotically growing $r$. The other is the regular Lyapounov central limit theorem, which pertains to the case of asymptotically fixed $r$.

Asymptotic power

To describe the power of the proposed test, we introduce a drifting sequence of local alternatives indexed by a deviation $\boldsymbol\delta$ from the null times $(\boldsymbol R\boldsymbol S_{xx}^{-1} \boldsymbol R')^{1/2}$, which specifies the precision the tested linear restrictions can be estimated with in the given sample. Thus, we consider alternatives of the form

align[align omitted — 178 chars of source]

for ${\boldsymbol\delta} \in \mathbb{R}^r$ satisfying the limiting condition

align[align omitted — 118 chars of source]

Below we show that the power of the test is monotone in $\Delta_\delta$, with power equal to size when $\Delta_\delta=0$ and power equal to one when $\Delta_\delta=\infty$.

The role of $(\boldsymbol R\boldsymbol S_{xx}^{-1}\boldsymbol R')^{1/2}$ in indexing the local alternatives is analogous to that of $n^{-1/2}$ often used in parametric problems. However, in settings with many regressors some linear restrictions may be estimated at rates that are substantially lower than the standard parametric one. Therefore, we index the deviations from the null by the actual rate of $(\boldsymbol R\boldsymbol S_{xx}^{-1}\boldsymbol R')^{1/2}$ instead of $n^{-1/2}$.

The alternative is additionally indexed by $\boldsymbol\delta$, which in standard parametric problems is typically fixed. However, fixed $\boldsymbol \delta$ is less natural here, as the dimension of $\boldsymbol\delta$ increases with sample size. Instead, we fix the limit of its Euclidean norm when scaled by $r^{1/4}$. This approach allows us to discuss different types of alternatives and how the numerosity of the tested restrictions affects the test's ability to detect deviations from the null. Specifically, note that when the deviation $\boldsymbol\delta$ is sparse, i.e., only a bounded number of its entries are non-zero, then the test has a non-trivial power against alternatives, {whose individual elements on average diverges} at a rate that is $r^{1/4}$ lower than when only a fixed number of restrictions is tested. This observation highlights the cost for the power of including many irrelevant restrictions in the hypothesis. On the other hand, if $\boldsymbol\delta$ is dense, e.g., with all entries bounded away from zero, then the test can detect local deviations, {in which an individual element on average shrinks} at a rate that is $r^{1/4}$ greater than the usual. This means that if the tested restrictions can be estimated at the parametric rate and they are all relevant, then the test can detect deviations from the null of order ${n^{-1/2} r^{-1/4}}$.

The following theorem states the asymptotic power under sequences of local alternatives of the form given in (ref) and discussed above.

theoremIf Assumptions (ref), (ref), (ref), and (ref) hold, then, under $H_\delta$, \begin{align} \lim_{n,r \rightarrow \infty}\mathbb{P}\left( F> \hat c_\alpha \right) - \Phi \left( \Phi ^{-1}\left( \alpha \right) + \Delta_{\delta}^2 \left( {\mathbb{V}_{0}\!\big[ \mathcal{F}-\hat{E}_{\cal F}\big]}/r\right)^{-1/2} \right) = 0 , \end{align} where $\Phi $ denotes the cumulative distribution function of the standard normal and $\Phi(\infty)=1$.
remIt is instructive to compare the power curve documented in Theorem (ref) with the asymptotic power curve of the exact F test when both tests are valid. When the individual error terms are homoskedastic normal with variance $\sigma^2$, the asymptotic power of the exact F test is the limit of anatolyev2012inference \begin{align} \Phi \left( \Phi ^{-1}\left( \alpha \right) + \Delta_{\delta}^2 \left(2\sigma^4 + 2\sigma^4r/(n-m)\right)^{-1/2} \right). \end{align} Thus, the relative asymptotic power of the proposed LO test and the exact F test is determined by the limiting ratio of $r^{-1}\mathbb{V}_{0}\big[ \mathcal{F}-\hat{E}_{\cal F}\big]$ to $2\sigma^4\!\left(1 + r/{(n-m)}\right)$. The Appendix shows that this ratio approaches one in large samples if the number of tested restriction is small relative to the sample size or if the limiting variability of $B_{ii}/M_{ii}$ is small kline2018leave. When neither of these conditions holds, the proposed test will, in general, have a slightly lower power than the exact F test, which we also document in the simulations in Section (ref).
remThe order of the numerosity of alternatives that can be detected with the proposed test is optimal in the minimax sense when the alternatives are moderately sparse to dense, i.e., when $O(\sqrt{r})$ or more of the tested restrictions are violated arias2011global. However, if the alternative is strongly sparse so that at most $o(\sqrt{r})$ tested restrictions are violated, a higher power can be achieved by tests that redirect their power towards those alternatives. Such tests typically focus their attention on a few largest t statistics (i.e., smallest p values) and are often described as multiple comparison procedures donoho2004higher,romano2010multiple. While such tests can control size when the error terms are homoskedastic normal, it is not clear whether they can do so in the current semiparametric framework with an unspecified error distribution. The issue is that the size control for multiple comparisons relies on knowing the (normal or t) distributions of individual t statistics, but in the current framework with many regressors those distributions are not necessarily known (even asymptotically).

If leave-three-out fails

This section extends the definition of the critical value $c_{\alpha}$ to settings where the design matrix may turn rank deficient after leaving certain pairs or triples of observations out of the sample. When Assumption (ref) fails in this way, $\hat E_{\cal F}$ is still an unbiased estimator of $\mathbb{E}_0[{\cal F}]$, but the unbiased variance estimator introduced in Section (ref) does not exist. For this reason, we propose an adjustment to the variance estimator that introduces a positive bias for pairs of observations where we are unable to construct an unbiased estimator of the variance product $\sigma_i^2 \sigma_j^2$ and for triples of observations where we are unable to construct an unbiased estimator of $\boldsymbol x_j'{\boldsymbol\beta} \,\boldsymbol x_k'{\boldsymbol\beta}\, \sigma_i^2$. This introduction of a positive bias to the variance estimator ensures asymptotic size control, even when Assumption (ref) fails.

Since this section considers a setup where Assumption (ref) may fail, we introduce a weaker version of the assumption, which only imposes the full rank of the design matrix after dropping any one observation.

customass{1'} $\sum_{j \neq i}\boldsymbol x_j\boldsymbol x_j'$ is invertible for every $i \in \{1,\dots,n\}$.

One can always satisfy this assumption by appropriately pruning the sample, the model, and the hypothesis of interest. For example, if $\boldsymbol S_{xx}$ does not have full rank, then one can remove unidentified parameters from both the model and hypothesis of interest, and proceed by testing the subset of restrictions in $H_0$ that are identified by the sample. Similarly, if $\sum_{j \neq i}\boldsymbol x_j\boldsymbol x_j'$ does not have full rank for some observation $i$, then there is a parameter in the model which is identified only by this observation. Therefore, one can proceed as in the case of rank deficiency of $\boldsymbol S_{xx}$, by dropping observation $i$ from the sample and by removing the parameter that determines the mean of this observation from the model and null hypothesis. When doing this for any observation $i$ such that $\sum_{j \neq i}\boldsymbol x_j\boldsymbol x_j'$ is non-invertible, one obtains a sample that satisfies Assumption (ref) and can be used to test the restrictions in $H_0$ that are identified by this leave-one-out sample.

Variance estimator

When Assumption (ref) fails, some of the unbiased estimators $\hat \sigma^2_{i,-jk}$ and $\widehat{\sigma_i^2 \sigma_j^2}$ cease to exist. For such cases, the variance estimator $\hat{V}_{\cal F}$ utilizes replacements that are either also unbiased or positively biased, depending on the cause of the failure. Assumption (ref) fails if $D_{ijk} =0$ for some triple of observations, and we say that this failure of full rank is caused by $i$ if $D_{jk}>0$ or $D_{ij}D_{ik}=0$, i.e., if the design retains full rank when only observations $j$ and $k$ are left out or if leaving out observations $(i,j)$ or $(i,k)$ leads to rank deficiency. Our replacement for $\hat \sigma^2_{i,-jk}$ is biased when $i$ causes $D_{ijk} =0$, while the replacement for $\widehat{\sigma_i^2 \sigma_j^2}$ is biased when both $i$ and $j$ cause $D_{ijk} =0$ for some $k$.

To introduce the replacement for $\hat \sigma^2_{i,-jk}$, we consider the case when it does not exist, or equivalently, when $D_{ijk}=0$. If $i$ causes this leave-three-out failure, then our replacement is the upward biased estimator $y_i^2$. When this failure of leave-three-out is not caused by $i$, the leave-two-out estimators $\hat \sigma_{i,-j}^2$ and $\hat \sigma_{i,-k}^2$ are equal and independent of both $y_j$ and $y_k$ (as shown in the Appendix). These properties imply that $y_j y_k \hat \sigma_{i,-j}^2$ is an unbiased estimator of $\boldsymbol x_j'{\boldsymbol\beta}\,\boldsymbol x_k'{\boldsymbol\beta}\, \sigma_i^2$, and we therefore use $\hat \sigma_{i,-j}^2$ as a replacement for $\hat \sigma^2_{i,-jk}$. To summarize, we let

align[align omitted — 225 chars of source]

When $j$ is equal to $k$, we consider pairs of observations, and the definition only involves the last two lines since $D_{ijj}=0$. In this case, we also write $\bar \sigma_{i,-j}^2$ for $\bar \sigma_{i,-jj}^2$.

For the replacement of $\widehat{\sigma_i^2 \sigma_j^2} = y_i \sum_{k \neq j} \check M_{ik,-ij} y_k \cdot \hat \sigma_{j,-ik}^2$, we similarly consider the case where this estimator does not exist, i.e., where $D_{ijk}=0$ for a $k$ not equal to $i$ or $j$. When any such rank deficiency is caused by both $i$ and $j$, we rely on the upward biased replacement $y_i^2 \bar \sigma_{j,-i}^2$. When none of the leave-three-out failures are caused by both $i$ and $j$, the replacement uses $\bar \sigma_{i,-jk}^2$ in place of $\hat \sigma_{i,-jk}^2$. To summarize, we define

align[align omitted — 290 chars of source]

This estimator is unbiased for $\sigma_i^2 \sigma_j^2$ when none of the leave-three-out failures are caused by both $i$ and $j$, i.e., when the first line of the definition applies. Unbiasedness holds because the presence of a bias in $\bar \sigma_{j,-ik}^2$ implies that $j$ is causing the leave-three-out failure. Therefore, $i$ cannot be the cause, which yields that $\hat \sigma^2_{i,-j}$ is independent of $y_k$, or equivalently, that $\check M_{ik,-ij} =0$.

Now, we describe how these replacement estimators enter the variance estimator $\hat{V}_{\cal F}$. When $\overline{\sigma_i^2 \sigma_j^2}$ or $\bar \sigma_{i,-jk}^2$ are biased and would enter the variance estimator with a negative weight, we remove these terms, as they would otherwise introduce a negative bias. For $\overline{\sigma_i^2 \sigma_j^2}$, the weight is $U_{ij} - V_{ij}^2$, so a biased variance product estimator is removed when $U_{ij} - V_{ij}^2 < 0$. For $\bar \sigma_{i,-jk}^2$, the weight is $V_{ij} y_j \cdot V_{ik} y_k$, but $\bar \sigma_{i,-jk}^2$ does not depend on $j$ and $k$ when it is biased, so we sum these weights across all such $j$ and $k$, and we remove the term if this sum is negative.

The following variance estimator extends the definition of $\hat{V}_{\cal F}$ to settings where leave-three-out may fail:

align[align omitted — 268 chars of source]

where the indicators $G_{ij}$ and $G_{i,-jk}$ remove biased estimators with negative weights:

align[align omitted — 450 chars of source]

Asymptotic size

In order to establish that the proposed test controls asymptotic size when there are some failures of leave-three-out, we replace the regularity condition in Assumption (ref) with an analogous version that allows for some of the determinants $D_{ij}$ and $D_{ijk}$ to be zero. Otherwise, the role of Assumption (ref) below is the same as Assumption (ref) in that it rules out denominators that are arbitrarily close to zero.

customass{3'} (i) $\max\limits_{1 \le i \le n} M_{ii}^{-1} = O_p(1)$, and (ii) $\max\limits_{i,j : D_{ij} \neq 0} D_{ij}^{-1} + \max\limits_{i,j,k : D_{ijk} \neq 0} D_{ijk}^{-1} = O_p(1)$.

When computing $\hat{V}_{\cal F}$, one must account for machine zero imperfections while comparing $D_{ij}$ and $D_{ijk}$ with zero in the definitions of $\bar \sigma_{i,-jk}^2$ and $\overline{\sigma_i^2 \sigma_j^2}$. Such imperfections are typically of order $10^{-15}$; however, we propose to compare $D_{ij}$ to $10^{-4}$ and $D_{ijk}$ to $10^{-6}$. Doing so will replace any potential case of a small denominator with an upward biased alternative and ensures that Assumption (ref)$(ii)$ is automatically satisfied.

The following theorem establishes the asymptotic validity of the proposed leave-out test in settings where Assumption (ref) fails. The theorem pertains to a nominal size below $0.31$, as the upward biased variance estimator may not ensure validity in cases where a nominal size above $0.31$ is desired. This happens because the quantile $q_{1-\alpha}(\bar F_{\hat{\boldsymbol w},n-m})$ may fall below $1$ when $\alpha$ is greater than $0.31$.

theoremIf $\alpha \in (0,0.31]$ and Assumptions (ref), (ref), (ref), and (ref) hold, then, under $H_0$, \begin{align} \limsup_{n \rightarrow \infty}\, \mathbb{P}\left( F > \hat c_\alpha \right) \le \alpha. \end{align}

An important difference between this result and that of Theorem (ref) is that the asymptotic size may be smaller than desired, which can happen when leave-three-out fails for a large fraction of possible triples. When such conservatism materializes, there will be a corresponding loss in power relative to the result in Theorem (ref). Otherwise, the power properties are analogous to those reported in Theorem (ref) and we therefore omit a formal result.

remBefore turning to a study of the finite sample performance of the proposed test, we describe an adjustment to the test which is based on finite sample considerations. This adjustment is to rely on demeaned outcome variables in the definitions of $\hat E_{\cal F}$, $ \hat V_{\cal F}$, and $\hat{\boldsymbol w}$. The benefit of relying on demeaned outcomes is that it makes the critical value invariant to the location of the outcomes. On the other hand, this adjustment removes the exact unbiasedness used to motivate the estimators of ${\mathbb{E}}_0[ {\cal F} ]$ and ${\mathbb{V}}_{0}\big[\mathcal{F} - \hat{E}_{\cal F}\big]$. However, one can show that the biases introduced by demeaning vanish at a rate that ensures asymptotic validity. Therefore, we deem the gained location invariance sufficiently desirable that we are willing to introduce a small finite sample bias to achieve it. We refer to the Appendix for exact mathematical details but note that this adjustment is used {in the simulations that follow; we also probe the version that uses non-demeaned outcomes (see the end of Section (ref)).}

Simulation evidence

{This section documents finite sample performance of the leave-out test and compares it with that of benchmark tests that could be used by a researcher in the present context:}

enumerate{ • The proposed leave-out test, which will be marked as LO in the resulting tables. • The exact F test, marked as EF, which uses critical values from the F distribution to reject when $F > q_{1-\alpha}(F_{r,n-m})$. This test has actual size equal to nominal size in finite samples under conditionally homoskedastic normal errors for any number of regressors and restrictions. It is also asymptotically valid with conditional homoskedasticity and non-normality under certain regressor homogeneity conditions anatolyev2012inference, but not under general regressor designs calhoun2011many. • Three Wald tests that reject when a heteroskedasticity-robust Wald statistic exceeds the $(1-\alpha)$-th quantile of a $\chi^2_r$ distribution, i.e., when $W > q_{1-\alpha}(\chi^2_r)$ for \begin{align} W = \big( \boldsymbol R\hat {\boldsymbol\beta}-\boldsymbol q \big)'\!\left( \boldsymbol R\boldsymbol S_{xx}^{-1} \left(\sum\nolimits_{i=1}^n \boldsymbol x_i \boldsymbol x_i' \tilde \sigma_i^2\right) \boldsymbol S_{xx}^{-1} \boldsymbol R' \right)^{-1} \!\big(\boldsymbol R\hat {\boldsymbol\beta} - \boldsymbol q \big). \end{align} The three Wald tests differ only by how one constructs variance estimates $\{\tilde \sigma_i^2\}_{i=1}^n$, and are only palliatives: \begin{enumerate} • $W_\text{1}$ most closely corresponds to the original Wald test, but with the degrees-of-freedom adjustment mackinnon2012hetero: $\tilde \sigma_i^2 = (y_i - \boldsymbol x_i' \hat{\boldsymbol \beta})^2 n/(n-m)$, • $W_\text{K}$ uses variance estimates of cattaneo2017inference, • $W_\text{L}$ uses leave-one-out estimates $\tilde \sigma_i^2 = \hat \sigma_i^2$ as in (ref). \end{enumerate} Asymptotically, the Wald tests $W_\text{L}$ and $W_\text{K}$ are valid with many regressors under arbitrary heteroskedasticity but not necessarily with many restrictions, while $W_\text{1}$ is valid only with few regressors and few restrictions under arbitrary heteroskedasticity.\footnote{{That the baseline version $W_\text{1}$ is invalid with many restrictions was noticed empirically in berndt1977conflict and shown in anatolyev2012inference under homoskedasticity; one can hardly expect that such measures as simply altering estimation of individual variances is able to solve the matters in a more complex heteroskedastic situation.}} By their comparison with the LO test one can see Wald test's potential to control size, and how much distortions are due to its wrong structure when restrictions are many. • The test based on the split-sample idea of kline2018leave is not going to be available for our simulations, because it requires regressor numerosity to be at most a half of the sample size, which is not satisfied in the simulation design.}

Simulation design

The simulation setup borrows elements of mackinnon2012hetero and adapts it to the case of many regressors as in richard2019manyboot but with richer heterogeneity in the design. The outcome equation is

equation*[equation* omitted — 100 chars of source]

where data is drawn i.i.d. across $i$. Following mackinnon2012hetero, the sample sizes take the values $80,\ 160,\ 320,\ 640,$ and $1280$. The number of unknown coefficients is $m=0.8n$ throughout to demonstrate the validity of the proposed test even with very many regressors. The null restricts the values of the last $r$ coefficients using $\boldsymbol R=\left[ \boldsymbol 0_{r\times ( m-r) },\boldsymbol I_{r}\right]$. We consider both a design that contains only continuous regressors and a mixed one that also includes some discrete regressors.

In the continuous design, the regressors $x_{i2}, \dots, x_{im}$ are products of independent standard log-normal random variables and a common multiplicative mean-unity factor drawn independently from a shifted standard uniform distribution, i.e., $0.5+u_{i}$ where $u_{i}$ is standard uniform. This common factor induces dependence among the regressors and rich heterogeneity in the statistical leverages of individual observations. For this design, we consider $r=3$ and $r=0.6n$.

When also including discrete regressors, we let $x_{i2}, \dots, x_{i,m-r}$ be as above and let the last $r$ regressors be group dummies. This mixed design corresponds to random assignment into $r+1$ groups with the last group effect removed due to the presence of an intercept in the model. The assigned group number is the integer ceiling of $(r+1)(u_i + u_i^2)/2$, where $u_i$ is the multiplicative factor used to generate dependence among the continuous regressors. By reusing $u_i$ we maintain dependence between all regressors, and by using a nonlinear transformation of $u_i$ we induce systematic variability among the $r+1$ expected group sizes. We let $r = 0.15 n$, which leads the expected group sizes to vary between $4$ and $13$ with an average group size of about $6.5$. The null corresponds to a hypothesis of equality of means across all groups.

Each regression error is a product of a standard normal random variable and an individual specific standard deviation $\sigma_i$. The standard deviation is generated by

align[align omitted — 85 chars of source]

where $s_{i}>0$ depends on the design and the multiplier $z_{\zeta }$ is such that the mean of $\sigma_i^2$ is unity. The parameter $\zeta \in \left[ 0,2\right] $ indexes the strength of heteroskedasticity, with $\zeta=0$ corresponding to homoskedasticity. We consider only the two extreme cases of $\zeta \in \left \{ 0,2\right \} $. In the continuous design, we let $s_i = \sum_{k=2}^{m}x_{ik}$, and in the mixed design, $s_i = \sum_{k=2}^{m-r}x_{ik} + z_u u_i$. The factor $z_u = 2r \exp(1/2)$ ensures that $s_i$ has the same mean in both designs.

Under the null, the coefficients on the continuous regressors are all equal to $\varrho$, where $\varrho$ is such that the coefficient of determination, $\mathrm{R}^{2}$, equals $0.16$. The coefficients on the included group dummies are zero, which correspond to the null of equality across all groups. The intercept is chosen such that the mean of the outcomes is unity. For the continuous design this yields an intercept of $1-(m-1)\varrho \exp(1/2)$, while the intercept is $1-(m-r-1)\varrho \exp(1/2)$ in the mixed design. With these parameter values, the null is $(\beta_{m-r},\dots,\beta_r)' =\boldsymbol q$, where $\boldsymbol q=(\varrho,\dots,\varrho)'\in \mathbb{R}^r$ in the continuous design, and $\boldsymbol q = (0,\dots,0)' \in \mathbb{R}^r$ in the mixed design.

To document power properties, we consider both a sparse and dense deviations from the null, and focus on the settings where $r$ is proportional to $n$. In parallel to the theoretical power analysis in Section (ref), we consider deviations for the last $r$ coefficients that are parameterized using

align[align omitted — 178 chars of source]

where we use the lower triangular square-root matrix. This choice of square-root implies that the alternative is sparse when only the last few entries of $\boldsymbol\delta$ are non-zero. As shown in Section (ref), asymptotic power is governed by the norm of $\boldsymbol\delta$ over $r^{1/4}$, but whether an alternative is fixed or local, additionally depends on the rate at which the tested coefficients are estimated. This rate is governed by $\mathbb{E}\boldsymbol [\boldsymbol S_{xx}]$, which is reported in the Appendix.

In the continuous design, the tested coefficients are estimated at the standard parametric rate of $n^{-1/2}$. To specify a fixed sparse alternative we therefore use ${\boldsymbol\delta} = 0.5n^{1/2} (0,\dots,0,1)' \in \mathbb{R}^{r}$, for which $\beta_m$ differs from the null value by approximately $0.2$ (here and hereafter, the scaling is chosen so that the power is bounded away from the size and away from unity for the sample sizes we consider). Since the norm of $\boldsymbol\delta$ grows faster than $r^{1/4}$, the power will be an increasing function of the sample size. For the dense alternative, we consider instead ${\boldsymbol\delta} = 0.5 n^{1/2} r^{-1/2} {\boldsymbol\iota}_r$ where ${\boldsymbol\iota}_r = (1,\dots,1)' \in \mathbb{R}^{r}$, for which all deviations between the tested coefficients and $\varrho$ shrink at the standard parametric rate of $n^{-1/2}$. Here, power is again increasing in the sample size due to numerous deviations from the null.

In the mixed design, the group effects are not estimated consistently as the group sizes are bounded. A possible fixed sparse alternative is then ${\boldsymbol\delta} = (0,\dots,0,6)' \in \mathbb{R}^{r}$, for which $\beta_m$ differs from the null value of zero by roughly $3$. In contrast to the continuous design, the power will decrease with sample size as the precision, with which $\beta_m$ can be estimated, does not increase with $n$. For the dense alternative, we use ${\boldsymbol\delta} = 1.5 {\boldsymbol\iota}_r$, which corresponds to a fixed alternative for every tested coefficient. Here, the power will be increasing in $n$ due to the numerosity of deviations.

Simulation results

We present rejection rates based on 10000 Monte-Carlo replications and consider tests with nominal sizes of $1\%,$ $5\%$ and $10\%$. Furthermore, we report the frequency with which the proposed variance estimate $\hat{V}_{\cal F}$ is negative and therefore replaced by the upward biased and positive alternative introduced in Remark (ref). For the design that includes discrete regressors, we also report the average fraction of observations that cause a failure of leave-three-out full rank, and for which we therefore rely on an upward biased estimator of the corresponding error variance. For all sample sizes, this fraction is around $7\%$ in the mixed design, which corresponds to the percentage of observations that belong to groups of size 2 or 3. The fraction is zero in the design that only involves continuous regressors.

table[table omitted — 6,084 chars of source]

Table (ref) contains the actual rejection rates under the null for both the continuous and mixed designs. In settings with many regressors and restrictions, the considered versions of the “heteroskedasticity-robust" Wald test fail to control size irrespective of the design, presence of heteroskedasticity, and nominal size. The failure of the conventional Wald test, $W_1$, is spectacular, with type I error rates close to one for the continuous design, but the two versions that are robust to many regressors, $W_K$ and $W_L$, also exhibit size well above the nominal level. With few restrictions, the Wald tests show a more moderate inability to match actual size with nominal size, and the table suggests that the leave-one-out version, $W_L$, can control size in samples that are somewhat larger than considered here. Under homoskedasticity, the table reports that the exact F test indeed has exact size. However, in the heteroskedastic environments with many restrictions the exact F test is oversized with a type I error rate that approaches unity as the sample size increases.

By contrast, the proposed leave-out test exhibits nearly flawless size control as it is oversized by at most one percent across nearly all designs, nominal sizes, and whether heteroskedasticity is present or not. In the smallest sample for the continuous design, the test is somewhat conservative, presumably due to the relatively high rate of negative variance estimates ($20\%$ with homoskedasticity and $13\%$ with heteroskedasticty) that are replaced by a strongly upward biased alternative. This rate diminishes quickly with sample size, and the fraction of negative variance estimates is already essentially zero in samples with 640 observations and 512 regressors. In the mixed design, negative variance estimates are even less prevalent, potentially due to the fact that the test uses some upward biased variance estimators for $7\%$ of observations. Perhaps somewhat surprisingly, having $7\%$ of observations causing failure of leave-three-out is not sufficient to bring about any discernible conservativeness in the leave-out test for this design.

Table (ref) contains simulated rejection rates for the continuous and mixed designs under alternatives where the parameters deviate from their null values in one of two ways -- either one tested coefficient deviates (sparse) or all tested coefficients deviate (dense). The table reports these power figures for tests with a nominal size of $5\%$ and $10\%$ that also control the size well, i.e., the LO and exact F tests under homoskedasticity and the LO test under heteroskedasticity.

table[table omitted — 2,772 chars of source]

For the continuous design, the power of the tests increases from slightly above nominal size to somewhat below unity as the number of observations increases from $80$ to $1280$. This pattern largely holds irrespective of the type of deviation and presence of heteroskedasticity, although the LO test is a bit more responsive to sparse deviations than to dense ones. Along this stretch of the power curve, the LO test exhibits a power loss that varies between $4$ and $16$ percentage points when compared to the exact F test, and in relative terms, this gap in power shrinks as the sample size grows. Given that the number of tested restrictions in this setting is above half of the sample size, we conjecture that these figures are towards the high end of the power loss that a typical practitioner would incur in order to be robust with respect to heteroskedasticity.

In the mixed design, the fixed dense alternative exhibits similar power figures as in the continuous design, while the fixed sparse deviation generates a power function that decreases with sample size. The reason for the latter is, as discussed in the previous subsection, that the deviating group effect is not estimated more precisely as additional groups are added to the data. Upon comparison of the LO and exact F tests, we see that the differences in the power figures are only $0$--$7$ percentage points. In light of Remark (ref), which explains that there is no power difference between the LO and exact F tests when ${r}/{n}$ is small, it is natural to attribute this almost non-existent power loss to the fact that there are four times fewer tested restrictions in this mixed design than in the continuous one.

We have also run additional simulation experiments with our baseline continuous regressor design, where we track the impact of the relative numerosities of regressors and restrictions $r/m$ and $m/n$ and of the coefficient of determination $\mathrm{R}^{2}$ on the positivity failure rate of $\hat{V}_{\cal F}$, the empirical size of the test, and its empirical power. The two numerosity ratios show the severity of deviations from the standard regression testing setup, while the coefficient of determination summarizes the magnitude of regression coefficients relative to the size of error variances. The results are relegated to the Appendix (Tables A1 and A2), and here we give a brief summary. The general observation is that the percentage of negative $\hat{V}_{\cal F}$ positively varies with all three parameters, with $\mathrm{R}^{2}$ having most pronounced impact and the ratio $r/m$ having smallest impact. The percentage, however, is still kept within 0.1-0.2% for all combinations when $n=640$, and the negativity issue is practically non-existent when $n=1280$. Next, while the three parameters do affect some of the wrongly sized tests from the existing literature, they do not influence the actual empirical size of our proposal, except for minor variation in very small samples. The empirical power of the proposed test, however, is non-trivially affected by all the three parameters, whose higher values imply somewhat smaller power. The coefficient of determination, in particular, has such an effect because a higher signal relative to noise increases the variability of the individual error variance estimators relative to their targets, and so the power tends to be negatively affected by large(r) coefficients. The numerosity ratios also have a negative effect on power because the signal gets dispersed across a larger number of regressors or restrictions as the ratios increase, which naturally reduces power. These tendencies are shared by the EF test when it is appropriately sized.

{Finally, we have examined the differences that result from the use of non-demeaned outcomes when estimating individual variances and their products (see Remark (ref)). The general impression from those simulations is, first, the use of non-demeaned outcomes makes size control less stable; in particular, for smaller sample sizes, the LO test is undersized. Second, it seriously decreases power at all sample sizes. In practice, we therefore recommend exploiting the version with demeaned outcomes.}

Concluding remarks

This paper develops an inference method for use in a linear regression with conditional heteroskedasticity where the objective is to test a hypothesis that imposes many linear restrictions on the regression coefficients. The proposed test rejects the null hypothesis if the conventional F statistic exceeds a linearly transformed quantile from the F-bar distribution. The central challenges for construction of the test is estimation of individual error variances and their products, which requires new ideas when the number of regressors is large. We overcome these challenges by using the idea of leaving up to three observations out when estimating individual error variances and their products. In some samples the variance estimate used for rescaling of the critical value may either be negative or cease to exist due to the presence of many discrete regressors. For both of these issues, we propose an automatic adjustment that relies on intentionally upward biased estimators which in turn leaves the resulting test somewhat conservative. Simulation experiments show that the test controls size in small samples, even in strongly heteroskedastic environments, and only exhibits very limited adjustment-induced conservativeness. The simulations additionally illustrate good power properties that signal a manageable cost in power from relying on a test that is robust to heteroskedasticity and many restrictions.

Bootstrapping and closely related resampling methods are often advocated as automatic approaches for construction of critical values. However, in the context of linear regression with proportionality between the number of regressors and sample size, multiple papers bickel1983bootstrapping,elkaroui2018hidimboot,cattaneo2017inference demonstrate invalidity of standard bootstrap schemes even when inferences are made on a single regression coefficient. Under additional assumptions of homoskedasticity and restrictions on the design, elkaroui2018hidimboot and richard2019manyboot show that problem-specific corrections to bootstrap methods can restore validity. We leave it to future research to determine whether bootstrap or other resampling methods can be corrected to ensure validity in our context of a heteroskedastic regression with many regressors and tested restrictions.\looseness=-1