Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,595 characters · 10 sections · 74 citation commands
Bootstrap Consistency for Quadratic Forms of Sample Averages with Increasing Dimension
\affil{UC Berkeley}
Since its introduction by efron1979 the bootstrap has been widely used as a method for approximating the distribution of statistics. Many papers have extended the original idea in terms, both, of the applicability (see Horowitz2001 and Hall1986 for excellent reviews) and of its methodology; of particular interest for us are the bootstrap procedures: \textquotedblleft wild bootstrap" (see mammen1993) and more generally the \textquotedblleft weighted bootstrap" (see Ma2005).
In this paper we attempt to expand the applicability of the weighted bootstrap procedure to quadratic forms with increasing dimensions. Namely, we study quadratic forms of the form
where $(Z_{1,n},...,Z_{n,n})$ are independent (among each other) $\mathbb{R}^{d}$-valued random variables with mean zero and general covariance matrix $\Sigma_{n}$. We show that its distribution is well-approximated (under the Kolmogorov distance) by the distribution of
where $(\omega_{1,n},...,\omega_{n,n})$ are independent bootstrap weights. The novelty in this paper is that we allow for $d=d(n)$ to increase with the sample size.
Studying the asymptotic behavior of quadratic forms, in particular establishing bootstrap consistency, is relevant since many statistics of interest can asymptotically be represented as quadratic forms of (scaled) sample averages. For instance, the likelihood ratio and Wald test statistics are asymptotically represented as quadratic forms of the scores; see VdV00 Ch. 16, and references therein. Portnoy_AOS88 establishes such representations for the likelihood ratio test statistics; there $d(n)$ is the dimension of the parameter of interest and is allowed to grow with $n$. HmcKVK_AOS09 uses Portnoy's results to show a quadratic approximation result for Owen's (Owen-90) empirical likelihood, allowing for $d(n)^{3}/n \rightarrow 0$; see also Peng-Schick-12. Therefore, by establishing the validity of the bootstrap for general quadratic forms, we propose an alternative method for inference for these statistics.
So as to further illustrate the applicability of our results, in Section (ref) we study a concrete application motivated by the work of DIN_JOE03 who consider model-specification tests for models defined by a diverging number of moment conditions (this quantity determines our $d(n)$). By applying our results, we establish bootstrap consistency results for the distribution of the model-specification test statistics of two ubiquitous estimators in econometrics and statistics: The generalized empirical likelihood (GEL; Smith-EJ97) estimator and The generalized method of moments (GMM; Hansen-ECMA82) estimator. By employing our bootstrap result we are able to perform inference for non-optimally weighted GMM estimators. To our knowledge these results are new.
By letting $d$ to increase with sample size in our general theory, we allow for different asymptotics, a \textquotedblleft large-$d$ and large-$n$" asymptotics, rather than the standard \textquotedblleft fixed-$d$ and large-$n$". The former type of asymptotics are more explicit about how the dimension, $d$, can affect the quality of the approximations. That is, even if the dimension does not literally grow with $n$, if, for instance, the model has a large number of parameters (or moment conditions as in our application), doing \textquotedblleft fixed-$d$ and large-$n$" asymptotics could be misleading, whereas doing \textquotedblleft large-$d$ and large-$n$" asymptotics could depict a more accurate picture of the behavior for fixed samples; see Mammen_AOS89 for discussion. Our results can also be applied in cases where there is literally a growing number of parameters. For instance, ChenPouzo14 study the asymptotic behavior of the quasi-likelihood ratio and Wald test statistics in a semi-parametric conditional moment setup; in particular they show that the statistics are asymptotically equivalent to quadratic forms ((ref)) under a null hypothesis of increasing dimensions (see Appendix A.4 in their paper); our results, in conjunction with theirs, could be applied to establish bootstrap-based inference for the quasi-likelihood ratio and Wald test statistics.\footnote{In Section (ref) we provide more concrete examples of these two cases in the context of our application.}
In order to establish our main result of bootstrap consistency, we use Lindeberg interpolation techniques (see Chatterjee_AOP06, Rollin2013 and references therein) to approximate the quadratic forms of $n^{-1/2}\sum_{i=1}^{n} \omega_{i,n} Z_{i,n}$ and $n^{-1/2}\sum_{i=1}^{n} Z_{i,n}$ by the ones for Gaussian random variables with zero mean and covariance $n^{-1}\sum_{i=1}^{n} Z_{i,n}Z_{i,n}^{T}$ and $E[Z_{1,n}Z_{1,n}^{T}]$, respectively.
By proceeding in this manner, we are able to reduce the original problem to a Gaussian approximation problem wherein we need to establish convergence of a Gaussian distribution with zero mean and variance $n^{-1} \sum_{i=1}^{n} Z_{i,n}Z_{i,n}^{T}$ to one with zero mean and variance $E[Z_{1,n}Z_{1,n}^{T}]$. We use Slepian interpolation (Slepian1962, Rollin2013, CCK2013 and references therein) to accomplish this.
Due to the interpolation techniques used here, we need certain restrictions on the higher moments of the random variables. In particular, we impose growth restrictions on the higher moments of the bootstrap weights and the Euclidean norm of $Z_{1,n}$. These conditions essentially restrict the growth rate of $d(n)$. Although the precise growth rate depends on such conditions, the dimensions cannot grow faster than $n^{1/4}$.
A number of papers develop large sample results allowing for increasing dimension. To name a few, Portnoy_AOS88 establishes the validity of the Wilks phenomenon for the likelihood ratio for exponential families when $d(n)^{3/2}/n \rightarrow 0$. He-Shao2000 derive the asymptotic distribution for M-estimators when the number of parameters is allowed to grow with the sample size. Recently, a few papers develop this type of results for quadratic forms of the form ((ref)) allowing for increasing dimensions. In particular, Peng-Schick-12 and Xu_Zhang_Wu_14 develop a central limit theorem for quadratic forms of sample averages of vectors, allowing for the dimension to grow with $n$; both papers discuss several applications and examples. The results on our paper offer an alternative, bootstrap-based, method for inference for these cases.
Our paper also contributes to the growing literature of bootstrap results allowing for increasing dimensions. Mammen_AOS89 derives asymptotic expansion for M-estimators in linear models allowing for increasing dimension and use them to show consistency of a weighted bootstrap. In a different context, Radulovic-98 uses Lindeberg interpolation methods allowing for increasing dimension to show that the functional bootstrap CLT holds under weaker conditions than equicontinuity; in his paper the restriction over the growth rate is $d(n)^{6}/n \rightarrow 0$. In CCK-AOS13, the authors derive a Gaussian weighted bootstrap approximation result for the maximum of the sum of high dimensional random vectors; in this specific setup the dimension is allowed to grow very fast, even at an exponential rate. Zhang-Cheng-14 provide an extension of CCK-AOS13 to time series. In our paper the object of interest is the $\ell^{2}$-norm of the sum of high dimensional random vectors (as opposed to the $\ell^{\infty}$-norm), so the results in these papers are not directly applicable. Finally, in a recent independent work, Spoikony-Zhilova-14 study the validity of the weighted bootstrap procedure for the likelihood ratio test statistics in finite samples and model misspecification; their results require $d(n)^{3}/n$ to be \textquotedblleft small".
Organization of the Paper. In Section (ref) we define the problem and impose the required assumptions. Section (ref) presents the main Theorem and a discussion of its implications. Section (ref) presents an application to model-specification tests. Section (ref) presents a numerical simulations. Section (ref) presents the proof of the main Theorem. Section (ref) presents some concluding remarks. In order to keep the paper short, the proofs of intermediate results are gathered in the appendix.
Notation. For any vector $x \in \mathbb{R}^{d}$, we use $||x||^{p}_{p}$ to denote $\sum_{l=1}^{d} |x_{l}|^{p}$ and $x_{[l]}$ to denote the $l$-th coordinate of the vector. $tr\{ A \}$ denotes the trace of matrix $A$. We use $E_{P}$ to denote the expectation with respect to the probability measure $P$; for conditional distributions $P(\cdot|X)$ we use $E_{P(\cdot|X)}[\cdot]$ or sometimes directly $E_{P}[\cdot |X]$. We use $ X_{n} \precsim Y_{n}$ to denote that $ X_{n} \leq C Y_{n}$ for some universal $C>0$. We use $\partial^{r} f$ to denote the $r$-th derivative of $f$; for the cases of $r=1$ and $r=2$ we use the more standard $f'$ and $f''$ notation. $wpa1-P$ means \textquotedblleft with probability approaching one under $P$".
Let $\{ Z_{i,n} \in \mathbb{R}^{d(n)} : i=1,...,n~and~n \in \mathbb{N} \}$ with $(d(n))_{n\in \mathbb{N}}$ being a non-decreasing integer-valued sequence; $d(n)$ could diverge to infinity. For all $n \in \mathbb{N}$, let $Z^{n} \equiv (Z_{1,n},...,Z_{n,n})$ be independent among themselves with $Z_{i,n} \sim \mathbf{P}_{n}$ and $E_{\mathbf{P}_{n}} [ (Z_{i,n}) ] = 0$ and $\Sigma_{n} \equiv E_{\mathbf{P}_{n}} [ (Z_{i,n}) (Z_{i,n})^{T} ] \in \mathbb{R}^{d(n)\times d(n)}$ positive definite and finite. Henceforth, we will typically omit the sub-index $n$ in $Z_{i,n}$.
Let $\mathbb{Z}_{n} \equiv n^{-1} \sum_{i=1}^{n} Z_{i}$, and
For a given matrix $A \in \mathbb{R}^{ d \times d }$ we denote its eigenvalues as $\{ \lambda_{1}(A),..., \lambda_{d}(A)\} $.
The assumption that $ c \leq \lambda_{l}(\Sigma_{n}) \leq C$ can be somewhat relaxed; for instance, it could be replaced by $\limsup_{n \rightarrow \infty }\frac{ tr \{ \Sigma^{3}_{n} \} }{ ( tr \{ \Sigma^{2}_{n} \} )^{3/2} } =0 $ and $\frac{ tr \{ \Sigma_{n} \} }{ tr \{ \Sigma^{2}_{n} \} } \leq C < \infty $. The rest of Assumption (ref) essentially imposed restrictions on the rate of growth of $d(n)$ relative to $n$. In order to provide sufficient conditions for this part of Assumption (ref), it is convenient to provide bounds in terms of $d(n)$ for the quantities $E_{\mathbf{P}_{n}}[||Z_{1}||_{2}^{q}]$ (for different $q$'s) and $E_{\mathbf{P}_{n}} [ ||Z_{1}||^{2(2+\kappa)}_{2+\kappa} ]$ in the assumption.
Clearly, if $|Z_{[l],1}| \leq C <\infty$ a.s-$\mathbf{P}_{n}$ for all $l = 1,...,d(n)$ and all $n \in \mathbb{N}$, then $E_{\mathbf{P}_{n}}[||Z_{1}||_{2}^{2q}] = O(d(n)^{q})$ for any $q>0$.\footnote{Recall that for a vector $x$, $x_{[l]}$ denotes the $l$-th component.} For example, such condition is imposed by Vershynin-JTP12 in the context of estimation and approximation of covariance matrices of high dimensional distributions.
The next lemma shows that the result still holds if we impose the following (milder) restriction: $E_{\mathbf{P}_{n}} \left[e^{\lambda Z_{[l],1}^{2} } \right] \leq C < \infty $ for some $\lambda>0$. For instance, if $(Z_{[l],1})^{2}$ is a sub-Gamma random variable (Boucheron-book13 p. 27), then the condition holds since $E_{\mathbf{P}_{n}} \left[e^{\lambda Z_{[l],1}^{2} } \right] \leq \exp \{ \frac{\lambda^{2} v }{2(1-c\lambda)} \}$ for any $\lambda \in (0,1/c)$ and some $c>0$. If $Z_{[l],1}$ is sub-Gaussian, then $(Z_{[l],1})^{2}$ is sub-exponential (see Vershynin-12 Lemma 5.14) and the condition holds by the same argument.
An appealing feature of this result is that it only imposes restrictions on the marginal behavior of the components of the vector $Z_{1}$ and not on its joint behavior.
Under the conditions in the lemma, Assumption (ref)(i) boils down to $\frac{d(n)^{4}}{n} = o(1) $. For Assumption (ref)(ii) is sufficient to impose $\frac{d(n)^{4+2\gamma}}{n^{\gamma}} = o(1)$; for $\gamma = 2$ it boils down to $\frac{d(n)^{4}}{n} = o(1)$ but for large $\gamma$ it (roughly) becomes $\frac{d(n)^{2}}{n} = o(1)$. Finally, for, say $\kappa = 0$, Assumption (ref)(iii) is reduced to $\frac{ d(n)^{2} }{ n } E_{\mathbf{P}_{n}} [ ||Z_{1}||^{4}_{2} ] \precsim \frac{ d(n)^{4} }{ n } \rightarrow 0$.
That is, under conditions that bound all (polynomial) moments of the individual components of $Z_{1}$, the dimension is allowed to grow slower than the 4th-root of the sample size.
The bootstrap weights are given by $\{ \omega_{in} \in \mathbb{R} : i=1,...,n~and~n \in \mathbb{N} \}$ where, for any $n \in \mathbb{N}$ and conditional on $Z^{n} = z^{n}$, $(\omega_{1n},...,\omega_{nn}) \sim \mathbf{P}^{\ast}_{n}(\cdot | z^{n}) $ for some $\mathbf{P}^{\ast}_{n}(\cdot | z^{n}) $.
Part (i) is standard. Part (ii) is mild considering that the weights are chosen by the researcher.\footnote{Of course, the technique of proof can be applied to the case where the following (stronger) restriction is imposed: $E_{\mathbf{P}^{\ast}_{n}(\cdot | Z^{n})}\left[ \exp \{ \omega_{in} \} \right] \leq C_{w} < \infty $.}
We now present the main result of the paper. In what follows, for any measurable function $z^{n} \mapsto f(z^{n})$ we use $|f(Z^{n})| = o_{\mathbf{P}_{n}}(1)$ to denote: For any $\varepsilon>0$, there exists a $N(\varepsilon)$ such that for all $n \geq N(\varepsilon)$, $\mathbf{P}_{n}( |f(Z^{n}) | \geq \varepsilon ) < \varepsilon$.
Let $\mathbb{Z}^{\ast}_{n} \equiv n^{-1} \sum_{i=1}^{n} \omega_{in} Z_{i}$ be the bootstrap analog of $\mathbb{Z}_{n}$.
We now present some remarks and discuss some implications of the preceding Theorem.
Heuristics. We postpone the somewhat long proof of the Theorem to Section (ref); here we present an heuristic argument. The first step of the proof is to apply Lindeberg interpolation techniques (see Chatterjee_AOP06 and Rollin2013 and references therein) to approximate $\sqrt{n} \mathbb{Z}^{\ast}_{n}$ by $\sqrt{n} \mathbb{U}_{n}$ and $\sqrt{n} \mathbb{Z}_{n}$ by $\sqrt{n} \mathbb{V}_{n}$, where $\mathbb{U}_{n}$ and $\mathbb{V}_{n}$ are Gaussian random variables with zero mean and covariances $n^{-1} \sum_{i=1}^{n} Z_{i}Z_{i}^{T}$ and $E[Z_{1,n}Z_{1,n}^{T}]$ respectively.
In order to do this, we first approximate the indicator function $x \mapsto 1\{ ||x||^{2}_{2} \geq t \}$ by \textquotedblleft smooth" functions $x \mapsto \mathcal{P}_{t,\delta,h}(||x||^{2}_{2})$; the exact expression for $\mathcal{P}_{t,\delta,h}$ is presented in Lemma (ref) and follows from the suggestion by Pollard_01 p. 247. The functions are indexed by $(h,\delta)$ where $h$ is \textquotedblleft small” compared to $\delta$, and the \textquotedblleft smaller" $\delta$ is, the closer the function $\mathcal{P}_{t,\delta,h}$ is to the indicator function; see Lemmas (ref), (ref) and (ref) in the Appendix (ref). It is worth to note that what we mean by $\delta$ to be “small” depends on how $||\sqrt{n}\mathbb{V}_{n}||^{2}_{2}$ concentrates mass. Lemma (ref) in the Appendix (ref) establishes an anti-concentration result, wherein we obtain that this random variable puts very little mass in any given interval. Therefore $\delta$ could actually be quite large, of the order of $\sqrt{tr\{ \Sigma^{2}_{n} \}}$.
Second, since $x \mapsto \mathcal{P}_{t,\delta,h}(||x||^{2}_{2})$ belongs to a class of \textquotedblleft smooth" functions, we show that it suffices to show consistency under the weak norm (as opposed to the norm implied in (ref)).\footnote{The formal definition of the norm is presented in Equation (ref) in Section (ref).} This is done in Lemmas (ref) and (ref). The relevant class of \textquotedblleft smooth" functions is given by $\mathcal{C}_{M}$, which is the class of functions $f : \mathbb{R} \rightarrow \mathbb{R}$ that are three times continuously differentiable and $ \sup_{x} | \partial^{r} f(x)| \leq (M)^{r}$ and $\sup_{x} |f(x)| \leq 1$.
The following Theorems formalize the aforementioned approximation of $\sqrt{n} \mathbb{Z}^{\ast}_{n}$ by $\sqrt{n} \mathbb{U}_{n}$ and $\sqrt{n} \mathbb{Z}_{n}$ by $\sqrt{n} \mathbb{V}_{n}$ and can be viewed of independent interest since they show that a \textquotedblleft generalized invariance principle" holds in our setup. Henceforth, we use $\boldsymbol{\Phi}^{\ast}_{n}(\cdot|Z^{n})$ and $\boldsymbol{\Phi}_{n}$ respectively, to denote their probability distributions.
By using Theorems (ref) and (ref) we have reduced the original problem to a Gaussian approximation problem. That is, we need to establish convergence (under the distance induced by $\mathcal{C}$) of a Gaussian distribution with zero mean and variance $n^{-1} \sum_{i=1}^{n} Z_{i}Z_{i}^{T}$ to one with zero mean and variance $E[Z_{1}Z_{1}^{T}]$. Lemma (ref) in Section (ref) --- which is based in the Slepian interpolation (see CCK-AOS13, CCK2013 and Rollin2013 and references therein)--- establishes that is enough to show that
In Section (ref), we show that, employing standard arguments, the expression (ref) holds under our assumptions. A similar result is obtained by CCK-AOS13 without the scaling factor of $d(n)$; their setup, however, is different since the object of interest is $\max_{1\leq j \leq d(n)} |n^{-1/2} \sum_{i=1}^{n} Z_{[j],i}|$ (as opposed to $||n^{-1/2} \sum_{i=1}^{n} Z_{i}||^{2}_{2}$). \footnote{An important consequence of this difference is that, as opposed to our case, CCK-AOS13 can use a \textquotedblleft smooth maximum function" to approximate their quantity of interest; the approximation error is only of order $\log d$. This, allows them to obtain faster rates for the approximation of the indicator functions with smooth functions. This, in turn, translates into a faster overall rate of convergence --- $d = o(\exp(n))$ in their case. See Wasserman_2014 for a discussion and a nice review of these results.}
Asymptotic Distribution of $||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2}$. An implication of the proof of Theorem (ref) and Theorem (ref) is that
That is, if $\Sigma_{n} = I_{d(n)}$ then this expression and a direct application of the CLT (when $d(n) \rightarrow \infty$) imply that $\frac{ ||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2} - d(n) }{\sqrt{2 d(n)}} \Rightarrow N(0,1)$ or, informally, $||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2}$ is approximately chi-square distributed with $d(n)$ degrees of freedom. When $\Sigma_{n} \ne I_{d(n)}$, the last claim is no longer true but it holds that $\frac{ ||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2} - tr\{ \Sigma_{n} \} }{\sqrt{2 tr\{\Sigma^{2}_{n} \}}}$ is approximately distributed as $\sum_{j=1}^{d(n)} \frac{\lambda_{j}(\Sigma_{n})(\chi_{j}-1)}{\sqrt{2\sum_{j=1}^{d(n)} \lambda^{2}_{j}(\Sigma_{n})}}$ with $\chi^{2}_{j}$ drawn from a chi-square with degree one; see Xu_Zhang_Wu_14 and Peng-Schick-12 for a discussion regarding these results.
We note that in Theorem (ref) no scaling (by $-d(n)$ and $1/\sqrt{2d(n)}$ or $-tr\{\Sigma_{n} \}$ and $1/\sqrt{2 tr\{ \Sigma^{2}_{n} \}}$) is needed. That is, although the mean and variance of $||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2}$ are \textquotedblleft drifting" to infinity, the bootstrap still provides a good approximation since the moments of $||\sqrt{n} \mathbb{Z}^{\ast}_{n}||^{2}_{2}$ are mimicking this behavior.
On the Lindeberg Interpolation. Theorems (ref) and (ref) are based on the following Lindeberg interpolation for quadratic forms.\footnote{This Lindeberg interpolation builds on the approach in Xu_Zhang_Wu_14. }
It is worth pointing out that the interpolation compares the quantities $\sum_{i=1}^{n} A_{i}$ with $\sum_{i=1}^{n} B_{i}$ by comparing \textquotedblleft one component at a time". This comparison is essentially divided into two parts. First, we compare $||\mathbb{S}_{i:n} + A_{i} ||^{2}_{2}$ and $||\mathbb{S}_{i:n}+ B_{i}||^{2}_{2}$, which are real-valued quantities. Second, we exploit the smoothness of the univariate function $f$ to bound its variation using Taylor's approximation. Loosely speaking, the first step reduces a $d(n)$-dimensional problem to an univariate one. An alternative approach would be to consider interpolations for multivariate functions (e.g. Chatterjee-Meckes-2008) of the form $g : \mathbb{R}^{d(n)} \rightarrow \mathbb{R}$ with $g(x) \equiv f(||x||^{2}_{2})$. As can be seen from the derivations in Chatterjee-Meckes-2008, the remainder term will also require bounds on higher derivatives of $g$ (and thus $f$), but of the form $\sup_{x \ne y} \frac{\left \Vert Hess(g)(x) - Hess(g)(y) \right \Vert_{op} }{||x-y||_{2}}$. \footnote{$Hess(g)$ is the Hessian of the function and $||.||_{op}$ is the operator norm. Other type of bounds could be found in Raic2014 based on Hilbert-Schmidt norm.} Which approach is better depends largely on what type of restrictions over the class of test functions are natural in the problem at hand. For us, $||\partial^{r} f||_{L^{\infty}} < \infty$ is a natural assumption, but in other applications it could be too strong.
More generally, this discussion illustrates the relationship between restrictions in the class of test functions ($\mathcal{C}$) and the bounds on higher order moments and ultimately the rate of growth of $d(n)$.
Bootstrap P-Value. For any $\alpha \in (0,1)$ and $Z^{n} \in \mathbb{R}^{d(n)}$, let $t_{n}(\alpha,Z^{n}) \equiv \inf\{ t : \mathbf{P}^{\ast}_{n} \left( ||\sqrt{n} \mathbb{Z}^{\ast}_{n}||^{2}_{2} \leq t \mid Z^{n} \right) \geq \alpha \}$. Due to the distribution consistency result proven in Theorem (ref), we can approximate the $\alpha$-th quantile of the distribution of $ ||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2} $ by $t_{n}(\alpha,Z^{n})$, in the sense that
for any $\eta>0$. If $t_{n}(\alpha,Z^{n})$ is a continuity point of $\mathbf{P}^{\ast}_{n} \left( \cdot | Z^{n} \right)$, then
and the first display becomes $\mathbf{P}_{n} \left( ||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2} \geq t_{n}(\alpha,Z^{n}) \right) = \alpha + o(1)$. Hence, Theorem (ref) can be used to construct valid p-values based on the bootstrap.
In this section we apply our results to construct bootstrap-based specification tests for models with increasing number of moment restrictions. We do this for two estimators: generalized method of moment (GMM; see Hansen-ECMA82) estimator and generalized empirical likelihood (GEL; see Smith-EJ97) estimator. Both estimators are widely used in econometrics and statistics and encompass a wide range of commonly used estimators such as Z-estimators (VdV00 Ch. 5), and empirical likelihood estimator (Owen-BIO88), respectively.\footnote{See Imbens-JBES02 for additional examples and a discussion. See also Hall for a review for GMM.}
In models characterized by moment conditions, model-specification tests (MST) allow us to check whether the moment conditions match the data well or not. In this setup with increasing moment restrictions, MST has been studied by DIN_JOE03 (DIN, henceforth); see also Djong-Bierens-ET94. They show that the MST statistic is asymptotically a quadratic form of scaled sample averages; however, they rely on inferential methods build on expressions akin to (ref). Instead, by applying our Theorem (ref), we can use the weighted bootstrap method to approximate the asymptotic distribution of MST statistics; thus complementing their results by providing an alternative way of constructing asymptotic p-values. Moreover, as explained below, by not relying on CLT-type results to approximate the limiting distribution, we are able to provide valid asymptotic inference for a larger class of GMM estimators than the one considered in DIN.
The setup closely follows that of DIN and is as follows. Suppose $(X_{i})_{i=1}^{n}$ is an i.i.d. sample of real-valued random variables with $X_{i} \sim \mathbf{P}_{n}=\mathbf{P}$. The model we consider is one where the true parameter of interest, $\theta_{0} \in Int(\Theta)$ --- with $\Theta$ a compact subset of $\mathbb{R}^{q}$--- is uniquely identified by the following set of moment conditions
where $g : \mathbb{R} \times \mathbb{R}^{q} \rightarrow \mathbb{R}^{d}$ is known to the researcher.
The main feature of this setup is that it allows $d\equiv d(n)$ to grow with the sample size. In many cases this departure from the standard theory is of relevance. For example, in many models the identifying condition is given by a conditional moment restriction, $E_{\mathbf{P}}[\rho(Y,\theta_{0})|W]$ --- where $\rho$ maps into $\mathbb{R}^{J}$ with $J$ fixed --- and the researcher converts it to a series of unconditional moment restrictions $E_{\mathbf{P}}[\rho(Y,\theta_{0}) \otimes q^{K(n)}(W)]$ where $q^{K(n)}(w) = (q_{1}(w),...,q_{K(n)}(w))$ are basis functions such as Fourier series, P-splines, etc; this is the case considered in DIN (see also Djong-Bierens-ET94 and references therein). For this case $x = (y,w)$, $d(n) = J K(n)$ and $g(x,\theta) = \rho(y,\theta) \otimes q^{K(n)}(w)$.
An alternative motivation to consider increasing $d$ would be cases where although the number of moments is fixed, it could be large relative to the sample size and thus treating it as a diverging sequence could deliver more accurate asymptotics. As pointed out by K-M_JOE93 one example of this could be the panel data model in AB-RESTUD91 where $x = (y_{1},...,y_{T})$ and the components of the vector $g(x,\theta)$ are given by $((y_{t}-y_{t-1}) - \theta (y_{t-1}-y_{t-2})) y_{t-s} $ for $s=1,...,t-1$ and $t=3,...,T$. Here, for a panel of length $T$, the number of instruments/moments is given by $d=(T-2)(T-1)/2$. \footnote{For instance for $T=4$, $d=d(n)=3$ and for $T=5$, $d=d(n)=6$. In cases where $d(n) = o(n^{1/4})$, these values imply that, roughly speaking, the number of observations should be larger than 82 and 1300, resp. It is also worth to point out that in case where $T$ is large, one can simply include fewer lags $y_{t-s}$ in $((y_{t}-y_{t-1}) - \theta (y_{t-1}-y_{t-2})) y_{t-s} $ and thus reduce $d$.}
The next assumptions impose some regularity conditions on $g$. These restrictions are standard in the literature and can be somewhat relaxed (e.g. see DIN_JOE03 and references therein).
For instance, for the case where $g = \rho \otimes q^{K(n)} $ (for simplicity, let $J=1$) it suffices to assume that $E_{\mathbf{P}}[\rho(Y,\theta_{0})^{2} | W ]$ and that the eigenvalues of $E_{\mathbf{P}}[q^{K(n)}(W,\theta_{0})q^{K(n)}(W,\theta_{0})^{T}]$ are both bounded bounded and bounded away from zero a.s.-$\mathbf{P}$ These assumptions are standard; see DIN_JOE03 for a discussion. \footnote{These assumptions are also standard in the context of series-based estimators; see Chen20075549.}
Let $\mathcal{N}$ be an open neighborhood of $\theta_{0}$.
For instance, for the case $g = \rho \otimes q^{K}$ for many basis functions such as splines and Fourier series it holds that $\sup_{w} ||q^{K}(w)||_{2} \precsim \sqrt{K}$.\footnote{Other series like power series typically present $\sup_{w}||q^{K}(w)||_{2} \precsim K$, or more generally one can think of $\sup_{w} ||q^{K}(w)||_{2} \precsim \zeta(K)$ for some function $\zeta$. These cases can be accommodated in our theory, at the expense of further restricting the rate of growth of $d(n)$. } Thus, the previous assumption holds provided that $E[\sup_{\theta \in \mathcal{N}} ||\rho(Y,\theta)||_{2}^{2(2+\gamma)}|W]$ and $E_{\mathbf{P}}[\sup_{\theta \in \mathcal{N}}||\nabla_{\theta} \rho(Y,\theta)||^{2\beta}_{2} | W]$ are bounded by a constant $C$, and $||\nabla_{\theta} \rho(Y,\theta) - \nabla_{\theta} \rho(Y,\theta_{0})||_{2} \precsim \delta(Y) ||\theta - \theta_{0}||_{2}$ with $E_{\mathbf{P}}[\delta(Y)^{2}|W] \leq C$, a.s.-$\mathbf{P}$, for some $C>0$,.\footnote{These restrictions are analogous to Assumptions 4-6 in DIN_JOE03.}
The GMM estimator is given by $\hat{\theta}_{GMM,n} = \arg\min_{\theta \in \Theta} \hat{Q}_{GMM,n}(\theta) $ where
with $\hat{W}_{n} \in \mathbb{R}^{d \times d}$ is a (possibly random) positive definite matrix. The following mild condition is required
The bootstrap analog is given by $\hat{\theta}^{\ast}_{GMM,n} = \arg\min_{\theta \in \Theta} \hat{Q}^{\ast}_{GMM,n}(\theta)$ where
These formulas give raise to the following MST statistic: $\hat{T}_{GMM,n} \equiv n \hat{Q}_{GMM,n}(\hat{\theta}_{GMM,n})$ and its bootstrap version $\hat{T}^{\ast}_{GMM,n} \equiv n \hat{Q}^{\ast}_{GMM,n}(\hat{\theta}^{\ast}_{GMM,n})$.
In order to simplify the exposition we directly impose that $(\omega_{i,n})_{i \leq n}$ satisfy Assumption (ref) and also that they are uniformly bounded; this last assumption is not necessary for the results but imposing it greatly simplifies the technical derivations in our proofs.
It is worth to point out that DIN only considers GMM estimators with $W = \Omega^{-1}$ because they rely on CLT-type approximations for inference (e.g., see their Theorem 6.3). Since our result allow us to focus on bootstrap-based inference, the weighting matrix $W$ does not need to coincide with $\Omega^{-1}$; in fact it can simply be chosen as $\hat{W} = W = I$. That is, our results provide valid asymptotic inference for MST statistics for a larger class of GMM estimator, one with $W \ne \Omega^{-1}$.
The GEL estimator is given by
where $s : \mathcal{V} \subseteq \mathbb{R} \mapsto \mathbb{R}$ is concave and twice continuously differentiable with Lipschitz second derivative, $\mathcal{V}$ includes a neighborhood of 0, and $\Lambda(\theta) \equiv \{ \lambda \in \mathbb{R}^{d} : \lambda^{T}g(X,\theta) \in \mathcal{V},~a.s.-\mathbf{P} \}$. The function $s$ can be chosen to encompass several estimators of interest such as empirical likelihood ($s(\cdot) = \ln(1-\cdot)$), exponential tilting ($s(\cdot) = -\exp(\cdot)$; ISJ_ECMA98 and KS_ECMA97) and continuously updating GMM ($s(\cdot) = -0.5(1 + \cdot)^{2}$; HHY_JBES96). Henceforth, to simplify the presentation we assume the following normalization $s'(0)=s''(0)=-1$.
Analogously to GMM, we have the following MST statistic for GEL: $\hat{T}_{GEL,n} \equiv 2 \left\{ \hat{Q}_{GEL,n}(\hat{\theta}_{GEL,n}) - ns(0) \right\}$ and its bootstrap version $\hat{T}^{\ast}_{GEL,n} \equiv 2 \left\{ \hat{Q}^{\ast}_{GEL,n}(\hat{\theta}^{\ast}_{GEL,n}) - n s(0) \right\}$, where $\hat{\theta}^{\ast}_{GEL,n} = \arg\min_{\theta \in \Theta} \hat{Q}^{\ast}_{GEL,n}(\theta)$ and $\hat{Q}^{\ast}_{GEL,n}$ is defined as $\hat{Q}_{GEL,n}$ but with $\omega_{i,n} g(x_{i},\cdot)$ instead of $g(x_{i},\cdot)$.\footnote{Abusing notation we still denote $\Lambda(\theta)$ as the set for the bootstrap case.}
The next assumption is a high level condition. Part (i) ensures existence of a minimizer for $\lambda$ and part (ii) imposes convergence rates on the GMM and GEL estimators. Because our main goal is to establish the asymptotic behavior of the MST statistics, we directly impose this assumption to ease the exposition.
The derivation of both parts of this assumption from more primitive conditions can be obtained from the results in DIN and references therein; in particular in Lemma A.10 and Theorems 5.4 and 5.6.
The following lemma establishes that the test statistics for both estimators are asymptotically equivalent to a quadratic form on sample averages of $g$.
This Theorem establishes that the test statistics, asymptotically, behave as quadratic forms of (properly scaled) sample averages. Thus, our result in Theorem (ref) can be applied to these cases with $Z_{i} \equiv g(X_{i},\theta_{0}) W^{1/2}$ or $Z_{i} \equiv g(X_{i},\theta_{0}) \Omega^{-1/2}$ . The next Theorem formalizes this claim in this particular setting.
This result allow us to compute bootstrap-based p-values for the MST statistics for the general classes of GMM and GEL estimators, even when the number of moment restrictions increases with the sample size (but not too fast). In particular, for $\gamma \geq 2$, our condition on rate imposes that $d(n)^{4}/n = o(1)$ which is the one required in Theorem 6.4 in DIN, but at the cost of imposing restrictions on some higher moments of $||g(\cdot,\theta_{0})||_{2}$ (see Assumption (ref)(i)).
In this section we present a Monte Carlo (MC) study to assess the finite sample behavior of our procedure. We perform $5000$ MC repetitions and in each draw we perform $5000$ bootstrap repetitions.
The design is as follows: In each MC repetition we draw $Z_{i} = V^{1/2} \sqrt{12} U_{i}$ with $U_{i} \sim U(-0.5,0.5)$ for $i=1,...,n$, and $V$ is a positive definite symmetric matrix specified below. Let
and the associated bootstrapped version is given by
Throughout the study we use $\omega \sim N(0,1)$.
We are interested in studying $\mathbb{K}^{B}_{n} = \sup_{a \in \mathbb{A}} \left| \mathbf{P}_{n} \left( Q_{n} \geq t^{B}_{n}(a,Z^{n} ) \right) - (1-a) \right|$ and, for comparison, $\mathbb{K}_{n} = \sup_{a \in \mathbb{A}} \left|\mathbf{P}_{n} \left( Q_{n} \geq t_{n}(a) \right) - (1-a) \right|$, where $t^{B}_{n}(a,Z^{n} ) $ is the $a$-th empirical percentile of $Q^{\ast}_{n}$ and $t_{n}(a)$ is the $a$-th percentile of a chi-square with degrees of freedom $d(n)$.\footnote{In both cases, we approximate $\mathbf{P}_{n}$ using the empirical cdf across MC repetitions.} \footnote{$\mathbb{K}^{B}_{n}$ is in fact the quantity of interest since, by construction, $1-a$ coincides with the empirical quantile of $Q^{\ast}_{n}$, thus $\mathbb{K}^{B}_{n}$ approximates $\sup_{a \in \mathbb{A}} \left| \mathbf{P}_{n} \left( Q_{n} \geq t^{B}_{n}(a,Z^{n} ) \right) - \mathbf{P}^{\ast}_{n}( Q^{\ast}_{n} \geq t^{B}_{n}(a,Z^{n} ) \mid Z^{n} ) \right| $. A similar observation holds for $\mathbb{K}_{n}$. } The set $\mathbb{A}$ is given by $\{ 0.900,0.950,0.975,0.990 \}$. The typical application for our results is testing --- like in the Section (ref) ---, and with this in mind $\mathbb{A}$ is designed to capture the relevant values of $a$ for which we would like to assess the performance of the approximation.
Approximation Error. Figure (ref) shows the $\log (\mathbb{K}_{n}/\mathbb{K}^{B}_{n})$ for different values of the weighting matrix $V$ and for $n=500$ and $d(n) = 3$. When $V=I$ both, the chi-squared-based and boostrap-based procedures yield correct approximations of the limiting distribution, and thus the value is close to one. As expected, for cases where $V = (1+\epsilon/\sqrt{n}) I$ with $\epsilon \ne 0$, the chi-squared-based approximation does not approximate the limiting distribution, whereas the bootstrap-based continues to do so. The simulations shows that even for small values of $\epsilon/\sqrt{n}$, the difference is non-negligible. We note also that the deviations from $V=I$ we consider are \textquotedblleft mild" and we expect that for more complex deviations the results will be even more stark.
Table (ref) shows the value of $100 \times \left| \mathbf{P}_{n} \left( Q_{n} \geq t^{B}_{n}(a,Z^{n} ) \right) - (1-a) \right|$ for each $a \in \mathbb{A}$. We can see that regardless of the value of $V$, the approximation error of our bootstrap procedure remains stable at low values, below 0.5%.
Table (ref) shows $\mathbb{K}^{B}_{n}/\mathbb{K}_{n}$ for $d(n) = n^{1/5}$. We see that for all $n$ under consideration the ratio is around one, and in almost all below one. These results suggest that, at least for the current design, the convergence rate of the bootstrap-based approximation is no worse than the one for the chi-squared-based.
Robustness to $d(n)$ and choice of weights. We now assess how robust our procedure is to the choice of $d(n)$. Recall that, for this specification, our theory predicts that is sufficient to have $d(n) = o(n^{1/4})$; for values higher than this our theory is silent about the validity of our bootstrap procedure. We are thus particularly interested on the performance of our procedure for the latter set of values. In this exercise, we set $V=I$ and consider different values of $n$ and $d(n)$.
Table (ref) columns 2-5 shows the value of $100 \times \mathbb{K}^{B}_{n}$ for different choices of $d(n)$ and $n$. For values of $n$ less than 1000, the procedure seems to be quite robust to larger choices of $d(n)$ in the range of $n^{1/4}$ to $n^{1/2}$, but not higher. For values of $n$ around 2000-3000, however, our procedure seems to deteriorate for values of $d(n)$ larger than $n^{1/2}$.
We now assess the robustness of our procedure to different choices of weights. We compare the Gaussian weights with two other weights: $\omega_{i} \sim U(-0.5,0.5)$ and $\omega_{i} \sim t-Student(3)$ (properly scaled to have unit variance). These choices are designed to study how different tail behavior of the weight's distribution affect the performance of our bootstrap procedure.
In order to ease the computational burden we lower the bootstrap repetitions to 2000 each. Table (ref) presents the results. The overall pattern seems to suggest that the Gaussian and Uniform weights have comparable performances, and perform better than the t-Student weights. This pattern illustrates the discussion in Section (ref) regarding desirable properties of weights.
Remarks. Overall, the simulations suggest that our procedure has a finite sample performance that is at least as good as, and in some cases better than, the \textquotedblleft standard" chi-squared approach. Weights with \textquotedblleft thin tails" such as Uniform and Gaussian seem to perform better than weights with heavier tails. Additionally, as also discussed in the context of our application in Section (ref), our bootstrap-based approximation can be applied in situations that go beyond those covered by the chi-square approach.
Recall that $x \in \mathbb{R}^{d(n)} \mapsto ||x||^{2}_{2} \equiv x^{T}x$ and that $\mathcal{C}_{M}$ is the class of functions $f : \mathbb{R} \rightarrow \mathbb{R}$ that are three times continuously differentiable and $ \sup_{x} | \partial^{r} f(x)| \leq (M)^{r}$.
All the proofs of the lemmas in this section are relegated to Appendix (ref).
For any two probability measures $Q$ and $P$, let
We want to establish the following: For any $\varepsilon'>0$, there exists a $N(\varepsilon')$ such that
for all $n \geq N(\varepsilon')$. Observe that
where $S_{n} \equiv \{ Z^{n} : n^{-1} \sum_{i=1}^{n} ||Z_{i}||^{2}_{2} \leq (0.5 \varepsilon')^{-1} tr \{ \Sigma_{n} \} \}$. By the Markov inequality $\mathbf{P}_{n} \left( S_{n}^{C} \right) \leq 0.5 \varepsilon'$. Thus, it suffices to show that
By the triangle inequality, for all $t \in \mathbb{R}$ and $Z^{n}$
where $\sqrt{n} \mathbb{V}_{n} \sim N(0,\Sigma_{n})$. We use $\boldsymbol{\Phi}_{n}$ to denote this probability.
Therefore, in order to obtain display (ref), it suffices to bound
and
The next two lemmas allow us to \textquotedblleft replace" the indicator functions by \textquotedblleft smooth" functions.
(Recall that, $\Delta_{h^{-1}}(\mathbf{P}_{n},\boldsymbol{\Phi}_{n}) = \sup_{f \in \mathcal{C}_{h^{-1}}} \left| E_{\mathbf{P}_{n}} \left[ f \left( ||\sqrt{n} \mathbb{Z}_{n}||^{2}_{2} \right) \right] - E_{\boldsymbol{\Phi}_{n}} \left[ f \left(||\sqrt{n} \mathbb{V}_{n}||^{2}_{2} \right) \right] \right|$). And
Therefore, by letting $\varepsilon$ in the lemmas be such that $\frac{\varepsilon}{1-\varepsilon} + 3 \varepsilon = 0.25 \varepsilon'$ we obtain
and
for all $n \geq N(\varepsilon)$ and all $h \leq h(\varepsilon,\delta_{n})$ (note that $\varepsilon$ is a function of $\varepsilon'$).
By the triangle inequality and straightforward algebra, it follows that
where $\boldsymbol{\Phi}^{\ast}_{n}(\cdot|Z^{n})$ denotes the conditional probability (given the original data $Z^{n}$ ) $\sqrt{n} \mathbb{U}_{n} \sim N(0,n^{-1} \sum_{i=1}^{n} Z_{i} Z^{T}_{i})$
Hence, by the previous display and Equations (ref), (ref)-(ref), (ref) and (ref), in order to show the desired result it suffices to show that: For all $\varepsilon'$, there exists a $N(\varepsilon')$ such that
for all $n \geq N(\varepsilon')$ and some $h \leq h(\varepsilon,\sqrt{tr\{ \Sigma^{2}_{n} \}} \gamma(\varepsilon) )$. Theorems (ref) and (ref) establish expressions (ref) and (ref).
We have thus reduced the original problem to a Gaussian approximation problem. That is, it remains to show that
Since $\sqrt{n} \mathbb{U}_{n} \sim N(0,\hat{\Sigma}_{n})$ (with $\hat{\Sigma}_{n} = n^{-1} \sum_{i=1}^{n} Z_{i}Z_{i}^{T}$) and $\sqrt{n} \mathbb{V}_{n} \sim N(0,\Sigma_{n})$, the previous display is equivalent to showing that
Essentially, this expression follows by the fact that $\hat{\Sigma}_{n}$ converges in probability to $\Sigma_{n}$ in a suitable norm. The following lemma formalizes this.
Observe that for any $ Z^{n} \in S_{n} = \{ Z^{n} : n^{-1} \sum_{i=1}^{n} ||Z_{i}||^{2}_{2} \leq (0.5 \varepsilon')^{-1} tr \{ \Sigma_{n} \} \}$, the RHS of the expression in the Lemma is bounded above by
Thus by Lemma (ref), in order to establish the desired result, it suffices to show that
for sufficiently large $n$. Henceforth, let $c_{n} \equiv \frac{ (\varepsilon')^{2}}{ d(n) h^{-2} tr\{ \Sigma_{n} \} } $ and let $\mathbf{A}_{i,n}[j,l] \equiv Z_{[j],i} Z_{[l],i} $, observe that
Let $\mathbf{A}_{i,n}[j,l] = \mathbf{A}^{L}_{i,n}[j,l] + \mathbf{A}^{U}_{i,n}[j,l] \equiv \mathbf{A}_{i,n}[j,l] 1\{ |\mathbf{A}_{i,n}[j,l]| \leq e_{n} \} + \mathbf{A}_{i,n}[j,l] 1\{ |\mathbf{A}_{i,n}[j,l]| \geq e_{n} \}$ where $(e_{n})_{n}$ with $e_{n} > 0$ is defined below. Clearly, $\mathbf{A}^{L}_{i,n}[j,l] \leq e_{n}$. So, by Hoeffding inequality (see Boucheron-book13 p. 34)
Therefore, by setting $e_{n} = c_{n} \sqrt{\frac{n 0.25}{\log(d(n))}} $, the previous display implies that
for sufficiently large $n$.
Second, by the Markov inequality and the fact that
for all $i\ne k$, it follows that
Therefore by the Markov inequality, for $p > 0$
Since $e_{n} = c_{n} \sqrt{\frac{n 0.25}{\log(d(n))}} $ and $c_{n} \equiv \frac{ (\varepsilon')^{2}}{ d(n) h^{-2} tr\{ \Sigma_{n} \} } $, it follows that
Since we can set $h \asymp \sqrt{tr\{ \Sigma^{2}_{n} \}}$, the RHS becomes\\ $\frac{(\log(d(n)))^{p/2} d(n)^{2+p} }{n^{1+p/2} } \left( \frac{ tr\{ \Sigma_{n} \} }{ tr\{ \Sigma_{n}^{2} \} } \right)^{2+p} E_{\mathbf{P}_{n}} \left[ \left( \sum_{j=1}^{d(n)} ( Z_{[j],1})^{2+p} \right)^{2} \right]$. By choosing $p = \kappa$, by Assumptions (ref)(i) and (ref)(iii), the term vanishes as $n \rightarrow \infty$.
Therefore, Equation (ref) is established and with that the proof of Theorem (ref).
Applicability of our Results. The example developed in Section (ref) illustrates a general feature present in several test statistics, namely that they behave asymptotically as quadratic forms of (properly scaled) sample averages. These are the main motivational examples to which we can apply our result in Theorem (ref).
This remark is best illustrated in the Wald statistic case. To formalize this, consider i.i.d. data $(X_{1},...,X_{n})$ drawn from $\mathbf{P} $ and parameter a $k \geq d=d(n)$ dimensional $\theta_{\mathbf{P}}$ and a \textquotedblleft smooth" function $\theta \mapsto c(\theta) \in \mathbb{R}^{d}$ which represent the hypothesis we want to test; i.e., the null hypothesis is $c(\theta_{\mathbf{P}}) = 0$. \footnote{The notation $\theta_{\mathbf{P}}$ stresses that the parameter is a (known) function of the probability distribution. Thus, an estimator can obtained by \textquotedblleft plugging in" the empirical distribution $P_{n}$. } The fact that $d$ grows with the sample size is of potential interest because in certain situations one could have that the dimension of the parameter, although fixed, is not \textquotedblleft small" relative to $n$. Also, in some other situations, one could have a more explicit model of increasing dimensionality like in the cases discussed in Section (ref) or in series or sieves estimators; see, for example, ChenPouzo14.
Suppose there exists an estimator $\theta_{P_{n}}$ ($P_{n}$ is the empirical distribution), then the Wald statistic is given by
where $V_{n} \in \mathbb{R}^{d \times d}$ is some (possibly random) matrix to be determined later.
Suppose $\theta_{P_{n}}$ admits an asymptotic linear representation (ALR) of the form \footnote{See VdV00 and references therein for a discussion regarding ALR and sufficient conditions for it. Here we follow VdV-Murphy00. }
with $E_{\mathbf{P}}[\psi(X,\theta_{0})] = 0$ and finite second moment. \footnote{Note that $E_{P_{n}}[ \psi(X,\theta_{\mathbf{P}}) ] = n^{-1} \sum_{i=1}^{n} \psi(X_{i},\theta_{\mathbf{P}}) $.}.
The bootstrap analog of the Wald statistic are of the \textquotedblleft plug-in" type, i.e., \footnote{In principle, one could also \textquotedblleft replace" $V_{n}$ --- which typically is a function of $P_{n}$, $V_{P_{n}}$ --- by $V_{P^{\ast}_{n}}$. Our results could be extended to this case too.}
where $P^{\ast}_{n}$ is given by $n^{-1} \sum_{i=1}^{n} \omega_{i,n} \delta_{X_{i}}$ (the dependence of $P_{n}^{\ast}$ on $X^{n}$ is omitted to ease the notational burden). The bootstrap ALR (B-ALR) is given by
Given the asymptotic linear representations, we can show that the Wald and Bootstrapped Wald statistics can be represented asymptotically as quadratic forms, and thus fall in the framework studied in this paper. The following proposition formalizes such representation, and thereby allow us to apply our Theorem (ref) with $Z_{i} \equiv \psi(X_{i},\theta_{\mathbf{P}}) V^{1/2}$ to approximate the limiting distribution of $\mathbb{W}_{n}(P_{n},\mathbf{P})$.
A few remarks are in order. First, and more importantly, we note that, our results can be applied to other test statistic provided that are asymptotically equivalent (up to $o(\sqrt{d(n)})$) to $\mathbb{W}(P_{n},\mathbf{P})$ or to a quadratic form as in the proposition. Typically this is the case for the Likelihood ratio and Lagrange Multiplier (or Score) test statistics; see Newey19942111 Section 9.
Second, for the Chi-square-based approximation to be valid, $V$ must coincide with $(E_{\mathbf{P}}[\psi (X ,\theta_{0} ) \psi (X ,\theta_{0} )^{T} ])^{-1}$. The bootstrap-based approximation, however, does not require this assumption. This situation may arise, for instance, in Likelihood ratio tests under model misspecification.
Choice of Weights. We now provide some heuristic discussion regarding the weights.
The bootstrap procedure studied in this paper uses independent weights. Such restriction has also been used in several papers; e.g. CCK-AOS13 and Ma2005190. This choice is largely due to the fact that the independent behavior of weights makes many of the proofs easier. It would be of interest still to extend our results to non-iid weights such as Multinomial weights --- which yield non-parametric and m-out-n bootstrap procedures. While such an extension is beyond the scope of the paper, we point out that the key step in order to do this is to extend Theorem (ref) and Lemma (ref) (in the Appendix) to allow for non-independent data;\footnote{At least for one sequence, either the $A$'s or $B$'s in the Theorem. Independence is, due to the technique of proof, particularly important for establishing Lemma (ref).}
Even within the class of independent weights, one could wonder what properties are desirable for the weights to have. Clearly, as indicated by our Assumption (ref), restrictions on the tail behavior of the weights are important for our results. We now present a discussion, which expands on the quantitative explorations in Section (ref), about what other properties might be desirable to have.
Heuristically, the Lindeberg interpolation result --- Theorem (ref) --- relies on \textquotedblleft matching" the first and second moments. By choosing the weights to match higher moment one could expect to improve the approximation rates.\footnote{These observations are related to the four moment Theorem of Tao and Vu in the context of random matrices; see TaoVu11.} More precisely, the bounds for $\mathbf{S}_{n}$ (and $\mathbf{R}_{n}$) obtained in Lemma (ref) in the Appendix only use restrictions impose the restrictions in the original data and the bootstrap weights present in Assumptions (ref)(i)(ii) and (ref). However, it is easy to see that if one would have additional information on the higher moments, one could obtain sharper bounds for $\mathbf{S}_{n}$. For instance, to show Theorem (ref), we apply Theorem (ref) with $A_{i} = n^{-1/2} \omega_{i,n} Z_{i} $ and $B_{i} = n^{-1/2} u_{i} Z_{i}$ with $u_{i} \sim N(0,1)$. If we would have that $(\omega_{i,n})_{i=1}^{n}$ were such that $E[|\omega_{i,n}|^{4}]=E[(Z)^{4}]$ with $z \sim N(0,1)$, then $\mathbf{S}_{1,n} = 0$. A similar observation applies to $\mathbf{S}_{2,n}$ but in this case the relevant moments are $E[\omega_{i,n}^{3}]$ and $E [Z^{3}]$.\footnote{The rate of convergence of the term $\mathbf{R}_{n}$ is regulated by $q$, which is link to the bound on higher moments of the data and weights (see Assumptions (ref) and (ref)).}
Finally, another extension that is linked to the previous discussion, is that of refinements (or lack thereof) of certain choices of bootstrap weights. We leave this for future research.