Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Inference on Functionals under First Order Degeneracy
bibunit\pdfbookmark[1]{Title}{title}
\begin{abstract}
This paper presents a unified second order asymptotic framework for conducting inference on parameters of the form $\phi(\theta_0)$, where $\theta_0$ is unknown but can be estimated by $\hat\theta_n$, and $\phi$ is a known map that admits null first order derivative at $\theta_0$. For a large number of examples in the literature, the second order Delta method reveals a nondegenerate weak limit for the plug-in estimator $\phi(\hat\theta_n)$. We show, however, that the “standard” bootstrap is consistent if and only if the second order derivative $\phi_{\theta_0}''=0$ under regularity conditions, i.e., the standard bootstrap is inconsistent if $\phi_{\theta_0}''\neq 0$, and provides degenerate limits unhelpful for inference otherwise. We thus identify a source of bootstrap failures distinct from that in Fang_Santos2014HDD because the problem (of consistently bootstrapping a {\it nondegenerate} limit) persists even if $\phi$ is differentiable. We show that the correction procedure in Babu1984bootstrap can be extended to our general setup. Alternatively, a modified bootstrap is proposed when the map is {\it in addition} second order nondifferentiable. Both are shown to provide local size control under some conditions. As an illustration, we develop a test of common conditional heteroskedastic (CH) features, a setting with both degeneracy and nondifferentiability -- the latter is because the Jacobian matrix is degenerate at zero and we allow the existence of multiple common CH features.
\end{abstract}
\begin{center}
Keywords: First order degeneracy, Second order Delta method, Bootstrap consistency, Babu correction, Common CH features, $J$-test.
\end{center}
{JEL Classification: C12, C15}
\section{Introduction}
There is a large number of inference problems in economics and statistics in which the parameter of interest is of the form $\phi(\theta_0)$, where $\theta_0$ is an unknown parameter depending on the underlying distribution of the data and $\phi$ is a known map. In these settings, it is common practice to employ the plug-in estimator $\phi(\hat\theta_n)$, where $\hat\theta_n$ is an estimator for $\theta_0$, as a building block for conducting inference on $\phi(\theta_0)$. The Delta method asserts that if $r_n\{\hat\theta_n-\theta_0\}\xrightarrow{L}\mathbb G$ for some sequence $r_n\uparrow\infty$, then
\begin{align}
r_n\{\phi(\hat\theta_n)-\phi(\theta_0)\}\xrightarrow{L}\phi_{\theta_0}'(\mathbb G) ,
\end{align}
provided $\phi$ is at least Hadamard directionally differentiable at $\theta_0$, where $\phi_{\theta_0}'$ is the derivative of $\phi$ at $\theta_0$ Shapiro1991, Dumbgen1993. As powerful as the Delta method has proven to be Vaart1998,Fang_Santos2014HDD, an implicit and yet crucial assumption for the convergence (ref) to be useful for inferential purposes is that $\phi_{\theta_0}'(\mathbb G)$ or $\phi_{\theta_0}'$ is nondegenerate, i.e., $\phi_{\theta_0}'\neq 0$. Unfortunately, such {\it first order degeneracy} arises frequently in asymptotic analysis, with applications including Wald tests or Wald type functionals Wald1943tests,Engle1984Handbook, unconditional and conditional moment inequality models AndrewsandSoares2010,Andrews_Shi2013CMI, Cram\'{e}r-von Mises functionals Darling1957KSCvM, the study of stochastic dominance Linton2010, and the $J$-test for overidentification in GMM settings Hall_Horowitz1996bootstrap.
In the presence of first order degeneracy, one may resort to a higher order analysis for the sake of a nondegenerate limiting distribution. Shapiro2000inference established that if $\phi$ is second order Hadamard directionally differentiable (see Definition (ref)) -- a feature shared by aforementioned examples, then
\begin{align}
r_n^2\{\phi(\hat\theta_n)-\phi(\theta_0)-\phi_{\theta_0}'(\hat\theta_n-\theta_0)\}\xrightarrow{L}\phi_{\theta_0}”(\mathbb G) ,
\end{align}
where $\phi_{\theta_0}''$ denotes the second order derivative of $\phi$ at $\theta_0$. Thus, when first order degeneracy occurs, (ref) suggests that we may base our asymptotic analysis on
\begin{align}
r_n^2\{\phi(\hat\theta_n)-\phi(\theta_0)\}\xrightarrow{L}\phi_{\theta_0}”(\mathbb G) .
\end{align}
Usefulness of the limiting distribution in (ref), however, relies on our ability to consistently estimate it. In this regard, Efron1979's bootstrap seems to be a potential option. Specifically, if $\hat\theta_n^*$ is a bootstrap analog of $\hat\theta_n$ that works for estimating the law of $\mathbb G$, then in view of (ref) one may hope that
\begin{align}
r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}
\end{align}
can be employed as an estimator for the law of $\phi_{\theta_0}''(\mathbb G)$, at least when $\phi$ is smooth. Unfortunately, there are simple examples where the law of (ref) conditional on the data, referred to as the standard bootstrap, fails to provide consistent estimates Babu1984bootstrap.
As the first contribution of this paper, we show that the standard bootstrap (ref) is consistent if and only if $\phi_{\theta_0}''=0$ under mild conditions. Thus, the standard bootstrap is necessarily inconsistent when $\phi_{\theta_0}''$ is nondegenerate, while when $\phi_{\theta_0}''$ is degenerate, the resulting asymptotic distribution is degenerate and hence not useful for inference. Therefore, the failure of the standard bootstrap is an inherent implication of first order degeneracy. It is worth noting that the failure of the standard bootstrap persists even when $\phi$ is differentiable. Hence, we identify a source of bootstrap inconsistency distinct from that in Fang_Santos2014HDD, i.e., nondifferentiability of the map $\phi$, as explained further towards the end of this section.
Heuristically, the reason why the standard bootstrap fails is that even though $r_n^2\phi_{\theta_0}'(\hat\theta_n-\theta_0)=0$ in the “real world”, its bootstrap counterpart is nondegenerate, i.e., $r_n^2\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)=O_p(1)$, echoing Efron1979's point that the bootstrap provides approximate frequency statements rather than approximate likelihood statements. This observation was picked up by Babu1984bootstrap who provided a consistent resampling procedure by including the first order correction term:
\begin{align}
r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)-\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)\} .
\end{align}
As the second contribution, we generalize the above modified bootstrap (ref), referred to as the Babu correction, to settings that accommodate infinite dimensional models and a wide range of bootstrap schemes for $\hat\theta_n^*$. However, we stress that the Babu correction is inappropriate when $\phi$ is only Hadamard directionally differentiable.
As the third contribution, we follow Fang_Santos2014HDD and provide a modified bootstrap which is consistent regardless of the presence of first order degeneracy and nondifferentiability of $\phi$. The insight we exploit is that the weak limit $\phi_{\theta_0}''(\mathbb G)$ in (ref) is a composition of the limit $\mathbb G$ and the derivative $\phi_{\theta_0}''$. Therefore, we may estimate the law of $\phi_{\theta_0}''(\mathbb G)$ by composing a suitable estimator $\hat\phi_n''$ for $\phi_{\theta_0}''$ with a bootstrap approximation $r_{n}\{\hat\theta_n^*-\hat\theta_n\}$ for $\mathbb G$. Since the conditions on $\hat\phi_n''$ proposed by Fang_Santos2014HDD in order for this approach to work are either demanding or hard to check in our setup, we provide a high level condition that is easy to verify. We further demonstrate that numerical differentiation provides a desirable estimator $\hat\phi_n''$ in general; alternatively, we show how to estimate $\phi_{\theta_0}''$ by exploiting its structure in particular examples. Our inference procedures are also shown to enjoy the local size control property under a key condition that is algebraically simple.
Finally, to further demonstrate the applicability of our framework, we develop a test of common conditional heteroskedastic (CH) features studied by Dovonon_Renault2013testing but under weaker assumptions that allow more than one common CH features. Thus, in addition to the first order identification failure they focused on, we further allow second order (and hence global) identification failures, which renders the functional involved highly (second-order) nondifferentiable as well as first order degenerate. Such a generalization is important because it is unknown {\it a priori} how many common features there are and in the context of asset pricing the number can be large Engle_Ng_Rothschild1990asset. Moreover, the linear normalization in Dovonon_Renault2013testing can falsely exclude the existence of common features even when there does exist a unique common CH feature, a deficiency which we avoid by the unit-length normalization. Monte Carlo simulations indicate our tests substantially alleviate size distortion and have good power performance. We stress that first order degeneracy is of a nature different from that of the degeneracy of Jacobian matrices which is the focus of Dovonon_Renault2013testing; see Section (ref) for details. Our approach may also be used to develop tests for other common features Engle_Kozicki1993CF.
There have been extensive studies on the bootstrap consistency Hall1992bootstrap,HorowitzBoot. It was realized soon after Efron1979 that the bootstrap is not always successful BickelandFreedman1981bootstrap; see also Andrews2000Bootstrap for a summary. Babu1984bootstrap provided a simple example of bootstrap failure due to first order degeneracy, and established the validity of the Babu correction for the special case studied there. Shao1994bootstrap and Bertail_Politis_Romano1999subsampling showed that $m$ out of $n$ resampling and subsampling can serve as alternative remedies. There are, however, three reasons we choose not to use these methods. First, they entail the choice of tuning parameters while our proposal can work without such nuisances when $\phi$ is differentiable. Second, when $\phi$ is nondifferentiable, both can lead to invalid tests due to lack of uniform approximations AndrewsandGuggen2010ET. We provide a simple algebraic condition which, together with regularity of $\hat\theta_n$, delivers local uniformity of our inferential procedure. Third, they have been shown to be dominated by other inferential methods, for example, in moment inequality models AndrewsandSoares2010 which our framework includes as special cases. Datta1995bootstrap revisited Babu's example and offered a bias correction procedure that depends on a first stage shrinkage type estimator. Somewhat similar methods were later proposed in Andrews2000Bootstrap and Giurcanu2012bootsrtap. These methods are not easily extendable to more general settings.
Bootstrap inconsistency due to nondifferentiability of $\phi$ was studied in Dumbgen1993 and recently in Fang_Santos2014HDD who formally established that (first order) differentiability of $\phi$ is a necessary as well as sufficient condition for the standard bootstrap to work under regularity conditions. Our work complements theirs by identifying a different source of bootstrap failure. Specifically, given bootstrap consistency of $\hat\theta_n^*$ and if $\phi$ is first order degenerate (and hence fully differentiable!), then Fang_Santos2014HDD implies that the standard bootstrap $r_n\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}$ is consistent for the law of $\phi_{\theta_0}'(\mathbb G)$ which is degenerate (and unhelpful for inference). We further show that the law of the second order limit $\phi_{\theta_0}''(\mathbb G)$ cannot be consistently estimated by the second order standard bootstrap (ref) unless $\phi_{\theta_0}''$ itself is degenerate -- {\it this remains true regardless of whether $\phi$ is (second order) differentiable or not!} Moreover, extra work is needed in order to show our bootstrap inferential procedures work well in the local uniformity sense. In applications, first order degeneracy and second order nondifferentiability are often mixed together, for example, in Romano_Shaikh2010, AndrewsandSoares2010, Linton2010, and Andrews_Shi2013CMI. The numerical differentiation approach of estimating derivatives was somewhat implicit in Dumbgen1993's rescaled bootstrap, recently employed by Song2014minimax and studied by Hong_Li2015numericaldelta. We provide a more general condition that may be used to verify “consistency” of derivative estimators (not necessarily constructed via numerical differentiation). Our theory has been utilized in ChenFang2016Rank to develop a rank test where, unlike previous studies, the true rank is potentially strictly less than the hypothesized value, a longstanding problem in the literature.
We now introduce some notation. For a set $T$, we let $\ell^\infty(T)$ denote the space of bounded real-valued functions defined on $T$ and $C(T)$ the space of real-valued continuous functions on a compact set $T$ (endowed with some topology). Both $\ell^\infty(T)$ and $C(T)$ are equipped with the uniform norm, i.e., $\|f\|_\infty\equiv\sup_{t\in T}|f(t)|$. For a normed space $\mathbb D$ endowed with norm $\|\cdot\|_{\mathbb D}$ and $m\in\mathbf N$, we equip the product space $\prod_{j=1}^m \mathbb D$ with the product norm $\max_{j=1}^m\|\theta^{(j)}-\vartheta^{(j)}\|_{\mathbb D}$, denoted $\|\cdot\|_{\mathbb D}$ with some abuse of notation, for $\theta,\vartheta\in \prod_{j=1}^m \mathbb D$, where $\theta^{(j)}$ and $\vartheta^{(j)}$ are the $j$th coordinates of $\theta$ and $\vartheta$ respectively. For a subset $A\subset T$, we write $1\{A\}$ for the indicator function of $A$.
The remainder of the paper is structured as follows. Section (ref) formalizes the general setup, shows the wide applicability of our framework by introducing related examples, and establishes the asymptotic framework by presenting a mild extension of the second order Delta method. Section (ref) characterizes the inherent difficulties caused by first order degeneracy, extends the Babu correction to our general setup, and offers a flexible modified bootstrap procedure. Section (ref) develops a test for common CH features that allows multiple common CH features, while Section (ref) concludes. Appendix (ref) demonstrates that our inferential procedure is robust to local perturbations of the distribution of the data under regularity conditions. The remaining appendices collect all the proofs and additional discussions.
\section{Setup and Background}
In this section, we formalize the general setup, introduce related examples, and review notions of differentiability based on which we present the second order Delta method.
\subsection{General Setup}
The treatment in this paper is general in the sense that we allow both the parameter $\theta_0$ and the map $\phi$ to take values in infinite dimensional spaces, though attention is confined to real-valued $\phi$ when studying tests. In particular, we assume $\theta_0\in\mathbb D_\phi\subset\mathbb D$ and $\phi: \mathbb D_\phi\to\mathbb E$, where $\mathbb D$ and $\mathbb E$ are normed spaces with norms $\|\cdot\|_{\mathbb D}$ and $\|\cdot\|_{\mathbb E}$ respectively. Moreover, the data generating process is general as well in that the model can be parametric, semiparametric and nonparametric and that the data $\{X_i\}_{i=1}^n$ need not be i.i.d.. However, we do impose i.i.d.\ assumption in our local analysis, but only for simplicity. The results there can presumably be extended to general asymptotically normal experiments Vaart_Wellner1990prohorov.
The common probability space on which all (random) maps are defined is the canonical one. For example, in the simplest i.i.d.\ setup, we think of the data $\{X_i\}_{i=1}^n$ as the coordinate projections on the first $n$ coordinates in the product probability space $(\prod_{i=1}^\infty\mathscr X, \bigotimes_{i=1}^\infty \mathcal A, \prod_{i=1}^\infty P)$ where $(\mathscr X,\mathcal A)$ is the sample space each $X_i$ lives in and $P$ is the common Borel probability measure that governs each $X_i$. In the presence of bootstrap weights, we further think of the product space as the “first $\infty$” coordinates of the even “larger” product space $\big((\prod_{i=1}^\infty\mathscr X)\times \mathscr W, (\bigotimes_{i=1}^\infty \mathcal A) \otimes\mathcal W, (\prod_{i=1}^\infty P)\times Q\big)$, where $(\mathscr W,\mathcal W,Q)$ governs the infinite sequence of bootstrap weights.
Given the generality of our setup, weak convergence throughout the paper is meant in the Hoffmann-J{\o}rgensen sense Vaart1996. Expectations and probabilities should therefore be interpreted as outer expectations and outer probabilities respectively defined relative to the canonical probability space, though we obviate the distinction in the notation. The notation is made explicit in the appendices whenever differentiating between inner and outer expectations is necessary.
\subsection{Related Examples}
To fix ideas, we now turn to related examples that serve to illustrate the wide applicability of our framework. The first example is taken from Babu1984bootstrap, which provides an easy illustration of bootstrap inconsistency in the presence of first order degeneracy even if the transformation $\phi$ is smooth.
\begin{ex}[Wald Functional: Squared Mean]
Let $X\in\mathbf R$ be a random variable, and suppose that we are interested in conducting inference on
\begin{align}
\phi(\theta_0)=(E[X])^2 .
\end{align}
Here, $\theta_0=E[X]$, $\mathbb D=\mathbb E=\mathbf R$, and $\phi: \mathbf R\to\mathbf R$ is defined by $\phi(\theta)=\theta^2$. In fact, $\phi$ is a special case of the more general quadratic functionals of the form $\|W\theta\|^2$ for $\theta\in\mathbf R^k$ and $W$ a $k\times k$ weighting matrix. This seemingly toy example also arises in VAR models for inference on impulse responses Benkwitz_Neumann_Lutekpohl2000 and in some nonseparable models with structural measurement errors Hoderlein_Winter2010. \qed
\end{ex}
The second example is a special case of the unconditional moment inequality models studied in CHT2007, Romano_Shaikh2008,Romano_Shaikh2010, AndrewsandGuggen2009ET, and AndrewsandSoares2010.
\begin{ex}[Unconditional Moment Inequalities]
Let $X\in\mathbf R$ be a scalar random variable and suppose we want to test the moment inequality $E[X]\le 0$. The modified method of moments approach is based on estimating the functional
\begin{align}
\phi(\theta_0)=(\max\{\theta_0,0\})^2 ,
\end{align}
where $\theta_0=E[X]$, $\mathbb D=\mathbb E=\mathbf R$, and $\phi:\mathbf R\to\mathbf R$ is defined by $\phi(\theta)=(\max\{\theta,0\})^2$. The functional $\phi$ can be easily adapted to handle general moment inequality models.\qed
\end{ex}
The third example concerns the classical Cram\'{e}r-von Mises functional employed to test goodness of fit Darling1957KSCvM,Vaart1998.
\begin{ex}[Cram\'{e}r-von Mises Functional]
Suppose that we are interested in testing if the distribution function of a random vector $X\in\mathbf R^{d_x}$ is a given function $F_0$. The Cram\'{e}r-von Mises approach considers the functional
\[
\phi(\theta_0)=\int(F-F_0)^2\,dF_0~.
\]
Here, $\theta_0=F$, $\mathbb D=\ell^\infty(\mathbf R^{d_x})$, $\mathbb E=\mathbf R$, and $\phi: \ell^\infty(\mathbf R^{d_x})\to\mathbf R$ is defined to be $\phi(\theta)=\int(\theta-F_0)^2\,dF_0$. More generally, it is possible to test if $F$ belongs to a parametric family $\{F_\gamma: \gamma\in\Gamma\}$ by studying $\phi(\theta_0)=\inf_{\gamma\in\Gamma}\int(\theta_0-F_\gamma)^2\,dF_\gamma$. \qed
\end{ex}
The fourth example, closely related to but significantly different from Example (ref), is based on Linton2010 for testing stochastic dominance.
\begin{ex}[Stochastic Dominance]
Let $X = (X^{(1)},X^{(2)})^\intercal \in \mathbf R^2$ be continuously distributed, and define the marginal cdfs $F^{(j)}(u) \equiv P(X^{(j)} \leq u)$ for $j \in \{1,2\}$. For a weighting function $w:\mathbf R\rightarrow \mathbf R^+\equiv \{x\in\mathbf R: x\ge 0\}$, Linton2010 estimate
\begin{equation}
\phi(\theta_0) = \int_{\mathbf R} \max\{F^{(1)}(u) - F^{(2)}(u),0\}^2w(u)du ,
\end{equation}
to construct a test of whether $X^{(1)}$ first order stochastically dominates $X^{(2)}$. In this example, we set $\theta_0 = (F^{(1)},F^{(2)})$, $\mathbb D = \ell^\infty(\mathbf R)\times \ell^\infty(\mathbf R)$, $\mathbb E = \mathbf R$ and $\phi(\theta) = \int \max\{\theta^{(1)}(u) - \theta^{(2)}(u),0\}^2w(u)du$ for any $\theta\equiv(\theta^{(1)},\theta^{(2)}) \in \ell^\infty(\mathbf R)\times \ell^\infty(\mathbf R)$. We note that the Cram\'{e}r-von Mises type functionals in Andrews_Shi2013CMI,Andrews_Shi2014CMI shares the common structure of the functional $\phi$ in (ref) and hence can be taken care of by our framework as well.\qed
\end{ex}
The fifth example is a special case of the Kolmogorov-Smirnov type functionals for inference on conditional moment inequalities studied by Andrews_Shi2013CMI.
\begin{ex}[Conditional Moment Inequalities]
Let $Z\in\mathbf R^2$ and $W\in\mathbf R^{d_w}$ be random vectors satisfying $E[Z^{(1)}|W]\le 0$ and $E[Z^{(2)}|W]=0$. For a suitably chosen class of nonnegative functions $\mathcal F$ on $\mathbf R^{d_w}$, the above conditional moment inequality is equivalent to $E[Z^{(1)}f(W)]\le 0$ and $E[Z^{(2)}f(W)]=0$ for all $f\in\mathcal F$. Andrews_Shi2013CMI propose testing the above restriction by estimating the functional
\begin{align}
\phi(\theta_0)=\sup_{f\in \mathcal F}\{[\max(E[Z^{(1)}f(W)],0)]^2+(E[Z^{(2)}f(W)])^2\} .
\end{align}
Here, $\theta_0\in\ell^\infty(\mathcal F)\times \ell^\infty(\mathcal F)$ satisfies $\theta_0(f)=E[Zf(W)]$ for all $f\in\mathcal F$, $\mathbb D=\ell^\infty(\mathcal F)\times \ell^\infty(\mathcal F)$, $\mathbb E=\mathbf R$, and $\phi: \mathbb D\to\mathbb E$ is given by $\phi(\theta)=\sup_{f\in\mathcal F}\{[\max(\theta^{(1)}(f),0)]^2+[\theta^{(2)}(f)]^2\}$. \qed
\end{ex}
Our final example is concerned with the $J$-test of overidentification in GMM settings proposed by Sargan1958iv,Sargan1959IV and further developed in Hansen1982.
\begin{ex}[Overidentification Test]
Let $X\in\mathbf R^{d_x}$ be a random vector and consider the model defined by the moment restriction $E[g(X,\gamma_0)]=0$ for some $\gamma_0\in\Gamma\subset \mathbf R^k$ where $g: \mathbf R^{d_x}\times\Gamma\to\mathbf R^m$ is a known function with $m>k$. The conventional $J$-test can be recast by estimating the functional $\phi$ defined as: for some known $m\times m$ symmetric positive definite matrix $W$,
\begin{align}
\phi(\theta_0)=\inf_{\gamma\in\Gamma}E[g(X,\gamma)]^\intercal WE[g(X,\gamma)] .
\end{align}
Here, $\theta_0\in \prod_{j=1}^m \ell^\infty(\Gamma)$ is defined by $\theta_0(\gamma)=E[g(X,\gamma)]$, $\mathbb D=\prod_{j=1}^m\ell^\infty(\Gamma)$, $\mathbb E=\mathbf R$, and $\phi: \prod_{j=1}^m\ell^\infty(\Gamma)\to\mathbf R$ is defined by $\phi(\theta)=\inf_{\gamma\in\Gamma}\theta(\gamma)^\intercal W\theta(\gamma)$. The bootstrap for the $J$ statistic has been studied by Hall_Horowitz1996bootstrap and Andrews2002higher. Note that $\theta_0$ is always identified even though $\gamma_0$ is potentially partially identified, which makes $\phi$ second order nondifferentiable as will be shown below. \qed
\end{ex}
\subsection{Concepts of Differentiability}
All examples in the previous subsection exhibit first order degeneracy, i.e., there exist points $\theta$ in $\mathbb D$ such that the first order derivative $\phi_\theta'$ is $0$ and in some cases $\phi$ is not even differentiable at $\theta$, which can be seen from Examples (ref) and (ref) respectively. As such, we resort to a second order expansion that handles first order degeneracy and meanwhile accommodates potential nondifferentiability of $\phi$. Let us proceed by recalling notions of first order differentiability Shapiro1990,Fang_Santos2014HDD
\begin{defn}
Let $\mathbb D$ and $\mathbb E$ be normed spaces equipped with norms $\|\cdot\|_{\mathbb D}$ and $\|\cdot\|_{\mathbb E}$ respectively, and $\phi:\mathbb D_\phi\subseteq \mathbb D\to\mathbb E$.
\begin{itemize}
• The map $\phi$ is said to be {\it Hadamard differentiable} at $\theta \in\mathbb D_\phi$ {\it tangentially} to a set $\mathbb D_0\subseteq\mathbb D$, if there is a continuous linear map $\phi_\theta':\mathbb D_0\to\mathbb E$ such that:
\begin{equation}
\lim_{n\rightarrow \infty}\| \frac{\phi(\theta +t_nh_n)-\phi(\theta)}{t_n} - \phi_\theta'(h)\|_{\mathbb E} = 0 ,
\end{equation}
for all sequences $\{h_n\}\subset\mathbb D$ and $\{t_n\}\subset\mathbf R$ such that $t_n\to 0$, $h_n\to h\in\mathbb D_0$ as $n\to\infty$ and $\theta+t_nh_n\in\mathbb D_\phi$ for all $n$.
• The map $\phi$ is said to be {\it Hadamard directionally differentiable} at $\theta \in\mathbb D_\phi$ {\it tangentially} to a set $\mathbb D_0\subseteq\mathbb D$, if there is a continuous map $\phi_\theta':\mathbb D\to\mathbb E$ such that:\footnote{We note that the “tangential set” in Shapiro1991 refers to the domain of $\phi$ (i.e., $\mathbb D_\phi$ in our context), whereas here it refers to the domain $\mathbb D_0$ of the derivative $\phi_\theta'$.}
\begin{equation}
\lim_{n\rightarrow \infty}\|\frac{\phi(\theta +t_n h_n)-\phi(\theta)}{t_n} -\phi_\theta'(h)\|_{\mathbb E} = 0 ,
\end{equation}
for all sequences $\{h_n\}\subset\mathbb D$ and $\{t_n\}\subset\mathbf R_+$ such that $t_n\downarrow 0$, $h_n\to h\in\mathbb D_0$ as $n\to\infty$ and $\theta+t_nh_n\in\mathbb D_\phi$ for all $n$.
\end{itemize}
\end{defn}
Inspecting Definition (ref), we see that the main difference between Hadamard differentiability and directional differentiability lies in the linearity of the derivative. This turns out to be the exact gap between these two notions of differentiability. In particular, (ref) ensures that the directional derivative $\phi_\theta'$ is necessarily continuous and positively homogeneous of degree one, though potentially nonlinear Shapiro1990.
Given the introduced notions of differentiability and in view of the remarkable fact that Delta method is valid under even Hadamard directional differentiability in terms of deriving asymptotic distributions Shapiro1991,Dumbgen1993, it seems a natural next step to invoke the Delta method. However, in the presence of first order degeneracy, the resulting limiting distribution is degenerate at zero, rendering substantial challenges for inferential purposes. In essence, the Delta method is a stochastic version of Taylor expansion. Therefore, one could go one step further to explore the quadratic term when the linear term is degenerate. We thus follow Shapiro2000inference and define
\begin{defn}
Let $\phi:\mathbb D_\phi\subseteq \mathbb D\to\mathbb E$ be a map as in Definition (ref).
\begin{itemize}
• Suppose that $\phi:\mathbb D_\phi\to\mathbb E$ is Hadamard differentiable tangentially to $\mathbb D_0\subset\mathbb D$ such that the derivative $\phi_{\theta}': \mathbb D_0\to\mathbb E$ is well defined on $\mathbb D$. We say that $\phi$ is {\it second order Hadamard differentiable} at $\theta\in\mathbb D_\phi$ {\it tangentially} to $\mathbb D_0$ if there is a bilinear map $\Phi_\theta'':\mathbb D_0\times \mathbb D_0\to\mathbb E$ such that: for $\phi_\theta''(h)\equiv\Phi_\theta''(h,h)$,
\begin{align}
\lim_{n\rightarrow \infty}\|\frac{\phi(\theta +t_n h_n)-\phi(\theta)-t_n\phi_\theta'(h_n)}{t_n^2} -\phi_\theta”(h)\|_{\mathbb E} = 0 ,
\end{align}
for all sequences $\{h_n\}\subset\mathbb D$ and $\{t_n\}\subset\mathbf R^+$ such that $t_n\to 0$, $h_n\to h\in\mathbb D_0$ as $n\to\infty$ and $\theta+t_nh_n\in\mathbb D_\phi$ for all $n$.
• Suppose that $\phi:\mathbb D_\phi\to\mathbb E$ is Hadamard directionally differentiable tangentially to $\mathbb D_0\subset\mathbb D$ such that the derivative $\phi_{\theta}': \mathbb D_0\to\mathbb E$ is well defined on $\mathbb D$. We say that $\phi$ is {\it second order Hadamard directionally differentiable} at $\theta\in\mathbb D_\phi$ {\it tangentially} to $\mathbb D_0$ if there is a map $\phi_\theta'':\mathbb D_0\to\mathbb E$ such that:\footnote{Compared with Shapiro2000inference, we omitted $\frac{1}{2}$ in the denominator for notational compactness.}
\begin{align}
\lim_{n\rightarrow \infty}\|\frac{\phi(\theta +t_n h_n)-\phi(\theta)-t_n\phi_\theta'(h_n)}{t_n^2} -\phi_\theta”(h)\|_{\mathbb E} = 0 ,
\end{align}
for all sequences $\{h_n\}\subset\mathbb D$ and $\{t_n\}\subset\mathbf R^+$ such that $t_n\downarrow 0$, $h_n\to h\in\mathbb D_0$ as $n\to\infty$ and $\theta+t_nh_n\in\mathbb D_\phi$ for all $n$.
\end{itemize}
\end{defn}
The second order derivative $\phi_\theta''$ in both cases is necessarily continuous on $\mathbb D_0$, which can be shown in a straightforward manner as in the proof of Proposition 3.1 in Shapiro1990. Similar in spirit to Definition (ref),
the key difference between the above two notions of second order differentiability is that the former is a quadratic form corresponding to a bilinear map while the latter is in general only positively homogeneous of degree two, i.e., $\phi_\theta''(th)=t^2\phi_\theta''(h)$ for all $t\ge 0$ and all $h\in\mathbb D_0$. Note that it is possible that $\phi$ is first order Hadamard differentiable but only second order Hadamard directionally differentiable (see Example (ref)). In all our examples, $\phi$ is first order Hadamard differentiable though $\phi_{\theta}'$ may be degenerate; see Subsection (ref). We stress that requiring $\phi_{\theta}'$ to be well defined on the entirety of $\mathbb D$ does not demand differentiability on $\mathbb D$. Instead, it just means that $\phi_{\theta}'$ can take elements potentially not in $\mathbb D_0$ as arguments. Finally, we note that first and second order (directional) derivatives share the same domain $\mathbb D_0$.
If $\phi_\theta''$ in turn is degenerate, one can go beyond the second order, a possibility we do not pursue at length in this paper; see Remark (ref).
\begin{rem}
Suppose that $\phi:\mathbb D_\phi\subseteq \mathbb D\to\mathbb E$ is $(p-1)$-th order Hadamard directionally differentiable tangentially to $\mathbb D_0\subset\mathbb D$ such that the derivative $\phi_{\theta}^{(j)}: \mathbb D_0\to\mathbb E$ is well defined on $\mathbb D$ for all $j=1,\ldots,p-1$, where $p\ge 2$. Then we say that $\phi$ is $p$th order Hadamard directionally differentiable at $\theta\in\mathbb D_\phi$ tangentially to $\mathbb D_0$ if there is a map $\phi_\theta^{(p)}:\mathbb D_0\to\mathbb E$ such that:
\begin{align}
\phi(\theta +t_n h_n)=\phi(\theta)+\sum_{j=1}^{p-1}t_n^j\phi_\theta^{(j)}(h_n)+t_n^p\phi_\theta^{(p)}(h)+o(t_n^p) ,
\end{align}
for all sequences $\{h_n\}\subset\mathbb D$ and $\{t_n\}\subset\mathbf R^+$ such that $t_n\downarrow 0$, $h_n\to h\in\mathbb D_0$ as $n\to\infty$ and $\theta+t_nh_n\in\mathbb D_\phi$ for all $n$. Note that, similar to the treatment of $\phi_\theta''$, the factors $1/j!$ are incorporated in the definition of the derivatives $\phi_\theta^{(j)}$ to reflect the nature of them as approximating maps. Demyanov1974Minimax established the above high order expansion for $\mathbb D=\mathbf R^k$ with $k\in\mathbf N$ and $\mathbb E=\mathbf E$;\footnote{We thank an anonymous referee for bringing this reference to our attention.} see also Demyanov2009Minimax. \qed
\end{rem}
\subsubsection{Examples Revisited}
From now on, we shall focus on Examples (ref) and (ref) exclusively for conciseness; Examples (ref), (ref), (ref) and (ref) will be treated in Appendix (ref).
\begin{exctd}[(ref)]
In this example, the functional involved is second order Hadamard differentiable. Trivially we have
\begin{align}
\phi_\theta'(h)=2\theta h ,\,\phi_\theta”(h)=h^2 .
\end{align}
Note that the first order derivative $\phi_\theta'$ is degenerate when $\theta=0$, whereas $\phi_\theta''$ is everywhere nondegenerate. The bilinear map $\Phi_\theta'': \mathbf R^2\to\mathbf R$ here is given by $\Phi_\theta''(h,g)=hg$. \qed
\end{exctd}
In Example (ref), the domain $\mathbb D_0$ of the derivative $\phi_{\theta_0}''$ is a strict subset of $\mathbb D$.
\begin{exctd}[(ref)]
Consider $\theta\in\prod_{j=1}^{m}\ell^\infty(\Gamma)$ such that $\theta(\gamma_0)=0$ for some $\gamma_0\in\Gamma$. Then $\phi$ is Hadamard differentiable at $\theta$ and $\phi_{\theta}'(h)=0$ for all $h\in\prod_{j=1}^m\ell^\infty(\Gamma)$. Suppose further that $\Gamma$ is compact and that $\Gamma_0(\theta)\equiv\{\gamma_0\in\Gamma: \theta(\gamma_0)=0\}$ is in the interior of $\Gamma$. For $C^1(\Gamma)$ the space of continuously differentiable functions on $\Gamma$, if $\theta\in \prod_{j=1}^m C^1(\Gamma)$, then by Lemma (ref), under additional regularity conditions, $\phi$ is second order Hadamard directionally differentiable at $\theta$ tangentially to $\prod_{j=1}^m C(\Gamma)$ with the derivative given by: for any $h\in \prod_{j=1}^m C(\Gamma)$,
\[
\phi_{\theta}''(h)=\min_{\gamma_0\in\Gamma_0(\theta)}h(\gamma_0)^\intercal W^{1/2}M(\gamma_0) W^{1/2} h(\gamma_0)~,
\]
where $M(\gamma_0)=I_m-W^{1/2}J(\gamma_0)[J(\gamma_0)^\intercal W J(\gamma_0)]^{-1}J(\gamma_0)^\intercal W^{1/2}$ with $J(\gamma_0)\equiv\frac{d\theta(\gamma)}{d\gamma^\intercal}\big|_{\gamma=\gamma_0}$ the Jacobian matrix and $I_m$ the identity matrix of size $m$. Here, invertibility of $J(\gamma_0)$ is an implied requirement in Lemma (ref); see Remark (ref). Note that if $\gamma_0$ is point identified, then $\phi$ becomes second order Hadamard differentiable with
\[
\phi_{\theta}''(h)=h(\gamma_0)^\intercal W^{1/2}M(\gamma_0) W^{1/2} h(\gamma_0)~,
\]
which in turn yields $\chi^2(m-k)$ as the asymptotic distribution of the $J$-statistic under optimal weighting. We emphasize that the regularity conditions in Lemma (ref) are sufficient for applying our framework but by no means necessary -- as explained in Section (ref), those sufficient conditions exclude the setup of Dovonon_Renault2013testing, and so we shall provide an alternative set of sufficient conditions there. \qed
\end{exctd}
\subsection{Second Order Delta Method}
The Delta method for potentially directionally differentiable maps as well as differentiable ones has proven powerful in asymptotic analysis Vaart1998,Shapiro1991,Fang_Santos2014HDD,Hansen2015regression. Unfortunately, it is insufficient to handle substantial challenges for inference arising from first order degeneracy. Heuristically, if $r_n\{\hat\theta_n-\theta_0\}\xrightarrow{L}\mathbb G$ and $\phi_{\theta_0}'=0$, then the Delta method implies that
\[
r_n\{\phi(\hat\theta_n)-\phi(\theta_0)\}\xrightarrow{L} \phi_{\theta_0}'(\mathbb G)\equiv 0~.
\]
For real-valued $\phi$, the usual confidence interval for $\phi(\theta_0)$ at asymptotic level $1-\alpha$ is
\begin{align}
[\phi(\hat\theta_n)-\frac{c_{1-\alpha/2}}{r_n},\phi(\hat\theta_n)-\frac{c_{\alpha/2}}{r_n}]=\{\phi(\hat\theta_n)\} ,
\end{align}
where the $c_\alpha$ is the $\alpha$-th quantile of $\phi_{\theta_0}'(\mathbb G)\equiv 0$ and is zero for all $\alpha\in(0,1)$. Clearly, $P(\phi(\theta_0)\in \{\phi(\hat\theta_n)\})=0$ if, for example, $\phi(\hat\theta_n)$ is a continuous random variable.
To circumvent the above difficulty, we impose the following conditions in order to obtain a suitable second order Delta method.
\begin{ass}
(i) $\mathbb D$ and $\mathbb E$ are normed spaces with norms $\|\cdot\|_{\mathbb D}$ and $\|\cdot\|_{\mathbb E}$ respectively; (ii) $\phi:\mathbb D_\phi\subset\mathbb D\to\mathbb E$ is second order Hadamard directionally differentiable at $\theta_0\in\mathbb D_\phi$ tangentially to $\mathbb D_0\subset\mathbb D$; (iii) $\phi_{\theta_0}'(h)=0$ for all $h\in\mathbb D_0$.
\end{ass}
\begin{ass}
(i) There is $\hat\theta_n:\{X_i\}_{i=1}^n\to\mathbb D_\phi$ such that $r_n\{\hat\theta_n-\theta_0\}\xrightarrow{L} \mathbb G$ in $\mathbb D$ for some $r_n\uparrow\infty$; (ii) $\mathbb G$ is tight and its support is in $\mathbb D_0$;\footnote{The support of $\mathbb G$ is the set of points in $\mathbb D$ all of whose open neighborhoods have positive probability.} (iii) $\mathbb D_0$ is closed under vector addition, i.e., $h_1+h_2\in\mathbb D_0$ whenever $h_1,h_2\in\mathbb D_0$.
\end{ass}
Assumption (ref) formalizes the requirement that the map $\phi:\mathbb D_\phi \rightarrow \mathbb E$ be second order Hadamard directionally differentiable at $\theta_0$, and the defining feature of this paper, namely, degeneracy of the first order derivative. Assumption (ref)(i) defines another key ingredient: there is an estimator $\hat\theta_n$ for $\theta_0$ that admits a weak limit $\mathbb G$ at a potentially non-$\sqrt n$ rate $r_n$; see Remark (ref). Assumption (ref)(ii) ensures that the support of $\mathbb G$ is included in the domain of the derivative $\phi_{\theta_0}''$ so that $\phi_{\theta_0}''(\mathbb G)$ is well defined, while tightness of $\mathbb G$ is only a minimal requirement. Assumption (ref)(iii) is a mild condition, which shall play a technical role in the proof of our bootstrap results.
Given Assumptions (ref) and (ref), we now present a second order Delta method building upon Shapiro2000inference and Romish2004delta but without requiring $\mathbb D_\phi$ to be convex.
\begin{thm}
If Assumptions (ref)(i)(ii) and (ref)(i)(ii) hold, then\footnote{The term $\phi_{\theta_0}''(r_n\{\hat\theta_n-\theta_0\})$ is interpreted as some continuous extension of $\phi_{\theta_0}''$ (which always exists in our setup) evaluated at $r_n\{\hat\theta_n-\theta_0\}$ whenever $r_n\{\hat\theta_n-\theta_0\}\notin \mathbb D_0$; see the comment preceding the proof of Theorem (ref). Since (ref) is an asymptotic result, the choice of the continuous extension is irrelevant.}
\begin{align}
r_n^2\{\phi(\hat\theta_n)-\phi(\theta_0)-\phi_{\theta_0}'(\hat\theta_n-\theta_0)\}=\phi_{\theta_0}”(r_n\{\hat\theta_n-\theta_0\})+o_p(1) .
\end{align}
and hence
\begin{align}
r_n^2\{\phi(\hat\theta_n)-\phi(\theta_0)-\phi_{\theta_0}'(\hat\theta_n-\theta_0)\}\xrightarrow{L} \phi_{\theta_0}”(\mathbb G) .
\end{align}
\end{thm}
The essence of Theorem (ref) is in complete accord with that underlying the first order Delta method. In particular, the definition of second order Hadamard directional differentiability is engineered so that the second order Delta method is nothing more than a stochastic version of the Taylor expansion of order two, i.e.,
\begin{align*}
\phi(\theta_0+t_nh_n)=\phi(\theta_0)+t_n\phi_{\theta_0}'(h_n)+t_n^2\phi_{\theta_0}”(h)+o(t_n^2) ,
\end{align*}
where $t_n$ corresponds to $r_n^{-1}$, and $h_n$ to $r_n\{\hat\theta_n-\theta_0\}$. Note that Theorem (ref) is valid regardless of the nature of the differentiability (i.e., fully differentiable or directionally differentiable) and the presence of first order degeneracy. When $\phi_{\theta_0}'$ is degenerate, the convergence (ref) simplifies to
\begin{align}
r_n^2\{\phi(\hat\theta_n)-\phi(\theta_0)\}\xrightarrow{L} \phi_{\theta_0}”(\mathbb G) .
\end{align}
Finally, we note that higher order versions of the Delta method can be developed along the lines of Remark (ref); see Remark (ref).
\begin{rem}
Suppose that Assumptions (ref)(i) and (ref)(i)(ii) hold and $\phi$ is $p$-th order Hadamard directionally differentiable at $\theta_0\in\mathbb D_\phi$ tangentially to $\mathbb D_0$. It follows that
\[
r_n^p\big[\phi(\hat\theta_n)-\{\phi(\theta_0)+\sum_{j=1}^{p-1}\phi_{\theta_0}^{(j)}(\hat\theta_n-\theta_0)\}\big]=
\phi_{\theta_0}^{(p)}(r_n\{\hat\theta_n-\theta_0\})+o_p(1)~,
\]
and hence
\[
r_n^p\big[\phi(\hat\theta_n)-\{\phi(\theta_0)+\sum_{j=1}^{p-1}\phi_{\theta_0}^{(j)}(\hat\theta_n-\theta_0)\}\big]\xrightarrow{L}\phi_{\theta_0}^{(p)}(\mathbb G)~. \tag*{$\qed$}
\]
\end{rem}
\section{The Bootstrap}
Establishing asymptotic distributions as in Theorem (ref) is the first step towards conducting statistical inference on $\phi(\theta_0)$, the usefulness of which relies on our ability to accurately estimate the limiting law. In this section, we discuss how first order degeneracy of $\phi$ can complicate inference using the standard bootstrap based on first and especially second order asymptotics, and provide alternative consistent resampling schemes.
\subsection{Bootstrap Setup}
Throughout, we let $\hat \theta_n^*$ denote a “bootstrapped version" of $\hat \theta_n$, which is defined as a function mapping the data $\{X_i\}_{i=1}^n$ and random weights $\{W_i\}_{i=1}^n$ that are independent of $\{X_i\}_{i=1}^n$ into the domain $\mathbb D_\phi$ of $\phi$. This general definition allows us to include diverse resampling schemes such as nonparametric, Bayesian, block, score, more generally multiplier and exchangeable bootstrap as special cases. Next, making sense of bootstrap consistency necessitates a metric that quantifies distances between probability measures. As is standard in the literature, we employ the bounded Lipschitz metric $d_{\operatorname{BL}}$ formalized by Dudley1966Baire,Dudley1968distance: for two Borel probability measures $L_1$ and $L_2$ on $\mathbb D$, define
\[
d_{\operatorname{BL}}(L_1,L_2)\equiv\sup_{f\in\operatorname{BL}_1(\mathbb D)}|\int f\,dL_1-\int f\,dL_2| ~,
\]
where we recall that $\operatorname{BL}_1(\mathbb D)$ denotes the set of Lipschitz functionals whose absolute level and Lipschitz constant are bounded by one, i.e.,
\begin{equation*}
BL_1(\mathbb D) \equiv \{f : \mathbb D \rightarrow \mathbf R : \sup_{t\in \mathbb D} |f(t)| + \sup_{t_1,t_2\in \mathbb D, t_1\neq t_2}\frac{|f(t_1) - f(t_2)|}{\|t_1-t_2\|_{\mathbb D}}\le 1\} .
\end{equation*}
Since weak convergence in the Hoffmann-J{\o}rgensen sense to separable limits can be metrized by $d_{\operatorname{BL}}$ Dudley1990nonlinear,Vaart_Wellner1990prohorov, we may now measure the distance between the “conditional law” of $\hat{\mathbb G}_n^*\equiv r_n\{\hat\theta_n^*-\hat\theta_n\}$ given $\{X_i\}$ and the limiting law of $r_n\{\hat\theta_n-\theta_0\}$ by
\begin{align}
d_{\operatorname{BL}}(\hat{\mathbb G}_n^*,\mathbb G)=\sup_{f\in\operatorname{BL}_1(\mathbb D)}|E_W[f(r_n\{\hat\theta_n^*-\hat\theta_n\})]-E[f(\mathbb G)]| ,
\end{align}
where $E_W$ denotes expectation with respect to the bootstrap weights $\{W_i\}_{i=1}^n$ holding the data $\{X_i\}_{i=1}^n$ fixed. Employing the distribution of $r_n\{\hat \theta_n^* - \hat \theta_n\}$ conditional on the data as an approximation to the distribution of $\mathbb G$ is then asymptotically justified if their distance, equivalently (ref), converges in probability to zero.
We formalize the above discussion by imposing the following assumptions on $\hat \theta_n^*$.
\begin{ass}
(i) $\hat \theta_n^* : \{X_i,W_i\}_{i=1}^n \rightarrow \mathbb D_\phi$ with $\{W_i\}_{i=1}^n$ independent of $\{X_i\}_{i=1}^n$; (ii) $\hat \theta_n^*$ satisfies $\sup_{f \in \text{BL}_1(\mathbb D)} |E_W[f(r_n\{\hat \theta_n^* - \hat \theta_n\})] - E[f(\mathbb G)]| = o_p(1)$.
\end{ass}
\begin{ass}
(i) $E[f(r_n\{\hat \theta_n^* - \hat \theta_n\})^*]-E[f(r_n\{\hat \theta_n^* - \hat \theta_n\})_*]\to 0$ for all $f\in\text{BL}_1(\mathbb D)$ where $f(r_n\{\hat \theta_n^* - \hat \theta_n\})^*$ and $f(r_n\{\hat \theta_n^* - \hat \theta_n\})_*$ denote minimal measurable majorant and maximal measurable minorant (with respect to $\{X_i,W_i\}_{i=1}^n$ jointly) respectively; (ii) $f(r_n\{\hat \theta_n^* - \hat \theta_n\})$ is a measurable function of $\{W_i\}_{i=1}^n$ outer almost surely in $\{X_i\}_{i=1}^n$ for any continuous and bounded $f:\mathbb D \rightarrow \mathbf R$.
\end{ass}
Assumption (ref)(i) formally defines the bootstrap analog $\hat \theta_n^*$ of $\hat\theta_n$, while Assumption (ref)(ii) simply imposes the consistency of the “law” of $r_n\{\hat \theta_n^* - \hat \theta_n\}$ conditional on the data for the law of $\mathbb G$, i.e., the bootstrap “works" for the estimator $\hat \theta_n$. Assumption (ref) is of technical concern. In particular, Assumption (ref)(i) can often be established as a result of bootstrap consistency Vaart1996, while Assumption (ref)(ii) is easy to verify for particular resampling schemes. For example, if $\{W_i\}_{i=1}^n\mapsto f(r_n\{\hat \theta_n^* - \hat \theta_n\})$ is continuous, then Assumption (ref)(ii) is fulfilled. When $\theta_0$ is Euclidean-valued, i.e., $\mathbb D=\mathbf R^k$ with $k\in\mathbf N$, one can dispense with Assumption (ref).
\subsection{Failures of the Standard Bootstrap}
We now turn to the challenges for inferences using the standard bootstrap caused by first order degeneracy. As is well known in the literature, the law of
\begin{align}
r_n\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}
\end{align}
conditional on the data provides a consistent estimator of the law of $\phi_{\theta_0}'(\mathbb G)$ provided $\phi$ is Hadamard differentiable Vaart1996, which in particular includes the case when $\phi_{\theta_0}'=0$. In other words, the standard bootstrap, meaning the law of (ref) conditional on the data, is consistent for the law of $\phi_{\theta_0}'(\mathbb G)$ regardless of the presence of first order degeneracy.
Substantial difficulties, however, arise from using (ref) for inferential purposes when first order degeneracy does occur. Ignoring the first order degeneracy or perhaps as a way to avoid ridiculous confidence intervals such as (ref), one might consider the following confidence interval for real-valued $\phi(\theta_0)$:
\begin{align}
[\phi(\hat\theta_n)-\frac{\tilde c_{1-\alpha/2}}{r_n}, \phi(\hat\theta_n)-\frac{\tilde c_{\alpha/2}}{r_n}] ,
\end{align}
where $\tilde c_{1-\alpha}$ is the $(1-\alpha)$-th bootstrapped quantile for $\alpha\in(0,1)$ defined as
\[
\tilde c_{1-\alpha}\equiv \inf\{c\in\mathbf R: P_W(r_n\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}\le c)\ge 1-\alpha\}~.
\]
However, establishing the validity of (ref) as a level $1-\alpha$ confidence interval for $\phi(\theta_0)$ is problematic because $\tilde c_{1-\alpha}\xrightarrow{p} 0$ for all $\alpha\in(0,1)$ and $0$ is a discontinuity point of the cdf of the limit (see Lemma (ref)).
In fact, simple algebra reveals that (ref) is numerically identical to
\begin{align}
[\phi(\hat\theta_n)-\frac{\bar c_{1-\alpha/2}}{r_n^2}, \phi(\hat\theta_n)-\frac{\bar c_{\alpha/2}}{r_n^2}] ,
\end{align}
where $\bar c_\alpha$ is defined as
\[
\bar c_{1-\alpha}\equiv \inf\{c\in\mathbf R: P_W(r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}\le c)\ge 1-\alpha\}~.
\]
In other words, $\bar c_\alpha$ is the $\alpha$-th bootstrapped quantile of the standard bootstrap based on second order asymptotics:
\begin{align}
r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\} .
\end{align}
As illustrated by Babu1984bootstrap for the squared mean example, the conditional law of (ref) is inconsistent for the law of $\phi_{\theta_0}''(\mathbb G)$ when $\theta_0=0$, the point at which first order degeneracy arises. We next demonstrate that the bootstrap failure in this simple example is a reflection of a deeper principle: the second order standard bootstrap is consistent if and only if $\phi_{\theta_0}''$ is degenerate, under regularity conditions.
\begin{thm}
Suppose that Assumptions (ref), (ref), (ref) and (ref) hold, and that $\mathbb G$ is centered Gaussian. Then $\phi_{\theta_0}''=0$ on the support of $\mathbb G$ if and only if
\begin{align}
\sup_{f\in\operatorname{BL}_1(\mathbb E)}|E_W[f(r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\})]-E[f(\phi_{\theta_0}”(\mathbb G))]|=o_p(1) .
\end{align}
If, in addition, $\phi$ is second order Hadamard differentiable, then the conclusion holds without requiring $\mathbb G$ to be centered Gaussian.
\end{thm}
The sufficiency part of the theorem is somewhat expected and not a deep result, while the necessity is perhaps surprising and has far-reaching implications for statistical inference as we shall detail shortly. The proof of the latter consists of two steps: in the first step, we show that bootstrap consistency as in (ref) implies existence of a bilinear map $\Phi_{\theta_0}''$ corresponding to $\phi_{\theta_0}''$, in similar fashion as the proof of Theorem 3.1 in Fang_Santos2014HDD; in the second step, we establish that $\Phi_{\theta_0}''$ and hence $\phi_{\theta_0}''$ is necessarily degenerate. Both steps involve the insights of equating distributions through their characteristic functionals as in Vaart1991differentibility and Hirano_Porter2012.
Theorem (ref) implies that, in the presence of first order degeneracy, if the second order derivative $\phi_{\theta_0}''$ is nondegenerate, then the standard bootstrap based on second order asymptotics is necessarily inconsistent whenever $\mathbb G$ is centered Gaussian. If $\phi_{\theta_0}''$ is degenerate, we have a degenerate limiting distribution that can not be directly used for inference. We thus conclude that bootstrap failure is an inherent implication of models with first order degeneracy.
Heuristically, the reason why the standard bootstrap fails is that even though $r_n^2\phi_{\theta_0}'(\hat\theta_n-\theta_0)=0$ in the “real world”, its bootstrap counterpart is non-negligible. To see this, consider the squared mean example. If $\theta_0=0$, then
\[
n\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)=n 2\hat\theta_n\cdot\{\hat\theta_n^*-\hat\theta_n\}=2\sqrt{n}\{\hat\theta_n-\theta_0\}\cdot\sqrt n\{\hat\theta_n^*-\hat\theta_n\}=O_p(1)~.
\]
This is an emphatic reflection of Efron1979's caveat that the bootstrap, as well as other resampling schemes, provides frequency approximations rather than likelihood approximations. These heuristics suggest that the standard bootstrap might work if the first order term $r_n^2\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)$ is included, which turns out to be true for sufficiently smooth maps; see Theorem (ref).
It is worth noting that Theorem (ref) holds even if $\phi$ is smooth. Consequently, first order degeneracy is a source of bootstrap inconsistency completely different from that discussed in Fang_Santos2014HDD, i.e., nondifferentiability of $\phi$. In addition, we note that, without the qualifier that $\mathbb G$ is centered Gaussian, bootstrap consistency (ref) holds if and only if $\phi_{\theta_0}''(\mathbb G+h)-\phi_{\theta_0}''(h)\overset{d}{=}\phi_{\theta_0}''(\mathbb G)$ for all $h\in\mathrm{Supp}(\mathbb G)$ under mild support conditions; see Theorem A.1 in Fang_Santos2014HDD.
Finally, to further articulate the relations between the current work and that of Fang_Santos2014HDD, we present a table that describes the scopes we work in.
\begin{table}[!htbp]
\begin{threeparttable}
\caption{Comparison with Fang_Santos2014HDD}
\begin{tabular}{cccc}
\hline\hline
\multirow{2}{*} & \multirow{2}{*} & \multicolumn{2}{c}{First Order Degeneracy (i.e.\ $\phi_{\theta_0}'=0$)} \\
\cmidrule{3-4}
& & Yes & No\\
\hline
\multirowcell{2}[-2ex]{ Nondifferentiability\\ (1st or 2nd order)} & Yes & \makecell{This paper\\ $r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}\overset{L^*}{\nrightarrow}\phi_{\theta_0}''(\mathbb G)$} & \makecell{Fang_Santos2014HDD\\ $r_n\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}\overset{L^*}{\nrightarrow}\phi_{\theta_0}'(\mathbb G)$} \\
\cmidrule{3-4}
& No & \makecell{This paper\\ $r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}\overset{L^*}{\nrightarrow}\phi_{\theta_0}''(\mathbb G)$} & \makecell{Standard\\ $r_n\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)\}\overset{L^*}{\to}\phi_{\theta_0}'(\mathbb G)$}\\
\hline\hline
\end{tabular}
\end{threeparttable}
\begin{tablenotes}
• 1. $L^*$ signifies conditional weak convergence (made precise by, for example, $d_{\mathrm{BL}}$).
• 2. It is assumed that $r_n\{\hat\theta_n-\theta_0\}\xrightarrow{L}\mathbb G$ and $r_n\{\hat\theta_n^*-\hat\theta_n\}\overset{L^*}{\rightarrow}\mathbb G$.
• 3. Since $\phi$ is first order differentiable when $\phi_{\theta_0}'=0$, the nondifferentiability is meant in the second order for the third column and the first order for the last column.
\end{tablenotes}
\end{table}
\subsection{The Babu Correction}
We now extend the Babu correction under our more general setup. We proceed by imposing the following assumption.
\begin{ass}
(i) The map $\phi: \mathbb D_\phi\subset\mathbb D\to\mathbb E$ is second order Hadamard differentiable at $\theta_0\in\mathbb D_\phi$ tangentially to $\mathbb D_0$; (ii) $\phi$ is first order Hadamard differentiable at every point in some neighborhood of $\theta_0$ tangentially to $\mathbb D_0$ such that \footnote{The appearance of the factor 2 is due to omission of the factor $1/2$ in Definition (ref).}
\begin{align}
\lim_{n\to\infty}\|\frac{\phi_{\theta_0+t_ng_n}'(h_n)-\phi_{\theta_0}'(h_n)}{t_n}-2\Phi_{\theta_0}”(g,h)\|_{\mathbb E}=0 ,
\end{align}
for all sequences $\{g_n,h_n\}\subset\mathbb D$ and $\{t_n\}\subset\mathbf R^+$ such that $t_n\downarrow 0$, $(g_n,h_n)\to (g,h)\in\mathbb D_0\times\mathbb D_0$ as $n\to\infty$ and $\theta+t_ng_n, \theta+t_nh_n\in\mathbb D_\phi$ for all sufficiently large $n$, where $\Phi_{\theta_0}'': \mathbb D_0\times\mathbb D_0\to\mathbb E$ is the bilinear map underlying $\phi_{\theta_0}''$.
\end{ass}
Assumption (ref)(i) defines the scope of the Babu correction: it shall be applied to smooth maps, which excludes, for example, the functional associated with the $J$-test in GMM settings when first order or global identification fails -- see Section (ref). Assumption (ref)(ii) is stronger than $\phi$ being simply second order Hadamard differentiable, in that it requires the existence of first order derivative at all points in a neighborhood of $\theta_0$ such that (ref) holds. Assumption (ref) is fulfilled for the setup considered in Babu1984bootstrap and for Examples (ref) and (ref), but violated for the remaining examples.
Under Assumption (ref), the corrected bootstrap
\begin{align}
r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)-\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)\}
\end{align}
is consistent for the law of $\phi_{\theta_0}''(\mathbb G)$ regardless of the degeneracy of $\phi_{\theta_0}'$.
\begin{thm}
If Assumptions (ref)(i)(ii), (ref), (ref), (ref) and (ref) hold, then
\begin{align}
\sup_{f\in\operatorname{BL}_1(\mathbb E)}|E_W[f(r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)-\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)\})]-E[f(\phi_{\theta_0}”(\mathbb G))]|=o_p(1) .
\end{align}
\end{thm}
Theorem (ref) generalizes Babu1984bootstrap considerably in that it accommodates semiparametric and nonparametric models, and allows wider resampling schemes beyond the nonparametric bootstrap of Efron1979. The Babu correction works nicely with smooth maps in the sense of Assumption (ref), but unfortunately is inadequate to handle nonsmooth ones. This is because when $\phi$ is only second order directionally differentiable, often times the derivative $\phi_{\theta_{0}}''$ is not “continuous” in $\theta_{0}$, implying that the Babu correction (ref) is unable to estimate $\phi_{\theta_{0}}''$ properly and in this way results in inconsistent estimates. For this reason, we next provide yet another resampling method which accommodates (second order) nondifferentiable maps.
\subsection{A Modified Bootstrap}
In this subsection, we shall present a modified bootstrap following Fang_Santos2014HDD that is consistent for the law of $\phi_{\theta_0}''(\mathbb G)$, and adaptive to both the presence of first order degeneracy and nondifferentiability of $\phi$.
The heuristics underlying our proposal, however, are connected to those in Fang_Santos2014HDD in a subtle way. In the context of first order asymptotics where $\phi$ is only directionally differentiable, inconsistency of the standard bootstrap arises from its inability to properly estimate the directional derivative $\phi_{\theta_0}'$. In our setup, however, there are examples in which the derivative $\phi_{\theta_0}''$ is a known map; see Examples (ref) and (ref) which are all differentiable maps. The standard bootstrap in these settings fails because there is a non-negligible term being neglected. However, in all other examples where $\phi$ is not smooth enough, Fang_Santos2014HDD's arguments will come into play as well.
In any case, the second order weak limit $\phi_{\theta_0}''(\mathbb G)$ is a composition of the derivative $\phi_{\theta_0}''$ and the limit $\mathbb G$ of $\hat\theta_n$, as is the first order limit $\phi_{\theta_0}'(\mathbb G)$. Thus, the law of $\phi_{\theta_0}''(\mathbb G)$ can be estimated by composing a suitable estimator $\hat\phi_n''$ for $\phi_{\theta_0}''$ with a consistent bootstrap approximation for the law of $\mathbb G$, in exactly the same fashion as the resampling scheme proposed by Fang_Santos2014HDD. That is, we propose employing the law of
\begin{align}
\hat\phi_n”(r_n\{\hat\theta_n^*-\hat\theta_n\})
\end{align}
conditional on the data as an approximation for the law of $\phi_{\theta_0}''(\mathbb G)$, where $\hat\phi_n'': \mathbb D\to\mathbb E$ is a suitable estimator of $\phi_{\theta_0}''$. Certainly, we would like $\hat\phi_n''$ to converge to $\phi_{\theta_0}''$ in some sense as $n\to\infty$. This can be made precise as follows.
\begin{ass}
$\hat \phi_n'' : \mathbb D \rightarrow \mathbb E$ is a function of $\{X_i\}_{i=1}^n$ satisfying that for every sequence $\{h_n\}\subset\mathbb D$ and every $h\in\mathbb D_0$ such that $h_n\to h$ as $n\to\infty$,
\begin{align}
\hat\phi_n”(h_n)\xrightarrow{p} \phi_{\theta_0}”(h) .
\end{align}
\end{ass}
Assumption (ref) says that $\hat\phi_n''$ converges in probability to $\phi_{\theta_0}''$ along any convergent sequence $h_n\to h$ as $n\to\infty$. In cases when $\phi_{\theta_0}''$ is a known map, we may simply set $\hat\phi_n''=\phi_{\theta_0}''$ for all $n\in\mathbf N$. It is worth noting that Assumption (ref) is equivalent to requiring: for every compact set $K\subset\mathbb D_0$ and every $\epsilon>0$,
\begin{align}
\lim_{\delta\downarrow 0}\limsup_{n\rightarrow \infty} P\Big( \sup_{h \in K^{\delta}} \| \hat \phi_n”(h) - \phi_{\theta_0}”(h)\|_{\mathbb E} > \epsilon\Big) = 0 ,
\end{align}
where $K^{\delta} \equiv \{a \in \mathbb D : \inf_{b \in K} \|a - b\|_{\mathbb D} < \delta\}$; see Lemma (ref). Condition (ref) was employed in Fang_Santos2014HDD who also provided several sufficient conditions for it to hold. For example, if $\hat \phi_n'' : \mathbb D \rightarrow \mathbb E$ is Lipschitz continuous, then pointwise consistency of $\hat \phi_n''$ suffices for (ref). Unfortunately, second order derivatives often lack uniform continuity and hence those sufficient conditions are inapplicable. Nonetheless, condition (ref) is straightforward to verify in all our examples.
Given the equivalence of conditions (ref) and (ref), consistency of our modified bootstrap (ref) follows from Theorem 3.2 in Fang_Santos2014HDD.
\begin{thm}
Under Assumptions (ref)(i)(ii), (ref), (ref), (ref) and (ref), it follows that
\begin{equation}
\sup_{f \in \operatorname{BL}_1(\mathbb E)} |E_W[f(\hat \phi_n”(r_n\{\hat \theta_n^* - \hat \theta_n\}))] - E[f(\phi_{\theta_0}”(\mathbb G))]| = o_p(1) .
\end{equation}
\end{thm}
Theorem (ref) shows that the law of $\hat \phi_n''(r_n\{\hat \theta_n^* - \hat \theta_n\})$ conditional on the data is indeed consistent for the law of $\phi_{\theta_0}''(\mathbb G)$, regardless of the degree of smoothness of $\phi$ and degeneracy of $\phi_{\theta_0}'$. Interestingly, the resampling scheme in Theorem (ref) is a mixture of the classical bootstrap and analytical asymptotic approximations. Finally, we note that Assumption (ref) allows us to think of Theorem (ref) as a variant of the extended continuous mapping theorem.
Theorems (ref) and (ref) are useful for hypothesis testing. Specifically, consider
\begin{align}
\mathrm H_0: \phi(\theta_0)=0 \mathrm H_1: \phi(\theta_0)>0 .
\end{align}
Under first order degeneracy, as is the case in all our examples, we employ the test of rejecting $\mathrm H_0$ if $r_n^2\phi(\hat\theta_n)>\hat c_{1-\alpha}$ where $\hat c_{1-\alpha}$ is the critical value constructed from the Babu correction or our proposed bootstrap, i.e.,
\begin{align}
\hat c_{1-\alpha} = \inf\{ c\in\mathbf R : P_W(r_n^2\{\phi(\hat\theta_n^*)-\phi(\hat\theta_n)-\phi_{\hat\theta_n}'(\hat\theta_n^*-\hat\theta_n)\}\leq c) \geq 1-\alpha\} ,
\end{align}
or
\begin{align}
\hat c_{1-\alpha} = \inf\{ c\in\mathbf R : P_W(\hat \phi_n”(r_n\{\hat \theta_n^* - \hat \theta_n\}) \leq c) \geq 1-\alpha\} .
\end{align}
Note that $\hat c_{1-\alpha}$ is generally infeasible but can be estimated by Monte Carlo simulations Efron1979,Hall1992bootstrap,HorowitzBoot. The pointwise size control of our test then follows according to Theorems (ref) and (ref). In fact, under additional restrictions, it can provide local size control. This property is particularly attractive because of the irregularity arising from nondifferentiability of $\phi$. In this case, pointwise asymptotic approximations can be misleading Imbens_Manski,AndrewsandGuggen2009ETA. Interestingly, it turns out that there is another source of irregularity due to the nature of first order degeneracy (see Lemma (ref)). We relegate the detailed discussions to Appendix (ref) in order to make our presentation concise.
We now briefly compare the Babu correction, the above composition procedure and the recentered bootstrap Hall_Horowitz1996bootstrap,HorowitzBoot. In some cases (for instance, Example (ref) and the regular $J$-test), they coincide with each other. However, the Babu correction applies to general smooth functionals, rather than just quadratic forms, and hence can be thought of as a generalization of the recentered bootstrap. The composition procedure, which works for an even larger class of functionals, is a direct approach by exploiting the structure of the limits, and hence is more tractable.
\begin{rem}
Examples where the convergence rate is not $\sqrt n$ include inference based on kernel estimators with undersmoothing Hall1992bootstrap, smoothed maximum score estimators Horowitz2002maxscore, and cointegration regressions ChangParkSong2006BootCoint. For nonstandard convergence rates, however, the bootstrap process $r_n\{\hat\theta_n^*-\hat\theta_n\}$ can fail to consistently estimate the law of $\mathbb G$, violating Assumption (ref)(ii). Fortunately, as far as Theorem (ref) is concerned, any consistent estimator, which need not satisfy Assumption (ref)(ii), will do. For example, in cube-root estimation problems, one could instead employ some smoothed bootstrap $r_n\{\tilde\theta_n^*-\tilde\theta_n\}$ where $\tilde\theta_n^*$ and $\tilde\theta_n$ are some smoothed estimators, or $m$ out of $n$ resampling (or subsampling) $m_n\{\hat\theta_{m_n}^*-\hat\theta_{n}\}$ where $\hat\theta_{m_n}^*$ is a bootstrap estimator based on subsamples of size $m_n$. In the context of estimating nonincreasing density functions, see Kosorok2008Grenander and Sen_Banerjee_Woodroofe2010; for bootstrapping the maximum score estimators, see Delgado_Poo_Wolf2001 and Patra_Seijo_Sen2015.\qed
\end{rem}
\subsection{Estimation of the Derivative}
Given the posited bootstrap consistency for the law of $\mathbb G$, the remaining crucial piece towards consistent bootstrap for the law of $\phi_{\theta_0}''(\mathbb G)$ based on Theorem (ref) is then an estimator $\hat\phi_n''$ of the derivative $\phi_{\theta_0}''$ that satisfies Assumption (ref). There are two general approaches for estimation of $\phi_{\theta_0}''$: one by exploiting the structure of $\phi_{\theta_0}''$, and the other one based on numerical differentiation as we describe now.
When first order degeneracy occurs, we have
\begin{align}
\phi_{\theta_0}”(h)=\lim_{n\to\infty}\frac{\phi(\theta_0+t_nh)-\phi(\theta_0)}{t_n^2} .
\end{align}
We may thus estimate $\phi_{\theta_0}''$ via numerical differentiation as follows: for any $h\in\mathbb D$,
\begin{align}
\hat\phi_n”(h)=\frac{\phi(\hat\theta_n+t_nh)-\phi(\hat\theta_n)}{t_n^2} .
\end{align}
If $t_n$ tends to zero at a suitable rate, the sense of which is made precise by the following assumption, then $\hat\phi_n''$ is a good estimator for $\phi_{\theta_0}''$ in the sense of Assumption (ref).
\begin{ass}
$\{t_n\}_{n=1}^\infty$ is a sequences of scalars such that $t_n\downarrow 0$ and $r_{n}t_{n}\to\infty$.
\end{ass}
Assumption (ref) allows a wide range of tuning parameters that can deliver first order validity of our method. The optimal choice of $t_n$ is challenging and beyond the scope of the present paper, which we hope to address in future. The next proposition confirms the validity of the numerical estimator (ref).
\begin{pro}[Hong_Li2015numericaldelta]
If Assumptions (ref), (ref)(i)(ii), and (ref) hold, then the numerical estimator $\hat\phi_n''$ in (ref) satisfies Assumption (ref).
\end{pro}
The numerical differentiation approach of estimating the derivatives, in the context of the Delta method, dates back to at least Dumbgen1993 in his proposal of the rescaled bootstrap. However, the way it was presented is quite implicit in revealing this, and so the bootstrap procedure is sometimes misunderstood as the $m$ out of $n$ resampling. Effectively, the rescaled bootstrap amounts to estimating the derivative numerically and the law of $\mathbb G$ using $n$ bootstrap samples; see Beare_Fang2016Grenander for more details. The recent work of Hong_Li2015numericaldelta provided a range of extensions of the numerical Delta method that have wide applications in econometrics.
Proposition (ref) provides a way of estimating the derivative $\phi_{\theta_0}''$ that is tractable in the sense that there is no need to explore the particular structures of $\phi$ or $\phi_{\theta_0}''$ as long as the tuning parameter $t_n$ is properly chosen. On the other hand, the expression of $\phi_{\theta_0}''$ itself often suggests an intuitive estimator as we elaborate in the next subsection.
\subsubsection{Examples Revisited}
Examples (ref) is trivial since $\phi_{\theta_0}''$ is a known map and hence one can simply set $\hat\phi_{n}''=\phi_{\theta_0}''$ for all $n\in\mathbf N$. Example (ref) is more complicated.
\begin{exctd}[(ref)]
In the classical case when $\Gamma_0(\theta)$ is singleton, we may estimate $\phi_{\theta_0}''$ based on the GMM estimator $\hat\gamma_n$ and the estimated Jacobian matrix $\hat J_n$. Generally, there are two unknown objects involved in the second order derivative: the identified set $\Gamma_0(\theta)$ and $J(\cdot)$. Let $\mathbf M^{m\times k}$ be the space of $m\times k$ matrices. Suppose that $\hat\Gamma_n\subset\Gamma$ is a $d_H$-consistent estimator for $\Gamma_0(\theta)$, and $\hat J_n: \Gamma\to\mathbf M^{m\times k}$ an estimator for $J: \Gamma\to\mathbf M^{m\times k}$ such that $\sup_{\gamma\in\Gamma}\|\hat J_n(\gamma)-J(\gamma)\|\xrightarrow{p} 0$. Then we may estimate $\phi_{\theta_0}''$ by
\begin{align}
\hat\phi_n”(h)=\min_{\gamma\in\hat\Gamma_n}\min_{v\in B_n}\{h(\gamma)-\hat J_n(\gamma)v\}^\intercal W\{h(\gamma)-\hat J_n(\gamma)v\} ,
\end{align}
where $B_n\equiv\{v\in\mathbf R^k: \|v\|\le t_n^{-1}\}$ for $t_n\downarrow 0$ satisfying $t_n\sqrt n\to\infty$. Consistency of $\hat\Gamma_n$ can be established by appealing to CHT2007, while uniform consistency of $\hat J_n$ can be derived using Glivenko-Cantelli type arguments. Following the proof of Lemma (ref), it is straightforward to show that $\hat\phi_n''$ satisfies Assumption (ref). \qed
\end{exctd}
\section{Application: Testing for Common CH Features}
In this section, we apply our framework to develop a robust test of common conditionally heteroskedastic (CH) factor structure by allowing multiple common CH features. Let $\{Y_t\}_{t=1}^T$ be a $k$-dimensional time series. According to Engle_Kozicki1993CF, a feature that is present in each component of $Y_t$ is said to be common to $Y_t$ if there exists a linear combination of $Y_t$ that fails to have the feature. A canonical example is the notion of cointegration developed by Engle_Granger1987Co-In in order to characterize the common feature of stochastic trend.
\subsection{The Setup}
Following Engle_Ng_Rothschild1990asset and Dovonon_Renault2013testing, suppose that the $k$-dimensional process $\{Y_t\}$ satisfies
\begin{align}
\operatorname{Var}(Y_{t+1}|\mathcal F_t)=\Lambda D_t\Lambda^\intercal+\Omega ,
\end{align}
where $\Lambda$ is a $k\times p$ matrix of full column rank with $p\le k$, $D_t$ a $p\times p$ diagonal matrix with diagonal (random) elements $\sigma _{jt}^{2}$ for $j=1,\ldots ,p$, $\Omega $ a $k\times k$ positive semidefinite matrix, and $\{\mathcal F_t\}_{t=1}^\infty$ a filtration to which $\{Y_t\}_{t=1}^\infty$ and $\{\sigma_{jt}^2: j=1,\ldots,p\}_{t=1}^\infty$ are adapted. By Engle_Kozicki1993CF, we say that $\{Y_t\}$ has a common CH feature if there exists some nonzero $\gamma_0\in\mathbf R^k$ such that $\operatorname{Var}(\gamma_0^\intercal Y_t|\mathcal F_t)$ is constant. The conditional covariance structure (ref) has some attractive properties that help to understand, for example, asset excess returns in a parsimonious way Engle_Ng_Rothschild1990asset. Thus, tests of common CH features can be used to detect the underlying common factor structures that simplify capturing interrelations of economic and financial variables under consideration.
With the help of instrumental variables, a common CH feature can be reformulated by unconditional moments that fit into the classical GMM framework. The following assumption is taken directly from Dovonon_Renault2013testing.
\begin{ass}
(i) $\Lambda$ is of full column rank; (ii) $\operatorname{Var}(\sigma_t^2)$ is nonsingular for $\sigma_t^2\equiv(\sigma_{1t}^2,\ldots,\sigma_{pt}^2)^\intercal$; (iii) $E[Y_{t+1}|\mathcal F_t]=0$; (iv) $Z_t$ is an $m\times 1$ $\mathcal F_t$-measurable random vector such that $\operatorname{Var}(Z_t)$ is nonsingular; (v) $\operatorname{Cov}(Z_t,\sigma_t^2)$ has full column rank $p$; (vi) $\{Y_t,Z_t\}$ is stationary and ergodic such that $E[\|Z_t\|^2]<\infty$ and $E[\|Y_t\|^4]<\infty$.
\end{ass}
Assumption (ref)(i)-(ii) ensure that there are exactly $k-p$ linearly independent vectors $\gamma_0$, spanning the null space of $\Lambda^\intercal$, such that $\operatorname{Var}(\gamma_0^\intercal Y_t|\mathcal F_t)$ is constant. In other words, the common CH features $\gamma_0$ are nonzero solutions of the equation $\Lambda^\intercal \gamma_0=0$.\footnote{If $\gamma_0$ is a common CH feature, so is $a\gamma_0$ for any nonzero $a\in\mathbf R$. For mathematical purpose, however, the number of common CH features is defined to the dimension of the null space of $\Lambda^\text{\scalebox{0.7}{$\intercal$}}$.} Assumption (ref)(iii) is a normalization condition that helps to simplify the exposition. Assumption (ref)(iv) defines the instrument $Z_t$ formed from the information set $\mathcal F_t$, while Assumption (ref)(v) implicitly requires that the number of instruments is no less than that of factors. Assumption (ref)(vi) further specifies the data generating process. We refer the readers to Dovonon_Renault2013testing for further details on Assumption (ref).
Assumption (ref) allows us to characterize common CH features as nonzero $\gamma_0$ satisfying the vector of unconditional moment equalities Dovonon_Renault2013testing:
\begin{align}
E[Z_{t}\{( \gamma_0^\intercal Y_{t+1}) ^{2}-c( \gamma_0 )\}]=0 ,
\end{align}
where $c(\gamma_0)=E[( \gamma_0^\intercal Y_{t+1})^2]$. It is then tempting to employ Hansen's $J$ statistic to test the existence of common CH features Engle_Kozicki1993CF. Unfortunately, as noted by Dovonon_Renault2013testing, the Jacobian matrix evaluated at the truth is degenerate at zero, rendering standard theory inapplicable. Though, as shall be illustrated, such degeneracy is of a nature different from first order degeneracy. By expanding the moment function to the second order, Dovonon_Renault2013testing showed that the asymptotic distribution of the $J$ statistic is highly nonstandard. Nonetheless, Dovonon_Goncalves2017bootstrapping developed a corrected bootstrap that can consistently estimate the limiting law when the bootstrap of Hall_Horowitz1996bootstrap fails to do so.
However, a key assumption in previous studies is that there exists a unique nonzero $\gamma_0$ such that (ref) is satisfied, ensured by exclusion restrictions and linear normalization $\sum_{j=1}^{k}\gamma_0^{(j)}=1$ Dovonon_Renault2013testing,Dovonon_Goncalves2017bootstrapping,Lee_Liao2017LocalIDfailure. This is undesirable for the following reasons. First, it is unknown {\it a priori} how many (linearly independent) CH features are common to the series under consideration. Second, as pointed out by Engle_Ng_Rothschild1990asset in the context of asset pricing, empirical work often considers large numbers of assets and the numbers of common CH features are expected to be large as well. Third, the linear normalization may in fact lead to no $\gamma_0$ satisfying (ref) (i.e.\ non-existence). For example, suppose $\Lambda=[1,1]^\intercal$. Then any common CH feature $\gamma_0$ must satisfy $\gamma_0^{(1)}+\gamma_0^{(2)}=0$, contradicting the linear normalization $\gamma_0^{(1)}+\gamma_0^{(2)}=1$ proposed in Dovonon_Renault2013testing. Fourth, in addition to the possibility that exclusion restrictions may be hard to form, the linear normalization is not susceptible of a unique common CH feature (i.e.\ non-uniqueness). To see this, suppose $\Lambda=[1,-1,-1]^\intercal$. Then for any common CH feature satisfying the normalization, we must have $\gamma_0^{(1)}-\gamma_0^{(2)}-\gamma_0^{(3)}=0$ and $\gamma_0^{(1)}+\gamma_0^{(2)}+\gamma_0^{(3)}=1$, which admit infinitely many solutions, i.e., the uniqueness is undermined in this case. These arguments motivate us to modify the $J$-test in a way that accommodates partial identification as well as degenerate Jacobian matrices. Such an extension is nontrivial because the second order (and hence global) identification,\footnote{Given first order identification failure, second order identification is equivalent to global identification in the current context because the moment function is quadratic in $\gamma_0$.} a condition that Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping heavily rely on, fails.
\subsection{A Modified $J$ Test}
To exclude the zero solution and avoid falsely excluding the existence of CH features, we employ the following normalization
\begin{align}
\gamma\in\mathbb S^k\equiv\{\gamma'\in\mathbf R^k: \|\gamma'\|=1\} .
\end{align}
Next, to map the current setup into our developed framework, we define a function $\phi: \prod_{j=1}^m \ell^\infty(\mathbb S^k)\to\mathbf R$ by: for any $\theta\in \prod_{j=1}^m \ell^\infty(\mathbb S^k)$,
\begin{align}
\phi(\theta)\equiv\inf_{\gamma\in\mathbb S^k}\|\theta(\gamma)\|^2 .
\end{align}
Then in view of the moment conditions (ref), the hypothesis that there exists at least one common CH feature can be reformulated as
\begin{align}
\mathrm H_0: \phi(\theta_0)=0\qquad \mathrm H_1: \phi(\theta_0)>0 ,
\end{align}
where $\theta_0: \mathbb S^k\to \mathbf R^m$ is defined as $\theta_0(\gamma)\equiv E[Z_{t}\{( \gamma ^\intercal Y_{t+1}) ^{2}-c( \gamma )\}]$. In this formulation, we have taken the identity matrix $I_m$ as the weighting matrix for simplicity.
Given our treatment of Example (ref), one might next try appealing to the results developed there. Unfortunately, they are not directly applicable. First, the parameter space $\Gamma$ of $\gamma_0$ is required to have nonempty interior (see Lemma (ref)), whereas in the current context $\Gamma=\mathbb S^k$ which has empty interior. Second, there is a technical condition there that prevents the Jacobian matrix from being degenerate even when there does exist a unique common CH feature; see Remark (ref) for details. Consequently, we have to re-verify the differentiability conditions for the map (ref). By Lemma (ref), under the null, $\phi$ is Hadamard differentiable with degenerate derivative, and second order Hadamard directionally differentiable at $\theta_0$ tangentially to $\prod_{j=1}^m C(\mathbb S^k)$ with the derivative
\begin{align}
\phi_{\theta_0}”(h)=\min_{\gamma_0\in\Gamma_0}\min_{v\in\mathbf R^k}\|h(\gamma_0)+G\operatorname{vec}(vv^\intercal)\|^2 ,
\end{align}
for any $h\in\prod_{j=1}^m C(\mathbb S^k)$, where $\Gamma_0=\{\gamma_0\in\mathbb S^k: \theta_0(\gamma_0)=0\}$ is the identified set of $\gamma_0$, and $G\in\mathbf M^{m\times k^2}$ with the $j$th row given by $\operatorname{vec}(\Delta_j)^\intercal$ and
\[
\Delta_j=E[Z_t^{(j)}(Y_{t+1}Y_{t+1}^\intercal-E[Y_{t+1}Y_{t+1}^\intercal])]~.
\]
We now make some remarks before proceeding further. First, we stress that first order degeneracy refers to the first order derivative $\phi_{\theta_0}'$ of the functional $\phi$, mapping from the function space $\prod_{j=1}^m \ell^\infty(\mathbb S^k)$ to $\mathbf R$, being degenerate, while the degeneracy Dovonon_Renault2013testing focused on refers to degeneracy of the Jacobian matrix $J(\gamma_0)\equiv\frac{d}{d\gamma}\theta_0(\gamma)|_{\gamma=\gamma_0}$ of the moment function $\theta_0$ that maps from the parameter space $\Gamma\subset\mathbf R^k$ of $\gamma_0$ to $\mathbf R^m$. Thus, the two types of degeneracy are conceptually different. Second, perhaps more importantly, they are also different in terms of the consequences. By Theorem (ref) and in view of (ref), $\phi$ being first order degenerate means that the second order standard bootstrap is inconsistent regardless of whether the Jacobian matrix is degenerate or not, while degeneracy of the Jacobian matrix generates the {\it additional} complication that $\phi$ is second order nondifferentiable as reflected by the inside minimization in (ref). Third, further allowing multiple (linearly independent) common CH features reinforces the nondifferentiability of $\phi$ as can be seen from the outside minimization in (ref).
Next, let the estimator $\hat\theta_T: \mathbb S^k\to\mathbf R^m$ be defined by $\hat\theta_T(\gamma)=\frac{1}{T}\sum_{t=1}^T Z_t\{(\gamma^\intercal Y_{t+1})^2-\hat c(\gamma)\}$ with $\hat c(\gamma)=\frac{1}{T}\sum_{t=1}^T(\gamma^\intercal Y_{t+1})^2$. Given the established differentiability of $\phi$, the asymptotic distribution of $\phi(\hat\theta_T)$ is then an immediate consequence of Theorem (ref) provided $\hat\theta_T$ converges weakly. Towards this end, we impose the following assumption as in Dovonon_Renault2013testing.
\begin{ass}
$Z_{t}$, $\operatorname{vec}(Y_{t}Y_{t}^\intercal)$ and $\operatorname{vec}(Y_{t}Y_{t}^\intercal) \otimes Z_{t}$ fulfill CLT.\footnote{The symbol $\otimes$ denotes Kronecker product.}
\end{ass}
Assumptions (ref) and (ref) together imply that
\begin{align}
\sqrt T \{\hat\theta_T-\theta_0\} \xrightarrow{L} \mathbb G in \prod_{j=1}^m \ell^\infty(\mathbb S^k) ,
\end{align}
where $\mathbb{G}$ is a zero mean Gaussian process with the covariance functional satisfying: for any $\gamma_{1}$, $\gamma_{2}\in\Gamma_{0}$ and $\mu_z\equiv E[Z_t]$,
\begin{align*}
E[\mathbb{G}(\gamma_{1})\mathbb{G}(\gamma_{2})]=E[(Z_{t}-\mu_{z})(Z_{t}-\mu_{z})^\intercal\{(\gamma_{1}^\intercal Y_{t+1})^{2}-c(\gamma_{1})\}\{(\gamma_{2}^\intercal Y_{t+1})^{2}-c(\gamma_{2})\}] .
\end{align*}
The proposition below delivers the limiting distribution of test statistic $T\phi(\hat\theta_T)$.
\begin{pro}
Let Assumptions (ref) and (ref) hold. Then we have under $\mathrm H_0$
\begin{equation}
T\min_{\gamma \in \mathbb S^k}\Vert\hat{\theta}_T(\gamma)\Vert^{2}\overset{L}{\rightarrow}\min_{\gamma_0 \in \Gamma_0}\min_{v\in \mathbf{R}^k}\Vert \mathbb{G}(\gamma_0) + G \operatorname{vec}(vv^\intercal)\Vert ^{2} .
\end{equation}
\end{pro}
The asymptotic distribution in (ref) is a highly nonlinear functional of the Gaussian process $\mathbb G$ in general, which turns out to be consistent with the limits obtained in Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping whenever their second order identification (and global) condition holds; see Remark (ref). In the latter setting, Dovonon_Goncalves2017bootstrapping showed that the recentered bootstrap of Hall_Horowitz1996bootstrap is inconsistent and thus proposed corrected versions of the standard GMM bootstrap. Unfortunately, their methods are not directly applicable to our setup that allows multiple common CH features (i.e.\ partial identification), because they crucially rely on the second order and global identification.
We next demonstrate how our bootstrap works. First, let $\{Y_{t+1}^*,Z_t^*\}_{t=1}^T$ be a bootstrap sample, which can be obtained by block bootstrap, nonoverlapping or overlapping Carlstein1986subseries,Kunsch1989Jackknife. Because the limiting process $\{\mathbb G(\gamma): \gamma\in\Gamma_0\}$ is determined by a martingale difference sequence indexed by $\gamma\in\Gamma_0$, the dependence structure of the data does not enter into the limit and we may thus employ Efron1979's nonparametric bootstrap or more general bootstrap schemes. In any case, we set
\begin{align}
\hat\theta_T^*(\gamma)=\frac{1}{T}\sum_{t=1}^T Z_t^*\{(\gamma^\intercal Y_{t+1}^*)^2-\hat c^*(\gamma)\} ,\,\hat c^*(\gamma)=\frac{1}{T}\sum_{t=1}^T(\gamma^\intercal Y_{t+1}^*)^2 .
\end{align}
To accommodate diverse resampling schemes, we simply impose the high level condition that $\hat\theta_T^*$ satisfies Assumptions (ref) and (ref) DehlingMikoschSorensen2002EPDep.
It remains to estimate the derivative (ref). The numerical differentiation approach can be implemented as in the beginning of Section (ref). That is, we estimate $\phi_{\theta_0}''$ by
\begin{align}
\hat\phi_T”(h)=\frac{\inf_{\gamma\in\mathbb S^k}\|\hat\theta_T(\gamma)+\kappa_Th(\gamma)\|^2- \min_{\gamma\in\mathbb S^k}\|\hat\theta_T(\gamma)\|^2}{\kappa_T^2} ,
\end{align}
where $\kappa_T$ satisfies Assumption (ref). We now describe how to estimate $\phi_{\theta_0}''$ by exploiting its structure. Let $B_T\equiv\{v\in\mathbf R^k: \|v\|\le \kappa_T^{-1/2}\}$ and $\hat\Gamma_T\equiv \{\gamma\in\mathbb S^k:\Vert\hat{\theta}_T(\gamma)\Vert^{2}-\phi(\hat{\theta}_T)\le \kappa_{T}^2\}$,\footnote{One can theoretically ignore $\phi(\hat\theta_T)$ in the expression of $\hat\Gamma_T$. As pointed out by CHT2007, however, such a modification helps avoid an empty set of solutions and improve power.} where $\kappa_T$ is to be specified. Then we may estimate $\phi_{\theta_0}''(h)$ by:
\begin{equation}
\hat{\phi} _{T }^{\prime \prime }(h) =\inf_{\gamma \in \hat\Gamma_T}\min_{v\in B_T}\Vert h(\gamma) + \hat{G} \operatorname{vec}(vv^\intercal)\Vert ^{2} ,
\end{equation}
where $\hat G\in\mathbf M^{m\times k^2}$ with its $j$th row given by $\operatorname{vec}(\hat\Delta_j)^\intercal$ for
\[
\hat\Delta_j=\frac{1}{T}\sum_{t=1}^T Z_t^{(j)}Y_{t+1}Y_{t+1}^\intercal-\frac{1}{T}\sum_{t=1}^T Z_t^{(j)}\frac{1}{T}\sum_{t=1}^T Y_{t+1}Y_{t+1}^\intercal~.
\]
In fact, we may further restrict the bounded set $B_T$ to reduce the computation burden for $\hat{\phi}_{T}^{\prime \prime }$; see Remark (ref). Clearly, the sequence $\{\kappa_T\}$ should tend to zero at a suitable rate as $T\to\infty$. This is made precise as follows.
\begin{ass}
$\{\kappa_T\}$ satisfies (i) $\kappa_T\downarrow 0$, and (ii) $\sqrt{T}\kappa_T\to\infty$.
\end{ass}
Assumption (ref) regulates the rates at which the tuning parameters $\kappa_T$ should approach zero, in order to deliver first order validity of our bootstrap inference procedure. The optimal choice of $\kappa_T$ is concerned with higher order accuracy of our method, which we do not touch in this paper. Combining the bootstrap $\hat\theta_T^*$ in (ref) and the derivative estimator, we are then able to consistently estimate the law of the weak limit in (ref) following Theorem (ref), which in turn allows us to construct critical values. Specifically, let $\hat c_{1-\alpha}$ be the $1-\alpha$ quantile of $\hat \phi_T''(\sqrt{T}\{\hat \theta_T^* - \hat \theta_T\})$ conditional on the data:\footnote{As usual, $P_W$ denotes the probability taken with respect to the bootstrap weights $\{W_T\}$, though in the current setup they are implicitly defined. Alternatively, one can think of $P_W$ as the probability with respect to the bootstrap sample $\{Z_t^*,Y_{t+1}^*\}$ holding data fixed.}
\begin{align}
\hat c_{1-\alpha} \equiv \inf\{ c\in\mathbf R : P_W(\hat \phi_T”(\sqrt{T}\{\hat \theta_T^* - \hat \theta_T\}) \leq c) \geq 1-\alpha\} .
\end{align}
The following proposition confirms that the test of rejecting the existence of common CH features when $T\phi(\hat\theta_T)>\hat c_{1-\alpha}$ is valid.
\begin{pro}
Suppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. If the cdf of the limit in (ref) is continuous and strictly increasing at its $1-\alpha$ quantile for $\alpha\in(0,1)$, then we have under $\mathrm H_0$,
\[\lim_{T\rightarrow\infty}P(T\min_{\gamma \in \mathbb S^k}\Vert\hat{\theta}_T(\gamma)\Vert^{2}> \hat c_{1-\alpha})=\alpha~.
\]
\end{pro}
Proposition (ref) implies our test has pointwise asymptotic exact size $\alpha$ and thus is not conservative (in the pointwise sense). Establishing local size control, unfortunately, is challenging in this case, because asymptotic distributions of the statistic under local perturbations do not have definitive relations (to us) to the corresponding pointwise limits in terms of first order dominance. It appears that the problem of developing (at least) locally valid {\it and} non-conservative overidentification tests is prevalent in the literature of partial identification CHT2007,AndrewsandSoares2010.
Finally, we stress that the quadratic structure of the moment function plays no essential roles in our framework. Building upon Example (ref), one may work with a general moment function that admits a zero Jacobian matrix, but without the requirement that the parameter space have nonempty interior. It is also possible to deal with GMM problems with a rank deficient but possibly nonzero Jacobian matrix. For example, consider testing whether a matrix $\Pi_0\in\mathbf M^{m\times k}$ with $m\ge k$ has rank $k$. This amounts to testing
\begin{align}
\mathrm H_0: \Pi_0 \gamma=0 for some \gamma\in\mathbb S^k\quad v.s. \quad \mathrm H_0: \Pi_0 \gamma\neq 0 for any \gamma\in\mathbb S^k .
\end{align}
Here, the moment function is $\gamma\mapsto \theta_0(\gamma)\equiv\Pi_0\gamma$ which is non-quadratic and whose Jacobian matrix, namely, $\Pi_0$, may have rank less than or equal to $k-1$. Note also that the parameter space $\mathbb S^k$ of $\gamma$ has empty interior. We refer the reader to ChenFang2016Rank for more detailed discussions.
\begin{rem}
The weak limit in Proposition (ref) is consistent with the one in Dovonon_Renault2013testing, when there does exist a unique common CH feature which satisfies their linear normalization and when the weighting matrix is the identity matrix (for reasons we have mentioned at the beginning of this section) -- otherwise the two are not comparable. At the first sight, our testing statistic is different from Dovonon_Renault2013testing's because we adopted a different normalization, resulting in a different parameter space.\footnote{Dovonon_Renault2013testing also recentered $Z_t$ in their construction, though this does not change the statistic numerically.} Close inspection, however, shows that the asymptotic distributions are in fact identical, up to a multiplicative constant. Specifically, let $\gamma_0$ be the (nonzero) unique CH feature such that $\sum_{j=1}^{k}\gamma_0^{(j)}=1$. Then $\Gamma_0=\{\pm\gamma_0/\|\gamma_0\|\}$ and so by Proposition 5.1, the asymptotic distribution of our $J$-statistic is simply the law of
\begin{multline}
\min_{v\in\mathbf R^k}\|\mathbb G(\pm\gamma_{0}/\|\gamma_0\|)+G\mathrm{vec}(vv^\intercal)\|^2\overset{d}{=}\|\gamma_0\|^{-4}\min_{v\in\mathbf R^k} \big\{\mathbb G(\gamma_{0})^\intercal \mathbb G(\gamma_{0})\\+\mathbb G(\gamma_0)^\intercal G\mathrm{vec}(vv^\intercal)+\frac{1}{4}(\mathrm{vec}(vv^\intercal))^\intercal G^\intercal G\mathrm{vec}(vv^\intercal)\big\} ,
\end{multline}
where we simply replaced $v$ with $v/(\sqrt{2}\|\gamma_0\|^2)$. By Theorem 3.1 and Corollary 3.1 in Dovonon_Renault2013testing-- see also Dovonon_Goncalves2017bootstrapping, their $J$-statistic (with $W$ being the identity matrix) converges in law to
\begin{align}
&\min_{u\in\mathbf R^{k-1}} \{\mathbb G(\gamma_0)^\intercal \mathbb G(\gamma_0)+\mathbb G(\gamma_0)^\intercal \bar G\mathrm{vec}(uu^\intercal)+\frac{1}{4}(\mathrm{vec}(uu^\intercal))^\intercal \bar G^\intercal \bar G\mathrm{vec}(uu^\intercal)\} ,
\end{align}
where $\bar G\in\mathbf M^{m\times (k-1)^2}$ with the $j$th row $\mathrm{vec}(A\Delta_jA^{\intercal})^\intercal$ for $A = [I_{k-1},-\jmath_{k-1}]$ and $\jmath_{k-1}$ the $(k-1)\times 1$ vector of ones. By Lemma (ref), however, the two limits in (ref) and (ref) differ only by the multiplicative constant $\|\gamma_0\|^{-4}$, establishing the claimed consistency. If the common CH feature also satisfies our normalization, i.e., $\|\gamma_0\|=1$, then the two limits are identical. We reiterate that the our main motivation is to build upon Dovonon_Renault2013testing by allowing multiple common CH features and adopting a normalization that would not falsely exclude the existence of any common features.\footnote{Any other linear normalization $c^\intercal \gamma_0=r$ for known $c\in\mathbf R^k$ and $r\in\mathbf R$ would share the same deficiency as the linear normalization, which includes, for example, $\gamma_0^{(1)}=1$ -- see our next section.} \qed
\end{rem}
\subsection{Simulation Studies}
In this section, we examine the finite sample performance of our framework based on Monte Carlo simulations, and show how the identification assumption in Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping may suffer from their linear normalization. One may then try the multiple testing versions of these tests by testing a few linearly independent linear restrictions, but we show they may be too conservative.
As in Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping, we consider the following CH factor model:
\begin{equation}
Y_{t}=\Lambda F_{t}+U_{t} ,
\end{equation}
where $Y_t$ is a $k\times 1$ vector that can be thought of asset returns, $F_t$ is a $p\times 1$ vector of CH factors, $\Lambda $ is a $k\times p$ matrix of factor loadings, and $U_t$ is a vector of idiosyncratic shocks independent of $F_t$. Following Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping, we let $\{U_t\}$ be an i.i.d.\ sequence from $N(0,I_k/2)$, and the $j$th component $f_{j,t+1}$ of $F_{t+1}$ follow a Gaussian-GARCH(1,1) model such that
\begin{equation*}
f_{j,t+1}=\sigma _{j,t}\epsilon _{j,t+1} ,\,\sigma _{j,t}^{2}=\omega_{j} +\alpha_{j} f_{j,t}^{2}+\beta_{j} \sigma _{j,t-1}^{2} ,
\end{equation*}
where $\omega_{j}, \alpha_{j}, \beta_{j} >0$, $\{\epsilon_{j,t}\}\sim N(0,1)$ i.i.d.\ across both $j$ and $t$, and $\{\sigma_{j0}\}$ are independent across $j$ and of $\{\epsilon_{j,t}\}$. It follows that $\{f_{j,t}\}$ are independent across $j$ for each $t$. The remaining specifications are detailed in Table (ref). Our designs are the same as those in Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping except that different values for $\Lambda$ are used to illustrate the restrictiveness of the linear normalization. Designs D1 and D2 generate two assets while Designs D3, D4 and D5 generate three assets. In Designs D1, D3 and D4, the factor loading matrices $\Lambda$ ensure the existence of common CH features and thus serves for investigation of size performance, while no common CH features exist in Designs D2 and D5, which help us inspect power performance.
\begin{table}[!htbp]
\caption{Simulation Designs}
\begin{small}
\begin{center}
\begin{tabular}{ccccc}
\hline\hline
Design & \# of Assets & \# of Factors & GARCH Parameters & Factor Loadings\\
\hline
D1 & $k=2$ & $p=1$ & $(\omega_1,\alpha_1,\beta_1)=(0.2,0.2,0.6)$ & $\Lambda=(1,1)^\intercal$\\
\multirow{2}{*}{D2} & \multirow{2}{*}{$k=2$} & \multirow{2}{*}{$p=2$} & $(\omega_1,\alpha_1,\beta_1)=(0.2,0.2,0.6)$ & \multirow{2}{*}{$\Lambda=I_2$}\\
& & & $(\omega_2,\alpha_2,\beta_2)=(0.2,0.4,0.4)$ & \\
D3 & $k=3$ & $p=1$ & $(\omega_1,\alpha_1,\beta_1)=(0.2,0.2,0.6)$ & $\Lambda=(1,1,1)^\intercal$\\
\multirow{2}{*}{D4} & \multirow{2}{*}{$k=3$} & \multirow{2}{*}{$p=2$} & $(\omega_1,\alpha_1,\beta_1)=(0.2,0.2,0.6)$ & \multirow{2}{*}{$\Lambda=\begin{bmatrix} 1 & 1 & 1\\ -1 & 0 & 1 \end{bmatrix}^\intercal$}\\
& & & $(\omega_2,\alpha_2,\beta_2)=(0.2,0.4,0.4)$ & \\
\multirow{3}{*}{D5} & \multirow{3}{*}{$k=3$} & \multirow{3}{*}{$p=3$} & $(\omega_1,\alpha_1,\beta_1)=(0.2,0.2,0.6)$ & \multirow{3}{*}{$\Lambda=I_3$}\\
& & & $(\omega_2,\alpha_2,\beta_2)=(0.2,0.4,0.4)$ & \\
& & & $(\omega_3,\alpha_3,\beta_3)=(0.1,0.1,0.8)$ & \\
\hline\hline
\end{tabular}
\end{center}
\end{small}
\end{table}
The tests are implemented with $m=2$ and instruments $Z_{t}=(Y_{1,t}^{2}, Y_{2,t}^{2})^\intercal$ for Designs D1 and D2, and with $m=3$ and $Z_{t}=(Y_{1,t}^{2}, Y_{2,t}^{2}, Y_{3,t}^{2})^\intercal$ for Designs D3, D4 and D5. For derivative estimation, we set the tuning parameters $\kappa_{T}=T^{-1/4}, T^{-1/3}, T^{-2/5}$ for both the derivative estimator in (ref) and the numerical derivative estimator as in (ref) respectively. These choices are meant to satisfy Assumption (ref). Again, we do not touch the issue of optimality in this paper, but instead hope to make the point that, even with these crude choices, our methods show substantial improvement over existing ones. The results corresponding to the two sets of choices are denoted as CF1 and CF2. To show the restrictiveness of the linear normalization $\gamma\in\{\gamma' \in \mathbf{R}^{k}: \sum_{i=1}^{k}\gamma _{i}'=1\}$ as in Dovonon_Renault2013testing, Dovonon_Goncalves2017bootstrapping and Lee_Liao2017LocalIDfailure, we report the results based on Dovonon_Goncalves2017bootstrapping's corrected and continuously-corrected bootstrap as well as those based on the asymptotic test of Dovonon_Renault2013testing, denoted as DG1, DG2 and DR respectively. The sample sizes are $T=1,000,$ $2,000$, $5,000,$ $10,000$, $20,000$, $40,000$ and $50,000$. To minimize the initial value effect, the data are obtained by generating $T+100$ samples and dropping the first $100$ samples. We conduct $2,000$ Monte Carlo replications with $200$ empirical bootstrap repetitions for each replication. The nominal level is $5\%$ throughout.
The results are summarized in Tables (ref)-(ref). As expected, Dovonon_Goncalves2017bootstrapping's resampling methods exhibit substantial size distortion, often close to or over $50\%$; so does the asymptotic test DR. This does not appear to be a finite sample issue as the distortion is especially severe in large samples. Rather, it is because the linear normalization excludes common CH features that actually exist in the data and in this way leads to wrong conclusions. Our tests considerably reduce the null rejection rates for all the chosen tuning parameters, though both CF1 and CF2 exhibit some degrees of over- and under-rejection, due to the issue of tuning parameters. Another interesting finding is that our bootstrap based on numerical differentiation (CF2) appears to be more sensitive to the choice of tuning parameters, which is somewhat expected because the structural method (CF1) exploits more information of the derivative. We leave a thorough comparison between these two methods for future study.
Alternatively, one may test a few linearly independent linear restrictions by adopting multiple testing versions of the DG and the DR tests, so as to avoid falsely excluding the existence of common CH features. One then rejects the existence of common CH features if {\it all} the restrictions are rejected at level $\alpha=5\%$.\footnote{Since the null is a union of “sub-nulls”, no Bonferroni-type correction is needed.} However, the resulting tests, though valid, may be too conservative. To illustrate, we test the null that $\gamma_0$ satisfies (i) $\gamma_0^{(1)}+\gamma_0^{(2)} = 1$ or (ii) $\gamma_0^{(1)} = 1$ for D1 and D2, and satisfies (i) $\gamma_0^{(1)}+\gamma_0^{(2)} + \gamma_0^{(3)} = 1$, (ii) $\gamma_0^{(1)} = 1$, or (iii) $\gamma_0^{(2)} = 1$ for D3, D4 and D5. We implement the multiple testing procedures based on Dovonon_Renault2013testing with optimal weighting matrix and Dovonon_Goncalves2017bootstrapping with the identity weighting matrix, and respectively label them as M-DG1, M-DG2 and M-DR. As expected, the M-DR test suffers from substantial under-rejection for D1, D3 and D4 even in large samples. M-DG1 and M-DG2 improve the situation somewhat, but the under-rejection is still significant for D3. Tables (ref) and (ref) indicate that our tests are more powerful than M-DG1, M-DG2 and M-DR in all cases. In particular, for D5 the rejection rates of our tests are close to one when $T$ is large while those of M-DG1 and M-DG2 are not. Results for multiple testing procedures based on Dovonon_Goncalves2017bootstrapping with optimal weighting matrix share similar patterns and are available upon request. We reiterate that the multiple testing procedure would not help with partial identification, and both Dovonon_Renault2013testing and Dovonon_Goncalves2017bootstrapping crucially rely on point identification.
{2pt}
\begin{table}[htbp]
\caption{Rejection rates under the null: Design D1}
\begin{footnotesize}
\begin{center}
\begin{tabular}{ccccccccccccc}
\hline\hline
\multirow{2}{*}{$T\backslash\text{Tests}$} & \multicolumn{3}{c}{CF1} & \multicolumn{3}{c}{CF2} & \multicolumn{4}{c}{DG}& \multicolumn{2}{c}{DR}\\
\cmidrule{2-13}
& $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & DG1 & DG2& M-DG1 & M-DG2 & DR & M-DR\\
\hline
$1000$ & 0.0850 & 0.0640 & $0.0420$& 0.0395 & $0.0185$ & 0.0100 & $0.3975$ & $0.4015$ &0.0140&0.0160&0.1740&0.0075\\
$2000$ & 0.0940 & 0.0715 & $0.0530$& 0.0550 & $0.0320$ & 0.0120 & $0.5060$ & $0.5045$ &0.0290&0.0315&0.2855&0.0125\\
$5000$ & 0.1010 & 0.0740 & $0.0515$& 0.0505 & $0.0290$ & 0.0075 & $0.6215$ & $0.6185$ &0.0485&0.0510&0.3805&0.0185\\
$10000$ & 0.1010 &0.0820 & $0.0585$& 0.0550 & $0.0285$ & 0.0090 & $0.6375$ & $0.6270$ &0.0480&0.0545&0.4005&0.0240\\
$20000$ & 0.1005 & 0.0725 & $0.0525$& 0.0495 & $0.0285$ & 0.0115 & $0.6750$ & $0.6705$ &0.0425&0.0550&0.4405&0.0225\\
$40000$ & 0.1180 & 0.0900 &$0.0670$& 0.0700 & $0.0410$ & 0.0165 & $0.6865$ & $0.6845$ &0.0635&0.0625&0.4710&0.0400\\
$50000$ & 0.1070 & 0.0830 & $0.0660$ & 0.0665 & $0.0410$ & 0.0145 & $0.6895$ & $0.6870$ &0.0425&0.0515&0.4430&0.0335\\
\hline\hline
\end{tabular}
\end{center}
\end{footnotesize}
\end{table}
{\setlength\intextsep{0.1in}
\begin{table}[htbp]
\caption{Rejection rates under the null: Design D3}
\begin{footnotesize}
\begin{center}
\begin{tabular}{ccccccccccccc}
\hline\hline
\multirow{2}{*}{$T\backslash\text{Tests}$} & \multicolumn{3}{c}{CF1} & \multicolumn{3}{c}{CF2} & \multicolumn{4}{c}{DG}& \multicolumn{2}{c}{DR}\\
\cmidrule{2-13}
& $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & DG1 & DG2& M-DG1 & M-DG2 & DR & M-DR\\
\hline
$1000$ & 0.0605 & 0.0390 & $0.0285$& 0.0660 & $0.0605$ & 0.0430 & $0.2300$&$0.2400$ &0.0025&0.0030 &0.0305 &0.0000 \\
$2000$ & 0.0645 & 0.0385 & $0.0280$& 0.0655 & $0.0570$ & 0.0380 & $0.3425$ &$0.3470$ &0.0040&0.0040 &0.0565 &0.0005 \\
$5000$ & 0.0520 & 0.0385 & $0.0315$& 0.0505 & $0.0455$ & 0.0275 & $0.3970$ &$0.3965$ &0.0025&0.0015 &0.0715 &0.0000 \\
$10000$ & 0.0690 & 0.0565 & $0.0450$& 0.0830 & $0.0665$ & 0.0320 & $0.4385$ & $0.4415$ &0.0030&0.0040 &0.0960 &0.0000 \\
$20000$ & 0.0660 & 0.0600 & $0.0490$& 0.0850 & $0.0660$ & 0.0335 & $0.4765$ &$0.4790$&0.0070&0.0065 &0.1145 &0.0005 \\
$40000$ & 0.0520 & 0.0460 &$0.0390$& 0.0645 & $0.0475$ & 0.0225 & $0.5030$ & $0.5065$ &0.0025&0.0040 &0.1175 &0.0000 \\
$50000$ & 0.0745 & 0.0670 & $0.0585$ & 0.0920 & $0.0635$ & 0.0395 & $ 0.5255$ &$0.5290$&0.0065&0.0040 &0.1540&0.0005 \\
\hline\hline
\end{tabular}
\end{center}
\end{footnotesize}
\end{table}
}
{\setlength\intextsep{0pt}
\begin{table}[htbp]
\caption{Rejection rates under the null: Design D4}
\begin{footnotesize}
\begin{center}
\begin{tabular}{ccccccccccccc}
\hline\hline
\multirow{2}{*}{$T\backslash\text{Tests}$} & \multicolumn{3}{c}{CF1} & \multicolumn{3}{c}{CF2} & \multicolumn{4}{c}{DG}& \multicolumn{2}{c}{DR}\\
\cmidrule{2-13}
& $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & DG1 & DG2& M-DG1 & M-DG2 & DR & M-DR\\
\hline
$1000$ & 0.0715 & 0.0445 & $0.0265$& 0.1305 & $0.0915$ & 0.0415 & $0.4795$ & $0.4870$ &0.0240&0.0240&0.1795&0.0010 \\
$2000$ & 0.0895 & 0.0515 & $0.0380$& 0.1485 & $0.0935$ & 0.0330 & $0.6380$ & $0.6515$ &0.0345&0.0335&0.3210&0.0055 \\
$5000$ & 0.1055 & 0.0720 & $0.0545$& 0.1590 & $0.0960$ & 0.0300 & $0.7810$ & $0.7820$ &0.0400&0.0400&0.4625&0.0075 \\
$10000$ & 0.1135 & 0.0615 & $0.0485$& 0.1440 & $0.0750$ & 0.0290 & $0.8055$ & $0.8030$ &0.0445&0.0370&0.4840&0.0080 \\
$20000$ & 0.1155 & 0.0715 & $0.0555$& 0.1530 & $0.0960$ & 0.0290 & $0.8495$ & $0.8485$ &0.0565&0.0460&0.5555&0.0170 \\
$40000$ & 0.1280 & 0.0810 &$0.0640$& 0.1655 & $0.0900$ & 0.0300 & $0.8650$ & $0.8670$ &0.0635&0.0700&0.5650&0.0145 \\
$50000$ & 0.1150 & 0.0775 & $0.0660$ & 0.1650 & $0.0855$ & 0.0260 & $0.8610$ & $0.8590$ &0.0535&0.0685&0.5980&0.0125 \\
\hline\hline
\end{tabular}
\end{center}
\end{footnotesize}
\end{table}
}
{\setlength\intextsep{0pt}
\begin{table}[!htbp]
\caption{Rejection rates under the alternative: Design D2}
\begin{footnotesize}
\begin{center}
\begin{tabular}{cccccccccc}
\hline\hline
\multirow{2}{*}{$T\backslash\text{Tests}$} & \multicolumn{3}{c}{CF1} & \multicolumn{3}{c}{CF2} & \multicolumn{2}{c}{DG}& \multicolumn{1}{c}{DR}\\
\cmidrule{2-10}
& $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$& M-DG1 & M-DG2 & M-DR\\
\hline
$1000$ &0.6450 & 0.5915 & $0.5050$& 0.7255 & $0.6890$ & 0.5570 &0.2420&0.2170&0.3740 \\
$2000$ &0.9410 & 0.9185 & $0.8805$& 0.9530 & $0.9365$ & 0.8785 &0.4935&0.3945&0.8325 \\
$5000$ &0.9975 & 0.9975 & $0.9960$& 0.9995 & $0.9990$ & 0.9950 &0.9070&0.9180&0.9940 \\
$10000$ &0.9980 & 0.9980 & $0.9975$& 0.9985 & $0.9985$ & 0.9985 &0.9995&0.9995&0.9985 \\
$20000$ &0.9985 & 0.9990 & $0.9985$& 0.9995 & $0.9995$ & 0.9985 &1.0000&1.0000&0.9985 \\
$40000$ &0.9995 & 0.9995 & $0.9995$& 1.0000 & $1.0000$ & 1.0000 &1.0000&1.0000&0.9950 \\
$50000$ &0.9995 & 0.9995 & $0.9995$& 0.9995 & $0.9995$ & 0.9995 &1.0000&1.0000&0.9995 \\
\hline\hline
\end{tabular}
\end{center}
\end{footnotesize}
\end{table}
}
{\setlength\intextsep{0pt}
\begin{table}[!htbp]
\caption{Rejection rates under the alternative: Design D5}
\begin{footnotesize}
\begin{center}
\begin{tabular}{cccccccccc}
\hline\hline
\multirow{2}{*}{$T\backslash\text{Tests}$} & \multicolumn{3}{c}{CF1} & \multicolumn{3}{c}{CF2} & \multicolumn{2}{c}{DG}& \multicolumn{1}{c}{DR}\\
\cmidrule{2-10}
& $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$ & $T^{-1/4}$ & $T^{-1/3}$ & $T^{-2/5}$& M-DG1 & M-DG2 & M-DR\\
\hline
$1000$ & 0.1240 & 0.0740 & $0.0630$& 0.3990 & $0.3645$ & 0.3000 &0.0385&0.0395&0.0140 \\
$2000$ & 0.3520 & 0.2710 & $0.2300$& 0.6975 & $0.6675$ & 0.5570 &0.1065&0.0870&0.1295 \\
$5000$ & 0.8250 & 0.7710 & $0.7255$& 0.9610 & $0.9460$ & 0.8885 &0.3470&0.3365&0.6675 \\
$10000$ & 0.9865 & 0.9850 & $0.9755$& 0.9995 & $0.9985$ & 0.9955 &0.5945&0.6765&0.9420 \\
$20000$ & 0.9980 & 0.9970 & $0.9955$& 1.0000 & $1.0000$ & 1.0000 &0.6385&0.6005&0.9665 \\
$40000$ & 1.0000 & 1.0000 & $0.9985$& 1.0000 & $1.0000$ & 1.0000 &0.7225&0.7135&0.9710 \\
$50000$ & 0.9995 & 0.9995 & $0.9990$& 1.0000 & $1.0000$ & 1.0000 &0.7755&0.7445&0.9765 \\
\hline\hline
\end{tabular}
\end{center}
\end{footnotesize}
\end{table}
}
\section{Conclusion}
In this paper, we developed a general statistical framework for conducting inference on functionals exhibiting first order degeneracy, i.e., the first order derivative of the parameter is zero. Our first contribution implies that the standard bootstrap necessarily fails to work in these settings. In light of this failure, we provided two general solutions: one generalizes the Babu correction, and the other one is a modified bootstrap following Fang_Santos2014HDD. Our framework includes many existing results as special cases. To further demonstrate the applicability of our theory, we developed a test of common CH features studied by Dovonon_Renault2013testing but under weaker assumptions that allow the existence of more than one common CH features.
\addcontentsline{toc}{section}{References}
\putbib
bibunit\setcounter{page}{1}