EconBase
← Back to paper

A Higher-Order Correct Fast Moving-Average Bootstrap for Dependent Data

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

93,615 characters

A Higher-Order Correct Fast Moving-Average Bootstrap for Dependent Data










\begin{frontmatter}

\title{\textbf{A Higher-Order Correct Fast Moving-Average Bootstrap \\ for Dependent Data}}


\author[mymainaddress]{Davide La Vecchia}

\author[mymainaddress]{Alban Moor}

\author[mysecondaryaddress]{Olivier Scaillet\corref{mycorrespondingauthor}}
\cortext[mycorrespondingauthor]{Corresponding author}
\ead{[email removed]}

\address[mymainaddress]{Research Center for Statistics, Geneva School of Economics and Management, University of Geneva,\\ Bd Pont-d’Arve 40, CH-1211 Geneva 4, Switzerland.}
\address[mysecondaryaddress]{Geneva Finance Research Institute, University of Geneva, and Swiss Finance Institute,\\ Bd Pont-d’Arve 40, CH-1211 Geneva 4, Switzerland.}


\begin{abstract}
We develop theory of a novel fast bootstrap for dependent data. Our scheme deploys
i.i.d.\ resampling of smoothed moment indicators.
We characterize the class of parametric and semiparametric estimation problems for which the method is valid.
We show the asymptotic refinements of the new procedure, proving that it is higher-order correct under mild assumptions on the
time series, the estimating functions, and the smoothing kernel. We illustrate the applicability and the advantages of our
procedure for M-estimation, generalized method of moments, and generalized empirical likelihood estimation. In a Monte Carlo study, we consider an
autoregressive conditional duration model and we compare our method with other extant, routinely-applied first- and higher-order correct methods. The results
provide numerical evidence that the novel bootstrap yields higher-order accurate confidence intervals, while remaining computationally lighter than its higher-order correct competitors. A real-data example on
dynamics of trading volume of US stocks illustrates the empirical relevance of our method.
\bigskip

\noindent \textit{JEL classification:} C12, C15, C22, C52, C58, G12.

\medskip

\end{abstract}


\begin{keyword}
 Fast bootstrap methods, Higher-order refinements, Generalized Empirical Likelihood, Confidence distributions, Mixing processes.
\end{keyword}


\end{frontmatter}

\newpage

\section{Introduction}

Inference based on first-order correct asymptotics can be misleading with confidence intervals having erratic probability coverage. It is especially true in the presence of serial dependence where first-order asymptotics often requires larger sample sizes than for i.i.d.\ data to apply. Resampling methods for time series help to obtain confidence intervals with better finite sample properties. Bootstrap methods for moment condition models have been extensively discussed under various dependence structures by, for example, \cite{hall_bootstrap_1996}, \cite{brown_generalized_2002}, \cite{inoue_bootstrapping_2006}, and \cite{davidson_bootstrap_2006}. If bootstrap methods for $m$-dependent and strongly mixing data can achieve higher-order correctness (\cite{hall_bootstrap_1996}, \cite{inoue_bootstrapping_2006}), they are computationally too intensive, once applied to heavy numerical estimation procedures. For a book-length review, see e.g.\ \cite{lahiri_resampling_2010}.

In this paper, we propose a novel fast bootstrap scheme, that we call the Fast Moving-average Bootstrap (FMB). The resampling method is computationally attractive while maintaining higher-order correctness of the inferential procedure for strongly mixing data. Our idea for building confidence regions for the parameter of interest is to realize that smoothing the moment indicators as in the Generalized Empirical Likelihood (GEL) literature permits to bootstrap them as if they were i.i.d.\ \cite{parente_generalised_2018a} study the first-order validity of GEL test statistics based on a similar bootstrapping scheme, the Kernel Block Bootstrap (henceforth KBB); see \cite{parente_kernel_2018b} and \cite{parente_quasi-maximum_2019}. Our approach differs from KBB in two significant aspects. First, our methodology does not require to solve the estimation problem at each bootstrap sample, lessening drastically the computational burden. Indeed, FMB is at least (except for a simple low-dimensional linear model) one thousand times faster, according to standard rules on bootstrap simulation errors (\cite{efron_better_1987} Section 9, \cite{davison_bootstrap_1997} Section 2.5.2). Second, we exploit an inversion technique to benefit from the kernel smoothing used in the studentization of our test statistic. The inversion is related to the standard percentile$-t$ bootstrap approach in the linear univariate case (see Example 1 below). The studentization relies on a simple sample variance of the smoothed moment
indicators, which turns out to be asymptotically equivalent to a HAC estimator for the
original moment indicators, as shown by \cite{smith_automatic_2005}.
Together these inversion and studentization make our FMB inference amenable to be shown higher-order correct. Our proof strategy is not directly applicable to KBB; its higher-order correctness remains a conjecture.

The already existing fast resampling methods usually hinge on the first-order von Mises expansion of the estimating function (\cite{shao_jackknife_1995}, \cite{davidson_bootstrap_1999}, \cite{andrews_higher-order_2002}, \cite{salibian-barrera_bootstrapping_2002}, \cite{goncalves_maximum_2004}, \cite{hong_fast_2006}, \cite{salibian-barrera_principal_2006}, \cite{salibian-barrera_fast_2008}, \cite{camponovo_robust_2012}, \cite{camponovo_predictability_2013}, \cite{armstrong_fast_2014}, and \cite{goncalves_bootstrapping_2019}).
 It yields a fast approximation, but its inherent construction does not ensure higher-order correctness. Instead, our fast method relies on inversion, namely we identify the level sets of test statistics under the null hypothesis to obtain confidence regions for the parameter of interest (see \cite{parzen_resampling_1994} and \cite{hu_estimating_2000} for i.i.d.\ data). Furthermore, the FMB confidence regions are invariant to monotonic reparameterization, due to studentization of the moment indicators. It ensures stability of our method across varying parameter scales (\cite{diciccio_bootstrap_1996}).

We design FMB for GEL estimator to exploit its intrinsic smoothing, and as it provides a considerably wide theoretical framework on semiparametric estimation (\cite{smith_gel_2011}). As a consequence, the higher-order refinements achieved by our method ensue for the Empirical Likelihood (see \cite{qin_empirical_1994}, \cite{imbens_one-step_1996}, \cite{kitamura_empirical_1997}), the Exponential Tilting (\cite{kitamura_information-theoretic_1997}, \cite{imbens_information_1998}), and the Continuously Updating Estimator (\cite{hansen_finite-sample_1996}). In addition to the KBB, other bootstrap methods already exist in the GEL literature. For instance, \cite{bravo_empirical_2004} shows the higher-order correctness of the bootstrap for inference based on empirical likelihood with i.i.d.\ data, while \cite{bravo_blockwise_2005} shows consistency of the block bootstrap for empirical entropy tests in times series regressions with strongly mixing data. However, to our knowledge, there is no proof of higher-order correctness of the bootstrap for GEL in the literature yet.

Clearly, we can also apply FMB in the setting of M-estimation (\cite{huber_robust_1964}) and Generalized Method of Moment (\cite{hansen_large_1982}), obtaining a fast version of the bootstrap methods derived in \cite{hall_bootstrap_1996} for $m$-dependent data and \cite{inoue_bootstrapping_2006} for strongly mixing data.

The structure of the paper is as follows. Section \ref{Sketch_Methodology} is a simple introduction to the FMB algorithm in the univariate case.
  There, we also discuss connections between FMB and already existing resampling schemes. In Section \ref{MParam}, we briefly present the GMM and GEL estimators for strongly mixing time series, using these frameworks as a tool to extend FMB to the multivariate setting. We itemize our assumptions and present the main theoretical results in Section \ref{Theory}.  In Section \ref{implement}, we give details on the implementation aspects of FMB, emphasizing the relation between the choice of the kernel and the properties of the long-run variance estimator. We also discuss connections with the recent literature on confidence distributions, that we use in our empirical application. We present our Monte Carlo experiments in Section \ref{Monte_Carlo}, and a real data example in Section \ref{empirical_application}. Finally, we prove our theorems in appendix. For some technical lemmas, we give the proofs in the Supplementary Material (available online).


\section{FMB methodology}\label{method}



\subsection{An introduction in the univariate case}\label{Sketch_Methodology}


Let $\left\{X_t\right\}_{t \in \mathbb{Z}}$ be a stationary strongly mixing process in $\mathbb{R}^d$, observed at $t = 1,...,T$. We assume that the time series of interest satisfies Assumptions \ref{a1}---\ref{EE5} in Section \ref{Theory}, which are standard in the bootstrap literature. Let $\mathcal{B}
\subset \mathbb{R}$ be the compact space of the parameter $\beta$ and $\mathbb{X}_t := \left\{ X_{t_1},...,X_{t_v} \right\}$ be a collection of vectors from the process $\left\{X_t\right\}_{t \in \mathbb{Z}}$. Consider the function $g: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R}$ such that:
\begin{align}
\mathbb { E } \left[ g \left( \mathbb{X} _ { t } , \beta _ { 0 } \right) \right] = 0, \label{moment}
\end{align}
where the expectation $\mathbb { E }$ is taken w.r.t.\ the true underlying distribution, unknown and depending on $\beta_0$. In the following, we use the shorthand notation $g_{t} \left( \beta \right) := g\left( \mathbb{X} _ {t}, \beta \right)$.

\noindent The function $g$ in \eqref{moment} can be the (conditional) likelihood in full parametric models, or it can be obtained using the (conditional) moments and/or may depend on instrumental variables in semiparametric models. The collection of vectors $\mathbb{X}_t$ typically contains information on the relation between the observations and the parameter characterizing the $q$-dimensional stationary distribution of a time series. More generally, we can exploit the knowledge in closed-form of the (conditional) moments to obtain (martingale) estimating functions for non-linear conditional autoregressive and  heteroscedastic
models or discretely observed diffusions. We refer
to \cite{godambe_quasi-likelihood_1987}, \cite{taniguchi_asymptotic_2000}, and \cite{kessler2012statistical} for
book-length presentations.

Each function of the sequence $\left\{ g_{t}\left( \beta \right)\right\}_{t=1}^{T}$ is often defined using the innovations, which can be i.i.d.\ random variables or more generally martingale differences. Thus, $\left\{ g_{t}\left( \beta \right)\right\}_{t=1}^{T}$ exhibits less dependence than the original process $\left\{ X_{t} \right\}_{t=1}^{T}$. Nevertheless, neglecting this temporal dependence has a serious impact on the performance of several inferential procedures, in particular it can affect the consistency of the bootstrap variance estimator and the higher-order accuracy of bootstrap confidence intervals.

To take automatically this aspect into account, we follow \cite{kitamura_information-theoretic_1997}, \cite{otsu_generalized_2006}, \cite{guggenberger_generalized_2008}, and \cite{smith_gel_2011}, and we perform
a convolution of the moment indicator $g$ with the kernel $k: \mathbb{R} \rightarrow \mathbb{R}$, obtaining:
\begin{align}
g_{T,t}\left(\beta\right) := {B_{T}}^{-1/2}\displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right) g_{t-s}\left(\beta\right), \label{convol}
\end{align}
\noindent  where $B_T$ is a bandwidth parameter, increasing in $T$ and such that $B_T / T \longrightarrow 0$. The convolution in \eqref{convol} induces a HAC-type modification, ensuring consistency of the long-run variance estimation of the mean over time
$\bar { g }_T ( \beta ) := T ^ { - 1 } \sum _ { t = 1 } ^ { T } g _ {T,t } ( \beta )$; see \cite{newey_simple_1987}, \cite{andrews_heteroscedasticity_1991}, and \cite{smith_automatic_2005}. Solving $\bar { g }_T ( \beta )=0$ gives the just-identified univariate estimator $\hat{\beta}$. Hence, the estimator $\hat{\beta}$ relies on a smoothed  moment condition. Below, we explain how we can further exploit the convolution in \eqref{convol}
to derive our bootstrap.

Let us first give the intuition of our methodology for the construction of confidence interval (CI) for $\beta_0$; more
technical aspects are available in Sections \ref{MParam} and \ref{sectionHAC}.
For ease of notation, we drop the subscript $T$ from any estimator, whenever its dependence on the sample size is clear from the context.

To keep the exposition as simple as possible, we assume temporarily a one-to-one relationship between the
parameter and the estimating function $\bar { g }_T ( \beta ),$ in an suitable subset of $\mathcal{B}$.
Even though the probabilistic validity of the FMB CI does not depend on this condition (see e.g. \cite{lehmann_testing_1959}, \cite{shao_mathematical_1999}, \cite{hansen_finite-sample_1996} and \cite{guggenberger_generalized_2008-1}), this assumption allows us to explain easily why our resampling scheme does not need to solve the estimating equation for each bootstrap sample.

Intuitively, the construction of the FMB CI goes as follows. First, our bootstrap scheme provides a higher-order correct approximation of the distribution of a statistic $\hat{S}(\beta_0)$. This statistic is an asymptotically pivotal version of the estimating function $T^{1/2}\bar { g }_T ( \beta ),$ evaluated at the true parameter $\beta_0.$ Second, the one-to-one relationship allows us to map the quantile estimates of $\hat{S}$ to CI limits in $\mathcal{B}.$ This mapping is crucial to gain computational efficiency. Indeed, we use the computationally intensive part of the FMB algorithm to compute the distribution of simple mean-type statistic $\hat{S}(\beta_0)$ which is much faster to compute than roots of $T^{1/2}\bar { g }_T ( \beta ),$ or numerical solutions to the estimating optimization problem (see Section \ref{MParam}). From an hypothesis testing point of view, FMB yields an approximation to the distribution of $\hat S(\beta)$ under $H_0:$ $\beta=\beta_0.$ Then, each $\beta$ in the CI is in the non-rejection region of $H_0$.

The next example is a widely-applied model where the one-to-one condition is satisfied, since the considered function is monotonic. Below, we explain in Remark 2 how to adapt FMB to deal with general estimating functions, which do not necessarily satisfy the monotonicity condition.
\begin{example}\label{exAR1}
We consider an $AR(1)$ process $\{Y_t\}_{t=1}^{T},$ $Y_t = \beta Y_{t-1} + \varepsilon_t,$  where $\mid \beta \mid < 1,$ and $\{\varepsilon_t\}_{t=1}^{T}$ is a white noise. The orthogonality of the innovations yields the moment indicators $g_t(\beta)=(Y_t-\beta Y_{t-1})Y_{t-1}.$ For a given kernel $k$, smoothing these moment indicators leads to
\begin{align}
g_{T,t}(\beta) = {B_{T}}^{-1/2}\displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right) Y_{t-s}Y_{t-s-1} -\beta {B_{T}}^{-1/2}\displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right)Y_{t-s-1}^2. \label{gTtAR}
\end{align}
Equation \eqref{gTtAR} is linear in the parameter of interest $\beta.$ Thus, neither taking the mean $T^{1/2}\bar{g}_T(\beta)$ nor rescaling $T^{1/2}\bar{g}_T(\beta)$ by a constant affect this linearity. As we build our statistic of interest by rescaling $T^{1/2}\bar{g}_T(\beta),$ the one-to-one condition is verified.
\end{example}

Now that the main principles of FMB are settled, we present the detailed algorithm underlying its numerical implementation. The statistic serving as basis for inference is the asymptotically pivotal quantity:
\begin{equation}
\hat{S} := T^{1/2}\bar{g}_T\left(\beta_0\right) / \hat{\sigma}, \label{Eq S_psi}
\end{equation}
where, for instance,
$
\hat{\sigma} := \kappa^2_1 (T\kappa_2)^{-1} \sum_{t=1}^{T} g^{2}_{T,t}( \hat{\beta} )
$ and $\kappa_j := \int k\left(u\right)^j du$ for $j=1,2$. This statistic is a particular value of the function $\hat{S}: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R},$ $\hat{S}(\beta):=T^{1/2}\bar{g}_T\left(\beta\right) / \hat{\sigma},$ that we suppose strictly increasing on $\mathcal{B}$. The studentization in $\hat{S}$ is crucial for FMB to be higher-order correct. In principle, we can apply other estimators of the long-run variance and we flag that each estimator $\hat{\sigma}$ has its own bias, which is going to affect the properties (e.g.\ the accuracy) of FMB. We refer to Section \ref{sectionHAC} for further discussion.

Considering an i.i.d.\ bootstrap sample drawn from $\{g_{T,t}( \hat{\beta})\}_{t=1}^{T},$ say $\{g^{\ast}_{T,t}\}_{t=1}^{T},$ the bootstrap version of $\hat{S}$ in \eqref{Eq S_psi} is:
\begin{equation}
S^{\ast}:= T^{1/2}\bar{g}_T^\ast / \hat{\sigma}^{\ast},
\label{Eq S_psi_ast}
\end{equation}
with $\bar{g}_T^\ast:=T^{-1} \sum_{t=1}^{T} g^{\ast}_{T,t}$ and $\hat{\sigma}^{\ast 2} := {T}^{-1}\sum_{t=1}^{T} g^{\ast 2}_{T,t}.$ In \eqref{Eq S_psi_ast}, both the computed numerator and denominator rely on $g^{\ast}_{T,t},$ and thus avoid re-estimating the parameter on each bootstrap sample in order to make it fast.

Then, the algorithm of our FMB is made of five steps (lines 2-4, 5, 6-13, 14-15, 16-17).
\begin{Algorithm}\label{AlgFMB}
\textnormal{\vspace{0.01cm}}
    \begin{algorithmic}[1]
            \Procedure{FMB}{$X_1, ..., X_T,$  $g,$ $\alpha$} \Comment{The FMB $(1-\alpha)$-CI for $\beta_0 \in \mathcal{B}$ s.t. $\mathbb { E } \left[ g \left( \mathbb{X} _ { t } , \beta _ { 0 } \right) \right] = 0.$}
            \For{$t = 1, ..., T$}
						    \State $g_{T,t}\left(\beta\right) \leftarrow {B_{T}}^{-1/2}\sum_{s = t-T}^{t-1} k\left(s/B_{T}\right) g_{t-s}(\beta)$
						\EndFor\label{AlgFMBconvol}
						\State \textbf{Solve}(\hspace{0.05cm}$\sum_{t=1}^T g_{T,t}( \beta) = 0$) $\rightarrow$ $\hat{\beta}$
						\For{$r = 1, ..., R$}

						    \For{$t = 1, ..., T$}
						         \State $g_{T,t}^{\ast} \leftarrow$ \textbf{Draw}($g_{T,1}(\hat{\beta}), ..., g_{T,T}(\hat{\beta})$)
						    \EndFor\label{AlgFMBSample}
						    \State $\bar{g}_{T,r}^\ast \leftarrow T^{-1} \sum_{t=1}^{T} g^{\ast}_{T,t}$
						    \State $\hat{\sigma}_r^{\ast} \leftarrow ({T}^{-1}\sum_{t=1}^{T} g^{\ast 2}_{T,t})^{1/2}$
						    \State $S^{\ast}_r \leftarrow T^{1/2}\bar{g}_{T,r}^\ast \hat{\sigma}_r^{\ast-1}$
						\EndFor\label{AlgFMBBoot}
						\State ConfidenceLevel $\leftarrow$ $1-\alpha$
						\State $q^{\ast} \leftarrow$ \textbf{Quantile}($S^{\ast}_1, ..., S^{\ast}_R,$ ConfidenceLevel)
						\State \textbf{Solve}($\hat S(\beta) =q^{\ast}$) $\rightarrow$ $UpperLimit$
            \State \textbf{return} UpperLimit \Comment{The upper limit of the one-sided FMB CI}
        \EndProcedure
    \end{algorithmic}
\end{Algorithm}
In Algorithm \ref{AlgFMB}, we exploit the monotonicity assumption only in Step 5, where we invert the studentized estimating function $\hat S(\beta)$; see \cite{hu_estimating_2000} for the use of a similar device.

Indeed, to derive the CI, the bootstrap procedure first provides $(1-\alpha)$-quantile
estimates of $\hat S(\beta_0)$, say $q^\ast_{1-\alpha}.$
Then, a numerical method (e.g. Newton-Raphson or secant methods) defines a one-sided $(1-\alpha)$-CI for $\beta_0$ as $[\beta_{\min}, \hat{q}_{1-\alpha}]$, where $\beta_{\min}:= \min\mathcal{B}$ and the upper limit $\hat{q}_{1-\alpha}$ solves $\hat S (q_{1-\alpha}) = q^\ast_{1-\alpha}$ in $q_{1-\alpha}.$

For a two-sided equal-tailed CI, we follow the same principles.
We consider two real numbers $s_1$ and $s_2$ such that $\mathbb{P}[\hat{S} \left( \beta_0 \right) \leq s_1]=\alpha/2$ and $\mathbb{P}[ \hat{S} \left( \beta_0 \right) > s_2 ]=\alpha/2$. From FMB, we obtain the approximation\footnote{As customary in the bootstrap literature, $\mathbb{P}^{\ast}[X^{\ast} \leq x ]$ denotes the empirical c.d.f.\ of any variable $X^{\ast}$ generated by the bootstrap scheme (here $S^\ast$). We give the general definition of the bootstrap probability measure $\mathbb{P}^{\ast}$ in \eqref{boot_distrib}, Section \ref{Theory}.} $\mathbb{P}^{\ast}[s_1 < S^{\ast} \leq s_2 ] = \mathbb{P}[s_1 < \hat{S} ( \beta_0 ) \leq s_2 ] + R_T$, where $R_T$ is an asymptotically  negligible remainder term. Hence, we can compute $s_1$ and $s_2$ such that $\mathbb{P}^{\ast}[s_1 < S^{\ast} \leq s_2 ] = 1 - \alpha$. Then, the CI for $\beta_0$ is $\mathcal{C}_{1-\alpha} := (c_1,c_2]$, with $c_1 := \hat{S}^{-1}\left(s_1\right)$ and $c_2 := \hat{S}^{-1}\left(s_2\right)$, ensuring that $\mathbb{P}\left[c_1 < \beta_0 \leq c_2 \right] = 1 - \alpha + R_T$. Under Assumptions \ref{a1}---\ref{EE5} (Section \ref{Theory}), we can get $R_T = o_p\left(T^{-1/2}\right),$ given a suitable choice of kernel $k$ and bandwidth $B_T$ (see Theorem \ref{higher_order} and discussion in Section \ref{Theory}). It implies that $\mathcal{C}_{1-\alpha}$ is correct up to a higher order.

From the studentization in $\hat S(\beta)$, the CI limits $c_1$ and $c_2$ derived in Step 5 remain invariant to monotonic transformation of the parameter. This property is crucial for the bootstrap CI (\cite{diciccio_bootstrap_1996}), ensuring stability of FMB across varying parameter scale.

\vspace{1cm}
\textbf{Example 1 [cont'd].}
\textit{ Let us see how the steps in Algorithm 1 specialize for the $AR(1).$ In Step 1, we have $g_{T,t}$ as in \eqref{gTtAR} and $\bar { g }_T ( \beta ) = T ^ { - 1 } \sum _ { t = 1 } ^ { T } {B_{T}}^{-1/2}\sum_{s = t-T}^{t-1} k({s}/{B_{T}}) (Y_{t-s}-\beta Y_{t-s-1})Y_{t-s-1}$. For Step 2, the estimator is available in closed form:
$$
\hat{\beta}=\left( \displaystyle \sum_{t=1}^{T} \displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right) Y_{t-s}Y_{t-s-1}\right) \left(\displaystyle \sum_{t=1}^{T} \displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right)Y_{t-s-1}^2 \right)^{-1}.
$$
For Step 3 and Step 4, we define $S^{\ast}$ using \eqref{Eq S_psi_ast}, and we use i.i.d.\ resampling of $\{g_{T,t}(\hat{\beta})\}_{t=1}^{T}$. Since $\hat S(\beta)$ is strictly decreasing, we switch the sign of $S^{\ast}$ and $\hat S(\beta)$ and proceed as in Step 5 of Algorithm 1 to build a one-sided CI for $\beta_0$. In this particular case of linear models, we can rewrite $\hat{S}(\beta_0)$ as $T^{1/2}(\hat{\beta}-\beta_0)/\hat{\varsigma},$ where $\hat{\varsigma} = \hat{W}^{-1}\hat{\sigma}$ and $\hat{W} := T^{-1}\sum_{t=1}^{T} \partial g_{T,t}(\hat{\beta})/ \partial \beta.$ Similarly, we can rewrite the bootstrap counterpart $S^{\ast}$ as $T^{1/2}(\beta^{\ast}-\hat{\beta})/\varsigma^{\ast},$ where $\beta^{\ast}$ is a bootstrap estimate, $\varsigma^{\ast} = W^{\ast-1}\sigma^{\ast }$ and $W^{\ast} := T^{-1}\sum_{t=1}^{T} \partial g_{T,t}^{\ast}(\hat{\beta})/ \partial \beta.$ Then, the FMB CI is equivalent to $[\hat{\beta}-q^{\ast}_{1-\alpha/2}\hat{\sigma} , \hat{\beta}-q^{\ast}_{\alpha/2}\hat{\sigma}],$
 where $q^{\ast}$ are quantiles of $S^{\ast}.$ In this representation, the FMB CI is similar to a percentile$-t$ bootstrap CI, up to our use of $\hat{\beta}$ instead of $\beta^{\ast}$ in the definition of $\varsigma^{\ast},$ avoiding to compute $\beta^{\ast}$, and in the multivariate case, to invert a matrix, for each bootstrap sample. This modification does not impact higher-order correctness, as shown in Section \ref{Theory}.}\\

\begin{remark}[Step 5, Lines 16-17]\label{nomono}


When the function $\hat S(\beta)$ is not monotonic, we can slightly modify the procedure if we want to obtain a simply connected CI, as opposed to a union of intervals. To this end, let us define $\hat{Q}(\beta) := \hat S(\beta)^2.$ As an alternative statistic, we take the third-order Taylor expansion of $\hat{Q}(\beta)$ around the root-$T$ consistent estimator $\hat{\beta}.$ Namely, we define $\tilde{Q}(\beta) := (\partial^2 \hat{Q}(\hat{\beta})/\partial \beta^2)(\beta-\hat{\beta})^2/2 + (\partial^3 \hat{Q}(\hat{\beta})/\partial \beta^3)(\beta-\hat{\beta})^3/6,$ the first two terms being zero.  If we make use of $\tilde{Q}(\beta)$ in a neighborhood $\{ \check{\beta} \in \mathcal{B} : \check{\beta} = \beta_0 + \bar{\delta} T^{-1/2} \}$, for  $\bar{\delta} \in \mathbb{R}$  (\cite{newey_notitle_1994}), we can show that $\tilde{Q}(\beta) = \hat{Q}(\beta) + O_p(T^{-1})$ in that neighborhood.
Hence, FMB allows us to approximate $\mathbb{P}[\tilde{Q}( \beta_0 ) \leq q^{\ast}_{1-\alpha} ]$ by $\mathbb{P}^{\ast}[S^{\ast 2} \leq q^{\ast}_{1-\alpha} ]$ with higher-order accuracy, as shown in Corollary \ref{HOtildeQ} (Section \ref{Theory}). Then, we compute
$\mathcal{C}_{1-\alpha} := \left\{ \beta \in \mathcal{B} : \tilde{Q}\left( \beta \right) \leq q^{\ast}_{1-\alpha} \right\}$ to get the desired higher-order correct $(1-\alpha)$-CI. From Section 9.1 in \cite{newey_notitle_1994}, it should be clear that $\bar \delta$ is not a tuning parameter to be chosen to apply FMB.

This modified FMB CI is simply connected with high probability when $T$ is large enough. Indeed, we can show that $\tilde{q}=\frac{4T}{27\sigma^2}\left[\frac{\partial\bar{g}_T(\hat{\beta})^2}{\partial \beta}/\frac{\partial^2\bar{g}_T(\hat{\beta})}{\partial \beta^2}\right]^2$ is the highest value such that the sublevel set $\left\{ \beta \in \mathcal{B} : \tilde{Q}\left( \beta \right) \leq \tilde{q} \right\}$ is still simply connected. It corresponds to the local maximum of a cubic polynomial (we provide a graphical illustration in Figure \ref{nomonoC}, Supplementary Material SM.9). Thus, the range of confidence level from which we can draw a simply connected set increases proportionally to the sample size $T.$ In practice, we recommend to try using the $\tilde{Q}$ approximation to guarantee the second-order correctness of the CI and connected CI with high probability. If this approach yields a disconnected CI because of a too small sample size $T$, the user can truncate the Taylor approximation at the quadratic term, which gives a simply connected CI w.p.1.\ This quadratic approximation does not guarantee higher-order correctness, but is prone to work better than the Gaussian approximation in practice. To summarise, we face three possibilities. If we use $\hat{Q}$, we always get higher-order correctness, but not necessarily connected CI when monotonicity is not satisfied. If we use the cubic approximation $\tilde{Q}$, we get higher-order correctness and connected CI with high probability. If we use a quadratic truncation, while being asymptotically correct, higher-order correctness might be lost, but we ensure connected CI.

Therefore, we conclude that, even if the monotonicity condition is violated (or is not easy to check), we can choose a convenient statistic based on a cubic approximation and preserving the asymptotic properties of FMB. We point out that the proposed derivation of the CI only involves the (potentially numerical) computation of $\hat{Q}'$s derivatives, whose evaluation is required only at the single point $\hat{\beta}$. As the shape of the CI is fully determined by these derivatives, it simplifies and speeds up the implementation of Step 5.
\end{remark}

Some further remarks on the other steps Algorithm \ref{AlgFMB} are in order. First, Step 1 --- Step 3 (Lines 2-13) hinge on bootstrapping the moment indicator evaluated at $\hat\beta$. It justifies the adjective ``fast'' in the name of our resampling scheme, and bears some similarities to the already existing  fast bootstrap  (henceforth FB) methods; see \cite{shao_jackknife_1995}, \cite{davidson_bootstrap_1999}, \cite{andrews_higher-order_2002}, \cite{salibian-barrera_bootstrapping_2002}, \cite{goncalves_maximum_2004}, \cite{salibian-barrera_principal_2006}, \cite{salibian-barrera_fast_2008}, \cite{camponovo_predictability_2013}, \cite{armstrong_fast_2014}, \cite{goncalves_bootstrapping_2019}), and to the estimating function bootstrap (\cite{parzen_resampling_1994}, \cite{hu_estimating_2000}). However, FB methods typically rely on a first order von Mises expansion, which approximation error prevents the FB to be higher-order accurate.

Second, Step 3 (Line 12) computes the bootstrap statistic $S^{\ast},$ where the kernel $k$ creates a block of moment indicators evaluated at $\hat\beta$. The block of $g_t$ induced by the kernel is similar to a moving-average, as we emphasize in the name of our resampling scheme. The Moving Block Bootstrap (henceforth MBB) is the state-of-the-art higher-order correct alternative to FMB (\cite{gotze_second-order_1996}, \cite{lahiri_edgeworth_1996}). For MBB, the blocks are defined at the level of the observations, whereas in our case the convolution is applied to the moment indicators.

In the same family of groupwise resampling schemes, FMB is even more reminiscent of the Tapered Block Bootstrap of \cite{paparoditis_tapered_2001} (hereafter TBB), in the sense that we can view their tapered block as our moving-average kernel. The main difference is that our kernel has unbounded support, in contradistinction with their block tapering window. It gives FMB an advantage in the studentization: it allows us to use the Quadratic Spectral (QS) kernel, which is optimal in terms of asymptotic mean squared error according to \cite{andrews_heteroscedasticity_1991}. \cite{parente_kernel_2018b} have already pointed out such an advantage for a KBB variance estimator. Yet, the TBB and the KBB approach of \cite{parente_generalised_2018a} both require $R$ bootstrap estimations. Thus, neither the TBB nor the KBB is fast and there is no result on their potential higher-order correctness. The higher-order correctness of the FMB approximation to the distribution of $\hat{S}$ comes from jointly considering two ingredients: the smoothing of moment indicators and the studentization.\footnote{The higher-order correctness of the FMB CI comes from the higher-order correctness of the latter FMB distribution and from the inversion step (Step 5 with the monotonicity condition or the cubic approximation of Remark \ref{nomono}).} Taken in isolation, each ingredient does not allow to show higher-order correctness of the FMB CI.

To summarize, we itemize in Table \ref{RM} the main features of the discussed bootstrap schemes. We only list methodologies that are specifically designed for dependent data.

\begin{table}[h]
\caption{Properties of related bootstrap schemes for dependent data.}
\begin{center}
\begin{tabular}{|l||*{2}{c|}}\hline
\backslashbox{Fast}{HOC}
&\makebox[3em]{Yes}&\makebox[4.5em]{No}\\\hline\hline
\makebox[3em]{Yes} &\makebox[3em]{FMB}&\makebox[4.5em]{FB}\\\hline
\makebox[3em]{No} &\makebox[3em]{MBB}&\makebox[4.5em]{TBB, KBB}\\\hline
\end{tabular}
\label{RM}
\caption*{

We distinguish the Fast Moving-average Bootstrap (FMB), the Fast Bootstrap (FB), the Moving Block Bootstrap (MBB), the Tapered Block Bootstrap (TBB), and the Kernel Block Bootstrap (KBB) with respect to two features, namely computational speed (Fast) and higher-order correctness (HOC).}
\end{center}
\end{table}

\subsection{Over-identified case with multivariate parameter}\label{MParam}

In this section, we explain how FMB can yield higher-order correct inference on a multivariate parameter. In Subsection \ref{Sec: GMM}, we consider the Generalized Method of Moments. Then, we extend the setting to Generalized Empirical Likelihood Estimation in Subsection \ref{Sec: GEL}. The asymptotic refinements (Section \ref{Theory}) also hold for the standard GMM case and are not tied to the use of GEL.

\subsubsection{Generalized Method of Moments} \label{Sec: GMM}

Assume we have to conduct inference on the multivariate parameter $\beta \in \mathcal{B} \subset \mathbb{R}^p$, where $\mathcal{B}$ is compact. We are given a random sample of $\mathbb{X}_t = \left\{ X_{t_1},...,X_{t_v} \right\}$ observed at $t=1,...,T,$ and we define a set of moment conditions $g: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R}^r$, with $r \geq p$, such that $\mathbb{E}\left[g(\mathbb{X}_t,\beta_0)\right] = 0$. To handle the serial dependence, we define the smoothed moment indicator $g_{T,t}\left(\beta\right)$ as in the univariate case
\begin{align*}
g_{T,t}\left(\beta\right) := {B_{T}}^{-1/2}\displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right) g_{t-s}\left(\beta\right),
\end{align*}
\noindent and $\bar{g}_T(\beta) := T^{-1}\sum_{t=1}^{T}g_{T,t}(\beta).$ Estimating $\beta_0$ via Generalized Method of Moments  (\cite{hansen_large_1982}) is the most popular approach in econometrics. In the next subsection, we discuss alternative estimators.

Step 1 --- Step 3 of the FMB Algorithm \ref{AlgFMB} remain conceptually
unchanged. As far as the bootstrap statistic is concerned, the principles of Step 4 and Step 5 stay the same as in Algorithm \ref{AlgFMB},
the main change being that the asymptotically pivotal statistic becomes:
\begin{align}
\hat Q  := T \bar{g}_{T} \left( \beta_0 \right) ^\intercal \hat{\Omega}^{-1} \bar{g}_{T} \left( \beta_0 \right),\label{Rao_type}
\end{align}
where $\hat{\Omega} = \kappa_1^2(\kappa_2 T)^{-1}\sum_{t = 1}^{T} \{ g_{T,t}( \hat \beta ) - \bar{g}_T( \hat \beta ) \} \{ g_{T,t}( \hat \beta ) - \bar{g}_T( \hat \beta )\}^{\intercal}$ is a consistent estimator of the long-run covariance matrix of $T^{1/2}\bar{g}_{T} \left( \beta_0 \right)$, of rank $\nu = r$;\footnote{If the rank is lower than $r$, the covariance matrix is not invertible anymore and we resort to the generalized inverse, adapting the degrees of freedom of the $\mathcal{X}^2$ distribution accordingly (\cite{moore_generalized_1977}).} see Section \ref{sectionHAC}. Standard results guarantee that $\hat Q$ is asymptotically $\mathcal{X}^2_{\nu}$. As in the univariate case, the statistic of interest is a particular value of a function, here $\hat{Q}: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R},$
\begin{equation}
\hat{Q}(\beta) :=  T \bar{g}_{T} \left( \beta \right) ^\intercal \hat{\Omega}^{-1} \bar{g}_{T} \left( \beta \right). \label{Eq. hatQ}
\end{equation}

\noindent We define the GMM estimator as $\hat{\beta} = \operatorname*{argmin}_{\beta \in \mathcal{B}} \hat{Q}(\beta).$ Similarly to the argument of Remark \ref{nomono}, we also define the cubic approximation centered on the root-$T$ consistent estimator $\hat{\beta}:$
\begin{align}
\tilde{Q}(\beta) := \hat{Q}(\hat{\beta}) + (\beta-\hat{\beta})^{\intercal} \hat{H} (\beta-\hat{\beta}) /2 + \left( (\beta-\hat{\beta}) \otimes (\beta-\hat{\beta}) \right)^{\intercal} \frac{\partial \operatorname*{vec} (\hat{H})}{\partial \beta^{\intercal}} (\beta-\hat{\beta}) /6,
\label{tildeQStat}
\end{align}
\noindent where the matrix $\hat{H} := \partial^2 \hat{Q}(\hat{\beta}) / \partial \beta^{\intercal} \partial \beta$. It allows us to build simply connected level sets, yielding higher-order correct Confidence Region (henceforth CR), as shown in Corollary \ref{HOtildeQ} (Section \ref{Theory}).

The bootstrap version of $\hat Q$ is
\begin{align}
Q^{\ast} := T^{-1} \sum_{t=1}^{T} \left\{ g^{\ast}_{T,t} - \bar{g}_T\left( \hat\beta \right) \right\} ^\intercal \hat{\Omega}^{\ast-1} \sum_{t=1}^{T} \left\{ g^{\ast}_{T,t} - \bar{g}_T\left( \hat\beta \right) \right\}\label{Rao_type_ast},
\end{align}
where
$
\hat{\Omega}^{\ast} := {T}^{-1}\sum_{t = 1}^{T} \left\{ g^{\ast}_{T,t} - \bar{g}_T\left( \hat\beta \right) \right\} \left\{ g^{\ast}_{T,t} -\bar{g}_T\left( \hat\beta \right)\right\}^{\intercal},
$
with the asterisk denoting the same i.i.d.\ resampling scheme as in Algorithm \ref{AlgFMB}. Now we are ready to state the algorithm of our FMB in the over-identified case, made of five steps (lines 2-4, 5, 6-13, 14-15, 16-17).
\begin{Algorithm}\label{Alg2FMB}
\textnormal{\vspace{0.01cm}}
    \begin{algorithmic}[1]
            \Procedure{FMB}{$X_1, ..., X_T,$  $g,$ $\alpha$} \Comment{The FMB $(1-\alpha)$-CR for $\beta_0 \in \mathcal{B}$ s.t. $\mathbb { E } \left[ g \left( \mathbb{X} _ { t } , \beta _ { 0 } \right) \right] = 0.$}
            \For{$t = 1, ..., T$}
						    \State $g_{T,t}\left(\beta\right) \leftarrow {B_{T}}^{-1/2}\sum_{s = t-T}^{t-1} k\left(s/B_{T}\right) g_{t-s}(\beta)$
						\EndFor\label{Alg2FMBconvol}
						\State \textbf{Argmin}$_{\beta \in \mathcal{B}} \hat{Q}(\beta)$ $\rightarrow$ $\hat{\beta}$ \Comment{Using the function $\hat{Q}(\beta)$ as in \eqref{Eq. hatQ}.}
						\For{$r = 1, ..., R$}

						    \For{$t = 1, ..., T$}
						         \State $g_{T,t}^{\ast} \leftarrow$ \textbf{Draw}($g_{T,1}(\hat{\beta}), ..., g_{T,T}(\hat{\beta})$)
						    \EndFor\label{Alg2FMBSample}
						    \State $\bar{g}_{T,r}^\ast \leftarrow T^{-1} \sum_{t=1}^{T} \left\{ g^{\ast}_{T,t} - \bar{g}_T\left( \hat\beta \right) \right\}$
						    \State $\hat{\Omega}_r^{\ast} \leftarrow {T}^{-1}\sum_{t = 1}^{T} \left\{ g^{\ast}_{T,t} - \bar{g}_T\left( \hat\beta \right) \right\} \left\{ g^{\ast}_{T,t} -\bar{g}_T\left( \hat\beta \right)\right\}^{\intercal}$
						    \State $Q^{\ast}_r := T \bar{g}_{T,r}^{\ast \intercal} \hat{\Omega}_r^{\ast-1} \bar{g}_{T,r}^\ast$
						\EndFor\label{Alg2FMBBoot}
						\State ConfidenceLevel $\leftarrow$ $1-\alpha$
						\State $q^{\ast} \leftarrow$ \textbf{Quantile}($Q^{\ast}_1, ..., Q^{\ast}_R,$ ConfidenceLevel)
            \State $\mathcal{C}$ $\leftarrow$ $\left\{ \beta \in \mathcal{B} : \hat{Q}\left( \beta \right) \leq q^{\ast} \right\}$
						\State \textbf{return} $\mathcal{C}$
        \EndProcedure
    \end{algorithmic}
\end{Algorithm}
A few remarks are in order. Step 3 uses \eqref{Rao_type_ast}, where we recenter the bootstrap statistic. Indeed, the bootstrap expectation $\mathbb{E}^{\ast}\left[g^{\ast}_{T,t}\right] = \bar{g}_T(\hat{\beta}) \neq 0$ in the over-identified case. Thus, we subtract its expectation from $g_{T,t}^{\ast}$ to recenter the bootstrap variable. This operation is crucial to achieve higher-order accurate CR in Step 5 of Algorithm \ref{Alg2FMB} (see e.g.\ \cite{hall_bootstrap_1996}).

Moreover, in contradistinction with the already existing FB methods, Step 3---4 mimic the variability of the covariance estimator $\hat{\Omega}$ (in \eqref{Rao_type}) to achieve higher-order refinements. To this end, we use $\hat{\Omega}^{\ast}$ instead of $\hat{\Omega}$ in the definition of $Q^{\ast}$ (in \eqref{Rao_type_ast}), such that we randomize the bootstrap covariance estimator across the different bootstrap samples, and do not keep it fixed at $\hat{\Omega}$. Similar comment applies to Algorithm \ref{AlgFMB}.

Finally, to define the CR for $\beta_0$, we proceed similarly to Step 5 of Algorithm \ref{AlgFMB}. We set $q^{\ast}_{1-\alpha}$ such that $\mathbb{P}^{\ast}[ Q^{\ast} \leq q^{\ast}_{1-\alpha} ] = 1-\alpha$ and compute the CR as the subset $\mathcal{C}_{1-\alpha} := \left\{ \beta \in \mathcal{B} : \hat{Q}\left( \beta \right) \leq q^{\ast}_{1-\alpha} \right\}$. Thus, we get $\mathbb{P}[ \beta_0 \in \mathcal{C}_{1-\alpha} ] = \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq q^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq q^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T$. Under Assumptions \ref{a1}---\ref{EE5} (Section \ref{Theory}), we show in Theorem \ref{higher_order} that the remainder $R_T$ can be at most of order $o_p\left(T^{-1/2}\right),$ given a suitable choice of kernel $k$ and bandwidth $B_T$, which implies that $\mathcal{C}_{1-\alpha}$ is correct up to a higher order.

\begin{remark}[Step 5, Lines 16-17]\label{nomono2}

If the higher-order correctness of FMB is guaranteed independently of the CR shape, they are generally not elliptical and they might fail to be simply connected (they can come in several pieces or contain holes). Although this irregularity is customary for small to moderate sample sizes, it can still make the results difficult to interpret. The monotonicity condition aforementioned is one of the possible ways out. As a more general solution, we propose to use the cubic approximation $\tilde{Q}$ (as in \eqref{tildeQStat}). Similarly to Remark \ref{nomono}, we have $\tilde{Q}(\beta) = \hat{Q}(\beta) + O_p(T^{-1})$ in a neighborhood $\{ \check{\beta}  \in \mathcal{B} : \check{\beta} = \beta_0 + \bar{\delta} T^{-1/2}\},$ for $ \bar{\delta} \in \mathbb{R}^p.$  Thus, defining the modified FMB CR as $\mathcal{C}_{1-\alpha} := \left\{ \beta \in \mathcal{B} : \tilde{Q}\left( \beta \right) \leq q^{\ast}_{1-\alpha} \right\}$ preserves higher-order correctness, as shown in Corollary \ref{HOtildeQ}. The resulting CR is not simply connected for all sample sizes, but we can show that the range of confidence level $(1-\alpha)$ leading to a simply connected CR is proportional to the sample size $T.$ It is not an asymptotic property and the desired CR can already be simply connected for a small sample size, depending on the second derivative of the moment indicator $g.$ If the sample size is too small for this modified CR to be simply connected and the user thoroughly needs this property, we recommend to truncate the $\tilde{Q}$ approximation at the second term. The resulting CR is always simply connected and elliptical, but, while being asymptotically correct, we cannot guarantee its higher-order properties.



\end{remark}

\subsubsection{Generalized Empirical Likelihood Estimation} \label{Sec: GEL}

FMB is not tied to a particular estimation method. In order to show its wide applicability, we consider a general estimation method, which includes (among others) a version of GMM. The Continuously Updating Estimator (CUE) (\cite{hansen_finite-sample_1996}), Empirical Likelihood (EL) (\cite{qin_empirical_1994}, \cite{imbens_one-step_1996}, \cite{kitamura_empirical_1997}), and the Exponential Tilting (ET) (\cite{kitamura_information-theoretic_1997}, \cite{imbens_information_1998}) are all asymptotically equivalent to the (2S)GMM, but they tend to be less biased for small to moderate sample sizes (see \cite{altonji_small-sample_1996} for a Monte Carlo exploration, and \cite{newey_higher_2004}, \cite{anatolyev_gmm_2005} for theoretical insights). Putting EL, ET, and CUE  under the same umbrella, \cite{smith_gel_2011} introduces the Generalized Empirical Likelihood (GEL) criterion for time series data. We briefly describe it before extending our FMB to this general setting.

Let $\rho \left( \nu \right)$ be a concave function on an open interval $\mathcal{V} \in \mathbb{R}$ containing $0$.
Writing $\rho_{\iota}(\nu) := \partial^{\iota} \rho (\nu) / \partial \nu^{\iota}$
and $\rho_\iota = \rho_{\iota}(0)$ for $\iota=0,1,$ the function
$\rho \left( \nu \right)$ is standardized such that $\rho_1 = -1$.
Defining a vector of auxiliary parameters $\lambda \in \Lambda_ { T } ( \beta )$ with $\Lambda_ { T } ( \beta ) := \left\{ \lambda \in
\mathbb{R} ^ { r } : \kappa B_{T}^{-1/2} \lambda ^ { \intercal } g _ { T,t } ( \beta ) \in \mathcal{V} \right\},$ and $\kappa := \kappa_1 / \kappa_2$, \cite{smith_gel_2011} defines the GEL criterion as:
\begin{align}
\hat{P}(\beta,\lambda) = T^{-1} \sum_{t=1}^{T} \left[ \rho \left( \kappa B_{T}^{-1/2} \lambda ^\intercal g_{T,t} \left( \beta \right) \right) - \rho_0 \right].\label{GEL_criterion}
\end{align}
\noindent To derive an estimator of $\beta$, we first optimize criterion \eqref{GEL_criterion} w.r.t.\ $\lambda$ for a given $\beta$, so that $\lambda \left( \beta \right) = \operatorname*{argsup}_{\lambda \in \Lambda_ { T } ( \beta )} \hat{P}(\beta,\lambda)$. Then, we define $\hat{\beta}$ as the solution to $\operatorname*{argmin}_{\beta \in \mathcal{B}} \hat{P}(\beta,\lambda\left( \beta \right))$.

Similarly to \cite{khundi_edgeworth_2012} and \cite{lee_asymptotic_2016}, we are going to use the first order condition of the GEL criterion as a just-identified representation of the estimation problem to explain how to build the FMB statistic. Differentiating \eqref{GEL_criterion} w.r.t.\ to $\lambda$ and $\beta$, we obtain:
\begin{align}
&T^{-1}\displaystyle\sum_{t=1}^{T} \rho_{1} \left( \kappa B_{T}^{-1/2} \lambda \left( \beta \right) ^\intercal g_{T,t}\left( \beta \right) \right) B_{T}^{-1/2} {g}_{T,t}\left( \beta \right) = 0,\label{GELFOC1}
\\
&T^{-1}\displaystyle\sum_{t=1}^{T} \rho_{1} \left( \kappa B_{T}^{-1/2} \lambda \left( \beta \right) ^\intercal g_{T,t} \left( \beta \right) \right) B_{T}^{-1/2}\frac{ \partial g_{T,t} \left( \beta \right)}{\partial \beta }^\intercal \lambda \left( \beta \right) = 0.\label{GELFOC2}
\end{align}
\noindent We can see from \eqref{GELFOC1} that $\rho_{1} \left( \kappa B_{T}^{-1/2} \lambda \left( \beta \right) ^\intercal g_{T,t}\left( \beta \right) \right)$ gives weights to the observations such that the moment conditions in $g$ are always enforced in a given sample. GEL estimators are equivalent to some minimum discrepancy estimators based on the power-divergence family (see \cite{cressie_multinomial_1984}), where the auxiliary vector parameter $\lambda$ corresponds to the Lagrange multiplier enforcing this empirical moment condition. Thus, if the original moment conditions in $g$ are correctly specified, the true (long-run) Lagrange multiplier is zero ($\lambda_0 = 0$). Applying our FMB to this setting only requires to define a quadratic statistic from the GEL first order conditions (\eqref{GELFOC1} and \eqref{GELFOC2}) evaluated at the true value of the parameters. Yet, replacing $\lambda = \lambda_0 = 0$ and $\beta = \beta_0$ in \eqref{GELFOC1} and \eqref{GELFOC2} boils down to the original $\bar{g}_T(\beta_0).$ Thus, the natural extension of FMB for GEL estimators requires to take the asymptotically pivotal statistic $\hat{Q}(\beta_0)$ as in the GMM case \eqref{Rao_type}. As a consequence, FMB in the GEL setting is exactly the same as in the GMM setting (see Algorithm \ref{Alg2FMB}), up to the initial estimator $\hat{\beta}.$  It is quite  intuitive since, in absence of misspecification, the first order conditions of the GEL criterion convey the same information on $\beta_0$ as the moment condition $g$.

\section{Theory}\label{Theory}

In the next theorems, we state that FMB CI and CR are higher-order correct. By construction, the higher-order correctness of the FMB CI and CR entirely hinges on our bootstrap approximation of the distribution for the test statistics $\hat{S}(\beta_0)$ (as in \eqref{Eq S_psi}) and $\hat{Q}(\beta_0)$ (as in \eqref{Rao_type}).

We start by itemizing the assumptions and regularity conditions. In the following, we keep using the shorthand notations $g_t = g_t\left(\beta_0\right)$ and $g_{T,t} = g_{T,t}(\beta_0).$ For any vector $V \in \mathbb{R}^n$, we write $\| V \| = (v_1^2 + ... + v_n^2)^{1/2}$, where $v_j$ is the $j$-th element of $V$. We make use of generic constants $C$, $\delta$ and $\epsilon$, whose value can differ from an expression to another. We define $\left\{g_t\right\}_{t \in \mathbb{Z}}$ on the probability space $\left(\Omega,\mathcal{A},\mathbb{P}\right)$. Let $\left\{ \mathcal{D}_t \right\} _ {t \in \mathbb{Z}}$ be a given sequence of sub-sigma-fields of $\mathcal{A}$, and $\mathcal{D}_{a}^{b} = \sigma\left\langle\left\{\mathcal{D}_{j} : a \leq j \leq b \right\}\right\rangle$. A straightforward example is to take $\mathcal{D}_t:=\sigma\left\langle g_t \right\rangle,$ but it is not always the most efficient choice to check the assumptions below (see \cite{gotze_asymptotic_1983} and \cite{gotze_asymptotic_1994} for practical examples). The higher-order correctness of FMB is subject to the following conditions (\cite{gotze_second-order_1996}, \cite{lahiri_resampling_2010}), which we assume to hold for $\left\{g_t\right\}_{t \in \mathbb{Z}}$:

\begin{asu}\label{a1}
$\mathbb{E}\left[g_{t}(\beta)\right] = 0, t=1,2,...$ only for $\beta = \beta_0$ on the compact parameter space $\mathcal{B}$. Moreover, $\hat{\beta} \overset{a.s.} \rightarrow \beta_0$ and $\|\hat{\beta}-\beta_0\|=O_p(T^{-1/2})$ as $T \rightarrow \infty$.
\end{asu}
\begin{asu}\label{a2}
$\mathbb{E}\left[\left\| g_{t} \right\|^{s+\delta}\right] < \infty \text{ for a positive } s \geq 8,\text{ }t=1,2,...,\text{ and } \delta > 0.$
\end{asu}
\begin{asu}\label{RV_approx}
There exists a constant $\delta > 0$ such that for $t,m = 1,2,...$ and $m > \delta^{-1}$, we can approximate $g_{t}$ by a $\mathcal{D}_{t-m}^{t+m}$-measurable random vector $g_{t,m}^{\ddagger}$, such that $\mathbb{E}\left\| g_t - g_{t,m}^{\ddagger} \right\| \leq \delta^{-1} \exp\left(-\delta m\right)$.
\end{asu}
\begin{asu}\label{Rosenblatt_mix}
There exists a constant $\delta > 0$ such that
$\alpha\left(m\right) = \sup \left\lvert \mathbb{P}\left[A \cap B\right] - \mathbb{P}\left[A\right]\mathbb{P}\left[B\right]\right\rvert \leq \delta^{-1} \exp\left(-\delta m\right)$, for all $t,m = 1,2,... $ and $A \in \mathcal{D}_{-\infty}^{t}, B \in \mathcal{D}_{t+m}^{\infty}$.
\end{asu}
\begin{asu}\label{Condition_Cramer}
There exists a constant $\delta > 0$ such that for all $t,m = 1,2,...$, $\delta^{-1} < m < t$, and all $\tau \in \mathbb{R}^r$ with $\| \tau \| \geq \delta$,
$\mathbb{E}\left[\left\lvert\mathbb{E}\left[\exp\left(i\tau ^{\intercal} \left(g_{t-m} + ... + g_{t+m}\right) \right) \middle| \mathcal{D}_k : k \neq t\right]\right\rvert\right] \leq \exp\left(-\delta\right)$, and
\begin{align}
\displaystyle\liminf_{T \rightarrow \infty} \frac{1}{T} \operatorname*{Var} \left[ \displaystyle\sum_{t=1}^{T} g_t \right] > 0.
\end{align}
\end{asu}
\begin{asu}\label{Markov_type}
There exists a constant $\delta > 0$ such that for all $t,m,p = 1,2,...$ and $A \in \mathcal{D}_{t-p}^{t+p}$, \newline
$\mathbb{E}\left[\left\lvert \mathbb{P}\left[A\middle|\mathcal{D}_k: k \neq t\right] - \mathbb{P}\left[A\middle|\mathcal{D}_k: 0 < \left\lvert t-k \right\rvert \leq m + p\right]\right\rvert\right] \leq \delta^{-1} \exp\left(-\delta m\right)$.
\end{asu}
\begin{asu}\label{EE5}
Defining $f(\beta):=\limsup_{T \rightarrow \infty} \sup_{b < \| \tau \| < e^{\delta T}} \lvert T^{-1} \displaystyle\sum_{t=1}^{T} \exp \left( i  \tau^{\intercal} g_{T,t}\left(\beta\right) \right) \rvert,$ there exist constants $b$ and $\delta>0$ such that $f(\beta_0)<1$ and is continuous at $\beta_0$ a.s.
\end{asu}
Assumption \ref{a1} is an identification condition. It is necessary also because we evaluate the moment conditions at $\hat{\beta}$ in the bootstrap samples. Since we are going to prove the validity and higher-order correctness of FMB using  the $(s-2)$-th order Edgeworth expansion for the mean of \cite{gotze_asymptotic_1994}, we require the moments in Assumption \ref{a2} to be defined. Assumption \ref{RV_approx} ensures that the process $\{g_t\}$ is close enough to another process $\{g_{t,m}^{\ddagger}\}$, measurable w.r.t.\ sub-sigma-fields belonging to the sequence $\{\mathcal{D}_t\}$, whose dependence structure is controlled by the mixing condition in Assumption \ref{Rosenblatt_mix}. Assumption \ref{Condition_Cramer} is the conditional Cramér condition of \cite{gotze_asymptotic_1994} for weakly dependent process. Assumption \ref{Markov_type} ensures that we can approximate the probability of $A \in \mathcal{D}_{t-p}^{t+p}$ given $\{ \mathcal{D}_k: k \neq t \}$ with an exponentially increasing accuracy, as the information in $\{ \mathcal{D}_k: 0 < \left\lvert t-k \right\rvert \leq m + p \}$ increases with $m$. We need Assumption \ref{EE5} on the regularity of the bootstrap characteristic function in order for the appropriate Cramér condition to hold for $S^{\ast}$ and $Q^{\ast}$. We refer to \cite{gotze_asymptotic_1983} for a general overview of the processes in agreement with our assumptions. As an example, the OLS moment indicators for the autoregressive parameters of a stationary AR($p$) process $Y_t = \sum_{i=1}^{p} \theta_i Y_{t-i} + e_t = \sum_{j=0}^\infty w_j e_{t-j}$ satisfy Assumptions \ref{a1}---\ref{EE5}, when $e_t \overset{i.i.d.} \sim \mathcal{N}(0,\sigma^2)$ and $\lvert w_j \rvert \leq \delta^{-1} \exp(- \delta j)$ $\forall j \in \mathbb{N},$ $\delta >0$. We verify the assumptions for this example in the Supplementary Material (SM.10). We can check along the same lines that an ACD($v$,0) model, and in particular the ACD(1,0) used in our Monte Carlo simulations, satisfies Assumptions \ref{a1}---\ref{EE5}.

Under the stated assumptions, we prove (see Appendix) the following two theorems, in which we denote by $\mathbb{P}^\ast$ the bootstrap probability measure given the data:
\begin{align}
\mathbb{P}^{\ast}(A) := T^{-1}\sum_{t=1}^{T}\mathbb{I}_{A}\{g_{T,t}(\hat{\beta})\},\label{boot_distrib}
\end{align}
where $A$ is a set and $\mathbb{I}_{A}\{V\} = 1$ if $V \in A$ and 0 otherwise.
\begin{theorem}\label{higher_order}
Under Assumptions \ref{a1}---\ref{EE5}, $\mathbb{E} \left[ \left\| g_t \right\|^{\bar q s + \delta} \right] < \infty$, for $\delta > 0$, $s \geq 8$, and $\bar q \geq 3$, $\log T = o(B_T)$, we have (i) for $S^{\ast}$ as in \eqref{Eq S_psi_ast} and $\hat{S}_{}$ as in \eqref{Eq S_psi}:
$$\sup_{x \in \mathbb{R}}\left\lvert \mathbb{P}^\ast\left[S^{\ast} \leq x\right] - \mathbb{P}\left[\hat{S}_{} \leq x\right]\right\rvert = o_p\left(T^{-1/2}\right) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right),
$$
and (ii) for $Q^{\ast}$ as in \eqref{Rao_type_ast} and $\hat{Q}$ as in \eqref{Rao_type}: $$\sup_{x \in \mathbb{R}^{+}}\left\lvert \mathbb{P}^\ast[ Q^{\ast} \leq x] - \mathbb{P}[\hat{Q} \leq x]\right\rvert = o_p(T^{-1/2}) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right),$$
where $q$ is the Parzen exponent of $k^{\ast}$,  where $k^{\ast}\left( a \right) := \kappa_2^{-1} \int{k\left( b-a \right)k\left( b \right)db}$.
\end{theorem}
It is now apparent that the kernel $k$ impacts the bootstrap accuracy through the variance estimators in $\hat{S}$ and $\hat{Q}.$ As discussed by \cite{parzen_consistent_1957} and \cite{andrews_heteroscedasticity_1991}, the bias of this estimator depends on the smoothness of the kernel $k^{\ast}$ at zero. The kernel $k^{\ast}\left( a \right) := \kappa_2^{-1} \int{k\left( b-a \right)k\left( b \right)db}$ is induced by the self-convolution of the smoothing kernel $k$. This bias is of order $O(B_T^{-q})$, where $q$ is the Parzen exponent of $k^{\ast}$, namely the maximal natural number such that $\lim_{a \rightarrow 0} \left(1-k^{\ast}(a)\right)/ \lvert a \rvert^{q}$ is finite. The bias is minimal for the rectangular kernel. However, as discussed in Section \ref{sectionHAC}, the resulting estimate $\hat{\Omega}$ is not necessarily positive semi-definite. In contrast, the QS kernel has an optimal Parzen exponent $q=2$ over all the kernels giving positive semi-definite estimators. Thus, the covariance matrix estimator with QS kernel has asymptotic MSE of order $O(B_T/T) + O\left( B_T^{-2} \right)$. Consequently, $B_T$ must grow faster than $T^{1/4}$ and slower than $T^{1/2}$, for the bootstrap error to be $o_p\left( T^{-1/2} \right)$.
In Section \ref{method}, we define the FMB CR by $\mathcal{C}_{1-\alpha} = \left\{ \beta \in \mathcal{B} : \hat{Q}\left( \beta \right) \leq {q}^{\ast}_{1-\alpha} \right\}$ (or alternatively using $\tilde{Q}(\beta)$ as in \eqref{tildeQStat}),
 where ${q}^{\ast}_{1-\alpha}$ is the approximation of $\hat{Q}(\beta_0)$ quantiles by FMB. Thus, $\mathbb{P}[ \beta_0 \in \mathcal{C}_{1-\alpha} ] = \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq {q}^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T$. Theorem \ref{higher_order} gives the order $R_T = o_p\left(T^{-1/2}\right) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right).$ If we select a kernel $k$ and a bandwidth $B_T$ such that $O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right) =  o_p\left(T^{-1/2}\right),$ it implies that $\mathcal{C}_{1-\alpha}$ is higher-order correct by construction, as the (first-order) Gaussian approximation is at best of order $O(T^{-1/2}).$
 If $B_T = C T^{\gamma}$, $C > 0,$ we get the conditions $q >1$ and $(2q)^{-1} <\gamma  < 1/2$. Similar arguments apply for Algorithm \ref{AlgFMB}. Part (ii) extends the higher-order correctness of the i.i.d.\ bootstrap of \cite{hu_estimating_2000} in the just-identified multivariate parameter case (see their Remark 10) to the over-identified case with time-dependent data.

When the monotonicity condition discussed in Section \ref{method} (Example \ref{exAR1}) is violated (or is uneasy to check), we may use the alternative FMB CI or CR defined as $\mathcal{C}_{1-\alpha} := \left\{ \beta \in \mathcal{B} : \tilde{Q}\left( \beta \right) \leq q^{\ast}_{1-\alpha} \right\}$ (see Remark \ref{nomono} and \eqref{tildeQStat}), which are simply connected when $T$ is large enough and centered at $\hat{\beta}$. The following corollary (see the Supplementary Material for its proof) states the higher-order correctness of this alternative version of FMB
CI and CR, based on the $\tilde{Q}$ statistic.
\begin{corollary}\label{HOtildeQ}
Under Assumptions \ref{a1}---\ref{EE5}, $\mathbb{E} \left[ \left\| g_t \right\|^{\bar q s + \delta} \right] < \infty$, for $\delta > 0$, $s \geq 8$, and $\bar q \geq 3$, $\log T = o(B_T)$, we have for $\tilde{Q}$ as in \eqref{tildeQStat} and $Q^{\ast}$ as in \eqref{Rao_type_ast}: $$\sup_{x \in \mathbb{R}^{+}}\left\lvert \mathbb{P}^\ast[ Q^{\ast} \leq x] - \mathbb{P}[\tilde{Q} \leq x]\right\rvert = o_p(T^{-1/2}) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right),$$
where $q$ is the Parzen exponent of $k^{\ast}$.
\end{corollary}
As a consequence, we have $\mathbb{P}[ \beta_0 \in \mathcal{C}_{1-\alpha} ] = \mathbb{P}[ \tilde{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq {q}^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T,$ with $R_T = o_p\left(T^{-1/2}\right) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right).$ Thus, similarly to the CI and CR based on $\hat{S}$ and $\hat{Q},$ the CI and CR based on $\tilde{Q}$ are higher-order correct if we select a kernel $k$ such that $O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right) =  o_p\left(T^{-1/2}\right)$.




\section{Implementation aspects}\label{implement}

\subsection{Consistent covariance matrix estimation} \label{sectionHAC}

In this section, we give the necessary details on the appropriate way to estimate the long-run variance in the statistics $\hat{S}$ (as in \eqref{Eq S_psi}) and $\hat{Q}$ (as in \eqref{Rao_type}). We treat the general case of $\hat{\Omega}$ (as in \eqref{Rao_type}), from which the univariate case can be easily deduced.
Following \cite{andrews_heteroscedasticity_1991} and \cite{smith_gel_2011}, we can obtain several estimators $\hat{\Omega}$ using different types of kernels. Let us give some examples of kernels useful in the implementation of FMB and their respective properties.


\begin{example}[Truncated kernel] The so-called truncated kernel is $k(x)=1$, if $|x| \leq 1$, and $0$, if $|x| > 1$. Considering $m_T$ such that $B_T=(2m_T+1)/2$, we write $g_{T,t}(\beta)=2(2m_T+1)^{-1} \sum_{s=\max{[t-T,-m_T]}}^{\min{[t-1,m_T]}} g_{t-s}(\beta).$ The corresponding long-run variance estimator follows directly from the definition of $\hat{\Omega}$ and has minimal asymptotic mean squared error (see \cite{andrews_heteroscedasticity_1991}), but does not guarantee the resulting variance estimator to be positive semi-definite.
The spectral window generator of the truncated kernel is its Fourier transform $K(\lambda)=\pi^{-1}[\sin(\lambda)/\lambda]$.
The corresponding induced kernel is the Bartlett kernel $k^{\ast}(x)=1-|x/2|$ for $|x| \leq 2$, $0$ for $|x| > 2$, as its spectral window  generator is $K^{\ast}(\lambda)=\pi^{-1}[\sin(\lambda)/\lambda]^2.$ According to \cite{andrews_heteroscedasticity_1991}, the corresponding optimal bandwidth parameter is $m^{opt}_T = O(T^{1/3}),$ and the standardizing constants are $\kappa_1 = 2$ and $\kappa_2 = 2.$
\end{example}


\begin{example}[Quadratic Spectral kernel] \label{QS kernel} Among the available kernels ensuring the long-run variance estimator to be positive semi-definite, \cite{andrews_heteroscedasticity_1991} points out the optimal QS kernel $k_{QS}^{\ast}$, as well as the respective optimal bandwidth $B_T^{opt} = O\left( T^{1/5} \right)$. From the relationship $K^{\ast}\left(\lambda\right)= \left(2\pi / \kappa_2\right) \left\lvert K\left( \lambda \right) \right\rvert^{2}$ and the inverse Fourier transform, \cite{smith_gel_2011} identifies the kernel
\begin{align*}
&k_{J}\left( x \right) := \begin{cases}
&\frac{1}{x}J_1 \left( \frac{6 \pi x}{5} \right)(\frac{5 \pi}{8})^{1/2}\text{ if }x \neq 0,
\\
&\frac{3\pi}{5}(\frac{5 \pi}{8})^{1/2}\text{ if }x = 0,
\end{cases} \nonumber
\\
&J_{\nu}\left( z \right) := \frac{z^{\nu}}{2^{\nu}} \displaystyle \sum_{j = 0}^{\infty} \left( -1\right)^{j} \frac{z^{2j}}{2^{2j}j ! \Gamma \left( \nu + j + 1 \right)},\nonumber
\end{align*}
\noindent inducing the QS kernel by self-convolution: $k^{\ast}_{QS} ( a ) = \left( 1 / \kappa _ { 2 } \right) \int k_{J} ( b - a ) k_{J} ( b ) d b$. The standardizing constants are $\kappa_1 = (5\pi/2)^{1/2}$ and $\kappa_2 = 2\pi.$
\end{example}

Thereby, we preferably use $k_{J}\left( x \right)$ of Example \ref{QS kernel} in \eqref{convol} to get the estimator $\hat{\Omega}$. Indeed, both from theory and simulations, the QS kernel is the optimal induced kernel in terms of asymptotic mean squared error (\cite{andrews_heteroscedasticity_1991}), among all kernels giving positive semi-definite long-run variance estimators. The optimal bandwidth $B_T = O (T^{1/5})$ minimising the mean-squared error of the variance estimator with $k_J(x)$ (see Example 5) does not satisfy the condition $(1-\epsilon)/q < \gamma < \epsilon$ for any $\epsilon \in (0,1/2]$ (see Section 3) since $q = 2$ for the QS kernel. Then, we can take $\gamma = 1/3$, so that we match the condition for $q = 2$.  \cite{wilhelm_optimal_2015}  has already exhibited a discrepancy between the optimal choice for the HAC variance estimator and the optimal choice for a GMM point estimator. In his case, the optimal bandwidth for the point estimate is of the same order as the one minimizing the mean-squared error of the nonparametric plugin estimate, while  the constants of proportionality are significantly different. In our case, the order is even different if we want to achieve higher-order correctness for FMB.

Alternatively, we may use the flat-top kernel version of the QS kernel, see \cite{politis_higher-order_2011}, to get an even faster rate of convergence for the estimated long-run variance. Unfortunately, self-convolution of a kernel $k$ cannot induce a flat-top kernel $k^{\ast}$. Indeed, we know that it cannot be the case that $U = X + Y$, where the random variable $U$ is uniformly distributed on $[0, 1]$ (the flat-top part) and the random variables $X$ and $Y$ are independent and identically distributed (see  Exercise 4.14.20 and its proof by contradiction in \cite{grimmett_thousand_2001}). As a consequence, if we want to benefit from the smaller bias of the flat-top kernel, we should use a different kernel for the original statistic and the bootstrap one. A potential modification of FMB is to decouple the kernel $k^{\ast}$ used in the variance estimator of the original statistic, say a flat-top kernel, and the kernel $k$ used for the smoothed moment indicators. This version of FMB also achieves higher-order correctness since we maintain the asymptotic pivotal nature of the test statistics.


\subsection{From confidence regions to confidence intervals} \label{CICR}

FMB user can manipulate the higher-order correct CR $\mathcal{C}_{1-\alpha}$ to obtain CR for a subset of parameters, or CI for a single parameter. In this section, we will consider separately the (possibly multivariate) parameter of interest $\beta^{(1)}_0,$ and the (possibly multivariate) nuisance parameter $\beta^{(2)}_0.$ Without loss of generality, we write the partition in the order $\beta = (\beta^{(1)\intercal},\beta^{(2)\intercal})^{\intercal}.$

We define the CR for $\beta^{(1)}_0$ as $\mathcal{C}^{(1)}_{1-\alpha} := \{ \beta^{(1)} : \hat{Q}(\beta^{(1)},\hat{\beta}^{(2)}) \leq q^{\ast}_{1-\alpha}\}.$ This manipulation preserves higher-order correctness when $\hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq \hat{Q} ( \beta_0 )$ $a.s.$, in the sense that it ensures $\mathbb{P}[ \beta^{(1)}_0 \in \mathcal{C}^{(1)}_{1-\alpha}] = \mathbb{P}[ \hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq {q}^{\ast}_{1-\alpha}] \geq \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq {q}^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T.$ This condition is satisfied, for instance, when $\hat{Q}$ is a monotonic function of the norm $\|\beta-\hat{\beta}\|,$ since $\| (\beta_0^{(1)\intercal}, \hat{\beta}^{(2)\intercal})^{\intercal} - \hat{\beta} \| \leq  \| \beta_0 - \hat{\beta} \|$. It is also satisfied when the two sets of parameters $\beta^{(1)}$ and $\beta^{(2)}$ can be estimated independently from each other. For instance, in a two-step OLS estimation, it is the case in our ACD(1,0) example of Section \ref{Monte_Carlo}, by the Frisch-Waugh-Lovell Theorem, since we can rewrite the ACD(1,0) as an AR(1) process with orthogonal regressors.

If the inequality $\hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq \hat{Q} ( \beta_0 )$ fails to be true almost surely, we cannot guarantee the higher-order correctness of $\mathcal{C}^{(1)}_{1-\alpha}$. Nevertheless, the inequality $\mathbb{P}[ \hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq {q}^{\ast}_{1-\alpha}] \geq \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}]$ remains true for large enough $T,$ as long as $\hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \overset{\mathcal{D}}\rightarrow \mathcal{X}_{\dim(\beta^{(1)}_0)}$ and $\hat{Q} ( \beta_0 ) \overset{\mathcal{D}}\rightarrow \mathcal{X}_{p}$. This general result implies  that $\mathcal{C}^{(1)}_{1-\alpha}$ contains $\beta^{(1)}_0$ at least with probability $1 - \alpha$ asymptotically, but not always with higher-order refinements.

The drawback of  no guarantee of higher-order refinements is inherent to the existing information on a multivariate parameter. The difficult task of reducing a CR for $\beta$ to a CR for $\beta^{(1)}$ is not directly entangled with the FMB, but more with the nature of dependence between $\beta^{(1)}$ and $\beta^{(2)}$ induced by the moment condition themselves. However, there exists a general way to build CR for $\beta_0^{(1)}$ while preserving higher-order refinements.  Indeed, defining the profile statistic $\bar{Q}(\beta^{(1)}):=\inf_{\beta^{(2)}}\hat{Q}(\beta^{(1)},\beta^{(2)}),$ we get $\bar{Q}(\beta^{(1)}_0)\leq \hat{Q} ( \beta_0 )$ $a.s.$  by construction. Thus, by the same argument, the alternative CR $\bar{\mathcal{C}}^{(1)}_{1-\alpha} := \{ \beta^{(1)} : \bar{Q}(\beta^{(1)}) \leq q^{\ast}_{1-\alpha} \}$ preserves the higher-order refinements. This property comes with a cost, as $\bar{\mathcal{C}}^{(1)}_{1-\alpha}$ is generally more conservative than $\mathcal{C}^{(1)}_{1-\alpha}$ and it might be heavy to compute in high dimension.



\subsection{From confidence intervals to a confidence curve} \label{CICV}


Let us now make connections to the concept of Confidence Distributions (CD) that we use in our empirical application. It aims at answering the following question: can we also use a distribution function, or a “distribution estimator”, to estimate or test for a parameter of interest
in frequentist inference in the style of a Bayesian posterior? (see the review paper by \cite{singh_confidence_2013}). Natural point estimators include the median, the mean, and the maximum of the CD density (\cite{singh_combining_2005}).
That “distribution estimator” is named CD in agreement with the terminology coined by \cite{efron_fisher_1998}, and traces back to the fiducial distribution of \cite{fisher_inverse_1930}, albeit being a purely frequentist concept.
It was introduced by \cite{hjort_confidence_2002} and its asymptotic extension by \cite{singh_combining_2005}
(see also \cite{singh_confidence_2011}, \cite{veronese_fiducial_2015}, and the book-length presentation of \cite{hjort_confidence_2016}).
Example 2.4 of \cite{singh_combining_2005} discusses how a bootstrap distribution can yield a valid asymptotic CD, and Section 2.3.3 of \cite{singh_confidence_2013} how studentization can transmit higher-order accuracy in the i.i.d.\ case. Paralleling these recent developments in fiducial inference theory (see also the review paper of \cite{hannig_generalized_2016}), we can exploit our FMB to produce a fast methodology to build an asymptotically higher-order correct CD as a by-product. Let us define the functions $H_{S}(\beta):=\mathbb{P}[\hat{S}\leq \hat{S}(\beta)]$ and its FMB counterpart $H^{\ast}_{S}(\beta):=\mathbb{P}^{\ast}[S^{\ast} \leq \hat{S}(\beta)]$, for $\beta \in \mathcal{B}$.

\begin{corollary}\label{CDS}
Under Assumptions \ref{a1}---\ref{EE5}, $\mathbb{E} \left[ \left\| g_t \right\|^{\bar q s + \delta} \right] < \infty$, for $\delta > 0$, $s \geq 8$, and $\bar q \geq 3$, $\log T = o(B_T)$, we have the uniform error bound: $\sup_{\beta \in \mathcal{B}}\left\lvert H^{\ast}_{S}(\beta) - H_{S}(\beta)\right\rvert = o_p\left(T^{-1/2}\right) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right),$ and $H^{\ast}_{S}(\beta)$ is an asymptotic CD.
\end{corollary}

We omit the proof since the uniform error bound follows immediately from the proof of Theorem \ref{higher_order}. The second statement comes from the two conditions of Definition 1.1. of \cite{singh_combining_2005} being met, namely $H^{\ast}_{S}(\beta)$ is a cdf, and $H^{\ast}_{S}(\beta_0)$ is uniformly distributed on the unit interval when $T$ goes to infinity. Here, as clarified by \cite{pitman_statistics_1957}, we follow indeed the frequentist view. In $H_{S}(\beta)$ and $H^{\ast}_{S}(\beta)$, randomness is not coming from the (non-random) parameter $\beta$, but from $\hat{S}$ and $S^{\ast}.$

As described in \cite{singh_combining_2005} (see also \cite{fraser_fiducial_1961}, \cite{singh_confidence_2013}), we can also use CD to get $p$-values. For example, the classical bootstrap $p$-value of $H_0 : \beta_0 \leq \beta$ versus $H_1: \beta_0 > \beta$ corresponds to $H^{\ast}_{S}(\beta)$, and the classical equal-tail bootstrap $p$-value of $H_0: \beta_0 = \beta$ versus $H_1: \beta_0 \neq \beta$ corresponds to $2 \min \{H^{\ast}_{S}(\beta), 1-H^{\ast}_{S}(\beta)\}$.

These $p$-values also benefit from higher-order correctness. Collecting them for different values of $\beta$ yields the so-called confidence curve $CV^{\ast}(\beta):=2 \min \{H^{\ast}_{S}(\beta), 1-H^{\ast}_{S}(\beta)\}$, introduced by \cite{birnbaum_confidence_1961} (see \cite{singh_confidence_2013} and \cite{hannig_generalized_2016}  for illustrations). We can view this graphical tool as a piled-up form of two-sided CI of equal tails, at all confidence levels. We provide an example of such a plot in Figure \ref{CC1} for our empirical application in Section \ref{empirical_application}, where we compare CI given by our FMB and first-order Gaussian asymptotics. Finally, \cite{coudin_finite-sample_2020} show how we can design a Hodges-Lehmann-type point estimator (\cite{hodges_estimates_1963}) when a CD is constructed from a hypothesis test.

There exist analogue multivariate CD, for instance Definitions 5.1 and 5.2 in \cite{singh_confidence_2007}. Similarly to the univariate case, we can apply FMB to achieve higher-order accuracy, as long as these multivariate confidence distributions are based on the test statistic $\hat{Q}(\beta_0)$.

\section{Monte Carlo experiments}\label{Monte_Carlo}

To illustrate the applicability of FMB, we consider a simulation exercise on constructing CR for the parameters of an Autoregressive Conditional Duration (ACD) model (\cite{engle_autoregressive_1998}). It is a model typically applied for the analysis of high-frequency data in finance, and more generally to model positive variables (e.g. volatility or volume) via a multiplicative error model; see e.g.\ \cite{hautsch_econometrics_2012} for a recent book-length presentation.

The duration is defined as the time lag  between two consecutive events occurrence, namely $x_\ell := t_\ell-t_{\ell-1}$. Clearly, $x_\ell >0 $, for any $\ell\in \mathbb{T}$. We model $\mathbb{E} \left( x _ { \ell } | x _ { \ell - 1 } , \ldots , x _ { 1 } \right) = m _ { \ell } \left( x _ { \ell - 1 } , \ldots , x _ { 1 } ; \beta \right) := m _ { \ell }$, assuming the
model $x_\ell = m_\ell \varepsilon_\ell$, with $\epsilon_\ell \overset{i.i.d.} \sim \mathcal{E}\left( 1 \right)$ for any $\ell$, with $\mathcal{E}\left( 1 \right)$ being an exponential random variable with mean one. Specifically, for the ACD(1, 1) specification, we have
\begin{equation} \label{Eq: ACD}
x_\ell = \epsilon_\ell m_{\ell}, \text{\ \ with \ \ } m _ { \ell } = \omega + \beta_1 x _ { \ell - 1 } + \beta_2 m _ { \ell - 1 }  ,\qquad \ell\in\mathbb{Z},
\end{equation}
for $\omega > 0$  $\beta_1 , \beta_2 \in \mathbb{R}^{+}$ and  $\beta_1 + \beta_2 < 1$.
When we take $\beta_2 = 0$, the ACD(1,0) model is in agreement with Assumptions \ref{a1}--\ref{EE5} (in Section \ref{Theory}). Thus, we start our numerical experiment with this specification, conducting inference on $\beta:=(\omega,\beta_1)^{\intercal}$. We apply the optimal estimating functions of \cite{li_semiparametric_2000} and a moment condition which does not assume any specific functional form for the underlying innovation density, but relies on the
unconditional expectation of $x_\ell$. Therefore, given a random sample of durations
$(x_1,...,x_\ell,...,x_T)$, the  vector of moment
conditions for the $\ell$-th observation is $g_\ell ( \beta ) := \left( g_{1,\ell}(  \beta), g_{2,\ell}
(\beta), g_{3,\ell}( \beta)\right)^{\intercal},$
with $g_{1,\ell}\left( \beta\right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 }  )(  \partial m _ { \ell } / \partial \omega  )$, $g_{2,\ell}\left( \beta \right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 }  )(  \partial m _ { \ell } / \partial \beta_1  )$, and $g_{3,\ell}\left( \beta\right) := x_\ell  - \left( 1-\beta_1 \right)^{-1}$ (\cite{li_semiparametric_2000}). Thus, we are in the over-identified case with $r=3$ for $p=2$.

In a second step, to get numerical insights on the applicability of our FMB, we extend our Monte Carlo experiment to a general ACD(1,1) model. The latter is non-markovian in the observations and does not fit our setting stricto sensu, since we assume the vector of observations in \eqref{moment} to be finite to ease notations and proofs. However, taking a large number $v$ of lagged durations in an ACD($v$,0) model is close to an ACD(1,1) model, and it meets our current theoretical framework.

We conduct inference on $\beta:=(\omega,\beta_1,\beta_2)^{\intercal}$ and our moment
conditions for the $\ell$-th observation is $g_\ell ( \beta ) := \left( g_{1,\ell}(  \beta), g_{2,\ell}
(\beta), g_{3,\ell}( \beta), g_{4,\ell}( \beta) \right)^{\intercal},$ with $g_{1,\ell}\left( \beta\right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 }  )(  \partial m _ { \ell } / \partial \omega  )$, $g_{2,\ell}\left( \beta \right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 }  )(  \partial m _ { \ell } / \partial \beta_1  )$,
$g_{3,\ell}\left( \beta \right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 }  ) (  \partial m _ { \ell } / \partial \beta_2  ),$
and $g_{4,\ell}\left( \beta\right) := x_\ell  - \left( 1-\beta_1-\beta_2 \right)^{-1}$. We stay in the over-identified case with $r=4$ for $p=3$.

In the following, we compute by Monte Carlo simulations the coverage of FMB CR. We label the results $\hat{Q}_{FMB}$ for the FMB using $\hat{Q},$ $\tilde{Q}_{FMB}$ for the FMB using the cubic approximation of Remark \ref{nomono2} and $\tilde{Q}_{FMB,2}$ for the FMB using the quadratic  approximation of Remark \ref{nomono2}. To validate numerically our theoretical results, we compare the coverage of these FMB CR to some standard first-order correct alternatives. The first one, labeled as $\hat{Q}_{\mathcal{X}^2_r}$, defines CR as contours of the same statistic than FMB, but making use of the (first-order correct) $\mathcal{X}^2_{r}$ asymptotic distribution to compute the rejection probabilities. The second one is the standard elliptical contour of an asymptotically $\mathcal{X}^2_{p}$ distributed Wald statistic (labeled as $W_{\mathcal{X}^2_{p}}$), whose covariance matrix is a HAC estimator with bandwidth $B_T$.

To compare with a state-of-the-art competitor of FMB in terms of higher-order correctness, we also show the coverage of CR yielded by MBB. We adapt the MBB of \cite{inoue_bootstrapping_2006} to a Wald statistic, which is, in turn, an adaptation of \cite{gotze_second-order_1996}.
As the statistic defining the CR has to be asymptotically pivotal to obtain higher-order correctness, we choose to apply MBB to the Wald statistic, which we label as $W_{MBB}.$ We show a detailed CPU time comparison between FMB and MBB in the Supplementary Material (Table \ref{Tab: CPU}). In general, FMB appears to be at least $10^3$ times faster than MBB as expected.

In Table \ref{Tab: 10_025COVER_3}, we display the empirical coverages for the ACD(1,0) model, for $B_T=3$ and $B_T=5.$ In Table \ref{Tab: 025COVER_3}, we show the same outputs for the ACD(1,1) specification.


\begin{table}[htbp]
\begin{center}
\caption{Empirical coverage of CR, a comparison between first and higher-order correct methods.}
\begin{tabular}{l l l l l l l l}
\hline
\hline

	\\


  &&\multicolumn{3}{c}{ $B_T=3$ } & \multicolumn{3}{c}{$B_T=5$} \\

  \multicolumn{2}{l}{Coverages:}  & 0.90 & 0.95 & 0.99   & 0.90 & 0.95 & 0.99 \\
	\\
  \multirow{6}{*}{\rotatebox[origin=c]{90}{$T=250$}} & $\hat{Q}_{FMB}$ & 0.92 & 0.94 & 0.98   & 0.89 & 0.93 & 0.96 \\

  &$\tilde{Q}_{FMB}$ & 0.92 & 0.95 & 0.98 & 0.90 & 0.94 & 0.97 \\
	&$\tilde{Q}_{FMB,2}$  & 0.92 & 0.95 & 0.98 & 0.90 & 0.93 & 0.97  \\
	&$\hat{Q}_{\mathcal{X}^2_r}$ & 0.91 & 0.93 & 0.97 & 0.88 & 0.90 & 0.96 \\
  &$W_{MBB}$  & 0.93 & 0.97 & 1.00 & 0.93 & 0.97 & 1.00 \\
  &$W_{\mathcal{X}^2_{p}}$  & 0.84 & 0.90 & 0.96 & 0.80 & 0.86 & 0.92 \\
	\\

   \multirow{6}{*}{\rotatebox[origin=c]{90}{$T=500$}} & $\hat{Q}_{FMB}$ & 0.93 & 0.96 & 0.98 & 0.92 & 0.95 & 0.98  \\

  & $\tilde{Q}_{FMB}$  & 0.93 & 0.95 & 0.98 & 0.92 & 0.95 & 0.98 \\
	& $\tilde{Q}_{FMB,2}$  & 0.92 & 0.96 & 0.99 & 0.91 & 0.95 & 0.99 \\
	& $\hat{Q}_{\mathcal{X}^2_r}$ & 0.92 & 0.95 & 0.98 & 0.91 & 0.94 & 0.97 \\
  & $W_{MBB}$ & 0.95 & 0.98 & 1.00  & 0.94 & 0.98 & 1.00 \\
  & $W_{\mathcal{X}^2_{p}}$ & 0.86 & 0.91 & 0.97  & 0.84 & 0.89 & 0.96 \\
	\\
  \hline
	\hline
\end{tabular}
\label{Tab: 10_025COVER_3}
\caption*{

The true values of the unknown parameters of the ACD(1,0) are $\omega = 1.5$ and $\beta_1 = 0.25$.}
\end{center}
\end{table}



\begin{table}[htbp]
\begin{center}
\caption{Empirical coverage of CR, a comparison between first and higher-order correct methods.}
\begin{tabular}{l l l l l l l l}
\hline
\hline
\\

&&\multicolumn{3}{c}{$B_T = 3$}&\multicolumn{3}{c}{$B_T = 5$} \\

\multicolumn{2}{l}{Coverages:} & 0.90 & 0.95 & 0.99 & 0.90 & 0.95 & 0.99 \\
\\


 \multirow{6}{*}{\rotatebox[origin=c]{90}{$T=250$}} & $\hat{Q}_{FMB}$  & 0.88 & 0.93 & 0.98  & 0.87 & 0.90 & 0.96\\
    & $\tilde{Q}_{FMB}$  & 0.89 & 0.92 & 0.95 & 0.87 & 0.90 & 0.94 \\
		& $\tilde{Q}_{FMB,2}$  & 0.82 & 0.87 & 0.94 & 0.80 & 0.85 & 0.93 \\
		& $\hat{Q}_{\mathcal{X}^2_r}$  & 0.86 & 0.91 & 0.96 & 0.85 & 0.89 & 0.95 \\
    & $W_{MBB}$  & 0.97 & 0.99 & 1.00 & 0.97 & 0.99 & 1.00 \\
    & $W_{\mathcal{X}^2_{p}}$  & 0.73 & 0.79 & 0.88 & 0.71 & 0.77 & 0.86 \\
		\\
   \multirow{6}{*}{\rotatebox[origin=c]{90}{$T=500$}} & $\hat{Q}_{FMB}$ & 0.91 & 0.95 & 0.99 & 0.90 & 0.94 & 0.98\\
    & $\tilde{Q}_{FMB}$ & 0.91 & 0.94 & 0.97 & 0.91 & 0.94 & 0.97 \\
		 & $\tilde{Q}_{FMB,2}$ & 0.85 & 0.90 & 0.96 & 0.84 & 0.89 & 0.95 \\
		& $\hat{Q}_{\mathcal{X}^2_r}$  & 0.90 & 0.93 & 0.97 & 0.89 & 0.93 & 0.97 \\
    & $W_{MBB}$  & 0.98 & 0.99 & 1.00  & 0.97 & 0.99 & 1.00 \\
    & $W_{\mathcal{X}^2_{p}}$ & 0.78 & 0.83 & 0.92 & 0.77 & 0.83 & 0.92 \\
	 \\
   \hline
	 \hline

\end{tabular}
\label{Tab: 025COVER_3}
\caption*{

The true values of the unknown parameters of the ACD(1,1) are $\omega=1.5,$ $\beta_1 = 0.25$ and $\beta_2 = 0.25$.}
\end{center}
\end{table}



\newpage

In line with our theoretical results, we observe that FMB performs well compared to its first-order correct competitors: the coverages are typically closer to their nominal level. It generally remains true for the $\tilde{Q}_{FMB}$ version of FMB CR, whose coverage is often very close to the one of the original FMB CR based on $\hat{Q}_{FMB}.$ The coverage of the quadratic $\tilde{Q}_{FMB,2}$ version of FMB CR is generally further away from the original FMB.


The Wald statistic seems to yield very erratic CR (which is additionally confirmed by unreported plots). The MBB version of the Wald statistic improves slightly on the asymptotic $\mathcal{X}^2_p$ distribution, without being convincing though.

Finally, the FMB presented here does not take advantage of all the potential fine-tuning, and this should leave room for practical improvement. First, we might improve FMB if the long-run variance $\Omega(\beta_0)$ is estimated with a less biased version of variance estimator, for instance carrying out a prewhitening step (\cite{andrews_improved_1992}), or using a flat-top kernel (\cite{politis_higher-order_2011}) as discussed in Section \ref{sectionHAC}. It should yield a smaller bootstrap error, as shown in Theorem \ref{higher_order}. Second, we stress that the moment indicators do not have necessarily the same dependence structure. Thus, smoothing the multivariate moment indicators with different bandwidths might further improve the coverage of FMB.

\section{Real data application}\label{empirical_application}

In this section, we illustrate how FMB performs on real data. We look at daily volumes of stock transaction (in millions), modeled with the same exponential ACD as in Section \ref{Monte_Carlo} (see \eqref{Eq: ACD}). We focus on data available online (\textit{Yahoo! Finance}), for five stocks
in three different sectors, namely bank, technology, and food. We compute the CR (for parameters $\omega$, $\beta_1$ and $\beta_2$), before  the subprime crisis (2005), during the crisis (2008), and the current period (2018). The sample size of each period corresponds to the number of trading days, namely $T=252$ up to negligible variations from year to year. Before diving into a deeper analysis, we briefly describe the data at hand in Table \ref{Tab: SUM} below.

\begin{table}[htbp]
\begin{center}
\caption{Summary statistics of volumes of transaction.}
\begin{tabular}{|lllllllll|}
\hline
    \textit{Year 2005} & Min & Max & Med & Mean & IQR & SD & SKN & KURT \\
\hline
 \rowcolor{mygray} BA & 4.52 & 42.05 & 10.67 & 11.24 & 4.98 & 4.47 & 2.36 & 11.18 \\

JPM & 3.77 & 25.47 & 9.93 & 10.6 & 4.11 & 3.25 & 0.94 & 1.31 \\

 \rowcolor{mygray} MSF & 27.21 & 187.38 & 63.53 & 66.61 & 20.94 & 20.23 & 1.78 & 6.40 \\

KO & 3.96 & 39.75 & 11.01 & 11.86 &  3.83 & 4.15 & 2.64 & 12.16\\

 \rowcolor{mygray} UL & 0.16 & 3.36 & 0.49 & 0.59 & 0.31 & 0.39 & 2.91 & 12.78\\

\hline \hline
\textit{Year 2008} & Min & Max & Med & Mean & IQR & SD & SKN & KURT \\
\hline

 \rowcolor{mygray} BA & 22.59 & 322.73 & 63.16  & 78.63 & 59.52 & 49.01 &  1.66 & 3.28\\

 JPM & 12.34 & 194.07   & 41.96 & 48.99 & 29.64 & 25.90 & 1.93 & 5.72\\

 \rowcolor{mygray} MSF & 16.88 & 291.14  & 78.50  & 84.17 & 38.81 & 35.51 & 1.55 & 4.73 \\

 KO & 5.32 & 79.21 & 23.35 & 25.26 & 12.77 & 10.63 & 1.62 & 4.20\\

 \rowcolor{mygray} UL & 0.26 & 5.19 &  0.79  & 1.08 & 0.73 & 0.83 & 2.09 & 4.85\\

\hline \hline
  \textit{Year 2018}  & Min & Max & Med & Mean & IQR & SD & SKN & KURT \\
\hline
 \rowcolor{mygray} BA & 22.97 & 165.88 & 62.21 & 67.91 & 29.16 & 24.15 & 1.29 & 2.02\\

 JPM & 6.49 & 41.31 & 13.90 & 15.17 & 6.43 & 5.53 & 1.60 & 3.63\\

 \rowcolor{mygray} MSF & 13.66 & 111.24 & 27.61 & 31.59 & 14.23 & 13.40 & 1.74 & 4.82\\

 KO &  4.79  & 32.48 & 11.91  & 12.52 & 4.32 & 4.13 & 1.32 & 2.67\\

 \rowcolor{mygray} UL & 0.33 & 4.88 & 0.90  & 1.05  & 0.54 & 0.64 & 2.96 & 11.54\\

\hline
\end{tabular}
\caption*{

We consider the companies Bank of America (BA), JP Morgan (JPM), Microsoft (MSF), Coca-Cola (KO) and Unilever (UL). In the summary, Med stands for the median, IQR for the inter-quantile range, SD for the standard deviation, SKN for skewness, and KURT for excess of kurtosis.}
\label{Tab: SUM}
\end{center}
\end{table}
Table \ref{Tab: SUM} illustrates the larger variability of the volumes of transaction during 2008, as measured by the standard deviation (SD) and the interquartile range (IQR). The high skewness (SKN) and excess of kurtosis (KURT) typically indicate that a higher-order correct inferential procedure might be required in finite samples.

To investigate further the impact that asymmetry and fat tails may have on the conducted inference, we compute the FMB and the asymptotic normal (Asy) CR of nominal coverage $(1-\alpha) = 95\%$, for the ACD(1,1) parameters $\omega,$ $\beta_1$ and $\beta_2$ at each period. As FMB yields higher-order accurate CR by inverting probabilities of the test statistic $\hat{Q}(\beta_0)=\hat{Q}(\omega_0,\beta_{1,0},\beta_{2,0})$, we represent these trivariate CR by slicing them at the estimates $\hat{\omega},$ $\hat{\beta}_1$ and $\hat{\beta}_2$ (Table \ref{Tab: Big1}). Namely, we cut the CR by fixing the parameters that are not of interest to their estimated values.
To keep Table \ref{Tab: Big1} concise, we do not report the estimate $\hat{\omega}$ and the intervals for the parameter $\omega$. We can deduce the former from Table \ref{Tab: SUM}.\footnote{Using Section \ref{Monte_Carlo} and volumes of transaction $\{x_t\}_{t=1}^{T}$ instead of durations, we have $\hat{\omega} \approx (1-\hat{\beta_1}-\hat{\beta_2})T^{-1}\sum_{\ell=1}^{T}x_{\ell},$ from the moment condition based on $g_{3,t}.$ We report the sample mean $T^{-1}\sum_{\ell=1}^{T}x_{\ell}$ in Table \ref{Tab: SUM}.}

\begin{table}[htbp]
\begin{center}
\small\addtolength{\tabcolsep}{-5pt}
\caption{Analysis of volumes of transaction.}
\begin{tabular}{|c|c|cc|cc|cc|}
\hline
\multicolumn{2}{|c|}{\multirow{2}{*}{Asset}} & \multicolumn{2}{|c|}{2005} & \multicolumn{2}{|c|}{2008} & \multicolumn{2}{|c|}{2018} \\ \cline{3-8}
\multicolumn{2}{|c|}{}                        & \multicolumn{1}{|c|}{$\hat{\beta_1}$}   & \multicolumn{1}{|c|}{$\hat{\beta_2}$}  & \multicolumn{1}{|c|}{$\hat{\beta_1}$}  & \multicolumn{1}{|c|}{$\hat{\beta_2}$}  & \multicolumn{1}{|c|}{$\hat{\beta_1}$}  & \multicolumn{1}{|c|}{$\hat{\beta_2}$} \\ \hline


\multirow{3}{*}{JPM}         & Est            &   0.402   &   0.132   &    0.691    &   0.122     &   0.459   &   0.345   \\
                             & Asy            &    [0.384,0.420]     &    [0.114,0.150]      &    [0.673,0.709]    &   [0.104,0.140]   &   [0.448,0.469]   &  [0.335,0.355]   \\
                             & FMB        &    [0.357,0.442]    &   [0.085,0.178]     &    [0.621,0.726]    &   [0.060,0.164]     & [0.429,0.483]    &    [0.318,0.372]  \\ \hline

\multirow{3}{*}{BA}          & Est         &     0.366    &  0.534    &    0.624    &   0.323   &  0.538  &   0.141  \\
                             & Asy           &     [0.361,0.371]    &    [0.529,0.540]    &   [0.618,0.629]   &   [0.318,0.328]   &  [0.521,0.554]   &   [0.125,0.157]  \\
                             & FMB           &     [0.345,0.381]     &   [0.512,0.549]     &    [0.594,0.641]    &   [0.295,0.344]    & [0.491,0.584]      &    [0.102,0.192]   \\ \hline


\multirow{3}{*}{MSF}         & Est          &   0.271   &  0.340     &   0.595  &   0.272      &  0.570   &  0.268  \\
                             & Asy         &   [0.261,0.281]   & [0.330,0.350]   &    [0.586,0.604]   &   [0.262,0.281]     &  [0.559,0.582]   & [0.257,0.280]      \\
                             & FMB          &  [0.242,0.297]   &  [0.307,0.373] &  [0.562,0.623]  &  [0.243,0.301]   &   [0.536,0.596]   & [0.238,0.294]    \\	\hline


\multirow{3}{*}{KO}         & Est         &    0.202    &   0.377      &   0.488    &    0.371     &   0.392    &   0.382   \\
                             & Asy           &     [0.190,0.213]       &  [0.365,0.389]    &   [0.479,0.496]    &   [0.363,0.380]     &   [0.384,0.399]   &    [0.375,0.390]    \\
                             & FMB         &   [0.166,0.233]  &   [0.344,0.426]     &  [0.458,0.513]    &  [0.343,0.400]    &    [0.363,0.420]    &    [0.358,0.413]   \\ \hline


\multirow{3}{*}{UL}         & Est          &   0.331   &  0.529    &   0.571  &   0.290      &  0.435  &   0.483 \\
                             & Asy          &    [0.320,0.343]     &  [0.518,0.541]    &    [0.555,0.586]   &   [0.275,0.305]     &   [0.428,0.441]  &   [0.477,0.490]  \\
                             & FMB         &  [0.297,0.365]  &   [0.509,0.571]    &  [0.513,0.599]  &  [0.236,0.321]   &   [0.413,0.456]   & [0.464,0.503]     \\	\hline


\end{tabular}
\label{Tab: Big1}
\caption*{

We consider the companies Bank of America (BA), JP Morgan (JPM), Microsoft (MSF), Coca-Cola (KO) and Unilever (UL). We use the Exponential Tilting estimator (Est), a particular case of GEL, and the benchmark intervals are the first-order asymptotic Gaussian (Asy). All the intervals are equal-tailed and have conditional nominal coverage of $95\%$ when we fix the other parameters at their estimated values.}
\end{center}
\end{table}
A few comments are in order. First of all, the different sectors exhibit very diverse reactions to the events happening in 2008. For instance, the food sector seems to be the most stable, while financial sector undergoes a huge variability, as we could expect. We can observe this  either comparing non-critical periods to the crisis, or comparing the estimates and their CI before and after the crisis. For instance, the estimates for Unilever are almost the same before and after the crisis, as if the company has recovered the same volume behaviour. Coca-Cola looks equally stable with respect to the parameter $\beta_2$, which is almost unchanged after the crisis. Second, the estimate $\hat{\beta_1},$ respectively $\hat{\beta_2},$ seems to be larger, respectively smaller, during the crisis period.  It is expected since $\beta_1$ reflects the sudden trading reactions due to changes in the expectations by the market participants during the crisis period. Thus, this feature of 2008 corresponds to an increase of the impact of news (shocks) on the volumes of transaction (via the parameter $\beta_1$), relative to persistence (via the parameter $\beta_2$). Finally, we see that the FMB CI are longer than the first-order correct Gaussian CI. It is in line with our Monte Carlo experiments, as available in Section \ref{Monte_Carlo}. Indeed, as the CR are defined by level sets of $\hat{Q}$, a longer CI corresponds to an adaptation of FMB to a skewed or fat-tailed distribution. Since the CI obtained by Gaussian approximation are typically shorter, we conclude that the routinely applied first-order asymptotic theory tends to underestimate the rejection probability, whereas FMB stays conservative.
Our experience underpinned by several Monte Carlo simulations makes us expect that the distribution of $\hat{Q}(\beta_0)$ is more skewed or fat-tailed than the chi-squared; see the comparison between $\hat{Q}_{\mathcal{X}^2_r}$ and FMB in Table \ref{Tab: 025COVER_3}.

Following our discussion on CD (Subsection \ref{CICV}), we illustrate here the link between our FMB CR and our previous definition of asymptotic confidence distribution $H^{\ast}_S(\beta)$, via the confidence curve $CV^{\ast}(\beta)$. Among the alternative ways to represent the former CR, marginalization allows us to build unconditional CI. Stacking the CR at different coverages $1-\alpha$ leads to a center-outward confidence curve for the multidimensional parameter $CV(\omega,\beta_1,\beta_2)$. For each CI, we integrate out the two parameters that are not of interest in $CV(\omega,\beta_1,\beta_2)$. This yields a different confidence curve $CV^{\ast}(\beta)$ for each $\beta \in \{\omega,\beta_1,\beta_2\},$ whose level sets give the equal-tailed CI. As an illustration of graphical use of these confidence curves (defined in Section \ref{Theory}), Figure \ref{CC1} reports a comparison between the FMB and Gaussian CI based on the FMB and Gaussian confidence curves. Again we observe that FMB is much more conservative.

\begin{figure}[htbp]
\begin{center}
\caption{Confidence Curves of the FMB and Gaussian approximation.}\label{CC1}
\includegraphics[width = 15cm, height = 11cm]{ULCC.pdf}
\caption*{Confidence curve for the parameter $\beta_1$ of Unilever in 2005, with point estimate $\hat{\beta}_1=0.33$. The sample size is $T=252$, and the nominal coverage of the confidence intervals is $(1-\alpha)=95\%$. The flat solid line is the rejection probability level $5\%$, and its crossing with the confidence curves gives the FMB $95\%$ equal-tailed CI $[-0.31, 0.75]$ and Gaussian $95\%$ equal-tailed CI $[-0.05, 0.72]$.}
\end{center}
\end{figure}

\newpage

\section*{Acknowledgement}
We would like to thank the editor, the co-editor, and the referee for constructive criticism and numerous
suggestions which have led to substantial improvements over the previous versions. We thank A.-P.\ Fortin, E.\ Paparoditis, D.\ Politis, and R.\ J.\ Smith for helpful comments and discussion, as well as participants in the North American Summer Meeting of the Econometric Society (Seattle, 2019),  Geneva Finance Research Institute seminar, Workshop on Statistical Learning (Geneva, 2019-20), online RCEA Time Series Workshop 2021, and online IAAE Conference 2021.



\section*{}
\bibliography{FMB_JOE}

\newpage