Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,625 characters · 15 sections · 145 citation commands
A Higher-Order Correct Fast Moving-Average Bootstrap for Dependent Data
Inference based on first-order correct asymptotics can be misleading with confidence intervals having erratic probability coverage. It is especially true in the presence of serial dependence where first-order asymptotics often requires larger sample sizes than for i.i.d.\ data to apply. Resampling methods for time series help to obtain confidence intervals with better finite sample properties. Bootstrap methods for moment condition models have been extensively discussed under various dependence structures by, for example, hall_bootstrap_1996, brown_generalized_2002, inoue_bootstrapping_2006, and davidson_bootstrap_2006. If bootstrap methods for $m$-dependent and strongly mixing data can achieve higher-order correctness (hall_bootstrap_1996, inoue_bootstrapping_2006), they are computationally too intensive, once applied to heavy numerical estimation procedures. For a book-length review, see e.g.\ lahiri_resampling_2010.
In this paper, we propose a novel fast bootstrap scheme, that we call the Fast Moving-average Bootstrap (FMB). The resampling method is computationally attractive while maintaining higher-order correctness of the inferential procedure for strongly mixing data. Our idea for building confidence regions for the parameter of interest is to realize that smoothing the moment indicators as in the Generalized Empirical Likelihood (GEL) literature permits to bootstrap them as if they were i.i.d.\ parente_generalised_2018a study the first-order validity of GEL test statistics based on a similar bootstrapping scheme, the Kernel Block Bootstrap (henceforth KBB); see parente_kernel_2018b and parente_quasi-maximum_2019. Our approach differs from KBB in two significant aspects. First, our methodology does not require to solve the estimation problem at each bootstrap sample, lessening drastically the computational burden. Indeed, FMB is at least (except for a simple low-dimensional linear model) one thousand times faster, according to standard rules on bootstrap simulation errors (efron_better_1987 Section 9, davison_bootstrap_1997 Section 2.5.2). Second, we exploit an inversion technique to benefit from the kernel smoothing used in the studentization of our test statistic. The inversion is related to the standard percentile$-t$ bootstrap approach in the linear univariate case (see Example 1 below). The studentization relies on a simple sample variance of the smoothed moment indicators, which turns out to be asymptotically equivalent to a HAC estimator for the original moment indicators, as shown by smith_automatic_2005. Together these inversion and studentization make our FMB inference amenable to be shown higher-order correct. Our proof strategy is not directly applicable to KBB; its higher-order correctness remains a conjecture.
The already existing fast resampling methods usually hinge on the first-order von Mises expansion of the estimating function (shao_jackknife_1995, davidson_bootstrap_1999, andrews_higher-order_2002, salibian-barrera_bootstrapping_2002, goncalves_maximum_2004, hong_fast_2006, salibian-barrera_principal_2006, salibian-barrera_fast_2008, camponovo_robust_2012, camponovo_predictability_2013, armstrong_fast_2014, and goncalves_bootstrapping_2019). It yields a fast approximation, but its inherent construction does not ensure higher-order correctness. Instead, our fast method relies on inversion, namely we identify the level sets of test statistics under the null hypothesis to obtain confidence regions for the parameter of interest (see parzen_resampling_1994 and hu_estimating_2000 for i.i.d.\ data). Furthermore, the FMB confidence regions are invariant to monotonic reparameterization, due to studentization of the moment indicators. It ensures stability of our method across varying parameter scales (diciccio_bootstrap_1996).
We design FMB for GEL estimator to exploit its intrinsic smoothing, and as it provides a considerably wide theoretical framework on semiparametric estimation (smith_gel_2011). As a consequence, the higher-order refinements achieved by our method ensue for the Empirical Likelihood (see qin_empirical_1994, imbens_one-step_1996, kitamura_empirical_1997), the Exponential Tilting (kitamura_information-theoretic_1997, imbens_information_1998), and the Continuously Updating Estimator (hansen_finite-sample_1996). In addition to the KBB, other bootstrap methods already exist in the GEL literature. For instance, bravo_empirical_2004 shows the higher-order correctness of the bootstrap for inference based on empirical likelihood with i.i.d.\ data, while bravo_blockwise_2005 shows consistency of the block bootstrap for empirical entropy tests in times series regressions with strongly mixing data. However, to our knowledge, there is no proof of higher-order correctness of the bootstrap for GEL in the literature yet.
Clearly, we can also apply FMB in the setting of M-estimation (huber_robust_1964) and Generalized Method of Moment (hansen_large_1982), obtaining a fast version of the bootstrap methods derived in hall_bootstrap_1996 for $m$-dependent data and inoue_bootstrapping_2006 for strongly mixing data.
The structure of the paper is as follows. Section (ref) is a simple introduction to the FMB algorithm in the univariate case. There, we also discuss connections between FMB and already existing resampling schemes. In Section (ref), we briefly present the GMM and GEL estimators for strongly mixing time series, using these frameworks as a tool to extend FMB to the multivariate setting. We itemize our assumptions and present the main theoretical results in Section (ref). In Section (ref), we give details on the implementation aspects of FMB, emphasizing the relation between the choice of the kernel and the properties of the long-run variance estimator. We also discuss connections with the recent literature on confidence distributions, that we use in our empirical application. We present our Monte Carlo experiments in Section (ref), and a real data example in Section (ref). Finally, we prove our theorems in appendix. For some technical lemmas, we give the proofs in the Supplementary Material (available online).
Let $\left\{X_t\right\}_{t \in \mathbb{Z}}$ be a stationary strongly mixing process in $\mathbb{R}^d$, observed at $t = 1,...,T$. We assume that the time series of interest satisfies Assumptions (ref)---(ref) in Section (ref), which are standard in the bootstrap literature. Let $\mathcal{B} \subset \mathbb{R}$ be the compact space of the parameter $\beta$ and $\mathbb{X}_t := \left\{ X_{t_1},...,X_{t_v} \right\}$ be a collection of vectors from the process $\left\{X_t\right\}_{t \in \mathbb{Z}}$. Consider the function $g: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R}$ such that:
where the expectation $\mathbb { E }$ is taken w.r.t.\ the true underlying distribution, unknown and depending on $\beta_0$. In the following, we use the shorthand notation $g_{t} \left( \beta \right) := g\left( \mathbb{X} _ {t}, \beta \right)$.
The function $g$ in (ref) can be the (conditional) likelihood in full parametric models, or it can be obtained using the (conditional) moments and/or may depend on instrumental variables in semiparametric models. The collection of vectors $\mathbb{X}_t$ typically contains information on the relation between the observations and the parameter characterizing the $q$-dimensional stationary distribution of a time series. More generally, we can exploit the knowledge in closed-form of the (conditional) moments to obtain (martingale) estimating functions for non-linear conditional autoregressive and heteroscedastic models or discretely observed diffusions. We refer to godambe_quasi-likelihood_1987, taniguchi_asymptotic_2000, and kessler2012statistical for book-length presentations.
Each function of the sequence $\left\{ g_{t}\left( \beta \right)\right\}_{t=1}^{T}$ is often defined using the innovations, which can be i.i.d.\ random variables or more generally martingale differences. Thus, $\left\{ g_{t}\left( \beta \right)\right\}_{t=1}^{T}$ exhibits less dependence than the original process $\left\{ X_{t} \right\}_{t=1}^{T}$. Nevertheless, neglecting this temporal dependence has a serious impact on the performance of several inferential procedures, in particular it can affect the consistency of the bootstrap variance estimator and the higher-order accuracy of bootstrap confidence intervals.
To take automatically this aspect into account, we follow kitamura_information-theoretic_1997, otsu_generalized_2006, guggenberger_generalized_2008, and smith_gel_2011, and we perform a convolution of the moment indicator $g$ with the kernel $k: \mathbb{R} \rightarrow \mathbb{R}$, obtaining:
where $B_T$ is a bandwidth parameter, increasing in $T$ and such that $B_T / T \longrightarrow 0$. The convolution in (ref) induces a HAC-type modification, ensuring consistency of the long-run variance estimation of the mean over time $\bar { g }_T ( \beta ) := T ^ { - 1 } \sum _ { t = 1 } ^ { T } g _ {T,t } ( \beta )$; see newey_simple_1987, andrews_heteroscedasticity_1991, and smith_automatic_2005. Solving $\bar { g }_T ( \beta )=0$ gives the just-identified univariate estimator $\hat{\beta}$. Hence, the estimator $\hat{\beta}$ relies on a smoothed moment condition. Below, we explain how we can further exploit the convolution in (ref) to derive our bootstrap.
Let us first give the intuition of our methodology for the construction of confidence interval (CI) for $\beta_0$; more technical aspects are available in Sections (ref) and (ref). For ease of notation, we drop the subscript $T$ from any estimator, whenever its dependence on the sample size is clear from the context.
To keep the exposition as simple as possible, we assume temporarily a one-to-one relationship between the parameter and the estimating function $\bar { g }_T ( \beta ),$ in an suitable subset of $\mathcal{B}$. Even though the probabilistic validity of the FMB CI does not depend on this condition (see e.g. lehmann_testing_1959, shao_mathematical_1999, hansen_finite-sample_1996 and guggenberger_generalized_2008-1), this assumption allows us to explain easily why our resampling scheme does not need to solve the estimating equation for each bootstrap sample.
Intuitively, the construction of the FMB CI goes as follows. First, our bootstrap scheme provides a higher-order correct approximation of the distribution of a statistic $\hat{S}(\beta_0)$. This statistic is an asymptotically pivotal version of the estimating function $T^{1/2}\bar { g }_T ( \beta ),$ evaluated at the true parameter $\beta_0.$ Second, the one-to-one relationship allows us to map the quantile estimates of $\hat{S}$ to CI limits in $\mathcal{B}.$ This mapping is crucial to gain computational efficiency. Indeed, we use the computationally intensive part of the FMB algorithm to compute the distribution of simple mean-type statistic $\hat{S}(\beta_0)$ which is much faster to compute than roots of $T^{1/2}\bar { g }_T ( \beta ),$ or numerical solutions to the estimating optimization problem (see Section (ref)). From an hypothesis testing point of view, FMB yields an approximation to the distribution of $\hat S(\beta)$ under $H_0:$ $\beta=\beta_0.$ Then, each $\beta$ in the CI is in the non-rejection region of $H_0$.
The next example is a widely-applied model where the one-to-one condition is satisfied, since the considered function is monotonic. Below, we explain in Remark 2 how to adapt FMB to deal with general estimating functions, which do not necessarily satisfy the monotonicity condition.
Now that the main principles of FMB are settled, we present the detailed algorithm underlying its numerical implementation. The statistic serving as basis for inference is the asymptotically pivotal quantity:
where, for instance, $ \hat{\sigma} := \kappa^2_1 (T\kappa_2)^{-1} \sum_{t=1}^{T} g^{2}_{T,t}( \hat{\beta} ) $ and $\kappa_j := \int k\left(u\right)^j du$ for $j=1,2$. This statistic is a particular value of the function $\hat{S}: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R},$ $\hat{S}(\beta):=T^{1/2}\bar{g}_T\left(\beta\right) / \hat{\sigma},$ that we suppose strictly increasing on $\mathcal{B}$. The studentization in $\hat{S}$ is crucial for FMB to be higher-order correct. In principle, we can apply other estimators of the long-run variance and we flag that each estimator $\hat{\sigma}$ has its own bias, which is going to affect the properties (e.g.\ the accuracy) of FMB. We refer to Section (ref) for further discussion.
Considering an i.i.d.\ bootstrap sample drawn from $\{g_{T,t}( \hat{\beta})\}_{t=1}^{T},$ say $\{g^{\ast}_{T,t}\}_{t=1}^{T},$ the bootstrap version of $\hat{S}$ in (ref) is:
with $\bar{g}_T^\ast:=T^{-1} \sum_{t=1}^{T} g^{\ast}_{T,t}$ and $\hat{\sigma}^{\ast 2} := {T}^{-1}\sum_{t=1}^{T} g^{\ast 2}_{T,t}.$ In (ref), both the computed numerator and denominator rely on $g^{\ast}_{T,t},$ and thus avoid re-estimating the parameter on each bootstrap sample in order to make it fast.
Then, the algorithm of our FMB is made of five steps (lines 2-4, 5, 6-13, 14-15, 16-17).
In Algorithm (ref), we exploit the monotonicity assumption only in Step 5, where we invert the studentized estimating function $\hat S(\beta)$; see hu_estimating_2000 for the use of a similar device.
Indeed, to derive the CI, the bootstrap procedure first provides $(1-\alpha)$-quantile estimates of $\hat S(\beta_0)$, say $q^\ast_{1-\alpha}.$ Then, a numerical method (e.g. Newton-Raphson or secant methods) defines a one-sided $(1-\alpha)$-CI for $\beta_0$ as $[\beta_{\min}, \hat{q}_{1-\alpha}]$, where $\beta_{\min}:= \min\mathcal{B}$ and the upper limit $\hat{q}_{1-\alpha}$ solves $\hat S (q_{1-\alpha}) = q^\ast_{1-\alpha}$ in $q_{1-\alpha}.$
For a two-sided equal-tailed CI, we follow the same principles. We consider two real numbers $s_1$ and $s_2$ such that $\mathbb{P}[\hat{S} \left( \beta_0 \right) \leq s_1]=\alpha/2$ and $\mathbb{P}[ \hat{S} \left( \beta_0 \right) > s_2 ]=\alpha/2$. From FMB, we obtain the approximation\footnote{As customary in the bootstrap literature, $\mathbb{P}^{\ast}[X^{\ast} \leq x ]$ denotes the empirical c.d.f.\ of any variable $X^{\ast}$ generated by the bootstrap scheme (here $S^\ast$). We give the general definition of the bootstrap probability measure $\mathbb{P}^{\ast}$ in (ref), Section (ref).} $\mathbb{P}^{\ast}[s_1 < S^{\ast} \leq s_2 ] = \mathbb{P}[s_1 < \hat{S} ( \beta_0 ) \leq s_2 ] + R_T$, where $R_T$ is an asymptotically negligible remainder term. Hence, we can compute $s_1$ and $s_2$ such that $\mathbb{P}^{\ast}[s_1 < S^{\ast} \leq s_2 ] = 1 - \alpha$. Then, the CI for $\beta_0$ is $\mathcal{C}_{1-\alpha} := (c_1,c_2]$, with $c_1 := \hat{S}^{-1}\left(s_1\right)$ and $c_2 := \hat{S}^{-1}\left(s_2\right)$, ensuring that $\mathbb{P}\left[c_1 < \beta_0 \leq c_2 \right] = 1 - \alpha + R_T$. Under Assumptions (ref)---(ref) (Section (ref)), we can get $R_T = o_p\left(T^{-1/2}\right),$ given a suitable choice of kernel $k$ and bandwidth $B_T$ (see Theorem (ref) and discussion in Section (ref)). It implies that $\mathcal{C}_{1-\alpha}$ is correct up to a higher order.
From the studentization in $\hat S(\beta)$, the CI limits $c_1$ and $c_2$ derived in Step 5 remain invariant to monotonic transformation of the parameter. This property is crucial for the bootstrap CI (diciccio_bootstrap_1996), ensuring stability of FMB across varying parameter scale.
Example 1 [cont'd]. Let us see how the steps in Algorithm 1 specialize for the $AR(1).$ In Step 1, we have $g_{T,t}$ as in (ref) and $\bar { g }_T ( \beta ) = T ^ { - 1 } \sum _ { t = 1 } ^ { T } {B_{T}}^{-1/2}\sum_{s = t-T}^{t-1} k({s}/{B_{T}}) (Y_{t-s}-\beta Y_{t-s-1})Y_{t-s-1}$. For Step 2, the estimator is available in closed form: $$ \hat{\beta}=\left( \displaystyle \sum_{t=1}^{T} \displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right) Y_{t-s}Y_{t-s-1}\right) \left(\displaystyle \sum_{t=1}^{T} \displaystyle\sum_{s = t-T}^{t-1} k\left(\frac{s}{B_{T}}\right)Y_{t-s-1}^2 \right)^{-1}. $$ For Step 3 and Step 4, we define $S^{\ast}$ using (ref), and we use i.i.d.\ resampling of $\{g_{T,t}(\hat{\beta})\}_{t=1}^{T}$. Since $\hat S(\beta)$ is strictly decreasing, we switch the sign of $S^{\ast}$ and $\hat S(\beta)$ and proceed as in Step 5 of Algorithm 1 to build a one-sided CI for $\beta_0$. In this particular case of linear models, we can rewrite $\hat{S}(\beta_0)$ as $T^{1/2}(\hat{\beta}-\beta_0)/\hat{\varsigma},$ where $\hat{\varsigma} = \hat{W}^{-1}\hat{\sigma}$ and $\hat{W} := T^{-1}\sum_{t=1}^{T} \partial g_{T,t}(\hat{\beta})/ \partial \beta.$ Similarly, we can rewrite the bootstrap counterpart $S^{\ast}$ as $T^{1/2}(\beta^{\ast}-\hat{\beta})/\varsigma^{\ast},$ where $\beta^{\ast}$ is a bootstrap estimate, $\varsigma^{\ast} = W^{\ast-1}\sigma^{\ast }$ and $W^{\ast} := T^{-1}\sum_{t=1}^{T} \partial g_{T,t}^{\ast}(\hat{\beta})/ \partial \beta.$ Then, the FMB CI is equivalent to $[\hat{\beta}-q^{\ast}_{1-\alpha/2}\hat{\sigma} , \hat{\beta}-q^{\ast}_{\alpha/2}\hat{\sigma}],$ where $q^{\ast}$ are quantiles of $S^{\ast}.$ In this representation, the FMB CI is similar to a percentile$-t$ bootstrap CI, up to our use of $\hat{\beta}$ instead of $\beta^{\ast}$ in the definition of $\varsigma^{\ast},$ avoiding to compute $\beta^{\ast}$, and in the multivariate case, to invert a matrix, for each bootstrap sample. This modification does not impact higher-order correctness, as shown in Section (ref).\\
Some further remarks on the other steps Algorithm (ref) are in order. First, Step 1 --- Step 3 (Lines 2-13) hinge on bootstrapping the moment indicator evaluated at $\hat\beta$. It justifies the adjective “fast” in the name of our resampling scheme, and bears some similarities to the already existing fast bootstrap (henceforth FB) methods; see shao_jackknife_1995, davidson_bootstrap_1999, andrews_higher-order_2002, salibian-barrera_bootstrapping_2002, goncalves_maximum_2004, salibian-barrera_principal_2006, salibian-barrera_fast_2008, camponovo_predictability_2013, armstrong_fast_2014, goncalves_bootstrapping_2019), and to the estimating function bootstrap (parzen_resampling_1994, hu_estimating_2000). However, FB methods typically rely on a first order von Mises expansion, which approximation error prevents the FB to be higher-order accurate.
Second, Step 3 (Line 12) computes the bootstrap statistic $S^{\ast},$ where the kernel $k$ creates a block of moment indicators evaluated at $\hat\beta$. The block of $g_t$ induced by the kernel is similar to a moving-average, as we emphasize in the name of our resampling scheme. The Moving Block Bootstrap (henceforth MBB) is the state-of-the-art higher-order correct alternative to FMB (gotze_second-order_1996, lahiri_edgeworth_1996). For MBB, the blocks are defined at the level of the observations, whereas in our case the convolution is applied to the moment indicators.
In the same family of groupwise resampling schemes, FMB is even more reminiscent of the Tapered Block Bootstrap of paparoditis_tapered_2001 (hereafter TBB), in the sense that we can view their tapered block as our moving-average kernel. The main difference is that our kernel has unbounded support, in contradistinction with their block tapering window. It gives FMB an advantage in the studentization: it allows us to use the Quadratic Spectral (QS) kernel, which is optimal in terms of asymptotic mean squared error according to andrews_heteroscedasticity_1991. parente_kernel_2018b have already pointed out such an advantage for a KBB variance estimator. Yet, the TBB and the KBB approach of parente_generalised_2018a both require $R$ bootstrap estimations. Thus, neither the TBB nor the KBB is fast and there is no result on their potential higher-order correctness. The higher-order correctness of the FMB approximation to the distribution of $\hat{S}$ comes from jointly considering two ingredients: the smoothing of moment indicators and the studentization.\footnote{The higher-order correctness of the FMB CI comes from the higher-order correctness of the latter FMB distribution and from the inversion step (Step 5 with the monotonicity condition or the cubic approximation of Remark (ref)).} Taken in isolation, each ingredient does not allow to show higher-order correctness of the FMB CI.
To summarize, we itemize in Table (ref) the main features of the discussed bootstrap schemes. We only list methodologies that are specifically designed for dependent data.
In this section, we explain how FMB can yield higher-order correct inference on a multivariate parameter. In Subsection (ref), we consider the Generalized Method of Moments. Then, we extend the setting to Generalized Empirical Likelihood Estimation in Subsection (ref). The asymptotic refinements (Section (ref)) also hold for the standard GMM case and are not tied to the use of GEL.
Assume we have to conduct inference on the multivariate parameter $\beta \in \mathcal{B} \subset \mathbb{R}^p$, where $\mathcal{B}$ is compact. We are given a random sample of $\mathbb{X}_t = \left\{ X_{t_1},...,X_{t_v} \right\}$ observed at $t=1,...,T,$ and we define a set of moment conditions $g: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R}^r$, with $r \geq p$, such that $\mathbb{E}\left[g(\mathbb{X}_t,\beta_0)\right] = 0$. To handle the serial dependence, we define the smoothed moment indicator $g_{T,t}\left(\beta\right)$ as in the univariate case
and $\bar{g}_T(\beta) := T^{-1}\sum_{t=1}^{T}g_{T,t}(\beta).$ Estimating $\beta_0$ via Generalized Method of Moments (hansen_large_1982) is the most popular approach in econometrics. In the next subsection, we discuss alternative estimators.
Step 1 --- Step 3 of the FMB Algorithm (ref) remain conceptually unchanged. As far as the bootstrap statistic is concerned, the principles of Step 4 and Step 5 stay the same as in Algorithm (ref), the main change being that the asymptotically pivotal statistic becomes:
where $\hat{\Omega} = \kappa_1^2(\kappa_2 T)^{-1}\sum_{t = 1}^{T} \{ g_{T,t}( \hat \beta ) - \bar{g}_T( \hat \beta ) \} \{ g_{T,t}( \hat \beta ) - \bar{g}_T( \hat \beta )\}^{\intercal}$ is a consistent estimator of the long-run covariance matrix of $T^{1/2}\bar{g}_{T} \left( \beta_0 \right)$, of rank $\nu = r$;\footnote{If the rank is lower than $r$, the covariance matrix is not invertible anymore and we resort to the generalized inverse, adapting the degrees of freedom of the $\mathcal{X}^2$ distribution accordingly (moore_generalized_1977).} see Section (ref). Standard results guarantee that $\hat Q$ is asymptotically $\mathcal{X}^2_{\nu}$. As in the univariate case, the statistic of interest is a particular value of a function, here $\hat{Q}: \mathbb{R}^{d v} \times \mathcal{B} \rightarrow \mathbb{R},$
We define the GMM estimator as $\hat{\beta} = \operatorname*{argmin}_{\beta \in \mathcal{B}} \hat{Q}(\beta).$ Similarly to the argument of Remark (ref), we also define the cubic approximation centered on the root-$T$ consistent estimator $\hat{\beta}:$
where the matrix $\hat{H} := \partial^2 \hat{Q}(\hat{\beta}) / \partial \beta^{\intercal} \partial \beta$. It allows us to build simply connected level sets, yielding higher-order correct Confidence Region (henceforth CR), as shown in Corollary (ref) (Section (ref)).
The bootstrap version of $\hat Q$ is
where $ \hat{\Omega}^{\ast} := {T}^{-1}\sum_{t = 1}^{T} \left\{ g^{\ast}_{T,t} - \bar{g}_T\left( \hat\beta \right) \right\} \left\{ g^{\ast}_{T,t} -\bar{g}_T\left( \hat\beta \right)\right\}^{\intercal}, $ with the asterisk denoting the same i.i.d.\ resampling scheme as in Algorithm (ref). Now we are ready to state the algorithm of our FMB in the over-identified case, made of five steps (lines 2-4, 5, 6-13, 14-15, 16-17).
A few remarks are in order. Step 3 uses (ref), where we recenter the bootstrap statistic. Indeed, the bootstrap expectation $\mathbb{E}^{\ast}\left[g^{\ast}_{T,t}\right] = \bar{g}_T(\hat{\beta}) \neq 0$ in the over-identified case. Thus, we subtract its expectation from $g_{T,t}^{\ast}$ to recenter the bootstrap variable. This operation is crucial to achieve higher-order accurate CR in Step 5 of Algorithm (ref) (see e.g.\ hall_bootstrap_1996).
Moreover, in contradistinction with the already existing FB methods, Step 3---4 mimic the variability of the covariance estimator $\hat{\Omega}$ (in (ref)) to achieve higher-order refinements. To this end, we use $\hat{\Omega}^{\ast}$ instead of $\hat{\Omega}$ in the definition of $Q^{\ast}$ (in (ref)), such that we randomize the bootstrap covariance estimator across the different bootstrap samples, and do not keep it fixed at $\hat{\Omega}$. Similar comment applies to Algorithm (ref).
Finally, to define the CR for $\beta_0$, we proceed similarly to Step 5 of Algorithm (ref). We set $q^{\ast}_{1-\alpha}$ such that $\mathbb{P}^{\ast}[ Q^{\ast} \leq q^{\ast}_{1-\alpha} ] = 1-\alpha$ and compute the CR as the subset $\mathcal{C}_{1-\alpha} := \left\{ \beta \in \mathcal{B} : \hat{Q}\left( \beta \right) \leq q^{\ast}_{1-\alpha} \right\}$. Thus, we get $\mathbb{P}[ \beta_0 \in \mathcal{C}_{1-\alpha} ] = \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq q^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq q^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T$. Under Assumptions (ref)---(ref) (Section (ref)), we show in Theorem (ref) that the remainder $R_T$ can be at most of order $o_p\left(T^{-1/2}\right),$ given a suitable choice of kernel $k$ and bandwidth $B_T$, which implies that $\mathcal{C}_{1-\alpha}$ is correct up to a higher order.
FMB is not tied to a particular estimation method. In order to show its wide applicability, we consider a general estimation method, which includes (among others) a version of GMM. The Continuously Updating Estimator (CUE) (hansen_finite-sample_1996), Empirical Likelihood (EL) (qin_empirical_1994, imbens_one-step_1996, kitamura_empirical_1997), and the Exponential Tilting (ET) (kitamura_information-theoretic_1997, imbens_information_1998) are all asymptotically equivalent to the (2S)GMM, but they tend to be less biased for small to moderate sample sizes (see altonji_small-sample_1996 for a Monte Carlo exploration, and newey_higher_2004, anatolyev_gmm_2005 for theoretical insights). Putting EL, ET, and CUE under the same umbrella, smith_gel_2011 introduces the Generalized Empirical Likelihood (GEL) criterion for time series data. We briefly describe it before extending our FMB to this general setting.
Let $\rho \left( \nu \right)$ be a concave function on an open interval $\mathcal{V} \in \mathbb{R}$ containing $0$. Writing $\rho_{\iota}(\nu) := \partial^{\iota} \rho (\nu) / \partial \nu^{\iota}$ and $\rho_\iota = \rho_{\iota}(0)$ for $\iota=0,1,$ the function $\rho \left( \nu \right)$ is standardized such that $\rho_1 = -1$. Defining a vector of auxiliary parameters $\lambda \in \Lambda_ { T } ( \beta )$ with $\Lambda_ { T } ( \beta ) := \left\{ \lambda \in \mathbb{R} ^ { r } : \kappa B_{T}^{-1/2} \lambda ^ { \intercal } g _ { T,t } ( \beta ) \in \mathcal{V} \right\},$ and $\kappa := \kappa_1 / \kappa_2$, smith_gel_2011 defines the GEL criterion as:
To derive an estimator of $\beta$, we first optimize criterion (ref) w.r.t.\ $\lambda$ for a given $\beta$, so that $\lambda \left( \beta \right) = \operatorname*{argsup}_{\lambda \in \Lambda_ { T } ( \beta )} \hat{P}(\beta,\lambda)$. Then, we define $\hat{\beta}$ as the solution to $\operatorname*{argmin}_{\beta \in \mathcal{B}} \hat{P}(\beta,\lambda\left( \beta \right))$.
Similarly to khundi_edgeworth_2012 and lee_asymptotic_2016, we are going to use the first order condition of the GEL criterion as a just-identified representation of the estimation problem to explain how to build the FMB statistic. Differentiating (ref) w.r.t.\ to $\lambda$ and $\beta$, we obtain:
We can see from (ref) that $\rho_{1} \left( \kappa B_{T}^{-1/2} \lambda \left( \beta \right) ^\intercal g_{T,t}\left( \beta \right) \right)$ gives weights to the observations such that the moment conditions in $g$ are always enforced in a given sample. GEL estimators are equivalent to some minimum discrepancy estimators based on the power-divergence family (see cressie_multinomial_1984), where the auxiliary vector parameter $\lambda$ corresponds to the Lagrange multiplier enforcing this empirical moment condition. Thus, if the original moment conditions in $g$ are correctly specified, the true (long-run) Lagrange multiplier is zero ($\lambda_0 = 0$). Applying our FMB to this setting only requires to define a quadratic statistic from the GEL first order conditions ((ref) and (ref)) evaluated at the true value of the parameters. Yet, replacing $\lambda = \lambda_0 = 0$ and $\beta = \beta_0$ in (ref) and (ref) boils down to the original $\bar{g}_T(\beta_0).$ Thus, the natural extension of FMB for GEL estimators requires to take the asymptotically pivotal statistic $\hat{Q}(\beta_0)$ as in the GMM case (ref). As a consequence, FMB in the GEL setting is exactly the same as in the GMM setting (see Algorithm (ref)), up to the initial estimator $\hat{\beta}.$ It is quite intuitive since, in absence of misspecification, the first order conditions of the GEL criterion convey the same information on $\beta_0$ as the moment condition $g$.
In the next theorems, we state that FMB CI and CR are higher-order correct. By construction, the higher-order correctness of the FMB CI and CR entirely hinges on our bootstrap approximation of the distribution for the test statistics $\hat{S}(\beta_0)$ (as in (ref)) and $\hat{Q}(\beta_0)$ (as in (ref)).
We start by itemizing the assumptions and regularity conditions. In the following, we keep using the shorthand notations $g_t = g_t\left(\beta_0\right)$ and $g_{T,t} = g_{T,t}(\beta_0).$ For any vector $V \in \mathbb{R}^n$, we write $\| V \| = (v_1^2 + ... + v_n^2)^{1/2}$, where $v_j$ is the $j$-th element of $V$. We make use of generic constants $C$, $\delta$ and $\epsilon$, whose value can differ from an expression to another. We define $\left\{g_t\right\}_{t \in \mathbb{Z}}$ on the probability space $\left(\Omega,\mathcal{A},\mathbb{P}\right)$. Let $\left\{ \mathcal{D}_t \right\} _ {t \in \mathbb{Z}}$ be a given sequence of sub-sigma-fields of $\mathcal{A}$, and $\mathcal{D}_{a}^{b} = \sigma\left\langle\left\{\mathcal{D}_{j} : a \leq j \leq b \right\}\right\rangle$. A straightforward example is to take $\mathcal{D}_t:=\sigma\left\langle g_t \right\rangle,$ but it is not always the most efficient choice to check the assumptions below (see gotze_asymptotic_1983 and gotze_asymptotic_1994 for practical examples). The higher-order correctness of FMB is subject to the following conditions (gotze_second-order_1996, lahiri_resampling_2010), which we assume to hold for $\left\{g_t\right\}_{t \in \mathbb{Z}}$:
Assumption (ref) is an identification condition. It is necessary also because we evaluate the moment conditions at $\hat{\beta}$ in the bootstrap samples. Since we are going to prove the validity and higher-order correctness of FMB using the $(s-2)$-th order Edgeworth expansion for the mean of gotze_asymptotic_1994, we require the moments in Assumption (ref) to be defined. Assumption (ref) ensures that the process $\{g_t\}$ is close enough to another process $\{g_{t,m}^{\ddagger}\}$, measurable w.r.t.\ sub-sigma-fields belonging to the sequence $\{\mathcal{D}_t\}$, whose dependence structure is controlled by the mixing condition in Assumption (ref). Assumption (ref) is the conditional Cramér condition of gotze_asymptotic_1994 for weakly dependent process. Assumption (ref) ensures that we can approximate the probability of $A \in \mathcal{D}_{t-p}^{t+p}$ given $\{ \mathcal{D}_k: k \neq t \}$ with an exponentially increasing accuracy, as the information in $\{ \mathcal{D}_k: 0 < \left\lvert t-k \right\rvert \leq m + p \}$ increases with $m$. We need Assumption (ref) on the regularity of the bootstrap characteristic function in order for the appropriate Cramér condition to hold for $S^{\ast}$ and $Q^{\ast}$. We refer to gotze_asymptotic_1983 for a general overview of the processes in agreement with our assumptions. As an example, the OLS moment indicators for the autoregressive parameters of a stationary AR($p$) process $Y_t = \sum_{i=1}^{p} \theta_i Y_{t-i} + e_t = \sum_{j=0}^\infty w_j e_{t-j}$ satisfy Assumptions (ref)---(ref), when $e_t \overset{i.i.d.} \sim \mathcal{N}(0,\sigma^2)$ and $\lvert w_j \rvert \leq \delta^{-1} \exp(- \delta j)$ $\forall j \in \mathbb{N},$ $\delta >0$. We verify the assumptions for this example in the Supplementary Material (SM.10). We can check along the same lines that an ACD($v$,0) model, and in particular the ACD(1,0) used in our Monte Carlo simulations, satisfies Assumptions (ref)---(ref).
Under the stated assumptions, we prove (see Appendix) the following two theorems, in which we denote by $\mathbb{P}^\ast$ the bootstrap probability measure given the data:
where $A$ is a set and $\mathbb{I}_{A}\{V\} = 1$ if $V \in A$ and 0 otherwise.
It is now apparent that the kernel $k$ impacts the bootstrap accuracy through the variance estimators in $\hat{S}$ and $\hat{Q}.$ As discussed by parzen_consistent_1957 and andrews_heteroscedasticity_1991, the bias of this estimator depends on the smoothness of the kernel $k^{\ast}$ at zero. The kernel $k^{\ast}\left( a \right) := \kappa_2^{-1} \int{k\left( b-a \right)k\left( b \right)db}$ is induced by the self-convolution of the smoothing kernel $k$. This bias is of order $O(B_T^{-q})$, where $q$ is the Parzen exponent of $k^{\ast}$, namely the maximal natural number such that $\lim_{a \rightarrow 0} \left(1-k^{\ast}(a)\right)/ \lvert a \rvert^{q}$ is finite. The bias is minimal for the rectangular kernel. However, as discussed in Section (ref), the resulting estimate $\hat{\Omega}$ is not necessarily positive semi-definite. In contrast, the QS kernel has an optimal Parzen exponent $q=2$ over all the kernels giving positive semi-definite estimators. Thus, the covariance matrix estimator with QS kernel has asymptotic MSE of order $O(B_T/T) + O\left( B_T^{-2} \right)$. Consequently, $B_T$ must grow faster than $T^{1/4}$ and slower than $T^{1/2}$, for the bootstrap error to be $o_p\left( T^{-1/2} \right)$. In Section (ref), we define the FMB CR by $\mathcal{C}_{1-\alpha} = \left\{ \beta \in \mathcal{B} : \hat{Q}\left( \beta \right) \leq {q}^{\ast}_{1-\alpha} \right\}$ (or alternatively using $\tilde{Q}(\beta)$ as in (ref)), where ${q}^{\ast}_{1-\alpha}$ is the approximation of $\hat{Q}(\beta_0)$ quantiles by FMB. Thus, $\mathbb{P}[ \beta_0 \in \mathcal{C}_{1-\alpha} ] = \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq {q}^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T$. Theorem (ref) gives the order $R_T = o_p\left(T^{-1/2}\right) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right).$ If we select a kernel $k$ and a bandwidth $B_T$ such that $O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right) = o_p\left(T^{-1/2}\right),$ it implies that $\mathcal{C}_{1-\alpha}$ is higher-order correct by construction, as the (first-order) Gaussian approximation is at best of order $O(T^{-1/2}).$ If $B_T = C T^{\gamma}$, $C > 0,$ we get the conditions $q >1$ and $(2q)^{-1} <\gamma < 1/2$. Similar arguments apply for Algorithm (ref). Part (ii) extends the higher-order correctness of the i.i.d.\ bootstrap of hu_estimating_2000 in the just-identified multivariate parameter case (see their Remark 10) to the over-identified case with time-dependent data.
When the monotonicity condition discussed in Section (ref) (Example (ref)) is violated (or is uneasy to check), we may use the alternative FMB CI or CR defined as $\mathcal{C}_{1-\alpha} := \left\{ \beta \in \mathcal{B} : \tilde{Q}\left( \beta \right) \leq q^{\ast}_{1-\alpha} \right\}$ (see Remark (ref) and (ref)), which are simply connected when $T$ is large enough and centered at $\hat{\beta}$. The following corollary (see the Supplementary Material for its proof) states the higher-order correctness of this alternative version of FMB CI and CR, based on the $\tilde{Q}$ statistic.
As a consequence, we have $\mathbb{P}[ \beta_0 \in \mathcal{C}_{1-\alpha} ] = \mathbb{P}[ \tilde{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq {q}^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T,$ with $R_T = o_p\left(T^{-1/2}\right) + O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right).$ Thus, similarly to the CI and CR based on $\hat{S}$ and $\hat{Q},$ the CI and CR based on $\tilde{Q}$ are higher-order correct if we select a kernel $k$ such that $O_p\left(B_T^{-q}\right) + O_p\left(B_T/T \right) = o_p\left(T^{-1/2}\right)$.
In this section, we give the necessary details on the appropriate way to estimate the long-run variance in the statistics $\hat{S}$ (as in (ref)) and $\hat{Q}$ (as in (ref)). We treat the general case of $\hat{\Omega}$ (as in (ref)), from which the univariate case can be easily deduced. Following andrews_heteroscedasticity_1991 and smith_gel_2011, we can obtain several estimators $\hat{\Omega}$ using different types of kernels. Let us give some examples of kernels useful in the implementation of FMB and their respective properties.
Thereby, we preferably use $k_{J}\left( x \right)$ of Example (ref) in (ref) to get the estimator $\hat{\Omega}$. Indeed, both from theory and simulations, the QS kernel is the optimal induced kernel in terms of asymptotic mean squared error (andrews_heteroscedasticity_1991), among all kernels giving positive semi-definite long-run variance estimators. The optimal bandwidth $B_T = O (T^{1/5})$ minimising the mean-squared error of the variance estimator with $k_J(x)$ (see Example 5) does not satisfy the condition $(1-\epsilon)/q < \gamma < \epsilon$ for any $\epsilon \in (0,1/2]$ (see Section 3) since $q = 2$ for the QS kernel. Then, we can take $\gamma = 1/3$, so that we match the condition for $q = 2$. wilhelm_optimal_2015 has already exhibited a discrepancy between the optimal choice for the HAC variance estimator and the optimal choice for a GMM point estimator. In his case, the optimal bandwidth for the point estimate is of the same order as the one minimizing the mean-squared error of the nonparametric plugin estimate, while the constants of proportionality are significantly different. In our case, the order is even different if we want to achieve higher-order correctness for FMB.
Alternatively, we may use the flat-top kernel version of the QS kernel, see politis_higher-order_2011, to get an even faster rate of convergence for the estimated long-run variance. Unfortunately, self-convolution of a kernel $k$ cannot induce a flat-top kernel $k^{\ast}$. Indeed, we know that it cannot be the case that $U = X + Y$, where the random variable $U$ is uniformly distributed on $[0, 1]$ (the flat-top part) and the random variables $X$ and $Y$ are independent and identically distributed (see Exercise 4.14.20 and its proof by contradiction in grimmett_thousand_2001). As a consequence, if we want to benefit from the smaller bias of the flat-top kernel, we should use a different kernel for the original statistic and the bootstrap one. A potential modification of FMB is to decouple the kernel $k^{\ast}$ used in the variance estimator of the original statistic, say a flat-top kernel, and the kernel $k$ used for the smoothed moment indicators. This version of FMB also achieves higher-order correctness since we maintain the asymptotic pivotal nature of the test statistics.
FMB user can manipulate the higher-order correct CR $\mathcal{C}_{1-\alpha}$ to obtain CR for a subset of parameters, or CI for a single parameter. In this section, we will consider separately the (possibly multivariate) parameter of interest $\beta^{(1)}_0,$ and the (possibly multivariate) nuisance parameter $\beta^{(2)}_0.$ Without loss of generality, we write the partition in the order $\beta = (\beta^{(1)\intercal},\beta^{(2)\intercal})^{\intercal}.$
We define the CR for $\beta^{(1)}_0$ as $\mathcal{C}^{(1)}_{1-\alpha} := \{ \beta^{(1)} : \hat{Q}(\beta^{(1)},\hat{\beta}^{(2)}) \leq q^{\ast}_{1-\alpha}\}.$ This manipulation preserves higher-order correctness when $\hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq \hat{Q} ( \beta_0 )$ $a.s.$, in the sense that it ensures $\mathbb{P}[ \beta^{(1)}_0 \in \mathcal{C}^{(1)}_{1-\alpha}] = \mathbb{P}[ \hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq {q}^{\ast}_{1-\alpha}] \geq \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}] = \mathbb{P}^{\ast} [ Q^{\ast} \leq {q}^{\ast}_{1-\alpha}] + R_T = 1-\alpha + R_T.$ This condition is satisfied, for instance, when $\hat{Q}$ is a monotonic function of the norm $\|\beta-\hat{\beta}\|,$ since $\| (\beta_0^{(1)\intercal}, \hat{\beta}^{(2)\intercal})^{\intercal} - \hat{\beta} \| \leq \| \beta_0 - \hat{\beta} \|$. It is also satisfied when the two sets of parameters $\beta^{(1)}$ and $\beta^{(2)}$ can be estimated independently from each other. For instance, in a two-step OLS estimation, it is the case in our ACD(1,0) example of Section (ref), by the Frisch-Waugh-Lovell Theorem, since we can rewrite the ACD(1,0) as an AR(1) process with orthogonal regressors.
If the inequality $\hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq \hat{Q} ( \beta_0 )$ fails to be true almost surely, we cannot guarantee the higher-order correctness of $\mathcal{C}^{(1)}_{1-\alpha}$. Nevertheless, the inequality $\mathbb{P}[ \hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \leq {q}^{\ast}_{1-\alpha}] \geq \mathbb{P}[ \hat{Q} ( \beta_0 ) \leq {q}^{\ast}_{1-\alpha}]$ remains true for large enough $T,$ as long as $\hat{Q} ( \beta^{(1)}_0, \hat{\beta}^{(2)} ) \overset{\mathcal{D}}\rightarrow \mathcal{X}_{\dim(\beta^{(1)}_0)}$ and $\hat{Q} ( \beta_0 ) \overset{\mathcal{D}}\rightarrow \mathcal{X}_{p}$. This general result implies that $\mathcal{C}^{(1)}_{1-\alpha}$ contains $\beta^{(1)}_0$ at least with probability $1 - \alpha$ asymptotically, but not always with higher-order refinements.
The drawback of no guarantee of higher-order refinements is inherent to the existing information on a multivariate parameter. The difficult task of reducing a CR for $\beta$ to a CR for $\beta^{(1)}$ is not directly entangled with the FMB, but more with the nature of dependence between $\beta^{(1)}$ and $\beta^{(2)}$ induced by the moment condition themselves. However, there exists a general way to build CR for $\beta_0^{(1)}$ while preserving higher-order refinements. Indeed, defining the profile statistic $\bar{Q}(\beta^{(1)}):=\inf_{\beta^{(2)}}\hat{Q}(\beta^{(1)},\beta^{(2)}),$ we get $\bar{Q}(\beta^{(1)}_0)\leq \hat{Q} ( \beta_0 )$ $a.s.$ by construction. Thus, by the same argument, the alternative CR $\bar{\mathcal{C}}^{(1)}_{1-\alpha} := \{ \beta^{(1)} : \bar{Q}(\beta^{(1)}) \leq q^{\ast}_{1-\alpha} \}$ preserves the higher-order refinements. This property comes with a cost, as $\bar{\mathcal{C}}^{(1)}_{1-\alpha}$ is generally more conservative than $\mathcal{C}^{(1)}_{1-\alpha}$ and it might be heavy to compute in high dimension.
Let us now make connections to the concept of Confidence Distributions (CD) that we use in our empirical application. It aims at answering the following question: can we also use a distribution function, or a “distribution estimator”, to estimate or test for a parameter of interest in frequentist inference in the style of a Bayesian posterior? (see the review paper by singh_confidence_2013). Natural point estimators include the median, the mean, and the maximum of the CD density (singh_combining_2005). That “distribution estimator” is named CD in agreement with the terminology coined by efron_fisher_1998, and traces back to the fiducial distribution of fisher_inverse_1930, albeit being a purely frequentist concept. It was introduced by hjort_confidence_2002 and its asymptotic extension by singh_combining_2005 (see also singh_confidence_2011, veronese_fiducial_2015, and the book-length presentation of hjort_confidence_2016). Example 2.4 of singh_combining_2005 discusses how a bootstrap distribution can yield a valid asymptotic CD, and Section 2.3.3 of singh_confidence_2013 how studentization can transmit higher-order accuracy in the i.i.d.\ case. Paralleling these recent developments in fiducial inference theory (see also the review paper of hannig_generalized_2016), we can exploit our FMB to produce a fast methodology to build an asymptotically higher-order correct CD as a by-product. Let us define the functions $H_{S}(\beta):=\mathbb{P}[\hat{S}\leq \hat{S}(\beta)]$ and its FMB counterpart $H^{\ast}_{S}(\beta):=\mathbb{P}^{\ast}[S^{\ast} \leq \hat{S}(\beta)]$, for $\beta \in \mathcal{B}$.
We omit the proof since the uniform error bound follows immediately from the proof of Theorem (ref). The second statement comes from the two conditions of Definition 1.1. of singh_combining_2005 being met, namely $H^{\ast}_{S}(\beta)$ is a cdf, and $H^{\ast}_{S}(\beta_0)$ is uniformly distributed on the unit interval when $T$ goes to infinity. Here, as clarified by pitman_statistics_1957, we follow indeed the frequentist view. In $H_{S}(\beta)$ and $H^{\ast}_{S}(\beta)$, randomness is not coming from the (non-random) parameter $\beta$, but from $\hat{S}$ and $S^{\ast}.$
As described in singh_combining_2005 (see also fraser_fiducial_1961, singh_confidence_2013), we can also use CD to get $p$-values. For example, the classical bootstrap $p$-value of $H_0 : \beta_0 \leq \beta$ versus $H_1: \beta_0 > \beta$ corresponds to $H^{\ast}_{S}(\beta)$, and the classical equal-tail bootstrap $p$-value of $H_0: \beta_0 = \beta$ versus $H_1: \beta_0 \neq \beta$ corresponds to $2 \min \{H^{\ast}_{S}(\beta), 1-H^{\ast}_{S}(\beta)\}$.
These $p$-values also benefit from higher-order correctness. Collecting them for different values of $\beta$ yields the so-called confidence curve $CV^{\ast}(\beta):=2 \min \{H^{\ast}_{S}(\beta), 1-H^{\ast}_{S}(\beta)\}$, introduced by birnbaum_confidence_1961 (see singh_confidence_2013 and hannig_generalized_2016 for illustrations). We can view this graphical tool as a piled-up form of two-sided CI of equal tails, at all confidence levels. We provide an example of such a plot in Figure (ref) for our empirical application in Section (ref), where we compare CI given by our FMB and first-order Gaussian asymptotics. Finally, coudin_finite-sample_2020 show how we can design a Hodges-Lehmann-type point estimator (hodges_estimates_1963) when a CD is constructed from a hypothesis test.
There exist analogue multivariate CD, for instance Definitions 5.1 and 5.2 in singh_confidence_2007. Similarly to the univariate case, we can apply FMB to achieve higher-order accuracy, as long as these multivariate confidence distributions are based on the test statistic $\hat{Q}(\beta_0)$.
To illustrate the applicability of FMB, we consider a simulation exercise on constructing CR for the parameters of an Autoregressive Conditional Duration (ACD) model (engle_autoregressive_1998). It is a model typically applied for the analysis of high-frequency data in finance, and more generally to model positive variables (e.g. volatility or volume) via a multiplicative error model; see e.g.\ hautsch_econometrics_2012 for a recent book-length presentation.
The duration is defined as the time lag between two consecutive events occurrence, namely $x_\ell := t_\ell-t_{\ell-1}$. Clearly, $x_\ell >0 $, for any $\ell\in \mathbb{T}$. We model $\mathbb{E} \left( x _ { \ell } | x _ { \ell - 1 } , \ldots , x _ { 1 } \right) = m _ { \ell } \left( x _ { \ell - 1 } , \ldots , x _ { 1 } ; \beta \right) := m _ { \ell }$, assuming the model $x_\ell = m_\ell \varepsilon_\ell$, with $\epsilon_\ell \overset{i.i.d.} \sim \mathcal{E}\left( 1 \right)$ for any $\ell$, with $\mathcal{E}\left( 1 \right)$ being an exponential random variable with mean one. Specifically, for the ACD(1, 1) specification, we have
for $\omega > 0$ $\beta_1 , \beta_2 \in \mathbb{R}^{+}$ and $\beta_1 + \beta_2 < 1$. When we take $\beta_2 = 0$, the ACD(1,0) model is in agreement with Assumptions (ref)--(ref) (in Section (ref)). Thus, we start our numerical experiment with this specification, conducting inference on $\beta:=(\omega,\beta_1)^{\intercal}$. We apply the optimal estimating functions of li_semiparametric_2000 and a moment condition which does not assume any specific functional form for the underlying innovation density, but relies on the unconditional expectation of $x_\ell$. Therefore, given a random sample of durations $(x_1,...,x_\ell,...,x_T)$, the vector of moment conditions for the $\ell$-th observation is $g_\ell ( \beta ) := \left( g_{1,\ell}( \beta), g_{2,\ell} (\beta), g_{3,\ell}( \beta)\right)^{\intercal},$ with $g_{1,\ell}\left( \beta\right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 } )( \partial m _ { \ell } / \partial \omega )$, $g_{2,\ell}\left( \beta \right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 } )( \partial m _ { \ell } / \partial \beta_1 )$, and $g_{3,\ell}\left( \beta\right) := x_\ell - \left( 1-\beta_1 \right)^{-1}$ (li_semiparametric_2000). Thus, we are in the over-identified case with $r=3$ for $p=2$.
In a second step, to get numerical insights on the applicability of our FMB, we extend our Monte Carlo experiment to a general ACD(1,1) model. The latter is non-markovian in the observations and does not fit our setting stricto sensu, since we assume the vector of observations in (ref) to be finite to ease notations and proofs. However, taking a large number $v$ of lagged durations in an ACD($v$,0) model is close to an ACD(1,1) model, and it meets our current theoretical framework.
We conduct inference on $\beta:=(\omega,\beta_1,\beta_2)^{\intercal}$ and our moment conditions for the $\ell$-th observation is $g_\ell ( \beta ) := \left( g_{1,\ell}( \beta), g_{2,\ell} (\beta), g_{3,\ell}( \beta), g_{4,\ell}( \beta) \right)^{\intercal},$ with $g_{1,\ell}\left( \beta\right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 } )( \partial m _ { \ell } / \partial \omega )$, $g_{2,\ell}\left( \beta \right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 } )( \partial m _ { \ell } / \partial \beta_1 )$, $g_{3,\ell}\left( \beta \right) := (\left( x_\ell - m_\ell \right) / m _ { \ell } ^ { 2 } ) ( \partial m _ { \ell } / \partial \beta_2 ),$ and $g_{4,\ell}\left( \beta\right) := x_\ell - \left( 1-\beta_1-\beta_2 \right)^{-1}$. We stay in the over-identified case with $r=4$ for $p=3$.
In the following, we compute by Monte Carlo simulations the coverage of FMB CR. We label the results $\hat{Q}_{FMB}$ for the FMB using $\hat{Q},$ $\tilde{Q}_{FMB}$ for the FMB using the cubic approximation of Remark (ref) and $\tilde{Q}_{FMB,2}$ for the FMB using the quadratic approximation of Remark (ref). To validate numerically our theoretical results, we compare the coverage of these FMB CR to some standard first-order correct alternatives. The first one, labeled as $\hat{Q}_{\mathcal{X}^2_r}$, defines CR as contours of the same statistic than FMB, but making use of the (first-order correct) $\mathcal{X}^2_{r}$ asymptotic distribution to compute the rejection probabilities. The second one is the standard elliptical contour of an asymptotically $\mathcal{X}^2_{p}$ distributed Wald statistic (labeled as $W_{\mathcal{X}^2_{p}}$), whose covariance matrix is a HAC estimator with bandwidth $B_T$.
To compare with a state-of-the-art competitor of FMB in terms of higher-order correctness, we also show the coverage of CR yielded by MBB. We adapt the MBB of inoue_bootstrapping_2006 to a Wald statistic, which is, in turn, an adaptation of gotze_second-order_1996. As the statistic defining the CR has to be asymptotically pivotal to obtain higher-order correctness, we choose to apply MBB to the Wald statistic, which we label as $W_{MBB}.$ We show a detailed CPU time comparison between FMB and MBB in the Supplementary Material (Table (ref)). In general, FMB appears to be at least $10^3$ times faster than MBB as expected.
In Table (ref), we display the empirical coverages for the ACD(1,0) model, for $B_T=3$ and $B_T=5.$ In Table (ref), we show the same outputs for the ACD(1,1) specification.
In line with our theoretical results, we observe that FMB performs well compared to its first-order correct competitors: the coverages are typically closer to their nominal level. It generally remains true for the $\tilde{Q}_{FMB}$ version of FMB CR, whose coverage is often very close to the one of the original FMB CR based on $\hat{Q}_{FMB}.$ The coverage of the quadratic $\tilde{Q}_{FMB,2}$ version of FMB CR is generally further away from the original FMB.
The Wald statistic seems to yield very erratic CR (which is additionally confirmed by unreported plots). The MBB version of the Wald statistic improves slightly on the asymptotic $\mathcal{X}^2_p$ distribution, without being convincing though.
Finally, the FMB presented here does not take advantage of all the potential fine-tuning, and this should leave room for practical improvement. First, we might improve FMB if the long-run variance $\Omega(\beta_0)$ is estimated with a less biased version of variance estimator, for instance carrying out a prewhitening step (andrews_improved_1992), or using a flat-top kernel (politis_higher-order_2011) as discussed in Section (ref). It should yield a smaller bootstrap error, as shown in Theorem (ref). Second, we stress that the moment indicators do not have necessarily the same dependence structure. Thus, smoothing the multivariate moment indicators with different bandwidths might further improve the coverage of FMB.
In this section, we illustrate how FMB performs on real data. We look at daily volumes of stock transaction (in millions), modeled with the same exponential ACD as in Section (ref) (see (ref)). We focus on data available online (Yahoo! Finance), for five stocks in three different sectors, namely bank, technology, and food. We compute the CR (for parameters $\omega$, $\beta_1$ and $\beta_2$), before the subprime crisis (2005), during the crisis (2008), and the current period (2018). The sample size of each period corresponds to the number of trading days, namely $T=252$ up to negligible variations from year to year. Before diving into a deeper analysis, we briefly describe the data at hand in Table (ref) below.
Table (ref) illustrates the larger variability of the volumes of transaction during 2008, as measured by the standard deviation (SD) and the interquartile range (IQR). The high skewness (SKN) and excess of kurtosis (KURT) typically indicate that a higher-order correct inferential procedure might be required in finite samples.
To investigate further the impact that asymmetry and fat tails may have on the conducted inference, we compute the FMB and the asymptotic normal (Asy) CR of nominal coverage $(1-\alpha) = 95\%$, for the ACD(1,1) parameters $\omega,$ $\beta_1$ and $\beta_2$ at each period. As FMB yields higher-order accurate CR by inverting probabilities of the test statistic $\hat{Q}(\beta_0)=\hat{Q}(\omega_0,\beta_{1,0},\beta_{2,0})$, we represent these trivariate CR by slicing them at the estimates $\hat{\omega},$ $\hat{\beta}_1$ and $\hat{\beta}_2$ (Table (ref)). Namely, we cut the CR by fixing the parameters that are not of interest to their estimated values. To keep Table (ref) concise, we do not report the estimate $\hat{\omega}$ and the intervals for the parameter $\omega$. We can deduce the former from Table (ref).\footnote{Using Section (ref) and volumes of transaction $\{x_t\}_{t=1}^{T}$ instead of durations, we have $\hat{\omega} \approx (1-\hat{\beta_1}-\hat{\beta_2})T^{-1}\sum_{\ell=1}^{T}x_{\ell},$ from the moment condition based on $g_{3,t}.$ We report the sample mean $T^{-1}\sum_{\ell=1}^{T}x_{\ell}$ in Table (ref).}
A few comments are in order. First of all, the different sectors exhibit very diverse reactions to the events happening in 2008. For instance, the food sector seems to be the most stable, while financial sector undergoes a huge variability, as we could expect. We can observe this either comparing non-critical periods to the crisis, or comparing the estimates and their CI before and after the crisis. For instance, the estimates for Unilever are almost the same before and after the crisis, as if the company has recovered the same volume behaviour. Coca-Cola looks equally stable with respect to the parameter $\beta_2$, which is almost unchanged after the crisis. Second, the estimate $\hat{\beta_1},$ respectively $\hat{\beta_2},$ seems to be larger, respectively smaller, during the crisis period. It is expected since $\beta_1$ reflects the sudden trading reactions due to changes in the expectations by the market participants during the crisis period. Thus, this feature of 2008 corresponds to an increase of the impact of news (shocks) on the volumes of transaction (via the parameter $\beta_1$), relative to persistence (via the parameter $\beta_2$). Finally, we see that the FMB CI are longer than the first-order correct Gaussian CI. It is in line with our Monte Carlo experiments, as available in Section (ref). Indeed, as the CR are defined by level sets of $\hat{Q}$, a longer CI corresponds to an adaptation of FMB to a skewed or fat-tailed distribution. Since the CI obtained by Gaussian approximation are typically shorter, we conclude that the routinely applied first-order asymptotic theory tends to underestimate the rejection probability, whereas FMB stays conservative. Our experience underpinned by several Monte Carlo simulations makes us expect that the distribution of $\hat{Q}(\beta_0)$ is more skewed or fat-tailed than the chi-squared; see the comparison between $\hat{Q}_{\mathcal{X}^2_r}$ and FMB in Table (ref).
Following our discussion on CD (Subsection (ref)), we illustrate here the link between our FMB CR and our previous definition of asymptotic confidence distribution $H^{\ast}_S(\beta)$, via the confidence curve $CV^{\ast}(\beta)$. Among the alternative ways to represent the former CR, marginalization allows us to build unconditional CI. Stacking the CR at different coverages $1-\alpha$ leads to a center-outward confidence curve for the multidimensional parameter $CV(\omega,\beta_1,\beta_2)$. For each CI, we integrate out the two parameters that are not of interest in $CV(\omega,\beta_1,\beta_2)$. This yields a different confidence curve $CV^{\ast}(\beta)$ for each $\beta \in \{\omega,\beta_1,\beta_2\},$ whose level sets give the equal-tailed CI. As an illustration of graphical use of these confidence curves (defined in Section (ref)), Figure (ref) reports a comparison between the FMB and Gaussian CI based on the FMB and Gaussian confidence curves. Again we observe that FMB is much more conservative.
We would like to thank the editor, the co-editor, and the referee for constructive criticism and numerous suggestions which have led to substantial improvements over the previous versions. We thank A.-P.\ Fortin, E.\ Paparoditis, D.\ Politis, and R.\ J.\ Smith for helpful comments and discussion, as well as participants in the North American Summer Meeting of the Econometric Society (Seattle, 2019), Geneva Finance Research Institute seminar, Workshop on Statistical Learning (Geneva, 2019-20), online RCEA Time Series Workshop 2021, and online IAAE Conference 2021.