EconBase
← Back to paper

Higher-order Expansions and Inference for Panel Data Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,148 characters · 12 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\def\spacingset#1{ {#1}} \spacingset{1}

titlepage\begin{center} { Higher-order Expansions and Inference\\for Panel Data Models} $^{\ast}${\sc Jiti Gao} and $^{\ast}${\sc Bin Peng} and $^{\ddag}${\sc Yayi Yan}\footnote{Emails: \url{[email removed]}; \url{[email removed]}; \url{[email removed]}.} $^{\ast}$Department of Econometrics and Business Statistics, Monash University\\ $^{\dag}$School of Statistics and Management, Shanghai University of Finance and Economics \today \end{center} \begin{abstract} In this paper, we propose a simple inferential method for a wide class of panel data models with a focus on such cases that have both serial correlation and cross-sectional dependence. In order to establish an asymptotic theory to support the inferential method, we develop some new and useful higher-order expansions, such as Berry-Esseen bound and Edgeworth Expansion, under a set of simple and general conditions. We further demonstrate the usefulness of these theoretical results by explicitly investigating a panel data model with interactive effects which nests many traditional panel data models as special cases. Finally, we show the superiority of our approach over several natural competitors using extensive numerical studies. Keywords: Dependent Wild Bootstrap, Edgeworth Expansion, Fund Performance Evaluation \end{abstract}

\spacingset{1.15}

Introduction

As we are embracing the era of data rich environment, the literature of panel data modelling starts shifting its focus to large $N$ and $T$ cases, and then nicely falls in the category of high-dimensional data analyses. Mathematically, we may denote a panel dataset as

eqnarray[eqnarray omitted — 149 chars of source]

in which $u_{it}$'s are scalers, and $[L] =\{1, 2,\ldots, L\}$ for a positive integer $L$. With a special focus, panel data modelling aims to capture the dependence along both dimensions of $u_{it}$ (e.g., HG2006,Petersen2009). However, how to establish valid inference has not been satisfactorily addressed to the best of our knowledge. In this article, we aim to contribute along this line of research. To formulate our concern, suppose that

eqnarray[eqnarray omitted — 96 chars of source]

where $\sum_{i,j=1}^{N} \sum_{t,s=1}^T |\alpha_{ij,ts}| =O(\mathbb{N})$ with $\mathbb{N}=NT$ for short. Note that (ref) does not only allow for correlation over both dimensions of $u_{it}$, but also permits heteroskedasticity. On this matter, Figure (ref) provides a graphical representation:

figure[figure omitted — 831 chars of source]

Here, CSD and TSA stand for cross-sectional dependence and time series autocorrelation respectively. In addition, one usually assumes the following Central Limit Theorem (CLT) holds:

eqnarray[eqnarray omitted — 184 chars of source]

where $\sigma_u^2 =\lim_{(N, T)\rightarrow (\infty, \infty)}\frac{1}{\mathbb{N}}\sum_{t,s=1}^T 1_{N}^\top E[U_{t} U_{s}^\top] 1_{N}>0$, and $1_{N}$ is a $N\times 1$ vector of ones. The conditions (ref) and (ref) are typical assumptions in the existing literature of panel studies (e.g., Chapter 2 of HG2006; Chapter 3 of FLW2011; Assumptions C and E of Bai). However, some fundamental questions have been left unanswered, e.g., (i) how to ensure (ref) and (ref) from a set of primitive conditions in view of the presence of correlation over both dimensions of $u_{it}$ ? (ii) does the convergence of (ref) achieve the usual Berry-Esseen bound ? and (iii) is the quantity $\sigma_u^2$ estimable ?

Among the aforementioned questions, the estimation of $\sigma_u^2$ is especially important for practical applications. Similar concerns have also been raised in the field of financial studies more than a decade ago. For example, Petersen2009 writes “Although the literature has used an assortment of methods to estimate standard errors in panel data sets, the chosen method is often incorrect and the literature provides little guidance to researchers as to which method should be used. In addition, some of the advice in the literature is simply wrong." Over the past couple of decades, although a variety of panel data models have been investigated, not much work has been done to improve the estimation of a quantity like $\sigma_u^2$. More often than not, one has to assume independence along at least one dimension of the dataset (e.g., Assumption 2 of Pesaran2006, Proposition 2 of Bai, Assumption 2.1 of Menzel2021, among others).

The studies sharing a similar concern with ours are goncalves_2011 and BAI2020. Specifically, goncalves_2011 studies a fixed effect panel data model, and proposes using the moving blocks bootstrap (MBB) technique, which allows for the error terms to have correlation over both dimensions. However, how to select the optimal block size is left unanswered, and to what extent CSD can be allowed is stated vaguely using some high level conditions (Assumptions 3 and 4 of the paper). BAI2020 consider a setting similar to goncalves_2011, and propose using a combination of the heteroskedasticity and autocorrelation consistent (HAC) covariance estimation approach and the thresholding technique, which then involves two tuning parameters --- one is the bandwidth of the HAC approach, and the other one is the threshold for penalizing entries of the covariance matrix. Such a procedure can be computationally expensive, and the optimal choices of the two tuning parameters are even more daunting. In addition, BAI2020 require the true covariance matrix of error terms to be sparse in a suitable sense, which is hard to justify in practical applications (giannone2021economic).

In this paper, we propose a simple inferential method for a wide class of panel data models, with a focus on the cases with correlation presenting in both dimensions. We first establish some higher-order expansions, such as Berry--Esseen bound and Edgeworth Expansion, under a set of simple and general conditions, which extend the results on time series (e.g., Jirak16, jirak2021sharp) to the panel data models. These results with sufficiently fast rates validate the use of distributional approximation for valid inference in finite sample studies. Then we develop a simple dependent wild bootstrap (DWB) procedure for valid inferences in practice. The DWB method is initially proposed by shao2010 to mimic the autocorrelation of time series, is easy to implement, and requires only one tuning parameter. Accordingly, we establish the necessary asymptotic properties, and conduct extensive numerical studies to examine the theoretical findings. In addition, we propose a data driven procedure to guide the selection of the optimal tuning parameter. It is noteworthy that the DWB covariance estimator is identical to a panel HAC covariance estimator, which is able to consistently estimate the true covariance matrix without requiring truncating cross-sectional dimension (e.g., p. 1252 in Bai) or penalising a large covariance matrix (e.g., BAI2020). Also, compared to the MBB, the DWB can better handle the missing values. The reason is that the MBB shuffles the blocks randomly, which then destroys the data structure. By contrast, the DWB method preserves the original information of the dataset in a natural manner. Last but not least, we demonstrate the usefulness of the DWB by explicitly investigating a panel data model with interactive effects which nests many traditional panel data models as special cases. As a by-product, we provide a solution to deal with bias correction and inference issues within one framework, which, to our knowledge, is the first result that has successfully addressed both issues.

The structure of the rest paper is as follows. Section (ref) provides the setup and the methodology, and establishes the necessary asymptotic properties. In Section (ref), we conduct extensive simulations to examine the finite sample properties of the theoretical findings. Section (ref) applies the proposed DWB approach to a real dataset by evaluating the aggregated mutual fund performance. Section (ref) concludes. Proofs, secondary results in nature, and extra simulations are given in the online supplementary appendices.

Before proceeding further, we introduce some mathematical symbols which will be repeatedly used in the article. $|\cdot|$ denotes the absolute value of a scalar or the spectral norm of a matrix; $\|\cdot\|$ denotes the Euclidean norm of a vector or the Frobenius norm of a matrix; for a matrix $A=\{a_{ij}\}$, let $|A|_1 = \max_{j}\sum_{i}|a_{ij}| $ and $|A|_\infty =\max_{i}\sum_{j}|a_{ij}|$; for a random variable $v$, let $\|v\|_q= (E|v|^q )^{1/q}$ for $q\ge 1$; $=_D$ denotes equality in distribution; $E^*[\cdot]$ and $\text{Pr}^*(\cdot)$ stand for the expectation and probability operations induced by the bootstrap procedure; $\to_P$ and $\to_D$ stand for convergence in probability and convergence in distribution respectively; $\lfloor q\rfloor$ stands for the largest integer not larger than $q$; for two numbers $a$ and $b$, $a\asymp b$ stands for $a=O(b)$ and $b=O(a)$; for two random variables $c$ and $d$, we write $c \simeq d$ if $c/d \to_P 1$; let $\Psi(\cdot)$ and $\psi(\cdot)$ be the cumulative distribution function (CDF) and the probability density function (PDF) of the standard normal distribution respectively; $M_A =I -A(A^\top A)^{-1}A^\top $ denotes the projection matrix for any matrix $A$ with full column rank; $1_N$ and $0_N$ stand for a $N\times 1$ vector of ones and a $N\times 1$ vector of zeros, respectively.

The Setup and Asymptotic Theory

In what follows, we present the basic framework and introduce some preliminary results without specifying any model in Section (ref), and propose the DWB method in Section (ref). In Section (ref), we demonstrate the usefulness of the newly proposed method by considering a panel date model with interactive effects. In Section (ref), we discuss how to handle the unbalanced panel dataset.

The Setup

We now focus on the dataset of (ref). Instead of assuming (ref) and (ref), we introduce a set of general conditions that can ensure the validity of both of them.

assumptionLet $\overline{U}_t = \frac{1}{\sqrt{N}}U_t^\top 1_{N}$, in which $U_t$ follows a process $U_t = g(\varepsilon_t, \varepsilon_{t-1}, \ldots)$ with $\varepsilon_t = (\varepsilon_{1t},\ldots, \varepsilon_{Nt})^\top$ being a sequence of independent and identically distributed (i.i.d.) random vectors, $E[U_t]=0_N$, and $g(\cdot)$ is a measurable function. In addition, let $ \overline{U}_t^* = \frac{1}{\sqrt{N}}U_t^{*\top} 1_{N}$, where $U_t^* = g(\varepsilon_t,\ldots,\varepsilon_{1}, \varepsilon_{0}^\prime,\varepsilon_{-1}^\prime,\ldots)$ is the coupled version of $U_t $, and $\{\varepsilon_t^\prime\}$ is an independent copy of $\{\varepsilon_t\}$. Suppose that $\sum_{t=0}^{\infty}t^2 \lambda_{t,\delta}^U < \infty$ for $\delta \ge 4$, where $\lambda_{t,\delta}^{U} = \|\overline{U}_t - \overline{U}_t^* \|_\delta$.

Assumption (ref) generalizes the nonlinear system discussed in Wu2005 to a panel data setting, and covers a wide range of data generating processes (DGPs) in the literature. Observe that $\lambda_{t,\delta}^{U}$ measures the dependence of $\overline{U}_t$ using the inputs $\varepsilon_0,\varepsilon_{-1},\ldots$ indicating that the cumulative impacts of $\varepsilon_0,\varepsilon_{-1},\ldots$ on future values are bounded in a suitable sense.

Below, we provide two examples to show Assumption (ref) is fulfilled by some well known DGPs, and the theoretical justification is given in the online appendices of the paper.

exampleConsider a high-dimensional MA($\infty$) process of the form: $U_t = \sum_{j=0}^{\infty} B_j\varepsilon_{t-j}$, where $B_j$'s are $N\times N$ matrices. Suppose that (a). $|B_j| = O(\rho^j)$ for some $|\rho|<1$; (b). $\{\varepsilon_{it}\}$ is independent over $i$; and (c). $E|\varepsilon_{it}|^\delta <\infty$. Then 1. $\{U_t\}$ fulfil Assumption (ref). 2. Suppose further that $|B_j|_1 = O(\rho^j)$, $|B_j|_\infty = O(\rho^j)$, and $E(\varepsilon_{it}^8)<\infty$. Then $\{U_t\}$ also satisfy Assumption C of Bai, which regulates the correlation along both dimensions of $u_{it}$.

Obviously, Example (ref).1 nests Assumption 2 of Pesaran2006 (i.e., $u_{it}= \sum_{j=0}^{\infty} b_{ij}\varepsilon_{i,t-j}$) as a special case, in which TSA presents, and heteroskedasticity over $i$ is allowed. Assuming $B_j\equiv 0$ for $j \ge 1$, the CSD matters only.

Example (ref).2 imposes more structure on the off-diagonal elements of $B_j$'s, as a result, both of CSD and TSA exist. Although it is not our focus, it is worth mentioning that some DGPs for network models (such as zhu2017network) are also covered by Assumption (ref). Intuitively, the condition on $|B_j|_{\infty}$ implies that the cumulative impacts of each $\varepsilon_{it}$ on $\{u_{1t},\ldots,u_{Nt}\}$ is bounded, which is consistent with the notion of weak CSD, while $|B_j|_1$ suggests that the cumulative impacts of $\{\varepsilon_{1t},\ldots,\varepsilon_{Nt}\}$ on each $u_{it}$ should be bounded.

exampleConsider a high-dimensional GARCH process: $U_t = \Omega^{1/2}V_{t}$, where $|\Omega|$ is bounded, $V_{t}=\left(v_{1t},\ldots,v_{Nt}\right)^\top$, $v_{it} = h_{it}^{1/2}\varepsilon_{it}$, and $h_{it} = c_i + \sum_{j=1}^{p}C_{ij}v_{i,t-j}^2 + \sum_{j=1}^{q} D_{ij}h_{i,t-j}$. For $\forall i\in [N]$, let (a). $c_i > 0, C_{i1},\ldots,C_{ip},D_{i1},\ldots,D_{iq}\geq 0$; (b). $\sum_{j=1}^{r}\|C_{ij} + D_{ij}\varepsilon_{i,0}^2\|_{\delta/2} < 1$ with $r = \max\{p,q\}$ for some $\delta \geq 4$. Then 1. $\{U_t\}$ fulfil Assumption (ref). 2. Suppose further that $|\Omega|_1 = O(1)$, $|\Omega|_\infty = O(1)$, and $E(\varepsilon_{it}^8)<\infty$. Then $\{U_t\}$ also satisfy Assumption C of Bai, which regulates the correlation along both dimensions of $u_{it}$.

Example (ref) presents an example involving the second moment induced by $\{U_t\}$, and infers that our framework can also be used to study those models focusing on conditional heteroskedasticity.

In fact, we can further prove that $U_t\otimes V_t=(U_{ti} V_{tj})_{i,j\le N}$ satisfies Assumption 1 provided that both $U_t$ and $V_t$ satisfy Assumption 1. By Minkowski inequality and Cauchy-Schwarz inequality, we have

eqnarray*[eqnarray* omitted — 290 chars of source]

which implies that $\sum_{t=1}^{\infty}t^2\left\|\frac{1}{N}1_{N^2}^\top (U_t\otimes V_t) - \frac{1}{N}1_{N^2}^\top (U_t^*\otimes V_t^*)\right\|_{\delta/2} < \infty$. In connection with Examples (ref) and (ref), we can conclude that squared linear processes, squared high dimensional GARCH processes, and their cross-product processes satisfy Assumption (ref), it therefore demonstrates the generality of our framework.

We are now ready to present the first theorem of this article.

theoremUnder Assumption (ref), as $(N,T)\to (\infty,\infty)$, 1. $S_{\mathbb{N}} \to_D N(0, \sigma_u^2)$, 2. $ \sup_{w\in \mathbb{R}}|\Pr(S_{\mathbb{N}} \le w) - \Psi_{\mathbb{N}}(w)| = O(T^{-1}(\log T)^5)$, where $ \Psi_{\mathbb{N}}(w) = \Psi\left(\frac{w}{s_{\mathbb{N}}}\right)+\frac{1}{6}\kappa_{\mathbb{N}}^3\left(1-\frac{w^2}{s_{\mathbb{N}}^2}\right)\psi \left(\frac{w}{s_{\mathbb{N}}}\right)$, $s_{\mathbb{N}}^2 = E[S_{\mathbb{N}}^2]$, and $\kappa_{\mathbb{N}}^3 = E[S_{\mathbb{N}}^3]$.

The first result of Theorem (ref) infers that the CLT holds under Assumption (ref), while the second result presents Edgeworth Expansion up to the second order. By Lemma (ref).3 of the online Appendix B (i.e., $E[S_{\mathbb{N}}^3] = O(1/\sqrt{T})$), we can simplify the second result to obtain Berry-Esseen bound:

eqnarray[eqnarray omitted — 171 chars of source]

By imposing more specific structures on the CSD, we can show $E[S_{\mathbb{N}}^3] = O(1/\sqrt{NT})$ and thus the rate of Berry-Esseen bound is given as follows.

corollarySuppose that (a). the conditions of Example (ref).2 hold; or (b). the conditions of Example (ref).2 hold. Then \begin{eqnarray} \sup_{w\in \mathbb{R}}\left|\Pr(S_{\mathbb{N}} \le w) -\Psi\left(\frac{w}{s_{\mathbb{N}}}\right)\right| = O\left(\max\left\{\frac{1}{\sqrt{NT}}, \frac{(\log T)^5}{T} \right\}\right). \end{eqnarray}

Looking at (ref) and (ref), in order to recover the distribution of $\Pr(S_{\mathbb{N}} \le w)$, one will need only to focus on the population quantity $E[S_{\mathbb{N}}^2]$. What's more, these results with sufficiently fast rates validate the use of the distributional approximation for valid inference in finite sample studies. In practice, the limiting distribution may not be useful if the approximation rate such as those involved in (ref) and (ref) is extremely slow (e.g., zhou2010simultaneous).

Up to this point, we have fully demonstrated the generality and applicability of Assumption (ref).

The DWB Method

In this subsection, we recover the asymptotic distribution $N(0,\sigma_u^2)$ of Theorem (ref) using the DWB approach. Specifically, for each bootstrap replication, we draw an $\ell$-dependent time series $\{\xi_t\, |\, t\in [T]\} $ satisfying the following condition.

assumptionLet $E[\xi_t ]=0$, $E|\xi_t |^2 =1$, $E|\xi_t|^4 <\infty$, $E[\xi_t \xi_s] =a (\frac{t-s}{\ell} )$, where $(\frac{1}{\ell}, \frac{\ell}{T})\to (0, 0)$ as $(\ell, T)\rightarrow (\infty, \infty)$, and $a(\cdot)$ is a symmetric kernel function defined on $[-1,1]$ satisfying that $a(\cdot)$ is Lipschitz continuous on $[-1,1]$, $a(0)=1$, and $K_a(x)=\int_{-\infty}^{\infty}a(u)e^{-iux}\, du \geq 0$ for $ x\in \mathbb{R}$.

The condition of $ K_a(x)$ ensures the semi-positive definiteness of the covariance matrix of $\{\xi_t\} $, while the restrictions on $a(\cdot)$ are satisfied by a number of commonly used kernels, such as the Bartlett and Parzen kernels. In practice, one may generate $\xi =(\xi _1,\ldots, \xi _T)^\top$ using $ N(0_T, \Sigma_{\xi})$ with $\Sigma_{\xi}= \{a\left(\frac{t-s}{\ell} \right) \}_{T\times T}$, while the normal distribution is not required in theory.

Accordingly, the bootstrap version of $S_{\mathbb{N}}$ is constructed as follows:

eqnarray[eqnarray omitted — 111 chars of source]

which, in connection with Assumptions (ref) and (ref), yields the following theorem of the paper.

theoremUnder Assumptions (ref) and (ref), as $(N,T)\to (\infty,\infty)$, we have 1. $\sup_{w\in \mathbb{R}}\left| \text{\normalfont Pr}^*(S_{\mathbb{N}}^* \le w) - \Pr(S_{\mathbb{N}}\le w) \right| = o_P(1)$, 2. $\sup_{w\in \mathbb{R}}\left| \text{\normalfont Pr}^*(S_{\mathbb{N}}^* \le w) - \Psi\left(\frac{w}{s_{\mathbb{N}}^*}\right) \right| = O_P\left(\sqrt{\frac{\ell}{T}}\right)$ with $s_{\mathbb{N}}^{*2} = E^*[S_{\mathbb{N}}^{*2}]$, 3. $\sup_{w\in \mathbb{R}^+}\left| \text{\normalfont Pr}^*(|S_{\mathbb{N}}^*| \le w) - \Pr(|S_{\mathbb{N}}| \le w) - 2\left(\Psi\left(\frac{w}{s_{\mathbb{N}}^*}\right) -\Psi\left(\frac{w}{s_{\mathbb{N}}}\right)\right) \right| = O_P\left(\frac{\ell}{T}\right)$.

The first result of Theorem (ref) shows that the bootstrap procedure can fully recover the asymptotic distribution $N(0,\sigma_u^2)$ of Theorem (ref). As a consequence, one can use the bootstrap draws to establish the corresponding confidence interval while allowing for both of CSD and TSA.

The second result of Theorem (ref) establishes the Berry--Esseen bound for the DWB procedure. To the best of our knowledge, the rate $\sqrt{\frac{\ell}{T}}$ is optimal and cannot be further improved (shergin1980convergence). This is because $\{\xi_t\}_{t=1}^T$ is a sequence of strongly $\ell$-dependent random variables (e.g., $E^*(\xi_{t}\xi_{t+\lfloor \ell/2\rfloor}) = a(1/2)$ as $\ell\to \infty$). As a result, the approximation rate of normal distribution is of the same order of $|s_{\mathbb{N}}^{*2} - s_{\mathbb{N}}^{2}|$ (see Theorem 2.3 below).

In the third result, by utilizing the fact that the first term of Edgeworth expansion is an even function of $w$, we provide a faster rate than that given in the second result. It then sheds light on how to select the optimal $\ell$. To see this point, note that the bootstrap draws offer a sample version of the form:

eqnarray[eqnarray omitted — 143 chars of source]

to consistently estimate $E[S_{\mathbb{N}}^2]$. Therefore, to select the optimal $\ell$ below, we minimise the mean squared error (MSE) between $E^*[S_{\mathbb{N}}^{*2}] $ and $E[S_{\mathbb{N}}^2]$, which is equivalent to minimise the MSE of coverage rates of the bootstrap confidence intervals\footnote{This point can be seen by setting $w = q_{\alpha}^*$, where $q_{\alpha}^*$ denotes the $\alpha$-th quantile of $|S_{\mathbb{N}}^*|$ such that $\text{\normalfont Pr}^*(|S_{\mathbb{N}}^*| \leq q_{\alpha}^*) = \alpha$. Hence, we have $\Pr(|S_{\mathbb{N}}| \le q_{\alpha}^*) = \alpha - \Psi\left(\frac{q_{\alpha}^*}{s_{\mathbb{N}}}\right)\frac{q_{\alpha}^*}{s_{\mathbb{N}}}(s_{\mathbb{N}}^{*2}-s_{\mathbb{N}}^2) + O_P(|s_{\mathbb{N}}^{*2}-s_{\mathbb{N}}^2|^2)+O_P(\ell/T)$.} according to the third result of Theorem 2.2, and is also a widely adopted criterion in the literature of HAC method (e.g., Andrews1991).

Before proceeding further, we impose one more condition on $a(\cdot)$.

assumptionFor $q \in [2]$, suppose that $\lim_{|x|\to 0}\frac{1 - a(x)}{|x|^q} = c_q$ for some real number $0 < c_q < \infty$.

Assumption (ref) is standard in the literature. For example, for the Bartlett kernel, $q=1$ and $c_1= 1$; for the Parzen, Tukey-Hanning, QS kernels, and the trapezoidal functions, $q=2$ and the values of $c_2$ vary but all satisfy $c_2<\infty$. We refer interested readers to KV2002 and PP2001 for the properties of Bartlett kernel and trapezoidal functions respectively, and to Andrews1991 for discussions on the other kernel functions.

theoremUnder Assumptions (ref)-(ref), as $(N,T)\to (\infty,\infty)$, {\normalfont {\bf Bias}:} $E\left(E^*[S_{\mathbb{N}}^{*2}]\right) - E[S_{\mathbb{N}}^2]= - \frac{c_q}{\ell^q} \Delta_1 + o(\ell^{-q})$, {\normalfont {\bf Variance}:} $\text{\normalfont Var}(E^*[S_{\mathbb{N}}^{*2}]) = \frac{2\ell}{T}\Delta_2 + o(\ell/T)$, where $\Delta_1= \sum_{k = -\infty}^{\infty}|k|^qE[\overline{U}_0\overline{U}_{k}]$ and $\Delta_2 = (E[S_{\mathbb{N}}^2])^2\, \int_{-1}^{1}a^2(x)\mathrm{d}x$.

By Theorem (ref), the MSE is minimized at

eqnarray[eqnarray omitted — 225 chars of source]

Theorem (ref) indicates that our DWB covariance estimator is a panel HAC covariance estimator, which is able to consistently estimate the true covariance matrix and does not require any cross-sectional parameter truncation (e.g., Bai) or regularization (e.g., BAI2020).

Up to this point, no specific model has been investigated. In what follows, we specifically apply the DWB method to a panel data model with interactive effects, which has attracted considerable attention since the seminal papers of Pesaran2006 and Bai, and nests lots of classic panel data models as special cases. Although a variety of extensions have been published in the past decade or so, to our knowledge, no work has successfully addressed the bias and inference issues simultaneously since Theorem 3 of Bai. That said, we shall tackle both problems together in the next subsection.

An Application of the DWB Method

From now on, we treat $u_{it}$'s as unobservable idiosyncratic errors, and consider a specific example to demonstrate the usefulness of the DWB method:

eqnarray[eqnarray omitted — 68 chars of source]

which is initially studied in Bai, and has been substantially extended since then (e.g., LiQianSu, Ando, among others). Here, only $\{Y_t, X_t \}$ are observable, while $\Gamma_0$ and $\{f_t\}$ can be correlated with $\{X_t\}$. As identifying the rank of $\Gamma_0$ is not the main focus here, we follow the aforementioned works to assume $f_t$ is a $p\times 1$ vector with $p$ being fixed and known.

Provided dependence along both dimensions of $u_{it}$, establishing valid inference for the model (ref) requires to tackle the following two challenging issues: (1). correct the estimation bias, and (2). estimate the asymptotic covariance matrix. Although the DWB method can address the latter one, one still needs to deal with the bias. That said, we provide a valid procedure to infer $\theta_0$ in what follows.

Consider the following objective function

eqnarray[eqnarray omitted — 121 chars of source]

where $\Gamma$ is a generic $N\times p$ matrix and satisfies $\frac{1}{N}\Gamma^\top \Gamma =I_p $ for the purpose of identification. Accordingly, we estimate $\theta_0$ and $\Gamma_0$ by minimizing (ref):

eqnarray[eqnarray omitted — 142 chars of source]

Also, we estimate $f_t$ and $U_t$ by $\widehat{f}_t=\frac{1}{N}\widehat{\Gamma}^\top (Y_t-X_t\widehat{\theta})$ and $\widehat{U}_t= Y_t-X_t\widehat{\theta}-\widehat{\Gamma}\widehat{f}_t$ respectively.

To proceed, we define a few extra notations. Let

eqnarray*[eqnarray* omitted — 180 chars of source]

where $\widetilde{X}_t = X_t - \frac{1}{T}\sum_{s=1}^TX_s f_s^\top ( \frac{F^\top F}{T} )^{-1} f_t $ and $F=(f_1,\ldots, f_T)^\top$. With these notations, we present the following lemma.

lemmaConsider the model (ref), and suppose (a). $E\|x_{it} \|^4<\infty$, and $\inf_{\Gamma} D(\Gamma) >0$ with $\Gamma^\top \Gamma /N =I_p$; (b). $\frac{1}{T}F^\top F\to_P \Sigma_F>0$ with $E\|f_t \|^4<\infty$, and $\frac{1}{N}\Gamma_0^\top \Gamma_0 \to_P\Sigma_{\Gamma}>0 $ with $E\| \gamma_{0i}\|^4<\infty$; (c). $\{U_t \}$ is independent of $\{X_t\}$, $\Gamma_0$ and $F$, where $x_{it}^\top$ and $\gamma_{0i}^\top $ stand for the $i^{th}$ rows of $X_t$ and $\Gamma_0$ respectively. In addition, let $\{U_t \}$ satisfy the conditions of Example (ref).2, and let $N/T\to \rho$ with $\rho$ being a positive constant. Then \begin{eqnarray*} \sqrt{\mathbb{N}}(\widehat{\theta} -\theta_0) \to_D N(\rho^{1/2}\mu_B+\rho^{-1/2}\mu_C, \Sigma_1^{-1}\Sigma_2\Sigma_1^{-1}), \end{eqnarray*} where $\mu_B = \operatorname*{p\!\lim} \mu_{\mathbb{N},B}$, $\mu_C = \operatorname*{p\!\lim} \mu_{\mathbb{N},C}$, and \begin{eqnarray*} \mu_{\mathbb{N},B} &=& - D(\Gamma_0 )^{-1} \frac{1}{\mathbb{N}}\sum_{t,s=1}^T \frac{\widetilde{X}_t ^\top \Gamma_0}{N} \left(\frac{\Gamma_0^\top \Gamma_0}{N} \right)^{-1} \left( \frac{F^\top F}{T} \right)^{-1} f_s \sum_{i =1}^N u_{it}u_{is},\nonumber \\ \mu_{\mathbb{N},C} &=& - D(\Gamma_0)^{-1} \frac{1}{\mathbb{N}}\sum_{t=1}^T X_t^\top M_{\Gamma_0} \Omega \Gamma_0 \left(\frac{\Gamma_0^\top \Gamma_0}{N} \right)^{-1} \left(\frac{F^\top F}{T} \right)^{-1}f_t. \end{eqnarray*}

Lemma (ref) repeats Theorem 3 of Bai, but interchanges the $i$ and $t$ dimensions. The first three conditions in the body of this lemma are equivalent to Assumptions A, B and D of Bai, while his Assumption C has been justified by Example (ref).2.

Next, we deal with the two biases before adopting the DWB method. For notational simplicity, we write

eqnarray[eqnarray omitted — 173 chars of source]

where $W_t$ is formed by $\{X_t\}$ and $\Gamma_0$ only, and $\frac{1}{\sqrt{\mathbb{N}}} \sum_{t=1}^T W_t^\top U_t\to_D N(0, \Sigma_1^{-1}\Sigma_2\Sigma_1^{-1})$.

Up to this point, it is worth commenting on the two-way half panel jackknife technique of Chen2021 that partitions sample along both dimensions. We would like to point out that, under the current context, it is probably not a good idea to split sample along the cross-sectional dimension, as it will change the structure of $\Omega$ internally in this case. Also, CSD usually depends on the “distance" among individuals implicitly which is unknown in general, so how to split the sample along the cross-sectional dimension is unclear. On the other hand, time series is naturally ordered, so it makes more sense to work with the time dimension when splitting sample.

We now propose the following procedure. Without loss of generality, let $T$ be an even number. For the time dimension, we create another two new sets, $S_1 = \{1,\ldots, T/2 \} $ and $S_2= \{T/2+1,\ldots, T \}$, and define a bias corrected estimator as follows:

eqnarray*[eqnarray* omitted — 119 chars of source]

where $\widehat{\theta}_{S_1} $ and $\widehat{\theta}_{S_2}$ are obtained using sample from $[N]\otimes S_1 $ and $[N]\otimes S_2$ respectively. However, $\widehat{\theta}_{\text{bc}}$ cannot fully remove two biases, which should be expected in view of Chen2021. By (ref) of the online supplementary appendix, we know that

eqnarray*[eqnarray* omitted — 176 chars of source]

Thus, we further provide the following estimator to deal with $\mu_{\mathbb{N},C}$:

eqnarray*[eqnarray* omitted — 232 chars of source]

where $\widehat{\Omega} = \frac{1}{T}\sum_{s=1}^T (Y_s- X_s\widehat{\theta})(Y_s- X_s\widehat{\theta})^\top$. Consequently, the final form of the bias corrected estimator is given below:

eqnarray*[eqnarray* omitted — 108 chars of source]

Finally, in order to infer $\theta_0$, the bootstrap procedure is as follows.

1. For each bootstrap replication, let $Y_{t}^* =X_{t}^\top \widehat{\theta} + \widehat{\Gamma}\widehat{f}_t+\widehat{U}_t \xi_t$.

2. Using the bootstrap sample $\{Y_T^*, X_t \}$, calculate $\widehat{\theta}^*$:

eqnarray*[eqnarray* omitted — 150 chars of source]

3. Repeat the first two steps $R$ times.

For the above procedure, the following theorem\ holds.

theoremSuppose that the conditions of Lemma (ref) and Assumption (ref) hold. As $(N,T)\to (\infty,\infty)$, 1. $\sqrt{\mathbb{N}}(\widetilde{\theta}_{\normalfont\text{bc}}-\theta_0) \to_D N(0,\Sigma_1^{-1}\Sigma_2\Sigma_1^{-1})$, 2. $\sup_w \left|\text{\normalfont Pr}^*(\sqrt{\mathbb{N}}(\widehat{\theta}^*-\widehat{\theta}) \le w) - \text{\normalfont Pr}(\sqrt{\mathbb{N}}(\widetilde{\theta}_{\normalfont\text{bc}}-\theta_0) \le w)\right| =o_P(1).$

In Theorem (ref), the first result provides an unbiased estimator for $\theta_0$, while the second result recovers the distribution $N(0,\Sigma_1^{-1}\Sigma_2\Sigma_1^{-1})$ using the bootstrap draws. Our investigation on (ref) is now completed.

On Unbalanced Dataset

To close our investigation on the DWB method, we consider the following unbalanced panel dataset:

eqnarray[eqnarray omitted — 97 chars of source]

where $N_t$ varies with respect to $t$, and $\mathbb{N}=\sum_{t=1}^T N_t$. The structure of (ref) is widely adopted in the literature (e.g., Chapter 4 of HG2006), and also suits the mutual fund dataset of Section (ref).

To accommodate the missing values, we can rewrite $\overline{U}_t$ of Assumption (ref) as

eqnarray[eqnarray omitted — 92 chars of source]

where $\mathscr{L}_t$ is a $N\times 1$ vector with elements being $1$ and $0$ only to represent non-missing and missing respectively. By (ref) and (ref), $\|\mathscr{L}_t\|=\sqrt{N_t}$. Under some trivial modification, one can show that the established results still hold. For example, we may adopt the following condition.

assumptionSuppose that $\frac{\overline{N}T}{\mathbb{N}}\to c\in (0,\infty) $, where $\overline{N} = \max_t N_t$ and $c$ is a constant.

Assumption (ref) allows $\underline{N}=\min_t N_t$ to be a fixed value, however, the number of $N_t$'s being finite has to be negligible.

corollaryUnder Assumptions (ref), (ref) and (ref), as $(N,T)\to (\infty,\infty)$, \begin{eqnarray*} \sup_{w\in \mathbb{R}} \left|\normalfont Pr^*(S_{\mathbb{N}}^* \le w) - \Pr(S_{\mathbb{N}}\le w) \right| =o_P(1), \end{eqnarray*} where $S_{\mathbb{N}}^*$ and $S_{\mathbb{N}}$ are defined in an obvious manner using (ref).

Compared to the block bootstrap based studies, one more advantage of DWB is that it can better handle the missing values. Note that the block bootstrap shuffles the blocks randomly, as a consequence the positions of missing values will be different for each bootstrap replication. In a sense, shuffling the blocks may destroy the data structure. By contrast, the DWB method preserves the original information of the dataset much better.

Up to this point, we have finished the theoretical investigation. In the next section, we examine the theoretical results using extensive simulation studies, and compare the DWB method with some existing ones.

Simulations

In this section, we conduct simulations to validate the theoretical findings of Section (ref). First, we explain how to select $\ell_{\text{opt}}$ practically in Section (ref). Then we evaluate Theorem (ref) of Section (ref) and the example of Section (ref) respectively. For the sake of space, we only report some selected results below, and provide the extra simulation results in the online supplementary Appendix (ref) of this paper.

Numerical Implementation

We now discuss how to calculate $\ell_{\text{opt}}$. The quantity $qc_q^2\Delta_1^2/\Delta_2$ in (ref) in fact can be consistently estimated, so there is a data-driven $\widehat{\ell}_{\text{opt}}$. To see this, note that $c_q$ is decided by the kernel function, and is therefore known. Thus, we need only to focus on $\Delta_1$ and $\Delta_2$.

By Theorem (ref), $\frac{1}{T}\sum_{t,s=1}^T\overline{U}_t \overline{U}_s a\left(\frac{t-s}{T^{\nu_q}} \right)\to_P E[S_{\mathbb{N}}^2]$, where $\nu_q = 1/3$ if $q=1$, and $\nu_q =1/5$ if $q=2$ by (ref). Thus, $\widehat{\Delta}_2 \equiv \left(\frac{1}{T}\sum_{t,s=1}^T\overline{U}_t \overline{U}_s a\left(\frac{t-s}{T^{\nu_q}} \right) \right)^2 \int_{-1}^{1}a^2(x)\mathrm{d}x \to_P\Delta_2.$

For $\Delta_1$, let $\widehat{\Delta}_1 \equiv 2\sum_{k=1}^{Q_T} \frac{k^q}{T}\sum_{t=1}^{T-k}\overline{U}_t\overline{U}_{t+k}$, where $Q_T \asymp T^{2/(4q+5)} $ is a truncation parameter. Since $\text{Var}\left( \frac{1}{T}\sum_{t=1}^{T-k}\overline{U}_t\overline{U}_{t+k} - \sigma(k)\right)=O(1/T)$ by Lemma (ref).4, we require $Q_T \asymp T^{2/(4q+5)} $ to ensure $\widehat{\Delta}_1\to_P \Delta_1$.

Finally, we recommend the following data-driven bandwidth: $\widehat{\ell}_{\text{opt}} = \widehat{\ell} \vee \ell_{\min},$ where $\widehat{\ell}=(qc_q^2\widehat{\Delta}_1^2/\widehat{\Delta}_2 )^{1/(2q+1)}T^{1/(2q+1)}$, and $\ell_{\min} $ is a fixed value (say, $\ell_{\min} = 10$). The reason for having $\ell_{\min} $ is to boost the finite sample performance when $T$ is relatively small. Note that even for $T=200$, $T^{1/5}$ only returns 2.89, and it is also not guaranteed that the term $(qc_q^2\widehat{\Delta}_1^2/\widehat{\Delta}_2 )^{1/(2q+1)}$ will return a value greater than 1. Therefore, to avoid an unreasonably small $ \widehat{\ell}$, we use $\ell_{\min}$ to bound $\widehat{\ell}_{\text{opt}} $ from below in the numerical implementation, which does not alter any aforementioned theoretical argument. For the model considered in Section (ref), we simply replace $\{U_t\}$ with $\{\widehat{U}_t\}$.

Examination of Theorem (ref)

We are now ready to conduct the simulation study. The DGP is as follows: $U_t^* = \rho_u U_{t-1}^*+ \epsilon_t,$ where we consider both light tail and heavy tail behaviour for $\epsilon_t$:

center[center omitted — 160 chars of source]

with $ \Sigma_N^\epsilon=\{\delta_\epsilon^{|i-j|} \}_{N\times N}$, and $t_5$ stands for the $t$-distribution with a degree freedom of 5. We let $\rho_u ,\rho_\epsilon \in\{0.25, 0.5 \}$. To introduce heteroscedasticity, we further let $\mathbb{U}_i =\sqrt{1+i/N} \widetilde{U}_i$, where $\mathbb{U}_i = (u_{i1},\ldots, u_{iT})^\top$, and $\widetilde{U}_i =(U_{i1}^*,\ldots, U_{iT}^*)^\top$ with $U_{it}^*$ being the $i^{th}$ element of $U_t^*$. Thus, $u_{it}$ has weak correlation over both dimensions, and also has heteroskedasticity over $i$.

To implement the bootstrap procedure, $\xi_t$'s are generated in the same way as mentioned under Assumption (ref). We specifically consider two kernels in the following simulations:

1. Bartlett kernel: $\psi(w) =(1-|w|)I(|w|\le 1)$,

2. A trapezoidal function: $a(x) = \frac{\int_{-1}^{1}w(u)w(u+|x|)\mathrm{d}u}{\int_{-1}^{1}w^2(u)\mathrm{d}u}$,

where $w(u)=\frac{u}{0.43}I\left(u\in[0,0.43)\right)+I\left(u\in[0.43,0.57]\right)+\frac{1-u}{0.43}I\left(u\in(0.57,1]\right).$

The Bartlett kernel is well adopted in the literature for its simplicity (e.g., Andrews1991, goncalves_2011, BAI2020; among others), while the specific form of the trapezoidal function can be seen in shao2010. Regarding (ref), both kernel functions represent the cases with $q=1$ and $q=2$ respectively. The bandwidths $\ell_B$ and $\ell_T$ of the Bartlett kernel and the trapezoidal function are selected as in Section (ref), and, for each kernel we further consider $0.8\ell_j$ and $1.2\ell_j$ for $j\in \{B,T\}$ to examine the sensitivity.

For every generated dataset, we record the value $S_{\mathbb{N}}$ of (ref) and the 95% confidence interval (CI) yielded by the 399 bootstrap draws. After $R$ replications, we report

eqnarray*[eqnarray* omitted — 95 chars of source]

where $S_{\mathbb{N},m}$ and $\text{CI}_m$ respectively stand for the value of $S_{\mathbb{N}}$ and the 95% CI from the $m^{th}$ replication.

For the purpose of comparison, we first consider three traditional methods to calculate the 95% CI. Specifically, we estimate the variances as follows:

eqnarray*[eqnarray* omitted — 224 chars of source]

where $s_1^2$ is a consistent estimator of the variance when $u_{it}$ is independent over $(i,t)$, and $s_2^2$ and $s_3^2$ are consistent estimators of the variance provided that the observed $u_{it}$ is independent over either the cross-sectional or time dimension.

The second method considered for comparison is the MBB method of goncalves_2011. The block length $\ell_M$ is generated in the same way as in goncalves_2011, so we omit the details here. We further consider $\lfloor 0.8\ell_M \rfloor$ and $\lfloor1.2\ell_M\rfloor$ to examine the sensitivity. For each dataset, we also do 399 bootstrap draws to obtain the 95% CI.

The third method included for comparison is the approach of BAI2020 (referred to BCL below). The implementation is identical to Section 2.1 of BAI2020. For the two tuning parameters $L$ (for HAC) and $M$ (for penalization), we use $L=3,7,11$ and $M =0.1, 0.15, 0.2, 0.25$ as in Section 3 of their paper. We do not further provide the details of their approach as it is quite lengthy.

We let $R=1000$. Due to space limit, we only report the results of $(\rho_u ,\rho_\epsilon) = (0.25,0.5)$ in Table (ref) and Table (ref) of the main text, and report the extra results in Tables (ref)-(ref) of the online supplementary appendices. For the traditional methods, the size is always greater than 5%, which is not surprising. As all of $s_j^2$ for $j=1,2,3$ just include a proportion of the asymptotic variance, we do expect the three traditional methods will over reject. As the sample size goes up, MBB seems to converge to the nominal rate (i.e., 5%) but slower than DWB. The BCL method tends to over reject, which might be due to the fact that many weak correlations get penalized by the thresholding method. The sizes of the DWB method are very close to the nominal one, and are quite robust in terms of the tail behaviour of $\epsilon_t$. Finally, it is noteworthy that the DWB method is not very sensitive to the choice of the kernel function, and is robust to different choices of the bandwidth. The findings are consistent across the tables.

center[center omitted — 75 chars of source]

Examination of Section (ref)

Having shown the superiority of the DWB, in this subsection we consider the model and the approach of Section (ref). When conducting inference, we focus on the DWB method only. The DGP is as follows: $Y_t =X_t\theta_0 +\Gamma_0 f_t + U_t,$ where $\theta_0=1$, and $U_t$ follows the identical DGP of Case 1 of Section (ref). For the factor structure, we let $\Gamma_0 = (\gamma_{01},\ldots, \gamma_{0N})^\top$ with $\gamma_{0i,\ell}\sim U(0.2, 2.2)$, and $f_t \sim N(0_{p}, I_{p})$, where $\gamma_{0i,\ell}$ stands for the $\ell^{th}$ element of $\gamma_{0i}$. We let $p=2$. To introduce a correlation between the regressors and the factor structure, we let $X_{t} = X_{t}^* + v_{t}$, where $X_{it}^* =|\gamma_{0i}'f_{t}|$, $X_{it}^* $ stands for the $i^{th}$ element of $X_{t}^* $, and $v_{t}\sim N(0_N,I_N)$. Based on the above DGP, $\{X_{t}\}$ are correlated with both $F=(f_1,\ldots, f_T)^\top$ and $\Gamma_0$.

The estimation procedure and the bootstrap draws are obtained in exactly the same way as documented above Theorem (ref). We calculate the size as follows:

eqnarray*[eqnarray* omitted — 140 chars of source]

where $\widetilde{\theta}_{\text{bc},m}$ and $\text{CI}_m$ stand for the bias corrected estimate and the 95% confidence interval based on the 399 bootstrap draws in the $m^{th}$ replication respectively.

After 1000 replications (i.e., $R=1000$), the results are reported in Table (ref). It is easy to see that as the sample size increases, the rejection rate approaches the nominal one, which infers two facts that the bias correction method works well, and the DWB method is able to recover the asymptotic covariance reasonably well. Due to the estimation errors, the performance is slightly worse than those in Tables (ref)-(ref) and Tables (ref)-(ref), which is acceptable.

center[center omitted — 55 chars of source]

An Empirical Study

In this section, we apply the proposed DWB method to a real dataset by evaluating the aggregated mutual fund performance.

A vast literature of financial economics has been devoted to evaluating the skills of the mutual fund managers. However, the existing results present many discrepancies berk2015measuring, which may be due to the fact that the analyses suffer from various modelling problems. For example, the traditional approach ignores the panel nature of the dataset, so the inter-fund information of the cross-sectional dimension has been largely ignored. In the same spirit, fama2010luck suggest that the TSA of the regression residuals may also alter the size of the usual fund alpha test. In this empirical study, we apply the DWB method of Section (ref), and aim to settle the discrepancies by accounting for the dependences along both dimensions of the dataset.

We obtain active U.S. equity mutual funds data from the Center for Research in Security Prices (CRSP) Survivor-Bias-Free Mutual Fund database for the period over Feb 1987 -- Sep 2017, and exclude the passive index funds (e.g., harvey2018detecting). As the data are monthly collected, the sample size is $T=368$. We only include the funds which have initial total net assets above 10 million, and have more than 80% of their holdings in equity markets. To mitigate degree of the unbalanced panel data structure, we consider three datasets by removing the funds with more than 20%, 25%, and 30% missing values\footnote{The thresholds 20%, 25%, and 30% are set arbitrarily. After different attempts, we note that the conclusion is not sensitive to the thresholds adopted here. In addition, we regard the three choices of the threshold as one type of robustness check.} during the entire period respectively, which leave us with 97, 114, and 132 mutual funds for different thresholds.

We consider the following unbalanced panel data model:

equation*[equation* omitted — 60 chars of source]

where $y_{it}$ is the net return (excluding fees and expenses) for fund $i$, $x_t$ includes the Fama-French-Carhart four-factor (including the market excess return factor, the Small-Minus-Big size factor, the High-Minus-Low value factor, the momentum factor), $\beta$ includes the slope coefficients, and $\alpha$ measures the abnormal performance of the mutual fund industry. We are interested in inferring $\alpha$, which is usually considered as an average indicator of the managerial skill of fund managers since the seminal work of jensen1968performance.

After running the OLS regression, we obtain the estimates of $\alpha $ and $\beta$ as follows:

eqnarray*[eqnarray* omitted — 381 chars of source]

The estimated residuals can then be calculated as follows:

eqnarray*[eqnarray* omitted — 90 chars of source]

To show the necessity of accounting for CSD and TSA, we first conduct the following two tests:

1. Examine TSA by conducting the Ljung-Box Q-test for each time series (i.e., $\{\widehat{u}_{i1},\ldots, \widehat{u}_{iT}\}$), and report the percentage of individuals having non-negligible autocorrelation;

2. Examine CSD by conducting the CD test\footnote{The asymptotic distribution of the CD test follows the standard normal distribution, so at the 5% significance level, the critical values are $\pm 1.96$. We refer interested readers to Pesaran2004 for more details.} of Pesaran2004 on $\widehat{u}_{it}$'s, and report the test statistics.

As shown in Table (ref), a non-negligible portion of individuals show evidence of TSA, while the CD test statistic always yields a significantly large value, which indicates the presence of CSD among the residuals. It is noteworthy that the computed value of the CD test increases, as the threshold (of removing individuals) becomes less restrictive, so it is a strong sign of CSD.

center[center omitted — 56 chars of source]

Below, we start reporting the 95% CI by using different methods. First, in Table (ref), we present the CIs using the three traditional methods as in Section (ref). It is clear the CIs generated by $s_1^2$ and $s_2^2$ indicate that the annualized aggregate mutual fund alpha is positively significant, which implies that the overall mutual fund industry can actually beat the market. However, the CIs generated by $s_3^2$ tell a different story. The results are not very surprising given that Table (ref) shows a reasonable amount of individuals fail to reject the null of the Ljung-Box Q-Test that assumes no time series autocorrelation.

center[center omitted — 74 chars of source]

In what follows, we consider the BCL, MBB, and DWB methods and focus on the CIs associated with the annualized alpha. The implementation of these methods is identical to that of Section (ref). The results are summarized in Table (ref). Note that in Table (ref) the BCL and MBB methods show mixed conclusions, while the DWB method consistently supports the result of $\alpha=0$ regardless the different combinations of the sample size, the bandwidth parameter, and the kernel function. Also, the consistent finding from the DWB method agrees with that of fama2010luck, in which they conclude that the mutual fund industry cannot beat the market.

Finally, in connection with the numerical results presented in Section (ref), we argue that the DWB method shows strong evidence of its superiority over some natural competitors in finite sample studies. We thus think the DWB method may yield more reliable results in practice.

Conclusion

Although a variety of panel data models have been investigated over the past couple of decades, not much work has been done to improve inferences associated with the estimation of the parameters-of-interest. In this paper, we have developed a simple dependent wild bootstrap procedure to establish inferences for a wide class of panel data models, including those with interactive effects. The proposed method allows for the error components to have CSD, TSA, and heteroskedasticity. The asymptotic properties, including Berry-Esseen bound and Edgeworth Expansion, have been established under a set of simple and general conditions. In addition, the newly proposed DWB method is easy to implement, and requires only one tuning parameter. We show the superiority of our approach over some natural competitors using extensive numerical studies. Last but not least, we demonstrate the usefulness of the DWB by explicitly investigating a panel data model with interactive effects which nests many traditional panel data models as special cases. As a by-product, we provide a solution to deal with bias correction and inference issue within one framework, which, to our knowledge, is the first result that has successfully addressed both issues.

In this paper, we have considered stationary time series for all individuals. We are aware of the growing literature on using bootstrap assisted methods to establish inferences for co-integrated time series models (e.g., PP2003, CNR2015, RJ2022). Along this line of research, Shao2015 provides a recent review on the bootstrap techniques frequently adopted. It would be interesting to investigate co-integrated panel data models (associated with certain cross-sectional dependence) using bootstrap methods. Such settings should be appealing in view of the increasing availability of large financial datasets over the past two decades. We leave possible extensions for future research.

{

}

{ {5pt}

table[table omitted — 6,343 chars of source]

}

{ {5pt}

table[table omitted — 6,346 chars of source]

}

{ {5pt}

table[table omitted — 6,450 chars of source]

}

{

table[table omitted — 336 chars of source]

}

{

table[table omitted — 1,725 chars of source]
table[table omitted — 2,113 chars of source]

}

{

\setcounter{page}{1}

center[center omitted — 134 chars of source]

This documents includes Appendix A and Appendix B. Overall, the structure is as follows.

In Appendix A,

itemize• Appendix (ref) discusses the case with $E[U_t]\ne 0_N$; • Appendix (ref) provides an extra example to show the usefulness of the DWB method; • Appendix (ref) provides some extra simulation results; • Appendix (ref) outlines the theoretical development, presents some notations which will be used throughout the theoretical development, and also provides some useful bounds; • Appendix (ref) presents the proofs of the main results.

In Appendix B,

itemize• Appendix (ref) introduces a few definitions to facilitate development of the preliminary lemmas; • Appendix (ref) summaries the preliminary lemmas; • Appendix (ref) provides the proofs of the preliminary lemmas.