EconBase
← Back to paper

Nickell Meets Stambaugh: A Tale of Two Biases in Panel Predictive Regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,176 characters · 19 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nickell Meets Stambaugh: A Tale of Two Biases in Panel Predictive Regressions

abstractIn panel predictive regressions with persistent covariates, coexistence of the Nickell bias and the Stambaugh bias imposes challenges for hypothesis testing. This paper introduces a new estimator, the IVX-X-Jackknife (IVXJ), which effectively removes this composite bias and reinstates standard inferential procedures. The IVXJ estimator is inspired by the IVX technique in time series. In panel data where the cross section is of the same order as the time dimension, the bias of the baseline panel IVX estimator can be corrected via an analytical formula by leveraging an innovative X-Jackknife scheme that divides the time dimension into the odd and even indices. IVXJ is the first procedure that achieves unified inference across a wide range of modes of persistence in panel predictive regressions, whereas such unified inference is unattainable for the popular within-group estimator. Extended to accommodate long-horizon predictions with multiple regressions, IVXJ is used to examine the impact of debt levels on financial crises by panel local projection. Our empirics provide comparable results across different categories of debt. \noindentKey words: bias correction, dynamic panel, local projection, persistence, macro-finance \noindentJEL code: C33, C53, E17

Introduction

Prediction has been one of the fundamental tasks of econometrics since its inception. The least squares (LS) is the default estimator for a linear predictive regression

equation[equation omitted — 96 chars of source]

where the regressor $x_{t}$ is uncorrelated with the future error term $e_{t+1}$. Nevertheless, LS incurs a finite sample bias in the time series context, which is known as the Stambaugh bias stambaugh1999predictive. The Stambaugh bias is particularly severe when $x_{t}$ is persistent, for example in the autoregression of order one (AR(1)) form

equation[equation omitted — 60 chars of source]

with the AR parameter $\rho^{*}$ close to 1. Here the bias becomes a first-order issue in that it will substantially distort the size of the conventional testing procedure based on the perceived asymptotic standard normal distribution $\mathcal{N}(0,1)$ of the $t$-statistic.

How to conduct valid hypothesis testing in predictive regressions has spanned into a large literature; see phillips2015halbert for a survey. phillips2009econometric's proposal, called IVX, is to use an instrumental variable (IV) generated only by $x_{t}$ and run a two-stage least squares estimation. IVX enjoys standard asymptotic distributions, allowing users to refer to the critical values based on $\mathcal{N}(0,1)$ or the $\chi^{2}$ distribution for hypothesis testing. The IVX method has witnessed many applications and extensions, for instance phillips2013predictive, Kostakis2015, Xu2020, hjalmarsson2022long, and Demetrescu2023, to name a few.

With the advent of rich economic datasets covering cross sections of countries and states, the time series predictive regression has been introduced into panel data, for example hjalmarsson2010predicting and westerlund2017testing. Panel data is featured by unobservable individual-specific fixed effects (FE). A rigorous theoretical development in dynamic panel data models is challenging when the number of individuals/states/countries $n$ over the cross section and the number of time periods $T$ are comparably large, where a non-vanishing bias spoils the standard asymptotic distribution of the $t$-statistic. nickell1981biases points out the bias of the within-group (WG) estimator in stationary dynamic panels, now well-known as the Nickell bias.

This paper focuses on the bias of commonly used estimators in panel predictive regressions with potentially persistent regressors. To fix ideas, we first consider a target variable $y_{i,t+1}$ generated by the following linear model

align[align omitted — 137 chars of source]

where $x_{i,t}$ follows an AR(1) process

equation[equation omitted — 76 chars of source]

The slope coefficient $\beta^{*}$ is the parameter of key interest, for it measures the predictability of $y_{i,t+1}$ using $x_{i,t}$. The AR(1) coefficient $\rho^{*}$ signifies the persistence of $x_{i,t}$. When $\rho^{*}\in(-1,1)$ is bounded away from unity, the vector $\left(y_{i,t},x_{i,t}\right)$ is stationary over time if the innovation vector $\left(e_{i,t},v_{i,t}\right)$ is stationary. This is a two-equation panel VAR system studied by holtz1988estimating, and the presence of the Nickell bias in the WG estimator and the analytical bias correction have been explored by hahn2002asymptotically under the “large-$n$-large-$T$” asymptotics.

In time series (a trivial “panel” with $n=1$), a widely used approach to characterizing the effect of a close-to-unity $\rho^{*}$ is modeling it as a deterministic function of the sample size $T$ in the form

equation[equation omitted — 82 chars of source]

where $c^{*}\in\mathbb{R}$ and $\gamma\in[0,1]$ are fixed constants, and the subscript $T$ is suppressed in $\rho^{*}$ when there is no ambiguity. Note that this setup also includes the familiar stationary case (Case I) when $\gamma=0$ and $c^{*}\in(-2,0)$, under which $\rho^{*}\in(-1,1)$ is a constant. As an asymptotic device, the representation in ((ref)) accommodates a persistent $x_{i,t}$, where “persistent” means $\rho^{*}\to1$ as $T\to\infty$, which is the source of non-trivial Stambaugh bias. The following modes of persistence are covered by ((ref)): Case II---mildly integrated (MI, $c^{*}<0$ and $\gamma\in\left(0,1\right)$); Case III---locally integrated (LI, $c^{*}<0$ and $\gamma=1$); Case IV---unit root (UR, $c^{*}=0$ and $\gamma=1$); and Case V---locally explosive (LE, $c^{*}>0$ and $\gamma=1$). For convenience, we wrap Cases III-V into the category of local unit root (LUR, $c^{*}\in\mathbb{R}$ and $\gamma=1$).

We keep an agnostic attitude about the degrees of persistence. In theoretical development, we adopt phillips1999linear's joint asymptotics to allow $n$ and $T$ simultaneously passing toward infinity. In panel data, the time series Stambaugh bias will be carried over and fused with the Nickell bias, resulting in a composite Nickell-Stambaugh bias in the $t$-statistic with an order enlarged from $1/\sqrt{T^{1-\gamma}}$ to $\sqrt{n/T^{1-\gamma}}$. The inflated bias shifts the center of the $t$-statistic much further away from 0, imposing challenges upon the standard statistical inference which refers to the critical values from $\mathcal{N}(0,1)$. To reinstate the standard inference based on the $t$-statistic, one strategy is to find the analytical formula of the bias and deduct it from the estimator. The bias formula for WG depends primarily on $\rho^{*}$. Were there an “oracle” to reveal the true $\rho^{*}$, bias correction would be straightforward. Unfortunately, bias correction for WG by plugging in a consistent estimator $\widehat{\rho}$ is infeasible under $\gamma=1$, because correcting this excessively large bias demands an impossibly fast rate of convergence of $\hat{\rho}$.

Faced with the intrinsic difficulty of WG, we turn to IVX. Our baseline estimator $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$ is a panel analog of the time series counterpart, which constructs a mildly integrated IV from the original predictor. IVX also incurs the Nickell-Stambaugh bias, with the expression of the bias again relying on $\rho^{*}$. The silver lining is that the mildly integrated IV reduces the order of the bias to $o_{p}\bigl(\sqrt{n/T^{1-\gamma}}\bigr)$, thereby providing a small niche for bias correction by plugging a $\sqrt{nT^{1+\gamma}}$-consistent estimator $\widehat{\rho}$ into IVX's analytical bias formula.

Which $\widehat{\rho}$ is qualified? We find that the WG estimator for $\rho^{*}$ in the panel AR fails the task due to its own bias. Alternatively, han2014x's X-differencing (XD) estimator removes the bias in the stationary case (Case I) and the exact unit root case (Case IV), but it does not cope with Cases II, III, and V, where $\rho^{*}\to1$ yet $\rho^{*}\neq1$. We are unaware of any existing estimator eligible for bias correction under all Cases I--V.

To this end, we propose an innovative solution that delivers valid statistical inference for all the types of regressor persistence under consideration. We look for $\rho^{*}$ by a new estimator, called X-Jackknife (XJ). XJ is a modified version of XD that eradicates the Nickell bias in the panel AR(1) model ((ref)) while retains signal strength; hence it enjoys the desirable rate of convergence. Specifically, XJ is devised by splitting the time dimension into odd indices and even indices, a unique jackknife scheme that is different from the common leave-one-out jackknife hahn2004jackknife or half-panel jackknife dhaene2015split. Plugging XJ into the bias formula of $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$, we produce a bias corrected IVX estimator named IVX-X-Jackknife (IVXJ). IVXJ restores the asymptotic normality centered at zero, and thus the standard inferential procedure follows.

The main contribution of this paper is to offer the new estimator IVXJ. It fuses two distinctive estimators in a creative way. In time series IVX's finite sample bias is negligible; however, its asymptotic bias surfaces and undermines statistical inference in panel data. XJ, on the other hand, is proposed to remove the Nickell bias in the panel AR model, to which the odd-even splitting scheme is tailored. To broaden its usability, we further extend IVXJ from the simple predictive model (((ref)) and ((ref))) to multivariate models and long-horizon prediction, which incorporate the important application of panel local projection jorda2005estimation.

The key theoretical insight lies in the delicate interplay of distinctive biases, in particular the bias of the plug-in estimator in the analytical formula. The lessons we learned from the drawbacks of other potential estimators culminate in the IVXJ estimator as a proper solution. To the best of our knowledge, IVXJ is the first and only available estimator that unifies the estimation and inferential procedure in Cases I-V from the stationary regime to the near-unity regime.

Literature review. This paper stands on strands of vast literature. In time series, the finite sample bias in the estimation of AR(1) is found by Hurwicz1950 and Kendall1954. For persistent time series, chan1987asymptotic study the asymptotics of LS when $\rho^{*}$ is local to unity, and phillips1987towards provides a comprehensive treatment. The bias in the LS estimator of $\rho^{*}$ carries over into that of the predictive coefficient $\beta^{*}$ stambaugh1999predictive. The impact of the persistent regressor on statistical inference has been explored by cavanagh1995inference, campbell2006efficient, and jansson2006optimal, and the effects on shrinkage estimation have been investigated by lee2022lasso and mei2022lasso. To avoid nonstandard inference, phillips2009econometric propose IVX, which serves as the baseline estimator of the procedure recommended by this paper.

The Nickell bias in panel data makes the WG estimator inconsistent under the large-$n$-small-$T$ asymptotics. Classical solutions are arellano1991some, arellano1995another, and many others. Anderson and Hsiao anderson1981estimation,anderson1982formulation consider the maximum likelihood under various orders of $n$ and $T$. In the large-$n$-large-$T$ asymptotics, hahn2002asymptotically and okui2010asymptotically investigate the analytical bias correction. Alternative proposals include split-panel jackknife dhaene2015split,chudik2018half, indirect inference gourieroux2010indirect,bao2023indirect, and X-differencing han2014x. While studies of the Nickell bias mainly focus on the panel AR, the bias is inherited by panel predictive regressions greenaway2012asymptotic.

Given persistent regressors in panel predictive regressions, Hjalmarsson hjalmarsson2008stambaugh,hjalmarsson2010predicting employs sequential asymptotics (first letting $T\to\infty$ and then $n\to\infty$, or conversely) to analyze the bias. The sequential limits are mathematically tractable, but they are insufficient in reasonably approximating the behaviors of the estimators under many relevant scenarios of large-$n$-and-large-$T$, say when $n$ and $T$ are of the same order. On the other hand, the joint asymptotics phillips1999linear allows $n$ and $T$ to go to infinity simultaneously. Though technically more challenging, the joint asymptotics paints a fuller picture of the finite sample behaviors --- particularly the biases --- than the sequential asymptotics does. Besides predictive models, biases are ubiquitous in large panel data; see hahn2004jackknife and fernandez2016individual for nonlinear models, and moon2015linear and moon2017dynamic for factor models.

The extension to multiple predictive horizons connects our study with the recent advancements of local projection jorda2005estimation, whose convenience in estimating the impulse response has drawn considerable research interest barnichon2019impulse,montiel2021local,plagborg2021local,herbst2024bias. While most empirical applications use the WG estimator for panel local projection, mei2023implicit show that the Nickell bias sustains asymptotically and invalidates the standard inference when $n$ and $T$ are comparably large. IVXJ further handles panel local projection with persistent regressors.

Layout. The rest of the paper is organized as follows. Section (ref) sets up the model, and uses a simple simulation to illustrate the empirical performance of WG and IVX with and without bias correction. The simulation makes it clear that IVXJ is the only estimator that delivers valid statistical inference in all the persistence cases under consideration. Section (ref) explains, in a heuristic manner, the impact of the biases in WG and IVX, respectively. The IVX-based estimator is followed by introducing the XJ estimator. Section (ref) elaborates the construction of XJ, formally develops the asymptotic properties of IVXJ with a univariate regressor, and further extends the theory into multivariate predictors and long-horizon prediction. This section is concluded by revealing the inconvenience of the WG-based estimators. An empirical example of predicting financial crises is carried out in Section (ref). An Appendix follows with the essential proofs of the main results. An Online Appendix includes more simulation results and technical proofs.

Notations. The symbols “$\to_p$” and “$\to_d$” signify convergence in probability and convergence in distribution, respectively. We use $\mathbb{E}_{s}(\cdot):=\mathbb{E}(\cdot|\mathcal{F}_{s})$ to denote conditional expectation with respect to a sigma-field $\mathcal{F}_{s}$. For a time series $x_{t}$, let $\Delta x_{t}:=x_{t}-x_{t-1}$ be its first-order difference. For a generic panel data random variable $x_{i,t}$ with $i=1,\dots,n$ and $t=1,\dots,T-1$, we use $\bar{x}_{i}:=\frac{1}{T-1}\sum_{t=1}^{T-1}x_{i,t}$ to denote the within-group average for the individual $i$, and $\tilde{x}_{i,t}:=x_{i,t}-\bar{x}_{i}$ is the within-group demeaned data. Let $a\wedge b:=\min\left\{ a,b\right\} $ and $a\vee b:=\max\left\{ a,b\right\} $. For a vector $\bm{a}$, we use $\|\bm{a}\|$ to denote its Euclidean norm.

Panel Predictive Regression

Setup

We start with a simple predictive regression of the two-equation system ((ref)) and ((ref)). Following phillips2019uniform, let the regressor $x_{i,t}$ follow a state space representation

equation[equation omitted — 157 chars of source]

It implies that $x_{i,t}$ admits the AR(1) form ((ref)) with the FE

equation[equation omitted — 65 chars of source]

under which ((ref)) can be rewritten as $x_{i,t+1}-\alpha_{i}=\rho^{*}\left(x_{i,t}-\alpha_{i}\right)+v_{i,t+1}.$ Such a specification of FE is standard in the literature han2014x.\footnote{Otherwise an unrestricted nonzero intercept in a local-to-unity process would accumulate and become a drift that dominates the stochastic trend, thus drastically complicating the asymptotic orders.} To enable theoretical analysis, we impose the following regularity conditions on the initial values $\delta_{i,0}$ and the drift $\alpha_{i}$.

assumption[Initial values and drift] Uniformly across all $i\leq n$, the conditional mean $\mathbb{E}(\delta_{i,0}|\alpha_{i})=0$, the unconditional variance $\mathbb{E}(\delta_{i,0}^{2})=O(|1-\rho^{*}|^{-1}\wedge T^{1-\varepsilon})$ where $\varepsilon>0$ is an arbitrarily small absolute constant, and $\mathbb{E}(\delta_{i,0}^{4}+\alpha_{i}^{4})=O(|1-\rho^{*}|^{-2}\wedge T^{2})$. (The convention $1/0=\infty$ is invoked if $\rho^{*}=1$).

In Assumption (ref), the zero conditional mean of the initial values facilitates the derivation of the analytical expression of the Nickell bias. The restriction on the second moment of the initial values implies that the unconditional variance of $\delta_{i,t}$ is of $O(T^{\gamma})$ when $\gamma<1$, and the small constant $\varepsilon$ takes effect only when $\gamma=1$ so that $\delta_{i,0}$ is of $o_{p}(\sqrt{T})$ and would not affect the asymptotics. The condition concerning the fourth moment is useful in showing consistency of the XJ estimator as well as consistent estimation of the variances when constructing the $t$-statistic. In essence, these assumptions are intended to ensure that the initial values and the drift are asymptotically negligible compared to the major components.

We impose the following regularity conditions on $\bm{w}_{i,t}:=(e_{i,t},v_{i,t})^{\prime}$, the underlying shocks of the observables $(y_{i,t},x_{i,t})$.

assumption[Innovations] \begin{enumerate} • The time series $\{\bm{w}_{i,t}\}_{t}$ are independently and identically distributed (i.i.d.) across $i$. • For each $i$, the sequence $\{\bm{w}_{i,t}\}_{t}$ is a strictly stationary and ergodic martingale difference sequence (m.d.s.) adaptive to the filtration $\{\mathcal{F}_{i,t}:=\sigma(\delta_{i,0},\alpha_{i},\bm{w}_{i,1},\dots,\bm{w}_{i,t})\}_{t\geq0}$, with the positive definite conditional covariance matrix \[ \mathbb{E}_{t-1}[\bm{w}_{i,t}\bm{w}_{i,t}']=\begin{bmatrix}\omega_{11}^{*} & \omega_{12}^{*}\\ \omega_{12}^{*} & \omega_{22}^{*} \end{bmatrix}=:\bm{\Omega}^{*} \] that is invariant over time, and $\mathbb{E}\|\boldsymbol{w}_{i,t}\|^{4}<C_{w}<\infty$ for some absolute constant $C_{w}$. \end{enumerate}

Assumption (ref) consists of simplifying conditions. Condition (i) allows for rigorous development of the joint asymptotics as in phillips1999linear. Condition (ii) implies that the dynamics of the observables is well specified so that the innovations are m.d.s., and the conditional homoskedasticity streamlines the derivation of the standard error in the $t$-statistic.

Algorithms: IVX versus WG

In this subsection we provide the estimation and inference procedure for $\beta^{*}$ based on IVX. For the time series models ((ref)) and ((ref)), many approaches to inferring $\beta^{*}$ have been proposed, such as the Bonferroni method cavanagh1995inference,campbell2006efficient and the conditional likelihood approach jansson2006optimal. However, these procedures are designed for a univariate regressor, and the respective limit distributions are nonstandard. The IVX method, on the other hand, achieves valid standard inference in time series multiple regressions and is robust to degrees of persistence magdalinos2009limit,Kostakis2015,phillips2013predictive. IVX leverages a mildly integrated IV by filtering the regressor. Specifically, in the time series context the IV is constructed by $z_{t}:=\sum_{s=1}^{t}\rho_{z}^{t-s}\Delta x_{t}$ where $\rho_{z}\in(0,1)$ is a user-specified parameter.

Our procedure is based on a panel version of IVX. First, we choose a parameter $\rho_{z}=\rho_{z,T}<1$ to produce the IV

equation[equation omitted — 91 chars of source]

The vanilla panel IVX estimator is

equation[equation omitted — 189 chars of source]

To correct the bias in $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$, we will estimate $\rho^{*}$ in ((ref)). The following algorithm describes the inferential procedure for a univariate panel predictive regression with a generic estimator $\widehat{\rho}$.

lyxalgorithm[IVX-based] \begin{description}[itemsep=0pt,leftmargin=*,parsep=0pt,topsep=0pt] • (slope coefficient estimation) Obtain $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$ in ((ref)), and a generic $\widehat{\rho}$ (to be specified) by estimating the model ((ref)). • (variance and covariance) Plug $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$ and $\widehat{\rho}$ from Step 1 into $\hat{\omega}_{12}\left(\beta,\rho\right)$ in ((ref)) and $\hat{\omega}_{11}(\beta)$ in ((ref)) to obtain $\widehat{\omega}_{12}$ and $\widehat{\omega}_{11}$, respectively. • (bias correction) Compute the bias corrected estimator \[ \hat{\beta}^{\mkern2mu{\mathrm{IVX}}}=\hat{\beta}^{\,\textup{IVX-bc}}+\widehat{\omega}_{12}\cdot b_{n,T}^{{\mathrm{IVX}}}\bigl(\widehat{\rho}\bigr) \] where $b_{n,T}^{{\mathrm{IVX}}}\left(\rho\right)$ is defined in ((ref)). • (confidence interval) Let $q$ be the desirable critical value according to the standard normal distribution, for example $q=1.96$ for a two-sided confidence interval with coverage probability 95%. The confidence interval is estimated as \[ \Bigl(\hat{\beta}^{\,\textup{IVX-bc}}-q\cdot\haat{\varsigma}^{\mathrm{IVX}}(\widehat{\omega}_{11},\widehat{\rho}),\ \hat{\beta}^{\,\textup{IVX-bc}}+q\cdot\haat{\varsigma}^{\mathrm{IVX}}(\widehat{\omega}_{11},\widehat{\rho})\Bigr) \] where the standard error $\haat{\varsigma}^{\mathrm{IVX}}(\omega_{11},\rho)$ is defined in ((ref)). \end{description}

In the proceeding simulation exercises, we will compare the empirical performance of the following four estimation methods: the vanilla IVX estimator $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$ in ((ref)), and the bias corrected $\hat{\beta}^{\,\textup{IVX-bc}}$ where the generic $\widehat{\rho}$ is substituted either by $\hat{\rho}^{\mkern2mu{\mathrm{WG}}}$ in ((ref)), or han2014x's $\hat{\rho}^{\mkern2mu{\mathrm{XD}}}$ in ((ref)), or our proposed XJ estimator $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$ in ((ref)). We refer to the three bias corrected estimators under the IVX-based algorithm as IVX-WG, IVX-XD, and IVXJ, respectively.

For comparison, we introduce the WG-based estimators. The vanilla WG estimator

equation[equation omitted — 177 chars of source]

is the most widely used estimator in panel data models where FE are present. The WG-based algorithm is parallel to Algorithm (ref). The difference lies in that in Steps 1--2 the main regression will be estimated by $\hat{\beta}^{\mkern2mu{\mathrm{WG}}}$, in Step 3 the bias formula $b_{n,T}^{{\mathrm{WG}}}(\rho)$ will follow ((ref)), and in Step 4 the standard error will be estimated by $\haat{\varsigma}^{\mathrm{WG}}(\omega_{11},\rho)$ in ((ref)). Similarly, besides the vanilla $\hat{\beta}^{\mkern2mu{\mathrm{WG}}}$, the biased corrected estimators are referred to as WG-WG, WG-XD, and WG-XJ, respectively.

Monte Carlo Simulations: An Illustration

figure[figure omitted — 580 chars of source]
figure[figure omitted — 581 chars of source]
figure[figure omitted — 413 chars of source]

In this subsection we provide Monte Carlo simulation evidence of the eight estimators mentioned above. For the data generating process (DGP) of the simple predictive regression ((ref)), we generate the dependent variable by setting the true coefficient $\beta^{*}=0$; that is, $x_{i,t}$ has no predictive power for $y_{i,t+1}$. We set the AR(1) coefficient in ((ref)) as $\rho^{*}\in\{0.60,0.95,0.99,1.00,1.01\}$ to reflect various degrees of the regressor's persistence, which are the finite sample embodiment of the stationary, MI, LI, UR, and LE regressors, respectively. In terms of the innovations, we normalize $\omega_{11}^{*}=\omega_{22}^{*}=1$ whereas the contemporaneous correlation $\omega_{12}^{*}\in\{0.70,0.95\}$ is specified to produce moderate and strong endogeneity, respectively.

For the IVX-based estimators, we adopt Kostakis2015's choices of $c_{z}=-1$ and $\theta=0.95$ in the user-specified persistence index $\rho_{z}=1+c_{z}/T^{\theta}$. The fixed effects $\mu_{y,i}$ are set as $\mu_{y,i}=T^{-1}\sum_{t}x_{i,t}$, to be correlated with the regressor; and the drift $\alpha_{i}$ and the initial value $\delta_{i,0}$ in ((ref)) are both independently drawn from $\mathcal{N}(0,1)$. We set the sample size $n=T\in\left\{ 30,50,100\right\} $.\footnote{We experimented with all the combinations of $n,T\in\left\{ 30,50,100\right\} $, and the results follow the theoretical predictions. Due to the limitation of spaces, we report the scenarios of $n=T$, and relegate the tabulated simulation results to Section (ref) of the Online Appendices. Table (ref) reports the empirical biases of the eight estimators and Table (ref) displays the corresponding empirical RMSEs. Table (ref) exhibits the coverage rates of confidence intervals corresponding to the results in Figure (ref). } The simulations are repeated for $5000$ times. The relative point estimation performance is clearly presented in Figures (ref) and (ref). The center of each bar is the empirical bias, and the height equals 2 times of the empirical standard deviation. It is obvious that the vanilla WG and IVX are severely biased, and the bias is more substantial when $\omega_{12}^{*}$ is bigger. All the bias correction methods are helpful when $\rho^{*}=0.60$, but only IVXJ is well centered around the true value under all the five cases of $\rho^{*}$ as the sample size gets large.

The empirical coverage probability based on the symmetric confidence interval is plotted in Figure (ref). Obviously, the vanilla WG estimator fails to work, and the distortion cannot be fixed by either the WG, XD, or XJ estimator of $\rho^{*}$ in the bias correction formula. IVXJ provides satisfactory coverage probability in all the cases. In contrast, IVX-WG and IVX-XD do not enjoy asymptotic validity. In particular, the coverage probability of IVX-XD strays far from 95% when $\rho^{*}=0.99$ and $1.01$, though it performs well in other cases. This phenomenon will be explained in Remark (ref).

In summary, IVXJ boasts competitive performance in terms of RMSE, and in terms of coverage probability it is the only estimator that demonstrates asymptotic validity in all the settings. In the following section, we will provide theoretical justifications for these observed phenomena.

Mechanisms

wrapfigure[wrapfigure omitted — 218 chars of source]

This section is a non-technical discussion of the mechanisms of the bias from the WG and IVX. Why each suffers from the Nickell bias in panel data? How to design a jackknife scheme to fix the latter? The heuristics will be followed by Section (ref) for a full development of asymptotic theory.

We will take phillips1999linear's joint asymptotics, which allows $n,T\to\infty$ simultaneously. We will pay special attention to the case when $n$ and $T$ grows to infinity in the same order, that is, $n/T\to r\in(0,1)$, which is the most relevant in practice. We refer to it as the leading asymptotic case.

Before we dive into asymptotics, we preview the theoretical results. The diagram of Figure (ref) depicts the estimation steps, the respective estimators, and the asymptotic validity under the leading asymptotic case. If either the main regression or the AR regression is estimated by WG, red lights flash in all the LUR cases. The amber lights under MI means that the validity depends on the user's choice of $\theta$ relative to $\gamma$ in the DGP, but since $\gamma$ is unknown there is no guarantee of success. IVX-XD is valid when $\rho^{*}$ is invariant with $T$, but it does not work with sample-size dependent $\rho^{*}$. IVXJ is the only procedure that secures green lights in all Cases I--V.

Bias in WG

We start with the time series context. Here the LS estimator for ((ref)) is $\hat{\beta}^{\mkern2mu\mathrm{LS}}=\sum_{t}\tilde{x}_{t}y_{t+1}/\sum_{t}\td{x}_{t}^{2}$, where we recall “tilde” means a within-group demeaned time series. The conditional homoskedasticity assumption (Assumption (ref)(ii) with $n=1$ where the time series is viewed as a trivial panel) admits a decomposition $e_{t}=(\omega_{12}^{*}/\omega_{22}^{*})v_{t}+u_{t}$, where by construction $u_{t}$ is orthogonal to $v_{t}$. It follows that

align[align omitted — 313 chars of source]

Throughout the discussion we assume $\omega_{12}^{*}\neq0$; otherwise there is no bias to be of concern. On the rightmost expression, the first term is unbiased as $u_{t+1}$ is uncorrelated with $\tilde{x}_{t}$ by the m.d.s.\ assumption (Assumption (ref)(ii)). Bias arises from the second term, as $v_{t+1}$ is correlated with $\tilde{x}_{t}$ via the latter's demeaning. Indeed, the bias is proportional to \[ \sum_{t}\tilde{x}_{t}v_{t+1}\Big/\sum_{t}\td{x}_{t}^{2}=\hat{\rho}^{\mkern2mu\mathrm{LS}}-\rho^{*}, \] that is, the bias of the LS estimator in time series AR model ((ref)). stambaugh1999predictive highlights the bias of $\hat{\beta}^{\mkern2mu\mathrm{LS}}$ in the time series predictive regression ((ref)), including the persistent cases when $\rho^{*}$ is close or equal to 1, and therefore the bias is called the Stambaugh bias. A well-known case is the exact unit root with $\rho^{*}=1$, where $T(\hat{\rho}^{\mkern2mu\mathrm{LS}}-1)$ converges in distribution to the Dicky-Fuller distribution --- a stable but nonstandard law. When $\rho^{*}\to1$ as $T\to\infty$, the nonstandard asymptotic distribution of $\hat{\rho}^{\mkern2mu\mathrm{LS}}$ translates to that of $\hat{\beta}^{\mkern2mu\mathrm{LS}}$, invalidating the standard inference based on the $t$-statistics.

Now we move to panel data, where the WG estimator

align*[align* omitted — 353 chars of source]

is the counterpart of the time series LS. When the regressor along the time dimension is stationary ($\gamma=0$), we have, in the leading asymptotic case where $n/T\to r\in\left(0,\infty\right)$,

equation[equation omitted — 198 chars of source]

where we can derive

align[align omitted — 163 chars of source]

On the left-hand side of ((ref)) the estimator is inflated by $\sqrt{nT}$, which we refer to as the (standard) panel factor. Equation ((ref)) highlights the fact that $\sqrt{nT}(\hat{\beta}^{\mkern2mu{\mathrm{WG}}}-\beta^{*})$ involves an asymptotic bias --- called the implicit Nickell bias by mei2023implicit --- that shifts the center of the asymptotic normal distribution away from 0.

When the regressor $x_{i,t}$ is nonstationary, that is, $\gamma\in(0,1]$, the order of the bias from \[ \sum_{i}\sum_{t}\tilde{x}_{i,t}v_{i,t+1}\Big/\sum_{i}\sum_{t}\td{x}_{i,t}^{2}=\hat{\rho}^{\mkern2mu{\mathrm{WG}}}-\rho^{*} \] depends on $\gamma$. This is the Nickell-Stambaugh bias to which the title of the paper alludes. We will show that the panel factor $\sqrt{nT}$ will be further multiplied by the persistence factor $\sqrt{T^{\gamma}}$ to deliver asymptotic normality: \[ \sqrt{nT^{1+\gamma}}\left[(\hat{\beta}^{\mkern2mu{\mathrm{WG}}}-\beta^{*})+\omega_{12}^{*}\cdot b_{n,T}^{\mathrm{WG}}\left(\rho^{*}\right)\right]\to_d\mathcal{N}(0,\Sigma^{\text{WG}}). \] What is worse, it is difficult to obtain a feasible estimator of $\rho^{*}$ to be plugged into the above expression. Again, we use $\hat{\rho}$ to denote a generic estimator of $\rho^{*}$. When $\hat{\rho}-\rho^{*}\to_p0$, a simple Taylor expansion gives

equation[equation omitted — 303 chars of source]

where “$h.o.t$” collects the higher order terms that we omit here in heuristic discussion. We can deduce that $\frac{\mathrm{d}}{\mathrm{d}\rho}b_{n,T}^{\mathrm{WG}}\left(\rho^{*}\right)=O_{p}(1)$ as $(n,T)\to\infty$. Therefore, to make the feasible estimator of the bias $b_{n,T}^{\mathrm{WG}}\left(\widehat{\rho}\right)$ asymptotically equivalent to $b_{n,T}^{\mathrm{WG}}\left(\rho^{*}\right)$, as to be detailed in Section (ref), it would request $\sqrt{nT^{1+\gamma}}(\hat{\rho}-\rho^{*})=o_{p}(1)$ as $(n,T)\to\infty$. This is mission impossible because a regular estimator $\hat{\rho}$ can at most achieve $\sqrt{nT^{1+\gamma}}(\hat{\rho}-\rho^{*})=O_{p}(1)$ in the panel AR regression, but not $o_{p}(1)$; see Remark (ref). The simulations in Section (ref) have provided evidence of the conspicuous bias in WG, and the undesirable performance of the bias-correction procedures when the baseline estimator is WG.

Bias in IVX

As mentioned in Section (ref), IVX generates an internal IV that is highly correlated with the original regressor yet orthogonal to the future error term. The internal IV requires no information outside of the system of observed variables. In asymptotic theory, the IV's persistence depends on \[ \rho_{z}=1+c_{z}/T^{\theta}, \] where $c_{z}<0$ and $\theta\in(0,1)$ are two hyperparameters to be decided by the user. The time series IVX estimator $\hat{\beta}^{\mkern2mu \text{ts-IVX}}=\sum_{t}\tilde{z}_{t}y_{t+1}\big/\sum_{t}\tilde{z}_{t}x_{t}$ admits the following decomposition

align*[align* omitted — 301 chars of source]

where the second equality again applies $e_{t}=(\omega_{12}^{*}/\omega_{22}^{*})v_{t}+u_{t}$. Since $u_{t}$ is orthogonal to $v_{t}$ by construction, the numerator of the first term $\sum_{t}\tilde{z}_{t}u_{t+1}$ has zero mean. Parallel to the analysis for the time series LS estimator ((ref)), the bias stems from the second term when $\omega_{12}^{*}\neq0$, since $v_{t+1}$ is correlated with the demeaned instrument $\tilde{z}_{t}$. The self-generated $z_{t}$ is mildly integrated by construction, and therefore circumvents the Stambaugh bias in the LS estimator. It can be shown that the bias caused by $\sum_{t}\tilde{z}_{t}v_{t+1}\big/\sum_{t}\tilde{z}_{t}x_{t}$ has an order $O_{p}(1/T^{2-(\theta\vee\gamma)})$. Since this order converges to zero sufficiently fast, the time series IVX is asymptotically unbiased and bias correction is thus unnecessary.

In panel data, however, IVX encounters additional challenges. The panel IVX estimator $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}=\sum_{i}\sum_{t}\tilde{z}_{i,t}y_{i,t+1}\big/\sum_{i}\sum_{t}\tilde{z}_{i,t}x_{i,t}$ can be decomposed as \[ \hat{\beta}^{\mkern2mu{\mathrm{IVX}}}-\beta^{*}=\frac{\sum_{i}\sum_{t}\tilde{z}_{i,t}e_{i,t+1}}{\sum_{i}\sum_{t}\tilde{z}_{i,t}x_{i,t}}=\frac{\sum_{i}\sum_{t}\tilde{z}_{i,t}u_{i,t+1}}{\sum_{i}\sum_{t}\tilde{z}_{i,t}x_{i,t}}+\frac{\omega_{12}^{*}}{\omega_{22}^{*}}\cdot\frac{\sum_{i}\sum_{t}\tilde{z}_{i,t}v_{i,t+1}}{\sum_{i}\sum_{t}\tilde{z}_{i,t}x_{i,t}}. \] Again, the bias comes from the second term due to the correlation between $v_{i,t+1}$ and $\tilde{z}_{i,t}$. In the oracle case when $\rho^{*}$ and $\omega_{12}^{*}$ are known, we have

equation[equation omitted — 241 chars of source]

where, after some calculation,

align[align omitted — 188 chars of source]

Shown in (ref) below, the bias $b_{n,T}^{{\mathrm{IVX}}}(\rho)$ remains $O_{p}(1/T^{2-(\theta\vee\gamma)})$, but the inflating factor on the left-hand side of ((ref)) is $\sqrt{nT^{1+(\theta\wedge\gamma)}}$ with an additional factor $\sqrt{n}$ from the large cross section. If $n/T\to r\in(0,\infty)$ and $\gamma+\theta+(\theta\vee\gamma)>2$ (e.g., $\gamma=1$ under which $x_{i,t}$ is LUR), then $O_{p}(1/T^{2-(\theta\vee\gamma)})$ cannot cancel out the inflating factor $\sqrt{nT^{1+(\theta\wedge\gamma)}}$, thereby moving the center of the asymptotic normal distribution of $\sqrt{nT^{1+(\theta\wedge\gamma)}}(\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}-\beta^{*})$ away from 0.

IVX keeps the standard panel factor $\sqrt{nT}$ on the left-hand side of ((ref)) and in the meantime its persistence factor is $\sqrt{T^{\theta\wedge\gamma}}$, in contrast to WG's $\sqrt{T^{\gamma}}$. When $\hat{\rho}-\rho^{*}\to_p0$, the Taylor expansion for a generic plug-in estimator of the bias becomes \[ \sqrt{nT^{1+(\theta\wedge\gamma)}}\left[b_{n,T}^{\mathrm{IVX}}\left(\widehat{\rho}\right)-b_{n,T}^{\mathrm{IVX}}\left(\rho^{*}\right)\right]=\frac{\mathrm{d}}{\mathrm{d}\rho}b_{n,T}^{\mathrm{IVX}}\left(\rho^{*}\right)\cdot\sqrt{nT^{1+(\theta\wedge\gamma)}}(\hat{\rho}-\rho^{*})+h.o.t. \] We can verify $\frac{\mathrm{d}}{\mathrm{d}\rho}b_{n,T}^{\mathrm{IVX}}\left(\rho^{*}\right)=O_{p}(1/T^{2-(\theta\vee\gamma)-\gamma})$, and therefore a desirable order

equation[equation omitted — 160 chars of source]

would be sufficiently small to remove the bias in $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$.

X-Jackknife for $\rho^{*}$

Which estimator satisfies ((ref)) under any $\gamma\in[0,1]$? Our response is the new XJ estimator, which we now introduce. Without loss of generality, we assume that $T$ is an even number. Define $\mathcal{O}:=\{3,5,\dots,T-1\}$ as the set of all odd numbers from $3$ to $T-1$, and $\mathcal{E}:=\{2,4,\dots,T-2\}$ as the set of all even numbers from $2$ to $T-2$. Both $\mathcal{O}$ and $\mathcal{E}$ contain $(T-2)/2$ elements. The XJ estimator is constructed as

equation[equation omitted — 436 chars of source]

where $\tilde{x}_{i,t}^{\mathcal{O}}:=x_{i,t}-\frac{2}{T-2}\sum_{t\in\mathcal{O}}x_{i,t}$, $\tilde{x}_{i,t}^{\mathcal{E}}:=x_{i,t}-\frac{2}{T-2}\sum_{t\in\mathcal{E}}x_{i,t}$, and

equation[equation omitted — 154 chars of source]

While jackknife has a long history in statistics, the interlacing jackknife scheme in $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$ is originally developed in this paper. The panel data literature has witnessed leave-one-out jackknife hahn2004jackknife, half-panel split along the time dimension dhaene2015split and along the cross section fernandez2016individual. To the best of our knowledge, we have not seen a jackknife that splits the sample with the odd-and-even indices.

Now that we have analyzed the bias of panel IVX, in the next section we will first show that the XJ estimator for $\rho^{*}$ achieves ((ref)), and then explain the ideas behind the design of XJ. Given the desirable convergence rate of XJ, we further verify the asymptotic normality of IVXJ and the validity of the standard inference, and extend it to accommodate multiple regressors and long prediction horizons. Finally, we discuss the impossibility of getting rid of the bias from the WG-based estimators.

Theory

Convergence Rate of X-Jackknife

Why does the XJ estimator in ((ref)) take such an unusual interlacing of the odd index set $\mathcal{O}$ and the even index set $\mathcal{E}$? Despite the distinctive appearance, it is nevertheless inspired by han2014x's X-Differencing (XD) estimator

align[align omitted — 236 chars of source]

A remark on the bias of the XD estimator under a persistent $x_{i,t}$ follows.

remThe unbiasedness of $\hat{\rho}^{\mkern2mu{\mathrm{XD}}}$ counts on the following orthogonality condition in han2014x: \begin{equation} \mathbb{E}\bigl[(x_{i,t-1}-x_{i,s+1})\left[(x_{i,t}-x_{i,s})-\rho^{*}(x_{i,t-1}-x_{i,s+1})\right]\bigr]=0\quadfor s+1<t, \end{equation} where the quantity inside the expectation operator appears in the numerator of \[ \hat{\rho}^{\mkern2mu{\mathrm{XD}}}-\rho^{*}=\frac{\sum_{i=1}^{n}\sum_{t=4}^{T}\sum_{s=1}^{t-3}(x_{i,t-1}-x_{i,s+1})\left[(x_{i,t}-x_{i,s})-\rho^{*}(x_{i,t-1}-x_{i,s+1})\right]}{\sum_{i=1}^{n}\sum_{t=4}^{T}\sum_{s=1}^{t-3}(x_{i,t-1}-x_{i,s+1})^{2}}. \] han2014x show that the left-hand side of ((ref)) equals $(\rho^{*t-s-2}-\rho^{*})\bigl[\mathbb{E}(x_{i,s+1}^{2})-\mathbb{E}(x_{i,s}^{2})\bigr]$, and thus ((ref)) holds if $x_{i,t}$ is an exact unit root with $\rho^{*}=1$, or $x_{i,s}$ is stationary in mean square such that $\mathbb{E}(x_{i,s+1}^{2})=\mathbb{E}(x_{i,s}^{2})$. However, this condition fails if $\rho^{*}$ varies with $T$.

The efficacy of the XJ estimator is manifested by its rate of convergence given by the following proposition.

propUnder \assumpref[s]{initval} and (ref), as $(n,T)\to\infty$ we have \[ \hat{\rho}^{\mkern2mu{\mathrm{XJ}}}-\rho^{*}=O_{p}\biggl(\frac{1}{\sqrt{nT^{1+\gamma}}}+\frac{1}{T^{2}}\biggr). \]

(ref) immediately implies that in the leading asymptotic case where $n/T\to r\in(0,\infty)$, XJ is a desirable estimator for IVX's bias correction since

equation[equation omitted — 268 chars of source]

satisfies the requirement ((ref)).

remUnder the extra assumption $\alpha_{i}=0$, the pooled OLS estimator \[ \hat{\rho}^{\mkern2mu\mathrm{LS}}=\frac{\sum_{i=1}^{n}\sum_{t=1}^{T-1}x_{i,t}x_{i,t+1}}{\sum_{i=1}^{n}\sum_{t=1}^{T-1}x_{i,t}^{2}} \] is $\sqrt{nT^{1+\gamma}}$-consistent. In reality when $\alpha_{i}\neq0$ this rate of convergence cannot be improved. (ref) indicates $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}-\rho^{*}=O_{p}\bigl(1/\sqrt{nT^{1+\gamma}}\bigr)$ in the leading asymptotic case. That is, the XJ estimator achieves the optimal rate of convergence.

To see the intuition why XJ works, we assume $x_{i,0}=\alpha_{i}=0$ to simplify the heuristic discussion, by which we can ignore $\eta_{i,T}$ in ((ref)). Following the decomposition of “$\hat{\rho}_{\mathit{fa}}$” in han2011uniform, which is a time series version of XD, we can deduce that when $\rho^{*}$ is close to one,

align*[align* omitted — 213 chars of source]

where $Q_{xx}$ is a deterministic positive number. The dominating term in the expectation of the terms in parentheses equals

equation[equation omitted — 152 chars of source]

The above expression hints a clue: the bias can be removed if we could somehow replace $\rho^{*(t-s)}$ in the first term with $\rho^{*2(t-s)}$, then ((ref)) becomes 0. Such a replacement is achieved by creative engineering. Observe that the indices in $\mathcal{E}$ are those of $\{1,2,\dots,(T-2)/2\}$ multiplied by 2. Therefore, as a prototype in the numerator of ((ref)) we can use

equation[equation omitted — 158 chars of source]

The first term in ((ref)) corresponds to \[ -\frac{\rho^{*}}{(T-2)/2}\sum_{t\in\mathcal{E}}\sum_{s\in\mathcal{E},s<t}\rho^{*(t-s)}=-\frac{\rho^{*}}{(T-2)/2}\sum_{t=1}^{(T-2)/2}\sum_{s<t}\rho^{*2(t-s)}, \] where the right-hand side is exactly compensated by the expectation of the second term in ((ref)) using observations with $t\in\{1,2,\dots,(T-2)/2\}$.

Furthermore, in principle the samples indexed by $\mathcal{E}$ and $\mathcal{O}$ are asymptotically equivalent. Hence, to reduce variance, we exploit both the odd and even indices, which explains the first three terms in the numerator of ((ref)). Finally, as for $\eta_{i,T}$, it serves as an adjustment for attenuating the impact of the initial values (if nonzero) caused by the first three terms. The above discussion makes it clear how the XJ estimator is constructed and why it attains a sharp convergence rate.

IVXJ

We first state a result of panel IVX under infeasible $\boldsymbol{\Omega}^{*}$ and $\rho^{*}$.

propUnder \assumpref[s]{initval} and (ref), as $(n,T)\to\infty$ we have \begin{gather*} \sqrt{nT^{1+(\theta\wedge\gamma)}}\bigl[\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}-\beta^{*}+\omega_{12}^{*}b_{n,T}^{{\mathrm{IVX}}}(\rho^{*})\bigr]\to_d\mathcal{N}\bigl(0,\Sigma^{{\mathrm{IVX}}}\bigr),\\ b_{n,T}^{{\mathrm{IVX}}}(\rho^{*})=O_{p}\left(1/T^{2-(\theta\vee\gamma)}\right), \end{gather*} where $b_{n,T}^{{\mathrm{IVX}}}(\rho)$ is defined in ((ref)) and $\Sigma^{{\mathrm{IVX}}}$ in ((ref)). Moreover, \[ \frac{\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}-\beta^{*}+\omega_{12}^{*}b_{n,T}^{{\mathrm{IVX}}}(\rho^{*})}{\haat{\varsigma}^{\mathrm{IVX}}(\omega_{11}^{*},\rho^{*})}\to_d\mathcal{N}(0,1), \] where $\haat{\varsigma}^{\mathrm{IVX}}(\omega_{11},\rho)$ is defined in ((ref)).

In reality, we must estimate $\rho^{*}$ and $\boldsymbol{\Omega}^{*}$ for feasible bias correction. We have highlighted the importance of an accurate estimator of $\rho^{*}$ with sufficiently small bias in Section (ref), and established the convergence rate of the XJ estimator $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$ in Section (ref). Although the estimation of $\boldsymbol{\Omega}^{*}$ also incurs errors, it does not undermine the key message: the XJ estimator $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$ enables valid bias correction in the leading asymptotic case where $n/T\to r\in\left(0,\infty\right)$.

The IVXJ estimator is

align*[align* omitted — 191 chars of source]

where $\hat{\omega}_{12}^{{\mathrm{IVX}}}:=\hat{\omega}_{12}(\hat{\beta}^{\mkern2mu{\mathrm{IVX}}})$ with $\hat{\omega}_{12}(\beta)$ given by ((ref)). Its $t$-statistic $t^{{\rm IVXJ}}$ is constructed by \[ t^{\mathrm{IVXJ}}:=\frac{\hat{\beta}^{\mkern2mu{\mathrm{IVXJ}}}-\beta^{*}}{\haat{\varsigma}^{\mathrm{IVX}}(\hat{\omega}_{11}^{{\mathrm{IVX}}},\hat{\rho}^{\mkern2mu{\mathrm{XJ}}})}, \] where $\hat{\omega}_{11}^{{\mathrm{IVX}}}:=\hat{\omega}_{11}(\hat{\beta}^{\mkern2mu{\mathrm{IVX}}})$ with $\hat{\omega}_{11}(\beta)$ given by ((ref)) and $\haat{\varsigma}^{\mathrm{IVX}}(\omega_{11},\rho)$ given by ((ref)). The following theorem is the key theoretical result for IVXJ.

thmUnder \assumpref[s]{initval} and (ref), if $(n,T)\to\infty$ and $n/T^{7-(\theta\vee\gamma)-\theta-3\gamma}\to0$, then \[ t^{{\rm IVXJ}}\to_d\mathcal{N}(0,1). \]
remConsider the leading asymptotic case. When $\gamma=1$, the Nickell-Stambaugh bias of $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$ is severe since $\sqrt{nT^{1+(\theta\wedge\gamma)}}b_{n,T}^{{\mathrm{IVX}}}(\rho^{*})\to\infty$, whereas in Theorem (ref) the rate condition $n/T^{7-(\theta\vee\gamma)-\theta-3\gamma}=n/T^{3-\theta}\to0$ ensures that IVXJ achieves standard and valid inference in the leading asymptotic case.

(ref) is a unified inference result under the polynomial rate $\rho^{*}=1+c^{*}/T^{\gamma}$. In fact, we can further enhance it to achieve asymptotic normality uniformly in $\rho^{*}$ using the drifting parameter sequence approach in andrews2020generic.

corFix three absolute constants $m_{1}^{*}\in(0,1)$, $m_{2}^{*}\in(0,\infty)$, and $\alpha\in[0,1]$. Under \assumpref[s]{initval} and (ref), as $(n,T)\to\infty$ and $n/T^{3-\theta}\to0$ we have \[ \sup_{-1+m_{1}^{*}\leq\rho^{*}\leq1+m_{2}^{*}/T}{}\left|\Pr\left\{ t^{{\rm IVXJ}}<\Phi^{-1}\left(\alpha\right)\right\} -\alpha\right|\to0 \] where $\Phi\left(\cdot\right)$ is the cumulative distribution function of $\mathcal{N}(0,1)$.
remThe interval for the admissible $\rho^{*}$ is a sequence of closed sets. The left-end is invariant and bounded away from $-1$, whereas the right-end exceeds but converges to 1 as in the LE case. Inside such sequence of closed sets $\rho^{*}$ can be an arbitrary sequence. This uniform result is more general and flexible than the convergent sequences specified in ((ref)).

We have shown that, in the simple predictive regression system ((ref)) and ((ref)) with one scalar regression $x_{i,t}$, IVXJ achieves unified and uniform inference for the parameter $\beta^{*}$ of interest. Admittedly, this model is a simplistic one in order to illustrate the ideas. Extending the feasible estimation and inference into multivariate predictive regressions is important, as in practice researchers mostly would like to include other control variables. This extension is discussed in the next subsection.

Mixed Roots and Long-Horizon Prediction

Suppose that the target variable of interest $y_{i,t+1}$ is linked with $k$ regressors $\boldsymbol{x}_{i,t}=(x_{j,i,t})_{j=1}^{k}$ in the linear form

equation[equation omitted — 143 chars of source]

The regressors are generated by a vector state space model:

align[align omitted — 170 chars of source]

where \[ \bm{R}^{*}=\bm{R}_{T}^{*}=\mathrm{diag}\bigl(\{\rho_{j}^{*}\}_{j=1}^{k}\bigr)\quad\text{with}\quad\rho_{j}^{*}=1+c_{j}^{*}/T^{\gamma_{j}},\quad c_{j}^{*}\in\mathbb{R}\text{ and }\gamma_{j}\in[0,1]. \] We suppress the subscript $T$ in $\bm{R}_{T}^{*}$ when there is no ambiguity. We allow the degree of persistence measured by $\gamma_{j}$ to be heterogeneous across regressors. Let $\boldsymbol{w}_{i,t}=(e_{i,t},\bm{v}_{i,t}^{\prime})^{\prime}$, which is assumed to be a homoskedastic m.d.s.\ with conditional variance\footnote{See (ref) in Online Appendix (ref) for the formal assumption.} \[ \mathbb{E}_{t-1}[\bm{w}_{i,t}\bm{w}_{i,t}']=

bmatrix[bmatrix omitted — 104 chars of source]

=:\bm{\Omega}^{*}. \]

Equations ((ref)) and ((ref)) make a generative model for the impulse response function, which is a key object of interest in empirical macroeconomic analysis. Let $H$ be the maximum prediction horizon of interest specified by the user. The generative model implies that the slope coefficient $\bm{\beta}^{(h)*}$ in the $h$-period-ahead predictive regression

equation[equation omitted — 164 chars of source]

happens to be the impulse response. Note that the effective number of time periods is $T_{h}:=T-h$. Indeed, ((ref)) is a panel version of jorda2005estimation's local projection. This paper is the first to consider panel local projection with persistent regressors, complementing mei2023implicit who cover panel with a stationary time dimension.

To implement IVX, we construct the self-generated IV as \[ \bm{z}_{i,t}:=\sum_{s=1}^{t}\rho_{z}^{t-s}\Delta\bm{x}_{i,s},\quad\text{for }t=1,\dots,T_{h}, \] where $\rho_{z}=1+c_{z}/T^{\theta}$, and $\Delta\bm{x}_{i,s}=\bm{x}_{i,s}-\bm{x}_{i,s-1}$. The IVX estimator for the $h$-period-ahead predictive regression is

equation[equation omitted — 262 chars of source]

After bias correction, we have the (feasible) IVXJ estimator \[ \hat{\bm{\beta}}^{(h){\rm IVXJ}}=\hat{\bm{\beta}}^{{(h)}{\mathrm{IVX}}}+\bm{b}_{n,T}^{{(h)}{\mathrm{IVX}}}(\hat{\bm{R}}^{{\rm XJ}},\hat{\boldsymbol{\omega}}_{12}^{{\rm IVXJ}},\hat{\bm{\Omega}}_{22}^{{\rm XJ}},\hat{\bm{\beta}}^{\mkern1.5mu{\mathrm{IVX}}}), \] where the bias function $\bm{b}_{n,T}^{{(h)}{\mathrm{IVX}}}$ is defined in ((ref)), and $\hat{\boldsymbol{\omega}}_{12}^{{\rm IVXJ}}$ and $\hat{\bm{\Omega}}_{22}^{{\rm XJ}}$ are given by ((ref)) and ((ref)), respectively.

Suppose we are interested in testing a linear joint null hypothesis $\mathbb{H}_{0}\colon\bm{A}\bm{\beta}^{(h)*}=\bm{q}$, where $\boldsymbol{A}$ is an $m\times k$ constant matrix of full row rank accommodating $m$ linear restrictions, and $\boldsymbol{q}$ is an $m\times1$ constant vector. We reject $\mathbb{H}_{0}$ under the significance level $\alpha$ if the Wald statistic

equation[equation omitted — 244 chars of source]

is greater than the $(1-\alpha)$-th quantile of a $\chi^{2}$ distribution of degree of freedom $m$, where $\hat{\bm{\Theta}}^{(h){\rm IVXJ}}$ is defined in ((ref)). In Section (ref) of the Online Appendices, we provide the formulae of the covariance estimators and the asymptotic theory, and supportive Monte Carlo evidence of the IVXJ estimator for multivariate and long-horizon regressions that allow for heterogeneous degrees of persistence across different regressors.

The Inconvenience of WG Estimation

This subsection makes it clear that if either the main regression ((ref)) or the AR regression ((ref)) is estimated by WG, then the unified inference is not achievable unless $n$ is much smaller than $T$. Since the theory here does not cover the leading asymptotic case and the simulation in Section (ref) illustrates the unsatisfactory performance of WG in finite samples, we do not recommend using WG for panel predictive regressions.

Throughout this subsection we focus on the simple regression with one regressor and a known $\boldsymbol{\Omega}^{*}$ for simplicity. If there is no asymptotic guarantee for a procedure when $\boldsymbol{\Omega}^{*}$ is known, it would be even worse when $\boldsymbol{\Omega}^{*}$ is unknown.

WG in AR Regression

We first look at the scenario when WG is used in the panel AR regression ((ref)). The WG estimator of $\rho^{*}$ is

equation[equation omitted — 185 chars of source]

With a known $\bm{\Omega}^{*}$, the IVX-WG estimator by plugging $\hat{\rho}^{\mkern2mu{\mathrm{WG}}}$ into the bias function is \[ \hat{\beta}^{\,\textrm{IVX-WG}}:=\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}+\omega_{12}^{*}\cdot b_{n,T}^{{\mathrm{IVX}}}\bigl(\hat{\rho}^{\mkern2mu{\mathrm{WG}}}\bigr). \] We immediately obtain the following result, as a corollary of Proposition (ref).

corUnder \assumpref[s]{initval} and (ref), as $(n,T)\to\infty$ we have \begin{equation} \hat{\rho}^{\mkern2mu{\mathrm{WG}}}-\rho^{*}=O_{p}\biggl(\frac{1}{\sqrt{nT^{1+\gamma}}}+\frac{1}{T}\biggr). \end{equation} If, in addition, $n/T^{5-(\theta\vee\gamma)-\theta-3\gamma}\to0$, then \[ \frac{\hat{\beta}^{\,\textup{IVX-WG}}-\beta^{*}}{\haat{\varsigma}^{\mathrm{IVX}}(\omega_{11}^{*},\hat{\rho}^{\mkern2mu{\mathrm{WG}}})}\to_d\mathcal{N}(0,1), \] where $\haat{\varsigma}^{\mathrm{IVX}}(\omega_{11},\rho)$ is defined in ((ref)).
remThe $1/T$ term in the convergence rate of $\hat{\rho}^{\mkern2mu{\mathrm{WG}}}$ arises from the Nickell-Stambaugh bias in panel AR, and is slower than that of $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$ in (ref). It leads to a narrower range of $n$ and $T$ for asymptotic normality. In particular, in the leading asymptotic case, when $\gamma=1$ the expansion rate condition for $\hat{\beta}^{\,\textup{IVX-WG}}$ is violated, since $n/T^{5-(\theta\vee\gamma)-\theta-3\gamma}=n/T^{1-\theta}$ diverges to infinity. This explains the unsatisfactory coverage probability of IVX-WG as we have observed in Figure (ref).

WG in Main Regression

We impose the following simplifying assumption in place of (ref) for simplicity.

assumptionprimep[Innovations]{\ref*{assump:innov}} The innovations $\{\bm{w}_{i,t}\}$ are i.i.d.\ across both $i$ and $t$, and have finite fourth moment, i.e., $\mathbb{E}\|\bm{w}_{i,t}\|^{4}<C_{w}<\infty$ for some positive constant $C_{w}$.

This assumption streamlines the construction of the standard error and guarantees a sufficient condition for joint central limit theorem when $\gamma=1$. If an estimator does not possess desirable properties under the simplifying (ref), it is expected to perform even worse under (ref).

The WG estimator $\hat{\beta}^{\mkern2mu{\mathrm{WG}}}$ in ((ref)) is the most widely used method of estimating the panel data model ((ref)). As explained in Section (ref), in panel predictive regressions the Nickell-Stambaugh bias is severe.

propUnder \assumpref[s]{initval} and (ref), as $(n,T)\to\infty$ we have \begin{gather*} \sqrt{nT^{1+\gamma}}\bigl[\hat{\beta}^{\mkern2mu{\mathrm{WG}}}-\beta^{*}+\omega_{12}^{*}\cdot b_{n,T}^{{\mathrm{WG}}}(\rho^{*})\bigr]\to_d\mathcal{N}\bigl(0,\Sigma^{{\mathrm{WG}}}\bigr),\\ b_{n,T}^{{\mathrm{WG}}}(\rho^{*})=O_{p}\left(1/T\right), \end{gather*} where $\Sigma^{{\mathrm{WG}}}$ is a positive constant. Furthermore, \[ \left[\hat{\beta}^{\mkern2mu{\mathrm{WG}}}-\beta^{*}+\omega_{12}^{*}\cdot b_{n,T}^{{\mathrm{WG}}}(\rho^{*})\right]\Big/\varsigma^{\mathrm{WG}}\to_d\mathcal{N}(0,1) \] as $(n,T)\to\infty$, where $\varsigma^{\mathrm{WG}}$ is defined in ((ref)).

(ref) characterizes the stochastic order of the bias, which is proportional to $b_{n,T}^{{\mathrm{WG}}}(\rho^{*})$. From this proposition we have $\hat{\beta}^{\mkern2mu{\mathrm{WG}}}-\beta^{*}=O_{p}\bigl(1/\sqrt{nT^{1+\gamma}}+1/T\bigr)$, which means that the WG estimator is consistent when both $n$ and $T$ pass to infinity. The main focus of predictive regressions lies in the inference to determine whether the variable $x_{i,t}$ retains predictive power to the targeted dependent variable. Mere consistency is insufficient for this purpose. The bias $\sqrt{nT^{1+\gamma}}b_{n,T}^{{\mathrm{WG}}}(\rho^{*})=O_{p}\bigl(\sqrt{n/T^{1-\gamma}}\bigr)$ may diverge to infinity and dominate the variance of $\hat{\beta}^{\mkern2mu{\mathrm{WG}}}$ when $\gamma=1$.

Eliminating the bias in WG is a challenging task, in particular when the regressor persistence is high. We try $\hat{\rho}^{\mkern2mu{\mathrm{WG}}}$ and $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$, and construct two respective estimators \[ \hat{\beta}^{\,\textup{WG-WG}}:=\hat{\beta}^{\mkern2mu{\mathrm{WG}}}+\omega_{12}^{*}\cdot b_{n,T}^{{\mathrm{WG}}}(\hat{\rho}^{\mkern2mu{\mathrm{WG}}})\quad\text{and}\quad\hat{\beta}^{\,\textup{WG-XJ}}:=\hat{\beta}^{\mkern2mu{\mathrm{WG}}}+\omega_{12}^{*}\cdot b_{n,T}^{{\mathrm{WG}}}(\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}). \] The following proposition summarizes the asymptotics of the two estimators.

propSuppose \assumpref[s]{initval} and (ref) hold, and $(n,T)\to\infty$. \begin{enumerate} • If $n/T^{3(1-\gamma)}\to0$, then $(\hat{\beta}^{\,\textup{WG-WG}}-\beta^{*})/\varsigma^{{\rm WG}}\to_d\mathcal{N}(0,1).$ • If $1/T^{1-\gamma}\to0$ and $n/T^{5-3\gamma}\to0$, then $(\hat{\beta}^{\,\textup{WG-XJ}}-\beta^{*})/\varsigma^{{\rm WG}}\to_d\mathcal{N}(0,1).$ \end{enumerate}

We have stated in ((ref)) that $\hat{\rho}^{\mkern2mu{\mathrm{WG}}}$'s rate of convergence is $O_{p}\bigl(1/\sqrt{nT^{1+\gamma}}+1/T\bigr)$, which reflects the Nickell-Stambaugh bias in the panel AR. The asymptotic bias vanishes when $n/T^{3(1-\gamma)}\to0$ for the $t$-statistics based on WG-WG, but this is not helpful for unified inference, as it obviously rules out the LUR cases with the persistence index $\gamma=1$.

WG-XJ tightens the valid asymptotic regime of the $t$-statistic from $n/T^{3(1-\gamma)}\to0$ in (ref)(i) to $(1/T^{1-\gamma})\wedge(n/T^{5-3\gamma})\to0$ in (ref)(ii) based on $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$. This is a substantial enhancement. However, it still rules out $\gamma=1$ in the leading asymptotic case. It is therefore also undesirable for statistical inference.

Summary of the Theory

This paper provides, for the first time, valid inference for panel predictive regressions using IVXJ. It highlights an undetected important issue in panel predictive regressions: the impact of the Nickell-Stambaugh bias on hypothesis testing is far more severe than we might have thought. Therefore, inferential procedures using popular estimators like WG and IVX are invalid for persistent regressors, and thus fail to meet the demand in empirical studies of macro-finance.

Specifically, in panel predictive regressions the bias for the asymptotic normality of $\hat{\beta}$ is inflated by both the standard panel factor $\sqrt{nT}$ and the persistence factor $\sqrt{T^{\gamma}}$. If the main regression is estimated by the WG estimator, only under very limited asymptotic regime that rules out $\gamma=1$ can we establish asymptotic normality, because the order of the Nickell-Stambaugh bias of $\hat{\beta}^{\mkern2mu{\mathrm{WG}}}$ is too large to be correctable. In contrast, the IVX estimator $\hat{\beta}^{\mkern2mu{\mathrm{IVX}}}$ in the main regression enables a niche for the construction of a feasible estimator in the panel AR regression, which effectively removes the bias and restores the validity of statistical inference based on the $t$-statistic. The contrast highlights the importance of identifying the sources of the bias --- the fusion of the Nickell bias from the panel data, and the Stambaugh bias from persistence regressors. The sources shed lights on devising the new solution $\hat{\rho}^{\mkern2mu{\mathrm{XJ}}}$ to cope with persistence.

As a final remark, this paper serves as a stepping stone toward a comprehensive theory for predictive regressions in practical and realistic empirical settings. The theoretical results here rely on the correctly specified AR(1) model ((ref)) for $x_{i,t}$. In view of the lengthy proofs in the current paper, general specifications like the AR of higher orders and heterogeneous AR coefficients across individuals are not covered. Furthermore, properties of the IVXJ estimator under model misspecification will be an important topic for future exploration.

Empirical Application

Understanding economic and financial crises is the holy grail of macroeconomic research. Since the 2008 Great Financial Recession, a plethora of new research papers look for precursors of crises from historical variables, among which are romer2017new, mian2018finance, and baron2021banking. The recent literature is surveyed by sufi2021financial.

table[table omitted — 6,330 chars of source]

One of the latest contributions by greenwood2022predictable uses a single predictor in the panel data predictive regression. Their baseline sample is an unbalanced panel that covers 42 countries from 1950 to 2016 annually. Their investigation is conducted via the panel local projection form

equation[equation omitted — 129 chars of source]

The dependent variable $\mathit{y_{i,t+h}}$ is a binary indicator of financial crisis occurring in country $i$ in any year between year $t+1$ and year $t+h$, and the regressor $x_{i,t}^{\mathrm{diff}}$ is the three-year debt growth relative to either GDP or CPI. Four measures of normalized debt are used: (i) ratio of private debt to GDP ($\mathit{Debt}^{\mathit{Priv}}/\mathit{GDP}$); (ii) business debt to GDP ($\mathit{Debt}^{\mathit{Bus}}/\mathit{GDP}$); (iii) household debt to GDP ($\mathit{Debt}^{\mathit{HH}}/\mathit{GDP}$); and (iv) the logarithm of real private debt ($\log(\mathit{Debt}^{\mathit{Priv}}/\mathit{CPI})$).

For IVX we set $c_{z}=-1$ and $\theta=0.95$, the same as what we did in Section (ref)'s simulations, thereby $\rho_{z}=1-1/\text{\ensuremath{\max_{i\leq n}}}T_{i}^{0.95}$, where $T_{i}$ is the time span of country $i$ in this unbalanced panel. Panel A of Table (ref) reports the IVX estimates for $\beta^{{(h)}*}$, which are similar in magnitude to those of greenwood2022predictable. In column (1), for example, the bias corrected coefficient estimate 2.38 means that a one-standard-deviation increase this year in debt growth measured by change in private debt to GDP is associated with a 2.38% increase in the likelihood of crisis occurring in the next year. The message in greenwood2022predictable is supported by the results: credit growth has predictive power in forecasting the onset of a financial crisis.

Our theory and method allow for highly persistent regressors to leverage faster convergence speed. Therefore we are interested in the estimates if we directly use the level of the debt, instead of the growth rate of the debt, to predict financial crises. Panel B of Tables (ref) presents the results of the regression

equation[equation omitted — 118 chars of source]

As we use different regressors, the coefficients in Panel A and B have different interpretation. For example, in column (9) of Panel B of Table (ref), the bias corrected estimate 11.37 means that a one-standard-deviation increase in this year's debt level measured by private debt to GDP is associated with a 11.37% increase in the probability of crisis occurring within the next three years. This effect is much more considerable than the one estimated from debt growth regression, which is only 5.18% (column (9) of Panel A). Furthermore, the debiased estimates in Panel B are generally larger than those without bias correction (except columns (3) and (7), the coefficients on household debt to GDP), and the bias correction is essential to provide asymptotically unbiased estimation for valid inference.

Panels A and B use the differenced variable $x_{i,t}^{\mathrm{diff}}$ and the level variable $x_{i,t}$, respectively. Which one is preferred? In our opinion, the level variable provides a more direct measurement.

table[table omitted — 632 chars of source]
exampleThe differenced specification in ((ref)) literally implies the chance of crises is related to the change of the debt-to-GDP ratio but the level is irrelevant. Consider an imagery country that takes a path of the ratios as shown in Table (ref). It starts in year 0 at 180% (close to Greek's level, highest in the European Union) and in Year 3 this country's policymakers carry out a debt-reduction policy and successfully lower the ratio to 60% (close to Germany's level). The differenced specification, in the first row, implies that although in Year 3 it enjoyed a reduction (it is expected that $\beta^{*}\geq0$), in the other years the probability of crisis remains the same. The level specification in the second row, in contrast, implies a high probability of crisis when the ratio is high, and a low probability when the ratio is low. The latter, in our view, is much more intuitive and natural, as the level is directly associated with the monetary interest that a country has to pay.

In academic research that uses descriptive statistics instead of regression techniques, the level debt-to-GDP ratio is mentioned much more frequently. For instance, in their influential book reinhart2009time collected historical data to analyze a wide range of crises, and concluded that crises often share common triggers, such as excessive debt accumulation, whether by the government, corporations, or consumers. reinhart2010growth further wrote: “When gross external debt reaches 60 percent of GDP, annual growth declines by about two percent; for levels of external debt in excess of 90 percent of GDP, growth rates are roughly cut in half.” This quotation links the debt-to-GDP ratio with economic growth via two thresholds in the level.

In policy debate and practice, austerity refers to economic policies implemented by governments to reduce the size of budget deficits (relative to GDP) through fiscal consolidation. A high debt-to-GDP ratio suggests that a country might have difficulties to repay external debt and may be at risk of default. For example, during the Greek debt crisis in late 2009, austerity played a central role as a condition for receiving bailout funds. The short-term bailout and the long-term economic reforms were both aimed at reducing the overall debt burden.

Returning to our empirical exercise, the numbers in Panel B of Table (ref) across different measurements of the debt are much closer than their counterparts in Panel A. It provides a unified message: the debts --- private, business, and household --- deliver consistent impact in triggering financial crises. Such homogeneity is not observed from the numbers in Panel A.

Conclusion

This paper investigates the problem of panel predictive regressions, with focus on valid inference based on the $t$-statistic. When $n$ and $T$ are both large, WG and IVX incur the Nickell-Stambaugh bias which distorts the size of the standard inferential procedure. We propose to use IVX to estimate $\beta^{*}$ in the panel predictive regression, and then plug in the novel XJ estimator for $\rho^{*}$ to correct the bias. We show that this procedure provides unified inference in various modes of dynamic regressors, including the stationary case, the mildly integrated case, and the local unit root case. The unified inference cannot be achieved if either the main regression or the AR regression is estimated by WG.

IVXJ appears to be the first and only available estimator that eliminates the asymptotic bias in panel predictive regressions. To automate the IVXJ procedure, an open-source Python package is made available at \url{https://github.com/metricshilab/ivxj}.

center[center omitted — 55 chars of source]