EconBase
← Back to paper

Inference in dynamic models for panel data using the moving block bootstrap

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

22,724 characters · 4 sections · 22 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

INFERENCE IN DYNAMIC MODELS FOR PANEL DATA USING THE MOVING BLOCK BOOTSTRAP

\def\spacingset#1 \spacingset{1}

\if11 \fi

\if01 {

center[center omitted — 43 chars of source]

} \fi

abstractInference in linear panel data models is complicated by the presence of fixed effects when (some of) the regressors are not strictly exogenous. Under asymptotics where the number of cross-sectional observations and time periods grow at the same rate, the within-group estimator is consistent but its limit distribution features a bias term. In this paper we show that a panel version of the moving block bootstrap, where blocks of adjacent cross-sections are resampled with replacement, replicates the limit distribution of the within-group estimator. Confidence ellipsoids and hypothesis tests based on the reverse-percentile bootstrap are thus asymptotically valid without the need to take the presence of bias into account.

{\bf JEL Classification:} C23

{\bf Keywords:} asymptotic bias, bootstrap, dynamic model, fixed effects, inference

\spacingset{1.45}

\setcounter{equation}{0}

Introduction

Econometric models for panel data almost invariably feature fixed effects. Their presence complicates estimation and inference as they introduce bias in the fixed-effect estimator of the parameter of interest, in general. A leading case is the estimation of the regression slopes in a linear model when the regressors are not strictly exogenous. The canonical example is an autoregressive model, where the bias of the within-group estimator was first derived in influential work by Nickell1981, but correlation between contemporaneous errors and both past and future regressors is the norm rather than the exception. The resulting bias remains important in large samples unless the number of time-series observations, $m$, is large relative to the cross-sectional sample size, $n$. This is, however, not a situation regularly encountered in practice.

Under asymptotics where $n$ and $m$ grow at the same rate the fixed-effect estimator is generally consistent but suffers from asymptotic bias. Several ways to estimate this bias to restore the validity of inference based on the limit distribution have been proposed; see ArellanoHahn2007 and DhaeneJochmans2015. An alternative recently explored by GoncalvesKaffo2015 and HigginsJochmans2024 is to use the bootstrap to replicate the distribution of the fixed-effect estimator, including its bias. This allows to perform inference based on the bootstrap distribution in the usual manner, i.e., no adjustment for the presence of bias needs to be made. The difficulty here lies in devising a bootstrap scheme that correctly reproduces the bias. HigginsJochmans2024 show that the parametric bootstrap does so in quite general nonlinear models. GoncalvesKaffo2015 focus on the linear autoregressive model and show that the wild bootstrap permits inference using the within-group estimator. They also demonstrate the failure of several other conventional bootstrap schemes to replicate bias, illustrating the difficulty in devising a successful procedure. An important limitation of the methods available to date is that they cannot handle unspecified feedback processes, leaving open the question of how to do so and, indeed, whether or not this is possible.

In this paper we study the linear model where the errors are allowed to be correlated with both past and future regressors. This is the workhorse model in applied economics. We show that a version of the moving block bootstrap of Kunsch1989, where blocks of adjacent cross-sections are resampled with replacement, correctly replicates the distribution of the within-group estimator under asymptotics where $\nicefrac{n}{m} \rightarrow c \in [0,+\infty)$. This bootstrap scheme was not considered in GoncalvesKaffo2015. Goncalves2011 did show that this procedure is capable of yielding correct inference in our model even under the additional complication of cross-sectional dependence. However, the assumptions under which she established this result include a condition on $n$ and $m$ that renders the bias of the estimator small relative to its standard deviation; in our context, this amounts to the rate condition $\nicefrac{n}{m}\rightarrow 0$. Here, instead, we focus on the ability of the moving block bootstrap to replicate the bias when $m$ need not be small relative to $n$ and we abstract from cross-sectional dependence. Nevertheless, as the expression of the bias would remain unchanged, our main findings should generalize to such a situation.

Our contribution should be interpreted against the backdrop of a recent interest in the validity of the bootstrap in the presence of (asymptotic) bias; see also CavaliereGoncalvesNielsenZanelli2024 in addition to the work already mentioned above. We conjecture that the validity of our bootstrap scheme extends to a large class of nonlinear problems. We have established this to be the case for the variance estimator in the classical many normal means problem of NeymanScott1948. A more general theory is left for future work.

The paper is structured as follows. In the next section we first formally state our model and assumptions, and derive the limit distribution of the within-group estimator. This yields a somewhat more general expression of its asymptotic bias than is available in the literature. The following section introduces the moving block bootstrap and states our main result---that is, the distribution of the bootstrap within-group estimator, centered around the within-group estimator and conditional on the data, is consistent for the limit distribution of the within-group estimator centered around the truth---together with its chief implications for estimation and inference. A final section reports on a numerical experiment that was conducted to support our claims. All proofs are collected in the Appendix.

Linear regression with fixed effects

We are interested in estimation of and inference on the slope vector $\beta$ in the linear model $$ y_{it} = \alpha_i + x_{it}^\prime \beta + \varepsilon_{it}, \qquad \mathbb{E}(\varepsilon_{it}) = 0, \qquad \mathbb{E}(x_{it} \varepsilon_{it}) = 0, $$ from $n\times m$ panel data, treating the $n$ intercepts $\alpha_1,\ldots, \alpha_n$ as unknown parameters to be estimated.

We will work under a set of three assumptions which we state next. These assumptions are standard in the literature. The first assumption contains moment requirements and mixing conditions.

assumption(i) The variables $x_{it}$ and $\varepsilon_{it}$ have uniformly bounded moments of order $2r$ for some $r>2$. (ii) The variables $x_{it} \varepsilon_{it}$ have uniformly bounded moments of order $3r$. (iii) The variables $x_{it}$ and $\varepsilon_{it}$ are independent across $i$. (iv) The processes $\lbrace (x_{it},\varepsilon_{it}) \rbrace$ are stationary mixing and their mixing coefficients, $a_i$, satisfy $$ \sup_{1\leq i \leq n} a_i(h) = O(h^{-s}) $$ \hphantom{(iii) }for some $s > \nicefrac{4r}{r-2}$.\footnote{We recall that $$ a_i(h) \coloneqq \sup_{1\leq t \leq m} \sup_{A\in\mathcal{A}_{it}} \sup_{B\in \mathcal{B}_{it+h}} \lvert \mathbb{P}(A\cap B) - \mathbb{P}(A) \, \mathbb{P}(B) \rvert, $$ for $\mathcal{A}_{it}$ and $\mathcal{B}_{it}$ the sigma algebras generated by the sequences $(x_{it},\varepsilon_{it}) ,(x_{it-1},\varepsilon_{it-1}),\ldots$ and $(x_{it},\varepsilon_{it}) ,(x_{it+1},\varepsilon_{it+1}),\ldots$, respectively.}

The second assumption states conventional rank conditions on certain covariance matrices. Here and in the sequel we let $z_{it} \coloneqq x_{it} - \mathbb{E}(x_{it})$.

assumption(i) The covariance matrix $ \varSigma_{n,m} \coloneqq \nicefrac{1}{nm} \sum_{i=1}^{n} \sum_{t=1}^{m} \mathbb{E} ( z_{it}^{\vphantom{\prime}} z_{it}^\prime ) $ is uniformly positive definite. (ii) The variance-covariance matrix of $ \nicefrac{1}{\sqrt{nm}} \sum_{i=1}^{n} \sum_{t=1}^{m} z_{it} \varepsilon_{it}, $ $$ \varOmega_{n,m} \coloneqq \nicefrac{1}{nm} \sum_{i=1}^{n} \sum_{t=1}^{m} \left( \mathbb{E}(z_{it}^{\vphantom{\prime}} z_{it}^\prime \varepsilon_{it}^{2}) + \sum_{\tau=1}^{m-1} \nicefrac{(m-\tau)}{m} \ \mathbb{E} ( (z_{it}^{\vphantom{\prime}} z_{it+\tau}^\prime + z_{it+\tau}^{\vphantom{\prime}} z_{it}^\prime) \, \varepsilon_{it}^{\vphantom{\prime}} \varepsilon_{it+\tau}^{\vphantom{\prime}} ) \right), $$ \hphantom{(ii) }is uniformly positive definite.

The third assumption states the asymptotic approximation under which we proceed.

assumption$n,m\rightarrow\infty$ with $\nicefrac{n}{m} \rightarrow c \in [0,+\infty)$.

Because we allow for $\mathbb{E}(x_{it} \varepsilon_{it+\tau})$ and $\mathbb{E}(x_{it+\tau} \varepsilon_{it})$ to be non-zero when $\tau\neq 0$, both the within-group least-squares estimator and generalized method-of-moment estimators as in ArellanoBond1991 will be inconsistent, in general, under asymptotics where $m$ is held fixed.\footnote{In the conventional first-order autoregressive model, for example, $x_{it}=y_{it-1}$ and $\varepsilon_{it}\sim\mathrm{i.i.d.}(0,\sigma^2)$, and so $\mathbb{E}(x_{it} \varepsilon_{it+\tau})=0$ but $\mathbb{E}(x_{it+\tau} \varepsilon_{it}) = \beta^{\tau-1}\sigma^2\neq 0$ for all $\tau>0$. This, then, leads to the well-known Nickell1981 bias in the within-group estimator. More generally, while moment-based strategies are available that can handle situations where $\mathbb{E}(x_{it+\tau} \varepsilon_{it}) \neq 0$, these approaches require that $\mathbb{E}(x_{it} \varepsilon_{it+\tau}) = 0$ at least for two known values of $\tau\leq m-t$ to yield valid moment conditions that can be exploited to construct an estimator; see Arellano2003a.} In fact, it would appear that $\beta$ is not point identified in such a framework without further conditions. Reversely, asymptotics where $n$ does not diverge are unsuitable for typical microeconometric applications.

The within-group estimator of $\beta$ is $$ \hat{\beta} \coloneqq \left(\sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it} - \bar{x}_i ) (x_{it} - \bar{x}_i )^\prime \right)^{-1} \left(\sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it} - \bar{x}_i ) (y_{it} - \bar{y}_i )^{\hphantom{\prime}} \right), $$ where $\bar{y}_i \coloneqq \nicefrac{1}{m} \sum_{t=1}^m y_{it}$ and $\bar{x}_i \coloneqq \nicefrac{1}{m} \sum_{t=1}^m x_{it}$. Under Assumptions (ref)-(ref) this estimator is consistent and asymptotically normally-distributed, but its limit distribution features an asymptotic bias term unless $\nicefrac{n}{m}\rightarrow 0$.

In the following theorem, we let $ \varUpsilon \coloneqq \mathrm{lim}_{n,m\rightarrow\infty} \varUpsilon_{n,m} $ for $ \varUpsilon_{n,m} \coloneqq \varSigma_{n,m}^{-1} \varOmega_{n,m}^{\vphantom{-1}} \varSigma_{n,m}^{-1}. $

theoremLet Assumptions (ref)-(ref) hold. Then $$ \sqrt{nm} \, (\hat{\beta} - \beta - \nicefrac{\varSigma_{n,m}^{-1} b_{n,m}^{\vphantom{-1}}}{m}) \overset{L}{\rightarrow} N(0,\varUpsilon), $$ where $ b_{n,m} \coloneqq - \nicefrac{1}{n} \sum_{i=1}^n \sum_{\tau=1}^{m-1} \nicefrac{(m-\tau)}{m} \left( \mathbb{E}(z_{it}\varepsilon_{it+\tau}) + \mathbb{E}(z_{it+\tau}\varepsilon_{it}) \right). $

Theorem (ref) generalizes results available for models with predetermined regressors generated through a specified process such as those derived in HahnKuersteiner2002 and ChudikPesaranYang2018.

Moving block bootstrap

Our bootstrap scheme consists of applying the moving block bootstrap of Kunsch1989 in the time-series dimension of the data, jointly for all cross-sectional units. Moreover, for integers $p$ and $q$ with $m = p\times q$, we randomly select $p$ blocks of $q$ consecutive cross-sections from the original data; the blocks may overlap. Our bootstrap sample is then obtained on concatenating these $p$ blocks. If we let $\varpi_1,\ldots \varpi_{p}$ be a random sample from the discrete uniform distribution on $\lbrace 0,\ldots, m-q \rbrace$, the bootstrap time series $\lbrace (y_{it}^*, x_{it}^*) \rbrace$ is generated as $$ y_{i \, (p^\prime-1)q+q^\prime}^* \coloneqq y_{i\, \varpi_{p^\prime}+q^\prime}, \qquad x_{i \, (p^\prime-1)q+q^\prime}^* \coloneqq x_{i\, \varpi_{p^\prime}+q^\prime}, $$ for $1\leq p^\prime \leq p$ and $1\leq q^\prime \leq q$. Here, the random variables $\varpi_1,\ldots, \varpi_p$ select starting points for the different blocks. When the block length, $q$, is set to one we recover the bootstrap as originally introduced by Efron1979 although, in our setting, we will require $q$ to grow with $m$.

The within-group estimator computed from the bootstrap sample so constructed equals $$ \hat{\beta}^* \coloneqq \left(\sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it}^* - \bar{x}_i^* ) (x_{it}^* - \bar{x}_i^* )^\prime \right)^{-1} \left(\sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it}^* - \bar{x}_i^* ) (y_{it}^* - \bar{y}_i^* )^{\hphantom{\prime}} \right), $$ where $\bar{y}_i^* \coloneqq \nicefrac{1}{m} \sum_{t=1}^m y_{it}^*$ and $\bar{x}_i^* \coloneqq \nicefrac{1}{m} \sum_{t=1}^m x_{it}^*$.

The following theorem is our main result. In it, as usual, we let $\mathbb{P}^*$ denote a probability computed with respect to the bootstrap measure, that is, conditional on the original data.

theoremLet Assumptions (ref)-(ref) hold and suppose that $q\rightarrow \infty$ with $q = o(\sqrt{m})$. Then $$ \mathbb{P} \left( \sup_a \left\lvert \mathbb{P}^*(\sqrt{nm}(\hat{\beta}^*-\hat{\beta})\leq a) - \mathbb{P}(\sqrt{nm}(\hat{\beta}-\beta)\leq a) \right\rvert > \epsilon \right) = o(1) $$ for any $\epsilon>0$.

This result generalizes the findings in Goncalves2011, where Assumption (ref) is replaced by the stronger requirement that $\nicefrac{n}{m}\rightarrow 0$. Because the bias in $\sqrt{nm}(\hat{\beta}-\beta)$ is of the order $\sqrt{\nicefrac{n}{m}}$, her rate condition ensures that the limit distribution of the within-group estimator does not feature a bias term.

Theorem (ref) has a number of useful implications. A first is that inference based on the bootstrap remains valid in the presence of asymptotic bias. To see this, say we wish to perform inference on linear contrasts of the form $\theta\coloneqq c^\prime \beta$, where $c$ is a chosen vector of conformable dimension. For $\alpha \in (0,1)$, let $$ \hat{Q}_\alpha \coloneqq \inf \lbrace Q: \alpha \leq \mathbb{P}^*(\hat{\theta}^*-\hat{\theta} \leq Q) \rbrace, $$ with $\hat{\theta}\coloneqq c^\prime \hat{\beta}$ and $\hat{\theta}^*\coloneqq c^\prime \hat{\beta}^*$. By \citet*[Lemma 23.3]{vanderVaart2000} we have that, under the conditions of Theorem (ref), $$ \lim_{n,m\rightarrow\infty} \mathbb{P} ( \hat{\theta} - \hat{Q}_\alpha \leq \theta ) \rightarrow \alpha , $$ which allows the construction of confidence intervals and decision rules to conduct inference on $\theta$.

Theorem (ref) also implies that the median of the bootstrap distribution converges to the median of the limit distribution which, by Theorem (ref), equals the asymptotic bias. Hence, $$ \check{\beta} \coloneqq \hat{\beta} - \mathrm{median}^*(\hat{\beta}^* - \hat{\beta}). $$ is a bootstrap-based bias-corrected estimator of $\beta$.

The variance of the limit distribution in Theorem (ref) can be estimated using conventional HAC methods or by means of resampling via the moving block bootstrap. To describe the latter way, consider $$ \hat{\varOmega}_{n,m}^* \coloneqq \nicefrac{1}{np} \sum_{i=1}^{n} \sum_{p^\prime =1 }^{p} V_{\varpi_{p^\prime}}^{\vphantom{\prime}} V_{\varpi_{p^\prime}}^\prime, \qquad V_{\varpi} \coloneqq \left( \nicefrac{1}{\sqrt{q}} \sum_{q^\prime=1}^{q} (x_{i\varpi+q^\prime} - \bar{x}_i^*) \, \hat{\epsilon}_{i\varpi+q^\prime}^* \right), $$ where $\hat{\epsilon}_{it}^* \coloneqq (y_{it}- \bar{y}_i^*) - (x_{it}- \bar{x}_i^*)\hat{\beta}^*$ are bootstrap residuals. This is an estimator of the (conditional) bootstrap variance $$ \hat{\varOmega}_{n,m} \coloneqq \mathrm{var}^* \left( \nicefrac{1}{\sqrt{nm}} \sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it}^* - \bar{x}^*_i) \, \hat{\varepsilon}_{it}^* \right), $$ where $\hat{\varepsilon}_{it}^* \coloneqq (y_{it}^*- \bar{y}_i^*) - (x_{it}^*- \bar{x}_i^*)\hat{\beta}$. The latter, in turn, is an estimator of $\varOmega_{n,m}$; it is well-known that it can be computed without resampling the data (see Kunsch1989). If we further introduce the shorthands $$ \hat{\varSigma}_{n,m}^* \coloneqq \nicefrac{1}{nm} \sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it}^*-\bar{x}_{i}^*) (x_{it}^*-\bar{x}_{i}^*)^\prime , \quad \hat{\varSigma}_{n,m} \coloneqq \nicefrac{1}{nm} \sum_{i=1}^{n} \sum_{t=1}^{m} (x_{it}-\bar{x}_{i}) (x_{it}-\bar{x}_{i})^\prime, $$ we can define $$ \hat{\varUpsilon}^* \coloneqq (\hat{\varSigma}_{n,m}^*)^{-1} \hat{\varOmega}_{n,m}^{*} (\hat{\varSigma}_{n,m}^*)^{-1}, \qquad \hat{\varUpsilon} \coloneqq \hat{\varSigma}_{n,m}^{-1} \hat{\varOmega}_{n,m}^{\vphantom{-1}} \hat{\varSigma}_{n,m}^{-1}. $$ By the same arguments as in Goncalves2011 we can verify that $ \hat{\varUpsilon}^* - \varUpsilon \overset{P^*}{\rightarrow} 0 $ and $ \hat{\varUpsilon} - \varUpsilon \overset{P}{\rightarrow} 0. $

Combined with our theorems, this yields two additional results. The first of these is that inference can be performed using the normal approximation to our bias-corrected estimator, using $\hat{\varUpsilon}$. The second is that the reverse-percentile bootstrap can be applied to studentized quantities to perform inference in the same way as before. In the context of linear contrasts of the form $\theta = c^\prime\beta$, for example, if we let ${\hat{\sigma}^*}^2 \coloneqq c^\prime \, \hat{\varUpsilon}^* c$ and $\hat{\sigma}^2 \coloneqq c^\prime \, \hat{\varUpsilon} c$, and redefine $$ \hat{Q}_\alpha \coloneqq \inf \lbrace Q: \alpha \leq \mathbb{P}^*(\nicefrac{(\hat{\theta}^*-\hat{\theta})}{\hat{\sigma}^*} \leq Q) \rbrace, $$ then $ \mathbb{P} ( \hat{\theta} - \hat{\sigma} \, \hat{Q}_\alpha \leq \theta ) \rightarrow \alpha $ under the assumptions of Theorem (ref). This implies that decision rules for null hypotheses on $\theta$ involving critical values obtained from the bootstrap distribution of $\nicefrac{(\hat{\theta}^*-\hat{\theta})}{{\hat{\sigma}^*}} $ deliver asymptotic size control without the need to correct for bias.

Simulations

table[table omitted — 887 chars of source]

To support Theorem (ref) we report some numerical results. Consider a stationary first-order autoregression with standard-normal innovations, that is $$ y_{it} = x_{it} \beta + \varepsilon_{it} $$ with $\varepsilon_{it}\sim\mathrm{i.i.d.}~N(0,1)$ for all $1\leq t \leq m$, $x_{i1} \sim \mathrm{i.i.d.}~N(0,\nicefrac{1}{(1-\beta^2)})$, and $x_{it}=y_{it-1}$ for $2\leq t \leq m$. The within-group estimator is invariant to the distribution of the fixed effects, so setting $\alpha_i = 0$ for all $1\leq i \leq n$ is without loss of generality. It is well-known that, here,

equation*[equation* omitted — 134 chars of source]

as $n,m\rightarrow\infty$. In Table (ref) we report the average (over 10,000 Monte Carlo replications) of the bootstrap distribution of $\sqrt{nm}(\hat{\beta}^* - \hat{\beta})$ (as computed using 1,999 bootstrap replications) evaluated at the percentiles of the normal distribution in the above limit statement (which is not centered at zero). This effectively yields a QQ-plot, although we provide it in tabular form. We do this for two different sample sizes (combinations of $n$ and $m$) and for various choices of the number and length of the bootstrap blocks ($p$ and $q$). The table concerns simulations under a data generating process where $\beta=0$. It shows that the quantiles of the bootstrap distribution are, on average, close to those of the limit distribution. The same phenomenon was equally observed in a wider set of simulation designs and also carries over to the bootstrap distribution of the studentized estimator using the variance estimators given above.