Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,162 characters · 12 sections · 0 citation commands
Pooled Bewley Estimator of Long Run Relationships in Dynamic Heterogenous Panels
\thispagestyle{empty}
\pagenumbering{arabic}
\doublespacing
Estimation of cointegrating relationships in panels with heterogeneous short-run dynamics is important for empirical research in open economy macroeconomics as well as in other fields in economics. Existing single-equation panel estimators in the literature are panel Fully Modified OLS (FMOLS) by \citeANP{Pedroni1996} (\citeyearNP{Pedroni1996}, \citeyearNP{Pedroni2001}, \citeyearNP{Pedroni2001ReStat}) , panel Dynamic OLS (PDOLS) by \citeN{MarkSul2003} , and the likelihood based Pooled Mean Group (PMG) estimator by \citeN{PesaranShinSmith1999} . Multi-equation (system) approach by \citeN{Breitung2005} , and the related system PMG approach by \citeN{ChudikPesaranSmith2022} are another contributions in the literature on estimating cointegrating vectors in a panel context. In this paper, we propose a pooled Bewley (PB) estimator of long-run relationships, relying on the Bewley transform of an autoregressive distributed lag (ARDL) model ( \citeANP{Bewley1979}, \citeyearNP{Bewley1979} ). See also \citeANP{WickensBreusch1988} (\citeyearNP{WickensBreusch1988}) for a discussion of the Bewley transform.
Our setting is the same as that of \citeN{PesaranShinSmith1999} . Under this setting, any short-run feedbacks between the outcome variable ($ y$) and regressors ($x$) are allowed, but the direction of the long run causality is assumed to go from $x$ to $y$. Hence, as with the PMG, the PB estimator allows for heterogeneity in short-run feedbacks, but restricts the direction of long run causality. The PB estimator is computed analytically, and does not rely on numerical maximization of the likelihood function that underlies the PMG estimation. We derive the asymptotic distribution of the PB estimator when the cross-section dimension ($n$) and the time dimension ($ T$) diverge to infinity jointly such that $n=\Theta \left( T^{\theta }\right) $, for $0<\theta <2$, where we use the notation $\Theta \left( .\right) $ to denote the same order of magnitude asymptotically, namely if $ \left\{ f_{s}\right\} _{s=1}^{\infty }$ and $\left\{ g_{s}\right\} _{s=1}^{\infty }$ are both positive sequences of real numbers, then $ f_{s}=\ominus \left( g_{s}\right) $ if there exists $S_{0}\geq 1$ and positive finite constants $C_{0}$ and $C_{1}$, such that $\inf_{s\geq S_{0}}\left( f_{s}/g_{s}\right) \geq C_{0},$ and $\sup_{s\geq S_{0}}\left( f_{s}/g_{s}\right) \leq C_{1}$. Our asymptotic analysis is an advance over the theoretical results currently available for PMG, PDOLS and FMOLS estimators where it is assumed that $n$ is small relative to $T$ (which corresponds to the case where $\theta $ is close to zero).
How well individual estimators work in samples of interest in practice where $n$ and/or $T$ are often less than 50 is a different matter, which we shed light on using Monte Carlo experiments. Monte Carlo evidence shows PB estimator can be superior to PMG, PDOLS and FMOLS, in terms of its overall precision as measured by the Root Mean Square Error (RMSE), and in terms of accuracy of inference as measured by size distortions. These experiments reveal PB is a useful addition to the literature.
Monte Carlo evidence also shows that the time dimension is very important for the performance of these estimators, and that all of the four estimators under consideration suffer from the same two drawbacks: small sample bias and size distortion. Although the size distortions are found to be less serious for the PB estimator in our experiments, all four estimators exhibit notable over-rejections in sample sizes relevant in practice. In addition, all four estimators (perhaps unsurprisingly) suffer from bias in finite samples, albeit a rather small one. Both drawbacks diminish as $T$ is increased relative to $n$.
To conduct reliable inference regardless of the cross-sectional dependence of errors, we make use of the sieve wild bootstrap procedure. To accommodate cross-sectional dependence, we resample the cross-section vectors of residuals, an idea that was originally proposed by \citeN{MaddalaWu1999} . We found the sieve wild bootstrap procedure to be remarkably effective for all four estimators, regardless of cross-sectional dependence of errors, and we therefore recommend using it in empirical research.
Regarding the small sample bias, we consider the application of two bias-correction methods taken from the literature, relying either on split-panel jackknife ( \citeANP{DhaeneJochmansy2015}, \citeyearNP{DhaeneJochmansy2015} ) or sieve wild bootstrap approaches. In contrast to split-panel approaches in panels without stochastic trends, as, for instance, considered by of \citeN{DhaeneJochmansy2015} or \citeN{ChudikPesaranYang2018} , in this paper we need to combine the full sample and half-panel subsamples using different weighting due to the fact that the rate of convergence of the estimators of long run coefficients is faster, at the rate of $T\sqrt{n}$ , as compared to the standard rate of $\sqrt{nT}$. We find that both of these approaches can be helpful in reducing the bias (for all four estimators). However, given that the bias is small to begin with, the value of bias correction methods is limited.
The relevance of choosing a particular estimation approach is illustrated in the context of a consumption function application for OECD economies taken from \citeN{PesaranShinSmith1999} . This application shows that quite a different conclusion would be reached when using PB estimator, which does not reject the zero long-run coefficient on inflation, in line with the long-run neutrality of monetary policy, whereas the PMG estimator results in a highly statistically significant negative long-run coefficient. Estimates of the long-run coefficient on real income are less diverse across estimators, but the inference on whether a unit long-run coefficient on real income (as suggested by balanced growth path models in the literature) can be rejected or not depends on the choice of a particular estimator.
The remainder of this paper is organized as follows. Section (ref) presents the model and assumptions, introduces the PB estimator, and provides asymptotic results. Application of bias correction methods and bootstrapping critical values are also discussed in Section (ref). Section (ref) presents Monte Carlo evidence. Section (ref) revisits the aggregate consumption function empirical application in \citeN{PesaranShinSmith1999} . Section (ref) concludes. Mathematical derivations and proofs are provided in Appendix A. Details on the implementation of individual estimators and bootstrapping, and additional Monte Carlo results are provided in Appendix B.
We adopt the same setting as in \citeN{PesaranShinSmith1999} , and consider the following illustrative model
for $i=1,2,...,n$, and $t=1,2,...,T$. For expositional clarity and notational simplicity, we focus on a single regressor and one lag, but it is understood that our analysis is applicable to multiple lags of $\Delta \mathbf{z}_{it}=\left( \Delta y_{it},\Delta x_{it}\right) ^{\prime }$ entering both equations ((ref))-((ref)), and the approach is also applicable to multiple $x_{it}$'s with a single long run relationship. We consider the following assumptions:
Substituting first ((ref)) for $u_{y,it}$ in ((ref)), and then substituting $u_{x,it}=\Delta x_{it}$, we obtain the following ARDL representation for $y_{it}$
The pooled Bewley estimator takes advantage of the Bewley transform ( \citeANP{Bewley1979}, \citeyearNP{Bewley1979} ). Subtracting $\left( 1-\alpha _{i}\right) y_{it}$ from both sides of ((ref)) and re-arranging, we have
or (noting that $\alpha _{i}>0$ for all $i$ and multiplying the equation above by $\alpha _{i}^{-1}$)
where $\Delta \mathbf{z}_{it}=\left( \Delta y_{it},\Delta x_{it}\right) ^{\prime }$, and $\mathbf{\psi }_{i}=\left( -\frac{1-\alpha _{i}}{\alpha _{i} },\frac{\delta _{i}}{\alpha _{i}}\right) ^{\prime }$. Further, stacking ((ref)) for $t=1,2,...,T$, we have
where $\mathbf{y}_{i}=\left( y_{i1},y_{i2},...,y_{iT}\right) ^{\prime }$, $ \mathbf{x}_{i}=\left( x_{i1},x_{i2},...,x_{iT}\right) ^{\prime }$, $\Delta \mathbf{Z}_{i}=\left( \Delta \mathbf{z}_{i1}^{\prime },\Delta \mathbf{z} _{i2}^{\prime },...,\Delta \mathbf{z}_{iT}^{\prime }\right) ^{\prime }$, $ \mathbf{v}_{i}=\left( v_{i,1},v_{i,2},...,v_{i,T}\right) ^{\prime }$, and $ \mathbf{\tau }_{T}$ is $T\times 1$ vector of ones. Define projection matrix $ \mathbf{M}_{\tau }=\mathbf{I}_{T}-T^{-1}\mathbf{\tau }_{T}\mathbf{\tau } _{T}^{\prime }$. This projection matrix subtracts the period average. Let $ \mathbf{\tilde{y}}_{i}=\left( \tilde{y}_{i1},\tilde{y}_{i2},...,\tilde{y} _{iT}\right) ^{\prime }=\mathbf{M}_{\tau }\mathbf{y}_{i}$, and similarly $ \mathbf{\tilde{x}}_{i}=\left( \tilde{x}_{i1},\tilde{x}_{i2},...,\tilde{x} _{iT}\right) ^{\prime }=\mathbf{M}_{\tau }\mathbf{x}_{i}$, $\Delta \mathbf{ \tilde{Z}}_{i}=\mathbf{M}_{\tau }\Delta \mathbf{Z}_{i}$, and $\mathbf{\tilde{ v}}_{i}=\mathbf{M}_{\tau }\mathbf{v}_{i}$. Multiplying ((ref)) by $\mathbf{ M}_{\tau }$, we have
Now consider the matrix of instruments
where $\mathbf{y}_{i,-1}=\left( y_{i,1},y_{i,1},...,y_{i,T-1}\right) ^{\prime }$ is the data vector on the first lag of $y_{it}$, similarly $ \mathbf{x}_{i,-1}=\left( x_{i,1},x_{i,1},...,x_{i,T-1}\right) ^{\prime }$. The PB estimator of $\beta $ is given by
where
and
is the projection matrix associated with $\mathbf{\tilde{H}}_{i}$.
In addition to Assumptions (ref)-(ref), we also require the following high-level conditions to hold in the derivations of the asymptotic distribution of the PB estimator under the joint asymptotics $n,T\rightarrow \infty $.
Substituting $\mathbf{\tilde{y}}_{i}=\mathbf{\tilde{x}}_{i}\beta +\Delta \mathbf{\tilde{Z}}_{i}\mathbf{\psi }_{i}+\alpha _{i}^{-1}\mathbf{\tilde{v}} _{i}$ in ((ref)), and using $\mathbf{M}_{i}\Delta \mathbf{\tilde{Z}}_{i}= \mathbf{0}$, we have
Consider the first term on the right side of ((ref)) first. Since $ \mathbf{M}_{i}$ is an orthogonal projection matrix, $\mathbf{\tilde{x}} _{i}^{\prime }\mathbf{M}_{i}\mathbf{\tilde{x}}_{i}/T^{2}$ is bounded by $ \mathbf{\tilde{x}}_{i}^{\prime }\mathbf{\tilde{x}}_{i}/T^{2}$. The second moments of $\mathbf{\tilde{x}}_{i}^{\prime }\mathbf{\tilde{x}}_{i}/T^{2}$ are bounded, and, in addition, $\mathbf{\tilde{x}}_{i}^{\prime }\mathbf{M} _{i}\mathbf{\tilde{x}}_{i}/T^{2}$ is cross-sectionally independent. It follows that $\frac{1}{n}\sum_{i=1}^{n}\mathbf{\tilde{x}}_{i}^{\prime } \mathbf{M}_{i}\mathbf{\tilde{x}}_{i}/T^{2}$ converges to a constant, which we denote by $\omega _{x}^{2}$, as $n,T\rightarrow \infty $. Lemma (ref) in Appendix A establishes the expression for $\omega _{x}^{2}=\sigma _{x}^{2}/6$, where $\sigma _{x}^{2}=\lim_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}\sigma _{xi}^{2}$, but the specific expression for $ \omega _{x}^{2}$ is not relevant for the inference approach that we adopt below. Consider the second term of ((ref)) next,
The term in the square brackets has zero mean and is independently distributed over $i$. For the asymptotic distribution to be correctly centered we need
as $n$ and $T\rightarrow \infty $. This condition holds so long as $n=\Theta \left( T^{\theta }\right) $ for some $0<\theta <2$. See Lemma (ref) for a proof. The asymptotic distribution of the first term in ((ref)) is in turn established by Lemma (ref), see ((ref)). The following theorem now follows for the asymptotic distribution of $\hat{\beta}$.
To conduct inference, let
and $\hat{\omega}_{v}^{2}=n^{-1}\sum_{i=1}^{n}\left( \frac{\mathbf{x} _{i}^{\prime }\mathbf{M}_{i}\mathbf{\hat{v}}_{i}^{\ast }}{T}\right) ^{2}$, where $\mathbf{\hat{v}}_{i}^{\ast }=\mathbf{M}_{i}\left( \mathbf{y}_{i}-\hat{ \beta}\mathbf{x}_{i}\right) $, and $\mathbf{M}_{i}$ is defined by ((ref) ). Accordingly, we propose the following estimator of $\Omega $:
When $n$ is sufficiently large relative to $T$, specifically when $\sqrt{n} /T\rightarrow K>0$, then $\sqrt{n}T\left( \hat{\beta}-\beta \right) $ is no longer asymptotically distributed with zero mean. The asymptotic bias is due to the nonzero mean of $T^{-1}\mathbf{\tilde{x}}_{i}^{\prime }\mathbf{M}_{i} \mathbf{\tilde{v}}_{i}$, and it can be of some relevance for finite sample performance, as the Monte Carlo evidence in Section (ref) illustrates. Monte Carlo evidence also reveals that the inference based on PB and other existing estimators in the literature can suffer from serious size distortions in finite samples. To deal with these problems, we consider bootstrapping critical values using sieve wild bootstrap for more accurate and more robust inference that allows for cross-sectional dependence of errors. In addition, we also consider two bias-correction techniques - a bootstrap one as well as the split-panel jackknife method. The same bias-correction methods are also applied to the three other estimators, namely PMG, PDOLS, and FMOLS, considered in the paper. In what follows we focus on the PB\ estimator. A description of bias-corrections applied to the other three estimators are given in Section (ref) of Appendix B.
Once an estimate of the bias of $\hat{\beta}$ is available, denoted as $\hat{ b}$, then the bias-corrected PB estimator is given by
One possibility of estimating the bias in the literature is by bootstrap. We adopt the following sieve wild bootstrap algorithm for generating simulated data.
Using simulated data with $R=10,000$, we compute an estimate of the bias $ \hat{b}_{R}=\left[ R^{-1}\sum_{r=1}^{R}\hat{\beta}^{\left( r\right) }-\hat{ \beta}\right] $. We then compute the $\alpha $ percent critical values using the $1-\alpha $ percent quantile of $\left\{ \left\vert t^{\left( r\right) }\right\vert \right\} _{r=1}^{R}$, where $t^{\left( r\right) }=\tilde{\beta} ^{\left( r\right) }/se\left( \tilde{\beta}^{\left( r\right) }\right) $, $ \tilde{\beta}^{\left( r\right) }=\hat{\beta}^{\left( r\right) }-\hat{b}$ is the bias-corrected estimate of $\beta $ using the $r$-$th$ draw of the simulated data, $se\left( \tilde{\beta}^{\left( r\right) }\right) =T^{-1}n^{-1/2}\hat{\Omega}^{\left( r\right) }$ is the corresponding standard error estimate, and $\hat{\Omega}^{\left( r\right) }$ is computed in the same way as $\hat{\Omega}$ in ((ref)) but using the simulated data.
The split-panel jackknife bias correction method is given by
where $\hat{\beta}$ is the full sample PB estimator, $\hat{\beta}_{a}$ and $ \hat{\beta}_{b}$ are the first and the second half sub-sample PB estimators, and $\kappa $ is a suitably chosen weighting parameter. In a stationary setting, where the bias is of order $O\left( T^{-1}\right) $, $\kappa $ is chosen to be one, so that $\frac{K}{T}-\kappa \cdot \left( \frac{K}{T/2}- \frac{K}{T}\right) =0,$ for any arbitrary choice of $K$. See, for example, \citeN{DhaeneJochmansy2015} and \citeN{ChudikPesaranYang2018} .
In general, when the bias is of order $O\left( T^{-\epsilon }\right) $ for some $\epsilon >0$, then $\kappa $ can be chosen to solve $\frac{K}{ T^{\epsilon }}-\kappa \cdot \left( \frac{K}{\left( T/2\right) ^{\epsilon }}- \frac{K}{T^{\epsilon }}\right) =0$, which yields $\kappa =1/\left( 2^{\epsilon }-1\right) $. Under our setup with I(1) variables, we need to correct $\hat{\beta}$ for its $O\left( T^{-2}\right) $ bias, namely $ \epsilon =2$, which yields $\kappa =1/3$.
Inference using $\tilde{\beta}^{jk}$ can be conducted based on ((ref)) but with $\hat{\omega}_{v}^{2}$ replaced by
where $\mathbf{\tilde{v}}_{i}^{\ast }=\mathbf{M}_{i}\left( \mathbf{y}_{i}- \tilde{\beta}^{jk}\mathbf{x}_{i}\right) $,
$\mathbf{x}_{a,i}^{\prime }$ $\left( \mathbf{x}_{b,i}^{\prime }\right) $ and $\mathbf{M}_{a,i}$ ($\mathbf{M}_{b,i}$) are defined in the same way as $ \mathbf{x}_{i}$, and $\mathbf{M}_{i}$ but using only the first (second) half of the sample.
We compute bootstrapped critical values to conduct more accurate and robust small sample inference. Specifically, the $\alpha $ percent critical value is computed as the $1-\alpha $ percent quantile of $\left\{ \left\vert t_{jk}^{\left( r\right) }\right\vert \right\} _{r=1}^{R}$, where $ t_{jk}^{\left( r\right) }=\tilde{\beta}_{jk}^{\left( r\right) }/se\left( \tilde{\beta}_{jk}^{\left( r\right) }\right) $, $\tilde{\beta}_{jk}^{\left( r\right) }$ is the jackknife estimate of $\beta $ using the $r$-$th$ draw of the simulated data generated using the algorithm described in Subsection (ref), $se\left( \tilde{\beta}_{jk}^{\left( r\right) }\right) $ is the corresponding standard error estimate, namely $se\left( \tilde{\beta} _{jk}^{\left( r\right) }\right) =T^{-1}n^{-1/2}\hat{\Omega}_{jk}^{\left( r\right) }$, $\hat{\Omega}_{jk}^{\left( r\right) }=\hat{\omega}_{x,\left( r\right) }^{-4}\tilde{\omega}_{v,\left( r\right) }^{2}$, in which $\tilde{ \omega}_{v,\left( r\right) }$ and $\hat{\omega}_{x,\left( r\right) }^{2}$ are computed using the simulated data, based on expressions ((ref)) and ((ref)), respectively.
The Data Generating Process (DGP) is given by ((ref))-((ref)), for $ i=1,2,...,n,$ $T=1,2,...,T$, with starting values satisfying Assumption (ref) with $\mathbf{\mu }_{i}\sim IIDN\left( \mathbf{\tau }_{2},\mathbf{I} _{2}\right) $, and $c_{i}=$ $\alpha _{i}\mu _{i,1}-\alpha _{i}\beta \mu _{i,2}$. We generate $\alpha _{i}\sim IIDU\left[ 0.2,0.3\right] $. We consider two DGPs based on the cross-sectional dependence of errors. In the cross-sectionally independent DGP, we generate $u_{y,it}=\sigma _{y,i}e_{y,it}$, $u_{x,it}=\sigma _{x,i}e_{x,it}$, $\sigma _{y,i}^{2},\sigma _{x,i}^{2}\sim IIDU\left[ 0.8,1.2\right] $,
In the DGP with cross-sectionally dependent errors, we generate $e_{y,it}$ to contain a factor structure including strong, semi-strong and weak factors:
where $\varepsilon _{y,it}\sim IIDN\left( 0,1\right) $, $f_{\ell ,t}\sim IIDN\left( 0,1\right) $, $\gamma _{\ell }\sim IIDU\left[ 0,\gamma _{\max ,\ell }\right] $, for $\ell =1,2,...,m$. We choose $m=5$ factors and $\gamma _{\max ,\ell }=2n^{\alpha _{\ell }-1}$ with $\alpha _{\ell }=1,0.9,0.8,0.7,0.6$, for $\ell =1,2,...,5$, respectively. Scaling constant $ \varkappa _{i}$ is set to ensure $E\left( e_{y,it}^{2}\right) =1$, namely $ \varkappa _{i}=\left( 1+\sum_{\ell =1}^{m}\gamma _{i\ell }^{2}\right) ^{-1/2} $. We generate $e_{x,it}$ to ensure unit variance and $cov\left( e_{y,it},e_{x,it}\right) =\rho _{i}$. Specifically, $e_{x,it}=\rho _{i}e_{y,it}+\sqrt{1-\rho _{i}^{2}}\varepsilon _{x,it}$, $\varepsilon _{x,it}\sim IIDN\left( 0,1\right) $. Both designs features heteroskedastic (over $i$) and correlated (over $y$ & $x$ equations) errors, namely $ E\left( u_{y,it}^{2}\right) =\sigma _{y,i}^{2}$, $E\left( u_{x,it}^{2}\right) =\sigma _{x,i}^{2}$, and $corr\left( u_{y,it},u_{x,it}\right) =\rho _{i}$. We consider $n,T=20,30,40,50$ and compute $R_{MC}=2000$ Monte Carlo replications.
We report bias, root mean square error (RMSE), size ($H_{0}:\beta =1$, 5% nominal level) and power ($H_{1}:\beta =0.9$, 5% nominal level) findings for the PB estimator $\hat{\beta}$ given by ((ref)), with variance estimated using ((ref)). Moreover, we also report findings for the two bias corrected versions of PB estimator as described in Subsection (ref) with bootstrapped critical values for inference robust to error cross-sectional dependence. We compare the performance of the PB estimator with the PMG estimator by \citeN{PesaranShinSmith1999} , panel dynamic OLS (PDOLS) estimator by \citeN{MarkSul2003} , and the group-mean fully modified OLS (FMOLS) estimator by \citeANP{Pedroni1996} (\citeyearNP{Pedroni1996}, \citeyearNP{Pedroni2001ReStat}) . Similarly to the PB estimator, we also consider jackknife and bootstrap based bias-corrected versions of the PMG, PDOLS and FMOLS estimators with cross-sectionally robust bootstrapped critical values, described in Appendix B. We use $R_{b}=10,000$ bootstrap replications (within each MC replication) for bootstrap bias correction and for computation of robust and more accurate bootstrapped critical values.
Table 1 report the results for the original (without bias-correction) estimators. PB estimator stands out as the most precise estimator in terms of having the lowest RMSE values among the four estimators. The second best is PMG estimator with RMSE values 1 to 21 percent larger compared with the PB estimator, the third is PDOLS with RMSE values 23 to 66 percent larger compared with PB, and the FMOLS comes last with RMSE\ values 95 to 180 percent larger compared with PB. In terms of the bias alone, the ordering of the estimators is slightly different with PMG and PB switching their places. For $T=20$, the bias of PMG estimator is -0.016 to -0.020, the bias of PB estimator is in the range -0.034 to -0.037, the bias of the PDOLS estimator is in the range -0.052 to -0.056 and the bias of the FMOLS estimator is in the range -0.104 to -0.110. For such a small value of $T$, the bias is not very large, and, as expected, it declines with an increase in $T$.
All four estimators suffer from varying degrees of size distortions. The inference based on the PB estimator is the most accurate. Specifically, the size distortions for the PB estimator are lowest among the four estimators - with reported size in the range between 9.9 and 25.2 percent, exceeding the chosen nominal value of 5 percent. Size distortions diminish with an increase in $T$.
We consider next the bias-corrected versions of the four estimators with inference carried out using robust bootstrap critical values. Upper panel of Table 2 reports findings for estimators corrected for bias using the jackknife procedure, and the bottom panel reports on bootstrap bias corrected estimators. Bias correction did not change the overall ranking of estimators -- PB continues to be the most precise (lowest RMSE). Both bias correction approaches are quite effective in reducing the bias. The bias of PB\ and PMG estimators for any of the two bias corrections are very low. In addition to reducing the bias, in many cases the bias-correction also resulted in reduced RMSE values. In the case of the PB estimator, using bootstrap bias correction resulted in improved RMSE performance for all choices of $n,T$ - by about 2 to 22 percent. Results in Table 2 also show notable improvement to inference comes from using bootstrapped critical values - with PB having virtually no size distortions and size distortions of the remaining estimators are relatively minor.
Last but not least, we consider the DGP with cross-sectionally correlated errors. The corresponding results, reported in Tables B1 and B2 in Appendix B, reveal the same ranking of the four estimators, and, importantly, the bootstrapped critical values continue to deliver correct size, despite the error cross-sectional dependence.
The Monte Carlo results show that PB\ estimator can perform better (in terms of overall precision as measured by RMSE, and in terms of accuracy of inference) than existing estimators (PMG, PDOLS, and FMOLS) in finite sample sizes of interest, whether or not bias correction is considered. Of' course, our results do not imply that PB estimator will always be better, but that it can be a useful addition to the existing literature as a complement to PMG, PDOLS, and FMOLS estimators. Bias corrections and bootstrapping critical values are helpful for all four estimators, resulting not only in reduced bias, but sometimes also in better RMSE. In all cases, they result in more accurate inference in our experiments.
This section revisits consumption function empirical application undertaken by \citeN{PesaranShinSmith1999} , hereafter PSS. The long-run consumption function is assumed to be given by
for country $i=1,2,...,n$, where $c_{it}$ is the logarithm of real consumption per capita, $y_{it}^{d}$ is the logarithm of real per capita disposable income, $\pi _{it}$ is the rate of inflation, and $\vartheta _{it} $ is an $I\left( 0\right) $ process. We take the dataset from PSS, which consists of $n=24$ countries and a slightly unbalanced time period covering 1960-1993. PSS estimate $\beta _{1}$ and $\beta _{2}$ using an ARDL(1,1,1) specification, which can be written as error-correcting panel regressions
for $i=1,2,...,n$, where all coefficients, except the long-run coefficients $ \beta _{1}$ and $\beta _{2}$ are country-specific.
Table 3 presents alternative estimates of the long-run coefficients. The upper panel presents findings for estimators without bias correction and standard confidence intervals. The middle and lower panels present jackknife and bootstrap\ bias-corrected estimates with confidence intervals based on bootstrapped critical values. Results differ widely across different approaches to estimation and inference. Depending on which bias correction approach is conducted, the PB\ estimates of the long-run coefficient on real income ($\beta _{1}$) is estimated to be 0.921 or 0.926, and the long-run coefficient on the inflation variable ($\beta _{2}$) is estimated to be -0.120 or -0.125. The null hypothesis that the coefficient on $y_{it}^{d}$ is unity cannot be rejected at the 5 percent nominal level, nor is the hypothesis that the long run coefficient on inflation is zero. From an economic perspective, unit long-run real income elasticity and no long-run effects of inflation on consumption seem both plausible - the former hypothesis is in line with balanced growth path models, and the latter in line with monetary policy neutrality in the long-run. A different conclusion would be reached according to PMG estimates - namely both the unit coefficients on the real income variable and zero coefficient on inflation would be rejected at the 5 percent nominal level. The results based on the PDOLS are in line with the PB estimates and do not reject unit real income and zero inflation long run coefficients. FMOLS estimates of $ \beta _{1}$ are larger than the other estimates, but the unit coefficient on the income variable still cannot be rejected. The FMOLS estimates of $\beta _{2}$ are also quite large. The choice of estimation method clearly matters in this empirical illustration.
\doublespacing
This paper proposes the pooled Bewley (PB)\ estimator of long-run relationships in heterogeneous dynamic panels. Relative to existing estimators in the literature -- namely PMG, PDOLS and FMOLS -- Monte Carlo evidence reveals that PB can perform well in small samples. While we developed the asymptotic theory of PB estimator under a similar setting to the PMG estimator, notably we assumed cross-sectionally independent errors, we have also shown the benefit of bootstrapping critical values for inference when errors are cross-sectionally correlated for all four estimators.
While the asymptotic distribution of the other estimators are derived for the case where $n$ is fixed and $T\rightarrow \infty ,$ we derive the joint $ \left( n,T\right) $ asymptotic distribution of the PB estimator, when both $ n $ and $T$ diverge to infinity jointly such that $n=\Theta \left( T^{\theta }\right) $, for $0<\theta <2.$ This covers a broader range of empirical applications where both $n$ and $T$ are large. The small sample and asymptotic results suggest that the PB estimator is a useful addition to estimators for long run effects in single equation dynamic heterogeneous panels, where the direction of long-run causality is known.