Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
174,590 characters · 0 sections · 49 citation commands
Estimation of Average Effects in Short $T$ Heterogeneous Panels
\thispagestyle{empty}
\setcounter{page}{1}
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Introduction}
\doublespacing Fixed effects estimation of average effects has been predominantly utilized for program and policy evaluation. For static panel data models where slope heterogeneity is uncorrelated with regressors, two-way fixed effects (TWFE) estimators that allow for unit-specific and time effects are $\sqrt{n}$ -consistent, and if used in conjunction with robust standard errors lead to valid inference in panels where the time dimension, $T$, is short and the cross section dimension, $n$, is sufficiently large. However, when the slope heterogeneity is correlated with the regressors, the TWFE estimators could become inconsistent even if both $n$ and $T\rightarrow \infty$.\footnote{ The concept of the correlated random coefficient model is due to HeckmanVytlacil1998. Wooldridge2005 shows that TWFE estimators continue to be consistent if slope heterogeneity is mean-independent of all the de-trended covariates. See also condition ((ref)) given below.} Such correlated heterogeneity arises endogenously in the case of dynamic panel data models, as originally noted by PesaranSmith1995, and more generally, when the regressors are weakly exogenous. In the case of static panels with strictly exogenous regressors, correlated (slope) heterogeneity can arise, for example, when there is a high degree of variation in treatments across units or when there are latent factors that influence the level and variability of treatments and their outcomes. For example, in estimation of returns to education, the choice of educational level is likely to be correlated with expected returns to education. Other examples include estimation of the effects of training programs on workers' productivity and earnings reviewed by CreponVandenberg2016, evaluation of the effectiveness of micro-credit programs discussed by BanerjeeEtal2015, and the analysis of the effectiveness of anti-poverty cash transfer programs considered by BastagliEtal2019.
In the presence of correlated heterogeneity, PesaranSmith1995 proposed to estimate the mean effects by simple averages of individual estimates, which they called the mean group (MG) estimator. They showed that the MG estimator is consistent for dynamic heterogeneous panels when $n$ and $T$ are both large. It was later shown that for panels with strictly exogenous regressors, the MG estimator is in fact $\sqrt{n}$-consistent in the presence of correlated heterogeneity even if $T$ is fixed as $ n\rightarrow \infty $, so long as $T$ is sufficiently large such that second-order moments of the individual estimates exist. However, such moment conditions need not hold when $T$ is very close to the number of estimated coefficients ($k$). In effect, we are faced with the problem of estimating the mean of random variables with fat tails.
This problem was originally recognized by Chamberlain1992, who showed that one needs $T$ to be strictly larger than $k$ for regular identification of average effects under correlated heterogeneity.\footnote{ An unknown parameter, ${\Greekmath 010C} _{0}$, is said to be regularly identified if there exists an estimator that converges to ${\Greekmath 010C} _{0}$ in probability at the rate of $\sqrt{n}$. Any estimator that converges to its true value at a rate slower than $\sqrt{n}$ is said to be irregularly identified.} In the statistics literature, CsorgoEtal1988b,CsorgoEtal1988a and GriffinPruitt1989, among others, consider a trimmed mean estimator whereby the observations are ordered, and those below and above a given threshold value are excluded. However, in general such trimmed estimators may not possess a limiting distribution if the observations are drawn from a heavy-tailed distribution with the tail index, ${\Greekmath 010B} _{p}<2$, and the threshold values are chosen in an ad hoc manner. Peng2001 proposes a new trimmed estimator where the thresholds are endogenized and the trimmed mean is augmented with mean estimates from the two tails. Peng's modified trimmed estimator is shown to have the same limiting normal distribution as the sample mean but with a slower convergence rate when $ {\Greekmath 010B} _{p}<2$ (see Theorem 1 and Remark 1 of Peng2001). For settings of inverse probability weighting where the denominator can be arbitrarily close to zero, leading to heavy-tailed sampling distributions, MaWang2020 propose trimmed estimators that exclude observations with small denominators and develop inference procedures for different trimming threshold choices. For our purposes, trimmed estimators proposed by Peng2001 and MaWang2020 are subject to two limitations. They consider only the scalar case and assume that the observations (in our application, the individual estimates) are identically and independently distributed. Extension of their estimators to a vector of estimates that are not identically distributed does not seem to be straightforward.
In the context of MG estimation, we have additional information about the precision of the individual estimates that are not used in the trimming approaches considered in the statistics literature. One example where such information is utilized is provided by GrahamPowell2012 (GP), who build on the pioneering work of Chamberlain1992. GP focus on panels with $T=k$, where identification issues of time effects and the mean coefficients arise especially when there are insufficient within-individual variations for some regressors. These authors derive an irregular estimator of the mean coefficients by excluding individual estimates from the estimation of the average effects if the sample variance of regressors in question is smaller than a given threshold value.
In this paper, we begin by providing conditions under which MG and fixed effects (FE) estimators of the average effects in heterogeneous panels are $ \sqrt{n}$-consistent, which serves as a basis for developing a diagnostic test of the validity of (two-way) fixed effects estimators commonly used in the literature. We then propose a trimmed mean group (TMG) estimator that does not exclude any of the individual estimates, but uses a threshold function similar to that of GP to shrink some of the estimates to overcome the fat-tailed nature of the distribution of the individual estimates when $ T $ is very close to $k$. In effect, we shrink rather than drop estimates as done by GP. The decision on whether unit $i$ is subject to shrinkage is made with respect to the determinant of the sample variance matrix of the regressors, denoted by $d_{i}(>0)$. Individual estimates become ill-conditioned when $d_{i}$ is close to zero. To characterize the trimming process, we assume that $1/d_{i}$ follows Pareto-type distributions with the tail index, ${\Greekmath 010B} _{p}$, and show that individual estimates have second-order moments when ${\Greekmath 010B} _{p}>2$. Shrinkage of the individual estimates is required only if ${\Greekmath 010B} _{p}\leq 2$. The literature on estimation of ${\Greekmath 010B} _{p}$ is well established, for example, EmbrechtsEtal1997, which can guide decisions on shrinkage/trimming before estimating the average effects. Based on extensive Monte Carlo experiments, we find that in general ${\Greekmath 010B} _{p}>2$ when $T\geq 2k+1$, which could be used as a practical rule of thumb when deciding whether to use the MG estimator or its trimmed version proposed in this paper.
We also consider heterogeneous panels with time effects and show that the TMG procedure can be applied after the time effects are eliminated. In the case where $T=k$, we require the dependence between heterogeneous slope coefficients and the regressors to be time-invariant. This assumption is not required when $T>k$, and the time effects can be eliminated using the approach first proposed by Chamberlain1992. We refer to these estimators as TMG-TE and derive their asymptotic distributions, bearing in mind that the way the time effects are eliminated depends on whether $T=k$ or $T>k$.
As noted above, slope heterogeneity by itself does not render the TWFE estimators inconsistent, which continue to have the regular convergence rate of $\sqrt{n}$. The problem arises when slope heterogeneity is correlated with the covariates. It is therefore important that before using the TWFE estimators, the null of uncorrelated heterogeneity is tested. To this end, we also propose Hausman tests of correlated heterogeneity by comparing the FE and TWFE estimators with the associated TMG estimators and derive their asymptotic distributions under fairly general conditions. The earlier Hausman test of slope homogeneity developed by PesaranEtal1996 is based on the difference between FE and MG estimators and does not apply when $T$ is close to $k$. The more recent dispersion-based test of slope homogeneity proposed by PesaranYamagata2008 is shown to be quite powerful as a test of slope homogeneity but does not distinguish correlated or uncorrelated heterogeneity and requires $\sqrt{n}/T^{2}\rightarrow 0$ as $ n$ and $T\rightarrow \infty $ jointly.
We also carry out an extensive set of Monte Carlo (MC) simulations to investigate the small sample properties of the TMG and TMG-TE estimators and how they compare with the trimmed estimator proposed by GP. The MC evidence on the size and empirical power of the Hausman tests of correlated heterogeneity in panel data models with and without time effects is provided, and the sensitivity of estimation results to the choice of the trimming threshold parameter, ${\Greekmath 010B} $, is also investigated. The MC and theoretical results of the paper are all in agreement. The TMG and TMG-TE estimators not only have the correct size but also achieve better finite sample properties compared with the other trimmed estimator across a number of experiments with different data generating processes, allowing for heteroskedasticity (random and correlated), error serial correlations, and regressors with heterogeneous dynamics and interactive effects. The simulation results also confirm that the Hausman tests based on the difference between FE (TWFE) and TMG (TMG-TE) estimators have the correct size and power against the alternative of correlated heterogeneity.
Finally, we illustrate the utility of our proposed trimmed estimators by re-examining the average effect of household expenditures on calorie demand using a balanced panel of $1,358$ households in poor rural communities in Nicaragua over the years 2001--2002 $(T=2)$ and 2000--2002 $(T=3)$.
The rest of the paper is organized as follows. Section (ref) sets out the heterogeneous panel data model and discusses the asymptotic properties of FE and MG estimators. Section (ref) considers ultra short $T$ panels (including the case of $T=k$) and introduces the proposed TMG estimator, with its asymptotic properties established in Section (ref) . Section (ref) extends the TMG estimation to panels with time effects, distinguishing between cases where $T>k$ and $T=k$. Section (ref) sets out the Hausman test of correlated heterogeneous slope coefficients. Section (ref) discusses how to apply the TMG approach to a subset of coefficients of interest. Section (ref) provides the main findings of the MC experiments. Section (ref) presents the empirical illustration. Section (ref) concludes. Mathematical proofs of propositions and theorems are given in a mathematical appendix. Supplementary materials covering additional mathematical derivations, MC experiments, as well as further empirical results, are provided in an online supplement.
Notations: Generic positive finite constants are denoted by $C$ when large, and $c$ when small. They can take different values at different instances. ${\Greekmath 0115} _{\max }\left( \boldsymbol{A}\right) $ and ${\Greekmath 0115} _{\min }\left( \boldsymbol{A}\right) $ denote the maximum and minimum eigenvalues of matrix $\boldsymbol{A}$. $\boldsymbol{A}\succ \boldsymbol{0}$ and $\boldsymbol{A}\succeq \boldsymbol{0}$ denote that matrix $\boldsymbol{A} $ is positive definite and is positive semi-definite, respectively. When matrix $\boldsymbol{A}$ is square, its adjugate (adjoint) and determinant are denoted by $\func{adj}(\boldsymbol{A})$ and $\det (\boldsymbol{A})$, respectively. If $\det (\boldsymbol{A})\neq 0$, then the inverse of $ \boldsymbol{A}$ is given by $\boldsymbol{A}^{-1}=\func{adj}(\boldsymbol{A)} /\det (\boldsymbol{A)}$. $\left\Vert \boldsymbol{A}\right\Vert ={\Greekmath 0115} _{\max }^{1/2}(\boldsymbol{A}^{\prime }\boldsymbol{A)}$ and $\left\Vert \boldsymbol{A}\right\Vert _{1}$ denote the spectral and column norms of matrix $\boldsymbol{A}$, respectively. $\left\Vert \boldsymbol{x}\right\Vert _{p}=\left[ E\left( \left\Vert \boldsymbol{x}\right\Vert ^{p}\right) \right] ^{1/p}$. If $\left\{ f_{n}\right\} _{n=1}^{\infty }$ is any real sequence and $\left\{ g_{n}\right\} _{n=1}^{\infty }$ is a sequence of positive real numbers, then $f_{n}=O(g_{n})$ if there exists $C$ such that $\left\vert f_{n}\right\vert /g_{n}\leq C$ for all $n$, and $f_{n}=o(g_{n})$ if $ f_{n}/g_{n}\rightarrow 0$ as $n\rightarrow \infty $. Similarly, $ f_{n}=O_{p}(g_{n})$ if $f_{n}/g_{n}$ is stochastically bounded, and $ f_{n}=o_{p}(g_{n})$, if $f_{n}/g_{n}\rightarrow _{p}0$. $f_{n}=\ominus (g_{n})$ if there exist $n_{0}\geq 1$ and positive finite constants $C_{0}$ and $C_{1}$, such that $\inf_{n\geq n_{0}}\left( \left\vert f_{n}\right\vert /g_{n}\right) \geq C_{0}$, and $\sup_{n\geq n_{0}}\left( \left\vert f_{n}\right\vert /g_{n}\right) \leq C_{1}$. The operator $\rightarrow _{p}$ denotes convergence in probability, and $\rightarrow _{d}$ denotes convergence in distribution. $IID$ stands for independently and identically distributed. $\boldsymbol{u}\perp \boldsymbol{v}$ is used to show that vectors of random variables $\boldsymbol{u}$ and $\boldsymbol{v}$ are independently distributed.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Heterogeneous linear panel data models}
Consider the following panel data model with individual fixed effects, $ {\Greekmath 010B} _{i}$, and heterogeneous slope coefficients, $\boldsymbol{{\Greekmath 010C} }_{i}$ ,
where $\boldsymbol{x}_{it}$ is a $k^{\prime }\times 1$ vector of regressors, and $u_{it}$ is the error term. $\left\{ \boldsymbol{{\Greekmath 010C} }_{i}\right\} _{i=1}^{n}$ follow the random coefficient model
where $\left\{ \boldsymbol{{\Greekmath 0111} }_{i}\right\} _{i=1}^{n}$ are the random components, and $\boldsymbol{{\Greekmath 010C} }_{0}$ is the $k^{\prime }\times 1$ vector of average effects. In matrix notations,
where $\boldsymbol{y}_{i}=(y_{i1},y_{i2},...,y_{iT})^{\prime }$, $ \boldsymbol{{\Greekmath 011C} }_{T}$ is a $T\times 1$ vector of ones, $\boldsymbol{X} _{i}=(\boldsymbol{x}_{i1},\boldsymbol{x}_{i2},...,\boldsymbol{x} _{iT})^{\prime }$, and $\boldsymbol{u}_{i}=(u_{i1},u_{i2},...,u_{iT})^{ \prime }$. The FE estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ is given by
where $\boldsymbol{M}_{T}=\boldsymbol{I}_{T}-T^{-1}\boldsymbol{{\Greekmath 011C} }_{T} \boldsymbol{{\Greekmath 011C} }_{T}^{\prime }$, and $\boldsymbol{I}_{T}$ is a $T\times T$ identity matrix. It is well known that for a fixed $T\geq k=k^{\prime }+1$, $ \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ is a $\sqrt{n}$-consistent estimator of $ \boldsymbol{{\Greekmath 010C} }_{0}$ and robust to possible correlations between ${\Greekmath 010B} _{i}$ and $\{\boldsymbol{x}_{it}\}_{t=1}^{T}$, even under slope heterogeneity so long as $\boldsymbol{{\Greekmath 0111} }_{i}$ are not correlated with $ \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}$. This condition is clearly satisfied when heterogeneity is exogenous and $E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{ {\Greekmath 0111} }_{i}\right) =\boldsymbol{0}$, for the majority of the units (to be formalized below).
In the presence of correlated slope heterogeneity, the MG estimator, initially proposed by PesaranSmith1995, is typically considered for consistent estimation of the average effects. When $T\geq k$, $\boldsymbol{ {\Greekmath 010C} }_{0}$ can be estimated by the MG estimator, $\boldsymbol{\hat{{\Greekmath 010C}}} _{MG}$, computed as a simple average of the least square estimates of $ \boldsymbol{{\Greekmath 010C} }_{i}$, namely
where
In contrast to the FE estimator, the MG estimator does not depend on $ E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i} \boldsymbol{{\Greekmath 0111} }_{i}\right) $ and is consistent irrespective of whether slope heterogeneity is correlated or not. As pointed out by an associate editor, $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ can also be written as a weighted average of $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$,
where $\boldsymbol{\mathcal{W}}_{i}=\left( n^{-1}\sum_{j=1}^{n}\boldsymbol{X} _{j}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{j}\right) ^{-1}\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) $ is the $k^{\prime }\times k^{\prime }$ weight matrix for unit $i$. The difference between $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ and $\boldsymbol{\hat{ {\Greekmath 010C}}}_{MG}$ lies in the choice of the weights. When heterogeneity is exogenous, both estimators converge to the same limit $\boldsymbol{{\Greekmath 010C} } _{0}$. However, when heterogeneity is correlated, as we shall see, $ \boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ continues to converge to $\boldsymbol{{\Greekmath 010C} } _{0}$, but $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ converges to $\boldsymbol{{\Greekmath 010C} } _{0}+\lim\limits_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{\mathcal{W}}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) $, where $ \lim\limits_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{ \mathcal{W}}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) \neq \boldsymbol{0}$.
To investigate the asymptotic properties of FE and MG estimators for a fixed $T\geq k$ as $n\rightarrow \infty $, we consider the following assumptions:
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Why does correlated slope heterogeneity matter?}
Correlated slope heterogeneity can arise in various contexts, with important implications for estimation and inference on the average effects. For example, return to education is likely to be positively correlated with latent factors such as ability, talent, and degree of self-belief. Similarly, propensity to save across households is often inversely related to their income volatility. FE estimation allows for possible correlation between $\boldsymbol{x}_{it}$ and ${\Greekmath 010B} _{i}$, but not between $ \boldsymbol{x}_{it}$ and $\boldsymbol{{\Greekmath 010C} }_{i}$. As a simple example, consider the panel data model
where $x_{it}$ and $y_{it}$ could, respectively, be years of schooling and return to schooling, or could be income and the saving rate of individual $i$ at time $t$. Suppose now that $x_{it}$ follows the following model with an interactive effect and idiosyncratic error heteroskedasticity:
where $f_{t}$ is a latent factor (could be talent in the context of return to education) and ${\Greekmath 011B} _{ix}u_{x,it}$ is the idiosyncratic innovation to the $x_{it}$ process (${\Greekmath 011B} _{ix}$ being income volatility in the saving rate example). A simple model of correlated slope heterogeneity is given by
where ${\Greekmath 0114} _{{\Greekmath 010D} }$ and ${\Greekmath 0114} _{{\Greekmath 011B} }$ measure the extent to which heterogeneity is correlated with $x_{it}$, and ${\Greekmath 010F} _{i}$ is the exogenous component of heterogeneity. FE estimators are robust to exogenous heterogeneity but can become badly biased if ${\Greekmath 0114} _{{\Greekmath 010D} }\neq 0$ and/or ${\Greekmath 0114} _{{\Greekmath 011B} }\neq 0$.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Bias of the fixed effects estimator under correlated heterogeneity}
It is known that fixed effects (with or without time effects) estimators are biased under correlated heterogeneity. For example, see Wooldridge2005 . Here we provide a minimal set of conditions required for the FE estimator to be $\sqrt{n}$-consistent in the presence of correlated heterogeneity. Under the heterogeneous specification ((ref)) and noting that $ \boldsymbol{M}_{T}\boldsymbol{{\Greekmath 011C} }_{T}=\boldsymbol{0}$, we have
Then by Assumption (ref), $E\left( \boldsymbol{u}_{i}\left\vert \boldsymbol{X}_{i}\right. \right) =\boldsymbol{0}$, and hence $E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{u}_{i}\right) = \boldsymbol{0}$. Under Assumptions (ref), (ref) and (ref),
Hence, $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ is a consistent estimator of the average effect, $\boldsymbol{{\Greekmath 010C} }_{0}$, only if
This condition is clearly met if
and has already been derived by Wooldridge2005. But it is too restrictive, since it is possible for the average condition in ((ref)) to hold even though condition ((ref)) is violated for some units as $n\rightarrow \infty $. Suppose the number of units that do not satisfy ((ref)) is given by $m_{n}=\ominus \left( n^{a_{{\Greekmath 0111} }}\right) $, namely $m_{n}$ rises in line with $n^{a_{{\Greekmath 0111} }}$. \footnote{ Note that $m_{n}=\ominus \left( n^{a_{{\Greekmath 0111} }}\right) $ differs from the familiar big O order, $m_{n}=O(n^{a_{{\Greekmath 0111} }})$. The latter requires $ n^{-a_{{\Greekmath 0111} }}m_{n}$ is bounded in $n$ and its limiting value could be zero. The former requires $n^{a_{{\Greekmath 0111} }}m_{n}$ to tend to a positive non-zero number.} Then $n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{X}_{i}^{\prime } \boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) =\ominus \left( n^{a_{{\Greekmath 0111} }-1}\right) $, and condition ((ref)) is met if $ a_{{\Greekmath 0111} }<1$.
But for $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ to be a $\sqrt{n}$-consistent estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$, a much more restrictive condition on $a_{{\Greekmath 0111} }$ is required. Using ((ref)),
where $\boldsymbol{{\Greekmath 0110} }_{iT}=\boldsymbol{X}_{i}^{\prime }\boldsymbol{M} _{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}$. Under Assumptions (ref) and (ref), as $n\rightarrow \infty $, $\boldsymbol{\bar{\Psi}} _{n}\rightarrow _{p}\boldsymbol{\bar{\Psi}}\succ \boldsymbol{0}$, and $ n^{-1/2}\sum_{i=1}^{n}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T} \boldsymbol{u}_{i}\rightarrow _{d}N\left( \boldsymbol{0},\boldsymbol{Q} _{FE}\right) $, where $\boldsymbol{Q}_{FE}=\lim\limits_{n\rightarrow \infty } \frac{1}{n}\sum_{i=1}^{n}E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M} _{T}\boldsymbol{H}_{i}\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) \succ \boldsymbol{0} $, for a fixed $T$. The first term in the brackets of ((ref)) can be decomposed as
where $\boldsymbol{s}_{nT}=n^{-1/2}\sum_{i=1}^{n}\left[ \boldsymbol{{\Greekmath 0110} } _{iT}-E\left( \boldsymbol{{\Greekmath 0110} }_{iT}\right) \right] $, and
which is bounded by Assumption (ref). It also follows that $ \boldsymbol{s}_{nT}$ tends to a limiting distribution with a zero mean by construction. Therefore, $\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}- \boldsymbol{{\Greekmath 010C} }_{0}\right) $ will converge to a distribution with a zero mean only if $n^{-1/2}\sum_{i=1}^{n}E\left( \boldsymbol{{\Greekmath 0110} }_{iT}\right) \rightarrow \boldsymbol{0}$. For this condition to hold, it is required that $m_{n}n^{-1/2}=\ominus \left( n^{a_{{\Greekmath 0111} }-1/2}\right) \rightarrow 0$, which occurs only if $a_{{\Greekmath 0111} }<1/2$. A formal statement of this result is summarized in the following proposition.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Mean group estimator}
In contrast to the FE estimator, the MG estimator given by ((ref)) continues to be $\sqrt{n}$-consistent, so long as certain moment conditions, to be discussed below, are met. Substituting ((ref)) in ((ref)),
where $\boldsymbol{{\Greekmath 0118} }_{iT}=\boldsymbol{R}_{i}^{\prime }\boldsymbol{u}_{i} $ and $\boldsymbol{R}_{i}=\boldsymbol{M}_{T}\boldsymbol{X}_{i}\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) ^{-1} $. Averaging both sides of ((ref)) over $i$ yields
where $\boldsymbol{\bar{{\Greekmath 010C}}}_{n}=\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{ {\Greekmath 010C} }_{i}$ and $\boldsymbol{\bar{{\Greekmath 0118}}}_{nT}=\frac{1}{n}\sum_{i=1}^{n} \boldsymbol{{\Greekmath 0118} }_{iT}$. By Assumption (ref), $E\left( \boldsymbol{ \bar{{\Greekmath 0118}}}_{nT}\right) =E\left( \frac{1}{n}\sum_{i=1}^{n}\boldsymbol{{\Greekmath 0118} } _{iT}\right) =\frac{1}{n}\sum_{i=1}^{n}E\left[ \boldsymbol{R}_{i}^{\prime }E\left( \boldsymbol{u}_{i}\left\vert \boldsymbol{X}_{i}\right. \right) \right] =\boldsymbol{0}$. Then using ((ref)), $E(\boldsymbol{\hat{{\Greekmath 010C}} }_{MG})=E(\boldsymbol{\bar{{\Greekmath 010C}}}_{n})+E\left( \boldsymbol{\bar{{\Greekmath 0118}}} _{nT}\right) =\boldsymbol{{\Greekmath 010C} }_{0}$, i.e., $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ is an unbiased estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ irrespective of the possible dependence of $\boldsymbol{{\Greekmath 010C} }_{i}$ on $\boldsymbol{X} _{i}$. However, the MG estimator is likely to have a large variance when $T$ is too small. This arises, for example, when the variance of $\boldsymbol{ \bar{{\Greekmath 0118}}}_{nT}$ does not exist or is very large. The conditions under which $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ converges to $\boldsymbol{{\Greekmath 010C} }_{0}$ at the regular $\sqrt{n}$ rate are given in the following proposition:
For a proof, see sub-section (ref) of the mathematical appendix.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Relative efficiency of FE and MG estimators}
Suppose now that conditions ((ref)) and ((ref)) hold and both FE and MG estimators are $\sqrt{n}$-consistent. The choice between the two estimators will then depend on their relative efficiency, which we measure in terms of their covariances conditional on $\boldsymbol{X}=\left( \boldsymbol{X}_{1},\boldsymbol{X}_{2},...,\boldsymbol{X}_{n}\right) $. We have
and
where $\boldsymbol{\Omega }_{{\Greekmath 010C} }=Var(\boldsymbol{{\Greekmath 010C} }_{i}|\boldsymbol{ X})\succeq \boldsymbol{0}$, $\boldsymbol{H}_{i}=E\left( \boldsymbol{u}_{i} \boldsymbol{u}_{i}^{\prime }|\boldsymbol{X}\right) $, and as before $ \boldsymbol{\Psi }_{i}=\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T} \boldsymbol{X}_{i}$ and $\boldsymbol{\bar{\Psi}}_{n}=n^{-1}\sum_{i=1}^{n} \boldsymbol{\Psi }_{i}$. Hence,
where
and
$\boldsymbol{A}_{n}$ and $\boldsymbol{B}_{n}$ capture the effects of two different types of heterogeneity, namely slope heterogeneity and regressors/errors heterogeneity. The superiority of the FE estimator over the MG estimator is readily established when the slope coefficients and error variances are homogeneous across $i$ and the errors are serially uncorrelated, namely if $\boldsymbol{\Omega }_{{\Greekmath 010C} }=\boldsymbol{0}$ and $ \boldsymbol{H}_{i}={\Greekmath 011B} ^{2}\boldsymbol{I}_{T}$ for all $i$. In this case, $\boldsymbol{A}_{n}=\boldsymbol{0}$, and we have
which is the difference between the harmonic mean of $\boldsymbol{\Psi }_{i}$ and the inverse of its arithmetic mean and is a positive semi-definite matrix.\footnote{ For a proof, see the Appendix to PesaranEtal1996.} However, this result may be reversed when we allow for heterogeneity, $\boldsymbol{\Omega } _{{\Greekmath 010C} }\succ \boldsymbol{0}$, and/or if $\boldsymbol{H}_{i}\neq {\Greekmath 011B} ^{2} \boldsymbol{I}_{T}$. The following proposition summarizes the results of the comparison between FE and MG estimators.
For a proof, see sub-section (ref) in the mathematical appendix.
In general, under uncorrelated heterogeneity, the relative efficiency of the MG and FE estimators depends on the relative magnitude of the two components in ((ref)). Since $\boldsymbol{A}_{n} \preceq \boldsymbol{0}$, the outcome depends on the sign and the magnitude of $\boldsymbol{B}_{n}$, which in turn depends on the heterogeneity of error variances, $\boldsymbol{H}_{i}( \boldsymbol{X}_{i}$), and $\boldsymbol{\Psi }_{i}$ over $i$.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Irregular mean group estimators}
So far, we have argued that the MG estimator is robust to correlated heterogeneity and its performance is comparable to the FE estimator even under uncorrelated heterogeneity. However, since the MG estimator is based on the individual estimates, $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$, $i=1,2,...,n$, its optimality and robustness critically depend on how well individual coefficients can be estimated. This is particularly important when $T$ is ultra short, which is the primary concern of this paper. In cases where $T$ is small and/or the observations on $\boldsymbol{x}_{it}$ are highly correlated or are slowly moving, $d_{i}=\func{det}\left( \boldsymbol{X} _{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) $ is likely to be close to zero for a large number of units $i=1,2,...,n$. As a result, $ \boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ is likely to be a poor estimate of $ \boldsymbol{{\Greekmath 010C} }_{i}$ for some $i$, and including such estimates when computing $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ could be problematic, rendering the MG estimator inefficient and unreliable.
However, as discussed above, $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ continues to be an unbiased estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$, even if $\boldsymbol{ {\Greekmath 010C} }_{i}$ are correlated with $\boldsymbol{X}_{i}$, so long as the stochastic component of $\boldsymbol{x}_{it}$ is strictly exogenous with respect to $u_{it}$. By averaging over $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ for $ i=1,2,...,n$, as $n\rightarrow \infty $, the MG estimator converges to $ \boldsymbol{{\Greekmath 010C} }_{0}$ if $T$ is sufficiently large such that $\boldsymbol{ \hat{{\Greekmath 010C}}}_{i}$ have at least second-order moments for all $i$. The existence of first-order moments of $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ is required for the MG estimator to be unbiased, and we need $\boldsymbol{\hat{ {\Greekmath 010C}}}_{i}$ to have second-order moments for $\sqrt{n}$-consistent estimation and valid inference about the average effects, $\boldsymbol{{\Greekmath 010C} }_{0}$.
When individual estimates do not have second-order moments, trimming is required. The question is how to trim the individual estimates $\boldsymbol{ \hat{{\Greekmath 010C}}}_{i}$. Peng2001 proposes to categorize the individual estimates and then use a weighted average of the estimates depending on their left and right tail shape parameters. He assumes the estimates are identically and independently distributed and considers only the case of a scalar parameter. GrahamPowell2012 propose to trim by exclusion, namely dropping individual estimates with $d_{i}$ below a threshold value. We establish that trimming and/or shrinkage is required only if the tail index, ${\Greekmath 010B} _{p}$, of the distribution of $1/d_{i}$ is below or equal to $ 2$, and propose to shrink only those estimates whose $d_{i}$ are below a threshold value, $a_{n}$, when ${\Greekmath 010B} _{p}\leq 2$. Moreover, it is shown by simulations that estimates of ${\Greekmath 010B} _{p}$ rise with $T$, and shrinkage is typically required when $T<2k+1$.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{A trimmed mean group estimator }
To formalize our proposed TMG estimator, we start with the least square estimator of $\boldsymbol{{\Greekmath 010C} }_{i}$, namely $\boldsymbol{\hat{{\Greekmath 010C}}} _{i}=(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})^{-1} \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{y}_{i}$, and shrink it using the threshold function $\boldsymbol{1}\{d_{i}>a_{n}\}$, which takes the value of unity if $d_{i}>a_{n}$ and zero otherwise, where $ d_{i}=\func{det}(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X} _{i})$. The threshold value is set as
where ${\Greekmath 010B} >0$ and $C_{n}$ is a positive constant bounded in $n$. The choice of ${\Greekmath 010B} $ and $C_{n}$ will be discussed below. The resultant shrinkage estimator is given by
where $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}^{\ast }$ is the shrinkage version of $ \boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ since $a_{n}^{-1}\leq d_{i}^{-1}$. Written more compactly, we have
where
We considered two versions of TMG estimators, depending on how individual trimmed estimators, $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}$, are combined. An obvious choice is to use a simple average of $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}$ , namely $\overline{\boldsymbol{\tilde{{\Greekmath 010C}}}}_{n}=n^{-1}\sum_{i=1}^{n} \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}=n^{-1}\sum_{i=1}^{n}(1+{\Greekmath 010E} _{i}) \boldsymbol{\hat{{\Greekmath 010C}}}_{i}$, which can also be viewed as a weighted average estimator with the weights $w_{i}=(1+{\Greekmath 010E} _{i})/n<1/n$. But it is easily seen that these weights do not add up to unity, and it might be desirable to use the scaled weights $w_{i}/(1+\bar{{\Greekmath 010E}} _{n})=n^{-1}(1+{\Greekmath 010E} _{i})/(1+\bar{{\Greekmath 010E}}_{n})$, where $\bar{{\Greekmath 010E}} _{n}=n^{-1}\sum_{i=1}^{n}{\Greekmath 010E} _{i}$. Using these modified weights, we propose the following TMG estimator
which can be written more compactly as
where
The TMG estimator differs in two important respects from the trimmed estimator proposed by GrahamPowell2012, which in the context of our setup (and abstracting from time effects for now) can be written as
where $d_{i,GP}=\det (\boldsymbol{W}_{i}^{\prime }\boldsymbol{W}_{i})$, and $ \boldsymbol{W}_{i}=\left( \boldsymbol{{\Greekmath 011C} }_{T},\boldsymbol{X}_{i}\right) $ . In the special case where $T=k$, $d_{i,GP}=\left\vert \func{det}\left( \boldsymbol{W}_{i}\right) \right\vert ^{2}$ and the threshold function considered by GP reduces to $\boldsymbol{1}\{\left\vert \func{det}\left( \boldsymbol{W}_{i}\right) \right\vert >h_{n}\}$.
The GP approach can be viewed as trimming by exclusion, which overlooks the information that might be contained in $\func{adj}(\boldsymbol{W} _{i}^{\prime }\boldsymbol{W}_{i})$ when $\left\vert \func{det}\left( \boldsymbol{W}_{i}\right) \right\vert \leq h_{n}$. More specifically, to relate $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ to the GP estimator given by ((ref)), using ((ref)) in ((ref)), we note that
where ${\Greekmath 0119} _{n}$ is the fraction of the estimates being trimmed given by
Compared to $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$, the GP estimator places zero weights on the estimates with $d_{i}\leq a_{n}$, and leaves out the scaling factor, $\left( 1+\bar{{\Greekmath 010E}}_{n}\right) ^{-1}$, which, as already noted, can play an important role in the small sample performance of the TMG estimator. Our proposed method also differs from GP in the way we motivate and calibrate the threshold function.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Asymptotic properties of the TMG estimator}
To investigate the asymptotic properties of the TMG estimator, $\boldsymbol{ \hat{{\Greekmath 010C}}}_{TMG}$, we introduce the following additional assumptions:
Using ((ref)) and ((ref)) in ((ref)), we have
and $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ defined by ((ref)) can be written as
((ref)) can be written equivalently as
where
Also, by Lemma (ref), $E({\Greekmath 010E} _{i})=O(a_{n}^{{\Greekmath 010B} _{p}})$, $ E\left( \bar{{\Greekmath 010E}}_{n}\right) =O(a_{n}^{{\Greekmath 010B} _{p}})$, $E\left( {\Greekmath 010E} _{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) =O(a_{n}^{{\Greekmath 010B} _{p}})$, and hence
Since under Assumptions (ref) and (ref), ${\Greekmath 010E} _{i}-E\left( {\Greekmath 010E} _{i}\right) $ is distributed independently over $i$ with a zero mean and bounded variance, then
Similarly, under Assumptions (ref), (ref) and (ref), $\boldsymbol{p}_{i}-E\left( \boldsymbol{p}_{i}\right) $ is distributed independently over $i$ with zero means and bounded variances, and we have
Consider now $\boldsymbol{\bar{q}}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{q} _{iT}$, where $\boldsymbol{q}_{iT}$ is defined by ((ref)), and note that
where $\boldsymbol{\bar{{\Greekmath 0118}}}_{{\Greekmath 010E} ,nT}=n^{-1}\sum_{i=1}^{n}\left( 1+{\Greekmath 010E} _{i}\right) \boldsymbol{{\Greekmath 0118} }_{iT}$. Using results of Lemma (ref), we have $E\left( \boldsymbol{\bar{q}}_{nT}\right) =\boldsymbol{0}$ , and for any ${\Greekmath 010B} _{p}>0$,
Hence, $\boldsymbol{\bar{q}}_{nT}=O_{p}\left( n^{-1/2+{\Greekmath 010B} /2}\right) $. Using the above results in ((ref)), we have
Thus, $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ converges to $\boldsymbol{{\Greekmath 010C} }_{0}$ asymptotically so long as $0<{\Greekmath 010B} <1$ as $n\rightarrow \infty $. The convergence rate of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ to $\boldsymbol{{\Greekmath 010C} } _{0}$ will depend on the trade-off between the bias and variance of $ \boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$. Though it is possible to reduce the bias of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ by choosing a value of ${\Greekmath 010B} $ close to unity, it will be at the expense of a larger variance. In what follows, we shed light on the choice of ${\Greekmath 010B} $ by considering the conditions under which the asymptotic distribution of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is centered around $\boldsymbol{{\Greekmath 010C} }_{0}$ and at the same time, its asymptotic variance tends to zero at a reasonably fast rate.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{The choice of the trimming threshold value }
We begin by assuming that the rate at which $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ converges to $\boldsymbol{{\Greekmath 010C} }_{0}$ is given by $n^{{\Greekmath 010D} }$, where $ {\Greekmath 010D} $ is set in relation to ${\Greekmath 010B} $. Since it is not guaranteed that the individual estimates, $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$, have second order moments when when $T$ is ultra short, we expect the rate, $n^{{\Greekmath 010D} }$, to be below the regular rate of $n^{1/2}$. Using ((ref)) and ((ref)) and noting that ${\Greekmath 010D} \leq 1/2$ (with equality holding only under regular convergence), we have
To ensure that the asymptotic distribution of $\boldsymbol{\hat{{\Greekmath 010C}}} _{TMG} $ is correctly centered, we must have $n^{{\Greekmath 010D} }\boldsymbol{b} _{n}\rightarrow \boldsymbol{0,}$ as $n\rightarrow \infty $. Since $n^{{\Greekmath 010D} }\boldsymbol{b}_{n}=O(n^{{\Greekmath 010D} }a_{n}^{{\Greekmath 010B} _{p}})=O(n^{{\Greekmath 010D} -{\Greekmath 010B} {\Greekmath 010B} _{p}})$, this condition is ensured if ${\Greekmath 010D} <{\Greekmath 010B} {\Greekmath 010B} _{p}$. Turning to the second term of the above, we note that to obtain a non-degenerate distribution, we also need to set ${\Greekmath 010D} =\left( 1-{\Greekmath 010B} \right) /2$. Combining these two requirements yields $\left( 1-{\Greekmath 010B} \right) /2<{\Greekmath 010B} {\Greekmath 010B} _{p}$, or
In view of ((ref)), the convergence rate of $\boldsymbol{\hat{{\Greekmath 010C}}} _{TMG}-\boldsymbol{{\Greekmath 010C} }_{0}$, namely $n^{-\frac{(1-{\Greekmath 010B} )}{2}}$, depends on ${\Greekmath 010B} _{p}$, which governs the tail property of the distribution of $1/d_{i}$. In the simple example (ref), assuming ${\Greekmath 011B} _{ix}^{2}={\Greekmath 011B} _{x}^{2}$, then $1/d_{i}=(1/{\Greekmath 011B} _{x}^{2})(1/{\Greekmath 011F} _{v}^{2})$, and the shape parameter of $1/d_{i}$ will equal the shape parameter of $1/{\Greekmath 011F} _{v}^{2}$, which is given by ${\Greekmath 010B} _{p}=v/2$. In the simple case where $k^{\prime }=1$, we have ${\Greekmath 010B}_{p}=(T-1)/2$. In the worse case scenario with $T=2=k$, ${\Greekmath 010B} _{p}=1/2$, and the convergence rate of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is at most be $n^{-1/4}$, well below the regular convergence rate of $n^{-1/2}$. The optimum choice of $ {\Greekmath 010B} $ depends on ${\Greekmath 010B} _{p}$, which in turn is determined by the rate at which $F_{d}(a_{n})$ tends to zero as $a_{n}\rightarrow 0$.
To ensure that the threshold function, $\boldsymbol{1}\{d_{i}>C_{n}n^{- {\Greekmath 010B} }\}$, does not depend on the scale of $\boldsymbol{x}_{it}$, we suggest setting $C_{n}=\bar{d}_{n}=n^{-1}\sum_{i=1}^{n}d_{i}$. With this choice, the threshold function can be written as $\boldsymbol{1} \{d_{i}>C_{n}n^{-{\Greekmath 010B} }\}=\boldsymbol{1}\{d_{i}/\bar{d}_{n}>n^{-{\Greekmath 010B} }\}$ , noting that $d_{i}/\bar{d}_{n}$ is scale free. The selection of ${\Greekmath 010B} $ is more complicated. In practice, a two-step procedure can be implemented, whereby in the first step the value of ${\Greekmath 010B} _{p}$ is estimated using the observations $\left\{ 1/d_{i}\right\} _{i=1}^{n}$.\footnote{ Asymptotically, the estimate of ${\Greekmath 010B} _{p}$ does not depend on the scale of $d_{i}$.} Then TMG estimation can be carried out using $\hat{{\Greekmath 010B}} =1/(1+2\hat{{\Greekmath 010B}}_{p})+{\Greekmath 010F} $, where $\hat{{\Greekmath 010B}}_{p}$ is a consistent estimator of ${\Greekmath 010B} _{p}$, and ${\Greekmath 010F} $ is a small positive constant. See also Remark (ref).
Given the uncertainty associated with the estimates of ${\Greekmath 010B} _{p}$, we also considered setting ${\Greekmath 010B} _{p}=1$, and using the threshold function $ \boldsymbol{1}\{d_{i}/\bar{d}_{n}>n^{-1/3}\}$. This approach has the advantage of being simple to implement. The small sample performance of these two approaches to setting ${\Greekmath 010B} $ are compared using MC experiments. See sub-section (ref) of the online supplement. The simulation results show that the simple threshold rule $\boldsymbol{1}\{d_{i}/\bar{d} _{n}>n^{-1/3}\}$ performs better when $T$ is ultra-short ($T=k$), and has a similar performance when the estimated value of $\hat{{\Greekmath 010B}}_{p}$ is used if $T>k$.
GP show that to correctly center the limiting distribution of their proposed estimator, $\boldsymbol{\hat{{\Greekmath 010C}}}_{GP}$ given by ((ref)), $h_{n}$ in their threshold function $\boldsymbol{1}\{|\det (\boldsymbol{W} _{i})|>h_{n}\} $ must be set as $h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}$, such that $ (nh_{n})^{1/2}h_{n}\rightarrow 0$, as $n\rightarrow \infty $, which implies that ${\Greekmath 010B} _{GP}>1/3$ (see p. 2125 and p. 2138 of GP). SU adopt the same choice and set $h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}n^{-L/(1+2L)}$, with $L\in \{1,2\}$ being the order of local polynomials used to infer the average effects of stayers (see sub-section 5.2 on p. 16 of SU).
For a proof, see sub-section (ref) of the mathematical appendix.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Robust estimation of the covariance matrix of the TMG estimator}
As with standard MG estimation, consistent estimation of $\boldsymbol{V} _{{\Greekmath 010C} }$ using ((ref)) requires knowledge of $\boldsymbol{H}_{i}$ which cannot be estimated consistently when $T$ is short. We follow the literature and propose a robust covariance estimator of $\boldsymbol{V} _{{\Greekmath 010C} }$, which is asymptotically unbiased for a wide class of error variances, $E\left( \boldsymbol{u}_{i}\boldsymbol{u}_{i}^{\prime }\left\vert \boldsymbol{X}_{i}\right. \right) =\boldsymbol{H}_{i}(\boldsymbol{X}_{i})$, thus allowing for serially correlated and conditionally heteroskedastic errors. The following theorem summarizes the main result.
For a proof, see sub-section (ref) of the mathematical appendix.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Panels with time effects}
The panel data model with time effects is given by
where ${\Greekmath 011E} _{t}$ for $t=1,2,...,T$ are the time effects. Without loss of generality, we adopt the normalization $\boldsymbol{{\Greekmath 011C} }_{T}^{\prime } \boldsymbol{{\Greekmath 011E} }=0$, where $\boldsymbol{{\Greekmath 011E} }=({\Greekmath 011E} _{1},{\Greekmath 011E} _{2},...,{\Greekmath 011E} _{T})^{\prime }$. To estimate $\boldsymbol{{\Greekmath 010C} }_{0}$, initially we suppose $\boldsymbol{{\Greekmath 011E} }$ is known, and the trimmed estimator of $\boldsymbol{{\Greekmath 010C} }_{i}$ is now given by $\boldsymbol{\tilde{ {\Greekmath 010C}}}_{i}(\boldsymbol{{\Greekmath 011E} })=\boldsymbol{Q}_{i}^{\prime }(\boldsymbol{y} _{i}-\boldsymbol{{\Greekmath 011E} })=\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{Q} _{i}^{\prime }\boldsymbol{{\Greekmath 011E} }$, where $\boldsymbol{Q}_{i}$ is defined by ( (ref)). The associated TMG-TE estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ then follows as
where $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is given by ((ref)), and
From our earlier analysis, it is clear that for a known $\boldsymbol{{\Greekmath 011E} }$ , $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG-TE}(\boldsymbol{{\Greekmath 011E} })$ has the same asymptotic distribution as $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ with $ \boldsymbol{y}_{i}$ replaced by $\boldsymbol{y}_{i}-\boldsymbol{{\Greekmath 011E} }$. When $T>k$, we can follow Chamberlain1992 and eliminate the time effects by the de-meaning transformation $\boldsymbol{M}_{i}=\boldsymbol{I} _{T}-\boldsymbol{M}_{T}\boldsymbol{X}_{i}(\boldsymbol{X}_{i}^{\prime } \boldsymbol{M}_{T}\boldsymbol{X}_{i})^{-1}\boldsymbol{X}_{i}^{\prime } \boldsymbol{M}_{T}$. Under the normalization $\boldsymbol{{\Greekmath 011C} }_{T}^{\prime }\boldsymbol{{\Greekmath 011E} }=0$, we have $\boldsymbol{M}_{T}\boldsymbol{{\Greekmath 011E} }= \boldsymbol{{\Greekmath 011E} }$, and$\ \boldsymbol{M}_{T}\boldsymbol{y}_{i}=\boldsymbol{M }_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 010C} }_{i}+\boldsymbol{{\Greekmath 011E} }+ \boldsymbol{M}_{T}\boldsymbol{u}_{i}$. Then $\boldsymbol{M}_{i}\boldsymbol{M} _{T}\boldsymbol{y}_{i}=\boldsymbol{M}_{i}\boldsymbol{{\Greekmath 011E} }+\boldsymbol{M} _{i}\boldsymbol{M}_{T}\boldsymbol{u}_{i}$, and averaging over $i$ we obtain
Hence, $\boldsymbol{{\Greekmath 011E} }$ can be estimated without knowing $\boldsymbol{ {\Greekmath 010C} }_{0}$, if $\boldsymbol{\bar{M}}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{M} _{i}$ is a positive definite matrix. This requires $T>k$, since $\boldsymbol{ \bar{M}}_{n}$ is singular if $T=k$. Therefore, to implement the Chamberlain's estimation approach, we require the following assumption:
Under this Assumption, $\boldsymbol{{\Greekmath 011E} }$ can be estimated by
and its asymptotic distribution follows straightforwardly. Specifically, using ((ref)) we have
and $\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 011E}}}_{C}-\boldsymbol{{\Greekmath 011E} } _{0}\right) \rightarrow _{d}N(\boldsymbol{0},\boldsymbol{V}_{{\Greekmath 011E} ,C})$, where $\boldsymbol{V}_{{\Greekmath 011E} ,C}=Avar\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 011E}}} _{C}\right) $ is given by
Since $\boldsymbol{M}_{i}\boldsymbol{M}_{T}\boldsymbol{u}_{i}=\boldsymbol{M} _{i}\boldsymbol{M}_{T}(\boldsymbol{y}_{i}-\boldsymbol{{\Greekmath 011E} })$, the asymptotic variance of $\boldsymbol{\hat{{\Greekmath 011E}}}_{C}$ can be consistently estimated by
Using $\boldsymbol{\hat{{\Greekmath 011E}}}_{C}$, the TMG-TE estimator of $\boldsymbol{ {\Greekmath 010C} }_{0}$ is now given by
and
where $\boldsymbol{\bar{Q}}_{n}^{\prime }\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 011E} }}_{C}-\boldsymbol{{\Greekmath 011E} }\right) =O_{p}(1)$. Hence
Also since ${\Greekmath 010B} >0$, then $n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}} }_{C,TMG-TE}-\boldsymbol{{\Greekmath 010C} }_{0}\right) $ has the same asymptotic distribution as $n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}( \boldsymbol{{\Greekmath 011E} })-\boldsymbol{{\Greekmath 010C} }_{0}\right) $, with $\boldsymbol{{\Greekmath 011E} }$ treated as known. This result follows since $\boldsymbol{\hat{{\Greekmath 011E}}}_{C}- \boldsymbol{{\Greekmath 011E} }\rightarrow _{p}\boldsymbol{0}$ at the faster rate of $ \sqrt{n}$, compared with the rate $n^{\frac{1-{\Greekmath 010B} }{2}}$ of $\boldsymbol{ \hat{{\Greekmath 010C}}}_{C,TMG-TE}-\boldsymbol{{\Greekmath 010C} }_{0}\rightarrow _{p}\boldsymbol{0} $. In short, $Avar\left( n^{(1-{\Greekmath 010B} )/2}\boldsymbol{\hat{{\Greekmath 010C}}} _{C,TMG-TE}\right) $ is not affected by the estimation uncertainty of the time effects when ${\Greekmath 010B} >0$. A consistent estimator is given by
where $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i,C}=\boldsymbol{Q}_{i}^{\prime }( \boldsymbol{y}_{i}-\boldsymbol{\hat{{\Greekmath 011E}}}_{C})$.
When $T=k$, Assumption (ref) does not hold and standard de-meaning techniques can not be used to eliminate $\boldsymbol{{\Greekmath 011E} }$. This in turn requires the correlation of slope coefficients with the regressors to be time-invariant, a condition that holds trivially under uncorrelated heterogeneity but could be restrictive when heterogeneity is correlated. The TMG-TE estimator for the case of $T= k$ and its asymptotic properties are derived in Section (ref) of the mathematical appendix.
For the Monte Carlo experiments and the empirical application, $\boldsymbol{ \hat{{\Greekmath 011E}}}_{C}$, given by ((ref)), is used as an estimator of $ \boldsymbol{{\Greekmath 011E} }$ for panels with time effects when $T>k$. When $T=k$, we use the estimator set out in Section (ref) of the mathematical appendix.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{A test of correlated heterogeneity}
As summarized by Proposition (ref), $\sqrt{n}$-consistency of the FE estimator requires the slope coefficients, $\boldsymbol{{\Greekmath 010C} }_{i}$, and the regressors, $\boldsymbol{X}_{i}=\left( \boldsymbol{x}_{i1}, \boldsymbol{x}_{i2},...,\boldsymbol{x}_{iT}\right) ^{\prime }$, to be independently distributed. As a diagnostic check, we consider testing the null hypothesis
where $\boldsymbol{{\Greekmath 0111} }_{i}=\boldsymbol{{\Greekmath 010C} }_{i}-\boldsymbol{{\Greekmath 010C} } _{0} $. To test $H_{0}$, we follow Hausman1978 and compare FE and TMG estimators given by ((ref)) and ((ref)), respectively. Under $H_{0}$ , both estimators are consistent. But, as required by Hausman tests, only the TMG estimator is consistent under the alternative
The Hausman test of $H_{0}$ is based on the difference $\sqrt{n}\boldsymbol{ \hat{\Delta}}_{{\Greekmath 010C} }=\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}- \boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}\right) $. Such a test has been considered by PesaranEtal1996 and PesaranYamagata2008, assuming the MG estimator has at least second-order moments.\footnote{ See pp. 160--162 of PesaranEtal1996 and p. 53 of PesaranYamagata2008.} Here we extend this test to cover cases when $T$ is ultra short and the moment condition ((ref)) is not met. Also, the earlier tests were derived under the null of homogeneity (namely $ \boldsymbol{{\Greekmath 0111} }_{i}=\boldsymbol{0}$ for all $i$), whilst the null that we consider is more general and covers the null of homogeneity as a special case.
In the development of his test, Hausman originally assumed that one of the estimators under consideration is asymptotically efficient and showed that in this case, the asymptotic variance of the difference between the two estimators is equal to the difference between the two variances. But in the present application, as shown in Proposition (ref) and illustrated by Example (ref), neither of the two estimators, $ \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ and $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$, are efficient, and the asymptotic variance of $\sqrt{n}\boldsymbol{\hat{\Delta}} _{{\Greekmath 010C} }$ will not be equal to the difference between $Avar\left( \sqrt{n} \boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}\right) $ and $Avar\left( \sqrt{n}\boldsymbol{ \hat{{\Greekmath 010C}}}_{FE}\right) $. See also Section 26.9.1 of Pesaran2015.
The following theorem summarizes our main result for the Hausman test applied to $\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{\hat{ {\Greekmath 010C}}}_{TMG}\right) $.
For a proof, see sub-section (ref) in the mathematical appendix.
To implement the $H_{{\Greekmath 010C} }$ test, the asymptotic variance, $\boldsymbol{V} _{\Delta }$, can be rewritten as $\boldsymbol{V}_{\Delta }=n^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}E\left( \boldsymbol{g}_{it}\boldsymbol{g}_{it^{\prime }}^{\prime }\tilde{{\Greekmath 0117}}_{it} \tilde{{\Greekmath 0117}}_{it^{\prime }}\right) $, where $\tilde{{\Greekmath 0117}}_{it}={\Greekmath 0117} _{it}-\bar{ {\Greekmath 0117}}_{i\circ }$, $\bar{{\Greekmath 0117}}_{i\circ } = \sum_{t=1}^{T} \sum_{t^{\prime }=1}^{T}{\Greekmath 0117}_{it}$, and $\boldsymbol{g}_{it}$ is the $t^{th}$ column of $ \boldsymbol{G}_{i}$, where (see also ((ref)) in the mathematical appendix)
For $T$ fixed, a consistent estimator of $\boldsymbol{V}_{\Delta }$, which is robust to the choices of $\boldsymbol{H}_{i}$ and $\boldsymbol{\Omega } _{{\Greekmath 010C} }$, is
where $\hat{\tilde{{\Greekmath 0117}}}_{it}=(y_{it}-\bar{y}_{i\circ })-\boldsymbol{\hat{ {\Greekmath 010C}}}_{FE}^{\prime }(\boldsymbol{x}_{it}-\boldsymbol{\bar{x}}_{i\circ })$, $\bar{y}_{i\circ }=\frac{1}{T}\sum_{t=1}^{T}y_{it}$, and $\boldsymbol{\bar{x} }_{i\circ }=\frac{1}{T}\sum_{t=1}^{T}\boldsymbol{x}_{it}$. Using $ \boldsymbol{\widehat{V}}_{\Delta }$, the Hausman test statistic for uncorrelated slope heterogeneity is given by
Under $H_{0}$, the consistency of $\boldsymbol{\widehat{V}}_{\Delta }$ as an estimator of $\boldsymbol{V}_{\Delta }$ follows from the $\sqrt{n}$ -consistency of the FE estimator.
An extension of the $\hat{H}_{{\Greekmath 010C} }$ test statistic to panel data models with time effects is provided in Section (ref) of the online supplement.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Trimmed mean group estimation for a subset of coefficients}
The TMG approach can be applied to a subset of coefficients of interest. Let $\boldsymbol{X}_{i}=(\boldsymbol{X}_{i1},\boldsymbol{X}_{i2})$, where $ \boldsymbol{X}_{i1}$ is the $T\times p$ matrix of observations on the focal (treatment) variables and $\boldsymbol{X}_{i2}$ is the $T\times (k^{\prime }-p)$ matrix of observations on the auxiliary (control) variables, and partition the coefficients accordingly as $\boldsymbol{{\Greekmath 010C} }_{i}=( \boldsymbol{{\Greekmath 010C} }_{i1}^{\prime },\boldsymbol{{\Greekmath 010C} }_{i2}^{\prime })^{\prime }$, where $\boldsymbol{{\Greekmath 010C} }_{i1}$ is the $p\times 1$ vector of the coefficients of interest. Using results from the partitioned regressions, we have
where $\boldsymbol{M}_{i2}=\boldsymbol{I}_{T}-\boldsymbol{X}_{i2}\left( \boldsymbol{X}_{i2}^{\prime }\boldsymbol{X}_{i2}\right) ^{-}\boldsymbol{X} _{i2}^{\prime }$, and $\left( \boldsymbol{X}_{i2}^{\prime }\boldsymbol{X} _{i2}\right) ^{-}$ denotes a generalized inverse of $\left( \boldsymbol{X} _{i2}^{\prime }\boldsymbol{X}_{i2}\right) $.\footnote{ Note that $\boldsymbol{\hat{{\Greekmath 010C}}}_{i1}$ is invariant to the choice of the generalized inverse, $\left( \boldsymbol{X}_{i2}^{\prime }\boldsymbol{X} _{i2}\right) ^{-}$.}
The TMG estimator of $\boldsymbol{{\Greekmath 010C} }_{01}=E\left( \boldsymbol{{\Greekmath 010C} } _{i1}\right) $ is given by
where ${\Greekmath 010E} _{i1}=\left( \frac{d_{i1}-a_{n1}}{a_{n1}}\right) \boldsymbol{1} \{d_{i1}\leq a_{n1}\}$, $\bar{{\Greekmath 010E}}_{n1}=\frac{1}{n}\sum_{i=1}^{n}{\Greekmath 010E} _{i1}$, and $d_{i1}=\det \left( \boldsymbol{X}_{i1}^{\prime }\boldsymbol{M} _{i2}\boldsymbol{X}_{i1}\right) $. Similarly, for the choice of threshold value, we consider $a_{n1}=\bar{d}_{n1}n^{-{\Greekmath 010B} _{1}}$, where ${\Greekmath 010B} _{1}> \frac{1}{1+2{\Greekmath 010B} _{1,p}}$, ${\Greekmath 010B} _{1,p}$ is the shape parameter of the tail probability distribution of $1/d_{i1}$ over $i$, and $\bar{d} _{n1}=n^{-1}\sum_{i=1}^{n}d_{i1}$.
The above partitioned formula can also be adapted for testing general linear restrictions $H_{0}:\boldsymbol{R{\Greekmath 010C} }_{0}=\boldsymbol{r}$ against $H_{1}: \boldsymbol{R{\Greekmath 010C} }_{0}\neq \boldsymbol{r}$, where $\boldsymbol{r}$ is a $ p\times 1$ vector and $\boldsymbol{R}$ is a $p\times k^{\prime }$ ($ p<k^{\prime }$) full rank matrix of fixed constants. Let $\boldsymbol{{\Greekmath 010D} }_{i}=$ $\boldsymbol{R{\Greekmath 010C} }_{i}-\boldsymbol{r}$ and partition $\boldsymbol{ R=(R}_{1},\boldsymbol{R}_{2})$ such that $\boldsymbol{R}_{1}$ is a $p\times p $ non-singular matrix.\footnote{ This can be achieved by a suitable reordering of the elements of $ \boldsymbol{{\Greekmath 010C} }_{i}$.} Then ((ref)) can be written as
where $\boldsymbol{\tilde{y}}_{i}=\boldsymbol{y}_{i}-\boldsymbol{X}_{i1} \boldsymbol{R}_{1}^{-1}\boldsymbol{r}$, $\boldsymbol{\tilde{X}}_{i1}= \boldsymbol{X}_{i1}\boldsymbol{R}_{1}^{-1}$, and $\boldsymbol{\tilde{X}} _{i2}=\boldsymbol{X}_{i2}-\boldsymbol{X}_{i1}\boldsymbol{R}_{1}^{-1} \boldsymbol{R}_{2}$. The TMG estimator of $\boldsymbol{{\Greekmath 010D} }_{0}=E\left( \boldsymbol{{\Greekmath 010D} }_{i}\right) $ can be computed using the unit-specific estimates of $\boldsymbol{{\Greekmath 010D} }_{i}$ as above.
The Hausman test proposed in the previous section can also be applied to a subset of the coefficients in a straightforward manner.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Monte Carlo experiments}
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Data generating processes (DGP) }
The outcome variable, $y_{it}$, is generated as
where the errors, $u_{it}$, are allowed to be serially correlated and heteroskedastic. Specifically, we set $u_{it}={\Greekmath 0114} {\Greekmath 011B} _{it}e_{it}$ and generate $e_{it}$ as AR(1) processes
In the baseline DGP, we set ${\Greekmath 011B} _{it}={\Greekmath 011B} _{iu}$ for all $t$ and generate ${\Greekmath 011B} _{iu}^{2}\sim IID\frac{1}{2}\left( 1+z_{iu}^{2}\right)$, with $z_{iu}\sim IIDN(0,1)$.\footnote{ More general heteroskedastic specifications for ${\Greekmath 011B} _{it}$ are considered in sub-section (ref) of the online supplement.} Both Gaussian and non-Gaussian errors are considered: ${\Greekmath 0126} _{it}\sim IIDN(0,1)$ and ${\Greekmath 0126} _{it}\sim IID\frac{1}{2}\left( {\Greekmath 011F} _{2}^{2}-2\right) $.
The regressors, $x_{j,it}$, for $j=1,2,....k^{\prime }$, $i=1,2,...,n$, and $ t=1,2,...,T$, are generated as factor-augmented AR(1) processes
where ${\Greekmath 010B} _{j,ix}\sim IIDN(1,1)$, $u_{xj,it}={\Greekmath 011B} _{j,ix}e_{xj,it}$, $ e_{xj,it}\sim IID(0,1)$, ${\Greekmath 011B} _{j,ix}^{2}=\frac{1}{2}\left( 1+z_{j,ix}^{2}\right) $, and $z_{j,ix}\sim IIDN(0,1)$. We consider panels with $k^{\prime }=1,2$ and $3$ regressors and investigate the extent to which trimming is required for different combinations of $T$ and $k^{\prime } $. The common factors are generated as $ f_{j,t}=0.9f_{j,t-1}+(1-0.9^{2})^{1/2}v_{j,t}$, for $ t=-49,-48,...,-1,0,1,...,T$, where $v_{j,t}\sim IIDN(0,1)$, and $f_{j,-50}=0 $. The factor loadings are generated as ${\Greekmath 010D} _{j,ix}\sim IIDU(0,2)$. When time effects are included in the model, we set ${\Greekmath 011E} _{t}=t$, for $ t=1,2,...,T-1$, and ${\Greekmath 011E} _{T}=-T(T-1)/2$, so that $\boldsymbol{{\Greekmath 011C} } _{T}^{\prime }\boldsymbol{{\Greekmath 011E} }=0$.
For the slope coefficients, we experiment with both correlated and uncorrelated effects in ${\Greekmath 010C} _{i1}$, while considering uncorrelated heterogeneous effects for the other slope coefficients when the respective regressors are included. Specifically, ${\Greekmath 010B} _{i}$ and ${\Greekmath 010C} _{i1}$ are generated as
where
$\boldsymbol{{\Greekmath 0120} =(}{\Greekmath 0120} _{{\Greekmath 010B} },{\Greekmath 0120} _{{\Greekmath 010C} _{1}})^{\prime }$, $ \boldsymbol{{\Greekmath 010F} }_{i}=({\Greekmath 010F} _{i{\Greekmath 010B} },{\Greekmath 010F} _{i{\Greekmath 010C} _{1}})^{\prime }\sim IIDN\left( \boldsymbol{0},\boldsymbol{V}_{{\Greekmath 010F} }\right) $, and $\boldsymbol{V}_{{\Greekmath 010F} }=Diag(\boldsymbol{{\Greekmath 011B} } _{{\Greekmath 010F} }^{2})$ with $\boldsymbol{{\Greekmath 011B} }_{{\Greekmath 010F} }^{2}=\left( {\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010B} }^{2},{\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010C} _{1}}^{2}\right) ^{\prime }$ . It follows that
The degree of correlated heterogeneity is determined by $\boldsymbol{{\Greekmath 0120} {\Greekmath 0120} }^{\prime }$, with $Cov({\Greekmath 010B} _{i},{\Greekmath 010C} _{i1})={\Greekmath 011B} _{{\Greekmath 010B} {\Greekmath 010C} _{1}}\neq 0$ when ${\Greekmath 0120} _{{\Greekmath 010B} }$ and ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ are both non-zero. Specifically, ${\Greekmath 011B} _{{\Greekmath 010B} }^{2}={\Greekmath 0120} _{{\Greekmath 010B} }^{2}+{\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010B} }^{2}$, ${\Greekmath 011B} _{{\Greekmath 010B} {\Greekmath 010C} _{1}}={\Greekmath 0120} _{{\Greekmath 010B} }{\Greekmath 0120} _{{\Greekmath 010C} _{1}}$, and ${\Greekmath 011B} _{{\Greekmath 010C} _{1}}^{2}={\Greekmath 0120} _{{\Greekmath 010C} _{1}}^{2}+{\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010C} _{1}}^{2}$. The coefficients of $x_{j,it}$, for $ j=2,3,...,k^{\prime }$ are generated as ${\Greekmath 010C} _{ij}={\Greekmath 010C} _{0j}+{\Greekmath 010F} _{i{\Greekmath 010C} _{j}}$ with ${\Greekmath 010F} _{i{\Greekmath 010C} _{j}}\sim IIDN(0,{\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010C} _{j}}^{2})$, which are not correlated with any of the regressors. The true values of the parameters of interest are set as follows: $E({\Greekmath 010B} _{i})={\Greekmath 010B} _{0}=1$ and $E({\Greekmath 010C} _{ij})={\Greekmath 010C} _{0j}=1$ for $ j=1,2,...,k^{\prime }$. We also set ${\Greekmath 011B} _{{\Greekmath 010B} }^{2}=0.5$, ${\Greekmath 011B} _{{\Greekmath 010C} _{1}}^{2}=0.75$, and ${\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010C} _{j}}^{2}=0.5$ for $ j=2,...,k^{\prime }$.
When $k=k^{\prime }+1=2$, for a fixed $T\geq k$, the asymptotic bias of the FE estimator of ${\Greekmath 010C} _{01}=E({\Greekmath 010C} _{i1})$ is given by (see ((ref)))
The size of the bias will depend on ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ and the parameters of $x_{1,it}$ process. The exact expression for this bias simplifies considerably if ${\Greekmath 011A} _{1,ix}=0$ (no dynamics in the $x_{1,it}$ equation). In this case
which, noting that $E\left( {\Greekmath 011B} _{1,ix}^{2}\right) =1$, $E\left( {\Greekmath 011B} _{1,ix}^{4}\right) =3/2$ and $E\left( {\Greekmath 010D} _{1,ix}^{2}\right) =4/3$, simplifies to
It is also worth noting that when $n$ and $T\rightarrow \infty $, jointly, then $\func{plim}_{n,T\rightarrow \infty }(\hat{{\Greekmath 010C}}_{1,FE}-{\Greekmath 010C} _{01})=0.303{\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ without interactive effects in $x_{1, it}$ process, and the bias of the FE estimator does not vanish even if $ T\rightarrow \infty $.
For the baseline DGP, we generate the errors in the outcome equation as chi-squared without serial correlation (${\Greekmath 011A} _{ie}=0$ in ((ref))). $ x_{j,it}$ are generated without dynamics or interactive effects (${\Greekmath 011A} _{j,ix}=0$ and ${\Greekmath 010D} _{j,ix}=0$ in ((ref))). We set ${\Greekmath 0120} _{{\Greekmath 010B} }=0.5$ for individual fixed effects. For the slope coefficient, we consider three possibilities: (a) uncorrelated heterogeneity, ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0$, (b) a medium level of correlated heterogeneity, ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.5$, generating a bias around $15\%$ when $k^{\prime }=1$, and (c) a high level, $ {\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.8$, leading to a bias of around $24\%$ for the FE estimator when $k^{\prime }=1$. For each choice of ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$, the scalar parameter ${\Greekmath 0114} $ in ((ref)) is set such that the pooled $ R^{2} $ ($PR^{2}$) of panel regressions is around $0.2$. This is achieved by stochastic simulation for each $T$, as described in sub-section (ref) of the online supplement. We also experiment with a medium level of fit by setting $PR^{2}=0.4$ when $k^{\prime }=1$.\footnote{ The scaling of the errors in the outcome equation, ${\Greekmath 0114} $, is set when the DGP contains one regressor. As a result, the $PR^{2}$ will be slightly higher than the target values of $0.2$ and $0.4$, when $k^{\prime }=2$ or $3$ .}
We consider different values of $T$ depending on $k^{\prime }$ up to $T=15$. We also carried out a number of robustness checks, detailed in Section (ref) of the online supplement.
\@startsection{subsection}{2}{\z@} {-1.5ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .2ex} {\normalfont}{Monte Carlo findings}
\@startsection{subsubsection}{3}{\z@} {-1ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .1ex} {\normalfont}{Estimates of the tail index of the distribution of $1/d_{i}$}
We first provide estimates of the shape parameter of Pareto distributions fitted to $z_{i}=1/\func{det}(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T} \boldsymbol{X}_{i})$, $i=1,2,...,n$, for different choices of $T$ and $ k^{\prime }$. We use the estimator of ${\Greekmath 010B} _{p}$ proposed by Hill1975, which is given by
where $z^{(1)}\geq z^{(2)}....\geq z^{(m)}$ are the first $m$ largest values of $z_{i}$, and $m$ is the cut-off point set such that $m$ and $ n/m\rightarrow \infty $, as $n\rightarrow \infty $.\footnote{ See also Pickands1975 and Section 6 of PesaranYang2020 for more recent literature on estimation of ${\Greekmath 010B} _{p}$.}
Table (ref) summarizes the estimates of ${\Greekmath 010B} _{p}$ for $n=5,000$ and two cut-off values: $m=n^{1/2}$ and $n^{1/3}$, in the case of panels with $k^{\prime }=1,2$ and $3$ regressors. As to be expected, the estimates of ${\Greekmath 010B} _{p}$ are larger for the lower cut-off value, although the differences between the two estimates are quite small when $T=k^{\prime }+1$ . It is also interesting that when $T=k^{\prime }+1$, the estimates of $ {\Greekmath 010B} _{p}$ lie in the narrow interval [$0.51,0.57$], irrespective of the value of $k^{\prime }=1,2$ and $3$; thus indicating lack of moments for the individual estimates and the need for trimming. This finding is not affected when we consider more general $\{\boldsymbol{x}_{it}\}$ processes. See Table (ref) in the online supplement for $k^{\prime }=1$. Also, as to be expected, the estimates of ${\Greekmath 010B} _{p}$ rise with $T-k^{\prime}$ and exceed the threshold value of $2$ for both choices of cut-off values when $ T>2k^{\prime }+3=2k+1$. $T=2k+1$ represents a borderline case where ${\Greekmath 010B} _{p}$ is estimated to be very close to $2$, but there is still a wide margin of uncertainty. For example, for $k^{\prime }=1$ and using the cut-off value of $n^{1/2}$, estimates of ${\Greekmath 010B} _{p}$ for $T=4$ and $5$ are given by $ 1.96 $ $(0.23)$ and $2.37$ $(0.28)$, respectively. The standard errors are in parentheses.
Based on these estimates, it seems plausible to conclude that trimming should be considered if $T<2k+1$.\footnote{ Additional estimation results for ${\Greekmath 010B} _{p}$ are provided in Tables (ref) and (ref) in Section (ref) of the online supplement.} This conclusion is further supported when we compare the performance of MG and TMG estimators for different choices of $k$ and $T$ using Monte Carlo experiments, to which we now turn.
\@startsection{subsubsection}{3}{\z@} {-1ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .1ex} {\normalfont}{Comparison of TMG, FE, and MG estimators }
We begin by comparing the performance of the TMG estimator with those of FE and MG estimators under both uncorrelated and correlated heterogeneity. We consider the sample size combinations $n=1,000,2,000,5,000$, $10,000$, $ T=2,3,...,15$ subject to $T\geq k^{\prime }+1$, for $k^{\prime }=1,2$ and $3$ . Recall that $k^{\prime }$ is the number of regressors and $k=k^{\prime }+1$ . The TMG estimator depends on the indicator, $\boldsymbol{1}\{d_{i}>a_{n}\}$ , where $a_{n}=C_{n}n^{-{\Greekmath 010B} }$. In view of the discussion in Section (ref) on the choice of ${\Greekmath 010B} $, we consider the values of ${\Greekmath 010B} =1/3,0.35$ and $1/2$, and as discussed earlier we set $C_{n}=\bar{d} _{n}=n^{-1}\sum_{i=1}^{n}d_{i}>0$, where $d_{i}=\func{det}(\boldsymbol{X} _{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})$. In what follows, we report the results for the TMG estimator with ${\Greekmath 010B} =1/3$, but discuss the sensitivity of the TMG estimator to the choice of ${\Greekmath 010B} $ in Section (ref) of the online supplement. All MC experiments are based on $ R=2,000 $ replications.
Table (ref) reports bias, root mean squared errors (RMSE) and size for estimation of $E({\Greekmath 010C} _{i1})={\Greekmath 010C} _{01}$ in the case of DGPs with one regressor, $k^{\prime }=1$. The left panel of the table provides results when heterogeneity is uncorrelated (i.e. ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0$), whilst the right panel of the table gives the results for the case of correlated heterogeneity with ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.5$. The fraction of the trimmed estimates, ${\Greekmath 0119} _{n}$, defined by ((ref)), tends to be quite large for the case where $T=k$, but falls quite rapidly as $T$ $-k$ is increased. For example, for $T=2$ and $n=1,000$, as many as $27.3$ per cent of the individual estimates are trimmed when computing the TMG estimates, but this fraction falls to $0.6$ per cent when the number of time periods is increased to $T=6$. However, recall that the TMG estimator continues to make use of the trimmed estimates, as can be seen from ((ref)), and the TMG estimator shows little bias compared to the (untrimmed) MG estimator. The TMG and MG estimators converge as $T$ is increased, and they are almost identical for the panels when $T\geq 8$. This is in line with the two estimates of ${\Greekmath 010B} _{p}$ reported in Table (ref) for $ k^{\prime }=1$ and $T=8$. Both estimates ($3.11$ and $3.59$) are well in excess of $2$, such that all individual estimates have second-order moments and therefore no trimming is required.
Comparing TMG and FE estimators, we first note that in line with the theory, the FE estimator performs very well under uncorrelated heterogeneity but is badly biased when heterogeneity is correlated. Further, this bias does not diminish if $n$ and $T$ are increased. The simulated bias of the FE estimator in the case where $T=k=2$ and $n=1,000$ amounts to $0.354$, which is close to the analytical result presented in Section (ref). When heterogeneity is correlated, the FE estimator also exhibits substantial size distortions, which tend to get accentuated as $n$ is increased for a given $ T $. In contrast, the TMG estimator is robust to the choice of ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ and delivers size very close to the assumed five per cent level. \footnote{ Increasing $PR^{2}$ from $0.2$ to $0.4$ does not affect the bias and RMSE of the FE estimator but results in a higher degree of size distortion under correlated heterogeneity. See the results summarized in the right panel of Table (ref) and Table (ref) in the online supplement.}
The empirical power functions for TMG and FE estimators in the case of a single regressor and for the sample sizes $n=10,000$ and $T=2$, $3$, and $4$ , are displayed in Figure (ref). As can be seen, under uncorrelated heterogeneity (the left panel with ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0$), both estimators are centered correctly around ${\Greekmath 010C} _{01}=1$, with the FE estimator having better power properties. But the differences between the power of FE and TMG estimators shrink rapidly and become negligible as $T$ is increased from $T=2$ to $T=4$.\footnote{ But as shown in Example (ref), it does not necessarily follow that the FE estimator will dominate the TMG estimator in terms of efficiency under uncorrelated heterogeneity. See the left panel of Table (ref) and Figure (ref) in the online supplement.} The right panel provides the power plots under correlated heterogeneity with ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.5$. In this case, the empirical power functions of the FE estimator now shift markedly to the right, away from the true value, an outcome that becomes more sharpened as $ T $ is increased. In contrast, the empirical power functions for the TMG estimator are always centered correctly and are robust to the choice of $ {\Greekmath 0120} _{{\Greekmath 010C} _{1}}$.
Turning to DGPs with more than one regressor, to save space, the results are summarized in Tables (ref) and (ref) in Section (ref) of the online supplement. These tables present estimation results for the baseline DGP with two and three regressors, respectively. The FE estimator continues to display substantial bias under correlated heterogeneity. For the TMG estimator, as the number of regressors, $k^{\prime }$, increases, the trimmed fraction rises for a given $T$. With $n=1,000$, it grows from $27.3$ per cent to $ 41.6 $ percent for $T=k=3$, and $50.1$ per cent for $T=k=4$. Moreover, as $ k^{\prime }$ increases, the trimmed fraction decreases with $T$ at a slower rate, in line with the estimates of ${\Greekmath 010B} _{p}$ in Table (ref).
To summarize, under uncorrelated heterogeneity, the FE estimator performs well despite the heterogeneity, and is more efficient than the TMG estimator in the case of the baseline DGP used in our MCs. But, in general, the relative efficiency of TMG and FE estimators depends on the underlying DGP. The situation is markedly different when heterogeneity is correlated, and the FE estimator can be badly biased, leading to incorrect inference, whilst the TMG estimator provides valid inference with size around the nominal five per cent level and reasonable power, irrespective of whether ${\Greekmath 010C} _{i1}$ is correlated with $x_{1,it}$ or not.
\@startsection{subsubsection}{3}{\z@} {-1ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .1ex} {\normalfont}{Comparison of TMG and GP estimators }
Focusing on the case of correlated heterogeneity, we now compare the relative performance of TMG and GP estimators. To implement the GP estimator, defined by ((ref)), for $T=k$, we follow GP and set $ h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}$, with ${\Greekmath 010B} _{GP}=1/3$ and $C_{GP}=\frac{1}{ 2}\min \left( \hat{{\Greekmath 011B}}_{D},\hat{r}_{D}/1.34\right) $, where $\hat{{\Greekmath 011B}} _{D}$ and $\hat{r}_{D}$ are the respective sample standard deviation and interquartile range of $|\func{det}\left( \boldsymbol{W}_{i}\right) |$ with $ \boldsymbol{W}_{i}=\left( \boldsymbol{{\Greekmath 011C} }_{T},\boldsymbol{X}_{i}\right) $ . For further details of GP's choice of ${\Greekmath 010B} _{GP}$ when $T=k$, see p. 2138 of GrahamPowell2012. The asymptotic variance of their estimator in the case of models with and without time effects is provided in equation (30) on p. 2126 of their paper. There is no clear guidance by GP as to the choice of $h_{n}$ when $T>k$.\footnote{ For $T=3$ with $k=2$, GP do not use the bandwidth parameter, $h_{n}$, but directly select the \textquotedblleft percent trimmed\textquotedblright , $ {\Greekmath 0119} _{n}$. In their empirical application, they report estimates with 4 per cent being trimmed for $T=3$ with $k=2$. See the last column of Table 3 on p. 2136 of GP.} For consistency, when $T>k$, for GP estimates we continue to use their bandwidth, $h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}$ with ${\Greekmath 010B} _{GP}=1/3$ , but set $C_{GP}=\left( n^{-1}\sum_{i=1}^{n}d_{i,GP}\right) ^{1/2}$ and trim if $d_{i,GP} = \func{det}\left( \boldsymbol{W}_{i}^{\prime} \boldsymbol{ W}_{i}\right) < h_{n}^{2}$.
The bias, RMSE, and size for the two estimators are summarized in Table (ref) for $T=2,3,4,5,6$, and $n=1000,2000,5000,10000$ . The associated empirical power functions are displayed in Figure (ref).
The fractions of the trimmed estimates, $\hat{{\Greekmath 0119}}$, differ markedly across the estimators. For example, when $T=2$ and $n=1,000$, the fraction of trimmed estimates for the TMG estimator is around $27.3$ per cent as compared to $4.0$ per cent for the GP estimator, and falls to $18.9$ per cent as $n$ is increased to $10,000$. Increasing $T$ from $2$ to $3$ with $ n=1,000$ reduces this fraction to $16.5$ per cent as compared to $2$ per cent for the GP estimator. The heavy trimming causes the TMG estimator to have a larger bias than the GP estimator, particularly when $T=2$ and $n$ is large. However, the TMG estimator continues to have better overall small sample performance due to its higher efficiency. Recall that the TMG estimator makes use of the trimmed estimates, as set out in the second term of ((ref)), but the trimmed estimates are not used in the GP estimator. This difference in the way trimmed estimates are treated is reflected in the lower RMSE of the TMG estimator as compared to MG and GP estimators for all $T$ and $n$ combinations. For example, when $T=2$ and $ n=1,000$, the RMSE of the TMG is $0.27$ as compared to $0.60$ for the GP estimator. The relative advantage of the TMG estimator continues when $T$ increases from $2$ to $3$. For $T=3$, the RMSE of the TMG estimator stands at $0.17$ compared to $0.21$ for the GP estimator. The larger the value of $ T $, the less important the trimming becomes.
The empirical power functions for TMG, GP and MG estimators are shown in Figure (ref). As can be seen, the TMG estimator is uniformly more powerful than the GP estimator. To save space, the MC results for models with time effects are summarized in sub-section (ref) of the online supplement. Similar outcomes are obtained when we consider DGPs with two or three regressors. See Tables (ref) and (ref), and the corresponding empirical power functions, Figures (ref) and (ref), in sub-section (ref) of the online supplement. This is particularly the case when $T=k\in \{3,4\}$, where the trimmed fraction of the GP estimator rises only slightly with the number of regressors, resulting in substantial declines in the empirical powers.
For the purpose of comparisons, in addition to the choice of ${\Greekmath 010B} _{GP}=1/3$ by GP, we also considered the threshold values $2{\Greekmath 010B} _{GP}={\Greekmath 010B} \in \{0.35,1/2\}$, so that the two threshold functions (ours and the one suggested by GP) share the same exponents. As reported in Table (ref), the TMG estimator has a lower RMSE for all choices of ${\Greekmath 010B} $ and ${\Greekmath 010B} _{GP}$, when $T=2$, and delivers better empirical powers, as shown in Figures (ref) and (ref) in the online supplement. Additional results for $ T=3$ are provided in sub-section (ref) of the online supplement, where the differences between TMG and GP estimators are much smaller.
\@startsection{subsubsection}{3}{\z@} {-1ex\@plus -1ex \@minus -.2ex} {0.5ex \@plus .1ex} {\normalfont}{MC evidence on the Hausman test of correlated heterogeneity}
Table (ref) reports empirical size and power of the Hausman test of correlated heterogeneity given by ((ref)) under three scenarios: homogeneity (left), uncorrelated heterogeneity (middle), and correlated heterogeneity (right). We continue to focus on the simple case where $k^{\prime }=1$. The size of the test is around the nominal level of 5 per cent. When $\boldsymbol{x}_{it}$ is strictly exogenous and $ \boldsymbol{{\Greekmath 010C} }_{i1}$ is distributed independently of $\boldsymbol{X} _{i} $, FE, MG and TMG estimators are all consistent under homogeneity and if heterogeneity is present but uncorrelated. In such a case, the Hausman test does not have power. However, in the case where slope coefficients are heterogeneous and correlated with the regressors, the TMG estimator is consistent while the FE estimator is biased for all $T$. In this case, we would expect the proposed test to have power, and this is indeed evident in the right panel of Table (ref). Also, the power of the test rises with increases in $n$ even when $T=2$, illustrating the (ultra) small $T$ consistency of the proposed test. We also obtain similar test results when we allow for time effects. See sub-section (ref) of the online supplement.
Hausman test results for DGPs with two or three regressors are summarized in Tables (ref) through (ref) of the online supplement. Recall that we allow for correlated heterogeneity only in the coefficients of the first regressor. Consequently, we observe a decline in the power of the Hausman test as we add regressors with uncorrelated heterogeneous coefficients. The empirical power of the test decreases with $k$, particularly when $T=k$.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Empirical application }
In this section, we re-visit the empirical application in GrahamPowell2012 who provide estimates of the average effect of household expenditures on calorie demand, based on a sample of households from poor rural communities in Nicaragua that participated in a conditional cash transfer program. The data set is a balanced panel with $n=1,358$ households observed from 2000 to 2002. We present estimates of the average effects using the following panel data model with time effects:
where $\ln (Cal_{it})$ denotes the logarithm of household calorie availability per capita in year $t$ of household $i$, and $\ln (Exp_{it})$ denotes the logarithm of real household expenditures per capita (in thousands of 2001 cordobas) of household $i$ in year $t$. The parameter of interest is the average effect defined by ${\Greekmath 010C} _{0}=E({\Greekmath 010C}_{i})$.
We first estimate ${\Greekmath 010B} _{p}$ as a diagnostic to see if trimming is needed. For the panel covering the period 2001--2002 $(T=2)$, the Hill's estimates of ${\Greekmath 010B} _{p}$ are 0.71 (0.12) and 0.59 (0.17) for the two cut-off values of $n^{1/2}$ and $n^{1/3}$, respectively, with standard errors in parentheses. Similarly, for the dataset 2000--2002 $(T=3)$, the estimates of ${\Greekmath 010B} _{p}$ are 0.94 (0.15) and 0.96 (0.28), respectively. All these estimates are well below two, and it is advisable that the trimmed MG estimator is used for consistent estimation of the average effects, $ {\Greekmath 010C} _{0}$, in case heterogeneity in ${\Greekmath 010C} _{i}$ is correlated.
Table (ref) reports results of the Hausman test of correlated heterogeneity in the effects of household expenditures on calorie demand. The null hypothesis of uncorrelated heterogeneity is rejected for both the panels and irrespective of whether time effects are included. Therefore, for this application, the FE and TWFE estimates of ${\Greekmath 010C} _{0}$ could be biased.
Table (ref) presents the estimates of ${\Greekmath 010C} _{0}$ based on the panel of 2001--2002 (with $T=2$) without time effects (left panel), and with time effects (right panel). The estimates are not affected by the inclusion of time effects but differ considerably across different methods. \footnote{ When $T=2$, $\hat{{\Greekmath 011E}}_{2002}$ is not significant, and adding time effects does not change the estimated average effect.} Turning to the trimmed estimators, we find that only the TMG estimator is heavily trimmed with 27.1 per cent of the estimates being trimmed, whilst the rate of trimming is only around 3.8 per cent for the GP estimator.\footnote{ For the 2001--2002 panel, $\hat{{\Greekmath 0119}}$ of the GP estimator is identical to the one reported in Table 3 of GrahamPowell2012. GP estimated a model with time-varying coefficients, $y_{it}={\Greekmath 010B} _{i}+{\Greekmath 011E} _{t}+\left( {\Greekmath 010C} _{i}+{\Greekmath 011E} _{t,{\Greekmath 010C} }\right) x_{it}+u_{it}$, where $\left( \boldsymbol{{\Greekmath 011E} } ^{\prime},\boldsymbol{{\Greekmath 011E} }_{{\Greekmath 010C} }^{\prime}\right)^{\prime}$ are identified by stayers but estimated by near stayers with $\boldsymbol{{\Greekmath 011E} } _{{\Greekmath 010C} } = ({\Greekmath 011E} _{1,{\Greekmath 010C} }, {\Greekmath 011E} _{2,{\Greekmath 010C} }, ..., {\Greekmath 011E} _{T,{\Greekmath 010C} })^{\prime}$. While $\boldsymbol{{\Greekmath 011E} }_{{\Greekmath 010C} }$ is not included in ((ref)), the GP estimates we compute are close to the trimmed estimates in Table 3 of GrahamPowell2012.} Focusing on the estimates without time effects, we find the FE estimate, 0.6568 (0.0287), is much larger and more precisely estimated than either the GP or TMG estimates, given by 0.4549 (0.1003) and 0.5623 (0.0425), respectively, with standard errors in brackets. Judging by the standard errors, it is also noticeable that the TMG is more precisely estimated than the GP estimate and lies somewhere between the FE and GP estimates. These estimates are in line with the MC results reported in the previous section, where we found that in the presence of correlated heterogeneity, FE estimates are biased with smaller standard errors (thus leading to incorrect inference), whilst GP and TMG estimators are correctly centered, with the TMG estimator being more efficient. Similar results are obtained when we use the extended panel with $T=3$ (2000--2002), presented in Section (ref) of the online supplement.
\@startsection {section}{1}{\z@} {-1.5ex \@plus -1ex \@minus -.2ex} {0.8ex \@plus.2ex} {\normalfont}{Conclusions }
This paper studies the estimation of average effects in panel data models with possibly correlated heterogeneous coefficients, when the number of cross-sectional units is large, but the number of time periods can be as small as the number of regression coefficients. We recall that the FE estimator is inconsistent under correlated heterogeneity, and the MG estimator could not have second-order (or even first-order) moments when applied to ultra short panels. The TMG estimator is therefore proposed to deal with the fat-tailed distributions of the individual estimates (which does arise if $T$ is very close to $k$) by shrinking (not trimming) individual estimates that are most likely to fail the second-order moment condition. The TMG estimator is shown to be consistent and asymptotically normally distributed, but at a slower rate than $\sqrt{n}$. The paper also proposes TMG estimators for panels with time effects, distinguishing between cases where $T=k$ and $T>k$. The TMG estimators play a crucial role in assessing the robustness of FE and TWFE estimators against slope correlated heterogeneity. The dispersion slope homogeneity tests by PesaranYamagata2008 require large $T$ and do not differentiate between uncorrelated heterogeneity and correlated heterogeneity.
We highlight the bias and size distortion properties of the FE and TWFE estimators under correlated heterogeneity. In contrast, the TMG and TMG-TE estimators are shown to have desirable finite sample performance under a number of different MC designs, allowing for Gaussian and non-Gaussian heteroskedastic error processes, dynamic heterogeneity and interactive effects in the covariates, different numbers of regressors, and different choices of the trimming threshold parameter, ${\Greekmath 010B} $. In particular, since the TMG and TMG-TE estimators exploit information on all available individual estimates, they have the smallest RMSE, and tests based on them have the correct size and are more powerful than the other trimmed estimators currently proposed in the literature. The Hausman tests based on TMG and TMG-TE estimators are also shown to have very good small sample properties, with their size controlled and their power rising strongly with $ n$ even when $T=k=2$.
{ \setstretch{1.02} }