The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
174,589 characters
Estimation of Average Effects in Short $T$ Heterogeneous Panels
\title{{\large {Estimation of Average Effects in Short $T$ Heterogeneous
Panels\thanks{{\small We are grateful to Cheng Hsiao, Oliver Linton, Ron
Smith and Hayun Song for helpful comments on an earlier version of this
paper. The current version has benefited greatly from constructive comments
and suggestions from Jiti Gao, who was the discussant of our paper at the
2025 Workshop in Honour of Professors Heather Anderson and Farshid Vahid at
Monash University, two anonymous referees and an Associate Editor. The
research on this paper was started when Liying Yang was a Ph.D. student at
the Department of Economics, University of Southern California.}}}} }
\author{ {\normalsize M. Hashem Pesaran\thanks{{\small Trinity College,
University of Cambridge, UK, and Department of Economics, University of
Southern California, US, Email: [email removed].}}} \and {\normalsize Liying
Yang\thanks{{\small Shenzhen Audencia Financial Technology~Institute,
Shenzhen University, China, Email: [email removed].}} }}
\date{{\normalsize \today }}
\maketitle
\begin{abstract}
The commonly used two-way fixed effects estimator is biased under correlated
heterogeneity and can lead to misleading inference. The mean group estimator
proposed by \cite{PesaranSmith1995} is robust to correlated heterogeneity
but requires the underlying individual estimates to have second-order
moments that could fail if the number of estimated coefficients ($k$) is too
close to the time dimension ($T$) of the panel. This paper focuses on panels
where $k$ is close to $T$ (including $k=T$), and proposes a trimmed mean
group (TMG) estimator that shrinks individual estimates most likely to fail
the second-order moment condition. The TMG estimator is shown to be $
n^{(1-{\Greekmath 010B} )/2}$-consistent and asymptotically normally distributed, where
${\Greekmath 010B}$ is determined by the degree to which individual estimates might not
have moments. The $\sqrt{n}$ convergence rate is achieved only if all
individual estimates have second-order moments. Extensions to panels with
time effects are provided, and a new Hausman test of correlated
heterogeneity is proposed. Small sample properties of the TMG estimator
(with and without time effects) are investigated by Monte Carlo experiments
and shown to be satisfactory. The proposed test of correlated heterogeneity
is also shown to have the correct size and satisfactory power. The utility
of the TMG approach is illustrated with an empirical application.
{\small \noindent \textbf{Keywords}: Correlated heterogeneity, irregular
estimators, two-way fixed effects, mean group estimation, tests of
correlated heterogeneity, calorie demand}
{\small \noindent \textbf{JEL Classification:} C21, C23}
\end{abstract}
\vspace{-5mm}
\vspace{-10mm}
\thispagestyle{empty}
\newpage \setcounter{page}{1}
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Introduction}
\doublespacing
Fixed effects estimation of average effects has been predominantly utilized
for program and policy evaluation. For static panel data models where slope
heterogeneity is uncorrelated with regressors, two-way fixed effects (TWFE)
estimators that allow for unit-specific and time effects are $\sqrt{n}$
-consistent, and if used in conjunction with robust standard errors lead to
valid inference in panels where the time dimension, $T$, is short and the
cross section dimension, $n$, is sufficiently large. However, when the slope
heterogeneity is correlated with the regressors, the TWFE estimators could
become inconsistent even if both $n$ and $T\rightarrow \infty$.\footnote{
The concept of the correlated random coefficient model is due to \cite
{HeckmanVytlacil1998}. \cite{Wooldridge2005} shows that TWFE estimators
continue to be consistent if slope heterogeneity is mean-independent of all
the de-trended covariates. See also condition (\ref{FEconsistency}) given
below.} Such correlated heterogeneity arises endogenously in the case of
dynamic panel data models, as originally noted by \cite{PesaranSmith1995},
and more generally, when the regressors are weakly exogenous. In the case of
static panels with strictly exogenous regressors, correlated (slope)
heterogeneity can arise, for example, when there is a high degree of
variation in treatments across units or when there are latent factors that
influence the level and variability of treatments and their outcomes. For
example, in estimation of returns to education, the choice of educational
level is likely to be correlated with expected returns to education. Other
examples include estimation of the effects of training programs on workers'
productivity and earnings reviewed by \cite{CreponVandenberg2016},
evaluation of the effectiveness of micro-credit programs discussed by \cite
{BanerjeeEtal2015}, and the analysis of the effectiveness of anti-poverty
cash transfer programs considered by \cite{BastagliEtal2019}.
In the presence of correlated heterogeneity, \cite{PesaranSmith1995}
proposed to estimate the mean effects by simple averages of individual
estimates, which they called the mean group (MG) estimator. They showed that
the MG estimator is consistent for dynamic heterogeneous panels when $n$ and
$T$ are both large. It was later shown that for panels with strictly
exogenous regressors, the MG estimator is in fact $\sqrt{n}$-consistent in
the presence of correlated heterogeneity even if $T$ is fixed as $
n\rightarrow \infty $, so long as $T$ is sufficiently large such that
second-order moments of the individual estimates exist. However, such moment
conditions need not hold when $T$ is very close to the number of estimated
coefficients ($k$). In effect, we are faced with the problem of estimating
the mean of random variables with fat tails.
This problem was originally recognized by \cite{Chamberlain1992}, who showed
that one needs $T$ to be strictly larger than $k$ for regular identification
of average effects under correlated heterogeneity.\footnote{
An unknown parameter, ${\Greekmath 010C} _{0}$, is said to be regularly identified if
there exists an estimator that converges to ${\Greekmath 010C} _{0}$ in probability at
the rate of $\sqrt{n}$. Any estimator that converges to its true value at a
rate slower than $\sqrt{n}$ is said to be irregularly identified.} In the
statistics literature, \cite{CsorgoEtal1988b,CsorgoEtal1988a} and \cite
{GriffinPruitt1989}, among others, consider a trimmed mean estimator whereby
the observations are ordered, and those below and above a given threshold
value are excluded. However, in general such trimmed estimators may not
possess a limiting distribution if the observations are drawn from a
heavy-tailed distribution with the tail index, ${\Greekmath 010B} _{p}<2$, and the
threshold values are chosen in an \textit{ad hoc} manner. \cite{Peng2001}
proposes a new trimmed estimator where the thresholds are endogenized and
the trimmed mean is augmented with mean estimates from the two tails. Peng's
modified trimmed estimator is shown to have the same limiting normal
distribution as the sample mean but with a slower convergence rate when $
{\Greekmath 010B} _{p}<2$ (see Theorem 1 and Remark 1 of \cite{Peng2001}). For settings
of inverse probability weighting where the denominator can be arbitrarily
close to zero, leading to heavy-tailed sampling distributions, \cite
{MaWang2020} propose trimmed estimators that exclude observations with small
denominators and develop inference procedures for different trimming
threshold choices. For our purposes, trimmed estimators proposed by \cite
{Peng2001} and \cite{MaWang2020} are subject to two limitations. They
consider only the scalar case and assume that the observations (in our
application, the individual estimates) are identically and independently
distributed. Extension of their estimators to a vector of estimates that are
not identically distributed does not seem to be straightforward.
In the context of MG estimation, we have additional information about the
precision of the individual estimates that are not used in the trimming
approaches considered in the statistics literature. One example where such
information is utilized is provided by \cite{GrahamPowell2012} (GP), who
build on the pioneering work of \cite{Chamberlain1992}. GP focus on panels
with $T=k$, where identification issues of time effects and the mean
coefficients arise especially when there are insufficient within-individual
variations for some regressors. These authors derive an irregular estimator
of the mean coefficients by excluding individual estimates from the
estimation of the average effects if the sample variance of regressors in
question is smaller than a given threshold value.
In this paper, we begin by providing conditions under which MG and fixed
effects (FE) estimators of the average effects in heterogeneous panels are $
\sqrt{n}$-consistent, which serves as a basis for developing a diagnostic
test of the validity of (two-way) fixed effects estimators commonly used in
the literature. We then propose a trimmed mean group (TMG) estimator that
does not exclude any of the individual estimates, but uses a threshold
function similar to that of GP to shrink some of the estimates to overcome
the fat-tailed nature of the distribution of the individual estimates when $
T $ is very close to $k$. In effect, we shrink rather than drop estimates as
done by GP. The decision on whether unit $i$ is subject to shrinkage is made
with respect to the determinant of the sample variance matrix of the
regressors, denoted by $d_{i}(>0)$. Individual estimates become
ill-conditioned when $d_{i}$ is close to zero. To characterize the trimming
process, we assume that $1/d_{i}$ follows Pareto-type distributions with the
tail index, ${\Greekmath 010B} _{p}$, and show that individual estimates have
second-order moments when ${\Greekmath 010B} _{p}>2$. Shrinkage of the individual
estimates is required only if ${\Greekmath 010B} _{p}\leq 2$. The literature on
estimation of ${\Greekmath 010B} _{p}$ is well established, for example, \cite
{EmbrechtsEtal1997}, which can guide decisions on shrinkage/trimming before
estimating the average effects. Based on extensive Monte Carlo experiments,
we find that in general ${\Greekmath 010B} _{p}>2$ when $T\geq 2k+1$, which could be
used as a practical rule of thumb when deciding whether to use the MG
estimator or its trimmed version proposed in this paper.
We also consider heterogeneous panels with time effects and show that the
TMG procedure can be applied after the time effects are eliminated. In the
case where $T=k$, we require the dependence between heterogeneous slope
coefficients and the regressors to be time-invariant. This assumption is not
required when $T>k$, and the time effects can be eliminated using the
approach first proposed by \cite{Chamberlain1992}. We refer to these
estimators as TMG-TE and derive their asymptotic distributions, bearing in
mind that the way the time effects are eliminated depends on whether $T=k$
or $T>k$.
As noted above, slope heterogeneity by itself does not render the TWFE
estimators inconsistent, which continue to have the regular convergence rate
of $\sqrt{n}$. The problem arises when slope heterogeneity is correlated
with the covariates. It is therefore important that before using the TWFE
estimators, the null of uncorrelated heterogeneity is tested. To this end,
we also propose Hausman tests of correlated heterogeneity by comparing the
FE and TWFE estimators with the associated TMG estimators and derive their
asymptotic distributions under fairly general conditions. The earlier
Hausman test of slope homogeneity developed by \cite{PesaranEtal1996} is
based on the difference between FE and MG estimators and does not apply when
$T$ is close to $k$. The more recent dispersion-based test of slope
homogeneity proposed by \cite{PesaranYamagata2008} is shown to be quite
powerful as a test of slope homogeneity but does not distinguish correlated
or uncorrelated heterogeneity and requires $\sqrt{n}/T^{2}\rightarrow 0$ as $
n$ and $T\rightarrow \infty $ jointly.
We also carry out an extensive set of Monte Carlo (MC) simulations to
investigate the small sample properties of the TMG and TMG-TE estimators and
how they compare with the trimmed estimator proposed by GP. The MC evidence
on the size and empirical power of the Hausman tests of correlated
heterogeneity in panel data models with and without time effects is
provided, and the sensitivity of estimation results to the choice of the
trimming threshold parameter, ${\Greekmath 010B} $, is also investigated. The MC and
theoretical results of the paper are all in agreement. The TMG and TMG-TE
estimators not only have the correct size but also achieve better finite
sample properties compared with the other trimmed estimator across a number
of experiments with different data generating processes, allowing for
heteroskedasticity (random and correlated), error serial correlations, and
regressors with heterogeneous dynamics and interactive effects. The
simulation results also confirm that the Hausman tests based on the
difference between FE (TWFE) and TMG (TMG-TE) estimators have the correct
size and power against the alternative of correlated heterogeneity.
Finally, we illustrate the utility of our proposed trimmed estimators by
re-examining the average effect of household expenditures on calorie demand
using a balanced panel of $1,358$ households in poor rural communities in
Nicaragua over the years 2001--2002 $(T=2)$ and 2000--2002 $(T=3)$.
The rest of the paper is organized as follows. Section \ref{HPM} sets out
the heterogeneous panel data model and discusses the asymptotic properties
of FE and MG estimators. Section \ref{IRMGE} considers ultra short $T$
panels (including the case of $T=k$) and introduces the proposed TMG
estimator, with its asymptotic properties established in Section \ref{AsyTMG}
. Section \ref{CRCTE} extends the TMG estimation to panels with time
effects, distinguishing between cases where $T>k$ and $T=k$. Section \ref
{Test} sets out the Hausman test of correlated heterogeneous slope
coefficients. Section \ref{subset} discusses how to apply the TMG approach
to a subset of coefficients of interest. Section \ref{MC} provides the main
findings of the MC experiments. Section \ref{APP} presents the empirical
illustration. Section \ref{conclusion} concludes. Mathematical proofs of
propositions and theorems are given in a mathematical appendix.
Supplementary materials covering additional mathematical derivations, MC
experiments, as well as further empirical results, are provided in an online
supplement.
\textbf{Notations:} Generic positive finite constants are denoted by $C$
when large, and $c$ when small. They can take different values at different
instances. ${\Greekmath 0115} _{\max }\left( \boldsymbol{A}\right) $ and ${\Greekmath 0115}
_{\min }\left( \boldsymbol{A}\right) $ denote the maximum and minimum
eigenvalues of matrix $\boldsymbol{A}$. $\boldsymbol{A}\succ \boldsymbol{0}$
and $\boldsymbol{A}\succeq \boldsymbol{0}$ denote that matrix $\boldsymbol{A}
$ is positive definite and is positive semi-definite, respectively. When
matrix $\boldsymbol{A}$ is square, its adjugate (adjoint) and determinant
are denoted by $\func{adj}(\boldsymbol{A})$ and $\det (\boldsymbol{A})$,
respectively. If $\det (\boldsymbol{A})\neq 0$, then the inverse of $
\boldsymbol{A}$ is given by $\boldsymbol{A}^{-1}=\func{adj}(\boldsymbol{A)}
/\det (\boldsymbol{A)}$. $\left\Vert \boldsymbol{A}\right\Vert ={\Greekmath 0115}
_{\max }^{1/2}(\boldsymbol{A}^{\prime }\boldsymbol{A)}$ and $\left\Vert
\boldsymbol{A}\right\Vert _{1}$ denote the spectral and column norms of
matrix $\boldsymbol{A}$, respectively. $\left\Vert \boldsymbol{x}\right\Vert
_{p}=\left[ E\left( \left\Vert \boldsymbol{x}\right\Vert ^{p}\right) \right]
^{1/p}$. If $\left\{ f_{n}\right\} _{n=1}^{\infty }$ is any real sequence
and $\left\{ g_{n}\right\} _{n=1}^{\infty }$ is a sequence of positive real
numbers, then $f_{n}=O(g_{n})$ if there exists $C$ such that $\left\vert
f_{n}\right\vert /g_{n}\leq C$ for all $n$, and $f_{n}=o(g_{n})$ if $
f_{n}/g_{n}\rightarrow 0$ as $n\rightarrow \infty $. Similarly, $
f_{n}=O_{p}(g_{n})$ if $f_{n}/g_{n}$ is stochastically bounded, and $
f_{n}=o_{p}(g_{n})$, if $f_{n}/g_{n}\rightarrow _{p}0$. $f_{n}=\ominus
(g_{n})$ if there exist $n_{0}\geq 1$ and positive finite constants $C_{0}$
and $C_{1}$, such that $\inf_{n\geq n_{0}}\left( \left\vert f_{n}\right\vert
/g_{n}\right) \geq C_{0}$, and $\sup_{n\geq n_{0}}\left( \left\vert
f_{n}\right\vert /g_{n}\right) \leq C_{1}$. The operator $\rightarrow _{p}$
denotes convergence in probability, and $\rightarrow _{d}$ denotes
convergence in distribution. $IID$ stands for independently and identically
distributed. $\boldsymbol{u}\perp \boldsymbol{v}$ is used to show that
vectors of random variables $\boldsymbol{u}$ and $\boldsymbol{v}$ are
independently distributed.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Heterogeneous linear panel data models\label{HPM}}
Consider the following panel data model with individual fixed effects, $
{\Greekmath 010B} _{i}$, and heterogeneous slope coefficients, $\boldsymbol{{\Greekmath 010C} }_{i}$
,
\begin{equation}
y_{it}={\Greekmath 010B} _{i}+\boldsymbol{{\Greekmath 010C} }_{i}^{\prime }\boldsymbol{x}
_{it}+u_{it}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }i=1,2,...,n\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }t=1,2,...,T, \label{eq2}
\end{equation}
where $\boldsymbol{x}_{it}$ is a $k^{\prime }\times 1$ vector of regressors,
and $u_{it}$ is the error term. $\left\{ \boldsymbol{{\Greekmath 010C} }_{i}\right\}
_{i=1}^{n}$ follow the random coefficient model
\begin{equation}
\boldsymbol{{\Greekmath 010C} }_{i}=\boldsymbol{{\Greekmath 010C} }_{0}+\boldsymbol{{\Greekmath 0111} }_{i},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
for }i=1,2,...,n, \label{RCM}
\end{equation}
where $\left\{ \boldsymbol{{\Greekmath 0111} }_{i}\right\} _{i=1}^{n}$ are the random
components, and $\boldsymbol{{\Greekmath 010C} }_{0}$ is the $k^{\prime }\times 1$
vector of average effects. In matrix notations,
\begin{equation}
\boldsymbol{y}_{i}={\Greekmath 010B} _{i}\boldsymbol{{\Greekmath 011C} }_{T}+\boldsymbol{X}_{i}
\boldsymbol{{\Greekmath 010C} }_{i}+\boldsymbol{u}_{i}, \label{m1}
\end{equation}
where $\boldsymbol{y}_{i}=(y_{i1},y_{i2},...,y_{iT})^{\prime }$, $
\boldsymbol{{\Greekmath 011C} }_{T}$ is a $T\times 1$ vector of ones, $\boldsymbol{X}
_{i}=(\boldsymbol{x}_{i1},\boldsymbol{x}_{i2},...,\boldsymbol{x}
_{iT})^{\prime }$, and $\boldsymbol{u}_{i}=(u_{i1},u_{i2},...,u_{iT})^{
\prime }$. The FE estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ is given by
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}=\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) ^{-1}\left(
n^{-1}\sum_{i=1}^{n}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{y}_{i}\right) , \label{fee}
\end{equation}
where $\boldsymbol{M}_{T}=\boldsymbol{I}_{T}-T^{-1}\boldsymbol{{\Greekmath 011C} }_{T}
\boldsymbol{{\Greekmath 011C} }_{T}^{\prime }$, and $\boldsymbol{I}_{T}$ is a $T\times T$
identity matrix. It is well known that for a fixed $T\geq k=k^{\prime }+1$, $
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ is a $\sqrt{n}$-consistent estimator of $
\boldsymbol{{\Greekmath 010C} }_{0}$ and robust to possible correlations between ${\Greekmath 010B}
_{i}$ and $\{\boldsymbol{x}_{it}\}_{t=1}^{T}$, even under slope
heterogeneity so long as $\boldsymbol{{\Greekmath 0111} }_{i}$ are not correlated with $
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}$. This
condition is clearly satisfied when heterogeneity is exogenous and $E\left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{
{\Greekmath 0111} }_{i}\right) =\boldsymbol{0}$, for the majority of the units (to be
formalized below).
In the presence of correlated slope heterogeneity, the MG estimator,
initially proposed by \cite{PesaranSmith1995}, is typically considered for
consistent estimation of the average effects. When $T\geq k$, $\boldsymbol{
{\Greekmath 010C} }_{0}$ can be estimated by the MG estimator, $\boldsymbol{\hat{{\Greekmath 010C}}}
_{MG}$, computed as a simple average of the least square estimates of $
\boldsymbol{{\Greekmath 010C} }_{i}$, namely
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}=\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\hat{{\Greekmath 010C}
}}_{i}, \label{mge}
\end{equation}
where
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}=\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M
}_{T}\boldsymbol{X}_{i}\right) ^{-1}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M
}_{T}\boldsymbol{y}_{i}. \label{betaihat}
\end{equation}
In contrast to the FE estimator, the MG estimator does not depend on $
E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}
\boldsymbol{{\Greekmath 0111} }_{i}\right) $ and is consistent irrespective of whether
slope heterogeneity is correlated or not. As pointed out by an associate
editor, $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ can also be written as a weighted
average of $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$,
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\mathcal{W}}
_{i}\boldsymbol{\hat{{\Greekmath 010C}}}_{i}, \label{fee2}
\end{equation}
where $\boldsymbol{\mathcal{W}}_{i}=\left( n^{-1}\sum_{j=1}^{n}\boldsymbol{X}
_{j}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{j}\right) ^{-1}\left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) $ is
the $k^{\prime }\times k^{\prime }$ weight matrix for unit $i$. The
difference between $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ and $\boldsymbol{\hat{
{\Greekmath 010C}}}_{MG}$ lies in the choice of the weights. When heterogeneity is
exogenous, both estimators converge to the same limit $\boldsymbol{{\Greekmath 010C} }
_{0}$. However, when heterogeneity is correlated, as we shall see, $
\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ continues to converge to $\boldsymbol{{\Greekmath 010C} }
_{0}$, but $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ converges to $\boldsymbol{{\Greekmath 010C} }
_{0}+\lim\limits_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left(
\boldsymbol{\mathcal{W}}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) $, where $
\lim\limits_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{
\mathcal{W}}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) \neq \boldsymbol{0}$.
To investigate the asymptotic properties of FE and MG estimators for a fixed
$T\geq k$ as $n\rightarrow \infty $, we consider the following assumptions:
\begin{assumption}[errors]
\label{errors} Conditional on $\boldsymbol{X}_{i}$, (a) the errors, $u_{it}$
, in (\ref{eq2}) are cross-sectionally independent, (b) $E(\boldsymbol{u}
_{i}\left\vert \boldsymbol{X}_{j}\right. )=\boldsymbol{0}$, for all $i$ and $
j$, and (c) $E(\boldsymbol{u}_{i}\boldsymbol{u}_{i}^{\prime }\left\vert
\boldsymbol{X}_{i}\right. )=\boldsymbol{H}_{i}(\boldsymbol{X}_{i})=
\boldsymbol{H}_{i}$, where $\boldsymbol{H}_{i}$ is a symmetric $T\times T$
matrix with $0<c<\func{inf}_{i}{\Greekmath 0115} _{min}\left( \boldsymbol{H}
_{i}\right) <\func{sup}_{i}{\Greekmath 0115} _{max}\left( \boldsymbol{H}_{i}\right) <C$
.
\end{assumption}
\begin{assumption}[coefficients]
\label{rcm} For $i=1,2,...,n$, the $k^{\prime }\times 1$ vector of
unit-specific slope coefficients, $\boldsymbol{{\Greekmath 010C} }_{i}$, follow the
random coefficient model
\begin{equation}
\boldsymbol{{\Greekmath 010C} }_{i}=\boldsymbol{{\Greekmath 010C} }_{0}+\boldsymbol{{\Greekmath 0111} }_{i},
\label{RC}
\end{equation}
where $\left\Vert \boldsymbol{{\Greekmath 010C} }_{0}\right\Vert <C$, and $\left\{
\boldsymbol{{\Greekmath 0111} }_{i}\right\}_{i=1}^{n}$ are independently distributed with
mean zero and a bounded variance, $Var(\boldsymbol{{\Greekmath 010C} }_{i}|\boldsymbol{X}
_{i})=\boldsymbol{\Omega }_{{\Greekmath 010C} }\succeq \boldsymbol{0}$, such that
\begin{equation}
\boldsymbol{\bar{{\Greekmath 010C}}}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{{\Greekmath 010C} }_{i}=
\boldsymbol{{\Greekmath 010C} }_{0}+O_{p}(n^{-1/2}). \label{Truebeta}
\end{equation}
\end{assumption}
\begin{assumption}[pooling]
\label{PoolA} (a) The pooled sample covariance matrix $\boldsymbol{\bar{\Psi}
}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\Psi }_{i}$, where $\boldsymbol{\Psi }
_{i}=\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}$,
tends to $\lim_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{
\Psi }_{i}\right) =\boldsymbol{\bar{\Psi}}\succ \boldsymbol{0}$ and $
\boldsymbol{\bar{\Psi}}_{n}^{-1}=\boldsymbol{\bar{\Psi}}^{-1}+o_{p}(1)$. (b)
The sample covariances $\left\{ \boldsymbol{\Psi }_{i}\right\}_{i=1}^{n} $
satisfy the condition $0<c<\func{inf}_{i}{\Greekmath 0115} _{min}\left( T^{-1}
\boldsymbol{\Psi }_{i}\right) <\func{sup}_{i}{\Greekmath 0115} _{max}\left( T^{-1}
\boldsymbol{\Psi }_{i}\right) <C$, for a fixed $T\geq k$.
\end{assumption}
\begin{assumption}[correlated heterogeneity]
\label{CrCorr} The $k^{\prime }\times 1$ vectors $\boldsymbol{{\Greekmath 0110} }_{iT}=
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{
{\Greekmath 0111} }_{i}$, for $i=1,2,...,n$, are weakly cross-correlated such that for a
fixed $T \geq k$,
\begin{equation}
\sup_{j}\sum_{i=1}^{n}\left\Vert Cov\left( \boldsymbol{{\Greekmath 0110} }_{iT},
\boldsymbol{{\Greekmath 0110} }_{jT}\right) \right\Vert <C. \label{WCR}
\end{equation}
\end{assumption}
\begin{remark}
Assumption \ref{errors} requires the regressors, $\boldsymbol{x}_{it}$, to
be strictly exogenous, but it allows the conditional variance of $
\boldsymbol{u}_{i}$ to depend on $\boldsymbol{X}_{i}$, and the errors, $
u_{it}$, to be serially correlated. Assumption \ref{rcm} implies that $
\boldsymbol{{\Greekmath 010C} }_{0}=\func{plim}_{n\rightarrow \infty }\left(
n^{-1}\sum_{i=1}^{n}\boldsymbol{{\Greekmath 010C} }_{i}\right) $. Part (a) of Assumption
\ref{PoolA} is standard in the literature on FE estimation, and part (b) is
required for estimation of individual estimates and can be relaxed if
particular linear combinations of $\boldsymbol{{\Greekmath 010C} }_{i}$ are of interest.
Assumption \ref{CrCorr} is required for the convergence of the FE estimator
under correlated heterogeneity. Condition (\ref{WCR}) trivially holds when
heterogeneity is exogenous and $\boldsymbol{{\Greekmath 0111} }_{i}$ are independently
distributed, as under Assumption \ref{rcm}.
\end{remark}
\begin{remark}
The above assumptions do not require $(y_{it},\boldsymbol{x}_{it}^{\prime
},u_{it})$ to be identically and independently distributed as often assumed
in micro econometric panel data analysis with $T$ fixed. See Example \ref
{ex1} below, and the MC designs used for evaluation of the small sample
performance of the proposed TMG estimators.
\end{remark}
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Why does correlated slope heterogeneity matter?}
\label{CREmatter}
Correlated slope heterogeneity can arise in various contexts, with important
implications for estimation and inference on the average effects. For
example, return to education is likely to be positively correlated with
latent factors such as ability, talent, and degree of self-belief.
Similarly, propensity to save across households is often inversely related
to their income volatility. FE estimation allows for possible correlation
between $\boldsymbol{x}_{it}$ and ${\Greekmath 010B} _{i}$, but not between $
\boldsymbol{x}_{it}$ and $\boldsymbol{{\Greekmath 010C} }_{i}$. As a simple example,
consider the panel data model
\begin{equation*}
y_{it}={\Greekmath 010B} _{i}+{\Greekmath 010C} _{i}x_{it}+u_{it},
\end{equation*}
where $x_{it}$ and $y_{it}$ could, respectively, be years of schooling and
return to schooling, or could be income and the saving rate of individual $i$
at time $t$. Suppose now that $x_{it}$ follows the following model with an
interactive effect and idiosyncratic error heteroskedasticity:
\begin{equation*}
x_{it}={\Greekmath 010B} _{ix}+{\Greekmath 010D} _{ix}f_{t}+{\Greekmath 011B} _{ix}u_{x,it},
\end{equation*}
where $f_{t}$ is a latent factor (could be talent in the context of return
to education) and ${\Greekmath 011B} _{ix}u_{x,it}$ is the idiosyncratic innovation to
the $x_{it}$ process (${\Greekmath 011B} _{ix}$ being income volatility in the saving
rate example). A simple model of correlated slope heterogeneity is given by
\begin{equation*}
{\Greekmath 010C} _{i}-{\Greekmath 010C} _{0}={\Greekmath 0114} _{{\Greekmath 010D} }\left[ {\Greekmath 010D} _{ix}-E\left( {\Greekmath 010D}
_{ix}\right) \right] +{\Greekmath 0114} _{{\Greekmath 011B} }\left[ {\Greekmath 011B} _{ix}^{2}-E\left(
{\Greekmath 011B} _{ix}^{2}\right) \right] +{\Greekmath 010F} _{i},
\end{equation*}
where ${\Greekmath 0114} _{{\Greekmath 010D} }$ and ${\Greekmath 0114} _{{\Greekmath 011B} }$ measure the extent to which
heterogeneity is correlated with $x_{it}$, and ${\Greekmath 010F} _{i}$ is the
exogenous component of heterogeneity. FE estimators are robust to exogenous
heterogeneity but can become badly biased if ${\Greekmath 0114} _{{\Greekmath 010D} }\neq 0$
and/or ${\Greekmath 0114} _{{\Greekmath 011B} }\neq 0$.
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Bias of the fixed effects estimator under correlated
heterogeneity}
It is known that fixed effects (with or without time effects) estimators are
biased under correlated heterogeneity. For example, see \cite{Wooldridge2005}
. Here we provide a minimal set of conditions required for the FE estimator
to be $\sqrt{n}$-consistent in the presence of correlated heterogeneity.
Under the heterogeneous specification (\ref{eq2}) and noting that $
\boldsymbol{M}_{T}\boldsymbol{{\Greekmath 011C} }_{T}=\boldsymbol{0}$, we have
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{{\Greekmath 010C} }_{0}=\boldsymbol{\bar{\Psi}}
_{n}^{-1}\left[ n^{-1}\sum_{i=1}^{n}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M
}_{T}\boldsymbol{X}_{i}(\boldsymbol{{\Greekmath 010C} }_{i}-\boldsymbol{{\Greekmath 010C} }_{0})
\right] +\boldsymbol{\bar{\Psi}}_{n}^{-1}\left( n^{-1}\sum_{i=1}^{n}
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{u}_{i}\right) .
\label{FE1}
\end{equation}
Then by Assumption \ref{errors}, $E\left( \boldsymbol{u}_{i}\left\vert
\boldsymbol{X}_{i}\right. \right) =\boldsymbol{0}$, and hence $E\left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{u}_{i}\right) =
\boldsymbol{0}$. Under Assumptions \ref{errors}, \ref{rcm} and \ref{PoolA},
\begin{equation*}
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{{\Greekmath 010C} }_{0}\rightarrow _{p}
\boldsymbol{\bar{\Psi}}^{-1}\lim_{n\rightarrow \infty
}n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) .
\end{equation*}
Hence, $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ is a consistent estimator of the
average effect, $\boldsymbol{{\Greekmath 010C} }_{0}$, only if
\begin{equation}
\lim_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }
_{i}\right) =\boldsymbol{0}. \label{AveFE}
\end{equation}
This condition is clearly met if
\begin{equation}
E\left[ \left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}
_{i}\right) \boldsymbol{{\Greekmath 0111} }_{i}\right] =\boldsymbol{0},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for all }i
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \label{FEconsistency}
\end{equation}
and has already been derived by \cite{Wooldridge2005}. But it is too
restrictive, since it is possible for the average condition in (\ref{AveFE})
to hold even though condition (\ref{FEconsistency}) is violated for some
units as $n\rightarrow \infty $. Suppose the number of units that \textit{do
not} satisfy (\ref{FEconsistency}) is given by $m_{n}=\ominus \left(
n^{a_{{\Greekmath 0111} }}\right) $, namely $m_{n}$ rises in line with $n^{a_{{\Greekmath 0111} }}$.
\footnote{
Note that $m_{n}=\ominus \left( n^{a_{{\Greekmath 0111} }}\right) $ differs from the
familiar big O order, $m_{n}=O(n^{a_{{\Greekmath 0111} }})$. The latter requires $
n^{-a_{{\Greekmath 0111} }}m_{n}$ is bounded in $n$ and its limiting value could be zero.
The former requires $n^{a_{{\Greekmath 0111} }}m_{n}$ to tend to a positive non-zero
number.} Then $n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) =\ominus
\left( n^{a_{{\Greekmath 0111} }-1}\right) $, and condition (\ref{AveFE}) is met if $
a_{{\Greekmath 0111} }<1$.
But for $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ to be a $\sqrt{n}$-consistent
estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$, a much more restrictive condition on
$a_{{\Greekmath 0111} }$ is required. Using (\ref{FE1}),
\begin{equation}
\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{{\Greekmath 010C} }_{0}\right) =
\boldsymbol{\bar{\Psi}}_{n}^{-1}\left( n^{-1/2}\sum_{i=1}^{n}\boldsymbol{
{\Greekmath 0110} }_{iT}+n^{-1/2}\sum_{i=1}^{n}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}
_{T}\boldsymbol{u}_{i}\right) , \label{AsyBiasFE}
\end{equation}
where $\boldsymbol{{\Greekmath 0110} }_{iT}=\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}
_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}$. Under Assumptions \ref{errors}
and \ref{PoolA}, as $n\rightarrow \infty $, $\boldsymbol{\bar{\Psi}}
_{n}\rightarrow _{p}\boldsymbol{\bar{\Psi}}\succ \boldsymbol{0}$, and $
n^{-1/2}\sum_{i=1}^{n}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{u}_{i}\rightarrow _{d}N\left( \boldsymbol{0},\boldsymbol{Q}
_{FE}\right) $, where $\boldsymbol{Q}_{FE}=\lim\limits_{n\rightarrow \infty }
\frac{1}{n}\sum_{i=1}^{n}E\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}
_{T}\boldsymbol{H}_{i}\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) \succ
\boldsymbol{0} $, for a fixed $T$. The first term in the brackets of (\ref
{AsyBiasFE}) can be decomposed as
\begin{equation*}
n^{-1/2}\sum_{i=1}^{n}\boldsymbol{{\Greekmath 0110} }_{iT}=\boldsymbol{s}
_{nT}+n^{-1/2}\sum_{i=1}^{n}E\left( \boldsymbol{{\Greekmath 0110} }_{iT}\right) ,
\end{equation*}
where $\boldsymbol{s}_{nT}=n^{-1/2}\sum_{i=1}^{n}\left[ \boldsymbol{{\Greekmath 0110} }
_{iT}-E\left( \boldsymbol{{\Greekmath 0110} }_{iT}\right) \right] $, and
\begin{equation*}
Var\left( \boldsymbol{s}_{nT}\right) =n^{-1}\left\Vert
\sum_{i=1}^{n}\sum_{j=1}^{n}Cov\left( \boldsymbol{{\Greekmath 0110} }_{iT},\boldsymbol{
{\Greekmath 0110} }_{jT}\right) \right\Vert \leq \sup_{j}\sum_{i=1}^{n}\left\Vert
Cov\left( \boldsymbol{{\Greekmath 0110} }_{iT},\boldsymbol{{\Greekmath 0110} }_{jT}\right)
\right\Vert ,
\end{equation*}
which is bounded by Assumption \ref{CrCorr}. It also follows that $
\boldsymbol{s}_{nT}$ tends to a limiting distribution with a zero mean by
construction. Therefore, $\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-
\boldsymbol{{\Greekmath 010C} }_{0}\right) $ will converge to a distribution with a zero
mean only if $n^{-1/2}\sum_{i=1}^{n}E\left( \boldsymbol{{\Greekmath 0110} }_{iT}\right)
\rightarrow \boldsymbol{0}$. For this condition to hold, it is required that
$m_{n}n^{-1/2}=\ominus \left( n^{a_{{\Greekmath 0111} }-1/2}\right) \rightarrow 0$, which
occurs only if $a_{{\Greekmath 0111} }<1/2$. A formal statement of this result is
summarized in the following proposition.
\begin{proposition}[Condition for $\protect\sqrt{n}-$consistency of the FE
estimator]
\label{prop_confee} Suppose that for $i=1,2,...,n$ and $t=1,2,...,T$, $
y_{it} $ are generated by the heterogeneous panel data model (\ref{m1}) and
Assumptions \ref{errors}--\ref{CrCorr} hold. Then the FE estimator given by (
\ref{fee}) is $\sqrt{n}$-consistent if
\begin{equation}
\lim_{n\rightarrow \infty }n^{-1/2}\sum_{i=1}^{n}E\left( \boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }
_{i}\right) =\boldsymbol{0}. \label{rootnFE}
\end{equation}
where $\boldsymbol{{\Greekmath 0111} }_{i}=\boldsymbol{{\Greekmath 010C} }_{i}-\boldsymbol{{\Greekmath 010C} }
_{0} $.
\end{proposition}
\begin{remark}
It is also interesting that the $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ continues to
be inconsistent even if both $n$ and $T\rightarrow \infty $, jointly. In
this case, using (\ref{FE1}) we have
\begin{equation*}
\sqrt{nT}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{{\Greekmath 010C} }_{0}\right)
=\boldsymbol{\bar{\Psi}}_{nT}^{-1}\left[ \sqrt{nT}\frac{1}{n}
\sum_{i=1}^{n}\left( \frac{1}{T}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}
_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) \right] +\boldsymbol{
\bar{\Psi}}_{nT}^{-1}\left( T^{-1/2}n^{-1/2}\sum_{i=1}^{n}\boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{u}_{i}\right) ,
\end{equation*}
where $\boldsymbol{\bar{\Psi}}_{nT}=\frac{1}{n}\sum_{i=1}^{n}\frac{1}{T}
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}$. Assuming $
\boldsymbol{\bar{\Psi}}_{nT}$ converges to a positive definite matrix as $n$
and $T\rightarrow \infty $ jointly, then for $\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$
to achieve the regular $\sqrt{nT}$ convergence rate, it is required that
\begin{equation*}
\sqrt{nT}n^{-1}\sum_{i=1}^{n}E\left( T^{-1}\boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right)
\rightarrow \boldsymbol{0}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, as }n\rightarrow \infty \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }
T\rightarrow \infty.
\end{equation*}
As before, suppose that $E\left( T^{-1}\boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) \neq
\boldsymbol{0}$ for $n^{a_{{\Greekmath 0111} }}$ of the $n$ cross-section units, and $
T=\ominus \left( n^{d}\right) $, for $d\geq 0$. Then the above condition is
met if $\sqrt{nT}n^{a_{{\Greekmath 0111} }-1}=\ominus \left( n^{a_{{\Greekmath 0111} }-1/2+d/2}\right)
\rightarrow 0$, i.e., if $a_{{\Greekmath 0111} }+d/2<1/2$. Thus, increasing the time
dimension of the panel does not help with bias reduction and as a matter of
fact accentuates it.
\end{remark}
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Mean group estimator}
In contrast to the FE estimator, the MG estimator given by (\ref{mge})
continues to be $\sqrt{n}$-consistent, so long as certain moment conditions,
to be discussed below, are met. Substituting (\ref{m1}) in (\ref{betaihat}),
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}=\boldsymbol{{\Greekmath 010C} }_{i}+\boldsymbol{{\Greekmath 0118} }_{iT},
\label{betaihat2}
\end{equation}
where $\boldsymbol{{\Greekmath 0118} }_{iT}=\boldsymbol{R}_{i}^{\prime }\boldsymbol{u}_{i}
$ and $\boldsymbol{R}_{i}=\boldsymbol{M}_{T}\boldsymbol{X}_{i}\left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right)
^{-1} $. Averaging both sides of (\ref{betaihat2}) over $i$ yields
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}=\boldsymbol{\bar{{\Greekmath 010C}}}_{n}+\boldsymbol{\bar{
{\Greekmath 0118}}}_{nT}, \label{mge2}
\end{equation}
where $\boldsymbol{\bar{{\Greekmath 010C}}}_{n}=\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{
{\Greekmath 010C} }_{i}$ and $\boldsymbol{\bar{{\Greekmath 0118}}}_{nT}=\frac{1}{n}\sum_{i=1}^{n}
\boldsymbol{{\Greekmath 0118} }_{iT}$. By Assumption \ref{errors}, $E\left( \boldsymbol{
\bar{{\Greekmath 0118}}}_{nT}\right) =E\left( \frac{1}{n}\sum_{i=1}^{n}\boldsymbol{{\Greekmath 0118} }
_{iT}\right) =\frac{1}{n}\sum_{i=1}^{n}E\left[ \boldsymbol{R}_{i}^{\prime
}E\left( \boldsymbol{u}_{i}\left\vert \boldsymbol{X}_{i}\right. \right)
\right] =\boldsymbol{0}$. Then using (\ref{mge2}), $E(\boldsymbol{\hat{{\Greekmath 010C}}
}_{MG})=E(\boldsymbol{\bar{{\Greekmath 010C}}}_{n})+E\left( \boldsymbol{\bar{{\Greekmath 0118}}}
_{nT}\right) =\boldsymbol{{\Greekmath 010C} }_{0}$, i.e., $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$
is an \textit{unbiased} estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ irrespective
of the possible dependence of $\boldsymbol{{\Greekmath 010C} }_{i}$ on $\boldsymbol{X}
_{i}$. However, the MG estimator is likely to have a large variance when $T$
is too small. This arises, for example, when the variance of $\boldsymbol{
\bar{{\Greekmath 0118}}}_{nT}$ does not exist or is very large. The conditions under which
$\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ converges to $\boldsymbol{{\Greekmath 010C} }_{0}$ at
the regular $\sqrt{n}$ rate are given in the following proposition:
\begin{proposition}[Sufficient conditions for $\protect\sqrt{n}$-consistency
of $\boldsymbol{\hat{\protect{\Greekmath 010C}}}_{MG}$]
\label{prop_mexist} Suppose that $y_{it}$ for $i=1,2,...,n$ and $t=1,2,...,T$
are generated by model (\ref{m1}) and Assumptions \ref{errors} and \ref{rcm}
hold. Then for a fixed $T$, as $n \rightarrow \infty$, the MG estimator
given by (\ref{mge}) is $\sqrt{n}$-consistent if
\begin{equation}
E\left( d_{i}^{-2}\right) <C\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }\func{sup}_{i}E\left[ \left\Vert
\func{adj}(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}
_{i})\right\Vert _{1}^{2}\right] <C, \label{sufficient}
\end{equation}
where $d_{i}=\func{det}(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{X}_{i})$, and $\func{adj}(\boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{X}_{i})$ is the adjugate (or adjoint) of $
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}$.
\end{proposition}
For a proof, see sub-section \ref{pf_prop_mexist} of the mathematical
appendix.
\begin{example}
\label{ex1} Consider the simple example $y_{it}={\Greekmath 010B} _{i}+{\Greekmath 010C}
_{i}x_{it}+u_{it}$, where $x_{it}={\Greekmath 010B} _{ix}+{\Greekmath 011B} _{ix}{\Greekmath 0122}
_{x,it}$, $\func{sup}_{i}\Vert {\Greekmath 010B} _{ix}\Vert <C$, $\inf_{i}{\Greekmath 011B}
_{ix}^{2}>c>0$, and ${\Greekmath 0122} _{x,it}\thicksim IIDN(0,1)$. Suppose that $
E(\boldsymbol{u}_{i}\boldsymbol{u}_{i}^{\prime }|\boldsymbol{x}_{i})={\Greekmath 011B}
_{i}^{2}\boldsymbol{I}_{T}$, for $i=1,2,...,n$. Then the individual OLS
estimator of the slope coefficient, $\hat{{\Greekmath 010C}}_{i}=(\boldsymbol{x}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{x}_{i})^{-1}\boldsymbol{x}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{y}_{i}$, has first- and
second-order moments if $E\left( u_{it}^{2}\right) <C$ and $E\left(
d_{i}^{-2}\right) <C$, where $d_{i}=\func{det}(\boldsymbol{x}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{x}_{i})$, $\boldsymbol{x}_{i}={\Greekmath 010B} _{ix}
\boldsymbol{{\Greekmath 011C} }_{T}+{\Greekmath 011B} _{ix}\boldsymbol{{\Greekmath 0122} }_{ix}$, and $
\boldsymbol{{\Greekmath 0122} }_{ix}=({\Greekmath 0122} _{x,i1},{\Greekmath 0122}
_{x,i2},...,{\Greekmath 0122} _{x,iT})^{\prime }$. In this case $1/d_{i}=(1/{\Greekmath 011B}
_{ix}^{2})\left( 1/\boldsymbol{{\Greekmath 0122} }_{ix}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{{\Greekmath 0122} }_{ix}\right) $, where $\boldsymbol{{\Greekmath 0122} }
_{ix}^{\prime }\boldsymbol{M}_{T}\boldsymbol{{\Greekmath 0122} }_{ix}\thicksim
{\Greekmath 011F} _{v}^{2}$, with $v=T-1$ degrees of freedom. Suppose further that $
{\Greekmath 011B} _{ix}^{2}$ and $\boldsymbol{{\Greekmath 0122} }_{ix}^{\prime }\boldsymbol{M}
_{T}\boldsymbol{{\Greekmath 0122} }_{ix}$ are independently distributed, then
\begin{equation*}
E\left( \frac{1}{d_{i}}\right) =E\left( \frac{1}{{\Greekmath 011B} _{ix}^{2}}\right)
E\left( \frac{1}{{\Greekmath 011F} _{v}^{2}}\right) =\frac{1}{v-2}E\left( \frac{1}{{\Greekmath 011B}
_{ix}^{2}}\right) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }v>2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.}
\end{equation*}
\begin{equation*}
E\left( \frac{1}{d_{i}^{2}}\right) =E\left( \frac{1}{{\Greekmath 011B} _{ix}^{4}}
\right) E\left( \frac{1}{{\Greekmath 011F} _{v}^{4}}\right) =\frac{1}{(v-2)(v-4)}E\left(
\frac{1}{{\Greekmath 011B} _{ix}^{4}}\right) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }v>4.
\end{equation*}
Since by assumption ${\Greekmath 011B} _{ix}^{-2}<1/c<\infty $, it follows that $
E\left( d_{i}^{-2}\right) $ exists if $T>5$. This example also illustrates
that $d_{i}$ could be random draws from a common distribution without
requiring $x_{it}$ to be IID. Note that no restrictions are placed on the
distribution of ${\Greekmath 010B} _{ix}$ over $i$.
\end{example}
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Relative efficiency of FE and MG estimators}
Suppose now that conditions (\ref{rootnFE}) and (\ref{sufficient}) hold and
both FE and MG estimators are $\sqrt{n}$-consistent. The choice between the
two estimators will then depend on their relative efficiency, which we
measure in terms of their covariances conditional on $\boldsymbol{X}=\left(
\boldsymbol{X}_{1},\boldsymbol{X}_{2},...,\boldsymbol{X}_{n}\right) $. We
have
\begin{equation*}
Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}\left\vert \boldsymbol{X}
\right. \right) =\boldsymbol{\Omega }_{{\Greekmath 010C} }+n^{-1}\sum_{i=1}^{n}
\boldsymbol{\Psi }_{i}^{-1}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{H}_{i}\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{\Psi }
_{i}^{-1},
\end{equation*}
and
\begin{equation*}
Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}\left\vert \boldsymbol{X}
\right. \right) =\boldsymbol{\bar{\Psi}}_{n}^{-1}\left( n^{-1}\sum_{i=1}^{n}
\boldsymbol{\Psi }_{i}\boldsymbol{\Omega }_{{\Greekmath 010C} }\boldsymbol{\Psi }
_{i}\right) \boldsymbol{\bar{\Psi}}_{n}^{-1}+\boldsymbol{\bar{\Psi}}
_{n}^{-1}\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M
}_{T}\boldsymbol{H}_{i}\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right)
\boldsymbol{\bar{\Psi}}_{n}^{-1},
\end{equation*}
where $\boldsymbol{\Omega }_{{\Greekmath 010C} }=Var(\boldsymbol{{\Greekmath 010C} }_{i}|\boldsymbol{
X})\succeq \boldsymbol{0}$, $\boldsymbol{H}_{i}=E\left( \boldsymbol{u}_{i}
\boldsymbol{u}_{i}^{\prime }|\boldsymbol{X}\right) $, and as before $
\boldsymbol{\Psi }_{i}=\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{X}_{i}$ and $\boldsymbol{\bar{\Psi}}_{n}=n^{-1}\sum_{i=1}^{n}
\boldsymbol{\Psi }_{i}$. Hence,
\begin{equation}
Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}\left\vert \boldsymbol{X}
\right. \right) -Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}\left\vert
\boldsymbol{X}\right. \right) =\boldsymbol{A}_{n}+\boldsymbol{B}_{n},
\label{vardif}
\end{equation}
where
\begin{equation}
\boldsymbol{A}_{n}=\boldsymbol{\Omega }_{{\Greekmath 010C} }-\boldsymbol{\bar{\Psi}}
_{n}^{-1}\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{\Psi }_{i}\boldsymbol{\Omega
}_{{\Greekmath 010C} }\boldsymbol{\Psi }_{i}\right) \boldsymbol{\bar{\Psi}}_{n}^{-1},
\label{An}
\end{equation}
and
\begin{equation}
\boldsymbol{B}_{n}=\left( \frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\Psi }
_{i}^{-1}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{H}_{i}
\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{\Psi }_{i}^{-1}\right) -
\boldsymbol{\bar{\Psi}}_{n}^{-1}\left( \frac{1}{n}\sum_{i=1}^{n}\boldsymbol{X
}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{H}_{i}\boldsymbol{M}_{T}
\boldsymbol{X}_{i}\right) \boldsymbol{\bar{\Psi}}_{n}^{-1}. \label{Bn}
\end{equation}
$\boldsymbol{A}_{n}$ and $\boldsymbol{B}_{n}$ capture the effects of two
different types of heterogeneity, namely slope heterogeneity and
regressors/errors heterogeneity. The superiority of the FE estimator over
the MG estimator is readily established when the slope coefficients and
error variances are homogeneous across $i$ and the errors are serially
uncorrelated, namely if $\boldsymbol{\Omega }_{{\Greekmath 010C} }=\boldsymbol{0}$ and $
\boldsymbol{H}_{i}={\Greekmath 011B} ^{2}\boldsymbol{I}_{T}$ for all $i$. In this case,
$\boldsymbol{A}_{n}=\boldsymbol{0}$, and we have
\begin{equation*}
\frac{Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}\left\vert \boldsymbol{X
}\right. \right) -Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}\left\vert
\boldsymbol{X}\right. \right) }{{\Greekmath 011B} ^{2}}=n^{-1}\sum_{i=1}^{n}\boldsymbol{
\Psi }_{i}^{-1}-\boldsymbol{\bar{\Psi}}_{n}^{-1},
\end{equation*}
which is the difference between the harmonic mean of $\boldsymbol{\Psi }_{i}$
and the inverse of its arithmetic mean and is a positive semi-definite
matrix.\footnote{
For a proof, see the Appendix to \cite{PesaranEtal1996}.} However, this
result may be reversed when we allow for heterogeneity, $\boldsymbol{\Omega }
_{{\Greekmath 010C} }\succ \boldsymbol{0}$, and/or if $\boldsymbol{H}_{i}\neq {\Greekmath 011B} ^{2}
\boldsymbol{I}_{T}$. The following proposition summarizes the results of the
comparison between FE and MG estimators.
\begin{proposition}[Relative efficiency of FE and MG estimators]
\label{prop_mgvsfe} Suppose that $y_{it}$ for $i=1,2,...,n$ and $t=1,2,...,T$
are generated by the heterogeneous panel data model (\ref{m1}), Assumptions
\ref{errors}--\ref{CrCorr} hold, and the uncorrelated heterogeneity
condition (\ref{rootnFE}) and second-order moment conditions (\ref
{sufficient}) are met. Then $Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}
_{MG}\left\vert \boldsymbol{X}\right. \right) -Var\left( \sqrt{n}\boldsymbol{
\hat{{\Greekmath 010C}}}_{FE}\left\vert \boldsymbol{X}\right. \right) =\boldsymbol{A}
_{n}+\boldsymbol{B}_{n}$, where $\boldsymbol{A}_{n}$ and $\boldsymbol{B}_{n}$
are given by (\ref{An}) and (\ref{Bn}), respectively. $\boldsymbol{A}_{n}$
is a negative semi-definite matrix, and the sign of$\ \boldsymbol{B}_{n}$ is
indeterminate. Under uncorrelated heterogeneity, the FE estimator, $
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$, is asymptotically more efficient than the MG
estimator if the benefit from pooling (i.e., when $\boldsymbol{B}_{n}\succ
\boldsymbol{0}$) outweighs the loss in efficiency due to slope heterogeneity
(since $\boldsymbol{A}_{n}\preceq \boldsymbol{0}$).
\end{proposition}
For a proof, see sub-section \ref{Proofmgvsfe} in the mathematical appendix.
\begin{example}
\label{ExampleMG-FE}Consider a simple case where $k^{\prime}=1$, $
\boldsymbol{\Psi }_{i}={\Greekmath 0120} _{i}$ and $\boldsymbol{\Omega }_{{\Greekmath 010C} }={\Greekmath 011B}
_{{\Greekmath 010C} }^{2}$ are scalars, and suppose that $\boldsymbol{H}_{i}(\boldsymbol{
X}_{i})=E\left( \boldsymbol{u}_{i}\boldsymbol{u}_{i}^{\prime }\left\vert
\boldsymbol{X}_{i}\right. \right) ={\Greekmath 011B} ^{2}{\Greekmath 0120} _{i}\boldsymbol{I}_{T}$,
then
\begin{equation*}
Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}\left\vert \boldsymbol{X}
\right. \right) -Var\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}\left\vert
\boldsymbol{X}\right. \right) =-\left( {\Greekmath 011B} _{{\Greekmath 010C} }^{2}+{\Greekmath 011B}
^{2}\right) \left[ n^{-1}\frac{\sum_{i=1}^{n}\left( {\Greekmath 0120} _{i}-\bar{{\Greekmath 0120}}
_{n}\right) ^{2}}{\bar{{\Greekmath 0120}}_{n}^{2}}\right] ,
\end{equation*}
where $\bar{{\Greekmath 0120}}_{n}=n^{-1}\sum_{i=1}^{n}{\Greekmath 0120} _{i}$. In this simple case,
the MG estimator is more efficient than the FE estimator even if ${\Greekmath 011B}
_{{\Greekmath 010C} }^{2}=0$.
\end{example}
In general, under uncorrelated heterogeneity, the relative efficiency of the
MG and FE estimators depends on the relative magnitude of the two components
in (\ref{vardif}). Since $\boldsymbol{A}_{n} \preceq \boldsymbol{0}$, the
outcome depends on the sign and the magnitude of $\boldsymbol{B}_{n}$, which
in turn depends on the heterogeneity of error variances, $\boldsymbol{H}_{i}(
\boldsymbol{X}_{i}$), and $\boldsymbol{\Psi }_{i}$ over $i$.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Irregular mean group estimators\label{IRMGE}}
So far, we have argued that the MG estimator is robust to correlated
heterogeneity and its performance is comparable to the FE estimator even
under uncorrelated heterogeneity. However, since the MG estimator is based
on the individual estimates, $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$, $i=1,2,...,n$,
its optimality and robustness critically depend on how well individual
coefficients can be estimated. This is particularly important when $T$ is
ultra short, which is the primary concern of this paper. In cases where $T$
is small and/or the observations on $\boldsymbol{x}_{it}$ are highly
correlated or are slowly moving, $d_{i}=\func{det}\left( \boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) $ is likely to be
close to zero for a large number of units $i=1,2,...,n$. As a result, $
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ is likely to be a poor estimate of $
\boldsymbol{{\Greekmath 010C} }_{i}$ for some $i$, and including such estimates when
computing $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ could be problematic, rendering
the MG estimator inefficient and unreliable.
However, as discussed above, $\boldsymbol{\hat{{\Greekmath 010C}}}_{MG}$ continues to be
an unbiased estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$, even if $\boldsymbol{
{\Greekmath 010C} }_{i}$ are correlated with $\boldsymbol{X}_{i}$, so long as the
stochastic component of $\boldsymbol{x}_{it}$ is strictly exogenous with
respect to $u_{it}$. By averaging over $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ for $
i=1,2,...,n$, as $n\rightarrow \infty $, the MG estimator converges to $
\boldsymbol{{\Greekmath 010C} }_{0}$ if $T$ is sufficiently large such that $\boldsymbol{
\hat{{\Greekmath 010C}}}_{i}$ have at least second-order moments for all $i$. The
existence of first-order moments of $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ is
required for the MG estimator to be unbiased, and we need $\boldsymbol{\hat{
{\Greekmath 010C}}}_{i}$ to have second-order moments for $\sqrt{n}$-consistent
estimation and valid inference about the average effects, $\boldsymbol{{\Greekmath 010C}
}_{0}$.
When individual estimates do not have second-order moments, trimming is
required. The question is how to trim the individual estimates $\boldsymbol{
\hat{{\Greekmath 010C}}}_{i}$. \cite{Peng2001} proposes to categorize the individual
estimates and then use a weighted average of the estimates depending on
their left and right tail shape parameters. He assumes the estimates are
identically and independently distributed and considers only the case of a
scalar parameter. \cite{GrahamPowell2012} propose to trim by exclusion,
namely dropping individual estimates with $d_{i}$ below a threshold value.
We establish that trimming and/or shrinkage is required only if the tail
index, ${\Greekmath 010B} _{p}$, of the distribution of $1/d_{i}$ is below or equal to $
2$, and propose to shrink only those estimates whose $d_{i}$ are below a
threshold value, $a_{n}$, when ${\Greekmath 010B} _{p}\leq 2$. Moreover, it is shown by
simulations that estimates of ${\Greekmath 010B} _{p}$ rise with $T$, and shrinkage is
typically required when $T<2k+1$.
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{A trimmed mean group estimator \label{TMGE}}
To formalize our proposed TMG estimator, we start with the least square
estimator of $\boldsymbol{{\Greekmath 010C} }_{i}$, namely $\boldsymbol{\hat{{\Greekmath 010C}}}
_{i}=(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})^{-1}
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{y}_{i}$, and
shrink it using the threshold function $\boldsymbol{1}\{d_{i}>a_{n}\}$,
which takes the value of unity if $d_{i}>a_{n}$ and zero otherwise, where $
d_{i}=\func{det}(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}
_{i})$. The threshold value is set as
\begin{equation}
a_{n}=C_{n}n^{-{\Greekmath 010B} }, \label{an}
\end{equation}
where ${\Greekmath 010B} >0$ and $C_{n}$ is a positive constant bounded in $n$. The
choice of ${\Greekmath 010B} $ and $C_{n}$ will be discussed below. The resultant
shrinkage estimator is given by
\begin{equation*}
\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}=\left\{
\begin{array}{ll}
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}=(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{X}_{i})^{-1}\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{y}_{i}, & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }d_{i}>a_{n}, \\
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}^{\ast }=a_{n}^{-1}\func{adj}(\boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})\boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{y}_{i}, & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }d_{i}\leq
a_{n},
\end{array}
\right.
\end{equation*}
where $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}^{\ast }$ is the shrinkage version of $
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$ since $a_{n}^{-1}\leq d_{i}^{-1}$. Written
more compactly, we have
\begin{equation}
\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}=\boldsymbol{1}\{d_{i}>a_{n}\}\boldsymbol{\hat{
{\Greekmath 010C}}}_{i}+\boldsymbol{1}\{d_{i}\leq a_{n}\}\boldsymbol{\hat{{\Greekmath 010C}}}
_{i}^{\ast }=(1+{\Greekmath 010E} _{i})\boldsymbol{\hat{{\Greekmath 010C}}}_{i}, \label{betai2}
\end{equation}
where
\begin{equation}
{\Greekmath 010E} _{i}=\left( \frac{d_{i}-a_{n}}{a_{n}}\right) \boldsymbol{1}
\{d_{i}\leq a_{n}\}\leq 0. \label{deltai}
\end{equation}
We considered two versions of TMG estimators, depending on how individual
trimmed estimators, $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}$, are combined. An
obvious choice is to use a simple average of $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}$
, namely $\overline{\boldsymbol{\tilde{{\Greekmath 010C}}}}_{n}=n^{-1}\sum_{i=1}^{n}
\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}=n^{-1}\sum_{i=1}^{n}(1+{\Greekmath 010E} _{i})
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$, which can also be viewed as a weighted
average estimator with the weights $w_{i}=(1+{\Greekmath 010E} _{i})/n<1/n$. But it is
easily seen that these weights do not add up to unity, and it might be
desirable to use the scaled weights $w_{i}/(1+\bar{{\Greekmath 010E}}
_{n})=n^{-1}(1+{\Greekmath 010E} _{i})/(1+\bar{{\Greekmath 010E}}_{n})$, where $\bar{{\Greekmath 010E}}
_{n}=n^{-1}\sum_{i=1}^{n}{\Greekmath 010E} _{i}$. Using these modified weights, we
propose the following TMG estimator
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}=n^{-1}\sum_{i=1}^{n}\left( \frac{1+{\Greekmath 010E} _{i}
}{1+\bar{{\Greekmath 010E}}_{n}}\right) \boldsymbol{\hat{{\Greekmath 010C}}}_{i}, \label{TMGb}
\end{equation}
which can be written more compactly as
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}=n^{-1}\sum_{i=1}^{n}\left( 1+\bar{{\Greekmath 010E}}
_{n}\right) ^{-1}\boldsymbol{Q}_{i}^{\prime }\boldsymbol{y}_{i},
\label{TMGc}
\end{equation}
where
\begin{equation}
\boldsymbol{Q}_{i}=\left( 1+{\Greekmath 010E} _{i}\right) \boldsymbol{R}_{i}=\left(
1+{\Greekmath 010E} _{i}\right) \boldsymbol{M}_{T}\boldsymbol{X}_{i}\left( \boldsymbol{X
}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) ^{-1}. \label{Qi}
\end{equation}
The TMG estimator differs in two important respects from the trimmed
estimator proposed by \cite{GrahamPowell2012}, which in the context of our
setup (and abstracting from time effects for now) can be written as
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{GP}=\frac{\sum_{i=1}^{n}\boldsymbol{1}
\{d_{i,GP}>h_{n}^{2}\}\boldsymbol{\hat{{\Greekmath 010C}}}_{i}}{\sum_{i=1}^{n}
\boldsymbol{1}\{d_{i,GP}>h_{n}^{2}\}}, \label{gpe}
\end{equation}
where $d_{i,GP}=\det (\boldsymbol{W}_{i}^{\prime }\boldsymbol{W}_{i})$, and $
\boldsymbol{W}_{i}=\left( \boldsymbol{{\Greekmath 011C} }_{T},\boldsymbol{X}_{i}\right) $
. In the special case where $T=k$, $d_{i,GP}=\left\vert \func{det}\left(
\boldsymbol{W}_{i}\right) \right\vert ^{2}$ and the threshold function
considered by GP reduces to $\boldsymbol{1}\{\left\vert \func{det}\left(
\boldsymbol{W}_{i}\right) \right\vert >h_{n}\}$.
The GP approach can be viewed as trimming by exclusion, which overlooks the
information that might be contained in $\func{adj}(\boldsymbol{W}
_{i}^{\prime }\boldsymbol{W}_{i})$ when $\left\vert \func{det}\left(
\boldsymbol{W}_{i}\right) \right\vert \leq h_{n}$. More specifically, to
relate $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ to the GP estimator given by (\ref
{gpe}), using (\ref{betai2}) in (\ref{TMGb}), we note that
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}=\frac{1-{\Greekmath 0119} _{n}}{1+\bar{{\Greekmath 010E}}_{n}}\left(
\frac{\sum_{i=1}^{n}\boldsymbol{1}\{d_{i}>a_{n}\}\boldsymbol{\hat{{\Greekmath 010C}}}_{i}
}{\sum_{i=1}^{n}\boldsymbol{1}\{d_{i}>a_{n}\}}\right) +\frac{{\Greekmath 0119} _{n}}{1+
\bar{{\Greekmath 010E}}_{n}}\left( \frac{\sum_{i=1}^{n}\boldsymbol{1}\{d_{i}\leq a_{n}\}
\boldsymbol{\hat{{\Greekmath 010C}}}_{i}^{\ast }}{\sum_{i=1}^{n}\boldsymbol{1}
\{d_{i}\leq a_{n}\}}\right), \label{betaTn}
\end{equation}
where ${\Greekmath 0119} _{n}$ is the fraction of the estimates being trimmed given by
\begin{equation}
{\Greekmath 0119} _{n}=\frac{\sum_{i=1}^{n}\boldsymbol{1}\{d_{i}\leq a_{n}\}}{n}.
\label{pin}
\end{equation}
Compared to $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$, the GP estimator places zero
weights on the estimates with $d_{i}\leq a_{n}$, and leaves out the scaling
factor, $\left( 1+\bar{{\Greekmath 010E}}_{n}\right) ^{-1}$, which, as already noted,
can play an important role in the small sample performance of the TMG
estimator. Our proposed method also differs from GP in the way we motivate
and calibrate the threshold function.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Asymptotic properties of the TMG estimator\label{AsyTMG}}
To investigate the asymptotic properties of the TMG estimator, $\boldsymbol{
\hat{{\Greekmath 010C}}}_{TMG}$, we introduce the following additional assumptions:
\begin{assumption}[moments]
\label{regressorsx} For $i=1,2,...,n$, denote by $d_{i}=\func{det}\left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) $,
where $\boldsymbol{X}_{i}=(\boldsymbol{x}_{i1},\boldsymbol{x}_{i2},...,
\boldsymbol{x}_{iT})^{\prime }$ is the $T\times k^{\prime }$ matrix of
observations on $\boldsymbol{x}_{it}$ in the heterogeneous panel data model
given by (\ref{m1}) and $\boldsymbol{M}_{T}=\boldsymbol{I}_{T}-\boldsymbol{
{\Greekmath 011C} }_{T}\boldsymbol{{\Greekmath 011C} }_{T}^{\prime }/T$. $\func{inf}_{i}\left(
d_{i}\right) >0$, almost surely, and $\func{sup}_{i}E\left[ \left\Vert \func{
adj}\left( \boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}
_{i}\right) \right\Vert ^{2}\right] <C$, where $\func{adj}\left( \boldsymbol{
X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) $ is the
adjugate of $\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}
_{i} $.
\end{assumption}
\begin{assumption}[distribution of $d_{i}$]
\label{distributiondi} For $i=1,2,...,n$, $d_{i}$ are random draws from the
probability distribution function, $F_{d}(u)$, with a continuously
differentiable density function, $f_{d}(u)$, over $u\in (0,\infty )$, such
that $F_{d}(0)=0$, and $F_{d}(a_{n})\thicksim C_{f}a_{n}^{{\Greekmath 010B} _{p}}$, as $
n\rightarrow \infty $, where $C_{f}$ is bounded in $n$, and ${\Greekmath 010B} _{p}>0$.
\end{assumption}
\begin{remark}
\label{Pareto}Parameter ${\Greekmath 010B} _{p}$ in Assumption \ref{distributiondi} is
related to the tail index of Pareto-type distributions of $z=1/d$ given by
\footnote{
Pareto-type distributions cover a wide class of distributions including
Pareto, Cauchy, Burr and Stable distributions with exponent ${\Greekmath 010B} _{p}<2$.
See Example 3.3.10 (p.133) in \cite{EmbrechtsEtal1997}.}
\begin{equation*}
\Pr \left( z>b_{n}\right) \thicksim C_{p}b_{n}^{-{\Greekmath 010B} _{p}}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, }
b_{n}\rightarrow \infty .
\end{equation*}
For $d>0$,
\begin{equation*}
\Pr \left( z>b_{n}\right) =\Pr \left( 1/d>b_{n}\right) =\Pr \left(
d<1/b_{n}\right) =F_{d}\left( b_{n}^{-1}\right) \thicksim
C_{f}b_{n}^{-{\Greekmath 010B} _{p}}.
\end{equation*}
Connecting ${\Greekmath 010B} _{p}$ to the shape parameter of the Pareto-type
distributions of $1/d$ allows us to utilize the rich and extensive
literature that already exists on estimation of the tail index. See, for
example, Section 6.4.2 of \cite{EmbrechtsEtal1997}
\end{remark}
\begin{remark}
Note that when ${\Greekmath 010B} _{p}>2$, $E\left( d_{i}^{-2}\right) <C$, condition (
\ref{sufficient}) will be met and as a result the MG estimator becomes $
\sqrt{n}$-consistent. Trimming is only required if ${\Greekmath 010B} _{p}\leq 2$.
\end{remark}
\begin{remark}
\label{Fu} By the mean value theorem, $F_{d}(a_{n})=
\int_{0}^{a_{n}}f_{d}(u)du=f_{d}(\bar{a}_{n})a_{n}$, where $\bar{a}_{n}$
lies on the segment $(0,a_{n})$. Hence, under Assumption \ref{distributiondi}
, it follows that $f_{d}(\bar{a}_{n})=O\left( a_{n}^{{\Greekmath 010B} _{p}-1}\right) $.
\end{remark}
\begin{remark}
Assumption \ref{distributiondi}, by imposing $F_{d}(0)=0$, rules out the
case where $d_{i}=0$ for some units. In practice, units with $d_{i}=0$ can
be excluded when computing the TMG estimates. This case arises, for example,
when $k^{\prime }=1$ and $T=2$, and $d_{i}=\func{det}(\boldsymbol{x}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{x}_{i})=\frac{1}{2}\left(
x_{i2}-x_{i1}\right) ^{2}=0$, for some units. Units with $x_{i2}=x_{i1}$ are
known as stayers, and those with very small values of $\left\vert
x_{i2}-x_{i1}\right\vert $ (below a given threshold value) are known as slow
movers and discussed by GP, and more recently by \cite{SasakiUra2026} (SU).
Our proposed TMG estimator does exploit the information on slow movers,
namely those with $a_{n}\geq d_{i}>0$, and excludes units with $d_{i}=0$
only. Due to the error cross-sectional independence, we conjecture that the
TMG estimator will perform well even if there are stayers that are excluded
from our analysis, so long as the fraction of stayers in the sample is not
too large. In cases where stayers form a large fraction of the units, the
TMG estimate can be augmented with an estimate for the stayers computed
using the extrapolation approach proposed by SU. Following such a mixed
strategy, the extrapolation error must be balanced against the loss of
efficiency associated with using TMG that excludes the stayers.
\end{remark}
\begin{assumption}[$\boldsymbol{\protect{\Greekmath 010C} }_{i}$ and $d_{i}$ dependence]
\label{CRE} Dependence of $\boldsymbol{{\Greekmath 010C} }_{i}=({\Greekmath 010C} _{i1},{\Greekmath 010C}
_{i2},...,{\Greekmath 010C} _{ik^{\prime }})^{\prime }$ on $d_{i}$ is characterized by
\begin{equation}
\boldsymbol{{\Greekmath 010C} }_{i}=\boldsymbol{{\Greekmath 010C} }_{0}+\boldsymbol{{\Greekmath 0111} }_{i},
\label{etai}
\end{equation}
where $\boldsymbol{{\Greekmath 0111} }_{i}$ are independently distributed over $i$,
\begin{equation}
\boldsymbol{{\Greekmath 0111} }_{i}=\boldsymbol{B}_{i}\left\{ \boldsymbol{g}(d_{i})-E
\left[ \boldsymbol{g}(d_{i})\right] \right\} +\boldsymbol{{\Greekmath 010F} }_{i},
\label{etai2}
\end{equation}
$E(\boldsymbol{{\Greekmath 010F} }_{i}|d_{i})=\boldsymbol{0}$, $\func{sup}
_{i}E\left\Vert \boldsymbol{{\Greekmath 010F} }_{i}\right\Vert ^{4}<C$, $\boldsymbol{g
}(u)=(g_{1}(u),g_{2}(u),...,g_{k^{\prime }}(u))^{\prime }$, and $g_{j}(u)$
for $j=1,2,...,k^{\prime }$ are bounded and continuously differentiable
functions of $u$ on $(0,\infty )$. $\boldsymbol{B}_{i}$ are bounded $
k^{\prime }\times k^{\prime }$ matrices of fixed constants with $\func{sup}
_{i}\left\Vert \boldsymbol{B}_{i}\right\Vert <C$.
\end{assumption}
\begin{remark}
Under Assumption \ref{distributiondi}, $d_{i}$ are distributed independently
over $i$, which implies that ${\Greekmath 010E} _{i}$, defined by (\ref{deltai}), are
also distributed independently over $i$. The cross-sectional independence
assumption is made to simplify the mathematical exposition. It can be
relaxed by requiring $\boldsymbol{{\Greekmath 0111} }_{i}$ and ${\Greekmath 010E} _{i}$ to be weakly
cross-sectionally correlated.
\end{remark}
Using (\ref{betaihat2}) and (\ref{etai}) in (\ref{betai2}), we have
\begin{equation*}
\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}=(1+{\Greekmath 010E} _{i})\boldsymbol{{\Greekmath 010C} }
_{0}+(1+{\Greekmath 010E} _{i})\left( \boldsymbol{{\Greekmath 0111} }_{i}+\boldsymbol{{\Greekmath 0118} }
_{iT}\right) ,
\end{equation*}
and $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ defined by (\ref{TMGb}) can be written
as
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}-\boldsymbol{{\Greekmath 010C} }_{0}=\left( \frac{1}{1+
\bar{{\Greekmath 010E}}_{n}}\right) n^{-1}\sum_{i=1}^{n}(1+{\Greekmath 010E} _{i})\left(
\boldsymbol{{\Greekmath 0111} }_{i}+\boldsymbol{{\Greekmath 0118} }_{iT}\right) . \label{TMG_egzi}
\end{equation}
(\ref{TMG_egzi}) can be written equivalently as
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}-\boldsymbol{{\Greekmath 010C} }_{0}=\left( \frac{
1+E\left( \bar{{\Greekmath 010E}}_{n}\right) }{1+\bar{{\Greekmath 010E}}_{n}}\right) \left(
\boldsymbol{b}_{n}+n^{-1}\sum_{i=1}^{n}\left[ \boldsymbol{p}_{i}-E\left(
\boldsymbol{p}_{i}\right) \right] +n^{-1}\sum_{i=1}^{n}\boldsymbol{q}
_{iT}\right) , \label{TMGgap1}
\end{equation}
where
\begin{equation}
\boldsymbol{b}_{n}=n^{-1}\sum_{i=1}^{n}E\left( \boldsymbol{p}_{i}\right)
=n^{-1}\sum_{i=1}^{n}\frac{E({\Greekmath 010E} _{i}\boldsymbol{{\Greekmath 0111} }_{i})}{1+E\left(
\bar{{\Greekmath 010E}}_{n}\right) },\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ } \label{bn}
\end{equation}
\begin{equation}
\boldsymbol{p}_{i}=\left[ 1+E\left( \bar{{\Greekmath 010E}}_{n}\right) \right]
^{-1}\left( 1+{\Greekmath 010E} _{i}\right) \boldsymbol{{\Greekmath 0111} }_{i} \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }
\boldsymbol{q}_{iT}=\left[ 1+E\left( \bar{{\Greekmath 010E}}_{n}\right) \right]
^{-1}\left( 1+{\Greekmath 010E} _{i}\right) \boldsymbol{{\Greekmath 0118} }_{iT}. \label{piqiT}
\end{equation}
Also, by Lemma \ref{deltaeta}, $E({\Greekmath 010E} _{i})=O(a_{n}^{{\Greekmath 010B} _{p}})$, $
E\left( \bar{{\Greekmath 010E}}_{n}\right) =O(a_{n}^{{\Greekmath 010B} _{p}})$, $E\left( {\Greekmath 010E}
_{i}\boldsymbol{{\Greekmath 0111} }_{i}\right) =O(a_{n}^{{\Greekmath 010B} _{p}})$, and hence
\begin{equation*}
\boldsymbol{b}_{n}=\frac{1}{1+E\left( \bar{{\Greekmath 010E}}_{n}\right) }\left[
n^{-1}\sum_{i=1}^{n}E({\Greekmath 010E} _{i}\boldsymbol{{\Greekmath 0111} }_{i})\right] =\frac{
O(a_{n}^{{\Greekmath 010B} _{p}})}{1+O(a_{n}^{{\Greekmath 010B} _{p}})}=O(a_{n}^{{\Greekmath 010B}
_{p}})=O(n^{-{\Greekmath 010B} {\Greekmath 010B} _{p}}).
\end{equation*}
Since under Assumptions \ref{regressorsx} and \ref{distributiondi}, ${\Greekmath 010E}
_{i}-E\left( {\Greekmath 010E} _{i}\right) $ is distributed independently over $i$ with
a zero mean and bounded variance, then
\begin{equation}
\frac{1+E\left( \bar{{\Greekmath 010E}}_{n}\right) }{1+\bar{{\Greekmath 010E}}_{n}}=1-\frac{\bar{
{\Greekmath 010E}}_{n}-E\left( \bar{{\Greekmath 010E}}_{n}\right) }{1+E\left( \bar{{\Greekmath 010E}}
_{n}\right) +\left( \bar{{\Greekmath 010E}}_{n}-E\left( \bar{{\Greekmath 010E}}_{n}\right) \right)
}=1+O_{p}\left( n^{-1/2-{\Greekmath 010B} {\Greekmath 010B} _{p}}\right) . \label{deltaOrder}
\end{equation}
Similarly, under Assumptions \ref{regressorsx}, \ref{distributiondi} and \ref
{CRE}, $\boldsymbol{p}_{i}-E\left( \boldsymbol{p}_{i}\right) $ is
distributed independently over $i$ with zero means and bounded variances,
and we have
\begin{equation*}
\frac{1}{n}\sum_{i=1}^{n}\left[ \boldsymbol{p}_{i}-E\left( \boldsymbol{p}
_{i}\right) \right] =\left[ 1+E\left( \bar{{\Greekmath 010E}}_{n}\right) \right]
^{-1}\left\{ \frac{1}{n}\sum_{i=1}^{n}\boldsymbol{{\Greekmath 0111} }_{i}+\frac{1}{n}
\sum_{i=1}^{n}\left[ {\Greekmath 010E} _{i}\boldsymbol{{\Greekmath 0111} }_{i}-E\left( {\Greekmath 010E} _{i}
\boldsymbol{{\Greekmath 0111} }_{i}\right) \right] \right\} =O_{p}(n^{-1/2}).
\end{equation*}
Consider now $\boldsymbol{\bar{q}}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{q}
_{iT}$, where $\boldsymbol{q}_{iT}$ is defined by (\ref{piqiT}), and note
that
\begin{equation}
\boldsymbol{\bar{q}}_{nT}=\left( \frac{1}{1+E\left( \bar{{\Greekmath 010E}}_{n}\right) }
\right) \boldsymbol{\bar{{\Greekmath 0118}}}_{{\Greekmath 010E} ,nT}, \label{qbarnt}
\end{equation}
where $\boldsymbol{\bar{{\Greekmath 0118}}}_{{\Greekmath 010E} ,nT}=n^{-1}\sum_{i=1}^{n}\left(
1+{\Greekmath 010E} _{i}\right) \boldsymbol{{\Greekmath 0118} }_{iT}$. Using results of Lemma \ref
{EVegzi}, we have $E\left( \boldsymbol{\bar{q}}_{nT}\right) =\boldsymbol{0}$
, and for any ${\Greekmath 010B} _{p}>0$,
\begin{equation*}
Var\left( \boldsymbol{\bar{q}}_{nT}\right) =\left( \frac{1}{1+E\left( \bar{
{\Greekmath 010E}}_{n}\right) }\right) ^{2}Var\left( \boldsymbol{\bar{{\Greekmath 0118}}}_{{\Greekmath 010E}
,nT\ }\right) =O(n^{-1}a_{n}^{-1})+O(n^{-1}a_{n}^{-1+{\Greekmath 010B}
_{p}/2})=O(n^{-1}a_{n}^{-1}).
\end{equation*}
Hence, $\boldsymbol{\bar{q}}_{nT}=O_{p}\left( n^{-1/2+{\Greekmath 010B} /2}\right) $.
Using the above results in (\ref{TMGgap1}), we have
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}-\boldsymbol{{\Greekmath 010C} }_{0}=O(n^{-{\Greekmath 010B} {\Greekmath 010B}
_{p}})+O_{p}\left( n^{-\frac{(1-{\Greekmath 010B} )}{2}}\right) . \label{TMGgap2}
\end{equation}
Thus, $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ converges to $\boldsymbol{{\Greekmath 010C} }_{0}$
asymptotically so long as $0<{\Greekmath 010B} <1$ as $n\rightarrow \infty $. The
convergence rate of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ to $\boldsymbol{{\Greekmath 010C} }
_{0}$ will depend on the trade-off between the bias and variance of $
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$. Though it is possible to reduce the bias of
$\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ by choosing a value of ${\Greekmath 010B} $ close to
unity, it will be at the expense of a larger variance. In what follows, we
shed light on the choice of ${\Greekmath 010B} $ by considering the conditions under
which the asymptotic distribution of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is
centered around $\boldsymbol{{\Greekmath 010C} }_{0}$ and at the same time, its
asymptotic variance tends to zero at a reasonably fast rate.
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{The choice of the trimming threshold value \label{Threshold}}
We begin by assuming that the rate at which $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$
converges to $\boldsymbol{{\Greekmath 010C} }_{0}$ is given by $n^{{\Greekmath 010D} }$, where $
{\Greekmath 010D} $ is set in relation to ${\Greekmath 010B} $. Since it is not guaranteed that
the individual estimates, $\boldsymbol{\hat{{\Greekmath 010C}}}_{i}$, have second order
moments when when $T$ is ultra short, we expect the rate, $n^{{\Greekmath 010D} }$, to
be below the regular rate of $n^{1/2}$. Using (\ref{TMGgap1}) and (\ref
{deltaOrder}) and noting that ${\Greekmath 010D} \leq 1/2$ (with equality holding only
under regular convergence), we have
\begin{equation}
n^{{\Greekmath 010D} }\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}-\boldsymbol{{\Greekmath 010C} }
_{0}\right) =n^{{\Greekmath 010D} }\boldsymbol{b}_{n}+n^{{\Greekmath 010D} -\frac{1-{\Greekmath 010B} }{2}}
\left[ n^{-\frac{1+{\Greekmath 010B} }{2}}\sum_{i=1}^{n}\left[ \boldsymbol{p}
_{i}-E\left( \boldsymbol{p}_{i}\right) \right] +n^{-\frac{1+{\Greekmath 010B} }{2}
}\sum_{i=1}^{n}\boldsymbol{q}_{iT}\right] +o_{p}(1). \label{Asy3}
\end{equation}
To ensure that the asymptotic distribution of $\boldsymbol{\hat{{\Greekmath 010C}}}
_{TMG} $ is correctly centered, we must have $n^{{\Greekmath 010D} }\boldsymbol{b}
_{n}\rightarrow \boldsymbol{0,}$ as $n\rightarrow \infty $. Since $n^{{\Greekmath 010D}
}\boldsymbol{b}_{n}=O(n^{{\Greekmath 010D} }a_{n}^{{\Greekmath 010B} _{p}})=O(n^{{\Greekmath 010D} -{\Greekmath 010B}
{\Greekmath 010B} _{p}})$, this condition is ensured if ${\Greekmath 010D} <{\Greekmath 010B} {\Greekmath 010B} _{p}$.
Turning to the second term of the above, we note that to obtain a
non-degenerate distribution, we also need to set ${\Greekmath 010D} =\left( 1-{\Greekmath 010B}
\right) /2$. Combining these two requirements yields $\left( 1-{\Greekmath 010B}
\right) /2<{\Greekmath 010B} {\Greekmath 010B} _{p}$, or
\begin{equation}
{\Greekmath 010B} >\frac{1}{1+2{\Greekmath 010B} _{p}}. \label{calpha}
\end{equation}
In view of (\ref{TMGgap2}), the convergence rate of $\boldsymbol{\hat{{\Greekmath 010C}}}
_{TMG}-\boldsymbol{{\Greekmath 010C} }_{0}$, namely $n^{-\frac{(1-{\Greekmath 010B} )}{2}}$,
depends on ${\Greekmath 010B} _{p}$, which governs the tail property of the
distribution of $1/d_{i}$. In the simple example \ref{ex1}, assuming ${\Greekmath 011B}
_{ix}^{2}={\Greekmath 011B} _{x}^{2}$, then $1/d_{i}=(1/{\Greekmath 011B} _{x}^{2})(1/{\Greekmath 011F}
_{v}^{2})$, and the shape parameter of $1/d_{i}$ will equal the shape
parameter of $1/{\Greekmath 011F} _{v}^{2}$, which is given by ${\Greekmath 010B} _{p}=v/2$. In the
simple case where $k^{\prime }=1$, we have ${\Greekmath 010B}_{p}=(T-1)/2$. In the
worse case scenario with $T=2=k$, ${\Greekmath 010B} _{p}=1/2$, and the convergence
rate of $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is at most be $n^{-1/4}$, well
below the regular convergence rate of $n^{-1/2}$. The optimum choice of $
{\Greekmath 010B} $ depends on ${\Greekmath 010B} _{p}$, which in turn is determined by the rate
at which $F_{d}(a_{n})$ tends to zero as $a_{n}\rightarrow 0$.
To ensure that the threshold function, $\boldsymbol{1}\{d_{i}>C_{n}n^{-
{\Greekmath 010B} }\}$, does not depend on the scale of $\boldsymbol{x}_{it}$, we
suggest setting $C_{n}=\bar{d}_{n}=n^{-1}\sum_{i=1}^{n}d_{i}$. With this
choice, the threshold function can be written as $\boldsymbol{1}
\{d_{i}>C_{n}n^{-{\Greekmath 010B} }\}=\boldsymbol{1}\{d_{i}/\bar{d}_{n}>n^{-{\Greekmath 010B} }\}$
, noting that $d_{i}/\bar{d}_{n}$ is scale free. The selection of ${\Greekmath 010B} $
is more complicated. In practice, a two-step procedure can be implemented,
whereby in the first step the value of ${\Greekmath 010B} _{p}$ is estimated using the
observations $\left\{ 1/d_{i}\right\} _{i=1}^{n}$.\footnote{
Asymptotically, the estimate of ${\Greekmath 010B} _{p}$ does not depend on the scale
of $d_{i}$.} Then TMG estimation can be carried out using $\hat{{\Greekmath 010B}}
=1/(1+2\hat{{\Greekmath 010B}}_{p})+{\Greekmath 010F} $, where $\hat{{\Greekmath 010B}}_{p}$ is a
consistent estimator of ${\Greekmath 010B} _{p}$, and ${\Greekmath 010F} $ is a small positive
constant. See also Remark \ref{Pareto}.
Given the uncertainty associated with the estimates of ${\Greekmath 010B} _{p}$, we
also considered setting ${\Greekmath 010B} _{p}=1$, and using the threshold function $
\boldsymbol{1}\{d_{i}/\bar{d}_{n}>n^{-1/3}\}$. This approach has the
advantage of being simple to implement. The small sample performance of
these two approaches to setting ${\Greekmath 010B} $ are compared using MC experiments.
See sub-section \ref{MChatalpha} of the online supplement. The simulation
results show that the simple threshold rule $\boldsymbol{1}\{d_{i}/\bar{d}
_{n}>n^{-1/3}\}$ performs better when $T$ is ultra-short ($T=k$), and has a
similar performance when the estimated value of $\hat{{\Greekmath 010B}}_{p}$ is used
if $T>k$.
GP show that to correctly center the limiting distribution of their proposed
estimator, $\boldsymbol{\hat{{\Greekmath 010C}}}_{GP}$ given by (\ref{gpe}), $h_{n}$ in
their threshold function $\boldsymbol{1}\{|\det (\boldsymbol{W}
_{i})|>h_{n}\} $ must be set as $h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}$, such that $
(nh_{n})^{1/2}h_{n}\rightarrow 0$, as $n\rightarrow \infty $, which implies
that ${\Greekmath 010B} _{GP}>1/3$ (see p. 2125 and p. 2138 of GP). SU adopt the same
choice and set $h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}n^{-L/(1+2L)}$, with $L\in
\{1,2\}$ being the order of local polynomials used to infer the average
effects of stayers (see sub-section 5.2 on p. 16 of SU).
\begin{theorem}[Asymptotic distribution of TMG estimator]
\label{thm_asytmg} Suppose for $i=1,2,...,n$ and $t=1,2,...,T$, $y_{it}$ are
generated by the heterogeneous panel data model (\ref{m1}), and Assumptions
\ref{errors}, \ref{rcm}, and \ref{regressorsx}--\ref{CRE} hold. Then for $
{\Greekmath 010B} >1/(1+2{\Greekmath 010B} _{p})$, where ${\Greekmath 010B} _{p}\in (0,2]$ is the shape
parameter of the tail probability distribution of $1/d_{i}$ over $i$,
\begin{equation*}
n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}-\boldsymbol{{\Greekmath 010C} }
_{0}\right) \rightarrow _{d}N\left( \boldsymbol{0}_{k^{\prime }},\boldsymbol{
V}_{{\Greekmath 010C} }\right) ,\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ as }n\rightarrow \infty ,
\end{equation*}
where $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is given by (\ref{TMGb}), and
\begin{equation}
\boldsymbol{V}_{{\Greekmath 010C} }=Avar\left( n^{(1-{\Greekmath 010B} )/2}\boldsymbol{\hat{{\Greekmath 010C}}}
_{TMG}\right) =C^{-1}\lim_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E\left[
a_{n}\boldsymbol{1}\{d_{i}>a_{n}\}\boldsymbol{R}_{i}^{\prime }\boldsymbol{H}
_{i}\boldsymbol{R}_{i}\right] , \label{Varbeta}
\end{equation}
where $\boldsymbol{H}_{i}=E\left( \boldsymbol{u}_{i}\boldsymbol{u}
_{i}^{\prime }\left\vert \boldsymbol{X}_{i}\right. \right) $, $\boldsymbol{R}
_{i}=\boldsymbol{M}_{T}\boldsymbol{X}_{i}\left( \boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) ^{-1}$, $d_{i}=\func{det}\left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) >0$,
$a_{n}=C_{n}n^{-{\Greekmath 010B} }$, and $C=\lim_{n\rightarrow \infty
}C_{n}=\lim_{n\rightarrow \infty }\bar{d}_{n}>0$.
\end{theorem}
For a proof, see sub-section \ref{Proofthm1} of the mathematical appendix.
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Robust estimation of the covariance matrix of the TMG estimator}
As with standard MG estimation, consistent estimation of $\boldsymbol{V}
_{{\Greekmath 010C} }$ using (\ref{Varbeta}) requires knowledge of $\boldsymbol{H}_{i}$
which cannot be estimated consistently when $T$ is short. We follow the
literature and propose a robust covariance estimator of $\boldsymbol{V}
_{{\Greekmath 010C} }$, which is asymptotically unbiased for a wide class of error
variances, $E\left( \boldsymbol{u}_{i}\boldsymbol{u}_{i}^{\prime }\left\vert
\boldsymbol{X}_{i}\right. \right) =\boldsymbol{H}_{i}(\boldsymbol{X}_{i})$,
thus allowing for serially correlated and conditionally heteroskedastic
errors. The following theorem summarizes the main result.
\begin{theorem}[Robust covariance matrix of TMG estimator]
\label{VarCon}Suppose Assumptions \ref{errors}, \ref{rcm}, and \ref
{regressorsx}-\ref{CRE} hold and $\boldsymbol{{\Greekmath 010C} }_{0}$ is estimated by $
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ given by (\ref{TMGb}). Then for ${\Greekmath 010B}
>1/(1+2{\Greekmath 010B} _{p})$, where ${\Greekmath 010B} _{p}\in (0,2]$ is the shape parameter of
the tail probability distribution of $1/d_{i}$ over $i$, as $n\rightarrow
\infty $,
\begin{equation*}
\boldsymbol{V}_{{\Greekmath 010C} }=\func{plim}_{n\rightarrow \infty }\left[
n^{-1}\sum_{i=1}^{n}\left( \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{\hat{
{\Greekmath 010C}}}_{TMG}\right) \left( \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{\hat{
{\Greekmath 010C}}}_{TMG}\right) ^{\prime }\right] ,
\end{equation*}
where $\boldsymbol{V}_{{\Greekmath 010C} }=Avar\left( n^{(1-{\Greekmath 010B} )/2}\boldsymbol{\hat{
{\Greekmath 010C}}}_{TMG}\right) $ is defined by (\ref{Varbeta}). Accordingly, $Var(
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG})$ can be consistently estimated by $
n^{-2}\sum_{i=1}^{n}\left( \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{\hat{
{\Greekmath 010C}}}_{TMG}\right) \left( \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{\hat{
{\Greekmath 010C}}}_{TMG}\right) ^{\prime }$, where $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}$ is
given by (\ref{betai2}).
\end{theorem}
For a proof, see sub-section \ref{ProofVarCon} of the mathematical appendix.
\begin{remark}
Following the literature on MG estimation, here we also consider the
following bias-adjusted and scaled version:
\begin{equation}
\widehat{Var(\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG})}=\frac{1}{n(n-1)(1+\bar{{\Greekmath 010E}}
_{n})^{2}}\sum_{i=1}^{n}\left( \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{
\hat{{\Greekmath 010C}}}_{TMG}\right) \left( \boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{
\hat{{\Greekmath 010C}}}_{TMG}\right) ^{\prime }. \label{varC}
\end{equation}
\end{remark}
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Panels with time effects\label{CRCTE}}
The panel data model with time effects is given by
\begin{equation}
y_{it}={\Greekmath 010B} _{i}+{\Greekmath 011E} _{t}+\boldsymbol{x}_{it}^{\prime }\boldsymbol{{\Greekmath 010C} }
_{i}+u_{it}, \label{panTE1}
\end{equation}
where ${\Greekmath 011E} _{t}$ for $t=1,2,...,T$ are the time effects. Without loss of
generality, we adopt the normalization $\boldsymbol{{\Greekmath 011C} }_{T}^{\prime }
\boldsymbol{{\Greekmath 011E} }=0$, where $\boldsymbol{{\Greekmath 011E} }=({\Greekmath 011E} _{1},{\Greekmath 011E}
_{2},...,{\Greekmath 011E} _{T})^{\prime }$. To estimate $\boldsymbol{{\Greekmath 010C} }_{0}$,
initially we suppose $\boldsymbol{{\Greekmath 011E} }$ is known, and the trimmed
estimator of $\boldsymbol{{\Greekmath 010C} }_{i}$ is now given by $\boldsymbol{\tilde{
{\Greekmath 010C}}}_{i}(\boldsymbol{{\Greekmath 011E} })=\boldsymbol{Q}_{i}^{\prime }(\boldsymbol{y}
_{i}-\boldsymbol{{\Greekmath 011E} })=\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}-\boldsymbol{Q}
_{i}^{\prime }\boldsymbol{{\Greekmath 011E} }$, where $\boldsymbol{Q}_{i}$ is defined by (
\ref{Qi}). The associated TMG-TE estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ then
follows as
\begin{equation*}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG-TE}(\boldsymbol{{\Greekmath 011E} })=n^{-1}\sum_{i=1}^{n}
\left( 1+\bar{{\Greekmath 010E}}_{n}\right) ^{-1}\boldsymbol{\tilde{{\Greekmath 010C}}}_{i}(
\boldsymbol{{\Greekmath 011E} })=\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}-\boldsymbol{\bar{Q}}
_{n}^{\prime }\boldsymbol{{\Greekmath 011E} },
\end{equation*}
where $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ is given by (\ref{TMGc}), and
\begin{equation}
\boldsymbol{\bar{Q}}_{n}=\frac{1}{1+\bar{{\Greekmath 010E}}_{n}}\left(
n^{-1}\sum_{i=1}^{n}\boldsymbol{Q}_{i}\right) . \label{Qn}
\end{equation}
From our earlier analysis, it is clear that for a known $\boldsymbol{{\Greekmath 011E} }$
, $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG-TE}(\boldsymbol{{\Greekmath 011E} })$ has the same
asymptotic distribution as $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$ with $
\boldsymbol{y}_{i}$ replaced by $\boldsymbol{y}_{i}-\boldsymbol{{\Greekmath 011E} }$.
When $T>k$, we can follow \cite{Chamberlain1992} and eliminate the time
effects by the de-meaning transformation $\boldsymbol{M}_{i}=\boldsymbol{I}
_{T}-\boldsymbol{M}_{T}\boldsymbol{X}_{i}(\boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{X}_{i})^{-1}\boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}$. Under the normalization $\boldsymbol{{\Greekmath 011C} }_{T}^{\prime
}\boldsymbol{{\Greekmath 011E} }=0$, we have $\boldsymbol{M}_{T}\boldsymbol{{\Greekmath 011E} }=
\boldsymbol{{\Greekmath 011E} }$, and$\ \boldsymbol{M}_{T}\boldsymbol{y}_{i}=\boldsymbol{M
}_{T}\boldsymbol{X}_{i}\boldsymbol{{\Greekmath 010C} }_{i}+\boldsymbol{{\Greekmath 011E} }+
\boldsymbol{M}_{T}\boldsymbol{u}_{i}$. Then $\boldsymbol{M}_{i}\boldsymbol{M}
_{T}\boldsymbol{y}_{i}=\boldsymbol{M}_{i}\boldsymbol{{\Greekmath 011E} }+\boldsymbol{M}
_{i}\boldsymbol{M}_{T}\boldsymbol{u}_{i}$, and averaging over $i$ we obtain
\begin{equation}
n^{-1}\sum_{i=1}^{n}\boldsymbol{M}_{i}\boldsymbol{M}_{T}\boldsymbol{y}
_{i}=\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{M}_{i}\right) \boldsymbol{{\Greekmath 011E} }
+n^{-1}\sum_{i=1}^{n}\boldsymbol{M}_{i}\boldsymbol{M}_{T}\boldsymbol{u}_{i}.
\label{TE-C}
\end{equation}
Hence, $\boldsymbol{{\Greekmath 011E} }$ can be estimated without knowing $\boldsymbol{
{\Greekmath 010C} }_{0}$, if $\boldsymbol{\bar{M}}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{M}
_{i}$ is a positive definite matrix. This requires $T>k$, since $\boldsymbol{
\bar{M}}_{n}$ is singular if $T=k$. Therefore, to implement the
Chamberlain's estimation approach, we require the following assumption:
\begin{assumption}[identification of time effects]
\label{invertm} For $T>k$, $\boldsymbol{\bar{M}}_{n}=n^{-1}\sum_{i=1}^{n}
\boldsymbol{M}_{i}\rightarrow _{p}\boldsymbol{M}\succ \boldsymbol{0}$, where
$\boldsymbol{M}_{i}=\boldsymbol{I}_{T}-\boldsymbol{M}_{T}\boldsymbol{X}_{i}(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})^{-1}
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}$.
\end{assumption}
Under this Assumption, $\boldsymbol{{\Greekmath 011E} }$ can be estimated by
\begin{equation}
\boldsymbol{\hat{{\Greekmath 011E}}}_{C}=\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{M}
_{i}\right) ^{-1}\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{M}_{i}\boldsymbol{M}
_{T}\boldsymbol{y}_{i}\right) , \label{TE2}
\end{equation}
and its asymptotic distribution follows straightforwardly. Specifically,
using (\ref{TE-C}) we have
\begin{equation}
\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 011E}}}_{C}-\boldsymbol{{\Greekmath 011E} }\right) =
\boldsymbol{\bar{M}}_{n}^{-1}\left( n^{-1/2}\sum_{i=1}^{n}\boldsymbol{M}_{i}
\boldsymbol{M}_{T}\boldsymbol{u}_{i}\right) , \label{TE-Ca}
\end{equation}
and $\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 011E}}}_{C}-\boldsymbol{{\Greekmath 011E} }
_{0}\right) \rightarrow _{d}N(\boldsymbol{0},\boldsymbol{V}_{{\Greekmath 011E} ,C})$,
where $\boldsymbol{V}_{{\Greekmath 011E} ,C}=Avar\left( \sqrt{n}\boldsymbol{\hat{{\Greekmath 011E}}}
_{C}\right) $ is given by
\begin{equation*}
\boldsymbol{V}_{{\Greekmath 011E} ,C}=\boldsymbol{M}^{-1}\lim_{n\rightarrow \infty
}E\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{M}_{i}\boldsymbol{M}_{T}\boldsymbol{
u}_{i}\boldsymbol{u}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{M}
_{i}\right) \boldsymbol{M}^{-1}.
\end{equation*}
Since $\boldsymbol{M}_{i}\boldsymbol{M}_{T}\boldsymbol{u}_{i}=\boldsymbol{M}
_{i}\boldsymbol{M}_{T}(\boldsymbol{y}_{i}-\boldsymbol{{\Greekmath 011E} })$, the
asymptotic variance of $\boldsymbol{\hat{{\Greekmath 011E}}}_{C}$ can be consistently
estimated by
\begin{equation}
\widehat{\boldsymbol{V}}_{{\Greekmath 011E} ,C}=\boldsymbol{\bar{M}}_{n}^{-1}\left[
n^{-1}\sum_{i=1}^{n}\boldsymbol{M}_{i}\boldsymbol{M}_{T}(\boldsymbol{y}_{i}-
\boldsymbol{\hat{{\Greekmath 011E}}}_{C})(\boldsymbol{y}_{i}-\boldsymbol{\hat{{\Greekmath 011E}}}
_{C})^{\prime }\boldsymbol{M}_{T}\boldsymbol{M}_{i}\right] \boldsymbol{\bar{M
}}_{n}^{-1}. \label{Varphi2}
\end{equation}
Using $\boldsymbol{\hat{{\Greekmath 011E}}}_{C}$, the TMG-TE estimator of $\boldsymbol{
{\Greekmath 010C} }_{0}$ is now given by
\begin{equation}
\boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}=\frac{1}{1+\bar{{\Greekmath 010E}}_{n}}\left[
n^{-1}\sum_{i=1}^{n}\boldsymbol{Q}_{i}^{\prime }(\boldsymbol{y}_{i}-
\boldsymbol{\hat{{\Greekmath 011E}}}_{C})\right] \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }T>k, \label{BetaTE2}
\end{equation}
and
\begin{equation}
n^{\frac{1-{\Greekmath 010B} }{2}}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}-
\boldsymbol{{\Greekmath 010C} }_{0}\right) =n^{\frac{1-{\Greekmath 010B} }{2}}\left( \boldsymbol{
\hat{{\Greekmath 010C}}}_{C,TMG-TE}(\boldsymbol{{\Greekmath 011E} })-\boldsymbol{{\Greekmath 010C} }_{0}\right)
-n^{-{\Greekmath 010B} /2}\boldsymbol{\bar{Q}}_{n}^{\prime }\sqrt{n}\left( \boldsymbol{
\hat{{\Greekmath 011E}}}_{C}-\boldsymbol{{\Greekmath 011E} }\right) , \label{AsyDTE-C1}
\end{equation}
where $\boldsymbol{\bar{Q}}_{n}^{\prime }\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 011E}
}}_{C}-\boldsymbol{{\Greekmath 011E} }\right) =O_{p}(1)$. Hence
\begin{equation}
n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}-\boldsymbol{
{\Greekmath 010C} }_{0}\right) =n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}}}
_{C,TMG-TE}(\boldsymbol{{\Greekmath 011E} })-\boldsymbol{{\Greekmath 010C} }_{0}\right)
+O_{p}(n^{-{\Greekmath 010B} /2}). \label{AsyDTE-C}
\end{equation}
Also since ${\Greekmath 010B} >0$, then $n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}}
}_{C,TMG-TE}-\boldsymbol{{\Greekmath 010C} }_{0}\right) $ has the same asymptotic
distribution as $n^{(1-{\Greekmath 010B} )/2}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}(
\boldsymbol{{\Greekmath 011E} })-\boldsymbol{{\Greekmath 010C} }_{0}\right) $, with $\boldsymbol{{\Greekmath 011E}
}$ treated as known. This result follows since $\boldsymbol{\hat{{\Greekmath 011E}}}_{C}-
\boldsymbol{{\Greekmath 011E} }\rightarrow _{p}\boldsymbol{0}$ at the faster rate of $
\sqrt{n}$, compared with the rate $n^{\frac{1-{\Greekmath 010B} }{2}}$ of $\boldsymbol{
\hat{{\Greekmath 010C}}}_{C,TMG-TE}-\boldsymbol{{\Greekmath 010C} }_{0}\rightarrow _{p}\boldsymbol{0}
$. In short, $Avar\left( n^{(1-{\Greekmath 010B} )/2}\boldsymbol{\hat{{\Greekmath 010C}}}
_{C,TMG-TE}\right) $ is not affected by the estimation uncertainty of the
time effects when ${\Greekmath 010B} >0$. A consistent estimator is given by
\begin{equation}
\widehat{Var\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}\right) }=\frac{1}{
n(n-1)(1+\bar{{\Greekmath 010E}}_{n})^{2}}\sum_{i=1}^{n}\left( \boldsymbol{\tilde{{\Greekmath 010C}}
}_{i,C}-\boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}\right) \left( \boldsymbol{\tilde{
{\Greekmath 010C}}}_{i,C}-\boldsymbol{\hat{{\Greekmath 010C}}}_{C,TMG-TE}\right) ^{\prime },
\label{VarbetaC}
\end{equation}
where $\boldsymbol{\tilde{{\Greekmath 010C}}}_{i,C}=\boldsymbol{Q}_{i}^{\prime }(
\boldsymbol{y}_{i}-\boldsymbol{\hat{{\Greekmath 011E}}}_{C})$.
When $T=k$, Assumption \ref{invertm} does not hold and standard de-meaning
techniques can not be used to eliminate $\boldsymbol{{\Greekmath 011E} }$. This in turn
requires the correlation of slope coefficients with the regressors to be
time-invariant, a condition that holds trivially under uncorrelated
heterogeneity but could be restrictive when heterogeneity is correlated. The
TMG-TE estimator for the case of $T= k$ and its asymptotic properties are
derived in Section \ref{Teqk} of the mathematical appendix.
For the Monte Carlo experiments and the empirical application, $\boldsymbol{
\hat{{\Greekmath 011E}}}_{C}$, given by (\ref{TE2}), is used as an estimator of $
\boldsymbol{{\Greekmath 011E} }$ for panels with time effects when $T>k$. When $T=k$, we
use the estimator set out in Section \ref{Teqk} of the mathematical appendix.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{A test of correlated heterogeneity\label{Test}}
As summarized by Proposition \ref{prop_confee}, $\sqrt{n}$-consistency of
the FE estimator requires the slope coefficients, $\boldsymbol{{\Greekmath 010C} }_{i}$,
and the regressors, $\boldsymbol{X}_{i}=\left( \boldsymbol{x}_{i1},
\boldsymbol{x}_{i2},...,\boldsymbol{x}_{iT}\right) ^{\prime }$, to be
independently distributed. As a diagnostic check, we consider testing the
null hypothesis
\begin{equation}
H_{0}:\boldsymbol{{\Greekmath 0111} }_{i}\perp \boldsymbol{X}_{i}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }E\left(
\boldsymbol{{\Greekmath 0111} }_{i}\right) =\boldsymbol{0}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for all }i,
\label{null}
\end{equation}
where $\boldsymbol{{\Greekmath 0111} }_{i}=\boldsymbol{{\Greekmath 010C} }_{i}-\boldsymbol{{\Greekmath 010C} }
_{0} $. To test $H_{0}$, we follow \cite{Hausman1978} and compare FE and TMG
estimators given by (\ref{fee}) and (\ref{TMGb}), respectively. Under $H_{0}$
, both estimators are consistent. But, as required by Hausman tests, only
the TMG estimator is consistent under the alternative
\begin{equation}
H_{1}:E\left( n^{-1}\sum_{i=1}^{n}\boldsymbol{\bar{\Psi}}_{n}^{-1}
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\boldsymbol{
{\Greekmath 0111} }_{i}\right) =\ominus \left( n^{a_{{\Greekmath 0111} }-1}\right). \label{H1}
\end{equation}
The Hausman test of $H_{0}$ is based on the difference $\sqrt{n}\boldsymbol{
\hat{\Delta}}_{{\Greekmath 010C} }=\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}\right) $. Such a test has been considered by
\cite{PesaranEtal1996} and \cite{PesaranYamagata2008}, assuming the MG
estimator has at least second-order moments.\footnote{
See pp. 160--162 of \cite{PesaranEtal1996} and p. 53 of \cite
{PesaranYamagata2008}.} Here we extend this test to cover cases when $T$ is
ultra short and the moment condition (\ref{sufficient}) is not met. Also,
the earlier tests were derived under the null of homogeneity (namely $
\boldsymbol{{\Greekmath 0111} }_{i}=\boldsymbol{0}$ for all $i$), whilst the null that we
consider is more general and covers the null of homogeneity as a special
case.
In the development of his test, Hausman originally assumed that one of the
estimators under consideration is asymptotically efficient and showed that
in this case, the asymptotic variance of the difference between the two
estimators is equal to the difference between the two variances. But in the
present application, as shown in Proposition \ref{prop_mgvsfe} and
illustrated by Example \ref{ExampleMG-FE}, neither of the two estimators, $
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}$ and $\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}$, are
efficient, and the asymptotic variance of $\sqrt{n}\boldsymbol{\hat{\Delta}}
_{{\Greekmath 010C} }$ will not be equal to the difference between $Avar\left( \sqrt{n}
\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}\right) $ and $Avar\left( \sqrt{n}\boldsymbol{
\hat{{\Greekmath 010C}}}_{FE}\right) $. See also Section 26.9.1 of \cite{Pesaran2015}.
The following theorem summarizes our main result for the Hausman test
applied to $\sqrt{n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{\hat{
{\Greekmath 010C}}}_{TMG}\right) $.
\begin{theorem}
\label{AsyDTest} Suppose for $i=1,2,...,n$ and $t=1,2,...,T$, $y_{it}$ are
generated by the heterogeneous panel data model (\ref{m1}) and Assumptions
\ref{errors}--\ref{CRE} hold. Then under $H_{0}$, defined by (\ref{null}),
and for a fixed $T\geq k$, $\sqrt{n}\boldsymbol{\hat{\Delta}}_{{\Greekmath 010C} }=\sqrt{
n}\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}\right)
\rightarrow _{d}N(\boldsymbol{0},\boldsymbol{V}_{\Delta })$, as $
n\rightarrow \infty $, where
\begin{equation*}
\boldsymbol{V}_{\Delta }=\lim_{n\rightarrow \infty }\frac{1}{n}
\sum_{i=1}^{n}E\left( \boldsymbol{\Gamma }_{i}\boldsymbol{X}_{i}^{\prime }
\boldsymbol{M}_{T}\boldsymbol{H}_{i}\boldsymbol{M}_{T}\boldsymbol{X}_{i}
\boldsymbol{\Gamma }_{i}\right) +\lim_{n\rightarrow \infty }\frac{1}{n}
\sum_{i=1}^{n}E\left( \boldsymbol{\Gamma }_{i}\boldsymbol{\Psi }_{i}
\boldsymbol{\Omega }_{{\Greekmath 010C} }\boldsymbol{\Psi }_{i}\boldsymbol{\Gamma }
_{i}\right) ,
\end{equation*}
$\boldsymbol{\Psi }_{i}=\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{X}_{i}$, $\boldsymbol{H}_{i}=$ $E(\boldsymbol{u}_{i}\boldsymbol{u
}_{i}^{\prime }\left\vert \boldsymbol{X}_{i}\right. )$, $\boldsymbol{\Gamma }
_{i}=\boldsymbol{\bar{\Psi}}_{n}^{-1}-\left( \frac{1+{\Greekmath 010E} _{i}}{1+\bar{
{\Greekmath 010E}}_{n}}\right) \boldsymbol{\Psi }_{i}^{-1}$, $\boldsymbol{\bar{\Psi}}
_{n}=\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\Psi }_{i}$, $\boldsymbol{\Omega }
_{{\Greekmath 010C} } = Var(\boldsymbol{{\Greekmath 010C} }_{i}|\boldsymbol{X}_{i})\succeq
\boldsymbol{0}$, and ${\Greekmath 010E} _{i}$ is defined by (\ref{deltai}). The
asymptotic variance matrix, $\boldsymbol{V}_{\Delta }$, is positive definite
if either $\lim\limits_{n\rightarrow \infty }\frac{1}{n}\sum_{i=1}^{n}E
\left( \boldsymbol{\Gamma }_{i}^{2}\right) \succ \boldsymbol{0}$, and/or $
\lim\limits_{n\rightarrow \infty }\frac{1}{n}\sum_{i=1}^{n}E\left(
\boldsymbol{\Gamma }_{i}\boldsymbol{\Psi }_{i}\boldsymbol{\Psi }_{i}
\boldsymbol{\Gamma }_{i}\right) \succ \boldsymbol{0}$ and ${\Greekmath 0115} _{\min
}\left( \boldsymbol{\Omega }_{{\Greekmath 010C} }\right) >c>0$. When one of these
conditions is met, under $H_{0}$ and for a fixed $T\geq k$, $H_{{\Greekmath 010C} }=n
\boldsymbol{\hat{\Delta}}_{{\Greekmath 010C} }^{\prime }\boldsymbol{V}_{\Delta }^{-1}
\boldsymbol{\hat{\Delta}}_{{\Greekmath 010C} }\rightarrow _{d}{\Greekmath 011F} _{k^{\prime }}^{2}$
as $n\rightarrow \infty $. Under the alternative hypothesis, $H_{1}$ defined
by (\ref{H1}), $H_{{\Greekmath 010C} }=\ominus (n^{2a_{{\Greekmath 0111} }-1})+O(n^{a_{{\Greekmath 0111} }-{\Greekmath 010B}
{\Greekmath 010B} _{p}})+O\left( n^{1-2{\Greekmath 010B} {\Greekmath 010B} _{p}}\right) $, and $H_{{\Greekmath 010C}
}\rightarrow \infty $ as $n\rightarrow \infty $, if $a_{{\Greekmath 0111} }>1/2$.
\end{theorem}
For a proof, see sub-section \ref{ProofAsyDTest} in the mathematical
appendix.
\begin{remark}
The proposed test of correlated heterogeneity is consistent, and its power
rises not only with $a_{{\Greekmath 0111} }$, but is also enhanced from the correlation
between $\boldsymbol{{\Greekmath 0111} }_{i}$ and ${\Greekmath 010E} _{i}$, induced from the
shrinkage of some of the individual estimators used in construction of the
TMG estimator (see equation (\ref{meuB}) in the mathematical appendix). But
under $H_{0}$, $\boldsymbol{{\Greekmath 0111} }_{i}$ and $\boldsymbol{X}_{i}$ are
independently distributed, which implies that $\boldsymbol{{\Greekmath 0111} }_{i}$ and $
{\Greekmath 010E} _{i}$ are also independently distributed, and shrinkage of individual
estimates does not lead to size distortions.
\end{remark}
To implement the $H_{{\Greekmath 010C} }$ test, the asymptotic variance, $\boldsymbol{V}
_{\Delta }$, can be rewritten as $\boldsymbol{V}_{\Delta
}=n^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}\sum_{t^{\prime }=1}^{T}E\left(
\boldsymbol{g}_{it}\boldsymbol{g}_{it^{\prime }}^{\prime }\tilde{{\Greekmath 0117}}_{it}
\tilde{{\Greekmath 0117}}_{it^{\prime }}\right) $, where $\tilde{{\Greekmath 0117}}_{it}={\Greekmath 0117} _{it}-\bar{
{\Greekmath 0117}}_{i\circ }$, $\bar{{\Greekmath 0117}}_{i\circ } = \sum_{t=1}^{T}
\sum_{t^{\prime }=1}^{T}{\Greekmath 0117}_{it}$, and $\boldsymbol{g}_{it}$ is the $t^{th}$ column of $
\boldsymbol{G}_{i}$, where (see also (\ref{GiApp}) in the mathematical
appendix)
\begin{equation}
\boldsymbol{G}_{i}=\boldsymbol{X}_{i}\left[ \left( n^{-1}\sum_{i=1}^{n}
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right)
^{-1}-\left( \frac{1+{\Greekmath 010E} _{i}}{1+\bar{{\Greekmath 010E}}_{n}}\right) \left(
\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i}\right) ^{-1}
\right]. \label{Gi}
\end{equation}
For $T$ fixed, a consistent estimator of $\boldsymbol{V}_{\Delta }$, which
is robust to the choices of $\boldsymbol{H}_{i}$ and $\boldsymbol{\Omega }
_{{\Greekmath 010C} }$, is
\begin{equation}
\boldsymbol{\widehat{V}}_{\Delta }=\frac{1}{n}\sum_{i=1}^{n}\sum_{t=1}^{T}
\sum_{t^{\prime }=1}^{T}\boldsymbol{g}_{it}\boldsymbol{g}_{it^{\prime
}}^{\prime }\hat{\tilde{{\Greekmath 0117}}}_{it}\hat{\tilde{{\Greekmath 0117}}}_{it^{\prime }},
\label{VarHbeta}
\end{equation}
where $\hat{\tilde{{\Greekmath 0117}}}_{it}=(y_{it}-\bar{y}_{i\circ })-\boldsymbol{\hat{
{\Greekmath 010C}}}_{FE}^{\prime }(\boldsymbol{x}_{it}-\boldsymbol{\bar{x}}_{i\circ })$,
$\bar{y}_{i\circ }=\frac{1}{T}\sum_{t=1}^{T}y_{it}$, and $\boldsymbol{\bar{x}
}_{i\circ }=\frac{1}{T}\sum_{t=1}^{T}\boldsymbol{x}_{it}$. Using $
\boldsymbol{\widehat{V}}_{\Delta }$, the Hausman test statistic for
uncorrelated slope heterogeneity is given by
\begin{equation}
\hat{H}_{{\Greekmath 010C} }=n\left( \boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{\hat{{\Greekmath 010C}
}}_{TMG}\right) ^{\prime }\boldsymbol{\widehat{V}}_{\Delta }^{-1}\left(
\boldsymbol{\hat{{\Greekmath 010C}}}_{FE}-\boldsymbol{\hat{{\Greekmath 010C}}}_{TMG}\right) .
\label{htest}
\end{equation}
Under $H_{0}$, the consistency of $\boldsymbol{\widehat{V}}_{\Delta }$ as an
estimator of $\boldsymbol{V}_{\Delta }$ follows from the $\sqrt{n}$
-consistency of the FE estimator.
An extension of the $\hat{H}_{{\Greekmath 010C} }$ test statistic to panel data models
with time effects is provided in Section \ref{TestTE} of the online
supplement.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Trimmed mean group estimation for a subset of coefficients}
\label{subset}
The TMG approach can be applied to a subset of coefficients of interest. Let
$\boldsymbol{X}_{i}=(\boldsymbol{X}_{i1},\boldsymbol{X}_{i2})$, where $
\boldsymbol{X}_{i1}$ is the $T\times p$ matrix of observations on the focal
(treatment) variables and $\boldsymbol{X}_{i2}$ is the $T\times (k^{\prime
}-p)$ matrix of observations on the auxiliary (control) variables, and
partition the coefficients accordingly as $\boldsymbol{{\Greekmath 010C} }_{i}=(
\boldsymbol{{\Greekmath 010C} }_{i1}^{\prime },\boldsymbol{{\Greekmath 010C} }_{i2}^{\prime
})^{\prime }$, where $\boldsymbol{{\Greekmath 010C} }_{i1}$ is the $p\times 1$ vector of
the coefficients of interest. Using results from the partitioned
regressions, we have
\begin{equation*}
\boldsymbol{\hat{{\Greekmath 010C}}}_{i1}=(\boldsymbol{X}_{i1}^{\prime }\boldsymbol{M}
_{i2}\boldsymbol{X}_{i1})^{-1}\boldsymbol{X}_{i1}^{\prime }\boldsymbol{M}
_{i2}\boldsymbol{y}_{i},
\end{equation*}
where $\boldsymbol{M}_{i2}=\boldsymbol{I}_{T}-\boldsymbol{X}_{i2}\left(
\boldsymbol{X}_{i2}^{\prime }\boldsymbol{X}_{i2}\right) ^{-}\boldsymbol{X}
_{i2}^{\prime }$, and $\left( \boldsymbol{X}_{i2}^{\prime }\boldsymbol{X}
_{i2}\right) ^{-}$ denotes a generalized inverse of $\left( \boldsymbol{X}
_{i2}^{\prime }\boldsymbol{X}_{i2}\right) $.\footnote{
Note that $\boldsymbol{\hat{{\Greekmath 010C}}}_{i1}$ is invariant to the choice of the
generalized inverse, $\left( \boldsymbol{X}_{i2}^{\prime }\boldsymbol{X}
_{i2}\right) ^{-}$.}
The TMG estimator of $\boldsymbol{{\Greekmath 010C} }_{01}=E\left( \boldsymbol{{\Greekmath 010C} }
_{i1}\right) $ is given by
\begin{equation*}
\boldsymbol{\hat{{\Greekmath 010C}}}_{1,TMG}=n^{-1}\sum_{i=1}^{n}\left( \frac{1+{\Greekmath 010E}
_{i1}}{1+\bar{{\Greekmath 010E}}_{n1}}\right) \boldsymbol{\hat{{\Greekmath 010C}}}_{i1},
\end{equation*}
where ${\Greekmath 010E} _{i1}=\left( \frac{d_{i1}-a_{n1}}{a_{n1}}\right) \boldsymbol{1}
\{d_{i1}\leq a_{n1}\}$, $\bar{{\Greekmath 010E}}_{n1}=\frac{1}{n}\sum_{i=1}^{n}{\Greekmath 010E}
_{i1}$, and $d_{i1}=\det \left( \boldsymbol{X}_{i1}^{\prime }\boldsymbol{M}
_{i2}\boldsymbol{X}_{i1}\right) $. Similarly, for the choice of threshold
value, we consider $a_{n1}=\bar{d}_{n1}n^{-{\Greekmath 010B} _{1}}$, where ${\Greekmath 010B} _{1}>
\frac{1}{1+2{\Greekmath 010B} _{1,p}}$, ${\Greekmath 010B} _{1,p}$ is the shape parameter of the
tail probability distribution of $1/d_{i1}$ over $i$, and $\bar{d}
_{n1}=n^{-1}\sum_{i=1}^{n}d_{i1}$.
The above partitioned formula can also be adapted for testing general linear
restrictions $H_{0}:\boldsymbol{R{\Greekmath 010C} }_{0}=\boldsymbol{r}$ against $H_{1}:
\boldsymbol{R{\Greekmath 010C} }_{0}\neq \boldsymbol{r}$, where $\boldsymbol{r}$ is a $
p\times 1$ vector and $\boldsymbol{R}$ is a $p\times k^{\prime }$ ($
p<k^{\prime }$) full rank matrix of fixed constants. Let $\boldsymbol{{\Greekmath 010D}
}_{i}=$ $\boldsymbol{R{\Greekmath 010C} }_{i}-\boldsymbol{r}$ and partition $\boldsymbol{
R=(R}_{1},\boldsymbol{R}_{2})$ such that $\boldsymbol{R}_{1}$ is a $p\times
p $ non-singular matrix.\footnote{
This can be achieved by a suitable reordering of the elements of $
\boldsymbol{{\Greekmath 010C} }_{i}$.} Then (\ref{m1}) can be written as
\begin{equation*}
\boldsymbol{\tilde{y}}_{i}={\Greekmath 010B} _{i}\boldsymbol{{\Greekmath 011C} }_{T}+\boldsymbol{
\tilde{X}}_{i1}\boldsymbol{{\Greekmath 010D} }_{i}+\boldsymbol{\tilde{X}}_{i2}
\boldsymbol{{\Greekmath 010C} }_{i2}+\boldsymbol{u}_{i},
\end{equation*}
where $\boldsymbol{\tilde{y}}_{i}=\boldsymbol{y}_{i}-\boldsymbol{X}_{i1}
\boldsymbol{R}_{1}^{-1}\boldsymbol{r}$, $\boldsymbol{\tilde{X}}_{i1}=
\boldsymbol{X}_{i1}\boldsymbol{R}_{1}^{-1}$, and $\boldsymbol{\tilde{X}}
_{i2}=\boldsymbol{X}_{i2}-\boldsymbol{X}_{i1}\boldsymbol{R}_{1}^{-1}
\boldsymbol{R}_{2}$. The TMG estimator of $\boldsymbol{{\Greekmath 010D} }_{0}=E\left(
\boldsymbol{{\Greekmath 010D} }_{i}\right) $ can be computed using the unit-specific
estimates of $\boldsymbol{{\Greekmath 010D} }_{i}$ as above.
The Hausman test proposed in the previous section can also be applied to a
subset of the coefficients in a straightforward manner.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Monte Carlo experiments\label{MC}}
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Data generating processes (DGP) \label{DGP}}
The outcome variable, $y_{it}$, is generated as
\begin{equation}
y_{it}={\Greekmath 010B} _{i}+\sum_{j=1}^{k^{\prime }}{\Greekmath 010C} _{ij}x_{j,it}+{\Greekmath 0114} {\Greekmath 011B}
_{it}e_{it}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, for }i=1,2,...,n\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }t=1,2,...,T, \label{ydgp}
\end{equation}
where the errors, $u_{it}$, are allowed to be serially correlated and
heteroskedastic. Specifically, we set $u_{it}={\Greekmath 0114} {\Greekmath 011B} _{it}e_{it}$ and
generate $e_{it}$ as AR(1) processes
\begin{equation}
e_{it}={\Greekmath 011A} _{ie}e_{i,t-1}+\left( 1-{\Greekmath 011A} _{ie}^{2}\right) ^{1/2}{\Greekmath 0126}
_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for }i=1,2,...,n. \label{yerror}
\end{equation}
In the baseline DGP, we set ${\Greekmath 011B} _{it}={\Greekmath 011B} _{iu}$ for all $t$ and
generate ${\Greekmath 011B} _{iu}^{2}\sim IID\frac{1}{2}\left( 1+z_{iu}^{2}\right)$,
with $z_{iu}\sim IIDN(0,1)$.\footnote{
More general heteroskedastic specifications for ${\Greekmath 011B} _{it}$ are
considered in sub-section \ref{dgprobust} of the online supplement.} Both
Gaussian and non-Gaussian errors are considered: ${\Greekmath 0126} _{it}\sim
IIDN(0,1)$ and ${\Greekmath 0126} _{it}\sim IID\frac{1}{2}\left( {\Greekmath 011F}
_{2}^{2}-2\right) $.
The regressors, $x_{j,it}$, for $j=1,2,....k^{\prime }$, $i=1,2,...,n$, and $
t=1,2,...,T$, are generated as factor-augmented AR(1) processes
\begin{equation}
x_{j,it}={\Greekmath 010B} _{j,ix}(1-{\Greekmath 011A} _{j,ix})+{\Greekmath 010D} _{j,ix}f_{j,t}+{\Greekmath 011A}
_{j,ix}x_{j,i,t-1}+\left( 1-{\Greekmath 011A} _{j,ix}^{2}\right) ^{1/2}u_{xj,it},
\label{xdgp}
\end{equation}
where ${\Greekmath 010B} _{j,ix}\sim IIDN(1,1)$, $u_{xj,it}={\Greekmath 011B} _{j,ix}e_{xj,it}$, $
e_{xj,it}\sim IID(0,1)$, ${\Greekmath 011B} _{j,ix}^{2}=\frac{1}{2}\left(
1+z_{j,ix}^{2}\right) $, and $z_{j,ix}\sim IIDN(0,1)$. We consider panels
with $k^{\prime }=1,2$ and $3$ regressors and investigate the extent to
which trimming is required for different combinations of $T$ and $k^{\prime
} $. The common factors are generated as $
f_{j,t}=0.9f_{j,t-1}+(1-0.9^{2})^{1/2}v_{j,t}$, for $
t=-49,-48,...,-1,0,1,...,T$, where $v_{j,t}\sim IIDN(0,1)$, and $f_{j,-50}=0
$. The factor loadings are generated as ${\Greekmath 010D} _{j,ix}\sim IIDU(0,2)$. When
time effects are included in the model, we set ${\Greekmath 011E} _{t}=t$, for $
t=1,2,...,T-1$, and ${\Greekmath 011E} _{T}=-T(T-1)/2$, so that $\boldsymbol{{\Greekmath 011C} }
_{T}^{\prime }\boldsymbol{{\Greekmath 011E} }=0$.
For the slope coefficients, we experiment with both correlated and
uncorrelated effects in ${\Greekmath 010C} _{i1}$, while considering uncorrelated
heterogeneous effects for the other slope coefficients when the respective
regressors are included. Specifically, ${\Greekmath 010B} _{i}$ and ${\Greekmath 010C} _{i1}$ are
generated as
\begin{equation}
({\Greekmath 010B} _{i},{\Greekmath 010C} _{i1})^{\prime }=\left( {\Greekmath 010B} _{0},{\Greekmath 010C} _{01}\right)
^{\prime }+\left( {\Greekmath 0111} _{i{\Greekmath 010B} },{\Greekmath 0111} _{i{\Greekmath 010C} _{1}}\right) ^{\prime },
\label{coefdgp}
\end{equation}
where
\begin{equation}
\boldsymbol{\tilde{{\Greekmath 0111}}}_{i}=\left( {\Greekmath 0111} _{i{\Greekmath 010B} },{\Greekmath 0111} _{i{\Greekmath 010C}
_{1}}\right) ^{\prime }=\left( \frac{{\Greekmath 011B} _{1,ix}^{2}-E\left( {\Greekmath 011B}
_{1,ix}^{2}\right) }{\sqrt{Var\left( {\Greekmath 011B} _{1,ix}^{2}\right) }}\right)
\boldsymbol{{\Greekmath 0120} }+\boldsymbol{{\Greekmath 010F} }_{i}=\sqrt{2}({\Greekmath 011B} _{1,ix}^{2}-1)
\boldsymbol{{\Greekmath 0120} }+\boldsymbol{{\Greekmath 010F} }_{i}, \label{eta_i}
\end{equation}
$\boldsymbol{{\Greekmath 0120} =(}{\Greekmath 0120} _{{\Greekmath 010B} },{\Greekmath 0120} _{{\Greekmath 010C} _{1}})^{\prime }$, $
\boldsymbol{{\Greekmath 010F} }_{i}=({\Greekmath 010F} _{i{\Greekmath 010B} },{\Greekmath 010F} _{i{\Greekmath 010C}
_{1}})^{\prime }\sim IIDN\left( \boldsymbol{0},\boldsymbol{V}_{{\Greekmath 010F}
}\right) $, and $\boldsymbol{V}_{{\Greekmath 010F} }=Diag(\boldsymbol{{\Greekmath 011B} }
_{{\Greekmath 010F} }^{2})$ with $\boldsymbol{{\Greekmath 011B} }_{{\Greekmath 010F} }^{2}=\left( {\Greekmath 011B}
_{{\Greekmath 010F} {\Greekmath 010B} }^{2},{\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010C} _{1}}^{2}\right) ^{\prime }$
. It follows that
\begin{equation*}
E(\boldsymbol{\tilde{{\Greekmath 0111}}}_{i})=\boldsymbol{0}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }\boldsymbol{V}_{
\boldsymbol{\tilde{{\Greekmath 0111}}}}=E(\boldsymbol{\tilde{{\Greekmath 0111}}}_{i}\boldsymbol{\tilde{
{\Greekmath 0111}}}_{i}^{\prime })=\left(
\begin{array}{cc}
{\Greekmath 011B} _{{\Greekmath 010B} }^{2} & {\Greekmath 011B} _{{\Greekmath 010B} {\Greekmath 010C} _{1}} \\
{\Greekmath 011B} _{{\Greekmath 010B} {\Greekmath 010C} _{1}} & {\Greekmath 011B} _{{\Greekmath 010C} _{1}}^{2}
\end{array}
\right) =\boldsymbol{{\Greekmath 0120} {\Greekmath 0120} }^{\prime }+\boldsymbol{V}_{{\Greekmath 010F} }.
\end{equation*}
The degree of correlated heterogeneity is determined by $\boldsymbol{{\Greekmath 0120}
{\Greekmath 0120} }^{\prime }$, with $Cov({\Greekmath 010B} _{i},{\Greekmath 010C} _{i1})={\Greekmath 011B} _{{\Greekmath 010B} {\Greekmath 010C}
_{1}}\neq 0$ when ${\Greekmath 0120} _{{\Greekmath 010B} }$ and ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ are both
non-zero. Specifically, ${\Greekmath 011B} _{{\Greekmath 010B} }^{2}={\Greekmath 0120} _{{\Greekmath 010B} }^{2}+{\Greekmath 011B}
_{{\Greekmath 010F} {\Greekmath 010B} }^{2}$, ${\Greekmath 011B} _{{\Greekmath 010B} {\Greekmath 010C} _{1}}={\Greekmath 0120} _{{\Greekmath 010B} }{\Greekmath 0120}
_{{\Greekmath 010C} _{1}}$, and ${\Greekmath 011B} _{{\Greekmath 010C} _{1}}^{2}={\Greekmath 0120} _{{\Greekmath 010C} _{1}}^{2}+{\Greekmath 011B}
_{{\Greekmath 010F} {\Greekmath 010C} _{1}}^{2}$. The coefficients of $x_{j,it}$, for $
j=2,3,...,k^{\prime }$ are generated as ${\Greekmath 010C} _{ij}={\Greekmath 010C} _{0j}+{\Greekmath 010F}
_{i{\Greekmath 010C} _{j}}$ with ${\Greekmath 010F} _{i{\Greekmath 010C} _{j}}\sim IIDN(0,{\Greekmath 011B} _{{\Greekmath 010F}
{\Greekmath 010C} _{j}}^{2})$, which are not correlated with any of the regressors. The
true values of the parameters of interest are set as follows: $E({\Greekmath 010B}
_{i})={\Greekmath 010B} _{0}=1$ and $E({\Greekmath 010C} _{ij})={\Greekmath 010C} _{0j}=1$ for $
j=1,2,...,k^{\prime }$. We also set ${\Greekmath 011B} _{{\Greekmath 010B} }^{2}=0.5$, ${\Greekmath 011B}
_{{\Greekmath 010C} _{1}}^{2}=0.75$, and ${\Greekmath 011B} _{{\Greekmath 010F} {\Greekmath 010C} _{j}}^{2}=0.5$ for $
j=2,...,k^{\prime }$.
When $k=k^{\prime }+1=2$, for a fixed $T\geq k$, the asymptotic bias of the
FE estimator of ${\Greekmath 010C} _{01}=E({\Greekmath 010C} _{i1})$ is given by (see (\ref
{AsyBiasFE}))
\begin{equation*}
\func{plim}_{n\rightarrow \infty }\left( \hat{{\Greekmath 010C}}_{1,FE}-{\Greekmath 010C}
_{01}\right) =\frac{\sum_{t=1}^{T}\lim_{n\rightarrow \infty
}n^{-1}\sum_{i=1}^{n}E\left[ (x_{1,it}-\bar{x}_{1,iT})^{2}{\Greekmath 0111} _{i{\Greekmath 010C} _{1}}
\right] }{\sum_{t=1}^{T}\lim_{n\rightarrow \infty }n^{-1}\sum_{i=1}^{n}E
\left[ (x_{1,it}-\bar{x}_{1,iT})^{2}\right] }.
\end{equation*}
The size of the bias will depend on ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ and the parameters
of $x_{1,it}$ process. The exact expression for this bias simplifies
considerably if ${\Greekmath 011A} _{1,ix}=0$ (no dynamics in the $x_{1,it}$ equation).
In this case
\begin{equation*}
\func{plim}_{n\rightarrow \infty }\left( \hat{{\Greekmath 010C}}_{1,FE}-{\Greekmath 010C}
_{01}\right) =\frac{\sqrt{2}\left[ E\left( {\Greekmath 011B} _{1,ix}^{4}\right)
-E\left( {\Greekmath 011B} _{1,ix}^{2}\right) \right] {\Greekmath 0120} _{{\Greekmath 010C} _{1}}\left( \frac{T-1
}{T}\right) }{\left[ T^{-1}\sum_{t=1}^{T}\left( f_{1,t}-\bar{f}_{1,T}\right)
^{2}\right] E\left( {\Greekmath 010D} _{1,ix}^{2}\right) +\left( \frac{T-1}{T}\right)
E\left( {\Greekmath 011B} _{1,ix}^{2}\right) },
\end{equation*}
which, noting that $E\left( {\Greekmath 011B} _{1,ix}^{2}\right) =1$, $E\left( {\Greekmath 011B}
_{1,ix}^{4}\right) =3/2$ and $E\left( {\Greekmath 010D} _{1,ix}^{2}\right) =4/3$,
simplifies to
\begin{equation*}
\func{plim}_{n\rightarrow \infty }\left( \hat{{\Greekmath 010C}}_{1,FE}-{\Greekmath 010C}
_{01}\right) =\frac{\left( \frac{T-1}{T}\right) \sqrt{2}/2{\Greekmath 0120} _{{\Greekmath 010C} _{1}}
}{\left[ T^{-1}\sum_{t=1}^{T}\left( f_{1,t}-\bar{f}_{1,T}\right) ^{2}\right]
\left( 4/3\right) +\left( \frac{T-1}{T}\right) }.
\end{equation*}
It is also worth noting that when $n$ and $T\rightarrow \infty $, jointly,
then $\func{plim}_{n,T\rightarrow \infty }(\hat{{\Greekmath 010C}}_{1,FE}-{\Greekmath 010C}
_{01})=0.303{\Greekmath 0120} _{{\Greekmath 010C} _{1}}$ without interactive effects in $x_{1, it}$
process, and the bias of the FE estimator does not vanish even if $
T\rightarrow \infty $.
For the baseline DGP, we generate the errors in the outcome equation as
chi-squared without serial correlation (${\Greekmath 011A} _{ie}=0$ in (\ref{yerror})). $
x_{j,it}$ are generated without dynamics or interactive effects (${\Greekmath 011A}
_{j,ix}=0$ and ${\Greekmath 010D} _{j,ix}=0$ in (\ref{xdgp})). We set ${\Greekmath 0120} _{{\Greekmath 010B}
}=0.5$ for individual fixed effects. For the slope coefficient, we consider
three possibilities: (a) uncorrelated heterogeneity, ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0$,
(b) a medium level of correlated heterogeneity, ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.5$,
generating a bias around $15\%$ when $k^{\prime }=1$, and (c) a high level, $
{\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.8$, leading to a bias of around $24\%$ for the FE
estimator when $k^{\prime }=1$. For each choice of ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}$, the
scalar parameter ${\Greekmath 0114} $ in (\ref{ydgp}) is set such that the pooled $
R^{2} $ ($PR^{2}$) of panel regressions is around $0.2$. This is achieved by
stochastic simulation for each $T$, as described in sub-section \ref{Simuk}
of the online supplement. We also experiment with a medium level of fit by
setting $PR^{2}=0.4$ when $k^{\prime }=1$.\footnote{
The scaling of the errors in the outcome equation, ${\Greekmath 0114} $, is set when
the DGP contains one regressor. As a result, the $PR^{2}$ will be slightly
higher than the target values of $0.2$ and $0.4$, when $k^{\prime }=2$ or $3$
.}
We consider different values of $T$ depending on $k^{\prime }$ up to $T=15$.
We also carried out a number of robustness checks, detailed in Section \ref
{dgprobust} of the online supplement.
\@startsection{subsection}{2}{\z@}
{-1.5ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .2ex}
{\normalfont\normalsize\bfseries}{Monte Carlo findings}
\@startsection{subsubsection}{3}{\z@}
{-1ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .1ex}
{\normalfont\normalsize\bfseries}{Estimates of the tail index of the distribution of $1/d_{i}$}
\label{MCalphap}
We first provide estimates of the shape parameter of Pareto distributions
fitted to $z_{i}=1/\func{det}(\boldsymbol{X}_{i}^{\prime }\boldsymbol{M}_{T}
\boldsymbol{X}_{i})$, $i=1,2,...,n$, for different choices of $T$ and $
k^{\prime }$. We use the estimator of ${\Greekmath 010B} _{p}$ proposed by \cite
{Hill1975}, which is given by
\begin{equation}
\hat{{\Greekmath 010B}}_{p,Hill}=\frac{(m+1)}{\sum_{j=1}^{m}\ln \left( z^{(j)}\right)
-m\ln \left( z^{(m+1)}\right) }, \label{aphill}
\end{equation}
where $z^{(1)}\geq z^{(2)}....\geq z^{(m)}$ are the first $m$ largest values
of $z_{i}$, and $m$ is the cut-off point set such that $m$ and $
n/m\rightarrow \infty $, as $n\rightarrow \infty $.\footnote{
See also \cite{Pickands1975} and Section 6 of \cite{PesaranYang2020} for
more recent literature on estimation of ${\Greekmath 010B} _{p}$.}
Table \ref{tab:ap_mc} summarizes the estimates of ${\Greekmath 010B} _{p}$ for $n=5,000$
and two cut-off values: $m=n^{1/2}$ and $n^{1/3}$, in the case of panels
with $k^{\prime }=1,2$ and $3$ regressors. As to be expected, the estimates
of ${\Greekmath 010B} _{p}$ are larger for the lower cut-off value, although the
differences between the two estimates are quite small when $T=k^{\prime }+1$
. It is also interesting that when $T=k^{\prime }+1$, the estimates of $
{\Greekmath 010B} _{p}$ lie in the narrow interval [$0.51,0.57$], irrespective of the
value of $k^{\prime }=1,2$ and $3$; thus indicating lack of moments for the
individual estimates and the need for trimming. This finding is not affected
when we consider more general $\{\boldsymbol{x}_{it}\}$ processes. See Table
\ref{tab:ap_mc_3} in the online supplement for $k^{\prime }=1$. Also, as to
be expected, the estimates of ${\Greekmath 010B} _{p}$ rise with $T-k^{\prime}$ and
exceed the threshold value of $2$ for both choices of cut-off values when $
T>2k^{\prime }+3=2k+1$. $T=2k+1$ represents a borderline case where ${\Greekmath 010B}
_{p}$ is estimated to be very close to $2$, but there is still a wide margin
of uncertainty. For example, for $k^{\prime }=1$ and using the cut-off value
of $n^{1/2}$, estimates of ${\Greekmath 010B} _{p}$ for $T=4$ and $5$ are given by $
1.96 $ $(0.23)$ and $2.37$ $(0.28)$, respectively. The standard errors are
in parentheses.
\begin{table}[h]
\caption{The estimates of $\protect{\Greekmath 010B} _{p}$, the tail index of the
distribution of $1/d_{i}$, where $d_{i}=\func{det}(\boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})$, by Hill's method for
two choices of cut-off values}\vspace{-6mm}
\par
\begin{center}
\scalebox{0.75}{
\begin{tabular}{cccccccccccccccccccc}
\hline \hline
\multicolumn{6}{c}{One regressor $(k^{\prime}=1)$} & & \multicolumn{6}{c}{Two regressors $(k^{\prime}=2)$} & & \multicolumn{6}{c}{Three regressors $(k^{\prime}=3)$} \\ \cline{1-6} \cline{8-13} \cline{15-20}
Cut-off & \multicolumn{2}{c}{$n^{1/2}$} && \multicolumn{2}{c}{$n^{1/3}$} & & & \multicolumn{2}{c}{$n^{1/2}$} && \multicolumn{2}{c}{$n^{1/3}$} & & & \multicolumn{2}{c}{$n^{1/2}$} && \multicolumn{2}{c}{$n^{1/3}$} \\ \cline{1-3} \cline{5-6} \cline{9-10} \cline{12-13} \cline{16-17} \cline{19-20}
$T$ & $\hat{{\Greekmath 010B}}_{p}$ & & & $\hat{{\Greekmath 010B}}_{p}$ & & & $T$ & $\hat{{\Greekmath 010B}}_{p}$ & & & $\hat{{\Greekmath 010B}}_{p}$ & & & $T$ & $\hat{{\Greekmath 010B}}_{p}$ & & & $\hat{{\Greekmath 010B}}_{p}$ & \\ \hline
& \multicolumn{19}{c}{$n=5,000$} \\ \hline
2 & 0.51 & (0.06) && 0.56 & (0.13) & & 3 & 0.51 & (0.06) && 0.56 & (0.13) & & 4 & 0.51 & (0.06) && 0.57 & (0.13) \\
3 & 1.02 & (0.12) && 1.13 & (0.27) & & 4 & 0.98 & (0.12) && 1.10 & (0.26) & & 5 & 0.95 & (0.11) && 1.08 & (0.25) \\
4 & 1.51 & (0.18) && 1.67 & (0.39) & & 5 & 1.39 & (0.16) && 1.59 & (0.37) & & 6 & 1.31 & (0.15) && 1.52 & (0.36) \\
5 & 1.96 & (0.23) && 2.19 & (0.52) & & 6 & 1.73 & (0.20) && 1.98 & (0.47) & & 7 & 1.59 & (0.19) && 1.87 & (0.44) \\
6 & 2.37 & (0.28) && 2.68 & (0.63) & & 7 & 2.03 & (0.24) && 2.37 & (0.56) & & 8 & 1.83 & (0.22) && 2.20 & (0.52) \\
8 & 3.11 & (0.37) && 3.59 & (0.85) & & 8 & 2.29 & (0.27) && 2.71 & (0.64) & & 9 & 2.02 & (0.24) && 2.42 & (0.57) \\
10 & 3.73 & (0.44) && 4.32 & (1.02) & & 10 & 2.76 & (0.33) && 3.29 & (0.77) & & 10 & 2.22 & (0.26) && 2.68 & (0.63) \\
15 & 5.10 & (0.60) && 6.04 & (1.42) & & 15 & 3.67 & (0.43) && 4.52 & (1.07) & & 15 & 2.95 & (0.35) && 3.66 & (0.86)
\\\hline\hline
\end{tabular}
}
\end{center}
\par
\vspace{-1mm}
\begin{spacing}{1}
{\footnotesize
Notes: The estimates of ${\Greekmath 010B}_{p}$ and their standard errors are computed using Hill's estimation procedure, with the cut-off values $n^{1/2}$ and $n^{1/3}$. $\boldsymbol{X}_{i} = (\boldsymbol{x}_{i1}, \boldsymbol{x}_{i2}, ..., \boldsymbol{x}_{iT})^{\prime}$, where the $k^{\prime} \times 1$ regressors, $\boldsymbol{x}_{it}$, are generated as specified in the baseline DGP. $\boldsymbol{M}_{T} = \boldsymbol{I}_{T} - \boldsymbol{{\Greekmath 011C}}_{T}\boldsymbol{{\Greekmath 011C}}_{T}^{\prime}/T$. The numbers in brackets are standard errors. }
\end{spacing}
\label{tab:ap_mc}
\end{table}
Based on these estimates, it seems plausible to conclude that trimming
should be considered if $T<2k+1$.\footnote{
Additional estimation results for ${\Greekmath 010B} _{p}$ are provided in Tables \ref
{tab:ap_mc_2} and \ref{tab:ap_mc_3} in Section \ref{MCap} of the online
supplement.} This conclusion is further supported when we compare the
performance of MG and TMG estimators for different choices of $k$ and $T$
using Monte Carlo experiments, to which we now turn.
\@startsection{subsubsection}{3}{\z@}
{-1ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .1ex}
{\normalfont\normalsize\bfseries}{Comparison of TMG, FE, and MG estimators \label{MCfe}}
We begin by comparing the performance of the TMG estimator with those of FE
and MG estimators under both uncorrelated and correlated heterogeneity. We
consider the sample size combinations $n=1,000,2,000,5,000$, $10,000$, $
T=2,3,...,15$ subject to $T\geq k^{\prime }+1$, for $k^{\prime }=1,2$ and $3$
. Recall that $k^{\prime }$ is the number of regressors and $k=k^{\prime }+1$
. The TMG estimator depends on the indicator, $\boldsymbol{1}\{d_{i}>a_{n}\}$
, where $a_{n}=C_{n}n^{-{\Greekmath 010B} }$. In view of the discussion in Section \ref
{Threshold} on the choice of ${\Greekmath 010B} $, we consider the values of ${\Greekmath 010B}
=1/3,0.35$ and $1/2$, and as discussed earlier we set $C_{n}=\bar{d}
_{n}=n^{-1}\sum_{i=1}^{n}d_{i}>0$, where $d_{i}=\func{det}(\boldsymbol{X}
_{i}^{\prime }\boldsymbol{M}_{T}\boldsymbol{X}_{i})$. In what follows, we
report the results for the TMG estimator with ${\Greekmath 010B} =1/3$, but discuss the
sensitivity of the TMG estimator to the choice of ${\Greekmath 010B} $ in Section \ref
{MChatalpha} of the online supplement. All MC experiments are based on $
R=2,000 $ replications.
\begin{sidewaystable}
\caption{Bias, RMSE and size of FE, MG and TMG estimators of ${\Greekmath 010C}_{01}$ $(E({\Greekmath 010C}_{i1}) = {\Greekmath 010C}_{01}=1)$ in the baseline DGP with one regressor, without time effects}
\label{tab:T_d1_c12_chi2_tex0}
\vspace{-7mm}
\begin{center}
\scalebox{0.65}{
\begin{tabular}{rrcrrrrrrrrrrrrrrrrrrrrrrrrr}
\hline\hline & \multicolumn{13}{c}{ Uncorrelated heterogeneity: ${\Greekmath 0120}_{{\Greekmath 010C}_{1}} = 0 $, $PR^{2}=0.2$ } & & \multicolumn{13}{c}{ Correlated heterogeneity: ${\Greekmath 0120}_{{\Greekmath 010C}_{1}}=0.5$, $PR^{2}=0.2$ } \\ \cline{2-14} \cline{16-28}
& $\hat{{\Greekmath 0119}}$ $(\times 100)$ & & \multicolumn{3}{c}{Bias} & & \multicolumn{3}{c}{RMSE} && \multicolumn{3}{c}{Size $(\times 100)$} && $\hat{{\Greekmath 0119}}$ $(\times 100)$ & & \multicolumn{3}{c}{Bias} & & \multicolumn{3}{c}{RMSE} && \multicolumn{3}{c}{Size $(\times 100)$} \\ \cline{2-2} \cline{4-6} \cline{8-10} \cline{12-14} \cline{16-16} \cline{18-20} \cline{22-24} \cline{26-28}
$T$ & TMG & & FE & MG & TMG && FE & MG & TMG && FE & MG & TMG && TMG & & FE & MG & TMG && FE & MG & TMG && FE & MG & TMG \\ \hline
& \multicolumn{27}{c}{$n=1,000$} \\ \hline
2 & 27.30 & ~ & 0.001 & 12.871 & -0.004 & ~ & 0.129 & 744.592 & 0.238 & ~ & 5.0 & 2.8 & 5.0 & ~ & 27.30 & ~ & 0.354 & 14.537 & 0.012 & ~ & 0.395 & 841.266 & 0.268 & ~ & 49.8 & 2.8 & 5.1 \\
3 & 12.00 & ~ & 0.002 & -0.002 & 0.001 & ~ & 0.096 & 0.341 & 0.147 & ~ & 5.7 & 4.2 & 5.2 & ~ & 12.00 & ~ & 0.350 & -0.007 & 0.006 & ~ & 0.371 & 0.386 & 0.165 & ~ & 77.4 & 4.2 & 5.2 \\
4 & 5.90 & ~ & -0.002 & -0.003 & -0.003 & ~ & 0.079 & 0.134 & 0.109 & ~ & 5.2 & 4.6 & 5.0 & ~ & 5.90 & ~ & 0.347 & -0.007 & -0.002 & ~ & 0.361 & 0.151 & 0.122 & ~ & 90.3 & 4.7 & 5.1 \\
5 & 3.10 & ~ & 0.001 & -0.001 & -0.001 & ~ & 0.069 & 0.100 & 0.092 & ~ & 4.3 & 5.3 & 5.4 & ~ & 3.10 & ~ & 0.352 & -0.005 & -0.002 & ~ & 0.363 & 0.111 & 0.102 & ~ & 95.0 & 5.2 & 5.2 \\
6 & 1.70 & ~ & 0.000 & 0.000 & 0.001 & ~ & 0.064 & 0.083 & 0.080 & ~ & 5.3 & 5.2 & 5.5 & ~ & 1.70 & ~ & 0.351 & -0.004 & -0.002 & ~ & 0.360 & 0.091 & 0.088 & ~ & 97.9 & 4.4 & 4.8 \\
8 & 0.60 & ~ & 0.000 & 0.001 & 0.001 & ~ & 0.056 & 0.064 & 0.063 & ~ & 4.8 & 4.9 & 5.0 & ~ & 0.60 & ~ & 0.351 & -0.003 & -0.003 & ~ & 0.357 & 0.069 & 0.068 & ~ & 99.6 & 4.5 & 4.3 \\
10 & 0.20 & ~ & 0.001 & 0.002 & 0.002 & ~ & 0.050 & 0.055 & 0.055 & ~ & 4.8 & 4.9 & 4.9 & ~ & 0.20 & ~ & 0.351 & -0.002 & -0.002 & ~ & 0.356 & 0.059 & 0.059 & ~ & 100.0 & 4.2 & 4.2 \\
15 & 0.00 & ~ & 0.000 & 0.002 & 0.002 & ~ & 0.045 & 0.046 & 0.046 & ~ & 5.1 & 4.7 & 4.7 & ~ & 0.00 & ~ & 0.351 & -0.003 & -0.003 & ~ & 0.354 & 0.047 & 0.047 & ~ & 100.0 & 3.8 & 3.8 \\
& \multicolumn{27}{c}{$n=2,000$} \\ \hline
2 & 24.50 & ~ & 0.002 & -32.239 & 0.001 & ~ & 0.094 & 1307.118 & 0.179 & ~ & 5.8 & 1.7 & 5.1 & ~ & 24.50 & ~ & 0.331 & -36.440 & 0.004 & ~ & 0.351 & 1476.827 & 0.202 & ~ & 80.3 & 1.8 & 5.2 \\
3 & 9.60 & ~ & -0.001 & -0.009 & -0.003 & ~ & 0.068 & 0.263 & 0.108 & ~ & 5.7 & 4.9 & 5.0 & ~ & 9.60 & ~ & 0.326 & -0.025 & -0.011 & ~ & 0.337 & 0.298 & 0.122 & ~ & 97.1 & 4.9 & 4.9 \\
4 & 4.30 & ~ & -0.001 & -0.001 & -0.001 & ~ & 0.056 & 0.097 & 0.081 & ~ & 4.9 & 4.8 & 5.2 & ~ & 4.30 & ~ & 0.328 & -0.016 & -0.012 & ~ & 0.335 & 0.109 & 0.091 & ~ & 99.7 & 5.1 & 5.4 \\
5 & 2.00 & ~ & 0.001 & 0.003 & 0.002 & ~ & 0.049 & 0.070 & 0.065 & ~ & 4.3 & 5.1 & 4.3 & ~ & 2.00 & ~ & 0.329 & -0.012 & -0.012 & ~ & 0.334 & 0.078 & 0.072 & ~ & 100.0 & 4.9 & 4.4 \\
6 & 1.00 & ~ & 0.000 & 0.000 & 0.000 & ~ & 0.045 & 0.059 & 0.058 & ~ & 4.8 & 5.2 & 5.1 & ~ & 1.00 & ~ & 0.328 & -0.015 & -0.014 & ~ & 0.333 & 0.067 & 0.065 & ~ & 100.0 & 5.7 & 5.6 \\
8 & 0.30 & ~ & 0.000 & 0.001 & 0.001 & ~ & 0.039 & 0.046 & 0.045 & ~ & 4.9 & 4.6 & 4.5 & ~ & 0.30 & ~ & 0.328 & -0.013 & -0.013 & ~ & 0.332 & 0.052 & 0.051 & ~ & 100.0 & 5.2 & 5.1 \\
10 & 0.10 & ~ & -0.001 & -0.001 & -0.001 & ~ & 0.036 & 0.041 & 0.041 & ~ & 4.5 & 5.1 & 5.0 & ~ & 0.10 & ~ & 0.329 & -0.016 & -0.016 & ~ & 0.331 & 0.046 & 0.046 & ~ & 100.0 & 6.2 & 6.2 \\
15 & 0.00 & ~ & -0.001 & -0.001 & -0.001 & ~ & 0.031 & 0.032 & 0.032 & ~ & 5.6 & 5.0 & 5.0 & ~ & 0.00 & ~ & 0.327 & -0.016 & -0.016 & ~ & 0.329 & 0.037 & 0.037 & ~ & 100.0 & 6.6 & 6.5 \\
& \multicolumn{27}{c}{$n=5,000$} \\ \hline
2 & 21.10 & ~ & -0.001 & -2.518 & -0.001 & ~ & 0.057 & 115.070 & 0.123 & ~ & 3.6 & 2.0 & 5.4 & ~ & 21.10 & ~ & 0.317 & -2.856 & 0.002 & ~ & 0.324 & 130.011 & 0.139 & ~ & 99.6 & 2.0 & 5.2 \\
3 & 7.20 & ~ & 0.000 & -0.003 & 0.001 & ~ & 0.043 & 0.192 & 0.073 & ~ & 5.5 & 4.2 & 5.4 & ~ & 7.20 & ~ & 0.319 & -0.015 & -0.005 & ~ & 0.323 & 0.217 & 0.082 & ~ & 100.0 & 4.5 & 5.3 \\
4 & 2.80 & ~ & 0.001 & 0.003 & 0.002 & ~ & 0.037 & 0.061 & 0.051 & ~ & 5.9 & 4.2 & 3.6 & ~ & 2.80 & ~ & 0.320 & -0.008 & -0.006 & ~ & 0.322 & 0.069 & 0.057 & ~ & 100.0 & 4.4 & 4.0 \\
5 & 1.20 & ~ & 0.000 & -0.001 & -0.001 & ~ & 0.032 & 0.046 & 0.044 & ~ & 5.2 & 5.1 & 5.2 & ~ & 1.20 & ~ & 0.318 & -0.013 & -0.012 & ~ & 0.320 & 0.052 & 0.050 & ~ & 100.0 & 5.7 & 5.9 \\
6 & 0.50 & ~ & -0.001 & 0.000 & 0.000 & ~ & 0.029 & 0.038 & 0.038 & ~ & 5.3 & 5.2 & 5.0 & ~ & 0.50 & ~ & 0.318 & -0.012 & -0.012 & ~ & 0.320 & 0.044 & 0.043 & ~ & 100.0 & 5.7 & 5.3 \\
8 & 0.10 & ~ & -0.001 & -0.001 & -0.001 & ~ & 0.025 & 0.030 & 0.030 & ~ & 5.4 & 5.1 & 5.2 & ~ & 0.10 & ~ & 0.318 & -0.013 & -0.013 & ~ & 0.319 & 0.035 & 0.035 & ~ & 100.0 & 6.6 & 6.6 \\
10 & 0.00 & ~ & -0.001 & -0.001 & -0.001 & ~ & 0.023 & 0.026 & 0.026 & ~ & 5.2 & 4.9 & 4.8 & ~ & 0.00 & ~ & 0.318 & -0.013 & -0.013 & ~ & 0.319 & 0.030 & 0.030 & ~ & 100.0 & 6.7 & 6.6 \\
15 & 0.00 & ~ & 0.000 & 0.000 & 0.000 & ~ & 0.020 & 0.021 & 0.021 & ~ & 5.0 & 5.1 & 5.1 & ~ & 0.00 & ~ & 0.319 & -0.011 & -0.011 & ~ & 0.320 & 0.024 & 0.024 & ~ & 100.0 & 6.0 & 6.0 \\
& \multicolumn{27}{c}{$n=10,000$} \\ \hline
2 & 18.90 & ~ & 0.000 & -0.063 & -0.003 & ~ & 0.042 & 73.476 & 0.093 & ~ & 4.9 & 2.4 & 5.0 & ~ & 18.90 & ~ & 0.333 & -0.078 & 0.005 & ~ & 0.336 & 83.016 & 0.105 & ~ & 100.0 & 2.4 & 4.8 \\
3 & 5.80 & ~ & -0.001 & -0.002 & -0.001 & ~ & 0.030 & 0.146 & 0.054 & ~ & 4.9 & 4.8 & 5.1 & ~ & 5.80 & ~ & 0.332 & -0.009 & -0.002 & ~ & 0.334 & 0.166 & 0.060 & ~ & 100.0 & 4.9 & 4.9 \\
4 & 2.00 & ~ & 0.000 & 0.000 & 0.000 & ~ & 0.025 & 0.045 & 0.039 & ~ & 4.9 & 5.2 & 5.6 & ~ & 2.00 & ~ & 0.333 & -0.006 & -0.004 & ~ & 0.334 & 0.051 & 0.044 & ~ & 100.0 & 5.2 & 5.6 \\
5 & 0.80 & ~ & 0.000 & -0.002 & -0.002 & ~ & 0.023 & 0.033 & 0.031 & ~ & 5.9 & 5.6 & 5.3 & ~ & 0.80 & ~ & 0.333 & -0.008 & -0.008 & ~ & 0.334 & 0.037 & 0.036 & ~ & 100.0 & 5.9 & 5.8 \\
6 & 0.30 & ~ & 0.000 & 0.000 & 0.000 & ~ & 0.020 & 0.026 & 0.026 & ~ & 4.9 & 4.5 & 4.7 & ~ & 0.30 & ~ & 0.333 & -0.006 & -0.005 & ~ & 0.333 & 0.029 & 0.029 & ~ & 100.0 & 4.5 & 4.4 \\
8 & 0.00 & ~ & 0.000 & -0.001 & -0.001 & ~ & 0.018 & 0.022 & 0.022 & ~ & 5.4 & 5.9 & 5.9 & ~ & 0.00 & ~ & 0.333 & -0.007 & -0.007 & ~ & 0.334 & 0.024 & 0.024 & ~ & 100.0 & 5.8 & 5.9 \\
10 & 0.00 & ~ & 0.000 & 0.000 & 0.000 & ~ & 0.017 & 0.018 & 0.018 & ~ & 5.7 & 4.7 & 4.7 & ~ & 0.00 & ~ & 0.333 & -0.006 & -0.006 & ~ & 0.333 & 0.021 & 0.020 & ~ & 100.0 & 5.9 & 5.8 \\
15 & 0.00 & ~ & 0.000 & 0.000 & 0.000 & ~ & 0.014 & 0.015 & 0.015 & ~ & 4.8 & 5.4 & 5.4 & ~ & 0.00 & ~ & 0.333 & -0.006 & -0.006 & ~ & 0.333 & 0.016 & 0.016 & ~ & 100.0 & 5.4 & 5.4 \\
\hline
\hline
\end{tabular}
}
\end{center}
\vspace{-3mm}
{\footnotesize
Notes:
(i) The baseline DGP is generated as $y_{it}={\Greekmath 010B}_{i} + {\Greekmath 010C}_{i1} x_{1,it} + u_{it}$, where the errors processes for $y_{it}$ and $x_{1,it}$ equations are chi-squared and Gaussian, respectively, $x_{1,it}$ are generated without autoregressions, ${\Greekmath 011A}_{1,ix}=0$, or interactive effects, ${\Greekmath 010D}_{1,ix}=0$, and ${\Greekmath 0120}_{{\Greekmath 010C}_{1}}$ (the degree of correlated heterogeneity) is defined by (\ref{eta_i}).
For further details see Section \ref{DGP}.
(ii) FE and MG estimators are given by (\ref{fee}) and (\ref{mge}), respectively.
The TMG estimator and its asymptotic variance estimator are given by (\ref{TMGb}) and (\ref{varC}).
(iii) The trimming threshold value for the TMG estimator is given by $a_{n}=\bar{d}_{n} n^{-{\Greekmath 010B}}$, where $\bar{d}_{n} =\frac{1}{n} \sum_{i}^{n} d_{i}$ and $d_{i} =\func{det}(\boldsymbol{X}_{i}^{\prime} \boldsymbol{M}_{T} \boldsymbol{X}_{i})$ with $\boldsymbol{X}_{i}=(\boldsymbol{x}_{i1},\boldsymbol{x}_{i2},...,\boldsymbol{x}_{iT})^{\prime}$, $\boldsymbol{M}_{T} = \boldsymbol{I}_{T} - \boldsymbol{{\Greekmath 011C}}_{T}\boldsymbol{{\Greekmath 011C}}_{T}^{\prime}/T$, $\boldsymbol{I}_{T}$ being a $T \times T$ identity matrix and $\boldsymbol{{\Greekmath 011C}}_{T}$ being a $T \times 1$ vector of of ones.
${\Greekmath 010B}$ is set to $1/3$.
$\hat{{\Greekmath 0119}}$ is the simulated fraction of individual estimates being trimmed, defined by (\ref{pin}).
}
\end{sidewaystable}
Table \ref{tab:T_d1_c12_chi2_tex0} reports bias, root mean squared errors
(RMSE) and size for estimation of $E({\Greekmath 010C} _{i1})={\Greekmath 010C} _{01}$ in the case
of DGPs with one regressor, $k^{\prime }=1$. The left panel of the table
provides results when heterogeneity is uncorrelated (i.e. ${\Greekmath 0120} _{{\Greekmath 010C}
_{1}}=0$), whilst the right panel of the table gives the results for the
case of correlated heterogeneity with ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.5$. The fraction
of the trimmed estimates, ${\Greekmath 0119} _{n}$, defined by (\ref{pin}), tends to be
quite large for the case where $T=k$, but falls quite rapidly as $T$ $-k$ is
increased. For example, for $T=2$ and $n=1,000$, as many as $27.3$ per cent
of the individual estimates are trimmed when computing the TMG estimates,
but this fraction falls to $0.6$ per cent when the number of time periods is
increased to $T=6$. However, recall that the TMG estimator continues to make
use of the trimmed estimates, as can be seen from (\ref{betaTn}), and the
TMG estimator shows little bias compared to the (untrimmed) MG estimator.
The TMG and MG estimators converge as $T$ is increased, and they are almost
identical for the panels when $T\geq 8$. This is in line with the two
estimates of ${\Greekmath 010B} _{p}$ reported in Table \ref{tab:ap_mc_3} for $
k^{\prime }=1$ and $T=8$. Both estimates ($3.11$ and $3.59$) are well in
excess of $2$, such that all individual estimates have second-order moments
and therefore no trimming is required.
Comparing TMG and FE estimators, we first note that in line with the theory,
the FE estimator performs very well under uncorrelated heterogeneity but is
badly biased when heterogeneity is correlated. Further, this bias does not
diminish if $n$ and $T$ are increased. The simulated bias of the FE
estimator in the case where $T=k=2$ and $n=1,000$ amounts to $0.354$, which
is close to the analytical result presented in Section \ref{DGP}. When
heterogeneity is correlated, the FE estimator also exhibits substantial size
distortions, which tend to get accentuated as $n$ is increased for a given $
T $. In contrast, the TMG estimator is robust to the choice of ${\Greekmath 0120} _{{\Greekmath 010C}
_{1}}$ and delivers size very close to the assumed five per cent level.
\footnote{
Increasing $PR^{2}$ from $0.2$ to $0.4$ does not affect the bias and RMSE of
the FE estimator but results in a higher degree of size distortion under
correlated heterogeneity. See the results summarized in the right panel of
Table \ref{tab:T_d1_c12_chi2_tex0} and Table \ref{tab:T_d1_c34_chi2_tex0} in
the online supplement.}
The empirical power functions for TMG and FE estimators in the case of a
single regressor and for the sample sizes $n=10,000$ and $T=2$, $3$, and $4$
, are displayed in Figure \ref{fig:fe_tmg_k2_base_1}. As can be seen, under
uncorrelated heterogeneity (the left panel with ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0$),
both estimators are centered correctly around ${\Greekmath 010C} _{01}=1$, with the FE
estimator having better power properties. But the differences between the
power of FE and TMG estimators shrink rapidly and become negligible as $T$
is increased from $T=2$ to $T=4$.\footnote{
But as shown in Example \ref{ExampleMG-FE}, it does not necessarily follow
that the FE estimator will dominate the TMG estimator in terms of efficiency
under uncorrelated heterogeneity. See the left panel of Table \ref
{tab:T_d1_c2_hk23_chi2_tex0} and Figure \ref{fig:fe_tmg_k2_hetrosk} in the
online supplement.} The right panel provides the power plots under
correlated heterogeneity with ${\Greekmath 0120} _{{\Greekmath 010C} _{1}}=0.5$. In this case, the
empirical power functions of the FE estimator now shift markedly to the
right, away from the true value, an outcome that becomes more sharpened as $
T $ is increased. In contrast, the empirical power functions for the TMG
estimator are always centered correctly and are robust to the choice of $
{\Greekmath 0120} _{{\Greekmath 010C} _{1}}$.
\begin{figure}[h!]
\caption{Empirical power functions for FE, MG, and TMG estimators of $
\protect{\Greekmath 010C} _{01}$ $(E(\protect{\Greekmath 010C}_{i1}) = \protect{\Greekmath 010C}_{01}=1)$ in the
baseline DGP with one regressor, without time effects, for $n=10,000$ and $
T=2,3,4$}
\label{fig:fe_tmg_k2_base_1}\vspace{-7mm}
\par
\begin{center}
\includegraphics[scale=0.25]{fe_tmg_k2_base_1.png}
\end{center}
\par
\vspace{-3mm}
\begin{spacing}{1}
{\footnotesize
Notes: See footnotes to Table \ref{tab:T_d1_c12_chi2_tex0}.}
\end{spacing}
\end{figure}
Turning to DGPs with more than one regressor, to save space, the results are
summarized in Tables \ref{tab:T_d1_c12_chi2_tex0_k3} and \ref
{tab:T_d1_c12_chi2_tex0_k4} in Section \ref{MCk34} of the online supplement.
These tables present estimation results for the baseline DGP with two and
three regressors, respectively. The FE estimator continues to display
substantial bias under correlated heterogeneity. For the TMG estimator, as
the number of regressors, $k^{\prime }$, increases, the trimmed fraction
rises for a given $T$. With $n=1,000$, it grows from $27.3$ per cent to $
41.6 $ percent for $T=k=3$, and $50.1$ per cent for $T=k=4$. Moreover, as $
k^{\prime }$ increases, the trimmed fraction decreases with $T$ at a slower
rate, in line with the estimates of ${\Greekmath 010B} _{p}$ in Table \ref{tab:ap_mc}.
To summarize, under uncorrelated heterogeneity, the FE estimator performs
well despite the heterogeneity, and is more efficient than the TMG estimator
in the case of the baseline DGP used in our MCs. But, in general, the
relative efficiency of TMG and FE estimators depends on the underlying DGP.
The situation is markedly different when heterogeneity is correlated, and
the FE estimator can be badly biased, leading to incorrect inference, whilst
the TMG estimator provides valid inference with size around the nominal five
per cent level and reasonable power, irrespective of whether ${\Greekmath 010C} _{i1}$
is correlated with $x_{1,it}$ or not.
\@startsection{subsubsection}{3}{\z@}
{-1ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .1ex}
{\normalfont\normalsize\bfseries}{Comparison of TMG and GP estimators \label{MCcorr}}
Focusing on the case of correlated heterogeneity, we now compare the
relative performance of TMG and GP estimators. To implement the GP
estimator, defined by (\ref{gpe}), for $T=k$, we follow GP and set $
h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}$, with ${\Greekmath 010B} _{GP}=1/3$ and $C_{GP}=\frac{1}{
2}\min \left( \hat{{\Greekmath 011B}}_{D},\hat{r}_{D}/1.34\right) $, where $\hat{{\Greekmath 011B}}
_{D}$ and $\hat{r}_{D}$ are the respective sample standard deviation and
interquartile range of $|\func{det}\left( \boldsymbol{W}_{i}\right) |$ with $
\boldsymbol{W}_{i}=\left( \boldsymbol{{\Greekmath 011C} }_{T},\boldsymbol{X}_{i}\right) $
. For further details of GP's choice of ${\Greekmath 010B} _{GP}$ when $T=k$, see p.
2138 of \cite{GrahamPowell2012}. The asymptotic variance of their estimator
in the case of models with and without time effects is provided in equation
(30) on p. 2126 of their paper. There is no clear guidance by GP as to the
choice of $h_{n}$ when $T>k$.\footnote{
For $T=3$ with $k=2$, GP do not use the bandwidth parameter, $h_{n}$, but
directly select the \textquotedblleft percent trimmed\textquotedblright , $
{\Greekmath 0119} _{n}$. In their empirical application, they report estimates with 4 per
cent being trimmed for $T=3$ with $k=2$. See the last column of Table 3 on
p. 2136 of GP.} For consistency, when $T>k$, for GP estimates we continue to
use their bandwidth, $h_{n}=C_{GP}n^{-{\Greekmath 010B} _{GP}}$ with ${\Greekmath 010B} _{GP}=1/3$
, but set $C_{GP}=\left( n^{-1}\sum_{i=1}^{n}d_{i,GP}\right) ^{1/2}$ and
trim if $d_{i,GP} = \func{det}\left( \boldsymbol{W}_{i}^{\prime} \boldsymbol{
W}_{i}\right) < h_{n}^{2}$.
The bias, RMSE, and size for the two estimators are summarized in Table \ref
{tab:beta_d1_c2_chi2_tex0_b} for $T=2,3,4,5,6$, and $n=1000,2000,5000,10000$
. The associated empirical power functions are displayed in Figure \ref
{fig:tmg_gp_k2_base}.
\begin{table}[h]
\caption{Bias, RMSE and size of TMG and GP estimators of $\protect{\Greekmath 010C}
_{01} $ $(E(\protect{\Greekmath 010C} _{i1})=\protect{\Greekmath 010C} _{01}=1)$ in the baseline DGP
with one regressor, without time effects, but with correlated heterogeneity,
$\protect{\Greekmath 0120} _{\protect{\Greekmath 010C}_{1}}=0.5$}
\label{tab:beta_d1_c2_chi2_tex0_b}\vspace{-6mm}
\par
\begin{center}
\scalebox{0.78}{
\begin{tabular}{rrrrrrrrrrrr}
\hline\hline
& \multicolumn{2}{c}{$\hat{{\Greekmath 0119}}$ $(\times 100)$} & & \multicolumn{2}{c}{Bias} & & \multicolumn{2}{c}{RMSE} & & \multicolumn{2}{c}{Size $(\times 100)$} \\ \cline{2-3} \cline{5-6} \cline{8-9} \cline{11-12}
$T$ & TMG & GP & & TMG & GP & & TMG & GP & & TMG & GP \\ \hline
& \multicolumn{11}{c}{$n=1,000$} \\ \hline
2 & 27.30 & 4.00 & & 0.012 & -0.004 & & 0.268 & 0.599 & & 5.1 & 5.1 \\
3 & 12.00 & 1.30 & & 0.006 & -0.003 & & 0.165 & 0.210 & & 5.2 & 5.0 \\
4 & 5.90 & 0.20 & & -0.002 & -0.007 & & 0.122 & 0.140 & & 5.1 & 4.6 \\
5 & 3.10 & 0.00 & & -0.002 & -0.005 & & 0.102 & 0.110 & & 5.2 & 5.7 \\
6 & 1.70 & 0.00 & & -0.002 & -0.004 & & 0.088 & 0.090 & & 4.8 & 4.6 \\
& \multicolumn{11}{c}{$n=2,000$} \\ \hline
2 & 24.50 & 3.20 & & 0.004 & -0.008 & & 0.202 & 0.474 & & 5.2 & 4.9 \\
3 & 9.60 & 0.80 & & -0.011 & -0.018 & & 0.122 & 0.164 & & 4.9 & 5.2 \\
4 & 4.30 & 0.10 & & -0.012 & -0.016 & & 0.091 & 0.105 & & 5.4 & 5.3 \\
5 & 2.00 & 0.00 & & -0.012 & -0.012 & & 0.072 & 0.078 & & 4.4 & 4.8 \\
6 & 1.00 & 0.00 & & -0.014 & -0.015 & & 0.065 & 0.067 & & 5.6 & 5.8 \\
& \multicolumn{11}{c}{$n=5,000$} \\ \hline
2 & 21.10 & 2.40 & & 0.002 & -0.012 & & 0.139 & 0.355 & & 5.2 & 4.4 \\
3 & 7.20 & 0.40 & & -0.005 & -0.010 & & 0.082 & 0.110 & & 5.3 & 5.1 \\
4 & 2.80 & 0.00 & & -0.006 & -0.008 & & 0.057 & 0.065 & & 4.0 & 3.8 \\
5 & 1.20 & 0.00 & & -0.012 & -0.013 & & 0.050 & 0.052 & & 5.9 & 5.6 \\
6 & 0.50 & 0.00 & & -0.012 & -0.012 & & 0.043 & 0.044 & & 5.3 & 5.7 \\
& \multicolumn{11}{c}{$n=10,000$} \\ \hline
2 & 18.90 & 1.90 & & 0.005 & -0.013 & & 0.105 & 0.281 & & 4.8 & 4.7 \\
3 & 5.80 & 0.30 & & -0.002 & -0.008 & & 0.060 & 0.082 & & 4.9 & 5.1 \\
4 & 2.00 & 0.00 & & -0.004 & -0.006 & & 0.044 & 0.049 & & 5.6 & 5.1 \\
5 & 0.80 & 0.00 & & -0.008 & -0.008 & & 0.036 & 0.037 & & 5.8 & 5.9 \\
6 & 0.30 & 0.00 & & -0.005 & -0.006 & & 0.029 & 0.029 & & 4.4 & 4.4 \\
\hline\hline
\end{tabular}
}
\end{center}
\par
\vspace{-1mm}
\begin{spacing}{1}
{\footnotesize
Notes:
(i) The GP estimator proposed by \cite{GrahamPowell2012} is given by (\ref{gpe}). For $T=k$, GP compare $d_{i,GP}^{1/2}$ with the bandwidth $h_{n} = C_{GP}n^{-{\Greekmath 010B}_{GP}}$, where $d_{i,GP} = \func{det}(\boldsymbol{W}_{i}^{\prime}\boldsymbol{W}_{i})$ and $\boldsymbol{W}_{i} = (\boldsymbol{{\Greekmath 011C}}_{T}, \boldsymbol{X}_{i})$. ${\Greekmath 010B}_{GP}$ is set to 1/3. $C_{GP}=\frac{1}{2}\min \left( \hat{{\Greekmath 011B}}_{D},\hat{r}_{D}/1.34\right) $, where $\hat{{\Greekmath 011B}}_{D}$ and $\hat{r}_{D}$ are the respective sample standard deviation and interquartile range of $d_{i,GP}^{1/2}$. For $T>k$, $C_{GP}=\left(n^{-1}\sum_{i=1}^{n}d_{i,GP}\right)^{1/2}$. See sub-section \ref{MCcorr} for details.
(ii) For details of the baseline DGP without time effects and the TMG estimator, see footnotes to Table \ref{tab:T_d1_c12_chi2_tex0}.
$\hat{{\Greekmath 0119}}$ is the simulated fraction of individual estimates being trimmed, defined by (\ref{pin}). }
\end{spacing}
\end{table}
The fractions of the trimmed estimates, $\hat{{\Greekmath 0119}}$, differ markedly across
the estimators. For example, when $T=2$ and $n=1,000$, the fraction of
trimmed estimates for the TMG estimator is around $27.3$ per cent as
compared to $4.0$ per cent for the GP estimator, and falls to $18.9$ per
cent as $n$ is increased to $10,000$. Increasing $T$ from $2$ to $3$ with $
n=1,000$ reduces this fraction to $16.5$ per cent as compared to $2$ per
cent for the GP estimator. The heavy trimming causes the TMG estimator to
have a larger bias than the GP estimator, particularly when $T=2$ and $n$ is
large. However, the TMG estimator continues to have better overall small
sample performance due to its higher efficiency. Recall that the TMG
estimator makes use of the trimmed estimates, as set out in the second term
of (\ref{betaTn}), but the trimmed estimates are not used in the GP
estimator. This difference in the way trimmed estimates are treated is
reflected in the lower RMSE of the TMG estimator as compared to MG and GP
estimators for all $T$ and $n$ combinations. For example, when $T=2$ and $
n=1,000$, the RMSE of the TMG is $0.27$ as compared to $0.60$ for the GP
estimator. The relative advantage of the TMG estimator continues when $T$
increases from $2$ to $3$. For $T=3$, the RMSE of the TMG estimator stands
at $0.17$ compared to $0.21$ for the GP estimator. The larger the value of $
T $, the less important the trimming becomes.
The empirical power functions for TMG, GP and MG estimators are shown in
Figure \ref{fig:tmg_gp_k2_base}. As can be seen, the TMG estimator is
uniformly more powerful than the GP estimator. To save space, the MC results
for models with time effects are summarized in sub-section \ref{MCte} of the
online supplement. Similar outcomes are obtained when we consider DGPs with
two or three regressors. See Tables \ref{tab:beta_d1_c2_chi2_tex0_b_k3} and
\ref{tab:beta_d1_c2_chi2_tex0_b_k4}, and the corresponding empirical power
functions, Figures \ref{fig:tmg_gp_k3_base} and \ref{fig:tmg_gp_k4_base}, in
sub-section \ref{MCk34} of the online supplement. This is particularly the case
when $T=k\in \{3,4\}$, where the trimmed fraction of the GP estimator rises
only slightly with the number of regressors, resulting in substantial
declines in the empirical powers.
\begin{figure}[h!]
\caption{Empirical power functions for TMG, GP, and MG estimators of $
\protect{\Greekmath 010C}_{01}$ $(E(\protect{\Greekmath 010C}_{i1}) = \protect{\Greekmath 010C}_{01}=1)$ in the
baseline DGP with one regressor, without time effects, but with correlated
heterogeneity, $\protect{\Greekmath 0120}_{\protect{\Greekmath 010C}_{1}}=0.5$, for $n=10,000$ and $
T=2,3,4,5$}
\label{fig:tmg_gp_k2_base}\vspace{-8mm}
\par
\begin{center}
\includegraphics[scale=0.14]{tmg_gp_k2_base.png}
\end{center}
\par
\vspace{-3mm}
\begin{spacing}{1}
{\footnotesize
Notes: See footnotes to Tables \ref{tab:T_d1_c12_chi2_tex0} and \ref{tab:beta_d1_c2_chi2_tex0_b}.
}
\end{spacing}
\end{figure}
For the purpose of comparisons, in addition to the choice of ${\Greekmath 010B}
_{GP}=1/3$ by GP, we also considered the threshold values $2{\Greekmath 010B}
_{GP}={\Greekmath 010B} \in \{0.35,1/2\}$, so that the two threshold functions (ours
and the one suggested by GP) share the same exponents. As reported in Table
\ref{tab:thresh_d1_c2_chi2_tex0_b_t2}, the TMG estimator has a lower RMSE
for all choices of ${\Greekmath 010B} $ and ${\Greekmath 010B} _{GP}$, when $T=2$, and delivers
better empirical powers, as shown in Figures \ref{fig:tmg_gp_alpha_1} and
\ref{fig:tmg_gp_alpha_2} in the online supplement. Additional results for $
T=3$ are provided in sub-section \ref{MCthresh} of the online supplement,
where the differences between TMG and GP estimators are much smaller.
\begin{table}[!htb]
\caption{Bias, RMSE and size of TMG and GP estimators of $\protect{\Greekmath 010C}_{01}$
$(E(\protect{\Greekmath 010C}_{i1})=\protect{\Greekmath 010C}_{01}=1)$ for different threshold
exponents, $\protect{\Greekmath 010B} $ and $\protect{\Greekmath 010B} _{GP}$, in the baseline DGP
with one regressor, without time effects, but with correlated heterogeneity,
$\protect{\Greekmath 0120} _{\protect{\Greekmath 010C}_{1}}=0.5$ ($T=2$)}
\label{tab:thresh_d1_c2_chi2_tex0_b_t2}\vspace{-6mm}
\par
\begin{center}
\scalebox{0.45}{
\LARGE
\begin{tabular}{lcrrcclrrcclrrcclrrcc}
\hline\hline
& & \multicolumn{4}{c}{$n=1,000$} & & \multicolumn{4}{c}{$n=2,000$} & & \multicolumn{4}{c}{$n=5,000$} & & \multicolumn{4}{c}{$n=10,000$} \\ \cline{3-6} \cline{8-11} \cline{13-16} \cline{18-21}
& ${\Greekmath 010B}$/${\Greekmath 010B}_{GP}$ & \multicolumn{1}{c}{$\hat{{\Greekmath 0119}}$} & \multicolumn{1}{c}{Bias} & RMSE & Size & & \multicolumn{1}{c}{$\hat{{\Greekmath 0119}}$} & \multicolumn{1}{c}{Bias} & RMSE & Size & & \multicolumn{1}{c}{$\hat{{\Greekmath 0119}}$} &\multicolumn{1}{c}{Bias} & RMSE & Size & & \multicolumn{1}{c}{$\hat{{\Greekmath 0119}}$} & \multicolumn{1}{c}{Bias} & RMSE & Size \\ \hline
TMG & 1/3 & 27.3 & 0.012 & 0.27 & 5.1 & & 24.5 & 0.004 & 0.20 & 5.2 & & 21.1 & 0.002 & 0.14 & 5.2 & & 18.9 & 0.005 & 0.11 & 4.8 \\
TMG & 0.35 & 25.9 & 0.011 & 0.28 & 5.1 & & 23.0 & 0.002 & 0.21 & 5.0 & & 19.7 & 0.001 & 0.14 & 5.2 & & 17.6 & 0.003 & 0.11 & 5.2 \\
TMG & 0.50 & 15.6 & 0.001 & 0.35 & 4.6 & & 13.1 & -0.009 & 0.27 & 5.3 & & 10.5 & -0.007 & 0.20 & 4.6 & & 8.9 & -0.004 & 0.15 & 4.9 \\
GP & 0.35/2 & 12.0 & 0.006 & 0.36 & 5.3 & & 10.7 & -0.006 & 0.27 & 4.9 & & 9.1 & -0.006 & 0.19 & 5.8 & & 8.1 & -0.001 & 0.14 & 4.8 \\
GP & 0.25 & 7.2 & -0.003 & 0.45 & 4.6 & & 6.0 & -0.015 & 0.35 & 4.4 & & 4.8 & -0.009 & 0.25 & 4.5 & & 4.1 & -0.007 & 0.19 & 5.1 \\
GP & 1/3 & 4.0 & -0.004 & 0.60 & 5.1 & & 3.2 & -0.008 & 0.47 & 4.9 & & 2.4 & -0.012 & 0.36 & 4.4 & & 1.9 & -0.013 & 0.28 & 4.7
\\
\hline\hline
\end{tabular}}
\end{center}
\par
\vspace{-2mm}
\begin{spacing}{1}
{\footnotesize
Notes: See footnotes to Table \ref{tab:beta_d1_c2_chi2_tex0_b}. $\hat{{\Greekmath 0119}}$ is the simulated fraction of individual estimates being trimmed, defined by (\ref{pin}); the reported values are $100 \times \hat{{\Greekmath 0119}}$. Size is reported in per cent.}
\end{spacing}
\end{table}
\@startsection{subsubsection}{3}{\z@}
{-1ex\@plus -1ex \@minus -.2ex}
{0.5ex \@plus .1ex}
{\normalfont\normalsize\bfseries}{MC evidence on the Hausman test of correlated heterogeneity}
Table \ref{tab:Test_d1_arx_chi2_tex0} reports empirical size and power of
the Hausman test of correlated heterogeneity given by (\ref{htest}) under
three scenarios: homogeneity (left), uncorrelated heterogeneity (middle),
and correlated heterogeneity (right). We continue to focus on the simple
case where $k^{\prime }=1$. The size of the test is around the nominal level
of 5 per cent. When $\boldsymbol{x}_{it}$ is strictly exogenous and $
\boldsymbol{{\Greekmath 010C} }_{i1}$ is distributed independently of $\boldsymbol{X}
_{i} $, FE, MG and TMG estimators are all consistent under homogeneity and
if heterogeneity is present but uncorrelated. In such a case, the Hausman
test does not have power. However, in the case where slope coefficients are
heterogeneous \textit{and} correlated with the regressors, the TMG estimator
is consistent while the FE estimator is biased for all $T$. In this case, we
would expect the proposed test to have power, and this is indeed evident in
the right panel of Table \ref{tab:Test_d1_arx_chi2_tex0}. Also, the power of
the test rises with increases in $n$ even when $T=2$, illustrating the
(ultra) small $T$ consistency of the proposed test. We also obtain similar
test results when we allow for time effects. See sub-section \ref{MCtest} of
the online supplement.
\begin{table}[h]
\caption{Empirical size and power of the Hausman test of correlated
heterogeneity in the baseline DGP with one regressor, without time effects}
\label{tab:Test_d1_arx_chi2_tex0}\vspace{-6mm}
\par
\begin{center}
\scalebox{0.8}{
\begin{tabular}{lccccccccccrrrr}
\hline\hline
& \multicolumn{9}{c}{Under $H_{0}$: Uncorrelated heterogeneity} & & \multicolumn{4}{c}{Under $H_{1}$: Correlated heterogeneity} \\ \cline{2-10} \cline{12-15}
& \multicolumn{4}{c}{${\Greekmath 011B}_{{\Greekmath 010C}_{1}}^{2}=0$} & & \multicolumn{4}{c}{${\Greekmath 011B}_{{\Greekmath 010C}_{1}}^{2}=0.75$, ${\Greekmath 0120}_{{\Greekmath 010C}_{1}} =0$} & & \multicolumn{4}{c}{${\Greekmath 011B}_{{\Greekmath 010C}_{1}}^{2}=0.75$, ${\Greekmath 0120}_{{\Greekmath 010C}_{1}} =0.5$ } \\ \cline{2-5} \cline{7-10} \cline{12-15}
$T/n$ & 1,000 & 2,000 & 5,000 & 10,000 & & 1,000 & 2,000 & 5,000 & 10,000 & & $\,\,\,\,\,$1,000 & $\,\,\,\,\,$2,000 & $\,\,\,\,\,$5,000 & $\,\,\,\,$10,000 \\ \hline
2 & 4.9 & 5.0 & 4.4 & 5.1 & & 5.2 & 4.6 & 4.9 & 5.0 & & 25.8 & 39.2 & 67.9 & 92.2 \\
3 & 5.4 & 4.9 & 5.8 & 5.4 & & 5.2 & 5.1 & 4.7 & 5.1 & & 58.9 & 86.5 & 99.5 & 100.0 \\
4 & 4.8 & 5.0 & 5.0 & 5.4 & & 4.7 & 5.2 & 4.4 & 4.7 & & 86.4 & 99.0 & 100.0 & 100.0 \\
5 & 5.3 & 5.0 & 4.8 & 4.6 & & 5.5 & 5.2 & 5.3 & 4.7 & & 95.9 & 99.9 & 100.0 & 100.0 \\
6 & 5.0 & 4.8 & 4.9 & 5.4 & & 4.9 & 5.2 & 5.7 & 4.6 & & 99.1 & 100.0 & 100.0 & 100.0 \\
8 & 4.9 & 5.4 & 4.7 & 4.7 & & 4.7 & 4.9 & 5.4 & 4.7 & & 99.9 & 100.0 & 100.0 & 100.0 \\
\hline\hline
\end{tabular}}
\end{center}
\par
\vspace{-3mm}
\begin{spacing}{1}
\footnotesize{
Notes:
(i) In the baseline DGP for the test, the outcome variable is generated as $y_{it}={\Greekmath 010B}_{i} + {\Greekmath 010C}_{i1} x_{1,it} + u_{it}$, with ${\Greekmath 010B}_{i}$ correlated with $x_{1,it}$ under both the null and alternative hypotheses. For details of the baseline DGP without time effects, see sub-section \ref{DGP}. (ii) The null hypothesis is given by (\ref{null}), including the case of homogenity with ${\Greekmath 011B}_{{\Greekmath 010C}_{1}}^{2}=0$ and the case of uncorrelated heterogeneity with ${\Greekmath 0120}_{{\Greekmath 010C}_{1}}=0$ (the degree of correlated heterogeneity defined by (\ref{eta_i})) and ${\Greekmath 011B}_{{\Greekmath 010C}_{1}}^{2}=0.75$. The alternative of correlated heterogeneity is generated with ${\Greekmath 0120}_{{\Greekmath 010C}_{1}}=0.5$ and ${\Greekmath 011B}_{{\Greekmath 010C}_{1}}^{2}=0.75$.
(iii) The test statistic is calculated based on the difference between FE and TMG estimators, given by (\ref{htest}).
Size and power are in per cent.}
\end{spacing}
\end{table}
Hausman test results for DGPs with two or three regressors are summarized in
Tables \ref{tab:Test_d1_arx_chi2_tex0_k3} through \ref
{tab:Test_te_d1_arx_chi2_tex0_k4} of the online supplement. Recall that we
allow for correlated heterogeneity only in the coefficients of the first
regressor. Consequently, we observe a decline in the power of the Hausman
test as we add regressors with uncorrelated heterogeneous coefficients. The
empirical power of the test decreases with $k$, particularly when $T=k$.
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Empirical application \label{APP}}
In this section, we re-visit the empirical application in \cite
{GrahamPowell2012} who provide estimates of the average effect of household
expenditures on calorie demand, based on a sample of households from poor
rural communities in Nicaragua that participated in a conditional cash
transfer program. The data set is a balanced panel with $n=1,358$ households
observed from 2000 to 2002. We present estimates of the average effects
using the following panel data model with time effects:
\begin{equation}
\ln (Cal_{it})={\Greekmath 010B} _{i}+{\Greekmath 011E} _{t}+{\Greekmath 010C} _{i}\ln (Exp_{it})+u_{it},
\label{calm1}
\end{equation}
where $\ln (Cal_{it})$ denotes the logarithm of household calorie
availability per capita in year $t$ of household $i$, and $\ln (Exp_{it})$
denotes the logarithm of real household expenditures per capita (in
thousands of 2001 cordobas) of household $i$ in year $t$. The parameter of
interest is the average effect defined by ${\Greekmath 010C} _{0}=E({\Greekmath 010C}_{i})$.
We first estimate ${\Greekmath 010B} _{p}$ as a diagnostic to see if trimming is
needed. For the panel covering the period 2001--2002 $(T=2)$, the Hill's
estimates of ${\Greekmath 010B} _{p}$ are 0.71 (0.12) and 0.59 (0.17) for the two
cut-off values of $n^{1/2}$ and $n^{1/3}$, respectively, with standard
errors in parentheses. Similarly, for the dataset 2000--2002 $(T=3)$, the
estimates of ${\Greekmath 010B} _{p}$ are 0.94 (0.15) and 0.96 (0.28), respectively.
All these estimates are well below two, and it is advisable that the trimmed
MG estimator is used for consistent estimation of the average effects, $
{\Greekmath 010C} _{0}$, in case heterogeneity in ${\Greekmath 010C} _{i}$ is correlated.
Table \ref{tab:RPS_cal_htest} reports results of the Hausman test of
correlated heterogeneity in the effects of household expenditures on calorie
demand. The null hypothesis of uncorrelated heterogeneity is rejected for
both the panels and irrespective of whether time effects are included.
Therefore, for this application, the FE and TWFE estimates of ${\Greekmath 010C} _{0}$
could be biased.
\begin{table}[h]
\caption{Hausman statistics for testing correlated heterogeneity in the
effects of household expenditures on calorie demand in Nicaragua}
\label{tab:RPS_cal_htest}
\begin{center}
\vspace{-6mm}
\scalebox{0.85}{
\begin{tabular}{lccccccc}
\hline\hline
& \multicolumn{3}{c}{Without time effects} & & \multicolumn{3}{c}{With time effects} \\ \cline{2-4} \cline{6-8}
& 2001--2002 & & 2000--2002 & & 2001--2002 & & 2000--2002 \\ \hline
Statistics & 5.918 & & 7.626 & & 5.959 & & 7.653 \\
$p$-value & 0.015 & & 0.006 & & 0.015 & & 0.006\\
$T$ & 2 & & 3 & & 2 & & 3\\
\hline\hline
\end{tabular}}
\end{center}
\par
\vspace{-2mm}{\footnotesize Notes: The test is applied to the average effect
${\Greekmath 010C}_{0}=E({\Greekmath 010C}_{i})$ in the model (\ref{calm1}) based on a panel of $
1,358$ households. The test statistic for panels without time effects is
described in the footnote (iii) to Table \ref{tab:Test_d1_arx_chi2_tex0}.
For panels with time effects, the test statistic is based on the difference
between the TWFE and TMG-TE estimators given by (\ref{htetest}) with $T=2$
and (\ref{htetest2}) with $T>2$. For further details see sub-section \ref{TestTE}
in the online supplement. }
\end{table}
Table \ref{tab:RPS_cal_T2} presents the estimates of ${\Greekmath 010C} _{0}$ based on
the panel of 2001--2002 (with $T=2$) without time effects (left panel), and
with time effects (right panel). The estimates are not affected by the
inclusion of time effects but differ considerably across different methods.
\footnote{
When $T=2$, $\hat{{\Greekmath 011E}}_{2002}$ is not significant, and adding time effects
does not change the estimated average effect.} Turning to the trimmed
estimators, we find that only the TMG estimator is heavily trimmed with 27.1
per cent of the estimates being trimmed, whilst the rate of trimming is only
around 3.8 per cent for the GP estimator.\footnote{
For the 2001--2002 panel, $\hat{{\Greekmath 0119}}$ of the GP estimator is identical to
the one reported in Table 3 of \cite{GrahamPowell2012}. GP estimated a model
with time-varying coefficients, $y_{it}={\Greekmath 010B} _{i}+{\Greekmath 011E} _{t}+\left( {\Greekmath 010C}
_{i}+{\Greekmath 011E} _{t,{\Greekmath 010C} }\right) x_{it}+u_{it}$, where $\left( \boldsymbol{{\Greekmath 011E} }
^{\prime},\boldsymbol{{\Greekmath 011E} }_{{\Greekmath 010C} }^{\prime}\right)^{\prime}$ are
identified by stayers but estimated by near stayers with $\boldsymbol{{\Greekmath 011E} }
_{{\Greekmath 010C} } = ({\Greekmath 011E} _{1,{\Greekmath 010C} }, {\Greekmath 011E} _{2,{\Greekmath 010C} }, ..., {\Greekmath 011E} _{T,{\Greekmath 010C}
})^{\prime}$. While $\boldsymbol{{\Greekmath 011E} }_{{\Greekmath 010C} }$ is not included in (\ref
{calm1}), the GP estimates we compute are close to the trimmed estimates in
Table 3 of \cite{GrahamPowell2012}.} Focusing on the estimates without time
effects, we find the FE estimate, 0.6568 (0.0287), is much larger and more
precisely estimated than either the GP or TMG estimates, given by 0.4549
(0.1003) and 0.5623 (0.0425), respectively, with standard errors in
brackets. Judging by the standard errors, it is also noticeable that the TMG
is more precisely estimated than the GP estimate and lies somewhere between
the FE and GP estimates. These estimates are in line with the MC results
reported in the previous section, where we found that in the presence of
correlated heterogeneity, FE estimates are biased with smaller standard
errors (thus leading to incorrect inference), whilst GP and TMG estimators
are correctly centered, with the TMG estimator being more efficient. Similar
results are obtained when we use the extended panel with $T=3$ (2000--2002),
presented in Section \ref{appt3} of the online supplement.
\begin{table}[h]
\caption{Alternative estimates of the average effect of household
expenditures on calorie demand in Nicaragua over the period 2001--2002 ($T=2$
) }
\label{tab:RPS_cal_T2}\vspace{-6mm}
\par
\begin{center}
\scalebox{0.85}{
\begin{tabular}{lccccccc}
\hline\hline & \multicolumn{3}{c}{Without time effects} & & \multicolumn{3}{c}{With time effects} \\ \cline{2-4} \cline{6-8}
& (1) & (2) & (3) & & (5) & (6) & (7) \\
& FE & GP & TMG & & TWFE & GP-TE & TMG-TE \\ \hline
$\hat{{\Greekmath 010C}}_{0}$ & 0.6568 & 0.4549 & 0.5623 & & 0.6554 & 0.4629 & 0.5612 \\
& (0.0287) & (0.1003) & (0.0425) & & (0.0284) & (0.1025) & (0.0424) \\
$\hat{{\Greekmath 011E}}_{2002}$ & ... & ... & ... & & 0.0172 & -0.0181 & 0.0178 \\
& ... & ... & ... & & (0.0063) & (0.0296) & (0.0064) \\
$\hat{{\Greekmath 0119}}$ $(\times 100)$ & ... & 3.8 & 27.1 & & ... & 3.8 & 27.1\\
\hline\hline
\end{tabular}}
\end{center}
\par
\vspace{-2mm} {\footnotesize Notes: (i) The estimates of ${\Greekmath 010C} _{0}=E({\Greekmath 010C}
_{i}) $ and ${\Greekmath 011E} _{2002}$ in the model (\ref{calm1}) are based on a panel
of $1,358$ households. (ii) FE and GP estimators are given by (\ref{fee})
and (\ref{gpe}), respectively. The TMG estimator is given by (\ref{TMGb}),
and its asymptotic variance is estimated by (\ref{varC}). (iii) The TWFE
estimator is given by (\ref{TWFEhat}) in the online supplement. The GP-TE
estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ and $\boldsymbol{{\Greekmath 011E}}_{0}$ is given
by equations (25) and (24) in \cite{GrahamPowell2012}. When $T=k=2$, the
TMG-TE estimator of $\boldsymbol{{\Greekmath 010C} }_{0}$ and $\boldsymbol{{\Greekmath 011E} }_{0}$
are given by (\ref{TMG-TE1}) and (\ref{phihat_b}), and their asymptotic
variances are estimated by (\ref{varbetaTE}) and (\ref{VarPhicombined}),
respectively, in the mathematical appendix. (iv) The trimming threshold
value for TMG and TMG-TE estimators is given by $a_{n}=\bar{d}_{n}
n^{-{\Greekmath 010B}}$, where $\bar{d}_{n} =\frac{1}{n} \sum_{i}^{n} d_{i}$, $d_{i} =
\func{det}(\boldsymbol{X}_{i}^{\prime} \boldsymbol{M}_{T} \boldsymbol{X}
_{i}) $, and $\boldsymbol{X}_{i}=(\boldsymbol{x}_{i1},\boldsymbol{x}
_{i2},...,\boldsymbol{x}_{iT})^{\prime}$ and $\boldsymbol{M}_{T} =
\boldsymbol{I}_{T} - \boldsymbol{{\Greekmath 011C}}_{T}\boldsymbol{{\Greekmath 011C}}_{T}^{\prime}/T$.
${\Greekmath 010B}$ is set to $1/3$. $\hat{{\Greekmath 0119}}$ is the estimated fraction of
individual estimates being trimmed given by (\ref{pin}). \textquotedblleft
...\textquotedblright denotes that the estimation algorithms are not applicable. The numbers
in brackets are standard errors. }
\end{table}
\@startsection {section}{1}{\z@}
{-1.5ex \@plus -1ex \@minus -.2ex}
{0.8ex \@plus.2ex}
{\normalfont\large\bfseries}{Conclusions \label{conclusion}}
This paper studies the estimation of average effects in panel data models
with possibly correlated heterogeneous coefficients, when the number of
cross-sectional units is large, but the number of time periods can be as
small as the number of regression coefficients. We recall that the FE
estimator is inconsistent under correlated heterogeneity, and the MG
estimator could not have second-order (or even first-order) moments when
applied to ultra short panels. The TMG estimator is therefore proposed to
deal with the fat-tailed distributions of the individual estimates (which
does arise if $T$ is very close to $k$) by shrinking (not trimming)
individual estimates that are most likely to fail the second-order moment
condition. The TMG estimator is shown to be consistent and asymptotically
normally distributed, but at a slower rate than $\sqrt{n}$. The paper also
proposes TMG estimators for panels with time effects, distinguishing between
cases where $T=k$ and $T>k$. The TMG estimators play a crucial role in
assessing the robustness of FE and TWFE estimators against slope correlated
heterogeneity. The dispersion slope homogeneity tests by \cite
{PesaranYamagata2008} require large $T$ and do not differentiate between
uncorrelated heterogeneity and correlated heterogeneity.
We highlight the bias and size distortion properties of the FE and TWFE
estimators under correlated heterogeneity. In contrast, the TMG and TMG-TE
estimators are shown to have desirable finite sample performance under a
number of different MC designs, allowing for Gaussian and non-Gaussian
heteroskedastic error processes, dynamic heterogeneity and interactive
effects in the covariates, different numbers of regressors, and different
choices of the trimming threshold parameter, ${\Greekmath 010B} $. In particular, since
the TMG and TMG-TE estimators exploit information on all available
individual estimates, they have the smallest RMSE, and tests based on them
have the correct size and are more powerful than the other trimmed
estimators currently proposed in the literature. The Hausman tests based on
TMG and TMG-TE estimators are also shown to have very good small sample
properties, with their size controlled and their power rising strongly with $
n$ even when $T=k=2$.
\newpage \medskip {\small \setstretch{1.02}
\bibliographystyle{chicago}
\bibliography{TMGref}
}
\newpage