EconBase
← Back to paper

Two-Way Mean Group Estimators for Heterogeneous Panel Models with Fixed T

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

108,950 characters

Two-Way Mean Group Estimators for Heterogeneous Panel Models with Fixed $T$



\title{Two-Way Mean Group Estimators for Heterogeneous Panel Models with
Fixed $T$\thanks{
We thank Laura Serlenga for sharing her data. Lu gratefully acknowledges the
support from Chinese University of Hong Kong for the start-up fund. Su
gratefully acknowledges the NSFC for financial support under the grant
number 72133002. Address correspondence to: Liangjun Su, School of Economics
and Management, Tsinghua University, Beijing, 100084, China; Phone: +86 10
6278 9506; E-mail: [email removed].}}
\author{Xun Lu$^{a},$ Liangjun Su$^{b}$ \\
$^{a}$Department of Economics, Chinese University of Hong Kong\\
$^{b}$School of Economics and Management, Tsinghua University, China}
\date{}
\maketitle

\begin{abstract}
We consider a correlated random coefficient panel data model with two-way
fixed effects and interactive fixed effects in a fixed T framework. We
propose a two-way mean group (TW-MG) estimator for the expected value of the
slope coefficient and propose a leave-one-out jackknife method for valid
inference. We also consider a pooled estimator and provide a Hausman-type
test for poolability. Simulations demonstrate the excellent performance of
our estimators and inference methods in finite samples. We apply our new
methods to two datasets to examine the relationship between health-care
expenditure and income, and estimate a production function.\medskip

\noindent \textbf{Key words:} Heterogeneity; Least squares dummy variable
estimator; Jackknife; Mean-group estimator; Random coefficient\medskip

\noindent \textbf{JEL Classification:}\ C23, C33, C51, C55.
\end{abstract}

\section{Introduction\label{Sec1}}

Heterogeneity is a fundamental characteristic of many economic data (see,
e.g., \cite{heckman2001micro}). The assumption\ of homogeneous slope
coefficients in traditional panel data models has been argued to be
unrealistic. Empirically, the homogeneity of slope coefficients is
frequently rejected; see, e.g., \cite{durlauf2001}, \cite{Browning2007},
\cite{phillips2007}, \cite{Su_Chen2013}, and \cite{lu2023}. In this paper,
we consider a heterogeneous panel data model with two-way fixed effects
(TWFEs) and interactive fixed effects (IFEs) in a fixed $T$ framework. This
type of general model usually requires a large $T,$ which is not applicable
to many panel data sets with only a modest number of time periods. With a
fixed $T,$ although we cannot consistently estimate individual slope
coefficients, we can still obtain a consistent estimator of the average
effects, which is usually of central interest in empirical work. Therefore,
we propose a mean-group (MG) type estimator for the expected value of the
slope coefficient based on the \textit{least squares dummy variable} (LSDV)
approach and a leave-one-out jackknife inference method.\footnote{
Note that there is no simple within transformation we can apply to remove
the TWFEs when the slope coefficients vary with individuals.} One main
advantage of the new estimation and inference methods is their ease of
implementation. Despite their simplicity, they remain effective for a large
class of data-generating processes (DGPs), accommodating an arbitrarily
correlated slope coefficient with the regressors and allowing for the
presence of IFEs. This combination of simplicity and broad applicability can
make our new methods appealing to empirical researchers.

Specifically, we consider the following\ heterogeneous panel data model with
standard TWFEs and IFEs\footnote{
Strictly speaking, the TWFEs can be absorbed into IFEs. Here we explicitly
allow the presence of both, as they play different roles in our estimation
procedure.}:
\begin{equation}
y_{it}={\Greekmath 010C} _{i}^{\prime }x_{it}+{\Greekmath 010B} _{i\cdot }^{o}+{\Greekmath 010B} _{\cdot
t}^{o}+{\Greekmath 0115} _{i}^{o\prime }f_{t}^{o}+u_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }i=1,...,N\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }
t=1,...,T,  \label{y_eqn1}
\end{equation}
where $x_{it}$ is a $K_{x}\times 1$ vector of observable regressors, $
f_{t}^{o}$ and ${\Greekmath 0115} _{i}^{o}$ are $R\times 1$ vectors of unobservable
factors and loadings, respectively, ${\Greekmath 010B} _{i\cdot }^{o}$ and ${\Greekmath 010B}
_{\cdot t}^{o}$ are the individual and time fixed effects, respectively, and
$u_{it}$ is the error term. Here ${\Greekmath 0115} _{i}^{\circ \prime }f_{t}^{\circ }$
in (\ref{y_eqn1}) is often referred to as IFEs in the model, which can be
correlated with $x_{it}.$ For the slope coefficient ${\Greekmath 010C} _{i},$ we assume
that
\begin{equation}
{\Greekmath 010C} _{i}={\Greekmath 010C} ^{0}+{\Greekmath 0111} _{i},  \label{beta_i}
\end{equation}
with $E\left( {\Greekmath 0111} _{i}\right) =0$ so that ${\Greekmath 010C} ^{0}$ represents the
population average marginal effect of $x_{it}$ on $y_{it}.$ The key
parameter of interest is ${\Greekmath 010C} ^{0}.$ Even though we are not interested in
the fixed effect parameters, viz. ${\Greekmath 010B} _{i\cdot }^{o},$ ${\Greekmath 010B} _{\cdot
t}^{o},$ ${\Greekmath 0115} _{i}^{o},$ and $f_{t}^{o},$ we find that it is convenient
to reformulate the model in (\ref{y_eqn1}) to ensure the unobserved factors
and loadings to have zero mean. For this purpose, we define
\begin{equation*}
{\Greekmath 0115} _{i}={\Greekmath 0115} _{i}^{o}-E({\Greekmath 0115} _{i}^{o})\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }
f_{t}=f_{t}^{o}-E(f_{t}^{o}),
\end{equation*}
and rewrite the models in (\ref{y_eqn1}) equivalently as
\begin{equation}
y_{it}={\Greekmath 010C} _{i}^{\prime }x_{it}+{\Greekmath 010B} _{i\cdot }+{\Greekmath 010B} _{\cdot
t}+u_{it}^{\dag },  \label{y_eqn2}
\end{equation}
where ${\Greekmath 010B} _{i\cdot }={\Greekmath 010B} _{i\cdot }^{o}+{\Greekmath 0115} _{i}^{o\prime
}E(f_{t}^{o}),$ ${\Greekmath 010B} _{\cdot t}={\Greekmath 010B} _{\cdot t}^{o}+E\left( {\Greekmath 0115}
_{i}^{o\prime }\right) f_{t},$ and $u_{it}^{\dag }={\Greekmath 0115} _{i}^{\prime
}f_{t}+u_{it}.$ We obtain the LSDV estimators of $\left\{ {\Greekmath 010B} _{i\cdot
}\right\} _{i=1}^{N},\left\{ {\Greekmath 010B} _{\cdot t}\right\} _{t=1}^{T}$ and $
\left\{ {\Greekmath 010C} _{i}\right\} _{i=1}^{N}$ based on (\ref{y_eqn2})$.$ Let $\hat{
{\Greekmath 010C}}_{i}$ denote the estimator of ${\Greekmath 010C} _{i},$ and our estimator of $
{\Greekmath 010C} ^{0}$ is simply
\begin{equation*}
\hat{{\Greekmath 010C}}^{MG}=\frac{1}{N}\sum_{i=1}^{N}\hat{{\Greekmath 010C}}_{i}.
\end{equation*}
For simplicity, we name our estimator $\hat{{\Greekmath 010C}}^{MG}$ two-way-mean-group
(TW-MG) estimator.

There are three main features in our model. First, we consider a fixed $T$
framework and typically only require that $T>K_{x}+1$. Therefore, many
estimators, such as \cite{bai2009}'s principal component analysis (PCA)
estimator and \cite{pesaran2006}'s common correlated effects (CCE)
estimator, which typically require a large $T,$ are not applicable here. \

Second, we consider a heterogeneous panel data model where ${\Greekmath 010C} _{i}$
varies across individuals and can be correlated with $x_{it},$ so we have a
\textit{correlated random coefficient} (CRC) panel in (\ref{y_eqn1}). For
our model, two popular estimators, namely, the pooled TWFE\ estimator and
\cite{pesaran_smith1995}'s standard MG estimator, are not applicable for the
reasons discussed below. The pooled TWFE estimator of ${\Greekmath 010C} _{0}$ is based
on the following regression
\begin{equation}
y_{it}={\Greekmath 010C} ^{0\prime }x_{it}+{\Greekmath 010B} _{i\cdot }+{\Greekmath 010B} _{\cdot t}+\mathring{
u}_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }i=1,...,N\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }t=1,...,T,  \label{TWFE}
\end{equation}
where $\mathring{u}_{it}={\Greekmath 0111} _{i}x_{it}+u_{it}^{\dag }$ is treated as a
composite error term. The reason why the pooled estimator does not work here
is that $x_{it}$ can be correlated with ${\Greekmath 0111} _{i}x_{it}$ (and thus $
\mathring{u}_{it}$), causing the typical endogeneity issue. For concrete
examples, see DGPs 3 and 4 in our simulations in Section \ref{Sim} where the
pooled TWFE\ estimator is inconsistent. See Remark \ref{Rmk3.1} for further
discussion on the correlated random coefficient. \cite{pesaran_smith1995}'s
MG estimator is based on the individual regression for each $i$ as follows,
\begin{equation}
y_{it}={\Greekmath 010C} _{i}^{\prime }x_{it}+{\Greekmath 010B} _{i\cdot }+\ddot{u}_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }
t=1,...,T,  \label{MG}
\end{equation}
where $\ddot{u}_{it}={\Greekmath 010B} _{\cdot t}+u_{it}^{\dag }$ is treated as a
composite error term. The issue here is that $x_{it}$ can be correlated with
${\Greekmath 010B} _{\cdot t}$ (and thus $\ddot{u}_{it}$)$,$ leading to an endogeneity
problem. For example, for all four DGPs in our simulations in Section \ref
{Sim}, the standard MG estimator is inconsistent. For further discussion on
the relationship between our TW-MG and the standard MG estimators, see
Remark \ref{Rmk2.3}. Our TW-MG estimator can be thought of as a generalized
version of the pooled TWFE\ and standard MG estimators, as estimation
equations (\ref{TWFE}) and (\ref{MG}) are restricted versions of (\ref
{y_eqn2}) with the restrictions of ${\Greekmath 010C} _{i}={\Greekmath 010C} ^{0}$ and ${\Greekmath 010B}
_{\cdot t}=0,$ respectively.

Third, we allow IFEs in our model. We argue that once we include TWFEs, the
demeaned IFEs can enter the error term under our key assumption that $
{\Greekmath 0115} _{i}$ and $x_{it}^{o}$ are uncorrelated, where $x_{it}^{o}$ denotes
the part of\ $x_{it}$ with its additive individual and time effects (if any)
removed. For a more precise statement of the key assumption, see Assumption
\ref{Assmp3.1}(i) below. This point has already been made in the literature
(see, e.g., \cite{westerlund2019estimation} and \cite{kapetanios2023testing}
for the large $T$ case). Here we provide a rigorous asymptotic analysis for
a heterogenous model with a fixed $T.$ To illustrate this point, consider
the following simple example where
\begin{equation*}
y_{it}={\Greekmath 010C} _{i}x_{it}+{\Greekmath 0115} _{i}^{o}f_{t}^{o}+u_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }
x_{it}={\Greekmath 010D} _{i}^{o}f_{t}^{o}+v_{it},
\end{equation*}
with $K_{x}=1$ and $R=1.$ Suppose that $E\left( {\Greekmath 0115} _{i}^{o}\right) =$ $
E\left( {\Greekmath 010D} _{i}^{o}\right) =1$ and ${\Greekmath 0115} _{i}^{o},$ ${\Greekmath 010D} _{i}^{o},$
$f_{t}^{o},$ $u_{it}$ and $v_{it}$ are mutually independent with $E\left(
u_{it}\right) =E\left( v_{it}\right) =0$. It appears that the IFEs, ${\Greekmath 0115}
_{i}^{o}f_{t}^{o},$ cause much trouble, as
\begin{equation*}
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Cov}\left( x_{it},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }{\Greekmath 0115} _{i}^{o}f_{t}^{o}\right)
=E[(f_{t}^{o})^{2}]\not=0,
\end{equation*}
which can lead to an endogeneity issue. Furthermore, ${\Greekmath 0115}
_{i}^{o}f_{t}^{o}$ can cause strong cross-sectional dependence in the $y$
-equation. However, we show that our TW-MG estimator still works in this
case after controlling TWFEs, as the demeaned version of ${\Greekmath 0115}
_{i}^{o}f_{t}^{o}$ does not cause the endogeneity issue and strong
cross-sectional dependence. Admittedly, we do rule out the case where $
{\Greekmath 0115} _{i}^{o}$ and $x_{it}^{o}$ are correlated. However, as we already
allow individual (${\Greekmath 010B} _{i\cdot })$ and time effects $({\Greekmath 010B} _{\cdot t})$
without any restrictions, the uncorrelated assumption between ${\Greekmath 0115}
_{i}^{o}$ and $x_{it}^{o}$ may not be that strong. For the simple example
above, essentially, we require that the factor loadings ${\Greekmath 0115} _{i}^{o}$
in the $y$-equation and ${\Greekmath 010D} _{i}^{o}$ in the $x$-equation be
uncorrelated.\footnote{
In fact, the classic paper of \cite{pesaran2006} for large $T$ maintains the
assumption that the factor loadings in the $y$-equation and the $x$-equation
are uncorrelated.} Recently, \cite{kapetanios2023testing} assume that both $
x_{it}$ and $y_{it}$ follow a factor structure as in our simple example (or
more generally, the DGPs in Remark \ref{Rmk2.1}) and test the null
hypothesis that the loadings of $x$- and $y$-equations are uncorrelated.
They find that the null hypothesis is not rejected in thirteen out of
fourteen datasets considered.\footnote{
Their test is based on the large $T$ framework and compares the TWFE
estimator and the PCA\ estimator.} Furthermore, to the best of our
knowledge, there is no well-accepted solution for estimating ${\Greekmath 010C} _{0}$ in
model (\ref{y_eqn1}) with a fixed $T$ when ${\Greekmath 0115} _{i}^{o}$ and $
x_{it}^{o} $ are correlated.

This paper makes several important methodological contributions. First, we
study the asymptotic properties of our TW-MG estimator for the general model
(\ref{y_eqn1}) and demonstrate its consistency and asymptotic mixture normal
distribution. In particular, we invoke the notion of stable convergence and
stable central limit theorem (CLT) (see, e.g., \cite{hausler2015stable} and
\cite{van2023weak}) to take into account the randomness of the factors.
Second, we propose an inference method based on the leave-one-out jackknife
approach and provide rigorous justification. Although this method is easy to
implement, its theoretical justification is challenging because the
parameters involved (${\Greekmath 010C} _{i}$'s) are not fixed when one cross-section
unit is deleted. We have to augment certain matrices and vectors in the
leave-one-out estimator to match those in the original estimator and show
the consistency of the variance estimator. See Remark \ref{Rmk3.2} below for
more details. Third, we examine the asymptotic properties of the pooled TWFE
estimator of ${\Greekmath 010C} _{0}$ for our model (\ref{y_eqn1}) under the additional
assumption that ${\Greekmath 010C} _{i}$ is random and uncorrelated with a demeaned
version of the regressor (Assumption \ref{Assmp4.1}(i) below). To the best
of our knowledge, this is also a new result for the pooled estimator, as we
allow for the presence of IFEs. Furthermore, we propose a Hausman-type test (
\cite{hausman1978specification}) to assess the poolability, that is, whether
the TW-MG and pooled TWFE estimators converge to the same probability limit.
The test is also based on the leave-one-out jackknife method.

Our paper contributes to the extensive literature on panel data models with
IFEs and random coefficients (RCs). For a review of the RC panel data
models, see, e.g., \cite{hsiao2008random} and \cite{Hsiao2022}. For the
review of the panel data models with IFEs, see, e.g., \cite{baiwang2016}. In
the large $T$ framework, \cite{pesaran2006} introduces the CCE estimator for
both homogeneous and heterogeneous static panel data models with IFEs. For
the PCA\ approach, \cite{bai2009}, \cite{moon2015} and \cite{lu_su2016},
among others, propose\ estimators for homogeneous model. \cite
{li2020efficient} and \cite{cao2024oracle} extend this by proposing a
PCA-type estimator for heterogeneous models. The literature for
heterogeneous panel data models with fixed $T$ is relatively sparse.
Although \cite{pesaran_smith1995}'s MG estimator is originally proposed for
the large $T$ case, it has been shown to be effective for fixed $T$ as well
(see, e.g., \cite{hsiao2019panel}). \cite{chamberlain1992efficiency}
proposes an estimator similar to ours but does not account for IFEs. More
recently, \cite{pesaran2024trimmed} introduce a trimmed version of
Chamberlain's estimator, which also excludes IFEs and employs a different
inference approach. \cite{westerlund2022cce} provide a pooled CCE estimator
for heterogeneous models.

The rest of the paper is structured as follows. In Section \ref{Sec2}, we
propose the estimation and inference methods. We aim to make Section \ref
{Sec2} accessible to applied researchers who are interested in applying our
new methods. Section \ref{Sec3} examines the asymptotics of our TW-MG
estimator and the jackknife inference procedure. In Section \ref{Sec4}, we
study the pooled estimator and propose a Hausman-type test for poolability.
Section \ref{Sim} reports Monte Carlo simulation results. In Section \ref
{App}, we apply our new methods to health-care expenditure data and
production data. Section \ref{Sec7} concludes. The appendix contains the
proofs of all main results. The online supplement contains the proofs of
some technical lemmas and additional simulation results.

\textit{Notation}. For a positive integer $n$, let $[n]\equiv \{1,2,\ldots
,n\},$ $n_{1}=n-1,$ $\mathbb{I}_{n}$ an $n\times n$ identity matrix, and $
\mathbf{{\Greekmath 0113} }_{n}$ an $n\times 1$ vector of ones. For a real $m\times n$
matrix $A$, we use $\Vert A\Vert ,$ $\Vert A\Vert _{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sp }}$and $
A^{\prime }$ to denote its Frobenius norm, spectral norm, and transpose,
respectively. When $A$ has full column rank, let $\mathbb{P}_{A}\equiv
A\left( A^{\prime }A\right) ^{-1}A^{\prime }$ and $\mathbb{M}_{A}\equiv
\mathbb{I}_{m}-\mathbb{P}_{A}$, where $\equiv $ signifies a definitional
relationship. Let ${\Greekmath 0116} _{\max }(\cdot )$ and ${\Greekmath 0116} _{\min }(\cdot )$ denote
the largest and smallest eigenvalue of a real symmetric matrix,
respectively. Let $\otimes $ denote the Kronecker product and $\mathbf{1}
\{\cdot \}$ the usual indicator function. Let $\mathbf{0}_{a\times b}$
denote an $a\times b$ matrix of zeros. Let bdiag$\left(
A_{1},...,A_{N}\right) $ denote the block diagonal matrix. The symbol $
\overset{p}{\rightarrow }$ denotes convergence in probability, $\overset{d}{
\rightarrow }$ convergence in distribution, and plim probability limit. We
use $a_{N}\lesssim b_{N}$ to denote that $a_{N}/b_{N}=O_{p}\left( 1\right) $
as $N\rightarrow \infty .$

\section{Model, Estimation and Inference Procedure \label{Sec2}}

In this section, we introduce the model and estimation and inference
procedure.

\subsection{Model \label{Sec2.1}}

Our original model is specified in (\ref{y_eqn1}), and it is reformulated in
(\ref{y_eqn2}). Note that $E(u_{it}^{\dag })=0$ in (\ref{y_eqn2}) under some
weak assumptions (e.g., Assumption \ref{Assmp3.1}(i) in Section \ref{Sec3.1}
or Assumption \ref{Assmp4.1}(i) in Section \ref{Sec4.1}). More importantly, $
u_{it}^{\dag }$ is also uncorrelated with a demeaned version of $x_{it}$ (by
removing its individual and time effects) under Assumption \ref{Assmp3.1},
which ensures that one can estimate ${\Greekmath 010C} _{i}$'s and ${\Greekmath 010C} ^{0}$ by the
usual least squares approach. We emphasize that the above reformulation is
solely for theoretical analysis and is not required for the implementation
of our estimation and inference procedure.

Since we focus on the large $N$ and fixed $T$ framework, it is hard to
consider a PCA estimate for ${\Greekmath 010C} ^{0}.$ Due to the presence of the time
effects ${\Greekmath 010B} _{\cdot t}$ in the $y$-equation, one cannot pool the $T$
observations for each $i$ to run a time series regression of $y_{it}$ on $
\left( x_{it}^{\prime },1\right) $ to estimate ${\Greekmath 010C} _{i}$ as in standard
MG estimation (see, e.g., \cite{pesaran_smith1995}). Below we propose a
TW-MG estimator of ${\Greekmath 010C} ^{0}$ via the LSDV approach. Even though we do not
formally address the missing value problem in this paper, it is well known
that the use of the LSDV approach can easily account for the missing value
issues in panel regressions.

\begin{remark}
\label{Rmk2.1}\textbf{(An example of }$x_{it}$\textbf{) }In the above setup,
we only specify the DGP for $y_{it}.$ If one follows the literature on the
CCE estimation (e.g., \cite{pesaran2006}), one can also specify a DGP for $
x_{it}$ as follows:
\begin{equation}
x_{it}={\Greekmath 0116} _{i\cdot }^{o}+{\Greekmath 0116} _{\cdot t}^{o}+{\Greekmath 010D} _{i}^{o\prime
}f_{t}^{o}+v_{it},  \label{x_eqn2a}
\end{equation}
where ${\Greekmath 010D} _{i}^{o}$ is an $R\times K_{x}$ matrix of factor loadings in
the $x$-equation, ${\Greekmath 0116} _{i\cdot }^{o}$ and ${\Greekmath 0116} _{\cdot t}^{o}$ are the
individual and time fixed effects, respectively, and $v_{it}$ is the
zero-mean error term. As above, one can reformulate ${\Greekmath 010D} _{i}^{o}$ and $
f_{t}^{o}$ to have zero mean such that
\begin{equation}
x_{it}={\Greekmath 0116} _{i\cdot }+{\Greekmath 0116} _{\cdot t}+{\Greekmath 010D} _{i}^{\prime }f_{t}+v_{it},
\label{x_eqn2}
\end{equation}
where ${\Greekmath 010D} _{i}={\Greekmath 010D} _{i}^{o}-E({\Greekmath 010D} _{i}^{o}),$ $
f_{t}=f_{t}^{o}-E(f_{t}^{o}),$ ${\Greekmath 0116} _{i\cdot }={\Greekmath 0116} _{i\cdot }^{o}+{\Greekmath 010D}
_{i}^{o\prime }E(f_{t}^{o})$ and ${\Greekmath 0116} _{\cdot t}={\Greekmath 0116} _{\cdot t}^{o}+E\left(
{\Greekmath 010D} _{i}^{o\prime }\right) f_{t}.$ In this paper we do not need to
specify a DGP for $x_{it}.$ Even so, it is worth noting that if $x_{it}$ is
indeed generated via ($\ref{x_eqn2}$), then only its double-demeaned
version, $x_{it}^{o},$ will enter the asymptotics of our TW-MG estimator
below.
\end{remark}

\subsection{Estimation procedure \label{Sec2.2}}

To present the estimation procedure, we introduce some notations. Since $
{\Greekmath 010B} _{i\cdot }$ and ${\Greekmath 010B} _{\cdot t}$ are not separately identified, we
need to impose a location restriction for the LSDV estimation. Below, we
impose the restriction that
\begin{equation*}
\sum_{t=1}^{T}{\Greekmath 010B} _{\cdot t}=0
\end{equation*}
in the estimation procedure without requiring the true values $\left\{
{\Greekmath 010B} _{i\cdot }\right\} $ to satisfy such a restriction. Let $\mathbf{
{\Greekmath 010B} }^{\left( 1\right) }\equiv \left( {\Greekmath 010B} _{1\cdot },...,{\Greekmath 010B}
_{N\cdot }\right) ^{\prime },$ $\mathbf{{\Greekmath 010B} }^{\left( 2\right) }\equiv
\left( {\Greekmath 010B} _{\cdot 1},...,{\Greekmath 010B} _{\cdot T-1}\right) ^{\prime },$ and $
\mathbf{{\Greekmath 010C} }\equiv ({\Greekmath 010C} _{1}^{\prime },...,{\Greekmath 010C} _{N}^{\prime
})^{\prime }.$ Define
\begin{equation}
\mathbf{D}^{\left( 1\right) }=\mathbb{I}_{N}\otimes \mathbf{{\Greekmath 0113} }_{T}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
and }\mathbf{D}^{\left( 2\right) }=\mathbf{{\Greekmath 0113} }_{N}\otimes \left(
\begin{array}{c}
\mathbb{I}_{T_{1}} \\
-\mathbf{{\Greekmath 0113} }_{T_{1}}^{\prime }
\end{array}
\right) ,  \label{DD1}
\end{equation}
where $T_{1}\equiv T-1.$ Let $\mathbf{y}_{i}=\left( y_{i1},...,y_{iT}\right)
^{\prime },$ \ $\mathbf{Y}=\left( \mathbf{y}_{1}^{\prime },...,\mathbf{y}
_{N}^{\prime }\right) ^{\prime },$ $\mathbf{x}_{i}=\left(
x_{i1},...,x_{iT}\right) ^{\prime }$ and $\mathbf{X}=$bdiag$(\mathbf{x}
_{1},...,\mathbf{x}_{N}).$ Define $\mathbf{u}_{i}=\left(
u_{i1,}...,u_{iT}\right) ^{\prime },$\ $\mathbf{U}=\left( \mathbf{u}
_{1}^{\prime },...,\mathbf{u}_{N}^{\prime }\right) ^{\prime },$ $F=\left(
f_{1},...,f_{T}\right) ^{\prime }$ and $\mathbf{{\Greekmath 0115} =(}{\Greekmath 0115}
_{1}^{\prime },...,{\Greekmath 0115} _{N}^{\prime }\mathbf{)}^{\prime }.$ Define $
\mathbf{u}_{i}^{\dag }$ and $\mathbf{U}^{\dag }$ analogously. Then we can
rewrite the model in (\ref{y_eqn2}) in matrix form:
\begin{equation}
\mathbf{Y}=\mathbf{X{\Greekmath 010C} }+\mathbf{D}^{\left( 1\right) }\mathbf{{\Greekmath 010B} }
^{\left( 1\right) }+\mathbf{D}^{\left( 2\right) }\mathbf{{\Greekmath 010B} }^{\left(
2\right) }+\mathbf{U}^{\dag }=\mathbf{X{\Greekmath 010C} }+\mathbf{D{\Greekmath 010B} }+\mathbf{U}
^{\dag },  \label{Y_eqn1}
\end{equation}
where $\mathbf{D}=\left( \mathbf{D}^{\left( 1\right) },\mathbf{D}^{\left(
2\right) }\right) ,$ $\mathbf{{\Greekmath 010B} }=(\mathbf{{\Greekmath 010B} }^{\left( 1\right)
\prime },\mathbf{{\Greekmath 010B} }^{\left( 2\right) \prime })^{\prime },$ and $
\mathbf{U}^{\dag }=(\mathbb{I}_{N}\otimes F)\mathbf{{\Greekmath 0115} }+\mathbf{U}.$

By partialling out $\mathbf{D}$ in (\ref{Y_eqn1}), we obtain the following
LSDV estimator of $\mathbf{{\Greekmath 010C} }$:
\begin{equation}
\mathbf{\hat{{\Greekmath 010C}}}\equiv (\hat{{\Greekmath 010C}}_{1}^{\prime },...,\hat{{\Greekmath 010C}}
_{N}^{\prime })^{\prime }\equiv \left( \mathbf{X}^{\prime }{\mathbb{M}}_{
\mathbf{D}}\mathbf{X}\right) ^{-1}\mathbf{X}^{\prime }{\mathbb{M}}_{\mathbf{D
}}\mathbf{Y.}  \label{betai_est}
\end{equation}
Then our TW-MG estimator of ${\Greekmath 010C} ^{0}$ is given by
\begin{equation*}
\hat{{\Greekmath 010C}}^{MG}=\frac{1}{N}\sum_{i=1}^{N}\hat{{\Greekmath 010C}}_{i}.
\end{equation*}
Let $\mathbf{S}_{i,N}=(\mathbf{0}_{K_{x}\times K_{x}},...,\mathbf{0}
_{K_{x}\times K_{x}},\mathbb{I}_{K_{x}},\mathbf{0}_{K_{x}\times K_{x}},...,
\mathbf{0}_{K_{x}\times K_{x}})$ be a $K_{x}\times NK_{x}$ selection matrix
that has $\mathbb{I}_{K_{x}}$ in its $i$th block. Then $\hat{{\Greekmath 010C}}^{MG}$
can also be written as
\begin{equation}
\hat{{\Greekmath 010C}}^{MG}=\frac{1}{N}\sum_{i=1}^{N}\mathbf{S}_{i,N}(\mathbf{X}
^{\prime }{\mathbb{M}}_{\mathbf{D}}\mathbf{X)}^{-1}\mathbf{X}^{\prime }{
\mathbb{M}}_{\mathbf{D}}\mathbf{Y}=\frac{1}{N}\mathbf{S}_{N}(\mathbf{X}
^{\prime }{\mathbb{M}}_{\mathbf{D}}\mathbf{X)}^{-1}\mathbf{X}^{\prime }{
\mathbb{M}}_{\mathbf{D}}\mathbf{Y,}  \label{betahat1}
\end{equation}
where $\mathbf{S}_{N}=\mathbf{{\Greekmath 0113} }_{N}^{\prime }\otimes \mathbb{I}
_{K_{x}}=(\mathbb{I}_{K_{x}},...,\mathbb{I}_{K_{x}}).$

\begin{remark}
\label{Rmk2.2}\textbf{(An alternative representation)} Note that $\mathbf{D}
=(\mathbf{D}^{\left( 1\right) },\mathbf{D}^{\left( 2\right) }),$ where $
\mathbf{D}^{\left( 1\right) }$ and $\mathbf{D}^{\left( 2\right) }$ are
defined in (\ref{DD1}). It is easy to verify that
\begin{eqnarray*}
\mathbf{D}^{\prime }\mathbf{D} &\mathbf{=}&\left(
\begin{array}{cc}
T\mathbb{I}_{N} & \mathbf{0}_{N\times T_{1}} \\
\mathbf{0}_{T_{1}\times N} & N\left( \mathbb{I}_{T_{1}}+\mathbf{{\Greekmath 0113} }
_{T_{1}}\mathbf{{\Greekmath 0113} }_{T_{1}}^{\prime }\right)
\end{array}
\right) , \\
\mathbb{P}_{\mathbf{D}} &=&\mathbf{D}\left( \mathbf{D}^{\prime }\mathbf{D}
\right) ^{-1}\mathbf{D}^{\prime }=\frac{\mathbf{{\Greekmath 0113} }_{N}\mathbf{{\Greekmath 0113} }
_{N}^{\prime }}{N}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}+\mathbb{I}
_{N}\otimes \frac{\mathbf{{\Greekmath 0113} }_{T}\mathbf{{\Greekmath 0113} }_{T}^{\prime }}{T},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
and} \\
\mathbb{M}_{\mathbf{D}} &=&\mathbb{I}_{NT}-\mathbb{P}_{\mathbf{D}}=\mathbf{D}
\left( \mathbf{D}^{\prime }\mathbf{D}\right) ^{-1}\mathbf{D}^{\prime }=
\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}.
\end{eqnarray*}
As a result, we can also write the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}$ as
follows:
\begin{equation}
\hat{{\Greekmath 010C}}^{MG}=\frac{1}{N}\mathbf{S}_{N}\left[ \mathbf{X}^{\prime }(
\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})
\mathbf{X}\right] ^{-1}\mathbf{X}^{\prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }
_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{Y.}  \label{betahat2}
\end{equation}
Clearly, $\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }
_{T}}$ is the double-demean (or two-way within group transformation)
operator that removes the individual and time fixed effects in both $y_{it}$
and $x_{it}.$
\end{remark}

\begin{remark}
\label{Rmk2.3}(\textbf{Comparison with standard MG\ estimator}) Since the
seminal work of \cite{pesaran_smith1995}, MG\ estimation has been widely
applied in panel data analyses to estimate the average marginal effect. To
the best of our knowledge, standard MG estimation is applied to heterogenous
panel data models with only individual effects. In this case, one can use
the $T$ observations for cross-section unit $i$ and estimate the
individual-specific slope coefficients ${\Greekmath 010C} _{i}$ via a time series
regression and then average the resulting estimates. Due to the presence of
time effects ${\Greekmath 010B} _{\cdot t}$ in (\ref{y_eqn2}), this approach is not
applicable. In contrast, in our TW-MG estimation procedure above, we first
pool all $NT$ observations across $i$ and $t$ and estimate $\left\{ {\Greekmath 010C}
_{i}\right\} _{i\in \left[ N\right] }$ jointly and then average the
resulting estimates.
\end{remark}

\begin{remark}
\label{Rmk2.4}\textbf{(Ridge-type estimator)} When $T$ is small, to improve
the finite sample performance, we can modify the estimator in (\ref{betahat2}
) as
\begin{equation}
\hat{{\Greekmath 010C}}^{MG,Ridge}=\frac{1}{N}\mathbf{S}_{N}\left[ \frac{1}{T}\mathbf{X}
^{\prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113}
}_{T}})\mathbf{X}+{\Greekmath 0114} _{N}\cdot \mathbb{I}_{NK_{x}}\right] ^{-1}\frac{1}{T
}\mathbf{X}^{\prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{
\mathbf{{\Greekmath 0113} }_{T}})\mathbf{Y,}  \label{ridge}
\end{equation}
where ${\Greekmath 0114} _{N}$ is a tuning parameter. This is a ridge-type estimator.
For the choice of ${\Greekmath 0114} _{N},$ we can set ${\Greekmath 0114} _{N}=c_{{\Greekmath 0114} }\frac{1}{
N},$ where $c_{{\Greekmath 0114} }$ is a constant. We can show that all our theoretical
results below hold for $\hat{{\Greekmath 010C}}^{MG,Ridge}$ as well with this choice of $
{\Greekmath 0114} _{N}.$ In the simulations and applications below, to account for the
scale of the data, we let $c_{{\Greekmath 0114} }$ be the median of $\left\{ \det
\left( \frac{1}{T}\sum_{t=1}^{T}\ddot{x}_{it}\ddot{x}_{it}^{\prime }\right)
\right\} _{i\in \lbrack N]},$ where $\det $ stands for determinant and $
\ddot{x}_{it}$ is the double-demeaned version of $x_{it}$: $\ddot{x}
_{it}=x_{it}-\frac{1}{N}\sum_{j=1}^{N}x_{jt}-\frac{1}{T}\sum_{s=1}^{T}x_{is}+
\frac{1}{NT}\sum_{j=1}^{N}\sum_{s=1}^{T}x_{js}.$
\end{remark}

\subsection{Inference procedure \label{Sec2.3}}

Below we show that $\hat{{\Greekmath 010C}}^{MG}$ is asymptotically mixture normal under
some regularity conditions. We propose a jackknife method to estimate the
asymptotic variance of $\hat{{\Greekmath 010C}}^{MG}$ and to conduct inference for $
{\Greekmath 010C} ^{0}.$ Following (\ref{betahat2}), we can obtain the leave-one-out
estimator of ${\Greekmath 010C} ^{0}$ in (\ref{y_eqn2}) by leaving out the observations
for the $i$th cross-sectional unit:
\begin{equation}
\hat{{\Greekmath 010C}}^{MG\left( -i\right) }=\frac{1}{N_{1}}\mathbf{S}_{N_{1}}\left[
\mathbf{X}^{\left( -i\right) \prime }\left( \mathbb{M}_{\mathbf{{\Greekmath 0113} }
_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}\right) \mathbf{X}^{\left(
-i\right) }\right] ^{-1}\mathbf{X}^{\left( -i\right) \prime }\left( \mathbb{M
}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}\right)
\mathbf{Y}^{\left( -i\right) },  \label{beta_leave_one_out1}
\end{equation}
where $N_{1}\equiv N-1,$ $\mathbf{X}^{\left( -i\right) \prime }$ is the
leave-one-out version\ of $\mathbf{X}^{\prime }$ by deleting its $i$th row
block: $\mathbf{X}^{\left( -i\right) \prime }=$bdiag$(\mathbf{x}_{1}^{\prime
},$ $...,\mathbf{x}_{i-1}^{\prime },\mathbf{x}_{i+1}^{\prime },...,\mathbf{x}
_{N}^{\prime }),$ and $\mathbf{Y}^{\left( -i\right) }=\left( \mathbf{y}
_{1}^{\prime },...,\mathbf{y}_{i-1}^{\prime },\mathbf{y}_{i+1}^{\prime },...,
\mathbf{y}_{N}^{\prime }\right) ^{\prime }.$

Given the jackknife estimates $\{\hat{{\Greekmath 010C}}^{MG\left( -i\right) }\}_{i\in
\left[ N\right] },$ we propose to estimate the asymptotic variance of $\sqrt{
N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})$ by
\begin{equation*}
\hat{\Omega}^{MG}=(N-1)\mathop{\displaystyle \sum }\limits_{i=1}^{N}\left( \hat{{\Greekmath 010C}}^{MG\left(
-i\right) }-\overline{\hat{{\Greekmath 010C}}}^{_{MG}}\right) \left( \hat{{\Greekmath 010C}}
^{MG\left( -i\right) }-\overline{\hat{{\Greekmath 010C}}}^{_{MG}}\right) ^{\prime },
\end{equation*}
where $\overline{\hat{{\Greekmath 010C}}}^{_{MG}}\equiv \frac{1}{N}\sum_{i=1}^{N}\hat{
{\Greekmath 010C}}^{MG\left( -i\right) }.$ Let $\mathbb{S}$ be a $K_{x}\times 1$
selection vector, which defines the scalar parameter of interest: $\mathbb{S}
^{\prime }{\Greekmath 010C} ^{0}.$ Let $z_{{\Greekmath 011C} /2}$ denote the upper ${\Greekmath 011C} /2$-quantile
of $N(0,1).$ Then we construct the $100(1-{\Greekmath 011C} )\%$ confidence interval (CI)
of $\mathbb{S}^{\prime }{\Greekmath 010C} ^{0}$ as
\begin{equation}
\left[ \mathbb{S}^{\prime }\hat{{\Greekmath 010C}}^{MG}-z_{{\Greekmath 011C} /2}\sqrt{\mathbb{S}
^{\prime }\hat{\Omega}^{MG}\mathbb{S}/N},\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ }\mathbb{S}^{\prime }
\hat{{\Greekmath 010C}}^{MG}+z_{{\Greekmath 011C} /2}\sqrt{\mathbb{S}^{\prime }\hat{\Omega}^{MG}
\mathbb{S}/N}\right] .  \label{CI1}
\end{equation}

\section{Asymptotic Properties of $\hat{\protect{\Greekmath 010C}}^{MG}$ \label{Sec3}}

In this section, we study the asymptotic properties of $\hat{{\Greekmath 010C}}^{MG}$.
We first state the assumptions and show that $\sqrt{N}(\hat{{\Greekmath 010C}}
^{MG}-{\Greekmath 010C} ^{0})$ follows a version of stable CLT. Then we show that the
above jackknife method can estimate the asymptotic variance of $\sqrt{N}(
\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})$ consistently.

\subsection{Basic assumptions \label{Sec3.1}}

To proceed, we introduce some notations. Let $\mathbf{\hat{Q}}=\frac{1}{T}
\mathbf{X}^{\prime }{\mathbb{M}}_{\mathbf{D}}\mathbf{X}\ $and $\mathbf{\ddot{
X}}={\mathbb{M}}_{\mathbf{D}}\mathbf{X}.$ Noting that $\mathbb{M}_{\mathbf{D}
}=\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}},$
we can write $\mathbf{\ddot{X}}=$bdiag$\left( \mathbf{\ddot{x}}_{1},...,
\mathbf{\ddot{x}}_{N}\right) ,$ where $\mathbf{\ddot{x}}_{1}=(\ddot{x}
_{i1},...,\ddot{x}_{iT})^{\prime }$ and $\ddot{x}_{it}=x_{it}-\frac{1}{N}
\sum_{j=1}^{N}x_{jt}-\frac{1}{T}\sum_{s=1}^{T}x_{is}+\frac{1}{NT}
\sum_{j=1}^{N}\sum_{s=1}^{T}x_{js}.$ In addition,
\begin{equation*}
\mathbf{\hat{Q}}=\frac{1}{T}\mathbf{\ddot{X}}^{\prime }\mathbf{\ddot{X}}=
\frac{1}{T}\mathbf{X}^{o\prime }{\mathbb{M}}_{\mathbf{D}}\mathbf{X}^{o},
\end{equation*}
where $\mathbf{X}^{o}=$bdiag$\left( \mathbf{x}_{1}^{o},...,\mathbf{x}
_{N}^{o}\right) ,$ $\mathbf{x}_{i}^{o}=(x_{i1}^{o},...,x_{iT}^{o})^{\prime
}, $ and $x_{it}^{o}$ denotes the part of\ $x_{it}$ with its additive
individual and time effects (if any) removed. For example, if $x_{it}$ is
generated from ($\ref{x_eqn2}$), then $x_{it}^{o}={\Greekmath 010D} _{i}^{\prime
}f_{t}+v_{it}$. In the case where $x_{it}$ does not contain any additive
term that is varying along either $i$ or $t$ but not both, we have that $
x_{it}^{o}=x_{it}$ and $\mathbf{x}_{i}^{o}=\mathbf{x}_{i}.$ Define
\begin{equation*}
\hat{{\Greekmath 0121}}_{j}=\sum_{i=1}^{N}[\mathbf{\hat{Q}}^{-1}]_{ij}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }{\Greekmath 011F}
_{j}^{\prime }=\frac{1}{T}\sum_{i=1}^{N}\hat{{\Greekmath 0121}}_{i}\mathbf{x}
_{i}^{o\prime }\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}m_{ij},
\end{equation*}
where $m_{ij}=\mathbf{1}\left\{ i=j\right\} -\frac{1}{N}$ denotes the $
\left( i,j\right) $th element of $\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}.$

Let $\mathcal{F}_{N,0}^{MG}={\Greekmath 011B} \left( \{\mathbf{x}_{i}^{o}\}_{1\leq
i\leq N},F\right) ,$ the minimal sigma-field generated by $\{\mathbf{x}
_{i}^{o}\}_{1\leq i\leq N}$ and $F.\footnote{
If $x_{it}$ is generated from ($\ref{x_eqn2}$), we can equivalently define $
\mathcal{F}_{N,0}^{MG}={\Greekmath 011B} \left( \{\left( \mathbf{{\Greekmath 010D} }_{i},\mathbf{v}
_{i}\right) \}_{1\leq i\leq N},F\right) .$}$ Let $\mathcal{F}
_{N,j}^{MG}={\Greekmath 011B} (\mathcal{F}_{N,0}^{MG},\left\{ \mathbf{u}_{i},{\Greekmath 0115}
_{i}\right\} _{i=1}^{j})$ for $j\in \left[ N\right] .$ Let $
\max_{i}=\max_{1\leq i\leq N},$ and $\min_{i}=\min_{1\leq i\leq N}$. \ Let $
\underline{c},$ $\overline{c}\in \left( 0,\infty \right) $ be some generic
constants that do not depend on $N$ and whose values may vary across places.
Let $E_{A}\left( \cdot \right) =E\left( \cdot |A\right) .$ Let $\tilde{q}
_{ii}=\frac{1}{T}\mathbf{x}_{i}^{o\prime }{\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}
\mathbf{x}_{i}^{o}.$ To study the asymptotic properties of $\hat{{\Greekmath 010C}}
^{MG}, $ we impose the following two assumptions.

\begin{assumption}
\label{Assmp3.1} (i) Conditional on $\mathcal{F}_{N,0}^{MG},$ $\left\{ (
\mathbf{u}_{i},{\Greekmath 0115} _{i})\right\} _{i\in \left[ N\right] }$ are
independent across $i$ with zero mean and $E_{\mathcal{F}_{N,0}^{MG}}[\left
\Vert \mathbf{u}_{i}\right\Vert ^{4+2{\Greekmath 010E} }+\left\Vert {\Greekmath 0115}
_{i}\right\Vert ^{4+2{\Greekmath 010E} }]\lesssim 1$ for some ${\Greekmath 010E} >0;$

(ii) $\left\{ {\Greekmath 0111} _{i}\right\} _{i\in \left[ N\right] }$ are independent
across $i$ with zero mean, $\frac{1}{N}\sum_{i=1}^{N}E({\Greekmath 0111} _{i}{\Greekmath 0111}
_{i}^{\prime })\rightarrow \Omega _{{\Greekmath 0111} },$ and $E\left\Vert {\Greekmath 0111}
_{i}\right\Vert ^{2+{\Greekmath 010E} }\leq \overline{c};$

(iii) $E_{\mathcal{F}_{N,i-1}^{MG}}(\mathbf{u}_{i}{\Greekmath 0111} _{i}^{\prime })=0;$

(iv) $E_{F}\left\Vert \mathbf{x}_{i}^{o}\right\Vert ^{4+2{\Greekmath 010E} }\lesssim 1.$
\
\end{assumption}

\begin{assumption}
\label{Assmp3.2}(i) $\tilde{q}_{ii}$ has full rank almost surely (a.s.) and $
\frac{1}{N}\sum_{i=1}^{N}\left\Vert \tilde{q}_{ii}^{-1}\right\Vert
^{2}=O_{p}(1)$;

(ii) $\frac{1}{N}\sum_{i=1}^{N}{\Greekmath 011F} _{i}^{\prime }E_{\mathcal{F}_{N,0}^{MG}}[
\mathbf{u}_{i}^{\dag }\mathbf{u}_{i}^{\dag \prime }]{\Greekmath 011F} _{i}\overset{p}{
\rightarrow }\Omega _{1}^{MG}$ and $\Omega _{1}^{MG}$ is positive definite
almost surely (a.s.).
\end{assumption}

Assumption \ref{Assmp3.1}(i)--(iii) impose some dependence and moment
conditions on the unobserved components, i.e., $\mathbf{u}_{i},{\Greekmath 0115} _{i},$
and ${\Greekmath 0111} _{i}.$ Noting that $\mathbf{X}^{o}$ is measurable with respect to $
\mathcal{F}_{N,0}^{MG},$ Assumption \ref{Assmp3.1}(i) implies the
independence of $(\mathbf{u}_{i},{\Greekmath 0115} _{i})$ across $i$ conditional on $
\mathbf{X}^{o}.$ The zero conditional mean condition ensures that the
composite error term $u_{it}^{\dag }$ ($={\Greekmath 0115} _{i}^{\prime }f_{t}+u_{it}$
) can be treated as a usual error term in least squares regression as the
individual and time effects can be partialled out via the usual partitioned
regressions. Clearly, Assumption \ref{Assmp3.1}(i) essentially requires that
the regressors $x_{it}$, after its individual and time effects are removed,
are strictly exogenous so that a dynamic panel is ruled out. Assumption \ref
{Assmp3.1}(ii) imposes standard conditions on the heterogeneous parameters $
{\Greekmath 0111} _{i}.$ Note that we allow $\Omega _{{\Greekmath 0111} }$ to be positive semidefinite
so that some elements in ${\Greekmath 010C} _{i}$ are allowed to be homogenous across $
i. $ Assumption \ref{Assmp3.1}(iii) requires the conditional
uncorrelatedness between $\mathbf{u}_{i}$ and ${\Greekmath 0111} _{i}$ because $E(\mathbf{
u}_{i}|\mathcal{F}_{N,i-1}^{MG})=0$ under Assumption \ref{Assmp3.1}(i).
Assumption \ref{Assmp3.1}(iv) imposes some moment conditions on $\mathbf{x}
_{i}^{o}$ that will be used to verify certain Lyapunov conditions. In case
where $x_{it}$ is generated according to (\ref{x_eqn2}), this moment
condition can be replaced by\medskip

\noindent \textbf{Assumption 3.1(iv*)} $E\left\Vert {\Greekmath 010D} _{i}\right\Vert
^{4+2{\Greekmath 010E} }+E\left\Vert \mathbf{v}_{i}\right\Vert ^{4+2{\Greekmath 010E} }\leq
\overline{c},$ where $\mathbf{v}_{i}=(v_{i1},...,v_{iT})^{\prime }.$\medskip

It is worth mentioning that we do not impose any conditions on the
individual and time effects, viz. ${\Greekmath 010B} _{i\cdot }$ and ${\Greekmath 010B} _{\cdot t}$
for the $y$-equation and ${\Greekmath 0116} _{i\cdot }$ and ${\Greekmath 0116} _{\cdot t}$ in the $x$
-equation if $x_{it}$ is generated according to (\ref{x_eqn2}) as they are
partialled out in the LSDV regression. Neither do we impose any condition on
the unobserved factor $f_{t}$ due to our focus on the fixed $T$ framework.
In other words, we allow various kinds of dependence, trending or
nonstationary behavior in ${\Greekmath 010B} _{\cdot t},$ ${\Greekmath 0116} _{\cdot t}$ and $f_{t}.$
In addition, we do not impose any dependence condition on $\left\{
f_{t},u_{it},x_{it}\right\} $ along the time dimension.

Assumption \ref{Assmp3.2}(i) specifies some rank and moment conditions on
the use of TW-MG\ estimator when a time effect is present in the $y$
-equation, which implicitly requires that $T>K_{x}+1.$ When $T\leq K_{x}$,
there is no way to consider the MG estimator even without the individual or
time effects in the $y$-equation. In this case, one can only consider the
pooled regression to estimate the population average marginal effect ${\Greekmath 010C}
^{0}$ as in the next section. When $T=K_{x}+1,$ we run into the irregular
case discussed in \cite{pesaran2024trimmed} even if $\tilde{q}_{ii}$ may
still be full rank a.s.. Note that in the fixed $T$-framework, $\tilde{q}
_{ii}$ can be of full rank but with minimum eigenvalue that may take values
with positive probability in the neighborhood of 0. A sufficient condition
for the second part of Assumption \ref{Assmp3.2}(i) to hold is $E\left\Vert
\tilde{q}_{ii}^{-1}\right\Vert ^{2}\leq \bar{c}.$ Assumption \ref{Assmp3.2}
(ii)\ specifies a convergence condition on the conditional second moments of
${\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dag },$ which appears in the Bahadur
representation of the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}.$ Due to the
presence of the unobserved factor component, $F,$ which enters both ${\Greekmath 011F}
_{i}$ and $\mathbf{u}_{i}^{\dag },$ the probability limit object $\Omega
_{1}^{MG}$ is random unless all factors inside $F$ are nonrandom. This
condition facilitates the application of a stable martingale CLT as
introduced in \cite{hausler2015stable}.

\begin{remark}
\label{Rmk3.1} \textbf{(Allowance of Correlated Random Coefficients)} We
allow correlation between ${\Greekmath 0111} _{i}$ and $x_{it}$ via various components in
$x_{it}$ including ${\Greekmath 0116} _{i\cdot },$ ${\Greekmath 010D} _{i}$ and $v_{it}$ if $x_{it}$
is generated according to (\ref{x_eqn2}). This largely broadens potential
applications of \textit{correlated random coefficient} (CRC) panel data
models. For a textbook treatment on random coefficient panels, see Chapter
13 in \cite{Hsiao2022}. For some recent applications and studies on CRC
panels, see \cite{arellano2012identifying}, \cite{graham2012identification},
and \cite{lu2023}, among others.
\end{remark}

\subsection{Asymptotic normality of $\hat{\protect{\Greekmath 010C}}^{MG}$ \label{Sec3.2}
}

To state the first main result in this paper, let $\mathcal{F}
_{0}^{MG}={\Greekmath 011B} \left( \{\mathbf{x}_{i}^{o}\}_{1\leq i\leq \infty
},F\right) .$ The following theorem reports the stable convergence of $\sqrt{
N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0}).$

\begin{theorem}
\label{Thm3.1} Suppose that Assumptions \ref{Assmp3.1}--\ref{Assmp3.2} hold.
Then
\begin{equation}
\sqrt{N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})=\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left(
{\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dagger }+{\Greekmath 0111} _{i}\right) \rightarrow
(\Omega ^{MG})^{1/2}\mathbb{Z}\mathcal{\ \ F}_{0}^{MG}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-stably as }
N\rightarrow \infty ,  \label{SCLT1}
\end{equation}
where $\Omega ^{MG}=\Omega _{1}^{MG}+\Omega _{{\Greekmath 0111} },\ \mathbb{Z}\backsim
N\left( 0,\mathbb{I}_{K_{x}}\right) ,$ and $\mathbb{Z}$ is independent of $
\mathcal{F}_{0}^{MG}.$
\end{theorem}

Theorem \ref{Thm3.1} indicates that the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}$
is $\sqrt{N}$-consistent and follows a stable CLT under some regularity
conditions. The sequence $\{\sqrt{N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})\}_{N\geq
1} $ converges $\mathcal{F}$-stably as $N\rightarrow \infty $ to a mixture
normal distribution whose variance is random but measurable with respect to
the limit sigma-field $\mathcal{F}_{0}^{MG}$. We refer the readers directly
to \cite{hausler2015stable} and \cite{van2023weak} for the notion of stable
convergence and stable CLT. For recent applications of the concept of stable
convergence in econometrics and statistics, see \cite{lee2019stable}, \cite
{cavaliere2020inference}, \cite{Jin_Miao_Su2021}, \cite{mikusheva2021many}
and \cite{hahn2024econometric}. It is worth mentioning that the usual weak
convergence concept, a standard tool for asymptotic analysis in
econometrics, is not adequate for our asymptotic analysis. Stable
convergence is a middle ground between convergence in probability and
convergence in distribution. It is weaker than convergence in probability
but stronger than weak convergence. Importantly, it allows for a stochastic
normalization in the central limit theorem. Intuitively, stable convergence
can be thought of as a weak convergence conditional on certain sigma-algebra
($\mathcal{F}_{0}^{MG}$ here) so that it can be recast as a convergence of
certain conditional distributions.

\subsection{Asymptotic inference on $\protect{\Greekmath 010C} ^{0}$ \label{Sec3.3}}

In this subsection, we address the problem of inference on $\mathbb{S}
^{\prime }{\Greekmath 010C} ^{0}$ based on the jackknife method for the TW-MG
estimation. We add the following condition.

\begin{assumption}
\label{Assmp3.3}(i)$\frac{1}{N}\sum_{i=1}^{N}\left\Vert \tilde{q}
_{ii}^{-1}\right\Vert ^{2}\left\Vert \mathbf{x}_{i}^{o}\right\Vert
^{2s}=O_{p}(1)$ for $s=0,1$.

(ii) ${\Greekmath 010E} >1.$
\end{assumption}

Assumption \ref{Assmp3.3}(i) strengthens Assumption \ref{Assmp3.2}(i). As
sufficient condition for it to hold is $\frac{1}{N}\sum_{i=1}^{N}E[\left
\Vert \tilde{q}_{ii}^{-1}\right\Vert ^{2}\left\Vert \mathbf{x}
_{i}^{o}\right\Vert ^{2s}]=O(1)$ for $s=0,1.$ Assumption \ref{Assmp3.3}(ii)
strengthens the moment condition in \ref{Assmp3.1} by requiring ${\Greekmath 010E} >1.$
Both parts of Assumption \ref{Assmp3.3} are reasonable because consistent
estimation of the variance-covariance matrix generally requires stronger
conditions than those needed for the asymptotic distribution result.

The following theorem establishes the consistency of $\hat{\Omega}^{MG}\ $
and the distributional limit for a normalized version of $\sqrt{N}\mathbb{S}
^{\prime }(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0}).$

\begin{theorem}
\label{Thm3.2}Suppose Assumptions \ref{Assmp3.1}--\ref{Assmp3.3} hold. Then

(i) $\hat{\Omega}^{MG}=\frac{1}{N}\sum_{j=1}^{N}({\Greekmath 011F} _{j}^{\prime }\mathbf{u
}_{j}^{\dagger }+{\Greekmath 0111} _{j})({\Greekmath 011F} _{j}^{\prime }\mathbf{u}_{j}^{\dagger
}+{\Greekmath 0111} _{j})^{\prime }+o_{p}\left( 1\right) =\Omega ^{MG}+o_{p}\left(
1\right) ;$

(ii) $(\mathbb{S}^{\prime }\hat{\Omega}^{MG}\mathbb{S})^{-1/2}\sqrt{N}
\mathbb{S}^{\prime }(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})\overset{d}{\rightarrow }
N\left( 0,1\right) \ $as $N\rightarrow \infty .$
\end{theorem}

Theorem \ref{Thm3.2}(i) indicates that we can estimate $\Omega ^{MG}$
consistently via the jackknife method and Theorem \ref{Thm3.2}(ii) justifies
the asymptotic validity of the jackknife confidence interval defined in (\ref
{CI1}).

\begin{remark}
\label{Rmk3.2} \textbf{(High dimensional jackknife)} The jackknife method is
a popular method to estimate the asymptotic variance of an estimator whose
dimension does not grow with the sample size; see Chapter 10.3 in \cite
{hansen2022econometrics} for a brief account. However, our TW-MG estimator $
\hat{{\Greekmath 010C}}^{MG}$ is based on $\mathbf{\hat{{\Greekmath 010C}}},$ whose row dimension
diverges to infinity linearly in the sample size $N.$ To study the
asymptotic expansion of $\hat{{\Greekmath 010C}}^{MG\left( -i\right) },$ we have to
study that of $\mathbf{\hat{{\Greekmath 010C}}}^{\left( -i\right) },$ which is an $
N_{1}K_{x}\times 1$ vector. The change of the row dimension of $\mathbf{\hat{
{\Greekmath 010C}}}$ to that of $\mathbf{\hat{{\Greekmath 010C}}}^{\left( -i\right) }$ substantially
complicates the asymptotic analysis. Let $\mathbf{\hat{Q}}^{\left( -i\right)
}\equiv \frac{1}{T}\mathbf{X}^{\left( -i\right) \prime }(\mathbb{M}_{\mathbf{
{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}^{\left(
-i\right) }=\frac{1}{T}\mathbf{X}^{o\left( -i\right) \prime }(\mathbb{M}_{
\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}
^{o\left( -i\right) },$ where $\mathbf{X}^{o\left( -i\right) \prime }$ is
the leave-one-out version of $\mathbf{X}^{o}.$ Noting that
\begin{equation*}
\mathbf{\hat{{\Greekmath 010C}}}=\mathbf{\hat{Q}}^{-1}\frac{1}{T}\mathbf{X}^{\prime }(
\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})
\mathbf{Y}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }\mathbf{\hat{{\Greekmath 010C}}}^{\left( -i\right) }=\left[
\mathbf{\hat{Q}}^{\left( -i\right) }\right] ^{-1}\frac{1}{T}\mathbf{X}
^{\left( -i\right) \prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes
\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{Y}^{\left( -i\right) },\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }
\end{equation*}
we make the following decompositions:
\begin{eqnarray*}
\mathbf{\hat{Q}} &\equiv &\frac{1}{T}\mathbf{X}^{o\prime }\left( {\mathbb{I}}
_{N}\otimes {\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\right) \mathbf{X}^{o}-\frac{1
}{T}\mathbf{X}^{o\prime }\left( {\mathbb{P}}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes {
\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\right) \mathbf{X}^{o}\equiv \mathbf{Q}
^{\left( d\right) }-\mathbf{CC}^{\prime }\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and} \\
\mathbf{\hat{Q}}^{\left( -i\right) } &=&\frac{1}{T}\mathbf{X}^{o\left(
-i\right) \prime }\left( {\mathbb{I}}_{N_{1}}\otimes {\mathbb{M}}_{\mathbf{
{\Greekmath 0113} }_{T}}\right) \mathbf{X}^{o\left( -i\right) }-\frac{1}{T}\mathbf{X}
^{o\left( -i\right) \prime }({\mathbb{P}}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes {
\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}^{o\left( -i\right) }\equiv
\mathbf{Q}^{\left( -i,d\right) }-\mathbf{C}^{\left( -i\right) }\mathbf{C}
^{\left( -i\right) \prime },
\end{eqnarray*}
where $\mathbf{C}=(NT)^{-1/2}\mathbf{X}^{o\prime }[$\textbf{${\Greekmath 0113} $}$
_{N}\otimes {\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}]$ and $\mathbf{C}^{\left(
-i\right) }=\left( N_{1}T\right) ^{-1/2}\mathbf{X}^{o\left( -i\right) \prime
}[$\textbf{${\Greekmath 0113} $}$_{N_{1}}\otimes {\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}].$
By the Sherman-Morrison Woodbury matrix identity (e.g., Fact 6.4.31 in \cite
{bernstein2005}), we have
\begin{equation*}
\left[ \mathbf{\hat{Q}}^{\left( -i\right) }\right] ^{-1}=\left[ \mathbf{Q}
^{\left( -i,d\right) }\right] ^{-1}\left\{ \mathbf{Q}^{\left( -i,d\right) }+
\mathbf{C}^{\left( -i\right) }[\mathbb{I}_{K_{x}T}+\mathbf{C}^{\left(
-i\right) \prime }\mathbf{C}^{\left( -i\right) }]^{-1}\mathbf{C}^{\left(
-i\right) \prime }\right\} \left[ \mathbf{Q}^{\left( -i,d\right) }\right]
^{-1}.
\end{equation*}
A similar expression holds for $\mathbf{\hat{Q}}^{-1}.$ Due to the mismatch
of the dimensions of various matrix pairs such as $\mathbf{Q}^{\left(
-i,d\right) }$ and $\mathbf{Q}^{\left( d\right) },$ $\mathbf{C}^{\left(
-i\right) }$ and $\mathbf{C},$ $\mathbf{X}^{\left( -i\right) }$ and $\mathbf{
X},$ and $\mathbf{Y}^{\left( -i\right) }$ and $\mathbf{Y},$ one key step in
our analysis is to augment the leave-one-out version of various
matrices/vectors to match the dimension of the original version. For
example, we augment the $N_{1}K_{x}\times N_{1}K_{x}$ block diagonal matrix $
\mathbf{Q}^{\left( -i,d\right) }$ to $\mathbf{\breve{Q}}^{\left( -i,d\right)
}$ (an $NK_{x}\times NK_{x}$ block diagonal matrix) such that
\begin{equation*}
\left[ \mathbf{\breve{Q}}^{\left( -i,d\right) }\right] _{jj}=\left\{
\begin{array}{ll}
\left[ \mathbf{Q}^{\left( -i,d\right) }\right] _{jj} & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }j<i \\
\mathbf{0}_{K_{x}\times K_{x}} & \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }j=i \\
\left[ \mathbf{Q}^{\left( -i,d\right) }\right] _{j-1,j-1}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ } &
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }j>i
\end{array}
\right. ,
\end{equation*}
where $\left[ \mathbf{Q}^{\left( -i,d\right) }\right] _{jj}$ denotes the $
\left( j,j\right) $th block of $\mathbf{Q}^{\left( -i,d\right) }$ of size $
K_{x}\times K_{x}.$ We do similar things for matrices $\mathbf{C}^{\left(
-i\right) },$ $\mathbf{X}^{\left( -i\right) },$ $\mathbf{Y}^{\left(
-i\right) },$ etc., to obtain their augmented versions $\mathbf{\breve{C}}
^{\left( -i\right) },$ $\mathbf{\breve{X}}^{\left( -i\right) },$ $\mathbf{
\breve{Y}}^{\left( -i\right) },$ etc. By using these augmented matrices
associated with some tedious matrix algebra, we can ultimately establish the
link between $\hat{{\Greekmath 010C}}^{MG\left( -i\right) }-\frac{1}{N_{1}}\sum_{j\neq
i}^{N}{\Greekmath 010C} _{j}$ and $\hat{{\Greekmath 010C}}^{MG}-\frac{1}{N}\sum_{j=1}^{N}{\Greekmath 010C} _{j}:$
\begin{equation}
\hat{{\Greekmath 010C}}^{MG\left( -i\right) }-\frac{1}{N_{1}}\sum_{j\neq i}^{N}{\Greekmath 010C}
_{j}=\frac{N}{N_{1}}\left( \hat{{\Greekmath 010C}}^{MG}-\frac{1}{N}\sum_{j=1}^{N}{\Greekmath 010C}
_{j}\right) +\Delta _{i2}^{MG}+\Delta _{i3}^{MG}+o_{p}(N^{-1}),
\label{jackknife0}
\end{equation}
where $\Delta _{i2}^{MG}\equiv -\frac{1}{N_{1}T}\hat{{\Greekmath 0121}}_{i}\mathbf{x}
_{i}^{o\prime }{\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\mathbf{u}_{i}^{\dag },$ $
\Delta _{i3}^{MG}\equiv \frac{1}{NN_{1}}\sum_{j=1}^{N}\hat{{\Greekmath 0121}}_{j}\frac{1
}{T}\mathbf{x}_{j}^{o\prime }{\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\mathbf{u}
_{i}^{\dag },\ $and $o_{p}(N^{-1})$ holds uniformly in $i\in \left[ N\right]
.$ $\Delta _{i2}^{MG}+\Delta _{i3}^{MG}$ plays the role of a summand in the
influence function of $\hat{{\Greekmath 010C}}^{MG}-\frac{1}{N}\sum_{j=1}^{N}{\Greekmath 010C} _{j}$
. The result in (\ref{jackknife0}) is a fundamental result that helps
deliver the main result in Theorem \ref{Thm3.2}(i).
\end{remark}

\begin{remark}
\label{Rmk3.3}\textbf{(Alternative approaches)} In the MG estimation
literature, one commonly uses the following formula to estimate the
asymptotic variance of $\hat{{\Greekmath 010C}}^{MG}$
\begin{equation*}
\frac{1}{N}\sum_{i=1}^{N}\left( \hat{{\Greekmath 010C}}_{i}-\hat{{\Greekmath 010C}}^{MG}\right)
\left( \hat{{\Greekmath 010C}}_{i}-\hat{{\Greekmath 010C}}^{MG}\right) ^{\prime }.
\end{equation*}
However, it turns out that this variance estimator is hard to justify in our
framework. The main reason is that $\hat{{\Greekmath 010C}}_{i}$ is not a consistent
estimator of ${\Greekmath 010C} _{i}$ in the fixed $T$ framework. Alternatively, we may
consider cross-sectional bootstrap by resampling the data along the
cross-sectional dimension. We find that the cross-sectional bootstrap and
our leave-one-out jackknife method perform similarly in simulations.
However, the asymptotic validity of the cross-sectional bootstrap is much
harder to justify theoretically; therefore, we propose the jackknife method
here.
\end{remark}

\section{Pooled Estimation and Inference\label{Sec4}}

In this section, we consider the pooled estimation for the model in (\ref
{y_eqn1}) and study the inference theory.

\subsection{Estimation and inference procedure \label{Sec4.1}}

We rewrite the model in (\ref{y_eqn1}) as follows:
\begin{equation}
y_{it}={\Greekmath 010C} ^{0\prime }x_{it}+{\Greekmath 010B} _{i\cdot }+{\Greekmath 010B} _{\cdot
t}+u_{it}^{\ddag },  \label{y_eqn3}
\end{equation}
where $u_{it}^{\ddag }=u_{it}+{\Greekmath 0115} _{i}^{\prime }f_{t}+{\Greekmath 0111} _{i}^{\prime
}x_{it}$. Let $X=\mathbf{(x}_{1}^{\prime },...,\mathbf{x}_{N}^{\prime
})^{\prime }$ and $\mathbf{{\Greekmath 0111} }=({\Greekmath 0111} _{1}^{\prime },...,{\Greekmath 0111} _{N}^{\prime
})^{\prime }.$ Then we can rewrite the model in (\ref{y_eqn3}) in matrix
form:
\begin{equation}
\mathbf{Y}=X{\Greekmath 010C} ^{0}+\mathbf{D{\Greekmath 010B} }+\mathbf{U}^{\ddag },  \label{Y_eqn2}
\end{equation}
where $\mathbf{U}^{\ddag }\equiv \mathbf{U}+(\mathbb{I}_{N}\otimes F)\mathbf{
{\Greekmath 0115} +X{\Greekmath 0111} }.$ By partialling out $\mathbf{D}$ in (\ref{Y_eqn2}), we
obtain the following pooled estimator of ${\Greekmath 010C} ^{0}$:
\begin{equation}
\hat{{\Greekmath 010C}}^{P}\equiv \left( X^{\prime }{\mathbb{M}}_{\mathbf{D}}X\right)
^{-1}X^{\prime }{\mathbb{M}}_{\mathbf{D}}\mathbf{Y.}  \label{B_tilde1}
\end{equation}

Below we show that $\hat{{\Greekmath 010C}}^{P}$ is asymptotically mixture normal under
some regularity conditions. We propose a jackknife method to estimate its
asymptotic variance. As in (\ref{B_tilde1}), we can obtain the leave-one-out
pooled estimator of ${\Greekmath 010C} ^{0}$ by leaving out the observations for the $i$
th cross-section unit:
\begin{equation}
\hat{{\Greekmath 010C}}^{P\left( -i\right) }=\left[ X^{\left( -i\right) \prime }(\mathbb{
M}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }
_{T}})X^{\left( -i\right) }\right] ^{-1}X^{\left( -i\right) \prime }(\mathbb{
M}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{
Y}^{\left( -i\right) },  \label{theta_leave_one_out3}
\end{equation}
where $X^{\left( -i\right) }$ is the leave-one-out version\ of $X:X^{\left(
-i\right) }=(\mathbf{x}_{1}^{\prime },...,\mathbf{x}_{i-1}^{\prime },\mathbf{
x}_{i+1}^{\prime },...,\mathbf{x}_{N}^{\prime })^{\prime }$. Given the
jackknife estimates $\{\hat{{\Greekmath 010C}}^{P\left( -i\right) }\}_{i\in \left[ N
\right] },$ we propose to estimate the asymptotic variance of $\sqrt{N}(\hat{
{\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0})$ by
\begin{equation*}
\hat{\Omega}^{P}=(N-1)\mathop{\displaystyle \sum }\limits_{i=1}^{N}\left( \hat{{\Greekmath 010C}}^{P\left(
-i\right) }-\overline{\hat{{\Greekmath 010C}}}^{_{P}}\right) \left( \hat{{\Greekmath 010C}}^{P\left(
-i\right) }-\overline{\hat{{\Greekmath 010C}}}^{_{P}}\right) ^{\prime },
\end{equation*}
where $\overline{\hat{{\Greekmath 010C}}}^{_{P}}=\frac{1}{N}\sum_{i=1}^{N}\hat{{\Greekmath 010C}}
^{P\left( -i\right) }.$ Then one can construct the confidence intervals for
element of $\mathbb{S}{\Greekmath 010C} ^{0}$ as in (\ref{CI1}) with $\hat{\Omega}^{MG}$
replaced by $\hat{\Omega}^{P}$.

\subsection{Asymptotic properties \label{Sec4.2}}

To study the asymptotic properties of $\hat{{\Greekmath 010C}}^{P},$ we add some
notations and assumptions. Let $\mathbf{\dot{x}}_{i}\equiv \mathbb{M}_{
\mathbf{{\Greekmath 0113} }_{T}}\mathbf{x}_{i}=(\dot{x}_{i1},...,\dot{x}_{iT})^{\prime }$
and $\mathbf{\dot{u}}_{i}=\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}\mathbf{u}_{i}=(
\dot{u}_{i1},...,\dot{u}_{iT})^{\prime }.$ Let $\mathbf{\dot{X}}\equiv {
\mathbb{({\mathbb{I}}_{N}\otimes M}}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}=$bdiag$
(\mathbf{\dot{x}}_{1},...,\mathbf{\dot{x}}_{N})$ and $\mathbf{\dot{U}}
^{\ddagger }=\mathbf{\dot{U}}+(\mathbb{I}_{N}\otimes \dot{F})\mathbf{{\Greekmath 0115}
+\dot{X}{\Greekmath 0111} }=(\mathbf{\dot{u}}_{1}^{\ddagger \prime },...,\mathbf{\dot{u}}
_{N}^{\ddagger \prime })^{\prime }$ where $\mathbf{\dot{u}}_{i}^{\ddagger }=(
\dot{u}_{i1}^{\ddagger },...,\dot{u}_{iT}^{\ddagger })^{\prime }$ and $\dot{u
}_{it}^{\ddagger }=\dot{u}_{it}+\dot{f}_{t}^{\prime }{\Greekmath 0115} _{i}+\dot{x}
_{it}^{\prime }{\Greekmath 0111} _{i}.$ Let
\begin{equation*}
\hat{Q}\equiv \frac{1}{NT}X^{\prime }{\mathbb{M}}_{\mathbf{D}}X=\frac{1}{NT}
\ddot{X}^{\prime }\ddot{X},
\end{equation*}
where $\ddot{X}={\mathbb{M}}_{\mathbf{D}}X.$ Let $\mathcal{F}_{0}^{P}={\Greekmath 011B}
\left( \{\mathbf{\dot{x}}_{i}\}_{1\leq i\leq \infty },F\right) ,$ $\mathcal{F
}_{N,0}^{P}={\Greekmath 011B} (\{\mathbf{\dot{x}}_{i}\}_{1\leq i\leq N},F)$ and $
\mathcal{F}_{N,j}^{P}={\Greekmath 011B} (\mathcal{F}_{N,0}^{P},\left\{ (\mathbf{u}
_{i},{\Greekmath 0115} _{i},{\Greekmath 0111} _{i})\right\} _{i=1}^{j})$ for $j\in \left[ N\right]
. $ Note that $\mathcal{F}_{N,0}^{P}$ (resp. $\mathcal{F}_{0}^{P}$) slightly
differs from $\mathcal{F}_{N,0}^{MG}$ (resp. $\mathcal{F}_{0}^{MG}$) because
$x_{it}^{\prime }{\Greekmath 0111} _{i}$ now enters the error term $u_{it}^{\ddag }$ in
the pooled regression and the time effects in $x_{it}$ (if any) cannot be
removed in the estimation procedure.

\begin{assumption}
\label{Assmp4.1} (i) Conditional on $\mathcal{F}_{N,0}^{P},$ $\left\{ (
\mathbf{u}_{i},{\Greekmath 0115} _{i},{\Greekmath 0111} _{i})\right\} _{i\in \left[ N\right] }$ are
independent across $i$ with zero mean and $E_{\mathcal{F}_{N,0}^{P}}[\left
\Vert \mathbf{u}_{i}\right\Vert ^{4+2{\Greekmath 010E} }+\left\Vert {\Greekmath 0115}
_{i}\right\Vert ^{4+2{\Greekmath 010E} }+\left\Vert {\Greekmath 0111} _{i}\right\Vert ^{4+2{\Greekmath 010E}
}]\lesssim 1$ for some ${\Greekmath 010E} >0;$

(ii) $E_{F}[\left\Vert \mathbf{\dot{x}}_{i}\right\Vert ^{4+2{\Greekmath 010E}
}]\lesssim 1.$
\end{assumption}

\begin{assumption}
\label{Assmp4.2}(i) $\hat{Q}\overset{p}{\rightarrow }Q$ and $Q$ is positive
definite a.s.;

(ii) $\frac{1}{NT^{2}}\sum_{i=1}^{N}\mathbf{\ddot{x}}_{i}^{\prime }E_{
\mathcal{F}_{N,0}^{P}}\left[ \mathbf{\dot{u}}_{i}^{\ddag }\mathbf{\dot{u}}
_{i}^{\ddag \prime }\right] \mathbf{\ddot{x}}_{i}=\Omega ^{P}+o_{p}\left(
1\right) \ $and $\Omega ^{P}$ is positive definite a.s..
\end{assumption}

Assumption \ref{Assmp4.1}(i) parallels Assumption \ref{Assmp3.1}(i)--(iii)
and it imposes some dependence and moment conditions on the unobserved
components, i.e., $\mathbf{u}_{i},{\Greekmath 0115} _{i},$ and ${\Greekmath 0111} _{i}.$ Since $
x_{it}^{\prime }{\Greekmath 0111} _{i}$ is a treated as part of the composite error term $
u_{it}^{\ddagger },$ we now need to assume that ${\Greekmath 0111} _{i}$ has finite $
\left( 4+2{\Greekmath 010E} \right) $th moment conditional on $\mathcal{F}_{N,0}^{P}.$
The condition that $E\left( {\Greekmath 0111} _{i}|\mathcal{F}_{N,0}^{P}\right) =\mathbf{0
}_{K_{x}}$ rules out potential correlations between ${\Greekmath 0111} _{i}$ and $\mathbf{
\dot{x}}_{i}$ via components in $x_{it}$ other than the individual effects.
That is, if there is any correlation between ${\Greekmath 0111} _{i}$ and $x_{it}$ as in
CRC\ panels, the correlation can only be caused by ${\Greekmath 0111} _{i}$ and a
time-invariant additive component in $x_{it}.$ When $x_{it}$ is generated
according to (\ref{x_eqn2}), ${\Greekmath 0111} _{i}$ can be correlated with ${\Greekmath 0116}
_{i\cdot }$ but not ${\Greekmath 0116} _{\cdot t},$ ${\Greekmath 010D} _{i}$ and $v_{it}.$ Assumption
\ref{Assmp4.1}(ii) parallels Assumption \ref{Assmp3.1}(iv). It is easy to
see that Assumption \ref{Assmp4.1}(ii) implies that $E_{F}[\left\Vert
\mathbf{\dot{x}}_{i}^{o}\right\Vert ^{4+2{\Greekmath 010E} }]\lesssim 1$ and $
E_{F}\left\Vert \mathbf{\ddot{x}}_{i}\right\Vert ^{4+2{\Greekmath 010E} }\lesssim 1.$

Assumption \ref{Assmp4.2}(i)-(ii) parallels Assumption \ref{Assmp3.2}
(i)--(ii). In comparison with Assumption \ref{Assmp3.2}(i), Assumption \ref
{Assmp4.2}(i) only requires the probability limit of $Q$ be positive
definite a.s., which may hold even when $T\leq K_{x}\ $so that the TW-MG
estimator $\hat{{\Greekmath 010C}}^{MG}$ is not well defined. Assumption \ref{Assmp4.2}
(ii)\ specifies a convergence condition on the conditional second moments of
$\mathbf{\ddot{x}}_{i}^{\prime }\mathbf{\dot{u}}_{i}^{\ddag },$ which
appears in the Bahadur representation of the pooled estimator $\hat{{\Greekmath 010C}}
^{P}.$ Due to the presence of the unobserved factor component, $F,$ which
surely enters $\mathbf{\dot{u}}_{i}^{\ddag }$ and may also enter $\mathbf{
\ddot{x}}_{i}$ if $x_{it}$ also contains a factor structure, the probability
limit object $\Omega ^{P}$ is random unless all factors inside $F$ are
nonrandom. This condition facilitates the application of a stable martingale
CLT as introduced in \cite{hausler2015stable}.

The following theorem reports the stable convergence of $\sqrt{N}(\hat{{\Greekmath 010C}}
^{P}-{\Greekmath 010C} ^{0}).$

\begin{theorem}
\label{Thm4.1} Suppose that Assumptions \ref{Assmp4.1}--\ref{Assmp4.2} hold.
Then
\begin{equation*}
\sqrt{N}(\hat{{\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0})=\hat{Q}^{-1}\frac{1}{\sqrt{N}}
\sum_{i=1}^{N}\frac{1}{T}\mathbf{\ddot{x}}_{i}^{\prime }\mathbf{\dot{u}}
_{i}^{\ddag }\rightarrow (Q^{-1}\Omega ^{P}Q^{-1})^{1/2}\mathbb{Z}\mathcal{\
\ F}_{0}^{P}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-stably as }N\rightarrow \infty ,
\end{equation*}
where $\mathbb{Z}\backsim N\left( 0,\mathbb{I}_{K_{x}}\right) $ and is
independent of $\mathcal{F}_{0}^{P}.$
\end{theorem}

Theorem \ref{Thm4.1} shows that the sequence $\{\sqrt{N}(\hat{{\Greekmath 010C}}
^{MG}-{\Greekmath 010C} ^{0})\}_{N\geq 1}$ converges $\mathcal{F}_{0}^{P}$-stably as $
N\rightarrow \infty $ to a mixture normal distribution whose variance is
random but measurable with respect to the limit sigma-field $\mathcal{F}
_{0}^{P}$.

The following theorem establishes the consistency of $\hat{\Omega}^{P}.$

\begin{theorem}
\label{Thm4.2}Suppose Assumptions \ref{Assmp4.1}--\ref{Assmp4.2} hold. Then

(i) $\hat{\Omega}^{P}=\frac{1}{N}\sum_{j=1}^{N}\mathbf{\ddot{x}}_{j}^{\prime
}\mathbf{\dot{u}}_{j}^{\ddag }\mathbf{\dot{u}}_{j}^{\ddag \prime }\mathbf{
\ddot{x}}_{j}+o_{p}\left( 1\right) =\Omega ^{P}+o_{p}\left( 1\right) ;$

(ii) $(\mathbb{S}^{\prime }\hat{\Omega}^{P}\mathbb{S})^{-1/2}\sqrt{N}\mathbb{
S}^{\prime }(\hat{{\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0})\overset{d}{\rightarrow }N\left(
0,1\right) .$
\end{theorem}

Theorem \ref{Thm4.2} indicates that we can estimate $\Omega ^{P}$
consistently via the jackknife method. It justifies the use of the normal
critical values for asymptotic inference based on the jackknife estimate of
the asymptotic variance.

\subsection{Specification tests for the poolability\label{Sec4.3}}

In this subsection, we propose a test for the poolability based on the
results in the early subsections. In general, $\hat{{\Greekmath 010C}}^{MG}$ is
consistent with ${\Greekmath 010C} ^{0}$ no matter whether one has a homogenous panel
with a slope coefficient ${\Greekmath 010C} ^{0}$ or a heterogenous panel with slope
coefficients whose mean is given by ${\Greekmath 010C} ^{0}$, and it also allows for
arbitrary correlations between ${\Greekmath 010C} _{i}$ and $x_{it}$ in a CRC panel. In
contrast, $\hat{{\Greekmath 010C}}^{P}$ is generally inconsistent when ${\Greekmath 010C} _{i}$ is
correlated with a time-varying component of $x_{it}$ so we allow $\hat{{\Greekmath 010C}}
^{P}$ to converge to something different from ${\Greekmath 010C} ^{0}\ $in this case.
For this reason, we consider the following null hypothesis:
\begin{equation}
\mathbb{H}_{0}:\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{MG}=\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}
_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{P}={\Greekmath 010C} ^{0}.  \label{H0}
\end{equation}
The alternative hypothesis is the negation of $\mathbb{H}_{0}:$
\begin{equation*}
\mathbb{H}_{1}:\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{MG}\neq \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
plim}_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{P}.
\end{equation*}
Below we shall maintain the condition that plim$_{N\rightarrow \infty }\hat{
{\Greekmath 010C}}^{MG}={\Greekmath 010C} ^{0}\ $and assume that the probability limit of $\hat{{\Greekmath 010C}
}^{P}$ is given by ${\Greekmath 010C} ^{P}$ under $\mathbb{H}_{1}.$ A failure to reject $
\mathbb{H}_{0}$ justifies the common practice of pooling all cross-sectional
units to estimate the slope coefficients in a panel despite the potential
presence of slope heterogeneity. If the panel is indeed homogeneous, the
common slope ${\Greekmath 010C} ^{0}$ indicates the marginal effect; otherwise, it
indicates the population average marginal effect.

Under $\mathbb{H}_{0},$ we have
\begin{eqnarray*}
\sqrt{N}(\hat{{\Greekmath 010C}}^{MG}-\hat{{\Greekmath 010C}}^{P}) &=&\sqrt{N}(\hat{{\Greekmath 010C}}
^{MG}-{\Greekmath 010C} ^{0})-\sqrt{N}(\hat{{\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0}) \\
&=&\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left( {\Greekmath 011F} _{i}^{\prime }\mathbf{u}
_{i}^{\dagger }+{\Greekmath 0111} _{i}\right) -\hat{Q}^{-1}\frac{1}{\sqrt{N}}
\sum_{i=1}^{N}\mathbf{\ddot{x}}_{i}^{\prime }\mathbf{\dot{u}}_{i}^{\ddag },
\end{eqnarray*}
which has a well behaved limit distribution if we assume that Assumptions
\ref{Assmp3.1}--\ref{Assmp3.2} and \ref{Assmp4.1}--\ref{Assmp4.2} are all
satisfied. In this case, we can estimate the asymptotic variance of $\sqrt{N}
(\hat{{\Greekmath 010C}}^{MG}-\hat{{\Greekmath 010C}}^{P})$ via the jackknife method:
\begin{equation*}
\hat{\Omega}^{\Delta }=(N-1)\mathop{\displaystyle \sum }\limits_{i=1}^{N}\left( \hat{\Delta}
^{\left( -i\right) }-\overline{\hat{\Delta}}\right) \left( \hat{\Delta}
^{\left( -i\right) }-\overline{\hat{\Delta}}\right) ^{\prime },
\end{equation*}
where $\hat{\Delta}^{\left( -i\right) }=\hat{{\Greekmath 010C}}^{MG\left( -i\right) }-
\hat{{\Greekmath 010C}}^{P\left( -i\right) }$ and $\overline{\hat{\Delta}}=\frac{1}{N}
\sum_{i=1}^{N}\hat{\Delta}^{\left( -i\right) }.$ Following the lead of \cite
{hausman1978specification}, we propose the following test statistic:
\begin{equation}
J_{N}=N(\hat{{\Greekmath 010C}}^{MG}-\hat{{\Greekmath 010C}}^{P})^{\prime }\left[ \hat{\Omega}
^{\Delta }\right] ^{-1}(\hat{{\Greekmath 010C}}^{MG}-\hat{{\Greekmath 010C}}^{P}).  \label{test}
\end{equation}

Below we only assume that Assumptions \ref{Assmp3.1}--\ref{Assmp3.3} are
always satisfied under $\mathbb{H}_{0}$ and $\mathbb{H}_{1}\ $so that both
estimators are well defined. Under $\mathbb{H}_{0},$ we also assume that
Assumptions \ref{Assmp4.1}--\ref{Assmp4.2} are satisfied so that the pooled
estimator $\hat{{\Greekmath 010C}}^{P}$ is consistent with ${\Greekmath 010C} ^{0}$. Under $\mathbb{H
}_{1},$ we dispense with Assumptions \ref{Assmp4.1}--\ref{Assmp4.2} so that $
\hat{{\Greekmath 010C}}^{P}$ does not converge to ${\Greekmath 010C} ^{0}$ in general.

Let $\mathcal{F}_{N}=\mathcal{{\Greekmath 011B} (}\{\mathbf{\dot{x}}_{i}^{o}\}_{i\leq
N},F).$ To study the asymptotic properties of $J_{N}$ under $\mathbb{H}_{0}$
and $\mathbb{H}_{1},$ we add the following two assumptions.

\begin{assumption}
\label{Assmp4.3}$\frac{1}{N}\sum_{i=1}^{N}E_{\mathcal{F}_{N}}\left\{ \left[
{\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dagger }+{\Greekmath 0111} _{i}-Q^{-1}\mathbf{\ddot{x}
}_{i}^{\prime }\mathbf{\dot{u}}_{i}^{\ddag }\right] \left[ {\Greekmath 011F} _{i}^{\prime
}\mathbf{u}_{i}^{\dagger }+{\Greekmath 0111} _{i}-Q^{-1}\mathbf{\ddot{x}}_{i}^{\prime }
\mathbf{\dot{u}}_{i}^{\ddag }\right] ^{\prime }\right\} \overset{p}{
\rightarrow }\Omega ^{\Delta },$ where $\Omega ^{\Delta }$ is positive
definite a.s.
\end{assumption}

\begin{assumption}
\label{Assmp4.4}(i) plim$_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{P}={\Greekmath 010C} ^{P}$
for some ${\Greekmath 010C} ^{P}\neq {\Greekmath 010C} ^{0};$

(ii) $\hat{\Omega}^{\Delta }\overset{p}{\rightarrow }\mathring{\Omega}
^{\Delta },$ where $\mathring{\Omega}^{\Delta }$ is positive definite a.s.
\end{assumption}

Assumption \ref{Assmp4.3} specifies a condition to ensure that the
asymptotic conditional variance of $\sqrt{N}(\hat{{\Greekmath 010C}}^{MG}$ $-\hat{{\Greekmath 010C}}
^{P})$ given $\mathcal{F}_{N}$ is well behaved under $\mathbb{H}_{0}.$
Assumption \ref{Assmp4.4}(i) restricts that $\hat{{\Greekmath 010C}}^{P}$ be a
consistent estimator of ${\Greekmath 010C} ^{P},$ which is different from ${\Greekmath 010C} ^{0}.$
Assumption \ref{Assmp4.4}(ii) imposes that $\hat{\Omega}^{\Delta }$ be
asymptotically nonsingular. We will use Assumptions \ref{Assmp4.3} and \ref
{Assmp4.4} for the study of the asymptotic properties of $J_{N}$ under $
\mathbb{H}_{0}\ $and $\mathbb{H}_{1},$ respectively.

The following theorem reports the asymptotic properties of $J_{N}$ under $
\mathbb{H}_{0}\ $and $\mathbb{H}_{1}.$

\begin{theorem}
\label{Thm4.3} Suppose that Assumptions \ref{Assmp3.1}--\ref{Assmp3.3} hold.

(i) If Assumptions \ref{Assmp4.1}--\ref{Assmp4.3} hold, then under $\mathbb{H
}_{0},$ $J_{N}\overset{d}{\rightarrow }{\Greekmath 011F} ^{2}\left( K_{x}\right) \mathcal{
\ }$as $N\rightarrow \infty .$

(ii) If Assumption \ref{Assmp4.4} holds, then under $\mathbb{H}_{1},$ $
P\left( J_{N}>c_{N}\right) \rightarrow 1$ for any $c_{N}=o(N)$.
\end{theorem}

Theorem \ref{Thm4.3} indicates that $J_{N}$ has the usual asymptotic ${\Greekmath 011F}
^{2}$-distribution under $\mathbb{H}_{0}$ and it diverges to infinity at
rate $N$ under $\mathbb{H}_{1}.\ $As a result, the test is consistent and
one can use the critical values from the ${\Greekmath 011F} ^{2}\left( K_{x}\right) $
distribution to make a decision: one rejects the null when $J_{N}>{\Greekmath 011F}
_{{\Greekmath 011C} }^{2}\left( K_{x}\right) ,$ the upper ${\Greekmath 011C} $-quantile of the ${\Greekmath 011F}
^{2}\left( K_{x}\right) $ distribution.

\section{Monte Carlo Simulations\label{Sim}}

In this section, we conduct some Monte Carlo simulations to evaluate the
proposed TW-MG\ estimator and jackknife method.

\subsection{DGPs\label{Sim_DGP} and implementations}

We consider four DGPs where $y_{it}$ and $x_{it}$ follow equations (\ref
{y_eqn1}) and (\ref{x_eqn2}), respectively. DGPs 1--3 are simple DGPs with
one regressor $(K_{x}=1)$, covering three different cases of ${\Greekmath 010C} _{i}$, \
to confirm our main theoretical results. In DGP 1, ${\Greekmath 010C} _{i}$ is
homogeneous; in DGP 2, ${\Greekmath 010C} _{i}$ is random and independent of $x_{it};$
and in DGP 3, ${\Greekmath 010C} _{i}$ enters the DGP\ of $x_{it}.$ With this design,
the pooled-type estimator is consistent for DGPs 1 and 2 and inconsistent
for DGP\ 3. In contrast, the MG-type estimator is consistent for all three
DGPs. The conventional MG estimator (\cite{pesaran_smith1995}), without
accounting for time fixed effects, is inconsistent for all three DGPs due to
the presence of time fixed effects. DGP 4 is more general and realistic with
two regressors ($K_{x}=2$), and auto-correlated and heteroskedastic error
terms as described below. In DGP 4, only the MG-type estimator is
consistent. In all four DGPs, we include both TWFEs and IFEs and also ensure
that the key Assumption \ref{Assmp3.1}(i) holds. In Section \ref{SecS2} of
the online supplement, we report the simulation results for two more DGPs
with TWFEs only.

Specifically, in DGPs 1-3, $K_{x}=R=1,$ ${\Greekmath 010C} ^{0}=1.$ Let $({\Greekmath 0115}
_{i}^{o},f_{t}^{o},{\Greekmath 010D} _{i}^{o},u_{it},{\Greekmath 0118} _{it},{\Greekmath 0111} _{i}^{\ast
},v_{it}^{\ast })$ be generated as $i.i.d.$ $N\left( (\mathbf{{\Greekmath 0113} }
_{3}^{\prime },\mathbf{0}_{4\times 1}^{\prime })^{\prime },\mathbb{I}
_{7}\right) $ random variables where $\mathbf{{\Greekmath 0113} }_{k}$ and $\mathbf{0}
_{k\times 1}$ denote a $k\times 1$ vector of ones and zeros, respectively.
We let ${\Greekmath 010B} _{i\cdot }^{o}={\Greekmath 0116} _{i\cdot }^{o}={\Greekmath 0115} _{i}^{o},$ and $
{\Greekmath 010B} _{\cdot t}^{o}={\Greekmath 0116} _{\cdot t}^{o}=f_{t}^{o}.$ Now we specify ${\Greekmath 0111}
_{i}$ and $v_{it}$ for DGPs 1-3. In DGP 1, ${\Greekmath 0111} _{i}=0,$ and in DGPs\ 2 and
3, ${\Greekmath 0111} _{i}={\Greekmath 0111} _{i}^{\ast }$. \ In DGPs 1 and 2, $v_{it}=v_{it}^{\ast },$
and in DGP$\ $3, $v_{it}={\Greekmath 010C} _{i}{\Greekmath 0118} _{it}+v_{it}^{\ast }.$

In DGP 4, $K_{x}=2$ and $R=1.$ Let $({\Greekmath 0115} _{i}^{o},f_{t}^{o},{\Greekmath 010D}
_{1,i}^{o},{\Greekmath 010D} _{2,i}^{o},u_{it}^{\ast \ast },{\Greekmath 0118} _{it},{\Greekmath 0111} _{1,i}^{\ast
},{\Greekmath 0111} _{2,i}^{\ast },v_{1,it}^{\ast },v_{2,it}^{\ast })$ be generated as $
i.i.d.$ $N\left( (\mathbf{{\Greekmath 0113} }_{4}^{\prime },\mathbf{0}_{6\times
1}^{\prime })^{\prime },\mathbb{I}_{10}\right) $ random variables. We let $
{\Greekmath 010B} _{i\cdot }^{o}={\Greekmath 0116} _{1,i\cdot }^{o}={\Greekmath 0116} _{2,i\cdot }^{o}={\Greekmath 0115}
_{i}^{o},$ and ${\Greekmath 010B} _{\cdot t}^{o}={\Greekmath 0116} _{1,\cdot t}^{o}={\Greekmath 0116} _{2,\cdot
t}^{o}=f_{t}^{o}.$ Let $x_{it}=\left( x_{1,it},x_{2,it}\right) $ which are
generated as $x_{1,it}={\Greekmath 0116} _{1,i\cdot }^{o}+{\Greekmath 0116} _{1,\cdot t}^{o}+{\Greekmath 010D}
_{1,i}^{o}f_{t}^{o}+v_{1,it}$ and $x_{2,it}={\Greekmath 0116} _{2,i\cdot }^{o}+{\Greekmath 0116}
_{2,\cdot t}^{o}+{\Greekmath 010D} _{2,i}^{o}f_{t}^{o}+v_{2,it}$ and let ${\Greekmath 010C}
_{i}=({\Greekmath 010C} _{1,i},{\Greekmath 010C} _{2,i})^{\prime }={\Greekmath 010C} ^{0}+({\Greekmath 0111} _{1,i},{\Greekmath 0111}
_{2,i})^{\prime },$ where ${\Greekmath 010C} ^{0}=(1,1)^{\prime }.$ For ${\Greekmath 0111} _{1,i},$ $
{\Greekmath 0111} _{2,i},$ $v_{1,it}$ and $v_{2,it},$ we let
\begin{equation*}
v_{1,it}={\Greekmath 010C} _{1,i}{\Greekmath 0118} _{it}+v_{1,it}^{\ast },\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }v_{2,it}=v_{2,it}^{
\ast },\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }{\Greekmath 0111} _{1,i}={\Greekmath 0111} _{1,i}^{\ast }\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }{\Greekmath 0111} _{2,i}={\Greekmath 010D}
_{2,i}^{o}-1+{\Greekmath 0111} _{2,i}^{\ast }.
\end{equation*}
$u_{it}$ is generated as
\begin{equation*}
u_{it}=\sqrt{1+0.25x_{1,it}^{2}}u_{it}^{\ast }\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }u_{it}^{\ast
}=0.25u_{i,t-1}^{\ast }+u_{it}^{\ast \ast }.
\end{equation*}
Note that ${\Greekmath 010C} _{1,i}$ enters the DGP of $x_{1,it}$ via $v_{1,it},$ and $
{\Greekmath 0111} _{2,i}$ (and ${\Greekmath 010C} _{2,i}$) is correlated with the factor loading $
{\Greekmath 010D} _{2,i}^{o}$ in the $x$-equation. Therefore, the pooled-type\
estimators of ${\Greekmath 010C} ^{0}$ are inconsistent.

For each DGP, we compare the six estimators: $(i)$ the TW-MG estimator in (
\ref{betahat1}), $(ii)$ the TW-MG ridge estimator in (\ref{ridge}), $(iii)$
the TW-pooled estimator in (\ref{B_tilde1}), $(iv)$ \cite{pesaran2006}'s
CCE-pooled estimator, $(v)$ \cite{pesaran2006}'s CCE-MG estimator, and $(vi)$
\cite{pesaran_smith1995}'s standard MG estimator. Table \ref{consistency}
below summarizes the consistency of the estimators for the four DGPs. So,
only the TW-MG, TW-MG ridge, and CCE-MG estimators are consistent for all
four DGPs. For the CCE-MG estimator, the consistency results are merely our
conjecture, and we are not aware of any rigorous studies on the CCE-MG
estimator for heterogeneous models with fixed $T.$\footnote{
We conjecture that theoretical analysis can be challenging. For example, we
will encounter the issue of the second moment matrix of the estimated
factors becoming asymptotically singular when the number of proxies exceeds
the number of factors, as pointed by \cite{karabiyik2017} for the large $T$
case (see, also \cite{andrews1987}). In addition, it is unclear how to
conduct valid inference for the CCE-MG estimators when $T$ is fixed.} Even
so, we observe below that our TW-MG and TW-MG ridge estimators substantially
outperform the CCE-MG estimator when $T$ is small. \cite{westerlund2022cce}
study the CCE-pooled estimator for heterogeneous models for fixed $T$, which
requires that the slope coefficients be uncorrelated with the regressors.
The choice of tuning parameter for the TW-MG ridge estimator is discussed in
Remark \ref{Rmk2.4}. For the CCE-pooled and CCE-MG estimators, we also
include the individual fixed effects.\footnote{
If we do not include the individual effects, the performance is worse in
general for these four DGPs.} We consider different combinations of $\left(
N,T\right) $ as shown in the tables below. For DGPs 1-3 where $K_{x}=1,$ $T$
starts with $\,T=3.$ For DGP 4 where $K_{x}=2,$ $T$ starts with $T=4.$ The
number of replications is 250.
\begin{table}[tph]
\caption{Consistency of the estimators}
\label{consistency}
\begin{center}
{\small $
\begin{tabular}{ccccccc}
\hline\hline
& TW-MG & TW-MG ridge & TW-pooled & CCE-MG & CCE-pooled & MG \\ \hline
DGP 1 & consistent & consistent & consistent & consistent & consistent &
inconsistent \\ \hline
DGP 2 & consistent & consistent & consistent & consistent & consistent &
inconsistent \\ \hline
DGP 3 & consistent & consistent & inconsistent & consistent & inconsistent &
inconsistent \\ \hline
DGP 4 & consistent & consistent & inconsistent & consistent & inconsistent &
inconsistent \\ \hline
\end{tabular}
$}
\end{center}
\par
{\small Notes: \textquotedblleft TW-MG\textquotedblright\ refers to the
two-way-mean-group estimator proposed in (\ref{betahat1}), \textquotedblleft
TW-MG ridge\textquotedblright\ refers to the ridge-type estimator proposed
in (\ref{ridge}), \textquotedblleft TW-pooled\textquotedblright\ refers to
the pooled two-way fixed effects estimator in (\ref{B_tilde1}),
\textquotedblleft CCE-pooled\textquotedblright\ refers to \cite{pesaran2006}
's CCE\ pooled estimator, and \textquotedblleft CCE-MG\textquotedblright\
refers to \cite{pesaran2006}'s CCE\ mean group estimator, and
\textquotedblleft MG\textquotedblright\ refers to \cite{pesaran_smith1995}'s
standard mean-group estimator. }
\end{table}

\subsection{Simulation results\label{Sim_results}}

Table \ref{estimation1} presents the performance of the six estimators
discussed above for the four DGPs. For each estimator, we report its bias$
\times 10$ and mean-squared-error\ (MSE)$\times 100$. To save space, the
variance is not reported but available upon request. For DGP 4 where there
are two regressors with ${\Greekmath 010C} ^{0}=({\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0})^{\prime
},$ we only report the results for ${\Greekmath 010C} _{1}^{0}$ to save space, as the
results for ${\Greekmath 010C} _{2}^{0}$ are similar. Among the four DGPs, the standard
MG estimator is inconsistent, so we mainly compare the other five
estimators. For DGP 1 where ${\Greekmath 010C} _{i}$ is homogeneous, all five estimators
are consistent. However, when $T$ is small ($T=3$), the TW-pooled estimator
performs best, as it is based on the most parsimonious model. The
second-best estimator is the TW-MG ridge estimator, followed by the TW-MG
estimator, CCE-pooled and CCE-MG estimators. The difference in the
performance is substantial. For example, when $N=50$ and $T=3,$ the MSEs$
\times 100$ of the TW-pooled, TW-MG ridge, TW-MG, CCE-pooled and CCE-MG
estimators are 1.31, 3.02, 7.35, 16.60 and 34.33, respectively. However when
$T$ is large ($T=10),$ the performances of the five estimators are all good
and similar. For DGP 2 where ${\Greekmath 010C} _{i}$ is heterogeneous but independent
of $x_{it},$ our TW-MG ridge estimator is the best when $T$ is small, by
achieving a smaller variance. When $T$ is large, the performances of the
TW-MG ridge, TW-MG and CCE-MG estimators are similar, and they outperform
the CCE-pooled and TW-pooled estimators. For DGPs 3 and 4, the pooled
estimators are inconsistent for the reason discussed above. It is clear that
its bias of the TW-pooled and CCE-pooled estimators does not converge to
zero when the sample size increases. In contrast, the bias of the TW-MG,
TW-MG ridge and CCE-MG estimators is all small. Again, when $T$ is small,
the TW-MG ridge estimator outperforms the TW-MG estimator, which in turn
outperforms the CCE-MG estimators. When $T$ is large, the three estimators
all perform well.

Table \ref{coverage1} shows the empirical coverage of the 95\% CI in (\ref
{CI1}) for ${\Greekmath 010C} ^{0}$ in DGPs 1-3 and ${\Greekmath 010C} _{1}^{0}$ in DGP\ 4 based on
the leave-one-out jackknife inference method proposed in Section \ref{Sec2.3}
. We consider both TW-MG and TW-MG ridge estimators. The actual coverages
are all close to the nominal 95\% for all sample sizes and DGPs. This
confirms the asymptotic validity of the leave-one-out jackknife method.
Table \ref{test_results} reports the rejection frequency of testing $\mathbb{
H}_{0}$ in (\ref{H0}) with a nominal size of $5\%$. We consider two types of
tests, one based on the TW-MG estimator and the other based on the TW-MG
ridge estimators. As discussed above, in DGPs 1 and 2, both the TW-MG (TW-MG
ridge) and pooled estimators are consistent, so these two DGPs are under the
null. Under the null, the rejection frequency of both tests is close to $5\%$
for all $N$ and $T$ considered. For DGPs 3 and 4 which are under the
alternative, the rejection frequency of both tests increases with $N$ and $
T, $ demonstrating the power of the tests. When $N$ is large, such as $N=$ $
200, $ the rejection frequency of both tests is close to $100\%$. We also
observe that when the sample size is small (e.g., $N=50$ and $T=3$), the
test based on the TW-MG ridge estimator is more powerful than that based on
the TW-MG estimator, as expected. \setlength{\tabcolsep}{2.0pt}

\begin{table}[tph]
\caption{Estimation Results}
\label{estimation1}
\begin{center}
{\small
\begin{tabular}{ccccccccccccccccc}
\hline\hline
&  &  & \multicolumn{6}{c}{bias$\times 10$} &  &  & \multicolumn{6}{c}{MSE$
\times 100$} \\ \cline{4-9}\cline{12-17}
&  &  & TW- & TW-MG & TW- & CCE- & CCE- &  & \ \ \ \  & \ \ \ \  & TW- &
TW-MG & TW- & CCE- & CCE- &  \\
$N$ & $T$ &  & MG & ridge & pooled & MG & pooled & MG &  &  & MG & ridge &
pooled & MG & pooled & MG \\ \cline{1-2}\cline{4-9}\cline{9-17}
\multicolumn{17}{c}{DGP 1} \\ \hline
50 & 3 &  & -0.12 & -0.39 & -0.07 & -0.04 & -0.24 & 7.42 &  &  & 7.35 & 3.02
& 1.31 & 34.33 & 16.60 & 72.78 \\
100 & 3 &  & -0.04 & -0.28 & 0.04 & 0.49 & -0.04 & 7.38 &  &  & 6.42 & 1.74
& 0.58 & 23.98 & 12.44 & 70.86 \\
200 & 3 &  & -0.25 & -0.28 & 0.01 & -0.27 & 0.03 & 7.07 &  &  & 2.08 & 0.99
& 0.31 & 13.10 & 11.35 & 63.13 \\
50 & 5 &  & -0.05 & -0.36 & 0.05 & 0.13 & 0.04 & 7.40 &  &  & 1.20 & 1.21 &
0.75 & 7.69 & 1.06 & 58.53 \\
100 & 5 &  & 0.01 & -0.16 & -0.04 & -0.13 & 0.00 & 7.38 &  &  & 0.53 & 0.53
& 0.36 & 5.07 & 0.52 & 58.05 \\
200 & 5 &  & -0.01 & -0.09 & -0.07 & 0.09 & 0.02 & 7.39 &  &  & 0.30 & 0.29
& 0.23 & 2.58 & 0.25 & 57.74 \\
50 & 10 &  & -0.06 & -0.32 & -0.03 & -0.02 & -0.02 & 7.41 &  &  & 0.62 & 0.69
& 0.74 & 0.43 & 0.36 & 56.80 \\
100 & 10 &  & 0.01 & -0.12 & 0.03 & -0.01 & -0.02 & 7.40 &  &  & 0.26 & 0.27
& 0.32 & 0.22 & 0.14 & 56.12 \\
200 & 10 &  & -0.01 & -0.07 & 0.02 & -0.01 & 0.01 & 7.10 &  &  & 0.11 & 0.11
& 0.15 & 0.11 & 0.07 & 51.65 \\ \hline
\multicolumn{17}{c}{DGP 2} \\ \hline
50 & 3 &  & -0.02 & -0.30 & 0.19 & -0.02 & -1.10 & 7.53 &  &  & 9.76 & 4.84
& 8.46 & 29.86 & 118.60 & 76.74 \\
100 & 3 &  & -0.01 & -0.25 & 0.07 & 0.30 & -0.62 & 7.41 &  &  & 7.78 & 2.79
& 4.29 & 18.66 & 95.24 & 73.19 \\
200 & 3 &  & -0.22 & -0.26 & -0.01 & 0.08 & -0.08 & 7.10 &  &  & 2.46 & 1.44
& 1.77 & 18.78 & 95.47 & 63.83 \\
50 & 5 &  & 0.06 & -0.26 & 0.27 & 0.13 & 0.11 & 7.50 &  &  & 3.38 & 3.20 &
7.19 & 9.18 & 4.77 & 61.78 \\
100 & 5 &  & 0.02 & -0.14 & -0.04 & -0.02 & -0.02 & 7.39 &  &  & 1.59 & 1.56
& 3.65 & 5.17 & 2.81 & 59.25 \\
200 & 5 &  & 0.03 & -0.06 & -0.04 & 0.17 & -0.03 & 7.42 &  &  & 0.76 & 0.75
& 1.29 & 6.17 & 1.30 & 58.81 \\
50 & 10 &  & 0.10 & -0.16 & 0.18 & 0.14 & 0.18 & 7.57 &  &  & 2.55 & 2.44 &
5.99 & 2.28 & 2.60 & 61.19 \\
100 & 10 &  & 0.02 & -0.12 & -0.02 & 0.01 & -0.02 & 7.41 &  &  & 1.35 & 1.33
& 3.30 & 1.35 & 1.64 & 57.38 \\
200 & 10 &  & 0.01 & -0.05 & 0.02 & 0.01 & -0.01 & 7.12 &  &  & 0.64 & 0.63
& 1.18 & 0.61 & 0.77 & 52.69 \\ \hline
\multicolumn{17}{c}{DGP 3} \\ \hline
50 & 3 &  & 0.11 & -0.24 & 5.05 & 0.44 & 11.91 & 5.55 &  &  & 5.43 & 3.43 &
37.61 & 35.99 & 264.24 & 43.69 \\
100 & 3 &  & 0.02 & -0.14 & 5.13 & 0.01 & 15.00 & 5.58 &  &  & 2.70 & 2.10 &
31.97 & 10.68 & 303.12 & 41.62 \\
200 & 3 &  & -0.06 & -0.17 & 5.31 & -0.05 & 16.98 & 5.62 &  &  & 1.36 & 1.02
& 31.38 & 7.85 & 352.15 & 41.30 \\
50 & 5 &  & 0.14 & -0.10 & 5.15 & 0.37 & 5.96 & 5.68 &  &  & 2.66 & 2.47 &
35.21 & 8.59 & 43.56 & 36.61 \\
100 & 5 &  & 0.04 & -0.08 & 4.97 & 0.03 & 6.25 & 5.59 &  &  & 1.36 & 1.33 &
30.12 & 3.98 & 44.11 & 35.43 \\
200 & 5 &  & 0.00 & -0.06 & 5.09 & 0.16 & 6.61 & 5.51 &  &  & 0.65 & 0.64 &
28.22 & 1.93 & 46.01 & 33.59 \\
50 & 10 &  & 0.11 & -0.08 & 5.00 & 0.30 & 6.35 & 5.71 &  &  & 2.20 & 2.10 &
31.39 & 2.07 & 45.15 & 35.66 \\
100 & 10 &  & -0.04 & -0.13 & 4.99 & 0.06 & 6.35 & 5.55 &  &  & 1.22 & 1.21
& 29.01 & 1.21 & 43.52 & 32.93 \\
200 & 10 &  & 0.02 & -0.03 & 5.15 & 0.04 & 6.56 & 5.29 &  &  & 0.55 & 0.54 &
28.00 & 0.54 & 44.42 & 29.81 \\ \hline
\multicolumn{17}{c}{DGP\ 4} \\ \hline
50 & 4 &  & -0.18 & -0.07 & 5.18 & -0.35 & 11.70 & 2.73 &  &  & 16.43 & 6.10
& 39.50 & 30.40 & 251.05 & 21.79 \\
100 & 4 &  & 0.09 & -0.10 & 5.00 & -0.01 & 12.20 & 2.80 &  &  & 7.86 & 3.71
& 32.63 & 14.16 & 286.55 & 16.15 \\
200 & 4 &  & 0.04 & 0.06 & 5.37 & -0.05 & 15.34 & 2.71 &  &  & 7.78 & 1.66 &
32.65 & 9.55 & 346.44 & 16.06 \\
50 & 5 &  & 0.11 & -0.08 & 5.29 & -0.65 & 6.27 & 2.86 &  &  & 5.43 & 3.99 &
39.86 & 42.57 & 55.73 & 13.41 \\
100 & 5 &  & 0.02 & -0.11 & 4.89 & -0.42 & 6.16 & 2.84 &  &  & 2.98 & 2.40 &
30.62 & 23.46 & 46.97 & 10.90 \\
200 & 5 &  & 0.04 & 0.00 & 5.16 & -0.33 & 6.58 & 2.77 &  &  & 1.56 & 1.29 &
29.69 & 25.54 & 48.73 & 9.39 \\
50 & 10 &  & 0.13 & -0.09 & 4.96 & 0.11 & 6.44 & 2.88 &  &  & 2.60 & 2.40 &
33.03 & 3.67 & 47.97 & 10.68 \\
100 & 10 &  & -0.04 & -0.14 & 4.95 & -0.06 & 6.31 & 2.70 &  &  & 1.66 & 1.61
& 29.69 & 2.17 & 43.53 & 8.79 \\
200 & 10 &  & 0.00 & -0.05 & 5.19 & 0.01 & 6.52 & 2.65 &  &  & 0.75 & 0.74 &
28.91 & 1.04 & 44.26 & 7.80 \\ \hline
\end{tabular}
}
\end{center}
\par
{\small Notes: see the notes to Table \ref{consistency}. }
\end{table}

\begin{table}[tph]
\caption{Inference Results}
\label{coverage1}
\begin{center}
{\small \ \tabcolsep 2 mm
\begin{tabular}{cccccccccccc}
\hline\hline
\multicolumn{12}{c}{Empirical coverage of 95\% CI for ${\Greekmath 010C} ^{0}$} \\ \hline
&  &  & \multicolumn{4}{c}{based on TW-MG estimator} & \ \ \ \ \  &
\multicolumn{4}{c}{based on TW-MG ridge estimator} \\
\cline{4-7}\cline{9-10}\cline{11-12}
$N$ & $T$ &  & DGP\ 1 & DGP\ 2 & DGP\ 3 & DGP\ 4 &  & DGP\ 1 & DGP\ 2 & DGP\
3 & DGP\ 4 \\ \cline{1-2}\cline{4-7}\cline{9-10}\cline{11-12}
50 & 3 &  & 0.96 & 0.96 & 0.95 & 0.96 &  & 0.96 & 0.94 & 0.95 & 0.96 \\
100 & 3 &  & 0.98 & 0.95 & 0.94 & 0.95 &  & 0.94 & 0.94 & 0.95 & 0.95 \\
200 & 3 &  & 0.97 & 0.96 & 0.96 & 0.98 &  & 0.94 & 0.95 & 0.96 & 0.97 \\
50 & 5 &  & 0.95 & 0.96 & 0.96 & 0.95 &  & 0.93 & 0.95 & 0.96 & 0.96 \\
100 & 5 &  & 0.97 & 0.94 & 0.94 & 0.96 &  & 0.95 & 0.96 & 0.93 & 0.95 \\
200 & 5 &  & 0.94 & 0.96 & 0.96 & 0.95 &  & 0.93 & 0.95 & 0.95 & 0.94 \\
50 & 10 &  & 0.94 & 0.95 & 0.95 & 0.96 &  & 0.92 & 0.94 & 0.96 & 0.95 \\
100 & 10 &  & 0.95 & 0.95 & 0.96 & 0.94 &  & 0.93 & 0.95 & 0.95 & 0.92 \\
200 & 10 &  & 0.97 & 0.96 & 0.96 & 0.94 &  & 0.96 & 0.96 & 0.97 & 0.95 \\
\hline
\end{tabular}
}
\end{center}
\par
{\small Notes: 95\% CI\ refers to 95\% confidence interval based on the
jackknife method. \textquotedblleft TW-MG\textquotedblright\ refers to the
two-way-mean-group estimator proposed in (\ref{betahat1}) and
\textquotedblleft TW-MG ridge\textquotedblright\ refers to the ridge-type
estimator proposed in (\ref{ridge}). For DGP\ 4 where $K_{x}=2$, the
shortest $T$ is 4 instead of $3$. }
\end{table}

\begin{table}[tph]
\caption{Testing Results}
\label{test_results}
\begin{center}
{\small \ \tabcolsep 1.8 mm
\begin{tabular}{cccccccccccccccc}
\hline\hline
\multicolumn{16}{c}{Rejection frequency of testing $\mathbb{H}_{0}$} \\
\hline
&  &  & \multicolumn{6}{c}{based on TW-MG estimator} & \ \ \ \ \  &
\multicolumn{6}{c}{based on TW-MG ridge estimator} \\
\cline{4-9}\cline{11-14}\cline{13-16}
&  &  & \multicolumn{2}{c}{size} &  &  & \multicolumn{2}{c}{power} &  &
\multicolumn{2}{c}{size} &  &  & \multicolumn{2}{c}{power} \\
\cline{4-5}\cline{8-9}\cline{11-12}\cline{15-16}
$N$ & $T$ &  & DGP\ 1 & DGP\ 2 &  &  & DGP\ 3 & DGP\ 4 &  & DGP\ 1 & DGP\ 2
&  &  & DGP\ 3 & DGP\ 4 \\
\cline{1-2}\cline{4-5}\cline{8-9}\cline{11-12}\cline{15-16}
50 & 3 &  & 0.03 & 0.04 &  &  & 0.44 & 0.51 &  & 0.02 & 0.05 &  &  & 0.56 &
0.68 \\
100 & 3 &  & 0.02 & 0.04 &  &  & 0.68 & 0.77 &  & 0.04 & 0.06 &  &  & 0.74 &
0.86 \\
200 & 3 &  & 0.03 & 0.03 &  &  & 0.87 & 0.99 &  & 0.04 & 0.02 &  &  & 0.91 &
1.00 \\
50 & 5 &  & 0.06 & 0.08 &  &  & 0.67 & 0.74 &  & 0.05 & 0.08 &  &  & 0.73 &
0.80 \\
100 & 5 &  & 0.02 & 0.04 &  &  & 0.84 & 0.95 &  & 0.04 & 0.06 &  &  & 0.86 &
0.97 \\
200 & 5 &  & 0.02 & 0.05 &  &  & 0.97 & 1.00 &  & 0.03 & 0.04 &  &  & 0.98 &
1.00 \\
50 & 10 &  & 0.06 & 0.06 &  &  & 0.83 & 0.91 &  & 0.09 & 0.07 &  &  & 0.84 &
0.94 \\
100 & 10 &  & 0.05 & 0.04 &  &  & 0.96 & 1.00 &  & 0.06 & 0.04 &  &  & 0.97
& 1.00 \\
200 & 10 &  & 0.04 & 0.05 &  &  & 1.00 & 1.00 &  & 0.06 & 0.06 &  &  & 1.00
& 1.00 \\ \hline
\end{tabular}
}
\end{center}
\par
{\small Notes: The test of $\mathbb{H}_{0}$ (\ref{H0}) is based on the test
statistic (\ref{test}). The nominal size is $5\%$. \textquotedblleft
TW-MG\textquotedblright\ refers to the two-way-mean-group estimator proposed
in (\ref{betahat1}) and \textquotedblleft TW-MG ridge\textquotedblright\
refers to the ridge-type estimator proposed in (\ref{ridge}). For DGP\ 4
where $K_{x}=2$, the shortest $T$ is 4 instead of $3$. }
\end{table}

\section{Empirical Applications\label{App}}

In this section, we apply our new TW-MG estimator and jackknife method to
two datasets. The first application examines the relationship between
health-care expenditure and income, while the second estimates a production
function. In the first application, we reject the null of poolability,
whereas in the second, we fail to reject it.

\subsection{Application 1: Health-care expenditure and income\label{App1}}

In this application, we use the data from \cite{baltagi2017health} and \cite
{kapetanios2023testing} to investigate the relationship between health-care
expenditure and income for a panel of 167 countries covering the period
1995--2012 $(N=167$ and $T=18).$ Here $y_{it}$ is the per-capita health-care
spending in the $i$th country at year $t.$ There are two regressors $\left(
K_{x}=2\right) $. The key regressor of interest is the income level measured
by per capita GDP. The second regressor is the rate of public expenditure
over total health expenditure as a control variable. For a more detailed
description of these two variables, see \cite{baltagi2017health}.

Let ${\Greekmath 010C} ^{0}=({\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0})^{\prime },$ where ${\Greekmath 010C}
_{1}^{0}$ and ${\Greekmath 010C} _{2}^{0}$ represent the expected values of slope
coefficients of income and public expenditure rate, respectively. Table \ref
{health_table} reports the estimation results with four estimation methods,
namely the TW-MG, TW-MG ridge, TW-pooled, and MG estimators. We find that
TW-MG and TW-MG ridge estimates are almost identical, which are
substantially different from the TW-pooled and MG estimates. Therefore, our
discussion below focuses on the TW-MG, TW-pooled and MG estimates. For $
{\Greekmath 010C} _{1}^{0},$ the three estimates are 0.94, 0.72, and 1.87, respectively.
For ${\Greekmath 010C} _{2}^{0},$ the three estimates are 0.30, 0.02 and 0.76,
respectively. If we treat our TW-MG estimator as a benchmark, the TW-pooled
estimator tends to under-estimate, while the standard MG estimator
over-estimates. For ${\Greekmath 010C} _{1}^{0},$ all three estimation methods indicate
a positive and statistically significant average effect of income, though
the magnitudes differ. For ${\Greekmath 010C} _{2}^{0},$ while all estimation methods
show a positive effect, only the TW-MG and MG estimates are statistically
significant at the 5\% level; the TW-pooled estimate is not. Figure \ref
{health_hist} plots the histogram of the estimates of ${\Greekmath 010C} _{i}=({\Greekmath 010C}
_{1i},{\Greekmath 010C} _{2i})^{\prime }$ based on (\ref{betai_est}). Although these
estimators are not consistent in the fixed $T$ framework, we observe
substantial variations, suggesting the heterogeneity. Notably, most
estimates of ${\Greekmath 010C} _{1i}$ are positive, demonstrating a robust positive
effect of income.

We conduct the test for the hypotheses of poolability introduced in Section
\ref{Sec4.3}. In addition to testing poolability in estimating both
parameters in (${\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0}$) together, we also test the
poolability for each individual parameter. Specifically, let $\hat{{\Greekmath 010C}}
^{P}=(\hat{{\Greekmath 010C}}_{1}^{P},\hat{{\Greekmath 010C}}_{2}^{P})^{\prime }$ and $\hat{{\Greekmath 010C}}
^{MG}=(\hat{{\Greekmath 010C}}_{1}^{MG},\hat{{\Greekmath 010C}}_{2}^{MG})^{\prime },$ and test the
following three hypotheses:
\begin{equation}
\mathbb{H}_{0\ell }:\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}_{N\rightarrow \infty }\hat{{\Greekmath 010C}}_{\ell
}^{MG}=\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}_{N\rightarrow \infty }\hat{{\Greekmath 010C}}_{\ell }^{P}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for }
\ell \in \left[ 2\right] ,\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }\mathbb{H}_{0}:\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}
_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{MG}=\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{plim}_{N\rightarrow \infty }
\hat{{\Greekmath 010C}}^{P}.  \label{H_app}
\end{equation}
As shown in Table \ref{test_health}, we reject $\mathbb{H}_{01},$ $\mathbb{H}
_{02}$ and $\mathbb{H}_{0}$ all with $p$-values of 0.027, 0.019 and 0.009,
respectively, based on the TW-MG estimator and the jackknife method. We also
report the step-down Holm adjusted $p$-values (\cite{holm1979simple})
accounting for multiple testing, which are all smaller than 0.05.\footnote{
Suppose that we have conducted the tests for $K^{\ }$individual hypotheses.
We order the individual \textit{p}-values from the smallest to the largest
as $p_{\left( 1\right) }\leq p_{\left( 2\right) }\leq ...\leq p_{\left(
K\right) }$ with their corresponding null hypotheses labeled accordingly as $
\mathbb{H}_{0\left( 1\right) },$ $\mathbb{H}_{0\left( 2\right) },...,$ $
\mathbb{H}_{0\left( K\right) }.$ The step-down Holm adjusted $p$-values for
testing $\mathbb{H}_{0\left( k\right) }$ is calculated as: adjusted-$
p_{\left( k\right) }=\min \left( \left( K-k+1\right) p_{\left( k\right)
},1\right) .$
\par
{}} Therefore, we conclude that we cannot pool the data to estimate the
average effects, confirming the substantial differences in the estimation
results discussed above.
\begin{table}[tph]
\caption{Estimation results for the health expenditure application $(N=167,$
$T=18)$}
\label{health_table}
\begin{center}
{\small
\begin{tabular}{ccccccccccccc}
\hline\hline
&  & \multicolumn{2}{c}{TW-MG} &  & \multicolumn{2}{c}{TW-MG ridge} &  &
\multicolumn{2}{c}{TW-pooled} &  & \multicolumn{2}{c}{MG} \\
\cline{3-7}\cline{9-13}
&  & estimates & 95\% CI &  & estimates & 95\% CI &  & estimates & 95\% CI &
& estimates & 95\% CI \\ \hline
\multicolumn{1}{l}{${\Greekmath 010C} _{1}^{0}$} &  & 0.94 & [0.68 1.20] &  & 0.94 &
[0.68 1.20] &  & 0.72 & [0.55 0.98] &  & 1.87 & [1.64 2.10] \\
\multicolumn{1}{l}{${\Greekmath 010C} _{2}^{0}$} &  & 0.30 & [0.06 0.54] &  & 0.30 &
[0.06 0.54] &  & 0.02 & [-0.03 0.26] &  & 0.76 & [0.33 1.19] \\ \hline
\end{tabular}
}
\end{center}
\par
{\small Notes: ${\Greekmath 010C} _{1}^{0}$ and ${\Greekmath 010C} _{2}^{0}$ are the expected values
of slope coefficients of income and public expenditure rate, respectively.
\textquotedblleft TW-MG\textquotedblright\ refers to the two-way-mean-group
estimator proposed in (\ref{betahat1}), \textquotedblleft TW-MG
ridge\textquotedblright\ refers to the ridge-type estimator proposed in (\ref
{ridge}), \textquotedblleft TW-pooled\textquotedblright\ refers to the
pooled two-way fixed effects estimator in (\ref{B_tilde1}), and
\textquotedblleft MG\textquotedblright\ refers to \cite{pesaran_smith1995}'s
standard mean-group estimator. 95\% CI\ refers to 95\% confidence interval
based on the proposed jackknife method. }
\end{table}
\begin{table}[tph]
\caption{Testing results for the health expenditure application $(N=167,$ $
T=18)$}
\label{test_health}
\begin{center}
{\small \ \tabcolsep 3 mm
\begin{tabular}{ccccccc}
\hline\hline
& Null hypothesis & $\ \ \ \ \mathbb{H}_{01}$ $\ \ \ \ $ &  & $\ \ \ \
\mathbb{H}_{02}$ $\ \ \ \ $ &  & $\ \ \ \ \mathbb{H}_{0}$ $\ \ \ \ $ \\
\hline
& test statistics & 4.90 &  & 5.53 &  & 9.38 \\
\multicolumn{1}{l}{} & $p$-value & 0.027 &  & 0.019 &  & 0.009 \\
\multicolumn{1}{l}{} & Holm adjusted\ $p$-values & 0.027 &  & 0.038 &  &
0.027 \\ \hline
\end{tabular}
}
\end{center}
\par
{\small Notes: The hypotheses $\mathbb{H}_{01},$ $\mathbb{H}_{02}$ and $
\mathbb{H}_{0}$ are defined in (\ref{H_app}). The test is based on the
two-way-mean-group (TW-MG) estimators and the jackknife method proposed in
Section \ref{Sec4.3}. }
\end{table}
{\normalsize
\begin{figure}[tbp]
{\normalsize \centering
\includegraphics[keepaspectratio=true,scale=0.50,trim={3cm 0cm
				0	0cm},clip]{health_his.eps}  }
\caption{Histogram of estimates of individual slope coefficients in the
health expenditure application}
\label{health_hist}
\end{figure}
}

\subsection{Application 2: Production function\label{App2}}

In this subsection, we apply our new methods to estimate a production
function using a balanced panel from \cite{blundell2000gmm}. The dataset
consists of 508 R\&D-performing U.S. manufacturing companies for 8 years,
1982--1988, so $N=508$ and $T=8.$\footnote{
The original sample consists of 509 firms. However, we drop one firm because
its labor input (the first regressor) remains unchanged over the sample
period.} Here, the dependent variable $y_{it}$ is the logarithms of the
sales of firm $i$ in year $t$ and there are two regressors $(K_{x}=2)$, the\
logarithms of its labor (employment) and the logarithm of its capital stock,
which are inputs to the production function. For the detailed definitions of
the variables, see \cite{blundell2000gmm} and \cite{mairessehall}.

Table \ref{estimation_prod} reports the estimation results based on the four
estimation methods. Let ${\Greekmath 010C} ^{0}=({\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0})^{\prime
},$ where ${\Greekmath 010C} _{1}^{0}$ and ${\Greekmath 010C} _{2}^{0}$ represent the average return
to labor and capital, respectively. For ${\Greekmath 010C} _{1}^{0},$ the TW-MG, TW-MG
ridge, TW-pooled and MG estimates are similar, with the values of 0.67,
0.67, 0.65 and 0.63, respectively, all of which are statistically
significant. For ${\Greekmath 010C} _{2}^{0},$ the TW-MG, TW-MG ridge, and TW-pooled
estimates are similar with the values of 0.21, 0.21, and 0.23, respectively,
which are smaller than the MG estimate, 0.32. Again, all four estimates are
statistically significant. Figure \ref{prod} plots the histogram of the
estimates of ${\Greekmath 010C} _{i}=({\Greekmath 010C} _{1i},{\Greekmath 010C} _{2i})^{\prime }$ based on (\ref
{betai_est}), where we also observe some variations, though most of them are
positive

Similar to the first application in Section \ref{App1}, we conduct the test
for poolability. We test the three hypotheses as in (\ref{H_app}). As shown
in Table \ref{testing_prod}, we fail to reject $\mathbb{H}_{01},$ $\mathbb{H}
_{02},$ and $\mathbb{H}_{0}$ all with $p$-values of 0.56, 0.67, and 0.84,
respectively. Therefore, for this application, we do not have strong
evidence against poolability.
\begin{table}[tph]
\caption{Estimation results for the production function application $(N=508,$
$T=8)$}
\label{estimation_prod}
\begin{center}
{\small
\begin{tabular}{ccccccccccccc}
\hline\hline
&  & \multicolumn{2}{c}{TW-MG} &  & \multicolumn{2}{c}{TW-MG ridge} &  &
\multicolumn{2}{c}{TW-pooled} &  & \multicolumn{2}{c}{MG} \\
\cline{3-7}\cline{9-13}
&  & estimates & 95\% CI &  & estimates & 95\% CI &  & estimates & 95\% CI &
& estimates & 95\% CI \\ \hline
\multicolumn{1}{l}{${\Greekmath 010C} _{1}^{0}$} &  & 0.67 & [0.61 0.74] &  & 0.67 &
[0.61 0.74] &  & 0.65 & [0.59 0.72] &  & 0.63 & [0.57 0.69] \\
\multicolumn{1}{l}{${\Greekmath 010C} _{2}^{0}$} &  & 0.21 & [0.12 0.31] &  & 0.21 &
[0.12 0.31] &  & 0.23 & [0.17 0.33] &  & 0.32 & [0.26 0.39] \\ \hline
\end{tabular}
}
\end{center}
\par
{\small Notes: ${\Greekmath 010C} _{1}^{0}$ and ${\Greekmath 010C} _{2}^{0}$ are the expected values
of slope coefficients of labor and capital, respectively. \textquotedblleft
TW-MG\textquotedblright\ refers to the two-way-mean-group estimator proposed
in (\ref{betahat1}), \textquotedblleft TW-MG ridge\textquotedblright\ refers
to the ridge-type estimator proposed in (\ref{ridge}), \textquotedblleft
TW-pooled\textquotedblright\ refers to the pooled two-way fixed effects
estimator in (\ref{B_tilde1}), and \textquotedblleft MG\textquotedblright\
refers to \cite{pesaran_smith1995}'s standard mean-group estimator. 95\% CI\
refers to 95\% confidence interval based on the proposed jackknife method. }
\end{table}
\begin{table}[tph]
\caption{Testing results for the production function application $(N=508,$ $
T=8)$}
\label{testing_prod}
\begin{center}
{\small \ \tabcolsep 3 mm
\begin{tabular}{ccccccc}
\hline\hline
& Null hypothesis & $\ \ \ \ \mathbb{H}_{01}$ $\ \ \ \ $ &  & $\ \ \ \
\mathbb{H}_{02}$ $\ \ \ \ $ &  & $\ \ \ \ \mathbb{H}_{0}$ $\ \ \ \ $ \\
\hline
& test statistics & 0.26 &  & 0.18 &  & 0.36 \\
\multicolumn{1}{l}{} & $p$-value & 0.56 &  & 0.67 &  & 0.84 \\
\multicolumn{1}{l}{} & Holm adjusted\ $p$-values & 1 &  & 1 &  & 0.84 \\
\hline
\end{tabular}
}
\end{center}
\par
{\small Notes: The hypotheses $\mathbb{H}_{01},$ $\mathbb{H}_{02}$ and $
\mathbb{H}_{0}$ are defined in (\ref{H_app}). The test is based on the
two-way-mean-group (TW-MG) estimators and the jackknife method proposed in
Section \ref{Sec4.3}. }
\end{table}
{\normalsize
\begin{figure}[tbp]
{\normalsize \centering
\includegraphics[keepaspectratio=true,scale=0.50,trim={3cm 0.5cm
				0	0.05cm},clip]{prod.eps}  }
\caption{Histogram of estimates of individual slope coefficients in the
production function application}
\label{prod}
\end{figure}
}

\section{Conclusion\label{Sec7}}

This paper considers estimation and inference for the population average of
slope coefficients in heterogeneous panels with two-way additive fixed
effects (TWFEs) and interactive fixed effects (IFEs). We propose two methods
to estimate the parameter of interest. The first method extends the mean
group (MG) approach to accommodate both individual and time effects, which
is valid for general correlated random coefficient (CRC) models but imposes
stronger rank conditions. The second method pools all observations to
estimate the parameter of interest, which may not be valid for some CRC
models but imposes weaker rank conditions. For both estimation methods, we
establish stable convergence for the estimators and propose a jackknife
method to estimate their asymptotic variance-covariance matrices. In
addition, we propose a Hausman-type test to assess poolability, which may
justify the commonly used pooled estimator in empirical research in the
presence of slope heterogeneity. Simulations demonstrate that our methods
perform well in finite samples.

There are many interesting topics for further research. First, we may extend
our methods to dynamic models and allow for weakly exogenous regressors.
Second, we may address the endogeneity issue in heterogeneous panels by
extending our methods. Third, we may consider nonlinear heterogeneous panels
with short $T$ and TWFEs and/or IFEs. These extensions are highly
challenging, but worth exploring.

{\footnotesize \setstretch{0.75} \setlength{\bibsep}{5pt}
\bibliographystyle{apalike}
\bibliography{Hetero_panel2025}
}

\newpage