EconBase
← Back to paper

Two-Way Mean Group Estimators for Heterogeneous Panel Models with Fixed T

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,959 characters · 20 sections · 70 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Two-Way Mean Group Estimators for Heterogeneous Panel Models with Fixed $T$

abstractWe consider a correlated random coefficient panel data model with two-way fixed effects and interactive fixed effects in a fixed T framework. We propose a two-way mean group (TW-MG) estimator for the expected value of the slope coefficient and propose a leave-one-out jackknife method for valid inference. We also consider a pooled estimator and provide a Hausman-type test for poolability. Simulations demonstrate the excellent performance of our estimators and inference methods in finite samples. We apply our new methods to two datasets to examine the relationship between health-care expenditure and income, and estimate a production function. Key words: Heterogeneity; Least squares dummy variable estimator; Jackknife; Mean-group estimator; Random coefficient JEL Classification:\ C23, C33, C51, C55.

Introduction

Heterogeneity is a fundamental characteristic of many economic data (see, e.g., heckman2001micro). The assumption\ of homogeneous slope coefficients in traditional panel data models has been argued to be unrealistic. Empirically, the homogeneity of slope coefficients is frequently rejected; see, e.g., durlauf2001, Browning2007, phillips2007, Su_Chen2013, and lu2023. In this paper, we consider a heterogeneous panel data model with two-way fixed effects (TWFEs) and interactive fixed effects (IFEs) in a fixed $T$ framework. This type of general model usually requires a large $T,$ which is not applicable to many panel data sets with only a modest number of time periods. With a fixed $T,$ although we cannot consistently estimate individual slope coefficients, we can still obtain a consistent estimator of the average effects, which is usually of central interest in empirical work. Therefore, we propose a mean-group (MG) type estimator for the expected value of the slope coefficient based on the least squares dummy variable (LSDV) approach and a leave-one-out jackknife inference method.\footnote{ Note that there is no simple within transformation we can apply to remove the TWFEs when the slope coefficients vary with individuals.} One main advantage of the new estimation and inference methods is their ease of implementation. Despite their simplicity, they remain effective for a large class of data-generating processes (DGPs), accommodating an arbitrarily correlated slope coefficient with the regressors and allowing for the presence of IFEs. This combination of simplicity and broad applicability can make our new methods appealing to empirical researchers.

Specifically, we consider the following\ heterogeneous panel data model with standard TWFEs and IFEs\footnote{ Strictly speaking, the TWFEs can be absorbed into IFEs. Here we explicitly allow the presence of both, as they play different roles in our estimation procedure.}:

equation[equation omitted — 354 chars of source]

where $x_{it}$ is a $K_{x}\times 1$ vector of observable regressors, $ f_{t}^{o}$ and ${\Greekmath 0115} _{i}^{o}$ are $R\times 1$ vectors of unobservable factors and loadings, respectively, ${\Greekmath 010B} _{i\cdot }^{o}$ and ${\Greekmath 010B} _{\cdot t}^{o}$ are the individual and time fixed effects, respectively, and $u_{it}$ is the error term. Here ${\Greekmath 0115} _{i}^{\circ \prime }f_{t}^{\circ }$ in ((ref)) is often referred to as IFEs in the model, which can be correlated with $x_{it}.$ For the slope coefficient ${\Greekmath 010C} _{i},$ we assume that

equation[equation omitted — 101 chars of source]

with $E\left( {\Greekmath 0111} _{i}\right) =0$ so that ${\Greekmath 010C} ^{0}$ represents the population average marginal effect of $x_{it}$ on $y_{it}.$ The key parameter of interest is ${\Greekmath 010C} ^{0}.$ Even though we are not interested in the fixed effect parameters, viz. ${\Greekmath 010B} _{i\cdot }^{o},$ ${\Greekmath 010B} _{\cdot t}^{o},$ ${\Greekmath 0115} _{i}^{o},$ and $f_{t}^{o},$ we find that it is convenient to reformulate the model in ((ref)) to ensure the unobserved factors and loadings to have zero mean. For this purpose, we define

equation*[equation* omitted — 198 chars of source]

and rewrite the models in ((ref)) equivalently as

equation[equation omitted — 151 chars of source]

where ${\Greekmath 010B} _{i\cdot }={\Greekmath 010B} _{i\cdot }^{o}+{\Greekmath 0115} _{i}^{o\prime }E(f_{t}^{o}),$ ${\Greekmath 010B} _{\cdot t}={\Greekmath 010B} _{\cdot t}^{o}+E\left( {\Greekmath 0115} _{i}^{o\prime }\right) f_{t},$ and $u_{it}^{\dag }={\Greekmath 0115} _{i}^{\prime }f_{t}+u_{it}.$ We obtain the LSDV estimators of $\left\{ {\Greekmath 010B} _{i\cdot }\right\} _{i=1}^{N},\left\{ {\Greekmath 010B} _{\cdot t}\right\} _{t=1}^{T}$ and $ \left\{ {\Greekmath 010C} _{i}\right\} _{i=1}^{N}$ based on ((ref))$.$ Let $\hat{ {\Greekmath 010C}}_{i}$ denote the estimator of ${\Greekmath 010C} _{i},$ and our estimator of $ {\Greekmath 010C} ^{0}$ is simply

equation*[equation* omitted — 99 chars of source]

For simplicity, we name our estimator $\hat{{\Greekmath 010C}}^{MG}$ two-way-mean-group (TW-MG) estimator.

There are three main features in our model. First, we consider a fixed $T$ framework and typically only require that $T>K_{x}+1$. Therefore, many estimators, such as bai2009's principal component analysis (PCA) estimator and pesaran2006's common correlated effects (CCE) estimator, which typically require a large $T,$ are not applicable here. \

Second, we consider a heterogeneous panel data model where ${\Greekmath 010C} _{i}$ varies across individuals and can be correlated with $x_{it},$ so we have a correlated random coefficient (CRC) panel in ((ref)). For our model, two popular estimators, namely, the pooled TWFE\ estimator and pesaran_smith1995's standard MG estimator, are not applicable for the reasons discussed below. The pooled TWFE estimator of ${\Greekmath 010C} _{0}$ is based on the following regression

equation[equation omitted — 309 chars of source]

where $\mathring{u}_{it}={\Greekmath 0111} _{i}x_{it}+u_{it}^{\dag }$ is treated as a composite error term. The reason why the pooled estimator does not work here is that $x_{it}$ can be correlated with ${\Greekmath 0111} _{i}x_{it}$ (and thus $ \mathring{u}_{it}$), causing the typical endogeneity issue. For concrete examples, see DGPs 3 and 4 in our simulations in Section (ref) where the pooled TWFE\ estimator is inconsistent. See Remark (ref) for further discussion on the correlated random coefficient. pesaran_smith1995's MG estimator is based on the individual regression for each $i$ as follows,

equation[equation omitted — 196 chars of source]

where $\ddot{u}_{it}={\Greekmath 010B} _{\cdot t}+u_{it}^{\dag }$ is treated as a composite error term. The issue here is that $x_{it}$ can be correlated with ${\Greekmath 010B} _{\cdot t}$ (and thus $\ddot{u}_{it}$)$,$ leading to an endogeneity problem. For example, for all four DGPs in our simulations in Section (ref), the standard MG estimator is inconsistent. For further discussion on the relationship between our TW-MG and the standard MG estimators, see Remark (ref). Our TW-MG estimator can be thought of as a generalized version of the pooled TWFE\ and standard MG estimators, as estimation equations ((ref)) and ((ref)) are restricted versions of ((ref)) with the restrictions of ${\Greekmath 010C} _{i}={\Greekmath 010C} ^{0}$ and ${\Greekmath 010B} _{\cdot t}=0,$ respectively.

Third, we allow IFEs in our model. We argue that once we include TWFEs, the demeaned IFEs can enter the error term under our key assumption that $ {\Greekmath 0115} _{i}$ and $x_{it}^{o}$ are uncorrelated, where $x_{it}^{o}$ denotes the part of\ $x_{it}$ with its additive individual and time effects (if any) removed. For a more precise statement of the key assumption, see Assumption (ref)(i) below. This point has already been made in the literature (see, e.g., westerlund2019estimation and kapetanios2023testing for the large $T$ case). Here we provide a rigorous asymptotic analysis for a heterogenous model with a fixed $T.$ To illustrate this point, consider the following simple example where

equation*[equation* omitted — 219 chars of source]

with $K_{x}=1$ and $R=1.$ Suppose that $E\left( {\Greekmath 0115} _{i}^{o}\right) =$ $ E\left( {\Greekmath 010D} _{i}^{o}\right) =1$ and ${\Greekmath 0115} _{i}^{o},$ ${\Greekmath 010D} _{i}^{o},$ $f_{t}^{o},$ $u_{it}$ and $v_{it}$ are mutually independent with $E\left( u_{it}\right) =E\left( v_{it}\right) =0$. It appears that the IFEs, ${\Greekmath 0115} _{i}^{o}f_{t}^{o},$ cause much trouble, as

equation*[equation* omitted — 238 chars of source]

which can lead to an endogeneity issue. Furthermore, ${\Greekmath 0115} _{i}^{o}f_{t}^{o}$ can cause strong cross-sectional dependence in the $y$ -equation. However, we show that our TW-MG estimator still works in this case after controlling TWFEs, as the demeaned version of ${\Greekmath 0115} _{i}^{o}f_{t}^{o}$ does not cause the endogeneity issue and strong cross-sectional dependence. Admittedly, we do rule out the case where $ {\Greekmath 0115} _{i}^{o}$ and $x_{it}^{o}$ are correlated. However, as we already allow individual (${\Greekmath 010B} _{i\cdot })$ and time effects $({\Greekmath 010B} _{\cdot t})$ without any restrictions, the uncorrelated assumption between ${\Greekmath 0115} _{i}^{o}$ and $x_{it}^{o}$ may not be that strong. For the simple example above, essentially, we require that the factor loadings ${\Greekmath 0115} _{i}^{o}$ in the $y$-equation and ${\Greekmath 010D} _{i}^{o}$ in the $x$-equation be uncorrelated.\footnote{ In fact, the classic paper of pesaran2006 for large $T$ maintains the assumption that the factor loadings in the $y$-equation and the $x$-equation are uncorrelated.} Recently, kapetanios2023testing assume that both $ x_{it}$ and $y_{it}$ follow a factor structure as in our simple example (or more generally, the DGPs in Remark (ref)) and test the null hypothesis that the loadings of $x$- and $y$-equations are uncorrelated. They find that the null hypothesis is not rejected in thirteen out of fourteen datasets considered.\footnote{ Their test is based on the large $T$ framework and compares the TWFE estimator and the PCA\ estimator.} Furthermore, to the best of our knowledge, there is no well-accepted solution for estimating ${\Greekmath 010C} _{0}$ in model ((ref)) with a fixed $T$ when ${\Greekmath 0115} _{i}^{o}$ and $ x_{it}^{o} $ are correlated.

This paper makes several important methodological contributions. First, we study the asymptotic properties of our TW-MG estimator for the general model ((ref)) and demonstrate its consistency and asymptotic mixture normal distribution. In particular, we invoke the notion of stable convergence and stable central limit theorem (CLT) (see, e.g., hausler2015stable and van2023weak) to take into account the randomness of the factors. Second, we propose an inference method based on the leave-one-out jackknife approach and provide rigorous justification. Although this method is easy to implement, its theoretical justification is challenging because the parameters involved (${\Greekmath 010C} _{i}$'s) are not fixed when one cross-section unit is deleted. We have to augment certain matrices and vectors in the leave-one-out estimator to match those in the original estimator and show the consistency of the variance estimator. See Remark (ref) below for more details. Third, we examine the asymptotic properties of the pooled TWFE estimator of ${\Greekmath 010C} _{0}$ for our model ((ref)) under the additional assumption that ${\Greekmath 010C} _{i}$ is random and uncorrelated with a demeaned version of the regressor (Assumption (ref)(i) below). To the best of our knowledge, this is also a new result for the pooled estimator, as we allow for the presence of IFEs. Furthermore, we propose a Hausman-type test ( hausman1978specification) to assess the poolability, that is, whether the TW-MG and pooled TWFE estimators converge to the same probability limit. The test is also based on the leave-one-out jackknife method.

Our paper contributes to the extensive literature on panel data models with IFEs and random coefficients (RCs). For a review of the RC panel data models, see, e.g., hsiao2008random and Hsiao2022. For the review of the panel data models with IFEs, see, e.g., baiwang2016. In the large $T$ framework, pesaran2006 introduces the CCE estimator for both homogeneous and heterogeneous static panel data models with IFEs. For the PCA\ approach, bai2009, moon2015 and lu_su2016, among others, propose\ estimators for homogeneous model. li2020efficient and cao2024oracle extend this by proposing a PCA-type estimator for heterogeneous models. The literature for heterogeneous panel data models with fixed $T$ is relatively sparse. Although pesaran_smith1995's MG estimator is originally proposed for the large $T$ case, it has been shown to be effective for fixed $T$ as well (see, e.g., hsiao2019panel). chamberlain1992efficiency proposes an estimator similar to ours but does not account for IFEs. More recently, pesaran2024trimmed introduce a trimmed version of Chamberlain's estimator, which also excludes IFEs and employs a different inference approach. westerlund2022cce provide a pooled CCE estimator for heterogeneous models.

The rest of the paper is structured as follows. In Section (ref), we propose the estimation and inference methods. We aim to make Section (ref) accessible to applied researchers who are interested in applying our new methods. Section (ref) examines the asymptotics of our TW-MG estimator and the jackknife inference procedure. In Section (ref), we study the pooled estimator and propose a Hausman-type test for poolability. Section (ref) reports Monte Carlo simulation results. In Section (ref), we apply our new methods to health-care expenditure data and production data. Section (ref) concludes. The appendix contains the proofs of all main results. The online supplement contains the proofs of some technical lemmas and additional simulation results.

Notation. For a positive integer $n$, let $[n]\equiv \{1,2,\ldots ,n\},$ $n_{1}=n-1,$ $\mathbb{I}_{n}$ an $n\times n$ identity matrix, and $ \mathbf{{\Greekmath 0113} }_{n}$ an $n\times 1$ vector of ones. For a real $m\times n$ matrix $A$, we use $\Vert A\Vert ,$ $\Vert A\Vert _{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sp }}$and $ A^{\prime }$ to denote its Frobenius norm, spectral norm, and transpose, respectively. When $A$ has full column rank, let $\mathbb{P}_{A}\equiv A\left( A^{\prime }A\right) ^{-1}A^{\prime }$ and $\mathbb{M}_{A}\equiv \mathbb{I}_{m}-\mathbb{P}_{A}$, where $\equiv $ signifies a definitional relationship. Let ${\Greekmath 0116} _{\max }(\cdot )$ and ${\Greekmath 0116} _{\min }(\cdot )$ denote the largest and smallest eigenvalue of a real symmetric matrix, respectively. Let $\otimes $ denote the Kronecker product and $\mathbf{1} \{\cdot \}$ the usual indicator function. Let $\mathbf{0}_{a\times b}$ denote an $a\times b$ matrix of zeros. Let bdiag$\left( A_{1},...,A_{N}\right) $ denote the block diagonal matrix. The symbol $ \overset{p}{\rightarrow }$ denotes convergence in probability, $\overset{d}{ \rightarrow }$ convergence in distribution, and plim probability limit. We use $a_{N}\lesssim b_{N}$ to denote that $a_{N}/b_{N}=O_{p}\left( 1\right) $ as $N\rightarrow \infty .$

Model, Estimation and Inference Procedure

In this section, we introduce the model and estimation and inference procedure.

Model

Our original model is specified in ((ref)), and it is reformulated in ((ref)). Note that $E(u_{it}^{\dag })=0$ in ((ref)) under some weak assumptions (e.g., Assumption (ref)(i) in Section (ref) or Assumption (ref)(i) in Section (ref)). More importantly, $ u_{it}^{\dag }$ is also uncorrelated with a demeaned version of $x_{it}$ (by removing its individual and time effects) under Assumption (ref), which ensures that one can estimate ${\Greekmath 010C} _{i}$'s and ${\Greekmath 010C} ^{0}$ by the usual least squares approach. We emphasize that the above reformulation is solely for theoretical analysis and is not required for the implementation of our estimation and inference procedure.

Since we focus on the large $N$ and fixed $T$ framework, it is hard to consider a PCA estimate for ${\Greekmath 010C} ^{0}.$ Due to the presence of the time effects ${\Greekmath 010B} _{\cdot t}$ in the $y$-equation, one cannot pool the $T$ observations for each $i$ to run a time series regression of $y_{it}$ on $ \left( x_{it}^{\prime },1\right) $ to estimate ${\Greekmath 010C} _{i}$ as in standard MG estimation (see, e.g., pesaran_smith1995). Below we propose a TW-MG estimator of ${\Greekmath 010C} ^{0}$ via the LSDV approach. Even though we do not formally address the missing value problem in this paper, it is well known that the use of the LSDV approach can easily account for the missing value issues in panel regressions.

remark(An example of $x_{it}$) In the above setup, we only specify the DGP for $y_{it}.$ If one follows the literature on the CCE estimation (e.g., pesaran2006), one can also specify a DGP for $ x_{it}$ as follows: \begin{equation} x_{it}={\Greekmath 0116} _{i\cdot }^{o}+{\Greekmath 0116} _{\cdot t}^{o}+{\Greekmath 010D} _{i}^{o\prime }f_{t}^{o}+v_{it}, \end{equation} where ${\Greekmath 010D} _{i}^{o}$ is an $R\times K_{x}$ matrix of factor loadings in the $x$-equation, ${\Greekmath 0116} _{i\cdot }^{o}$ and ${\Greekmath 0116} _{\cdot t}^{o}$ are the individual and time fixed effects, respectively, and $v_{it}$ is the zero-mean error term. As above, one can reformulate ${\Greekmath 010D} _{i}^{o}$ and $ f_{t}^{o}$ to have zero mean such that \begin{equation} x_{it}={\Greekmath 0116} _{i\cdot }+{\Greekmath 0116} _{\cdot t}+{\Greekmath 010D} _{i}^{\prime }f_{t}+v_{it}, \end{equation} where ${\Greekmath 010D} _{i}={\Greekmath 010D} _{i}^{o}-E({\Greekmath 010D} _{i}^{o}),$ $ f_{t}=f_{t}^{o}-E(f_{t}^{o}),$ ${\Greekmath 0116} _{i\cdot }={\Greekmath 0116} _{i\cdot }^{o}+{\Greekmath 010D} _{i}^{o\prime }E(f_{t}^{o})$ and ${\Greekmath 0116} _{\cdot t}={\Greekmath 0116} _{\cdot t}^{o}+E\left( {\Greekmath 010D} _{i}^{o\prime }\right) f_{t}.$ In this paper we do not need to specify a DGP for $x_{it}.$ Even so, it is worth noting that if $x_{it}$ is indeed generated via ($\ref{x_eqn2}$), then only its double-demeaned version, $x_{it}^{o},$ will enter the asymptotics of our TW-MG estimator below.

Estimation procedure

To present the estimation procedure, we introduce some notations. Since $ {\Greekmath 010B} _{i\cdot }$ and ${\Greekmath 010B} _{\cdot t}$ are not separately identified, we need to impose a location restriction for the LSDV estimation. Below, we impose the restriction that

equation*[equation* omitted — 61 chars of source]

in the estimation procedure without requiring the true values $\left\{ {\Greekmath 010B} _{i\cdot }\right\} $ to satisfy such a restriction. Let $\mathbf{ {\Greekmath 010B} }^{\left( 1\right) }\equiv \left( {\Greekmath 010B} _{1\cdot },...,{\Greekmath 010B} _{N\cdot }\right) ^{\prime },$ $\mathbf{{\Greekmath 010B} }^{\left( 2\right) }\equiv \left( {\Greekmath 010B} _{\cdot 1},...,{\Greekmath 010B} _{\cdot T-1}\right) ^{\prime },$ and $ \mathbf{{\Greekmath 010C} }\equiv ({\Greekmath 010C} _{1}^{\prime },...,{\Greekmath 010C} _{N}^{\prime })^{\prime }.$ Define

equation[equation omitted — 367 chars of source]

where $T_{1}\equiv T-1.$ Let $\mathbf{y}_{i}=\left( y_{i1},...,y_{iT}\right) ^{\prime },$ \ $\mathbf{Y}=\left( \mathbf{y}_{1}^{\prime },...,\mathbf{y} _{N}^{\prime }\right) ^{\prime },$ $\mathbf{x}_{i}=\left( x_{i1},...,x_{iT}\right) ^{\prime }$ and $\mathbf{X}=$bdiag$(\mathbf{x} _{1},...,\mathbf{x}_{N}).$ Define $\mathbf{u}_{i}=\left( u_{i1,}...,u_{iT}\right) ^{\prime },$\ $\mathbf{U}=\left( \mathbf{u} _{1}^{\prime },...,\mathbf{u}_{N}^{\prime }\right) ^{\prime },$ $F=\left( f_{1},...,f_{T}\right) ^{\prime }$ and $\mathbf{{\Greekmath 0115} =(}{\Greekmath 0115} _{1}^{\prime },...,{\Greekmath 0115} _{N}^{\prime }\mathbf{)}^{\prime }.$ Define $ \mathbf{u}_{i}^{\dag }$ and $\mathbf{U}^{\dag }$ analogously. Then we can rewrite the model in ((ref)) in matrix form:

equation[equation omitted — 322 chars of source]

where $\mathbf{D}=\left( \mathbf{D}^{\left( 1\right) },\mathbf{D}^{\left( 2\right) }\right) ,$ $\mathbf{{\Greekmath 010B} }=(\mathbf{{\Greekmath 010B} }^{\left( 1\right) \prime },\mathbf{{\Greekmath 010B} }^{\left( 2\right) \prime })^{\prime },$ and $ \mathbf{U}^{\dag }=(\mathbb{I}_{N}\otimes F)\mathbf{{\Greekmath 0115} }+\mathbf{U}.$

By partialling out $\mathbf{D}$ in ((ref)), we obtain the following LSDV estimator of $\mathbf{{\Greekmath 010C} }$:

equation[equation omitted — 306 chars of source]

Then our TW-MG estimator of ${\Greekmath 010C} ^{0}$ is given by

equation*[equation* omitted — 99 chars of source]

Let $\mathbf{S}_{i,N}=(\mathbf{0}_{K_{x}\times K_{x}},...,\mathbf{0} _{K_{x}\times K_{x}},\mathbb{I}_{K_{x}},\mathbf{0}_{K_{x}\times K_{x}},..., \mathbf{0}_{K_{x}\times K_{x}})$ be a $K_{x}\times NK_{x}$ selection matrix that has $\mathbb{I}_{K_{x}}$ in its $i$th block. Then $\hat{{\Greekmath 010C}}^{MG}$ can also be written as

equation[equation omitted — 369 chars of source]

where $\mathbf{S}_{N}=\mathbf{{\Greekmath 0113} }_{N}^{\prime }\otimes \mathbb{I} _{K_{x}}=(\mathbb{I}_{K_{x}},...,\mathbb{I}_{K_{x}}).$

remark(An alternative representation) Note that $\mathbf{D} =(\mathbf{D}^{\left( 1\right) },\mathbf{D}^{\left( 2\right) }),$ where $ \mathbf{D}^{\left( 1\right) }$ and $\mathbf{D}^{\left( 2\right) }$ are defined in ((ref)). It is easy to verify that \begin{eqnarray*} \mathbf{D}^{\prime }\mathbf{D} &\mathbf{=}&\left( \begin{array}{cc} T\mathbb{I}_{N} & \mathbf{0}_{N\times T_{1}} \\ \mathbf{0}_{T_{1}\times N} & N\left( \mathbb{I}_{T_{1}}+\mathbf{{\Greekmath 0113} } _{T_{1}}\mathbf{{\Greekmath 0113} }_{T_{1}}^{\prime }\right) \end{array} \right) , \\ \mathbb{P}_{\mathbf{D}} &=&\mathbf{D}\left( \mathbf{D}^{\prime }\mathbf{D} \right) ^{-1}\mathbf{D}^{\prime }=\frac{\mathbf{{\Greekmath 0113} }_{N}\mathbf{{\Greekmath 0113} } _{N}^{\prime }}{N}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}+\mathbb{I} _{N}\otimes \frac{\mathbf{{\Greekmath 0113} }_{T}\mathbf{{\Greekmath 0113} }_{T}^{\prime }}{T},\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and} \\ \mathbb{M}_{\mathbf{D}} &=&\mathbb{I}_{NT}-\mathbb{P}_{\mathbf{D}}=\mathbf{D} \left( \mathbf{D}^{\prime }\mathbf{D}\right) ^{-1}\mathbf{D}^{\prime }= \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}. \end{eqnarray*} As a result, we can also write the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}$ as follows: \begin{equation} \hat{{\Greekmath 010C}}^{MG}=\frac{1}{N}\mathbf{S}_{N}\left[ \mathbf{X}^{\prime }( \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}) \mathbf{X}\right] ^{-1}\mathbf{X}^{\prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} } _{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{Y.} \end{equation} Clearly, $\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} } _{T}}$ is the double-demean (or two-way within group transformation) operator that removes the individual and time fixed effects in both $y_{it}$ and $x_{it}.$
remark(Comparison with standard MG\ estimator) Since the seminal work of pesaran_smith1995, MG\ estimation has been widely applied in panel data analyses to estimate the average marginal effect. To the best of our knowledge, standard MG estimation is applied to heterogenous panel data models with only individual effects. In this case, one can use the $T$ observations for cross-section unit $i$ and estimate the individual-specific slope coefficients ${\Greekmath 010C} _{i}$ via a time series regression and then average the resulting estimates. Due to the presence of time effects ${\Greekmath 010B} _{\cdot t}$ in ((ref)), this approach is not applicable. In contrast, in our TW-MG estimation procedure above, we first pool all $NT$ observations across $i$ and $t$ and estimate $\left\{ {\Greekmath 010C} _{i}\right\} _{i\in \left[ N\right] }$ jointly and then average the resulting estimates.
remark(Ridge-type estimator) When $T$ is small, to improve the finite sample performance, we can modify the estimator in ((ref) ) as \begin{equation} \hat{{\Greekmath 010C}}^{MG,Ridge}=\frac{1}{N}\mathbf{S}_{N}\left[ \frac{1}{T}\mathbf{X} ^{\prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}+{\Greekmath 0114} _{N}\cdot \mathbb{I}_{NK_{x}}\right] ^{-1}\frac{1}{T }\mathbf{X}^{\prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{ \mathbf{{\Greekmath 0113} }_{T}})\mathbf{Y,} \end{equation} where ${\Greekmath 0114} _{N}$ is a tuning parameter. This is a ridge-type estimator. For the choice of ${\Greekmath 0114} _{N},$ we can set ${\Greekmath 0114} _{N}=c_{{\Greekmath 0114} }\frac{1}{ N},$ where $c_{{\Greekmath 0114} }$ is a constant. We can show that all our theoretical results below hold for $\hat{{\Greekmath 010C}}^{MG,Ridge}$ as well with this choice of $ {\Greekmath 0114} _{N}.$ In the simulations and applications below, to account for the scale of the data, we let $c_{{\Greekmath 0114} }$ be the median of $\left\{ \det \left( \frac{1}{T}\sum_{t=1}^{T}\ddot{x}_{it}\ddot{x}_{it}^{\prime }\right) \right\} _{i\in \lbrack N]},$ where $\det $ stands for determinant and $ \ddot{x}_{it}$ is the double-demeaned version of $x_{it}$: $\ddot{x} _{it}=x_{it}-\frac{1}{N}\sum_{j=1}^{N}x_{jt}-\frac{1}{T}\sum_{s=1}^{T}x_{is}+ \frac{1}{NT}\sum_{j=1}^{N}\sum_{s=1}^{T}x_{js}.$

Inference procedure

Below we show that $\hat{{\Greekmath 010C}}^{MG}$ is asymptotically mixture normal under some regularity conditions. We propose a jackknife method to estimate the asymptotic variance of $\hat{{\Greekmath 010C}}^{MG}$ and to conduct inference for $ {\Greekmath 010C} ^{0}.$ Following ((ref)), we can obtain the leave-one-out estimator of ${\Greekmath 010C} ^{0}$ in ((ref)) by leaving out the observations for the $i$th cross-sectional unit:

equation[equation omitted — 511 chars of source]

where $N_{1}\equiv N-1,$ $\mathbf{X}^{\left( -i\right) \prime }$ is the leave-one-out version\ of $\mathbf{X}^{\prime }$ by deleting its $i$th row block: $\mathbf{X}^{\left( -i\right) \prime }=$bdiag$(\mathbf{x}_{1}^{\prime },$ $...,\mathbf{x}_{i-1}^{\prime },\mathbf{x}_{i+1}^{\prime },...,\mathbf{x} _{N}^{\prime }),$ and $\mathbf{Y}^{\left( -i\right) }=\left( \mathbf{y} _{1}^{\prime },...,\mathbf{y}_{i-1}^{\prime },\mathbf{y}_{i+1}^{\prime },..., \mathbf{y}_{N}^{\prime }\right) ^{\prime }.$

Given the jackknife estimates $\{\hat{{\Greekmath 010C}}^{MG\left( -i\right) }\}_{i\in \left[ N\right] },$ we propose to estimate the asymptotic variance of $\sqrt{ N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})$ by

equation*[equation* omitted — 303 chars of source]

where $\overline{\hat{{\Greekmath 010C}}}^{_{MG}}\equiv \frac{1}{N}\sum_{i=1}^{N}\hat{ {\Greekmath 010C}}^{MG\left( -i\right) }.$ Let $\mathbb{S}$ be a $K_{x}\times 1$ selection vector, which defines the scalar parameter of interest: $\mathbb{S} ^{\prime }{\Greekmath 010C} ^{0}.$ Let $z_{{\Greekmath 011C} /2}$ denote the upper ${\Greekmath 011C} /2$-quantile of $N(0,1).$ Then we construct the $100(1-{\Greekmath 011C} )\%$ confidence interval (CI) of $\mathbb{S}^{\prime }{\Greekmath 010C} ^{0}$ as

equation[equation omitted — 379 chars of source]

Asymptotic Properties of $\hat{\protect{\Greekmath 010C}}^{MG}$

In this section, we study the asymptotic properties of $\hat{{\Greekmath 010C}}^{MG}$. We first state the assumptions and show that $\sqrt{N}(\hat{{\Greekmath 010C}} ^{MG}-{\Greekmath 010C} ^{0})$ follows a version of stable CLT. Then we show that the above jackknife method can estimate the asymptotic variance of $\sqrt{N}( \hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})$ consistently.

Basic assumptions

To proceed, we introduce some notations. Let $\mathbf{\hat{Q}}=\frac{1}{T} \mathbf{X}^{\prime }{\mathbb{M}}_{\mathbf{D}}\mathbf{X}\ $and $\mathbf{\ddot{ X}}={\mathbb{M}}_{\mathbf{D}}\mathbf{X}.$ Noting that $\mathbb{M}_{\mathbf{D} }=\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}},$ we can write $\mathbf{\ddot{X}}=$bdiag$\left( \mathbf{\ddot{x}}_{1},..., \mathbf{\ddot{x}}_{N}\right) ,$ where $\mathbf{\ddot{x}}_{1}=(\ddot{x} _{i1},...,\ddot{x}_{iT})^{\prime }$ and $\ddot{x}_{it}=x_{it}-\frac{1}{N} \sum_{j=1}^{N}x_{jt}-\frac{1}{T}\sum_{s=1}^{T}x_{is}+\frac{1}{NT} \sum_{j=1}^{N}\sum_{s=1}^{T}x_{js}.$ In addition,

equation*[equation* omitted — 163 chars of source]

where $\mathbf{X}^{o}=$bdiag$\left( \mathbf{x}_{1}^{o},...,\mathbf{x} _{N}^{o}\right) ,$ $\mathbf{x}_{i}^{o}=(x_{i1}^{o},...,x_{iT}^{o})^{\prime }, $ and $x_{it}^{o}$ denotes the part of\ $x_{it}$ with its additive individual and time effects (if any) removed. For example, if $x_{it}$ is generated from ($\ref{x_eqn2}$), then $x_{it}^{o}={\Greekmath 010D} _{i}^{\prime }f_{t}+v_{it}$. In the case where $x_{it}$ does not contain any additive term that is varying along either $i$ or $t$ but not both, we have that $ x_{it}^{o}=x_{it}$ and $\mathbf{x}_{i}^{o}=\mathbf{x}_{i}.$ Define

equation*[equation* omitted — 321 chars of source]

where $m_{ij}=\mathbf{1}\left\{ i=j\right\} -\frac{1}{N}$ denotes the $ \left( i,j\right) $th element of $\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}.$

Let $\mathcal{F}_{N,0}^{MG}={\Greekmath 011B} \left( \{\mathbf{x}_{i}^{o}\}_{1\leq i\leq N},F\right) ,$ the minimal sigma-field generated by $\{\mathbf{x} _{i}^{o}\}_{1\leq i\leq N}$ and $F.\footnote{ If $x_{it}$ is generated from ($(ref)$), we can equivalently define $ \mathcal{F}_{N,0}^{MG}={\Greekmath 011B} \left( \{\left( \mathbf{{\Greekmath 010D} }_{i},\mathbf{v} _{i}\right) \}_{1\leq i\leq N},F\right) .$}$ Let $\mathcal{F} _{N,j}^{MG}={\Greekmath 011B} (\mathcal{F}_{N,0}^{MG},\left\{ \mathbf{u}_{i},{\Greekmath 0115} _{i}\right\} _{i=1}^{j})$ for $j\in \left[ N\right] .$ Let $ \max_{i}=\max_{1\leq i\leq N},$ and $\min_{i}=\min_{1\leq i\leq N}$. \ Let $ \underline{c},$ $\overline{c}\in \left( 0,\infty \right) $ be some generic constants that do not depend on $N$ and whose values may vary across places. Let $E_{A}\left( \cdot \right) =E\left( \cdot |A\right) .$ Let $\tilde{q} _{ii}=\frac{1}{T}\mathbf{x}_{i}^{o\prime }{\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}} \mathbf{x}_{i}^{o}.$ To study the asymptotic properties of $\hat{{\Greekmath 010C}} ^{MG}, $ we impose the following two assumptions.

assumption(i) Conditional on $\mathcal{F}_{N,0}^{MG},$ $\left\{ ( \mathbf{u}_{i},{\Greekmath 0115} _{i})\right\} _{i\in \left[ N\right] }$ are independent across $i$ with zero mean and $E_{\mathcal{F}_{N,0}^{MG}}[\left \Vert \mathbf{u}_{i}\right\Vert ^{4+2{\Greekmath 010E} }+\left\Vert {\Greekmath 0115} _{i}\right\Vert ^{4+2{\Greekmath 010E} }]\lesssim 1$ for some ${\Greekmath 010E} >0;$ (ii) $\left\{ {\Greekmath 0111} _{i}\right\} _{i\in \left[ N\right] }$ are independent across $i$ with zero mean, $\frac{1}{N}\sum_{i=1}^{N}E({\Greekmath 0111} _{i}{\Greekmath 0111} _{i}^{\prime })\rightarrow \Omega _{{\Greekmath 0111} },$ and $E\left\Vert {\Greekmath 0111} _{i}\right\Vert ^{2+{\Greekmath 010E} }\leq \overline{c};$ (iii) $E_{\mathcal{F}_{N,i-1}^{MG}}(\mathbf{u}_{i}{\Greekmath 0111} _{i}^{\prime })=0;$ (iv) $E_{F}\left\Vert \mathbf{x}_{i}^{o}\right\Vert ^{4+2{\Greekmath 010E} }\lesssim 1.$ \
assumption(i) $\tilde{q}_{ii}$ has full rank almost surely (a.s.) and $ \frac{1}{N}\sum_{i=1}^{N}\left\Vert \tilde{q}_{ii}^{-1}\right\Vert ^{2}=O_{p}(1)$; (ii) $\frac{1}{N}\sum_{i=1}^{N}{\Greekmath 011F} _{i}^{\prime }E_{\mathcal{F}_{N,0}^{MG}}[ \mathbf{u}_{i}^{\dag }\mathbf{u}_{i}^{\dag \prime }]{\Greekmath 011F} _{i}\overset{p}{ \rightarrow }\Omega _{1}^{MG}$ and $\Omega _{1}^{MG}$ is positive definite almost surely (a.s.).

Assumption (ref)(i)--(iii) impose some dependence and moment conditions on the unobserved components, i.e., $\mathbf{u}_{i},{\Greekmath 0115} _{i},$ and ${\Greekmath 0111} _{i}.$ Noting that $\mathbf{X}^{o}$ is measurable with respect to $ \mathcal{F}_{N,0}^{MG},$ Assumption (ref)(i) implies the independence of $(\mathbf{u}_{i},{\Greekmath 0115} _{i})$ across $i$ conditional on $ \mathbf{X}^{o}.$ The zero conditional mean condition ensures that the composite error term $u_{it}^{\dag }$ ($={\Greekmath 0115} _{i}^{\prime }f_{t}+u_{it}$ ) can be treated as a usual error term in least squares regression as the individual and time effects can be partialled out via the usual partitioned regressions. Clearly, Assumption (ref)(i) essentially requires that the regressors $x_{it}$, after its individual and time effects are removed, are strictly exogenous so that a dynamic panel is ruled out. Assumption (ref)(ii) imposes standard conditions on the heterogeneous parameters $ {\Greekmath 0111} _{i}.$ Note that we allow $\Omega _{{\Greekmath 0111} }$ to be positive semidefinite so that some elements in ${\Greekmath 010C} _{i}$ are allowed to be homogenous across $ i. $ Assumption (ref)(iii) requires the conditional uncorrelatedness between $\mathbf{u}_{i}$ and ${\Greekmath 0111} _{i}$ because $E(\mathbf{ u}_{i}|\mathcal{F}_{N,i-1}^{MG})=0$ under Assumption (ref)(i). Assumption (ref)(iv) imposes some moment conditions on $\mathbf{x} _{i}^{o}$ that will be used to verify certain Lyapunov conditions. In case where $x_{it}$ is generated according to ((ref)), this moment condition can be replaced by

Assumption 3.1(iv*) $E\left\Vert {\Greekmath 010D} _{i}\right\Vert ^{4+2{\Greekmath 010E} }+E\left\Vert \mathbf{v}_{i}\right\Vert ^{4+2{\Greekmath 010E} }\leq \overline{c},$ where $\mathbf{v}_{i}=(v_{i1},...,v_{iT})^{\prime }.$

It is worth mentioning that we do not impose any conditions on the individual and time effects, viz. ${\Greekmath 010B} _{i\cdot }$ and ${\Greekmath 010B} _{\cdot t}$ for the $y$-equation and ${\Greekmath 0116} _{i\cdot }$ and ${\Greekmath 0116} _{\cdot t}$ in the $x$ -equation if $x_{it}$ is generated according to ((ref)) as they are partialled out in the LSDV regression. Neither do we impose any condition on the unobserved factor $f_{t}$ due to our focus on the fixed $T$ framework. In other words, we allow various kinds of dependence, trending or nonstationary behavior in ${\Greekmath 010B} _{\cdot t},$ ${\Greekmath 0116} _{\cdot t}$ and $f_{t}.$ In addition, we do not impose any dependence condition on $\left\{ f_{t},u_{it},x_{it}\right\} $ along the time dimension.

Assumption (ref)(i) specifies some rank and moment conditions on the use of TW-MG\ estimator when a time effect is present in the $y$ -equation, which implicitly requires that $T>K_{x}+1.$ When $T\leq K_{x}$, there is no way to consider the MG estimator even without the individual or time effects in the $y$-equation. In this case, one can only consider the pooled regression to estimate the population average marginal effect ${\Greekmath 010C} ^{0}$ as in the next section. When $T=K_{x}+1,$ we run into the irregular case discussed in pesaran2024trimmed even if $\tilde{q}_{ii}$ may still be full rank a.s.. Note that in the fixed $T$-framework, $\tilde{q} _{ii}$ can be of full rank but with minimum eigenvalue that may take values with positive probability in the neighborhood of 0. A sufficient condition for the second part of Assumption (ref)(i) to hold is $E\left\Vert \tilde{q}_{ii}^{-1}\right\Vert ^{2}\leq \bar{c}.$ Assumption (ref) (ii)\ specifies a convergence condition on the conditional second moments of ${\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dag },$ which appears in the Bahadur representation of the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}.$ Due to the presence of the unobserved factor component, $F,$ which enters both ${\Greekmath 011F} _{i}$ and $\mathbf{u}_{i}^{\dag },$ the probability limit object $\Omega _{1}^{MG}$ is random unless all factors inside $F$ are nonrandom. This condition facilitates the application of a stable martingale CLT as introduced in hausler2015stable.

remark(Allowance of Correlated Random Coefficients) We allow correlation between ${\Greekmath 0111} _{i}$ and $x_{it}$ via various components in $x_{it}$ including ${\Greekmath 0116} _{i\cdot },$ ${\Greekmath 010D} _{i}$ and $v_{it}$ if $x_{it}$ is generated according to ((ref)). This largely broadens potential applications of correlated random coefficient (CRC) panel data models. For a textbook treatment on random coefficient panels, see Chapter 13 in Hsiao2022. For some recent applications and studies on CRC panels, see arellano2012identifying, graham2012identification, and lu2023, among others.

Asymptotic normality of $\hat{\protect{\Greekmath 010C}}^{MG}$

To state the first main result in this paper, let $\mathcal{F} _{0}^{MG}={\Greekmath 011B} \left( \{\mathbf{x}_{i}^{o}\}_{1\leq i\leq \infty },F\right) .$ The following theorem reports the stable convergence of $\sqrt{ N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0}).$

theoremSuppose that Assumptions (ref)--(ref) hold. Then \begin{equation} \sqrt{N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})=\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\left( {\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dagger }+{\Greekmath 0111} _{i}\right) \rightarrow (\Omega ^{MG})^{1/2}\mathbb{Z}\mathcal{\ \ F}_{0}^{MG}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-stably as } N\rightarrow \infty , \end{equation} where $\Omega ^{MG}=\Omega _{1}^{MG}+\Omega _{{\Greekmath 0111} },\ \mathbb{Z}\backsim N\left( 0,\mathbb{I}_{K_{x}}\right) ,$ and $\mathbb{Z}$ is independent of $ \mathcal{F}_{0}^{MG}.$

Theorem (ref) indicates that the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}$ is $\sqrt{N}$-consistent and follows a stable CLT under some regularity conditions. The sequence $\{\sqrt{N}(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})\}_{N\geq 1} $ converges $\mathcal{F}$-stably as $N\rightarrow \infty $ to a mixture normal distribution whose variance is random but measurable with respect to the limit sigma-field $\mathcal{F}_{0}^{MG}$. We refer the readers directly to hausler2015stable and van2023weak for the notion of stable convergence and stable CLT. For recent applications of the concept of stable convergence in econometrics and statistics, see lee2019stable, cavaliere2020inference, Jin_Miao_Su2021, mikusheva2021many and hahn2024econometric. It is worth mentioning that the usual weak convergence concept, a standard tool for asymptotic analysis in econometrics, is not adequate for our asymptotic analysis. Stable convergence is a middle ground between convergence in probability and convergence in distribution. It is weaker than convergence in probability but stronger than weak convergence. Importantly, it allows for a stochastic normalization in the central limit theorem. Intuitively, stable convergence can be thought of as a weak convergence conditional on certain sigma-algebra ($\mathcal{F}_{0}^{MG}$ here) so that it can be recast as a convergence of certain conditional distributions.

Asymptotic inference on $\protect{\Greekmath 010C} ^{0}$

In this subsection, we address the problem of inference on $\mathbb{S} ^{\prime }{\Greekmath 010C} ^{0}$ based on the jackknife method for the TW-MG estimation. We add the following condition.

assumption(i)$\frac{1}{N}\sum_{i=1}^{N}\left\Vert \tilde{q} _{ii}^{-1}\right\Vert ^{2}\left\Vert \mathbf{x}_{i}^{o}\right\Vert ^{2s}=O_{p}(1)$ for $s=0,1$. (ii) ${\Greekmath 010E} >1.$

Assumption (ref)(i) strengthens Assumption (ref)(i). As sufficient condition for it to hold is $\frac{1}{N}\sum_{i=1}^{N}E[\left \Vert \tilde{q}_{ii}^{-1}\right\Vert ^{2}\left\Vert \mathbf{x} _{i}^{o}\right\Vert ^{2s}]=O(1)$ for $s=0,1.$ Assumption (ref)(ii) strengthens the moment condition in (ref) by requiring ${\Greekmath 010E} >1.$ Both parts of Assumption (ref) are reasonable because consistent estimation of the variance-covariance matrix generally requires stronger conditions than those needed for the asymptotic distribution result.

The following theorem establishes the consistency of $\hat{\Omega}^{MG}\ $ and the distributional limit for a normalized version of $\sqrt{N}\mathbb{S} ^{\prime }(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0}).$

theoremSuppose Assumptions (ref)--(ref) hold. Then (i) $\hat{\Omega}^{MG}=\frac{1}{N}\sum_{j=1}^{N}({\Greekmath 011F} _{j}^{\prime }\mathbf{u }_{j}^{\dagger }+{\Greekmath 0111} _{j})({\Greekmath 011F} _{j}^{\prime }\mathbf{u}_{j}^{\dagger }+{\Greekmath 0111} _{j})^{\prime }+o_{p}\left( 1\right) =\Omega ^{MG}+o_{p}\left( 1\right) ;$ (ii) $(\mathbb{S}^{\prime }\hat{\Omega}^{MG}\mathbb{S})^{-1/2}\sqrt{N} \mathbb{S}^{\prime }(\hat{{\Greekmath 010C}}^{MG}-{\Greekmath 010C} ^{0})\overset{d}{\rightarrow } N\left( 0,1\right) \ $as $N\rightarrow \infty .$

Theorem (ref)(i) indicates that we can estimate $\Omega ^{MG}$ consistently via the jackknife method and Theorem (ref)(ii) justifies the asymptotic validity of the jackknife confidence interval defined in ((ref)).

remark(High dimensional jackknife) The jackknife method is a popular method to estimate the asymptotic variance of an estimator whose dimension does not grow with the sample size; see Chapter 10.3 in hansen2022econometrics for a brief account. However, our TW-MG estimator $ \hat{{\Greekmath 010C}}^{MG}$ is based on $\mathbf{\hat{{\Greekmath 010C}}},$ whose row dimension diverges to infinity linearly in the sample size $N.$ To study the asymptotic expansion of $\hat{{\Greekmath 010C}}^{MG\left( -i\right) },$ we have to study that of $\mathbf{\hat{{\Greekmath 010C}}}^{\left( -i\right) },$ which is an $ N_{1}K_{x}\times 1$ vector. The change of the row dimension of $\mathbf{\hat{ {\Greekmath 010C}}}$ to that of $\mathbf{\hat{{\Greekmath 010C}}}^{\left( -i\right) }$ substantially complicates the asymptotic analysis. Let $\mathbf{\hat{Q}}^{\left( -i\right) }\equiv \frac{1}{T}\mathbf{X}^{\left( -i\right) \prime }(\mathbb{M}_{\mathbf{ {\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}^{\left( -i\right) }=\frac{1}{T}\mathbf{X}^{o\left( -i\right) \prime }(\mathbb{M}_{ \mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X} ^{o\left( -i\right) },$ where $\mathbf{X}^{o\left( -i\right) \prime }$ is the leave-one-out version of $\mathbf{X}^{o}.$ Noting that \begin{equation*} \mathbf{\hat{{\Greekmath 010C}}}=\mathbf{\hat{Q}}^{-1}\frac{1}{T}\mathbf{X}^{\prime }( \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}) \mathbf{Y}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }\mathbf{\hat{{\Greekmath 010C}}}^{\left( -i\right) }=\left[ \mathbf{\hat{Q}}^{\left( -i\right) }\right] ^{-1}\frac{1}{T}\mathbf{X} ^{\left( -i\right) \prime }(\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes \mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{Y}^{\left( -i\right) },\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi \end{equation*} we make the following decompositions: \begin{eqnarray*} \mathbf{\hat{Q}} &\equiv &\frac{1}{T}\mathbf{X}^{o\prime }\left( {\mathbb{I}} _{N}\otimes {\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\right) \mathbf{X}^{o}-\frac{1 }{T}\mathbf{X}^{o\prime }\left( {\mathbb{P}}_{\mathbf{{\Greekmath 0113} }_{N}}\otimes { \mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\right) \mathbf{X}^{o}\equiv \mathbf{Q} ^{\left( d\right) }-\mathbf{CC}^{\prime }\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and} \\ \mathbf{\hat{Q}}^{\left( -i\right) } &=&\frac{1}{T}\mathbf{X}^{o\left( -i\right) \prime }\left( {\mathbb{I}}_{N_{1}}\otimes {\mathbb{M}}_{\mathbf{ {\Greekmath 0113} }_{T}}\right) \mathbf{X}^{o\left( -i\right) }-\frac{1}{T}\mathbf{X} ^{o\left( -i\right) \prime }({\mathbb{P}}_{\mathbf{{\Greekmath 0113} }_{N_{1}}}\otimes { \mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}^{o\left( -i\right) }\equiv \mathbf{Q}^{\left( -i,d\right) }-\mathbf{C}^{\left( -i\right) }\mathbf{C} ^{\left( -i\right) \prime }, \end{eqnarray*} where $\mathbf{C}=(NT)^{-1/2}\mathbf{X}^{o\prime }[$${\Greekmath 0113} $$ _{N}\otimes {\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}]$ and $\mathbf{C}^{\left( -i\right) }=\left( N_{1}T\right) ^{-1/2}\mathbf{X}^{o\left( -i\right) \prime }[$${\Greekmath 0113} $$_{N_{1}}\otimes {\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}].$ By the Sherman-Morrison Woodbury matrix identity (e.g., Fact 6.4.31 in bernstein2005), we have \begin{equation*} \left[ \mathbf{\hat{Q}}^{\left( -i\right) }\right] ^{-1}=\left[ \mathbf{Q} ^{\left( -i,d\right) }\right] ^{-1}\left\{ \mathbf{Q}^{\left( -i,d\right) }+ \mathbf{C}^{\left( -i\right) }[\mathbb{I}_{K_{x}T}+\mathbf{C}^{\left( -i\right) \prime }\mathbf{C}^{\left( -i\right) }]^{-1}\mathbf{C}^{\left( -i\right) \prime }\right\} \left[ \mathbf{Q}^{\left( -i,d\right) }\right] ^{-1}. \end{equation*} A similar expression holds for $\mathbf{\hat{Q}}^{-1}.$ Due to the mismatch of the dimensions of various matrix pairs such as $\mathbf{Q}^{\left( -i,d\right) }$ and $\mathbf{Q}^{\left( d\right) },$ $\mathbf{C}^{\left( -i\right) }$ and $\mathbf{C},$ $\mathbf{X}^{\left( -i\right) }$ and $\mathbf{ X},$ and $\mathbf{Y}^{\left( -i\right) }$ and $\mathbf{Y},$ one key step in our analysis is to augment the leave-one-out version of various matrices/vectors to match the dimension of the original version. For example, we augment the $N_{1}K_{x}\times N_{1}K_{x}$ block diagonal matrix $ \mathbf{Q}^{\left( -i,d\right) }$ to $\mathbf{\breve{Q}}^{\left( -i,d\right) }$ (an $NK_{x}\times NK_{x}$ block diagonal matrix) such that \begin{equation*} \left[ \mathbf{\breve{Q}}^{\left( -i,d\right) }\right] _{jj}=\left\{ \begin{array}{ll} \left[ \mathbf{Q}^{\left( -i,d\right) }\right] _{jj} & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }j<i \\ \mathbf{0}_{K_{x}\times K_{x}} & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }j=i \\ \left[ \mathbf{Q}^{\left( -i,d\right) }\right] _{j-1,j-1}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ } & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }j>i \end{array} \right. , \end{equation*} where $\left[ \mathbf{Q}^{\left( -i,d\right) }\right] _{jj}$ denotes the $ \left( j,j\right) $th block of $\mathbf{Q}^{\left( -i,d\right) }$ of size $ K_{x}\times K_{x}.$ We do similar things for matrices $\mathbf{C}^{\left( -i\right) },$ $\mathbf{X}^{\left( -i\right) },$ $\mathbf{Y}^{\left( -i\right) },$ etc., to obtain their augmented versions $\mathbf{\breve{C}} ^{\left( -i\right) },$ $\mathbf{\breve{X}}^{\left( -i\right) },$ $\mathbf{ \breve{Y}}^{\left( -i\right) },$ etc. By using these augmented matrices associated with some tedious matrix algebra, we can ultimately establish the link between $\hat{{\Greekmath 010C}}^{MG\left( -i\right) }-\frac{1}{N_{1}}\sum_{j\neq i}^{N}{\Greekmath 010C} _{j}$ and $\hat{{\Greekmath 010C}}^{MG}-\frac{1}{N}\sum_{j=1}^{N}{\Greekmath 010C} _{j}:$ \begin{equation} \hat{{\Greekmath 010C}}^{MG\left( -i\right) }-\frac{1}{N_{1}}\sum_{j\neq i}^{N}{\Greekmath 010C} _{j}=\frac{N}{N_{1}}\left( \hat{{\Greekmath 010C}}^{MG}-\frac{1}{N}\sum_{j=1}^{N}{\Greekmath 010C} _{j}\right) +\Delta _{i2}^{MG}+\Delta _{i3}^{MG}+o_{p}(N^{-1}), \end{equation} where $\Delta _{i2}^{MG}\equiv -\frac{1}{N_{1}T}\hat{{\Greekmath 0121}}_{i}\mathbf{x} _{i}^{o\prime }{\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\mathbf{u}_{i}^{\dag },$ $ \Delta _{i3}^{MG}\equiv \frac{1}{NN_{1}}\sum_{j=1}^{N}\hat{{\Greekmath 0121}}_{j}\frac{1 }{T}\mathbf{x}_{j}^{o\prime }{\mathbb{M}}_{\mathbf{{\Greekmath 0113} }_{T}}\mathbf{u} _{i}^{\dag },\ $and $o_{p}(N^{-1})$ holds uniformly in $i\in \left[ N\right] .$ $\Delta _{i2}^{MG}+\Delta _{i3}^{MG}$ plays the role of a summand in the influence function of $\hat{{\Greekmath 010C}}^{MG}-\frac{1}{N}\sum_{j=1}^{N}{\Greekmath 010C} _{j}$ . The result in ((ref)) is a fundamental result that helps deliver the main result in Theorem (ref)(i).
remark(Alternative approaches) In the MG estimation literature, one commonly uses the following formula to estimate the asymptotic variance of $\hat{{\Greekmath 010C}}^{MG}$ \begin{equation*} \frac{1}{N}\sum_{i=1}^{N}\left( \hat{{\Greekmath 010C}}_{i}-\hat{{\Greekmath 010C}}^{MG}\right) \left( \hat{{\Greekmath 010C}}_{i}-\hat{{\Greekmath 010C}}^{MG}\right) ^{\prime }. \end{equation*} However, it turns out that this variance estimator is hard to justify in our framework. The main reason is that $\hat{{\Greekmath 010C}}_{i}$ is not a consistent estimator of ${\Greekmath 010C} _{i}$ in the fixed $T$ framework. Alternatively, we may consider cross-sectional bootstrap by resampling the data along the cross-sectional dimension. We find that the cross-sectional bootstrap and our leave-one-out jackknife method perform similarly in simulations. However, the asymptotic validity of the cross-sectional bootstrap is much harder to justify theoretically; therefore, we propose the jackknife method here.

Pooled Estimation and Inference

In this section, we consider the pooled estimation for the model in ((ref)) and study the inference theory.

Estimation and inference procedure

We rewrite the model in ((ref)) as follows:

equation[equation omitted — 149 chars of source]

where $u_{it}^{\ddag }=u_{it}+{\Greekmath 0115} _{i}^{\prime }f_{t}+{\Greekmath 0111} _{i}^{\prime }x_{it}$. Let $X=\mathbf{(x}_{1}^{\prime },...,\mathbf{x}_{N}^{\prime })^{\prime }$ and $\mathbf{{\Greekmath 0111} }=({\Greekmath 0111} _{1}^{\prime },...,{\Greekmath 0111} _{N}^{\prime })^{\prime }.$ Then we can rewrite the model in ((ref)) in matrix form:

equation[equation omitted — 116 chars of source]

where $\mathbf{U}^{\ddag }\equiv \mathbf{U}+(\mathbb{I}_{N}\otimes F)\mathbf{ {\Greekmath 0115} +X{\Greekmath 0111} }.$ By partialling out $\mathbf{D}$ in ((ref)), we obtain the following pooled estimator of ${\Greekmath 010C} ^{0}$:

equation[equation omitted — 172 chars of source]

Below we show that $\hat{{\Greekmath 010C}}^{P}$ is asymptotically mixture normal under some regularity conditions. We propose a jackknife method to estimate its asymptotic variance. As in ((ref)), we can obtain the leave-one-out pooled estimator of ${\Greekmath 010C} ^{0}$ by leaving out the observations for the $i$ th cross-section unit:

equation[equation omitted — 427 chars of source]

where $X^{\left( -i\right) }$ is the leave-one-out version\ of $X:X^{\left( -i\right) }=(\mathbf{x}_{1}^{\prime },...,\mathbf{x}_{i-1}^{\prime },\mathbf{ x}_{i+1}^{\prime },...,\mathbf{x}_{N}^{\prime })^{\prime }$. Given the jackknife estimates $\{\hat{{\Greekmath 010C}}^{P\left( -i\right) }\}_{i\in \left[ N \right] },$ we propose to estimate the asymptotic variance of $\sqrt{N}(\hat{ {\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0})$ by

equation*[equation* omitted — 297 chars of source]

where $\overline{\hat{{\Greekmath 010C}}}^{_{P}}=\frac{1}{N}\sum_{i=1}^{N}\hat{{\Greekmath 010C}} ^{P\left( -i\right) }.$ Then one can construct the confidence intervals for element of $\mathbb{S}{\Greekmath 010C} ^{0}$ as in ((ref)) with $\hat{\Omega}^{MG}$ replaced by $\hat{\Omega}^{P}$.

Asymptotic properties

To study the asymptotic properties of $\hat{{\Greekmath 010C}}^{P},$ we add some notations and assumptions. Let $\mathbf{\dot{x}}_{i}\equiv \mathbb{M}_{ \mathbf{{\Greekmath 0113} }_{T}}\mathbf{x}_{i}=(\dot{x}_{i1},...,\dot{x}_{iT})^{\prime }$ and $\mathbf{\dot{u}}_{i}=\mathbb{M}_{\mathbf{{\Greekmath 0113} }_{T}}\mathbf{u}_{i}=( \dot{u}_{i1},...,\dot{u}_{iT})^{\prime }.$ Let $\mathbf{\dot{X}}\equiv { \mathbb{({\mathbb{I}}_{N}\otimes M}}_{\mathbf{{\Greekmath 0113} }_{T}})\mathbf{X}=$bdiag$ (\mathbf{\dot{x}}_{1},...,\mathbf{\dot{x}}_{N})$ and $\mathbf{\dot{U}} ^{\ddagger }=\mathbf{\dot{U}}+(\mathbb{I}_{N}\otimes \dot{F})\mathbf{{\Greekmath 0115} +\dot{X}{\Greekmath 0111} }=(\mathbf{\dot{u}}_{1}^{\ddagger \prime },...,\mathbf{\dot{u}} _{N}^{\ddagger \prime })^{\prime }$ where $\mathbf{\dot{u}}_{i}^{\ddagger }=( \dot{u}_{i1}^{\ddagger },...,\dot{u}_{iT}^{\ddagger })^{\prime }$ and $\dot{u }_{it}^{\ddagger }=\dot{u}_{it}+\dot{f}_{t}^{\prime }{\Greekmath 0115} _{i}+\dot{x} _{it}^{\prime }{\Greekmath 0111} _{i}.$ Let

equation*[equation* omitted — 121 chars of source]

where $\ddot{X}={\mathbb{M}}_{\mathbf{D}}X.$ Let $\mathcal{F}_{0}^{P}={\Greekmath 011B} \left( \{\mathbf{\dot{x}}_{i}\}_{1\leq i\leq \infty },F\right) ,$ $\mathcal{F }_{N,0}^{P}={\Greekmath 011B} (\{\mathbf{\dot{x}}_{i}\}_{1\leq i\leq N},F)$ and $ \mathcal{F}_{N,j}^{P}={\Greekmath 011B} (\mathcal{F}_{N,0}^{P},\left\{ (\mathbf{u} _{i},{\Greekmath 0115} _{i},{\Greekmath 0111} _{i})\right\} _{i=1}^{j})$ for $j\in \left[ N\right] . $ Note that $\mathcal{F}_{N,0}^{P}$ (resp. $\mathcal{F}_{0}^{P}$) slightly differs from $\mathcal{F}_{N,0}^{MG}$ (resp. $\mathcal{F}_{0}^{MG}$) because $x_{it}^{\prime }{\Greekmath 0111} _{i}$ now enters the error term $u_{it}^{\ddag }$ in the pooled regression and the time effects in $x_{it}$ (if any) cannot be removed in the estimation procedure.

assumption(i) Conditional on $\mathcal{F}_{N,0}^{P},$ $\left\{ ( \mathbf{u}_{i},{\Greekmath 0115} _{i},{\Greekmath 0111} _{i})\right\} _{i\in \left[ N\right] }$ are independent across $i$ with zero mean and $E_{\mathcal{F}_{N,0}^{P}}[\left \Vert \mathbf{u}_{i}\right\Vert ^{4+2{\Greekmath 010E} }+\left\Vert {\Greekmath 0115} _{i}\right\Vert ^{4+2{\Greekmath 010E} }+\left\Vert {\Greekmath 0111} _{i}\right\Vert ^{4+2{\Greekmath 010E} }]\lesssim 1$ for some ${\Greekmath 010E} >0;$ (ii) $E_{F}[\left\Vert \mathbf{\dot{x}}_{i}\right\Vert ^{4+2{\Greekmath 010E} }]\lesssim 1.$
assumption(i) $\hat{Q}\overset{p}{\rightarrow }Q$ and $Q$ is positive definite a.s.; (ii) $\frac{1}{NT^{2}}\sum_{i=1}^{N}\mathbf{\ddot{x}}_{i}^{\prime }E_{ \mathcal{F}_{N,0}^{P}}\left[ \mathbf{\dot{u}}_{i}^{\ddag }\mathbf{\dot{u}} _{i}^{\ddag \prime }\right] \mathbf{\ddot{x}}_{i}=\Omega ^{P}+o_{p}\left( 1\right) \ $and $\Omega ^{P}$ is positive definite a.s..

Assumption (ref)(i) parallels Assumption (ref)(i)--(iii) and it imposes some dependence and moment conditions on the unobserved components, i.e., $\mathbf{u}_{i},{\Greekmath 0115} _{i},$ and ${\Greekmath 0111} _{i}.$ Since $ x_{it}^{\prime }{\Greekmath 0111} _{i}$ is a treated as part of the composite error term $ u_{it}^{\ddagger },$ we now need to assume that ${\Greekmath 0111} _{i}$ has finite $ \left( 4+2{\Greekmath 010E} \right) $th moment conditional on $\mathcal{F}_{N,0}^{P}.$ The condition that $E\left( {\Greekmath 0111} _{i}|\mathcal{F}_{N,0}^{P}\right) =\mathbf{0 }_{K_{x}}$ rules out potential correlations between ${\Greekmath 0111} _{i}$ and $\mathbf{ \dot{x}}_{i}$ via components in $x_{it}$ other than the individual effects. That is, if there is any correlation between ${\Greekmath 0111} _{i}$ and $x_{it}$ as in CRC\ panels, the correlation can only be caused by ${\Greekmath 0111} _{i}$ and a time-invariant additive component in $x_{it}.$ When $x_{it}$ is generated according to ((ref)), ${\Greekmath 0111} _{i}$ can be correlated with ${\Greekmath 0116} _{i\cdot }$ but not ${\Greekmath 0116} _{\cdot t},$ ${\Greekmath 010D} _{i}$ and $v_{it}.$ Assumption (ref)(ii) parallels Assumption (ref)(iv). It is easy to see that Assumption (ref)(ii) implies that $E_{F}[\left\Vert \mathbf{\dot{x}}_{i}^{o}\right\Vert ^{4+2{\Greekmath 010E} }]\lesssim 1$ and $ E_{F}\left\Vert \mathbf{\ddot{x}}_{i}\right\Vert ^{4+2{\Greekmath 010E} }\lesssim 1.$

Assumption (ref)(i)-(ii) parallels Assumption (ref) (i)--(ii). In comparison with Assumption (ref)(i), Assumption (ref)(i) only requires the probability limit of $Q$ be positive definite a.s., which may hold even when $T\leq K_{x}\ $so that the TW-MG estimator $\hat{{\Greekmath 010C}}^{MG}$ is not well defined. Assumption (ref) (ii)\ specifies a convergence condition on the conditional second moments of $\mathbf{\ddot{x}}_{i}^{\prime }\mathbf{\dot{u}}_{i}^{\ddag },$ which appears in the Bahadur representation of the pooled estimator $\hat{{\Greekmath 010C}} ^{P}.$ Due to the presence of the unobserved factor component, $F,$ which surely enters $\mathbf{\dot{u}}_{i}^{\ddag }$ and may also enter $\mathbf{ \ddot{x}}_{i}$ if $x_{it}$ also contains a factor structure, the probability limit object $\Omega ^{P}$ is random unless all factors inside $F$ are nonrandom. This condition facilitates the application of a stable martingale CLT as introduced in hausler2015stable.

The following theorem reports the stable convergence of $\sqrt{N}(\hat{{\Greekmath 010C}} ^{P}-{\Greekmath 010C} ^{0}).$

theoremSuppose that Assumptions (ref)--(ref) hold. Then \begin{equation*} \sqrt{N}(\hat{{\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0})=\hat{Q}^{-1}\frac{1}{\sqrt{N}} \sum_{i=1}^{N}\frac{1}{T}\mathbf{\ddot{x}}_{i}^{\prime }\mathbf{\dot{u}} _{i}^{\ddag }\rightarrow (Q^{-1}\Omega ^{P}Q^{-1})^{1/2}\mathbb{Z}\mathcal{\ \ F}_{0}^{P}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-stably as }N\rightarrow \infty , \end{equation*} where $\mathbb{Z}\backsim N\left( 0,\mathbb{I}_{K_{x}}\right) $ and is independent of $\mathcal{F}_{0}^{P}.$

Theorem (ref) shows that the sequence $\{\sqrt{N}(\hat{{\Greekmath 010C}} ^{MG}-{\Greekmath 010C} ^{0})\}_{N\geq 1}$ converges $\mathcal{F}_{0}^{P}$-stably as $ N\rightarrow \infty $ to a mixture normal distribution whose variance is random but measurable with respect to the limit sigma-field $\mathcal{F} _{0}^{P}$.

The following theorem establishes the consistency of $\hat{\Omega}^{P}.$

theoremSuppose Assumptions (ref)--(ref) hold. Then (i) $\hat{\Omega}^{P}=\frac{1}{N}\sum_{j=1}^{N}\mathbf{\ddot{x}}_{j}^{\prime }\mathbf{\dot{u}}_{j}^{\ddag }\mathbf{\dot{u}}_{j}^{\ddag \prime }\mathbf{ \ddot{x}}_{j}+o_{p}\left( 1\right) =\Omega ^{P}+o_{p}\left( 1\right) ;$ (ii) $(\mathbb{S}^{\prime }\hat{\Omega}^{P}\mathbb{S})^{-1/2}\sqrt{N}\mathbb{ S}^{\prime }(\hat{{\Greekmath 010C}}^{P}-{\Greekmath 010C} ^{0})\overset{d}{\rightarrow }N\left( 0,1\right) .$

Theorem (ref) indicates that we can estimate $\Omega ^{P}$ consistently via the jackknife method. It justifies the use of the normal critical values for asymptotic inference based on the jackknife estimate of the asymptotic variance.

Specification tests for the poolability

In this subsection, we propose a test for the poolability based on the results in the early subsections. In general, $\hat{{\Greekmath 010C}}^{MG}$ is consistent with ${\Greekmath 010C} ^{0}$ no matter whether one has a homogenous panel with a slope coefficient ${\Greekmath 010C} ^{0}$ or a heterogenous panel with slope coefficients whose mean is given by ${\Greekmath 010C} ^{0}$, and it also allows for arbitrary correlations between ${\Greekmath 010C} _{i}$ and $x_{it}$ in a CRC panel. In contrast, $\hat{{\Greekmath 010C}}^{P}$ is generally inconsistent when ${\Greekmath 010C} _{i}$ is correlated with a time-varying component of $x_{it}$ so we allow $\hat{{\Greekmath 010C}} ^{P}$ to converge to something different from ${\Greekmath 010C} ^{0}\ $in this case. For this reason, we consider the following null hypothesis:

equation[equation omitted — 312 chars of source]

The alternative hypothesis is the negation of $\mathbb{H}_{0}:$

equation*[equation* omitted — 282 chars of source]

Below we shall maintain the condition that plim$_{N\rightarrow \infty }\hat{ {\Greekmath 010C}}^{MG}={\Greekmath 010C} ^{0}\ $and assume that the probability limit of $\hat{{\Greekmath 010C} }^{P}$ is given by ${\Greekmath 010C} ^{P}$ under $\mathbb{H}_{1}.$ A failure to reject $ \mathbb{H}_{0}$ justifies the common practice of pooling all cross-sectional units to estimate the slope coefficients in a panel despite the potential presence of slope heterogeneity. If the panel is indeed homogeneous, the common slope ${\Greekmath 010C} ^{0}$ indicates the marginal effect; otherwise, it indicates the population average marginal effect.

Under $\mathbb{H}_{0},$ we have

eqnarray*[eqnarray* omitted — 452 chars of source]

which has a well behaved limit distribution if we assume that Assumptions (ref)--(ref) and (ref)--(ref) are all satisfied. In this case, we can estimate the asymptotic variance of $\sqrt{N} (\hat{{\Greekmath 010C}}^{MG}-\hat{{\Greekmath 010C}}^{P})$ via the jackknife method:

equation*[equation* omitted — 245 chars of source]

where $\hat{\Delta}^{\left( -i\right) }=\hat{{\Greekmath 010C}}^{MG\left( -i\right) }- \hat{{\Greekmath 010C}}^{P\left( -i\right) }$ and $\overline{\hat{\Delta}}=\frac{1}{N} \sum_{i=1}^{N}\hat{\Delta}^{\left( -i\right) }.$ Following the lead of hausman1978specification, we propose the following test statistic:

equation[equation omitted — 207 chars of source]

Below we only assume that Assumptions (ref)--(ref) are always satisfied under $\mathbb{H}_{0}$ and $\mathbb{H}_{1}\ $so that both estimators are well defined. Under $\mathbb{H}_{0},$ we also assume that Assumptions (ref)--(ref) are satisfied so that the pooled estimator $\hat{{\Greekmath 010C}}^{P}$ is consistent with ${\Greekmath 010C} ^{0}$. Under $\mathbb{H }_{1},$ we dispense with Assumptions (ref)--(ref) so that $ \hat{{\Greekmath 010C}}^{P}$ does not converge to ${\Greekmath 010C} ^{0}$ in general.

Let $\mathcal{F}_{N}=\mathcal{{\Greekmath 011B} (}\{\mathbf{\dot{x}}_{i}^{o}\}_{i\leq N},F).$ To study the asymptotic properties of $J_{N}$ under $\mathbb{H}_{0}$ and $\mathbb{H}_{1},$ we add the following two assumptions.

assumption$\frac{1}{N}\sum_{i=1}^{N}E_{\mathcal{F}_{N}}\left\{ \left[ {\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dagger }+{\Greekmath 0111} _{i}-Q^{-1}\mathbf{\ddot{x} }_{i}^{\prime }\mathbf{\dot{u}}_{i}^{\ddag }\right] \left[ {\Greekmath 011F} _{i}^{\prime }\mathbf{u}_{i}^{\dagger }+{\Greekmath 0111} _{i}-Q^{-1}\mathbf{\ddot{x}}_{i}^{\prime } \mathbf{\dot{u}}_{i}^{\ddag }\right] ^{\prime }\right\} \overset{p}{ \rightarrow }\Omega ^{\Delta },$ where $\Omega ^{\Delta }$ is positive definite a.s.
assumption(i) plim$_{N\rightarrow \infty }\hat{{\Greekmath 010C}}^{P}={\Greekmath 010C} ^{P}$ for some ${\Greekmath 010C} ^{P}\neq {\Greekmath 010C} ^{0};$ (ii) $\hat{\Omega}^{\Delta }\overset{p}{\rightarrow }\mathring{\Omega} ^{\Delta },$ where $\mathring{\Omega}^{\Delta }$ is positive definite a.s.

Assumption (ref) specifies a condition to ensure that the asymptotic conditional variance of $\sqrt{N}(\hat{{\Greekmath 010C}}^{MG}$ $-\hat{{\Greekmath 010C}} ^{P})$ given $\mathcal{F}_{N}$ is well behaved under $\mathbb{H}_{0}.$ Assumption (ref)(i) restricts that $\hat{{\Greekmath 010C}}^{P}$ be a consistent estimator of ${\Greekmath 010C} ^{P},$ which is different from ${\Greekmath 010C} ^{0}.$ Assumption (ref)(ii) imposes that $\hat{\Omega}^{\Delta }$ be asymptotically nonsingular. We will use Assumptions (ref) and (ref) for the study of the asymptotic properties of $J_{N}$ under $ \mathbb{H}_{0}\ $and $\mathbb{H}_{1},$ respectively.

The following theorem reports the asymptotic properties of $J_{N}$ under $ \mathbb{H}_{0}\ $and $\mathbb{H}_{1}.$

theoremSuppose that Assumptions (ref)--(ref) hold. (i) If Assumptions (ref)--(ref) hold, then under $\mathbb{H }_{0},$ $J_{N}\overset{d}{\rightarrow }{\Greekmath 011F} ^{2}\left( K_{x}\right) \mathcal{ \ }$as $N\rightarrow \infty .$ (ii) If Assumption (ref) holds, then under $\mathbb{H}_{1},$ $ P\left( J_{N}>c_{N}\right) \rightarrow 1$ for any $c_{N}=o(N)$.

Theorem (ref) indicates that $J_{N}$ has the usual asymptotic ${\Greekmath 011F} ^{2}$-distribution under $\mathbb{H}_{0}$ and it diverges to infinity at rate $N$ under $\mathbb{H}_{1}.\ $As a result, the test is consistent and one can use the critical values from the ${\Greekmath 011F} ^{2}\left( K_{x}\right) $ distribution to make a decision: one rejects the null when $J_{N}>{\Greekmath 011F} _{{\Greekmath 011C} }^{2}\left( K_{x}\right) ,$ the upper ${\Greekmath 011C} $-quantile of the ${\Greekmath 011F} ^{2}\left( K_{x}\right) $ distribution.

Monte Carlo Simulations

In this section, we conduct some Monte Carlo simulations to evaluate the proposed TW-MG\ estimator and jackknife method.

DGPs and implementations

We consider four DGPs where $y_{it}$ and $x_{it}$ follow equations ((ref)) and ((ref)), respectively. DGPs 1--3 are simple DGPs with one regressor $(K_{x}=1)$, covering three different cases of ${\Greekmath 010C} _{i}$, \ to confirm our main theoretical results. In DGP 1, ${\Greekmath 010C} _{i}$ is homogeneous; in DGP 2, ${\Greekmath 010C} _{i}$ is random and independent of $x_{it};$ and in DGP 3, ${\Greekmath 010C} _{i}$ enters the DGP\ of $x_{it}.$ With this design, the pooled-type estimator is consistent for DGPs 1 and 2 and inconsistent for DGP\ 3. In contrast, the MG-type estimator is consistent for all three DGPs. The conventional MG estimator (pesaran_smith1995), without accounting for time fixed effects, is inconsistent for all three DGPs due to the presence of time fixed effects. DGP 4 is more general and realistic with two regressors ($K_{x}=2$), and auto-correlated and heteroskedastic error terms as described below. In DGP 4, only the MG-type estimator is consistent. In all four DGPs, we include both TWFEs and IFEs and also ensure that the key Assumption (ref)(i) holds. In Section (ref) of the online supplement, we report the simulation results for two more DGPs with TWFEs only.

Specifically, in DGPs 1-3, $K_{x}=R=1,$ ${\Greekmath 010C} ^{0}=1.$ Let $({\Greekmath 0115} _{i}^{o},f_{t}^{o},{\Greekmath 010D} _{i}^{o},u_{it},{\Greekmath 0118} _{it},{\Greekmath 0111} _{i}^{\ast },v_{it}^{\ast })$ be generated as $i.i.d.$ $N\left( (\mathbf{{\Greekmath 0113} } _{3}^{\prime },\mathbf{0}_{4\times 1}^{\prime })^{\prime },\mathbb{I} _{7}\right) $ random variables where $\mathbf{{\Greekmath 0113} }_{k}$ and $\mathbf{0} _{k\times 1}$ denote a $k\times 1$ vector of ones and zeros, respectively. We let ${\Greekmath 010B} _{i\cdot }^{o}={\Greekmath 0116} _{i\cdot }^{o}={\Greekmath 0115} _{i}^{o},$ and $ {\Greekmath 010B} _{\cdot t}^{o}={\Greekmath 0116} _{\cdot t}^{o}=f_{t}^{o}.$ Now we specify ${\Greekmath 0111} _{i}$ and $v_{it}$ for DGPs 1-3. In DGP 1, ${\Greekmath 0111} _{i}=0,$ and in DGPs\ 2 and 3, ${\Greekmath 0111} _{i}={\Greekmath 0111} _{i}^{\ast }$. \ In DGPs 1 and 2, $v_{it}=v_{it}^{\ast },$ and in DGP$\ $3, $v_{it}={\Greekmath 010C} _{i}{\Greekmath 0118} _{it}+v_{it}^{\ast }.$

In DGP 4, $K_{x}=2$ and $R=1.$ Let $({\Greekmath 0115} _{i}^{o},f_{t}^{o},{\Greekmath 010D} _{1,i}^{o},{\Greekmath 010D} _{2,i}^{o},u_{it}^{\ast \ast },{\Greekmath 0118} _{it},{\Greekmath 0111} _{1,i}^{\ast },{\Greekmath 0111} _{2,i}^{\ast },v_{1,it}^{\ast },v_{2,it}^{\ast })$ be generated as $ i.i.d.$ $N\left( (\mathbf{{\Greekmath 0113} }_{4}^{\prime },\mathbf{0}_{6\times 1}^{\prime })^{\prime },\mathbb{I}_{10}\right) $ random variables. We let $ {\Greekmath 010B} _{i\cdot }^{o}={\Greekmath 0116} _{1,i\cdot }^{o}={\Greekmath 0116} _{2,i\cdot }^{o}={\Greekmath 0115} _{i}^{o},$ and ${\Greekmath 010B} _{\cdot t}^{o}={\Greekmath 0116} _{1,\cdot t}^{o}={\Greekmath 0116} _{2,\cdot t}^{o}=f_{t}^{o}.$ Let $x_{it}=\left( x_{1,it},x_{2,it}\right) $ which are generated as $x_{1,it}={\Greekmath 0116} _{1,i\cdot }^{o}+{\Greekmath 0116} _{1,\cdot t}^{o}+{\Greekmath 010D} _{1,i}^{o}f_{t}^{o}+v_{1,it}$ and $x_{2,it}={\Greekmath 0116} _{2,i\cdot }^{o}+{\Greekmath 0116} _{2,\cdot t}^{o}+{\Greekmath 010D} _{2,i}^{o}f_{t}^{o}+v_{2,it}$ and let ${\Greekmath 010C} _{i}=({\Greekmath 010C} _{1,i},{\Greekmath 010C} _{2,i})^{\prime }={\Greekmath 010C} ^{0}+({\Greekmath 0111} _{1,i},{\Greekmath 0111} _{2,i})^{\prime },$ where ${\Greekmath 010C} ^{0}=(1,1)^{\prime }.$ For ${\Greekmath 0111} _{1,i},$ $ {\Greekmath 0111} _{2,i},$ $v_{1,it}$ and $v_{2,it},$ we let

equation*[equation* omitted — 472 chars of source]

$u_{it}$ is generated as

equation*[equation* omitted — 192 chars of source]

Note that ${\Greekmath 010C} _{1,i}$ enters the DGP of $x_{1,it}$ via $v_{1,it},$ and $ {\Greekmath 0111} _{2,i}$ (and ${\Greekmath 010C} _{2,i}$) is correlated with the factor loading $ {\Greekmath 010D} _{2,i}^{o}$ in the $x$-equation. Therefore, the pooled-type\ estimators of ${\Greekmath 010C} ^{0}$ are inconsistent.

For each DGP, we compare the six estimators: $(i)$ the TW-MG estimator in ( (ref)), $(ii)$ the TW-MG ridge estimator in ((ref)), $(iii)$ the TW-pooled estimator in ((ref)), $(iv)$ pesaran2006's CCE-pooled estimator, $(v)$ pesaran2006's CCE-MG estimator, and $(vi)$ pesaran_smith1995's standard MG estimator. Table (ref) below summarizes the consistency of the estimators for the four DGPs. So, only the TW-MG, TW-MG ridge, and CCE-MG estimators are consistent for all four DGPs. For the CCE-MG estimator, the consistency results are merely our conjecture, and we are not aware of any rigorous studies on the CCE-MG estimator for heterogeneous models with fixed $T.$\footnote{ We conjecture that theoretical analysis can be challenging. For example, we will encounter the issue of the second moment matrix of the estimated factors becoming asymptotically singular when the number of proxies exceeds the number of factors, as pointed by karabiyik2017 for the large $T$ case (see, also andrews1987). In addition, it is unclear how to conduct valid inference for the CCE-MG estimators when $T$ is fixed.} Even so, we observe below that our TW-MG and TW-MG ridge estimators substantially outperform the CCE-MG estimator when $T$ is small. westerlund2022cce study the CCE-pooled estimator for heterogeneous models for fixed $T$, which requires that the slope coefficients be uncorrelated with the regressors. The choice of tuning parameter for the TW-MG ridge estimator is discussed in Remark (ref). For the CCE-pooled and CCE-MG estimators, we also include the individual fixed effects.\footnote{ If we do not include the individual effects, the performance is worse in general for these four DGPs.} We consider different combinations of $\left( N,T\right) $ as shown in the tables below. For DGPs 1-3 where $K_{x}=1,$ $T$ starts with $\,T=3.$ For DGP 4 where $K_{x}=2,$ $T$ starts with $T=4.$ The number of replications is 250.

table[table omitted — 1,324 chars of source]

Simulation results

Table (ref) presents the performance of the six estimators discussed above for the four DGPs. For each estimator, we report its bias$ \times 10$ and mean-squared-error\ (MSE)$\times 100$. To save space, the variance is not reported but available upon request. For DGP 4 where there are two regressors with ${\Greekmath 010C} ^{0}=({\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0})^{\prime },$ we only report the results for ${\Greekmath 010C} _{1}^{0}$ to save space, as the results for ${\Greekmath 010C} _{2}^{0}$ are similar. Among the four DGPs, the standard MG estimator is inconsistent, so we mainly compare the other five estimators. For DGP 1 where ${\Greekmath 010C} _{i}$ is homogeneous, all five estimators are consistent. However, when $T$ is small ($T=3$), the TW-pooled estimator performs best, as it is based on the most parsimonious model. The second-best estimator is the TW-MG ridge estimator, followed by the TW-MG estimator, CCE-pooled and CCE-MG estimators. The difference in the performance is substantial. For example, when $N=50$ and $T=3,$ the MSEs$ \times 100$ of the TW-pooled, TW-MG ridge, TW-MG, CCE-pooled and CCE-MG estimators are 1.31, 3.02, 7.35, 16.60 and 34.33, respectively. However when $T$ is large ($T=10),$ the performances of the five estimators are all good and similar. For DGP 2 where ${\Greekmath 010C} _{i}$ is heterogeneous but independent of $x_{it},$ our TW-MG ridge estimator is the best when $T$ is small, by achieving a smaller variance. When $T$ is large, the performances of the TW-MG ridge, TW-MG and CCE-MG estimators are similar, and they outperform the CCE-pooled and TW-pooled estimators. For DGPs 3 and 4, the pooled estimators are inconsistent for the reason discussed above. It is clear that its bias of the TW-pooled and CCE-pooled estimators does not converge to zero when the sample size increases. In contrast, the bias of the TW-MG, TW-MG ridge and CCE-MG estimators is all small. Again, when $T$ is small, the TW-MG ridge estimator outperforms the TW-MG estimator, which in turn outperforms the CCE-MG estimators. When $T$ is large, the three estimators all perform well.

Table (ref) shows the empirical coverage of the 95% CI in ((ref)) for ${\Greekmath 010C} ^{0}$ in DGPs 1-3 and ${\Greekmath 010C} _{1}^{0}$ in DGP\ 4 based on the leave-one-out jackknife inference method proposed in Section (ref) . We consider both TW-MG and TW-MG ridge estimators. The actual coverages are all close to the nominal 95% for all sample sizes and DGPs. This confirms the asymptotic validity of the leave-one-out jackknife method. Table (ref) reports the rejection frequency of testing $\mathbb{ H}_{0}$ in ((ref)) with a nominal size of $5\%$. We consider two types of tests, one based on the TW-MG estimator and the other based on the TW-MG ridge estimators. As discussed above, in DGPs 1 and 2, both the TW-MG (TW-MG ridge) and pooled estimators are consistent, so these two DGPs are under the null. Under the null, the rejection frequency of both tests is close to $5\%$ for all $N$ and $T$ considered. For DGPs 3 and 4 which are under the alternative, the rejection frequency of both tests increases with $N$ and $ T, $ demonstrating the power of the tests. When $N$ is large, such as $N=$ $ 200, $ the rejection frequency of both tests is close to $100\%$. We also observe that when the sample size is small (e.g., $N=50$ and $T=3$), the test based on the TW-MG ridge estimator is more powerful than that based on the TW-MG estimator, as expected. {2.0pt}

table[table omitted — 4,697 chars of source]
table[table omitted — 1,621 chars of source]
table[table omitted — 1,970 chars of source]

Empirical Applications

In this section, we apply our new TW-MG estimator and jackknife method to two datasets. The first application examines the relationship between health-care expenditure and income, while the second estimates a production function. In the first application, we reject the null of poolability, whereas in the second, we fail to reject it.

Application 1: Health-care expenditure and income

In this application, we use the data from baltagi2017health and kapetanios2023testing to investigate the relationship between health-care expenditure and income for a panel of 167 countries covering the period 1995--2012 $(N=167$ and $T=18).$ Here $y_{it}$ is the per-capita health-care spending in the $i$th country at year $t.$ There are two regressors $\left( K_{x}=2\right) $. The key regressor of interest is the income level measured by per capita GDP. The second regressor is the rate of public expenditure over total health expenditure as a control variable. For a more detailed description of these two variables, see baltagi2017health.

Let ${\Greekmath 010C} ^{0}=({\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0})^{\prime },$ where ${\Greekmath 010C} _{1}^{0}$ and ${\Greekmath 010C} _{2}^{0}$ represent the expected values of slope coefficients of income and public expenditure rate, respectively. Table (ref) reports the estimation results with four estimation methods, namely the TW-MG, TW-MG ridge, TW-pooled, and MG estimators. We find that TW-MG and TW-MG ridge estimates are almost identical, which are substantially different from the TW-pooled and MG estimates. Therefore, our discussion below focuses on the TW-MG, TW-pooled and MG estimates. For $ {\Greekmath 010C} _{1}^{0},$ the three estimates are 0.94, 0.72, and 1.87, respectively. For ${\Greekmath 010C} _{2}^{0},$ the three estimates are 0.30, 0.02 and 0.76, respectively. If we treat our TW-MG estimator as a benchmark, the TW-pooled estimator tends to under-estimate, while the standard MG estimator over-estimates. For ${\Greekmath 010C} _{1}^{0},$ all three estimation methods indicate a positive and statistically significant average effect of income, though the magnitudes differ. For ${\Greekmath 010C} _{2}^{0},$ while all estimation methods show a positive effect, only the TW-MG and MG estimates are statistically significant at the 5% level; the TW-pooled estimate is not. Figure (ref) plots the histogram of the estimates of ${\Greekmath 010C} _{i}=({\Greekmath 010C} _{1i},{\Greekmath 010C} _{2i})^{\prime }$ based on ((ref)). Although these estimators are not consistent in the fixed $T$ framework, we observe substantial variations, suggesting the heterogeneity. Notably, most estimates of ${\Greekmath 010C} _{1i}$ are positive, demonstrating a robust positive effect of income.

We conduct the test for the hypotheses of poolability introduced in Section (ref). In addition to testing poolability in estimating both parameters in (${\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0}$) together, we also test the poolability for each individual parameter. Specifically, let $\hat{{\Greekmath 010C}} ^{P}=(\hat{{\Greekmath 010C}}_{1}^{P},\hat{{\Greekmath 010C}}_{2}^{P})^{\prime }$ and $\hat{{\Greekmath 010C}} ^{MG}=(\hat{{\Greekmath 010C}}_{1}^{MG},\hat{{\Greekmath 010C}}_{2}^{MG})^{\prime },$ and test the following three hypotheses:

equation[equation omitted — 744 chars of source]

As shown in Table (ref), we reject $\mathbb{H}_{01},$ $\mathbb{H} _{02}$ and $\mathbb{H}_{0}$ all with $p$-values of 0.027, 0.019 and 0.009, respectively, based on the TW-MG estimator and the jackknife method. We also report the step-down Holm adjusted $p$-values (holm1979simple) accounting for multiple testing, which are all smaller than 0.05.\footnote{ Suppose that we have conducted the tests for $K^{\ }$individual hypotheses. We order the individual p-values from the smallest to the largest as $p_{\left( 1\right) }\leq p_{\left( 2\right) }\leq ...\leq p_{\left( K\right) }$ with their corresponding null hypotheses labeled accordingly as $ \mathbb{H}_{0\left( 1\right) },$ $\mathbb{H}_{0\left( 2\right) },...,$ $ \mathbb{H}_{0\left( K\right) }.$ The step-down Holm adjusted $p$-values for testing $\mathbb{H}_{0\left( k\right) }$ is calculated as: adjusted-$ p_{\left( k\right) }=\min \left( \left( K-k+1\right) p_{\left( k\right) },1\right) .$ } Therefore, we conclude that we cannot pool the data to estimate the average effects, confirming the substantial differences in the estimation results discussed above.

table[table omitted — 1,519 chars of source]
table[table omitted — 830 chars of source]

{

figure[figure omitted — 265 chars of source]

}

Application 2: Production function

In this subsection, we apply our new methods to estimate a production function using a balanced panel from blundell2000gmm. The dataset consists of 508 R&D-performing U.S. manufacturing companies for 8 years, 1982--1988, so $N=508$ and $T=8.$\footnote{ The original sample consists of 509 firms. However, we drop one firm because its labor input (the first regressor) remains unchanged over the sample period.} Here, the dependent variable $y_{it}$ is the logarithms of the sales of firm $i$ in year $t$ and there are two regressors $(K_{x}=2)$, the\ logarithms of its labor (employment) and the logarithm of its capital stock, which are inputs to the production function. For the detailed definitions of the variables, see blundell2000gmm and mairessehall.

Table (ref) reports the estimation results based on the four estimation methods. Let ${\Greekmath 010C} ^{0}=({\Greekmath 010C} _{1}^{0},{\Greekmath 010C} _{2}^{0})^{\prime },$ where ${\Greekmath 010C} _{1}^{0}$ and ${\Greekmath 010C} _{2}^{0}$ represent the average return to labor and capital, respectively. For ${\Greekmath 010C} _{1}^{0},$ the TW-MG, TW-MG ridge, TW-pooled and MG estimates are similar, with the values of 0.67, 0.67, 0.65 and 0.63, respectively, all of which are statistically significant. For ${\Greekmath 010C} _{2}^{0},$ the TW-MG, TW-MG ridge, and TW-pooled estimates are similar with the values of 0.21, 0.21, and 0.23, respectively, which are smaller than the MG estimate, 0.32. Again, all four estimates are statistically significant. Figure (ref) plots the histogram of the estimates of ${\Greekmath 010C} _{i}=({\Greekmath 010C} _{1i},{\Greekmath 010C} _{2i})^{\prime }$ based on ((ref)), where we also observe some variations, though most of them are positive

Similar to the first application in Section (ref), we conduct the test for poolability. We test the three hypotheses as in ((ref)). As shown in Table (ref), we fail to reject $\mathbb{H}_{01},$ $\mathbb{H} _{02},$ and $\mathbb{H}_{0}$ all with $p$-values of 0.56, 0.67, and 0.84, respectively. Therefore, for this application, we do not have strong evidence against poolability.

table[table omitted — 1,503 chars of source]
table[table omitted — 819 chars of source]

{

figure[figure omitted — 258 chars of source]

}

Conclusion

This paper considers estimation and inference for the population average of slope coefficients in heterogeneous panels with two-way additive fixed effects (TWFEs) and interactive fixed effects (IFEs). We propose two methods to estimate the parameter of interest. The first method extends the mean group (MG) approach to accommodate both individual and time effects, which is valid for general correlated random coefficient (CRC) models but imposes stronger rank conditions. The second method pools all observations to estimate the parameter of interest, which may not be valid for some CRC models but imposes weaker rank conditions. For both estimation methods, we establish stable convergence for the estimators and propose a jackknife method to estimate their asymptotic variance-covariance matrices. In addition, we propose a Hausman-type test to assess poolability, which may justify the commonly used pooled estimator in empirical research in the presence of slope heterogeneity. Simulations demonstrate that our methods perform well in finite samples.

There are many interesting topics for further research. First, we may extend our methods to dynamic models and allow for weakly exogenous regressors. Second, we may address the endogeneity issue in heterogeneous panels by extending our methods. Third, we may consider nonlinear heterogeneous panels with short $T$ and TWFEs and/or IFEs. These extensions are highly challenging, but worth exploring.

{ \setstretch{0.75} {5pt} }