The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
72,421 characters
Model Selection in Panel Data Models: A Generalization of the Vuong Test
\title{Model Selection in Panel Data Models: A Generalization of the Vuong
Test}
\author{Jinyong Hahn\thanks{
Department of Economics, UCLA, Los Angeles, CA 90095, USA. Email:
[email removed].} \\
UCLA \and Zhipeng Liao\thanks{
Department of Economics, UCLA, Los Angeles, CA 90095, USA. Email:
[email removed].} \\
UCLA \and Konrad Menzel\thanks{
Department of Economics, NYU, New York, NY 10012, USA. Email: [email removed].}
\\
NYU \and Quang Vuong\thanks{
Department of Economics, NYU, New York, NY 10012, USA. Email: [email removed].
} \\
NYU}
\date{\today }
\maketitle
\begin{abstract}
This paper generalizes the classical \cite{Vuong1989} test to panel data
models by employing modified profile likelihoods and the Kullback--Leibler
information criterion. Unlike the standard likelihood function, the profile
likelihood lacks certain regular properties, making modification necessary.
We adopt a generalized panel data framework that incorporates group fixed
effects for time and individual pairs, rather than traditional individual
fixed effects. Applications of our approach include linear models with
non-nested specifications of individual-time effects.
\bigskip
\noindent JEL Classification: C14, C31, C32\bigskip
\noindent\textit{Keywords: } Panel data models, Vuong test, Bias correction,
Grouped heterogeneity.
\end{abstract}
\section{Introduction\label{sec:intro}}
With the advent of big data, modern empirical research increasingly adopts
more complex models. These models often involve high-dimensional nuisance
parameters, which can introduce substantial biases in standard estimators -
manifesting as variations of the incidental parameters problem. Classical
references include \cite{NeymanScott1948} and \cite{Nickell1981} among
others. Numerous studies have been conducted to address and mitigate these
issues. It is now well established that these issues represent a form of
higher-order bias, and a variety of methodological approaches have been
proposed to overcome them. See \cite{ArellanoHahn2007}, \cite
{Arellano-Bonhomme}, \cite{BesterHansen2009}, \cite{CARRO2007503}, \cite
{Dhaene-Jochmans}, \cite{FERNANDEZVAL200971}, \cite{Fernandez-Val-Weidner-1}
, \cite{Fernandez-Val-Weidner-2}, \cite{HahnKuersteiner2002}, and \cite
{HahnNewey2004}, for example.
Although the literature is well developed, its generalizations have only
recently found application in substantively meaningful but technically
complex contexts, such as network-type models. See \cite
{Bonhomme-Lamadon-Manresa-1}, \cite{Jochmans-Weidner}, \cite
{bonhomme2025neymanorthogonalizationapproachincidentalparameter}. Motivated
by these developments, we propose to investigate issues of model
specification in nonlinear panel data models. Like most econometric
frameworks, panel data models often rely on potentially restrictive
assumptions. However, to the best of our knowledge, there is currently no
literature that provides a systematic approach for comparing potentially
non-nested models in this context. We would like to close this gap in the
literature by proposing a panel generalization of the classical \cite
{Vuong1989} test.
Specifically, we consider a model characterized by the joint log-likelihood
of the data $\{z_{i,t}\}_{i\leq n,t\leq T}$ specified as:
\begin{equation}
L_{n,T}(\Greekmath 011E )\equiv \sum_{i\leq n}\sum_{t\leq T}\log f\left( z_{i,t};\Greekmath 0112
,\Greekmath 010D _{g(i),m(t)}\right) , \label{general_spec}
\end{equation}
where $f\left( \cdot \right) $ is a known function, $\Greekmath 0112 $ represents
unknown parameters that are individual/time-invariant, $\Greekmath 010D _{g(i),m(t)}$
denotes the individual and time fixed effect which has a grouping structure
determined by cluster/group assignment functions $g(\cdot )$ and $m(\cdot )$
with ranges $\mathcal{G}\equiv \{1,\ldots ,G\}$ and $\mathcal{M}\equiv
\{1,\ldots ,M\}$, respectively, and $\Greekmath 011E \equiv (\Greekmath 0112 ^{\top },(\Greekmath 010D
_{g,m}^{\top })_{g\in \mathcal{G},m\in \mathcal{M}})^{\top }$ include all
unknown parameters specified in the model. For ease of notations, we let $
I_{g}$ for $g\in \mathcal{G}$ and $I_{m}$ for $m\in \mathcal{M}$ denote the
partitions of $\{1,\ldots ,n\}$ and $\{1,\ldots ,T\}$ according to the group
assignment functions $g(\cdot )$ and $m(\cdot )$, respectively. The pseduo
true value $\Greekmath 011E ^{\ast }$ is defined as the maximizer of $\mathbb{E}\left[
L_{n,T}(\Greekmath 011E )\right] $ over the parameter space, and it is often estimated
by the maximizer $\hat{\Greekmath 011E }$ of $L_{n,T}(\Greekmath 011E )$ over the same space. Our
group structure is inspired the analyses by \cite{Bonhomme-Manresa} or \cite
{Bonhomme-Lamadon-Manresa-2}, yet it is general enough to include the
classical panel models with individual fixed effects.
In practice, different ways of specifying the likelihood function $f\left(
\cdot \right) $ and the group assignment functions $g(\cdot )$ and $m(\cdot
) $ lead to different modeling strategies in panel data models, and we often
have the situation to compare competing modeling strategies and hope to find
the one which is closer (or the closest) to the true data generating process
in terms of the Kullback--Leibler (KL) distance. It would be reasonable to
compare models through the \cite{Vuong1989} test, which tests the null
hypothesis:
\begin{equation*}
H_{0}:\mathbb{E}\left[ L_{1,n,T}(\Greekmath 011E _{1}^{\ast })\right] =\mathbb{E}\left[
L_{2,n,T}(\Greekmath 011E _{2}^{\ast })\right] ,
\end{equation*}
where $L_{j,n,T}(\cdot )$ represents the joint log-likelihood specified by
modeling strategy $j$ with $\Greekmath 011E _{j}^{\ast }$ denoting the pseudo true
value for $j=1,2$. The alternative hypothesis can be
\begin{equation*}
H_{1}^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{2-sided}}:\mathbb{E}\left[ L_{1,n,T}(\Greekmath 011E _{1}^{\ast })\right]
\neq \mathbb{E}\left[ L_{2,n,T}(\Greekmath 011E _{2}^{\ast })\right] \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ or \ }
H_{1}^{\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{1-sided}}:\mathbb{E}\left[ L_{1,n,T}(\Greekmath 011E _{1}^{\ast })\right] >
\mathbb{E}\left[ L_{2,n,T}(\Greekmath 011E _{2}^{\ast })\right] .
\end{equation*}
The classical Vuong test employs the quasi-likelihood ratio (QLR), defined
as:
\begin{equation}
QLR_{n,T}\equiv (nT)^{-1/2}(L_{1,n,T}(\hat{\Greekmath 011E }_{1})-L_{2,n,T}(\hat{\Greekmath 011E }
_{2})), \label{G_QLR}
\end{equation}
to construct the test statistic, where $\hat{\Greekmath 011E }_{j}$ denotes the
estimator of $\Greekmath 011E _{j}^{\ast }$ in model $j$. It turns out that in panel
models, the log-likelihoods $L_{1,n,T}$ and $L_{2,n,T}$ are biased due to
the incidental parameter problems. Panel adaptation of the Vuong test thus
requires working with bias-corrected versions of the log likelihoods, which
we will call $LM_{1,n,T}^{\ast }$ and $LM_{2,n,T}^{\ast }$. We also show
that there is additional complication because the order of magnitude of the
standard error of the test statistic (after such modification) can depend on
the nature of the hypothesis. Our technical contribution consists of
estimation of $LM_{1,n,T}^{\ast }$ and $LM_{2,n,T}^{\ast }$ to make the
modification feasible as well as estimation of the standard error while
avoiding the uniformity issue.
We extend the classical \cite{Vuong1989} test to panel models, contributing
to the understanding of model selection complexities in high-dimensional
settings. The Vuong test originally compares two non-nested parametric
models and formally tests the hypothesis that their KL distances from the
true data distribution are equal. Originally designed for parsimoniously
parameterized models, this test has been expanded in various studies, see,
\cite{RiversVuong2002}, \cite{ChenHongShum}, and \cite{Shi2015b}, for
example. In this paper, we demonstrate that the incidental parameters
problem in panel data contexts invalidates the classical Vuong test, which
compares maximized likelihood ratios between two models. We present a
modified test statistic that ensures validity for panel data analysis.
Our modification follows a precedent in \cite{LeePhillips2015} and \cite
{Liao&Shi2020}. \cite{LeePhillips2015} explored model selection in panel
data settings using nested hypothesis testing. Our work complements theirs
by introducing approaches for testing non-nested hypotheses. \cite
{Liao&Shi2020} examined the Vuong test for semiparametric models in a
cross-sectional setting, whereas we extend the analysis to a panel data
context with fixed effects. Our paper and \cite{LeePhillips2015} share the
common feature in the sense that both analyses utilize application of
profiling to the Kullback--Leibler information criterion (KLIC). It is
well-known that profile likelihood does not share the standard properties of
the genuine likelihood function. Modification of profile likelihood was
discussed a way of overcoming this problem. See \cite{ArellanoHahn2007},
\cite{ArellanoHahn2016}, \cite{CoxReid1987}, \cite{DiCiccioStern1993}, \cite
{DiCiccioMartinSternYoung}, \cite{CoxFergusonReid}, and \cite{PaceSalvan},
for example. Following this tradition, our paper as well as \cite
{LeePhillips2015} propose a modified profile likelihood that provides
approximations to the standard KLIC, and provided a valid method of model
selection in panel data models with individual fixed effects. Interestingly,
we find some close connection between the modification in \cite{Liao&Shi2020}
and the panel type modification following \cite{ArellanoHahn2007}, \cite
{ArellanoHahn2016}, and \cite{LeePhillips2015}.
Recently developed methods in panel data allow applied researchers to choose
among a broad range of alternative models for separable, interactive, or
discretely supported unobserved heterogeneity across units and time. It is
widely understood that parameter estimates are often sensitive to the chosen
specification, however there is often no clear theoretical guidance on how
to choose among alternative models. Our paper makes further contribution in
this regard by considering the cases where the definition of ``individuals''
may not be agreed across different models. In empirical practice, the
group/cluster structure may not be obviously agreed among researchers. The
competing group/cluster structures may be nested in the sense that one
structure is ``finer'' than the other one, but they may be non-nested. In
this situation, extension of the insights in previous papers such as \cite
{LeePhillips2015} or \cite{Liao&Shi2020} may not be straightforward. We
establish that the classical Vuong test is valid after some modification. In
doing so, we also make a technical contribution that may be of some
independent interest. \cite{FernandezValWeidener2016} analyzed the
asymptotic distribution of the (pseudo) MLE in nonlinear panel data models
with individual and time effects. Under the assumption that the objective
function is globally concave in all the parameters, they established the
asymptotic bias formula. Our paper contains a result\footnote{
See Theorem \ref{Rate_gamma} in the Online Appendix and Theorem \ref
{Rep_Theta} in Appendix \ref{Sec:AP2}, along with the related discussion
therein, for further details.} that relaxes this concavity assumption,
although it is achieved at the cost of restricting the group structure.
Therefore, our result can be argued to complement \cite
{FernandezValWeidener2016}.
The remainder of the paper is organized as follows. In Section \ref{Sec:
Simple-Vuong},\ we introduce the Vuong test for comparing classical panel
data models that allow for potential group heterogeneity. Section \ref
{sec:TWE} presents a test for comparing the two-way fixed effects (TWFE)
model against a heterogeneous time fixed effects model. Section \ref
{sec:conclusion} concludes.\ The proofs of the main results, along with
auxiliary lemmas, are provided in the Appendix. Additionally, Appendix \ref
{Sec:AP2} develops a general asymptotic theory for estimators in panel
models with individual and time fixed effects, as well as the asymptotic
distribution of the statistic $QLR_{n,T}$ in this broader setting -- results
that are of independent interest.\footnote{
The technical proofs for this general theory, together with the proofs of
the auxiliary lemmas used to establish the main results in Sections \ref
{Sec: Simple-Vuong} and \ref{sec:TWE}, are provided in the Online Appendix.}
The following notation will be adopted throughout the paper.\ We use $K$ to
denote a generic strictly positive constant that may vary from one instance
to another but remains independent of the panel dimensions $n$ {and }$T$.\
We adopt the convention that a summation over an empty set equals zero. We
use $a\equiv b$ to indicate that $a$ is defined as $b$. For real numbers $
a_{1},\ldots ,a_{m}$, $(a_{j})_{j\leq m}\equiv (a_{1},\ldots ,a_{m})^{\top }$
. For any matrix $A$, $A^{\top }$\ denotes the transpose of $A$,\ and $
\left\Vert A\right\Vert $ denotes the Euclidean norm of $A$. For any doubly
indexed sequence $a_{i,t}$ (where $i=1,\ldots ,n$ and $t=1,\ldots ,T$), we
define\ $\bar{a}_{i}\equiv T^{-1}\sum_{t\leq T}a_{i,t}$, $\bar{a}_{t}\equiv
n^{-1}\sum_{i\leq n}a_{i,t}$ and $\bar{a}\equiv (nT)^{-1}\sum_{t\leq
T}\sum_{i\leq n}a_{i,t}$. The summation $\sum_{i^{\prime }\neq i}$ is taken
over all $i^{\prime }$ except $i$, which means $\sum_{i^{\prime }\neq
i}a_{i^{\prime },t}=\sum_{i^{\prime }=1}^{i-1}a_{i^{\prime
},t}+\sum_{i^{\prime }=i+1}^{n}a_{i^{\prime },t}$.\ For two sequences of
positive numbers $a_{n}$ and $b_{n}$, we write $a_{n}\succ b_{n}$ if $
a_{n}\geq c_{n}b_{n}$ for some strictly positive sequence $c_{n}\rightarrow
\infty $. The notation $\left\Vert \cdot \right\Vert _{p}$ denotes the $
L_{p} $-norm. Finally, for any function $g(\cdot )$, define $\mathbb{E}_{T}
\left[ g\left( z_{i,t}\right) \right] \equiv T^{-1}\sum_{t\leq T}\mathbb{E}
\left[ g\left( z_{i,t}\right) \right] $ and $\widehat{\mathbb{E}}_{T}\left[
g\left( z_{i,t}\right) \right] \equiv T^{-1}\sum_{t\leq T}g\left(
z_{i,t}\right) $.
\section{Vuong Test for Classical Panel Models\ \label{Sec: Simple-Vuong}}
In this section, we investigate the Vuong test for panel models with
time-invariant individual effects.\footnote{
A simple version of the test appeared in \cite{hahn-liu}, which later became
a chapter in Liu's UCLA dissertation. Our current paper substantially
extends and generalizes the earlier simple result.} With a dataset\ $\left\{
z_{i,t}\right\} _{i\leq n,t\leq T}$, we compare two models, denoted as model
1 and model 2 respectively. Model 1 represents the classical panel model
with individual effects. In contrast, model 2 may specify a different joint
likelihood function and/or adopt a distinct clustering or grouping structure
for the individual effects.
For each model $j$ ($j=1,2$), the joint log-likelihood function simplifies
from the general specification in (\ref{general_spec}) to the following
form:
\begin{equation*}
L_{j,n,T}(\Greekmath 011E _{j})\equiv \sum_{i\leq n}\sum_{t\leq T}\log
f_{j}(z_{i,t};\Greekmath 0112 _{j},\Greekmath 010D _{g_{j}(i)})=\sum_{g\in \mathcal{G}
_{j}}\sum_{i\in I_{j,g}}\sum_{t\leq T}\Greekmath 0120 _{j}\left( z_{i,t};\Greekmath 011E
_{j,g}\right) ,
\end{equation*}
where\ $\Greekmath 011E _{j}\equiv (\Greekmath 0112 _{j}^{\top },(\Greekmath 010D _{j,g})_{g\in \mathcal{G
}_{j}}^{\top })^{\top }$, $\Greekmath 0120 _{j}\left( z_{i,t};\Greekmath 011E _{j,g}\right) \equiv
\log f_{j}(z_{i,t};\Greekmath 0112 _{j},\Greekmath 010D _{j,g})$, $\Greekmath 011E _{j,g}\equiv (\Greekmath 0112
_{j}^{\top },\Greekmath 010D _{j,g})^{\top }$,\ $\mathcal{G}_{j}\equiv \{1,\ldots
,G_{j}\}$ defines the range of the grouping function $g_{j}(\cdot )$, and $
I_{j,g}$ ($g\in \mathcal{G}_{j}$) represents the partition of $\{1,\ldots
,n\}$ induced by $g_{j}(\cdot )$. For model 1, it follows that $\mathcal{G}
_{1}=\{1,\ldots ,n\}$, implying that each individual has its own group. For
model 2, however, the partitions $\{I_{2,g}\}_{g\in \mathcal{G}_{2}}$ may
form any arbitrary grouping of $\{1,\ldots ,n\}$\ allowing for greater
flexibility in the specification of individual effects.\footnote{
This means that model 1 is nested within model 2 in terms of the grouping
structure, although the likelihood specifications may differ. In many
applications, the grouping structure is expected to be identical across
models.}
As this section is relatively long, we provide a roadmap to facilitate
reading. Sections \ref{Sec: Incid_Para} - \ref{section:lm-variance} provides
the theoretical foundation of the infeasible test, while Sections \ref
{section-biasestimation} - \ref{section:varianceestimation} presents the
actual algorithm for implementation of the feasible version of the test. In
Section \ref{Sec: Incid_Para}, we provide an intuitive overview of the
adjustment to the maximized joint likelihood and discuss its importance for
model comparison using the Vuong test in the panel data models considered in
this paper. In Section \ref{section:log-likelihoods}, we present the log
likelihood of each model. In Section \ref{section:lm-bias}, we present an
infeasible adjustment that corrects for the bias in the likelihood, and in
Section \ref{section:lm-variance}, we present a theoretical analysis for the
asymptotic variance of the modified likelihood. Asymptotic properties of the
(infeasible) modified test statistic is presented in Theorem ~\ref{C_L0}. In
Sections \ref{section-biasestimation} and \ref{section:varianceestimation},
we present actual algorithms for estimation of the modified log likelihoods
and the asymptotic variances. Section \ref{section:feasiblestatistic}
presents Theorem~\ref{C_T1} that characterizes the asymptotic properties of
the feasible test statistic.
\subsection{Incidental Parameter Problem in Panel Data Models \label{Sec:
Incid_Para}}
To streamline notation and improve clarity, we focus here on a simplified
version of (\ref{general_spec}) that includes only individual fixed effects,
specified as:
\begin{equation}
L_{n,T}(\Greekmath 011E )\equiv \sum_{i\leq n}\sum_{t\leq T}\log f\left( z_{i,t};\Greekmath 0112
,\Greekmath 010D _{i}\right) =\sum_{i\leq n}\sum_{t\leq T}\Greekmath 0120 \left( z_{i,t};\Greekmath 011E
_{i}\right) , \label{Fixed_Eff_Model}
\end{equation}
where $\Greekmath 0120 \left( z_{i,t};\Greekmath 011E _{i}\right) \equiv \log f\left(
z_{i,t};\Greekmath 0112 ,\Greekmath 010D _{i}\right) $, $\Greekmath 011E _{i}\equiv (\Greekmath 0112 ^{\top
},\Greekmath 010D _{i})^{\top }$\ and $\Greekmath 011E \equiv (\Greekmath 0112 ^{\top },(\Greekmath 010D
_{i})_{i\leq n}^{\top })^{\top }$.\ Since QLR statistic defined in (\ref
{G_QLR}) involves the maximized joint likelihood\ $L_{n,T}(\hat{\Greekmath 011E })$,
understanding the QLR's properties requires investigating the estimation
error of $L_{n,T}(\hat{\Greekmath 011E })$.
Let $\Greekmath 011E ^{\ast }\equiv (\Greekmath 0112 ^{\ast \top },(\Greekmath 010D _{i}^{\ast })_{i\leq
n}^{\top })^{\top }$ denote the pseudo-true parameter which maximizes $
\mathbb{E}\left[ L_{n,T}(\Greekmath 011E )\right] $.\footnote{
Potential dependence of the pseudo-true parameters $\Greekmath 011E ^{\ast }$ on the
sample size is not made explicit for notational simplicity throughout the
paper.} Define:
\begin{equation}
\Greekmath 0120 \left( z_{i,t}\right) \equiv \Greekmath 0120 \left( z_{i,t};\Greekmath 011E _{i}^{\ast
}\right) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, \ }\Greekmath 0120 _{\Greekmath 010D }\left( z_{i,t}\right) \equiv \partial
\Greekmath 0120 \left( z_{i,t};\Greekmath 011E _{i}^{\ast }\right) /\partial \Greekmath 010D _{i}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \
and \ }\Greekmath 0120 _{\Greekmath 010D \Greekmath 010D }\left( z_{i,t}\right) \equiv \partial ^{2}\Greekmath 0120
\left( z_{i,t};\Greekmath 011E _{i}^{\ast }\right) /\partial \Greekmath 010D _{i}^{2},
\label{derv_1}
\end{equation}
where $\Greekmath 011E _{i}^{\ast }\equiv (\Greekmath 0112 ^{\ast \top },\Greekmath 010D _{i}^{\ast
})^{\top }$. It can be shown\footnote{
See Theorem \ref{Rep_L} in the Appendix, specialized to the case $\mathcal{G}
=\{1,\ldots ,n\}$ and $\mathcal{M}=\{1\}$, so that $G=n$ and $M=1$}\ that
under the asymptotics with $nT^{-1}\rightarrow \Greekmath 011A $\ where $\Greekmath 011A \in
(0,\infty )$:
\begin{equation}
L_{n,T}(\hat{\Greekmath 011E })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right]
=L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right]
+2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}+o_{p}(n^{1/2}),
\label{L_nT_exp}
\end{equation}
where $L_{n,T}(\Greekmath 011E ^{\ast })=\sum_{i\leq n}\sum_{t\leq T}\Greekmath 0120 \left(
z_{i,t}\right) $, and for any $i\leq n$,
\begin{equation}
\tilde{\Psi}_{\Greekmath 010D ,i}\equiv T^{-1/2}\sum_{t\leq T}\tilde{\Greekmath 0120 }_{\Greekmath 010D
}^{\ast }\left( z_{i,t}\right) ,\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ and \ \ }\tilde{\Greekmath 0120 }_{\Greekmath 010D
}^{\ast }\left( z_{i,t}\right) \equiv \frac{\Greekmath 0120 _{\Greekmath 010D }\left(
z_{i,t}\right) -\mathbb{E}\left[ \Greekmath 0120 _{\Greekmath 010D }\left( z_{i,t}\right) \right]
}{(\mathbb{E}_{T}\left[ -\Greekmath 0120 _{\Greekmath 010D \Greekmath 010D }\left( z_{i,t}\right) \right]
)^{1/2}}. \label{derv_2}
\end{equation}
The expansion in (\ref{L_nT_exp}) shows that the estimation error of $
L_{n,T}(\hat{\Greekmath 011E })$ has two leading components. The first component, $
L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $,
results from estimating\ $\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $
at the given pseudo-true parameter $\Greekmath 011E ^{\ast }$. When scaled by a factor
of $(nT)^{-1/2}$, this component converges in distribution to a normal
distribution under some regularity conditions. The second component, $
2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$, arises from estimating
the unknown pseudo-true parameter $\Greekmath 011E ^{\ast }$ and is determined by the
estimation error of the incidental parameters $(\Greekmath 010D _{i}^{\ast })_{i\leq
n}$.\footnote{
Since the dimension of $\Greekmath 0112 ^{\ast }$ is fixed, the contribution of its
estimation error to $L_{n,T}(\hat{\Greekmath 011E })$ is asymptotically negligible
compared with the two leading components, and is therefore included in the
remainder term of (\ref{L_nT_exp}), represented as $o_{p}(n^{1/2})$.}
From the expansion in (\ref{L_nT_exp}), it is evident that in panel data
models, the $L_{n,T}(\hat{\Greekmath 011E })$ has a first-order bias, stemming from $
2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$. To evaluate this bias
explicitly, suppose that $\{\tilde{\Greekmath 0120 }_{\Greekmath 010D }\left( z_{i,t}\right)
\}_{i,t}$ are i.i.d. across $t\leq T$ for any $i$. Since $\mathbb{E}[\tilde{
\Greekmath 0120 }_{\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ]=0$, it follows that $\mathbb{E
}[2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}]=2^{-1}\sum_{i\leq
n}\Greekmath 011B _{\Greekmath 010D ,i}^{2}$, where $\Greekmath 011B _{\Greekmath 010D ,i}^{2}\equiv \mathbb{E}[
\tilde{\Greekmath 0120 }_{\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ^{2}]$. Assuming that $
\Greekmath 011B _{\Greekmath 010D ,i}^{2}$ is bounded above and below, the bias from the mean
of $2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$ is of the same order
as $L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $.
This provides the explanation why the maximized joint likelihood $L_{n,T}(
\hat{\Greekmath 011E })$ should be modified with a bias correction based on the mean of\
$2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$.
The above discussion demonstrates that the maximized joint likelihood\ $
L_{n,T}(\hat{\Greekmath 011E })$ requires a first-order bias correction in the form of a
consistent estimator of $2^{-1}\sum_{i\leq n}\mathbb{E}[\tilde{\Psi}
_{\Greekmath 010D ,i}^{2}]$. As a result, the QLR statistic in the classical Vuong
test must also be adjusted with a bias correction to ensure valid inference.
This bias correction for the QLR statistic is given by the difference in the
bias of $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$ ($j=1,2$) from the two competing models.
There is an additional complication. Depending on whether the models are
nested or overlapping, the first components, $L_{j,n,T}(\Greekmath 011E _{j}^{\ast })-
\mathbb{E}[L_{j,n,T}(\Greekmath 011E _{j}^{\ast})]$, from the two models may cancel out.
See \cite{Vuong1989} for discussion in the cross section context. In such
cases, the second components, denoted as $2^{-1}\sum_{i\leq n}\tilde{\Psi }
_{j,\Greekmath 010D ,i}^{2}$ ($j=1,2$),\ can become the primary drivers of the
asymptotic distribution of the bias-corrected QLR statistic.\footnote{
In the classical nested case, where the functional form of the likelihood $f$
is identical across the two competing models, the second component of the
test statistic cancels out. In this setting, bias correction of the
estimator itself, along with further refinements of (\ref{L_nT_exp}), is
needed. This case was previously studied by \cite{LeePhillips2015}, and we
do not pursue it further in this paper.}\ Consequently, the estimation error
of the incidental parameters influences not only the bias but also the
variance of the QLR statistic.
\subsection{Quasi-likelihood Estimator\label{section:log-likelihoods}}
The pseudo true parameter $\Greekmath 011E _{j}^{\ast }$ for model $j$ is estimated as
the maximizer $\hat{\Greekmath 011E }_{j}$ of $L_{j,n,T}(\Greekmath 011E _{j})$ where $\Greekmath 011E
_{j}\equiv (\Greekmath 0112 _{j}^{\top },(\Greekmath 010D _{j,g})_{g\in \mathcal{G}_{j}}^{\top
})^{\top }$. The estimator\ $\hat{\Greekmath 011E }_{j}$\ satisfies the following
first-order conditions:
\begin{equation}
\sum_{g\in \mathcal{G}_{j}}\sum_{i\in I_{j,g}}\sum_{t\leq T}\Greekmath 0120 _{j,\Greekmath 0112
}(z_{i,t};\hat{\Greekmath 011E }_{j,g})=0_{d_{\Greekmath 0112 }\times 1}, \label{C_foc_1}
\end{equation}
and for any $g\in \mathcal{G}_{j}$,
\begin{equation}
\sum_{i\in I_{j,g}}\sum_{t\leq T}\Greekmath 0120 _{j,\Greekmath 010D }(z_{i,t};\hat{\Greekmath 011E }
_{j,g})=0, \label{C_foc_2}
\end{equation}
where $\Greekmath 011E _{j,g}\equiv (\Greekmath 0112 _{j}^{\top },\Greekmath 010D _{j,g})^{\top }$\ and $
\Greekmath 0120 _{j,a}(z_{i,t};\Greekmath 011E _{j,g})\equiv \partial \Greekmath 0120 _{j}(z_{i,t};\Greekmath 011E
_{j,g})/\partial a$ for $a\in \{\Greekmath 0112 _{j},\Greekmath 010D _{j,g}\}$.\ The estimator
$\hat{\Greekmath 011E }_{j}$ is subsequently used to compute the maximized joint
likelihood $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$\ for model $j$.
Following the discussion in the previous subsection, the maximized joint
quasi-likelihood for each model must be adjusted to account for the
respective first-order bias, which can be derived from the asymptotic
expansion of $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$. To present this expansion in a
more general setting than previously considered, we first update the
notations introduced earlier. Let $\Greekmath 0120 _{j}\left( z_{i,t}\right) $, $\Greekmath 0120
_{j,\Greekmath 010D }\left( z_{i,t}\right) $, $\Greekmath 0120 _{j,\Greekmath 010D \Greekmath 010D }\left(
z_{i,t}\right) $ and $\tilde{\Psi}_{j,\Greekmath 010D ,i}$ represent the counterparts
to their definitions in (\ref{derv_1}) and (\ref{derv_2}), with $\Greekmath 011E
_{j,i}\equiv \Greekmath 011E _{j,g}$,
\begin{equation*}
\Greekmath 0120 _{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) \equiv \Greekmath 0120 _{j,\Greekmath 010D
}\left( z_{i,t}\right) -\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ and \ }\tilde{\Greekmath 0120 }_{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) \equiv
\frac{\Greekmath 0120 _{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) }{(n_{j,g}^{-1}\sum_{i
\in I_{j,g}}\mathbb{E}_{T}\left[ -\Greekmath 0120 _{j,\Greekmath 010D \Greekmath 010D }\left(
z_{i,t}\right) \right] )^{1/2}},
\end{equation*}
for any $i\in I_{j,g}$ and any $g\in \mathcal{G}_{j}$.\footnote{
The demeaned version of $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $, i.e., $
\Greekmath 0120 _{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) $ is introduced here because
the first-order condition for $\Greekmath 011E _{j}^{\ast }$ only ensures that $
\sum_{i\in I_{j,g}}\sum_{t\leq T}\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left(
z_{i,t}\right) ]=0$. This condition does not guarantee that $\mathbb{E}[\Greekmath 0120
_{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. If $\Greekmath 0120 _{j,\Greekmath 010D }\left(
z_{i,t}\right) $ is stationry across $t$ for each $i$, the first-order
condition for $\Greekmath 011E _{j}^{\ast }$ is simplified to $\sum_{i\in I_{j,g}}
\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. In this case, $
\mathbb{E}[\Greekmath 0120 _{1,\Greekmath 010D }\left( z_{i,t}\right) ]=0$ holds for model 1,
but is still not guaranteed for model 2.}
\subsection{Infeasible Modified QLR Statistic - Bias\label{section:lm-bias}}
With the updated notations, we can obtain\footnote{
See Theorem \ref{Rep_L} together with the decomposition in (\ref{Exp_2nd_1})
in the Appendix.} the following expansion:
\begin{equation}
L_{j,n,T}(\hat{\Greekmath 011E }_{j})=T^{1/2}\sum_{i\leq n}(\tilde{\Psi}_{j,i}+\tilde{V}
_{j,i}+\tilde{U}_{j,i})+o_{p}(G_{j}^{1/2}), \label{C_Exp_1}
\end{equation}
where the components are defined\footnote{
The $\tilde{\Psi}_{j,\Greekmath 010D ,i}$ is not the derivative of $\tilde{\Psi}
_{j,i} $, as can be seen from its definition (\ref{derv_2}).} as:
\begin{equation*}
\tilde{\Psi}_{j,i}\equiv T^{-1/2}\sum_{t\leq T}\Greekmath 0120 _{j}(z_{i,t};\Greekmath 011E
_{j,i}^{\ast })\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, \ \ }\tilde{V}_{j,i}\equiv \frac{\tilde{\Psi}
_{j,\Greekmath 010D ,i}^{2}}{2n_{j,g}T^{1/2}}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ and \ \ }\tilde{U}
_{j,i}\equiv \sum_{\{i^{\prime }\in I_{j,g}:i^{\prime }<i\}}\frac{\tilde{\Psi
}_{j,\Greekmath 010D ,i}\tilde{\Psi}_{j,\Greekmath 010D ,i^{\prime }}}{n_{j,g}T^{1/2}}.
\end{equation*}
It is reasonable to expect $\mathbb{E}[\tilde{U}_{j,i}]=0$ for $j=1,2$ and
for any $i\leq n$,\footnote{
For example, it is implied by Assumption \ref{A1}\ in the Appendix.}\textbf{
\ }which implies that the first-order bias arises from the mean of $\tilde{V}
_{j,i}$. To account for this bias, we define the (infeasible) modified
maximized joint likelihood as:
\begin{equation}
LM_{j,n,T}^{\ast }(\hat{\Greekmath 011E }_{j})\equiv L_{j,n,T}(\hat{\Greekmath 011E }
_{j})-T^{1/2}\sum_{i\leq n}\mathbb{E}[\tilde{V}_{j,i}]. \label{C_Exp_2}
\end{equation}
Using this expression, we define the (infeasible) modified QLR statistic,
whose asymptotic expansion follows from (\ref{C_Exp_1}):
\begin{equation}
MQLR_{n,T}^{\ast }\equiv (nT)^{-1/2}\left( LM_{1,n,T}^{\ast }(\hat{\Greekmath 011E }
_{1})-LM_{2,n,T}^{\ast }(\hat{\Greekmath 011E }_{2})\right) =n^{-1/2}\sum_{i\leq n}(
\tilde{\Psi}_{i}+\tilde{V}_{i}^{\ast }+\tilde{U}_{i})+o_{p}(T^{-1/2}),
\label{C_Exp_3}
\end{equation}
where $\tilde{\Psi}_{i}\equiv \tilde{\Psi}_{1,i}-\tilde{\Psi}_{2,i}$, $
\tilde{V}_{i}^{\ast }\equiv \tilde{V}_{1,i}-\tilde{V}_{2,i}-\mathbb{E}[
\tilde{V}_{1,i}-\tilde{V}_{2,i}]$ and $U_{i}\equiv \tilde{U}_{1,i}-\tilde{U}
_{2,i}$.
\subsection{Infeasible Modified QLR Statistic - Asymptotic Distribution\label
{section:lm-variance}}
After bias correction, the main remaining components arising from the
estimation errors of the incidental parameters, i.e., $\tilde{V}_{i}^{\ast }+
\tilde{U}_{i}$ primarily affect the variance of the modified QLR statistic\ $
MQLR_{n,T}^{\ast}$.\ This effect becomes particularly significant when the
two models are overlapping. In such cases, the first component on the
right-hand side of (\ref{C_Exp_1}), i.e., $n^{-1/2}\sum_{i\leq n}\tilde{\Psi
}_{i}$, can become arbitrarily small or even vanish, leaving the asymptotic
distribution of $MQLR_{n,T}^{\ast}$ predominantly determined by the
estimation errors of the incidental parameters, specifically $
n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast}+\tilde{U}_{i})$.
To facilitate the discussion of the variance and asymptotic distribution of $
MQLR_{n,T}^{\ast}$, we follow the literature (see, e.g., \cite{Vuong1989},
\cite{Shi2015b} and \cite{Liao&Shi2020}) and formally define the
relationships between models in the panel setting below.
\begin{definition}
(i) Model 1 and model 2 are \textbf{strictly non-nested} if there do not
exist $\Greekmath 011E _{1,i}\in \Phi _{1}$ and $\Greekmath 011E _{2,i}\in \Phi _{12}$ such that
\begin{equation}
\Greekmath 0120 _{1}\left( z;\Greekmath 011E _{1,i}\right) =\Greekmath 0120 _{2}\left( z;\Greekmath 011E _{2,i}\right)
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \label{Def_1}
\end{equation}
for any $i$ and any $z$ in the support of $z_{i,t}$; (ii) the two models are
\textbf{overlapping} if they are not strictly non-nested; (iii) model 1 and
model 2 are said to be \textbf{nested }if, for any $\Greekmath 011E _{j,i}\in \Phi _{j}$
, there exists a $\Greekmath 011E _{j^{\prime },i}\in \Phi _{j^{\prime }}$ (where $
j,j^{\prime }=1,2$ and $j\neq j^{\prime }$) such that (\ref{Def_1}) holds.
\end{definition}
When the two models are strictly non-nested, the variance of $
n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$ is bounded away from zero. In this
case, the component $n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast }+\tilde{U}
_{i})$ on the right-hand side of (\ref{C_Exp_1}) becomes asymptotically
negligible,\ as it is of order $O_{p}(T^{-1/2})$ under certain regularity
conditions.\footnote{
This order follows from an approximation of\ $n^{-1}\sum_{i\leq n}(\mathrm{
Var}(\tilde{V}_{i})+\mathrm{Var}(\tilde{U}_{i}))$ (see (\ref
{P_C_L_Auxillary9_11}) in the Online Appendix, which appears in the proof of
Lemma \ref{C_L_Auxillary9}), together with an upper bound on the variance of
$\tilde{\Greekmath 0120 }_{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) $ for $j=1,2$.}\
Consequently, the asymptotic distribution of $MQLR_{n,T}^{\ast }$ is
primarily determined by $n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$.
Conversely, when the two models are nested, $\tilde{\Psi}_{i}=0$ under the
null hypothesis. In this scenario, the asymptotic distribution of $
MQLR_{n,T}^{\ast }$ is determined by\ $n^{-1/2}\sum_{i\leq n}(\tilde{V}
_{i}^{\ast }+\tilde{U}_{i})$. As a result, deriving the asymptotic
distribution of $MQLR_{n,T}^{\ast }$ is relatively straightforward, when the
models are known to be either strictly non-nested or nested. In both cases,
the dominating term on the right-hand side of (\ref{C_Exp_1}) is
identifiable, and the asymptotic distribution of $MQLR_{n,T}^{\ast }$ can be
deduced from that of the dominating term.
The asymptotic distribution of $QLR_{n,T}^{\ast }$ is more challenging to
derive when the two models are overlapping. Because in this case, it is in
general unclear which components between $n^{-1/2}\sum_{i\leq n}\tilde{\Psi}
_{i}$ and $n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast }+\tilde{U}_{i})$
could be the dominating term, and both could be equally important. To
address this issue, we follow \cite{Liao&Shi2020} to consider both terms
together when deriving the asymptotic distribution of $MQLR_{n,T}^{\ast }$.
\begin{theorem}
\label{C_L0} Under Assumptions \ref{A1} and \ref{C_A1}\ in the Appendix, we
have
\begin{equation}
\frac{MQLR_{n,T}^{\ast }-\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}}\rightarrow
_{d}N(0,1)\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \label{C_P1_1}
\end{equation}
where
\begin{equation}
\Greekmath 0121 _{n,T}^{2}\equiv n^{-1}\sum_{i\leq n}\mathrm{Var}(\tilde{\Psi}_{i}+
\tilde{V}_{i}+\tilde{U}_{i})\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ and \ }\overline{QLR}_{n,T}\equiv
(nT)^{-1/2}\mathbb{E}[L_{1,n,T}(\Greekmath 011E _{1}^{\ast })-L_{1,n,T}(\Greekmath 011E _{2}^{\ast
})]. \label{def_omega}
\end{equation}
\end{theorem}
\begin{remark}
Theorem \ref{C_L0} provides the asymptotic distribution of $MQLR_{n,T}^{\ast
}$ under both the null and alternative hypotheses. Since $\overline{QLR}
_{n,T}=0$ under the null hypothesis, we can construct a test for the null by
using a consistent estimator of $\Greekmath 0121 _{n,T}$, along with a feasible
version of the modified QLR statistic. This feasible statistic replaces the
unknown bias correction terms in $MQLR_{n,T}^{\ast }$ with their consistent
estimators.\footnote{
We impose Assumption \ref{C_A2} to simplify the estimation of the bias
correction and the variance of $MQLR_{n,T}^{\ast }$.\ All assumptions are
collected in the Appendix.}
\end{remark}
\subsection{Estimation of Bias \label{section-biasestimation}}
We first discuss the estimation of the bias correction terms. Under
Assumption \ref{C_A2}\ in the Appendix, the definition of $\tilde{V}_{j,i}$
implies that:
\begin{equation}
T^{1/2}\sum_{i\leq n}\mathbb{E}[\tilde{V}_{j,i}]=\sum_{g\in \mathcal{G}
_{j}}\sum_{i\in I_{j,g}}(2n_{j,g})^{-1}\Greekmath 011B _{j,\Greekmath 010D ,i}^{2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, \ \
where }\Greekmath 011B _{j,\Greekmath 010D ,i}\equiv (\mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{j,\Greekmath 010D
}^{\ast }\left( z_{i,t}\right) ^{2}])^{1/2}. \label{C_Bias_1}
\end{equation}
This motivates a straightforward plug-in estimator of the bias correction
term for model $j$, relying on a consistent estimator of $\Greekmath 011B _{j,\Greekmath 010D
,i}^{2}$\ for any $i\in I_{j,g}$ and any $g\in \mathcal{G}_{j}$:
\begin{equation}
\hat{\Greekmath 011B }_{j,\Greekmath 010D ,i}^{2}\equiv \frac{\widehat{\mathbb{E}}_{T}[\hat{\Greekmath 0120
}_{j,\Greekmath 010D }(z_{i,t})^{2}]-(\widehat{\mathbb{E}}_{T}[\hat{\Greekmath 0120 }_{j,\Greekmath 010D
}(z_{i,t})])^{2}}{|\hat{\Psi}_{j,\Greekmath 010D \Greekmath 010D }(\hat{\Greekmath 011E }_{j,g})|},
\label{C_Var_1}
\end{equation}
where
\begin{equation}
\hat{\Greekmath 0120 }_{j,\Greekmath 010D }(z_{i,t})\equiv \Greekmath 0120 _{j,\Greekmath 010D }(z_{i,t};\hat{\Greekmath 011E }
_{j,g})\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ and \ \ }\hat{\Psi}_{j,\Greekmath 010D \Greekmath 010D }(\Greekmath 011E _{j,g})\equiv
n_{j,g}^{-1}\sum_{i\in I_{j,g}}\widehat{\mathbb{E}}_{T}[\Greekmath 0120 _{j,\Greekmath 010D
\Greekmath 010D }(z_{i,t};\Greekmath 011E _{j,g})]. \label{C_Var_2}
\end{equation}
The bias correction term for model $j$ is then estimated as
\begin{equation}
R_{j,n,T}(\hat{\Greekmath 011E }_{j})\equiv \sum_{g\in \mathcal{G}_{j}}\sum_{i\in
I_{j,g}}(2n_{j,g})^{-1}\hat{\Greekmath 011B }_{j,\Greekmath 010D ,i}^{2}. \label{C_B_est}
\end{equation}
Using this, the feasible modified QLR statistic is defined as:
\begin{equation}
MQLR_{n,T}\equiv (nT)^{-1/2}\left( LM_{1,n,T}(\hat{\Greekmath 011E }_{1})-LM_{2,n,T}(
\hat{\Greekmath 011E }_{2})\right) , \label{C_MQLR}
\end{equation}
where
\begin{equation*}
LM_{j,n,T}(\hat{\Greekmath 011E }_{j})\equiv L_{j,n,T}(\hat{\Greekmath 011E }_{j})-R_{j,n,T}(\hat{
\Greekmath 011E }_{j})
\end{equation*}
denotes the modified maximized joint likelihood for model $j$.
\begin{theorem}
\textit{\label{C_L1} }Under Assumptions \ref{A1}, \ref{C_A1} and \ref{C_A2}
in the Appendix, we have
\begin{equation*}
\frac{R_{j,n,T}(\hat{\Greekmath 011E }_{j})-T^{1/2}\sum_{i\leq n}\mathbb{E}[\tilde{V}
_{j,i}]}{(nT)^{1/2}\Greekmath 0121 _{n,T}}=O_{p}(T^{-1/2}).
\end{equation*}
\end{theorem}
Theorem \ref{C_L1} establishes the consistency of $R_{j,n,T}(\hat{\Greekmath 011E }_{j})$
as an estimator of the bias of the maximized joint likelihood $L_{j,n,T}(
\hat{\Greekmath 011E }_{j})$ in model $j$. Combined with the result in (\ref{C_P1_1}),
this directly leads to the asymptotic normality of the modified QLR
statistic $QLR_{n,T}$:
\begin{equation}
\frac{MQLR_{n,T}-\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}}\rightarrow_{d}N(0,1)
\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \label{C_L1_2}
\end{equation}
To apply this result for inference under the null hypothesis, it remains to
construct a consistent estimator for $\Greekmath 0121 _{n,T}^{2}$, which we now
proceed to address.
\subsection{Estimation of Variance\ \label{section:varianceestimation}}
A consistent estimator of $\Greekmath 0121 _{n,T}^{2}$, which was defined in (\ref
{def_omega}), can be constructed using the sample analog of the variance of $
n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$, and a variance correction term to
account for estimation errors of the incidental parameters. Define $\Greekmath 011B
_{n,T}^{2}\equiv n^{-1}\sum_{i\leq n}\mathrm{Var}(\tilde{\Psi}_{i})$ and $
\Delta \Greekmath 0120 (z_{i,t},\Greekmath 011E _{i})\equiv \Greekmath 0120 _{1}(z_{i,t};\Greekmath 011E _{1,i})-\Greekmath 0120
_{2}(z_{i,t};\Greekmath 011E _{2,i})$. Under Assumption \ref{C_A2}, a natural estimator
for $\Greekmath 011B _{n,T}^{2}$ is given by:
\begin{equation}
\hat{\Greekmath 011B }_{n,T}^{2}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\Delta
\Greekmath 0120 (z_{i,t};\hat{\Greekmath 011E }_{i})^{2}-(nT)^{-1}\left( MQLR_{n,T}\right) ^{2}.
\label{Var_Est_1}
\end{equation}
The theoretical property of $\hat{\Greekmath 011B }_{n,T}^{2}$ is provided in the
lemma below.\footnote{
The sample variance\ $\hat{\Greekmath 011B }_{n,T}^{2}$ will be used below as a
building block for variance estimation and therefore requires $\{\Delta \Greekmath 0120
(z_{i,t},\Greekmath 011E _{i}^{\ast })\}_{t}$ to be serially uncorrelated, as assumed
in Assumption \ref{C_A2}. If serial correlation is present, an
autocorrelation-consistent estimator is required. Since variance estimation
is mainly needed under the null for size control, Assumption \ref{C_A2} is
imposed only under the null; Lemma \ref{C_T2} in the Appendix shows that
consistency of the test does not rely on this assumption.}
\begin{theorem}
\textit{\label{C_L2}\ }Under Assumptions \ref{A1}, \ref{C_A1} and \ref{C_A2}
in the Appendix, we have:
\begin{equation}
\frac{\hat{\Greekmath 011B }_{n,T}^{2}-(\Greekmath 011B _{n,T}^{2}+2\Greekmath 011B _{S,n,T}^{2})}{
\Greekmath 0121 _{n,T}^{2}}=O_{p}(T^{-1/2}), \label{C_L2_1}
\end{equation}
where
\begin{equation}
\Greekmath 011B _{S,n,T}^{2}\equiv (2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}\sum_{i\in
I_{2,g}}\left( \Greekmath 011B _{1,\Greekmath 010D ,i}^{4}+n_{2,g}^{-2}\Greekmath 011B _{2,\Greekmath 010D
,i}^{2}\sum_{i^{\prime }\in I_{2,g}}s_{2,\Greekmath 010D ,i^{\prime
}}^{2}-2n_{2,g}^{-1}\Greekmath 011B _{12,\Greekmath 010D ,i}^{2}\right) , \label{C_Var_3}
\end{equation}
$s_{2,\Greekmath 010D ,i}^{2}\equiv \mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{2,\Greekmath 010D }\left(
z_{i,t}\right) ^{2}]$ and $\Greekmath 011B _{12,\Greekmath 010D ,i}\equiv \mathbb{E}_{T}[
\tilde{\Greekmath 0120 }_{1,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) \tilde{\Greekmath 0120 }_{2,\Greekmath 010D
}^{\ast }\left( z_{i,t}\right) ]$ for any $i\in I_{2,g}$ and any $g\in
\mathcal{G}_{2}$.\footnote{
Since $\mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{1,\Greekmath 010D }(z_{i,t})]=0$ by the
first-order condition for $\Greekmath 011E _{1}^{\ast }$, we have $\Greekmath 011B _{12,\Greekmath 010D
,i}=\mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{1,\Greekmath 010D }\left( z_{i,t}\right) \tilde{\Greekmath 0120 }
_{2,\Greekmath 010D }\left( z_{i,t}\right) ]$ for any $i\leq n$.}
\end{theorem}
Theorem \ref{C_L2} shows that $\hat{\Greekmath 011B }_{n,T}^{2}$ tends to overestimate
the true variance\ $\Greekmath 011B _{n,T}^{2}$ by $2\Greekmath 011B _{S,n,T}^{2}$. It is
reasonable to expect the latter to be of order $O(T^{-1})$, which is
formalized by Assumptions \ref{C_A1}(iii) and \ref{A4}. On the other hand,\
it can be shown\footnote{
See Lemma \ref{C_L_Auxillary9} in the Online Appendix.} that
\begin{equation}
\frac{\Greekmath 0121 _{n,T}^{2}-(\Greekmath 011B _{n,T}^{2}+\Greekmath 011B _{U,n,T}^{2})}{\Greekmath 0121
_{n,T}^{2}}=O(T^{-1/2}), \label{C_Var_4}
\end{equation}
where
\begin{equation}
\Greekmath 011B _{U,n,T}^{2}\equiv (2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}\sum_{i\in
I_{2,g}}\left( \Greekmath 011B _{1,\Greekmath 010D ,i}^{4}+n_{2,g}^{-2}\Greekmath 011B _{2,\Greekmath 010D
,i}^{2}\sum_{i^{\prime }\in I_{2,g}}\Greekmath 011B _{2,\Greekmath 010D ,i^{\prime
}}^{2}-2n_{2,g}^{-1}\Greekmath 011B _{12,\Greekmath 010D ,i}^{2}\right) , \label{C_Var_5}
\end{equation}
which represents the leading term in the variance of $n^{-1/2}\sum_{i\leq n}(
\tilde{V}_{i}^{\ast }+\tilde{U}_{i})$.
Comparing their expressions in (\ref{C_Var_3}) and (\ref{C_Var_5}), it is
evident that $\Greekmath 011B _{S,n,T}^{2}\geq \Greekmath 011B _{U,n,T}^{2}$ in general, with
the inequality becoming strict when $\mathbb{E}[\tilde{\Greekmath 0120 }_{j,\Greekmath 010D
}\left( z_{i,t}\right) ]\neq 0$. From (\ref{C_L2_1}) and (\ref{C_Var_4}), it
follows that $\hat{\Greekmath 011B }_{n,T}^{2}$ is a consistent estimator of $\Greekmath 0121
_{n,T}^{2}$ when $\Greekmath 011B _{n,T}^{2}$ dominates $\Greekmath 011B _{S,n,T}^{2}$. Since $
\Greekmath 011B _{S,n,T}^{2}=O(T^{-1})$, this dominance is guaranteed when the two
models are strictly non-nested, as $\Greekmath 011B _{n,T}^{2}$ is bounded away from
zero in such cases. Conversely, if $\Greekmath 011B _{n,T}^{2}$ is of the same or
even smaller order than\ $\Greekmath 011B _{S,n,T}^{2}$, a scenario that arises when
the two models are (nearly) nested or overlapping, the result in (\ref
{C_L2_1}) implies that $\hat{\Greekmath 011B }_{n,T}^{2}$ overestimates $\Greekmath 0121
_{n,T}^{2}$ by $2\Greekmath 011B _{S,n,T}^{2}-\Greekmath 011B _{U,n,T}^{2}$, and becomes an
inconsistent estimator of $\Greekmath 0121 _{n,T}^{2}$.
The over-estimation issue in $\hat{\Greekmath 011B }_{n,T}^{2}$ can be adjusted
through consistent estimators of $\Greekmath 011B _{U,n,T}^{2}$ and $\Greekmath 011B
_{S,n,T}^{2}$. From its expression in (\ref{C_Var_5}), $\Greekmath 011B _{U,n,T}^{2}$
can be estimated by replacing $\Greekmath 011B _{j,\Greekmath 010D ,i}^{2}$ with $\hat{\Greekmath 011B }
_{j,\Greekmath 010D ,i}^{2}$ defined in (\ref{C_Var_1}), and replacing $\Greekmath 011B
_{12,\Greekmath 010D ,i}^{2}$\ by
\begin{equation*}
\hat{\Greekmath 011B }_{12,\Greekmath 010D ,i}^{2}\equiv \frac{(\widehat{\mathbb{E}}_{T}[\hat{
\Greekmath 0120 }_{1,\Greekmath 010D }(z_{i,t})\hat{\Greekmath 0120 }_{2,\Greekmath 010D }(z_{i,t})])^{2}}{|\hat{\Psi}
_{1,\Greekmath 010D \Greekmath 010D }(\hat{\Greekmath 011E }_{1,g})\hat{\Psi}_{2,\Greekmath 010D \Greekmath 010D }(\hat{\Greekmath 011E }
_{2,g})|}.
\end{equation*}
Specifically, the estimator of $\Greekmath 011B _{U,n,T}^{2}$ is defined as
\begin{equation}
\hat{\Greekmath 011B }_{U,n,T}^{2}\equiv (2nT)^{-1}\sum_{g\in \mathcal{G}
_{2}}\sum_{i\in I_{2,g}}\left( \hat{\Greekmath 011B }_{1,\Greekmath 010D ,i}^{4}+n_{2,g}^{-2}
\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i}^{2}\sum_{i^{\prime }\in I_{2,g}}\hat{\Greekmath 011B }
_{2,\Greekmath 010D ,i^{\prime }}^{2}-2n_{2,g}^{-1}\hat{\Greekmath 011B }_{12,\Greekmath 010D
,i}^{2}\right) . \label{Var_Est_U}
\end{equation}
Similarly, we define the estimator of $\Greekmath 011B _{S,n,T}^{2}$ as
\begin{equation}
\hat{\Greekmath 011B }_{S,n,T}^{2}\equiv (2nT)^{-1}\sum_{g\in \mathcal{G}
_{2}}\sum_{i\in I_{2,g}}\left( \hat{\Greekmath 011B }_{1,\Greekmath 010D ,i}^{4}+n_{2,g}^{-2}
\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i}^{2}\sum_{i^{\prime }\in I_{2,g}}\hat{s}_{2,\Greekmath 010D
,i^{\prime }}^{2}-2n_{2,g}^{-1}\hat{\Greekmath 011B }_{12,\Greekmath 010D ,i}^{2}\right) .
\label{Var_Est_S}
\end{equation}
Here, for any $i\in I_{2,g}$ and any $g\in \mathcal{G}_{2}$:
\begin{equation}
\hat{s}_{2,\Greekmath 010D ,i}^{2}\equiv \frac{\widehat{\mathbb{E}}_{T}[\hat{\Greekmath 0120 }
_{j,\Greekmath 010D }(z_{i,t})^{2}]}{|\hat{\Psi}_{j,\Greekmath 010D \Greekmath 010D }(\hat{\Greekmath 011E }_{j,g})|
}. \label{Var_Est_3}
\end{equation}
Thus, $\Greekmath 0121 _{n,T}^{2}$ can be consistently estimated in a general setting
as $\hat{\Greekmath 011B }_{n,T}^{2}+\hat{\Greekmath 011B }_{U,n,T}^{2}-2\hat{\Greekmath 011B }
_{S,n,T}^{2} $, as indicated by Theorem \ref{C_L3} below.
\begin{theorem}
\textit{\label{C_L3}\ }Under Assumptions \ref{A1}, \ref{C_A1} and \ref{C_A2}
in the Appendix, we have:
\begin{equation}
\frac{\hat{\Greekmath 011B }_{U,n,T}^{2}-\Greekmath 011B _{U,n,T}^{2}}{\Greekmath 0121 _{n,T}^{2}}
=O_{p}(n^{-1/2})\ \ \ \ \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and \ \ }\frac{\hat{\Greekmath 011B }_{S,n,T}^{2}-
\Greekmath 011B _{S,n,T}^{2}}{\Greekmath 0121 _{n,T}^{2}}=O_{p}(n^{-1/2}). \label{C_L3_1}
\end{equation}
\end{theorem}
One drawback of using $\hat{\Greekmath 011B }_{n,T}^{2}-2\hat{\Greekmath 011B }_{S,n,T}^{2}+\hat{
\Greekmath 011B }_{U,n,T}^{2}$ as a variance estimator is that it may take negative
values in finite samples. To address this issue, we follow \cite
{Liao&Shi2020} and propose the following hybrid variance estimator:
\begin{equation}
\hat{\Greekmath 0121 }_{n,T}^{2}=\max \left\{ \hat{\Greekmath 011B }_{n,T}^{2}+\hat{\Greekmath 011B }
_{U,n,T}^{2}-2\hat{\Greekmath 011B }_{S,n,T}^{2},\hat{\Greekmath 011B }_{U,n,T}^{2}\right\} .
\label{Var_Est_5}
\end{equation}
\begin{remark}
By definition, $\hat{\Greekmath 011B }_{U,n,T}^{2}$ can be expressed as
\begin{align}
\hat{\Greekmath 011B }_{U,n,T}^{2}& =(2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}\sum_{i\in
I_{2,g}}(\hat{\Greekmath 011B }_{1,\Greekmath 010D ,i}^{4}-2n_{2,g}^{-1}\hat{\Greekmath 011B }_{12,\Greekmath 010D
,i}^{2}+n_{2,g}^{-2}\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i}^{4}) \notag \\
& +(2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}n_{2,g}^{-2}\sum_{i\in
I_{2,g}}\sum_{i^{\prime }\in I_{2,g},i^{\prime }\neq i}\hat{\Greekmath 011B }
_{2,\Greekmath 010D ,i}^{2}\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i^{\prime }}^{2}.
\label{Var_Est_6}
\end{align}
Since\ $\hat{\Greekmath 011B }_{12,\Greekmath 010D ,i}^{2}\leq \hat{\Greekmath 011B }_{1,\Greekmath 010D ,i}^{2}
\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i}^{2}$ for any $i$, the first term on the
right-hand side of (\ref{Var_Est_6}) is non-negative. Additionally, the
second term after the equality in (\ref{Var_Est_6}) is also non-negative by
definition. Therefore, $\hat{\Greekmath 011B }_{U,n,T}^{2}$ is guaranteed to be
non-negative, and by construction, $\hat{\Greekmath 0121 }_{n,T}^{2}$ is also
non-negative. This ensures the practical reliability of $\hat{\Greekmath 0121 }
_{n,T}^{2}$ as a variance estimator.
\end{remark}
\subsection{Feasible Modified QLR Statistic and Generalized Vuong Test\label
{section:feasiblestatistic}}
Given the variance estimator $\hat{\Greekmath 0121 }_{n,T}^{2}$, the two-sided and
one-sided Vuong tests are defined as:
\begin{eqnarray}
\Greekmath 0127 _{n,T}^{2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p) &\equiv &1\left\{ \left\vert
MQLR_{n,T}\right\vert >\hat{\Greekmath 0121 }_{n,T}z_{1-p/2}\right\} \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \notag
\\
\Greekmath 0127 _{n,T}^{1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p) &\equiv &1\left\{ MQLR_{n,T}>\hat{\Greekmath 0121 }
_{n,T}z_{1-p}\right\} , \label{C_Vuong_test}
\end{eqnarray}
respectively, where $z_{1-p}$ is the $1-p$ quantile of the standard normal
distribution, and $p\in (0,1)$ denotes the significance level.
\begin{theorem}
\label{C_T1}\ Suppose that Assumptions \ref{A1}, \ref{C_A1} and \ref{C_A2}
in the Appendix hold. Further suppose that as $n,T\rightarrow \infty $,
\begin{equation}
\frac{\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}}\rightarrow c, \label{C_T1_1}
\end{equation}
for some constant $c\ $in the extended real line. Then as $n,T\rightarrow
\infty $,
\begin{equation}
\mathbb{E}[\Greekmath 0127 _{n,T}^{2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 2-\Phi
(z_{1-p/2}-c)-\Phi (z_{1-p/2}+c), \label{C_T1_2}
\end{equation}
and
\begin{equation}
\mathbb{E}[\Greekmath 0127 _{n,T}^{1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 1-\Phi (z_{1-p}-c),
\label{C_T1_3}
\end{equation}
where $\Phi (\cdot )$ denotes the cumulative distribution function of the
standard normal.
\end{theorem}
\begin{remark}
Theorem \ref{C_T1} establishes both the size and power properties of the
Vuong test defined in (\ref{C_Vuong_test}). Under the null hypothesis where $
\overline{QLR}_{n,T}=0$, (\ref{C_T1_1}) holds with $c=0$. In this scenario, (
\ref{C_T1_2}) and (\ref{C_T1_3}) imply that
\begin{equation}
\mathbb{E}[\Greekmath 0127 _{n,T}^{2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 2(1-\Phi
(z_{1-p/2}))=p\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ \ and \ \ \ }\mathbb{E}[\Greekmath 0127 _{n,T}^{1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side
}}(p)]\rightarrow p, \label{C_Vuong_Size}
\end{equation}
which establishes the size control over data generating processes in which
the two models being compared are either nested or overlapping. Furthermore,
(\ref{C_T1_2}) demonstrates that the two-sided Vuong test has power against
local alternatives with $\left\vert c\right\vert \in (0,\infty )$, and is
consistent against any fixed alternatives, while the one-sided Vuong test
share the similar power properties with $c\in (0,\infty )$.
\end{remark}
\section{Testing Heterogeneous Time Effects vs. TWFE\label{sec:TWE}}
The general result in the previous section was based on Theorems \ref
{Rep_Theta} -- \ref{QLR_Stat} in Appendix \ref{Sec:AP2}. As discussed in the
introduction, the result established in Appendix \ref{Sec:AP2} is a
technical contribution that relaxes the concavity assumption in \cite
{FernandezValWeidener2016}, but it was done at the cost of restricting the
group structure. In this section, we consider potentially more complicated
group structure but we do so in the context of the linear model, which
satisfies the concavity assumption. Because of substantive economic interest
associated with the linear model, in particular, the literature initiated by
\cite{Bonhomme-Manresa}, we believe that the linear model may deserve a
special attention.
\cite{Bonhomme-Manresa} have recently proposed a new model that extends
beyond the scope of traditional panel data models. In light of its
computational demands, one may naturally question whether a standard TWFE
specification could offer a practical alternative. Obviously the two models
are not nested, so it is of interest to find a Vuong test for this
comparison.
We consider the linear panel model
\begin{equation*}
y_{i,t}=x_{i,t}^{\top }\Greekmath 010C +\Greekmath 010D _{g(i),m(t)}+\Greekmath 0122 _{i,t},
\end{equation*}
where $y_{i,t}$ is the dependent variable, $x_{i,t}$ are the observed
regressors, and $\Greekmath 0122 _{i,t}$ is i.i.d. across $i$ and $t$ and with
mean zero and variance $\Greekmath 011B ^{2}$. Here, the $g(\cdot )$ and $m(\cdot )$
are \emph{known} cluster/group assignment functions with range $\mathcal{G}
=\{1,\ldots ,G\}$ and $\mathcal{M}=\{1,\ldots ,M\}$ respectively, and $
\Greekmath 010D _{g(i),m(t)}$ denotes the unknown fixed effect.\footnote{
The $\Greekmath 010D _{g(i),m(t)}$ takes different forms under various model
specifications. For example, in models involving time invariant individual
fixed effects only, we have $\Greekmath 010D _{g(i),m(t)}=\Greekmath 010D _{i}$, e.g., with $
\mathcal{G}=\{1,\ldots ,n\}$, and $\mathcal{M}=\{1\}$. In models involving
only time fixed effects, we have $\Greekmath 010D _{g(i),m(t)}=\Greekmath 010D _{t}$, e.g.,
with $\mathcal{G}=\{1\}$, and $\mathcal{M}=\{1,\ldots ,T\}$. In the Section
\ref{sec:linear-section} of the Online Appendix, we consider non-nested
hypothesis testing comparing several variants of the linear model.} In the
main text of the paper, we focus on the Vuong test comparing the
heterogeneous time effects model (as considered by \cite{Bonhomme-Manresa})
and the TWFE model, denoted as model 1 and model 2, respectively.
Specifically, model 1 defines the joint log-likelihood as
\begin{equation}
\sum_{g\in \mathcal{G}_{1}}\sum_{i\in I_{g}}\sum_{t\leq T}\Greekmath 0120 \left(
z_{i,t};\Greekmath 0112 _{1},\Greekmath 010D _{1,g,t}\right) , \label{M_1}
\end{equation}
where $\mathcal{G}_{1}\equiv \{1,\ldots ,G_{1}\}$ and $G_{1}$ is a positive
integer. The function $\Greekmath 0120 \left( z_{i,t};\Greekmath 0112 ,\Greekmath 010D \right) $ is given
by
\begin{equation}
\Greekmath 0120 \left( z_{i,t};\Greekmath 0112 ,\Greekmath 010D \right) =(-2)^{-1}(y_{i,t}-x_{i,t}^{\top
}\Greekmath 0112 -\Greekmath 010D )^{2}. \label{M_Form}
\end{equation}
In contrast, model 2 assumes that the joint log-likelihood is
\begin{equation}
\sum_{i\leq n}\sum_{t\leq T}\Greekmath 0120 \left( z_{i,t};\Greekmath 0112 _{2},\Greekmath 010D
_{2,i}+\Greekmath 010D _{2,t}\right) , \label{M_2}
\end{equation}
where $\Greekmath 010D _{2,i}$ represents the time-invariant individual effect, and $
\Greekmath 010D _{2,t}$ captures the homogeneous time effect. We impose the following
normalization condition on model 2:
\begin{equation}
\sum_{t\leq T}\Greekmath 010D _{2,t}=0, \label{Norm_C}
\end{equation}
to ensure unique identification of the pseudo-true parameters.
In the following, we introduce the algorithm for calculating the modified
QLR statistic and explain its asymptotic properties. Even though the
technical details are more involved and use more complex notation, the
results in this section are presented in a way quite similar to those in the
previous one. To make things easier to follow, we have organized this
section in a way that mirrors the structure of the (second half of the)
previous section.
\subsection{Estimation under Group-Time Effects and TWFE}
The pseudo true parameters $\Greekmath 0112 _{1}^{\ast }$ and $\Greekmath 010D _{1,g,t}^{\ast
} $ (for $g\in \mathcal{G}_{1}$ and $t\leq T$) for model 1, defined as the
maximizers of the population version of the objective in (\ref{M_1}), take
the following forms:
\begin{equation*}
\Greekmath 0112 _{1}^{\ast }=\frac{\Sigma _{\dot{x}}^{-1}}{nT}\sum_{i\leq
n}\sum_{t\leq T}\mathbb{E}[\dot{x}_{i,t}^{\ast }y_{i,t}]\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ \ \ and \
\ \ \ }\Greekmath 010D _{1,g,t}^{\ast }=\mathbb{E}[\bar{y}_{g,t}]-\mathbb{E}[\bar{x}
_{g,t}^{\top }]\Greekmath 0112 _{1}^{\ast },
\end{equation*}
where $\Sigma _{\dot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\mathbb{E
}[\dot{x}_{i,t}^{\ast }\dot{x}_{i,t}^{\ast \top }]$, $\dot{x}_{i,t}^{\ast
}=x_{i,t}-\mathbb{E}[\bar{x}_{g,t}]$, and for any $i\in I_{g}$ and any $g\in
\mathcal{G}_{1}$, $\bar{x}_{g,t}\equiv n_{g}^{-1}\sum_{i\in I_{g}}x_{i,t}$
and $\bar{y}_{g,t}\equiv n_{g}^{-1}\sum_{i\in I_{g}}y_{i,t}$.\ The
estimators of $\Greekmath 0112 _{1}^{\ast }$ and $\Greekmath 010D _{1,g,t}^{\ast }$ are their
sample analogs:
\begin{equation}
\hat{\Greekmath 0112 }_{1}=\frac{\hat{\Sigma}_{\dot{x}}^{-1}}{nT}\sum_{i\leq
n}\sum_{t\leq T}\dot{x}_{i,t}y_{i,t}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ \ \ and \ \ \ \ }\hat{\Greekmath 010D }
_{1,g,t}=\bar{y}_{g,t}-\bar{x}_{g,t}\hat{\Greekmath 0112 }_{1}, \label{HTFE_1}
\end{equation}
where $\dot{x}_{i,t}\equiv x_{i,t}-\bar{x}_{g,t}$ and $\hat{\Sigma}_{\dot{x}
}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\dot{x}_{i,t}\dot{x}
_{i,t}^{\top }$.
Applying (\ref{Norm_C}), we can solve the pseudo true parameters $\Greekmath 0112
_{2}^{\ast }$, $(\Greekmath 010D _{2,i}^{\ast })_{i\leq n}$\ and\ $(\Greekmath 010D
_{2,t}^{\ast })_{t\leq T}$ in model 2 as
\begin{equation*}
\Greekmath 0112 _{2}^{\ast }=\frac{\Sigma _{\ddot{x}^{\ast }}^{-1}}{nT}\sum_{i\leq
n}\sum_{t\leq T}\mathbb{E}[\ddot{x}_{i,t}^{\ast }y_{i,t}]\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, \ }\Greekmath 010D
_{2,i}^{\ast }=\mathbb{E}[\bar{y}_{i}]-\mathbb{E}[\bar{x}_{i}^{\top }]\Greekmath 0112
_{2}^{\ast }\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ and \ }\Greekmath 010D _{2,t}^{\ast }=\mathbb{E}[\bar{y}_{t}-
\bar{y}]-\mathbb{E}[\bar{x}_{t}^{\top }-\bar{x}^{\top }]\Greekmath 0112 _{2}^{\ast },
\end{equation*}
where $\Sigma _{\ddot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\mathbb{
E}[\ddot{x}_{i,t}^{\ast }\ddot{x}_{i,t}^{\ast \top }]$ and $\ddot{x}
_{i,t}^{\ast }\equiv x_{i,t}-\mathbb{E}[\bar{x}_{i}-\bar{x}_{t}+\bar{x}]$.
The estimators of the pseudo true parameters are their empirical
counterparts:
\begin{equation}
\hat{\Greekmath 0112 }_{2}=\frac{\hat{\Sigma}_{\ddot{x}}^{-1}}{nT}\sum_{i\leq
n}\sum_{t\leq T}\ddot{x}_{i,t}y_{i,t}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, \ }\hat{\Greekmath 010D }_{2,i}=\bar{y}
_{i}-\bar{x}_{i}^{\top }\hat{\Greekmath 0112 }_{2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ and \ }\hat{\Greekmath 010D }_{2,t}=(
\bar{y}_{t}-\bar{y})-(\bar{x}_{t}-\bar{x})^{\top }\hat{\Greekmath 0112 }_{2},
\label{TFE_2}
\end{equation}
where\ $\ddot{x}_{i,t}\equiv x_{i,t}-\bar{x}_{i}-\bar{x}_{t}+\bar{x}$ and $
\hat{\Sigma}_{\ddot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\ddot{x}
_{i,t}\ddot{x}_{i,t}^{\top }$.
\subsection{Asymptotic Distribution of the QLR Statistic}
Using the estimators of the pseudo true parameters from both models, we
obtain the QLR statistic for model comparison as
\begin{equation*}
QLR_{n,T}=(nT)^{-1/2}\left( \sum_{i\leq n}\sum_{t\leq T}\frac{\hat{
\Greekmath 0122 }_{1,i,t}^{2}}{-2}-\sum_{i\leq n}\sum_{t\leq T}\frac{\hat{
\Greekmath 0122 }_{2,i,t}^{2}}{-2}\right)
\end{equation*}
where $\hat{\Greekmath 0122 }_{j,i,t}\equiv y_{i,t}-x_{i,t}^{\top }\hat{\Greekmath 0112 }
_{j}-\hat{\Greekmath 010D }_{j,i,t}$\ for $j=1,2$, $\hat{\Greekmath 010D }_{1,i,t}\equiv \hat{
\Greekmath 010D }_{1,g,t}$ for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, and $
\hat{\Greekmath 010D }_{2,i,t}\equiv \hat{\Greekmath 010D }_{2,i}+\hat{\Greekmath 010D }_{2,t}$ for any $
i\leq n$ and any $t\leq T$.
We next derive the asymptotic distribution of $QLR_{n,T}$. Some notation are
needed. Let $\Greekmath 0122 _{j,i,t}\equiv y_{i,t}-x_{i,t}^{\top }\Greekmath 0112
_{j}^{\ast }-\Greekmath 010D _{j,i,t}^{\ast }$ and $\Greekmath 0122 _{j,i,t}^{\ast
}\equiv \Greekmath 0122 _{j,i,t}-\mathbb{E}[\Greekmath 0122 _{j,i,t}]$ for $j=1,2$,
where $\Greekmath 010D _{1,i,t}^{\ast }\equiv \Greekmath 010D _{1,g,t}^{\ast }$ for any $i\in
I_{g}$ and any $g\in \mathcal{G}_{1}$, and $\Greekmath 010D _{2,i,t}^{\ast }\equiv
\Greekmath 010D _{2,i}^{\ast }+\Greekmath 010D _{2,t}^{\ast }$ for any $i\leq n$ and any $
t\leq T$. Without loss of generality,\ we suppose that for any $g\in G_{1}$,
$I_{g}=\{N_{g}+1,\ldots ,N_{g}+n_{g}\}$ where $N_{g}\equiv \sum_{g^{\prime
}<g}n_{g^{\prime }}$.
For model 1, we define
\begin{equation*}
\tilde{V}_{1,i}=\frac{\sum_{t\leq T}\Greekmath 0122 _{1,i,t}^{\ast 2}}{
2n_{g}T^{1/2}}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ \ \ \ and \ \ \ \ \ }\tilde{U}_{1,i}=\frac{
\sum_{i^{\prime }=N_{g}+1}^{i-1}\sum_{t\leq T}\Greekmath 0122 _{1,i,t}^{\ast
}\Greekmath 0122 _{1,i^{\prime }t}^{\ast }}{n_{g}T^{1/2}},
\end{equation*}
for any $g\in \mathcal{G}_{1}$ and any $i\in I_{g}$. For model 2, we let
\begin{equation*}
\tilde{V}_{2,i}=\frac{(\sum_{t\leq T}\Greekmath 0122 _{2,i,t}^{\ast })^{2}}{
2T^{3/2}}+\frac{\sum_{t\leq T}\Greekmath 0122 _{2,i,t}^{\ast 2}}{2nT^{1/2}}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
\ \ \ \ and \ \ \ \ }\tilde{U}_{2,i}=\frac{\sum_{i^{\prime
}=1}^{i-1}\sum_{t\leq T}\Greekmath 0122 _{2,i,t}^{\ast }\Greekmath 0122
_{2,i^{\prime }t}^{\ast }}{nT^{1/2}}.
\end{equation*}
It can be shown\footnote{
See Lemma \ref{TWFE_Est2} and Lemma \ref{GTFE_Est2} in the Online Appendix.}
that
\begin{equation}
QLR_{n,T}=n^{-1/2}\sum_{i\leq n}(\tilde{\Psi}_{i}+\tilde{V}_{i}+\tilde{U}
_{i})+o_{p}(n^{-1/2}), \label{QLR-approx-linear}
\end{equation}
where\ $\tilde{\Psi}_{i}=(2T^{1/2})^{-1}\sum_{t\leq T}(\Greekmath 0122
_{2,i,t}^{2}-\Greekmath 0122 _{1,i,t}^{2})$, $\tilde{V}_{i}=\tilde{V}_{1,i}-
\tilde{V}_{2,i}$ and $\tilde{U}_{i}=\tilde{U}_{1,i}-\tilde{U}_{2,i}$. The
variance $\Greekmath 0121 _{n,T}^{2}$ of the $QLR_{n,T}$ is defined as in (\ref
{def_omega}), with $\tilde{\Psi}_{i}$, $\tilde{V}_{i}$ and $\tilde{U}_{i}$
taking the specific forms given above. Below is the asymptotic property of
the infeasible test statistic that mirrors Theorem \ref{C_L0}:
\begin{theorem}
\textit{\label{TWFE_Vuong}\ Under\ Assumptions \ref{A1},\ }\ref{A8} and \ref
{A9} in the Appendix, we have
\begin{equation*}
\frac{QLR_{n,T}-\mathbb{E}[S_{n,T}]-\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}}
\rightarrow _{d}N(0,1),
\end{equation*}
where $\mathbb{E}[S_{n,T}]=(2nT)^{-1/2}\sum_{g\in \mathcal{G}_{1}}\sum_{i\in
I_{g}}\mathbb{E}\left[ \sum_{t\leq T}(n_{g}^{-1}\Greekmath 0122 _{1,i,t}^{\ast
2}-n^{-1}\Greekmath 0122 _{2,i,t}^{\ast 2})-T^{-1}(\sum_{t\leq T}\Greekmath 0122
_{2,i,t}^{\ast })^{2}\right] $.
\end{theorem}
\subsection{Bias/Variance Estimation and the Feasible Test}
Following the discussion in Section \ref{Sec: Simple-Vuong}, we present a
feasible modified QLR procedure by estimating the bias $\mathbb{E}[S_{n,T}]$
and the variance $\Greekmath 0121 _{n,T}^{2}$. Consistent estimation of these
quantities is required under the null hypothesis to ensure correct size
control of the test. For this purpose, we assume that $(\Greekmath 0122
_{1,i,t},\Greekmath 0122 _{2,i,t})$ are independent across $t$ with
time-invariant first and second moments; this condition is imposed in
Assumption \ref{A10}.\footnote{
The independence of $(\Greekmath 0122 _{1,i,t},\Greekmath 0122 _{2,i,t})$ over $t$
is imposed primarily to simplify the form of $\Greekmath 0121 _{n,T}^{2}$ and its
consistent estimator. This condition is only required under the null
hypothesis; see Lemma \ref{TWFE_F_Vuong_Power} in the Appendix for
consistency of the test without this assumption.}
Under this condition, we have
\begin{equation*}
\mathbb{E}[S_{n,T}]=(2nT)^{-1/2}\sum_{g\in \mathcal{G}_{1}}\sum_{i\in I_{g}}
\left[ Tn_{g}^{-1}\Greekmath 011B _{1,i}^{2}-(1+Tn^{-1})\Greekmath 011B _{2,i}^{2}\right]
\end{equation*}
where $\Greekmath 011B _{j,i}^{2}\equiv \mathrm{Var}(\Greekmath 0122 _{j,i,t})$ can be
estimated by $\hat{\Greekmath 011B }_{j,i}^{2}\equiv T^{-1}\sum_{t\leq T}\hat{
\Greekmath 0122 }_{j,i,t}^{2}$. Therefore, the bias term can be estimated by
\begin{equation}
\widehat{\mathbb{E}[S_{n,T}]}=(2nT)^{-1/2}\sum_{g\in \mathcal{G}
_{1}}\sum_{i\in I_{g}}\left[ Tn_{g}^{-1}\hat{\Greekmath 011B }_{1,i}^{2}-(1+Tn^{-1})
\hat{\Greekmath 011B }_{2,i}^{2}\right] , \label{TWE_Best}
\end{equation}
which leads to the modified QLR statistic:
\begin{equation}
MQLR_{n,T}=QLR_{n,T}-\widehat{\mathbb{E}[S_{n,T}]}. \label{def_MQLR}
\end{equation}
It can be shown\footnote{
See Lemma \ref{TWFE_Var} in the Appendix.}\ that $\Greekmath 0121 _{n,T}^{2}$ can be
approximated by\ $\Greekmath 011B _{n,T}^{2}+\Greekmath 011B _{U,n,T}^{2}$, where
\begin{equation}
\Greekmath 011B _{n,T}^{2}\equiv (4n)^{-1}\sum_{i\leq n}\mathrm{Var}(\Greekmath 0122
_{2,i,t}^{2}-\Greekmath 0122 _{1,i,t}^{2}) \label{def__Variance_1}
\end{equation}
and
\begin{align}
\Greekmath 011B _{U,n,T}^{2}& \equiv (2nT)^{-1}\sum_{i\leq n}\Greekmath 011B
_{2,i}^{4}+(2n)^{-1}\sum_{g\in \mathcal{G}_{1}}n_{g}^{-2}\left( \sum_{i\in
I_{g}}\Greekmath 011B _{1,i}^{2}\right) ^{2} \notag \\
& +(2n^{3})^{-1}\left( \sum_{i\leq n}\Greekmath 011B _{2,i}^{2}\right)
^{2}-n^{-2}\sum_{g\in \mathcal{G}_{1}}n_{g}^{-1}\left( \sum_{i\in
I_{g}}\Greekmath 011B _{1,2,i}\right) ^{2}. \label{def__Variance_2}
\end{align}
We consider the sample variance
\begin{equation}
\hat{\Greekmath 011B }_{n,T}^{2}\equiv (4nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}(\hat{
\Greekmath 0122 }_{2,i,t}^{2}-\hat{\Greekmath 0122 }_{1,i,t}^{2})^{2}-(nT)^{-1}\left(
MQLR_{n,T}\right) ^{2}, \label{def__Variance_Est_1}
\end{equation}
as an estimator of $\Greekmath 011B _{n,T}^{2}$. The variance $\Greekmath 011B _{U,n,T}^{2}$
from estimating the incidental parameters is then estimated by
\begin{align}
\hat{\Greekmath 011B }_{U,n,T}^{2}& \equiv (2nT)^{-1}\sum_{i\leq n}\hat{\Greekmath 011B }
_{2,i}^{4}+(2n)^{-1}\sum_{g\in \mathcal{G}_{1}}n_{g}^{-2}\left( \sum_{i\in
I_{g}}\hat{\Greekmath 011B }_{1,i}^{2}\right) ^{2} \notag \\
& +(2n^{3})^{-1}\left( \sum_{i\leq n}\hat{\Greekmath 011B }_{2,i}^{2}\right)
^{2}-n^{-2}\sum_{g\in \mathcal{G}_{1}}n_{g}^{-1}\left( \sum_{i\in I_{g}}\hat{
\Greekmath 011B }_{1,2,i}\right) ^{2}, \label{def__Variance_Est_2}
\end{align}
where\ $\hat{\Greekmath 011B }_{1,2,i}\equiv T^{-1}\sum_{t\leq T}\hat{\Greekmath 0122 }
_{1,i,t}\hat{\Greekmath 0122 }_{2,i,t}$.
Proceeding as in the previous section, our estimate of $\Greekmath 0121 _{n,T}^{2}$
is thus
\begin{equation}
\hat{\Greekmath 0121 }_{n,T}^{2}\equiv \max \left\{ \hat{\Greekmath 011B }_{n,T}^{2}-\hat{\Greekmath 011B }
_{U,n,T}^{2},\hat{\Greekmath 011B }_{U,n,T}^{2}\right\} , \label{def_omega-hat}
\end{equation}
and the Vuong test follows (\ref{C_Vuong_test}) with $MQLR_{n,T}$ and $\hat{
\Greekmath 0121 }_{n,T}^{2}$ constructed in (\ref{def_MQLR}) and (\ref{def_omega-hat}
), respectively.\footnote{
In Lemma \ref{TWFE_Var_U} of the Appendix, we show that\ $\hat{\Greekmath 011B }
_{U,n,T}^{2}$ is positive in finite samples, which implies that the
estimator of $\Greekmath 0121 _{n,T}^{2}$ constructed in (\ref{def_omega-hat}) below
is also positive.} Mirroring Theorem \ref{C_T1}, we have:
\begin{theorem}
\textit{\label{TWFE_F_Vuong} }Suppose that\ \textit{Assumptions \ref{A1},\ }
\ref{A8}, \ref{A9} and \ref{A10} in the Appendix hold. Further suppose that
as $n,T\rightarrow \infty $,
\begin{equation}
\frac{(nT)^{-1/2}\sum_{i\leq n}\sum_{t\leq T}\mathbb{E}[\Greekmath 0122
_{2,i,t}^{2}-\Greekmath 0122 _{1,i,t}^{2}]}{\Greekmath 0121 _{n,T}}\rightarrow c.
\label{TWFE_F_Vuong_1}
\end{equation}
Then we have as $n,T\rightarrow \infty $,
\begin{equation*}
\mathbb{E}[\Greekmath 0127 _{n,T}^{2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 2-\Phi
(z_{1-p/2}-c)-\Phi (z_{1-p/2}+c),
\end{equation*}
and
\begin{equation*}
\mathbb{E}[\Greekmath 0127 _{n,T}^{1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 1-\Phi (z_{1-p}-c),
\end{equation*}
where $\Greekmath 0127 _{n,T}^{2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}$ and $\Greekmath 0127 _{n,T}^{1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}$
follow the identical form as (\ref{C_Vuong_test}).
\end{theorem}
\section{Conclusion\label{sec:conclusion}}
This paper extends the classical \cite{Vuong1989} test to panel data models
with fixed effects, addressing challenges that arise from high-dimensional
nuisance parameters. By modifying the profile likelihood and applying bias
correction techniques, we develop a valid procedure for comparing non-nested
panel data models. Our analysis emphasizes the need for variance adjustments
in the presence of incidental parameter problems and shows how these
modifications support more reliable model selection in high-dimensional
settings.
We contribute to the literature by generalizing the Vuong test to
accommodate grouped heterogeneity in both individual and time effects.
Furthermore, we propose a methodological framework for model selection when
competing specifications define cross-sectional units differently---a common
issue in empirical work. The theoretical results in this paper highlight the
critical role of bias correction in panel data analysis and complement
existing research on non-nested hypothesis testing.
\bigskip