EconBase
← Back to paper

Model Selection in Panel Data Models: A Generalization of the Vuong Test

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,432 characters · 14 sections · 58 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Model Selection in Panel Data Models: A Generalization of the Vuong Test

abstractThis paper generalizes the classical Vuong1989 test to panel data models by employing modified profile likelihoods and the Kullback--Leibler information criterion. Unlike the standard likelihood function, the profile likelihood lacks certain regular properties, making modification necessary. We adopt a generalized panel data framework that incorporates group fixed effects for time and individual pairs, rather than traditional individual fixed effects. Applications of our approach include linear models with non-nested specifications of individual-time effects. JEL Classification: C14, C31, C32 \noindentKeywords: Panel data models, Vuong test, Bias correction, Grouped heterogeneity.

Introduction

With the advent of big data, modern empirical research increasingly adopts more complex models. These models often involve high-dimensional nuisance parameters, which can introduce substantial biases in standard estimators - manifesting as variations of the incidental parameters problem. Classical references include NeymanScott1948 and Nickell1981 among others. Numerous studies have been conducted to address and mitigate these issues. It is now well established that these issues represent a form of higher-order bias, and a variety of methodological approaches have been proposed to overcome them. See ArellanoHahn2007, Arellano-Bonhomme, BesterHansen2009, CARRO2007503, Dhaene-Jochmans, FERNANDEZVAL200971, Fernandez-Val-Weidner-1 , Fernandez-Val-Weidner-2, HahnKuersteiner2002, and HahnNewey2004, for example.

Although the literature is well developed, its generalizations have only recently found application in substantively meaningful but technically complex contexts, such as network-type models. See Bonhomme-Lamadon-Manresa-1, Jochmans-Weidner, bonhomme2025neymanorthogonalizationapproachincidentalparameter. Motivated by these developments, we propose to investigate issues of model specification in nonlinear panel data models. Like most econometric frameworks, panel data models often rely on potentially restrictive assumptions. However, to the best of our knowledge, there is currently no literature that provides a systematic approach for comparing potentially non-nested models in this context. We would like to close this gap in the literature by proposing a panel generalization of the classical Vuong1989 test.

Specifically, we consider a model characterized by the joint log-likelihood of the data $\{z_{i,t}\}_{i\leq n,t\leq T}$ specified as:

equation[equation omitted — 176 chars of source]

where $f\left( \cdot \right) $ is a known function, $\Greekmath 0112 $ represents unknown parameters that are individual/time-invariant, $\Greekmath 010D _{g(i),m(t)}$ denotes the individual and time fixed effect which has a grouping structure determined by cluster/group assignment functions $g(\cdot )$ and $m(\cdot )$ with ranges $\mathcal{G}\equiv \{1,\ldots ,G\}$ and $\mathcal{M}\equiv \{1,\ldots ,M\}$, respectively, and $\Greekmath 011E \equiv (\Greekmath 0112 ^{\top },(\Greekmath 010D _{g,m}^{\top })_{g\in \mathcal{G},m\in \mathcal{M}})^{\top }$ include all unknown parameters specified in the model. For ease of notations, we let $ I_{g}$ for $g\in \mathcal{G}$ and $I_{m}$ for $m\in \mathcal{M}$ denote the partitions of $\{1,\ldots ,n\}$ and $\{1,\ldots ,T\}$ according to the group assignment functions $g(\cdot )$ and $m(\cdot )$, respectively. The pseduo true value $\Greekmath 011E ^{\ast }$ is defined as the maximizer of $\mathbb{E}\left[ L_{n,T}(\Greekmath 011E )\right] $ over the parameter space, and it is often estimated by the maximizer $\hat{\Greekmath 011E }$ of $L_{n,T}(\Greekmath 011E )$ over the same space. Our group structure is inspired the analyses by Bonhomme-Manresa or Bonhomme-Lamadon-Manresa-2, yet it is general enough to include the classical panel models with individual fixed effects.

In practice, different ways of specifying the likelihood function $f\left( \cdot \right) $ and the group assignment functions $g(\cdot )$ and $m(\cdot ) $ lead to different modeling strategies in panel data models, and we often have the situation to compare competing modeling strategies and hope to find the one which is closer (or the closest) to the true data generating process in terms of the Kullback--Leibler (KL) distance. It would be reasonable to compare models through the Vuong1989 test, which tests the null hypothesis:

equation*[equation* omitted — 155 chars of source]

where $L_{j,n,T}(\cdot )$ represents the joint log-likelihood specified by modeling strategy $j$ with $\Greekmath 011E _{j}^{\ast }$ denoting the pseudo true value for $j=1,2$. The alternative hypothesis can be

equation*[equation* omitted — 527 chars of source]

The classical Vuong test employs the quasi-likelihood ratio (QLR), defined as:

equation[equation omitted — 137 chars of source]

to construct the test statistic, where $\hat{\Greekmath 011E }_{j}$ denotes the estimator of $\Greekmath 011E _{j}^{\ast }$ in model $j$. It turns out that in panel models, the log-likelihoods $L_{1,n,T}$ and $L_{2,n,T}$ are biased due to the incidental parameter problems. Panel adaptation of the Vuong test thus requires working with bias-corrected versions of the log likelihoods, which we will call $LM_{1,n,T}^{\ast }$ and $LM_{2,n,T}^{\ast }$. We also show that there is additional complication because the order of magnitude of the standard error of the test statistic (after such modification) can depend on the nature of the hypothesis. Our technical contribution consists of estimation of $LM_{1,n,T}^{\ast }$ and $LM_{2,n,T}^{\ast }$ to make the modification feasible as well as estimation of the standard error while avoiding the uniformity issue.

We extend the classical Vuong1989 test to panel models, contributing to the understanding of model selection complexities in high-dimensional settings. The Vuong test originally compares two non-nested parametric models and formally tests the hypothesis that their KL distances from the true data distribution are equal. Originally designed for parsimoniously parameterized models, this test has been expanded in various studies, see, RiversVuong2002, ChenHongShum, and Shi2015b, for example. In this paper, we demonstrate that the incidental parameters problem in panel data contexts invalidates the classical Vuong test, which compares maximized likelihood ratios between two models. We present a modified test statistic that ensures validity for panel data analysis.

Our modification follows a precedent in LeePhillips2015 and Liao&Shi2020. LeePhillips2015 explored model selection in panel data settings using nested hypothesis testing. Our work complements theirs by introducing approaches for testing non-nested hypotheses. Liao&Shi2020 examined the Vuong test for semiparametric models in a cross-sectional setting, whereas we extend the analysis to a panel data context with fixed effects. Our paper and LeePhillips2015 share the common feature in the sense that both analyses utilize application of profiling to the Kullback--Leibler information criterion (KLIC). It is well-known that profile likelihood does not share the standard properties of the genuine likelihood function. Modification of profile likelihood was discussed a way of overcoming this problem. See ArellanoHahn2007, ArellanoHahn2016, CoxReid1987, DiCiccioStern1993, DiCiccioMartinSternYoung, CoxFergusonReid, and PaceSalvan, for example. Following this tradition, our paper as well as LeePhillips2015 propose a modified profile likelihood that provides approximations to the standard KLIC, and provided a valid method of model selection in panel data models with individual fixed effects. Interestingly, we find some close connection between the modification in Liao&Shi2020 and the panel type modification following ArellanoHahn2007, ArellanoHahn2016, and LeePhillips2015.

Recently developed methods in panel data allow applied researchers to choose among a broad range of alternative models for separable, interactive, or discretely supported unobserved heterogeneity across units and time. It is widely understood that parameter estimates are often sensitive to the chosen specification, however there is often no clear theoretical guidance on how to choose among alternative models. Our paper makes further contribution in this regard by considering the cases where the definition of “individuals” may not be agreed across different models. In empirical practice, the group/cluster structure may not be obviously agreed among researchers. The competing group/cluster structures may be nested in the sense that one structure is “finer” than the other one, but they may be non-nested. In this situation, extension of the insights in previous papers such as LeePhillips2015 or Liao&Shi2020 may not be straightforward. We establish that the classical Vuong test is valid after some modification. In doing so, we also make a technical contribution that may be of some independent interest. FernandezValWeidener2016 analyzed the asymptotic distribution of the (pseudo) MLE in nonlinear panel data models with individual and time effects. Under the assumption that the objective function is globally concave in all the parameters, they established the asymptotic bias formula. Our paper contains a result\footnote{ See Theorem (ref) in the Online Appendix and Theorem (ref) in Appendix (ref), along with the related discussion therein, for further details.} that relaxes this concavity assumption, although it is achieved at the cost of restricting the group structure. Therefore, our result can be argued to complement FernandezValWeidener2016.

The remainder of the paper is organized as follows. In Section (ref),\ we introduce the Vuong test for comparing classical panel data models that allow for potential group heterogeneity. Section (ref) presents a test for comparing the two-way fixed effects (TWFE) model against a heterogeneous time fixed effects model. Section (ref) concludes.\ The proofs of the main results, along with auxiliary lemmas, are provided in the Appendix. Additionally, Appendix (ref) develops a general asymptotic theory for estimators in panel models with individual and time fixed effects, as well as the asymptotic distribution of the statistic $QLR_{n,T}$ in this broader setting -- results that are of independent interest.\footnote{ The technical proofs for this general theory, together with the proofs of the auxiliary lemmas used to establish the main results in Sections (ref) and (ref), are provided in the Online Appendix.}

The following notation will be adopted throughout the paper.\ We use $K$ to denote a generic strictly positive constant that may vary from one instance to another but remains independent of the panel dimensions $n$ {and }$T$.\ We adopt the convention that a summation over an empty set equals zero. We use $a\equiv b$ to indicate that $a$ is defined as $b$. For real numbers $ a_{1},\ldots ,a_{m}$, $(a_{j})_{j\leq m}\equiv (a_{1},\ldots ,a_{m})^{\top }$ . For any matrix $A$, $A^{\top }$\ denotes the transpose of $A$,\ and $ \left\Vert A\right\Vert $ denotes the Euclidean norm of $A$. For any doubly indexed sequence $a_{i,t}$ (where $i=1,\ldots ,n$ and $t=1,\ldots ,T$), we define\ $\bar{a}_{i}\equiv T^{-1}\sum_{t\leq T}a_{i,t}$, $\bar{a}_{t}\equiv n^{-1}\sum_{i\leq n}a_{i,t}$ and $\bar{a}\equiv (nT)^{-1}\sum_{t\leq T}\sum_{i\leq n}a_{i,t}$. The summation $\sum_{i^{\prime }\neq i}$ is taken over all $i^{\prime }$ except $i$, which means $\sum_{i^{\prime }\neq i}a_{i^{\prime },t}=\sum_{i^{\prime }=1}^{i-1}a_{i^{\prime },t}+\sum_{i^{\prime }=i+1}^{n}a_{i^{\prime },t}$.\ For two sequences of positive numbers $a_{n}$ and $b_{n}$, we write $a_{n}\succ b_{n}$ if $ a_{n}\geq c_{n}b_{n}$ for some strictly positive sequence $c_{n}\rightarrow \infty $. The notation $\left\Vert \cdot \right\Vert _{p}$ denotes the $ L_{p} $-norm. Finally, for any function $g(\cdot )$, define $\mathbb{E}_{T} \left[ g\left( z_{i,t}\right) \right] \equiv T^{-1}\sum_{t\leq T}\mathbb{E} \left[ g\left( z_{i,t}\right) \right] $ and $\widehat{\mathbb{E}}_{T}\left[ g\left( z_{i,t}\right) \right] \equiv T^{-1}\sum_{t\leq T}g\left( z_{i,t}\right) $.

Vuong Test for Classical Panel Models\

In this section, we investigate the Vuong test for panel models with time-invariant individual effects.\footnote{ A simple version of the test appeared in hahn-liu, which later became a chapter in Liu's UCLA dissertation. Our current paper substantially extends and generalizes the earlier simple result.} With a dataset\ $\left\{ z_{i,t}\right\} _{i\leq n,t\leq T}$, we compare two models, denoted as model 1 and model 2 respectively. Model 1 represents the classical panel model with individual effects. In contrast, model 2 may specify a different joint likelihood function and/or adopt a distinct clustering or grouping structure for the individual effects.

For each model $j$ ($j=1,2$), the joint log-likelihood function simplifies from the general specification in ((ref)) to the following form:

equation*[equation* omitted — 284 chars of source]

where\ $\Greekmath 011E _{j}\equiv (\Greekmath 0112 _{j}^{\top },(\Greekmath 010D _{j,g})_{g\in \mathcal{G }_{j}}^{\top })^{\top }$, $\Greekmath 0120 _{j}\left( z_{i,t};\Greekmath 011E _{j,g}\right) \equiv \log f_{j}(z_{i,t};\Greekmath 0112 _{j},\Greekmath 010D _{j,g})$, $\Greekmath 011E _{j,g}\equiv (\Greekmath 0112 _{j}^{\top },\Greekmath 010D _{j,g})^{\top }$,\ $\mathcal{G}_{j}\equiv \{1,\ldots ,G_{j}\}$ defines the range of the grouping function $g_{j}(\cdot )$, and $ I_{j,g}$ ($g\in \mathcal{G}_{j}$) represents the partition of $\{1,\ldots ,n\}$ induced by $g_{j}(\cdot )$. For model 1, it follows that $\mathcal{G} _{1}=\{1,\ldots ,n\}$, implying that each individual has its own group. For model 2, however, the partitions $\{I_{2,g}\}_{g\in \mathcal{G}_{2}}$ may form any arbitrary grouping of $\{1,\ldots ,n\}$\ allowing for greater flexibility in the specification of individual effects.\footnote{ This means that model 1 is nested within model 2 in terms of the grouping structure, although the likelihood specifications may differ. In many applications, the grouping structure is expected to be identical across models.}

As this section is relatively long, we provide a roadmap to facilitate reading. Sections (ref) - (ref) provides the theoretical foundation of the infeasible test, while Sections (ref) - (ref) presents the actual algorithm for implementation of the feasible version of the test. In Section (ref), we provide an intuitive overview of the adjustment to the maximized joint likelihood and discuss its importance for model comparison using the Vuong test in the panel data models considered in this paper. In Section (ref), we present the log likelihood of each model. In Section (ref), we present an infeasible adjustment that corrects for the bias in the likelihood, and in Section (ref), we present a theoretical analysis for the asymptotic variance of the modified likelihood. Asymptotic properties of the (infeasible) modified test statistic is presented in Theorem (ref). In Sections (ref) and (ref), we present actual algorithms for estimation of the modified log likelihoods and the asymptotic variances. Section (ref) presents Theorem (ref) that characterizes the asymptotic properties of the feasible test statistic.

Incidental Parameter Problem in Panel Data Models

To streamline notation and improve clarity, we focus here on a simplified version of ((ref)) that includes only individual fixed effects, specified as:

equation[equation omitted — 261 chars of source]

where $\Greekmath 0120 \left( z_{i,t};\Greekmath 011E _{i}\right) \equiv \log f\left( z_{i,t};\Greekmath 0112 ,\Greekmath 010D _{i}\right) $, $\Greekmath 011E _{i}\equiv (\Greekmath 0112 ^{\top },\Greekmath 010D _{i})^{\top }$\ and $\Greekmath 011E \equiv (\Greekmath 0112 ^{\top },(\Greekmath 010D _{i})_{i\leq n}^{\top })^{\top }$.\ Since QLR statistic defined in ((ref)) involves the maximized joint likelihood\ $L_{n,T}(\hat{\Greekmath 011E })$, understanding the QLR's properties requires investigating the estimation error of $L_{n,T}(\hat{\Greekmath 011E })$.

Let $\Greekmath 011E ^{\ast }\equiv (\Greekmath 0112 ^{\ast \top },(\Greekmath 010D _{i}^{\ast })_{i\leq n}^{\top })^{\top }$ denote the pseudo-true parameter which maximizes $ \mathbb{E}\left[ L_{n,T}(\Greekmath 011E )\right] $.\footnote{ Potential dependence of the pseudo-true parameters $\Greekmath 011E ^{\ast }$ on the sample size is not made explicit for notational simplicity throughout the paper.} Define:

equation[equation omitted — 669 chars of source]

where $\Greekmath 011E _{i}^{\ast }\equiv (\Greekmath 0112 ^{\ast \top },\Greekmath 010D _{i}^{\ast })^{\top }$. It can be shown\footnote{ See Theorem (ref) in the Appendix, specialized to the case $\mathcal{G} =\{1,\ldots ,n\}$ and $\mathcal{M}=\{1\}$, so that $G=n$ and $M=1$}\ that under the asymptotics with $nT^{-1}\rightarrow \Greekmath 011A $\ where $\Greekmath 011A \in (0,\infty )$:

equation[equation omitted — 294 chars of source]

where $L_{n,T}(\Greekmath 011E ^{\ast })=\sum_{i\leq n}\sum_{t\leq T}\Greekmath 0120 \left( z_{i,t}\right) $, and for any $i\leq n$,

equation[equation omitted — 599 chars of source]

The expansion in ((ref)) shows that the estimation error of $ L_{n,T}(\hat{\Greekmath 011E })$ has two leading components. The first component, $ L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $, results from estimating\ $\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $ at the given pseudo-true parameter $\Greekmath 011E ^{\ast }$. When scaled by a factor of $(nT)^{-1/2}$, this component converges in distribution to a normal distribution under some regularity conditions. The second component, $ 2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$, arises from estimating the unknown pseudo-true parameter $\Greekmath 011E ^{\ast }$ and is determined by the estimation error of the incidental parameters $(\Greekmath 010D _{i}^{\ast })_{i\leq n}$.\footnote{ Since the dimension of $\Greekmath 0112 ^{\ast }$ is fixed, the contribution of its estimation error to $L_{n,T}(\hat{\Greekmath 011E })$ is asymptotically negligible compared with the two leading components, and is therefore included in the remainder term of ((ref)), represented as $o_{p}(n^{1/2})$.}

From the expansion in ((ref)), it is evident that in panel data models, the $L_{n,T}(\hat{\Greekmath 011E })$ has a first-order bias, stemming from $ 2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$. To evaluate this bias explicitly, suppose that $\{\tilde{\Greekmath 0120 }_{\Greekmath 010D }\left( z_{i,t}\right) \}_{i,t}$ are i.i.d. across $t\leq T$ for any $i$. Since $\mathbb{E}[\tilde{ \Greekmath 0120 }_{\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ]=0$, it follows that $\mathbb{E }[2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}]=2^{-1}\sum_{i\leq n}\Greekmath 011B _{\Greekmath 010D ,i}^{2}$, where $\Greekmath 011B _{\Greekmath 010D ,i}^{2}\equiv \mathbb{E}[ \tilde{\Greekmath 0120 }_{\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ^{2}]$. Assuming that $ \Greekmath 011B _{\Greekmath 010D ,i}^{2}$ is bounded above and below, the bias from the mean of $2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$ is of the same order as $L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $. This provides the explanation why the maximized joint likelihood $L_{n,T}( \hat{\Greekmath 011E })$ should be modified with a bias correction based on the mean of\ $2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$.

The above discussion demonstrates that the maximized joint likelihood\ $ L_{n,T}(\hat{\Greekmath 011E })$ requires a first-order bias correction in the form of a consistent estimator of $2^{-1}\sum_{i\leq n}\mathbb{E}[\tilde{\Psi} _{\Greekmath 010D ,i}^{2}]$. As a result, the QLR statistic in the classical Vuong test must also be adjusted with a bias correction to ensure valid inference. This bias correction for the QLR statistic is given by the difference in the bias of $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$ ($j=1,2$) from the two competing models.

There is an additional complication. Depending on whether the models are nested or overlapping, the first components, $L_{j,n,T}(\Greekmath 011E _{j}^{\ast })- \mathbb{E}[L_{j,n,T}(\Greekmath 011E _{j}^{\ast})]$, from the two models may cancel out. See Vuong1989 for discussion in the cross section context. In such cases, the second components, denoted as $2^{-1}\sum_{i\leq n}\tilde{\Psi } _{j,\Greekmath 010D ,i}^{2}$ ($j=1,2$),\ can become the primary drivers of the asymptotic distribution of the bias-corrected QLR statistic.\footnote{ In the classical nested case, where the functional form of the likelihood $f$ is identical across the two competing models, the second component of the test statistic cancels out. In this setting, bias correction of the estimator itself, along with further refinements of ((ref)), is needed. This case was previously studied by LeePhillips2015, and we do not pursue it further in this paper.}\ Consequently, the estimation error of the incidental parameters influences not only the bias but also the variance of the QLR statistic.

Quasi-likelihood Estimator

The pseudo true parameter $\Greekmath 011E _{j}^{\ast }$ for model $j$ is estimated as the maximizer $\hat{\Greekmath 011E }_{j}$ of $L_{j,n,T}(\Greekmath 011E _{j})$ where $\Greekmath 011E _{j}\equiv (\Greekmath 0112 _{j}^{\top },(\Greekmath 010D _{j,g})_{g\in \mathcal{G}_{j}}^{\top })^{\top }$. The estimator\ $\hat{\Greekmath 011E }_{j}$\ satisfies the following first-order conditions:

equation[equation omitted — 205 chars of source]

and for any $g\in \mathcal{G}_{j}$,

equation[equation omitted — 147 chars of source]

where $\Greekmath 011E _{j,g}\equiv (\Greekmath 0112 _{j}^{\top },\Greekmath 010D _{j,g})^{\top }$\ and $ \Greekmath 0120 _{j,a}(z_{i,t};\Greekmath 011E _{j,g})\equiv \partial \Greekmath 0120 _{j}(z_{i,t};\Greekmath 011E _{j,g})/\partial a$ for $a\in \{\Greekmath 0112 _{j},\Greekmath 010D _{j,g}\}$.\ The estimator $\hat{\Greekmath 011E }_{j}$ is subsequently used to compute the maximized joint likelihood $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$\ for model $j$.

Following the discussion in the previous subsection, the maximized joint quasi-likelihood for each model must be adjusted to account for the respective first-order bias, which can be derived from the asymptotic expansion of $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$. To present this expansion in a more general setting than previously considered, we first update the notations introduced earlier. Let $\Greekmath 0120 _{j}\left( z_{i,t}\right) $, $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $, $\Greekmath 0120 _{j,\Greekmath 010D \Greekmath 010D }\left( z_{i,t}\right) $ and $\tilde{\Psi}_{j,\Greekmath 010D ,i}$ represent the counterparts to their definitions in ((ref)) and ((ref)), with $\Greekmath 011E _{j,i}\equiv \Greekmath 011E _{j,g}$,

equation*[equation* omitted — 615 chars of source]

for any $i\in I_{j,g}$ and any $g\in \mathcal{G}_{j}$.\footnote{ The demeaned version of $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $, i.e., $ \Greekmath 0120 _{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) $ is introduced here because the first-order condition for $\Greekmath 011E _{j}^{\ast }$ only ensures that $ \sum_{i\in I_{j,g}}\sum_{t\leq T}\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. This condition does not guarantee that $\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. If $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $ is stationry across $t$ for each $i$, the first-order condition for $\Greekmath 011E _{j}^{\ast }$ is simplified to $\sum_{i\in I_{j,g}} \mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. In this case, $ \mathbb{E}[\Greekmath 0120 _{1,\Greekmath 010D }\left( z_{i,t}\right) ]=0$ holds for model 1, but is still not guaranteed for model 2.}

Infeasible Modified QLR Statistic - Bias

With the updated notations, we can obtain\footnote{ See Theorem (ref) together with the decomposition in ((ref)) in the Appendix.} the following expansion:

equation[equation omitted — 165 chars of source]

where the components are defined\footnote{ The $\tilde{\Psi}_{j,\Greekmath 010D ,i}$ is not the derivative of $\tilde{\Psi} _{j,i} $, as can be seen from its definition ((ref)).} as:

equation*[equation* omitted — 543 chars of source]

It is reasonable to expect $\mathbb{E}[\tilde{U}_{j,i}]=0$ for $j=1,2$ and for any $i\leq n$,\footnote{ For example, it is implied by Assumption (ref)\ in the Appendix.} \ which implies that the first-order bias arises from the mean of $\tilde{V} _{j,i}$. To account for this bias, we define the (infeasible) modified maximized joint likelihood as:

equation[equation omitted — 174 chars of source]

Using this expression, we define the (infeasible) modified QLR statistic, whose asymptotic expansion follows from ((ref)):

equation[equation omitted — 272 chars of source]

where $\tilde{\Psi}_{i}\equiv \tilde{\Psi}_{1,i}-\tilde{\Psi}_{2,i}$, $ \tilde{V}_{i}^{\ast }\equiv \tilde{V}_{1,i}-\tilde{V}_{2,i}-\mathbb{E}[ \tilde{V}_{1,i}-\tilde{V}_{2,i}]$ and $U_{i}\equiv \tilde{U}_{1,i}-\tilde{U} _{2,i}$.

Infeasible Modified QLR Statistic - Asymptotic Distribution

After bias correction, the main remaining components arising from the estimation errors of the incidental parameters, i.e., $\tilde{V}_{i}^{\ast }+ \tilde{U}_{i}$ primarily affect the variance of the modified QLR statistic\ $ MQLR_{n,T}^{\ast}$.\ This effect becomes particularly significant when the two models are overlapping. In such cases, the first component on the right-hand side of ((ref)), i.e., $n^{-1/2}\sum_{i\leq n}\tilde{\Psi }_{i}$, can become arbitrarily small or even vanish, leaving the asymptotic distribution of $MQLR_{n,T}^{\ast}$ predominantly determined by the estimation errors of the incidental parameters, specifically $ n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast}+\tilde{U}_{i})$.

To facilitate the discussion of the variance and asymptotic distribution of $ MQLR_{n,T}^{\ast}$, we follow the literature (see, e.g., Vuong1989, Shi2015b and Liao&Shi2020) and formally define the relationships between models in the panel setting below.

definition(i) Model 1 and model 2 are strictly non-nested if there do not exist $\Greekmath 011E _{1,i}\in \Phi _{1}$ and $\Greekmath 011E _{2,i}\in \Phi _{12}$ such that \begin{equation} \Greekmath 0120 _{1}\left( z;\Greekmath 011E _{1,i}\right) =\Greekmath 0120 _{2}\left( z;\Greekmath 011E _{2,i}\right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \end{equation} for any $i$ and any $z$ in the support of $z_{i,t}$; (ii) the two models are overlapping if they are not strictly non-nested; (iii) model 1 and model 2 are said to be nested if, for any $\Greekmath 011E _{j,i}\in \Phi _{j}$ , there exists a $\Greekmath 011E _{j^{\prime },i}\in \Phi _{j^{\prime }}$ (where $ j,j^{\prime }=1,2$ and $j\neq j^{\prime }$) such that ((ref)) holds.

When the two models are strictly non-nested, the variance of $ n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$ is bounded away from zero. In this case, the component $n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast }+\tilde{U} _{i})$ on the right-hand side of ((ref)) becomes asymptotically negligible,\ as it is of order $O_{p}(T^{-1/2})$ under certain regularity conditions.\footnote{ This order follows from an approximation of\ $n^{-1}\sum_{i\leq n}(\mathrm{ Var}(\tilde{V}_{i})+\mathrm{Var}(\tilde{U}_{i}))$ (see ((ref)) in the Online Appendix, which appears in the proof of Lemma (ref)), together with an upper bound on the variance of $\tilde{\Greekmath 0120 }_{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) $ for $j=1,2$.}\ Consequently, the asymptotic distribution of $MQLR_{n,T}^{\ast }$ is primarily determined by $n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$. Conversely, when the two models are nested, $\tilde{\Psi}_{i}=0$ under the null hypothesis. In this scenario, the asymptotic distribution of $ MQLR_{n,T}^{\ast }$ is determined by\ $n^{-1/2}\sum_{i\leq n}(\tilde{V} _{i}^{\ast }+\tilde{U}_{i})$. As a result, deriving the asymptotic distribution of $MQLR_{n,T}^{\ast }$ is relatively straightforward, when the models are known to be either strictly non-nested or nested. In both cases, the dominating term on the right-hand side of ((ref)) is identifiable, and the asymptotic distribution of $MQLR_{n,T}^{\ast }$ can be deduced from that of the dominating term.

The asymptotic distribution of $QLR_{n,T}^{\ast }$ is more challenging to derive when the two models are overlapping. Because in this case, it is in general unclear which components between $n^{-1/2}\sum_{i\leq n}\tilde{\Psi} _{i}$ and $n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast }+\tilde{U}_{i})$ could be the dominating term, and both could be equally important. To address this issue, we follow Liao&Shi2020 to consider both terms together when deriving the asymptotic distribution of $MQLR_{n,T}^{\ast }$.

theoremUnder Assumptions (ref) and (ref)\ in the Appendix, we have \begin{equation} \frac{MQLR_{n,T}^{\ast }-\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}}\rightarrow _{d}N(0,1)\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \end{equation} where \begin{equation} \Greekmath 0121 _{n,T}^{2}\equiv n^{-1}\sum_{i\leq n}\mathrm{Var}(\tilde{\Psi}_{i}+ \tilde{V}_{i}+\tilde{U}_{i})\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ and \ }\overline{QLR}_{n,T}\equiv (nT)^{-1/2}\mathbb{E}[L_{1,n,T}(\Greekmath 011E _{1}^{\ast })-L_{1,n,T}(\Greekmath 011E _{2}^{\ast })]. \end{equation}
remarkTheorem (ref) provides the asymptotic distribution of $MQLR_{n,T}^{\ast }$ under both the null and alternative hypotheses. Since $\overline{QLR} _{n,T}=0$ under the null hypothesis, we can construct a test for the null by using a consistent estimator of $\Greekmath 0121 _{n,T}$, along with a feasible version of the modified QLR statistic. This feasible statistic replaces the unknown bias correction terms in $MQLR_{n,T}^{\ast }$ with their consistent estimators.\footnote{ We impose Assumption (ref) to simplify the estimation of the bias correction and the variance of $MQLR_{n,T}^{\ast }$.\ All assumptions are collected in the Appendix.}

Estimation of Bias

We first discuss the estimation of the bias correction terms. Under Assumption (ref)\ in the Appendix, the definition of $\tilde{V}_{j,i}$ implies that:

equation[equation omitted — 422 chars of source]

This motivates a straightforward plug-in estimator of the bias correction term for model $j$, relying on a consistent estimator of $\Greekmath 011B _{j,\Greekmath 010D ,i}^{2}$\ for any $i\in I_{j,g}$ and any $g\in \mathcal{G}_{j}$:

equation[equation omitted — 350 chars of source]

where

equation[equation omitted — 479 chars of source]

The bias correction term for model $j$ is then estimated as

equation[equation omitted — 189 chars of source]

Using this, the feasible modified QLR statistic is defined as:

equation[equation omitted — 154 chars of source]

where

equation*[equation* omitted — 138 chars of source]

denotes the modified maximized joint likelihood for model $j$.

theoremUnder Assumptions (ref), (ref) and (ref) in the Appendix, we have \begin{equation*} \frac{R_{j,n,T}(\hat{\Greekmath 011E }_{j})-T^{1/2}\sum_{i\leq n}\mathbb{E}[\tilde{V} _{j,i}]}{(nT)^{1/2}\Greekmath 0121 _{n,T}}=O_{p}(T^{-1/2}). \end{equation*}

Theorem (ref) establishes the consistency of $R_{j,n,T}(\hat{\Greekmath 011E }_{j})$ as an estimator of the bias of the maximized joint likelihood $L_{j,n,T}( \hat{\Greekmath 011E }_{j})$ in model $j$. Combined with the result in ((ref)), this directly leads to the asymptotic normality of the modified QLR statistic $QLR_{n,T}$:

equation[equation omitted — 184 chars of source]

To apply this result for inference under the null hypothesis, it remains to construct a consistent estimator for $\Greekmath 0121 _{n,T}^{2}$, which we now proceed to address.

Estimation of Variance\

A consistent estimator of $\Greekmath 0121 _{n,T}^{2}$, which was defined in ((ref)), can be constructed using the sample analog of the variance of $ n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$, and a variance correction term to account for estimation errors of the incidental parameters. Define $\Greekmath 011B _{n,T}^{2}\equiv n^{-1}\sum_{i\leq n}\mathrm{Var}(\tilde{\Psi}_{i})$ and $ \Delta \Greekmath 0120 (z_{i,t},\Greekmath 011E _{i})\equiv \Greekmath 0120 _{1}(z_{i,t};\Greekmath 011E _{1,i})-\Greekmath 0120 _{2}(z_{i,t};\Greekmath 011E _{2,i})$. Under Assumption (ref), a natural estimator for $\Greekmath 011B _{n,T}^{2}$ is given by:

equation[equation omitted — 214 chars of source]

The theoretical property of $\hat{\Greekmath 011B }_{n,T}^{2}$ is provided in the lemma below.\footnote{ The sample variance\ $\hat{\Greekmath 011B }_{n,T}^{2}$ will be used below as a building block for variance estimation and therefore requires $\{\Delta \Greekmath 0120 (z_{i,t},\Greekmath 011E _{i}^{\ast })\}_{t}$ to be serially uncorrelated, as assumed in Assumption (ref). If serial correlation is present, an autocorrelation-consistent estimator is required. Since variance estimation is mainly needed under the null for size control, Assumption (ref) is imposed only under the null; Lemma (ref) in the Appendix shows that consistency of the test does not rely on this assumption.}

theorem\ Under Assumptions (ref), (ref) and (ref) in the Appendix, we have: \begin{equation} \frac{\hat{\Greekmath 011B }_{n,T}^{2}-(\Greekmath 011B _{n,T}^{2}+2\Greekmath 011B _{S,n,T}^{2})}{ \Greekmath 0121 _{n,T}^{2}}=O_{p}(T^{-1/2}), \end{equation} where \begin{equation} \Greekmath 011B _{S,n,T}^{2}\equiv (2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}\sum_{i\in I_{2,g}}\left( \Greekmath 011B _{1,\Greekmath 010D ,i}^{4}+n_{2,g}^{-2}\Greekmath 011B _{2,\Greekmath 010D ,i}^{2}\sum_{i^{\prime }\in I_{2,g}}s_{2,\Greekmath 010D ,i^{\prime }}^{2}-2n_{2,g}^{-1}\Greekmath 011B _{12,\Greekmath 010D ,i}^{2}\right) , \end{equation} $s_{2,\Greekmath 010D ,i}^{2}\equiv \mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{2,\Greekmath 010D }\left( z_{i,t}\right) ^{2}]$ and $\Greekmath 011B _{12,\Greekmath 010D ,i}\equiv \mathbb{E}_{T}[ \tilde{\Greekmath 0120 }_{1,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) \tilde{\Greekmath 0120 }_{2,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ]$ for any $i\in I_{2,g}$ and any $g\in \mathcal{G}_{2}$.\footnote{ Since $\mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{1,\Greekmath 010D }(z_{i,t})]=0$ by the first-order condition for $\Greekmath 011E _{1}^{\ast }$, we have $\Greekmath 011B _{12,\Greekmath 010D ,i}=\mathbb{E}_{T}[\tilde{\Greekmath 0120 }_{1,\Greekmath 010D }\left( z_{i,t}\right) \tilde{\Greekmath 0120 } _{2,\Greekmath 010D }\left( z_{i,t}\right) ]$ for any $i\leq n$.}

Theorem (ref) shows that $\hat{\Greekmath 011B }_{n,T}^{2}$ tends to overestimate the true variance\ $\Greekmath 011B _{n,T}^{2}$ by $2\Greekmath 011B _{S,n,T}^{2}$. It is reasonable to expect the latter to be of order $O(T^{-1})$, which is formalized by Assumptions (ref)(iii) and (ref). On the other hand,\ it can be shown\footnote{ See Lemma (ref) in the Online Appendix.} that

equation[equation omitted — 169 chars of source]

where

equation[equation omitted — 388 chars of source]

which represents the leading term in the variance of $n^{-1/2}\sum_{i\leq n}( \tilde{V}_{i}^{\ast }+\tilde{U}_{i})$.

Comparing their expressions in ((ref)) and ((ref)), it is evident that $\Greekmath 011B _{S,n,T}^{2}\geq \Greekmath 011B _{U,n,T}^{2}$ in general, with the inequality becoming strict when $\mathbb{E}[\tilde{\Greekmath 0120 }_{j,\Greekmath 010D }\left( z_{i,t}\right) ]\neq 0$. From ((ref)) and ((ref)), it follows that $\hat{\Greekmath 011B }_{n,T}^{2}$ is a consistent estimator of $\Greekmath 0121 _{n,T}^{2}$ when $\Greekmath 011B _{n,T}^{2}$ dominates $\Greekmath 011B _{S,n,T}^{2}$. Since $ \Greekmath 011B _{S,n,T}^{2}=O(T^{-1})$, this dominance is guaranteed when the two models are strictly non-nested, as $\Greekmath 011B _{n,T}^{2}$ is bounded away from zero in such cases. Conversely, if $\Greekmath 011B _{n,T}^{2}$ is of the same or even smaller order than\ $\Greekmath 011B _{S,n,T}^{2}$, a scenario that arises when the two models are (nearly) nested or overlapping, the result in ((ref)) implies that $\hat{\Greekmath 011B }_{n,T}^{2}$ overestimates $\Greekmath 0121 _{n,T}^{2}$ by $2\Greekmath 011B _{S,n,T}^{2}-\Greekmath 011B _{U,n,T}^{2}$, and becomes an inconsistent estimator of $\Greekmath 0121 _{n,T}^{2}$.

The over-estimation issue in $\hat{\Greekmath 011B }_{n,T}^{2}$ can be adjusted through consistent estimators of $\Greekmath 011B _{U,n,T}^{2}$ and $\Greekmath 011B _{S,n,T}^{2}$. From its expression in ((ref)), $\Greekmath 011B _{U,n,T}^{2}$ can be estimated by replacing $\Greekmath 011B _{j,\Greekmath 010D ,i}^{2}$ with $\hat{\Greekmath 011B } _{j,\Greekmath 010D ,i}^{2}$ defined in ((ref)), and replacing $\Greekmath 011B _{12,\Greekmath 010D ,i}^{2}$\ by

equation*[equation* omitted — 386 chars of source]

Specifically, the estimator of $\Greekmath 011B _{U,n,T}^{2}$ is defined as

equation[equation omitted — 418 chars of source]

Similarly, we define the estimator of $\Greekmath 011B _{S,n,T}^{2}$ as

equation[equation omitted — 401 chars of source]

Here, for any $i\in I_{2,g}$ and any $g\in \mathcal{G}_{2}$:

equation[equation omitted — 253 chars of source]

Thus, $\Greekmath 0121 _{n,T}^{2}$ can be consistently estimated in a general setting as $\hat{\Greekmath 011B }_{n,T}^{2}+\hat{\Greekmath 011B }_{U,n,T}^{2}-2\hat{\Greekmath 011B } _{S,n,T}^{2} $, as indicated by Theorem (ref) below.

theorem\ Under Assumptions (ref), (ref) and (ref) in the Appendix, we have: \begin{equation} \frac{\hat{\Greekmath 011B }_{U,n,T}^{2}-\Greekmath 011B _{U,n,T}^{2}}{\Greekmath 0121 _{n,T}^{2}} =O_{p}(n^{-1/2})\ \ \ \ \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{and \ \ }\frac{\hat{\Greekmath 011B }_{S,n,T}^{2}- \Greekmath 011B _{S,n,T}^{2}}{\Greekmath 0121 _{n,T}^{2}}=O_{p}(n^{-1/2}). \end{equation}

One drawback of using $\hat{\Greekmath 011B }_{n,T}^{2}-2\hat{\Greekmath 011B }_{S,n,T}^{2}+\hat{ \Greekmath 011B }_{U,n,T}^{2}$ as a variance estimator is that it may take negative values in finite samples. To address this issue, we follow Liao&Shi2020 and propose the following hybrid variance estimator:

equation[equation omitted — 229 chars of source]
remarkBy definition, $\hat{\Greekmath 011B }_{U,n,T}^{2}$ can be expressed as \begin{align} \hat{\Greekmath 011B }_{U,n,T}^{2}& =(2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}\sum_{i\in I_{2,g}}(\hat{\Greekmath 011B }_{1,\Greekmath 010D ,i}^{4}-2n_{2,g}^{-1}\hat{\Greekmath 011B }_{12,\Greekmath 010D ,i}^{2}+n_{2,g}^{-2}\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i}^{4}) \notag \\ & +(2nT)^{-1}\sum_{g\in \mathcal{G}_{2}}n_{2,g}^{-2}\sum_{i\in I_{2,g}}\sum_{i^{\prime }\in I_{2,g},i^{\prime }\neq i}\hat{\Greekmath 011B } _{2,\Greekmath 010D ,i}^{2}\hat{\Greekmath 011B }_{2,\Greekmath 010D ,i^{\prime }}^{2}. \end{align} Since\ $\hat{\Greekmath 011B }_{12,\Greekmath 010D ,i}^{2}\leq \hat{\Greekmath 011B }_{1,\Greekmath 010D ,i}^{2} \hat{\Greekmath 011B }_{2,\Greekmath 010D ,i}^{2}$ for any $i$, the first term on the right-hand side of ((ref)) is non-negative. Additionally, the second term after the equality in ((ref)) is also non-negative by definition. Therefore, $\hat{\Greekmath 011B }_{U,n,T}^{2}$ is guaranteed to be non-negative, and by construction, $\hat{\Greekmath 0121 }_{n,T}^{2}$ is also non-negative. This ensures the practical reliability of $\hat{\Greekmath 0121 } _{n,T}^{2}$ as a variance estimator.

Feasible Modified QLR Statistic and Generalized Vuong Test

Given the variance estimator $\hat{\Greekmath 0121 }_{n,T}^{2}$, the two-sided and one-sided Vuong tests are defined as:

eqnarray[eqnarray omitted — 497 chars of source]

respectively, where $z_{1-p}$ is the $1-p$ quantile of the standard normal distribution, and $p\in (0,1)$ denotes the significance level.

theorem\ Suppose that Assumptions (ref), (ref) and (ref) in the Appendix hold. Further suppose that as $n,T\rightarrow \infty $, \begin{equation} \frac{\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}}\rightarrow c, \end{equation} for some constant $c\ $in the extended real line. Then as $n,T\rightarrow \infty $, \begin{equation} \mathbb{E}[\Greekmath 0127 _{n,T}^{2\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 2-\Phi (z_{1-p/2}-c)-\Phi (z_{1-p/2}+c), \end{equation} and \begin{equation} \mathbb{E}[\Greekmath 0127 _{n,T}^{1\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 1-\Phi (z_{1-p}-c), \end{equation} where $\Phi (\cdot )$ denotes the cumulative distribution function of the standard normal.
remarkTheorem (ref) establishes both the size and power properties of the Vuong test defined in ((ref)). Under the null hypothesis where $ \overline{QLR}_{n,T}=0$, ((ref)) holds with $c=0$. In this scenario, ( (ref)) and ((ref)) imply that \begin{equation} \mathbb{E}[\Greekmath 0127 _{n,T}^{2\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 2(1-\Phi (z_{1-p/2}))=p\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ \ \ \ and \ \ \ }\mathbb{E}[\Greekmath 0127 _{n,T}^{1\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side }}(p)]\rightarrow p, \end{equation} which establishes the size control over data generating processes in which the two models being compared are either nested or overlapping. Furthermore, ((ref)) demonstrates that the two-sided Vuong test has power against local alternatives with $\left\vert c\right\vert \in (0,\infty )$, and is consistent against any fixed alternatives, while the one-sided Vuong test share the similar power properties with $c\in (0,\infty )$.

Testing Heterogeneous Time Effects vs. TWFE

The general result in the previous section was based on Theorems (ref) -- (ref) in Appendix (ref). As discussed in the introduction, the result established in Appendix (ref) is a technical contribution that relaxes the concavity assumption in FernandezValWeidener2016, but it was done at the cost of restricting the group structure. In this section, we consider potentially more complicated group structure but we do so in the context of the linear model, which satisfies the concavity assumption. Because of substantive economic interest associated with the linear model, in particular, the literature initiated by Bonhomme-Manresa, we believe that the linear model may deserve a special attention.

Bonhomme-Manresa have recently proposed a new model that extends beyond the scope of traditional panel data models. In light of its computational demands, one may naturally question whether a standard TWFE specification could offer a practical alternative. Obviously the two models are not nested, so it is of interest to find a Vuong test for this comparison.

We consider the linear panel model

equation*[equation* omitted — 112 chars of source]

where $y_{i,t}$ is the dependent variable, $x_{i,t}$ are the observed regressors, and $\Greekmath 0122 _{i,t}$ is i.i.d. across $i$ and $t$ and with mean zero and variance $\Greekmath 011B ^{2}$. Here, the $g(\cdot )$ and $m(\cdot )$ are known cluster/group assignment functions with range $\mathcal{G} =\{1,\ldots ,G\}$ and $\mathcal{M}=\{1,\ldots ,M\}$ respectively, and $ \Greekmath 010D _{g(i),m(t)}$ denotes the unknown fixed effect.\footnote{ The $\Greekmath 010D _{g(i),m(t)}$ takes different forms under various model specifications. For example, in models involving time invariant individual fixed effects only, we have $\Greekmath 010D _{g(i),m(t)}=\Greekmath 010D _{i}$, e.g., with $ \mathcal{G}=\{1,\ldots ,n\}$, and $\mathcal{M}=\{1\}$. In models involving only time fixed effects, we have $\Greekmath 010D _{g(i),m(t)}=\Greekmath 010D _{t}$, e.g., with $\mathcal{G}=\{1\}$, and $\mathcal{M}=\{1,\ldots ,T\}$. In the Section (ref) of the Online Appendix, we consider non-nested hypothesis testing comparing several variants of the linear model.} In the main text of the paper, we focus on the Vuong test comparing the heterogeneous time effects model (as considered by Bonhomme-Manresa) and the TWFE model, denoted as model 1 and model 2, respectively.

Specifically, model 1 defines the joint log-likelihood as

equation[equation omitted — 175 chars of source]

where $\mathcal{G}_{1}\equiv \{1,\ldots ,G_{1}\}$ and $G_{1}$ is a positive integer. The function $\Greekmath 0120 \left( z_{i,t};\Greekmath 0112 ,\Greekmath 010D \right) $ is given by

equation[equation omitted — 182 chars of source]

In contrast, model 2 assumes that the joint log-likelihood is

equation[equation omitted — 167 chars of source]

where $\Greekmath 010D _{2,i}$ represents the time-invariant individual effect, and $ \Greekmath 010D _{2,t}$ captures the homogeneous time effect. We impose the following normalization condition on model 2:

equation[equation omitted — 72 chars of source]

to ensure unique identification of the pseudo-true parameters.

In the following, we introduce the algorithm for calculating the modified QLR statistic and explain its asymptotic properties. Even though the technical details are more involved and use more complex notation, the results in this section are presented in a way quite similar to those in the previous one. To make things easier to follow, we have organized this section in a way that mirrors the structure of the (second half of the) previous section.

Estimation under Group-Time Effects and TWFE

The pseudo true parameters $\Greekmath 0112 _{1}^{\ast }$ and $\Greekmath 010D _{1,g,t}^{\ast } $ (for $g\in \mathcal{G}_{1}$ and $t\leq T$) for model 1, defined as the maximizers of the population version of the objective in ((ref)), take the following forms:

equation*[equation* omitted — 360 chars of source]

where $\Sigma _{\dot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\mathbb{E }[\dot{x}_{i,t}^{\ast }\dot{x}_{i,t}^{\ast \top }]$, $\dot{x}_{i,t}^{\ast }=x_{i,t}-\mathbb{E}[\bar{x}_{g,t}]$, and for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, $\bar{x}_{g,t}\equiv n_{g}^{-1}\sum_{i\in I_{g}}x_{i,t}$ and $\bar{y}_{g,t}\equiv n_{g}^{-1}\sum_{i\in I_{g}}y_{i,t}$.\ The estimators of $\Greekmath 0112 _{1}^{\ast }$ and $\Greekmath 010D _{1,g,t}^{\ast }$ are their sample analogs:

equation[equation omitted — 319 chars of source]

where $\dot{x}_{i,t}\equiv x_{i,t}-\bar{x}_{g,t}$ and $\hat{\Sigma}_{\dot{x} }\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\dot{x}_{i,t}\dot{x} _{i,t}^{\top }$.

Applying ((ref)), we can solve the pseudo true parameters $\Greekmath 0112 _{2}^{\ast }$, $(\Greekmath 010D _{2,i}^{\ast })_{i\leq n}$\ and\ $(\Greekmath 010D _{2,t}^{\ast })_{t\leq T}$ in model 2 as

equation*[equation* omitted — 563 chars of source]

where $\Sigma _{\ddot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\mathbb{ E}[\ddot{x}_{i,t}^{\ast }\ddot{x}_{i,t}^{\ast \top }]$ and $\ddot{x} _{i,t}^{\ast }\equiv x_{i,t}-\mathbb{E}[\bar{x}_{i}-\bar{x}_{t}+\bar{x}]$. The estimators of the pseudo true parameters are their empirical counterparts:

equation[equation omitted — 487 chars of source]

where\ $\ddot{x}_{i,t}\equiv x_{i,t}-\bar{x}_{i}-\bar{x}_{t}+\bar{x}$ and $ \hat{\Sigma}_{\ddot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\ddot{x} _{i,t}\ddot{x}_{i,t}^{\top }$.

Asymptotic Distribution of the QLR Statistic

Using the estimators of the pseudo true parameters from both models, we obtain the QLR statistic for model comparison as

equation*[equation* omitted — 201 chars of source]

where $\hat{\Greekmath 0122 }_{j,i,t}\equiv y_{i,t}-x_{i,t}^{\top }\hat{\Greekmath 0112 } _{j}-\hat{\Greekmath 010D }_{j,i,t}$\ for $j=1,2$, $\hat{\Greekmath 010D }_{1,i,t}\equiv \hat{ \Greekmath 010D }_{1,g,t}$ for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, and $ \hat{\Greekmath 010D }_{2,i,t}\equiv \hat{\Greekmath 010D }_{2,i}+\hat{\Greekmath 010D }_{2,t}$ for any $ i\leq n$ and any $t\leq T$.

We next derive the asymptotic distribution of $QLR_{n,T}$. Some notation are needed. Let $\Greekmath 0122 _{j,i,t}\equiv y_{i,t}-x_{i,t}^{\top }\Greekmath 0112 _{j}^{\ast }-\Greekmath 010D _{j,i,t}^{\ast }$ and $\Greekmath 0122 _{j,i,t}^{\ast }\equiv \Greekmath 0122 _{j,i,t}-\mathbb{E}[\Greekmath 0122 _{j,i,t}]$ for $j=1,2$, where $\Greekmath 010D _{1,i,t}^{\ast }\equiv \Greekmath 010D _{1,g,t}^{\ast }$ for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, and $\Greekmath 010D _{2,i,t}^{\ast }\equiv \Greekmath 010D _{2,i}^{\ast }+\Greekmath 010D _{2,t}^{\ast }$ for any $i\leq n$ and any $ t\leq T$. Without loss of generality,\ we suppose that for any $g\in G_{1}$, $I_{g}=\{N_{g}+1,\ldots ,N_{g}+n_{g}\}$ where $N_{g}\equiv \sum_{g^{\prime }<g}n_{g^{\prime }}$.

For model 1, we define

equation*[equation* omitted — 356 chars of source]

for any $g\in \mathcal{G}_{1}$ and any $i\in I_{g}$. For model 2, we let

equation*[equation* omitted — 408 chars of source]

It can be shown\footnote{ See Lemma (ref) and Lemma (ref) in the Online Appendix.} that

equation[equation omitted — 139 chars of source]

where\ $\tilde{\Psi}_{i}=(2T^{1/2})^{-1}\sum_{t\leq T}(\Greekmath 0122 _{2,i,t}^{2}-\Greekmath 0122 _{1,i,t}^{2})$, $\tilde{V}_{i}=\tilde{V}_{1,i}- \tilde{V}_{2,i}$ and $\tilde{U}_{i}=\tilde{U}_{1,i}-\tilde{U}_{2,i}$. The variance $\Greekmath 0121 _{n,T}^{2}$ of the $QLR_{n,T}$ is defined as in ((ref)), with $\tilde{\Psi}_{i}$, $\tilde{V}_{i}$ and $\tilde{U}_{i}$ taking the specific forms given above. Below is the asymptotic property of the infeasible test statistic that mirrors Theorem (ref):

theorem\ Under\ Assumptions (ref),\ (ref) and (ref) in the Appendix, we have \begin{equation*} \frac{QLR_{n,T}-\mathbb{E}[S_{n,T}]-\overline{QLR}_{n,T}}{\Greekmath 0121 _{n,T}} \rightarrow _{d}N(0,1), \end{equation*} where $\mathbb{E}[S_{n,T}]=(2nT)^{-1/2}\sum_{g\in \mathcal{G}_{1}}\sum_{i\in I_{g}}\mathbb{E}\left[ \sum_{t\leq T}(n_{g}^{-1}\Greekmath 0122 _{1,i,t}^{\ast 2}-n^{-1}\Greekmath 0122 _{2,i,t}^{\ast 2})-T^{-1}(\sum_{t\leq T}\Greekmath 0122 _{2,i,t}^{\ast })^{2}\right] $.

Bias/Variance Estimation and the Feasible Test

Following the discussion in Section (ref), we present a feasible modified QLR procedure by estimating the bias $\mathbb{E}[S_{n,T}]$ and the variance $\Greekmath 0121 _{n,T}^{2}$. Consistent estimation of these quantities is required under the null hypothesis to ensure correct size control of the test. For this purpose, we assume that $(\Greekmath 0122 _{1,i,t},\Greekmath 0122 _{2,i,t})$ are independent across $t$ with time-invariant first and second moments; this condition is imposed in Assumption (ref).\footnote{ The independence of $(\Greekmath 0122 _{1,i,t},\Greekmath 0122 _{2,i,t})$ over $t$ is imposed primarily to simplify the form of $\Greekmath 0121 _{n,T}^{2}$ and its consistent estimator. This condition is only required under the null hypothesis; see Lemma (ref) in the Appendix for consistency of the test without this assumption.}

Under this condition, we have

equation*[equation* omitted — 185 chars of source]

where $\Greekmath 011B _{j,i}^{2}\equiv \mathrm{Var}(\Greekmath 0122 _{j,i,t})$ can be estimated by $\hat{\Greekmath 011B }_{j,i}^{2}\equiv T^{-1}\sum_{t\leq T}\hat{ \Greekmath 0122 }_{j,i,t}^{2}$. Therefore, the bias term can be estimated by

equation[equation omitted — 225 chars of source]

which leads to the modified QLR statistic:

equation[equation omitted — 85 chars of source]

It can be shown\footnote{ See Lemma (ref) in the Appendix.}\ that $\Greekmath 0121 _{n,T}^{2}$ can be approximated by\ $\Greekmath 011B _{n,T}^{2}+\Greekmath 011B _{U,n,T}^{2}$, where

equation[equation omitted — 171 chars of source]

and

align[align omitted — 431 chars of source]

We consider the sample variance

equation[equation omitted — 238 chars of source]

as an estimator of $\Greekmath 011B _{n,T}^{2}$. The variance $\Greekmath 011B _{U,n,T}^{2}$ from estimating the incidental parameters is then estimated by

align[align omitted — 462 chars of source]

where\ $\hat{\Greekmath 011B }_{1,2,i}\equiv T^{-1}\sum_{t\leq T}\hat{\Greekmath 0122 } _{1,i,t}\hat{\Greekmath 0122 }_{2,i,t}$.

Proceeding as in the previous section, our estimate of $\Greekmath 0121 _{n,T}^{2}$ is thus

equation[equation omitted — 204 chars of source]

and the Vuong test follows ((ref)) with $MQLR_{n,T}$ and $\hat{ \Greekmath 0121 }_{n,T}^{2}$ constructed in ((ref)) and ((ref) ), respectively.\footnote{ In Lemma (ref) of the Appendix, we show that\ $\hat{\Greekmath 011B } _{U,n,T}^{2}$ is positive in finite samples, which implies that the estimator of $\Greekmath 0121 _{n,T}^{2}$ constructed in ((ref)) below is also positive.} Mirroring Theorem (ref), we have:

theoremSuppose that\ Assumptions (ref),\ (ref), (ref) and (ref) in the Appendix hold. Further suppose that as $n,T\rightarrow \infty $, \begin{equation} \frac{(nT)^{-1/2}\sum_{i\leq n}\sum_{t\leq T}\mathbb{E}[\Greekmath 0122 _{2,i,t}^{2}-\Greekmath 0122 _{1,i,t}^{2}]}{\Greekmath 0121 _{n,T}}\rightarrow c. \end{equation} Then we have as $n,T\rightarrow \infty $, \begin{equation*} \mathbb{E}[\Greekmath 0127 _{n,T}^{2\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 2-\Phi (z_{1-p/2}-c)-\Phi (z_{1-p/2}+c), \end{equation*} and \begin{equation*} \mathbb{E}[\Greekmath 0127 _{n,T}^{1\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}(p)]\rightarrow 1-\Phi (z_{1-p}-c), \end{equation*} where $\Greekmath 0127 _{n,T}^{2\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}$ and $\Greekmath 0127 _{n,T}^{1\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{-side}}$ follow the identical form as ((ref)).

Conclusion

This paper extends the classical Vuong1989 test to panel data models with fixed effects, addressing challenges that arise from high-dimensional nuisance parameters. By modifying the profile likelihood and applying bias correction techniques, we develop a valid procedure for comparing non-nested panel data models. Our analysis emphasizes the need for variance adjustments in the presence of incidental parameter problems and shows how these modifications support more reliable model selection in high-dimensional settings.

We contribute to the literature by generalizing the Vuong test to accommodate grouped heterogeneity in both individual and time effects. Furthermore, we propose a methodological framework for model selection when competing specifications define cross-sectional units differently---a common issue in empirical work. The theoretical results in this paper highlight the critical role of bias correction in panel data analysis and complement existing research on non-nested hypothesis testing.