Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,432 characters · 14 sections · 58 citation commands
Model Selection in Panel Data Models: A Generalization of the Vuong Test
With the advent of big data, modern empirical research increasingly adopts more complex models. These models often involve high-dimensional nuisance parameters, which can introduce substantial biases in standard estimators - manifesting as variations of the incidental parameters problem. Classical references include NeymanScott1948 and Nickell1981 among others. Numerous studies have been conducted to address and mitigate these issues. It is now well established that these issues represent a form of higher-order bias, and a variety of methodological approaches have been proposed to overcome them. See ArellanoHahn2007, Arellano-Bonhomme, BesterHansen2009, CARRO2007503, Dhaene-Jochmans, FERNANDEZVAL200971, Fernandez-Val-Weidner-1 , Fernandez-Val-Weidner-2, HahnKuersteiner2002, and HahnNewey2004, for example.
Although the literature is well developed, its generalizations have only recently found application in substantively meaningful but technically complex contexts, such as network-type models. See Bonhomme-Lamadon-Manresa-1, Jochmans-Weidner, bonhomme2025neymanorthogonalizationapproachincidentalparameter. Motivated by these developments, we propose to investigate issues of model specification in nonlinear panel data models. Like most econometric frameworks, panel data models often rely on potentially restrictive assumptions. However, to the best of our knowledge, there is currently no literature that provides a systematic approach for comparing potentially non-nested models in this context. We would like to close this gap in the literature by proposing a panel generalization of the classical Vuong1989 test.
Specifically, we consider a model characterized by the joint log-likelihood of the data $\{z_{i,t}\}_{i\leq n,t\leq T}$ specified as:
where $f\left( \cdot \right) $ is a known function, $\Greekmath 0112 $ represents unknown parameters that are individual/time-invariant, $\Greekmath 010D _{g(i),m(t)}$ denotes the individual and time fixed effect which has a grouping structure determined by cluster/group assignment functions $g(\cdot )$ and $m(\cdot )$ with ranges $\mathcal{G}\equiv \{1,\ldots ,G\}$ and $\mathcal{M}\equiv \{1,\ldots ,M\}$, respectively, and $\Greekmath 011E \equiv (\Greekmath 0112 ^{\top },(\Greekmath 010D _{g,m}^{\top })_{g\in \mathcal{G},m\in \mathcal{M}})^{\top }$ include all unknown parameters specified in the model. For ease of notations, we let $ I_{g}$ for $g\in \mathcal{G}$ and $I_{m}$ for $m\in \mathcal{M}$ denote the partitions of $\{1,\ldots ,n\}$ and $\{1,\ldots ,T\}$ according to the group assignment functions $g(\cdot )$ and $m(\cdot )$, respectively. The pseduo true value $\Greekmath 011E ^{\ast }$ is defined as the maximizer of $\mathbb{E}\left[ L_{n,T}(\Greekmath 011E )\right] $ over the parameter space, and it is often estimated by the maximizer $\hat{\Greekmath 011E }$ of $L_{n,T}(\Greekmath 011E )$ over the same space. Our group structure is inspired the analyses by Bonhomme-Manresa or Bonhomme-Lamadon-Manresa-2, yet it is general enough to include the classical panel models with individual fixed effects.
In practice, different ways of specifying the likelihood function $f\left( \cdot \right) $ and the group assignment functions $g(\cdot )$ and $m(\cdot ) $ lead to different modeling strategies in panel data models, and we often have the situation to compare competing modeling strategies and hope to find the one which is closer (or the closest) to the true data generating process in terms of the Kullback--Leibler (KL) distance. It would be reasonable to compare models through the Vuong1989 test, which tests the null hypothesis:
where $L_{j,n,T}(\cdot )$ represents the joint log-likelihood specified by modeling strategy $j$ with $\Greekmath 011E _{j}^{\ast }$ denoting the pseudo true value for $j=1,2$. The alternative hypothesis can be
The classical Vuong test employs the quasi-likelihood ratio (QLR), defined as:
to construct the test statistic, where $\hat{\Greekmath 011E }_{j}$ denotes the estimator of $\Greekmath 011E _{j}^{\ast }$ in model $j$. It turns out that in panel models, the log-likelihoods $L_{1,n,T}$ and $L_{2,n,T}$ are biased due to the incidental parameter problems. Panel adaptation of the Vuong test thus requires working with bias-corrected versions of the log likelihoods, which we will call $LM_{1,n,T}^{\ast }$ and $LM_{2,n,T}^{\ast }$. We also show that there is additional complication because the order of magnitude of the standard error of the test statistic (after such modification) can depend on the nature of the hypothesis. Our technical contribution consists of estimation of $LM_{1,n,T}^{\ast }$ and $LM_{2,n,T}^{\ast }$ to make the modification feasible as well as estimation of the standard error while avoiding the uniformity issue.
We extend the classical Vuong1989 test to panel models, contributing to the understanding of model selection complexities in high-dimensional settings. The Vuong test originally compares two non-nested parametric models and formally tests the hypothesis that their KL distances from the true data distribution are equal. Originally designed for parsimoniously parameterized models, this test has been expanded in various studies, see, RiversVuong2002, ChenHongShum, and Shi2015b, for example. In this paper, we demonstrate that the incidental parameters problem in panel data contexts invalidates the classical Vuong test, which compares maximized likelihood ratios between two models. We present a modified test statistic that ensures validity for panel data analysis.
Our modification follows a precedent in LeePhillips2015 and Liao&Shi2020. LeePhillips2015 explored model selection in panel data settings using nested hypothesis testing. Our work complements theirs by introducing approaches for testing non-nested hypotheses. Liao&Shi2020 examined the Vuong test for semiparametric models in a cross-sectional setting, whereas we extend the analysis to a panel data context with fixed effects. Our paper and LeePhillips2015 share the common feature in the sense that both analyses utilize application of profiling to the Kullback--Leibler information criterion (KLIC). It is well-known that profile likelihood does not share the standard properties of the genuine likelihood function. Modification of profile likelihood was discussed a way of overcoming this problem. See ArellanoHahn2007, ArellanoHahn2016, CoxReid1987, DiCiccioStern1993, DiCiccioMartinSternYoung, CoxFergusonReid, and PaceSalvan, for example. Following this tradition, our paper as well as LeePhillips2015 propose a modified profile likelihood that provides approximations to the standard KLIC, and provided a valid method of model selection in panel data models with individual fixed effects. Interestingly, we find some close connection between the modification in Liao&Shi2020 and the panel type modification following ArellanoHahn2007, ArellanoHahn2016, and LeePhillips2015.
Recently developed methods in panel data allow applied researchers to choose among a broad range of alternative models for separable, interactive, or discretely supported unobserved heterogeneity across units and time. It is widely understood that parameter estimates are often sensitive to the chosen specification, however there is often no clear theoretical guidance on how to choose among alternative models. Our paper makes further contribution in this regard by considering the cases where the definition of “individuals” may not be agreed across different models. In empirical practice, the group/cluster structure may not be obviously agreed among researchers. The competing group/cluster structures may be nested in the sense that one structure is “finer” than the other one, but they may be non-nested. In this situation, extension of the insights in previous papers such as LeePhillips2015 or Liao&Shi2020 may not be straightforward. We establish that the classical Vuong test is valid after some modification. In doing so, we also make a technical contribution that may be of some independent interest. FernandezValWeidener2016 analyzed the asymptotic distribution of the (pseudo) MLE in nonlinear panel data models with individual and time effects. Under the assumption that the objective function is globally concave in all the parameters, they established the asymptotic bias formula. Our paper contains a result\footnote{ See Theorem (ref) in the Online Appendix and Theorem (ref) in Appendix (ref), along with the related discussion therein, for further details.} that relaxes this concavity assumption, although it is achieved at the cost of restricting the group structure. Therefore, our result can be argued to complement FernandezValWeidener2016.
The remainder of the paper is organized as follows. In Section (ref),\ we introduce the Vuong test for comparing classical panel data models that allow for potential group heterogeneity. Section (ref) presents a test for comparing the two-way fixed effects (TWFE) model against a heterogeneous time fixed effects model. Section (ref) concludes.\ The proofs of the main results, along with auxiliary lemmas, are provided in the Appendix. Additionally, Appendix (ref) develops a general asymptotic theory for estimators in panel models with individual and time fixed effects, as well as the asymptotic distribution of the statistic $QLR_{n,T}$ in this broader setting -- results that are of independent interest.\footnote{ The technical proofs for this general theory, together with the proofs of the auxiliary lemmas used to establish the main results in Sections (ref) and (ref), are provided in the Online Appendix.}
The following notation will be adopted throughout the paper.\ We use $K$ to denote a generic strictly positive constant that may vary from one instance to another but remains independent of the panel dimensions $n$ {and }$T$.\ We adopt the convention that a summation over an empty set equals zero. We use $a\equiv b$ to indicate that $a$ is defined as $b$. For real numbers $ a_{1},\ldots ,a_{m}$, $(a_{j})_{j\leq m}\equiv (a_{1},\ldots ,a_{m})^{\top }$ . For any matrix $A$, $A^{\top }$\ denotes the transpose of $A$,\ and $ \left\Vert A\right\Vert $ denotes the Euclidean norm of $A$. For any doubly indexed sequence $a_{i,t}$ (where $i=1,\ldots ,n$ and $t=1,\ldots ,T$), we define\ $\bar{a}_{i}\equiv T^{-1}\sum_{t\leq T}a_{i,t}$, $\bar{a}_{t}\equiv n^{-1}\sum_{i\leq n}a_{i,t}$ and $\bar{a}\equiv (nT)^{-1}\sum_{t\leq T}\sum_{i\leq n}a_{i,t}$. The summation $\sum_{i^{\prime }\neq i}$ is taken over all $i^{\prime }$ except $i$, which means $\sum_{i^{\prime }\neq i}a_{i^{\prime },t}=\sum_{i^{\prime }=1}^{i-1}a_{i^{\prime },t}+\sum_{i^{\prime }=i+1}^{n}a_{i^{\prime },t}$.\ For two sequences of positive numbers $a_{n}$ and $b_{n}$, we write $a_{n}\succ b_{n}$ if $ a_{n}\geq c_{n}b_{n}$ for some strictly positive sequence $c_{n}\rightarrow \infty $. The notation $\left\Vert \cdot \right\Vert _{p}$ denotes the $ L_{p} $-norm. Finally, for any function $g(\cdot )$, define $\mathbb{E}_{T} \left[ g\left( z_{i,t}\right) \right] \equiv T^{-1}\sum_{t\leq T}\mathbb{E} \left[ g\left( z_{i,t}\right) \right] $ and $\widehat{\mathbb{E}}_{T}\left[ g\left( z_{i,t}\right) \right] \equiv T^{-1}\sum_{t\leq T}g\left( z_{i,t}\right) $.
In this section, we investigate the Vuong test for panel models with time-invariant individual effects.\footnote{ A simple version of the test appeared in hahn-liu, which later became a chapter in Liu's UCLA dissertation. Our current paper substantially extends and generalizes the earlier simple result.} With a dataset\ $\left\{ z_{i,t}\right\} _{i\leq n,t\leq T}$, we compare two models, denoted as model 1 and model 2 respectively. Model 1 represents the classical panel model with individual effects. In contrast, model 2 may specify a different joint likelihood function and/or adopt a distinct clustering or grouping structure for the individual effects.
For each model $j$ ($j=1,2$), the joint log-likelihood function simplifies from the general specification in ((ref)) to the following form:
where\ $\Greekmath 011E _{j}\equiv (\Greekmath 0112 _{j}^{\top },(\Greekmath 010D _{j,g})_{g\in \mathcal{G }_{j}}^{\top })^{\top }$, $\Greekmath 0120 _{j}\left( z_{i,t};\Greekmath 011E _{j,g}\right) \equiv \log f_{j}(z_{i,t};\Greekmath 0112 _{j},\Greekmath 010D _{j,g})$, $\Greekmath 011E _{j,g}\equiv (\Greekmath 0112 _{j}^{\top },\Greekmath 010D _{j,g})^{\top }$,\ $\mathcal{G}_{j}\equiv \{1,\ldots ,G_{j}\}$ defines the range of the grouping function $g_{j}(\cdot )$, and $ I_{j,g}$ ($g\in \mathcal{G}_{j}$) represents the partition of $\{1,\ldots ,n\}$ induced by $g_{j}(\cdot )$. For model 1, it follows that $\mathcal{G} _{1}=\{1,\ldots ,n\}$, implying that each individual has its own group. For model 2, however, the partitions $\{I_{2,g}\}_{g\in \mathcal{G}_{2}}$ may form any arbitrary grouping of $\{1,\ldots ,n\}$\ allowing for greater flexibility in the specification of individual effects.\footnote{ This means that model 1 is nested within model 2 in terms of the grouping structure, although the likelihood specifications may differ. In many applications, the grouping structure is expected to be identical across models.}
As this section is relatively long, we provide a roadmap to facilitate reading. Sections (ref) - (ref) provides the theoretical foundation of the infeasible test, while Sections (ref) - (ref) presents the actual algorithm for implementation of the feasible version of the test. In Section (ref), we provide an intuitive overview of the adjustment to the maximized joint likelihood and discuss its importance for model comparison using the Vuong test in the panel data models considered in this paper. In Section (ref), we present the log likelihood of each model. In Section (ref), we present an infeasible adjustment that corrects for the bias in the likelihood, and in Section (ref), we present a theoretical analysis for the asymptotic variance of the modified likelihood. Asymptotic properties of the (infeasible) modified test statistic is presented in Theorem (ref). In Sections (ref) and (ref), we present actual algorithms for estimation of the modified log likelihoods and the asymptotic variances. Section (ref) presents Theorem (ref) that characterizes the asymptotic properties of the feasible test statistic.
To streamline notation and improve clarity, we focus here on a simplified version of ((ref)) that includes only individual fixed effects, specified as:
where $\Greekmath 0120 \left( z_{i,t};\Greekmath 011E _{i}\right) \equiv \log f\left( z_{i,t};\Greekmath 0112 ,\Greekmath 010D _{i}\right) $, $\Greekmath 011E _{i}\equiv (\Greekmath 0112 ^{\top },\Greekmath 010D _{i})^{\top }$\ and $\Greekmath 011E \equiv (\Greekmath 0112 ^{\top },(\Greekmath 010D _{i})_{i\leq n}^{\top })^{\top }$.\ Since QLR statistic defined in ((ref)) involves the maximized joint likelihood\ $L_{n,T}(\hat{\Greekmath 011E })$, understanding the QLR's properties requires investigating the estimation error of $L_{n,T}(\hat{\Greekmath 011E })$.
Let $\Greekmath 011E ^{\ast }\equiv (\Greekmath 0112 ^{\ast \top },(\Greekmath 010D _{i}^{\ast })_{i\leq n}^{\top })^{\top }$ denote the pseudo-true parameter which maximizes $ \mathbb{E}\left[ L_{n,T}(\Greekmath 011E )\right] $.\footnote{ Potential dependence of the pseudo-true parameters $\Greekmath 011E ^{\ast }$ on the sample size is not made explicit for notational simplicity throughout the paper.} Define:
where $\Greekmath 011E _{i}^{\ast }\equiv (\Greekmath 0112 ^{\ast \top },\Greekmath 010D _{i}^{\ast })^{\top }$. It can be shown\footnote{ See Theorem (ref) in the Appendix, specialized to the case $\mathcal{G} =\{1,\ldots ,n\}$ and $\mathcal{M}=\{1\}$, so that $G=n$ and $M=1$}\ that under the asymptotics with $nT^{-1}\rightarrow \Greekmath 011A $\ where $\Greekmath 011A \in (0,\infty )$:
where $L_{n,T}(\Greekmath 011E ^{\ast })=\sum_{i\leq n}\sum_{t\leq T}\Greekmath 0120 \left( z_{i,t}\right) $, and for any $i\leq n$,
The expansion in ((ref)) shows that the estimation error of $ L_{n,T}(\hat{\Greekmath 011E })$ has two leading components. The first component, $ L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $, results from estimating\ $\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $ at the given pseudo-true parameter $\Greekmath 011E ^{\ast }$. When scaled by a factor of $(nT)^{-1/2}$, this component converges in distribution to a normal distribution under some regularity conditions. The second component, $ 2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$, arises from estimating the unknown pseudo-true parameter $\Greekmath 011E ^{\ast }$ and is determined by the estimation error of the incidental parameters $(\Greekmath 010D _{i}^{\ast })_{i\leq n}$.\footnote{ Since the dimension of $\Greekmath 0112 ^{\ast }$ is fixed, the contribution of its estimation error to $L_{n,T}(\hat{\Greekmath 011E })$ is asymptotically negligible compared with the two leading components, and is therefore included in the remainder term of ((ref)), represented as $o_{p}(n^{1/2})$.}
From the expansion in ((ref)), it is evident that in panel data models, the $L_{n,T}(\hat{\Greekmath 011E })$ has a first-order bias, stemming from $ 2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$. To evaluate this bias explicitly, suppose that $\{\tilde{\Greekmath 0120 }_{\Greekmath 010D }\left( z_{i,t}\right) \}_{i,t}$ are i.i.d. across $t\leq T$ for any $i$. Since $\mathbb{E}[\tilde{ \Greekmath 0120 }_{\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ]=0$, it follows that $\mathbb{E }[2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}]=2^{-1}\sum_{i\leq n}\Greekmath 011B _{\Greekmath 010D ,i}^{2}$, where $\Greekmath 011B _{\Greekmath 010D ,i}^{2}\equiv \mathbb{E}[ \tilde{\Greekmath 0120 }_{\Greekmath 010D }^{\ast }\left( z_{i,t}\right) ^{2}]$. Assuming that $ \Greekmath 011B _{\Greekmath 010D ,i}^{2}$ is bounded above and below, the bias from the mean of $2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$ is of the same order as $L_{n,T}(\Greekmath 011E ^{\ast })-\mathbb{E}\left[ L_{n,T}(\Greekmath 011E ^{\ast })\right] $. This provides the explanation why the maximized joint likelihood $L_{n,T}( \hat{\Greekmath 011E })$ should be modified with a bias correction based on the mean of\ $2^{-1}\sum_{i\leq n}\tilde{\Psi}_{\Greekmath 010D ,i}^{2}$.
The above discussion demonstrates that the maximized joint likelihood\ $ L_{n,T}(\hat{\Greekmath 011E })$ requires a first-order bias correction in the form of a consistent estimator of $2^{-1}\sum_{i\leq n}\mathbb{E}[\tilde{\Psi} _{\Greekmath 010D ,i}^{2}]$. As a result, the QLR statistic in the classical Vuong test must also be adjusted with a bias correction to ensure valid inference. This bias correction for the QLR statistic is given by the difference in the bias of $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$ ($j=1,2$) from the two competing models.
There is an additional complication. Depending on whether the models are nested or overlapping, the first components, $L_{j,n,T}(\Greekmath 011E _{j}^{\ast })- \mathbb{E}[L_{j,n,T}(\Greekmath 011E _{j}^{\ast})]$, from the two models may cancel out. See Vuong1989 for discussion in the cross section context. In such cases, the second components, denoted as $2^{-1}\sum_{i\leq n}\tilde{\Psi } _{j,\Greekmath 010D ,i}^{2}$ ($j=1,2$),\ can become the primary drivers of the asymptotic distribution of the bias-corrected QLR statistic.\footnote{ In the classical nested case, where the functional form of the likelihood $f$ is identical across the two competing models, the second component of the test statistic cancels out. In this setting, bias correction of the estimator itself, along with further refinements of ((ref)), is needed. This case was previously studied by LeePhillips2015, and we do not pursue it further in this paper.}\ Consequently, the estimation error of the incidental parameters influences not only the bias but also the variance of the QLR statistic.
The pseudo true parameter $\Greekmath 011E _{j}^{\ast }$ for model $j$ is estimated as the maximizer $\hat{\Greekmath 011E }_{j}$ of $L_{j,n,T}(\Greekmath 011E _{j})$ where $\Greekmath 011E _{j}\equiv (\Greekmath 0112 _{j}^{\top },(\Greekmath 010D _{j,g})_{g\in \mathcal{G}_{j}}^{\top })^{\top }$. The estimator\ $\hat{\Greekmath 011E }_{j}$\ satisfies the following first-order conditions:
and for any $g\in \mathcal{G}_{j}$,
where $\Greekmath 011E _{j,g}\equiv (\Greekmath 0112 _{j}^{\top },\Greekmath 010D _{j,g})^{\top }$\ and $ \Greekmath 0120 _{j,a}(z_{i,t};\Greekmath 011E _{j,g})\equiv \partial \Greekmath 0120 _{j}(z_{i,t};\Greekmath 011E _{j,g})/\partial a$ for $a\in \{\Greekmath 0112 _{j},\Greekmath 010D _{j,g}\}$.\ The estimator $\hat{\Greekmath 011E }_{j}$ is subsequently used to compute the maximized joint likelihood $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$\ for model $j$.
Following the discussion in the previous subsection, the maximized joint quasi-likelihood for each model must be adjusted to account for the respective first-order bias, which can be derived from the asymptotic expansion of $L_{j,n,T}(\hat{\Greekmath 011E }_{j})$. To present this expansion in a more general setting than previously considered, we first update the notations introduced earlier. Let $\Greekmath 0120 _{j}\left( z_{i,t}\right) $, $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $, $\Greekmath 0120 _{j,\Greekmath 010D \Greekmath 010D }\left( z_{i,t}\right) $ and $\tilde{\Psi}_{j,\Greekmath 010D ,i}$ represent the counterparts to their definitions in ((ref)) and ((ref)), with $\Greekmath 011E _{j,i}\equiv \Greekmath 011E _{j,g}$,
for any $i\in I_{j,g}$ and any $g\in \mathcal{G}_{j}$.\footnote{ The demeaned version of $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $, i.e., $ \Greekmath 0120 _{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) $ is introduced here because the first-order condition for $\Greekmath 011E _{j}^{\ast }$ only ensures that $ \sum_{i\in I_{j,g}}\sum_{t\leq T}\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. This condition does not guarantee that $\mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. If $\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) $ is stationry across $t$ for each $i$, the first-order condition for $\Greekmath 011E _{j}^{\ast }$ is simplified to $\sum_{i\in I_{j,g}} \mathbb{E}[\Greekmath 0120 _{j,\Greekmath 010D }\left( z_{i,t}\right) ]=0$. In this case, $ \mathbb{E}[\Greekmath 0120 _{1,\Greekmath 010D }\left( z_{i,t}\right) ]=0$ holds for model 1, but is still not guaranteed for model 2.}
With the updated notations, we can obtain\footnote{ See Theorem (ref) together with the decomposition in ((ref)) in the Appendix.} the following expansion:
where the components are defined\footnote{ The $\tilde{\Psi}_{j,\Greekmath 010D ,i}$ is not the derivative of $\tilde{\Psi} _{j,i} $, as can be seen from its definition ((ref)).} as:
It is reasonable to expect $\mathbb{E}[\tilde{U}_{j,i}]=0$ for $j=1,2$ and for any $i\leq n$,\footnote{ For example, it is implied by Assumption (ref)\ in the Appendix.} \ which implies that the first-order bias arises from the mean of $\tilde{V} _{j,i}$. To account for this bias, we define the (infeasible) modified maximized joint likelihood as:
Using this expression, we define the (infeasible) modified QLR statistic, whose asymptotic expansion follows from ((ref)):
where $\tilde{\Psi}_{i}\equiv \tilde{\Psi}_{1,i}-\tilde{\Psi}_{2,i}$, $ \tilde{V}_{i}^{\ast }\equiv \tilde{V}_{1,i}-\tilde{V}_{2,i}-\mathbb{E}[ \tilde{V}_{1,i}-\tilde{V}_{2,i}]$ and $U_{i}\equiv \tilde{U}_{1,i}-\tilde{U} _{2,i}$.
After bias correction, the main remaining components arising from the estimation errors of the incidental parameters, i.e., $\tilde{V}_{i}^{\ast }+ \tilde{U}_{i}$ primarily affect the variance of the modified QLR statistic\ $ MQLR_{n,T}^{\ast}$.\ This effect becomes particularly significant when the two models are overlapping. In such cases, the first component on the right-hand side of ((ref)), i.e., $n^{-1/2}\sum_{i\leq n}\tilde{\Psi }_{i}$, can become arbitrarily small or even vanish, leaving the asymptotic distribution of $MQLR_{n,T}^{\ast}$ predominantly determined by the estimation errors of the incidental parameters, specifically $ n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast}+\tilde{U}_{i})$.
To facilitate the discussion of the variance and asymptotic distribution of $ MQLR_{n,T}^{\ast}$, we follow the literature (see, e.g., Vuong1989, Shi2015b and Liao&Shi2020) and formally define the relationships between models in the panel setting below.
When the two models are strictly non-nested, the variance of $ n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$ is bounded away from zero. In this case, the component $n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast }+\tilde{U} _{i})$ on the right-hand side of ((ref)) becomes asymptotically negligible,\ as it is of order $O_{p}(T^{-1/2})$ under certain regularity conditions.\footnote{ This order follows from an approximation of\ $n^{-1}\sum_{i\leq n}(\mathrm{ Var}(\tilde{V}_{i})+\mathrm{Var}(\tilde{U}_{i}))$ (see ((ref)) in the Online Appendix, which appears in the proof of Lemma (ref)), together with an upper bound on the variance of $\tilde{\Greekmath 0120 }_{j,\Greekmath 010D }^{\ast }\left( z_{i,t}\right) $ for $j=1,2$.}\ Consequently, the asymptotic distribution of $MQLR_{n,T}^{\ast }$ is primarily determined by $n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$. Conversely, when the two models are nested, $\tilde{\Psi}_{i}=0$ under the null hypothesis. In this scenario, the asymptotic distribution of $ MQLR_{n,T}^{\ast }$ is determined by\ $n^{-1/2}\sum_{i\leq n}(\tilde{V} _{i}^{\ast }+\tilde{U}_{i})$. As a result, deriving the asymptotic distribution of $MQLR_{n,T}^{\ast }$ is relatively straightforward, when the models are known to be either strictly non-nested or nested. In both cases, the dominating term on the right-hand side of ((ref)) is identifiable, and the asymptotic distribution of $MQLR_{n,T}^{\ast }$ can be deduced from that of the dominating term.
The asymptotic distribution of $QLR_{n,T}^{\ast }$ is more challenging to derive when the two models are overlapping. Because in this case, it is in general unclear which components between $n^{-1/2}\sum_{i\leq n}\tilde{\Psi} _{i}$ and $n^{-1/2}\sum_{i\leq n}(\tilde{V}_{i}^{\ast }+\tilde{U}_{i})$ could be the dominating term, and both could be equally important. To address this issue, we follow Liao&Shi2020 to consider both terms together when deriving the asymptotic distribution of $MQLR_{n,T}^{\ast }$.
We first discuss the estimation of the bias correction terms. Under Assumption (ref)\ in the Appendix, the definition of $\tilde{V}_{j,i}$ implies that:
This motivates a straightforward plug-in estimator of the bias correction term for model $j$, relying on a consistent estimator of $\Greekmath 011B _{j,\Greekmath 010D ,i}^{2}$\ for any $i\in I_{j,g}$ and any $g\in \mathcal{G}_{j}$:
where
The bias correction term for model $j$ is then estimated as
Using this, the feasible modified QLR statistic is defined as:
where
denotes the modified maximized joint likelihood for model $j$.
Theorem (ref) establishes the consistency of $R_{j,n,T}(\hat{\Greekmath 011E }_{j})$ as an estimator of the bias of the maximized joint likelihood $L_{j,n,T}( \hat{\Greekmath 011E }_{j})$ in model $j$. Combined with the result in ((ref)), this directly leads to the asymptotic normality of the modified QLR statistic $QLR_{n,T}$:
To apply this result for inference under the null hypothesis, it remains to construct a consistent estimator for $\Greekmath 0121 _{n,T}^{2}$, which we now proceed to address.
A consistent estimator of $\Greekmath 0121 _{n,T}^{2}$, which was defined in ((ref)), can be constructed using the sample analog of the variance of $ n^{-1/2}\sum_{i\leq n}\tilde{\Psi}_{i}$, and a variance correction term to account for estimation errors of the incidental parameters. Define $\Greekmath 011B _{n,T}^{2}\equiv n^{-1}\sum_{i\leq n}\mathrm{Var}(\tilde{\Psi}_{i})$ and $ \Delta \Greekmath 0120 (z_{i,t},\Greekmath 011E _{i})\equiv \Greekmath 0120 _{1}(z_{i,t};\Greekmath 011E _{1,i})-\Greekmath 0120 _{2}(z_{i,t};\Greekmath 011E _{2,i})$. Under Assumption (ref), a natural estimator for $\Greekmath 011B _{n,T}^{2}$ is given by:
The theoretical property of $\hat{\Greekmath 011B }_{n,T}^{2}$ is provided in the lemma below.\footnote{ The sample variance\ $\hat{\Greekmath 011B }_{n,T}^{2}$ will be used below as a building block for variance estimation and therefore requires $\{\Delta \Greekmath 0120 (z_{i,t},\Greekmath 011E _{i}^{\ast })\}_{t}$ to be serially uncorrelated, as assumed in Assumption (ref). If serial correlation is present, an autocorrelation-consistent estimator is required. Since variance estimation is mainly needed under the null for size control, Assumption (ref) is imposed only under the null; Lemma (ref) in the Appendix shows that consistency of the test does not rely on this assumption.}
Theorem (ref) shows that $\hat{\Greekmath 011B }_{n,T}^{2}$ tends to overestimate the true variance\ $\Greekmath 011B _{n,T}^{2}$ by $2\Greekmath 011B _{S,n,T}^{2}$. It is reasonable to expect the latter to be of order $O(T^{-1})$, which is formalized by Assumptions (ref)(iii) and (ref). On the other hand,\ it can be shown\footnote{ See Lemma (ref) in the Online Appendix.} that
where
which represents the leading term in the variance of $n^{-1/2}\sum_{i\leq n}( \tilde{V}_{i}^{\ast }+\tilde{U}_{i})$.
Comparing their expressions in ((ref)) and ((ref)), it is evident that $\Greekmath 011B _{S,n,T}^{2}\geq \Greekmath 011B _{U,n,T}^{2}$ in general, with the inequality becoming strict when $\mathbb{E}[\tilde{\Greekmath 0120 }_{j,\Greekmath 010D }\left( z_{i,t}\right) ]\neq 0$. From ((ref)) and ((ref)), it follows that $\hat{\Greekmath 011B }_{n,T}^{2}$ is a consistent estimator of $\Greekmath 0121 _{n,T}^{2}$ when $\Greekmath 011B _{n,T}^{2}$ dominates $\Greekmath 011B _{S,n,T}^{2}$. Since $ \Greekmath 011B _{S,n,T}^{2}=O(T^{-1})$, this dominance is guaranteed when the two models are strictly non-nested, as $\Greekmath 011B _{n,T}^{2}$ is bounded away from zero in such cases. Conversely, if $\Greekmath 011B _{n,T}^{2}$ is of the same or even smaller order than\ $\Greekmath 011B _{S,n,T}^{2}$, a scenario that arises when the two models are (nearly) nested or overlapping, the result in ((ref)) implies that $\hat{\Greekmath 011B }_{n,T}^{2}$ overestimates $\Greekmath 0121 _{n,T}^{2}$ by $2\Greekmath 011B _{S,n,T}^{2}-\Greekmath 011B _{U,n,T}^{2}$, and becomes an inconsistent estimator of $\Greekmath 0121 _{n,T}^{2}$.
The over-estimation issue in $\hat{\Greekmath 011B }_{n,T}^{2}$ can be adjusted through consistent estimators of $\Greekmath 011B _{U,n,T}^{2}$ and $\Greekmath 011B _{S,n,T}^{2}$. From its expression in ((ref)), $\Greekmath 011B _{U,n,T}^{2}$ can be estimated by replacing $\Greekmath 011B _{j,\Greekmath 010D ,i}^{2}$ with $\hat{\Greekmath 011B } _{j,\Greekmath 010D ,i}^{2}$ defined in ((ref)), and replacing $\Greekmath 011B _{12,\Greekmath 010D ,i}^{2}$\ by
Specifically, the estimator of $\Greekmath 011B _{U,n,T}^{2}$ is defined as
Similarly, we define the estimator of $\Greekmath 011B _{S,n,T}^{2}$ as
Here, for any $i\in I_{2,g}$ and any $g\in \mathcal{G}_{2}$:
Thus, $\Greekmath 0121 _{n,T}^{2}$ can be consistently estimated in a general setting as $\hat{\Greekmath 011B }_{n,T}^{2}+\hat{\Greekmath 011B }_{U,n,T}^{2}-2\hat{\Greekmath 011B } _{S,n,T}^{2} $, as indicated by Theorem (ref) below.
One drawback of using $\hat{\Greekmath 011B }_{n,T}^{2}-2\hat{\Greekmath 011B }_{S,n,T}^{2}+\hat{ \Greekmath 011B }_{U,n,T}^{2}$ as a variance estimator is that it may take negative values in finite samples. To address this issue, we follow Liao&Shi2020 and propose the following hybrid variance estimator:
Given the variance estimator $\hat{\Greekmath 0121 }_{n,T}^{2}$, the two-sided and one-sided Vuong tests are defined as:
respectively, where $z_{1-p}$ is the $1-p$ quantile of the standard normal distribution, and $p\in (0,1)$ denotes the significance level.
The general result in the previous section was based on Theorems (ref) -- (ref) in Appendix (ref). As discussed in the introduction, the result established in Appendix (ref) is a technical contribution that relaxes the concavity assumption in FernandezValWeidener2016, but it was done at the cost of restricting the group structure. In this section, we consider potentially more complicated group structure but we do so in the context of the linear model, which satisfies the concavity assumption. Because of substantive economic interest associated with the linear model, in particular, the literature initiated by Bonhomme-Manresa, we believe that the linear model may deserve a special attention.
Bonhomme-Manresa have recently proposed a new model that extends beyond the scope of traditional panel data models. In light of its computational demands, one may naturally question whether a standard TWFE specification could offer a practical alternative. Obviously the two models are not nested, so it is of interest to find a Vuong test for this comparison.
We consider the linear panel model
where $y_{i,t}$ is the dependent variable, $x_{i,t}$ are the observed regressors, and $\Greekmath 0122 _{i,t}$ is i.i.d. across $i$ and $t$ and with mean zero and variance $\Greekmath 011B ^{2}$. Here, the $g(\cdot )$ and $m(\cdot )$ are known cluster/group assignment functions with range $\mathcal{G} =\{1,\ldots ,G\}$ and $\mathcal{M}=\{1,\ldots ,M\}$ respectively, and $ \Greekmath 010D _{g(i),m(t)}$ denotes the unknown fixed effect.\footnote{ The $\Greekmath 010D _{g(i),m(t)}$ takes different forms under various model specifications. For example, in models involving time invariant individual fixed effects only, we have $\Greekmath 010D _{g(i),m(t)}=\Greekmath 010D _{i}$, e.g., with $ \mathcal{G}=\{1,\ldots ,n\}$, and $\mathcal{M}=\{1\}$. In models involving only time fixed effects, we have $\Greekmath 010D _{g(i),m(t)}=\Greekmath 010D _{t}$, e.g., with $\mathcal{G}=\{1\}$, and $\mathcal{M}=\{1,\ldots ,T\}$. In the Section (ref) of the Online Appendix, we consider non-nested hypothesis testing comparing several variants of the linear model.} In the main text of the paper, we focus on the Vuong test comparing the heterogeneous time effects model (as considered by Bonhomme-Manresa) and the TWFE model, denoted as model 1 and model 2, respectively.
Specifically, model 1 defines the joint log-likelihood as
where $\mathcal{G}_{1}\equiv \{1,\ldots ,G_{1}\}$ and $G_{1}$ is a positive integer. The function $\Greekmath 0120 \left( z_{i,t};\Greekmath 0112 ,\Greekmath 010D \right) $ is given by
In contrast, model 2 assumes that the joint log-likelihood is
where $\Greekmath 010D _{2,i}$ represents the time-invariant individual effect, and $ \Greekmath 010D _{2,t}$ captures the homogeneous time effect. We impose the following normalization condition on model 2:
to ensure unique identification of the pseudo-true parameters.
In the following, we introduce the algorithm for calculating the modified QLR statistic and explain its asymptotic properties. Even though the technical details are more involved and use more complex notation, the results in this section are presented in a way quite similar to those in the previous one. To make things easier to follow, we have organized this section in a way that mirrors the structure of the (second half of the) previous section.
The pseudo true parameters $\Greekmath 0112 _{1}^{\ast }$ and $\Greekmath 010D _{1,g,t}^{\ast } $ (for $g\in \mathcal{G}_{1}$ and $t\leq T$) for model 1, defined as the maximizers of the population version of the objective in ((ref)), take the following forms:
where $\Sigma _{\dot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\mathbb{E }[\dot{x}_{i,t}^{\ast }\dot{x}_{i,t}^{\ast \top }]$, $\dot{x}_{i,t}^{\ast }=x_{i,t}-\mathbb{E}[\bar{x}_{g,t}]$, and for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, $\bar{x}_{g,t}\equiv n_{g}^{-1}\sum_{i\in I_{g}}x_{i,t}$ and $\bar{y}_{g,t}\equiv n_{g}^{-1}\sum_{i\in I_{g}}y_{i,t}$.\ The estimators of $\Greekmath 0112 _{1}^{\ast }$ and $\Greekmath 010D _{1,g,t}^{\ast }$ are their sample analogs:
where $\dot{x}_{i,t}\equiv x_{i,t}-\bar{x}_{g,t}$ and $\hat{\Sigma}_{\dot{x} }\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\dot{x}_{i,t}\dot{x} _{i,t}^{\top }$.
Applying ((ref)), we can solve the pseudo true parameters $\Greekmath 0112 _{2}^{\ast }$, $(\Greekmath 010D _{2,i}^{\ast })_{i\leq n}$\ and\ $(\Greekmath 010D _{2,t}^{\ast })_{t\leq T}$ in model 2 as
where $\Sigma _{\ddot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\mathbb{ E}[\ddot{x}_{i,t}^{\ast }\ddot{x}_{i,t}^{\ast \top }]$ and $\ddot{x} _{i,t}^{\ast }\equiv x_{i,t}-\mathbb{E}[\bar{x}_{i}-\bar{x}_{t}+\bar{x}]$. The estimators of the pseudo true parameters are their empirical counterparts:
where\ $\ddot{x}_{i,t}\equiv x_{i,t}-\bar{x}_{i}-\bar{x}_{t}+\bar{x}$ and $ \hat{\Sigma}_{\ddot{x}}\equiv (nT)^{-1}\sum_{i\leq n}\sum_{t\leq T}\ddot{x} _{i,t}\ddot{x}_{i,t}^{\top }$.
Using the estimators of the pseudo true parameters from both models, we obtain the QLR statistic for model comparison as
where $\hat{\Greekmath 0122 }_{j,i,t}\equiv y_{i,t}-x_{i,t}^{\top }\hat{\Greekmath 0112 } _{j}-\hat{\Greekmath 010D }_{j,i,t}$\ for $j=1,2$, $\hat{\Greekmath 010D }_{1,i,t}\equiv \hat{ \Greekmath 010D }_{1,g,t}$ for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, and $ \hat{\Greekmath 010D }_{2,i,t}\equiv \hat{\Greekmath 010D }_{2,i}+\hat{\Greekmath 010D }_{2,t}$ for any $ i\leq n$ and any $t\leq T$.
We next derive the asymptotic distribution of $QLR_{n,T}$. Some notation are needed. Let $\Greekmath 0122 _{j,i,t}\equiv y_{i,t}-x_{i,t}^{\top }\Greekmath 0112 _{j}^{\ast }-\Greekmath 010D _{j,i,t}^{\ast }$ and $\Greekmath 0122 _{j,i,t}^{\ast }\equiv \Greekmath 0122 _{j,i,t}-\mathbb{E}[\Greekmath 0122 _{j,i,t}]$ for $j=1,2$, where $\Greekmath 010D _{1,i,t}^{\ast }\equiv \Greekmath 010D _{1,g,t}^{\ast }$ for any $i\in I_{g}$ and any $g\in \mathcal{G}_{1}$, and $\Greekmath 010D _{2,i,t}^{\ast }\equiv \Greekmath 010D _{2,i}^{\ast }+\Greekmath 010D _{2,t}^{\ast }$ for any $i\leq n$ and any $ t\leq T$. Without loss of generality,\ we suppose that for any $g\in G_{1}$, $I_{g}=\{N_{g}+1,\ldots ,N_{g}+n_{g}\}$ where $N_{g}\equiv \sum_{g^{\prime }<g}n_{g^{\prime }}$.
For model 1, we define
for any $g\in \mathcal{G}_{1}$ and any $i\in I_{g}$. For model 2, we let
It can be shown\footnote{ See Lemma (ref) and Lemma (ref) in the Online Appendix.} that
where\ $\tilde{\Psi}_{i}=(2T^{1/2})^{-1}\sum_{t\leq T}(\Greekmath 0122 _{2,i,t}^{2}-\Greekmath 0122 _{1,i,t}^{2})$, $\tilde{V}_{i}=\tilde{V}_{1,i}- \tilde{V}_{2,i}$ and $\tilde{U}_{i}=\tilde{U}_{1,i}-\tilde{U}_{2,i}$. The variance $\Greekmath 0121 _{n,T}^{2}$ of the $QLR_{n,T}$ is defined as in ((ref)), with $\tilde{\Psi}_{i}$, $\tilde{V}_{i}$ and $\tilde{U}_{i}$ taking the specific forms given above. Below is the asymptotic property of the infeasible test statistic that mirrors Theorem (ref):
Following the discussion in Section (ref), we present a feasible modified QLR procedure by estimating the bias $\mathbb{E}[S_{n,T}]$ and the variance $\Greekmath 0121 _{n,T}^{2}$. Consistent estimation of these quantities is required under the null hypothesis to ensure correct size control of the test. For this purpose, we assume that $(\Greekmath 0122 _{1,i,t},\Greekmath 0122 _{2,i,t})$ are independent across $t$ with time-invariant first and second moments; this condition is imposed in Assumption (ref).\footnote{ The independence of $(\Greekmath 0122 _{1,i,t},\Greekmath 0122 _{2,i,t})$ over $t$ is imposed primarily to simplify the form of $\Greekmath 0121 _{n,T}^{2}$ and its consistent estimator. This condition is only required under the null hypothesis; see Lemma (ref) in the Appendix for consistency of the test without this assumption.}
Under this condition, we have
where $\Greekmath 011B _{j,i}^{2}\equiv \mathrm{Var}(\Greekmath 0122 _{j,i,t})$ can be estimated by $\hat{\Greekmath 011B }_{j,i}^{2}\equiv T^{-1}\sum_{t\leq T}\hat{ \Greekmath 0122 }_{j,i,t}^{2}$. Therefore, the bias term can be estimated by
which leads to the modified QLR statistic:
It can be shown\footnote{ See Lemma (ref) in the Appendix.}\ that $\Greekmath 0121 _{n,T}^{2}$ can be approximated by\ $\Greekmath 011B _{n,T}^{2}+\Greekmath 011B _{U,n,T}^{2}$, where
and
We consider the sample variance
as an estimator of $\Greekmath 011B _{n,T}^{2}$. The variance $\Greekmath 011B _{U,n,T}^{2}$ from estimating the incidental parameters is then estimated by
where\ $\hat{\Greekmath 011B }_{1,2,i}\equiv T^{-1}\sum_{t\leq T}\hat{\Greekmath 0122 } _{1,i,t}\hat{\Greekmath 0122 }_{2,i,t}$.
Proceeding as in the previous section, our estimate of $\Greekmath 0121 _{n,T}^{2}$ is thus
and the Vuong test follows ((ref)) with $MQLR_{n,T}$ and $\hat{ \Greekmath 0121 }_{n,T}^{2}$ constructed in ((ref)) and ((ref) ), respectively.\footnote{ In Lemma (ref) of the Appendix, we show that\ $\hat{\Greekmath 011B } _{U,n,T}^{2}$ is positive in finite samples, which implies that the estimator of $\Greekmath 0121 _{n,T}^{2}$ constructed in ((ref)) below is also positive.} Mirroring Theorem (ref), we have:
This paper extends the classical Vuong1989 test to panel data models with fixed effects, addressing challenges that arise from high-dimensional nuisance parameters. By modifying the profile likelihood and applying bias correction techniques, we develop a valid procedure for comparing non-nested panel data models. Our analysis emphasizes the need for variance adjustments in the presence of incidental parameter problems and shows how these modifications support more reliable model selection in high-dimensional settings.
We contribute to the literature by generalizing the Vuong test to accommodate grouped heterogeneity in both individual and time effects. Furthermore, we propose a methodological framework for model selection when competing specifications define cross-sectional units differently---a common issue in empirical work. The theoretical results in this paper highlight the critical role of bias correction in panel data analysis and complement existing research on non-nested hypothesis testing.