The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
49,294 characters
The Panel Information Matrix Test
\title{The Panel Information Matrix Test}
\author{Martin Schumann\\[0.5em]
{\normalsize Maastricht University}\\[0.25em]
{\normalsize \texttt{[email removed]}}}
\date{\today}
\maketitle
\begin{abstract}
\noindent
The Information Matrix (IM) test is a natural specification check for likelihood-based models, yet cross-sectional implementations are often badly sized, confound neglected heterogeneity with distributional misspecification, and require third derivatives. We develop the panel information matrix (PIM) test for models with fixed effects. Explicit fixed effects isolate functional-form misspecification from time-invariant unobserved heterogeneity. Although the profile maximum likelihood estimator inherits incidental-parameter bias of order $O(1/T)$, the leading bias of the profile PIM is of order $\sqrt{n}/T$, so it is asymptotically negligible whenever $n/T^2\to 0$, which includes the rectangular regime $n/T\to\mathrm{const}$, and stabilizes under parabolic asymptotics $n/T^2\to\rho\in(0,\infty)$. This robustness arises because the indicator fluctuates at rate $\sqrt{n}$ rather than $\sqrt{nT}$. The same rate argument makes third-derivative corrections asymptotically unnecessary, so the test uses only first- and second-order likelihood derivatives. Simulations confirm that asymptotic PIM critical values suffice in moderately long panels, even when Wald tests for $\theta$ are badly sized, while a parametric bootstrap largely removes size distortions when $\rho>0$.
\end{abstract}
\noindent\textit{JEL classification:} C12, C15, C23, C52.\\
\noindent\textit{Keywords:} Information matrix test; panel data; fixed effects; incidental parameters; misspecification; parametric bootstrap.
\section{Introduction}\label{sec:introduction}
Likelihood-based inference requires correct specification of the conditional density of the outcome. Under a correctly specified model, counterfactual and policy analysis can identify estimands that are not identified from experimental or quasi-experimental variation alone \citep{heckman-vytlacil-2005}. The Information Matrix (IM) test of \citet{white1982maximum} is a natural check of that hypothesis, yet it is rarely used. Three difficulties are well known. The test has poor size in small samples \citep{taylor1987size,orme1990small,chesher1991asymptotic}. It often rejects because of neglected heterogeneity rather than because the functional form of the density is wrong \citep{chesher1984testing}. In cross-sections, its asymptotic variance depends on the estimation error of the parameters, so the statistic requires third derivatives of the log-likelihood \citep{chesher1983information,lancaster1984covariance}, which are cumbersome to derive and often numerically unstable.
This paper develops the asymptotic theory of the IM test in panel models with fixed effects. The resulting procedure is the panel information matrix (PIM) test. A central feature of the panel setting is that the information-matrix indicator is an average across individuals of within-individual moments, so that its sampling variability is of order $n^{-1/2}$ rather than $(nT)^{-1/2}$. Relative to the classical cross-sectional IM test, this $\sqrt{n}$ scaling has two main consequences: the profile PIM is first-order robust to incidental parameters, and third-derivative corrections in the asymptotic variance become redundant.
Including fixed effects isolates functional-form misspecification of the idiosyncratic density from time-invariant unobserved heterogeneity and thereby addresses the critique in \citet{chesher1984testing}. For estimation of the common parameter $\theta$, however, profiling out the fixed effects creates an incidental parameter problem \citep{neyman-scott}, as the profile maximum likelihood estimator has a first-order bias of order $O(1/T)$, and Wald-type inference based on that estimator is misleading under rectangular asymptotics $n/T\to\mathrm{const}$ \citep{hahn-newey-nonlinear-panels,hahn-kuersteiner-nonlinear-panels}. The PIM statistic behaves differently. Incidental-parameter effects enter the profile indicator through the score and information biases of the profile likelihood, both of order $O(T^{-1})$ \citep{lee-phillips2015modelselection,schumann-info}. After $\sqrt{n}$ scaling, the leading bias of the profile indicator is therefore of order $\sqrt{n}/T$. This term is $o(1)$ whenever $n/T^2\to 0$, which includes the rectangular regime. Consequently, panels that are too short for reliable inference on $\theta$ can still support a first-order valid specification test. Moreover, the infeasible target likelihood, and feasible alternatives that are first-order equivalent to it, yield a PIM statistic whose null limit is central $\chi^2$ without a restriction on $n/T^2$.
When $n/T^2\to\rho\in(0,\infty)$, the profile PIM converges to a noncentral chi-square whose noncentrality is proportional to $\rho$. In that regime, and more generally in short panels, we recommend a parametric bootstrap along the lines of \citet{horowitz1994bootstrap} and \citet{higgins2024bootstrap}. Under high-level conditions stated below, the bootstrap replicates the limiting law of the PIM statistic. Monte Carlo experiments for static logit, static probit, and dynamic probit show that asymptotic critical values already work well for moderate $T$, even when Wald tests for $\theta$ remain badly sized, while the bootstrap further improves size when $T$ is small. Power against fixed alternatives is reported in the appendix.
The same $\sqrt{n}$ scaling also implies that third-derivative corrections for parameter uncertainty are asymptotically unnecessary. In cross-sections, plug-in estimation error enters the asymptotic variance of the IM indicator on the same scale as the indicator itself \citep{lancaster1984covariance}. In panels, the estimation error for $\theta$ shrinks with both $n$ and $T$, while the indicator continues to fluctuate at rate $\sqrt{n}$. The PIM can therefore be computed from first- and second-order likelihood derivatives only, which is especially attractive when the likelihood must be evaluated numerically, as in many structural panel models \citep{rust1987,aguirregabiria-mira-2007,dearing-blevins-2025}.
Recent work has revived the IM test by characterizing and decomposing its moment restrictions in specific models. \citet{amengual2024multivariate} relate the multivariate-normality IM statistic to Hermite polynomial moment tests and provide a directional decomposition. Related results for multinomial logit and incomplete-data models are developed in \citet{amengual2025logit} and \citet{amengual2026mixtures}. These contributions are cross-sectional. To the best of our knowledge, the asymptotic properties of the IM test in panel models with fixed effects, and in particular its robustness to incidental parameters, have not been studied.
The remainder of the paper is organized as follows. Section~\ref{sec:pseudo-lik} introduces the panel pseudo-likelihoods and the associated score- and information-bias expansions. Section~3 develops the limiting distributions of the target and profile PIM statistics. Section~\ref{sec:bootstrap} develops the parametric bootstrap. Section~\ref{sec:simulations} reports Monte Carlo evidence. Section~\ref{sec:conclusion} concludes. Assumptions, intermediate results, and proofs are collected in Sections~\ref{sec-ass}--\ref{sec-proofs}. Supplementary simulation results, a Neyman--Scott illustration, implementation details, and the illustrative applications of Section~\ref{sec:illustrative-example} are deferred to the appendices.
\section{Pseudo-likelihoods in panel data models with fixed effects}\label{sec:pseudo-lik}
Let $Y_{it}$ denote outcomes and $X_{it}$ a (column) vector of explanatory variables for $i=1, \ldots, n$ individuals and $t=1, \ldots, T$ time periods,
where $n,T \ge 2$.
The time-invariant fixed effect, denoted by $\alpha_{i0}$, is an individual-specific unobserved random variable
whose distribution has known support $\mathcal{J}\vcentcolon= \operatorname{supp}(\alpha_{i0}) \subset \mathbb{R}$, but is otherwise unknown. By modeling unobserved heterogeneity through $\alpha_{i0}$, the PIM test developed here isolates misspecification of the idiosyncratic density from neglected heterogeneity of the type emphasized by \citet{chesher1984testing}. Let further $\mathcal{X}_{i}=(X_{i1},\ldots,X_{iT})$ denote the time series of the observed covariates and $\mathcal{Y}_{i0}=(Y_{i,1-L},\ldots,Y_{i0})$ for a fixed lag length $L\in\mathbb{N}$.
Following \citet{higgins2024bootstrap}, we assume that the likelihood for individual $i$ is the density, conditional on $\mathcal{X}_{i}$ and $\mathcal{Y}_{i0}$, with respect to some appropriate dominating measure. Here, $\theta\in\Theta\subset\mathbb{R}^p$ is the common parameter of interest with true value $\theta_0$ in the interior of $\Theta$, and $\alpha_i$ is a candidate value for the incidental parameter $\alpha_{i0}$. The period log-likelihood contribution is
\[
\ell_{it}(\theta,\alpha_i)=\log f_{it}(Y_{it}|Y_{it-1},\ldots,Y_{it-L},\mathcal{X}_{i};\theta,\alpha_i),
\]
and the scaled log-likelihood for the $i$-th stratum is $\ell_{i}(\theta,\alpha_i)=T^{-1}\sum_{t=1}^T\ell_{it}(\theta,\alpha_i)$. The joint distribution of $(\mathcal{X}_{i},\alpha_{i0})$ is unrestricted, which allows for arbitrary correlation between the explanatory variables and the fixed effect. The operator $\mathbb{E}_0$ refers to integration with respect to the product of the conditional densities in the sample, $\prod_{i=1}^n\prod_{t=1}^T f_{it}(Y_{it}|Y_{it-1},\ldots,Y_{it-L},\mathcal{X}_{i};\theta_0,\alpha_{i0})$, whereas unconditional expectations are denoted by $\mathbb{E}$.\footnote{When the likelihood for individual $i$ is evaluated at non-random parameters, $\mathbb{E}_0$ reduces to an integral with respect to $\prod_{t=1}^T f_{it}(Y_{it}|Y_{it-1},\ldots,Y_{it-L},\mathcal{X}_{i};\theta,\alpha_i)$.} Precise regularity conditions on derivatives and moments are collected in Appendix~\ref{sec-ass}.
\subsection{Profile and target likelihoods}\label{sec:score-info-bias}
Since $\alpha_i$ is unobserved, it must be eliminated from the likelihood, resulting in a pseudo-likelihood that only depends on the common parameter. The two pseudo-likelihoods considered in this paper are the concentrated or ``profile'' likelihood and the infeasible ``target'' likelihood, given as $\ell_i^{\mathfrak{p}}(\theta)\vcentcolon= \ell_{i}(\theta,\hat{\alpha}_{i}(\theta))$ and $\ell_i^{\mathfrak{t}}(\theta)\vcentcolon= \ell_{i}(\theta,\alpha_i^{\mathfrak{t}}(\theta))$, where $\hat{\alpha}_{i}(\theta)=\operatorname*{argmax}_{\alpha_i\in\mathcal{J}}\ell_{i}(\theta,\alpha_i)$ and $\alpha_i^{\mathfrak{t}}(\theta)\vcentcolon=\operatorname*{argmax}_{\alpha_i\in\mathcal{J}}\mathbb{E}_0\ell_{i}(\theta,\alpha_i)$. Their sample averages are $\ell^{\mathfrak{p}}(\theta)=n^{-1}\sum_{i=1}^n\ell_i^{\mathfrak{p}}(\theta)$ and $\ell^{\mathfrak{t}}(\theta)=n^{-1}\sum_{i=1}^n\ell_i^{\mathfrak{t}}(\theta)$, with maximizers $\hat{\theta}^{\mathfrak{p}}=\operatorname*{argmax}_{\theta\in\Theta}\ell^{\mathfrak{p}}(\theta)$ and $\hat{\theta}^{\mathfrak{t}}=\operatorname*{argmax}_{\theta\in\Theta}\ell^{\mathfrak{t}}(\theta)$. Superscripts $^{\mathfrak{p}}$ and $^{\mathfrak{t}}$ are also attached to the corresponding IM indicators, variance estimators, and test statistics introduced below. When a result applies to either pseudo-likelihood, we omit the superscript and write $\ell_i(\theta)$, $\hat{\theta}$, and so on. Derivatives of a pseudo-likelihood with respect to $\theta$ are indexed numerically: $\ell_{i,1}(\theta)=\partial_\theta\ell_i(\theta)$ is a $p$-vector and $\ell_{i,2}(\theta)=\partial_{\theta\theta'}^2\ell_i(\theta)$ is a $p\times p$ matrix.
The target likelihood is score unbiased and information unbiased at $\theta_0$ \citep{pace-salvan-JSPI}. While the target likelihood itself is not available in practical applications, it serves here as a placeholder for feasible likelihoods that are fixed-$T$ consistent, or first-order equivalent to the target as $T\to\infty$, under additional structure---for example conditional likelihoods when a sufficient statistic for $\alpha_i$ exists, and integrated likelihoods based on a correctly specified or robust prior for the fixed effects \citep{arellano-bonhomme}. Results for $\ell^{\mathfrak{t}}$ therefore cover these cases.
Unlike the target likelihood, the profile likelihood suffers from score- and information bias of order $O_{\mathrm{p}}(T^{-1})$ due to the presence of the unobserved fixed effects \citep{schumann-info}. In particular, the score bias drives the incidental-parameter bias of the fixed-effects MLE $\hat{\theta}^{\mathfrak{p}}$ \citep{neyman-scott,hahn-newey-nonlinear-panels}. Because of these first-order biases, the profile likelihood does not behave like a ``genuine'' likelihood even under correct specification \citep{lee-phillips2015modelselection}, and inference based on the MLE is misleading under rectangular asymptotics. A large literature develops analytical, jackknife, and related corrections that remove the leading first-order bias of the fixed-effects MLE and the associated score- and information biases; see the surveys of \citet{lancaster-survey-joe}, \citet{arellano-hahn-world-congress}, and \citet{fernandez2018fixed}, and the developments in \citet{hahn-kuersteiner-nonlinear-panels}, \citet{fernandez-val-probit}, \citet{dhaene-jochmans-restud}, \citet{diciccio-martin-stern-young}, and \citet{schumann2020joe}. Despite the fixed-effects MLE entering the panel IM statistic, we show in the next section that such corrections are asymptotically unnecessary for the PIM test, as the leading incidental-parameter bias term vanishes under rectangular asymptotics.
\section{The IM test in panel data models}
In this section, we establish the asymptotic properties of the PIM test. We demonstrate two main results. First, unlike in the pure cross-sectional case, the PIM statistic does not require third derivatives of the underlying pseudo-likelihood. Second, under parabolic asymptotics, i.e., $\rho=\lim n/T^2\in[0,\infty)$, incidental-parameter bias affecting the profile PIM vanishes when $\rho=0$ and stabilizes when $\rho\in(0,\infty)$. This contrasts with Wald-type inference for parameters, which is biased under rectangular asymptotics.
To define the PIM statistic we adapt the IM test of \citet{white1982maximum} to panel data. Recall that $\ell_{i,1}(\theta)$ and $\ell_{i,2}(\theta)$ are the score and Hessian of the $T$-averaged pseudo-likelihood $\ell_i(\theta)$. The corresponding information-matrix indicator for individual $i$ is based on
\[
T\ell_{i,1}(\theta)\ell_{i,1}(\theta)'+\ell_{i,2}(\theta),
\]
where the factor $T$ appears because $\mathbb{E}_0[\ell_{i,1}(\theta_0)\ell_{i,1}(\theta_0)']=O(T^{-1})$ while $\mathbb{E}_0[\ell_{i,2}(\theta_0)]=O(1)$, so both terms are of the same order. As in \citet{white1982maximum}, the test may use a subset of the $p(p+1)/2$ distinct entries of this matrix. Let $D_i(\theta)\in\mathbb{R}^q$ collect the chosen components, where $1\le q\le p(p+1)/2$, and set
\[
D(\theta)=\frac{1}{n}\sum_{i=1}^n D_i(\theta).
\]
The following lemma shows that $\operatorname{var}_0(\sqrt{n}D(\theta_0))$ remains of order one as $n$ and $T$ grow. The subsequent asymptotic properties of the PIM rest on this result.
\begin{lem}\label{lemma-variance}
Let Assumptions~\ref{ass-sampling} and~\ref{ass-moments} hold. Then,
\begin{equation*}
\operatorname{var}_0(\sqrt{n}D(\theta_0))=O_{\mathrm{p}}(1).
\label{eq:variance}
\end{equation*}
\end{lem}
In the cross-sectional IM test of \citet{white1982maximum}, the averaged indicator has variance of order $1/n$, so the statistic is scaled by $\sqrt{n}$. The same scaling is the right one in the panel. Under the target likelihood the individual indicators are mean zero, but each contains products of scores from different periods. The variance of those products does not vanish as $T\to\infty$, even without unobserved individual effects, so $\operatorname{var}_0(D_i(\theta_0))$ remains of order one and the sample indicator is scaled by $\sqrt{n}$.
The PIM test is constructed based on the indicator evaluated at the pseudo-likelihood maximizer, e.g.\ $D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}})$ for the profile case. As is well-known, the profile estimator $\hat{\theta}^{\mathfrak{p}}$ suffers from an incidental-parameter bias of order $O_{\mathrm{p}}(T^{-1})$ in the presence of fixed effects, which is driven by the score-bias of the same order of the profile likelihood (\citealp{schumann-info}). As a consequence, the asymptotic distribution of $\sqrt{nT}(\hat{\theta}^{\mathfrak{p}}-\theta_0)$ is biased under rectangular asymptotics (see, e.g., \citealp{hahn-newey-nonlinear-panels,hahn-kuersteiner-nonlinear-panels}), and hypothesis tests based on $\hat{\theta}^{\mathfrak{p}}$ are severely misleading in short panels. The score- and information bias of the profile likelihood also affect the bias of the profile indicator function when evaluated at the MLE, since $\mathbb{E}_0[D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}})]=O_{\mathrm{p}}(T^{-1})$. However, because of Lemma \ref{lemma-variance}, the indicator is scaled by $\sqrt{n}$. As is shown in the following Lemma, this implies that $\sqrt{n}D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}})$ remains asymptotically unbiased if $n/T^2\to 0$, whereas the incidental-parameter bias of the scaled profile indicator stabilizes if $n/T^2\to\rho\in (0,\infty)$. Moreover, it is shown that the target likelihood indicator is asymptotically unbiased without the need to impose relative rates on $n$ and $T$. To state the Lemma, let $V\vcentcolon=\lim_{n,T\to\infty}\operatorname{var}_0(\sqrt{n}D^{\mathfrak{t}}(\theta_0))$, which exists and is nonsingular (see Assumption~\ref{ass-variance}).
\begin{lem}\label{lemma-scaled-indicator}
Let Assumptions~\ref{ass-sampling}--\ref{ass-estimators}, and~\ref{ass-variance} hold.
\begin{enumerate}
\item
\[
\sqrt{n}D^{\mathfrak{t}}(\hat{\theta}^{\mathfrak{t}})=\sqrt{n}D^{\mathfrak{t}}(\theta_0)+O_{\mathrm{p}}(n^{-1/2})+O_{\mathrm{p}}(T^{-1/2})
\]
and
\[
\sqrt{n}D^{\mathfrak{t}}(\theta_0)\overset{d}{\to} \mathrm{N}(0,V).
\]
\item
If $n,T\to\infty$,
\[
\sqrt{n}D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}})=
\sqrt{n}D^{\mathfrak{t}}(\theta_0)+O_{\mathrm{p}}(n^{-1/2})+O_{\mathrm{p}}(T^{-1/2})+O_{\mathrm{p}}(\frac{\sqrt{n}}{T}).
\]
Moreover, if Assumption~\ref{ass-n-T-rate}\textup{(ii)} holds, i.e.\ $n/T^2\to\rho\in[0,\infty)$, then
\[
\sqrt{n}D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}})\overset{d}{\to} \mathrm{N}(\sqrt{\rho}\,\kappa(\theta_0),V),
\]
where $\kappa(\theta_0)$ is a non-random limit containing the first-order score and information bias of the profile likelihood (see the proof).
\end{enumerate}
\end{lem}
\begin{rem}
Although the target estimator can be consistent for \(\theta_0\) at fixed panel length, the first part of Lemma \ref{lemma-scaled-indicator} still requires \(T\to\infty\). To see this, consider the first term in a Taylor expansion \(\sqrt{n}\,D_1^{\mathfrak{t}}(\theta_0)(\hat{\theta}^{\mathfrak{t}}-\theta_0)\). By Lemma~\ref{lemma-rates}, \(\mathbb{E}_0[D_{i,1}^{\mathfrak{t}}(\theta_0)]=O_{\mathrm{p}}(1)\) and \(\operatorname{var}_0(D_{i,1}^{\mathfrak{t}}(\theta_0))=O_{\mathrm{p}}(T)\), so this remainder is \(O_{\mathrm{p}}(n^{-1/2})+O_{\mathrm{p}}(T^{-1/2})\). The \(O_{\mathrm{p}}(T^{-1/2})\) piece, which forces \(T\to\infty\), comes from the nonzero mean of \(D_1^{\mathfrak{t}}(\theta_0)\). Bartlett identities keep \(D_i^{\mathfrak{t}}(\theta_0)\) mean zero with \(O_{\mathrm{p}}(1)\) variance only at the true parameter, where the score is mean zero. Since they do not transfer to \(\theta\)-derivatives of the indicator, higher-order likelihood derivatives retain nonzero means, and \(D_{i,1}^{\mathfrak{t}}\) therefore does not inherit the rates of \(D_i^{\mathfrak{t}}\) itself.
\end{rem}
\begin{rem}
The asymptotic bias term $\sqrt{\rho}\kappa(\theta_0)$ that appears in the asymptotic distribution of the profile indicator when $\rho>0$ is a function of both the score- and information bias of the profile likelihood. The score bias enters when expanding the indicator around the true parameter, as the first-order bias of the fixed effects MLE is driven by the first-order score bias of the profile likelihood. The information bias enters when replacing the profile indicator at $\theta_0$ with the target indicator at the true parameter.
\end{rem}
As shown in Lemma \ref{lemma-variance}, the order of the variance does not depend on the panel length. In contrast, both the bias and the variance of $\hat{\theta}^{\mathfrak{p}}$ and $\hat{\theta}^{\mathfrak{t}}$ vanish as $n,T\to\infty$. Combined with the plug-in rates in Lemma \ref{lemma-scaled-indicator}, this fact implies that the estimator of the variance of $D(\hat{\theta})$ does not require third derivatives of the pseudo-likelihood $\ell$, even in the presence of unobserved fixed effects:
\begin{lem}\label{lemma-variance-estimator}
Let Assumptions~\ref{ass-sampling} -- \ref{ass-estimators}, and~\ref{ass-variance} hold. Define
\[
\widehat{V}(\theta)\vcentcolon= \frac{1}{n}\sum_{i=1}^n D_i(\theta)D_i(\theta)'.
\]
Then,
\[
\widehat{V}(\hat{\theta})=V+o_{\mathrm{p}}(1)
\]
for $\ell\in\{\ell^{\mathfrak{p}},\ell^{\mathfrak{t}}\}$.
\end{lem}
Given that the estimator of the variance only requires first- and second derivatives of the pseudo-likelihood, we define the PIM test statistics based on the target and profile likelihoods as
\begin{equation}
\mathcal{W}^{\mathfrak{t}}\vcentcolon= n D^{\mathfrak{t}}(\hat{\theta}^{\mathfrak{t}})' (\widehat{V}^{\mathfrak{t}})^{-1} D^{\mathfrak{t}}(\hat{\theta}^{\mathfrak{t}}),
\qquad
\mathcal{W}^{\mathfrak{p}}\vcentcolon= n D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}})' (\widehat{V}^{\mathfrak{p}})^{-1} D^{\mathfrak{p}}(\hat{\theta}^{\mathfrak{p}}).
\label{eq:def-Whatp}
\end{equation}
Lemmas \ref{lemma-scaled-indicator} and \ref{lemma-variance-estimator} directly lead to the main result of this section, which derives the asymptotic distribution of the PIM test for the target and the profile likelihood.
\begin{thm}\label{main-theorem}
Let Assumptions~\ref{ass-sampling}--\ref{ass-estimators} and~\ref{ass-variance} hold.
\begin{enumerate}
\item As $n,T\to\infty$,
\begin{equation}
\mathcal{W}^{\mathfrak{t}} \overset{d}{\to} \chi^2_q.
\label{eq:main-thm-target}
\end{equation}
\item If $n,T\to\infty$ such that $n/T^2\to\rho\in[0,\infty)$, then
\begin{equation}
\mathcal{W}^{\mathfrak{p}} \overset{d}{\to} \chi^2_q(\lambda_\rho),
\qquad
\lambda_\rho=\rho\,\kappa(\theta_0)'V^{-1}\kappa(\theta_0).
\label{eq:main-thm-profile}
\end{equation}
\end{enumerate}
\end{thm}
Theorem~\ref{main-theorem} highlights two practical consequences.
The more surprising one concerns incidental parameters. Under rectangular
asymptotics the profile MLE $\hat{\theta}^{\mathfrak{p}}$ is biased at order $O(1/T)$,
and Wald tests for $\theta$ based on it are unreliable. In contrast,
when $n/T^2\to 0$, $\mathcal{W}^{\mathfrak{p}}$ has a central $\chi^2_q$ limit even
though estimation uses the same biased $\hat{\theta}^{\mathfrak{p}}$. When
$n/T^2\to\rho\in(0,\infty)$, the distortion is confined to the explicit
noncentrality $\lambda_\rho$ in \eqref{eq:main-thm-profile}, which contrasts with the explosive bias of the Wald test statistic. A specification
check can therefore be first-order valid in panels that are still too short
for conventional asymptotic inference on $\theta$.
The second consequence is computational. Because plug-in estimation error is
asymptotically negligible for $D(\hat{\theta})$, the feasible statistic uses only
the sample second-moment matrix of the individual indicators
$D_i(\hat{\theta})$. Third derivatives of the log-likelihood, which have long
impeded routine use of the cross-sectional IM test
\citep{chesher1983information,lancaster1984covariance,orme1990small}, are
not required. For the target (and target-equivalent) PIM, the null limit is
central $\chi^2_q$ without any restriction on $n/T^2$.
\begin{rem}
The IM test is the leading example of a likelihood specification test based on Bartlett identities. Related procedures include conditional moment diagnostics that exploit the same identities \citep{newey1985,tauchen1985} and tests of higher-order Bartlett restrictions. The rate argument developed here is not special to the second identity, since whenever the indicator is an average across individuals of within-individual likelihood moments whose $\sqrt{n}$-scaled variance remains of order one as $T\to\infty$, plug-in of $\hat{\theta}$ becomes negligible and incidental-parameter bias enters only through $\sqrt{n}/T$. The same qualitative conclusions---redundancy of the next order of derivatives in the variance estimator, and first-order robustness of the profile statistic under rectangular asymptotics---should therefore extend to other Bartlett-identity tests in panels with fixed effects. A full treatment is left for future work.
\end{rem}
\begin{rem}
The results in Theorem \ref{main-theorem} crucially depend on the definition of the individual indicator.
Write $\ell_{it,1}(\theta)$ and $\ell_{it,2}(\theta)$ for the period-$t$ contributions to the score and Hessian of $\ell_i(\theta)$, so that $\ell_{i,1}(\theta)=T^{-1}\sum_{t=1}^T\ell_{it,1}(\theta)$ and $\ell_{i,2}(\theta)=T^{-1}\sum_{t=1}^T\ell_{it,2}(\theta)$. Then,
\[
D_i(\theta)
=
T^{-1}\sum_{t=1}^T\bigl(\ell_{it,1}(\theta)\ell_{it,1}(\theta)'+\ell_{it,2}(\theta)\bigr)
+
T^{-1}\sum_{t\neq s}\ell_{it,1}(\theta)\ell_{is,1}(\theta)',
\]
where the second (cross-period) sum is why $\operatorname{var}_0(D_i(\theta_0))$ remains bounded as $T\to\infty$, and why the sample indicator must be scaled by $\sqrt{n}$. Because those products are mean-zero at $\theta_0$ under independence across time, an apparently natural alternative is to retain only the period-level terms
\[
\widetilde{D}_i(\theta)
=
\frac{1}{T}\sum_{t=1}^T\bigl(\ell_{it,1}(\theta)\ell_{it,1}(\theta)'+\ell_{it,2}(\theta)\bigr),
\qquad
\widetilde{D}=n^{-1}\sum_{i=1}^n\widetilde{D}_i.
\]
For this choice, $\operatorname{var}_0(\widetilde{D}(\theta_0))=O((nT)^{-1})$, so the corresponding statistic must be scaled by $\sqrt{nT}$. The plug-in error replacing $\theta_0$ by $\hat{\theta}$ is then of the same order as the fluctuation of the indicator, and, as in the pure cross-sectional case, the variance estimator requires third derivatives of $\ell$. In addition, the $O(T^{-1})$ incidental-parameter bias of the profile likelihood indicator is multiplied by $\sqrt{nT}$ rather than $\sqrt{n}$, resulting in an asymptotic bias under rectangular asymptotics.
\end{rem}
\section{Parametric bootstrap}\label{sec:bootstrap}
While Theorem~\ref{main-theorem} shows that the PIM is reliable even when $n$ is substantially larger than $T$, size distortions can remain when asymptotic inference for the profile statistic uses central $\chi^2_q$ critical values in the stabilizing regime $\rho\in(0,\infty)$. Residual finite-sample distortions of that feasible statistic may also persist in short panels. In these cases we recommend a parametric bootstrap of the profile PIM. \citet{horowitz1994bootstrap} motivates a parametric bootstrap for the cross-sectional information-matrix test. The construction follows \citet{higgins2024bootstrap}. Formal bootstrap validity is not inherited from \citet{higgins2024bootstrap}, because their result applies to the MLE and the likelihood-ratio statistic under rectangular asymptotics, and their delta method (Theorem~2) does not apply directly to the PIM statistic. We therefore establish a separate first-order bootstrap validity result for the profile statistic under neighborhood-uniform versions of the $\mathbb{E}_0$ primitives in Section~\ref{sec:ass-bootstrap}.
Following \citet{higgins2024bootstrap}, bootstrap outcomes are drawn recursively from the fitted profile transition densities at $(\hat{\theta}^{\mathfrak{p}},\{\hat\alpha_i\})$, holding covariates and initial conditions fixed. On each bootstrap panel one recomputes the profile bootstrap PIM statistic $\mathcal{W}^{{\mathfrak{p}}*}$. Write $P^*$ for the bootstrap probability measure conditional on the original sample, and $P_0$ for the sampling probability under the true conditional density with parameter $(\theta_0',\alpha_{10},...,\alpha_{n0})'$. We write $\overset{d^*}{\to}$ for weak convergence under $P^*$, understood throughout as holding in probability under the sampling measure $P_0$. Explicit transition densities and implementation details are given in Appendix~\ref{sec:mc-implementation}.
\begin{thm}\label{thm:bootstrap}
Let Assumptions~\ref{ass-sampling}--\ref{ass-estimators},~\ref{ass-variance}, and~\ref{ass-bootstrap-uniform}--\ref{ass-bootstrap-bias} hold, and suppose $n/T^2\to\rho\in[0,\infty)$. Conditionally on the sample, under the parametric bootstrap for the profile pseudo-likelihood,
\[
\mathcal{W}^{{\mathfrak{p}}*}\overset{d^*}{\to} \chi^2_q(\lambda_\rho)
\quad\text{in probability},
\]
with $\lambda_\rho=\rho\,\kappa(\theta_0)'V^{-1}\kappa(\theta_0)$ as in Theorem~\ref{main-theorem}. Consequently,
\[
\sup_{x\in\mathbb{R}}\bigl|P^*(\mathcal{W}^{{\mathfrak{p}}*}\le x)-P_0(\mathcal{W}^{\mathfrak{p}}\le x)\bigr|\to_p 0,
\]
so that bootstrap critical values are consistent for the quantiles of the limiting law of the profile PIM statistic.
\end{thm}
Theorem~\ref{thm:bootstrap} is proved under the neighborhood-uniform bootstrap conditions in Section~\ref{sec:ass-bootstrap}. When $\rho>0$, the profile bootstrap replicates the noncentral limit $\chi^2_q(\lambda_\rho)$ rather than a central $\chi^2_q$ law. Consequently, the bootstrap yields critical values that track the sampling distribution of $\mathcal{W}^{\mathfrak{p}}$. While the target PIM does not suffer from size distortions as $n,T\to\infty$ with $n/T^2\to\rho>0$, we find that a similar parametric bootstrap procedure outlined in Appendix~\ref{sec:mc-implementation} improves the size of the target PIM in small panel data sets. This is in line with the findings of \cite{horowitz1994bootstrap}, who finds that the parametric bootstrap improves the empirical size of the cross-sectional IM test for correctly specified likelihoods.
\section{Simulations}\label{sec:simulations}
In this section, we analyze the small-sample properties of the PIM test in static logit, static probit, and dynamic probit under the designs in Section~\ref{sec:sim-setup}. Neyman--Scott simulation results are deferred to Subsection~\ref{sec:add-sim-ns}.
\subsection{Simulation Setup}\label{sec:sim-setup}
\noindent
We conduct Monte Carlo experiments for three panel binary-choice models: static logit, static probit, and dynamic probit. In each design, we report results for a balanced panel with $n\in\{125,250,500,1000\}$ and $T\in\{5,10,15,20,25,30,40,50\}$.
\noindent
Let $m_{it}(\theta)=\theta_1 X_{it,1}+\theta_2 X_{it,2}$. In the correctly-specified design, the outcomes are generated as
\begin{equation}\label{eq:assumed-model-simulation}
\Pr(Y_{it}=1|\mathcal{X}_{i},\alpha_{i0};\theta_0)=F(m_{it}(\theta_0)+\alpha_{i0}),
\end{equation}
where $F$ corresponds to the standard logistic CDF in static logit and to the standard normal CDF in static probit, respectively. The covariates $X_{it,1}$ and $X_{it,2}$ are standard normal with $\mathrm{Corr}(X_{it,1},X_{it,2})=0.5$. The individual effects satisfy $\alpha_{i0}\sim\mathrm{N}(\bar{X}_{i,1},1)$ with $\bar{X}_{i,1}=T^{-1}\sum_{t=1}^T X_{it,1}$, and we set $\theta_0=(\theta_{0,1},\theta_{0,2})'=(1,-1)'$. Estimation always assumes model \eqref{eq:assumed-model-simulation}. We further consider two data generating processes that lead to a misspecified likelihood. In the heteroskedastic design, each unit has a latent scale $\sigma_i=\exp(0.75\,\xi_i)$ with $\xi_i\stackrel{\mathrm{iid}}{\sim}\mathrm{N}(0,1)$, and outcomes follow
\begin{equation*}
\Pr(Y_{it}=1\mid\mathcal{X}_{i},\alpha_{i0},\sigma_i;\theta_0)
=F\!\left(\frac{m_{it}(\theta_0)+\alpha_{i0}}{\sigma_i}\right),
\end{equation*}
whereas in the omitted-interaction design \citep[Section~3.2, eq.~(7)]{horowitz1994bootstrap},
\begin{equation*}
\Pr(Y_{it}=1\mid\mathcal{X}_{i},\alpha_{i0};\theta_0)
=F\!\left(m_{it}(\theta_0)+1.5\,X_{it,1}X_{it,2}+\alpha_{i0}\right).
\end{equation*}
In dynamic probit, $m_{it}(\theta)$ is replaced by
\[
m_{it}^{\mathrm{dyn}}(\theta,\gamma)
=\theta_{1} X_{it,1}+\theta_{2} X_{it,2}+\gamma Y_{i,t-1},
\]
with $\gamma_0=0.5$.
In each of the 5000 Monte Carlo iterations, we compute the PIM test based on the target and the profile likelihood as defined in \eqref{eq:def-Whatp}. For the asymptotic PIM, we compare the statistic with the $95\%$ quantile of $\chi^2_q$, where $q=K(K+1)/2$ is the number of distinct entries of the information-matrix indicator ($K=2$ and $q=3$ in static logit and probit, $K=3$ and $q=6$ in dynamic probit). The Wald test uses $\chi^2_K$. The bootstrap is implemented as in Appendix~\ref{sec:mc-implementation} and uses $B=99$ bootstrap draws. We further report the rejection rates of the Wald test, which is asymptotically biased under rectangular asymptotics. The comparison of the Wald test and the PIM test highlights the practical relevance of the asymptotic regime since the PIM test is expected to perform well under rectangular asymptotics, whereas the Wald test is expected to suffer from severe size distortions. The rejection rates under the misspecified designs are reported in Appendix~\ref{sec:add-sim-results}.
\subsection{Simulation Results}\label{sec:sim-results}
\noindent
Tables~\ref{tab:logitrejectionrates}--\ref{tab:dynamicprobitrejectionrates} report rejection rates under the correctly specified DGP. Entries give the fraction of rejections at nominal size $5\%$ for the asymptotic IM test (IM), the bootstrap IM test (Boot), and the Wald test (Wald).
Several patterns recur across the three designs. The infeasible target likelihood behaves like a correctly specified cross-sectional IM test, since under the null its limiting law is central $\chi^2$ without requiring restrictions on the relative growth rates of $n$ and $T$. Asymptotic critical values therefore work well once $n$ is large enough. For instance, at $n=1000$, the target PIM rejection rates are close to the nominal level at every $T$ in the static designs. In the dynamic probit, the rates remain mildly elevated, which is consistent with the remainder terms in Lemma \ref{lemma-scaled-indicator}. In short panels the target test can still be liberal (up to $18.5\%$ in static probit and $28.6\%$ in dynamic probit at $T=5$ and $n=125$). However, the parametric bootstrap restores size close to $5\%$ throughout (e.g.\ $5.6\%$, $5.2\%$, and $6.6\%$ at that grid point in static logit, static probit, and dynamic probit), as observed for pure cross-sections in \citet{horowitz1994bootstrap}. In static logit, the conditional PIM tracks the target at every $(n,T)$ in Table~\ref{tab:logitrejectionrates}, consistent with the role of $\ell^{\mathfrak{t}}$ as a placeholder for feasible pseudo-likelihoods in Section~\ref{sec:pseudo-lik}.
Turning to the profile PIM, asymptotic $\chi^2$ critical values can be severely oversized when $\rho>0$ is plausible, as in static logit with $T=5$, where rejection rises from $10.4\%$ at $n=125$ to $100\%$ at $n=1000$. Size distortions vanish much faster as $T$ increases than for Wald inference on $\theta$ based on the same profile estimator and asymptotic $\chi^2$ critical values. Holding $n=1000$ fixed in static logit, asymptotic profile rejection falls from $100\%$ at $T=5$ to $5.8\%$ at $T=50$, whereas Wald rejection remains $53\%$ at $T=30$ and $34\%$ at $T=50$. Moderately long panels can therefore support a reliable specification test before conventional asymptotic inference on $\theta$ becomes trustworthy.
When $T$ is still too short for asymptotic profile critical values, the parametric bootstrap removes most of the incidental-parameter size distortion (static logit, $T=5$, $n=500$: $98.9\%$ to $8.7\%$). At $T=5$ it does not fully restore the nominal level when $n$ is large. For instance, in static probit, profile bootstrap rejection rises with $n$ from $8.3\%$ at $n=125$ to $20.6\%$ at $n=1000$, and in dynamic probit it is $11.8\%$ at $n=1000$. By $T=10$ the same bootstrap is close to $5\%$ in all three designs, while asymptotic profile rejection at $n=1000$ is still $83.9\%$ in static logit.
\noindent
Power against misspecification is reported in Appendix~\ref{sec:add-sim-results}. Under the unit-scale heteroskedastic design of Subsection~\ref{sec:add-sim-hetero}, the PIM has strong power once $T$ is only moderately large. In static logit, bootstrap target rejection already reaches $98.2\%$ at $T=10$ and $n=500$, and $100\%$ at $T=10$ and $n=1000$; at the same grid points the Wald test based on $\hat{\theta}^{\mathfrak{p}}$ remains far weaker ($38.7\%$ and $60.6\%$). Even at $T=5$ and $n=1000$, bootstrap target rejection is $70.1\%$. Dynamic probit follows the same pattern once $T$ is moderate (profile Boot $92.7\%$ and target Boot $80.2\%$ at $T=10$ and $n=500$), though short-$T$ power is weaker than in the static designs. The omitted-interaction design of \citet{horowitz1994bootstrap} (Subsection~\ref{sec:add-sim-omitted}) is likewise powerful at moderate $T$: in static logit with $T=15$ and $n=1000$, bootstrap target rejection is $80.0\%$, and in static probit with $T=10$ and $n=500$ it is $72.4\%$. The wrong-link design (Subsection~\ref{sec:add-sim-link}) is harder to detect: for completed cells with $T\ge 10$, bootstrap rejection rates for the target and conditional PIM are moderate, typically between $10\%$ and $16\%$ at $n=1000$, reflecting that a link-function error is a relatively subtle deviation from the fitted logit model. Across these alternatives, the conditional likelihood again tracks the target closely in static logit, and correcting incidental-parameter size distortions through the bootstrap does not by itself create power where the correctly centered indicator has little signal.
\begin{landscape}
\begin{table}[p]
\centering
\caption{Rejection Rates for Static Logit Model}
\label{tab:logitrejectionrates}
\small
\setlength{\tabcolsep}{4.5pt}
\begin{tabular}{llcccccccccccccccc}
\toprule
& & \multicolumn{3}{c}{$n=125$} & & \multicolumn{3}{c}{$n=250$} & & \multicolumn{3}{c}{$n=500$} & & \multicolumn{3}{c}{$n=1000$} \\
\cmidrule(lr){3-5} \cmidrule(lr){7-9} \cmidrule(lr){11-13} \cmidrule(lr){15-17}
$T$ & Method & IM & Boot & Wald & & IM & Boot & Wald & & IM & Boot & Wald & & IM & Boot & Wald \\
\midrule
5 & Profile & 0.1040 & 0.0876 & 0.4277 & & 0.5968 & 0.0896 & 0.7120 & & 0.9894 & 0.0872 & 0.9432 & & 1.0000 & 0.0967 & 0.9978 \\
& Conditional & 0.0868 & 0.0586 & 0.0432 & & 0.0562 & 0.0564 & 0.0492 & & 0.0454 & 0.0606 & 0.0512 & & 0.0358 & 0.0627 & 0.0477 \\
& Target & 0.0966 & 0.0556 & 0.0450 & & 0.0646 & 0.0546 & 0.0498 & & 0.0564 & 0.0594 & 0.0488 & & 0.0429 & 0.0597 & 0.0491 \\
\addlinespace
10 & Profile & 0.0436 & 0.0668 & 0.2436 & & 0.1086 & 0.0594 & 0.4274 & & 0.3816 & 0.0578 & 0.7202 & & 0.8394 & 0.0494 & 0.9542 \\
& Conditional & 0.0930 & 0.0574 & 0.0524 & & 0.0742 & 0.0608 & 0.0490 & & 0.0520 & 0.0588 & 0.0506 & & 0.0436 & 0.0536 & 0.0516 \\
& Target & 0.0968 & 0.0574 & 0.0530 & & 0.0796 & 0.0654 & 0.0464 & & 0.0540 & 0.0568 & 0.0512 & & 0.0498 & 0.0592 & 0.0518 \\
\addlinespace
15 & Profile & 0.0458 & 0.0658 & 0.1648 & & 0.0516 & 0.0574 & 0.2910 & & 0.1468 & 0.0616 & 0.5308 & & 0.4082 & 0.0540 & 0.8330 \\
& Conditional & 0.0952 & 0.0566 & 0.0508 & & 0.0798 & 0.0618 & 0.0548 & & 0.0574 & 0.0602 & 0.0526 & & 0.0500 & 0.0580 & 0.0524 \\
& Target & 0.0986 & 0.0612 & 0.0500 & & 0.0804 & 0.0642 & 0.0550 & & 0.0570 & 0.0602 & 0.0526 & & 0.0472 & 0.0566 & 0.0502 \\
\addlinespace
20 & Profile & 0.0544 & 0.0630 & 0.1310 & & 0.0460 & 0.0574 & 0.2368 & & 0.0764 & 0.0534 & 0.4092 & & 0.2080 & 0.0558 & 0.7140 \\
& Conditional & 0.1024 & 0.0604 & 0.0464 & & 0.0674 & 0.0550 & 0.0512 & & 0.0500 & 0.0502 & 0.0524 & & 0.0548 & 0.0622 & 0.0512 \\
& Target & 0.1032 & 0.0600 & 0.0456 & & 0.0674 & 0.0544 & 0.0502 & & 0.0486 & 0.0500 & 0.0528 & & 0.0544 & 0.0614 & 0.0490 \\
\addlinespace
25 & Profile & 0.0592 & 0.0590 & 0.1178 & & 0.0482 & 0.0626 & 0.1984 & & 0.0592 & 0.0558 & 0.3630 & & 0.1326 & 0.0642 & 0.6142 \\
& Conditional & 0.0970 & 0.0590 & 0.0480 & & 0.0772 & 0.0616 & 0.0564 & & 0.0582 & 0.0546 & 0.0478 & & 0.0556 & 0.0608 & 0.0512 \\
& Target & 0.0952 & 0.0600 & 0.0468 & & 0.0770 & 0.0608 & 0.0564 & & 0.0594 & 0.0534 & 0.0478 & & 0.0546 & 0.0628 & 0.0518 \\
\addlinespace
30 & Profile & 0.0686 & 0.0638 & 0.1142 & & 0.0514 & 0.0636 & 0.1626 & & 0.0528 & 0.0616 & 0.2870 & & 0.0914 & 0.0546 & 0.5256 \\
& Conditional & 0.1050 & 0.0610 & 0.0534 & & 0.0784 & 0.0626 & 0.0502 & & 0.0552 & 0.0558 & 0.0536 & & 0.0486 & 0.0556 & 0.0518 \\
& Target & 0.1038 & 0.0618 & 0.0546 & & 0.0760 & 0.0610 & 0.0486 & & 0.0560 & 0.0560 & 0.0534 & & 0.0480 & 0.0552 & 0.0520 \\
\addlinespace
40 & Profile & 0.0712 & 0.0564 & 0.0932 & & 0.0574 & 0.0630 & 0.1396 & & 0.0492 & 0.0600 & 0.2160 & & 0.0598 & 0.0542 & 0.4208 \\
& Conditional & 0.0992 & 0.0626 & 0.0492 & & 0.0796 & 0.0632 & 0.0518 & & 0.0620 & 0.0570 & 0.0408 & & 0.0514 & 0.0582 & 0.0522 \\
& Target & 0.0996 & 0.0578 & 0.0494 & & 0.0780 & 0.0648 & 0.0514 & & 0.0596 & 0.0568 & 0.0424 & & 0.0514 & 0.0568 & 0.0512 \\
\addlinespace
50 & Profile & 0.0768 & 0.0582 & 0.0780 & & 0.0530 & 0.0576 & 0.1118 & & 0.0454 & 0.0598 & 0.1954 & & 0.0582 & 0.0622 & 0.3424 \\
& Conditional & 0.1000 & 0.0606 & 0.0494 & & 0.0756 & 0.0580 & 0.0518 & & 0.0618 & 0.0578 & 0.0482 & & 0.0562 & 0.0618 & 0.0506 \\
& Target & 0.1004 & 0.0624 & 0.0486 & & 0.0772 & 0.0600 & 0.0508 & & 0.0614 & 0.0572 & 0.0476 & & 0.0580 & 0.0602 & 0.0510 \\
\bottomrule
\end{tabular}
\end{table}
\end{landscape}
\begin{landscape}
\begin{table}[p]
\centering
\caption{Rejection Rates for Static Probit Model}
\label{tab:probitrejectionrates}
\small
\setlength{\tabcolsep}{4.5pt}
\begin{tabular}{llcccccccccccccccc}
\toprule
& & \multicolumn{3}{c}{$n=125$} & & \multicolumn{3}{c}{$n=250$} & & \multicolumn{3}{c}{$n=500$} & & \multicolumn{3}{c}{$n=1000$} \\
\cmidrule(lr){3-5} \cmidrule(lr){7-9} \cmidrule(lr){11-13} \cmidrule(lr){15-17}
$T$ & Method & IM & Boot & Wald & & IM & Boot & Wald & & IM & Boot & Wald & & IM & Boot & Wald \\
\midrule
5 & Profile & 0.0224 & 0.0826 & 0.7476 & & 0.2210 & 0.1232 & 0.9640 & & 0.9058 & 0.1640 & 0.9998 & & 1.0000 & 0.2062 & 1.0000 \\
& Target & 0.1852 & 0.0522 & 0.0456 & & 0.1252 & 0.0550 & 0.0474 & & 0.0894 & 0.0618 & 0.0532 & & 0.0560 & 0.0550 & 0.0456 \\
\addlinespace
10 & Profile & 0.0272 & 0.0616 & 0.4930 & & 0.0738 & 0.0630 & 0.7852 & & 0.4148 & 0.0522 & 0.9748 & & 0.9386 & 0.0542 & 0.9998 \\
& Target & 0.1426 & 0.0572 & 0.0504 & & 0.0948 & 0.0516 & 0.0474 & & 0.0654 & 0.0538 & 0.0534 & & 0.0462 & 0.0518 & 0.0490 \\
\addlinespace
15 & Profile & 0.0378 & 0.0658 & 0.3390 & & 0.0452 & 0.0608 & 0.6076 & & 0.1624 & 0.0586 & 0.8934 & & 0.5738 & 0.0490 & 0.9946 \\
& Target & 0.1278 & 0.0604 & 0.0514 & & 0.0858 & 0.0588 & 0.0464 & & 0.0632 & 0.0530 & 0.0550 & & 0.0490 & 0.0576 & 0.0530 \\
\addlinespace
20 & Profile & 0.0470 & 0.0604 & 0.2602 & & 0.0352 & 0.0584 & 0.4840 & & 0.0838 & 0.0526 & 0.7758 & & 0.3030 & 0.0508 & 0.9802 \\
& Target & 0.1176 & 0.0566 & 0.0504 & & 0.0860 & 0.0572 & 0.0540 & & 0.0612 & 0.0556 & 0.0484 & & 0.0552 & 0.0576 & 0.0436 \\
\addlinespace
25 & Profile & 0.0526 & 0.0590 & 0.2212 & & 0.0400 & 0.0610 & 0.3938 & & 0.0614 & 0.0546 & 0.6904 & & 0.1900 & 0.0586 & 0.9300 \\
& Target & 0.1102 & 0.0606 & 0.0540 & & 0.0884 & 0.0584 & 0.0516 & & 0.0618 & 0.0586 & 0.0498 & & 0.0596 & 0.0638 & 0.0476 \\
\addlinespace
30 & Profile & 0.0568 & 0.0576 & 0.1922 & & 0.0394 & 0.0584 & 0.3272 & & 0.0520 & 0.0572 & 0.5950 & & 0.1304 & 0.0586 & 0.8910 \\
& Target & 0.1172 & 0.0632 & 0.0542 & & 0.0826 & 0.0584 & 0.0480 & & 0.0658 & 0.0572 & 0.0514 & & 0.0604 & 0.0648 & 0.0520 \\
\addlinespace
40 & Profile & 0.0614 & 0.0536 & 0.1556 & & 0.0412 & 0.0546 & 0.2504 & & 0.0420 & 0.0540 & 0.4636 & & 0.0848 & 0.0598 & 0.7600 \\
& Target & 0.1062 & 0.0540 & 0.0566 & & 0.0816 & 0.0552 & 0.0504 & & 0.0636 & 0.0588 & 0.0468 & & 0.0536 & 0.0578 & 0.0492 \\
\addlinespace
50 & Profile & 0.0766 & 0.0612 & 0.1230 & & 0.0488 & 0.0580 & 0.2050 & & 0.0464 & 0.0606 & 0.3800 & & 0.0654 & 0.0592 & 0.6724 \\
& Target & 0.1136 & 0.0634 & 0.0516 & & 0.0842 & 0.0608 & 0.0538 & & 0.0678 & 0.0638 & 0.0492 & & 0.0534 & 0.0566 & 0.0470 \\
\bottomrule
\end{tabular}
\end{table}
\end{landscape}
\begin{landscape}
\begin{table}[p]
\centering
\caption{Rejection Rates for the Dynamic Probit Model}
\label{tab:dynamicprobitrejectionrates}
\small
\setlength{\tabcolsep}{4.5pt}
\begin{tabular}{llcccccccccccccccc}
\toprule
& & \multicolumn{3}{c}{$n=125$} & & \multicolumn{3}{c}{$n=250$} & & \multicolumn{3}{c}{$n=500$} & & \multicolumn{3}{c}{$n=1000$} \\
\cmidrule(lr){3-5} \cmidrule(lr){7-9} \cmidrule(lr){11-13} \cmidrule(lr){15-17}
$T$ & Method & IM & Boot & Wald & & IM & Boot & Wald & & IM & Boot & Wald & & IM & Boot & Wald \\
\midrule
5 & Profile & 0.0472 & 0.0470 & 0.9794 & & 0.0468 & 0.0848 & 1.0000 & & 0.4094 & 0.1092 & 1.0000 & & 0.9940 & 0.1182 & 1.0000 \\
& Target & 0.2864 & 0.0662 & 0.0478 & & 0.1924 & 0.0552 & 0.0490 & & 0.1256 & 0.0526 & 0.0466 & & 0.0812 & 0.0550 & 0.0430 \\
\addlinespace
10 & Profile & 0.0668 & 0.0680 & 0.8582 & & 0.0500 & 0.0706 & 0.9962 & & 0.1496 & 0.0638 & 1.0000 & & 0.7298 & 0.0446 & 1.0000 \\
& Target & 0.2130 & 0.0602 & 0.0476 & & 0.1428 & 0.0620 & 0.0426 & & 0.0990 & 0.0582 & 0.0532 & & 0.0674 & 0.0566 & 0.0470 \\
\addlinespace
15 & Profile & 0.0760 & 0.0620 & 0.7000 & & 0.0566 & 0.0686 & 0.9538 & & 0.0844 & 0.0632 & 0.9994 & & 0.3276 & 0.0564 & 1.0000 \\
& Target & 0.1802 & 0.0572 & 0.0436 & & 0.1312 & 0.0620 & 0.0476 & & 0.0856 & 0.0572 & 0.0528 & & 0.0620 & 0.0576 & 0.0470 \\
\addlinespace
20 & Profile & 0.0902 & 0.0626 & 0.5702 & & 0.0526 & 0.0620 & 0.8872 & & 0.0642 & 0.0598 & 0.9954 & & 0.1702 & 0.0590 & 1.0000 \\
& Target & 0.1788 & 0.0608 & 0.0510 & & 0.1216 & 0.0574 & 0.0486 & & 0.0874 & 0.0570 & 0.0496 & & 0.0600 & 0.0586 & 0.0488 \\
\addlinespace
25 & Profile & 0.0884 & 0.0550 & 0.4708 & & 0.0588 & 0.0606 & 0.7924 & & 0.0536 & 0.0594 & 0.9838 & & 0.1124 & 0.0586 & 1.0000 \\
& Target & 0.1688 & 0.0532 & 0.0490 & & 0.1190 & 0.0628 & 0.0486 & & 0.0842 & 0.0572 & 0.0550 & & 0.0710 & 0.0666 & 0.0498 \\
\addlinespace
30 & Profile & 0.0970 & 0.0532 & 0.4152 & & 0.0712 & 0.0654 & 0.7038 & & 0.0514 & 0.0580 & 0.9558 & & 0.0890 & 0.0612 & 0.9998 \\
& Target & 0.1600 & 0.0526 & 0.0548 & & 0.1260 & 0.0686 & 0.0546 & & 0.0822 & 0.0600 & 0.0488 & & 0.0648 & 0.0596 & 0.0442 \\
\addlinespace
40 & Profile & 0.1044 & 0.0546 & 0.3158 & & 0.0668 & 0.0558 & 0.5666 & & 0.0520 & 0.0596 & 0.8778 & & 0.0664 & 0.0628 & 0.9956 \\
& Target & 0.1558 & 0.0538 & 0.0492 & & 0.1064 & 0.0586 & 0.0546 & & 0.0832 & 0.0612 & 0.0488 & & 0.0704 & 0.0678 & 0.0418 \\
\addlinespace
50 & Profile & 0.1176 & 0.0598 & 0.2494 & & 0.0756 & 0.0598 & 0.4728 & & 0.0462 & 0.0528 & 0.7996 & & 0.0594 & 0.0578 & 0.9810 \\
& Target & 0.1500 & 0.0576 & 0.0500 & & 0.1152 & 0.0610 & 0.0576 & & 0.0774 & 0.0552 & 0.0504 & & 0.0604 & 0.0556 & 0.0496 \\
\bottomrule
\end{tabular}
\end{table}
\end{landscape}
\section{Conclusion}\label{sec:conclusion}
This paper extends the Information Matrix test to panel data models with fixed effects. Explicitly accounting for fixed effects allows us to separate functional misspecification from time-invariant unobserved heterogeneity. We show that the indicator fluctuates at rate $\sqrt{n}$ while estimation error vanishes as both $n$ and $T$ increase. As a consequence, testing the likelihood specification is much more robust to incidental parameter bias than tests about the common parameter, as the panel information matrix test statistic is asymptotically unbiased under rectangular asymptotics whereas Wald tests remain biased and thus unreliable. Our Monte Carlo evidence confirms these theoretical predictions and indicates that the PIM is reliable in moderately long panels, even when Wald tests remain badly sized. Finally, our statistic can be computed from first- and second-order likelihood derivatives only, removing some of the computational burden that made the information matrix test impractical in cross-sectional designs.
For future research, we note that the rate argument developed here is not special to the second Bartlett identity. Hence, other likelihood-based specification diagnostics that aggregate within-individual moment conditions at rate $\sqrt{n}$ should inherit the same incidental-parameter robustness, potentially unifying a class of panel specification tests. Moreover, the computational simplicity of the PIM and its robustness to incidental parameters make it a promising tool for assessing the credibility of structural panel models, especially in cases where likelihoods are simulated or otherwise costly to differentiate \citep{rust1987,aguirregabiria-mira-2007,dearing-blevins-2025}.
\newpage