EconBase
← Back to paper

A further look at Modified ML estimation of the panel AR(1) model with fixed effects and arbitrary initial conditions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,782 characters · 8 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A further look at Modified ML estimation of the panel AR(1) model with fixed effects and arbitrary initial conditions.

center[center omitted — 31 chars of source]

In this paper we consider two generalizations of Lancaster's (Review of Economic Studies, 2002) Modified Maximum Likelihood estimator (MMLE) for the panel AR(1) model with fixed effects, arbitrary initial conditions and strictly exogenous covariates when the time dimension of the panel, $T$, is fixed. When the autoregressive parameter ${\Greekmath 011A} =1,$ the limiting modified profile log-likelihood function for this model has a stationary point of inflection and ${\Greekmath 011A} $ is first-order underidentified but second-order identified. We show that, unlike the Random Effects and Transformed MLEs for this type of model, the generalized MMLEs are uniquely defined in finite samples w.p.1. for any value of $\left\vert {\Greekmath 011A} \right\vert \leq 1$. When $ {\Greekmath 011A} =1,$ the rate of convergence of the MMLEs is $N^{1/4},$ where $N$ is the cross-sectional dimension of the panel. We derive the limiting distributions of the MMLEs when ${\Greekmath 011A} =1$. They are generally asymmetric. We also show that Quasi LM tests that are based on the modified profile log-likelihood function and use its expected rather than observed Hessian for hypotheses that include a restriction on ${\Greekmath 011A} $, and confidence sets that are based on inverting these tests have correct asymptotic size in a uniform sense when $\left\vert {\Greekmath 011A} \right\vert \leq 1$. Finally, we investigate the finite sample properties of the MMLEs and the QLM test in a Monte Carlo study.

JEL\ classification: C11, C13, C23.

Keywords: asymptotic size, confidence set, dynamic panel data, expected Hessian, Modified Maximum Likelihood, Quasi Lagrange Multiplier (LM) test, rate of convergence, second-order identification, stationary point of inflection, uniform.

\baselineskip=18.8pt

\setcounter{page}{0} \thispagestyle{empty}

Introduction

In this paper we consider generalized Modified ML (cf. Neyman and Scott, 1948) estimators and identification robust inference methods for panel AR(1) models with fixed effects (FE), arbitrary initial conditions and strictly exogenous covariates when $T$ is fixed.

It is well known that the FE ML estimator for the autoregressive parameter $ {\Greekmath 011A} $ that is equal to the LSDV estimator is inconsistent when $T$ is fixed, cf. Nickell (1981).\footnote{ FE estimators only use data in differences.} To obtain a consistent FE\ estimator for ${\Greekmath 011A} $ (or for ${\Greekmath 0112} _{0}=({\Greekmath 011A} $ ${\Greekmath 011B} ^{2}$ ${\Greekmath 010C} ^{\prime })^{\prime }$, where ${\Greekmath 011B} ^{2}$ is the error variance and ${\Greekmath 010C} $ is the vector of coefficients of the covariates) based on the likelihood function for the model, Lancaster (2002) proposed a Bayesian approach that involves reparametrizing the fixed effects and integrating the new effects from the likelihood function using a uniform prior density. He defined his estimator for ${\Greekmath 011A} $ (or for ${\Greekmath 0112} _{0}$) as a local rather than a global maximizer of the resulting marginal (or joint) posterior density because this posterior density is improper and has a "global maximum" at $r=-\infty $ or $r=\infty $ for any sample size, cf. Dhaene and Jochmans (2016). Bun and Carree (2005) took a different route and proposed a bias-corrected LSDV estimator for ${\Greekmath 0112} _{0}$ with the correction based on formulae for the asymptotic biases of the LSDV estimators for ${\Greekmath 011A} $ and ${\Greekmath 010C} $. However, a version of their estimator is equal to Lancaster's estimator for ${\Greekmath 0112} _{0}$, cf. Dhaene and Jochmans (2016), and both of them can be viewed as a Modified ML estimator (MMLE). Bun and Carree (2005) investigated the finite sample properties of their estimator using various Monte Carlo experiments. They reported non-convergence of their estimator in about 40% of the replications in some experiments where $N=100,$ $T=6$ and ${\Greekmath 011A} =0.8.$ The possible non-existence of the MMLE is also related to the fact that the posterior density is improper. Specifically, when ${\Greekmath 011A} =1,$ the limiting modified profile log-likelihood function of $r$ has a stationary point of inflection at $r=1$, cf. Ahn and Thomas (2023). Dhaene and Jochmans (2016) addressed the non-existence problem by generalizing the MMLE for ${\Greekmath 0112} _{0} $: they defined their Adjusted Likelihood estimator for ${\Greekmath 011A} $ as the minimizer of the 2-norm of the gradient of modified profile log-likelihood function of $r$ subject to some constraints: two bounds on $r$ that depend on the LSDV estimate for ${\Greekmath 011A} $ and were imposed to ensure uniqueness of the estimator asymptotically, and a condition for a maximum. ${\Greekmath 011B} ^{2}$ and ${\Greekmath 010C} $ are estimated by solving the likelihood equations for ${\Greekmath 011B} ^{2}$ and ${\Greekmath 010C} $ and replacing $r$ by $\widehat{{\Greekmath 011A} }$ ($\widehat{ {\Greekmath 011A} }_{ADJ}$).

In this paper we discuss two generalized MMLEs for ${\Greekmath 0112} _{0}$ that exist with probability approaching one (w.p.a.1) as $N$ increases for any $ \left\vert {\Greekmath 011A} \right\vert \leq 1$.\footnote{ Note that w.p.a.1. means with probability approaching one, i.e., w.p.1 asymptotically.} The first generalized MMLE minimizes a quadratic form in the modified profile score vector for ${\Greekmath 0112} _{0}$ subject to $r\in \lbrack -1,\infty )$ and a condition for a maximum, while the second one minimizes the 2-norm of the modified profile score for ${\Greekmath 011A} $ on $[-1,\infty )$ subject to a condition for a maximum. The former MMLE depends on a weight matrix, while the latter MMLE\ depends on different constraints on $r$ than the generalized MMLE of Dhaene and Jochmans (2016) does.

While the likelihood functions of the Transformed MLE of Hsiao et al. (2002) and the Random Effects (RE) MLE of Chamberlain (1980) and Anderson and Hsiao (1982) may each have two local maxima, see Bun et al. (2017), we show that the modified profile likelihood function of $r$ has at most one local maximum on the interval $[-1,\infty ),$ and that if $\left\vert {\Greekmath 011A} \right\vert \leq 1,$ then both types of generalized MMLEs are uniquely defined w.p.1. and consistent. If the latter function has a local maximum on $[-1,\infty )$, then all the generalized MMLEs will be equal to each other. However, if that function has no local maximum on $[-1,\infty )$, then these estimators are different from each other and the value of the first type of generalized MMLE will depend on the choice of the weight matrix.

We also derive the limiting distributions of two generalized MMLEs. Similar to the cases of the REMLE and the Transformed MLE, which we will hereafter refer to as the FEMLE, if ${\Greekmath 011A} =1,$ then ${\Greekmath 011A} $ is only second-order identified by their objective functions and as a result the rate of convergence of the MMLEs for ${\Greekmath 011A} $ is $N^{1/4}$, cf. Ahn and Thomas (2023) and Kruiniger (2013). Our analysis for ${\Greekmath 011A} =1$ is closely related to Sargan (1983) for instrumental variable and ML estimators and also to Rotnitzky et al. (2000) for MLEs when a parameter is only second-order identified, although there are some important differences. We view the MMLEs as GMM estimators in order to derive their limiting distributions when ${\Greekmath 011A} =1$. Using an appropriate reparametrization of the modified profile likelihood, we find that if ${\Greekmath 011A} =1$ and the data are i.i.d. and normal, then the limiting distributions of the MMLEs are generally asymmetric unlike those of the RE- and FEMLE and other MLEs for parameters that are only second-order identified.

We also discuss inference methods related to the modified profile likelihood function. Wald tests, some versions of (Quasi) LM tests, and (Quasi) LR tests that are used for testing hypotheses involving ${\Greekmath 011A} $ and are based on the (reparametrized) modified\ profile likelihood function do not uniformly converge to their fixed parameter first-order limiting distributions when ${\Greekmath 011A} $ is close or equal to one, cf. Rotnitzky et al. (2000) and Bottai (2003). As a consequence these tests do not asymptotically have correct size in a uniform sense when $\left\vert {\Greekmath 011A} \right\vert \leq 1 $. Similarly to Kruiniger (2025a) in the case of (Quasi) LM tests related to the RE- and the FE(Q)MLE, we show that (Q)LM test statistics that are based on the modified profile log-likelihood function and use its expected rather than observed Hessian for hypotheses that include a restriction on ${\Greekmath 011A} $, and confidence sets that are based on inverting these tests have correct asymptotic size in a uniform sense when $\left\vert {\Greekmath 011A} \right\vert \leq 1$.

Monte Carlo results confirm that the QLM\ tests have correct size and show that when the data are i.i.d. and normal and $\left\vert {\Greekmath 011A} \right\vert <1$ , the MMLEs for ${\Greekmath 011A} $ can have a significantly smaller RMSE than the asymptotically efficient REMLE in panels as large as $T=9$ and $N=500$. When the data are not i.i.d. and normal, it is generally not possible to rank the Quasi MMLEs, the RE- and the FEQMLE in terms of asymptotic efficiency.

Both types of generalized MMLEs are also useful for estimating other models with parameters that may correspond to stationary points of inflection of the profile likelihood function. Examples of such models are the sample selection model and the stochastic production frontier model for a cross-section of units that are discussed in Lee and Chesher (1986) and models with skew-normal distributions, see e.g. Hallin and Ley (2014).

Dhaene and Jochmans (2016) discuss several alternative approaches to constructing modified (profile) objective functions for the nonstationary panel AR(1) model that yield estimators similar to Lancaster's MMLE. They have shown that their Adjusted Likelihood estimator for the nonstationary panel AR(1) model is uniquely defined asymptotically. However, they have not proven uniqueness of their estimator in finite samples nor have they derived its limiting distribution when ${\Greekmath 011A} =1$.\thinspace Furthermore, our paper is the first paper that provides MMLE-based inference methods that have correct uniform asymptotic size.

Hahn and Kuersteiner (2002) modified the LSDV\ estimator to remove bias up to order $O(T^{-1}).$ Other FE estimators for dynamic panel models include the first-difference (FD) instrumental variable estimator of Anderson and Hsiao (1982), the FE GMM estimators of Kruiniger (2001), the Maximum Invariant Likelihood estimator of Moreira (2009), the FDMLE of Kruiniger (2008) and the Panel Fully Aggregated Estimator of Han et al. (2014). The latter two estimators rely on covariance stationarity of the data when $\left\vert {\Greekmath 011A} \right\vert <1.$

Dovonon and Hall (2018) present a limiting distribution theory for GMM estimators when first-order identification fails but second-order identification holds. However, as shown in Kruiniger (2025b), their theory, specifically their Theorem 1(c), depends on a condition, namely $\Pr ( \mathbb{R} _{1}=0)=0$, that is only satisfied in the overidentified case and therefore cannot be used to derive the results in this paper. Kruiniger (2025b) completes the theory in Dovonon and Hall (2018) by adding a limiting distribution theory for the exactly identified case.

The paper is organised as follows. Section 2 presents the panel AR(1) model and the assumptions. Section 3 discusses existence, uniqueness and consistency of the generalized MMLEs as well as their asymptotic distributions. Section 4 discusses inference methods that have correct asymptotic size in a uniform sense. Section 5 studies the finite sample properties of the MMLEs and a (Q)LM test. Finally, section 6 offers some concluding remarks. Derivations and proofs can be found in the appendix.

The panel AR(1) model

We consider ML-type estimators for the panel AR(1) model with $K$ strictly exogenous covariates $x_{i,t,k},$ $k=1,...,K:\vspace{-0.12in}$

equation[equation omitted — 445 chars of source]

for $i=1,...,N$ and $t=1,...,T,$ where $x_{i,t}^{\prime }$ is the $t-th$ row of the $T\times K$ matrix $X_{i},$ ${\Greekmath 010B} _{i}$ is a fixed effect and $ {\Greekmath 0122} _{i,t}$ is an error term. We can also allow for time effects in the model.

Let $y_{i}=(y_{i,1}$ $...$ $y_{i,T})^{\prime },$ $y_{i,-1}=(y_{i,0}$ $...$ $ y_{i,T-1})^{\prime },$ ${\Greekmath 0122} _{i}=({\Greekmath 0122} _{i,1}$ $...$ $ {\Greekmath 0122} _{i,T})^{\prime }$ and $\overline{x}_{i}^{\prime }=T^{-1}{\Greekmath 0113} ^{\prime }X_{i}$, with ${\Greekmath 0113} $ equal to a $T-$vector of ones. If we let $ v_{i}=({\Greekmath 011A} -1)y_{i,0}+{\Greekmath 010B} _{i}+\overline{x}_{i}^{\prime }{\Greekmath 010C} $ for $ i=1,...,N,$ then the model in ((ref)) can also be written as $ y_{i}-y_{i,0}{\Greekmath 0113} ={\Greekmath 011A} (y_{i,-1}-y_{i,0}{\Greekmath 0113} )+QX_{i}{\Greekmath 010C} +v_{i}{\Greekmath 0113} +{\Greekmath 0122} _{i}$ for $i=1,...,N$, where $Q=I_{T}-T^{-1}{\Greekmath 0113} {\Greekmath 0113} ^{\prime }$ and $I_{T}$ is an identity matrix with dimension $T,$ cf. Lancaster (2002). We make the following assumption:

assumptionThe variable $y_{i,t}$ is generated by ((ref)) with (i) $T\geq 2$; (ii) $-1\leq {\Greekmath 011A} \leq 1$;\newline (iii) $\{({\Greekmath 0122} _{i}^{\prime },v_{i},(vech(QX_{i}))^{\prime })^{\prime }\}_{i=1}^{N}$ is a sequence of $i.i.d.$ random vectors with $E(v_{i})=0,$ $\qquad Var(v_{i})={\Greekmath 011B} _{v}^{2}<\infty $ and $E(X_{i}^{\prime }QX_{i})$ is a finite and positive definite matrix; and\newline (iv) ${\Greekmath 0122} _{i}\perp (v_{i},(vech(QX_{i}))^{\prime })^{\prime },$ $ E({\Greekmath 0122} _{i})=0$ and $Var({\Greekmath 0122} _{i})={\Greekmath 011B} ^{2}I_{T}<\infty ,$ $i=1,...,N$.

Thus we assume cross-sectional independence, strict exogeneity of the regressors in first-differences, homoskedasticity and no multicollinearity. On the other hand, we allow for ARCH and non-normality of the error terms, the ${\Greekmath 0122} _{i,t}.$

We require that $T\geq 2$ and ${\Greekmath 011A} \geq -1$ for identification. In economics the assumption ${\Greekmath 011A} \geq -1$ can reasonably be expected to hold when the covariates are strictly exogenous. The restrictive parametrization $ {\Greekmath 010B} _{i}=(1-{\Greekmath 011A} ){\Greekmath 0116} _{i}$ and ${\Greekmath 010C} =(1-{\Greekmath 011A} )\check{{\Greekmath 010C}}$ prevents the fixed effects and the means of the individual regressors from turning into trends at ${\Greekmath 011A} =1$ and thereby avoids a discontinuity in the data generating process at ${\Greekmath 011A} =1$. These restrictions and the restriction $ {\Greekmath 011A} \leq 1$ are only imposed on the DGP but not in estimation.

We are interested in consistent estimation of the common parameters ${\Greekmath 011A} ,$ ${\Greekmath 011B} ^{2}$ and ${\Greekmath 010C} $ under large $N$, fixed $T$ asymptotics. We will treat the individual effects as nuisance parameters. We will work with a Gaussian homoskedastic (quasi-)likelihood but we note that consistency of the MMLEs (for ${\Greekmath 011A} $ and ${\Greekmath 010C} $) does not depend on normality or cross-sectional homoskedasticity of the errors.

Modified ML estimation of the panel AR(1) model

Conditional on $y_{i,0}$ and $X_{i},$ $i=1,...,N$ and normalized by $N$, the Gaussian FE log-likelihood function for the model in ((ref)) is, up to an additive constant, given by:

equation[equation omitted — 203 chars of source]

To obtain a consistent FE estimator for ${\Greekmath 0112} _{0}$ based on ((ref) ), Lancaster (2002) proposed a Bayesian approach that involves using a reparametrization of the fixed effects, which aims to achieve information orthogonality (but fails to do so when covariates are present), and integrating the new effects from the likelihood function using a uniform prior density. He defines his estimator for ${\Greekmath 0112} _{0}$ as a local maximum of the joint posterior density. Letting ${\Greekmath 0112} =(r$ $s^{2}$ $b)^{\prime }$, his joint posterior log-density for the model in ((ref)), normalized by $N$, which can be interpreted as a (normalized) modified profile log-likelihood function, is given by:

align[align omitted — 558 chars of source]

and the corresponding modified profile likelihood equations are given by:

eqnarray[eqnarray omitted — 544 chars of source]

Note that the joint posterior density is not proper.

Let $\widehat{{\Greekmath 0112} }_{LAN}$ denote Lancaster's estimator for ${\Greekmath 0112} _{0}$ and let $\Theta _{N}$ be the set of roots of $\frac{\partial \widetilde{l} _{N}}{\partial {\Greekmath 0112} }=0$ corresponding to local maxima of $\widetilde{l} _{N}$ on $\Omega $ which is an open subset of $ \mathbb{R} \times \mathbb{R} ^{+}\times \mathbb{R} ^{K}.$ Thus $\widehat{{\Greekmath 0112} }_{LAN}\in \Theta _{N}$ unless $\Theta _{N}$ is empty, in which case (we will say that) $\widehat{{\Greekmath 0112} }_{LAN}$ does not exist. In that case Lancaster effectively puts $\widehat{{\Greekmath 0112} }_{LAN}= \mathbf{0,}$ see his consistency proof. This `trick' ensures that $\widehat{ {\Greekmath 0112} }_{LAN}$ always exists so that one can consider whether $\widehat{ {\Greekmath 0112} }_{LAN}$ is a consistent estimator for ${\Greekmath 0112} _{0}.$ Note that none of the roots of $\frac{\partial \widetilde{l}_{N}}{\partial {\Greekmath 0112} }=0$ correspond to the global maxima that can occur at $r=\infty $ and, if $T$ is odd, at $r=-\infty .$

Lancaster showed that $\widetilde{l}_{N}({\Greekmath 0112} )$ converges uniformly in probability to a nonstochastic differentiable function of ${\Greekmath 0112} ,$ say $ \widetilde{l}({\Greekmath 0112} ),$ and that $\frac{\partial \widetilde{l}({\Greekmath 0112} )}{ \partial {\Greekmath 0112} }|_{{\Greekmath 0112} _{0}}=0.$ Next we derive necessary and sufficient conditions for negative definiteness of the Hessian of $\widetilde{l}({\Greekmath 0112} )$ at ${\Greekmath 0112} _{0},$ viz.:

equation[equation omitted — 603 chars of source]

where $\Sigma _{zqz}=$ p$\lim_{N\rightarrow \infty }N^{-1}\sum_{i=1}^{N} \widetilde{Z}_{i}Q\widetilde{Z}_{i},$ $\Sigma _{xqx}=$ p$\lim_{N\rightarrow \infty }N^{-1}\sum_{i=1}^{N}X_{i}^{\prime }QX_{i}$ and $\Sigma _{xqz}=$ p$ \lim_{N\rightarrow \infty }N^{-1}\sum_{i=1}^{N}X_{i}^{\prime }Q\widetilde{Z} _{i}$ with $\widetilde{Z}_{i}={\Greekmath 0127} v_{i}+\Phi QX_{i}{\Greekmath 010C} ,$

equation[equation omitted — 630 chars of source]

It follows from lemma 4.1 in Dhaene and Jochmans (2016) that if $T=2$ and $ \Sigma _{zqz}>0$ (so that ${\Greekmath 011A} \neq 1$) or if $T>2$ and ${\Greekmath 011A} \neq 1$, then $MH$ is negative definite so that $\widetilde{l}({\Greekmath 0112} )$ has a local maximum at ${\Greekmath 0112} _{0}$.\footnote{ Their lemma 4.1 implies that ${\Greekmath 0118} ^{\prime \prime }({\Greekmath 011A} )-(T-1)^{-1}tr(\Phi ^{\prime }Q\Phi )+2({\Greekmath 0118} ^{\prime }({\Greekmath 011A} ))^{2}\leq 0$ with equality if and only if $T=2$ or ${\Greekmath 011A} =1$.} Kruiniger (2001) had already shown that if $ {\Greekmath 011A} =1$ and $T\geq 2,$ then $MH$ is singular. Moreover, Ahn and Thomas (2023) have shown that $\widetilde{l}({\Greekmath 0112} )$ actually has a stationary point of inflection when ${\Greekmath 011A} =1$ rather than a local maximum. This property is related to the fact that the posterior density is not proper. Later on, in the context of Theorem 1 below, we will show that if ${\Greekmath 011A} =1,$ $\widetilde{l}_{N}$ may not have any local maximum on $\widetilde{\Omega } =[-1,\infty )\times (0,\infty )\times \mathbb{R} ^{K}$ asymptotically, so that $\widehat{{\Greekmath 0112} }_{LAN}$ is inconsistent. \footnote{ Lancaster's model is $y_{i}={\Greekmath 011A} y_{i,-1}+X_{i}{\Greekmath 010C} +{\Greekmath 010B} _{i}{\Greekmath 0113} +{\Greekmath 0122} _{i}$ without the restrictions ${\Greekmath 010C} =(1-{\Greekmath 011A} )\check{{\Greekmath 010C}}$ and ${\Greekmath 010B} _{i}=(1-{\Greekmath 011A} ){\Greekmath 0116} _{i}.$ Therefore, if ${\Greekmath 011A} =1$ and ${\Greekmath 010C} \neq 0,$ then the probability limit of the Hessian of his modified log-likelihood function at ${\Greekmath 0112} _{0}$ is still negative definite and his estimator is consistent. However, if ${\Greekmath 011A} =1,$ ${\Greekmath 010C} =0$ and ${\Greekmath 010B} _{i}=0 $ for $i=1,...,N,$ then his estimator is inconsistent.} $\widehat{ {\Greekmath 0112} }_{LAN}$ has two more drawbacks. Firstly, $\widetilde{l}_{N}({\Greekmath 0112} )$ may not have any local maximum in small samples, in which case $\widehat{ {\Greekmath 0112} }_{LAN}$ does not exist. This may happen when ${\Greekmath 011A} $ is close or equal to unity. Secondly, Lancaster did not rule out that $\widetilde{l} _{N}({\Greekmath 0112} )$ and $\widetilde{l}({\Greekmath 0112} )$ have multiple local maxima on $ \Omega $ and he did not explain how to find the consistent estimator if that were the case.

Generalized Modified ML estimators

We will now introduce two generalizations of $\widehat{{\Greekmath 0112} }_{LAN}$. We have assumed that $\left\vert {\Greekmath 011A} \right\vert \leq 1.$ Under this assumption we will be able to show below that $\widetilde{l}_{N}({\Greekmath 0112} )$ can have one local maximum on $\widetilde{\Omega }$ at most. To ensure that the MMLE for ${\Greekmath 0112} _{0}$ is also defined in most cases where $\Theta _{N}\cap \widetilde{\Omega }=\varnothing ,$ we will generalize its definition as follows:

equation[equation omitted — 656 chars of source]

where $W_{N}$ is a positive definite (PD) symmetric weight matrix and plim$ _{N\rightarrow \infty }W_{N}=W$ where $W$ is PD. Thus our MMLE is defined as the minimizer of a quadratic form in the modified profile score vector, $ \frac{\partial \widetilde{l}_{N}}{\partial {\Greekmath 0112} },$ subject to the Hessian of $\widetilde{l}_{N}$ being negative semi-definite. If $\widetilde{l} _{N}({\Greekmath 0112} )$ has a local maximum, then our MMLE for ${\Greekmath 0112} _{0}$ does not depend on $W_{N}$ and is equal to $\widehat{{\Greekmath 0112} }_{LAN}$. Theorem 1 below asserts that $\widehat{{\Greekmath 0112} }_{W}$ exists w.p.a.1, is uniquely defined (given $W_{N}$) w.p.1 and is consistent for any ${\Greekmath 0112} _{0}\in \widetilde{ \Omega }$.

Note that among the likelihood equations in ((ref)) only the one for $r$ is modified. Hence, when solving $\Psi _{{\Greekmath 010C} }({\Greekmath 0112} )=0$ for $b$ we obtain the unique solution $\widehat{{\Greekmath 010C} }(r)=(\sum_{i=1}^{N}X_{i}^{ \prime }QX_{i})^{-1}\times $ $\sum_{i=1}^{N}X_{i}^{\prime }Q(y_{i}-ry_{i,-1}) $ and when solving $\Psi _{{\Greekmath 011B} ^{2}}({\Greekmath 0112} )=0$ for $ s^{2}$ we obtain the unique solution $\widehat{{\Greekmath 011B} } ^{2}(r,b)=(T-1)^{-1}N^{-1}\sum_{i=1}^{N}(y_{i}-ry_{i,-1}-X_{i}b)^{\prime }Q(y_{i}-ry_{i,-1}-X_{i}b).$ Let $\widehat{{\Greekmath 0112} }(r)=(r,\widehat{{\Greekmath 011B} } ^{2}(r,\widehat{{\Greekmath 010C} }(r)),\widehat{{\Greekmath 010C} }(r))^{\prime },$ then the (normalized) modified profile log-likelihood function of $r$, $\widetilde{l} _{N}^{c}(r),$ is defined by the equality $\widetilde{l}_{N}^{c}(r)= \widetilde{l}_{N}(\widehat{{\Greekmath 0112} }(r)),$ i.e., $\widetilde{l}_{N}^{c}(r)= \widetilde{l}_{N}(r,\widehat{{\Greekmath 011B} }^{2}(r,\widehat{{\Greekmath 010C} }(r)),\widehat{ {\Greekmath 010C} }(r)).$

An alternative generalized MMLE for ${\Greekmath 0112} _{0}$, which is based on $ \widetilde{l}_{N}^{c}(r)$, is given by $\widehat{{\Greekmath 0112} }_{C}$ with \footnote{ One can also define a class of MMLEs where only $s^{2}$ is profiled out but not $b$.}

gather[gather omitted — 727 chars of source]

The Adjusted Likelihood estimator of Dhaene and Jochmans (2016), viz. $ \widehat{{\Greekmath 0112} }_{ADJ}$, is defined similarly to $\widehat{{\Greekmath 0112} }_{C}$: $ \widehat{{\Greekmath 011A} }_{ADJ}=\arg \min_{r\in \mathcal{E}}\left( \frac{\partial \widetilde{l}_{N}^{c}(r)}{\partial r}\right) ^{2}$ s.t. $\frac{\partial ^{2} \widetilde{l}_{N}^{c}(r)}{\partial r^{2}}\leq 0$ with $\mathcal{E=\{}r:$ $ \mathcal{(}r\mathcal{-}\widehat{{\Greekmath 011A} }_{LSDV})^{2}(-\frac{\partial ^{2}l_{N}^{c}(r)}{\partial r^{2}}|_{\widehat{{\Greekmath 011A} }_{LSDV}})\leq 1\}$ where $ l_{N}^{c}(r)=l_{N}(r,\widehat{{\Greekmath 011B} }^{2}(r,\widehat{{\Greekmath 010C} }(r)),\widehat{ {\Greekmath 010C} }(r))$. $\widehat{{\Greekmath 010C} }_{ADJ}=\widehat{{\Greekmath 010C} }(\widehat{{\Greekmath 011A} } _{ADJ}) $ and $\widehat{{\Greekmath 011B} }_{ADJ}^{2}=\widehat{{\Greekmath 011B} }^{2}(\widehat{ {\Greekmath 011A} }_{ADJ},\widehat{{\Greekmath 010C} }_{ADJ}).$ However, they have only shown that $ \widehat{{\Greekmath 0112} }_{ADJ}$ is uniquely defined asymptotically. On the other hand, Theorem 1 below states that when $\left\vert {\Greekmath 011A} \right\vert \leq 1,$ $\widehat{{\Greekmath 0112} }_{C}$ is uniquely defined in finite samples. It is unclear how rare the event $\widehat{{\Greekmath 011A} }_{C}\notin \mathcal{E}$ is in finite samples when $\left\vert {\Greekmath 011A} \right\vert \leq 1$, but if $\widehat{{\Greekmath 011A} } _{C}\notin \mathcal{E}$, then $\widehat{{\Greekmath 011A} }_{C}$ is likely to be a better estimate than $\widehat{{\Greekmath 0112} }_{ADJ}$ because either $\left\vert \frac{ \partial \widetilde{l}_{N}^{c}(r)}{\partial r}|_{\widehat{{\Greekmath 011A} } _{C}}\right\vert <\left\vert \frac{\partial \widetilde{l}_{N}^{c}(r)}{ \partial r}|_{\widehat{{\Greekmath 011A} }_{ADJ}}\right\vert $ or $\widehat{{\Greekmath 0112} } _{ADJ}<-1.$

There exists no $W_{N}$ such that the $\widehat{{\Greekmath 0112} }_{W}$ estimator always equals the $\widehat{{\Greekmath 0112} }_{C}$ estimator: if $\frac{\partial \widetilde{l}_{N}^{c}(r)}{\partial r}|_{\widehat{{\Greekmath 011A} }_{C}}=0,$ then $\frac{ \partial \widetilde{l}_{N}({\Greekmath 0112} )}{\partial {\Greekmath 0112} }|_{\widehat{{\Greekmath 0112} } _{W}}=0$\ and both estimates of ${\Greekmath 0112} $ are equal, but if $\frac{\partial \widetilde{l}_{N}^{c}(r)}{\partial r}|_{\widehat{{\Greekmath 011A} }_{C}}$ $\neq 0,$ then $\frac{\partial \widetilde{l}_{N}({\Greekmath 0112} )}{\partial {\Greekmath 0112} }|_{\widehat{ {\Greekmath 0112} }_{W}}\neq 0$\ and the two estimates of ${\Greekmath 0112} $ are unequal although the value of $\widehat{{\Greekmath 0112} }_{W}$ will be close to that of $ \widehat{{\Greekmath 0112} }_{C}$ for $W_{N}$ that give relatively little weight to $ \frac{\partial \widetilde{l}_{N}({\Greekmath 0112} )}{\partial r}.$ We also consider a variation on $\widehat{{\Greekmath 0112} }_{W}$ with the first element of $\frac{ \partial \widetilde{l}_{N}({\Greekmath 0112} )}{\partial {\Greekmath 0112} }$ replaced by $\frac{ \partial \widetilde{l}_{N}^{c}(r)}{\partial r}.$ We call this MMLE $\widehat{ {\Greekmath 0112} }_{F}.$ If $W_{N}=diag(\infty ,\underline{W}_{N,2,2})$ and the elements of $\underline{W}_{N,2,2}$ are finite, then $\widehat{{\Greekmath 0112} } _{F}=\widehat{{\Greekmath 0112} }_{C}.$

In the appendix we show that $\widetilde{l}_{N}^{c}(r)$ converges uniformly in probability to a nonstochastic differentiable function of $r,$ say $ \widetilde{l}^{c}(r),$ that $\frac{\partial \widetilde{l}^{c}(r)}{\partial r} |_{{\Greekmath 011A} }=0$ and that $\frac{\partial ^{2}\widetilde{l}^{c}(r)}{\partial r^{2}}|_{{\Greekmath 011A} }\leq 0,$ with equality holding if ${\Greekmath 011A} =1$ or if $T=2$ and $ {\Greekmath 011B} _{v}^{2}={\Greekmath 010C} =0$ (i.e., $\Sigma _{zqz}=0$). Thus, similar to $ \widetilde{l}({\Greekmath 0112} ),$ $\widetilde{l}^{c}(r)$ has a local maximum at ${\Greekmath 011A} $ when ${\Greekmath 011A} \neq 1$ and, in case $T=2$, $\Sigma _{zqz}>0$. In the appendix we also show that $\widetilde{l}^{c}(r)$ has a stationary point of inflection at ${\Greekmath 011A} $ when ${\Greekmath 011A} =1$. To simplify the exposition we assume in the remainder of this paper that if $T=2$ and ${\Greekmath 011A} \neq 1,$ then either $ {\Greekmath 011B} _{v}^{2}>0$ or ${\Greekmath 010C} \neq 0$ so that $\Sigma _{zqz}>0.$

Note that $\widehat{{\Greekmath 0112} }_{C}$ would only fail to exist in the extremely unlikely case that $\frac{\partial ^{2}\widetilde{l}_{N}^{c}(r)}{\partial r^{2}}>0$ on the entire interval $[-1,\infty ).$ Similarly, $\widehat{{\Greekmath 0112} }_{W}$ and $\widehat{{\Greekmath 0112} }_{F}$ would only fail to exist in the extremely unlikely case that for no ${\Greekmath 0112} \in \widetilde{\Omega },$ $x^{\prime }\left( \frac{\partial ^{2}\widetilde{l}_{N}({\Greekmath 0112} )}{\partial {\Greekmath 0112} \partial {\Greekmath 0112} ^{\prime }}\right) x\leq 0$ $\forall x\in \mathbb{R} ^{2+K}.$ \footnote{ One could ensure that $\widehat{{\Greekmath 0112} }_{W},$ $\widehat{{\Greekmath 0112} }_{F}$ and $ \widehat{{\Greekmath 0112} }_{C}$ are always defined by replacing them by $\widehat{ {\Greekmath 0112} }(\widehat{{\Greekmath 011A} }_{ML}+\frac{3}{T+1})$ in these improbable cases, where $-\frac{3}{T+1}$ is the asymptotic bias of $\widehat{{\Greekmath 011A} }_{ML}$ when ${\Greekmath 011A} =1$.\ The rationale for this proposed solution is that the non-existence problem most likely only occurs (if ever) when the sample size is very small and ${\Greekmath 011A} $ is close or equal to unity.} The second-order conditions $\frac{\partial ^{2}\widetilde{l}_{N}^{c}(r)}{\partial r^{2}}\leq 0$ and $x^{\prime }\left( \frac{\partial ^{2}\widetilde{l}_{N}({\Greekmath 0112} )}{ \partial {\Greekmath 0112} \partial {\Greekmath 0112} ^{\prime }}\right) x\leq 0$ $\forall x\in \mathbb{R} ^{2+K}$ are a crucial part of the definitions of $\widehat{{\Greekmath 0112} }_{C},$ $ \widehat{{\Greekmath 0112} }_{W}$ and $\widehat{{\Greekmath 0112} }_{F}$ because $\widetilde{l} _{N}^{c}(r)$ and $\widetilde{l}_{N}(r)$ may attain a minimum on $[-1,\infty ) $ and $\widetilde{\Omega }$, respectively, see lemma 1 in the appendix.

The next theorem asserts uniqueness and consistency of $\widehat{{\Greekmath 0112} } _{W},$ $\widehat{{\Greekmath 0112} }_{F}$ and $\widehat{{\Greekmath 0112} }_{C}$:

theoremLet Assumption 1 hold. Then the Modified MLEs $\widehat{{\Greekmath 0112} }_{W},$ $ \widehat{{\Greekmath 0112} }_{F}$ and $\widehat{{\Greekmath 0112} }_{C}$ for ${\Greekmath 0112} _{0}$ are uniquely defined w.p.1 when they exist, exist w.p.a.1 and are consistent.

If $-1\leq {\Greekmath 011A} <1,$ $\lim_{N\rightarrow \infty }\Pr (\Theta _{N}\cap \widetilde{\Omega }=\varnothing )=0$, i.e., $\widehat{{\Greekmath 0112} }_{LAN}$ exists w.p.a.1. In this case $\widehat{{\Greekmath 0112} }_{LAN}$ is also unique w.p.1. (if it exists) and consistent. However, if ${\Greekmath 011A} =1,$ $\lim_{N\rightarrow \infty }\Pr (\Theta _{N}\cap \widetilde{\Omega }=\varnothing )>0$ by lemma 4 in the appendix (and ${\Greekmath 0112} _{0}\neq \mathbf{0}$), i.e., $\widehat{{\Greekmath 0112} }_{LAN}$ may not exist even asymptotically, which implies that $\widehat{{\Greekmath 0112} } _{LAN}$ is inconsistent.

When $-1\leq {\Greekmath 011A} <1,$ the first-order, fixed parameter asymptotic distributions of $\widehat{{\Greekmath 0112} }_{W},$ $\widehat{{\Greekmath 0112} }_{F},$ $\widehat{ {\Greekmath 0112} }_{C}$ and $\widehat{{\Greekmath 0112} }_{LAN}$ are the same and given by (cf. Kruiniger, 2001):$\vspace{-0.14in}$

equation[equation omitted — 216 chars of source]

where $MH$ is given in ((ref)) and under normality of the ${\Greekmath 0122} _{i}$ $MIM$ (Modified Information Ma- trix) equals:\footnote{ To derive ((ref)) we have used that if ${\Greekmath 0122} _{i}|(v_{i},QX_{i})\sim N(0,{\Greekmath 011B} ^{2}I_{T}),$ then for any constant $ T\times T$ matrices $M_{1}$ and $M_{2},$ $E({\Greekmath 0122} _{i}^{\prime }M_{1}{\Greekmath 0122} _{i}{\Greekmath 0122} _{i}^{\prime }M_{2}{\Greekmath 0122} _{i})={\Greekmath 011B} ^{4}(tr(M_{1})tr(M_{2})+tr(M_{1}M_{2}+M_{1}^{\prime }M_{2}))$.}

equation[equation omitted — 583 chars of source]

It can easily be checked that $tr(Q\Phi Q\Phi )\neq -(T-1){\Greekmath 0118} ^{\prime \prime }({\Greekmath 011A} )$ and hence $MH\neq -MIM$.

If $T=2$, $\widehat{{\Greekmath 011A} }_{LAN}$ is equal to the FEMLE for ${\Greekmath 011A} $ that has been proposed by Hsiao et al. (2002), henceforth $\widehat{{\Greekmath 011A} }_{FEML}$, but if $T>2,$ the data are i.i.d. and normal and $\left\vert {\Greekmath 011A} \right\vert <1,$ $\widehat{{\Greekmath 011A} }_{LAN}$ is asymptotically less efficient than $\widehat{{\Greekmath 011A} }_{FEML}$, see Ahn and Thomas (2023); when the data are not i.i.d and normal, $\widehat{{\Greekmath 011A} }_{LAN}$ may be asymptotically more efficient than $\widehat{{\Greekmath 011A} }_{FEML}$.

If ${\Greekmath 011A} =1,$ $\det (MIM)\neq 0$ but $\frac{\partial ^{2}\widetilde{l}^{c}(r) }{\partial r^{2}}|_{{\Greekmath 011A} }=0$ and $\det (MH)=0.$ Thus ${\Greekmath 011A} $ and ${\Greekmath 0112} $ are first- order underidentified when ${\Greekmath 011A} =1$. Although we cannot directly apply the results of Rot- nitzky et al. (2000), who developed an asymptotic theory for MLEs when the information matrix is singular, to $ \widehat{{\Greekmath 0112} }_{W},$ $\widehat{{\Greekmath 0112} }_{F}$ and $\widehat{{\Greekmath 0112} }_{C}$ when ${\Greekmath 011A} =1$, because they are Modified MLEs and $\det (MIM)\neq 0$, arguments similar to theirs suggest that these MMLEs have a slower than $\sqrt{N}$ rate of convergence and that their limiting distributions are non-standard. When deriving the limiting distributions of $ \widehat{{\Greekmath 0112} }_{C}$ and $\widehat{{\Greekmath 0112} }_{F}$ for ${\Greekmath 011A} =1$ below, we will view the MMLEs as GMM estimators.\footnote{ The limiting distribution of $\widehat{{\Greekmath 0112} }_{W}$ for ${\Greekmath 011A} =1$ can be derived using results in Kruiniger (2025b).} If ${\Greekmath 011A} $ is close to 1, $\det (MH)$\ and $\frac{\partial ^{2}\widetilde{l}^{c}(r)}{\partial r^{2}}|_{{\Greekmath 011A} } $ are close to zero and the MMLEs will have a "weak moment conditions" problem.

The limiting distributions of $\protect\widehat{\protect{\Greekmath 0112} } _{C}$ and $\protect\widehat{\protect{\Greekmath 0112} }_{F}$ when $\protect{\Greekmath 011A} =1$

W.p.a.1 $\widehat{{\Greekmath 011A} }_{C}$ is a solution of the first-order condition (f.o.c.) $G_{N}^{c}(r)\equiv \frac{\partial ^{2}\widetilde{l}_{N}^{c}(r)}{ \partial r^{2}}\frac{\partial \widetilde{l}_{N}^{c}(r)}{\partial r}=0.$ Using a Taylor expansion of $G_{N}^{c}(\widehat{{\Greekmath 011A} }_{C})$ around $r=1,$ we show in the appendix that when ${\Greekmath 011A} =1,$ $N^{1/4}(\widehat{{\Greekmath 011A} } _{C}-1)=O_{p}(1),$ i.e., the rate of convergence of $\widehat{{\Greekmath 011A} }_{C}$ is at least $N^{1/4}$. This quartic root rate of convergence reflects the fact that $\frac{\partial ^{2}\widetilde{l}^{c}(1)}{\partial r^{2}}=0$ and $\frac{ \partial ^{3}\widetilde{l}^{c}(1)}{\partial r^{3}}=\frac{T(T-1)(T+1)}{12} \neq 0$, which means that ${\Greekmath 011A} $ is second-order identified when ${\Greekmath 011A} =1$, and is in line with results in Sargan (1983), Rotnitzky et al. (2000), Ahn and Thomas (2023), Madsen (2009), Dovonon and Renault (2013) and Kruiniger (2013) who also study estimation when a parameter is only second-order identified. Note that this rate is faster than the $N^{1/6}$\thinspace -rate of the MLEs of the parameters that correspond to the inflection point of the likelihood functions of the sample selection model and the stochastic production frontier model for a cross-section that are discussed in Lee and Chesher (1986) and the models with skew-normal distributions that are discussed in Hallin and Ley (2014).

Next we discuss the derivation of the limiting distribution of $\widehat{ {\Greekmath 0112} }_{C}$ when ${\Greekmath 011A} =1.$ Let $M_{N}^{c}(r)=N\left( \frac{\partial \widetilde{l}_{N}^{c}(r)}{\partial r}\right) ^{2}.$ Analogously to Sargan (1983) and Rotnitzky et al. (2000) consider the following Taylor expansion of $M_{N}^{c}(r)$ around $r=1\vspace{-0.1in}:$

equation[equation omitted — 178 chars of source]

where $P_{3,N}(N^{1/4}(r-1))$ is a polynomial in $N^{1/4}(r-1)$ with coefficients that are $o_{p}(1)$. Let $\widehat{{\Greekmath 011A} }=\widehat{{\Greekmath 011A} }_{C}.$ Substituting $\widehat{{\Greekmath 011A} }$ for $r$ in ((ref)) we obtain

eqnarray[eqnarray omitted — 506 chars of source]

where $R_{1,N}^{c}(N^{1/4}(\widehat{{\Greekmath 011A} }-1))=o_{p}(1).$

Let $Z_{1,N}=\left( -\frac{1}{2}\frac{\partial ^{3}\widetilde{l}_{N}^{c}(1)}{ \partial r^{3}}\right) ^{-1}N^{1/2}\left( \frac{\partial \widetilde{l} _{N}^{c}(1)}{\partial r}\right) .$ In the proof of Theorem 2 we show that $ Z_{1,N}=O_{p}(1)$ and that there exists a sequence $\{U_{N}\}$ with $ U_{N}=O_{p}(N^{-1/2})$ such that if $Z_{1,N}+U_{N}>0,$ then $M_{N}^{c}(r)$\ has two local minima attained at values $\widetilde{{\Greekmath 011A} }$ such that $ N^{1/2}(\widetilde{{\Greekmath 011A} }-1)^{2}=Z_{1,N}+o_{p}(1),$ whereas if $ Z_{1,N}+U_{N}<0,$ then $M_{N}^{c}(r)$\ has one local minimum attained at $r= \widehat{{\Greekmath 011A} }$ with $N^{1/2}(\widehat{{\Greekmath 011A} }-1)^{2}=o_{p}(1).$ Furthermore, when $Z_{1,N}+U_{N}>0,$ the sign of $N^{1/4}(\widehat{{\Greekmath 011A} }-1)$ is determined by the remainder $R_{1,N}^{c}(N^{1/4}(\widehat{{\Greekmath 011A} }-1))$.

To obtain the limiting distribution of $\widehat{{\Greekmath 0112} }_{C}$ when ${\Greekmath 011A} =1$ we use the following new parametrization (indicated by the subscript $n$), cf. Kruiniger (2013): ${\Greekmath 0112} _{n}=(r_{n},s_{n}^{2},b_{n}^{\prime })^{\prime }$ where $r_{n}=r,$ $s_{n}^{2}=s^{2}/r$ and $b_{n}=b.$ Noting that we can express the elements of ${\Greekmath 0112} $ as functions of the elements of ${\Greekmath 0112} _{n},$ viz. ${\Greekmath 0112} ={\Greekmath 0112} ({\Greekmath 0112} _{n})=(r_{n},s_{n}^{2}r_{n},b_{n}^{\prime })^{\prime },$ the reparameterized modified profile log-likelihood function is given by $\widetilde{l} _{n,N}({\Greekmath 0112} _{n})=\widetilde{l}_{N}({\Greekmath 0112} ({\Greekmath 0112} _{n})).$ Similarly to Lancaster (2002), it can be shown that $\widetilde{l}_{n,N}({\Greekmath 0112} _{n})$ converges uniformly in probability to a nonstochastic continuous function of ${\Greekmath 0112} _{n},$ i.e. $\widetilde{l}_{n}({\Greekmath 0112} _{n})=\widetilde{l}({\Greekmath 0112} ({\Greekmath 0112} _{n})).$ The reparametrization is such that the elements of the first row and the first column of the Hessian of $\widetilde{l}_{n}({\Greekmath 0112} _{n})$ at ${\Greekmath 0112} _{0,n}=({\Greekmath 011A} _{n},{\Greekmath 011B} _{n}^{2},{\Greekmath 010C} _{n}^{\prime })^{\prime }={\Greekmath 0112} _{\ast }\equiv (1,{\Greekmath 011B} ^{2},0^{\prime })^{\prime }$ are equal to zero. Note that if ${\Greekmath 011A} =1$, then ${\Greekmath 0112} _{0}={\Greekmath 0112} _{0,n}={\Greekmath 0112} _{\ast }$ for some ${\Greekmath 011B} ^{2}$.

We also need to introduce some additional notation. Let $\widehat{{\Greekmath 0112} }= \widehat{{\Greekmath 0112} }_{C}$ and $\widehat{{\Greekmath 0112} }_{n}=\widehat{{\Greekmath 0112} }_{n,C}=$ $(\widehat{{\Greekmath 011A} }_{C},$ $\widehat{{\Greekmath 011B} }_{n,C}^{2},$ $\widehat{ {\Greekmath 010C} }_{C}^{\prime })^{\prime }$ with $\widehat{{\Greekmath 011B} }_{n,C}^{2}=\widehat{ {\Greekmath 011B} }_{C}^{2}/\widehat{{\Greekmath 011A} }_{C}.$ Furthermore, let $Z_{2,N}=N^{1/2}( \widehat{{\Greekmath 011B} }^{2}(1,\widehat{{\Greekmath 010C} }(1))-{\Greekmath 011B} ^{2})$, $Z_{3,N}=N^{1/2}( \widehat{{\Greekmath 010C} }-{\Greekmath 010C} )$ and $Z_{N}=(Z_{1,N},Z_{2,N},Z_{3,N}^{\prime })^{\prime }.$ Then we have the following results:

theoremLet Assumption 1 hold, ${\Greekmath 0122} _{i}\sim N(0,{\Greekmath 011B} ^{2}I),$ $i=1,...,N,$ and ${\Greekmath 011A} =1.$ Then $\vspace{0.08in}$\newline (i) $Z_{N}\overset{d}{\rightarrow }Z=(Z_{1},Z_{2},Z_{3}^{\prime })^{\prime }\sim N(0,\Sigma _{Z}),$ where $E(Z_{1}Z_{2})=0,$ $E(Z_{1}Z_{3})=0,$ $ E(Z_{2}Z_{3})=0,$ \newline $Var(Z_{1})=48T^{-2}((T-1)(T+1))^{-1},$ $Var(Z_{2})=2{\Greekmath 011B} ^{4}(T-1)^{-1}$ and $Var(Z_{3})={\Greekmath 011B} ^{2}(\Sigma _{xqx})^{-1};\vspace{0.08in}$ (ii) letting $K_{+}={\Greekmath 011B} ^{2}(T+1)/6$ and $B^{c}=\mathbf{1}(R^{c}>0)$ with the r.v. $R^{c}$ defined in ((ref))$,\vspace{0.09in}\newline \left[ \begin{array}{c} N^{1/4}(\widehat{{\Greekmath 011A} }_{C}-1) \\ N^{1/2}(\widehat{{\Greekmath 011B} }_{n,C}^{2}-{\Greekmath 011B} ^{2}) \\ N^{1/2}\widehat{{\Greekmath 010C} }_{C} \end{array} \right] \overset{d}{\rightarrow }\left[ \begin{array}{c} (-1)^{B^{c}}Z_{1}^{1/2} \\ Z_{2}+K_{+}Z_{1} \\ Z_{3} \end{array} \right] \mathbf{1}\{Z_{1}>0\}+\left[ \begin{array}{c} 0 \\ Z_{2} \\ Z_{3} \end{array} \right] \mathbf{1}\{Z_{1}\leq 0\}.$

Comments: In the proof of Theorem 2 we show that the sign of $N^{1/4}(\widehat{{\Greekmath 011A} }_{C}-1)$ depends on $\frac{\partial ^{5} \widetilde{l}_{N}^{c}(1)}{\partial r^{5}}$, whereas it follows from Kruiniger (2013) and corollary 1 in Rotnitzky et al. (2000) that the sign of $N^{1/4}(\widehat{{\Greekmath 011A} }_{FEML}-1)$ only depends on the second and third derivatives of the FE\ log-likelihood. The latter is generally true for MLEs of parameters that are only second-order identified, cf. Rotnitzky et al. (2000);

Relaxing the assumption of normality of the ${\Greekmath 0122} _{i}$ affects $ \Sigma _{Z}$ and the conditional distribution of $B^{c}$ given $Z$ but otherwise does not change Theorem 2;

The limiting distribution of $\widehat{{\Greekmath 011A} }_{C}$ is asymmetric unlike that of $\widehat{{\Greekmath 011A} }_{FEML}$ and other MLEs of parameters that are only second-order identified, cf. Rotnitzky et al. (2000);

From $\widehat{{\Greekmath 0112} }_{C}={\Greekmath 0112} (\widehat{{\Greekmath 0112} }_{n,C})$ we have $ \widehat{{\Greekmath 011B} }_{C}^{2}=\widehat{{\Greekmath 011B} }_{n,C}^{2}\widehat{{\Greekmath 011A} }_{C}.$ Hence the rate of convergence of $\widehat{{\Greekmath 011B} }_{C}^{2}$ is also $ N^{1/4} $ and $N^{1/4}(\widehat{{\Greekmath 011B} }_{C}^{2}-{\Greekmath 011B} ^{2})=N^{1/4}( \widehat{{\Greekmath 011A} }_{C}-1){\Greekmath 011B} ^{2}+o_{p}(1)$;

Finally, the following result implies the sign of the asymptotic bias of $ \widehat{{\Greekmath 011A} }_{C}$\ and $\widehat{{\Greekmath 011B} }_{C}^{2}$:

corollaryLet Assumption 1 hold, ${\Greekmath 0122} _{i}\sim N(0,{\Greekmath 011B} ^{2}I),$ $i=1,...,N,$ and ${\Greekmath 011A} =1.$ Then if $T\geq 4,$ $E((-1)^{B^{c}}Z_{1}^{1/2}|Z_{1}>0)>0$ whereas if $T=2$ or $T=3,$ $E((-1)^{B^{c}}Z_{1}^{1/2}|Z_{1}>0)<0.$

We now consider the minimum rate of convergence of $\widehat{{\Greekmath 011A} }=\widehat{ {\Greekmath 011A} }_{F}$ and the limiting distribution of $\widehat{{\Greekmath 0112} }_{F}$ when $ {\Greekmath 011A} =1$. Details of the derivations of these properties of $\widehat{{\Greekmath 011A} } _{F}$ and $\widehat{{\Greekmath 0112} }_{F}$ are given in the appendix. There we show that $N^{1/4}(\widehat{{\Greekmath 011A} }-1)=O_{p}(1),$ cf. Lemma 5.

Let $\Psi _{N,n}({\Greekmath 0112} _{n})=(\frac{\partial \widetilde{l}_{N}^{c}(r)}{ \partial r},s_{n}^{2}r\frac{\partial \widetilde{l}_{n,N}({\Greekmath 0112} _{n})}{ \partial s_{n}^{2}},s_{n}^{2}r\frac{\partial \widetilde{l}_{n,N}({\Greekmath 0112} _{n}) }{\partial b^{\prime }})^{\prime },$ $\widehat{\underline{{\Greekmath 0121} }}_{n}=(( \widehat{{\Greekmath 011B} }_{n,F}^{2}-{\Greekmath 011B} ^{2}),\widehat{{\Greekmath 010C} }_{F}^{\prime })^{\prime }$ and $\underline{w}_{n}=(s_{n}^{2},b^{\prime })^{\prime }$. Then we have the following results:

theoremLet Assumption 1 hold, ${\Greekmath 0122} _{i}\sim N(0,{\Greekmath 011B} ^{2}I),$ $i=1,...,N,$ ${\Greekmath 011A} =1,$ and let $W_{N}$ be a PD matrix. Then $\vspace{0.08in}\vspace{ 0.08in}\newline \left[ \begin{array}{c} N^{1/4}(\widehat{{\Greekmath 011A} }_{F}-1) \\ N^{1/2}\widehat{\underline{{\Greekmath 0121} }}_{n} \end{array} \right] \overset{d}{\rightarrow }\left[ \begin{array}{c} (-1)^{B}Z_{1}^{1/2} \\ \underline{{\Greekmath 0121} }_{+} \end{array} \right] \mathbf{1}\{Z_{1}>0\}+\left[ \begin{array}{c} 0 \\ \underline{{\Greekmath 0121} }_{+}+K_{-}Z_{1} \end{array} \right] \mathbf{1}\{Z_{1}\leq 0\},\vspace{0.06in}\vspace{0.08in}\newline $where $(Z_{1},$ $\underline{{\Greekmath 0121} }_{+}^{\prime })^{\prime }\sim N(0,\Sigma _{{\Greekmath 0121} }),$ $B=\mathbf{1}(R>0)$ and the r.v. $R,$ the matrix $ \Sigma _{{\Greekmath 0121} }$ and the constant vector $K_{-}$ are implicitly defined in the proof.

Comments: In the proof of Theorem 3 we see that the sign of $N^{1/4}(\widehat{{\Greekmath 011A} }_{F}-1)$ depends on $\frac{\partial ^{5} \widetilde{l}_{N}^{c}(1)}{\partial r^{5}}$. The order of that derivative is the same as what Kruiniger (2013) found for Quasi MLEs of second-order identified parameters but different from what Rotnitzky et al. (2000) found for MLEs;

Relaxing the assumption of normality of the ${\Greekmath 0122} _{i}$ affects $ \Sigma _{{\Greekmath 0121} }$ and the conditional distributions of $B$ and $R$ given $ (Z_{1},$ $\underline{{\Greekmath 0121} }_{+}^{\prime })^{\prime }$ but otherwise does not fundamentally change the results in Theorem 3;

Like $\widehat{{\Greekmath 011A} }_{C}$ and $\widehat{{\Greekmath 011B} }_{C}^{2},$ when ${\Greekmath 011A} =1,$ $ \widehat{{\Greekmath 011A} }_{F}$ and $\widehat{{\Greekmath 011B} }_{F}^{2}$ converge at a rate of at least $N^{1/4}$ to ${\Greekmath 011A} $ and ${\Greekmath 011B} ^{2}$, whereas $\widehat{{\Greekmath 010C} } _{F}$ converges at a rate of $N^{1/2}$ to ${\Greekmath 010C} $ just like $\widehat{{\Greekmath 010C} }_{C}$;

For any $W,$ $(\widehat{{\Greekmath 011A} }_{F}-1)^{2}$ is first-order asymptotically equivalent to $(\widehat{{\Greekmath 011A} }_{C}-1)^{2}$ and\ hence the RMSEs of $ \widehat{{\Greekmath 011A} }_{F}$ and $\widehat{{\Greekmath 011A} }_{C}$ are asymptotically the same. However, the limiting distribution of $B$ and hence that of $N^{1/4}( \widehat{{\Greekmath 011A} }_{F}-1)$ depends on $W.$ The limiting distributions of $ \widehat{{\Greekmath 011B} }_{F}^{2}$ and $\widehat{{\Greekmath 010C} }_{F}$ also depend on $W$ and are different from those of $\widehat{{\Greekmath 011B} }_{C}^{2}$ and $\widehat{{\Greekmath 010C} } _{C}$ unless $W_{N}=diag(W_{N,1,1},\underline{W}_{N,2,2})$ where $W_{N,1,1}$ is a scalar. In the latter case $\underline{{\Greekmath 0121} } _{+}+K_{-}Z_{1}=(Z_{2},Z_{3}^{\prime })^{\prime }\ $and $K_{-}=(-K_{+},0)^{ \prime }.$ If in addition $W_{N,1,1}=\infty $ while the elements of $ \underline{W}_{N,2,2}$ are finite, then the limiting distributions of $ N^{1/4}(\widehat{{\Greekmath 011A} }_{F}-1)$ and $N^{1/4}(\widehat{{\Greekmath 011A} }_{C}-1)$ are also the same;

It can be expected that the MMLEs also have non-standard asymptotic properties close to the singularity point, ${\Greekmath 0112} _{\ast }$. Rotnitzky et al. (2000) informally discuss a richness of possibilities for the MLEs close to the singularity point and one can expect several possibilities for the MMLEs too. To save space we don't explore them here. Nonetheless they are a warning of the care needed in conducting inference close to ${\Greekmath 0112} _{\ast }$ . Finally, we note that the local-to-unity asymptotic behaviour of various GMM estimators for the panel AR(1) model discussed in Kruiniger (2009) is unrelated to second-order identification.

Modified likelihood based inference

Wald tests, some versions of (Quasi) LM tests, and (Quasi) LR tests that are used for testing hypotheses involving a parameter that is only second-order identified do not have correct asymptotic size in a uniform sense, cf. Bottai (2003), who explains this in the setting of one-dimensional parametric models. Generalizing the LM-type testing approach in Bottai (2003) that has correct uniform asymptotic size in that setting to a multiple parameter setting, Kruiniger (2025a) has shown that (Quasi) LM tests that are related to the RE- and the FE(Q)MLE for the panel AR(1) model and standardised by using (a sandwich formula involving) the expected rather than the observed average Hessian have correct uniform asymptotic size when $\left\vert {\Greekmath 011A} \right\vert \leq 1$. However, the situation is somewhat different in the case of the Quasi LM (QLM) tests that are based on the (normalized) reparametrized modified profile log-likelihood function $\widetilde{l}_{n,N}({\Greekmath 0112} _{n})$ and used for testing hypotheses that include a restriction on ${\Greekmath 011A} $. Firstly, in this case the singularity point, ${\Greekmath 0112} _{\ast }$, corresponds to a stationary point of inflection of $\widetilde{l}({\Greekmath 0112} )$ and $\widetilde{l}_{n}({\Greekmath 0112} _{n})$ rather than a maximum. Secondly, although $\det (MH)=0$ when ${\Greekmath 0112} _{0}={\Greekmath 0112} _{\ast }$ for any ${\Greekmath 011B} ^{2},$ $\det (MIM)\neq 0$ in this case.$\linebreak $As a result in finite samples $\widetilde{l}_{n,N}({\Greekmath 0112} _{n})$ may not even have a local maximum when ${\Greekmath 011A} $ is close$\linebreak $to one. Nevertheless, the expected average Hessian of $\widetilde{l}_{n,N}({\Greekmath 0112} _{n})$ at $ \underline{\mathcal{{\Greekmath 0112} }}_{n}=\underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n},$ viz. $\overline{H}(\underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n})\equiv E_{ \underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n}}(\frac{\partial ^{2}\widetilde{l} _{n,N}({\Greekmath 0112} _{n})}{\partial {\Greekmath 0112} _{n}\partial {\Greekmath 0112} _{n}^{\prime }}|_{ \underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n}})$, where $\underline{\mathcal{ {\Greekmath 0112} }}_{n}=({\Greekmath 0112} _{n}^{\prime }$ $s_{v,n}^{2})^{\prime }$ with $ s_{v,n}^{2}=s_{v}^{2}/s^{2}-(1-r)$, is negative definite and hence nonsingular for any value of $\underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n}$ that differs from the singularity point $\underline{\mathcal{{\Greekmath 0112} }}_{\ast }=\linebreak ({\Greekmath 0112} _{\ast }^{\prime }$ $0)^{\prime }$.\footnote{ Note that $\overline{H}(\underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n})$ depends on $\underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n}=(\check{{\Greekmath 0112}}_{n}^{\prime }$ $ \check{s}_{v,n}^{2})^{\prime }$, whereas the observed Hessian $\partial ^{2} \widetilde{l}_{N,n}({\Greekmath 0112} _{n})/\partial {\Greekmath 0112} _{n}\partial {\Greekmath 0112} _{n}^{\prime }|_{\check{{\Greekmath 0112}}_{n}}$ only depends on $\check{{\Greekmath 0112}}_{n}$.} \footnote{ This reparametrization is the same as the one used in Kruiniger (2013) for the FE(Q)MLE.} Note that the values of the elements of $\overline{H}( \underline{\mathcal{\breve{{\Greekmath 0112}}}}_{n})$ do not depend on the true distribution of the data.\ We will now introduce the QLM test-statistic $QLM( \widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})$ for testing $H_{0}:$ $A \mathcal{{\Greekmath 0112} }_{0,n}=a$, which includes a restriction on ${\Greekmath 011A} $, (i.e., $ {\Greekmath 011A} =a_{1},$) where $\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n}$ is a restricted estimate of $\underline{\mathcal{{\Greekmath 0112} }}_{0,n}$ such that $A \widetilde{\mathcal{{\Greekmath 0112} }}_{n}=a,$ $A$ is a $J\times \dim (\mathcal{ {\Greekmath 0112} })$ constant matrix of rank $J,$ and $a\ $is a constant vector. Define the average information matrix $\overline{\mathcal{J}}\mathcal{( \mathcal{\breve{{\Greekmath 0112}}}}_{n}\mathcal{)}=N^{-1}\mathop{\textstyle \sum }_{i=1}^{N}\mathcal{J}_{i} \mathcal{(\mathcal{\breve{{\Greekmath 0112}}}}_{n}\mathcal{)}$ with $\mathcal{J}_{i} \mathcal{(\breve{{\Greekmath 0112}}}_{n}\mathcal{)}=\left( \frac{\partial \widetilde{l} _{n,i}({\Greekmath 0112} _{n})}{\partial {\Greekmath 0112} _{n}}|_{\mathcal{\breve{{\Greekmath 0112}}} _{n}}\right) \left( \frac{\partial \widetilde{l}_{n,i}({\Greekmath 0112} _{n})}{ \partial {\Greekmath 0112} _{n}^{\prime }}|_{\mathcal{\breve{{\Greekmath 0112}}}_{n}}\right) ,$ where $\widetilde{l}_{n,i}({\Greekmath 0112} _{n})$ is the contribution to the modified profile log-likelihood function $N\times \widetilde{l}_{n,N}({\Greekmath 0112} _{n})$ by individual $i$. If $\widetilde{\mathcal{{\Greekmath 0112} }}_{n}\neq \mathcal{{\Greekmath 0112} }_{\ast }$ for all ${\Greekmath 011B} ^{2}>0$ or $J=1,$ then $QLM(\widetilde{\underline{ \mathcal{{\Greekmath 0112} }}}_{n})$ is given by:

eqnarray[eqnarray omitted — 755 chars of source]

The parameter ${\Greekmath 011B} _{v,n}^{2}$ can be estimated by the restricted FE(Q)MLE, cf. Kruiniger (2025a).

If $H_{0}$ is true and $\widetilde{\mathcal{{\Greekmath 0112} }}_{n}\neq \mathcal{ {\Greekmath 0112} }_{\ast }$ (for all ${\Greekmath 011B} ^{2}>0$) or $J=1$, then $QLM(\widetilde{ \underline{\mathcal{{\Greekmath 0112} }}}_{n})\sim {\Greekmath 011F} ^{2}(J).$ To test $H_{0}:$ $ {\Greekmath 011A} =a$ for some known value of $a\in (-1,1]$, one can use $QLM(\widetilde{ \underline{\mathcal{{\Greekmath 0112} }}}_{n})$ in ((ref)) with $A=(1$ $\mathbf{0} ^{\prime })$ and $\frac{\partial \widetilde{l}_{n,N}(\widetilde{\mathcal{ {\Greekmath 0112} }}_{n})}{\partial {\Greekmath 0112} _{n}}=A^{\prime }\frac{\partial \widetilde{l} _{n,N}(\widetilde{\mathcal{{\Greekmath 0112} }}_{n})}{\partial r}$. Note that the value of $QLM(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})$ remains the same when $\overline{H}^{-1}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})$ is replaced by $adj(\overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}} _{n})) $. Furthermore, $\overline{\mathcal{J}}(\widetilde{\mathcal{{\Greekmath 0112} }} _{n})$ and p$\lim_{N\rightarrow \infty }\overline{\mathcal{J}}(\widetilde{ \mathcal{{\Greekmath 0112} }}_{n})$ are positive definite.

If $\widetilde{\mathcal{{\Greekmath 0112} }}_{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ for some $ {\Greekmath 011B} ^{2}>0,$ $rk(\overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}} _{n}))=\dim (\overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n}))-1$ and hence $rk(adj(\overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}} _{n})))=1.$ Furthermore, $\overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})_{i,j}=0$ iff $i=1$ and/or $j=1,$ $adj(\overline{H}(\widetilde{ \underline{\mathcal{{\Greekmath 0112} }}}_{n}))_{i,j}\neq 0$ iff $i=j=1,$ and $adj( \overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n}))(1$ $\mathbf{0} ^{\prime })^{\prime }\propto (1$ $\mathbf{0}^{\prime })^{\prime }.$ Thus, if $\widetilde{\mathcal{{\Greekmath 0112} }}_{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ (for some $ {\Greekmath 011B} ^{2}>0$), the QLM test statistic in ((ref)) for testing $H_{0}:$ $ {\Greekmath 011A} =a$ equals $N\times (\frac{\partial \widetilde{l}_{n,N}(\mathcal{{\Greekmath 0112} }_{n})}{\partial r}|_{{\Greekmath 0112} _{\ast }})^{2}/(N^{-1}\mathop{\textstyle \sum }_{i=1}^{N}(\frac{ \partial \widetilde{l}_{n,i}(\mathcal{{\Greekmath 0112} }_{n})}{\partial r}|_{{\Greekmath 0112} _{\ast }})^{2}).$ As p$\lim_{N\rightarrow \infty }N^{-1}\mathop{\textstyle \sum }_{i=1}^{N}(\frac{\partial \widetilde{l}_{n,i}(\mathcal{{\Greekmath 0112} } _{n})}{\partial r}|_{{\Greekmath 0112} _{\ast }})^{2}>0,$ it follows that one can still use ((ref)) for testing $H_{0}:$ ${\Greekmath 011A} =a$ when $\widetilde{\mathcal{ {\Greekmath 0112} }}_{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ (for some ${\Greekmath 011B} ^{2}>0$). However, if $\widetilde{\mathcal{{\Greekmath 0112} }}_{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ (for some ${\Greekmath 011B} ^{2}>0$) and $J\geq 2,$ then p$\lim_{N\rightarrow \infty }\det (A\,adj(\overline{H}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})) \overline{\mathcal{J}}(\widetilde{\mathcal{{\Greekmath 0112} }}_{n})adj(\overline{H}( \widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n}))A^{\prime })=0\mathbf{.}$ To test $H_{0}:$ $A\mathcal{{\Greekmath 0112} }_{0,n}=a$ when $\widetilde{\mathcal{{\Greekmath 0112} } }_{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ (for some ${\Greekmath 011B} ^{2}>0$) and $J\geq 2,$ one can use the following QLM test statistic, cf. Bottai (2003) and Kruiniger (2025a):

eqnarray[eqnarray omitted — 759 chars of source]

$\vspace{-0.08in}$with$\vspace{-0.08in}$

eqnarray*[eqnarray* omitted — 2,042 chars of source]

where we have partitioned ${\Greekmath 0112} _{n}$ as ${\Greekmath 0112} _{n}=(r_{n},d_{n}^{\prime })^{\prime }$ and used $\widetilde{l}_{n,N}$ and $\widetilde{l}_{n,i}$ as short for $\widetilde{l}_{n,N}({\Greekmath 0112} _{n})$ and $\widetilde{l}_{n,i}({\Greekmath 0112} _{n})$. Note that $\frac{\partial \widetilde{l}_{n,i}(\widetilde{\mathcal{ {\Greekmath 0112} }}_{n})}{\partial r}$ and $\overline{H}(\widetilde{\underline{ \mathcal{{\Greekmath 0112} }}}_{n})$ in ((ref)) have been replaced by $S_{i,1}$ and $\widetilde{\mathcal{H}}(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})$ in ( (ref)). Furthermore, to derive the QLM test-statistic $QLM(\widetilde{ \underline{\mathcal{{\Greekmath 0112} }}}_{n})$ given in ((ref)), the hypothesis $ {\Greekmath 011A} =1$ had to be reformulated as $({\Greekmath 011A} -1)^{2}=0$. As a result $ \widetilde{A}_{1,1}$ is equal to $\frac{1}{2}\frac{\partial ^{2}({\Greekmath 011A} -1)^{2} }{\partial {\Greekmath 011A} ^{2}}=1$. If $H_{0}$ is true, $\widetilde{\mathcal{{\Greekmath 0112} }} _{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ (for some ${\Greekmath 011B} ^{2}>0$) and $J\geq 2$, then $QLM(\widetilde{\underline{\mathcal{{\Greekmath 0112} }}}_{n})\sim {\Greekmath 011F} ^{2}(J).$

theoremUnder regularity conditions (A1)-(A7) given in the appendix, the Quasi LM test for $H_{0}:$ $A\mathcal{{\Greekmath 0112} }_{0,n}=a,$ where $A_{1,.}=(1$ $\mathbf{0 }^{\prime }),$ that is based on ((ref)) if $\widetilde{\mathcal{{\Greekmath 0112} }} _{n}\neq \mathcal{{\Greekmath 0112} }_{\ast }$ for all ${\Greekmath 011B} ^{2}>0,$ and on ((ref)) if $\widetilde{\mathcal{{\Greekmath 0112} }}_{n}=\mathcal{{\Greekmath 0112} }_{\ast }$ for some ${\Greekmath 011B} ^{2}>0,$ has correct asymptotic size in a uniform sense.

Confidence sets (CSs) that are obtained by inverting the tests based on ((ref)) and ((ref)) have correct asymptotic size in a uniform sense. Other tests (and CSs) for ${\Greekmath 011A} $ that have correct asymptotic size include (CSs based on) the GMM LM test(-statistic)s of Newey and West (1987) that exploit the moments conditions of the System GMM and the nonlinear Ahn-Schmidt (AS) GMM estimator, respectively, see Kruiniger (2009) for the System version and Bun and Kleibergen (2022) for the AS version of the test, and identication-robust test(-statistics)s such as the GMM AR test of Stock and Wright (2000) and the KLM and GMM-CLR tests of Kleibergen (2005) that exploit System and AS moments conditions, cf. Bun and Kleibergen (2022). Kruiniger (2025a) has shown that if the data are i.i.d. and normal, then the QLM test for testing an hypothesis about ${\Greekmath 011A} $ that is based on the FEMLE and uses the expected Hessian\ shares the optimal power properties of the KLM\ test in a worst case scenario. To test $H_{0}:$ ${\Greekmath 011A} =1$ one could also use a Wald test based on $\sqrt{N}(\widehat{{\Greekmath 011A} } _{C}-1)^{2}. $ Under $H_{0}\ \sqrt{N}(\widehat{{\Greekmath 011A} }_{C}-1)^{2}\overset{d}{ \rightarrow }Z_{1}\mathbf{1}\{Z_{1}>0\},$ cf. Theorem 2. Recall that $ Z_{1,N}=\left( -\frac{1}{2}\frac{\partial ^{3}\widetilde{l}_{N}^{c}(1)}{ \partial r^{3}}\right) ^{-1}N^{1/2}\left( \frac{\partial \widetilde{l} _{N}^{c}(1)}{\partial r}\right) \overset{d}{\rightarrow }Z_{1},$ with p$ \lim\nolimits_{N\rightarrow \infty }\frac{\partial ^{3}\widetilde{l} _{N}^{c}(1)}{\partial r^{3}}=\frac{\partial ^{3}\widetilde{l}^{c}(1)}{ \partial r^{3}}=\frac{T(T-1)(T+1)}{12}$ and $\frac{\partial \widetilde{l} _{N}^{c}(1)}{\partial r}$ given in ((ref)). When the data are heterogeneous and/or non-normal, one can bootstrap the distribution of $ N^{1/2}\left( \frac{\partial \widetilde{l}_{N}^{c}(1)}{\partial r}\right) $ or estimate the averages of the second and the fourth moments of the $ {\Greekmath 0122} _{i,t}$ by using that under $H_{0}$\ ${\Greekmath 0122} _{i}=y_{i}-y_{i,-1}$ for $i=1,...,N.$ To test $H_{0}:$ ${\Greekmath 011A} =1$ one could also use any other panel unit root test, e.g. the test of Harris and Tzavalis (1999) that is based on the bias-corrected LSDV estimator for ${\Greekmath 011A} ,$ i.e., $\widehat{{\Greekmath 011A} }_{ML}+\frac{3}{T+1}$, where $-\frac{3}{T+1}$ is the asymptotic bias of $\widehat{{\Greekmath 011A} }_{ML}$ when ${\Greekmath 011A} =1.$ The rate of convergence of $\widehat{{\Greekmath 011A} }_{ML}$ is $N^{1/2}$ which is faster than $ N^{1/4},$ the rate of $\widehat{{\Greekmath 011A} }_{C}.$ Hence if $N$ is large enough inference based on $\widehat{{\Greekmath 011A} }_{ML}$ is better in terms of power and size. Finally, to test a hypothesis that only involves ${\Greekmath 010C} $, one can use a Wald test based on $\widehat{{\Greekmath 010C} }_{C}$.

The finite sample performance of the Modified ML estimators and the Quasi LM test

In this section we compare through Monte Carlo simulations the finite sample properties of three estimators in various panel AR(1) models without covariates: $\widehat{{\Greekmath 011A} }_{C}$; the REMLE for ${\Greekmath 011A} $ that has been proposed by both Chamberlain (1980) and Anderson and Hsiao (1982), henceforth $\widehat{{\Greekmath 011A} }_{REML}$; and the FEMLE for ${\Greekmath 011A} $ (i.e., $ \widehat{{\Greekmath 011A} }_{FEML}$) that has been proposed by Hsiao et al. (2002). We study how the properties of these estimators are affected if we change (1) the distributions of the $v_{i}=y_{i,0}-{\Greekmath 0116} _{i}$ or (2) the ratio of the variances of the error components, i.e. ${\Greekmath 011B} _{{\Greekmath 0116} }^{2}/{\Greekmath 011B} ^{2}$. We conducted the simulation experiments for $(T,N)=(4,100),$ $(9,100),$ $ (4,500) $ or $(9,500)$ and ${\Greekmath 011A} =0.5,$ $0.8,$ $0.9,$ $0.95,$ $0.98$ or $1.$

In all simulation experiments the error components have been drawn from normal distributions with zero means. We assumed that ${\Greekmath 011B} _{{\Greekmath 0116} }^{2}=0,$ $1$ or $25.$ For the ${\Greekmath 0122} _{i,t}$ we assumed homoskedasticity and no autocorrelation: $E({\Greekmath 0122} _{i}{\Greekmath 0122} _{i}^{\prime })={\Greekmath 011B} ^{2}I$ with ${\Greekmath 011B} ^{2}=1.$

In order to assess how the assumptions with respect to $y_{i,0}-{\Greekmath 0116} _{i}$, $ i=1,...,N,$ affect the properties of the estimators, we conducted two different sets of experiments, which are identified by a capital: in one set, labeled NS, the initial observations are non-stationary, i.e., $ y_{i,0}-{\Greekmath 0116} _{i}=0$, $i=1,...,N,$ whereas in the other set, labeled S, the initial observations are drawn from stationary distributions when $ \left\vert {\Greekmath 011A} \right\vert <1$, i.e., $(y_{i,0}-{\Greekmath 0116} _{i})\sim N(0,{\Greekmath 011B} _{i,0}^{2}/(1-{\Greekmath 011A} ^{2}))$ with ${\Greekmath 011B} _{i,0}^{2}={\Greekmath 011B} ^{2},$ although $ y_{i,0}-{\Greekmath 0116} _{i}=0$, $i=1,...,N,$ when ${\Greekmath 011A} =1$.

Note that all four estimators suffer from a weak moment conditions problem when ${\Greekmath 011A} $ is close to one, cf. Kruiniger (2013).

In the cases of the RE- and FEMLE $(1-{\Greekmath 011A} ){\Greekmath 0116} _{i}+{\Greekmath 0122} _{i}$ is decomposed as $(1-{\Greekmath 011A} ){\Greekmath 0119} y_{i,0}-(1-{\Greekmath 011A} )v_{i}+{\Greekmath 0122} _{i}=(1-{\Greekmath 011A} ){\Greekmath 0119} y_{i,0}+u_{i}$ with ${\Greekmath 0119} =1$ for the FE case. In the experiments we imposed homoskedasticity on their likelihood functions and added the restrictions ${\Greekmath 011B} ^{2}>0$ and $(T-1)(1-{\Greekmath 011A} )^{2}{\Greekmath 011B} _{v}^{2}+{\Greekmath 011B} ^{2}>0$ to ensure that the estimates of $E(u_{i}u_{i}^{\prime })$ were PD.

We allowed for time effects by subtracting cross-sectional averages from the data.

To speed up the computations, we computed $\widehat{{\Greekmath 011A} }_{C}$ by maximizing $\widetilde{l}_{N}({\Greekmath 0112} )$ subject to\thinspace $-1\leq r\leq 1.4$ rather than $-1\leq r<\infty .$ (We also tried using $-1\leq r\leq 2,$ but never found an internal local maximum between $1.4$ and $2$.) If no internal local maximum was found, we computed $\widehat{{\Greekmath 011A} }_{C}$ by solving ((ref)) s.t.\thinspace $-1\leq r\leq 1.4$ using grid search.

Tables 1-6 report the simulation results in terms of the biases and root mean squared errors (RMSEs) of the estimators and the relative frequencies that $\widehat{{\Greekmath 011A} }_{LAN}$ did not exist (NM). The tables differ with respect to the dimensions of the panel and the assumptions made about the $ y_{i,0}-{\Greekmath 0116} _{i}$, $i=1,...,N$. Inspection of the\thinspace results leads\thinspace to\thinspace the\thinspace following conclusions:\thinspace \footnote{ Dhaene and Jochmans (2016) report simulations results on the finite sample properties of their Adjusted Likelihood estimator ($\widehat{{\Greekmath 011A} }_{ADJ}$), the bias corrected LSDV estimator of Hahn and Kuersteiner (2002) ($\widehat{ {\Greekmath 011A} }_{HK}$) and the 1-step\ GMM estimator of Arellano and Bond (1991) ($ \widehat{{\Greekmath 011A} }_{AB}$). Some of their simulation experiments are equal to some of our experiments. The results for these experiments show that $ \widehat{{\Greekmath 011A} }_{ADJ}$ and $\widehat{{\Greekmath 011A} }_{C}$ are very similar and that $ \widehat{{\Greekmath 011A} }_{HK}$ has a large bias when $T$ is small. $\widehat{{\Greekmath 011A} } _{AB}$ has poor properties when ${\Greekmath 011A} $ is close to 1 due to weak instruments.}

enumerate• In almost all experiments (the exception is design NS with $N=100$ and ${\Greekmath 011A} =.0.5$) $\widehat{{\Greekmath 011A} }_{REML}$ is superior in terms of RMSE for `smaller' values of ${\Greekmath 011A} $ (i.e., values closer to 0), $\widehat{{\Greekmath 011A} } _{FEML}$ is superior for `larger' values of ${\Greekmath 011A} $ (i.e., values closer to 1), while $\widehat{{\Greekmath 011A} }_{C}$ is superior on an interval of `intermediate' values of ${\Greekmath 011A} $, which includes ${\Greekmath 011A} =0.8$ when $T=4$ and $N=100,$ and $ {\Greekmath 011A} =0.9$ when $T=4$ and $N=500.$ In most experiments $\widehat{{\Greekmath 011A} } _{REML}$ is superior when ${\Greekmath 011A} =0.5,$ while $\widehat{{\Greekmath 011A} }_{FEML}$ is superior when ${\Greekmath 011A} $ is near/equals $1.$\ When ${\Greekmath 011A} $ is near $ 1, $ the bias of $\widehat{{\Greekmath 011A} }_{C}$ is larger than the biases of $ \widehat{{\Greekmath 011A} }_{FEML}$ and $\widehat{{\Greekmath 011A} }_{REML}$. • When $T$ or $N$ increases, the values of the bounds of the interval for ${\Greekmath 011A} $ on which $\widehat{{\Greekmath 011A} }_{C}$ is superior increase. When $T=9$ and $N=500,$ $\widehat{{\Greekmath 011A} }_{C}$ is superior around ${\Greekmath 011A} =0.95.$ Furthermore, when ${\Greekmath 011A} =0.50$ and $T=9$ or $N=500,$ $\widehat{{\Greekmath 011A} }_{FEML}$ is often the most efficient estimator after $\widehat{{\Greekmath 011A} }_{REML}$. • When ${\Greekmath 011B} _{{\Greekmath 0116} }^{2}/{\Greekmath 011B} ^{2}$ increases, the RMSE of $\widehat{ {\Greekmath 011A} }_{REML}$ increases and hence the value of the lowerbound of the interval of values of ${\Greekmath 011A} $ on which $\widehat{{\Greekmath 011A} }_{C}$ is superior decreases. • When $Var(y_{i,0}-{\Greekmath 0116} _{i})/{\Greekmath 011B} ^{2}$ decreases, the bias and the RMSE of $\widehat{{\Greekmath 011A} }_{C}$ and the RMSE of $\widehat{{\Greekmath 011A} }_{REML}$ increase and the value of the upperbound of the interval of values of ${\Greekmath 011A} $ on which $\widehat{{\Greekmath 011A} }_{C}$ is superior decreases. • When $T=4$ and $N=100,$ $NM>0.35$ for ${\Greekmath 011A} \geq 0.8;$ when $T=4$ and $ N=500,$ $NM>0.29$ for ${\Greekmath 011A} \geq 0.8;$ when $T=9$ and $N=100,$ $NM>0.35$ for ${\Greekmath 011A} \geq 0.9;$ and when $T=9$ and $N=500,$ $NM>0.25$ for ${\Greekmath 011A} \geq 0.9.$ Generally, the higher the value of ${\Greekmath 011A} ,$ the higher the value of $NM$. When ${\Greekmath 011A} =1,$ $NM\approx 0.50$ for all panels considered, which supports the idea that even asymptotically $\widehat{{\Greekmath 011A} }_{LAN}$ may not exist when ${\Greekmath 011A} =1$. If the value of $Var(y_{i,0}-{\Greekmath 0116} _{i})/{\Greekmath 011B} ^{2}$ decreases, the value of $NM$ increases. Under design NS, when $T=4$, $N=100$ and ${\Greekmath 011A} =0.5,$ we still have $NM>0.3.$

We have also investigated the size and power properties of the modified likelihood based QLM-test for testing $H_{0}:$ ${\Greekmath 011A} =a$, that is, $QLM( \mathcal{{\Greekmath 011A} }).$ To this end, we conducted three types of Monte Carlo experiments. The designs of two of them, labelled S-Normal and NS-Normal, were similar to designs S and NS described above. The designs of the third kind of experiments, labelled S-ChiSq., were also similar to S with one difference: the ${\Greekmath 0122} _{i,t}$ were i.i.d. $({\Greekmath 011F} ^{2}(1)-1)/\sqrt{2}$ instead of i.i.d. $N(0,1)$ so that $(y_{i,0}-{\Greekmath 0116} _{i})\sim ({\Greekmath 011F} ^{2}(1)-1)/ \sqrt{2(1-{\Greekmath 011A} ^{2})}$ instead of $N(0,1/(1-{\Greekmath 011A} ^{2})).$ In all experiments ${\Greekmath 0116} _{i}\sim N(0,1).$ We used various true values for ${\Greekmath 011A} $ including $ 0.5,$ $0.9,$ $0.95$ and $0.99$. The results for the power of $QLM(\mathcal{ {\Greekmath 011A} })$ were based on testing $H_{0}:$ ${\Greekmath 011A} =0.8$. In all experiments $T=9$ and $N\in \{100,500\}$.

$QLM(\mathcal{{\Greekmath 011A} })$ depends on $\mathcal{H}(\widetilde{\underline{ \mathcal{{\Greekmath 0112} }}}_{n})$, i.e., an estimate of the expected Hessian that is based on the restricted estimate $\widetilde{\underline{\mathcal{{\Greekmath 0112} }}} _{n}$. One of the parameters in $\mathcal{H}(\underline{\mathcal{{\Greekmath 0112} }} _{0,n})$ is ${\Greekmath 011B} _{v,n}^{2}$. However, the latter is not estimated by a MMLE. Instead we used the restricted FE(Q)MLE for ${\Greekmath 011B} _{v,n}^{2}$.

Tables 7 and 8 report the simulation results for the size and the power of $ QLM(\mathcal{{\Greekmath 011A} })$, respectively.\footnote{ Dhaene and Jochmans (2016) report simulations results on the finite sample properties of confidence intervals for ${\Greekmath 011A} $ based on $\widehat{{\Greekmath 011A} } _{ADJ}$, $\widehat{{\Greekmath 011A} }_{HK}$ and $\widehat{{\Greekmath 011A} }_{AB},$ respectively, and their first-order asymptotic standard errors. Their results show that none of these intervals have correct size when ${\Greekmath 011A} $ is close to one, with the size distortions being particularly large for the confidence intervals based on $\widehat{{\Greekmath 011A} }_{HK}$ and $\widehat{{\Greekmath 011A} }_{AB}$.} Table 7 shows that the empirical size of the test is very close to the nominal size of 5% in all experiments, including those where ${\Greekmath 011A} $ is close to one. Finally, table 8 shows that the power properties of $QLM(\mathcal{{\Greekmath 011A} })$ do not change much across the three types of experiments and also that its power is still high when (true) ${\Greekmath 011A} =0.99$ despite weak identification in that case.

Concluding remarks

Alvarez and Arellano (2022) and Juodis (2013) have extended the MMLE of Lancaster to panel AR(1) models that allow for time-series heteroskedasticity. Their estimators suffer from the same problems as Lancaster's MMLE, namely a weak moment conditions problem if the parameter values are close to the unit root and time-series homoskedasticity, cf. Alvarez and Arellano (2022) and Kruiniger (2013); the related problem of possible non-existence; and the possibility of non-uniqueness of local maxima of the modified profile likelihood function. The non-existence problem can be solved by generalizing their estimators in a similar way as Lancaster's estimator has been generalized to ((ref)) or ((ref)). However, it is unclear whether the modified profile likelihood function has at most one local maximum even when the parameter space for ${\Greekmath 011A} $ is restricted to $[-1,1]$.\footnote{ The modified profile likelihood equation for ${\Greekmath 011A} $ is a polynomial in $r$. If the model has no covariates, then the coefficients of this polynomial are functions of ${\Greekmath 011A} ,$ ${\Greekmath 011B} _{v}^{2}$ and $T$ variance parameters instead of one.} If uniqueness would not hold, then one could select a local maximum that is (plausible and) closest to the value of $\widehat{{\Greekmath 011A} }_{FEML}$ (or $\widehat{{\Greekmath 011A} }_{REML}$), which is a consistent estimator, as the MMLE. \footnote{ Note that this method of selecting the MMLE is also \textquotedblleft sensible\textquotedblright\ in finite samples.}

Alvarez and Arellano (2022) and Dhaene and Jochmans (2016) have also extended the MMLE of Lancaster to panel AR(p) models, while Juodis (2013) has also extended the MMLE of Lancaster to panel VARX(1) models. Comments similar to those made in the previous paragraph apply to these extensions. The MMLEs discussed in section 3 are inconsistent for models with endogenous or predetermined covariates. However, in some cases these models can be replaced by VAR models.

It seems reasonable to expect that the aforementioned extensions of the MMLEs to more general models may also outperform the RE- and FEMLEs for those models in panels of realistic dimensions for some parts of the parameter space.\ However, a comprehensive Monte Carlo study of their finite sample properties is left for future research.

Finally, we note that Bester and Hansen (2007) and Arellano and Bonhomme (2009) have proposed priors that result in first-order unbiased Bayesian estimators for ${\Greekmath 011A} $ in a version of model ((ref)) that does not include the $K$ exogenous covariates.