Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Individual and Time Effects in Nonlinear Panel Models with Large $N$, $T$
\abstract{
center[center omitted — 657 chars of source]
}
{
{ Keywords:} Panel data, nonlinear model, dynamic model, asymptotic bias correction, fixed effects, time effects.\\
{ JEL:} C13, C23.
}
Introduction
Fixed effects estimators of nonlinear panel data models can be severely biased
because of the incidental parameter problem Neyman:1948p881.
A growing literature, surveyed in Arellano and Hahn ArellanoHahn2007,
shows that the leading term of an asymptotic expansion of the bias as both the cross-sectional dimension $N$ and time series dimension $T$ of the panel
grow, can be characterized and corrected for. In models with individual effects, the leading bias term is of order $1/T$ and comes from the estimation of the individual effects.
This result, however, does not apply to models with individual and time effects, where both of these effects are treated as parameters to be estimated. In this paper we show that the estimation of the time effects causes an additional incidental parameter bias of order $1/N$. Thus, if $N$ and $T$ are similarly large, the bias produced by the estimation of the time effects is of similar order of magnitude to the bias produced by the estimation of the individual effects, and both biases need to be corrected. We provide the corresponding analytical and jackknife bias corrections.
The asymptotic approximation to the fixed effects estimators that lets the two dimensions of the panel grow with the sample size is motivated by the recent availability of long panels and other large pseudo-panel data structures where the indexes might not correspond to individuals and time periods. Examples of these datasets include traditional microeconomic panel surveys with a long history of data such as the PSID and NLSY, international cross-country panels such as the Penn World Table, U.S. state level panels over time such as the CPS, and square pseudo-panels of trade flows across countries such as the Feenstra's World Trade Flows and CEPII, where the indexes correspond to the same countries indexed as importers and exporters.
We focus on semi-parametric models with log-likelihood functions that are concave in all parameters, and where each individual effect $\alpha_i$ and time effect $\gamma_t$ enter the log-likelihood for observation $(i,t)$ additively as $\alpha_i + \gamma_t$.
This is the most common specification for the individual and time effects in linear models and is also a natural specification in the nonlinear models that we consider. Imposing concavity of the log-likelihood function
greatly facilitates showing consistency in our setting where the dimension of the parameter space
grows with the sample size. The
most popular limited dependent variable models, including logit,
probit, ordered probit, Tobit and Poisson models have concave log-likelihood functions, possibly after reparametrization (Olsen Olsen:1978p3375, and Pratt Pratt:1981p654). We note here that the general expansion that we derive in Appendix B do not impose additivity and concavity, but we use these restrictions to apply the expansion to fixed effects estimators.
The models that we consider are semi-parametric because the joint distribution of the explanatory variables and the unobserved effects is left unspecified. The explanatory variables can be either strictly exogenous or predetermined.
We derive bias expansions and corrections for fixed effects estimators of common parameters $\beta$ and average partial effects (APEs). The vector $\beta$ includes all the unknown parameters that enter the log-likelihood function other than the individual and time effects, such as index coefficients in a probit model. The APEs are functions of the data, the common parameters, and the individual and time effects in nonlinear models. We find that the properties of the fixed effects estimators of $\beta$ and the APEs are different. For $\beta$, the order of the bias is $1/T + 1/N$, which is of the same as the rate of convergence $1/\sqrt{NT}$ under sequences where $N/T$ converge to a constant. For the APEs, we uncover that the incidental parameter problem is negligible asymptotically
because the order of the bias, $1/N + 1/T$, is smaller than the rate of convergence, which is $1/\sqrt{N} + 1/\sqrt{T}$, slower than for model parameters. To the best of our knowledge, this rate result is new for fixed effects estimators of average partial effects in nonlinear panel models with individual and time effects.\footnote{Galvao and Kato GalvaoKato:2013 also found slow rates of convergence for fixed effects estimators in linear models with individual effects under misspecification. Fernandez-Val and Lee FL13 pointed out this issue in nonlinear models with only individual effects.} In numerical examples we find that the bias corrections, while not necessary to center the asymptotic distribution of APE estimators, do improve their finite-sample properties, specially in dynamic models.
The bias correction eliminates the bias terms of orders $1/T$ and $1/N$ from the fixed effects estimators. We considerer two methods to implement the correction: an analytical bias correction similar to Hahn
and Newey Hahn:2004p882 and Hahn and Kuersteiner HahnKuersteiner2011, and a suitable modification of the split panel jackknife of Dhaene and Jochmans DhaeneJochmans2015.\footnote{A similar split panel jackknife bias correction method was outlined in Hu Hu2002.}
However, the theory of the previous papers does not cover the models that we consider, because, in addition to not allowing for
time effects, it
assumes either identical distribution or stationarity over time for the processes of the observed variables,
conditional on the unobserved effects. These assumptions are violated
in our models due to the presence of the time effects,
so we need to adjust the asymptotic theory accordingly.
The individual and time effects introduce strong correlation in both dimensions of the panel.
Conditional on the unobserved effects, we impose cross-sectional independence and weak time-serial dependence, and we allow for heterogeneity in both dimensions.
Simulation evidence indicates that our corrections improve the estimation and inference performance of the fixed effects estimators of parameters and average effects. The analytical corrections dominate the jackknife corrections in a probit model for sample sizes that are relevant for empirical practice.
In the online supplement, Fern{\'a}ndez-Val and Weidner Supp2015, we illustrate the corrections with an empirical application on the relationship between competition and innovation using a panel of U.K. industries, following Aghion, Bloom, Blundell, Griffith and Howitt AghionBloomBlundellGriffithHowitt2005. We find that the inverted-U pattern relationship found by Aghion et al is robust to relaxing the strict exogeneity assumption of competition with respect to the innovation process and to the inclusion of innovation dynamics. We also uncover substantial state dependence in the innovation process.
\paragraph{Literature review.} The Neyman and Scott incidental parameter problem has been extensively discussed in the econometric literature; see, for example, Heckman Heckman:1981p2940,
Lancaster Lancaster:2000p879, and Greene Greene:2004p3125. There is also a vast literature that shows how to tackle the problem in specific models under asymptotic sequences where $T$ is fixed and $N$ grows to infinity. However, there are results,
e.g. from Honor{\'e} and Tamer HonoreTamer2006,
Chamberlain Chamberlain2010,
and Chernozhukov, Fern{\'a}ndez-Val, Hahn and Newey CFHN13, showing
that model parameters and APEs are not point identified in important nonlinear panel data models under fixed-$T$ asymptotic sequences, implying that no fixed-$T$ consistent point estimators exist in these models.
A recent response to the incidental parameter problem is to adopt an alternative asymptotic approximation where both $N$ and $T$ grow with the sample size. Under these large-$T$ sequences, the fixed effects estimator is consistent but has bias in the asymptotic distribution. This asymptotic bias is the large-$T$ version of the incidental parameter problem and has motivated the development of bias corrections. Examples of papers that use this approximation include Phillips and Moon Phillips:1999p733, Hahn and Kuersteiner Hahn:2002p717,
Lancaster Lancaster:2002p875, Woutersen Woutersen:2002p3683,
Alvarez and Arellano AlvarezArellano2003,
Hahn and Newey Hahn:2004p882,
Carro Carro:2007p3601,
Arellano and Bonhomme ArellanoBonhomme2009,
Fernandez-Val FernandezVal:2009p3313,
Hahn and Kuersteiner HahnKuersteiner2011,
Fernandez-Val and Vella FernandezValVella2011,
and Kato, Galvao and Montes-Rojas KatoGalvaoMontes-Rojas2012.
This previous work, however, does not cover models with time effects.\footnote{An
exception is
Woutersen Woutersen:2002p3683, which considers a special type of grouped time effects whose number is fixed with $T$. We instead consider an unrestricted set of $T$ time effects, one for each time period.} Our contribution to this literature is to extend the large-$T$ bias corrections to models with two-way unobserved effects such as the individual and time effects commonly included in linear models.
The large-$T$ panel literature on models with both individual and time effects is
sparse. Pesaran Pesaran2006, Bai Bai:2009p3321, and Moon and Weidner MoonWeidner2015a,MoonWeidner2015b study linear regression models with interactive individual and time fixed effects. The fixed effects estimators in these models also have asymptotic bias of order $1/T + 1/N$, but the methods used to derive this bias rely on linearity and therefore cannot be applied to the nonlinear models that we consider. Hahn and Moon HahnMoon2006 consider bias corrected fixed effects estimators in panel linear autoregressive models with additive individual and
time effects. Regarding non-linear models, there is
independent and contemporaneous work by Charbonneau Charbonneau2011,Charbonneau2014, which extends the conditional fixed effects estimators to logit and Poisson models
with individual and time effects. She differences out the individual and time effects
by conditioning on sufficient statistics. The conditional approach
completely eliminates the asymptotic bias coming from the estimation of the incidental parameters,
but it does not permit estimation of average partial effects and has not been developed for models with predetermined regressors. We instead consider
estimators of model parameters and average partial effects in nonlinear models with predetermined
regressors. The two approaches can therefore be considered as complementary.
\paragraph{Outline of the paper.} The rest of the paper is organized as follows. Section (ref) introduces the model and fixed effects estimators. Section (ref) describes the bias corrections to deal with the incidental parameters problem and illustrates how the bias corrections work through an example. Section (ref) provides
the asymptotic theory.
Section (ref) presents Monte Carlo results. The Appendix collects the proofs of the main results, and an online supplement to the paper contains additional technical derivations, numerical examples, and an empirical application Supp2015.
Model and Estimators
Model
The data consist of $N \times T$ observations $\{(Y_{it}, X'_{it})': 1 \leq i \leq N, 1 \leq t \leq T \},$ for a scalar outcome variable of interest $Y_{it}$ and a vector
of explanatory variables $X_{it}$. We assume that the outcome for
individual $i$ at time $t$ is generated by the sequential process:
equation*[equation* omitted — 151 chars of source]
where $X^t_i = (X_{i1}, \ldots, X_{it}),$ $\alpha = (\alpha_1, \ldots, \alpha_N)$, $\gamma = (\gamma_1, \ldots, \gamma_T)$,
$f_{Y}$ is a known probability function, and $\beta$ is a finite dimensional parameter vector.
The variables $\alpha_i$ and $\gamma_t$ are unobserved individual and time effects that in economic applications capture individual heterogeneity and aggregate shocks, respectively. The model is semiparametric because we do not specify the distribution of these effects nor their relationship with the explanatory variables. The conditional distribution $f_{Y}$ represents the parametric part of the model.
The vector $X_{it}$ contains predetermined variables
with respect to $Y_{it}$. Note that $X_{it}$ can include lags of $Y_{it}$ to accommodate dynamic models.
We consider two running examples throughout the analysis:
example[Binary response model] Let $Y_{it}$ be a binary outcome and $F$ be a cumulative
distribution function, e.g. the standard normal or standard logistic distribution. We can model the conditional distribution of $Y_{it}$ using the single-index specification with individual and time effects
\begin{equation*}
f_{Y}(y \mid X_{it}, \alpha_i, \gamma_t, \beta) = F(X_{it}'\beta+ \alpha_i + \gamma_t)^y[1 - F(X_{it}'\beta+ \alpha_i + \gamma_t)]^{1-y}, \ \ y \in \{0,1\}.
\end{equation*}
In a labor economics application, $Y$ can be an indicator for female labor force participation and $X$ can include fertility indicators and other socio-economic characteristics.
example[Poisson model] Let $Y_{it}$ be a non-negative interger-valued outcome, and $f(\cdot; \lambda)$ be the
probability mass function of a Poisson random variable with mean $\lambda > 0$. We can model the conditional distribution of $Y_{it}$ using the single index specification with individual and time effects
\begin{equation*}
f_{Y }(y \mid X_{it}, \alpha_i, \gamma_t, \beta) = f(y; \exp[X_{it}'\beta + \alpha_i + \gamma_t ]), \ \ y \in \{0, 1, 2, .... \}.
\end{equation*}
In an industrial organization application, $Y$ can be the number of patents that a firm produces and $X$ can include investment in R&D and other firm characteristics.
For estimation, we adopt a fixed effects approach, treating the realization of the unobserved individual and time effects as parameters to be estimated. We collect all these effects in the vector $\phi_{NT} = (\alpha_1, ..., \alpha_N, \gamma_1, ..., \gamma_T)'$. The model parameter $\beta$ usually includes regression coefficients of interest, while the vector $\phi_{NT}$ is treated as a nuisance parameter.
The true values of the parameters, denoted by $\beta^0$
and $\phi_{NT}^0 = ({\alpha^{0}_1}, ..., {\alpha^{0}_N}, {\gamma^{0}_1}, ..., {\gamma^{0}_T})'$, are the solution to the population conditional maximum likelihood problem
align[align omitted — 358 chars of source]
for every $N,T$, where $\mathbb{E}_{\phi}$ denotes the expectation with respect to
the distribution of the data conditional on the unobserved effects and initial conditions including strictly exogenous variables, $b > 0$ is an arbitrary constant,
$v_{NT} = ( 1_N' ,- 1_T' )'$, and $1_N$ and $1_T$ denote vectors of ones with dimensions $N$ and $T$.
Existence and uniqueness of the solution to the population problem will be
guaranteed by our assumptions in Section (ref) below, including concavity
of the objective function in all parameters.
The second term of $\mathcal{L}_{NT}$ is a penalty that imposes a normalization needed to identify $\phi_{NT}$ in
models with scalar individual and time effects that enter
additively into the log-likelihood function
as $\alpha_i + \gamma_t$.\footnote{
In Appendix (ref) we derive asymptotic expansions that apply to general models with multiple unobseved effects.
In order to use these expansions to obtain the asymptotic distribution of the panel fixed effects estimators, we need to derive the properties of the expected Hessian of the incidental parameters, a matrix with increasing dimension, and to show the consistency of the estimator of the incidental parameter vector. The additive specification $\alpha_i + \gamma_t$ is useful to characterize the Hessian and we impose strict concavity of the objective function to show the consistency.
}
In this case, adding a constant to all $\alpha_i$, while
subtracting it from all $\gamma_t$, does not change $\alpha_{i} + \gamma_t$.
To eliminate this ambiguity, we normalize
$\phi^0_{NT}$ to satisfy $v_{NT}' \phi^0_{NT} = 0$, i.e. $\sum_i \alpha_i^0 = \sum_t
\gamma_t^0$. The penalty produces a maximizer of ${\cal L}_{NT}$ that is automatically normalized.
We could equivalently impose $v_{NT}' \phi_{NT} = 0$ as a constraint, but for technical reasons we prefer to work with an unconstrained optimization problem. There are other possible normalizations for $\phi_{NT}$, such as $\alpha_1 = 0$. The model parameter $\beta$ is invariant to the choice of normalization, that is, our asymptotic results on the estimator for $\beta$
are independent of this choice of normalization.
Our choice
is convenient for certain intermediate results that involve the incidental parameter $\phi_{NT}$, its score vector and its Hessian matrix.
The pre-factor $(NT)^{-1/2}$ in $\mathcal{L}_{NT}(\beta,\phi_{NT})$ is just a rescaling.
Other quantities of interest involve averages over the data and unobserved effects
equation[equation omitted — 203 chars of source]
where $\mathbb{E}$ denotes the expectation with respect to
the joint distribution of the data and the unobserved effects, provided that the expectation exists.
$\delta^0_{NT} $ is indexed by $N$ and $T$ because the marginal distribution of $\{(X_{it}, \alpha_i, \gamma_t) : 1 \leq i \leq N, 1 \leq t \leq T\}$ can be heterogeneous across $i$ and/or $t$; see Section (ref). These averages
include average partial effects (APEs), which are often the ultimate
quantities of interest in nonlinear models. The APEs are invariant to the choice of normalization for $\phi_{NT}$ if $\alpha_i$ and $\gamma_t$ enter $\Delta(X_{it}, \beta, \alpha_i, \gamma_t)$ as $\alpha_i + \gamma_t$.
Some examples of partial effects that satisfy this condition are the following:
Example (ref) (Binary response model). If $X_{it,k}$, the $k$th element of $X_{it}$, is
binary, its partial effect on the conditional probability of $Y_{it}$ is
equation[equation omitted — 198 chars of source]
where $\beta_k$ is the $k$th element of $\beta$, and $X_{it,-k}$ and $\beta_{-k}$ include all elements of $X_{it}$ and $\beta$ except for the $k$th element. If $X_{it,k}$ is continuous and $F$ is differentiable, the partial effect of $X_{it,k}$
on the conditional probability of $Y_{it}$ is
equation[equation omitted — 145 chars of source]
where $\partial F$ is the derivative of $F$.
}
Example (ref) (Poisson model). If $X_{it}$ includes $Z_{it}$
and some known transformation $H(Z_{it})$ with coefficients $\beta_k$ and $\beta_j$, the partial effect of $Z_{it}$
on the conditional expectation of $Y_{it}$ is
equation[equation omitted — 172 chars of source]
}
Fixed effects estimators
We estimate the parameters by solving the sample analog of problem (ref), i.e.
equation[equation omitted — 142 chars of source]
As in the population case, we shall impose conditions guaranteeing that the solution to this maximization problem
exists and is unique with probability approaching one as $N$ and $T$ become large.
For computational purposes, we note that the solution to the program (ref) for $\beta$ is the same as the solution to the program that imposes $v_{NT}' \phi_{NT} = 0$ directly as a constraint in the optimization, and is invariant to the normalization.
In our numerical examples we impose either $\alpha_1 = 0$ or $\gamma_1 = 0$ directly by dropping the first individual or time effect. This constrained program has good computational properties because its objective function is concave and smooth in all the parameters. We have developed the commands probitfe and logitfe in Stata to implement the methods of the paper for probit and logit models Stata2015.\footnote{We refer to this companion work for computational details.} When $N$ and $T$ are large, e.g., $N > 2,000$ and $T > 50$, we recommend the use of optimization routines that exploit the sparsity of the design matrix of the model to speed up computation such as the package Speedglm in R Speedglm2012. For a probit model with $N=2,000$ and $T = 52$, Speedglm computes the fixed effects estimator in less than 2 minutes with a 2 x 2.66 GHz 6-Core Intel Xeon processor, more than 7.5 times faster than our \texttt{Stata} command \texttt{probitfe} and more than 30 times faster than the \texttt{R} command \texttt{glm}.\footnote{Additional comparisons of computational times are available from the authors upon request.}
To analyze the statistical properties of the estimator of $\beta$ it is
convenient to first concentrate out the nuisance parameter $\phi_{NT}$. For
given $\beta$, we define the optimal $\widehat \phi_{NT}(\beta)$ as
equation[equation omitted — 182 chars of source]
The fixed effects estimators of $\beta^0$ and $\phi_{NT}^0$ are
equation[equation omitted — 264 chars of source]
Estimators of APEs can be
formed by plugging-in the estimators of the model parameters in the sample version of (ref), i.e.
equation[equation omitted — 108 chars of source]
Again, $\widehat \delta_{NT}$ is invariant to the normalization chosen for $\phi_{NT}$ if $\alpha_i$ and $\gamma_t$ enter $\Delta(X_{it}, \beta, \alpha_i, \gamma_t)$ as $\alpha_i + \gamma_t$.
Incidental parameter problem and bias corrections
In this section we give a heuristic discussion of the main results, leaving the technical details to Section (ref). We illustrate the analysis with numerical calculations based on a variation of the classical Neyman and Scott Neyman:1948p881 variance example.
Incidental parameter problem
Fixed effects estimators in nonlinear models
suffer from the
incidental parameter problem Neyman:1948p881. The source of the problem is that the dimension of the nuisance parameter $\phi_{NT}$ increases with the sample size under asymptotic approximations where either $N$ or $T$ pass to infinity.
To describe the problem let
equation[equation omitted — 210 chars of source]
The fixed effects estimator is inconsistent under the traditional Neyman and Scott asymptotic sequences where $N \to \infty$ and $T$ is fixed, i.e., $\operatorname*{plim}_{N \to \infty} \overline{\beta}_{NT} \neq \beta^0$. Similarly, the fixed effects estimator is inconsistent under asymptotic sequences where $T \to \infty$ and $N$ is fixed, i.e., $\operatorname*{plim}_{T \to \infty} \overline{\beta}_{NT} \neq \beta^0$. Note that $\overline \beta_{NT} = \beta^0$ if $\widehat \phi_{NT}(\beta)$ is replaced by $\phi_{NT}(\beta) = \operatorname*{argmax}_{\phi_{NT} \in \mathbb{R}^{\dim \phi_{NT}}}
\, \mathbb{E}_{\phi}[ {{\cal L}}_{NT}(\beta,\,\phi_{NT})]$. Under asymptotic approximations where either $N$ or $T$ are fixed, there is only a fixed number of observations to estimate some of the components of $\phi_{NT}$, $T$ for each individual effect or $N$ for each time effect, rendering the estimator $\widehat \phi_{NT}(\beta)$ inconsistent for $\phi_{NT}(\beta)$. The nonlinearity of the model propagates the inconsistency to the estimator of $\beta$.
A key insight of the large-$T$ panel data literature is that the incidental parameter problem becomes an asymptotic bias problem under an asymptotic approximation where $N \to \infty$ and $T \to \infty$ (e.g., Arellano and Hahn, ArellanoHahn2007). For models with only individual effects, this literature derived the expansion $\overline{\beta}_{NT} = \beta^0 + B/T + o_P(T^{-1})$ as $N,T \to \infty$, for some constant $B$. The fixed effects estimator is consistent because $\operatorname*{plim}_{N,T \to \infty} \overline{\beta}_{NT} = \beta^0$, but has bias in the asymptotic distribution if $B/T$ is not negligible relative to $1/\sqrt{NT}$,
the order of the standard deviation of the estimator.
This asymptotic bias problem, however, is easier to correct than the
inconsistency problem that arises under the traditional Neyman and Scott asymptotic approximation. We show that the same insight still applies to models with individual and time effects, but with a different expansion for $\overline{\beta}_{NT}$. We characterize the expansion and develop bias corrections.
Bias Expansions and Bias Corrections
Some expansions can be used to explain our corrections.
For smooth likelihoods and under appropriate regularity conditions, as $N,T \to \infty$,
equation[equation omitted — 162 chars of source]
for some $\overline B_{\infty}^{\beta}$ and $\overline D_{\infty}^{\beta}$ that we characterize in Theorem (ref) and explain in Remark (ref),
where $a \vee b := \max(a,b)$.
Unlike in nonlinear models without incidental parameters, the order of the bias is higher than the inverse of the sample size $(NT)^{-1}$ due to the slow rate of convergence of $\widehat \phi_{NT}$. Note also that by the properties of the maximum likelihood estimator
$$
\sqrt{NT}(\widehat \beta_{NT} - \overline{\beta}_{NT}) \to_d \mathcal{N}(0, \overline V_{\infty}),
$$
for some $\overline V_{\infty}$ that we also characterize in Theorem (ref).
Under asymptotic sequences where $N/T \to \kappa^2$ as $N,T \to \infty$, the fixed effects estimator is asymptotically
biased because
align[align omitted — 375 chars of source]
Relative to fixed effects estimators with only individual effects, the presence of time effects
introduces additional asymptotic bias through $\overline D_{\infty}^{\beta}$.
This asymptotic result predicts that the fixed effects estimator can have significant bias relative to its dispersion. Moreover, confidence intervals constructed around the fixed effects estimator can severely undercover the true value of the parameter even in large samples. We show that these predictions provide a good approximations to the finite sample behavior of the fixed effects estimator through analytical and simulation examples in Sections (ref) and (ref).
The analytical bias correction consists of subtracting estimates of the leading terms of the bias from the fixed effect estimator
of $\beta^0$. Let $\widehat B_{NT}^{\beta}$ and $\widehat D_{NT}^{\beta}$ be estimators of $\overline B_{\infty}^{\beta}$ and $\overline D_{\infty}^{\beta}$ as defined in (ref). The bias corrected estimator can be formed as
equation*[equation* omitted — 134 chars of source]
If $N/T \to \kappa^2$, $\widehat B_{NT}^{\beta} \to_P \overline B_{\infty}^{\beta}$, and $\widehat D_{NT}^{\beta} \to_P \overline D_{\infty}^{\beta},$ then
equation*[equation* omitted — 105 chars of source]
The analytical correction therefore centers the asymptotic distribution at the true value of the parameter, without increasing asymptotic variance. This asymptotic result predicts that in large samples the corrected estimator has small bias relative to dispersion, the correction does not increase dispersion, and the confidence intervals constructed around the corrected estimator have coverage probabilities close to the nominal levels. We show that these predictions provide a good approximations to the behavior of the corrections in Sections (ref) and (ref) even in small panels with $N < 60$ and $T < 15$.
We also consider a jackknife bias correction method
that does not require explicit estimation of the bias.
This method is based on the split panel jackknife (SPJ) of Dhaene and Jochmans DhaeneJochmans2015 applied to the
time and cross-section dimension of the panel.
Alternative jackknife corrections based on the leave-one-observation-out panel jackknife (PJ) of Hahn and Newey Hahn:2004p882 and combinations of PJ and SPJ are also possible. We do not consider corrections based on PJ because they are theoretically justified by second-order expansions of $\overline{\beta}_{NT}$ that are beyond the scope of this paper.
To describe our generalization of the SPJ, define the fixed effects estimator of $\beta$ in the subpanel with cross sectional indexes $A$ and time series indexes $B$ as
$$
\widehat \beta_{A,B} \in \operatorname*{argmax}_{\beta \in \mathbb{R}^{\dim \beta} } \;
\max_{\alpha(A) \in \mathbb{R}^{|A|}} \;
\max_{\gamma(B) \in \mathbb{R}^{|B|}}
\; \sum_{i,t} d_{it}(A,B) \, \log f_{Y}(Y_{it} \mid X_{it}, \alpha_i, \gamma_t, \beta) ,
$$
where $\alpha(A)=( \alpha_i : i \in A )$,
$\gamma(B)=( \gamma_t : t \in B )$, and
$d_{it}(A,B) = 1(i \in A) \times 1(t \in B)$.
Let $\widetilde{\beta}_{N,T/2}$ be the average of the 2 split jackknife
estimators in the subpanels with $A = \{1,2,\ldots,N\}$, and $B = \{1,2,\ldots,T/2\}$ or $B = \{T/2+1, T/2+2, \ldots, T\}$, i.e. including all the individuals and leaving out the first and second halves of the time
periods. Let $\widetilde{\beta}_{N/2,T}$ be the average of the 2 split jackknife
estimators in the subpanels with $B = \{1,2,\ldots,T\}$, and $A = \{1,2,\ldots,N/2\}$ or $A = \{N/2+1,N/2+2,\ldots,N\}$, i.e. including all the time periods and leaving out half of the individuals of the panel.\footnote{When $T$ is odd we define $\widetilde{\beta}_{N,T/2}$ as the average of the 2 split jackknife
estimators that use overlapping subpanels with $B = \{1,2,\ldots,(T+1)/2\}$ and $B = \{(T+1)/2, (T+1)/2+1, \ldots,T\}$. We define $\widetilde{\beta}_{N/2,T}$ similarly when $N$ is odd.}
In choosing the cross sectional
indexing of the panel, one might want to take into account individual clustering structures and other dependencies to
preserve them in the SPJ. For example, all the individuals belonging to the same cluster should be indexed such that they remain in the same subpanel after the cross sectional split.
If there are no cross sectional dependencies, the indexing of the individuals is unrestricted. We recommend to construct $\widetilde{\beta}_{N/2,T}$ as the average
of the estimators obtained from all possible partitions of $N/2$ individuals to avoid ambiguity and arbitrariness in the choice of the division.\footnote{There are $P = {N \choose N/2}$ different cross sectional partitions with $N/2$ individuals. When $N$
is large, we can approximate the average over all possible partitions by the average over $S \ll P$ randomly chosen partitions to speed up computation.}
The bias corrected estimator is
equation[equation omitted — 146 chars of source]
To give some intuition about how the corrections works, note that
equation*[equation* omitted — 195 chars of source]
where
$ \widetilde{\beta}_{N,T/2} - \widehat \beta_{NT} = \overline B_{\infty}^{\beta}/T + o_P(T^{-1} \vee N^{-1})$ and $\widetilde{\beta}_{N/2,T} - \widehat \beta_{NT} = \overline D_{\infty}^{\beta}/N + o_P(T^{-1} \vee N^{-1}).$ Relative to $\widehat \beta_{NT}$, $ \widetilde{\beta}_{N,T/2} $ has double the bias coming from the estimation of the individual effects because it is based on subpanels with half of the time periods, and $ \widetilde{\beta}_{N/2,T} $ has double the bias coming from the estimation of the time effects because it is based on subpanels with half of the individuals. The time series split removes the bias term $\overline B_{\infty}^{\beta}$ and the cross sectional split removes the bias term $\overline D_{\infty}^{\beta}.$
Illustrative Example
To illustrate how the bias corrections work in finite samples, we consider a simple model where the solution to the population program (ref)
has closed form. This model corresponds to a variation of the classical Neyman and Scott Neyman:1948p881 variance example that includes both individual and time effects, $Y_{it} \mid \alpha, \gamma, \beta \sim \mathcal{N}(\alpha_i + \gamma_t, \beta)$.
It is well-know that in this case
$$
\widehat \beta_{NT} = (NT)^{-1} \sum_{i,t} \left( Y_{it} - \bar Y_{i.} - \bar Y_{.t} + \bar Y_{..} \right)^2,
$$
where $\bar Y_{i.} = T^{-1} \sum_t Y_{it}$, $\bar Y_{.t} = N^{-1} \sum_i Y_{it},$ and $\bar Y_{..} = (NT)^{-1} \sum_{i,t} Y_{it}.$ Moreover,
from the well-known results on the degrees of freedom adjustment of the estimated variance
$$
\overline{\beta}_{NT} = \mathbb{E}_{\phi} [\widehat \beta_{NT}] = \beta^0 \frac{(N-1)(T-1)}{NT} = \beta^0 \left(1 - \frac{1}{ T} - \frac{1}{ N} + \frac{1}{NT} \right),
$$
so that $\overline{B}_{\infty}^{\beta} = -\beta^0$ and $\overline{D}_{\infty}^{\beta} = -\beta^0$.\footnote{Okui Okui2013 derived the bias of fixed effects estimators of autocovariances and autocorrelations in this model.}
To form the analytical bias correction we can set $\widehat{B}_{NT}^{\beta} = - \widehat \beta_{NT}$ and $\widehat{D}_{NT}^{\beta} = - \widehat \beta_{NT}$. This yields $\widetilde \beta^A_{NT} = \widehat \beta_{NT} (1 + 1/T + 1/N)$ with
$$
\overline{\beta}^A_{NT} = \mathbb{E}_{\phi}[\widetilde \beta^A_{NT}] = \beta^0 \left( 1 - \frac{1}{T^2} - \frac{1}{N^2} - \frac{1}{NT} + \frac{1}{NT^2} + \frac{1}{N^2T}\right).
$$
This correction reduces the order of the bias from $(T^{-1} \vee N^{-1})$ to $(T^{-2} \vee N^{-2}),$ and introduces additional higher order
terms. The analytical correction increases finite-sample variance because the factor $(1 + 1/T + 1/N) > 1$. We compare the biases and standard deviations of the fixed effects estimator and the corrected estimator in a numerical example below.
For the Jackknife correction, straightforward calculations give
$$
\overline{\beta}^{J}_{NT} = \mathbb{E}_{\phi}[\widetilde \beta^{J}_{NT}] = 3 \overline{\beta}_{NT} - \overline{\beta}_{N,T/2} - \overline{\beta}_{N/2,T} = \beta^0 \left( 1 - \frac{1}{NT} \right).$$
The correction therefore reduces the order of the bias from $(T^{-1} \vee N^{-1})$ to $(TN)^{-1}.$\footnote{ In this example it is possible to develop higher-order jackknife corrections that completely eliminate the bias because we know the entire expansion of $\overline{\beta}_{NT}$. For example, $\mathbb{E}_{\phi}[ 4 \widehat \beta_{NT} - 2 \widetilde \beta_{N,T/2} - 2 \widetilde \beta_{N/2,T} + \widetilde \beta_{N/2,T/2}] = \beta^0,$ where $\widetilde \beta_{N/2,T/2}$ is the average of the four split jackknife estimators that leave out half of the individuals and the first or the second halves of the time periods. See Dhaene and Jochmans DhaeneJochmans2015 for a discussion on higher-order bias corrections of panel fixed effects estimators.}
Table (ref) presents numerical results for the bias and standard deviations of the fixed effects and bias corrected estimators in finite samples. We consider panels with $N,T \in \{10,25, 50\},$ and only report the results for $T \leq N$ since all the expressions are symmetric in $N$ and $T$. All the numbers in the table
are in percentage of the true parameter value, so we do not need to specify the value of $\beta^0$. We find that the analytical and jackknife corrections
offer substantial improvements over the fixed effects estimator in terms of bias. The first and fourth row of the table show that the bias of the fixed effects estimator is of the same order of magnitude as the standard deviation, where $\overline{V}_{NT} = {\rm Var}[\widehat \beta_{NT}] = 2 (N-1)(T-1) (\beta^0)^2 / (NT)^2$ under independence of $Y_{it}$ over $i$ and $t$ conditional on the unobserved effects. The fifth row shows the increase in standard deviation due to analytical bias correction is small compared to the bias reduction, where $\overline{V}_{NT}^A = {\rm Var}[\widetilde \beta_{NT}^A] = (1+1/N+1/T)^2 \overline{V}_{NT}$. The last row shows that the jackknife yields less precise estimates than the analytical correction when $T=10$.
table[table omitted — 1,091 chars of source]
Table (ref)
illustrates the effect of the bias on the inference based on the asymptotic distribution. It shows the
coverage probabilities of 95% asymptotic confidence intervals for $\beta^0$ constructed in the usual way as $$\text{CI}_{.95}(\widehat \beta) = \widehat \beta \pm 1.96 \widehat V_{NT}^{1/2} =
\widehat \beta (1 \pm 1.96 \sqrt{2/(NT)}),$$
where $\widehat \beta = \{\widehat \beta_{NT}, \widetilde \beta_{NT}^{A}, \widetilde \beta_{NT}^{J}\}$ and $\widehat V_{NT} = 2 \widehat \beta^2/(NT)$ is an estimator of the asymptotic variance $\overline{V}_{\infty}/(NT) = 2 (\beta^0)^2/(NT)$. To find the coverage probabilities, we use that $NT \widehat \beta_{NT}/ \beta^0 \sim \chi^2_{(N-1)(T-1)}$ and $\widetilde \beta_{NT}^{A} = (1 + 1/N + 1/T) \widehat \beta_{NT}$. These probabilities do not depend on the value of $\beta^0$ because the limits of the intervals are proportional to $\widehat \beta$. For the Jackknife we compute the probabilities numerically by simulation with $\beta^0=1$. As a benchmark of comparison, we also consider confidence intervals constructed from the unbiased estimator $\widetilde \beta_{NT} = NT \widehat \beta_{NT}/[(N-1)(T-1)]$. Here we find that the confidence intervals based on the fixed effect estimator display severe undercoverage for all the sample sizes. The confidence intervals based on the corrected estimators have high coverage probabilities, which approach the nominal level as the sample size grows. Moreover, the bias corrected estimators produce confidence intervals with very similar coverage probabilities to the ones from the unbiased estimator.
table[table omitted — 1,021 chars of source]
Asymptotic Theory for Bias Corrections
In nonlinear panel data models the population problem ((ref)) generally does not have closed form solution, so we need to
rely on asymptotic arguments to characterize the terms in the expansion of the bias ((ref)) and to justify the validity of the corrections.
Asymptotic distribution of model parameters
We consider panel models with scalar individual and time effects that enter the likelihood function
additively through $\pi_{it} = \alpha_i + \gamma_t$. In these models the dimension of the incidental parameters is $\dim \phi_{NT} = N+T$.
The leading cases are single index models, where the dependence of the likelihood function on the parameters is through an index
$ X_{it}'\beta + \alpha_i + \gamma_t$. These models cover the
probit and Poisson specifications of Examples (ref)
and (ref). The additive structure
only applies to the unobserved effects, so we can allow for scale parameters to cover the Tobit and negative
binomial models.
We focus on these additive models for computational tractability and because we can establish
the consistency of the fixed effects estimators under a concavity assumption in the log-likelihood function with respect to all the parameters.
The parametric part of our panel models takes the form
equation[equation omitted — 140 chars of source]
We denote the derivatives of the log-likelihood function $\ell_{it}$ by $\partial_\beta \ell_{it}(\beta, \pi) := \partial \ell_{it}(\beta,\pi)/\partial \beta$,
$\partial_{\beta \beta'} \ell_{it}(\beta, \pi) := \partial^2 \ell_{it}(\beta,\pi)/(\partial \beta \partial \beta')$,
$\partial_{\pi^q} \ell_{it}(\beta, \pi) := \partial^q \ell_{it}(\beta,\pi)/\partial \pi^q$, $q = 1,2,3$, etc.
We drop the arguments $\beta$ and $\pi$ when the derivatives are evaluated at the true
parameters $\beta^0$ and $\pi^0_{it} := \alpha_i^0 + \gamma_t^0$, e.g.
$\partial_{\pi^q} \ell_{it} := \partial_{\pi^q}\ell_{it}(\beta^0,\pi^0_{it})$. We also drop the dependence on
$NT$ from all the sequences of functions and parameters, e.g. we use $\mathcal{L}$ for $\mathcal{L}_{NT}$ and $\phi$
for $\phi_{NT}$.
We make the following assumptions:
assumption[Panel models]
Let $\nu > 0$ and $\mu>4 (8+\nu)/\nu$.
Let $\varepsilon>0$ and let
${\cal B}^0_{\varepsilon}$ be a subset of $\mathbb{R}^{\dim \beta+1}$
that contains an $\varepsilon$-neighbourhood of $(\beta^0,\pi^0_{it})$
for all $i,t,N,T$.\footnote{
For example, ${\cal B}^0_{\varepsilon} $ can be chosen to be the Cartesian product of the $\varepsilon$-ball around $\beta^0$
and the interval $[ \pi_{\min}, \pi_{\max}]$, with $\pi_{\min} \leq \pi_{it} - \varepsilon$ and $\pi_{\max} \geq \pi_{it} + \varepsilon$
for all $i,t,N,T$. We can have $\pi_{\min} = -\infty$ and $\pi_{\max} = \infty$, as long as this is compatible with
Assumption (ref) (iv) and (v).
}
\begin{itemize}
• Asymptotics: we consider limits of sequences where $N/T \rightarrow \kappa^2$, $0<\kappa<\infty$, as $N,T \rightarrow \infty$.
• Sampling: conditional on $\phi$, $\{(Y_i^T, X_i^T) : 1 \leq i \leq N \}$ is independent across $i$
and, for each $i$, $\{ (Y_{it}, X_{it}) : 1 \leq t \leq T\}$ is $\alpha$-mixing with mixing coefficients
satisfying $\sup_i a_i(m) = {\cal O}( m^{-\mu} )$ as $m \rightarrow \infty$, where
\begin{equation*}
a_i(m) := \sup_t \sup_{A \in \mathcal{A}_{t}^i, B \in \mathcal{B}_{t+m}^i} |P(A \cap B) - P(A) P(B)|,
\end{equation*}
and for $Z_{it} = (Y_{it}, X_{it})$, $\mathcal{A}_{t}^i$ is the sigma field generated by $(Z_{it}, Z_{i,t-1}, \ldots)$, and
$\mathcal{B}_{t}^i$ is the sigma field generated by $(Z_{it}, Z_{i,t+1}, \ldots)$.
• Model: for $X^t_i = \{X_{is}: s = 1, ..., t\}$, we assume that for all $i,t,N,T,$
\begin{equation*}
Y_{it} \mid X^t_i, \phi, \beta \sim \exp[\ell_{it}(\beta,\alpha_{i} + \gamma_t)].
\end{equation*}
The realizations of the parameters and unobserved effects that generate the observed data are denoted by
$\beta^0$ and $\phi^0$.\
• Smoothness and moments:
We assume that
$(\beta,\pi) \mapsto \ell_{it}(\beta,\pi)$ is
four times continuously differentiable over ${\cal B}^0_{\varepsilon}$ a.s.
The partial derivatives of $\ell_{it}(\beta,\pi)$
with respect to the elements of $(\beta,\pi)$
up to fourth order
are bounded in absolute value uniformly over $(\beta,\pi) \in {\cal B}^0_{\varepsilon}$
by a function $M(Z_{it})>0$ a.s.,
and $\max_{i,t} \mathbb{E}_{\phi}[M(Z_{it})^{8+\nu}]$
is a.s. uniformly bounded over $N,T$.
• Concavity:
For all $N,T,$
$(\beta,\phi) \mapsto \mathcal{L}(\beta,\phi) = (NT)^{-1/2} \{\sum_{i,t} \ell_{it} (\beta,\alpha_i + \gamma_t) - b (v'\phi)^2/2\}$
is strictly concave over $\mathbb{R}^{\dim \beta + N + T}$ a.s.
Furthermore, there exist constants $b_{\min}$
and $b_{\max}$ such that for all $(\beta,\pi) \in {\cal B}^0_{\varepsilon}$,
$0< b_{\min} \leq - \mathbb{E}_{\phi}\left[ \partial_{\pi^2} \ell_{it}(\beta,\pi) \right] \leq b_{\max}$ a.s. uniformly over $i,t,N,T$.
\end{itemize}
remark[Assumption (ref)]
Assumption (ref)$(i)$ defines the large-$T$ asymptotic framework and is the same as in Hahn and Kuersteiner HahnKuersteiner2011. The relative rate of $N$ and $T$ exactly balances the order of the bias and variance producing a non-degenerate asymptotic distribution.
Assumption (ref)$(ii)$ does not impose identical distribution nor stationarity over the time series dimension, conditional on the unobserved effects, unlike most of the large-$T$ panel literature, e.g.,
Hahn and Newey Hahn:2004p882 and
Hahn and Kuersteiner HahnKuersteiner2011.
These assumptions are violated by the presence of the time effects, because they are treated as parameters.
The mixing condition is used to bound covariances and moments in the application of laws of large numbers and central limit theorems -- it could replaced by other conditions that guarantee the applicability of these results.
Assumption (ref)$(iii)$ is the parametric part of the panel model. We rely on this assumption to guarantee that $\partial_\beta \ell_{it}$ and $\partial_\pi \ell_{it}$ have martingale difference properties. Moreover, we use certain Bartlett identities implied by this assumption to simplify some expressions, but those simplifications are not crucial for our results.
We provide expressions for the asymptotic bias and variance that do not apply these simplifications in Remark (ref) below.
Assumption (ref)$(iv)$ imposes smoothness and moment conditions in the log-likelihood function and its derivatives. These conditions
guarantee that the higher-order stochastic expansions of the fixed effect estimator that we use to characterize the asymptotic bias are well-defined, and that the remainder terms of these expansions are bounded.
The most commonly used nonlinear models in applied economics such as logit, probit, ordered probit, Poisson, and Tobit models have
smooth log-likelihoods functions that satisfy the concavity condition of Assumption (ref)$(v)$, provided that all the elements of $X_{it}$ have cross sectional and time series variation.
Assumption (ref)$(v)$ guarantees that $\beta^0$ and $\phi^0$ are the unique solution to the
population problem (ref), that is all the parameters are point identified.
To describe the asymptotic distribution of the fixed effects estimator $\widehat \beta,$ it is convenient to introduce some additional
notation. Let $\overline{\cal H}$ be the $(N+T) \times (N+T)$ expected Hessian matrix of
the log-likelihood with respect to the nuisance parameters evaluated at the true parameters, i.e.
align[align omitted — 379 chars of source]
where $\overline{\mathcal{H}}_{(\alpha\alpha)}^* = \text{diag}( \sum_{t} \mathbb{E}_{\phi}[- \partial_{\pi^2} \ell_{it}])/\sqrt{NT}$,
$\overline{\mathcal{H}}_{(\alpha\gamma)it}^* = \mathbb{E}_{\phi}[- \partial_{\pi^2} \ell_{it}]/\sqrt{NT}$,
and
$\overline{\mathcal{H}}_{(\gamma\gamma)} ^*= \linebreak \text{diag}( \sum_{i} \mathbb{E}_{\phi}[- \partial_{\pi^2} \ell_{it}])/\sqrt{NT}$.
Furthermore, let
$\overline{\cal H}^{-1}_{(\alpha\alpha)}$, $\overline{\cal H}^{-1}_{(\alpha\gamma)}$,
$\overline{\cal H}^{-1}_{(\gamma\alpha)}$,
and $\overline{\cal H}^{-1}_{(\gamma\gamma)}$ denote the $N\times N$, $N\times T$, $T\times N$
and $T\times T$ blocks of the inverse $\overline{\cal H}^{-1}$ of $\overline{\cal H}$.
We define the $\dim \beta$-vector $\Xi_{it}$ and the operator $D_{\beta \pi^q}$
as
align[align omitted — 520 chars of source]
with $q=0,1,2$. The $k$-th component of $\Xi_{it}$ corresponds to the population least squares projection of
$\mathbb{E}_{\phi}(\partial_{\beta_k \pi} \ell_{it})/\mathbb{E}_{\phi}( \partial_{\pi^2} \ell_{it} )$
on the space spanned by the incidental parameters under a
metric given by $ \mathbb{E}_{\phi}( - \partial_{\pi^2} \ell_{it})$, i.e.
align*[align* omitted — 457 chars of source]
The operator $D_{\beta \pi^q}$ partials out individual and time effects in nonlinear models. It corresponds to individual and time differencing when the model is linear. To see this, consider the normal linear model $Y_{it} \mid X_{i}^t, \alpha_i,\gamma_t \sim {\cal N}(X_{it}'\beta + \alpha_i + \gamma_t, 1).$ Then, $\Xi_{it} = T^{-1} \sum_{t=1}^T \mathbb{E}_{\phi}[X_{it}] + N^{-1} \sum_{i=1}^N \mathbb{E}_{\phi}[X_{it}] - (NT)^{-1} \sum_{i=1}^N \sum_{t=1}^T \mathbb{E}_{\phi}[X_{it}]$, $D_{\beta} \ell_{it} = - \tilde X_{it} \varepsilon_{it}, $ $D_{\beta \pi} \ell_{it} = - \tilde X_{it},$ and $D_{\beta \pi^2} \ell_{it} = 0$, where $\varepsilon_{it} = Y_{it} - X_{it}'\beta - \alpha_i - \gamma_t$ and $\tilde X_{it} = X_{it} - \Xi_{it}$ is the individual and time demeaned explanatory variables.
The following theorem establishes the asymptotic distribution of the fixed effects estimator $\widehat \beta.$
theorem[Asymptotic distribution of $\widehat \beta$]
Suppose that Assumption (ref) holds, that the following limits exist
\begin{align*}
\overline B_{\infty} &=
\overline{\mathbb{E}} \left[ - \frac {1} {N} \sum_{i=1}^{N}
\frac{ \sum_{t=1}^T \sum_{\tau=t}^T
\mathbb{E}_{\phi}\left(
\partial_{\pi} \ell_{it} D_{\beta \pi} \ell_{i\tau}
\right)
+ \frac 1 2 \sum_{t=1}^T
\mathbb{E}_{\phi} ( D_{\beta \pi^2} \ell_{it} ) }
{ \sum_{t=1}^T \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) } \right] ,
\nonumber \\
\overline D_{\infty} &= \overline{\mathbb{E}} \left[ -
\frac {1} {T} \sum_{t=1}^{T}
\frac{ \sum_{i=1}^N
\mathbb{E}_{\phi}\left(
\partial_{\pi} \ell_{it} D_{\beta \pi} \ell_{it}
+ \frac 1 2 D_{\beta \pi^2} \ell_{it} \right) }
{ \sum_{i=1}^N \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) } \right], \\
\overline W_{\infty} &= \overline{\mathbb{E}} \left[ - \frac 1 {NT} \sum_{i=1}^N
\sum_{t=1}^T \mathbb{E}_{\phi} \left(
\partial_{\beta \beta'} \ell_{it}
- \partial_{\pi^2} \ell_{it} \Xi_{it} \Xi'_{it} \right) \right],
\end{align*}
and that $\overline W_{\infty}>0$.
Then,
\begin{align*}
\sqrt{NT} \left( \widehat \beta - \beta^0 \right)
\; \to_d \;
\overline{W}_{\infty}^{-1} {\cal N}( \kappa \overline B_{\infty}
+ \kappa^{-1} \overline D_{\infty} ,
\;\overline W_{\infty}),
\end{align*}
so that $\overline B_{\infty}^{\beta} = \overline W_{\infty}^{-1} \overline B_{\infty}$, $\overline D_{\infty}^{\beta} = \overline W_{\infty}^{-1} \overline D_{\infty}$,
and $\overline V_{\infty} = \overline W_{\infty}^{-1}$
in (ref) and (ref).
remarkThe complete proof of Theorem (ref) is provided in the Appendix. Here we point out why the argument for the consistency proof in models with only individual effects does not apply to our setting, give a heuristic derivation of the asymptotic distribution, and highlight where some of the assumptions are used in the proof.
\begin{itemize}
• The consistency proof for models with only individual effects relies on partitioning the log-likelihood in the sum of individual log-likelihoods that depend on a fixed number of parameters, the model parameter $\beta$ and the corresponding individual effect $\alpha_i$. The maximizers of the individual log-likelihood are then consistent estimators of all the parameters as $T$ becomes large by standard arguments. This approach does not work in models with individual and time effects because there is no partition of the data that is only affected by a fixed number of parameters, and whose size grows with the sample size.
• In the following we give a heuristic discussion of the asymptotic distribution result for $\widehat \beta$. A first-order Taylor series expansion to approximate the first order conditions of (ref) around $\beta^0$ gives
\begin{equation}
0 = \partial_\beta {\cal L}(\widehat \beta,\widehat \phi(\widehat \beta)) \approx \partial_\beta {\cal L}(\beta^0,\widehat \phi^0) - \overline W_{\infty} \sqrt{NT}(\widehat \beta - \beta^0),
\end{equation}
where $\widehat \phi^0 = \widehat \phi(\beta^0)$.
A second-order Taylor series expansion to approximate $\partial_\beta {\cal L}(\beta^0,\widehat \phi^0)$ around $\phi^0$ yields
$$
\partial_{\beta} {\cal L}(\beta^0, \widehat \phi^0) \approx \partial_{\beta} {\cal L}(\beta^0,\phi^0) + \partial_{\beta \phi'} {\cal L}(\beta^0,\phi^0) [\widehat \phi^0 - \phi^0] + \sum_{g=1}^{\dim \phi} \partial_{\beta \phi' \phi_g} {\cal L}(\beta^0,\phi^0) [\widehat \phi^0 - \phi^0][\widehat \phi_g^0 - \phi_g^0]/2,
$$
where the first term has zero mean and determines the asymptotic variance, and the second and third term determine the asymptotic bias. Thus, by the central limit theorem and the information equality,
$$
\partial_{\beta} {\cal L}(\beta^0,\phi^0) \to_d \mathcal{N}(0, \overline W_{\infty}).
$$
The second and third terms satisfy
\begin{equation*}
\partial_{\beta \phi'} {\cal L}(\beta^0,\phi^0) [\widehat \phi^0 - \phi^0] + \sum_{g=1}^{\dim \phi} \partial_{\beta \phi' \phi_g} {\cal L}(\beta^0,\phi^0) [\widehat \phi^0 - \phi^0][\widehat \phi_g^0 - \phi_g^0]/2 \approx \sqrt{NT}(\overline B_{\infty}/T + \overline D_{\infty}/N),
\end{equation*}
where $\overline B_{\infty}$ and $\overline D_{\infty}$ are characterized from a second-order Taylor series expansion to approximate $\widehat \phi^0$ around $\phi^0$. We refer to the Appendix for the details of this derivation. There we show that $\overline B_{\infty}$ and $\overline D_{\infty}$ originate from the elements of $\widehat \phi^0$ corresponding to the individual effects and time effects, respectively. Plugging those results into (ref), and solving for $\sqrt{NT}(\widehat{\beta} - \beta^0)$ yields
$$
\sqrt{NT}(\widehat{\beta} - \beta^0) \approx \overline W_{\infty}^{-1} [\partial_{\beta} {\cal L}(\beta^0,\phi^0) + \overline B_{\infty}\sqrt{N/T} + \overline D_{\infty} \sqrt{T/N} ] \to_d \overline W_{\infty}^{-1} \mathcal{N}(\kappa \overline B_{\infty} + \kappa^{-1} \overline D_{\infty}, W_{\infty}).
$$
This derivation shows that the source of the bias is that the score $\partial_{\beta} {\cal L}(\beta,\widehat \phi)$ is not centered at zero when $\beta = \beta^0$. This problem arises from the substitution of the incidental parameter $\phi$ by the sample analog $\widehat \phi^0$ that has a rate of convergence slower than $\sqrt{NT}$. Thus, $\overline B_{\infty}$ originates from the estimators of the individual effects in $\phi$, which have rate of convergence $\sqrt{T}$; whereas $\overline D_{\infty}$ originates from the estimators of the time effects in $\phi$, which have convergence rate $\sqrt{N}$.
• The two key assumptions in the derivation of the asymptotic distribution are the additive separability of $\alpha_i$ and $\gamma_t$ in Assumption (ref)(iii) and the concavity in Assumption (ref)(v). We resort to concavity to prove consistency of $\widehat \beta$ and to bound the remainder terms in all the expansions. Additive separability is convenient to characterize the order of the inverse average Hessian, $\overline{\mathcal{H}}$, defined in (ref). This inverse Hessian features prominently in the second-order Taylor series expansion of $\widehat \phi$ around $\phi^0$ used to characterize $\overline B_{\infty}$ and $\overline D_{\infty}$.
\end{itemize}
It is instructive to evaluate the expressions of the bias in our running examples.
Example (ref) (Binary response model). In this case
$$\ell_{it}(\beta,\pi) = Y_{it} \log F(X_{it}'\beta + \pi) + (1 - Y_{it}) \log [1 - F(X_{it}'\beta + \pi)],$$ so that $\partial_{\pi} \ell_{it} = H_{it} (Y_{it} - F_{it}),$ $\partial_{\beta} \ell_{it} = \partial_{\pi} \ell_{it} X_{it},$ $ \partial_{{\pi}^2} \ell_{it} = - H_{it} \partial F_{it} + \partial H_{it} (Y_{it} - F_{it})$, $\partial_{\beta \beta'} \ell_{it} = \partial_{\pi^2} \ell_{it} X_{it} X_{it}'$, $\partial_{\beta \pi} \ell_{it} = \partial_{\pi^2} \ell_{it} X_{it},$ $\partial_{{\pi}^3} \ell_{it} = - H_{it} \partial^2 F_{it} - 2 \partial H_{it} \partial F_{it} + \partial^2 H_{it} (Y_{it} - F_{it})$, and $\partial_{\beta {\pi}^2} \ell_{it} = \partial_{{\pi}^3} \ell_{it} X_{it}$, where $H_{it} = \partial F_{it} / [F_{it}(1-F_{it})], and $ $\partial^{j} G_{it} := \partial^{j} G(Z)|_{Z = X_{it}'\beta^0 + \pi_{it}^0}$ for any function $G$ and $j = 0,1,2$. Substituting these values in the expressions of the bias of Theorem (ref) yields
eqnarray*[eqnarray* omitted — 905 chars of source]
where $\omega_{it} = H_{it} \partial F_{it}$ and $\tilde X_{it}$ is the residual of the population projection of $X_{it}$ on the space spanned by the incidental parameters under a metric weighted by $\mathbb{E}_{\phi}(\omega_{it})$. For the probit model where all the components of $X_{it}$ are strictly exogenous,
$$
\overline B_{\infty} = \overline{\mathbb{E}} \left[
\frac {1} {2N} \sum_{i=1}^{N}
\frac{ \sum_{t=1}^T \mathbb{E}_{\phi}[\omega_{it} \tilde{X}_{it} \tilde{X}_{it}'] } { \sum_{t=1}^T \mathbb{E}_{\phi}\left( \omega_{it} \right) }\right] \beta^0 , \ \ \overline D_{\infty} = \overline{\mathbb{E}} \left[
\frac {1} {2T} \sum_{t=1}^{T}
\frac{ \sum_{i=1}^N \mathbb{E}_{\phi}[\omega_{it} \tilde{X}_{it} \tilde{X}_{it}'] } { \sum_{i=1}^N \mathbb{E}_{\phi}\left( \omega_{it} \right) } \right] \beta^0.
$$
The asymptotic bias is therefore a positive definite matrix weighted average of the true parameter value as in the case of the probit model with only individual effects FernandezVal:2009p3313.}
Example (ref) (Poisson model). In this case
$$\ell_{it}(\beta,\pi) = (X_{it}'\beta + \pi) Y_{it} - \exp(X_{it}'\beta + \pi) - \log Y_{it}!,$$
so that $\partial_{\pi} \ell_{it} = Y_{it} - \omega_{it},$ $\partial_{\beta} \ell_{it} = \partial_{\pi} \ell_{it} X_{it},$ $\partial_{{\pi}^2} \ell_{it} = \partial_{{\pi}^3} \ell_{it}= - \omega_{it} $, $\partial_{\beta \beta'} \ell_{it} = \partial_{{\pi}^2} \ell_{it} X_{it} X_{it}',$ and $\partial_{\beta {\pi}} \ell_{it} = \partial_{\beta {\pi}^2} \ell_{it} = \partial_{{\pi}^3} \ell_{it} X_{it},$ where $\omega_{it} = \exp(X_{it}'\beta^0 + \pi_{it}^0)$. Substituting these values in the expressions of the bias of Theorem (ref) yields
eqnarray*[eqnarray* omitted — 558 chars of source]
and $\overline D_{\infty} = 0$, where $\tilde X_{it}$ is the residual of the population projection of $X_{it}$ on the space spanned by the incidental parameters under a metric weighted by $\mathbb{E}_{\phi}(\omega_{it})$. If in addition
all the components of $X_{it}$ are strictly exogenous, then we get the no asymptotic bias result $\overline B_{\infty} = \overline D_{\infty} = 0$.}
remark[Bias and Variance expressions for Conditional Moment Models]
In the derivation of the asymptotic distribution, we apply Bartlett identities implied by Assumption (ref)$(iii)$ to simplify the expressions. The following expressions of the asymptotic bias and variance do not make use of these identities and therefore remain valid in conditional moment models that do not specify the entire conditional distribution of $Y_{it}$:
\begin{align*}
\overline B_{\infty} &=
\overline{\mathbb{E}} \left[ - \frac {1} {N} \sum_{i=1}^{N}
\frac{ \sum_{t=1}^T \sum_{\tau=t}^T
\mathbb{E}_{\phi}\left(
\partial_\pi \ell_{it} D_{\beta \pi} \ell_{i\tau} \right) }
{ \sum_{t=1}^T \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) } \right] \nonumber \\
& \quad + \, \frac 1 2 \, \overline{\mathbb{E}} \left[ \,
\frac 1 {N} \sum_{i=1}^N
\frac{ \sum_{t=1}^T
\mathbb{E}_{\phi} [ ( \partial_{\pi} \ell_{it} )^2 ]
\sum_{t=1}^T \mathbb{E}_{\phi} ( D_{\beta \pi^2 } \ell_{it} ) }
{\left[ \sum_{t=1}^T \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) \right]^2} \right] \; ,
\nonumber \\
\overline D_{\infty} &= \overline{\mathbb{E}} \left[ -
\frac {1} {T} \sum_{t=1}^{T}
\frac{ \sum_{i=1}^N
\mathbb{E}_{\phi}\left[
\partial_{\pi} \ell_{it} D_{\beta \pi} \ell_{it} \right] }
{ \sum_{i=1}^N \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) } \right]
\nonumber \\
& \quad + \, \frac 1 2 \, \overline{\mathbb{E}} \left[
\frac 1 {T} \sum_{t=1}^T
\frac{ \sum_{i=1}^N
\mathbb{E}_{\phi} [ ( \partial_{\pi} \ell_{it} )^2 ]
\sum_{i=1}^N \mathbb{E}_{\phi} ( D_{\beta \pi^2} \ell_{it}
) }
{\left[ \sum_{i=1}^N \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) \right]^2} \right],
\nonumber \\
\overline V_{\infty} &=
\overline W_{\infty}^{-1}
\; \overline \Omega_{\infty} \overline W_{\infty}^{-1}, \\
\overline \Omega_{\infty} &=
\overline{\mathbb{E}} \left[ \frac 1 {NT} \sum_{i=1}^N
\sum_{t=1}^T \sum_{\tau = 1}^T
\mathbb{E}_{\phi} \left[
D_{\beta} \ell_{it} (D_{\beta} \ell_{i\tau})' \right] \right],
\end{align*}
and $ \overline W_{\infty}$ is the same as in Theorem (ref).
For example, consider the Poisson fixed effects estimator in the conditional mean model $ \mathbb{E} [Y_{it} \mid X_{i}^t, \phi, \beta] = \omega_{it} = \exp( X_{it}'\beta + \alpha_i + \gamma_t)$. Applying the previous expressions to $\ell_{it}(\beta,\pi) = (X_{it}'\beta + \pi) Y_{it} - \exp(X_{it}'\beta + \pi) - \log Y_{it}!$ yields the same expressions for $\overline B_{\infty}$, $\overline D_{\infty}$, $\overline W_{\infty}$ as in Example 2, and
$$
\overline \Omega_{\infty} =
\overline{\mathbb{E}} \left[ \frac 1 {NT} \sum_{i=1}^N
\sum_{t=1}^T
\mathbb{E}_{\phi} \left[
(Y_{it} - \omega_{it})^2 \tilde{X}_{it} \tilde{X}_{it}' \right] \right],
$$
where $\tilde{X}_{it}$ is defined as in Example 2. If all the components of $X_{it}$ are strictly exogenous, then we get again the no asymptotic bias result $\overline B_{\infty} = \overline D_{\infty} = 0$.
Asymptotic distribution of APEs
In nonlinear models we are often interested in APEs, in addition to model parameters. These effects
are averages of the data, parameters and unobserved effects; see expression
(ref). For the panel models of Assumption (ref) we specify the partial effects as $\Delta(X_{it}, \beta, \alpha_i, \gamma_t) = \Delta_{it}(\beta, \pi_{it})$.
The restriction that the partial effects depend on $\alpha_i$ and $\gamma_t$ through $\pi_{it}$ is natural in our panel models since
$$
\mathbb{E} [Y_{it} \mid X^t_i, \alpha_i, \gamma_t, \beta] = \int y \exp [\ell_{it}(\beta, \, \pi_{it} )] dy,
$$
and the partial effects are usually defined as differences or derivatives of this conditional expectation with respect to the components of $X_{it}$. For example, the partial effects for the binary response and Poisson
models described in Section (ref) satisfy this restriction.
The distribution of the unobserved individual and time effects is not ancillary for the APEs, unlike for model parameters. We therefore need to make assumptions on this distribution to define and interpret the APEs, and to derive the asymptotic distribution of their estimators. We control the heterogeneity of the partial effects assuming that the individual effects and explanatory variables are identically distributed cross sectionally and/or stationary over time. If $(X_{it}, \alpha_i, \gamma_t)$ is identically distributed over $i$ and can be heterogeneously distributed over $t$, $\mathbb{E}[\Delta_{it}] = \delta_t^0$ and $\delta_{NT}^0 = T^{-1} \sum_{t=1}^T \delta_t^0$ changes only with $T$. If $(X_{it}, \alpha_i, \gamma_t)$ is stationary over $t$ and can be heterogeneously distributed over $i$, $\mathbb{E}[\Delta_{it}] = \delta_i^0$ and $\delta_{NT}^0 = N^{-1} \sum_{i=1}^N \delta_i^0$ changes only with $N$. Finally, if $(X_{it}, \alpha_i, \gamma_t)$ is identically distributed over $i$ and stationary over $t$, $\mathbb{E}[\Delta_{it}] = \delta_{NT}^0$ and $\delta_{NT}^0 = \delta^0$ does not change with $N$ and $T.$
We also impose smoothness and moment conditions on the function $\Delta$ that defines the partial effects. We use these conditions to derive higher-order stochastic expansions for the fixed effect estimator of the APEs and to bound the remainder terms in these expansions. Let $\{\alpha_i\}_N := \{\alpha_i : 1 \leq i \leq N\}$, $\{\gamma_t\}_T := \{\gamma_t : 1 \leq t \leq T\},$ and $\{X_{it}, \alpha_i, \gamma_t \}_{NT} := \{(X_{it}, \alpha_i, \gamma_t) : 1 \leq i \leq N, 1 \leq t \leq T\}.$
assumption[Partial effects]
Let $\nu>0$, $\epsilon>0$, and ${\cal B}^0_{\varepsilon}$ all be as in
Assumption (ref).
\begin{itemize}
• Sampling: for all $N,T,$ $\{X_{it}, \alpha_i, \gamma_t \}_{NT}$ is identically distributed across $i$ and/or stationary across $t$.\footnote{In the working paper version,
Fern\'andez-Val and Weidner ThisWorkingPaper2015, we also consider inference conditional on the unobserved effects by assuming that $\{\alpha_i\}_N$ and $\{\gamma_t\}_T$ are deterministic sequences.}
• Model: for all $i,t,N,T,$
the partial effects depend on $\alpha_i$ and $\gamma_t$ through $\alpha_i + \gamma_t$:
\begin{equation*}
\Delta(X_{it}, \beta, \alpha_i, \gamma_t) = \Delta_{it}(\beta, \alpha_i + \gamma_t).
\end{equation*}
The realizations of the partial effects are denoted by $ \Delta_{it} := \Delta_{it}(\beta^0, \alpha_i^0 + \gamma_t^0).$
• Smoothness and moments: The function
$(\beta,\pi) \mapsto \Delta_{it}(\beta,\pi)$
is
four times continuously differentiable over ${\cal B}^0_{\varepsilon}$ a.s.
The partial derivatives of $\Delta_{it}(\beta,\pi)$
with respect to the elements of $(\beta,\pi)$
up to fourth order
are bounded in absolute value uniformly over $(\beta,\pi) \in {\cal B}^0_{\varepsilon}$
by a function $M(Z_{it})>0$ a.s.,
and $\max_{i,t} \mathbb{E}_{\phi}[M(Z_{it})^{8+\nu}]$
is a.s. uniformly bounded over $N,T$.
• Non-degeneracy and moments: $0 < \min_{i,t} [\mathbb{E} (\Delta_{it}^2) - \mathbb{E}(\Delta_{it})^2] \leq \max_{i,t} [\mathbb{E} (\Delta_{it}^2) - \mathbb{E}(\Delta_{it})^2] < \infty,$ uniformly over $N,T.$
\end{itemize}
Analogous to $\Xi_{it}$ and $D_{\beta \pi^q} \ell_{it} $ in equation (ref) we define
align[align omitted — 477 chars of source]
for $q \in \{1,2\}$.
Here, $\Psi_{it}$ is the population projection of
$ \partial_{\pi} \Delta_{it} /
\mathbb{E}_{\phi}[ \partial_{\pi^2} \ell_{it} ]$
on the space spanned by the incidental parameters under the
metric given by $ \mathbb{E}_{\phi}[ - \partial_{\pi^2} \ell_{it} ]$.
We use analogous notation to the previous section for the derivatives with respect to $\beta$ and higher order derivatives with respect to $\pi$.
Let $\delta_{NT}^0$ and $\widehat \delta$ be the APE and its fixed effects estimator, defined as in equations (ref) and (ref) with $\Delta(X_{it}, \beta, \alpha_i, \gamma_t) = \Delta_{it}(\beta, \alpha_i + \gamma_t).$\footnote{We keep the dependence of $\delta_{NT}^0$ on $NT$ to distinguish $\delta_{NT}^0$ from $\delta^0 = \lim_{N,T \to \infty} \delta_{NT}^0$.}
The following theorem establishes the asymptotic distribution of $\widehat \delta.$
theorem[Asymptotic distribution of $\widehat \delta$]
Suppose that the assumptions of Theorem (ref)
and Assumption (ref) hold, and that the following limits exist:\footnote{We thank Fa Wang for pointing out errors in the expressions
for $ \overline B_{\infty}^{\delta} $, $ \overline D_{\infty}^{\delta} $, and
$\overline{V}_{\infty}^{\delta}$ in the published version of the paper.}
\begin{align*}
\overline {(D_{\beta} \Delta)}_{\infty} &= \overline{\mathbb{E}} \left[
\frac 1 {NT} \sum_{i=1}^N \sum_{t=1}^T
\mathbb{E}_{\phi}( \partial_{\beta} \Delta_{it} - \Xi_{it} \partial_{\pi} \Delta_{it} ) \right],
\nonumber \\
\overline B_{\infty}^{\delta} &=
\overline {(D_{\beta} \Delta)}_{\infty}' \overline W_\infty^{-1} \overline B_{\infty}
- \overline{\mathbb{E}} \left[ \frac {1} {N} \sum_{i=1}^{N}
\frac{ \sum_{t=1}^T \sum_{\tau=t}^T
\mathbb{E}_{\phi}\left(
\partial_{\pi} \ell_{it} D_{\pi} \Delta_{i\tau}
\right)
+ \frac 1 2 \sum_{t=1}^T
\mathbb{E}_{\phi} ( D_{\pi^2} \Delta_{it} ) }
{ \sum_{t=1}^T \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) } \right] ,
\nonumber \\
\overline D_{\infty}^{\delta} &= \overline {(D_{\beta} \Delta)}_{\infty}' \overline W_\infty^{-1} \overline D_{\infty} -
\overline{\mathbb{E}} \left[
\frac {1} {T} \sum_{t=1}^{T}
\frac{ \sum_{i=1}^N
\mathbb{E}_{\phi}\left(
\partial_{\pi} \ell_{it} D_{\pi} \Delta_{it}
+ \frac 1 2 D_{\pi^2} \Delta_{it} \right) }
{ \sum_{i=1}^N \mathbb{E}_{\phi}\left( \partial_{\pi^2} \ell_{i t} \right) } \right],
\nonumber \\
\overline{V}_{\infty}^{\delta} &= \overline{\mathbb{E}} \left\{
\frac {r_{NT}^2} {N^2T^2}
\mathbb{E} \left[ \left(\sum_{i=1}^N \sum_{t = 1}^T \widetilde \Delta_{it} \right)\left(\sum_{i=1}^N \sum_{t = 1}^T \widetilde \Delta_{it} \right)' + \sum_{i=1}^N \sum_{t=1}^T \Gamma_{it} \Gamma_{it}' + 2 \sum_{i=1}^N \left(\sum_{t = 1}^T \widetilde \Delta_{it}\sum_{s = t+1}^T \Gamma_{is}' \right) \right] \right\},
\end{align*}
for some deterministic sequence $r_{NT} \to \infty$ such that $r_{NT} = \mathcal{O}(\sqrt{NT})$ and $ \overline{V}_{\infty}^{\delta} > 0,$
where $ \widetilde \Delta_{it} = \Delta_{it} - \mathbb{E}(\Delta_{it})$ and $\Gamma_{it}= \overline {(D_{\beta} \Delta)}_{\infty}' \overline W_\infty^{-1} D_{\beta} \ell_{it}
- \mathbb{E}_{\phi}( \Psi_{it} )
\partial_{\pi} \ell_{it}$.
Then,
\begin{equation*}
r_{NT} (\widehat \delta - \delta_{NT}^0 - T^{-1} \overline B_{\infty}^{\delta}
- N^{-1} \overline D_{\infty}^{\delta}) \to_d \mathcal{N}(0 ,
\;\overline V_{\infty}^{\delta}) .
\end{equation*}
remark[Convergence rate, bias and variance] To understand the asymptotic distribution of $\widehat \delta$ is useful to decompose
$$
\widehat \delta - \delta_{NT}^0 = [\widehat \delta - \delta] + [\delta - \delta_{NT}^0],
$$
where $\delta := (NT)^{-1} \sum_{i=1}^N \sum_{t=1}^T \Delta_{it}$. In this decomposition the first term captures variation due to parameter estimation, whereas the second term captures variation due to estimation of a population mean by a sample mean. Under Assumption (ref)(iv) the convergence rate $r_{NT}$ is determined by the convergence rate of $\delta - \delta_{NT}^0$, which depends on the sampling properties of the unobserved effects. For example, if $\{\alpha_i\}_N$ and $\{\gamma_t\}_T$ are independent sequences, and $\alpha_i$ and $\gamma_t$ are independent for all $i,t$, then $r_{NT} = \sqrt{NT/(N+T-1)}$, and
$$
\overline{V}_{\infty}^{\delta} = \overline{\mathbb{E}} \left\{
\frac {r_{NT}^2} {N^2 T^2}
\sum_{i=1}^N \left[ \sum_{t, \tau = 1}^T \mathbb{E}( \widetilde \Delta_{it} \widetilde \Delta_{i \tau}' ) + \sum_{j \neq i}\sum_{t = 1}^T \mathbb{E}( \widetilde \Delta_{it} \widetilde \Delta_{j t} ') + \sum_{t=1}^T \mathbb{E}(\Gamma_{it} \Gamma_{it}') + 2 \sum_{s>t} \mathbb{E}(\widetilde \Delta_{it} \Gamma_{is}')\right] \right\}.
$$
In the expression of $ \overline{V}_{\infty}^{\delta}$, the first two terms come from $\delta - \delta_{NT}^0$, the third term comes from $\widehat \delta - \delta$, and the last term is the asymptotic covariance between $\delta - \delta_{NT}^0$ and $\widehat \delta - \delta$. The last term drops out when all the components of $X_{it}$ are strictly exogenous.
The first two terms of $ \overline{V}_{\infty}^{\delta}$ are of order $NT(T + N - 1)r_{NT}^2/(NT)^2 = \mathcal{O}(1)$ by construction, the last term of $ \overline{V}_{\infty}^{\delta}$ is of order $NT r_{NT}^2/(NT)^2 = \mathcal{O}(T^{-1}+N^{-1})$, and the asymptotic bias $ r_{NT} (T^{-1} \overline B_{\infty}^{\delta} + N^{-1} \overline D_{\infty}^{\delta})$ is of order $r_{NT}(T^{-1} + N^{-1}) = \mathcal{O}(T^{-1/2}+N^{-1/2})$. Thus, the bias and variance coming from parameter estimation are asymptotically negligible relative to the variances coming from the estimation of a population mean by a sample mean. In numerical examples, however, we find that correcting the mean and variance for parameter estimation improves the finite-sample estimation and inference properties of the APE estimators.
remark[Average effects from bias corrected estimators] The first term in the expressions of the biases $ \overline B_{\infty}^{\delta}$ and $ \overline D_{\infty}^{\delta}$ comes from the bias of the estimator of $\beta$. It drops out when the APEs are
constructed from asymptotically unbiased or bias corrected estimators of the parameter $\beta$, i.e.
\begin{equation*}
\widetilde \delta = \Delta( \widetilde
\beta, \widehat \phi(\widetilde \beta)),
\end{equation*}
where $\widetilde \beta$ is such that $\sqrt{NT}(\widetilde \beta - \beta^0) \to_d N(0, \overline W_{\infty}^{-1})$.
The asymptotic variance of $\widetilde \delta$ is the same as in Theorem (ref).
In the following examples we assume that the APEs are constructed from asymptotically unbiased estimators of the model parameters.
Example (ref) (Binary response model). Consider the partial effects defined in ((ref)) and ((ref)) with
$$
\Delta_{it}(\beta , \pi) = F(\beta_k +
X_{it,-k}'\beta_{-k} + \pi) -
F(X_{it,-k}'\beta_{-k} + \pi) \text{ and } \Delta_{it}(\beta, \pi) = \beta_k \partial F(X_{it}'\beta +
\pi).
$$
Using the notation previously introduced for this example, the components of the asymptotic bias of $\widetilde \delta$ are
align*[align* omitted — 894 chars of source]
where $\tilde \Psi_{it}$ is the residual of the population regression of
$ - \partial_{\pi} \Delta_{it} /
\mathbb{E}_{\phi}[ \omega_{it}]$
on the space spanned by the incidental parameters under the
metric given by $\mathbb{E}_{\phi}[\omega_{it}]$. If all the components of $X_{it}$ are strictly exogenous,
the first term of $ \overline B_{\infty}^{\delta}$ is zero.}
Example (ref) (Poisson model). Consider the partial effect
$$
\Delta_{it}(\beta, \pi) = g_{it}(\beta) \exp(X_{it}'\beta + \pi),
$$
where $g_{it}$ does not depend on $\pi$. For example, $g_{it}(\beta) = \beta_k + \beta_j
h(Z_{it})$ in ((ref)). Using the notation previously introduced for this example, the components of the asymptotic bias are
equation*[equation* omitted — 366 chars of source]
and $\overline D_{\infty}^{\delta} = 0$, where $\tilde g_{it}$ is the residual of the population projection of $g_{it}$ on the space spanned by the incidental parameters under a metric weighted by $\mathbb{E}_{\phi}[\omega_{it}]$. The asymptotic bias is zero if all the components of $X_{it}$ are strictly exogenous or $g_{it}(\beta)$ is constant. The latter arises in the leading case of the partial effect of the $k$-th component of $X_{it}$ since $g_{it}(\beta) = \beta_k$. This no asymptotic bias result applies to any type of regressor, strictly exogenous or predetermined.}
Bias corrected estimators
The results of the previous sections show that the asymptotic
distributions of the fixed effects estimators of the model parameters and APEs can have biases of the same
order as the variances under sequences where $T$ grows at
the same rate as $N$. This is the large-$T$ version of the
incidental parameters problem that invalidates
any inference based on the fixed effect estimators even in large samples.
In this section we describe how to construct analytical and jackknife bias corrections for
the fixed effect estimators and give conditions for the asymptotic validity of
these corrections.
The jackknife correction for the model parameter $\beta$ in equation (ref) is generic and applies to the panel model. For the APEs, the jackknife correction is formed similarly as
equation*[equation* omitted — 150 chars of source]
where $\widetilde{\delta}_{N,T/2}$ is the average of the 2 split jackknife
estimators of the APE that use all the individuals and leave out the first and second halves of the time
periods, and $\widetilde{\delta}_{N/2,T}$ is the average of the 2 split jackknife
estimators of the APE that use all the time periods and leave out half of the individuals.
The analytical corrections are constructed using sample analogs of the expressions in Theorems (ref) and (ref), replacing the true values of $\beta$ and $\phi$ by the fixed effects estimators. To describe these corrections, we introduce some additional notation. For any function of the data, unobserved effects and parameters $g_{itj}(\beta,\alpha_i + \gamma_t,\alpha_i + \gamma_{t-j})$ with $0 \leq j < t$, let $\widehat g_{itj} = g_{it}(\widehat \beta, \widehat \alpha_i + \widehat \gamma_t, \widehat \alpha_i + \widehat \gamma_{t-j})$ denote the fixed effects estimator,
e.g., $\widehat{\mathbb{E}_{\phi}[\partial_{\pi^2} \ell_{it}]}$ denotes the fixed effects estimator of $\mathbb{E}_{\phi}[\partial_{\pi^2} \ell_{it}].$
Let $\widehat{\cal H}^{-1}_{(\alpha\alpha)}$, $\widehat{\cal H}^{-1}_{(\alpha\gamma)}$,
$\widehat{\cal H}^{-1}_{(\gamma\alpha)}$,
and $\widehat{\cal H}^{-1}_{(\gamma\gamma)}$ denote the
blocks of the matrix $\widehat{\cal H}^{-1}$, where
$$
\widehat{\cal H} =
\left(
array[array omitted — 194 chars of source]
\right)
+ \frac{b} {\sqrt{NT}} \, vv' ,$$
$\widehat{\mathcal{H}}_{(\alpha\alpha)}^* = diag( - \sum_{t} \widehat{\mathbb{E}_{\phi}[\partial_{\pi^2} \ell_{it}]})/\sqrt{NT}$, $\widehat{\mathcal{H}}_{(\alpha\alpha)} ^*= diag( - \sum_{i} \widehat{\mathbb{E}_{\phi}[\partial_{\pi^2} \ell_{it}]})/\sqrt{NT}$, and $\widehat{\mathcal{H}}_{(\alpha\gamma)it}^* =
- \widehat{\mathbb{E}_{\phi}[\partial_{\pi^2} \ell_{it}]}/\sqrt{NT}$.
Let
align*[align* omitted — 389 chars of source]
The $k$-th component of $\widehat \Xi_{it}$ corresponds to a least squares regression of
$ \widehat{\mathbb{E}_{\phi} \left( \partial_{\beta_k \pi} \ell_{it} \right)}/\widehat{\mathbb{E}_{\phi}( \partial_{\pi^2} \ell_{it} )}$ on the space spanned by the incidental parameters weighted by $ \widehat{\mathbb{E}_{\phi}( - \partial_{\pi^2} \ell_{it})}.$
The analytical bias corrected estimator of $\beta^0$ is
equation[equation omitted — 135 chars of source]
where $\widehat B_{NT}^{\beta} = \widehat W^{-1} \widehat B$, $\widehat D_{NT}^{\beta} = \widehat W^{-1} \widehat D$,
align[align omitted — 1,248 chars of source]
and $L$ is a trimming parameter for estimation of spectral expectations such that $L \to \infty$ and $L/T \to 0$ HahnKuersteiner2011. Here we use truncation instead of kernel smoothing in the estimation of spectral expectations following Hahn and Kuersteiner HK2007. Note that, unlike for variance estimation, a kernel is not needed to ensure that the bias estimator be positive. Instead of choosing a value of $L$, our recommendation for practice is to conduct a sensitivity analysis by reporting estimates for multiple values of $L$ starting from $L=1$.
From our experience based on extensive Monte Carlo simulations, we do not recommend values of $L$ greater than $4$, because the finite-sample dispersion of the estimator quickly increases with $L$. We refer to Section (ref) for an example of sensitivity analysis with respect to $L$. The factor $T/(T-j)$ is a degrees of freedom adjustment that rescales the time series averages $T^{-1} \sum_{t=j+1}^T$ by the number of observations instead of by $T$. Similar corrections for conditional mean models can be formed using the sample analogs of the expressions of $\overline{B}_{\infty}$ and $\overline{D}_{\infty}$ in Remark (ref). We do not spell out these estimators for the sake of brevity.
Asymptotic $(1-p)$--confidence intervals for the components of $\beta^0$ can be formed as
$$
\widetilde \beta_k^A \pm z_{1-p} \sqrt{\widehat W_{kk}^{-1} / (NT)}, \ \ k = \{1, ..., \dim \beta^0\},
$$
where $z_{1-p}$ is the $(1-p)$--quantile of the standard normal distribution, and $\widehat W_{kk}^{-1} $ is the $(k,k)$-element of the
matrix $\widehat W^{-1}$. In conditional moment models we replace $\widehat W_{kk}$ by the $(k,k)$-element of the matrix $\widehat W^{-1} \widehat \Omega \widehat W^{-1}$, where
$$
\widehat \Omega = \frac 1 {NT} \sum_{i=1}^N
\sum_{t=1}^T \sum_{\tau=1}^T
\widehat{\mathbb{E}_{\phi} \left[
D_{\beta} \ell_{it} (D_{\beta} \ell_{i\tau})' \right]}.
$$
We have implemented the analytical correction at the level of the estimator. Alternatively, we can implement the correction at the level of the score or first order conditions by solving
equation[equation omitted — 129 chars of source]
for $\beta$.
Global concavity of the objective function guarantees that the solution to (ref) is unique. Other possible extensions such as continuously updated score corrections where $\overline B_{\infty}$ and $\overline D_{\infty}$ are estimated together with $\beta$, corrections at the level of the objective function, or iterative corrections are left to future research.
The analytical bias corrected estimator of $\delta^0_{NT}$ is
$$
\widetilde \delta^A = \widehat \delta - \widehat B^\delta/T - \widehat D^\delta/N,
$$
where $\widetilde \delta$ is the APE constructed from a bias corrected estimator of $\beta$. Let
align*[align* omitted — 367 chars of source]
The fixed effects estimators of the components of the asymptotic bias are
align*[align* omitted — 936 chars of source]
The estimator of the asymptotic variance depends on the sampling properties of the unobserved effects. Under the independence assumption of Remark (ref) with all the components of $X_{it}$ strictly exogenous,
equation[equation omitted — 396 chars of source]
where $ \widehat{\tilde \Delta}_{it} = \widehat \Delta_{it} - N^{-1} \sum_{i=1}^N \widehat \Delta_{it}$ under identical distribution over $i$, $ \widehat{\tilde \Delta}_{it} = \widehat \Delta_{it} - T^{-1} \sum_{t=1}^T \widehat \Delta_{it}$ under stationarity over $t$, and $ \widehat{\tilde \Delta}_{it} = \widehat \Delta_{it} - \widehat \delta$ under both.
Note that we do not need to specify the convergence rate $r_{NT}$ to make inference because the standard errors $\sqrt{\widehat{V}^{\delta}}/r_{NT}$ do not depend on $r_{NT}$. Bias corrected estimators and confidence intervals can be constructed in the same fashion as for the model parameter.
We use the following homogeneity assumption to show the validity of the jackknife corrections for the model parameters and APEs.
It implies that $ \widetilde{\beta}_{N,T/2} - \widehat \beta_{NT} = \overline B_{\infty}^{\beta}/T + o_P(T^{-1} \vee N^{-1})$ and $\widetilde{\beta}_{N/2,T} - \widehat \beta_{NT} = \overline D_{\infty}^{\beta}/N + o_P(T^{-1} \vee N^{-1})$, which are weaker but higher level sufficient conditions for the validity of the jackknife for the model parameter. For APEs, Assumption (ref) also ensures that these effects do not change with $T$ and $N$, i.e. $\delta_{NT}^0 = \delta^0$. The analytical corrections
{\it do not} require this assumption.
assumption[Unconditional homogeneity] The sequence $\{(Y_{it}, X_{it}, \alpha_i, \gamma_t): 1 \leq i \leq N, 1 \leq t \leq T\}$ is identically distributed across $i$ and strictly stationary
across $t,$ for each $N,T.$
This assumption might seem restrictive for dynamic models where $X_{it}$ includes lags of the dependent variable because in this case it restricts the unconditional distribution of the initial conditions of $Y_{it}$. Note, however, that Assumption (ref) allows the initial conditions to depend on the unobserved effects. In other words, it does not impose that the initial conditions are generated from the stationary distribution of $Y_{it}$ conditional on $X_{it}$ and $\phi$. Assumption (ref) rules out time trends and structural breaks in the processes for the unobserved effects and observed variables.
remark[Test of homogeneity] Assumption (ref) is a sufficient condition for the validity of the jackknife corrections. It has the testable implications that the probability limits of the fixed effects estimator are the same in all the partitions of the panel. For example, it implies that $\beta_{N,T/2}^1 =\beta_{N,T/2}^2$, where $\beta_{N,T/2}^1$ and $\beta_{N,T/2}^2$ are the probability limits of the fixed effects estimators of $\beta$ in the subpanels that include all the individuals and the first and second halves of the time periods, respectively. These implications can be tested using variations of the Chow-type test proposed in Dhaene and Jochmans DhaeneJochmans2015. We provide an example of the application of these tests to our setting in Section (ref) of the supplemental material.
The following theorems are the main result of this section. They show that the analytical and jackknife bias corrections eliminate the bias from the asymptotic
distribution of the fixed effects estimators of the model parameters and APEs without increasing variance, and that the estimators of the asymptotic variances are consistent.
theorem[Bias corrections for $\widehat \beta$] Under the conditions of Theorems (ref),
$$
\widehat W \to_P \overline{W}_{\infty},
$$
and, if $L \to \infty$ and $L/T \to 0,$
$$
\sqrt{NT}(\widetilde \beta^A - \beta^0) \to_d \mathcal{N}(0,
\overline{W}_{\infty}^{-1}).
$$
Under the conditions of Theorems (ref) and Assumption (ref),
$$
\sqrt{NT}(\widetilde \beta^J - \beta^0) \to_d \mathcal{N}(0,
\overline{W}_{\infty}^{-1}).
$$
theorem[Bias corrections for $\widehat \delta$] Under the conditions of Theorems (ref) and (ref),
$$
\widehat V^{\delta} \to_P \overline{V}^{\delta}_{\infty},
$$
and, if $L \to \infty$ and $L/T \to 0,$
$$
r_{NT}(\widetilde \delta^A - \delta_{NT}^0) \to_d \mathcal{N}(0,
\overline V_{\infty}^{\delta}).
$$
Under the conditions of Theorems (ref) and (ref), and Assumption (ref),
$$
r_{NT}(\widetilde \delta^{J} - \delta^0) \to_d \mathcal{N}(0,
\overline V_{\infty}^{\delta}).
$$
remark[Rate of convergence] The rate of convergence $r_{NT}$ depends on the properties of the sampling process for the explanatory variables and unobserved effects (see remark (ref)).
Monte Carlo Experiments
This section reports evidence on the finite sample behavior
of fixed effects estimators of model parameters and APEs in static models with strictly exogenous regressors and dynamic models with predetermined regressors such as lags of the dependent variable. We analyze
the performance of uncorrected and bias-corrected
fixed effects estimators in terms of bias and inference accuracy of
their asymptotic distribution.
In particular we compute the biases, standard deviations, and root mean squared errors of the estimators, the ratio of
average standard errors to the simulation standard deviations (SE/SD); and the empirical coverages of confidence intervals with 95% nominal value (p; .95).\footnote{The standard errors are computed using the expressions ((ref)) and ((ref)) with $ \widehat{\tilde \Delta}_{it} = \widehat \Delta_{it} - \widehat \delta$, evaluated at uncorrected estimates of the parameters. We find little difference in performance of constructing standard errors based on corrected estimates.} Overall, we find that the analytically corrected estimators dominate the uncorrected and jackknife corrected estimators.\footnote{Kristensen and Salani{\'e} KS2013 also found that analytical corrections dominate jackknife corrections to reduce the bias of approximate estimators.} A possible explanation for the better finite-sample performance of the analytical over the jackknife corrections is that the jackknife increases dispersion because the components of the bias are estimated from subsamples that include half of the observations of the panel. We observe this variance increase in all our numerical examples, specially in short panels.
The jackknife corrections are also more sensitive than the analytical corrections to Assumption (ref). All the results are based on 500 replications.
The designs correspond to static and dynamic probit models.
As in the analytical example of Section (ref), we find that our large $T$ asymptotic approximations capture well the behavior of the fixed effects estimator and the bias corrections in moderately long panels with $N=56$ and $T=14$.
Static probit model
The data generating process is
equation*[equation* omitted — 145 chars of source]
where $\alpha_{i} \sim \mathcal{N}(0,1/16)$, $\gamma_{t} \sim
\mathcal{N}(0,1/16)$, $\varepsilon_{it} \sim \mathcal{N}(0,1)$, and
$\beta = 1$. We consider two alternative designs for $X_{it}$:
autoregressive process and linear trend process both with individual and time effects. In
the first design, $X_{it} = X_{i,t-1} / 2 + \alpha_{i} +
\gamma_{t} + \upsilon_{it}$, $\upsilon_{it} \sim
\mathcal{N}(0,1/2)$, and $X_{i0} \sim \mathcal{N}(0,1)$. In the
second design, $X_{it} = 2 t / T + \alpha_{i} +
\gamma_{t} + \upsilon_{it}$,
$\upsilon_{it} \sim \mathcal{N}(0,3/4)$, which violates Assumption (ref). In both designs $X_{it}$ is strictly exogenous with respect to $\varepsilon_{it}$ conditional on the individual and time effects. The variables $\alpha_i$, $\gamma_t$,
$\varepsilon_{it}$, $\upsilon_{it}$, and $X_{i0}$ are independent
and $i.i.d.$ across individuals and time periods. We generate
panel data sets with $N=56$ individuals and three different numbers
of time periods $T$: 14, 28 and 56.\footnote{Following a suggestion from an anonymous referee, we obtained results for panel data sets with $T=56$ and $N$ in $\{14, 28, 56\}$. These results are similar to the results reported and are available from the authors upon request.}
Table 3 reports the results for the probit coefficient $\beta$, and the APE of $X_{it}$. We compute the APE using ((ref)).
Throughout the table, MLE-FETE corresponds to the probit maximum
likelihood estimator with individual and time fixed effects, Analytical is the bias corrected estimator that uses the analytical correction,
and Jackknife is the bias corrected estimator that
uses SPJ in both the individual and time
dimensions. The cross-sectional division in the jackknife follows the
order of the observations. All the results are
reported in percentage of the true parameter value.
We find that the bias is of the same order of magnitude as the standard deviation for the uncorrected estimator of the probit coefficient causing severe undercoverage of the confidence intervals. This result holds for both designs and all the sample sizes considered. The bias corrections, specially Analytical, remove the bias without increasing dispersion, and produce substantial improvements in rmse and coverage probabilities. For example, Analytical reduces rmse by 50% and increases coverage by 26% in the first design with $T=14$. As in Hahn and Newey Hahn:2004p882 and Fernandez-Val FernandezVal:2009p3313, we find very little bias in the uncorrected estimates of the APE, despite the large bias in the probit coefficients. Jackknife performs relatively worse in the second design that does not satisfy Assumption (ref).
Dynamic probit model
The data generating process is
eqnarray*[eqnarray* omitted — 274 chars of source]
where $\alpha_{i} \sim \mathcal{N}(0,1/16)$, $\gamma_{t} \sim
\mathcal{N}(0,1/16)$, $\varepsilon_{it} \sim \mathcal{N}(0,1)$,
$\beta_Y = 0.5$, and $\beta_Z = 1$. We consider two alternative designs for $Z_{it}$:
autoregressive process and linear trend process both with individual and time effects. In
the first design, $Z_{it} = Z_{i,t-1} / 2 + \alpha_{i} +
\gamma_{t} + \upsilon_{it}$, $\upsilon_{it} \sim
\mathcal{N}(0,1/2)$, and $Z_{i0} \sim \mathcal{N}(0,1)$. In the
second design, $Z_{it} = 1.5 t / T + \alpha_{i} +
\gamma_{t} + \upsilon_{it}$,
$\upsilon_{it} \sim \mathcal{N}(0,3/4)$, which violates Assumption (ref).
The variables $\alpha_i$, $\gamma_t$,
$\varepsilon_{it}$, $\upsilon_{it}$, and $Z_{i0}$ are independent
and $i.i.d.$ across individuals and time periods. We generate
panel data sets with $N=56$ individuals and three different numbers
of time periods $T$: 14, 28 and 56.
Table 4 reports the simulation results for the probit coefficient $\beta_Y$ and the APE of $Y_{i,t-1}$. We compute the partial effect of $Y_{i,t-1}$ using the
expression in equation ((ref)) with $X_{it,k} = Y_{i,t-1}$.
This effect is commonly reported as a measure of state dependence for dynamic binary processes. Table 5 reports the simulation results for the estimators of the probit coefficient $\beta_Z$ and the APE of $Z_{it}$. We compute the partial effect using ((ref)) with $X_{it,k} = Z_{it}$.
Throughout the tables, we compare the same estimators as for the static model. For the analytical correction we consider two versions, Analytical (L=1) sets the trimming parameter to estimate spectral expectations $L$ to one, whereas Analytical (L=2) sets $L $ to two.\footnote{In results not reported for brevity, we find little difference in performance of increasing the trimming parameters to $L=3$ and $L=4$. These results are available from the authors upon request.}
Again, all the results in the tables are
reported in percentage of the true parameter value.
The results in table 4 show important biases toward zero for both the probit coefficient and the APE of $Y_{i,t-1}$ in the two designs. This bias can indeed be substantially larger than the corresponding standard deviation for short panels yielding coverage probabilities below 70% for $T=14$. The analytical corrections significantly reduce biases and rmse, bring coverage probabilities close to their nominal level, and have little sensitivity to the trimming parameter $L$. The jackknife corrections reduce bias but increase dispersion, producing less drastic improvements in rmse and coverage than the analytical corrections. The results for the APE of $Z_{it}$ in table 5 are similar to the static probit model. There are significant bias and undercoverage of confidence intervals for the coefficient $\beta_Z$, which are removed by the corrections, whereas there are little bias and undercoverage in the APE. As in the static model, Jackknife performs relatively worse in the second design.
table[table omitted — 90 chars of source]
table[table omitted — 91 chars of source]
table[table omitted — 91 chars of source]
Concluding remarks
In this paper we develop analytical and jackknife corrections for fixed effects estimators of model parameters and APEs in semiparametric nonlinear panel models with additive individual and time effects. Our analysis applies to conditional maximum likelihood estimators with concave log-likelihood functions, and therefore covers logit, probit, ordered probit, ordered logit, Poisson, negative binomial, and Tobit estimators, which are the most popular nonlinear estimators in empirical economics.
We are currently developing similar corrections for nonlinear models with interactive individual and time effects (Chen, Fern\'andez-Val, and Weidner CFW2014). Another interesting avenue of future research is to derive higher-order expansions for fixed effects estimators with individual and time effects. These expansions are needed to justify theoretically the validity of alternative corrections based on the leave-one-observation-out panel jackknife method of Hahn and Newey Hahn:2004p882.
appendix\begin{center}
\begin{LARGE}
Appendix
\end{LARGE}
\end{center}
\section{Notation and Choice of Norms}
We write $A'$ for the transpose of a matrix or vector $A$.
We use $\mathbbm{1}_n$ for the $n\times n$ identity matrix,
and $1_n$ for the column vector of length $n$ whose entries are all
unity.
For
square $n \times n$ matrices $B$, $C$, we use $B>C$ (or $B\geq C$) to indicate
that $B-C$ is positive (semi) definite.
We write wpa1 for “with probability approaching one” and wrt for “with respect to”. All the limits are taken as
$N,T \to \infty$ jointly.
As in the main text, we usually suppress the dependence on $NT$
of all the sequences of functions and parameters to lighten the notation, e.g. we write ${\cal L}$ for ${\cal L}_{NT}$
and $\phi$ for $\phi_{NT}$.
Let
\begin{align*}
{\cal S}(\beta,\phi) &= \partial_{\phi} {\cal L}(\beta,\,\phi), &
{\cal H}(\beta,\phi) &= - \partial_{\phi\phi'} {\cal L}(\beta,\, \phi),
\end{align*}
where $\partial_x f$ denotes the partial derivative of $f$ with respect to $x$, and additional subscripts denote higher-order partial derivatives.
We refer to the $\dim \phi$-vector ${\cal S}(\beta,\phi)$
as the incidental parameter score, and to
the $\dim \phi \times \dim \phi$ matrix ${\cal H}(\beta,\phi)$
as the incidental parameter Hessian.
We omit the
arguments of the functions when they are evaluated at the true
parameter values $(\beta^0, \, \phi^0)$, e.g.
${\cal H}={\cal H}(\beta^0, \phi^0)$.
We use a bar to indicate expectations conditional on $\phi$,
e.g. $\partial_{\beta} \overline {\cal L} =\mathbb{E}_\phi[ \partial_{\beta} {\cal L}]$,
and a tilde to denote variables in deviations with respect to expectations, e.g.
$\partial_{\beta} \widetilde{\cal L} = \partial_{\beta} {\cal L} - \partial_{\beta} \overline {\cal L}$.
We use the Euclidian norm $\|.\|$ for vectors of dimension $\dim \beta$, and we use the norm induced by the Euclidian norm for the corresponding matrices and tensors,
which we also denote by $\|.\|$. For matrices of dimension $\dim \beta \times \dim \beta$ this induced
norm is the spectral norm. The generalization of the spectral norm to higher order tensors
is straightforward, e.g. the induced norm of the $\dim \beta \times \dim \beta \times \dim \beta$ tensor
of third partial derivatives of ${\cal L}(\beta,\phi)$ wrt $\beta$ is given by
\begin{align*}
\left\| \partial_{\beta \beta \beta} {\cal L}(\beta,\phi) \right\|
&= \max_{\left\{ u,v \in \mathbb{R}^{\dim \beta}, \, \|u\|=1, \, \|v\|=1
\right\}}
\left\| \sum_{k,l=1}^{\dim \beta}
u_k \, v_l \,
\partial_{\beta \beta_k \beta_l} {\cal L}(\beta,\phi)
\right\| .
\end{align*}
This choice of norm is immaterial for the asymptotic analysis because $\dim \beta$ is fixed with the sample size.
In contrast, it is important what norms we choose for vectors of dimension $\dim \phi$, and their corresponding matrices and tensors, because $\dim \phi$ is increasing with the sample size. For vectors of dimension $\dim \phi$,
we use the $\ell_q$-norm
\begin{align*}
\| \phi \|_q = \left( \sum_{g=1}^{\dim \phi} | \phi_g |^q \right)^{1/q} ,
\end{align*}
where $2 \leq q \leq \infty$.\footnote{We use the letter $q$ instead of $p$ to avoid confusion with the use of $p$ for probability.}
The particular value $q=8$ will be chosen later.\footnote{The main reason not to choose $q= \infty$ is the assumption
$\| \widetilde {\cal H} \|_q = o_P(1)$ below, which is used to guarantee that
$\| {\cal H}^{-1} \|_q$ is of the same order as $\| \overline {\cal H}^{-1} \|_q$.
If we assume $\| {\cal H}^{-1} \|_q = {\cal O}_P(1)$ directly instead of $\| {\overline{\cal H}}^{-1} \|_q = {\cal O}_P(1)$, then we
can set $q=\infty$.} We use the norms that are induced by the $\ell_q$-norm for the corresponding matrices and tensors, e.g. the induced $q$-norm of the $\dim \phi \times \dim \phi \times \dim \phi$ tensor
of third partial derivatives of ${\cal L}(\beta,\phi)$ wrt $\phi$ is
\begin{align}
\left\| \partial_{\phi \phi \phi} {\cal L}(\beta,\phi) \right\|_q
&= \max_{\left\{ u,v \in \mathbb{R}^{\dim \phi}, \, \|u\|_q=1, \, \|v\|_q=1
\right\}}
\left\| \sum_{g,h=1}^{\dim \phi}
u_g \, v_h \,
\partial_{\phi \phi_g \phi_h} {\cal L}(\beta,\phi)
\right\|_q .
\end{align}
Note that in general the ordering of the indices of the tensor would matter in the definition of this norm,
with the first index having a special role. However, since partial derivatives like
$\partial_{\phi_g \phi_h \phi_l} {\cal L}(\beta,\phi)$ are fully symmetric in the indices $g$, $h$, $l$, the ordering is
not important in their case.
For mixed partial derivatives of ${\cal L}(\beta,\phi)$ wrt $\beta$ and $\phi$, we use the norm that is induced
by the Euclidian norm on $\dim \beta$-vectors and the $q$-norm on $\dim \phi$-indices,
e.g.
\begin{align}
\left\| \partial_{\beta \beta \phi \phi \phi} {\cal L}(\beta,\phi) \right\|_q
&=
\max_{\left\{ u,v \in \mathbb{R}^{\dim \beta}, \, \|u\|=1, \, \|v\|=1
\right\}}
\max_{\left\{ w,x \in \mathbb{R}^{\dim \phi}, \, \|w\|_q=1, \, \|x\|_q=1
\right\}}
\nonumber \\ & \qquad \qquad \qquad
\left\| \sum_{k,l=1}^{\dim \beta} \sum_{g,h=1}^{\dim \phi}
u_k \, v_l \, w_g \, x_h \,
\partial_{\beta_k \beta_l \phi \phi_g \phi_h} {\cal L}(\beta,\phi)
\right\|_q ,
\end{align}
where we continue to use the notation $\|.\|_q$, even though this is a mixed norm.
Note that for $ w,x \in \mathbb{R}^{\dim \phi}$ and $q \geq 2$,
\begin{align*}
|w' x| \leq \| w \|_q \|x\|_{q/(q-1)} \leq (\dim \phi)^{(q-2)/q} \| w \|_q \| x \|_q.
\end{align*}
Thus, whenever we bound a scalar product of vectors, matrices and tensors in terms of the above
norms we have to account for this additional factor
$(\dim \phi)^{(q-2)/q}$.
For example,
\begin{align*}
& \left| \sum_{k,l=1}^{\dim \beta} \sum_{f,g,h=1}^{\dim \phi}
u_k \, v_l \, w_f \, x_h \, y_f \,
\partial_{\beta_k \beta_l \phi_f \phi_g \phi_h} {\cal L}(\beta,\phi)
\right|
\leq
(\dim \phi)^{(q-2)/q}
\| u \| \,
\| v \| \,
\| w \|_q \,
\| x \|_q \,
\| y \|_q \,
\left\| \partial_{ \beta \beta \phi \phi \phi} {\cal L}(\beta,\phi) \right\|_q .
\end{align*}
For higher-order tensors, we use the notation
$ \partial_{\phi \phi \phi} {\cal L}(\beta,\phi)$ inside the $q$-norm $\|.\|_q$ defined above, while we rely on
standard index and matrix notation for all other expressions involving those partial derivatives, e.g.
$\partial_{\phi \phi' \phi_g} {\cal L}(\beta,\phi)$ is a $\dim \phi \times \dim \phi$ matrix for every
$g=1,\ldots, \dim \phi$.
Occasionally, e.g. in Assumption (ref)$(vi)$ below, we use the
Euclidian norm for $\dim \phi$-vectors, and the spectral
norm for $\dim \phi \times \dim \phi$-matrices, denoted by $\|.\|$,
and defined as $\|.\|_q$ with $q=2$. Moreover, we employ the
matrix infinity norm $\left\| A \right\|_\infty
= \max_i \sum_j |A_{ij}|$, and the matrix maximum norm
$\left\| A \right\|_{\max} = \max_{ij} |A_{ij}|$ to characterize the properties of the inverse of the expected Hessian of the incidental parameters in Section (ref).
For $r \geq 0$, we define the sets
${\cal B}(r,\beta^0)=
\left\{ \beta: \|\beta-\beta^0\| \leq r \right\}$,
and
${\cal B}_q(r, \phi^0)=
\left\{ \phi :
\|\phi-\phi^0\|_q \leq r \right\}$,
which are closed balls of radius $r$ around the true parameter values
$\beta^0$ and $\phi^0$, respectively.
\section{Asymptotic Expansions}
In this section, we derive asymptotic expansions for the score of
the profile objective function, ${\cal L}(\beta,\widehat \phi(\beta)),$ and for the fixed effects estimators
of the parameters and APEs, $\widehat \beta$ and $\widehat \delta$. We do not employ the panel
structure of the model, nor the particular form of the objective
function given in Section (ref). Instead, we consider the
estimation of an unspecified model based on a sample of size $NT$
and a generic objective function ${\cal L}(\beta,\phi)$, which
depends on the parameter of interest $\beta$ and the incidental
parameter $\phi$. The estimators $\widehat \phi(\beta)$ and $\widehat \beta$ are
defined in (ref) and (ref). The proof of all the results in this Section are given in the supplementary material.
We make the following
high-level assumptions. These assumptions might appear somewhat
abstract, but will be justified by more primitive conditions in
the context of panel models.
\begin{assumption}[Regularity conditions for asymptotic expansion of $\widehat \beta$]
Let $q>4$ and $0 \leq \epsilon < 1/8 - 1/(2q)$.
Let $r_\beta = r_{\beta,NT} >0$,
$r_\phi = r_{\phi,NT}>0$,
with $r_\beta = o\left[ (NT)^{-1/(2q)-\epsilon} \right]$
and $r_\phi = o\left[ (NT)^{ -\epsilon} \right]$.
We assume that
\begin{itemize}
• $\frac{\dim \phi} {\sqrt{NT}} \rightarrow a$, $0<a<\infty$.
• $(\beta,\phi) \mapsto {\cal L}(\beta,\, \phi)$ is four times continuously
differentiable in ${\cal B}(r_\beta, \beta^0) \times {\cal B}_q(r_\phi, \phi^0)$, wpa1.
• $\displaystyle \sup_{\beta \in {\cal B}(r_\beta ,\beta^0)}
\left\| \widehat \phi(\beta) - \phi^0 \right\|_q
= o_P(r_\phi)$.
• $\overline {\cal H} > 0$, and
$\left\| \overline {\cal H}^{-1} \right\|_q
= {\cal O}_P\left( 1\right)$.
•
For the $q$-norm defined in Appendix (ref),
\begin{align*}
\| {\cal S} \|_q &= {\cal O}_P \left( (NT)^{-1/4 + 1/(2q)} \right) ,
&
\| \partial_{\beta} {\cal L} \| &= {\cal O}_P(1) ,
&
\| \widetilde {\cal H} \|_q &= o_P(1) ,
\\
\left\| \partial_{\beta \phi'} {\cal L} \right\|_q &= {\cal O}_P \left( (NT)^{1/(2q)} \right) ,
&
\left\| \partial_{\beta \beta'} {\cal L} \right\| &= {\cal O}_P(\sqrt{NT}) ,
&
\left\| \partial_{\beta \phi \phi} {\cal L} \right\|_q &= {\cal O}_P( (NT)^{\epsilon} ) ,
\\
\left\| \partial_{\phi \phi \phi} {\cal L} \right\|_q &= {\cal O}_P\left( (NT)^{\epsilon} \right) ,
\end{align*}
and
\begin{align*}
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\beta \beta \beta} {\cal L}(\beta,\, \phi) \right\|
&= {\cal O}_P\left( \sqrt{NT} \right) ,
\\
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\beta \beta \phi} {\cal L}(\beta,\, \phi) \right\|_q
&= {\cal O}_P\left( (NT)^{1/(2q)} \right) ,
\\
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\beta \beta \phi \phi} {\cal L}(\beta,\, \phi) \right\|_q
&= {\cal O}_P\left( (NT)^{\epsilon} \right) ,
\\
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\beta \phi \phi \phi} {\cal L}(\beta,\, \phi) \right\|_q
&= {\cal O}_P\left( (NT)^{\epsilon} \right) ,
\\
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)} \left\| \partial_{\phi \phi \phi \phi} {\cal L}(\beta,\phi) \right\|_q &= {\cal O}_P\left( (NT)^{\epsilon} \right) .
\end{align*}
• For the spectral norm $\|.\|=\|.\|_2$,
\begin{eqnarray*}
\| \widetilde {\cal H} \| = o_P \left( (NT)^{-1/8} \right) , \ \
\left\| \partial_{\beta \beta'} \widetilde {\cal L} \right\| = o_P( \sqrt{NT} ) , \ \
\left\| \partial_{\beta \phi \phi} \widetilde {\cal L} \right\|
= o_P \left( (NT)^{-1/8} \right) ,\\
\left\| \partial_{\beta \phi'} \widetilde {\cal L} \right\|
= {\cal O}_P \left( 1 \right) , \ \
\left\| \sum_{g,h=1}^{\dim \phi}
\partial_{\phi \phi_g \phi_h} \widetilde {\cal L} \,
[ \overline {\cal H}^{-1} {\cal S} ]_g
[ \overline {\cal H}^{-1} {\cal S} ]_h
\right\|
= o_P \left( (NT)^{-1/4} \right) .
\end{eqnarray*}
\end{itemize}
\end{assumption}
Let
$\partial_{\beta} \mathcal{L}(\beta,\widehat \phi(\beta))$ be the score of the profile objective function.\footnote{Note that $\frac{d} {d \beta} \mathcal{L}(\beta,\widehat \phi(\beta)) = \partial_{\beta} \mathcal{L}(\beta,\widehat \phi(\beta))$ by the envelope theorem.}
The following theorem is the main result of this appendix.
\begin{theorem}[Asymptotic expansions of $\widehat \phi(\beta)$ and $\partial_{\beta} \mathcal{L}(\beta,\widehat \phi(\beta))$]
Let Assumption (ref) hold. Then
\begin{align*}
\widehat \phi(\beta) - \phi^0 &= {\cal H}^{-1} {\cal S}
+ {\cal H}^{-1}
[\partial_{\phi \beta'} {\cal L}] (\beta-\beta^0)
+ \ft 1 2 {\cal H}^{-1} \sum_{g=1}^{\dim \phi}
[\partial_{\phi \phi' \phi_g} {\cal L} ] {\cal H}^{-1} {\cal S}
[ {\cal H}^{-1} {\cal S} ]_g
+ R^{\phi}(\beta) ,
\end{align*}
and
\begin{align*}
\partial_{\beta} {\cal L}(\beta, \widehat \phi(\beta))
&= U
- \overline W \, \sqrt{NT} (\beta-\beta^0)
+ R(\beta) ,
\end{align*}
where $U= U^{(0)}
+ U^{(1)}$, and
\begin{align*}
\overline W &= - \, \frac 1 {\sqrt{NT}} \,
\left( \partial_{\beta \beta'} \overline {\cal L}
+ [\partial_{\beta \phi'} \overline {\cal L}] \; \overline {\cal H}^{-1} \;
[\partial_{\phi \beta'} \overline {\cal L}] \right) ,
\nonumber \\
U^{(0)} &=
\partial_{\beta} {\cal L}
+ [\partial_{\beta \phi'} \overline {\cal L} ]\, \overline {\cal H}^{-1} {\cal S} ,
\nonumber \\
U^{(1)} &=
[\partial_{\beta \phi'} \widetilde{\cal L}] \overline {\cal H}^{-1} {\cal S}
- [\partial_{\beta \phi'} \overline {\cal L}] \,
\overline {\cal H}^{-1} \, \widetilde {\cal H} \,
\overline {\cal H}^{-1} \, {\cal S}
+
\frac 1 2 \, \sum_{g=1}^{\dim \phi}
\left( \partial_{\beta \phi' \phi_g} \overline {\cal L}
+ [\partial_{\beta \phi'} \overline {\cal L}] \, \overline {\cal H}^{-1}
[\partial_{\phi \phi' \phi_g} \overline {\cal L}] \right)
[\overline {\cal H}^{-1} {\cal S}]_g
\overline {\cal H}^{-1} {\cal S} .
\end{align*}
The remainder terms of the expansions satisfy
\begin{align*}
\sup_{\beta \in {\cal B}(r_\beta ,\beta^0)}
\frac{ (NT)^{1/2-1/(2q)} \, \left\| R^\phi(\beta) \right\|_q}
{1 + \sqrt{NT} \|\beta-\beta^0\|}
&= o_P \left( 1 \right) \; ,
&
\sup_{\beta \in {\cal B}(r_\beta ,\beta^0)}
\frac{\| R(\beta) \|} {1 + \sqrt{NT} \|\beta-\beta^0\|} &= o_P(1) \; .
\end{align*}
\end{theorem}
\begin{remark}
The result for $\widehat \phi(\beta) - \phi^0$ does not rely on Assumption (ref)$(vi)$. Without this assumption we can also show that
\begin{align*}
\partial_\beta {\cal L}(\beta,\widehat \phi(\beta))
&= \partial_\beta {\cal L} +
\left[ \partial_{\beta \beta'} {\cal L}
+ ( \partial_{\beta \phi'} {\cal L} ) {\cal H}^{-1}
( \partial_{\phi' \beta} {\cal L} ) \right] (\beta-\beta^0)
+ (\partial_{\beta \phi'} {\cal L}) {\cal H}^{-1} {\cal S}
\nonumber \\ & \quad
+ \frac 1 2 \sum_g \left( \partial_{\beta \phi' \phi_g} {\cal L}
+ [\partial_{\beta \phi'} {\cal L}] \, {\cal H}^{-1}
[\partial_{\phi \phi' \phi_g} {\cal L}] \right)
[ {\cal H}^{-1} {\cal S}]_g
{\cal H}^{-1} {\cal S}
+ R_1 (\beta) ,
\end{align*}
with $R_1 (\beta)$ satisfying the same bound as $R (\beta)$.
Thus, the spectral norm bounds in Assumption (ref)$(vi)$ for $\dim \phi$-vectors, matrices and tensors
are only used after separating expectations from deviations of expectations for certain partial derivatives.
Otherwise, the derivation of the bounds is purely based on the $q$-norm for
$\dim \phi$-vectors, matrices and tensors.
\end{remark}
The proofs are given in Section (ref) of the supplementary material.
Theorem (ref) characterizes asymptotic expansions
for the incidental parameter estimator
and the score of the profile objective function in the incidental parameter score ${\cal S}$ up to quadratic order.
The theorem provides bounds on the
the remainder terms $R^{\phi}(\beta)$ and $R(\beta)$,
which make the expansions applicable to estimators of
$\beta$ that take values within a shrinking $r_{\beta}$-neighborhood of $\beta^0$ wpa1.
Given such an $r_{\beta}$-consistent estimator $\widehat \beta$ that solves the first order condition $\partial_{\beta} {\cal L}(\beta, \widehat \phi(\beta)) = 0$, we can use the expansion
of the profile objective score to obtain an asymptotic expansion for $\widehat \beta$. This gives rise
to the following corollary of Theorem (ref) . Let $\overline W_{\infty} := \lim_{N,T \to \infty} \overline W$.
\begin{corollary}[Asymptotic expansion of $\widehat \beta$]
Let Assumption (ref) be satisfied. In addition, let $U = {\cal O}_P(1)$, let
$\overline W_{\infty}$ exist with $\overline W_{\infty} > 0$,
and let
$\| \widehat \beta - \beta^0\| = o_P(r_\beta)$.
Then
$$\sqrt{NT} (\widehat \beta - \beta^0) = \overline W_{\infty}^{-1} U+ o_P(1).$$
\end{corollary}
The following theorem states that for strictly concave objective functions no separate consistency proof
is required for $\widehat \phi(\beta)$ and for $\widehat \beta$.
\begin{theorem}[Consistency under Concavity]
Let Assumption (ref)$(i)$, $(ii)$, $(iv)$, $(v)$ and $(vi)$ hold, and let $(\beta,\phi) \mapsto {\cal L}(\beta,\phi)$ be strictly
concave over $(\beta,\phi) \in \mathbb{R}^{\dim \beta + \dim \phi}$, wpa1.
Assume furthermore that
$(NT)^{-1/4+1/(2q)} = o_P(r_\phi)$
and $(NT)^{1/(2q)} r_\beta = o_P(r_\phi)$.
Then,
$$\displaystyle \sup_{\beta \in {\cal B}(r_\beta ,\beta^0)}
\left\| \widehat \phi(\beta) - \phi^0 \right\|_q
= o_P(r_\phi),$$
i.e. Assumption (ref)$(iii)$ is satisfied.
If, in addition,
$\overline W_{\infty}$ exists with $\overline W_{\infty} > 0$,
then $\| \widehat \beta - \beta^0\| = {\cal O}_P\left( (NT)^{-1/4} \right)$.
\end{theorem}
In the application of Theorem (ref) to panel models, we focus on estimators with strictly concave objective functions. By Theorem (ref),
we only need to check Assumption (ref)$(i)$, $(ii)$, $(iv)$, $(v)$ and $(vi)$,
as well as $U={\cal O}_P(1)$ and $\overline W_{\infty} > 0$, when we apply
Corollary (ref) to derive the limiting distribution of $\widehat \beta$.
We give the proofs of Corollary (ref) and Theorem (ref)
in Section (ref).
\subsubsection*{Expansion for Average Effects}
We invoke the following
high-level assumption, which is verified under more primitive
conditions for panel data models in the next section.
\begin{assumption}[Regularity conditions for asymptotic expansion of $\widehat \delta$]
Let $q$, $\epsilon$, $r_\beta$ and $r_\phi$
be defined as in Assumption (ref).
We assume that
\begin{itemize}
• $(\beta,\phi) \mapsto \Delta(\beta,\, \phi)$ is three times continuously
differentiable in ${\cal B}(r_\beta, \beta^0) \times {\cal B}_q(r_\phi, \phi^0)$, wpa1.
• $\left\| \partial_{\beta} \Delta \right\| = {\cal O}_P(1),$
$\left\| \partial_{\phi} \Delta \right\|_q =
{\cal O}_P \left( (NT)^{1/(2q)-1/2} \right),$ $\left\| \partial_{\phi \phi} \Delta \right\|_q = {\cal O}_P( (NT)^{\epsilon-1/2} ),$
and
\begin{align*}
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\beta \beta} \Delta(\beta,\, \phi) \right\|
&= {\cal O}_P\left( 1 \right) ,
\\
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\beta \phi'} \Delta(\beta,\, \phi) \right\|_q
&= {\cal O}_P\left( (NT)^{1/(2q)-1/2} \right) ,
\\
\sup_{\beta \in {\cal B}(r_\beta, \beta^0)} \sup_{\phi \in {\cal B}_q(r_\phi, \phi^0)}
\left\| \partial_{\phi \phi \phi} \Delta(\beta,\, \phi) \right\|_q
&= {\cal O}_P\left( (NT)^{\epsilon-1/2} \right) .
\end{align*}
• $\left\| \partial_{\beta} \widetilde \Delta \right\| = o_P( 1 ),$
$\left\| \partial_{\phi} \widetilde \Delta \right\|
= {\cal O}_P \left( (NT)^{-1/2} \right),$ and $\left\| \partial_{\phi \phi} \widetilde \Delta \right\|
= o_P \left( (NT)^{-5/8} \right).$
\end{itemize}
\end{assumption}
The following result gives the asymptotic expansion for the estimator,
$\widehat \delta = \Delta(\beta,\widehat \phi(\beta)),$
wrt $\delta = \Delta(\beta^0,\phi^0)$.
\begin{theorem}[Asymptotic expansion of $\hat \delta$]
Let Assumptions (ref) and (ref) hold
and let
$\| \widehat \beta - \beta^0 \| = {\cal O}_P\left( (NT)^{-1/2} \right)
= o_P\left( r_\beta \right)$.
Then
\begin{align*}
\widehat \delta - \delta
&= \left[ \partial_{\beta'} \overline {\Delta}
+
(\partial_{\phi'} \overline \Delta)
\overline {\cal H}^{-1} (\partial_{\phi \beta'} \overline {\cal L})
\right] (\widehat \beta - \beta^0)
+ U^{(0)}_\Delta
+ U^{(1)}_\Delta + o_P\left( 1/ \sqrt{NT} \right),
\end{align*}
where
\begin{align*}
U^{(0)}_\Delta &= (\partial_{\phi'} \overline \Delta)
\overline {\cal H}^{-1} {\cal S} ,
\\
U^{(1)}_\Delta
&= (\partial_{\phi'} \widetilde \Delta) \overline {\cal H}^{-1} {\cal S}
- (\partial_{\phi'} \overline \Delta)
\overline {\cal H}^{-1} \widetilde {\cal H}
\overline {\cal H}^{-1} {\cal S}
\nonumber \\ & \quad
+ \ft 1 2 \, {\cal S}' \overline {\cal H}^{-1}
\left[\partial_{\phi \phi'} \overline \Delta +
\sum_{g=1}^{\dim \phi}
\left[\partial_{\phi \phi' \phi_g} \overline {\cal L} \right]
\left[ \overline {\cal H}^{-1}
(\partial_{\phi} \overline \Delta) \right]_g \right]
\overline {\cal H}^{-1} {\cal S}.
\end{align*}
\end{theorem}
\begin{remark}
The expansion of the profile score
$\partial_{\beta_k} {\cal L}(\beta,\widehat \phi(\beta))$ in
Theorem (ref)
is a special case of the expansion in
Theorem (ref), for
$\Delta(\beta,\phi) = \frac 1 {\sqrt{NT}} \partial_{\beta_k} {\cal L}(\beta,\phi)$.
Assumptions (ref) also exactly match with the corresponding
subset of Assumption (ref).
\end{remark}
\section{Proofs of Section (ref)}
\subsection{Application of General Expansion to Panel Estimators}
We now apply the
general expansion of appendix (ref) to the panel fixed effects estimators considered in the main text.
For the objective function specified in (ref)
and (ref), the
incidental parameter score evaluated at the true parameter value is
\begin{align*}
{\cal S} &= \left( \begin{array}{c}
\left[ \frac 1 {\sqrt{NT}} \sum_{t=1}^T \, \partial_{\pi} \ell_{it} \right]_{i=1,\ldots,N} \\
\left[ \frac 1 {\sqrt{NT}} \sum_{i=1}^N \, \partial_{\pi} \ell_{it} \right]_{t=1,\ldots,T}
\end{array}
\right) .
\end{align*}
The penalty term in the objective function does not
contribute to ${\cal S}$, because at the true parameter value
$v'\phi^0 = 0$. The corresponding expected incidental parameter Hessian $\overline {\cal H}$
is given in (ref). Section (ref) discusses
the structure of $\overline {\cal H}$ and $\overline {\cal H}^{-1}$ in more detail.
Define
\begin{align}
\Lambda_{it} &:= - \frac 1 {\sqrt{NT}} \sum_{j=1}^N \sum_{\tau=1}^T \left(
\overline {\cal H}^{-1}_{(\alpha\alpha)ij}
+ \overline {\cal H}^{-1}_{(\gamma\alpha)tj}
+ \overline {\cal H}^{-1}_{(\alpha\gamma)i\tau}
+ \overline {\cal H}^{-1}_{(\gamma\gamma)t\tau} \right) \partial_{\pi} \ell_{j\tau},
\end{align}
and the operator $D_{\beta} \Delta_{it} := \partial_{\beta} \Delta_{it} - \partial_{\pi} \Delta_{it} \Xi_{it}$, which are similar to $\Xi_{it}$ and $D_{\beta} \ell_{it}$ in equation (ref).
The following theorem shows that Assumption (ref) and Assumption (ref) for the panel model are sufficient for Assumption (ref) and Assumption (ref) for the general expansion, and particularizes the terms of the expansion to the panel estimators. The proof is given in the supplementary material.
\begin{theorem}
Consider an estimator with objective function given by (ref)
and (ref).
Let Assumption (ref) be satisfied
and suppose that the limit $\overline W_\infty$ defined in
Theorem (ref) exists and is positive definite.
Let $q=8$, $\epsilon=1/(16+2 \nu)$,
$r_{\beta,NT} = \log(NT) (NT)^{-1/8}$
and
$r_{\phi,NT} = (NT)^{-1/16}$. Then,
\begin{itemize}
• Assumption (ref)
holds and
$\| \widehat \beta - \beta^0\| = {\cal O}_P( (NT)^{-1/4} )$.
• The approximate Hessian and the terms of the score defined in Theorem (ref)
can be written as
\begin{align*}
\overline W &= - \frac 1 {NT} \sum_{i=1}^N
\sum_{t=1}^T \mathbb{E}_\phi \left(
\partial_{\beta \beta'} \ell_{it}
- \partial_{\pi^2} \ell_{it} \Xi_{it} \Xi'_{it} \right) ,
\nonumber \\
U^{(0)} &= \frac 1 {\sqrt{NT}}
\sum_{i=1}^N \sum_{t=1}^T
D_{\beta} \ell_{it} ,
\nonumber \\
U^{(1)} &= \frac 1 {\sqrt{NT}}
\sum_{i=1}^N \sum_{t=1}^T \,\left\{ - \Lambda_{it} \,
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] +
\frac 1 2 \Lambda_{it}^2 \, \mathbb{E}_\phi(
D_{\beta \pi^2} \ell_{it}) \right\}.
\end{align*}
• In addition, let Assumption (ref) hold.
Then, Assumption (ref)
is satisfied for the partial effects defined in (ref).
By Theorem (ref),
$$\sqrt{NT}
\left( \widehat \delta - \delta \right) = V^{(0)}_\Delta + V^{(1)}_\Delta
+ o_P(1),$$
where
\begin{align*}
V^{(0)}_\Delta &=
\left[ \frac 1 {NT}
\sum_{i,t}
\mathbb{E}_\phi( D_{\beta} {\Delta_{it}} ) \right]'
\overline W_\infty^{-1} U^{(0)}
- \frac 1 {\sqrt{NT}} \sum_{i,t}
\mathbb{E}_\phi( \Psi_{it} )
\partial_{\pi} \ell_{it} ,
\\
V^{(1)}_\Delta &=
\left[ \frac 1 {NT}
\sum_{i,t}
\mathbb{E}_\phi( D_{\beta} {\Delta_{it}} ) \right]'
\overline W_\infty^{-1} U^{(1)}
+ \frac 1 {\sqrt{NT}}\sum_{i,t} \Lambda_{it}
\left[ \mathbb{E}_\phi( \Psi_{it} ) \partial_{\pi^2} \ell_{it}
- \Psi_{it} \mathbb{E}_\phi( \partial_{\pi^2} \ell_{it} )
\right]
\nonumber \\ & \qquad
+
\frac 1 {2 \, \sqrt{NT}} \sum_{i,t} \Lambda_{it}^2 \left[
\mathbb{E}_\phi( \partial_{\pi^2} \Delta_{it} )
-
\mathbb{E}_\phi( \partial_{\pi^3} \ell_{it} )
\mathbb{E}_\phi( \Psi_{it} )
\right].
\end{align*}
\end{itemize}
\end{theorem}
\subsection{Proofs of Theorems (ref) and (ref)}
\begin{proof}[\bf Proof of Theorem (ref)]
\# First, we want to show that $U^{(0)} \to_d {\cal N}( 0 ,\; \overline W_{\infty})$.
In our likelihood setting,
$\mathbb{E}_\phi \partial_{\beta} {\cal L} = 0$, $\mathbb{E}_\phi {\cal S} = 0$,
and,
by the Bartlett identities,
$\mathbb{E}_\phi( \partial_{\beta} {\cal L} \partial_{\beta'} {\cal L} )
= - \frac 1 {\sqrt{NT}} \partial_{\beta \beta'} \overline {\cal L}$,
$\mathbb{E}_\phi( \partial_{\beta} {\cal L} {\cal S}' ) = -
\frac 1 {\sqrt{NT}} \partial_{\beta \phi'} \overline {\cal L}$
and $\mathbb{E}_\phi( {\cal S} {\cal S}' )
= \frac 1 {\sqrt{NT}} \left( \overline {\cal H} - \frac b {\sqrt{NT}} v v' \right)$.
Furthermore, ${\cal S}' v=0$ and
$ \partial_{\beta \phi'} \overline {\cal L} v =0$.
Then, by definition of
$\overline W = - \, \frac 1 {\sqrt{NT}} \,
\left( \partial_{\beta \beta'} \overline {\cal L}
+ [\partial_{\beta \phi'} \overline {\cal L}] \; \overline {\cal H}^{-1} \;
[\partial_{\phi \beta'} \overline {\cal L}] \right)$
and
$U^{(0)} = \partial_{\beta} {\cal L}
+ [\partial_{\beta \phi'} \overline {\cal L} ]\, \overline {\cal H}^{-1} {\cal S},$
\begin{align*}
\mathbb{E}_\phi\left( U^{(0)} \right) &= 0 , &
{\rm Var} \left( U^{(0)} \right) &= \overline W,
\end{align*}
which implies that $\lim_{N,T\rightarrow \infty} {\rm Var} \left( U^{(0)} \right) =
\lim_{N,T\rightarrow \infty} \overline W = \overline W_{\infty}$. Moreover,
part $(ii)$
of Theorem (ref) yields
\begin{align*}
U^{(0)} &= \frac 1 {\sqrt{NT}}
\sum_{i=1}^N \sum_{t=1}^T
D_{\beta} \ell_{it},
\end{align*}
where
$D_{\beta} \ell_{it} = \partial_{\beta} \ell_{it} - \partial_{\pi} \ell_{it} \Xi_{it}$
is a martingale difference sequence for each $i$ and independent across $i$, conditional on $\phi$.
Thus, by Lemma (ref)
and the Cramer-Wold device we conclude that
\begin{align*}
U^{(0)} \to_d {\cal N}\left[ 0 ,\; \lim_{N,T\rightarrow \infty} {\rm Var} \left( U^{(0)} \right)
\right]
\sim {\cal N}( 0 ,\; \overline W_{\infty}) .
\end{align*}
\# Next, we show that $U^{(1)} \to_P
\kappa \overline B_{\infty} + \kappa^{-1} \overline D_{\infty}$. Part $(ii)$
of Theorem (ref) gives $U^{(1)} = U^{(1a)} + U^{(1b)} $, with
\begin{align*}
U^{(1a)} &= - \frac 1 {\sqrt{NT}}
\sum_{i=1}^N \sum_{t=1}^T \, \Lambda_{it} \,
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] ,
\nonumber \\
U^{(1b)} & = \frac 1 {2 \, \sqrt{NT}}
\sum_{i=1}^N \sum_{t=1}^T \Lambda_{it}^2 \, \mathbb{E}_\phi(
D_{\beta \pi^2} \ell_{it}).
\end{align*}
Plugging-in the definition of $\Lambda_{it}$, we decompose
$U^{(1a)} = U^{(1a,1)} + U^{(1a,2)} + U^{(1a,3)} + U^{(1a,4)}$, where
\begin{align*}
U^{(1a,1)} &= \frac 1 {NT}
\sum_{i,j} \overline {\cal H}^{-1}_{(\alpha\alpha)ij}
\left( \sum_{\tau} \partial_{\pi} \ell_{j\tau} \right)
\sum_t
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] ,
\nonumber \\
U^{(1a,2)} &= \frac 1 {NT}
\sum_{j,t} \overline {\cal H}^{-1}_{(\gamma\alpha)tj}
\left( \sum_{\tau} \partial_{\pi} \ell_{j\tau} \right)
\sum_i
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] ,
\nonumber \\
U^{(1a,3)} &= \frac 1 {NT}
\sum_{i,\tau} \overline {\cal H}^{-1}_{(\alpha\gamma)i\tau}
\left( \sum_j \partial_{\pi} \ell_{j\tau} \right)
\sum_t
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] ,
\nonumber \\
U^{(1a,4)} &= \frac 1 {NT}
\sum_{t,\tau} \overline {\cal H}^{-1}_{(\gamma\gamma)t\tau}
\left( \sum_j \partial_{\pi} \ell_{j\tau} \right)
\sum_i \left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] .
\end{align*}
By the Cauchy-Schwarz inequality applied to the sum over $t$ in $U^{(1a,2)},$
\begin{align*}
\left( U^{(1a,2)} \right)^2
&\leq \frac {1} {(NT)^2}
\left[ \sum_{t}
\left( \sum_{j,\tau} \overline {\cal H}^{-1}_{(\gamma\alpha)tj} \partial_{\pi} \ell_{j\tau}
\right)^2 \right]
\left[ \sum_t \left( \sum_i
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right]
\right)^2 \right] .
\end{align*}
By Lemma (ref), $ \overline {\cal H}^{-1}_{(\gamma\alpha)tj} = {\cal O}_P(1/\sqrt{NT})$,
uniformly over $t,j$.
Using that both $\sqrt{NT} \, \overline {\cal H}^{-1}_{(\gamma\alpha)tj} \partial_{\pi} \ell_{j\tau}$
and $D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} )$ are mean zero,
independence across $i$ and Lemma (ref) in the supplementary material
across $t$, we obtain
\begin{align*}
\mathbb{E}_\phi \left(
\frac 1 {\sqrt{NT}} \sum_{j,\tau}
[\sqrt{NT} \, \overline {\cal H}^{-1}_{(\gamma\alpha)tj}] \partial_{\pi} \ell_{j\tau}
\right)^2 = {\cal O}_P(1) ,
\ \ \mathbb{E}_\phi \left( \frac 1 {\sqrt{N}} \sum_i
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right]
\right)^2 = {\cal O}_P(1) ,
\end{align*}
uniformly over $t$. Thus,
$ \sum_{t}
\left( \sum_{j,\tau} \overline {\cal H}^{-1}_{(\gamma\alpha)tj} \partial_{\pi} \ell_{j\tau}
\right)^2 = {\cal O}_P(T)$
and
$ \sum_t \left( \sum_i
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right]
\right)^2 = {\cal O}_P(NT)$. We conclude that
\begin{align*}
\left( U^{(1a,2)} \right)^2
&= \frac {1} {(NT)^2} {\cal O}_P(T)
{\cal O}_P(NT)
= {\cal O}_P(1/N) = o_P(1),
\end{align*}
and therefore that $U^{(1a,2)} = o_P(1)$. Analogously one can show that
$U^{(1a,3)} = o_P(1)$.
By Lemma (ref),
$\overline {\cal H}^{-1}_{(\alpha\alpha)} =
- {\rm diag}
\left[ \left( \frac 1 {\sqrt{NT}} \sum_{t=1}^T \, \mathbb{E}_\phi( \partial_{\pi^2} \ell_{it} \right)^{-1}
\right]
+ {\cal O}_P(1/\sqrt{NT})$. Analogously to the proof of $U^{(1a,2)} = o_P(1),$
one can show that the ${\cal O}_P(1/\sqrt{NT})$ part of
$\overline {\cal H}^{-1}_{(\alpha\alpha)}$ has an
asymptotically negligible contribution to $U^{(1a,1)}$.
Thus,
\begin{align*}
U^{(1a,1)} &= - \frac 1 {\sqrt{NT}}
\sum_{i}
\underbrace{
\frac{
\left( \sum_{\tau} \partial_{\pi} \ell_{i\tau} \right)
\sum_t
\left[ D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} ) \right] }
{ \sum_t \, \mathbb{E}_\phi( \partial_{\pi^2} \ell_{it} )}
}_{=: U^{(1a,1)}_i}
+o_P(1) .
\end{align*}
Our assumptions guarantee that
$\mathbb{E}_\phi\left[ \left(U^{(1a,1)}_i \right)^2 \right] = {\cal O}_P(1)$,
uniformly over $i$.
Note that both the denominator and the numerator
of $U^{(1a,1)}_i$ are of order $T$. For the denominator this is obvious because of
the sum over $T$. For the numerator there are two sums over $T$, but both
$\partial_{\pi} \ell_{i\tau}$ and $D_{\beta \pi} \ell_{it} - \mathbb{E}_\phi( D_{\beta \pi} \ell_{it} )$
are mean zero weakly correlated processes, so that their sums are of order
$\sqrt{T}$. By the WLLN over $i$
(remember that we have cross-sectional independence, conditional on $\phi$, and we
assume finite moments),
$N^{-1} \sum_i U^{(1a,1)}_i = N^{-1} \sum_i \mathbb{E}_\phi U^{(1a,1)}_i + o_P(1)$,
and therefore
\begin{align*}
U^{(1a,1)} &= \underbrace{ - \sqrt{ \frac N T } \frac 1 N
\sum_{i=1}^{N}
\frac{ \sum_{t=1}^T \sum_{\tau=t}^T
\mathbb{E}_\phi\left(
\partial_{\pi} \ell_{it} D_{\beta \pi} \ell_{i\tau}
\right)
}
{ \sum_{t=1}^T \mathbb{E}_\phi\left( \partial_{\pi^2} \ell_{i t} \right) }
}_{ =: \sqrt{ \frac N T } \overline B^{(1)}}
+o_P(1) .
\end{align*}
Here, we use that
$\mathbb{E}_\phi\left(
\partial_{\pi} \ell_{it} D_{\beta \pi} \ell_{i\tau}
\right)=0$ for $t>\tau$.
Analogously,
\begin{align*}
U^{(1a,4)} &= - \underbrace{ \sqrt{ \frac T N } \frac 1 T
\sum_{t=1}^{T}
\frac{ \sum_{i=1}^N
\mathbb{E}_\phi\left(
\partial_{\pi} \ell_{it} D_{\beta \pi} \ell_{it}
\right) }
{ \sum_{i=1}^N \mathbb{E}_\phi\left( \partial_{\pi^2} \ell_{i t} \right) }
}_{=: \sqrt{ \frac T N } \overline D^{(1)}}
+o_P(1) .
\end{align*}
We conclude that
$U^{(1a)} = \kappa \overline B^{(1)} + \kappa^{-1} \overline D^{(1)} + o_P(1)$.
Next, we analyze $U^{(1b)}$.
We decompose $\Lambda_{it}=\Lambda_{it}^{(1)}+\Lambda_{it}^{(2)}
+ \Lambda_{it}^{(3)} + \Lambda_{it}^{(4)}$, where
\begin{align*}
\Lambda_{it}^{(1)} &= - \frac 1 {\sqrt{NT}} \sum_{j=1}^N
\overline {\cal H}^{-1}_{(\alpha\alpha)ij}
\sum_{\tau=1}^T \partial_{\pi} \ell_{j\tau},
&
\Lambda_{it}^{(2)} &= - \frac 1 {\sqrt{NT}} \sum_{j=1}^N
\overline {\cal H}^{-1}_{(\gamma\alpha)tj}
\sum_{\tau=1}^T \partial_{\pi} \ell_{j\tau},
\nonumber \\
\Lambda_{it}^{(3)} &= - \frac 1 {\sqrt{NT}} \sum_{\tau=1}^T
\overline {\cal H}^{-1}_{(\alpha\gamma)i\tau}
\sum_{\tau=1}^T \partial_{\pi} \ell_{j\tau},
&
\Lambda_{it}^{(4)} &= - \frac 1 {\sqrt{NT}} \sum_{\tau=1}^T
\overline {\cal H}^{-1}_{(\gamma\gamma)t\tau} \sum_{\tau=1}^T \partial_{\pi} \ell_{j\tau}.
\end{align*}
This decomposition of $\Lambda_{it}$ induces the following decomposition
of $U^{(1b)}$
\begin{align*}
U^{(1b)} &= \sum_{p,q=1}^4 U^{(1b,p,q)} , &
U^{(1b,p,q)} &= \frac 1 {2 \, \sqrt{NT}}
\sum_{i=1}^N \sum_{t=1}^T
\Lambda_{it}^{(p)} \Lambda_{it}^{(q)}
\mathbb{E}_\phi(
D_{\beta \pi^2} \ell_{it}).
\end{align*}
Due to the symmetry $U^{(1b,p,q)} = U^{(1b,q,p)},$ this decomposition
has 10 distinct terms. Start with $U^{(1b,1,2)}$ noting that
\begin{align*}
U^{(1b,1,2)} &= \frac 1 {\sqrt{NT}}
\sum_{i=1}^N U^{(1b,1,2)}_i ,
\nonumber \\
U^{(1b,1,2)}_i &= \frac 1 {2T} \sum_{t=1}^T
\mathbb{E}_\phi(
D_{\beta \pi^2} \ell_{it})
\frac{1} {N^2}
\sum_{j_1,j_2=1}^N
\left[ NT \overline {\cal H}^{-1}_{(\alpha\alpha)ij_1}
\overline {\cal H}^{-1}_{(\gamma\alpha)tj_2} \right]
\left( \frac 1 {\sqrt{T}} \sum_{\tau=1}^T \partial_{\pi} \ell_{j_1\tau} \right)
\left( \frac 1 {\sqrt{T}} \sum_{\tau=1}^T \partial_{\pi} \ell_{j_2\tau} \right) .
\end{align*}
By $\mathbb{E}_\phi( \partial_{\pi} \ell_{it} ) =0$,
$\mathbb{E}_\phi( \partial_{\pi} \ell_{it} \partial_{\pi} \ell_{j\tau}) =0$ for $(i,t) \neq (j,\tau)$,
and the properties of the inverse expected Hessian from Lemma (ref),
$\mathbb{E}_\phi\left[ U^{(1b,1,2)}_i \right] = {\cal O}_P(1/N)$, uniformly over $i$,
$\mathbb{E}_\phi\left[ \left(U^{(1b,1,2)}_i \right)^2 \right] = {\cal O}_P(1)$,
uniformly over $i$,
and $\mathbb{E}_\phi\left[ U^{(1b,1,2)}_i U^{(1b,1,2)}_j \right] = {\cal O}_P(1/N)$,
uniformly over $i \neq j$.
This implies that $\mathbb{E}_\phi \, U^{(1b,1,2)} = {\cal O}_P(1/N)$ and
$\mathbb{E}_\phi\left[ \left(U^{(1b,1,2)} - \mathbb{E}_\phi \, U^{(1b,1,2)} \right)^2 \right]
= {\cal O}_P(1/\sqrt{N})$,
and therefore $U^{(1b,1,2)} = o_P(1)$.
By similar arguments one obtains
$U^{(1b,p,q)} = o_P(1)$ for all combinations of $p,q=1,2,3,4$, except
for $p=q=1$ and $p=q=4$.
For $p=q=1$,
\begin{align*}
U^{(1b,1,1)} &= \frac 1 {\sqrt{NT}}
\sum_{i=1}^N U^{(1b,1,1)}_i ,
\nonumber \\
U^{(1b,1,1)}_i &= \frac 1 {2T} \sum_{t=1}^T
\mathbb{E}_\phi(
D_{\beta \pi^2} \ell_{it})
\frac{1} {N^2}
\sum_{j_1,j_2=1}^N
\left[ NT \overline {\cal H}^{-1}_{(\alpha\alpha)ij_1}
\overline {\cal H}^{-1}_{(\alpha\alpha)ij_2} \right]
\left( \frac 1 {\sqrt{T}} \sum_{\tau=1}^T \partial_{\pi} \ell_{j_1\tau} \right)
\left( \frac 1 {\sqrt{T}} \sum_{\tau=1}^T \partial_{\pi} \ell_{j_2\tau} \right) .
\end{align*}
Analogous to the result for $U^{(1b,1,2)},$
$\mathbb{E}_\phi\left[ \left(U^{(1b,1,1)} - \mathbb{E}_\phi \, U^{(1b,1,1)} \right)^2 \right]
= {\cal O}_P(1/\sqrt{N})$, and therefore
$U^{(1b,1,1)} = \mathbb{E}_\phi \, U^{(1b,1,1)} + o(1)$.
Furthermore,
\begin{align*}
\mathbb{E}_\phi \, U^{(1b,1,1)}
&= \frac 1 {2 \sqrt{NT}}
\sum_{i=1}^N \frac{ \sum_{t=1}^T
\mathbb{E}_\phi( D_{\beta \pi^2} \ell_{it})
\sum_{\tau=1}^T \mathbb{E}_\phi\left[ \left( \partial_{\pi} \ell_{i\tau} \right)^2 \right] }
{ \left[ \sum_{t=1}^T \mathbb{E}_\phi\left( \partial_{\pi^2} \ell_{it} \right) \right]^2}
+ o(1)
\nonumber \\ &
= \underbrace{ - \sqrt{ \frac N T } \frac 1 {2 N}
\sum_{i=1}^N \frac{ \sum_{t=1}^T
\mathbb{E}_\phi( D_{\beta \pi^2} \ell_{it}) }
{ \sum_{t=1}^T \mathbb{E}_\phi\left( \partial_{\pi^2} \ell_{it} \right) }
}_ { =: \sqrt{ \frac N T } \overline B^{(2)}}
+ o(1) .
\end{align*}
Analogously,
\begin{align*}
U^{(1b,4,4)}
&= \mathbb{E}_\phi \, U^{(1b,4,4)} + o_P(1)
= \underbrace{ - \sqrt{ \frac T N } \frac 1 {2 T}
\sum_{t=1}^T \frac{ \sum_{i=1}^N
\mathbb{E}_\phi( D_{\beta \pi^2} \ell_{it}) }
{ \sum_{i=1}^N \mathbb{E}_\phi\left( \partial_{\pi^2} \ell_{it} \right) }
}_ { =: \sqrt{ \frac T N } \overline D^{(2)}}
+ o(1) .
\end{align*}
We have thus shown that
$U^{(1b)} = \kappa \overline B^{(2)} + \kappa^{-1} \overline D^{(2)} + o_P(1)$.
Since $\overline B_{\infty} = \lim_{N,T \rightarrow \infty}[ \overline B^{(1)} + \overline B^{(2)}]$
and $\overline D_{\infty} = \lim_{N,T \rightarrow \infty}[ \overline D^{(1)} + \overline D^{(2)}]$
we thus conclude $U^{(1)} =
\kappa \overline B_{\infty} + \kappa^{-1} \overline D_{\infty} + o_P(1)$.
\# We have shown
$U^{(0)} \to_d {\cal N}( 0 ,\; \overline W_{\infty})$, and
$U^{(1)} \to_P
\kappa \overline B_{\infty} + \kappa^{-1} \overline D_{\infty}$. Then, part $(ii)$ of Theorem (ref) yields
$\sqrt{NT} ( \widehat \beta - \beta^0 )
\; \to_d \;
\overline{W}_{\infty}^{-1} {\cal N}( \kappa \overline B_{\infty}
+ \kappa^{-1} \overline D_{\infty} ,
\;\overline W_{\infty})$.
\end{proof}
\begin{proof}[\bf Proof of Theorem (ref)]
We consider the case of scalar $\Delta_{it}$ to simplify the notation. Decompose
$$
r_{NT} (\widehat \delta - \delta_{NT}^0 - \overline{B}_{\infty}^{\delta}/T - \overline{D}_{\infty}^{\delta}/N) = r_{NT} (\delta - \delta_{NT}^0) + \frac{r_{NT}}{\sqrt{NT}} \sqrt{NT} (\widehat \delta - \delta - \overline{B}_{\infty}^{\delta} /T - \overline{D}_{\infty}^{\delta}/N).
$$
\# Part (1): Limit of $\sqrt{NT} (\widehat \delta - \delta - \overline{B}_{\infty}^{\delta}/T - \overline{D}_{\infty}^{\delta}/N)$. An argument analogous to to the proof of Theorem (ref) using Theorem (ref)$(iii)$ yields
$$
\sqrt{NT} (\widehat \delta - \delta) \to_d \mathcal{N}\left(\kappa \overline{B}_{\infty}^{\delta} + \kappa^{-1} \overline{D}_{\infty}^{\delta}, \overline{V}_{\infty}^{\delta(1)} \right),
$$
where $\overline{V}_{\infty}^{\delta(1)} = \overline{\mathbb{E}}\left\{ (NT)^{-1} \sum_{i,t}\mathbb{E}_{\phi}[ \Gamma_{it}^2] \right\},$ for the expressions of $ \overline{B}_{\infty}^{\delta}$, $ \overline{D}_{\infty}^{\delta}$, and $\Gamma_{it}$ given in the statement of the theorem. Then, by Mann-Wald theorem
\begin{equation*}
\sqrt{NT} (\widehat \delta - \delta - \overline{B}_{\infty}^{\delta}/T - \overline{D}_{\infty}^{\delta}/N) \to_d \mathcal{N}\left(0, \overline{V}_{\infty}^{\delta(1)} \right).
\end{equation*}
\# Part (2): Limit of $r_{NT}(\delta - \delta_{NT}^0)$. Here we show that $r_{NT}(\delta - \delta_{NT}^0) \to_d \mathcal{N}(0,\overline{V}^{\delta(2)}_{\infty})$
for the convergence rate $r_{NT}$ given in Remark (ref), and characterize the asymptotic variance $\overline{V}^{\delta(2)}_{\infty}$.
We determine $r_{NT}$ through $\mathbb{E}[(\delta - \delta_{NT}^0)^2] = \mathcal{O}(r_{NT}^{-2})$ and $r_{NT}^{-2} = \mathcal{O}(\mathbb{E}[(\delta - \delta_{NT}^0)^2])$, where
\begin{equation}
\mathbb{E}[(\delta - \delta_{NT}^0)^2] = \mathbb{E} \left[
\left(\frac {1} {NT} \sum_{i,t} \widetilde \Delta_{it} \right)^2 \right] =
\frac {1} {N^2T^2} \sum_{i,j,t,s}\mathbb{E} \left[ \widetilde \Delta_{it} \widetilde \Delta_{js} \right],
\end{equation}
for $\widetilde \Delta_{it} = \Delta_{it} - \mathbb{E} ( \Delta_{it} )$. Then, we characterize $\overline{V}^{\delta(2)}_{\infty}$ as $\overline{V}_{\infty}^{\delta(2)} = \overline{\mathbb{E}}\{r_{NT}^2 \mathbb{E}[(\delta - \delta_{NT}^0)^2]\},$ because $\mathbb{E}[\delta - \delta_{NT}^0] = 0$. The order of $\mathbb{E}[(\delta - \delta_{NT}^0)^2]$ is equal to the number of terms of the sums in equation (ref) that are non zero, which it is determined by the sample properties of $\{(X_{it}, \alpha_i, \gamma_t) : 1 \leq i \leq N, 1 \leq t \leq T) \}$.
Under Assumption (ref)$(i)$, if $\{\alpha_i\}_N$ and $\{\gamma_t\}_T$ are independent sequences, and $\alpha_i$ and $\gamma_t$ are independent for all $i,t$, then $\mathbb{E} [ \widetilde \Delta_{it} \widetilde \Delta_{js} ] = \mathbb{E} [ \widetilde \Delta_{it} ] \mathbb{E} [\widetilde \Delta_{js} ] = 0$ if $i \neq j$ and $t \neq s,$ so that
$$
\mathbb{E}[(\delta - \delta_{NT}^0)^2] = \frac {1} {N^2T^2} \left\{ \sum_{i,t,s}\mathbb{E} \left[ \widetilde \Delta_{it} \widetilde \Delta_{is} \right] + \sum_{i,j,t}\mathbb{E} \left[ \widetilde \Delta_{it} \widetilde \Delta_{jt} \right] - \sum_{i,t} \mathbb{E} \left[ \widetilde \Delta_{it}^2 \right] \right\}= \mathcal{O}\left(\frac{N+T-1}{NT} \right),
$$
because $\mathbb{E} [ \widetilde \Delta_{it} \widetilde \Delta_{is} ] \leq \mathbb{E} [ \mathbb{E}_{\phi} (\widetilde \Delta_{it}^2)]^{1/2} \mathbb{E}[ \mathbb{E}_{\phi}( \widetilde \Delta_{is}^2) ]^{1/2} < C$ by the Cauchy-Schwarz inequality and Assumption (ref)$(ii)$. We conclude that $r_{NT} = \sqrt{NT/(N+T-1)}$ and
$$
\overline{V}^{\delta(2)} = \overline{\mathbb{E}} \left\{\frac {r_{NT}^2} {N^2T^2} \left( \sum_{i,t,s}\mathbb{E} \left[ \widetilde \Delta_{it} \widetilde \Delta_{is} \right] + \sum_{i\neq j,t}\mathbb{E} \left[ \widetilde \Delta_{it} \widetilde \Delta_{jt} \right] \right) \right\}.
$$
Note that $r_{NT} \to \infty$ and $r_{NT} = \mathcal{O}(\sqrt{NT}).$
\# Part (3): Asymptotic covariance between $r_{NT}(\delta - \delta_{NT}^0)$ and $\sqrt{NT} (\widehat \delta - \delta - T^{-1} \overline{B}_{\infty}^{\delta} - N^{-1} \overline{D}_{\infty}^{\delta})$. Note that
$$
\mathbb{E}\left[(\delta - \delta_{NT}^0) \frac{1}{NT} \sum_{i,t} \Gamma_{it} \right] = \frac{1}{N^2T^2} \sum_{i,s>t}\mathbb{E} \left[ \widetilde \Delta_{it} \Gamma_{is} \right] = \mathcal{O}\left(\frac{1}{N} \right)
$$
since $\Gamma_{it}$ is a martingale difference over $t$ and independent over $i$ conditional on the unobserved effects. Let
$$
\overline{C}^{\delta(1,2)} = \overline{\mathbb{E}} \left\{\frac {1} {NT^2} \sum_{i,s>t}\mathbb{E} \left[ \widetilde \Delta_{it} \Gamma_{is} \right] \right\}.
$$
\# Part (4): limit of $r_{NT} (\widehat \delta - \delta_{NT}^0 - T^{-1} \overline{B}_{\infty}^{\delta} - N^{-1} \overline{D}_{\infty}^{\delta})$. The conclusion of the Theorem follows because $\overline{V}_{\infty}^{\delta} = \overline{V}^{\delta(2)} + \overline{V}^{\delta(1)} \lim_{N,T \to \infty} (r_{NT}/\sqrt{NT})^2 + 2 \overline{C}^{\delta(1,2)} \lim_{N,T \to \infty} (r_{NT}^2/N).$
\end{proof}
\section{Properties of the Inverse Expected Incidental Parameter Hessian}
The expected incidental parameter Hessian evaluated at the true parameter values is
\begin{align*}
\overline{\cal H} = \mathbb{E}_{\phi}[ - \partial_{\phi \phi'} {\cal L}] =
\left(\begin{array}{cc} \overline{\mathcal{H}}_{(\alpha\alpha)}^* & \overline{\mathcal{H}}_{(\alpha\gamma)}^* \\ {[\overline{\mathcal{H}}_{(\alpha\gamma)}^*]}' & \overline{\mathcal{H}}_{(\gamma\gamma)}^*
\end{array}\right)
+ \frac{b} {\sqrt{NT}} \, vv' ,
\end{align*}
where $v= v_{NT} = ( 1_N' ,- 1_T')'$,
$\overline{\mathcal{H}}_{(\alpha\alpha)}^* = \text{diag}(\frac 1 {\sqrt{NT}} \sum_{t} \mathbb{E}_{\phi}[- \partial_{\pi^2} \ell_{it}])$,
$\overline{\mathcal{H}}_{(\alpha\gamma)it}^* = \frac 1 {\sqrt{NT}} \mathbb{E}_{\phi}[-\partial_{\pi^2} \ell_{it}]$,
and
$\overline{\mathcal{H}}_{(\gamma\gamma)} ^*= \text{diag}(\frac 1 {\sqrt{NT}} \sum_{i} \mathbb{E}_{\phi}[- \partial_{\pi^2} \ell_{it}])$.
In panel models with only individual effects, it is straightforward to determine the order of magnitude of $\overline {\cal H}^{-1}$ in Assumption (ref)$(iv)$, because $\overline {\cal H}$ contains only the diagonal matrix $ \overline {\cal H}_{(\alpha \alpha)}^*$.
In our case,
$\overline {\cal H}$ is no longer diagonal, but it has a special structure. The diagonal terms
are of order 1, whereas the off-diagonal terms are of order $(NT)^{-1/2}$. Moreover,
$\left\| \overline {\cal H} - {\rm diag}( \overline {\cal H}_{(\alpha \alpha)}^*, \overline {\cal H}_{(\gamma \gamma)}^*) \right\|_{\max}
= \mathcal{O}_P((NT)^{-1/2})$.
These observations, however, are not sufficient to establish the order of $\overline {\cal H}^{-1}$ because the number of non-zero off-diagonal terms is of much larger order than the number of diagonal
terms; compare $\mathcal{O}(NT)$ to $\mathcal{O}(N+T)$.
Note also that the expected
Hessian without penalty term $\overline{\cal H}^*$ has the same structure
as $\overline{\cal H}$ itself, but is not even invertible, i.e. the observation on the relative
size of diagonal vs. off-diagonal terms is certainly not sufficient to make statements about the
structure of $\overline {\cal H}^{-1}$.
The result of the following lemma is therefore not obvious.
It shows
that the diagonal terms of $\overline{{\cal H}}$ also dominate in determining
the order of $\overline{{\cal H}}^{-1}$.
\begin{lemma}
Under Assumptions (ref),
\begin{equation*}
\left\| \overline {\cal H}^{-1} -
{\rm diag} \left( \overline {\cal H}_{(\alpha \alpha)}^*, \overline {\cal H}_{(\gamma \gamma)}^*
\right)^{-1}
\right\|_{\max}
= \mathcal{O}_P\left( (NT)^{-1/2} \right) .
\end{equation*}
\end{lemma}
The proof of Lemma (ref) is provided in the supplementary material.
The lemma result establishes that $\overline {\cal H}^{-1}$ can be uniformly approximated
by a diagonal matrix,
which is given by
the inverse of the diagonal terms of $\overline {\cal H}$ without the penalty.
The diagonal elements of
$ {\rm diag}( \overline {\cal H}_{(\alpha \alpha)}^*, \overline {\cal H}_{(\gamma \gamma)}^*)^{-1} $
are of order 1, i.e. the order of the difference established
by the lemma is relatively small.
Note that the choice of penalty in the objective function
is important to obtain Lemma (ref).
Different penalties, corresponding to other
normalizations (e.g. a penalty proportional to $\alpha_1^2$, corresponding
to the normalization $\alpha_1=0$), would fail to deliver
Lemma (ref). However, these alternative choices do not affect
the estimators $\widehat \beta$ and $\widehat \delta$, i.e. which normalization
is used to compute $\widehat \beta$ and $\widehat \delta$ in practice is irrelevant
(up to numerical precision errors).
thebibliography\bibitem[\astroncite{Aghion
et al.}{2005}]{AghionBloomBlundellGriffithHowitt2005}
Aghion, P., Bloom, N., Blundell, R., Griffith, R., and Howitt, P. (2005).
\newblock Competition and innovation: an inverted-{U} relationship.
\newblock {\em The Quarterly Journal of Economics}, 120(2):701--728.
\bibitem[\astroncite{Alvarez and Arellano}{2003}]{AlvarezArellano2003}
Alvarez, J. and Arellano, M. (2003).
\newblock The time series and cross-section asymptotics of dynamic panel data
estimators.
\newblock {\em Econometrica}, 71(4):1121--1159.
\bibitem[\astroncite{Arellano and
Bonhomme}{2009}]{ArellanoBonhomme2009}
Arellano, M. and Bonhomme, S. (2009).
\newblock Robust priors in nonlinear panel data models.
\newblock {\em Econometrica}, 77(2):489--536.
\bibitem[\astroncite{Arellano and Hahn}{2007}]{ArellanoHahn2007}
Arellano, M. and Hahn, J. (2007).
\newblock Advances in economics and econometrics. theory and applications.
volume 3.
\newblock In {\em Ninth World Congress, Econometric Society Monographs,
Cambridge University Press, chapter “Understanding Bias in Nonlinear Panel
Models: Some Recent Developments”}, pages 381--409.
\bibitem[\astroncite{Bai}{2009}]{Bai:2009p3321}
Bai, J. (2009).
\newblock Panel data models with interactive fixed effects.
\newblock {\em Econometrica}, 77(4):1229--1279.
\bibitem[\astroncite{Carro}{2007}]{Carro:2007p3601}
Carro, J. (2007).
\newblock Estimating dynamic panel data discrete choice models with fixed
effects.
\newblock {\em Journal of Econometrics}, 140(2):503--528.
\bibitem[\astroncite{Chamberlain}{2010}]{Chamberlain2010}
Chamberlain, G. (2010).
\newblock Binary response models for panel data: Identification and
information.
\newblock {\em Econometrica}, 78(1):159--168.
\bibitem[\astroncite{Charbonneau}{2012}]{Charbonneau2011}
Charbonneau, K. (2012).
\newblock Multiple fixed effects in nonlinear panel data models.
\newblock {\em Unpublished manuscript}.
\bibitem[\astroncite{Charbonneau}{2014}]{Charbonneau2014}
Charbonneau, K. (2014).
\newblock Multiple fixed effects in binary response panel data models.
\newblock {\em Bank of Canada Working Paper 2014-17}.
\bibitem[\astroncite{{Chen} et al.}{2014}]{CFW2014}
{Chen}, M., {Fernandez-Val}, I., and {Weidner}, M. (2014).
\newblock {Nonlinear Panel Models with Interactive Effects}.
\newblock {\em ArXiv e-prints}.
\bibitem[\astroncite{Chernozhukov et al.}{2013}]{CFHN13}
Chernozhukov, V., Fern{\'a}ndez-Val, I., Hahn, J., and Newey, W. (2013).
\newblock Average and quantile effects in nonseparable panel models.
\newblock {\em Econometrica}, 81(2):535--580.
\bibitem[\astroncite{Cox and Kim}{1995}]{CoxKim1995}
Cox, D. D. and Kim, T. Y. (1995).
\newblock Moment bounds for mixing random variables useful in nonparametric
function estimation.
\newblock {\em Stochastic processes and their applications}, 56(1):151--158.
\bibitem[\astroncite{Cruz-Gonz{\'a}lez et al.}{2015}]{Stata2015}
Cruz-Gonz{\'a}lez, M., Fern{\'a}ndez-Val, I., and Weidner, M. (2015).
\newblock probitfe and logitfe: Bias corrections for fixed effects estimators
of probit and logit panel models.
\newblock {\em Unpublished manuscript}.
\bibitem[\astroncite{Dhaene and Jochmans}{2015}]{DhaeneJochmans2015}
Dhaene, G. and Jochmans, K. (2015).
\newblock Split-panel jackknife estimation of fixed-effect models.
\newblock {\em The Review of Economic Studies}, 82(3):991--1030.
\bibitem[\astroncite{Enea}{2012}]{Speedglm2012}
Enea, M. (2012).
\newblock {\em speedglm: Fitting Linear and Generalized Linear Models to large
data sets.}
\newblock R package version 0.1.
\bibitem[\astroncite{Fan and Yao}{2003}]{FanYao2003}
Fan, J. and Yao, Q. (2003).
\newblock Nonlinear time series: nonparametric and parametric methods.
\bibitem[\astroncite{Fern{\'a}ndez-Val}{2009}]{FernandezVal:2009p3313}
Fern{\'a}ndez-Val, I. (2009).
\newblock Fixed effects estimation of structural parameters and marginal
effects in panel probit models.
\newblock {\em Journal of Econometrics}, 150:71--85.
\bibitem[\astroncite{Fern{\'a}ndez-Val and Lee}{2013}]{FL13}
Fern{\'a}ndez-Val, I. and Lee, J. (2013).
\newblock {Panel data models with nonadditive unobserved heterogeneity:
Estimation and inference}.
\newblock {\em Quantitative Economics}, 4(3):453--481.
\bibitem[\astroncite{Fern{\'a}ndez-Val and
Vella}{2011}]{FernandezValVella2011}
Fern{\'a}ndez-Val, I. and Vella, F. (2011).
\newblock Bias corrections for two-step fixed effects panel data estimators.
\newblock {\em Journal of Econometrics}, 163(2):144--162.
\bibitem[\astroncite{Fern{\'a}ndez-Val and
Weidner}{2015a}]{ThisWorkingPaper2015}
Fern{\'a}ndez-Val, I. and Weidner, M. (2015a).
\newblock Individual and time effects in nonlinear panel models with large {N},
{T}.
\newblock {\em cemmap working paper, Centre for Microdata Methods and
Practice}.
\bibitem[\astroncite{Fern{\'a}ndez-Val and Weidner}{2015b}]{Supp2015}
Fern{\'a}ndez-Val, I. and Weidner, M. (2015b).
\newblock Supplement to `{I}ndividual and time effects in nonlinear panel
models with large {N},{T}'.
\newblock {\em Unpublished manuscript}.
\bibitem[\astroncite{Galvao and Kato}{2014}]{GalvaoKato:2013}
Galvao, A. F. and Kato, K. (2014).
\newblock Estimation and inference for linear panel data models under
misspecification when both {N} and {T} are large.
\newblock {\em Journal of Business & Economic Statistics}, 32(2):285--309.
\bibitem[\astroncite{Greene}{2004}]{Greene:2004p3125}
Greene, W. (2004).
\newblock The behavior of the fixed effects estimator in nonlinear models.
\newblock {\em The Econometrics Journal}, 7(1):98--119.
\bibitem[\astroncite{Hahn and Kuersteiner}{2002}]{Hahn:2002p717}
Hahn, J. and Kuersteiner, G. (2002).
\newblock Asymptotically unbiased inference for a dynamic panel model with
fixed effects when both n and {T} are large.
\newblock {\em Econometrica}, 70(4):1639--1657.
\bibitem[\astroncite{Hahn and Kuersteiner}{2007}]{HK2007}
Hahn, J. and Kuersteiner, G. (2007).
\newblock Bandwidth choice for bias estimators in dynamic nonlinear panel
models.
\newblock {\em Working paper}.
\bibitem[\astroncite{Hahn and Kuersteiner}{2011}]{HahnKuersteiner2011}
Hahn, J. and Kuersteiner, G. (2011).
\newblock Bias reduction for dynamic nonlinear panel models with fixed effects.
\newblock {\em Econometric Theory}, 27(06):1152--1191.
\bibitem[\astroncite{Hahn and Moon}{2006}]{HahnMoon2006}
Hahn, J. and Moon, H. (2006).
\newblock Reducing bias of {MLE} in a dynamic panel model.
\newblock {\em Econometric Theory}, 22(03):499--512.
\bibitem[\astroncite{Hahn and Newey}{2004}]{Hahn:2004p882}
Hahn, J. and Newey, W. (2004).
\newblock Jackknife and analytical bias reduction for nonlinear panel models.
\newblock {\em Econometrica}, 72(4):1295--1319.
\bibitem[\astroncite{Heckman}{1981}]{Heckman:1981p2940}
Heckman, J. (1981).
\newblock The incidental parameters problem and the problem of initial
conditions in estimating a discrete time-discrete data stochastic process.
\newblock {\em Structural analysis of discrete data with econometric
applications}, pages 179--195.
\bibitem[\astroncite{Higham}{1992}]{Higham1992}
Higham, N. J. (1992).
\newblock Estimating the matrix p-norm.
\newblock {\em Numerische Mathematik}, 62(1):539--555.
\bibitem[\astroncite{Honor{\'e} and Tamer}{2006}]{HonoreTamer2006}
Honor{\'e}, B. E. and Tamer, E. (2006).
\newblock Bounds on parameters in panel dynamic discrete choice models.
\newblock {\em Econometrica}, pages 611--629.
\bibitem[\astroncite{Horn and Johnson}{1985}]{HornJohnson1985}
Horn, R. A. and Johnson, C. R. (1985).
\newblock {\em Matrix analysis}.
\newblock Cambridge university press.
\bibitem[\astroncite{Hu}{2002}]{Hu2002}
Hu, L. (2002).
\newblock Estimation of a censored dynamic panel data model.
\newblock {\em Econometrica}, 70(6):2499--2517.
\bibitem[\astroncite{Kato et al.}{2012}]{KatoGalvaoMontes-Rojas2012}
Kato, K., Galvao, A., and Montes-Rojas, G. (2012).
\newblock Asymptotics for panel quantile regression models with individual
effects.
\newblock {\em Journal of Econometrics}, 170(1):76--91.
\bibitem[\astroncite{Kristensen and Salani{\'e}}{2013}]{KS2013}
Kristensen, D. and Salani{\'e}, B. (2013).
\newblock Higher-order properties of approximate estimators.
\newblock {\em cemmap working paper, Centre for Microdata Methods and
Practice}.
\bibitem[\astroncite{Lancaster}{2000}]{Lancaster:2000p879}
Lancaster, T. (2000).
\newblock The incidental parameter problem since 1948.
\newblock {\em Journal of Econometrics}, 95(2):391--413.
\bibitem[\astroncite{Lancaster}{2002}]{Lancaster:2002p875}
Lancaster, T. (2002).
\newblock Orthogonal parameters and panel data.
\newblock {\em The Review of Economic Studies}, 69(3):647--666.
\bibitem[\astroncite{McLeish}{1974}]{Mcleish1974}
McLeish, D. (1974).
\newblock Dependent central limit theorems and invariance principles.
\newblock {\em the Annals of Probability}, pages 620--628.
\bibitem[\astroncite{Moon and Weidner}{2015a}]{MoonWeidner2015a}
Moon, H. and Weidner, M. (2015a).
\newblock {Dynamic Linear Panel Regression Models with Interactive Fixed
Effects}.
\newblock {\em forthcoming in Econometric Theory}.
\bibitem[\astroncite{Moon and Weidner}{2015b}]{MoonWeidner2015b}
Moon, H. R. and Weidner, M. (2015b).
\newblock Linear regression for panel with unknown number of factors as
interactive fixed effects.
\newblock {\em Econometrica}, 83(4):1543--1579.
\bibitem[\astroncite{Neyman and Scott}{1948}]{Neyman:1948p881}
Neyman, J. and Scott, E. (1948).
\newblock Consistent estimates based on partially consistent observations.
\newblock {\em Econometrica}, 16(1):1--32.
\bibitem[\astroncite{Okui}{2013}]{Okui2013}
Okui, R. (2013).
\newblock Asymptotically unbiased estimation of autocovariances and
autocorrelations with panel data in the presence of individual and time
effects.
\newblock {\em Journal of Time Series Econometrics}, pages 1--53.
\bibitem[\astroncite{Olsen}{1978}]{Olsen:1978p3375}
Olsen, R. (1978).
\newblock Note on the uniqueness of the maximum likelihood estimator for the
tobit model.
\newblock {\em Econometrica: Journal of the Econometric Society}, pages
1211--1215.
\bibitem[\astroncite{Pesaran}{2006}]{Pesaran2006}
Pesaran, M. H. (2006).
\newblock Estimation and inference in large heterogeneous panels with a
multifactor error structure.
\newblock {\em Econometrica}, 74(4):967--1012.
\bibitem[\astroncite{Phillips and Moon}{1999}]{Phillips:1999p733}
Phillips, P. C. B. and Moon, H. (1999).
\newblock Linear regression limit theory for nonstationary panel data.
\newblock {\em Econometrica}, 67(5):1057--1111.
\bibitem[\astroncite{Pratt}{1981}]{Pratt:1981p654}
Pratt, J. W. (1981).
\newblock Concavity of the log likelihood.
\newblock {\em Journal of the American Statistical Association},
76(373):103--106.
\bibitem[\astroncite{White}{2001}]{White2001}
White, H. (2001).
\newblock {\em Asymptotic theory for econometricians}.
\newblock Academic press New York.
\bibitem[\astroncite{Woutersen}{2002}]{Woutersen:2002p3683}
Woutersen, T. (2002).
\newblock Robustness against incidental parameters.
\newblock {\em Unpublished manuscript}.
\setcounter{section}{0}
\setcounter{footnote}{0}
\setcounter{page}{1}
center[center omitted — 124 chars of source]
\abstract{
This supplemental material contains five appendices. Appendix (ref) presents the results of an empirical application and a Monte Carlo simulation calibrated to the application. Following Aghion et al. AghionBloomBlundellGriffithHowitt2005, we use a panel of U.K. industries to estimate Poisson models with industry and time effects for the relationship between innovation and competition. Appendix (ref) gives the proofs of Theorems (ref) and (ref).
Appendices (ref), (ref), and (ref) contain the proofs of Appendices (ref), (ref), and (ref), respectively. Appendix (ref) collects some useful intermediate results that are used in the proofs of the main results.}
Relationship between Innovation and Competition
Empirical Example
To illustrate the bias corrections with real data, we revisit the empirical application of Aghion, Bloom, Blundell, Griffith and Howitt AghionBloomBlundellGriffithHowitt2005 (ABBGH) that estimated a count data model to analyze the relationship between innovation and competition. They used an unbalanced panel of seventeen U.K. industries followed over the 22 years between 1973 and 1994.\footnote{We assume that the observations are missing at random conditional on the explanatory variables and unobserved effects and apply the corrections without change since the level of attrition is low in this application.} The dependent variable, $Y_{it},$ is innovation as measured by a citation-weighted number of patents, and the explanatory variable of interest, $Z_{it},$ is competition as measured by one minus the Lerner index in the industry-year.
Following ABBGH we consider a quadratic static Poisson model with industry and year effects where
$$
Y_{it} \mid Z_i^T, \alpha_i,\gamma_t \sim \mathcal{P}(\exp[\beta_1 Z_{it} + \beta_2 Z_{it}^2 + \alpha_i + \gamma_t]),
$$
for $(i = 1,...,17; t = 1973, ..., 1994),$ and extend the analysis to a dynamic Poisson model with industry and year effects where
$$
Y_{it} \mid Y_i^{t-1}, Z_i^t, \alpha_i, \gamma^t \sim \mathcal{P}(\exp[\beta_{Y} \log(1 + Y_{i,t-1}) + \beta_1 Z_{it} + \beta_2 Z_{it}^2 + \alpha_i + \gamma_t]),
$$
for $(i = 1,...,17; t = 1974, ..., 1994).$ In the dynamic model we use the year 1973 as the initial condition for $Y_{it}$.
table[table omitted — 89 chars of source]
Table S1 reports the results of the analysis. Columns (2) and (3) for the static model replicate the empirical results of Table I in ABBGH (p. 708), adding estimates of the APEs. Columns (4) and (5) report estimates of the analytical corrections that do not assume that competition is strictly exogenous with $L=1$ and $L=2$, and column (6) reports estimates of the jackknife bias corrections described in equation ((ref)) of the paper. Note that we do not need to report separate standard errors for the corrected estimators, because the standard errors of the uncorrected estimators are consistent for the corrected estimators under the asymptotic approximation that we consider.\footnote{In numerical examples, we find very little gains in terms of the ratio SE/SD and coverage probabilities when we reestimate the standard errors using bias corrected estimates.} Overall, the corrected estimates, while numerically different from the uncorrected estimates in column (3), agree with the inverted-U pattern in the relationship between innovation and competition found by ABBGH. The close similarity between the uncorrected and bias corrected estimates gives some evidence in favor of the strict exogeneity of competition with respect to the innovation process.
\setcounter{table}{1}
table[table omitted — 581 chars of source]
The results for the dynamic model show substantial positive state dependence in the innovation process that is not explained by industry heterogeneity.
Uncorrected fixed effects underestimates the coefficient and APE of lag patents relative to the bias corrections, specially relative to the jackknife. The pattern of the differences between the estimates is consistent with the biases that we find in the numerical example in Table S4. Accounting for state dependence does not change the inverted-U pattern, but flattens the relationship between innovation and competition.
Table (ref) implements Chow-type homogeneity tests for the validity of the jackknife corrections. These tests compare the uncorrected fixed effects estimators of the common parameters within the elements of the cross section and time series partitions of the panel. Under time homogeneity, the probability limit of these estimators is the same, so that a standard Wald test can be applied based on the difference of the estimators in the sub panels within the partition. For the static model, the test is rejected at the 1% level in both the cross section and time series partitions. Since the cross sectional partition is arbitrary, these rejection might be a signal of model misspecification. For the dynamic model, the test is rejected at the 1% level in the time series partition, but it cannot be rejected at conventional levels in the cross section partition. The rejection of the time homogeneity might explain the difference between the jackknife and analytical corrections in the dynamic model.
Calibrated Monte Carlo Simulations
We conduct a simulation that mimics the empirical example. The designs correspond to static and dynamic Poisson models with additive individual and time effects. We calibrate all the parameters and exogenous variables using the dataset from ABBGH.
Static Poisson model
The data generating process is
equation*[equation* omitted — 166 chars of source]
where $\mathcal{P}$ denotes the Poisson distribution. The variable $Z_{it}$ is fixed to the values of
the competition variable in the dataset and all the parameters are set to the fixed effect estimates of the model. We generate
unbalanced panel data sets with $T=22$ years and three different numbers
of industries $N$: 17, 34, and 51. In the second (third) case, we double (triple) the cross-sectional size by merging two (three) independent realizations of the panel.
Table S3 reports the simulation results for the coefficients $\beta_1$ and $\beta_2$, and the APE of $Z_{it}$. We compute the APE using the expression ((ref)) with $H(Z_{it}) = Z_{it}^2$.
Throughout the table, MLE corresponds to the pooled Poisson maximum likelihood estimator (without individual and time effects), MLE-TE corresponds to the Poisson estimator with only time effects, MLE-FETE corresponds to the Poisson maximum
likelihood estimator with individual and time fixed effects, Analytical (L=l) is the bias corrected estimator that uses the analytical correction with $L=l$,
and Jackknife is the bias corrected estimator that
uses SPJ in both the individual and time
dimensions. The analytical corrections are different from the uncorrected estimator because they do not use that the regressor $Z_{it}$ is strictly exogenous. The cross-sectional division in the jackknife follows the
order of the observations. The choice of these estimators is motivated by the empirical analysis of ABBGH. All the results in the table are
reported in percentage of the true parameter value.
The results of the table agree with the no asymptotic bias result for the Poisson model with exogenous regressors. Thus, the bias of MLE-FETE for the coefficients and APE is negligible relative to the standard deviation and the coverage probabilities get close to the nominal level as $N$ grows. The analytical corrections preserve the performance of the estimators and have very little sensitivity to the trimming parameter. The jackknife correction increases dispersion and rmse, specially for the small cross-sectional size of the application. The estimators that do not control for individual effects are clearly biased.
Dynamic Poisson model
The data generating process is
equation*[equation* omitted — 210 chars of source]
The competition variable $Z_{it}$ and the initial condition for the number of patents $Y_{i0}$ are fixed to the values
in the dataset and all the parameters are set to the fixed effect estimates of the model. To generate panels, we first impute values
to the missing observations of $Z_{it}$ using forward and backward predictions from a panel AR(1) linear model with individual and time effects. We then draw
panel data sets with $T=21$ years and three different numbers
of industries $N$: 17, 34, and 51. As in the static model, we double (triple) the cross-sectional size by merging two (three) independent realizations of the panel. We make the generated panels unbalanced by dropping the values corresponding to the missing observations in the original dataset.
Table S4 reports the simulation results for the coefficient $\beta^0_Y$ and the APE of $Y_{i,t-1}$. The estimators considered are the same as for the static Poisson model above. We compute the partial effect of $Y_{i,t-1}$ using ((ref)) with $Z_{it} = Y_{i,t-1}$, $H(Z_{it}) = \log (1+Z_{it}),$ and dropping the linear term.
Table S5 reports the simulation results for the coefficients $\beta_1^0$ and $\beta_2^0$, and the APE of $Z_{it}$. We compute the partial effect using ((ref)) with $H(Z_{it}) = Z_{it}^2$.
Again, all the results in the tables are
reported in percentage of the true parameter value.
The results in table S4 show biases of the same order of magnitude as the standard deviation for the fixed effects estimators of the coefficient and APE of $Y_{i,t-1}$, which cause severe undercoverage of confidence intervals. Note that in this case the rate of convergence for the estimator of the APE is $r_{NT} = \sqrt{NT}$, because the individual and time effects are hold fixed across the simulations. The analytical corrections reduce bias by more than half without increasing dispersion, substantially reducing rmse and bringing coverage probabilities closer to their nominal levels. The jackknife corrections reduce bias and increase dispersion leading to lower improvements in rmse and coverage probability than the analytical corrections. The results for the coefficient of $Z_{it}$ in table 8 are similar to the static model. The results for the APE of $Z_{it}$ are imprecise, because the true value of the effect is close to zero.
Proofs of Theorems (ref) and (ref)
We start with a lemma that shows the consistency of the fixed effects estimators of averages of the data and parameters. We will use this result to show the validity of the analytical bias corrections and the consistency of the variance estimators.
lemmaLet $G(\beta,\phi) := [N(T-j)]^{-1} \sum_{i,t \geq j+1} g(X_{it}, X_{i,t-j}, \beta, \alpha_{i} + \gamma_t, \alpha_{i} + \gamma_{t-j})$ for $0 \leq j < T,$ and $\mathcal{B}_{\varepsilon}^0$ be a subset of $\mathbb{R}^{\dim \beta + 2}$ that contains an $\varepsilon$-neighborhood of $(\beta, \pi_{it}^0, \pi_{i,t-j}^0)$ for all $i,t,j,N, T$, and for some $\varepsilon > 0$. Assume that $(\beta, \pi_1, \pi_2) \mapsto g_{itj}(\beta, \pi_1, \pi_2) := g(X_{it}, X_{i,t-j}, \beta, \pi_1, \pi_2)$ is Lipschitz continuous over $\mathcal{B}_{\varepsilon}^0$ a.s, i.e. $|g_{itj}(\beta_1, \pi_{11}, \pi_{21}) - g_{itj}(\beta_0, \pi_{10}, \pi_{20})| \leq M_{itj} \|(\beta_1, \pi_{11}, \pi_{21}) - (\beta, \pi_{10}, \pi_{20}) \|$ for all $(\beta_0, \pi_{10}, \pi_{20}) \in \mathcal{B}_{\varepsilon}^0$, $(\beta_1, \pi_{11}, \pi_{21}) \in \mathcal{B}_{\varepsilon}^0$, and some $M_{itj} = \mathcal{O}_P(1)$ for all $i,t,j,N, T$. Let $(\widehat \beta, \widehat \phi)$ be an estimator of $(\beta, \phi)$ such that $\|\widehat \beta - \beta^0\| \to_P 0$ and $\|\widehat \phi - \phi^0 \|_{\infty} \to_P 0.$ Then,
$$
G(\widehat \beta, \widehat \phi) \to_P \overline{\mathbb{E}}[G(\beta^0, \phi^0)],
$$
provided that the limit exists.
proof[\bf Proof of Lemma (ref)] By the triangle inequality
$$
|G(\widehat \beta, \widehat \phi) - \overline{\mathbb{E}}[G(\beta^0, \phi^0)]| \leq |G(\widehat \beta, \widehat \phi) - G(\beta^0, \phi^0)| + o_P(1),
$$
because $|G(\beta^0, \phi^0) - \overline{\mathbb{E}}[G(\beta^0, \phi^0)]| = o_P(1)$. By the local Lipschitz continuity of $g_{itj}$ and the consistency of $(\widehat \beta, \widehat \phi)$,
\begin{multline*}
|G(\widehat \beta, \widehat \phi) - G(\beta^0, \phi^0)| \leq \frac{1}{N(T-j)} \sum_{i,t \geq j+1} M_{itj} \|(\widehat \beta, \widehat \alpha_i + \widehat \gamma_t , \widehat \alpha_i + \widehat \gamma_{t-j}) - ( \beta^0, \alpha_i^0 + \gamma_t^0 , \alpha_i^0 + \gamma_{t-j}^0) \| \\ \leq \frac{1}{N(T-j)} \sum_{i,t \geq j+1} M_{itj} (\|\widehat \beta - \beta^0 \| + 4 \| \widehat \phi - \phi^0 \|_{\infty} )
\end{multline*}
wpa1. The result then follows because $[N(T-j)]^{-1} \sum_{i,\tau \geq t} M_{it\tau} = \mathcal{O}_P(1)$ and $(\|\widehat \beta - \beta^0 \| + 4 \| \widehat \phi - \phi^0 \|_{\infty} ) = o_P(1)$ by assumption.
proof[\bf Proof of Theorem (ref)] We separate the proof in three parts corresponding to the three statements of the theorem.
Part I: Proof of $\widehat W \to_P \overline{W}_{\infty}$. The asymptotic variance and its fixed effects estimators can be expressed as $\overline{W}_{\infty} = \overline{\mathbb{E}}[W(\beta^0, \phi^0)]$ and $\widehat W = W(\widehat \beta, \widehat \phi),$ where $W(\beta,\phi)$ has a first order representation as a continuously differentiable transformation of terms that have the form of $G(\beta,\phi)$ in Lemma (ref). The result then follows by the continuous mapping theorem noting that $\|\widehat \beta - \beta^0\| \to_P 0$ and $\|\widehat \phi - \phi^0 \|_{\infty} \leq \|\widehat \phi - \phi^0 \|_{q} \to_P 0$ by Theorem (ref).
Part II: Proof of $\sqrt{NT}(\widetilde \beta^A - \beta^0) \to_d \mathcal{N}(0,
\overline W_{\infty}^{-1}).$ By the argument given after equation (ref) in the text, we only need to show that $\widehat B \to_P \overline{B}_{\infty}$ and $\widehat D \to_P \overline{D}_{\infty}$. These asymptotic biases and their fixed effects estimators are either time-series averages of fractions of cross-sectional averages,
or vice versa. The nesting of the averages makes the analysis a bit more cumbersome than the analysis of $\widehat W$, but the result follows by similar standard arguments, also using that
$L \to \infty$ and $L/T \to 0$ guarantee that the trimmed estimator in $\widehat B$ is
also consistent for the spectral expectations; see Lemma 6 in Hahn and Kuersteiner HahnKuersteiner2011.
Part III: Proof of $\sqrt{NT}(\widetilde \beta^J - \beta^0) \to_d \mathcal{N}(0,
\overline W_{\infty}^{-1}).$ For $\mathcal{T}_1 = \{1, \ldots, \lfloor (T+1)/2 \rfloor \}$, $\mathcal{T}_2 = \{\lfloor T/2 \rfloor + 1, \ldots, T \}$, $\mathcal{T}_0 = \mathcal{T}_1 \cup \mathcal{T}_2$, $\mathcal{N}_1 = \{1, \ldots, \lfloor (N+1)/2 \rfloor \}$, $\mathcal{N}_2 = \{\lfloor N/2 \rfloor + 1, \ldots, N \}$, and $\mathcal{N}_0 = \mathcal{N}_1 \cup \mathcal{N}_2$, let $\widehat \beta^{(jk)}$ be the fixed effect estimator of $\beta$ in the subpanel defined by $i \in \mathcal{N}_j$ and $t \in \mathcal{T}_k$.\footnote{Note that this definition of the subpanels covers all the cases regardless of whether $N$ and $T$ are even or odd.} In this notation,
$$
\widetilde \beta^J = 3 \widehat \beta^{(00)} - \widehat \beta^{(10)} / 2 - \widehat \beta^{(20)} /2 - \widehat \beta^{(01)} /2 - \widehat \beta^{(02)}/2.
$$
We derive the asymptotic distribution of $\sqrt{NT}(\widetilde \beta^J -\beta^0)$ from the joint asymptotic distribution of the vector $\widehat{ \mathbb{B}} = \sqrt{NT}(\widehat \beta^{(00)} -\beta^0,\widehat \beta^{(10)} -\beta^0,\widehat \beta^{(20)} -\beta^0,\widehat \beta^{(01)} -\beta^0,\widehat \beta^{(02)} -\beta^0)$ with dimension $5 \times \dim \beta$. By Theorem (ref),
$$
\sqrt{NT}(\widehat \beta^{(jk)} -\beta^0) = \frac{2^{1(j>0)}2^{1(k>0)}}{\sqrt{NT}} \sum_{i \in N_j, t \in T_k} \left[\psi_{it} + b_{it} + d_{it} \right] + o_P(1),
$$
for $\psi_{it} = \overline{W}_{\infty}^{-1} D_{\beta} \ell_{it}$, $b_{it} = \overline{W}_{\infty}^{-1} [U_{it}^{(1a,1)} + U_{it}^{(1b,1,1)}]$, and $d_{it} =\overline{W}_{\infty}^{-1} [U_{it}^{(1a,4)} + U_{it}^{(1b,4,4)}]$, where the $U_{it}^{(\cdot)}$ is implicitly defined by $U^{(\cdot)} = (NT)^{-1/2} \sum_{i,t} U_{it}^{(\cdot)}$.
Here, none of the terms carries a superscript $(jk)$ by Assumption (ref). The influence function $\psi_{it}$ has zero mean and determines the asymptotic variance $\overline{W}_{\infty}^{-1}$, whereas $b_{it} $ and $d_{it}$ determine the asymptotic biases $\overline{B}_{\infty}$ and $\overline{D}_{\infty},$ but do not affect the asymptotic variance. By this representation,
$$
\widehat{ \mathbb{B}} \to_d \mathcal{N}
\left(\kappa \left[
\begin{array}{c}
1 \\
1 \\
1 \\
2 \\
2
\end{array}
\right] \otimes \overline{B}_{\infty} + \kappa^{-1} \left[
\begin{array}{c}
1 \\
2 \\
2 \\
1 \\
1
\end{array}
\right] \otimes \overline{D}_{\infty}, \left[
\begin{array}{ccccc}
1 & 1 & 1 & 1 & 1 \\
1 & 2 & 0 & 1 & 1 \\
1 & 0 & 2 & 1 & 1 \\
1 & 1 & 1 & 2 & 0 \\
1 & 1 &1 & 0 & 2
\end{array}
\right] \otimes \overline{W}_{\infty}^{-1} \right),
$$
where we use that $\{\psi_{it}: 1 \leq i \leq N, 1 \leq t \leq T\}$ is independent across $i$ and martingale difference across $t$ and Assumption (ref).
The result follows by writing $\sqrt{NT}(\widetilde \beta^J - \beta^0) = (3,-1/2,-1/2,-1/2,-1/2) \widehat{ \mathbb{B}}$ and using the properties of the multivariate normal distribution.
proof[\bf Proof of Theorem (ref)] We separate the proof in three parts corresponding to the three statements of the theorem.
Part I: $\widehat V^{\delta} \to_P \overline{V}^{\delta}_{\infty}$. $ \overline{V}^{\delta}_{\infty}$ and $\widehat V^{\delta}$ have a similar structure to $\overline{W}_{\infty}$ and $\widehat W$ in part I of the proof of Theorem (ref), so that the consistency follows by an analogous argument.
Part II: $
\sqrt{NT} (\widetilde \delta^A - \delta_{NT}^0) \to_d \mathcal{N}(0, \overline{V}_{\infty}^{\delta})
$. As in the proof of Theorem (ref),
we decompose
$$
r_{NT} (\widetilde \delta^A - \delta_{NT}^0) = r_{NT} (\delta - \delta_{NT}^0) + \frac{r_{NT}}{\sqrt{NT}} \sqrt{NT} (\widetilde \delta^A - \delta).
$$
Then, by Mann-Wald theorem,
$$
\sqrt{NT} (\widetilde \delta^A - \delta) = \sqrt{NT} (\widehat \delta - \widehat B^{\delta}/T - \widehat D^{\delta}/N - \delta) \to_d \mathcal{N}(0, \overline{V}_{\infty}^{\delta(1)}),
$$
provided that $ \widehat B^{\delta} \to_P \overline{B}_{\infty}^{\delta}$ and $ \widehat D^{\delta} \to_P \overline{D}_{\infty}^{\delta}$, and $ r_{NT} (\delta - \delta_{NT}^0) \to_d \mathcal{N}(0, \overline{V}_{\infty}^{\delta(2)})$, where $\overline{V}_{\infty}^{\delta(1)}$ and $\overline{V}_{\infty}^{\delta(2)}$ are defined as in the proof of Theorem (ref). The statement thus follows by using a similar argument to part II of the proof of Theorem (ref) to show the consistency of $ \widehat B^{\delta}$ and $ \widehat D^{\delta}$, and because $ (\delta - \delta_{NT}^0)$ and $(\widetilde \delta^A - \delta)$
are asymptotically independent, and $\overline{V}_{\infty}^{\delta} = \overline{V}^{\delta(2)} + \overline{V}^{\delta(1)} \lim_{N,T \to \infty} (r_{NT}/\sqrt{NT})^2.$
Part III: $
\sqrt{NT} (\widetilde \delta^J - \delta_{NT}^0) \to_d \mathcal{N}(0, \overline{V}_{\infty}^{\delta})
$. As in part II,
we decompose
$$
r_{NT} (\widetilde \delta^J - \delta_{NT}^0) = r_{NT} (\delta - \delta_{NT}^0) + \frac{r_{NT}}{\sqrt{NT}} \sqrt{NT} (\widetilde \delta^J - \delta).
$$
Then, by an argument similar to part III of the proof of Theorem (ref),
$$
\sqrt{NT} (\widetilde \delta^J - \delta) \to_d \mathcal{N}(0, \overline{V}_{\infty}^{\delta(1)}),
$$
and $ r_{NT} (\delta - \delta_{NT}^0) \to_d \mathcal{N}(0, \overline{V}_{\infty}^{\delta(2)})$, where $\overline{V}_{\infty}^{\delta(1)}$ and $\overline{V}_{\infty}^{\delta(2)}$ are defined as in the proof of Theorem (ref). The statement follows because $ (\delta - \delta_{NT}^0)$ and $(\widetilde \delta^J - \delta)$
are asymptotically independent, and $\overline{V}_{\infty}^{\delta} = \overline{V}^{\delta(2)} + \overline{V}^{\delta(1)} \lim_{N,T \to \infty} (r_{NT}/\sqrt{NT})^2.$