Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,589 characters · 16 sections · 85 citation commands
Indirect Inference for Nonlinear Panel Models with Fixed Effects
Panel data refers to data on multiple entities (e.g., individuals, firms, etc.) observed at two or more time periods. Unobserved heterogeneity across entities often accounts for a large fraction of the variation in panel data. When this heterogeneity is correlated with the explanatory variables in the regression specifications, the resulting omitted variable bias renders point estimates inconsistent.
Adding individual fixed effects, $\alpha_{i0}$'s, is the main approach to control for time--invariant unobserved heterogeneity in panel data models. Compared to other approaches like random effects and correlated random effects, the fixed effect approach does not impose distributional assumptions on $\alpha_{i0}$'s or restrict their relationships with other explanatory variables. Instead, each $\alpha_{i0}$ is treated as a parameter to be estimated. However, because the number of $\alpha_{i0}$'s increases with the sample size and each $\alpha_{i0}$ is estimated using only entity $i$'s time series observations, adding fixed effects introduces the incidental parameter problem in estimating the vector of parameters of interest $\theta_{0}$. It has two consequences for applied research: (1) point estimates are subject to large biases, and (2) confidence intervals have incorrect coverages.
This paper proposes a new method to debias fixed effect estimators in a class of nonlinear panel models. The method is named indirect fixed effect estimation and features two main steps: the first one is to simulate data by using estimated individual effects $\widehat{\alpha}_{i}$'s from the observed data. The second step is to find the vector of parameters that matches the fixed effect estimators using observed and simulated data.
The method has two advantages: first, it does not require an explicit characterization of the bias term, which can be hard to derive in complex models. Instead, the method finds the solution by automatically correcting the bias because the vector of parameter values that is the closest to $\theta_{0}$ renders similar bias in fixed effect estimations. Second, standard errors can be derived using the delta method, so there is no need to use the bootstrap, which is computationally intensive.
The two properties are inherited from a precedent simulation--based estimation approach called indirect inference, which was first developed by GourierouxMonfortRenault1993 and Smith1993. In a nutshell, indirect inference uses an auxiliary model to summarize the statistical properties of the observed data and simulated data, and finds values of model parameters that match the parameters of the auxiliary model, estimated using the observed and simulated data, in terms of a minimum--distance criterion function. Because the same regression is run on observed and simulated data, matched estimators have the same bias structure and thus the bias gets cancelled.
The theory of indirect inference, however, is not directly applicable to nonlinear panel models, which are widely used in various fields of economics like industrial organization and labor. Because the individual effects cannot be differenced out, data simulations seem infeasible without imposing a parametric specification on their distributions, and the bias term is a complicated function of $\theta_{0}$ and $\alpha_{i0}$'s.
To simulate data, this paper proposes using the estimated individual effects $\widehat{\alpha}_{i}$'s. These are informative proxies for the unknown individual effects $\alpha_{i0}$'s because they become more accurate estimates when each individual's number of time series observations $T$ grows large. Intuitively speaking, although data simulated using $\widehat{\alpha}_{i}$'s do not perfectly mimic the observed data, such a difference vanishes when $T$ increases.
The indirect fixed effect estimator then debiases by matching the fixed effect estimates using observed and simulated data. This brings two advantages for the implementation and theoretical analysis of the new estimator. First, the minimum--distance criterion function for matching is just--identified because the dimensions of the fixed effect estimates are identical. Therefore, there is no need to consider an estimation of an optimal weighting matrix. It further implies that the matching can be made as exact as machine precision permits. The second advantage is with respect to the relationship between the vector of parameters of interest $\theta_{0}$ and the unique maximizer of the limiting log--likelihood function for fixed effect estimation. To back out point estimates of $\theta_{0}$ from fixed effect estimators using simulated data, this relationship should be invertible. Because the unique maximizer is $\theta_{0}$, the relation turns out to be an identity function. Therefore, invertibility is satisfied trivially.
This paper derives consistency, bias correction and asymptotic normality results for the indirect fixed effect estimator. As usual in the indirect inference literature, consistency requires that the fixed effect estimates using observed and simulated data converge to the unique maximizer of the limiting log likelihood. Although the pointwise convergence of $\widehat{\theta}$ to $\theta_{0}$ is a standard result in the large--$T$ panel literature, three important differences arise in the analysis of fixed effect estimates using simulated data and pose theoretical challenges.
First, the simulated data are generated using $\widehat{\alpha}_{i}$'s instead of $\alpha_{i0}$'s. To justify this practice, the corresponding log likelihood function should uniformly well--approximate the one rendered by data simulated using the true individual effects. Otherwise, simulated fixed effect estimator could not be pointwise convergent. The proof of this statement, however, is complicated by the fact that the log likelihood function using simulated data is typically nonsmooth for important types of nonlinear panel models, with binary choice models as leading examples. Intuitively speaking, when the dependent variable is discrete, a small change in the parameter values can lead to discrete changes in the simulated data. As a result, the sample log likelihood function using simulated data is discontinuous.
Simulations often generate discontinuous objective functions McFadden1989, PakesPollard1989, but this paper confronts a second difference: the fixed effect estimator using simulated data is nonsmooth with respect to the parameters of the data generating process (DGP). Therefore, standard proof strategies in the panel literature HahnNewey2004, HahnKuersteiner2011 cannot be directly applied to characterize its limiting behavior.
Empirical process theory provides ample tools to handle nonsmoothness functions and moments in econometrics Andrews1994, but the analysis of a nonsmooth fixed effect estimator is further complicated by the third difference: the presence of incidental parameters, whose number increases with the sample size $n$.
To prove uniform convergence with nonsmoothness, this paper follows Newey1991 by establishing pointwise convergence and stochastic equicontinuity of the fixed effect estimator in the simulation world. Intuitively speaking, pointwise convergence is equivalent to uniform convergence for any finite number of grid points, but without smoothness, the gap between any two grids can behave rather erratically. The stochastic equicontinuity condition is hence required to restrict such behaviors in probability.
The theoretical analysis of the indirect fixed effect estimator relies on some key structures of the panel data and the log likelihood function. Under the assumption that panel data are independent along the cross section dimension, this paper first justifies data simulation with $\widehat{\alpha}_{i}$'s by proving that the corresponding log likelihood function uniformly approximates the one from simulated data generated by $\alpha_{i0}$'s. As such, a uniform law of large number can be established and pointwise convergence in the simulation world follows from the standard consistency argument NeweyMcFadden1994. To verify the stochastic equicontinuity condition of fixed effect estimators using simulated data, this paper uses the concavity property of the profiled log likelihood to verify one of the primitive conditions for stochastic equicontinuity in Andrews1994. The proof strategy might be of independent interest.
Regarding the asymptotic unbiasedness and normality of the new estimator, due to non--smoothness in the simulation world, the conventional strategy in indirect inference that relies on the implicit function theorem GourierouxMonfortRenault1993 is not directly applicable. A regularity conditions is thus imposed, which, combined with consistency, allows to explore bias correction and asymptotic normality through the lens of fixed effect estimates. More specifically, the fixed effect estimators in both worlds have the same structures regarding the bias terms and influence functions. The difference is that the ones using the real data are functions of $\theta_{0}$ and $\alpha_{i0}$'s while those using the simulated data are functions of $\theta_{0}$ and $\widehat{\alpha}_{i}$'s.
This paper currently imposes two high--level conditions to ensure that the bias term and the influence function from data simulated using $\theta_{0}$ and $\widehat{\alpha}_{i}$'s is uniformly close to their infeasible counterparts from data simulated using $\theta_{0}$ and $\alpha_{i0}$'s with asymptotically negligible approximation errors. The infeasible bias converges to the same probability limit as does the bias obtained from observed data, while the infeasible influence function converges to the same normal distribution as does the influence function from observed data. Therefore, the theory of indirect inference can be invoked to establish bias cancellation and asymptotic normality.
Like other simulation--based estimation methods, the asymptotic variance of the new estimator is inflated by the inverse of the number of simulation draws. The result can also be intepreted as a reflection of the classic bias--variance tradeoff. As shown in the application and Monte Carlo simulations, however, the finite--sample performance of the indirect fixed effect estimator is comparable to the leading methods in terms of bias correction and outperforms half--panel bias correction methods in terms of standard errors.
The indirect fixed effect estimator presented in this paper combines four strands of literature, of which this section provides a non--exhaustive overview. The incidental parameter problem is first discussed by NeymanScott1948. When $T$ is fixed, fixed effect estimators of nonlinear models are in general inconsistent because estimation errors of $\widehat{\alpha}_{i}$'s do not vanish even when the cross--section sample size $n$ is very large Chamberlain1984, Lancaster2000. Only some special models like static linear and logit specifications feature fixed--$T$ consistent estimators Anderson1970. A key insight of the large--$T$ panel data literature is that the incidental parameter problem becomes an asymptotic bias problem when $T$ grow with the sample size $n$. When $n$ and $T$ grow at the same rate, fixed effect estimators are consistent and asymptotically normal, but they have a bias comparable to standard errors.
In the search for asymptotically unbiased estimators, there are two leading approaches. For certain types of models, the bias terms have been characterized analytically and corrected using a plug--in approach HahnKuersteiner2002, HahnNewey2004, Fernandez-Val2009, HahnKuersteiner2011. However, such terms can be hard to derive for complicated models. Under further sampling and regularity conditions, bias terms can be automatically corrected using jackknife. For example, HahnNewey2004 proposed leave--one--out panel jackknife for data that do not have dependencies among observations of the same unit. DhaeneJochmans2015 relaxed the assumption to stationarity along the time series, and proposed a half--panel method. Under an unconditional homogeneity assumption, Fernandez-ValWeidner2016 allowed for two--way fixed effects and propose a jackknife method that corrects biases from both dimensions. See ArellanoHahn2007 and Fernandez-ValWeidner2018 for recent surveys. Standard errors are typically obtained using panel bootstrap, which can be computationally intensive.\footnote{For example, to obtain one debiased point estimate, fixed effect estimations are run three times: one for the whole sample, and twice for the two split samples. If the number of bootstraps is set to be 500, then the total number of fixed effects estimations becomes 1500. In addition, in practice it is often recommended to use multiple sample splits to improve the finite--sample performance.}
Another popular simulation--based method that can achieve bias correction is bootstrap Horowitz2001, Horowitz2019. GoncalvesKaffo2015 proposed the bootstrap bias correction methods for dynamic linear panel models without covariates. KimSun2016 proposed a parametric bootstrap bias correction (BBC) method for the nonlinear panel models considered in this paper. Compared to the indirect fixed effect estimator, the BBC estimator mimics the bias term and removes it from the fixed effect estimate explicitly, and thus the proof strategies are very different.
Second, this paper extends the existing theory and practice of indirect inference. Since the introduction of the method, its asymptotic theory has mainly been focused on times series data GourierouxMonfortRenault1993, Smith1993, GallantTauchen1996. Some recent papers explore asymptotic properties in panel data with discrete dependent variables, but there are two key differences with this paper. First, their settings hold time series dimension fixed and study different types of models. For example, GII2018 did not consider models with fixed effects, FrazierOkaZhu2019 imposed normality on individual effects, and TaberSauer2021 assumed a bivarite normal distribution on the types of individuals. Second, they deal with nonsmoothness by smoothing the discontinuous parts and showing that the resultant bias can be corrected.
GourierouxPhillipsYu2010 is the first paper that establishes theoretical properties of indirect inference for a class of large--$T$ panel models. They applied indirect inference to dynamic panel linear models, whose fixed effect estimators are known to be biased Nickell1981. The linear structure allows them to eliminate the individual fixed effects $\alpha_{i0}$'s by first--difference. As such, $\alpha_{i0}$'s do not show up in the bias term, and data can be simulated without information on them. However, first difference does not remove the $\alpha_{i0}$'s to nonlinear panel models, and this paper fills the gap by extending the theory to handle the presence of $\alpha_{i0}$'s in data simulation and the bias term.
Indirect inference is popular in various fields of economics, including empirical industrial organization CollardWexler2013, labor economics AltonjiSmithVidangos2013 and macroeconomics GuvenenSmith2014, BergerVavra2019. However, finding an informative auxiliary model is not a trivial task, and researchers often have to assume invertibility of the limiting relationship between auxiliary parameters and parameters of interest. This paper provides an alternative choice, namely the log likelihood function from the nonlinear panel model, for researchers that employ panel data with fixed effects. The estimation procedures are simple to implement as fixed effect estimation schemes are available in free software like R and Julia.
Nonsmooth objective functions are common in econometrics, and empirical process methods are standard tools for asymptotic analysis Andrews1994, NeweyMcFadden1994, VanDerVaartWellner1996. The seminal work on simulation--based methods by PakesPollard1989 is predicated on the independence assumption of cross section data and therefore is not suitable for panel data, which feature dependence for each individual time series. DedeckerLouhichi2002 provided an overview of maximal inequalities for empirical central limit theorems for dependent data. KatoGalvaoMontesRojas2012 provided new stochastic inequalities for mixing sequences and also established stochastic equicontinuity in the presence of nuisance parameters, but their analysis focused on a different class of nonlinear models, namely panel quantile regression models.
Simulation--based methods like simulated method of moments McFadden1989, PakesPollard1989, LeeIngram1991, DuffieSingleton1993 and indirect inference are widely used to estimate models that do not render tractable moments or likelihood functions. See GourierouxMonfort1997 for an overview. These methods typically require models to be fully specified, but economic theory does not always provide guidance on functional forms, distributions of shocks or measurement error of observed data. Therefore, the resultant estimators can be subject to misspecification.
This paper considers a class of nonlinear panel models that do not impose distributional assumptions on the individual effects, and hence contributes to a burgeoning literature that considers simulations for models that relax the full parametric specifications in various ways. DridiRenault2000 and DridiGuayRenault2007 embedded the semiparametric structural model into a full model for data simulation, and proposed an encompassing principle where parameters of interest are consistently estimated even though nuisance parameters are inconsistently estimated due to misspecification of the full model. Schennach2014 considered parameters estimation in moment conditions that contain unobservable variables, and proposed a simulation--based method that constructs equivalent moments involving only observable variables. GospodinovKomunjerNg2017 considered parameter estimation of autoregressive distributed lag models in which covariates are contaminated by serially correlated measurement errors. They proposed a method such that simulated covariates preserve the dependence structure observed in the data even though the dynamics of latent covariates or measurement errors are not specified. Forneron2020 approximated the distribution of shocks by sieves and proposes a sieve--SMM estimator that jointly estimates structural parameters and the distribution of shocks.
The rest of the paper proceeds as follows: Section (ref) introduces the model and describes the fixed effect estimator and incidental parameter problem. Section (ref) provides an overview of the indirect fixed effect estimator and its implementation. Section (ref) presents the theoretical properties of the estimator. Section (ref) applies the method to a dataset on labor force participation to illustrate the finite--sample properties of the estimator. Section (ref) uses numerical simulations to compare the new estimator with other bias correction methods. Section (ref) concludes and discusses open questions. Appendices (ref), (ref) and (ref) consist of proofs and computation details.
This section starts with a description of nonlinear panel models with fixed effects. Let the data observations be denoted by $\{z_{it}=(y_{it}, x_{it})$: $i=1,\dots,n; t=1,\dots,T\}$, where $y_{it}$ is the dependent variable and $x_{it}$ is a $p\times 1$ vector of explanatory variables. The observations are independent across entity $i$ and weakly dependent across time $t$. The DGP of outcome $y_{it}$ takes the following form:
where $x_{i}^{T}:=(x_{i1},\dots,x_{iT})$, $\theta$ is a $p\times 1$ vector of model parameters and $\alpha_{i}$ is a scalar individual effect. The explanatory variable $x_{it}$ is strictly exogenous. The model is semiparametric in that neither the distribution of $\alpha_{i}$ nor its relationship with $x_{it}$ is specified. The conditional density $f$ denotes the parametric part of the model and its form depends on the parametric family of distributions $\{u_{it}\}$. Depending on the specification of $f$, this type of models have been used to study many different questions of economic interest.
Model ((ref)) admits a log likelihood function. The true values of the parameters, denoted by $\theta_{0}$ and $\alpha_{i0}$'s, are one solution to the population conditional maximum likelihood problem
where $\Theta$ and $\Gamma_{\alpha}$ are the parameter spaces for $\theta$ and $\alpha_{i}$ respectively, the expectation is with respect to the distribution of the observed data, conditional on the unobserved effects and initial conditions. Section (ref) discusses assumptions under which the log likelihood function is concave in all parameters and the solution uniquely exists. The indirect fixed effect estimator relies on the uniqueness condition for consistency.
The fixed effect estimator of $\theta$ is obtained by doing maximum likelihood estimation on the sample analog of the population problem ((ref)), treating each $\alpha_{i}$ as a parameter to be estimated.
To facilitate theoretical analysis, this equation is rewritten such that the individual effects are profiled out. More specifically, given $\theta$, the optimal $\widehat{\alpha}_{i}(\theta)$ for each $i$ is defined as
The estimators $\widehat{\theta}$ and $\widehat{\alpha}_{i}$ are then
Section (ref) discusses assumptions under which these estimators exist and are unique with probability approaching one as $n$ and $T$ become large.
In panel models, the individual effects are incidental parameters, i.e., nuisance parameters whose dimension grows with the number of cross sectional observations $n$. As equation ((ref)) shows, the fixed effect estimator $\widehat{\theta}$ cannot generally be separated from the estimator of individual effects $\widehat{\alpha}_{i}$'s. Because each $\widehat{\alpha}_{i}$ is only estimated using the $T$ observations for entity $i$, its estimation error does not vanish if $T$ is fixed, even as $n$ grows. These estimation errors in turn contaminate $\widehat{\theta}$. This is the incidental parameter problem for fixed effects estimation. Mathematically,
whereas $$\theta_{0}:=\operatorname*{arg\,max}_{ \theta\in\Theta}\operatorname*{plim}_{n\rightarrow\infty}\frac{1}{n}\sum^{n}_{i=1}\Big(\frac{1}{T}\sum^{T}_{t=1}\mathbb{E}\ln f(y_{it}\mid x_{it}, \theta, \alpha_{i}(\theta))\Big),$$ where $$\alpha_{i}(\theta)=\operatorname*{arg\,max}_{\alpha\in\Gamma_{\alpha}}\frac{1}{T}\sum^{T}_{t=1}\mathbb{E}(\ln f(y_{it}\mid x_{it}, \theta, \alpha)).$$ With fixed $T$, $\widehat{\alpha}_{i}(\theta)\neq \alpha_{i}(\theta)$ in general. Therefore, $\theta_{T}\neq \theta_{0}$.
To illustrate the problem, suppose $y_{it}$ has the normal distribution with mean $\alpha_{i0}$ and variance $\theta_{0}$, and the goal is to estimate $\theta_{0}$. The fixed effect estimator is $\widehat{\theta}=\frac{1}{nT}\sum^{n}_{i=1}\sum^{T}_{t=1}(X_{it}-\widehat{\alpha}_{i})^{2},$ where $\widehat{\alpha}_{i}=\frac{1}{T}\sum^{T}_{t=1}X_{it}.$ When $T$ is fixed and $n$ approaches infinity, NeymanScott1948 show that $$\widehat{\theta}\xrightarrow{p}\theta_{0}-\frac{\theta_{0}}{T}.$$ On the other hand, when $T$ also grows to infinity, the bias term $-\frac{\theta_{0}}{T}$ approaches zero. The large--$T$ panel literature generalizes this insight and shows that the incidental parameter problem becomes an asymptotic bias problem when $n$ and $T$ grow at the same rate.
The key feature of the indirect fixed effect estimator is to match $\widehat{\theta}$ with a fixed effect estimator from simulated data generated by $\widehat{\alpha}_{i}$'s and a given $\theta$. To avoid confusion, it is necessary to introduce notations to distinguish parameters in the simulation world from those in Section (ref). More specifically, this paper uses $\beta$ and $\gamma_{i}$ to denote the vector of parameters of interest and individual effects in the log likelihood function using simulated data.
To clarify the notations and introduce the implementation of indirect fixed effect estimator, this section first revisits the Neyman--Scott example. Using the panel model as a concrete example, this section then illustrates the challenges associated with the presence of nonsmoothness and discusses the general estimation procedures.
If $y_{it}\mid \alpha_{i0} \sim \mathcal{N}(\alpha_{i0}, \theta_{0})$ is i.i.d over $n$ and $t$, then the DGP of the observed data is:
This equation cannot be simulated without information on the distribution of $\alpha_{i0}$'s. The indirect fixed effect estimator uses $\widehat{\alpha}_{i}$'s instead, and the simulated data have the following DGP:
where the superscript $h$ denotes a simulation path. The fixed effect estimator using $\{y^{h}_{it}(\widehat{\alpha}_{i}, \theta)\}$ is
where $\boldsymbol{\widehat{\alpha}} := (\widehat{\alpha}_{1},\dots,\widehat{\alpha}_{n})$ and $\widehat{\gamma}_{i}:=\frac{1}{T}\sum^{T}_{t=1}y^{h}_{it}(\widehat{\alpha}_{i}, \theta)$. The interpretation of $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$ is that the estimator changes if a different value of $\theta$ is used to simulate the data. Note that the $\widehat{\alpha}_{i}$'s are fixed throughout the simulation process. The indirect fixed effect estimator $\widetilde{\theta}$ is the solution to $$\widehat{\theta}=\widehat{\beta}^{h}(\widetilde{\theta}, \boldsymbol{\widehat{\alpha}}).$$ Figure ((ref)) illustrates the issues with $\widehat{\theta}$ and the performance of $\widetilde{\theta}$ in this example. The density of the fixed effect estimator $\widehat{\theta}$ conveys two messages: (1) fixed effect estimator is subject to a large bias and (2) a confidence interval around $\widehat{\theta}$ would not have the correct coverage. The density of $\widetilde{\theta}$ illustrates that (1) the new estimator corrects the bias significantly and (2) a confidence interval around $\widetilde{\theta}$ would have a much larger coverage than the one around $\widehat{\theta}$.
Consider the binary choice panel model as a concrete example. Given $\theta$, $\widehat{\alpha}_{i}$'s and $x_{it}$, the simulated dependent variable is $$y_{it}^{h}(\theta, \widehat{\alpha}_{i})=\boldsymbol{1}(x'_{it}\theta+\widehat{\alpha}_{i}>u^{h}_{it}), \quad u^{h}_{it}\sim\mathcal{N}(0, 1),$$ where $u^{h}_{it}$ are simulation draws from the standard normal distribution. The corresponding log likelihood function is
It illustrates the three different aspects of simulated fixed effect estimator $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$. Because $y^{h}_{it}(\theta, \widehat{\alpha}_{i})$ is discontinuous in $\theta$ and $\widehat{\alpha}_{i}$'s, equation ((ref)) is discontinuous, which carries over to its maximizer $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$. In addition, estimating $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$ involves incidental parameters $\gamma_{i}$'s. The population version of equation ((ref)) does not have randomness due to data sampling, use of simulations and $\widehat{\alpha}_{i}$'s, and takes the form
where the expectation is over $u^{h}_{it}$ and $x_{it}$, and $\widehat{\alpha}_{i}$'s are replaced by $\alpha_{i0}$'s.
\setcounter{example}{0} From the known distribution $F_{u}$, the simulated unobservables $\{u^{h}_{it}\}$ are independently drawn for $h=1,\dots,H$, where $H$ denotes the number of simulated panel data sets. For a given value of $\theta$, let $y^{h}_{it}(\theta, \widehat{\alpha}_{i})$ denote the simulated dependent variable for simulation path $h$, then the sample log likelihood function using the $h$--th simulated data is
where $\beta$ and $\gamma_{i}$ respectively denote the finite--dimensional parameter and incidental parameter in the simulation world. The fixed effect estimator to this problem is
where $$\widehat{\gamma}_{i}(\beta, \theta, \widehat{\alpha}_{i})=\operatorname*{arg\,max}_{\gamma\in\mathbb{R}}\frac{1}{T}\sum^{T}_{t=1}\ln f(y^{h}_{it}(\theta, \widehat{\alpha}_{i}) \mid x_{it}; \beta, \gamma).$$ Repeating this estimation for all the simulated panel data, the following average can be computed: $$\widehat{\beta}_{H}(\theta, \boldsymbol{\widehat{\alpha}}):=\frac{1}{H}\sum^{H}_{h=1}\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}}),$$ and the indirect fixed effect estimator $\widetilde{\theta}^{H}$ is the solution to
where the superscript $H$ stresses that the finite--sample performance depends on the number of simulations conducted. The box below summarizes the steps required to compute the estimator.
This section starts with a discussion of the main assumptions that lead to theoretical properties of $\widehat{\theta}$ and $\widehat{\alpha}_{i}$'s. These assumptions are standard in large--$T$ panel data models HahnKuersteiner2011, and they also impose certain structures that help establish the asymptotic properties of the indirect fixed effect estimator. Additional assumptions are imposed to ensure the simulations do not affect the panel data structure.
Assumption (ref) requires that the time series dimension grows at the same rate as the cross section dimension. The assumption defines the large--$T$ asymptotics framework and allows to transform the incidental parameter problem from a consistency to a bias problem, the latter of which can be quantified. From a practitioner's point of view, if the ratio $T/n$ is not negligible, then it is reasonable to consider the large T asymptotics.
Assumption (ref) imposes independence along the cross--section dimension. Assumption (ref) imposes a weak temporal dependence on each individual time series. The quantity $\alpha_{i}(m)$ measures for each $i$ how much dependence exists between data separated by at least $m$ time periods, and a uniform bound is imposed so as to bound covariances and moments when using law of large numbers (LLN) and central limit theorem (CLT). Interested readers can refer to Section 3.4 in White2000 for definitions and properties. Note that Assumption (ref) rules out non--stationary explanatory variables such as time effects and linear trends.
Assumption (ref) is a sufficient condition that ensures the log likelihood function admits a unique maximizer based on the time series variation. This assumption allows to prove the consistency of fixed effect estimators under large--T asymptotics. The indirect fixed effect estimator also requires this assumption for consistency.
Assumption (ref) imposes compactness of parameter space, which is standard for establishing asymptotic properties of extremum estimators NeweyMcFadden1994. Compactness is convenient for proving uniform convergence with nonsmooth criterion functions Newey1991. Assumption (ref) imposes a Lipschitz condition on the log likelihood function and a moment condition on the envelope function. This allows to establish uniform law of large number (ULLN) of sample log likelihood function and hence the pointwise consistency of $\widehat{\theta}$.
Under these assumptions and some regularity conditions on the Hessian matrix, HahnKuersteiner2011 established the following two results:
where $\boldsymbol{\alpha_{0}}:=(\alpha_{10},\dots,\alpha_{n0})$. Equation ((ref)) states that the maximal deviation of $\widehat{\alpha}_{i}$ from $\alpha_{i0}$ converges to zero. This uniform consistency result is crucial for the theory of indirect fixed effect estimator because it justifies the usage of $\widehat{\alpha}_{i}$'s for data simulations. Equation ((ref)) characterizes the asymptotic relationship between $\widehat{\theta}$ and $\theta_{0}$. The term $A(\theta_{0}, \boldsymbol{\alpha_{0}})$ is the influence function that satisfies the central limit theorem (CLT) with mean zero. The term $B(\theta_{0}, \boldsymbol{\alpha_{0}})$ converges to its expected value. Therefore, $\widehat{\theta}$ is consistent, asymptotically normal, but biased. HahnKuersteiner2011 derived the analytical forms of both terms, which are complicated functions of $\theta_{0}$ and $\boldsymbol{\alpha_{0}}$.
Because the new estimator involves simulations, the following regularity condition is required so that the simulated data still maintain the mixing properties. Another regularity condition is that the parameter space in the simulation world is compact. Because $(\beta, \gamma)$ is just a change of notation from $(\theta, \alpha)$, this assumption is natural.
In sum, Assumption (ref) allows for an asymptotic representation of simulated fixed effect estimator $\widehat{\beta}^{h}(\theta, \widehat{\boldsymbol{\alpha}})$ that resembles the one for $\widehat{\theta}$, i.e., equation ((ref)).
In order for the indirect inference--type estimator to be consistent, three conditions should be satisfied GourierouxMonfortRenault1993: an invertible relationship between $\theta$ and $\beta(\theta, \boldsymbol{\alpha_{0}})$, pointwise convergence of $\widehat{\theta}$ to $\beta(\theta_{0}, \boldsymbol{\alpha_{0}})$, and uniform convergence of $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$ to $\beta(\theta, \boldsymbol{\alpha_{0}})$ over the compact parameter space $\Theta$. The first condition is satisfied because $\beta(\theta, \boldsymbol{\alpha_{0}})$ is an identity, and equation ((ref)) gives the second condition. The following proposition states the uniform convergence condition.
The current proof specializes to panel models, but it is generalizable to other models that feature concavity and smoothness in $(\beta, \gamma_{i})$. Details are available in Appendix (ref), and here the main ideas are discussed.
Proving the uniform convergence condition with nonsmoothness requires two steps: pointwise convergence of $\widehat{\beta}^{h}(\theta,\boldsymbol{\widehat{\alpha}})$ to $\beta(\theta, \boldsymbol{\alpha_{0}})$, and a stochastic equicontinuity condition as follows:
where $C$ is a constant and $\delta$ is a positive scalar.
Following the standard argument in NeweyMcFadden1994, pointwise convergence requires a ULLN result of log likelihood function using simulated data ((ref)) to the limiting log likelihood ((ref)). The log likelihood ((ref)) has two sources of randomness: the first source comes from sampling variation of observed data, and the other is from simulations of unobservables. The non--standard part, however, is that data are simulated using $\widehat{\alpha}_{i}$'s. Therefore, it is necessary to first show that ((ref)) uniformly well approximates the log likelihood using data generated by $\alpha_{i0}$'s:
where the integration is with respect to the distribution of simulation draws $u^{h}_{it}$ to eliminate randomness from simulations. The details are available in Lemma (ref), and intuition is provided here. Because panel data are independent along the cross section, it suffices to show that each individual's log likelihood: $$ \frac{1}{T}\sum^{T}_{t=1}y_{it}^{h}(\theta, \widehat{\alpha}_{i})\log\Big(\Phi(x'_{it}\beta+\gamma_{i})\Big)+(1-y_{it}^{h}(\theta, \widehat{\alpha}_{i}))\log\Big(1-\Phi(x'_{it}\beta+\gamma_{i})\Big),$$ satisfies this property. Given $\theta$, this individual log likelihood is an additive and multiplicative combination of indicator functions of scalar $\widehat{\alpha}_{i}$ and smooth functions of $(\beta, \gamma_{i})$, which belongs to classes of functions that satisfy stochastic equicontinuity VanDerVaartWellner1996. Therefore, its empirical process:
is stochastic equicontinuous. Combined with uniform consistency result of $\widehat{\alpha}_{i}$'s and LLN of $\nu_{T}(\alpha_{i0})$, an application of the triangular inequality leads to the uniform approximation result. Now that ((ref)) only has randomness from observed data, its uniform convergence to the limiting log likelihood ((ref)) follows the argument as in HahnKuersteiner2011. As such, the pointwise convergence of $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$ follows through.\footnote{Details are available in Lemma (ref).} To verify the stochastic equicontinuity condition ((ref)), note that the profiled log likelihood:
is concave in $\beta$. By definition, $\widehat{\beta}^{h}(\theta_{1}, \boldsymbol{\widehat{\alpha}})$ satisfies $\partial\widehat{Q}(\widehat{\beta}^{h}(\theta_{1}, \boldsymbol{\widehat{\alpha}}); \theta_{1})/\partial\beta=0.$ A first--order Taylor expansion with respect to $\widehat{\beta}^{h}(\theta_{1}, \boldsymbol{\widehat{\alpha}})$ around $\widehat{\beta}^{h}(\theta_{2}, \boldsymbol{\widehat{\alpha}})$ and positive--definiteness of the Hessian shows that $\widehat{\beta}^{h}(\theta_{1}, \boldsymbol{\widehat{\alpha}})-\widehat{\beta}^{h}(\theta_{2}, \boldsymbol{\widehat{\alpha}})$ is bounded by $$\Big\lvert\frac{\partial\widehat{Q}(\widehat{\beta}^{h}(\theta_{2}, \boldsymbol{\widehat{\alpha}}); \theta_{2})}{\partial\beta}-\frac{\partial\widehat{Q}(\widehat{\beta}^{h}(\theta_{2}, \boldsymbol{\widehat{\alpha}}); \theta_{1})}{\partial\beta}\Big\rvert,$$ which, by the Cauchy--Schwarz inequality, is bounded by the product of two terms: a smooth function of $(\beta, \gamma_{i})$ and
Therefore, it suffices to bound the two terms in expectation. The technical challenge mainly comes from proving this for equation ((ref)) that features nonsmooth components. Although indicator functions are well--known to have controlled complexities Andrews1994, and a similar result on the difference of indicator functions with univariate variable is given in ChenLintonVanKeilegom2003, here $\theta$'s can be multi--dimensional. It turns out that the expectation of ((ref)) satisfies the $L^{2}$--smoothness regularity condition in Andrews1994.\footnote{See proof of Proposition (ref) for details.} Therefore, the stochastic equicontinuity condition for $\widehat{\beta}^{h}(\theta, \boldsymbol{\widehat{\alpha}})$ is verified.
Armed with Proposition (ref), the consistency of $\widetilde{\theta}^{H}$ follows the arguments as in GourierouxMonfortRenault1993. The proof is straightforward because there is no need to consider a weighting matrix.
Recall that the indirect fixed effect estimator using $H$ simulations $\widetilde{\theta}^{H}$ is the solution to $\widehat{\theta}=\widehat{\beta}_{H}(\widetilde{\theta}^{H}, \boldsymbol{\widehat{\alpha}})$. Non--differentiability of $\theta\mapsto\widehat{\beta}_{H}(\theta, \boldsymbol{\widehat{\alpha}})$ means that the techniques in the indirect inference literature GourierouxMonfortRenault1993 are not applicable. The following stochastic equicontinuity assumption is imposed.
Assumption (ref) requires that the difference between $\widehat{\beta}_{H}(\theta_{1}, \boldsymbol{\widehat{\alpha}})$ and $\widehat{\beta}_{H}(\theta_{2}, \boldsymbol{\widehat{\alpha}})$ can be approximated by its expectation at a $\sqrt{nT}$ rate. Combined with consistency of $\widetilde{\theta}^{H}$ and the mean value theorem, it allows to analyze the asymptotic normality of $\widetilde{\theta}^{H}$ through the lens of fixed effect estimators as follows:
Recall that equation ((ref)) characterizes the representation of $\widehat{\theta} - \theta_{0}$. Because the same regression is run on simulated data $h$ and the likelihood is smooth in $(\beta, \gamma_{i})$, the same structure of representation arises, namely that
The terms $A^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}})$ and $B^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}})$ reflect that the data are generated using $\theta_{0}$, $\boldsymbol{\widehat{\alpha}}$ and simulated unobservables $\{u^{h}_{it}\}$. A combination of ((ref)), ((ref)) and ((ref)) therefore leads to
This equation reflects two observations. First, $\widehat{\theta}$ is unbiased if $B(\theta_{0}, \boldsymbol{\alpha_{0}})$ and $B^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}})$ both converge to the same limit. Second, $\widehat{\theta}$ is asymptotically normal if $A(\theta_{0}, \boldsymbol{\alpha_{0}})$ and $A^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}})$ converge to the same limiting distribution, but the variance is inflated by a factor of $1/H$. The rest of the section provides the main ideas of the proof.
The intuition can be gained by setting $H=1$ and considering an infeasible fixed effect estimator $\widehat{\beta}_{H}(\theta_{0}, \boldsymbol{\alpha_{0}})$, which is obtained from data simulated by $(\theta_{0}, \alpha_{0})$. Then the representation of $\widehat{\beta}_{H}(\theta_{0}, \boldsymbol{\alpha_{0}})-\theta_{0}$ takes the form
The theory of indirect inference implies that $B(\theta_{0}, \boldsymbol{\alpha_{0}})$ and $B^{h}(\theta_{0}, \boldsymbol{\alpha_{0}})$ converge to the same probability limit. Because the actual simulated data are generated by $\widehat{\alpha}_{i}$'s, it suffices to show that $B^{h}(\theta_{0}, \widehat{\alpha})$ uniformly well approximates $B^{h}(\theta_{0}, \boldsymbol{\alpha_{0}})$ such that the approximation error is asymptotically negligible. More specifically, the bias term using simulated data takes the following form, $$B^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}}) = -\Big[\frac{1}{n}\sum^{n}_{i=1}\mathcal{I}_{i}(\theta_{0}, \widehat{\alpha}_{i})\Big]^{-1} \frac{1}{n}\sum^{n}_{i=1}B^{h}_{i}(\theta_{0}, \widehat{\alpha}_{i}),$$ where $\mathcal{I}_{i}(\theta_{0}, \widehat{\alpha}_{i})$ is individual $i$'s information matrix, and it is a smooth function of all its arguments. Therefore, $\mathcal{I}_{i}(\theta_{0}, \widehat{\alpha}_{i})\xrightarrow{p}\mathcal{I}_{i}(\theta_{0}, \alpha_{i0})$ for each $i$. Each $B^{h}_{i}(\theta_{0}, \widehat{\alpha}_{i})$ is nonsmooth in $\widehat{\alpha}_{i}$, and the following assumption is imposed such that $B^{h}(\theta_{0}; \boldsymbol{\widehat{\alpha}})$ replaces $B^{h}(\theta_{0}, \boldsymbol{\alpha_{0}})$ with negligible errors:
As such, the indirect fixed effect estimator corrects the bias. HahnKuersteiner2011 derived the analytical expression of the term $A(\theta_{0}, \boldsymbol{\alpha_{0}})$. The term $A^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}})$ has the same structure, namely $$A^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}}) = \Big[\frac{1}{n}\sum^{n}_{i=1}\mathcal{I}^{h}_{i}(\theta_{0}, \widehat{\alpha}_{i})\Big]^{-1}\frac{1}{\sqrt{nT}}\sum^{n}_{i=1}\sum^{T}_{t=1}U^{h}_{it}(\theta_{0}, \widehat{\alpha}_{i}),$$ where $A^{h}_{it}(\theta_{0}, \widehat{\alpha}_{i})$ is a combination of high--order derivatives of the log likelihood. The following high--level assumption is imposed:
Under this assumption, $A^{h}(\theta_{0}, \boldsymbol{\widehat{\alpha}})$ can uniformly well approximate $A^{h}(\theta_{0}, \boldsymbol{\alpha_{0}})$ with negligible errors, and the asymptotic normality result in indirect inference literature follows through GourierouxMonfortRenault1993. Combined with Proposition (ref), the indirect fixed effect estimator is asymptotically unbiased and normal.
The estimation of the variance--covariance matrix uses the estimated Hessian matrix of the sample log likelihood function from the real data. As previously discussed in Remark (ref), the number of simulations $H$ shows up as a factor that inflates the asymptotic variance. There are two interpretations. The first is in line with other simulation--based methods: using simulations introduces an additional source of uncertainty and it is manifested through an increase in variance. The other interpretation is related to the trade--off between bias and variance. Because the indirect inference estimator debiases fixed effects estimator, the variance is larger, and it is quantified by the number of panel data simulated.
DhaeneJochmans2015 proposed a half--panel method that removes the leading bias. Intuitively, the method splits in half the panel along the time series dimension, obtains the fixed effect estimators for the half samples and applies a linear combination with respect to the full--sample fixed effect estimator. Theoretically, the method does not change the asymptotic variance because the influence function is linear. However, in finite samples the variance is inflated due to an inefficient use of data. The indirect fixed effect method is explicit about the bias--variance tradeoff in that the asymptotic variance is multiplied by $H$. However, as will be shown in the simulations and applications, this method leads to a smaller standard error compared to the half--panel bias correction method.
Research on the relationship between female labor force participation and fertility is complicated by the presence of unobserved factors that affect both decisions. Following hyslop99, this paper addresses the omitted variable issue by including individual fixed effects into the binary response panel model for the female labor force participation.
The data come from the Panel Study of Income Dynamics (PSID) and constitute a nine--year longitudinal sample spanning from 1979 to 1988. The sample includes 664 women aged 18–-60 in 1985 who were continuously married with husbands in the labor force in each of the sample periods and changed their labor force participation statuses. Consider the following static specification:
where $y_{it}$ denotes the labor force participation indicator for woman $i$ at time $t$, and $x_{it}$ denotes a vector of time--varying covariates. These covariates include numbers of children of at most 2 years of age, between 3 and 5 years of age, between 6 and 17 years of age; log of the husband's income,\footnote{This variable serves as a proxy for permanent nonlabor income hyslop99.} age and age squared. The individual effects $\alpha_{i}$'s are included to control for time--invariant unobserved heterogeneity such as willingness to work or ability.
Table ((ref)) reports estimates of index coefficients using different methods. The standard errors are reported in parentheses. The standard errors for the fixed effects are computed from the Hessian of the profiled log likelihood. IFE--1, IFE--10 and IFE--20 denote indirect fixed effect estimators with $H$ being 1, 10 and 20 respectively, and their respective standard errors are computed by multiplying the FE standard errors by $(1+\frac{1}{H})$. For comparisons, the table includes results using half--panel jackknife method (HBC) DhaeneJochmans2015, analytical bias correction (ABC) Fernandez-Val2009 and the leave--one--out jackknife method (BC--HN) HahnNewey2004. The ABC has the same standard errors as the uncorrected fixed effect estimators, while the standard error computation for BC--HN and HBC follow the descriptions in HahnNewey2004 and DhaeneJochmans2015 respectively. The results show that the uncorrected estimates of index coefficients are about 15% larger (in absolute value) than their bias-corrected counterparts, indirect fixed effect estimators are closely comparable to ABC and BC--HN, and HBC produces estimates that are larger in magnitude. Because HBC achieves bias correction through sample splitting, the standard errors are larger.
This section considers Monte Carlo simulations calibrated to the same PSID data. The details of calibration procedures are available in Appendix (ref). The indirect inference fixed effect estimator is compared with the fixed effect estimation, the ABC and two jackknife bias correction methods. All simulations are done 1000 times and $H$ is set to 10. The coverage reports the proportion of the times that $\theta_{0}$ falls within the 95% confidence interval. All of the other statistics are relative to the true parameters and multiplied by 100.
Table ((ref)) reports the simulation results of fixed effects and indirect fixed effect estimators. Fixed effect estimators are subject to a bias that is of the same order of magnitude as the standard deviation. This leads to severe under--coverage of the confidence intervals. The indirect fixed effect estimators, on the other hand, reduce bias by a margin without much inflation in the standard deviation. Therefore, the empirical coverage is close to the nominal value of $95\%$.
Table ((ref)) tabulates the simulation results of ABC and two jackknife bias correction methods. Compared with IFE, ABC features smaller biases as it removes the bias term based on a plugged--in estimate, but standard deviations are comparable. Turning to the other two methods that automatically correct bias, first note that BC--HN admits smaller biases and standard deviations than HBC. The simulation results are in line with those reported in HughesHahn2020, who theoretically showed that HBC has a larger higher--order variance and remaining bias than BC--HN.\footnote{Therefore, in practice it is recommended to use panel bootstrap to obtain standard errors for HBC.} On the other hand, IFE is comparable with BC--HN in terms of both bias and standard deviation. A theoretical exploration is left for future work.
The current theory is restricted to strictly exogenous explanatory variables, but Monte Carlo simulations in Appendix (ref) shows that the method can accommodate lagged dependent variables as well. Naturally, the next step is to extend the current theory to allow for dynamics in the DGP.
Like other bias correction methods, the theoretical properties of the indirect fixed effect estimator are predicated on the large--$T$ assumption. Therefore, to compare how the estimator performs against other methods under varying lengths of time periods, this paper follows the literature HahnNewey2004, Fernandez-Val2009, HughesHahn2020 and considers the following simulation design:
The numerical experiments consider panels with $n = \{100, 200\}$ and $T = \{4, 8, 12\}$.
Table ((ref)) reports the simulation results. The fixed effect estimators are subject to large biases, even when the time periods is 12. Because their biases are comparable to the standard deviations, fixed effect estimators exhibit undercoverage, which implies under--rejections. HBC has a poor performance when $T=4$, and this is because each of the split sample only uses 2 time periods for estimation. HBC substantially reduces the bias when $T$ is 8 or 12, but it is subject to a large dispersion, which reflects a larger confidence interval. As a result, the empirical coverage of HBC is larger than the nominal value of 95%.
The indirect fixed effect estimator has a better bias reduction performance compared to HBC, especially when $T=4$. This means that the new estimator is less sensitive to the time periods, and thus can be potentially useful for short--panel applications as well. The number of simulation paths affects the dispersion. When $H=1$, the estimator has a larger dispersion compared to the fixed effect estimator. When $H=10$, the dispersion is reduced. The trade--off for setting a large $H$ in practice is an increase in computation time. For example, when $n=200$ and $T=12$, it took roughly sixteen minutes to obtain the results for $H=10$ with ten cores on the MacBook Pro (M1, 2020). For $H=1$, it took about two minutes. It is interesting to explore algorithms that efficiently search for solutions to non--smoonth functions, but this is beyond the scope of the paper.
It is worth noting that both the application and the simulation design feature non--stationary regressors. The results provide suggestive evidence that the indirect fixed effect estimator can accommodate these variables, which do not satisfy the stationarity condition in Assumption (ref). However, the theoretical exploration is beyond the scope of this paper. In another ongoing project, the author proposes a new method called crossover jackknife that deals with non--stationarity explicitly.
Fixed effect estimations of nonlinear panel models are subject to large biases of point estimates and incorrect coverages of confidence intervals. This paper proposes a new estimator that reduces the bias and obtains standard errors without bootstrap.
There are at least three other questions for further explorations. First, average partial effects are often the quantities of interest in nonlinear models. This paper establishes theoretical properties of finite dimensional parameters, and it could be interesting to explore if they can be extended to handle average partial effects, which is a function of explanatory variables, parameters of interest and incidental parameters.
Second, this paper directly works with non--smooth log likelihood function and establishes the asymptotic properties of the new estimator. However, a practical concern of non--smoothness is that gradient--based optimization schemes cannot be used for estimation, and gradient--free schemes like Nelder--Mead face computational difficulty in high--dimensional problems. The indirect fixed effect estimator might benefit from approaches like kernel smoothing, but the theoretical justification can be nontrivial as smoothing can introduce an additional bias.
Finally, incorporating unobserved heterogeneity into the dynamic discrete choice (DDC) models is an active area of research. One popular approach treats unobserved heterogeneity as an unobserved state variable and assumes individuals can be categorized into a finite number of types KasaharaShimotsu2009, ArcidiaconoMiller2011. Introducing fixed effects circumvents the need to take a stand on the number of types, but can potentially complicate identification and estimation: the individual effects show up in both the current payoff and the continuation value, the latter of which has to be solved using a fixed--point algorithm. It would be exciting to investigate whether some of the ideas in this paper can be applied to incorporate fixed effects into DDC models.