Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
117,617 characters · 20 sections · 115 citation commands
Moment Restrictions for Nonlinear Panel Data Models with Feedback
\thispagestyle{empty}
{JEL Codes:} C23, C33
{Keywords:} {Sequential Exogeneity, Feedback, Panel Data, Duration Models, Incidental Parameters, Semiparametric Efficiency Bounds.}
\pagenumbering{arabic} \onehalfspacing
\pagenumbering{arabic} \onehalfspacing
An econometrican randomly samples units from a population of interest. For each sampled unit $i=1,\ldots,N$, let $Y_{it}$ denote a period $t=1,\ldots,T$ outcome, $X_{it}$ a corresponding vector of covariates, and $A_{i}$, a latent variable representing unmeasured unit-specific attributes. Importantly, $A_i$ is constant over time and may freely covary with the regressors, $X_{i1},\ldots,X_{iT}$. An initial condition, $Y_{i0}$, is also observed. Panel data of this type feature prominently in empirical research in economics and other fields. That panel data offers the possibility to “control for" the correlated heterogeneity, $A_{i}$, is a key attraction.
While “fixed effects” panel data methods place no restrictions (beyond mild regularity conditions) on the joint distribution of the initial condition $Y_{i0}$ and latent heterogeneity $A_i$, they generally do place strong restrictions on how the outcomes $Y_{i1},\ldots,Y_{iT}$ and regressors $X_{i1},\ldots,X_{iT}$ relate to each other. These restrictions involve more than substantive modeling assumptions; they also constrain what we will call the feedback process, whereby past outcomes and covariates $Y_{it-1},Y_{it-2},\ldots,Y_{i0},X_{it-1},X_{it-2},\ldots,X_{i1}$, as well as heterogeneity $A_i$, influence current covariates $X_{it}$.
Feedback arises naturally in many dynamic economic problems. For example, a firm's optimal investment rule typically varies with its current capital stock (and hence past investment decisions) as well as past productivity shocks (and hence its output history), see, e.g., Olley_Pakes_EM96 and Blundell_Bond_ER00. A doctor may adjust a patient's treatment protocol in a way which depends on her perceptions of their health response to past treatments robins1986new. A worker's decision to participate in job training may, as is typically the focus in evaluation studies, influence their future labor market outcomes, but participation may also depend on their past labor market experiences Ashenfelter_RESTAT1978,Ashenfelter_Card_RESTAT1985.
One approach to handling feedback, indeed the leading one in empirical work, rules it out a priori. This corresponds to maintaining the strict exogeneity assumption formulated by Chamberlain_EM1982,Chamberlain_JE82. Strict exogeneity assumptions underpin, albeit generally implicitly, many panel data based approaches to program evaluation (see ghanem2022selection on difference-in-differences methods). The overwhelming majority of nonlinear panel data estimators also require strict exogeneity arellano2001panel,arellano2011nonlinear. Strict exogeneity, while a convenient assumption for estimation, is restrictive in many economic applications. Ironically, although Chamberlain_JE82,Chamberlain_HBE84 emphasized the testable implications of strict exogeneity, today the assumption is so common as to often go unmentioned in applied work.
A different approach, pioneered by robins1986new, assumes that the feedback process is homogeneous. By homogeneous we mean that the mapping from past outcomes, $Y_{it-1},Y_{it-2},\ldots,Y_{i0}$ and covariates $X_{it-1},X_{it-2},\ldots,X_{i1}$ to the current covariate, $X_{it}$, does not vary with $A_i$: it is identical across agents. This is a powerful simplification, leading to feasible nonparametric and semiparametric estimators robins2000marginal. However, the restriction to homogeneous feedback, like strict exogeneity, is a strong assumption. It rules out, for example, a firm's investment rule varying with its persistent productivity level. Robin's robins1986new setup is often plausible in environments where the researcher controls $X_{it}$, such as in a dynamic experiment. Strict exogeneity and homogeneous feedback are non-nested assumptions; but both restrictions correspond to a subset of the data generating processes covered by our results.
In this paper we study nonlinear panel data models with unrestricted heterogeneous feedback and correlated heterogeneity. Almost 25 years ago, surveying the then extant work on nonlinear panel data analysis, arellano2001panel observed:
Arellano and Honor\'e's arellano2001panel observation remains largely true today. Chamberlain_JOE2022, in a paper first circulated in the early 1990s, studied a class of panel data models with multiplicative heterogeneity defined by sequential moment restrictions. Certain panel data count models are covered by his results Chamberlain_JBES92,wooldridge1997multiplicative, Blundell_Griffith_Windmeijer_JOE2002,Windmeijer_EPD2008. Feedback in dynamic linear panel data models with sequential moment restrictions is also well understood arellano1991some,arellano1995another,Chamberlain_JBES92,Hahn_ET1997,ai2012semiparametric. However, outside the setting studied by Chamberlain_JOE2022 that includes linear models as a special case, very little is known about panel data models with unrestricted feedback.\footnote{Buchinsky_et_al_EL2010 show how to compute semiparametric efficiency bounds in dynamic discrete choice models that feature feedback. Recently, in independent work, botosaru2024adversarialapproachidentification and chesher2024robust propose innovative approaches to static and dynamic nonlinear panel data models.}
We characterize the set of feedback and heterogeneity robust (FHR) moment conditions in nonlinear panel data models with feedback. We work in a likelihood setting where the outcome density depends on a finite-dimensional parameter, and both the feedback process and the unobserved heterogeneity distribution are unrestricted. Our results cover both the finite-dimensional parameter indexing the parametric part of the model, as well as estimands which involve averages over the distribution of unobserved heterogeneity and feedback (e.g., average partial effects, average treatment effects and other average effects). We also characterize semiparametric efficiency bounds for the common parameter and average effects. We demonstrate how these results may be used to find feasible estimating equations with good efficiency properties in practice.
To illustrate the power of our approach, we include a complete characterization of all FHR moments in the multi-spell mixed proportional hazards (MPH) model with both feedback and lagged duration dependence. heckman1980does emphasized the importance of incorporating these phenomena into duration analysis, although we are not aware of methods for doing so beyond those appearing in the unpublished dissertation of woutersen2000essays. hahn1994efficiency studied efficiency bounds in the multi-spell MPH model under strict exogeneity and no lagged duration dependence (see also Ridder_Woutersen_EM2003 for related work on estimation of MPH models).
In the next section we formally define the class of semiparametric nonlinear panel data models with feedback. Section (ref) presents our first main result: a complete characterization of the set of all feedback and heterogeneity robust (FHR) moment conditions. This extends the characterization obtained by functional differencing (Bonhomme_EM12), which requires strict exogeneity, to models with unrestricted feedback. Section (ref) provides a similar characterization for average effects. Section (ref) presents the semiparametric efficiency bound analysis. There we demonstrate that the orthogonal complement of the nuisance tangent set coincides with the set of all FHR estimating equations. This result has important implications for efficient estimation. Specifically, we show how to construct FHR moment functions that have good efficiency properties, leading to locally efficient estimators in the sense of newey1990semiparametric. Throughout we use the MPH model to illustrate key results in a concrete setting. Our MPH results are novel and of independent interest. Finally, in Section (ref), we touch on a number of important additional issues, including existence of FHR moments, regularization, and additional examples (several of which are novel). A version of the paper that combines the main text, appendices, and all supplementary materials in a single file is available [\href{https://kevindano.github.io/assets/files/BDG_feedback.pdf}{here}].
\paragraph{Notation.}
In what follows we generally suppress the $i$ subscript when referring to a single random draw from the cross-sectional population. Hence, for example, $Y_t$ denotes the period $t$ outcome of a randomly sampled unit and $A$ its unobserved, time-invariant, attribute. We let $Z^{t}=\left(Z_{t},Z_{t-1},\ldots\right)$ denote the entire observed history of $Z_{t}$ and $Z^{s:t}=\left(Z_{s},\ldots,Z_{t}\right)$ its history from periods $s\leq t$ to $t$.
In this section we introduce the semiparametric panel data model with feedback, connect this model to the more restrictive one which maintains strict exogeneity, and formally state our main research questions. Throughout this section, and those that follow, we illustrate key results in the context of a multi-spell mixed proportional hazards (MPH) model with lagged duration dependence and feedback.
Let $\left\{ \left(X_{i1},\ldots,X_{iT},Y_{i0},Y_{i1},\ldots,Y_{iT},A_{i}\right)\right\} _{i=1}^{\infty}$ be an independently and identically distributed random sequence drawn from some distribution function $F$. The sole prior restriction on $F$ is that the conditional density of $Y_{t}$ at $y_{t}$ given the past $y^{t-1}$, regressor history $x^{t}$, and latent unit-specific heterogeneity $a$, belongs to a known parametric family indexed by the unknown parameter $\theta\in\Theta\subset\mathbb{R}^{K}:$
for some $\theta\in\Theta$. The density $f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)$ is the parametric component of our setup.\footnote{While ((ref)) imposes that only the contemporaneous regressor and the first lag of the outcome matter, additional lags could be easily accommodated in what follows.}
Familiar nonlinear examples of (ref) include binary choice logit models Chamberlain_ReStud80,chamberlain2010binary,bonhomme2023identification, honore2024moment and count models Chamberlain_JBES92,Chamberlain_JOE2022, wooldridge1997multiplicative. Specific forms for $f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)$ also arise in the context of dynamic structural models Aguirregabiria_Mira_JOE2010. Observe that the parametric families, $f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)$ and $f_{\theta}\left(\left.y_{s}\right|y_{s-1},x_{s},a\right)$ need not coincide for $s \neq t$; this allows for time effects and other forms of nonstationarity.
Returning to our general setup, sequentially factorizing the joint density of $Y_{0},Y_{1},\ldots,Y_{T},X_{1},\ldots,X_{T},A$ at $y_{0},y_{1},\ldots,y_{T},x_{1},\ldots,x_{T},a$ yields the following expression for the likelihood contribution of a single unit:
where the second equality follows by imposing the parametric assumption (ref) and establishing the following notations:
While $f_{\theta}\left(\left.y_{t}\right|y^{t-1},x^{t},a\right)$ belongs to a parametric family, the remaining components, items (i) to (iii) above, are all unrestricted. We call this model the semiparametric panel data model with feedback.
In what follows we call the $T-1$ densities $g\left(\left.x_{T}\right|y^{T-1},x^{T-1},a\right)$, $g\left(\left.x_{T-1}\right|y^{T-2},x^{T-2},a\right)$, ..., $g\left(\left.x_{2}\right|y^{1},x_{1},a\right)$ the feedback process. This process describes how past values of the outcome influence current regressor values. Because it depends on $A$, the feedback process is heterogeneous across units. We do not impose any stationarity over $t=1,\ldots,T$, so $g\left(\left.x_{t}\right|y^{t-1},x^{t-1},a\right)$ and $g\left(\left.x_{s}\right|y^{s-1},x^{s-1},a\right)$ for $s \neq t$ may differ arbitrarily. When $X_t$ is a policy variable, such flexibility accommodates un-modeled regime shifts (e.g., as when a change in tax policy alters firm investment behavior). We use the notation $g$ to denote a generic element of the set of all allowable feedback processes $\mathcal{G}$. We call $\pi\left(\left.a\right|y_{0},x_{1}\right)$ the heterogeneity distribution; this density describes the distribution of unobserved heterogeneity, $A$, across units as well as any dependence of this heterogeneity on the {initial condition}, $\left(y_{0},x_{1}\right)$. Let $\pi$ denote a generic element of the set of all allowable heterogeneity distributions, $\Pi$. Finally we let $\nu\in\mathcal{N}$ denote an element of the set of all possible initial condition densities.
Although not emphasized in our exposition, it is straightforward to incorporate strictly exogenous regressors into the feedback model. Similarly, additional sources of heterogeneity, beyond $A$, can enter the feedback process. This generality, while important in some applications, involves no new issues and clutters notation.\footnote{To be specific: let $W^{T}$ contain all leads and lags of a vector of strictly exogeneous regressors and $B$ an additional source of heterogeneity. The results that follow are easily modified to accommodate the richer model:
where we assume that only the contemporaneous value of $W_{t}$ enters the parametric part of the model for simplicity.} We also note that, although not emphasized in most of our examples, $Y_t$ may be vector-valued in some settings.
Panel data models with feedback and heterogeneity arise frequently in economic applications. In structural dynamic choice models, agents' dynamic optimization typically leads to a likelihood function of the form ((ref)), where $Y_t$ contains choice variables of interest, as well as payoff variables, and $X_t$ contains dynamic state variables and un-modeled choice variables; see Aguirregabiria_Mira_JOE2010 for an exposition in the case of models with discrete outcomes. The moment conditions we derive are robust to any possible process for the dynamic state variables and distribution of unobserved heterogeneity.
Feedback also arises in program evaluation settings. Let $Y_t$ denote earnings and $X_t$ recent past participation in a job-training program. Ashenfelter_RESTAT1978 observed that Manpower Development and Training Act (MDTA) trainees had unusually low earnings in the year prior to undertaking training, consistent with a behavioral model where poor labor market outcomes (low values of $Y_{t-1}$) induced agents to seek out training in the next period ($X_{t}=1$). In contrast, maintaining no feedback would require that, conditional on the latent attribute, $A$, a worker's labor market history has no bearing on the decision to undertake training; a rather strong assumption Ashenfelter_Card_RESTAT1985.
In this paper we study estimation of the common parameter, $\theta$, in panel data models when (ref) is the only prior restriction on $F$ (except for some mild regularity conditions). We wish to construct estimators that are consistent irrespective of the precise instances of the feedback process, $g$, heterogeneity distribution, $\pi$, and initial condition, $\nu$, which describe the sampled data. In what follows we call such estimators feedback and heterogeneity robust (FHR). Because strict exogeneity obtains as a special case, any FHR estimator remains valid under strict exogeneity. FHR estimators are natural extensions of familiar “fixed effects” approaches to estimation in panel data models without feedback. We fully characterize the set of moment-based FHR estimators. We also derive semiparametric efficiency bounds for $\theta$. In particular, our results allow for a precise quantification of the information loss associated with accommodating unrestricted feedback relative to maintaining strict exogeneity.
One alternative to FHR estimation involves assuming that both the feedback process, $g$, and heterogeneity density, $\pi$, belong to parametric families indexed by some parameter vector $\eta$. With these additional maintained assumptions, the econometrician can then maximize the resulting likelihood, conditional on $Y_{0}$ and $X_{1}$, with respect to both $\theta$ and $\eta$. This transforms the problem from a semiparametric to a parametric one. Consistency of such parametric “random effects” estimators typically requires that the additionally maintained parametric restrictions on $g$ and $\pi$ hold in the sampled population. Consequently, such estimators are not generally feedback and heterogeneity robust.
We next formally define the FHR property. Let $\theta\in\Theta$, and $\omega=\left(g,\pi,\nu\right)\in\Omega$ collect all the nuisance parameters in our model. We assume that all elements $\omega\in\Omega$ have a common support, known to the econometrician (which may be unbounded and include the full real line, for example, for the support of the heterogeneity $A$). Let $\mathbb{E}_{\theta,\omega}\left[\cdot\right]$ denote an expectation taken under the DGP at $\left(\theta,\omega\right)$. Let $\phi_{\theta}\left(y_{0},y_{1},\ldots,y_{T},x_{1},\ldots,x_{T}\right)$ be a function of the observed data indexed by $\theta$. We say that $\phi_{\theta}$ is a FHR moment function if, for all $\omega$ such that $\phi_{\theta}$ is absolutely integrable under DGP $\left(\theta,\omega\right)$, we have
The main goal of the paper is to derive moment restrictions $\phi_{\theta}$ that have the FHR property. We will first provide a complete characterization of FHR moment restrictions for $\theta$ in Section (ref).\footnote{In fact, our characterization of FHR moment functions remains valid if $\theta$ is infinite-dimensional (for example, if $\theta$ contains nonparametric elements such as functions). However, the finite-dimensional $\theta$ case contains many important models, and all the examples we will use as illustrations feature a finite-dimensional $\theta$ vector. We also focus on this case in our analysis of efficiency.} Then, in Section (ref) we will show how essentially the same analysis can be used to provide a characterization of FHR moment functions for average effects of the form $$\mu(\theta,\omega)=\mathbb{E}_{\theta,\omega}\left[h_{\theta}\left(Y^{T},X^{T},A\right)\right],$$ where $h_{\theta}\left(Y^{T},X^{T},A\right)$ is a known function of $(Y^{T},X^{T},A)$ indexed by $\theta$. Average effects include a variety of causal or structural parameters, such as average partial effects, as in Chamberlain_HBE84, and average structural functions, as in Blundel_Powell_WC03.
Since any FHR moment function is also valid under strict exogeneity, the set of all FHR moment conditions for a given panel data model will be a subset of the corresponding set of moment conditions derived under strict exogeneity (and characterized in Bonhomme_EM12). In some cases the latter set may be non-trivial and the former empty. For example, chamberlain2010binary showed point identification of the panel logit model under strict exogeneity, while bonhomme2023identification show a failure of identification in this model under feedback (see also Section (ref) below).
This section presents our first main result: a characterization of all FHR moment conditions for $\theta$. As an application, we also specialize our results to provide a constructive characterization of the set of all possible FHR moment conditions for the MPH model.
We begin by characterizing the set of FHR moments for $\theta$.
\vskip .3cm
As an implication of Part (A) of Theorem (ref), suppose the data is generated according to some $(\theta_0,\omega_0)$. Then any function $\phi_{\theta}$ that satisfies (ref) and (ref) and is absolutely integrable under the population DGP ($\theta_0,\omega_0$) has zero mean at the true $\theta_0$, that is, $$\mathbb{E}_{\theta_0,\omega_0}[\phi_{\theta_0}(Y^T,X^T)]=0.$$ To see this, simply apply Part (A) of Theorem (ref) with $\theta=\theta_0$ and $\omega=\omega_0$.
The first condition for the FHR property, restriction (ref), ensures robustness of $\phi_\theta$ to the presence of heterogeneity of an unknown form. This condition is highlighted in Bonhomme_EM12, in a setting with strictly exogenous covariates, as the key condition ensuring valid moment functions for $\theta$. Under strict exogeneity, (ref) ensures that $\phi_{\theta}(Y^T,X^T)$ is conditionally mean zero given $X^T$ and $A$. By the law of iterated expectations, this suffices to ensure that it is unconditionally mean zero as well. The functional differencing approach then provides a general recipe for constructing functions $\phi_{\theta}$ satisfying (ref).
To illustrate how the presence of feedback modifies the interpretation of condition (ref), consider the $T=2$ setting. In this case, the condition is $$\int\left[\int \phi_{\theta}\left(y_0,y_1,y_2,x_1,x_2\right)f_{\theta}(y_2\,|\, y_1,x_2,a)dy_2\right]f_{\theta}(y_1\,|\, y_0,x_1,a)dy_1=0,$$ which coincides with
Here the inner expectation corresponds to an average over $y_2$ with respect to its model density, $f_{\theta}(y_{2}\,|\, y_{1},x_{2},a)$. This inner expectation conditions on $X_2=x_2$, while the outer expectation, which averages over $y_1$ with respect to its model density, $f_{\theta}(y_{1}\,|\, y_{0},x_{1},a)$, does not condition on $x_2$. This is a key difference between the heterogeneous feedback case, considered here, and the setting with strict exogeneity studied by Bonhomme_EM12. Under strict exogeneity, $Y_1$ is independent of $X_2$ conditional on $(Y_0,X_1,A )$, so ((ref)) coincides with $$\mathbb{E}\left[\mathbb{E}\left[\left.\phi_{\theta}\left(Y_0,Y_1,Y_2,X_1,X_2\right)\right|Y_0,Y_1,X_1,X_2,A\right]\,|\,Y_0,X_1,X_2,A \right]=0,$$ with both the inner and outer expectations conditioning on $X_2$. By iterated expectations, this conditional expectation implies a zero mean condition given $A$, initial conditions, and the entire sequence of covariates, $$\mathbb{E}\left[\left.\phi_{\theta}\left(Y_0,Y_1,Y_2,X_1,X_2\right)\right|Y_0,X_1,X_2,A\right]=0.$$ However, under feedback it is not generally the case that ((ref)) can be written as a zero mean condition given the covariates sequence $(X_1,X_2)$.
In the presence of feedback, Theorem (ref) additionally requires the $T-1$ conditions (ref), to ensure that $\phi_\theta$ is a valid moment function. To explain this additional requirement, consider again the $T=2$ setting, in which case (ref) reads $$\int \phi_{\theta}\left(y_0,y_1,y_2,x_1,x_2\right)f_{\theta}(y_2\,|\, y_1,x_2,a)dy_2\quad \text{does not depend on }x_2,$$ which corresponds to the single mean independence restriction
Equation (ref) implies that, at the the population parameter, $X_2$ does not predict $\phi_{\theta}\left(Y_0,Y_1,Y_2,X_1,X_2\right)$ conditional on $Y_0,Y_1,X_1,A$. Under this condition, which we interpret as “feedback robustness” of $\phi_{\theta}$, the conditioning on $X_2=x_2$ disappears in ((ref)), which implies
ensuring, by iterated expectations, that $\phi_{\theta}$ is a valid moment function. Importantly, $\phi_\theta$ is then a valid moment under both unrestricted heterogeneity and unrestricted feedback.
Lastly, Theorem (ref) also establishes that (ref) and (ref) are, under suitable conditions, necessary for the FHR property. To show the necessity property, we assume quadratic mean differentiability at $(\theta,\omega^*)$ as stated in Part (B) Condition (i). This is a standard regularity condition in semiparametric estimation, see for example van2000asymptotic. Part (B) Condition (ii), which imposes square-integrability of the moment function in a neighborhood of $\omega^*$, is similarly standard; e.g., see the assumptions for the generalized information equality in Lemma 5.4 of newey1994large.\footnote{One can show that, if $\phi_{\theta}$ is bounded, then the necessity of (ref) and (ref) obtains directly without reference to $\omega^*$ or differentiability in quadratic mean.} We will rely on differentiability in quadratic mean in our analysis of efficiency in Section (ref).
The following corollary to Theorem (ref) facilitates the construction of FHR moment functions in practice.
\vskip .3cm
To understand Corollary (ref), it is helpful to return to the $T=2$ setting. In this case, letting $$\psi_{\theta,1}(Y_0,Y_1,X_1,A)=\mathbb{E}[\phi_{\theta}(Y_0,Y_1,Y_2,X_1,X_2,A)\,|\, Y_0,Y_1,X_1,A],$$ it follows from ((ref)) that $$\psi_{\theta,1}(Y_0,Y_1,X_1,A)=\mathbb{E}[\phi_{\theta}(Y_0,Y_1,Y_2,X_1,X_2,A)\,|\, Y_0,Y_1,X_1,X_2,A],$$ which implies ((ref)), whereas it follows from ((ref)) that $$\mathbb{E}[\psi_{\theta,1}(Y_0,Y_1,X_1,A)\,|\, Y_0,X_1,A]=0$$ (a requirement for $\psi_{\theta,1}$ given in the corollary).
The representation provided by Corollary (ref) suggests a systematic recipe for constructing FHR moment functions. The first step involves finding functions $\psi_{\theta,t}(Y^{t},X^t,A)$ that are mean zero conditional on $Y^{t-1},X^t,A$ for $t=1,\ldots,T-1$. This is straightforward since any function of $(Y^{t},X^t,A)$, suitably de-meaned, automatically fulfills this requirement (see Section (ref) for several examples). Then, given a collection of $\psi_{\theta,t}(Y^{t},X^t,A)$ functions, the second step requires solving a linear integral equation to recover a valid moment function $\phi_{\theta}$. Indeed, ((ref)) can equivalently be written as
which is known as an inhomogeneous Fredholm equation of the first kind. Note that the integral operator on the left-hand side of (ref) is known given the parameter $\theta$. Solution methods for this type of linear integral equation are the subject of a large literature engl1996regularization,carrasco2007linear. In Section (ref) we will show how to construct FHR moment functions in some specific models using Corollary (ref), and discuss when such functions exist.
A special case of Corollary (ref) obtains when one is able to find functions $\eta_{\theta,t}$ such that
for some function $b$ of the heterogeneity and initial conditions. Then, the instrumented first difference
satisfies ((ref)), for $\psi_{\theta,T-1}(y^{T-1},x^{T-1},a)=[b\left(y_0,x_1,a\right)-\eta_{\theta,T-1}\left(y^{T-1},x^{T-1}\right)]\cdot m(y^{T-2},x^{T-1})$ and $\psi_{\theta,t}=0$ for all $t<T-1$ (for an arbitrary function $m$). This provides a FHR moment function on $\theta$. More generally, one can check that
all satisfy ((ref)), thus providing additional moments.
We now illustrate this particular recipe with the MPH model.
Moments of the form ((ref)) were considered by arellano1991some in the linear model context, by Chamberlain_JBES92,Chamberlain_JOE2022 and wooldridge1997multiplicative for Poisson models, and by Al-Sadoon_et_al_ER2017 for certain binary choice models. However, this family of estimating equations does not exhaust all available FHR moments, and may lead to estimators with low levels of asymptotic precision in practice. By comparison, Theorem (ref) and Corollary (ref) characterize all available FHR moments, as we will now illustrate in the case of the MPH model. Moreover, as we will establish in later sections, our characterization can be used to derive efficient estimators (i.e., we extend to nonlinear models efficiency arguments which appear in the linear panel data literature such as in arellano1995another,arellano2001panel, Arellano_RIE2016).
For specific models it is sometimes possible to use Theorem (ref) to provide a direct characterization of all FHR moment functions. We show how this can be done in the MPH model with feedback in Lemma (ref) below. Our result provides insight into the MPH model and simplifies the derivation of new FHR moments.
Our characterization makes use of several special features of the MPH model, which we introduce first (details can be found in Supplemental Appendix (ref)). Let $P_{\theta,t}=\rho_{\theta}\left(Z_t\right)$. As indicated in (ref), conditional on $Y_{0},X_{1},A$, the $P_{\theta,t}$ for $t=1,\ldots,T$ are independent exponential random variables; each with a common rate parameter of $e^{A}$.
Next, we define a one-to-one transformation of the vector $P_{\theta}^T$ into a “forward orthogonal deviations” part and a “between” part, as follows:
Observe that $\widetilde{P}_{\theta,t}$ involves the ratio of $\rho_{\theta}\left(Z_{t}\right)$ to the sum of itself and the future values of $\rho_{\theta}\left(Z_{s}\right)$ for $s=t+1,\ldots,T$. Lemma (ref) in Supplemental Appendix (ref) establishes that, conditionally on $Y_{0},X_{1},A$:
where, throughout, $\theta$ is the parameter indexing the DGP. Lemma (ref) additionally establishes that the elements of $\left(\widetilde{P}_{\theta,1},\ldots,\widetilde{P}_{\theta,T-1},\overline{P}_{\theta}\right)$ are mutually independent of one another.
Transformation (ref) can be thought of as a MPH-specific analog of the forward orthogonal deviations transformation used by arellano1995another in the context of linear panel data models with predetermined regressors. It has the property that the random variables $\widetilde{P}_{\theta,t}$ are independent of both contemporaneous and lagged predetermined covariates as well as lagged values of the spell outcomes $\{(X_{is},Y_{is-1})\}_{s=1}^t$ (see part (i) of Lemma (ref)). Appealing to the same analogy, we can think of $\overline{P}_{\theta}$ as containing the “between” variation in $P_{\theta}^T$.
With these preliminaries in place, we can state the following FHR moment characterization for the MPH model with feedback.
\vskip .3cm
The proof of Lemma (ref) is available in Supplemental Appendix (ref). It represents a useful simplification, induced by the special structure of the MPH model, relative to the general result of Theorem (ref). Note in particular that the latent heterogeneity $A$ does not appear in the characterization of Lemma (ref). To illustrate, consider the $T=2$ setting. In this case the lemma implies that all FHR moment functions are of the form $$\phi_{\theta}\left(Y_{0},Y_1,Y_2,X_1,X_2\right)=\psi_{\theta,1}(Y_{0},\widetilde{P}_{\theta,1},\overline{P}_{\theta},X_1),$$ where $\mathbb{E}\left[\left.\psi_{\theta,1}(Y_{0},\widetilde{P}_{\theta,1},\overline{P}_{\theta},X_1)\right|Y_{0},\overline{P}_{\theta},X_1\right]=0$. We seek functions $\psi_{\theta,1}$ of the forward orthogonal deviations $\widetilde{P}_{\theta,1}$, that are mean zero conditional on the between variation $\overline{P}_{\theta}$, as well as the initial condition $\left(Y_{0},X_{1}\right)$. It is straightforward to construct such functions by de-meaning, and we will exploit this property in our analysis of efficiency.
For the general case of an arbitrary number $T$ of periods, $\phi_{\theta}$ is the sum of functions $\psi_{\theta,t}\left(Y_{0}, \widetilde{P}_{\theta}^{t},\overline{P}_{\theta},X^{t}\right)$, for $t=1,\ldots,T-1$, that are mean independent of the between variation, $\overline{P}_{\theta}$, current and past values of the predetermined regressors, $X^{t}$, and past values of the spell outcomes, $Y^{t-1}$. This structure is reminiscent of how moment conditions are typically constructed for linear panel data models with predetermined regressors (see, especially, arellano1995another).
In this section we characterize all FHR moment conditions for average effects, $$\mu(\theta,\omega)=\mathbb{E}_{\theta,\omega}\left[h_{\theta}\left(Y^{T},X^{T},A\right)\right],$$ where $h_{\theta}\left(Y^{T},X^{T},A\right)$ is a known function of $Y^{T},X^{T},A$ indexed by $\theta$.
In nonlinear panel data models, knowledge of $\theta$ does not suffice to identify the effect of an external manipulation of a regressor's value on the probability distribution of the outcome. This follows from the nonseparable way in which the unobserved heterogeneity enters such models. In contrast, estimands which average over the marginal distribution of $A$, such as (ref), do provide easy-to-interpret summaries of such effects.
Until recently the identifiability of such averages was not well understood. Recent work by honore2006bounds, chernozhukov2013average, aguirregabiria2021identification, davezies2021identification, dobronyi2021identification, pakel2023bounds and others, however, has shown that average effects are (partially) identified in several specific settings of interest. The study of average effects in more general settings, such as those with feedback, as we consider here, remains underdeveloped bonhomme2023identification.
Our first result is an analog of Theorem (ref) for $\mu(\theta,\omega)$.
\vskip .3cm
Note the strong parallel between this theorem and Theorem (ref). As in the case of $\theta$, (ref) and (ref) imply that, under absolute integrability, the true value $\mu_0=\mu(\theta_0,\omega_0)$ satisfies a moment condition. Indeed, if $\mathbb{E}_{\theta_0,\omega_0}\left[\abs{\varphi_{\theta_0}(Y^T,X^T)}\right]<\infty$ and $\mathbb{E}_{\theta_0,\omega_0}\left[\abs{h_{\theta_0}(Y^T,X^T,A)}\right]<\infty$, then Part (A) implies $$\mathbb{E}_{\theta_0,\omega_0}[\varphi_{\theta_0}(Y^T,X^T)]=\mu_0.$$ This moment condition has the FHR property: (ref) ensures that $\varphi_{\theta}$ is robust to an unknown distribution of unobserved heterogeneity, whereas the $T-1$ conditions (ref) endow $\varphi_{\theta}$ with robustness to heterogeneous feedback of an unknown form.
We additionally state the following corollary, which mimics Corollary (ref) and suggests a systematic recipe for constructing functions $\varphi_{\theta}$.
\vskip .3cm
The conditional hazard function \[ \lambda\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)=\lambda_{\alpha}(y_{t})e^{x_{t}'\beta+\gamma y_{t-1}+a} \] gives the instantaneous exit rate of a unit, at duration $y_{t}$, with lagged duration $y_{t-1}$, beginning-of-spell covariate $x_{t}$, and latent attribute $a$. Unfortunately, although easily identified, the observed hazard function \[ \lambda\left(\left.y_{t}\right|y_{t-1},x_{t}\right)=\lambda_{\alpha}(y_{t})e^{x_{t}'\beta+\gamma y_{t-1}}\mathbb{E}\left[\left.e^{A}\right|Y_{t}>y_{t},Y_{t-1}=y_{t-1},X_{t}=x_{t}\right] \] suffers from spurious duration dependence (see Lancaster_EATD1990,Heckman_AER1991): units with higher values of $A$ will exit earlier, implying that the mean $\mathbb{E}\left[\left.e^{A}\right|Y_{t}>y_{t},Y_{t-1}=y_{t-1},X_{t}=x_{t}\right]$ declines with $y_{t}$. In contrast, the average structural hazard (ASH),
which equals the (expected) hazard function for a randomly sampled unit when externally assigned lagged duration $y_{t-1}$ and covariate $x_{t}$, does not suffer from heterogeneity bias.
Since $\left.\overline{P}_{\theta}\right|Y_{0},X_{1},A\sim\mathrm{Gamma}\left(T,e^{A}\right)$ we have
for $\delta>-T$. Using (ref) with $\delta=-1$ shows that $\mathbb{E}\left[e^{A}\right]=\mathbb{E}\left[\frac{T-1}{\overline{P}_{\theta}}\right]$, so
We have thus found one FHR moment function for the ASH. Now, for any other FHR moment function for the ASH, $\varphi_{\theta}$, we have that $$\phi_{\theta}(Y^T,X^T)=\varphi_{\theta}(Y^T,X^T)-\lambda_{\alpha}(y_{t})e^{x_{t}'\beta+\gamma y_{t-1}}\frac{T-1}{\overline{P}_{\theta}}$$ is a FHR moment function for $\theta$ (where here $y_t,y_{t-1},x_t$ are fixed values at which the ASH is evaluated). Since all such moment functions $\phi_{\theta}$ are fully characterized in the MPH case by Lemma (ref), this characterizes all FHR moment functions for the ASH. It turns out, as we will shortly see, that estimation based upon the sample analog of (ref) is efficient when $\theta$ is replaced by an efficient estimate.
Theorems (ref) and (ref) characterize the complete set of moment conditions available for estimating $\theta$ and $\mu(\theta,\omega)$. When this set is nonempty, estimation at parametric rates may be (and often is) feasible. Moreover, many valid moment restrictions may be available (see Lemma (ref) for the case of the MPH model). In such settings semiparametric efficiency bound theory provides a useful tool for selecting specific moments for estimation purposes.
In this section we derive the form of the efficient scores for $\theta$ and $\mu(\theta,\omega)$ and, consequently, their corresponding semiparametric efficiency bounds. As we shall see, the characterizations presented earlier facilitate the derivation of these bounds. We illustrate the application of our results via a detailed analysis of efficiency bounds for the MPH model.\footnote{hahn1994efficiency derived the information bound for $\theta$ in the single-spell case studied by elbers1982true and Heckman_Singer_ReStud1984. He also derived the bound for the multi-spell case under strict exogeneity (see also chamberlain1985heterogeneity). Relative to prior work, our analysis covers the case with feedback and, additionally, considers average effects.}
We begin with a brief review of basic concepts in semiparametric efficiency theory; see, for example, van2000asymptotic and newey1990semiparametric. Let $\theta_0$, $g_0$, $\pi_0$, and $\nu_0$ denote, respectively, the common parameter, feedback process, heterogeneity distribution, and initial condition prevailing in the sampled population. Let $\omega=(g,\pi,\nu)\in\Omega$, with true value $\omega_0=(g_0,\pi_0,\nu_0)$, denote the nonparametric components of the model.\footnote{To characterize the efficiency bound for $\theta$, by ancillarity it is sufficient to consider the conditional density given $(y_{0},x_{1})$ (see, e.g., hahn1994efficiency).} A regular parametric submodel is defined by a likelihood function for a single random draw, $\ell(\theta,\omega_{\eta}\,|\,y^T,x^{T})$, where $\omega_{\eta_0}=\omega_0$ for some scalar $\eta_0$. The likelihood satisfies mean-square differentiability of its square root with respect to $(\theta,\eta)$, with its information matrix nonsingular. The semiparametric variance bound is the supremum of the Cramer Rao bounds for $\theta$ over all such regular parametric submodels.
Let $S^{\theta}$ denote the score for $\theta$, for a submodel evaluated at $\theta=\theta_0$ and $\eta=\eta_0$: $$S^{\theta}(Y^T,X^T)=\frac{\partial \ln \ell(\theta_0,{\omega}_0\,|\, Y^T,X^T)}{\partial \theta},$$ where we leave the dependence of $S^{\theta}$ on $(\theta_0,\omega_0)$ implicit. Likewise, let $S^{\eta}$ denote the score for $\eta$. The nonparametric tangent set $\mathcal{T}_{\theta_0,\omega_0,K}$ is the mean-square closure of the $K\times 1$ linear combinations of scores $S^{\eta}$ across all regular parametric submodels (where $K$ is the dimension of $\theta$). Let $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$ denote the orthocomplement of $\mathcal{T}_{\theta_0,\omega_0,K}$ in the Hilbert space of square-integrable mean-zero functions with inner product $\left<s_1,s_2\right>=\mathbb{E}_{\theta_0,\omega_0}\left[s_1(Y^T,X^T)'s_2(Y^T,X^T)\right]$. By definition this set consists of all $K\times 1$ elements $\phi_{\theta}$ such that $\left<\phi,s\right>=0$ for all $s\in \mathcal{T}_{\theta_0,\omega_0,K}$.
The next theorem provides a characterization of the orthocomplement of the tangent set, $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$, which is key to the analysis of efficiency in our context.
\vskip .3cm
An implication of Theorem (ref) is that, if $\phi_{\theta_0} \in \mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$, then it is also an element of the orthogonal complement of the nuisance tangent set associated with any other data generating process $(\theta_0, \omega_*)$ with $\omega_* \neq \omega_0$ (i.e., we also have $\phi_{\theta_0} \in \mathcal{T}_{\theta_0,\omega_*,K}^{\perp}$). This follows from the fact that $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ consists of the set of FHR moments characterized in Theorem (ref) earlier; a set which does not vary with $\omega_0$. Indeed, it is precisely this feature of the model which makes feedback and heterogeneity robust estimation possible. Knowledge, whether a priori or up to sampling error, of the form of the feedback process and/or heterogeneity distribution is not required for consistent estimation. This is a crucial feature of the class of models we study in this paper.
One subtlety is that, although the elements of $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ do not depend on the precise instance of $\omega_{0}$, the definition of the reference Hilbert space does depend on it. Let $\phi_{\theta_{0}}$ be a valid FHR moment that is absolutely integrable under $(\theta_0,\omega^*)$. Then, while $\phi_{\theta_{0}}(Y^{T},X^{T})$ has zero mean under both $(\theta_0,\omega_0)$ and $(\theta_0,\omega^*)$, in general its variance differs under $\omega_0$ and $\omega^*$. Consequently, the ability to precisely estimate $\theta_{0}$ using a particular $\phi_{\theta}$ generally varies with the population feedback process and heterogeneity distribution, although the validity of $\phi_{\theta}$ as a moment function does not. This connects to the discussion of locally efficient estimation below.
To understand Theorem (ref) it is helpful to consider the implications of restricting $\omega_0$ such that it belongs to a parametric family (indexed by, say, $\eta$). This is the approach taken in, for example, parametric random-effects analysis Chamberlain_ReStud80, chamberlain1985heterogeneity. In that setting the residualized score, $\widetilde{S}^{\theta}=S^{\theta}-\mathbb{E}\left[S^{\theta}S^{\eta\prime}\right]\times \mathbb{E}\left[S^{\eta}S^{\eta\prime}\right]^{-1}S^{\eta}$, will generally vary with $\eta$: consistent estimation of $\theta$ requires knowledge of $\eta$ (up to sampling error), and it requires correct specification of the parametric models of feedback and heterogeneity. This is not required in our approach; indeed (elements of) $\omega$ may even be unidentified, while $\theta$ remains $\sqrt{N}$-estimable.
Theorem (ref) is reminiscent of the situation which arises in average treatment effect (ATE) estimation under unconfoundedness with a known propensity score robins1994estimation,hahn1994efficiency. In that setting the set of consistent estimating equations for the ATE does not depend on the form of the conditional distribution of the potential outcomes given covariates. Theorem (ref) extends prior work for the case of strictly exogenous regressors showing that $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$ is characterized by moment functions that have zero means conditional on all covariates, initial conditions, and heterogeneity (see, e.g., hahn1994efficiency for the MPH model, and dano2023transition for dynamic logit models).
Of course, in many semiparametric estimation problems, consistent estimation of (features of) the nonparametric model component is a requirement for consistent estimation of $\theta$. Examples include the binary choice model with random utility components drawn from an unknown distribution independent of the regressors newey1990semiparametric and ATE estimation under unconfoundedness with an unknown propensity score.
\paragraph{Efficient score for $\theta$.} Under regularity conditions,\footnote{Namely that $\ell(\theta,\omega|y^T,x^{T})$ is smooth in a neighborhood of $(\theta_0,\omega_0)$, and that the information matrix is nonsingular (see for example Theorem 3.2 in newey1990semiparametric).} the semiparametric variance bound for $\theta_0$ is equal to the inverse of the variance of the efficient score
where $\Pi\left(.\,|\,\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}\right)$ denotes the orthogonal projection onto $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$. Note that the projection is well defined since $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ is closed and linear.\footnote{An equivalent and more frequent formulation of the efficient score is $\phi_{\theta_0}^{\rm eff}=S^{\theta}-\Pi\left(S^{\theta}\,|\,\mathcal{T}_{\theta_0,\omega_0,K}\right)$ (e.g., hahn1994efficiency), which is interpreted as the residual from the population regression of $S^{\theta}$ on the nuisance tangent set. The equivalent formulation based on $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$ is convenient to work with in our setting; see VanDerLaan_Robin_Book2003.} As a result, the efficient moment restriction for $\theta_0$ is
The efficient score $\phi_{\theta_0,\omega_0}^{\rm eff}$ in ((ref)) is defined through a projection onto the orthocomplement $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$, which we have fully characterized in Theorem (ref). Below we show, via examples, how to use this observation to calculate -- whether analytically or numerically -- efficient scores in models with unknown heterogeneity and feedback.
\paragraph{Efficient score for $\mu$.} We now turn to an analysis of efficiency for average effects. Suppose there exists a moment function $\varphi_{\theta_0}$ that identifies an $L\times 1$ average effect of interest $\mu(\theta_0,\omega_0)$ given by\footnote{Standard implicit smoothness conditions are required, namely that, for a regular parametric submodel, $\sup_{(\theta,\eta) \in \mathcal{N}} \mathbb{E}_{\theta,\omega_{\eta}}\left[\norm{\varphi_{\theta}(Y^T,X^T)}^2\right]<\infty$ in a neighborhood $\mathcal{N}$ of $(\theta_0,\eta_0)$, such that a generalized information equality holds (see brown1998efficient).}
and suppose that $\theta_0$ is identified from (ref). Let $${\varphi}^{\rm eff}_{\theta_0,\omega_0}(Y^T,X^T)=\varphi_{\theta_0}(Y^T,X^T)-\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L}^{\perp}),$$ where the orthocomplement of the tangent set, $\mathcal{T}_{\theta_0,\omega_0,L}^{\perp}$, is given by Theorem (ref), except for the fact that the relevant dimension is $L$ instead of $K$.
Theorem 1 in brown1998efficient shows that the joint efficient moment restrictions for $\theta_0$ and $\mu_0=\mu(\theta_0,\omega_0)$ are given by (ref) and
The semiparametric variance bound for $\mu_0$ equals the lower $L\times L$ block of the asymptotic variance of the joint GMM estimator $(\widehat{\theta},\widehat{\mu})$ based on (ref) and (ref).
The construction of $\varphi^{\rm eff}_{\theta_0,\omega_0}$ relies on a function $\varphi_{\theta_0}$ that satisfies (ref). In some models one can find such a function (a Riesz representer), which does not depend on the nonparametric component $\omega_0$. This can be done by exploiting Theorem (ref). See, for example, the discussion of the average structural hazard function in the MPH earlier. Moreover, $\varphi^{\rm eff}_{\theta_0,\omega_0}$ is not affected by the particular choice of $\varphi_{\theta_0}$.\footnote{This follows from noting that $\varphi_{\theta_0}(Y^T,X^T)-\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L}^{\perp})=\mu_{0}+\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L})$, and that, if $\varphi_{\theta_0}^{1}$ and $\varphi_{\theta_0}^{2}$ both satisfy (ref), then $\Pi(\varphi_{\theta_0}^{1}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L})=\Pi(\varphi_{\theta_0}^{2}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L})$ as we show in Supplemental Appendix Lemma (ref).} Of course not all average effects will have non-zero efficiency bounds, but when a Riesz representer for an effect of interest is available, the bound can be calculated using extant results about expectations brown1998efficient.
Although the characterization of feasible moment conditions provided by Theorem (ref) is invariant to the specific instance of $\omega_0$ indexing the sampled population, the projection (ref) generally does vary with $\omega_0$. Consequently, although knowledge of $\omega_0$ is not required for consistent estimation, it is generally valuable for improving asymptotic precision. Moreover, constructing an estimator which attains the semiparametric efficiency bound for all possible feedback processes, $g_0$, heterogeneity distributions, $\pi_0$, and initial conditions, $\nu_0$ (i.e., for all $\omega_0 \in \Omega$) generally requires nonparametrically estimating features of these model components. This may be practically difficult, or even impossible, in short panels as considered here.
An alternative approach involves constructing locally efficient estimators newey1990semiparametric, Graham_Pinto_Egel_ReStud12. Let $\widetilde{\omega}=(\widetilde{g},\widetilde{\pi},\widetilde{\nu})\in\Omega$ denote candidate “working models" for the feedback process, heterogeneity distribution, and initial condition. We show how to construct method-of-moments estimators that (i) attain the bound for $\theta_0$ (or $\mu_0$) when these working models “happen to characterize the sampled population" (i.e., $\widetilde{\omega}=\omega_0$, but this is not part of the prior restriction) and (ii) remain $\sqrt{N}$-consistent irrespective of the true $\omega_0$ characterizing the sampled population (i.e., when $\widetilde{\omega}\neq\omega_0$). A key property in our setting is that consistency holds for arbitrary $\widetilde{\omega}$ (i.e., our working models may be misspecified), only subject to regularity conditions.
Given working models $\widetilde{\omega}$, let $$\widetilde{S}^{\theta}(Y^T,X^T)=\frac{\partial \ln \ell(\theta_0,\widetilde{\omega}\,|\, Y^T,X^T)}{\partial \theta}$$ denote the score for $\theta_0$. Next define the counterpart, under the working models, to the efficient score $\phi_{\theta_0}^{\rm eff}$ for $\theta_0$ as $$\widetilde{\phi}_{\theta_0,\widetilde{\omega}}^{\rm eff}(Y^T,X^T)=\widetilde{\Pi}\left(\widetilde{S}^{\theta}(Y^T,X^T)\,|\, \mathcal{T}_{\theta_0,\widetilde{\omega},K}^{\perp}\right),$$ where $\widetilde{\Pi}$ denotes the projection operator under the working models, that is,
We next proceed similarly for $\mu(\theta,\omega)$: the counterpart to the efficient score $\varphi_{\theta_0,\omega_0}^{\rm eff}$ is
Finally, consider the following moment restrictions for $\theta_0$ and $\mu_0=\mu(\theta_0,\omega_0)$:
where we note that the expectations are taken under the true DGP $(\theta_0,\omega_0)$.
We can now state the following result.
Theorem (ref) articulates a locally efficient approach to estimation. The moment functions $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ have the FHR property, irrespective of whether the working models used to derive them actually characterize the sampled population. However, if $\widetilde{\omega}=\omega_0$ “happens to hold" in the sampled population, then (ref) and (ref) coincide with the efficient moment restrictions for $\theta_0$ and $\mu_{0}$ (when $\widetilde{\omega}=\omega_0$ is not part of the prior restriction used to calculate the efficiency bound).
The functions $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ are FHR because calculation (ref) returns an element in the orthogonal complement of the nuisance tangent set by construction. Calculation (ref) provides a principled way to select a particular FHR moment, one that is optimal if the working model happens to hold in the sampled population. Heuristically, the method-of-moments estimator based upon (ref) and (ref) will be more precise when the working models are “approximately true", but this -- to repeat -- is not required for consistency.
In practice $\widetilde{\omega}$ may be a fixed set of working models chosen by the researcher. Alternatively the researcher may posit that these models belong to parametric families $\omega_{\eta}$ indexed by an unknown finite-dimensional (not necessarily scalar) parameter $\eta$. These models may be misspecified, in the sense that there may not exist any $\eta_{0}$ such that $\omega_{\eta_0}=\omega_0$. Nevertheless the moments (ref) and (ref) will be valid for any $\eta$. There are different strategies for picking a particular $\eta$. First, the researcher may choose a particular (non-stochastic) $\eta$ via introspection. Second, she might maximize the likelihood under the working models with respect to $\theta$ and $\eta$. While the resulting estimate of $\theta$ will generally be inconsistent, the corresponding estimate of $\eta$ can be used to define the working models $\widetilde{\omega}$ under which $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ are calculated.
A third approach is to select $\eta$ by maximizing an empirical counterpart to the information for $\theta_0$, as in lindsay1985using. Let $\widehat{\eta}$ be such an estimator of $\eta$, and let $\eta^*$ be its large-$N$ probability limit. If one defines $\widetilde{\omega}=\omega_{{\eta}^*}$, then Theorem (ref) implies that ((ref))-((ref)) are satisfied at true parameter values $\theta_0$ and $\mu_{0}$ (under absolute integrability of the functions). This suggests that the GMM estimators $\widehat{\theta}$ and $\widehat{\mu}$ based on (ref) and (ref) that uses $\omega_{\widehat\eta}$ in lieu of $\widetilde{\omega}$ is consistent and asymptotically normal under standard conditions. We leave details about efficient estimation, using working models of increasing dimensions (i.e., “sieves”), to future work.
Lastly, it is important to stress that our approach based on working models is fundamentally different from (parametric) random-effects maximum likelihood estimation. Indeed, plugging in misspecified parametric models $\omega_{\eta}$ into the likelihood function ((ref)), and maximizing that likelihood with respect to $\theta$ and $\eta$, generally results in an inconsistent estimator of $\theta_0$. In contrast, the moment restrictions (ref) and (ref) remain valid even when the working models are (globally) misspecified.
In this section we specialize the general results presented above in order to analyze semiparametric efficiency in the MPH model. We focus on the $T=2$ special case in what follows.
Efficient score for $\mathbf{\theta}$. Using Lemma (ref) and Theorem (ref), we show in Supplemental Appendix (ref) that, for \( T = 2 \), the efficient score for $\theta$ in the MPH model with feedback equals
With some additional algebra we show that the efficient score for the \( \beta \) subvector is
The first term in (ref) does not involve the feedback process and resembles the efficient score for $\beta$ under strict exogeneity originally derived by hahn1994efficiency:
The second and third terms of (ref), in contrast, are specific to the feedback case, involving averages over the second-period covariate $X_{2}$. More generally, the efficient score for $\theta$ under strict exogeneity equals:
In the presence of feedback and latent heterogeneity, $X_2$ is an endogenous variable and cannot be conditioned on. \\
Efficient estimation of the ASH. An interesting average effect in the context of the MPH is the average structural hazard (ASH) defined in (ref). The latter is identified by the FHR moment function in (ref), namely $\varphi_{\theta_0}(Y^T,X^T)=\lambda_{\alpha_{0}}(y_{t})e^{x_{t}'\beta_{0}+\gamma_{0} y_{t-1}}\frac{T-1}{\overline{P}_{\theta_{0}}}$. Applying Lemma (ref) in the Supplemental Appendix, which characterizes projections onto the orthocomplement of the tangent set in the MPH model, one can readily show that $\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L}^{\perp})=0$. It then follows from the discussion in Section (ref) that the efficient moment function for the ASH is
In turn, a semiparametrically efficient estimator of the ASH is
where $\widehat{\theta}=(\widehat{\alpha},\widehat{\beta}',\widehat{\gamma})'$ is semiparametrically efficient for $\theta$.\footnote{Note that this coincides with the method of moments estimator of brown1998efficient when the condition \( \varphi_{\theta_0}(Y^T, X^T) - \mu_0 \in \mathcal{T}_{\theta_0,\omega_0,L} \) holds, where \( \mu_0 =\mu(\theta_0,\omega_0)\) denotes the average effect of interest and \( \varphi_{\theta_0} \) is an identifying moment function. This condition is satisfied in the case of the ASH for the choice $\varphi_{\theta_0}(Y^T,X^T)=\lambda_{\alpha_{0}}(y_{t})e^{x_{t}'\beta_{0}+\gamma_{0} y_{t-1}}\frac{T-1}{\overline{P}_{\theta_{0}}}$.}\\
In this subsection we summarize our findings from two numerical experiments designed to (i) illustrate the efficiency loss associated with accommodating feedback, and (ii) assess the performance of locally efficient estimators based on working models. Our context is the MPH model. In both experiments we impose a Weibull baseline hazard, set $T=2$, and fix the common parameter at $\theta_{0}=(\alpha_{0},\gamma_{0},\beta_{0})=(\frac{3}{4},\frac{3}{4}\ln 2, -\frac{1}{10})$. Our experiments are meant to numerically approximate the asymptotic precision of various methods of estimation, not to assess the accuracy of such approximations in small samples. Details on implementation can be found in Supplemental Appendix (ref).
The initial duration is drawn from an exponential distribution: $Y_{0}\sim \text{Exponential}(\frac{3}{2})$, and the first-period covariate is a randomized binary treatment: $X_{1}\sim \text{Bernoulli}(\frac{1}{2})$. The heterogeneity distribution equals $V= e^{A}\sim \text{Gamma}(5,5)$, independent of $Y_0,X_1$. The second-period covariate, $X_2$, is a Bernoulli random variable with success probability $p(Y_0,Y_1,X_1,V)$, specified differently across the two experiments to reflect alternative assumptions about the DGP:
One goal of our experiments is to assess the efficiency loss that arises when a researcher wishes to accommodate the possibility of heterogeneous feedback, but no such feedback is actually present in the sampled population. Put differently, this exercise gives us a sense of the benefit, in terms of asymptotic precision, of using the strict exogeneity assumption when it is valid (as is the case for the DGP in Experiment (A)).
We compare the asymptotic standard errors of two GMM estimators: the first uses the efficient score under strict exogeneity (ref), while the second uses the efficient score under feedback (ref).
The close connection between our FHR moment characterization (Theorem (ref)) and the relevant semiparametric efficiency bound theory (Theorem (ref)), raises interesting practical questions regarding estimation. While many possible FHR moments are available in the MPH setting, the precision with which they recover $\theta$ varies with the population values of the feedback process and heterogeneity distribution.
The complex form of the efficient score (those for the baseline hazard parameter, $\alpha$, and the coefficient on the lagged duration, $\gamma$ -- both reported in Supplemental Appendix (ref) -- are particularly complicated), suggests that crafting a globally efficient estimator would be difficult. As a principled, yet practical, alternative, we instead explore the properties of a locally efficient estimator based upon simple working models for $\widetilde{\omega}=(\widetilde{g},\widetilde{\pi})$ (a model for $\nu$ is not needed in this case).
Our chosen working models are deliberately rudimentary, intended to illustrate how a researcher might build parsimonious working models while retaining favorable efficiency propertie in more realistic settings. Specifically, in contrast to what prevails in the sampled population, the working model for the feedback process maintains that $X_{2}\sim\mathrm{Bernoulli}(p)$, for a constant $p$. Observe that the working model for the feedback process involves no feedback. For the heterogeneity distribution, we calculate the locally efficient score under a vague Gamma prior of $\widetilde{\pi}(v) = \frac{1}{v} \mathds{1}\{v > 0\}$, independent of initial conditions.
Putting all these pieces together yields the following locally efficient score for $\beta$:
which is simpler than its optimal counterpart (ref) yet, of course, still feedback and heterogeneity robust. Additional details, along with the full expression for the score $\widetilde{\phi}_{\theta_0,\widetilde{\omega}}^{\rm eff}(Y_0,\widetilde{P}_{\theta_0,1}, \overline{P}_{\theta_0}, X_1)$, are provided in Supplemental Appendix (ref). \\ As a final point of comparison, we also report the limiting standard errors of a just-identified GMM estimator that employs the moment function:
The first entry in (ref) is a mean-zero function that exploits the fact that $\widetilde{P}_{\theta_0,1}\sim U[0,1]$ and leverages symmetry (it is also, coincidentally, a component of the efficient score for $\alpha$ under the Weibull baseline hazard, derived and presented in Supplemental Appendix (ref)). The second and third entries correspond to the log-transformed analogs of (ref). While this set of moment conditions lacks an overt efficiency justification, it reflects a common strategy of choosing a small set of “simple" moments for estimation purposes. We include it primarily to benchmark the efficiency gains provided by the locally efficient approach based on working models.
Table (ref) reports the asymptotic standard errors for each estimate of $\theta_0$ in Experiment (A). We make several observations. First, comparing the first and second rows reveals that accommodating feedback -- when it is, in fact, absent from the DGP -- results in some efficiency loss for the slope coefficient $\beta$, with a 29% increase in standard error. Strict exogeneity is a strong assumption and imposing it, when it is valid to do so, improves asymptotic precision.
Although allowing for feedack degrades the precision with which we can learn $\beta$, this is not really the case for $\alpha$ (the Weibull baseline hazard parameter) and $\gamma$ (the lagged duration dependence parameter) in design (A). This finding is consistent with the structure of the efficient scores: those for $\alpha$ and $\gamma$ are very similar under both strict exogeneity and feedback.\footnote{Compare (ref) to (ref) and (ref) to (ref) in Supplemental Appendix (ref).} By contrast, the efficient score for $\beta$ differs markedly across the two settings.
A second observation, inspecting the third row of the table, is that the precision loss associated with using the locally efficient estimator based on the working models $\widetilde{\omega}$ is moderate. Recall that our working models do not characterize the sampled population in design (A). For $\beta$ and $\gamma$, respectively, we observe a 28% and 26% increase in standard error when comparing rows 2 and 3. The efficiency losses are concentrated on the coefficients for the predetermined covariate and the lagged dependent variable, with only minimal deterioration for the parameter of the baseline hazard $\alpha$.
Finally, the approach based on working models leads to large improvements relative to using the “simple" moment functions (ref). As seen in row 4, the standard errors associated with the simple GMM estimator are substantially larger for all three parameters. For example, the standard error for $\gamma$ is nearly 7 times higher than that obtained using the estimator based on the working models $\widetilde{\omega}$ (compare rows 3 and 4, column 3).
In Experiment (B), we repeat our analysis, but for a population where heterogeneous feedback is, in fact, present. Accordingly, Table (ref) compares the asymptotic standard errors of the locally efficient estimator based on $\widetilde{\phi}_{\theta_0,\widetilde{\omega}}^{\rm eff,\beta}$ to that of the “simple” GMM estimator based on ((ref)) under the second DGP described above (an estimate based on the efficient score under strict exogeneity would not be consistent in this design). As a benchmark, we use the globally semiparametrically efficient estimator based upon the true efficient score that uses (ref). The fourth column of Table (ref) also reports the corresponding asymptotic standard errors for the average structural hazard (ASH) when using the efficient moment function presented in the previous subsection. The key takeaways mirror those of Table (ref). First, the efficiency loss from relying on locally efficient scores is modest, both for common parameters and for average effects. Second, and perhaps more importantly, locally efficient estimators significantly outperform the alternative of using a set of “simple" moments, reaffirming the practical advantages of approaches based on working models.
Our characterizations can be used to find FHR moment functions for many models. We have already analyzed the MPH model in detail. In this section we provide additional analytical examples and discuss how to obtain moment functions more generally. Given that the mathematical structure for model parameters and average effects is similar, in this section we focus our discussion on $\theta$.
For certain models, it may be that the only solution to the system of equations in Theorem (ref), or equivalently in Corollary (ref), is the degenerate one, $\phi_{\theta}=0$. We now provide two examples where only trivial moment functions exist. For simplicity we focus on the $T=2$ case.
It is important to note that, even when there exist non-zero functions $\phi_{\theta}$, $\theta$ may fail to be identified. For example, in the MPH model the covariates may be collinear, in which case identification fails. This is of course not specific to our setting. As in any nonlinear GMM problem, identification needs to be verified on a case-by-case basis, and while rank conditions for local identification of $\theta$ are available, verifying global identification may be difficult. Conversely, it may also be that the only $\phi_{\theta}$ satisfying the conditions of Theorem (ref) is $\phi_{\theta}=0$, yet $\theta$ is point-identified.\footnote{As an example, let $T=2$ and let $Y_{t}=\theta+A X_{t}+\varepsilon_{t}$ with $X_{t}$ continuously distributed on $\mathbb{R}$ and $\varepsilon_{t}\,|\, Y^{t-1},X^t,A \sim {\cal{N}}(0,1)$. Applying a similar logic to that in model (ref), one can show that $\phi_{\theta}=0$ since the assumptions imply that $P(X_{2}=0)=P(X_{1}=0)=0$. However, this is a case where the parameter of interest $\theta$ is identified at 0 since $\lim_{x \to 0} \mathbb{E}[Y_{1}|X_{1}=x]=\theta$ (Graham_Powell_EM12).} However, in that case an implication of our analysis in Section (ref) is that such identification is necessarily irregular and the semiparametric efficiency bound for $\theta$ is zero Chamberlain_JE1986. Lastly, even if point-identification fails identified sets may be informative, as shown in lee2020identification and bonhomme2023identification.
We now illustrate how researchers can apply the two-step procedure underlying Corollary (ref) to derive new moment conditions. In the first step, we construct a function $\psi_{\theta}=\sum_{t=1}^{T-1}\psi_{\theta,t}$, for $\psi_{\theta,t}$ such that
In the second step, in the final time period, we invert the linear integral equation
to get a FHR moment function $\phi_{\theta}(y^T,x^T)$. Naturally, this last task requires the function $\psi_{\theta}(y^{T-1},x^{T-1},a)$ to lie in the range of the integral operator induced by the parametric model. In many examples of interest, equation ((ref)) can actually be inverted in closed form, yielding explicit expressions for functions $\phi_{\theta}$. As an initial example, consider the Poisson regression model introduced earlier.
Example (ref) illustrates how, using Corollary (ref), one can derive moment functions by operator inversion. When suitable functions $\psi_{\theta}$ exist, closed-form inversion delivers explicit moment functions. In other settings, it may be that the inverse is not available in closed form, and numerical inversion techniques need to be used (see, e.g., engl1996regularization).
In this section we have described an approach, based on Corollary (ref), to construct moment functions $\phi_{\theta}$ when those are available. However, for those functions to be helpful for estimation they need to be sufficiently regular. In this last part we provide examples that show how irregularity may arise, and how regularization techniques can help.
The regularization strategy we have outlined in the context of the MPH model can be useful in more complex models. A general strategy, when solving for $\phi_{\theta}$ in the integral equation ((ref)), is to use a regularized inverse of the relevant integral operator as in carrasco2007linear. We now describe an example where this strategy can be successfully applied.