EconBase
← Back to paper

Moment Restrictions for Nonlinear Panel Data Models with Feedback

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,617 characters · 20 sections · 115 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Moment Restrictions for Nonlinear Panel Data Models with Feedback

\thispagestyle{empty}

abstract{Many panel data methods, while allowing for general dependence between covariates and time-invariant agent-specific heterogeneity, place strong a priori restrictions on feedback: how past outcomes, covariates, and heterogeneity map into future covariate levels. Ruling out feedback entirely, as often occurs in practice, is unattractive in many dynamic economic settings. We provide a general characterization of all feedback and heterogeneity robust (FHR) moment conditions for nonlinear panel data models and present constructive methods to derive feasible moment-based estimators for specific models. We also use our moment characterization to compute semiparametric efficiency bounds, allowing for a quantification of the information loss associated with accommodating feedback, as well as providing insight into how to construct estimators with good efficiency properties in practice. Our results apply both to the finite dimensional parameter indexing the parametric part of the model as well as to estimands that involve averages over the distribution of unobserved heterogeneity. We illustrate our methods by providing a complete characterization of all FHR moment functions in the multi-spell mixed proportional hazards model. We compute efficient moment functions for both model parameters and average effects in this setting.}

{JEL Codes:} C23, C33

{Keywords:} {Sequential Exogeneity, Feedback, Panel Data, Duration Models, Incidental Parameters, Semiparametric Efficiency Bounds.}

\pagenumbering{arabic} \onehalfspacing

\pagenumbering{arabic} \onehalfspacing

An econometrican randomly samples units from a population of interest. For each sampled unit $i=1,\ldots,N$, let $Y_{it}$ denote a period $t=1,\ldots,T$ outcome, $X_{it}$ a corresponding vector of covariates, and $A_{i}$, a latent variable representing unmeasured unit-specific attributes. Importantly, $A_i$ is constant over time and may freely covary with the regressors, $X_{i1},\ldots,X_{iT}$. An initial condition, $Y_{i0}$, is also observed. Panel data of this type feature prominently in empirical research in economics and other fields. That panel data offers the possibility to “control for" the correlated heterogeneity, $A_{i}$, is a key attraction.

While “fixed effects” panel data methods place no restrictions (beyond mild regularity conditions) on the joint distribution of the initial condition $Y_{i0}$ and latent heterogeneity $A_i$, they generally do place strong restrictions on how the outcomes $Y_{i1},\ldots,Y_{iT}$ and regressors $X_{i1},\ldots,X_{iT}$ relate to each other. These restrictions involve more than substantive modeling assumptions; they also constrain what we will call the feedback process, whereby past outcomes and covariates $Y_{it-1},Y_{it-2},\ldots,Y_{i0},X_{it-1},X_{it-2},\ldots,X_{i1}$, as well as heterogeneity $A_i$, influence current covariates $X_{it}$.

Feedback arises naturally in many dynamic economic problems. For example, a firm's optimal investment rule typically varies with its current capital stock (and hence past investment decisions) as well as past productivity shocks (and hence its output history), see, e.g., Olley_Pakes_EM96 and Blundell_Bond_ER00. A doctor may adjust a patient's treatment protocol in a way which depends on her perceptions of their health response to past treatments robins1986new. A worker's decision to participate in job training may, as is typically the focus in evaluation studies, influence their future labor market outcomes, but participation may also depend on their past labor market experiences Ashenfelter_RESTAT1978,Ashenfelter_Card_RESTAT1985.

One approach to handling feedback, indeed the leading one in empirical work, rules it out a priori. This corresponds to maintaining the strict exogeneity assumption formulated by Chamberlain_EM1982,Chamberlain_JE82. Strict exogeneity assumptions underpin, albeit generally implicitly, many panel data based approaches to program evaluation (see ghanem2022selection on difference-in-differences methods). The overwhelming majority of nonlinear panel data estimators also require strict exogeneity arellano2001panel,arellano2011nonlinear. Strict exogeneity, while a convenient assumption for estimation, is restrictive in many economic applications. Ironically, although Chamberlain_JE82,Chamberlain_HBE84 emphasized the testable implications of strict exogeneity, today the assumption is so common as to often go unmentioned in applied work.

A different approach, pioneered by robins1986new, assumes that the feedback process is homogeneous. By homogeneous we mean that the mapping from past outcomes, $Y_{it-1},Y_{it-2},\ldots,Y_{i0}$ and covariates $X_{it-1},X_{it-2},\ldots,X_{i1}$ to the current covariate, $X_{it}$, does not vary with $A_i$: it is identical across agents. This is a powerful simplification, leading to feasible nonparametric and semiparametric estimators robins2000marginal. However, the restriction to homogeneous feedback, like strict exogeneity, is a strong assumption. It rules out, for example, a firm's investment rule varying with its persistent productivity level. Robin's robins1986new setup is often plausible in environments where the researcher controls $X_{it}$, such as in a dynamic experiment. Strict exogeneity and homogeneous feedback are non-nested assumptions; but both restrictions correspond to a subset of the data generating processes covered by our results.

In this paper we study nonlinear panel data models with unrestricted heterogeneous feedback and correlated heterogeneity. Almost 25 years ago, surveying the then extant work on nonlinear panel data analysis, arellano2001panel observed:

quoteThe main limitation of much of the literature on nonlinear panel data methods is that it is assumed that the explanatory variables are strictly exogenous in the sense that some assumptions will be made on the errors conditional on all (including future) values of the explanatory variables.

Arellano and Honor\'e's arellano2001panel observation remains largely true today. Chamberlain_JOE2022, in a paper first circulated in the early 1990s, studied a class of panel data models with multiplicative heterogeneity defined by sequential moment restrictions. Certain panel data count models are covered by his results Chamberlain_JBES92,wooldridge1997multiplicative, Blundell_Griffith_Windmeijer_JOE2002,Windmeijer_EPD2008. Feedback in dynamic linear panel data models with sequential moment restrictions is also well understood arellano1991some,arellano1995another,Chamberlain_JBES92,Hahn_ET1997,ai2012semiparametric. However, outside the setting studied by Chamberlain_JOE2022 that includes linear models as a special case, very little is known about panel data models with unrestricted feedback.\footnote{Buchinsky_et_al_EL2010 show how to compute semiparametric efficiency bounds in dynamic discrete choice models that feature feedback. Recently, in independent work, botosaru2024adversarialapproachidentification and chesher2024robust propose innovative approaches to static and dynamic nonlinear panel data models.}

We characterize the set of feedback and heterogeneity robust (FHR) moment conditions in nonlinear panel data models with feedback. We work in a likelihood setting where the outcome density depends on a finite-dimensional parameter, and both the feedback process and the unobserved heterogeneity distribution are unrestricted. Our results cover both the finite-dimensional parameter indexing the parametric part of the model, as well as estimands which involve averages over the distribution of unobserved heterogeneity and feedback (e.g., average partial effects, average treatment effects and other average effects). We also characterize semiparametric efficiency bounds for the common parameter and average effects. We demonstrate how these results may be used to find feasible estimating equations with good efficiency properties in practice.

To illustrate the power of our approach, we include a complete characterization of all FHR moments in the multi-spell mixed proportional hazards (MPH) model with both feedback and lagged duration dependence. heckman1980does emphasized the importance of incorporating these phenomena into duration analysis, although we are not aware of methods for doing so beyond those appearing in the unpublished dissertation of woutersen2000essays. hahn1994efficiency studied efficiency bounds in the multi-spell MPH model under strict exogeneity and no lagged duration dependence (see also Ridder_Woutersen_EM2003 for related work on estimation of MPH models).

In the next section we formally define the class of semiparametric nonlinear panel data models with feedback. Section (ref) presents our first main result: a complete characterization of the set of all feedback and heterogeneity robust (FHR) moment conditions. This extends the characterization obtained by functional differencing (Bonhomme_EM12), which requires strict exogeneity, to models with unrestricted feedback. Section (ref) provides a similar characterization for average effects. Section (ref) presents the semiparametric efficiency bound analysis. There we demonstrate that the orthogonal complement of the nuisance tangent set coincides with the set of all FHR estimating equations. This result has important implications for efficient estimation. Specifically, we show how to construct FHR moment functions that have good efficiency properties, leading to locally efficient estimators in the sense of newey1990semiparametric. Throughout we use the MPH model to illustrate key results in a concrete setting. Our MPH results are novel and of independent interest. Finally, in Section (ref), we touch on a number of important additional issues, including existence of FHR moments, regularization, and additional examples (several of which are novel). A version of the paper that combines the main text, appendices, and all supplementary materials in a single file is available [\href{https://kevindano.github.io/assets/files/BDG_feedback.pdf}{here}].

\paragraph{Notation.}

In what follows we generally suppress the $i$ subscript when referring to a single random draw from the cross-sectional population. Hence, for example, $Y_t$ denotes the period $t$ outcome of a randomly sampled unit and $A$ its unobserved, time-invariant, attribute. We let $Z^{t}=\left(Z_{t},Z_{t-1},\ldots\right)$ denote the entire observed history of $Z_{t}$ and $Z^{s:t}=\left(Z_{s},\ldots,Z_{t}\right)$ its history from periods $s\leq t$ to $t$.

Panel data models with feedback

In this section we introduce the semiparametric panel data model with feedback, connect this model to the more restrictive one which maintains strict exogeneity, and formally state our main research questions. Throughout this section, and those that follow, we illustrate key results in the context of a multi-spell mixed proportional hazards (MPH) model with lagged duration dependence and feedback.

Setup

Let $\left\{ \left(X_{i1},\ldots,X_{iT},Y_{i0},Y_{i1},\ldots,Y_{iT},A_{i}\right)\right\} _{i=1}^{\infty}$ be an independently and identically distributed random sequence drawn from some distribution function $F$. The sole prior restriction on $F$ is that the conditional density of $Y_{t}$ at $y_{t}$ given the past $y^{t-1}$, regressor history $x^{t}$, and latent unit-specific heterogeneity $a$, belongs to a known parametric family indexed by the unknown parameter $\theta\in\Theta\subset\mathbb{R}^{K}:$

equation[equation omitted — 246 chars of source]

for some $\theta\in\Theta$. The density $f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)$ is the parametric component of our setup.\footnote{While ((ref)) imposes that only the contemporaneous regressor and the first lag of the outcome matter, additional lags could be easily accommodated in what follows.}

Familiar nonlinear examples of (ref) include binary choice logit models Chamberlain_ReStud80,chamberlain2010binary,bonhomme2023identification, honore2024moment and count models Chamberlain_JBES92,Chamberlain_JOE2022, wooldridge1997multiplicative. Specific forms for $f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)$ also arise in the context of dynamic structural models Aguirregabiria_Mira_JOE2010. Observe that the parametric families, $f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)$ and $f_{\theta}\left(\left.y_{s}\right|y_{s-1},x_{s},a\right)$ need not coincide for $s \neq t$; this allows for time effects and other forms of nonstationarity.

example{(Mixed Proportional Hazards (MPH) Model)} Let $\left\{Y_{t}\right\}_{t=1}^{T}$ denote a sequence of durations, or spell lengths, with $X_{t}$ a corresponding vector of beginning-of-spell covariates; $Y_{0}$ is an “initial duration”. For example, $Y_{t}$ might equal time to re-arrest following an agent's $t^{th}$ release from prison while $X_{t}$ might include measures of their post-release support and supervision. Since support and supervision, $X_{t}$, might depend on the previous time to re-arrest, $Y_{t-1}$, as well as unmeasured agent attributes, $A$, heterogeneous feedback is plausible. The MPH model was introduced by Lancaster_EM1979 and Nickell_EM1979 for single-spell data. chamberlain1985heterogeneity studied identification and estimation with multi-spell data under strict exogeneity. hahn1994efficiency derived the semiparametric efficiency bound for both the single- and multi-spell case (the latter under strict exogeneity). heckman1980does emphasize the relevance of lagged duration dependence and feedback in labor market applications. The instantaneous conditional hazard rate is given by \begin{equation} \lambda\left(\left.y_{t}\right|Y^{t-1}=y^{t-1},X^{t}=x^{t},A=a\right)=\lambda_{\alpha}\left(y_{t}\right)\exp\left(\gamma y_{t-1}+x_{t}'\beta+a\right), \end{equation} with $\lambda_{\alpha}\left(y_{t}\right)$ a known parametric family of baseline hazard functions indexed by $\alpha$ (we emphasize the Weibull case with $\lambda_{\alpha}\left(y_{t}\right)=\alpha y_{t}^{\alpha-1}$ below). Under (ref) the conditional density at $Y_{t}=y_{t}$ is \begin{equation} f_{\theta}\left(\left.y_{t}\right|y^{t-1},x^{t},a\right)=\lambda_{\alpha}\left(y_{t}\right)\exp\left(\gamma y_{t-1}+x_{t}'\beta+a\right)\exp\left(-\rho_{\theta}\left(z_{t}\right)e^a\right), \end{equation} for $z_t=(y_t,y_{t-1},x_t')'$, $\rho_{\theta}\left(z_{t}\right)=\Lambda_{\alpha}\left(y_{t}\right)\exp\left(\gamma y_{t-1}+x_{t}'\beta\right)$, $\theta=\left(\alpha',\beta',\gamma\right)'$, and $\Lambda_{\alpha}\left(y_{t}\right)=\int_{0}^{y_{t}}\lambda_{\alpha}\left(u\right)\mathrm{d}u$ the integrated baseline hazard.
example{(Poisson Model)} Let $\left\{ Y_{t}\right\} _{t=1}^{T}$ denote a sequence of counts, for example the number of patents awarded to a firm in a year, as in Blundell_Griffith_Windmeijer_JOE2002, and $X_{t}$ a vector of time-varying regressors. For $t=1,\ldots,T$ we have \begin{equation} \left.Y_{t}\right|Y^{t-1},X^{t},A\sim\mathrm{Poisson}\left(\exp\left(\gamma Y_{t-1}+X_{t}'\beta+A\right)\right), \end{equation} with $Y_{0}$ an initial count. Chamberlain_JBES92 and wooldridge1997multiplicative proposed GMM estimators for $\theta=\left(\beta',\gamma \right)'$ in this model. Windmeijer_EPD2008 reviews extant results.

Returning to our general setup, sequentially factorizing the joint density of $Y_{0},Y_{1},\ldots,Y_{T},X_{1},\ldots,X_{T},A$ at $y_{0},y_{1},\ldots,y_{T},x_{1},\ldots,x_{T},a$ yields the following expression for the likelihood contribution of a single unit:

align[align omitted — 805 chars of source]

where the second equality follows by imposing the parametric assumption (ref) and establishing the following notations:

enumerate$g\left(\left.x_{t}\right|y^{t-1},x^{t-1},a\right)$, denotes the density of $X_{t}$ at $x_{t}$ given the past $\left(y^{t-1},x^{t-1}\right)$ and heterogeneity $a$; • $\pi\left(\left.a\right|y_{0},x_{1}\right)$, the conditional density of $A$ at $a$ given the initial condition $\left(y_{0},x_{1}\right)$; and • $\nu\left(y_{0},x_{1}\right)$, the initial condition density.

While $f_{\theta}\left(\left.y_{t}\right|y^{t-1},x^{t},a\right)$ belongs to a parametric family, the remaining components, items (i) to (iii) above, are all unrestricted. We call this model the semiparametric panel data model with feedback.

In what follows we call the $T-1$ densities $g\left(\left.x_{T}\right|y^{T-1},x^{T-1},a\right)$, $g\left(\left.x_{T-1}\right|y^{T-2},x^{T-2},a\right)$, ..., $g\left(\left.x_{2}\right|y^{1},x_{1},a\right)$ the feedback process. This process describes how past values of the outcome influence current regressor values. Because it depends on $A$, the feedback process is heterogeneous across units. We do not impose any stationarity over $t=1,\ldots,T$, so $g\left(\left.x_{t}\right|y^{t-1},x^{t-1},a\right)$ and $g\left(\left.x_{s}\right|y^{s-1},x^{s-1},a\right)$ for $s \neq t$ may differ arbitrarily. When $X_t$ is a policy variable, such flexibility accommodates un-modeled regime shifts (e.g., as when a change in tax policy alters firm investment behavior). We use the notation $g$ to denote a generic element of the set of all allowable feedback processes $\mathcal{G}$. We call $\pi\left(\left.a\right|y_{0},x_{1}\right)$ the heterogeneity distribution; this density describes the distribution of unobserved heterogeneity, $A$, across units as well as any dependence of this heterogeneity on the {initial condition}, $\left(y_{0},x_{1}\right)$. Let $\pi$ denote a generic element of the set of all allowable heterogeneity distributions, $\Pi$. Finally we let $\nu\in\mathcal{N}$ denote an element of the set of all possible initial condition densities.

Although not emphasized in our exposition, it is straightforward to incorporate strictly exogenous regressors into the feedback model. Similarly, additional sources of heterogeneity, beyond $A$, can enter the feedback process. This generality, while important in some applications, involves no new issues and clutters notation.\footnote{To be specific: let $W^{T}$ contain all leads and lags of a vector of strictly exogeneous regressors and $B$ an additional source of heterogeneity. The results that follow are easily modified to accommodate the richer model:

align*[align* omitted — 395 chars of source]

where we assume that only the contemporaneous value of $W_{t}$ enters the parametric part of the model for simplicity.} We also note that, although not emphasized in most of our examples, $Y_t$ may be vector-valued in some settings.

Panel data models with feedback and heterogeneity arise frequently in economic applications. In structural dynamic choice models, agents' dynamic optimization typically leads to a likelihood function of the form ((ref)), where $Y_t$ contains choice variables of interest, as well as payoff variables, and $X_t$ contains dynamic state variables and un-modeled choice variables; see Aguirregabiria_Mira_JOE2010 for an exposition in the case of models with discrete outcomes. The moment conditions we derive are robust to any possible process for the dynamic state variables and distribution of unobserved heterogeneity.

Feedback also arises in program evaluation settings. Let $Y_t$ denote earnings and $X_t$ recent past participation in a job-training program. Ashenfelter_RESTAT1978 observed that Manpower Development and Training Act (MDTA) trainees had unusually low earnings in the year prior to undertaking training, consistent with a behavioral model where poor labor market outcomes (low values of $Y_{t-1}$) induced agents to seek out training in the next period ($X_{t}=1$). In contrast, maintaining no feedback would require that, conditional on the latent attribute, $A$, a worker's labor market history has no bearing on the decision to undertake training; a rather strong assumption Ashenfelter_Card_RESTAT1985.

remark{(Strict exogeneity)} In applications, researchers routinely maintain strict exogeneity of the regressors, $X_t$, conditional on the latent attribute, $A$. Strict exogeneity imposes independence of $Y_t$ and $\left(X_{t+1},\ldots,X_{T}\right)$ conditional on $\left(Y_{0},X_{1},\ldots,X_{t},A\right)$. Chamberlain_EM1982 shows that strict exogeneity is equivalent to the following no feedback condition: \begin{equation} g\left(\left.x_{t}\right|y^{t-1},x^{t-1},a\right)=g\left(\left.x_{t}\right|y_{0},x^{t-1},a\right),\ t=2,\ldots,T. \end{equation} While ((ref)) allows $Y_{s}$ to influence $X_{t}$ for $s<t$, (ref) rules out such feedback (we allow dependence on the initial condition throughout). Restriction (ref) is a counterpart of Granger's Granger_EM1969 definition of “$Y$ does not cause $X$" in the panel data setting with latent heterogeneity. If $X_{t}$ is a choice variable, and $Y^{t-1}$ is a component of the agent's begining-of-period $t$ information set, then (ref) typically implies strong restrictions on economic behavior Ashenfelter_Card_RESTAT1985 and/or the structure of agent's information sets Chamberlain_HBE84, chamberlain1985heterogeneity.\footnote{Under strict exogeneity, the likelihood function becomes \begin{equation}\ell^{\mathrm{SE}}\left(\left.\theta,\pi,\nu\right|y^{T},x^{T}\right)=\left\{ \int\left[\prod_{t=1}^{T}f_{\theta}\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)\right]\pi\left(\left.a\right|x_{1}\ldots,x_{T},y_{0}\right)\mathrm{d}a\right\} \nu\left(y_{0},x^{T}\right), \end{equation}where now the distribution of unobserved heterogeneity $\pi\left(\left.a\right|x_{1}\ldots,x_{T},y_{0}\right)$ is conditional on covariates in all periods.} Our approach, by accommodating unrestricted feedback and heterogeneity, allows the researcher to proceed without maintaining such assumptions.

The goal: feedback and heterogeneity robust estimation

In this paper we study estimation of the common parameter, $\theta$, in panel data models when (ref) is the only prior restriction on $F$ (except for some mild regularity conditions). We wish to construct estimators that are consistent irrespective of the precise instances of the feedback process, $g$, heterogeneity distribution, $\pi$, and initial condition, $\nu$, which describe the sampled data. In what follows we call such estimators feedback and heterogeneity robust (FHR). Because strict exogeneity obtains as a special case, any FHR estimator remains valid under strict exogeneity. FHR estimators are natural extensions of familiar “fixed effects” approaches to estimation in panel data models without feedback. We fully characterize the set of moment-based FHR estimators. We also derive semiparametric efficiency bounds for $\theta$. In particular, our results allow for a precise quantification of the information loss associated with accommodating unrestricted feedback relative to maintaining strict exogeneity.

One alternative to FHR estimation involves assuming that both the feedback process, $g$, and heterogeneity density, $\pi$, belong to parametric families indexed by some parameter vector $\eta$. With these additional maintained assumptions, the econometrician can then maximize the resulting likelihood, conditional on $Y_{0}$ and $X_{1}$, with respect to both $\theta$ and $\eta$. This transforms the problem from a semiparametric to a parametric one. Consistency of such parametric “random effects” estimators typically requires that the additionally maintained parametric restrictions on $g$ and $\pi$ hold in the sampled population. Consequently, such estimators are not generally feedback and heterogeneity robust.

We next formally define the FHR property. Let $\theta\in\Theta$, and $\omega=\left(g,\pi,\nu\right)\in\Omega$ collect all the nuisance parameters in our model. We assume that all elements $\omega\in\Omega$ have a common support, known to the econometrician (which may be unbounded and include the full real line, for example, for the support of the heterogeneity $A$). Let $\mathbb{E}_{\theta,\omega}\left[\cdot\right]$ denote an expectation taken under the DGP at $\left(\theta,\omega\right)$. Let $\phi_{\theta}\left(y_{0},y_{1},\ldots,y_{T},x_{1},\ldots,x_{T}\right)$ be a function of the observed data indexed by $\theta$. We say that $\phi_{\theta}$ is a FHR moment function if, for all $\omega$ such that $\phi_{\theta}$ is absolutely integrable under DGP $\left(\theta,\omega\right)$, we have

equation[equation omitted — 155 chars of source]

The main goal of the paper is to derive moment restrictions $\phi_{\theta}$ that have the FHR property. We will first provide a complete characterization of FHR moment restrictions for $\theta$ in Section (ref).\footnote{In fact, our characterization of FHR moment functions remains valid if $\theta$ is infinite-dimensional (for example, if $\theta$ contains nonparametric elements such as functions). However, the finite-dimensional $\theta$ case contains many important models, and all the examples we will use as illustrations feature a finite-dimensional $\theta$ vector. We also focus on this case in our analysis of efficiency.} Then, in Section (ref) we will show how essentially the same analysis can be used to provide a characterization of FHR moment functions for average effects of the form $$\mu(\theta,\omega)=\mathbb{E}_{\theta,\omega}\left[h_{\theta}\left(Y^{T},X^{T},A\right)\right],$$ where $h_{\theta}\left(Y^{T},X^{T},A\right)$ is a known function of $(Y^{T},X^{T},A)$ indexed by $\theta$. Average effects include a variety of causal or structural parameters, such as average partial effects, as in Chamberlain_HBE84, and average structural functions, as in Blundel_Powell_WC03.

Since any FHR moment function is also valid under strict exogeneity, the set of all FHR moment conditions for a given panel data model will be a subset of the corresponding set of moment conditions derived under strict exogeneity (and characterized in Bonhomme_EM12). In some cases the latter set may be non-trivial and the former empty. For example, chamberlain2010binary showed point identification of the panel logit model under strict exogeneity, while bonhomme2023identification show a failure of identification in this model under feedback (see also Section (ref) below).

FHR moment restrictions for common parameters

This section presents our first main result: a characterization of all FHR moment conditions for $\theta$. As an application, we also specialize our results to provide a constructive characterization of the set of all possible FHR moment conditions for the MPH model.

Main characterization of FHR moments for $\theta$

We begin by characterizing the set of FHR moments for $\theta$.

theorem{(FHR Moment Characterization).} Let $\theta\in\Theta$. (A) Suppose that the following $T$ restrictions hold: (i) \begin{equation}\int \phi_{\theta}(y^T,x^T){{\prod_{t=1}^T f_{\theta}(y_t\,|\, y_{t-1},x_{t},a)}}\mathrm{d}y^{1:T}=0,\end{equation}and (ii), for all $s=2,...,T$, \begin{align} &\int \phi_{\theta}(y^T,x^T)\prod_{t=s}^Tf_{\theta}(y_{t}\,|\, y_{t-1},x_{t},a)\mathrm{d}y^{s:T} does not depend on x^{s:T}. \end{align} Then, for all $\omega=(g,\pi,\nu)\in\Omega$ such that $\mathbb{E}_{\theta,\omega}\left[\abs{\phi_{\theta}(Y^T,X^T)}\right]<\infty$, we have $$\mathbb{E}_{\theta,\omega}[\phi_{\theta}(Y^T,X^T)]=0.$$ (B) Suppose that, (i) the root density $\omega\mapsto \ell^{1/2}\left(\left.\theta,\omega\right|y^{T},x^{T}\right)$ is differentiable in quadratic mean at $\omega^*$ for some $\omega^*=(g^*,\pi^*,\nu^*)\in\Omega$, and (ii) there is a neighborhood $\mathcal{N}$ of $\omega^{*}$ such that $\sup_{\omega \in \mathcal{N}}\mathbb{E}_{\theta,\omega}[\lvert{\phi_{\theta}(Y^T,X^T)}\rvert^2]<\infty$. Then, the converse is also true.

\vskip .3cm

As an implication of Part (A) of Theorem (ref), suppose the data is generated according to some $(\theta_0,\omega_0)$. Then any function $\phi_{\theta}$ that satisfies (ref) and (ref) and is absolutely integrable under the population DGP ($\theta_0,\omega_0$) has zero mean at the true $\theta_0$, that is, $$\mathbb{E}_{\theta_0,\omega_0}[\phi_{\theta_0}(Y^T,X^T)]=0.$$ To see this, simply apply Part (A) of Theorem (ref) with $\theta=\theta_0$ and $\omega=\omega_0$.

The first condition for the FHR property, restriction (ref), ensures robustness of $\phi_\theta$ to the presence of heterogeneity of an unknown form. This condition is highlighted in Bonhomme_EM12, in a setting with strictly exogenous covariates, as the key condition ensuring valid moment functions for $\theta$. Under strict exogeneity, (ref) ensures that $\phi_{\theta}(Y^T,X^T)$ is conditionally mean zero given $X^T$ and $A$. By the law of iterated expectations, this suffices to ensure that it is unconditionally mean zero as well. The functional differencing approach then provides a general recipe for constructing functions $\phi_{\theta}$ satisfying (ref).

To illustrate how the presence of feedback modifies the interpretation of condition (ref), consider the $T=2$ setting. In this case, the condition is $$\int\left[\int \phi_{\theta}\left(y_0,y_1,y_2,x_1,x_2\right)f_{\theta}(y_2\,|\, y_1,x_2,a)dy_2\right]f_{\theta}(y_1\,|\, y_0,x_1,a)dy_1=0,$$ which coincides with

equation[equation omitted — 234 chars of source]

Here the inner expectation corresponds to an average over $y_2$ with respect to its model density, $f_{\theta}(y_{2}\,|\, y_{1},x_{2},a)$. This inner expectation conditions on $X_2=x_2$, while the outer expectation, which averages over $y_1$ with respect to its model density, $f_{\theta}(y_{1}\,|\, y_{0},x_{1},a)$, does not condition on $x_2$. This is a key difference between the heterogeneous feedback case, considered here, and the setting with strict exogeneity studied by Bonhomme_EM12. Under strict exogeneity, $Y_1$ is independent of $X_2$ conditional on $(Y_0,X_1,A )$, so ((ref)) coincides with $$\mathbb{E}\left[\mathbb{E}\left[\left.\phi_{\theta}\left(Y_0,Y_1,Y_2,X_1,X_2\right)\right|Y_0,Y_1,X_1,X_2,A\right]\,|\,Y_0,X_1,X_2,A \right]=0,$$ with both the inner and outer expectations conditioning on $X_2$. By iterated expectations, this conditional expectation implies a zero mean condition given $A$, initial conditions, and the entire sequence of covariates, $$\mathbb{E}\left[\left.\phi_{\theta}\left(Y_0,Y_1,Y_2,X_1,X_2\right)\right|Y_0,X_1,X_2,A\right]=0.$$ However, under feedback it is not generally the case that ((ref)) can be written as a zero mean condition given the covariates sequence $(X_1,X_2)$.

In the presence of feedback, Theorem (ref) additionally requires the $T-1$ conditions (ref), to ensure that $\phi_\theta$ is a valid moment function. To explain this additional requirement, consider again the $T=2$ setting, in which case (ref) reads $$\int \phi_{\theta}\left(y_0,y_1,y_2,x_1,x_2\right)f_{\theta}(y_2\,|\, y_1,x_2,a)dy_2\quad \text{does not depend on }x_2,$$ which corresponds to the single mean independence restriction

equation[equation omitted — 261 chars of source]

Equation (ref) implies that, at the the population parameter, $X_2$ does not predict $\phi_{\theta}\left(Y_0,Y_1,Y_2,X_1,X_2\right)$ conditional on $Y_0,Y_1,X_1,A$. Under this condition, which we interpret as “feedback robustness” of $\phi_{\theta}$, the conditioning on $X_2=x_2$ disappears in ((ref)), which implies

equation*[equation* omitted — 170 chars of source]

ensuring, by iterated expectations, that $\phi_{\theta}$ is a valid moment function. Importantly, $\phi_\theta$ is then a valid moment under both unrestricted heterogeneity and unrestricted feedback.

Lastly, Theorem (ref) also establishes that (ref) and (ref) are, under suitable conditions, necessary for the FHR property. To show the necessity property, we assume quadratic mean differentiability at $(\theta,\omega^*)$ as stated in Part (B) Condition (i). This is a standard regularity condition in semiparametric estimation, see for example van2000asymptotic. Part (B) Condition (ii), which imposes square-integrability of the moment function in a neighborhood of $\omega^*$, is similarly standard; e.g., see the assumptions for the generalized information equality in Lemma 5.4 of newey1994large.\footnote{One can show that, if $\phi_{\theta}$ is bounded, then the necessity of (ref) and (ref) obtains directly without reference to $\omega^*$ or differentiability in quadratic mean.} We will rely on differentiability in quadratic mean in our analysis of efficiency in Section (ref).

Alternative characterization of FHR moments for $\theta$

The following corollary to Theorem (ref) facilitates the construction of FHR moment functions in practice.

corollary{(Alternative FHR Moment Characterization)} Let $\omega^*\in\Omega$. An absolutely-integrable function $\phi_{\theta}$ under $(\theta,\omega^*)$ satisfies (ref) and (ref) if and only if there exist absolutely-integrable functions $\psi_{\theta,t}$ under $(\theta,\omega^*)$, for $t=1,\ldots,T-1$, such that: \begin{align}\mathbb{E}\left[\left.\phi_{\theta}(Y^T,X^T)\right|Y^{T-1},X^T,A\right]=\sum_{t=1}^{T-1}\psi_{\theta,t}(Y^t,X^t,A), \end{align} with $\mathbb{E}\left[\left.\psi_{\theta,t}\left(Y^{t},X^{t},A\right)\right|Y^{t-1},X^{t},A\right]=0$ for $t=1,\ldots,T-1$.

\vskip .3cm

To understand Corollary (ref), it is helpful to return to the $T=2$ setting. In this case, letting $$\psi_{\theta,1}(Y_0,Y_1,X_1,A)=\mathbb{E}[\phi_{\theta}(Y_0,Y_1,Y_2,X_1,X_2,A)\,|\, Y_0,Y_1,X_1,A],$$ it follows from ((ref)) that $$\psi_{\theta,1}(Y_0,Y_1,X_1,A)=\mathbb{E}[\phi_{\theta}(Y_0,Y_1,Y_2,X_1,X_2,A)\,|\, Y_0,Y_1,X_1,X_2,A],$$ which implies ((ref)), whereas it follows from ((ref)) that $$\mathbb{E}[\psi_{\theta,1}(Y_0,Y_1,X_1,A)\,|\, Y_0,X_1,A]=0$$ (a requirement for $\psi_{\theta,1}$ given in the corollary).

The representation provided by Corollary (ref) suggests a systematic recipe for constructing FHR moment functions. The first step involves finding functions $\psi_{\theta,t}(Y^{t},X^t,A)$ that are mean zero conditional on $Y^{t-1},X^t,A$ for $t=1,\ldots,T-1$. This is straightforward since any function of $(Y^{t},X^t,A)$, suitably de-meaned, automatically fulfills this requirement (see Section (ref) for several examples). Then, given a collection of $\psi_{\theta,t}(Y^{t},X^t,A)$ functions, the second step requires solving a linear integral equation to recover a valid moment function $\phi_{\theta}$. Indeed, ((ref)) can equivalently be written as

align[align omitted — 170 chars of source]

which is known as an inhomogeneous Fredholm equation of the first kind. Note that the integral operator on the left-hand side of (ref) is known given the parameter $\theta$. Solution methods for this type of linear integral equation are the subject of a large literature engl1996regularization,carrasco2007linear. In Section (ref) we will show how to construct FHR moment functions in some specific models using Corollary (ref), and discuss when such functions exist.

A special case of Corollary (ref) obtains when one is able to find functions $\eta_{\theta,t}$ such that

equation[equation omitted — 155 chars of source]

for some function $b$ of the heterogeneity and initial conditions. Then, the instrumented first difference

equation*[equation* omitted — 149 chars of source]

satisfies ((ref)), for $\psi_{\theta,T-1}(y^{T-1},x^{T-1},a)=[b\left(y_0,x_1,a\right)-\eta_{\theta,T-1}\left(y^{T-1},x^{T-1}\right)]\cdot m(y^{T-2},x^{T-1})$ and $\psi_{\theta,t}=0$ for all $t<T-1$ (for an arbitrary function $m$). This provides a FHR moment function on $\theta$. More generally, one can check that

equation[equation omitted — 189 chars of source]

all satisfy ((ref)), thus providing additional moments.

We now illustrate this particular recipe with the MPH model.

example[continues=ex: mph_intro] (Simple Moments for the MPH) It is convenient to first develop some additional implications of the MPH model under feedback. Recall the notations $z_t=(y_t,y_{t-1},x_t')'$ and $\rho_{\theta}\left(z_{t}\right)=\Lambda_{\alpha}\left(y_{t}\right)\exp\left(\gamma y_{t-1}+x_{t}'\beta\right)$. Adapting arguments appearing in Lancaster_EATD1990, hahn1994efficiency and Ridder_Woutersen_EM2003, it is straightforward to show that \begin{equation} \left.\rho_{\theta}\left(Z_{t}\right)\right|Y^{t-1},X^{t},A\sim\mathrm{Exponential}\left(e^{A}\right),\thinspace t=1,\ldots,T, \end{equation} (see Lemma (ref) in Supplemental Appendix (ref)). From this observation we have \begin{equation} \mathbb{E}\left[\left.\rho_{\theta}\left(Z_{t}\right)\right|Y^{t-1},X^{t},A\right]=e^{-A},\thinspace t=1,\ldots,T, \end{equation} which coincides with (ref) after setting $\eta_{\theta,t}(Y^t,X^t)=\rho_{\theta}(Z_t)$ and $b(Y_0,X_1,A)=e^{-A}$. This yields the FHR moment functions \begin{equation} \phi_{\theta}\left(Y^{t},X^{t}\right)=\left[\rho_{\theta}\left(Z_{t}\right)-\rho_{\theta}\left(Z_{t-1}\right)\right]\cdot m\left(Y^{t-2},X^{t-1}\right),\quad t=2,...,T, \end{equation} where $m$ is an arbitrary function. In unpublished dissertation research, woutersen2000essays presented moments similar to (ref).

Moments of the form ((ref)) were considered by arellano1991some in the linear model context, by Chamberlain_JBES92,Chamberlain_JOE2022 and wooldridge1997multiplicative for Poisson models, and by Al-Sadoon_et_al_ER2017 for certain binary choice models. However, this family of estimating equations does not exhaust all available FHR moments, and may lead to estimators with low levels of asymptotic precision in practice. By comparison, Theorem (ref) and Corollary (ref) characterize all available FHR moments, as we will now illustrate in the case of the MPH model. Moreover, as we will establish in later sections, our characterization can be used to derive efficient estimators (i.e., we extend to nonlinear models efficiency arguments which appear in the linear panel data literature such as in arellano1995another,arellano2001panel, Arellano_RIE2016).

FHR moments for $\theta$ in the MPH model

For specific models it is sometimes possible to use Theorem (ref) to provide a direct characterization of all FHR moment functions. We show how this can be done in the MPH model with feedback in Lemma (ref) below. Our result provides insight into the MPH model and simplifies the derivation of new FHR moments.

Our characterization makes use of several special features of the MPH model, which we introduce first (details can be found in Supplemental Appendix (ref)). Let $P_{\theta,t}=\rho_{\theta}\left(Z_t\right)$. As indicated in (ref), conditional on $Y_{0},X_{1},A$, the $P_{\theta,t}$ for $t=1,\ldots,T$ are independent exponential random variables; each with a common rate parameter of $e^{A}$.

Next, we define a one-to-one transformation of the vector $P_{\theta}^T$ into a “forward orthogonal deviations” part and a “between” part, as follows:

equation[equation omitted — 281 chars of source]

Observe that $\widetilde{P}_{\theta,t}$ involves the ratio of $\rho_{\theta}\left(Z_{t}\right)$ to the sum of itself and the future values of $\rho_{\theta}\left(Z_{s}\right)$ for $s=t+1,\ldots,T$. Lemma (ref) in Supplemental Appendix (ref) establishes that, conditionally on $Y_{0},X_{1},A$:

equation[equation omitted — 224 chars of source]

where, throughout, $\theta$ is the parameter indexing the DGP. Lemma (ref) additionally establishes that the elements of $\left(\widetilde{P}_{\theta,1},\ldots,\widetilde{P}_{\theta,T-1},\overline{P}_{\theta}\right)$ are mutually independent of one another.

Transformation (ref) can be thought of as a MPH-specific analog of the forward orthogonal deviations transformation used by arellano1995another in the context of linear panel data models with predetermined regressors. It has the property that the random variables $\widetilde{P}_{\theta,t}$ are independent of both contemporaneous and lagged predetermined covariates as well as lagged values of the spell outcomes $\{(X_{is},Y_{is-1})\}_{s=1}^t$ (see part (i) of Lemma (ref)). Appealing to the same analogy, we can think of $\overline{P}_{\theta}$ as containing the “between” variation in $P_{\theta}^T$.

With these preliminaries in place, we can state the following FHR moment characterization for the MPH model with feedback.

lemma(Characterization of FHR moment restrictions for the MPH model) Consider the MPH model with $T\geq2$. Let $\omega^*\in\Omega$. Then an absolutely-integrable $\phi_{\theta}$ under $(\theta,\omega^*)$ satisfies (ref) and (ref) if and only if there exist absolutely-integrable $\psi_{\theta,t}$ under $(\theta,\omega^*)$, for $t=1,...,T$, such that: \begin{align*} \phi_{\theta}\left(Y^{T},X^{T}\right)=&\sum_{t=1}^{T-1}\psi_{\theta,t}(Y_{0},\widetilde{P}_{\theta}^{t},\overline{P}_{\theta},X^{t}), \end{align*} where, for $t=1,\ldots,T-1$, \[ \mathbb{E}\left[\left.\psi_{\theta,t}(Y_{0},\widetilde{P}_{\theta}^{t},\overline{P}_{\theta},X^{t})\right|Y_{0},\widetilde{P}_{\theta}^{t-1},\overline{P}_{\theta},X^{t}\right]=0. \]

\vskip .3cm

The proof of Lemma (ref) is available in Supplemental Appendix (ref). It represents a useful simplification, induced by the special structure of the MPH model, relative to the general result of Theorem (ref). Note in particular that the latent heterogeneity $A$ does not appear in the characterization of Lemma (ref). To illustrate, consider the $T=2$ setting. In this case the lemma implies that all FHR moment functions are of the form $$\phi_{\theta}\left(Y_{0},Y_1,Y_2,X_1,X_2\right)=\psi_{\theta,1}(Y_{0},\widetilde{P}_{\theta,1},\overline{P}_{\theta},X_1),$$ where $\mathbb{E}\left[\left.\psi_{\theta,1}(Y_{0},\widetilde{P}_{\theta,1},\overline{P}_{\theta},X_1)\right|Y_{0},\overline{P}_{\theta},X_1\right]=0$. We seek functions $\psi_{\theta,1}$ of the forward orthogonal deviations $\widetilde{P}_{\theta,1}$, that are mean zero conditional on the between variation $\overline{P}_{\theta}$, as well as the initial condition $\left(Y_{0},X_{1}\right)$. It is straightforward to construct such functions by de-meaning, and we will exploit this property in our analysis of efficiency.

For the general case of an arbitrary number $T$ of periods, $\phi_{\theta}$ is the sum of functions $\psi_{\theta,t}\left(Y_{0}, \widetilde{P}_{\theta}^{t},\overline{P}_{\theta},X^{t}\right)$, for $t=1,\ldots,T-1$, that are mean independent of the between variation, $\overline{P}_{\theta}$, current and past values of the predetermined regressors, $X^{t}$, and past values of the spell outcomes, $Y^{t-1}$. This structure is reminiscent of how moment conditions are typically constructed for linear panel data models with predetermined regressors (see, especially, arellano1995another).

FHR moment restrictions for average effects

In this section we characterize all FHR moment conditions for average effects, $$\mu(\theta,\omega)=\mathbb{E}_{\theta,\omega}\left[h_{\theta}\left(Y^{T},X^{T},A\right)\right],$$ where $h_{\theta}\left(Y^{T},X^{T},A\right)$ is a known function of $Y^{T},X^{T},A$ indexed by $\theta$.

example[continues=ex: mph_intro] (Average Effects in the MPH Model) Consider the average structural hazard (ASH) function \begin{align} \overline{\lambda}(y_{t}|x_t,y_{t-1})&=\mathbb{E}_{\theta,\omega}\left[\lambda_{\alpha}(y_{t})e^{x_t'\beta+\gamma y_{t-1}+A}\right]. \end{align} The ASH appears to be a new estimand, albeit a natural one given concerns about spurious duration dependence Lancaster_EATD1990, Heckman_AER1991. In the context of the MPH model, the ASH corresponds to the mean survival time for a unit exogenously assigned to policy $X_t=x_t$ and history $Y_{t-1}=y_{t-1}.$ Like other quantities relevant to causal analysis, the ASH depends on the (marginal) distribution of unobserved heterogeneity $A$.\footnote{A related average effect of interest is the average structural function (ASF) (blundell2004endogeneity): \begin{align} \mu(x_t,y_{t-1})&=\mathbb{E}_{\theta,\omega}\left[m(x_t,y_{t-1},A;\theta)\right], \end{align} where $m(x_t,y_{t-1},a;\theta)=\mathbb{E}\left[Y_t|X_t=x_t,Y_{t-1}=y_{t-1},A=a\right]$.}

In nonlinear panel data models, knowledge of $\theta$ does not suffice to identify the effect of an external manipulation of a regressor's value on the probability distribution of the outcome. This follows from the nonseparable way in which the unobserved heterogeneity enters such models. In contrast, estimands which average over the marginal distribution of $A$, such as (ref), do provide easy-to-interpret summaries of such effects.

Until recently the identifiability of such averages was not well understood. Recent work by honore2006bounds, chernozhukov2013average, aguirregabiria2021identification, davezies2021identification, dobronyi2021identification, pakel2023bounds and others, however, has shown that average effects are (partially) identified in several specific settings of interest. The study of average effects in more general settings, such as those with feedback, as we consider here, remains underdeveloped bonhomme2023identification.

Characterization of FHR moments for $\mu(\theta,\omega)$

Our first result is an analog of Theorem (ref) for $\mu(\theta,\omega)$.

theorem{(Characterization of moment restrictions for average effects)} Let $\theta\in\Theta$. (A) Suppose that the following $T$ restrictions hold: (i) \begin{equation}\int \left(\varphi_{\theta}(y^T,x^T)-h_{\theta}(y^T,x^T,a)\right){{\prod_{t=1}^T f_{\theta}(y_t\,|\, y_{t-1},x_{t},a)}}\mathrm{d}y^{1:T}=0,\end{equation}and (ii), for all $s=2,...,T$, \begin{align} &\int \left(\varphi_{\theta}(y^T,x^T)-h_{\theta}(y^T,x^T,a)\right)\prod_{t=s}^Tf_{\theta}(y_{t}\,|\, y_{t-1},x_{t},a)\mathrm{d}y^{s:T} does not depend on x^{s:T}. \end{align} Then, for all $\omega\in\Omega$ such that $\mathbb{E}_{\theta,\omega}\left[\abs{\varphi_{\theta}(Y^T,X^T)}\right]<\infty$ and $\mathbb{E}_{\theta,\omega}\left[\abs{h_{\theta}(Y^T,X^T,A)}\right]<\infty$ we have $$\mathbb{E}_{\theta,\omega}[\varphi_{\theta}(Y^T,X^T)]=\mu(\theta,\omega).$$ (B) Suppose (i) the root density $\omega\mapsto \ell^{1/2}\left(\left.\theta,\omega\right|y^{T},x^{T}\right)$ is differentiable in quadratic mean at $\omega^*$ for some $\omega^*=(g^*,\pi^*,\nu^*)\in\Omega$, (ii) there is a neighborhood $\mathcal{N}$ of $\omega^{*}$ such that $\sup_{\omega \in \mathcal{N}}\mathbb{E}_{\theta,\omega}[\norm{\phi_{\theta}(y^T,x^T)}^2]<\infty$ and $\sup_{\omega \in \mathcal{N}}\mathbb{E}_{\theta,\omega}[\|h_{\theta}(Y^T,X^T,A)\|^2]<\infty$. Then, the converse is also true.

\vskip .3cm

Note the strong parallel between this theorem and Theorem (ref). As in the case of $\theta$, (ref) and (ref) imply that, under absolute integrability, the true value $\mu_0=\mu(\theta_0,\omega_0)$ satisfies a moment condition. Indeed, if $\mathbb{E}_{\theta_0,\omega_0}\left[\abs{\varphi_{\theta_0}(Y^T,X^T)}\right]<\infty$ and $\mathbb{E}_{\theta_0,\omega_0}\left[\abs{h_{\theta_0}(Y^T,X^T,A)}\right]<\infty$, then Part (A) implies $$\mathbb{E}_{\theta_0,\omega_0}[\varphi_{\theta_0}(Y^T,X^T)]=\mu_0.$$ This moment condition has the FHR property: (ref) ensures that $\varphi_{\theta}$ is robust to an unknown distribution of unobserved heterogeneity, whereas the $T-1$ conditions (ref) endow $\varphi_{\theta}$ with robustness to heterogeneous feedback of an unknown form.

We additionally state the following corollary, which mimics Corollary (ref) and suggests a systematic recipe for constructing functions $\varphi_{\theta}$.

corollary{(Alternative characterization of moment restrictions for average effects)} Let $\omega^*\in\Omega$. Suppose that $h_{\theta}$ is absolutely integrable under $(\theta,\omega^*)$. An absolutely-integrable $\varphi_{\theta}$ satisfies ((ref))-((ref)) if and only if there exist absolutely-integrable $\zeta_{\theta,t}$, for $t=1,\ldots,T-1$, such that: \begin{align} \mathbb{E}\left[\varphi_{\theta}(Y^T,X^T)-h_{\theta}(Y^T,X^T,A)\,|\, Y^{T-1},X^T,A\right]=\sum_{t=1}^{T-1}\zeta_{\theta,t}(Y^{t},X^{t},A), \end{align} with $\mathbb{E}\left[\zeta_{\theta,t}(Y^{t},X^{t},A)\,|\,Y^{t-1},X^{t},A\right]=0$ for $t=1,\ldots,T-1$.

\vskip .3cm

FHR moments for the average structural hazard

The conditional hazard function \[ \lambda\left(\left.y_{t}\right|y_{t-1},x_{t},a\right)=\lambda_{\alpha}(y_{t})e^{x_{t}'\beta+\gamma y_{t-1}+a} \] gives the instantaneous exit rate of a unit, at duration $y_{t}$, with lagged duration $y_{t-1}$, beginning-of-spell covariate $x_{t}$, and latent attribute $a$. Unfortunately, although easily identified, the observed hazard function \[ \lambda\left(\left.y_{t}\right|y_{t-1},x_{t}\right)=\lambda_{\alpha}(y_{t})e^{x_{t}'\beta+\gamma y_{t-1}}\mathbb{E}\left[\left.e^{A}\right|Y_{t}>y_{t},Y_{t-1}=y_{t-1},X_{t}=x_{t}\right] \] suffers from spurious duration dependence (see Lancaster_EATD1990,Heckman_AER1991): units with higher values of $A$ will exit earlier, implying that the mean $\mathbb{E}\left[\left.e^{A}\right|Y_{t}>y_{t},Y_{t-1}=y_{t-1},X_{t}=x_{t}\right]$ declines with $y_{t}$. In contrast, the average structural hazard (ASH),

equation[equation omitted — 183 chars of source]

which equals the (expected) hazard function for a randomly sampled unit when externally assigned lagged duration $y_{t-1}$ and covariate $x_{t}$, does not suffer from heterogeneity bias.

Since $\left.\overline{P}_{\theta}\right|Y_{0},X_{1},A\sim\mathrm{Gamma}\left(T,e^{A}\right)$ we have

align[align omitted — 190 chars of source]

for $\delta>-T$. Using (ref) with $\delta=-1$ shows that $\mathbb{E}\left[e^{A}\right]=\mathbb{E}\left[\frac{T-1}{\overline{P}_{\theta}}\right]$, so

equation[equation omitted — 216 chars of source]

We have thus found one FHR moment function for the ASH. Now, for any other FHR moment function for the ASH, $\varphi_{\theta}$, we have that $$\phi_{\theta}(Y^T,X^T)=\varphi_{\theta}(Y^T,X^T)-\lambda_{\alpha}(y_{t})e^{x_{t}'\beta+\gamma y_{t-1}}\frac{T-1}{\overline{P}_{\theta}}$$ is a FHR moment function for $\theta$ (where here $y_t,y_{t-1},x_t$ are fixed values at which the ASH is evaluated). Since all such moment functions $\phi_{\theta}$ are fully characterized in the MPH case by Lemma (ref), this characterizes all FHR moment functions for the ASH. It turns out, as we will shortly see, that estimation based upon the sample analog of (ref) is efficient when $\theta$ is replaced by an efficient estimate.

Efficient moment restrictions

Theorems (ref) and (ref) characterize the complete set of moment conditions available for estimating $\theta$ and $\mu(\theta,\omega)$. When this set is nonempty, estimation at parametric rates may be (and often is) feasible. Moreover, many valid moment restrictions may be available (see Lemma (ref) for the case of the MPH model). In such settings semiparametric efficiency bound theory provides a useful tool for selecting specific moments for estimation purposes.

In this section we derive the form of the efficient scores for $\theta$ and $\mu(\theta,\omega)$ and, consequently, their corresponding semiparametric efficiency bounds. As we shall see, the characterizations presented earlier facilitate the derivation of these bounds. We illustrate the application of our results via a detailed analysis of efficiency bounds for the MPH model.\footnote{hahn1994efficiency derived the information bound for $\theta$ in the single-spell case studied by elbers1982true and Heckman_Singer_ReStud1984. He also derived the bound for the multi-spell case under strict exogeneity (see also chamberlain1985heterogeneity). Relative to prior work, our analysis covers the case with feedback and, additionally, considers average effects.}

Efficiency for common parameters and average effects

We begin with a brief review of basic concepts in semiparametric efficiency theory; see, for example, van2000asymptotic and newey1990semiparametric. Let $\theta_0$, $g_0$, $\pi_0$, and $\nu_0$ denote, respectively, the common parameter, feedback process, heterogeneity distribution, and initial condition prevailing in the sampled population. Let $\omega=(g,\pi,\nu)\in\Omega$, with true value $\omega_0=(g_0,\pi_0,\nu_0)$, denote the nonparametric components of the model.\footnote{To characterize the efficiency bound for $\theta$, by ancillarity it is sufficient to consider the conditional density given $(y_{0},x_{1})$ (see, e.g., hahn1994efficiency).} A regular parametric submodel is defined by a likelihood function for a single random draw, $\ell(\theta,\omega_{\eta}\,|\,y^T,x^{T})$, where $\omega_{\eta_0}=\omega_0$ for some scalar $\eta_0$. The likelihood satisfies mean-square differentiability of its square root with respect to $(\theta,\eta)$, with its information matrix nonsingular. The semiparametric variance bound is the supremum of the Cramer Rao bounds for $\theta$ over all such regular parametric submodels.

Let $S^{\theta}$ denote the score for $\theta$, for a submodel evaluated at $\theta=\theta_0$ and $\eta=\eta_0$: $$S^{\theta}(Y^T,X^T)=\frac{\partial \ln \ell(\theta_0,{\omega}_0\,|\, Y^T,X^T)}{\partial \theta},$$ where we leave the dependence of $S^{\theta}$ on $(\theta_0,\omega_0)$ implicit. Likewise, let $S^{\eta}$ denote the score for $\eta$. The nonparametric tangent set $\mathcal{T}_{\theta_0,\omega_0,K}$ is the mean-square closure of the $K\times 1$ linear combinations of scores $S^{\eta}$ across all regular parametric submodels (where $K$ is the dimension of $\theta$). Let $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$ denote the orthocomplement of $\mathcal{T}_{\theta_0,\omega_0,K}$ in the Hilbert space of square-integrable mean-zero functions with inner product $\left<s_1,s_2\right>=\mathbb{E}_{\theta_0,\omega_0}\left[s_1(Y^T,X^T)'s_2(Y^T,X^T)\right]$. By definition this set consists of all $K\times 1$ elements $\phi_{\theta}$ such that $\left<\phi,s\right>=0$ for all $s\in \mathcal{T}_{\theta_0,\omega_0,K}$.

The next theorem provides a characterization of the orthocomplement of the tangent set, $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$, which is key to the analysis of efficiency in our context.

theorem{(Orthocomplement of tangent set)} $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ is the linear span of square-integrable, $K$-dimensional moment functions $\phi_{\theta_0}$ that satisfy (ref) and (ref).

\vskip .3cm

An implication of Theorem (ref) is that, if $\phi_{\theta_0} \in \mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$, then it is also an element of the orthogonal complement of the nuisance tangent set associated with any other data generating process $(\theta_0, \omega_*)$ with $\omega_* \neq \omega_0$ (i.e., we also have $\phi_{\theta_0} \in \mathcal{T}_{\theta_0,\omega_*,K}^{\perp}$). This follows from the fact that $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ consists of the set of FHR moments characterized in Theorem (ref) earlier; a set which does not vary with $\omega_0$. Indeed, it is precisely this feature of the model which makes feedback and heterogeneity robust estimation possible. Knowledge, whether a priori or up to sampling error, of the form of the feedback process and/or heterogeneity distribution is not required for consistent estimation. This is a crucial feature of the class of models we study in this paper.

One subtlety is that, although the elements of $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ do not depend on the precise instance of $\omega_{0}$, the definition of the reference Hilbert space does depend on it. Let $\phi_{\theta_{0}}$ be a valid FHR moment that is absolutely integrable under $(\theta_0,\omega^*)$. Then, while $\phi_{\theta_{0}}(Y^{T},X^{T})$ has zero mean under both $(\theta_0,\omega_0)$ and $(\theta_0,\omega^*)$, in general its variance differs under $\omega_0$ and $\omega^*$. Consequently, the ability to precisely estimate $\theta_{0}$ using a particular $\phi_{\theta}$ generally varies with the population feedback process and heterogeneity distribution, although the validity of $\phi_{\theta}$ as a moment function does not. This connects to the discussion of locally efficient estimation below.

To understand Theorem (ref) it is helpful to consider the implications of restricting $\omega_0$ such that it belongs to a parametric family (indexed by, say, $\eta$). This is the approach taken in, for example, parametric random-effects analysis Chamberlain_ReStud80, chamberlain1985heterogeneity. In that setting the residualized score, $\widetilde{S}^{\theta}=S^{\theta}-\mathbb{E}\left[S^{\theta}S^{\eta\prime}\right]\times \mathbb{E}\left[S^{\eta}S^{\eta\prime}\right]^{-1}S^{\eta}$, will generally vary with $\eta$: consistent estimation of $\theta$ requires knowledge of $\eta$ (up to sampling error), and it requires correct specification of the parametric models of feedback and heterogeneity. This is not required in our approach; indeed (elements of) $\omega$ may even be unidentified, while $\theta$ remains $\sqrt{N}$-estimable.

Theorem (ref) is reminiscent of the situation which arises in average treatment effect (ATE) estimation under unconfoundedness with a known propensity score robins1994estimation,hahn1994efficiency. In that setting the set of consistent estimating equations for the ATE does not depend on the form of the conditional distribution of the potential outcomes given covariates. Theorem (ref) extends prior work for the case of strictly exogenous regressors showing that $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$ is characterized by moment functions that have zero means conditional on all covariates, initial conditions, and heterogeneity (see, e.g., hahn1994efficiency for the MPH model, and dano2023transition for dynamic logit models).

Of course, in many semiparametric estimation problems, consistent estimation of (features of) the nonparametric model component is a requirement for consistent estimation of $\theta$. Examples include the binary choice model with random utility components drawn from an unknown distribution independent of the regressors newey1990semiparametric and ATE estimation under unconfoundedness with an unknown propensity score.

\paragraph{Efficient score for $\theta$.} Under regularity conditions,\footnote{Namely that $\ell(\theta,\omega|y^T,x^{T})$ is smooth in a neighborhood of $(\theta_0,\omega_0)$, and that the information matrix is nonsingular (see for example Theorem 3.2 in newey1990semiparametric).} the semiparametric variance bound for $\theta_0$ is equal to the inverse of the variance of the efficient score

equation[equation omitted — 161 chars of source]

where $\Pi\left(.\,|\,\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}\right)$ denotes the orthogonal projection onto $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$. Note that the projection is well defined since $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$ is closed and linear.\footnote{An equivalent and more frequent formulation of the efficient score is $\phi_{\theta_0}^{\rm eff}=S^{\theta}-\Pi\left(S^{\theta}\,|\,\mathcal{T}_{\theta_0,\omega_0,K}\right)$ (e.g., hahn1994efficiency), which is interpreted as the residual from the population regression of $S^{\theta}$ on the nuisance tangent set. The equivalent formulation based on $\mathcal{T}^{\perp}_{\theta_0,\omega_0,K}$ is convenient to work with in our setting; see VanDerLaan_Robin_Book2003.} As a result, the efficient moment restriction for $\theta_0$ is

equation[equation omitted — 123 chars of source]

The efficient score $\phi_{\theta_0,\omega_0}^{\rm eff}$ in ((ref)) is defined through a projection onto the orthocomplement $\mathcal{T}_{\theta_0,\omega_0,K}^{\perp}$, which we have fully characterized in Theorem (ref). Below we show, via examples, how to use this observation to calculate -- whether analytically or numerically -- efficient scores in models with unknown heterogeneity and feedback.

\paragraph{Efficient score for $\mu$.} We now turn to an analysis of efficiency for average effects. Suppose there exists a moment function $\varphi_{\theta_0}$ that identifies an $L\times 1$ average effect of interest $\mu(\theta_0,\omega_0)$ given by\footnote{Standard implicit smoothness conditions are required, namely that, for a regular parametric submodel, $\sup_{(\theta,\eta) \in \mathcal{N}} \mathbb{E}_{\theta,\omega_{\eta}}\left[\norm{\varphi_{\theta}(Y^T,X^T)}^2\right]<\infty$ in a neighborhood $\mathcal{N}$ of $(\theta_0,\eta_0)$, such that a generalized information equality holds (see brown1998efficient).}

equation[equation omitted — 193 chars of source]

and suppose that $\theta_0$ is identified from (ref). Let $${\varphi}^{\rm eff}_{\theta_0,\omega_0}(Y^T,X^T)=\varphi_{\theta_0}(Y^T,X^T)-\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L}^{\perp}),$$ where the orthocomplement of the tangent set, $\mathcal{T}_{\theta_0,\omega_0,L}^{\perp}$, is given by Theorem (ref), except for the fact that the relevant dimension is $L$ instead of $K$.

Theorem 1 in brown1998efficient shows that the joint efficient moment restrictions for $\theta_0$ and $\mu_0=\mu(\theta_0,\omega_0)$ are given by (ref) and

equation[equation omitted — 133 chars of source]

The semiparametric variance bound for $\mu_0$ equals the lower $L\times L$ block of the asymptotic variance of the joint GMM estimator $(\widehat{\theta},\widehat{\mu})$ based on (ref) and (ref).

The construction of $\varphi^{\rm eff}_{\theta_0,\omega_0}$ relies on a function $\varphi_{\theta_0}$ that satisfies (ref). In some models one can find such a function (a Riesz representer), which does not depend on the nonparametric component $\omega_0$. This can be done by exploiting Theorem (ref). See, for example, the discussion of the average structural hazard function in the MPH earlier. Moreover, $\varphi^{\rm eff}_{\theta_0,\omega_0}$ is not affected by the particular choice of $\varphi_{\theta_0}$.\footnote{This follows from noting that $\varphi_{\theta_0}(Y^T,X^T)-\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L}^{\perp})=\mu_{0}+\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L})$, and that, if $\varphi_{\theta_0}^{1}$ and $\varphi_{\theta_0}^{2}$ both satisfy (ref), then $\Pi(\varphi_{\theta_0}^{1}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L})=\Pi(\varphi_{\theta_0}^{2}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L})$ as we show in Supplemental Appendix Lemma (ref).} Of course not all average effects will have non-zero efficiency bounds, but when a Riesz representer for an effect of interest is available, the bound can be calculated using extant results about expectations brown1998efficient.

Moment restrictions based on working models

Although the characterization of feasible moment conditions provided by Theorem (ref) is invariant to the specific instance of $\omega_0$ indexing the sampled population, the projection (ref) generally does vary with $\omega_0$. Consequently, although knowledge of $\omega_0$ is not required for consistent estimation, it is generally valuable for improving asymptotic precision. Moreover, constructing an estimator which attains the semiparametric efficiency bound for all possible feedback processes, $g_0$, heterogeneity distributions, $\pi_0$, and initial conditions, $\nu_0$ (i.e., for all $\omega_0 \in \Omega$) generally requires nonparametrically estimating features of these model components. This may be practically difficult, or even impossible, in short panels as considered here.

An alternative approach involves constructing locally efficient estimators newey1990semiparametric, Graham_Pinto_Egel_ReStud12. Let $\widetilde{\omega}=(\widetilde{g},\widetilde{\pi},\widetilde{\nu})\in\Omega$ denote candidate “working models" for the feedback process, heterogeneity distribution, and initial condition. We show how to construct method-of-moments estimators that (i) attain the bound for $\theta_0$ (or $\mu_0$) when these working models “happen to characterize the sampled population" (i.e., $\widetilde{\omega}=\omega_0$, but this is not part of the prior restriction) and (ii) remain $\sqrt{N}$-consistent irrespective of the true $\omega_0$ characterizing the sampled population (i.e., when $\widetilde{\omega}\neq\omega_0$). A key property in our setting is that consistency holds for arbitrary $\widetilde{\omega}$ (i.e., our working models may be misspecified), only subject to regularity conditions.

Given working models $\widetilde{\omega}$, let $$\widetilde{S}^{\theta}(Y^T,X^T)=\frac{\partial \ln \ell(\theta_0,\widetilde{\omega}\,|\, Y^T,X^T)}{\partial \theta}$$ denote the score for $\theta_0$. Next define the counterpart, under the working models, to the efficient score $\phi_{\theta_0}^{\rm eff}$ for $\theta_0$ as $$\widetilde{\phi}_{\theta_0,\widetilde{\omega}}^{\rm eff}(Y^T,X^T)=\widetilde{\Pi}\left(\widetilde{S}^{\theta}(Y^T,X^T)\,|\, \mathcal{T}_{\theta_0,\widetilde{\omega},K}^{\perp}\right),$$ where $\widetilde{\Pi}$ denotes the projection operator under the working models, that is,

align[align omitted — 307 chars of source]

We next proceed similarly for $\mu(\theta,\omega)$: the counterpart to the efficient score $\varphi_{\theta_0,\omega_0}^{\rm eff}$ is

align[align omitted — 244 chars of source]

Finally, consider the following moment restrictions for $\theta_0$ and $\mu_0=\mu(\theta_0,\omega_0)$:

align[align omitted — 305 chars of source]

where we note that the expectations are taken under the true DGP $(\theta_0,\omega_0)$.

We can now state the following result.

theoremFor any working models $\widetilde{\omega}\in\Omega$ such that $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ are absolutely integrable under DGP $(\theta_0,\omega_0)$, the moment restrictions (ref) and (ref) hold. Moreover, if $\widetilde{\omega}=\omega_0$, then (ref) and (ref) coincide with the efficient moment restrictions for $\theta_0$ and $\mu_{0}$.

Theorem (ref) articulates a locally efficient approach to estimation. The moment functions $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ have the FHR property, irrespective of whether the working models used to derive them actually characterize the sampled population. However, if $\widetilde{\omega}=\omega_0$ “happens to hold" in the sampled population, then (ref) and (ref) coincide with the efficient moment restrictions for $\theta_0$ and $\mu_{0}$ (when $\widetilde{\omega}=\omega_0$ is not part of the prior restriction used to calculate the efficiency bound).

The functions $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ are FHR because calculation (ref) returns an element in the orthogonal complement of the nuisance tangent set by construction. Calculation (ref) provides a principled way to select a particular FHR moment, one that is optimal if the working model happens to hold in the sampled population. Heuristically, the method-of-moments estimator based upon (ref) and (ref) will be more precise when the working models are “approximately true", but this -- to repeat -- is not required for consistency.

In practice $\widetilde{\omega}$ may be a fixed set of working models chosen by the researcher. Alternatively the researcher may posit that these models belong to parametric families $\omega_{\eta}$ indexed by an unknown finite-dimensional (not necessarily scalar) parameter $\eta$. These models may be misspecified, in the sense that there may not exist any $\eta_{0}$ such that $\omega_{\eta_0}=\omega_0$. Nevertheless the moments (ref) and (ref) will be valid for any $\eta$. There are different strategies for picking a particular $\eta$. First, the researcher may choose a particular (non-stochastic) $\eta$ via introspection. Second, she might maximize the likelihood under the working models with respect to $\theta$ and $\eta$. While the resulting estimate of $\theta$ will generally be inconsistent, the corresponding estimate of $\eta$ can be used to define the working models $\widetilde{\omega}$ under which $\widetilde{\phi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ and $\widetilde{\varphi}^{\rm eff}_{\theta_0,\widetilde{\omega}}$ are calculated.

A third approach is to select $\eta$ by maximizing an empirical counterpart to the information for $\theta_0$, as in lindsay1985using. Let $\widehat{\eta}$ be such an estimator of $\eta$, and let $\eta^*$ be its large-$N$ probability limit. If one defines $\widetilde{\omega}=\omega_{{\eta}^*}$, then Theorem (ref) implies that ((ref))-((ref)) are satisfied at true parameter values $\theta_0$ and $\mu_{0}$ (under absolute integrability of the functions). This suggests that the GMM estimators $\widehat{\theta}$ and $\widehat{\mu}$ based on (ref) and (ref) that uses $\omega_{\widehat\eta}$ in lieu of $\widetilde{\omega}$ is consistent and asymptotically normal under standard conditions. We leave details about efficient estimation, using working models of increasing dimensions (i.e., “sieves”), to future work.

Lastly, it is important to stress that our approach based on working models is fundamentally different from (parametric) random-effects maximum likelihood estimation. Indeed, plugging in misspecified parametric models $\omega_{\eta}$ into the likelihood function ((ref)), and maximizing that likelihood with respect to $\theta$ and $\eta$, generally results in an inconsistent estimator of $\theta_0$. In contrast, the moment restrictions (ref) and (ref) remain valid even when the working models are (globally) misspecified.

Efficiency in the multi-spell MPH with unrestricted feedback

In this section we specialize the general results presented above in order to analyze semiparametric efficiency in the MPH model. We focus on the $T=2$ special case in what follows.

Efficiency bounds analysis

Efficient score for $\mathbf{\theta}$. Using Lemma (ref) and Theorem (ref), we show in Supplemental Appendix (ref) that, for \( T = 2 \), the efficient score for $\theta$ in the MPH model with feedback equals

align[align omitted — 309 chars of source]

With some additional algebra we show that the efficient score for the \( \beta \) subvector is

align[align omitted — 620 chars of source]

The first term in (ref) does not involve the feedback process and resembles the efficient score for $\beta$ under strict exogeneity originally derived by hahn1994efficiency:

align[align omitted — 295 chars of source]

The second and third terms of (ref), in contrast, are specific to the feedback case, involving averages over the second-period covariate $X_{2}$. More generally, the efficient score for $\theta$ under strict exogeneity equals:

align[align omitted — 325 chars of source]

In the presence of feedback and latent heterogeneity, $X_2$ is an endogenous variable and cannot be conditioned on. \\

Efficient estimation of the ASH. An interesting average effect in the context of the MPH is the average structural hazard (ASH) defined in (ref). The latter is identified by the FHR moment function in (ref), namely $\varphi_{\theta_0}(Y^T,X^T)=\lambda_{\alpha_{0}}(y_{t})e^{x_{t}'\beta_{0}+\gamma_{0} y_{t-1}}\frac{T-1}{\overline{P}_{\theta_{0}}}$. Applying Lemma (ref) in the Supplemental Appendix, which characterizes projections onto the orthocomplement of the tangent set in the MPH model, one can readily show that $\Pi(\varphi_{\theta_0}(Y^T,X^T)\,|\,\mathcal{T}_{\theta_0,\omega_0,L}^{\perp})=0$. It then follows from the discussion in Section (ref) that the efficient moment function for the ASH is

align*[align* omitted — 198 chars of source]

In turn, a semiparametrically efficient estimator of the ASH is

align*[align* omitted — 232 chars of source]

where $\widehat{\theta}=(\widehat{\alpha},\widehat{\beta}',\widehat{\gamma})'$ is semiparametrically efficient for $\theta$.\footnote{Note that this coincides with the method of moments estimator of brown1998efficient when the condition \( \varphi_{\theta_0}(Y^T, X^T) - \mu_0 \in \mathcal{T}_{\theta_0,\omega_0,L} \) holds, where \( \mu_0 =\mu(\theta_0,\omega_0)\) denotes the average effect of interest and \( \varphi_{\theta_0} \) is an identifying moment function. This condition is satisfied in the case of the ASH for the choice $\varphi_{\theta_0}(Y^T,X^T)=\lambda_{\alpha_{0}}(y_{t})e^{x_{t}'\beta_{0}+\gamma_{0} y_{t-1}}\frac{T-1}{\overline{P}_{\theta_{0}}}$.}\\

Numerical illustrations

In this subsection we summarize our findings from two numerical experiments designed to (i) illustrate the efficiency loss associated with accommodating feedback, and (ii) assess the performance of locally efficient estimators based on working models. Our context is the MPH model. In both experiments we impose a Weibull baseline hazard, set $T=2$, and fix the common parameter at $\theta_{0}=(\alpha_{0},\gamma_{0},\beta_{0})=(\frac{3}{4},\frac{3}{4}\ln 2, -\frac{1}{10})$. Our experiments are meant to numerically approximate the asymptotic precision of various methods of estimation, not to assess the accuracy of such approximations in small samples. Details on implementation can be found in Supplemental Appendix (ref).

The initial duration is drawn from an exponential distribution: $Y_{0}\sim \text{Exponential}(\frac{3}{2})$, and the first-period covariate is a randomized binary treatment: $X_{1}\sim \text{Bernoulli}(\frac{1}{2})$. The heterogeneity distribution equals $V= e^{A}\sim \text{Gamma}(5,5)$, independent of $Y_0,X_1$. The second-period covariate, $X_2$, is a Bernoulli random variable with success probability $p(Y_0,Y_1,X_1,V)$, specified differently across the two experiments to reflect alternative assumptions about the DGP:

itemize• In Experiment (A), we set $p(Y_0,Y_1,X_1,V)=1-\exp(-(Y_{0}+X_{1})V)$, which produces a DGP with strictly exogenous covariates but correlated unobserved heterogeneity. • In Experiment (B), we instead let $p(Y_0,Y_1,X_1,V)=1-\exp(-(Y_{0}+X_{1}+Y_{1})V)$, which introduces feedback from past outcomes to future covariates, while continuing to include correlated heterogeneity.

One goal of our experiments is to assess the efficiency loss that arises when a researcher wishes to accommodate the possibility of heterogeneous feedback, but no such feedback is actually present in the sampled population. Put differently, this exercise gives us a sense of the benefit, in terms of asymptotic precision, of using the strict exogeneity assumption when it is valid (as is the case for the DGP in Experiment (A)).

We compare the asymptotic standard errors of two GMM estimators: the first uses the efficient score under strict exogeneity (ref), while the second uses the efficient score under feedback (ref).

The close connection between our FHR moment characterization (Theorem (ref)) and the relevant semiparametric efficiency bound theory (Theorem (ref)), raises interesting practical questions regarding estimation. While many possible FHR moments are available in the MPH setting, the precision with which they recover $\theta$ varies with the population values of the feedback process and heterogeneity distribution.

The complex form of the efficient score (those for the baseline hazard parameter, $\alpha$, and the coefficient on the lagged duration, $\gamma$ -- both reported in Supplemental Appendix (ref) -- are particularly complicated), suggests that crafting a globally efficient estimator would be difficult. As a principled, yet practical, alternative, we instead explore the properties of a locally efficient estimator based upon simple working models for $\widetilde{\omega}=(\widetilde{g},\widetilde{\pi})$ (a model for $\nu$ is not needed in this case).

Our chosen working models are deliberately rudimentary, intended to illustrate how a researcher might build parsimonious working models while retaining favorable efficiency propertie in more realistic settings. Specifically, in contrast to what prevails in the sampled population, the working model for the feedback process maintains that $X_{2}\sim\mathrm{Bernoulli}(p)$, for a constant $p$. Observe that the working model for the feedback process involves no feedback. For the heterogeneity distribution, we calculate the locally efficient score under a vague Gamma prior of $\widetilde{\pi}(v) = \frac{1}{v} \mathds{1}\{v > 0\}$, independent of initial conditions.

Putting all these pieces together yields the following locally efficient score for $\beta$:

align*[align* omitted — 281 chars of source]

which is simpler than its optimal counterpart (ref) yet, of course, still feedback and heterogeneity robust. Additional details, along with the full expression for the score $\widetilde{\phi}_{\theta_0,\widetilde{\omega}}^{\rm eff}(Y_0,\widetilde{P}_{\theta_0,1}, \overline{P}_{\theta_0}, X_1)$, are provided in Supplemental Appendix (ref). \\ As a final point of comparison, we also report the limiting standard errors of a just-identified GMM estimator that employs the moment function:

equation[equation omitted — 386 chars of source]

The first entry in (ref) is a mean-zero function that exploits the fact that $\widetilde{P}_{\theta_0,1}\sim U[0,1]$ and leverages symmetry (it is also, coincidentally, a component of the efficient score for $\alpha$ under the Weibull baseline hazard, derived and presented in Supplemental Appendix (ref)). The second and third entries correspond to the log-transformed analogs of (ref). While this set of moment conditions lacks an overt efficiency justification, it reflects a common strategy of choosing a small set of “simple" moments for estimation purposes. We include it primarily to benchmark the efficiency gains provided by the locally efficient approach based on working models.

Table (ref) reports the asymptotic standard errors for each estimate of $\theta_0$ in Experiment (A). We make several observations. First, comparing the first and second rows reveals that accommodating feedback -- when it is, in fact, absent from the DGP -- results in some efficiency loss for the slope coefficient $\beta$, with a 29% increase in standard error. Strict exogeneity is a strong assumption and imposing it, when it is valid to do so, improves asymptotic precision.

Although allowing for feedack degrades the precision with which we can learn $\beta$, this is not really the case for $\alpha$ (the Weibull baseline hazard parameter) and $\gamma$ (the lagged duration dependence parameter) in design (A). This finding is consistent with the structure of the efficient scores: those for $\alpha$ and $\gamma$ are very similar under both strict exogeneity and feedback.\footnote{Compare (ref) to (ref) and (ref) to (ref) in Supplemental Appendix (ref).} By contrast, the efficient score for $\beta$ differs markedly across the two settings.

A second observation, inspecting the third row of the table, is that the precision loss associated with using the locally efficient estimator based on the working models $\widetilde{\omega}$ is moderate. Recall that our working models do not characterize the sampled population in design (A). For $\beta$ and $\gamma$, respectively, we observe a 28% and 26% increase in standard error when comparing rows 2 and 3. The efficiency losses are concentrated on the coefficients for the predetermined covariate and the lagged dependent variable, with only minimal deterioration for the parameter of the baseline hazard $\alpha$.

Finally, the approach based on working models leads to large improvements relative to using the “simple" moment functions (ref). As seen in row 4, the standard errors associated with the simple GMM estimator are substantially larger for all three parameters. For example, the standard error for $\gamma$ is nearly 7 times higher than that obtained using the estimator based on the working models $\widetilde{\omega}$ (compare rows 3 and 4, column 3).

table[table omitted — 592 chars of source]

In Experiment (B), we repeat our analysis, but for a population where heterogeneous feedback is, in fact, present. Accordingly, Table (ref) compares the asymptotic standard errors of the locally efficient estimator based on $\widetilde{\phi}_{\theta_0,\widetilde{\omega}}^{\rm eff,\beta}$ to that of the “simple” GMM estimator based on ((ref)) under the second DGP described above (an estimate based on the efficient score under strict exogeneity would not be consistent in this design). As a benchmark, we use the globally semiparametrically efficient estimator based upon the true efficient score that uses (ref). The fourth column of Table (ref) also reports the corresponding asymptotic standard errors for the average structural hazard (ASH) when using the efficient moment function presented in the previous subsection. The key takeaways mirror those of Table (ref). First, the efficiency loss from relying on locally efficient scores is modest, both for common parameters and for average effects. Second, and perhaps more importantly, locally efficient estimators significantly outperform the alternative of using a set of “simple" moments, reaffirming the practical advantages of approaches based on working models.

table[table omitted — 601 chars of source]

FHR moment restrictions in other models

Our characterizations can be used to find FHR moment functions for many models. We have already analyzed the MPH model in detail. In this section we provide additional analytical examples and discuss how to obtain moment functions more generally. Given that the mathematical structure for model parameters and average effects is similar, in this section we focus our discussion on $\theta$.

Existence of moment functions

For certain models, it may be that the only solution to the system of equations in Theorem (ref), or equivalently in Corollary (ref), is the degenerate one, $\phi_{\theta}=0$. We now provide two examples where only trivial moment functions exist. For simplicity we focus on the $T=2$ case.

example{(Binary Choice Model)} Consider a binary choice logit model with continuous heterogeneity and sequentially exogenous covariates: $$\Pr(Y_{t}=1\,|\, Y_{t-1}=y_{t-1},X_{t}=x_t,A=a;\theta)=\frac{\exp(\gamma y_{t-1}+\beta'x_t+a)}{1+\exp(\gamma y_{t-1}+\beta'x_t+a)},\quad t=1,2.$$ Any valid moment function of $\theta=(\gamma,\beta')'$ needs to satisfy ((ref)), that is, \begin{align*} \sum_{y_2=0}^1\phi_{\theta}(y_0,y_1,y_2,x_1,x_2)\frac{\exp(\gamma y_1+\beta'x_2+a)^{y_2}}{1+\exp(\gamma y_1+\beta'x_2+a)} does not depend on $x_2$. \end{align*} Suppose that $\beta\neq 0$. The only functions $\phi_{\theta}$ satisfying this restriction do not depend on $y_2$ or $x_2$. Hence, by ((ref)) we obtain \begin{align*} \sum_{y_1=0}^1\phi_{\theta}(y_0,y_1,x_1)\frac{\exp(\gamma y_0+\beta'x_1+a)^{y_1}}{1+\exp(\gamma y_0+\beta'x_1+a)}=0, \end{align*} from which we obtain that $\phi_{\theta}=0$. This shows the absence of non-trivial moment restrictions for $\theta$ in this model. Building on Chamberlain_JOE2022,chamberlain2023identification, bonhomme2023identification study the failure of point-identification in binary choice models with sequentially exogenous covariates, and show how to compute identified sets on the parameters and average effects. For this reason, our examples in the next subsections will focus on continuous outcomes.\footnote{As bonhomme2023identification note, imposing assumptions on the feedback process, such as Markovianity, may lead to non-trivial moment restrictions in discrete choice and other models where an approach allowing for fully unrestricted feedback is uninformative. Extending our approach to accommodate such additional assumptions is an interesting topic for future work.}
example{(Random Coefficients Model)} Consider the Gaussian linear random coefficients model \begin{equation} Y_{t}=\gamma Y_{t-1}+B'X_{t}+C+\varepsilon_{t},\quad \varepsilon_{t}\,|\, Y^{t-1},X^t,A \sim {\cal{N}}(0,\sigma^2), \end{equation} where $A=(B',C)'$ is multidimensional. Any moment function on $\theta=(\gamma,\sigma^2)'$ needs to satisfy \begin{equation} \int \phi_{\theta}(y_0,y_1,y_2,x_1,x_2)\exp\left(-\frac{1}{2\sigma^2}\left(y_2-\gamma y_1-b'x_2-c\right)^2\right)\mathrm{d}y_2 does not depend on $x_2$. \end{equation} Note that, if ((ref)) holds, then, for all $b$ and $x_2,\widetilde{x}_2$, $$ \phi_{\theta}(y_0,y_1,y_2+b'x_2,x_1,x_2)=\phi_{\theta}(y_0,y_1,y_2+b'\widetilde{x}_2,x_1,\widetilde{x}_2).$$ This implies that $\phi_{\theta}$ does not depend on $y_2$ or $x_2$. Then ((ref)) implies \begin{equation*} \int \phi_{\theta}(y_0,y_1,x_1)\exp\left(-\frac{1}{2\sigma^2}\left(y_1-\gamma y_0-b'x_1-c\right)^2\right)\mathrm{d}y_1 =0, \end{equation*} from which it follows that $\phi_{\theta}=0$. This shows the only FHR moment function on $\theta$ in this model is the null function. This negative result echoes an example in Chamberlain_JOE2022. For this reason, our examples in the next subsections, which all involve scalar outcomes, will feature one-dimensional unobserved heterogeneity.

It is important to note that, even when there exist non-zero functions $\phi_{\theta}$, $\theta$ may fail to be identified. For example, in the MPH model the covariates may be collinear, in which case identification fails. This is of course not specific to our setting. As in any nonlinear GMM problem, identification needs to be verified on a case-by-case basis, and while rank conditions for local identification of $\theta$ are available, verifying global identification may be difficult. Conversely, it may also be that the only $\phi_{\theta}$ satisfying the conditions of Theorem (ref) is $\phi_{\theta}=0$, yet $\theta$ is point-identified.\footnote{As an example, let $T=2$ and let $Y_{t}=\theta+A X_{t}+\varepsilon_{t}$ with $X_{t}$ continuously distributed on $\mathbb{R}$ and $\varepsilon_{t}\,|\, Y^{t-1},X^t,A \sim {\cal{N}}(0,1)$. Applying a similar logic to that in model (ref), one can show that $\phi_{\theta}=0$ since the assumptions imply that $P(X_{2}=0)=P(X_{1}=0)=0$. However, this is a case where the parameter of interest $\theta$ is identified at 0 since $\lim_{x \to 0} \mathbb{E}[Y_{1}|X_{1}=x]=\theta$ (Graham_Powell_EM12).} However, in that case an implication of our analysis in Section (ref) is that such identification is necessarily irregular and the semiparametric efficiency bound for $\theta$ is zero Chamberlain_JE1986. Lastly, even if point-identification fails identified sets may be informative, as shown in lee2020identification and bonhomme2023identification.

Obtaining new moment conditions

We now illustrate how researchers can apply the two-step procedure underlying Corollary (ref) to derive new moment conditions. In the first step, we construct a function $\psi_{\theta}=\sum_{t=1}^{T-1}\psi_{\theta,t}$, for $\psi_{\theta,t}$ such that

equation[equation omitted — 130 chars of source]

In the second step, in the final time period, we invert the linear integral equation

align[align omitted — 150 chars of source]

to get a FHR moment function $\phi_{\theta}(y^T,x^T)$. Naturally, this last task requires the function $\psi_{\theta}(y^{T-1},x^{T-1},a)$ to lie in the range of the integral operator induced by the parametric model. In many examples of interest, equation ((ref)) can actually be inverted in closed form, yielding explicit expressions for functions $\phi_{\theta}$. As an initial example, consider the Poisson regression model introduced earlier.

example[continues=ex: count_intro] (Moments for the Poisson model) Recall, for $t=1,\ldots,T$, \begin{align*} \left.Y_{t}\right|Y^{t-1},X^{t},A\sim\mathrm{Poisson}\left(\exp\left(\gamma Y_{t-1}+X_{t}'\beta+A\right)\right), \end{align*} with both feedback and heterogeneity unrestricted. For simplicity, consider the case $T=2$ and denote $Z_{t}=(X_t',Y_{t-1})'$ and $\theta=(\beta',\gamma)'$. Following the logic laid out above, we start out by picking a moment function $\psi_{\theta}$ such that $\mathbb{E}\left[\psi_{\theta}(Y_{0},Y_{1},X_{1},A)\,|\,Y_{0},X_{1},A\right]=0$. An example of such a function is provided by the score for period $t=1$ in the likelihood which conditions on $A$, $\psi_{\theta}(y_{0},y_{1},x_{1},a)=z_1\left(y_1-\exp(z_1'\theta+a)\right)$.\\ Next, the challenge is to solve for $\phi_{\theta}$ in ((ref)); here corresponding to finding a solution to \begin{align*} \sum_{y_2=0}^{\infty} \phi_{\theta}(y_{0},y_{1},y_2,x_{1},x_{2})\exp(-\exp(z_2'\theta+a))\frac{\exp(z_2'\theta+a)^{y_2}}{y_2!}=\psi_{\theta}(y_{0},y_{1},x_{1},a). \end{align*} After multiplying by $\exp(\exp(z_2'\theta+a))$ and letting $v=\exp(a)$, this is equivalent to \begin{align*} \sum_{y_2=0}^{\infty} \phi_{\theta}(y_{0},y_{1},y_2,x_{1},x_{2})\exp(y_2z_2'\theta) \frac{v^{y_2}}{y_2!}=\psi_{\theta}(y_{0},y_{1},x_{1},\ln v)e^{v\exp(z_2'\theta)}. \end{align*} This formulation reveals that, for a solution to exist, we require that $v\mapsto \psi_{\theta}(y_{0},y_{1},x_{1},\ln v)e^{v\exp(z_2'\theta)}$ admits a Taylor series at $v=0$, with coefficients given by $\phi_{\theta}(y_{0},y_{1},y_2,x_{1},x_{2})\exp(y_2z_2'\theta)$. Appealing to the uniqueness of the Taylor series, we then infer that: \begin{align} \phi_{\theta}(y_{0},y_{1},y_{2},x_{1},x_{2})=\frac{\partial^{y_2}}{\partial v^{y_2}}\bigg|_{v=0}\, \left[{\psi_{\theta}}(y_{0},y_{1},x_{1},\ln v)e^{v\exp(z_2'\theta)}\right]\exp(-y_2z_2'\theta). \end{align} By Corollary (ref), all FHR moment functions for $\theta$ take the form ((ref)), for some appropriate mean-zero function $\psi_{\theta}$. For instance, in the case where $\psi_{\theta}$ is the score in period $t=1$, we obtain \begin{align*} \phi_{\theta}(y_{0},y_{1},y_{2},x_{1},x_{2}) &=z_1\left(y_1-y_2e^{(z_1-z_2)'\theta}\right), \end{align*} which is proportional to the moment function of Chamberlain_JBES92 and wooldridge1997multiplicative. However, using (ref) provides many additional valid moment restrictions on $\theta$. For example, by the second-moment properties of the Poisson distribution, \begin{equation} \psi_{\theta}(y_{0},y_{1},x_{1},a)=\left[y_{1}\left(y_{1}-1\right)-\exp\left(z_{1}'\theta+a\right)^{2}\right]\cdot m\left(z_1\right) \end{equation} is analytic in $v=\exp(a)$ and also satisfies (ref), from which we get the moment \begin{equation} \phi_{\theta}(y_{0},y_{1},y_{2},x_{1},x_{2})=\left[y_{1}\left(y_{1}-1\right)-\frac{y_{2}\left(y_{2}-1\right)\exp\left(z_{1}'\theta\right)^{2}}{\exp\left(z_{2}'\theta\right)^{2}}\right]\cdot m\left(z_{1}\right). \end{equation} Additional FHR moment functions, based upon higher order moments of the Poisson distribution, are straightforward to construct.\footnote{Moreover, by Theorem (ref), the functions $\phi_{\theta}$ in (ref) span the orthocomplement of the tangent set. This suggests that one could compute the efficient score for $\theta$ by projecting the $\theta$-score on that set of functions, although we leave the derivation of the precise form of the efficient score to future work.}
example{(Mixed Interactive Hazard (MIH) Model)} As an example of a new model, one for which no valid FHR moment conditions are known, consider the following Mixed Interactive Hazards (MIH) model. The MIH model relaxes the proportionality assumption of the MPH model; the conditional density of the $t$-th spell equals: \begin{align}f_{\theta}(y_{t}\,|\, y_{t-1},x_t,a)&=\exp(\gamma y_{t-1}+x_t'\beta+a+\left(x_t'\delta\right)\cdot a)\lambda_{\alpha}(y_t)\notag\\ &\quad \times \exp\left(-\exp(\gamma y_1+x_t'\beta+a+\left(x_t'\delta\right)\cdot a)\Lambda_{\alpha}(y_t)\right),\end{align} where $\theta=(\alpha',\gamma,\beta',\delta')'$. Note that (ref) simplifies to (ref) when $\delta=0$. However, $\delta\neq 0 $ allows for more general interaction effects between the covariate and unobserved heterogeneity. The MIH model is an example of a “generalized hazards” model (bonev2020nonparametric). Let $\psi_{\theta}(y_0,y_1,x_1,a)$ be a function satisfying (ref). The integral equation ((ref)), after employing the change of variable $y_2\mapsto p_2$, equals \begin{align*} \int_0^{\infty} \phi_{\theta}(y_0,y_1,p_2,x_1,x_2)e^{(1+x_2'\delta) a}\exp\left(-e^{(1+x_2'\delta) a}p_2\right)\mathrm{d}p_2=\psi_{\theta}(y_0,y_1,x_1,a). \end{align*} Multiplying both sides by $e^{-(1+x_2'\delta) a}$, this is equivalent to \begin{align*} {\cal{L}}\left[\phi_{\theta}(y_0,y_1,\cdot,x_1,x_2)\right]\left(e^{(1+x_2'\delta)a}\right)=e^{-(1+x_2'\delta) a}\psi_{\theta}(y_0,y_1,x_1,a), \end{align*} where ${\cal{L}}[g](s)=\int_0^{+\infty} g(z)\exp(-sz)\mathrm{d}z$ denotes the Laplace transform operator. Letting $s=e^{(1+x_2'\delta)a}$, we effectively wish to solve $${\cal{L}}\left[\phi_{\theta}(y_0,y_1,\cdot,x_1,x_2)\right](s)=s^{-1}\psi_{\theta}\left(y_0,y_1,x_1,\frac{\ln s}{1+x_2'\delta}\right),$$ and, provided $s \mapsto s^{-1}\psi_{\theta}\left(y_0,y_1,x_1,\frac{\ln s}{1+x_2'\delta}\right)$ lies in the range of $\cal{L}$, we can back out $\phi_{\theta}$ using the inverse Laplace transform. To illustrate, take $$\psi_{\theta}(y_0,y_1,x_1,a)=p_1-\exp(-(1+x_1'\delta)a),$$ from which we obtain $${\cal{L}}\left[\phi_{\theta}(y_0,y_1,\cdot,x_1,x_2)\right](s)=s^{-1}p_1-s^{-\frac{1+x_1'\delta}{1+x_2'\delta}-1},$$ which, provided $\frac{1+\delta'x_1}{1+\delta'x_2}>-1$,\footnote{A similar restriction on predetermined covariates features in the moment restrictions of the censored regression model of honore2004estimation.} admits the solution $$\phi_{\theta}(y_0,y_1,y_2,x_1,x_2)=p_1-\frac{p_2^{\frac{1+x_1'\delta}{1+x_2'\delta}}}{\Gamma\left(1+\frac{1+x_1'\delta}{1+x_2'\delta}\right)}.$$ where $\Gamma$ is the Gamma function. This gives the FHR moment function \begin{equation} \phi_{\theta}(y_{0},y_{1},y_{2},x_{1},x_{2})=\left[p_{1}-\frac{p_{2}^{\frac{1+x_{1}'\delta}{1+x_{2}'\delta}}}{\Gamma\left(1+\frac{1+x_{1}'\delta}{1+x_{2}'\delta}\right)}\right]\cdot m(y_0,x_1). \end{equation} More generally, one can obtain closed-form expressions if we choose $\psi_{\theta}$ as a polynomial function of $p_1$.\footnote{For example, we can use for any $b>0$, $$\psi_{\theta}(y_0,y_1,x_1,a)=p_1^b-\exp(-b(1+x_1'\delta)a)\Gamma(1+b),$$ which has zero mean, and gives the FHR moment functions $$\phi_{\theta}(y_{0},y_{1},y_{2},x_{1},x_{2})=\left[p_{1}^b-\frac{\Gamma(1+b)}{\Gamma\left(1+b\frac{1+x_{1}'\delta}{1+x_{2}'\delta}\right)}p_{2}^{b\frac{1+x_{1}'\delta}{1+x_{2}'\delta}}\right]\cdot m(y_0,x_1),$$ which provides a continuum of possible moment functions on $\theta$.}

Example (ref) illustrates how, using Corollary (ref), one can derive moment functions by operator inversion. When suitable functions $\psi_{\theta}$ exist, closed-form inversion delivers explicit moment functions. In other settings, it may be that the inverse is not available in closed form, and numerical inversion techniques need to be used (see, e.g., engl1996regularization).

Irregular moment conditions

In this section we have described an approach, based on Corollary (ref), to construct moment functions $\phi_{\theta}$ when those are available. However, for those functions to be helpful for estimation they need to be sufficiently regular. In this last part we provide examples that show how irregularity may arise, and how regularization techniques can help.

example[continues=ex: mph_intro] (Irregularity in the MPH Model) As an example, consider applying Corollary (ref) to the MPH model, where we again focus on the two-period case for simplicity. We wish to solve for $\phi_{\theta}$ in $${\cal{L}}\left[\phi_{\theta}(y_0,y_1,\cdot,x_1,x_2)\right]\left(e^{a}\right)=e^{- a}\psi_{\theta}(y_0,y_1,x_1,a),$$ that is, letting $s=e^{a}$, in $${\cal{L}}\left[\phi_{\theta}(y_0,y_1,\cdot,x_1,x_2)\right](s)=s^{-1}\psi_{\theta}\left(y_0,y_1,x_1,\ln s\right).$$ Suppose in this case that we take $\psi_{\theta}$ to be the score of the parametric model with respect to $\theta$, that is, $$\psi_{\theta}\left(y_0,y_1,x_1,a\right)=z_1(1-p_1e^a).$$ Now, the solution to $${\cal{L}}\left[\phi_{\theta}(y_0,y_1,\cdot,x_1,x_2)\right](s)=s^{-1}z_1(1-p_1s)$$ is $$\phi_{\theta}(y_0,y_1,y_2,x_1,x_2)=z_1\left(1-p_1 \cdot\delta(p_2)\right),$$ where $\delta(\cdot)$ is Dirac's delta. While $\phi_{\theta}$ has zero expectation, it is a highly irregular function that cannot be directly used in GMM estimation. A possible strategy to address this issue is to regularize $\phi_{\theta}$ by replacing $\delta(p_2)$ with $h^{-1}\kappa(p_2/h)$, where $\kappa$ is a nonparametric kernel and $h>0$ a bandwidth parameter. However, for fixed $h$, the regularized function $\phi_{\theta}$ is no longer mean-zero, necessitating $h$ to shrink to zero as the sample size increases to ensure consistent estimation of $\theta$. In the MPH model, it turns out that these difficulties can be entirely avoided. A rich set of regular moment functions exists, and, in fact, we have characterized the efficient moment function for this model.\footnote{These issues are not unique to the feedback setting. In the context of panel logit models with strictly exogenous covariates, the conditional likelihood estimator of honore2000panel can also be interpreted as relying on an irregular moment condition and requires kernel methods. Only recently did honore2024moment demonstrate the existence of regular moment functions for this class of models. We are grateful to Manuel Arellano for this observation.}

The regularization strategy we have outlined in the context of the MPH model can be useful in more complex models. A general strategy, when solving for $\phi_{\theta}$ in the integral equation ((ref)), is to use a regularized inverse of the relevant integral operator as in carrasco2007linear. We now describe an example where this strategy can be successfully applied.

example{(Nonlinear Regression Model)} Consider the nonlinear panel data regression model \begin{equation} Y_{t}=m_{\beta}(Y_{t-1},X_{t},A)+\varepsilon_{t},\quad \varepsilon_{t}\,|\, Y^{t-1},X^t,A\sim {\cal{N}}(0,\sigma^2), \end{equation} where $a\mapsto m_{\beta}(y_{t-1},x_t,a)$ is differentiable and strictly increasing. Here $\theta=(\beta',\sigma^2)'$. Let $\lambda>0$, and let \begin{align*} K_{\lambda}(z)=\frac{1}{2\pi }\int \lambda\kappa(\lambda \tau) \exp\left(\lambda \boldsymbol{i}\tau z+\frac{1}{2}\sigma^2\tau^2\right)\mathrm{d}\tau, \end{align*} where $\kappa$ is the Fourier transform of a kernel function, satisfying $\kappa(0)=1$ and $\kappa'(0)=0$, $\abs{\kappa}$ is integrable, and where $\boldsymbol{i}$ is the imaginary number. A possible choice is $\kappa(\tau)=\boldsymbol{1}\{\tau\in(-1,1)\}$, which is the Fourier transform of the sinc kernel. $K_{\lambda}(z)$ corresponds to the deconvolution kernel introduced by stefanski1990deconvolving in the context of a deconvolution problem with normal measurement error. Given any mean-zero function $\psi_{\theta}(y_0,y_1,x_1,a)$, we define \begin{equation}\phi_{\theta}^{\lambda}(y_0,y_1,y_2,x_1,x_2)=\int \psi_{\theta}(y_0,y_1,x_1,a)\frac{\partial m_{\beta}(y_1,x_2,a)}{\partial a}K_{\lambda}\left(\frac{m_{\beta}(y_1,x_2,a)-y_2}{\lambda}\right)\mathrm{d}a.\end{equation} Applying Corollary (ref) and a regularization strategy, we show in Supplemental Appendix (ref) that \begin{equation}\mathbb{E}_{\theta,\omega}\left[\phi^{\lambda}_{\theta}(Y_0,Y_1,Y_2,X_1,X_2)\right]\rightarrow 0 \, as \,\lambda\rightarrow 0.\end{equation} In this sense, $\phi_{\theta}^{\lambda}$ provides an approximately valid moment function for $\theta$. As a special case, consider the linear Gaussian model where $m_{\beta}(y_{t-1},x_t,a)=\gamma y_{t-1}+x_t'\beta+a$, and take $\psi_{\theta}(y_0,y_1,x_1,a)=z_1(y_1-\gamma y_0-x_1'\beta-a)$. Then (ref) simplifies to \begin{align*}\phi_{\theta}^{\lambda}(y_0,y_1,y_2,x_1,x_2)&=\int z_1(y_1-\gamma y_0-x_1'\beta-a)K_{\lambda}\left(\frac{\gamma y_1+x_2'\beta+a-y_2}{\lambda}\right)\mathrm{d}a\\ &=z_1\left[(y_1-\gamma y_0-x_1'\beta)-(y_2-\gamma y_1-x_2'\beta)\right],\end{align*} which corresponds to the Arellano-Bond moment function for the case $T=2$. Note that $\phi_{\theta}^{\lambda}$ does not depend on $\lambda $ in this case, so the regularization is immaterial.