EconBase
← Back to paper

Back to Feedback: Dynamics and Heterogeneity in Panel Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

123,977 characters · 35 sections · 159 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Back to Feedback Dynamics and Heterogeneity in Panel Data

\vskip 1cm

abstractMany popular estimation methods in panel data rely on the assumption that the covariates of interest are strictly exogenous. However, this assumption is empirically restrictive in a wide range of settings. In this paper I argue that credible empirical work requires meaningfully relaxing strict exogeneity assumptions. Econometricians have developed methods that allow for sequential exogeneity, which in contrast with strict exogeneity allows for the presence of feedback from past outcomes to future covariates or treatments. I review some of the classic work on linear models with constant coefficients, and then describe some approaches that allow for coefficient heterogeneity in models with feedback. Finally, in the last two parts of the paper I review recent work that allows for sequential exogeneity in nonlinear panel data models, and mention possible extensions to network settings. JEL codes: C10, C50. Keywords: Dynamic Feedback, Panel Data, Difference in Differences.

\global\long\global\long\global\long\global\long\global\long

Introduction

A central concern during the past thirty years in applied (micro) economics has been the presence of un-modeled heterogeneity. Some of the most influential methodological advances, such as work on local average treatment effects (imbens1994estimation) and difference-in-differences (abadie2005semiparametric, goodman2021difference) have been motivated by the desire to allow for rich individual heterogeneity in treatment responses.

However, if one looks back to the pre-“empirical credibility revolution” period, a chief concern of econometricians then was to correctly specify dynamic processes of outcome and covariates. David Hendry's textbook summarizes many of the advances made during that previous era (hendry1995dynamic). As it stands, the focus on unobserved heterogeneity appears to have shifted attention away from earlier concerns about dynamics.

This situation is unfortunate. Presumably, both dynamics and heterogeneity are present in many of the empirical settings that applied economists study. Overlooking one misspecification concern to solely focus on another one hardly seems to be a recipe for credible research. Yet, today's applied work, though often based on methods that are mindful of heterogeneity, overwhelmingly relies on restrictive assumptions about un-modeled dynamics of covariates and outcomes.

A key concept when introducing dynamics is the feedback process -- the way past outcomes and covariates influence (i.e., “feed back” into) later covariates. When a covariate is a treatment of interest, this process is a dynamic analog to the propensity score, which depends on past values of outcome and treatment. In panel data applications, the feedback process may additionally vary between individuals.

Allowing for feedback appears a priori reasonable: after all, why should a covariate today (for example, the value of the minimum wage) not depend on past outcomes (such as employment or earnings)? Yet, a large body of empirical work relies on an assumption that rules out feedback entirely. This assumption, known as strict exogeneity (SE), is at the heart of the textbook justification for two-way fixed effects estimators, and it is either explicitly or implicitly assumed in most applications of difference-in-differences.

Several arguments are commonly mentioned to support an assumption of strict exogeneity. One argument is based on the use of pre-trend checks. Another argument is that feedback is less likely when the treatment is aggregate (such as a policy reform) and outcomes are individual-specific. I will argue that neither arguments are fully compelling, and that feedback bias is likely present in many difference-in-differences settings. Rather than dismissing outright the possibility that the treatment may in fact respond to past outcomes, it seems reasonable to study how one can address the presence of feedback empirically.

The goal of this article is to highlight existing research on un-modeled dynamics and feedback. The working assumption is that of sequential exogeneity (SeqE). Under sequential exogeneity, the errors in the regression are assumed to be uncorrelated with past and current covariates, however they are allowed to correlate with future covariates. Thus, while SeqE may still be restrictive since it rules out contemporaneous selection, it is less restrictive than SE since it allows future treatments to be chosen based on past outcomes. This is a crucial, yet often under-appreciated, difference.

After introducing the main concepts and describing their implications for common estimators, I survey some “classic” approaches in the literature on dynamic panels. This includes the popular estimators of arellano1991some and blundell1998initial; see arellano2003panel for a textbook treatment. However, the methods developed in this literature suffer from two main drawbacks. A first challenge is the proliferation of instruments in dynamic settings and the resulting instability of estimators. I review several approaches that have been proposed to improve over the original dynamic panel methods. A second challenge is that classic dynamic panel methods are vulnerable to the presence of un-modeled heterogeneity. Although heterogeneity in the model is limited to the intercept (the so-called “fixed effect”, which is differenced out), heterogeneity in coefficients -- reflecting treatment effects heterogeneity -- may be empirically prevalent, as emphasized by marx2024heterogeneous.

An important goal of the paper is to cover some of the “modern” approaches to models with feedback, including approaches that allow for coefficient heterogeneity. The starting point is a seminal paper by Gary Chamberlain (chamberlain2022feedback), which was first circulated in the 1990s and remains highly relevant to today's methodological and empirical work. This article establishes a negative result, showing that, in a simple model with a binary sequentially exogenous covariate and two periods, averages of coefficients (i.e., average treatment effects on subpopulations) are not identified. This stands in contrast with the strictly exogenous case, where average effects for individuals whose covariate changes over time (the so-called “movers”) are identified (e.g., chamberlain1992efficiency).

There are two possible reactions to Chamberlain's negative result. The first one is to focus on quantities that are identified, as in bonhomme2025unrestricted. The second one is to relax the identification requirement and bound the estimand of interest, as exemplified by lee2020identification. Since addressing dynamic aspects is in my view key for credible inference, more work is needed in this area, with the aim to account for heterogeneity similarly to the “new difference-in-differences literature” (e.g., de2023two, sun2021estimating, callaway2021difference, borusyak2024revisiting) while relaxing the often implausible assumption of strict exogeneity.

Another “modern” question in the literature is to relax the assumption that the model is linear in parameters. Deriving valid moment restrictions on parameters is challenging in nonlinear settings, due to the twin presence of heterogeneity and feedback. Solutions exist in specific models, such as Poisson regressions and multiplicative proportional hazard models. I also review recent work by bonhomme2025feedback that extends the functional differencing approach introduced in bonhomme2012functional to models with feedback. Following bonhomme2023identification, I further explain how the identified set of parameters can be characterized using linear programming, generalizing the analysis in honore2006bounds to models with feedback.

Lastly, an important avenue is the study of dynamic network settings. Popular estimators such as the AKM estimators of abowd1999high have been used in a variety of settings to understand the sources of wage dispersion, among many other questions. However, the validity of this approach hinges on strict exogeneity. Allowing for feedback in network contexts may be as or more relevant than in single-agent panel data contexts. For example, the assumption in AKM that job mobility does not depend on past wages is controversial. Although the literature on the topic is still in its infancy, I conclude the article by mentioning some recent research and possible approaches.

Sequential exogeneity and feedback

Some definitions

Strict and sequential exogeneity are central concepts in panel data. To introduce them, it is useful to start with a linear model for a single time series (that is, focusing on a single individual in the panel). Specifically, consider a linear time-series model of the form

equation[equation omitted — 83 chars of source]

where $Y_t$ is a scalar outcomes and $X_t$ is a vector of covariates.

definition{(strict and sequential exogeneity, linear models)}$\quad$ $X_t$ is strictly exogenous (SE) if \begin{equation}\mathbb{E}[U_t\,|\, X_1,...,X_T]=0, \quad t=1,...,T.\end{equation} $X_t$ is sequentially exogenous (SeqE) if \begin{equation}\mathbb{E}[U_t\,|\, X_1,...,X_t]=0, \quad t=1,...,T.\end{equation}

The two types of exogeneity introduced in Definition (ref) are classical concepts in time series analysis, and they can be found under various names in the literature. Strict exogeneity is often referred to as “strong exogeneity”, and contrasted with “weak” (i.e., sequential) exogeneity, see for example engle1983exogeneity. Another common term for sequential exogeneity is “predeterminedness”.

Strict exogeneity requires covariates $X_s$ in all periods $s=1,...,T$ to be unrelated to the outcome disturbances $U_t$. In contrast, sequential exogeneity only requires covariates $X_s$ in the past and current periods $s=1,...,t$ to be unrelated to $U_t$. The key difference between the two assumptions is that, unlike SE, SeqE allows for feedback: outcome realizations associated with some shocks $U_t$ are allowed to influence future realizations of the covariates $X_{t+1},...,X_T$.

Consider next the case where the model includes a lagged outcome,\footnote{As a convention, I will assume throughout that $Y_0$ is part of the observation sample.}

equation[equation omitted — 110 chars of source]

In this dynamic model, strict and sequential exogeneity of $X_t$ are typically assumed to hold conditional on the history of past outcomes. Hence, SE becomes

equation[equation omitted — 99 chars of source]

while SeqE becomes

equation[equation omitted — 99 chars of source]

Note that ((ref)) and ((ref)) imply that $U_t$ are serially uncorrelated.

It is possible to generalize these definitions to models that are nonlinear and feature arbitrary heterogeneity. Doing so is useful given the widespread use of strict exogeneity assumptions for causal inference in difference-in-differences settings. Let $Y_t(x)$ denote the potential outcome for a given value of the covariate $X_t=x$. For example, in a nonlinear dynamic model $Y_t=m( Y_{0},...,Y_{t-1},X_1,...,X_t,U_t)$, the potential outcome is $Y_t(x)=m( Y_{0},...,Y_{t-1},X_1,...,x,U_t)$.\footnote{Here $X_t$ could indicate the history of the covariate to accommodate lag effects. Moreover, one could explicitly introduce the lagged outcome $Y_{t-1}$ using the potential outcome notation $Y_t(x_t,y_{t-1})$ so as to model state dependence, as in torgovitsky2019nonparametric.}

In this general setup, one can write sequential exogeneity as the following conditional independence assumption:

equation[equation omitted — 412 chars of source]

That is, the treatment is assumed to be independent of future potential outcomes conditional on the history of treatment and past outcomes. In a time-series setting, Condition ((ref)) coincides with the assumption of sequential exchangeability (robins1986new, marx2024heterogeneous). However, in panel data, we will assume sequential exogeneity to hold conditional on latent heterogeneity (i.e., “fixed effects”) as well.

Condition ((ref)) is helpful to understand the economic content of sequential exogeneity. Under ((ref)), the agent chooses $X_t$ given her information set, which includes her past treatment and outcome realizations (as well as heterogeneity and possibly other state variables included in $X_t$), but does not include current or future potential outcomes. This precludes advance information about returns, for example. In contrast, in this general setup one can write strict exogeneity as

equation[equation omitted — 387 chars of source]

where the information set does not include any of the potential outcomes $Y_{t+s}(x)$, past, present or future. This precludes agents responding to past outcomes when choosing $X_t$, and requires the choice problem to be completely unrelated to outcome realizations.

In this setting, the feedback process is defined as follows.

definitionThe feedback process is the conditional density of $$X_t\,|\, Y_0,...,Y_{t-1},X_{1},...,X_{t-1}.$$

In words, the feedback process is the conditional density of current covariates given past outcomes and covariates (and, in panel data, latent heterogeneity). Under SE, the feedback process is simply the conditional density of the current covariate given the history of covariates. From a time-series perspective, strict exogeneity thus rules out Granger causality from past $Y_s$, for $s=1,...,t-1$, to future $X_t$.\footnote{ chamberlain1982general shows that the classic definition of strict exogeneity through conditional independence,

equation*[equation* omitted — 216 chars of source]

is equivalent to the following no feedback condition:

equation*[equation* omitted — 220 chars of source]

}

While more general than strict exogeneity, sequential exogeneity is a potentially restrictive assumption as well. Importantly, sequential exogeneity cannot accommodate the presence of simultaneity (when $X_t$ and $Y_t$ simultaneously determined) or serially-correlated time-varying counfounders (e.g., some unobserved covariate $V_t$ in ((ref)) that is correlated with $X_t$). In both these situations, instruments $Z_t$ external to the model would be required for identification. Nonetheless, the appeal of sequential exogeneity is that it captures any empirically plausible mechanisms through which past disturbances $U_{t-1},U_{t-2},...$ may correlate with current covariates $X_t$.

figure[figure omitted — 1,920 chars of source]

Figure (ref) represents the strict and sequential exogeneity assumptions graphically. In the left panel, the solid arrows indicate that $X_t$ has a direct effect on $Y_t$. In addition, the lagged outcome $Y_{t-1}$ may also directly affect $Y_t$, for example if the model includes a lag as in ((ref)). The graph in the left panel also features dashed arrows, i.e., possible relationships between covariates $X_t$ over time. However, it does not feature any link between past outcomes, such as $Y_{t-1}$, and current covariates $X_t$. In contrast, the right panel in Figure (ref), which represents the sequential exogeneity assumption, does allow for relationships from past outcomes $Y_{t-1}$ to current covariates $X_t$, indicated by dashed arrows.

The perils of strict exogeneity

Under strict exogeneity, the feedback process is conditionally independent of past outcomes. That is, SE rules out the presence of feedback. However, assuming no feedback is restrictive in many economic settings.

Consider a model of job training, as described in arellano2003panel, Chapter 8. Focusing on a single worker, let the earnings of the worker be $Y_t=Y_t^*+{\Greekmath 010C}_t X$, for $Y_t^*$ the earnings in the absence of training, and $X$ a time-invariant binary training indicator. In this setting, strict exogeneity is likely to be violated since there is evidence of a dip in the earnings of training participants before training happens (ashenfelter1985using). One way to account for this dip is to model earnings in the absence of training $Y_t^*$ as a dynamic (e.g., autoregressive) process while allowing for pre-training $Y_{s}^*$ to affect the training indicator $X$. This type of approach based on sequential exogeneity will be the topic on Section (ref) onward.

In fact, feedback is a central mechanism in most dynamic economic models. Consider as an example a consumption model where $Y_t$ denotes household consumption and $X_t$ include household assets. Through the budget constraint, assets today depend on how much the household consumed last period, which mechanically implies the presence of feedback and violates strict exogeneity. More generally, feedback is central to dynamic economic models with forward-looking agents which are popular in industrial organizations (rust1987optimal), labor economics (keane1997career), and many other fields.

Yet, while feedback is economically plausible in many settings, strict exogeneity is routinely assumed in applications. A case in point is the popularity of difference-in-differences methods, which explicitly or implicitly rely on the assumption that the treatment of interest is strictly exogenous (chabe2015analysis, de2022not, ghanem2022selection, marx2024parallel). The next section will describe how strict exogeneity, when assumed in a setting where covariates are only sequentially exogenous, can create important biases in estimation.

Fixed effects

Panel data has become the leading data format used in empirical micro-economics. A common specification accounts for a latent individual-specific intercept (a “fixed effect”), as in the model

equation[equation omitted — 106 chars of source]

In applications, the model is often augmented with a lagged outcome and other covariates, as in

equation[equation omitted — 159 chars of source]

where $W_{it}$ may include time effects and other strictly exogenous or sequentially exogenous covariates.

There are two common versions of sequential exogeneity in this context (abstracting from lagged outcomes $Y_{i,t-1}$ and additional covariates $W_{it}$ for simplicity). The first one is not conditional on the latent effect $A_i$:

equation[equation omitted — 73 chars of source]

while the second one is conditional on the latent effect $A_i$:

equation[equation omitted — 77 chars of source]

One can similarly define unconditional and conditional versions of strict exogeneity.\footnote{That is, $\mathbb{E}[U_{it}\,|\, X_{i1},...,X_{iT}]=0$ (unconditional) and $\mathbb{E}[U_{it}\,|\, X_{i1},...,X_{iT},A_i]=0$ (conditional).} In addition, in dynamic models such as ((ref)), strict and sequential exogeneity are typically assumed to hold conditional on the history of past outcomes.

These notions can be extended to models with coefficient heterogeneity, such as

equation[equation omitted — 98 chars of source]

In model ((ref)), the unconditional and conditional versions of SeqE and SE are defined exactly as above, with $A_i=(B_i',C_i)'$. Section (ref) will consider such extensions allowing for coefficient heterogeneity. This will reveal that, in models with coefficient heterogeneity, the unconditional and conditional versions of SeqE ((ref)) and ((ref)) have -- perhaps unexpectedly -- different implications for (partial) identification of average effects.

In panel data models such as ((ref)), it is common to leave the feedback process $$X_{it}\,|\, Y_{i0},...,Y_{i,t-1},X_{i1},...,X_{i,t-1},A_i$$ unrestricted. This is conceptually appealing since this allows for arbitrary forms of feedback, based on the history of covariates and outcomes (which are often state variables in the economic model), in a way that accommodates arbitrary individual heterogeneity (for example, consistent with heterogeneous technology or preferences).

However, in some settings it may be appealing to impose some assumptions on the feedback process. Markovian feedback, through which $X_{it}$ are independent of the history conditional on $Y_{i,t-1}$, $X_{it}$, and $A_i$, say, is commonly assumed in structural models. Homogeneous feedback, through which $X_{it}$ are independent of $A_i$ given the history $Y_{i0},...,Y_{i,t-1},X_{i1},...,X_{i,t-1}$, is often assumed in sequential experiments (as in robins1986new) and may be plausible in some economic models, such as models with a homogeneous state transition equation. Section (ref) will return to the case of restricted feedback.

Feedback bias

An important statistical consequence of the presence of feedback is that the OLS estimator is biased for finite $T$. This bias affects popular fixed effects regressions in panel data, where the bias of the estimator in the time series translates into panel data inconsistency as the cross-sectional size of the sample $N$ grows while the number of time periods $T$ remains fixed.

Bias in time series

Consider first the time-series model ((ref)). The OLS estimator is given by (assuming that $\sum_{t=1}^T X_tX_t'$ is non-singular): $$\widehat{{\Greekmath 010C}}=\left(\sum_{t=1}^T X_tX_t'\right)^{-1}\sum_{t=1}^T X_tY_t.$$ The OLS estimator $\widehat{{\Greekmath 010C}}$ is unbiased under strict exogeneity, since in this case $$\mathbb{E}[\widehat{{\Greekmath 010C}}]={\Greekmath 010C}+\mathbb{E}\left[\left(\sum_{t=1}^T X_tX_t'\right)^{-1}\sum_{t=1}^T X_t\underset{=0}{\underbrace{\mathbb{E}\left[U_t\,|\, X_1,...,X_T\right]}}\right]={\Greekmath 010C}.$$ However, $\widehat{{\Greekmath 010C}}$ is generally biased under sequential exogeneity.

To see the source of bias, consider the case where there are $T=2$ periods. We have $$\widehat{{\Greekmath 010C}}={\Greekmath 010C}+\underset{(I)}{\underbrace{(X_1X_1'+X_2X_2')^{-1}X_1U_1}}+\underset{(II)}{\underbrace{(X_1X_1'+X_2X_2')^{-1}X_2U_2}}.$$ Under SE, both terms (I) and (II) have zero mean. Under SeqE, (II) still has zero mean, since $\mathbb{E}[U_2\,|\, X_1,X_2]=0$, so $$\mathbb{E}[(II)]=\mathbb{E}\left[(X_1X_1'+X_2X_2')^{-1}X_2\underset{=0}{\underbrace{\mathbb{E}[U_2\,|\, X_1,X_2]}}\right]=0.$$ However, $U_1$ may be correlated with $X_2$ under SeqE because of the presence of feedback. As a result, we generally have $$\mathbb{E}[(I)]=\mathbb{E}\left[(X_1X_1'+X_2X_2')^{-1}X_2\underset{\neq 0}{\underbrace{\mathbb{E}[U_1\,|\, X_1,X_2]}}\right]\neq 0.$$

In a time-series setting, if sequential exogeneity holds, then $\widehat{{\Greekmath 010C}}$ is consistent as $T$ tends to infinity under standard conditions, despite the presence of bias. A key property for consistency of OLS is that a lack of contemporaneous correlation property is satisfied, $$\mathbb{E}[X_tU_t]=0,\quad t=1,...,T.$$ See for example hayashi2011econometrics, Chapter 2.\footnote{Although hayashi2011econometrics refers to $\mathbb{E}[X_tU_t]=0$ as regressors being “predetermined”, this terminology differs from what the panel data literature defines as “predetermined” (i.e., sequentially exogenous) regressors.}

However, the bias can be large for moderate $T$. Recently, mikusheva2023linear study the properties of OLS in linear regression models with a large number of sequentially exogenous covariates relative to sample size. Focusing on an asymptotic regime where the number of regressors increases with $T$, they document that biases may be so large as to render the estimator inconsistent.

Inconsistency in short panel data

The bias of OLS for short $T$ is particularly problematic in empirical applications to short panel data. Consider the linear panel data model with fixed effects

equation[equation omitted — 109 chars of source]

under the assumption that the regressors $X_{it}$ are sequentially exogenous\footnote{Note here that here the sequential exogeneity assumption is not conditional on the fixed effect $A_i$, but the conditional SeqE assumption ((ref)) has the same bias implications. }

equation[equation omitted — 73 chars of source]

Let $\widehat{{\Greekmath 010C}}$ denote the OLS estimator in ((ref)) -- with fixed effects, so $\widehat{{\Greekmath 010C}}$ coincides with the so-called “within-group” estimator. Typically, one can show that, as $N\rightarrow \infty$ for $T$ fixed we have $$\underset{N\rightarrow \infty}{\limfunc{plim}} \, \widehat{{\Greekmath 010C}}={\Greekmath 010C}+C_T\neq {\Greekmath 010C},$$ where $C_T$ is a non-zero constant. To provide additional intuition on the bias, it is informative to expand $C_T$ as $T$ tends to infinity. Under suitable regularity conditions, we obtain the following expansion:

equation[equation omitted — 167 chars of source]

for some constant $C$. This shows that, to some approximation, the bias of $\widehat{{\Greekmath 010C}}$ decays as $T$ increases in a way that is inversely proportional to $T$. Hence, one expects that, at least when $T$ is not too small, the bias will be approximately divided by two when the number of time periods doubles. However, in short panels, the bias can be substantial.\footnote{The OLS estimator in first differences is similarly inconsistent for fixed $T$ in the presence of feedback. However, in contrast with the within-group estimator, it remains inconsistent as $T$ tends to infinity in general (e.g., alvarez2003time).}

An important example is the so-called “Nickell bias” in an autoregressive panel data model where $X_{it}$ is scalar and coincides with the lagged outcome $Y_{i,t-1}$. nickell1981biases derives an explicit expression for the probability limit of $\widehat{{\Greekmath 010C}}$ as $N$ tends to infinity for fixed $T$, and alvarez2003time derive the asymptotic properties of $\widehat{{\Greekmath 010C}}$ as $N$ and $T$ tend to infinity jointly. Under asssumptions that include i.i.d. homoskedastic errors and stationary initial conditions, alvarez2003time show that $$\underset{N\rightarrow \infty}{\limfunc{plim}} \, \widehat{{\Greekmath 010C}}={\Greekmath 010C}-\frac{1+{\Greekmath 010C}}{T}+o\left(\frac{1}{T}\right).$$ More generally, the presence of feedback bias represents a major challenge in panel data models with feedback and heterogeneity.

Feedback bias and practice

The presence of feedback bias has important implications for the practice of difference-in-differences, as this section and the next illustrate.

Dynamic bias in difference-in-differences and event studies

Consider the two-way fixed effects model popular in the current applied literature

equation[equation omitted — 115 chars of source]

in the case of a scalar $X_{it}$ and abstracting from other covariates for simplicity. Consider the OLS estimator with unit and time fixed effects, which coincides with the standard two-way fixed-effects estimator (here for a balanced panel), $$\widehat{{\Greekmath 010C}}^{\rm TWFE}= \frac{\sum_{i=1}^N\sum_{t=1}^T \overset{..}{X}_{it}\overset{..}{Y}_{it}}{\sum_{i=1}^N\sum_{t=1}^T\overset{..}{X}_{it}^2},$$ where $\overset{..}{Z}_{it}=Z_{it}-\frac{1}{T}\sum_{s=1}^TZ_{is}-\frac{1}{N}\sum_{j=1}^NZ_{jt}+\frac{1}{NT}\sum_{j=1}^N\sum_{s=1}^TZ_{js}$ denotes the double difference of $Z_{it}$. Under strict exogeneity, in model ((ref)) with constant coefficients, $\widehat{{\Greekmath 010C}}^{\rm TWFE}$ is unbiased for ${\Greekmath 010C}$, and consistent in short panels under mild additional conditions. However, $\widehat{{\Greekmath 010C}}^{\rm TWFE}$ is generally inconsistent in short panels when $X_{it}$ are sequentially exogenous.

It is important to point out that the nature of the bias is different from the issues highlighted by the recent literature on difference-in-differences (e.g., goodman2021difference, de2023two, sun2021estimating, callaway2021difference). In that literature, the presence of un-modeled heterogeneity in coefficients affects the interpretation the probability limit of popular estimators such as two-way fixed effects. In contrast, here the (homogeneous) coefficient of interest is biased due to the presence of un-modeled dynamics, through feedback. Presumably both forces may be simultaneously at play in the data, and Section (ref) will study models with both feedback and coefficient heterogeneity.

As ashenfelter1985using point out, strict exogeneity may be restrictive in the context of estimating the effect of training programs using difference-in-differences methods. Participation to the program, indicated by $X_{it}=1$, may respond to realizations of past outcomes. In the presence of sequential exogeneity, the two-way fixed-effects estimator is generally inconsistent, with a bias $C_T$ inversely related to the panel length as in ((ref)). However, today's applied work often provides too little discussion of how the treatment was determined and whether it is affected by past outcomes through feedback, which are key to argue for the plausibility of strict exogeneity.

The exact same issue affects event study regressions. Consider as an example the model

equation[equation omitted — 116 chars of source]

which controls for leads and lags of $X_{it}$. This is a popular model for policy evaluation. While the OLS estimators of ${\Greekmath 010C}_{-a},...,{\Greekmath 010C}_b$ (with unit and time fixed effects) are unbiased when $X_{it}$ are strictly exogenous, they are generally biased -- and inconsistent in short panels -- when $X_{it}$ are sequentially but not strictly exogenous. This shows that event study estimates, similarly to two-way fixed-effects estimates, crucially rely on strict exogeneity.

A cautionary note on pre-trend checks

In difference-in-differences, researchers routinely run pre-trend checks to validate their empirical approaches. In some settings, such checks suggest clear violations of strict exogeneity. As a recent example, consider the analysis in acemoglu2019democracy, who study how the spread of democracy during the period 1960-2010 has affected economic growth. Using a panel of countries, they regress GDP per capita on a democracy indicator, lags of GDP per capita, and country and year fixed effects. The inclusion of a country effect accounts for permanent GDP differences between countries, and lags of GDP are included to make the error term serially uncorrelated. The authors document graphically that democratization episodes are, on average, preceded by a temporary dip in GDP, and argue this shows a clear violation of the parallel trends assumption that underlies difference-in-differences and other panel data methods based on strict exogeneity.

However, there is growing evidence against using pre-trend checks as a formal validation of strict exogeneity (or parallel trends) assumptions. roth2022pretest points out two main issues with the current practice of pre-trend checks. First, pre-trend tests often have low power. Second, they are only indirect test of model validity. The limitations of pre-trend checks have been noted by a number of authors, and have motivated some recent methodological developments in difference-in-differences settings (e.g., freyaldenhoven2019pre, rambachan2023more).

Recently, ghanem2022selection revisit the lalonde1986evaluating analysis of the effect of job training on earnings. They find that, while a pre-trend test does not reject strict exogeneity, two-way fixed effects and other difference-in-differences estimators differ substantially from the experimental benchmark. This suggests that, while clear rejections such as suggested by the graphical evidence in acemoglu2019democracy are informative, non-rejections of strict exogeneity should be interpreted with caution. We will return to the limitations of pre-trend checks in the next section, in the context of an example.

Dynamic bias when treatments are aggregate

A common argument against strict exogeneity assumptions is that, when covariates are choice variables, it is natural to expect that they may respond to realizations of lagged outcomes. However, many applications of difference-in-differences and event studies feature aggregate covariates such as state-level policy variables. It seems unlikely that an aggregate treatment, chosen by, say, a local government, will respond to idiosyncratic shocks to individual outcomes. Does this imply that feedback bias is less of an issue in such settings?

To examine this question, consider the following model with an aggregate treatment,

equation[equation omitted — 122 chars of source]

where $j(i)\in\{1,...,J\}$ is, say, the state where $i$ lives, and the treatment $X_{j(i),t}$ is determined at the state level (such as a minimum wage policy). Suppose that $$U_{it}={\Greekmath 0110} V_{j(i),t}+{\Greekmath 0122}_{it},$$ where $V_{j(i),t}$ are shocks common to all individuals in state $j(i)$ at time $t$, ${\Greekmath 0122}_{it}$ are purely idiosyncratic shocks, i.i.d. over $i$ and $t$, and ${\Greekmath 0110}$ is a constant that measures the relative importance of the two components.

It may be realistic to assume that idiosyncratic shocks ${\Greekmath 0122}_{it}$ are mean independent of the aggregate treatment $X_{j(i),s}$ in all periods $s=1,...,T$, i.e., $$\mathbb{E}[{\Greekmath 0122}_{it}\,|\, X_{j(1),1},...,X_{j(i),T}]=0.$$ Hence, in a world without common shocks (i.e., ${\Greekmath 0110}=0$), strict exogeneity of an aggregate treatment appears a priori plausible, and standard difference-in-differences techniques may be justified.

However, common shocks are present in many applications, and it is often restrictive to rule out all types of dynamic dependence of the treatment on them. To see this, suppose that there are many individual observations within a state, and write state-level averages based on ((ref)):

equation[equation omitted — 158 chars of source]

where $\overline{Y}_{jt}$ is the population average outcome in state $j$ at time $t$, and $\overline{A}_{j}$ is the population average of $A_i$ in state $j$. Note that the state-level average of the idiosyncratic shocks ${\Greekmath 0122}_{it}$ is equal to zero. In ((ref)), both treatment and outcome are at the same level of aggregation (the state). Hence, the usual concerns with strict exogeneity assumptions apply.

While it may be unrealistic to think of the decision of a local government to be based on individual outcome realizations $Y_{i,t-s}$, it is natural to expect that average realizations in the state $\overline{Y}_{j,t-s}$ (for example, how state-level employment evolved in the past) may influence how the state sets the minimum wage at time $t$. Now, if sequential exogeneity holds at the state level, $$\mathbb{E}[V_{jt}\,|\, X_{j1},...,X_{jt}]=0,$$ but strict exogeneity does not hold, $$\mathbb{E}[V_{jt}\,|\, X_{j1},...,X_{jT}]\neq 0,$$ and if ${\Greekmath 0110}\neq 0$, then standard difference-in-differences estimators such as two-way fixed-effects will be biased and inconsistent in general.

This discussion highlights that concerns with strict exogeneity apply equally to aggregate treatments, unless the researcher can convincingly rule out the presence of dynamic selection based on aggregate shocks.

Feedback in two- and three-period models

The two-period model with a binary treatment is often presented as the canonical difference-in-differences setting (abadie2005semiparametric). Here we first analyze how the presence of feedback modifies the usual conclusions based on this model, and then discuss how to interpret a lack of pre-trends in a setting with an additional pre-treatment period.

Bias in a two-period model with a binary treatment

Suppose $X_{it}$ is binary, $X_{i1} = 0$ for all $i$, and $X_{i2} = 1$ for a subset of units (the treated group) while the other units (the control group) have $X_{i2} = 0$. Suppose that $$Y_{it} = B_{it}X_{it}+A_i+F_t+U_{it},\quad i=1,...,N,\quad t=1,2,$$ where note that the treatment effect $B_{it}$ is heterogeneous across units and over time. Lastly, suppose that the treatment is sequentially exogenous (but may not be strictly exogenous), so that $$\mathbb{E}[U_{i2}\,|\, X_{i1}=0,X_{i2}=1]=\mathbb{E}[U_{i2}\,|\, X_{i1}=0,X_{i2}=0].$$

In this model, we can write

align*[align* omitted — 539 chars of source]

where the bias term is only zero under strict exogeneity. Indeed, in this setting, SE coincides with the assumption of parallel trends (PT), since, denoting as $Y_{it}(0)$ potential outcomes in the absence of treatment (assuming that the treatment is not anticipated), we have

align[align omitted — 304 chars of source]

Under SE/PT, the bias term is equal to zero, and the difference-in-differences (DID) estimand is equal to the average treatment effect on the treated (ATT). This recovers the classic result in abadie2005semiparametric. However, when SE/PT fails the bias term is generally non-zero, and the difference-in-differences estimand differs from the ATT in general.\footnote{Exceptions include the case where $U_{it}$ follows a random walk with an innovation that is independent of the treatment, as discussed in ghanem2022selection.}

Pre-trends in a three-period model with a binary treatment

Next, maintaining the setup from the previous subsection, suppose there is an additional initial period “$0$”, where no-one is treated. It is common in applications to verify that pre-trends are parallel in support of the SE/PT assumption, as discussed in Subsection (ref). In the present model, pre-trends are parallel if

align[align omitted — 250 chars of source]

Does the absence of pre-trends in ((ref)) imply that the SE/PT assumption holds, i.e., that the bias is zero in ((ref))?

To illustrate the difference between parallel pre-trends and parallel trends/SE, suppose $X_{i0}=X_{i1}$ is exogenously set to zero in periods $0$ and $1$, and treatment in period $2$ satisfies $$X_{i2}=\boldsymbol{1}\left\{Z_{i}\geq 0\right\},$$ where selection into treatment depends on an index $Z_{i}$ that may include past shocks to outcomes $U_{i0}$ and $U_{i1}$, and some other unrelated factors, but not the current shocks $U_{i2}$ (consistently with the assumption of sequential exogeneity). In addition, suppose that the conditional means $\mathbb{E}[U_{i1}\,|\, Z_{i}]={\Greekmath 0115}_1(Z_{i}-\mathbb{E}[Z_i])$ and $\mathbb{E}[U_{i0}\,|\, Z_{i}]={\Greekmath 0115}_0(Z_{i}-\mathbb{E}[Z_i])$ are linear.

In this case, pre-trends are parallel if and only if ${\Greekmath 0115}_1={\Greekmath 0115}_0$ (by ((ref))), while the bias due to failure of SE/PT is (by ((ref)))

align*[align* omitted — 510 chars of source]

that is,

equation[equation omitted — 195 chars of source]

There are many instances where pre-trends are parallel, yet SE/PT does not hold and the bias in ((ref)) is substantial. A simple example is when $Z_i=U_{i0}+U_{i1}+{\Greekmath 0118}_i$, where $U_{i0}$ and $U_{i1}$ have the same variances and are independent of ${\Greekmath 0118}_i$ (and conditional means are linear). In fact, it is clear in this example that there is no reason at all for ${\Greekmath 0115}_1={\Greekmath 0115}_0$ to imply that the bias in ((ref)) is zero.

The discussion in this section underscores that failure of strict exogeneity causes bias and inconsistency even in the canonical two-period difference-in-differences design. ghanem2022selection and marx2024parallel provide in-depth analyses of selection issues in DID settings. Moreover, the simple three-period model considered here highlights that verifying that pre-trends are parallel does not suffice to support the assumption of strict exogeneity/parallel trends.

The remainder of this paper presents methods that, unlike two-way fixed effects estimators and related methods, are robust to failure of strict exogeneity and the presence of feedback. These methods rely on using lagged covariates as instrumental variables. Such methods are ineffective in the simple two- and three-period models with an absorbing treatment considered in this section since $X_{i0}$ and $X_{i1}$ do not vary at all. However, in less stylized settings where $X_{it}$ exhibits more variation, these methods provide ways to meaningfully relax strict exogeneity by allowing for feedback.

Solutions and pitfalls in linear panel data models with constant coefficients

Sequential moment restrictions

A classic approach to estimation in linear models with sequentially exogenous covariates is based on sequential moment restrictions. arellano2003panel provides a comprehensive review of many available approaches. Consider the panel data model ((ref)) under the assumption that covariates $X_{it}$ are sequentially exogenous as in ((ref)). In applications, $X_{it}$ may include lagged outcomes in addition to other sequentially exogenous covariates.\footnote{Including outcome lags is often crucial in applications, as recently highlighted in klosin2024dynamic.}

Differencing between periods $t$ and $t-1$ yields

equation[equation omitted — 120 chars of source]

where in the rest of the paper the notation $Z^t$ will indicate the history of $Z_t$ up to period $t$. Equation ((ref)) represents a set of sequential conditional moment restrictions, for all periods $t$. As a result, any function $g(X_i^{t-1})$ of the history of the covariate can be used as an instrument in first differences, giving the unconditional moment restrictions

equation[equation omitted — 137 chars of source]

Different choices of instrument functions are used in practice. One example is the use of first differences as instruments, such as $g(X_{i}^{t-1})=X_{i,t-1}-X_{i,t-2}$. Another example is the use of instruments in levels, such as $g(X_{i}^{t-1})=X_{i,t-1}$. The popular GMM estimator from arellano1991some (ABond hereafter) is based on the moment restrictions ((ref)), see also holtz1988estimating. The optimally-weighted GMM estimator achieves the semi-parametric efficiency bound of the model under restrictions ((ref)) (see chamberlain1992comment).

Issues with common dynamic panel data estimators

Despite their popularity in applied work, estimators based on first differences and instrumental variables suffer from several shortcomings. Two main issues have been extensively discussed in the literature, see Chapter 11 in wooldridge2010econometric for example.

The first issue is that instruments may be weak. For example, with $T=2$ and a scalar covariate we have $$Y_{i2}-Y_{i1}={\Greekmath 010C}(X_{i2}-X_{i1})+U_{i2}-U_{i1},$$ and the instrument $X_{i1}$ is weak when $X_{i2}-X_{i1}$ and $X_{i1}$ are weakly correlated. As another example, consider a first-order autoregressive model

equation[equation omitted — 99 chars of source]

where $\mathbb{E}[U_{it}\,|\, Y_{i}^{t-1}]=0$. The relevance of $Y_{i,t-1}$ as an instrument hinges on there being a non-zero correlation between $Y_{i,t-1}-Y_{i,t-2}$ and $Y_{i,t-2}$. One thus expects the instrument to be weak when ${\Greekmath 011A}$ is close to 1.

A second issue with panel data estimators based on IV and first differences is that, when $T$ is not small relative to $N$, the number of instruments is large relative to the sample size. Indeed, the fact that any lagged $X_{it}$ beyond the first order is a valid instrument because of ((ref)) implies that the number of instruments increases with $T$. Specifically, linear moment restrictions imply $T(T-1)/2$ valid moments (for a single scalar covariate). This causes an issue of proliferation of instruments. In simulations, researchers have found that method-of-moment estimators exhibit bias in such cases (e.g., kiviet1995bias).

Considering the autoregressive model ((ref)) in an asymptotic regime where both $N$ and $T$ increase, alvarez2003time show that the ABond estimator is biased when $N/T$ tends to a constant. Specifically, they show that, while the OLS estimator with fixed effects suffers from a bias of order 1/T, the ABond estimator suffers from a bias of order 1/N. When $N$ and $T$ have comparable magnitudes, the bias is not negligible relative to the standard deviation of the estimator (which is proportional to $1/\sqrt{NT}$). This distorts confidence intervals and affects the validity of inference.

The analysis in alvarez2003time implies that, when $N$ and $T$ have comparable magnitudes, IV estimators may suffer from biases, similarly to OLS with fixed effects. This issue is due to the fact that the estimator is based on a set of moment restrictions that grows with $T$. If instead one uses a single lag $Y_{i,t-2}$ in model ((ref)), for example, then the GMM estimator does not suffer from an asymptotic bias in the asymptotic regime analyzed in alvarez2003time.

Additional moment restrictions

A first reaction to the weak instruments issue highlighted in the previous subsection has been to impose additional moment restrictions. One popular assumption is that the association between the instruments and the unobserved effects $A_i$ is constant. For example, when using lags of $X_{it}$ as instruments this requires assuming that

equation[equation omitted — 153 chars of source]

If the stationarity condition ((ref)) holds, then $$\mathbb{E}[(X_{it}-X_{i,t-1})(A_i+U_{it})]=\underset{=0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ by }(\ref{eq_stat})}{\underbrace{\mathbb{E}[X_{it}A_i]-\mathbb{E}[X_{i,t-1}A_i]}}+\mathbb{E}[(X_{it}-X_{i,t-1})\underset{=0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ by } (\ref{eq_seq})}{\underbrace{\mathbb{E}[U_{it}\,|\, X_i^t]}}] =0,$$ so $X_{it}-X_{i,t-1}$ is a valid instrument for the equations in levels, leading to the moment conditions

equation[equation omitted — 100 chars of source]

arellano1995another introduced such stationarity restrictions and analyzed efficiency properties in a setting that combines restrictions of the form ((ref)) with the ABond restrictions ((ref)), corresponding to the “system GMM” estimator. blundell1998initial adapted and applied this strategy for production function estimation.

Another type of moment restrictions was introduced by ahn1995efficient. Suppose that the sequential exogeneity assumption ((ref)) is replaced by

equation[equation omitted — 75 chars of source]

This requires adding lagged outcomes as regressors to the model, which is often not restrictive since dynamic models typically include lagged outcomes. In addition, ((ref)) requires mean independence to hold conditional on the unobserved effect $A_i$. While this conditional sequential exogeneity assumption is formally more restrictive that its unconditional version $\mathbb{E}[U_{it}\,|\, X_i^t,Y_{i}^{t-1}]=0$, it still allows for unrestricted forms of feedback. Now, if ((ref)) holds, then $$\mathbb{E}[(A_i+U_{it})(U_{i,t-1}-U_{i,t-2})]=0,$$ which implies quadratic moment restrictions on the parameters. Such restrictions may be combined with the ones derived by arellano1991some and arellano1995another.

However, while one may hope that adding informative moment restrictions may alleviate weak instruments issues, by construction doing so increases the number of moment restrictions used in estimation and may thus be subject to the “many instruments” problem. The methods reviewed next are not based on additional moment restrictions and are thus potentially more robust to this last issue.

Quasi-likelihood approaches

Another reaction to the issues with the ABond and other GMM estimators based on sequential moment restrictions has been to develop alternative estimation strategies. A strategy proposed by several authors is to adopt a likelihood approach based on a Gaussian specification. The motivation for this approach is that the likelihood implicitly “weighs” the sequential moments implied by the model in a way that would be optimal under the Gaussian parametric assumptions, yet may perform better than GMM even when the parametric assumptions are violated (hence the name “quasi” likelihood).

hsiao2002maximum introduce a transformed maximum likelihood approach for equations in first differences. alvarez2003time consider and analyze a random-effects estimation approach in levels that additionally requires a normal specification for the distribution of the individual effect $A_i$ given initial conditions, see also alvarez2004robust. This correlated random-effects specification is reminiscent of the mundlak1978pooling approach for static panel data models. moral2013likelihood and moral2019dynamic develop likelihood approaches for general models with sequentially exogenous regressors. bun2017maximum study the finite-sample performance of some of these estimators in simulations.

Consider the autoregressive model in alvarez2004robust,

equation[equation omitted — 111 chars of source]

Suppose that $A_i$, given the initial condition $Y_{i0}$, is drawn from a Gaussian distribution with a mean that is linear in $Y_{i0}$ and a constant variance, and suppose that $U_{it}$ are Gaussian i.i.d. over time with constant variance. Maximizing the log-likelihood yields the maximum likelihood estimator $\widehat{{\Greekmath 011A}}$, which is consistent and asymptotically normal as $N$ tend to infinity under Gaussianity. Moreover, $\widehat{{\Greekmath 011A}}$ remains consistent and asymptotically normal even if the true data generating process does not satisfy Gaussian assumptions, albeit with a different asymptotic variance. Moreover, alvarez2003time show that, as $N$ and $T$ tend to infinity at the same rate, $$\sqrt{NT}(\widehat{{\Greekmath 011A}}-{\Greekmath 011A})\overset{d}{\rightarrow}{\cal{N}}(0,1-{\Greekmath 011A}^2),$$ showing that, unlike the ABond estimator, the quasi-maximum likelihood estimator $\widehat{{\Greekmath 011A}}$ does not suffer from asymptotic bias.

In more general models with sequentially exogenous covariates, moral2013likelihood proposes a likelihood approach that relies on a specification of the feedback process, assuming that $$X_{it}\,|\, Y_{i}^{t-1},X_{i}^{t-1},A_i$$ is Gaussian with a mean that depends linearly on $Y_{i}^{t-1},X_{i}^{t-1},A_i$ and a constant variance matrix. As in the case of the likelihood specification for the autoregressive model, the resulting quasi-maximum likelihood estimator remains consistent and asymptotically normal even when the Gaussian assumptions are violated.

Large-T perspective and bias correction

The large-T perspective introduced in hahn2002asymptotically can be used to develop novel estimators with improved properties in panels of moderate size (i.e., where the time dimension is not too short). This class of estimators is based on an asymptotic approach where $N$ and $T$ tend to infinity jointly. The key idea is, starting from an estimator such as OLS, to then remove the bias (or, in practice, an estimate of the bias) from the estimator.

To illustrate the idea, consider the simple case of the OLS estimator with fixed effects $\widehat{{\Greekmath 011A}}$ based on the autoregressive model ((ref)). As $N,T$ tend to infinity at the same rate, hahn2002asymptotically show that

equation[equation omitted — 197 chars of source]

This suggests that the bias-corrected estimator $$\widehat{{\Greekmath 011A}}^{\rm BC}=\widehat{{\Greekmath 011A}}+\frac{1+\widehat{{\Greekmath 011A}}}{T}$$ satisfies

equation[equation omitted — 186 chars of source]

In other words, the bias-corrected estimator $\widehat{{\Greekmath 011A}}^{\rm BC}$ is free from (asymptotic) bias, and the asymptotic distribution of $\sqrt{NT}\left(\widehat{{\Greekmath 011A}}^{\rm BC}-{\Greekmath 011A}\right)$ is correctly centered at zero, enabling the usual construction of asymptotically valid confidence intervals.

More generally, in models with general sequentially exogenous covariates, and starting with an estimator $\widehat{{\Greekmath 010C}}$ satisfying an expansion of the form ((ref)), the bias-correction approach is based on a consistent estimator $\widehat{C}$ of the constant $C$, and on the resulting bias-corrected estimator $$\widehat{{\Greekmath 010C}}^{\rm BC}=\widehat{{\Greekmath 010C}}-\frac{\widehat{C}}{T}.$$ Typically, the asymptotic distribution of $\sqrt{NT}(\widehat{{\Greekmath 010C}}^{\rm BC}-{\Greekmath 010C})$ is correctly centered, similar to ((ref)). See arellano2007understanding for a survey of various approaches, and dhaene2015split for an approach based on split-panel jackknife.

The price to pay for relying on large-T asymptotic arguments is that bias-corrected estimators such as $\widehat{{\Greekmath 011A}}^{\rm BC}$ and $\widehat{{\Greekmath 010C}}^{\rm BC}$ are generally not consistent in short panels, in the sense that they do not converge to the population parameter as $N$ tends to infinity while $T$ is fixed, only as $N$ and $T$ tend to infinity jointly. In some settings, it is possible to completely remove the bias and achieve fixed-$T$ consistency, as bun2005bias do in an autoregressive model, however this situation appears to be the exception rather than the rule.

Misspecification issues in dynamic panel models

It is important to note that, except for methods based on large-T arguments, all the methods reviewed in this section rely on conditional moment restrictions such as ((ref)) being satisfied. That is, the credibility of the methods hinges on the validity of instrumental variables that, unlike in the quasi-experimental literature, are internal to the model. As a result, a threat to validity is that the model be misspecified in some dimensions. One form of misspecification, which was already highlighted in arellano1991some, is related to the dependence structure of errors when lagged outcomes are present in the equation. Indeed, considering the autoregressive model ((ref)) and assuming that $\mathbb{E}[U_{it}\,|\, Y_i^{t-1}]=0$ requires $U_{it}$ to be independent over time. arellano1991some propose a test of this hypothesis, against a more general alternative that $U_{it}$ follow a moving-average process of a given order.

Empirically, there are other reasons why the model, and the implied moment restrictions, may be misspecified. One reason is that, unlike what model ((ref)) postulates, the coefficients ${\Greekmath 010C}$ may be heterogeneous across individuals and possibly over time as well (marx2024heterogeneous). This concern with treatment effects heterogeneity has taken a central place in modern applied econometrics, as reviewed in the introduction, and it is the topic of the next section. Another possible source of misspecification is that, in contrast with the linear model ((ref)), the true relationship between outcomes $Y_{it}$ and covariates $X_{it}$ may be nonlinear. Such a concern is prevalent in settings with continuous treatments or discrete outcomes, and it is the topic of Section (ref).

Heterogeneity and feedback

Chamberlain's negative result

Allowing for coefficient heterogeneity is of paramount importance in many applications. In policy evaluation settings, how to handle the presence of unobserved treatment effects heterogeneity is often a central part of the analysis. In models with strictly exogenous covariates, many methods have been developed to estimate features of the distribution of heterogeneous coefficients, see, e.g., chamberlain1992efficiency, arellano2012identifying, graham2012identification, and the survey by bonhomme2025fixed. Those methods have also been applied to difference-in-differences settings, see borusyak2024revisiting, botosaru2025time, de2025treatment, and the survey by arkhangelsky2024causal.

However, identification raises important new challenges when covariates are sequentially exogenous and treatment effects are heterogeneous. An important contribution by Gary Chamberlain (chamberlain2022feedback) establishes a negative identification result in a model with a binary sequentially exogenous covariate where both the coefficient of the covariate and the intercept vary unrestrictedly across individuals.

To understand the identification challenge, consider the following random coefficients model

equation[equation omitted — 104 chars of source]

where it is assumed that the scalar covariate $X_{it}$ is sequentially exogenous in the sense of

equation[equation omitted — 96 chars of source]

and observations are i.i.d. over $i$. Suppose one wishes to characterize the identified set for the average treatment effect parameter ${\Greekmath 0116}=\mathbb{E}[B_i]$. Establishing identification requires showing that the set is in fact a singleton.

To provide intuition, consider first model ((ref)) without time effects (i.e., $F_t=0$) and $T=2$, and write first differences as

equation[equation omitted — 64 chars of source]

Clearly, the mean of $B_i$ may only be identified in subpopulations such that $X_{i2}-X_{i1}$ is non-zero. In other words, the average treatment effect (ATE) is never identified, and as is common in the literature we will aim to identify some form of average treatment effect “on the treated” (ATT). In the present discussion I will assume that the support of $X_{i2}-X_{i1}$ is bounded away from zero.\footnote{graham2012identification study the case where $X_{i2}-X_{i1}$ is close to zero and identification may be irregular.}

We then have

equation[equation omitted — 133 chars of source]

where $\widehat{B}_i$ is the OLS estimator of $B_i$ in model ((ref)) for the single individual $i$. Since there are only two periods, we expect $\widehat{B}_i$ to be very noisy. Nevertheless, under strict exogeneity, the average OLS (or “mean group”) estimator is unbiased for ${\Greekmath 0116}$, since $$\mathbb{E}\left[\frac{1}{N}\sum_{i=1}^N \widehat{B}_i\right]=\mathbb{E}[B_i]+\frac{1}{N}\sum_{i=1}^N\mathbb{E}\left[\frac{\overset{=0}{\overbrace{\mathbb{E}[U_{i2}-U_{i1}\,|\, X_{i1},X_{i2}]}}}{X_{i2}-X_{i1}}\right]={\Greekmath 0116}.$$ However, this argument breaks down when $X_{it}$ are not strictly exogenous and $$\mathbb{E}[U_{i2}-U_{i1}\,|\, X_{i1},X_{i2}]\neq 0.$$

chamberlain2022feedback considers model ((ref)) with a binary treatment in the presence of time effects, and he takes $T=2$. That is, imposing the normalization $F_1=0$ and denoting $F_2=F$,

align[align omitted — 197 chars of source]

Obviously, in this model there is no way to identify $\mathbb{E}[B_i\,|\, X_{i1}=1,X_{i2}=1]$ and $\mathbb{E}[B_i\,|\, X_{i1}=0,X_{i2}=0]$ (“average treatment effects on stayers”) due to the presence of the unrestricted intercept $A_i$. However, remarkably, chamberlain2022feedback shows that $\mathbb{E}[B_i\,|\, X_{i1}=0,X_{i2}=1]$ and $\mathbb{E}[B_i\,|\, X_{i1}=1,X_{i2}=0]$ (“average treatment effects on movers”) are not identified either, and nor is the time effect $F$. This situation stands in sharp contrast with settings under strict exogeneity, where these last three quantities are identified, although average effects on stayers are not (chamberlain1992efficiency).

Point-identified averages

There have been two responses to Chamberlain's negative result. The first one has been to characterize and study quantities that remain identified under sequential heterogeneity. This is the topic of this subsection and the next. The second response, which I will turn to in Subsection (ref), has been to adopt a partial identification approach and bound the quantities that are not point-identified.

Consider model ((ref))-((ref)), with the normalization $F_T=0$, and focus for concreteness on the case where $X_{it}$ is binary. As in bonhomme2025unrestricted, we ask what weighted averages of $B_i$, $${\Greekmath 0116}=\mathbb{E}\left[c(X_{i1},...,X_{iT})B_i\right],$$ are identified, where $c(X_{i1},...,X_{iT})$ are possibly positive or negative (scalar) weights.

A sufficient condition for ${\Greekmath 0116}$ to be identified is that there exist functions ${\Greekmath 0127}_t:\{0,1\}^t\mapsto \mathbb{R}$, for $t=1,...,T-1$, such that

align[align omitted — 194 chars of source]

Here, ${\Greekmath 0127}_t(X_{i}^t)$ are functions of the history of the covariate, which are effectively used as instrumental variables (as is often the case in dynamic panel data settings).

To see why ((ref)) and ((ref)) imply that ${\Greekmath 0116}$ is identified, note that

align*[align* omitted — 910 chars of source]

where the next-to-last line uses that, by ((ref)) and the law of iterated expectations, $$\mathbb{E}\left[{\Greekmath 0127}_t(X_{i}^t)(U_{it}-U_{iT})\right]=\mathbb{E}\left[{\Greekmath 0127}_t(X_{i}^t)\mathbb{E}\left(U_{it}-U_{iT}\,|\, X_i^t\right)\right]=0,$$ and the expression in the last line is a population mean of observed variables, hence identified. Note that ((ref)) is only needed due to the presence of the time effects $F_t$.

In fact, it can be shown using the strategy in bonhomme2025unrestricted that, if no functions ${\Greekmath 0127}_t$ satisfy ((ref))-((ref)), then ${\Greekmath 0116}$ is not identified. Moreover, in that case the identified set for ${\Greekmath 0116}$ is the whole real line. Hence, the class of weights $c$ leading to point-identified weighted averages is fully characterized as the solution to a linear system of equations whose parameters are the values of the functions ${\Greekmath 0127}_t$. This allows for simple and exhaustive characterizations in cases that have been previously considered in the literature.

As a first example, consider the model without time effects and $T=3$, as studied in arellano2012identifying. In this model, ((ref)) is not needed for identification (due to the absence of time effects), and ((ref)) reads

align[align omitted — 132 chars of source]

Hence, enumerating the $2^3=8$ support points of $(X_{i1},X_{i2},X_{i3})$, the identification condition can equivalently be written as

align*[align* omitted — 366 chars of source]

which defines a six-dimensional linear space. arellano2012identifying highlight that $X_{i1}(X_{i2}-X_{i1})$, $X_{i1}(X_{i3}-X_{i1})$, and $X_{i2}(X_{i3}-X_{i2})$ all lead to point identification. In fact, as the above derivation shows, additional weighted averages are identified.

As a second example, consider model ((ref))-((ref)) analyzed by chamberlain2022feedback, with time effects and $T=2$. In this case ((ref)) and ((ref)) read, denoting ${\Greekmath 0119}_1=\mathbb{E}[X_{i1}]$,

align*[align* omitted — 229 chars of source]

Hence, if ${\Greekmath 0119}_1\neq 0$, ${\Greekmath 0116}$ is point-identified if and only if $c(X_{i1},X_{i2}) $ is proportional to

equation[equation omitted — 125 chars of source]

This defines a one-dimensional linear space that does not include either the average effects on movers nor the ones on stayers, consistently with the analysis in chamberlain2022feedback.

Best identified approximation

When point-identification of ${\Greekmath 0116}$ fails, its identified set is unbounded. In this case, following azriel2020estimation, bonhomme2025unrestricted proposes to focus on the weighted average ${\Greekmath 0116}^*$ that is closest to ${\Greekmath 0116}$ while being point-identified. The quantity ${\Greekmath 0116}^*$ is the best identified approximation to ${\Greekmath 0116}$.

To illustrate this approach, consider again model ((ref))-((ref)), and suppose that the researcher is interested in the average treatment effect ${\Greekmath 0116}=\mathbb{E}[B_i]$, which corresponds to $c$ being a vector of ones. Although ${\Greekmath 0116}$ is not point-identified is that setting, one can consider $${\Greekmath 0116}^*=\mathbb{E}\left[{\Greekmath 0115}^* {c}_0(X_{i1},X_{i2})B_i\right],$$ where $c_0$ is given by ((ref)), and $${\Greekmath 0115}^*=\underset{{\Greekmath 0115}}{\limfunc{argmin}}\, \mathbb{E}\left[\left(1-{\Greekmath 0115}{c}_0(X_{i1},X_{i2})\right)^2\right].$$ Note the presence of the constant weights $c=1$ in this formula, which correspond to the average effect of interest $\mathbb{E}[B_i]=\mathbb{E}[1\times B_i]$.

I plot the weights $c$ and $c^*$ in Figure (ref), for the case where $X_{i1}$ and $X_{i2}$ follow independent Bernoulli distributions: $X_{i1}\sim \mbox{Ber}({\Greekmath 0119}_1)$, where I vary ${\Greekmath 0119}_1$ on the x-axis, and $X_{i2}\sim \mbox{Ber}(1/2)$. The original weights $c$ are identically equal to one, and are shown in black in the Figure. The weights $c^*$ corresponding to the best identified approximation are shown in blue (for $c^*(0,1)$) and red (for $c^*(1,0)$). The best identified weights for stayers (not shown in the Figure) are $c^*(0,0)=c^*(1,1)=0$.

figure[figure omitted — 493 chars of source]

Several conclusions can be drawn from Figure (ref). First, $c^*$ and $c$ are different, which reflects the lack of point-identification of ${\Greekmath 0116}=\mathbb{E}[B_i]$. Second, despite these differences, the weights $c^*$ corresponding to the best identified approximation are all non-negative, consistent with a “weakly causal” interpretation of ${\Greekmath 0116}^*$. Third, as the probability ${\Greekmath 0119}_1=\Pr(X_{i1}=1)$ increases, ${\Greekmath 0116}^*$ puts more weight on the combination $(X_{i1}=0,X_{i2}=1)$ relative to $(X_{i1}=1,X_{i2}=0)$. Lastly, when ${\Greekmath 0119}_1=\frac{1}{2}$ we have $c^*(0,1)=c^*(1,0)=1$, which identifies the following average treatment effect on movers: $$\mathbb{E}[B_i\,|\, (X_{i1}=0,X_{i2}=1) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ OR }(X_{i1}=1,X_{i2}=0)].$$

Characterizing further the properties of best identified approximations in models with sequential exogeneity is a promising avenue for future work.

Partial identification: the (Wooyong) Lee bounds

To see how to obtain finite bounds on ${\Greekmath 0116}$ even when it is not point-identified, consider again model ((ref)) -- without time effects for simplicity -- and assume that sequential exogeneity holds conditional on the latent individual effects,

equation[equation omitted — 96 chars of source]

Note that this conditional sequential exogeneity assumption differs from the unconditional one, ((ref)), which we assumed in the previous two subsections.

When ((ref)) holds, the characterization of the identified set is more complicated than in ((ref))-((ref)). The reason is that one needs to work with the entire distributions of the latent variables $(A_i,B_i)$, as their conditional expectations no longer exhaust all the information in the model. Nevertheless, one can show that point identification fails in many cases. For example, focusing on an autoregressive model with $X_{it}=Y_{i,t-1}$, lee2020identification shows that $\mathbb{E}[B_i]$ is not point-identified. Hence, imposing the stronger condition ((ref)) is not enough to restore point-identification.

Nevertheless, under ((ref)) one can bound the average coefficient $\mathbb{E}[B_i]$, as shown by lee2020identification. To see why, consider the simple case $T=2$, and note the model implies the following unconditional moment restrictions

equation[equation omitted — 98 chars of source]

It turns out that ((ref)) is enough to obtain finite bounds on ${\Greekmath 0116}=\mathbb{E}[B_i]$.

To develop the argument, observe that, for all scalar ${\Greekmath 0115}$, $$ {\Greekmath 0116}=\mathbb{E}\left[B_i+{\Greekmath 0115}\sum_t(A_i+B_iX_{it})(Y_{it}-B_iX_{it}-A_i)\right]. $$ Hence, denoting $$ Q_i(A_i,B_i,{\Greekmath 0115})=B_i+{\Greekmath 0115}\sum_t(A_i+B_iX_{it})(Y_{it}-B_iX_{it}-A_i), $$ we have $$ {\Greekmath 0116}=\mathbb{E}[Q_i(A_i,B_i,{\Greekmath 0115})]. $$

Next, note that $Q_i(A_i,B_i,{\Greekmath 0115})$ is a second-order polynomial in $A_i,B_i$. Taking ${\Greekmath 0115}>0$, we have $$ \underset{A_i}{\limfunc{max}}\, Q_i(A_i,B_i,{\Greekmath 0115}) = B_i+{\Greekmath 0115}\sum_t\!\left(\tfrac{1}{2}\overline{Y}_i+B_i(X_{it}-\overline{X}_i)\right) \!\left(Y_{it}-\tfrac{1}{2}\overline{Y}_i-B_i(X_{it}-\overline{X}_i)\right)=P_i(B_i,{\Greekmath 0115}), $$ where $P_i(B_i,{\Greekmath 0115})$ is a second-order polynomial in $B_i$. The coefficient of $B_i^2$ in $P_i(B_i,{\Greekmath 0115})$ is $-{\Greekmath 0115}\sum_t(X_{it}-\overline{X}_i)^2$. If $\sum_t(X_{it}-\overline{X}_i)^2>0$ then the polynomial is strictly concave and it has a finite maximum: $$ Z_i({\Greekmath 0115})=\underset{B_i}{\limfunc{max}}\, P_i(B_i,{\Greekmath 0115})=\underset{A_i,B_i}{\limfunc{max}}\, Q_i(A_i,B_i,{\Greekmath 0115}). $$ Hence, if $Z_i({\Greekmath 0115})$ has finite mean then

equation[equation omitted — 157 chars of source]

This provides a finite upper bound on ${\Greekmath 0116}=\mathbb{E}[B_i]$. The argument for the lower bound is similar, taking ${\Greekmath 0115}<0$.

This basic bound can be improved in many ways. One approach is to optimize with respect to ${\Greekmath 0115}$. Another approach is to consider additional moment restrictions implied by the model.\footnote{In fact, the model implies a continuum of those, since one can take arbitrary functions of $A_i$, $B_i$, and $X_i^t$ as instruments.} See lee2020identification for a detailed analysis of identification and inference on the identified set.

It is interesting to note that the existence of a finite bound on ${\Greekmath 0116}$ contrasts with the discussion following ((ref))-((ref)) that, under ((ref)), the identified set of ${\Greekmath 0116}$ is either a singleton or the whole real line. The reason is that the bound in ((ref)) is based on ((ref)), which imposes sequential exogeneity conditional on $A_i,B_i$. Suppose instead that sequential exogeneity holds unconditionally, as in ((ref)). In this case, the only moment restrictions implied from the model are of the form $$ \mathbb{E}[{\Greekmath 0120}(X_i^t)(Y_{it}-B_iX_{it}-A_i)]=0, $$ for arbitrary functions ${\Greekmath 0120}$. Then, defining analogously to before $$ Q_i(A_i,B_i,{\Greekmath 0115})=B_i+{\Greekmath 0115}{\Greekmath 0120}(X_i^t)(Y_{it}-B_iX_{it}-A_i), $$ one notes that $Q_i(A_i,B_i,{\Greekmath 0115})$ is a first-order polynomial in $A_i,B_i$ that has no finite maximum (except in trivial cases). Hence, indeed, ${\Greekmath 0116}$ is either point-identified or its identified set is the whole real line in this case.

Restoring point-identification?

The failure of point-identification in model ((ref)) arises since heterogeneity $(A_i,B_i)$ is bivariate. A possibility to restore point-identification is to assume that heterogeneity is scalar. Consider as an example the model

equation[equation omitted — 77 chars of source]

where one assumes that $$B_i={\Greekmath 0112}_1 A_i+{\Greekmath 0112}_0.$$ This implies $$Y_{it}={\Greekmath 0112}_0X_{it}+A_i(1+{\Greekmath 0112}_1X_{it})+U_{it}.$$

Hence we have (assuming that the denominator is non-zero) $$\frac{Y_{it}-{\Greekmath 0112}_0X_{it}}{1+{\Greekmath 0112}_1X_{it}}=A_i+V_{it},$$ where $V_{it}=\frac{U_{it}}{1+{\Greekmath 0112}_1X_{it}}$ is such that $\mathbb{E}[V_{it}\,|\, X_i^t]=0$. Taking first differences then implies the conditional moment restrictions $$\mathbb{E}\left[\frac{Y_{it}-{\Greekmath 0112}_0X_{it}}{1+{\Greekmath 0112}_1X_{it}}-\frac{Y_{i,t-1}-{\Greekmath 0112}_0X_{i,t-1}}{1+{\Greekmath 0112}_1X_{i,t-1}}\,\bigg|\, X_i^{t-1}\right]=0.$$ Instruments in levels and differences can be used to identify and estimate ${\Greekmath 0112}_0$ and ${\Greekmath 0112}_1$. Once those have been recovered, one obtains $$\mathbb{E}[A_i]=\mathbb{E}\left[\frac{Y_{it}-{\Greekmath 0112}_0X_{it}}{1+{\Greekmath 0112}_1X_{it}}\right],\quad \mathbb{E}[B_i]={\Greekmath 0112}_1 \mathbb{E}[A_i]+{\Greekmath 0112}_0.$$ chamberlain2022feedback characterizes efficient estimators in this setup. However, assuming that $(A_i,B_i)$ is uni-dimensional is a strong assumption, which restricts the nature of heterogeneity substantially.

Another possibility to guarantee point-identification in models with coefficient heterogeneity and sequential heterogeneity is to follow a large-$T$ approach, as outlined in Subsection (ref). In this approach, estimates based on models with coefficient heterogeneity and sequentially exogenous covariates typically satisfy bias expansions of the form ((ref)). Given this, bias correction approaches can be applied to such settings as well, see fernandez2013panel for details. A drawback of large-T approaches is that they do not achieve fixed-T consistency in general, and thus may not perform well in short panels.

Nonlinearity and feedback

Economic models often imply nonlinear relationships between covariates and outcomes. Decreasing returns, or curvature of preferences, generate nonlinearities that have implications for policy predictions. However, a challenge to extend the analysis in the previous sections to nonlinear settings comes from the presence of latent heterogeneity (i.e., “fixed effects”). Unlike in linear models, it is typically not possible to difference out the heterogeneity except in special cases (arellano2001panel). Under strict exogeneity, solutions exist in general semi-parametric likelihood models (bonhomme2012functional). However, the presence of un-modeled dynamics and feedback creates additional challenges.

After mentioning some model-specific approaches in certain nonlinear settings, this section focuses on semi-parametric likelihood models with feedback, reviewing results on the existence of valid moment restrictions (which are useful when the parameters are point-identified) and on the characterization of identified sets more generally. The last part of the section incorporates restrictions on the feedback process and shows how this affects the analysis.

Moment conditions in some specific models

Moment restrictions allowing for feedback have been derived in certain specific nonlinear models with a multiplicative structure. Blundell_Griffith_Windmeijer_JOE2002 consider count data regression models, see also wooldridge1997multiplicative. al2017exponential derive moment conditions in a binary choice model with an exponential structure. The efficiency analysis in chamberlain2022feedback covers those cases.

To see how to handle the presence of multiplicative heterogeneity and feedback, consider a panel data model with a multiplicative structure, $$Y_{it}=\exp({\Greekmath 010C} X_{it}+A_i)U_{it},$$ where $$\mathbb{E}[U_{it}\,|\, X_i^t,A_i]=1.$$ Note that here we require a moment restriction that is conditional on the individual heterogeneity $A_i$. Left-multiplying by $\exp(-{\Greekmath 010C} X_{it})$ gives $$\mathbb{E}[\exp(-{\Greekmath 010C} X_{it})Y_{it}\,|\, X_i^t,A_i]=\exp(A_i).$$ Hence, taking differences between periods $t$ and $t-1$, we obtain $$\mathbb{E}[\exp(-{\Greekmath 010C} X_{it})Y_{it}-\exp(-{\Greekmath 010C} X_{i,t-1})Y_{i,t-1}\,|\, X_i^{t-1},A_i]=\exp(A_i)-\exp(A_i)=0.$$

Generalizing this approach to models without a multiplicative structure has proven difficult. In unpublished dissertation work, woutersen2000essays derive moment conditions in a multiple-spell mixed proportional hazards model of duration. Recently, bonhomme2023identification study identification in binary choice models with feedback. They find that lack of point-identification is pervasive in these models. For example, studying a logit binary choice model $$Y_{it}=\boldsymbol{1}\{{\Greekmath 010C} X_{it}+A_i+U_{it}\geq 0\},$$ where $X_{it}$ are binary and $$U_{it}\,|\, X_i^t,Y_{i}^{t-1},A_i\sim\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Logistic},$$ they find that ${\Greekmath 010C}$ is never point-identified, irrespective of how large $T$ is. This finding contrasts with the case with a strictly exogenous $X_{it}$, where point-identification can be achieved when $T=2$ (georg1960probabilistic), and with the autoregressive case $X_{it}=Y_{i,t-1}$, where point-identification can be achieved when $T=3$ (chamberlain2023identification).

An approach for likelihood models

bonhomme2025feedback introduce an approach to find moment conditions in nonlinear models with feedback that have a semi-parametric likelihood structure, such that the outcome follows a parametric distribution indexed by a parameter vector ${\Greekmath 0112}$, $$Y_{it}\,|\, Y_{i,t-1},X_{it},A_i\sim f_{{\Greekmath 0112}},$$ and the feedback process and the distribution of heterogeneity given initial conditions are both unrestricted. This setup mimics, in a nonlinear setting, linear dynamic models with sequential exogeneity where the model restricts the mean (and possibly the variance) of outcomes given covariates and heterogeneity while leaving feedback and heterogeneity unrestricted.

bonhomme2025feedback provide necessary and sufficient conditions for a moment function ${\Greekmath 011E}_{{\Greekmath 0112}}$ to be feedback and heterogeneity robust (FHR), in the sense that

equation[equation omitted — 95 chars of source]

irrespective of the form of the feedback process and the distribution of heterogeneity. Their characterization, which takes the form of integral equations, extends the characterization under strict exogeneity in bonhomme2012functional to models with feedback.

To provide intuition about the conditions ensuring the FHR property, consider first the case $T=2$ under the assumption that covariates are strictly exogenous. bonhomme2012functional points out that ${\Greekmath 011E}_{{\Greekmath 0112}}$ satisfies ((ref)) if and only if

equation[equation omitted — 149 chars of source]

That ((ref)) implies ((ref)) is obvious by the law of iterated expectations, and the converse holds due to the fact that the distribution of $A_i$ given covariates is unrestricted. Now, note that ((ref)) solely depends on ${\Greekmath 0112}$. Indeed, ((ref)) is equivalent to

equation[equation omitted — 5,269 chars of source]

Hence, finding moment functions ${\Greekmath 011E}_{{\Greekmath 0112}}$ amounts to solving the integral equation ((ref)). bonhomme2012functional, and subsequently honore2024moment, honore2025dynamic, and dano2023transition, among others, use this approach to find moment restrictions in a variety of models with strictly exogenous covariates.

bonhomme2025feedback observe that, in models with sequentially exogenous covariates, ((ref)) is no longer sufficient for ${\Greekmath 011E}_{{\Greekmath 0112}}$ to provide a valid moment restriction in general. This is because, in the presence of feedback, $X_{i2}$ depends on $Y_{i1}$ and ((ref)) is no longer equivalent to ((ref)). The authors show that, for ${\Greekmath 011E}_{{\Greekmath 0112}}$ to be a valid moment function, it needs to satisfy another condition, which in integral form reads

equation[equation omitted — 244 chars of source]

They show that the FHR moment functions are those that satisfy both ((ref)) and ((ref)). The second condition ((ref)) can be interpreted as ensuring robustness to the presence of an unknown feedback process. It can be equivalently stated as

equation*[equation* omitted — 254 chars of source]

which requires that $X_{i2}$ does not predict ${\Greekmath 011E}_{{\Greekmath 0112}}$, conditional on its other arguments and the individual effect $A_i$.

bonhomme2025feedback show that FHR moment functions span the ortho-complement of the nuisance tangent set of the model, and are thus helpful to construct efficient moment functions using sequential projection arguments, as in the literature on linear dynamic models initiated by chamberlain1992comment and arellano1995another. In addition, they provide an analogous characterization of FHR moment functions for average effects of the form

equation[equation omitted — 99 chars of source]

for a known function $h$, hence allowing one to obtain estimators of average partial effects and other average effects of interest in economic applications.

Identified sets in nonlinear models

In certain models, feedback and heterogeneity robust (FHR) moment functions as in ((ref)) may fail to exist. In such cases, a natural approach is to resort to partial identification and bound the quantity of interest. Fortunately, identified sets in likelihood models with feedback have a tractable representation. To see this, let $f_0(y^T,x^T)$ denote the joint density of $Y_t,X_t$ across periods. Focusing again on the case $T=2$ for simplicity, let $$g(x_2\,|\, y_1,y_0,x_1,a)$$ denote the feedback process (i.e., the conditional density of $X_{i2}$), let $${\Greekmath 0119}(a\,|\,y_0,x_1)$$ denote the density of heterogeneity $A_i$, and let $${\Greekmath 0117}(y_0,x_1)$$ denote the density of initial conditions $Y_{i0},X_{i1}$. Consider as before a setting where $g$, ${\Greekmath 0119}$, and ${\Greekmath 0117}$ are all left unrestricted. The joint likelihood of $(Y_{i2},Y_{i1},Y_{i0},X_{i2},X_{i1},A_i)$ is

align*[align* omitted — 181 chars of source]

and the identified set for ${\Greekmath 0112}$ is the set of ${\Greekmath 0112}$ values such that, for some $g,{\Greekmath 0119},{\Greekmath 0117}$,

equation[equation omitted — 229 chars of source]

In words,the identified set is the set of ${\Greekmath 0112}$ values that are consistent with the data density for some values of the feedback process, the heterogeneity density, and the density of initial conditions.

bonhomme2023identification observe that ((ref)) is equivalent to the existence of a density $p(y_0,y_1,y_2,x_1,x_2,a)$ such that the following conditions hold:

align[align omitted — 10,489 chars of source]

where ((ref)) expresses that the model and data density are consistent with each other, while ((ref)) and ((ref)) express that the model admits $f_{{\Greekmath 0112}}$ as the (parametric) conditional densities of outcomes in both periods.

Since ((ref)), ((ref)) and ((ref)) are linear in the density $p$, one can check whether any given value ${\Greekmath 0112}$ belongs to the identified set by verifying whether a linear program has a solution. This feature was first noticed by honore2006bounds in likelihood models with strictly exogenous covariates, and it is here extended to models with feedback. Likewise, bonhomme2023identification show that, for ${\Greekmath 0112}$ fixed, the identified set of an average effect such as ((ref)) can be computed by relying on linear programming.

figure[figure omitted — 944 chars of source]

Figure (ref), reproduced from bonhomme2023identification, shows the identified sets for ${\Greekmath 0112}$ in the binary choice model $$Y_{it}=\boldsymbol{1}\{{\Greekmath 0112} X_{it}+A_i+U_{it}\geq 0\},$$ where $X_{it}$ are sequentially exogenous. The Figure was obtained under a particular data generating process, for which the true ${\Greekmath 0112}$ parameter is indicated on the x-axis. In the logit case (top panel), one sees that the identified set is a singleton under strict exogeneity (dashed line), yet that it is an interval with non-empty interior under sequential exogeneity (solid line). Moreover, going from $T=2$ to $T=4$ shrinks the size of the identified set substantially, to the point that it is essentially a singleton and ${\Greekmath 0112}$ is close to being point-identified. This observation suggests that, despite the failure of point-identification, identified sets may be informative in applications to (not too short) panels. In the probit case (bottom panel), the Figure shows similar conclusions, with the difference that point-identification fails under both strict and sequential exogeneity.

Partial identification is a promising approach in nonlinear panel data models with strictly or sequentially exogenous covariates (with or without a likelihood structure). See botosaru2024adversarial and chesher2024robust for recent proposals.

Restricted feedback

Restrictions on the feedback process can have identifying power. Consider first the case of homogeneous feedback,

equation[equation omitted — 221 chars of source]

under which the feedback process does not depend on $A_i$ and is identical across individuals. robins1986new, and subsequent work, develop approaches based on such sequential exchangeability assumptions. Condition ((ref)) is plausible in dynamic experiments, where the researcher controls the treatment $X_{it}$ and can adjust it depending on past outcome realizations. In non-experimental settings, it is a priori less plausible, although it may be a natural assumption in some economic settings.\footnote{An example is when $X_{it}$ is a dynamic state variable (e.g., wealth) whose law of motion does not vary across individuals (e.g., if returns to wealth do not vary across individuals).}

Consider again the case $T=2$ for illustration. Under homogeneous feedback, the feedback process is $$g(x_2\,|\, y_1,y_0,x_1),$$ independent of $a$. Hence, the characterization of the identified set for ${\Greekmath 0112}$ becomes (assuming that the denominator on the left-hand side is non-zero)

equation[equation omitted — 243 chars of source]

As bonhomme2023identification point out, the right-hand side in ((ref)) coincides with the likelihood of a model with strictly exogenous covariates (and $A_i$ independent of $X_{i2}$ conditional on $Y_{i0},X_{i1}$). This shows that, once the density of the data is properly weighted, one can use methods for the analysis of models with strictly exogenous covariates for identification and estimation.

Another possible restriction on the feedback process is Markovian feedback, such as the first-order Markov restriction $$X_{it}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ independent of }Y_i^{t-2},X_i^{t-2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ given } Y_{i,t-1},X_{i,t-1},A_i.$$ This type of restriction is common in dynamic structural models, and it has identifying content too. arellano2016nonlinear study identification under such Markovian assumptions (building on previous work by hu2008instrumental and hu2012nonparametric), and propose an estimation approach based on quantile regressions.

Looking ahead: feedback in networks

Most of the panel data literature focuses on single-agent models where individuals do not interact with each other. However, in many empirical applications, interactions and spillovers are of economic interest. The concept of feedback is similarly relevant to network settings.

Let $\boldsymbol{Y}_t=(Y_{1t},...,Y_{Nt})$ and $\boldsymbol{X}_t=(X_{1t},...,X_{Nt})$. Let also $\boldsymbol{A}=(A_1,...,A_N)$. The feedback process is the density of $$ \boldsymbol{X}_t\,|\, \boldsymbol{Y}^{t-1},\boldsymbol{X}^t,\boldsymbol{A}.$$ It is appealing to allow for general forms of feedback. For example, changes in network structure (i.e., $\boldsymbol{X}_t$) could be partly driven by shocks to outcomes (kuersteiner2020dynamic). An illustration of feedback in networks is given by Figure (ref). Models with unrestricted feedback allow past outcomes of all units, $Y_{i,t-1}$, to affect future unit-specific covariates $X_{i',t}$ (such as network links).

figure[figure omitted — 2,699 chars of source]

As an illustration, consider a model on a {bipartite network},

equation[equation omitted — 111 chars of source]

under the sequential exogeneity assumption

equation[equation omitted — 75 chars of source]

where $X^t$ denotes the set of all $X_{ijs}$ for $i=1,...,N$, $j=1,...,J$, and $s\leq t$.

For example, $i$ and $j$ may denote workers and firms, respectively, as in the AKM model of wage determination (abowd1999high), and $X_{ijt}$ denote the indicator that $i$ works in $j$ at $t$. In this model, $A_i$ and $B_j$ are worker and firm effects, respectively. However, model ((ref))-((ref)) allows for feedback, since it does not impose the strict exogeneity assumption of the AKM model that stipulates

equation[equation omitted — 74 chars of source]

While the so-called “exogenous mobility” condition ((ref)) rules out job mobility to be influenced by previous wage shocks $U_{it}$, ((ref)) allows for such dependence.\footnote{Model ((ref))-((ref)) also differs from a “distributed lags” specification that allows past firms to affect future wages, yet still relies on strict exogeneity (di2023ain).}

bonhomme2019distributional show that allowing for feedback is essential for the model to be compatible with {structural models} of sorting and wage determination. Here, the feedback process represents a dynamic model of {network} {formation} with heterogeneous workers and firms. An important advantage of leaving the feedback process unrestricted is that it is not necessary to specify a dynamic network formation model.

Estimation raises a number of issues, in part related to the existence of many moment restrictions. See mikusheva2025estimation, kuersteiner2020dynamic, and also bonhomme2019distributional for a setting that allows for nonlinear relationships between wages and (discrete) worker and firm heterogeneity. The literature on feedback in networks is still in its infancy, and more work is needed given the potential relevance of these methods for economic applications.

Concluding remarks

Many popular estimation methods rely on (explicit or implicit) {restrictive assumptions} about feedback. Feedback is central to many economic models, and often empirically plausible. Yet, many popular methods hinge on the assumption of strict exogeneity that rules out feedback entirely. It is important for applied researchers to be {aware of the dynamic} {restrictions} they impose, since those are often key for identification.

Allowing simultaneously for feedback and heterogeneity raises difficult challenges. While the recent literature has made some progress on these and other vexing issues, feedback should become (once again) a {central topic} for econometric research.