EconBase
← Back to paper

On Local Overidentification and Efficiency Gains in Modern Causal Inference and Data Combination

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

57,666 characters · 9 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On Local Overidentification and Efficiency Gains in Modern Causal Inference and Data Combination

abstract\singlespacing This paper studies nonparametric local (over-)identification and the semiparametric efficiency in modern causal frameworks. We develop a unified approach that begins by translating structural models with latent variables into their induced statistical models of observables and then analyzes local overidentification through conditional moment restrictions. We apply this approach to three popular classes of causal models: (1) the general treatment model under unconfoundedness; (2) the negative control model, and (3) the long-term causal inference model under unobserved confounding. The first model yields a locally just-identified statistical model, implying that all regular asymptotically linear estimators of the treatment effect have the same asymptotic variance, which equals the (trivial) semiparametric efficient variance bound. In contrast, the latter two models involve nonparametric endogeneity and are naturally locally overidentified; consequently, some doubly robust orthogonal moment estimators of the average treatment effect are inefficient. Whereas existing work typically imposes strong conditions to restore local just-identification to justify the efficiency of their doubly robust orthogonal moment estimators, we characterize the semiparametric efficient variance bounds, along with efficient estimators, for the (locally) overidentified models (2) and (3). A small real data application, along with a simulation study, illustrates the semiparametric efficiency gains in model (3). {\bf Keywords:} Causal Inference, Local Just Identification, Local Overidentification, Long-Term Treatment Effect, Negative Controls, Semiparametric Efficiency.

Introduction

In the era of generative artificial intelligence (AI) and abundant off-the-shelf machine-learning (ML) tools, it has never been easier to fit flexible models and report “black-box” causal effects. A common workflow estimates a nonparametric object, such as a conditional mean, quantile, or density, using whatever ML packages in a first stage, which is then plugged into an unconditional moment condition to estimate a causal parameter in the second stage. Hidden in this popular workflow is a fundamental econometric point that has been largely overlooked in the modern ML causal literature: local (over-)identification of the model.

Following chen2018overidentification, a statistical model is locally just identified at a data distribution when the model’s tangent space spans all valid score directions at that data distribution. Intuitively, this means that near the true data-generating process, any small change might be observed in the data can be matched by the model, so there is no additional information to make estimators more efficient and no locally testable restrictions.

Modern causal models, however, are typically structural rather than purely statistical: they involve unobserved variables—such as potential outcomes, negative controls, or latent confounders—and identification assumptions formulated as restrictions on the joint distribution of these unobserved variables. Our contribution is to make this distinction explicit and to provide a unified approach to studying local (over-)identification in causal models. We achieve this by translating each structural model into an observationally equivalent statistical model of observables and then analyzing local identification of the observable distribution. This structural-to-statistical translation, combined with the semiparametric efficiency calculation method for general sequential conditional moment restriction models in ai2012semiparametric, yields a unified framework for determining efficiency bounds and constructing efficient estimators, thereby avoiding the need for case-by-case analyses across different causal designs.

We use this framework to study three representative modern causal inference models: (1) the general treatment model under unconfoundedness ai2021unified, (2) the negative control model miao2018identifying, and (3) the long-term causal inference model under unobserved confounding imbens2025long. The first model is without nonparametric endogeneity: under unconfoundedness, treatment assignment is as good as random given observed covariates, so the causal effects are determined by comparing observable conditional distributions. In contrast, the negative control model and the long-term causal inference feature nonparametric endogeneity: the treatment effect is identified through solving a nonparametric instrumental variables (NPIV) type inverse problem, in which the nuisance functions appear inside a conditional expectation operator. As shown by chen2018overidentification, nonparametric endogeneity typically implies local over-identification, and efficiency considerations become essential in this case.

The unconfoundedness design, we show, is typically locally just identified. Two consequences follow, regardless of the sophistication of the ML first stage: (i) all regular, asymptotically linear estimators of the treatment effect are first-order equivalent, so the semiparametric efficiency bound is trivial; and (ii) there is no nontrivial specification test for the causal model. In contrast, designs with nonparametric endogeneity are naturally locally overidentified. Rather than imposing extra structure to force local just-identification as in the literature, we directly characterize the efficiency bound in the overidentified case and construct the corresponding efficient estimators. To this end, we formulate these designs as sequential moment condition models and apply ai2012semiparametric to obtain efficient influence functions.

Beyond the semiparametric efficiency results in ai2003efficient,ai2012semiparametric, a few other recent studies examine the case of overidentification and semiparametric efficiency. navjeevan2023identification studies the identification and semiparametric efficiency in a general class of instrumental variables model with effect heterogeneity. hahn2024overidentification analyze the overidentifying restrictions that arise in the shift-share (Bartik) instrument designs. chen2025efficient derives semiparametric efficient estimators in multi-period difference-in-differences design.

Overidentification can arise when nuisance functions are specified parametrically or treated as known. Imposing functional form restrictions can lead to local overidentification of the joint distribution of observables, allowing for the existence of estimators that are strictly more efficient than others. A well-known example is the inefficiency of the inverse probability weighting estimator that uses the true propensity score in estimating the average treatment effect hirano2003efficient, as well as in more general GMM models chen2004semiparametric,chen2008semiparametric. More recently, CarlsonDell2025 develop a unified framework for robust and efficient estimation with unstructured data that relies on a known propensity score (which they term the “annotation score”). These examples demonstrate that nontrivial efficiency bounds can emerge from functional form restrictions on nuisance functions. At the same time, in practice—especially in the current era of AI and ML—researchers often prefer to estimate nuisance functions nonparametrically using flexible, data-adaptive methods.

Local overidentification can also arise in semiparametric two-stage GMM settings, where the target parameter is defined by potentially overidentified unconditional moment conditions, while the first-stage nonparametric nuisance functions are just identified. In this case, ackerberg2014asymptotic show that the semiparametric two-step GMM estimator achieves efficiency.

The remainder of the paper is organized as follows. Section (ref) formally introduces the concept of local (over-)identification and its connection to semiparametric efficiency gains. Sections (ref), (ref), and (ref) examine the general treatment model, negative control, and long-term causal inference models, respectively. Section (ref) presents an empirical illustration along with a simulation study. Section (ref) concludes.

Review: Model over-identification and efficiency gains

In this section, we first introduce the concepts of structural and statistical models. We then review the key concept of local just-identification and over-identification of a statistical model introduced by chen2018overidentification. We also summarize its implications for the existence of efficiency gains in estimation and testing of the regular linear parameter of interest.

\paragraph{Structural and statistical model.} In economic applications, structural variables describe aspects of the data-generating process (DGP) and reflect the researcher’s view of underlying mechanisms. These variables may be hypothetical and exist only in economic theory. In modern causal inference, structural variables encompass potential outcomes under different treatments, potential treatments under different instruments, and possibly other latent variables.

Let $W^*$ denote the vector of structural variables, taking values in $\mathcal{W}^*$. The structural model $\mathbf{P}^*$ is the collection of joint distributions of $W^*$ consistent with the imposed structural assumptions, which capture the researcher’s economic intuition about the setting. The structural variables are not fully observed. We instead observe $W = s(W^*)$, where $s:\mathcal{W}^* \to \mathcal{W}$ is a transformation into the observable space. The statistical model is the set of distributions of $W$ induced by $\mathbf{P}^*$: \[ \mathbf{P} = \{ P = P^* \circ s^{-1} : P^* \in \mathbf{P}^* \}, \] where $P^* \circ s^{-1}$ is the pushforward measure of $P^*$ by the function $s$. Although the structural model encodes theoretical restrictions, estimation and inference are always carried out in the induced statistical model.

\paragraph{Model just-identification and over-identification.}

For any Euclidean set $\mathcal{W}$, denote $\mathbf{M}(\mathcal{W})$ as the set of all probability measures on $\mathcal{W}$, where we always consider the Borel sigma-algebra associated with the Euclidean space.

definitionA statistical model $\mathbf{P}$ is globally just identified if it is fully unrestricted, that is, $\mathbf{P} = \mathbf{M}(\mathcal{W})$. Conversely, $\mathbf{P}$ is globally overidentified if $\mathbf{P}$ is a strict subset of $\mathbf{M}(\mathcal{W})$.

The concept of global just identification is based on whether there is any restriction imposed on the statistical model $\mathbf{P}$. Although seemingly stringent, the structural assumptions may impose no observable restrictions, so the statistical model remains globally just identified. In Section (ref), we show that the unconfoundedness model is globally just identified even though assumptions are imposed in the structural model.

The concept of local identification requires the notion of tangent space. Let $L^2_0(P)$ denote the set of mean zero and $P$-square integrable functions, that is,

align*[align* omitted — 117 chars of source]

Take $g \in L^2_0(P)$ and $P \in \mathbf{P}$. A path is a mapping $\theta \mapsto P_{\theta,g}$ from $[0,\bar{\theta})$ to $\mathbf{M}(\mathcal{W})$ such that $P_{0,g} = P$, and

align[align omitted — 222 chars of source]

where $\mu_\theta$ is a $\sigma$-finite positive measure dominating $P_{\theta,g} + P$, and $p_{\theta,g}$ and $p$ denote the respective densities of $P_{\theta,g}$ and $P$.

For the statistical model $\mathbf{P}$, the tangent space at $P$ is defined as the closed linear span of the set of feasible scores within $\mathbf{P}$:

align*[align* omitted — 210 chars of source]

where $\textit{cl}$ denotes the closed linear span of a set. By definition, the tangent space $\bar{\mathscr{T}}(P)$ is a subset of $L^2_0(P)$. Whenever this inclusion holds as equality, we say the distribution $P$ is locally just identified by $\mathbf{P}$.

definitionA distribution $P \in \mathbf{P}$ is locally just identified by $\mathbf{P}$ if $\bar{\mathscr{T}}(P) = L^2_0(P)$. Conversely, $P \in \mathbf{P}$ is locally overidentified by $\mathbf{P}$ if $\bar{\mathscr{T}}(P) \subsetneqq L^2_0(P)$.

Under global just-identification, $P$ can be approach from any score direction, and hence local just-identification holds given regularity of the model. However, global over-identification does not imply local over-identification, while local over-identification implies global over-identification. See chen2018overidentification for details.

A parameter of interest is represented as a functional $\mu:\mathbf{P}\to\mathbb{R}^{d_\mu}$. Following AndrewsChenTecchio2025EstimatorPurpose, we refer to $(\mu,\mathbf{P})$ as an econometric model. A parameter $\mu$ is said to be regular at $P$ if there exists $\psi \in L^2_0(P)$ such that for every path $\theta \mapsto P_{\theta,g}$ with score $g \in T(P)$,

align*[align* omitted — 101 chars of source]

\paragraph{Regular asymptotically linear estimator.} An estimator $\hat\mu$ maps the independent and identically distributed (iid) sample $W_1,\ldots,W_n$ into $\mathbb{R}^{d_\mu}$. For a path $\theta \mapsto P_{\theta,g}$, denote $\overset{L_{n,g}}{\rightarrow}$ as convergence in law under $\otimes_{i=1}^n P_{1/\sqrt{n},g}$ and $\overset{L}{\rightarrow}$ as convergence in law under $P^n$. We use $o_P(1)$ to denote a term that converges in probability to zero under $P$. The estimator $\hat{\mu}$ is said to be a regular estimator of $\mu(P)$ if there is a tight random variable $\zeta$ such that

align*[align* omitted — 107 chars of source]

for any path $\theta \mapsto P_{\theta,g} \in \mathbf{P}$. The estimator $\hat{\mu}$ is asymptotically linear if

align*[align* omitted — 108 chars of source]

for some $\psi \in L^2_0(P)$. The function $\psi$ is the influence function of the estimator. The efficient influence function is the projection of any influence function onto the tangent space, attaining the minimal covariance matrix among all influence functions.

We emphasize some results from chen2018overidentification that we use in this paper: the presence of more efficient estimators of a regular parameter $\mu(P)$ is equivalent to the data distribution $P$ being locally over-identified by the model $\mathbf{P}$. Equivalently, in locally just-identified models, all regular, asymptotically linear estimators are first-order equivalent. In locally just-identified models, any specification test has trivial local power: along any path, the local asymptotic power of a level-$\alpha$ test cannot exceed $\alpha$.

\paragraph{Graphical demonstration.} To build intuition, we illustrate these concepts with a simple two-dimensional graphical demonstration. Consider the statistical variable $W$ whose support contains three points, $\mathcal{W}=\{a,b,c\}$. The set of distributions on $\mathcal{W}$ is equal to

align*[align* omitted — 99 chars of source]

which can be equivalently represented as a two-dimensional triangle:\footnote{This triangle of probabilities is often used for demonstration purposes in the literature. It is sometimes referred to as the Marschak-Machina triangle.}

align*[align* omitted — 90 chars of source]

For a specific distribution $P = (p_a,p_b,p_c)$, the set of all scores is equal to

align*[align* omitted — 80 chars of source]

which is equivalent to the set of all directions $(g_a,g_b)$ on the two-dimensional plane.

We consider three statistical models in this space:

align*[align* omitted — 218 chars of source]

By definition, the model $\mathbf{P}_1$ is globally just identified while the models $\mathbf{P}_2$ and $\mathbf{P}_3$ are globally overidentified. For local identification, we consider a distribution $P = (p_a=0.4,p_b=0.4)$ that belongs to all three statistical models. The distribution $P$ is locally just identified by the models $\mathbf{P}_1$ and $\mathbf{P}_2$ while locally overidentified by the model $\mathbf{P}_3$.

figure[figure omitted — 3,272 chars of source]

Figure (ref) illustrates the three statistical models. Both $\mathbf{P}_1$ and $\mathbf{P}_2$ are two-dimensional with nonempty interiors, while $\mathbf{P}_3$ is one-dimensional with no interior: the line segment representing $\mathbf{P}_3$ is the path $\theta \mapsto (p_a,p_b)(1+\theta)$. Intuitively, $P$ is locally just identified if and only if it is an interior point of the model. In Figures (ref)(a)–(b), $P$ lies in the interior, so a parametric submodel can approach $P$ from any direction. In Figure (ref)(c), $P$ is on the boundary, so submodels can approach only from two directions.

To illustrate the connection to efficiency gains, consider estimating $p_a = \mathbb{P}_P(W=a)$ from an iid sample $W_1,\ldots,W_n$. In models $\mathbf{P}_1$ or $\mathbf{P}_2$, any regular, asymptotically linear estimator has the same asymptotic variance as the sample mean $\tfrac{1}{n}\sum_{i=1}^n \mathbf{1}\{W_i=a\}$. In contrast, under $\mathbf{P}_3$, the additional restriction $p_a=p_b$ allows construction of a strictly more efficient estimator, $\tfrac{1}{2n}\sum_{i=1}^n \mathbf{1}\{W_i\in\{a,b\}\}$.

General treatment model under unconfoundedness

In this section, we introduce the general treatment model with the unconfoundedness assumption, derive the corresponding structural and statistical model, and show that the statistical model is globally just identified.

In the general treatment model, the treatment variable $T$ has support $\mathcal{T} \subset \mathbb{R}$, which may be discrete, continuous, or mixed. Let $Y(t)$ denote the potential outcome when treatment is assigned at $t \in \mathcal{T}$. For simplicity, we assume that the support of $Y(t)$, $\mathcal{Y} \subset \mathbb{R}$, is a bounded set so that the moments of $Y(t)$ exist. The conditioning covariate is $X \in \mathcal{X}$. The unconfoundedness assumption states that the treatment choice is conditionally independent of the potential outcomes, that is, $Y(t) \perp T | X$. Formally, we define the structural model to be

align*[align* omitted — 206 chars of source]

where the subscript UC denotes “unconfoundedness.” Any element $P^* \in \mathbf{P}^*_{\operatorname{UC}}$ of the structural model is a plausible joint distribution of $(\{Y(t)\}_{t \in \mathcal{T}},T,X)$ such that the unconfoundedness assumption is satisfied.

The potential outcomes are not observed. Instead, we observe the realized outcome $Y$ based on the treatment choice $T$: $ Y = Y(T).$ The statistical model is the set of distributions of the observable data $(Y,T,X)$ induced by the structural model. Let $s:\mathcal{Y}^\mathcal{T} \times \mathcal{T} \times \mathcal{X} \rightarrow \mathcal{Y} \times \mathcal{T} \times \mathcal{X}$ denote the transformation from structural variables to observed variables, that is,

align*[align* omitted — 67 chars of source]

The statistical model is formally defined as

align*[align* omitted — 113 chars of source]

where $P^* \circ s^{-1}$ denotes the pushforward measure of $P^*$ by the function $s$. That is, the statistical model $\mathbf{P}_{\operatorname{UC}}$ consists of all probability measures $P$ of the observable data $(Y,T,X)$ such that $P$ can be induced by some structural distribution $P^* \in \mathbf{P}^*_{\operatorname{UC}}$.

ai2021unified consider the following implicitly defined parameter $\mu(P)$:

align*[align* omitted — 157 chars of source]

where $L$ is a convex, differentiable loss function and $g$ is a parametric causal effect function known up to the parameter $\mu$. The terms $dP_{\cdot}$ and $dP_{\cdot|\cdot}$ denote the marginal and conditional density functions, respectively, and the ratio $dP_T/dP_{T| X}$ is defined where $dP_{T|X}>0$ and set to zero otherwise.

This general treatment model encompasses many prominent models in the causal inference literature as special cases: the average treatment effect (ATE) under binary treatment hahn1998role,hirano2003efficient, the quantile treatment effect firpo2007efficient, and multivalued treatments cattaneo2010efficient. See ai2021unified and the references therein. See also chen2024causal for efficient estimation of general treatment effects with a diverging number of confounders using modern neural network techniques.

Our analysis focuses on finite-dimensional causal parameters. For inference on fully nonparametric objects such as the conditional average treatment effect (CATE) function, see, for example, chang2015nonparametric,lee2017doubly.

The following theorem states our result regarding the general treatment model under unconfoundedness.

theoremThe statistical model $\mathbf{P}_{\operatorname{UC}}$ is globally just identified, that is, $\mathbf{P}_{\operatorname{UC}} = \mathbf{M}(\mathcal{Y}^{\mathcal{T}} \times \mathcal{T} \times \mathcal{X})$. Hence, the model is locally just identified.

Theorem (ref) implies that for estimating the same causal effect in the general treatment model, different regular, asymptotically linear estimators are first-order equivalent and attain the efficiency bound given by Theorem 1 in ai2021unified.

A subtle point is that ensuring the treatment effect parameter is $\sqrt{n}$-estimable requires a regularity condition that the efficient influence function has finite variance. \footnote{For example, as shown in chen2004semiparametric, $\sqrt{n}$-estimability can be achieved under a mild condition on the propensity score, which is weaker than assuming it is bounded away from zero.} This makes the model globally overidentified. However, as shown in the proof of Theorem (ref), imposing such a regularity condition does not affect the tangent space, and hence the model remains locally just identified.

When the propensity score is known or parametrically specified, the model is locally overidentified. In this case, there exist regular, asymptotically linear estimators that are strictly inefficient; for example, the inverse probability weighting estimator for the average treatment effect using the true propensity score is inefficient hirano2003efficient.

Negative control model

The previous model does not involve nonparametric endogeneity and is locally just identified. In this section and the next, we study models with nonparametric endogeneity, represented by NPIV-type conditional moment restrictions. Such models are naturally locally overidentified, as analyzed by chen2018overidentification.

Here, we focus on the negative control model studied by miao2018identifying and subsequent work. We follow tchetgen2024introduction to formulate the potential outcome framework. Let $X$ denote baseline covariates, $D\in\{0,1\}$ the treatment, $V$ the negative-control outcome, and $Z$ the negative-control exposure. For each treatment level $d\in\{0,1\}$ and exposure value $z$ in the support $\mathcal{Z}$ of $Z$, define the potential outcomes $Y(d,z)$ and $V(d,z)$ as the values $Y$ and $V$ would take if, possibly contrary to fact, we set $D=d$ and $Z=z$; likewise define the potential treatment $D(z)$ as the treatment that would be realized if $Z$ were set to $z$. The observed treatment and outcomes are generated according to

align*[align* omitted — 73 chars of source]

Thus, the observed variables are $(Y,D,V,Z,X)$.

The following structural assumptions are maintained for the identification of the average treatment effect.

enumerate[label=(\arabic*)] • Negative–control outcome: $V(d,z)=V$ for all $d,z$. • Negative–control exposure: $Y(d,z)=Y(d,z')=Y(d)$ for all $d$ and all $z,z'$. • Latent ignorability: there exists an unobserved latent variable $U$ such that, for all $d,z$, $(Y(d),V) \perp (D,Z) | (U,X).$ • Outcome bridge function: there exists a function $h$ such that \begin{align} \mathbb{E}\left[Y | D,X,Z\right]=\mathbb{E}\left[h(V,D,X)| D,X,Z\right]. \end{align} • Completeness: Given $D$ and $X$, the distribution of $U$ is complete for $Z$.

The estimand function can be written as

align*[align* omitted — 61 chars of source]

which identifies the average treatment effect under the standard overlap condition.

We denote the statistical model induced by the above identification assumptions as $\mathbf{P}_{\operatorname{NC}}$. In the following, we show that $\mathbf{P}_{\operatorname{NC}}$ is fully characterized by the NPIV-type moment ((ref)).

lemmaThe family of probability distributions of $(\{(Y(d,z),V(d,z),D(z)):d=0,1,z\in\mathcal{Z}\},U,Z,X)$ satisfying the above identification assumptions is observationally equivalent to the family of probability distributions of $(Y,D,V,Z,X)$ satisfying the moment condition ((ref)).

Let \( T : L^2(V,D,X) \to L^2(Z,D,X) \) be the conditional expectation operator given by \( (Th) \equiv \mathbb{E}[ h(V,D,X) \mid Z,D,X ] \), and let the adjoint \( T^* : L^2(Z,D,X) \to L^2(V,D,X) \) be \( (T^*g) \equiv \mathbb{E}[ g(Z,D,X) \mid V,D,X ] \).

theoremA distribution $P$ is locally just identified by the negative control model $\mathbf{P}_{\operatorname{NC}}$ if and only if the closure of the range space of $T$ is the full space: \begin{align*} \bar{\mathcal{R}} = cl(\{ f \in L^2(Z,D,X) : f = Th, \exists h \in L^2(V,D,X) \}) = L^2(Z,D,X), \end{align*} which is equivalent to the adjoint operator $T^*$ being injective.

cui2024semiparametric study the semiparametric negative-control model and propose a doubly robust estimator. They show that this estimator achieves semiparametric efficiency when both $T$ and $T^{*}$ are surjective, which in particular implies that $T^{*}$ is injective. Our Theorem (ref) provides a complementary perspective: when $T^{*}$ is injective (equivalently, when $T$ is surjective), the model is locally just identified, so any regular, asymptotically linear estimator attains the efficiency bound.

As noted by cui2024semiparametric, requiring both $T$ and $T^*$ to be surjective can be a demanding condition in practice. When $V$ and $Z$ are finitely valued, this requirement implies that their sample spaces must have the same cardinality (as in the case studied by shi2020multiply); similarly, when $V$ and $Z$ are continuous, it requires that they have the same dimension. If these requirements are not satisfied, the doubly robust estimator they propose may fail to achieve semiparametric efficiency. In what follows, we build on ai2012semiparametric to derive the efficient influence function for the estimand $\mu(P)$ without imposing such restrictions.

The estimation of ATE can be formulated as the following sequential moment restriction model:

align*[align* omitted — 173 chars of source]

where $\mu_1$ and $\mu_0$ are, respectively, the mean treated and untreated potential outcomes, and with a slight abuse of notation, we denote the nuisance function as $h=(h_1,h_0)'$, where $h_d(V,X)$ is shorthand for $h(V,d,X),d\in{0,1}$. The parameter ATE is equal to $\mu = \mu_1 - \mu_0$.

We introduce some notation. Let $e_1 = (1,0)$ and $e_0 = (0,1)$. Define the conditional moment function $\rho(h;Y,V,D,Z,X) = (D(Y - h_1(V,X)),(1-D)(Y - h_0(V,X)))'$. Let $\Sigma(Z, X) = \mathbb{E}[\rho \rho' | Z, X]$ denote the conditional variance matrix, and write $\Sigma^{-1}(Z,X)$ for its inverse. For $d = 0,1$, define $ \Gamma_d(Z, X) = \mathbb{E}[(h_d(V, X) - \mu_d)\rho' \mid Z, X] \Sigma^{-1}(Z,X), $ which captures the conditional covariance between the moment functions, and is used for orthogonalization. Finally, let $L(Z, X)$ be the diagonal matrix with diagonal entries $-p(Z, X)$ and $-(1 - p(Z, X))$, where $p(Z, X) = \mathbb{E}[D \mid Z, X]$.

theoremThe efficient influence functions for $\mu_1$ and $\mu_0$ in the negative control model are given by \begin{align*} \mathbb{EIF}(\mu_1) & = h_1(V,X) - \mu_1 - \Gamma_1 \rho - (LTH^{-1}T^*(e_1 - \Gamma_1 L)')'\Sigma^{-1}\rho, \end{align*} and \begin{align*} \mathbb{EIF}(\mu_0) & = h_0(V,X) - \mu_0 - \Gamma_0 \rho - (LTH^{-1}T^*(e_0 - \Gamma_0 L)')'\Sigma^{-1}\rho, \end{align*} respectively, where $H = L(Z,X)T^* \Sigma^{-1}(Z,X,D) TL(Z,X)$, and $H^{-1}$ is the generalized inverse of $H$. The efficient influence function for ATE is \begin{align*} \mathbb{EIF}(\mu) = \mathbb{EIF}(\mu_1) - \mathbb{EIF}(\mu_0). \end{align*} The semiparametric efficiency bound for ATE is given by the second moment of the efficient influence function.

In the proof of Theorem (ref), we formally demonstrate that the efficient influence function reduces to the one in cui2024semiparametric when $T$ and $T^*$ are bijective.

For efficient estimation, one may use either the optimally weighted minimum distance estimator proposed by ai2012semiparametric or the influence function–based efficient estimator developed in chen2023efficient. As the construction of estimators closely parallels that in the following section, we defer the detailed exposition and refer the reader to that discussion.

Long-term causal inference via data combination

Data combination provides a versatile toolkit for addressing measurement error, missing data, and treatment effect problems. See, e.g., chen2004semiparametric,chen2008semiparametric for a unified framework. Recent work has focused on identifying and estimating long-term causal effects by combining experimental datasets that report only short-term outcomes with observational datasets that contain both short and long-term outcomes. For example, chen2023semiparametric and athey2025experimental explicitly show that their models are locally just identified. We now focus on the framework studied by imbens2025long.

Following imbens2025long, let the treatment indicator be binary, $D\in\{0,1\}$. Each unit has a long-term potential outcome $Y(d)$ and a vector of short-term potential outcomes $S(d)=\bigl(S_{1}(d),S_{2}(d),S_{3}(d)\bigr)$, where the subscript indicates the time of measurement. The realized outcomes are $Y = Y(D)$ and $S = S(D)$.

Data come from two sources: an experimental sample and an observational sample. Let $G\in\{E,O\}$ indicate the source. In the experimental sample $(G=E)$, we observe only the short-term outcomes, whereas in the observational sample $(G=O),$ we observe both long- and short-term outcomes. Thus, the observed variables are $(Y\,\mathbf 1\{G=O\},\, S,\, D,\, G).$ For clarity, we suppress the covariates from the presentation.

Besides the potential outcomes, there is another variable $U$ that denotes the latent confounders that account for the association between treatment and potential outcomes in the observational data.

The following structural assumptions are imposed by imbens2025long:

enumerate[label=(\arabic*)] • Observational data: the latent $U$ accounts for all confounding in the observational data: $(Y(d),S(d))\perp D | U, G=O$. • Experiment: the treatment is independent of the potential outcomes and latent $U$ in the experimental data: $(Y(d),S(d),U) \perp D | G=E$. • External validity: the short-term potential outcomes and latent $U$ are independent with the data source $G$: $(S(d),U) \perp G$. • Sequential outcomes: the long-term $Y(d)$ and last period $S_3(d)$ are independent with the first period $S_1(d)$ conditional on the second period $S_2(d)$ and latent $U$ in the observational data: $(Y(d),S_3(d)) \perp S_1(d) | S_2(d),U, G=O$. • Completeness: given $S_2,D,$ and $G=O$, the distribution of latent $U$ is complete for $S_1$. • Outcome bridge function: there exists a function $h$ such that \begin{align} \mathbb{E}[Y|S_2,D,U,G=O] = \mathbb{E}[h(S_3,S_2,D)|S_2,D,U,G=O]. \end{align}

Under the maintained assumptions, imbens2025long show that the bridge function also satisfies the following observable conditional moment:

align[align omitted — 117 chars of source]

Denote the above conditional expectation operator by $K$. That is, for any $h \in L^2(S_3,S_2,D)$, $(Kh)(S_2,S_1,D) = \mathbb{E}[\mathbf{1}\{G=O\}h(S_3,S_2,D)|S_2,S_1,D]$. Denote its adjoint operator as $K^*$, i.e., $(K^*g)(S_3,S_2,D) = \mathbb{E}[\mathbf{1}\{G=O\}g(S_2,S_1,D)|S_3,S_2,D]$. In fact, any $h$ satisfying ((ref)) satisfies ((ref)), which leads to the identification of the Average Long-Term Treatment Effect (ALTTE) for the observational population:\footnote{In addition to the outcome-bridge approach, imbens2025long introduce an alternative identification strategy, the selection-bridge, for identifying the long-term treatment effect. When the two bridge functions are used jointly, the resulting identification is “doubly robust.” Since our paper focuses on semiparametric efficiency rather than the identification of causal parameters, we omit the analysis of the selection bridge function to maintain focus.}

align*[align* omitted — 99 chars of source]

We denote the statistical model induced by the above identification assumptions as $\mathbf{P}_{\operatorname{LT}}$. In the following, we show that $\mathbf{P}_{\operatorname{LT}}$ is fully characterized by the nonparametric-IV-type moment ((ref)). This result can be used to characterize when the model is locally just identified.

lemmaThe family of probability distributions of $(\{Y(d),S(d):d=0,1\},U,D,G)$ satisfying the above identification assumptions is observationally equivalent to the family of probability distributions of $(Y\mathbf{1}\{G=O\},S,D,G)$ satisfying the moment condition ((ref)).
theoremA distribution $P$ is locally just identified by the long-term causal inference model $\mathbf{P}_{\operatorname{LT}}$ if and only if the closure of the range space of $K$ is the full space: \begin{align*} \bar{\mathcal{R}} = cl(\{ f \in L^2(S_2,S_1,D) : f = Kh, \exists h \in L^2(S_3,S_2,D) \}) = L^2(S_2,S_1,D), \end{align*} which is equivalent to the adjoint operator $K^*$ being injective.

Theorem 7 in imbens2025long characterizes the semiparametric efficiency bound and proposes an efficient estimator for ALTTE under the condition that the operator $K$ is bijective. In this case, the adjoint operator $K^{*}$ is injective, so the model is locally just identified according to our Theorem (ref), and any regular, asymptotically linear estimator attains the efficiency bound. In more practical settings, however, $K^{*}$ may fail to be injective—e.g., when $S_{3}$ is not sufficiently rich to capture the full variation in $S_{1}$—leading the model to be locally overidentified. Our results highlight that in such cases, additional efficiency gains are possible beyond the estimator of imbens2025long.

To obtain the semiparametric efficiency bound in the general setting in which $K$ is not necessarily bijective, we formulate the estimation of ALTTE as the following sequential moment restriction model:

align*[align* omitted — 278 chars of source]

where $\mu_1$ and $\mu_0$ are respectively the mean treated/untreated potential outcome. The parameter ALTTE is equal to $\mu = \mu_1 - \mu_0$.

theoremThe efficient influence functions for $\mu_1$ and $\mu_0$ in the long term causal inference model are given by \begin{align*} \mathbb{EIF}(\mu_1) & = \frac{\mathbf{1}\{G=E\} D\left( h(S_3,S_2,D) - \mu_1 \right)}{\mathbb{P}(G=E,D=1)} \\ & + \frac{ \mathbf{1}\{G=O\}[KM^{-1} \pi(S_3,S_2,D)] D \Sigma_2^{-1}(Y - h(S_3,S_2,D))}{\mathbb{P}(G=E,D=1)} \end{align*} and \begin{align*} \mathbb{EIF}(\mu_0) & = \frac{\mathbf{1}\{G=E\} (1-D)\left( h(S_3,S_2,D) - \mu_0 \right)}{\mathbb{P}(G=E,D=0)} \\ & + \frac{ \mathbf{1}\{G=O\}[KM^{-1} \pi(S_3,S_2,D)] (1-D)\Sigma_2^{-1}(Y - h(S_3,S_2,D))}{\mathbb{P}(G=E,D=0)}, \end{align*} respectively, where $M = (K^* \Sigma_2^{-1}(S_2,S_1,D) K)$, $M^{-1}$ is the generalized inverse of $M$, $\Sigma_2$ is the conditional variance of the conditional moment: \begin{align*} \Sigma_2(S_2,S_1,D) = \mathbb{E}[\mathbf{1}\{G=O\}(Y - h(S_3,S_2,D))^2 | S_2,S_1,D], \end{align*} and $\pi(S_3,S_2,D) = \mathbb{P}(G=E|S_3,S_2,D)$. The efficient influence function for ALTTE is \begin{align*} \mathbb{EIF}(\mu) = \mathbb{EIF}(\mu_1) - \mathbb{EIF}(\mu_0). \end{align*} The semiparametric efficiency bound for ALTTE is given by the second moment of the efficient influence function: \begin{align} & \frac{1}{\mathbb{P}(G=E)^2} \mathbb{E}\left[ \mathbf{1}\{G=E\} \left( \frac{D - p_E}{p_E(1-p_E)} (h(S_3,S_2,D) - \mu(D))\right)^2 \right] \nonumber \\ + & \frac{1}{\mathbb{P}(G=O)^2} \left\lVert M^{-1/2} \left( \frac{D - p_O}{p_O(1-p_O)} \frac{\mathbb{P}(G=O|D)}{\mathbb{P}(G=E|D)} \pi(S_3,S_2,D) \right) \right\lVert^2, \end{align} where $p_E = \mathbb{P}(D=1|G=E)$, $p_O = \mathbb{P}(D=1|G=O)$, $\mu(D) = D\mu_1 + (1-D)\mu_0$, $\lVert \cdot \rVert$ is the $L^2$-norm in the Hilbert space.

The first component of the efficiency bound in ((ref)) coincides with the corresponding term in Theorem 7 of imbens2025long, whereas the second component generally differs. However, when $K$ is bijective, this second term reduces to the corresponding expression in Theorem 7 of imbens2025long.\footnote{A formal proof of this statement is given in the proof of Theorem (ref).}

Below, we introduce two efficient estimators for the ALTTE. First, we consider the efficient influence function-based estimator proposed by chen2023efficient, which automatically satisfies the Neyman orthogonality condition with respect to the nuisance parameters. We focus on the efficient estimator for $\mu_1$; the estimator for $\mu_0$ is symmetric. Taking the difference between the two estimators yields an efficient estimator for the ALTTE. In the proof of Theorem (ref), we have shown that the efficient influence function can be written in the following form:

align*[align* omitted — 261 chars of source]

where $v^*$ is the Riesz representer defined as

align[align omitted — 185 chars of source]

with $r^*$ being the minimizer of the following objective function

align[align omitted — 322 chars of source]

In implementation, given an iid sample $(Y_i \mathbf{1}\{G_i=O\}, S_i, D_i, G_i)$, one can first obtain a preliminary consistent estimator $\hat{h}$ for $h$, for example, using the conditional moment condition with identity weighting. Then, we can construct the following efficient influence function-based estimator:

align*[align* omitted — 362 chars of source]

where $\hat{v}^*$ is estimated as the sample analogue of ((ref)) and ((ref)) over the sieve space.

Second, we construct an estimator based on the optimally weighted minimum distance method. The analysis of the estimator as a minimum distance problem is a specialization of ai2003efficient,ai2007estimation,ai2012semiparametric,chenpouzo2015sieve. We first obtain a consistent estimator $\hat{\Sigma}_2$ using a preliminary estimate of $h$, and then estimate $\hat{h}$ by minimizing

align*[align* omitted — 167 chars of source]

over the sieve space, where $\hat{\mathbb{E}}[\cdot \mid S_2, S_1, D]$ denotes a consistent nonparametric estimator of the corresponding conditional expectation. The resulting estimator for $\mu_1$ is

align*[align* omitted — 137 chars of source]

Numerical studies

Empirical application: long-term effect of job training

In this section, we apply the efficient estimator developed in Section (ref) to study the long-term effect of job training on earnings. We use the Job Corps (JC) dataset as the experimental sample and the Survey of Income and Program Participation (SIPP) as the observational sample.

The JC dataset is obtained from the R package causalweight causalweight and originates from the Job Corps National Study, a large-scale randomized evaluation conducted by the US Department of Labor schochet2008job. It reports random assignment to the training program, which we take as the treatment $D$, as well as post-treatment log monthly earnings measured one, two, and three years after assignment, which form the short-term outcomes $S_1, S_2, S_3$. The covariates include years of schooling, age, race, and sex at birth.

The SIPP data are publicly available from the US Census Bureau.\footnote{See the Survey of Income and Program Participation, US Census Bureau, \url{https://www.census.gov/programs-surveys/sipp.html}.} We focus on the 1996 to 2000 panel to ensure close temporal alignment with the JC study. This panel consists of twelve waves (each wave corresponds to a four-month period), with training-related information reported in the topical module of Wave 2. We define the treatment $D$ as participation in any training intended to help train for a new job or improve skills in the current job during the past year. Short-term outcomes $(S_1,S_2,S_3)$ are defined as earnings measured in Waves 2, 5, and 8. The long-term outcome $Y$ is earnings measured in the final wave of the panel. The set of covariates is chosen to match those used in the JC dataset. In the end, the observational sample contains 28{,}232 untreated and 9{,}629 treated observations, while the experimental sample contains 3{,}663 untreated and 5{,}577 treated observations.

We implement both the estimator of imbens2025long and the efficient estimator proposed in Section (ref). The sieve spaces for the nuisance function $h$ are constructed using fourth-order spline bases. The probability $\pi$ is estimated using a sieve logit with spline basis, and the conditional covariance matrix is estimated via $K$-nearest neighbors (KNN). The results are reported in Table (ref).

table[table omitted — 614 chars of source]

Both estimators indicate that job training increases long term earnings by nearly 20%, a magnitude that is broadly consistent with existing evidence in the literature schochet2008job. The efficient estimator achieves a reduction in the standard error exceeding 20%, reflecting a substantial gain in statistical precision. The Hausman test yields no evidence against the model specification of this empirical study.\footnote{In cases of model misspecification, one can utilize the theory developed in ai2007estimation.}

Simulations

We run a simulation study with data-generating processes (DGP) that follows the simulation study in imbens2025long. We first generate four covariates $\tilde X$ from a multivariate normal distribution with mean zero and correlation coefficients calibrated to match those estimated in the empirical application. We also generate latent unobservables $U$ from a multivariate normal distribution with mean zero, variance $0.5$, and zero covariance across components.

Potential intermediate and final outcomes are generated according to the recursive system

align*[align* omitted — 345 chars of source]

where $\tau_y$ and $\tau_j$ are scalars and $(\alpha_y,\beta_y,\gamma_y)$ and $(\alpha_j,\beta_j,\gamma_j)$ are scalars, vectors, or matrices of conformable sizes. Following imbens2025long, we draw the entries of $\tau_y$, $(\tau_j,\alpha_y,\beta_y,\gamma_y)$, and $(\alpha_j,\beta_j,\gamma_j)$ from the uniform distribution on $[0,1]$ and rescale them so that the $\ell_2$ norms of the vectors $(\tau_j,\alpha_y,\beta_y,\gamma_y)$ and the columns of $(\alpha_j,\beta_j,\gamma_j)$ are equal to $0.5$. The disturbances $\epsilon_j$ are independent mean zero Gaussian variables with variance $0.5$. The error term $\epsilon_y$ has standard deviation $\exp(|\tilde X_1|)$ to introduce heteroskedasticity.

In the experimental sample, treatment is randomly assigned with probability one half. In the observational sample, treatment follows a logistic specification $\mathbb P(D=1 \mid \tilde X,U,G=O) = \left(1+\exp(\kappa_1^{\top}\tilde X+\kappa_2^{\top}U)\right)^{-1},$ where $(\kappa_1,\kappa_2)$ are generated using the same sampling and rescaling scheme. This induces persistent confounding through the dependence of treatment on both observed covariates and latent factors.

To introduce nonlinear bridge functions, we apply the elementwise transformation

align[align omitted — 86 chars of source]

to each component of $\tilde X,\tilde S_1,\tilde S_2,\tilde S_3$, yielding the observed variables $(X,S_1,S_2,S_3)$. This transformation is invertible and preserves the one to one correspondence between latent and observed variables. When $q=1$, the bridge functions are linear in $(X,S_1,S_2,S_3)$, whereas for $q=2$ they become nonlinear, allowing us to evaluate performance under both linear and nonlinear specifications. We consider $\lambda\in\{0.33, 0.50, 0.67\}$ as the share of experimental data in the sample.

The nuisance functions are estimated using spline bases with four knots and cubic degree. Ridge regularization is applied to the outcome regression to incorporate modern machine learning techniques, and the conditional variances are estimated via KNN.\footnote{Other choices of the tuning parameters deliver similar results.} Table (ref) reports the simulation results based on 500 Monte Carlo replications. As shown in the table, the efficient estimator achieves a reduction in variance across all designs. The dispersion of both estimators declines with the sample size at the root-$n$ rate, consistent with the asymptotic theory.

table[table omitted — 1,604 chars of source]

Conclusion

This paper develops a unified framework for analyzing local identification in modern causal inference by explicitly linking structural models to their corresponding observable statistical models. We show that the general nonparametric treatment model under unconfoundedness is locally just identified; consequently, all regular, asymptotically linear estimators are first-order equivalent, and the influence functions are all first-order equivalent to the efficient influence function. In contrast, models that are locally overidentified—such as those involving negative controls or long-term causal inference—allow for efficiency improvements and enable nontrivial specification tests.

These results highlight the central role of local overidentification in modern causal analysis. When developing new causal models, it is essential to carry out an explicit model identification analysis. If the model is locally just identified, effort should focus on higher-order properties of regular estimators as they are all first-order equivalent. When feasible, it is preferable to formulate locally overidentified models, which permit the construction of semiparametric efficient estimators, support specification testing with non-trivial powers, and yield richer empirical content.

For future research, the framework could be extended to accommodate dependent data structures and, for example, to investigate event study approaches to causal inference in time series settings linton2007quantilogram,han2016cross.