EconBase
← Back to paper

Influence Function: Local Robustness and Efficiency

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

110,596 characters · 22 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Influence Function: Local Robustness and Efficiency

{1.0\baselineskip} \hyphenpenalty=5000 \tolerance=1000

abstract\singlespace This paper introduces a direct differentiation-based framework that unifies the derivation of influence functions across parametric, nonparametric, and semiparametric models. We show that the Riesz representer of the functional derivative is obtained by orthogonally projecting the identification function onto the subspace of mean-zero functions. Consequently, the influence function emerges as a linear transformation of this centered moment function. The approach extends seamlessly to infinite-dimensional parameters, revealing a common algebraic form for influence functions across both finite- and infinite-dimensional parameters. Applied to semiparametric multi-step plug-in estimation, our method automatically yields locally robust moment functions and provides an explicit closed-form expression for the adjustment term. Finally, we leverage this framework to revisit the joint versus plug-in estimation debate, establishing verifiable sufficient conditions for their semiparametric efficiency equivalence even when nuisance parameters are over-identified. JEL classification: C13, C14. Keywords: Semiparametrics, influence function, locally robust, efficiency.

Introduction

The influence function is a foundational object in modern statistics and econometrics, characterizing the first-order effect of an infinitesimal perturbation to the data-generating distribution on a target parameter or estimator. It plays a central role in establishing asymptotic linearity, semiparametric efficiency, and local robustness. Seminal work by BKRW:1993 uses it to characterize semiparametric efficiency bounds, while recent literature CCDDHN:2018double,LocalRobust:2022 highlights its role in constructing locally robust moment functions.

Despite its theoretical centrality, deriving influence functions remains technically demanding. Stochastic expansions of the estimation error are often algebraically cumbersome. While tangent-space method tsiatis2006semiparametric is mathematical elegant, it can be abstract and opaque for complex models. As Hahn:1998 notes, the general frameworks of Newey:1990, Newey:1994 and BKRW:1993 may require one to posit a candidate influence function and verify its properties ex post. Hahn&Ridder:2013 extend Newey:1994 to three-step settings but restrict the analysis to differentiable functions. The method by Ichimura&Newey:2022 ultimately require solving a least square projection problem, where the nuisance moment function is scalar-valued. A unified, explicit, and calculus-driven recipe remains absent.

This work establishes a cohesive analytical framework, grounded in functional differentiation, for the unified derivation of influence functions across parametric and nonparametric specifications. We define the target parameter as a statistical functional of the probability distribution $P$ and compute its functional derivative under local perturbations of the form $P^\epsilon \!=\! (1-\epsilon)P + \epsilon Q$. For regular parameters, we show that the Riesz representer of this derivative is obtained by orthogonally projecting the identification function onto the subspace of mean-zero square-integrable functions. Crucially, this projection reduces to a simple centering operation: subtracting the unconditional or conditional expectation from the identification function (see Example (ref)).

The framework extends seamlessly to infinite-dimensional parameters. While the convergence rates differ fundamentally between finite- and infinite-dimensional settings, we demonstrate that the influence functions share a common algebraic structure: in both cases, they arise as linear transformations of the moment function (see Section (ref)). The transformation operator, rather than the moment function, governs the convergence rate. To operationalize this insight, we provide multiple application strategies and distill the derivation into a set of transparent calculus rules (Lemmas (ref), (ref), and (ref)).

In semiparametric multi-step plug-in estimation, our approach immediately yields the locally robust moment function. Deriving the associated adjustment term reduces to computing conditional expectations of partial derivatives. Crucially, this procedure does not require classical differentiability. By employing distributional derivatives, the method accommodates non-smooth moment conditions, with the conditional expectation serves as smoothing operator.

Using this framework, we derive easily verifiable conditions for local robustness in both parametric and semiparametric models and clarify its role in adaptive estimation. We revisit the classical question of Stein:1956: under what conditions can a finite-dimensional parameter be estimated as efficiently without knowledge of an infinite-dimensional nuisance parameter as with full knowledge of it? We show that local robustness ensures the asymptotic variance remains invariant to whether the moment function is evaluated at the true or a consistent estimator of the nuisance parameter, providing a direct answer to Stein's question.

We also resolve the classic debate between joint and plug-in estimation. As ACHL:2014 observe, plug-in procedures are computationally tractable but exploits moment conditions sequentially, potentially discarding information. Joint estimation, while theoretically more efficient, typically entails high-dimensional nonlinear optimization over both finite- and infinite-dimensional spaces. We employ an orthogonal decomposition of the moment function to decouple the joint optimization problem, rendering it structurally comparable to sequential plug-in estimation. In a general setting where the nuisance parameter may be over-identified, we identify an oracle-based locally robust moment function for the plug-in estimator that weakly dominates the locally robust version of the original moment funciton in terms of asymptotic variance. More importantly, we also establish easily verifiable sufficient conditions for the efficiency equivalence between the joint and the plug-in approaches.

This paper is organized as follows. Section (ref) introduces the formal setup and key definitions. Section (ref) develops the differentiation-based derivation across parametric, nonparametric, and semiparametric models. Section (ref) establishes the link between local robustness, adaptive estimation, and efficiency, and applies the framework to a range of treatment effect estimators. Section (ref) concludes. All technical proofs and supplementary results are collected in the online supplement.

Notation and Definitions

Parameter: identification and estimation

Let $Z$ be a random vector with distribution $P$. Following BKRW:1993, all parameters are viewed as functionals of $P$, e.g., $\nu\!=\!\nu(P)$ for a generic parameter $\nu$. This makes it more straightforward to define its influence function $\dot{\nu}$ as a functional derivative in the next subsection. Expectations are written as $P[f(Z)] \!\coloneqq\! \int f(z) dP(z)$, where $z$ is a real vector. Whenever possible, we will omit $Z$ or $z$ for simplicity.

We consider two types of identification for the parameter of interest $\beta$:

align[align omitted — 372 chars of source]

where $h_\beta$ and $m_\beta$ are known functions and $\gamma = (\gamma_1^\intercal, \ldots, \gamma_l^\intercal)^\intercal$ is the vector containing all the nuisance parameters. The direct case is a special instance of the indirect one, upon noticing that one can set $m_\beta \!=\! h_\beta - \beta$. When there are no nuisance parameters, an estimator of $\beta$ can be obtained by simply replacing $P$ with the empirical measure $\mathbb{P}_n$, i.e., $\hat{\beta}\!=\!\beta(\mathbb{P}_n)$. The asymptotic properties of $\hat{\beta}$ then readily follow from the well-established empirical process theory EmpiricalProcess.

Each $\gamma_j$ is assumed to be identified from $P_{(:|{j-1})}[m_{\gamma_j}]\!=\!0$, which may or may not reduce to a direct identification. The conditional distribution $P_{(:|{j-1})}$ is associated with the decomposition $Z \!=\! (Z^{(1)\intercal}, \ldots, Z^{(l)\intercal})^\intercal$. For $i \leq j$, let $Z^{(j:i)} \!=\! (Z^{(i)\intercal}, \ldots, Z^{(j)\intercal})^\intercal$. For $j \!=\! 2, \ldots, l$, let $P_{(:|{j-1})} \!\coloneqq\! P_{(l:j|j-1:1)}$ denote the conditional distribution of $Z^{(l:j)}$ given $Z^{(j-1:1)}$. When $j\!=\!1$, $P_{(:|{0})} \!=\! P$ reduces to the unconditional distribution. The decomposition is based on the levels of “exogeneity” of each $Z^{(j)}$. For example, we have $P_{(2|1)} \!=\! P_{(Y|X)}$ in linear regression model. In instrumental variables (IV) estimation, we get $P_{(2|1)} \!=\! P_{(Y,X|W)}$. In treatment effect analysis, $P_{(:|2)} \!=\! P_{(Y|T, X)}$ and $P_{(:|1)} \!=\! P_{(Y,T|X)}$.

In econometric literature, conditional and unconditional moment restrictions, as well as finite- and infinite-dimensional parameters, are often treated separately. Here, we try to find some symbolic similarities by adopting a generalized function perspective. More specifically, define the Dirac delta function $\delta_z$ as the distributional derivative of the Heaviside function $H_z(z') \!=\! 1_{\{z' \geq z\}}$, which satisfies $H_z[f] \!=\! f(z)$ in the Lebesgue–Stieltjes integral sense. With some abuse of notation, we symbolically write $H_z[f] \!=\! \int f(z') \delta_z(z') \mu(dz')$. Let $Z^{(1)}=X$ and consider the identification of the value of an unknown function $\gamma$ at $x$ (denoted $\gamma_x$):

align[align omitted — 232 chars of source]

where $\delta_{x}^\dag \!\coloneqq\! P[\delta_{x}]^{-1} \delta_{x}$. Note that the introduction of the generalized function $\delta_x^\dag$ symbolically transforms the conditional moment condition into an unconditional one.

remarkWe can write $\gamma_x \!=\! \gamma(P, \delta_x^\dag)$ following (ref). Section (ref) demonstrates that different nonparametric methods use different approximations to the generalized function $\delta_{x}^\dag$. Consider a kernel function $k_{b,x}$ with bandwidth $b$ and a growing linear sieve basis $\boldsymbol{\phi}_J(X) \!=\! (\phi_1(X), \dots, \phi_J(X))^\intercal$. Define \begin{align*} k_{b,x}^\dag(\cdot,P) \coloneqq (P[k_{b,x}(\cdot)])^{-1} k_{b,x}(\cdot) \quad and \quad \boldsymbol{\phi}_{J,x}^\dag(\cdot, P) \coloneqq \boldsymbol{\phi}_J(x)^\intercal (P[\boldsymbol{\phi}_J \boldsymbol{\phi}_J^\intercal])^{-1} \boldsymbol{\phi}_J(\cdot). \end{align*} The approximation step introduces the biased parameter $\gamma_x^{\texttt{B}} \!=\! \gamma(P, k_{b,x}^\dag)$ or $\gamma_x^{\texttt{B}} \!=\! \gamma(P, \boldsymbol{\phi}_{J,x}^\dag)$, respectively. Then we obtain the Nadaraya-Watson (NW) estimator and the linear sieve estimator by replacing $P$ with the empirical measure $\mathbb{P}_n$ in the biase parameters. Refer to Section (ref) for more details. Consistency eventually leads to a slower nonparametric convergence rate. Apart from this key difference, the structure of $\gamma_x^{\texttt{B}}$ is very similar to that of $\beta$ without nuisance parameters. The calculation of their influence functions follows the general principles we found.

When (ref) holds for all $x$ in the support of $X$, we can obtain a conditional moment condition $P[m_\gamma(Z, \gamma) \!\mid\! X] \!=\! 0$, without the appearance of generalized functions. Intuitively, this identifies the infinite-dimensional parameter $\gamma(\cdot)$, which is a function of $x$, as $\gamma(\cdot, P_{(Z|X)})$. We show in Section (ref) that computing the influence function of $\gamma(\cdot)$, denoted by $\dot{\gamma}$, is significantly more straightforward than computing $\dot{\gamma}_x$, where $\gamma_x \!\coloneqq\! \gamma(x)$. More importantly, the estimator of $\beta$ typically requires the entire function $\gamma(\cdot)$ as input. Consequently, it is $\dot{\gamma}$ rather than $\dot{\gamma}_x$ that appears in the influence function of $\beta$ (see Section (ref)).

Throughout the paper, we maintain the strong identification assumption and defer the analysis of weak identification to future work.

assumption[Strong Identification] Assume there exist positive semi-definite weighting matrices $\mathcal{W}_{\beta\beta}$ and $\mathcal{W}_{j}(z^{(j-1:1)})$ such that \begin{align*} P[\partial_\beta m_\beta]^\intercal \mspace{1mu} \mathcal{W}_{\beta\beta} \mspace{1mu} P[\partial_\beta m_\beta] \quad and \quad P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}]^\intercal \mspace{1mu} \mathcal{W}_{j}(Z^{(j-1:1)}) \mspace{1mu} P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}] \end{align*} are invertible almost surely for all $j=1,\ldots,l$.

We require only positive semi-definiteness of the weighting matrices to accommodate potential over-identification of $\beta$ and/or the nuisance parameters $\gamma_j$.

Influence function: functional derivative

Motivated by the structural parallels between finite- and infinite-dimensional parameters, we analyze a generic statistical functional $\nu\!=\!\nu(P)$. Let $\mathcal{L}_2(P)$ denote the space of square-integrable functions with respect to $P$, and let $\mathcal{L}_2^0(P) \!\subseteq\! \mathcal{L}_2(P)$ be the subspace of mean-zero functions. We equip $\mathcal{L}_2(P)$ with the inner product $\langle \cdot, \cdot \rangle$. Following standard semiparametric theory, we define functional derivatives via pathwise differentiability Hampel:1974, Huber:1984, BKRW:1993.

definition[Functional Derivative] The functional derivative $\dot{\nu}(P; Q-P)$ of a generic parameter $\nu$ at $P$ along the direction $Q-P$ is defined as \begin{align} \dot{\nu}(P; Q-P) \coloneq \frac{\partial} {\partial \epsilon} \nu\big( P^\epsilon \big) \big|_{\epsilon=0} = \frac{\partial} {\partial \epsilon} \nu\big( P+ \epsilon (Q-P) \big) \big|_{\epsilon=0}. \end{align} Let $\mathcal{S}$ be the collection of directions $Q-P$ such that the above derivative exists. Then different types of functional derivatives (G\^{a}teaux, Hadamard, or Fr\'{e}chet) correspond to different structures of the set $\mathcal{S}$ (refer to Appendix B.1 of the supplement and Section A.5 of BKRW:1993 for more details).\footnote{We do not specify the type of functional differentiability mainly because the asymptotic negligibility of the remainder terms is guaranteed by the stochastic equicontinuity condition Andrews:1994, Newey:1994, Newey&McFadden:1994.}

The influence function is defined similarly across methods, but the challenge lies in deriving its explicit form. For example, Newey:1994 uses a score-based expression, which is extended to three-step settings by Hahn&Ridder:2013; tsiatis2006semiparametric relies on tangent space methods and Ichimura&Newey:2022 propose a least squares-based approach.

Our method begins with the integral representation of the derivative $\dot{\nu}(P; Q - P)$. To ensure tractability, it is typically assumed that $\dot{\nu}(P; Q - P)$ is linear and continuous in $Q - P$. Under G\^{a}teaux differentiability, Huber:1984 shows that (Chapter 2.5 therein):

align[align omitted — 106 chars of source]

where $\dot{\nu}(P) \!\in\! \mathcal{L}_2^0(P)$. Note that the second equality holds only if $P[\dot{\nu}] \!=\! 0$, as otherwise infinitely many functions could satisfy the first equality (e.g., $\dot{\nu} + \mathbf{v}$ for any finite vector $\mathbf{v}$). By the Riesz representation theorem (e.g., Rudin:1987), the function $\dot{\nu}(P)$ is unique and is referred to as the influence function or influence curve, as in Hampel:1968, Hampel:1974. In other words, the Riesz representation here is a simple projection from $\mathcal{L}_2(P)$ onto $\mathcal{L}_2^0(P)$ by subtracting the mean. BKRW:1993 suggest taking $Q$ as a point mass at $z$ to obtain $\dot{\nu}(z, P)$ (pp. 19 therein). We will take an alternative approach, which applies to any admissible $Q$.

Regularity

The regularity of parameters is closely related to the functional derivative $\dot{\nu}(P; Q-P)$. The following definition is originally applicable to finite-dimensional parameters. However, a slight modification (see Section (ref)) extends this to infinite-dimensional parameters.

definition[Regular Parameter] A parameter $\nu(P)$, mapping $\mathcal{P} \rightarrow \mathbb{B}$ (where $\mathbb{B}$ is a Banach space), is regular at $P \in \mathcal{P}$ if: (i) Differentiability: The functional derivative $\dot{\nu}(P; Q-P)$, defined in Definition (ref), exists for all $Q - P \in \mathcal{S}$, where $\mathcal{S}$ is a sufficiently rich set. (ii) Linearity: $\dot{\nu}(P; Q-P)$ is continuous and linear in $Q - P$, implying the existence of a unique $\dot{\nu}(P) \in \mathcal{L}_2^0(P)$ satisfying (ref). (iii) Square Integrability and Invertibility: $\dot{\nu}(P)$ is square-integrable under $P$, and the variance-covariance matrix $\langle \dot{\nu}, \dot{\nu}^\intercal \rangle$ is invertible. If $\nu$ is regular at all $P \in \mathcal{P}$ and $\dot{\nu}(P; Q-P)$ is continuous in $P$, $\nu$ is termed a regular parameter.

Intuitively, differentiability ensures smooth variation of $\nu$ with respect to $P$. Linearity guarantees the existence of asymptotically linear estimators. Square integrability ensures finite variance for estimators with the parametric convergence rate; for nonparametric estimators, which have slower convergence rates, this may be adjusted (see Section (ref)).

Influence Function

Parametric case: prototype examples

We illustrate the derivation of the influence function $\dot{\nu}(P)$ from its functional derivative $\dot{\nu}(P; Q-P)$ using two foundational examples.

example[Direct identification] A standard example of direct identification is $\nu(P) \!=\! P[h(Z)]$, where $h \in \mathcal{L}_2(P)$ is a known function. Because $\nu$ is linear in $P$, its influence function follows directly from projecting $h$ onto the mean-zero subspace $\mathcal{L}_2^0(P)$: \begin{align*} \dot{\nu}(P; Q-P) = (Q-P)[h] \quad \Longrightarrow \quad \dot{\nu}(P) = h - P[h] = h - \nu. \end{align*}
remark[The Key Step: Integral Representation] The core of our approach is to express the functional derivative $\dot{\nu}(P; Q - P)$ in the form $(Q - P)[h]$ for some integrable function $h$. The influence function is then recovered by centering $h$, i.e., subtracting $P[h]$. This representation parallels the relationship between a directional derivative $\nabla_{\mathbf{v}} f(x) \!=\! \nabla f(x) \cdot \mathbf{v}$ and the gradient $\nabla f(x)$. When $Q$ is absolutely continuous with respect to $P$ with density $g \!=\! dQ/dP$, we have $\dot{\nu}(P; Q - P) \!=\! \langle \dot{\nu}(P), g \rangle$, where $\dot{\nu}(P)$ acts as the functional gradient. As in finite-dimensional calculus, identifying the inner product of $\dot{\nu}(P)$ with a sufficiently rich class of perturbation densities uniquely determines the influence function, even in infinite-dimensional settings. We will extend this identification argument below in the context of sequential plug-in estimation under both unconditional and conditional constraints.
example[Indirect identification] Consider a parameter $\nu$ identified indirectly via the moment restriction $P[m(Z, \nu)] \!=\! 0$, with $d_m \geq d_\nu$. Let $A$ be a $d_\nu \!\times\! d_m$ matrix of full row rank, and assume $A \, P[\partial_\nu m]$ is invertible, which constitutes a strong identification condition. Although the implicit function $\nu(P)$ typically lacks a closed-form expression in this case, its influence function $\dot{\nu}(P)$ admits an explicit analytic representation. Define $F(\nu, A, P) \!\coloneq\! P[A \, m(\nu)]$. The moment condition implies $F(\nu, A, P) \!\equiv\! 0$. We can then use the implicity function theorem to solve $\dot{\nu}(P)$. Differentiating with respect to $P$ yields (noting that $\partial_\nu F$ is a fixed invertible matrix): \begin{gather*} \frac{\partial F}{\partial \nu} \dot{\nu}(P; Q-P) + (Q-P)\big[ A \, m\big] = 0 \\ \Longleftrightarrow \dot{\nu}(P; Q-P) = - (\partial_\nu F)^{-1} (Q-P)[ A \, m] = - (Q-P)[ (\partial_\nu F)^{-1} A \, m]. \end{gather*} There exists a symmetric weighting matrix $\mathcal{W}$ such that $A \!=\! P[\partial_\nu m]^\intercal \mathcal{W}$. Using $P[m] \!=\! 0$, the influence function simplifies to: \begin{align} \dot{\nu}(P) = - \big( P[\partial_\nu m]^\intercal \mathcal{W} \, P[\partial_\nu m] \big)^{-1} P[\partial_\nu m]^\intercal \mathcal{W} \mspace{1mu} m \eqqcolon - P[\partial_\nu m]_{L}^{-1} \mspace{1mu} m, \end{align} where $P[\partial_\nu m]_{L}^{-1} \!\coloneqq\! \big( P[\partial_\nu m]^\intercal \mathcal{W} \, P[\partial_\nu m] \big)^{-1} P[\partial_\nu m]^\intercal \mathcal{W}$ denotes the left inverse of $P[\partial_\nu m]$. Consequently, for any admissible weighting matrix $\mathcal{W}$, the influence function satisfies the fundamental normalization condition $P[\partial_\nu \dot{\nu}] \!=\! -I$. A natural candidate for the optimal weighting matrix is the Moore--Penrose inverse $V_\nu^+$, where $V_\nu \!\coloneqq\! \langle m, m^\intercal \rangle$. The condition $(I - V_\nu V_\nu^+) P[\partial_\nu m] \!=\! 0$ is sufficient to ensure that this generalized inverse is indeed the optimal choice. When this condition holds, the efficient asymptotic variance simplifies to $\Sigma_{\nu\nu}^{-1} \coloneqq \{ P[\partial_\nu m]^\intercal \langle m, m^\intercal \rangle^+ P[\partial_\nu m] \}^{-1}$. The influence function $\dot{\nu}(P)$ admits a multiplicative representation that extends naturally to nonparametric settings (see Section (ref)). This structure underpins our comparison of joint and sequential plug-in estimation in Section (ref), where we partition $\nu^\intercal \!=\! (\beta, \gamma)$ and $m_\nu^\intercal \!=\! (m_\beta, m_\gamma)$.

The derivation of the linear system $(\partial_\nu F) \dot{\nu}(P; Q - P)$ follows from a standard perturbation argument. Let $\nu^\epsilon \!\coloneqq\! \nu(P^\epsilon)$. Intuitively, the pertubated parameter $\nu^\epsilon$ satisfies the moment condition up to the order $o(\epsilon)$ (or zero): $P^\epsilon[m(\nu^\epsilon)] \!=\! o(\epsilon)$. Applying a first-order Taylor expansion around $(P, \nu)$ yields:

align*[align* omitted — 225 chars of source]

For this equality to hold for all admissible directions $Q-P$, the $O(\epsilon)$ terms must vanish, which yields the linear system used to solve for $\dot{\nu}(P; Q-P)$. Evaluating such expressions requires handling terms of the form $P[w(Z) \dot{\nu}(P; Q - P)]$, which we resolve via the following key representation.

lemma[The Key Representation: Unconditional Case] For a regular parameter $\nu$ and a function $w$ such that the following expectations are all well-defined, we have \begin{align} P[w(Z) \dot{\nu}(P; Q-P)] = P[w(Z)] \cdot (Q-P)[\dot{\nu}(P)] = (Q-P)\big[ P[w(Z)] \mspace{1mu} \dot{\nu}(P) \big]. \end{align} This can also be written as $P[w(Z) \dot{\nu}(P)] = P[w(Z)] \dot{\nu}(P)$.

The proof relies on the richness of the admissible perturbation class. Because the functional derivative is linear in $Q-P$ and the set of admissible directions spans a dense subspace of $\mathcal{L}_2^0(P)$, we may treat the perturbation density as separable from the baseline measure for the purpose of integration. Formally, applying Fubini’s theorem to the double integral yields:

align*[align* omitted — 251 chars of source]

Lemma (ref) below extends this representation to the conditional case, which is essential for deriving influence functions in nonparametric and semiparametric models.

remarkThe representation extends directly when multiple infinite-dimensional parameters are identified under the same conditional law. The primary technical difficulty in semiparametric settings arises when the target parameter $\beta$ and nuisance parameter $\gamma$ are identified under different probability measures, requiring careful accounting of cross-measure dependencies.

Notably, the moment function $m$ needs not be classically differentiable with respect to $\beta$, contrasting with Hahn&Ridder:2013, who impose standard smoothness conditions (see the statement preceding their Lemma 1). Instead, we employ distributional (Sobolev) derivatives Sobolev, which extend differentiation to functions with discontinuities and preserve the validity of integration-by-parts and Fubini-type interchanges.

example[Non-smooth moment] Consider the $q$-th quantile $\nu_q$ of $Z$, with $q \in (0,1)$. The moment function $m(z, \nu_q) \!=\! \mathds{1}_{\{z \leq \nu_q\}} - q$ is non-differentiable in the classical sense, as it corresponds to a shifted Heaviside function. Interpreting the derivative in the distributional sense, $\partial_{\nu_q} m$ corresponds to the Dirac delta $\delta(\nu_q - z)$, yielding $P[\partial_{\nu_q} m] \!=\! P[\delta(\nu_q - Z)] \!=\! \text{\Fontskrivan{p}}(\nu_q)$, where $\text{\Fontskrivan{p}}(\cdot)$ denotes the density of $Z$. Provided $\text{\Fontskrivan{p}}(\nu_q) \!>\! 0$, the influence function is: \begin{align*} \dot{\nu}_q(P) = - \Big( P \Big[ \frac{\partial m}{\partial \nu_q} \Big] \Big)^{-1} m(Z, \nu_q) = - \frac{m(Z,\nu_q)}{\Fontskrivan{p}(\nu_q)} = \frac{q - \mathds{1}_{\{ Z \leq \nu_q \}}}{\Fontskrivan{p}(\nu_q)} . \end{align*}

Nonparametric influence function

We illustrate our method using the standard nonparametric regression model, where $Z \!=\! (X^\intercal, Y)^\intercal$:

align*[align* omitted — 98 chars of source]

A critical distinction must be drawn between the functional $\gamma(X) \!=\! P_{(Y|X)}[Y]$, which depends on the random variable $X$, and its point evaluation $\gamma_x \!\coloneqq\! \gamma(x)$ at any given $x$.

Expanding the perturbed joint measure to first order in $\epsilon$ yields

align*[align* omitted — 248 chars of source]

Algebraic rearrangement gives the corresponding perturbation of the conditional law:

align[align omitted — 181 chars of source]

Crucially, the identification $\gamma(X) \!=\! P_{(Y|X)}[Y]$ depends only on the conditional measure and is invariant to perturbations of the marginal $P_{(X)}$. This invariance permits a direct derivation of $\dot{\gamma}(P_{(Y|X)})$ (refer to the Proof of Theorem (ref) for a general treatment). Setting $Q_{(X)} \!=\! P_{(X)}$ (so that $dQ_{(X)}/dP_{(X)}^\epsilon \!\equiv\! 1$) and extending Definition (ref) to conditional measures yields

align[align omitted — 240 chars of source]

This mirrors the direct identification structure of Example (ref). Under the standard assumption that $\text{Var}(Y \!\mid\! X)$ is finite, the infinite-dimensional parameter $\gamma(\cdot)$ satisfies the regularity conditions outlined in Definition (ref).

By contrast, the point evaluation $\gamma_x$ depends explicitly on the marginal density $\text{\Fontskrivan{p}}(x) \!=\! P[\delta_x(X)]$. Specifically, $\gamma_x$ is a nonlinear functional of both $P$ and the Dirac measure $\delta_x$:

align*[align* omitted — 160 chars of source]

Constructing an estimator for $\gamma_x$ requires approximating the singular functional $\delta_x$ with a sequence of regular functionals. Consider the Nadaraya–Watson (NW) kernel estimator and the linear sieve estimator with a growing basis $\boldsymbol{\phi}_J(X) = (\phi_1(X), \dots, \phi_J(X))^\intercal$:

align*[align* omitted — 321 chars of source]

The approximation of $\delta_x$ induces asymptotic bias, whereas the substitution of $P$ with $\mathbb{P}_n$ governs the stochastic fluctuation. To formalize this decomposition, define a smoothed (or approximating) functional $\gamma^{\texttt{B}}_x(P)$ for each case, satisfying $\hat{\gamma}_x = \gamma_x^{\texttt{B}}(\mathbb{P}_n)$:

align*[align* omitted — 267 chars of source]

This yields the standard bias–variance decomposition:

align[align omitted — 275 chars of source]

The bias component originates from the regularization (or approximation) of $\delta_x$, whereas the first term captures the sampling variability induced by the empirical measure.

In the parametric setting, direct identification yields a functional that is linear in the underlying distribution (Example (ref)). By contrast, the biased nonparametric functional $\gamma_x^{\texttt{B}}(P)$ is inherently nonlinear due to the normalization induced by kernel or sieve approximations. Nevertheless, the functional derivative $\dot{\gamma}_x^{\texttt{B}}(P; \mathbb{P}_n - P)$ continues to characterize the dominant stochastic term in the von Mises expansion of $\gamma_x^{\texttt{B}}(\mathbb{P}_n) - \gamma_x^{\texttt{B}}(P)$. Specifically, the derivative isolates the first-order empirical process fluctuation, while higher-order remainders are of smaller orders.

Applying a suitably modified version of Definition (ref) to $\gamma_x^{\texttt{B}}$, the functional derivative $\dot{\gamma}_x^{\texttt{B}}$ serves as the nonparametric influence function. While its derivation requires careful handling of the approximation sequence, it yields transparent closed-form expressions that foreshadow the treatment of sequential plug-in semiparametric models. Below, we present two complementary approaches to unify the kernel and sieve specifications within our differentiation framework.

A non-composite function perspective

The first approach is to view $\gamma^{\texttt{B}}_x$ as a special case of a generic parameter $\nu_x(P) \!=\! c^\intercal (P[g])^{-1} P[h]$:

align*[align* omitted — 251 chars of source]

Although the linear sieve case involves matrix-valued functions, the derivation of influence function remains rather intuitive. Taking differential on both sides of $M M^{-1} \!\equiv\! I$ yields (where $M$ is any invertible matrix-valued function):

align*[align* omitted — 107 chars of source]

Based on this, we have

align*[align* omitted — 186 chars of source]

Plugging in the corresponding values for $c, g, h$, the nonparametric influence functions are

align[align omitted — 238 chars of source]

where both $k_{b,x}^\dag(\cdot,P)$ and $\boldsymbol{\phi}_{J,x}^\dag(\cdot, P)$ are approximations of $\delta_x^\dag(\cdot,P) \!\coloneqq\! (P[\delta_x(\cdot)])^{-1} \delta_x(\cdot)$:

align[align omitted — 291 chars of source]

The structure of the influence function becomes clearer in this context, mirroring (ref), where $\dot{\nu}(P)$ is linear in the moment function $m(Z, \nu)$. Here, the moment function is $m(Z, \gamma) \!=\! Y - \gamma(X)$, and $\dot{\gamma}^{\texttt{B}}_x$ is linear in $m(Z, \gamma^{\texttt{B}}_x)$. The “loading” term, such as $k_{b,x}^\dag$ or $\boldsymbol{\phi}_{J,x}^\dag$, reflects the estimation method and links the moment function to the influence function. For fixed $b$ or $J$, the variance of $\dot{\nu}_x(P)$ is finite, but to ensure $\gamma^{\texttt{B}}_x - \gamma_x \!=\! o_p(1)$, we need $b \to 0$ or $J \to \infty$, leading to diverging variances in the loading terms. The convergence rate of the nonparametric estimator is root-$n$ adjusted by this divergence rate. As long as the rate-adjusted variance of $\dot{\gamma}^{\texttt{B}}_x(P)$ is finite, $\gamma_x$ can be viewed as a (rate-adjusted) regular parameter.

A composite function perspective

Using the normalized weights defined in (ref), we can express the estimators as plug-in functionals:

align*[align* omitted — 289 chars of source]

Accordingly, the smoothed functional $\gamma^{\texttt{B}}_x(P)$ is a special case of the composite functional:

align[align omitted — 168 chars of source]

This representation isolates a key technical feature of our approach: $\nu_x(P)$ depends on $P$ both through the expectation operator and through the weight function $f_x^\dag(\cdot, P)$. Applying the chain rule for functional derivatives yields

align*[align* omitted — 99 chars of source]

The second term requires careful handling because $\dot{f}_x^\dag(X,P; Q-P)$ depends linearly on the perturbation $(Q-P)$ but retains a stochastic dependence on $X$. When integrating with respect to $P$, components independent of $(Q-P)$ are grouped with $Y$ and evaluated under the baseline measure. Formally, this interchange relies on the linearity of expectation and the separability of the perturbation direction. Parallel to the previous calculation, we obtain:

gather*[gather* omitted — 368 chars of source]

In the NW specification, the derivation simplifies considerably because $\nu_x(P)$ is a scalar constant with respect to $X$:

align*[align* omitted — 229 chars of source]

For the linear sieve estimator, the coefficient $\boldsymbol{\phi}_J(x)^\intercal P[g]^{-1}$ is deterministic conditional on $x$:

align*[align* omitted — 479 chars of source]

Combining these results with the first term $(Q-P)[f_x^\dag Y]$ recovers the influence function derived in the previous subsection, confirming the internal consistency of the two perspectives.

A third angle: toward semiparametric models

Incorporating $\gamma(X) \!=\! P_{(Y|X)}[Y]$ yields a third representation of $\gamma_x(P)$ that mirrors the structure of semiparametric plug-in estimation in (ref):

align*[align* omitted — 126 chars of source]

If the weighting function $f_x^\dag(\cdot, P)$ were square-integrable rather than a regularized approximation to a Dirac measure, this setup would correspond to a standard semiparametric model.

More generally, consider $\gamma_x(P)$ as a special case of the composite functional $\nu_x(P) \!=\! P[ w(Z, P) \nu(P_{(Y|X)}) ]$, where $w$ is a measurable function of $Z$ and $\nu$ is a functional of the conditional distribution $P_{(Y|X)}$. Applying the product rule for functional derivatives yields

align*[align* omitted — 206 chars of source]

where $\dot{\nu}(P_{(Y|X)}; Q_{(Y|X)} - P_{(Y|X)}) \equiv \dot{\nu}(P_{(Y|X)}; Q - P)$ since $\nu(P_{(Y|X)})$ is immunte to any pertubation of $P_{(X)}$.

The first term is already in the canonical $(Q-P)[\cdot]$ form (cf. Example (ref)). The second term follows from the composite function perspective developed above, combined with Lemma (ref). The third term captures the perturbation of the inner conditional functional and is defined by

align*[align* omitted — 208 chars of source]

Using the perturbation expansion in (ref) and the zero-mean property $P_{(Y|X)}[\dot{\nu}(P_{(Y|X)})] \!=\! 0$, we evaluate the limit as follows:

align*[align* omitted — 812 chars of source]

The limit

align*[align* omitted — 99 chars of source]

formalizes the change-of-measure step: the outer integration with respect to $P_{(X)}$ is replaced by integration against $Q_{(X)}$. The subsequent calculation eventually replaces the expectation of $w$ under $P$ with that under $P_{(Y|X)}$. Notably, even when $w$ is a regularized approximation to a Dirac measure, the smoothing parameters (e.g., bandwidth or sieve dimension) remain fixed at this stage of the derivation. Thus, $w$ stays bounded and square-integrable, ensuring that Fubini's theorem applies without requiring passage to the singular limit.

By replacing $\nu$ with $\gamma$ and $w$ with $f_x^\dag$, this representation recovers the same influence function as the previous two approaches, but explicitly isolates the conditional perturbation component.

lemma[The Key Representation: Conditional Case] Suppose $\nu$ is a functional of the conditional distribution $P_{(Y|X)}$. Then for any measurable $w$, \begin{align} \begin{gathered} P[w(Z) \mspace{1mu} \dot{\nu}(P_{(Y|X)}; Q_{(Y|X)} - P_{(Y|X)})] = P[w(Z) \mspace{1mu} \dot{\nu}(P_{(Y|X)}; Q - P)] \\ = P\big[w(Z) \mspace{1mu} (Q-P)[\dot{\nu}(P_{(Y|X)})] \big] = (Q-P)\big[ P_{(Y|X)}[w(Z)] \dot{\nu}(P_{(Y|X)}) \big] \\ \Longrightarrow\,\, P[w(Z) \dot{\nu}(P_{(Y|X)})] = P_{(Y|X)}[w(Z)] \dot{\nu}(P_{(Y|X)}), \end{gathered} \end{align} where $w$ may represent a regular function or a regularized approximation to a Dirac measure. An analogous result holds with $P_{(Y|X)}$ repalced by any conditional distribution in the previous decomposition.

Lemmas (ref) and (ref) share a common algebraic structure. Let $P_{(:|{j-1})}$ denote the conditional distributions defined in Section (ref). Both results are unified by the following general rule.

lemma[Expectation-Substitution Rule] Assume $1 \!\leq\! j' \!\leq\! j'' \!\leq\! l$. Define $P' \!\coloneqq\! P_{(:|{j'-1})}$ and $P'' \!\coloneqq\! P_{(:|{j''-1})}$. For any measurable weight $w$ and functional $\nu$ identified under $P''$, we have \begin{align} P'[w(Z) \dot{\nu}(P”)] = P”[w(Z)] \mspace{1mu} \dot{\nu}(P”). \end{align} When $j'\!=\!1$ and $j''\!>\!1$, $P'\!=\!P$ and the rule coincides with Lemma (ref). When $j'\!=\!j''\!=\!1$, $P'\!=\!P''\!=\!P$ and it reduces to Lemma (ref).

This general rule bypasses model-specific tangent-space characterizations and guess-and-verify constructions, reducing the derivation to a straightforward expecation.

More examples and discussions.

example[Functional coefficient] Consider the exogenous functional coefficient model: \begin{align*} Y = \beta(W)^\intercal X + \epsilon, \quad P_{(Y|W,X)}[Y - \beta(W)^\intercal X] = 0. \end{align*} The local linear estimator of $\nu_w \coloneqq (\beta(w)^\intercal, \text{vec}(\partial_w \beta)^\intercal )^\intercal$ at a fixed $w$ is \begin{gather*} \hat{\nu}_w = \nu_w(\mathbb{P}_n) = \mathbb{P}_n[f_w^\dag(W, X, \mathbb{P}_n) Y], where f_w^\dag(W, X, P) = ( P[ \tilde{X}_w k_{b,w} \tilde{X}_w^\intercal ] )^{-1} \tilde{X}_w k_{b,w}. \end{gather*} The term $\tilde{X}(w) \!=\! E_w \otimes X $ with $E_w \!=\! (1, (W - w)^\intercal)^\intercal$. Applying the differentiation rule yields the influence function \begin{align*} \dot{\nu}_w^{B}(P) = f_w^\dag(W, X, P) (Y - \beta(W)^\intercal X). \end{align*} Extending the basis $E_w$ yields the standard local polynomial estimator Fan&Gijbels:1996.
example[Endogeneity] Next, consider a nonparametric instrumental variable model with $Z^{(1)} \!=\! W$ and $Z^{(2)} \!=\! (X, Y)$: \begin{align*} Y = \gamma(X) + \epsilon,\quad P_{(Y,X|W)}[Y-\gamma(X)] = P[\delta_w^\dag(W, P) (Y-\gamma(X))] = 0. \end{align*} This is a well-known ill-posed inverse problem (see, e.g., Newey&Powell:2003, Hall&Horowitz:2005, Che&Reiss:2011, and DFFR:2011 among others). The sieve method can be viewed as a projection onto a subspace, which regularizes the problem by reducing its dimensionality. One can also add an explicit penalty term to further regularize the problem. In particular, a ridge-type penalty can yield a closed-form solution for the estimator $\hat{\gamma}_x \!=\! \gamma_x(\mathbb{P}_n) \!=\! \mathbb{P}_n[f_x^\dag(W, \mathbb{P}_n) Y]$ with \begin{align*} f_x^\dag(W, P) &= \phi_J(x) (P[ \phi_J \psi_K^\intercal] (P[\psi_K \psi_K^\intercal])^{-1} P[\psi_K \phi_J^\intercal] + \lambda )^{-1} \\ &\qquad \times P[\phi_J \psi_K^\intercal] (P[\psi_K \psi_K^\intercal])^{-1} \psi_K(W). \end{align*} Here, $\psi_K(W)$ denotes a growing sieve basis for the instruments. When $\lambda\!=\!0$, identification requires $K\!\geq\! J$. The smoothed functional and its influence function are \begin{align*} \gamma_x^{B}(P) = P[f_x^{\dag}(W, P) Y] and \dot{\gamma}_x^{B}(P) = f_x^{\dag}(W, P) (Y - \gamma(X)). \end{align*} Despite the added complexity from endogeneity, which is absorbed into the weighting function $f_x^\dag(W, P)$, the multiplicative structure of the influence function remains identical to the exogenous case.

In multivariate settings, the curse of dimensionality primarily impacts the approximation of the singular functional $\delta_x^\dag$ by $f_{x}^\dag$. Structural restrictions (e.g., additivity or low-dimensional interaction kernels) are commonly imposed to mitigate this. While such constraints modify the explicit form of $f_{x}^\dag$, they preserve the multiplicative structure of the influence function in (ref), which is dictated solely by the underlying identification condition in (ref).

Alternative approximation schemes (e.g. local polynomials, neural networks, or other machine learning estimators) generate distinct weighting functions $f_{x}^\dag$ that may depend on $P$ in a more complex manner. Even when closed-form expressions are unavailable, the influence function intuitively should retain the canonical form $\dot{\gamma}_x^{\texttt{B}} \!=\! f_{x}^\dag \mspace{1mu} \dot{\gamma}(P_{(Y,X|W)})$, where one may need to use the implicit function theorem to compute $\dot{\gamma}(P_{(Y,X|W)})$ as in Example (ref). Consequently, deriving the influence function remains a systematic application of the chain rule and the above expectation-substitution rule. Quantifying the associated nonparametric bias, which requires precise control over the approximation error of $\delta_x^\dag$, involves technically demanding mathematical analysis and goes beyond the scope of the current paper.

Semiparametric Multi-Step Plug-in Estimation

Lemma (ref) provides a direct route to deriving influence functions in semiparametric settings. We first establish the result for directly identified parameters.

theoremAssumption (ref) holds true. Let $\beta$ be a regular finite-dimensional parameter identified via \begin{align} \beta(P) = P\big[h_\beta\big(\gamma_1(P), \gamma_2(P_{(:|{1})}), \ldots, \gamma_l(P_{(:|{l-1})}) \big)\big], \end{align} where $h_\beta$ is a known function and each $\gamma_j$ is an unknown functional of $Z^{(j-1:1)}$. For $j'<j$, the function $\gamma_{j}$ could be an input to $\gamma_{j'}$. The influence function of $\beta$ is (recall $P_{(:|{j-1})} \!\coloneqq\! P_{(l:j|j-1:1)}$): \begin{align} \dot{\beta} = h_\beta + \sum_{j=1}^{l} P_{(:|{j-1})} [ \partial_{\gamma_j} h_\beta ] \times \dot{\gamma}_j, \end{align} where $\partial_{\gamma_j} h_\beta$ could involve generalized derivative(s). If each $\gamma_j$ is directly identified as $\gamma_j(P_{(:|{j-1})}) \!=\! P_{(:|{j-1})}[h_{\gamma,j}(Z^{(j:1)})]$ for a known function $h_{\gamma,j}$, then $\dot{\gamma}_j \!=\! h_{\gamma,j} - P_{(:|{j-1})}[h_{\gamma,j}]$. If $\gamma_j$ is indirectly identified via a known moment function $m_{\gamma,j}$ satisfying $P_{(:|{j-1})}[m_{\gamma,j}(Z, \gamma_j)] \!=\! 0$, then (following Example (ref)) \begin{align} \begin{split} \dot{\gamma}_j &= - \big( P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}]^\intercal \mspace{1mu} \mathcal{W}_j \mspace{1mu} P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}] \big)^{-1} P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}]^\intercal \mspace{1mu} \mathcal{W}_j \mspace{1mu} m_{\gamma,j} \\ &\eqqcolon - P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}]_{L}^{-1} \mspace{1mu} m_{\gamma,j}, \end{split} \end{align} where $\mathcal{W}_j$ is a symmetric positive semi-definite matrix function of $Z^{(j-1:1)}$.

The derived expression depends exclusively on the influence functions $\dot{\gamma}_j$ rather than their regularized approximations $\dot{\gamma}_j^{\texttt{B}}$. Consequently, the influence function $\dot{\beta}$ is invariant to the specific nonparametric estimator employed for the nuisance parameters. This invariance extends to the indirect identification case (Theorem (ref)) and aligns with the classic result Newey:1994, ACHL:2014: provided nuisance estimators are consistent and converge at appropriate rates, the asymptotic variance of $\hat{\beta}$ remains unchanged regardless of the estimation method for $\gamma$. This property also clarifies the choice of our examples in the previous subsection. While complex nonparametric or machine learning estimators affect higher-order remainders and bias terms, they do not alter the first-order influence function. A detailed analysis of these higher-order properties lies outside the scope of this paper.

example[Propensity score and conditional outcomes] In treatment effect analysis, the propensity score is directly identified as $\pi(X) \!\coloneqq\! P_{(T|X)}[T]$. The conditional potential outcomes are defined as \begin{align*} \tau_1 \coloneqq P_{(Y(1) | X)}[Y(1)] \quadand\quad \tau_0 \coloneqq P_{(Y(0) | X)}[Y(0)]. \end{align*} Under unconfoundedness, these reduce to observed conditional expectations: \begin{align*} \tau_1 = P_{(Y|T=1, X)}[Y] = \frac{P_{(Y,T|X)}[TY]}{P_{(T|X)}[T]} \quad \tau_0 = P_{(Y|T=0, X)}[Y] = \frac{P_{(Y,T|X)}[(1-T)Y]}{P_{(T|X)}[1-T]}. \end{align*} Both parameters can thus be treated as directly identified, but the fact that $P_{(T|X)}$ is part of $P_{(Y,T|X)}$ complicates the calculation (details can be found in the supplement). An easier way is to recognize that $\tau_1$ satisfies $P_{(Y,T|X)}[TY - T \tau_1] = 0$, for example. The corresponding moment functions are \begin{align} m_\pi = T - \pi(X), \quad m_{\tau_1} = TY - T \tau_1(X) \quad m_{\tau_0} = (1-T)(Y - \tau_0(X)). \end{align} Under standard overlap conditions, the following derivatives are almost surely invertible: \begin{align*} P_{(T|X)}[\partial_\pi m_\pi] = -1, \quad P_{(Y,T|X)}[\partial_{\tau_1} m_{\tau_1}] = \pi(X), \quad P_{(Y,T|X)}[\partial_{\tau_0} m_{\tau_0}] = 1- \pi(X). \end{align*} Applying (ref) yields \begin{align} \dot{\pi} = T - \pi(X), \quad \dot{\tau}_1 = \frac{T}{\pi(X)} (Y - \tau_1(X)), \quad \dot{\tau}_0 = \frac{1-T}{1- \pi(X)} (Y - \tau_0(X)). \end{align}
example[Average Treatment Effect (ATE)] Under unconfoundedness, the identifying functionals for the inverse probability weighting (IPW) and regression-based estimators are \begin{align*} h_{\mathrel{\scriptscriptstyle ATE}}^{\mathrel{\scriptscriptstyle IPW}} = \frac{TY}{\pi(X)} - \frac{(1-T)Y}{1-\pi(X)} \,\,\,and\,\,\, h_{\mathrel{\scriptscriptstyle ATE}}^{\mathrel{\scriptscriptstyle Reg}} = \tau_1(X) - \tau_0(X). \end{align*} Direct calculation shows that $P_{(Y,T|X)}[\partial_{\tau_1} h_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}}] \!=\! 1$, $P_{(Y,T|X)}[\partial_{\tau_0} h_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}}] \!=\! -1$, and \begin{align*} P_{(Y,T|X)} [\partial_\pi h_{\mathrel{\scriptscriptstyle ATE}}^{\mathrel{\scriptscriptstyle \texttt{IPW}}}] = - \frac{\tau_1(X)}{\pi(X)} - \frac{\tau_0(X)}{1-\pi(X)}. \end{align*} Substituting these into (ref) and using (ref) yields \begin{align*} \dot{\tau}_{\text{IPW}}=\dot{\tau}_{\text{Reg}}= \dot{\tau}_1 - \dot{\tau}_0 + \tau_1(X) - \tau_0(X) - \tau. \end{align*} This expression coincides with the influence function of the doubly robust (AIPW) estimator.

Existing approaches for semiparametric moment models Ai&Chen:2003, Ichimura&Newey:2022 typically derive influence functions by solving abstract projection or minimum-distance/least-squares problems Ai&Chen:2003 Ichimura&Newey:2022. Our differentiation-based framework bypasses these steps, yielding closed-form expressions without explicit tangent-space characterizations. The following theorem extends the result to indirect identification.

theoremSuppose the conditions of Theorem (ref) hold, except that $\beta$ satisfies the unconditional moment condition \begin{align} P\big[m_\beta\big(\beta, \gamma_1(P), \gamma_2(P_{(:|{1})}), \ldots, \gamma_l(P_{(:|{l-1})}) \big)\big] = 0, \end{align} and let $\mathcal{W}_{\beta\beta}$ be a conformable symmetric positive semi-definite matrix. The influence function of $\beta$ is \begin{align} \dot{\beta} = \big\{ P[\partial_\beta m_\beta]^\intercal \mathcal{W}_{\beta\beta} \, P[\partial_\beta m_\beta] \big\}^{-1} P[\partial_\beta m_\beta]^\intercal \mathcal{W}_{\beta\beta} \Big( m_\beta + \sum_{j=1}^l P_{(:|{j-1})} [\partial_{\gamma_j} m_\beta ] \dot{\gamma}_j \Big). \end{align}

The expression enclosed in parentheses in (ref) is precisely the locally robust moment function LocalRobust:2022:

align*[align* omitted — 154 chars of source]

The asymptotic variance of $\dot{\beta}$ depends jointly on the nuisance weighting functions $\mathcal{W}_j$ in (ref) and the target weighting matrix $\mathcal{W}_{\beta\beta}$. This raises a central efficiency question: does the choice of $\mathcal{W}_j$ that optimally estimates the nuisance parameters $\gamma$ also minimize the asymptotic variance of $\dot{\beta}$?

The term $v_\rho$ in Ichimura&Newey:2022 corresponds to $P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}]$ in (ref). However, their least-squares formulation relies on the ratio $-v_{m}(X)/v_\rho(X)$, which suggests that $v_\rho$ is scalar-valued. Consequently, the selection of an optimal nuisance weighting matrix $\mathcal{W}_j$ does not arise in their framework. We characterize the optimal choice of $\mathcal{W}_j$ in the general matrix-valued case in Section (ref), where we will construct a weakly better moment function to plug-in.

Local Robustness and Efficiency

Local robustness

Early robustness literature primarily addressed sensitivity to outliers Tukey:1960. With the development of sequential plug-in estimation, focus shifted toward robustness against specification errors and first-step estimation errors in nuisance parameters. Key concepts include double robustness Robins&Rotnitzky:2001Comment and local robustness LocalRobust:2022. While double robustness relies on second-order influence functions, local robustness is a first-order property. This paper focuses on local robustness, leaving second-order analysis for future work.

definition[Local Robustness to First-Step Parameter] Let $\beta$ be a regular parameter dependent on a nuisance parameter $\gamma$. Suppose $\beta$ is identified by the moment condition $P[m_\beta(\beta, \gamma(P))] \!=\! 0$. We say the moment function $m_\beta$ is locally robust to the first-step parameter $\gamma$ if \begin{align*} \lim_{\epsilon \to 0} \frac{1}{\epsilon} P\left[ m_\beta\big(\beta, \gamma(P + \epsilon(Q - P))\big) - m_\beta\big(\beta, \gamma(P)\big) \right] = 0 \end{align*} for all admissible directions $Q - P \in \mathcal{S}$. For brevity, we may simply state that $m_\beta$ is locally robust. This definition extends naturally to conditional probability measures defining $\beta$ and $\gamma$, respectively.

Here, we treat local robustness as a property of the moment function $m_\beta$ rather than the parameter $\beta$ itself. This distinction is necessary because $\beta$ may admit multiple identifying moment conditions (e.g., ATE), each with distinct sensitivity to nuisance estimation error. The following theorem provides necessary and sufficient conditions for local robustness under our framework.

theoremSuppose that $m_\beta$ is continuously differentiable with respect to $\gamma$, which is itself a regular parameter. (i) If $\gamma$ is identified under $P$, the moment function $m_\beta$ is locally robust to $\gamma$ if and only if (by Lemma (ref)) \begin{align*} P[ (\partial_\gamma m_\beta) \mspace{1mu} \dot{\gamma}(P; Q - P) ] = 0 \,\, \forall Q - P \in \mathcal{S} \,\, \Longleftrightarrow \,\, P[\partial_\gamma m_\beta] \dot{\gamma}(P) = 0 \,\, \Longleftrightarrow \,\, P[\partial_\gamma m_\beta] = 0. \end{align*} (ii) If $\gamma \!=\! (\gamma_1^\intercal, \ldots, \gamma_l^\intercal)^\intercal$ and each $\gamma_j$ is identified under $P_{(:|{j-1})}$, then $m_\beta$ is locally robust to $\gamma_j$ if and only if (by Lemma (ref) and Theorem (ref)) \begin{align*} P[(\partial_{\gamma_j} m_\beta) \mspace{1mu} \dot{\gamma}_j( P_{(:|{j-1})}; Q_{(:|{j-1})} - P_{(:|{j-1})})] = 0 \,\, \forall (Q_{(:|{j-1})} - P_{(:|{j-1})}) \in \mathcal{S}, \end{align*} which is equivalent to $P_{(:|{j-1})}[\partial_{\gamma_j} m_\beta] \!=\! 0$ with probability one.

Case (i) is a special case of Case (ii) with $l\!=\!1$ and $\gamma\!=\!\gamma_1$ identified under $P_{(:|{0})} \!=\! P$. For $j\!\geq\! 2$, the condition $P_{(:|{j-1})}[\partial_{\gamma_j} m_\beta] \!=\! 0$ is strictly stronger than $P[\partial_{\gamma_j} m_\beta] \!=\! 0$, as discussed in the ATT example below.

Existing approaches for sequential plug-in models typically derive the locally robust moment function by solving projection or least-squares problems Hahn&Ridder:2013, Ichimura&Newey:2022. Hahn&Ridder:2013 adopt a generated regressors perspective, yielding an adjustment term that often involves second-order derivatives via integration by parts. Ichimura&Newey:2022 propose a least-squares projection alternative. Our differentiation-based framework bypasses these constructions, yielding the adjustment term directly.

The results in Section (ref) show that the locally robust moment function corresponds to the term in parentheses in (ref) (recalling (ref)):

align[align omitted — 290 chars of source]

Note that $\eta_j$ may depend on $\gamma_j$ and additional nuisance components. By construction, $m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$ is locally robust to $\eta_j$:

align*[align* omitted — 225 chars of source]

Assuming $\eta_{j'}$ ($j'\geq j$) may depend on $\gamma_{j}$, applying the law of iterated expectations yields

align*[align* omitted — 481 chars of source]

This confirms that $m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$ is locally robust to all $\gamma_j$ and the auxiliary parameters $\eta_j$. The adjustment term $\sum_{j=1}^l \eta_j m_{\gamma,j}$ coincides with the first-order correction derived in Newey:1994 and LocalRobust:2022.

exampleFor Case (i) in Theorem (ref), both the variance $\sigma^2 \!\coloneqq\! P[(Z- P[Z])^2]$ and the third central moment $\mu_3 \!\coloneq\! P[(Z- \mu_1)^3]$ depend on the mean $\mu_1\!\coloneqq\! P[Z]$. The variance is locally robust to the mean, while $\mu_3$ is not. For Case (ii), the moment functions for the IPW and regression-based estimators of the ATE are \begin{align*} m_{\mathrel{\scriptscriptstyle ATE}}^{\mathrel{\scriptscriptstyle IPW}} = \frac{TY}{\pi(X)} - \frac{(1-T)Y}{1-\pi(X)} - \tau_{\mathrel{\scriptscriptstyle ATE}} \quadand\quad m_{\mathrel{\scriptscriptstyle ATE}}^{\mathrel{\scriptscriptstyle Reg}} = \tau_1(X) - \tau_0(X) - \tau_{\mathrel{\scriptscriptstyle \texttt{ATE}}}. \end{align*} Neither satisfies the local robustness condition: \begin{align*} P_{(Y,T|X)}[\partial_{\pi} m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{IPW}}}] = - \frac{\tau_1(X)}{\pi(X)} - \frac{\tau_0(X)}{1-\pi(X)} \neq 0 \quad P_{(Y,T|X)}[\partial_{\tau_t} m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}}] = (-1)^{t+1}. \end{align*} In contrast, the AIPW moment function is locally robust. This illustrates why local robustness is a property of the identifying moment condition rather than the parameter itself: different strategies for the same parameter may exhibit distinct sensitivity to first-step estimation error.

When $m_\beta \!\neq\! m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$, researchers may choose between the raw and locally robust identification strategies. The corresponding optimal asymptotic variances are:

align*[align* omitted — 496 chars of source]

These variances coincide under the standard condition that the nuisance moment functions are locally robust to $\beta$, i.e., $P_{(:|{j-1})}[\partial_\beta m_{\gamma, j}] \!=\! 0$ for all $j$. This condition is mild: in most semiparametric two-step models, $m_{\gamma,j}$ does not depend on $\beta$ at all. Otherwise, a consistent estimator of $\beta$ would be needed to estimate $\gamma_j$, leading to a circular dependency. Given this condition, (ref) implies $P[\partial_\beta m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta] \!=\! P[\partial_\beta m_\beta]$, as the derivative of the adjustment term with respect to $\beta$ has zero expectation. Consequently, using $m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$ does not lead to a more efficient estimator. Its primary role is to mitigate first-order bias arising from nuisance estimation error LocalRobust:2022.

Adaptive estimation

Bickel:1982Adaptive formalizes the theory of adaptive estimation, addressing Stein's classical question: “when can a Euclidean parameter be estimated as efficiently without knowledge of an infinite-dimensional nuisance parameter as with full knowledge of it?” Stein:1956.

Within the moment-based framework, an affirmative answer for a given identification strategy corresponds to the following property, which, to the best of our knowledge, has not been previously highlighted in the literature and plays a crucial role in the subsequent discussion.

definition[Plug-in Neutral Variance] Let $m_\beta$ be a moment function for $\beta$ that depends on nuisance parameter $\gamma$. Compare two estimation scenarios: (i) the infeasible oracle case, where the true $\gamma$ is known and plugged into $m_\beta$, and (ii) the feasible plug-in case, where a consistent estimator $\hat{\gamma}$ is used. We say $m_\beta$ has a plug-in neutral variance if the optimal asymptotic variance of $\hat{\beta}$ is identical in both scenarios.
propositionIf a moment function $m_\beta$ is locally robust with respect to all nuisance parameters on which it depends, then it has a plug-in neutral variance. When $P[\partial_\beta m_\beta]$, $\langle m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta, (m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta)^\intercal \rangle$, and $\langle m_\beta, m_\beta^\intercal \rangle$ are all invertible, if $P_{(:|{j-1})}[m_{\gamma,j} \mspace{1mu} m_\beta^\intercal] \!=\! 0$ and $P_{(:|{j\vee j'-1})}[m_{\gamma,j} \mspace{1mu} m_{\gamma,j'}^\intercal] \!=\! 0$ for all $j\!\neq\! j'$, then $m_\beta$ having a plug-in neutral variance implies it is also locally robust.

When the true $\gamma$ is known, estimation reduces to the parametric prototype in Example (ref) with $\nu \!=\! \beta$. The feasible plug-in case follows from Theorem (ref). Their respective optimal asymptotic variances are

align*[align* omitted — 351 chars of source]

The sufficiency direction follows immediately: local robustness implies $m_\beta \!=\! m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$, so the two variances coincide.

To examine necessity, assume the relevant covariance matrices are invertible. Using $\eta_j$ introduced in (ref), the difference in the covariance kernels is

align*[align* omitted — 520 chars of source]

Under the orthogonality conditions noted in the proposition, the expression simplifies to

align*[align* omitted — 286 chars of source]

Equality of the variances requires this difference to vanish, which implies $\eta_j \!=\! 0$ almost surely, and consequently $P_{(:|{j-1})} [\partial_{\gamma_j} m_\beta ] \!=\! 0$ with probability one. While necessity may hold under weaker orthogonality conditions, a complete characterization is left for future work.

exampleFor the IPW estimator, it is well documented that estimating the propensity score yields a smaller asymptotic variance than plugging in the true $\pi(X)$. Because this specification is directly identified, the invertibility conditions in Proposition (ref) hold. Consequently, the variance reduction is equivalent to the absence of local robustness for the IPW estimator, a well-known result in treatment effect literature. For the regression-based estimator, if the conditional outcome functions $\tau_1(X)$ and $\tau_0(X)$ were known, the asymptotic variance would reduce to $\text{Var}(\tau_1(X) - \tau_0(X) - \tau_{\mathrel{\scriptscriptstyle \texttt{ATE}}})$. This quantity is strictly smaller than $\text{Var}(Y(1) - Y(0))$, the variance that would obtain if both potential outcomes were directly observable for each unit. This strict inequality aligns with the fact that the regression-based moment function is not local robustness. It also raises a question on what information an oracle is allowed to access. We will delineate oracle's admissible information structure in Section (ref). In contrast, the AIPW moment function is locally robust and therefore exhibits plug-in neutral variance.

Example (ref) highlights an interesting scenario: one may have more than one moment function to identify $\beta$. Furthermore, even when there is only one $m_\beta$ to use when one needs to estimate $\gamma$, one can have multiple $\tilde{m}_\beta \!\coloneqq\! m_\beta + B m_\gamma$ in the hypothetical case where $\gamma$ were known. This leads to the oracle moment funciton that will be discussed in the next subsection.

Oracle Moment Functions and Plug-in Efficiency

ACHL:2014 raise the question of whether sequential two-step estimation achieves the same efficiency as joint estimation. They note that, although the plug-in method often offers substantial computational advantages, it may suffer from “limited information,” in the sense that the moment conditions are not jointly exploited. By contrast, the joint procedure fully utilizes all available moment restrictions simultaneously and may therefore deliver efficiency gains. However, it typically requires a high-dimensional nonlinear search over both finite- and infinite-dimensional parameter spaces.

ACHL:2014 derive the influence function for the plug-in estimator under exact identification of the nuisance parameter and note a methodological distinction from the orthogonalization approach of Ai&Chen:2012. Specifically, ACHL:2014 construct the adjustment term using derivatives of the moment function, whereas Ai&Chen:2012 employ a variance-covariance projection. However, they do not comment on whether the two constructions coincide (see Section 3.4 of ACHL:2014).

To compare plug-in and joint procedures rigorously, we first identify the optimal moment function for plug-in estimation. The locally robust moment function $m_{\beta}^{\mathrel{\scriptstyle \texttt{LR}}}$ is a natural candidate. When all nuisance parameters are exactly identified, the weighting matrices $\mathcal{W}_j$ do not affect the asymptotic variance, leaving $m_{\beta}^{\mathrel{\scriptstyle \texttt{LR}}}$ as the only choice.

When nuisance parameters are over-identified, however, the optimal choice of $\mathcal{W}_j$ becomes nontrivial. Besides, it is theoretically possible that another moment function may weakly dominate $m_{\beta}^{\mathrel{\scriptstyle \texttt{LR}}}$. Motivated by the structure of joint estimation, we characterize this optimal specification through an oracle perspective.

To clarify the construction, first consider the case where $\beta$ and $\gamma$ are identified under the same distribution $P$. If $\gamma$ were known, any moment function of the form $m_{\beta} + B m_{\gamma}$ identifies $\beta$. The optimal oracle moment function is obtained by projecting $m_\beta$ onto the orthogonal complement of the nuisance moment space:

align*[align* omitted — 272 chars of source]

The linear family $m_\beta + B m_\gamma$ is assumed to span the space of all admissible moment conditions for $\beta$. The rationale is that any zero-mean transformation of the existing moments should already be included in them.

In the general setting, $\beta$ and the nuisance components $\gamma_j$ are identified under distinct probability measures. The following orthogonal decomposition for any $f \in \mathcal{L}_2(P)$ proves instrumental:

align*[align* omitted — 166 chars of source]

Each component $f_j$ depends on the data only through $Z^{(j:1)}$ and satisfies $f_j \!\in\! \mathcal{L}_2^0(P_{({j}|:)})$, where $P_{({j}|:)} \!\coloneqq\! P_{(j|j-1:1)}$. Applying this decomposition to $m_\beta$ yields $m_{\beta,0} \!=\! 0$.

To streamline the analysis, we assume that $m_{\beta}$ does not lie entirely within any single subspace $\mathcal{L}_2^0(P_{({j}|:)})$ and that the ordering of $Z$ and $\gamma$ can be arranged so that each $m_{\gamma,j} \!\in\! \mathcal{L}_2^0(P_{({j}|:)})$ for $j\!=\!1,\ldots,l$. Standard IPW and AIPW specifications satisfy this structure, whereas the regression-based estimator does not (see Section (ref)). Under this arrangement, $m_{\gamma,j}$ and $m_{\gamma,j'}$ are orthogonal for $j\!\neq\! j'$, and $m_\beta$ can correlate with $m_{\gamma,j}$ only through its component $m_{\beta,j}$.

Since all second-moment matrices are positive semi-definite, $V_{\beta\gamma,j} (I - V_{\gamma\gamma,j}^+ V_{\gamma\gamma,j} ) \!=\! 0$ almost surely (refer to Theorem 16.1 of Gallier:2011), where

align*[align* omitted — 297 chars of source]

This permits an orthogonal projection within each conditional subspace:

align*[align* omitted — 266 chars of source]

Aggregating across components yields the optimal oracle moment function for $\beta$:

align*[align* omitted — 210 chars of source]

Unless $V_{\beta\gamma,j} \!=\! 0$ almost surely, the oracle moment function $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ differs from $m_{\beta,j}$. This yields two candidates for plug-in estimation. We construct and compare the locally robust versions of both:

align*[align* omitted — 394 chars of source]

Define the Jacobian components (recall that $P_{({j}|:)}[\partial_{\beta} m_{\gamma,j}] \!=\! 0$ is required for valid plug-in estimation):

align*[align* omitted — 348 chars of source]

The terms $\Delta_j$ and $P_{({j}|:)}[\partial_{\gamma_j} m_{\beta,j}]$ may introduce auxiliary nuisance parameters, but the resulting moment functions remain locally robust to them because $P_{({j}|:)}[\dot{\gamma}_j] \!=\! 0$ by construction.

With the notation, the influence funciton $\dot{\gamma}_j $ given in (ref) can be written as

align*[align* omitted — 145 chars of source]

The conditional variances of $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ and $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{LR}}}$ depend explicitly on $\mathcal{W}_j$. Denote these by $\text{Var}_j^{\mathrel{\scriptstyle \texttt{orc-LR}}}(\mathcal{W}_j)$ and $\text{Var}_j^{\mathrel{\scriptstyle \texttt{LR}}}(\mathcal{W}_j)$, respectively.

theorem[Optimal Plug-in Moment Functions] Suppose that $m_{\gamma,j} \!\in\! \mathcal{L}_2^0(P_{({j}|:)})$ for $j\!=\!1,\ldots,l$. Assume $P_{({j}|:)}[\partial_{\beta_j} m_{\gamma,j}] \!=\! 0$ to ensure valid plug-in estimation, and suppose $(I - V_{\gamma\gamma,j} V_{\gamma\gamma,j}^+) J_{\gamma,j} \!=\! 0$ almost surely. Then \begin{align*} \min_{\mathcal{W}_j} Var_j^{\mathrel{ LR}}(\mathcal{W}_j) \succeq \min_{\mathcal{W}_j} Var_j^{\mathrel{ orc-LR}}(\mathcal{W}_j) = S_{\beta\beta,j} + \Delta_j (J_{\gamma,j}^\intercal V_{\gamma\gamma,j}^+ J_{\gamma,j})^{-1} \Delta_j^\intercal, \end{align*} where $S_{\beta\beta,j} \!\coloneqq\! \langle m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}, (m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j})^\intercal \rangle_j$ and $A \!\succeq\! B$ denotes that $A-B$ is positive semi-definite. Consequently, the aggregated moment function \begin{align} m_\beta^{\mathrel{ orc-LR}} = \sum_{j=1}^l m_{\beta,j}^{\mathrel{ orc-LR}} = m^{\mathrel{ \texttt{orc}}}_\beta - \sum_{j=1}^l \Delta_j (J_{\gamma,j}^\intercal V_{\gamma\gamma,j}^+ J_{\gamma,j})^{-1} J_{\gamma,j}^\intercal V_{\gamma\gamma,j}^+ m_{\gamma,j} \end{align} weakly dominates the optimally weighted locally robust moment function $m_\beta^{\mathrel{\scriptstyle \texttt{LR}}}$.

The condition $(I - V_{\gamma\gamma,j} V_{\gamma\gamma,j}^+) J_{\gamma,j} \!=\! 0$ is stronger than the invertibility of $J_{\gamma,j}^\intercal V_{\gamma\gamma,j}^+ J_{\gamma,j}$ required in Assumption (ref). Geometrically, it requires that the column space of $J_{\gamma,j}$ lies entirely within the column space of $V_{\gamma\gamma,j}$, rather than merely having no column contained in the nullspace. This condition guarantees that the natural weighting choice $\mathcal{W}_j \!=\! V_{\gamma\gamma,j}^+$ is optimal for estimating $\gamma_j$, making the assumption mild in standard over-identified settings.

Because $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc}}}$ is constructed to be orthogonal to $\dot{\gamma}_j$, the same weighting matrix $\mathcal{W}_j \!=\! V_{\gamma\gamma,j}^+$ minimizes $\text{Var}_j^{\mathrel{\scriptstyle \texttt{orc-LR}}}(\mathcal{W}_j)$. The optimal weighting for $\text{Var}_j^{\mathrel{\scriptstyle \texttt{LR}}}(\mathcal{W}_j)$ may differ from $V_{\gamma\gamma,j}^+$ (see the supplement for techinical details). Theorem (ref) implies that $m_\beta^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ achieves the minimal attainable variance among all valid plug-in moment specifications, and thus serves as the appropriate benchmark for comparing plug-in and joint estimation efficiency.

Efficiency: Joint versus sequential plug-in

The orthogonal decomposition of $m_{\beta}$ facilitates a direct comparison between joint and sequential plug-in estimation. Specifically, the condition $P_{({j}|:)}[m_{\beta,j}] \!=\! 0$ identifies a parameter $\beta_j$, which we term a companion parameter. When $j\!>\!1$, this parameter is infinite-dimensional.

Before proceeding to the general analysis, we exclude two degenerate configurations:

align*[align* omitted — 145 chars of source]

The first implies $\beta_j \!=\! 0$; the second implies $\beta_j$ is a known linear transformation of $\gamma_j$. In both instances, there is no distinction between joint and plug-in estimation. Both configurations arise in the AIPW estimator for the ATE (see Section (ref) for details).

In the non-degenerate case, we group $\beta_j$ and $\gamma_j$ into a joint parameter $\nu_j$ and define the stacked moment vector and its conditional covariance:

align*[align* omitted — 528 chars of source]

Assuming $(I - V_{\nu\nu,j} V_{\nu\nu,j}^+) P_{({j}|:)}[\partial_\nu m_j] \!=\! 0$, the asymptotic efficiency of the joint estimator is governed by the optimal variance matrix

align*[align* omitted — 152 chars of source]

Under the identifying restriction $P_{({j}|:)}[\partial_{\beta_j} m_{\gamma,j}] \!=\! 0$, the covariance matrix $V_{\nu\nu,j}$ admits a block decomposition (cf. Gallier:2011):

align*[align* omitted — 334 chars of source]

where $S_{\beta\beta,j} \!\coloneqq\! V_{\beta\beta,j} - V_{\beta\gamma,j} V_{\gamma\gamma,j}^+ V_{\gamma\beta,j}$ is the Schur complement of $V_{\gamma\gamma,j}$ in $V_{\nu\nu,j}$. Notably, $S_{\beta\beta,j} \!=\! \langle m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}, (m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j})^\intercal \rangle_j$, linking this decomposition directly to the oracle construction in Theorem (ref).

Standard results on the Moore--Penrose inverse (e.g., MPinverse:2023) imply that if $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) V_{\beta\gamma,j} \!=\! 0$ almost surely, then $V_{\nu\nu,j}^+$ admits the closed-form block structure:

align*[align* omitted — 343 chars of source]

Substituting this into the variance formula yields

align*[align* omitted — 441 chars of source]

This representation establishes that joint estimation is asymptotically equivalent to using the oracle moment $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc}}}$ and the nuisance moment $m_{\gamma,j}$ simultaneously.

To streamline the efficiency comparison, define the following matrix components:

gather*[gather* omitted — 716 chars of source]

Here, $\Sigma_{\beta\beta,j}^{-1}$ represents the optimal oracle variance for $\beta_j$ when $\gamma_j$ is known, while $\tilde{\Sigma}_{\gamma\gamma,j}^{-1}$ is the optimal variance for estimating $\gamma_j$ alone when $(I - V_{\gamma\gamma,j} V_{\gamma\gamma,j}^+) J_{\gamma,j} \!=\! 0$. The local robustness of $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ is central to the following characterization.

theorem[Joint Variance] Suppose both $m_{\beta,j}$ and $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ are non-zero, so that $P_{({j}|:)}[m_{\beta,j}] \!=\! 0$ identifies a companion parameter $\beta_j$ that is neither degenerate nor a conditional linear transformation of $\gamma_j$. Suppose $(I - V_{\nu\nu,j} V_{\nu\nu,j}^+) P_{({j}|:)}[\partial_\nu m_j] \!=\! 0$. Assume $\Omega_{\beta\beta,j}$ and $\Omega_{\gamma\gamma,j}$ are invertible. If $P_{({j}|:)}[\partial_{\beta_j} m_{\gamma,j}] \!=\! 0$ and $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) V_{\beta\gamma,j} \!=\! 0$ almost surely, then the asymptotic variance of the joint estimator $\hat{\nu}_j$ admits the block representation \begin{align*} \Sigma_{\nu\nu,j}^{-1} = \left( \begin{matrix} \Sigma_{\beta\beta,j} & \Sigma_{\beta\gamma,j} \\ \Sigma_{\beta\gamma,j}^\intercal & \Sigma_{\gamma\gamma,j} \end{matrix} \right)^{-1} = \left( \begin{matrix} \Omega_{\beta\beta,j}^{-1} & - \Sigma_{\beta\beta,j}^{-1} \Sigma_{\beta\gamma,j} \Omega_{\gamma\gamma,j}^{-1} \\ - \Sigma_{\gamma\gamma,j}^{-1} \Sigma_{\beta\gamma,j}^\intercal \Omega_{\beta\beta,j}^{-1} & \Omega_{\gamma\gamma,j}^{-1} \end{matrix} \right), \end{align*} where the symmetry condition $\Sigma_{\beta\beta,j}^{-1} \Sigma_{\beta\gamma,j} \Omega_{\gamma\gamma,j}^{-1} \!=\! \Omega_{\beta\beta,j}^{-1} \Sigma_{\beta\gamma,j} \Sigma_{\gamma\gamma,j}^{-1}$ holds by construction. Furthermore, $\Sigma_{\nu\nu,j}^{-1}$ reduces to the block diagonal form $\textrm{diag}( \Sigma_{\beta\beta,j}^{-1}, \tilde{\Sigma}_{\gamma\gamma,j}^{-1})$ if and only if $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ is locally robust, i.e., $\Delta_j \!=\! 0$.

Direct inspection yields the matrix inequality

align*[align* omitted — 369 chars of source]

Theorem (ref) implies that the optimal oracle variance $\Sigma_{\beta\beta,j}^{-1}$ is attainable by joint estimation only when $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ is locally robust ($\Delta_j \!=\! 0$). Under this same condition, Theorem (ref) guarantees that sequential plug-in estimation also achieves the optimal oracle variance. This characterizes the ideal adaptive setting Bickel:1982Adaptive, in which both procedures attain the same asymptotic variance as the infeasible oracle that knows the nuisance parameter.

The remaining question is whether plug-in and joint estimation yield identical asymptotic variance when $\Delta_j \!\neq\! 0$. Plug-in estimation of $\beta_j$ using $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ produces the variance

align*[align* omitted — 157 chars of source]

Comparing this to the joint variance $\Omega_{\beta\beta,j}^{-1}$:

align*[align* omitted — 216 chars of source]

If $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) \Delta_j \!=\! 0$, standard Moore--Penrose identities verify that

align*[align* omitted — 225 chars of source]
theorem[Joint versus plug-in] Suppose the assumptions of Theorem (ref) hold, so that the comparison may be restricted to plug-in estimation based on $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$. If (i) $m_{\beta,j} \!=\! 0$, (ii) $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j} \!=\! 0$, or (iii) $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) V_{\beta\gamma,j} \!=\! 0$ and $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) \Delta_{j} \!=\! 0$ for each $j$, then the plug-in procedure with $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ attains the same asymptotic variance as the joint procedure. Moreover, both procedures achieve adaptive estimation and the oracle variance if and only if $\Delta_j \!=\! 0$ for all $j$.

The conditions in Theorem (ref) are independent of the identification status of the nuisance parameters, as the optimality of $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ (Theorem (ref)) already accounts for over-identification. When $S_{\beta\beta,j}$ is invertible, the conditions in part (iii) are automatically satisfied. Thus, the invertibility of the Schur complement $S_{\beta\beta,j}$ is the decisive factor for plug-in--joint equivalence, regardless $V_{\gamma\gamma,j}$ is invertible or not.

In practice, researchers may apply the following sequential procedure:

itemize• Compute the orthogonal decomposition of $m_\beta$ when $\beta$ and $\{\gamma_j\}$ are identified under distinct measures. • Verify the conditions of Theorem (ref). Likely scenarios include: $m_{\beta,j} \!=\! 0$, or $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j} \!=\! 0$, or $S_{\beta\beta,j}$ is invertible. • If (2) holds, implement sequential plug-in estimation using $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$, which may reduce to $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{LR}}}$ or $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc}}}$.

One may additionally check whether $\Delta_j \!=\! 0$ to determine if the optimal oracle variance is attainable. Violations are common, as illustrated below.

example[Third-order moment] The third-order moment $\beta \!\coloneqq\! P[(Z-\gamma)^3]$ is not locally robust to the mean $\gamma \!\coloneqq\! P[Z]$. The Schur complement $S_{\beta\beta}$ is invertible, so the rank condition $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) \Delta_{j} \!=\! 0$ holds trivially, ruling out any efficiency gain from joint estimation. Let $\mu_4$ denote the fourth central moment and $\sigma^2$ the variance. Direct calculation yields \begin{align*} m_\beta^{\mathrel{ LR}} = (Z-\gamma)^3 - \beta + 3\sigma^2(Z-\gamma) \,\, and \,\, m_\beta^{\mathrel{ orc}} = (Z-\gamma)^3 - \beta - \frac{\mu_4}{\sigma^2}(Z-\gamma). \end{align*} It is easy to check that $P[\partial_\gamma m^{\texttt{orc}}_\beta] \!=\! (\mu_4 - 3 \sigma^4) / \sigma^2$. Unless the distribution of $Z$ has zero excess kurtosis, $m_\beta^{\mathrel{\scriptstyle \texttt{orc}}}$ is not locally robust, and the oracle variance is unattainable.

Example: Treatment Effect Model

This subsection applies the preceding methodology to two canonical treatment effect estimands: the average treatment effect (ATE), which admits adaptive estimation, and the average treatment effect on the treated (ATT), which does not.

Average Treatment Effect (ATE)

First consider the ATE case. The moment function for the IPW estimator has the following orthogonal decomposition:

align*[align* omitted — 799 chars of source]

The sole nuisance parameter is the propensity score $\pi(X)$, identified by $P_{(T|X)}[m_\pi] \!=\! P_{(T|X)}[T-\pi(X)] \!=\! 0$, which corresponds to $j\!=\!2$. It is easy to see that $m_{\mathrel{\scriptscriptstyle \texttt{ATE}},2} \neq 0$ while its oracle projection vanishes ($m_{\mathrel{\scriptscriptstyle \texttt{ATE}},2}^{\mathrel{\scriptstyle \texttt{orc}}} \!=\! 0$). By Theorem (ref), joint estimation yields no asymptotic efficiency gain over sequential plug-in estimation.

The oracle moment function $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}} = m_{\mathrel{\scriptscriptstyle \texttt{ATE}},3} + m_{\mathrel{\scriptscriptstyle \texttt{ATE}},1}$ satisfies the local robustness condition, confirming that the ATE is adaptively estimable. Because $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\scriptscriptstyle\texttt{IPW}} \!\neq\! m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}}$, plugging the true propensity score into the IPW moment is suboptimal for the oracle.

The regression-based estimator presents a more subtle structure. Its moment function lies entirely in the marginal subspace $\mathcal{L}_2^0(P_{(X)})$, rather than $\mathcal{L}_2^0(P)$:

align*[align* omitted — 230 chars of source]

This aligns with Example (ref), which shows $\text{Var}(m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}}) \!<\! \text{Var}(Y(1) - Y(0))$. The discrepancy arises because $\tau_1(X)$ and $\tau_0(X)$ are companion parameters: they are not of direct interest, but should not be in the oracle's admissible information set.

The reason is best illustrated through a directly identified model $\beta \!=\! P[h_\beta]$. Define the hierarchical projections $h_{\beta,j} \!\coloneqq\! P_{(:|{j})}[h_\beta]$. The induced companion parameters $\beta_j \!=\! P_{({j}|:)}[h_{\beta,j}]$ depend only on $Z^{(j-1:1)}$ and satisfy $\beta \!=\! P[\beta_j]$ for all $j$. Permitting the oracle to observe companion parameters $\beta_j$ ($j\!>\!1$) would artificially deflate the benchmark variance, since $\text{Var}(\beta_j) \!<\! \text{Var}(h_\beta)$. Extending this logic recursively would permit the oracle to observe $\beta_1 \!=\! \beta$, yielding a degenerate zero-variance benchmark. To maintain a well-defined efficiency comparison, the oracle's information set must exclude all companion parameters.

In the regression specification, all nuisance components are companion parameters, placing it in the degenerate regime where $\beta_j \equiv \gamma_j$. Consequently, plug-in and joint estimation are the same, and the locally robust and oracle moment functions coincide: $m_\beta^{\mathrel{\scriptstyle \texttt{LR}}} \equiv m_\beta^{\mathrel{\scriptstyle \texttt{orc}}}$. Thus, the regression-based oracle moment is again $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}} = m_{\mathrel{\scriptscriptstyle \texttt{ATE}},3} + m_{\mathrel{\scriptscriptstyle \texttt{ATE}},1}$.

Lastly, consider the doubly robust AIPW moment function

align*[align* omitted — 442 chars of source]

which is locally robust by construction. The only non-companion nuisance parameter is $\pi(X)$. Because $m_{\mathrel{\scriptscriptstyle \texttt{ATE}},2} \!=\! 0$ in this decomposition, its projection onto the propensity score moment vanishes. Furthermore, because $\tau_1(X)$ and $\tau_0(X)$ are companion parameters ($\beta_j \!\equiv\! \gamma_j$), the oracle construction excludes projection onto their identifying moments $m_{\tau_1}$ and $m_{\tau_0}$ (cf. (ref)). Consequently, the oracle moment function coincides exactly with the AIPW specification: $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}} \!=\! m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{AIPW}}}$.

Average Treatment Effect on the Treated (ATT)

Consider the average treatment effect on the treated. The IPW moment function admits the orthogonal decomposition:

align*[align* omitted — 893 chars of source]

This specification yields the plug-in estimator

align*[align* omitted — 399 chars of source]

Hirano&Imbens&Ridder:2003 construct an alternative ATT estimator by formulating ATT as a weighted ATE. Their moment function differs from $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptscriptstyle \texttt{IPW}}, \scriptscriptstyle T}$ only in the normalization of the target parameter:

align*[align* omitted — 530 chars of source]

The corresponding estimator is

align*[align* omitted — 247 chars of source]

Two regression-based moment functions follow analogously:

align*[align* omitted — 425 chars of source]

Substituting the identifying conditions for $\tau_1(X)$ and $\tau_0(X)$ into $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}, \scriptscriptstyle T}$ and replacing $T$ with $\pi(X)$ recovers $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptscriptstyle \texttt{IPW}}, \scriptscriptstyle \pi}$. The associated plug-in estimators are

align*[align* omitted — 513 chars of source]

The first estimator, $\hat{\tau}_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\scriptscriptstyle \texttt{Reg/T}}$, corresponds to the specification in Hahn:1998.

Unlike the ATE case, the ATT admits distinct locally robust and oracle moment functions:

align*[align* omitted — 912 chars of source]

These specifications generate two augmented estimators:

align*[align* omitted — 606 chars of source]

where the adjustment component is

align*[align* omitted — 266 chars of source]

As established by Hahn:1998, the infeasible efficiency bound with known propensity score is

align*[align* omitted — 387 chars of source]

When $\pi(X)$ is estimated, the asymptotic variance increases to $\text{Var}(m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{LR}}}) / P[T]^2$. The efficiency loss equals

align*[align* omitted — 277 chars of source]

By Theorem (ref), this variance increse occurs precisely because $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{orc}}}$ is not locally robust. Direct computation confirms that

align*[align* omitted — 198 chars of source]

which is non-zero almost surely unless the conditional average treatment effect is constant across treated units. Furthermore, applying the local robustness correction to $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{orc}}}$ recovers $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{LR}}}$, implying $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{orc-LR}}} \!=\! m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{LR}}}$.

Conclusion

In this paper, we introduce a direct differentiation-based framework for deriving influence functions across parametric, nonparametric, and semiparametric models. By grounding the derivation in the functional functional derivative and its integral representation, we demonstrate that the influence function is obtained by centering the identification function, that is, subtracting its conditional or unconditional expectation. This algebraic characterization replaces the traditional reliance on tangent-space projections, score-based Riesz representation arguments, and ad hoc guess-and-verify constructions. We formalize this insight through a unified expectation-substitution rule that seamlessly handles both unconditional and conditional perturbations, yielding closed-form influence functions and revealing a common structural form across finite- and infinite-dimensional settings.

Building on this foundation, we derive transparent conditions for local robustness and clarify its precise role in adaptive estimation. A central contribution is our rigorous efficiency comparison between sequential plug-in and joint estimation. By exploiting the orthogonal decomposition of moment conditions along the hierarchy of conditioning variables, we establish verifiable sufficient conditions under which plug-in estimation attains the same asymptotic variance as joint estimation, even when nuisance parameters are over-identified. Furthermore, we construct an oracle-based locally robust moment function that weakly dominates the local robsut version of the original moment function, providing a principled benchmark for multi-step inference. Applied to treatment effect models, the framework precisely diagnoses when plug-in procedures remain adaptive (as in the ATE) and when they incur efficiency losses due to failure of local robustness (as in the ATT).