Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,596 characters · 22 sections · 61 citation commands
Influence Function: Local Robustness and Efficiency
{1.0\baselineskip} \hyphenpenalty=5000 \tolerance=1000
The influence function is a foundational object in modern statistics and econometrics, characterizing the first-order effect of an infinitesimal perturbation to the data-generating distribution on a target parameter or estimator. It plays a central role in establishing asymptotic linearity, semiparametric efficiency, and local robustness. Seminal work by BKRW:1993 uses it to characterize semiparametric efficiency bounds, while recent literature CCDDHN:2018double,LocalRobust:2022 highlights its role in constructing locally robust moment functions.
Despite its theoretical centrality, deriving influence functions remains technically demanding. Stochastic expansions of the estimation error are often algebraically cumbersome. While tangent-space method tsiatis2006semiparametric is mathematical elegant, it can be abstract and opaque for complex models. As Hahn:1998 notes, the general frameworks of Newey:1990, Newey:1994 and BKRW:1993 may require one to posit a candidate influence function and verify its properties ex post. Hahn&Ridder:2013 extend Newey:1994 to three-step settings but restrict the analysis to differentiable functions. The method by Ichimura&Newey:2022 ultimately require solving a least square projection problem, where the nuisance moment function is scalar-valued. A unified, explicit, and calculus-driven recipe remains absent.
This work establishes a cohesive analytical framework, grounded in functional differentiation, for the unified derivation of influence functions across parametric and nonparametric specifications. We define the target parameter as a statistical functional of the probability distribution $P$ and compute its functional derivative under local perturbations of the form $P^\epsilon \!=\! (1-\epsilon)P + \epsilon Q$. For regular parameters, we show that the Riesz representer of this derivative is obtained by orthogonally projecting the identification function onto the subspace of mean-zero square-integrable functions. Crucially, this projection reduces to a simple centering operation: subtracting the unconditional or conditional expectation from the identification function (see Example (ref)).
The framework extends seamlessly to infinite-dimensional parameters. While the convergence rates differ fundamentally between finite- and infinite-dimensional settings, we demonstrate that the influence functions share a common algebraic structure: in both cases, they arise as linear transformations of the moment function (see Section (ref)). The transformation operator, rather than the moment function, governs the convergence rate. To operationalize this insight, we provide multiple application strategies and distill the derivation into a set of transparent calculus rules (Lemmas (ref), (ref), and (ref)).
In semiparametric multi-step plug-in estimation, our approach immediately yields the locally robust moment function. Deriving the associated adjustment term reduces to computing conditional expectations of partial derivatives. Crucially, this procedure does not require classical differentiability. By employing distributional derivatives, the method accommodates non-smooth moment conditions, with the conditional expectation serves as smoothing operator.
Using this framework, we derive easily verifiable conditions for local robustness in both parametric and semiparametric models and clarify its role in adaptive estimation. We revisit the classical question of Stein:1956: under what conditions can a finite-dimensional parameter be estimated as efficiently without knowledge of an infinite-dimensional nuisance parameter as with full knowledge of it? We show that local robustness ensures the asymptotic variance remains invariant to whether the moment function is evaluated at the true or a consistent estimator of the nuisance parameter, providing a direct answer to Stein's question.
We also resolve the classic debate between joint and plug-in estimation. As ACHL:2014 observe, plug-in procedures are computationally tractable but exploits moment conditions sequentially, potentially discarding information. Joint estimation, while theoretically more efficient, typically entails high-dimensional nonlinear optimization over both finite- and infinite-dimensional spaces. We employ an orthogonal decomposition of the moment function to decouple the joint optimization problem, rendering it structurally comparable to sequential plug-in estimation. In a general setting where the nuisance parameter may be over-identified, we identify an oracle-based locally robust moment function for the plug-in estimator that weakly dominates the locally robust version of the original moment funciton in terms of asymptotic variance. More importantly, we also establish easily verifiable sufficient conditions for the efficiency equivalence between the joint and the plug-in approaches.
This paper is organized as follows. Section (ref) introduces the formal setup and key definitions. Section (ref) develops the differentiation-based derivation across parametric, nonparametric, and semiparametric models. Section (ref) establishes the link between local robustness, adaptive estimation, and efficiency, and applies the framework to a range of treatment effect estimators. Section (ref) concludes. All technical proofs and supplementary results are collected in the online supplement.
Let $Z$ be a random vector with distribution $P$. Following BKRW:1993, all parameters are viewed as functionals of $P$, e.g., $\nu\!=\!\nu(P)$ for a generic parameter $\nu$. This makes it more straightforward to define its influence function $\dot{\nu}$ as a functional derivative in the next subsection. Expectations are written as $P[f(Z)] \!\coloneqq\! \int f(z) dP(z)$, where $z$ is a real vector. Whenever possible, we will omit $Z$ or $z$ for simplicity.
We consider two types of identification for the parameter of interest $\beta$:
where $h_\beta$ and $m_\beta$ are known functions and $\gamma = (\gamma_1^\intercal, \ldots, \gamma_l^\intercal)^\intercal$ is the vector containing all the nuisance parameters. The direct case is a special instance of the indirect one, upon noticing that one can set $m_\beta \!=\! h_\beta - \beta$. When there are no nuisance parameters, an estimator of $\beta$ can be obtained by simply replacing $P$ with the empirical measure $\mathbb{P}_n$, i.e., $\hat{\beta}\!=\!\beta(\mathbb{P}_n)$. The asymptotic properties of $\hat{\beta}$ then readily follow from the well-established empirical process theory EmpiricalProcess.
Each $\gamma_j$ is assumed to be identified from $P_{(:|{j-1})}[m_{\gamma_j}]\!=\!0$, which may or may not reduce to a direct identification. The conditional distribution $P_{(:|{j-1})}$ is associated with the decomposition $Z \!=\! (Z^{(1)\intercal}, \ldots, Z^{(l)\intercal})^\intercal$. For $i \leq j$, let $Z^{(j:i)} \!=\! (Z^{(i)\intercal}, \ldots, Z^{(j)\intercal})^\intercal$. For $j \!=\! 2, \ldots, l$, let $P_{(:|{j-1})} \!\coloneqq\! P_{(l:j|j-1:1)}$ denote the conditional distribution of $Z^{(l:j)}$ given $Z^{(j-1:1)}$. When $j\!=\!1$, $P_{(:|{0})} \!=\! P$ reduces to the unconditional distribution. The decomposition is based on the levels of “exogeneity” of each $Z^{(j)}$. For example, we have $P_{(2|1)} \!=\! P_{(Y|X)}$ in linear regression model. In instrumental variables (IV) estimation, we get $P_{(2|1)} \!=\! P_{(Y,X|W)}$. In treatment effect analysis, $P_{(:|2)} \!=\! P_{(Y|T, X)}$ and $P_{(:|1)} \!=\! P_{(Y,T|X)}$.
In econometric literature, conditional and unconditional moment restrictions, as well as finite- and infinite-dimensional parameters, are often treated separately. Here, we try to find some symbolic similarities by adopting a generalized function perspective. More specifically, define the Dirac delta function $\delta_z$ as the distributional derivative of the Heaviside function $H_z(z') \!=\! 1_{\{z' \geq z\}}$, which satisfies $H_z[f] \!=\! f(z)$ in the Lebesgue–Stieltjes integral sense. With some abuse of notation, we symbolically write $H_z[f] \!=\! \int f(z') \delta_z(z') \mu(dz')$. Let $Z^{(1)}=X$ and consider the identification of the value of an unknown function $\gamma$ at $x$ (denoted $\gamma_x$):
where $\delta_{x}^\dag \!\coloneqq\! P[\delta_{x}]^{-1} \delta_{x}$. Note that the introduction of the generalized function $\delta_x^\dag$ symbolically transforms the conditional moment condition into an unconditional one.
When (ref) holds for all $x$ in the support of $X$, we can obtain a conditional moment condition $P[m_\gamma(Z, \gamma) \!\mid\! X] \!=\! 0$, without the appearance of generalized functions. Intuitively, this identifies the infinite-dimensional parameter $\gamma(\cdot)$, which is a function of $x$, as $\gamma(\cdot, P_{(Z|X)})$. We show in Section (ref) that computing the influence function of $\gamma(\cdot)$, denoted by $\dot{\gamma}$, is significantly more straightforward than computing $\dot{\gamma}_x$, where $\gamma_x \!\coloneqq\! \gamma(x)$. More importantly, the estimator of $\beta$ typically requires the entire function $\gamma(\cdot)$ as input. Consequently, it is $\dot{\gamma}$ rather than $\dot{\gamma}_x$ that appears in the influence function of $\beta$ (see Section (ref)).
Throughout the paper, we maintain the strong identification assumption and defer the analysis of weak identification to future work.
We require only positive semi-definiteness of the weighting matrices to accommodate potential over-identification of $\beta$ and/or the nuisance parameters $\gamma_j$.
Motivated by the structural parallels between finite- and infinite-dimensional parameters, we analyze a generic statistical functional $\nu\!=\!\nu(P)$. Let $\mathcal{L}_2(P)$ denote the space of square-integrable functions with respect to $P$, and let $\mathcal{L}_2^0(P) \!\subseteq\! \mathcal{L}_2(P)$ be the subspace of mean-zero functions. We equip $\mathcal{L}_2(P)$ with the inner product $\langle \cdot, \cdot \rangle$. Following standard semiparametric theory, we define functional derivatives via pathwise differentiability Hampel:1974, Huber:1984, BKRW:1993.
The influence function is defined similarly across methods, but the challenge lies in deriving its explicit form. For example, Newey:1994 uses a score-based expression, which is extended to three-step settings by Hahn&Ridder:2013; tsiatis2006semiparametric relies on tangent space methods and Ichimura&Newey:2022 propose a least squares-based approach.
Our method begins with the integral representation of the derivative $\dot{\nu}(P; Q - P)$. To ensure tractability, it is typically assumed that $\dot{\nu}(P; Q - P)$ is linear and continuous in $Q - P$. Under G\^{a}teaux differentiability, Huber:1984 shows that (Chapter 2.5 therein):
where $\dot{\nu}(P) \!\in\! \mathcal{L}_2^0(P)$. Note that the second equality holds only if $P[\dot{\nu}] \!=\! 0$, as otherwise infinitely many functions could satisfy the first equality (e.g., $\dot{\nu} + \mathbf{v}$ for any finite vector $\mathbf{v}$). By the Riesz representation theorem (e.g., Rudin:1987), the function $\dot{\nu}(P)$ is unique and is referred to as the influence function or influence curve, as in Hampel:1968, Hampel:1974. In other words, the Riesz representation here is a simple projection from $\mathcal{L}_2(P)$ onto $\mathcal{L}_2^0(P)$ by subtracting the mean. BKRW:1993 suggest taking $Q$ as a point mass at $z$ to obtain $\dot{\nu}(z, P)$ (pp. 19 therein). We will take an alternative approach, which applies to any admissible $Q$.
The regularity of parameters is closely related to the functional derivative $\dot{\nu}(P; Q-P)$. The following definition is originally applicable to finite-dimensional parameters. However, a slight modification (see Section (ref)) extends this to infinite-dimensional parameters.
Intuitively, differentiability ensures smooth variation of $\nu$ with respect to $P$. Linearity guarantees the existence of asymptotically linear estimators. Square integrability ensures finite variance for estimators with the parametric convergence rate; for nonparametric estimators, which have slower convergence rates, this may be adjusted (see Section (ref)).
We illustrate the derivation of the influence function $\dot{\nu}(P)$ from its functional derivative $\dot{\nu}(P; Q-P)$ using two foundational examples.
The derivation of the linear system $(\partial_\nu F) \dot{\nu}(P; Q - P)$ follows from a standard perturbation argument. Let $\nu^\epsilon \!\coloneqq\! \nu(P^\epsilon)$. Intuitively, the pertubated parameter $\nu^\epsilon$ satisfies the moment condition up to the order $o(\epsilon)$ (or zero): $P^\epsilon[m(\nu^\epsilon)] \!=\! o(\epsilon)$. Applying a first-order Taylor expansion around $(P, \nu)$ yields:
For this equality to hold for all admissible directions $Q-P$, the $O(\epsilon)$ terms must vanish, which yields the linear system used to solve for $\dot{\nu}(P; Q-P)$. Evaluating such expressions requires handling terms of the form $P[w(Z) \dot{\nu}(P; Q - P)]$, which we resolve via the following key representation.
The proof relies on the richness of the admissible perturbation class. Because the functional derivative is linear in $Q-P$ and the set of admissible directions spans a dense subspace of $\mathcal{L}_2^0(P)$, we may treat the perturbation density as separable from the baseline measure for the purpose of integration. Formally, applying Fubini’s theorem to the double integral yields:
Lemma (ref) below extends this representation to the conditional case, which is essential for deriving influence functions in nonparametric and semiparametric models.
Notably, the moment function $m$ needs not be classically differentiable with respect to $\beta$, contrasting with Hahn&Ridder:2013, who impose standard smoothness conditions (see the statement preceding their Lemma 1). Instead, we employ distributional (Sobolev) derivatives Sobolev, which extend differentiation to functions with discontinuities and preserve the validity of integration-by-parts and Fubini-type interchanges.
We illustrate our method using the standard nonparametric regression model, where $Z \!=\! (X^\intercal, Y)^\intercal$:
A critical distinction must be drawn between the functional $\gamma(X) \!=\! P_{(Y|X)}[Y]$, which depends on the random variable $X$, and its point evaluation $\gamma_x \!\coloneqq\! \gamma(x)$ at any given $x$.
Expanding the perturbed joint measure to first order in $\epsilon$ yields
Algebraic rearrangement gives the corresponding perturbation of the conditional law:
Crucially, the identification $\gamma(X) \!=\! P_{(Y|X)}[Y]$ depends only on the conditional measure and is invariant to perturbations of the marginal $P_{(X)}$. This invariance permits a direct derivation of $\dot{\gamma}(P_{(Y|X)})$ (refer to the Proof of Theorem (ref) for a general treatment). Setting $Q_{(X)} \!=\! P_{(X)}$ (so that $dQ_{(X)}/dP_{(X)}^\epsilon \!\equiv\! 1$) and extending Definition (ref) to conditional measures yields
This mirrors the direct identification structure of Example (ref). Under the standard assumption that $\text{Var}(Y \!\mid\! X)$ is finite, the infinite-dimensional parameter $\gamma(\cdot)$ satisfies the regularity conditions outlined in Definition (ref).
By contrast, the point evaluation $\gamma_x$ depends explicitly on the marginal density $\text{\Fontskrivan{p}}(x) \!=\! P[\delta_x(X)]$. Specifically, $\gamma_x$ is a nonlinear functional of both $P$ and the Dirac measure $\delta_x$:
Constructing an estimator for $\gamma_x$ requires approximating the singular functional $\delta_x$ with a sequence of regular functionals. Consider the Nadaraya–Watson (NW) kernel estimator and the linear sieve estimator with a growing basis $\boldsymbol{\phi}_J(X) = (\phi_1(X), \dots, \phi_J(X))^\intercal$:
The approximation of $\delta_x$ induces asymptotic bias, whereas the substitution of $P$ with $\mathbb{P}_n$ governs the stochastic fluctuation. To formalize this decomposition, define a smoothed (or approximating) functional $\gamma^{\texttt{B}}_x(P)$ for each case, satisfying $\hat{\gamma}_x = \gamma_x^{\texttt{B}}(\mathbb{P}_n)$:
This yields the standard bias–variance decomposition:
The bias component originates from the regularization (or approximation) of $\delta_x$, whereas the first term captures the sampling variability induced by the empirical measure.
In the parametric setting, direct identification yields a functional that is linear in the underlying distribution (Example (ref)). By contrast, the biased nonparametric functional $\gamma_x^{\texttt{B}}(P)$ is inherently nonlinear due to the normalization induced by kernel or sieve approximations. Nevertheless, the functional derivative $\dot{\gamma}_x^{\texttt{B}}(P; \mathbb{P}_n - P)$ continues to characterize the dominant stochastic term in the von Mises expansion of $\gamma_x^{\texttt{B}}(\mathbb{P}_n) - \gamma_x^{\texttt{B}}(P)$. Specifically, the derivative isolates the first-order empirical process fluctuation, while higher-order remainders are of smaller orders.
Applying a suitably modified version of Definition (ref) to $\gamma_x^{\texttt{B}}$, the functional derivative $\dot{\gamma}_x^{\texttt{B}}$ serves as the nonparametric influence function. While its derivation requires careful handling of the approximation sequence, it yields transparent closed-form expressions that foreshadow the treatment of sequential plug-in semiparametric models. Below, we present two complementary approaches to unify the kernel and sieve specifications within our differentiation framework.
The first approach is to view $\gamma^{\texttt{B}}_x$ as a special case of a generic parameter $\nu_x(P) \!=\! c^\intercal (P[g])^{-1} P[h]$:
Although the linear sieve case involves matrix-valued functions, the derivation of influence function remains rather intuitive. Taking differential on both sides of $M M^{-1} \!\equiv\! I$ yields (where $M$ is any invertible matrix-valued function):
Based on this, we have
Plugging in the corresponding values for $c, g, h$, the nonparametric influence functions are
where both $k_{b,x}^\dag(\cdot,P)$ and $\boldsymbol{\phi}_{J,x}^\dag(\cdot, P)$ are approximations of $\delta_x^\dag(\cdot,P) \!\coloneqq\! (P[\delta_x(\cdot)])^{-1} \delta_x(\cdot)$:
The structure of the influence function becomes clearer in this context, mirroring (ref), where $\dot{\nu}(P)$ is linear in the moment function $m(Z, \nu)$. Here, the moment function is $m(Z, \gamma) \!=\! Y - \gamma(X)$, and $\dot{\gamma}^{\texttt{B}}_x$ is linear in $m(Z, \gamma^{\texttt{B}}_x)$. The “loading” term, such as $k_{b,x}^\dag$ or $\boldsymbol{\phi}_{J,x}^\dag$, reflects the estimation method and links the moment function to the influence function. For fixed $b$ or $J$, the variance of $\dot{\nu}_x(P)$ is finite, but to ensure $\gamma^{\texttt{B}}_x - \gamma_x \!=\! o_p(1)$, we need $b \to 0$ or $J \to \infty$, leading to diverging variances in the loading terms. The convergence rate of the nonparametric estimator is root-$n$ adjusted by this divergence rate. As long as the rate-adjusted variance of $\dot{\gamma}^{\texttt{B}}_x(P)$ is finite, $\gamma_x$ can be viewed as a (rate-adjusted) regular parameter.
Using the normalized weights defined in (ref), we can express the estimators as plug-in functionals:
Accordingly, the smoothed functional $\gamma^{\texttt{B}}_x(P)$ is a special case of the composite functional:
This representation isolates a key technical feature of our approach: $\nu_x(P)$ depends on $P$ both through the expectation operator and through the weight function $f_x^\dag(\cdot, P)$. Applying the chain rule for functional derivatives yields
The second term requires careful handling because $\dot{f}_x^\dag(X,P; Q-P)$ depends linearly on the perturbation $(Q-P)$ but retains a stochastic dependence on $X$. When integrating with respect to $P$, components independent of $(Q-P)$ are grouped with $Y$ and evaluated under the baseline measure. Formally, this interchange relies on the linearity of expectation and the separability of the perturbation direction. Parallel to the previous calculation, we obtain:
In the NW specification, the derivation simplifies considerably because $\nu_x(P)$ is a scalar constant with respect to $X$:
For the linear sieve estimator, the coefficient $\boldsymbol{\phi}_J(x)^\intercal P[g]^{-1}$ is deterministic conditional on $x$:
Combining these results with the first term $(Q-P)[f_x^\dag Y]$ recovers the influence function derived in the previous subsection, confirming the internal consistency of the two perspectives.
Incorporating $\gamma(X) \!=\! P_{(Y|X)}[Y]$ yields a third representation of $\gamma_x(P)$ that mirrors the structure of semiparametric plug-in estimation in (ref):
If the weighting function $f_x^\dag(\cdot, P)$ were square-integrable rather than a regularized approximation to a Dirac measure, this setup would correspond to a standard semiparametric model.
More generally, consider $\gamma_x(P)$ as a special case of the composite functional $\nu_x(P) \!=\! P[ w(Z, P) \nu(P_{(Y|X)}) ]$, where $w$ is a measurable function of $Z$ and $\nu$ is a functional of the conditional distribution $P_{(Y|X)}$. Applying the product rule for functional derivatives yields
where $\dot{\nu}(P_{(Y|X)}; Q_{(Y|X)} - P_{(Y|X)}) \equiv \dot{\nu}(P_{(Y|X)}; Q - P)$ since $\nu(P_{(Y|X)})$ is immunte to any pertubation of $P_{(X)}$.
The first term is already in the canonical $(Q-P)[\cdot]$ form (cf. Example (ref)). The second term follows from the composite function perspective developed above, combined with Lemma (ref). The third term captures the perturbation of the inner conditional functional and is defined by
Using the perturbation expansion in (ref) and the zero-mean property $P_{(Y|X)}[\dot{\nu}(P_{(Y|X)})] \!=\! 0$, we evaluate the limit as follows:
The limit
formalizes the change-of-measure step: the outer integration with respect to $P_{(X)}$ is replaced by integration against $Q_{(X)}$. The subsequent calculation eventually replaces the expectation of $w$ under $P$ with that under $P_{(Y|X)}$. Notably, even when $w$ is a regularized approximation to a Dirac measure, the smoothing parameters (e.g., bandwidth or sieve dimension) remain fixed at this stage of the derivation. Thus, $w$ stays bounded and square-integrable, ensuring that Fubini's theorem applies without requiring passage to the singular limit.
By replacing $\nu$ with $\gamma$ and $w$ with $f_x^\dag$, this representation recovers the same influence function as the previous two approaches, but explicitly isolates the conditional perturbation component.
Lemmas (ref) and (ref) share a common algebraic structure. Let $P_{(:|{j-1})}$ denote the conditional distributions defined in Section (ref). Both results are unified by the following general rule.
This general rule bypasses model-specific tangent-space characterizations and guess-and-verify constructions, reducing the derivation to a straightforward expecation.
In multivariate settings, the curse of dimensionality primarily impacts the approximation of the singular functional $\delta_x^\dag$ by $f_{x}^\dag$. Structural restrictions (e.g., additivity or low-dimensional interaction kernels) are commonly imposed to mitigate this. While such constraints modify the explicit form of $f_{x}^\dag$, they preserve the multiplicative structure of the influence function in (ref), which is dictated solely by the underlying identification condition in (ref).
Alternative approximation schemes (e.g. local polynomials, neural networks, or other machine learning estimators) generate distinct weighting functions $f_{x}^\dag$ that may depend on $P$ in a more complex manner. Even when closed-form expressions are unavailable, the influence function intuitively should retain the canonical form $\dot{\gamma}_x^{\texttt{B}} \!=\! f_{x}^\dag \mspace{1mu} \dot{\gamma}(P_{(Y,X|W)})$, where one may need to use the implicit function theorem to compute $\dot{\gamma}(P_{(Y,X|W)})$ as in Example (ref). Consequently, deriving the influence function remains a systematic application of the chain rule and the above expectation-substitution rule. Quantifying the associated nonparametric bias, which requires precise control over the approximation error of $\delta_x^\dag$, involves technically demanding mathematical analysis and goes beyond the scope of the current paper.
Lemma (ref) provides a direct route to deriving influence functions in semiparametric settings. We first establish the result for directly identified parameters.
The derived expression depends exclusively on the influence functions $\dot{\gamma}_j$ rather than their regularized approximations $\dot{\gamma}_j^{\texttt{B}}$. Consequently, the influence function $\dot{\beta}$ is invariant to the specific nonparametric estimator employed for the nuisance parameters. This invariance extends to the indirect identification case (Theorem (ref)) and aligns with the classic result Newey:1994, ACHL:2014: provided nuisance estimators are consistent and converge at appropriate rates, the asymptotic variance of $\hat{\beta}$ remains unchanged regardless of the estimation method for $\gamma$. This property also clarifies the choice of our examples in the previous subsection. While complex nonparametric or machine learning estimators affect higher-order remainders and bias terms, they do not alter the first-order influence function. A detailed analysis of these higher-order properties lies outside the scope of this paper.
Existing approaches for semiparametric moment models Ai&Chen:2003, Ichimura&Newey:2022 typically derive influence functions by solving abstract projection or minimum-distance/least-squares problems Ai&Chen:2003 Ichimura&Newey:2022. Our differentiation-based framework bypasses these steps, yielding closed-form expressions without explicit tangent-space characterizations. The following theorem extends the result to indirect identification.
The expression enclosed in parentheses in (ref) is precisely the locally robust moment function LocalRobust:2022:
The asymptotic variance of $\dot{\beta}$ depends jointly on the nuisance weighting functions $\mathcal{W}_j$ in (ref) and the target weighting matrix $\mathcal{W}_{\beta\beta}$. This raises a central efficiency question: does the choice of $\mathcal{W}_j$ that optimally estimates the nuisance parameters $\gamma$ also minimize the asymptotic variance of $\dot{\beta}$?
The term $v_\rho$ in Ichimura&Newey:2022 corresponds to $P_{(:|{j-1})}[\partial_{\gamma_j} m_{\gamma,j}]$ in (ref). However, their least-squares formulation relies on the ratio $-v_{m}(X)/v_\rho(X)$, which suggests that $v_\rho$ is scalar-valued. Consequently, the selection of an optimal nuisance weighting matrix $\mathcal{W}_j$ does not arise in their framework. We characterize the optimal choice of $\mathcal{W}_j$ in the general matrix-valued case in Section (ref), where we will construct a weakly better moment function to plug-in.
Early robustness literature primarily addressed sensitivity to outliers Tukey:1960. With the development of sequential plug-in estimation, focus shifted toward robustness against specification errors and first-step estimation errors in nuisance parameters. Key concepts include double robustness Robins&Rotnitzky:2001Comment and local robustness LocalRobust:2022. While double robustness relies on second-order influence functions, local robustness is a first-order property. This paper focuses on local robustness, leaving second-order analysis for future work.
Here, we treat local robustness as a property of the moment function $m_\beta$ rather than the parameter $\beta$ itself. This distinction is necessary because $\beta$ may admit multiple identifying moment conditions (e.g., ATE), each with distinct sensitivity to nuisance estimation error. The following theorem provides necessary and sufficient conditions for local robustness under our framework.
Case (i) is a special case of Case (ii) with $l\!=\!1$ and $\gamma\!=\!\gamma_1$ identified under $P_{(:|{0})} \!=\! P$. For $j\!\geq\! 2$, the condition $P_{(:|{j-1})}[\partial_{\gamma_j} m_\beta] \!=\! 0$ is strictly stronger than $P[\partial_{\gamma_j} m_\beta] \!=\! 0$, as discussed in the ATT example below.
Existing approaches for sequential plug-in models typically derive the locally robust moment function by solving projection or least-squares problems Hahn&Ridder:2013, Ichimura&Newey:2022. Hahn&Ridder:2013 adopt a generated regressors perspective, yielding an adjustment term that often involves second-order derivatives via integration by parts. Ichimura&Newey:2022 propose a least-squares projection alternative. Our differentiation-based framework bypasses these constructions, yielding the adjustment term directly.
The results in Section (ref) show that the locally robust moment function corresponds to the term in parentheses in (ref) (recalling (ref)):
Note that $\eta_j$ may depend on $\gamma_j$ and additional nuisance components. By construction, $m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$ is locally robust to $\eta_j$:
Assuming $\eta_{j'}$ ($j'\geq j$) may depend on $\gamma_{j}$, applying the law of iterated expectations yields
This confirms that $m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$ is locally robust to all $\gamma_j$ and the auxiliary parameters $\eta_j$. The adjustment term $\sum_{j=1}^l \eta_j m_{\gamma,j}$ coincides with the first-order correction derived in Newey:1994 and LocalRobust:2022.
When $m_\beta \!\neq\! m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$, researchers may choose between the raw and locally robust identification strategies. The corresponding optimal asymptotic variances are:
These variances coincide under the standard condition that the nuisance moment functions are locally robust to $\beta$, i.e., $P_{(:|{j-1})}[\partial_\beta m_{\gamma, j}] \!=\! 0$ for all $j$. This condition is mild: in most semiparametric two-step models, $m_{\gamma,j}$ does not depend on $\beta$ at all. Otherwise, a consistent estimator of $\beta$ would be needed to estimate $\gamma_j$, leading to a circular dependency. Given this condition, (ref) implies $P[\partial_\beta m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta] \!=\! P[\partial_\beta m_\beta]$, as the derivative of the adjustment term with respect to $\beta$ has zero expectation. Consequently, using $m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$ does not lead to a more efficient estimator. Its primary role is to mitigate first-order bias arising from nuisance estimation error LocalRobust:2022.
Bickel:1982Adaptive formalizes the theory of adaptive estimation, addressing Stein's classical question: “when can a Euclidean parameter be estimated as efficiently without knowledge of an infinite-dimensional nuisance parameter as with full knowledge of it?” Stein:1956.
Within the moment-based framework, an affirmative answer for a given identification strategy corresponds to the following property, which, to the best of our knowledge, has not been previously highlighted in the literature and plays a crucial role in the subsequent discussion.
When the true $\gamma$ is known, estimation reduces to the parametric prototype in Example (ref) with $\nu \!=\! \beta$. The feasible plug-in case follows from Theorem (ref). Their respective optimal asymptotic variances are
The sufficiency direction follows immediately: local robustness implies $m_\beta \!=\! m^{\mathrel{\scriptstyle \texttt{LR}}}_\beta$, so the two variances coincide.
To examine necessity, assume the relevant covariance matrices are invertible. Using $\eta_j$ introduced in (ref), the difference in the covariance kernels is
Under the orthogonality conditions noted in the proposition, the expression simplifies to
Equality of the variances requires this difference to vanish, which implies $\eta_j \!=\! 0$ almost surely, and consequently $P_{(:|{j-1})} [\partial_{\gamma_j} m_\beta ] \!=\! 0$ with probability one. While necessity may hold under weaker orthogonality conditions, a complete characterization is left for future work.
Example (ref) highlights an interesting scenario: one may have more than one moment function to identify $\beta$. Furthermore, even when there is only one $m_\beta$ to use when one needs to estimate $\gamma$, one can have multiple $\tilde{m}_\beta \!\coloneqq\! m_\beta + B m_\gamma$ in the hypothetical case where $\gamma$ were known. This leads to the oracle moment funciton that will be discussed in the next subsection.
ACHL:2014 raise the question of whether sequential two-step estimation achieves the same efficiency as joint estimation. They note that, although the plug-in method often offers substantial computational advantages, it may suffer from “limited information,” in the sense that the moment conditions are not jointly exploited. By contrast, the joint procedure fully utilizes all available moment restrictions simultaneously and may therefore deliver efficiency gains. However, it typically requires a high-dimensional nonlinear search over both finite- and infinite-dimensional parameter spaces.
ACHL:2014 derive the influence function for the plug-in estimator under exact identification of the nuisance parameter and note a methodological distinction from the orthogonalization approach of Ai&Chen:2012. Specifically, ACHL:2014 construct the adjustment term using derivatives of the moment function, whereas Ai&Chen:2012 employ a variance-covariance projection. However, they do not comment on whether the two constructions coincide (see Section 3.4 of ACHL:2014).
To compare plug-in and joint procedures rigorously, we first identify the optimal moment function for plug-in estimation. The locally robust moment function $m_{\beta}^{\mathrel{\scriptstyle \texttt{LR}}}$ is a natural candidate. When all nuisance parameters are exactly identified, the weighting matrices $\mathcal{W}_j$ do not affect the asymptotic variance, leaving $m_{\beta}^{\mathrel{\scriptstyle \texttt{LR}}}$ as the only choice.
When nuisance parameters are over-identified, however, the optimal choice of $\mathcal{W}_j$ becomes nontrivial. Besides, it is theoretically possible that another moment function may weakly dominate $m_{\beta}^{\mathrel{\scriptstyle \texttt{LR}}}$. Motivated by the structure of joint estimation, we characterize this optimal specification through an oracle perspective.
To clarify the construction, first consider the case where $\beta$ and $\gamma$ are identified under the same distribution $P$. If $\gamma$ were known, any moment function of the form $m_{\beta} + B m_{\gamma}$ identifies $\beta$. The optimal oracle moment function is obtained by projecting $m_\beta$ onto the orthogonal complement of the nuisance moment space:
The linear family $m_\beta + B m_\gamma$ is assumed to span the space of all admissible moment conditions for $\beta$. The rationale is that any zero-mean transformation of the existing moments should already be included in them.
In the general setting, $\beta$ and the nuisance components $\gamma_j$ are identified under distinct probability measures. The following orthogonal decomposition for any $f \in \mathcal{L}_2(P)$ proves instrumental:
Each component $f_j$ depends on the data only through $Z^{(j:1)}$ and satisfies $f_j \!\in\! \mathcal{L}_2^0(P_{({j}|:)})$, where $P_{({j}|:)} \!\coloneqq\! P_{(j|j-1:1)}$. Applying this decomposition to $m_\beta$ yields $m_{\beta,0} \!=\! 0$.
To streamline the analysis, we assume that $m_{\beta}$ does not lie entirely within any single subspace $\mathcal{L}_2^0(P_{({j}|:)})$ and that the ordering of $Z$ and $\gamma$ can be arranged so that each $m_{\gamma,j} \!\in\! \mathcal{L}_2^0(P_{({j}|:)})$ for $j\!=\!1,\ldots,l$. Standard IPW and AIPW specifications satisfy this structure, whereas the regression-based estimator does not (see Section (ref)). Under this arrangement, $m_{\gamma,j}$ and $m_{\gamma,j'}$ are orthogonal for $j\!\neq\! j'$, and $m_\beta$ can correlate with $m_{\gamma,j}$ only through its component $m_{\beta,j}$.
Since all second-moment matrices are positive semi-definite, $V_{\beta\gamma,j} (I - V_{\gamma\gamma,j}^+ V_{\gamma\gamma,j} ) \!=\! 0$ almost surely (refer to Theorem 16.1 of Gallier:2011), where
This permits an orthogonal projection within each conditional subspace:
Aggregating across components yields the optimal oracle moment function for $\beta$:
Unless $V_{\beta\gamma,j} \!=\! 0$ almost surely, the oracle moment function $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ differs from $m_{\beta,j}$. This yields two candidates for plug-in estimation. We construct and compare the locally robust versions of both:
Define the Jacobian components (recall that $P_{({j}|:)}[\partial_{\beta} m_{\gamma,j}] \!=\! 0$ is required for valid plug-in estimation):
The terms $\Delta_j$ and $P_{({j}|:)}[\partial_{\gamma_j} m_{\beta,j}]$ may introduce auxiliary nuisance parameters, but the resulting moment functions remain locally robust to them because $P_{({j}|:)}[\dot{\gamma}_j] \!=\! 0$ by construction.
With the notation, the influence funciton $\dot{\gamma}_j $ given in (ref) can be written as
The conditional variances of $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ and $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{LR}}}$ depend explicitly on $\mathcal{W}_j$. Denote these by $\text{Var}_j^{\mathrel{\scriptstyle \texttt{orc-LR}}}(\mathcal{W}_j)$ and $\text{Var}_j^{\mathrel{\scriptstyle \texttt{LR}}}(\mathcal{W}_j)$, respectively.
The condition $(I - V_{\gamma\gamma,j} V_{\gamma\gamma,j}^+) J_{\gamma,j} \!=\! 0$ is stronger than the invertibility of $J_{\gamma,j}^\intercal V_{\gamma\gamma,j}^+ J_{\gamma,j}$ required in Assumption (ref). Geometrically, it requires that the column space of $J_{\gamma,j}$ lies entirely within the column space of $V_{\gamma\gamma,j}$, rather than merely having no column contained in the nullspace. This condition guarantees that the natural weighting choice $\mathcal{W}_j \!=\! V_{\gamma\gamma,j}^+$ is optimal for estimating $\gamma_j$, making the assumption mild in standard over-identified settings.
Because $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc}}}$ is constructed to be orthogonal to $\dot{\gamma}_j$, the same weighting matrix $\mathcal{W}_j \!=\! V_{\gamma\gamma,j}^+$ minimizes $\text{Var}_j^{\mathrel{\scriptstyle \texttt{orc-LR}}}(\mathcal{W}_j)$. The optimal weighting for $\text{Var}_j^{\mathrel{\scriptstyle \texttt{LR}}}(\mathcal{W}_j)$ may differ from $V_{\gamma\gamma,j}^+$ (see the supplement for techinical details). Theorem (ref) implies that $m_\beta^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ achieves the minimal attainable variance among all valid plug-in moment specifications, and thus serves as the appropriate benchmark for comparing plug-in and joint estimation efficiency.
The orthogonal decomposition of $m_{\beta}$ facilitates a direct comparison between joint and sequential plug-in estimation. Specifically, the condition $P_{({j}|:)}[m_{\beta,j}] \!=\! 0$ identifies a parameter $\beta_j$, which we term a companion parameter. When $j\!>\!1$, this parameter is infinite-dimensional.
Before proceeding to the general analysis, we exclude two degenerate configurations:
The first implies $\beta_j \!=\! 0$; the second implies $\beta_j$ is a known linear transformation of $\gamma_j$. In both instances, there is no distinction between joint and plug-in estimation. Both configurations arise in the AIPW estimator for the ATE (see Section (ref) for details).
In the non-degenerate case, we group $\beta_j$ and $\gamma_j$ into a joint parameter $\nu_j$ and define the stacked moment vector and its conditional covariance:
Assuming $(I - V_{\nu\nu,j} V_{\nu\nu,j}^+) P_{({j}|:)}[\partial_\nu m_j] \!=\! 0$, the asymptotic efficiency of the joint estimator is governed by the optimal variance matrix
Under the identifying restriction $P_{({j}|:)}[\partial_{\beta_j} m_{\gamma,j}] \!=\! 0$, the covariance matrix $V_{\nu\nu,j}$ admits a block decomposition (cf. Gallier:2011):
where $S_{\beta\beta,j} \!\coloneqq\! V_{\beta\beta,j} - V_{\beta\gamma,j} V_{\gamma\gamma,j}^+ V_{\gamma\beta,j}$ is the Schur complement of $V_{\gamma\gamma,j}$ in $V_{\nu\nu,j}$. Notably, $S_{\beta\beta,j} \!=\! \langle m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}, (m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j})^\intercal \rangle_j$, linking this decomposition directly to the oracle construction in Theorem (ref).
Standard results on the Moore--Penrose inverse (e.g., MPinverse:2023) imply that if $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) V_{\beta\gamma,j} \!=\! 0$ almost surely, then $V_{\nu\nu,j}^+$ admits the closed-form block structure:
Substituting this into the variance formula yields
This representation establishes that joint estimation is asymptotically equivalent to using the oracle moment $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc}}}$ and the nuisance moment $m_{\gamma,j}$ simultaneously.
To streamline the efficiency comparison, define the following matrix components:
Here, $\Sigma_{\beta\beta,j}^{-1}$ represents the optimal oracle variance for $\beta_j$ when $\gamma_j$ is known, while $\tilde{\Sigma}_{\gamma\gamma,j}^{-1}$ is the optimal variance for estimating $\gamma_j$ alone when $(I - V_{\gamma\gamma,j} V_{\gamma\gamma,j}^+) J_{\gamma,j} \!=\! 0$. The local robustness of $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ is central to the following characterization.
Direct inspection yields the matrix inequality
Theorem (ref) implies that the optimal oracle variance $\Sigma_{\beta\beta,j}^{-1}$ is attainable by joint estimation only when $m^{\mathrel{\scriptstyle \texttt{orc}}}_{\beta,j}$ is locally robust ($\Delta_j \!=\! 0$). Under this same condition, Theorem (ref) guarantees that sequential plug-in estimation also achieves the optimal oracle variance. This characterizes the ideal adaptive setting Bickel:1982Adaptive, in which both procedures attain the same asymptotic variance as the infeasible oracle that knows the nuisance parameter.
The remaining question is whether plug-in and joint estimation yield identical asymptotic variance when $\Delta_j \!\neq\! 0$. Plug-in estimation of $\beta_j$ using $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ produces the variance
Comparing this to the joint variance $\Omega_{\beta\beta,j}^{-1}$:
If $(I-S_{\beta\beta,j} S_{\beta\beta,j}^+) \Delta_j \!=\! 0$, standard Moore--Penrose identities verify that
The conditions in Theorem (ref) are independent of the identification status of the nuisance parameters, as the optimality of $m_{\beta,j}^{\mathrel{\scriptstyle \texttt{orc-LR}}}$ (Theorem (ref)) already accounts for over-identification. When $S_{\beta\beta,j}$ is invertible, the conditions in part (iii) are automatically satisfied. Thus, the invertibility of the Schur complement $S_{\beta\beta,j}$ is the decisive factor for plug-in--joint equivalence, regardless $V_{\gamma\gamma,j}$ is invertible or not.
In practice, researchers may apply the following sequential procedure:
One may additionally check whether $\Delta_j \!=\! 0$ to determine if the optimal oracle variance is attainable. Violations are common, as illustrated below.
This subsection applies the preceding methodology to two canonical treatment effect estimands: the average treatment effect (ATE), which admits adaptive estimation, and the average treatment effect on the treated (ATT), which does not.
First consider the ATE case. The moment function for the IPW estimator has the following orthogonal decomposition:
The sole nuisance parameter is the propensity score $\pi(X)$, identified by $P_{(T|X)}[m_\pi] \!=\! P_{(T|X)}[T-\pi(X)] \!=\! 0$, which corresponds to $j\!=\!2$. It is easy to see that $m_{\mathrel{\scriptscriptstyle \texttt{ATE}},2} \neq 0$ while its oracle projection vanishes ($m_{\mathrel{\scriptscriptstyle \texttt{ATE}},2}^{\mathrel{\scriptstyle \texttt{orc}}} \!=\! 0$). By Theorem (ref), joint estimation yields no asymptotic efficiency gain over sequential plug-in estimation.
The oracle moment function $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}} = m_{\mathrel{\scriptscriptstyle \texttt{ATE}},3} + m_{\mathrel{\scriptscriptstyle \texttt{ATE}},1}$ satisfies the local robustness condition, confirming that the ATE is adaptively estimable. Because $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\scriptscriptstyle\texttt{IPW}} \!\neq\! m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}}$, plugging the true propensity score into the IPW moment is suboptimal for the oracle.
The regression-based estimator presents a more subtle structure. Its moment function lies entirely in the marginal subspace $\mathcal{L}_2^0(P_{(X)})$, rather than $\mathcal{L}_2^0(P)$:
This aligns with Example (ref), which shows $\text{Var}(m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}}) \!<\! \text{Var}(Y(1) - Y(0))$. The discrepancy arises because $\tau_1(X)$ and $\tau_0(X)$ are companion parameters: they are not of direct interest, but should not be in the oracle's admissible information set.
The reason is best illustrated through a directly identified model $\beta \!=\! P[h_\beta]$. Define the hierarchical projections $h_{\beta,j} \!\coloneqq\! P_{(:|{j})}[h_\beta]$. The induced companion parameters $\beta_j \!=\! P_{({j}|:)}[h_{\beta,j}]$ depend only on $Z^{(j-1:1)}$ and satisfy $\beta \!=\! P[\beta_j]$ for all $j$. Permitting the oracle to observe companion parameters $\beta_j$ ($j\!>\!1$) would artificially deflate the benchmark variance, since $\text{Var}(\beta_j) \!<\! \text{Var}(h_\beta)$. Extending this logic recursively would permit the oracle to observe $\beta_1 \!=\! \beta$, yielding a degenerate zero-variance benchmark. To maintain a well-defined efficiency comparison, the oracle's information set must exclude all companion parameters.
In the regression specification, all nuisance components are companion parameters, placing it in the degenerate regime where $\beta_j \equiv \gamma_j$. Consequently, plug-in and joint estimation are the same, and the locally robust and oracle moment functions coincide: $m_\beta^{\mathrel{\scriptstyle \texttt{LR}}} \equiv m_\beta^{\mathrel{\scriptstyle \texttt{orc}}}$. Thus, the regression-based oracle moment is again $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}} = m_{\mathrel{\scriptscriptstyle \texttt{ATE}},3} + m_{\mathrel{\scriptscriptstyle \texttt{ATE}},1}$.
Lastly, consider the doubly robust AIPW moment function
which is locally robust by construction. The only non-companion nuisance parameter is $\pi(X)$. Because $m_{\mathrel{\scriptscriptstyle \texttt{ATE}},2} \!=\! 0$ in this decomposition, its projection onto the propensity score moment vanishes. Furthermore, because $\tau_1(X)$ and $\tau_0(X)$ are companion parameters ($\beta_j \!\equiv\! \gamma_j$), the oracle construction excludes projection onto their identifying moments $m_{\tau_1}$ and $m_{\tau_0}$ (cf. (ref)). Consequently, the oracle moment function coincides exactly with the AIPW specification: $m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptstyle \texttt{orc}}} \!=\! m_{\mathrel{\scriptscriptstyle \texttt{ATE}}}^{\mathrel{\scriptscriptstyle \texttt{AIPW}}}$.
Consider the average treatment effect on the treated. The IPW moment function admits the orthogonal decomposition:
This specification yields the plug-in estimator
Hirano&Imbens&Ridder:2003 construct an alternative ATT estimator by formulating ATT as a weighted ATE. Their moment function differs from $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptscriptstyle \texttt{IPW}}, \scriptscriptstyle T}$ only in the normalization of the target parameter:
The corresponding estimator is
Two regression-based moment functions follow analogously:
Substituting the identifying conditions for $\tau_1(X)$ and $\tau_0(X)$ into $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptscriptstyle \texttt{Reg}}, \scriptscriptstyle T}$ and replacing $T$ with $\pi(X)$ recovers $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptscriptstyle \texttt{IPW}}, \scriptscriptstyle \pi}$. The associated plug-in estimators are
The first estimator, $\hat{\tau}_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\scriptscriptstyle \texttt{Reg/T}}$, corresponds to the specification in Hahn:1998.
Unlike the ATE case, the ATT admits distinct locally robust and oracle moment functions:
These specifications generate two augmented estimators:
where the adjustment component is
As established by Hahn:1998, the infeasible efficiency bound with known propensity score is
When $\pi(X)$ is estimated, the asymptotic variance increases to $\text{Var}(m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{LR}}}) / P[T]^2$. The efficiency loss equals
By Theorem (ref), this variance increse occurs precisely because $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{orc}}}$ is not locally robust. Direct computation confirms that
which is non-zero almost surely unless the conditional average treatment effect is constant across treated units. Furthermore, applying the local robustness correction to $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{orc}}}$ recovers $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{LR}}}$, implying $m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{orc-LR}}} \!=\! m_{\mathrel{\scriptscriptstyle \texttt{ATT}}}^{\mathrel{\scriptstyle \texttt{LR}}}$.
In this paper, we introduce a direct differentiation-based framework for deriving influence functions across parametric, nonparametric, and semiparametric models. By grounding the derivation in the functional functional derivative and its integral representation, we demonstrate that the influence function is obtained by centering the identification function, that is, subtracting its conditional or unconditional expectation. This algebraic characterization replaces the traditional reliance on tangent-space projections, score-based Riesz representation arguments, and ad hoc guess-and-verify constructions. We formalize this insight through a unified expectation-substitution rule that seamlessly handles both unconditional and conditional perturbations, yielding closed-form influence functions and revealing a common structural form across finite- and infinite-dimensional settings.
Building on this foundation, we derive transparent conditions for local robustness and clarify its precise role in adaptive estimation. A central contribution is our rigorous efficiency comparison between sequential plug-in and joint estimation. By exploiting the orthogonal decomposition of moment conditions along the hierarchy of conditioning variables, we establish verifiable sufficient conditions under which plug-in estimation attains the same asymptotic variance as joint estimation, even when nuisance parameters are over-identified. Furthermore, we construct an oracle-based locally robust moment function that weakly dominates the local robsut version of the original moment function, providing a principled benchmark for multi-step inference. Applied to treatment effect models, the framework precisely diagnoses when plug-in procedures remain adaptive (as in the ATE) and when they incur efficiency losses due to failure of local robustness (as in the ATT).