EconBase
← Back to paper

Ill-Conditioned Orthogonal Scores in Double Machine Learning

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,747 characters · 22 sections · 42 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Ill-Conditioned Orthogonal Scores in Double Machine Learning

abstract\begin{singlespace} Double Machine Learning is often justified by nuisance-rate conditions, yet finite-sample reliability also depends on the conditioning of the orthogonal-score Jacobian. This conditioning is typically assumed rather than tracked. When residualized treatment variance is small, the Jacobian is ill-conditioned and small systematic nuisance errors can be amplified, so nominal confidence intervals may look precise yet systematically under-cover. Our main result is an exact identity for the cross-fitted PLR-DML estimator, with no Taylor approximation. From this identity, we derive a stochastic-order bound that separates oracle noise from a conditioning-amplified nuisance remainder and yields a sufficiency condition for $\sqrt{n}$-inference. We further connect the amplification factor to semiparametric efficiency geometry via the Riesz representer and use a triangular-array framework to characterize regimes as residual treatment variation weakens. These results motivate an out-of-fold diagnostic that summarizes the implied amplification scale. We do not propose universal thresholds. Instead, we recommend reporting the diagnostic alongside cross-learner sensitivity summaries as a fragility assessment, illustrated in simulation and an empirical example. \end{singlespace}

\noindentKeywords: Neyman orthogonality; overlap/positivity; Riesz representer; semiparametric efficiency; weak identification

\noindentJEL Codes: C14, C21, C55

\thispagestyle{plain}

Introduction

Double Machine Learning chernozhukov2018dml permits the application of machine learning methods for nuisance estimation, while preserving $\sqrt{n}$-inference for low-dimensional treatment effects. This holds because Neyman-orthogonal scores reduce first-order sensitivity to nuisance estimation errors and cross-fitting avoids own-observation bias chernozhukov2018dml, kennedy2023semipar. As a result, the corresponding sufficient conditions are typically stated as rate conditions, rather than as parametric smoothness restrictions. In the canonical partially linear regression (PLR) model, for instance, the classical product-rate requirement is $r_n^m r_n^\ell = o_P(n^{-1/2})$, where $r_n^m,r_n^\ell$ denote $L^2$ nuisance-estimation rates chernozhukov2018dml.

These conditions control the entry of nuisance estimation error into the orthogonal score. However, these rate conditions do not track how residual treatment variation after conditioning (the PLR analogue of overlap/positivity) directly impacts finite-sample reliability khanTamer2010overlap,damour2021overlap. Standard DML theory often relies on regularity conditions to keep this residual variation bounded away from zero. We instead make the finite-sample sensitivity induced by low residual treatment variation explicit in the canonical PLR-DML setting through an exact identity and a stochastic bound.\footnote{Our main results are for PLR-DML, where the score is affine in $\theta$ and the identity is exact. Appendix A.1 shows the generic orthogonal-score expansion for nonlinear scores, where a Taylor/linearization step is required.} A simple reportable out-of-fold diagnostic follows as a reporting implication.

An amplification mechanism is transparent in PLR. In this case, the sensitivity of the empirical orthogonal score to the treatment effect $ \theta $ depends on the empirical Jacobian $\widehat{J}_\theta = -\,\widehat{\sigma}_{V}^{2}$, where $V$ is the residualized treatment. If $\widehat{\sigma}_{ V }^{ 2 } $ is small, the score is nearly flat in the $\theta$ direction. Small perturbations of the empirical score equation therefore create large perturbations of $\widehat{\theta}$. This corresponds to ill-conditioning in the numerical sense golubvanloan2013. We capture this channel through $\kappa := { \sigma }_{ D }^{ 2 }/{\sigma}_{ V }^{ 2 } = 1/(1-R^{ 2 }(D\mid X))$. It scales the impact of any residual bias left after orthogonalization and cross-fitting. While standard errors reflect increased sampling variability when $\sigma_{ V }^{ 2 }$ is small, they do not consider the bias-amplification induced by large $ \kappa $. Coverage can fail as a result, even when reported standard errors are only moderately inflated.

Our first main result is an exact finite-sample decomposition for DML in PLR: \[ \widehat{\theta}-\theta_0 \;=\; \widehat{\kappa}\,(S_n' + B_n'). \] Here $\widehat{\kappa}$ is the sample analogue of the condition number, $S_n'$ is the standardized oracle sampling term, and $B_n'$ aggregates the nuisance-driven bias components that remain after orthogonalization and cross-fitting. The identity is exact. In other words, there is no Taylor approximation, because the PLR score is affine in $\theta$. This makes the role of $\widehat{\kappa}$ as a finite-sample amplification factor explicit.

Our second main result converts the exact identity into a stochastic-order bound: \[ \widehat{\theta}-\theta_0 = O_P\!\left(\frac{\sqrt{\kappa_n}}{\sqrt{n}} \;+\; \kappa_n\cdot \mathrm{Rem}_n\right), \qquad \mathrm{Rem}_n := r_n^m r_n^\ell + (r_n^m)^2 + \frac{r_n^m + r_n^\ell}{\sqrt{n}}. \] The first term is the oracle sampling component under conditioning. The second term shows that conditioning enters multiplicatively with the full nuisance remainder. As a consequence, a sufficient condition for the usual $\sqrt{n}$-approximation (oracle term dominates the remainder) is $\kappa_n\cdot \mathrm{Rem}_n = o(n^{-1/2})$. When overlap is stable so that $\kappa_n$ is bounded, this reduces to familiar remainder-rate restrictions. If overlap is weak and $\kappa_n$ is large, the same nuisance errors can be amplified into first-order bias.

Two theoretical connections improve interpretation. First, we link conditioning to semiparametric efficiency geometry by proving $\kappa = \sigma_D^2 \|\alpha_0\|_{L^2}^2$, where $\alpha_0$ is the Riesz representer chernozhukov2022riesz. Second, we formalize weakening overlap through a triangular-array framework in which $\kappa_n\to\infty$. This shows that standard $\sqrt{n}$-asymptotics may fail even when nuisance rates are favorable. The theory also yields an estimable reporting implication. We propose reporting the condition number, \[ \widehat{\kappa}_{\mathrm{oof}} = \frac{1}{1-\widehat{R}^2_{+}(D\mid X)}, \qquad \widehat{R}^2_{+} := \max\{0,\widehat{R}^2_{\mathrm{oof}}\}, \] a scale-invariant measure of residual treatment variation, where $\widehat{R}^2_{\mathrm{oof}}$ is an out-of-fold predictive $R^2$. When multiple first-stage learners are compared, $\widehat{\kappa}_{\mathrm{oof}}$ is computed learner-by-learner.\footnote{Out-of-fold $R^2$ avoids in-sample overfitting bias and aligns the conditioning diagnostic with the cross-fitting used in DML. In PLR, the amplification factor in the bound satisfies $\kappa=\sigma_D^2/\sigma_V^2=1/\{1-R^2(D\mid X)\}$, so $\widehat{\kappa}_{\mathrm{oof}}=1/\{1-\widehat R^2_{+}\}$ is a plug-in proxy for the same conditioning object.} This operationalizes the overlap/conditioning component that enters the bound and helps interpret cross-learner sensitivity predicted by the theory. Importantly, $\widehat{\kappa}_{\mathrm{oof}}$ concerns the conditioning of the orthogonalized score, but it is not a test of unconfoundedness and does not by itself validate causal identification. Analogous to weak instruments, where weak first-stage amplifies bias and motivates strength diagnostics staigerstock1997weakiv,stockyogo2005weakiv,moreira2003conditional, limited residual treatment variation in DML implies large $\kappa$, which amplifies residual nuisance bias in the orthogonal score equation.

The paper proceeds as follows. Section (ref) situates our contribution in the literature. Section (ref) develops the exact decomposition and stochastic-order bound. Section (ref) characterizes conditioning regimes and rate implications. Section (ref) presents Monte Carlo evidence validating the theoretical predictions. Section (ref) provides an empirical illustration. Section (ref) concludes.

Related Work

chernozhukov2018dml formalize DML as Neyman-orthogonal, cross-fitted score estimation, building on Robinson's partialling-out construction for PLR robinson1988plr and influence-function theory newey1990semipar. Subsequent work develops finite-sample guarantees for orthogonal-score estimators chernozhukov2023simple. kennedy2023semipar reviews influence-function--based semiparametric inference (including cross-fitted one-step/TMLE constructions and doubly robust structure) and chernozhukov2021rieszregression develops a geometric debiasing view via Riesz representers. We build on this literature, but shift attention to a typically implicit PLR regularity.

Standard orthogonal-score DML analyses pair nuisance-rate conditions with a nondegeneracy condition for the score equation. In PLR, this amounts to requiring residualized treatment variation so the score Jacobian is not close to singular chernozhukov2018dml. When this curvature deteriorates, inference becomes a weak-identification/ill-conditioning problem. Estimating equations flatten and small singular values act as error amplifiers kaji2021weakid,hanmccloskey2019singular,breunig2020illposed. Yet in applied DML this stability component is rarely tracked or reported alongside the final estimate. We make conditioning an explicit channel in PLR-DML by isolating a condition-number factor $\kappa$ that governs amplification of the orthogonal-score remainder as residual treatment variation shrinks. This links overlap-driven fragility in DML to classical collinearity diagnostics belsley1980diagnostic.

Overlap matters through residualized treatment variation, which governs both identification strength and finite-sample sensitivity. For binary treatment, efficiency theory implies sensitivity increases as propensity scores approach 0 or 1 hahn1998, motivating strict-overlap conditions and practical remedies such as trimming crumpetal2009 and overlap weighting lietal2018overlap. When overlap fails, effects may be only partially identified khanTamer2010overlap, and overlap can deteriorate with covariate dimension, creating tension between flexible adjustment and identification strength damour2021overlap. We use this overlap literature primarily to clarify the mechanism by which weak overlap undermines PLR-DML, while leaving the analysis of overlap-adjusted estimators to future work.

Modern causal estimators rely on regularized, data-adaptive learners for nuisance components, so finite-sample orthogonal-score remainders need not be negligible even with cross-fitting. High-dimensional and debiased approaches for learned nuisances are developed in bellonietal2014highdim, and sufficient conditions for valid semiparametric inference with flexible learners are studied in farrellliangmisra2021. Related robustness ideas are central in doubly robust and targeted learning methods kangschafer2007,bangrobins2005,kennedy2023semipar. These results primarily govern how small nuisance error must be. We complement them by tracking how that error is scaled by the sample geometry of the PLR score equation. Because this scaling depends on the particular out-of-fold nuisance fits, our diagnostic is inherently specification-specific for each learner.

Theoretical Results

table[table omitted — 1,603 chars of source]

This section introduces notation and assumptions for PLR--DML and develops our finite-sample characterization of how limited residual treatment variation affects inference.\footnote{We keep the main bound in stochastic-order form because fully nonasymptotic Gaussian approximation requires additional constants and concentration tools, such as finite-sample analyses for sample-splitting/DML-type estimators.} We present (i) an explicit link between overlap strength and a condition number $\kappa$, (ii) an exact finite-sample identity for $\widehat{\theta}-\theta_0$, and (iii) a stochastic-order bound that yields an operational sufficiency condition for $\sqrt{n}$-valid inference. We write $\kappa$ for the population condition number under a fixed law $P$, and $\kappa_n$ under a triangular-array sequence $(P_n)_{n\ge1}$.

Data Structure

assumption[Data Structure] For each $n$, we observe $\{W_{i,n}\}_{i=1}^n$ i.i.d.\ under $P_n$, where $W_{i,n} = (Y_{i,n}, D_{i,n}, X_{i,n})$ consists of outcome $Y \in \mathbb{R}$, treatment $D \in \mathbb{R}$, and covariates $X \in \mathbb{R}^p$. The classical i.i.d.\ setting is $P_n \equiv P$. We suppress index $n$ when unambiguous.
definition[Function Spaces] For measurable $f: \mathcal{X} \to \mathbb{R}$, define $\|f\|_{L^2} := (\mathbb{E}[|f(X)|^2])^{1/2}$ and $\|f\|_n := (n^{-1}\sum_{i=1}^n f(X_i)^2)^{1/2}$. We employ standard stochastic order notation: $Z_n = O_P(a_n)$ if $|Z_n/a_n|$ is bounded in probability; $Z_n=o_P(a_n)$ if $Z_n/a_n \xrightarrow{p} 0$.

Assumption map. Lemma (ref) uses only $\mathbb{E}[V\varepsilon] = 0$ and $\sigma_V^2 > 0$. The exact decomposition (Theorem (ref)) is algebraic under PLR and does not require rate conditions. Finite-sample bounds (Theorem (ref)) additionally use moment bounds, nuisance $L^2$ rates, and residual-variance stability. Regime analysis allows $\sigma_{V,n}^2 \to 0$ via the triangular array setup.

The Partially Linear Regression Model

The PLR model robinson1988plr specifies:

align[align omitted — 105 chars of source]

where $\theta_0 \in \mathbb{R}$ is the target parameter, $g_0: \mathcal{X} \to \mathbb{R}$ is a nuisance function, $m_0(X) := \mathbb{E}[D \mid X]$ is the treatment regression (conditional mean), and $V := D - m_0(X)$ is the treatment residual satisfying $\mathbb{E}[V \mid X] = 0$ by construction.

remark[Terminology] When $D$ is binary, $m_0(X) = \mathbb{E}[D \mid X]$ equals the propensity score $e(X)$ rosenbaumrubin1983. When $D$ is continuous, $m_0(X)$ is the treatment regression (conditional mean), and our overlap notion is expressed through $\mathrm{Var}(D \mid X)$.

Target Parameter and Identification

definition[Target Parameter] The target parameter is the constant marginal effect $\theta_0$ in the partially linear model $Y = \theta_0 D + g_0(X) + \varepsilon$.
lemma[Identification of $\theta_0$] Under the PLR model (ref)--(ref) and Assumption (ref), the target parameter $\theta_0$ is identified as the unique solution to \[ \mathbb{E}\big[V\{Y - \theta D - g_0(X)\}\big] = 0, \] equivalently, \[ \theta_0 = \frac{\mathbb{E}[VY]}{\mathbb{E}[V^2]} = \frac{\mathbb{E}[(D - m_0(X))Y]}{\mathbb{E}[(D - m_0(X))^2]}, \] provided $\sigma_V^2 = \mathbb{E}[V^2] > 0$. This “partialling-out” identification strategy dates to frischwaugh1933 and robinson1988plr.
proofFrom (ref), $Y = \theta_0 D + g_0(X) + \varepsilon$. Multiply by $V = D - m_0(X)$ and take expectations: \[ \mathbb{E}[VY] = \theta_0 \mathbb{E}[VD] + \mathbb{E}[Vg_0(X)] + \mathbb{E}[V\varepsilon]. \] Now $\mathbb{E}[VD] = \mathbb{E}[V(m_0(X) + V)] = \mathbb{E}[V^2] = \sigma_V^2$, and $\mathbb{E}[Vg_0(X)] = \mathbb{E}[\mathbb{E}[V \mid X]g_0(X)] = 0$ since $\mathbb{E}[V \mid X] = 0$. Finally, Assumption (ref) implies $\mathbb{E}[V\varepsilon] = 0$. Rearranging yields the formula.

Causal Setup and Identification

assumption[Causal PLR and Conditional Mean Independence] There exist potential outcomes $\{Y(d): d \in \mathcal{D}\}$ rubin1974,rubin2005 such that \[ Y(d) = \theta_0 d + g_0(X) + \varepsilon, \qquad \mathbb{E}[\varepsilon \mid D, X] = 0, \] and consistency holds: $Y = Y(D)$.
remark[Causal interpretation and orthogonality] Under Assumption (ref), $\theta_0$ is a constant causal marginal effect (a constant CATE). All results below rely on the residual orthogonality moment $\mathbb{E}[V\varepsilon]=0$, which is implied by the causal restriction $\mathbb{E}[\varepsilon\mid D,X]=0$ because $\mathbb{E}[V\varepsilon]=\mathbb{E}\{\mathbb{E}[(D-\mathbb{E}[D\mid X])\varepsilon\mid X]\} =\mathbb{E}\{\mathbb{E}[D\varepsilon\mid X]-\mathbb{E}[D\mid X]\mathbb{E}[\varepsilon\mid X]\}=0$.

Causal interpretation. The estimand $\theta_0$ is causal under standard potential-outcome conditions: consistency ($Y = Y(D)$), conditional exogeneity ($\mathbb{E}[\varepsilon \mid D, X] = 0$), and overlap ($\mathrm{Var}(D \mid X) > 0$). In the PLR framework, these conditions imply the orthogonal moment $\mathbb{E}[V \cdot (Y - g_0(X) - \theta_0(D - m_0(X)))] = 0$ chernozhukov2018dml. Our analysis takes this causal target as given and studies how finite-sample inference behaves as overlap weakens (via $\sigma_V$) and $\kappa$ grows. Thus, $\kappa$ measures inferential fragility, not identification validity.

We distinguish (i) strong overlap (fixed-$\kappa$) and (ii) weakening overlap (triangular array):

assumption[Strong Overlap] There exists $\underline{\sigma}^2 > 0$ such that $\mathrm{Var}(D \mid X = x) \geq \underline{\sigma}^2$ for $P_X$-almost all $x$. This is the continuous-treatment analogue of the positivity condition rosenbaumrubin1983. damour2021overlap establish that such strict overlap assumptions become more restrictive as covariate dimension grows.
assumption[Bounded Treatment Variance] $\sigma_{D,n}^2 \asymp 1$; i.e., $0 < c \leq \sigma_{D,n}^2 \leq C < \infty$ for some constants $c, C$ and all $n$.
remark[Bounded $\kappa$ under Strong Overlap] Assumption (ref) implies $\sigma_{V,n}^2 \geq \underline{\sigma}^2 > 0$. Under bounded treatment variance (Assumption (ref)), it follows that $\kappa_n \leq (\sup_n \sigma_{D,n}^2)/\underline{\sigma}^2 < \infty$, so $\kappa_n = O(1)$.
remark[Relation to Positivity / Overlap in the Binary Case] If $D \in \{0,1\}$ and $e(X) := \mathbb{P}(D = 1 \mid X)$, then \[ \mathrm{Var}(D \mid X) = e(X)\{1 - e(X)\}, \qquad \sigma_V^2 = \mathbb{E}[e(X)\{1 - e(X)\}]. \] Thus, the usual positivity condition $\epsilon \leq e(X) \leq 1 - \epsilon$ implies $\mathrm{Var}(D \mid X) \geq \epsilon(1 - \epsilon)$, which is a special case of Assumption (ref).

To analyze regimes where $\kappa$ may grow, we introduce a triangular array framework:

assumption[Triangular Array with Weakening Overlap] Consider a sequence of DGPs indexed by $n$. There exists a sequence $\underline{\sigma}_n^2 \downarrow 0$ such that for each $n$: \[ \mathrm{Var}(D_n \mid X_n = x) \geq \underline{\sigma}_n^2 \quad \text{for } P_{X,n}\text{-almost all } x. \] Define $\sigma_{V,n}^2 := \mathbb{E}[\mathrm{Var}(D_n \mid X_n)]$ and $\kappa_n := \sigma_{D,n}^2/\sigma_{V,n}^2$.
remark[Reconciling Fixed and Growing $\kappa$] Assumption (ref) covers the fixed-$\kappa$ case (bounded condition number). Assumption (ref) covers the growing-$\kappa$ case: as $\underline{\sigma}_n^2 \to 0$, we may have $\sigma_{V,n}^2 \to 0$ and thus $\kappa_n \to \infty$. The rate $\kappa_n = O(n^\gamma)$ for $\gamma \geq 0$ determines the conditioning regime. When presenting results under Assumption (ref), $\kappa$ is treated as fixed; when presenting asymptotic regime analysis, Assumption (ref) applies.

Define $\ell_0(X) := \mathbb{E}[Y \mid X]$ and outcome residual $U := Y - \ell_0(X)$. Under Assumption (ref), $\ell_0(X) = \theta_0 m_0(X) + g_0(X)$.

lemma[Residual Decomposition] Under PLR, $U = \theta_0 V + \varepsilon$.
proofBy definition, $U = Y - \ell_0(X)$. Substituting $Y = \theta_0 D + g_0(X) + \varepsilon$ and $\ell_0(X) = \theta_0 m_0(X) + g_0(X)$: \begin{align*} U &= [\theta_0 D + g_0(X) + \varepsilon] - [\theta_0 m_0(X) + g_0(X)] \\ &= \theta_0 D - \theta_0 m_0(X) + \varepsilon \\ &= \theta_0(D - m_0(X)) + \varepsilon = \theta_0 V + \varepsilon. \qedhere \end{align*}
remark[Interpretation] Lemma (ref) shows the outcome residual equals the causal effect times the treatment residual plus noise. The treatment residual $V$ contains all identifying variation. The precision of identifying $\theta_0$ depends on $\mathrm{Var}(V) = \sigma_V^2$.

Variance Components and the Condition Number

assumption[Second Moments] $\mathbb{E}[D^2] < \infty$ and $\mathbb{E}[Y^2] < \infty$ for all $n$.

Define the variance components:

align[align omitted — 208 chars of source]
lemma[Law of Total Variance] The law of total variance yields $\sigma_D^2 = \sigma_V^2 + \sigma_m^2$.
proofBy the law of total variance: \begin{align*} \mathrm{Var}(D) &= \mathbb{E}[\mathrm{Var}(D \mid X)] + \mathrm{Var}(\mathbb{E}[D \mid X]) \\ &= \mathbb{E}[(D - m_0(X))^2] + \mathrm{Var}(m_0(X)) \\ &= \sigma_V^2 + \sigma_m^2. \qedhere \end{align*}

The population $R^2$ for treatment explained by covariates is:

equation[equation omitted — 112 chars of source]
definition[Condition Number and $R^2$ Representation] The condition number is: \begin{equation} \kappa := \frac{\sigma_D^2}{\sigma_V^2} = \frac{1}{1 - R^2(D\mid X)}. \end{equation} The empirical analogue uses cross-fitted residuals and defines the sample ratio: \[ \widehat{\sigma}_V^2 := \frac{1}{n}\sum_{i=1}^n \widehat{V}_i^2, \quad \widehat{\sigma}_D^2 := \frac{1}{n}\sum_{i=1}^n (D_i - \bar{D})^2, \quad \bar{D} := \frac{1}{n}\sum_{i=1}^n D_i, \] and $\widehat{\kappa} := \widehat{\sigma}_D^2/\widehat{\sigma}_V^2$.
remark[VIF Connection] The condition number $\kappa = 1/(1 - R^2(D \mid X))$ has the same functional form as the classical Variance Inflation Factor belsley1980diagnostic, and reduces to the classical VIF when $R^2(D \mid X)$ is interpreted as the $R^2$ from the linear projection of $D$ on $X$. Our $R^2(D \mid X)$ uses the nonparametric conditional mean $m_0(X) = \mathbb{E}[D \mid X]$, making $\kappa$ a nonparametric generalization.
remark[Properties of $\kappa$] By Lemma (ref), $\sigma_D^2 \geq \sigma_V^2$, so $\kappa \geq 1$ with equality when $m_0(X)$ is constant (treatment is unpredictable). As $R^2(D \mid X) \to 1$, we have $\sigma_V^2 \to 0$ and $\kappa \to \infty$.
remark[Binary-Treatment Specialization of $\kappa$] If $D \in \{0,1\}$ with $\mathbb{P}(D = 1) = p$, then $\sigma_D^2 = p(1-p)$ and $\sigma_V^2 = \mathbb{E}[e(X)\{1 - e(X)\}]$, so \[ \kappa = \frac{p(1-p)}{\mathbb{E}[e(X)\{1 - e(X)\}]}, \] which grows as overlap weakens.

Cross-Fitting and the DML Estimator

definition[Cross-Fitting] A $K$-fold partition $\{I_1, \ldots, I_K\}$ of $\{1,\ldots,n\}$ with disjoint folds. For $i \in I_k$, nuisance estimates $\widehat{m}^{(-k)}, \widehat{\ell}^{(-k)}$ are trained on $\{W_j : j \notin I_k\}$. We take $K$ fixed as $n \to \infty$. This sample-splitting strategy originates in schick1986 and is central to modern semiparametric estimation chernozhukov2018dml. Throughout, we assume the relevant second moments exist under $P_n$ so that $\sigma_{D,n}^2$, $\sigma_{V,n}^2$, and $L^2$ norms are well-defined.

Cross-fitted residuals: $\widehat{V}_i := D_i - \widehat{m}^{(-k)}(X_i)$, $\widehat{U}_i := Y_i - \widehat{\ell}^{(-k)}(X_i)$ for $i \in I_k$. Define errors:

align[align omitted — 165 chars of source]
remark[Residual Decomposition] With these sign conventions, the cross-fitted residuals decompose as: \[ \widehat{V}_i = D_i - \widehat{m}^{(-k)}(X_i) = (D_i - m_0(X_i)) + (m_0(X_i) - \widehat{m}^{(-k)}(X_i)) = V_i + \Delta_i^m, \] and similarly $\widehat{U}_i = U_i + \Delta_i^\ell$. This decomposition is central. Estimated residuals equal true residuals plus nuisance error.

The PLR score is:

equation[equation omitted — 102 chars of source]

where $\eta = (\ell, m)$. At true values, $\psi(W; \theta_0, \eta_0) = V\varepsilon$.

lemma[Neyman Orthogonality] The pathwise derivative of $\mathbb{E}[\psi(W; \theta_0, \eta)]$ with respect to $\eta$ vanishes at $\eta_0$. This property is the cornerstone of debiased machine learning chernozhukov2018dml.

The proof is in Appendix A.2. Neyman orthogonality implies the first-order effect of nuisance perturbations on the population moment vanishes. Combined with cross-fitting, the leading sample terms involving $\Delta^m, \Delta^\ell$ are mean-zero and of order $(r_n^m + r_n^\ell)/\sqrt{n}$, while the systematic remainder is second order (e.g., $r_n^m r_n^\ell$ and $(r_n^m)^2$).

definition[DML Estimator] \begin{equation} \widehat{\theta} := \frac{\sum_{i=1}^n \widehat{V}_i \widehat{U}_i}{\sum_{i=1}^n \widehat{V}_i^2}. \end{equation}
definition[Empirical Score Map] For any $(\theta, \eta) = (\theta, \ell, m)$ define \[ \Psi_n(\theta, \eta) := \frac{1}{n}\sum_{i=1}^n (D_i - m(X_i))\{Y_i - \ell(X_i) - \theta(D_i - m(X_i))\}. \] With cross-fitting, $\Psi_n(\theta, \widehat{\eta})$ is computed using fold-specific $\widehat{\ell}^{(-k)}, \widehat{m}^{(-k)}$ for $i \in I_k$.

The Score Jacobian

definition[Empirical Jacobian] \begin{equation} \widehat{J}_\theta := \frac{\partial}{\partial \theta}\left[\frac{1}{n}\sum_i \widehat{V}_i(\widehat{U}_i - \theta\widehat{V}_i)\right] = -\frac{1}{n}\sum_i \widehat{V}_i^2 = -\widehat{\sigma}_V^2. \end{equation}
remark[Jacobian Interpretation] The Jacobian magnitude $|\widehat{J}_\theta| = \widehat{\sigma}_V^2$ measures the score's curvature. In other words, it measures how quickly the score changes as $\theta$ varies. When $\widehat{\sigma}_V^2$ is small, the score is nearly flat in the $\theta$ direction, and small score perturbations cause large parameter shifts. This is the classical numerical-analysis insight that condition numbers govern error propagation golubvanloan2013.
lemma[Jacobian-Kappa Relationship] \begin{equation} |\widehat{J}_\theta|^{-1} = \frac{\widehat{\kappa}}{\widehat{\sigma}_D^2}. \end{equation}
proofFrom Definition (ref), $|\widehat{J}_\theta| = \widehat{\sigma}_V^2$. From Definition (ref), $\widehat{\kappa} = \widehat{\sigma}_D^2/\widehat{\sigma}_V^2$. Therefore: \[ |\widehat{J}_\theta|^{-1} = \frac{1}{\widehat{\sigma}_V^2} = \frac{\widehat{\kappa}}{\widehat{\sigma}_D^2}. \qedhere \]
remark[] When we invert the Jacobian to solve for estimator error, the condition number $\widehat{\kappa}$ appears as a scale-invariant amplification factor. This is the mechanism through which ill-conditioning affects inference.

Connection to the Riesz Representer

definition[Minimal-Norm Representer] For the PLR moment, we define the minimal-norm representer $\alpha_0: \mathcal{W} \to \mathbb{R}$ as the unique function satisfying: \begin{enumerate}[label=(\roman*)] • $\mathbb{E}[\alpha_0(W) \cdot f(X)] = 0$ for all $f \in L^2(P_X)$; • $\mathbb{E}[\alpha_0(W) \cdot D] = 1$; • $\|\alpha_0\|_{L^2}^2 = \min\{\|\alpha\|_{L^2}^2 : \alpha \text{ satisfies (i)--(ii)}\}$. \end{enumerate} This is the Riesz representer for the functional $\theta \mapsto \mathbb{E}[\alpha(W) \cdot \theta D]$ restricted to the space orthogonal to $L^2(P_X)$, expressed as a constrained minimal-norm problem chernozhukov2022riesz.
theorem[Condition Number as Riesz Norm] In the PLR model, $\alpha_0(W) = V/\sigma_V^2$. Also: \begin{equation} \kappa = \sigma_D^2 \|\alpha_0\|_{L^2}^2, \qquad \|\alpha_0\|_{L^2}^2 = \sigma_V^{-2}. \end{equation}
proofWe show $\alpha_0(W) = V/\sigma_V^2$ satisfies Definition (ref) and is the unique minimizer. Step 1 (Orthogonality): For any $f \in L^2(P_X)$: \begin{align*} \mathbb{E}[\alpha_0(W) \cdot f(X)] &= \frac{1}{\sigma_V^2}\mathbb{E}[V \cdot f(X)] \\ &= \frac{1}{\sigma_V^2}\mathbb{E}[\mathbb{E}[V \mid X] \cdot f(X)] = 0, \end{align*} since $\mathbb{E}[V \mid X] = 0$ by construction. Step 2 (Normalization): \begin{align*} \mathbb{E}[\alpha_0(W) \cdot D] &= \frac{1}{\sigma_V^2}\mathbb{E}[V \cdot (m_0(X) + V)] \\ &= \frac{1}{\sigma_V^2}(\mathbb{E}[V m_0(X)] + \mathbb{E}[V^2]) = \frac{\sigma_V^2}{\sigma_V^2} = 1. \end{align*} Step 3 (Norm computation): \[ \|\alpha_0\|_{L^2}^2 = \mathbb{E}\left[\frac{V^2}{\sigma_V^4}\right] = \frac{\sigma_V^2}{\sigma_V^4} = \frac{1}{\sigma_V^2}. \] Step 4 (Minimality and uniqueness): Let $\alpha$ be any function satisfying (i)--(ii). Define $h := \alpha - \alpha_0$. Then $h$ satisfies: \begin{itemize} • $\mathbb{E}[h(W) f(X)] = 0$ for all $f \in L^2(P_X)$ (orthogonality inherited); • $\mathbb{E}[h(W) D] = 0$ (normalization difference). \end{itemize} Since $D = m_0(X) + V$ and $\mathbb{E}[h(W) m_0(X)] = 0$ by orthogonality, we have $\mathbb{E}[h(W) V] = 0$. Now, $\alpha_0 = V/\sigma_V^2$ is a scalar multiple of $V$, so $\mathbb{E}[h \cdot \alpha_0] = \mathbb{E}[h V]/\sigma_V^2 = 0$. By the Pythagorean identity: \[ \|\alpha\|_{L^2}^2 = \|\alpha_0 + h\|_{L^2}^2 = \|\alpha_0\|_{L^2}^2 + \|h\|_{L^2}^2 \geq \|\alpha_0\|_{L^2}^2, \] with equality if and only if $h = 0$ a.s. Thus $\alpha_0$ is the unique minimizer. Step 5 (Final result): \[ \sigma_D^2 \cdot \|\alpha_0\|_{L^2}^2 = \sigma_D^2 \cdot \frac{1}{\sigma_V^2} = \kappa. \qedhere \]
remark[Semiparametric Interpretation] Theorem (ref) grounds our diagnostic in modern semiparametric theory newey1990semipar, kennedy2016semipar,chernozhukov2022riesz. The minimal-norm representer $\alpha_0$ is the correction weight for double robustness. Its norm measures how much correction is required. A large $\|\alpha_0\|_{L^2}^2$ signals both large variance and large sensitivity to nuisance bias.
remark[Hilbert-Space View] Constraint (i) is equivalent to $\mathbb{E}[\alpha(W) \mid X] = 0$, i.e.\ $\alpha$ lies in the orthogonal complement of $L^2(P_X)$ inside $L^2(P)$. The minimization in Definition (ref) therefore selects the minimum-$L^2$ instrument in the residual variation direction.

The Exact Finite-Sample Decomposition

Define the oracle sampling term:

equation[equation omitted — 78 chars of source]
lemma[Bias Decomposition] Define the nuisance bias term $B_n := \Psi_n(\theta_0,\widehat{\eta})-S_n$. Then \begin{equation} B_n = B_n^{(1)}+B_n^{(2)}+B_n^{(3)}+B_n^{(4)}+B_n^{(5)}, \end{equation} with components \begin{equation} \begin{aligned} B_n^{(1)} &:= \frac{1}{n}\sum_i V_i \Delta_i^\ell, \qquad B_n^{(2)} := -\theta_0 \frac{1}{n}\sum_i V_i \Delta_i^m, \qquad B_n^{(3)} := \frac{1}{n}\sum_i \Delta_i^m \varepsilon_i, \\ B_n^{(4)} &:= \frac{1}{n}\sum_i \Delta_i^m \Delta_i^\ell, \qquad B_n^{(5)} := -\theta_0 \frac{1}{n}\sum_i (\Delta_i^m)^2. \end{aligned} \end{equation}
proofExpand $\widehat{V}_i(\widehat{U}_i - \theta_0\widehat{V}_i)$ using $\widehat{V}_i = V_i + \Delta_i^m$ and $\widehat{U}_i = U_i + \Delta_i^\ell = \theta_0 V_i + \varepsilon_i + \Delta_i^\ell$: \begin{align*} \widehat{V}_i(\widehat{U}_i - \theta_0\widehat{V}_i) &= (V_i + \Delta_i^m)[(\theta_0 V_i + \varepsilon_i + \Delta_i^\ell) - \theta_0(V_i + \Delta_i^m)] \\ &= (V_i + \Delta_i^m)[\varepsilon_i + \Delta_i^\ell - \theta_0\Delta_i^m]. \end{align*} Expanding: \begin{align*} &= V_i\varepsilon_i + V_i\Delta_i^\ell - \theta_0 V_i\Delta_i^m \\ &\quad + \Delta_i^m\varepsilon_i + \Delta_i^m\Delta_i^\ell - \theta_0(\Delta_i^m)^2. \end{align*} Averaging over $i$ and subtracting $S_n = n^{-1}\sum_i V_i\varepsilon_i$ yields the five terms. \qedhere
remark[Interpretation of Bias Components] Terms $B_n^{(1)}$--$B_n^{(3)}$ are “first-order” in nuisance error: under cross-fitting, they have conditional mean zero and contribute $O_P(n^{-1/2})$ variance. Term $B_n^{(4)}$ is the product term driving the product-rate requirement. Term $B_n^{(5)}$ is always negative (for $\theta_0 > 0$) and scales as $(r_n^m)^2$.
theorem[Exact Decomposition] Under PLR, the DML estimator satisfies: \begin{equation} \widehat{\theta} - \theta_0 = \widehat{\kappa}\,(S_n' + B_n'). \end{equation} where $S_n' := S_n/\widehat{\sigma}_D^2$ and $B_n' := B_n/\widehat{\sigma}_D^2$. This is an exact algebraic identity. It does not rely on any Taylor expansion or linearization.
proofStep 1 (Solve for estimator): The DML estimator solves $\Psi_n(\widehat{\theta}, \widehat{\eta}) = 0$. Since the score is affine in $\theta$: \[ \Psi_n(\theta, \widehat{\eta}) = \frac{1}{n}\sum_i \widehat{V}_i\widehat{U}_i - \theta \cdot \widehat{\sigma}_V^2. \] Setting $\Psi_n = 0$: $\widehat{\theta} = \frac{1}{n}\sum_i \widehat{V}_i\widehat{U}_i / \widehat{\sigma}_V^2$. Step 2 (Express error): Subtracting $\theta_0$: \[ \widehat{\theta} - \theta_0 = \frac{\frac{1}{n}\sum_i \widehat{V}_i(\widehat{U}_i - \theta_0\widehat{V}_i)}{\widehat{\sigma}_V^2} = \frac{\Psi_n(\theta_0, \widehat{\eta})}{\widehat{\sigma}_V^2}. \] Step 3 (Decompose score): By Lemma (ref): \[ \Psi_n(\theta_0, \widehat{\eta}) = S_n + B_n. \] Step 4 (Standardize): \begin{align*} \widehat{\theta} - \theta_0 &= \frac{S_n + B_n}{\widehat{\sigma}_V^2} \\ &= \frac{\widehat{\sigma}_D^2}{\widehat{\sigma}_V^2} \cdot \frac{S_n + B_n}{\widehat{\sigma}_D^2} \\ &= \widehat{\kappa}(S_n' + B_n'). \qedhere \end{align*}
remark[Exactness of the decomposition] The decomposition is exact because the PLR score is affine in $\theta$. There is no Taylor expansion, hence no remainder term. This exactness holds at any sample size. Appendix A.1 records the corresponding generic identity for cross-fitted score estimators where the score need not be affine in $\theta$.

Variance Inflation versus Bias Amplification

The exact decomposition shows that $\widehat{\kappa}$ multiplies both the oracle sampling component and the nuisance-induced bias, but the inferential consequences are different. On the one hand, variance inflation arises through the oracle term $\widehat{\kappa}\,S_n'$, since $S_n'$ scales with $\sigma_V$ and, under the efficiency bound in Theorem (ref), the relevant benchmark variance is $V_{\mathrm{eff}}=\sigma_\varepsilon^2/\sigma_V^2$. In this case, the usual standard errors track the increased sampling variability induced by limited residual treatment variation, so coverage can remain approximately nominal. On the other hand, bias amplification arises through $\widehat{\kappa}\,B_n'$, because any remaining nuisance error is multiplied by $\kappa$.

Stochastic-Order Implications of the Finite-Sample Identity

assumption[Bounded Moments] $\mathbb{E}[D^4], \mathbb{E}[Y^4] < \infty$.
assumption[Moment Bounds] There exists a constant $C < \infty$ such that: \begin{enumerate}[label=(\roman*)] • $\mathop{\mathrm{ess\,sup}}_x \mathbb{E}[V^2 \mid X = x] \leq C$; • $\mathop{\mathrm{ess\,sup}}_{d,x} \mathbb{E}[\varepsilon^2 \mid D = d, X = x] \leq C$. \end{enumerate}
remark[Role of Moment Bounds] Assumption (ref) (i) ensures that $V^2$ has uniformly bounded conditional expectation, enabling the bound $\mathbb{E}[V^2 (\Delta^\ell)^2] \leq C \|\Delta^\ell\|_{L^2}^2$. Assumption (ref) (ii) controls the oracle term variance: $\mathbb{E}[V^2 \varepsilon^2] = \mathbb{E}[V^2 \mathbb{E}[\varepsilon^2 \mid D, X]] \leq C \sigma_{V,n}^2$. Conditional homoskedasticity $\mathbb{E}[\varepsilon^2 \mid D, X] = \sigma_\varepsilon^2$ is a special case.
assumption[Nuisance Rates] $\|\widehat{m}^{(-k)} - m_0\|_{L^2} = O_P(r_n^m)$, $\|\widehat{\ell}^{(-k)} - \ell_0\|_{L^2} = O_P(r_n^\ell)$. Here $r_n^m$ and $r_n^\ell$ are deterministic rate sequences (so $\mathrm{Rem}_n$ is deterministic).
assumption[Residual-Variance Stability] $\sigma_{V,n}^2 / \widehat{\sigma}_V^2 = O_P(1)$. Equivalently, there exists $c > 0$ with $\mathbb{P}(\widehat{\sigma}_V^2 \geq c \sigma_{V,n}^2) \to 1$.
remark[Role of Variance Stability] Assumption (ref) ensures that the empirical residual variance $\widehat{\sigma}_V^2$ does not collapse relative to the population variance $\sigma_{V,n}^2$. This is needed to control the oracle term scaling $S_n/\widehat{\sigma}_V^2$. Under strong overlap with consistent estimation, this holds automatically. Under weakening overlap, it becomes a requirement.
lemma[Sufficient Condition for Variance Stability] Suppose $\mathbb{E}[D^4] < \infty$ and let $\widehat{\sigma}_V^2 := n^{-1}\sum_{i=1}^n (D_i - \widehat{m}^{(-k(i))}(X_i))^2$ denote the cross-fitted residual second moment. If \[ \|\widehat{m}^{(-k)} - m_0\|_{L^2} = o_P(\sigma_{V,n}) \quad \text{uniformly over } k, \] then $\widehat{\sigma}_V^2/\sigma_{V,n}^2 \xrightarrow{p} 1$, hence Assumption (ref) holds.
proofWrite $\widehat{V}_i = V_i + \Delta_i^m$. Then \[ n^{-1}\sum_i \widehat{V}_i^2 = n^{-1}\sum_i V_i^2 + 2n^{-1}\sum_i V_i \Delta_i^m + n^{-1}\sum_i (\Delta_i^m)^2. \] By Cauchy--Schwarz, $|n^{-1}\sum_i V_i \Delta_i^m| \leq \|V\|_n \|\Delta^m\|_n = O_P(\sigma_{V,n}) \cdot o_P(\sigma_{V,n}) = o_P(\sigma_{V,n}^2)$. Also $n^{-1}\sum_i (\Delta_i^m)^2 = \|\Delta^m\|_n^2 = o_P(\sigma_{V,n}^2)$. Finally, $n^{-1}\sum_i V_i^2 \xrightarrow{p} \sigma_{V,n}^2$ under finite fourth moments. \qed
remark[Interpretation of Sufficient Condition] Lemma (ref) shows that, under weakening overlap, variance stability is ensured when the first-stage error is small relative to the residual scale $\sigma_{V,n}$. This prevents $\widehat{\sigma}_V^2$ from collapsing due to first-stage error.
lemma[Foldwise Empirical-to-Population Norm] Let $\{I_1, \ldots, I_K\}$ be the cross-fitting folds (Definition (ref)). For each fold $k$, let $\Delta^{(-k)}: \mathcal{X} \to \mathbb{R}$ be any (possibly random) function measurable with respect to the training sigma-field generated by $\{W_j : j \notin I_k\}$. Define the fold empirical norm \[ \|\Delta^{(-k)}\|_{n,k}^2 := \frac{1}{|I_k|}\sum_{i \in I_k} \Delta^{(-k)}(X_i)^2. \] Then, conditional on the training sample for fold $k$, \[ \mathbb{E}\!\left[\|\Delta^{(-k)}\|_{n,k}^2 \mid \{W_j : j \notin I_k\}\right] = \|\Delta^{(-k)}\|_{L^2(P_X)}^2. \] Consequently, by Markov's inequality, $\|\Delta^{(-k)}\|_{n,k} = O_P(\|\Delta^{(-k)}\|_{L^2})$ uniformly over $k$. If $\Delta_i := \Delta^{(-k)}(X_i)$ for $i \in I_k$, then the full-sample empirical norm satisfies \[ \|\Delta\|_n^2 = \frac{1}{n}\sum_{k=1}^K |I_k|\, \|\Delta^{(-k)}\|_{n,k}^2, \] so in particular $\|\Delta\|_n = O_P\!\left(\max_{1 \leq k \leq K} \|\Delta^{(-k)}\|_{L^2}\right)$.
theorem[Stochastic-order bound] Under Assumptions (ref), (ref), (ref), (ref), (ref), and (ref): \begin{equation} \widehat{\theta} - \theta_0 = O_P\!\left(\frac{\sqrt{\kappa_n}}{\sqrt{n}} + \kappa_n \cdot \mathrm{Rem}_n\right). \end{equation} where the remainder term is: \begin{equation} \mathrm{Rem}_n := r_n^m r_n^\ell + (r_n^m)^2 + \frac{r_n^m + r_n^\ell}{\sqrt{n}}. \end{equation} The oracle term $O_P(\sqrt{\kappa_n}/\sqrt{n})$ follows from $O_P(1/(\sigma_{V,n}\sqrt{n}))$ and Assumption (ref).
proofFrom Theorem (ref): $\widehat{\theta} - \theta_0 = \widehat{\kappa}(S_n' + B_n')$. Oracle term: The oracle term $S_n = n^{-1}\sum_i V_i\varepsilon_i$ has conditional mean zero and \[ \mathrm{Var}(S_n) = \frac{1}{n}\mathbb{E}[V^2\varepsilon^2]. \] By Assumption (ref)(ii) and iterated expectations: \[ \mathbb{E}[V^2\varepsilon^2] = \mathbb{E}[V^2 \mathbb{E}[\varepsilon^2 \mid D, X]] \leq C \cdot \mathbb{E}[V^2] = C \sigma_{V,n}^2. \] Thus $S_n = O_P(\sigma_{V,n}/\sqrt{n})$. By Assumption (ref), $\sigma_{V,n}^2/\widehat{\sigma}_V^2 = O_P(1)$, so: \[ \widehat{\kappa} S_n' = \frac{S_n}{\widehat{\sigma}_V^2} = O_P\!\left(\frac{\sigma_{V,n}}{\sqrt{n}} \cdot \frac{1}{\sigma_{V,n}^2}\right) = O_P\!\left(\frac{1}{\sigma_{V,n}\sqrt{n}}\right). \] Bias term: By Lemma (ref), $B_n = \sum_{j=1}^5 B_n^{(j)}$. Terms $B_n^{(1)}$--$B_n^{(3)}$: Under cross-fitting, these have conditional mean zero. For $B_n^{(1)} = n^{-1}\sum_i V_i \Delta_i^\ell$, by Assumption (ref)(i) and iterated expectations: \[ \mathbb{E}[V^2 (\Delta^\ell)^2] = \mathbb{E}[(\Delta^\ell)^2 \mathbb{E}[V^2 \mid X]] \leq C \|\Delta^\ell\|_{L^2}^2 = O_P((r_n^\ell)^2). \] Thus $\mathrm{Var}(B_n^{(1)}) = O_P((r_n^\ell)^2/n)$ and $B_n^{(1)} = O_P(r_n^\ell/\sqrt{n})$. Similarly for $B_n^{(2)}, B_n^{(3)}$. Term $B_n^{(4)}$: By sample Cauchy--Schwarz: \[ |B_n^{(4)}| = \left|\frac{1}{n}\sum_i \Delta_i^m \Delta_i^\ell\right| \leq \|\Delta^m\|_n \|\Delta^\ell\|_n. \] By Lemma (ref) applied foldwise to $\Delta^{m,(-k)}(x) := m_0(x) - \widehat{m}^{(-k)}(x)$ and $\Delta^{\ell,(-k)}(x) := \ell_0(x) - \widehat{\ell}^{(-k)}(x)$, we have $\|\Delta^m\|_n = O_P(r_n^m)$ and $\|\Delta^\ell\|_n = O_P(r_n^\ell)$. Hence $|B_n^{(4)}| = O_P(r_n^m r_n^\ell)$. Term $B_n^{(5)}$: Similarly, $|B_n^{(5)}| = |\theta_0| \|\Delta^m\|_n^2 = O_P((r_n^m)^2)$. Combining: $B_n = O_P(r_n^m r_n^\ell + (r_n^m)^2 + (r_n^m + r_n^\ell)/\sqrt{n}) = O_P(\mathrm{Rem}_n)$. Therefore: $\widehat{\kappa} B_n' = O_P(\kappa_n \cdot \mathrm{Rem}_n)$. \qedhere
remark[Complete Remainder] The remainder (ref) includes $(r_n^m)^2$ from $B_n^{(5)}$, which can dominate $r_n^m r_n^\ell$ when $r_n^m \gg r_n^\ell$. The cross-fitting terms contribute $(r_n^m + r_n^\ell)/\sqrt{n}$, negligible when $r_n^m, r_n^\ell = o(1)$.
corollary[Critical Rate] Under Assumption (ref) (strong overlap) and Assumption (ref), a sufficient condition for valid $\sqrt{n}$-inference beyond oracle noise is $\mathrm{Rem}_n = o(n^{-1/2})$, equivalently $\kappa_n \cdot \mathrm{Rem}_n = o(n^{-1/2})$.

Efficiency Bound Connection

assumption[Conditional Homoskedasticity] $\mathbb{E}[\varepsilon^2 \mid D, X] = \sigma_\varepsilon^2$ almost surely.
remark[Why Condition on $(D,X)$] Assumption (ref) (conditional homoskedasticity) is stronger than the second-moment condition $\mathbb{E}[\varepsilon^2\mid X]<\infty$. It is imposed here only to obtain a clean link between conditioning and the efficiency bound, since it delivers the identity $\mathbb{E}[V^2\varepsilon^2]=\sigma_\varepsilon^2\,\sigma_V^2$.
theorem[Efficiency and Condition Number] Under Assumption (ref), the semiparametric efficiency bound is: \begin{equation} V_{\mathrm{eff}} = \frac{\sigma_\varepsilon^2}{\sigma_V^2} = \frac{\sigma_\varepsilon^2 \kappa}{\sigma_D^2}. \end{equation} Without Assumption (ref), the bound is $V_{\mathrm{eff}} = \mathbb{E}[V^2\varepsilon^2]/\sigma_V^4$. This efficiency result connects to the classical semiparametric variance bound literature newey1990semipar.
proofThe efficiency bound for moment $\mathbb{E}[\psi] = 0$ is $V_{\mathrm{eff}} = \mathbb{E}[\psi^2]/(\partial_\theta \mathbb{E}[\psi])^2$. For $\psi = V\varepsilon$: $\mathbb{E}[\psi^2] = \mathbb{E}[V^2\varepsilon^2]$ and $\partial_\theta \mathbb{E}[\psi] = -\sigma_V^2$. Under Assumption (ref): \begin{align*} \mathbb{E}[V^2\varepsilon^2] &= \mathbb{E}[\mathbb{E}[V^2\varepsilon^2 \mid D, X]] \\ &= \mathbb{E}[V^2 \mathbb{E}[\varepsilon^2 \mid D, X]] \quad ($V^2$ is $(D,X)$-measurable)\\ &= \mathbb{E}[V^2 \sigma_\varepsilon^2] = \sigma_\varepsilon^2 \sigma_V^2. \end{align*} Therefore: $V_{\mathrm{eff}} = \sigma_\varepsilon^2 \sigma_V^2 / \sigma_V^4 = \sigma_\varepsilon^2/\sigma_V^2 = \sigma_\varepsilon^2 \kappa/\sigma_D^2$. \qedhere

Conditioning Regimes and Rate Implications

The regimes below classify how overlap weakens along the triangular array. Plugging regime-specific $\kappa_n$ growth into Theorem (ref) yields the corresponding rate implications. Proofs for the regime-specific rate statements are provided in Appendix A.3. From this point, we analyze sequences of DGPs under Assumption (ref), replacing Assumption (ref) with $\mathrm{Var}(D_n \mid X_n = x) \geq \underline{\sigma}_n^2$ where $\underline{\sigma}_n^2 \downarrow 0$.

definition[Conditioning Regimes] \leavevmode \begin{enumerate}[label=(\roman*), leftmargin=*] • Well-conditioned: $\kappa_n = O(1)$. Standard $\sqrt{n}$\nobreakdash-asymptotics apply. • Moderately ill-conditioned: $\kappa_n = O(n^\gamma)$, $0<\gamma<1/2$. Slower convergence rates. • Severely ill-conditioned: $\kappa_n \asymp \sqrt{n}$. Standard $\sqrt{n}$\nobreakdash-asymptotics fail. \end{enumerate}
theorem[Rates by Regime] Suppose $\mathrm{Rem}_n = O(n^{-\alpha})$ and $\sigma_{D,n}^2 \asymp 1$: \begin{enumerate}[label=(\roman*)] • Well-conditioned: $\widehat{\theta} - \theta_0 = O_P(n^{-1/2})$. • Moderately ill-conditioned: $\widehat{\theta} - \theta_0 = O_P(n^{\gamma/2 - 1/2} + n^{\gamma-\alpha})$. • Severely ill-conditioned: $\widehat{\theta} - \theta_0 = O_P(n^{-1/4} + n^{1/2-\alpha})$. \end{enumerate}

We interpret $\widehat{\kappa}_{\mathrm{oof}}$ through qualitative conditioning regimes that guide reporting practices rather than impose strict cutoffs. In well-conditioned scenarios, the orthogonal score exhibits sufficient curvature to ensure that the estimating equation is well-posed. In this case, residual nuisance error is not substantially amplified. As conditioning deteriorates, the same nuisance error becomes increasingly magnified. Consequently, the sensitivity of cross-learner and nuisance specifications serves as direct, observable indicator of estimator fragility and should be reported. In severely ill-conditioned contexts, the sampling term may remain small while an amplified systematic component dominates the estimator. Under such circumstances, conventional confidence intervals tend to under-cover.

A recommended workflow involves computing $\widehat{\kappa}_{\mathrm{oof}}$ for each learner and reporting these values alongside a multi-learner sensitivity summary, as demonstrated in the empirical illustration. Notably, $\widehat{\kappa}_{\mathrm{oof}}$ functions as a diagnostic for the conditioning of the orthogonalized score, rather than as evidence for unconfoundedness or causal validity. When $\widehat{\kappa}_{\mathrm{oof}}$ is elevated, estimates should be regarded as fragile. Selection of a single preferred learner should be avoided. Overlap adjustments represent natural complementary responses in applied research, although a formal analysis of these remedies is beyond the scope of this paper.

Monte Carlo Evidence

We validate the theoretical predictions via a corrupted oracle design that isolates the amplification mechanism. Predictions P1--P3 follow from Theorems (ref)--(ref):

description[leftmargin=1em,labelindent=0em] • For fixed nuisance-error $\delta$, $|\text{Bias}|$ increases proportionally with $\kappa$. • $|\text{Bias}|/\text{SE}$ is approximately a function of the single index $\kappa \times \delta$. • Coverage fails when $|\text{Bias}|/\text{SE}$ increases

We generate $n = 2{,}000$ observations from the PLR model with $\theta_0 = 1$, $X \in \mathbb{R}^{10}$ drawn from $N(0, \Sigma)$ with AR(1) correlation $\Sigma_{jk} = 0.5^{|j-k|}$ and $U, \varepsilon \stackrel{\text{i.i.d.}}{\sim} N(0, 1)$. The treatment equation is $D = m_0(X) + \sigma_U U$ where $m_0(X) = \beta^\top X$ with $\beta_j = 0.7^j$ and $\sigma_U$ calibrated to target $R^2(D|X) \in {0.50, 0.75, 0.90, 0.95, 0.97, 0.99}$, yielding $\kappa \in {2, 4, 10, 20, 33, 100}$. The corrupted oracle injects multiplicative bias: $\widehat{m}(X) = m_0(X)(1 + \delta)$ and $\widehat{\ell}(X) = \ell_0(X)(1 + \delta)$ for $\delta \in {0, 0.02, 0.05, 0.10, 0.20}$.

figure[figure omitted — 487 chars of source]

Figure (ref) visualizes the bias-amplification mechanism. For fixed nuisance error magnitude, increasing conditioning (higher $\kappa$) magnifies the resulting estimation error. Panel (a) confirms Prediction P1. For each fixed $\delta$, bias increases monotonically with $\kappa$, and the approximately parallel lines across $\delta$ levels illustrate the multiplicative structure. Each unit increase in $\log \kappa$ shifts $\log|\text{Bias}|$ by roughly the same amount regardless of $\delta$. Panel (b) confirms Prediction P2. When bias is plotted against the single index $\kappa\times\delta$, the points align along a single monotone trend, and the pooled log--log regression

equation[equation omitted — 120 chars of source]

captures this organization. The estimated slope below one is consistent with the finite-sample theory. The remainder term in Theorem (ref) contains both linear and quadratic components (e.g., $r_n^m r_n^\ell+(r_n^m)^2$), and the presence of a nonzero noise floor at $\delta=0$ mechanically flattens the log--log relationship for small values of $\kappa\times\delta$. Table (ref) further reports separate elasticities, reinforcing the finite-sample amplification mechanism.

Figure (ref) shows that coverage remains near nominal in well-conditioned regimes, but collapses sharply as $\kappa$ increases for fixed $\delta$, consistent with P3. At $\delta = 0$, coverage ranges from 0.94 to 0.97 across $\kappa$ (see Table (ref) for the full grid). Table (ref) reports the structural $\kappa$ values associated with each $R^2$ regime and confirms $\kappa$ is invariant to injected bias. Coverage decreases primarily because $|\text{Bias}|/\text{SE}$ increases with $\kappa \times \delta$, not because SE inflates. This is a silent failure where CIs are narrow but miss $\theta_0$.

figure[figure omitted — 557 chars of source]

Panel (b) reveals why coverage fails: the ratio $|\text{Bias}|/\text{SE}$ crosses 1.0 precisely when coverage drops below 80%. This confirms that undercoverage is due to bias amplification, not variance inflation. Sample-size sensitivity is summarized in Table (ref).

The Monte Carlo evidence confirms all three predictions. A sign-structure experiment in Table (ref) validates the cross-term structure of Lemma (ref). Opposite-sign nuisance errors produce $3.7\times$ larger bias and near-zero coverage at high $\kappa$, confirming that asymmetric regularization is more damaging than symmetric bias. Robustness to nonlinear DGPs appears in Table (ref).

When $\kappa$ is large, small systematic nuisance errors can dominate sampling noise. In this case, bias is amplified beyond what standard errors capture. The single-index $\kappa \times \delta$ organizes when coverage collapses, making $\kappa$ a fragility predictor. Instability across learner implementations is itself informative about estimation sensitivity. For practice, the main implication is reporting-oriented rather than threshold-based. Large estimated $\kappa$ should be interpreted as a fragility flag that motivates transparent sensitivity analysis.

Empirical Application

We demonstrate our proposed reporting workflow in the LaLonde lalonde1986 benchmark, using the NSW experimental sample and the NSW--PSID observational comparison design studied by dehejiawahba1999. The purpose of this section is diagnostic and interpretation rather than causal adjudication. In particular, the observational benchmark cannot disentangle implementation fragility induced by weak residualized treatment variation (a conditioning problem) from genuine violations of causal assumptions (an identification problem). Our goal is instead to show how the out-of-fold conditioning diagnostic $\hat\kappa_{\mathrm{oof}}$ behaves in a standard applied PLR--DML pipeline. Throughout we use $K=5$ cross-fitting and run PLR--DML across six learners (OLS, Lasso, Ridge, RF, GBM, MLP).\footnote{ For simplicity, we use the same learner class for both $\widehat m(X)$ and $\widehat \ell(X)$. Mixed learner choices are straightforward but not pursued here.}

table[table omitted — 1,768 chars of source]

Table (ref) reports $\hat\theta$, standard errors, confidence intervals, the out-of-fold first-stage fit $\hat R^2_{\mathrm{oof}}(D\mid X)$ and the implied conditioning diagnostic $\hat\kappa_{\mathrm{oof}} = 1/(1-\hat R^2_+)$ with $\hat R^2_+=\max\{0,\hat R^2_{\mathrm{oof}}(D\mid X)\}$. In the experimental NSW sample, $\hat R^2_{\mathrm{oof}}(D\mid X)$ is near zero, so $\hat\kappa_{\mathrm{oof}}\approx 1$ for all learners and estimates are comparatively stable across nuisance specifications. In the observational NSW--PSID sample, more flexible learners achieve higher out-of-fold first-stage fit, implying larger $\hat\kappa_{\mathrm{oof}}$. These same specifications exhibit some larger cross-learner dispersion, including sign reversals, despite confidence interval widths that do not mechanically expand one-for-one with dispersion. This is the qualitative empirical pattern predicted by our exact PLR identity and bound.\footnote{In many observational settings $\hat R^2_{\mathrm{oof}}(D\mid X)$ can be materially higher than in Table (ref), so $\hat\kappa_{\mathrm{oof}}$ may be much larger and the amplification mechanism correspondingly more pronounced. We use the LaLonde benchmark because it admits comparison to an experimental benchmark estimate. As noted above, the observational comparison does not disentangle conditioning-driven fragility from identification failures.} As residualized treatment variation weakens, amplification can make otherwise second-order implementation differences (learner class, tuning, splitting, and regularization) economically consequential for $\hat\theta$.

figure[figure omitted — 683 chars of source]

Accordingly, we interpret $\hat\kappa_{\mathrm{oof}}$ narrowly as a learner-specific proxy for orthogonal-score conditioning. Our recommendation is procedural and reporting-oriented: (a) pre-specify a menu of reasonable nuisance learners and tuning rules (including cross-fitting and tuning conventions); (b) compute $\hat\kappa_{\mathrm{oof}}$ learner-by-learner from the same out-of-fold residuals that enter the cross-fitted score; and (c) report $\hat\kappa_{\mathrm{oof}}$ together with a compact multi-learner sensitivity summary, such as a forest plot (see Figure (ref)). When $\hat\kappa_{\mathrm{oof}}$ is elevated and dispersion is large, nominal standard errors can understate uncertainty because amplification may dominate sampling noise.

Conclusion

We show that, in PLR-DML, conditioning of the orthogonal score is a first-order amplification channel for finite-sample error. The exact identity $\hat\theta-\theta_0=\widehat{\kappa}(S_n' + B_n')$ makes the multiplicative role of the condition number explicit, and the stochastic bound $\hat\theta-\theta_0 = O_P(\sqrt{\kappa_n/n}+\kappa_n\mathrm{Rem}_n)$ yields an operational sufficiency condition: $\kappa_n\mathrm{Rem}_n=o(n^{-1/2})$. For practice, we recommend reporting $\hat\kappa_{\mathrm{oof}}:=1/\{1-\max(0,\hat R^2_{\mathrm{oof}}(D\mid X))\}$ alongside estimates. A large $\hat\kappa_{\mathrm{oof}}$ indicates fragility. In this case, narrow CIs should be interpreted cautiously and multi-learner sensitivity summaries are recommended.

Our results are exact for PLR. In more general orthogonal-score problems, the diagnostic framework extends by linearizing the score equation and tracking both (i) the score Jacobian (the stability object) and (ii) the associated linearization remainder. Appendix A.1 provides the generic orthogonal-score identity that anchors this template. Developing weak-conditioning-robust inference procedures, and extending the analysis to heterogeneous treatment effects and other orthogonal-score settings, are promising directions.

Disclosure Statement

The author has no conflicts of interest to declare.

Data Availability Statement

Replication code and data are available at \url{https://github.com/gsaco/dml-diagnostic}.

Acknowledgments

I used Grammarly for language editing. All remaining errors are my own.

thebibliography{99} \bibitem[Bang and Robins, 2005]{bangrobins2005} Bang, H., & Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics, 61(4), 962--973. \url{https://doi.org/10.1111/j.1541-0420.2005.00377.x} \bibitem[Belloni et al., 2014]{bellonietal2014highdim} Belloni, A., Chernozhukov, V., & Hansen, C. (2014). Inference on treatment effects after selection among high-dimensional controls. Review of Economic Studies, 81(2), 608--650. \url{https://doi.org/10.1093/restud/rdt044} \bibitem[Belsley et al., 1980]{belsley1980diagnostic} Belsley, D. A., Kuh, E., & Welsch, R. E. (1980). Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. John Wiley & Sons. \url{https://doi.org/10.1002/0471725153} \bibitem[Breunig et al.(2020)]{breunig2020illposed} Breunig, C., E. Mammen, and A. Simoni (2020). Ill-posed estimation in high-dimensional models with instrumental variables. Journal of Econometrics 219(1), 171--200. \url{https://doi.org/10.1016/j.jeconom.2020.04.043} \bibitem[Chernozhukov et al., 2018]{chernozhukov2018dml} Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., & Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21(1), C1--C68. \url{https://doi.org/10.1111/ectj.12097} \bibitem[Chernozhukov et al., 2021]{chernozhukov2021rieszregression} Chernozhukov, V., Newey, W. K., Quintas-Martinez, V., & Syrgkanis, V. (2021). Automatic debiased machine learning via Riesz regression. arXiv preprint arXiv:2104.14737. \url{https://doi.org/10.48550/arXiv.2104.14737} \bibitem[Chernozhukov et al., 2022]{chernozhukov2022riesz} Chernozhukov, V., Newey, W. K., & Singh, R. (2022). Automatic debiased machine learning of causal and structural effects. \emph{Econometrica}, 90(3), 967--1027. \url{https://doi.org/10.3982/ECTA18515} \bibitem[Chernozhukov et al., 2023]{chernozhukov2023simple} Chernozhukov, V., Newey, W. K., & Singh, R. (2023). A simple and general debiased machine learning theorem with finite-sample guarantees. \emph{Biometrika}, 110(1), 257--264. \url{https://doi.org/10.1093/biomet/asac033} \bibitem[Crump et al., 2009]{crumpetal2009} Crump, R. K., Hotz, V. J., Imbens, G. W., & Mitnik, O. A. (2009). Dealing with limited overlap in estimation of average treatment effects. \emph{Biometrika}, 96(1), 187--199. \url{https://doi.org/10.1093/biomet/asn055} \bibitem[D'Amour et al., 2021]{damour2021overlap} D'Amour, A., Ding, P., Feller, A., Lei, L., & Sekhon, J. (2021). Overlap in observational studies with high-dimensional covariates. \emph{Journal of Econometrics}, 221(2), 644--654. \url{https://doi.org/10.1016/j.jeconom.2019.10.014} \bibitem[Dehejia and Wahba(1999)]{dehejiawahba1999} Dehejia, R. H. and Wahba, S. (1999). Causal Effects in Nonexperimental Studies: Reevaluating the Evaluation of Training Programs. \emph{Journal of the American Statistical Association} 94(448), 1053--1062. \url{https://doi.org/10.1080/01621459.1999.10473858} \bibitem[Farrell et al., 2021]{farrellliangmisra2021} Farrell, M. H., Liang, T., & Misra, S. (2021). Deep neural networks for estimation and inference. \emph{Econometrica}, 89(1), 181--213. \url{https://doi.org/10.3982/ECTA16901} \bibitem[Frisch and Waugh, 1933]{frischwaugh1933} Frisch, R., & Waugh, F. V. (1933). Partial time regressions as compared with individual trends. \emph{Econometrica}, 1(4), 387--401. \url{https://doi.org/10.2307/1907330} \bibitem[Golub and Van Loan, 2013]{golubvanloan2013} Golub, G. H., & Van Loan, C. F. (2013). \emph{Matrix Computations} (4th ed.). Johns Hopkins University Press. \url{https://doi.org/10.56021/9781421407944} \bibitem[Hahn, 1998]{hahn1998} Hahn, J. (1998). On the role of the propensity score in efficient semiparametric estimation of average treatment effects. \emph{Econometrica}, 66(2), 315--331. \url{https://doi.org/10.2307/2998560} \bibitem[Han and McCloskey(2019)]{hanmccloskey2019singular} Han, S. and A. McCloskey (2019). Estimation and inference with a (nearly) singular Jacobian. \emph{Quantitative Economics} 10(3), 1019--1068. \url{https://doi.org/10.3982/QE989} \bibitem[Kaji(2021)]{kaji2021weakid} Kaji, T. (2021). Theory of weak identification in semiparametric models. \emph{Econometrica} 89(2), 733--763. \url{https://doi.org/10.3982/ECTA16413}. \bibitem[Kang and Schafer, 2007]{kangschafer2007} Kang, J. D. Y., & Schafer, J. L. (2007). Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data. \emph{Statistical Science}, 22(4), 523--539. \url{https://doi.org/10.1214/07-STS227} \bibitem[Kennedy, 2016]{kennedy2016semipar} Kennedy, E. H. (2016). Semiparametric theory and empirical processes in causal inference. In H. He, P. Wu, & D.-G. Chen (Eds.), \emph{Statistical Causal Inferences and Their Applications in Public Health Research} (pp. 141--167). Springer. \url{https://doi.org/10.1007/978-3-319-41259-7_8} \bibitem[Kennedy, 2024]{kennedy2023semipar} Kennedy, E. H. (2024). Semiparametric doubly robust targeted double machine learning: A review. In E. B. Laber, M. J. Meyer, B. J. Reich, & R. Wang (Eds.), \emph{Handbook of Statistical Methods for Precision Medicine} (Chap. 10, pp. 207--236). Chapman & Hall/CRC. \url{https://doi.org/10.1201/9781003216223}. \bibitem[Khan and Tamer, 2010]{khanTamer2010overlap} Khan, S., & Tamer, E. (2010). Irregular identification, support conditions, and inverse weight estimation. \emph{Econometrica}, 78(6), 2021--2042. \url{https://doi.org/10.3982/ECTA7372} \bibitem[LaLonde, 1986]{lalonde1986} LaLonde, R. J. (1986). Evaluating the econometric evaluations of training programs with experimental data. \emph{American Economic Review}, 76(4), 604--620. \bibitem[Li et al., 2018]{lietal2018overlap} Li, F., Morgan, K. L., & Zaslavsky, A. M. (2018). Balancing covariates via propensity score weighting. \emph{Journal of the American Statistical Association}, 113(521), 390--400. \url{https://doi.org/10.1080/01621459.2016.1260466} \bibitem[Moreira, 2003]{moreira2003conditional} Moreira, M. J. (2003). A conditional likelihood ratio test for structural models. \emph{Econometrica}, 71(4), 1027--1048. \url{https://doi.org/10.1111/1468-0262.00438} \bibitem[Newey, 1990]{newey1990semipar} Newey, W. K. (1990). Semiparametric efficiency bounds. \emph{Journal of Applied Econometrics}, 5(2), 99--135. \url{https://doi.org/10.1002/jae.3950050202} \bibitem[Robinson, 1988]{robinson1988plr} Robinson, P. M. (1988). Root-$N$-consistent semiparametric regression. \emph{Econometrica}, 56(4), 931--954. \url{https://doi.org/10.2307/1912705} \bibitem[Rosenbaum and Rubin, 1983]{rosenbaumrubin1983} Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. \emph{Biometrika}, 70(1), 41--55. \url{https://doi.org/10.1093/biomet/70.1.41} \bibitem[Rubin, 1974]{rubin1974} Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. \emph{Journal of Educational Psychology}, 66(5), 688--701. \url{https://doi.org/10.1037/h0037350} \bibitem[Rubin, 2005]{rubin2005} Rubin, D. B. (2005). Causal inference using potential outcomes: Design, modeling, decisions. \emph{Journal of the American Statistical Association}, 100(469), 322--331. \url{https://doi.org/10.1198/016214504000001880} \bibitem[Schick, 1986]{schick1986} Schick, A. (1986). On asymptotically efficient estimation in semiparametric models. \emph{The Annals of Statistics}, 14(3), 1139--1151. \url{https://doi.org/10.1214/aos/1176350055} \bibitem[Staiger and Stock, 1997]{staigerstock1997weakiv} Staiger, D., & Stock, J. H. (1997). Instrumental variables regression with weak instruments. \emph{Econometrica}, 65(3), 557--586. \url{https://doi.org/10.2307/2171753} \bibitem[Stock and Yogo, 2005]{stockyogo2005weakiv} Stock, J. H., & Yogo, M. (2005). Testing for weak instruments in linear IV regression. In D. W. K. Andrews & J. H. Stock (Eds.), \emph{Identification and Inference for Econometric Models: Essays in Honor of Thomas Rothenberg} (pp. 80--108). Cambridge University Press. \url{https://doi.org/10.1017/CBO9780511614491.006}