Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,086 characters · 22 sections · 68 citation commands
-2.5cm Identification and Inference with Machine-Learned Instruments
\pagenumbering{roman}
\bgroup \let\footnoterule\relax
\thispagestyle{empty}
\egroup \pagenumbering{arabic} \setcounter{page}{1}
Instrumental variables estimation increasingly runs through machine learning. To exploit many or high-dimensional instruments, a flexible first stage pools them into a single predicted signal, and a rich set of controls is partialled out nonparametrically. The resulting estimator is the partially linear IV coefficient of the “flexible IV” routines in DoubleML and ddml. It is now common in judge- and examiner-leniency, shift-share, and quarter-of-birth returns-to-schooling designs angrist2022machine, and chen2021mostly analyze it as an optimal-instrument problem. This estimator is the object of the paper. Once treatment effects are heterogeneous in a way correlated with selection, two questions about it remain open. What does the estimator identify, and are its standard errors valid? We answer both.
Two strands of the literature focus on the same estimator without connecting. The weights literature shows that under unobserved effect heterogeneity, linear IV estimands can lose their causal reading. Two-stage least squares with multiple instruments can place negative weights on local effects mogstad2021causal,mogstad2024policy, and so can covariate misspecification blandhol2022tsls,sloczynski2024interpret. That literature offers no estimation theory for a machine-learned signal. The double/debiased machine learning literature chernozhukov2018double delivers $\sqrt{N}$ inference for the partially linear IV coefficient, but only under effect homogeneity or heterogeneity in observed covariates emmenegger2021regularizing,scheidegger2025machine. The unobserved, selection-correlated heterogeneity we treat is what makes the estimand signal-weighted and the naive moment non-orthogonal, and neither feature arises in those settings. The question is therefore what a machine-learned IV estimand identifies, and how valid inference can be conducted on it.
Our first contribution answers the identification question. Write the outcome as $Y = Y(0) + X\Delta$, with $\Delta$ the heterogeneous, unobserved effect of $X$ that may be chosen based on $\Delta$ itself. For any measurable signal $g(Z,W)$ of instruments $Z$ and covariates $W$, and in particular a cross-fitted machine-learning first stage, applied work computes the partialled-out IV estimand
with covariates removed by Robinson-style partialling robinson1988root. This is the flexible “rich covariate” adjustment that blandhol2022tsls and sloczynski2024interpret show is needed to read linear IV as weakly causal. Under conditional exogeneity and relevance alone, this estimand is a signal-weighted average treatment effect (SWATE),
whose weights integrate to one. Whatever the algorithm outputs, $\beta_{IV}$ is the corresponding signal-weighted average of effects, which gives a structural meaning to an opaque first stage. These weights are defined on the effect scale and need no single-index structure. The marginal-treatment-effect weights of heckman2005structural,heckman2006understanding are instead defined on the resistance scale and require such structure. And unlike the complier-effect estimators built from imbens1994identification and angrist1995two, the object we interpret is the very signal our estimation focuses on.
The outcome model $Y = Y(0) + X\Delta$ is the correlated random coefficients (CRC) model garen1984,wooldridge1997two,heckman1998instrumental,card2001estimating. It naturally nests the canonical binary-treatment case of local average treatment effect (LATE), and the SWATE becomes a weighted combination of LATEs. It also extends the problem to continuous $X$ with linear effects. Two caveats worth noting. With continuous $X$, the linear effect form is a restriction adopted for tractability, while with binary $X$ it is vacuous. The weight function $\omega(\cdot)$ is not identified, only $\beta_{IV}$ and its convexity are, and we do not attempt to estimate the weights. Classical CRC results give conditions under which IV recovers the average effect. The SWATE characterizes what IV recovers absent those conditions, which reappear as the special case $\omega = 1$.
The weights integrate to one but need not be non-negative, which is the concern the multiple-instrument critique raises. Our second contribution is to address this concern. In a scalar threshold-crossing model they are automatically non-negative. For the optimal first stage $g_X = \mathbb{E}[X\mid Z,W]$ pooling combined instruments, we show that the vector monotonicity of goff2024vector---equivalently the actual monotonicity of mogstad2021causal---together with a positive-association condition on the instruments delivers a new sufficient condition. We call it Conditional Covariance Monotonicity (CCM), under which $\beta_{IV}$ is a convex combination with no single-index restriction on choice behavior. We share this monotone choice model with goff2024vector but take a different route to convex weights: rather than the tailored complier-parameter estimator, we add a positive-association restriction on the observable distribution of $Z$ given $W$ and read the convex signal-weighted average off the workhorse projection. The threshold-crossing model that vytlacil2002independence showed equivalent to instrument monotonicity is its single-index special case, and CCM is strictly weaker than monotone selection.
Our Interpretation holds for any signal. Estimating the optimal signal is where standard theory breaks, and our third contribution addresses this issue. The naive debiased moment, implemented in the standard DML routines such as DoubleML package in R and Python, is not Neyman-orthogonal in $g_X$ under heterogeneity. Its first-order bias consists of the component of the instrument's effect on outcomes that the linear signal fails to absorb, and the first-stage estimation error of $g_X$. Because this bias is the covariance of those two pieces, it has a geometric consequence. Rescalings and all covariate-measurable recalibrations of the signal leave the estimate unchanged, and only “shape” change aligned with that unabsorbed component moves it. This is why well-tuned learners in the standard DML routines frequently appear valid, and why that appearance cannot be relied upon.
The bias is a drift of the estimand rather than noise. The naive moment performs valid inference on the SWATE generated by the learner's own signal, $\beta_{IV}(\hat{g}_X)$. That object is a legitimate signal-weighted estimand, convex whenever CCM holds for $\hat{g}_X$. The naive procedure targets a well-defined estimand, but only a learner-dependent one. It cannot attain valid inference for the fixed target uniformly.
To recover the fixed target, we construct a CRC-robust orthogonal score, adding one new nuisance $q_Y = \mathbb{E}[Y\mid Z,W]$ for debiasing purpose. It is Neyman-orthogonal in all four nuisances and $\sqrt{N}$-consistent under a product-rate condition that tolerates a slowly learned $q_Y$. It is also the efficient influence function for the target, and identical to the naive score under homogeneity. Around it we build a practice toolkit. The first is a Hausman-type diagnostic, the studentized naive-versus-robust contrast. The second is a standard error result that explains why projection first stages corrects the point estimate but not the reported standard error, the semiparametric analogue of heterogeneity-invalid two-stage least squares standard errors kolesar2013estimation,lee2018consistent. The third is an Anderson--Rubin confidence set. The optimal instrument $g_X = \mathbb{E}[X\mid Z,W]$ is classical under homogeneity chamberlain1987asymptotic,newey1990semiparametric,belloni2012sparse, and our contribution is to identify the orthogonal and efficient moment under heterogeneity. The identification-robust set guards a different axis than the classical and many weak instrument literatures staiger1997instrumental,dufour1997some,andrews2019weak,mikusheva2022inference, namely weakness of the residual signal left after nonparametric partialling. All proofs are collected in Appendix.
The remainder of the paper is organized as follows. Section (ref) defines the estimand, proving the representation theorems, micro-founding CCM, and relating the framework to LATE. Sections (ref) and (ref) then develop the moving target that naive DML estimates and the CRC-robust score that recovers the fixed, learner-invariant one; and Section (ref) supplies the alignment diagnostic and identification-robust inference under a weak residual signal. Sections (ref)--(ref) present the Monte Carlo and empirical evidence, and Section (ref) concludes.
Let $(\Omega, \mathcal{F}, \mathbb{P})$ be a probability space. We observe an independent and identically distributed sample of random vectors $(Y, X, Z, W)$. For notational clarity, we suppress individual subscripts. Throughout, “almost surely” (a.s.) means outside a $\mathbb{P}$-null set; for a relation stated at a fixed realization of a conditioning variable---as in $\mathbb{E}[\,\cdot\mid\Delta=\delta]\ge 0$ or $\operatorname{Cov}(\,\cdot\mid W=w)\ge 0$---it means the relation holds for almost every value of that variable.
We formalize the underlying causal data-generating process within the canonical Correlated Random Coefficient (CRC) framework wooldridge1997two, heckman1998instrumental, angrist2000interpretation. Let $Y(0)$ denote the untreated potential outcome, and let $\Delta$ denote the heterogeneous, unobserved causal treatment effect. The structural outcome equation is
Following robinson1988root and the Frisch--Waugh--Lovell theorem, we partial out the covariates $W$. Let $\tilde{A} = A - \mathbb{E}[A \mid W]$ denote the conditional residual of any random variable $A$. The generalized partialled-out instrumental variable (IV) estimand is defined as
When $g$ is a fixed instrument and only the covariate projections $\mathbb{E}[Y\mid W], \mathbb{E}[X\mid W]$ are learned, the partialled-out IV moment is already Neyman-orthogonal, so the double/debiased machine-learning theory of chernozhukov2018double delivers valid $\sqrt{N}$ inference regardless of effect heterogeneity, with Theorem (ref) supplying its structural interpretation. When the signal itself is learned, with $g$ a cross-fitted first stage pooling many instruments, orthogonality breaks under essential heterogeneity. The identification results of this section hold for any measurable $g$ and cover both regimes. The estimation theory of Sections (ref)--(ref) is what the second regime requires.
Having fixed the estimand, we now state the identifying assumptions and prove the two representation theorems that give $\beta_{IV}$ its structural meaning.
Because $\beta_{IV}$ is invariant to rescaling the signal, and in particular to its sign, formally shown in Lemma (ref) below, we adopt throughout the harmless normalization $\mathbb{E}[\tilde{X}\tilde{g}]>0$. This is a convention rather than an additional restriction, as it fixes the denominator's sign so that the following discussion of non-negative weights and convex combination read literally.
Read within an effect stratum, CCM requires that the signal not, on average, push treatment the wrong way for units sharing a given effect. The failure case is a high effect subgroup whose treatment responds negatively to the signal, which can then receive negative SWATE weight. Because CCM constrains only a conditional covariance rather than individual response functions, it sits strictly below the monotonicity notions of the literature. We show in following sections that monotonicity in imbens1994identification implies it in the scalar threshold model, and vector monotonicity goff2024vector together with positive association of the instruments implies it for $g_X$ with multiple instruments.
Note that CCM is not directly testable, and $\omega(\cdot)$ is not identified. Our defense is instead micro-foundations from primitive choice behavior in Section (ref), which restricts only the observable distribution of $Z$ given $W$.
A third point concerns the denominator. As $\mathbb{E}[\tilde{X}\tilde{g}]\to 0$ the weights $\omega$ lose meaning even when every conditional covariance is well behaved, because the SWATE is a ratio whose denominator vanishes. This failure is directly related to the weak instrument regime, and it is what motivates the identification-robust confidence sets of Section (ref).
The representation draws only on conditional exogeneity and relevance. It asks nothing of the signal $g$ beyond measurability, with no monotonicity, no correct specification, and no first-stage optimality. It therefore holds for any measurable $g$, in particular a black-box, cross-fitted machine-learning first stage. This is what equips otherwise opaque machine-learning IV estimands with an exact structural meaning.
In the scalar one-instrument case this weighting scheme is not new. huntingtonklein2020instruments shows that IV under first-stage heterogeneity identifies the first stage effect weighted average $\mathbb{E}[\gamma\beta]/\mathbb{E}[\gamma]$, with $\gamma$ the individual first-stage response. This is Theorem (ref) for a single raw instrument, with his $\gamma$ in the role of our $\mathbb{E}[\tilde{X}\tilde{g}\mid\Delta]$. Theorem (ref) generalizes it to any measurable signal $g(Z,W)$. Our weights act on the structural effects $\Delta$, whereas the Heckman--Vytlacil identification invoked by chen2021mostly places weights on marginal treatment effects, convex only under a single-index structure that Theorem (ref) does not require.
Textbook just-identified IV is already covered. Taking the signal to be a fixed scalar instrument, $g = Z$, Theorem (ref) gives \[ \beta_{IV} = \frac{\mathbb{E}[\tilde{Y}\tilde{Z}]}{\mathbb{E}[\tilde{X}\tilde{Z}]} = \mathbb{E}_\Delta\big[\Delta\,\omega(\Delta)\big], \qquad \omega(\delta) = \frac{\mathbb{E}[\tilde{X}\tilde{Z}\mid\Delta=\delta]}{\mathbb{E}[\tilde{X}\tilde{Z}]}, \] the Wald/2SLS estimand with covariates partialled out.
As mentioned in the previous section, CCM is not directly testable, and its warrant is that primitive models of treatment choice imply it. It holds automatically in the single-index threshold model $X = h(g(Z), U, W)$ with $h$ weakly increasing in $g$. The map $g \mapsto \mathbb{E}[X \mid g, W, \Delta=\delta]$ is then non-decreasing and covaries non-negatively with $g$ by Chebyshev's association inequality, so CCM holds by construction. But routing the whole instrument vector through one scalar index forces choice behavior to be effectively homogeneous across instruments, which mogstad2021causal show is restrictive precisely when $Z$ collects multiple instruments. Their response is to relax single-index imbens1994identification monotonicity to a multidimensional monotone choice model. mogstad2021causal work with partial monotonicity, under which each instrument moves treatment in a common direction across individuals while that direction may itself depend on the levels of the other instruments. The stronger actual monotonicity, in which each instrument's direction is fixed globally, coincides with the vector monotonicity of goff2024vector. We adopt this assumption and show that under it CCM survives, given a positive dependence condition on the instruments, delivering non-negative SWATE weights without any single-index restriction.
Let the treatment be generated by an arbitrary measurable structural map
where $V \in \mathcal{V}$ is an unobserved, possibly infinite-dimensional vector of choice heterogeneity. This nests the single-index model $h(g(Z), U, W)$ but allows the relative responses to different coordinates of $Z$ to vary across individuals. We take the signal to be the optimal first stage \[ g_X(Z,W) := \mathbb{E}[X \mid Z, W]. \] The same projection is used for estimation in Section (ref), and write $m_X(W) := \mathbb{E}[X \mid W] = \mathbb{E}[g_X \mid W]$, so that $\tilde{g}_X = g_X - m_X$.
Assumption (ref) strengthens Assumption (ref) only by naming the selection unobservable $V$. VM imposes the vector monotonicity of goff2024vector, which is also the nonparametric form of the actual monotonicity of mogstad2021causal. It assumes each individual's choice map is coordinate-wise non-decreasing in a fixed direction. This is strictly stronger than the partial monotonicity of mogstad2021causal, yet imposes no single-index restriction on how individuals weight the instruments. The fixed direction is precisely what Proposition (ref) needs, since its covariance is signed by associating two coordinate-wise non-decreasing functions. Each instrument is accordingly sign-normalized so that its directional effect is non-decreasing, and PA must be oriented to match. PA is the multivariate generalization of non-negative dependence esary1967association. It holds automatically when the coordinates of $Z$ are conditionally independent given $W$, and more generally under affiliation milgrom1982theory or, in the Gaussian case, whenever all conditional correlations are non-negative pitt1982positively. It is stronger than non-negative pairwise correlation, because for $d \ge 2$ the inequality $\operatorname{Cov}(Z_i, Z_j \mid W) \ge 0$ does not imply association.
Two practical points attach to PA. First, unlike CCM itself it is refutable. It restricts only the observable law of $Z$ given $W$, so its pairwise implications (positive quadrant dependence of each $(Z_i,Z_j)$ given $W$) can be assessed with the data, even if full multivariate association is harder to test. Second, association is directional, so the instruments must be sign-normalized first. Orient each coordinate of $Z$ in the direction of its own first-stage effect on $X$ (estimable from $\mathbb{E}[X\mid Z,W]$) so that VM reads coordinate-wise non-decreasing, and impose PA on the oriented vector. This is a design step for the analyst rather than an extra assumption on nature.
Note that negative association is structural in fixed-pool categorical designs. In judge- or examiner-leniency settings, case-assignment balancing renders two judges' leniencies mechanically negatively associated within a court $\times$ period cell, and blocks of mutually exclusive dummies are negatively associated within block by construction. These are the configurations Remark (ref) isolates, where sign-normalization cannot help and the projection can carry negative weight. Since PA is sufficient but not necessary for CCM, its failure does not by itself imply negative weights. But the route to convexity discussed in this section is then unavailable, and a convexity guarantee falls back on the single-index homogeneity this section set out to relax.
We share the monotone choice model of goff2024vector; our departure lies not in the behavioral assumption but in what we do with it. Where goff2024vector identifies complier parameters under vector monotonicity through a tailored estimator, and mogstad2021causal characterize the sign of the raw two-stage least squares weights, we add a design restriction on the instruments and read a convex signal-weighted average directly off the workhorse projection $g_X$. Proposition (ref) thus delivers not a new estimator but a coherence: it requires no re-estimation of complier-specific parameters and no discrete-support requirement, and it gives validity for the very signal $g_X$ that Section (ref) learns.
The SWATE representation incorporates the LATE framework. When multiple instruments are combined, the aggregate two-stage least squares estimand can fracture into a “causal salad” of local effects with negative weights mogstad2021causal. Projecting the instruments into one scalar signal $g(Z,W)$ restores a single interpretable estimand. That estimand is convex only under CCM (Assumption (ref)), which the threshold-crossing structure below delivers under the extended exogeneity condition stated next.
First fix a covariate value $w$. Define the conditional first-stage covariance and, whenever it is positive, the corresponding within-stratum covariance ratio by \[ D(w):=\mathbb{E}[\tilde X\tilde g\mid W=w], \qquad \beta_{IV}(w):=\frac{\mathbb{E}[\tilde Y\tilde g\mid W=w]}{D(w)}. \] Until Corollary (ref), all expectations, probabilities, and distributions are taken under the conditional law given $W=w$, and the dependence of the objects below on $w$ is suppressed. The paper's unconditional estimand $\beta_{IV}$ is not an unweighted average of the ratios $\beta_{IV}(w)$; the corollary gives the exact denominator-weighted aggregation.
Consider a binary treatment $X\in\{0,1\}$ generated by a canonical threshold-crossing model \[ X \;=\; \mathbf{1}\{\,U \le g(Z,W)\,\}, \] where $U$ is unobserved treatment resistance. We use Assumption (ref) with the selection unobservable specialized to $V=U$. Because $g$ is measurable in $(Z,W)$, this condition implies \[ g(Z,W)\perp\!\!\!\perp\bigl(Y(0),\Delta,U\bigr)\mid W. \] Suppose the signal $g$ takes $L$ ordered values $g_1<g_2<\dots<g_L$ with $\mathbb{P}(g=g_\ell)=p_\ell$, write $\bar g:=\mathbb{E}[g]=\sum_\ell p_\ell g_\ell$, and recall $\tilde g = g-\bar g$, $\tilde X = X-\mathbb{E}[X]$. For $\ell=2,\dots,L$ define the margin-$\ell$ complier event \[ C_\ell \;:=\; \{\, g_{\ell-1} < U \le g_\ell \,\}, \] so that an individual is treated once the signal reaches level $g_\ell$ but not at $g_{\ell-1}$. Units with $U\le g_1$ are always-takers and units with $U>g_L$ are never-takers; for both groups, $X$ is constant in $g$. Write $\pi_\ell(\delta):=\mathbb{P}(C_\ell\mid\Delta=\delta)$ and $\bar\pi_\ell:=\mathbb{P}(C_\ell)$, and define \[ \mathcal L_+ := \{\ell\in\{2,\dots,L\}:\bar\pi_\ell>0\}, \qquad \mathrm{LATE}_\ell:=\mathbb{E}[\Delta\mid C_\ell],\quad \ell\in\mathcal L_+, \] where $\mathrm{LATE}_\ell$ is the local average treatment effect imbens1994identification for margin-$\ell$ compliers. Zero-probability margins are omitted throughout. Define the covariance increments \[ S_\ell \;:=\; \mathbb{E}\!\big[(g-\bar g)\,\mathbf{1}\{g\ge g_\ell\}\big] \;=\; \mathbb{P}(g\ge g_\ell)\,\big(\mathbb{E}[g\mid g\ge g_\ell]-\mathbb{E}[g]\big) \;=\; \operatorname{Cov}\!\big(\mathbf{1}\{g\ge g_\ell\},\,g\big)\;\ge\;0 , \] the inequality holding because $\mathbf{1}\{g\ge g_\ell\}$ and $g$ are both nondecreasing in $g$.
Corollary (ref) makes the covariate aggregation precise. A stratum is weighted by its conditional first-stage covariance $D(w)$ as well as its population frequency. A stratum with $D(w)=0$ contributes zero to both unconditional moments, even though its conditional ratio is undefined.
Unlike the aggregate two-stage least squares estimand, whose multiple-instrument weights can be negative mogstad2021causal, the scalar-signal projection has non-negative weights both within covariate strata and unconditionally. The non-negative covariance increments $S_\ell$ generate the threshold-model CCM kernel, so collapsing the instruments into $g$ turns the “causal salad” back into a convex combination of margin-specific local effects. The $\lambda_\ell$ are the ordered-instrument analogue of angrist1995two. Where mogstad2021causal obtain a signed decomposition of the raw 2SLS estimand and goff2024vector identify complier parameters under vector monotonicity via a tailored estimator, we characterize what the workhorse projection estimand is, and Section (ref) shows how to make its inference robust.
Part (ii) shows where the within-stratum weight comes from. Equation (ref) expresses $\omega_w$ as a $\lambda$-convex average of complier likelihood ratios, and the ratio $dF_{\Delta\mid C_\ell}/dF_\Delta$ measures how over- or under-represented an effect $\delta$ is among margin-$\ell$ compliers in that stratum, so $\omega_w(\delta)>1$ when $\delta$ is over-represented. An effect $\delta^\ast$ occurring only among always- or never-takers has $\pi_\ell(\delta^\ast)=0$, hence $\omega_w(\delta^\ast)=0$. This is the CRC generalization of LATE's always-/never-taker zero-weighting, and the formal counterpart of huntingtonklein2020instruments's scalar observation that units with $\gamma_i\approx 0$ receive vanishing weight.
Everything to this point concerns the population signal. We now show the fact that the optimal signal $g_X=\mathbb{E}[X\mid Z,W]$ must itself be learned, and that learning it breaks the standard moment.
The sample analogue of the covariance ratio is not robust to the first-order estimation error of a machine-learned signal. Under essential heterogeneity the naive moment is not Neyman-orthogonal, so plugging in a slowly converging $\hat{g}$ contaminates $\hat{\beta}_{IV}$ with regularization bias. This section formalizes that obstruction and quantifies it, showing that the naive plug-in validly estimates a moving target that drifts with the learner. Section (ref) then constructs a locally robust score in the sense of chernozhukov2022locally, which restores $\sqrt{N}$-consistent and asymptotically normal estimation of the fixed target.
The closest precursor is chen2021mostly, who learn the optimal instrument for a linear IV model by predicting treatment from instruments and covariates with cross-fitted machine learning. They establish its asymptotic normality and Chamberlain efficiency, interpreted under threshold selection as a convex average of marginal treatment effects. Their argument that the learned signal's error has no first-order effect rests on mean-independence of the structural error from the exogenous variables, which essential heterogeneity breaks. The effective residual is correlated with the instrument through selection, so the first-order error of a learned $g_X$ enters the moment and the moment is not Neyman-orthogonal for any fixed target. Their normality is validity for the learner-dependent estimand the plug-in converges to, whereas inference on the fixed $\beta_0$ requires the orthogonal score below. The estimand is thus mostly harmless for what it estimates, though not for how one infers it.
We adopt one notational convention for the whole section. The symbol $\Delta$ (standalone) always denotes the structural effect of (ref), whereas a nuisance estimation error always carries a nuisance operand, as in $\Delta l_Y := \hat{l}_Y - l_Y$ and likewise $\Delta m_X, \Delta q_Y, \Delta g_X$, collected as $\Delta\eta := \hat{\eta} - \eta_0$. The direction $\delta_g \in L_2(\sigma(Z,W))$ denotes a perturbation of the signal.
To render the estimand operational with off-the-shelf machine learning, we fix the generic signal to the structurally optimal, predictively estimable quantity
This is the projection of the (possibly high-dimensional) instrument $Z$ onto the treatment, which is what a cross-fitted first-stage learner targets. The four nuisance functions are collected in the vector $\eta = (l_Y, m_X, q_Y, g_X)$, \[
\] Note that $\mathbb{E}[g_X\mid W] = m_X$, and the residualized signal is $\tilde{g} = g_X - m_X$. The estimand (ref) can be written as
where the last equality uses $\mathbb{E}[(X - m_X)(g_X - m_X)] = \mathbb{E}[(g_X - m_X)^2]$, which holds because $\mathbb{E}[X - m_X\mid Z,W] = g_X - m_X$. As a specialization of $\beta_{IV}$, the estimand $\beta_0$ inherits the SWATE interpretation of Theorem (ref). It is the signal-weighted average of the heterogeneous effects $\Delta$, with the signal taken to be the optimal first stage $g_X$.
Three properties recommend $g_X$ as the signal. First, it is the maximal-relevance choice as the projection maximizes covariance with $\tilde{X}$. Second, $\beta_{IV}$ is invariant to rescaling and to $W$-measurable recalibration of the signal (Lemma (ref)). Third, $g_X$ is the classical optimal instrument under homoskedastic homogeneity chamberlain1987asymptotic,newey1990semiparametric.
To read this definition, define the heterogeneity residual
Definition (ref) is exactly the statement $\mathbb{P}(\mathcal{H}\neq 0) > 0$. When effects are constant, i.e., $\Delta = \beta_0$, one has $\mathbb{E}[X\Delta\mid Z,W] = \beta_0\,g_X$ and $\mathbb{E}[X\Delta\mid W] = \beta_0\,m_X$, so $\mathcal{H} = 0$. Under essential heterogeneity, $\mathcal{H}$ is the component of the instrument's effect on $\mathbb{E}[X\Delta\mid Z,W]$ that the linear signal $\beta_0(g_X - m_X)$ fails to absorb. We therefore read Definition (ref) throughout as “heterogeneity along the signal.”
The four nuisances $\eta$ are unknown and are estimated from the sample by $K$-fold cross-fitting chernozhukov2018double. Partition the observations into $K \ge 2$ folds $\mathcal{I}_1, \dots, \mathcal{I}_K$ of equal size $n_k = N/K$, with $K$ fixed as $N \to \infty$, and write $\mathcal{D}_{-k}$ for the complementary subsample of observations outside $\mathcal{I}_k$. On each fold we fit the nuisance vector $\hat{\eta}_k = (\hat{l}_{Y,k}, \hat{m}_{X,k}, \hat{q}_{Y,k}, \hat{g}_{X,k})$ using only $\mathcal{D}_{-k}$, by any machine learner (such as lasso, random forests, gradient boosting, or an ensemble), and evaluate it on the held-out fold $\mathcal{I}_k$. Its $j$-th coordinate is written $\hat{\eta}_{j,k}$. Because $\hat{\eta}_k$ is trained without the observations at which it is used, it is independent of them and may be held fixed conditional on $\mathcal{D}_{-k}$, so overfitting does not enter the moment. When no fold index is displayed, $\hat{\eta}$ refers to this out-of-fold estimator, and fold by fold nuisance error is denoted as $\Delta\eta_k := \hat{\eta}_k - \eta_0$. The realized precision of the first stage is measured by the fold-wise $L_2$ errors \[ \rho_{j,k} := \|\hat{\eta}_{j,k} - \eta_{0,j}\|_{\mathbb{P},2} \ \ (j \in \{l_Y, m_X, q_Y, g_X\}), \qquad \rho_N := \max_{j,k}\rho_{j,k}. \] The next two assumptions place regularity and rate conditions on this construction.
Part (ii) is innocuous. Truncating each fitted nuisance at $\bar{C} \ge \max_j\|\eta_{0,j}\|_\infty$ preserves cross-fitting and weakly reduces every $L_2$ error, so it can always be imposed by the analyst.
Assumption (ref)(ii) holds in particular under the single-rate $\|\hat{\eta}_{j,k} - \eta_{0,j}\|_{\mathbb{P},2} = o_{\mathbb{P}}(N^{-1/4})$ for each nuisance, which is attainable for lasso, random forests, and gradient boosting under standard sparsity or smoothness. The product form relaxes this single-rate condition. The outcome-side projections $l_Y, q_Y$ may converge arbitrarily slowly provided the treatment-side projections $m_X, g_X$ compensate, which is valuable because $q_Y = \mathbb{E}[Y \mid Z, W]$ is typically the hardest nuisance to learn.
The standard estimator in DML routines such as DoubleML package in R and Python sets the empirical average of the canonical partialled-out moment
to zero. This moment identifies $\beta_0$, since its expectation vanishes at $(\beta_0,\eta_0)$ by construction, but it is not robust to estimation error in the signal $g_X$.
With a fixed instrument and only the covariate nuisances $(l_Y,m_X)$ estimated, the partially linear IV score of chernozhukov2018double is Neyman-orthogonal irrespective of heterogeneity. The obstruction below is specific to the learned-signal variant, where $g_X=\mathbb{E}[X\mid Z,W]$ is itself estimated, and it arises only under heterogeneity along the signal in Definition (ref).
A black-box first-stage learner $\hat{g}_X$ converges at a non-parametric rate slower than $N^{-1/2}$, and Proposition (ref) shows its error enters the moment at first order. The first-order influence of the estimated signal is $\mathbb{E}[\mathcal{H}\,\Delta g_X]$. Because it is a covariance, only the component of the first-stage error aligned with $\mathcal{H}$ can move the estimate. In the following section, we will show that this first-order consequence for the estimator is a drift of the estimand rather than noise.
This section characterizes the first-order bias identified in Proposition (ref). We first describe the geometry of $\mathcal{H}$, which governs the first-stage errors that the naive moment can and cannot detect. We then show that the bias is a drift of the estimand (Theorem (ref)), so that the naive plug-in performs valid inference on a moving target. Throughout, each perturbation direction inherits the measurability of the nuisance it perturbs: $W$-measurable for $l_Y$ and $m_X$, and $(Z,W)$-measurable for $q_Y$ and $g_X$. We write $\|\cdot\| := \|\cdot\|_{\mathbb{P},2}$.
Part (iii) shows the operator of Proposition (ref) eliminates every perturbation in $\mathcal{G}$, a rescaled signal plus a $W$-measurable shift, leaving only the residual shape component $(I-\Pi_{\mathcal{G}})\delta_g$ to move the naive moment at first order. The obstruction is broader than essential heterogeneity. Decomposing $\mathcal{H} = (\bar{\Delta}_W - \beta_0)\tilde{g}_X + \tilde{\kappa}$ into an observable-heterogeneity term ($\bar{\Delta}_W := \mathbb{E}[\Delta\mid W]$) and a selection-on-gains term ($\tilde{\kappa} := \kappa - \mathbb{E}[\kappa\mid W]$, $\kappa := \operatorname{Cov}(X,\Delta\mid Z,W)$), Lemma (ref) in Appendix (ref) shows observable heterogeneity alone makes $\mathcal{H}\neq 0$.
We now formalize the drifting estimand of naive DML. We write the naive moment as
The leading term is the drift loading, and the second is a product of nuisance errors, hence higher order.
Define $A_k := \mathbb{E}\big[(Y - \hat{l}_{Y,k})(\hat{g}_{X,k} - \hat{m}_{X,k}) \mid \mathcal{D}_{-k}\big]$ and $B_k := \mathbb{E}\big[(X - \hat{m}_{X,k})(\hat{g}_{X,k} - \hat{m}_{X,k}) \mid \mathcal{D}_{-k}\big]$,
Let $\bar{D}_N := \sum_{k=1}^K (n_k/N)\, \mathbb{E}\big[\mathcal{H}\,\Delta g_{X,k} \mid \mathcal{D}_{-k}\big]$ denote the average realized drift loading, where $\Delta g_{X,k} := \hat{g}_{X,k} - g_X$.
The Fréchet derivative in (ii) equals the G\^{a}teaux derivative of the naive moment divided by $J_0$, so the “regularization bias” is, to first order, an estimand drift. The naive plug-in tracks $\beta_{IV}(\hat{g}_X)$, the SWATE of the learner's own signal, itself legitimate and convex only if CCM holds for $\hat{g}_X$. Combining parts (iii) and (iv), the naive DML estimator has the first-order expansion \[ \hat{\beta}_{\mathrm{naive}} - \beta_0 = \frac{1}{N}\sum_{i=1}^N J_0^{-1}\,\psi_{\mathrm{naive}}(O_i;\beta_0,\eta_0) \;+\; J_0^{-1}\,\mathbb{E}\big[\mathcal{H}(Z,W)\,\Delta g_X(Z,W)\big] \;+\; o_{\mathbb{P}}\big(N^{-1/2}\big), \] whose regularization bias term $J_0^{-1}\mathbb{E}[\mathcal{H}\,\Delta g_X]$ is $o_{\mathbb{P}}(N^{-1/4})$ but not, in general, $o_{\mathbb{P}}(N^{-1/2})$. Hence $\sqrt{N}(\hat{\beta}_{\mathrm{naive}}-\beta_0)$ generally has no mean-zero limit and conventional intervals for $\beta_0$ are invalid.
We now build the score whose target is the fixed, learner-invariant $\beta_0$, which is an orthogonalized moment that eliminates the first-order influence of the learned signal, restores $\sqrt{N}$ inference, attains the efficiency bound, and costs nothing under homogeneity.
Neyman-orthogonality is restored by augmenting the numerator and denominator moments with correction terms that eliminate the first-order influence of every nuisance. Define
and the CRC-robust score
Both augmentation terms carry the first-stage residual $(X - g_X)$, whose conditional mean given $(Z,W)$ is zero. They therefore leave the estimand unchanged, so that $\mathbb{E}[\psi_{\mathrm{CRC}}(O;\beta_0,\eta_0)] = 0$ still identifies $\beta_0$, while cancelling the pathwise derivative of the moment with respect to the signal $g_X$ and the new outcome projection $q_Y$.
The orthogonalization thus re-injects the first-stage residual weighted by the unabsorbed heterogeneity. Three consequences follow from the identity. First, the correction is mean zero at every $\beta$, so both scores identify $\beta_0$. Second, when $\sigma_X^2 > 0$ it vanishes when the essential heterogeneity of Definition (ref) fails, which explains why the naive score is orthogonal under homogeneity. Third, the correction is the influence contribution of the estimated signal that the naive sandwich variance omits (Proposition (ref) in Appendix (ref)).
Write $\beta_0 = \theta_Y / J_0$ with $\theta_Y = \mathbb{E}[(Y - l_Y)(g_X - m_X)]$ and $J_0 = \mathbb{E}[(g_X - m_X)^2]$. The CRC-robust score restores $\sqrt{N}$ inference on the fixed target, attains the efficiency bound, and reduces to the naive score under homogeneity. Orthogonality, together with cross-fitting at the rates of Assumption (ref), reduces the impact of nuisance estimation to an asymptotically negligible second-order remainder.
The estimation of $\hat{\beta}_{\mathrm{CRC}}$ is the standard DML problem for a user-defined linear score, which can be directly handled by the DoubleML libraries in R and Python.\footnote{Here we provide code for using CRC-robust score in R and Python:\newline\usebox{\crcRbox}\newline\usebox{\crcPYbox}} The standard library routine returns the same estimator and standard error of Theorem (ref).
DML's validity is uniform, and it must hold across the whole ball of nuisance realizations a learner of a given accuracy could return, not just at one. Cross-fitting controls the magnitude of the first-stage error but not its direction, and the practitioner never observes which realization the learner draws, so a guarantee is only as good as its worst point in the ball. Coverage that holds on average, or at a favorable realization, need not extend to the least favorable one. Lemma (ref) makes the comparison of the two scores over such sets exact.
Proposition (ref) is more than a bound. An explicit deterministic sequence $\hat{g}_X^{(N)} = g_X + N^{-\gamma}\mathcal{H}$ with $\gamma\in(1/4,1/2)$ drives naive coverage of $\beta_0$ to zero while the CRC-robust interval stays nominal (Corollary (ref), Appendix (ref)).
There are two remaining tasks for practice. First, telling whether the naive number is on target and inferring $\beta_0$ when partialling out $W$ leaves little residual signal. This section supplies an alignment diagnostic and gives an identification-robust confidence set that stays valid when the instruments are weak.
Corollary (ref) makes the practitioner's question precise. Is the realized drift negligible, or is the naive estimator answering the question about $\beta_0$? Lemma (ref) supplies a test. With $\hat{\mathcal{H}}_k(Z,W) := (\hat{q}_{Y,k} - \hat{l}_{Y,k}) - \hat{\beta}_{\mathrm{CRC}}\,(\hat{g}_{X,k} - \hat{m}_{X,k})$, define \[ \widehat{\mathcal{A}}_N := \frac{1}{N}\sum_{k=1}^{K}\sum_{i\in\mathcal{I}_k} \big(X_i - \hat{g}_{X,k,i}\big)\,\hat{\mathcal{H}}_{k,i}, \quad \widehat{\Xi}_N := \frac{1}{N}\sum_{k=1}^{K}\sum_{i\in\mathcal{I}_k} \big(X_i - \hat{g}_{X,k,i}\big)^2\,\hat{\mathcal{H}}_{k,i}^2, \quad T_N := \frac{\sqrt{N}\,\widehat{\mathcal{A}}_N}{\widehat{\Xi}_N^{1/2}} . \]
A rejection says that the learner's first-stage error is aligned with $\mathcal{H}$, so the naive number answers a different, learner-dependent question. Because the CRC-robust estimator is valid regardless, the test has power against the naive estimator's failure mode, in the specification test pattern of hausman1978specification. Note that it does not test $\mathcal{H}=0$. Unlike the CRC tests of heckman2010testing, $T_N$ tests the realized alignment of the first-stage error with $\mathcal{H}$, while $\widehat{\Xi}_N$ points to detectable heterogeneity along the signal. We recommend reporting the pair $(T_N,\widehat{\Xi}_N)$. As mentioned in the pre-testing discussion of Section (ref), the pair is an interpretive report rather than a selector between estimators.
A distinct failure concerns the standard error rather than the point estimate. Even when the naive point estimator is centered at $\beta_0$, its plug-in sandwich variance, which treats $\hat{g}_X$ as known, is generally inconsistent, because the drift loading carries a mean-zero $O_{\mathbb{P}}(N^{-1/2})$ component inherited from the training folds. Under a first-stage asymptotic linearity condition, the naive estimator is $\sqrt{N}$-normal around $\beta_0$ with variance $\Omega_S = \Omega_0^{\mathrm{nv}} + 2\,\mathbb{E}[UVea] + \mathbb{E}[e^2a^2]$, of which the sandwich captures only $\Omega_0^{\mathrm{nv}}$ (Proposition (ref)). A least-squares first stage whose dictionary spans $g_X$ and $\mathcal{H}$ is the leading case, satisfying Assumption (ref) with $a=\mathcal{H}$ (Proposition (ref)). Such projection first stages correct the point estimate implicitly, in the tradition of newey1994asymptotic, yet leave the standard error inconsistent. Outside the projection family (such as trees, boosting, heavily penalized fits) even the centering is unprotected. This is the semiparametric analogue of heterogeneity-invalid two-stage least squares standard errors kolesar2013estimation,evdokimov2018inference,lee2018consistent.
Finally, because $\psi_{\mathrm{CRC}}$ is affine in $\beta$, inference can avoid the Jacobian. This is valuable when partialling $W$ out of a strongly $W$-predictable instrument leaves little residual signal ($J_0 \approx 0$), the regime in which the Wald interval of Theorem (ref) degrades. Define
$\mathrm{AR}_N(\beta) := \sqrt{N}\,\bar{\psi}(\beta)/\hat{\sigma}(\beta)$, and the confidence set $\mathcal{C}_{1-\alpha} := \{\beta \in \mathbb{R} : |\mathrm{AR}_N(\beta)| \le z_{1-\alpha/2}\}$.
This section examines the finite-sample performance of the naive and CRC-robust estimators and of the accompanying diagnostics. The experiments have three aims: to show that the drift of Theorem (ref) distorts naive inference at conventional sample sizes; to assess the coverage and efficiency of the CRC-robust estimator, including its behavior under homogeneity; and to show the finite-sample size and power of the diagnostics of Section (ref) in designs where the target is known.
The data-generating process is a threshold-crossing CRC model, in which a binary treatment is selected on its own gain
with $Y = Y(0) + X\Delta$ as in equation (ref). The shocks $(U,\xi,e)$ are standard normal, $Z\sim\mathcal{N}(0,I_{p_Z})$, $W\sim\mathcal{N}(0,I_{p_W})$, and all five are mutually independent, so conditional exogeneity (Assumption (ref)) holds by construction. Because the selection shock $U$ enters both the treatment equation and the return, the coefficient $b$ measures essential heterogeneity in the sense of Definition (ref), with $b=0$ the homogeneous benchmark, and $a_W$ measures heterogeneity in observables through $\mathbb{E}[\Delta\mid W]$. All four nuisance functions are available in closed form. In particular, $g_X=\mathbb{E}[X\mid Z,W]=\Phi(\nu)$.
We consider two main configurations. The first, the oracle design, has $p_Z=50$ instruments whose first-stage coefficient vector $\alpha_Z$ is approximately sparse, $p_W=10$ covariates, strong essential heterogeneity ($b=2.4$), and no heterogeneity in observables ($a_W=0$). Its nuisances are evaluated at their closed-form values rather than estimated, which separates the behavior of the estimators from first-stage estimation error. An adversarial variant perturbs the signal by an error of known size and direction. The second, the learnable design, is lower-dimensional and less noisy ($p_Z=8$, $p_W=4$, $a_W=1.2$, $b=0.6$), and its nuisances are cross-fitted with the lasso and gradient boosting. The configuration is chosen so that the product-rate condition of Assumption (ref) is plausible for both learners. The two configurations imply the same target, $\beta_0\approx0.50$. For each sample size $N\in\{500,1000,2000,4000,8000\}$ we generate $R=1000$ samples. Throughout, coverage refers to the frequency with which a nominal $95\%$ confidence interval covers the fixed target $\beta_0$, and the interval is the Wald interval unless stated otherwise. Two auxiliary designs, constructed to evaluate the diagnostics of Section (ref), are described together with the corresponding results below.
Figure (ref) reports the behavior of naive inference when the first-stage error is aligned with the heterogeneity residual. Starting from the oracle design, we replace the signal with $\hat{g}_X = g_X + c_N\,\mathcal{H}$ under two schemes for $c_N$. The fast error, $c_N = 2\,N^{-1/3}$, is of the form in Corollary (ref) and drives $\mu_N$ in Corollary (ref) to infinity. The slow error, $c_N = 2\sqrt{\log p_Z/N}$, holds $\mu_N$ at a nonzero constant. The two schemes realize branches (c) and (b) of Corollary (ref), respectively. Under the fast error, coverage of the naive Wald interval falls from $0.847$ at $N=500$ to $0.681$ at $N=8000$, the scaled bias $\sqrt{N}(\hat\beta_{\mathrm{naive}}-\beta_0)$ grows from $4.69$ to $7.08$, and the alignment test of Section (ref) rejects in every replication. This pattern is consistent with the drift of Theorem (ref). Under the slow error, coverage is near $0.90$ and does not improve as $N$ grows. The CRC-robust interval has coverage between $0.918$ and $0.950$ in both cases (Theorem (ref)).
Table (ref) shows the result of the CRC-robust estimator. In the first panel effects are homogeneous, so the naive and CRC-robust scores coincide and the correction has no efficiency loss. The simulated coverage rates of the two estimators are both close to the nominal rate. The second panel reports $N$ times the mean squared error of $\hat\beta_{\mathrm{CRC}}$ about $\beta_0$, denoted $N\widehat{\mathrm{Var}}$, in the oracle design. The semiparametric bound $J_0^{-2}\Omega_0$ equals $20.63$, and the simulated values lie between $19.30$ and $21.17$ across the sample-size grid, in line with the efficiency statement of Theorem (ref)(iii). The third panel repeats the exercise in the learnable design, where the bound equals $2.86$ and the nuisances are cross-fitted. With lasso nuisances the scaled variance is near the bound from moderate $N$ onward, between $2.63$ and $3.16$ across the grid, and coverage is close to nominal at every sample size, between $0.941$ and $0.959$. Because the design has a low-dimensional sparse linear index, the lasso recovers the nuisances at a near-parametric rate, so the second-order remainder that orthogonality leaves is already negligible at $N=500$ and the Wald interval is both centered and correctly scaled. With gradient boosting the scaled variance approaches the bound from above, from $5.98$ to $3.74$, and coverage rises from $0.867$ at $N=500$ to $0.947$ at $N=8000$, the pattern implied by Theorem (ref)(iii) when a nuisance converges more slowly while Assumption (ref) holds.
Table (ref) shows the inference when partialling out $W$ leaves little residual signal, a failure mode that does not operate through bias in the point estimate. We scale down $J_0$ from $0.075$ to $0.0013$. The identification-robust set of Section (ref) has coverage $0.944$ and $0.947$ at the two values of $J_0$, whereas the plug-in Wald interval has coverage $0.949$ under the strong signal and $1.000$ under the weak one. In the weak-signal design the AR set is unbounded in $62\%$ of replications, which is the expected behavior of an identification-robust set when $J_0$ is near zero (Proposition (ref), Remark (ref)).
We revisit the quarter-of-birth (QOB) design for the returns to schooling, the natural experiment of angrist1991does and the machine-learning-first-stage setting of angrist2022machine. The instruments are numerous and individually weak. The first stage is exactly the high-dimensional problem of pooling many weak instruments into a single signal, and the returns to schooling are an example of essential heterogeneity, since individuals plausibly select schooling on their own unobserved return. The design is therefore a test of whether the machine-learned-signal estimand drifts with the first-stage learner.
We use the angrist1991does data from the 1980 U.S.\ Census\footnote{The data is downloaded from the Angrist Data Archive, https://economics.mit.edu/people/faculty/josh-angrist/angrist-data-archive}, restricted to men born 1930--1939, giving $N=329{,}509$. The outcome $Y$ is log weekly wage and the treatment $X$ is years of completed schooling. Following angrist2022machine, the excluded instruments $Z$ are the three quarter-of-birth indicators fully interacted with year and state of birth, in total $186$ raw instruments. The covariates $W$ partialled out are a saturated set of year-of-birth and state-of-birth indicators together with race, marital status, and SMSA residence. The covariate projections $l_Y = \mathbb{E}[Y\mid W]$ and $m_X = \mathbb{E}[X\mid W]$ are estimated by saturated least squares on the discrete $W$ values, which is nonparametric and consistent because $W$ takes finitely many values. The signal $g_X = \mathbb{E}[X\mid Z,W]$ and the outcome projection $q_Y = \mathbb{E}[Y\mid Z,W]$ are the two functions estimated by the machine learner. All four nuisances are $5$-fold cross-fitted, and the procedure is repeated over $11$ independent random partitions of the sample; we report the median of the $11$ estimates with the median-adjusted variance, following the median aggregation of chernozhukov2018double.
We estimate the signal with four first stages: an ordinary least-squares projection on the full instrument set (the “optimal instrument” 2SLS of chen2021mostly), the lasso, random forest, and gradient boosting. For each we compute $\hat\beta_{\mathrm{naive}}$ and $\hat\beta_{\mathrm{CRC}}$, the alignment diagnostic pair $(T_N,\widehat\Xi_N)$, and the identification-robust set $\mathcal{C}_{0.95}$.
Table (ref) and Figure (ref) report the results. The naive $\hat\beta_{\mathrm{naive}}$ estimates a learner-dependent moving SWATE $\beta_{IV}(\hat g_X)$, while $\hat\beta_{\mathrm{CRC}}$ estimates the fixed, learner-invariant $\beta_0$, so the gap between the two estimates measures the drift of the naive estimand learner by learner. The alignment diagnostic $T_N$ tests, for each learner, whether the naive estimate targets $\beta_0$ or a drifted alternative.
The diagnostic behaves as Proposition (ref)(iv) predicts, which shows $T_N$ to be the studentized contrast between the naive and robust estimates. It rejects alignment for gradient boosting ($T_N=10.6$), whose naive and robust estimates differ by $0.044$, and does not reject for the lasso ($T_N=0.0$), whose two estimates coincide near $0.11$. The random forest lies just inside the band ($T_N=1.6$). The heterogeneity measure $\widehat\Xi_N$ is largest for the random forest, so heterogeneity along the signal is most detectable for the tree learners.
The CRC-robust estimates are more dispersed across learners than the naive ones, inflated by the OLS-projection case, which is also the most instructive case. Its naive estimate, $0.083\,(0.024)$, coincides with the classic angrist1991does two-stage least squares estimate. But partialling the saturated $W$ out of the $186$ weak and mechanically collinear QOB interactions leaves little residual signal with $\widehat J_0 = 0.012$, and the identification-robust set is the whole real line. This is the weak residual signal regime of Proposition (ref) and Remark (ref), and it is the semiparametric counterpart of the bound1995problems critique of the QOB design: the unbounded $\mathcal{C}_{0.95}$ reports that $\beta_0$ is not identified by unregularized pooling of these instruments, which the naive Wald interval does not reveal. The alignment statistic indicates the same degeneracy, with its large value $T_N=-4.7$ reflecting the distance between the naive and robust estimates ($0.083$ against $-0.008$).
Regularizing the first stage restores a bounded target. The lasso, random forest, and gradient boosting all yield bounded sets, though only the two tree learners deliver a strong residual signal. The lasso is an intermediate case. Its residual signal $\widehat J_0 = 0.017$ is barely above the OLS projection's, and its naive Wald interval is too wide to be informative, yet the CRC-robust score returns a tight estimate, $0.108\,(0.012)$, inside the bounded set $[0.084,\,0.134]$. The orthogonal score and its identification-robust set thus recover an interpretable, convex signal-weighted average of returns to schooling from a first stage whose naive interval would be uninformative.
The three reports give the empirical recommendation. We report $\hat\beta_{\mathrm{CRC}}$ as the estimate, a signal-weighted average of the heterogeneous returns to schooling that is convex under covariance monotonicity (Theorem (ref)). We read $(T_N,\widehat\Xi_N)$ as interpretation rather than as a rule for selecting a preferred learner. A rejecting $T_N$ is an informative warning, but a non-rejecting $T_N$ does not certify naive validity, which cannot be verified without estimating $\mathcal{H}$. Where the residual signal is weak, we read the unbounded $\mathcal{C}_{0.95}$ as a statement about the interpretability and validity of the estimand.
We have given the machine-learned-signal IV estimand a structural meaning and a valid inference theory under correlated random coefficients. The partialled-out estimand built from any signal of the instruments is a signal-weighted average of the heterogeneous effects, convex under a covariance-monotonicity condition, which we micro-found from vector monotonicity and positive association. When the signal is learned, the standard debiased moment is not Neyman-orthogonal for the fixed target, and its first-order bias is a drift of the estimand. Therefore, naive DML validly answers a moving, learner-dependent question while inference on the fixed, learner-invariant target requires the heterogeneity-robust orthogonal score we construct, which is the efficient influence function for the target.
Our recommended estimator is $\hat{\beta}_{\mathrm{CRC}}$, which stays valid where the naive one is inconsistent even when the point estimate is centered. Alongside it we report the pair $(T_N,\widehat{\Xi}_N)$, read as interpretation and never as a selector, since pre-testing distorts coverage. When the residual signal is weak, we add the identification-robust set $\mathcal{C}_{1-\alpha}$, whose unbounded realizations indicate that the estimand's interpretability is questionable. Together, the estimator and its diagnostics allow a flexible machine-learned first stage to be used for instrumental-variables estimation while the target remains a fixed, interpretable parameter on which valid inference is available. As such first stages become routine, our results give the resulting estimands a structural meaning and a way to tell whether a reported estimate answers this fixed question or instead tracks the learner that produced it.