The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
122,323 characters
Pragmatic DML with AI-Learned Representations
\setstretch{1.25}
\begin{abstract}
Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis.
We study when this approach is valid and develop a practical framework for causal inference with learned representations.
For a broad class of estimands, an imperfect representation distorts the target causal parameter by the product of two representation errors: one in the outcome regression and one in the balancing weight (or Riesz representer).
This yields three constructive results.
First, cross-fitted double machine learning (DML) provides valid Wald inference for the representation-dependent target. When representation errors are small, the same interval covers the causal parameter, and it can even attain the semiparametric efficiency bound.
Second, fold-wise representation learning (or fine-tuning) is compatible with DML inference for the causal parameter.
To this end, we develop convex- and star-aggregation pipelines for learning and combining representations.
Third, when representation errors are substantial, we can provide interpretable sensitivity regions and root-$n$ inference for their endpoints.
In a multi-modal demand application, seven representation-specific estimates and their star aggregate all imply a negative near-unit elasticity for rank-based price response, and the result remains robust over the reported sensitivity grid.
\end{abstract}
\maketitle
\section{Introduction}
Modern datasets rarely arrive in the tidy form imagined by classical regression theory.
Covariates are often complex objects: a document describing an individual or a firm, a product image, a sequence of transactions, or a panel of high-dimensional attributes.
To use these objects in empirical analysis, researchers often compress them into learned representations---embeddings, model predictions, or vectors of features produced by large foundation models---and then use those representations as controls.
This workflow
opens the door to new possibilities for research,
and it has become increasingly prominent as embeddings and large prediction models have improved \cite{mikolov2013efficient,pennington2014glove,devlin2019bert,mann2020language,athey2019machine}.
To formalize this workflow, write $Y$ for an outcome, $D$ for a treatment or policy variable, and $Z$ for the rich covariate object, and let $X_\phi=x(Z,\phi)$ denote its learned representation.
The subsequent analysis follows a familiar pattern: one chooses a target such as an average treatment effect (ATE), fits flexible models for the relevant nuisance objects using $(D,X_\phi)$ as controls, and reports a debiased machine learning (DML) estimate and its standard error.
Despite this familiar pattern, replacing $Z$ with $X_\phi$ changes two objects at once: the regression structure used for adjustment and the population parameter around which the DML procedure is centered.
To make the discussion concrete, consider the binary-treatment ATE.
Let $g^\star(d,z)=\mathrm{E}[Y\mid D=d,Z=z]$ denote the full-information regression function and let $g_\phi(d,x)=\mathrm{E}[Y\mid D=d,X_\phi=x]$ be its representation-based counterpart.
Under unconfoundedness, the ATE satisfies $\theta^\star=\mathrm{E}[g^\star(1,Z)-g^\star(0,Z)]$.
The corresponding representation-based estimand is $\theta_\phi=\mathrm{E}[g_\phi(1,X_\phi)-g_\phi(0,X_\phi)]$.
We call $\theta_\phi$ the \emph{average predictive effect} (APE): it is the conditional average treatment contrast when the analyst conditions only on $(D,X_\phi)$.
The central question, aimed at applied work, asks when DML inference computed using $(D,X_\phi)$ can be interpreted as valid inference for the full-information target $\theta^\star$.
Put more plainly, when working with embeddings coming from a learned representation, there could be substantial information loss. Imagine compressing an image of a dog over and over again. After compressing it enough, it may no longer be recognizable as a dog. In the same way, it could be that the representation learned via some machine learning algorithm loses crucial components contributing to the true ATE, leaving the APE that is retrieved to not be valid for the purposes of inference. We therefore seek conditions under which we may call the representation \textit{information-lossless}, in a way we make precise.
A central idea we exploit is to view representation learning as a problem of omitted information.
For a large class of causal estimands, this lets us decompose the gap between the full-information and representation-based targets as a product of two representation errors: one in the outcome regression and one in the balancing weight.
Precisely, this decomposition applies to targets that are continuous linear functionals of a regression, as in \cite{chernozhukov2022riesz}; this class includes the ATE and many other causal and policy parameters.
Let $W=(D,Z)$ denote the full information and $W_\phi=(D,X_\phi)$ the information retained by the representation.
Let $g^\star$ and $g_\phi$ be the corresponding regressions, and let $\alpha^\star$ and $\alpha_\phi$ be their Riesz representers (balancing weights).
A version of the omitted-information identity in \cite{chernozhukov2024long} then gives
\begin{equation}
\theta^\star-\theta_\phi
=
\mathrm{E}\!\left[\bigl(g^\star-g_\phi\bigr)
\bigl(\alpha^\star-\alpha_\phi\bigr)\right].
\label{eq:intro-ovb}
\end{equation}
Here and below, we suppress arguments: the full-information objects are evaluated at $(D,Z)$, while their representation-based counterparts are evaluated at $(D,X_\phi)$.
Thus $g^\star-g_\phi$ measures the regression information lost through the representation, and $\alpha^\star-\alpha_\phi$ measures the corresponding loss in the balancing weight.
Here, the target shifts only to the extent that the representation distorts both objects in common directions.
Moreover, Cauchy--Schwarz gives
\[
|\theta^\star-\theta_\phi|
\le
\left\lVert g^\star-g_\phi \right\rVert_{P,2}\,
\left\lVert \alpha^\star-\alpha_\phi \right\rVert_{P,2}.
\]
The identity \eqref{eq:intro-ovb} leads to a simple applied dichotomy.
Standard DML theory delivers asymptotically normal inference for $\theta_\phi$ under familiar conditions \cite{chernozhukov2018}.
If the representation is \emph{target-adaptive} for the estimand, meaning that $\sqrt n\,|\theta^\star-\theta_\phi|=o(1)$, then the target gap is dominated by sampling noise.
In that case, the usual DML expansion around $\theta_\phi$ automatically holds around $\theta^\star$, so a standard DML interval computed using $(D,X_\phi)$ may be reported as inference for the full-information target.
This answer is deliberately target-specific: the representation need not reconstruct $Z$ globally; it need only preserve, at the relevant rate, the outcome-regression and balancing components that jointly determine the estimand.
A further refinement concerns efficiency.
Even when target adaptivity holds, the relevant asymptotic variance is that of the \emph{short} score built on $X_\phi$, which may exceed the semiparametric efficiency bound for $\theta^\star$.
If, in addition, the representation is \emph{adaptively efficient}---meaning that both the regression $g_\phi$ and the representer $\alpha_\phi$ approach their full-information counterparts---then the short score converges to the efficient influence function and inference becomes semiparametrically efficient.
When target adaptivity and adaptive efficiency hold jointly, we say the representation is \emph{information-lossless} for the target.
The mixed-bias identity also extends the analysis beyond target adaptivity, producing sensitivity bounds for $\theta^\star$ under transparent, domain-motivated restrictions on the magnitude and alignment of representation loss, and
the same DML machinery delivers inference for its endpoints.
A second practical issue is that $\phi$ is rarely a classical parameter.
Embeddings are often produced by black-box pipelines---fine-tuning a large model, selecting prompts, or training a representation jointly with a prediction head.
Our analysis treats this entire representation step as part of the first-stage learning problem and therefore avoids a differentiability assumption or first-order expansion for $\hat\phi$.
This requires that representations be tuned fold-wise, with fixed embeddings as a special case.
Furthermore, by doing this, our results can be thought of being broadly without loss of generality of the particular method for learning the embedding.
Importantly, fold-wise tuning does not require stability or differentiability of the representation parameter $\phi$ itself.
Inference instead relies on Neyman-orthogonal scores and suitable prediction rates for the resulting nuisance learners.
Section~\ref{sec:bounds} uses the same idea to develop sensitivity bounds under learned representations.
For a reader who wants to use the paper as a guide to practice, the preceding results suggest a simple workflow.
\begin{enumerate}[leftmargin=*,label=(\roman*)]
\item Compute the usual cross-fitted DML point estimate and standard error using the representation $X_\phi$ (or a fold-wise tuned $X_{\hat\phi}$).
\item Interpret this output as inference for the representation-based target, or for the corresponding full-information target $\theta^\star$ under target adaptivity.
\item If target adaptivity is not comfortable to defend, replace the single-number interpretation with a sensitivity analysis that reports an identified interval for $\theta^\star$ together with inference for its endpoints.
\end{enumerate}
All three steps can be carried out with a fixed embedding or with fold-wise fine-tuning; the fixed-embedding case is obtained by taking the representation learner to be constant.
Section~2 formalizes the adaptive interpretation, Section~3 discusses how to design regression and Riesz learners over the common information space $\mathcal H$ so that the product-rate condition is plausible, and Section~4 develops sensitivity bounds and endpoint inference.
\subsection*{Contributions and Related Work}
Our contribution is pragmatic and constructive.
First, we derive an omitted-information (mixed-bias) identity for a broad class of Riesz-linear targets, which isolates the two essential representation errors.
Second, we establish DML inference for end-to-end foldwise pipelines and for the estimable components of the resulting sensitivity bounds without requiring a first-order expansion of the representation learner.
Finally, we develop a learner-design perspective over the induced information space $\mathcal H$ and give two complementary routes---proper learning over star-shaped approximations and improper learning via star aggregation---that deliver first-stage rates compatible with the DML product-rate requirements.
We also show that the aggregation weights may be fitted to pooled out-of-fold predictions rather than inside each training fold, at the cost of a factor $\sqrt{\log M}$ for star weights and $\sqrt{M\log n}$ for convex weights in one remainder term; this removes the need for nested sample splitting.
These results connect the orthogonal-score and cross-fitting literature \cite{chernozhukov2018,belloni2014high} with the representer calculus underlying Riesz regression \cite{chernozhukov2022riesz} and with empirical work that uses text and other learned representations in causal analysis \cite{wooddoughty2018challenges,sridhar2022causal}.
A multimodal demand application illustrates how the framework can be used to compare, aggregate, and assess the sensitivity of alternative representations.
We also pursue an influence-function correction approach in a companion paper; the present paper focuses on omitted-information identities, adaptive inference, and sensitivity analysis for learned representations.
Together, the results provide a target-aware way to choose, tune, aggregate, and evaluate embeddings while retaining familiar DML inference.
\subsection*{Notation} We use the following notation.
We observe an i.i.d.\ sample $O_1,\dots,O_n$ of a generic observation $O=(Y,D,Z)$ with law $P$ on a measurable space $(\mathcal O,\mathcal A)$.
For any measurable function $f:\mathcal O\to\mathbb{R}$ such that $Pf:=\int f(o)\,dP(o)$ is finite, we write $Pf$ for its expectation.
For $q\ge 1$, we write $\left\lVert f \right\rVert_{P,q}:=(P|f|^q)^{1/q}$ and $\left\lVert f \right\rVert_\infty:=\sup_{o\in\mathcal O}|f(o)|$ when the supremum exists.
We write $a_n=o_\mathrm{P}(b_n)$ if $a_n/b_n\to 0$ in probability and $a_n=O_\mathrm{P}(b_n)$ if $a_n/b_n$ is tight; all stochastic order statements are with respect to $P$.
We use $\overset{p}{\to}$ for convergence in probability and $\Rightarrow$ for weak convergence.
These conventions follow standard usage in asymptotic statistics \cite{vanderVaart1998}.
\section{Adaptive Inference with Learned Representations}
In applications, the analyst often freezes a representation and then runs cross-fitted DML using $(D,X_\phi)$ as controls.
We start by recalling the standard DML expansion for the corresponding representation-based target $\theta_\phi$.
We then turn to the full-information target $\theta^\star$ and show how it differs from $\theta_\phi$ through an omitted-information identity.
That identity motivates two conditions---target adaptivity and adaptive efficiency---under which the representation-based DML output can be interpreted as valid and efficient inference for $\theta^\star$.
Finally, we address the practically important case where the representation is tuned using the sample.
We show that as long as tuning is confined to the training folds, the same first-order DML theory applies to the population target defined by the complete tuning and estimation procedure.
\subsection{Static setting: fixed representations}
\label{sec:warmup}
\subsubsection{Setup: full covariates and learned representations}
\label{sec:setup}
Let $O=(Y,D,Z)\sim P$, where $Y\in\mathbb{R}$ is an outcome, $D$ is a treatment or policy variable, and $Z$ is a potentially high-dimensional object (text, images, panels, or other structured data).
We observe $n$ i.i.d.\ draws $O_1,\dots,O_n$.
Let
\[
X_\phi := x(Z,\phi)\in\mathbb{R}^p
\]
be a learned representation of $Z$, where $\phi$ indexes the embedding rule via a finite (possibly very high) dimensional parameter.
Throughout Sections~\ref{sec:warmup} and \ref{sec:omitted} we condition on the embedding rule and treat it as fixed for the purpose of inference.
This view is appropriate when the embedding model is frozen (for example, a pre-trained foundation model) or when it is trained on data that are independent of the sample used for estimation and inference.
While our formalization allows for a drifting sequence $\phi=\phi_n$ of parameters, we do treat the resulting map $z\mapsto x(z,\phi_n)$ as nonrandom, at least to start the discussion.
Section~\ref{sec:swap} returns to the case where the representation is tuned or learned on the sample.
It is helpful to keep track of the two information sets:
\[
W := (D,Z)
\qquad\text{and}\qquad
W_\phi := (D,X_\phi).
\]
The full-information regression is
\[
g^\star(w) := \mathrm{E}[Y\mid W=w],
\]
and the representation-based regression is
\[
g_\phi(w_\phi) := \mathrm{E}[Y\mid W_\phi=w_\phi].
\]
Since $W_\phi$ is a function of $W$, the representation-based regression is the $L^2(P)$ projection of $g^\star(W)$ onto $\sigma(W_\phi)$, i.e.\ $g_\phi(W_\phi)=\mathrm{E}[g^\star(W)\mid W_\phi]$.
We focus on targets that can be written as continuous linear functionals of the regression.
Concretely, let $m(\cdot,\cdot)$ be measurable and linear in its second argument, and define the \emph{full-information} (``long'') target as
\begin{equation}
\theta^\star := \mathrm{E}\!\left[m\!\left(O,g^\star\right)\right].
\label{eq:theta-long}
\end{equation}
After replacing $Z$ by $X_\phi$, the corresponding \emph{representation-based} (``short'') target is
\begin{equation}
\theta_\phi := \mathrm{E}\!\left[m\!\left(O,g_\phi\right)\right].
\label{eq:theta-short}
\end{equation}
To keep this comparison meaningful, we assume the following mild compatibility condition: whenever $h$ is a square-integrable function of $W_\phi$, the random variable $m(O,h)$ is $\sigma(W_\phi)$-measurable.
In leading causal examples this holds because $m$ depends on $W$ only through the values of $h$ at observable arguments.\footnote{For example, for the ATE we have $m((y,d,z),h)=h(1,z)-h(0,z)$, and if $h$ is a function of $(d,X_\phi)$ then $m((Y,D,Z),h)$ is a function of $(D,X_\phi)$ as well.}
The class \eqref{eq:theta-long} includes many common causal and policy parameters, as previously discussed in \cite{chernozhukov2022riesz} and other references.
For instance, if $D\in\{0,1\}$ and we target the average treatment effect, then the full-information target is the ATE,
$\theta^\star=\mathrm{E}[g^\star(1,Z)-g^\star(0,Z)]$,
whereas the representation-based target is the corresponding average predictive effect (APE),
$\theta_\phi=\mathrm{E}[g_\phi(1,X_\phi)-g_\phi(0,X_\phi)]$.
These two agree (for inference) precisely in the target-adaptive regime discussed below.
Weighted average treatment effects, average potential outcomes, and many policy effects defined as averages of $g^\star(d,z)$ against a known weight are also covered.
\subsubsection{Riesz representers and an omitted-information identity}
\label{sec:identity}
To compare $\theta^\star$ and $\theta_\phi$, it is convenient to express each linear functional as an inner product.
In $L^2$ language, this inner product is provided by the Riesz representer.
Suppose the map $h\mapsto \mathrm{E}[m(O,h)]$ is a continuous linear functional on the Hilbert space $L^2(P_W)$.
Then there exists a unique $\alpha^\star\in L^2(P_W)$ such that
\begin{equation}
\mathrm{E}[m(O,h)] = \mathrm{E}[h(W)\alpha^\star(W)]
\qquad\text{for all } h\in L^2(P_W).
\label{eq:rr-long}
\end{equation}
We call $\alpha^\star$ the \emph{Riesz representer} (or balancing weight) for the full-information functional.
Likewise, under the analogous continuity condition on $L^2(P_{W_\phi})$, there exists a unique $\alpha_\phi\in L^2(P_{W_\phi})$ such that
\begin{equation}
\mathrm{E}[m(O,h)] = \mathrm{E}[h(W_\phi)\alpha_\phi(W_\phi)]
\qquad\text{for all } h \in L^2(P_{W_\phi}).
\label{eq:rr-short}
\end{equation}
This representer-based view is standard in modern treatments of debiasing for generic linear functionals; see, for example, \cite{chernozhukov2022riesz}.
In many familiar causal problems, $\alpha^\star$ and $\alpha_\phi$ reduce to recognizable weights.
For the binary-treatment ATE (full information) and its representation-based analog (APE),
\[
\alpha^\star(D,Z)=\frac{D}{e(Z)}-\frac{1-D}{1-e(Z)},
\qquad
\alpha_\phi(D,X_\phi)=\frac{D}{e_\phi(X_\phi)}-\frac{1-D}{1-e_\phi(X_\phi)},
\]
where $e(z)=\mathrm{P}(D=1\mid Z=z)$ and $e_\phi(x)=\mathrm{P}(D=1\mid X_\phi=x)$.
The next result is the key identity.
It says that the difference between the long and short targets is driven by the interaction of two approximation errors: the error from compressing $Z$ into $X_\phi$ in the outcome regression, and the corresponding error in the balancing weight.
More precisely, Proposition~\ref{prop:ovb} is the learned-representation specialization of Theorem~2 in \cite{chernozhukov2024long}.\footnote{
Their long and short objects become
\[
(\theta,\theta_s,g,g_s,\alpha,\alpha_s)
=
(\theta^\star,\theta_\phi,g^\star,g_\phi,\alpha^\star,\alpha_\phi),
\]
with the short information set $\sigma(W_\phi)$ nested in the long information set $\sigma(W)$.
Under this identification, \eqref{eq:ovb-identity} reads $\theta-\theta_s=\mathrm{E}[(g-g_s)(\alpha-\alpha_s)]$.}
\begin{proposition}[Omitted-information identity]
\label{prop:ovb}
Assume \eqref{eq:rr-long}--\eqref{eq:rr-short} hold and $Y\in L^2(P)$.
Write $g^\star=g^\star(W)$, $g_\phi=g_\phi(W_\phi)$, $\alpha^\star=\alpha^\star(W)$, and $\alpha_\phi=\alpha_\phi(W_\phi)$.
Then
\begin{equation}
\theta^\star-\theta_\phi
=
\mathrm{E}\!\left[(g^\star-g_\phi)(\alpha^\star-\alpha_\phi)\right].
\label{eq:ovb-identity}
\end{equation}
Consequently,
\begin{equation}
\bigl|\theta^\star-\theta_\phi\bigr|
\le
\left\lVert g^\star-g_\phi \right\rVert_{P,2}\;
\left\lVert \alpha^\star-\alpha_\phi \right\rVert_{P,2}.
\label{eq:cs-bound}
\end{equation}
\end{proposition}
Proposition~\ref{prop:ovb} gives a target-specific measure of representation quality.
A representation preserves the target whenever it preserves either the relevant outcome regression or the relevant balancing weight, and the quantitative bound rewards partial preservation of both.
Thus a rich object $Z$ may be compressed substantially without changing the estimand: only the product of the two representation errors matters.
\subsubsection{Debiased scores and $K$-fold cross-fitted DML}
\label{sec:dml}
We now define the debiased score and the associated $K$-fold cross-fitted DML estimator for the short target $\theta_\phi$.
This makes the downstream inference problem concrete: in practice, the analyst computes a single point estimate and a standard error from a cross-fitted score, and the question is how to interpret that output.
For a fixed embedding rule $\phi$, define the score
\begin{equation}
\psi_\phi(O;g,\alpha)
:=
m(O,g)+\alpha(W_\phi)\{Y-g(W_\phi)\}.
\label{eq:score-short}
\end{equation}
At the truth $(g,\alpha)=(g_\phi,\alpha_\phi)$ we have $\mathrm{E}[\psi_\phi(O;g_\phi,\alpha_\phi)]=\theta_\phi$, and the score is Neyman-orthogonal with respect to perturbations of $(g,\alpha)$; see \cite{chernozhukov2018,chernozhukov2022riesz}. Here $\alpha_\phi$ is the Riesz representer on $L^2(P_{W_\phi})$ introduced in Section~\ref{sec:identity}.
Orthogonality has an especially concrete implication for the score \eqref{eq:score-short}.
Because $m$ is linear in its second argument and $g_\phi$ is the conditional mean of $Y$ given $W_\phi$, the population drift of the score away from its target can be written exactly as a \emph{mixed bias} term.
For every square-integrable pair $(g,\alpha)$ that is measurable with $W_\phi$,
\begin{equation}
\mathrm{E}\bigl[\psi_\phi(O;g,\alpha)\bigr]-\theta_\phi
=
-\mathrm{E}\!\left[(g-g_\phi)(\alpha-\alpha_\phi)\right].
\label{eq:mixed-bias}
\end{equation}
Moreover, a single Cauchy--Schwarz step yields
\[
\bigl|\mathrm{E}[\psi_\phi(O;g,\alpha)]-\theta_\phi\bigr|
\le
\left\lVert g-g_\phi \right\rVert_{P,2}\,\left\lVert \alpha-\alpha_\phi \right\rVert_{P,2},
\]
so mean score drift is negligible at $\sqrt{n}$ scale under the usual product-rate condition.
This is the same algebraic pattern as in Proposition~\ref{prop:ovb}: in both cases, the relevant discrepancy is an interaction of two first-stage approximation errors.
For comparison, the full-information (``long'') score for $\theta^\star$ is
\begin{equation}
\psi^\star(O;g,\alpha)
:=
m(O,g)+\alpha(W)\{Y-g(W)\}.
\label{eq:score-long}
\end{equation}
At the truth $(g,\alpha)=(g^\star,\alpha^\star)$, the centered score $\psi^\star(O;g^\star,\alpha^\star)-\theta^\star$ is the efficient influence function for $\theta^\star$ in the nonparametric model \cite{bickel1993,vanderVaart1998,newey1994}. The representer $\alpha^\star$ is defined in Section~\ref{sec:identity}.
Fix an integer $K\ge 2$.
Let $\{I_k\}_{k=1}^K$ be a partition of $\{1,\dots,n\}$ into $K$ folds of (approximately) equal size, and write $I_k^c:=\{1,\dots,n\}\setminus I_k$. For each fold $k$, use only the training sample $\{O_i:i\in I_k^c\}$ (and the corresponding representations $W_{\phi,i}=(D_i,X_{\phi,i})$) to construct nuisance estimators
\[
\hat g_{\phi}^{(-k)},
\qquad
\hat\alpha_{\phi}^{(-k)}.
\]
The superscript $(-k)$ indicates that the fold-$k$ observations are excluded from the training step.
For each observation $i\in I_k$, evaluate the score using the fold-$k$ nuisance estimates:
\[
\hat\psi_{\phi,i}:=\psi_\phi\!\left(O_i;\hat g_{\phi}^{(-k)},\hat\alpha_{\phi}^{(-k)}\right)
\qquad\text{for } i\in I_k.
\]
The $K$-fold cross-fitted DML estimator of $\theta_\phi$ is the fold-averaged score,
\begin{equation}
\hat\theta_\phi
:=
\frac{1}{n}\sum_{k=1}^K\sum_{i\in I_k} \hat\psi_{\phi,i}.
\label{eq:dml-est}
\end{equation}
A natural standard error is based on the empirical variance of the cross-fitted scores:
\begin{equation}
\hat\sigma_\phi^2
:=
\frac{1}{n}\sum_{k=1}^K\sum_{i\in I_k}\bigl(\hat\psi_{\phi,i}-\hat\theta_\phi\bigr)^2.
\label{eq:var-est}
\end{equation}
To interpret \eqref{eq:dml-est} as a root-$n$ estimator, we need two ingredients: (i) the score must be orthogonal, and (ii) the nuisance learners must converge fast enough in $L^2(P)$ so that the product of their errors is $o_\mathrm{P}(n^{-1/2})$.
The next assumption records a standard set of sufficient conditions.
\begin{assumption}[DML regularity for the short target]
\label{ass:dml}
Fix $K\ge 2$ and consider a (possibly drifting) sequence $\phi=\phi_n$.
Let $\hat g_{\phi}^{(-k)}$ and $\hat\alpha_{\phi}^{(-k)}$ be the cross-fitted nuisance estimators defined above.
\begin{enumerate}[label=(\roman*)]
\item \emph{Moments and bounded linearity.}
There exists $q>4$ and a constant $C_m<\infty$ such that $Y\in L^q(P)$, $\sup_n \left\lVert \alpha_{\phi_n} \right\rVert_{P,q}<\infty$, and
\begin{equation}
\left\lVert m(O,h) \right\rVert_{P,2}\le C_m \left\lVert h \right\rVert_{P,2}
\qquad\text{for all } h \in L^2(P_{W_{\phi_n}}).
\label{eq:m-L2-bounded}
\end{equation}
(When $\phi_n\equiv\phi$ is fixed, the supremum over $n$ is unnecessary.)
\item \emph{Nuisance consistency and moment control.}
Uniformly over folds,
\[
\max_{1\le k\le K}\left\lVert \hat g_{\phi_n}^{(-k)}-g_{\phi_n} \right\rVert_{P,2}=o_\mathrm{P}(1),
\qquad
\max_{1\le k\le K}\left\lVert \hat\alpha_{\phi_n}^{(-k)}-\alpha_{\phi_n} \right\rVert_{P,2}=o_\mathrm{P}(1),
\]
and the nuisance estimates have uniformly bounded $q$th moments:
\[
\max_{1\le k\le K}\left\lVert \hat g_{\phi_n}^{(-k)} \right\rVert_{P,q}
+
\max_{1\le k\le K}\left\lVert \hat\alpha_{\phi_n}^{(-k)} \right\rVert_{P,q}
=
O_\mathrm{P}(1).
\]
\item \emph{Rate condition.} Uniformly over folds,
\begin{equation}
\max_{1\le k\le K}\left\lVert \hat g_{\phi_n}^{(-k)}-g_{\phi_n} \right\rVert_{P,2}\;
\max_{1\le k\le K}\left\lVert \hat\alpha_{\phi_n}^{(-k)}-\alpha_{\phi_n} \right\rVert_{P,2}
=
o_\mathrm{P}(n^{-1/2}).
\label{eq:dml-rate}
\end{equation}
\item \emph{Nondegeneracy.}
Let $\sigma_{\phi_n}^2:=\mathrm{Var}\!\left(\psi_{\phi_n}(O;g_{\phi_n},\alpha_{\phi_n})\right)$.
Assume $0<\inf_n\sigma_{\phi_n}\le \sup_n\sigma_{\phi_n}<\infty$.
\item \emph{Extra moment when drifting.} For a possibly drifting sequence $\phi=\phi_n$, assume moreover that there exists $\delta>0$ such that
\[\sup_n\left\lVert \psi_{\phi_n}(O;g_{\phi_n},\alpha_{\phi_n})-\theta_{\phi_n} \right\rVert_{P,2+\delta}<\infty.\]
When $\phi_n\equiv\phi$ is fixed, this additional uniform moment condition is unnecessary.
\end{enumerate}
\end{assumption}
Under Assumption~\ref{ass:dml}, the DML estimator admits the usual asymptotically linear representation for the representation-based target $\theta_\phi$.
The next theorem specializes the standard DML expansion to our Riesz score \eqref{eq:score-short}.
We allow for a (possibly drifting) sequence of representations $\phi=\phi_n$ (treated as nonrandom for the sampling distribution).
A self-contained proof is given in Appendix~\ref{app:proofs-sec2}; see \cite{chernozhukov2018} for the general DML framework and \cite{chernozhukov2022riesz} for the Riesz-functional setup.
\begin{theorem}[DML inference for the representation-based target $\theta_\phi$]
\label{thm:dml-expansion}
Suppose Assumption~\ref{ass:dml} holds for a (possibly drifting) sequence $\phi=\phi_n$.
Then
\begin{equation}
\sqrt{n}\bigl(\hat\theta_{\phi_n}-\theta_{\phi_n}\bigr)
=
\frac{1}{\sqrt{n}}\sum_{i=1}^n
\Bigl(\psi_{\phi_n}(O_i;g_{\phi_n},\alpha_{\phi_n})-\theta_{\phi_n}\Bigr)
+o_\mathrm{P}(1).
\label{eq:dml-expansion}
\end{equation}
Let $\sigma_{\phi_n}^2:=\mathrm{Var}(\psi_{\phi_n}(O;g_{\phi_n},\alpha_{\phi_n}))$.
Then
\[
\sigma^{-1}_{\phi_n}\sqrt{n}(\hat\theta_{\phi_n}-\theta_{\phi_n})
\Rightarrow
N(0,1).
\]
If $\phi_n\equiv \phi$ is fixed, then $\sqrt{n}(\hat\theta_{\phi}-\theta_{\phi})\Rightarrow N(0,\sigma_\phi^2)$.
Moreover, the cross-fitted variance estimator $\hat\sigma_{\phi_n}^2$ defined in \eqref{eq:var-est} is consistent for $\sigma_{\phi_n}^2$.
\end{theorem}
Theorem~\ref{thm:dml-expansion} supports the familiar inferential operations for the representation-specific population truth.
For any $\gamma\in(0,1)$, let $z_{1-\gamma/2}$ denote the corresponding standard-normal quantile and form
\begin{equation}
\mathrm{CI}_{\phi_n}(1-\gamma)
:=
\left[
\hat\theta_{\phi_n}
\mathbin{\pm}
z_{1-\gamma/2}\frac{\hat\sigma_{\phi_n}}{\sqrt n}
\right].
\label{eq:ci-short}
\end{equation}
Then
$\mathrm{P}\{\theta_{\phi_n}\in\mathrm{CI}_{\phi_n}(1-\gamma)\}=1-\gamma+o(1)$.
The same studentized statistic yields asymptotically valid one- and two-sided tests of hypotheses about $\theta_{\phi_n}$ and the usual local-power calculations.
This result identifies and estimates the precise effect defined by the information retained in $X_{\phi_n}$, without yet imposing a link to the full-information target.
\subsection{Omitted information bias, target adaptivity, and adaptive efficiency}
\label{sec:omitted}
Proposition~\ref{prop:ovb} makes the \emph{target gap} $\theta^\star-\theta_{\phi_n}$ explicit as an interaction of two approximation errors. For inference, the key question is whether that gap is negligible relative to sampling noise, and the next definition records the minimal condition that makes this precise.
\begin{definition}[Target adaptivity]
\label{def:target-adaptivity}
Let $\{\phi_n\}_{n\ge 1}$ be a (possibly drifting) sequence of embedding rules, and write $\theta_{\phi_n}$ for the induced representation-based target.
We say that $\{X_{\phi_n}\}_{n\ge 1}$ is \emph{target-adaptive} for $\theta^\star$ if
\begin{equation}
\sqrt{n}\bigl(\theta^\star-\theta_{\phi_n}\bigr)=o_\mathrm{P}(1).
\label{eq:target-adaptivity}
\end{equation}
\end{definition}
Since $\theta^\star-\theta_{\phi_n}=\mathrm{E}[(g^\star-g_{\phi_n})(\alpha^\star-\alpha_{\phi_n})]$, a simple sufficient condition is that the two approximation errors are small enough in $L^2(P)$ that their product is negligible at $\sqrt{n}$ scale:
\begin{equation}
\sqrt{n}\,\left\lVert g^\star-g_{\phi_n} \right\rVert_{P,2}\,\left\lVert \alpha^\star-\alpha_{\phi_n} \right\rVert_{P,2}=o_\mathrm{P}(1).
\label{eq:target-adaptivity-suff}
\end{equation}
Theorem~\ref{thm:dml-expansion} delivers inference for the representation-based target $\theta_{\phi_n}$.
When the representation is target-adaptive in the sense of Definition~\ref{def:target-adaptivity}, the same output can be read as inference for the full-information target.
Indeed,
\[
\sqrt{n}(\hat\theta_{\phi_n}-\theta^\star)
=
\sqrt{n}(\hat\theta_{\phi_n}-\theta_{\phi_n})
+\sqrt{n}(\theta_{\phi_n}-\theta^\star),
\]
and the second term is $o_\mathrm{P}(1)$ by \eqref{eq:target-adaptivity}.
Combining this with \eqref{eq:dml-expansion} yields
\[
\sqrt{n}\bigl(\hat\theta_{\phi_n}-\theta^\star\bigr)
=
\frac{1}{\sqrt{n}}\sum_{i=1}^n
\Bigl(\psi_{\phi_n}(O_i;g_{\phi_n},\alpha_{\phi_n})-\theta_{\phi_n}\Bigr)
+o_\mathrm{P}(1),
\]
so the usual Wald intervals computed from the cross-fitted score are valid for $\theta^\star$.
Target adaptivity provides the first inferential enhancement of our results, as it changes the referent of the interval from $\theta_{\phi_n}$ to $\theta^\star$.
The precision of that interval is governed by a second property.
To recover the full-information efficiency bound as well as the full-information target, the short score must approach the efficient long score.
\begin{definition}[Adaptive efficiency]
\label{def:adaptive-efficiency}
We say that $\{X_{\phi_n}\}_{n\ge 1}$ is \emph{adaptively efficient} for $\theta^\star$ if
\begin{equation}
\left\lVert g^\star-g_{\phi_n} \right\rVert_{P,2}\to 0,
\qquad
\left\lVert \alpha^\star-\alpha_{\phi_n} \right\rVert_{P,2}\to 0,
\qquad
\sup_n\left\lVert \alpha_{\phi_n} \right\rVert_{P,2}<\infty.
\label{eq:adaptive-efficiency}
\end{equation}
\end{definition}
When both target adaptivity \eqref{eq:target-adaptivity} and adaptive efficiency \eqref{eq:adaptive-efficiency} hold, we say the representation is \emph{information-lossless} for the target. Target adaptivity makes the target gap negligible at $\sqrt n$ scale. Under the operator bound and score-product conditions in Lemma~\ref{lem:score-conv}, adaptive efficiency also ensures that the short score approaches the efficient full-information score.
Under these additional conditions, the two properties answer different questions: target adaptivity determines what parameter is covered, while adaptive efficiency determines whether that coverage uses the smallest attainable first-order variance.
In many applications the representation rule is itself obtained by fine-tuning on the training folds.
Section~\ref{sec:swap} explains how to incorporate this step directly into the nuisance learners, so that the same DML logic applies to the population target defined by the end-to-end pipeline.
\begin{lemma}[Short score convergence]
\label{lem:score-conv}
Assume $Y\in L^2(P)$ and \eqref{eq:rr-long}--\eqref{eq:rr-short} hold.
Assume furthermore that $m$ is $L^2$-bounded in the sense that there exists $C_m<\infty$ such that
$\left\lVert m(O,h) \right\rVert_{P,2}\le C_m \left\lVert h \right\rVert_{P,2}$ for all $h\in L^2(P_W)$.
Suppose adaptive efficiency \eqref{eq:adaptive-efficiency} holds, and, in addition,
\begin{equation}
\left\lVert (\alpha^\star-\alpha_{\phi_n})(Y-g^\star) \right\rVert_{P,2}=o(1),
\qquad
\left\lVert \alpha_{\phi_n}(g_{\phi_n}-g^\star) \right\rVert_{P,2}=o(1).
\label{eq:adaptive-efficiency-score-products}
\end{equation}
Define the centered long score
$
\varphi^\star(O):=\psi^\star(O;g^\star,\alpha^\star)-\theta^\star
$
and the centered short score
$
\varphi_{\phi_n}(O):=\psi_{\phi_n}(O;g_{\phi_n},\alpha_{\phi_n})-\theta_{\phi_n}.
$
Then
\begin{equation}
\left\lVert \varphi_{\phi_n}-\varphi^\star \right\rVert_{P,2}=o(1).
\label{eq:score-conv}
\end{equation}
In particular, $\sigma_{\phi_n}^2\to \sigma_\star^2$, where $\sigma_\star^2:=\mathrm{Var}(\psi^\star(O;g^\star,\alpha^\star))$.
\end{lemma}
When we combine target adaptivity with short-score convergence, the representation-based estimator behaves as if it were built from the full-information efficient influence function.
\begin{theorem}[Adaptive efficiency under information-losslessness]
\label{thm:strong-adaptive}
Suppose Assumption~\ref{ass:dml} holds for a (possibly drifting) sequence $\phi=\phi_n$.
Assume also that there exists $C_m<\infty$ such that
\[
\left\lVert m(O,h) \right\rVert_{P,2}\le C_m\left\lVert h \right\rVert_{P,2}
\qquad\text{for all }h\in L^2(P_W).
\]
This bound controls the score contribution of the omitted regression component $g^\star-g_{\phi_n}$, which need not belong to a short information space.
If, in addition, target adaptivity \eqref{eq:target-adaptivity}, the adaptive score products \eqref{eq:adaptive-efficiency-score-products}, and adaptive efficiency \eqref{eq:adaptive-efficiency} hold, then
\begin{equation}
\sqrt{n}\bigl(\hat\theta_{\phi_n}-\theta^\star\bigr)
=
\frac{1}{\sqrt{n}}\sum_{i=1}^n
\Bigl(\psi^\star(O_i;g^\star,\alpha^\star)-\theta^\star\Bigr)
+o_\mathrm{P}(1).
\label{eq:efficient-expansion}
\end{equation}
Moreover,
$\sigma_\star^{-1}\sqrt{n}\bigl(\hat\theta_{\phi_n}-\theta^\star\bigr)\Rightarrow N(0,1)$,
and the cross-fitted variance estimator $\hat\sigma_{\phi_n}^2$ is consistent for $\sigma_\star^2$.
In particular, $\hat\theta_{\phi_n}$ attains the semiparametric efficiency bound for $\theta^\star$.
\end{theorem}
Theorem~\ref{thm:strong-adaptive} permits direct inference about the full-information truth.
In particular, the same interval in \eqref{eq:ci-short} now satisfies
\[
\mathrm{P}\{\theta^\star\in\mathrm{CI}_{\phi_n}(1-\gamma)\}
=1-\gamma+o(1),
\]
and $\hat\sigma_{\phi_n}\overset{p}{\to}\sigma_\star$; equivalently, the reported standard error is $\sigma_\star/\sqrt n+o_\mathrm{P}(n^{-1/2})$.
This distinguishes the theorem from Theorem~\ref{thm:dml-expansion} in two steps.
Theorem~\ref{thm:dml-expansion} alone covers the representation-specific truth $\theta_{\phi_n}$; adding target adaptivity makes that interval valid for $\theta^\star$; adding adaptive efficiency, as in Theorem~\ref{thm:strong-adaptive}, also makes the estimator first-order equivalent to one constructed from the efficient full-information influence function.
\subsection{Inference with fine-tuned representations}
\label{sec:fine-tuned}
In many empirical workflows the representation rule is not literally fixed.
Analysts often fine-tune an embedding model, select a prompt, or otherwise adapt the representation to the downstream task.
Cross-fitting supplies a direct route to inference in this setting.
On each training fold, the analyst may run the complete tuning procedure using $(Y,D,Z)$ and then carry only the fitted nuisance functions to the held-out fold.
Conditional on the training sample, those functions are fixed when the score is evaluated, preserving the independence structure used by DML.
The resulting theory is deliberately output-based.
The tuning map $\hat\phi$ may be high-dimensional, nonsmooth, and fold-specific; first-order inference is governed by the prediction errors of the fitted regression and Riesz representer rather than by an expansion of $\hat\phi$ itself.
We therefore treat \emph{the entire first-stage pipeline}---representation learning together with the two nuisance heads---as a learning procedure that outputs fitted functions.
The joint-class view below identifies the population objects learned by this procedure and places the standard DML product-rate condition directly on its two outputs.
\subsubsection{A joint-class view of representation learning}
\label{sec:swap}
Let $\Phi$ be a collection of representation rules.
For $\phi\in\Phi$, write $X_\phi=x(Z,\phi)$ and $W_\phi=(D,X_\phi)$.
Let $\Gamma_g$ be a class of real-valued maps on the range of $W_\phi$ that we use to fit the outcome regression, and let $\Gamma_\alpha$ be an analogous class used to fit the Riesz representer.
These primitive classes induce composed function classes on the original covariate space:
\[
\mathcal G
:=
\Bigl\{\,w\mapsto \gamma\!\bigl(W_\phi(w)\bigr): \gamma\in\Gamma_g,\ \phi\in\Phi\,\Bigr\},
\qquad
\mathcal A
:=
\Bigl\{\,w\mapsto a\!\bigl(W_\phi(w)\bigr): a\in\Gamma_\alpha,\ \phi\in\Phi\,\Bigr\}.
\]
It is often convenient to work on a linear space that contains all functions the pipeline can express (up to linear combination and $L^2(P)$ limits).
Accordingly, define the closed linear span
\[
\mathcal H
:=
\overline{\mathrm{span}}\bigl(\mathcal G\cup \mathcal A\bigr)
\subseteq L^2(P_W),
\]
where the closure is taken in $L^2(P)$.
We interpret $\mathcal H$ as the \emph{population information structure} of the end-to-end pipeline.
The next definition formalizes the population objects that this pipeline targets.
These objects play the same conceptual role as $(g_\phi,\alpha_\phi)$ for a fixed representation, but they allow the representation to be selected or fine-tuned as part of the learning rule.
\begin{definition}[Population pipeline targets]
\label{def:pipeline-target}
Let $g^\star(w)=\mathrm{E}[Y\mid W=w]$ and let $\alpha^\star$ be the full-information Riesz representer from \eqref{eq:rr-long}.
Define
\[
g_0 := \arg\min_{g\in\mathcal H}\mathrm{E}\!\left[(Y-g(W))^2\right],
\qquad
\alpha_0 := \arg\min_{\alpha\in\mathcal H}\mathrm{E}\!\left[\alpha(W)^2-2m(O,\alpha)\right].
\]
The corresponding \emph{pipeline target} is
\[
\theta_0 := \mathrm{E}\!\left[m(O,g_0)\right].
\]
\end{definition}
This definition is motivated by the two quadratic objectives that are standard in DML.
The first is ordinary least squares risk, whose minimizer is the $L^2(P)$ projection of $g^\star$ onto $\mathcal H$.
The second is the Riesz loss from \cite{chernozhukov2022riesz}, whose minimizer is the $L^2(P)$ projection of $\alpha^\star$ onto the same space.
Working with a common space $\mathcal H$ ensures that the resulting score remains orthogonal and admits the same mixed-bias algebra that underlies double robustness.
\begin{proposition}[Mixed-bias identities for the pipeline target]
\label{prop:pipeline-mixed-bias}
Assume the compatibility condition in Section~\ref{sec:setup} and \eqref{eq:rr-long}.
Let $(g_0,\alpha_0,\theta_0)$ be as in Definition~\ref{def:pipeline-target}.
Then the following hold.
\smallskip
\noindent (i) (\emph{Restricted Riesz representation.})
For every $h\in\mathcal H$,
\[
\mathrm{E}[m(O,h)] = \mathrm{E}\!\left[h(W)\,\alpha_0(W)\right].
\]
\smallskip
\noindent (ii) (\emph{Mixed-bias property.})
For any $g,\alpha\in\mathcal H$,
\begin{equation}
\mathrm{E}\!\left[\psi^\star(O;g,\alpha)\right]-\theta_0
=
-\,\mathrm{E}\!\left[(g-g_0)(\alpha-\alpha_0)\right].
\label{eq:mixed-bias-pipeline}
\end{equation}
\smallskip
\noindent (iii) (\emph{Target gap.})
The gap between the full-information and the pipeline targets satisfies
\begin{equation}
\theta^\star-\theta_0
=
\mathrm{E}\!\left[(g^\star-g_0)(\alpha^\star-\alpha_0)\right],
\qquad
|\theta^\star-\theta_0|
\le
\left\lVert g^\star-g_0 \right\rVert_{P,2}\,\left\lVert \alpha^\star-\alpha_0 \right\rVert_{P,2}.
\label{eq:ovb-pipeline}
\end{equation}
\end{proposition}
Part (ii) is the key operational fact: the population drift of the orthogonal score is a \emph{mixed bias} term.
In particular, if either nuisance is correct (or nearly correct) in $L^2(P)$, the bias is second order.
Part (iii) shows that the same mixed-bias structure also governs how far the pipeline target $\theta_0$ can lie from the full-information target $\theta^\star$.
\subsection{Subordinated tuning and information loss}
The joint-class construction above answers a population question: what regression, representer, and target are induced by an end-to-end pipeline?
Subordinated tuning answers a distinct design question: what information-loss criterion is optimized when a representation is selected by prediction or Riesz risk?
The joint-class analysis accommodates a broad, possibly nonsmooth pipeline, whereas the profiling argument below explains the geometry of two especially natural tuning objectives.
To make that geometry explicit, we ``profile out'' the nuisance learners and view the embedding index as a tuning parameter that controls how much information is discarded when $W=(D,Z)$ is replaced by $W_\phi=(D,X_\phi)$.
Concretely, imagine a two-stage (subordinated) choice: for each $\phi$, take the population best predictor of $Y$ given $W_\phi$, and then choose $\phi$ to optimize the resulting profiled risk.
An analogous profiling construction applies to the Riesz representer by optimizing the population Riesz loss over $\phi$.
The key point is geometric: because both losses are quadratic, profiling removes the nuisance functions and leaves (up to constants) exactly the $L^2(P)$ information-loss quantities that enter Proposition~\ref{prop:ovb} and Proposition~\ref{prop:pipeline-mixed-bias}.
\textit{Outcome-side tuning targets $\left\lVert \Delta_g(\phi) \right\rVert_{P,2}^2$.}
Let $g^\star(W)=\mathrm{E}[Y\mid W]$ and, for each $\phi$, let $g_\phi(W_\phi)=\mathrm{E}[Y\mid W_\phi]$.
Define the regression discrepancy
\[
\Delta_g(\phi):=g^\star(W)-g_\phi(W_\phi).
\]
Since $g_\phi(W_\phi)$ is the $L^2(P)$ projection of $Y$ onto $\sigma(W_\phi)$, the Pythagorean identity for orthogonal projections yields
\begin{equation}
\mathrm{E}\!\left[(Y-g_\phi(W_\phi))^2\right]
=
\mathrm{E}\!\left[(Y-g^\star(W))^2\right]
+\mathrm{E}\!\left[\Delta_g(\phi)^2\right].
\label{eq:profile-reg}
\end{equation}
The first term does not depend on $\phi$, so minimizing population prediction risk over $\phi$ is equivalent to minimizing
$\mathrm{E}[\Delta_g(\phi)^2]=\left\lVert \Delta_g(\phi) \right\rVert_{P,2}^2=\left\lVert g^\star-g_\phi \right\rVert_{P,2}^2$.
In this sense, subordinated outcome-side tuning chooses the representation that best preserves the full-information conditional mean of $Y$ (given $D$) in $L^2(P)$.
\textit{Representer-side tuning targets $\left\lVert \Delta_\alpha(\phi) \right\rVert_{P,2}^2$.}
Let $\alpha^\star(W)$ be the full-information Riesz representer from \eqref{eq:rr-long}, and let $\alpha_\phi(W_\phi)$ be the representation-based representer from \eqref{eq:rr-short}.
In the unrestricted setup of Proposition~\ref{prop:ovb}, we have $\alpha_\phi(W_\phi)=\mathrm{E}[\alpha^\star(W)\mid W_\phi]$, so it is again an $L^2(P)$ projection.
Define
\[
\Delta_\alpha(\phi):=\alpha^\star(W)-\alpha_\phi(W_\phi).
\]
Consider the population Riesz loss for a square-integrable candidate weight $a(W_\phi)$,
\[
\mathcal L_\phi(a):=\mathrm{E}\!\left[a(W_\phi)^2\right]-2\,\mathrm{E}\!\left[m(O,a)\right].
\]
The minimizer of $\mathcal L_\phi(\cdot)$ over $L^2(P_{W_\phi})$ is $a=\alpha_\phi$, and the representer normal equation implies
$\inf_a \mathcal L_\phi(a)=\mathcal L_\phi(\alpha_\phi)=-\mathrm{E}[\alpha_\phi(W_\phi)^2]$.
Moreover, the projection identity gives
\begin{equation}
\left\lVert \Delta_\alpha(\phi) \right\rVert_{P,2}^2
=
\mathrm{E}\!\left[\alpha^\star(W)^2\right]-\mathrm{E}\!\left[\alpha_\phi(W_\phi)^2\right].
\label{eq:profile-riesz}
\end{equation}
Therefore, minimizing the profiled Riesz loss over $\phi$ is equivalent to minimizing the representer information loss
$\left\lVert \Delta_\alpha(\phi) \right\rVert_{P,2}^2=\left\lVert \alpha^\star-\alpha_\phi \right\rVert_{P,2}^2$.
Finally, since Proposition~\ref{prop:ovb} can be rewritten as $\theta^\star-\theta_\phi=\mathrm{E}[\Delta_g(\phi)\Delta_\alpha(\phi)]$, these profiled objectives directly control the two components whose interaction governs both the target gap and the sensitivity scalings in Section~\ref{sec:bounds}.
In the joint-class view developed above, the same geometry holds with $(g_0,\alpha_0)$ in place of $(g_\phi,\alpha_\phi)$ because both are $L^2(P)$ projections onto the linear space $\mathcal H$.
\subsection{DML inference after foldwise tuning}
\label{sec:pipeline-inference}
The profiling calculation characterizes what particular population tuning criteria preserve.
Inference requires a different and more general statement: given any foldwise end-to-end learner, can we conduct inference on the population target induced by its function class?
The answer below does not require the implemented pipeline to solve the subordinated profiling problems, nor does it require the regression and representer learners to select the same intermediate representation.
It requires only that their final fitted functions converge to the two projections on the common space $\mathcal H$.
We estimate $\theta_0$ exactly as in the warm-up, except that the first-stage learners are allowed to tune or fine-tune the representation internally.
Split the sample into folds $I_1,\dots,I_K$.
On each training fold $I_k^c$, run the end-to-end pipeline to obtain fitted nuisances
$\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$, intended to approximate the population projections $g_0$ and $\alpha_0$,
where both estimators may depend on a fold-specific tuned representation $\hat\phi^{(-k)}$.
We then evaluate the long score \eqref{eq:score-long} on the hold-out fold using these fitted functions:
\[
\hat\psi_i
:=
\psi^\star\!\left(O_i;\hat g^{(-k(i))},\hat\alpha^{(-k(i))}\right),
\qquad
\hat\theta_0 := \frac{1}{n}\sum_{i=1}^n \hat\psi_i,
\qquad
\hat\sigma_0^2 := \frac{1}{n}\sum_{i=1}^n(\hat\psi_i-\hat\theta_0)^2.
\]
The point of cross-fitting is that, conditional on the training folds, the functions $\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$ are fixed, so the empirical process arguments used in Theorem~\ref{thm:dml-expansion} still apply.
\begin{assumption}[DML regularity for an end-to-end pipeline]
\label{ass:pipeline-dml}
Let $\mathcal H$ and $(g_0,\alpha_0,\theta_0)$ be as in Definition~\ref{def:pipeline-target}.
Assume that, for some $q>4$, $Y,g_0,\alpha_0\in L^q(P)$, and that $\mathrm{Var}(\psi^\star(O;g_0,\alpha_0))$ is bounded away from zero and infinity.
Assume that $m$ is $L^2$-bounded on $\mathcal H$ in the sense of Lemma~\ref{lem:score-conv}.
For each fold $k$, require that both $\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$ belong to $\mathcal H$ and are constructed using only $I_k^c$. Also assume
\[
\max_{1\le k\le K}\left\lVert \hat g^{(-k)} \right\rVert_{P,q}
+\max_{1\le k\le K}\left\lVert \hat\alpha^{(-k)} \right\rVert_{P,q}=O_\mathrm{P}(1).
\]
These moment bounds prevent rare extreme predictions from dominating the score and allow $L^2(P)$ consistency to imply $L^4(P)$ consistency. The moment requirement on $g_0$ is explicit because projection onto a general linear space need not preserve the moments of $Y$.
For each fold $k$, the cross-fitted nuisance learners satisfy
\[
\begin{aligned}
& \left\lVert \hat g^{(-k)}-g_0 \right\rVert_{P,2} = o_\mathrm{P}(1), \quad
\left\lVert \hat\alpha^{(-k)}-\alpha_0 \right\rVert_{P,2} = o_\mathrm{P}(1),\\
& \sqrt{n}\,\left\lVert \hat g^{(-k)}-g_0 \right\rVert_{P,2}\,\left\lVert \hat\alpha^{(-k)}-\alpha_0 \right\rVert_{P,2}
= o_\mathrm{P}(1),
\end{aligned}
\]
uniformly over $k\in\{1,\dots,K\}$.
\end{assumption}
\begin{theorem}[DML inference for the population pipeline target]
\label{thm:pipeline-dml}
Suppose Assumption~\ref{ass:pipeline-dml} holds.
Then
\begin{equation}
\sqrt{n}\bigl(\hat\theta_0-\theta_0\bigr)
=
\frac{1}{\sqrt{n}}\sum_{i=1}^n
\Bigl(\psi^\star(O_i;g_0,\alpha_0)-\theta_0\Bigr)
+o_\mathrm{P}(1).
\label{eq:pipeline-expansion}
\end{equation}
Moreover,
$\sigma_0^{-1}\sqrt{n}\bigl(\hat\theta_0-\theta_0\bigr)\Rightarrow N(0,1)$,
where $\sigma_0^2:=\mathrm{Var}(\psi^\star(O;g_0,\alpha_0))$, and the cross-fitted variance estimator $\hat\sigma_0^2$ is consistent for $\sigma_0^2$.
\end{theorem}
Theorem~\ref{thm:pipeline-dml} extends Theorem~\ref{thm:dml-expansion} from a fixed or drifting representation to a fold-specific learned pipeline.
For
\[
\mathrm{CI}_{0}(1-\gamma)
:=
\left[
\hat\theta_0
\mathbin{\pm}
z_{1-\gamma/2}\frac{\hat\sigma_0}{\sqrt n}
\right],
\]
we have $\mathrm{P}\{\theta_0\in\mathrm{CI}_{0}(1-\gamma)\}=1-\gamma+o(1)$.
Thus the tuned pipeline defines a precise population estimand $\theta_0$, and the usual DML point estimate, test statistic, and confidence interval are valid for that estimand without a stability or differentiability condition on the tuning parameter itself.
Proposition~\ref{prop:pipeline-mixed-bias} then supplies the bridge from the pipeline truth to the full-information truth.
In particular, the sufficient condition
\[
\sqrt{n}\,\left\lVert g^\star-g_0 \right\rVert_{P,2}\,\left\lVert \alpha^\star-\alpha_0 \right\rVert_{P,2}=o(1).
\]
makes the same interval valid for $\theta^\star$.
More generally, Section~\ref{sec:bounds} uses the exact target-gap identity to enlarge inference from the point $\theta_0$ to a sensitivity region for $\theta^\star$.
We now turn to a complementary question: how to construct regression and Riesz learners whose errors satisfy the $L^2(P)$ product-rate condition, once the representation-and-head pipeline is viewed as inducing a joint prediction space~$\mathcal H$.
\section{Learning the pipeline targets over $\mathcal H$}
\label{sec:learn-H}
Assumption~\ref{ass:pipeline-dml} is phrased directly in terms of the $L^2(P)$ errors of the fitted nuisances $(\hat g^{(-k)},\hat\alpha^{(-k)})$, and that is exactly the level at which a practitioner can usually reason about a modern pipeline.
In an end-to-end workflow, an embedding rule, a fine-tuning choice, and a head architecture are merely intermediate artifacts; what enters the orthogonal score is the resulting pair of predictions.
This section explains how one can justify the product-rate requirement in Assumption~\ref{ass:pipeline-dml} by reducing it to familiar learning-theoretic bounds for two quadratic empirical objectives, one for the regression and one for the Riesz representer.
Because we cross-fit, all learning is foldwise.
For each fold $k$, we train on the complement $I_k^c$ and evaluate on the held-out fold $I_k$.
To keep formulas readable, we adopt the standard cross-fitting convention that, whenever we write an empirical objective $\mathbb{P}_n(\cdot)$ used to \emph{train} $\hat g^{(-k)}$ or $\hat\alpha^{(-k)}$, it is understood as the empirical measure over the part of $I_k^c$ used for the objective (with effective size $n_{\mathrm{tr}}\asymp n$). For a learned candidate menu, this part is independent of the data used to fit the candidates.
Thus, when a bound below displays $1/n$ or $\sqrt{1/n}$, it should be interpreted as $1/n_{\mathrm{tr}}$ or $\sqrt{1/n_{\mathrm{tr}}}$.
For fixed $K$, this distinction affects only constants, and it keeps the discussion aligned with the population condition \eqref{eq:product-rate-template}.
Throughout this section, we write $a\mathrel{\lesssim} b$ to mean that $a\le Cb$ for a numerical constant $C$ that does not depend on $n$, $M$, or the confidence level $\delta$ (though it may depend on fixed envelope constants that we display explicitly).
\subsection{From quadratic objectives to $L^2(P)$ rates}
\label{sec:learn-H-quadratic}
The defining feature of the pipeline targets $(g_0,\alpha_0)$ is that both are projections onto the \emph{same} closed linear space $\mathcal H$ (Definition~\ref{def:pipeline-target}).
This shared geometry is not a technicality: it is what makes it possible to control both nuisances with a single kind of argument and then combine them via a simple product condition.
Once we work on $\mathcal H$, both learning problems become quadratic minimization problems, and their excess risks coincide with squared $L^2(P)$ errors.
To make this explicit, define the population regression risk
\[
\mathcal R(g):=\mathrm{E}\!\left[(Y-g(W))^2\right],
\qquad g\in L^2(P_W),
\]
and the population Riesz loss (as in \cite{chernozhukov2022riesz})
\[
\mathcal L(\alpha):=\mathrm{E}\!\left[\alpha(W)^2-2m(O,\alpha)\right],
\qquad \alpha\in L^2(P_W).
\]
By Definition~\ref{def:pipeline-target}, $g_0$ minimizes $\mathcal R(g)$ over $\mathcal H$ and $\alpha_0$ minimizes $\mathcal L(\alpha)$ over $\mathcal H$.
\begin{lemma}[Excess-risk identities on $\mathcal H$]
\label{lem:excess-risk-identities}
Let $(g_0,\alpha_0)$ be the pipeline projections from Definition~\ref{def:pipeline-target}.
Then, for every $g\in\mathcal H$,
\begin{equation}
\mathcal R(g)-\mathcal R(g_0)=\left\lVert g-g_0 \right\rVert_{P,2}^2,
\label{eq:excess-risk-reg}
\end{equation}
and for every $\alpha\in\mathcal H$,
\begin{equation}
\mathcal L(\alpha)-\mathcal L(\alpha_0)=\left\lVert \alpha-\alpha_0 \right\rVert_{P,2}^2.
\label{eq:excess-risk-riesz}
\end{equation}
\end{lemma}
The lemma says that, on $\mathcal H$, there is no gap between “risk control” and “prediction error control”: they are the same statement.
This is the bridge from learning theory to the high-level requirements in Assumption~\ref{ass:pipeline-dml}.
A proof is given in Appendix~\ref{app:proofs-sec3}; it is an application of orthogonality for $L^2(P)$ projections and the restricted Riesz representation in Proposition~\ref{prop:pipeline-mixed-bias}(i).
A convenient template is to assume foldwise oracle inequalities of the form
\begin{align}
\mathcal R(\hat g^{(-k)})-\mathcal R(g_{0,n})
&= O_\mathrm{P}(r_{g,n}^2),
\label{eq:oracle-reg}\\
\mathcal L(\hat\alpha^{(-k)})-\mathcal L(\alpha_{0,n})
&= O_\mathrm{P}(r_{\alpha,n}^2),
\label{eq:oracle-riesz}
\end{align}
where $g_{0,n}$ and $\alpha_{0,n}$ are population minimizers over approximation spaces $\mathcal H_{g,n}\subseteq\mathcal H$ and $\mathcal H_{\alpha,n}\subseteq\mathcal H$, and $r_{g,n},r_{\alpha,n}$ are estimation (complexity) terms.
Writing $a_{g,n}:=\left\lVert g_{0,n}-g_0 \right\rVert_{P,2}$ and $a_{\alpha,n}:=\left\lVert \alpha_{0,n}-\alpha_0 \right\rVert_{P,2}$ for approximation errors, Lemma~\ref{lem:excess-risk-identities} yields
\[
\left\lVert \hat g^{(-k)}-g_0 \right\rVert_{P,2}=O_\mathrm{P}(r_{g,n}+a_{g,n}),
\qquad
\left\lVert \hat\alpha^{(-k)}-\alpha_0 \right\rVert_{P,2}=O_\mathrm{P}(r_{\alpha,n}+a_{\alpha,n}),
\]
uniformly over folds.
Given individual nuisance consistency and the remaining regularity conditions in Assumption~\ref{ass:pipeline-dml}, its product-rate requirement is verified by the design constraint
\begin{equation}
\sqrt{n}\,(r_{g,n}+a_{g,n})\,(r_{\alpha,n}+a_{\alpha,n})\to 0.
\label{eq:product-rate-template}
\end{equation}
The structure is the same as in standard DML with fixed covariates: each nuisance error of order $o_\mathrm{P}(n^{-1/4})$ suffices, and faster learning of one nuisance can compensate for slower learning of the other. The product condition makes the remaining bias smaller than the sampling noise.
What is new in the pipeline setting is that the approximation and complexity terms are governed by the induced information space $\mathcal H$.
The remainder of the section develops two complementary implementations of this template.
The first is a \emph{proper-learning route}: the analyst specifies a controlled star-shaped or convex approximation space inside $\mathcal H$ and minimizes each quadratic loss over that space.
This route is attractive when the pipeline admits an explicit sieve or convex aggregation class whose approximation and critical-radius terms can be evaluated.
The second is an \emph{improper-learning route}: the analyst begins with a finite, possibly nonconvex menu of fitted pipelines and applies star aggregation.
This route is designed for modular empirical workflows in which prompts, embeddings, and heads arrive as discrete candidates, and its complexity depends logarithmically on the menu size.
Both routes lead to the same destination---foldwise $L^2(P)$ rates that can be inserted into \eqref{eq:product-rate-template}---but they organize the first-stage search differently.
\subsection{Proper learning on controlled approximations to $\mathcal H$}
\label{sec:route-i}
When we can write down controlled approximation spaces inside $\mathcal H$, a natural approach is to learn each nuisance by (regularized) ERM over such a space on each training fold.
The practical advantage is transparency: once $\mathcal H_{g,n}$ and $\mathcal H_{\alpha,n}$ are chosen, we can read off $r_{g,n}$ and $r_{\alpha,n}$ from standard complexity calculations, and then check \eqref{eq:product-rate-template}.
For the regression nuisance, we can take $\hat g^{(-k)}$ to be a (regularized) empirical minimizer of the square loss over $\mathcal H_{g,n}$,
\[
\hat g^{(-k)}\in\arg\min_{g\in\mathcal H_{g,n}} \mathbb{P}_n\!\left[(Y-g(W))^2\right],
\]
and for the representer nuisance we analogously minimize the empirical Riesz loss over $\mathcal H_{\alpha,n}$,
\[
\hat\alpha^{(-k)}\in\arg\min_{\alpha\in\mathcal H_{\alpha,n}} \mathbb{P}_n\!\left[\alpha(W)^2-2m(O,\alpha)\right].
\]
For square loss on a star-shaped class, the sharp way to express $r_{g,n}$ is through a localized fixed point, often called the \emph{critical radius}.
The next proposition records a representative form; it is standard and we cite it rather than reprove it.
\begin{proposition}[Square-loss ERM on a star-shaped class]
\label{prop:crit-radius-reg}
Let $\mathcal F\subseteq L^2(P_W)$ be fixed, or constructed independently of the sample used in $\mathbb{P}_n$, and star-shaped around
\[
f_0\in\arg\min_{f\in\mathcal F}\mathrm{E}[(Y-f(W))^2].
\]
Thus $f_0$ is the population best predictor in the class; star-shapedness and this optimality supply the quadratic margin. For an independently constructed random class, the conclusion is conditional on that class.
Assume bounded envelopes: $|Y|\le B_Y$ and $\sup_{f\in\mathcal F}|f(W)|\le B_F$ almost surely.
Let $\hat f\in\arg\min_{f\in\mathcal F}\mathbb{P}_n[(Y-f(W))^2]$.
Define the localized Rademacher functional
\[
\mathcal R_n(r):=\mathrm{E}\Bigl[\sup_{\substack{f\in\mathcal F\\ \left\lVert f-f_0 \right\rVert_{P,2}\le r}}
\frac{1}{n}\sum_{i=1}^n\varepsilon_i\{f(W_i)-f_0(W_i)\}\Bigr],
\]
and the corresponding critical radius $r_n:=\inf\{r>0:\ \mathcal R_n(r)\le r^2/8\}$.
Then there is a universal constant $C=C(B_Y,B_F)$ such that, for any $\delta\in(0,1)$, with probability at least $1-\delta$,
\begin{equation}
\left\lVert \hat f-f_0 \right\rVert_{P,2}\ \le\ C\Bigl(r_n+\sqrt{\tfrac{\log(1/\delta)}{n}}\Bigr).
\label{eq:crit-radius-concl}
\end{equation}
\end{proposition}
Results of this kind can be proved cleanly using offset complexity and localization; see \cite{liang2015offset} and the references therein.
In practice, the point is that once we know a bound on $r_n$ for the induced approximation space, \eqref{eq:crit-radius-concl} converts it directly into an $L^2(P)$ learning rate.
For the representer nuisance, an analogous learning-theoretic statement is available via recent work on Riesz regression.
Theorem~2.1 of \cite{sridhar2022} provides high-probability rates for empirical minimizers of the Riesz loss under bounded-class assumptions (rather than sub-Gaussian tails), which is well matched to the kinds of envelope conditions that are natural when the primitive candidates are produced by a finite pipeline menu.
To keep the discussion concrete, we record the rate form in our notation.
\begin{proposition}[Riesz regression rates for bounded representer classes]
\label{prop:riesz-theorem21}
Let $\mathcal A\subseteq\mathcal H$ be a representer class fixed or constructed independently of the sample used in $\mathbb{P}_n$; in the latter case the statement is conditional on that class. Let
\[
\hat\alpha\in\arg\min_{\alpha\in\mathcal A}\mathbb{P}_n\!\left[\alpha(W)^2-2m(O,\alpha)\right].
\]
Let $\alpha_{0,n}\in\arg\min_{\alpha\in\mathcal A}\mathcal L(\alpha)$ be the population minimizer over $\mathcal A$.
Assume the bounded-class conditions of Theorem~2.1 of \cite{sridhar2022}.
Two components are worth emphasizing.
First, $m$ is mean-square continuous on $\mathrm{span}(\mathcal A)$: for some $C_m>0$, $\mathrm{E}[m(O,h)^2]\le C_m\,\left\lVert h \right\rVert_{P,2}^2$ for all $h$ in the relevant span.
Second, the deviation sets appearing in the proof (star-shaped hulls of $\mathcal A-\alpha_{0,n}$ and of $m\circ\mathcal A-m\circ\alpha_{0,n}$) admit bounded envelopes.
Let $\delta_n$ denote an upper bound on the critical radius in that theorem, and let $\zeta\in(0,1)$.
Then, with probability at least $1-\zeta$,
\begin{equation}
\left\lVert \hat\alpha-\alpha_0 \right\rVert_{P,2}^2
\ \mathrel{\lesssim}\
\left\lVert \alpha_{0,n}-\alpha_0 \right\rVert_{P,2}^2
\;+\;
C_m\,\delta_n^2
\;+\;
C_m\,\frac{\log(1/\zeta)}{n}.
\label{eq:riesz-theorem21}
\end{equation}
\end{proposition}
Propositions~\ref{prop:crit-radius-reg} and \ref{prop:riesz-theorem21} clarify what one needs to check in order to validate \eqref{eq:oracle-reg}--\eqref{eq:oracle-riesz}: one needs a complexity bound (a critical radius) for the regression class and for the representer class, together with approximation control for the chosen sieves.
A convenient worked example is the finite-dictionary setting, which is the one most directly aligned with “try a handful of representations and heads and keep what works.”
Suppose an independent part of the training fold yields a finite menu of bounded candidate predictors $\mathcal F=\{f_1,\dots,f_M\}\subseteq\mathcal H$. The aggregation objective uses the remaining part, so its observations are independent of the fitted menu.
For a radius $B>0$, consider the signed $\ell_1$-hull
\[
\mathcal F_B
:=
\left\{f=\sum_{j=1}^M b_j f_j:\ \sum_{j=1}^M |b_j|\le B\right\}.
\]
This is a convex (hence star-shaped) approximation to $\overline{\mathrm{span}}(\mathcal F)$ that is controlled by a single tuning parameter $B$.
A localized complexity calculation (recorded in Appendix~\ref{app:proofs-sec3}) yields the schematic critical-radius bound
\begin{equation}
r_n\ \mathrel{\lesssim}\
\min\Biggl\{\sqrt{\frac{M}{n}},\ \sqrt{B\max_{1\le j\le M}\left\lVert f_j \right\rVert_\infty}\Bigl(\frac{\log(2M)}{n}\Bigr)^{1/4}\Biggr\}.
\label{eq:l1-radius}
\end{equation}
Thus, under the boundedness assumptions of Proposition~\ref{prop:crit-radius-reg}, ERM over $\mathcal F_B$ yields
\[
\left\lVert \hat f-f_{0,B} \right\rVert_{P,2}=O_\mathrm{P}(r_n),
\]
where $f_{0,B}$ is the $L^2(P)$ projection onto $\mathcal F_B$.
For representer learning, apply the same $\ell_1$ geometry separately to the bounded dictionaries $a_j$ and $m(O,a_j)$; mean-square continuity enters the Riesz learning theorem. Appendix~\ref{app:proofs-sec3} records this reduction.
The practical takeaway is that the product condition \eqref{eq:product-rate-template} becomes a transparent constraint linking $(M,B)$ choices for the regression and the representer learners.
Proper convex aggregation borrows strength across candidates by fitting a genuinely mixed predictor over an $\ell_1$ hull or another controlled convex class.
Its estimation term grows with the effective size of that class: in \eqref{eq:l1-radius}, the first regime behaves like $\sqrt{M/n}$, so attaining an $n^{-1/4}$ nuisance rate typically calls for $M=o(\sqrt n)$, up to approximation terms and fixed constants.
The next subsection gives the complementary construction for much larger, nonconvex menus.
\subsection{Improper learning by star aggregation}
\label{sec:route-ii}
Many applied pipelines naturally produce a \emph{menu} of candidates rather than a single convex class.
Varying prompts or checkpoints, searching a grid of hyperparameters, or swapping heads yields a nonconvex collection of fitted predictors in $\mathcal H$ on each training fold.
Star aggregation converts this discrete menu into a predictor that remains in the same linear information space while achieving the optimal $\log(M)/n$ model-selection-aggregation remainder; see \cite{audibert2007progressive}.
It can also approximate the population projections $(g_0,\alpha_0)$ more closely than a rule constrained to output a single candidate.
Star aggregation is exactly this kind of cheap improper refinement.
We start from the empirical winner and then perform a one-dimensional refit along each segment that connects the winner to a competitor.
Because every segment lies in $\mathcal H$ whenever the endpoints do, this refinement does not enlarge the information space made available by the pipeline; it merely exploits the curvature of the quadratic objective.
The payoff is that, even when $M$ is very large, the star aggregate enjoys high-probability oracle inequalities with a $\log(M)/n$ remainder, as we show below for both the regression and Riesz losses.
First, we make explicit what we mean when we refer to \textit{star estimators}.
\begin{definition}[Star estimator]
\label{def:star-estimator}
Let $\mathcal C\subseteq\mathcal H$ be a candidate set and let
$Q_n(c):=\mathbb{P}_n[\ell(O,c)]$ be an empirical objective.
First choose an empirical-best center
\[
\tilde c\in\arg\min_{c\in\mathcal C}Q_n(c).
\]
Its star hull is
\[
\mathrm{star}(\mathcal C,\tilde c)
:=
\{\lambda\tilde c+(1-\lambda)c:
c\in\mathcal C,\ \lambda\in[0,1]\}.
\]
Any empirical minimizer
\[
\hat c\in
\arg\min_{c\in\mathrm{star}(\mathcal C,\tilde c)}Q_n(c)
\]
is called a \emph{star estimator} over $\mathcal C$.
\end{definition}
For the regression nuisance, take $\mathcal C=\mathcal F$ and use squared loss.
Let $\mathcal F\subseteq\mathcal H$ be a candidate class on a given training fold.
We first compute an empirical risk minimizer
\[
\tilde f\in\arg\min_{f\in\mathcal F}\mathbb{P}_n\!\left[(Y-f(W))^2\right],
\]
and then re-optimize over the star hull around this center,
\[
\hat f\in\arg\min_{f\in\mathrm{star}(\mathcal F,\tilde f)}\mathbb{P}_n\!\left[(Y-f(W))^2\right],
\qquad
\mathrm{star}(\mathcal F,\tilde f)
:=
\{\lambda\tilde f+(1-\lambda)f:\ f\in\mathcal F,\ \lambda\in[0,1]\}.
\]
The same construction applies to the representer objective: for $\mathcal A\subseteq\mathcal H$, compute
\[
\tilde\alpha\in\arg\min_{\alpha\in\mathcal A}\mathbb{P}_n\!\left[\alpha(W)^2-2m(O,\alpha)\right],
\qquad
\hat\alpha\in\arg\min_{\alpha\in\mathrm{star}(\mathcal A,\tilde\alpha)}\mathbb{P}_n\!\left[\alpha(W)^2-2m(O,\alpha)\right].
\]
The key operational fact is that the second step is a one-dimensional quadratic minimization along each segment, so it has a closed form.
If we fix $\tilde f$ and a candidate $f\in\mathcal F$ and write $f_\lambda:=\lambda\tilde f+(1-\lambda)f$, then the empirical square loss is quadratic in $\lambda$ and the unconstrained minimizer is
\begin{equation}
\lambda^\ast(f)
=
\frac{\mathbb{P}_n[(Y-f(W))(\tilde f(W)-f(W))]}{\mathbb{P}_n[(\tilde f(W)-f(W))^2]},
\label{eq:lambda-reg}
\end{equation}
projected onto $[0,1]$ (with the convention that any $\lambda$ is optimal if the denominator is zero).
Thus, for a finite list $\mathcal F=\{f_1,\dots,f_M\}$, star aggregation can be computed by scanning $j=1,\dots,M$, evaluating \eqref{eq:lambda-reg}, and picking the best segment fit.
For the empirical Riesz loss, if we fix $\tilde\alpha$ and $\alpha\in\mathcal A$ and set $\alpha_\lambda:=\lambda\tilde\alpha+(1-\lambda)\alpha$, then linearity of $m$ again yields a quadratic in $\lambda$, whose unconstrained minimizer is
\begin{equation}
\lambda^\ast(\alpha)
=
-\frac{\mathbb{P}_n\!\left[\alpha(W)\,(\tilde\alpha-\alpha)(W)-m(O,\tilde\alpha-\alpha)\right]}{\mathbb{P}_n[(\tilde\alpha(W)-\alpha(W))^2]},
\label{eq:lambda-riesz}
\end{equation}
projected onto $[0,1]$.
The finite-menu regime has a direct economic interpretation.
In a multimodal demand analysis, for example, an analyst may fit outcome regressions and Riesz representers using text embeddings, image embeddings, tabular product attributes, and their combinations, perhaps crossed with a small set of heads or regularization choices.
These pipelines form a discrete, nonconvex menu because changing a foundation-model checkpoint or feature modality changes the fitted function rather than a coefficient in a common parametric class.
Such menus are common in applied work: researchers typically compare a finite collection of economically defensible specifications, while the regression and representer objectives may favor different members of that collection.
Star aggregation retains this modular workflow, combines candidates when the data support a mixture, and provides a rate that remains useful as the menu grows.
For finite menus, Definition~\ref{def:star-estimator} yields
explicit $\log(M)/n$ oracle inequalities for a menu fixed independently of the aggregation sample. No additional split between selecting the star center and fitting its weights is required. The next two results establish the guarantees
for the outcome and Riesz objectives used by DML.
\begin{proposition}[Star aggregation oracle inequality: square loss]
\label{prop:star-finite-reg}
Let $(W_i,Y_i)_{i=1}^n$ be i.i.d.\ with $|Y|\le B_Y$ a.s.,
and let $\mathcal F=\{f_1,\dots,f_M\}\subseteq\mathcal H$
with $\max_{j}\left\lVert f_j \right\rVert_\infty\le B_F$, fixed or constructed independently of these observations. In the latter case, the probability statement is conditional on the menu.
Let $\hat f$ be the star estimator in Definition~\ref{def:star-estimator}
for the empirical squared loss
$R_n(f):=\mathbb{P}_n[(Y-f(W))^2]$.
Then there exists a universal constant $C>0$ such that for every
$\delta\in(0,1)$, with probability at least $1-\delta$,
\begin{equation}
\label{eq:star-finite-reg-oracle}
\left\lVert \hat f-g_0 \right\rVert_{P,2}^2
\ \le\
\min_{1\le j\le M}\left\lVert f_j-g_0 \right\rVert_{P,2}^2
\;+\;
C\,(B_Y+B_F)^2\,\frac{\log(M/\delta)}{n},
\end{equation}
where $g_0$ is the $L^2(P)$ projection of $Y$ onto $\mathcal H$.
\end{proposition}
\begin{proposition}[Star aggregation oracle inequality: Riesz loss]
\label{prop:star-finite-riesz}
Let $O_1,\dots,O_n$ be i.i.d.\ and let
$\mathcal A=\{a_1,\dots,a_M\}\subseteq\mathcal H$
with $\max_{j}\left\lVert a_j \right\rVert_\infty\le B_A$, fixed or constructed independently of these observations. In the latter case, the probability statement is conditional on the menu.
Assume that $m(O,h)$ is linear in $h$ with:
\begin{enumerate}[label=(\roman*),nosep]
\item \emph{mean-square continuity:}
$\mathrm{E}[m(O,h)^2]\le C_m\,\left\lVert h \right\rVert_{P,2}^2$
for all $h\in\mathrm{span}(\mathcal A)$;
\item \emph{bounded envelope:}
$\sup_{a\in\mathrm{star}(\mathcal A,a_k)}|m(O,a)|\le B_m$
a.s.\ for every center $k\in\{1,\ldots,M\}$.
\end{enumerate}
Let $\hat a$ be the star estimator in Definition~\ref{def:star-estimator}
for the empirical Riesz loss
$L_n(a):=\mathbb{P}_n[a(W)^2-2\,m(O,a)]$.
Then there exists a universal constant $C>0$ such that for every
$\delta\in(0,1)$, with probability at least $1-\delta$,
\begin{equation}
\label{eq:star-finite-riesz-oracle}
\left\lVert \hat a-\alpha_0 \right\rVert_{P,2}^2
\ \le\
\min_{1\le j\le M}\left\lVert a_j-\alpha_0 \right\rVert_{P,2}^2
\;+\;
C\,(B_A^2+C_m+B_m)\,\frac{\log(M/\delta)}{n},
\end{equation}
where $\alpha_0$ minimizes $\mathcal L(\alpha)$ over $\mathcal H$.
\end{proposition}
The two propositions make the practical message quite concrete.
When each nuisance is learned by star aggregation over a finite
foldwise menu, we obtain transparent high-probability $L^2(P)$
bounds of order $\sqrt{\log(M)/n}$.
These bounds can be plugged directly into
\eqref{eq:product-rate-template}, provided the menu is independent of the sample used to aggregate it. This independence prevents candidates from fitting the same noise used to assess them. When candidates are trained within a fold, use separate data for aggregation; Section~\ref{sec:cross-fold} treats pooled aggregation without this internal split.
\medskip
Stepping back, \eqref{eq:product-rate-template} turns learner design into an explicit approximation--estimation tradeoff.
Enriching $\mathcal H$ increases approximation power, while the proper and improper constructions above quantify the accompanying statistical complexity.
The product-rate structure is especially useful because faster learning of either nuisance can compensate for slower learning of the other.
Thus Section~\ref{sec:learn-H} supplies two concrete ways to verify the high-level DML condition from the geometry of the pipeline itself.
\subsection{Aggregation across folds}
\label{sec:cross-fold}
Propositions~\ref{prop:star-finite-reg} and~\ref{prop:star-finite-riesz} treat the menu of candidates as fixed.
An honest implementation inside a training fold must therefore train the candidates and fit the star weights on disjoint data.
That calls for a second split within each training fold, or for nested cross-fitting, and either way it multiplies the number of candidate fits.
When a candidate is a fine-tuned foundation model, each fit is expensive.
Practice often follows a cheaper design.
Train every candidate once on each training fold, as cross-fitting requires anyway, and fit a single set of star weights to the pooled out-of-fold predictions.
Each prediction in the aggregation criterion then comes from a model that never saw that observation, as in stacking \cite{wolpert1992stacked,breiman1996stacked} and the super learner \cite{vanderlaan2007super}.
The weights, however, depend on every fold, including the fold on which the score is later evaluated.
Cross-fitting then holds for the candidates but not for the aggregate.
This subsection shows that the breach is cheap: it adds one remainder term of order $(r_g+r_\alpha)\sqrt{\log M}$ for star weights and $(r_g+r_\alpha)\sqrt{M\log n}$ for convex weights, where $r_g$ and $r_\alpha$ are the nuisance rates.
\medskip
\noindent\emph{Construction.}
For each fold $k$, train on $I_k^c$ a regression menu
$\mathcal{F}^{(-k)}=\{\hat f_1^{(-k)},\dots,\hat f_M^{(-k)}\}\subseteq\mathcal{H}$
and a representer menu
$\mathcal A^{(-k)}=\{\hat a_1^{(-k)},\dots,\hat a_M^{(-k)}\}\subseteq\mathcal{H}$.
Candidate~$j$ is one recipe---an encoder, a prompt, a head---refitted on each training sample.
Index the segments of the star hull by
\[
\mathcal T:=\{1,\dots,M\}^2\times[0,1],
\qquad
f_\tau^{(-k)}:=\lambda\hat f_a^{(-k)}+(1-\lambda)\hat f_b^{(-k)}
\quad\text{for }\tau=(a,b,\lambda)\in\mathcal T,
\]
and define $a_\tau^{(-k)}$ from $\mathcal A^{(-k)}$ in the same way.
The index $\tau$ is common to all folds; the functions it selects are not.
The pooled out-of-fold objectives are
\begin{align}
\bar R_n(\tau)
&:=\frac1n\sum_{k=1}^K\sum_{i\in I_k}\bigl\{Y_i-f_\tau^{(-k)}(W_i)\bigr\}^2,
\label{eq:pooled-reg}\\
\bar L_n(\tau)
&:=\frac1n\sum_{k=1}^K\sum_{i\in I_k}\Bigl\{a_\tau^{(-k)}(W_i)^2-2\,m\bigl(O_i,a_\tau^{(-k)}\bigr)\Bigr\}.
\label{eq:pooled-riesz}
\end{align}
\begin{definition}[Pooled star estimator]
\label{def:pooled-star}
Choose a center $\tilde a\in\arg\min_{1\le j\le M}\bar R_n(j,j,1)$ and a segment
\[
\hat\tau_g\in\arg\min\bigl\{\bar R_n(\tau):\ \tau=(\tilde a,b,\lambda)\in\mathcal T\bigr\}.
\]
Set $\hat g^{(-k)}:=f^{(-k)}_{\hat\tau_g}$ for every $k$.
Define $\hat\tau_\alpha$ from $\bar L_n$ in the same way, and set $\hat\alpha^{(-k)}:=a^{(-k)}_{\hat\tau_\alpha}$.
The estimators $\hat\theta_0$ and $\hat\sigma_0^2$ are then computed from these nuisances exactly as in Section~\ref{sec:pipeline-inference}.
\end{definition}
The segment search keeps its closed form: \eqref{eq:lambda-reg} and \eqref{eq:lambda-riesz} apply with $\mathbb{P}_n$ replaced by the pooled out-of-fold average.
Equivalently, the pooled star estimator is the star estimator of Definition~\ref{def:star-estimator}, computed from the augmented observations $(O_i,k(i))$ and the stacked candidates $(o,k)\mapsto\hat f_j^{(-k)}(o)$, $j=1,\dots,M$.
Write $p_k:=n_k/n$.
The natural benchmark is the best single recipe, measured by its fold-averaged error:
\begin{equation}
e_g^2:=\min_{1\le j\le M}\sum_{k=1}^K p_k\left\lVert \hat f_j^{(-k)}-g_0 \right\rVert_{P,2}^2,
\qquad
e_\alpha^2:=\min_{1\le j\le M}\sum_{k=1}^K p_k\left\lVert \hat a_j^{(-k)}-\alpha_0 \right\rVert_{P,2}^2.
\label{eq:oracle-errors}
\end{equation}
\begin{proposition}[Pooled star aggregation: oracle inequalities]
\label{prop:pooled-oracle}
Let $K$ be fixed, and let $n_k\ge n/(2K)$ for every $k$.
Suppose that, for every $k$ and conditional on $\{O_i:i\in I_k^c\}$, the menu $\mathcal{F}^{(-k)}$ satisfies the conditions of Proposition~\ref{prop:star-finite-reg} and the menu $\mathcal A^{(-k)}$ satisfies those of Proposition~\ref{prop:star-finite-riesz}, with constants $(B_Y,B_F,B_A,C_m,B_m)$ that do not depend on $k$.
Then there is a universal constant $C$ such that, for every $\delta\in(0,1)$, each of the following bounds holds with probability at least $1-\delta$:
\begin{align}
\sum_{k=1}^K p_k\left\lVert \hat g^{(-k)}-g_0 \right\rVert_{P,2}^2
&\le
e_g^2+C\,K\,(B_Y+B_F)^2\,\frac{\log(KM/\delta)}{n},
\label{eq:pooled-oracle-reg}\\
\sum_{k=1}^K p_k\left\lVert \hat\alpha^{(-k)}-\alpha_0 \right\rVert_{P,2}^2
&\le
e_\alpha^2+C\,K\,(B_A^2+C_m+B_m)\,\frac{\log(KM/\delta)}{n}.
\label{eq:pooled-oracle-riesz}
\end{align}
\end{proposition}
The remainder matches that of Propositions~\ref{prop:star-finite-reg} and~\ref{prop:star-finite-riesz} up to the factor $K$, which comes from treating each fold separately.
Unlike those propositions, Proposition~\ref{prop:pooled-oracle} needs no independence between the menu and the data used to aggregate it.
The proof uses only that each fold's candidates are independent of that fold's observations, the same structure that underlies oracle inequalities for cross-validated selectors \cite{vandervaart2006oracle,vanderlaan2007super}.
Inference needs one more ingredient.
With strict cross-fitting, the empirical-process remainder in the DML expansion vanishes by conditioning alone.
Here the weights see the evaluation fold, so the remainder must be controlled uniformly over the star segments that the weights could select.
\begin{assumption}[Regularity for pooled star aggregation]
\label{ass:pooled-star}
Let $K$ be fixed, let $n_k\ge n/(2K)$ for every $k$, and let $M=M_n\ge2$.
The constants $B_Y$, $B$, $B_m$, and $C_m$ below do not depend on $n$.
\begin{enumerate}[label=(\roman*)]
\item \emph{Boundedness.}
$|Y|\le B_Y$ almost surely; every element of $\mathcal{F}^{(-k)}\cup\mathcal A^{(-k)}$, as well as $g_0$ and $\alpha_0$, has sup norm at most $B$; and $|m(O,f)|\le B_m$ almost surely for every $f\in\mathcal{F}^{(-k)}\cup\mathcal A^{(-k)}\cup\{g_0\}$ and every $k$.
\item \emph{Mean-square continuity.}
$\mathrm{E}[m(O,h)^2]\le C_m\left\lVert h \right\rVert_{P,2}^2$ for all $h\in\mathcal{H}$.
\item \emph{Rates.}
$e_g=O_\mathrm{P}(a_{g,n})$ and $e_\alpha=O_\mathrm{P}(a_{\alpha,n})$ for deterministic sequences such that, with
$r_{g,n}:=a_{g,n}+\sqrt{\log M/n}$ and $r_{\alpha,n}:=a_{\alpha,n}+\sqrt{\log M/n}$,
\begin{equation}
\sqrt n\,r_{g,n}\,r_{\alpha,n}\to0
\qquad\text{and}\qquad
\sqrt{\log M}\,\bigl(r_{g,n}+r_{\alpha,n}\bigr)\to0.
\label{eq:pooled-rates}
\end{equation}
\item \emph{Nondegeneracy.}
$\sigma_0^2:=\mathrm{Var}\bigl(\psi^\star(O;g_0,\alpha_0)\bigr)$ is bounded away from zero.
\end{enumerate}
\end{assumption}
\begin{theorem}[DML inference with pooled star aggregation]
\label{thm:pooled-star}
Suppose Assumption~\ref{ass:pooled-star} holds, and let $\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$ be the pooled star estimators of Definition~\ref{def:pooled-star}.
Then $\hat\theta_0$ satisfies the expansion \eqref{eq:pipeline-expansion},
\[
\sigma_0^{-1}\sqrt n\bigl(\hat\theta_0-\theta_0\bigr)\Rightarrow N(0,1),
\]
and $\hat\sigma_0^2$ is consistent for $\sigma_0^2$.
Consequently, $\mathrm{P}\{\theta_0\in\mathrm{CI}_0(1-\gamma)\}=1-\gamma+o(1)$.
\end{theorem}
The first condition in \eqref{eq:pooled-rates} is the product-rate template \eqref{eq:product-rate-template} with the star remainder of Proposition~\ref{prop:pooled-oracle}.
The second condition is the price of pooling.
It controls the empirical-process remainder uniformly over the star segments whose errors are of order $r_{g,n}$ and $r_{\alpha,n}$.
That set consists of at most $M^4$ pieces of dimension two, one for each pair of segments, so its complexity grows like $\sqrt{\log M}$.
If $a_{g,n}a_{\alpha,n}=o(n^{-1/2})$ and $a_{g,n}+a_{\alpha,n}=O(n^{-1/4})$, both conditions reduce to $\log M=o(\sqrt n)$.
The menu may therefore grow faster than any power of $n$, as it may for star aggregation inside the folds.
\begin{remark}[What must stay inside the folds]
\label{rem:inside-folds}
Only the aggregation step crosses folds.
Each candidate must still be trained on $I_k^c$ alone; otherwise the candidates themselves, and not merely a low-dimensional choice among them, would depend on the evaluation fold. We can think of this as a data contamination issue.
Two consequences follow.
First, the empirical-process part of the proof bounds a supremum over the localized star hulls, so it applies to any rule that selects a star segment, provided the selected aggregates attain the rates in Assumption~\ref{ass:pooled-star}(iii).
Second, a discrete choice made on the full sample, such as a prompt or a checkpoint, is covered by adding the rejected alternatives to the menu, at the cost of a larger $M$.
\end{remark}
\medskip
\noindent\emph{Convex weights.}
The same design applies to convex aggregation.
Let $\Delta_M:=\{w\in[0,1]^M:\sum_jw_j=1\}$ and, for $w\in\Delta_M$, write $f_w^{(-k)}:=\sum_jw_j\hat f_j^{(-k)}$ and $a_w^{(-k)}:=\sum_jw_j\hat a_j^{(-k)}$.
Let $\bar R_n(w)$ and $\bar L_n(w)$ be the pooled objectives \eqref{eq:pooled-reg}--\eqref{eq:pooled-riesz} with $f_w^{(-k)}$ and $a_w^{(-k)}$ in place of $f_\tau^{(-k)}$ and $a_\tau^{(-k)}$.
\begin{definition}[Pooled convex estimator]
\label{def:pooled-convex}
Choose $\hat w_g\in\arg\min_{w\in\Delta_M}\bar R_n(w)$ and $\hat w_\alpha\in\arg\min_{w\in\Delta_M}\bar L_n(w)$, and set $\hat g^{(-k)}:=f^{(-k)}_{\hat w_g}$ and $\hat\alpha^{(-k)}:=a^{(-k)}_{\hat w_\alpha}$ for every $k$.
\end{definition}
Both problems are convex quadratic programs over the simplex.
The benchmark is now the best convex combination, measured by its fold-averaged error:
\begin{equation}
\bar e_g^2:=\min_{w\in\Delta_M}\sum_{k=1}^K p_k\left\lVert f_w^{(-k)}-g_0 \right\rVert_{P,2}^2,
\qquad
\bar e_\alpha^2:=\min_{w\in\Delta_M}\sum_{k=1}^K p_k\left\lVert a_w^{(-k)}-\alpha_0 \right\rVert_{P,2}^2.
\label{eq:oracle-errors-convex}
\end{equation}
\begin{proposition}[Pooled convex aggregation: oracle inequalities]
\label{prop:pooled-convex-oracle}
Let $K$ be fixed, let $n_k\ge n/(2K)$ for every $k$, and suppose Assumption~\ref{ass:pooled-star}(i)--(ii) hold.
Then there is a universal constant $C$ such that, for every $\delta\in(0,1)$, each of the following bounds holds with probability at least $1-\delta$:
\begin{align}
\sum_{k=1}^K p_k\left\lVert \hat g^{(-k)}-g_0 \right\rVert_{P,2}^2
&\le
\bar e_g^2+C\,K\,(B_Y+B)^2\,\frac{M\log(3n)+\log(K/\delta)}{n},
\label{eq:pooled-convex-reg}\\
\sum_{k=1}^K p_k\left\lVert \hat\alpha^{(-k)}-\alpha_0 \right\rVert_{P,2}^2
&\le
\bar e_\alpha^2+C\,K\,(B^2+C_m+B_m)\,\frac{M\log(3n)+\log(K/\delta)}{n}.
\label{eq:pooled-convex-riesz}
\end{align}
\end{proposition}
\begin{theorem}[DML inference with pooled convex aggregation]
\label{thm:pooled-convex}
Suppose Assumption~\ref{ass:pooled-star}(i), (ii), and~(iv) hold, and let $\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$ be the pooled convex estimators of Definition~\ref{def:pooled-convex}.
Suppose $\bar e_g=O_\mathrm{P}(a_{g,n})$ and $\bar e_\alpha=O_\mathrm{P}(a_{\alpha,n})$ for deterministic sequences such that, with
$\bar r_{g,n}:=a_{g,n}+\sqrt{M\log n/n}$ and $\bar r_{\alpha,n}:=a_{\alpha,n}+\sqrt{M\log n/n}$,
\begin{equation}
\sqrt n\,\bar r_{g,n}\,\bar r_{\alpha,n}\to0
\qquad\text{and}\qquad
\sqrt{M\log n}\,\bigl(\bar r_{g,n}+\bar r_{\alpha,n}\bigr)\to0.
\label{eq:pooled-convex-rates}
\end{equation}
Then the conclusions of Theorem~\ref{thm:pooled-star} hold.
\end{theorem}
Convex aggregation pays for its richer class twice.
Its estimation term grows like $M\log n/n$ rather than $\log M/n$, and pooling costs a factor $\sqrt{M\log n}$ rather than $\sqrt{\log M}$.
If $a_{g,n}a_{\alpha,n}=o(n^{-1/2})$ and $a_{g,n}+a_{\alpha,n}=O(n^{-1/4})$, both conditions in \eqref{eq:pooled-convex-rates} reduce to $M\log n=o(\sqrt n)$.
Up to the logarithm, this is the condition $M=o(\sqrt n)$ that proper convex aggregation already requires after \eqref{eq:l1-radius}; pooling the weights across folds costs nothing further.
In return, $\bar e_g\le e_g$ and $\bar e_\alpha\le e_\alpha$: a convex combination can approximate the projections better than any single recipe.
\begin{remark}[Computation]
\label{rem:computation}
Star aggregation inside the folds with an internal split trains each candidate on part of each training fold and aggregates on the rest; nested cross-fitting with $K'$ inner folds requires $KK'$ fits of every candidate.
Pooled aggregation requires $K$ fits, the fewest that cross-fitting allows, and fits the weights to all $n$ out-of-fold predictions.
With fine-tuned foundation models, the saving can decide whether the analysis is feasible.
\end{remark}
\section{Sensitivity inference for full-information targets}
\label{sec:bounds}
\subsection{Identified regions from representation loss}
Target adaptivity gives a pointwise route from the pipeline target to $\theta^\star$.
The omitted-information identity gives a complementary route that remains informative when representation loss is material, as one can specify transparent bounds on its strength and alignment, and the theory maps those restrictions into an identified region for the full-information target.
This section develops that region and shows how to attach sampling uncertainty to its endpoints.
Throughout this section we focus on the fine-tuned setting of Section~\ref{sec:fine-tuned}.
The relevant ``short'' target is the population pipeline target $\theta_0$ from Definition~\ref{def:pipeline-target}.
For the fixed-embedding specialization, take $\Phi=\{\phi\}$ and require
\[
\{g_\phi,\alpha_\phi\}\subseteq\mathcal H.
\]
This requires the head space to represent the two relevant short nuisances; it need not contain every square-integrable function of $W_\phi$. Since $\mathcal H\subseteq L^2(P_{W_\phi})$, the projection identities then give $(\theta_0,g_0,\alpha_0)=(\theta_\phi,g_\phi,\alpha_\phi)$.
The starting point is the omitted-information identity for the pipeline target.
Proposition~\ref{prop:pipeline-mixed-bias}(iii) gives
\[
\theta^\star-\theta_0=\mathrm{E}\bigl[(g^\star-g_0)(\alpha^\star-\alpha_0)\bigr].
\]
Following \cite{chernozhukov2024long}, we adopt their long/short notation and write
\[
(\theta,\theta_s,g,g_s,\alpha,\alpha_s)
:=
(\theta^\star,\theta_0,g^\star,g_0,\alpha^\star,\alpha_0).
\]
Then the identity above reads $\theta-\theta_s=\mathrm{E}[(g-g_s)(\alpha-\alpha_s)]$.
To convert the exact identity into an estimable sensitivity analysis, we use the factorization in \cite{chernozhukov2024long}: an identified scaling term is multiplied by interpretable sensitivity parameters that describe the unobserved differences $(g-g_s)$ and $(\alpha-\alpha_s)$.
One convenient factorization is
\begin{equation}
\theta-\theta_s = \rho\, C_Y\, C_D\, S,
\label{eq:bias-factorization}
\end{equation}
where
\[
S^2 := \mathrm{E}\!\left[(Y-g_s(W))^2\right]\;\mathrm{E}\!\left[\alpha_s(W)^2\right],
\qquad
\rho:=\frac{\mathrm{E}[(g-g_s)(\alpha-\alpha_s)]}{\left\lVert g-g_s \right\rVert_{P,2}\,\left\lVert \alpha-\alpha_s \right\rVert_{P,2}}\in[-1,1],
\]
with $\rho:=0$ if either norm vanishes. Here $\rho$ measures the normalized $L^2(P)$ alignment of the omitted components; it coincides with their correlation when both components are centered and have nonzero variance. Centering is not required for this definition.
The sensitivity parameters $(C_Y,C_D)$ use the following uncentered projection $R^2$ ratios:
\[
C_Y^2 := R^2_{Y-g_s\sim g-g_s}
=
\frac{\mathrm{E}[(g-g_s)^2]}{\mathrm{E}[(Y-g_s)^2]},
\qquad
C_D^2 :=
\frac{1-R^2_{\alpha\sim \alpha_s}}{R^2_{\alpha\sim \alpha_s}}
=
\frac{\mathrm{E}[(\alpha-\alpha_s)^2]}{\mathrm{E}[\alpha_s^2]}.
\]
The term $S$ depends only on the short objects $(g_s,\alpha_s)$ and is therefore identified; by contrast, $(\rho,C_Y,C_D)$ summarize the magnitude and direction of information loss.
Given user-chosen bounds $|\rho|\le\bar\rho\in[0,1]$, $C_Y\le\bar C_Y$, and $C_D\le\bar C_D$ with $\bar C_Y,\bar C_D\ge0$, equation \eqref{eq:bias-factorization} yields the interval of feasible full-information targets
\begin{equation}
\theta
\in
\Bigl[\theta_s-\bar\rho\,\bar C_Y\,\bar C_D\, S,\;\;
\theta_s+\bar\rho\,\bar C_Y\,\bar C_D\, S\Bigr].
\label{eq:bounds}
\end{equation}
Translating back to our earlier notation, this is an interval for $\theta^\star$ centered at the population target of the end-to-end pipeline, namely $\theta_0$.
In the fixed-embedding special case $\Phi=\{\phi\}$ with $\{g_\phi,\alpha_\phi\}\subseteq\mathcal H$, this center reduces to $\theta_\phi$.
Once $(\bar\rho,\bar C_Y,\bar C_D)$ are fixed, the only remaining uncertainty in \eqref{eq:bounds} is sampling uncertainty in the estimable components $(\theta_s,S)$.
We now develop DML inference for these bound endpoints, following the approach of \cite{chernozhukov2024long}.
\subsection{DML inference for sensitivity bounds in the fine-tuned case}\label{sec:bounds-components}
Once a sensitivity scenario $(\bar\rho,\bar C_Y,\bar C_D)$ is fixed, \eqref{eq:bounds} shows that the only remaining sampling uncertainty is in the estimable components $(\theta_s,S)$.
In the fine-tuned setting we take $\theta_s=\theta_0$ and $(g_s,\alpha_s)=(g_0,\alpha_0)$ from Definition~\ref{def:pipeline-target}.
The fixed-embedding case is included by the specialization $\Phi=\{\phi\}$ with $\{g_\phi,\alpha_\phi\}\subseteq\mathcal H$, under which $(\theta_0,g_0,\alpha_0)=(\theta_\phi,g_\phi,\alpha_\phi)$.
Define the estimable bound components:
\[
\sigma_0^2 := \mathrm{E}\!\left[(Y-g_0(W))^2\right],
\qquad
\nu_0^2 := \mathrm{E}\!\left[\alpha_0(W)^2\right],
\qquad
S_0 := (\sigma_0^2\nu_0^2)^{1/2}.
\]
These are the fine-tuned analogues of the $(\sigma_s^2,\nu_s^2,S)$ components in \cite{chernozhukov2024long}.
In addition to the short-target score used to estimate $\theta_0$ in Section~\ref{sec:fine-tuned}, we use the following auxiliary scores:
\begin{align}
\psi_{\theta_0}(O;\theta_0,g,\alpha)
&:= m(O,g)+\alpha(W)\{Y-g(W)\}-\theta_0,
\label{eq:score-theta-0}
\\
\psi_{\sigma_0^2}(O;\sigma_0^2,g)
&:= \{Y-g(W)\}^2-\sigma_0^2,
\label{eq:score-sigma-0}
\\
\psi_{\nu_0^2}(O;\nu_0^2,\alpha)
&:= \{2m(O,\alpha)-\alpha(W)^2\}-\nu_0^2.
\label{eq:score-nu-0}
\end{align}
As in \cite{chernozhukov2024long}, these scores are Neyman-orthogonal with respect to perturbations of $(g,\alpha)$ at $(g_0,\alpha_0)$.
Orthogonality for $\psi_{\sigma_0^2}$ uses that $g_0$ is an $L^2(P)$ projection, and orthogonality for $\psi_{\nu_0^2}$ uses linearity of $m$ together with the restricted Riesz property in Proposition~\ref{prop:pipeline-mixed-bias}(i).
Using the same fold partition $\{I_k\}_{k=1}^K$ as in Section~\ref{sec:fine-tuned}, run the end-to-end pipeline on each training sample $I_k^c$ to obtain fitted functions
$\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$, intended to approximate the population projections $g_0$ and $\alpha_0$,
where the pipeline may internally fine-tune a fold-specific representation $\hat\phi^{(-k)}$.
Define the cross-fitted score evaluations
\[
\hat\psi_i:=\psi^\star\bigl(O_i;\hat g^{(-k(i))},\hat\alpha^{(-k(i))}\bigr),
\qquad
\hat\theta_0:=\frac{1}{n}\sum_{i=1}^n\hat\psi_i.
\]
To estimate $(\sigma_0^2,\nu_0^2,S_0)$, set
\begin{align}
\hat\sigma_0^2
&:= \frac{1}{n}\sum_{k=1}^K\sum_{i\in I_k}\Bigl\{Y_i-\hat g^{(-k)}(W_i)\Bigr\}^2,
\label{eq:sigmahat-pipeline}
\\
\hat\nu_0^2
&:= \frac{1}{n}\sum_{k=1}^K\sum_{i\in I_k}\Bigl\{2m(O_i,\hat\alpha^{(-k)})-\hat\alpha^{(-k)}(W_i)^2\Bigr\},
\label{eq:nuhat-pipeline}
\\
\hat S_0
&:= (\hat\sigma_0^2\hat\nu_0^2)^{1/2}.
\label{eq:Shat-pipeline}
\end{align}
Fix a sensitivity scenario $(\bar\rho,\bar C_Y,\bar C_D)$ and write
$
c:=\bar\rho\,\bar C_Y\,\bar C_D.$
The population endpoints in \eqref{eq:bounds} can then be written as
$\theta_0^{\pm}(c):=\theta_0\pm cS_0,$
with plug-in estimators
\begin{equation}
\hat\theta_0^{\pm}(c):=\hat\theta_0\pm c\hat S_0.
\label{eq:theta-pm-hat-pipeline}
\end{equation}
To obtain root-$n$ inference for these endpoints, we require $L^2(P)$ rates for the fitted pipeline nuisances.
As in Theorem~\ref{thm:pipeline-dml}, the condition is placed directly on the end-to-end fitted functions $(\hat g^{(-k)},\hat\alpha^{(-k)})$, allowing the intermediate tuning parameters $\hat\phi^{(-k)}$ to remain fold-specific and nonsmooth.
\begin{assumption}[Regularity for sensitivity bounds with a fine-tuned pipeline]
\label{ass:pipeline-bounds-dml}
Let $\mathcal H$ and $(g_0,\alpha_0,\theta_0)$ be as in Definition~\ref{def:pipeline-target}.
Assume that, for some $q>4$, $Y,g_0,\alpha_0\in L^q(P)$.
Assume that $m$ is $L^2$-bounded on $\mathcal H$ in the sense of Lemma~\ref{lem:score-conv}.
For each fold $k$, require that both $\hat g^{(-k)}$ and $\hat\alpha^{(-k)}$ belong to $\mathcal H$ and are constructed using only $I_k^c$. Also assume
\[
\max_{1\le k\le K}\left\lVert \hat g^{(-k)} \right\rVert_{P,q}
+\max_{1\le k\le K}\left\lVert \hat\alpha^{(-k)} \right\rVert_{P,q}=O_\mathrm{P}(1).
\]
These moment bounds prevent rare extreme predictions from dominating the score and allow $L^2(P)$ consistency to imply $L^4(P)$ consistency. The moment requirement on $g_0$ is explicit because projection onto a general linear space need not preserve the moments of $Y$.
Uniformly over folds,
\begin{equation}
\max_{1\le k\le K}\left\lVert \hat g^{(-k)}-g_0 \right\rVert_{P,2}=o_\mathrm{P}(n^{-1/4}),
\qquad
\max_{1\le k\le K}\left\lVert \hat\alpha^{(-k)}-\alpha_0 \right\rVert_{P,2}=o_\mathrm{P}(n^{-1/4}).
\label{eq:pipeline-bounds-rate}
\end{equation}
Finally, assume nondegeneracy: $\sigma_0^2:=\mathrm{E}[(Y-g_0(W))^2]$ and $\nu_0^2:=\mathrm{E}[\alpha_0(W)^2]$ are bounded away from zero.
\end{assumption}
\begin{lemma}[DML for the bound components]
\label{lem:dml-pipeline-bound-components}
Suppose Assumption~\ref{ass:pipeline-bounds-dml} holds.
Define the auxiliary scores
\begin{align*}
\psi_{\theta_0}(O;\theta_0,g,\alpha)
&:= m(O,g)+\alpha(W)\{Y-g(W)\}-\theta_0,\\
\psi_{\sigma_0^2}(O;\sigma_0^2,g)
&:= \{Y-g(W)\}^2-\sigma_0^2,\\
\psi_{\nu_0^2}(O;\nu_0^2,\alpha)
&:= \{2m(O,\alpha)-\alpha(W)^2\}-\nu_0^2.
\end{align*}
Then
\begin{align}
\sqrt{n}(\hat\theta_0-\theta_0)
&=
\frac{1}{\sqrt{n}}\sum_{i=1}^n\psi_{\theta_0}(O_i;\theta_0,g_0,\alpha_0)+o_\mathrm{P}(1),
\label{eq:asylin-theta-pipeline}
\\
\sqrt{n}(\hat\sigma_0^2-\sigma_0^2)
&=
\frac{1}{\sqrt{n}}\sum_{i=1}^n\psi_{\sigma_0^2}(O_i;\sigma_0^2,g_0)+o_\mathrm{P}(1),
\label{eq:asylin-sigma-pipeline}
\\
\sqrt{n}(\hat\nu_0^2-\nu_0^2)
&=
\frac{1}{\sqrt{n}}\sum_{i=1}^n\psi_{\nu_0^2}(O_i;\nu_0^2,\alpha_0)+o_\mathrm{P}(1).
\label{eq:asylin-nu-pipeline}
\end{align}
Consequently, $(\hat\theta_0,\hat\sigma_0^2,\hat\nu_0^2)$ is jointly asymptotically normal, and its asymptotic covariance matrix is consistently estimated by the empirical covariance of the corresponding cross-fitted score vector.
\end{lemma}
\begin{theorem}[DML inference for sensitivity bounds in the fine-tuned case]
\label{thm:bounds-pipeline}
Suppose Assumption~\ref{ass:pipeline-bounds-dml} holds and fix $c\ge 0$.
Let $S_0:=(\sigma_0^2\nu_0^2)^{1/2}$ and define the population endpoints $\theta_0^{\pm}(c):=\theta_0\pm cS_0$.
If $S_0>0$, then
\begin{equation}
\sqrt{n}\bigl(\hat\theta_0^{\pm}(c)-\theta_0^{\pm}(c)\bigr)
=
\frac{1}{\sqrt{n}}\sum_{i=1}^n\varphi_{0,\pm}(O_i)+o_\mathrm{P}(1)
\quad\Rightarrow\quad
N\bigl(0,\;\mathrm{E}[\varphi_{0,\pm}(O)^2]\bigr),
\label{eq:asylin-bounds-pipeline}
\end{equation}
where
\begin{equation}
\varphi_{0,\pm}(O)
:=
\psi_{\theta_0}(O;\theta_0,g_0,\alpha_0)
\pm
\frac{c}{2S_0}\Bigl\{\nu_0^2\,\psi_{\sigma_0^2}(O;\sigma_0^2,g_0)+\sigma_0^2\,\psi_{\nu_0^2}(O;\nu_0^2,\alpha_0)\Bigr\}.
\label{eq:if-bounds-pipeline}
\end{equation}
Moreover, a consistent variance estimator is obtained by replacing unknown objects in \eqref{eq:if-bounds-pipeline} with their cross-fitted estimates and taking the empirical variance.
\end{theorem}
For the Wald intervals and confidence envelope below, additionally require
\[
\mathrm{E}[\varphi_{0,-}(O)^2]>0,
\qquad
\mathrm{E}[\varphi_{0,+}(O)^2]>0.
\]
This requirement means that each endpoint has nonzero first-order sampling variation, so division by its estimated standard error is valid. Positivity of $S_0$ alone does not ensure this. Under these conditions, Theorem~\ref{thm:bounds-pipeline} provides feasible inference for the full-information sensitivity region.
Let $\hat s_-^2$ and $\hat s_+^2$ be the empirical variances of the two estimated influence functions in \eqref{eq:if-bounds-pipeline}.
Besides pointwise Wald intervals for each endpoint, a simple asymptotic $(1-\gamma)$ confidence envelope for the entire sensitivity region is
\begin{equation}
\mathcal C_{1-\gamma}(c)
:=
\left[
\hat\theta_0^-(c)-z_{1-\gamma/2}\frac{\hat s_-}{\sqrt n},
\quad
\hat\theta_0^+(c)+z_{1-\gamma/2}\frac{\hat s_+}{\sqrt n}
\right].
\label{eq:confidence-envelope-bounds}
\end{equation}
Bonferroni's inequality and the two endpoint central limit theorems imply
\[
\mathrm{P}\!\left\{
[\theta_0^-(c),\theta_0^+(c)]
\subseteq \mathcal C_{1-\gamma}(c)
\right\}
\ge 1-\gamma+o(1).
\]
Consequently, whenever the maintained sensitivity restrictions place $\theta^\star$ in the identified region, \eqref{eq:confidence-envelope-bounds} is also an asymptotically valid confidence set for $\theta^\star$.
A critical value computed from the estimated joint Gaussian law of the two endpoint scores can replace the Bonferroni value for a shorter simultaneous envelope.
The adaptive and sensitivity results therefore support two complementary reporting modes.
When target adaptivity is maintained, the usual DML estimate and standard error provide point inference on $\theta^\star$.
For a calibrated sensitivity scenario, the analyst reports the pipeline target $\theta_0$, the identified interval \eqref{eq:bounds}, and the confidence envelope \eqref{eq:confidence-envelope-bounds}.
The fixed-representation case uses the same procedure with $\theta_0=\theta_\phi$ under the containment condition $\{g_\phi,\alpha_\phi\}\subseteq\mathcal H$.
\section{Multimodal demand with AI-learned representations}
\label{sec:empirical}
We apply the framework to product demand, where text, images, and conventional attributes provide economically distinct descriptions of the same good.
The exercise uses product-level data scraped from Keepa.com,\footnote{\url{https://keepa.com/}} covering clothing and shoes sold on Amazon.
The raw data combine price histories with three sources of product information: text, product images, and conventional tabular attributes.
First differencing removes time-invariant additive product heterogeneity, while seven modality combinations reveal how the information supplied to the pipeline changes both the outcome regression and its Riesz representer.
The application demonstrates that there is a stable negative response of the inverse-rank demand proxy to price and shows how representation-specific targets, separate nuisance learning, star aggregation, and sensitivity inference fit together in one workflow.
\subsection{A continuous-treatment first-difference specification}
\label{sec:empirical-specification}
Let $R_{it}$ denote the numerical sales rank of product $i$ at date $t$.
Because a smaller rank indicates stronger sales, we orient the outcome so that larger values indicate stronger demand:
\[
Y_{it}:=-\log R_{it}=\log(1/R_{it}).
\]
Let $D_{it}$ be log price, and form adjacent changes
\[
\Delta Y_i:=Y_{it}-Y_{i,t-1},
\qquad
\Delta D_i:=D_{it}-D_{i,t-1}.
\]
Thus the outcome is a change in the log inverse-sales-rank proxy and the treatment is a change in log price.
Equivalently, one may formulate the model using $\log R_{it}$ itself, in which case all signs below are reversed.
We index product-period changes by $i$ to simplify notation.
If a product contributes several adjacent changes, all of its observations should remain in the same training or validation fold, and the variance calculation should allow for within-product dependence.
We use ``difference-in-differences'' in the broad sense of removing a product-level difference before comparing changes in outcomes and treatment intensity.
More precisely, because treatment is continuous and there is no separate treated--control indicator, the estimating equation is a first-difference panel regression:
\begin{equation}
\Delta Y_i=g_j(X_i^{(j)},\Delta D_i)+\varepsilon_i,
\qquad
\mathrm{E}[\varepsilon_i\mid X_i^{(j)},\Delta D_i]=0.
\label{eq:empirical-fd-model}
\end{equation}
Here $X_i^{(j)}$ denotes one of seven information sets,
\[
j\in\{\mathrm{txt},\mathrm{img},\mathrm{tab},
\mathrm{txt{+}img},\mathrm{txt{+}tab},\mathrm{img{+}tab},\mathrm{full}\}.
\]
Text and image features are AI-learned embeddings; the tabular representation contains conventional product attributes.
The full model uses all three modalities.
This notation maps directly into the main text.
Take the observation to be $O_i=(\Delta Y_i,\Delta D_i,Z_i)$, where $Z_i$ collects the raw product information, and let $X_i^{(j)}=x_j(Z_i)$ be the representation generated by pipeline $j$.
For a function $h(d,Z_i)$, the relevant Riesz-linear functional is
\[
m(O_i,h)
:=
\frac{\partial}{\partial d}h(d,Z_i)
\bigg|_{d=\Delta D_i}.
\]
When $h$ factors through $x_j(Z_i)$, the resulting short target is precisely the representation-specific target $\theta_{\phi_j}$ of Section~\ref{sec:warmup}.
The target for information set $j$ is the average partial derivative
\begin{equation}
\theta_j
:=
\mathrm{E}\!\left[
\frac{\partial}{\partial d}g_j(X_i^{(j)},d)
\bigg|_{d=\Delta D_i}
\right].
\label{eq:empirical-apd}
\end{equation}
Since both the outcome and price are in logarithmic units, $\theta_j$ is an elasticity-like response of the inverse-rank demand proxy.
An additional mapping from sales rank to units sold would translate this rank elasticity into a conventional unit-demand elasticity.
The label ``full'' used below means that all three \emph{observed} modalities enter the pipeline.
Accordingly, it supplies the richest observed-information benchmark, while $\theta^\star$ continues to denote the target based on all economically relevant information.
First differencing removes time-invariant product effects.
Suppose, in addition, that price changes are conditionally mean-independent of unobserved outcome changes given $X_i^{(j)}$.
Under this condition, $\theta_j$ has a causal average-response interpretation.
This is the continuous-treatment analogue of a conditional parallel-trends restriction: differencing absorbs fixed product quality, and the multimodal controls account for observed sources of heterogeneous demand evolution.
The sensitivity analysis in Section~\ref{sec:empirical-sensitivity} then quantifies how remaining time-varying demand shocks, promotions, or inventory changes could alter the conclusion.
\subsection{Outcome and Riesz learners}
\label{sec:empirical-learners}
For each information set, let $b_j(X_i^{(j)})$ denote the corresponding learned feature vector.
We use the same varying-coefficient architecture for both nuisance functions:
\begin{align}
g_j(X_i^{(j)},\Delta D_i)
&=
\Delta D_i\,\beta_j^\top b_j(X_i^{(j)})
+\delta_j^\top b_j(X_i^{(j)}),
\label{eq:empirical-outcome-learner}
\\
\alpha_j(X_i^{(j)},\Delta D_i)
&=
\Delta D_i\,\gamma_j^\top b_j(X_i^{(j)})
+\eta_j^\top b_j(X_i^{(j)}).
\label{eq:empirical-riesz-learner}
\end{align}
The matching architectures do not impose equal coefficients.
The outcome learner is trained by squared-error loss for $\Delta Y_i$, whereas the Riesz learner is fitted independently by minimizing
\begin{equation}
\mathrm{E}\!\left[\alpha_j(X_i^{(j)},\Delta D_i)^2\right]
-2\mathrm{E}\!\left[
\frac{\partial}{\partial d}\alpha_j(X_i^{(j)},d)
\bigg|_{d=\Delta D_i}
\right].
\label{eq:empirical-riesz-loss}
\end{equation}
For the average-derivative functional in \eqref{eq:empirical-apd}, the debiased score is
\begin{equation}
\psi_{ij}
=
\frac{\partial}{\partial d}\hat g_j(X_i^{(j)},d)
\bigg|_{d=\Delta D_i}
+\hat\alpha_j(X_i^{(j)},\Delta D_i)
\left\{\Delta Y_i-\hat g_j(X_i^{(j)},\Delta D_i)\right\}.
\label{eq:empirical-dr-score}
\end{equation}
Data are split into a training and validation fold.
Candidate learners, star centers, and aggregation weights are selected on the training sample, while $\hat\theta_j$ and its score variance are computed only from held-out validation observations.
The fully cross-fitted estimator studied in Sections~\ref{sec:warmup} and \ref{sec:fine-tuned} repeats the same construction across folds and averages the held-out scores, thereby using every observation for evaluation once.
By Theorem~\ref{thm:pooled-star}, the star weights in that construction may be fitted once to the pooled out-of-fold predictions, so each of the seven candidates is trained only once per fold.
\FloatBarrier
\subsection{Estimates and star aggregation}
\label{sec:empirical-results}
Table~\ref{tab:empirical-modalities} reports the average-derivative estimates.
All seven modality-specific estimates are negative, tightly estimated, and have 95\% intervals bounded away from zero; in fact every interval lies strictly below $-1$. They range from $-1.098$ for text plus image to $-1.034$ for the tabular representation.
The full multimodal estimate is $-1.069$ with standard error $0.015$.
\begin{table}[htbp]
\centering
\caption{Average partial derivatives by information set}
\label{tab:empirical-modalities}
\small
\setlength{\tabcolsep}{7pt}
\begin{tabular}{@{}lrrrr@{}}
\toprule
Information set & Estimate & Standard error & \multicolumn{2}{c}{95\% confidence interval} \\
\cmidrule(l){4-5}
& & & Lower & Upper \\
\midrule
Text & $-1.092$ & $0.015$ & $-1.120$ & $-1.063$ \\
Image & $-1.092$ & $0.014$ & $-1.120$ & $-1.065$ \\
Tabular & $-1.034$ & $0.016$ & $-1.064$ & $-1.003$ \\
Text $+$ image & $-1.098$ & $0.015$ & $-1.127$ & $-1.069$ \\
Text $+$ tabular & $-1.039$ & $0.018$ & $-1.074$ & $-1.004$ \\
Image $+$ tabular & $-1.056$ & $0.015$ & $-1.085$ & $-1.027$ \\
Full multimodal & $-1.069$ & $0.015$ & $-1.099$ & $-1.040$ \\
\addlinespace
Star aggregate & $-1.053$ & $0.016$ & $-1.086$ & $-1.021$ \\
\bottomrule
\end{tabular}
\begin{minipage}{0.92\textwidth}
\footnotesize
\emph{Note:} The outcome is the change in the oriented log-sales-rank proxy, and treatment is the change in log price. The final row aggregates the seven modality-specific candidates. The intervals are normal approximations computed from the held-out validation scores in the reported train--validation split.
\end{minipage}
\end{table}
We next apply the star-aggregation construction of Section~3.
For the outcome regression, the image-plus-tabular learner attains the smallest training loss and is selected as the center; no other candidate improves on it, so every fitted star weight equals one and the aggregate reduces to the center:
\[
\hat g_{\mathrm{star}}
=
1.0000\,\hat g_{\mathrm{img+tab}}.
\]
For the Riesz representer, the text-plus-tabular learner is the center, and the star step mixes in the full multimodal learner:
\[
\hat\alpha_{\mathrm{star}}
=0.8433\,\hat\alpha_{\mathrm{txt+tab}}
+0.1567\,\hat\alpha_{\mathrm{full}}.
\]
The two losses are allowed to select different candidates because $g$ and $\alpha$ solve different prediction problems.
Nevertheless, both selected functions lie in the shared ambient space $\mathcal H$ generated by the seven pipelines, exactly as required by the learner-design argument in Section~\ref{sec:learn-H}.
The aggregate estimates the pipeline parameter associated with $\mathcal H$ under the approximation, learning-rate, and regularity conditions stated in the theory. Common-space membership supplies the mixed-bias identity; convergence to its projections is additionally needed for inference.
Combining these two separately selected nuisances in \eqref{eq:empirical-dr-score} gives
\[
\hat\theta_{\mathrm{star}}=-1.053,
\qquad
\widehat{\mathrm{se}}(\hat\theta_{\mathrm{star}})=0.016,
\qquad
\mathrm{CI}_{0.95}=[-1.086,-1.021].
\]
At face value, a 10\% price increase is associated locally with roughly a 10.5\% decline in the inverse-rank demand index.
The narrow range across modalities is itself informative: the negative demand response is stable whether product information enters through text, images, tabular attributes, or their combinations.
Each information set still defines its own representation-based target, while, under those approximation, learning-rate, and regularity conditions, the star estimate summarizes the population pipeline target associated with their common span $\mathcal H$.
The aggregate therefore complements the modality-specific estimates by learning strong nuisances over the full candidate menu without selecting an economic estimand after inspecting the coefficient.
\FloatBarrier
\subsection{Sensitivity to omitted multimodal information}
\label{sec:empirical-sensitivity}
To calibrate the abstract sensitivity parameters of Section~\ref{sec:bounds}, we treat the full multimodal model as a richer \emph{observed-information benchmark} and the text-only model as the short specification.
This deliberate deletion exercise measures how much the two nuisance functions change when image and tabular information, whose empirical scale is observable, are withheld.
It thereby anchors the abstract sensitivity parameters to an economically recognizable loss of product information; the subsequent grid uses that scale to examine information beyond all observed modalities.
Omitting image and tabular controls changes the validation $R^2$ of the outcome learner only from $0.0314$ to $0.0297$.
The corresponding validation Riesz loss changes from $-254.24$ to $-252.02$, where a smaller loss is better.
For nested population learners, the excess-risk identities \eqref{eq:excess-risk-reg}--\eqref{eq:excess-risk-riesz} motivate the validation analogues
\[
\widehat C_Y^{2}
=
\frac{\widehat R^2_{\mathrm{full}}-\widehat R^2_{\mathrm{text}}}
{1-\widehat R^2_{\mathrm{text}}}
\approx 0.0017,
\qquad
\widehat C_D^{2}
=
\frac{\widehat{\mathcal L}_{\mathrm{text}}-\widehat{\mathcal L}_{\mathrm{full}}}
{-\widehat{\mathcal L}_{\mathrm{text}}}
\approx 0.0088.
\]
These diagnostics quantify the observed deletion and provide a benchmark for the sensitivity scenarios.
Together with the reported residual alignment $\hat\rho=0.8654$, they yield Table~\ref{tab:empirical-calibration}.
\begin{table}[htbp]
\centering
\caption{Full-versus-text calibration}
\label{tab:empirical-calibration}
\small
\begin{tabular}{@{}lrrl@{}}
\toprule
Validation criterion & Full & Text only & Calibration contrast \\
\midrule
Outcome-model $R^2$ & $0.0314$ & $0.0297$ & $R^2_{\mathrm{full}}-R^2_{\mathrm{text}}=0.0017$ \\
Riesz loss & $-254.24$ & $-252.02$ & $\mathcal L_{\mathrm{text}}-\mathcal L_{\mathrm{full}}=2.22$ \\
\midrule
\multicolumn{4}{@{}l}{Sensitivity calibration: $\widehat C_Y^{2}=0.0017$, $\widehat C_D^{2}=0.0088$, and $\hat\rho=0.8654$.} \\
\multicolumn{4}{@{}l}{Equal-strength benchmark: $r^2=(\widehat C_Y^{2}\widehat C_D^{2})^{1/2}=0.0039$.} \\
\bottomrule
\end{tabular}
\end{table}
The benchmark comparison shows that image and tabular information add little predictive content beyond text for the differenced outcome, while contributing more visibly to the representer.
This is precisely the mixed-bias logic in action: the high estimated alignment, $\hat\rho=0.8654$, is multiplied by small residual strengths, so the induced target shift remains modest.
Table~\ref{tab:empirical-sensitivity} extends the observed benchmark through a grid of scenarios for additional information.
The two comparisons play complementary roles: the deletion exercise treats the full observed pipeline as long relative to text only to calibrate an empirical scale, whereas the sensitivity grid treats that same full observed pipeline as short relative to the full-information target $\theta^\star$.
The center is the full-model estimate $\hat\theta_s=-1.069$.
For compactness, the scenarios set $C_Y^2=C_D^2=r^2$ at $1\%$, $2.5\%$, and $5\%$ and vary the alignment parameter over $\rho\in\{1,0.5,0.2\}$.
The columns $\theta^-$ and $\theta^+$ are the sensitivity endpoints from \eqref{eq:bounds}; the outer columns reproduce the reported 10th- and 90th-percentile sampling limits for those endpoints, complementing the 95\% intervals for the modality-specific targets in Table~\ref{tab:empirical-modalities}.
\begin{table}[htbp]
\centering
\caption{Sensitivity of the full-model estimate}
\label{tab:empirical-sensitivity}
\footnotesize
\setlength{\tabcolsep}{4.3pt}
\begin{tabular}{@{}clccccc@{}}
\toprule
$\rho$ & Scenario & 10th-percentile limit & $\theta^-$ & $\hat\theta_s$ & $\theta^+$ & 90th-percentile limit \\
\midrule
1.0 & Mild: $r^2=0.010$ & $-1.156$ & $-1.137$ & $-1.069$ & $-1.001$ & $-0.982$ \\
1.0 & Medium: $r^2=0.025$ & $-1.260$ & $-1.240$ & $-1.069$ & $-0.898$ & $-0.879$ \\
1.0 & Strong: $r^2=0.050$ & $-1.436$ & $-1.415$ & $-1.069$ & $-0.723$ & $-0.705$ \\
\addlinespace
0.5 & Mild: $r^2=0.010$ & $-1.122$ & $-1.103$ & $-1.069$ & $-1.035$ & $-1.016$ \\
0.5 & Medium: $r^2=0.025$ & $-1.174$ & $-1.154$ & $-1.069$ & $-0.984$ & $-0.965$ \\
0.5 & Strong: $r^2=0.050$ & $-1.262$ & $-1.242$ & $-1.069$ & $-0.896$ & $-0.877$ \\
\addlinespace
0.2 & Mild: $r^2=0.010$ & $-1.102$ & $-1.083$ & $-1.069$ & $-1.055$ & $-1.036$ \\
0.2 & Medium: $r^2=0.025$ & $-1.123$ & $-1.103$ & $-1.069$ & $-1.035$ & $-1.016$ \\
0.2 & Strong: $r^2=0.050$ & $-1.158$ & $-1.138$ & $-1.069$ & $-1.000$ & $-0.981$ \\
\bottomrule
\end{tabular}
\begin{minipage}{0.94\textwidth}
\footnotesize
\emph{Note:} In each row, $C_Y^2=C_D^2=r^2$, using the notation of Section~\ref{sec:bounds}. The interval $[\theta^-,\theta^+]$ measures uncertainty from omitted information; the outer limits additionally reflect sampling uncertainty through the reported 10th and 90th percentiles.
\end{minipage}
\end{table}
\FloatBarrier
The leading empirical conclusion is stable throughout the analysis.
Every modality-specific 95\% interval, the star-aggregate interval, and every full-information sensitivity interval in the reported grid lies below zero.
Even under the strong scenario with perfect alignment, the upper sensitivity endpoint is $-0.723$ and its reported 90th-percentile limit is $-0.705$.
The grid also reveals which conclusion is sharper: the sign is highly robust, while the magnitude responds to the assumed strength of omitted information.
Under $\rho=1$, moving from the mild to the strong scenario expands the sensitivity interval from $[-1.137,-1.001]$ to $[-1.415,-0.723]$.
Thus the multimodal estimates give strong evidence of a negative rank-based price response under the maintained differenced design, and the sensitivity region quantifies the remaining uncertainty about its economic magnitude.
\subsection{Empirical lessons}
\label{sec:empirical-takeaway}
The application implements the four main steps of the paper.
Each modality map produces a representation-specific short target; the regression and Riesz representer are trained by their own quadratic objectives; star aggregation learns the two nuisances over a common menu without forcing them to select the same candidate; and the mixed-bias identity turns a long--short benchmark comparison into a transparent sensitivity scale.
The different star choices---image-plus-tabular candidate alone for $g$, but text plus tabular mixed with the full multimodal candidate for $\alpha$---show concretely why the two nuisances should be learned separately.
The calibration gives the corresponding diagnostic lesson: image and tabular information barely change the outcome $R^2$, yet they produce a more visible change in the Riesz loss, so prediction quality alone does not summarize target preservation.
Under the maintained first-difference design, the central empirical finding is a quantitatively stable negative response across the observed representation menu and across the reported sensitivity grid.
The richest multimodal pipeline is an observed-information benchmark, and the sensitivity scenarios explicitly carry the analysis from that benchmark toward the full-information target.
Likewise, the train--validation implementation provides honest split-sample inference under the stated sampling conditions, while the cross-fitted construction in the theory offers the natural sample-efficient extension.
These scope conditions make clear how additional modalities, repeated-split analysis, or richer controls can sharpen the magnitude without obscuring the robust sign already present in the data.
\section{Conclusion}
\label{sec:conclusion}
AI-learned representations make text, images, and other rich objects usable in causal and policy analysis.
This paper shows that the resulting DML procedure has a precise, target-aware interpretation.
For Riesz-linear functionals, replacing the original covariates by a representation changes the target through one mixed-bias term: the inner product of the discarded component of the outcome regression and the discarded component of the Riesz representer.
This identity is useful because it focuses attention on the information that matters for the estimand, rather than on global reconstruction of the original covariates.
The inferential results form a natural hierarchy.
Standard cross-fitted DML provides valid confidence intervals and tests for the representation-specific target.
Target adaptivity makes the same output valid for the full-information truth, and adaptive efficiency makes it first-order equivalent to inference based on the full-information efficient influence function.
For representations tuned on the sample, the population pipeline target $\theta_0$ supplies the corresponding estimand: foldwise tuning is absorbed into the nuisance learners, so inference depends on their prediction rates rather than on a linear expansion of the tuning algorithm.
The learning results make those rate requirements operational.
Proper learning over controlled star-shaped approximations is well suited to explicit sieves and convex aggregation classes, while star aggregation provides a computationally simple and statistically sharp method for large, nonconvex menus of prompts, embeddings, and heads.
Both aggregation schemes may fit their weights to pooled out-of-fold predictions, which avoids nested sample splitting at the cost of a factor $\sqrt{\log M}$ for star weights, or $\sqrt{M\log n}$ for convex weights, in one remainder term.
Because the outcome regression and Riesz representer solve different learning problems, they may select different candidates; the shared space $\mathcal H$ preserves the orthogonal-score geometry that combines them.
Sensitivity inference completes the framework by translating economically meaningful restrictions on representation loss into an identified region for $\theta^\star$, together with confidence statements for its endpoints and for the region as a whole.
The multimodal demand application illustrates the complete workflow.
All seven representation-specific estimates and the star aggregate imply a negative inverse-rank demand response to price, and that sign remains negative across the reported sensitivity grid; the widening intervals show transparently how uncertainty about omitted time-varying information affects the magnitude.
Thus learned representations can support affirmative empirical conclusions while making the target, the efficiency claim, and the sensitivity to discarded information explicit.
\subsection*{Acknowledgments}
We are enormously grateful to Claude Opus~5.5 (Anthropic) for its help with this paper, and particularly with Section~\ref{sec:cross-fold} and the accompanying proofs in Appendix~\ref{app:proofs-cross-fold}. GPT-6 Astra was also used for proof audits.
\newpage
\bibliographystyle{plainnat}
\bibliography{AI_Embeddings_revision_2026-07-31_v3_full_authors_references}
\newpage