Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Double Machine Learning of Continuous Treatment Effects with General Instrumental Variables
\if11
\fi
\if01
{
center[center omitted — 122 chars of source]
} \fi
abstractEstimating causal effects of continuous treatments is a common problem in practice, for example, in studying average dose-response functions. Classical analyses typically assume that all confounders are fully observed, whereas in real-world applications, unmeasured confounding often persists.
In this article, we propose a novel framework for the identification of average dose-response functions using instrumental variables, thereby mitigating bias induced by unobserved confounders.
We introduce the concept of a uniform regular weighting function and consider covering the treatment space with a finite collection of open sets. On each of these sets, such a weighting function exists, allowing us to identify the average dose-response function locally within the corresponding region.
For estimation, we propose an augmented inverse probability weighted score for continuous treatments with instrumental variables under a debiased machine learning framework, \textcolor{black}{and provide practical guidance to adaptively establish regular weighting functions from the data}.
We further establish the asymptotic properties when the average dose-response function is estimated via kernel regression or empirical risk minimization.
Finally, we conduct both simulation and empirical studies to assess the finite-sample performance of the proposed methods.
{\it Keywords:}
Average dose-response function,
Continuous treatment,
Debiased machine learning,
Finite open covering,
Instrumental variable,
Kernel regression.
bibunit\section{Introduction}
Estimating causal effects of continuous treatments is a common problem in practice, for example, in studying dose-response relationships. This problem has been extensively investigated in the literature. For instance, hirano2004propensity,ImaiVanDyk2004,GalvaoWang2015 adapted generalized propensity score-based approaches to handle continuous treatments. DiazVanderLaan2013 considered global modeling of the average dose-response function (ADRF) and developed a doubly robust substitution estimator of the risk.
More recently, kennedy2017non propose a doubly robust influence function for estimating the ADRF using kernel regression techniques, whereas SemenovaChernozhukov2021,Colangelo16072025 develop doubly robust estimators for ADRFs within the debiased machine learning (DML) framework Chernozhukov2018.
Furthermore, bonvini2022fastconvergenceratesdoseresponse summarize and generalize these approaches, allowing full exploitation of the smoothness of the target curve and achieving the oracle convergence rate even when the nuisance functions must be estimated.
Moreover, FongHazlettImai2018,HulingGreiferChen2023 study covariate balancing and independence weights for estimating the ADRF,
ai2022estimation estimate the counterfactual distribution and quantile functions in continuous treatment models,
and huang2023nonparametric estimate the continuous treatment effect with measurement error.
All the above discussions on identifying continuous treatment effects rely on the assumption of no unmeasured confounders (NUC). Instrumental variable (IV) methods offer an effective approach to address unmeasured confounding. These methods rely on exogenous treatment variation generated by instruments that affect treatment assignment while remaining conditionally unrelated to latent confounding. Under the monotonicity assumption, IV methods recover causal effects for the subpopulation of compliers imbens1994identification,angrist1995two. Building on these works, vytlacil2002independence,kennedy2019robust propose estimators for the local IV effect curve, which captures treatment effects among individuals who comply when the instrument crosses a given threshold.
On the other hand, wang2018bounded,Hartwig2023 identify the average treatment effect (ATE) by imposing no-interaction assumptions between the IV and latent confounders in the treatment model, in settings where both the IV and the treatment are binary. Building on this framework, tchetgen2018marginal,michael2024instrumental extend the approach to longitudinal data, while richardson2017nonparametric,wang2023instrumental,cui2023instrumental apply this IV framework to identify and estimate causal Cox hazard ratios.
Recently, motivated by nonparametric instrumental variable methods, chen2025identificationdebiasedlearningcausal propose an additive IV condition to identify mean potential outcomes with multi-categorical treatments and general IVs, and further identify the average treatment effect under a no unmeasured common effect modifier assumption cui2021semiparametric; Independently and concurrently, similar results are obtained by dong2025marginalcausaleffectestimation.
Additionally, liu2025multiplicativeinstrumentalvariablemodel propose an alternative identification condition that rules out any multiplicative interaction between the IV and latent confounders in the treatment model,
syrgkanis2019machine estimate heterogeneous treatment effects using IVs within the DML framework,
and ye2023instrumented estimate the ATE under a difference-in-difference IV framework.
However, there is little work on leveraging IVs or other auxiliary variables to estimate the ADRF nonparametrically. Motivated by this gap, we propose a general IV framework for identifying the ADRF with continuous treatments. Specifically, we propose the definitions of the regular weighting function (RWF) and additive IV, and further show that a finite collection of RWFs suffices for identification over any compact subset of the treatment space.
We employ semiparametric theory vanderLaan2003,tsiatis2006semiparametric to derive an augmented inverse probability weighting (AIPW) score function with the mixed-bias property rotnitzky2021characterization, whose conditional expectation given the treatment exactly equals the ADRF.
Moreover, inspired by SemenovaChernozhukov2021, bonvini2022fastconvergenceratesdoseresponse, we develop a general cross-fitting algorithm under the DML framework to compute the AIPW score. Leveraging local linear kernel regression (LLKR) methods fan1996local,wasserman2006all,tsybakov2009introduction and empirical risk minimization BartlettBousquetMendelson2005,foster2023orthogonalstatisticallearning,bonvini2022fastconvergenceratesdoseresponse, we estimate the ADRF nonparametrically and locally with the assistance of an additive IV.
Overall, our approach accommodates unmeasured confounding, recovers the dose-response relationship nonparametrically, and extends the influence function of the ADRF of the NUC setting to the IV framework.
We outline the structure of this article as follows.
Section (ref) introduces the foundational assumptions under the IV framework for continuous treatments, including a high-level IV relevance condition based on the $\chi^2$-divergence, the concepts of an RWF with local stability, a uniform RWF, and an additive IV condition for continuous treatments.
Section (ref) derives an AIPW score for the ADRF, describes a cross-fitting procedure for computing the AIPW score, and employs LLKR methods for nonparametric and local estimation of the ADRF. \textcolor{black}{In addition, a hypothesis testing procedure is established for detecting whether the RWF condition is violated. In Section (ref), we provide practical guidance on how to construct coverage for a set such that each subset is associated with a corresponding uniform RWF.}
In Section (ref), we establish the convergence rate and asymptotic normality of the proposed estimator. Section (ref) reports simulation results assessing the finite-sample performance of our method. Finally, Section (ref) applies the proposed approach to investigate the effect of years of education on total annual earnings using data from the Job Training Partnership Act study.
All scripts and data required to replicate the numerical results in our paper are available in \href{https://github.com/chensy123-sys/Continuous_treatment_IV}{https://github.com/chensy123-sys/Continuous_treatment_IV}.
\section{The model framework}
\subsection{Assumptions and notation}
We begin by introducing the basic notation used throughout the paper. \textcolor{black}{Let $d_Z$, $d_L$, and $d_U$ denote three positive integers. Let
$Y \in \mathcal{Y} \subseteq \mathbb{R}$ be the outcome of interest,
$A \in \mathcal{A} \subseteq \mathbb{R}$ a continuous treatment,
$Z \in \mathcal{Z} \subseteq \mathbb{R}^{d_Z}$ an IV,
which may be discrete or continuous,
$L \in \mathcal{L} \subseteq \mathbb{R}^{d_L}$ the observed confounders,
and $U \in \mathcal{U} \subseteq \mathbb{R}^{d_U}$ the unobserved confounders.}
We assume that \(\mathcal{A}\), the support of \(A\),
is a closed subset of \(\mathbb{R}\).
Denote \(O := [L, Z, A, Y]\) as the observed data.
The observations consist of independent and identically distributed samples
\(\{O_i\}_{i=1}^n\).
Under the potential outcomes framework,
let \(Y(a)\) denote the potential outcome that would be observed if \(A = a\).
Let \(p_{A \mid Z, U, L}(a \mid z, u, l)\) denote the conditional probability density function
of \(A\) given \(Z, U, L\), \(p_A(a)\) its marginal density, and \(P_L(l)\) the marginal distribution of \(L\).
We use similar notation for other conditional and marginal distributions as well.
In addition, define \(\mathcal{B}(a,h) := (a-h, a+h)\). For any subset $\mathcal{N}\subseteq\mathcal{A}$,
Let \(\mathring{\mathcal{N}}\) denote the interior of \(\mathcal{N}\), \(\overline{\mathcal{N}}\) denote the closure of \(\mathcal{N}\), and \(f'(a)\) denote the derivative of \(f(a)\).
Throughout this paper, we assume that $p_{A|L}(a|L)>0$ for all $a\in\mathring{\mathcal{A}}$ almost surely, so that it shares the same support as the marginal distribution $p_A(a)$.
Next, we introduce several fundamental assumptions for the continuous treatment setting with IVs.
\begin{assumption}[Consistency]
$Y=Y(A)$.
\end{assumption}
\begin{assumption}[Latent ignorability]
For any $a\in\mathcal{A}$, $Y(a)\perp\!\!\!\perp \{A,Z\}\mid U,L$.
\end{assumption}
\begin{assumption}[IV independence]
$Z\perp\!\!\!\perp U\mid L$.
\end{assumption}
\begin{assumption}[Continuity and uniform boundedness]
For all $a\in\mathring{\mathcal{A}}$, $p_{A| Z,L}(a| Z,L)$ is continuous with respect to $a$ almost surely. In addition, there exists $M(a)$, such that $p_{A| Z,L}(a| Z,L)\leq M(a)$ almost surely.
\end{assumption}
Assumption (ref) is a standard consistency condition in causal inference, stating that the observed outcome equals the potential outcome corresponding to the received treatment.
Assumption (ref) states that, conditional on observed covariates and unmeasured confounders, \(Y(a)\) is independent of both the treatment and the IVs, implying that \(Z\) affects \(Y\) only through \(A\).
Assumption (ref) requires that the IV is independent of the unmeasured confounders.
Finally, Assumption (ref) assumes that the densities associated with the IVs and measured confounders do not vary sharply within the \(\mathring{\mathcal{A}}\).
Moreover, identification of the ADRF also requires that the treatment and instrument exhibit sufficient variability across the covariate space.
The following two assumptions ensure that the treatment is sufficiently positive across covariates and that the IV is relevant and informative for the treatment.
\begin{assumption}[Positivity]
For any $a\in\mathring{\mathcal{A}}$, there exists a $\epsilon_1(a)>0$ such that
$p_{A\mid L}(a\mid L)\geq \epsilon_1(a) \text{ almost surely.}$
\end{assumption}
\begin{assumption}[IV relevance]
Define the $\chi^2$-divergence between two distributions $q(Z)$ and $p(Z)$ as $\chi^2[q(\cdot)\| p(\cdot)]:=\int (\frac{q(z)}{p(z)} -1)^2\text p(z)d z$.
For any $a\in\mathring{\mathcal{A}}$, there exists a $\epsilon_2(a)>0$ such that
$\chi^2[p_{Z|A,L}(\cdot|a,L)\| p_{Z|L}(\cdot|L)]\geq \epsilon_2(a)$ almost surely.
\end{assumption}
\begin{remark}
We impose uniform positivity and IV relevance conditions given $L$.
For identification, it suffices that for all $a \in \mathring{\mathcal{A}}$,
both $\chi^2\!\left[p_{Z\mid A,L}(\cdot\mid a,L)\,\|\, p_{Z\mid L}(\cdot\mid L)\right]$
and $p_{A\mid L}(a\mid L)$ are almost surely nonzero.
\end{remark}
Assumption (ref) states that $p_{A\mid L}(a\mid L)$ is uniformly bounded away from zero almost surely, ensuring the stability of the proposed estimator.
Assumption (ref) can be viewed as an IV relevance condition, which requires that the instrument exerts a non-negligible influence on the treatment at any $a\in\mathring{\mathcal{A}}$.
Notably, we just set $\epsilon_0(a):=\min\{\epsilon_1(a),\epsilon_2(a)\}$ when it causes no misunderstanding.
We next provide an equivalent characterization of Assumption (ref), which provides an intuitive visualization tool to understand Proposition (ref).
\begin{proposition}
Under Assumptions (ref) and (ref), Assumption (ref) holds if and only if for any $a\in\mathring{\mathcal{A}}$, there exists a $\epsilon_3(a)>0$ such that
$\mathrm{Var}[p_{A|Z,L}(a\mid Z,L) \mid L] \geq \epsilon_3(a) \text{ almost surely.}$
\end{proposition}
\textcolor{black}{Interestingly, in the continuous treatment setting, binary IVs generally fail to satisfy Assumption (ref), which is fundamentally different from the discrete treatment case chen2025identificationdebiasedlearningcausal. This limitation is formalized in the following proposition.}
\begin{proposition}[Weakness of a binary IV]
Under Assumptions (ref) and (ref), let $Z$ be a binary IV and $\mathcal{A}$ a closed interval.
Further assume that, for any $z\in\mathcal{Z}$ and $l\in\mathcal{L}$, the density functions $p_{A\mid Z,L}(a\mid z,l)$ share the same support $\mathcal{A}$.
Then, for any \(l \in \mathcal{L}\), there exists \(a_0 \in \mathring{\mathcal{A}}\) such that $p_{A\mid Z,L}(a_0 \mid 0, l) = p_{A\mid Z,L}(a_0 \mid 1, l),$ thereby violating Assumption (ref) at \(A = a_0\).
\end{proposition}
\begin{figure}
\begin{subfigure}{0.33\textwidth}
\caption
\end{subfigure}
\begin{subfigure}{0.33\textwidth}
\caption
\end{subfigure}
\caption{
(a) Density curves of $p_{A| Z}(a|0),p_{A| Z}(a|1)$ when $Z$ is binary.
(b) Density curves of $p_{A| Z}(a|0),p_{A| Z}(a|1),p_{A| Z}(a|2)$ when $Z$ has three classes.
}
\end{figure}
This proposition shows that, under the conditions of Proposition (ref), a binary instrument \(Z\) can not satisfy the IV relevance condition.
For illustration, consider the case \(L = \emptyset\). In Figure (ref)(a), where \(\mathcal{A} = [0,1]\), the conditional densities \(p_{A\mid Z}(a \mid Z)\) corresponding to the two values of \(Z\) must intersect at least once within \(\mathring{\mathcal{A}} = (0,1)\), leading to a violation of Assumption (ref).
In contrast, when \(Z\) takes three categories, as shown in Figure (ref)(b), the intersection points need not coincide, making Assumption (ref) more likely to hold.
\subsection{Regular weighting functions}
In this subsection, we define a concept central to ADRF identification, closely tied to the IV relevance condition in Assumption (ref).
For clarity, all results in this subsection rely only on Assumptions (ref)--(ref).
First, for any measurable function $\pi(Z,L)$, define
$\kappa_{\pi}^o(A,L) := \mathbb{E}[\pi(Z,L) \mid A,L] - \mathbb{E}[\pi(Z,L) \mid L].$
\textcolor{black}{
We view $\kappa_{\pi}^o$ as a functional mapping $\pi$ to $\kappa_{\pi}^o(A,L)$, capturing how effectively $\pi(Z,L)$ leverages information in $Z$ to predict $A$. Values near zero indicate that $\pi$ uses little information from $Z$, while larger magnitudes suggest that $\pi$ is more informative. As this quantity later appears in the denominator of the identification formula, it must remain bounded away from zero. We thus use $\kappa_{\pi}$ to define a key concept for identifying the ADRF.}
\begin{definition}
A measurable function $\pi(Z, L): \mathcal{Z} \times \mathcal{L} \rightarrow \mathbb{R}$ is said to be a regular weighting function (RWF) for $A = a\in\mathring{\mathcal{A}}$ if it is uniformly bounded and there exists a constant $\epsilon_\pi(a) > 0$ such that
$|\kappa_{\pi}^o(a,L)| \geq \epsilon_\pi(a)$ almost surely.
\end{definition}
\textcolor{black}{
Intuitively, the existence of an RWF for $A = a$ implies that variation in the IVs has a non-negligible effect on the treatment at that level. Thus, an RWF can also be interpreted as a relevant weighting function.}
Next, it is therefore natural to ask whether such an RWF exists at a single interior point. We present a necessary and sufficient condition for its existence.
\begin{proposition}[Existence of RWFs]
Under Assumption (ref), for any $a \in \mathring{\mathcal{A}}$, an RWF for $A=a$ exists if and only if
$\kappa_{\pi_a}^o(a,L)=\mathrm{Var}[\pi_a(Z,L)\mid L]$ is uniformly bounded above and below almost surely,
where $\pi_a(Z,L) := p_{A \mid Z,L}(a \mid Z,L) / p_{A \mid L}(a \mid L)$.
Moreover, if there exists an RWF for $A=a$, then $\pi_a(Z, L)$ must be an RWF for $A=a$.
\end{proposition}
Actually, one can verify that
$\chi^2[p_{Z|A,L}(\cdot|a,L)\| p_{Z|L}(\cdot|L)]=\mathrm{Var}\left[\pi_a(Z,L)\middle| L\right].$
Proposition (ref) indicates that the IV relevance condition in Assumption (ref) is equivalent to the statement that, for each $a\in\mathring{\mathcal{A}}$, there exists a corresponding RWF. Notably, the subscript $a$ indicates that $\pi_a$ is “optimal” for a fixed $a\in\mathring{\mathcal{A}}$.
That is, if $\pi_a(Z, L)$ is not an RWF for $A = a$, then there is no other RWF at this point. This observation simplifies the search for an adoptable RWF in practice.
In fact, there exists a family of “optimal” weighting functions that perform well when used as RWFs specifically for $A=a$.
This family can be generated by multiplying $\pi_a(Z, L)$ by any function $f(L)$ that is uniformly bounded above and below.
For instance, under Assumption (ref), one may simply take
$p_{A\mid Z,L}(a\mid Z,L)$ as an “optimal” weighting function,
since it belongs to the same family of “optimal” functions as $\pi_a(Z, L)$.
\subsection{Uniform regular weighting function}
\textcolor{black}{
So far, we have derived a family of RWFs for each $a \in \mathring{\mathcal{A}}$.
However, to identify the ADRF uniformly over a subset $\mathcal{N} \subseteq \mathring{\mathcal{A}}$, it is impractical to specify separate RWFs for each $a\in\mathcal{N}$.
Instead, it is desirable that all $a \in \mathcal{N}$ share a common RWF.
Hence, we adopt a uniform RWF defined on $\mathcal{N}$ instead of pointwise RWFs for every $a \in \mathcal{N}$.
}
\begin{definition}
A measurable function $\pi(Z, L): \mathcal{Z} \times \mathcal{L} \rightarrow \mathbb{R}$ is said to be a uniform RWF (URWF) for a subset $\mathcal{N} \subseteq \mathring{\mathcal{A}}$ if it is uniformly bounded and there exists a constant $\epsilon_\pi(\mathcal{N}) > 0$ such that, for all $a \in \mathcal{N}$,
$|\kappa_{\pi}^o(a,L)|\geq \epsilon_\pi(\mathcal{N})$ almost surely.
\end{definition}
\textcolor{black}{
The definition of a URWF is necessary when identifying the ADRF over a set $\mathcal{N}\subseteq\mathring{\mathcal{A}}$ that contains an uncountable number of points.}
Once a URWF is specified, it is important to establish the conditions under which it exists.
Intuitively, if $\pi(Z, L)$ is an RWF for $A = a$, it can be extended to a URWF in a neighborhood of $a$.
\begin{proposition}[Local stability of RWFs]
Under Assumption (ref), let $\pi(Z,L)$ be an RWF for $A = a_0$ with $a_0 \in \mathring{\mathcal{A}}$, and assume that $\kappa_{\pi}^o(a,l)$ is equicontinuous at $A = a_0$ over $L$, i.e., for any $\epsilon(a_0)>0$, there exists a positive $r_0$ such that, for any $l \in \mathcal{L},\;a\in \mathcal{B}(a_0, r_0)$,
$|\kappa_{\pi}^o(a,l) - \kappa_{\pi}^o(a_0,l)| \leq \epsilon(a_0).$
Then, there exists a small $h > 0$ such that $\pi(Z,L)$ is a URWF for $\mathcal{B}(a_0, r_0)$.
\end{proposition}
Proposition (ref) implies that if $\pi(Z, L)$ serves as an RWF at $A = a_0 \in \mathring{\mathcal{A}}$, then it can also serve as a URWF for a neighborhood around $a_0$.
This demonstrates the local stability of an RWF when the corresponding $\kappa_{\pi}^o(a,l)$ is not too sharp or irregular.
So far, Proposition (ref) has established that a URWF can be constructed within a sufficiently small neighborhood under Assumptions (ref) and (ref).
A natural theoretical question is whether there exists a global URWF over the entire interior $\mathring{\mathcal{A}}$. Unfortunately, such a global URWF does not exist when $\mathcal{A}$ is a closed interval.
\begin{proposition}[Nonexistence of a URWF for $\mathring{\mathcal{A}}$]
Under Assumption (ref), let $\mathcal{A}$ be a closed interval and $\pi(Z,L)$ a measurable function such that $\kappa_{\pi}^o(a,l)$ is continuous in $a$.
Then, for any $l_0 \in \mathcal{L}$, there exists $a(l_0) \in \mathring{\mathcal{A}}$ such that
$\kappa_{\pi}^o(a(l_0), l_0) = 0.$
Moreover, if $l_0 \in \mathring{\mathcal{L}}$ and the directional derivative of $\kappa_{\pi}^o(a,l)$ with respect to $l$ is not identically zero at $(a(l_0), l_0)$, then there exists a neighborhood $\mathcal{B}(a(l_0), h) \subseteq \mathring{\mathcal{A}}$ such that $\pi(Z,L)$ is not an RWF for any $A = a \in \mathcal{B}(a(l_0), h)$.
\end{proposition}
\begin{remark}
\textcolor{black}{In the discrete case, the definitions of RWF and URWF do not differ in any essential way. The limitation described in Proposition (ref) arises only for continuous treatments and does not occur in the discrete setting, a distinction that stems from the completeness of the real line.}
\end{remark}
\begin{figure}[t]
\begin{subfigure}{0.33\textwidth}
\caption
\end{subfigure}
\begin{subfigure}{0.33\textwidth}
\caption
\end{subfigure}
\caption{
(a) Illustration of $\kappa_{\pi}^o(A,L)$ for different treatment values $A=a$ with $\pi(Z,L)=Z$.
(b) Illustration of $\kappa_{\pi}^o(A,L)$ for different treatment values $A=a$ with $\pi(Z,L)=Z^2$.
}
\end{figure}
Proposition (ref) implies that different values of $a \in \mathring{\mathcal{A}}$ actually require different choices of RWFs. For instance,
Figures (ref)(a), (b) graphically illustrate the phenomena described in Propositions (ref) and (ref).
In both figures, each curve represents the value of $\kappa_{\pi}^o(a, L)$ for a fixed $a$, and each color corresponds to a specific treatment level $a \in \mathring{\mathcal{A}}$. The concrete data-generating process is demonstrated in Appendix (ref).
In Figure (ref)(a), we set $\pi(Z, L) = Z$.
As shown, $\pi(Z, L) = Z$ acts as a URWF for $[-3, -2] \cup [2, 3]$, since all corresponding curves of $\kappa_{\pi}^o(a, L)$ remain strictly above or below the $x$-axis.
However, it fails to be an RWF for all $A = a \in [-0.2, 0.2]$, where the curves intersect the $x$-axis.
In contrast, in Figure (ref)(b), we set $\pi(Z, L) = Z^2$.
In this case, $\pi(Z, L) = Z^2$ serves as a URWF for $[-0.5, 0.5]$, but not for $A = a \in [-2.5, -2] \cup [2, 2.5]$.
Since a global URWF over $\mathring{\mathcal{A}}$ may fail to exist, it is natural to instead cover $\mathring{\mathcal{A}}$ (or its subsets) by finitely many regions, each admitting a URWF. Within each region, local identification can then be achieved using a single URWF.
In fact, based on the local existence and stability of URWFs, it is feasible to cover any compact subset of $\mathring{\mathcal{A}}$ by a finite collection of neighborhoods, each admitting a URWF.
\begin{proposition}[Finite covering number]
Under Assumptions (ref)--(ref),
for any fixed $a \in \mathring{\mathcal{A}}$ with $\pi_a(Z,L) := p_{A\mid Z,L}(a \mid Z,L)/p_{A\mid L}(a\mid L)$,
suppose that $\kappa_{\pi_a}^o(a,l)$ is equicontinuous at $a$ over $\mathcal{L}$ for all $a \in \mathring{\mathcal{A}}$.
Then, for any compact subset $\mathcal{A}_c \subseteq \mathring{\mathcal{A}}$,
there exists a finite collection of open intervals $\{\mathcal{B}(a_{m}, h_m)\}_{m=1}^M$ such that
$\mathcal{A}_c \subseteq \bigcup_{m=1}^M \mathcal{B}(a_{m}, h_m)$.
Moreover, for each $m = 1, \ldots, M$, the function $\pi_{a_{m}}(Z,L)$ serves as a URWF for $\mathcal{B}(a_{m}, h_m)$.
\end{proposition}
\textcolor{black}{This proposition essentially illustrates the main motivation behind our practical guidance for constructing URWFs. See the discussion in Section (ref) for further details.}
\subsection{Additive instrumental variables}
\textcolor{black}{While we have established the IV relevance condition and its intrinsic link to the proposed RWF and URWF framework, these components alone do not guarantee the identification of the ADRF.}
In this subsection, we introduce a key condition that plays a central role in the identification of the ADRF $\theta(a):=\mathbb{E}[Y(a)]$.
\begin{definition}[Additive IV]
For each fixed $a \in \mathring{\mathcal{A}}$, we say that $Z$ is an additive IV (AIV) for $A = a$ if there exist functions $b_a(U,L)$ and $c_a(Z,L)$ such that
$p_{A\mid Z,U,L}(a \mid Z, U, L) = b_a(U, L) + c_a(Z, L).$
Moreover, $Z$ is said to be an AIV for $A$ if, for every $a \in \mathcal{A}$, $Z$ is an AIV for $A=a$.
\end{definition}
The AIV definition is closely related to the no-interaction condition between $Z$ and $U$ in the treatment model wang2018bounded, Hartwig2023, wang2023instrumental, michael2024instrumental, and similar formulations appear in the invalid IV literature Tchetgen_Tchetgen_2021, sun2022selective, ye2024geniusmawiirobustmendelianrandomization.
Since these works focus on discrete treatments, they do not directly extend to continuous treatments. To clarify the proposed AIV condition, we derive an equivalent characterization.
\begin{proposition}[Equivalent characterization of AIV]
Under Assumptions (ref)--(ref), for each fixed $a \in \mathring{\mathcal{A}}$ and any RWF $\pi(Z,L)$ for $A=a$, define a weighting function
\begin{align}
\omega_{a,\pi}(U,L) := \frac{\mathrm{Cov}\{ p_{A|Z,U,L}(a\mid Z,U,L),\, \pi(Z,L) \mid U,L \}}
{\mathrm{Cov}\{ p_{A|Z,U,L}(a\mid Z,U,L),\, \pi(Z,L) \mid L \}}.
\end{align}
Then $\mathbb{E}[\omega_{a,\pi}(U,L)\mid L] \equiv 1$. Moreover, $Z$ is an AIV for $A = a$ if and only if, for any RWF $\pi(Z,L)$,
$\omega_{a,\pi}(U,L) \equiv 1.$
\end{proposition}
As we will see later, $\omega_{a,\pi}(U,L)$ plays an important role in the identification of ADRFs.
\begin{remark}
When the treatment variable \(A\) is discrete, we modify the definition of the weighting function \( \omega_{a,\pi}(U, L) \) by replacing \( p_{A\mid Z, U, L}(a\mid Z, U, L)\) with \(\Pr(A = a \mid Z, U, L) \), and the weighting function degenerates to that in chen2025identificationdebiasedlearningcausal:
\[\omega_{a,\pi}(U,L)= \frac{\mathrm{Cov}\{ I(A = a),\, \pi(Z,L) \mid U,L \}}
{\mathrm{Cov}\{ I(A = a),\, \pi(Z,L) \mid L \}}.\]
\end{remark}
\begin{remark}[Generating an AIV from a mixture model]
\textcolor{black}{We illustrate a setting satisfying the AIV assumption: with probability \(p_0(L)\), \(A \sim \tilde{p}_{A|Z, L}(a \mid Z, L)\) independent of \(U\); otherwise, \(A \sim \tilde{p}_{A|U, L}(a \mid U, L)\) independent of \(Z\). By the law of total probability,
$p_{A|Z,U,L}(a\mid Z,U,L) = p_0(L)\, \tilde p_{A|Z,L}(a\mid Z,L) + (1-p_0(L))\,
\tilde p_{A|U,L}(a\mid U,L)$,
so \(Z\) satisfies the AIV structure for any \(A=a\).}
\end{remark}
\textcolor{black}{Finally, we present a proposition that illustrates the generality of AIVs.
\begin{proposition}[AIV transformation]
If $Z$ is an AIV for $A$, then for any monotone and differentiable function $q(\cdot)$, $Z$ is also an AIV for $q(A)$.
\end{proposition}}
\section{Methodology}
\subsection{Identification}
In this subsection, we identify the ADRF $\theta(a):=\mathbb{E}[Y(a)]$ with the help of an AIV for each $a\in\mathring{\mathcal{A}}$. First, for a weighting function $\pi(Z,L)$, define $Z_{\pi} := \pi(Z,L)$, and the nuisance function
$$
\mu_{\pi}^o(a,l) := \frac{\mathbb{E}[Y Z_{\pi} \mid A=a,L=l] - \mathbb{E}[Y \mid A=a,L=l]\,\mathbb{E}[Z_{\pi} \mid L=l]}{\mathbb{E}[Z_{\pi} \mid A=a,L=l] - \mathbb{E}[Z_{\pi} \mid L=l]}.
$$
We formulate our main identification theorem as follows.
\begin{theorem}[Identification]
Under Assumptions (ref)--(ref), for a fixed \( a\in\mathring{\mathcal{A}} \),
suppose that $\pi(Z,L)$ is an RWF for $A=a$. Then
\begin{align}
\mathbb{E}\left[ \mu_{\pi}^o(a,L)\right]
=& \mathbb{E}[\mathbb{E}\{Y(a)\mid U,L\}\,\omega_{a,\pi}(U,L)],
\\
\textcolor{black}{\mathbb{E}\left[ \mu_{\pi}^o(a,L)\right] - \mathbb{E}[Y(a)]
=}& \textcolor{black}{\mathbb{E}\left[\mathrm{Cov}\{\mathbb{E}\{Y(a)\mid U,L\},\,\omega_{a,\pi}(U,L)\mid L\}\right].}
\end{align}
Furthermore, if $Z$ is an AIV for $A=a$, \textcolor{black}{we have $\theta(a):=\mathbb{E}[Y(a)] = \mathbb{E}[\mu_{\pi}^o(a,L)]$.}
\end{theorem}
Note that by Proposition (ref), the existence of an RWF at $A=a$ implies that $\mathrm{Var}[\pi_a(Z,L)\mid L]$ is bounded away from zero at $A=a$ almost surely. Thus, in Theorem (ref), Assumption (ref) is implicitly required for $A=a$ \textcolor{black}{rather than uniformly over $a\in \mathring{\mathcal{A}}$. Similarly, Assumption (ref) only needs to hold at $A=a$ as well.}
\begin{remark}\textcolor{black}{
For any function $\phi$, Theorem (ref) still holds if $Y$, $Y(a)$ is replaced by $\phi(Y)$, $\phi(Y(a))$ respectively. For example, setting $\phi(Y) = I(Y \leq t)$ allows us to identify the distribution function of $Y(a)$, namely $\Pr(Y(a) \leq t)$. This demonstrates that the identification framework may be extended to more general dose-response functions, such as quantile dose-response functions ai2026datadrivenuniforminferencegeneral.}
\end{remark}
\begin{remark}[Sensitivity analysis]
When $Z$ is an AIV for $A=a$, the right-hand side of Equation (ref) remains invariant across all choices of RWFs.
This invariance provides a potential foundation for testing the AIV condition under Assumptions (ref)--(ref), although such an investigation is beyond the scope of this article.
\end{remark}
\textcolor{black}{In certain applications, the AIV assumption may fail to hold. Below, we provide three remarks to discuss the interpretation of the estimand $\mathbb{E}[\mu_{\pi}^o(a, L)]$ in such cases.
\begin{remark}[Contrast identification]
Under Assumptions (ref)--(ref), for any $a_1,a_2\in\mathring{\mathcal{A}}$, assume that $\pi_1$, $\pi_2$ are two RWFs for $A=a_1,a_2$, respectively. Then, from Equation (ref),
\begin{align*}
&\mathbb{E}[\mu_{\pi_1}^o(a_1,L)] -\mathbb{E}[\mu_{\pi_2}^o(a_2,L)]
- \{\mathbb{E}[Y(a_1)]-\mathbb{E}[Y(a_2)]\}\\
=&\mathbb{E}\left[\mathrm{Cov}\left\{\mathbb{E}[Y(a_1)-Y(a_2)|U,L], \omega_{a_1,\pi_1}(U,L) \mid L\right\}\right]\\
&-\mathbb{E}\left[\mathrm{Cov}\left\{\mathbb{E}[Y(a_2)|U,L], \omega_{a_2,\pi_2}(U,L) -\omega_{a_1,\pi_1}(U,L)\mid L\right\}\right].
\end{align*}
If $\mathbb{E}[Y(a_1)-Y(a_2)|U,L]$ and $\omega_{a_2,\pi_2}(U,L)-\omega_{a_1,\pi_1}(U,L)$ do not depend on $U$, $\mathbb{E}[Y(a_1)-Y(a_2)]$ remains identifiable. This can be viewed as a natural generalization of the no unmeasured common effect modifier assumption in cui2021semiparametric. See Appendix (ref) for further discussion on a novel weak AIV assumption, i.e., $\omega_{a,\pi}(U,L)$ does not depend on $(a,\pi)$, which implies $\omega_{a_2,\pi_2}(U,L)-\omega_{a_1,\pi_1}(U,L)$ do not depend on $U$. Moreover, as shown in Corollary (ref), we note that this idea can further be extended to identify the derivative of the ADRF, \(\nabla_a \mathbb{E}[Y(a)]\).
\end{remark}
\begin{remark}[Covariate shift]
The right-hand side of Equation (ref) admits a weighted-average interpretation of $\mathbb{E}[Y(a)\mid U,L]$ when $\omega_{a,\pi}(U,L)\ge 0$, linking our framework to the assumption-lean perspective of vansteelandt2022assumption,vansteelandt2024assumption.
In particular, by Proposition (ref),
\(
p_{a,\pi}(U\mid L) := \omega_{a,\pi}(U,L)\, p_{U\mid L}(U\mid L)
\)
defines a shifted conditional distribution since $\mathbb{E}[\omega_{a,\pi}(U,L)\mid L]=1$. Let $\mathbb{E}_{a,\pi}$ denote expectation under this distribution. Then
\(
\mathbb{E}_{a,\pi}[Y(a)]
= \mathbb{E}_{a,\pi}\!\left[\mathbb{E}\{Y(a)\mid U,L\}\right]
= \mathbb{E}\!\left[\mathbb{E}\{Y(a)\mid U,L\}\,\omega_{a,\pi}(U,L)\right]
= \mathbb{E}\!\left[\mu_{\pi}^o(a,L)\right].
\)
Notably, $\omega_{a,\pi}(U,L)$ is likely to remain positive as long as the influence of $U$ on the numerator of Equation (ref) is not too large.
\end{remark}
\begin{remark}[Partial identification]
The right-hand side of Equation (ref) characterizes the gap between the estimand $\mathbb{E}[\mu_{\pi}^o(a,L)]$ and the true ADRF $\mathbb{E}[Y(a)]$. Even without the AIV assumption,
\begin{align*}
\left|\mathbb{E}[\mu_{\pi}^o(a,L)] - \mathbb{E}[Y(a)]\right|
\leq
\mathbb{E}\!\left[
\mathrm{Var}\{g_a(U,L)\mid L\}^{1/2}\right]\cdot
\mathbb{E}\!\left[
\mathrm{Var}\{\omega_{a,\pi}(U,L)\mid L\}^{1/2}\right].
\end{align*}
Hence, partial identification remains possible provided these variance terms are controlled well.
\end{remark}}
Next, we show two degenerate examples to illustrate Theorem (ref).
\begin{example}[NUC]
When $U=\emptyset$, set $Z\equiv A$ and $\pi(Z,L)=f(A)$. Equation (ref) degenerates:
\begin{align*}
\mathbb{E}[Y(a)]
= &\mathbb{E}\left[
\dfrac{ \{f(a)-\mathbb{E}[f(A)\mid L]\}\mathbb{E}[Y\mid A=a,L]}{ f(a) - \mathbb{E}[f(A)\mid L]}
\right]
= \mathbb{E}\left[\mathbb{E}[Y\mid A=a,L]\right].
\end{align*}
This demonstrates that the identification formula degenerates when there do not exist any unmeasured confounders $U$.
\end{example}
\begin{example}[Binary IV]
Assume that \( Z \) is binary, and $\pi(Z,L)=Z$ is an RWF for \( A = a \). Then,
if \( Z \) is an AIV for \( A \),
Theorem (ref) implies that
\begin{align*}
\mathbb{E}[Y(a)] =
\mathbb{E}\left[\dfrac{\mathbb{E}[YZ \mid A=a,L] - \mathbb{E}[Y \mid A=a,L]\mathbb{E}[Z \mid L]}
{\mathbb{E}[Z \mid A=a,L] - \mathbb{E}[Z \mid L]}\right].
\end{align*}
\end{example}
From Proposition (ref), when identifying the ADRF over a small subset $\mathcal{N}\subseteq\mathring{\mathcal{A}}$, one may rely on a single fixed $\pi(Z,L)$ as the URWF for $\mathcal{N}$. However, Proposition (ref) cautions against using a single fixed $\pi(Z,L)$ for the entire $\mathring{\mathcal{A}}$, as doing so may lead to critical errors.
As a compromise, by Proposition (ref), it is preferable to select different RWFs for different values of $A$.
For a graphical illustration, consider Figure (ref) again. When $A$ is close to zero, Figure (ref)(b) suggests that $\pi(Z,L)=Z^2$ is an appropriate RWF for identification. In contrast, when $A$ is far from zero, Figure (ref)(a) indicates that $\pi(Z,L)=Z$ can be used as an RWF.
Consequently, in practice, we should fix an RWF $\pi(Z,L)$ for $A = a_0$, to identify the ADRF only within a neighborhood of $a_0 \in \mathring{\mathcal{A}}$.
\subsection{Semiparametric theory}
In practice, the true form of $\mu_{\pi}^o(A,L)$ is unknown. Moreover, in the absence of parametric assumptions, the functional $\mathbb{E}[\mu_{\pi}^o(a,L)]$ lacks pathwise differentiability, which implies that root-$n$ consistent estimators for it are not readily available. To proceed, we follow the approach of kennedy2017non and leverage semiparametric theory to establish an AIPW score function $\varphi_{\pi}(O)$ that satisfies $\mathbb{E}\left[\varphi_{\pi}(O) \middle| A=a\right] = \theta(a)$ when $Z$ is an AIV for $a$.
First, we define an estimand that represents the average outcome under an intervention that randomly assigns treatment based on the marginal density $q(a)p_A(a)$:
\begin{align}
\psi_{\pi,q}^o := \int_{\mathcal{A}} \int_{\mathcal{L}} \mu_{\pi}^o(a, l)\, q(a)\, \mathrm{d}P_L(l)\, \mathrm{d}P_A(a).
\end{align}
The quantity $\psi_{\pi,q}^o$ can be interpreted as a covariate-shifted causal estimand, corresponding to the scenario where the conditional treatment distribution $p_{A|L}(a \mid l)$ is replaced by a rescaled version $q(a)p_A(a)$.
Before giving the semiparametric theory for $\psi_{\pi,q}^o$, we define nuisance functions as
$\rho_{\pi}^o(L):=\mathbb{E}[Z_{\pi}\mid L]$,
$\kappa_{\pi}^o(A,L) := \mathbb{E}[Z_{\pi}\mid A,L] - \mathbb{E}[Z_{\pi}\mid L]$,
$\eta^o(A,L):= \mathbb{E}[Y\mid A,L]$,
$\delta^o(A,L) := p_{A}(A)/p_{A|L}(A\mid L).$
We unify the nuisance functions into a nuisance vector
$\alpha_{\pi}^o(A,L) = [\mu_{\pi}^o(A,L),\rho_{\pi}^o(L), \kappa_{\pi}^o(A,L), \eta^o(A,L), \delta^o(A,L)].$
Next, we derive the efficient influence function (EIF) for $\psi_{\pi,q}^o$ in the following theorem.
\begin{theorem}[EIF of $\psi_{\pi,q}^o$]
Suppose that $\pi(Z,L)$ is a URWF over the support of $q(a)$. Then, the EIF for $\psi_{\pi,q}^o$ in the nonparametric model is given by
\begin{align*}
&\delta^o(A,L)(Z_{\pi}-\rho_{\pi}^o(L))
\left\{
\dfrac{Y-\mu_{\pi}^o(A,L)}{\kappa_{\pi}^o(A,L)}q(A)
-\int \dfrac{\eta^o(a,L) - \mu_{\pi}^o(a,L)}
{\kappa_{\pi}^o(a,L)} q(a)\mathrm d P_A(a)\right\}\\
&+q(A)\int \mu_{\pi}^o(A,l)\mathrm d P_L(l)
+\int \mu_{\pi}^o(a,L) q(a)\mathrm d P_A(a) - 2\psi_{\pi,q}^o.
\end{align*}
\end{theorem}
\begin{remark}
In kennedy2017non, they simply set $q(A) \equiv 1$. Here, we derive the EIF for a general $q(a)$ because, as shown in Proposition (ref), a URWF may not exist over $\mathring{\mathcal{A}}$. In such cases, introducing $q(a)$ effectively restricts the range of values, so that $\pi(Z,L)$ only needs to be a URWF over the support of $q(a)$, rather than over the entire $\mathring{\mathcal{A}}$.
\end{remark}
Motivated by the EIF of $\psi_{\pi,q}^o$, we establish an AIPW score function as
\begin{equation}
\begin{aligned}
&\varphi_{\pi}(O;\alpha_{\pi},P_O) := \delta(A,L)
\dfrac{(Z_{\pi}- \rho_{\pi}(L))}{\kappa_{\pi}(A,L)}\{Y-\mu_{\pi}(A,L)\}\\
&+\int \mu_{\pi}(A,l) - (z_{\pi}-\rho_{\pi}(l))\dfrac{\eta(A,l)-\mu_{\pi}(A,l)}{\kappa_{\pi}(A,l)}\mathrm d P_{O}(o).
\end{aligned}
\end{equation}
It can be readily verified that
$\mathbb{E}[\varphi_{\pi}(O; \alpha_{\pi}^o, P_L)| A = a] = \mathbb{E}[\mu_{\pi}^o(a, L)].$
\textcolor{black}{Akin to kennedy2017non, \(\varphi_{\pi}(O; \alpha_{\pi}, P_O)\) is not derived directly from standard semiparametric theory; it involves a degree of heuristic reasoning based on the EIF of \(\psi_{\pi,q}^o\).}
Notably, when $A$ is discrete, provided that $\delta^o(a,L)$ is redefined as $\Pr(A=a)/\Pr(A=a \mid L)$, we still have $\mathbb{E}[\varphi_{\pi}(O;\alpha_{\pi}, P_O) \mid A=a] = \mathbb{E}[Y(a)]$.
This indicates that, although the proposed AIPW score is designed for continuous treatments, it also applies when $A$ is discrete.
A detailed comparison with chen2025identificationdebiasedlearningcausal is provided in Appendix (ref).
Moreover, the construction naturally accommodates the degenerate case where $L$ is empty.
In this setting, the AIPW score differs slightly from Equation (ref); see Appendix (ref) for more details.
\subsection{General cross-fitting procedures}
In this subsection, we construct the AIPW score vector for \(\theta(a)\) using a general cross-fitting procedure. This procedure has previously appeared in the estimation of heterogeneous treatment effects kennedy2023optimaldoublyrobustestimation and ADRFs bonvini2022fastconvergenceratesdoseresponse.
Concretely, we randomly partition the samples into \(K\) folds, denoted by \(\{I_k\}_{k=1}^K\), and let \(I_{-k} := \bigcup_{j \neq k} I_j\) denote the complement of \(I_k\). When training the nuisance functions \(\hat\alpha_{\pi}^{(n,-k^1)}\) and \(\hat P_O^{(n,-k^2)}\) for the \(k\)-th fold, we further divide \(I_{-k}\) into \(I_{-k^1}\) and \(I_{-k^2}\). In particular, the sizes of \(I_{-k^1}\) and \(I_{-k^2}\) can be freely adjusted.
For notational simplicity, let \(\mathbb{E}_{nk}[O] := \sum_{i \in I_k} O_i/|I_k|\) denote the sample average within the \(k\)-th fold.
We further define conditional expectations based on the folds:
\(\mathbb{E}_k[O] := \mathbb{E}[O \mid O_i, i \notin I_k]\),
\(\mathbb{E}_{-k}[O] := \mathbb{E}[O \mid O_i, i \notin I_{-k}]\), and
\(\mathbb{E}_{-kl}[O] := \mathbb{E}[O \mid O_i, i \notin I_{-kl}]\) for \(l = 1, 2\).
The samples in $I_{-k^1}$ are employed to train $\hat\alpha_{\pi}^{(n,-k^1)}$, whereas those in $I_{-k^2}$ are used to train $\hat P_O^{(n,-k^2)}$.
This design guarantees that the training sets for $\hat\alpha_{\pi}^{(n,-k^1)}$ and $\hat P_O^{(n,-k^2)}$ are mutually independent.
Subsequently, we compute the AIPW score vector $[\varphi_{\pi}(O_i; \hat\alpha_{\pi}^{(n,k_i^1)}, \hat P_O^{(n,k_i^2)})]_{i=1}^n$.
The entire cross-fitting procedure is summarized in Algorithm (ref).
Following the same logic, an alternative cross-fitting strategy is provided in Appendix (ref).
\begin{algorithm}[t]
\caption{Cross-fitting procedure}
\begin{algorithmic}[1]
\State Input: Number of folds $K$ and $\{O_i\}_{i=1}^n$.
\State Randomly split the sample into $K$ folds $\{I_k\}_{k=1}^K$.
\For{$k = 1, \ldots, K$}
\State Split $I_{-k}$ into two subsamples $I_{-k^1}$ and $I_{-k^2}$.
\State Train nuisance estimators on $I_{-k^1}$:
$\hat \alpha_{\pi}^{(n,-k^1)} :=
\big[\hat\mu_{\pi}^{(n,-k^1)}, \hat\rho_{\pi}^{(n,-k^1)},
\hat\kappa_{\pi}^{(n,-k^1)}, \hat\eta^{(n,-k^1)},
\hat\delta^{(n,-k^1)}\big].$
\State Estimate the empirical distribution $\hat P_O^{(n,-k^2)}$ on $I_{-k^2}$.
\EndFor
\State Output: AIPW score vector $[\varphi_{\pi}(O_i; \hat\alpha_{\pi}^{(n,-k_i^1)}, \hat P_O^{(n,-k_i^2)}) ]_{i=1}^n$, where $O_i\in I_{k_i}$ for each $i$.
\end{algorithmic}
\end{algorithm}
\subsection{Local linear kernel regression}
In this subsection, we estimate the ADRF locally using LLKR by regressing the AIPW scores on the treatment variable.
Specifically, we select a kernel weighting function $K(a)$, bandwidth \(h\), and a target point \(a \in \mathring{\mathcal{A}}\). Let $K_h(a):=K(a/h)/h$, \(g(a) := [1,a]^T\) and \(e_{2,1} := [1,0]^T\). We then calculate the AIPW estimator \(\hat{\theta}_{\pi,h}^{(n)}(a)\) by solving:
\begin{align}
e_{2,1}^T \underset{\beta \in \mathbb{R}^2}{\arg\min}
\frac{1}{n}\sum_{k=1}^{K}
\sum_{i\in I_k}
K_h(A_i - a)
\left\{\varphi_{\pi}(O_i; \hat\alpha_{\pi}^{(n,-k^1)}, \hat P_O^{(n,-k^2)})
- g(\dfrac{A_i-a}{h})^T\beta\right\}^2.
\end{align}
Notably, the kernel \(K_h(A - a)\) restricts attention to a local neighborhood around \(a\).
As a result, the outputs of Algorithm (ref) are required to be bounded and regular only for values of \(A\) within this neighborhood, reflecting the inherently local nature of the estimation.
For bandwidth selection, classical criteria such as leave-one-out cross-validation (LOOCV),
generalized cross-validation (GCV), and the $\text{C}_{\text{p}}$-criterion must also be localized, since a URWF may not exist over $\mathring{\mathcal{A}}$ (Proposition (ref)).
Concretely, define a localized LOOCV for bandwidth selection as:
\begin{align*}
\hat h_{opt}^{(n)}(\mathcal{N})=\underset{h>0}{\arg\min} \dfrac{1}{n}\sum_{k=1}^K\sum_{i\in I_k}\left\{
\dfrac{\varphi_{\pi}(O_i;\hat\alpha_{\pi}^{(n,-k^1)}, \hat P_O^{(n,-k^2)})-\hat{\theta}_{\pi,h}^{(n)}(A_i)}
{1-w_{ni}(A_i,h)}
\right\}^2I\{A_i\in \mathcal{N}\},
\end{align*}
where $w_{ni}(A_i,h)$, defined in Appendix (ref), is the $i$-th diagonal element in the smoothing matrix.
This criterion minimizes the localized residual mean squared error (RMSE).
The localized GCV and $\text{C}_{\text{p}}$-criterion can be formulated similarly, just by focusing only on the samples in $\mathcal{N}$.
\subsection{Regular weighting function testing}
\textcolor{black}{In practice, once a weighting function $\pi(Z, L)$ is prespecified, it is natural to ask whether it serves as an RWF for any treatment level $A=a$. This motivates the development of a hypothesis testing procedure for the definition of RWF. To solve this challenging issue, we first consider a simpler hypothesis test that captures the key aspects of RWF. Concretely, we define the estimand $\zeta_{\pi}^o(a) = \mathbb{E}\!\big[p_{A\mid L}(a\mid L)\,\kappa_{\pi}^o(a,L)\big].$ We then test the null hypothesis $H_0: \zeta_{\pi}^o(a) = 0$ versus $ H_1: \zeta_{\pi}^o(a) \neq 0$.}
\textcolor{black}{Intuitively, under $H_0$, when $L$ is continuous and $\mathcal{L}$ is a connected set, and $\kappa_{\pi}^o(a, L)$ are continuous in $L$, the mean value theorem for integrals implies that there exists some $l_0 \in \mathcal{L}$ such that $p_{A|L}(a| l_0)\kappa_{\pi}^o(a,l_0) = 0.$ Under Assumption (ref), this indicates that $\pi(Z,L)$ is not an RWF for $A=a$.
Notably, $\zeta_{\pi}^o(a) = 0$ is equivalent to the definition of RWF only when $L$ is empty. In general, even if $\pi$ is not an RWF for $A=a$, it may still happen that $\zeta_{\pi}^o(a) \neq 0$.}
\textcolor{black}{The test procedure is given in Algorithm (ref), where $\Phi(\cdot)$ denotes the cumulative distribution function of the standard normal distribution, and $\gamma_{a,h}^o(L):= \mathbb{E}[K_h(A-a)\mid L]$. Intuitively, $\hat \zeta_{\pi, h}^{(n)}(a)$ estimate $\zeta_{\pi}^o(a)$ consistently. The validity of $p_{a, \pi}$ and the asymptotic property of $\hat{\zeta}_{\pi,h}^{(n)}(a)$ will be further discussed in Theorem (ref). Note that unless $L$ is empty, a small $p_{a, \pi}$ only gives evidence that $\zeta_{\pi}^o(a)\neq 0$, but does not necessarily imply that $\pi$ is a valid RWF at $A = a$.}
\begin{algorithm}[t]
\caption{Hypothesis testing procedure for $H_0: \zeta_{\pi}^o(a) = 0$.}
\begin{algorithmic}[1]
\State Input:
A prespecified weighting function $\pi$; a target point $a$; number of crossing folds $K$; testing samples $\{O_i\}_{i=1}^n$.
\State Set the bandwidth $h = \, n^{-1/4}\hat{\sigma}_{A}$, where $\hat{\sigma}_{A}$ is the sample standard deviation of $\{A_i\}_{i=1}^n$.
\State Allocate a score vector $\{S_i\}_{i=1}^n$ and randomly partition the samples into $K$ folds $\{I_k\}_{k=1}^K$.
\For{$k = 1, \ldots, K$}
\State Using the samples in $I_{-k}$, fit nuisance functions
$$\hat{\gamma}_{h,a}^{(n,-k)}(L):=\hat{\mathbb{E}}^{(n,-k)}[K_h(A-a)\mid L],\,\hat{\rho}_{\pi}^{(n,-k)}(L):=\hat{\mathbb{E}}^{(n,-k)}[\pi(Z,L)\mid L].$$
\State For each $i\in I_k$, calculate $S_{i}=(K_h(A_i-a) - \hat{\gamma}_{h,a}^{(n,-k)}(L_i))(\pi(Z_i,L_i) - \hat{\rho}_{\pi}^{(n,-k)}(L_i)).$
\EndFor
\State Calculate $\hat \zeta_{\pi, h}^{(n)}(a)= \frac{1}{n}\sum\limits_{i=1}^nS_{i}$ and $\hat \sigma_{\pi,\zeta}^{(n)}(a)^2:= \frac{h}{n}\sum\limits_{i=1}^n (S_{i}-\hat \zeta_{\pi, h}^{(n)}(a)) ^2$.
\State Output: A p-value $p_{a, \pi} := 2(1 - \Phi\{\sqrt{nh}|\hat \zeta_{\pi, h}^{(n)}(a)|/\hat \sigma_{\pi,\zeta}^{(n)}(a)\})$.
\end{algorithmic}
\end{algorithm}
\begin{algorithm}[t]
\caption{Hypothesis testing procedure for $\widetilde{H}_0$: $\pi$ is not an RWF for $A=a$.}
\begin{algorithmic}[1]
\State \textbf{Input:} A prespecified weighting function $\pi$; a target point $a$; number of partitions $D$; testing samples $\{O_i\}_{i=1}^n$.
\State Partition the samples according to the confounder $L$ into $D$ folds, denoted by $\{I_d\}_{d=1}^{D}$.
\State For each partition $I_d$, apply Algorithm (ref) using only the samples in $I_d$. Denote the resulting p-value by $p_{a,\pi}^{(d)}$.
\State \textbf{Output:} A p-value $\tilde{p}_{a,\pi}:=\max_{d=1,\ldots, D} p_{a,\pi}^{(d)}$.
\end{algorithmic}
\end{algorithm}
\textcolor{black}{Next, we consider testing the RWF assumption for a fixed $a$ and prespecified $\pi$. When $L$ is discrete, define $\zeta_{\pi}^o(a,l):= p_{A|L}(a|l)\kappa_{\pi}^o(a,l)$.
We test the null hypothesis $H_0^{(l)}: \zeta_\pi^o(a,l)=0$ versus $H_1^{(l)}: \zeta_\pi^o(a,l)\neq0$ for each $l\in\mathcal{L}$ adopting Algorithm (ref) separately, yielding $|\mathcal{L}|$ p-values $\{p_{a,\pi}^{(l)}\}_{l=1}^{|\mathcal{L}|}$. Let $\widetilde H_0 = \bigcup_l H_0^{(l)}$, exactly representing the null hypothesis that the $\pi$ is not an RWF for $A=a$.
The corresponding p-value is then defined as $\tilde p_{a,\pi} = \max_l p_{a,\pi}^{(l)}$, constituting an intersection-union test roger1996bioequivalence.}
\textcolor{black}{For continuous $L$, the RWF assumption is not directly testable due to the uncountable support of $L$. One practical approach is to partition its support into several connected clusters and apply the testing procedure in Algorithm (ref) to each cluster. We summarize this idea in Algorithm (ref), which reduces to Algorithm (ref) when $D=1$.
The procedure can thus be viewed as a test for $\widetilde{H}_1:$ $\pi$ is an RWF for $A=a$, approximating the truth when $D$ increases. Since the power decreases as $D$ increases, this approach is unsuitable for highly complex $L$, and we recommend $D \leq 10$ in practice.
Notably, the p-values $\tilde{p}_{a,\pi}$ produced by this procedure are theoretically valid only when $L$ is discrete; for continuous $L$, our partitioning approach is a practical approximation, but formal validity is not guaranteed.}
\section{Practical guidance}
\textcolor{black}{All analyses above assume a fixed RWF. In practice, to estimate the ADRF over a compact set $\mathcal{A}_c \subseteq \mathring{\mathcal{A}}$, one must first prespecify a collection of URWFs $\{\pi_{[m]}\}_{m=1}^M$ for subsets $\{\mathcal{N}_m\}_{m=1}^M$ covering $\mathcal{A}_c$.
Motivated by Propositions (ref), one practical approach is to select points $\{a_{m}\}_{m=1}^M$ uniformly over $\mathcal{A}_c$, and estimate $\pi_{[m]}(Z,L) := \hat{\pi}_{a_{m}}(Z,L)$, serving as the RWF at $A=a_{m}$.
In practice, $\pi_a(Z, L):= p_{A|Z, L}(a | Z, L) / p_{A|L}(a | L)$ can be estimated by fitting smoothing spline regressions of $K_h(A-a)$ on $(Z, L)$ and $L$ respectively to obtain estimates of the two conditional densities, and then compute their ratio to obtain $\hat{\pi}_a(Z, L)$.}
\textcolor{black}{Next, we introduce a graphical method for allocating a set $\mathcal{N}_m$ to each weighting function $\pi_{[m]}$.
First, select a sequence of points $\{a_j^*\}_{j=1}^J$ evenly spaced over $\mathcal{A}_c$, and apply the RWF testing procedure in Algorithm (ref) for each prespecified $\pi_{[m]}$ and $a_j^*$.
Then, plot $M$ curves, each displaying the p-values $\{\tilde{p}_{a_j^*,\pi_{[m]}}\}_{j=1}^J$ for a fixed $\pi_{[m]}$ (see, e.g., Figures (ref), (ref)).
Based on the resulting p-value plot, one may treat any $\{\pi_{[m]}, \mathcal{N}_m\}_{m=1}^M$ as a valid construction, provided that
(1) $\bigcup_{m=1}^M \mathcal{N}_m$ covers $\mathcal{A}_c$, and
(2) for each $m=1,\ldots,M$, the p-values $\tilde{p}_{a_j^*,\pi_{[m]}}$ are significant for all $a_j^* \in \mathcal{N}_m$.}
\textcolor{black}{In principle, $\{\pi_{[m]}, \mathcal{N}_m\}_{m=1}^M$ can be constructed from the p-value plot, provided the required conditions are satisfied.
For convenience, we also provide a simple but useful method to facilitate this construction.
Specifically, when $\pi_{a_m}$ is approximated well, by Proposition (ref), $\pi_{[m]}$ can serve as the URWF over $\mathcal{N}_m=\overline{\mathcal B(a_m,r_0)}$ for sufficiently small $r_0$, thus satisfying the second condition.
On the other side, $\bigcup_{m=1}^M\mathcal{N}_m$ can cover $\mathcal{A}_c$ when $\{a_m\}_{m=1}^M$ is sufficiently dense in $\mathcal{A}_c$.}
\textcolor{black}{With $\{\pi_{[m]},\mathcal{N}_m\}_{m=1}^M$ treated as fixed, Algorithm (ref) yields the AIPW score vector, based on which the LLKR estimator in Equation (ref) is constructed.
Finally, following bonvini2022fastconvergenceratesdoseresponse, we note that the AIPW score can also be used to estimate the ADRF via empirical risk minimization (ERM) methods, rather than solely through LLKR. See Appendix (ref) for further discussion.}
\section{Asymptotic theory}
We first establish the asymptotic theory for the LLKR estimator $\hat{\theta}_{\pi,h}^{(n)}(a)$. For any measurable function $g(L)$, $f(A,L)$, and subset $\mathcal{N}\subseteq\mathring{\mathcal{A}}$, define the norm
$\|g\|_2^2:= \int_{\mathcal{L}}|g(l)|^2\text d P_L(l)$,
$\|f\|_{\mathcal{N},2}^2:=\int_{\mathcal{L}} \sup_{a\in \mathcal{N}}|f(a,l)|^2\text d P_L(l),$
where $P_L(l)$ is the marginal distribution of $L$. Define the norm for the nuisance vector $\alpha_{\pi}$ as
$\|\alpha_{\pi}\|_{\mathcal{N},2}^2:=
\|\rho_{\pi}\|_{2}^2
+\|\mu_{\pi}\|_{\mathcal{N},2}^2
+\|\kappa_{\pi}\|_{\mathcal{N},2}^2
+\|\eta\|_{\mathcal{N},2}^2
+\|\delta\|_{\mathcal{N},2}^2.$
Next, we introduce several common regularity conditions in nonparametric analysis.
\begin{assumption}
Let $\mathcal{N}\subseteq \mathring{\mathcal{A}}$ be an interval. Assume the following conditions hold:
\begin{enumerate}[label=(\alph*), itemsep=0pt, topsep=0pt,leftmargin=0.75cm]
• $\mathcal{Z}$ and $\mathcal{Y}$ are bounded subsets of the Euclidean space.
• The bandwidth $h=h_n$ satisfies $h \to 0$ and $n h \to +\infty$ as $n \to +\infty$.
• $K(s)$ is a positive, continuous, symmetric function with support $[-1,1]$, satisfying
$\int_{\mathbb{R}} K(s)\, \mathrm{d}s = 1$,
$\int_{\mathbb{R}} K(s) s\, \mathrm{d}s = 0$,
and $\int_{\mathbb{R}} K(s) s^2\, \mathrm{d}s > 0$.
• $\theta(a)$ is twice continuously differentiable for all $a \in \mathcal{N}$.
• The conditional variance $\mathrm{Var}[\varphi_{\pi}(O;\alpha_{\pi}^o,P_O) \mid A=a]$ exists for all $a \in \mathcal{N}$.
• For each $k=1,\ldots,K$ in Algorithm (ref), $|I_{-k^2}| \to \infty$ as $n \to \infty$.
• For each $k=1,\ldots,K$, there exists a constant $\epsilon_0>0$ such that for all $n$,
$|\hat\kappa_{\pi}^{(n,-k^1)}(A,L)|\ge \epsilon_0$ for all $(A,L) \in \mathcal{N}\times \mathcal{L}$ almost surely. Moreover, all functions in $\hat\alpha_{\pi}^{(n,-k^1)}(A,L)$ are uniformly bounded across all $n,k$ pairs almost surely.
• For each $k=1,\ldots,K$ and any fixed $a \in \mathcal{N}$, $\mathbb{E}\big[\|\hat\alpha_{\pi}^{(n,-k^1)} - \alpha_{\pi}^o\|_{\mathcal{B}(a;h),2}^2\big] = o(1)$. Moreover, there exists a sequence $c(n)$ such that
$\|\hat\rho_{\pi}^{(n,-k^1)}-\rho_{\pi}^o\|_2\|\hat\eta^{(n,-k^1)}-\eta^o\|_{\mathcal{B}(a;h),2}$,
$\|\hat\rho_{\pi}^{(n,-k^1)}-\rho_{\pi}^o\|_2\|\hat\delta^{(n,-k^1)}-\delta^o\|_{\mathcal{B}(a;h),2}$,
$\|\hat\mu_{\pi}^{(n,-k^1)}-\mu_{\pi}^o\|_{\mathcal{B}(a;h),2} \|\hat\kappa_{\pi}^{(n,-k^1)}-\kappa_{\pi}^o\|_{\mathcal{B}(a;h),2}$,
$\|\hat\mu_{\pi}^{(n,-k^1)}-\mu_{\pi}^o\|_{\mathcal{B}(a;h),2} \|\hat\delta^{(n,-k^1)}-\delta^o\|_{\mathcal{B}(a;h),2}=o_p(c(n))$.
\end{enumerate}
\end{assumption}
We provide a brief overview of the regularity conditions in Assumption (ref). Conditions (ref)--(ref) are standard in the kernel regression literature tsybakov2009introduction. Condition (ref) encodes the classical bias-variance trade-off: $h \to 0$ ensures asymptotic unbiasedness, while $n h \to \infty$ ensures vanishing variance of order $1/(n h)$. Condition (ref) imposes smoothness on $\theta(a)$, controlling the estimation bias.
Condition (ref) controls the variance of $\hat{\theta}_{\pi,h}^{(n)}(a)$, which is a relatively mild requirement.
Condition (ref) requires that the sample size used for training $\hat P_O^{(n,-k^2)}$ grows sufficiently fast to ensure consistency.
Condition (ref) imposes stability and regularity on the nuisance functions in $\hat{\alpha}_{\pi}^{(n,-k^1)}$.
Finally, Condition (ref) specifies the convergence rate of $\hat{\alpha}_{\pi}^{(n,-k^1)}$ towards $\alpha_{\pi}^o$.
Next, we clarify the convergence rates of $\hat{\theta}_{\pi,h}^{(n)}(a)$ as follows.
\begin{theorem}[Convergence rate]
Under Assumptions (ref)--(ref), suppose that $\pi(Z,L)$ is a URWF for an interval $\mathcal{N}$ satisfying Assumption (ref).
Then, if $Z$ is an AIV for all $A=a \in \mathcal{N}$, the estimator $\hat{\theta}_{\pi,h}^{(n)}(a)$ obtained from Algorithm (ref) satisfies for all $a\in\mathcal{N}$,
\begin{align}
\hat{\theta}_{\pi,h}^{(n)}(a) - \theta(a) = O_p\Big(\frac{1}{\sqrt{nh}}\Big) + O_p(h^2) + o_p(c(n)).
\end{align}
\end{theorem}
We explain each term in Equation (ref) as follows.
The first term corresponds to the variance term in LLKR; the second term is the bias term in LLKR; the third term represents the mixed bias caused by plugging in $\hat\alpha_{\pi}^{(n,-k^1)}$ into the estimation.
Intuitively, by taking $h = n^{-1/5}$ and $c(n) = O(n^{-2/5})$, one obtains
$|\hat{\theta}_{\pi,h}^{(n)}(a) - \theta(a)| = O_p(n^{-2/5}).$
In fact, this convergence rate achieves the oracle minimax lower bound for LLKR.
In particular, a sufficient condition for $c(n)=O(n^{-2/5})$ is that $\|\hat\alpha_{\pi}^{(n,-k^1)}-\alpha_{\pi}^o\|_{\mathcal{B}(a;h),2}=o_p(n^{-1/5})$, which is a slightly weaker condition than the converging rate $o_p(n^{-1/4})$ in DML literature Chernozhukov2018.
\begin{theorem}[Asymptotic normality of $\hat{\theta}_{\pi,h}^{(n)}(a)$]
Under the conditions of Theorem (ref), if $c(n) = O(1/\sqrt{nh})$, then for any $a \in \mathcal{N}$ that satisfies Assumption (ref), we have
\begin{equation}
\begin{aligned}
& \sqrt{nh} \left\{\hat{\theta}_{\pi,h}^{(n)}(a) - \theta(a) - \mathrm{bias}(a) \right\}
\overset{d}{\longrightarrow} N(0, \sigma_{\pi,\theta}^2(a)), \quad \text{where} \\
& \sigma_{\pi,\theta}^2(a) := \frac{\int K(s)^2 \, \mathrm{d}s}{p_A(a)} \, \mathrm{Var}\left[\varphi_{\pi}(O;\alpha_{\pi}^o, P_L) \mid A=a\right].
\end{aligned}
\end{equation}
In Equation (ref), the bias term $\mathrm{bias}(a)$, defined in Appendix (ref), satisfies
\begin{align*}
\mathrm{bias}(a) = \frac{h^2}{2} \theta”(a) \int K(s) s^2 \, \mathrm{d}s + o_p(h^2).
\end{align*}
\end{theorem}
This theorem provides a valid variance representation $\sigma_{\pi,\theta}^2(a)$ for $\hat{\theta}_{\pi,h}^{(n)}(a)$.
Specifically, when there exist consistent estimators for $p_A(a)$ and $\mathbb{E}[\varphi_{\pi}(O;\alpha_{\pi}^o, P_L)^2 | A=a]$, one can construct
a consistent variance estimator $\hat{\sigma}_{\pi,\theta}^{(n)}(a)$ for $\sigma_{\pi,\theta}(a)$.
If $h = o(n^{-1/5})$ and $h^{-1} = o(n)$, the bias term in Equation (ref) can be neglected. In this setting, a $95\%$ confidence interval for $\theta(a)$ can be formulated as $\hat{\theta}_{\pi,h}^{(n)}(a) \pm 1.96 \, \hat{\sigma}_{\pi,\theta}^{(n)}(a)/\sqrt{nh}.$
\textcolor{black}{Unfortunately, this guarantee does not hold when the bandwidth is selected adaptively via LOOCV. Relatedly, takatsu2023debiasedinferencecovariateadjustedregression propose a debiased inference approach to mitigate the $\mathrm{bias}(a)$, which is beyond the scope of this article.}
\textcolor{black}{
Next, we give a theoretically analogous analysis for the estimator $\hat \zeta_{\pi, h}^{(n)}(a)$. This will guarantee the asymptotic validity of the p-value $p_{a, \pi}$ established in Algorithm (ref).
\begin{theorem}[Asymptotic normality of $\hat \zeta_{\pi, h}^{(n)}(a)$]
Under Assumptions (ref)(a)--(c), assume that $p_{A|Z,L}(a| z,l)$ is twice continuously differentiable at $a$ for any $(z,l)\in\mathcal{Z}\times\mathcal{L}$.
Moreover, assume that
$\mathbb{E}[\|\hat{\gamma}_{h,a}^{(n,-k)}-\gamma_{h,a}^o\|_2^2]$,
$\mathbb{E}[\|\hat{\rho}_{\pi}^{(n,-k)}-\rho_{\pi}^o\|_2^2] = o(1)$, and
$\|\hat{\gamma}_{h,a}^{(n,-k)}-\gamma_{h,a}^o\|_2 \, \|\hat{\rho}_{\pi}^{(n,-k)}-\rho_{\pi}^o\|_2 = o_p(1/\sqrt{nh})$.
Then, we have
$\hat \zeta_{\pi, h}^{(n)}(a) = \zeta_{\pi}^o(a) + O_p(1/\sqrt{nh} + h^2).$
Define $\mathrm{bias}_{\zeta}(a) := \frac{h^2}{2} \, \mathbb{E}\!\big[p''_{A|Z,L}(a| Z,L) \, (\pi(Z,L)-\rho_{\pi}^o(L))\big] \int K(x) x^2 \, \mathrm dx.$ If $h = O(n^{-1/5})$, then
\begin{align*}
&\sqrt{nh} \big\{\hat \zeta_{\pi, h}^{(n)}(a) - \zeta_{\pi}^o(a) - \mathrm{bias}_{\zeta}(a)\big\} \overset{d}{\rightarrow} N(0,\sigma_{\pi,\zeta}^o(a)^2),
\quad\text{where}\\
&\sigma_{\pi,\zeta}^o(a)^2 := p_A(a)\int K(x)^2\mathrm dx \mathrm{Var}[\pi(Z,L)|A=a].
\end{align*}
\end{theorem}
In Theorem (ref), if $h=o_p(n^{-1/5})$, the bias term $\mathrm{bias}_{\zeta}(a)$ is negligible. Therefore, in Algorithm (ref) we set $h=n^{-1/4}\hat{\sigma}_A$ for computational simplicity. We estimate $\sigma_{\pi,\zeta}^o(a)$ using the sample standard deviation of the scores $\{S_i\}_{i=1}^n$ in Algorithm (ref). Consequently, $p_{a,\pi}$ can be treated as a valid p-value for testing $H_0:\zeta_{\pi}^o(a)= 0$ versus $H_1:\zeta_{\pi}^o(a)\neq 0$, and $\tilde{p}_{a,\pi}$ also serves as a valid p-value for detecting violations of the RWF assumption when $L$ is discrete and $D = |\mathcal{L}|$ in Algorithm (ref).
}
\section{Simulations}
We conduct simulation studies to investigate the finite-sample performance of the proposed methods.
Specifically, we simulate independent and identically distributed random variables
$\varepsilon_L, \varepsilon_U, \varepsilon_Z\sim \mathrm{Unif}(0,1)$, $\varepsilon_{AZ},\varepsilon_{AU}\sim N(0,1)$, and $\varepsilon_A \sim \mathrm{Ber}(0.7)$.
We then generate the observed and latent variables according to
$L = \varepsilon_L - 0.5$,
$U = 3(\varepsilon_U - 0.5)$,
$Z = -0.5L + 3(\varepsilon_Z - 0.5)$,
$\mathrm{logit}(p_z) = 2Z+\varepsilon_{AZ}$,
$\mathrm{logit}(p_u) = -2U+\varepsilon_{AU}$,
$A = 2\varepsilon_A p_z+ 2(1-\varepsilon_A)p_u - 1$,
$Y = A + U - 0.5 L.$
By construction, $\mathcal{A} = [-1,1]$, and $Z$ serves as an AIV
for each treatment level $a \in \mathcal{A}$. We focus on a target subinterval
$\mathcal{A}_c := [-0.75, 0.75]$.
\textcolor{black}{As described in Section (ref), we construct a cover of $\mathcal{A}_c$ by defining $\mathcal{N}_m:= [0.5m - 1.25, 0.5m - 0.75]$ for $m = 1,2,3$. For each $m$, we generate 10{,}000 additional samples to construct $\pi_{[m]} = \hat{\pi}_{a_{m}}(Z,L)$ at $a_{m} = 0.5m - 1$.
We present an analysis of RWF testing procedure (Algorithm (ref)) in Appendix (ref), demonstrating that $\pi_{[m]}$ provides a suitable choice of URWF for $\mathcal{N}_m$, for $m = 1, 2, 3$.}
\begin{figure}[t]
\begin{subfigure}{0.48\textwidth}
\caption{$\mathcal{N}_1 = [-0.75,-0.25],\,n=5000$}
\end{subfigure}
\begin{subfigure}{0.48\textwidth}
\caption{$\mathcal{N}_1 = [-0.75,-0.25],\,n=10000$}
\end{subfigure}
\begin{subfigure}{0.48\textwidth}
\caption{$\mathcal{N}_2 = [-0.25,0.25],\,n=5000$}
\end{subfigure}
\begin{subfigure}{0.48\textwidth}
\caption{$\mathcal{N}_2 = [-0.25,0.25],\,n=10000$}
\end{subfigure}
\begin{subfigure}{0.48\textwidth}
\caption{$\mathcal{N}_3 = [0.25,0.75],\,n=5000$}
\end{subfigure}
\begin{subfigure}{0.48\textwidth}
\caption{$\mathcal{N}_3 = [0.25,0.75],\,n=10000$}
\end{subfigure}
\caption{Empirical mean and pointwise $95\%$ \textcolor{black}{uncertainty bands} of estimated ADRF for $n=5000,\,10000$ with 400 replications.}
\end{figure}
To demonstrate the effectiveness of our proposed IV-based framework, we compare our estimators with those introduced in kennedy2017non.
For completeness, the formal definitions of the inverse probability weighting (IPW) and outcome regression (OR) estimators in the IV setting and their counterparts under the NUC framework are provided in Appendix (ref).
We set the sample size to be 5,000 or 10,000. The bandwidth is adaptively selected using the localized LOOCV criterion.
\textcolor{black}{We set the crossing folds \(K=5\) folds for cross-fitting, a common choice that balances computational cost and the accuracy of nuisance parameter estimation Chernozhukov2018.}
We use the Epanechnikov kernel $K(a)=0.75\max\{1 - a^2,0\}$. We train the nuisance functions using the splines method \texttt{gam} provided in \texttt{R} package \texttt{mgcv} and the kernel density estimators \texttt{npcdens}, \texttt{npcdensbw} in \texttt{R} package \texttt{np}.
For graphical illustration, Figure (ref) shows plots of curves estimated from the six methods, displaying the average estimates and the pointwise $95\%$ uncertainty bands of the ADRFs obtained from 400 repetitions.
\textcolor{black}{Notably, we report 95% uncertainty bands rather than confidence bands, as adaptively selecting the weighting function via the LOOCV procedure may introduce non-negligible bias.}
When adopting the NUC framework (left panel of Figure (ref)(a)--(f)), noticeable deviations from the true curve are observed. In contrast, the IV framework (right panel of Figure (ref)(a)--(f)) effectively reduces the estimation bias, albeit with a slight increase in variance compared to NUC framework.
See Appendix (ref) for further discussion.
\section{Real data application}
We use the dataset provided in the supplementary material of han2024optimal for empirical illustration.
It combines the Job Training Partnership Act (JTPA) dataset with data from the US
Census and the National Center for Education Statistics (NCES).
This dataset contains the number of high schools per square mile as an IV, the years of education as the treatment, pre-program earnings as the outcome of interest, measured in dollars and representing total annual incomes, and sex as a confounder between $A$ and $Y$.
We are interested in estimating the effect of years of education on the pre-program earnings.
\textcolor{black}{
Intuitively, local variation in the number of high schools per square mile affects educational level and thus the treatment, and it is unlikely to directly influence pre-program earnings. Hence, it constitutes a reasonable IV in this setting.}
The dataset consists of 9,223 individuals. Since there are only approximately 500 individuals with fewer than 8 or more than 15 years of education, we grouped those with less than 8 years into the “8 years” category and those with more than 15 years into the “15 years” category.
Notably, although years of education are observed discretely, they can be treated as realizations of a continuous treatment variable defined on the real line. See the discussion in Section (ref) and Appendix (ref) for further discussion.
\textcolor{black}{For illustration purposes, we conduct our analyses under two specifications: one that sets $L = \emptyset$ and one that includes sex as $L$, treating $Z$ as a fully exogenous IV in both cases. In particular, when $L = \emptyset$, sex can be regarded as part of the unmeasured confounder $U$. Consequently, regardless of whether we condition on sex, Assumptions (ref)--(ref) are expected to hold, and $Z$ serves as an AIV for $A$. For illustration, we compare these two settings and report results under both specifications.}
Before presenting the estimated ADRFs, we first estimate $\hat{\pi}_{a}(Z,L)$ for $a$ ranging from 8 to 15, where the density functions are fitted using the \texttt{R} package \texttt{mgcv}.
We conduct the RWF testing procedure in Algorithm (ref) for all eight prespecified weighting functions. Figures (ref)(a) and (b) present the logarithm of the p-values across years of education. Figure (ref) (a) adjusts for sex as a confounder (with $D=2$ in Algorithm (ref)), while Figure (ref) (b) does not.
At each $A = a$, smaller p-values indicate that the corresponding weighting function is more likely to be the RWF. \textcolor{black}{Overall, $\hat{\pi}_9$ is a good candidate for $A = 8,9,10$, $\hat{\pi}_{12}$ for $A = 12$, and $\hat{\pi}_{15}$ for $A = 13,14,15$. Accordingly, we select $\hat{\pi}_9$, $\hat{\pi}_{12}$, and $\hat{\pi}_{15}$ as the RWFs for the corresponding values of $A$.
For the 1,167 subjects with $A = 11$, the estimated p-values are all close to or exceed $0.05$. To mitigate the resulting instability, we first compute the AIPW scores via Algorithm (ref) and then exclude the subjects with $A=11$.} In addition, we set $K=5$ in Algorithm (ref), in which the nuisance functions are fitted using the \texttt{R} package \texttt{mgcv}.
\begin{figure}[t]
\begin{subfigure}{0.32\textwidth}
\caption{Adjust covariate: sex.}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{Adjust covariate: empty.}
\end{subfigure}
\begin{subfigure}{0.32\textwidth}
\caption{Final results.}
\end{subfigure}
\caption{The first two plots show the RWF testing results for eight prespecified weighting functions $\{\hat{\pi}_a\}_{a=8}^{15}$, with the dashed line indicating the significance threshold at $p = 0.05$. The third plot shows the estimated curve (solid line) and the 95% uncertainty band based on 400 bootstrap replications.}
\end{figure}
\textcolor{black}{To estimate the ADRF globally, we merge the AIPW scores corresponding to the three selected RWFs. Specifically, for each treatment level $A$, we use the AIPW score estimated with the corresponding RWF for that level (e.g., at $A=8,9,10$, we use the AIPW score of $\hat{\pi}_9$). Then, we fit a smoothing spline regression with the smoothing parameter adaptively selected via the generalized cross-validation criterion.
The use of spline methods is theoretically justified, as their properties underpin the error bounds of the ERM framework discussed in Appendix (ref).}
\textcolor{black}{We conduct 400 bootstrap replications to calculate the standard deviation and construct 95% uncertainty bands for all treatment levels. To assess the potential impact of unmeasured confounders, we also report the corresponding estimates under the NUC framework, where sex is included as a confounder.
The final results are shown in Figure (ref)(c).
First, the results consistently highlight the positive effect of education, and those obtained under the IV framework further amplify this pattern for $A\leq 12$, while the NUC method estimates are much more stable than the IV method, providing a conservative depiction of the overall trend.
Second, the IV method shows that income slightly decreases when $A \geq 12$, indicating that further increases in education beyond a certain level may no longer have a positive effect on wages. In contrast, this pattern is not reflected in the results obtained using NUC.
Third, the estimated pre-program earnings are less stable when $L=\{\text{sex}\}$ than when $L$ is empty, although the overall pattern remains similar. This reflects the increased variability introduced by adjusting for additional confounders.}
\section{Discussion}
In this article, we propose a novel IV framework for the identification of ADRFs.
We demonstrate that it is not only feasible but also necessary to employ the idea of a finite open cover, whereby any compact subset of the treatment domain can be covered by a collection of open sets. For each of these open sets, we can construct a corresponding URWF, enabling the identification of the ADRFs within that region. \textcolor{black}{We propose an RWF testing procedure and provide practical guidance for constructing open covers, each of which admits a URWF.}
We propose the corresponding AIPW score functions and implement a general cross-fitting procedure within the DML framework to compute them.
Moreover, we establish the convergence rate of the proposed estimator with kernel regression or empirical risk minimization.
Several directions for future research remain to be explored.
\textcolor{black}{
First, while we have established a hypothesis test for a necessary condition of the RWF assumption with theoretical guarantees, we have yet to develop theoretically valid tests for settings with continuous covariates.
Second, developing uniformly valid confidence bands for the estimated ADRF remains an interesting direction for future research. Recent work, such as takatsu2023debiasedinferencecovariateadjustedregression, doss2026doubly,ai2026datadrivenuniforminferencegeneral, may provide useful guidance.
Third, it would be of great interest to develop constancy tests under our setting following recent developments of westling2022nonparametric,doss2024nonparametric.
Fourth, following branson2023causal, future work could focus on developing methods that are robust to violations of positivity and IV relevance conditions under the continuous treatments setting.}
Fifth, our framework also sheds light on developing personalized dose-finding strategies Chen01102016, kallus2018policyevaluationoptimizationcontinuous, ai2024datadrivenpolicylearningcontinuous within the proposed IV framework.
Finally, designing a hypothesis test to detect potential violations of the AIV condition represents another promising avenue for future research.
\putbib[bibfile_main]
bibunit