EconBase
← Back to paper

Orthogonal Integrated Conditional Moment Tests for Treatment Effect Heterogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,435 characters · 0 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Orthogonal Integrated Conditional Moment Tests for Treatment Effect Heterogeneity

bibunit\begin{spacing}{1.2} \begin{abstract} We propose a nonparametric integrated conditional moment (ICM) test for treatment effect heterogeneity across subpopulations defined by a given covariate subvector. Under unconfoundedness, the null is recast as a conditional moment restriction based on a Neyman-orthogonal score, which reduces the first-order sensitivity of the empirical process to nuisance parameter estimation. The test statistics are constructed as continuous functionals of a marked empirical process. We establish a uniform feasible-to-oracle approximation and derive the asymptotic properties of these test statistics under the null and fixed alternatives. We further show that the test has nontrivial power against local alternatives converging to the null at the \(n^{-1/2}\) rate, and develop an easy-to-implement multiplier bootstrap for feasible inference. We also develop extensions to tests of parametric CATE specifications and to settings with endogenous treatment and a binary instrument. Finally, we apply the proposed testing approach to study whether the effect of maternal smoking during pregnancy on infant birth weight varies with maternal age. \noindentKeywords: Doubly Robust; Empirical Processes; Integrated Conditional Moments; Neyman Orthogonality; Treatment Effect Heterogeneity. \noindentJEL Classification Number: C12; C14; C15. \end{abstract} \end{spacing} \thispagestyle{empty} \section{Introduction} In program evaluation, it has long been recognized that treatment effects may vary systematically across individuals and subpopulations with different observed characteristics; see, for example, heckman1985alternative, heckman1997making, and imbens2009recent. Such heterogeneity matters because an intervention that is beneficial on average may be less effective, ineffective, or even harmful for particular groups. Average treatment effects remain useful summaries, but they do not reveal how effects vary with economically meaningful characteristics such as age, income, education, or baseline risk. Understanding these differences is therefore important for interpreting empirical findings and assessing their external validity, and may also inform subsequent treatment-assignment decisions; see, for example, manski2004statistical. A natural way to study such heterogeneity is through the conditional average treatment effect (CATE), which characterizes how the average effect of a treatment varies across subpopulations defined by observed characteristics. In observational studies, however, the covariates needed for causal identification need not coincide with those along which treatment-effect heterogeneity is substantively of interest. Let \(X\) denote the full vector of pretreatment covariates used to adjust for treatment selection, and let \(Z\subseteq X\) denote the prespecified covariates of interest that define the relevant subpopulations. Thus, \(X\) is chosen to support a credible selection-on-observables strategy, whereas \(Z\) reflects the substantive heterogeneity question under study. This distinction is common in empirical work: researchers may require a relatively rich set of controls to account for treatment selection, while being primarily interested in how the treatment effect varies with a particular characteristic such as age, education, income, prior earnings, or baseline risk. In our empirical application, for example, maternal and pregnancy characteristics are used for selection adjustment, while the heterogeneity of interest is with respect to mother's age. In this setting, a growing literature has developed methods for estimating and conducting inference on the CATE indexed by covariates of interest. abrevaya2015estimating propose inverse-probability-weighted estimators and pointwise inference when the covariates defining the CATE form a subset of those used for selection adjustment. lee2017doubly develop doubly robust estimators and uniform confidence bands, while fan2022estimation accommodate high-dimensional first-step estimation through orthogonal scores and provide both pointwise and uniform inference. These methods allow researchers to estimate and visualize how treatment effects vary with \(Z\), together with the associated statistical uncertainty. A nonflat estimated CATE profile, however, need not by itself imply statistically significant heterogeneity. This observation motivates a direct specification question: whether the CATE is constant over \(Z\). Beyond constancy, one may further ask whether systematic variation in the CATE can be adequately summarized by a parsimonious functional form. In this paper, we develop a nonparametric test for CATE heterogeneity with respect to prespecified covariates of interest. We first rewrite the homogeneity null as a conditional moment restriction. Following the integrated conditional moment (ICM) approach pioneered by bierens1982consistent, we transform this restriction into a continuum of unconditional moment conditions indexed by the covariates of interest. The sample analog of these moments defines a marked empirical process, and we construct Kolmogorov--Smirnov and Cramér--von Mises test statistics as continuous functionals of this process. Importantly, the proposed procedure works directly with the moment restrictions and does not require nonparametrically estimating or smoothing the CATE function itself. The main econometric difficulty arises because the conditional moment restriction depends on unknown nuisance functions, including the propensity score and the two conditional outcome regressions. We construct the test using an augmented inverse-probability-weighted, or doubly robust, score robins1994estimation,hahn1998role,bang2005doubly and show that the resulting population moments are Neyman-orthogonal with respect to all three nuisance functions. They are therefore locally insensitive to first-order perturbations of these functions. We estimate the propensity score using the series logit estimator of hirano2003efficient and the outcome regressions by sieve least squares. Under primitive smoothness and series-rate conditions, we establish a uniform feasible-to-oracle approximation: the feasible process based on the estimated nuisance functions is uniformly asymptotically equivalent to an oracle process constructed from the population doubly robust score. Building on this approximation, we derive the weak limit of the test process under the null, establish consistency against fixed alternatives, and obtain nontrivial power against alternatives approaching homogeneity at the \(n^{-1/2}\) rate. The same oracle representation also supports a simple and computationally efficient multiplier bootstrap. Because the effects of first-step estimation are uniformly asymptotically negligible, the bootstrap can be implemented by resampling the estimated score residuals while holding the nuisance estimates fixed, thereby avoiding repeated estimation of the propensity score and outcome regressions across bootstrap draws. To accommodate related specification and identification questions, we consider two extensions. First, we develop tests of whether the CATE belongs to a prespecified parametric family, which may include, for example, linear or quadratic forms. Second, building on the local average treatment effect framework of imbens1994identification, we adapt the proposed approach to test homogeneity of the conditional LATE when treatment is endogenous and a binary instrument is available. Monte Carlo experiments show that the tests have rejection frequencies close to nominal levels under the null and increasing power against a range of heterogeneous alternatives as the sample size grows. We then apply the method to North Carolina vital-statistics data to examine whether the effect of maternal smoking on infant birth weight varies with mother's age. The homogeneity null is not rejected for Black mothers but is strongly rejected for White mothers. For White mothers, a linear CATE in age is not rejected at the $5\%$ level, suggesting that the detected heterogeneity can be summarized approximately by a linear age profile. The proposed procedures are related to a growing literature on formal tests of treatment-effect restrictions. A natural starting point is crump2008nonparametric, an early and influential contribution that develops direct nonparametric tests of zero and homogeneous conditional average treatment effects. Their hypotheses concern the CATE conditional on the full covariate vector \(X\), and their statistics are constructed from series estimates of the treated and untreated conditional outcome regressions. Our procedure differs in both the conditioning structure and the construction of the test statistic. We use the full vector \(X\) for selection adjustment while studying heterogeneity with respect to prespecified covariates \(Z\subseteq X\), and construct a marked empirical process from a doubly robust score rather than directly from estimated outcome regression functions. Also working with restrictions conditional on the full covariate vector, chang2015nonparametric develop kernel-based tests for a broad class of conditional equality and inequality restrictions, while delgado2013conditional study bootstrap tests of conditional stochastic dominance. When the heterogeneity of interest is indexed by selected covariates \(Z\subseteq X\), hsu2017consistent tests whether the CATE is nonnegative for every value of \(Z\), formulating the null as a conditional moment inequality and drawing on andrews2013inference to conduct inference through a collection of unconditional moment inequalities. sant2021nonparametric develops tests of zero and homogeneous treatment effects for possibly right-censored duration outcomes, while cai2025nonparametric study homogeneity of partially conditional quantile treatment effects using nonparametric estimators of the corresponding quantile-effect functions. More recently, dukes2026nonparametric propose nonparametric tests of quantitative and qualitative CATE heterogeneity, motivated by applications in precision medicine and by the policy consequences of individualized treatment rules. Their framework evaluates heterogeneity through contrasts over prespecified classes of treatment rules, with the choice of rule class determining the alternatives to which the tests are most sensitive. We instead formulate constancy and more general prespecified functional restrictions on the mean CATE as conditional moment restrictions and use the ICM principle to construct Kolmogorov--Smirnov and Cramér--von Mises tests that are consistent against fixed alternatives. The remainder of the paper is organized as follows. Section (ref) formulates the testing problem, introduces the orthogonal ICM process, and defines the KS and CvM statistics. Section (ref) establishes the feasible-to-oracle approximation and the null, fixed-alternative, and local-alternative limits. Section (ref) develops the multiplier bootstrap. Section (ref) presents the parametric-form and conditional-LATE extensions. Sections (ref) and (ref) report the Monte Carlo study and empirical illustration, respectively. Proofs and additional simulation results are collected in the Online Appendix. Notation. For an index set \(\mathcal A\), let \(\ell^\infty(\mathcal A)\) denote the Banach space of bounded real-valued functions on \(\mathcal A\), equipped with the supremum norm \(\|f\|_\infty:=\sup_{a\in\mathcal A}|f(a)|\). We write \(X_n\rightsquigarrow X\) for weak convergence of random elements in \(\ell^\infty(\mathcal A)\) in the Hoffmann--J{\o}rgensen sense; see, for example, Definition 1.3.3 of van1996weak. Convergence in probability and in distribution are denoted by \(\xrightarrow{p}\) and \(\xrightarrow{d}\), respectively. For a deterministic sequence \(a_n>0\) and random elements \(X_n\) in a normed space, \(X_n=o_p(a_n)\) means that \(\|X_n\|/a_n\xrightarrow{p}0\), whereas \(X_n=O_p(a_n)\) means that \(\|X_n\|/a_n\) is bounded in probability. \section{Testing Framework} We first introduce the framework considered in this paper. We work with the potential-outcome setup of rubin1974estimating. Let \(D\in\{0,1\}\) denote the treatment indicator, where \(D=1\) if an individual receives the treatment and \(D=0\) otherwise. Let \(Y(1)\) and \(Y(0)\) denote the potential outcomes under treatment and control, respectively. The observed outcome is \(Y=DY(1)+(1-D)Y(0)\). Let \(X\) be a vector of pretreatment covariates. Throughout the paper, we observe an i.i.d. sample \(W_i:=(Y_i,D_i,X_i)\), \(i=1,\ldots,n\). Let \(Z\) denote a subvector of \(X\) containing the covariates of interest. In many applications, \(Z\) is low-dimensional even when \(X\) is relatively rich. Define the conditional average treatment effect given \(Z=z\) as $\tau(z):=\mathbb{E}[Y(1)-Y(0)\mid Z=z]$. In this section, we are interested in testing whether the treatment effect is heterogeneous with respect to \(Z\). Formally, we test \begin{equation} \mathbb H_0:\ \exists \tau\in\mathcal T such that \tau(z)=\tau for a.e. z\in\mathcal Z, \quad versus\quad \mathbb H_1:\ \mathbb{P}\{\tau(Z)\neq \tau\}>0, \forall \tau\in\mathbb R. \end{equation} To identify \(\tau(z)\) from the observed data, we impose the following standard conditions. \begin{assumption}[Unconfoundedness] \((Y(0),Y(1))\perp D\mid X\). \end{assumption} \begin{assumption}[Overlap] Let \(p(x):=\mathbb{P}(D=1\mid X=x)\) denote the propensity score. There exists a constant \(\epsilon>0\) such that $\epsilon\le p(X)\le 1-\epsilon$ a.s. \end{assumption} Assumption (ref) is the selection-on-observables restriction introduced by rosenbaum1983central: conditional on the vector of pre-treatment covariates \(X\), treatment assignment is independent of the potential outcomes. Assumption (ref) rules out regions of the covariate support in which treatment status is deterministic. Together, these assumptions allow us to express the CATE in terms of observed-data nuisance functions. Let \(\mu_d(x):=\mathbb{E}[Y\mid D=d,X=x]\), \(d=0,1\). To reduce the first-order impact of nuisance estimation on the subsequent testing procedure, we work with the doubly robust score representation (also known as the augmented inverse-probability-weighted form); see, e.g., robins1994estimation, hahn1998role, and bang2005doubly. Its Neyman-orthogonality property will be formalized in Proposition (ref) below. For a generic nuisance value \(\eta=(q,m_1,m_0)\), define \[ \psi(W;\eta) := m_1(X)-m_0(X) +\frac{D\{Y-m_1(X)\}}{q(X)} -\frac{(1-D)\{Y-m_0(X)\}}{1-q(X)}. \] Let \(\eta_0=(p,\mu_1,\mu_0)\) denote the true nuisance value. Under Assumptions (ref) and (ref), the CATE is identified by \[ \tau(z) = \mathbb{E}[\psi(W;\eta_0)\mid Z=z]. \] Consequently, the average treatment effect is identified as \[ \tau_0:=\mathbb{E}[Y(1)-Y(0)]=\mathbb{E}[\tau(Z)]=\mathbb{E}[\psi(W;\eta_0)]. \] Under the null in (ref), the constant must equal \(\tau_0\) by the law of iterated expectations. Therefore, the testing problem in (ref) can be equivalently represented as the following conditional moment restriction: \begin{equation} \mathbb H_0:\mathbb{E}[\psi(W;\eta_0)-\tau_0\mid Z]=0 \quad \text{a.s.},\quad \text{versus}\quad \mathbb H_1:\mathbb{P}\!\left( \mathbb{E}[\psi(W;\eta_0)-\tau_0\mid Z]=0 \right)<1 . \end{equation} Following bierens1982consistent, bierens1997asymptotic, and stute1997nonparametric, we adopt an integrated conditional moment (ICM) approach. Specifically, we test the conditional moment restriction in (ref) through the family of unconditional moments \begin{equation} \mathbb{E}\!\left[\{\psi(W;\eta_0)-\tau_0\}w(Z,z)\right]=0, \qquad z\in\Pi, \end{equation} where \(\Pi\) is an indexing set and \(w(\cdot,z)\) is a measurable weighting function indexed by \(z\). Under suitable conditions on the class \(\{w(\cdot,z):z\in\Pi\}\), the family of unconditional moments in (ref) is equivalent to the conditional moment restriction in (ref); see bierens1997asymptotic and escanciano2006goodness. In this paper, following stute1997nonparametric, we use the lower-orthant indicator weight \(w(Z,z)=\mathbf{1}\{Z\le z\}\) with \(\Pi=\mathcal Z\).\footnote{If \(Z\) is vector-valued, the inequality is understood componentwise, i.e., \(\mathbf{1}\{Z\le z\}=\mathbf{1}\{Z_1\le z_1,\ldots,Z_{d_z}\le z_{d_z}\}\).} We now explain why we use the doubly robust score in the ICM moments. The key motivation is Neyman orthogonality. Moment conditions with this property have reduced first-order sensitivity to nuisance perturbations. Formally, Neyman orthogonality means that the moment condition has zero G\^ateaux derivative with respect to the nuisance functions. In the present setting, define, for each \(z\in\mathcal Z\), $M(z;\eta,\tau) := \mathbb{E}\!\left[ \{\psi(W;\eta)-\tau\}\mathbf{1}\{Z\le z\} \right].$ The result is presented in the following proposition. \begin{proposition}[Neyman orthogonality of the ICM moments] Suppose Assumptions (ref) and (ref) hold. Let \(\eta=(q,m_1,m_0)\) be an arbitrary nuisance value such that \(q\) is bounded away from zero and one. For \(r\in[0,1]\), define the path \(\eta^r:=\eta_0+r(\eta-\eta_0)\), where \(\eta_0=(p,\mu_1,\mu_0)\). Suppose that, for each \(z\in\mathcal Z\), \(\sup_{r\in[0,1]}\left|\partial[\{\psi(W;\eta^r)-\tau\}\mathbf{1}\{Z\le z\}]/\partial r\right|\) is integrable. Then \[ \left. \frac{\partial}{\partial r} M(z;\eta^r,\tau) \right|_{r=0} =0, \qquad z\in\mathcal Z, \] where \(\tau\) need not equal its true value. Moreover, for any fixed nuisance value \(\eta\) and any scalar direction \(h_\tau\in\mathbb R\), \[ \left. \frac{d}{dt} M(z;\eta,\tau_0+t h_\tau) \right|_{t=0} = -h_\tau F_Z(z), \qquad z\in\mathcal Z. \] \end{proposition} The first part of Proposition (ref) gives Neyman orthogonality with respect to the infinite-dimensional nuisance functions. Thus, the ICM moments based on the doubly robust score are locally insensitive to first-order perturbations of \(p\), \(\mu_1\), and \(\mu_0\). The second part shows that estimation of the scalar parameter \(\tau_0\) contributes the projection term \(F_Z(z)\). This term will appear in the oracle asymptotic representation, although the test itself is constructed using the uncentered lower-orthant indicator. We now construct the feasible test process. Let \(\hat\eta=(\hat p,\hat\mu_1,\hat\mu_0)\) be estimators of \(\eta_0=(p,\mu_1,\mu_0)\). Define the feasible doubly robust score \begin{equation} \hat\psi_i := \psi(W_i;\hat\eta) = \hat\mu_1(X_i)-\hat\mu_0(X_i) +\frac{D_i\{Y_i-\hat\mu_1(X_i)\}}{\hat p(X_i)} -\frac{(1-D_i)\{Y_i-\hat\mu_0(X_i)\}}{1-\hat p(X_i)}. \end{equation} The average treatment effect is estimated by \begin{equation} \hat\tau:=\frac1n\sum_{i=1}^n \hat\psi_i. \end{equation} Multiplying the sample analog of (ref) by \(\sqrt n\), we obtain the feasible ICM process \begin{equation} \widehat R_n(z) := \frac1{\sqrt n}\sum_{i=1}^n (\hat\psi_i-\hat\tau)\mathbf{1}\{Z_i\le z\}, \qquad z\in\mathcal Z. \end{equation} Because \(n^{-1}\sum_{i=1}^n(\hat\psi_i-\hat\tau)=0\), (ref) is algebraically equivalent to the empirical-centered form \begin{equation} \widehat R_n(z) =\frac1{\sqrt n}\sum_{i=1}^n \hat\psi_i \{\mathbf{1}\{Z_i\le z\}-\widehat F_{Z,n}(z)\}= \frac1{\sqrt n}\sum_{i=1}^n (\hat\psi_i-\hat\tau) \{\mathbf{1}\{Z_i\le z\}-\widehat F_{Z,n}(z)\}, \end{equation} where \(\widehat F_{Z,n}(z):=n^{-1}\sum_{i=1}^n\mathbf{1}\{Z_i\le z\}\). Hence the process is implemented with the intuitive uncentered indicator, while the centering induced by \(\hat\tau\) is built in automatically. By the equivalence between the conditional moment restriction in (ref) and the ICM moments in (ref), departures from \(\mathbb H_0\) are reflected in systematic deviations of \(\widehat R_n\) from zero. We therefore use the Kolmogorov--Smirnov and Cram\'er--von Mises functionals \[ KS_n:=\sup_{z\in\mathcal Z}|\widehat R_n(z)|, \qquad CvM_n:=\int_{\mathcal Z}\widehat R_n(z)^2\,d\widehat F_{Z,n}(z) = \frac1n\sum_{i=1}^n \widehat R_n(Z_i)^2. \] The null hypothesis is rejected for sufficiently large values of either statistic. The construction above is stated for generic first-step estimators. In the following section, we specify the sieve estimators used to construct \(\hat\eta\), substitute them into (ref)--(ref), and establish the feasible-to-oracle approximation that underlies the asymptotic theory. \section{Asymptotic Theory} We now specify the first-step estimators and derive the asymptotic properties of the feasible ICM process. The propensity score is estimated by the series logit estimator of hirano2003efficient. This estimator has been widely used in semiparametric treatment-effect analysis, both for estimation and for testing. It combines flexible nonparametric approximation with the logistic link, thereby keeping the estimated propensity score in the unit interval; see also firpo2007efficient, firpo2016identification, hsu2017consistent, sant2021nonparametric, and cai2025nonparametric. The outcome regressions are estimated by sieve least squares on the same series space. Formally, let \(r\) denote the dimension of \(X\), and let \(\varphi=(\varphi_1,\ldots,\varphi_r)'\in\mathbb Z_+^r\) be a vector of nonnegative integers. Write \(|\varphi|=\sum_{j=1}^r\varphi_j\), and let \(\{\varphi(k)\}_{k=1}^\infty\) be a sequence containing all distinct elements of \(\mathbb Z_+^r\), ordered so that \(|\varphi(k)|\) is nondecreasing in \(k\). For \(x^\varphi=\prod_{j=1}^r x_j^{\varphi_j}\), define $R^K(x):=\bigl(x^{\varphi(1)},\ldots,x^{\varphi(K)}\bigr)'$, where \(K\) denotes the number of terms in the series. Let \(\Lambda(a)=\exp(a)/(1+\exp(a))\). The series logit estimator of the propensity score is $\hat p(x)=\Lambda\{R^K(x)'\hat\pi_K\}$, where \[ \hat\pi_K = \arg\max_{\pi_K} \frac1n\sum_{i=1}^n \left[ D_i\log\Lambda\{R^K(X_i)'\pi_K\} + (1-D_i)\log\{1-\Lambda(R^K(X_i)'\pi_K)\} \right]. \] Given \(\hat p\), we estimate the outcome regressions by sieve least squares. Specifically, \[ \hat\mu_1(x) = R^K(x)' \left(\sum_{i=1}^n R^K(X_i)R^K(X_i)'\right)^{-1} \left(\sum_{i=1}^n \frac{D_iY_i}{\hat p(X_i)}R^K(X_i)\right), \] and \[ \hat\mu_0(x) = R^K(x)' \left(\sum_{i=1}^n R^K(X_i)R^K(X_i)'\right)^{-1} \left(\sum_{i=1}^n \frac{(1-D_i)Y_i}{1-\hat p(X_i)}R^K(X_i)\right). \] These estimators define \(\hat\eta=(\hat p,\hat\mu_1,\hat\mu_0)\), and the feasible score, ATE estimator, and ICM process are then obtained from (ref)--(ref). To establish the asymptotic properties of the feasible ICM process under the sieve estimators above, we impose the following additional regularity conditions. \begin{assumption}[Distribution and outcome smoothness] \quad \begin{itemize} {0pt} {0pt} {0pt} • The support of the \(r\)-dimensional covariate vector \(X\) is a Cartesian product of compact intervals, \(\mathcal X=\prod_{j=1}^{r}[x_{lj},x_{uj}]\). Moreover, \(X\) admits a density \(f\) on \(\mathcal X\) such that \(0<\underline c\le f(x)\le \bar c<\infty\) for all \(x\in\mathcal X\). • For \(d=0,1\), \(\mathbb{E}[|Y(d)|^2]<\infty\), and \(\sup_{x\in\mathcal X}\mathrm{Var}(Y(d)\mid X=x)<\infty\). • The conditional mean functions \(\mu_d(x):=\mathbb{E}[Y\mid D=d,X=x]\), \(d=0,1\), are \(s_\mu\)-times continuously differentiable on \(\mathcal X\), with \(s_\mu>r/2\). \end{itemize} \end{assumption} \begin{assumption}[Propensity-score smoothness] The propensity score \(p(x)\) is \(s\)-times continuously differentiable on \(\mathcal X\), with \(s>5r\). \end{assumption} \begin{assumption}[Rates for the series estimator] The number of terms in the series satisfies \(K=a\cdot n^\nu\) for some \(a>0\) and $\max\left\{ {r}/{(2s-4r)}, {r}{/(s+2s_\mu)} \right\} <\nu<1/6 $. \end{assumption} Similar assumptions have been adopted by hirano2003efficient, crump2008nonparametric, hsu2017consistent, and sant2021nonparametric, among others. Assumption (ref) requires the covariates in \(X\) to be continuous. This restriction is imposed mainly to keep the notation and arguments simple. At the expense of additional notation, the framework can accommodate covariates with both continuous and discrete components. In that case, the same procedure can be applied within cells defined by the discrete covariates, or equivalently by interacting the continuous power-series basis with indicators for the discrete covariate values. Similarly, if the covariate of interest \(Z\) contains discrete components, the lower-orthant instrument class can be augmented by indicators for the corresponding discrete cells. See Remark 2 of hirano2003efficient and Remark 3.2 of hsu2017consistent for related discussions. The smoothness requirements for \(\mu_d(\cdot)\) and \(p(\cdot)\) in Assumptions (ref) and (ref) reflect the doubly robust structure of the feasible process. In hirano2003efficient, the estimator is IPW-based, and the primitive smoothness requirement is concentrated on the propensity score: \(p(\cdot)\) is assumed to be \(s\)-times continuously differentiable with \(s\ge 7r\), while the outcome regressions are only required to be continuously differentiable. By contrast, crump2008nonparametric construct their tests from series estimators of the two outcome regression functions; they require \(\mu_w(\cdot)\) to be \(s\)-times continuously differentiable with \(s/r>25/4\), while the propensity score enters only through overlap. Our feasible ICM process uses both \(\hat p\) and \(\hat\mu_d\) through the doubly robust score. By Proposition (ref), the DR score removes their leading first-order effects in the ICM moments. The remaining plug-in effects are controlled uniformly through the convergence rates of the first-step estimators, which motivates the smoothness conditions \(s>5r\) for \(p(\cdot)\) and \(s_\mu>r/2\) for \(\mu_d(\cdot)\). Finally, the restrictions in Assumptions (ref) and (ref) guarantee the existence of a \(\nu\) satisfying the conditions in Assumption (ref). The next proposition establishes the key feasible-to-oracle approximation: under the maintained regularity conditions, the feasible ICM process based on estimated nuisance functions is uniformly asymptotically equivalent to its oracle counterpart based on the population doubly robust score. \begin{proposition} Suppose Assumptions (ref)-(ref) hold. Then, $$ \sup_{z\in\mathcal Z} \left| \widehat R_n(z) - \frac{1}{\sqrt n}\sum_{i=1}^n \left(\psi(W_i;\eta_0)-\tau_0\right)\bigl(\mathbf{1}\{Z_i\le z\}-F_Z(z)\bigr) \right| =o_p(1). $$ \end{proposition} \begin{remark} Proposition (ref) is the uniform sample-level realization of the population orthogonality in Proposition (ref). The latter shows that the ICM moments based on the doubly robust score are locally insensitive to first-order perturbations of \(p,\mu_1,\mu_0\); the former shows that this orthogonality removes the leading effect of nuisance estimation uniformly over \(z\in\mathcal Z\). Consequently, \(\widehat R_n\) has the same first-order representation as the oracle process based on \(\psi(W;\eta_0)\). The centering term \(F_Z(z)\) has a different source: it reflects the estimation of \(\tau_0\) by \(\hat\tau\). Although \(\widehat R_n\) is defined with the uncentered lower-orthant indicator, (ref) shows that it is equivalently an empirical-centered process; hence the population-centered weight \(\mathbf{1}\{Z\le z\}-F_Z(z)\) appears naturally in the oracle representation. It is also useful to relate the DR formulation to the IPW representation used in hirano2003efficient. Under unconfoundedness and overlap, the CATE is also identified by \[ \tau(z)=\mathbb{E}[\phi(W,p)\mid Z=z], \qquad \phi(W,p):=\frac{Y\{D-p(X)\}}{p(X)\{1-p(X)\}}. \] However, if one constructs the plug-in process using \(\phi(W,\hat p)\), the estimation error in \(\hat p\) contributes a non-negligible first-order term. Its linearization produces the correction \(-\{\mu_1(X)/p(X)+\mu_0(X)/(1-p(X))\}\{D-p(X)\}\), which is exactly the augmentation component embedded in the DR score. Consequently, \[ \sup_{z\in\mathcal Z} \left| \frac1{\sqrt n}\sum_{i=1}^n \{\phi(W_i,\hat p)-\hat\tau_{\mathrm{IPW}}\}\mathbf{1}\{Z_i\le z\} - \frac1{\sqrt n}\sum_{i=1}^n \{\psi(W_i;\eta_0)-\tau_0\}\{\mathbf{1}\{Z_i\le z\}-F_Z(z)\} \right| =o_p(1), \] where \(\hat\tau_{\mathrm{IPW}}=n^{-1}\sum_{i=1}^n\phi(W_i,\hat p)\). Thus the feasible DR process and the appropriately linearized IPW process share the same oracle limit. A formal statement and proof of this uniform IPW linearization are provided in Appendix (ref). This result may be of independent interest: it extends the estimated-propensity-score expansion of hirano2003efficient from the ATE estimator to the present ICM process, uniformly over \(z\in\mathcal Z\). It also shows that the first-order propensity-score correction required by the plug-in IPW process coincides with the augmentation component of the DR score. We therefore use the DR formulation as our baseline, since it incorporates the required correction directly into the score. \end{remark} Building on Proposition (ref), the next theorem gives the weak limit of the feasible ICM process under the null hypothesis. \begin{theorem} Suppose Assumptions (ref)-(ref) hold. Then, under $\mathbb H_0$, $$ \widehat R_n \rightsquigarrow \mathbb{G} \qquad \text{in } \ell^\infty(\mathcal Z), $$ where $\mathbb{G}$ is a mean-zero Gaussian process with covariance function $$ K(z_1,z_2) = \mathbb{E}\!\left[ \sigma_\psi^2(Z) \bigl(\mathbf{1}\{Z\le z_1\}-F_Z(z_1)\bigr) \bigl(\mathbf{1}\{Z\le z_2\}-F_Z(z_2)\bigr) \right], \qquad z_1,z_2\in\mathcal Z, $$ with $\sigma_\psi^2(Z) := \mathrm{Var}\!\left(\psi(W;\eta_0)\mid Z\right).$ \end{theorem} Theorem (ref), together with the continuous mapping theorem, gives the asymptotic null distributions of the KS and CvM statistics. For the CvM statistic, the replacement of \(F_Z\) by \(\widehat F_{Z,n}\) is asymptotically negligible; the argument is provided in Appendix. \begin{corollary} Suppose Assumptions (ref)--(ref) hold. Then, under $\mathbb H_0$, $$ KS_n \xrightarrow{d} \sup_{z\in\mathcal Z} |\mathbb G(z)|, \qquad CvM_n \xrightarrow{d} \int_{\mathcal Z} \mathbb G(z)^2\,d F_Z(z), $$ where $\mathbb G(\cdot)$ is the centered Gaussian process defined in Theorem (ref). \end{corollary} We next examine the behavior of the test under fixed alternatives. The following theorem shows that the feasible process has a deterministic drift of order \(\sqrt n\), which yields consistency of the proposed tests. \begin{theorem} Suppose Assumptions (ref)--(ref) hold. Then, under the fixed alternative $\mathbb H_1$, $$ \sup_{z\in\mathcal Z} \left| \frac{1}{\sqrt n}\widehat R_n(z)-\Gamma(z) \right| =o_p(1), $$ where $\Gamma(z):=\mathbb{E}\left[\left(\psi(W;\eta_0)-\tau_0\right)\mathbf{1}\{Z\le z\}\right]$. \end{theorem} Under \(\mathbb H_1\), the conditional moment restriction in (ref) fails. By the equivalence between the conditional moment restriction and the ICM moments, this implies \(\Gamma\not\equiv0\). Hence \(\sup_{z\in\mathcal Z}|\Gamma(z)|>0\), and, for the CvM statistic, \(\int_{\mathcal Z}\Gamma(z)^2\,dF_Z(z)>0\). Consequently, \(KS_n\) and \(CvM_n\) diverge in probability under fixed alternatives, and the proposed tests are consistent. We next consider local alternatives that approach the null at the \(n^{-1/2}\) rate. Specifically, let \(\mathbb H_{1n}\) be a sequence of alternatives of the form \[ \tau_n(z)=\tau_0+n^{-1/2}\lambda(z), \qquad z\in\mathcal Z, \] where \(\lambda:\mathcal Z\to\mathbb R\) is a measurable, non-a.s.-constant function satisfying \(\mathbb{E}[\lambda(Z)^2]<\infty\).\footnote{If \(\lambda(Z)\) is a.s. constant, then \(\tau_n(z)\) remains constant in \(z\), so \(\mathbb H_{1n}\) collapses to the null.} \begin{theorem} Suppose Assumptions (ref)--(ref) hold. Under \(\mathbb H_{1n}\), \[ \widehat R_n \rightsquigarrow \mathbb G+\Lambda \qquad \text{in } \ell^\infty(\mathcal Z), \] where \(\mathbb G\) is the centered Gaussian process in Theorem (ref) and $\Lambda(z):= \mathbb{E}\!\left[\lambda(Z)\{\mathbf{1}\{Z\le z\}-F_Z(z)\}\right]$. Moreover, $KS_n \xrightarrow{d}\sup_{z\in\mathcal Z}|\mathbb G(z)+\Lambda(z)|$ and $CvM_n \xrightarrow{d} \int_{\mathcal Z}\{\mathbb G(z)+\Lambda(z)\}^2\,dF_Z(z)$. \end{theorem} Since \(\lambda(Z)\) is not a.s. constant, \(\Lambda\not\equiv0\). Thus, under alternatives drifting toward the null at the parametric rate \(n^{-1/2}\), the test statistics have nondegenerate shifted limits, and the proposed tests exhibit nontrivial local power against \(\mathbb H_{1n}\). From the above theorems, the limiting null distributions of the continuous functionals of \(\widehat R_n\), including \(KS_n\) and \(CvM_n\), depend on the unknown covariance structure of \(\mathbb G\). The limiting distributions are nonpivotal and therefore do not yield closed-form critical values. To address this issue, we use an easy-to-implement multiplier bootstrap procedure, described next. \section{Computation of Critical Values} In this section, we introduce the multiplier bootstrap used to approximate the null distributions of \(KS_n\) and \(CvM_n\). The procedure exploits the feasible first-order representation in Proposition (ref); it is computationally simple and avoids re-estimating the first-step nuisance functions across bootstrap replications. The procedure is as follows: \begin{enumerate} {0pt} {0pt} {0pt} • Compute the feasible doubly robust scores $\hat\psi_i$, the average $\hat\tau=n^{-1}\sum_{i=1}^n\hat\psi_i$, and the empirical distribution function $\widehat F_{Z,n}(z)=n^{-1}\sum_{i=1}^n\mathbf{1}\{Z_i\le z\}$. • Generate i.i.d. multipliers $\{V_i\}_{i=1}^n$, independent of the data, with zero mean, unit variance, and bounded support. A popular choice is the Mammen two-point distribution, under which $V_i$ takes the values $1-\kappa$ and $\kappa$ with probabilities $\kappa/\sqrt{5}$ and $1-\kappa/\sqrt{5}$, respectively, where $\kappa=(\sqrt{5}+1)/2$, as suggested by mammen1993bootstrap. • Compute the bootstrap process $$ \widehat R_n^*(z) := \frac1{\sqrt n}\sum_{i=1}^n V_i(\hat\psi_i-\hat\tau)\{\mathbf{1}\{Z_i\le z\}-\widehat F_{Z,n}(z)\}, \qquad z\in\mathcal Z, $$ together with $$ KS_n^*:=\sup_{z\in\mathcal Z}|\widehat R_n^*(z)|, \qquad CvM_n^*:=\int_{\mathcal Z}\widehat R_n^*(z)^2\,d\widehat F_{Z,n}(z) = \frac1n\sum_{i=1}^n\widehat R_n^*(Z_i)^2. $$ • Repeat Steps 2--3 for $B$ times to obtain $\{KS_{n,b}^*,CvM_{n,b}^*\}_{b=1}^B$. • For a given significance level $\alpha\in(0,1)$, let $c_{KS,1-\alpha}^*$ and $c_{CvM,1-\alpha}^*$ denote the empirical $(1-\alpha)$-quantiles of $\{KS_{n,b}^*\}_{b=1}^B$ and $\{CvM_{n,b}^*\}_{b=1}^B$, respectively. Reject $\mathbb H_0$ whenever $KS_n>c_{KS,1-\alpha}^*$ or $CvM_n>c_{CvM,1-\alpha}^*$. \end{enumerate} Denote by “$\rightsquigarrow^*$ in probability” the weak convergence under the bootstrap law, that is, conditional on the original sample $\{W_i\}_{i=1}^n$. The next theorem establishes the asymptotic validity of the proposed multiplier bootstrap procedure. \begin{theorem} Suppose Assumptions (ref)--(ref) hold. Under $\mathbb H_0$ or the sequence of local alternatives $\mathbb H_{1n}$, $$ \widehat R_n^{*}\rightsquigarrow^{*}\mathbb G \qquad \text{in probability, in } \ell^\infty(\mathcal Z), $$ where $\mathbb G$ is the centered Gaussian process in Theorem (ref). Under the fixed alternative $\mathbb H_1$, $$ \widehat R_n^{*}\rightsquigarrow^{*}\mathbb G_1 \qquad \text{in probability, in } \ell^\infty(\mathcal Z), $$ where $\mathbb G_1$ is a centered Gaussian process with covariance function $$ K_1(z_1,z_2) = \mathbb{E}\!\left[ \{\psi(W;\eta_0)-\tau_0\}^2 \{\mathbf{1}\{Z\le z_1\}-F_Z(z_1)\} \{\mathbf{1}\{Z\le z_2\}-F_Z(z_2)\} \right]. $$ \end{theorem} On the one hand, under $\mathbb H_0$, Theorem (ref) shows that the multiplier bootstrap process $\widehat R_n^*$ converges weakly to the same Gaussian process $\mathbb G$ as the observed process $\widehat R_n$. Hence, the bootstrap critical values deliver asymptotically correct size for tests based on $KS_n$ and $CvM_n$. Moreover, under the local alternatives $\mathbb H_{1n}$, the bootstrap process still mimics the null Gaussian component, while Theorem (ref) shows that the observed process is shifted by the deterministic drift $\Lambda$. Therefore, the multiplier bootstrap preserves the local power properties characterized in Theorem (ref). On the other hand, under fixed alternatives $\mathbb H_1$, Theorem (ref) implies that $\widehat R_n^*$ remains stochastically bounded, although its covariance kernel generally differs from the null kernel because $\mathbb{E}[\psi(W;\eta_0)-\tau_0\mid Z]\neq0$. By contrast, Theorem (ref) shows that $KS_n$ and $CvM_n$ diverge in probability. Consequently, tests based on multiplier-bootstrap critical values are consistent against fixed alternatives. Therefore, the multiplier bootstrap procedure is asymptotically valid. \section{Extensions} The preceding sections focus on testing whether the CATE is constant with respect to the covariate of interest. We now discuss two extensions. The first concerns parametric restrictions on the CATE function. When the constant-effect null is rejected, researchers may further ask whether the detected heterogeneity can be summarized by a parsimonious and interpretable functional form, such as a linear or quadratic dependence on the covariate of interest. Such information may be useful for subsequent policy evaluation and statistical decision problems, where policy design may depend on how treatment effects vary across observable characteristics; see, for example, manski2004statistical, kitagawa2018should, and athey2021policy. The second extension considers settings in which treatment assignment may be endogenous but a binary instrument is available. In this case, building on the local average treatment effect (LATE) framework pioneered by imbens1994identification and its covariate-adjusted extensions, the proposed ICM approach can be adapted to test heterogeneity of a conditional local average treatment effect. \subsection{Testing Parametric CATE Specifications} Our main analysis tests whether the CATE function \(\tau(z)\) is constant in \(Z\). Once treatment-effect homogeneity is rejected, a natural follow-up question is whether \(\tau(z)\) can be described by a prespecified low-dimensional form. This subsection develops such a specification test. Let $\mathcal H:=\{h(\cdot,\theta):\theta\in\Theta\}$, where \(\Theta\subset\mathbb R^{d_\theta}\) is the parameter space and \(h:\mathcal Z\times\Theta\to\mathbb R\) is known up to \(\theta\). We consider the null hypothesis \[ \mathbb H_0^\dagger:\ \exists\,\theta_0\in\Theta\ \text{such that}\ \tau(z)=h(z,\theta_0)\ \text{for all }z\in\mathcal Z. \] The null \(\mathbb H_0\) in the main text is the special case \(h(z,\theta)=\theta\). Under Assumptions (ref) and (ref), $\mathbb H_0^\dagger$ is equivalently expressed as $$ \mathbb{E}\!\left[\psi(W;\eta_0)-h(Z,\theta_0)\mid Z\right]=0 \quad\text{a.s.} $$ Using the feasible doubly robust scores $\hat\psi_i$ defined in Section (ref), we estimate $\theta_0$ by nonlinear least squares: $$ Q_n(\theta) := \frac1n\sum_{i=1}^n \{\hat\psi_i-h(Z_i,\theta)\}^2, \qquad \hat\theta:=\arg\min_{\theta\in\Theta}Q_n(\theta). $$ This leads to the empirical process $$ \widehat R_n^\dagger(z) := \frac1{\sqrt n}\sum_{i=1}^n \{\hat\psi_i-h(Z_i,\hat\theta)\}\mathbf{1}\{Z_i\le z\}, \qquad z\in\mathcal Z. $$ Large values of suitable functionals of \(\widehat R_n^\dagger\) provide evidence against \(\mathbb H_0^\dagger\). Analogously to Proposition (ref), we next derive a uniform first-order representation for \(\widehat R_n^\dagger\). For a function \(h(z,\theta)\) differentiable with respect to the finite-dimensional parameter \(\theta\), we write \(\dot h_\theta(z,\theta)\) and \(\ddot h_{\theta\theta}(z,\theta)\) for its first and second derivatives, respectively. We impose the following additional conditions. \begin{assumption}[Parameter space and smoothness] The parameter space $\Theta\subset\mathbb R^{d_\theta}$ is compact. For each $z\in\mathcal Z$, the map $\theta\mapsto h(z,\theta)$ is twice continuously differentiable on a neighborhood of $\Theta$. \end{assumption} \begin{assumption}[Identification] Under \(\mathbb H_0^\dagger\), there exists a unique \(\theta_0\) in the interior of \(\Theta\) such that \(\tau(z)=h(z,\theta_0)\) for all \(z\in\mathcal Z\). \end{assumption} \begin{assumption}[Moment bounds and nonsingularity] There exist measurable envelopes $H_0(Z)$, $H_1(Z)$, and $H_2(Z)$ such that $\sup_{\theta\in\Theta}|h(Z,\theta)|\le H_0(Z)$, $\sup_{\theta\in\Theta}\|\dot h_\theta(Z,\theta)\|\le H_1(Z)$, and $\sup_{\theta\in\Theta}\|\ddot h_{\theta\theta}(Z,\theta)\|\le H_2(Z)$ almost surely, with $\mathbb{E}[H_0(Z)^2]<\infty$, $\mathbb{E}[H_1(Z)^2]<\infty$, and $\mathbb{E}[H_2(Z)^2]<\infty$. In addition, $H:=\mathbb{E}[\dot h_\theta(Z,\theta_0)\dot h_\theta(Z,\theta_0)']$ is finite and nonsingular. \end{assumption} Assumptions (ref)--(ref) are standard regularity conditions for finite-dimensional extremum estimation; see, for example, newey1994large. They ensure that the parametric projection is locally identified, sufficiently smooth, and nonsingular, so that the effect of estimating \(\theta_0\) can be accounted for by the usual first-order projection term. The next proposition gives the resulting uniform feasible-to-oracle approximation for \(\widehat R_n^\dagger\). \begin{proposition} Suppose Assumptions (ref)--(ref) hold. Under $\mathbb H_0^\dagger$, $$ \sup_{z\in\mathcal Z} \left| \widehat R_n^\dagger(z) - \frac1{\sqrt n}\sum_{i=1}^n \{\psi(W_i;\eta_0)-h(Z_i,\theta_0)\} \Bigl\{\mathbf{1}(Z_i\le z)-G(z)'H^{-1}\dot h_\theta(Z_i,\theta_0)\Bigr\} \right| =o_p(1), $$ where $G(z):=\mathbb{E}[\dot h_\theta(Z,\theta_0)\mathbf{1}\{Z\le z\}]$, $H:=\mathbb{E}[\dot h_\theta(Z,\theta_0)\dot h_\theta(Z,\theta_0)']$. \end{proposition} Proposition (ref) implies that, under \(\mathbb H_0^\dagger\), \[ \widehat R_n^\dagger \rightsquigarrow \mathbb G^\dagger \qquad\text{in }\ell^\infty(\mathcal Z), \] where \(\mathbb G^\dagger\) is a centered Gaussian process with covariance kernel \[ K^\dagger(z_1,z_2)= \mathbb{E}\Big[ \sigma_\psi^2(Z) \Bigl\{\mathbf{1}(Z\le z_1)-G(z_1)'H^{-1}\dot h_\theta(Z,\theta_0)\Bigr\} \Bigl\{\mathbf{1}(Z\le z_2)-G(z_2)'H^{-1}\dot h_\theta(Z,\theta_0)\Bigr\} \Big], \] with \(\sigma_\psi^2(Z)=\mathrm{Var}\{\psi(W;\eta_0)\mid Z\}\). The KS and CvM statistics constructed from \(\widehat R_n^\dagger\) therefore converge to the corresponding continuous functionals of \(\mathbb G^\dagger\). Critical values can be computed by the multiplier bootstrap process \begin{equation} \widehat R_n^{\dagger,*}(z) := \frac1{\sqrt n}\sum_{i=1}^n V_i\{\hat\psi_i-h(Z_i,\hat\theta)\} \Bigl\{\mathbf{1}(Z_i\le z)-\hat G(z)'\hat H^{-1}\dot h_\theta(Z_i,\hat\theta)\Bigr\}, \end{equation} where \(\hat G(z):=n^{-1}\sum_{i=1}^n \dot h_\theta(Z_i,\hat\theta)\mathbf{1}\{Z_i\le z\}\) and \(\hat H:=n^{-1}\sum_{i=1}^n\dot h_\theta(Z_i,\hat\theta)\dot h_\theta(Z_i,\hat\theta)'\). The term \(\hat G(z)'\hat H^{-1}\dot h_\theta(Z_i,\hat\theta)\) is the projection adjustment for estimating \(\theta_0\). In the constant-effect case \(h(z,\theta)=\theta\), \(\dot h_\theta(Z_i,\theta)=1\), \(\hat H=1\), and \(\hat G(z)=\widehat F_{Z,n}(z)\); hence this adjustment reduces to the empirical centering in (ref). Under \(\mathbb H_0^\dagger\) and \(\mathbb H_{1n}^\dagger\), \[ \widehat R_n^{\dagger,*}\rightsquigarrow^* \mathbb G^\dagger \qquad \text{in probability, in }\ell^\infty(\mathcal Z). \] Therefore, under \(\mathbb H_0^\dagger\), the bootstrap critical values yield asymptotically correct size for the KS and CvM tests. Under \(n^{-1/2}\)-local alternatives \(\mathbb H_{1n}^\dagger\), the bootstrap continues to approximate the null Gaussian component, whereas the observed test process is shifted by a nonzero deterministic projected drift; hence the tests have nontrivial local power. Under fixed alternatives, the observed statistics diverge while the bootstrap process remains stochastically bounded, so the tests are consistent. \subsection{Testing without Unconfoundedness} The ICM framework can also be adapted to settings in which treatment assignment may be endogenous but a binary instrument is available. We follow the local average treatment effect framework of imbens1994identification; see also abadie2003semiparametric and frolich2007nonparametric for covariate-adjusted and nonparametric formulations. Let \(W=(Y,D,A,X)\), where \(D\in\{0,1\}\) is the treatment, \(A\in\{0,1\}\) is a binary instrument, \(X\) is the full covariate vector, and \(Z\) is the covariate of interest, with \(Z\) a subvector of \(X\). Let \(D(a)\) denote the potential treatment status when the instrument is set to \(a\), and let \(Y(d,a)\) denote the potential outcome when the treatment and instrument are set to \((d,a)\). Under the exclusion restriction below, we write \(Y(d)\equiv Y(d,1)=Y(d,0)\). Let \(\mathcal C:=\{D(1)>D(0)\}\) denote the complier group. The conditional local average treatment effect with respect to \(Z\) is \[ \tau_L(z):=\mathbb{E}[Y(1)-Y(0)\mid \mathcal C,Z=z]. \] We test whether the conditional LATE is homogeneous in \(Z\): \begin{equation} \mathbb H_0^L:\ \exists \tau\in\mathcal T \text{ such that } \tau_L(z)=\tau \text{ for a.e. } z\in\mathcal Z, \quad \mathbb H_1^L:\ \mathbb{P}\{\tau_L(Z)\neq \tau\}>0,\ \forall \tau\in\mathcal T. \end{equation} We impose the following standard conditional LATE assumptions. \begin{assumption}[Conditional LATE model] \quad \begin{itemize} {0pt} {0pt} {0pt} • Instrument independence: \(A\perp \bigl(Y(1,1),Y(1,0),Y(0,1),Y(0,0),D(1),D(0)\bigr)\mid X\). • Exclusion restriction: \(Y(d,1)=Y(d,0)\) for \(d=0,1\). • Monotonicity: \(\mathbb{P}\{D(1)\ge D(0)\mid X\}=1\) a.s. • First stage: for \(q(x):=\mathbb{P}(A=1\mid X=x)\), there exists \(\epsilon>0\) such that \(\epsilon\le q(X)\le 1-\epsilon\) a.s.; moreover, \(\mathbb{P}(\mathcal C\mid Z=z)>0\) for \(z\) in the region of interest. \end{itemize} \end{assumption} Define \(\mu_{Y,a}(x):=\mathbb{E}[Y\mid A=a,X=x]\) and \(\mu_{D,a}(x):=\mathbb{E}[D\mid A=a,X=x]\), \(a=0,1\). The reduced-form and first-stage effects are identified by the doubly robust scores \[ \psi_Y^L(W;\eta_0) := \mu_{Y,1}(X)-\mu_{Y,0}(X) + \frac{A\{Y-\mu_{Y,1}(X)\}}{q(X)} - \frac{(1-A)\{Y-\mu_{Y,0}(X)\}}{1-q(X)}, \] and \[ \psi_D^L(W;\eta_0) := \mu_{D,1}(X)-\mu_{D,0}(X) + \frac{A\{D-\mu_{D,1}(X)\}}{q(X)} - \frac{(1-A)\{D-\mu_{D,0}(X)\}}{1-q(X)}. \] Under Assumption (ref), the conditional LATE is identified by the conditional Wald ratio \[ \tau_L(z) = \frac{\mathbb{E}[\psi_Y^L(W;\eta_0)\mid Z=z]} {\mathbb{E}[\psi_D^L(W;\eta_0)\mid Z=z]}. \] Under \(\mathbb H_0^L\), the common value of \(\tau_L(z)\) is invariant to the distribution used to average over \(Z\). For identification and interpretation, we average over the distribution of \(Z\) among compliers, rather than the marginal distribution of \(Z\). This yields the global LATE \[ \tau_0 := \mathbb{E}[Y(1)-Y(0)\mid \mathcal C] = \frac{\mathbb{E}[\psi_Y^L(W;\eta_0)]}{\mathbb{E}[\psi_D^L(W;\eta_0)]} = \int_{\mathcal Z}\tau_L(z)\,dF_{Z\mid\mathcal C}(z). \] Thus the testing problem in (ref) is equivalently written as the conditional moment restriction \begin{equation} \mathbb{E}\!\left[ \psi_Y^L(W;\eta_0)-\tau_0\psi_D^L(W;\eta_0) \mid Z \right] =0 \quad\text{a.s.} \end{equation} To construct the feasible process, estimate \(q\), \(\mu_{Y,a}\), and \(\mu_{D,a}\) by the same sieve procedures as in the main analysis, with the instrument \(A\) replacing the treatment indicator $D$. Let \(\widehat\psi_Y^L\) and \(\widehat\psi_D^L\) denote the resulting feasible reduced-form and first-stage scores, and estimate \(\tau_0\) by the sample Wald ratio \[ \widehat\tau_L := \frac{n^{-1}\sum_{i=1}^n \widehat\psi_Y^L(W_i)} {n^{-1}\sum_{i=1}^n \widehat\psi_D^L(W_i)}. \] The corresponding empirical process is \begin{equation} \widehat R_n^L(z) := \frac1{\sqrt n}\sum_{i=1}^n \left\{ \widehat\psi_Y^L(W_i)-\widehat\tau_L\widehat\psi_D^L(W_i) \right\} \mathbf{1}\{Z_i\le z\}, \qquad z\in\mathcal Z. \end{equation} Large values of KS or CvM functionals of \(\widehat R_n^L\) provide evidence against \(\mathbb H_0^L\). The next proposition gives the uniform feasible-to-oracle approximation for this process. \begin{proposition} Suppose Assumption (ref) holds, together with the analogues of Assumptions (ref)--(ref) for the instrument propensity score \(q\) and the nuisance regression functions entering \(\psi_Y^L\) and \(\psi_D^L\). Then, under \(\mathbb H_0^L\), \[ \sup_{z\in\mathcal Z} \left| \widehat R_n^L(z) - \frac1{\sqrt n}\sum_{i=1}^n \left\{ \psi_Y^L(W_i;\eta_0) - \tau_0\psi_D^L(W_i;\eta_0) \right\} \left\{ \mathbf{1}(Z_i\le z)-F_{Z\mid \mathcal C}(z) \right\} \right| =o_p(1), \] where \(F_{Z\mid \mathcal C}(z):=\mathbb{P}(Z\le z\mid \mathcal C)\). \end{proposition} Proposition (ref) is the IV analogue of Proposition (ref). The reduced-form residual $\psi_Y^L(W;\eta_0)-\tau_0\psi_D^L(W;\eta_0)$ plays the role of the oracle score. The centering term is \(F_{Z\mid\mathcal C}(z)\), rather than \(F_Z(z)\), because \(\tau_0\) is the global LATE averaged over the complier distribution. Under \(\mathbb H_0^L\), Proposition (ref) implies \[ \widehat R_n^L \rightsquigarrow \mathbb G^L \qquad\text{in } \ell^\infty(\mathcal Z), \] where \(\mathbb G^L\) is a centered Gaussian process with covariance kernel \[ K^L(z_1,z_2) = \mathbb{E}\!\left[ \sigma_L^2(Z) \left\{\mathbf{1}(Z\le z_1)-F_{Z\mid\mathcal C}(z_1)\right\} \left\{\mathbf{1}(Z\le z_2)-F_{Z\mid\mathcal C}(z_2)\right\} \right], \] with $\sigma_L^2(Z) := \mathrm{Var}\!\left( \psi_Y^L(W;\eta_0)-\tau_0\psi_D^L(W;\eta_0) \mid Z \right)$. Therefore, the KS and CvM statistics constructed from \(\widehat R_n^L\) converge to the corresponding continuous functionals of \(\mathbb G^L\). In implementation, critical values can be computed by the multiplier bootstrap process \begin{equation} \widehat R_n^{L,*}(z) := \frac1{\sqrt n}\sum_{i=1}^n V_i \{\widehat\psi_Y^L(W_i)-\widehat\tau_L\widehat\psi_D^L(W_i)\} \{\mathbf{1}\{Z_i\le z\}-\widehat F_{Z\mid\mathcal C}(z)\}, \end{equation} where $\widehat F_{Z\mid\mathcal C}(z):=\sum_i \widehat\psi_D^L(W_i)\mathbf{1}\{Z_i\le z\}/\sum_i \widehat\psi_D^L(W_i)$. The validity of this bootstrap procedure follows by the same arguments as in the unconfounded and parametric-extension cases. We omit the details. \section{Monte Carlo Simulation Study} In this section, we conduct a Monte Carlo study of the finite-sample performance of the proposed heterogeneity test. The designs examine empirical size when the CATE is constant and power when the CATE varies with the covariate of interest. Throughout, the covariate of interest is \(Z=X_1\), while the full covariate vector is \(X=(X_1,\ldots,X_r)'\) with \(r\in\{3,5,10\}\). We consider sample sizes \(n\in\{500,1000,2000\}\). All results are based on \(5{,}000\) Monte Carlo replications, and bootstrap \(p\)-values are computed using \(B=2{,}000\) multiplier bootstrap draws. We report empirical rejection frequencies at the 1%, 5%, and 10% nominal levels. In each replication, the propensity score is estimated by series logit, and the nuisance regressions \(\mu_1(X)\) and \(\mu_0(X)\) are estimated by sieve least squares using the same type of basis. To keep the number of series terms moderate as the dimension increases, we use a total-degree cubic polynomial basis when \(r=3\), a total-degree quadratic basis when \(r=5\), and an additive quadratic basis when \(r=10\). The additive quadratic basis contains an intercept, all linear terms, and all squared terms, but excludes interactions. Across all designs, outcomes are generated by \[ Y=\mu_0(X)+\tau(X_1)D+\varepsilon, \qquad \tau_0=0.5. \] Under the null, \(\tau(X_1)\equiv\tau_0\). Under the alternatives, \(\tau(X_1)\) varies with \(X_1\). We consider both homoskedastic disturbances, \(\varepsilon\sim \mathcal N(0,\sigma_u^2)\) with \(\sigma_u=0.25\), and heteroskedastic disturbances, \(\varepsilon\sim \mathcal N(0,\sigma_u^2\{1+|X_1-0.5|\}^2)\). The three data-generating processes differ in the covariate distribution, the treatment assignment mechanism, the baseline outcome function, and the form of treatment-effect heterogeneity, as described below. \paragraph{Case I: Smooth logit design.} Let \(X_j\stackrel{iid}{\sim}\mathrm{Unif}[0,1]\), \(j=1,\ldots,r\), and write \(U_j=X_j-0.5\). Treatment is generated from \(D\mid X\sim\mathrm{Bernoulli}(\Lambda\{m(X)\})\). When \(r=3\), \[ m(X)= -0.2+1.2U_1+0.8U_2+0.5U_3+0.6U_1U_2+0.4U_1^2+1.5U_1^3+0.4e^{-6U_2^2}, \] and $\mu_0(X)=0.5U_2+0.3U_2U_3+0.2U_1^2+0.15U_1^3$. For \(r>3\), we add \(0.2\bar U+0.2\bar E\) to \(m(X)\) and replace the last term in \(\mu_0(X)\) by \(0.15\bar U\), where \(\bar U=(r-3)^{-1}\sum_{j=4}^r U_j\) and \(\bar E=(r-3)^{-1}\sum_{j=4}^r e^{-6U_j^2}\). Under the alternative, $\tau(X_1)=\tau_0+0.2(X_1-0.5)$. \paragraph{Case II: Threshold assignment design.} Let \(X_j\stackrel{iid}{\sim}\mathrm{Unif}[0,1]\), \(j=1,\ldots,r\). Treatment is generated by \(D=\mathbf{1}\{U<p(X)\}\), where \(U\sim\mathrm{Unif}[0,1]\) is independent of \(X\), \(p(X)=0.1+0.4S(X)\), and \[ S(X)= \begin{cases} (X_1+X_2+X_3)/3, & r=3,\\ 0.8(X_1+X_2+X_3)/3+0.2(r-3)^{-1}\sum_{j=4}^r X_j, & r>3. \end{cases} \] The baseline outcome is \(\mu_0(X)=0.5(X_2-0.5)+0.5(X_3-0.5)\) when \(r=3\), and adds \(0.15\bar U\) when \(r>3\), where \(\bar U=(r-3)^{-1}\sum_{j=4}^r(X_j-0.5)\). Under the alternative, $\tau(X_1)=\tau_0+0.4(X_1-0.5)\mathbf{1}\{X_1>0.5\}$. \paragraph{Case III: Fractional latent-index design.} Let \(X_j\stackrel{iid}{\sim}\mathrm{Beta}(2,2)\), \(j=1,\ldots,r\), and write \(U_j=X_j-0.5\). Treatment is generated by \(D=\mathbf{1}\{m(X)-e>0\}\), where \(e\sim \mathcal N(0,1)\), so that \(p(X)=\Phi\{m(X)\}\). The latent index is \[ m(X)= \frac{-0.10+0.10\sum_{j=1}^r X_j^2+0.50U_1+0.40U_2+0.30U_1U_2} {\exp\{-0.10\sum_{j=1}^r X_j^2\}}, \] and the baseline outcome is $\mu_0(X)=0.4U_1+0.3U_2+0.2U_3^2+0.10r^{-1}\sum_{j=1}^r U_j$. Under the alternative, $\tau(X_1)=\tau_0+0.8(X_1-0.4)^2$. \begin{comment} \begin{table}[!htbp] \caption{Empirical size under homoskedastic errors.} \begin{tabular*}{0.85\textwidth}{@{\extracolsep{\fill}} llcccccc @} \toprule & & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} \\ \cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8} $r$ & $\alpha$ & KS & CvM & KS & CvM & KS & CvM \\ \midrule \multicolumn{8}{l}{\textit{Case I: Smooth logit design}} \\ \midrule \multirow{3}{*}{3} & 1% & 1.02 & 0.94 & 1.18 & 1.32 & 1.26 & 1.24 \\ & 5% & 5.42 & 4.98 & 5.40 & 5.18 & 5.98 & 5.72 \\ & 10% & 10.52 & 10.36 & 10.06 & 10.40 & 11.50 & 11.00 \\ \addlinespace \multirow{3}{*}{5} & 1% & 1.18 & 1.34 & 1.32 & 1.12 & 1.24 & 1.00 \\ & 5% & 5.30 & 5.16 & 5.38 & 5.28 & 5.30 & 5.22 \\ & 10% & 10.50 & 10.32 & 10.62 & 10.56 & 10.46 & 10.14 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.42 & 1.46 & 0.96 & 1.12 & 1.16 & 0.88 \\ & 5% & 5.88 & 5.62 & 5.64 & 5.50 & 4.98 & 5.22 \\ & 10% & 11.08 & 10.46 & 10.52 & 10.42 & 10.60 & 10.40 \\ \midrule \multicolumn{8}{l}{\textit{Case II: Threshold-assignment design}} \\ \midrule \multirow{3}{*}{3} & 1% & 0.94 & 0.88 & 0.84 & 0.80 & 0.96 & 0.80 \\ & 5% & 4.34 & 4.38 & 4.28 & 4.54 & 4.66 & 4.18 \\ & 10% & 9.30 & 9.50 & 9.64 & 9.42 & 9.82 & 9.74 \\ \addlinespace \multirow{3}{*}{5} & 1% & 0.62 & 0.66 & 0.98 & 0.88 & 1.02 & 1.14 \\ & 5% & 3.94 & 3.94 & 4.38 & 4.52 & 5.04 & 5.00 \\ & 10% & 8.70 & 8.72 & 9.34 & 9.70 & 10.18 & 9.78 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.08 & 0.92 & 1.22 & 1.22 & 0.98 & 1.10 \\ & 5% & 5.34 & 5.34 & 5.34 & 5.42 & 5.52 & 5.10 \\ & 10% & 10.96 & 11.42 & 11.26 & 10.44 & 10.48 & 9.94 \\ \midrule \multicolumn{8}{l}{\textit{Case III: Fractional latent-index design}} \\ \midrule \multirow{3}{*}{3} & 1% & 0.94 & 0.94 & 1.06 & 1.16 & 1.08 & 0.92 \\ & 5% & 4.96 & 4.96 & 5.04 & 5.22 & 4.80 & 4.46 \\ & 10% & 10.80 & 10.32 & 10.32 & 10.34 & 9.30 & 9.72 \\ \addlinespace \multirow{3}{*}{5} & 1% & 1.08 & 1.06 & 0.98 & 1.08 & 1.24 & 1.24 \\ & 5% & 5.48 & 5.42 & 5.30 & 4.92 & 4.94 & 4.88 \\ & 10% & 10.72 & 10.30 & 10.58 & 10.54 & 9.94 & 10.16 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.32 & 1.28 & 1.00 & 0.78 & 1.02 & 1.08 \\ & 5% & 5.90 & 6.38 & 4.96 & 4.74 & 5.18 & 5.26 \\ & 10% & 11.70 & 11.56 & 10.16 & 9.66 & 10.56 & 10.14 \\ \bottomrule \end{tabular*} \begin{minipage}{0.8\textwidth} \textit{Notes.} Entries report rejection frequencies (in percent) over 5,000 Monte Carlo replications. The nominal significance levels are 1%, 5%, and 10%. Bootstrap critical values are based on 2,000 multiplier bootstrap draws. \end{minipage} \end{table} \begin{table}[!htbp] \caption{Empirical size under heteroskedastic errors.} \begin{tabular*}{0.85\textwidth}{@{\extracolsep{\fill}} llcccccc @} \toprule & & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} \\ \cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8} $r$ & $\alpha$ & KS & CvM & KS & CvM & KS & CvM \\ \midrule \multicolumn{8}{l}{\textit{Case I: Smooth logit design}} \\ \midrule \multirow{3}{*}{3} & 1% & 1.04 & 1.08 & 1.14 & 1.24 & 1.30 & 1.22 \\ & 5% & 5.50 & 5.20 & 5.28 & 4.98 & 5.86 & 5.80 \\ & 10% & 10.90 & 10.64 & 10.12 & 10.58 & 11.20 & 10.86 \\ \addlinespace \multirow{3}{*}{5} & 1% & 1.18 & 1.36 & 1.22 & 1.14 & 1.06 & 0.90 \\ & 5% & 5.28 & 5.40 & 5.66 & 5.06 & 5.32 & 5.00 \\ & 10% & 10.86 & 10.46 & 10.76 & 10.48 & 10.50 & 10.34 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.46 & 1.54 & 1.06 & 1.16 & 1.06 & 0.84 \\ & 5% & 5.68 & 5.46 & 5.40 & 5.42 & 5.06 & 5.20 \\ & 10% & 10.82 & 10.52 & 10.72 & 10.58 & 10.40 & 10.24 \\ \midrule \multicolumn{8}{l}{\textit{Case II: Threshold-assignment design}} \\ \midrule \multirow{3}{*}{3} & 1% & 1.06 & 1.00 & 0.74 & 0.86 & 0.82 & 0.66 \\ & 5% & 4.78 & 4.40 & 4.32 & 4.46 & 4.34 & 4.12 \\ & 10% & 9.06 & 9.68 & 9.12 & 8.96 & 9.10 & 9.40 \\ \addlinespace \multirow{3}{*}{5} & 1% & 0.70 & 0.88 & 1.02 & 1.00 & 1.00 & 1.00 \\ & 5% & 4.34 & 4.56 & 4.72 & 4.82 & 4.94 & 5.00 \\ & 10% & 9.20 & 9.08 & 9.68 & 9.50 & 9.74 & 9.66 \\ \addlinespace \multirow{3}{*}{10} & 1% & 0.96 & 0.86 & 1.24 & 1.24 & 1.00 & 1.16 \\ & 5% & 4.84 & 4.98 & 5.46 & 5.24 & 5.26 & 4.86 \\ & 10% & 10.42 & 10.62 & 10.94 & 10.40 & 10.16 & 9.70 \\ \midrule \multicolumn{8}{l}{\textit{Case III: Fractional latent-index design}} \\ \midrule \multirow{3}{*}{3} & 1% & 0.94 & 1.06 & 1.20 & 1.12 & 1.00 & 0.94 \\ & 5% & 4.92 & 5.04 & 5.20 & 5.26 & 4.86 & 4.60 \\ & 10% & 10.78 & 10.52 & 10.50 & 10.34 & 9.88 & 9.94 \\ \addlinespace \multirow{3}{*}{5} & 1% & 1.08 & 1.12 & 1.00 & 1.08 & 1.20 & 1.26 \\ & 5% & 5.44 & 5.30 & 5.26 & 5.34 & 4.90 & 5.10 \\ & 10% & 10.70 & 10.76 & 10.68 & 10.52 & 9.98 & 9.78 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.22 & 1.26 & 0.94 & 0.72 & 1.10 & 1.10 \\ & 5% & 5.88 & 6.10 & 4.90 & 4.78 & 5.50 & 5.12 \\ & 10% & 11.48 & 11.58 & 9.68 & 9.76 & 10.36 & 10.28 \\ \bottomrule \end{tabular*} \begin{minipage}{0.8\textwidth} \textit{Notes.} Entries report rejection frequencies (in percent) over 5,000 Monte Carlo replications. The nominal significance levels are 1%, 5%, and 10%. Bootstrap critical values are based on 2,000 multiplier bootstrap draws. \end{minipage} \end{table} \begin{table}[!htbp] \caption{Empirical power under homoskedastic errors.} \begin{tabular*}{0.85\textwidth}{@{\extracolsep{\fill}} llcccccc @} \toprule & & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} \\ \cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8} $r$ & $\alpha$ & KS & CvM & KS & CvM & KS & CvM \\ \midrule \multicolumn{8}{l}{\textit{Case I: Smooth logit design}} \\ \midrule \multirow{3}{*}{3} & 1% & 28.32 & 34.88 & 60.42 & 68.94 & 91.94 & 95.88 \\ & 5% & 52.88 & 59.14 & 81.20 & 86.24 & 98.12 & 98.98 \\ & 10% & 65.22 & 70.62 & 88.54 & 91.94 & 99.22 & 99.52 \\ \addlinespace \multirow{3}{*}{5} & 1% & 32.74 & 39.72 & 64.56 & 72.90 & 93.96 & 96.78 \\ & 5% & 56.78 & 62.86 & 84.20 & 88.42 & 98.70 & 99.38 \\ & 10% & 69.46 & 73.86 & 90.68 & 93.68 & 99.44 & 99.70 \\ \addlinespace \multirow{3}{*}{10} & 1% & 35.58 & 42.32 & 68.58 & 76.34 & 95.46 & 97.66 \\ & 5% & 60.28 & 66.42 & 86.78 & 90.44 & 99.14 & 99.62 \\ & 10% & 72.22 & 77.06 & 92.08 & 94.34 & 99.64 & 99.84 \\ \midrule \multicolumn{8}{l}{\textit{Case II: Threshold-assignment design}} \\ \midrule \multirow{3}{*}{3} & 1% & 17.30 & 17.54 & 46.72 & 46.62 & 87.70 & 85.86 \\ & 5% & 36.94 & 38.88 & 70.86 & 71.08 & 96.12 & 95.72 \\ & 10% & 49.04 & 51.58 & 80.64 & 81.42 & 98.24 & 97.86 \\ \addlinespace \multirow{3}{*}{5} & 1% & 16.68 & 18.52 & 49.06 & 50.18 & 90.32 & 89.94 \\ & 5% & 35.16 & 38.10 & 70.96 & 73.02 & 97.06 & 97.16 \\ & 10% & 46.48 & 50.28 & 80.14 & 82.42 & 98.64 & 98.60 \\ \addlinespace \multirow{3}{*}{10} & 1% & 25.90 & 28.38 & 63.72 & 66.70 & 95.40 & 96.20 \\ & 5% & 48.98 & 52.98 & 84.10 & 86.14 & 98.84 & 99.08 \\ & 10% & 61.96 & 65.40 & 90.48 & 92.02 & 99.46 & 99.60 \\ \midrule \multicolumn{8}{l}{\textit{Case III: Fractional latent-index design}} \\ \midrule \multirow{3}{*}{3} & 1% & 13.44 & 14.48 & 33.32 & 35.36 & 72.68 & 75.02 \\ & 5% & 32.54 & 34.88 & 59.50 & 62.14 & 90.04 & 91.58 \\ & 10% & 45.72 & 48.62 & 72.14 & 75.54 & 94.98 & 96.02 \\ \addlinespace \multirow{3}{*}{5} & 1% & 16.72 & 17.78 & 37.42 & 39.04 & 76.22 & 76.38 \\ & 5% & 37.76 & 40.08 & 64.04 & 66.02 & 92.20 & 93.44 \\ & 10% & 50.98 & 53.04 & 75.54 & 77.96 & 96.42 & 97.24 \\ \addlinespace \multirow{3}{*}{10} & 1% & 21.54 & 21.92 & 42.14 & 42.44 & 77.26 & 76.90 \\ & 5% & 43.66 & 45.60 & 68.92 & 69.64 & 92.22 & 92.54 \\ & 10% & 57.40 & 58.98 & 79.52 & 80.92 & 95.88 & 96.90 \\ \bottomrule \end{tabular*} \begin{minipage}{0.8\textwidth} \textit{Notes.} Entries report rejection frequencies (in percent) over 5,000 Monte Carlo replications. The nominal significance levels are 1%, 5%, and 10%. Bootstrap critical values are based on 2,000 multiplier bootstrap draws. \end{minipage} \end{table} \begin{table}[!htbp] \caption{Empirical power under heteroskedastic errors.} \begin{tabular*}{0.85\textwidth}{@{\extracolsep{\fill}} llcccccc @} \toprule & & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} \\ \cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8} $r$ & $\alpha$ & KS & CvM & KS & CvM & KS & CvM \\ \midrule \multicolumn{8}{l}{\textit{Case I: Smooth logit design}} \\ \midrule \multirow{3}{*}{3} & 1% & 15.96 & 18.96 & 36.38 & 40.88 & 70.54 & 76.04 \\ & 5% & 36.34 & 38.80 & 60.80 & 65.20 & 87.86 & 90.36 \\ & 10% & 47.76 & 51.52 & 71.68 & 74.98 & 93.10 & 94.30 \\ \addlinespace \multirow{3}{*}{5} & 1% & 18.66 & 21.16 & 38.72 & 43.68 & 73.34 & 77.96 \\ & 5% & 38.60 & 42.00 & 63.26 & 67.52 & 89.20 & 92.10 \\ & 10% & 50.54 & 53.94 & 73.60 & 77.46 & 93.76 & 95.54 \\ \addlinespace \multirow{3}{*}{10} & 1% & 18.92 & 21.56 & 41.60 & 45.98 & 76.84 & 80.68 \\ & 5% & 39.56 & 43.52 & 65.20 & 70.22 & 91.14 & 93.08 \\ & 10% & 52.00 & 56.16 & 76.06 & 79.44 & 94.82 & 96.08 \\ \midrule \multicolumn{8}{l}{\textit{Case II: Threshold-assignment design}} \\ \midrule \multirow{3}{*}{3} & 1% & 10.84 & 11.08 & 29.12 & 28.26 & 67.34 & 63.42 \\ & 5% & 26.22 & 27.50 & 52.48 & 53.44 & 86.10 & 83.90 \\ & 10% & 37.82 & 39.72 & 64.74 & 65.68 & 92.04 & 91.46 \\ \addlinespace \multirow{3}{*}{5} & 1% & 10.40 & 11.48 & 30.94 & 31.54 & 69.78 & 68.32 \\ & 5% & 25.72 & 27.80 & 54.00 & 54.54 & 87.88 & 86.94 \\ & 10% & 36.44 & 38.84 & 65.42 & 67.26 & 92.98 & 92.92 \\ \addlinespace \multirow{3}{*}{10} & 1% & 15.46 & 16.36 & 38.66 & 38.86 & 76.82 & 76.32 \\ & 5% & 33.96 & 36.14 & 63.18 & 64.78 & 91.62 & 91.60 \\ & 10% & 46.32 & 48.56 & 74.46 & 76.00 & 95.08 & 95.76 \\ \midrule \multicolumn{8}{l}{\textit{Case III: Fractional latent-index design}} \\ \midrule \multirow{3}{*}{3} & 1% & 9.44 & 8.94 & 22.20 & 21.00 & 53.96 & 49.80 \\ & 5% & 24.70 & 24.26 & 45.64 & 44.24 & 76.84 & 76.52 \\ & 10% & 36.68 & 36.52 & 58.24 & 58.28 & 85.86 & 85.84 \\ \addlinespace \multirow{3}{*}{5} & 1% & 11.14 & 10.94 & 25.32 & 22.96 & 55.98 & 51.22 \\ & 5% & 28.10 & 28.14 & 47.92 & 47.12 & 78.30 & 76.56 \\ & 10% & 40.48 & 40.68 & 61.22 & 60.92 & 87.02 & 86.74 \\ \addlinespace \multirow{3}{*}{10} & 1% & 13.64 & 13.00 & 26.58 & 23.76 & 55.88 & 50.68 \\ & 5% & 32.02 & 30.94 & 50.76 & 48.20 & 77.88 & 75.30 \\ & 10% & 44.04 & 43.48 & 63.88 & 63.06 & 85.92 & 85.62 \\ \bottomrule \end{tabular*} \begin{minipage}{0.8\textwidth} \textit{Notes.} Entries report rejection frequencies (in percent) over 5,000 Monte Carlo replications. The nominal significance levels are 1%, 5%, and 10%. Bootstrap critical values are based on 2,000 multiplier bootstrap draws. \end{minipage} \end{table} \FloatBarrier \end{comment} The main findings are as follows. Table (ref) reports empirical rejection frequencies under the constant-CATE null. The results show accurate size control across a broad range of designs. This includes different treatment assignment mechanisms, from a smooth logit propensity score to a threshold assignment rule and a nonlinear latent-index model; different baseline outcome functions; dimensions \(r\in\{3,5,10\}\); and both homoskedastic and heteroskedastic disturbances. Across these settings, the rejection frequencies of both KS and CvM statistics remain close to the nominal 1%, 5%, and 10% levels. Thus, the multiplier bootstrap appears to provide a stable finite-sample approximation to the null distribution, even when the treatment allocation rule and the outcome equation differ substantially across designs. \begin{landscape} \begingroup \captionof{table}{Empirical size under homoskedastic and heteroskedastic errors.} \scriptsize {2.5pt} \begin{tabular*}{\linewidth}{@{\extracolsep{\fill}} ll*{12}{c} @} \toprule & & \multicolumn{6}{c}{Homoskedastic errors} & \multicolumn{6}{c}{Heteroskedastic errors} \\ \cmidrule(lr){3-8}\cmidrule(lr){9-14} & & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} \\ \cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8} \cmidrule(lr){9-10}\cmidrule(lr){11-12}\cmidrule(lr){13-14} $r$ & $\alpha$ & KS & CvM & KS & CvM & KS & CvM & KS & CvM & KS & CvM & KS & CvM \\ \midrule \multicolumn{14}{l}{\textit{Case I: Smooth logit design}} \\ \midrule \multirow{3}{*}{3} & 1% & 1.02 & 0.94 & 1.18 & 1.32 & 1.26 & 1.24 & 1.04 & 1.08 & 1.14 & 1.24 & 1.30 & 1.22 \\ & 5% & 5.42 & 4.98 & 5.40 & 5.18 & 5.98 & 5.72 & 5.50 & 5.20 & 5.28 & 4.98 & 5.86 & 5.80 \\ & 10% & 10.52 & 10.36 & 10.06 & 10.40 & 11.50 & 11.00 & 10.90 & 10.64 & 10.12 & 10.58 & 11.20 & 10.86 \\ \addlinespace \multirow{3}{*}{5} & 1% & 1.18 & 1.34 & 1.32 & 1.12 & 1.24 & 1.00 & 1.18 & 1.36 & 1.22 & 1.14 & 1.06 & 0.90 \\ & 5% & 5.30 & 5.16 & 5.38 & 5.28 & 5.30 & 5.22 & 5.28 & 5.40 & 5.66 & 5.06 & 5.32 & 5.00 \\ & 10% & 10.50 & 10.32 & 10.62 & 10.56 & 10.46 & 10.14 & 10.86 & 10.46 & 10.76 & 10.48 & 10.50 & 10.34 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.42 & 1.46 & 0.96 & 1.12 & 1.16 & 0.88 & 1.46 & 1.54 & 1.06 & 1.16 & 1.06 & 0.84 \\ & 5% & 5.88 & 5.62 & 5.64 & 5.50 & 4.98 & 5.22 & 5.68 & 5.46 & 5.40 & 5.42 & 5.06 & 5.20 \\ & 10% & 11.08 & 10.46 & 10.52 & 10.42 & 10.60 & 10.40 & 10.82 & 10.52 & 10.72 & 10.58 & 10.40 & 10.24 \\ \midrule \multicolumn{14}{l}{\textit{Case II: Threshold-assignment design}} \\ \midrule \multirow{3}{*}{3} & 1% & 0.94 & 0.88 & 0.84 & 0.80 & 0.96 & 0.80 & 1.06 & 1.00 & 0.74 & 0.86 & 0.82 & 0.66 \\ & 5% & 4.34 & 4.38 & 4.28 & 4.54 & 4.66 & 4.18 & 4.78 & 4.40 & 4.32 & 4.46 & 4.34 & 4.12 \\ & 10% & 9.30 & 9.50 & 9.64 & 9.42 & 9.82 & 9.74 & 9.06 & 9.68 & 9.12 & 8.96 & 9.10 & 9.40 \\ \addlinespace \multirow{3}{*}{5} & 1% & 0.62 & 0.66 & 0.98 & 0.88 & 1.02 & 1.14 & 0.70 & 0.88 & 1.02 & 1.00 & 1.00 & 1.00 \\ & 5% & 3.94 & 3.94 & 4.38 & 4.52 & 5.04 & 5.00 & 4.34 & 4.56 & 4.72 & 4.82 & 4.94 & 5.00 \\ & 10% & 8.70 & 8.72 & 9.34 & 9.70 & 10.18 & 9.78 & 9.20 & 9.08 & 9.68 & 9.50 & 9.74 & 9.66 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.08 & 0.92 & 1.22 & 1.22 & 0.98 & 1.10 & 0.96 & 0.86 & 1.24 & 1.24 & 1.00 & 1.16 \\ & 5% & 5.34 & 5.34 & 5.34 & 5.42 & 5.52 & 5.10 & 4.84 & 4.98 & 5.46 & 5.24 & 5.26 & 4.86 \\ & 10% & 10.96 & 11.42 & 11.26 & 10.44 & 10.48 & 9.94 & 10.42 & 10.62 & 10.94 & 10.40 & 10.16 & 9.70 \\ \midrule \multicolumn{14}{l}{\textit{Case III: Fractional latent-index design}} \\ \midrule \multirow{3}{*}{3} & 1% & 0.94 & 0.94 & 1.06 & 1.16 & 1.08 & 0.92 & 0.94 & 1.06 & 1.20 & 1.12 & 1.00 & 0.94 \\ & 5% & 4.96 & 4.96 & 5.04 & 5.22 & 4.80 & 4.46 & 4.92 & 5.04 & 5.20 & 5.26 & 4.86 & 4.60 \\ & 10% & 10.80 & 10.32 & 10.32 & 10.34 & 9.30 & 9.72 & 10.78 & 10.52 & 10.50 & 10.34 & 9.88 & 9.94 \\ \addlinespace \multirow{3}{*}{5} & 1% & 1.08 & 1.06 & 0.98 & 1.08 & 1.24 & 1.24 & 1.08 & 1.12 & 1.00 & 1.08 & 1.20 & 1.26 \\ & 5% & 5.48 & 5.42 & 5.30 & 4.92 & 4.94 & 4.88 & 5.44 & 5.30 & 5.26 & 5.34 & 4.90 & 5.10 \\ & 10% & 10.72 & 10.30 & 10.58 & 10.54 & 9.94 & 10.16 & 10.70 & 10.76 & 10.68 & 10.52 & 9.98 & 9.78 \\ \addlinespace \multirow{3}{*}{10} & 1% & 1.32 & 1.28 & 1.00 & 0.78 & 1.02 & 1.08 & 1.22 & 1.26 & 0.94 & 0.72 & 1.10 & 1.10 \\ & 5% & 5.90 & 6.38 & 4.96 & 4.74 & 5.18 & 5.26 & 5.88 & 6.10 & 4.90 & 4.78 & 5.50 & 5.12 \\ & 10% & 11.70 & 11.56 & 10.16 & 9.66 & 10.56 & 10.14 & 11.48 & 11.58 & 9.68 & 9.76 & 10.36 & 10.28 \\ \bottomrule \end{tabular*} \begin{minipage}{0.96\linewidth} \scriptsize \textit{Notes.} Entries report rejection frequencies (in percent) over 5,000 Monte Carlo replications under the constant-CATE null. The nominal significance levels are 1%, 5%, and 10%. Bootstrap critical values are based on 2,000 multiplier bootstrap draws. \end{minipage} \endgroup \begingroup \captionof{table}{Empirical power under homoskedastic and heteroskedastic errors.} \scriptsize {2.5pt} \begin{tabular*}{\linewidth}{@{\extracolsep{\fill}} ll*{12}{c} @} \toprule & & \multicolumn{6}{c}{Homoskedastic errors} & \multicolumn{6}{c}{Heteroskedastic errors} \\ \cmidrule(lr){3-8}\cmidrule(lr){9-14} & & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} & \multicolumn{2}{c}{$n=500$} & \multicolumn{2}{c}{$n=1000$} & \multicolumn{2}{c}{$n=2000$} \\ \cmidrule(lr){3-4}\cmidrule(lr){5-6}\cmidrule(lr){7-8} \cmidrule(lr){9-10}\cmidrule(lr){11-12}\cmidrule(lr){13-14} $r$ & $\alpha$ & KS & CvM & KS & CvM & KS & CvM & KS & CvM & KS & CvM & KS & CvM \\ \midrule \multicolumn{14}{l}{\textit{Case I: Smooth logit design}} \\ \midrule \multirow{3}{*}{3} & 1% & 28.32 & 34.88 & 60.42 & 68.94 & 91.94 & 95.88 & 15.96 & 18.96 & 36.38 & 40.88 & 70.54 & 76.04 \\ & 5% & 52.88 & 59.14 & 81.20 & 86.24 & 98.12 & 98.98 & 36.34 & 38.80 & 60.80 & 65.20 & 87.86 & 90.36 \\ & 10% & 65.22 & 70.62 & 88.54 & 91.94 & 99.22 & 99.52 & 47.76 & 51.52 & 71.68 & 74.98 & 93.10 & 94.30 \\ \addlinespace \multirow{3}{*}{5} & 1% & 32.74 & 39.72 & 64.56 & 72.90 & 93.96 & 96.78 & 18.66 & 21.16 & 38.72 & 43.68 & 73.34 & 77.96 \\ & 5% & 56.78 & 62.86 & 84.20 & 88.42 & 98.70 & 99.38 & 38.60 & 42.00 & 63.26 & 67.52 & 89.20 & 92.10 \\ & 10% & 69.46 & 73.86 & 90.68 & 93.68 & 99.44 & 99.70 & 50.54 & 53.94 & 73.60 & 77.46 & 93.76 & 95.54 \\ \addlinespace \multirow{3}{*}{10} & 1% & 35.58 & 42.32 & 68.58 & 76.34 & 95.46 & 97.66 & 18.92 & 21.56 & 41.60 & 45.98 & 76.84 & 80.68 \\ & 5% & 60.28 & 66.42 & 86.78 & 90.44 & 99.14 & 99.62 & 39.56 & 43.52 & 65.20 & 70.22 & 91.14 & 93.08 \\ & 10% & 72.22 & 77.06 & 92.08 & 94.34 & 99.64 & 99.84 & 52.00 & 56.16 & 76.06 & 79.44 & 94.82 & 96.08 \\ \midrule \multicolumn{14}{l}{\textit{Case II: Threshold-assignment design}} \\ \midrule \multirow{3}{*}{3} & 1% & 17.30 & 17.54 & 46.72 & 46.62 & 87.70 & 85.86 & 10.84 & 11.08 & 29.12 & 28.26 & 67.34 & 63.42 \\ & 5% & 36.94 & 38.88 & 70.86 & 71.08 & 96.12 & 95.72 & 26.22 & 27.50 & 52.48 & 53.44 & 86.10 & 83.90 \\ & 10% & 49.04 & 51.58 & 80.64 & 81.42 & 98.24 & 97.86 & 37.82 & 39.72 & 64.74 & 65.68 & 92.04 & 91.46 \\ \addlinespace \multirow{3}{*}{5} & 1% & 16.68 & 18.52 & 49.06 & 50.18 & 90.32 & 89.94 & 10.40 & 11.48 & 30.94 & 31.54 & 69.78 & 68.32 \\ & 5% & 35.16 & 38.10 & 70.96 & 73.02 & 97.06 & 97.16 & 25.72 & 27.80 & 54.00 & 54.54 & 87.88 & 86.94 \\ & 10% & 46.48 & 50.28 & 80.14 & 82.42 & 98.64 & 98.60 & 36.44 & 38.84 & 65.42 & 67.26 & 92.98 & 92.92 \\ \addlinespace \multirow{3}{*}{10} & 1% & 25.90 & 28.38 & 63.72 & 66.70 & 95.40 & 96.20 & 15.46 & 16.36 & 38.66 & 38.86 & 76.82 & 76.32 \\ & 5% & 48.98 & 52.98 & 84.10 & 86.14 & 98.84 & 99.08 & 33.96 & 36.14 & 63.18 & 64.78 & 91.62 & 91.60 \\ & 10% & 61.96 & 65.40 & 90.48 & 92.02 & 99.46 & 99.60 & 46.32 & 48.56 & 74.46 & 76.00 & 95.08 & 95.76 \\ \midrule \multicolumn{14}{l}{\textit{Case III: Fractional latent-index design}} \\ \midrule \multirow{3}{*}{3} & 1% & 13.44 & 14.48 & 33.32 & 35.36 & 72.68 & 75.02 & 9.44 & 8.94 & 22.20 & 21.00 & 53.96 & 49.80 \\ & 5% & 32.54 & 34.88 & 59.50 & 62.14 & 90.04 & 91.58 & 24.70 & 24.26 & 45.64 & 44.24 & 76.84 & 76.52 \\ & 10% & 45.72 & 48.62 & 72.14 & 75.54 & 94.98 & 96.02 & 36.68 & 36.52 & 58.24 & 58.28 & 85.86 & 85.84 \\ \addlinespace \multirow{3}{*}{5} & 1% & 16.72 & 17.78 & 37.42 & 39.04 & 76.22 & 76.38 & 11.14 & 10.94 & 25.32 & 22.96 & 55.98 & 51.22 \\ & 5% & 37.76 & 40.08 & 64.04 & 66.02 & 92.20 & 93.44 & 28.10 & 28.14 & 47.92 & 47.12 & 78.30 & 76.56 \\ & 10% & 50.98 & 53.04 & 75.54 & 77.96 & 96.42 & 97.24 & 40.48 & 40.68 & 61.22 & 60.92 & 87.02 & 86.74 \\ \addlinespace \multirow{3}{*}{10} & 1% & 21.54 & 21.92 & 42.14 & 42.44 & 77.26 & 76.90 & 13.64 & 13.00 & 26.58 & 23.76 & 55.88 & 50.68 \\ & 5% & 43.66 & 45.60 & 68.92 & 69.64 & 92.22 & 92.54 & 32.02 & 30.94 & 50.76 & 48.20 & 77.88 & 75.30 \\ & 10% & 57.40 & 58.98 & 79.52 & 80.92 & 95.88 & 96.90 & 44.04 & 43.48 & 63.88 & 63.06 & 85.92 & 85.62 \\ \bottomrule \end{tabular*} \begin{minipage}{0.96\linewidth} \scriptsize \textit{Notes.} Entries report rejection frequencies (in percent) over 5,000 Monte Carlo replications under heterogeneous alternatives. The nominal significance levels are 1%, 5%, and 10%. Bootstrap critical values are based on 2,000 multiplier bootstrap draws. \end{minipage} \endgroup \end{landscape} \FloatBarrier Table (ref) reports rejection frequencies under heterogeneous alternatives. The alternatives differ not only in the assignment and outcome equations but also in the form of treatment-effect heterogeneity: the smooth logit design uses a linear departure from homogeneity, the threshold-assignment design uses a one-sided threshold-type departure, and the fractional latent-index design uses a quadratic departure. In all cases, power increases clearly with the sample size, consistent with the fixed-alternative theory. Under homoskedastic errors, the tests have high power by \(n=2000\) in all three designs. Under heteroskedastic errors, power is lower, reflecting the reduced signal-to-noise ratio, but it still rises substantially with \(n\). Hence the tests remain informative across different forms of treatment allocation, outcome response, and treatment-effect heterogeneity. Comparing the two statistics, CvM is often slightly more powerful than KS for smooth alternatives, whereas the two statistics behave similarly in the threshold-type design. Overall, the results support the finite-sample reliability of the bootstrap approximation and the ability of the proposed tests to detect heterogeneous treatment effects across a range of allocation rules, outcome models, and heterogeneity patterns. \section{Empirical Illustration} In this section, we apply the proposed tests to examine whether the effect of maternal smoking during pregnancy on infant birth weight varies with mother's age. We use North Carolina vital-statistics records from 1988 to 2002. These data have been widely used to study treatment-effect heterogeneity in this setting; see, among others, abrevaya2015estimating, lee2017doubly, fan2022estimation, and cai2025nonparametric. Following the existing literature, we restrict the sample to first-time mothers and analyze Black and White mothers separately. We further follow cai2025nonparametric in restricting attention to mothers aged 20--30, yielding samples of 79,412 Black mothers and 274,495 White mothers. Low infant birth weight is associated with adverse health and human-capital outcomes throughout the life cycle black2007cradle,almond2011human, and maternal smoking during pregnancy is widely regarded as an important preventable cause of low birth weight kramer1987intrauterine. These concerns have motivated a growing literature on whether the effect of maternal smoking varies with mother's age. abrevaya2015estimating and lee2017doubly estimate age-specific CATEs, with their estimated profiles generally suggesting that the adverse effect of smoking becomes more pronounced at older ages, although the magnitude and statistical precision of the heterogeneity differ across studies. fan2022estimation revisit the same question using high-dimensional methods for nuisance estimation and also reports evidence of age-related variation. Focusing instead on partially conditional quantile treatment effects, cai2025nonparametric find significant age variation for White mothers but not for Black mothers. Motivated by these findings, we apply the proposed tests to assess, separately by race, whether the CATE is constant over mother's age. A rejection of the homogeneity null naturally raises the further question of whether the resulting age variation can be summarized by a parsimonious functional form. We therefore also test whether the CATE is linear in mother's age. In this way, our analysis complements the existing estimation and confidence-band evidence by providing direct tests of both CATE homogeneity and the linear specification. Here, the outcome $Y$ is infant birth weight measured in grams, the treatment indicator $D$ equals one if the mother smoked during pregnancy, and the covariate of interest $Z$ is mother's age. Following abrevaya2015estimating, the covariate vector $X$ contains 13 observed characteristics: mother's age and education, the month of the first prenatal visit, the number of prenatal visits, previous terminated pregnancies, and indicators for a male infant, marital status, missing father's age, gestational diabetes, hypertension, amniocentesis, ultrasound examinations, and alcohol use. We maintain unconfoundedness conditional on $X$, so that potential birth-weight outcomes are independent of maternal smoking status given these covariates; see abrevaya2015estimating for a detailed discussion of this identifying assumption and the choice of controls. We estimate the propensity score by series logit and the two conditional outcome regressions by sieve least squares. Our baseline nuisance basis contains an intercept, the 13 covariates, squared age, and interactions between age and the remaining covariates, for a total of 27 terms. This covariate expansion follows the empirical propensity-score implementations in abrevaya2015estimating and cai2025nonparametric. As a robustness check, we also use a 59-term age-mixed basis that adds interactions between the eight binary variables and the five continuous or count variables, together with additional centered-age terms. For each implementation, the same collection of basis terms is used in the propensity-score and outcome regressions. Overlap is limited in this application. Among mothers aged 20--30, smokers account for only 8.3% of the Black sample and 15.5% of the White sample. Under the baseline basis, 40.3% and 28.1% of observations, respectively, have estimated propensity scores below 0.05. Following abrevaya2015estimating, we therefore apply the data-driven trimming rule of crump2009dealing. For each race and basis, we first estimate the propensity score, select the cutoff $\widehat{\alpha}$, and retain observations satisfying $\widehat{\alpha}\leq \widehat{p}(X)\leq 1-\widehat{\alpha}$. The baseline basis yields cutoffs of 0.036 for Black mothers and 0.070 for White mothers, close to those reported by abrevaya2015estimating. We then re-estimate all nuisance functions on the retained sample before constructing the doubly robust scores and test statistics. As summarized in Table (ref), the two bases retain approximately 55,000--57,000 Black mothers and 166,000--182,000 White mothers, corresponding to about 70--72% and 61--66% of the respective pre-trimming samples. \begin{table}[!htbp] \caption{Analysis samples and overlap trimming} \begin{threeparttable} \begin{tabular}{llrrrr} \toprule Race & Nuisance basis & Analysis sample & Smokers (%) & $\widehat\alpha$ & Retained sample \\ \midrule Black & Baseline 27 & 79,412 & 8.3 & 0.036 & 56,838 \\ Black & Age-mixed 59 & 79,412 & 8.3 & 0.037 & 55,309 \\ White & Baseline 27 & 274,495 & 15.5 & 0.070 & 182,146 \\ White & Age-mixed 59 & 274,495 & 15.5 & 0.077 & 166,233 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft] • \textit{Notes.} The analysis sample contains first-time mothers aged 20--30 before overlap trimming. The smoking share is computed before trimming. The retained sample contains observations with estimated propensity scores in $[\widehat\alpha,1-\widehat\alpha]$. All nuisance functions are re-estimated after trimming. \end{tablenotes} \end{threeparttable} \end{table} Within each trimmed sample, we construct the doubly robust score and implement the KS and CvM tests using the Mammen multiplier bootstrap described in Section (ref). We begin with the CATE homogeneity null \(\mathbb H_0\). Panel A of Table (ref) reports the bootstrap \(p\)-values. Under the baseline basis, the homogeneity test yields KS and CvM \(p\)-values of 0.392 and 0.621 for Black mothers, so the null of a constant CATE is not rejected. The conclusion is unchanged under the age-mixed basis, for which the corresponding \(p\)-values are 0.292 and 0.443. For White mothers, by contrast, both \(p\)-values equal 0.0005 under the baseline basis, and the same result is obtained under the age-mixed basis. The homogeneity null is therefore strongly rejected for White mothers under both bases. These separate-sample results should not be interpreted as a formal test that the Black and White CATE functions differ. The qualitative pattern is nevertheless consistent with cai2025nonparametric, who find significant age variation in partially conditional quantile treatment effects for White mothers but not for Black mothers, although their estimand differs from the CATE considered here. \begin{table}[!htbp] \caption{Bootstrap \(p\)-values for CATE heterogeneity and linearity tests} \begin{threeparttable} \begin{tabular*}{0.96\textwidth}{ @{\extracolsep{\fill}} llrcccc @ } \toprule Race & Nuisance basis & \(n\) & KS & CvM & \(\widehat{\theta}_0\) & \(\widehat{\theta}_1\) \\ \midrule \multicolumn{7}{l}{\textit{Panel A: CATE heterogeneity test}} \\ \addlinespace[2pt] \multirow{2}{*}{Black} & Baseline 27 & 56,838 & 0.392 & 0.621 & -- & -- \\ & Age-mixed 59 & 55,309 & 0.292 & 0.443 & -- & -- \\ \addlinespace[2pt] \multirow{2}{*}{White} & Baseline 27 & 182,146 & 0.0005 & 0.0005 & -- & -- \\ & Age-mixed 59 & 166,233 & 0.0005 & 0.0005 & -- & -- \\ \midrule \multicolumn{7}{l}{ \textit{Panel B: Linear CATE test and coefficient estimates} } \\ \addlinespace[2pt] \multirow{2}{*}{Black} & Baseline 27 & 56,838 & 0.238 & 0.382 & -85.74 & -2.31 \\ & Age-mixed 59 & 55,309 & 0.119 & 0.157 & -58.71 & -3.37 \\ \addlinespace[2pt] \multirow{2}{*}{White} & Baseline 27 & 182,146 & 0.088 & 0.120 & -68.36 & -6.29 \\ & Age-mixed 59 & 166,233 & 0.137 & 0.154 & -70.64 & -6.05 \\ \bottomrule \end{tabular*} \begin{tablenotes}[flushleft] • \textit{Notes.} Panel A reports bootstrap \(p\)-values for the CATE homogeneity null, \(\tau(z)=\tau_0\). Panel B reports coefficient estimates and bootstrap \(p\)-values for the linear CATE restriction, \(\tau(z)=\theta_0+\theta_1 z\). The \(p\)-values are based on 2,000 Mammen multiplier bootstrap draws; 0.0005 is the smallest attainable value. \end{tablenotes} \end{threeparttable} \end{table} The rejection of CATE homogeneity for White mothers motivates a more structured question: whether the age variation can be summarized by a linear function. We therefore apply the projected ICM test developed in Section (ref) to $\mathbb H_0^{\dagger}:\tau(z)=\theta_0+\theta_1 z$. Panel B of Table (ref) reports the corresponding coefficient estimates and bootstrap \(p\)-values. For White mothers, the baseline KS and CvM \(p\)-values are 0.088 and 0.120, respectively, while the age-mixed basis yields \(p\)-values of 0.137 and 0.154. Thus, the linear CATE restriction is not rejected at the 5% level under either basis, although the baseline KS result is marginal at the 10% level. Together with the strong rejection of homogeneity, these results are consistent with an approximately linear age profile over the retained age range. The estimated slopes are \(-6.29\) and \(-6.05\) grams per year under the baseline and age-mixed bases, respectively, implying that the estimated adverse effect of smoking becomes about 6 grams larger in magnitude for each additional year of mother's age. For completeness, the linear restriction is also not rejected for Black mothers, for whom the homogeneity null was already not rejected. \begin{figure}[H] \begin{minipage}{0.88\textwidth} \begin{subfigure}[t]{0.485\linewidth} \caption{Black mothers} \end{subfigure} \begin{subfigure}[t]{0.485\linewidth} \caption{White mothers} \end{subfigure} \end{minipage} \caption{Doubly robust local-linear CATEF estimates of the effect of maternal smoking on birth weight, with 95% uniform confidence bands following lee2017doubly. The horizontal line in each panel denotes the estimated ATE; the additional line for White mothers is the fitted linear CATE profile from the projected ICM procedure, corresponding to a specification not rejected at the 5% level.} \end{figure} To complement the formal tests, Figure (ref) plots doubly robust local-linear estimates of the CATEF, together with 95% uniform confidence bands constructed using the method proposed by lee2017doubly. The estimates use the baseline nuisance basis and the same overlap-trimmed samples as the formal tests. For Black mothers, the estimated CATEF fluctuates around the ATE without a clear systematic age profile, and the horizontal ATE line remains within the uniform confidence band over the displayed age range. For White mothers, the estimated effect becomes more negative with age, while the horizontal ATE line does not capture this movement and lies outside the uniform confidence band over parts of the age range. The additional line in the White-mother panel is the fitted linear CATE profile obtained from our projected ICM procedure. It corresponds to the linear restriction that is not rejected at the 5% level and closely tracks the overall downward pattern of the nonparametric estimate. These graphical patterns are consistent with the non-rejection of CATE homogeneity for Black mothers and, for White mothers, the rejection of homogeneity together with the non-rejection of the linear restriction. \section{Conclusion} This paper develops orthogonal integrated conditional moment tests for CATE heterogeneity with respect to covariates of interest, while allowing a richer set of controls to be used for treatment-selection adjustment. Based on an augmented inverse-probability-weighted score, we construct Kolmogorov--Smirnov and Cramér--von Mises tests. Our main technical result establishes a uniform feasible-to-oracle approximation for the indexed process when the propensity score and outcome regressions are estimated by series logit and sieve least squares. This approximation yields the null asymptotic distributions, consistency against fixed alternatives, nontrivial power against \(n^{-1/2}\)-local alternatives, and a computationally simple multiplier bootstrap. The framework also accommodates projected specification tests of prespecified parametric CATE restrictions and conditional-LATE heterogeneity. The simulations show accurate size and increasing power against heterogeneous alternatives. In the maternal-smoking application, CATE homogeneity with respect to mother's age is not rejected for Black mothers but is strongly rejected for White mothers. For White mothers, a linear age profile is not rejected at the 5% level. These findings illustrate how direct specification tests can complement CATE estimation and visualization by assessing whether apparent heterogeneity is statistically significant and whether it admits a parsimonious representation. \putbib
bibunit