EconBase
← Back to paper

Robust inference for the treatment effect variance in experiments using machine learning

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

122,680 characters · 19 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust inference for the treatment effect variance in experiments using machine learning

\thispagestyle{empty}

abstract\singlespacing Experimenters often collect baseline data to study heterogeneity. I propose the first valid confidence intervals for the VCATE, the treatment effect variance explained by observables. Conventional approaches yield incorrect coverage when the VCATE is zero. As a result, practitioners could be prone to detect heterogeneity even when none exists. The reason why coverage worsens at the boundary is that all efficient estimators have a locally-degenerate influence function and may not be asymptotically normal. I solve the problem for a broad class of multistep estimators with a predictive first stage. My confidence intervals account for higher-order terms in the limiting distribution and are fast to compute. I also find new connections between the VCATE and the problem of deciding whom to treat. The gains of targeting treatment are (sharply) bounded by half the square root of the VCATE. Finally, I document excellent performance in simulation and reanalyze an experiment from Malawi. \noindentKeywords: Debiased machine learning, treatment effect heterogeneity, experiments, non-standard inference, variance decomposition \\ \noindentJEL Codes C14, C21, C55, C90

\onehalfspacing \selectfont

Introduction

In recent years, there has been a rapid expansion of experiments to evaluate public policy programs and corporate initiatives. There is also more evidence that the effectiveness of a program can vary across individuals. For instance, dizon2019parents studies a population of low-income parents in Malawi with large misperceptions about their children's school performance. She finds that a simple intervention can bridge these gaps and that the Conditional Average Treatment Effect (CATE) varies by children's initial scores. In practice, even though researchers collect baseline surveys with many characteristics, the CATE is typically estimated via regressions with one or two interactions, thus underutilizing the full set of variables. The promise of leveraging the vast and readily available baseline data has sparked more applications of supervised machine learning crepon2021cream,davis2020rethinking,deryugina2019mortality. Methods such as LASSO, neural networks, random forests, or boosting are data-driven and allow for more variables and flexibility.

In this paper, I focus on the unconditional variance of the CATE, the VCATE, which measures the dispersion of treatment effects predicted by a set of baseline characteristics. The VCATE has a clear interpretation, even if the CATE is nonlinear or depends on many characteristics. chernozhukov2020generic and ding2019decomposing separately propose estimators with a misspecification robust interpretation, whereas levy2021fundamental propose an efficient estimator. Despite these recent advances, there are currently no valid confidence intervals for the VCATE. In fact, levy2021fundamental study the performance of confidence intervals based on the efficient influence function, i.e., the conventional way. They find that coverage degrades near the boundary, reaching a low of 32% in simulations. They speculate that poor coverage is due to the degeneracy of the efficient influence function when the VCATE is zero. As a result, conventional guarantees for $\sqrt{n}-$asymptotic normality do not apply to this part of the parameter space. Moreover, a VCATE close to zero is economically meaningful because it could reflect null effects, low effect heterogeneity, or irrelevant covariates. For such situations, which are common in practice, conventional approaches to inference could be misleading.

This paper provides fresh insights regarding the VCATE and proposes a solution for inference in experiments. I propose (a) novel ways to interpret the VCATE for decision-making by deriving sharp bounds on the population gains of personalized treatment assignment, (b) novel estimators of the VCATE that are both misspecification-robust and efficient, combining the best features of previous approaches, and (c) novel confidence intervals that are shape-adaptive and fast to compute.

I break down why conventional confidence intervals have incorrect size. I show that the boundary inference problem can manifest even when the CATE is linear and univariate. I solve the problem for the linear case by proposing adaptive confidence intervals that meet the high-level conditions outlined in andrews2020generic. I then show that these can be readily extended to the class of nonlinear models in chernozhukov2020generic that combine regression adjustments and a machine learning first stage. In addition to providing new confidence intervals, my paper has novel implications for estimation by showing that only a subset of these multi-step estimators is efficient. Adaptive inference and regression adjustments work well for various predictive models under weak assumptions.\footnote{While I specialize my results to the VCATE, my inference approach relies on more general principles: I use knowledge of the limiting distribution function, conditional on the cross-fitted estimates. Correct coverage follows from verifying high-level assumptions that can be satisfied by a wide array of machine learning method used in the first step. In principle, my approach could extend to other non-standard inference problems.} In the fully nonparametric case, I use a conservative procedure with valid coverage over multiple sample splits.\footnote{This adjustment is in the spirit of chernozhukov2020generic, who propose robust t-tests assuming a conditionally normal distribution. Their results are not directly applicable here due to the boundary inference problem. However, I apply the principles behind their “median-parameter” confidence intervals to my adaptive intervals.} I derive the local power curve for the associated tests of homogeneity and their relationship to the tests in crump2008nonparametric and ding2019decomposing. I also propose confidence intervals for settings with cluster dependence.

I document excellent root mean square error (RMSE) performance and coverage in simulations using LASSO, even in high dimensions. I benchmark my multi-step approach against a two-step debiased machine learning estimator. As predicted by theory, all approaches are asymptotically normal, efficient, and have good coverage in highly heterogeneous designs. However, when the VCATE is zero or close to zero, coverage of two-step alternatives can be as low as 45%. By contrast, my adaptive intervals produce coverage at the intended 95% level and better RMSE at all regions of the parameter space. I study the robustness of the multi-step approach in both the theory and simulations. I consider situations where the predictive component is misspecified or slow to converge. I also discuss issues related to uniform vs. pointwise coverage.

I apply my approach to data from dizon2019parents, an information experiment with low-income parents in Malawi who had at least two school-age children. The intervention redesigned the way in which parents received information about their children's school performance. The endline survey measured parental beliefs about student grades and asked parents to allocate tickets to a scholarship between their children. dizon2019parents presented graphs with a non-parametric CATE by baseline test scores, which had an approximately linear shape, and separately tested for significance using a regression with an interaction. I use LASSO to compute the VCATE for the two outcomes (parental beliefs and lottery allocations) by different characteristics of students, parents, and households. To make the results interpretable, I focus on the standard deviation of the CATE, i.e., $\sqrt{VCATE}$, and normalize it by the standard deviation of the outcome in the control group.

My approach allows us to quantify the magnitude of effect heterogeneity. I find that the treatment effect heterogeneity explained by test scores is equivalent to 40% of the standard deviation (SD) of the beliefs of the control group, and 16% of the SD of the control group lottery allocation. I also find that the effect heterogeneity collectively explained by other student variables (grade, age, gender, attendance, and educational expenditures) is comparable to 11% of the SD of beliefs in the control group. The VCATE of beliefs by student variables is significant at the 5% level, but the VCATE of lottery outcomes by student variables is not. The combined VCATE associated with student scores and 12 other key characteristics has a similar value to the VCATE with only scores. Despite being conservative, the intervals for the VCATE are short in length in this empirical example. Using my new welfare bounds $(-|ATE| + \sqrt{VCATE + ATE^2})/2$, I predict that targeted interventions using the baseline covariates have a maximum added benefit of 7.9% SD and 7.4% SD (standard deviations of the outcome for the control group) on beliefs and lottery allocations, respectively.

Researchers should focus on the VCATE because it is a model-free quantity with good properties: it is well-defined even if the CATE is continuous or discrete, and it weakly increases when researchers add more covariates to their analysis. Researchers can test for homogeneity by evaluating whether confidence intervals for the VCATE include zero. In addition to testing, by quantifying the VCATE, researchers can compare the magnitude of heterogeneity relative to a benchmark, such as the variance at baseline, the VCATE for different covariates, or experiments in other sites.

Contribution

My first main contribution is to show that the VCATE provides a bound for the welfare gains of policy targeting. A policymaker might decide to use the information from the CATE to design personalized treatment recommendations manski2004statistical,athey2021policy,kitagawa2018should,mbakop2021model. One can measure utilitarian welfare by computing the expected outcome under different policies. Such policies can be further constrained to a class that respects budget limits, incentive compatibility, or fairness considerations viviano2020fair,sun2021empirical. I show that the difference in mean outcomes between a targeted policy and a non-targeted policy using only the average treatment effect (ATE), is bounded by $\sqrt{VCATE}/2$. For instance, under homogeneity ($VCATE = 0$), there are no gains from targeting. I show that this bound holds in the population regardless of the choice of policy class and the underlying distribution. Furthermore, the bound is sharp in the sense that it holds exactly for at least one policy and distribution.

The proposed bound on utilitarian welfare communicates information to practitioners about whether a targeting exercise is even worth pursuing, without needing to solve the targeting problem itself. The VCATE can be a supplemental quantity reported in regression analyses, or a benchmark for analysts choosing the optimal policy. If the VCATE is very low, practitioners may consider expanding the set of covariates in the analysis. To derive the bound, I use a constructive approach to solve the most adversarial distribution. I also prove a more general bound $(-|ATE| + \sqrt{VCATE + ATE^2})/2$, and show that the distribution that leads to a maximum welfare gain is one where the CATE has binary support and mean zero. The gains from targeting easily diminish if the value of the ATE is relatively higher than the VCATE.

My second contribution is related to efficient estimation and robust inference. New theory is required here because of a unique feature of the VCATE: the efficient influence function is degenerate when the CATE is homogeneous levy2021fundamental. Classical results by newey1990semiparametric show that any regular, efficient estimator can be decomposed as $\frac{1}{n}\sum_{i=1}^n \varphi_i + R_n$, where $n$ is the sample size, $\{\varphi_i\}_{i=1}^{n}$ are a set of i.i.d. mean-zero influence functions, and $R_n$ is a residual with higher-order terms that are $o_p(n^{-1/2})$.\footnote{Many standard estimators can achieve this property, e.g., the “debiased machine learning” estimator chernozhukov2018double or the targeted maximum likelihood estimator in levy2021fundamental.} Conventional approaches assume that $\mathbb{V}(\varphi) > 0$ and in this case, the estimation error converges at $\sqrt{n}$ to $\mathcal{N}(0,\mathbb{V}(\varphi_i))$ by the CLT. However, when the VCATE is zero, $\mathbb{V}(\varphi_i) = 0$ as well. Hence, the limiting distribution is dominated by the higher terms in $R_n$, which may not be asymptotically normal. Therefore, while $\sqrt{n}-$estimation is still possible, t-tests that plug in an estimate of $\mathbb{V}(\varphi)$ may have incorrect coverage. By contrast, other common quantities such as the average treatment effect (ATE) or the local average treatment effect (LATE) do not have this problem because they satisfy $\mathbb{V}(\varphi_i) > 0$ uniformly chernozhukov2018double.

I start by analyzing a simple two step estimator, assuming that the CATE is linear in covariates and can be estimated from a regression. I show that the limiting distribution of the VCATE estimator can be written as a linear combination of a Chi-square that converges at $n$-rate and a normal distribution that converges at $\sqrt{n}-$rate. The weights are determined by the value of the VCATE, which means that the shape of the distribution changes depending on the region of the parameter space. At the boundary, it behaves like a rescaled chi-square, is $O_p(n^{-1})$, and confidence intervals with normal critical values will have incorrect coverage. For values of the VCATE bounded away from zero, the distribution is asymptotically normal as in the classical results.

In the linear case, I construct adaptive confidence intervals that account for the higher terms of the distribution. I apply the framework of andrews2020generic to show that this produces uniform, exact coverage when the linear model is correctly specified.\footnote{This type of strategy has proven effective to deal with other non-standard problems where the shape of the limiting distribution depends on an unknown parameter, such as the AR coefficient in a time series, the effect parameter under weak instruments, or the quasi-likelihood ratio test for nonlinear regression andrews2020generic.} The intervals are fast to compute because the expressions are all analytic. When there is a single covariate, I also show that a homogeneity test that evaluates whether zero is contained in the confidence intervals is algebraically identical to (i) a test of whether the interaction in the regression model is equal to zero, and (ii) the single-covariate homogeneity test of crump2008nonparametric. The test is also asymptotically equivalent to ding2019decomposing. However, these other tests only apply to series estimators and are not nested with mine in the multivariate and non-parametric cases.

I extend my results to the class of nonlinear models proposed by chernozhukov2020generic. chernozhukov2020generic showed that in experiments with known assignment probabilities, their models produce a meaningful pseudo-VCATE even if the functional form is misspecified. The pseudo-VCATE is non-negative, weakly lower than the VCATE, and converges to the true value under mild conditions on the estimated CATE.\footnote{This monotonicity property means that in experiments the pseudo-VCATE will not falsely detect heterogeneity, even if the machine learning stage is misspecified.} chernozhukov2020generic argue that the pseudo-VCATE might be of independent interest as a measure of model fit.\footnote{ ding2019decomposing also define a similar pseudo-VCATE based on randomization inference.} They describe a three-step estimator with a machine learning/prediction first-stage, a regression second-stage, and a sample-variance third stage. Related multi-step estimators have also been considered in other work guo2021machine.

To the best of my knowledge, there are no existing asymptotic results for the multi-step VCATE estimator proposed in chernozhukov2020generic. I fill in that gap by proving two sets of results. First, I show that all the estimators in their class converge to the true VCATE at least at $\sqrt{n}$-rate, are $o_p(n^{-1/2})$ at the boundary (as in the simple linear model), and have the convenient property that they are always non-negative. This builds on the asymptotic expansion for the linear case I introduced above. Second, I prove that only a subset of the chernozhukov2020generic estimators are efficient, i.e. converge at $\sqrt{n}-$ to an average of i.i.d. efficient influence functions. The key ingredient is to prove a novel finite-sample equivalence result. I find that the first order conditions of the regression step and the bias-correction component of the VCATE influence function are in fact identical, given a particular decomposition of the nuisance functions. The asymptotic results follow from fairly standard assumptions on convergence rates chernozhukov2018double,belloni2017program. To get the limiting distribution, the only meaningful extra assumption is that the estimated CATE has bounded kurtosis (thin tails).

I show that extending the adaptive confidence intervals (CIs) to multi-step estimators is straightforward. The procedure randomly splits the data into subsets or folds and estimates the nuisance functions and the VCATE on different folds. To compute the confidence intervals for a particular fold, the researcher can treat the second-stage regression as if the variables were given, and then construct the CIs as in the simple case. I construct median confidence intervals (CIs) to aggregate information across multiple folds. I show that the single fold procedure produces uniform, exact coverage for the pseudo-VCATE and point-wise, exact coverage for the VCATE for all points in the parameter space, at a nominal level $1-\alpha$. The multifold CIs have pointwise conservative coverage.

Furthermore, the probability that the true VCATE is below the confidence interval bounds is uniformly bounded by $\alpha$ in large samples. This result applies to the single and multifold CIs and does not require that the first-stage estimates to converge. Instead it relies on the fact that in experiments the pseudo-VCATE is weakly lower than the true VCATE. Tests for homogeneity (whether zero is contained in the CI) belong to this broader class of tests. Having uniform size control for this class of one-sided tests means that my tests of homogeneity are robust.

This paper is also related to a growing literature on debiased-machine learning chernozhukov2016locally,chernozhukov2022automatic,belloni2014inference,belloni2017program,chernozhukov2018double, semiparametric efficiency newey1990semiparametric, uniform inference for non-standard problems andrews2020generic, and tests of treatment effect homogeneity ding2019decomposing,crump2008nonparametric,heckman1997making,bitler2017can. My approach combines results from these literatures by addressing a boundary inference problem with a machine learning stage, and applying techniques of uniform inference. A related literature also focuses on confidence intervals around point-predictions of the CATE athey2019generalized,semenova2021debiased, rather than overall measures of dispersion.

Section (ref) provides key definitions, introduces the welfare bound, and presents a version of the adaptive confidence intervals for the univariate regression case. Section (ref) frames the inference problem in a more general setting, and extends the adaptive confidence intervals for VCATE estimation with a machine learning first stage. Section (ref) presents the large sample theory. Section (ref) introduces the simulations. Section (ref) applies my approach to an empirical example from Malawi. Section (ref) concludes.

Overview of framework

Consider a program evaluation setting in which an individual is assigned to either a treatment $(D=1)$ or a control group $(D= 0)$. The outcome of interest $Y$ depends on the treatment status. I denote the potential outcome under treatment and control status as $Y_1$ and $Y_0$, respectively, and the treatment effect as $Y_1 - Y_0$. The conditional average treatment effect (CATE) given covariates $X$ is defined as $$ \tau(x) := \mathbb{E}[Y_1 - Y_0 \mid X = x],$$ and the average treatment effect (ATE) is defined as $\tau_{av} := \mathbb{E}[Y_1 - Y_0]$. This paper proposes an estimator of the variance of the CATE (VCATE) defined as $$ V_{\tau} := \mathbb{V}(\tau(X)).$$ The variance $V_\tau$ measures the dispersion of treatment effects that can be attributed to observable characteristics $X$. The value of $V_{\tau}$ depends on the choice of covariates. To understand how different covariates might impact the VCATE, let $V_{\tau}' = \mathbb{V}(\mathbb{E}[Y_1 - Y_0 \mid X'])$ be the VCATE for a different set of covariates $X'$.

lemIf $X$ is $X'$-measurable, then $V_{\tau} \le V_{\tau}' \le \mathbb{V}(Y_1-Y_0)$.

Lemma (ref) shows that the VCATE has the following monotonicity property: if the researcher adds more covariates to the analysis, or breaks down an existing covariate into more categories, then the VCATE will be weakly larger.

The propensity score, $p(x)$, is defined as follows

equation[equation omitted — 70 chars of source]

I restrict attention to experimental settings where $p(x)$ is known. The CATE can be identified under further assumptions.

assump(i) Stable unit treatment value assumption (SUTVA), $Y = Y_1D + (1-D)Y_0$ (ii) Strong overlap, there is a constant $\delta \in (0,1/2)$ such that $\mathbb{P}(\delta < p(X) < 1-\delta) = 1$, (iii) Selection on observables, $Y_1,Y_0 \ \indep \ D \mid X$.

Assumption (ref).(i) formalizes the idea that the researcher can only observe either $Y_1$ or $Y_0$, but not both, for any particular individual. Assumption (ii) holds in randomized controlled trials with treatment probabilities bounded away from $\{0,1\}$. Assumption (iii) states that an individual's treatment probability depends on $X$ but not their potential outcomes. Let $\mu_d(x)$ be the conditional mean of $Y$ given $X$ and a fixed value of $d \in \{0,1\}$,

equation[equation omitted — 87 chars of source]

Under Assumption (ref), $\mathbb{E}[Y_d \mid X =x ] = \mu_d(x)$, and hence $\tau(x) = \mu_1(x) - \mu_0(x)$. This means that the VCATE is identified, with $V_{\tau} = \mathbb{V}(\mu_1(X)-\mu_0(X))$.

The VCATE and policy targeting

Practitioners can use estimates of $\tau(X)$ to decide whom to treat in future interventions athey2021policy,kitagawa2018should,manski2004statistical. Program managers can target the treatment recipients based on their initial covariates. However, whether targeting can substantially improve average outcomes depends on the dispersion of $\tau(x)$. I show that a simple function of the VCATE bounds the marginal gains of targeting.

Let $\gamma$ denote the joint distribution of $(X,Y_1,Y_0)$, $\mathcal{X}$ a set containing the support of $X$, $\tau_\gamma(x) := \mathbb{E}_\gamma[Y_1 - Y_0 \mid X =x]$ the CATE given $\gamma$. A function which maps $x$ to a probability of treatment $\pi(x)$ is known as a statistical allocation rule manski2004statistical. Furthermore, I denote the set of all possible allocation rules by $\Pi$, which contains all functions $\{\pi:\mathcal{X} \to [0,1]\}$. The set $\Pi$ includes many well-known assignment rules. For instance, it includes the “non-targeted” policy which assigns everyone to treatment if $\mathbb{E}_\gamma[Y_1] > \mathbb{E}_\gamma[Y_0]$ and to the control group otherwise. Moreover, the average outcome under rule $\pi$ is $\mathbb{E}_\gamma[\pi(X)Y_1 + (1-\pi(X))Y_0]$, and the marginal benefit compared to the non-targeted policy is defined as $\mathcal{U}_\gamma(\pi) := \mathbb{E}_\gamma[\pi(X)Y_1 + (1-\pi(X))Y_0] - \max\{\mathbb{E}_\gamma[Y_1],\mathbb{E}_\gamma[Y_0]\}$.

thmLet $\Gamma$ denote the set of distributions such that $\mathbb{V}_\gamma[\tau_\gamma(X)] = V_{\tau}$. For all $\gamma \in \Gamma$ and $\pi \in \Pi$, $$ \mathcal{U}_\gamma(\pi) \le \underbrace{\left(\begin{array}{c} \text{Welfare} \\ \text{Optimal} \\ \text{Targeting} \end{array}\right)}_{\sup_{\pi \in \Pi} \mathbb{E}_\gamma[\pi(X)Y_1 + (1-\pi(X))Y_0]} - \underbrace{\left(\begin{array}{c} \text{Welfare} \\ \text{No} \\ \text{Targeting} \end{array}\right)}_{\max\{\mathbb{E}_\gamma[Y_1],\mathbb{E}_\gamma[Y_0]\}} \le \frac{1}{2}\sqrt{V_{\tau}}.$$ The bound is sharp in the sense that $\mathcal{U}_\gamma(\pi) = \frac{1}{2}\sqrt{V_{\tau}}$ for at least one $\gamma \in \Gamma$ and $\pi \in \Pi$.
thmConsider distributions where $\mathbb{E}_\gamma[\tau_\gamma(X)]= \tau_{av}$ and $\mathbb{V}_\gamma[\tau_\gamma(X)] = V_{\tau}$, then $U_\gamma(\pi)\le \frac{1}{2}\left(-|\tau_{av}| + \sqrt{V_{\tau}+\tau_{av}^2}\right)$. This bound is sharp over this subset of distributions.

Theorem (ref) shows the VCATE provides a welfare bound over the superset of policy classes. Furthermore, any type of restrictions on $\Pi$ such as budget constraints or incentive compatibility will achieve utilitarian welfare gains that are weakly lower than $\frac{1}{2}\sqrt{V_\tau}$. The bound in Theorem (ref) and the generalization in Theorem (ref) provide simple bounds on the prospective gains of targeting, without needing to solve for $\pi(x)$. The bounds are most informative when $V_{\tau}$ is low. For instance, when $V_{\tau} = 0$ there is no heterogeneity explained by the observables $X$ and therefore there are no gains from targeting. However, the fact that the bound is sharp does not imply that it is always achievable for every $\gamma \in \Gamma$, and when $V_{\tau}$ is high it is still be necessary to optimize $\pi(x)$ to determine whether personalized offers are worthwhile.

Finding the bound in Theorem (ref) relies on two important insights. On one hand, the optimal policy in $\Pi$ treats an individual if and only if $\tau(x) \ge 0$ kitagawa2018should. Substituting the optimal policy, $\sup_{\pi \in \Pi} \mathcal{U}_\gamma(\pi)$ is equal to $ \mathbb{E}_\gamma[\max\{\tau_\gamma(X),0\}] - \max\{\mathbb{E}_\gamma[\tau_\gamma(X)],0\}$. On the other hand, to avoid optimizing over all $\gamma \in \Gamma$, I break the problem down into equivalence classes based on the moments of the negative, zero, and positive components of the CATE. I use a constructive approach to derive the most “adversarial” distribution. The upper bound is achieved when the CATE has a binary support, which is partly why the bounds in Theorems (ref) and (ref) have simple closed forms.

corLet $\kappa_1,\kappa_2 \in \mathbb{R}$ and define a new outcome $\tilde{Y} = \kappa_1 + \kappa_2 Y $. The maximum welfare gain for the transformed outcome is $\frac{|\kappa_2|}{2}(-|\tau_{av}| + \sqrt{V_{\tau} + (\tau_{av})^2})$.

Corollary (ref) shows that the welfare bound is invariant to location shifts in the outcome, and grows linearly with scale shifts. This result implies that transformations that change the sign, e.g. $\kappa_2 = -1$, do not change the value of the welfare bound. Consequently, the bound applies regardless of whether the welfare objective is to increase a desirable outcome or to decrease an undesirable outcome.

Inference using regressions

Consider a simple situation where $X$ is real valued, the treatment $D$ is experimentally assigned with constant probability, and $U$ is a mean zero error term. The researcher runs the following linear regression,

equation[equation omitted — 120 chars of source]

Define the auxiliary quantities $\tau^*(x) := \beta_1 + \beta_2 x$ and $V_x := \mathbb{V}(X)$. The pseudo-VCATE is defined as

equation[equation omitted — 67 chars of source]

The pseudo-VCATE has a close connection to the VCATE. If the linear model describes the conditional mean $\mu_d(x)$, then $\tau^*(x) = \tau(x)$ and $V_{\tau} = V_{\tau}^*$. For instance, in models with binary $X$, the functional form is correctly specified and $V_{\tau} = V_{\tau}^*$. For now, assume that the pseudo-VCATE and the VCATE coincide. In later sections, I analyze models that allow for misspecification.

Consider a sequence of distributions $\{\gamma_n\}_{n=1}^\infty \in \Gamma^{\infty}$. I index the regression coefficients and model variances by the sample size $n$ as $\beta_{2n}$ and $(V_{xn}, V_{\tau n})$, respectively. Define an estimator of the VCATE as $\widehat{V}_{\tau n} = \widehat{\beta}_{2n}^2\widehat{V}_{xn}$, where $\widehat{\beta}_{2n}$ is the least squares estimator of (ref) and $\widehat{V}_{xn} :=\frac{1}{n}\sum_i X_i^2 -\left[\frac{1}{n}\sum_i X_i \right]^2$. With some algebraic manipulations the estimation error can be decomposed as

equation[equation omitted — 342 chars of source]

To derive the asymptotic distribution we can apply the central limit theorem to individual components. For generality, I state joint convergence to a normal distribution as an assumption. This holds as a special case if the observations are i.i.d. and key moments of the distribution are bounded, but may also hold under other forms of dependence. I defer stating primitive conditions until Section (ref).

assumpThere is a sequence of distributions $\{\gamma_n\}_{n=1}^\infty \in \Gamma^{\infty}$ with associated quantities $\{V_{xn},V_{\tau n}^*,\beta_{2n},\Omega_n \}_{n=1}^{\infty}$, which are related by the identity $V_{\tau n}^* = \beta_{2n}^2V_{xn}$, and satisfy the following properties: (i) $V_{xn} > 0$, (ii) $V_{\tau n}^*$ is contained in a bounded subset of $[0,\infty)$, and (iii) $\Omega_n$ is a positive definite matrix with eigenvalues bounded way from zero and a finite upper bound. There is a sequence of estimators $\{\widehat{V}_{xn},\widehat{V}_{\tau n},\widehat{\beta}_{2n},\widehat{\Omega}_n \}_{n=1}^{\infty}$ which satisfy $\widehat{V}_{\tau n} = \widehat{\beta}_{2n}^2\widehat{V}_{xn}$. As $n \to \infty$, $\widehat{\Omega}_n \to^p \Omega_n$, and \begin{equation} \Omega_n^{-1/2}\sqrt{n}\begin{pmatrix} \sqrt{V_{xn}}(\widehat{\beta}_{2n} - \beta_{2n}) \\ \frac{\widehat{V}_{xn}}{V_{xn}} - 1 \end{pmatrix} \to^d Z_n \sim \mathcal{N}(0,I_{2 \times 2}). \end{equation}

The normalization by $V_{xn}$ is intended to align with the decomposition in (ref). The $2 \times 2$ matrix $\Omega_n$ is an estimator of the covariance matrix. I present Assumption (ref) as a triangular array because it makes it easier to formalize discussions of uniform coverage over the parameter space. Assumption (ref) allows for cases where $V_{\tau n}^*$ is arbitrarily close to or includes zero. Let $\Omega^{1/2}$ denote the Cholesky decomposition of a matrix $\Omega$. The estimator of the VCATE converges to the empirical process $G$, defined as

equation[equation omitted — 215 chars of source]

where $z \in \mathbb{R}^2$, $\zeta \in \{-1,1\}$, $e_1 = [1,0]'$ and $e_2 = [0,1]'$.

lemSuppose that Assumption (ref) holds, then $\widehat{V}_{\tau n}-V_{\tau n}^* = O_p\left( \max\left\{\frac{1}{n},\sqrt{\frac{V_{\tau n}^*}{n}}\right\} \right)$, and there exists a sequence of $\zeta_n \in \{-1,1\}$, such that \begin{equation} \widehat{V}_{\tau n}-V_{\tau n}^* = G(n,V_{\tau n}^*,\Omega_n,Z_n,\zeta_n) + o_p\left(\frac{1}{n}\right) + o_p\left(\sqrt{\frac{V_{\tau n}^*}{n}}\right)+ o_p\left(\frac{V_{\tau n}^*}{\sqrt{n }}\right). \end{equation}

Lemma (ref) shows that the limiting distribution of $\widehat{V}_{\tau n}$ is a linear combination of a Chi-square and a normal, whose weights depend on the value of $V_{\tau n}^*$. The relative magnitude of $V_{\tau n}^*$ determines the fit of the normal approximation. In the heterogeneous case, $V_{\tau n}^* \ge \delta > 0$, $\sqrt{n}G$ converges to a normal as $n \to \infty$ because the first term in (ref) is asymptotically negligible. However, when $V_{\tau n}^* = 0$, only the first term remains and $nG$ converges to a non-central Chi-Square distribution, which is asymmetric. Using normal critical values here (even if everything else was known) would produce distorted coverage. Furthermore, when $V_{\tau n}^* = 0$ the rate of convergence is $n$, which is faster than $\sqrt{n}$, and hence the estimator is “super consistent” near the boundary. The error is dominated by the first stage sampling uncertainty in estimating the nuisance parameter $\beta_{2n}^2$, which converges at $n$ rate.

In practice, all three components in (ref) contribute to the limiting distribution, and this information can be used for inference. I propose an analytic approach based on the quantiles of the empirical process that can deliver exact coverage. Let $F_{n,V_{\tau}^*,\Omega,\zeta}(v)$ be the conditional CDF of the empirical process, defined as

equation[equation omitted — 171 chars of source]

Based on this CDF we can construct a test statistic, $$ F_{n,V_{\tau}^*,\widehat{\Omega}_{n},\zeta}(\widehat{V}_{\tau n}-V_{\tau}^*), $$ indexed by unknown values of $(V_{\tau}^*,\zeta)$ and substituting the estimated covariance matrix $\widehat{\Omega}_{n}$. By construction, the test statistic is contained in $[0,1]$. Similarly, I construct critical values as functions of the parameters for a nominal level $\alpha$, as follows

equation[equation omitted — 306 chars of source]

The difference in the critical values is $(1-\alpha)$ to achieve the desired coverage. The lower critical value is the minimum of the $\alpha/2$ percentile and $0$. This adjustment is meant to increase the power of tests of homogeneity (see Remark (ref)). I propose an adaptive confidence interval by substituting the $\widehat{\Omega}_{n}$, $n$, and $\widehat{V}_{\tau n}$ into the following formula

align[align omitted — 375 chars of source]

The set $\widehat{CI}_{\alpha n}$ can be constructed via a grid search between $0$ and an arbitrarily high value, to test whether a particular $V_{\tau}^*$ satisfies the inequality constraints. The procedure achieves correct asymptotic size because the test statistic converges to a uniform random variable in $[0,1]$ for each value of $V_{\tau}^*$. In general, the distribution in (ref) depends on the value of $\zeta$ and I obtain a conservative interval in (ref) by considering the union of intervals with different values of $\zeta$. Moreover, if the off-diagonal element of $\Omega_n$ is zero, then the distribution of the empirical process in (ref) does not depend on the value of $\zeta$. This property is plausible and I introduce primitive conditions that satisfy it in Section (ref). Under those conditions the confidence interval has exact asymptotic coverage .

The procedure is fast because at each point in the grid the researcher evaluates the condition in (ref), using the same estimate of $(\widehat{V}_{\tau n},\widehat{\Omega}_n)$. The critical values can be computed numerically from the quantiles of a generalized Chi-square with distribution $F$, which are available in most statistical software packages.

rem[Equivalence of homogeneity test, $\beta_2 = 0$] Researchers can test for homogeneity by evaluating whether $0 \in \widehat{CI}_{\alpha n}$. By definition, $e_1'\widehat{\Omega}_n^{1/2}Z = \sqrt{\Omega_{n,11}}Z_1$, where $\Omega_{n,11}$ is the upper-left entry. Under the null, the test statistic is $F_{n,0,\widehat{\Omega}_n,\zeta}(v) = \mathbb{P}(\widehat{\Omega}_{n,11}Z_1^2/n \le v )$, which is the CDF of a rescaled Chi-square distribution with one degree of freedom. Furthermore, the critical values are $\{0,1-\alpha\}$, given that $F_{n,0,\widehat{\Omega}_n,\zeta}(0) = 0$. Neither quantity depends on the choice of $\zeta$. Because of the normalization in (ref), we can choose $\widehat{\Omega}_{n,11} = \widehat{V}_{xn} \widehat{\mathbb{V}}(\widehat{\beta}_{2n})$, where $\widehat{\mathbb{V}}(\widehat{\beta}_{2n})$ is an estimate of the asymptotic variance of $\widehat{\beta}_{2n}$ such as the robust sandwich estimator. Therefore, evaluating $F_{n,0,\widehat{\Omega}_n,\zeta}(\widehat{V}_{\tau n}) \in [0,1-\alpha]$ is algebraically equivalent to a test of whether $n(\widehat{\beta}_{2n}^2\widehat{V}_{xn}/(\widehat{V}_{x n}\widehat{\mathbb{V}}(\widehat{\beta}_{2n})) = n\widehat{\beta}_{2n}^2/\widehat{\mathbb{V}}(\widehat{\beta}_{2n})$ exceeds the $1-\alpha$ quantile of a Chi-square with one degree of freedom. This is identical to a test of $\beta_2 = 0$ in the regression in (ref).
rem[Adjusting critical values] The critical value $q_{\alpha/2}(V_{\tau}^*,\Omega,n,\zeta)$ in (ref) is constructed to guarantee that $\widehat{V}_{\tau n} \in \widehat{CI}_{\alpha n}$. In this case, $\widehat{V}_{\tau n}$ belongs to the CI if and only if $F_{n,\widehat{V}_{\tau n},\widehat{\Omega}_{n},\zeta}(0)$ is contained in the critical region for some $\zeta \in \{-1,1\}$. The unadjusted CI with critical values $\left\{ \alpha/2, 1-\alpha/2\right\}$ is not guaranteed to contain the test statistic.\footnote{For example, suppose that $\widehat{V}_{\tau n} = 0$. Then the empirical process has a Chi-square distribution for $V_{\tau} = 0$. Since the unadjusted critical value is bounded away from zero, $\widehat{V}_{\tau n}$ would not be contained in the unadjusted CI.} Another rationale for doing the adjustment in (ref), is to increase the power of the test of homogeneity, $0 \in \widehat{CI}_{\alpha n}$, relative to a test based on the unadjusted CI. The unadjusted test has correct size but the rejection region is discontinuous: it rejects when the test statistics is very close to zero or when it exceeds a threshold. Instead, the adjusted test shifts the critical region left and has the form of a Chi-squared test. It only rejects the null if the test statistic is larger than $1-\alpha$, which is a threshold that is smaller than $1-\alpha/2$ for the unadjusted CI.
rem[Comparison to other tests of homogeneity] crump2008nonparametric suggest estimating $\widehat{\mu}_d(x)$ by a series estimator with $K$ terms, for subsamples $D = d \in \{0,1\}$. They propose a bias-corrected Wald statistic, which takes the form $\widehat{T}_n^{series} := \{[\widehat{\xi}_1 - \widehat{\xi}_0 ]'[\widehat{\mathbb{V}}(\widehat{\xi}_1-\widehat{\xi}_0)]^{-1}[\widehat{\xi}_1 - \widehat{\xi}_0]-(K-1)\}/\sqrt{2(K-1)}$, where $(\widehat{\xi}_1,\widehat{\xi}_0)$ are non-intercept coefficients associated with $\widehat{\mu}_1(x)$ and $\widehat{\mu}_0(x)$, respectively and $K$ is the number of covariates. For regressions with univariate $X$ as in (ref), $K=2$ and $\widehat{T}_n^{series} = (n\widehat{\beta}_{2n}^2/\widehat{\mathbb{V}}(\widehat{\beta}_{2n}) - 1) / \sqrt{2}$. Essentially this is just a transformation of the test statistic proposed above, which will produce the same acceptance/rejection result for significance level $\alpha$ (using the critical values in their equation 3.11). ding2019decomposing study a framework with a fixed population where the only source of randomness is the experimental assignment of offers. They propose a similar Wald estimator, but replace estimates of $(\widehat{\xi}_1,\widehat{\xi}_0)$ and the asymptotic variance with randomization inference counterparts. In samples with large $n$, this leads to very similar test statistics, but may produce slightly different results in small samples. The approach that I introduce in the following section differs substantially in the way that I handle multivariate cases. For $K > 2$, the approaches are non-nested because I use sample splitting and consider a wider range of methods to estimate $\widehat{\mu}_d(x)$ than series estimators.

Inference for nonparametric CATE

In this section I provide an overview of the inference problems associated with efficient estimators of the VCATE and how to solve them for the nonlinear/high-dimensional case. Let $\{Y_i,D_i,X_i\}$ be i.i.d.. As shown in newey1990semiparametric, efficient estimators can be decomposed as

equation[equation omitted — 193 chars of source]

where $\varphi_i$ is an i.i.d. realization from the efficient influence function with mean zero, and the residual becomes asymptotically negligible as $n \to \infty$. The semiparametric lower bound is $\mathbb{V}(\varphi_i)$. Let $\eta(\cdot)$ be a set of nuisance functions defined as

equation[equation omitted — 96 chars of source]

levy2021fundamental showed that the efficient influence function for the VCATE is equal to $\varphi_i = \varphi(Y_i,D_i,X_i,\eta) - V_{\tau n}$, where $\varphi$ is defined as

equation[equation omitted — 200 chars of source]

By (ref), all efficient estimators --regardless of their form-- are $\sqrt{n}$--asymptotically equivalent to $\frac{1}{n}\sum_{i=1}^n\varphi_i$. Let $\varphi_{i} = \varphi(Y_i,D_i,X_i,\eta(X_i))$ be a realization of the efficient influence function. If $\mathbb{V}(\varphi_i) > 0$, then $$\mathbb{V}(\varphi_i)^{-1}\times \sqrt{n}(\widehat{V}_{\tau n}-V_{\tau n}) = \mathbb{V}(\varphi_i)^{-1}\left[\frac{1}{\sqrt{n}}\sum_{i=1}^n \varphi_i \right]+ o_p(1) \to^p \mathcal{N}(0,1). $$ In this case, any confidence interval based on normal critical values and a consistent estimator of $\mathbb{V}(\varphi_i)$, produces valid coverage. For common functionals such as the ATE, $\mathbb{V}(\varphi_i) > 0$. However, this cannot be guaranteed for the VCATE.

lemLet $\sigma_d^2(x) := \mathbb{V}(Y \mid D =d,X=x)$. \begin{equation} \mathbb{V}(\varphi_i) = \mathbb{V}((\tau(X)-\tau_{av})^2) + 4\mathbb{E}\left[(\tau(X)-\tau_{av})^2\left(\frac{\sigma_1^2(X)}{p(X)} + \frac{\sigma_0^2(X)}{p(X)}\right) \right]. \end{equation}

Both of the inner terms in (ref) are multiplied by $(\tau(x) - \tau)$. When the VCATE is zero, $\tau(x) = \tau$ almost surely, and the influence function is degenerate. The condition that $\mathbb{V}(\varphi_i) > 0$ does not hold uniformly over all $V_{\tau}$ in the parameter space. In this case, the distribution of $\sqrt{n}(\widehat{V}_{\tau n}-V_{\tau n})$ is dominated by the higher order terms of the residual (ref), and the CLT cannot be applied to guarantee normality near the boundary. The linear estimator discussed in the previous section is just one example. Moreover, if the tails of $\tau(x)$ are thin, then the value of $\mathbb{V}(\varphi_i)$ is also small near the boundary.

corIf $\mathbb{E}[(\tau(X)-\tau)^4] \le \kappa^2 V_{\tau}^2$ for $\kappa \in \mathbb{R}_+$, then $$ \mathbb{V}(\varphi_i) \le \kappa^2 V_{\tau}^2+ 4\kappa V_{\tau}\sqrt{\mathbb{E}\left[ \left( \frac{\sigma_1^2(X)}{p(X)}+\frac{\sigma_0^2(X)}{1-p(X)} \right) \right]}. $$

Corollary (ref) shows that the variance of the efficient influence function is bounded by a quantity that scales up or down proportional to the value of $V_{\tau}$. Consequently, when $\mathbb{V}(\varphi_i)$ is relatively small, the higher order terms in the residual may still dominate.

Pseudo-VCATE, regressions, and efficiency

A robust way to introduce nonlinearity is to consider a regression with real-valued basis functions $M(x)$ and $S(x)$. For now, I will leave these unspecified but in the next section I will show how they can be estimated non-parametrically in a first stage.

equation[equation omitted — 130 chars of source]

with weights $\lambda(X) := [p(X)(1-p(X))]^{-1}$, regressors $W(X,D)$, and parameters $\theta = [c_0,c_1,\beta_1,\beta_2]$. This specification accommodates experiments with heterogeneous assignment probabilities.\footnote{When $p(x) = 1/2$ this produces exactly the same coefficients as a regression of $Y$ on $(1,M(X),D,D\times S(X))$, but differs when the probabilities are heterogeneous. If $M(X) = S(X) = X$ as well, this reduces to (ref).} Consider the following minimizers:

equation[equation omitted — 130 chars of source]

chernozhukov2020generic showed that if $\mathbb{V}(S(X)) > 0 $, $\mathbb{E}[S(X)] = 0$, and the vector $(\beta_1,\beta_2)$ are part of the solution to (ref), then $(\beta_1,\beta_2)$ are also the intercept and slope of the best linear projection of $\tau(X)$ on $S(X)$.\footnote{When $\mathbb{V}(S(X)) = 0$, $\beta_2$ does not have a unique solution in (ref), but $V_{\tau}^* = \beta_2^2\mathbb{V}(S(X)) = 0 \le V_{\tau}$ is still the best linear projection, regardless of the value of $\beta_2$.} Hence the pseudo-VCATE has an upper bound, $V_{\tau}^* = \beta_2^2\mathbb{V}(S(X)) \le V_{\tau}$. Because of this bound, if $V_{\tau} = 0$, then the pseudo-VCATE $(V_{\tau}^*)$ will not falsely detect heterogeneity even if $S(X)$ is misspecified. If anything, poor choices of $S(X)$ will possibly understate the amount of heterogeneity. When $\tau(X)$ is spanned by $S(X)$, $V_{\tau}^* = V_{\tau}$ and the two notions coincide.

To obtain a feasible estimator we define $\widehat{S}(x) := S(x) - \frac{1}{n}\sum_{i=1}^n S(X_i)$, and compute $\widehat{W}(x,d)$ by substituting $\widehat{S}(x)$ in (ref). Now consider a value of $\widehat{\theta}_n$ that minimizes $\frac{1}{n}\sum_{i=1}^n \lambda(X_i)(Y_i-\widehat{W}(X_i,D_i)'\theta)^2$, by solving the first order condition

equation[equation omitted — 180 chars of source]

The regression parameters can be used to construct the CATE and other nuisance functions. For a given $\theta \in \mathbb{R}^4$,

equation[equation omitted — 375 chars of source]

Lemma (ref) shows that the sample variance of the estimated CATE can be interpreted as an estimator that plugs-in (ref) to the efficient influence function in (ref).

lemDefine $\widehat{V}_{xn} = \frac{1}{n}\sum_{i=1}^n\widehat{S}(X_i)^2$. Let $\varphi$, $\tilde{\eta}_\theta$ and $\mathcal{Q}(\theta)$ be defined as in (ref), (ref), and (ref), respectively. If $\widehat{\theta}_n = (\widehat{c}_{1n},\widehat{c}_{2n},\widehat{\beta}_{1n},\widehat{\beta}_{2n})$ solves $\mathcal{Q}(\widehat{\theta}_n) = 0$, then $\widehat{V}_{\tau n} := \widehat{\beta}_{2n}^2\widehat{V}_{xn} = \frac{1}{n}\sum_{i=1}^n \varphi(Y_i,D_i,X_i,\tilde{\eta}_{\widehat{\theta}_n})$.

Mechanically, the influence function can be decomposed into primary and bias-correction components. As an intermediate step for Lemma (ref), I show that the bias-correction terms and the fourth component of (ref) are proportional to each other. The optimal $\widehat{\theta}_n$ implicitly sets the average bias correction to zero. Intuitively, the linear model minimizes the covariate imbalances between the treatment and control group in-sample. Lemma (ref) suggests that $\widehat{V}_{\tau n}$ could be asymptotically efficient if $\tilde{\eta}_{\widehat{\theta}_n}$ is sufficiently close to $\eta$. In Section (ref), I show that my proposed semiparametric estimator can indeed achieve this.

As a preliminary step, it is necessary to determine which $S(x)$ and $M(x)$ ensure that $\tilde{\eta}_\theta = \eta$. Not all choices achieve this property.\footnote{This point highlights that while all regressions of the form in (ref) proposed by chernozhukov2020generic estimate an interpretable $V_{\tau}^*$ --regardless of the choice of $M(x)$--, not every regression in this class is efficient. The functions $S(x)$, and $M(x)$ in particular, both affect efficiency.} However, if they are chosen in such a way that $W(x,d)'\theta = \mu_d(x)$ for some $\theta \in \mathbb{R}^4$, then that's sufficient to guarantee that $\eta_\theta = \eta$. Lemma (ref) shows that any $\theta$ with this property is also a solution to the regression problem, and provides guidance on the choice of $S(x)$ and $M(x)$.

lemLet $ \Theta^*$ be the optimizer set defined in (ref). If (i) $\mathbb{E}[S(X)] = 0$ and (ii) $W(x,d)'\theta = \mu_d(x)$ for some $\theta \in \mathbb{R}^4$, then $\theta \in \Theta^*$. Conditions (i) can be satisfied by setting $S(x) = \tau(x) - \mathbb{E}[\tau]$. Condition (ii) can be satisfied by setting $M(x) = \mu_0(x) + p(x)\tau(x)$. In this special case, $\theta = (0,1,\mathbb{E}[\tau(X)],1)' \in \Theta^*$.

Lemma (ref) provides efficient choices of $S(x)$ and $M(x)$ that can be expressed in terms of conditional moments, and that for this choice, the optimal $\theta$ has a known, simple form. In practice, $S(x)$ and $M(x)$ can be estimated non-parametrically.

Multi-step approach

My proposed procedure randomly partitions the observations $\mathcal{I}_n := \{1,\ldots,n\}$ into $K$ folds of equal size $n_k := n/K$. Denote the observations in each fold by $\mathcal{I}_{nk}$, so that $\bigcup_{k=1}^{K} \mathcal{I}_{nk} = \mathcal{I}_n$, and let $\mathcal{I}_{-nk} := \mathcal{I}_n \backslash \mathcal{I}_{nk}$ be the set of observations that are not in fold $k$. In a slight abuse of notation, I use $\mathcal{I}_{-nk}$ when defining conditional expectations, to denote the full set of random variables associated with observations not included in fold $k$. For simplicity, I also label the fold of observation $i \in \mathcal{I}_{nk}$, by $k_i$.

Let $\widehat{\eta}_{-k}(x) := (\widehat{\tau}_{-k}(x),\widehat{\mu}_{0,-k}(x), p(x),\widehat{\tau}_{-k,av})$ denote a prediction of the nuisance function $\eta(x)$ over the set $\mathcal{I}_{-nk}$, using the researcher's preferred prediction algorithm. This could include traditional methods such as linear regression, or more modern “machine learning” approaches such as LASSO, neural networks, or random forests. The only function that is known in advance is the propensity score, since I restrict attention to randomized experiments. Guided by Lemma (ref), define

align[align omitted — 393 chars of source]

Consider a regression with weights $\lambda(X_i)$, parameters $\theta := (c_1,c_2,\beta_1,\beta_2)$, and

equation[equation omitted — 127 chars of source]

In practice, $\mathbb{E}[\widehat{\tau}_{-k}(X_i) \mid \mathcal{I}_{-nk}]$ needs to be estimated, and I use a sample analog:

align[align omitted — 554 chars of source]

Let $\widehat{\theta}_{nk} = (\widehat{c}_{1nk},\widehat{c}_{2nk},\widehat{\beta}_{1nk},\widehat{\beta}_{2nk})$ be the estimator over the subsample $\mathcal{I}_{nk}$. The fold-specific variance of $\widehat{S}_{-k_i}(x) $ is defined as

equation[equation omitted — 127 chars of source]

The estimator of the VCATE for fold $k$ is

equation[equation omitted — 107 chars of source]

In this case $\widehat{V}_{xnk}$ can be viewed as a preliminary estimate of the VCATE using the data in $\mathcal{I}_{-nk}$, whereas $\widehat{V}_{\tau nk}$ is a regression-adjusted estimator that fits the sample $\mathcal{I}_{nk}$. This adjustment will produce better results, with a pseudo-VCATE interpretation even if the first step $\widehat{S}_{-k_i}(x)$ function is noisy, misspecified, or slow to converge to $\tau(x)$. The estimator in (ref) belongs to the class of multi-step estimators defined in chernozhukov2020generic. I add a restriction on the choice of $M_{-k}(x)$, guided by Lemma (ref), to ensure asymptotic efficiency.

To quantify the uncertainty in $(\widehat{\theta}_{nk},\widehat{V}_{xnk})$ I compute a robust (sandwich) estimator. I start by defining two auxiliary residuals, $\widehat{T}_i := \widehat{V}_{x n k}^{-1}\widehat{S}_{-k_i}(X_i)^2 - 1$ and $\widehat{U}_i := Y_i - \widehat{W}_i'\widehat{\theta}_{nk}$. Let $\widehat{\Pi}_{nk}$ be a $4 \times 4$ diagonal matrix with diagonal entries $(1,1,1,\widehat{V}_{xnk}^{-1/2})$. Researchers can compute estimators of the individual components of the sandwich form $\widehat{J}_{nk}$, $\widehat{H}_{nk}$, and a selection matrix $\Upsilon$ defined as follows

equation[equation omitted — 322 chars of source]
equation[equation omitted — 460 chars of source]

The sandwich covariance estimator is

equation[equation omitted — 146 chars of source]

The population covariance matrix is

equation[equation omitted — 211 chars of source]

When the VCATE is zero, then $\widehat{V}_{xnk}$ (as a consistent estimator of $V_{\tau n}$) should converge to zero along the asymptotic sequence. To prevent asymptotic degeneracy, we need to rescale the estimands along the lines of Assumption (ref). The random variable $V_{xnk}^{-1/2}S_{-k}(X_i)$ is normalized to (conditionally) have variance one by design, even if $S_{-k}(X_i)$ converges to zero. This requires two much weaker conditions: (i) that $V_{xnk} > 0$, i.e. there is some noise in estimating the CATE;\footnote{I also propose an extension that allows for $V_{xnk} = 0$ in Remark (ref).} (ii) $\Omega_{nk}$ has eigenvalues bounded away from zero. To ensure this, the tails of $S_{-k}(X_i)$ need to be thin.\footnote{One sufficient additional restriction is that $\mathbb{E}[U_{i} \mid X_{i},D_{i},\mathcal{I}_{-nk}] = 0$ (the model is correctly specified), $\mathbb{V}(U_{i} \mid X_{i},D_{i} \mid \mathcal{I}_{-nk})$ is bounded away from zero, and $S_{-k}(X_i)$ has bounded kurtosis. In that case the off-diagonal elements of $\Omega_{nk}$ are zero and the diagonals are uniformly bounded. Positive-definiteness may also hold in a neighborhood where the nuisance functions are close to the true value and $\mathbb{E}[U_{i} \mid X_{i},D_{i},\mathcal{I}_{-nk}] \approx 0$.}

assump[Moment Bounds] Suppose that there exists a constant $\delta \in (0,1)$ such that for each fold $k$, almost surely, (i) $\mathbb{E}[M_{-k}(X_i)^2 \lambda(X_i) \mid \mathcal{I}_{-nk}] - \mathbb{E}[M_{-k}(X_i)\lambda(X_i)\mid \mathcal{I}_{-nk}]\mathbb{E}[\lambda(X_i) \mid \mathcal{I}_{-nk}] \ge \delta$, (ii) $\mathbb{E}[M_{-k}(X_i)^4 \mid \mathcal{I}_{-nk}] \le 1/\delta$, (iii) $\mathbb{E}[U_{i}^4 \mid \mathcal{I}_{-nk}] \le (1/\delta)$, and (iv) $\mathbb{E}[S_{-k}(X_i)^4 \mid \mathcal{I}_{-nk}] \le (1/\delta) \mathbb{E}[S_{-k}^2 \mid \mathcal{I}_{-nk}]^2$, (v) $\mathbb{E}[\widehat{\tau}_{-k}(X)^2 \mid \mathcal{I}_{-nk}] < 1/\delta$.

Assumption (ref).(i) is a rank condition that ensures that the auxiliary regressor $M_{-k}(X_i)$ is not degenerate. Assumption (ref).(ii) ensures that the second-moment of the candidate regressor $M_{-k}(X_i)$ is bounded. Assumption (ref).(iii) is a standard condition indicating that the fourth moment of the residuals are bounded. Assumption (ref).(iv) is a bounded kurtosis condition indicating that the out-of-sample, machine learning predictions of $\tau(x)$ have thin tails. Finally, Assumption (ref).(v) is a bound on the variance of the first-stage VCATE.

assump[Non-degeneracy] The following properties hold almost surely over sequences of random data realizations $\{\mathcal{I}_{n1},\ldots,\mathcal{I}_{nK}\}_{n=1}^{\infty}$. Conditional on $\mathcal{I}_{-nk}$: (i) $V_{xn k} > 0$, (ii) $V_{xnk}$ has a finite upper bound, (iii) $V_{\tau nk}^* := \beta_{2nk}^2V_{xnk}$ is contained in a bounded subset of $[0,\infty)$, and (iv) $\Omega_{nk}$ defined in (ref) is a positive definite matrix with bounded eigenvalues.
assump[Random Sampling] The observations $\{Y_{0i},Y_{1i},D_{i},X_{i}\}_{i}^n$ are i.i.d. across $i$ for fixed $n$, and drawn from a sequence of data generating processes $\{\gamma_n\}_{n=1}^{\infty}$.

Theorem (ref) shows how these primitive conditions imply an analog of Assumption (ref) for the cross-fitted case.

thmConsider a sequence of random data realizations $\{\mathcal{I}_{n1},\ldots,\mathcal{I}_{nK}\}_{n=1}^{\infty}$ with associated quantities $\{V_{xnk},V_{\tau n k}^*,\beta_{2nk},\Omega_{nk} \}_{n=1}^{\infty}$ for each $k$, as well as a sequence of estimators $\{\widehat{\beta}_{2nk},\widehat{V}_{xnk},\widehat{V}_{\tau nk}^*,\widehat{\Omega}_{nk} \}_{n=1}^{\infty}$ computed from (ref), (ref), (ref), and (ref), respectively. Suppose that these quantities satisfy Assumptions (ref).(ii), (ref), (ref), and (ref). Then as $n_k \to \infty$, for all $k \in \{1,\ldots, K\}$, Conditional on a sequence of $\mathcal{I}_{-nk}$, \begin{enumerate}[(i)] {0pt} • \quad $ \Omega_{nk}^{-1/2}\sqrt{n_k}\begin{pmatrix} \sqrt{V_{xnk}}(\widehat{\beta}_{2nk} - \beta_{2nk}) \\ \frac{\widehat{V}_{xnk}}{V_{xnk}} - 1 \end{pmatrix}\mid \mathcal{I}_{-nk} \to Z_{nk} \sim \mathcal{N}(0,I_{2 \times 2}).$\quad $\widehat{\Omega}_{nk} \to^p \Omega_{nk}$. \end{enumerate}

Theorem (ref) presents a central limit theorem for the components of $\widehat{V}_{\tau nk}^* = \widehat{\beta}_{2nk}^2\widehat{V}_{xnk}$, properly rescaled and conditional on $\mathcal{I}_{-nk}$. This result holds regardless of whether the nuisance parameters are properly specified and primarily relies on the independence of the folds. By Lemma (ref), conditional on $\mathcal{I}_{-nk}$,

equation[equation omitted — 282 chars of source]

Then it is possible to construct adaptive confidence intervals, substituting the sample size $n_k$ and estimated statistics $(\widehat{V}_{\tau nk},\widehat{\Omega}_{nk})$.

align[align omitted — 399 chars of source]

The confidence intervals take the same form as in the regression case in (ref), except that now the inputs are obtained from the cross-fitted regression step. The confidence interval is fast to compute because $(\widehat{V}_{\tau nk},\widehat{\Omega}_{nk})$ only needs to be computed once. It is worth noting that because the confidence interval only uses information in fold $\mathcal{I}_{nk}$, the effective sample size is $n_k$. While this does not affect the nominal asymptotic size of the confidence interval, it may affect the power of tests against specific alternatives.

Ensemble estimator

We can construct an “ensemble” to aggregate across folds, defined as follows

equation[equation omitted — 139 chars of source]

In Section (ref), I show that this ensemble estimator is efficient.

Splitting uncertainty and median intervals

So far in this section we have used the data from a single split or fold of the data. However, the choice of fold $k$ or the particular split may lead to different values of $\widehat{V}_{\tau nk}$ and hence distinct confidence intervals. chernozhukov2020generic propose an aggregation procedure based on “median parameter” confidence intervals, inspired by false-discovery rate adjustments. Their proposed conditional t-tests are not directly applicable here because $\widehat{V}_{\tau nk}$ conditionally converges to a generalized Chi-square. However, I show that the basic idea can still be adapted.

Let $K$ be the total number of folds, obtained across one or more splits of the data. For instance, a 2-fold sample with 10 splits would have $K=20$, Let $\inf \widehat{CI}_{\alpha nk}$ and $\sup \widehat{CI}_{\alpha nk}$ denote the lower and upper bounds of $\widehat{CI}_{\alpha nk}$, respectively, and $\text{Med}_K\{\cdots\}$ denote the median over a set indexed by $k = \{1,\ldots,K\}$. If $K$ is even then two quantities might be tied for the median, and in that case I compute their midpoint. The multifold confidence interval is defined as

equation[equation omitted — 245 chars of source]

Intuitively, the $K$ fold-specific intervals “vote” to include a particular value, and $V_{\tau}^* \in \widehat{CI}_{\alpha n}^{\text{multifold}}$ only if there is a majority vote. The “median” interval $\widehat{CI}_{\alpha n}^{\text{multifold}}$ contains values within the median lower bound and the median upper bound across folds. To control the overall false discovery rate, I adjust the nominal size to $\alpha/2$. This adjustment produces a conservative interval because it assumes a worst-case dependence structure between the folds and the splits, regardless of the size of $K$. In some instances, the asymptotic coverage probability may be strictly higher than $(1-\alpha)$, particularly when there is a lot of heterogeneity.\footnote{For instance, given a single split, Theorem (ref) implies that the $\{\sqrt{n_k}(\widehat{V}_{\tau nk}-V_\tau)\}_{k=1}^K$ are asymptotically uncorrelated. However, near the boundary, the estimators converge at a rate faster than $\sqrt{n}$ and their relative dependence structure at that rate is unclear.} At the boundary, with low effect heterogeneity or none at all, it is much harder to asses the dependence structure between the fold-specific estimators. One of the benefits of using a worst-case approach is that it provides coverage guarantees under weak assumptions. Moreover, the empirical example illustrates that even though these intervals are conservative, they may have a short length in practice.

Large Sample Theory

$\sqrt{n}$-Consistency, Efficiency, and Boundary Rates

Let $\gamma \in \Gamma$ denote a probability distribution over i.i.d observations $(Y_{1i},Y_{0i},D_i,X_i)$. I use the notation $\mathbb{E}_{\gamma}[\cdot]$ and $\mathbb{P}_{\gamma}(\cdot)$ to denote the expectation and probability under $\gamma$, respectively. Let $S_{-k}(x)$ be the function defined in (ref). The true value of the CATE and VCATE is given by $\tau_\gamma$ and $V_{\tau}(\gamma)$, respectively. The pseudo-VCATE is given by

equation[equation omitted — 237 chars of source]

Define the estimation error of the CATE in the $L_2$ norm as

equation[equation omitted — 125 chars of source]

We can bound the difference between the pseudo-VCATE and its true value:

thm[Bias of the pseudo-VCATE] Under the distribution $\gamma \in \Gamma$, \begin{equation} \mathbb{E}_\gamma[| V_{\tau}(\gamma)-V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) |] \le \min\left\{16 \times \omega(\gamma)^2 ,V_{\tau}(\gamma)\right\}. \end{equation}

Theorem (ref) derives a non-asymptotic bound for the VCATE as the minimum of two key quantities: (i) the conditional $L_2$ error between the candidate function and the true CATE, and (ii) the true value of the VCATE. This proof only relies on the definition in (ref). For instance, when $V_{\tau}(\gamma) = 0$, then $V_{\tau}(\gamma)-V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) = 0$, regardless of whether $\widehat{\tau}_{-k}(\cdot)$ is properly specified. The difference between the two quantities is also small if $\omega(\gamma)$ is sufficiently close to zero. In the multi-step approach, $\omega(\gamma)$ captures the first-stage uncertainty from estimating the CATE, which decreases with sample size. I consider the following convergence condition.

assump[Convergence CATE] $\sqrt{n_k}\omega(\gamma_n)^2 = o(1)$ as $n \to \infty$.

Assumption (ref) imposes an $L_2$ consistency condition on the CATE. A large class of machine learning models can meet this requirement. For example, bickel2009simultaneous and belloni2014inference evaluate rates of convergence under sparse models, chen1999improved for neural networks, and wager2015adaptive for regression trees and random forest.

thm[Faster than $\sqrt{n}$ convergence near boundary] Consider a sequence of data generating processes $\{\gamma_n\}_{n=1}^{\infty}$ where $V_{\tau}(\gamma_n) \to 0$ as $n \to \infty$ and Assumptions (ref), (ref), (ref), (ref) and (ref) hold. Define $\Delta_{nk} := \left( \widehat{V}_{\tau n k} - V_{\tau}(\gamma_n) \right)$ for the estimator defined in (ref). Then (i) $\sqrt{n_k}\Delta_{nk} = o_p(1)$, and (ii) if in addition $n_k^{1/2+\rho}V_{\tau}(\gamma_n)=o(1) $ for $\rho \in [0,1/2)$, then $n_k^{1/2+\rho}\Delta_{nk} = o_p(1)$.

Theorem (ref) shows that multi-step estimators of the VCATE converge to zero faster than $\sqrt{n_k}$ near the boundary. I formalize “near” by considering sequences of distributions where the VCATE approaches zero. Theorem (ref) relies on the non-asymptotic bound in Theorem (ref), the normal approximation in Theorem (ref), and the empirical process in Lemma (ref). There is no requirement on the rate of convergence of $\widehat{\mu}_{-0k}(\cdot)$ (and consequently on the generated regressor $M_{-k_i}(\cdot)$), only an assumption that $p(x)$ is known and that the CATE is estimated at a sufficiently fast rate. Furthermore, if the true CATE is nearly flat in the sense that for $\rho \in [0,1/2)$, then $n_k^{1/2+\rho}V_{\tau}(\gamma_n) = o(1)$ (or even exactly equal to zero), then the estimator has a faster rate guarantee.

To prove efficiency we have the stronger requirement that all the nuisance functions converge to their true value in the $L_4$ norm and at $n_k^{1/4}$ rate in the $L_2$ norm.

assump[Regularity conditions] Define the residuals $U_i = Y_i - \mathbb{E}_{\gamma_n}[Y_i \mid D_i,X_i]$. (i) $\mathbb{E}_{\gamma_n}[\Vert Y_i \Vert^4]$, $\mathbb{E}_{\gamma_n}[\Vert U_i\Vert^4]$, $\mathbb{E}_{\gamma_n}[\Vert \eta(X_i) \Vert^4]$, (ii) $\mathbb{E}_{\gamma_n}[\Vert \widehat{\eta}_{-k}(X_i) \Vert^4]$ are uniformly bounded, (iii) $\mathbb{E}_{\gamma_n}\left[ \Vert \widehat{\eta}_{-k}(X_) - \eta(X_i) \Vert^4\right] \to 0$, (iv) $\sqrt{n_k}\mathbb{E}_{\gamma_n}[\Vert\widehat{\eta}_{-k}(X_i)-\eta(X_i)\Vert^2] = o(1)$ for all $k \in \{1,\ldots,K\}$.

The next step is to show that the estimation error of the fold-specific VCATE converges at $\sqrt{n_k}$ to an average of efficient influence functions.

thm[$\sqrt{n}$ Consistency and Efficiency] Consider a sequence of data generating processes $\{\gamma_n\}_{n=1}^{\infty}$ where $V_{\tau}(\gamma_n) \to 0$ as $n \to \infty$ and Assumptions (ref), (ref), (ref), (ref), (ref), and (ref) hold. Then $$ \sqrt{n_k}(\widehat{V}_{\tau nk}-V_{\tau}(\gamma_n)) = \frac{1}{\sqrt{n_k}}\sum_{i \in \mathcal{I}_{nk}} \varphi_i + o_p(1). $$

Theorem (ref) shows that the fold-specific estimator converges at $\sqrt{n_k}$-rate to an average of i.i.d influence function. This requires standard regularity conditions. The proof of Theorem (ref) is non-standard due to the multi-step nature of the procedure. I start by applying Lemma (ref), which shows hows to write $\widehat{V}_{\tau nk}$ as an average of estimated influence functions. I break down the proof into sequences where $V_{\tau}(\gamma_n)$ converges to zero and those where it's bounded away from zero. For the first part, I leverage (a) the boundary convergence result in Theorem (ref), (b) the bound for $\mathbb{V}(\varphi_i)$ in Lemma (ref). For the second part, I provide a novel decomposition of regression adjusted nuisance functions. The key is to prove that the regression parameters $\widehat{\theta}_{nk}$ converge at $n_k^{1/4}$ rate to the values in Lemma (ref) for sequences where $V_{\tau}(\gamma_n) \to V_{\tau} > 0$. Once in this form, the rest of the proof relies on a traditional Taylor expansion argument.

The ensemble estimator $\widehat{V}_{\tau n}$ combines information from the whole sample. By definition $n = n_k \times K$ and $K$ is finite, which means that algebraically $\sqrt{n}(\widehat{V}_{\tau n}-V_{\tau}(\gamma_n)) = \frac{\sqrt{n_k}}{\sqrt{K}}\sum_{k=1}^{K}(\widehat{V}_{\tau nk}-V_{\tau}(\gamma_n))$, and by Theorem (ref), $$ \sqrt{n}(\widehat{V}_{\tau n}-V_{\tau}(\gamma_n)) = \left[ \frac{1}{\sqrt{n_k K}}\sum_{k=1}^K\sum_{i \in \mathcal{I}_{nk}} \varphi_i + \frac{1}{\sqrt{K}}\sum_{k=1}^K o_p(1)\right] = \frac{1}{\sqrt{n}}\sum_{i=1}^n \varphi_i + o_p(1). $$ This means that aggregating the estimators restores full efficiency, satisfying the property described in (ref).

Asymptotic Coverage

I start by showing that the single fold confidence interval has uniform coverage for the pseudo-VCATE, and exact coverage under an additional assumption.

assump[Exact coverage condition] Let $\Omega_{nk,12}$ be the off-diagonal element of $\Omega_{nk}$. For each $t > 0$, $\underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}(\sqrt{V_{\tau}^*(\gamma,\mathcal{I}_{-nk})}|\Omega_{nk,12}| > t) = 0$.

Assumption (ref) states that the product of the pseudo-VCATE and the off-diagonal element of the limiting covariance matrix in (ref) needs to converge to zero uniformly.

thm[Uniform Coverage of Pseudo-VCATE] Let $\Gamma$ denote a set of distributions, constrained in such a way that Assumptions (ref), (ref), (ref), and (ref) hold. Let $\widehat{CI}_{\alpha n k} $ and $V_{\tau}^*(\gamma,\mathcal{I}_{-nk})$ be defined as in (ref) and (ref), respectively. Then \begin{equation} 1-\alpha \le \underset{n \to \infty}{\lim\inf}\inf_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \in \widehat{CI}_{\alpha n k} \right) \end{equation} If Assumption (ref) also holds, then \begin{equation} \underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \in \widehat{CI}_{\alpha n k} \right) \le 1-\alpha. \end{equation}

Theorem (ref) shows that the confidence intervals always have uniform coverage of the pseudo-VCATE of at least $(1-\alpha)$.\footnote{The theorem only uses Assumptions (ref), (ref), (ref), and (ref) to verify normality in Assumption (ref). A broad class of confidence intervals of the form in (ref) constructed from regression adjusted estimators will satisfy these uniformity properties.} The key is to prove that the confidence intervals yield coverage under arbitrary sequences of distributions, which includes cases where $V_{\tau n}^*$ is either equal to zero or approaches zero as $n \to \infty$. The proof builds on the approximation of Lemma (ref) and shows that for every sequence, the test statistic for a particular $\zeta_{n_k} \in \{-1,1\}$ converges to a uniform distribution. This sequential characterization suffices to apply generic results in andrews2020generic, which guarantee uniform coverage even in non standard cases like this one. Coverage over the pseudo-VCATE holds regardless of whether the nuisance functions are slow to converge or even misspecified.

The intervals are in general conservative because we're not plugging in the unknown $\zeta$, and instead define a robust confidence interval as the union of CIs with given $\zeta \in \{-1,1\}$. However, the key insight is that $\zeta$ only affects the coverage when the pseudo-VCATE is bounded away from zero. If Assumption (ref) holds, the value of $\zeta$ doesn't enter the asymptotic distribution of the estimator. I show that this condition holds automatically if the nuisance functions converge to their true value at a sufficiently fast rate.

lem[Verify Exact Coverage] Let $\Gamma$ denote a set of distributions that satisfy Assumptions (ref), (ref), (ref), (ref), (ref), and (ref). Then Assumption (ref) also holds.

As a special case, when the model is correctly specified, i.e. $W_i'\theta = \mu_d(x)$ for some $\theta \in \mathbb{R}^4$, then $\Omega_{nk,12} = 0$ by construction. Lemma (ref) states that we only need a model that is correctly specified asymptotically, given the rates in Assumptions (ref) and (ref). Then for non-boundary cases, $\Omega_{nk}$ converges to the population analog under correct specification. These conditions also imply point-wise coverage of the true VCATE.

thm[Pointwise, Exact Coverage of VCATE] Let $\Gamma$ denote a set of distributions that satisfy Assumptions (ref), (ref), (ref), (ref), (ref), and (ref). Then $$ \inf_{\gamma \in \Gamma} \underset{n \to \infty}{\lim\inf} \ \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) \in \widehat{CI}_{\alpha nk} \right) = \sup_{\gamma \in \Gamma} \underset{n \to \infty}{\lim\sup} \ \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) \in \widehat{CI}_{\alpha nk} \right) = 1-\alpha. $$

Theorem (ref) shows that if the nuisance functions converge at a sufficiently fast rate, then the proposed intervals achieve point-wise exact coverage. The confidence intervals provide correct size coverage for all regions of the parameter space, including $V_{\tau}(\gamma) = 0$.

Proving uniform coverage of the VCATE (rather than the pseudo-VCATE) is more challenging in the non-parametric case without much stronger conditions on the convergence rates of the nuisance functions. The lack of uniformity stems from a difficulty in controlling the ratio $\sqrt{n_k}\omega(\gamma_n) / \sqrt{V_{\tau}(\gamma_n)}$, which measures the relative error in estimating the CATE vs. the overall level of the VCATE. By the bound in (ref), this ratio is easy to control when $n_kV_{\tau}(\gamma_n) = o(1)$ (near homogeneity) or $V_{\tau}(\gamma_n) \to V_\tau > 0$ (strong heterogeneity). However, it is possible to construct sequences, e.g., $n_kV_{\tau}(\gamma_n) \to v > 0$, where $(\widehat{V}_{\tau nk} - V_{\tau}^*(\gamma_n,\mathcal{I}_{-nk}))$ converges to zero at a faster or comparable rate to the error of the pseudo-VCATE. There may be distortions in coverage in smaller samples. I illustrate this issue in the simulations.

rem[Uniform inference for one-sided tests] Uniform inference is only challenging for two-sided tests. If instead, the researcher is only interested in left-sided tests, then uniform inference is still possible. To do so, we can make explicit use of the inequality $V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \le V_{\tau}(\gamma)$. If $V_{\tau}(\gamma) < \inf_{V_{\tau}^*} \widehat{CI}_{\alpha nk}$ (the lower bound of the CI), then $V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \notin \widehat{CI}_{\alpha nk}$. Therefore, for all $\gamma \in \Gamma$, \begin{equation} \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) \ge \inf \widehat{CI}_{\alpha n k} \right) \le \mathbb{P}_{\gamma}\left( V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \notin \widehat{CI}_{\alpha n k} \right). \end{equation} I prove a weaker uniformity result for one-sided tests building on Theorem (ref). \begin{cor} If Assumptions (ref), (ref), (ref), and (ref) hold, then $$ \underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) < \inf \widehat{CI}_{\alpha n k} \right) \le \underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \notin \widehat{CI}_{\alpha n k} \right) \le \alpha. $$ \end{cor}

Corollary (ref) is empirically relevant for interpreting confidence intervals that do not include zero. It states that the asymptotic probability of having $V_{\tau}(\gamma) \in \left[0,\inf \widehat{CI}_{\alpha nk}\right)$ is uniformly less than $\alpha$. Tests of homogeneity belong to this class and therefore have the correct size when $V_\tau(\gamma) = 0$. Moreover, the result in Corollary (ref) is much stronger because it guarantees that a broader class of one-sided tests also has the correct size. It is important to emphasize that I do not impose any assumptions on rates of convergence of $(\widehat{\eta}_{-k}(x)-\eta(x))$, but only the inequality on the pseudo-VCATE. Consequently, while estimating $\mu(x)$ and $\tau(x)$ may be important for increasing the power of tests of homogeneity, it is not necessary for controlling their size.

Multifold Coverage

The multi-fold confidence interval covers the VCATE asymptotically.

thmLet $\Gamma$ be a set of distributions that satisfy Assumptions (ref), (ref), (ref), (ref). Then $$ \underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) < \inf \widehat{CI}_{\alpha n k}^{multifold} \right) \le \alpha. $$ If Assumptions (ref), and (ref) also hold, then $$ \sup_{\gamma \in \Gamma} \underset{n \to \infty}{\lim\sup} \ \mathbb{P}_{\gamma}\left(V_{\tau}(\gamma) \notin \widehat{CI}_{\alpha n}^{\text{multifold}}\right) \le \alpha.$$

The first part of Theorem (ref) shows that the multifold CI uniformly controls the size of one-sided tests. The second part shows that if the nuisance functions converge to their true value asymptotically, then the multifold confidence interval provides point-wise size-control for two-sided tests. Coverage of the true parameter will be weakly larger that $(1-\alpha)$ asymptotically.

Power

The test of homogeneity has power against local alternatives.

lemConsider a sequence of distributions $\{\gamma_n\}_{n=1}^\infty$ and $\{\mathcal{I}_{-nk}\}_{n=1}^{\infty}$, where $\Omega_{nk} \to \Omega_{\infty}$ and $n_kV_{\tau}^*(\gamma_n,\mathcal{I}_{-nk}) = v + o(1)$, for $v \in [0,\infty)$. Assume that (ref), (ref), (ref), and (ref) hold. Let $\Omega_{\infty,11}$ be the upper-left entry of $\Omega_{\infty}$, $\Phi(\cdot)$ be the standard normal CDF, and $z_{1-\alpha}$ be the $(1-\alpha)-$quantile. Then $$ \lim_{n \to \infty} \mathbb{P}_{\gamma_n}(0 \notin \widehat{CI}_{\alpha nk}\mid \mathcal{I}_{-nk}) = 1-\Phi\left(z_{1-\alpha}-\frac{\sqrt{v}}{\sqrt{\Omega_{\infty,11}}}\right) + \Phi\left( -\frac{\sqrt{v}}{\sqrt{\Omega_{\infty,11}}} \right). $$

Lemma (ref) computes the power curve for a sequence of local alternatives. When $v = 0$ the power is equal to $\alpha$, whereas when $v \to \infty$ the power tends to one. This shows that tests of homogeneity have local power the null. When the pseudo-VCATE is bounded away from zero, the test rejects with probability approaching one.

Extensions

rem[Clustered Standard Errors] In some cases, assuming that units $i$ are independent may be strong. For example, in dizon2019parents units are randomized at the household level, and it is reasonable to expects that units within a household have correlated outcomes and covariates. To deal with this dependence structure, suppose that the sample can be partitioned into $C$ clusters, $c \in \{1,\ldots,C\}$, which are independent and identically distributed. The researcher can compute $\widehat{\beta}_{2nk}$, $\widehat{V}_{\tau nk}$ and $\widehat{CI}_{\alpha nk}$ via cross-fitting by randomly partitioning entire clusters rather than the individual observations. \begin{lem} Let $\{r_{nk}\}_{n=1}^{\infty}$ be a sequence of positive scalars. Suppose that $V_{xnk} > 0$, $\Omega_{nk}$ is positive definite with positive eigenvalues, and that conditional on a sequence $\{\mathcal{I}_{-nk}\}_{n=1}^{\infty}$, $r_{nk}^{-1}\widehat{\Omega}_{nk} \to^p \Omega_{nk}$, $n_k/r_{nk} \to \infty$, and \begin{equation} \Omega_{nk}^{-1/2}\sqrt{\frac{n_k}{r_{nk}}}\begin{pmatrix} \sqrt{V_{xn}}(\widehat{\beta}_{2nk} - \beta_{2nk}) \\ \frac{\widehat{V}_{xnk}}{V_{xnk}} - 1 \end{pmatrix} \mid \mathcal{I}_{-nk} \to^d Z_n \sim \mathcal{N}(0,I_{2 \times 2}). \end{equation} Then $\widehat{CI}_{\alpha n k}$, substituting the arguments $(n_k,\widehat{V}_{\tau n k},\widehat{\Omega}_{nk})$, satisfies Theorem (ref). \end{lem} Lemma (ref) proposes high-level conditions that ensure that confidence intervals have correct coverage. The quantity $\sqrt{n_k/r_{nk}}$ is the effective rate of convergence, which features prominently in problems with cluster dependence mackinnon2022cluster. For example, if the observations are fully correlated within clusters and the clusters have equal size, then $r_{nk}$ is the cluster size, $n_k/r_{nk} = C$, and the estimators in (ref) converge at $\sqrt{C}$ rate (the total number of clusters). The analyst does not need to specify the quantity $r_{nk}$ to apply the procedure, but merely specify an estimator of the covariance matrix that meets the rate requirement. Under minor modifications to the existing proofs, we can also prove analogs of Theorems (ref) and (ref). We can construct estimators that satisfy Lemma (ref). Let $\mathcal{I}_{nkc}$ be the set of units in fold $k$ and cluster $c$, and let $\mathcal{C}_{nk}$ be the indexes of the clusters selected for fold $k$. $$ \widehat{H}_{nk}^{cluster} := \frac{1}{n_k}\sum_{c \in \mathcal{C}_{nk}}\left( \sum_{i \in \mathcal{I}_{nkc}} \begin{bmatrix} \lambda(X_i)\widehat{U}_i\widehat{W}_{i} \\ \widehat{T}_{i} \end{bmatrix} \right)\left( \sum_{i \in \mathcal{I}_{nkc}} \begin{bmatrix} \lambda(X_i)\widehat{U}_i\widehat{W}_{i} \\ \widehat{T}_{i} \end{bmatrix} \right)'.$$ The clustered standard errors are $\widehat{\Omega}_{nk}^{cluster} = \widehat{\Upsilon}_{nk}\widehat{J}_{nk}^{-1}\widehat{H}_{nk}^{cluster}\widehat{J}_{nk}^{-1}\widehat{\Upsilon}_{nk}'$, where $\widehat{J}_{nk},\widehat{\Upsilon}_{nk}$ are computed as outlined in (ref).
rem[Confidence intervals when $V_{xnk} = 0$] When the conditional mean is constant, i.e. $\mu_d(x) = \mathbb{E}[Y_d]$, prediction models with corner solutions like LASSO may estimate a constant conditional mean, i.e. $\widehat{\mu}_{d,-k}(x) = \widehat{\mu}_{d,av}$, $\widehat{\tau}_{-k}(x) = \widehat{\mu}_{1,-k}(x) - \widehat{\mu}_{0,-k}(x) $ is constant, and consequently $V_{xnk} = 0$.\footnote{It is still possible to have $V_{xnk} > 0$ almost surely even if $V_{\tau} = 0$, as long as $\mu_0(x)$ is not constant.} This violates Assumption (ref).(i), and it is challenging to construct a confidence interval with exact coverage. One alternative is to construct an ensemble of sparse and non-sparse estimators of the CATE in the first-stage. Another alternative is to use degenerate confidence intervals: \begin{equation} \widehat{CI}_{\alpha n k}^{0} = \begin{cases} \widehat{CI}_{\alpha n k} & if V_{xnk} \ne 0, \\ [0,0] &if V_{xnk} = 0. \end{cases} \end{equation} The confidence intervals collapse to zero when the $\widehat{\tau}_{-k}(x)$ prediction is degenerate. For example, in LASSO researchers can check whether the coefficients are zero, in tree-based methods when there are no splits, or whether $\widehat{V}_{xnk} = 0$. We can also define an analogous multifold confidence interval. \begin{equation} \widehat{CI}_{\alpha n}^{0,multifold} = \left[ Med_{K}\left\{ \inf \widehat{CI}_{\frac{\alpha}{2} nk} \right\}, Med_{K}\left\{ \sup \widehat{CI}_{\frac{\alpha}{2} nk} \right\} \right]. \end{equation} I study the asymptotic properties of these confidence intervals. \begin{lem} Let $\Gamma$ denote a set of distributions that satisfy Assumptions (ref), (ref), and (ref). Suppose that Assumption (ref) holds, except for the requirement that $V_{xnk} = 0$. Then (i) \begin{equation} \underset{n \to \infty}{\lim\inf}\inf_{\gamma \in \Gamma} \ \mathbb{P}_{\gamma}\left( V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) \in \widehat{CI}_{\alpha n k}^{0} \right) \ge 1-\alpha \end{equation} (ii) If Assumptions (ref), and (ref) also hold, then \begin{equation} \inf_{\gamma \in \Gamma} \underset{n \to \infty}{\lim\inf} \ \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) \in \widehat{CI}_{\alpha n k}^{0} \right) \ge 1-\alpha \end{equation} \begin{equation} \inf_{\gamma \in \Gamma} \underset{n \to \infty}{\lim\inf} \ \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) \in \widehat{CI}_{\alpha n k}^{0,multifold} \right) \ge 1-\alpha. \end{equation} \end{lem} To prove this result I focus on the coverage for subsequences where $V_{xnk} = 0$ and $V_{xnk} > 0$, and apply the results for conservative coverage results in andrews2020generic. In subsequences where $V_{xnk} = 0$, then $V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) = 0$ which means that coverage of the pseudo-VCATE is equal to one. In subsequences where $V_{xnk} > 0$ and assuming that $\Omega_{nk}$ has eigenvalues bounded away from zero, then we can apply similar arguments as before to prove $1-\alpha$ coverage. To prove point-wise coverage, I separate the cases where $V_{\tau}(\gamma) = 0$ and $V_{\tau}(\gamma) > 0$. In the latter case, I show that $V_{xnk}$ is point-wise bounded away from zero, though not uniformly. The proof of Lemma (ref) does not rely on the i.i.d. assumption, and can also accommodate cluster dependence. In the empirical example, I compute confidence intervals with clustered standard errors and degenerate CATE predictions. When there's more heterogeneity and the nuisance functions are estimated accurately, then $V_{xnk} > 0$ with high probability. However, when $V_{\tau n} \approx 0$ and $\mathbb{V}(\mu_0(X)) \approx 0$, then procedures like LASSO may imply $V_{xnk} = 0$ fu2000asymptotics, which means that marginally heterogeneous CATEs could be estimated as homogeneous. This could be impact the power of tests of homogeneity. The size for two-sided tests is not uniformly bounded. Furthermore, the multifold confidence interval allows for some quantification of uncertainty across folds/splits: the CI is degenerate only if more than half the fold/split-specific CIs are degenerate. Furthermore, the degenerate CI has correct size control for one-sided tests. \begin{cor} Under the assumptions of Lemma (ref).(i), \begin{equation} \underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) < \inf \widehat{CI}_{\alpha n k}^{0} \right) \le \alpha. \end{equation} \begin{equation} \underset{n \to \infty}{\lim\sup}\sup_{\gamma \in \Gamma} \mathbb{P}_{\gamma}\left( V_{\tau}(\gamma) < \inf \widehat{CI}_{\alpha n k}^{0,multifold} \right) \le \alpha. \end{equation} \end{cor} The tests of homogeneity have the correct size when $V_\tau(\gamma) = 0$. Corollary (ref) guarantees that the probability of falsely rejecting a class of one-sided test is uniformly bounded in large samples.
rem[Monotonic Transformations] It may be useful to report the standard deviation of the CATE, which is $\sqrt{VCATE}$. I propose the following confidence interval: \begin{equation} \widehat{CI}_{\alpha n}^{0,multifold,sqrt} = \left\{\sqrt{V_{\tau}^*}: V_{\tau}^* \in \widehat{CI}_{\alpha n}^{0,multifold}\right\}. \end{equation} Since the square root is a strictly increasing transformation and the VCATE is non-negative, then $\sqrt{V_{\tau}(\gamma)} \in \widehat{CI}_{\alpha n}^{0,multifold,sqrt} $ if and only $V_{\tau}(\gamma) \in \widehat{CI}_{\alpha n}^{0,multifold}$. Since the events are equivalent, the transformed confidence interval preserves the coverage probabilities and will have valid coverage by Lemma (ref).

Simulations

I use a simulation design to study the properties of the VCATE estimators. The baseline covariates are distributed as $[X_0,X_1] \in \mathcal{N}(0,\Sigma_x)$, where $\rho = 0.5$ and $$ \Sigma_x =

bmatrix[bmatrix omitted — 92 chars of source]

.$$ The random variables $X_0$ and $X_1$ are standard normal vectors of dimension $J$. The covariance between pairs of components $X_{0j}$ and $X_{1j'}$ is equal to $\rho = 0.5$ for when $j = j'$, but zero otherwise. The outcome is generated from a model where $Y = DY_0 + (1-D)$, $D$ is generated by a Bernoulli draw with probability $0.5$, and

align[align omitted — 280 chars of source]

where $c,\tau \in \mathbb{R}$, $\beta_0,\beta_\tau,\kappa_0,\kappa_1 \in \mathbb{R}^p$. The errors $(U_0,U_1)$ are independent of the covariates $(U_0,U_1) \indep (X_0,X_1)$, and distributed as standard normals $[U_0,U_1]' \in \mathcal{N}(0_{2 \times 1},I_2)$. The key model quantities have closed-form expressions. The conditional means at baseline and the CATE are given by $\mu_1(x) = \alpha + \beta_0'x_0$ and $\tau(x) = \tau + \beta_\tau'Z_1$, respectively. The conditional variances are $\sigma_d^2(x) = \kappa_d'x_dx_d'\kappa_d$ for $d \in \{0,1\}$. This formulation incorporates heteroskedasticity. Covariates that influence the outcomes at baseline may also affect the treatment effects.

The regressors are constructed in such a way that $\mathbb{E}[X_dX_d'] = I_p$ for $d \in \{0,1\}$. This implies simple expressions for the variances of the model, $\mathbb{V}(U_d) = \tilde{\sigma}_d^2 + \kappa_d'\kappa_d$, $$ V_{\tau} = \beta_\tau'\beta_\tau, \qquad \mathbb{V}(Y_0) = \beta_0'\beta_0 + \tilde{\sigma}_d^2 + \kappa_0'\kappa_0, $$ $$ \mathbb{V}(Y_1) = \beta_0'\beta_0 + \beta_\tau'\beta_\tau + 2(1-\rho)\beta_0'\beta_\tau + \tilde{\sigma}_1^2 + \kappa_1'\kappa_1, $$

I choose an approximately sparse specification for (ref) where the coefficients decay exponentially at a rate of decay of $\lambda = 0.7$. Let $\ell_j = \sqrt{\left(\frac{1-\lambda}{1-\lambda^J}\right)\lambda^{1-j}}$ be a geometric sequence, which satisfies $\sum_{j=1}^J \left(\frac{1-\lambda}{1-\lambda^J}\right)\lambda^{1-j} = 1$. Given user-specified parameters $(V_{\mu},V_{\tau},\sigma_0^2,\sigma_1^2)$, the coefficients for the entries $j \in \{1,\ldots,J \}$ are determined by $\beta_{0,j} = \ell_j\sqrt{V_{\mu}}$, $\beta_{\tau,j} = \ell_j \sqrt{V_\tau}$, $\kappa_{d,j} = \ell_j \sqrt{\sigma_d^2-\tilde{\sigma}_d^2}$, for $d \in \{0,1\}$. Since $\sum_{j=1}^J \left(\frac{1-\lambda}{1-\lambda^J}\right)\lambda^{1-j} = 1$, then $\beta_0'\beta_0 = V_{\mu}$, $\beta_\tau\beta_\tau = V_{\tau}$, and $\beta_0'\beta_\tau = \sqrt{V_\mu V_\tau}$. We can obtain analogous expressions for the variances of the unobserved components, so that $\tilde{\sigma}_d^2 + \kappa_0'\kappa_0 = \sigma_d^2$ for $d \in \{0,1\}$.

I choose an average effect size of $\tau = 0.15$, that is coherent with the recent meta-analyses of economic experiments in vivalt2015heterogeneous. To make sure that the magnitudes are interpretable, I normalize the coefficients so that the variance for the control group is $\mathbb{V}(Y_0) = 1$, by setting set $c = 1$, $\sigma_d = 0.7$, $\tilde{\sigma}_d = 0.21$, and $V_{\mu} = 0.3$. The design is easy to scale for different values of $V_{\tau}$ and $J$. My design is similar to that in belloni2014inference but I choose $\Sigma_x$ and the sparsity structure in such a way that $V_{\tau}$ has a closed form expression. I use LASSO to estimate $\mu_1(x)$ and $\mu_0(x)$, tuned via cross-validation. The coefficients of this model are consistent given this sparse linear structure, even in high dimensions. I randomly simulate 2000 datasets to compute each of the estimators, and split them into $K=2$ folds.

figure[figure omitted — 741 chars of source]

Figure (ref) considers a simulation with $n = 2500$. The figure displays a density plot for the multi-step estimator, $\widehat{V}_{\tau n}$ defined in (ref), and a two-step debiased machine learning estimator computed as:

equation[equation omitted — 155 chars of source]

where $\widehat{\eta}_{-k}(x) = (\widehat{\tau}_{-k}(x),\widehat{\mu}_{0,-k}(x),p(x),\widehat{\tau}_{n,av})$ and $\widehat{\tau}_{n,av} = \frac{1}{n}\sum_{i=1}^n \widehat{\tau}_{-k_i}(X_i)$. When $V_{\tau} > 0$ the efficient influence function is non-degenerate. In high-heterogeneity regimes both converge to the same limiting distribution.\footnote{This is shown in Theorem (ref) for the multistep approach and can be shown for the two-step using standard arguments, e.g. chernozhukov2018double.} However, when $V_{\tau} = 0$, the influence function is degenerate and they may converge at different rates. We see that the multi-step approach is much more precise. This can be explained by the fast boundary convergence rates derived in Theorem (ref). The two step approach can also produce negative estimates of $V_{\tau}$, which is an undesirable feature, whereas the multi-step estimator is always non-negative. Both estimators have higher bias when the dimension increases because there is more first-stage noise.

figure[figure omitted — 861 chars of source]

Figure (ref) plots the root mean-square error (RMSE) of $\widehat{V}_{\tau n}$ and $\widehat{V}_{\tau n}^{two-step}$ for different sample sizes. I compute the semiparametric efficiency bound by computing the RMSE of an “oracle” estimator that substitutes $\widehat{\eta}_{-k}(X_i) = \eta(X_i)$ in (ref). The results show that as the sample size increases, both estimators achieve a higher level of accuracy and their variance approaches the semi-parametric lower bound (the RMSE of the oracle). As expected by Corollary (ref), the semiparametric lower bound is zero at the boundary. The differences in RMSE shorten with higher $V_{\tau}$ and in lower dimensional settings $(2J = 10)$ (which have lower first-stage noise).

figure[figure omitted — 1,036 chars of source]

Figure (ref) shows the coverage of $V_{\tau}$ for the different proposed confidence intervals (CIs). For the multi-step approach, I consider the single splits CIs in (ref) and the conservative multi-fold CIs from and (ref). The two-step CIs are constructed as $\frac{1}{n}\sum_{i=1}^n \widehat{\varphi}_i \pm 1.96\sqrt{\widehat{V}_{\varphi}/n}$, where $\widehat{\varphi}_i$ is the summand in (ref) and $\widehat{V}_{\varphi}$ is an estimate of its sample variance. The coverage of the two-step approach is very low under homogeneity, and there is no improvement as sample size increases when $V_{\tau} = 0$. The coverage of the two-step estimator only improves with higher $n$, in high heterogeneity designs. By contrast, both multi-step approaches cover the parameter at the intended level, and coverage improves with higher sample size. For fixed $n$, coverage degrades for both cases when the number of covariates is higher.

Figure (ref) explores the differences in covering the VCATE vs the pseudo-VCATE when $n = 2500$ and $2J = 10$ for a fine-grained set of values of $V_{\tau}$. Panel (a) reflects a dip in coverage close to the boundary. My theory predicts that the multistep CIs have exact coverage when $V_{\tau} = 0$, but may not cover uniformly close to the boundary (see discussion after Theorem (ref)). The mulit-fold CIs have conservative coverage. Conversely, Figure (ref), Panel (b) shows the multi-step CIs always uniformly cover the pseudo-VCATE, as predicted by theory. This provides a robustness guarantee for how to interpret the CIs. The two-step approach has much lower coverage and no guarantees when $V_{\tau} = 0$ in either panel.

figure[figure omitted — 923 chars of source]
figure[figure omitted — 883 chars of source]

Figure (ref).(left) shows the power of tests of homogeneity in a simulation with $n = 2500$. The multi-step, single fold approach has correct size control and has local power, in line with the result of Lemma (ref). The power of the test using the multifold approach is similar to using a single fold. The right panel shows the probability that the VCATE is strictly below the CI bounds. As predicted by theory, this probability is uniformly bounded by $\alpha = 0.05$ for the single and multifold approaches (see Corollaries (ref) and (ref), and Theorem (ref)). The two-step approach has a non-monotonic power curve with incorrect size. The size of one-sided tests in the right panel is uniformly bounded by $\alpha = 0.05$, though this may be partly the fact that $\widehat{V}_{\tau n}^{two-step}$ can take negative values and has a negative bias (see Figure (ref)).

Empirical Example

In this section, I illustrate my approach using data from a large-scale information experiment conducted by dizon2019parents. The study, which covered 39 school districts, involved an intervention to provide low-income parents of at least two children with information about their children's school performance. Half the households where assigned to the information intervention and the rest were assigned to the control group. dizon2019parents showed that, at baseline, parents faced large information gaps regarding their children's grades and class ranking. Even though schools produced a report card, 60% of parents were unaware of their child's performance. Many parents reported that they did not receive the report card (children either lost them or did not take them home), or had trouble interpreting the report card structure, primarily due to low literacy levels.

The intervention was designed to present details of their children's school performance in an easily accessible way. dizon2019parents showed that the information gaps (the difference between believed and true test scores) went down as a result of the intervention, and the amount of updating varied depending on students' initial test scores. dizon2019parents also introduced a real-stakes scenario where parents received a series of lottery tickets for a scholarship paying for four years of high school. Parents had to decide how to allocate tickets between two siblings. If there were more than two siblings residing in the household, the survey team selected two at random. The results showed that parents allocated tickets towards their better performing child.

To test for heterogeneity, dizon2019parents ran a linear regression of parental beliefs on initial scores, treatment, and an interaction as in (ref), and reported estimates $\bar{X} = 46.8$ (on a scale of 100) and $(\widehat{\beta}_1,\widehat{\beta}_2) = (-25.9,0.40)$ in their Tables 1 and 2, respectively. The coefficient $\widehat{\beta}_2$ captures how much the treatment effects vary (on a $0$ to $100$ scale) for a each additional point in students initial scores. The VCATE combines information about the coefficient and the initial variability in scores. From the data, I also estimate the variance of the control $\widehat{V}_x = 305.53$, and estimate the VCATE as $\widehat{\beta}_2^2\widehat{V}_x = 48.89$. Taking the square root and normalizing by the standard deviation of the outcome for the control group ($\widehat{V}_{(Y \mid D= 0)} = 311.72$), produces $0.40$. This means that the magnitude of treatment effect heterogeneity explained by scores is comparable to 40% of the standard deviation of beliefs in the control group.

The VCATE can also help us understand the magnitude of treatment effect heterogeneity in the experiment using multiple covariates. I use LASSO for the first-stage predictions, two folds per split with 20 splits, and estimate $\{\widehat{\Omega}_{nk}\}_{k=1}^K$ using clustered standard errors at the household level and the formulas for the confidence intervals defined in (ref). The results are not very sensitive to the number of splits. For ease of exposition, I report point-estimates and CIs for $\sqrt{\frac{V_\tau}{\mathbb{V}(Y_0)}}$, which is the standard deviation of the CATE divided by the standard deviation of the outcome for the control group. I compute confidence intervals for the square root via the transformation proposed in (ref).

table[table omitted — 3,139 chars of source]

Table (ref) computes the ATE and the $\sqrt{VCATE}$ for two outcomes (parental beliefs and lottery allocations) and 8 different sets of covariates. Panel (a) shows that, on average, parents downgrade their beliefs about test scores by 42% of the standard deviation (SD) of the beliefs of the control group. The treatment effect heterogeneity explained by test scores is equivalent to 40% of the standard deviation (SD) of the beliefs of the control group. This is statistically significant at the 5% level and has a comparable magnitude to the ATE. The confidence intervals are relatively short in length. However, applying the bounds from Theorem (ref) and Corollary (ref) shows that differentiating treatment offers based on scores could further lower beliefs by at most 8.1% SDs of the beliefs in the control group. In this case the ATE is already fairly high compared to $\sqrt{V_{\tau}}$, so in spite of the large heterogeneity, the marginals gains from targeting would be modest. Panel (b) presents the results for the secondary school lottery. The ATE is estimated precisely at zero, because the lottery tickets had to be divided as a zero sum between the siblings. The VCATE measures how much the dispersion in the allocation depends on the covariates. The standard deviation of the VCATE explained by initial scores is 16% of the SD of the control group lottery allocation, and the maximum welfare gains are around 7.8% SD.

The student variables (grade, age, gender, attendance, and educational expenditures) collectively explain 11% of the SD of parental beliefs. This is statistically significant at the 5% level. The magnitude is around a fourth of the variation for test scores, and the maximum welfare gain from targeting is 0.7% of the SD of parental beliefs in the control group. I find that other subsets of covariates do not produce statistically significant estimates of the VCATE at the 5% level. The added welfare of personalizing treatment assignment using these covariates is also very low.

The estimates that use all the covariates are computed over a smaller subsample with non-missing values across all variables. Despite the large number of variables and the smaller sample, the estimates of the VCATE remain relatively stable across specifications. The VCATE computed from a rich set of respondent, household, and student covariates has a comparable magnitude to the VCATE that only includes student scores. The confidence intervals are also similar. The estimates of the maximum welfare gains from targeting using all covariates are 7.9% SD for beliefs and 7.4% SD for lottery outcomes, respectively, which are similar to the welfare gains computed using only scores.

Conclusion

I propose an efficient estimator of the variance of treatment effects that can be attributed to baseline characteristics and propose novel adaptive confidence intervals that produce valid coverage. I analyze issues of non-standard inference that arise in this context, and how to address them. I also explore the economic significance of the VCATE for policymakers and researchers, by showing that the $\sqrt{VCATE}/2$ bounds the marginal gains of targeted policies. Overall, this paper proposes a broadly applicable approach to measure treatment effect heterogeneity in experiments.

\counterwithin{assump}{section}