Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
122,680 characters · 19 sections · 65 citation commands
Robust inference for the treatment effect variance in experiments using machine learning
\thispagestyle{empty}
\onehalfspacing \selectfont
In recent years, there has been a rapid expansion of experiments to evaluate public policy programs and corporate initiatives. There is also more evidence that the effectiveness of a program can vary across individuals. For instance, dizon2019parents studies a population of low-income parents in Malawi with large misperceptions about their children's school performance. She finds that a simple intervention can bridge these gaps and that the Conditional Average Treatment Effect (CATE) varies by children's initial scores. In practice, even though researchers collect baseline surveys with many characteristics, the CATE is typically estimated via regressions with one or two interactions, thus underutilizing the full set of variables. The promise of leveraging the vast and readily available baseline data has sparked more applications of supervised machine learning crepon2021cream,davis2020rethinking,deryugina2019mortality. Methods such as LASSO, neural networks, random forests, or boosting are data-driven and allow for more variables and flexibility.
In this paper, I focus on the unconditional variance of the CATE, the VCATE, which measures the dispersion of treatment effects predicted by a set of baseline characteristics. The VCATE has a clear interpretation, even if the CATE is nonlinear or depends on many characteristics. chernozhukov2020generic and ding2019decomposing separately propose estimators with a misspecification robust interpretation, whereas levy2021fundamental propose an efficient estimator. Despite these recent advances, there are currently no valid confidence intervals for the VCATE. In fact, levy2021fundamental study the performance of confidence intervals based on the efficient influence function, i.e., the conventional way. They find that coverage degrades near the boundary, reaching a low of 32% in simulations. They speculate that poor coverage is due to the degeneracy of the efficient influence function when the VCATE is zero. As a result, conventional guarantees for $\sqrt{n}-$asymptotic normality do not apply to this part of the parameter space. Moreover, a VCATE close to zero is economically meaningful because it could reflect null effects, low effect heterogeneity, or irrelevant covariates. For such situations, which are common in practice, conventional approaches to inference could be misleading.
This paper provides fresh insights regarding the VCATE and proposes a solution for inference in experiments. I propose (a) novel ways to interpret the VCATE for decision-making by deriving sharp bounds on the population gains of personalized treatment assignment, (b) novel estimators of the VCATE that are both misspecification-robust and efficient, combining the best features of previous approaches, and (c) novel confidence intervals that are shape-adaptive and fast to compute.
I break down why conventional confidence intervals have incorrect size. I show that the boundary inference problem can manifest even when the CATE is linear and univariate. I solve the problem for the linear case by proposing adaptive confidence intervals that meet the high-level conditions outlined in andrews2020generic. I then show that these can be readily extended to the class of nonlinear models in chernozhukov2020generic that combine regression adjustments and a machine learning first stage. In addition to providing new confidence intervals, my paper has novel implications for estimation by showing that only a subset of these multi-step estimators is efficient. Adaptive inference and regression adjustments work well for various predictive models under weak assumptions.\footnote{While I specialize my results to the VCATE, my inference approach relies on more general principles: I use knowledge of the limiting distribution function, conditional on the cross-fitted estimates. Correct coverage follows from verifying high-level assumptions that can be satisfied by a wide array of machine learning method used in the first step. In principle, my approach could extend to other non-standard inference problems.} In the fully nonparametric case, I use a conservative procedure with valid coverage over multiple sample splits.\footnote{This adjustment is in the spirit of chernozhukov2020generic, who propose robust t-tests assuming a conditionally normal distribution. Their results are not directly applicable here due to the boundary inference problem. However, I apply the principles behind their “median-parameter” confidence intervals to my adaptive intervals.} I derive the local power curve for the associated tests of homogeneity and their relationship to the tests in crump2008nonparametric and ding2019decomposing. I also propose confidence intervals for settings with cluster dependence.
I document excellent root mean square error (RMSE) performance and coverage in simulations using LASSO, even in high dimensions. I benchmark my multi-step approach against a two-step debiased machine learning estimator. As predicted by theory, all approaches are asymptotically normal, efficient, and have good coverage in highly heterogeneous designs. However, when the VCATE is zero or close to zero, coverage of two-step alternatives can be as low as 45%. By contrast, my adaptive intervals produce coverage at the intended 95% level and better RMSE at all regions of the parameter space. I study the robustness of the multi-step approach in both the theory and simulations. I consider situations where the predictive component is misspecified or slow to converge. I also discuss issues related to uniform vs. pointwise coverage.
I apply my approach to data from dizon2019parents, an information experiment with low-income parents in Malawi who had at least two school-age children. The intervention redesigned the way in which parents received information about their children's school performance. The endline survey measured parental beliefs about student grades and asked parents to allocate tickets to a scholarship between their children. dizon2019parents presented graphs with a non-parametric CATE by baseline test scores, which had an approximately linear shape, and separately tested for significance using a regression with an interaction. I use LASSO to compute the VCATE for the two outcomes (parental beliefs and lottery allocations) by different characteristics of students, parents, and households. To make the results interpretable, I focus on the standard deviation of the CATE, i.e., $\sqrt{VCATE}$, and normalize it by the standard deviation of the outcome in the control group.
My approach allows us to quantify the magnitude of effect heterogeneity. I find that the treatment effect heterogeneity explained by test scores is equivalent to 40% of the standard deviation (SD) of the beliefs of the control group, and 16% of the SD of the control group lottery allocation. I also find that the effect heterogeneity collectively explained by other student variables (grade, age, gender, attendance, and educational expenditures) is comparable to 11% of the SD of beliefs in the control group. The VCATE of beliefs by student variables is significant at the 5% level, but the VCATE of lottery outcomes by student variables is not. The combined VCATE associated with student scores and 12 other key characteristics has a similar value to the VCATE with only scores. Despite being conservative, the intervals for the VCATE are short in length in this empirical example. Using my new welfare bounds $(-|ATE| + \sqrt{VCATE + ATE^2})/2$, I predict that targeted interventions using the baseline covariates have a maximum added benefit of 7.9% SD and 7.4% SD (standard deviations of the outcome for the control group) on beliefs and lottery allocations, respectively.
Researchers should focus on the VCATE because it is a model-free quantity with good properties: it is well-defined even if the CATE is continuous or discrete, and it weakly increases when researchers add more covariates to their analysis. Researchers can test for homogeneity by evaluating whether confidence intervals for the VCATE include zero. In addition to testing, by quantifying the VCATE, researchers can compare the magnitude of heterogeneity relative to a benchmark, such as the variance at baseline, the VCATE for different covariates, or experiments in other sites.
My first main contribution is to show that the VCATE provides a bound for the welfare gains of policy targeting. A policymaker might decide to use the information from the CATE to design personalized treatment recommendations manski2004statistical,athey2021policy,kitagawa2018should,mbakop2021model. One can measure utilitarian welfare by computing the expected outcome under different policies. Such policies can be further constrained to a class that respects budget limits, incentive compatibility, or fairness considerations viviano2020fair,sun2021empirical. I show that the difference in mean outcomes between a targeted policy and a non-targeted policy using only the average treatment effect (ATE), is bounded by $\sqrt{VCATE}/2$. For instance, under homogeneity ($VCATE = 0$), there are no gains from targeting. I show that this bound holds in the population regardless of the choice of policy class and the underlying distribution. Furthermore, the bound is sharp in the sense that it holds exactly for at least one policy and distribution.
The proposed bound on utilitarian welfare communicates information to practitioners about whether a targeting exercise is even worth pursuing, without needing to solve the targeting problem itself. The VCATE can be a supplemental quantity reported in regression analyses, or a benchmark for analysts choosing the optimal policy. If the VCATE is very low, practitioners may consider expanding the set of covariates in the analysis. To derive the bound, I use a constructive approach to solve the most adversarial distribution. I also prove a more general bound $(-|ATE| + \sqrt{VCATE + ATE^2})/2$, and show that the distribution that leads to a maximum welfare gain is one where the CATE has binary support and mean zero. The gains from targeting easily diminish if the value of the ATE is relatively higher than the VCATE.
My second contribution is related to efficient estimation and robust inference. New theory is required here because of a unique feature of the VCATE: the efficient influence function is degenerate when the CATE is homogeneous levy2021fundamental. Classical results by newey1990semiparametric show that any regular, efficient estimator can be decomposed as $\frac{1}{n}\sum_{i=1}^n \varphi_i + R_n$, where $n$ is the sample size, $\{\varphi_i\}_{i=1}^{n}$ are a set of i.i.d. mean-zero influence functions, and $R_n$ is a residual with higher-order terms that are $o_p(n^{-1/2})$.\footnote{Many standard estimators can achieve this property, e.g., the “debiased machine learning” estimator chernozhukov2018double or the targeted maximum likelihood estimator in levy2021fundamental.} Conventional approaches assume that $\mathbb{V}(\varphi) > 0$ and in this case, the estimation error converges at $\sqrt{n}$ to $\mathcal{N}(0,\mathbb{V}(\varphi_i))$ by the CLT. However, when the VCATE is zero, $\mathbb{V}(\varphi_i) = 0$ as well. Hence, the limiting distribution is dominated by the higher terms in $R_n$, which may not be asymptotically normal. Therefore, while $\sqrt{n}-$estimation is still possible, t-tests that plug in an estimate of $\mathbb{V}(\varphi)$ may have incorrect coverage. By contrast, other common quantities such as the average treatment effect (ATE) or the local average treatment effect (LATE) do not have this problem because they satisfy $\mathbb{V}(\varphi_i) > 0$ uniformly chernozhukov2018double.
I start by analyzing a simple two step estimator, assuming that the CATE is linear in covariates and can be estimated from a regression. I show that the limiting distribution of the VCATE estimator can be written as a linear combination of a Chi-square that converges at $n$-rate and a normal distribution that converges at $\sqrt{n}-$rate. The weights are determined by the value of the VCATE, which means that the shape of the distribution changes depending on the region of the parameter space. At the boundary, it behaves like a rescaled chi-square, is $O_p(n^{-1})$, and confidence intervals with normal critical values will have incorrect coverage. For values of the VCATE bounded away from zero, the distribution is asymptotically normal as in the classical results.
In the linear case, I construct adaptive confidence intervals that account for the higher terms of the distribution. I apply the framework of andrews2020generic to show that this produces uniform, exact coverage when the linear model is correctly specified.\footnote{This type of strategy has proven effective to deal with other non-standard problems where the shape of the limiting distribution depends on an unknown parameter, such as the AR coefficient in a time series, the effect parameter under weak instruments, or the quasi-likelihood ratio test for nonlinear regression andrews2020generic.} The intervals are fast to compute because the expressions are all analytic. When there is a single covariate, I also show that a homogeneity test that evaluates whether zero is contained in the confidence intervals is algebraically identical to (i) a test of whether the interaction in the regression model is equal to zero, and (ii) the single-covariate homogeneity test of crump2008nonparametric. The test is also asymptotically equivalent to ding2019decomposing. However, these other tests only apply to series estimators and are not nested with mine in the multivariate and non-parametric cases.
I extend my results to the class of nonlinear models proposed by chernozhukov2020generic. chernozhukov2020generic showed that in experiments with known assignment probabilities, their models produce a meaningful pseudo-VCATE even if the functional form is misspecified. The pseudo-VCATE is non-negative, weakly lower than the VCATE, and converges to the true value under mild conditions on the estimated CATE.\footnote{This monotonicity property means that in experiments the pseudo-VCATE will not falsely detect heterogeneity, even if the machine learning stage is misspecified.} chernozhukov2020generic argue that the pseudo-VCATE might be of independent interest as a measure of model fit.\footnote{ ding2019decomposing also define a similar pseudo-VCATE based on randomization inference.} They describe a three-step estimator with a machine learning/prediction first-stage, a regression second-stage, and a sample-variance third stage. Related multi-step estimators have also been considered in other work guo2021machine.
To the best of my knowledge, there are no existing asymptotic results for the multi-step VCATE estimator proposed in chernozhukov2020generic. I fill in that gap by proving two sets of results. First, I show that all the estimators in their class converge to the true VCATE at least at $\sqrt{n}$-rate, are $o_p(n^{-1/2})$ at the boundary (as in the simple linear model), and have the convenient property that they are always non-negative. This builds on the asymptotic expansion for the linear case I introduced above. Second, I prove that only a subset of the chernozhukov2020generic estimators are efficient, i.e. converge at $\sqrt{n}-$ to an average of i.i.d. efficient influence functions. The key ingredient is to prove a novel finite-sample equivalence result. I find that the first order conditions of the regression step and the bias-correction component of the VCATE influence function are in fact identical, given a particular decomposition of the nuisance functions. The asymptotic results follow from fairly standard assumptions on convergence rates chernozhukov2018double,belloni2017program. To get the limiting distribution, the only meaningful extra assumption is that the estimated CATE has bounded kurtosis (thin tails).
I show that extending the adaptive confidence intervals (CIs) to multi-step estimators is straightforward. The procedure randomly splits the data into subsets or folds and estimates the nuisance functions and the VCATE on different folds. To compute the confidence intervals for a particular fold, the researcher can treat the second-stage regression as if the variables were given, and then construct the CIs as in the simple case. I construct median confidence intervals (CIs) to aggregate information across multiple folds. I show that the single fold procedure produces uniform, exact coverage for the pseudo-VCATE and point-wise, exact coverage for the VCATE for all points in the parameter space, at a nominal level $1-\alpha$. The multifold CIs have pointwise conservative coverage.
Furthermore, the probability that the true VCATE is below the confidence interval bounds is uniformly bounded by $\alpha$ in large samples. This result applies to the single and multifold CIs and does not require that the first-stage estimates to converge. Instead it relies on the fact that in experiments the pseudo-VCATE is weakly lower than the true VCATE. Tests for homogeneity (whether zero is contained in the CI) belong to this broader class of tests. Having uniform size control for this class of one-sided tests means that my tests of homogeneity are robust.
This paper is also related to a growing literature on debiased-machine learning chernozhukov2016locally,chernozhukov2022automatic,belloni2014inference,belloni2017program,chernozhukov2018double, semiparametric efficiency newey1990semiparametric, uniform inference for non-standard problems andrews2020generic, and tests of treatment effect homogeneity ding2019decomposing,crump2008nonparametric,heckman1997making,bitler2017can. My approach combines results from these literatures by addressing a boundary inference problem with a machine learning stage, and applying techniques of uniform inference. A related literature also focuses on confidence intervals around point-predictions of the CATE athey2019generalized,semenova2021debiased, rather than overall measures of dispersion.
Section (ref) provides key definitions, introduces the welfare bound, and presents a version of the adaptive confidence intervals for the univariate regression case. Section (ref) frames the inference problem in a more general setting, and extends the adaptive confidence intervals for VCATE estimation with a machine learning first stage. Section (ref) presents the large sample theory. Section (ref) introduces the simulations. Section (ref) applies my approach to an empirical example from Malawi. Section (ref) concludes.
Consider a program evaluation setting in which an individual is assigned to either a treatment $(D=1)$ or a control group $(D= 0)$. The outcome of interest $Y$ depends on the treatment status. I denote the potential outcome under treatment and control status as $Y_1$ and $Y_0$, respectively, and the treatment effect as $Y_1 - Y_0$. The conditional average treatment effect (CATE) given covariates $X$ is defined as $$ \tau(x) := \mathbb{E}[Y_1 - Y_0 \mid X = x],$$ and the average treatment effect (ATE) is defined as $\tau_{av} := \mathbb{E}[Y_1 - Y_0]$. This paper proposes an estimator of the variance of the CATE (VCATE) defined as $$ V_{\tau} := \mathbb{V}(\tau(X)).$$ The variance $V_\tau$ measures the dispersion of treatment effects that can be attributed to observable characteristics $X$. The value of $V_{\tau}$ depends on the choice of covariates. To understand how different covariates might impact the VCATE, let $V_{\tau}' = \mathbb{V}(\mathbb{E}[Y_1 - Y_0 \mid X'])$ be the VCATE for a different set of covariates $X'$.
Lemma (ref) shows that the VCATE has the following monotonicity property: if the researcher adds more covariates to the analysis, or breaks down an existing covariate into more categories, then the VCATE will be weakly larger.
The propensity score, $p(x)$, is defined as follows
I restrict attention to experimental settings where $p(x)$ is known. The CATE can be identified under further assumptions.
Assumption (ref).(i) formalizes the idea that the researcher can only observe either $Y_1$ or $Y_0$, but not both, for any particular individual. Assumption (ii) holds in randomized controlled trials with treatment probabilities bounded away from $\{0,1\}$. Assumption (iii) states that an individual's treatment probability depends on $X$ but not their potential outcomes. Let $\mu_d(x)$ be the conditional mean of $Y$ given $X$ and a fixed value of $d \in \{0,1\}$,
Under Assumption (ref), $\mathbb{E}[Y_d \mid X =x ] = \mu_d(x)$, and hence $\tau(x) = \mu_1(x) - \mu_0(x)$. This means that the VCATE is identified, with $V_{\tau} = \mathbb{V}(\mu_1(X)-\mu_0(X))$.
Practitioners can use estimates of $\tau(X)$ to decide whom to treat in future interventions athey2021policy,kitagawa2018should,manski2004statistical. Program managers can target the treatment recipients based on their initial covariates. However, whether targeting can substantially improve average outcomes depends on the dispersion of $\tau(x)$. I show that a simple function of the VCATE bounds the marginal gains of targeting.
Let $\gamma$ denote the joint distribution of $(X,Y_1,Y_0)$, $\mathcal{X}$ a set containing the support of $X$, $\tau_\gamma(x) := \mathbb{E}_\gamma[Y_1 - Y_0 \mid X =x]$ the CATE given $\gamma$. A function which maps $x$ to a probability of treatment $\pi(x)$ is known as a statistical allocation rule manski2004statistical. Furthermore, I denote the set of all possible allocation rules by $\Pi$, which contains all functions $\{\pi:\mathcal{X} \to [0,1]\}$. The set $\Pi$ includes many well-known assignment rules. For instance, it includes the “non-targeted” policy which assigns everyone to treatment if $\mathbb{E}_\gamma[Y_1] > \mathbb{E}_\gamma[Y_0]$ and to the control group otherwise. Moreover, the average outcome under rule $\pi$ is $\mathbb{E}_\gamma[\pi(X)Y_1 + (1-\pi(X))Y_0]$, and the marginal benefit compared to the non-targeted policy is defined as $\mathcal{U}_\gamma(\pi) := \mathbb{E}_\gamma[\pi(X)Y_1 + (1-\pi(X))Y_0] - \max\{\mathbb{E}_\gamma[Y_1],\mathbb{E}_\gamma[Y_0]\}$.
Theorem (ref) shows the VCATE provides a welfare bound over the superset of policy classes. Furthermore, any type of restrictions on $\Pi$ such as budget constraints or incentive compatibility will achieve utilitarian welfare gains that are weakly lower than $\frac{1}{2}\sqrt{V_\tau}$. The bound in Theorem (ref) and the generalization in Theorem (ref) provide simple bounds on the prospective gains of targeting, without needing to solve for $\pi(x)$. The bounds are most informative when $V_{\tau}$ is low. For instance, when $V_{\tau} = 0$ there is no heterogeneity explained by the observables $X$ and therefore there are no gains from targeting. However, the fact that the bound is sharp does not imply that it is always achievable for every $\gamma \in \Gamma$, and when $V_{\tau}$ is high it is still be necessary to optimize $\pi(x)$ to determine whether personalized offers are worthwhile.
Finding the bound in Theorem (ref) relies on two important insights. On one hand, the optimal policy in $\Pi$ treats an individual if and only if $\tau(x) \ge 0$ kitagawa2018should. Substituting the optimal policy, $\sup_{\pi \in \Pi} \mathcal{U}_\gamma(\pi)$ is equal to $ \mathbb{E}_\gamma[\max\{\tau_\gamma(X),0\}] - \max\{\mathbb{E}_\gamma[\tau_\gamma(X)],0\}$. On the other hand, to avoid optimizing over all $\gamma \in \Gamma$, I break the problem down into equivalence classes based on the moments of the negative, zero, and positive components of the CATE. I use a constructive approach to derive the most “adversarial” distribution. The upper bound is achieved when the CATE has a binary support, which is partly why the bounds in Theorems (ref) and (ref) have simple closed forms.
Corollary (ref) shows that the welfare bound is invariant to location shifts in the outcome, and grows linearly with scale shifts. This result implies that transformations that change the sign, e.g. $\kappa_2 = -1$, do not change the value of the welfare bound. Consequently, the bound applies regardless of whether the welfare objective is to increase a desirable outcome or to decrease an undesirable outcome.
Consider a simple situation where $X$ is real valued, the treatment $D$ is experimentally assigned with constant probability, and $U$ is a mean zero error term. The researcher runs the following linear regression,
Define the auxiliary quantities $\tau^*(x) := \beta_1 + \beta_2 x$ and $V_x := \mathbb{V}(X)$. The pseudo-VCATE is defined as
The pseudo-VCATE has a close connection to the VCATE. If the linear model describes the conditional mean $\mu_d(x)$, then $\tau^*(x) = \tau(x)$ and $V_{\tau} = V_{\tau}^*$. For instance, in models with binary $X$, the functional form is correctly specified and $V_{\tau} = V_{\tau}^*$. For now, assume that the pseudo-VCATE and the VCATE coincide. In later sections, I analyze models that allow for misspecification.
Consider a sequence of distributions $\{\gamma_n\}_{n=1}^\infty \in \Gamma^{\infty}$. I index the regression coefficients and model variances by the sample size $n$ as $\beta_{2n}$ and $(V_{xn}, V_{\tau n})$, respectively. Define an estimator of the VCATE as $\widehat{V}_{\tau n} = \widehat{\beta}_{2n}^2\widehat{V}_{xn}$, where $\widehat{\beta}_{2n}$ is the least squares estimator of (ref) and $\widehat{V}_{xn} :=\frac{1}{n}\sum_i X_i^2 -\left[\frac{1}{n}\sum_i X_i \right]^2$. With some algebraic manipulations the estimation error can be decomposed as
To derive the asymptotic distribution we can apply the central limit theorem to individual components. For generality, I state joint convergence to a normal distribution as an assumption. This holds as a special case if the observations are i.i.d. and key moments of the distribution are bounded, but may also hold under other forms of dependence. I defer stating primitive conditions until Section (ref).
The normalization by $V_{xn}$ is intended to align with the decomposition in (ref). The $2 \times 2$ matrix $\Omega_n$ is an estimator of the covariance matrix. I present Assumption (ref) as a triangular array because it makes it easier to formalize discussions of uniform coverage over the parameter space. Assumption (ref) allows for cases where $V_{\tau n}^*$ is arbitrarily close to or includes zero. Let $\Omega^{1/2}$ denote the Cholesky decomposition of a matrix $\Omega$. The estimator of the VCATE converges to the empirical process $G$, defined as
where $z \in \mathbb{R}^2$, $\zeta \in \{-1,1\}$, $e_1 = [1,0]'$ and $e_2 = [0,1]'$.
Lemma (ref) shows that the limiting distribution of $\widehat{V}_{\tau n}$ is a linear combination of a Chi-square and a normal, whose weights depend on the value of $V_{\tau n}^*$. The relative magnitude of $V_{\tau n}^*$ determines the fit of the normal approximation. In the heterogeneous case, $V_{\tau n}^* \ge \delta > 0$, $\sqrt{n}G$ converges to a normal as $n \to \infty$ because the first term in (ref) is asymptotically negligible. However, when $V_{\tau n}^* = 0$, only the first term remains and $nG$ converges to a non-central Chi-Square distribution, which is asymmetric. Using normal critical values here (even if everything else was known) would produce distorted coverage. Furthermore, when $V_{\tau n}^* = 0$ the rate of convergence is $n$, which is faster than $\sqrt{n}$, and hence the estimator is “super consistent” near the boundary. The error is dominated by the first stage sampling uncertainty in estimating the nuisance parameter $\beta_{2n}^2$, which converges at $n$ rate.
In practice, all three components in (ref) contribute to the limiting distribution, and this information can be used for inference. I propose an analytic approach based on the quantiles of the empirical process that can deliver exact coverage. Let $F_{n,V_{\tau}^*,\Omega,\zeta}(v)$ be the conditional CDF of the empirical process, defined as
Based on this CDF we can construct a test statistic, $$ F_{n,V_{\tau}^*,\widehat{\Omega}_{n},\zeta}(\widehat{V}_{\tau n}-V_{\tau}^*), $$ indexed by unknown values of $(V_{\tau}^*,\zeta)$ and substituting the estimated covariance matrix $\widehat{\Omega}_{n}$. By construction, the test statistic is contained in $[0,1]$. Similarly, I construct critical values as functions of the parameters for a nominal level $\alpha$, as follows
The difference in the critical values is $(1-\alpha)$ to achieve the desired coverage. The lower critical value is the minimum of the $\alpha/2$ percentile and $0$. This adjustment is meant to increase the power of tests of homogeneity (see Remark (ref)). I propose an adaptive confidence interval by substituting the $\widehat{\Omega}_{n}$, $n$, and $\widehat{V}_{\tau n}$ into the following formula
The set $\widehat{CI}_{\alpha n}$ can be constructed via a grid search between $0$ and an arbitrarily high value, to test whether a particular $V_{\tau}^*$ satisfies the inequality constraints. The procedure achieves correct asymptotic size because the test statistic converges to a uniform random variable in $[0,1]$ for each value of $V_{\tau}^*$. In general, the distribution in (ref) depends on the value of $\zeta$ and I obtain a conservative interval in (ref) by considering the union of intervals with different values of $\zeta$. Moreover, if the off-diagonal element of $\Omega_n$ is zero, then the distribution of the empirical process in (ref) does not depend on the value of $\zeta$. This property is plausible and I introduce primitive conditions that satisfy it in Section (ref). Under those conditions the confidence interval has exact asymptotic coverage .
The procedure is fast because at each point in the grid the researcher evaluates the condition in (ref), using the same estimate of $(\widehat{V}_{\tau n},\widehat{\Omega}_n)$. The critical values can be computed numerically from the quantiles of a generalized Chi-square with distribution $F$, which are available in most statistical software packages.
In this section I provide an overview of the inference problems associated with efficient estimators of the VCATE and how to solve them for the nonlinear/high-dimensional case. Let $\{Y_i,D_i,X_i\}$ be i.i.d.. As shown in newey1990semiparametric, efficient estimators can be decomposed as
where $\varphi_i$ is an i.i.d. realization from the efficient influence function with mean zero, and the residual becomes asymptotically negligible as $n \to \infty$. The semiparametric lower bound is $\mathbb{V}(\varphi_i)$. Let $\eta(\cdot)$ be a set of nuisance functions defined as
levy2021fundamental showed that the efficient influence function for the VCATE is equal to $\varphi_i = \varphi(Y_i,D_i,X_i,\eta) - V_{\tau n}$, where $\varphi$ is defined as
By (ref), all efficient estimators --regardless of their form-- are $\sqrt{n}$--asymptotically equivalent to $\frac{1}{n}\sum_{i=1}^n\varphi_i$. Let $\varphi_{i} = \varphi(Y_i,D_i,X_i,\eta(X_i))$ be a realization of the efficient influence function. If $\mathbb{V}(\varphi_i) > 0$, then $$\mathbb{V}(\varphi_i)^{-1}\times \sqrt{n}(\widehat{V}_{\tau n}-V_{\tau n}) = \mathbb{V}(\varphi_i)^{-1}\left[\frac{1}{\sqrt{n}}\sum_{i=1}^n \varphi_i \right]+ o_p(1) \to^p \mathcal{N}(0,1). $$ In this case, any confidence interval based on normal critical values and a consistent estimator of $\mathbb{V}(\varphi_i)$, produces valid coverage. For common functionals such as the ATE, $\mathbb{V}(\varphi_i) > 0$. However, this cannot be guaranteed for the VCATE.
Both of the inner terms in (ref) are multiplied by $(\tau(x) - \tau)$. When the VCATE is zero, $\tau(x) = \tau$ almost surely, and the influence function is degenerate. The condition that $\mathbb{V}(\varphi_i) > 0$ does not hold uniformly over all $V_{\tau}$ in the parameter space. In this case, the distribution of $\sqrt{n}(\widehat{V}_{\tau n}-V_{\tau n})$ is dominated by the higher order terms of the residual (ref), and the CLT cannot be applied to guarantee normality near the boundary. The linear estimator discussed in the previous section is just one example. Moreover, if the tails of $\tau(x)$ are thin, then the value of $\mathbb{V}(\varphi_i)$ is also small near the boundary.
Corollary (ref) shows that the variance of the efficient influence function is bounded by a quantity that scales up or down proportional to the value of $V_{\tau}$. Consequently, when $\mathbb{V}(\varphi_i)$ is relatively small, the higher order terms in the residual may still dominate.
A robust way to introduce nonlinearity is to consider a regression with real-valued basis functions $M(x)$ and $S(x)$. For now, I will leave these unspecified but in the next section I will show how they can be estimated non-parametrically in a first stage.
with weights $\lambda(X) := [p(X)(1-p(X))]^{-1}$, regressors $W(X,D)$, and parameters $\theta = [c_0,c_1,\beta_1,\beta_2]$. This specification accommodates experiments with heterogeneous assignment probabilities.\footnote{When $p(x) = 1/2$ this produces exactly the same coefficients as a regression of $Y$ on $(1,M(X),D,D\times S(X))$, but differs when the probabilities are heterogeneous. If $M(X) = S(X) = X$ as well, this reduces to (ref).} Consider the following minimizers:
chernozhukov2020generic showed that if $\mathbb{V}(S(X)) > 0 $, $\mathbb{E}[S(X)] = 0$, and the vector $(\beta_1,\beta_2)$ are part of the solution to (ref), then $(\beta_1,\beta_2)$ are also the intercept and slope of the best linear projection of $\tau(X)$ on $S(X)$.\footnote{When $\mathbb{V}(S(X)) = 0$, $\beta_2$ does not have a unique solution in (ref), but $V_{\tau}^* = \beta_2^2\mathbb{V}(S(X)) = 0 \le V_{\tau}$ is still the best linear projection, regardless of the value of $\beta_2$.} Hence the pseudo-VCATE has an upper bound, $V_{\tau}^* = \beta_2^2\mathbb{V}(S(X)) \le V_{\tau}$. Because of this bound, if $V_{\tau} = 0$, then the pseudo-VCATE $(V_{\tau}^*)$ will not falsely detect heterogeneity even if $S(X)$ is misspecified. If anything, poor choices of $S(X)$ will possibly understate the amount of heterogeneity. When $\tau(X)$ is spanned by $S(X)$, $V_{\tau}^* = V_{\tau}$ and the two notions coincide.
To obtain a feasible estimator we define $\widehat{S}(x) := S(x) - \frac{1}{n}\sum_{i=1}^n S(X_i)$, and compute $\widehat{W}(x,d)$ by substituting $\widehat{S}(x)$ in (ref). Now consider a value of $\widehat{\theta}_n$ that minimizes $\frac{1}{n}\sum_{i=1}^n \lambda(X_i)(Y_i-\widehat{W}(X_i,D_i)'\theta)^2$, by solving the first order condition
The regression parameters can be used to construct the CATE and other nuisance functions. For a given $\theta \in \mathbb{R}^4$,
Lemma (ref) shows that the sample variance of the estimated CATE can be interpreted as an estimator that plugs-in (ref) to the efficient influence function in (ref).
Mechanically, the influence function can be decomposed into primary and bias-correction components. As an intermediate step for Lemma (ref), I show that the bias-correction terms and the fourth component of (ref) are proportional to each other. The optimal $\widehat{\theta}_n$ implicitly sets the average bias correction to zero. Intuitively, the linear model minimizes the covariate imbalances between the treatment and control group in-sample. Lemma (ref) suggests that $\widehat{V}_{\tau n}$ could be asymptotically efficient if $\tilde{\eta}_{\widehat{\theta}_n}$ is sufficiently close to $\eta$. In Section (ref), I show that my proposed semiparametric estimator can indeed achieve this.
As a preliminary step, it is necessary to determine which $S(x)$ and $M(x)$ ensure that $\tilde{\eta}_\theta = \eta$. Not all choices achieve this property.\footnote{This point highlights that while all regressions of the form in (ref) proposed by chernozhukov2020generic estimate an interpretable $V_{\tau}^*$ --regardless of the choice of $M(x)$--, not every regression in this class is efficient. The functions $S(x)$, and $M(x)$ in particular, both affect efficiency.} However, if they are chosen in such a way that $W(x,d)'\theta = \mu_d(x)$ for some $\theta \in \mathbb{R}^4$, then that's sufficient to guarantee that $\eta_\theta = \eta$. Lemma (ref) shows that any $\theta$ with this property is also a solution to the regression problem, and provides guidance on the choice of $S(x)$ and $M(x)$.
Lemma (ref) provides efficient choices of $S(x)$ and $M(x)$ that can be expressed in terms of conditional moments, and that for this choice, the optimal $\theta$ has a known, simple form. In practice, $S(x)$ and $M(x)$ can be estimated non-parametrically.
My proposed procedure randomly partitions the observations $\mathcal{I}_n := \{1,\ldots,n\}$ into $K$ folds of equal size $n_k := n/K$. Denote the observations in each fold by $\mathcal{I}_{nk}$, so that $\bigcup_{k=1}^{K} \mathcal{I}_{nk} = \mathcal{I}_n$, and let $\mathcal{I}_{-nk} := \mathcal{I}_n \backslash \mathcal{I}_{nk}$ be the set of observations that are not in fold $k$. In a slight abuse of notation, I use $\mathcal{I}_{-nk}$ when defining conditional expectations, to denote the full set of random variables associated with observations not included in fold $k$. For simplicity, I also label the fold of observation $i \in \mathcal{I}_{nk}$, by $k_i$.
Let $\widehat{\eta}_{-k}(x) := (\widehat{\tau}_{-k}(x),\widehat{\mu}_{0,-k}(x), p(x),\widehat{\tau}_{-k,av})$ denote a prediction of the nuisance function $\eta(x)$ over the set $\mathcal{I}_{-nk}$, using the researcher's preferred prediction algorithm. This could include traditional methods such as linear regression, or more modern “machine learning” approaches such as LASSO, neural networks, or random forests. The only function that is known in advance is the propensity score, since I restrict attention to randomized experiments. Guided by Lemma (ref), define
Consider a regression with weights $\lambda(X_i)$, parameters $\theta := (c_1,c_2,\beta_1,\beta_2)$, and
In practice, $\mathbb{E}[\widehat{\tau}_{-k}(X_i) \mid \mathcal{I}_{-nk}]$ needs to be estimated, and I use a sample analog:
Let $\widehat{\theta}_{nk} = (\widehat{c}_{1nk},\widehat{c}_{2nk},\widehat{\beta}_{1nk},\widehat{\beta}_{2nk})$ be the estimator over the subsample $\mathcal{I}_{nk}$. The fold-specific variance of $\widehat{S}_{-k_i}(x) $ is defined as
The estimator of the VCATE for fold $k$ is
In this case $\widehat{V}_{xnk}$ can be viewed as a preliminary estimate of the VCATE using the data in $\mathcal{I}_{-nk}$, whereas $\widehat{V}_{\tau nk}$ is a regression-adjusted estimator that fits the sample $\mathcal{I}_{nk}$. This adjustment will produce better results, with a pseudo-VCATE interpretation even if the first step $\widehat{S}_{-k_i}(x)$ function is noisy, misspecified, or slow to converge to $\tau(x)$. The estimator in (ref) belongs to the class of multi-step estimators defined in chernozhukov2020generic. I add a restriction on the choice of $M_{-k}(x)$, guided by Lemma (ref), to ensure asymptotic efficiency.
To quantify the uncertainty in $(\widehat{\theta}_{nk},\widehat{V}_{xnk})$ I compute a robust (sandwich) estimator. I start by defining two auxiliary residuals, $\widehat{T}_i := \widehat{V}_{x n k}^{-1}\widehat{S}_{-k_i}(X_i)^2 - 1$ and $\widehat{U}_i := Y_i - \widehat{W}_i'\widehat{\theta}_{nk}$. Let $\widehat{\Pi}_{nk}$ be a $4 \times 4$ diagonal matrix with diagonal entries $(1,1,1,\widehat{V}_{xnk}^{-1/2})$. Researchers can compute estimators of the individual components of the sandwich form $\widehat{J}_{nk}$, $\widehat{H}_{nk}$, and a selection matrix $\Upsilon$ defined as follows
The sandwich covariance estimator is
The population covariance matrix is
When the VCATE is zero, then $\widehat{V}_{xnk}$ (as a consistent estimator of $V_{\tau n}$) should converge to zero along the asymptotic sequence. To prevent asymptotic degeneracy, we need to rescale the estimands along the lines of Assumption (ref). The random variable $V_{xnk}^{-1/2}S_{-k}(X_i)$ is normalized to (conditionally) have variance one by design, even if $S_{-k}(X_i)$ converges to zero. This requires two much weaker conditions: (i) that $V_{xnk} > 0$, i.e. there is some noise in estimating the CATE;\footnote{I also propose an extension that allows for $V_{xnk} = 0$ in Remark (ref).} (ii) $\Omega_{nk}$ has eigenvalues bounded away from zero. To ensure this, the tails of $S_{-k}(X_i)$ need to be thin.\footnote{One sufficient additional restriction is that $\mathbb{E}[U_{i} \mid X_{i},D_{i},\mathcal{I}_{-nk}] = 0$ (the model is correctly specified), $\mathbb{V}(U_{i} \mid X_{i},D_{i} \mid \mathcal{I}_{-nk})$ is bounded away from zero, and $S_{-k}(X_i)$ has bounded kurtosis. In that case the off-diagonal elements of $\Omega_{nk}$ are zero and the diagonals are uniformly bounded. Positive-definiteness may also hold in a neighborhood where the nuisance functions are close to the true value and $\mathbb{E}[U_{i} \mid X_{i},D_{i},\mathcal{I}_{-nk}] \approx 0$.}
Assumption (ref).(i) is a rank condition that ensures that the auxiliary regressor $M_{-k}(X_i)$ is not degenerate. Assumption (ref).(ii) ensures that the second-moment of the candidate regressor $M_{-k}(X_i)$ is bounded. Assumption (ref).(iii) is a standard condition indicating that the fourth moment of the residuals are bounded. Assumption (ref).(iv) is a bounded kurtosis condition indicating that the out-of-sample, machine learning predictions of $\tau(x)$ have thin tails. Finally, Assumption (ref).(v) is a bound on the variance of the first-stage VCATE.
Theorem (ref) shows how these primitive conditions imply an analog of Assumption (ref) for the cross-fitted case.
Theorem (ref) presents a central limit theorem for the components of $\widehat{V}_{\tau nk}^* = \widehat{\beta}_{2nk}^2\widehat{V}_{xnk}$, properly rescaled and conditional on $\mathcal{I}_{-nk}$. This result holds regardless of whether the nuisance parameters are properly specified and primarily relies on the independence of the folds. By Lemma (ref), conditional on $\mathcal{I}_{-nk}$,
Then it is possible to construct adaptive confidence intervals, substituting the sample size $n_k$ and estimated statistics $(\widehat{V}_{\tau nk},\widehat{\Omega}_{nk})$.
The confidence intervals take the same form as in the regression case in (ref), except that now the inputs are obtained from the cross-fitted regression step. The confidence interval is fast to compute because $(\widehat{V}_{\tau nk},\widehat{\Omega}_{nk})$ only needs to be computed once. It is worth noting that because the confidence interval only uses information in fold $\mathcal{I}_{nk}$, the effective sample size is $n_k$. While this does not affect the nominal asymptotic size of the confidence interval, it may affect the power of tests against specific alternatives.
We can construct an “ensemble” to aggregate across folds, defined as follows
In Section (ref), I show that this ensemble estimator is efficient.
So far in this section we have used the data from a single split or fold of the data. However, the choice of fold $k$ or the particular split may lead to different values of $\widehat{V}_{\tau nk}$ and hence distinct confidence intervals. chernozhukov2020generic propose an aggregation procedure based on “median parameter” confidence intervals, inspired by false-discovery rate adjustments. Their proposed conditional t-tests are not directly applicable here because $\widehat{V}_{\tau nk}$ conditionally converges to a generalized Chi-square. However, I show that the basic idea can still be adapted.
Let $K$ be the total number of folds, obtained across one or more splits of the data. For instance, a 2-fold sample with 10 splits would have $K=20$, Let $\inf \widehat{CI}_{\alpha nk}$ and $\sup \widehat{CI}_{\alpha nk}$ denote the lower and upper bounds of $\widehat{CI}_{\alpha nk}$, respectively, and $\text{Med}_K\{\cdots\}$ denote the median over a set indexed by $k = \{1,\ldots,K\}$. If $K$ is even then two quantities might be tied for the median, and in that case I compute their midpoint. The multifold confidence interval is defined as
Intuitively, the $K$ fold-specific intervals “vote” to include a particular value, and $V_{\tau}^* \in \widehat{CI}_{\alpha n}^{\text{multifold}}$ only if there is a majority vote. The “median” interval $\widehat{CI}_{\alpha n}^{\text{multifold}}$ contains values within the median lower bound and the median upper bound across folds. To control the overall false discovery rate, I adjust the nominal size to $\alpha/2$. This adjustment produces a conservative interval because it assumes a worst-case dependence structure between the folds and the splits, regardless of the size of $K$. In some instances, the asymptotic coverage probability may be strictly higher than $(1-\alpha)$, particularly when there is a lot of heterogeneity.\footnote{For instance, given a single split, Theorem (ref) implies that the $\{\sqrt{n_k}(\widehat{V}_{\tau nk}-V_\tau)\}_{k=1}^K$ are asymptotically uncorrelated. However, near the boundary, the estimators converge at a rate faster than $\sqrt{n}$ and their relative dependence structure at that rate is unclear.} At the boundary, with low effect heterogeneity or none at all, it is much harder to asses the dependence structure between the fold-specific estimators. One of the benefits of using a worst-case approach is that it provides coverage guarantees under weak assumptions. Moreover, the empirical example illustrates that even though these intervals are conservative, they may have a short length in practice.
Let $\gamma \in \Gamma$ denote a probability distribution over i.i.d observations $(Y_{1i},Y_{0i},D_i,X_i)$. I use the notation $\mathbb{E}_{\gamma}[\cdot]$ and $\mathbb{P}_{\gamma}(\cdot)$ to denote the expectation and probability under $\gamma$, respectively. Let $S_{-k}(x)$ be the function defined in (ref). The true value of the CATE and VCATE is given by $\tau_\gamma$ and $V_{\tau}(\gamma)$, respectively. The pseudo-VCATE is given by
Define the estimation error of the CATE in the $L_2$ norm as
We can bound the difference between the pseudo-VCATE and its true value:
Theorem (ref) derives a non-asymptotic bound for the VCATE as the minimum of two key quantities: (i) the conditional $L_2$ error between the candidate function and the true CATE, and (ii) the true value of the VCATE. This proof only relies on the definition in (ref). For instance, when $V_{\tau}(\gamma) = 0$, then $V_{\tau}(\gamma)-V_{\tau}^*(\gamma,\mathcal{I}_{-nk}) = 0$, regardless of whether $\widehat{\tau}_{-k}(\cdot)$ is properly specified. The difference between the two quantities is also small if $\omega(\gamma)$ is sufficiently close to zero. In the multi-step approach, $\omega(\gamma)$ captures the first-stage uncertainty from estimating the CATE, which decreases with sample size. I consider the following convergence condition.
Assumption (ref) imposes an $L_2$ consistency condition on the CATE. A large class of machine learning models can meet this requirement. For example, bickel2009simultaneous and belloni2014inference evaluate rates of convergence under sparse models, chen1999improved for neural networks, and wager2015adaptive for regression trees and random forest.
Theorem (ref) shows that multi-step estimators of the VCATE converge to zero faster than $\sqrt{n_k}$ near the boundary. I formalize “near” by considering sequences of distributions where the VCATE approaches zero. Theorem (ref) relies on the non-asymptotic bound in Theorem (ref), the normal approximation in Theorem (ref), and the empirical process in Lemma (ref). There is no requirement on the rate of convergence of $\widehat{\mu}_{-0k}(\cdot)$ (and consequently on the generated regressor $M_{-k_i}(\cdot)$), only an assumption that $p(x)$ is known and that the CATE is estimated at a sufficiently fast rate. Furthermore, if the true CATE is nearly flat in the sense that for $\rho \in [0,1/2)$, then $n_k^{1/2+\rho}V_{\tau}(\gamma_n) = o(1)$ (or even exactly equal to zero), then the estimator has a faster rate guarantee.
To prove efficiency we have the stronger requirement that all the nuisance functions converge to their true value in the $L_4$ norm and at $n_k^{1/4}$ rate in the $L_2$ norm.
The next step is to show that the estimation error of the fold-specific VCATE converges at $\sqrt{n_k}$ to an average of efficient influence functions.
Theorem (ref) shows that the fold-specific estimator converges at $\sqrt{n_k}$-rate to an average of i.i.d influence function. This requires standard regularity conditions. The proof of Theorem (ref) is non-standard due to the multi-step nature of the procedure. I start by applying Lemma (ref), which shows hows to write $\widehat{V}_{\tau nk}$ as an average of estimated influence functions. I break down the proof into sequences where $V_{\tau}(\gamma_n)$ converges to zero and those where it's bounded away from zero. For the first part, I leverage (a) the boundary convergence result in Theorem (ref), (b) the bound for $\mathbb{V}(\varphi_i)$ in Lemma (ref). For the second part, I provide a novel decomposition of regression adjusted nuisance functions. The key is to prove that the regression parameters $\widehat{\theta}_{nk}$ converge at $n_k^{1/4}$ rate to the values in Lemma (ref) for sequences where $V_{\tau}(\gamma_n) \to V_{\tau} > 0$. Once in this form, the rest of the proof relies on a traditional Taylor expansion argument.
The ensemble estimator $\widehat{V}_{\tau n}$ combines information from the whole sample. By definition $n = n_k \times K$ and $K$ is finite, which means that algebraically $\sqrt{n}(\widehat{V}_{\tau n}-V_{\tau}(\gamma_n)) = \frac{\sqrt{n_k}}{\sqrt{K}}\sum_{k=1}^{K}(\widehat{V}_{\tau nk}-V_{\tau}(\gamma_n))$, and by Theorem (ref), $$ \sqrt{n}(\widehat{V}_{\tau n}-V_{\tau}(\gamma_n)) = \left[ \frac{1}{\sqrt{n_k K}}\sum_{k=1}^K\sum_{i \in \mathcal{I}_{nk}} \varphi_i + \frac{1}{\sqrt{K}}\sum_{k=1}^K o_p(1)\right] = \frac{1}{\sqrt{n}}\sum_{i=1}^n \varphi_i + o_p(1). $$ This means that aggregating the estimators restores full efficiency, satisfying the property described in (ref).
I start by showing that the single fold confidence interval has uniform coverage for the pseudo-VCATE, and exact coverage under an additional assumption.
Assumption (ref) states that the product of the pseudo-VCATE and the off-diagonal element of the limiting covariance matrix in (ref) needs to converge to zero uniformly.
Theorem (ref) shows that the confidence intervals always have uniform coverage of the pseudo-VCATE of at least $(1-\alpha)$.\footnote{The theorem only uses Assumptions (ref), (ref), (ref), and (ref) to verify normality in Assumption (ref). A broad class of confidence intervals of the form in (ref) constructed from regression adjusted estimators will satisfy these uniformity properties.} The key is to prove that the confidence intervals yield coverage under arbitrary sequences of distributions, which includes cases where $V_{\tau n}^*$ is either equal to zero or approaches zero as $n \to \infty$. The proof builds on the approximation of Lemma (ref) and shows that for every sequence, the test statistic for a particular $\zeta_{n_k} \in \{-1,1\}$ converges to a uniform distribution. This sequential characterization suffices to apply generic results in andrews2020generic, which guarantee uniform coverage even in non standard cases like this one. Coverage over the pseudo-VCATE holds regardless of whether the nuisance functions are slow to converge or even misspecified.
The intervals are in general conservative because we're not plugging in the unknown $\zeta$, and instead define a robust confidence interval as the union of CIs with given $\zeta \in \{-1,1\}$. However, the key insight is that $\zeta$ only affects the coverage when the pseudo-VCATE is bounded away from zero. If Assumption (ref) holds, the value of $\zeta$ doesn't enter the asymptotic distribution of the estimator. I show that this condition holds automatically if the nuisance functions converge to their true value at a sufficiently fast rate.
As a special case, when the model is correctly specified, i.e. $W_i'\theta = \mu_d(x)$ for some $\theta \in \mathbb{R}^4$, then $\Omega_{nk,12} = 0$ by construction. Lemma (ref) states that we only need a model that is correctly specified asymptotically, given the rates in Assumptions (ref) and (ref). Then for non-boundary cases, $\Omega_{nk}$ converges to the population analog under correct specification. These conditions also imply point-wise coverage of the true VCATE.
Theorem (ref) shows that if the nuisance functions converge at a sufficiently fast rate, then the proposed intervals achieve point-wise exact coverage. The confidence intervals provide correct size coverage for all regions of the parameter space, including $V_{\tau}(\gamma) = 0$.
Proving uniform coverage of the VCATE (rather than the pseudo-VCATE) is more challenging in the non-parametric case without much stronger conditions on the convergence rates of the nuisance functions. The lack of uniformity stems from a difficulty in controlling the ratio $\sqrt{n_k}\omega(\gamma_n) / \sqrt{V_{\tau}(\gamma_n)}$, which measures the relative error in estimating the CATE vs. the overall level of the VCATE. By the bound in (ref), this ratio is easy to control when $n_kV_{\tau}(\gamma_n) = o(1)$ (near homogeneity) or $V_{\tau}(\gamma_n) \to V_\tau > 0$ (strong heterogeneity). However, it is possible to construct sequences, e.g., $n_kV_{\tau}(\gamma_n) \to v > 0$, where $(\widehat{V}_{\tau nk} - V_{\tau}^*(\gamma_n,\mathcal{I}_{-nk}))$ converges to zero at a faster or comparable rate to the error of the pseudo-VCATE. There may be distortions in coverage in smaller samples. I illustrate this issue in the simulations.
Corollary (ref) is empirically relevant for interpreting confidence intervals that do not include zero. It states that the asymptotic probability of having $V_{\tau}(\gamma) \in \left[0,\inf \widehat{CI}_{\alpha nk}\right)$ is uniformly less than $\alpha$. Tests of homogeneity belong to this class and therefore have the correct size when $V_\tau(\gamma) = 0$. Moreover, the result in Corollary (ref) is much stronger because it guarantees that a broader class of one-sided tests also has the correct size. It is important to emphasize that I do not impose any assumptions on rates of convergence of $(\widehat{\eta}_{-k}(x)-\eta(x))$, but only the inequality on the pseudo-VCATE. Consequently, while estimating $\mu(x)$ and $\tau(x)$ may be important for increasing the power of tests of homogeneity, it is not necessary for controlling their size.
The multi-fold confidence interval covers the VCATE asymptotically.
The first part of Theorem (ref) shows that the multifold CI uniformly controls the size of one-sided tests. The second part shows that if the nuisance functions converge to their true value asymptotically, then the multifold confidence interval provides point-wise size-control for two-sided tests. Coverage of the true parameter will be weakly larger that $(1-\alpha)$ asymptotically.
The test of homogeneity has power against local alternatives.
Lemma (ref) computes the power curve for a sequence of local alternatives. When $v = 0$ the power is equal to $\alpha$, whereas when $v \to \infty$ the power tends to one. This shows that tests of homogeneity have local power the null. When the pseudo-VCATE is bounded away from zero, the test rejects with probability approaching one.
I use a simulation design to study the properties of the VCATE estimators. The baseline covariates are distributed as $[X_0,X_1] \in \mathcal{N}(0,\Sigma_x)$, where $\rho = 0.5$ and $$ \Sigma_x =
.$$ The random variables $X_0$ and $X_1$ are standard normal vectors of dimension $J$. The covariance between pairs of components $X_{0j}$ and $X_{1j'}$ is equal to $\rho = 0.5$ for when $j = j'$, but zero otherwise. The outcome is generated from a model where $Y = DY_0 + (1-D)$, $D$ is generated by a Bernoulli draw with probability $0.5$, and
where $c,\tau \in \mathbb{R}$, $\beta_0,\beta_\tau,\kappa_0,\kappa_1 \in \mathbb{R}^p$. The errors $(U_0,U_1)$ are independent of the covariates $(U_0,U_1) \indep (X_0,X_1)$, and distributed as standard normals $[U_0,U_1]' \in \mathcal{N}(0_{2 \times 1},I_2)$. The key model quantities have closed-form expressions. The conditional means at baseline and the CATE are given by $\mu_1(x) = \alpha + \beta_0'x_0$ and $\tau(x) = \tau + \beta_\tau'Z_1$, respectively. The conditional variances are $\sigma_d^2(x) = \kappa_d'x_dx_d'\kappa_d$ for $d \in \{0,1\}$. This formulation incorporates heteroskedasticity. Covariates that influence the outcomes at baseline may also affect the treatment effects.
The regressors are constructed in such a way that $\mathbb{E}[X_dX_d'] = I_p$ for $d \in \{0,1\}$. This implies simple expressions for the variances of the model, $\mathbb{V}(U_d) = \tilde{\sigma}_d^2 + \kappa_d'\kappa_d$, $$ V_{\tau} = \beta_\tau'\beta_\tau, \qquad \mathbb{V}(Y_0) = \beta_0'\beta_0 + \tilde{\sigma}_d^2 + \kappa_0'\kappa_0, $$ $$ \mathbb{V}(Y_1) = \beta_0'\beta_0 + \beta_\tau'\beta_\tau + 2(1-\rho)\beta_0'\beta_\tau + \tilde{\sigma}_1^2 + \kappa_1'\kappa_1, $$
I choose an approximately sparse specification for (ref) where the coefficients decay exponentially at a rate of decay of $\lambda = 0.7$. Let $\ell_j = \sqrt{\left(\frac{1-\lambda}{1-\lambda^J}\right)\lambda^{1-j}}$ be a geometric sequence, which satisfies $\sum_{j=1}^J \left(\frac{1-\lambda}{1-\lambda^J}\right)\lambda^{1-j} = 1$. Given user-specified parameters $(V_{\mu},V_{\tau},\sigma_0^2,\sigma_1^2)$, the coefficients for the entries $j \in \{1,\ldots,J \}$ are determined by $\beta_{0,j} = \ell_j\sqrt{V_{\mu}}$, $\beta_{\tau,j} = \ell_j \sqrt{V_\tau}$, $\kappa_{d,j} = \ell_j \sqrt{\sigma_d^2-\tilde{\sigma}_d^2}$, for $d \in \{0,1\}$. Since $\sum_{j=1}^J \left(\frac{1-\lambda}{1-\lambda^J}\right)\lambda^{1-j} = 1$, then $\beta_0'\beta_0 = V_{\mu}$, $\beta_\tau\beta_\tau = V_{\tau}$, and $\beta_0'\beta_\tau = \sqrt{V_\mu V_\tau}$. We can obtain analogous expressions for the variances of the unobserved components, so that $\tilde{\sigma}_d^2 + \kappa_0'\kappa_0 = \sigma_d^2$ for $d \in \{0,1\}$.
I choose an average effect size of $\tau = 0.15$, that is coherent with the recent meta-analyses of economic experiments in vivalt2015heterogeneous. To make sure that the magnitudes are interpretable, I normalize the coefficients so that the variance for the control group is $\mathbb{V}(Y_0) = 1$, by setting set $c = 1$, $\sigma_d = 0.7$, $\tilde{\sigma}_d = 0.21$, and $V_{\mu} = 0.3$. The design is easy to scale for different values of $V_{\tau}$ and $J$. My design is similar to that in belloni2014inference but I choose $\Sigma_x$ and the sparsity structure in such a way that $V_{\tau}$ has a closed form expression. I use LASSO to estimate $\mu_1(x)$ and $\mu_0(x)$, tuned via cross-validation. The coefficients of this model are consistent given this sparse linear structure, even in high dimensions. I randomly simulate 2000 datasets to compute each of the estimators, and split them into $K=2$ folds.
Figure (ref) considers a simulation with $n = 2500$. The figure displays a density plot for the multi-step estimator, $\widehat{V}_{\tau n}$ defined in (ref), and a two-step debiased machine learning estimator computed as:
where $\widehat{\eta}_{-k}(x) = (\widehat{\tau}_{-k}(x),\widehat{\mu}_{0,-k}(x),p(x),\widehat{\tau}_{n,av})$ and $\widehat{\tau}_{n,av} = \frac{1}{n}\sum_{i=1}^n \widehat{\tau}_{-k_i}(X_i)$. When $V_{\tau} > 0$ the efficient influence function is non-degenerate. In high-heterogeneity regimes both converge to the same limiting distribution.\footnote{This is shown in Theorem (ref) for the multistep approach and can be shown for the two-step using standard arguments, e.g. chernozhukov2018double.} However, when $V_{\tau} = 0$, the influence function is degenerate and they may converge at different rates. We see that the multi-step approach is much more precise. This can be explained by the fast boundary convergence rates derived in Theorem (ref). The two step approach can also produce negative estimates of $V_{\tau}$, which is an undesirable feature, whereas the multi-step estimator is always non-negative. Both estimators have higher bias when the dimension increases because there is more first-stage noise.
Figure (ref) plots the root mean-square error (RMSE) of $\widehat{V}_{\tau n}$ and $\widehat{V}_{\tau n}^{two-step}$ for different sample sizes. I compute the semiparametric efficiency bound by computing the RMSE of an “oracle” estimator that substitutes $\widehat{\eta}_{-k}(X_i) = \eta(X_i)$ in (ref). The results show that as the sample size increases, both estimators achieve a higher level of accuracy and their variance approaches the semi-parametric lower bound (the RMSE of the oracle). As expected by Corollary (ref), the semiparametric lower bound is zero at the boundary. The differences in RMSE shorten with higher $V_{\tau}$ and in lower dimensional settings $(2J = 10)$ (which have lower first-stage noise).
Figure (ref) shows the coverage of $V_{\tau}$ for the different proposed confidence intervals (CIs). For the multi-step approach, I consider the single splits CIs in (ref) and the conservative multi-fold CIs from and (ref). The two-step CIs are constructed as $\frac{1}{n}\sum_{i=1}^n \widehat{\varphi}_i \pm 1.96\sqrt{\widehat{V}_{\varphi}/n}$, where $\widehat{\varphi}_i$ is the summand in (ref) and $\widehat{V}_{\varphi}$ is an estimate of its sample variance. The coverage of the two-step approach is very low under homogeneity, and there is no improvement as sample size increases when $V_{\tau} = 0$. The coverage of the two-step estimator only improves with higher $n$, in high heterogeneity designs. By contrast, both multi-step approaches cover the parameter at the intended level, and coverage improves with higher sample size. For fixed $n$, coverage degrades for both cases when the number of covariates is higher.
Figure (ref) explores the differences in covering the VCATE vs the pseudo-VCATE when $n = 2500$ and $2J = 10$ for a fine-grained set of values of $V_{\tau}$. Panel (a) reflects a dip in coverage close to the boundary. My theory predicts that the multistep CIs have exact coverage when $V_{\tau} = 0$, but may not cover uniformly close to the boundary (see discussion after Theorem (ref)). The mulit-fold CIs have conservative coverage. Conversely, Figure (ref), Panel (b) shows the multi-step CIs always uniformly cover the pseudo-VCATE, as predicted by theory. This provides a robustness guarantee for how to interpret the CIs. The two-step approach has much lower coverage and no guarantees when $V_{\tau} = 0$ in either panel.
Figure (ref).(left) shows the power of tests of homogeneity in a simulation with $n = 2500$. The multi-step, single fold approach has correct size control and has local power, in line with the result of Lemma (ref). The power of the test using the multifold approach is similar to using a single fold. The right panel shows the probability that the VCATE is strictly below the CI bounds. As predicted by theory, this probability is uniformly bounded by $\alpha = 0.05$ for the single and multifold approaches (see Corollaries (ref) and (ref), and Theorem (ref)). The two-step approach has a non-monotonic power curve with incorrect size. The size of one-sided tests in the right panel is uniformly bounded by $\alpha = 0.05$, though this may be partly the fact that $\widehat{V}_{\tau n}^{two-step}$ can take negative values and has a negative bias (see Figure (ref)).
In this section, I illustrate my approach using data from a large-scale information experiment conducted by dizon2019parents. The study, which covered 39 school districts, involved an intervention to provide low-income parents of at least two children with information about their children's school performance. Half the households where assigned to the information intervention and the rest were assigned to the control group. dizon2019parents showed that, at baseline, parents faced large information gaps regarding their children's grades and class ranking. Even though schools produced a report card, 60% of parents were unaware of their child's performance. Many parents reported that they did not receive the report card (children either lost them or did not take them home), or had trouble interpreting the report card structure, primarily due to low literacy levels.
The intervention was designed to present details of their children's school performance in an easily accessible way. dizon2019parents showed that the information gaps (the difference between believed and true test scores) went down as a result of the intervention, and the amount of updating varied depending on students' initial test scores. dizon2019parents also introduced a real-stakes scenario where parents received a series of lottery tickets for a scholarship paying for four years of high school. Parents had to decide how to allocate tickets between two siblings. If there were more than two siblings residing in the household, the survey team selected two at random. The results showed that parents allocated tickets towards their better performing child.
To test for heterogeneity, dizon2019parents ran a linear regression of parental beliefs on initial scores, treatment, and an interaction as in (ref), and reported estimates $\bar{X} = 46.8$ (on a scale of 100) and $(\widehat{\beta}_1,\widehat{\beta}_2) = (-25.9,0.40)$ in their Tables 1 and 2, respectively. The coefficient $\widehat{\beta}_2$ captures how much the treatment effects vary (on a $0$ to $100$ scale) for a each additional point in students initial scores. The VCATE combines information about the coefficient and the initial variability in scores. From the data, I also estimate the variance of the control $\widehat{V}_x = 305.53$, and estimate the VCATE as $\widehat{\beta}_2^2\widehat{V}_x = 48.89$. Taking the square root and normalizing by the standard deviation of the outcome for the control group ($\widehat{V}_{(Y \mid D= 0)} = 311.72$), produces $0.40$. This means that the magnitude of treatment effect heterogeneity explained by scores is comparable to 40% of the standard deviation of beliefs in the control group.
The VCATE can also help us understand the magnitude of treatment effect heterogeneity in the experiment using multiple covariates. I use LASSO for the first-stage predictions, two folds per split with 20 splits, and estimate $\{\widehat{\Omega}_{nk}\}_{k=1}^K$ using clustered standard errors at the household level and the formulas for the confidence intervals defined in (ref). The results are not very sensitive to the number of splits. For ease of exposition, I report point-estimates and CIs for $\sqrt{\frac{V_\tau}{\mathbb{V}(Y_0)}}$, which is the standard deviation of the CATE divided by the standard deviation of the outcome for the control group. I compute confidence intervals for the square root via the transformation proposed in (ref).
Table (ref) computes the ATE and the $\sqrt{VCATE}$ for two outcomes (parental beliefs and lottery allocations) and 8 different sets of covariates. Panel (a) shows that, on average, parents downgrade their beliefs about test scores by 42% of the standard deviation (SD) of the beliefs of the control group. The treatment effect heterogeneity explained by test scores is equivalent to 40% of the standard deviation (SD) of the beliefs of the control group. This is statistically significant at the 5% level and has a comparable magnitude to the ATE. The confidence intervals are relatively short in length. However, applying the bounds from Theorem (ref) and Corollary (ref) shows that differentiating treatment offers based on scores could further lower beliefs by at most 8.1% SDs of the beliefs in the control group. In this case the ATE is already fairly high compared to $\sqrt{V_{\tau}}$, so in spite of the large heterogeneity, the marginals gains from targeting would be modest. Panel (b) presents the results for the secondary school lottery. The ATE is estimated precisely at zero, because the lottery tickets had to be divided as a zero sum between the siblings. The VCATE measures how much the dispersion in the allocation depends on the covariates. The standard deviation of the VCATE explained by initial scores is 16% of the SD of the control group lottery allocation, and the maximum welfare gains are around 7.8% SD.
The student variables (grade, age, gender, attendance, and educational expenditures) collectively explain 11% of the SD of parental beliefs. This is statistically significant at the 5% level. The magnitude is around a fourth of the variation for test scores, and the maximum welfare gain from targeting is 0.7% of the SD of parental beliefs in the control group. I find that other subsets of covariates do not produce statistically significant estimates of the VCATE at the 5% level. The added welfare of personalizing treatment assignment using these covariates is also very low.
The estimates that use all the covariates are computed over a smaller subsample with non-missing values across all variables. Despite the large number of variables and the smaller sample, the estimates of the VCATE remain relatively stable across specifications. The VCATE computed from a rich set of respondent, household, and student covariates has a comparable magnitude to the VCATE that only includes student scores. The confidence intervals are also similar. The estimates of the maximum welfare gains from targeting using all covariates are 7.9% SD for beliefs and 7.4% SD for lottery outcomes, respectively, which are similar to the welfare gains computed using only scores.
I propose an efficient estimator of the variance of treatment effects that can be attributed to baseline characteristics and propose novel adaptive confidence intervals that produce valid coverage. I analyze issues of non-standard inference that arise in this context, and how to address them. I also explore the economic significance of the VCATE for policymakers and researchers, by showing that the $\sqrt{VCATE}/2$ bounds the marginal gains of targeted policies. Overall, this paper proposes a broadly applicable approach to measure treatment effect heterogeneity in experiments.
\counterwithin{assump}{section}