Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,502 characters · 9 sections · 98 citation commands
Predicting the Distribution of Treatment Effects via Covariate-Adjustment, with an Application to Microcredit
Keywords: Distribution of treatment effects, heterogeneous treatment effects, machine learning, causal inference, microcredit.
JEL classification codes: C12, C14, O12.
\sloppy Important questions in impact evaluation cannot be answered by average treatment effects alone. What proportion of individuals are harmed by a treatment? Does a policy help many by a little? Or a few by a lot? The distribution of treatment effects carries crucial information with important equity and efficiency implications. For example, if a policy with positive average effect is found to harm a significant portion of the population, it may be crucial to prioritize mitigating the harm or reconsider its implementation. If a few individuals have large benefits, it may be important to understand why and how to efficiently target the policy recipients. However, learning the distribution of treatment effects is challenging because of the fundamental problem of causal inference: one can never observe individual counterfactuals. But one may be able to predict them.
The main contribution of this paper is to provide a new approach to inference on points of the distribution of treatment effects, $$\theta = P (Y(1) - Y(0) \le \delta),$$ for any value of $\delta$, by leveraging covariate-adjustment and predictive models. The setting assumes a binary treatment $D$ and that the researcher has access to a representative sample of $Y(1)$ and $Y(0)$, the potential outcomes under treatment and no treatment, respectively, and pre-treatment covariates $X$.\footnote{In Online Appendix (ref), I provide an alternative setting with a known propensity score.} This is typically the setting in Randomized Controlled Trials (RCTs). I denote $\theta$ by Distributional Treatment Effect (DTE) and the method by CAIDE (Covariate Adjustment for Inference on Distributional Effects). All results are pointwise in $\delta$.
Relative to previous literature, my first contribution is to provide a new characterization of the identified set of $\theta$ which is more amenable to inference. Based on this result, I develop two novel methods for inference. The first is valid in finite samples and requires arguably weak conditions. For example, the conditions are satisfied if the data come from an RCT. To my knowledge, this is the first finite-sample inference result on the DTE that incorporates covariates. The second inference method is asymptotically valid and more powerful than the finite-sample approach. Both inference methods rely on estimating a covariate-adjustment term, and I show they are robust to its misspecification, in the sense that a misspecified term leads to a valid identified set, although wider. If the covariate-adjustment term is correctly specified, the identified set is sharp. As a consequence of this robustness, the covariate-adjustment term can be estimated using any method, including machine learning algorithms. The relevance of my contribution is demonstrated in an application to microcredit, where I find important distributional impacts, specifically statistically significant bounds on the proportion of individuals helped and hurt by increased credit access. While academics and policymakers have hypothesized and debated the theory underlying such impacts, their empirical estimation has been hampered by methodological limitations.
My results can be contrasted with those in fan2010sharp, who characterized the sharp identified set of $\theta$ in the presence of covariates. Their approach to inference focuses on the case without covariates since it is challenging to implement inference based on their characterization of the identified set, as discussed in Remark (ref). In contrast, my characterization of the sharp identified set is more amenable to inference and satisfies the robustness property mentioned above.
Asymptotic inference on $\theta$ in the presence of covariates was studied by two recent and independent working papers, semenova2023adaptive and ji2023model. Both papers consider the more general problem of learning averages of intersection bounds, which includes the DTE as a particular case. Their approach is based on the dual of an optimization problem, and the confidence intervals of both papers coincide for the DTE. By leveraging the specific structure of the problem of learning the DTE, my main advantage relative to these alternatives is that CAIDE always estimates bounds that are (weakly) narrower, leading to more power under suitable conditions (Appendix (ref)). In practice, this difference can lead to different conclusions regarding statistical significance, as shown in a Monte Carlo study in Section (ref) and in the application in Section (ref). Another difference relative to the two papers is that I allow for using the proportion of treated units in-sample instead of the true propensity score, which further reduces the confidence interval length (Appendix (ref)). Finally, I show asymptotic exactness under continuous outcomes, while ji2023model restrict the support of the outcomes to a finite set to derive exactness. semenova2023adaptive and ji2023model do not provide finite-sample inference.
To understand my approach, consider a scenario where a researcher aims to predict $Y(0)$ using a function of covariates $s(X)$. An example of such a function is the conditional expectation, $s(x) = \mathbb{E}[Y(0) | X = x]$. Ideally, $s(X)$ would be close to $Y(0)$, and the distribution of $Y(1) - s(X)$ would approximate the distribution of treatment effects. However, $s(X)$ is generally an imperfect predictor for two main reasons. First, because $Y(0)$ is rarely a deterministic function of observable covariates, so one can typically only approximately predict it. Second, because $s(X)$ needs to be calculated from the data, and is subject to random sampling error. CAIDE is designed to address both of these challenges. In this example, the idea is that even if $s(X)$ is an imperfect predictor, examining the distribution of $Y(0) - s(X)$ (the prediction error) can help correct for the difference between $Y(1) - s(X)$ and the actual treatment effect. Intuitively, if $|Y(0) - s(X)|$ is small (e.g., less than 0.1), the error between $Y(1) - s(X)$ and the treatment effect is also small (at most $0.1$). In general, correcting for prediction error using the entire distribution of $Y(0) - s(X)$ provides more information than simply using an upper bound.
The intuition above does not rely on the form of $s$. In fact, by combining the information from the distributions of both $Y(1) - s(X)$ and $Y(0) - s(X)$, any function $s$ can be used to learn $\theta$. However, when correcting for prediction error, different choices of $s$ can lead to different approximations of the distribution of treatment effects. CAIDE proceeds by estimating an optimal form of $s$, defined in Theorem (ref), which depends on the distributions of both potential outcomes conditional on covariates --- that is, it is not the prediction of just one potential outcome. In this case, the function $s$ can be interpreted as a covariate-adjustment term, since the same function is subtracted from both $Y(1)$ and $Y(0)$. In Section (ref), I show that the optimal function is different for learning the lower or upper bound on $\theta$, and the procedure can be repeated with different functions $s_L^*$ and $s_U^*$ when one is interested in two-sided inference.
The fact that any function $s$ is valid motivates a sample-splitting strategy: one portion of the sample is used to estimate $s^*$ ($s^*_L$ or $s^*_U$) with some estimator $\hat{s}$ using any machine learning algorithm; in the remaining sample, I calculate a confidence interval for $\theta$ treating $\hat{s}$ as fixed. Once this function is fixed, estimation becomes a low-dimensional problem, similar to fan2010sharp, and valid inference can be achieved regardless of the statistical properties of $\hat{s}$. In fact, finite-sample inference on $\theta$ (Theorem (ref)) is valid when data come from an RCT, without requiring further regularity conditions. Moreover, cross-fitting, a more efficient way of sample-splitting, can be used for asymptotic inference under suitable conditions (Theorem (ref)). In both cases, if $\hat{s}$ is a poor estimate of $s^*$, the resulting confidence interval is wider but still valid compared to an estimator $\hat{s}$ that is close to $s^*$. I compare both approaches to inference in Section (ref) through a Monte Carlo study.
I apply CAIDE to study heterogeneous effects of microcredit (Section (ref)). Microcredit has reached more than 170 million borrowers worldwide in 2022 Convergences2023ImpactFinance and has been the subject of hundreds of academic papers, including several RCTs. Theory suggests that microborrowing can be either beneficial or detrimental (banerjee2013microcredit, garz2021consumer), and most empirical evidence indicates that average effects are in general mild if not negligible (banerjee2015six, meager2019understanding). Analyses of heterogeneous effects have mostly been limited to Quantile Treatment Effects (QTEs), suggesting mainly positive impacts for some and rarely negative impacts meager2022aggregating. I revisit five studies that randomized microloans at the individual level and find evidence of heterogeneity, including negative treatment effects for some households. The estimated lower bound on the proportion of individuals who would be better off without the loan varies from $2.8\%$ to $26.0\%$, and the lower bound on the proportion who benefited ranges from $2.9\%$ to $26.1\%$.
The paper proceeds as follows. Section (ref) reviews the related literature. Section (ref) describes the setup and shows the main identification results of the paper. Section (ref) provides finite-sample inference, and Section (ref) asymptotic inference and exactness using cross-fitting. Section (ref) presents a Monte Carlo study, and Section (ref) the application to microcredit. Finally, Section (ref) concludes. Proofs not displayed in the main text are deferred to the Appendix.
Early work on bounding the distribution of treatment effects are heckman1997making and manski1997monotone, who propose assumptions that restrict the joint distribution of potential outcomes in order to achieve informative bounds on $\theta$. makarov1982estimates and williamson1990probabilistic characterize the sharp bounds on $\theta$ given only knowledge of the marginal distributions of $Y(1)$ and $Y(0)$.\footnote{Both papers consider the generic problem of bounding the distribution of the sum of two random variables given only knowledge of their marginals.} fan2010sharp use these bounds to provide asymptotic inference on $\theta$, and ruiz2022non provide finite-sample inference.
fan2010sharp also characterize sharp bounds on $\theta$ in the presence of pre-treatment covariates, but they do not approach inference based on these bounds. It is challenging to provide inference from their characterization, as discussed in Remark (ref). In contrast, by incorporating covariates via a covariate-adjustment term, the characterization of the sharp bounds that I provide is amenable to inference.
kallus2022s studies the sharp bounds in the presence of covariates in the case of binary outcomes and proposes inference under high-level assumptions on the data-generating process. Beyond the binary case, two recent and independent working papers studied inference using the sharp bounds on $\theta$ in the presence of covariates, semenova2023adaptive and ji2023model. The two papers study the more general problem of averages of intersection bounds, which includes $\theta$ as a particular case. semenova2023adaptive provides inference when the conditional distributions of potential outcomes can be estimated uniformly at the $n^{-1/4}$ rate and potential outcomes are discrete. ji2023model uses optimal transport theory to show that solving the dual leads to valid bounds even if conditional distributions are misspecified. In contrast, my approach is based on covariate adjustment. The main differences are that I provide finite-sample inference, asymptotic inference that is more powerful under suitable conditions (Appendix (ref)), a further reduction in the length of confidence intervals by using the proportion of treated units instead of the propensity score (Appendix (ref)), and asymptotic exactness under continuous outcomes.
This paper also contributes to the literature on using machine learning for estimating heterogeneous treatment effects (athey2016recursive; chernozhukov2018double; wager2018estimation; athey2019generalized; yadlowsky2021evaluating; knaus2022heterogeneous; athey2023machine; chernozhukov2023fisherschultz). Most papers in this literature focus on estimating the point-identified conditional average treatment effect (CATE) and use the estimated CATE to predict individual treatment effects. CAIDE is complementary to these approaches. I focus on the partially-identified DTE $\theta$, which answers a different type of question. For example, if the CATE is zero for all values of $X$, many individuals could still be harmed, and many could benefit. In this case, the DTE could reveal, for example, the proportion of individuals who are harmed and who benefit.
This paper is also related to the literature on conformal prediction (vovk2005algorithmic; lei2018distribution). lei2021conformal propose a prediction interval to the individual treatment effect, i.e., an interval to cover the random variable $Y(1) - Y(0)$ for a specific individual with a certain probability. In contrast, I focus on the population parameter $\theta = P (Y(1) - Y(0) \le \delta)$. Also related is the literature on covariate adjustment in the context of estimating the Average Treatment Effect (ATE) in RCTs (e.g., lin2013agnostic, wager2016high, rothe2018flexible, wu2018loop; review in van2023use; comparisons in kahan2016comparison, tackney2023comparison). Most of this literature focuses on adjusting outcomes for covariates to improve precision. I use covariate adjustment to improve identification (i.e., to get narrower bounds) of a different estimand, the DTE.
I add to the development economics literature on evaluating the impacts of microcredit. Metanalyses suggest that average impacts are small if not negligible\footnote{breza2021measuring find negative effects of the shutdown of microfinance institutions in Andhra Pradesh, which suggests that microcredit might have large and positive equilibrium effects.} (banerjee2015six, meager2019understanding), however, there is growing evidence that impact is heterogeneous (banerjee2019can, bryan2021big, meager2022aggregating, cueva2024microfinance). Analysis of heterogeneity suggested mostly positive impact in the upper tail of income/profits distributions (angelucci2015microcredit, augsburg2015impacts, banerjee2015miracle, meager2022aggregating). Yet, evidence of negative effects has been scarce\footnote{crepon2015estimating, for example, finds negative quantile treatment effect on profits at the 10th percentile. However, they interpret the evidence as a “reduced form”: “they do not necessarily mean that the impact of getting credit itself has the same heterogeneity”, since take-up is low (randomization was conducted at the village level) and “we do not know where the compliers lie in the distribution of outcomes”. They also suggest the negative effect “might be partially due to long-term investments misclassified as current expenses”, supported by a positive estimated QTE on consumption at the 10th percentile.}. For example, through a metanalysis of seven studies, meager2022aggregating finds “no generalizable negative quantile treatment effects”. I revisit five RCTs in microcredit and find novel evidence of both positive and negative effects.
In this section, I define the setup of the paper and present the identification results of $\theta = P ( Y(1) - Y(0) \le \delta )$ that motivate the confidence intervals proposed in Sections (ref) and (ref). There are two sources of ambiguity when learning $\theta$ from a sample and predicted counterfactuals. The first is the usual statistical uncertainty due to random sampling error, the usual uncertainty in estimating a population parameter from a sample. The second is due to the prediction of the counterfactuals. Even with a large sample size, the predicted counterfactuals can generally approximate the true counterfactuals, but they do so with error. The consequence of this ambiguity is that, even with access to population data, in general, one can at best find an interval (or bounds) that contain $\theta$ (see, e.g., heckman1997making; manski1997monotone).
This section lays out the strategy I use to deal with the latter. The first result is a new identifying equation for $\theta$ in ((ref)), that is, an interval, denoted identified set, that contains $\theta$ expressed in terms of the distribution of the observable data. This equation incorporates the information on predicted counterfactuals through a scalar function $s(x)$, which acts as a covariate-adjustment term. The main benefit of ((ref)) is that the interval contains $\theta$ for any function $s$. That is, ((ref)) might be an outer identified set, but it is always a valid identified set. This property is crucial so that the confidence intervals proposed in Sections (ref) and (ref) are valid regardless of the statistical properties of how $s$ is estimated.
The second result is Theorem (ref), which characterizes optimal forms of $s$ that make the identified set in ((ref)) sharp, that is, the narrowest possible interval that has the guarantee to contain $\theta$. As it turns out, in general, no single choice of $s$ makes both lower and upper bounds sharp; instead, each bound requires a different function $s$, denoted $s^*_L$ and $s^*_U$ respectively for the lower and upper bounds. The sharp identified set in Theorem (ref) is an alternative characterization to the one provided in fan2010sharp. I argue in Remark (ref) that my representation leads to two major advantages for inference. Most importantly, I argue that inference using my characterization is tractable under arguably mild conditions (Theorems (ref) and (ref)) when the bounds are estimated nonparametrically. In contrast, I show that it is hard to characterize the asymptotic distribution of the plug-in estimator based on fan2010sharp's characterization, even when using sample-splitting, when the bounds are estimated nonparametrically. Note that fan2010sharp do not address inference on $\theta$ in the presence of covariates.
Finally, in subsection (ref) I discuss cases when $\theta$ is point-identified, which happens when covariates are informative enough to predict the sign of the difference of counterfactuals.
Without loss of generality and to simplify notation, I focus on $\delta = 0$. For a different value of $\delta$, all results hold by redefining $\tilde{Y}(0) = Y(0) + \delta$. Consider the standard potential outcomes setup, where $Y \in \mathcal{Y} \subseteq \mathbb{R}$ is the outcome, $D \in \{0,1\}$ is the treatment assignment indicator, $Y(1)$ and $Y(0)$ are potential outcomes connected by $Y = D Y(1) + (1 - D) Y(0)$, and $X \in \mathcal{X}$ denotes a set of pre-treatment covariates. The setting is formalized in Assumption (ref).
Assumption A(ref) of bounded $\mathcal{Y}$ is technical, and is discussed in Remark (ref). A(ref) is the usual unconfoundedness assumption, and it is used for identification of the marginals of $Y(1)$ and $Y(0)$. It typically holds, for example, in Randomized Controlled Trials (RCTs), where $D$ is randomized. The constant propensity score assumption facilitates exposition. In the asymptotic analysis, this assumption can be relaxed to allow for an arbitrary, known propensity score, as demonstrated in Online Appendix (ref). Extending my results to accommodate unknown propensity scores that must be estimated from the data is an interesting avenue for future research.
The starting point for CAIDE is the fact that for any measurable function $s: \mathcal{X} \rightarrow \mathcal{Y}$, it holds that
Therefore, I propose to bound $\theta$ using the marginal distributions of covariate-adjusted outcomes $Y(1) - s(X)$ and $Y(0) - s(X)$. This strategy is in contrast to fan2010sharp, which considered the marginals of $Y(1)$ and $Y(0)$ for inference. A major advantage of my approach is that it is flexible in the covariate adjustment term, in the sense that any $s$ induces bounds that contain $\theta$. Such flexibility is crucial to ensure that the estimation of $s$, regardless of the method or number of covariates used, leads to valid inference in Sections (ref) and (ref). The Makarov bounds (makarov1982estimates; williamson1990probabilistic), used in fan2010sharp, always contain $\theta$:
where $F_j(t) = P ( Y(j) \le t )$, $j \in \{0,1\}$. Hence, by replacing the random variables $Y(1)$ and $Y(0)$ by $Y(1) - s(X)$ and $Y(0) - s(X)$, the bounds induced by $s$ contain $\theta$:
where
Because of the equivalence in ((ref)), and the validity of the Makarov bounds ((ref)), ((ref)) contains $\theta$ for any $s$. If covariates are not used, e.g. if $s(x)=0$, then ((ref)) is equivalent to ((ref)). Different functions $s$ lead to different bounds, with varying lengths. For specific functions defined in Theorem (ref), the induced bounds are sharp, that is, the smallest possible valid bounds given the information from the data. fan2010sharp derived one characterization for these bounds:
where $F_j(t | x) = P ( Y(j) \le t | X = x )$, $j=0,1$. In Remark (ref), I compare ((ref)) and Theorem (ref).
The result of Theorem (ref) is new. It shows that the sharp bounds on $\theta$ in the presence of covariates can be obtained via covariate adjustment. That is, for specific functions $s^*_L(x)$ and $s^*_U(x)$, the induced bounds in ((ref)) are equal to the sharp bounds in ((ref)). Note that $s^*_L(x)$ and $s^*_U(x)$ need not be unique and are generally different. Hence, in general, there is no single function $s$ that induces both lower and upper bounds to be sharp in ((ref)), and in Sections (ref) and (ref) I propose to estimate both $s^*_L(x)$ and $s^*_U(x)$.
From equation ((ref)), a sufficient condition for point identification of $\theta$ is that $\max_t \{ F_1(t | x) - F_0(t | x) \} = 1 + \min_t \{ F_1(t | x) - F_0(t | x) \}$ for all $x$. Figure (ref) gives intuition for when this condition holds, focusing on a single value of $X = x$ for ease of exposition. All panels show the conditional distributions of $Y(1)$ (in blue) and $Y(0)$ (in red), which represent the prediction of the potential outcomes given $X = x$. The upper panels show the cumulative distribution and the lower panels show the probability density. Note that predictions are not a single point since potential outcomes are random variables even conditional on $X = x$. In both scenarios $\min_t \{ F_1(t | x) - F_0(t | x) \} = 0$, for example for $t = 0$, and $F_1(t | x) - F_0(t | x)$ is maximized at $t = 0.75$ (green dashed line).
In the lower left panel, the distributions of $Y(1)$ and $Y(0)$ are partially overlapping, and $\max_t \{ F_1(t | x) - F_0(t | x) \} < 1 + \min_t \{ F_1(t | x) - F_0(t | x) \} = 1$ (upper left panel). Since there is overlap, one cannot accurately predict if $Y(1) - Y(0) < 0$ or not. For example, if an individual is at the right tail of $Y(1)$ (blue) but at the left tail of $Y(0)$ (red), then $Y(1) - Y(0) > 0$. If the opposite happens, then $Y(1) - Y(0) < 0$. However, since the overlap is small, the probability that $Y(1) - Y(0) < 0$ is large, and the fraction harmed (for this value of $X=x$) is close to point identified. In the lower right panel, the distributions are disjoint, and $\max_t \{ F_1(t | x) - F_0(t | x) \} = 1 + \min_t \{ F_1(t | x) - F_0(t | x) \} = 1$ (upper right panel). In this case, even though the covariates are not informative enough to get a point prediction of $Y(1)$ and $Y(0)$, it is informative enough to learn that $Y(1) < Y(0)$ for this value of $X = x$, and thus all individuals with $X = x$ are harmed.
Note that the interval in ((ref)) is an average of $\max_t \{ F_1(t | x) - F_0(t | x) \}$ and $1 + \min_t \{ F_1(t | x) - F_0(t | x) \}$ over all values of $X$. Hence, the figure illustrates that the informativeness of covariates in predicting the outcomes is crucial for having a narrow interval for $\theta$. In fact, if the predictions (conditional distributions) of $Y(1)$ and $Y(0)$ are disjoint for all values of $X$, then $\theta$ is point-identified. Note that in Figure (ref), the green line represents the value of $s^*_L(x)$, which is the point $t$ that maximizes the distance $F_1(t | x) - F_0(t | x)$. Hence, $s^*_L(x)$ is the point that makes the two distributions as separated as possible. This fact is what motivates the notation $s$, for separator function.
In this section, I propose sample-splitting estimators for $(\theta_L^*, \theta_U^*)$ and a confidence interval (CI) for $\theta$ that provides a finite sample guarantee of coverage. Given an iid sample $(Y_i, D_i, X_i)_{i=1}^n \sim P$ that satisfies A(ref), the estimation strategy consists of randomly splitting the sample into two, using the first set to estimate $(s^*_L, s^*_U)$ with $(\hat{s}_L, \hat{s}_U)$, and the second set to estimate the bounds $[\max_t \Delta(t, \hat{s}_L), 1 + \min_t \Delta(t, \hat{s}_U)]$ (as in (ref)) with a sample analogue. Conditional on the first sample, these bounds are valid due to ((ref)), and finite sample inference is achieved by applying the DKW inequality dvoretzky1956asymptotic, massart1990tight.
I propose sample-splitting estimators for the sharp bounds $(\theta_L^*, \theta_U^*)$:
$F_1(t | X)$ and $F_0(t | X)$ can be estimated with a variety of methods, including machine learning algorithms or other nonparametric estimators (see, e.g., kneib2023rage for a review). I discuss guidelines for estimating $\hat{s}$ in Appendix (ref). Although any method can be used to estimate $F_1(t | X)$ and $F_0(t | X)$, estimates $(\hat{s}_L, \hat{s}_U)$ closer to $(s^*_L, s^*_U)$ yield bounds $[\max_t \Delta(t, \hat{s}_L), 1 + \min_t \Delta(t, \hat{s}_U)]$ that are closer to the sharp $[\max_t \Delta(t, s^*_L), 1 + \min_t \Delta(t, s^*_U)]$, and poor estimators can lead to wider intervals.
A default choice is to make the main and auxiliary samples roughly the same size, but that is unnecessary. In general, a larger auxiliary sample helps estimate $(\hat{s}_L, \hat{s}_U)$ closer to $(s^*_L, s^*_U)$, which improves the length of the target bounds $[\max_t \Delta(t, \hat{s}_L), 1 + \min_t \Delta(t, \hat{s}_U)]$. On the other hand, a larger auxiliary sample implies a smaller main sample, which leads to larger variance of $(\widetilde{\theta}_L, \widetilde{\theta}_U)$ around the target bounds.
Theorem (ref) provides confidence intervals that covers $\theta$ with probability at least $1 - \alpha$ for any size of the main and auxiliary samples. Since $\theta \in [0,1]$, the one-sided CIs are particularly useful for testing $\theta=0$ or $\theta=1$. The proof of the theorem exploits sample splitting to take $\hat{s}_L, \hat{s}_U$ as given in the main sample. Then, since by ((ref)) $\theta$ must lie within $[\max_t \Delta(t, \hat{s}_L), 1 + \min_t \Delta(t, \hat{s}_U)]$, the CIs are built to cover this interval.
The finite-sample guarantee comes from the DKW inequality dvoretzky1956asymptotic, massart1990tight, a concentration inequality for empirical cdfs. A similar approach was implemented in ruiz2022non in the context of finite-sample inference on $\theta$ without covariates. My proof differs in that I consider the splitting of the sample, and I use the one-sided version of the DKW inequality in massart1990tight to get a smaller critical value for the one-sided CIs.
Note that a large main sample leads to $c_\alpha$ being close to zero, and a large auxiliary sample makes a suitable choice of $(\hat{s}_L, \hat{s}_U)$ closer to $(s^*_L, s^*_U)$. Therefore, under regularity conditions, the CIs in Theorem (ref) collapse into the sharp bounds in ((ref)) as both sample sizes grow to infinity. The next section investigates the large sample behavior of a similar estimator, which allows for cross-fitting instead of sample-splitting.
The approach of the previous section uses sample-splitting to provide a strong guarantee of coverage under weak conditions. In this section, I propose a new approach that uses cross-fitting to improve power. Unlike sample-splitting, cross-fitting uses the whole sample for estimating both the covariate-adjustment term and the final estimator.
In order to derive asymptotically valid confidence intervals, I define new regularity conditions (A(ref)) and show in Theorem (ref) an intermediate result, that the estimators are asymptotically Gaussian. The main result is then established in Theorem (ref), which shows uniform validity of CIs for $\theta$. Finally, Theorem (ref) shows conditions under which the CIs are exact.
Again, several methods can be used to estimate $(\hat{s}_{L,k}, \hat{s}_{U,k})$, including machine learning algorithms or other nonparametric approaches (see, e.g., kneib2023rage for a review). $K$ is assumed fixed as $n \to \infty$, and typical choices are $K=5$ or $K=10$.
In order to derive the asymptotic properties of the cross-fitting estimators, I define new regularity conditions. The results are uniform over a set of probability functions $\mathcal{P}$. This is because the literature on inference in partially identified models argues that asymptotic results have to be uniform in order to obtain approximations that accurately represent the finite-sample behavior (e.g., imbens2004confidence, andrews2010inference). To this end, I introduce a slight revision in notation, incorporating an additional subscript $P$ for $P \in \mathcal{P}$. Consequently, $\Delta$ in ((ref)) is denoted as $\Delta_P$, while the sharp transformations in ((ref)) and ((ref)) are represented by $s^*_{L,P}$ and $s^*_{U,P}$, and sharp bounds in ((ref)) are $(\theta_{L,P}^*, \theta_{U,P}^*)$.
A(ref) restricts the possibility of point masses in the distribution of the outcome. A(ref) requires the estimators $(\hat{s}_{L}, \hat{s}_{U})$ to have any limit in probability, at any rate of convergence, and allows them to be inconsistent for the sharp $(s^*_{L,P}, s^*_{U,P})$. A(ref) is assumed to simplify exposition. It ensures the asymptotic distribution is first-order insensitive to the choice of the optimizers in Definition (ref), and it is a generalization for $P \in \mathcal{P}$ of the assumption used in fan2010sharp in their approach to inference on Makarov bounds without covariates. I discuss two approaches for relaxing A(ref) in the presence of multiple optimizers in the Online Appendix ((ref)). One of them allows dropping A(ref) by using the bootstrap for inference instead of relying on normal approximations to $\widehat{\theta}_{L}$ and $\widehat{\theta}_{U}$. Example (ref) illustrates a data generating process where A(ref) holds. A(ref) is used to ensure that the problem is not trivial in the sense that $Var[\sqrt{n} \widehat{\theta}_{L}]$ and $Var[\sqrt{n} \widehat{\theta}_{U}]$ do not converge to zero. It is only used to apply the result from stoye2009more to get asymptotic validity of the two-sided CI in Theorem (ref).
Under A(ref), the presence of the optimizers in Definition (ref) is asymptotically irrelevant, and $(\widehat{\theta}_{L}, \widehat{\theta}_{U})$ behave as sample averages. It follows that their asymptotic distribution will be $\sqrt{n}$-Gaussian, as shown in Theorem (ref). The presence of the estimated $(\hat{s}_{L, k}, \hat{s}_{U, k})$ will only affect where $(\widehat{\theta}_{L}, \widehat{\theta}_{U})$ are centered at, but these target bounds will always contain $\theta$, since any $s$ induces valid bounds ((ref)).
Definitions of $\sigma^2_{L,P}$, $\sigma^2_{U,P}$ and $\sigma_{L,U,P}$ are provided in Appendix (ref). The cross-fitting estimators are asymptotically normal at the usual $\sqrt{n}$ rate, centered at bounds $[\bar{\theta}_{L, P}, \bar{\theta}_{U, P}]$ that may be wider than the sharp $[\theta_L^*, \theta_U^*]$, but which always contain $\theta$. The nonparametric rates of convergence of $\hat{s}_L, \hat{s}_U$ only affect the target bounds $[\bar{\theta}_{L, P}, \bar{\theta}_{U, P} ]$, which always contain $\theta$. The resulting $\sqrt{n}$ rate is made possible through the splitting of the sample. Once conditioning on the estimated $(\hat{s}_{L, k}, \hat{s}_{U, k})$ and under A(ref), $(\widehat{\theta}_{L}, \widehat{\theta}_{U})$ behave as sample averages. Conditions under which $(\bar{\theta}_{L, P}, \bar{\theta}_{U, P}) = (\theta_{L,P}^*, \theta_{U,P}^*)$ are given in Theorem (ref).
Confidence intervals for $\theta$ follow from consistent sample-analogue estimators for $\sigma_{L}^2$, $\sigma_{U}^2$ and $\sigma_{L,U}$. I provide expressions in Appendix (ref). One-sided confidence intervals follow directly from the normal approximation of Theorem (ref). Since $\theta \in [0,1]$, the one-sided CIs are particularly useful for testing $\theta=0$ or $\theta=1$. For two-sided inference on $\theta$, I suggest using the confidence interval $\text{CI}^3_\alpha$ proposed in stoye2009more, here denoted by $\widehat{CI}_\alpha$. A precise definition is reproduced in Online Appendix (ref).
Finally, I show that the target bounds $(\bar{\theta}_{L, P}, \bar{\theta}_{U, P})$ converge to the sharp $(\theta^*_{L, P}, \theta^*_{U, P})$ if the estimators of the covariate adjustment terms are consistent to the optimal at any rate. Moreover, I show that the confidence intervals are asymptotically exact under a smoothness and rate condition.
Note that a simple sufficient condition for (ii) in Theorem (ref) is that $|\hat{s}_{L,k}(X) - s^*_{L, P}(X)|$ converges to zero at the semiparametric rate $o_{P}(n^{-1/4})$.
I illustrate the results of Theorems (ref) and (ref) with a simulation study. I also compare the performance of the sample-splitting and cross-fitting estimators, as well as the estimator proposed by semenova2023adaptive and ji2023model (denoted SJLS for the first letter of the authors' names). Note that both estimators are equivalent in the case of the DTE and that CAIDE by construction estimates bounds that are always (weakly) narrower than SJLS (Appendix (ref)).
For $n \in \{500, 2000\}$, I simulate $10,000$ times from the following data-generating process (DGP):
For this DGP, $\theta = \theta_0 \approx 0.43$. I fit with $K=5$ and $\delta=0$ the models Quantile Neural Networks, Support Vector Machine (Regression), and No Covariates ($s(x)=0$). I also calculate an Oracle model that uses the sharp transformations defined in Theorem (ref), which are generally unknown if the DGP is unknown. Technical details are delayed to Appendix (ref). I consider two scenarios: (i) $p=10$, where only covariates $X_1$ through $X_{10}$ are observed; (ii) $p=20$, where all $20$ covariates are observed. Note that when $p=20$, $\theta$ is point identified since potential outcomes are uniquely determined by all $20$ covariates.
Figure (ref) shows the distribution of the estimators CAIDE-CF (Cross Fitting) and SJLS for the lower and upper bounds on $\theta$, for $n=500,2000$ and scenarios $p=10,20$. It illustrates that the estimators are centered around bounds determined by the different models for estimating $\hat{s}$. Some models perform better than others in yielding narrower bounds. For example, in most cases, the Regr. SVM model led to the narrowest bounds, close to the sharp bounds in the $p=10$ case. A larger sample size leads to smaller variation in the estimators. Observing all $p=20$ covariates, instead of $p=10$, can lead to narrower bounds (e.g., Regr. SVM and $n=2000$) but also to wider bounds if the sample is smaller (Quant. NNet and $n=500$). Throughout, the estimates of CAIDE are narrower than those of SJLS.
In Table (ref), I show the probability of rejecting the hypotheses $H_0: \theta = 0$ and $H_0: \theta \ge \theta_0$, using the one-sided confidence interval based on the lower bound in Theorem (ref). Here, $\theta_0 \approx 0.43$ is the true value of $\theta$ for this DGP. A high rejection rate for the first hypothesis indicates power since $0$ is not in the sharp identified set of $\theta$. The second hypothesis's rejection rate indicates size, and it should ideally be less than or equal to the nominal $0.05$. I also measure the average length of the confidence intervals, defined as the distance between the lower and upper bounds calculated from the one-sided CIs in Theorem (ref). I compare $n=500,2000$ and $p=10,20$ for different models and estimators. CAIDE-SS is the sample-splitting estimator of Section (ref), and CAIDE-CF is the cross-fitting estimator of Section (ref). The table shows that including covariates, having a larger sample size, or observing all $20$ covariates increases power while preserving size. The sample-splitting estimators are conservative and exhibit the largest average length for Regr. SVM. Still, it often demonstrates some power. For example, it excludes zero in $99.4\%$ of the cases when $p=20$, $n=2000$, and the model is Regr. SVM. The cross-fitting estimator performs better than SJLS in terms of smaller confidence intervals while preserving coverage. It achieves the nominal coverage rate for the oracle model when $p=20$.
In 2022, microcredit reached more than 170 million borrowers worldwide, with a total loan portfolio of over \$140 billion Convergences2023ImpactFinance. Theory suggests that microborrowing can be beneficial or harmful (banerjee2013microcredit, garz2021consumer), and evidence suggests that average effects are small if not negligible (banerjee2015six, meager2019understanding). Analysis of heterogeneity mostly relied on rank preservation and suggested mainly positive impact in the upper tail of income/profits distributions (e.g., meager2022aggregating). Yet, evidence of negative effects has been scarce. I revisit five studies that randomized offering a microloan at the individual level\footnote{The only exception is Egypt, in which the treated group was offered a loan size twice as large as the one offered in the control group.}, and find evidence of both positive and negative treatment effects of microcredit. Table (ref) presents a brief description of the datasets.
CAIDE is able to provide informative bounds, even in this challenging context. First, the setting is challenging because the distributions of income/profits are similar between the treated and control groups. This similarity is reflected in not statistically significant ATEs in Table (ref), and estimated $[\theta_L, \theta_U]$ close to $[0,1]$ if covariates are not considered (Table (ref)). Second, because the sample sizes are relatively small (ranging from 548 to 1964), with a large number of covariates (from 24 to 97). Finally, because predicting business outcomes is itself a difficult task mckenzie2019predicting.
Using $K = 10$, I fit the estimators of Section (ref)\footnote{The sample-splitting estimators of Section (ref) with a 50-50 split lead all confidence intervals to be $[0,1]$ at $10\%$ significance level.}. For each fold, I fit another layer of sample splitting to pick via $10-$fold cross-validation one of the models: Quantile Forests, Quantile Neural Networks, Random Forest (Regression), Neural Networks (Reg.), Elastic Net (Reg.), Extreme Gradient Boosting (Reg.), and Constant ($s(X) = 0)$. Technical details are provided in Appendix (ref).
Table (ref) shows the estimates for the lower and upper bounds on $\theta$ for $\delta = 0$ using CAIDE with and without covariates\footnote{Note that the estimators and one-sided CIs without covariates ($s(x)=0$) are equivalent to the ones in fan2010sharp.}. The outcome is the logarithm of annual household income. Considering that the outcome is continuous, it follows that $\theta = P(Y(1) - Y(0) \le 0) = P(Y(1) - Y(0) < 0)$, the lower bound on $\theta$ is the minimum proportion of households that have a negative effect from the microloan, and the upper bound is the maximum proportion. Note that one minus the upper bound denotes the minimum proportion of households with a positive effect. Results for $\delta = -0.05$ and $\delta = 0.05$ are similar and are presented in Table (ref). In Bosnia, for example, I estimate that at least $10.8\%$ of the households would be better off without the microloan (p-value $ < 10^{-3}$), and at most $94.6\%$ have negative effects (p-value $0.028$). That is, at least $6.4\%$ are better off ($1 - 0.946$). Overall, I find significant lower bounds at the $10\%$ level for all five datasets and at the $5\%$ level for four of them. As for the upper bound, I find statistical significance at $5\%$ in four datasets. These results reveal an important presence of heterogeneity, even when the ATEs are not statistically significant (Table (ref)).
Incorporating covariates makes a substantial difference in many of the point estimates. In Egypt, for example, the lower bound without covariates is $0.009$ (p-value $0.264$), whereas incorporating covariates improves the estimate to $0.260$ (p-value $<10^{-3}$). The differences between the point estimates with and without covariates are significant at the $5\%$ level for both the lower and upper bounds for Egypt, the Philippines (2) and All.
Table (ref) also illustrates how CAIDE can outperform SJLS (semenova2023adaptive, ji2023model) in terms of narrower confidence intervals. In the case of the lower bound for Bosnia, for example, it is estimated $10.8\%$ with CAIDE vs. $2.4\%$ with SJLS, leading to a substantial difference in p-values ($<10^{-3}$ vs. $0.254$). It also illustrates how SJLS can estimate negative lower bounds (South Africa), or upper bounds greater than one (Philippines (1)). In both cases, the estimates with CAIDE are statistically significant at the $10\%$ level.
Table (ref) in Online Appendix (ref) shows results for business profits. Again, including covariates improves the bounds. Averaging estimates from all datasets, I find that at least $10.8\%$ (p-value $< 10^{-3}$) of the population was harmed, and at least $9.8\%$ (p-value $< 10^{-3}$) benefited from microcredit.
An important note is that this evidence does not inform the nature of the negative effects or their magnitude. For example, the negative outcomes could be relatively small and a natural consequence of investments being risky. In that case, the decision to take a microloan could still be rational if the prospect of gains is better than the risk of loss.
I propose two novel inference approaches to the distributional treatment effect $\theta = P ( Y(1) - Y(0) \le \delta )$ for any fixed $\delta$. A new characterization of the sharp identified set of $\theta$ is provided, based on covariate adjustment. It has the desirable property of being robust to the choice of the covariate-adjustment term $s(x)$. If this term is misspecified, the identified set is wider but still valid. As a consequence, any method can be used to estimate $s$, including machine learning algorithms. I provide finite-sample valid inference using sample-splitting and asymptotic inference using cross-fitting, which is also asymptotically exact under additional conditions.
I illustrate the practical relevance of the method in a simulation study and an application to microcredit. The application reveals an important presence of heterogeneity, with statistically significant negative and positive treatment effects. The results show that the method is able to provide informative bounds even in challenging contexts, such as when the ATE is not statistically significant, the sample size is small, and the number of covariates is large.
By using CAIDE, an applied researcher can learn more about heterogeneous effects not only by gathering larger samples but also by collecting information on variables that help predict potential outcomes.