EconBase
← Back to paper

Identification and Inference on Treatment Effects under Covariate-Adaptive Randomization and Imperfect Compliance

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,622 characters · 13 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Inference on Treatment Effects under Covariate-Adaptive Randomization and Imperfect Compliance

\relax \hypersetup{pageanchor=false} \hypersetup{pageanchor=true}

tabular[tabular omitted — 221 chars of source]

}

\thispagestyle{empty}

spacing{1.2} \begin{abstract} Randomized controlled trials (RCTs) frequently utilize covariate-adaptive randomization (CAR) (e.g., stratified block randomization) and commonly suffer from imperfect compliance. This paper studies the identification and inference for the average treatment effect (ATE) and the average treatment effect on the treated (ATT) in such RCTs with a binary treatment. We first develop characterizations of the identified sets for both estimands. Since data are generally not i.i.d.\ under CAR, these characterizations do not follow from existing results. We then provide consistent estimators of the identified sets and asymptotically valid confidence intervals for the parameters. Our asymptotic analysis leads to concrete practical recommendations regarding how to estimate the treatment assignment probabilities that enter the estimated bounds. For the ATE bounds, using sample analog assignment frequencies is more efficient than relying on the true assignment probabilities. For the ATT bounds, the most efficient approach is to use the true assignment probability for the probabilities in the numerator and the sample analog for those in the denominator. \end{abstract}

KEYWORDS: Randomized controlled trials, covariate-adaptive randomization, stratified block randomization, imperfect compliance, partial identification.

JEL classification codes: C12, C14

\thispagestyle{empty}

Introduction

This paper considers identification and inference in a randomized controlled trial (RCT) that features covariate-adaptive randomization (CAR) and imperfect compliance. CAR refers to randomization schemes that first stratify according to baseline covariates and then assign treatment status to achieve “balance” within each stratum (bugni/canay/shaikh:2018,bugni/canay/shaikh:2019,bugni/gao:2023, among many others). These randomization schemes are widely utilized to assign treatment status in RCTs across various scientific disciplines; see rosenberger/lachin:2016 for a textbook treatment of this topic focused on clinical trials and duflo/glennerster/kremer:2007 and bruhn/mckenzie:2008 for reviews focused on development economics. CAR complicates the econometric analysis, as it typically leads to treatment assignments, treatment decisions, and outcomes that are not independent and identically distributed (i.i.d.). In particular, observations may be dependent, and their distributions may differ across units.

We allow for imperfect compliance in the sense that the RCT participants endogenously decide their treatment, which may or may not coincide with the treatment status assigned by the CAR mechanism. This is empirically relevant, as imperfect compliance is a common occurrence in many RCTs. In this context, bugni/gao:2023 study inference on the local average treatment effect (LATE). For more results on inference about the LATE under CAR, see ansel/hong/li:2018,bai/guo/shaikh/tabord-meehan:2024,jiang/linton/tang/zhang:2024.

This paper studies the identification and inference for the average treatment effect (ATE) and the average treatment effect on the treated (ATT) in such RCTs with a binary treatment. We extend our analysis to the average treatment effect on the untreated (ATU) in the Appendix. These parameters are of substantial interest to researchers conducting RCTs. Under i.i.d.\ assumptions, the literature has shown that these parameters are partially identified under imperfect compliance; e.g., manski:1989,manski:1990,manski:1994,balke/pearl:1997,heckman:2001,huber:2017,kitagawa:2021. Under CAR, however, the data are typically not i.i.d., which renders these results inapplicable.

Our paper makes several contributions. First, we provide the identified set of our estimands of interest under CAR assumptions. Since the data are generally not i.i.d., our identification analysis is based on the joint data-generating process (DGP) of the individual data and the treatment assignment mechanism. To the extent of our knowledge, this is the first characterization of these identified sets in the CAR literature with imperfect compliance. Second, we provide consistent estimators of the bounds of the identified set. Finally, we provide asymptotically valid confidence intervals for our partially identified parameters. Our asymptotic results are uniformly valid in a large class of probability distributions, which is necessary to obtain approximations that adequately represent the finite sample experiment of interest; see imbens/manski:2004,andrews/soares:2010.

Our asymptotic analysis delivers concrete and practical recommendations on how to estimate the identified sets, particularly on how to estimate the treatment assignment probabilities. In the context of an RCT, the treatment assignment probabilities are often known, and so the question here is whether the researcher should estimate the identified sets using the true assignment probabilities or their sample analogs. As it turns out, our recommendations depend on the estimand under consideration. In the case of the ATE, using sample analog treatment assignment probabilities is (weakly) more efficient than using the true (limiting) treatment assignment probabilities, which translates into a more powerful inference about the ATE. In the case of the ATT (and the ATU), we recommend using the true (limiting) treatment assignment probabilities in the numerator and the sample analogs in the denominator. This approach yields more powerful inference for the ATT (and the ATU) than alternative bounds based on other configurations of treatment assignment probability estimators. To our knowledge, these asymptotic findings appear to be new in the context of analysis of RCTs with imperfect compliance (with or without CAR). In the context of perfect compliance (i.e., point identified estimands) and i.i.d.\ samples (i.e., without CAR), analogous results have been obtained in hahn:1998,hirano/imbens/ridder:2003.

The rest of the paper is organized as follows. Section (ref), we describe the setup. Section (ref) studies the identification and inference for the ATE, and Section (ref) does the same for the ATT. In Section (ref), we explore the finite sample behavior of our methods via Monte Carlo simulations. Section (ref) illustrates our results in an empirical application based on the RCT in dupas/karlan/robinson/ubfal:2018. Section (ref) provides concluding remarks. All proofs and several intermediate results are collected in the appendix. For brevity, several auxiliary results have been placed in the Appendix.

Setup

We consider an RCT with $n$ participants. For each participant $i=1,\dots,n$, $Y_i \in \mathbb{R}$ is the observed outcome of interest, $Z_i \in \mathcal{Z}$ is a vector of observed baseline covariates, $A_i \in \{0,1\}$ is the treatment assignment, and $D_i \in \{0,1\}$ is the treatment decision.

We propose potential outcome models for outcomes and treatment decisions. We use $Y_i(D)$ to denote the potential outcome of participant $i$ if he/she makes treatment decision $D$, and we use $D_i(A)$ to denote the potential treatment decision of participant $i$ if he/she has assigned treatment $A$. These are related to their observed counterparts in the usual manner:

align[align omitted — 117 chars of source]

Following the usual classification in the LATE framework in angrist/imbens:1994, each participant in the RCT can only be one of four types: complier, always taker, never taker, or a defier. An individual $i$ is said to be a complier if $\{D_i(0)=0,D_i(1)=1\}$, an always taker if $\{D_i(0)=D_i(1)=1\}$, a never taker if $\{D_i(0)=D_i(1)=0\}$, and a defier if $\{D_i(0)=1,D_i(1)=0\}$. As is common in the literature, we will later impose that there are no defiers in our population.

We use ${\bf Q}$ to denote the distribution of the underlying random variables, given by

align*[align* omitted — 78 chars of source]

We use ${\bf G}$ to denote the treatment assignment distribution, i.e., the distribution of $\{ A^{( n) }|W^{(n)}\} $. Finally, we use ${\bf P}$ to denote the observed data distribution, given by

align*[align* omitted — 61 chars of source]

Note that ${\bf P}$ is jointly determined by (ref) and the distributions ${\bf Q}$ and ${\bf G}$. We state our assumptions below in terms of restrictions on ${\bf Q}$ and ${\bf G}$. These distributions can vary with the sample size, but we keep this dependence implicit for the simplicity of exposition. Throughout the paper, we use $P$, $E$, and $V$ to denote probability, expectation, and variance, likewise keeping the underlying probability measures ${\bf P}$, ${\bf Q}$, or ${\bf G}$ implicit.

Strata are constructed from the observed baseline covariates $Z_i$ using a prespecified function $S:\mathcal{Z} \to \mathcal{S}$, where $\mathcal{Z}$ denotes the support of $Z_i$ and $\mathcal{S}$ is a finite set. For each participant $i=1,\dots,n$, let $S_{i} \equiv S(Z_{i})$ and let $S^{(n)} = \{S_{i}:i=1,\dots,n\}$. Our assumptions on ${\bf Q}$ are as follows.

assumption{\rm \bf (Underlying distribution)} For constants $Y_L$, $Y_H$, $\xi$, $W^{(n)}$ is an i.i.d.\ sample from ${\bf Q}$ that satisfies \begin{enumerate}[(a)] • $Y_{i}(d) \in [Y_L,Y_H]$ for all $d\in \{0,1\}$, where $Y_L$ and $Y_H$ are known constants. Also, $p(s) \equiv P(S_i=s)\geq \xi >0$ for all $s \in \mathcal{S}$. • $ P(D_{i}( 0) =1,D_{i}(1) =0) =0$ or, equivalently, $ P( D_{i}(1)\geq D_{i}(0)) =1$. • $V[ Y_{i}( d) |D_{i}( a) =d,S_{i}=s] P(D_{i}( a) =d,S_{i}=s) \geq \xi >0$ for some $((d,a),s)\in \{(0,0),(1,1)\} \times \mathcal{S}$. • $P(D_i(1)=1,D_i(0)=0)\geq \xi >0$ or, equivalently, $ P( D_{i}(1)> D_{i}(0)) \geq \xi >0$. \end{enumerate}

Assumption (ref) says that the underlying data distribution is i.i.d., i.e., ${\bf Q}$ is the product of $n$ identical marginal distributions of $(Y_i(1),Y_i(0),D_i(1),D_i(0),Z_i)$. In addition, Assumption (ref)(a) imposes that $Y_L$ and $Y_H$ are known logical bounds for the potential outcomes and that all strata are relevant. In principle, one could entertain the case with $Y_L = -\infty$ or $Y_H=\infty$, but our later results reveal that either of these generates a non-informative, trivial identified set for our parameters of interest. Assumption (ref)(b) corresponds to the “no defiers” or “monotonicity” condition that is standard in treatment effect analysis under imperfect compliance. Assumption (ref)(c) places a nontrivial bound on the conditional variance for at least one subgroup. This implies that our inference is non-degenerate, and it is arguably mild, as it is only required to hold for one $((d,a),s)\in \{(0,0),(1,1)\} \times \mathcal{S}$. Finally, Assumption (ref)(d) implies that the proportion of compliers in the population is non-trivial, ensuring that $P(D_i=1)\in(0,1)$, making ATT well-defined. It is not required for the ATE.

Under Assumption (ref), we introduce the following notation for population objects: for each $(d,a,s) \in \{0,1\}\times \{0,1\}\times \mathcal{S}$,

align[align omitted — 523 chars of source]

We note that the expressions in (ref) are all features of the distribution ${\bf Q}$ but we omit this dependence for ease of notation. It is also convenient to introduce notation for sample objects: for each $s \in \mathcal{S}$,

align*[align* omitted — 147 chars of source]

With this notation in place, we can state our assumptions regarding the treatment assignment mechanism ${\bf G}$.

assumption{\rm \bf (Assignment Mechanism)} For a constant $\varepsilon>0$, ${\bf G}$ satisfies \begin{enumerate}[{(a)}] • $W^{(n)}\perp A^{(n)}~|~S^{(n)}$. • $P(A_{i}=1 | S^{(n)}) \in (0,1)$ for all $i=1,\dots,n$. • For all $s \in \mathcal{S}$, $n_{A}(s)/n(s) = \pi _{A}(s) + o_{p}(1)$, where $\pi _{A}(s) \in (\varepsilon,1-\varepsilon)$. • $P(A_{i}=1 | S_i ) = P(A_{j}=1 | S_j )$ for all $i,j=1,\dots,n$. • $\sqrt{n} E[|\pi_A(S_i) - P( A_i=1|S_i)|] = o(1)$ for all $i=1,\ldots ,n$. • $\{\{\sqrt{n}(n_{A}(s)/n(s)-\pi_A(s)):s\in S\}|S^{(n)}\}\overset{d}{\to }N(\mathbf{0},{{\Sigma}}_{A})$ w.p.a.1, with ${{\Sigma}} _{A}=diag({\tau}(s)(1-\pi _{A}(s))\pi _{A}(s)/p(s):s\in \mathcal{S})$ and ${\tau}(s)\in [0,1]$. \end{enumerate}

Assumption (ref) represents the main departure relative to the i.i.d.\ setup, as it allows the treatment assignment vector $A^{(n)}$ to be non-independent and non-identically distributed. Assumption (ref)(a) requires that the treatment assignment vector $A^{(n)}$ is a function of the strata vector $S^{(n)}$ and a randomization device that is conditionally independent of the underlying sample $W^{(n)}$. Assumption (ref)(b) imposes that the conditional treatment assignment probability for any unit is between zero and one. This is a mild requirement that is satisfied by any of the treatment assignment mechanisms typically used in practice. Assumption (ref)(c) imposes that the fraction of units assigned to treatment in the stratum $s$, $n_A(s)/n(s)$, converges in probability to {\it some limit} $\pi_A(s) \in (0,1)$. In practice, $\pi_A(s)$ is typically the strata-specific treatment assignment probability desired by the researcher; see Examples (ref) and (ref). For this reason, we refer to $\{\pi_A(s):s\in \mathcal{S}\}$ as the {\it target probabilities}. Assumption (ref)(d) requires that the probability of assignment of any unit conditional only on its own strata is independent of the identity of the unit. We note that $P(A_{i}=1 | S_i) =E[P(A_i=1|S^{(n)})|S_i]$, which is not necessarily equal to $\pi_A(S_i)$.\footnote{In the typical CAR method, the reason for the difference between $P(A_i=1|S^{(n)})$ and $\pi_A(S_i)$ is that the sample size of strata $s=S_i$ multiplied by $\pi_A(s)$ may not be an integer. In turn, this results in $P(A_{i}=1 | S_i)$ and $\pi_A(S_i)$ being different.} While $P(A_i=1|S_i)$ and $\pi_A(S_i)$ need not coincide, Assumption (ref)(e) states that these need to converge to each other in expectation at a fast rate. Finally, Assumption (ref)(f) strengthens Assumption (ref)(c) to require that $\sqrt{n}(n_A(s)/n(s) - \pi_A(s))$ is asymptotically normal conditional on the strata information $S^{(n)}$. For each stratum $s \in \mathcal{S}$, the parameter $\tau(s) \in [0,1]$ determines the amount of dispersion that the CAR mechanism allows on the fraction of units assigned to the treatment. A lower value of $\tau(s)$ implies that the CAR mechanism imposes a higher degree of “balance” or “control” of the treatment assignment proportion relative to its desired target value. Assumption (ref) is satisfied by a wide array of CAR schemes, including stratified block randomization (SBR) and simple random sampling (SRS), among others.\footnote{Additional examples of CAR methods include Efron's biased coin design (efron:1971) and the so-called minimization methods (pocock/simon:1975,hu/hu:2012). Under suitable conditions, these can also be shown to satisfy Assumption (ref). See bugni/canay/shaikh:2018,bugni/gao:2023 for details.} SBR deserves special focus, as it is a frequently used CAR method in RCTs in economics.

example[Stratified Block Randomization (SBR)] This is sometimes also referred to as blocking, block randomization, or permuted blocks within strata. To implement this method, the researcher proposes a vector of desired assignment probabilities for each stratum, which we denote by $\{\pi_A(s):s \in \mathcal{S}\}$. Within every stratum $s \in \mathcal{S}$, SBR assigns exactly $\lfloor n(s)\pi_A(s)\rfloor$ of the $n(s)$ participants in stratum $s$ to treatment and the remaining $n(s) -\lfloor n(s)\pi_A(s)\rfloor $ to control, where all possible $$\binom{n(s)}{\lfloor n(s){\pi}_A(s)\rfloor}$$ assignments are equally likely. Then, $P(A_i = 1| S^{(n)}) ={\lfloor n(S_i)\pi_A(S_i)\rfloor}/{n(S_i)} $. SBR can be shown to satisfy Assumption (ref). In particular, bugni/canay/shaikh:2018 show that Assumptions (ref)(a)-(c) and (f) hold with $\tau(s)=0$ for all $s \in \mathcal{S}$. To verify Assumption (ref)(d), note that for all $s \in \mathcal{S}$ and $i=1,\dots,n$, $$P(A_i = 1|S_i = s) ~=~ E[P(A_i=1|S^{(n)})|S_i = s] ~=~ E[{\lfloor n(s)\pi_A(s)\rfloor}/{n(s)}].$$ Finally, Assumption (ref)(e) follows from this derivation: \begin{align*} \sqrt{n}E\left[ \left\vert \pi _{A}\left( S_{i}\right) -P( A_{i}=1|S_i) \right\vert \right] & = \sqrt{n}E\left[ \left\vert \pi _{A}\left( S_{i}\right) -E[\lfloor n( S_{i}) \pi _{A}( S_{i}) \rfloor /n( S_{i}) | S_i] \right\vert \right] \\ & \overset{(1)}{\leq} \sqrt{n}E\left[ 1\left[ n\left( S_{i}\right) \geq 1\right]/n\left( S_{i}\right) \right] \\ & = \sqrt{n}\sum\nolimits_{s\in \mathcal{S}}E\left[ 1/(\vartheta \left( s\right) +1)\right] p\left( s\right) \\ & \overset{(2)}{=} ({1}/{\sqrt{n}})\sum\nolimits_{s\in \mathcal{S}}\left( 1-\left( 1-p\left( s\right) \right) ^{n}\right) \overset{(3)}{\to} 0, \end{align*} for $\vartheta \left( s\right) \sim Bi\left( n-1,p\left( s\right) \right) $ for all $s\in \mathcal{S}$, where (1) holds by $n\left( S_{i}\right) \pi _{A}\left( S_{i}\right) -1\leq \left\lfloor n\left( S_{i}\right) \pi _{A}\left( S_{i}\right) \right\rfloor$ $ \leq n\left( S_{i}\right) \pi _{A}\left( S_{i}\right) $, (2) by direct computation, and (3) by $\left\vert \mathcal{S}\right\vert <\infty $.
example[Simple Random Sampling (SRS)] This refers to a treatment assignment mechanism in which $A^{(n)}$ satisfies \begin{equation} P(A^{(n)}=(a_i: i=1,\dots,n)|S^{(n)},W^{(n)}) = \prod_{i=1}^{n} \pi_A(S_i)^{a_i} (1-\pi_A(S_i))^{1-a_i}. \end{equation} In other words, SRS assigns each participant in stratum $s$ to treatment with probability $\pi_A(s)$ and to control with probability $(1-\pi_A(s))$, independent of the rest of the sample information. It is easy to see that SRS satisfies Assumption (ref). In particular, bugni/canay/shaikh:2018 show that Assumptions (ref)(a)-(c) and (f) hold with $\tau(s)=1$ for all $s \in \mathcal{S}$. Finally, Assumptions (ref)(d)-(e) hold by $P(A_i = 1|S^{(n)} ) =P(A_i = 1|S_i) = \pi_A(S_i)$ for all $i=1,\dots,n$.

Average treatment effect (ATE)

This section studies the identification and inference of the ATE, defined as

equation[equation omitted — 60 chars of source]

By definition, $\theta$ is the expectation of the treatment effect when the treatment is mandated across the entire population.

We note that the ATE is solely determined by the underlying data distribution ${\bf Q}$, i.e., the treatment assignment mechanism ${\bf G}$ plays no role in determining $ \theta$.

Identification

The next result characterizes the identified set of the ATE.

theoremUnder Assumptions (ref)(a)-(c) and (ref)(a)-(c), the identified set for the ATE is $\Theta _{I}({\bf P})= [\theta_{L}({\bf P}),\theta_{H}({\bf P})],$ where \begin{align} \theta _{L}({\bf P}) & = E\bigg[ \frac{(Y_{i}D_{i}+Y_{L}( 1-D_{i}) ) A_{i}}{P(A_{i}=1|S_{i}) }-\frac{( Y_{i}(1-D_{i}) +Y_{H}D_{i}) ( 1-A_{i}) }{1-P(A_{i}=1|S_{i}) }\bigg] , \notag\\ \theta _{H}({\bf P}) & = E\bigg[ \frac{(Y_{i}D_{i}+Y_{H}( 1-D_{i}) ) A_{i}}{P(A_{i}=1|S_{i}) }-\frac{( Y_{i}(1-D_{i}) +Y_{L}D_{i}) ( 1-A_{i}) }{1-P(A_{i}=1|S_{i})}\bigg] . \end{align}

A few remarks are in order. As already mentioned, our setup allows for treatment assignment using CAR methods, which implies that the data may not be i.i.d. Hence, Theorem (ref) does not follow from the existing results on the identification of the ATE under i.i.d.\ sampling, e.g., manski:1990,heckman:2001,huber:2017,kitagawa:2021. In fact, to our knowledge, Theorem (ref) is the first characterization of the identified set of the ATE under the CAR framework with imperfect compliance.

Theorem (ref) shows that our bounds coincide with the sharp bounds on the ATE originally derived by manski:1990 under the monotonicity condition stated in Assumption (ref)(b) and when $\mathcal{S}$ is a singleton. This implies that the departure from the i.i.d.\ conditions allowed by the treatment assignment mechanism under CAR does not affect the structure of the sharp bounds. While this might suggest that our proof is merely a by-product of existing results, this is not the case. Prior work focuses on the i.i.d.\ setting and therefore relies on identifying information contained in the marginal distribution of a single individual to construct the bounds. In contrast, our approach works with the joint distribution of all $n$ individuals in the sample, allowing us to account for the dependence and heterogeneity permitted by the treatment assignment mechanisms under our CAR framework.

It is worth highlighting that the equalities in (ref) hold for any $i=1,\dots,n$. That is, the identified set for the ATE is the same for all of the individuals in our sample. This observation would be straightforward in the context of an i.i.d.\ sample, but it requires proof in the CAR framework, where data need not be i.i.d. It is also notable that Theorem (ref) would not change if we restrict attention to Assumptions (ref)(a)-(b) and (ref)(a)-(b). In other words, Assumptions (ref)(c) and (ref)(c) do not provide any additional identifying power for the ATE, though they will become relevant when conducting inference in the next section.\footnote{This also connects to our earlier point: without Assumption (ref)(d), the probability $P(A_i = 1|S_i)$ may vary across individuals, and yet, under our assumptions, the identified set for the ATE does not vary with $i$.}

Finally, the proof of Theorem (ref) presented in Appendix (ref) reveals that $\theta _{L}({\bf P})$ and $\theta _{H}({\bf P})$ only depend on the underlying data distribution ${\bf Q}$. That is, the treatment assignment mechanism ${\bf G}$ plays no role in determining the sharp bounds on the ATE.

Inference

Our goal in this section is to estimate the identified set for the ATE and provide a confidence interval for its true value. Following Theorem (ref), we propose the following estimators of the sharp bounds in (ref),

align[align omitted — 471 chars of source]

Note that (ref) is the sample analog of (ref), where the treatment assignment probabilities $\{P(A_{i}=1|S_{i}):i=1,\dots,n\} $ are estimated by the sample treatment frequencies $\{n_{A}( S_{i}) /n( S_{i}):i=1,\dots,n \}$.\footnote{The definitions in (ref) require that $n_{A}(S_{i}) /n( S_{i}) \in (0,1)$ for all $i=1,\dots,n$. Lemma (ref) in the appendix shows that this occurs with probability approaching one (w.p.a.1).} An alternative bounds estimator could be obtained by estimating the treatment assignment probabilities with their target values $\{\pi_{A}( S_{i}) :i=1,\dots,n \}$. In Section (ref), we argue that the bounds estimator in (ref) should be preferred, as it generates inference for the ATE that is relatively more robust, efficient, and powerful (all in a sense made precise in that section).

We now derive the asymptotic properties of the bounds estimator in (ref). The literature on inference in partially identified models convincingly argues that uniform asymptotic results are necessary to obtain approximations that adequately represent the finite-sample experiment of interest; e.g., see imbens/manski:2004,andrews/soares:2010. With this motivation in mind, we establish the uniform asymptotic distribution of our estimated bounds. Let $\mathcal{P}_1 =\mathcal{P}(Y_L,Y_H,\xi,\varepsilon) $ denote the set of probabilities ${\bf P}$ generated by $( {\bf Q},{\bf G})$ that satisfy Assumptions (ref)(a)-(c) and (ref)(a)-(c),(e). Theorem (ref) in the appendix shows that

equation[equation omitted — 288 chars of source]

Two aspects of (ref) are worth highlighting. First, the asymptotic variance $\Sigma_{\theta}({\bf P})$ accounts for the use of CAR in treatment assignment and typically differs from the asymptotic variance obtained under SRS. Second, the uniformity of our results contrasts with the typical asymptotic approximations in the CAR literature, which are pointwise in nature. As discussed, this aspect of our derivations is crucial in the context of partial identification.

As a corollary of (ref), we now establish that $\hat{\Theta} _{I} \equiv [\hat{\theta}_{L},\hat{\theta}_{H}]$ is a uniformly consistent estimator of the identified set for the ATE.

theoremFor any $\varepsilon >0$, \begin{equation} \underset{n\to \infty }{\lim} \inf_{{\bf P} \in \mathcal{P}_1 } {\bf P}( d_{H}(\hat{\Theta}_{I},\Theta _{I}({\bf P}))\leq \varepsilon ) = 1, \end{equation} where $d_H(U,V)$ denotes the Hausdorff distance between sets $U,V \subset \mathbb{R}$.

Our next result in this section proposes a CI for the ATE under our assumptions. To this end, Theorem (ref) provides a uniformly consistent estimator of the asymptotic variance $\Sigma_{\theta}({\bf P})$ in (ref), given by

equation[equation omitted — 199 chars of source]

where $(\hat\sigma _{L}^{2} ,\hat\sigma _{H}^{2}, \hat\sigma _{HL})$ are defined in (ref). Following our earlier comments, we note that these estimators account for the use of CAR in treatment assignment. With these estimators in place, we propose using the CI for the ATE given by $CI_{\alpha }^{2}$ in stoye:2009, i.e.,

equation[equation omitted — 229 chars of source]

where $( \hat{c}_{L},\hat{c}_{H}) $ are the minimizers of $\hat{\sigma}_{L}c_{L}+\hat{\sigma}_{H}c_{H}$ subject to

align[align omitted — 656 chars of source]

with $(Z_{1},Z_{2}) \sim N({\bf 0}_2,{\bf I}_2)$. Under our assumptions, stoye:2009 shows that $CI_{\alpha }^{2}$ is the shortest confidence interval with the correct asymptotic nominal size. The next result establishes that the CI in (ref) is asymptotically uniformly valid and exact.

theoremThe CI in (ref) satisfies \begin{equation} \underset{n\to \infty }{\lim} \inf_{{\bf P}\in \mathcal{P}_1 } \inf_{\theta \in \Theta_{I}({\bf P})} {\bf P}( \theta \in \hat{C}_{\theta}( 1-\alpha )) = 1-\alpha , \end{equation} where $\Theta_{I}({\bf P})$ is as in (ref).

The proof of Theorem (ref) follows from applying stoye:2009 to our CAR setup. The main ingredients for this result are the uniform convergence in distribution of our estimated bounds (Theorem (ref)) and the uniform consistency of our estimators of the asymptotic variance (Theorem (ref)). As previously emphasized, the uniform convergence results are novel in the CAR framework, yet essential for inference on partially identified parameters such as the ATE.

Discussion

As already mentioned, the bounds in (ref) estimate the treatment assignment probabilities $\{P(A_{i}=1|S_{i}):i=1,\dots,n\} $ in (ref) with the sample treatment frequencies $\{n_{A}( S_{i}) /n( S_{i}):i=1,\dots,n \}$. In an RCT, it is not uncommon to know the target probabilities $\{\pi_{A}(s):s \in \mathcal{S} \}$. In principle, this information would also allow us to construct {\it alternative estimated bounds} that estimate the treatment assignment probabilities in (ref) with their target probabilities $\{\pi_{A}( S_i):i=1,\dots,n \}$. This section explains that the estimated bounds in (ref) are demonstrably better than either of these alternative estimated bounds in several dimensions.

The first step to analyzing the limiting properties of the alternative estimated bounds is to derive their uniform asymptotic distribution. Under Assumptions (ref)(a)-(c) and (ref)(a)-(c),(e)-(f), Theorem (ref) in the Appendix derives the uniform asymptotic distribution of the alternative estimated bounds. This result reveals that the alternative estimated bounds have a uniform asymptotic distribution as in (ref), but with ${\Sigma}_{\theta}({\bf P})$ replaced by a (weakly) larger asymptotic variance $\tilde{\Sigma}_{\theta}({\bf P})$ in the sense that $\tilde{\Sigma}_{\theta}({\bf P}) - {\Sigma}_{\theta}({\bf P})$ is positive semidefinite. In particular, $\tilde{\Sigma}_{\theta}({\bf P})$ depends on the parameter $\{\tau(s):s \in \mathcal{S}\}$ in Assumption (ref)(f) characterizing the amount of dispersion of the CAR mechanism.\footnote{In fact, $\tilde{\Sigma}_{\theta}({\bf P}) - {\Sigma}_{\theta}({\bf P})$ can be positive definite if $\tau(s)>0$ for some $s\in \mathcal{S}$.} This result implies that the estimated bounds in (ref) are better than the alternative estimated bounds in three ways.

First, the asymptotic analysis of the estimated bounds in (ref) requires fewer assumptions than that of the alternative estimated bounds. In particular, the analysis of the former is established uniformly for $P \in \mathcal{P}_1$, which does not require Assumption (ref)(f). This implies that the asymptotic behavior of the estimated bounds in (ref) is robust to the details of the CAR mechanism. On the other hand, the asymptotic analysis of the alternative estimated bounds requires Assumption (ref)(f), and so they do not enjoy this robustness property.

Second, the estimated bounds in (ref) are (weakly) more efficient than the alternative estimated bounds. This follows from the fact that the estimated bounds in (ref) are asymptotically distributed according to $N({\bf 0}_2,{\Sigma}_{\theta}({\bf P}))$, the alternative estimated bounds are asymptotically distributed according to $N({\bf 0}_2,\tilde{\Sigma}_{\theta}({\bf P}))$, and $\tilde{\Sigma}_{\theta}({\bf P}) - {\Sigma}_{\theta}({\bf P})$ is positive semidefinite. That is, estimating treatment probability using sample treatment frequencies is more efficient than plugging in the target probabilities. This finding resembles the semiparametric efficiency results for point-identified ATE estimator in hahn:1998 and hirano/imbens/ridder:2003, but is novel in the context of partially identified inference for the ATE under imperfect compliance.

Third, inference based on the estimated bounds in (ref) is more powerful than the inference based on the alternative estimated bounds. This finding is based on results developed in our related work in bugni/gao/obradovic/velez:2024b. In that paper, we study the asymptotic power properties of stoye:2009's CIs for sequences of local alternative hypotheses and show that asymptotically normal bounds estimators with higher variance-covariance matrices (in a positive semidefinite sense) have (weakly) lower rejection rates.\footnote{This result is not obvious in stoye:2009's framework, as the CI based on less efficient estimated bounds does not necessarily include the CI based on the more efficient estimated bounds.} In combination with the efficiency results described in the last paragraph, this allows us to conclude that for all sequences of local alternative hypotheses, the limiting rejection rate of the CI in (ref) is (weakly) larger than that of the CI related to the alternative estimated bounds.

Average treatment effect on the treated (ATT)

This section studies identification and inference on the ATT, given by

align[align omitted — 76 chars of source]

By definition, $\upsilon$ represents the treatment effect for the sub-population that (endogenously) adopted the treatment. This sub-population is composed of compliers assigned to treatment and always takers assigned to control.

By definition, the ATT is the average treatment effect conditional on the decision $D_i=1$. Since $D_i = D_i(A_i)$, this reveals that the ATT depends on both ${\bf Q}$ and ${\bf G}$. This contrasts with the ATE, which only depends on ${\bf Q}$.

There are a few subtleties to the definition of ATT in the context of CAR. First, if we allow CAR mechanisms that determine arbitrary treatment assignment probabilities $\{ P( A_i=1|S^{( n) })\}_{i=1}^{n}$, it is possible for the ATT to vary with $i=1,\dots,n$. Assumption (ref)(d) avoids this possibility by assuming that $P( A_i=1|S_i)$ does not depend on $i$. Second, the definition of the ATT requires $P(D_{i}=1)>0$ for all $i=1,\dots,n$. This result can be shown based on Assumptions (ref)(d) and (ref)(b).

Identification

The following result characterizes the identified set of the ATT.

theoremUnder Assumptions (ref) and (ref)(a)-(d), the identified set for the ATT is $\Upsilon _{I}({\bf P})= [\upsilon_{L}({\bf P}),\upsilon_{H}({\bf P})],$ where { \begin{align} \upsilon _{L}( {\bf P}) & \equiv \frac{1}{G} E\Bigg[ \left( \frac{Y_i A_i}{P(A_i =1 | S_i)} - \frac{Y_i(1-A_i)}{1- P(A_i = 1 | S_i)} \right) P(A_i =1 | S_i) + \frac{(Y_i-Y_H)D_i(1-A_i)}{1 - P(A_i = 1 | S_i)} \Bigg], \notag \\ \upsilon_{H}( {\bf P}) & \equiv \frac{1}{G} E\Bigg[ \left( \frac{Y_i A_i}{P(A_i =1 | S_i)} - \frac{Y_i(1-A_i)}{ 1 - P(A_i = 1 | S_i)} \right) P(A_i =1 | S_i) + \frac{(Y_i-Y_L)D_i(1-A_i)}{1 - P(A_i = 1 | S_i)} \Bigg], \end{align} } and { \begin{align} G = E \left[ \left( \frac{D_i A_i}{ P(A_i =1 | S_i) } - \frac{D_i (1-A_i)}{1 - P(A_i = 1 | S_i) } \right) P(A_i =1 | S_i) + \frac{D_i (1-A_i)}{1 - P(A_i = 1 | S_i)} \right]. \end{align} }

At this point, we reiterate several remarks made after Theorem (ref). Theorem (ref) accommodates treatment assignment via CAR methods, implying that the data need not be i.i.d. Consequently, Theorem (ref) does not follow from existing results on the identification of the ATT under i.i.d.\ sampling, e.g., heckman:2000,huber:2017. In fact, Theorem (ref) is the first characterization of the identified set of the ATT under the CAR framework with imperfect compliance.

When data are indeed i.i.d., the bounds in Theorem (ref) coincide with the sharp bounds on the ATT. As in the case of the ATE, this shows that the relaxation of i.i.d.\ sampling permitted by CAR treatment assignment does not alter the structure of the sharp bounds. We emphasize that this result does not follow from existing literature, which derives such bounds exclusively under i.i.d.\ assumptions. Accommodating CAR requires a distinct proof based on the joint data distribution.

The equalities in (ref) hold for any $i=1,\dots,n$. As with the ATE, all individuals in the sample have the same the identified set for the ATT, and establishing this requires proof when the data need not be i.i.d. We also note that Theorem (ref) remains unchanged if we restrict attention to Assumptions (ref)(a)–(b),(d) and (ref)(a)–(b),(d). In other words, Assumptions (ref)(c) and (ref)(c),(e) do not contribute to the identification of the ATT, but play a role when conducting inference in the next section. Compared to the result for the ATE, the sharp bounds for the ATT additionally require Assumptions (ref)(d) and (ref)(d). Assumption (ref)(d) ensures the presence of compliers, which guarantees that $G$ in (ref) is positive. In turn, we require Assumption (ref)(d) to obtain that all individuals share a common ATT.

Finally, the proof of Theorem (ref) presented in Appendix (ref) reveals that $\upsilon _{L}( {\bf P}) $ and $\upsilon _{H}( {\bf P}) $ depend on both ${\bf Q}$ and ${\bf G}$. In contrast to the ATE, the treatment assignment mechanism ${\bf G}$ plays a role in determining the sharp bounds on the ATT.

Inference

We now turn to constructing an estimator of the identified set for the ATT and a confidence interval for its true value. The bounds in (ref) motivate the following:

align[align omitted — 851 chars of source]

and

equation[equation omitted — 378 chars of source]

It is clear that the expression in (ref) are the sample analog of those in (ref), where the treatment assignment probabilities $\{P(A_{i}=1|S_i):i=1,\dots,n\} $ that appear in the denominator are replaced by the sample treatment frequencies $\{n_{A}( S_{i})/n(S_i):i=1,\dots,n \}$ and the ones in the numerator are replaced by $\{\pi_{A}( S_{i}):i=1,\dots,n \}$. This construction requires the researcher to know $\{\pi_{A}( s):s \in \mathcal{S}\}$, which is reasonable in the context of an RCT. If these were unknown, one would replace them with $\{n_{A}( S_{i})/n(S_i):i=1,\dots,n \}$ but, as we explain in Section (ref), this leads to inference for the ATT that is relatively less robust, efficient, and powerful (in a sense made precise in that section).

We now derive the uniform asymptotic properties of the bounds estimator in (ref). To this end, let $\mathcal{P}_2=\mathcal{P}(Y_{L},Y_{H},\xi ,\varepsilon )$ denote the set of probabilities $\mathbf{P}$ generated by $(\mathbf{Q},\mathbf{G})$ that satisfy Assumptions (ref) and (ref)(a)-(e).\footnote{Since $\mathcal{P}_2 \subset \mathcal{P}_1$, our analysis for ATT requires stronger assumptions than that of the ATE.} Theorem (ref) shows that

equation[equation omitted — 312 chars of source]

As in the previous section, we note that $\Sigma_{\upsilon}(\mathbf{P})$ accounts for the use of CAR in treatment assignment and generally does not coincide with the asymptotic variance under i.i.d.\ sampling. Based on (ref), we conclude that $\hat{\Upsilon}_{I}\equiv [\hat{\upsilon}_{L},\hat{\upsilon}_{H}]$ is a uniformly consistent estimate of the identified set of the ATT.

theoremFor any $\varepsilon >0$, \begin{equation} \underset{n\to \infty }{\lim} \inf_{{\bf P} \in \mathcal{P}_2 } {\bf P}( d_{H}(\hat{\Upsilon}_{I},\Upsilon _{I}(\mathbf{P}))\leq \varepsilon ) = 1. \end{equation} where $d_H(A,B)$ denotes the Hausdorff distance between sets $A, B\subset \mathbb{R}$.

Next, we propose a CI for the ATT. For this purpose, Theorem (ref) proposes a uniformly consistent estimator for the variance $\Sigma _{\upsilon }(\mathbf{P})$ in (ref), given by

equation[equation omitted — 197 chars of source]

where $\left(\hat\varpi _{L},\hat\varpi _{H},\hat\varpi _{HL}\right) $ are as in (ref) and $\hat{G}$ is as in (ref). We can then construct a CI for the ATT as in (ref) but with $( \hat{\theta}_{L},\hat{\theta}_{H},\hat{\Sigma}_{\theta }) $ replaced by $( \hat{\upsilon}_{L},\hat{\upsilon}_{H},\hat{\Sigma}_{\upsilon }) $, i.e.,

equation[equation omitted — 251 chars of source]

where $(\hat{c}_L,\hat{c}_H)$ are the minimizers of $\hat\varpi _{L} c_{L} + \hat\varpi _{H} c_{H}$ subject to

align*[align* omitted — 651 chars of source]

and $Z_{1},Z_{2}$ are i.i.d.\ $N( 0,1) $. The next result establishes that this CI is asymptotically uniformly valid.

theoremThe CI in (ref) satisfies \begin{equation} \underset{n\to \infty }{\lim} \inf_{\mathbf{P}\in \mathcal{P}_2} \inf_{\upsilon \in \Upsilon _{I}(\mathbf{P})} {\bf P}(\upsilon \in \hat{C}_{\upsilon}(1-\alpha )) = 1-\alpha , \end{equation} where $\Upsilon _{I}(\mathbf{P})$ is as in (ref).

The argument for this result is analogous to that of the ATE. The main ingredients for this proof are the uniform convergence in distribution of the ATT bounds (Theorem (ref)) and the uniform consistency of our estimators of the asymptotic variance (Theorem (ref)). We reiterate that the uniformity of our results is crucial for partially identified analysis and is new in the context of CAR.

Discussion

As previously pointed out, the bounds in (ref) are constructed by replacing the treatment assignment probabilities $\{P(A_{i}=1|S_i):i=1,\dots,n\}$ in the denominator of (ref) with the sample treatment frequencies $\{n_{A}(S_{i})/n(S_i):i=1,\dots,n\}$, and those in the numerator of (ref) with $\{\pi_{A}(S_{i}):i=1,\dots,n\}$. In principle, one could consider alternative bounds that substitute either of these probabilities with different estimators. Our results show that all of these alternative bounds lead to inference for the ATT that is worse relative to that based on the bounds in (ref). For brevity, we focus on the alternative bounds that estimate both the probabilities in the numerator and denominator of (ref) using the sample treatment frequencies. Compared to the bounds in (ref), these alternative bounds have the advantage of not requiring knowledge of $\{P(A_{i}=1|S_i):i=1,\dots,n\}$ or $\{\pi_{A}(S_{i}):i=1,\dots,n\}$. However, we now argue that using the alternative bounds can be costly in terms of efficiency.

Under Assumptions (ref) and (ref), Theorem (ref) in the Appendix provides the uniform limiting distribution of the alternative estimated bounds where the probabilities in the numerator and denominator are both replaced by the sample treatment frequencies. This result shows that the alternative estimated bounds have a uniform asymptotic distribution as in (ref), but with ${\Sigma}_{\upsilon}({\bf P})$ replaced by with a (weakly) larger asymptotic variance $\tilde{\Sigma}_{\upsilon}({\bf P})$ in a positive semidefinite sense. We can use this result to argue that the bounds in (ref) are better than the alternative estimated bounds along the three dimensions discussed in Section (ref). We describe these briefly at the risk of some repetition.

First, the asymptotic analysis of the estimated bounds in (ref) requires fewer assumptions than that of the alternative estimated bounds. The former is derived for $P \in \mathcal{P}_2$, while the latter also invokes Assumption (ref)(f). Thus, the asymptotic behavior of the estimated bounds in (ref) is robust to the details of the CAR mechanism, while the asymptotic behavior of the alternative estimated bounds is not.

Second, the estimated bounds in (ref) are more efficient than either of the alternative estimated bounds. This follows directly from comparing the variance-covariance matrix of the asymptotic distribution of the two sets of bounds. That is, for the estimated bounds on the ATT, estimating treatment probability using target probabilities is more efficient than plugging in the sample treatment frequencies. Notice that this recommendation is the exact opposite of the one obtained in Section (ref) for the estimated bounds on the ATE. We also note that these findings resemble the semiparametric efficiency results for point-identified ATT estimator in hirano/imbens/ridder:2003, but are novel in our partially identified setting.

Finally, inference based on the estimated bounds in (ref) is more powerful than the inference based on the alternative estimated bounds. That is, for all sequences of local alternative hypotheses, the limiting rejection rate of the CI in (ref) is (weakly) larger than the limiting rejection rate of the CI related to the alternative estimated bounds. This result is a consequence of the relative efficiency comparison described in the previous paragraph and our related work in bugni/gao/obradovic/velez:2024b.

Monte Carlo simulations

In this section, we illustrate the finite-sample performance of our inference methods using Monte Carlo simulations. This exercise has two goals. First, we aim to demonstrate that our asymptotic results are accurate in finite samples. This ensures that confidence intervals cover each point of the identified sets with a minimum prespecified probability. We particularly focus on the extremes of the identified set, where coverage is more challenging. Second, we seek to confirm that our recommended estimated bounds are preferable to alternative estimated bounds in terms of the statistical power of the related confidence intervals.

We consider three simulation designs. Each simulated dataset has $n=500$ i.i.d.\ individuals. All designs have four strata, i.e., $\mathcal{S} = \{1,2,3,4\}$, and we set $P(S_i=s) = 0.25$ for all $s \in \mathcal{S}$. As explained in Section (ref), each individual can be a compiler (C), an always taker (AT), or a never taker (NT). We set $P(C|S_i=s) = 0.85$, $P(AT| S_i=s) = 0.05$, and $P(NT | S_i=s) = 0.10$ for all $s \in \mathcal{S}$. For each decision $d\in \{0,1\}$, type $t \in \{C, AT,NT\}$, and strata $s \in \mathcal{S}$, the potential outcome satisfies

equation[equation omitted — 130 chars of source]

where $\{\alpha_{d,t,s}: (d,t,s) \in \{0,1\} \times \{NT, C, AT\} \times\mathcal{S}\} $ varies with the design. Note that (ref) implies $Y_i(d) \in [0,1]$, and so we set $Y_L =0$ and $Y_H=1$. Given the realized strata, we simulate treatment assignment according to SRS or SBR, with target probability $(\pi_A(s):s \in \mathcal{S})$ that depends on the design. The description of the designs is completed as follows:

itemize• Design 1: For all $s \in \mathcal{S}$, $\alpha_{0,C,s} = 2$, $\alpha_{1,C,s} = 8$, $\alpha_{1,AT,s} = 3$, $\alpha_{0,NT,s} = 5$, and $\pi_A(s) = 0.5$. • Design 2: For all $s \in \mathcal{S}$, $\alpha_{0,C,s} = 3 + (s-1)/3$, $\alpha_{1,C,s} = 4+(s-1)/3$, $\alpha_{1,AT,s} = 5+(s-1)/3$, $\alpha_{0,NT,s} = 2+(s-1)/3$, and $\pi_A(s) = 0.5$. • Design 3: As in Design 2, but $(\pi_A(s): s\in \mathcal{S}) = (0.3,0.7,0.6,0.8)$.

It is not hard to show that all designs satisfy Assumptions (ref) and (ref).

For each design, we compute the identified sets for the ATE and ATT. For the ATE, the identified sets are $[0.425, 0.575]$ for Design 1 and $[0.040, 0.190]$ for Designs 2 and 3. For the ATT, the identified sets are $[0.463, 0.568]$ for Design 1, $[0.047, 0.153]$ for Design 2, and $[0.055, 0.145]$ for Design 3. We note that the LATE is 0.6 for Design 1 and 0.1 for Designs 2 and 3, indicating that this parameter may or may not be included in the identified sets of the ATE and ATT.

Tables (ref) and (ref) provide results with $n=500$, $\alpha = 5\%$, and $5,000$ replications. We begin with Table (ref), which presents the ATE results. In this case, recall that we recommend estimating the bounds using sample analog treatment assignment probabilities (i.e., “sample”) rather than target probabilities (i.e., “target”). Our asymptotic theory predicts that the rejection rates for $H: \theta = \theta_0$ for any $\theta_0$ in the identified set should not exceed $5\%$ as the sample size grows. Rejection rates in the interior of the identified set are typically much smaller than $5\%$, so we focus on the boundary, i.e., $\theta_0 = \theta_L$ or $\theta_0 = \theta_H$. The results show that the rejection rate at either boundary point is close to $5\%$, regardless of which CAR method is used or how treatment assignment probabilities are estimated. This is indicative that our asymptotic analysis is accurate with $n=500$. Next, we turn to the rejection rates for points outside of the identified set, such as $\theta_0 = \theta_L \times 0.9$ or $\theta_0 = \theta_H \times 1.1$. As expected, our simulations show that the rejection rate at either of these points is much larger than $5\%$. Notably, our recommended bounds estimator (i.e., “sample”) delivers a higher rejection rate than the alternative bounds estimator (i.e., “target”), especially with SRS.\footnote{Our results imply that both bounds estimators have equal asymptotic distribution under SBR.} Relatedly, the confidence interval of our recommended bounds estimator is, on average, shorter than that of the alternative bounds estimator.

table[table omitted — 2,194 chars of source]

We now turn to Table (ref), which provides results for the ATT. In this case, recall that we recommend estimating the bounds for the ATT by using the target probabilities for the treatment assignment probabilities in the numerator, and the sample analogs for those in the denominator (i.e., the “target/sample” combination). Any other choice of estimators for the treatment assignment probabilities may result in lower-quality inference. To verify this, one could consider alternative combinations of estimators for the treatment assignment probabilities. For brevity, we focus on the most natural alternative, resulting from using the sample analog to estimate both probabilities (i.e., the “sample/sample” combination).

The empirical findings for the ATT are qualitatively similar to those for the ATE. The rejection rate at either boundary point is close to $5\%$, regardless of how we implement the CAR or how we estimate the treatment assignment probabilities. The recommended bounds estimator produces confidence intervals that are slightly shorter and have a marginally higher rejection probability outside the identified set compared to the alternative bounds estimator. The main quantitative distinction with Table (ref) is that the differences between the recommended bounds estimator and the alternative one are very small. By inspecting our formal results, we can see that the differences between the asymptotic variances of the two bounds estimators are indeed small in magnitude across simulation designs.

table[table omitted — 2,403 chars of source]

Empirical application

In this section, we revisit dupas/karlan/robinson/ubfal:2018, who conducted an RCT to assess the economic impact of expanding basic bank account access in three countries: Malawi, Uganda, and Chile.\footnote{The data are available at \url{https://www.aeaweb.org/articles?id=10.1257/app.20160597}.} Recently, bugni/gao:2023 focused on the RCT in Uganda and conducted inference on the (point-identified) LATE. We now use the analysis in this paper to perform inference on the ATE and ATT for this RCT.

We now briefly summarize the empirical setting of the RCT in Uganda; see dupas/karlan/robinson/ubfal:2018 for a more detailed description. dupas/karlan/robinson/ubfal:2018 randomly selected 2,159 Ugandan households without bank accounts in 2011 and assigned them to treatment or control groups. Treated households received a voucher for a free savings account and assistance with the paperwork, while control group households did not receive these vouchers. We use the binary variable $A_i \in \{0,1\}$ to indicate if household $i$ was given the voucher for the free savings account. Households were stratified by gender, occupation, and bank branch, creating 41 strata, i.e., $s \in \mathcal{S} = \{1, \ldots, 41\}$, and were assigned to treatment or control using SBR with $\pi_A(s) = 1/2$ for all $s \in \mathcal{S}$. We define the binary variable $D_i = D_i(A_i) \in \{0,1\}$ to indicate if household $i$ has opened and used the free savings account in this RCT.

This RCT featured one-sided noncompliance. dupas/karlan/robinson/ubfal:2018 reveals that none of the 1,079 households in the control group accessed a free savings account. In contrast, among the 1,080 treated households, only 54% actually opened a bank account, and only 42% made at least one deposit during the RCT. We use three outcomes measured in 2010 US dollars: savings in formal financial institutions, savings in cash at home or in a secret place, and expenditures in the last month.

There are two relevant aspects of this empirical exercise worth discussing. First, due to one-sided noncompliance and the constant target assignment probability, the ATT and the LATE coincide (both point identified). As a corollary, our results for the ATT should align with those obtained for the LATE by bugni/gao:2023. Second, our inference methods require knowledge of the logical lower and upper bounds for the outcomes of interest. Assuming that the outcome variables are bounded, we consistently estimate these bounds using the sample minimum and maximum. The minimum value of all variables is zero dollars. The maximum values for savings in formal financial institutions, savings in cash at home or in a secret place, and expenditures in the last month are 293.3, 366.6, and 256.6 dollars, respectively.

Table (ref) presents the empirical results for the ATE and ATT. For the sake of comparison, we also include the results for the LATE obtained by bugni/gao:2023. As expected, the identified set of the ATT is estimated to be a point, which coincides with the LATE estimator obtained by bugni/gao:2023.\footnote{The differences between the CIs are caused by slight variations in the estimators for the standard errors and are asymptotically negligible.} In contrast, the ATE is partially identified. We estimate that the ATE of opening and using these savings accounts on the amount saved in formal financial institutions is between 4.03 and 167.86 dollars, with a 95% CI between 1.36 and 175.18 dollars. Additionally, we estimate that the ATE of opening and using these savings accounts on the amount saved in cash or at home is between -14.54 and 190.25 dollars, with a CI spanning between -17.51 and 199.70 dollars. Finally, we estimate that the ATE of opening and using these savings accounts on expenditures is between -18.26 and 125.09 dollars, with a CI spanning between -20.92 and 131.21 dollars. While these intervals are admittedly wide, it is relevant to note that the corresponding identified sets are sharp, meaning they cannot be restricted further without adding additional information. Put differently, the only way to shrink the identified set—and consequently the resulting estimates and confidence sets—is to impose additional assumptions on the model.

table[table omitted — 1,286 chars of source]

Conclusions

This paper studied the identification and inference of the ATE and ATT in RCTs that have CAR and imperfect compliance. The interplay between these two aspects of the RCT makes our analysis novel relative to the existing literature. The imperfect compliance implies that our estimands are partially identified, contrasting with the CAR literature focusing on point-identified parameters. In fact, to our knowledge, ours is the first paper to consider partial identification analysis in RCTs with CAR. In turn, the treatment assignment using CAR implies that our data may not be i.i.d., which is non-standard in the identification analysis of the ATE and ATT.

We derive the identified set of the ATE and ATT. Due to CAR, this characterization does not follow from the existing literature under i.i.d.\ assumptions. In fact, our results appear to be the first characterization of these identified sets in the CAR literature with imperfect compliance. Our identified sets are intervals with endpoints given by sharp bounds. We provide estimators for these sharp bounds and demonstrate that they are asymptotically normally distributed, uniformly across a relevant class of probability distributions. This result enables us to (i) estimate the identified sets of the ATE and ATT uniformly consistently and (ii) provide uniformly valid confidence sets for the ATE and ATT.

An important aspect of our inference is the estimator of the treatment assignment probabilities used to construct the estimated sharp bounds. The two natural options are (i) the target values of the treatment assignment probabilities (known in an RCT) or (ii) the sample analog of the treatment assignment probabilities. Our asymptotic analysis yields concrete practical recommendations. For the ATE, we recommend using the sample analog estimators for these probabilities. In the case of the ATT, we recommend using the target values in the numerator and the sample analogs in the denominator. There are three reasons behind these recommendations. First, our recommended option produces estimated bounds that are relatively more efficient. Second, and relatedly, our recommended option leads to confidence sets that reject local alternative hypotheses with relatively higher probability than the alternative option. Finally, our recommended estimated bounds can be obtained under weaker assumptions than the alternative options.

Using Monte Carlo simulations, we confirm that our asymptotic predictions and recommendations are relevant in finite samples. Finally, we illustrate our methodology with an empirical application to the RCT implemented in dupas/karlan/robinson/ubfal:2018.

A natural extension of our results is considering parameters of interest beyond the ATE and the ATT. In line with this, we study the identification and inference for the average treatment effect on the untreated (ATU). Qualitatively speaking, the results resemble those obtained for the ATT. For brevity, we place them in Section (ref) of the Appendix. An interesting avenue for future research is to generalize the framework to encompass parameters representable via marginal treatment effects of bjorklund/moffit:1987 and heckman/vytlacil:1999. We are actively pursuing research in this direction.

appendix