Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
87,926 characters · 16 sections · 89 citation commands
Bootstrap Inference for Quantile Treatment Effects in Randomized Experiments with Matched Pairs
Matched-pairs designs (MPDs) have recently seen widespread and increasing use in various randomized experiments conducted by economists. By MPD we mean a randomization scheme that first pairs units based on the closeness of their baseline covariates and then randomly assigns one unit in the pair to be treated. In development economics, researchers routinely pair villages, neighborhoods, microenterprises, or townships in their experiments (BDGK15, BDGK15; crepon2015, crepon2015; glewwe2016, glewwe2016; groh2016, groh2016). In labor economics, especially in the field of education, researchers pair schools or students to evaluate the effects of various education interventions (angrist2009, angrist2009; beuermann2015, beuermann2015; fryer2017management, fryer2017management; fryer2017, fryer2017; bold2018, bold2018; fryer2018, fryer2018). B09 surveyed leading experts in development field experiments and reported that 56% of them explicitly match pairs of observations on baseline characteristics.
Researchers often use randomized experiments to estimate quantile treatment effects (QTEs) as well as average treatment effects (ATEs). Quantile effects can capture heterogeneity in both the sign and magnitude of treatment effects, which may vary according to position within the distribution of outcomes. A common practice in conducting inference on QTEs is to use bootstrap rather than analytical methods because the latter usually require tuning parameters in implementation. However, in MPDs, the treatment statuses are negatively dependent within pairs because exactly half of the units are treated; covariates across pairs are dependent as well because of pair matching. Neither the standard multiplier bootstrap nor bootstrapping the pairs mimics such dependence structure. This difficulty raises the question of how to conduct bootstrap inference for QTEs in MPDs in a manner that mitigates these shortcomings.
To tackle these shortcomings we propose two bootstrap inference methods: the gradient bootstrap and the inverse propensity score weighted (IPW) multiplier bootstrap. We first show that the gradient bootstrap can consistently approximate the limit distribution of the QTE estimator under MPDs uniformly over a compact set of quantile indexes. h17 proposed using the gradient bootstrap for the cluster-robust inference in linear quantile regression models. Like h17, we rely on the gradient bootstrap to avoid estimating the Hessian matrix that involves the infinite-dimensional nuisance parameters. The gradient bootstrap procedure is therefore free of tuning parameters. On the other hand and differing from h17, we construct a specific perturbation of the score based on pair and adjacent pairs of observations, which can capture the dependence structure in the original data.
To implement our gradient bootstrap method, researchers need to know the identities of pairs. Such information may not be available when they are using an experiment that was run by other investigators in the past and the randomization procedure may not have been fully described. For example, publicly available datasets for papers such as panagopoulos2008 and butler2010 contain no information on pair identities.\footnote{\doublespacing Both datasets are available in the data archive of the Institute for Social and Policy Studies at Yale University (https://isps.yale.edu/research/data).} B09 also pointed out that many papers in existing experiments do not describe the randomization procedure in detail.
To address these issues, we next propose an IPW multiplier bootstrap, which can be implemented without the knowledge of pair identities. We show that such a bootstrap can consistently approximate the limit distribution of the QTE estimator under MPDs. There is a cost to not using information about pair identities as the method requires one tuning parameter for the nonparametric estimation of the propensity score. In spite of this additional cost, this multiplier bootstrap method still has an advantage over direct analytic inference because practical implementation of the latter requires more than one tuning parameter.
The contributions in the present paper relate to other recent research. BRS19 first pointed out that in MPDs the two-sample $t$-test for the null hypothesis that the ATE equals a pre-specified value is conservative. They then proposed adjusting the standard error of the estimator and studied the validity of the permutation test. This paper complements those results by considering the QTEs and by developing new methods of bootstrap inference. Unlike the permutation test, our methods of bootstrap inference do not require studentization, which is cumbersome in the QTE context. In addition, our multiplier bootstrap method complements their results by providing a way to perform inference relating to both ATEs and QTEs when pair identities are unknown. In other work, B19 investigated the optimality of MPDs in randomized experiments. ZZ20 considered bootstrap inference under covariate-adaptive randomization. A key difference in our contribution is that in MPDs the number of strata is proportional to the sample size, whereas in covariate-adaptive randomization that number is fixed. In consequence, the present work uses fundamentally different asymptotic arguments and bootstrap methods from those employed by ZZ20. The present paper also fits within a growing literature that studies inference in randomized experiments (e.g., HHK11, athey2017, abadie2018, BCS17, T18, and BCS18, among others).
The remainder of the paper is organized as follows. Section (ref) describes the model setup and notation. Section (ref) develops the asymptotic properties of our QTE estimator. In Section (ref) we study the naive multiplier bootstrap, the naive multiplier bootstrap of the pairs, the gradient bootstrap, and the IPW multiplier bootstrap. Section (ref) provides computational details and recommendations for practitioners. Section (ref) reports simulation results. Section (ref) provides an empirical application of our methods of bootstrap inference to the data in groh2016, examining both the ATEs and QTEs of macroinsurance on consumption and profits. Section (ref) concludes. Proofs of all results and additional simulations are in the Online Supplement.
Denote the potential outcomes for treated and control groups as $Y(1)$ and $Y(0)$, respectively. Treatment status is written as $A$, where $A=1$ is treated and $A=0$ is untreated. The researcher only observes $\{Y_i,X_i,A_i\}_{i=1}^{2n}$ where $Y_i = Y_i(1)A_i + Y_i(0)(1-A_i)$, and $X_i \in \Re^{d_x}$ is a collection of baseline covariates, where $d_x$ is the dimension of $X$. The parameter of interest is the $\tau$th QTE, denoted as
where $q_1(\tau)$ and $q_0(\tau)$ are the $\tau$th quantiles of $Y(1)$ and $Y(0)$, respectively. The testing problems of interest involve single, multiple, or even a continuum of quantile indexes, as in the following null hypotheses
for some pre-specified value $\underline{q}$ or function $\underline{q}(\tau)$, where $\Upsilon$ is some compact subset of $(0,1)$.
The units are grouped into pairs based on the closeness of their baseline covariates, which is now made clear. Pairs of units are denoted
where $[n] = \{1,\cdots,n\}$ and $\pi$ is a permutation of $2n$ units based on $\{X_i\}^{2n}_{i=1}$ as specified in Assumption (ref)(iv) below. Within a pair, one unit is randomly assigned to treatment and the other to control. Specifically, we make the following assumption on the data generating process (DGP) and the treatment assignment rule.
Assumption (ref) is used in BRS19 to which we refer readers for more discussion. In Assumption (ref)(iv), $||\cdot||_2$ denotes Euclidean distance. However, all our results hold if $||\cdot||_2$ is replaced by any distance that is equivalent to it, such as $L_\infty$ distance, $L_1$ distance, and the Mahalanobis distance when all the eigenvalues of the covariance matrix are bounded and bounded away from zero.
Let $\hat{q}_1(\tau)$ and $\hat{q}_0(\tau)$ be the $\tau$th percentiles of outcomes in the treated and control groups, respectively. Then, the $\tau$th QTE estimator we consider is just
For ease of notation, dependence of $\hat{q}(\tau)$, $\hat{q}_1(\tau)$, $\hat{q}_0(\tau)$ and all the other estimators on $n$ is suppressed throughout the rest of the paper. To facilitate further analysis and motivate our bootstrap procedure, we note that $\hat{q}(\tau)$ can be equivalently computed by direct quantile regression. Let
where $\dot{A}_i = (1,A_i)^\top$ and $\rho_\tau(u) = u(\tau - 1\{u\leq 0\})$. Then, $\hat{q}(\tau) = \hat{\beta}_1(\tau)$ and $\hat{q}_0(\tau) = \hat{\beta}_0(\tau)$.
Assumption (ref)(i) is a standard regularity condition widely assumed in quantile estimation. The Lipschitz conditions in Assumptions (ref)(ii) and (ref)(iii) are similar in spirit to those assumed in BRS19 and ensure that units that are “close" in terms of their baseline covariates are suitably comparable. For $a = 0,1$, let $m_{a,\tau}(x,q) = \mathbb{E}(\tau - 1\{Y(a) \leq q\}|X=x)$ and $m_{a,\tau}(x) = m_{a,\tau}(x,q_a(\tau))$.
Several remarks are in order. First, the asymptotic variance of $\hat{q}(\tau)$ under MPDs is
Note further that the asymptotic variance of $\hat{q}(\tau)$ under simple random sampling (SRS)\footnote{By simple random sample, we mean the treatment status is assigned independently with probability $1/2$.} is
It is clear that
Equality in the last expression holds when both $m_{1,\tau}(X)$ and $m_{0,\tau}(X)$ are zero, which implies that $X$ is irrelevant to the $\tau$th quantiles of $Y(0)$ and $Y(1)$.
Second, note that $\hat{q}(\tau)$ has the same asymptotic variance as that for the QTE estimators studied by F07 and DH14 under SRS.
Third, to provide an analytic estimate of the asymptotic variance $\Sigma(\tau,\tau)$ it is necessary at least to estimate the infinite dimensional nuisance parameters $f_1(q_1(\tau))$ and $f_0(q_0(\tau))$, which requires two tuning parameters. Hence, if a researcher is interested in testing a null hypothesis that involves $G$ quantile indexes, $2G$ tuning parameters are needed to estimate $2G$ densities, a cumbersome task in practical work; and to construct a uniform confidence band for the QTE analytically, two tuning parameters are needed at each grid point of the quantile indexes. Moreover, if pair identities are unknown, analytic methods of inference potentially require nonparametric estimation of the quantities $m_{a,\tau}(\cdot)$ for $a = 0,1$ as well. There are other practical difficulties. Nonparametric estimation is sometimes sensitive to the choice of tuning parameters and rule-of-thumb tuning parameter selection may not be appropriate for every data generating process (DGP) or every quantile. Use of cross-validation in selecting the tuning parameters is possible in principle but in practice time-consuming. These practical difficulties of analytic methods of inference provide a strong motivation to investigate bootstrap inference procedures that are much less reliant on tuning parameters.
This section examines four bootstrap inference procedures for the QTEs in MPDs. We first show that both the naive multiplier bootstrap and the naive multiplier bootstrap of the pairs fail to approximate the limit distribution of the QTE estimator derived in Section (ref). We then propose two bootstrap methods that can consistently approximate the limit distribution of the QTE estimator.
Consider first the naive multiplier bootstrap estimators of $\hat{\beta}_0(\tau)$ and $\hat{\beta}_1(\tau)$, defining
where $\xi_i$ is the bootstrap weight defined in the next assumption.
In practice, we generate $\xi_{i}$ independently from the standard exponential distribution. Denote $\hat{q}^m(\tau) = \hat{\beta}_1^m(\tau)$ and recall that $\hat{q}(\tau) = \hat{\beta}_1(\tau)$.
Three remarks are in order. First, $\sqrt{n}(\hat{q}^m(\tau) - \hat{q}(\tau))$ and $\mathcal{B}^m(\tau)$ are viewed as processes indexed by $\tau \in \Upsilon$ and denoted by $G_n$ and $G$, respectively. Then, following VW96, we say $G_n$ weakly converges to $G$ conditionally on data and uniformly over $\tau \in \Upsilon$ if
where $\text{BL}_1$ is the set of all functions $h:\ell^\infty(\Upsilon) \mapsto [0,1]$ such that $|h(z_1)-h(z_2)| \leq ||z_1-z_2||_\infty$ for every $z_1,z_2 \in \ell^\infty(\Upsilon)$, and $\mathbb{E}_\xi$ denotes expectation with respect to the bootstrap weights $\{\xi\}_{i=1}^n$.\footnote{Asymptotic measurability holds in our setting from VW96, which requires asymptotic tightness of the bootstrap process. The latter has been established in the proof of Theorem (ref). For notational simplicity, we ignore the issue of asymptotic measurability.} The same remark applies to Theorems (ref), (ref), and (ref) below.
Second, $\Sigma^\dagger(\tau,\tau')$ is just the covariance kernel of the QTE estimator when simple random sampling (instead of the MPD) is used as the treatment assignment rule. It follows that the naive multiplier bootstrap fails to approximate the limit distribution of $\hat{q}(\tau)$ ($\hat{\beta}_1(\tau)$). The intuition is straightforward. Given the data, the bootstrap weights are i.i.d. and thus unable to mimic the cross-sectional dependence in the original sample.
Third, it is possible to consider the conventional nonparametric bootstrap in which the bootstrap sample is generated from the empirical distribution of the data. If the observations are i.i.d., VW96 showed that the conventional bootstrap is first-order equivalent to a multiplier bootstrap with Poisson(1) weights. However, in the current setting, $\{A_i\}_{i \in [2n]}$ are dependent. It is technically challenging to show rigorously that the above equivalence still holds and this challenge is left for future research.
Next consider the naive multiplier bootstrap of the pairs which uses the same bootstrap multiplier for the observations within the pair. Let
where $\xi_i^p$ is the bootstrap weight defined in the next assumption.
Because the units in the same pair share the same multiplier, we call this the naive multiplier bootstrap of the pairs. Denote $\hat{q}^p(\tau) = \hat{\beta}_{1}^p(\tau)$.
Three remarks are in order. First, Theorem (ref) implies that bootstrapping pairs of observations alone is unable to mimic the dependence structure in the original sample. In MPDs, covariates across pairs are dependent. However, bootstrapping pairs misses such dependence by assigning independent multipliers to pairs.\footnote{We provide an illustrative example in Section A of the supplement to show the dependence of covariates across pairs. We thank the editor for an inspiring and helpful discussion on the failure of bootstrapping pairs.}
Second, the variances of the original estimator $\hat{q}(\tau)$ under MPD and $\hat{q}^p(\tau)$ conditional on data have the following relationship:
Third, the gradient bootstrap procedure proposed below is based on a similar idea and uses the same weight for the observations within the pair to construct the score $S_{n,1}^*$ defined in (ref). But in order to construct a final score that exactly mimics the dependence in the data, an extra score component, $S_{n,2}^*$ defined in (ref) below, is needed. This component is constructed based on adjacent pairs of observations.
We develop an approximation for the asymptotic distribution of the QTE estimator via the gradient bootstrap. Let $u = \sqrt{n}(b - \beta(\tau))$ be a localizing estimation error parameter. From the derivations in Theorem (ref), we see that
where
and
The gradient bootstrap proposes to perturb the objective function by some random error $S_n^*(\tau)$, which will be specified later. This error in turn perturbs the score function $S_n(\tau)$. The corresponding bootstrap estimator $\hat{\beta}^*(\tau)$ solves the following optimization problem
We can then show that
Taking the difference between (ref) and (ref) gives
The second element of $\hat{\beta}^*(\tau)$ in (ref) is the bootstrap version of the QTE estimator, which is denoted $\hat{q}^*(\tau)$. By solving (ref) we avoid estimating the Hessian $Q(\tau)$, which involves infinite-dimensional nuisance parameters. Then, for the gradient bootstrap to consistently approximate the limit distribution of the original estimator $\hat{\beta}(\tau)$, we need only construct $S_n^*(\tau)$ in such a way that its weak limit given the data coincides with that of the original score $S_n(\tau)$.
Accordingly, we now show how to specify $S_n^*(\tau)$. Let $\{\eta_j\}_{j=1}^{n}$ and $\{\hat{\eta}_k\}_{k=1}^{\lfloor n/2 \rfloor}$ be two mutually independent i.i.d. sequences of standard normal random variables. Use the indexes $(j,1),(j,0)$ to denote the indexes in $(\pi(2j-1),\pi(2j))$ with $A=1$ and $A=0$, respectively. For example, if $A_{\pi(2j)} = 1$ and $A_{\pi(2j-1)}=0$, then $(j,1) = \pi(2j)$ and $(j,0) = \pi(2j-1)$. Similarly, use indexes $(k,1),\cdots, (k,4)$ to denote the first index in $(\pi(4k-3),\cdots,\pi(4k))$ with $A=1$, the first index with $A=0$, the second index with $A=1$, and the second index with $A=0$, respectively. Now let
where
and
The perturbation $S_{n}^*$ consists of two parts: $S_{n,1}^*$ and $S_{n,2}^*$. The first part $S_{n,1}^*$ is constructed by pairs of observations, based on the idea of bootstrapping the pairs. But bootstrapping the pairs alone cannot fully capture the dependence structure in the MPD, as shown in Section (ref). So $S_{n,1}^*$ is adjusted by adding a term capturing the remaining dependence. This second term, $S_{n,2}^*$, is motivated by the idea of using adjacent pairs to adjust the standard error of the ATE estimator under MPDs in BRS19.
In Section (ref) we show how to compute the bootstrap estimator $\hat{\beta}^*(\tau)$ directly from the sub-gradient condition of (ref). This method avoids the optimization inherent in (ref) and computation is fast. The following assumption imposes the condition that baseline covariates in adjacent pairs are also `close'.
Assumption (ref) and Assumption (ref)(iv) are jointly equivalent to BRS19. We refer readers to BRS19 for further discussion of this assumption. In particular, BRS19 show that it is possible to implement the matching algorithm to re-order pairs so that both Assumption (ref) and Assumption (ref)(iv) hold automatically. We provide more detail in Section (ref).
Define $\hat{q}^*(\tau) = \hat{\beta}_1^*(\tau)$ and recall that $\hat{q}(\tau) = \hat{\beta}_1(\tau)$. We have the following result.
Two remarks on Theorem (ref) are in order. First, the bootstrap estimator $\hat{q}^*(\tau)$ has the following objectives: (i) to avoid estimating densities; and (ii) to mimic the distribution of the original estimator $\hat{\beta}(\tau)$ under MPDs. Objective (i) relates to the Hessian ($Q$) and (ii) to the score ($S_n$) of the quantile regression. The gradient bootstrap provides a flexible approach to achieve both goals.
Second, to implement the gradient bootstrap, researchers need to know identities of pairs. This information may not be available when the experiment was run by others and the randomization procedure was not fully detailed. In such cases, we propose IPW multiplier bootstrap inference for the QTE, whose validity is established in the next section.
In empirical research, researchers may not know the identities of pairs when they are using an experiment that was run by other investigators in the past and the randomization procedure may not have been fully described. For example, publicly available datasets for papers such as panagopoulos2008 and butler2010 contain no information on pair identities. B09 also pointed out that many papers in existing experiments do not describe the randomization procedure in detail. In this section we establish the validity of IPW multiplier bootstrap inference for the QTE, showing that the procedure can be implemented without the knowledge of pair identities.
We use the sieve method to nonparametrically estimate the propensity score. Let $b(X)$ be the $K$-dimensional sieve basis on $X$ and $\hat{A}_i$ be the estimated propensity score for the $i$th individual. Then,
where $\hat{\theta} = \operatorname*{arg\,min}_\theta \sum_{i=1}^{2n}\xi_i(A_i - b^\top(X_i)\theta)^2$ and $\xi_i$ is the bootstrap weight defined in Assumption (ref).
Because the true propensity score is $1/2$, by setting the first component of $b(X)$ to unity, we have $1/2 = b^\top(X)\theta_0$ where $\theta_0 = (0.5,0,\cdots,0)^\top$. The linear probability model for the propensity score is correctly specified. It is possible to use sieve logistic regression to compute the propensity score, as done by HIR03, F07, and DH14. The main benefit of using logistic regression is to guarantee that the estimated propensity score lies between zero and one. However, in MPDs, the estimated propensity score is always very close to 0.5. Therefore, for simplicity, we use a linear sieve regression here.
The IPW multiplier bootstrap estimator can be computed as
where
Two remarks are in order. First, requiring $X$ to have a compact support is common in nonparametric sieve estimation. Second, the quantity $\zeta(K)$ depends on the choice of basis functions. For example, $\zeta(K) = O(K^{1/2})$ for splines and $\zeta(K) = O(K)$ for power series.\footnote{\doublespacing See c07 for a full discussion of the sieve method.} Taking splines as an example, Assumption (ref)(iii) requires $K = o(n^{1/3})$. Assumption (ref)(iv) is standard because $K \ll n$. Assumption (ref)(v) requires that the approximation error of $m_{a,\tau}(x)$ via a linear sieve function is sufficiently small. For instance, suppose $m_{a,\tau}(x)$ is s-times continuously differentiable in $x$ with all derivatives uniformly bounded by some constant $\overline{C}$, then $\sup_{a=0,1,\tau\in \Upsilon, x \in \text{Supp}(X)}|B_{a,\tau}(x)| = O(K^{-s/d_x})$. Assumptions (ref)(iii) and (ref)(v) imply that $K = n^{h}$ for some $h \in (d_x/(2s), 1/3)$, which implicitly requires $s>3d_x/2$. The choice of $K$ reflects the usual bias-variance trade-off and is the only tuning parameter that researchers need to specify when implementing this bootstrap method.
To understand the need to nonparametrically estimate the propensity score in the bootstrap sample, note that there are two stages to statistical inference in randomized experiments: the design stage and the analysis stage. In the design stage, researchers can use either simple random sampling (SRS) or matched-pairs design (MPD).\footnote{Simple random sampling means that treatment status is assigned independently with probability $1/2$. Note that SRS and MPD share the same true propensity score $1/2$. But MPD achieves the strong balance that exactly half of the units are treated whereas SRS does not.} In the analysis stage, researchers can choose either the true ($1/2$ in MPDs) or the nonparametrically estimated propensity score to construct the estimator.\footnote{Specifically, we have $\hat{q}_{ipw,1}(\tau) = \operatorname*{arg\,min}_q \sum_{i \in [2n]} \frac{A_i}{\hat{\pi}(X_i)}\rho_\tau(Y_i - q)$ and $\hat{q}_{ipw,0}(\tau) = \operatorname*{arg\,min}_q \sum_{i \in [2n]} \frac{1-A_i}{1-\hat{\pi}(X_i)}\rho_\tau(Y_i - q)$, where $\hat{\pi}(X_i)$ is an estimator of the propensity score. When we use the true score, i.e., $\hat{\pi}(X_i) = 1/2$, we have $\hat{q}_{ipw,a} = \hat{q}_a$ as defined in the paper for $a=0,1$. However, we can also let $\hat{\pi}(X_i)$ be the nonparametric estimator of the propensity score.} The asymptotic variances of the QTE estimator with true and estimated propensity scores under SRS are $\Sigma^\dagger(\tau,\tau)$ defined in (ref) and $\Sigma(\tau,\tau)$ defined in (ref), respectively, which are derived by H98 and F07; the asymptotic variance of the QTE estimator with the true propensity score under MPD is $\Sigma(\tau,\tau)$, which is shown in Theorem (ref). These results are summarized in the following table.
Note that the asymptotic variances for the QTE estimator under MPD with the true score and that under SRS with the nonparametrically estimated score are the same. If we conduct multiplier bootstrap inference, conditionally on data, the bootstrap sample of observations is independent. Therefore, in order for the multiplier bootstrap estimator to mimic the asymptotic behavior of the original estimator under MPD with the true score, we need to nonparametrically estimate the propensity score in the bootstrap sample.
The benefit of the IPW multiplier bootstrap is that it does not require knowledge of the pair identities. The cost is that we have to nonparametrically estimate the propensity score, which requires one tuning parameter and is subject to the usual curse of dimensionality. Nonetheless, we still prefer this bootstrap method of inference to the analytic approach. Analytic estimation of the standard error of the QTE estimator without the knowledge of pair identities requires nonparametric estimation of $\{m_{a,\tau}(X),f_a(q_a(\tau))\}_{a=0,1}$, which involves four tuning parameters. The number of tuning parameters further increases with the number of quantile indexes involved in the null hypothesis. To construct uniform confidence bands for QTE over $\tau$, we require $4G$ tuning parameters for grid size $G$. By contrast, implementation of the IPW multiplier bootstrap requires estimation of the propensity score only once, and thus, the use of a single tuning parameter.
Inference concerning the ATE in MPDs can also be accomplished via a similar IPW multiplier bootstrap procedure. We can show that such a bootstrap can consistently approximate the asymptotic distribution of the ATE estimator under MPDs. This result complements that established by BRS19 because it provides a way to make inferences about the ATE in MPDs when information on pair identities is unavailable. That pair identity information is required by BRS19 in computing standard errors for their adjusted $t$-test.
In practice, the order of pairs in the dataset is usually arbitrary and does not satisfy Assumption (ref). To apply the gradient bootstrap, researchers first need to re-order the pairs. For the $j$th pair with units indexed by $(j,1)$ and $(j,0)$ in the treatment and control groups, let $\overline{X}_j = \frac1{2}\{X_{(j,1)} + X_{(j,0)}\}$. Then, let $\overline{\pi}$ be any permutation of $n$ elements that minimizes
The pairs are re-ordered by indexes $\overline{\pi}(1),\cdots, \overline{\pi}(n)$. With an abuse of notation, we still index the pairs after re-ordering by $1,\cdots,n$. Note that the original QTE estimator $\hat{q}(\tau) = \hat{q}_1(\tau)-\hat{q}_0(\tau)$ is invariant to the re-ordering.
For the bootstrap sample, we directly compute $\hat{\beta}^*(\tau)$ from the sub-gradient condition of (ref). Specifically, we compute $\hat{\beta}_0^*(\tau)$ as $Y^0_{(h_0)}$ and $\hat{q}^*(\tau) \equiv \hat{\beta}_1^*(\tau)$ as $Y^1_{(h_1)} - Y^0_{(h_0)}$, where $Y^0_{(h_0)}$ and $Y^1_{(h_1)}$ are the $h_0$th and $h_1$th order statistics of outcomes in the treatment and control groups, respectively,\footnote{\doublespacing We assume $Y^a_{(1)} \leq \cdots \leq Y^a_{(n)}$ for $a = 0,1$.} and $h_0$ and $h_1$ are two integers satisfying
with
As the probability of $n\tau + T_{n,a}^*(\tau)$ being an integer is zero, $h_a$ is uniquely defined with probability one.\footnote{The sub-gradient condition of (ref) is $\hat{q}^*(\tau) = Y_{i_1}- Y_{i_0}$ such that $i_i,i_0 \in [2n]$ are two indexes, $A_{i_1}= 1$, $A_{i_0}= 0$,
By letting $h_1 = \sum_{i \in [2n]} A_i 1\{ Y_i < Y_{i_1}\}+1$ and $h_0 = \sum_{i \in [2n]} (1-A_i) 1\{ Y_i < Y_{i_0}\}+1$, we have $Y_{i_1} = Y_{(h_1)}^1$ and $Y_{i_0} = Y_{(h_0)}^0$. }
We summarize the steps in the bootstrap procedure as follows.
We first provide more details on the sieve bases. Let $b(x) \equiv (b_{1}(x), \cdots, b_{K}(x))^\top$, where $\{b_{k }(\cdot)\}_{k=1}^{K}$ are $K$ basis functions of a linear sieve space $\mathcal{B}$. Given that all $d_x$ elements of $X$ are continuously distributed, the sieve space $\mathcal{B}$ can be constructed as follows.
In practice, we suggest not using all the tensor products. Otherwise the dimension $J_n^{d_x}$ can be too large even for moderate sample sizes. Instead, when the number of pairs is $50$ or $100$, we suggest using splines with one node (usually the median) for each dimension and one interaction term across each pair of dimensions. In the appendix, we also propose a cross-validation method to select the basis functions.
Given the sieve bases, we can estimate the propensity score following (ref). We then obtain $\hat{q}_{ipw,1}^w(\tau)$ and $\hat{q}_{ipw,0}^w(\tau)$ by solving the sub-gradient conditions for the two optimizations in (ref). Specifically, we have $\hat{q}_{ipw,1}^w(\tau) = Y_{h_1'}$ and $\hat{q}_{ipw,0}^w(\tau) = Y_{h_0'}$, where the indexes $h_0'$ and $h_1'$ satisfy $A_{h_a'} = a$, $a=0,1$,
and
In practical implementation we set $\{\xi_i\}_{i \in [2n]}$ as i.i.d. standard exponential random variables. In this case, all the equalities in (ref) and (ref) hold with probability zero. Thus, $h_1'$ and $h_0'$ are uniquely defined with probability one.
The IPW multiplier bootstrap can also be used to infer the ATE when the pairs identities are unknown. The point estimator of ATE under MPD is just the difference in mean estimator: $\hat{\Delta} = \frac{1}{n}\sum_{i =1}^{2n}(A_iY_i - (1-A_i)Y_i)$. In order to mimic its limit distribution, we propose to an IPW multiplier bootstrap for the ATE estimator
We summarize the bootstrap procedure as follows.
For comparison, we also consider the naive multiplier bootstrap and the naive multiplier bootstrap of the pairs in our simulations. The computation of the naive multiplier bootstrap follows a procedure similar to the above with only one difference: the nonparametric estimate $\hat{A}_i$ of the propensity score is replaced by the truth, that is, $1/2$. The computation of the naive multiplier bootstrap of the pairs follows a similar procedure to that of the naive multiplier bootstrap except that the units in the same pair share the same multiplier.
Given the bootstrap estimates, we discuss how to conduct bootstrap inference for the null hypotheses with single, multiple, and a continuum of quantile indexes. We take the gradient bootstrap as an example. If the IPW multiplier bootstrap is used, one can just replace $\{\hat{q}^{*b}(\tau)\}_{b \in [B],\tau \in \mathcal{G}}$ by $\{\hat{q}_{ipw}^{w,b}(\tau)\}_{b \in [B],\tau \in \mathcal{G}}$ in the following cases. The same procedure applies to the bootstrap inference of ATE as well.
Case (1). We aim to test the single null hypothesis that $\mathcal{H}_0: q(\tau) = \underline{q}$ vs. $q(\tau) \neq \underline{q}$. Let $\mathcal{G} = \{\tau\}$ in the procedures described above. Further denote $\mathcal{Q}(\nu)$ as the $\nu$th empirical quantile of the sequence $\{\hat{q}^{*b}(\tau)\}_{b \in [B]}$. Let $\alpha \in (0,1)$ be the significance level. We suggest using the bootstrap estimator to construct the standard error of $\hat{q}(\tau)$ as $\hat{\sigma} = \frac{\mathcal{Q}(0.975)- \mathcal{Q}(0.025)}{C_{0.975} - C_{0.025}}$, where $C_{\mu}$ is the $\mu$th standard normal critical value. Then a valid confidence interval and Wald test using this standard error are
and $1\{\left|\frac{\hat{q}(\tau) - \underline{q}}{\hat{\sigma}}\right| \geq C_{1-\alpha/2}\}$, respectively.\footnote{It is asymptotically valid to use standard and percentile bootstrap confidence intervals. In our simulations, we found that the confidence interval proposed in the paper has better finite sample performance.}
Case (2). We aim to test the null hypothesis that $\mathcal{H}_0: q(\tau_1) - q(\tau_2) = \underline{q}$ vs. $q(\tau_1) - q(\tau_2) \neq \underline{q}$. In this case, let $\mathcal{G} = \{\tau_1,\tau_2\}$. Further, let $\mathcal{Q}(\nu)$ denote the $\nu$th empirical quantile of the sequence $\{\hat{q}^{*b}(\tau_1) - \hat{q}^{*b}(\tau_2)\}_{b \in [B]}$, and let $\alpha \in (0,1)$ be the significance level. We suggest using the bootstrap standard error to construct a valid confidence interval and Wald test as
and $1\{\left|\frac{\hat{q}(\tau_1)-\hat{q}(\tau_2) - \underline{q}}{\hat{\sigma}}\right| \geq C_{1-\alpha/2}\}$, respectively, where $\hat{\sigma} = \frac{\mathcal{Q}(0.975)- \mathcal{Q}(0.025)}{C_{0.975} - C_{0.025}}$.
Case (3). We aim to test the null hypothesis that $$\mathcal{H}_0: q(\tau) = \underline{q}(\tau)~\forall \tau \in \Upsilon~\text{vs.}~q(\tau) \neq \underline{q}(\tau)~ \exists \tau \in \Upsilon.$$ In theory, we should let $\mathcal{G} = \Upsilon$. In practice, we let $\mathcal{G} = \{\tau_1,\cdots,\tau_G\}$ be a fine grid of $\Upsilon$ where $G$ should be as large as computationally possible. Further, let $\mathcal{Q}_\tau(\nu)$ denote the $\nu$th empirical quantile of the sequence $\{\hat{q}^{*b}(\tau)\}_{b \in [B]}$ for $\tau \in \mathcal{G}$. Compute the standard error of $\hat{q}(\tau)$ as
The uniform confidence band with an $\alpha$ significance level is constructed as
where the critical value $\mathcal{C}_\alpha$ is computed as
and $\tilde{q}(\tau)$ is first-order equivalent to $\hat{q}(\tau)$ in the sense that $\sup_{\tau \in \Upsilon}|\tilde{q}(\tau) - \hat{q}(\tau)|= o_p(1/\sqrt{n})$. We suggest choosing $\tilde{q}(\tau) = \frac1{2}\{\mathcal{Q}_\tau(0.975) + \mathcal{Q}_\tau(0.025)\}$ over other choices such as $\tilde{q}(\tau) =\mathcal{Q}_\tau(0.5)$ and $\tilde{q}(\tau) = \hat{q}(\tau)$ due to its better finite sample performance. We reject $\mathcal{H}_0$ at an $\alpha$ significance level if $\underline{q}(\cdot) \notin CB(\alpha).$
Our practical recommendations are straightforward. If pair identities are known, we suggest using the gradient bootstrap for inference. Otherwise, we suggest using the IPW multiplier bootstrap with a nonparametrically estimated propensity score for inference.
In this section, we assess the finite sample performance of the methods discussed in Section (ref) with a Monte Carlo simulation study. In all cases, potential outcomes for $a\in\{0,1\}$ and $1\leq i\leq2n$ are generated as
where $\mu_{a},m_{a}\left(X_{i}\right),\sigma_{a}\left(X_{i}\right)$, and $\varepsilon_{a,i}$ are specified as follows. In each of the specifications below, $n\in\{50,100\}$ and ($X_{i},\varepsilon_{0,i},\varepsilon_{1,i}$) are i.i.d. The number of replications is 10,000. For bootstrap replications we set $B=5,000$.
Pairs are determined similarly to those in BRS19. Specifically, if $X_i$ is a scalar, then pairs are determined by sorting $\{X_i\}_{i \in [2n]}$. If $X_i$ is multi-dimensional, then the pairs are determined by the permutation $\pi$ computed using the R package nbpMatching. We refer interested readers to BRS19 for more detail. After forming the pairs, we assign treatment status within each pair through a random draw from the uniform distribution over $\{(0,1),(1,0)\}$.
We examine the performance of various tests for ATEs and QTEs at the nominal level $\alpha=5\%$. For the ATE, we consider the hypothesis that
For the QTE, we consider the hypotheses that
for $\tau = 0.25$, $0.5$, and $0.75$,
and
To illustrate size and power of the tests, we set $\mathcal{H}_0: \Delta = 0$ and $\mathcal{H}_1: \Delta = 1/2$. The true value for the ATE is $0$, whereas the true values for the QTEs are simulated with a $10,000$ sample size and replications. The computational procedures described in Section (ref) are followed to perform the bootstrap and calculate the test statistics. To test the single null hypothesis involving one or two quantile indexes, we use the Wald tests specified in Section (ref). To test the null hypothesis involving a continuum of quantile indexes, we use the uniform confidence band $CB(\alpha)$ defined in Case (3) in the same section.
\newcolumntype{L}{>{\arraybackslash}X} \newcolumntype{C}{>{\arraybackslash}X}
The results for the ATEs appear in Table (ref). Each row presents a different model and each column reports the rejection probabilities for the various methods. The column `Naive' refers to the two-sample $t$-test and `Adj' refers to the adjusted $t$-test in BRS19; the column `Naive pair' corresponds to the $t$-test using the standard errors estimated by the naive multiplier bootstrap of the pairs; the column `IPW' corresponds to the $t$-test using the standard errors estimated by the IPW multiplier bootstrap ATE estimator.
We make several observations on these findings. First, the two-sample $t$-test has rejection probability under $\mathcal{H}_{0}$ far below the nominal level and is the least powerful test among the four. Second, the adjusted $t$-test has rejection probability under $\mathcal{H}_{0}$ close to the nominal level and is not conservative. This result is consistent with those in BRS19. Third, the $t$-test using the standard error estimated by bootstrapping the pairs of units alone is conservative under the null and lacks power under the alternative. Fourth, the IPW $t$-test proposed in this paper has performance similar to the adjusted $t$-test.\footnote{\doublespacing Throughout this section, for both ATE and QTE estimation, we use splines to nonparametrically estimate the propensity score in the IPW multiplier bootstrap. If dim($X_i$)=1, we choose the bases $\{1,X,X^2,[\max(X-qx_{0.5},0)]^2 \}$ where $qx_{0.5}$ is the quantile of $X$ at 50%; if dim($X_i$)=2, we choose the bases $\{1,X_{1},X_{2},\max(X_{1}-qx_{1,0.5},0), \max(X_{2}-qx_{2,0.5},0), X_{1}X_{2}\}$, where for $j=1,2$, $qx_{j,\alpha}$ is the $\alpha$th sample percentile of $X_j$. Results are similar using the cross-validation method proposed in the appendix to choose the sieve bases: for details on the cross-validation method and its simulation results, see Section H in the supplement.} Under $\mathcal{H}_{0}$, the test has rejection probability close to 5%; under $\mathcal{H}_{1}$, it is more powerful than the `naive' and `naive pair' methods and has power similar to the adjusted $t$-test. These findings indicate that the IPW $t$-test provides an alternative to the adjusted $t$-test when pair identities are unknown.
The results for QTEs are summarized in Tables (ref) and (ref). Each table has four panels (Models 1-4). Each row in the panel displays the rejection probabilities for the tests using the standard errors estimated by various bootstrap methods. Specifically, the rows `Naive', `Naive pair', `Gradient', and `IPW' respectively correspond to the results of the naive multiplier bootstrap, the naive multiplier bootstrap of the pairs, the gradient bootstrap, and the IPW multiplier bootstrap.
Table (ref) reports the empirical size and power of the tests with a single null hypothesis involving one or two quantile indexes. Columns `0.25', `0.50', and `0.75' correspond to tests with quantiles at 25%, 50%, and 75%. Column `Dif' corresponds to the test with null hypothesis (ref). As expected given Theorem (ref), the test with standard errors estimated by two naive methods performs poorly in all cases. It is conservative under $\mathcal{H}_{0}$ and lacks power under $\mathcal{H}_{1}$. In contrast, the test using the standard errors estimated by either the gradient bootstrap or the IPW multiplier bootstrap method has a rejection probability under $\mathcal{H}_{0}$ that is close to the nominal level in almost all specifications. When the number of pairs is 50, the tests in the `Dif' column constructed based on either the gradient or the IPW multiplier bootstrap method are slightly conservative. Sizes approach the nominal level when $n$ increases to $100$.
Table (ref) reports empirical size and power of the uniform confidence bands for the hypothesis specified in (ref) with a grid $\mathcal{G} = \{0.25,0.27,\cdots,0.47,0.49,0.5,0.51,0.53,\cdots,0.73,0.75\}$. The test using standard errors estimated by two naive methods has rejection probabilities under $\mathcal{H}_{0}$ far below the nominal level in all specifications. In Models 1-2, the test using standard errors estimated by either the gradient bootstrap or the IPW multiplier bootstrap yields a rejection probability under $\mathcal{H}_{0}$ that is very close to the nominal level even when the number of pairs is as small as 50. Nonetheless, in Models 3-4, the tests constructed based on both methods are conservative when the number of pairs equals 50. When the number of pairs increases to 100, both tests perform much better and have rejection probabilities under $\mathcal{H}_{0}$ that are close to the nominal level. Under $\mathcal{H}_{1}$, the tests based on both the gradient and IPW methods are more powerful than those based on the naive methods.
In summary, the simulation results in Tables (ref) and (ref) are consistent with the results in Theorems (ref) and (ref): both the gradient bootstrap and the IPW multiplier bootstrap provide valid pointwise and uniform inference for QTEs under MPDs. The findings also show that when the information on pair identities is unavailable, the IPW multiplier bootstrap continues to provide a sound basis for inference.
Policy and macroeconomic uncertainty are considered to be two major constraints to firm growth in developing countries (bloom2014, bloom2014; bloom2018, bloom2018; worldbank2004, worldbank2004). groh2016 conducted a randomized experiment with a MPD to explore the treatment effect of providing insurance against macroeconomic and policy shocks to microenterprise owners. In this section, we apply the bootstrap methods developed in this paper to their data and examine both the ATEs and QTEs of macroinsurance on business owners' monthly consumption and their firms' monthly profits.\footnote{\doublespacing Data are available at https://microdata.worldbank.org/index.php/catalog/2063.}
\newcolumntype{B}{>{\hsize=1.3\hsize \arraybackslash}X} \newcolumntype{S}{>{\hsize=.9\hsize \arraybackslash}X}
The sample consists of 2824 business owners, who were the clients of Egypt's largest microfinance institution -- Alexandria Business Association (ABA). After an exact match on gender and microfinance branch code within ABA, the business owners were grouped into pairs by using an optimal greedy algorithm to minimize the Mahalanobis distance between the values of additional 13 matching variables (See groh2016, groh2016 for the definitions of these 13 variables). This segmentation gives 1412 pairs in the sample; one business owner in each pair was randomly assigned to the treatment group and the other to the control group. In the treatment group, a macroinsurance product was offered. Groh and McKenzie (2016) then examined the impacts of the access to macroinsurance on various outcome variables.
Here we focus on the impacts of macroinsurance on two outcome variables: the business owners' monthly consumption and their firms' monthly profits. Table (ref) gives descriptive statistics (means and standard deviations) of these two outcome variables as well as all the matching variables used by groh2016 to form the pairs in their experiments.\footnote{\doublespacing We filter out 137 unbalanced observations (less than 5% of the total observations in groh2016, groh2016) to keep a balanced data in both the pairs and the pairs of the pairs. The summary statistics in Table (ref) are almost exactly the same as those in Table 1 of groh2016.}
Table (ref) reports the ATE estimators of macroinsurance on the consumption and profits with the standard errors (in parentheses) calculated by four methods. Specifically, the columns `Naive' and `Adj' correspond to the two-sample $t$-test and the adjusted $t$-test in BRS19, respectively; the column `Naive pair' corresponds to the $t$-test using standard errors estimated by the naive multiplier bootstrap of the pairs; the column `IPW' corresponds to the $t$-test using standard errors estimated by IPW multiplier bootstrap.\footnote{\doublespacing Throughout this section, to nonparametrically estimate the propensity score in the IPW multiplier bootstrap, we first standardize all the continuous matching variables to have mean zero and variance one. There are only three continuous matching variables; the rest of the matching variables are all dummy variables. We then conduct sieve estimation by choosing the bases $\{1,\max(X_{1}-qx_{1,0.3},0),\max(X_{1}-qx_{1,0.5},0), \max(X_{2}-qx_{2,0.3},0),\max(X_{2}-qx_{2,0.5},0), \max(X_{3}-qx_{3,0.3},0),\max(X_{3}-qx_{3,0.5},0), X_{1}X_{2}, X_{2}X_{3}, X_{1}X_{3}, DV \}$, where $(X_1,X_2,X_3)$ denote three standardized continuous matching variables, $qx_{j,0.3}$ and $qx_{j,0.5}$ are $0.3$ and $0.5$th quantiles of $X_j$ for $j = 1,2,3$, and DV denotes the dummy matching variables except for the variable “missing Feb 2012 profits" as it is collinear with the variable “missing Jan 2012 profits." Results reported in this section are similar to those when the sieve basis functions are selected via cross-validation. For more details on the cross-validation method and the related empirical results, see Section H in the supplement.} The results lead to the following observations. First, consistent with the findings in groh2016, the naive two-sample $t$-tests show that expanding access to macroinsurance has no significant average effects on monthly consumption and profits. Second, the standard errors in the adjusted $t$-test are lower than those in the naive $t$-test, which is consistent with the finding in BRS19. Compared to the standard errors estimated by the naive multiplier bootstrap of the pairs, they are modestly lower for the ATE estimates of macroinsurance on the consumption and slightly larger for those on the profits. More importantly, the standard errors estimated by the IPW multiplier bootstrap are lower than those estimated by two naive methods. These results corroborate our earlier finding that the IPW multiplier bootstrap is an alternative to the approach adopted in BRS19, especially when the information on pair identities is unavailable. In addition, the results for the firms' profits also highlight the importance of accounting for the dependence structure within the pairs when estimating the standard errors of the ATE estimates. The ATE of macroinsurance on profits is statistically insignificant based on the naive $t$-test, but is significant at $10\%$ significance level if adjusted or IPW $t$-test is used instead.
\newcolumntype{B}{>{\hsize=1.44\hsize \arraybackslash}X} \newcolumntype{S}{>{\hsize=.89\hsize \arraybackslash}X}
Table (ref) presents the QTE estimates at quantile indexes 0.25, 0.5, and 0.75 with the standard errors (in parentheses) estimated by four different methods. Specifically, the columns “Naive,"“Naive pair," “Gradient," and “IPW" correspond to the results of the naive multiplier bootstrap, the naive multiplier bootstrap of the pairs, the gradient bootstrap,\footnote{\doublespacing Using the original pair identities and all the matching variables in groh2016, we can re-order the pairs according to the procedure described in Section (ref). We thank David McKenzie for providing us the Stata code to implement the matching algorithm.} and the IPW multiplier bootstrap, respectively. These results lead to the following three observations.
First, consistent with the theoretical results in Section (ref), all the standard errors estimated by the gradient bootstrap or the IPW multiplier bootstrap are lower than those estimated by the naive multiplier bootstrap. For example in Panel A, at the 75th percentile, compared with the naive multiplier bootstrap, the gradient and IPW multiplier bootstraps reduce the standard error by 9.3% and 8.2%, respectively.
Second, the standard errors estimated by the gradient or IPW multiplier bootstrap are mostly lower than those estimated by the naive multiplier bootstrap of the pairs as well. The magnitude of reduction is modest, which may be because the two terms, $\frac{m_{1,\tau}(X_i)}{f_1(q_1(\tau))}$ and $\frac{m_{0,\tau}(X_i)}{f_0(q_0(\tau))}$ in (ref), are almost equal in the current dataset, and thus, their effects on the standard error estimates are canceled.
Third, there is considerable heterogeneity in the effects of macroinsurance on the firm profits. Specifically, the magnitude of the treatment effects of macroinsurance rises as the quantile indexes increase. For example, in Panel B, the treatment effects on the monthly profits at the 25th percentile and the median are negative but not statistically significantly different from zero, whereas the effects at the 75th percentile are statistically significant. The magnitude of the treatment effect increases by over 100% from the 25th percentile to the median and by about 200% from the median to the 75th percentile. These findings may imply that expanding access to macroinsurance has small but negative effects on the firm profits in the lower tail of the distribution, and that these negative effects become stronger for upper-ranked microenterprises.
The third observation in Table (ref) indicates that the heterogeneous effects of macroinsurance on the firms' profits are economically substantial. To assess whether these are statistically significant too, Table (ref) reports statistical tests for the heterogeneity of the QTEs. Specifically, we test the null hypotheses that $q(0.50)- q(0.25) = 0$, $q(0.75)- q(0.50) = 0$, and $q(0.75)- q(0.25) = 0$. We find that only the difference between the 75th and 25th QTEs in Panel B is statistically significant at the 10% significance level. This finding implies that the statistical evidence of heterogeneous treatment effects of macroinsurance on the firm profits is strong only in comparisons of the microenterprises between the lower and upper tails of the distribution.
This paper has studied estimation and inference of QTEs under MPDs and developed new bootstrap methods to improve statistical performance. Derivation of the limit distribution of QTE estimators under MPDs reveals that analytic methods of inference based on asymptotic theory requires estimation of two infinite-dimensional nuisance parameters for every quantile index of interest. A further limitation is that both the naive multiplier bootstrap and the naive multiplier bootstrap of the pairs fail to approximate the limit distribution of the QTE estimator as they do not preserve the dependence structure in the original sample. Instead, we propose a gradient bootstrap approach that can consistently approximate the limit distribution of the original estimator and is free of tuning parameters. Implementation of the gradient bootstrap requires knowledge of pair identities. So when such information is unavailable we propose an IPW multiplier bootstrap and show that it consistently approximates the limit distribution of the original QTE estimator. Simulations provide finite sample evidence of these procedures that support the asymptotic findings. An empirical application of these bootstrap methods to the real dataset in groh2016 shows considerable evidence of heterogeneity in the effects of macroinsurance on firm profits. In both the simulations and the empirical application, the two recommended bootstrap methods of inference perform well in the sense that they usually provide smaller standard errors and greater inferential accuracy than those obtained by naive bootstrap methods.
Two directions for future research are especially evident from the present findings. First, it would be interesting to study inference of the QTEs when data are independent but not identically distributed. Such an assumption is adopted in the linear quantile regression literature to address issues regarding data heteroskedasticity and clustering (see, for example, chen2015, chen2015 and h17, h17). Second, it would be useful to incorporate data-driven methods such as cross-validation to the selection of the sieve basis functions when implementing the IPW multiplier bootstrap and then develop a procedure for inference orthogonal to the model selection bias introduced by data-driven methods.