Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
103,658 characters · 18 sections · 65 citation commands
Regression-Adjusted Estimation of Quantile Treatment Effects under Covariate-Adaptive Randomizations
{-1.5cm}
\affil[a]{ Fanhai International School of Finance, Fudan University. 220 Handan Rd, Shanghai, China 200437.} \affil[b]{ School of Economics, Singapore Management University. 90 Stamford Rd, Singapore 178903.} \affil[c]{ Yale University, New Haven, CT 06520-8281, United States of America.} \affil[d]{ University of Auckland, 12 Grafton Rd, Auckland Central, Auckland, 1010, New Zealand. } \affil[e]{ University of Southampton, University Rd, Southampton SO17 1BJ, United Kingdom. } \affil[f]{ Department of Economics, University of Macau. Avenida da Universidade, Taipa, Macao SAR, China.}
Covariate-adaptive randomizations (CARs) have recently seen growing use in a wide variety of randomized experiments in economic research. Examples include CCFNT16, greaney2016, jakiela2016, burchardi2019, anderson2021, among many others. In CAR modeling, units are first stratified using some baseline covariates, and then, within each stratum, the treatment status is assigned (independent of covariates) to achieve the balance between the numbers of treated and control units.
In many empirical studies, apart from the average treatment effect (ATE), researchers are often interested in using randomized experiments to estimate quantile treatment effects (QTEs). The QTE has a useful role as a robustness check for the ATE and characterizes any heterogeneity that may be present in the sign and magnitude of the treatment effects according to their position within the distribution of outcomes. See, for example, B06, MS11, DGPR13, BDGK15, CDDP15, and CF17.
Two practical issues arise in estimation and inference concerning QTEs under CARs. First, other covariates in addition to the strata indicators are collected during the experiment. It is possible to incorporate these covariates in the estimation of treatment effects to reduce variance and improve efficiency. In the estimation of ATE, the usual practice is to run a simple ordinary least squares (OLS) regression of the outcome on treatment status, strata indicators, and additional covariates as in the analysis of covariance (ANCOVA). F08b, F081 pointed out that such an OLS regression adjustment can degrade the precision of the ATE estimator. L13 reexamined Freedman's critique and showed that, in order to improve efficiency, the linear regression adjustment should include a full set of interactions between the treatment status and covariates. However, because the quantile function is a nonlinear operator, even when the assignment of treatment status is completely random, a similar linear quantile regression with a full set of interaction terms is unable to provide a consistent estimate of the unconditional QTE, not to mention the improvement of estimation efficiency. Second, in order to achieve balance in the respective number of treated and control units within each stratum, treatment statuses under CARs usually exhibit a (negative) cross-sectional dependence. Standard inference procedures that rely on cross-sectional independence are therefore conservative and lack power. These two issues raise questions of how to use the additional covariates to consistently and more efficiently estimate QTE in CAR settings and how to conduct valid statistical procedures that mitigate conservatism in inference.\footnote{For example, BCS17 and ZZ20 have shown that the usual two-sample $t$-test for inference concerning ATE and multiplier bootstrap inference concerning QTE are in general conservative under CARs.}
The present paper addresses these issues by proposing a regression-adjusted estimator of the QTE, deriving its limit theory, and establishing the validity of multiplier bootstrap inference under CARs. Even under potential misspecification of the auxiliary regressions, the proposed QTE estimator is shown to maintain its consistency, and the multiplier bootstrap procedure is shown to have an asymptotic size equal to the nominal level under the null. When the auxiliary regression is correctly specified, the QTE estimator achieves minimum asymptotic variance.
We further investigate efficiency gains that materialize from the regression adjustments in three scenarios: (1) parametric regressions, (2) nonparametric regressions, and (3) regressions with regularization in high-dimensional settings. Specifically, for parametric regressions with a potentially misspecified linear probability model, we propose to compute the optimal linear coefficient by minimizing the variance of the QTE estimator. Such an adjustment is optimal within the class of linear adjustments but does not necessarily achieve the global minimum asymptotic variance. However, as no adjustment is a special case of the linear regression with all the coefficients being zero, our optimal linear adjustment is guaranteed to be weakly more efficient than the QTE estimator with no adjustments, which addresses Freedman's critique. We also consider a potentially misspecified logistic regression with fixed-dimensional regressors and strata- and quantile-specific regression coefficients, which is then estimated by the quasi maximum likelihood estimation (QMLE). Although the QMLE does not necessarily minimize the asymptotic variance of the QTE, such a flexible logistic model can closely approximate the true specification. Therefore, in practice, the corresponding regression-adjusted QTE estimator usually has a smaller variance than that with no adjustments. Last, we propose to treat the logistic QMLE adjustments as new linear regressors and re-construct the corresponding optimal linear adjustments. We then show the QTE estimator with the new adjustments are weakly more efficient than that with both the original logistic QMLE adjustments and no adjustments.
In nonparametric regressions, we further justify the QMLE by letting the regressors in the logistic regression be a set of sieve basis functions with increasing dimension and show how such a nonparametric regression-adjusted QTE estimator can achieve the global minimum asymptotic variance. For high-dimensional regressions with regularization, we consider logistic regression under $\ell_1$ penalization, an approach that also achieves the global minimum asymptotic variance. All the limit theories hold uniformly over a compact set of quantile indices, implying that our multiplier bootstrap procedure can be used to conduct inference on QTEs involving single, multiple, or a continuum of quantile indices.
These results, including the limit distribution of the regression-adjusted QTE estimator and the validity of the multiplier bootstrap, provide novel contributions to the literature in three respects. First, the data generated under CARs are different from observational data as the observed outcomes and treatment statuses are cross-sectionally dependent due to the randomization schemes. Recently BCS17 established a rigorous asymptotic framework to study the ATE estimator under CARs and pointed out the conservatism of the two-sample t-test except for some special cases. (See BCS17 for more detail.) Our analysis follows this new framework, which departs from the literature of causal inference under an i.i.d. treatment structure.
Second, we contribute to the literature on causal inference under CARs by developing a new methodology that includes additional covariates in the estimation of the unconditional QTE and by establishing a general theory for regression adjustments that allow for parametric, nonparametric, and regularized estimations of the auxiliary regressions. As mentioned earlier, unlike ATE estimation, the naive linear quantile regression with additional covariates cannot even produce a consistent estimator of the QTE. Instead, we propose a new way to incorporate additional covariates based on the Neyman orthogonal moment and investigate the asymptotic properties and the efficiency gains of the proposed regression-adjusted estimator under CARs. This new machinery allows us to study the QTE regression, which is nonparametrically specified, with both linear (linear probability model) and nonlinear (logit and probit models) regression adjustments. To clarify this contribution to the literature we note that HH12,MHZ15,MQLH18,O21,SY13,ZZ20,Y18,YS20 considered inference of various causal parameters under CARs but without taking into account additional covariates. BCS17, BCS18, and BG21 considered saturated regressions for ATE and local ATE, which can be viewed as regression-adjustments where strata indicators are interacted with the treatment or instrument. SYZ10 showed that if a test statistic is constructed based on the correctly specified model between outcome and additional covariates and the covariates used for CAR are functions of additional covariates, then the test statistic is valid conditional on additional covariates. BL16,F18,L13,L16,LD20,LiD20,LTM20,LY20,NW20,YYS20,ZD20 studied various estimation methods based on regression adjustments, but these studies all focused on ATE estimation. Specifically, LTM20 considered linear adjustments for ATE under CARs in which the covariates can be high-dimensional and the adjustments can be estimated by Lasso. AHL18 considered regression adjustment using additional covariates for ATE and Local ATE. We differ from them by considering QTE with nonlinear adjustments such as logistic Lasso.
Third, we establish the validity of the multiplier bootstrap inference for the regression-adjusted QTE estimator under CARs. To the best of our knowledge, SYZ10 and ZZ20 are the only works in the literature that studied bootstrap inference under CARs. SYZ10 considered the covariate-adaptive bootstrap for the linear regression model. ZZ20 proposed to bootstrap inverse propensity score weighted (IPW) QTE estimator with the estimated target fraction of treatment even when the truth is known. They showed that the asymptotic variance of the IPW estimator is the same under various CARs. Thus, even though the bootstrap sample ignores the cross-sectional dependence and behaves as if the randomization scheme is simple, the asymptotic variance of the bootstrap analogue is still the same. We complement this research by studying the validity of multiplier bootstrap inference for our regression-adjusted QTE estimator. We establish analytically that the multiplier bootstrap with the estimated fraction of treatment is not conservative in the sense that it can achieve an asymptotic size equal to the nominal level under the null even when the auxiliary regressions are misspecified.
The present paper also comes under the umbrella of a growing literature that has addressed estimation and inference in randomized experiments. In this connection, we mention the studies of HHK11, athey2017, abadie2018, T18, BRS19, B19, JL20 among many others. B19 showed an `optimal' matched-pair design can minimize the mean-squared error of the difference-in-means estimator for ATE, conditional on covariates. T18 designed an adaptive randomization procedure which can minimize the variance of the weighted estimator for ATE. Both works rely on a pilot experiment to design the optimal randomization. In contrast, we take the randomization scheme (i.e., CARs) as given and search for new estimators (other than difference-in-quantile and weighted estimators) for QTE that have smaller variance. In addition, our approach does not require a pilot experiment. Therefore, our and their methods are applied to different scenarios depending on the definition of `optimality' and the data available, and thus, complement to each other.
From a practical perspective, our estimation and inferential methods have four advantages. First, they allow for common choices of auxiliary regressions such as linear probability, logit, and probit regressions, even though these regressions may be misspecified. Second, the methods can be implemented without tuning parameters. Third, our (bootstrap) estimator can be directly computed via the subgradient condition, and the auxiliary regressions need not be re-estimated in the bootstrap procedure, both of which save considerable computation time. Last, our estimation and inference methods can be implemented without the knowledge of the exact treatment assignment rule used in the experiment. This advantage is especially useful in subsample analysis, where sub-groups are defined using variables other than those to form the strata and the treatment assignment rule for each sub-group becomes unknown. See, for example, the anemic subsample analysis in CCFNT16 and ZZ20. These last three points carry over from ZZ20 and are logically independent of the regression adjustments. One of our contributions is to show these results still hold for our regression-adjusted estimator.
The remainder of the paper is organized as follows. Section (ref) describes the model setup and notation. Section (ref) develops the asymptotic properties of our regression-adjusted QTE estimator. Section (ref) studies the validity of the multiplier bootstrap inference. Section (ref) considers parametric, nonparametric, and regularized estimation of the auxiliary regressions. Section (ref) reports simulation results, and an empirical application of our methods to the impact of expanding access to basic bank accounts on savings is provided in Section (ref). Section (ref) concludes. Proofs of all results and some additional simulation results are given in the Online Supplement.
Potential outcomes for treated and control groups are denoted by $Y(1)$ and $Y(0)$, respectively. Treatment status is denoted by $A$, with $A=1$ indicating treated and $A=0$ untreated. The stratum indicator is denoted by $S$, based on which the researcher implements the covariate-adaptive randomization. The support of $S$ is denoted by $\mathcal{S}$, a finite set. After randomization, the researcher can observe the data $\{Y_i,S_i,A_i,X_i\}_{i \in [n]}$ where $[n]=\{1,2,...n\}$, $Y_i = Y_i(1)A_i + Y_i(0)(1-A_i)$ is the observed outcome, and $X_i$ contains extra covariates besides $S_i$ in the dataset. The support of $X$ is denoted $\text{Supp}(X)$. In this paper, we allow $X_i$ and $S_i$ to be dependent. For $i \in [n]$, let $p(s) = \mathbb{P}(S_i = s)$, $n(s) = \sum_{i \in [n]}1\{S_i = s\}$, $n_1(s) = \sum_{i \in [n]}A_i1\{S_i=s\}$, and $n_0(s) = n(s) - n_1(s)$. We make the following assumptions on the data generating process (DGP) and the treatment assignment rule.
Several remarks are in order. First, Assumption (ref)(i) allows for cross-sectional dependence among treatment statuses ($\{A_i\}_{i \in [n]}$), thereby accommodating many covariate-adaptive randomization schemes as discussed below. Second, although treatment statuses are cross-sectionally dependent, they are independent of the potential outcomes and additional covariates conditional on the stratum indicator $S$. Therefore, data are still experimental rather than observational. Third, Assumption (ref)(iii) requires the size of each stratum to be proportional to the sample size. Fourth, we can view $\pi(s)$ as the target fraction of treated units in stratum $s$. Similar to BCS18, we allow the target fractions to differ across strata. Just as for the overlapping support condition in an observational study, the target fractions are assumed to be bounded away from zero and one. In randomized experiments, this condition usually holds because investigators can determine $\pi(s)$ in the design stage. In fact, in most CARs, $\pi(s) = 0.5$ for $s \in \mathcal{S}$. Fifth, $D_n(s)$ represents the degree of imbalance between the real and target factions of treated units in the $s$th stratum. BCS17 show that Assumption (ref)(iv) holds under several covariate-adaptive treatment assignment rules such as simple random sampling (SRS), biased-coin design (BCD), adaptive biased-coin design (WEI), and stratified block randomization (SBR). For completeness, we briefly repeat their descriptions below. Note we only require $D_n(s)/n(s) = o_p(1)$, which is weaker than the assumption imposed by BCS17 but the same as that imposed by BCS18 and ZZ20.
Denote the $\tau$th quantile of $Y(a)$ by $q_a(\tau)$ for $a=0,1$. We are interested in estimating and inferring the $\tau$th quantile treatment effect defined as $q(\tau) = q_1(\tau) - q_0(\tau)$. The testing problems of interest involve single, multiple, or even a continuum of quantile indices, as in the following null hypotheses
for some pre-specified value $\underline{q}$ or function $\underline{q}(\tau)$, where $\Upsilon$ is some compact subset of $(0,1)$. We can also test constant QTE by letting $\underline{q}(\tau)$ in the last hypothesis be a constant $\underline{q}$.
Define $m_a(\tau,s,x) = \tau - \mathbb{P}(Y_i(a)\leq q_a(\tau)|S_i=s,X_i=x)$ for $a=0,1$ which are the true specifications but unknown to researchers. Instead, researchers specify working models $\{\overline{m}_a(\tau,s,x)\}_{a=0,1}$\footnote{We view $\overline{m}_a(\cdot)$ as some function with inputs $\tau,s,x$. For example, researchers can specify a linear probability model with $\overline{m}_a(\tau,s,x) = \tau - x^\top \beta_{a,s}$, where $\beta_{a,s}$ is the linear coefficient that varies across treatment status $a$ and stratum $s$.} for the true specification, which can be misspecified. Last, the researchers estimate the (potentially misspecified) working models via some forms of regression, and the estimators are denoted as $\{\widehat{m}_a(\tau,s,x)\}_{a=0,1}$. We also refer to $\overline{m}_a(\cdot)$ as the auxiliary regression.
Our regression-adjusted estimator of $q_1(\tau)$, denoted as $\hat{q}_1^{adj}(\tau)$, can be defined as
where $\rho_\tau(u) = u(\tau-1\{u \leq 0\})$ is the usual check function and $\hat{\pi}(s) = n_1(s)/n(s)$. We emphasize that $\widehat{m}_1(\cdot)$ may not consistently estimate the true specification $m_1(\cdot)$. Similarly, we can define
Then, our regression adjusted QTE estimator is
Several remarks are in order. First, in observational studies with i.i.d. data and $A_i \perp\!\!\!\perp X_i|S_i$, F07, BCFH13, and KMU19 showed that the doubly robust moment for $q_1(\tau)$ is
where $\overline{\pi}(s)$ and $\overline{m}_1(\tau,s,x)$ are the working models for the target fraction ($\pi(s)$) and conditional probability ($m_1(\tau,s,x)$), respectively. Our estimator is motivated by this doubly robust moment, but our analysis differs from that for the observational data as CARs introduces cross-sectional dependence among observations. Second, as our target fraction estimator $\hat{ \pi}(s) = n_1(s)/n(s)$ is consistent, it means $\overline{\pi}(s)$ is correctly specified as $\pi(s)$. Then, due to the double robustness, our regression adjusted estimator is consistent even when $\overline{m}_a(\cdot)$ is misspecified and $\widehat{m}_a(\cdot)$ is an inconsistent estimator of $m_a(\cdot)$. Third, we use the estimated target fraction $\hat{ \pi}(s)$ even when $\pi(s)$ is known because this guarantees that the bootstrap inference is not conservative. Further discussion is provided after Theorem (ref).
Several remarks are in order. First, Assumption (ref) is standard in the quantile regression literature. We do not need $f_a(y|x,s)$ to be bounded away from zero because we are interested in the unconditional quantile $q_a(\tau)$, which is uniquely defined as long as the unconditional density $f_a(q_a(\tau))$ is positive. Second, Assumption (ref)(i) is high-level. If we consider a linear probability model such that $\overline{m}_a(\tau,s,X_i) = \tau-X_i^\top \theta_{a,s}(\tau)$ and $\widehat{m}_{a}(\tau,s,X_i)= \tau-X_i^\top \hat{ \theta}_{a,s}(\tau)$, then Assumption (ref)(i) is equivalent to
which is similar to LTM20 and holds intuitively if $\hat{ \theta}_{a,s}(\tau)$ is a consistent estimator of the pseudo true value $ \theta_{a,s}(\tau)$. Third, Assumptions (ref)(ii) and (ref)(iii) impose mild regularity conditions on $\overline{m}_a(\cdot)$. Assumption (ref)(ii) holds automatically if $\Upsilon$ is a finite set. In general, both Assumption (ref)(ii) and (ref)(iii) hold if
for some constant $L>0$. Such Lipschitz continuity holds for the true specification ($\overline{m}_a(\cdot) = m_a(\cdot)$) under Assumption (ref). Fourth, we provide primitive sufficient conditions for Assumption (ref) in Section (ref).
Three remarks are in order. First, the expression for the asymptotic variance of $\hat{q}^{adj}(\tau)$ can be found in the proof of Theorem (ref). It is the same whether the randomization scheme achieves strong balance\footnote{We refer readers to BCS17 for the definition of strong balance.} or not. This robustness is due to the use of the estimated target fraction ($\hat{\pi}(s)$). The same phenomenon was discovered in the simplified setting by ZZ20. Second, although our estimator is still consistent and asymptotically normal when the auxiliary regression is misspecified, it is meaningful to pursue the correct specification as it achieves the minimum variance. As the estimator with no adjustments can be viewed as a special case of our estimator with $\overline{m}_a(\cdot) = 0$, Theorem (ref) implies that the adjusted estimator with the correctly specified auxiliary regression is more efficient than that with no adjustments. If the auxiliary regression is misspecified, the adjusted estimator can sometimes be less efficient than the unadjusted one, which is known as the Freedman's critique. In Section (ref), we discuss how to make adjustments that do not harm the precision of the QTE estimator. Third, the asymptotic variance of $\hat{q}^{adj}(\tau)$ depends on $(f_a(q_a(\tau)),m_a(\tau,s,x))_{a = 0,1}$, which are infinite-dimensional nuisance parameters. To conduct analytic inference, it is necessary to nonparametrically estimate these nuisance parameters, which requires tuning parameters. Nonparametric estimation can be sensitive to the choice of tuning parameters and rule-of-thumb tuning parameter selection may not be appropriate for every DGP or every quantile. The use of cross-validation in selecting the tuning parameters is possible in principle but, in practice, time-consuming. These practical difficulties of analytic methods of inference provide strong motivation to investigate bootstrap inference procedures that are much less reliant on tuning parameters.
We approximate the asymptotic distributions of $\hat{q}^{adj}(\tau)$ via the multiplier bootstrap. Let $\{\xi_i\}_{i \in [n]}$ be a sequence of bootstrap weights which will be specified later. Define $n_1^w(s) = \sum_{i \in [n]}\xi_iA_i1\{S_i=s\}$, $n_0^w(s) = \sum_{i \in [n]}\xi_i(1-A_i)1\{S_i=s\}$, $n^w(s) = \sum_{i \in [n]} \xi_i 1\{S_i=s\} = n_1^w(s) + n_0^w(s)$, and $\hat{\pi}^w(s) = n_1^w(s)/n^w(s)$. The multiplier bootstrap counterpart of $\hat{q}^{adj}(\tau)$ is denoted by $\hat{q}^{w}(\tau)$ and defined as
where
and
Two comments on implementation are noted here: (i) we do not re-estimate $\widehat{m}_a(\cdot)$ in the bootstrap sample, which is similar to the multiplier bootstrap procedure proposed by BCFH13; and (ii) in Section (ref) of the Online Supplement we propose a way to directly compute $(\hat{q}_a^{w}(\tau))_{a=0,1}$ from the subgradient conditions of (ref) and (ref), thereby avoiding the optimization. Both features considerably reduce computation time of our bootstrap procedure.
Next, we specify the bootstrap weights.
We require the bootstrap weights to be nonnegative so that the objective functions in (ref) and (ref) are convex. In practice, we generate $\xi_i$ independently from the standard exponential distribution. Assumption (ref) is the bootstrap counterpart of Assumption (ref). Continuing with the linear model example considered after Assumption (ref), Assumption (ref) requires
which holds if $\hat{ \theta}_{a,s}(\tau)$ is a uniformly consistent estimator of $\theta_{a,s}(\tau)$.
Two remarks are in order. First, Theorem (ref) shows the limit distribution of the bootstrap estimator conditional on data can approximate that of the original estimator uniformly over $\tau \in \Upsilon$. This is the theoretical foundation for the bootstrap confidence intervals and bands described in Section (ref) in the Online Supplement. Specifically, denote $\{\hat{q}^{w,b}(\tau)\}_{b \in [B]}$ as the bootstrap estimates where $B$ is the number of bootstrap replications. Let $\widehat{\mathcal{C}}(\nu)$ and $\mathcal{C}(\nu)$ be the $\nu$th empirical quantile of the sequence $\{\hat{q}^{w,b}(\tau)\}_{b \in [B]}$ and the $\nu$th standard normal critical value, respectively. Then, we suggest using the bootstrap estimator to construct the standard error of $\hat{q}^{adj}(\tau)$ as $\hat{\sigma} = \frac{\widehat{\mathcal{C}}(0.975)- \widehat{\mathcal{C}}(0.025)}{\mathcal{C}(0.975) - \mathcal{C}(0.025)}$. Note that, unlike HL21, our bootstrap standard error is not conservative. In our context, the bootstrap estimator of $\sigma^2$ considered by HL21 is $\mathbb{E}^*(\sqrt{n}(\hat{q}^w(\tau) - \hat{q}^{adj}(\tau))^2)$, where $\mathbb{E}^*$ is the conditional expectation given data. It is well-known that weak convergence does not imply convergence in $L_2$-norm, which explains why they can show their estimator is in general conservative. Instead, we use a different estimator of the standard error and can show it is consistent given weak convergence. Second, such a bootstrap approximation is consistent under CAR. ZZ20 showed that for the QTE estimation without regression adjustment, bootstrapping the IPW QTE estimator with the estimated target fraction results in non-conservative inference, while bootstrapping the IPW estimator with the true fraction is conservative under CARs. As the estimator considered by ZZ20 is a special case of our regression-adjusted estimator with $\widehat{m}_a(\cdot) =0$, we conjecture that the same conclusion holds. A proof of conservative bootstrap inference with the true target fraction is not included in the paper due to the space limit.\footnote{Full statements and proofs are lengthy because we need to derive the limit distributions of not only the bootstrap but also the original estimator with the true target fraction. Although the negative result is theoretically interesting, we are not aware of any empirical papers using the true target fraction while making regression adjustments. Moreover, our method is shown to have better performance than the one with the true target fraction in simulations. So the practical value of proving the negative result is limited.} Our simulations confirm both the correct size coverage of our inference method using the bootstrap with the estimated target fraction and the conservatism of the bootstrap with the true target fraction. The standard error of the QTE estimator is found to be 34.9% larger on average by using the true rather than the estimated target fraction in the simulations (see Tables (ref) below and (ref) in the Online Supplement).
In this section, we consider two approaches to estimation for the auxiliary regressions: (1) a parametric method and (2) a nonparametric method. In Section (ref) of the Online Supplement, we further consider a regularization method. For the parametric method, we do not require the model to be correctly specified. We propose ways to estimate the pseudo true value of the auxiliary regression. For the other two methods, we (nonparametrically) estimate the true model so that the asymptotic variance of $\hat{q}^{adj}(\tau)$ achieves its minimum based on Theorem (ref). For all three methods, we verify Assumptions (ref) and (ref).
In this section, we consider the case where $X_i$ is finite-dimensional. Recall $m_a(\tau,s,x) \equiv \tau - \mathbb{P}\left(Y_i(a) \leq q_a(\tau)|X_i=x,S_i=s\right)$ for $a=0,1$. We propose to model $\mathbb{P}\left(Y_i(a) \leq q_a(\tau)|X_i,S_i=s\right)$ as $\Lambda_{\tau,s}(X_i,\theta_{a,s}(\tau))$, where $\theta_{a,s}(\tau)$ is a finite dimensional parameter that depends on $(a,s,\tau)$ so that our model for $m_a(\tau,s,X_i)$ is
We note that, as we allow for misspecification, the researchers have the freedom to choose any functional forms for $\Lambda_{\tau,s}(\cdot)$ and any pseudo true values for $\theta_{a,s}(\tau)$, both of which can vary with respect to $(\tau,s)$. For example, if we assume a logistic regression with $\Lambda_{\tau,s}(X_i,\theta_{a,s}(\tau)) = \lambda(X_i^\top\theta_{a,s}(\tau))$, where $\lambda(\cdot)$ is the logistic CDF, then there are various choices of $\theta_{a,s}(\tau)$ such as the maximizer of the population pseudo likelihood, the maximizer of the population version of the least squares objective function, or the minimizer of the asymptotic variance of the adjusted QTE estimator. As the logistic model is potentially misspecified, these three pseudo true values are not necessarily the same and can lead to different adjustments, and thus, different asymptotic variances of the corresponding adjusted QTE estimators.
Next, we first state a general result for generic choices of $\Lambda_{\tau,s}(\cdot)$ and $\theta_{a,s}(\tau)$. Suppose we estimate $\theta_{a,s}(\tau)$ by $\hat{\theta}_{a,s}(\tau)$. Then, the corresponding $\widehat{m}_a(\tau,s,X_i)$ can be written as
Three remarks are in order. First, common choices for auxiliary regressions are linear probability, logistic, and probit regressions, corresponding to $\Lambda_{\tau,s}(X_i,\theta_{a,s}(\tau)) = X_i^\top\theta_{a,s}(\tau)$, $\lambda(\vec{X}_i^\top\theta_{a,s}(\tau))$, and $\Phi(\vec{X}_i^\top\theta_{a,s}(\tau))$, respectively, where $\Phi(\cdot)$ is the standard normal CDF and $\vec{X}_i = (1,X_i^\top)^\top$. For these models, the functional form $\Lambda_{\tau,s}(\cdot)$ does not depend on $(\tau,s)$, and Assumption (ref)(i) holds automatically. For the linear regression case, we do not include the intercept because our regression adjusted estimators ((ref) and (ref)) and their bootstrap counterparts ((ref) and (ref)) are numerically invariant to location shift of the auxiliary regressions. Second, it is also important to allow the functional form $\Lambda_{\tau,s}(\cdot)$ to vary across $\tau$ to incorporate the case in which the regressor $X_i$ in the linear, logistic, and probit regressions is replaced by $W_{i,s}(\tau)$, a function of $X_i$ that depends on $(\tau,s)$. We give a concrete example for this situation in Section (ref). Third, Assumption (ref)(ii) also holds automatically if $\Upsilon$ is finite. When $\Upsilon$ is infinite, this condition is still mild.
Theorem (ref) shows that, as long as the estimator of the pseudo true value ($\hat{\theta}_{a,s}(\tau)$) is uniformly consistent, under mild regularity conditions, all the general estimation and bootstrap inference results established in Sections (ref) and (ref) hold.
In this section, we consider linear adjustment with parameter $t_{a,s}(\tau)$ such that
where the regressor $W_{i,s}(\tau)$ is a function of $X_i$ but the functional form may vary across $s,\tau$. For example, we can consider $W_{i,s}(\tau) = X_i$, the transformations of $X_i$ such as quadratic and interaction terms, and some prediction of $(1\{Y_i(1) \leq q_1(\tau)\},1\{Y_i(0) \leq q_0(\tau)\})$ given $X_i$ and $S_i = s$. The last example is further explained in Section (ref).
We note that the asymptotic variance (denoted as $\sigma^2$) of the $\hat{q}^{adj}(\tau)$ is a function of the working model ($\overline{m}_{a}(\tau,s,\cdot)$), which is further indexed by its parameters (denoted as $\{t_{a,s}(\tau)\}_{a = 0,1, s \in \mathcal{S}}$), i.e., $\sigma^2 = \sigma^2(\{\overline{m}_{a}(\tau,s,\cdot;t_{a,s})\}_{a = 0 ,1, s \in \mathcal{S}})$. Our optimal linear adjustment corresponds to parameter value $\theta_{a,s}(\tau)$ such that it minimizes $\sigma^2(\{\overline{m}_{a}(\tau,s,\cdot;t_{a,s})\}_{a = 0 ,1, s \in \mathcal{S}})$, i.e.,
The next theorem derives the closed-form expression for the optimal linear coefficient.
Four remarks are in order. First, the optimal linear coefficients $\{\theta_{a,s}(\tau)\}_{a = 0 ,1, s \in \mathcal{S}}$ are not uniquely defined. In order to achieve the minimal variance, we only need to consistently estimate one of the minimizers. We choose
as this choice avoids estimation of the densities $f_1(q_1(\tau))$ and $f_0(q_0(\tau))$. In Theorem (ref) below, we propose estimators of $\theta_{1,s}^\textit{LP}(\tau)$ and $\theta_{0,s}^\textit{LP}(\tau)$ and show they are consistent uniformly over $s$ and $\tau$. Second, note that no adjustment is nested by our linear adjustment with zero coefficients. Due to the optimality result established in Theorem (ref), our regression-adjusted QTE estimator with (consistent estimators of) $\{\theta_{a,s}^\textit{LP}(\tau)\}_{a = 0 ,1, s \in \mathcal{S}}$ is more efficient than that with no adjustments. Third, we also need to clarify that the optimality of $\{\theta_{a,s}^\textit{LP}(\tau)\}_{a = 0 ,1, s \in \mathcal{S}}$ is only within the class of linear regressions. It is possible that the QTE estimator with some nonlinear adjustments are more efficient than that with the optimal linear adjustments, especially because the linear probability model is likely misspecified. Fourth, the optimal linear coefficients $\{\theta_{a,s}^\textit{LP}(\tau)\}_{a = 0 ,1, s \in \mathcal{S},\tau \in \Upsilon}$ minimize (over the class of linear models) not only the asymptotic variance of $\hat{q}^{par}(\tau)$ but also the covariance matrix of $(\hat{q}^{par}(\tau_1),\cdots,\hat{q}^{par}(\tau_K))$ for any finite-dimension quantile indices $(\tau_1,\cdots,\tau_K)$. This implies we can use the same (estimators of) optimal linear coefficients for hypothesis testing involving single, multiple, or even a continuum of quantile indices.
In the rest of this subsection, we focus on the estimation of $\{\theta_{a,s}^\textit{LP}(\tau)\}_{a = 0,1,s\in \mathcal{S}}$. Note that $\theta_{a,s}^\textit{LP}(\tau)$ is the projection coefficient of $1\{Y_i \leq q_a(\tau)\}$ on $\tilde{W}_{i,s}(\tau)$ for the sub-population with $S_i=s$ and $A_i = a$. We estimate them by sample analog. Specifically, the parameter $q_a(\tau)$ is unknown and is replaced by some $\sqrt{n}$-consistent estimator denoted by $\hat{q}_a(\tau)$.
In practice, we compute $\{\hat{q}_a(\tau)\}_{a=0,1}$ based on (ref) and (ref) by setting $\widehat{m}_a(\tau,S_i,X_i) \equiv 0$. Then, Assumption (ref) holds automatically by Theorem (ref) with $\widehat{m}_a(\tau,S_i,X_i) = \overline{m}_a(\tau,S_i,X_i)=0$. Analysis throughout this section takes into account that the estimator $\hat{q}_a(\tau)$ is used in place of $q_a(\tau)$.
Next, we define the estimator of $\theta_{a,s}^\textit{LP}(\tau)$. Recall $I_a(s)$ is defined in Assumption (ref). Let
and
We note that Assumption (ref) holds automatically if the regressor $W_{i,s}(\tau)$ does not depend on $\tau$.
We refer to the QTE estimator adjusted by this linear probability model with optimal linear coefficients $\theta_{a,s}^\textit{LP}(\tau)$ and estimators $\hat{\theta}_{a,s}^\textit{LP}(\tau)$ as the LP estimator and denote it and its bootstrap counterpart as $\hat{q}^\textit{LP}(\tau)$ and $\hat{q}^\textit{LP,w}(\tau)$, respectively. Theorem (ref) verifies Assumption (ref) for the proposed estimator of the optimal linear coefficient. Then, by Theorem (ref), Theorems (ref) and (ref) hold for $\hat{q}^\textit{LP}(\tau)$ and $\hat{q}^\textit{LP,w}(\tau)$, which implies all the estimation and inference methods established in the paper are valid for the LP estimator. Theorem (ref) further shows $\hat{q}^\textit{LP}(\tau)$ is the estimator with the optimal linear adjustment and weakly more efficient than the QTE estimator with no adjustments.
It is also common to consider the logistic regression as the adjustment and estimate the model by maximum likelihood (ML). The main goal of the working model is to approximate the true model as closely as possible. It is, therefore, useful to include additional technical regressors such as interactions in the logistic regression. The set of regressors used is defined as $H_i = H(X_i)$, which is allowed to contain the intercept. Let $\hat{\theta}_{a,s}^\textit{ML}(\tau)$ and $\theta_{a,s}^\textit{ML}(\tau)$ be the quasi-ML estimator and its corresponding pseudo true value, respectively, i.e.,
and
We then define
In addition to the inclusion of technical regressors, we allow the pseudo true value ($\theta_{a,s}^\textit{ML}(\tau)$) to vary across quantiles $\tau$, giving another layer of flexibility to the model. Such a model is called the distribution regression and was first proposed by CFM13. We emphasize here that, although we aim to make the regression model as flexible as possible, our theory and results do not require the model to be correctly specified.
Four remarks are in order. First, we refer to the QTE estimator adjusted by the logistic model with QMLE as the ML estimator and denote it and its bootstrap counterpart as $\hat{q}^\textit{ML}(\tau)$ and $\hat{q}^\textit{ML,w}(\tau)$, respectively. Assumptions (ref)(i) holds automatically for the logistic regression. If we further impose Assumption (ref)(ii), then Theorem (ref) implies that all the estimation and bootstrap inference methods established in the paper are valid for the ML estimator. Second, we take into account that $\hat{\theta}_{a,s}^\textit{ML}(\tau)$ is computed when the true $q_a(\tau)$ is replaced by its estimator $\hat{q}_a(\tau)$ and derive the results in Theorem (ref) under Assumption (ref). Third, the ML estimator is not guaranteed to be optimal or be more efficient than QTE estimator with no adjustments. On the other hand, as we can include additional technical terms in the regression and allow the regression coefficients to vary across $\tau$, the logistic model can be close to the true model $m_a(\tau,s,X_i)$, which achieves the global minimum asymptotic variance based on Theorem (ref). Fourth, in Section (ref), we further justify the use of the ML estimator with a flexible logistic model by letting the number of technical terms (or equivalently, the dimension of $H_i$) diverge to infinity, showing by this means that the ML estimator can indeed consistently estimate the true model and thereby achieve the global minimum covariance matrix of the adjusted QTE estimator.
Although in simulations, we cannot find a DGP in which the QTE estimator with logistic adjustment is less efficient than that with no adjustments, theoretically such a scenario still exists. In this section, we follow the idea of CF21 and construct an estimator which is weakly more efficient than both the ML estimator and the estimator with no adjustments. We denote $W_{i,s}(\tau) = (\lambda(H_i^\top \theta_{1,s}^\textit{ML}(\tau)),\lambda(H_i^\top \theta_{0,s}^\textit{ML}(\tau)))^\top$ and treat it as the regressor in a linear adjustment, i.e., define $\overline{m}_{a}(\tau,s,X_i) = \tau - W_{i,s}^\top(\tau) t_{a,s}(\tau)$. Then, the logistic adjustment in Section (ref) and no adjustments correspond to $t_{a,s}(\tau) = a(1,0)^\top + (1-a)(0,1)^\top$ and $t_{a,s}(\tau) = (0,0)^\top$ for $a=0,1$, respectively. However, following Theorem (ref), the optimal linear coefficient with regressor $W_{i,s}(\tau)$ is
where $\tilde{W}_{i,s}(\tau) = W_{i,s}(\tau) - \mathbb{E}(W_{i,s}(\tau)|S_i=s)$. Using the adjustment term $\overline{m}_{a}(\tau,s,X_i) = \tau - W_{i,s}^\top(\tau) t_{a,s}(\tau)$ with $t_{a,s}(\tau) = \theta_{a,s}^\textit{LPML}(\tau)$ is asymptotically weakly more efficient than any other choices of $t_{a,s}(\tau)$. In practice, we do not observe $W_{i,s}(\tau)$, but can replace it by its feasible version $\hat{W}_{i,s}(\tau) = (\lambda(H_i^\top \hat{\theta}_{1,s}^\textit{ML}(\tau)),\lambda(H_i^\top \hat{\theta}_{0,s}^\textit{ML}(\tau)))^\top$. We then define
and
In practice when $n$ is small $\breve{W}_{i,a,s}(\tau)$ may be nearly multicollinear within some stratum, which can lead to size distortion in inference concerning QTE. We therefore suggest first normalizing $\breve{W}_{i,a,s}(\tau)$ by its standard deviation (denoting the normalized $\breve{W}_{i,a,s}(\tau)$ as $\ddot{W}_{i,a,s}(\tau)$) and then running a ridge regression
where $I_2$ is the two-dimensional identity matrix and $\delta_n = 1/n$. Then, the final regression adjustment is
Given Assumption (ref), such a ridge penalty is asymptotically negligible and all the results in Theorem (ref) still hold.\footnote{In unreported simulations, we find that when $n\geq 800$, the ridge regularization is unnecessary and the original adjustment (i.e., (ref)) has no size distortion, implying that near-multicollinearity is indeed just a finite-sample issue.}
This section considers nonparametric estimation of $m_a(\tau,s,X_i)$ when the dimension of $X_i$ is fixed as $d_x$. For ease of notation, we assume all coordinates of $X_i$ are continuously distributed. If in an application some elements of $X$ are discrete, the dimension $d_x$ is interpreted as the dimension of the continuous covariates. All results in this section can then be extended in a conceptually straightforward manner by using the continuous covariates only within samples that are homogeneous in discrete covariates.
As $m_a(\tau,s,X_i)$ is nonparametrically estimated, we have $ \overline{m}_a(\tau,s,X_i) = m_a(\tau,s,X_i)=\tau- \mathbb{P}(Y_i(a)\leq q_a(\tau)|S_i=s,X_i)$. We estimate $\mathbb{P}(Y_i(a)\leq q_a(\tau)|S_i=s,X_i)$ by the sieve method of fitting a logistic model, as studied by HIR03. Specifically, recall $\lambda(\cdot)$ is the logistic CDF and denote the number of sieve bases by $h_n$, which depends on the sample size $n$ and can grow to infinity as $n \rightarrow \infty$. Let $H_{h_n}(x) = (b_{1n}(x),\cdots,b_{h_nn}(x))^\top$ where $\{b_{hn}(x)\}_{h \in [h_n]}$ is an $h_n$ dimensional basis of a linear sieve space. More details on the sieve space are given in Section (ref) of the Online Supplement. Denote
We refer to the QTE estimator with the nonparametric adjustment as the NP estimator. Note that we use the estimator $\hat{q}_a(\tau)$ of $q_a(\tau)$ in (ref), where $\hat{q}_a(\tau)$ satisfies Assumption (ref). All the analysis in this section takes account of the fact that $\hat{q}_a(\tau)$ instead of $q_a(\tau)$ is used.
Four remarks are in order. First, Assumption (ref)(i) is standard in the sieve literature. Second, Assumption (ref)(ii) means the approximation error of the sieve logistic model vanishes asymptotically, which holds given sufficient smoothness of $\mathbb{P}(Y_i(a)\leq q_a(\tau)|S_i=s,X_i=x)$ in $x$. Third, Assumption (ref)(iii) usually holds when $\text{Supp}(X)$ is compact. This condition is also assumed by HIR03. Fourth, the quantity $\zeta(h_n)$ in Assumption (ref)(iv) depends on the choice of basis functions. For example, $\zeta(h_n) = O(h_n^{1/2})$ for splines and $\zeta(h_n) = O(h_n)$ for power series. Taking splines as an example, Assumption (ref)(iv) requires $h_n = o(n^{1/2})$.
Three remarks are in order. First, as the nonparametric regression consistently estimates the true specifications $\{m_a(\cdot)\}_{a=0,1}$, the QTE estimator adjusted by the nonparametric regression achieves the global minimum asymptotic variance, and thus is weakly more efficient than QTE estimation with linear and logistic adjustments studied in the previous section. Second, the practical implementation of NP and ML methods are the same, given that they share the same set of covariates (basis functions). Therefore, even if we include a small number of basis functions so that $h_n$ is better treated as fixed, the proposed estimation and inference methods for the regression-adjusted QTE estimator are still valid, although they may not be optimal. Third, in Section (ref) in the Online Supplement, we consider computing $\widehat{m}_a(\tau,s,x)$ via an $\ell_1$ penalized logistic regression when the dimension of the regressors can be comparable or even higher than the sample size. We then provide primitive conditions under which we verify Assumptions (ref) and (ref).
Two DGPs are used to assess the finite sample performance of the estimation and inference methods introduced in the paper. We consider the outcome equation
where $\gamma = 4$ for all cases while $\alpha(X_{i})$, $\mu(X_{i})$, and $\eta_{i}$ are separately specified as follows.
For each DGP, we consider the following four randomization schemes as in ZZ20 with $\pi(s) = 0.5$ for $s \in \mathcal{S}$:
We assess the empirical size and power of the tests for $n = 200$ and $n = 400$. We compute the true QTEs and their differences by simulations with 10,000 sample size and 1,000 replications. To compute power, we perturb the true values by $\Delta=1.5$. We examine three null hypotheses:
For the pointwise test, we report the results for the median ($\tau = 0.5$) in the main text and give the cases $\tau = 0.25$ and $\tau = 0.75$ in the Online Supplement.
We consider the following estimation methods of the auxiliary regression.
Table (ref) presents the empirical size and power for the pointwise test with $\tau = 0.5$ under DGPs (i) and (ii). We make six observations. First, none of the auxiliary regressions is correctly specified, but test sizes are all close to the nominal level 5%, confirming that estimation and inference are robust to misspecification. Second, the inclusion of auxiliary regressions improves the efficiency of the QTE estimator as the powers for method “NA" are the lowest among all the methods for both DGPs and all randomization schemes. This finding is consistent with theory because methods “LP", “LPML", “LPMLX", “NP" are guaranteed to be weakly more efficient than “NA". Third, the powers of methods “LPML" and “LPMLX" are higher than those of methods “ML" and “MLX", respectively. This is consistent with our theory that methods “LPML" and “LPMLX" further improve “ML" and “MLX", respectively. In addition, methods “MLX" and “LPMLX" fit a flexible distribution regression that can approximate the true DGP well. Therefore, the powers of “MLX" and “LPMLX" are respectively much larger than those of “ML" and “LPML". For the same reason we observe that the power of “LPMLX" is close to “NP".\footnote{The results in Section (ref) of the Online Supplement show that “LPMLX" has much smaller bias than “NP" and its variance is similar to “NP", which make “LPMLX" preferable in practice.} Fourth, the powers of method “NP" are the best because it estimates the true specification and achieves the minimum asymptotic variance of $\hat{q}^{adj}(\tau)$ as shown in Theorem (ref). Fifth, when the sample size is 200, the method “NP" slightly over-rejects but size becomes closer to nominal when the sample size increases to 400. Sixth, the improvement of power of “LPMLX" estimator upon “NA" (i.e., with no adjustments) is due to about 12-15% reduction of the standard error of the QTE estimator on average.\footnote{The bias and standard errors are reported in the Section (ref) in the Online Supplement.}
Tables (ref) and (ref) present sizes and powers of inference on $q(0.75) - q(0.25)$ and on $q(\tau)$ uniformly over $\tau \in [0.25,0.75]$, respectively, for DGPs (i) and (ii) and four randomization schemes. All the observations made above apply to these results. The improvement in power of the “LPMLX" estimator upon “NA" (i.e., with no adjustments) is due to a 9% reduction of the standard error of the difference of the QTE estimators on average. In Section (ref) in the Online Supplement, we provide additional simulation results such as the empirical sizes and powers for the pointwise test with $\tau = 0.25$ and $0.75$, the bootstrap inference with the true target fraction, and the adjusted QTE estimator when the DGP contains high-dimensional covariates and the adjustments are computed via logistic Lasso. We also report the biases and standard errors of the adjusted QTE estimators.
When $X$ is finite-dimensional, we suggest using the LPMLX adjustment in which the logistic model includes interaction terms and the regression coefficients are allowed to depend on $(\tau,a,s)$. When $X$ is high-dimensional, we suggest using the logistic Lasso to estimate the regression adjustment.\footnote{The relevant theory and simulation results on high-dimensional covariates are provided in Section (ref) of the Online Supplement.}
Undersaving has been found to have important individual and social welfare consequences karlan2014. Does expanding access to bank accounts for the poor lead to an overall increase in savings? To answer the question, dupas2018 conducted a covariate-adaptive randomized experiment in Uganda, Malawi, and Chile to study the impact of a bank account subsidy on savings. In their paper, the authors examined the ATEs as well as the QTEs of the subsidy. This section reports an application of our methods to the same dataset to examine the QTEs of the subsidy on household total savings in Uganda.
The sample consists of 2160 households in Uganda.\footnote{We filter out observations with missing values. Our final sample contains 1952 households.} Within each of 41 strata by gender, occupation, and bank branch, 50 percent of households in the sample were randomly assigned to receive the bank account subsidy and the rest of the sample were in the control group. This is a block stratified randomization design with 41 strata, which satisfies Assumption (ref) in Section (ref). The target fraction of the treated units is 1/2. It is trivial to see that statements (i), (ii), and (iii) in Assumption (ref) are satisfied. Because $\max_{s\in\mathcal{S}}|\frac{D_{n}(s)}{n(s)}|\approx 0.056$, it is reasonable to claim that Assumption (ref)(iv) is also satisfied in our analysis.
After the randomization and the intervention, the authors conducted 3 rounds of follow-up surveys in Uganda (See dupas2018 for a detailed description). In this section, we focus on the first round follow up survey to examine the impact of the bank account subsidy on total savings.
Tables (ref) and (ref) present the QTE estimates and their standard errors (in parentheses) estimated by different methods at quantile indices 0.25, 0.5, and 0.75. The description of these estimators is similar to that in Section (ref).\footnote{Specifically, we have:
} In the analysis, we focus on two sets of additional baseline variables: baseline value of total savings only (one auxiliary regressor) and baseline value of total savings, household size, age, and married female dummy (four auxiliary regressors). The first set of regressors follows dupas2018. The second one is used to illustrate all the methods discussed in the paper. Tables (ref) and (ref) report the results with one and four auxiliary regressors, respectively.
\newcolumntype{L}{>{\arraybackslash}X} \newcolumntype{C}{>{\arraybackslash}X}
\newcolumntype{B}{>{\hsize=1.44\hsize \arraybackslash}X} \newcolumntype{S}{>{\hsize=.89\hsize \arraybackslash}X}
The results in Tables (ref)-(ref) prompt two observations. First, consistent with the theoretical and simulation results, the standard errors for the regression-adjusted QTEs are mostly lower than those for the QTE estimate without adjustment. This observation holds for most of the specification and estimation methods of the auxiliary regressions.\footnote{The efficiency gain from the “NP" adjustment is not the only reason for its small standard error at the 25% QTE. Another reason is that the treated outcomes around this percentile themselves do not have much variation.} For example, in Table (ref), the standard errors for the “LPML” QTE estimates are 16.7% less than those for the QTE estimate without adjustment at the 25th percentile, respectively. For another example in Table (ref), at the 25th percentile, the standard error for the “LPMLX” QTE estimates is 43.4% less than that for the QTE estimate without adjustment. At the median, the standard error for the “LPML” QTE estimates is 7% less than that for the QTE estimate without adjustment.
Second, there is substantial heterogeneity in the impact of the subsidy on total savings. In particular, we observe larger effects as the quantile indexes increase, which is consistent with the findings in dupas2018. For example, Table (ref) shows that, although the treatment effects are all positive and significantly different from zero at quantiles 25%, 50%, and 75%, the magnitude of the effects increases by over 200% from the 25th percentile to the median and by around 100% from the median to the 75th percentile.
The second observation suggests that the heterogeneous effects of the subsidy on savings are sizable economically. To evaluate whether these effects are statistically significant, we report statistical tests for the heterogeneity of the QTEs in Table (ref). Specifically, we test the null hypotheses: $H_0: q(0.5) - q(0.25)=0$, $H_0: q(0.75) - q(0.5)=0$, and $H_0: q(0.75) - q(0.25)=0$. Table (ref) shows that only the difference between the 50th and 25th QTEs is statistically significant at the 5% significance level.
How does the variation in the impact of the subsidy appear across the distribution of total savings? The QTEs on the distribution of savings are plotted in Figure (ref), where the shaded areas represent the 95% confidence region. The figure shows that the QTEs seem insignificantly different from zero below about the 20th percentile. At higher levels to near the 80th percentile, the treatment group savings exceed the control group savings at an accelerated rate, yielding increasingly significantly positive QTEs. Beyond the 80th percentile, the QTEs again become insignificantly different from zero. These findings point to notable distributional heterogeneity in the impact of the subsidy on savings.
This paper proposes the use of auxiliary regression to incorporate additional covariates into estimation and inference relating to unconditional QTEs under CARs. The auxiliary regression model may be estimated parametrically, nonparametrically, or via regularization if there are high-dimensional covariates. Both estimation and bootstrap inference methods are robust to potential misspecification of the auxiliary model and do not suffer from the conservatism due to the CAR. It is shown that efficiency can be improved when including extra covariates. When the auxiliary regression is correctly specified, the regression-adjusted estimator further achieves the minimum asymptotic variance. In both the simulations and the empirical application, the proposed regression-adjusted QTE estimator performs well. These results and the robustness of the methods to auxiliary model misspecification reflect the aphorism widespread in scientific modeling that all models may be wrong, but some are useful.\footnote{The aphorism “all models are wrong, but some are useful" is often attributed to the statistician George Box1976. But the notion has many antecedents, including a particularly apposite remark made in 1947 by John Neumann2019 in an essay on the empirical origins of mathematical ideas to the effect that “truth ... is much too complicated to allow anything but approximations”.}
We thank the Managing Editor, Elie Tamer, the Associate Editor and three anonymous referees for many useful comments that helped to improve this paper. We are also grateful to Michael Qingliang Fan and seminar participants from the 2022 Econometric Society Australasian Meeting, the 2021 Nanyang Econometrics Workshop, University of California, Irvine, and University of Sydney for their comments.
\paragraph{Funding:} Yichong Zhang acknowledges financial support from the Singapore Ministry of Education under Tier 2 grant No. MOE2018-T2-2-169, the NSFC under the grant No. 72133002, and a Lee Kong Chian fellowship. Peter C. B. Phillips acknowledges support from NSF Grant No. SES 18-50860, a Kelly Fellowship at the University of Auckland, and a Lee Kong Chian Fellowship. Yubo Tao acknowledges the financial support from the Start-up Research Grant of University of Macau (SRG2022-00016-FSS). Liang Jiang acknowledges support from MOE (Ministry of Education in China) Project of Humanities and Social Sciences (Project No.18YJC790063).