EconBase
← Back to paper

Semiparametric Estimation of Treatment Effects in Randomized Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

144,295 characters · 36 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Estimation of Treatment Effects in Randomized Experiments

abstractWe develop new semiparametric methods for estimating treatment effects. We focus on settings where the outcome distributions may be thick tailed, where treatment effects may be small, where sample sizes are large and where assignment is completely random. This setting is of particular interest in recent online experimentation. We propose using parametric models for the treatment effects, leading to semiparametric models for the outcome distributions. We derive the semiparametric efficiency bound for the treatment effects for this setting, and propose efficient estimators. In the leading case with constant quantile treatment effects one of the proposed efficient estimators has an interesting interpretation as a weighted average of quantile treatment effects, with the weights proportional to minus the second derivative of the log of the density of the potential outcomes. Our analysis also suggests an extension of Huber's model and trimmed mean to include asymmetry.

Keywords: Potential Outcomes, Average Treatment Effects, Quantile Treatment Effects, Semiparametric Efficiency Bound

\baselineskip=15pt \setcounter{page}{1}

Introduction

Historically, randomized experiments were often carried out in medical and agricultural settings. In these settings, sample sizes were often modest, typically on the order of hundreds or (more rarely) thousands of units. Outcomes commonly studied included mortality or crop yield, and were characterized by relatively well-behaved distributions with thin tails. Standard analyses in those settings typically involved estimating the average effect of the treatment using the difference in average outcomes by treatment group, followed by constructing confidence intervals using Normal distribution based approximations. These methods originated in the 1920s, {\it e.g.}, neyman1923, fisher1937design, but they continue to be the standard in modern applications. See wu2011experiments for a recent discussion.

More recently many experiments are conducted online (see kohavi2020trustworthy for an overview), leading to substantially different settings. gupta2019top claim that “Together these organizations [Airbnb, Amazon, Booking.com, Facebook, Google, LinkedIn, Lyft, Microsoft, Netflix, Twitter, Uber, and Yandex] tested more than one hundred thousand experimental treatments last year." The settings for these online experiments are substantially different from those in biomedical and agricultural settings. First, the experiments are often on a vastly different scale, with the number of units on the order of millions to tens of millions. Second, the outcomes of interest, variables such as time spent by a consumer, sales per consumer or payments per service provider, are characterized by distributions with extremely thick tails. Third, the treatment effects are often extremely small relative to the standard deviation of the outcomes, even if their magnitude remains substantively important. For example, lewis2015unfavorable analyze challenges with statistical power in experiments designed to measure the effect of digital advertising on consumer expenditures. They discuss a hypothetical experiment where the average expenditure per potential customer is \$7 with a standard deviation of \$75, and where an average treatment effect of \$0.35 (0.005 of a standard deviation) would be substantial in the sense of being highly profitable for the company given the cost of advertising. In the Lewis and Rao example, an experiment with power 0.8 for a treatment effect of \$0.35, and a significance level for the two-sided test of means of 0.05, would require a sample size of 1.4 million customers. As a result confidence intervals for the average treatment effect are likely to include zero even if the true effects were substantively important and samples are large. Even if a confidence interval for the average treatment effect includes zero there may be evidence about the presence of causal effects of the treatment. Using Fisher exact p-value calculations fisher1937design with well-chosen statistics ({\it e.g.}, the Hodges-Lehman difference in average ranks, rosenbaum1993hodges), one may well be able to establish conclusively that treatment effects are present. However, the magnitude of the treatment effect, rather than its presence, is typically important for decision makers.

This sets the stage for the problem we address in this paper. In the absence of additional information, there exists no estimator for the average treatment effect that is more efficient than the difference in means. To obtain more precise estimates, we either need to change the focus away from the average treatment effect, or we need to make additional assumptions. One approach to changing the question, at least slightly, is to transform the outcome ({\it e.g.} taking logarithms or winsorizing) followed by a standard analysis estimating the average effect of the treatment on the transformed outcome. In this paper, like taddy2016scalable, tripuraneni2021meta, we choose a different approach, namely making additional assumptions on the joint distribution of the outcomes and treatment indicator.

The key conceptual contribution is that we postulate a semi-parametric model for the outcome distributions by treatment group. The leading example of this semi-parametric model corresponds to restricting the quantile treatment effects to be identical across quantiles, thus assuming that the two conditional outcome distributions differ only by a shift. We do not directly use parametric models for the outcome distributions by treatment group, because specifying such a model that well approximates the full outcome distribution is more challenging than postulating a model for the treatment effects. Unlike outcomes, treatment effects tend to be small and often have little variation. For this semiparametric set-up ({\it e.g.}, bickel1993efficient, bickel2015mathematical), we derive the influence function, the semiparametric efficiency bound, and we propose semiparametrically efficient estimators.

It turns out that the parametrization of the treatment effect can be very informative, potentially making the asymptotic variance for the corresponding semiparametric estimators substantially smaller than the asymptotic variance for the difference-in-means estimator. For example, if the potential outcomes have Cauchy distributions, the variance bound for the average treatment effect is infinite because the moments of the Cauchy distribution do not exist. However, under the constant additive treatment effect assumption (implying that the quantile treatment effects are identical), the semiparametric variance bound for the treatment effect is finite.

In addition, even if this model for the treatment effect is misspecified, the estimand corresponding to proposed estimators continue to have a causal interpretation, as a weighted average of quantile treatment effects, making it an easy-to-implement and attractive choice in practice.

The remainder of the paper is organized as follows. First, in Section (ref) we consider the leading case where we assume the two potential outcome distributions differ only by a shift, so that the quantile treatment effects are all identical. This is implied by, but does not require, the assumption that the treatment effect is additive and constant. In Section (ref) we consider the case where we have more flexible parametric models linking the two conditional outcome distributions. In Section (ref) we provide some simulation evidence regarding the finite sample properties of the proposed methods in controlled settings and provide real data illustrations. Section (ref) concludes. A software implementation for R is available at \url{https://github.com/michaelpollmann/parTreat}.

Constant Quantile Treatment Effects

In this section we focus on a special case with constant quantile treatment effects. After setting up the problem formally we discuss robust estimation in the one-sample case to motivate a class of weighted quantile treatment effect estimators. We then discuss the formal semiparametric problem and show adaptivity of the proposed estimators. Finally we consider partial adaptivity and robustness.

This case is closely related to the classical two-sample problem, as discussed in hodges1963estimates, and to problems considered in the literature on robust descriptive statistics as in bickel1975descriptiveI, bickel1975descriptiveII, doksum1974empirical, doksum1976plotting, and in particular jaeckel1971robust, jaeckel1971some. Section (ref) can be interpreted as an extension of Jaeckel's work, in a causal inference framework, to the two sample context in the setting of semiparametric theory. In the process of doing so, we generalize Huber's model huber1964 and the estimator based on trimmed means to include asymmetry, and present a simplified version of the results of chernoff1967asymptotic (see also bickel1967some, govindarajulu1967generalizations and stigler1974linear on linear combinations of order statistics). In particular, we exhibit efficient M (maximum-likelihood type) and L (linear combination of order statistics) estimates for outcome distributions that are known up to a shift. We then analyze fully adaptive estimates of both types, as discussed in bickel1993efficient, and partially adaptive estimates, in particular flexible trimmed means (jaeckel1971robust). In this setting the problem is closely related to the literature on robust estimation of locations ({\it e.g.,} bickel1975descriptiveI, bickel1975descriptiveII, hampel2011robust, huber2011robust, bickel1976descriptive, bickel2012descriptiveIV).

Set Up

We consider a set up with a randomized experiment with $n$ observations drawn randomly from a large population. With probability $p\in(0,1)$ a unit is assigned to the treatment group. Let $n_1$ and $n_0=n-n_1$ denote the number of units assigned to the treatment and control group. Following neyman1923, rubin1974estimating, imbens2015causal, let $Y_i(0)$ and $Y_i(1)$ denote the two potential outcomes for unit $i$, and let the treatment be denoted by $Z_i\in\{0,1\}$. We assume that the treatment assignment for one unit does not affect the outcomes for any other unit. For all units in the sample we observe the pair $(Z_i,Y_i)$, where $Y_i\equiv Y_i(Z_i)$. The cumulative distribution functions for the two potential outcomes are $F_0(y)$ and $F_1(y)$ with inverses $F^{-1}_{0}(u)$ and $F^{-1}_{1}(u)$, and means and variances $\mu_0$, $\mu_1$, $\sigma_0^2$ and $\sigma_1^2$. Note that by the random assignment assumption the distribution of the potential outcome $Y_i(z)$ is identical to the conditional distribution of the realized outcome $Y_i$ conditional on $Z_i=z$: $F_z(y)\equiv \Pr(Y_i(z)\leq y)=\Pr(Y_i\leq y|Z_i=z)$.

We are interested in the average treatment effect in the population,

equation[equation omitted — 52 chars of source]

The natural estimator for this average treatment effect is the difference in sample averages

equation[equation omitted — 197 chars of source]

are the averages of the observed outcomes by treatment group. Under standard conditions $\frac{n_{1}}{n}\stackrel{P}{\rightarrow}p$, $\frac{n_{0}}{n}\stackrel{P}{\rightarrow}(1-p)$, and

align[align omitted — 149 chars of source]

The concern is that this conventional estimator $\hat\tau$ may be imprecise. In particular in settings where the outcome distribution is thick tailed, sometimes extremely so, confidence intervals may be wide. We address this issue in this paper by imposing some restrictions on the two potential outcome distributions. Following the semiparametric literature bickel1975descriptiveI, bickel1975descriptiveII, bickel1976descriptive, bickel2012descriptiveIV we exploit these restrictions to develop new estimators.

Weighted Average Quantile Treatment Effects

It is useful to start with quantile treatment effects lehmann1975nonparametrics, which play an important role in our setup. For quantile $u\in(0,1)$, define

equation[equation omitted — 97 chars of source]

These quantile treatment effects are closely related to what doksum1974empirical and doksum1976plotting label the {\it response} function: $R(y) \equiv F^{-1}_{1}\left( F_{0}(y)\right)-y=\Delta(F_0(y)) .$ The natural estimate for the quantile treatment effect is the empirical plug-in, $ \hat{{\Delta}}(u)\equiv\hat{F}_{1}^{-1}(u)-\hat{F}_{0}^{-1}(u) $ where $\hat{F}_{1}^{-1}(u)$ equals $Y_{([n_1 u])}^{(1)}$, defined as the $[n_1u]^{\rm th}$ order statistic of $Y_{i}|Z_{i}=1$, $i=1,\dots,n_{1}$, where $n_{1}=\sum_{i=1}^{n}Z_{i}$ and similarly for $\hat{F}_{0}^{-1}(u)$.

A natural class of parameters summarizing the difference between the $Y_i(1)$ and $Y_i(0)$ distributions consists of weighted averages of the quantile treatment effects:

align*[align* omitted — 70 chars of source]

where the weights integrate to one, $W(0)=0$, $W(1)=1$. Different choices for the weight function correspond to different estimands. The constant weight case, $W'(u)\equiv 1$, corresponds to the population average treatment effect $\tau=\mathbb{E}[Y_i(1)-Y_i(0)]$. The median corresponds to the case where $W(\cdot)$ puts all its mass at $1/2$. We thus allow $W(\cdot)$ to permit point masses.

For a given weight function $W(\cdot)$ we can estimate the parameter $\tau(F_{0},F_{1};W)$ using a {\it weighted average quantile} estimator:

equation[equation omitted — 149 chars of source]

\[\hskip1cm=\frac{1}{n_{1}}\sum_{i=1}^{n_{1}}w_{i}^{(1)}Y_{(i)}^{(1)}-\frac{1}{n_{0}}\sum_{i=1}^{n_{0}}w_{i}^{(0)}Y_{(i)}^{(0)}, \] where

align*[align* omitted — 97 chars of source]

and $Y_{(i)}^{(z)}$ again are the order statistics in treatment group $z$.

Efficient estimation of Weighted Average Quantile Treatment Effects Using Influence Functions

To understand the properties of the weighted average quantile estimator $\hat\tau_W$, we begin by considering the nonparametric model for a single sample, with cumulative distribution $F(\cdot)$ where the interest is in the weighted quantile $\int_{-\infty}^{\infty}F^{-1}(u)dW(u)$ for a given weight funtion $W(\cdot)$. For this one-sample case, jaeckel1971robust,jaeckel1971some, building on chernoff1967asymptotic, shows that under simple conditions on $W$ and $F$, for a sample size of $n$,

align[align omitted — 149 chars of source]

where the {\it influence function} $\psi$ is related to the weight function $W$ by

align[align omitted — 135 chars of source]

The last term ensures that $\int_{-\infty}^{\infty}\psi(x,F,W)dF(x)=0$.

Note that if the derivatives $\psi'(\cdot)$ and $W'(\cdot)=w(\cdot)$ exist, by (ref),

align*[align* omitted — 93 chars of source]

so that $\psi'(x,F,W)=w(F(x))$. Note that for the median, (ref) yields, $\psi(x,F,W)=\frac{sign(x-F^{-1}(\frac{1}{2}))}{2f(F^{-1}(\frac{1}{2}))}$. Our formula (ref) is slightly more general than Jaeckel's, and in the appendix we establish sufficient conditions on the cumulative distribution function $F(\cdot)$ and the weight function $W(\cdot)$ for our version of his result to hold.

Expression ((ref)) in turn implies that

align*[align* omitted — 127 chars of source]

where the variance equals the expectation of the square of the influence function: \[ \sigma^{2}(F,\omega)=\int_{-\infty}^{\infty}\psi(x,F,W)^{2}dF(x). \] The results in jaeckel1971robust,jaeckel1971some for the one-sample case extend in the following way to the two-sample setting that is our primary focus. If $\tau(F_{0},F_{1},W)$ is estimated by $\hat{\tau}_W$ in ((ref)), then, under regularity conditions given in the appendix (Theorem A.1),

align[align omitted — 191 chars of source]

where $\psi(x,F,W)$ is given by (ref).

Constant Quantile Treatment Effects

Now let us return to the primary focus of this section, the estimation of the average treatment effect under the constant quantile treatment effect assumption. Our key assumption in this section is that the quantile treatment effects are all equal:

align[align omitted — 77 chars of source]

and thus, for any weight functions $W(\cdot)$,

align[align omitted — 41 chars of source]

Later, in Section 3, we generalize this to allow for a more general parametric function linking the quantile treatment effects. One way to motivate the constant quantile treatment effect assumption is to assume that the unit-level treatment effects are all constant, $Y_i(1)-Y_i(0)=\tau$ for all units $i=1,\ldots,n$. This implies, but is not implied by, the assumption that all the quantile treatment effects are identical. The assumption of constant unit-level treatment effects is very strong, implying rank-invariance, which is in fact stronger than what we need.

In this section, for expository reasons we further assume that we know the control outcome distribution $F_{0}(\cdot)$ up to a shift. That is, $F_{0}(x)=F(x-\eta)$ where $F(\cdot)$ (with derivative $f$) is known and $\eta$ unknown. Because of the constant quantile treatment effect assumption the treated potential outcome distribution is also known up to a shift, $F_{1}(x)=F(x-\eta-\tau)$. Assuming that $F_0(\cdot)$ is known up to a shift is unrealistic in practice, and we remove this assumption below in Section 2.5, but it allows us to focus in this section on some key insights.

For this fully parametric model (with unknown parameters $\eta$ and $\tau$), if the Fisher information $I(f)=\int\left(\frac{f'}{f}\right)^{2}(x)f(x)dx=\int \left(-\frac{f'}{f}\right)'(x)f(x)dx$ satisfies $0<I(f)<\infty$, the maximum likelihood estimator of $\tau$, suitably regularized ({\it e.g.,} lecam-yang-1988), has influence function,

align[align omitted — 178 chars of source]

There is an interesting alternative efficient estimator in this known $f(\cdot)$ case. Suppose $f'/f$ is absolutely continuous. Then the weight function

equation[equation omitted — 176 chars of source]

provides an efficient $L$ estimate when substituted appropriately in (ref), leading to

align[align omitted — 134 chars of source]

It is interesting to inspect the form of the weights $w_f(u)$. These weights are proportional to minus the second derivative of the logarithm of the density function. In other words, we can approximate the efficient estimator by first estimating a large number of quantile treatment effects. Under the model these quantile treatment effects are all identical. To efficiently estimate that common treatment effect we can simply use a weighted average of the estimated quantile treatment effects. It turns out the optimal weights simplify to minus the second derivative of the logarithm of the density. For the Normal distribution, that means the weights are constant. For the double exponential distribution the weights put point mass at the median. For the Cauchy distribution the weights are proportional to $-\cos(2\pi u) \sin(\pi u)^{2}$. Interestingly these weights are negative for some quantiles. One can of course see this by inspecting the estimated weights. If one is concerned by the negative weights one can also modify them by restricting them to be nonnegative. Finally, note that implicitly the influence function estimator also has the negative weights in such cases because the two estimators are first order equivalent.

A final comment connects this to common methods for dealing with thick tailed distributions. In practice many researchers use winsorizing to deal with these problems. This can be interpreted as using a weighted average quantile estimator with a particular set of weights. Specifically, with winsorizing at the $q$ and $1-q$ quantiles, the implicit weights are constant on the interval $(q,1-q)$, and then put additional point mass $q$ on the $q$th and $(1-q)$th quantiles. As discussed in bickel1965some, the asymptotic properties of the winsorizing estimator depend delicately on the density at the winsorizing quantiles. In our simulations this estimator does not perform particularly well. Like other settings, there is tension here between having an interpretable target that may not be precisely estimable ({\it e.g.,} the average effect of the treatment), versus a precisely estimable estimand whose interpretation is more complex ({\it e.g.} the weighted average quantile effect). This tension arises also in other settings. An example is the estimation of average treatment effects under unconfoundedness where weighting by the confounders may affect both the interpretation of the estimand and the precision with which we can estimate it crump2009dealing, li2018balancing. Another setting is that discussed in vansteelandt2022assumption. The use of quantile methods for estimating treatment effects in thick-tailed settings has been studied in firpo2007efficient, firpo2009unconditional, but unlike in those papers, our focus is on the overall treatment effect, rather than the effect at specific quantiles.

Fully adaptive estimation

As stated earlier, in practice we do not know the density $f(\cdot)$ up to location. However, in this case with constant quantile treatment effects this knowledge does not matter up to first order. Because of the orthogonality of the tangent space with respect to $f$, it follows from semiparametric theory bickel1993efficient that even if the density $f$ is unknown, substituting a suitable estimate of $f$ (and $\eta$) in (ref) or ((ref)), will yield estimators with influence functions given by (ref), or equivalently by

align[align omitted — 177 chars of source]

where $f_0(\cdot) \equiv F_0'(\cdot) \equiv f(\cdot - \eta)$.

For our proposed estimator we split the data randomly into two parts, with the two subsamples denoted by $A$, corresponding to $\{(Z_{i},Y_{i}): 1\leq i \leq \frac{n}{2}\}$, and $B$, corresponding to $\{(Z_{i},Y_{i}): \frac{n}{2} < i\leq n\}$. The $M$ estimate using the estimated $\hat f(\cdot)$ is of the form,

align[align omitted — 275 chars of source]

where $\hat{f}_{0(A)}$ is an estimate of $f_0$ using $\{(Z_{i},Y_{i}): 1\leq i \leq \frac{n}{2}\}$, and $\hat{f}_{0(B)}$ using $\{(Z_{i},Y_{i}): \frac{n}{2} < i\leq n\}$, a one step estimate using the sample splitting technique klaassen1987consistent. $\tilde\tau_{(A)}$ and $\tilde\tau_{(B)}$ are initial $\sqrt n$ consistent estimates based on the two subsamples, for example based on the difference in medians or other quantiles. Algorithm (ref) shows the key steps; additional details are given in the supplementary materials.

algorithm[algorithm omitted — 2,026 chars of source]

We can also construct an $L$ estimate based on an average of the quantile differences. This estimator is obtained by first estimating $F_0(\cdot)$, $f_0(\cdot)$, and $f_0'(\cdot)$, substituting that for $F(\cdot)$, $f(\cdot)$, and $f'(\cdot)$ into $w_f(u)$ in equation ((ref)), followed by using this estimated set of weights in ((ref)), leading to

align[align omitted — 139 chars of source]

Formally we would use the same sample splitting as above. Details are in Algorithm (ref) and the supplementary materials.

algorithm[algorithm omitted — 2,202 chars of source]

Formally we have for the unknown $f(\cdot)$ case:

theoremFor all $f$ such that $f'$ exists and $0<I(f)<\infty$: \begin{enumerate} • There exist a $\sqrt{n}$-consistent estimator $\hat{\tau}$. • Under mild conditions (see (ref) and (ref) in the appendix), we can construct an $M$ estimate $\hat{\tau}^{if}$ such that \begin{align} \sqrt{n}(\hat{\tau}^{if}-\tau) & \stackrel{d}{\Rightarrow}\mathcal{N}\left(0,\frac{1}{p(1-p)I(f)}\right). \end{align} • Under mild conditions (see Lemma (ref) in the appendix), we can construct an $L$ estimate $\hat{\tau}^{waq}$ (weighted average quantile) such that $\hat{\tau}^{waq}$ also satisfies (ref). \end{enumerate}

The main insight is that the asymptotic variance for the proposed estimator is the same as the variance for the maximum likelihood estimator in the case where $F_0$ is known up to a shift. Our conditions are not optimal (see stone1975adaptive for minimal ones in the one sample case). The $\sqrt n$-consistent estimator for $\tau$ can be based on any quantile treatment effect estimator.

Thus, both the estimators $\hat{\tau}^{waq}$ and $\hat{\tau}^{if}$ are adaptive to models in which the distribution of control and treated potential outcomes is known up to location. Theorem (ref) implies that the influence function based estimator is as efficient as the maximum likelihood estimator based on the true distribution function. For instance, if potential outcomes are normally distributed, the maximum likelihood estimator is the difference in means, and the influence function estimator has the same limiting distribution. If, however, the potential outcomes follow a double exponential distribution, the difference in medians is the efficient estimator. Under this distribution, the influence function based estimator adapts and has the same limiting distribution as the difference in medians. For the Cauchy distribution the optimal weights are more complicated, \(w_f(u) \propto - \cos(2 \pi u) \sin(\pi u)^2\), but the influence function based estimator has the same limiting distribution, without requiring a priori knowledge about the distribution. This influence function based estimator $\hat{\tau}^{if}$ is a special case of the estimator developed by cuzick1992efficient,cuzick1992semiparametric for the partial linear regression model setting.

Although these estimators are efficient under the constant additive treatment model, as we shall see in Section 3, if the the constant quantile treatment effect assumption is violated, estimates of the types derived from $f$ known continue to estimate at rate $n^{-1/2}$, meaningful measures of the treatment effect as discussed in Section 2.1. This is unfortunately not the case for $\hat{\tau}^{if}$ and $\hat{\tau}^{waq}$ because estimation of $f'$ and $f''$ introduces components of variance of order larger than $n^{-1/2}$. However, there is a partial remedy, that we discuss next.

Partial Adaptation

An interesting alternative to fully adaptative estimation in the closely related one-sample symmetric case was studied by jaeckel1971some. See also yu2017robust. The estimator proposed by jaeckel1971some is

align*[align* omitted — 101 chars of source]

for a sample from $F(x-\mu)$ with $f$ symmetric. We can interpret this as restricting the class of weight functions to one indexed by a scalar parameter $\alpha$: \[ w_{\alpha}(u) = \left\{

array[array omitted — 109 chars of source]

\right. \] This weight function yields the $\alpha$-trimmed mean. We can then choose the value of $\alpha$ that minimizes the asymptotic variance of $\hat{\mu}_{\alpha}$. This asymptotic variance is equal to

align*[align* omitted — 244 chars of source]

with

align*[align* omitted — 117 chars of source]

which can be estimated by replacing $F$ with the empirical distribution, denoted as $\hat{\sigma}^2_{\alpha}$.

Let $\hat\alpha=\arg\min_\alpha\hat\sigma^2_\alpha$, and let $\hat{\mu}_{\hat\alpha}$ be the corresponding estimator for $\mu$. jaeckel1971some shows that $\hat{\mu}_{\hat\alpha}$ is adaptive for estimating $\mu$ over a Huber family of densities. In the Huber family $(X-\mu)/\sigma$ has density $f$ for varying $\mu$, $\sigma>0$: \[ \log f(x)= \left\{

array[array omitted — 116 chars of source]

\right. \] where $c(k)$ makes $\int f(x)dx=1$, and $k=-F^{-1}(\alpha)$. Adaptivity here means that using the trimming proportion optimizing the variance estimate, in fact yields an estimate which is efficient for the member of the Huber family generating the data. He optimizes $0<\alpha_0 \leq \alpha \leq \alpha_1 <\frac{1}{2}$.

Because this family includes among others the Gaussian ($k\rightarrow \infty$) and double exponential ($k=0$), this family is very flexible. For more properties, see huber2011robust.

In the two-sample case, it is reasonable to consider asymmetric weight functions leading to the natural generalization,

align[align omitted — 156 chars of source]

This estimator is partially adaptive, in a similar way to the symmetric trimmed mean in the one sample problem. In the online appendix, we extend Jaeckel's result on partial adaptation for our two sample problem to a generalization of the Huber family whose members are symmetric iff $k_{1}=k_{2}$, defined by $(X-\mu)/\sigma \sim f$:

align[align omitted — 260 chars of source]

where $c(k_{1},k_{2}) \equiv \log \Bigl(2\left(\frac{e^{-k_{1}/2}}{k_{1}} + \frac{e^{-k_{2}/2}}{k_{2}}\right) + \sqrt{2\pi}(\Phi(k_{2}) - \Phi(-k_{1}))\Bigr)$ and $\Phi$ is the CDF of $\mathcal{N}(0, 1)$. See Figure (ref) for illustration. This family can be equivalently parametrized by $F(-k_{1})$ and $F(k_{2})$. $f(\cdot)$ is symmetric if $k_1=k_2$.

figure[figure omitted — 193 chars of source]

In the supplementary materials we also discuss how inference can proceed in this setting.

The General Parametric Treatment Effect Case

In some settings, the assumption of an additive model may be too restrictive. In this section, we develop estimators given a general parametric model for this difference.

The starting point is a model governing the relation between the two potential outcomes:

assumption[Parametric Model Quantile Treatment Effects] The potential outcome distributions satisfy \begin{equation*} F_1(h(y,\theta)) = F_0(y). \end{equation*}

The constant quantile treatment effect case is a special case of this with $h(y,\theta)=y+\theta.$ Another important special case is the proportional treatment effect case, $h(y,\theta)=\theta y$. For the general case the weighted average quantile estimator does not directly generalize, so we focus on the influence-function-based estimator. For the general case the influence function is more complex.

This approach of modelling treatment effects has connections to the literature on structural nested models, which also imposes modeling restrictions on treatment effects, although for different reasons, largely based on the challenges in dynamic settings. See robins1986new for an early paper, and vansteelandt2014structural for a review.

As before we initially assume $F_0$ known. In terms of the quantile treatment effects $\tau(u)$ Assumption (ref) implies the restriction \[ \tau(u)=F_1^{-1}(u)-F^{-1}_0(u)=h(F_0^{-1}(u),\theta)-F^{-1}_0(u) .\] Given Assumption (ref), the population average treatment effect can be characterized as

equation*[equation* omitted — 90 chars of source]

In practice, however, estimating $\tau^{\rm pop}$ may still be subject to substantial sampling variance, even if $h(\cdot)$ is known. For example, suppose that $h(y,\theta)=\theta y$, so that the treatment effect is proportional. The average treatment effect is then $\theta E[Y_i|Z_i=0]$. Even if $\theta$ is known, estimating the population mean $E[Y_i|Z_i=0]$ could lead to a large standard error. As an alternative, we therefore focus on a different estimand. Specifically, we suggest to estimate the in-sample, as opposed to population, average treatment effect. This is still a well-defined average causal effect that is useful for decision makers. It is in the spirit of the typical analysis of randomized experiments based on convenience samples where the focus is on the average effect for the particular sample. A key insight is that estimators of this object can have a much lower variance. We define the in-sample average treatment effect as

equation[equation omitted — 236 chars of source]

When \(h\) is not just an additive function, \(\tau^{\rm is}\) is sample-dependent, and thus stochastic. In particular when the variance of \(Y_i\) is large because of thick tails for the potential outcome distributions, the variance of \(\tau^{\rm is}\) over repeated samples can be large, too. To give some intuition for this, suppose that $\theta$ is known. Then the variance of $\hat\tau-\tau^{\rm is}$ is zero, but the variance of $\tau^{\rm is} - \tau^{\rm pop}$ over repeated samples can be large. We therefore focus on the variance of estimators $\hat\tau$ relative to \(\tau^{\rm is}\) for the particular sample at hand, rather than on the variance of $\hat\tau$ relative to the population average \(\tau^{\rm pop}\).

If $F_0$ was known, we could estimate $\theta$ efficiently by some version of maximum likelihood to get an estimate $\hat{\theta}$ and

equation*[equation* omitted — 104 chars of source]

as an estimate of $\tau$. The density of $Y$ given $Z$ is

equation*[equation* omitted — 85 chars of source]

By Assumption (ref)

align[align omitted — 124 chars of source]

and the score function is

equation*[equation* omitted — 111 chars of source]

yielding,

equation*[equation* omitted — 120 chars of source]

where $I=E\left(\dot{\ell}(Y,Z,\theta)\right)^2$.

If $F_0$ is assumed unknown, to obtain an efficient influence function we must

itemize• Compute the tangent plane as $f_0$ varies with $\theta$ fixed. The tangent plane is \begin{equation*} \begin{aligned} \dot{P}_f = \Bigl\{ & \, u(Y,Z) = (1-Z) v(Y,\theta) + Z v(h^{-1}(Y,\theta),\theta): \\ & \qquad \int v^2(y,\theta)f_0(y)dy <\infty, \int v(y,\theta) f_0(y)dy=0 \Bigr\}. \end{aligned} \end{equation*} (Note that both factors of the likelihood must be varied treating $\theta$ as fixed.) • Project $\dot{\ell}$ on the orthocomplement of the tangent plane to get \begin{equation*} \dot{\ell}^*(Y,Z,\theta) = \dot{\ell}(Y,Z,\theta) - (Z Q(h^{-1}(Y,\theta), \theta) + (1-Z) Q(Y, \theta)) \end{equation*} where $Z Q(h^{-1}(Y,\theta), \theta) + (1-Z) Q(Y, \theta)$ is the projection of $\dot{\ell}$ on $ \dot{P}_f$. • The efficient influence function is given by \begin{equation*} \psi(Y,Z,\theta) = \frac{\dot{\ell}^*(Y,Z,\theta)}{ E\left(\dot{\ell}^*(Y,Z,\theta)\right)^2.} \end{equation*}
lemmaThe efficient influence function for $\theta$ is \begin{equation*} \psi_{f_0}(y, z, \theta )=I^{-1} \left\{ \frac{z}{p}\cdot g(y,\theta ) - \frac{1-z}{1-p} \cdot g(h(y,\theta ), \theta ) \right\}, \end{equation*} where \begin{align*} g(y, \theta ) &= \frac{\partial}{\partial \theta } \log f_1(y,\theta)= \frac{\partial}{\partial \theta } \log\left(f_0(h^{-1}(y,\theta))\cdot \frac{\partial h^{-1}(y,\theta)}{\partial y}\right), \end{align*} and \begin{align*} I = \int g^2(h(y,\theta ),\theta )f_0(y)dy. \end{align*}

To use the influence function approach we need a $\sqrt n$-consistent initial estimator $\tilde\theta$. We can do so by a fixed number of quantiles, $u_1,\ldots,u_d$, where $d$ is the dimension of $\theta$, and find the $\theta$ that solves \[ \hat F^{-1}_1(u)=h(\hat F_0^{-1}(u),\theta), \] for $u=u_1,\ldots,u_d$. For simplicity, we suggest using evenly-spaced quantiles, $u=\frac{1}{1+d},\frac{2}{1+d},\dots,\frac{d}{1+d}$. Let $\tilde\theta$ denote the solution to this system of equations. Then:

theoremUnder mild conditions on the estimation of density and its derivative, the estimator $\hat{\theta}^{if}$ below is efficient for $\theta$, i.e. $\sqrt{n}(\hat{\theta} - \theta) \stackrel{d}{\Rightarrow} \mathcal{N}(0,\frac{1}{p(1-p)} I^{-1})$: \begin{align*} \hat{\theta}^{if} & \equiv\tilde{\theta}+\frac{1}{n}\left\{\sum_{i=1}^{n/2}\psi_{\hat{f}_{0(2)}}(Y_i,Z_{i};\tilde{\theta})+\sum_{i=1+n/2}^{n}\psi_{\hat{f}_{0(1)}}(Y_i,Z_{i}; \tilde{\theta})\right\} \end{align*} where $\hat{f}_{0(1)}$ is the estimate of $f_0$ using $\{{(Y_i,Z_i): 1\leq i \leq \frac{n}{2}}\}$, and $\hat{f}_{0(2)}$ is the estimate of $f_0$ using $\{(Y_i,Z_i): \frac{n}{2} < i \leq n\}$, again a one step estimate using the sample splitting technique klaassen1987consistent, similar to (ref).

Note that, unlike the constant treatment effect setting, the efficient influence function is not necessarily the one corresponding to $F_0$ known.

Given inference for $\hat\theta$, inference for $\hat\tau$ as an estimator of the in-sample average treatment effect, is straightforward based on the Delta method and the representation in ((ref)), taken as given the potential outcomes. The asymptotic variance for $\hat\tau$ is equal to the variance for $\hat\theta$, pre and post multiplied by the derivative of the expression in ((ref)) with respect to $\theta.$

Simulations

We evaluate the performance of the proposed estimators and conventional estimators in a Monte Carlo study. Throughout most of these simulations, the true unit-level treatment effects are all zero. We estimate the treatment effect using the proposed efficient estimators based on an additive model. We consider seven estimators: The (standard) difference in means, the difference in medians, the Hodges-Lehman (hodges1963estimates) estimator,\footnote{The Hodges-Lehmann estimator is equal to the median of all pairwise differences between treated and control observations.} the adaptively trimmed mean, the adaptively winsorized mean,\footnote{We apply the ideas of jaeckel1971some for the optimal trimmed mean to choose the parameters for the estimators in Section (ref), see Theorem A.3 in the Appendix. We allow anywhere between no trimming (difference in means) and the extreme of trimming all but the medians (difference in medians). While including the extremes is not covered by the theory, this approach appeared to work best in our simulations.} the estimator based on the efficient influence function (eif), and the weighted average quantiles (waq) estimator. For the latter two we report only results without sample splitting. Results for the case with sample splitting are very similar and are available in the supplementary materials. Although in our illustrations we use relatively simple estimators for the densities and their derivatives based on variable bandwidth kernels, an alternative would be to use methods directy aimed at estimating derivatives of the logarithm of the density as in pinkse2021estimates. Detailed descriptions of how the new estimators are implemented, as well as R code implementing all estimators with performance optimizations for these simulations, are available in the supplemental materials.\footnote{The fully documented R package is available at \url{https://github.com/michaelpollmann/parTreat}. Details on the empirical implementation of our estimators are in the supplementary materials. For a sample of 1,000 treated and 1,000 control observations, the R package computes estimates and standard errors practically instantaneously. With very large samples, the derivatives of the log density can be precomputed on a random subsample of the data for similarly fast computation.} We present results for three sets of simulations, one with a range of known distributions for the potential outcomes, so we can directly assess the ability of the proposed methods to adapt to different distributions, and two with simulations based on real data: one based on housing prices and one based on medical expenditures, both with thick tailed distribution.

Simulations with Known Distributions

We simulate samples of \(n=\)20,000 observations, half of which are treated, and report summary statistics based on 10,001 simulated samples (using an odd number so that the median is unique). We repeat the simulation study for standardized Normal, Double Exponential (Laplace), and Cauchy distributions for the potential outcomes. The difference in means is the maximum likelihood estimator for Normally distributed data, and so will do well there, but may perform poorly for thicker tailed distributions such as the Double Exponential distribution and in particular the Cauchy distribution. The difference in medians is the maximum likelihood estimator for the Double Exponential distribution, and relatively robust to thick tails and outliers, and so is expected to perform reasonably well across all specifications, but not as well as the efficient estimators for the Normal.

For the simulations with known distributions we can derive the functional form for the optimal weights for the quantile-based estimator. The optimal weights for the waq estimator are proportional to the (estimated) second derivative of the log density. For the Normal distribution, \(\frac{\partial^{2} \ln f}{\partial y^{2}}(y) = - \frac{1}{\sigma^2}\), implying the optimal weights are constant. The density of the double exponential distribution is such that the optimal weights asymptotically place all weight close to the median. For the standard Cauchy distribution the efficient weights \(w\) on the difference in \(u\in(0,1)\) quantiles of treated and control distributions are \(w_f(u) \propto - \cos(2 \pi u) \sin(\pi u)^2\), shown in Figure (ref). Most of the weight is concentrated around the median, with strictly {negative} weights outside the $[0.25,0.75]$ quantile range.

figure[figure omitted — 640 chars of source]

The efficient estimators perform well across distributions, and confidence intervals based either on estimates of the analytic variance formulas or on the bootstrap achieve their nominal coverage levels, as shown in Table (ref). Their standard deviations are close to the theoretical efficiency bound, as shown in column 4, labelled relative efficiency, where values larger than one imply standard deviations of the estimator in excess of the efficiency bound. The efficient influence function and weighted average quantiles estimators are close to the most efficient estimator for the normal, double exponential, and Cauchy distribution. The last columns show that the confidence intervals are close to their nominal coverage for each distribution. For computational convenience in the simulations, the confidence intervals based on the bootstrap variance use the \(m\)-out-of-\(n\) bootstrap (bickel2012resampling), with \(m= \)2,000 (half treated, half control), to estimate the variance of the estimators. Even with these smaller sample sizes, the density estimates calculated within each bootstrap sample appear to be sufficiently good to yield reasonable confidence intervals for the estimators.

table[table omitted — 2,863 chars of source]

Simulations with House Price Data

In the second set of simulations we use house price data from the replication files of linden2008estimates available at linden2019replication. They obtained property sales data for Mecklenburg County, North Carolina, between January 1994 and December 2004. They dropped sales below \$5,000 and above \$1,000,000, such that 170,239 observations remain, which we take as our population of interest. Despite the trimming the distribution is noticeably skewed (skewness \(2.2\)) and thick tailed (kurtosis \(9.5\)). Even after taking logs, the distribution is heavy-tailed with kurtosis equal to \(5.1\). Figure (ref) plots a histogram for house prices, both in levels and in logs, along with the estimated optimal weights (minus the second derivative of the log density) based on all 170,239 observations.

figure[figure omitted — 625 chars of source]

We base simulations on this data by drawing samples of size \(n=\)20,000, and randomly assigning exactly half of each sample to the treatment group and the remaining half to the control group, with a zero treatment effect. Within each sample, observations are drawn from the population without replacement, but sampling is independent across samples, such that observations may appear in multiple samples. We estimate the efficiency bound using density estimates based on all observations. We compute the same estimators as in the simulations of the previous section, with a small adjustment to the adaptively trimmed and winsorized means where we fix the trimming and winsorizing percentiles on the left to 0% (no trimming/winsorizing), and only adaptively choose the threshold on the right.

Table (ref) summarizes the simulation results based on 10,001 simulated samples. When the house prices are in levels, the standard deviation of the difference in means estimator is twice as large as that of efficient estimators. For the difference in medians and the Hodges-Lehman estimators, which are less affected by outliers in the data, the standard deviation is larger by approximately 30% and 20%, respectively. Confidence intervals, based on estimated variances and asymptotic normal approximations, have close to nominal coverage throughout, and are meaningfully shorter for the efficient estimators we propose.

In Figure (ref) we show the root mean squared error and coverage of 95% confidence intervals both relative to the average treatment effect under deviations from the constant treatment effect model.\footnote{The design of these simulations was kindly suggested by one of the referees.} In the top panel, the unit-level treatment effects are independent draws from a normal distribution with mean equal to 0.1 standard deviations of the (population) standard deviation of the potential outcomes in the absence of treatment. On the horizontal axis, we vary the standard deviation of the normal distribution as a fraction \(q\) of the (population) standard deviation of the control potential outcomes, simulating 10,001 samples for each value. When \(q>0\), the constant treatment effect model is misspecified. In the bottom panel, the unit-level treatment effects are \(0\) with probability \(q\) and \(t\) with probability \(1-q\), where \(t\) is chosen as a function of \(q\) such that the average treatment effect is constant across all simulations and the same as in the top panel. The difference in means estimator is unbiased for the average treatment effect regardless of the value of \(q\), so its large root mean squared error is due to its variance. In both panels, the influence function-based and weighted average quantile estimators are only (asymptotically) unbiased for the average treatment effect when \(q=0\).

figure[figure omitted — 648 chars of source]

We also estimate a proportional treatment effect (multiplicative) model and translate the estimated coefficients into level effects. Under the multiplicative model, \(Y_{i}(1) = \theta Y_{i}(0)\). When outcomes are strictly positive, this is identical to an additive model for log outcomes; \(\log(Y_{i}(1)) = \log(\theta) + \log(Y_{i}(0))\). Given an estimate \(\hat{\tau}_{\rm log}\) of the additive model with the outcome in logs, the estimate of the level treatment effect is then

equation*[equation* omitted — 194 chars of source]

For estimates of the in-sample treatment effect, we therefore apply estimates \(\hat{\tau}_{\rm log}\) to a population of interest with known means \(\mu_{Y_{0}}\) and \(\mu_{Y_{1}}\) and fixed treatment probability \(p\) as

equation*[equation* omitted — 158 chars of source]

Using the Delta method, if \(V\) is the asymptotic variance of \(\hat{\tau}_{\rm log}\), then the asymptotic variance of \(\hat{\tau}\), holding \(\mu_{Y_{0}}\), \(\mu_{Y_{1}}\), and \(p\) fixed, is

equation*[equation* omitted — 112 chars of source]

which we estimate by replacing \(\tau_{\rm log}\) with \(\hat{\tau}_{\rm log}\) and \(V\) by the estimate of the variance, \(\hat{V}\). For the purpose of these simulations, we set \(\mu_{Y_{0}} = \mu_{Y_{1}}\) equal to the population mean of house prices, and \(p = 1/2\).

When treatment effects are assumed to be proportional to potential outcomes, the proposed estimators for this multiplicative model are still more efficient than alternative estimators, but the gains are smaller. The middle panel of Table (ref) shows the quality of estimates of the multiplicative parameter obtained by transforming outcomes into logs, \(\log(\theta)\). As can be seen in panel (b) of Figure (ref), the distribution of log house prices appears closer to the normal distribution with fewer “outliers.” Consequently, the difference in means estimator, which is efficient for the normal distribution, comes noticeably closer to the efficiency bound than when outcomes are in levels. Nevertheless, the efficient estimators further reduce the variance. The bottom panel of Table (ref) shows the same summary statistics when the multiplicative parameter is translated back into a level effect. The efficient estimators of the multiplicative parameter lead to treatment effects with smaller variance and shorter confidence intervals than the difference in means, irrespective of whether the latter is estimated in levels (top panel) or in logs and then translated into levels (bottom panel).

table[table omitted — 3,087 chars of source]

Medical Expenditures Data

Next, we present simulation results based on confidential medical expenditure data from the IBM MarketScan Research Database, following the sample construction of koenecke2021alpha. We restrict the sample to males, age 45--64, with pneumonia inpatient diagnosis and at least 1 year of continuous medical enrollment. For each patient, we consider the first inpatient admission only to abstract away from any dynamics. We focus on medical expenditure as the outcome variable. For each patient, we sum the payments recorded by MarketScan for this admission. In total, we use data on 103,662 admissions.\footnote{The sample size differs slightly from that reported by koenecke2021alpha due to missing expenditure data for a small number of admissions.} Figure (ref) plots a histogram for medical expenditure in levels and in logs along with the estimated optimal weights (minus the second derivative of the log density). We also observe a treatment variable in this data set, the (prior) use of alpha blockers, which koenecke2021alpha find may improve health outcomes during respiratory distress by preventing hyperinflammation.

figure[figure omitted — 1,146 chars of source]

We design a simulation study similar to those of the previous sections, treating the receipt of alpha blockers as randomly assigned in our population. This allows us to study (coverage) properties of the estimators and inference procedures in settings where the parametric treatment effect model is not (necessarily) correctly specified. For these simulations, each sample is a draw, without replacement, of 200 of the 5,507 observations in the treatment group and 3,565 of the 98,155 observations in the control group. While the control group is smaller than in our other simulations, it remains sufficiently large to estimate the density and its derivatives required for our estimators. In this simulation design, \(F_{0}\) is given by the empirical distribution of the control group, and \(F_{1}\) is given by the empirical distribution of the treatment group, of the full sample. Although it is not necessarily correct, the additive treatment effect model may offer a reasonable approximation, and we are interested in the performance of inference methods when the conditions for our theoretical results are not (quite) met in this particular application. The simulation results reported in Table (ref), for the same estimators as in Section (ref), suggest that the proposed estimators perform reasonably well in this setting despite mis-specification.

table[table omitted — 2,790 chars of source]

Conclusion

In many modern settings where randomized experiments are used to estimate treatment effects the presence of heavy-tailed distributions can lead to larger standard errors. Often researchers use winsorizing with {\it ad hoc} thresholds to address this. Here we develop systematic methods for obtaining more precise inferences using parametric models for the treatment effects, while avoiding the specification of models for the potential outcomes. We present results for semiparametric effiency bounds, suggest efficient estimators, and show in simulations that these methods can be effective in realistic settings.

In particular we recommend the semiparametrically efficient estimator under the constant additive treatment effect model. Although one may not think the constant additive treatment effect assumption holds exactly, the fact that the estimator can be interpreted as estimating a weighted average of the quantile treatment effects make this an attractive choice.

In this discussion we do not incorporate covariates or pretreatment variables. One could combine the ideas exposited in the current paper with models for the control potential outcome. One may also wish to incorporate covariates in the model for the quantile treatment effects and thus allow for heterogenous treatment effects {\it e.g., } chen2022robust, wager2018estimation. Another interesting avenue is to consider alternative estimators for the current model by first estimating the unrestricted quantile functions, followed by minimum distance methods, as in alvarez2022inference, alvarezsemiparametric.

Conflict of interest: We have no conflicts of interest to disclose.

Data availability

The source for the house price data is linden2019replication. The relevant columns of the data are also included in the supplementary materials of this paper. The medical data are proprietary and confidential. Researchers at some institutions, in particular medical schools, may have access to the IBM MarketScan Research Database from which the data is drawn following koenecke2021alpha.

Funding

Generous support from the Office of Naval Research through ONR grants N00014-17-1-2131 and N00014-19-1-2468 is gratefully acknowledged. Data for the health application were accessed using the Stanford Center for Population Health Sciences Data Core.

thebibliography\bibitem[\citeauthoryear{Alvarez, Chiann, and Morettin}{Alvarez et al.}{2022}]{alvarez2022inference} Alvarez, L., C. Chiann, and P. Morettin (2022). \newblock Inference in parametric models with many l-moments. \newblock {\em arXiv preprint arXiv:2210.04146\/}. \bibitem[\citeauthoryear{Alvarez and Biderman}{Alvarez and Biderman}{2022}]{alvarezsemiparametric} Alvarez, L. A. and C. Biderman (2022). \newblock Semiparametric analysis of randomised experiments using l-moments. \newblock {\em Working Paper\/}. \bibitem[\citeauthoryear{Batir}{Batir}{2017}]{Batir2017} Batir, N. (2017). \newblock Bounds for the gamma function. \newblock {\em Results in Mathematics\/} {\em 72}, 865--874. \bibitem[\citeauthoryear{Bickel}{Bickel}{1965}]{bickel1965some} Bickel, P. J. (1965). \newblock On some robust estimates of location. \newblock {\em The Annals of Mathematical Statistics\/} {\em 36\/}(3), 847--858. \bibitem[\citeauthoryear{Bickel}{Bickel}{1967}]{bickel1967some} Bickel, P. J. (1967). \newblock Some contributions to the theory of order statistics. \newblock In {\em Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics}, pp.\ 575--591. University of California Press. \bibitem[\citeauthoryear{Bickel}{Bickel}{1982}]{bickel1982adaptive} Bickel, P. J. (1982). \newblock On adaptive estimation. \newblock {\em The Annals of Statistics\/} {\em 10\/}(3), 647--671. \bibitem[\citeauthoryear{Bickel and Doksum}{Bickel and Doksum}{2015}]{bickel2015mathematical} Bickel, P. J. and K. A. Doksum (2015). \newblock {\em Mathematical statistics: basic ideas and selected topics}, Volume 2. \newblock CRC Press. \bibitem[\citeauthoryear{Bickel, G{\"o}tze, and van Zwet}{Bickel et al.}{2012}]{bickel2012resampling} Bickel, P. J., F. G{\"o}tze, and W. R. van Zwet (2012). \newblock Resampling fewer than n observations: gains, losses, and remedies for losses. \newblock In {\em Selected works of Willem van Zwet}, pp.\ 267--297. Springer. \bibitem[\citeauthoryear{Bickel, Klaassen, Ritov, and Wellner}{Bickel et al.}{1993}]{bickel1993efficient} Bickel, P. J., C. A. Klaassen, Y. Ritov, and J. A. Wellner (1993). \newblock {\em Efficient and adaptive estimation for semiparametric models}, Volume 4. \newblock Johns Hopkins University Press Baltimore. \bibitem[\citeauthoryear{Bickel and Lehmann}{Bickel and Lehmann}{1975a}]{bickel1975descriptiveI} Bickel, P. J. and E. L. Lehmann (1975a). \newblock Descriptive statistics for nonparametric models i. introduction. \newblock {\em The Annals of Statistics\/} {\em 3\/}(5), 1038--1044. \bibitem[\citeauthoryear{Bickel and Lehmann}{Bickel and Lehmann}{1975b}]{bickel1975descriptiveII} Bickel, P. J. and E. L. Lehmann (1975b). \newblock Descriptive statistics for nonparametric models ii. location. \newblock {\em The Annals of Statistics\/} {\em 3\/}(5), 1045--1069. \bibitem[\citeauthoryear{Bickel and Lehmann}{Bickel and Lehmann}{1976}]{bickel1976descriptive} Bickel, P. J. and E. L. Lehmann (1976). \newblock Descriptive statistics for nonparametric models. iii. dispersion. \newblock {\em Annals of Statistics\/} {\em 4\/}(6), 1139--1158. \bibitem[\citeauthoryear{Bickel and Lehmann}{Bickel and Lehmann}{2012}]{bickel2012descriptiveIV} Bickel, P. J. and E. L. Lehmann (2012). \newblock Descriptive statistics for nonparametric models iv. spread. \newblock In {\em Selected Works of EL Lehmann}, pp.\ 519--526. Springer. \bibitem[\citeauthoryear{Chen and Au}{Chen and Au}{2022}]{chen2022robust} Chen, A. and T. C. Au (2022). \newblock Robust causal inference for incremental return on ad spend with randomized paired geo experiments. \newblock {\em The Annals of Applied Statistics\/} {\em 16\/}(1), 1--20. \bibitem[\citeauthoryear{Chernoff, Gastwirth, and Johns}{Chernoff et al.}{1967}]{chernoff1967asymptotic} Chernoff, H., J. L. Gastwirth, and M. V. Johns (1967). \newblock Asymptotic distribution of linear combinations of functions of order statistics with applications to estimation. \newblock {\em The Annals of Mathematical Statistics\/} {\em 38\/}(1), 52--72. \bibitem[\citeauthoryear{Crump, Hotz, Imbens, and Mitnik}{Crump et al.}{2009}]{crump2009dealing} Crump, R. K., V. J. Hotz, G. W. Imbens, and O. A. Mitnik (2009). \newblock Dealing with limited overlap in estimation of average treatment effects. \newblock {\em Biometrika\/} {\em 96\/}(1), 187--199. \bibitem[\citeauthoryear{Csorgo and Revesz}{Csorgo and Revesz}{1978}]{Csorgo-Revesz-1978} Csorgo, M. and P. Revesz (1978). \newblock Strong approximations of the quantile process. \newblock {\em The Annals of Statistics\/} {\em 6\/}(4), 882--894. \bibitem[\citeauthoryear{Cuzick}{Cuzick}{1992a}]{cuzick1992efficient} Cuzick, J. (1992a). \newblock Efficient estimates in semiparametric additive regression models with unknown error distribution. \newblock {\em The Annals of Statistics\/} {\em 20\/}(2), 1129--1136. \bibitem[\citeauthoryear{Cuzick}{Cuzick}{1992b}]{cuzick1992semiparametric} Cuzick, J. (1992b). \newblock Semiparametric additive regression. \newblock {\em Journal of the Royal Statistical Society. Series B (Methodological)\/} {\em 45\/}(3), 831--843. \bibitem[\citeauthoryear{Doksum}{Doksum}{1974}]{doksum1974empirical} Doksum, K. (1974). \newblock Empirical probability plots and statistical inference for nonlinear models in the two-sample case. \newblock {\em The annals of statistics\/} {\em 2\/}(2), 267--277. \bibitem[\citeauthoryear{Doksum and Sievers}{Doksum and Sievers}{1976}]{doksum1976plotting} Doksum, K. A. and G. L. Sievers (1976). \newblock Plotting with confidence: Graphical comparisons of two populations. \newblock {\em Biometrika\/} {\em 63\/}(3), 421--434. \bibitem[\citeauthoryear{Firpo}{Firpo}{2007}]{firpo2007efficient} Firpo, S. (2007). \newblock Efficient semiparametric estimation of quantile treatment effects. \newblock {\em Econometrica\/} {\em 75\/}(1), 259--276. \bibitem[\citeauthoryear{Firpo, Fortin, and Lemieux}{Firpo et al.}{2009}]{firpo2009unconditional} Firpo, S., N. M. Fortin, and T. Lemieux (2009). \newblock Unconditional quantile regressions. \newblock {\em Econometrica\/} {\em 77\/}(3), 953--973. \bibitem[\citeauthoryear{Fisher}{Fisher}{1937}]{fisher1937design} Fisher, R. A. (1937). \newblock {\em The design of experiments}. \newblock Oliver And Boyd; Edinburgh; London. \bibitem[\citeauthoryear{Govindarajulu, Le Cam, and Raghavachari}{Govindarajulu et al.}{1967}]{govindarajulu1967generalizations} Govindarajulu, Z., L. Le Cam, and M. Raghavachari (1967). \newblock Generalizations of theorems of {Chernoff} and {Savage} on the asymptotic normality of test statistics. \newblock In {\em Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics}, pp.\ 609--638. University of California Press. \bibitem[\citeauthoryear{Gupta, Kohavi, Tang, Xu, Andersen, Bakshy, Cardin, Chandran, Chen, Coey, et al.}{Gupta et al.}{2019}]{gupta2019top} Gupta, S., R. Kohavi, D. Tang, Y. Xu, R. Andersen, E. Bakshy, N. Cardin, S. Chandran, N. Chen, D. Coey, et al. (2019). \newblock Top challenges from the first practical online controlled experiments summit. \newblock {\em ACM SIGKDD Explorations Newsletter\/} {\em 21\/}(1), 20--35. \bibitem[\citeauthoryear{Hampel, Ronchetti, Rousseeuw, and Stahel}{Hampel et al.}{2011}]{hampel2011robust} Hampel, F. R., E. M. Ronchetti, P. J. Rousseeuw, and W. A. Stahel (2011). \newblock {\em Robust statistics: the approach based on influence functions}, Volume 196. \newblock John Wiley & Sons. \bibitem[\citeauthoryear{Hodges Jr and Lehmann}{Hodges Jr and Lehmann}{1963}]{hodges1963estimates} Hodges Jr, J. L. and E. L. Lehmann (1963). \newblock Estimates of location based on rank tests. \newblock {\em The Annals of Mathematical Statistics\/} {\em 34\/}(2), 598--611. \bibitem[\citeauthoryear{Huber}{Huber}{1964}]{huber1964} Huber, P. J. (1964). \newblock Robust estimation of a location parameter. \newblock {\em The Annals of Mathematical Statistics\/} {\em 35\/}(1), 73--101. \bibitem[\citeauthoryear{Huber}{Huber}{2011}]{huber2011robust} Huber, P. J. (2011). \newblock {\em Robust statistics}. \newblock Springer. \bibitem[\citeauthoryear{Imbens and Rubin}{Imbens and Rubin}{2015}]{imbens2015causal} Imbens, G. W. and D. B. Rubin (2015). \newblock {\em Causal Inference in Statistics, Social, and Biomedical Sciences}. \newblock Cambridge University Press. \bibitem[\citeauthoryear{Jaeckel}{Jaeckel}{1971a}]{jaeckel1971robust} Jaeckel, L. A. (1971a). \newblock Robust estimates of location: Symmetry and asymmetric contamination. \newblock {\em The Annals of Mathematical Statistics\/} {\em 42\/}(3), 1020--1034. \bibitem[\citeauthoryear{Jaeckel}{Jaeckel}{1971b}]{jaeckel1971some} Jaeckel, L. A. (1971b). \newblock Some flexible estimates of location. \newblock {\em The Annals of Mathematical Statistics\/} {\em 45\/}(2), 1540--1552. \bibitem[\citeauthoryear{Klaassen}{Klaassen}{1987}]{klaassen1987consistent} Klaassen, C. A. (1987). \newblock Consistent estimation of the influence function of locally asymptotically linear estimators. \newblock {\em The Annals of Statistics\/} {\em 15\/}(4), 1548--1562. \bibitem[\citeauthoryear{Koenecke, Powell, Xiong, Shen, Fischer, Huq, Khalafallah, Trevisan, Sparen, Carrero, Nishimura, Caffo, Stuart, Bai, Staedtke, Thomas, Papadopoulos, Kinzler, Vogelstein, Zhou, Bettegowda, Konig, Mensh, Vogelstein, and Athey}{Koenecke et al.}{2021}]{koenecke2021alpha} Koenecke, A., M. Powell, R. Xiong, Z. Shen, N. Fischer, S. Huq, A. M. Khalafallah, M. Trevisan, P. Sparen, J. J. Carrero, A. Nishimura, B. Caffo, E. A. Stuart, R. Bai, V. Staedtke, D. L. Thomas, N. Papadopoulos, K. W. Kinzler, B. Vogelstein, S. Zhou, C. Bettegowda, M. F. Konig, B. D. Mensh, J. T. Vogelstein, and S. Athey (2021). \newblock Alpha-1 adrenergic receptor antagonists to prevent hyperinflammation and death from lower respiratory tract infection. \newblock {\em eLife\/} {\em 10}, e61700. \bibitem[\citeauthoryear{Kohavi, Tang, and Xu}{Kohavi et al.}{2020}]{kohavi2020trustworthy} Kohavi, R., D. Tang, and Y. Xu (2020). \newblock {\em Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing}. \newblock Cambridge University Press. \bibitem[\citeauthoryear{Le Cam and Yang}{Le Cam and Yang}{1988}]{lecam-yang-1988} Le Cam, L. and G. L. Yang (1988). \newblock On the preservation of local asyptotic normality under information loss. \newblock {\em The Annals of Statistics\/} {\em 16\/}(2), 483--520. \bibitem[\citeauthoryear{Lehmann and D'Abrera}{Lehmann and D'Abrera}{1975}]{lehmann1975nonparametrics} Lehmann, E. L. and H. J. D'Abrera (1975). \newblock {\em Nonparametrics: statistical methods based on ranks.} \newblock Holden-day. \bibitem[\citeauthoryear{Lewis and Rao}{Lewis and Rao}{2015}]{lewis2015unfavorable} Lewis, R. A. and J. M. Rao (2015). \newblock The unfavorable economics of measuring the returns to advertising. \newblock {\em The Quarterly Journal of Economics\/} {\em 130\/}(4), 1941--1973. \bibitem[\citeauthoryear{Li, Morgan, and Zaslavsky}{Li et al.}{2018}]{li2018balancing} Li, F., K. L. Morgan, and A. M. Zaslavsky (2018). \newblock Balancing covariates via propensity score weighting. \newblock {\em Journal of the American Statistical Association\/} {\em 113\/}(521), 390--400. \bibitem[\citeauthoryear{Linden and Rockoff}{Linden and Rockoff}{2008}]{linden2008estimates} Linden, L. and J. E. Rockoff (2008). \newblock Estimates of the impact of crime risk on property values from megan's laws. \newblock {\em American Economic Review\/} {\em 98\/}(3), 1103--27. \bibitem[\citeauthoryear{Linden and Rockoff}{Linden and Rockoff}{2019}]{linden2019replication} Linden, L. and J. E. Rockoff (2019). \newblock Replication data for: Estimates of the impact of crime risk on property values from megan's laws. \newblock {\em American Economic Association [publisher], Inter-university Consortium for Political and Social Research [distributor]\/}. \newblock \url{https://doi.org/10.3886/E113243V1}. \bibitem[\citeauthoryear{Neyman}{Neyman}{1990}]{neyman1923} Neyman, J. (1923/1990). \newblock On the application of probability theory to agricultural experiments. essay on principles. section 9. \newblock {\em Statistical Science\/} {\em 5\/}(4), 465--472. \bibitem[\citeauthoryear{Pinkse and Schurter}{Pinkse and Schurter}{2021}]{pinkse2021estimates} Pinkse, J. and K. Schurter (2021). \newblock Estimates of derivatives of (log) densities and related objects. \newblock {\em Econometric Theory\/}, 1--36. \bibitem[\citeauthoryear{Robins}{Robins}{1986}]{robins1986new} Robins, J. (1986). \newblock A new approach to causal inference in mortality studies with a sustained exposure period application to control of the healthy worker survivor effect. \newblock {\em Mathematical modelling\/} {\em 7\/}(9-12), 1393--1512. \bibitem[\citeauthoryear{Rosenbaum}{Rosenbaum}{1993}]{rosenbaum1993hodges} Rosenbaum, P. R. (1993). \newblock Hodges-lehmann point estimates of treatment effect in observational studies. \newblock {\em Journal of the American Statistical Association\/} {\em 88\/}(424), 1250--1253. \bibitem[\citeauthoryear{Rubin}{Rubin}{1974}]{rubin1974estimating} Rubin, D. B. (1974). \newblock Estimating causal effects of treatments in randomized and nonrandomized studies. \newblock {\em Journal of educational Psychology\/} {\em 66\/}(5), 688. \bibitem[\citeauthoryear{Stigler}{Stigler}{1974}]{stigler1974linear} Stigler, S. M. (1974). \newblock Linear functions of order statistics with smooth weight functions. \newblock {\em The Annals of Statistics\/} {\em 2\/}(4), 676--693. \bibitem[\citeauthoryear{Stone}{Stone}{1975}]{stone1975adaptive} Stone, C. J. (1975). \newblock Adaptive maximum likelihood estimators of a location parameter. \newblock {\em The Annals of Statistics\/} {\em 3\/}(2), 267--284. \bibitem[\citeauthoryear{Taddy, Lopes, and Gardner}{Taddy et al.}{2016}]{taddy2016scalable} Taddy, M., H. F. Lopes, and M. Gardner (2016). \newblock Scalable semiparametric inference for the means of heavy-tailed distributions. \newblock {\em arXiv preprint arXiv:1602.08066\/}. \bibitem[\citeauthoryear{Tripuraneni, Madeka, Foster, Perrault-Joncas, and Jordan}{Tripuraneni et al.}{2021}]{tripuraneni2021meta} Tripuraneni, N., D. Madeka, D. Foster, D. Perrault-Joncas, and M. I. Jordan (2021). \newblock Meta-analysis of randomized experiments with applications to heavy-tailed response data. \newblock {\em arXiv e-prints\/}, arXiv--2112. \bibitem[\citeauthoryear{Vansteelandt and Dukes}{Vansteelandt and Dukes}{2022}]{vansteelandt2022assumption} Vansteelandt, S. and O. Dukes (2022). \newblock Assumption-lean inference for generalised linear model parameters. \newblock {\em Journal of the Royal Statistical Society Series B: Statistical Methodology\/} {\em 84\/}(3), 657--685. \bibitem[\citeauthoryear{Vansteelandt and Joffe}{Vansteelandt and Joffe}{2014}]{vansteelandt2014structural} Vansteelandt, S. and M. Joffe (2014). \newblock Structural nested models and g-estimation: the partially realized promise. \newblock {\em Statistical Science\/} {\em 29\/}(4), 707--731. \bibitem[\citeauthoryear{Wager and Athey}{Wager and Athey}{2018}]{wager2018estimation} Wager, S. and S. Athey (2018). \newblock Estimation and inference of heterogeneous treatment effects using random forests. \newblock {\em Journal of the American Statistical Association\/} {\em 113\/}(523), 1228--1242. \bibitem[\citeauthoryear{Wu and Hamada}{Wu and Hamada}{2011}]{wu2011experiments} Wu, C. J. and M. S. Hamada (2011). \newblock {\em Experiments: planning, analysis, and optimization}, Volume 552. \newblock John Wiley & Sons. \bibitem[\citeauthoryear{Yu and Yao}{Yu and Yao}{2017}]{yu2017robust} Yu, C. and W. Yao (2017). \newblock Robust linear regression: A review and comparison. \newblock {\em Communications in Statistics-Simulation and Computation\/} {\em 46\/}(8), 6261--6282.

List of Figure Legends

\paragraph{Figure 1.} Example members of the generalized Huber family of distributions.

\paragraph{Figure 2.} The efficient weight function for the Cauchy distribution. The weights, normalized to have mean 1, are plotted against the quantile (left) and against the value of the observations (right). For the figure on the right, we only plot the range from \(-10\) to \(10\), which corresponds to approximately the \(0.03\) to \(0.97\) quantile. At this point, the weight is approximately \(-0.04\), and weights for more extreme quantiles are closer to 0.

\paragraph{Figure 3.} Histogram of house prices, in levels and logs, as well as estimated optimal weights (minus the second derivative of the log density) based on all 170,239 observations. The weights are normalized to be mean 1 (black horizontal line). Some weights are below 0 (blue horizontal line). Vertical lines indicate the \(0.0001\), \(0.001\), \(0.01\), \(0.1\), \(0.9\), \(0.99\), \(0.999\), \(0.9999\) quantiles.

\paragraph{Figure 4.} Root mean squared error and coverage of 95% confidence intervals relative to the average treatment effect (fixed at 0.1 standard deviations of the outcome in the absence of treatment) in simulations with heterogeneous treatment effects in the house price data. The horizontal axis varies the amount of heterogeneity. When \(q=0\), there is no heterogeneity such that the constant additive treatment effect model is correctly specified.

\paragraph{Figure 5.} Histogram of medical expenditures per admission, in levels and logs, as well as estimated optimal weights (minus the second derivative of the log density), based on the 98,155 observations in the control group. For the figure in levels, vertical lines indicate the \(0.001\), \(0.01\), \(0.1\), and \(0.9\) quantiles; the figure is limited to below \$200,000, such that the \(0.99\) and higher quantiles do not appear. For the figure in logs, vertical lines indicate the \(0.001\), \(0.01\), \(0.1\), \(0.99\), \(0.999\), \(0.9999\) quantiles. The weights are normalized to be mean 1 (black horizontal line). Some weights are below 0 (blue horizontal line).

\parindent 0cm\setcounter{equation}{0} {A.\arabic{equation}} \setcounter{lemma}{0}{A.\arabic{lemma}} \setcounter{theorem}{0}{A.\arabic{theorem}}

\setcounter{page}{1}\setcounter{section}{0}

\centerline{\sc Supplementary Materials for:}

\vskip0.5cm

\centerline{\sc Semiparametric Estimation of Treatment Effects } \vskip0.5cm

\centerline{\sc in Randomized Experiments}

\vskip1cm

\centerline{\sc Athey, Bickel, Chen, Imbens & Pollmann}

\vskip0.5cm

\centerline{\sc Appendix}

We first provide a general theorem on linear combination of order statistics, which simplifies the results of chernoff1967asymptotic, see also bickel1967some, govindarajulu1967generalizations and stigler1974linear, and then provide the technical details for proving the main results.

A general theorem on linear combination of order statistics

Let $F$ be a distribution function, twice differentiable, and $f=F'$. Let $W:[0,1]\rightarrow\mathcal{R}$ be a function s.t. $W(0)=0$ and $W(1)=1$. The C-R condition was first proposed by Csorgo-Revesz-1978.

description• There exists $C_{0}>0$ and $\epsilon\in(0,\frac{1}{2})$ such that
enumerate$F(x)(1-F(x))\frac{|f'(x)|}{f^{2}(x)}\leq C_{0}$ if $F(x)\in(0,\epsilon)$ or $F(x)\in(1-\epsilon,1)$, and $C_{0}$ can be replaced by another constant if $F(x)\in[\epsilon,1-\epsilon]$. • $f'(x)\geq0$ for $x<F^{-1}(\epsilon)$ and $f'(x)\leq0$ for $x>F^{-1}(1-\epsilon)$.
description• There exists $\epsilon_{n}=o(1)$ such that \begin{align} \int_{0}^{\epsilon_{n}}(1+t^{1-C_{0}})|dW(t)| & =o(n^{-1/2}) and \int_{1-\epsilon_{n}}^{1}(1+(1-t)^{1-C_{0}})|dW(t)|=o(n^{-1/2})\\ \int_{\epsilon_{n}}^{1-\epsilon_{n}}\frac{1}{f(F^{-1}(t))}dW(t) & =O(\frac{n^{1/4}}{\log n}). \end{align}
theoremLet $X_{1},\cdots,X_{n}$ be i.i.d. from $F$, and let $\hat{F}$ and $\hat{F}^{-1}$ be the empirical distribution function and empirical quantile function respectively. Assume that the C-R condition and the above W condition hold and let \begin{align} \psi(x) & =-\int_{0}^{1}\frac{1(F(x)\leq t)-t}{f(F^{-1}(t))}dW(t). \end{align} Then \begin{align*} \sqrt{n}\int_{0}^{1}(\hat{F}^{-1}(t)-F^{-1}(t))dW(t) & =\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\psi(X_{i})+o_{P}(1) \end{align*} which converges in distribution to $\mathcal{N}(0, \sigma_f^2)$ if further $E\psi^2(X_1)<\infty$, where \begin{align} \sigma_f^2 = \int_0^1\int_0^1 \frac{\min(s,t) - st}{f(F^{-1}(s))f(F^{-1}(t))}dW(s)dW(t). \end{align} When $\psi'$ exists, we may also represent $W$ as below if it is well defined: \begin{align} W(t) & =\frac{\int_{0}^{t}\psi'(F^{-1}(u))du}{\int_{0}^{1}\psi'(F^{-1}(u))du}. \end{align}
remarkJustification of the conditions for a few common scenarios.
enumerate• Gaussian: $f(x)=\frac{1}{\sqrt{2\pi}}e^{-x^{2}/2}$, then $\epsilon=0.5$ and $C_{0}=1$ meet the C-R condition. • Cauchy: $f(x)=\frac{1}{\pi(1+x^{2})}$, then $\epsilon=0.5$ and $C_{0}=2$ meet the C-R condition. • Equal-weight: $W'(t)=1$, then (ref) implies $\epsilon_{n}=o(n^{-1/2})$ for $C_{0}\leq1$ and $\epsilon_{n}=o(n^{-\frac{1}{2(2-C_{0})}})$ for $C_{0}\in(1,2)$. So the equal weight applies to Gaussian. This does not apply to Cauchy, but note that its sample mean does not converge. • Median: $W$ puts all mass at $t=\frac{1}{2}$. The W condition holds for any $f$.

A lemma on the C-R condition

lemmaIf the C-R condition holds, then for any $t\in(0,\epsilon)$ \begin{align} \frac{1}{f(F^{-1}(t))} & \leq K_{0}t^{-C_{0}}\\ |F^{-1}(t)| & \leq K_{1}t^{1-C_{0}}+K_{2} \end{align} where $K_{0},K_{1}$ and $K_{2}$ are positive constants, only dependent on $C_{0}$ and $\epsilon$. Similar results hold for $t\in(1-\epsilon,1)$, omitted for conciseness.
proofBy the C-R condition, $f'(x)\geq0$ for any $x<F^{-1}(\epsilon)$, and for any $t\in(0,\epsilon)$, \begin{align*} \int_{t}^{\epsilon}\frac{f'}{f^{2}}(F^{-1}(u))du & \leq\int_{t}^{\epsilon}\frac{C_{0}}{u(1-u)}du=C_{0}\log\frac{\epsilon}{1-\epsilon}-C_{0}\log\frac{t}{1-t}\leq C-C_{0}\log t \end{align*} where $C=C_{0}\log\frac{\epsilon}{1-\epsilon}$. Since \begin{align*} \frac{f'}{f^{2}}(F^{-1}(u)) & =\frac{\partial}{\partial u}\log f(F^{-1}(u)), \end{align*} then \begin{align*} \log\frac{f(F^{-1}(\epsilon))}{f(F^{-1}(t))} & \leq C-C_{0}\log t, \end{align*} which implies (ref) with $K_{0}=\frac{e^{C}}{f(F^{-1}(\epsilon))}$. Next, (ref) follows from the inequality below: \begin{align*} F^{-1}(\epsilon)-F^{-1}(t)=\int_{t}^{\epsilon}\frac{1}{f(F^{-1}(u))}du & \leq\int_{t}^{\epsilon}K_{0}u^{-C_{0}}du=\frac{K_{0}}{1-C_{0}}(\epsilon^{1-C_{0}}-t^{1-C_{0}}). \end{align*}

Proof of Theorem (ref)

proofFirst of all, \begin{align*} \int_{\epsilon_{n}}^{1-\epsilon_{n}}\big|(\hat{F}^{-1}(t)-F^{-1}(t))+\frac{\hat{F}(F^{-1}(t))-t}{f(F^{-1}(t))}\big|\cdot|dW(t)| & \leq\Delta_{n}\cdot\int_{\epsilon_{n}}^{1-\epsilon_{n}}\frac{|dW(t)|}{f(F^{-1}(t))} \end{align*} where \begin{align*} \Delta_{n} & =\sup_{0<t<1}\big|f(F^{-1}(t))(\hat{F}^{-1}(t)-F^{-1}(t))+(\hat{F}(F^{-1}(t))-t)\big|. \end{align*} Under the C-R condition, by Theorem 4 of Csorgo-Revesz-1978, we have \begin{align*} \Delta_{n} & =o_{P}(n^{-3/4}\log n). \end{align*} Under the condition (ref), \begin{align*} \int_{\epsilon_{n}}^{1-\epsilon_{n}}\frac{|dW(t)|}{f(F^{-1}(t))} & =O_{P}(\frac{n^{1/4}}{\log n}). \end{align*} Thus \begin{align*} \int_{\epsilon_{n}}^{1-\epsilon_{n}}(\hat{F}^{-1}(t)-F^{-1}(t))dW(t) & =-\int_{\epsilon_{n}}^{1-\epsilon_{n}}\frac{\hat{F}(F^{-1}(t))-t}{f(F^{-1}(t))}dW(t)+o_{P}(n^{-1/2}). \end{align*} Next, since \begin{align*} E\big|\int_{0}^{\epsilon_{n}}\frac{\hat{F}(F^{-1}(t))-t}{f(F^{-1}(t))}dW(t)\big| & \leq2\int_{0}^{\epsilon_{n}}\frac{t}{f(F^{-1}(t))}|dW(t)|\\ & \leq2K_{0}\int_{0}^{\epsilon_{n}}t^{1-C_{0}}|dW(t)|\\ & =o(n^{-1/2}) \end{align*} where the last inequality is due to Lemma (ref) and the last equality is due to (ref), then \begin{align*} \int_{0}^{\epsilon_{n}}\frac{\hat{F}(F^{-1}(t))-t}{f(F^{-1}(t))}dW(t) & =o_{P}(n^{-1/2}). \end{align*} Furthermore, by Lemma (ref), \begin{align*} \left|\int_{0}^{\epsilon_{n}}F^{-1}(t)dW(t)\right| & \leq K_{1}\int_{0}^{\epsilon_{n}}t^{1-C_{0}}|dW(t)|+K_{2}\int_{0}^{\epsilon_{n}}|dW(t)|=o(n^{-1/2}) \end{align*} where the last equality is due to (ref). Similarly, we have \begin{align*} \int_{1-\epsilon_{n}}^{1}\frac{\hat{F}(F^{-1}(t))-t}{f(F^{-1}(t))}dW(t) & =o_{P}(n^{-1/2}) \end{align*} and \begin{align*} \int_{1-\epsilon_{n}}^{1}F^{-1}(t)dW(t) & =o(n^{-1/2}). \end{align*} Hence, to prove that \begin{align*} \int_{0}^{1}(\hat{F}^{-1}(t)-F^{-1}(t))dW(t) & =-\int_{0}^{1}\frac{\hat{F}(F^{-1}(t))-t}{f(F^{-1}(t))}dW(t)+o_{P}(n^{-1/2}), \end{align*} it remains to show that \begin{align} \left(\int_{0}^{\epsilon_{n}}+\int_{1-\epsilon_{n}}^{1}\right)\hat{F}^{-1}(t)dW(t) & =o_{P}(n^{-1/2}). \end{align} Let $X_{(1)}\leq\cdots\leq X_{(n)}$ be the order statistics and $U_{(i)}\equiv F(X_{(i)})$, then $\hat{F}^{-1}(t)=F^{-1}(U_{(\lceil nt\rceil)})$. If $C_{0}>1$, then by Lemma (ref), \begin{align*} \left|\int_{0}^{\epsilon_{n}}\hat{F}^{-1}(t)dW(t)\right|=\left|\int_{0}^{\epsilon_{n}}F^{-1}(U_{(\lceil nt\rceil)})dW(t) \right| & \leq(K_{1}+K_{2})\int_{0}^{\epsilon_{n}}(U_{(\lceil nt\rceil)})^{1-C_{0}}|dW(t)|. \end{align*} Since $U_{(i)}$ follows the $\text{Beta}(i,n+1-i)$ distribution, one can verify that \begin{align*} E(U_{(i)}^{a}) & =\frac{\Gamma(n+1)\Gamma(i+a)}{\Gamma(n+1+a)\Gamma(i)}. \end{align*} Using the bounds for gamma functions Batir2017, i.e. for $x\geq1,$ \begin{align*} \sqrt{2\pi}x^{x}e^{-x}(x^{2}+\frac{x}{3}+0.0496)^{1/4} & <\Gamma(x+1)<\sqrt{2\pi}x^{x}e^{-x}(x^{2}+\frac{x}{3}+0.056)^{1/4}, \end{align*} one can get for $i\geq2-a$, \begin{align*} C_{1}\cdot(i/n)^{a}\leq\mathbb{E}U_{(i)}^{a} & \leq C_{2}\cdot(i/n)^{a} \end{align*} where $C_{1},C_{2}>0$ are some constants. For $1\leq i<2-a$, $EU_{(i)}^{a}=\frac{\Gamma(i+a)}{\Gamma(i)}n^{-a}(1+o(1))$. These imply that $EU_{(\lceil nt\rceil)})^{1-C_{0}}\leq C_{3}\cdot t^{1-C_{0}}$ for $C_{3}$ large enough, and hence \begin{align*} E\left(\left|\int_{0}^{\epsilon_{n}}\hat{F}^{-1}(t)dW(t)\right|\right) & \leq(K_{1}+K_{2})\int_{0}^{\epsilon_{n}}EU_{(\lceil nt\rceil)})^{1-C_{0}}\big|dW(t)\big|\\ & \leq C_{3}(K_{1}+K_{2})\int_{0}^{\epsilon_{n}}t^{1-C_{0}}|dW(t)|. \end{align*} If $0<C_{0}\leq1$, then \begin{align*} \mathbb{E}\left(\left|\int_{0}^{\epsilon_{n}}\hat{F}^{-1}(t)dW(t)\right|\right) & \leq(K_{1}+K_{2})\int_{0}^{\epsilon_{n}}|dW(t)|. \end{align*} Therefore, due to (ref), $\mathbb{E}\big|\int_{0}^{\epsilon_{n}}\hat{F}^{-1}(t)dW(t)\big|=o(n^{-1/2})$ for both $C_{0}>1$ and $C_{0}\in(0,1]$, which implies $\int_{0}^{\epsilon_{n}}\hat{F}^{-1}(t)dW(t)=o_{P}(n^{-1/2})$. Similarly, it can be shown that $\int_{1-\epsilon_{n}}^{1}\hat{F}^{-1}(t)dW(t)=o_P(n^{-1/2})$. Thus (ref) holds, which concludes the proof.

Proof of Theorem (ref)

Our proof is based on the sample splitting argument, see klaassen1987consistent.

General conditions for the density estimation

Let $X_{1},\cdots,X_{n}$ be i.i.d. $F$ with $f\equiv F'$ and $\hat{F}$ be the empirical distribution. Let $\hat{f}$ be an estimate of $f$ from data independent of $\{X_{1},\cdots,X_{n}\}$ such that $I(\hat{f})=I(f)+o_P(1)$ and the following conditions hold for any $\delta_{n}=O_{P}(n^{-1/2})$:

align[align omitted — 324 chars of source]
remarkThese conditions are satisfied by the kernel density estimator under mild conditions, see stone1975adaptive and bickel1982adaptive.

Let

align*[align* omitted — 78 chars of source]

where the dependency on $F$ is suppressed for notational simplicity.

Let $\hat{f}$ and $\hat{F}$ be the estimates of $f$ and $F$ respectively, and let

align*[align* omitted — 160 chars of source]

be the estimate of $w_{f}$. Then $\int_0^1 w_{\hat{f}}(u)du = 1$.

We assume that the following conditions hold for $w_{\hat{f}}$:

align[align omitted — 124 chars of source]
align[align omitted — 179 chars of source]
remarkConditions (ref) can be achieved by kernel density estimation with proper tail truncation in the spirit of Lemma 6.1 in bickel1982adaptive.
remarkIf $X\sim F$ and $var(X)<\infty$, then $\int_{0}^{1}\int_{0}^{1}\frac{\min(u,v)-uv}{f(F^{-1}(u))f(F^{-1}(v))}dudv=var(X)$. See Bickel (1967). This suggests that the requirement (ref) is not strong.
lemmaLet $X_{1},\cdots,X_{n}$ be i.i.d. $F$ with $f\equiv F'$ and $\hat{F}$ the empirical distribution. Let $w_{\hat{f}}$ be an estimate of $w_{f}$ independent of $\{X_{1},\cdots,X_{n}\}$ and satisfy the above conditions (ref) and (ref). Then under the C-R condition, \begin{align*} \int_{0}^{1}(\hat{F}^{-1}(u)-F^{-1}(u))\cdot w_{\hat{f}}(u)du & =-\frac{1}{nI(f)}\sum_{i=1}^{n}\frac{f'}{f}(X_{i}) + o_{P}(n^{-1/2}). \end{align*}
proofNote that \begin{align*} & \left|\int_{0}^{1}\left(\hat{F}^{-1}(u)-F^{-1}(u)+\frac{\hat{F}(F^{-1}(u))-u}{f(F^{-1}(u))}\right)\cdot w_{\hat{f}}(u)du\right|\\ \leq & \sup_{0<u<1}\big|f(F^{-1}(u))\cdot(\hat{F}^{-1}(u)-F^{-1}(u))+(\hat{F}(F^{-1}(u))-u)\big|\cdot\int_{0}^{1}\frac{\big|w_{\hat{f}}(u)\big|}{f(F^{-1}(u))}du\\ = & o_{P}(n^{-3/4}\log n)\cdot\int_{0}^{1}\frac{\big|w_{\hat{f}}(u)\big|}{f(F^{-1}(u))}du \end{align*} where the last equality is due to Theorem 4 of Csorgo-Revesz-1978. Thus under the further condition (ref), we have \begin{align*} \int_{0}^{1}\left(\hat{F}^{-1}(u)-F^{-1}(u)+\frac{\hat{F}(F^{-1}(u))-u}{f(F^{-1}(u))}\right)\cdot w_{\hat{f}}(u)du & =o_{P}(n^{-1/2}). \end{align*} Note that \begin{align*} \int_{0}^{1}\frac{\hat{F}(F^{-1}(u))-u}{f(F^{-1}(u))}w_{f}(u)du & =-\frac{1}{nI(f)}\sum_{i=1}^{n}\frac{f'}{f}(X_{i}). \end{align*} Hence it is sufficient to prove that \begin{align} \int_{0}^{1}\frac{\hat{F}(F^{-1}(u))-u}{f(F^{-1}(u))}\cdot(w_{\hat{f}}(u)-w_{f}(u))du & =o_{P}(n^{-1/2}). \end{align} Recall that $w_{\hat{f}}$ is independent of $\hat{F}$, so \begin{align*} & E\left(\int_{0}^{1}\frac{\hat{F}(F^{-1}(u))-u}{f(F^{-1}(u))}\cdot(w_{\hat{f}}(u)-w_{f}(u))du\right)^{2}\\ = & E\int_{0}^{1}\int_{0}^{1}\frac{\hat{F}(F^{-1}(u))-u}{f(F^{-1}(u))}\cdot\frac{\hat{F}(F^{-1}(v))-v}{f(F^{-1}(v))}\cdot(w_{\hat{f}}-w_{f})(u)\cdot(w_{\hat{f}}-w_{f})(v)dudv\\ = & \frac{1}{n}\int_{0}^{1}\int_{0}^{1}\frac{\min(u,v)-uv}{f(F^{-1}(u))f(F^{-1}(v))}\cdot E\left((w_{\hat{f}}(u)-w_{f}(u))(w_{\hat{f}}(v)-w_{f}(v))\right)dudv\\ \leq & \frac{1}{n}\int_{0}^{1}\int_{0}^{1}\frac{\min(u,v)-uv}{f(F^{-1}(u))f(F^{-1}(v))}\cdot E\left((w_{\hat{f}}(u)-w_{f}(u))^{2}\right)dudv \end{align*} which is $o(\frac{1}{n})$ due to (ref). Therefore (ref) holds, which concludes the proof.

Proof of Theorem (ref){\it (i)}

The median difference, i.e. $\hat{F}_{1}^{_{-1}}(\frac{1}{2})-\hat{F}_{0}^{-1}(\frac{1}{2})$, is a $\sqrt{n}$ consistent estimate of $\tau$, where $\hat{F}_{1}^{-1}$ and $\hat{F}_{0}^{-1}$ are the empirical quantile functions for $Y$ conditional on $Z=1$ and $Z=0$, respectively. This follows directly from Theorem A.1.

Proof of Theorem (ref){\it (ii)}

Let $\mathcal{D}_{1}=\{(Z_{i},Y_{i}):1\leq i\leq\frac{n}{2}\}$ and $\mathcal{D}_{2}=\{(Z_{i},Y_{i}):\frac{n}{2}+1\leq i\leq n\}$ be a sample split. For simplicity, assume that $\sum_{i=1}^{n/2}Z_{i}=\frac{1}{2}np$ and $\sum_{i=\frac{n}{2}+1}^{n}Z_{i}=\frac{1}{2}np$. Let $\hat{f}_{0(j)}$ be the estimate of $f_{0}\equiv F'_{0}$ based on $\mathcal{D}_{j}$, $j\in\{1,2\}$, which satisfy the conditions (ref) and (ref).

By the first order Taylor expansion at $\tau$, there exists $\overline{\tau}$ between $\tilde{\tau}$ and $\tau$ such that

align*[align* omitted — 255 chars of source]

where

align*[align* omitted — 403 chars of source]

Since $\hat{f}_{0(2)}$ is independent of $\{(Z_{i},Y_{i}):1\leq i\leq\frac{n}{2}\}$, then

align*[align* omitted — 607 chars of source]

which is $o_{P}(1)$ according to (ref). Thus

align*[align* omitted — 198 chars of source]

Furthermore, $I(\hat{f}_{0(2)})=I(f_{0})+o_{P}(1)$, hence

align*[align* omitted — 165 chars of source]

Similarly, we have

align*[align* omitted — 169 chars of source]

Therefore,

align*[align* omitted — 187 chars of source]

From (ref),

align*[align* omitted — 145 chars of source]

By Slutsky's theorem,

align*[align* omitted — 170 chars of source]

since $\sqrt{n}(\tilde{\tau}-\tau)=O_{P}(1)$ and $1+B=o_{P}(1)$. This concludes the proof.

Proof of Theorem (ref){\it (iii)}

Let $\mathcal{D}_j$, $j\in\{1,2\}$ be the sample split as above. Let $\hat{F}_{1(j)}$ and $\hat{F}_{0(j)}$ be the empirical distributions of $Y|Z=1$ and $Y|Z=0$ respectively based on $\mathcal{D}_{j}$, for $j\in\{1,2\}$. Let $w_{\hat{f}_{j}}$ be as defined earlier and assume that $w_{\hat{f}_{j}}$ satisfy the conditions of Lemma (ref).

Then the $L$ estimate of $\tau$ can be written as:

align[align omitted — 274 chars of source]

According to Lemma (ref),

align*[align* omitted — 365 chars of source]

where the second equality uses the fact that $F_{1}(\cdot)\equiv F_{0}(\cdot-\tau)$ and $F_{1}^{-1}(u)\equiv F_{0}^{-1}(u)+\tau$. Thus,

align*[align* omitted — 279 chars of source]

A similar approximation holds for the other term, then by Slutsky's theorem,

align*[align* omitted — 211 chars of source]

Hence $\sqrt{n}(\hat{\tau}^{waq}-\tau) \stackrel{d}{\Rightarrow} \mathcal{N}\Bigl(0,\frac{1}{p(1-p)I(f_0)}\Bigr)$.

Proof of partial adaptation with \texorpdfstring{$(\alpha,\beta)$}{(alpha,beta)}-trimmed mean

Main results

theoremSuppose that $f_0$ falls into the extended Huber family as defined in Section (ref). If $\alpha$ and $\beta$ are chosen properly, then for the constant additive treatment effect model, $\hat{\tau}_{\alpha,\beta}$ as defined in ((ref)) is an efficient estimate of $\tau$.
proofAccording to Theorem 1, the information bound is given by $\frac{1}{p(1-p)}I^{-1}$. According to Theorem (ref), we have \begin{align*} \sqrt{n}(\hat{\tau}_{\alpha,\beta}-\tau) \stackrel{d}{\Rightarrow} \mathcal{N}\Bigl(0, p^{-1}\sigma_{f_1}^2 + (1-p)^{-1}\sigma_{f_0}^2\Bigr), \end{align*} where $\sigma_{f_j}^2$ is defined by (ref) for $j\in \{0, 1\}$, i.e. \begin{align} \sigma_{f_j}^2 = \frac{1}{(1-\alpha-\beta)^2}\int_{\alpha}^{\beta}\int_{\alpha}^{\beta} \frac{\min(s,t) - st}{f_j(F_j^{-1}(s))f_j(F_j^{-1}(t))}dsdt. \end{align} Since $F_{1}^{-1}(t)-F_{0}^{-1}(t)\equiv\tau$, then $\sigma_{f_0}^2 = \sigma_{f_1}^2$. Therefore it is sufficient to prove that $\sigma_{f_0}^2 = I^{-1}$. For notational simplicity, below we restrict the proof to $\mu=0$ and $\sigma=1$, that is $f_{0}$ is parameterized by $(k_{1},k_{2})$ as (ref) in Section 2.4 and $f_{1}(x)\equiv f_{0}(x-\tau)$. Let $\alpha=F_{0}(-k_{1})$ and $1-\beta=F_{0}(k_{2})$. Then $I=-\int(\frac{f_0'}{f_0})^{'}(x)dF_{0}(x)=\int_{-k_{1}}^{k_{2}}dF_{0}(x)=1-(\alpha+\beta).$ First of all, with integral change of variable, \begin{align*} \sigma_{f_0}^2 = & \frac{1}{(1-\alpha-\beta)^{2}}\int_{-k_{1}}^{k_{2}}\int_{-k_{1}}^{k_{2}}(F_{0}(\min(x,y))-F_{0}(x)F_{0}(y))dxdy. \end{align*} Furthermore, it is not hard to verify that $F_{0}(-k_{1})=f_{0}(-k_{1})/k_{1}$ and $1-F_{0}(k_{2})=f_{0}(k_{2})/k_{2}$. Using the fact that $f_{0}'(x)=-xf_{0}(x)$ for $x\in(-k_{1},k_{2})$, then \begin{align*} \int_{-k_{1}}^{k_{2}}F_{0}(x)dx & =xF_{0}(x)\big|_{-k_{1}}^{k_{2}}-\int_{-k_{1}}^{k_{2}}xf_{0}(x)dx\\ & =k_{2}\cdot(1-\beta)+k_{1}\alpha+\int_{-k_{1}}^{k_{2}}df_{0}(x)\\ & =k_{2}\cdot(1-\beta)+k_{1}\alpha+f_{0}(k_{2})-f_{0}(-k_{1})\\ & =k_{2}\cdot(1-\beta)+k_{1}\alpha+k_{2}\beta-k_{1}\alpha\\ & =k_{2} \end{align*} and \begin{align*} & \int_{-k_{1}}^{k_{2}}\int_{-k_{1}}^{k_{2}}(F_{0}(\min(x,y))dxdy\\ = & 2\int_{-k_{1}}^{k_{2}}\int_{-k_{1}}^{k_{2}}F_{0}(x)\cdot1(x\leq y)dxdy\\ = & 2\int_{-k_{1}}^{k_{2}}F_{0}(x)\cdot(k_{2}-x)dx\\ = & -\int_{-k_{1}}^{k_{2}}F(x)d(x-k_{2})^{2}\\ = & \int_{-k_{1}}^{k_{2}}(x-k_{2})^{2}f(x)dx+(k_{2}+k_{1})^{2}\alpha\\ = & \int_{-k_{1}}^{k_{2}}\left(k_{2}^{2}f(x)+2k_{2}f'(x)-xf'(x)\right)dx+(k_{2}+k_{1})^{2}\alpha\\ = & k_{2}^{2}(1-\alpha-\beta)+2k_{2}(k_{2}\beta-k_{1}\alpha)+(k_{1}+k_{2})^{2}\alpha-\int_{-k_{1}}^{k_{2}}xf'(x)dx\\ = & k_{2}^{2}+\alpha k_{1}^{2}+\beta k_{2}^{2}-\int_{-k_{1}}^{k_{2}}xdf(x)\\ = & k_{2}^{2}+\int_{-k_{1}}^{k_{2}}f(x)dx\\ = & k_{2}^{2}+1-(\alpha+\beta). \end{align*} Thus \begin{align*} \sigma_{f_0}^2 & =\frac{1}{(1-(\alpha+\beta))^{2}}\left((k_{2}^{2}+1-(\alpha+\beta))-k_{2}^{2}\right)\\ & = I^{-1}. \end{align*} which completes the proof.
theoremLet \begin{align*} (\hat{\alpha},\hat{\beta}) & =\arg\min_{\alpha_{0}\leq\alpha,\beta\leq\alpha_{1}}(p^{-1}\hat{\sigma}_{1}^{2}(\alpha,\beta)+(1-p)^{-1}\hat{\sigma}_{0}^{2}(\alpha,\beta)) \end{align*} where $p^{-1}\hat{\sigma}_{1}^{2}(\alpha,\beta)+(1-p)^{-1}\hat{\sigma}_{0}^{2}(\alpha,\beta)$ is the estimate of the asymptotic variance as derived in (ref) and $\hat{\sigma}_j^2$ is the estimate of $\sigma_{f_j}^2$ as given in Proposition (ref). If $0<\alpha_{0},\alpha_{1}<1$ and $\alpha_{0}+\alpha_{1}<1$, and $\sigma_{1}^{2}(\alpha,\beta)$ has a unique solution w.r.t. $(\alpha,\beta)$, which falls into $(\alpha_{0},\alpha_{1})\times(\alpha_{0},\alpha_{1})$, then $(\hat{\alpha},\hat{\beta})$ is a consistent estimate of the optimal $(\alpha,\beta)$.
proofFrom Proposition (ref), uniform convergence below holds for $j\in\{0,1\}$: \begin{align*} \sup_{\alpha_{0}\leq(\alpha,\beta)\leq\alpha_{1}}|\hat{\sigma}_{j}^{2}(\alpha,\beta)-\sigma_{j}^{2}(\alpha,\beta)| & =o_{P}(1). \end{align*} The conclusion holds.
remarkBy Proposition (ref), the limiting objective function $p^{-1}\sigma_{1}^{2}(\alpha,\beta)+(1-p)^{-1}\sigma_{0}^{2}(\alpha,\beta)$ is strictly convex at the local neighborhood of the optimal $(\alpha,\beta)$.

Supporting results

lemmaThe influence function (ref), specialized for the $(\alpha,\beta)$-trimmed mean, can be simplified to \begin{align*} \psi(x) = & \frac{1}{1-\alpha-\beta}(x - \mu(\alpha,\beta)), & if x \in (F^{-1}(\alpha), F^{-1}(1-\beta)) \\ = & \frac{1}{1-\alpha-\beta}(F^{-1}(1-\beta) - \mu(\alpha,\beta)), & if x \geq F^{-1}(1-\beta) \\ = & \frac{1}{1-\alpha-\beta}(F^{-1}(\alpha) - \mu(\alpha,\beta)), & if x \leq F^{-1}(\alpha) \end{align*} where \begin{align*} \mu(\alpha,\beta) & =\int_{\alpha}^{1-\beta}F^{-1}(u)du + \alpha F^{-1}(\alpha) + \beta F^{-1}(1-\beta). \end{align*} The asymptotic variance $\sigma_f^2$, denoted as $\sigma^2(\alpha,\beta)$ for the $(\alpha,\beta)$-trimmed mean, can be written as \begin{align*} \sigma^{2}(\alpha,\beta) & = \int_0^1 \psi^2(F^{-1}(u))du. \end{align*}
proofThe proof follows from direct calculation and is omitted.
propositionAssume that $f$ is continuous and that $f(x)>0$ for any $x\in \mathcal{R}$. Let $\hat{\mu}(\alpha, \beta)$ and $\hat{\sigma}^2(\alpha,\beta)$ be the estimate of $\mu(\alpha,\beta)$ and $\hat{\sigma}^2(\alpha,\beta)$ respectively as defined in Lemma (ref), by replacing $F$ with $\hat{F}$. Assume that $\sigma^{2}(\alpha,\beta)$ has a unique minimum w.r.t $(\alpha,\beta)$ which falls into $(\alpha_{0},\alpha_{1})\times(\alpha_{0},\alpha_{1})$, where $(\alpha_{0},\alpha_{1})\in(0,1)\times(0,1)$ with $\alpha_{0}+\alpha_{1}<1$. Let \begin{align*} (\hat{\alpha},\hat{\beta}) & =\arg\min_{0<\alpha_{0}\leq\alpha,\beta\leq\alpha_{1}<1}\hat{\sigma}^{2}(\alpha,\beta) \end{align*} be the estimate of the optimal $(\alpha,\beta)$. Then $(\hat{\alpha},\hat{\beta})$ is a consistent estimate of the optimal $(\alpha,\beta)$.
proofLet $\psi(x)$ be the same as in Lemma (ref), and let $\hat{\psi}$ be an estimate of $\psi$ by replacing $F$ with $\hat{F}$. For $u\in (0,1)$, define \begin{align*} e(u) = \psi(F^{-1}(u)) and \hat{e}(u) = \hat{\psi}(\hat{F}^{-1}(u)). \end{align*} Then \begin{align*} |\hat{\sigma}^{2}(\alpha,\beta)-\sigma^{2}| & =|\int_{0}^{1}\hat{e}^{2}(u)du-\int_{0}^{1}e^{2}(u)du|\\ & =|\int_{0}^{1}(\hat{e}-e)^{2}(u)du+2\int_{0}^{1}e(u)(\hat{e}(u)-e(u))du|\\ & \leq|\hat{e}-e|_{\infty}^{2}+2|e|_{\infty}\cdot|\hat{e}-e|_{\infty}, \end{align*} where \begin{align*} |e|_{\infty} & =\sup_{\alpha_{0}\leq u\leq1-\alpha_{1}}|e(u)|<\infty\\ |\hat{e}-e|_{\infty} & =\sup_{\alpha_{0}\leq u\leq1-\alpha_{1}}|\hat{e}(u)-e(u)|\\ & \leq \frac{2}{1-\alpha_0-\alpha_1}\sup_{\alpha_{0}\leq u\leq1-\alpha_{1}}|\hat{F}^{-1}(u)-F^{-1}(u)|\\ & =O_{P}(n^{-1/2}). \end{align*} Therefore the uniform convergence holds: \begin{align*} \sup_{0<\alpha_{0}\leq\alpha,\beta\leq\alpha_{1}<1}|\hat{\sigma}^{2}(\alpha,\beta)-\sigma^{2}(\alpha,\beta)| & =o_{P}(1). \end{align*} The conclusion follows since the optimal $(\alpha,\beta)$ is unique.
propositionSuppose that $f$ is from the extended Huber family as given by (ref) in Section 2.4. Then $\sigma^2(\alpha,\beta)$ is strictly convex at the local neighborhood of $(\alpha,\beta)=(F(-k_{1}),1-F(k_{2}))$.
proofNote that \begin{align*} (1-\alpha-\beta)^{2}\sigma^2(\alpha, \beta) & =\int_{\alpha}^{1-\beta}\int_{\alpha}^{1-\beta}\frac{\min(s,t)-st}{f(F^{-1}(s))f(F^{-1}(t))}dsdt\\ & \equiv RHS \end{align*} then \begin{align*} \frac{\partial RHS}{\partial\alpha} & =-2\int_{\alpha}^{1-\beta}\frac{\alpha(1-s)}{f(F^{-1}(s))f(F^{-1}(\alpha))}ds\\ & =-\frac{2\alpha}{f(F^{-1}(\alpha))}\int_{\alpha}^{1-\beta}\frac{1-s}{f(F^{-1}(s))}ds\\ \end{align*} and \begin{align*} \frac{\partial RHS}{\partial\beta} & =-\frac{2\beta}{f(F^{-1}(1-\beta))}\int_{\alpha}^{1-\beta}\frac{s}{f(F^{-1}(s))}ds.\\ \end{align*} Hence \begin{align*} \frac{\partial^{2}RHS}{\partial\alpha^{2}} & =\frac{2\alpha}{f(F^{-1}(\alpha))}\frac{1-\alpha}{f(F^{-1}(\alpha))}-2\left(\frac{1-\alpha\frac{f'}{f^{2}}(F^{-1}(\alpha))}{f(F^{-1}(\alpha))}\right)\int_{\alpha}^{1-\beta}\frac{1-s}{f(F^{-1}(s))}ds\\ & =\frac{2\alpha(1-\alpha)}{f^{2}(F^{-1}(\alpha))}-\frac{2(f-\alpha f'/f)}{f^{2}}(F^{-1}(\alpha))\cdot\int_{\alpha}^{1-\beta}\frac{1-s}{f(F^{-1}(s))}ds \end{align*} \begin{align*} \frac{\partial^{2}RHS}{\partial\beta^{2}} & =\frac{2\beta(1-\beta)}{f^{2}(F^{-1}(1-\beta))}-2\left(\frac{f+\beta f'/f}{f^{2}}\right)(F^{-1}(1-\beta))\cdot\int_{\alpha}^{1-\beta}\frac{s}{f(F^{-1}(s))}ds\\ \frac{\partial RHS}{\partial\alpha\partial\beta} & =\frac{2\alpha\beta}{f(F^{-1}(\alpha))f(F^{-1}(1-\beta))}. \end{align*} Now to evaluate the values at \begin{align*} \alpha & =F(-k_{1})=f(-k_{1})k_{1}^{-1}\\ \beta & =1-F(k_{2})=f(k_{2})k_{2}^{-1}, \end{align*} we have \begin{align*} \frac{f'}{f}(F^{-1}(\alpha)) & =k_{1}\\ \frac{f'}{f}(F^{-1}(1-\beta)) & =-k_{2}\\ \alpha k_{1} & =f(-k_{1})\\ \beta k_{2} & =f(k_{2}). \end{align*} Thus \begin{align*} \int_{\alpha}^{1-\beta}\frac{s}{f(F^{-1}(s))}ds & =\int_{-k_{1}}^{k_{2}}F(x)dx=xF(x)|_{-k_{1}}^{k_{2}}-\int_{-k_{1}}^{k_{2}}xf(x)dx=k_{2}\\ \int_{\alpha}^{1-\beta}\frac{1-s}{f(F^{-1}(s))}ds & =(k_{2}+k_{1})-k_{2}=k_{1} \end{align*} and \begin{align*} RHS & =1-\alpha-\beta\\ \frac{\partial RHS}{\partial\alpha} & =-2\\ \frac{\partial RHS}{\partial\beta} & =-2 \end{align*} and \begin{align*} \frac{\partial^{2}RHS}{\partial\alpha^{2}} & =\frac{2(1-\alpha)}{k_{1}^{2}\alpha}\\ \frac{\partial^{2}RHS}{\partial\beta^{2}} & =\frac{2(1-\beta)}{k_{2}^{2}\beta}\\ \frac{\partial RHS}{\partial\alpha\partial\beta} & =\frac{2}{k_{1}k_{2}}. \end{align*} Hence \begin{align*} \frac{\partial\sigma^2(\alpha, \beta)}{\partial\alpha} & =\frac{\partial}{\partial\alpha}\frac{RHS}{(1-\alpha-\beta)^{2}}\\ & =\frac{-2}{(1-\alpha-\beta)^{2}}+RHS\cdot\frac{2}{(1-\alpha-\beta)^{3}}=0\\ \frac{\partial\sigma^2(\alpha, \beta)}{\partial\beta} & =\frac{\partial}{\partial\beta}\frac{RHS}{(1-\alpha-\beta)^{2}}=\frac{-2}{(1-\alpha-\beta)^{2}}+RHS\cdot\frac{2}{(1-\alpha-\beta)^{3}}=0 \end{align*} and the Hessian matrix can be calculated as below: \begin{align*} 2\cdot(1,1)^{T}(1,1)\cdot\sigma^2(\alpha,\beta)+(1-\alpha-\beta)^{2}\frac{\partial^{2}\sigma^2(\alpha, \beta)}{\partial(\alpha,\beta)\partial(\alpha,\beta)} & =\left(\begin{array}{cc} \frac{2(1-\alpha)}{k_{1}^{2}\alpha} & \frac{2}{k_{1}k_{2}}\\ \frac{2}{k_{1}k_{2}} & \frac{2(1-\beta)}{k_{2}^{2}\beta} \end{array}\right)\\ \end{align*} so \begin{align*} \frac{\partial^{2}\sigma^2(\alpha, \beta)}{\partial(\alpha,\beta)\partial(\alpha,\beta)} & =\frac{2}{(1-\alpha-\beta)^{2}}\left(\begin{array}{cc} \frac{(1-\alpha)}{k_{1}^{2}\alpha}-\frac{1}{1-\alpha-\beta} & \frac{1}{k_{1}k_{2}}-\frac{1}{1-\alpha-\beta}\\ \frac{1}{k_{1}k_{2}}-\frac{1}{1-\alpha-\beta} & \frac{(1-\beta)}{k_{2}^{2}\beta}-\frac{1}{1-\alpha-\beta} \end{array}\right).\\ \end{align*} Note that \begin{align*} \int_{-k_{1}}^{k_{2}}x^{2}f(x)dx & =\int_{-k_{1}}^{k_{2}}-xf'(x)dx\\ & =\int_{-k_{1}}^{k_{2}}f(x)dx-xf(x)|_{-k_{1}}^{k_{2}}\\ & =1-\alpha-\beta-(k_{2}^{2}\beta+k_{1}^{2}\alpha) \end{align*} that is, \begin{align*} 1-\alpha-\beta & =k_{1}^{2}\alpha+k_{2}^{2}\beta+\int_{-k_{1}}^{k_{2}}x^{2}f(x)dx\\ & =k_{1}^{2}\alpha+k_{2}^{2}\beta+(1-\alpha-\beta)ES^{2} \end{align*} where $S$ is defined on $x\in(-k_{1},k_{2})$ with density $\frac{f(x)}{1-\alpha-\beta}$. Then \begin{align*} ES & =(1-\alpha-\beta)^{-1}\int_{-k_{1}}^{k_{2}}xf(x)dx\\ & =(1-\alpha-\beta)^{-1}\int_{-k_{1}}^{k_{2}}-f'(x)dx\\ & =\frac{k_{1}\alpha-k_{2}\beta}{1-\alpha-\beta} \end{align*} So the determinant of $\frac{\partial^{2}\sigma^2(\alpha,\beta)}{\partial(\alpha,\beta)\partial(\alpha,\beta)}\times\frac{(1-\alpha-\beta)^{2}}{2}$ is \begin{align*} & (\frac{(1-\alpha)}{k_{1}^{2}\alpha}-\frac{1}{1-\alpha-\beta})(\frac{(1-\beta)}{k_{2}^{2}\beta}-\frac{1}{1-\alpha-\beta})-(\frac{1}{k_{1}k_{2}}-\frac{1}{1-\alpha-\beta})^{2}\\ = & \frac{(1-\alpha)}{k_{1}^{2}\alpha}\frac{(1-\beta)}{k_{2}^{2}\beta}-\frac{1}{k_{1}^{2}k_{2}^{2}}+\frac{1}{1-\alpha-\beta}(\frac{2}{k_{1}k_{2}}-\frac{(1-\alpha)}{k_{1}^{2}\alpha}-\frac{(1-\beta)}{k_{2}^{2}\beta})\\ = & \frac{k_{1}^{2}\alpha+k_{2}^{2}\beta+(1-\alpha-\beta)ES^{2}}{k_{1}^{2}k_{2}^{2}\alpha\beta}-\frac{1}{1-\alpha-\beta}\cdot\frac{k_{1}^{2}\alpha(1-\beta)+k_{2}^{2}\beta(1-\alpha)-2k_{1}k_{2}\alpha\beta}{k_{1}^{2}k_{2}^{2}\alpha\beta}\\ = & \frac{(1-\alpha-\beta)^{2}ES^{2}-(k_{1}^{2}\alpha^{2}+k_{2}^{2}\beta^{2}-2k_{1}k_{2}\alpha\beta)}{(1-\alpha-\beta)k_{1}^{2}k_{2}^{2}\alpha\beta}\\ = & \frac{1-\alpha-\beta}{k_{1}^{2}k_{2}^{2}\alpha\beta}\cdot var(S) \end{align*} which is greater than 0 since $\alpha+\beta<1$. This completes the proof.

Proof of Theorem (ref)

Proof of Lemma (ref)

Let $g(y,\theta)=\frac{\partial}{\partial\theta}\log f_{1}(y,\theta)$, where $f_{1}(y,\theta)$ is the same as in (ref).

In order to find the projection of $\dot{\ell}$ on the orthocomplement of the $\dot{P}_{f}$, we need to find $\text{Q}$ which minimizes

align*[align* omitted — 122 chars of source]

Since $Z\in\{0,1\}$ and $\dot{\ell}(y,z,\theta)=z\cdot g(y,\theta)$, we have

align*[align* omitted — 371 chars of source]

where the last equality is due to $P(Y\leq h(y,\theta)|Z=1)=P(Y\leq y|Z=0)$. Therefore,

align*[align* omitted — 99 chars of source]

where $const$ does not depend on $Q$. This leads to

align*[align* omitted — 58 chars of source]

Note that this $Q$ satisfies $\int Q(y,\theta)f_{0}(y)dy=0$ since

align*[align* omitted — 166 chars of source]

must be 0. Thus

align*[align* omitted — 204 chars of source]

and

align*[align* omitted — 183 chars of source]

Proof of Theorem (ref)

The proof is similar to the sample-splitting argument for Claim 3 of Theorem 1 and thus is omitted for conciseness.

Details on the Empirical Implementation

Density Estimation

Our estimators rely on estimates of the density and its first and second derivative.

Using data \(X\) (for instance half the sample of control observations for the estimator with sample splitting), we estimate \(f^{(k)}(x)\) at \(x \in \mathbb{X} \equiv \{X_{(\lceil n_{d} \frac{k}{1000} \rceil)} \; | \; k=1,\dots,999\}\) where \(n_{d}\) is the number of observations in the data \(X\) used for estimating densities. Estimating the density only at points in the density-estimation sample avoids estimating densities of \(0\) in the tails where there are few observations especially when using sample-splitting. We use adaptive kernel estimates (Silverman 1986, ch. 5.3) and a triweight kernel, which allows estimation of the first and second derivative of the density. An alternative based on direct estimation of the logarithm of the density and its derivatives is pinkse2021estimates.

Density estimation proceeds in two steps. First, we obtain an initial estimate of the density at each point \(x \in \mathbb{X}\) as

equation*[equation* omitted — 91 chars of source]

where \(K(u)\) is the triweight kernel

equation*[equation* omitted — 71 chars of source]

and \(h\) is the bandwidth, chosen as

equation*[equation* omitted — 67 chars of source]

where \(3.15\) is the Silverman rule-of-thumb constant for the triweight kernel and

equation*[equation* omitted — 108 chars of source]

is an estimate of the standard deviation of \(X\) if \(X\) is normally distributed but less affected by outliers than the conventional unbiased estimator of the variance.

Second, we obtain our final estimate of the density by calculating an adaptive bandwidth. Intuitively, in regions of high density, we can choose a small bandwidth that lowers bias without incurring much of a cost in terms of variance. In regions of low density, however, a larger bandwidth is needed to have sufficiently many data points near \(x\) to reign in the variance. Following Silverman (1986, ch. 5.3.1), define the local bandwidth factors \(\lambda_{i}\) for \(X_{i}\) as

equation*[equation* omitted — 80 chars of source]

where \(g\) is the geometric mean of the \(\tilde{f}(X_{i})\):

equation*[equation* omitted — 80 chars of source]

Then the estimate of the density is

equation*[equation* omitted — 124 chars of source]

To estimate the first and second derivative of the density, we take \(\lambda_{i}\) as above, and estimate

equation*[equation* omitted — 301 chars of source]

with \(h_{1}\) and \(h_{2}\) the Silverman rule-of-thumb bandwidth for estimating the derivatives of the density using a triweight kernel

equation*[equation* omitted — 166 chars of source]

and derivatives of the triweight kernel \(K\)

equation*[equation* omitted — 221 chars of source]

Then our estimates of the derivatives of the log density are, for \(x \in \mathbb{X}\),

equation*[equation* omitted — 262 chars of source]

When we need to evaluate the second derivative at some other point \(x \not\in \mathbb{X}\), we use linear interpolation and nearest-neighbor extrapolation of the estimated derivative of the log density at the points in \(\mathbb{X}\). In practice, this piecewise-linear approximation causes negligible error because areas where \(\mathbb{X}\) is not tightly spaced have low density by construction and typically very little curvature.

Efficient Influence Function Estimator

The M-estimator given in equation (ref) of the main text, substituting equation (ref) for \(\psi\) and without sample splitting for notational convenience, is:

equation*[equation* omitted — 197 chars of source]

The densities can be estimated as described in the previous section. Rather than this one-step estimate using a single initial estimate \(\tilde{\tau}\) (for instance the difference in medians), we implement the estimator that is the fixed point of the equation above; that is, \(\hat{\tau}^{if} = \tilde{\tau}\). At this fixed point, the equation simplifies, such that the efficient influence function estimator is the solution to the equation:

equation[equation omitted — 262 chars of source]

where the estimates \(\widehat{\frac{\partial \ln f}{\partial x}}\) either come from the full sample or we use sample-splitting as described in Section 2. Standard root-finding algorithms appear to work well for finding the solution \(\hat{\theta}^{if}\) to equation (ref) when initialized at reasonable guesses such as the weighted average quantiles estimator (see below) or the difference in medians.

We estimate standard errors as follows: First, calculate

equation*[equation* omitted — 128 chars of source]

where \(Y_{i}^{(0)}\) for \(i = 0,\dots,n_{0}\) are the control observations. Then an estimate of the approximate finite sample variance of \(\hat{\theta}^{if}\) is

equation*[equation* omitted — 113 chars of source]

where as before \(p = \frac{n_{1}}{n_{0}+n_{1}}\) is the fraction of treated observations. This estimator of the variance relies on the correct specification of the additive shift model.

Weighted Average Quantile Estimator

For the weighted average quantiles estimator, we calculate, assuming $n_0 \geq n_1$

equation[equation omitted — 187 chars of source]

where for $u\in \{i/(n_0+1): i=1,\cdots,n_0\}$

equation*[equation* omitted — 453 chars of source]

Here $c$ is for normalization such that $\sum_{i=1}^{n_0}w(\frac{i}{n_0+1})=1$, and \(\text{quantile}(Y^{(0)},u)\) is the \(u\) quantile of all the control observations for the full-sample estimator. \(\hat{f}(\hat{F}^{-1}(u))\) is an ad-hoc estimate of the density at the \(u\) quantile calculated as \(\hat{f}(\hat{F}^{-1}(i/(n_{0}+1))) = (1/n_{0}) / \max(\Delta Y_{(i)}^{(0)},\Delta Y_{(\lceil n_{1}\frac{i}{n_{0}+1} \rceil)}^{(1)})\) with \(\Delta\) the first difference operator. The factor $\text{MAD}(\hat{f})$, i.e. the median absolute deviation of the data (multiplied by $1.4826$ such that it estimates the standard deviation if the data are normally distributed), is used so that the truncation condition is both shift-invariant and scale-invariant. The truncation of the weights \(\tilde{w}\) addresses violations of condition (ref) in the tails by truncating observations with excessive (normalized) weight relative to the density at the quantile. In the simulations, it improves the performance for the Cauchy distribution for some samples with extreme outliers, and otherwise has no effect.

For a sample-splitting estimator, we randomly split the samples of treated and control in half, use the first half to estimate \(w(u)\) (the densities and \(\text{quantile}(Y^{(0)},u)\)), and then evaluate equation (ref) for the second half of the sample given \(w\). Reverse the roles of the sample splits and average the two estimators, as described in the main text.

Under correct specification, \(\hat{\theta}^{waq}\) and \(\hat{\theta}^{if}\) are asymptotically equivalent, so the standard errors based on \(\hat{I}\) are appropriate for \(\hat{\theta}^{waq}\) as well.

Results With Sample Splitting

Table (ref) reports results for the influence function-based and weighted average quantile estimators with and without sample splitting for the simulations shown in Tables (ref), (ref), and (ref).

table[table omitted — 4,369 chars of source]