EconBase
← Back to paper

How Much Weak Overlap Can Doubly Robust T-Statistics Handle?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,014 characters · 16 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

How Much Weak Overlap Can Doubly Robust T-Statistics Handle?

\doublespacing

abstractIn the presence of sufficiently weak overlap, it is known that no regular root-n-consistent estimators exist and standard estimators may fail to be asymptotically normal. This paper shows that a thresholded version of the standard doubly robust estimator is asymptotically normal with well-calibrated Wald confidence intervals even when constructed using nonparametric estimates of the propensity score and conditional mean outcome. The analysis implies a cost of weak overlap in terms of black-box nuisance rates, borne when the semiparametric bound is infinite, and the contribution of outcome smoothness to the outcome regression rate, which is incurred even when the semiparametric bound is finite. As a byproduct of this analysis, I show that under weak overlap, the optimal global regression rate is the same as the optimal pointwise regression rate, without the usual polylogarithmic penalty. The high-level conditions yield new rules of thumb for thresholding in practice. In simulations, thresholded AIPW can exhibit moderate overrejection in small samples, but I am unable to reject a null hypothesis of exact coverage in large samples. In an empirical application, the clipped AIPW estimator that targets the standard average treatment effect yields similar precision to a heuristic 10% fixed-trimming approach that changes the target sample.

Introduction

In observational data, it is common to find at least some treated observations with small propensity scores, in which case overlap is weak. Under weak overlap, outcome regression is more difficult and inverse propensity estimators can divide by small numbers.

This paper studies the extent to which standard asymptotic results for doubly robust estimators continue to apply when strict overlap does not hold, but there is a tail bound on inverse propensity scores. The paper aims to answer two questions: \\[-1.8ex]

enumerate[label=(\arabic*),topsep=0pt,itemsep=-1ex] • Black-box rates. Under what conditions on nuisance errors do Wald confidence intervals constructed using doubly robust estimators have well-calibrated coverage? • Feasibility. When are the black-box rates in (ref) feasible under weak overlap? \\[-1.8ex]

Overlap weakness is measured with a tail-bound parameter $\gamma_0$. It is known that there is a phase transition when $\gamma_0$ crosses two, which corresponds to a uniform propensity density.

Values of the tail parameter $\gamma_0$ above two correspond to stronger overlap assurances, a case which I call somewhat weak overlap. Under somewhat weak overlap, it is known that the semiparametric bound for estimation of $\psi_0$ is finite and can be achieved by the doubly robust Augmented Inverse Propensity Weighted (AIPW) estimator with known conditional mean outcome and propensity functions. I show that semiparametric efficiency holds even when using nonparametric estimates of the two AIPW nuisance functions, provided the nuisance estimates are constructed by cross-fitting and achieve error rates $r_{\mu,n}$ and $r_{e,n}$ that satisfy the product condition that $n^{1/2} r_{\mu,n} r_{e,n}$ tends to zero in probability. The one modification I make relative to the usual analysis is that I require the nuisance rates to apply uniformly in regions with small propensity scores, which is immediate under $L_\infty$ bounds.

Values of $\gamma_0$ below two allow the density of propensity scores to be unbounded at zero, a case which I call very weak overlap. Under very weak overlap, it is known that the semiparametric efficiency bound is infinite and Inverse Propensity Weighting (IPW) estimators fail to be asymptotically normal, even if the propensity function is known. XinweiMaRobustIPW show that thresholding strategies that clip (Winsorize) or trim (discard) observations with small propensity scores at a sequence of $b_n$ tending to zero can restore asymptotic normality, but at a slower-than-$\sqrt{n}$ rate and with first-order bias. Previous work has proposed constructing confidence intervals for this case using debiasing strategies and self-normalized subsampling. I show that if thresholding is applied to AIPW, then there is no first-order bias and simple Wald confidence have exact asymptotic coverage. The analysis provides sufficient black-box conditions for AIPW with estimated nuisance functions to achieve these asymptotic results: first, that the threshold $b_n$ tends to zero more slowly than the propensity error $r_{e,n}$; and second, that the product $n^{1/2} r_{\mu,n} b_n^{\gamma_0 / 2}$ tends to zero in probability.

That analysis characterizes the feasibility of AIPW Wald confidence intervals under black-box conditions on the propensity estimates and outcome regression. An added difficulty is that in regions with weak overlap, there cannot be many treated observations to use in outcome regression. I quantify the added difficulty for the case of regression within a Hölder smoothness class of order $\beta_{\mu} > 0$. I show that weak overlap scales the effective outcome smoothness by $1 - 1 / \gamma_0$, exhibiting a cost of weak overlap even in the somewhat weak overlap regime in which the other asymptotic results carry through nearly unchanged. I show that in this setting, the optimal global rate is equal to the optimal pointwise rate, without the usual polylogarithmic factor separating the optimal pointwise and global rates Stone1982. The argument leverages a novel construction partitioning the covariate space into regions with increasingly strong overlap assurances in such a way as the contribution of each region to the worst-case error is half as large as the previous contribution. In this way, the contribution of infinitely many regions to the global error is controlled uniformly.

Taken together, these results provide a precise answer to the question posed by this work's title: doubly-robust t-statistics can handle weak overlap of tail bound $\gamma_0$, provided the outcome and propensity nuisance functions are in Hölder smoothness classes of order $\beta_{\mu}$ and $\beta_{e}$ and

align*[align* omitted — 172 chars of source]

In this case, the threshold $b_n = n^{-\beta_e / (2 \beta_e + d)} \log(n)^{(3 \beta_e + d) / (2 \beta_e + d)}$ suffices, regardless of the weak overlap parameter $\gamma_0$. When the outcome and propensity smoothness orders are the same $\beta > 0$, then thresholded AIPW can handle weak overlap of order $\gamma_0$, so long as:

align*[align* omitted — 141 chars of source]

Under Lipschitz continuity of both nuisance functions in one dimension, doubly robust t-statistics can handle weak overlap of order $\gamma_0 > \frac{5}{3}$. In higher dimensions, there is always some sufficient smoothness order that yields valid t-statistics for any fixed tail bound.

The conditions here yield new rules of thumb for thresholding in applied work. In my favored regime, the econometrician is willing to posit a minimal consistency rate for one of the two nuisance function estimates. Given such a minimal rate, a simple plug-in procedure predicts the threshold with the laxest restriction on the other nuisance function needed to achieve well-calibrated Wald confidence intervals. In the absence of any such information, a third rule of thumb derives a threshold that imposes the laxest equal minimal consistency rate on both nuisance estimates. None of these rules of thumb depend directly on knowledge of the tail bound parameter $\gamma_0$.

In simulations, I find that clipped AIPW achieves the promised properties asymptotically. I consider a setting of very weak overlap with nonparametric outcome regression and propensity estimates. Unthresholded IPW and AIPW estimators perform poorly, with large errors and nonnormal asymptotic distributions. In this setting, clipped IPW displays its known first-order bias, and clipped AIPW displays the second-order bias justified by the theoretical analysis. With access to 1,000 or 10,000 observations, I find that p-values based on clipped AIPW t-statistics exhibit moderate overrejection. In large samples with 100,000 observations, a Kolmogorov-Smirnov test based on 5,000 simulations is unable to reject a null hypothesis that clipped AIPW p-values on the true causal effect are exactly uniformly distributed.

I apply the clipped AIPW estimator to data on right heart catheterization. I consider the setting of ConnorsEtAl1996, which has become a canonical setting with weak overlap, including providing the empirical application for CrumpEtAlOptimizePrecision's paper proposing a 10% fixed-trimming rule of thumb. I compare the clipped AIPW estimator that targets the full-population effect to estimators that apply AIPW to a sample trimmed based on a fixed rule. I find that by including observations with small estimated propensities, the clipped AIPW strategy increases the estimated harm of the procedure by 0.17 standard errors relative to the 10 percent fixed-trimming rule, while increasing the estimated standard error by only 5.1%. These results show that targeting the full-population treatment effect does not need to introduce a major efficiency loss, and show that thresholded AIPW can easily be added as a robustness test when practitioners apply a fixed-trimming rule.

Weak overlap is a common phenomenon in practice and in theory. The dominant response to weak overlap in inverse propensity score practice is trimming: dropping samples with small propensity estimates in order to estimate average effects within a more precise population CurrieWalker, BaileyGoodmanBacon, GalianiEtAl, typically following the 10% rule of thumb from CrumpEtAlOptimizePrecision. Other proposals targeting new samples include reweighting towards higher-precision populations YangAndDing, LiEtAlOverlapWeights or clipping strategies that Winsorize weights above LeeEtAlWeightTrimming, IonidesClipping.\footnote{Awkwardly, the epidemiological literature sometimes refers to the Winsorization strategy as “trimming." My results hold for both dropping or Winsorizing extreme propensities, so the confused reader can view this as a work deriving simple asymptotics for trimmed AIPW regardless of their preferred meaning of “trim."} DAmourEtAlOverlap argue that weak overlap is likely to be prevalent in modern settings with high-dimensional covariates. imbens_unconfoundedness_review argues that changing the target estimand may be necessary in the absence of sufficient precision.

There is theoretical work characterizing estimation of nonstandard estimands under weak overlap. KhanAndTamer show that very weak overlap yields an irregularly identified parameter, an infinite semiparametric efficiency bound, slower-than-$\sqrt{n}$ estimation rates for the traditional average causal effects, and no clear notion of best estimator. An important theoretical literature has proposed novel point and confidence interval estimators with desirable properties under weak overlap RotheLimitedOverlap, ArmstrongKolesarBandwidthSnooping, ArmstrongKolesarFiniteSampleOptimal, SasakiUra2022, ma2023doubly, ChaudhuriHill. Many of these procedures have favorable properties relative to the simpler thresholded estimator I consider here, but to my knowledge there has been little take-up by practitioners. XinweiMaRobustIPW and KhanUganderDoublyRobustTrimming show that sufficiently trimmed AIPW and IPW can remain asymptotically normal, but at the cost of introducing first-order bias for the standard causal effects that often calls for a nonstandard debiasing strategy that can enable laxer regression conditions than the standard smoothness conditions I explore. CrumpEtAlOptimizePrecision, YangAndDing, LiEtAlOverlapWeights, and ContaminationBiasInLinearRegressions propose changing estimands in response to weak overlap, which introduces a discontinuous estimand based on whether or not the econometrician detects meaningful overlap weakness.

Other theoretical literature so far has either proposed a nonstandard causal estimator, required nonstandard techniques to construct confidence intervals, or done both. XinweiMaRobustIPW and HeilerKazak propose using self-normalized subsampling methods that enable valid statistical inference for standard estimands without clipping or trimming, but empirical practice has favored simple t-tests. XinweiMaRobustIPW also propose a debiasing procedure. HeilerKazak also find that estimated untrimmed AIPW is first-order equivalent to the oracle AIPW estimator with known nuisance functions, and the associated asymptotic distribution is alpha-stable, if the product of nuisance estimation rates is of a lower order than the oracle standard deviation; I find this result does not extend to the thresholded AIPW estimator that I show is asymptotically normal. MaSasakiWang and LihuaLeiTesting propose statistical tests under a null of sufficient or strict overlap, respectively, presumably in the hopes of avoiding these complications. My analysis of nonparametric regression rates is also relevant to the literature on regression with degenerate designs, although to my knowledge the possibility of a global rate with no polylogarithmic penalty is new Stone1982, HallEtAlLocalLinear, GaiffasPointwise, mou2023kernelbased.

The plan of the paper is as follows. (ref) presents the setting and main theoretical results. (ref) interprets these results as minimal black-box consistency rates and as minimal smoothness rates. (ref) considers implications for parametric estimators and derives some rules of thumb for empirical use. (ref) presents numerical results for simulations and the empirical application to right-heart catheterization. (ref) concludes.

Notation. I follow HeilerKazak and use “strict overlap" to refer to the case in which the propensity score is bounded away from zero almost surely; I use “weak overlap" to refer to the case in which the infimum of the support of the propensity score is zero, which is sometimes called “limited overlap" KhanAndTamer, ChaudhuriHill. I focus my attention on distributions with weak overlap that may possess subexponential tails. I use “very weak overlap" to refer to case in which the associated heavy tails can fail to generate inverse propensity moments, a class which is sometimes called “heavy tailed" ChaudhuriHill. I use “somewhat weak overlap" to refer to the case in which I allow only subexponential tails that are sufficiently light, a class which is sometimes said to satisfy “strict overlap" or “overlap" HeilerKazak, brunssmith2024augmentedbalancingweightslinear. I write $\tilde{\psi}_{(Oracle)}^{AIPW}(b_n) = \frac{1}{n} \sum \phi(Z \mid b_n, \eta)$ for the oracle AIPW estimate with pseudo-outcome $\phi(Z \mid b_n, \{ \bar{e}(X), \bar{\mu}(X) \}) = \bar{\mu}(X) + \frac{D (Y - \bar{\mu}(X))}{\max\{\bar{e}(X), b\}}$ and clipping threshold $b_n$ and $\sigma_n = n^{-1/2} \sqrt{ \frac{1}{n} \sum \phi(Z \mid b_n, \eta)^2 - \tilde{\psi}_{(Oracle)}^{AIPW}(b_n)^2}$ and $\hat{\sigma}_n = n^{-1/2} \sqrt{ \frac{1}{n} \sum \phi(Z \mid b_n, \hat{\eta} )^2 - \left( \frac{1}{n} \sum \phi(Z \mid b_n, \hat{\eta}) \right)^2}$ for the associated oracle and estimated sample standard deviation, respectively. I refer to regions of the covariate space in which the propensity can be arbitrarily close to zero as singularities. I use the notation $E_{P}[ \cdot ]$ and $E[ \cdot ]$ to refer to the expectation under the maintained distribution $P$, and I use the notation $\psi(P)$ to refer to the statistical average potential outcome $E_{P}[E_{P}[Y \mid X, D=1]]$ where the right-hand side is well-defined under $P$. I abuse notation and write $\psi = \psi(P)$ and use $\sup_{P \in A} B$ to refer to the supremum of $B$ over distributions $P$ in $A$ under any maintained restrictions on the distribution and nuisance functions. I write that a set of nuisance functions are cross-fit if the data is partitioned into $K$ folds and the nuisance functions in fold $k$ are independent of the data in fold $k$. I write $A_n \leq_{P} B_n$ to refer to the case that for all $\epsilon > 0$, $P( A_n > B_n + \epsilon) \to 0$. I write $P\left( E_n \right)$ for the probability of event $E_n$ occurring under the distribution $P$, with the number of draws $n$ sometimes left implicit. I use the notation $c_n \ll d_n$ for nonnegative sequences $c_n, d_n$ to indicate that $d_n > 0$ for all $n$ large enough and $c_n / d_n \to 0$. I use the notation $c_n \precsim d_n$ and $d_n \succsim c_n$ to indicate that there is some $\delta > 0$ such that $d_n \geq \delta c_n$ for all $n$ large enough. I write $c_n = o_P(d_n)$ for sequence of $d_n > 0$ to indicate that for all $\delta > 0$, $P( | c_n | / d_n > \delta ) \to 0$; if there is only one distribution in a statement, $c_n = o(d_n)$ should be understood to mean $c_n = o_{P}(d_n)$. I use $\log$ to refer to the natural logarithm and $a \vee b$ to indicate $\max\{ a, b \}$. I define Hölder smoothness using a multivariate version of the notation of TsybakovBook2009: a function $f$ is in the Hölder smoothness class $\Sigma(\beta, L)$ if the $\lfloor \beta \rfloor$-order multivariate derivatives $D^{\alpha} f = \frac{\partial \| \alpha \|}{\partial x_{1}^{\alpha_1} \partial x_{2}^{\alpha_2} \hdots} f$ satisfy $\| D^{\alpha} f(x) - D^{\alpha} f(x') \| \leq L \| x - x' \|^{\beta - \lfloor \beta \rfloor}$, where I write $D^{\alpha} f(x)$ for $D^{\alpha} f$ evaluated at $x$. For simplicity, I use local polynomial regression to refer to specifically kernel regression with uniform bandwidth: $\hat{\mu}^{(NW)}(x \mid h) = \frac{\sum D 1\{ \| X - x \| \leq h \} Y}{\sum D 1\{ \| X - x \| \leq h \}}$ when feasible and $\hat{\mu}^{(NW)}(x \mid h) = 0$ when no nearby treated observations are available. For an estimator $\hat{\eta}$ of $\eta$, I use $\| \hat{\eta} - \eta \|_{\infty}$ to refer to the sup norm $\sup_{x \in Support(P)} | \hat{\eta}(x) - \eta(x) |$.

Setting, Consistency, and Asymptotic Normality

This section presents asymptotic results under black-box nuisance conditions.

Setting

I derive uniform convergence rates under lower bounds on overlap weakness. I follow XinweiMaRobustIPW, who provide important building blocks in my analysis, and parameterize overlap weakness through a tail parameter $\gamma_0$. Unlike their analysis, the results will be uniform over a model family $\mathscr{P}$ satisfying certain restrictions, including some basic regularity conditions.

assumptionLet $\mathscr{P}$ be a nonempty family of distributions, and write $e(X) = P(D = 1 \mid X)$ and $\mu(X) = E_{P}[Y \mid X, D=1]$. Then every $P \in \mathscr{P}$ satisfies the following conditions for some $q > 3, M, \sigma_{\min}, C > 0$ and $\gamma_0 > 1$: \begin{enumerate}[label=(\alph*), itemsep=-0.5ex, topsep=-0.5ex] • Conditional moments. $\mathbb{E}[|Y - \mu(X) |^q \mid X, D=1] \leq M^q < \infty$ almost surely. • Unconditional moments. $Var( \mu(X) ) \leq M$. • Residuals. $\textup{Var}(Y \mid X, D) \geq \sigma_{\min}^2$. • Propensity tail. $P(e(X) \leq \pi) \leq C \pi^{\gamma_0 - 1}$ for all $\pi \in [0, 1]$. \end{enumerate}

Definition (ref) generalizes XinweiMaRobustIPW's slowly varying tails assumption. Assumptions (ref)(ref) through (ref)(ref) are regularity conditions that rule out cases like perfectly predictable outcomes. (ref)(ref) provides the substantial restriction on $\mathscr{P}$: overlap may be weak in the sense that $\gamma_0$ is finite, but there is some minimal $\gamma_0$ and $C$ that provides a lower bound on the propensity's tail behavior. Under strict overlap, (ref)(ref) holds for any finite $\gamma_0 > 1$, and most results here hold after replacing $\gamma_0$ with infinity. Under weak overlap, (ref)(ref) may only hold for some values of $\gamma_0$, in which case the inverse propensity distribution may be heavy-tailed. I refer to the case $\gamma_0 > 2$ as “somewhat weak overlap" and refer to the case of $\gamma_0 < 2$ as “very weak overlap." As $\gamma_0$ shrinks below $2$, overlap is permitted to be increasingly weak. $\gamma_0 \leq 1$ corresponds to no bound on the propensity distribution.

figure[figure omitted — 292 chars of source]

(ref) illustrates behavior for simulated data with various values of $\gamma_0$. When the propensity score $e(X)$ has a well-defined density, $\gamma_0 = 2$ corresponds to a roughly uniform distribution of propensity scores XinweiMaRobustIPW. When $\gamma_0$ is above two, the density of propensity scores tends to zero at zero; when $\gamma_0$ is below two, the density of propensity scores can tend to infinity at zero. Heuristically, there are never too many treated observations with very small propensity scores, but $\gamma_0$ governs the degree to which there can be many untreated observations with very small propensity scores.

A phase transition occurs when $\gamma_0$ crosses two.

proposition(i) Suppose (ref) holds for some $\gamma_0 > 2$. Then the semiparametric bound is finite for all $P \in \mathscr{P}$. (ii) Suppose (ref) holds for some $\gamma_0 \in (1, 2)$, and there is a $P \in \mathscr{P}$ and $C' > 0$ such that $P( e(X) \leq \pi ) \geq C' \pi^{\gamma_0 - 1}$ for all $\pi \in (0, 1]$. Then the semiparametric bound is infinite for $P$.

I will require certain rates on the nuisance functions $e(X)$ and $\mu(X)$. I write the worst-case rates as $r_{e,n}$ and $r_{\mu,n}$.

assumption[Cross-fitting and $L_\infty$ rates] The nuisances $\hat{\mu}$ and $\hat{e}$ are estimated with cross-fitting with a fixed number of folds $K$. If $n_k$ is the number of observations per fold, then $\inf_k n_k / \sup_k n_k \to 1$. Further, for all $k \in 1, \hdots, K$ and all $P \in \mathscr{P}$, the cross-fit nuisances satisfy the uniform consistency rates $\mathbb{E}_{P}[\| \hat{\mu}^{(-k)}_n - \mu \|_{\infty}] \leq r_{\mu,n}$ and $\mathbb{E}_{P}[\| \hat{e}_n^{(-k)} - e \|_{\infty}] \leq r_{e,n}$ where $r_{\mu,n}, r_{e,n}$ are uniformly bounded above.

Cross-fitting is a common strategy for simplifying the analysis of Neyman-orthogonal estimators like AIPW doubleML. In practice, nuisances satisfying (ref) may only be achieved with arbitrarily high probability. I impose uniformity to ensure these rates hold in regions of $x$ with weak overlap, but can be bypassed with $L_2$ error conditions in other regions. Such uniformity assumptions are standard in studying semiparametric estimators under irregular identification semenova2024aggregatedintersectionboundsaggregated.

Estimator and Consistency

My formal analysis considers the clipped AIPW estimator with cross-fit nuisance function estimates. Results for the other standard thresholding procedure, trimming, generally follow by the same arguments. I begin by providing sufficient conditions for consistency.

For simplicity, I focus the theoretical analysis on estimating the average potential outcome $\psi = E[D Y / e(X)]$. The average treatment effect follows as a corollary. The clipped AIPW estimator of $\psi$ is:

align[align omitted — 285 chars of source]

In that equation, $\mathcal{F}^k$ is the set of observations $i$ randomly partitioned in fold $k$, $\hat{\eta}^{(-k)}$ is the nuisance function estimates constructed only on observations in folds other than $k$. The unthresholded AIPW estimator is the special case of $b_n = 0$. I analyze the clipped AIPW estimator because results for the trimmed AIPW estimator follow somewhat more easily.

A standard result for the unthresholded AIPW estimator is double robustness: when $e(X)$ is bounded away from zero, unthresholded AIPW is consistent for $\psi$ if either $r_{e,n}$ or $r_{\mu,n}$ tends to zero. The existence of weak overlap introduces a subtlety to double robustness.

proposition[Consistency] Suppose $b_n$ satisfies $n^{-1/2} \ll b_n \ll 1$, the conditions of (ref) hold, and either (i) $r_{e,n} b_n^{\min\{\gamma_0-2, 0\}} \to 0$ or (ii) $r_{\mu,n} \frac{r_{e,n} + b_n}{b_n} \to 0$. Then for all $\epsilon > 0$, $$\sup_{P \in \mathscr{P}} P\left( \left| \hat{\psi}_{clip}^{AIPW}(b_n) - \psi(P) \right| > \epsilon \right) \to 0.$$

Under very weak overlap, condition (i) is stronger than the classic strict overlap condition that $r_{e,n}$ consistency implies estimator consistency. It requires that $r_{e,n}$ go to zero faster than $b_n^{2-\gamma_0}$, so that as overlap is allowed to be weaker, the propensity consistency rate may need to be as fast as $b_n$ itself. Under even somewhat weak overlap, condition (ii) is stronger than the classic $r_{\mu,n} \to 0$ condition. With an inconsistent propensity score, a meaningful fraction of the data may be clipped even asymptotically, in which case the outcome regression error rate must offset the positive probability of incorrectly assigning an inverse propensity weight of $b_n^{-1}$.

Asymptotic Normality and Confidence Intervals

This subsection presents the main theoretical claims of the paper. It shows that under suitable rate restrictions, the clipped AIPW estimator is first-order equivalent to an oracle clipped AIPW estimator, both estimators are consistent and asymptotically normal, and simple Wald confidence intervals are well-calibrated.

A common strategy for deriving confidence intervals for unthresholded AIPW under strict overlap is Neyman orthogonality. In those classic settings, the difference between the feasible AIPW estimator with estimated nuisance functions and the hypothetical oracle AIPW estimator with known nuisance functions is

align*[align* omitted — 234 chars of source]

Intuitively, the regression errors $\hat{\mu} - \mu$ are debiased by the inverse propensity estimates $\frac{D}{\hat{e}}$ of the number one. As a result, in classical settings, slowly consistent nuisance estimates can yield quickly consistent causal estimates. When all nuisances are consistent at $o(n^{-1/4})$ rates and $e(X)$ is bounded away from zero, inverse propensity errors are of the same order as propensity errors, classical AIPW estimates are first-order equivalent to oracle estimates with known nuisances, and simple Wald confidence intervals cover the true causal effect by appeal to the asymptotically normal oracle AIPW estimator.

Under very weak overlap, thresholded AIPW does not obtain the standard debiasing benefit. The analogous decomposition for clipped AIPW is

align*[align* omitted — 276 chars of source]

Above the clipping threshold $b_n$, thresholded AIPW's nuisance estimation error enjoys a product-of-errors character that is similar to classical settings, albeit with multiplication by weights as large as $b_n^{-1}$. Below the clipping threshold, the regression errors are not debiased by any inverse propensity estimate.

Thresholded AIPW enjoys a subtly different form of debiasing: as the threshold tends to zero, increasingly little mass remains. If the threshold $b_n$ tends to zero quickly enough, the bias in the thresholded region is debiased by the threshold itself. If the threshold $b_n$ tends to zero slowly enough, the product of errors condition can remain feasible. My formal contribution is to show that there can be a Goldilocks range where $b_n$ tends to zero neither too slowly nor too quickly, so that thresholded AIPW is first-order equivalent to oracle AIPW and is asymptotically normal by appeal to the asymptotically normal thresholded oracle AIPW estimator.

My results for asymptotic normality and statistical inference will proceed under the following rate requirements.

assumption[Minimal rates] Assumption (ref) holds, with the following rates on the regression error $r_{\mu,n}$ and the propensity error $r_{e,n}$: \begin{enumerate}[label=(\alph*), itemsep=-0.5ex, topsep=-0.5ex] • Consistency. $r_{\mu,n}, r_{e,n} \to 0$. • Product of errors. $r_{\mu,n} r_{e,n} \left( 1 + b_n^{(\gamma_0 - 2) / 2} \log(1 / b_n)^{1\{ \gamma_0 = 2 \} / 2} \right) \ll n^{-1/2} $. • Regression error near singularities. $r_{\mu,n} b_n^{\gamma_0 / 2} \ll n^{-1/2}$. • Asymptotically known thresholding. $r_{e,n} \ll b_n$. \end{enumerate}

Under very weak overlap, conditions (ref) and (ref) are stronger than the standard product-of-errors condition $r_{\mu,n} r_{e,n} \ll n^{-1/2}$. For example, when $\gamma_0 \geq 1.5$ and $r_{\mu,n} = n^{-1/4}$, then $r_{e,n} \ll n^{-1/3}$ will suffice, provided $r_{e,n} \ll b_n \ll n^{-1/3}$. However, these conditions never require parametric $n^{-1/2}$ consistency rates: shared regression rates of $n^{-1/3}$ will always suffice for these conditions, provided the clipping threshold $b_n$ goes to zero at a rate sufficiently close to $n^{-1/3}$. I discuss these conditions further in (ref).

Certain technical possibilities call for one of two alternative further assumptions: a distributional smoothness assumption or a stronger rate assumption.

assumption[Nongeneracy or faster rates] One of the following two conditions hold: \begin{enumerate}[label=(\roman*)] • Nondegenerate overlap. There exists some $\rho > 0$ such that for all $P \in \mathscr{P}$ and $\pi \in [0, 1]$, $P( e(X) \leq \pi / 2 ) \leq (1-\rho) P( e(X) \leq \pi)$. • Faster rates. $r_{\mu,n} b_n^{(\gamma_0 - 1) 2 / \gamma_0} \ll n^{-1/2}$. \end{enumerate}

(ref)(ref) is a uniform version of the requirement that $P( e(X) \leq x ) = c(x) x^{\gamma_0 - 1}$ for $c(x)$ tending to a constant at zero. The definition formalizes the notion that a distribution may place some propensity mass near zero, but it may not place mass adversarially within the region of the origin. When $\gamma_0 < 2$, (ref)(ref) is stronger than (ref)(ref). As $\gamma_0$ tends to one, the condition approaches the parametric requirement $r_{\mu,n} = O(n^{-1/2})$.

I now provide the main theoretical result.

theorem[(Slow) Asymptotic Normality] Suppose $b_n$ satisfies $n^{-1/2} \ll b_n \ll 1$, and Assumptions (ref), (ref), (ref), and (ref) hold. Then the clipped AIPW estimator is oracle-equivalent: $$\lim_{n \to \infty} \sup_{P \in \mathscr{P}} \sigma_n^{-2} E_{P}\left[ \left( \hat{\psi}_{clip}^{AIPW}(b_n) - \tilde{\psi}_{(Oracle)}^{AIPW}(b_n) \right)^2 \right] = 0.$$ Further, clipped AIPW is asymptotically normal: $$\lim_{n \to \infty} \sup_{P \in \mathscr{P}} \sup_{t \in \mathbb{R}} \left| P \left( \frac{\hat{\psi}_{clip}^{AIPW}(b_n) - \psi(P)}{\hat{\sigma}_n} \leq t \right) - \Phi(t) \right| = 0.$$

(ref) is the core theoretical claim of this paper. The first result shows that thresholded AIPW is first-order equivalent to an oracle estimator with known nuisances: the effect of nuisance estimation error on the treatment effect estimate tends to zero faster than the standard deviation of the oracle estimator. The second result leverages this first-order equivalence to characterize the asymptotic distribution of the clipped AIPW estimates and t-statistics: the estimator is asymptotically normal, and estimated t-statistics are asymptotically standard normal.

Both results are standard for AIPW under strict overlap, but substantial care is required to handle unbounded inverse propensities under weak overlap. The argument for normality builds on XinweiMaRobustIPW's proof that aggressively-trimmed oracle IPW with known propensities achieves asymptotic normality with first-order bias. I extend their argument to a uniform family of distributions using the Berry-Esseen Theorem and note that oracle AIPW must have zero finite-sample bias.

The main task of (ref) is to show that replacing the true nuisances with estimated nuisances has a second-order effect on clipped AIPW estimates under appropriate conditions. This is nontrivial, because under weak overlap, there is an asymptotically unbounded number of observations with arbitrarily large inverse propensities with even known nuisance functions. Nevertheless, by taking appropriate care and leveraging that clipping introduces bias by reducing inverse propensities, I am able to show that the effect of nuisance estimation is second-order even under the very weak overlap case in which unthresholded AIPW fails to be asymptotically normal and no regular root-n estimators exist.

Next, I show that (ref) yields the natural result for inference: simple t-tests based on Wald confidence intervals are well-calibrated.

corollary[T-tests are well-calibrated] Suppose the conditions of (ref) hold. Consider the Wald confidence interval $\hat{\mathcal{C}}_n(\alpha) = \left[ \hat{\psi}_{clip}^{AIPW}(b_n) + z_{\alpha/2} \hat{\sigma}_n, \hat{\psi}_{clip}^{AIPW}(b_n) + z_{1-\alpha/2} \hat{\sigma}_n \right]$. Then for all $\alpha \in (0, 1/2)$, $$\limsup_{n \to \infty} \sup_{P \in \mathscr{P}} \left| P( \psi(P) \in \hat{\mathcal{C}}_n(\alpha)) - (1-\alpha) \right| = 0.$$

Under somewhat weak overlap, unthresholded AIPW is semiparametrically efficient. I now show thresholding is also unnecessary.

corollary[Thresholding is second-order under somewhat weak overlap] Suppose Assumption (ref) holds for some $\gamma_0 > 2$, $r_{e,n}$ and $r_{\mu,n} \to 0$, and $r_{e,n} r_{\mu,n} \ll n^{-1/2}$. Then the feasible AIPW estimator with $b_n = 0$ is semiparametrically efficient, and the associated Wald confidence interval $\hat{\mathcal{C}}_n(\alpha)$ satisfies $$\limsup_{n \to \infty} \sup_{P \in \mathscr{P}} \left| P( \psi(P) \in \hat{\mathcal{C}}_n(\alpha)) - (1-\alpha) \right| = 0.$$

The logic of (ref) is to show that any sequence of $b_n \to 0$ has a second-order effect on estimation, so that unthresholded and thresholded AIPW are first-order equivalent, and there is some $b_n \to 0$ slowly enough satisfying the conditions of (ref), so that there is a thresholded AIPW estimator achieving (ref).

Taken together, this subsection yields a remarkable result for practice. The distribution $P$ may place so much propensity mass near the origin that the semiparametric efficiency bound is infinite, the lower bound on the density of propensity mass near the origin can be so weak that identification nearly fails, and the nuisance estimator may be so poorly designed that it pushes all observations' estimated propensities towards the origin at a slower-than-parametric rate. Nevertheless, Neyman orthogonality is sufficiently powerful to ensure the validity of the simple t-test.

The next section interprets the rate requirements of (ref).

Interpretation of Nuisance Requirements

table[table omitted — 2,529 chars of source]

This section interprets the rate requirements for Wald confidence intervals to cover asymptotically. I summarize the results in (ref). Under somewhat weak overlap, thresholded and unthresholded AIPW remain semiparametrically efficient and $\sqrt{n}$-consistent, and the traditional product of errors nuisance condition remains in place with a modification to an $L_\infty$ on errors. Under very weak overlap, clipped AIPW achieves a slower consistency rate and the required black box nuisance rates are more stringent. Both cases make outcome regression more difficult, but never so difficult as to require parametric assumptions. As a byproduct of this analysis, I show that the optimal pointwise and global regression rates under weak overlap are the same, without the usual polylogarithmic factor in the global rate.

Degradation of Consistency Rate

The previous analysis suggests that smaller values of $b_n$ are preferable because they admit weaker black-box requirements. However, under very weak overlap, larger values of $b_n$ correspond to faster AIPW rates.

I characterize the consistency rate of any oracle-equivalent estimator as follows.

proposition[Consistency rate] Suppose the assumptions of (ref) hold. Then there exist positive constants $c_{\min}$ and $c_{\max}$ such that $c_{\min} n^{-1} \mathbb{E}_{P}\left[ \frac{D}{\max\{e(X), b_n\}^2} \right] \leq \sigma_n^2 \leq c_{\max} n^{-1} \mathbb{E}_{P}\left[ \frac{D}{\max\{e(X), b_n\}^2} \right]$ for all $P \in \mathscr{P}$, where $\sigma_n^2 = n^{-1} \left( \frac{1}{n} \sum \phi(Z \mid b_n, \eta)^2 - \tilde{\psi}_{(Oracle)}^{AIPW}(b_n)^2 \right)$ is the oracle sample variance.

If the estimator were trimmed instead of clipped, $\mathbb{E}_{P}\left[ \frac{D}{\max\{e(X), b_n\}^2} \right]$ would be replaced by a term like $\mathbb{E}_{P}\left[ \frac{D 1\{ e(X) \geq b_n \}}{e(X)} \right]$. Weaker overlap corresponds to larger values of $\mathbb{E}_{P}[D / \max\{e(X), b_n\}^2]$ and slower consistency rates. Conditional on $P$, larger values of $b_n$ correspond to a smaller value of $\mathbb{E}_{P}[D / \max\{e(X), b_n\}^2]$, faster oracle consistency, and greater asymptotic power.

(ref) implies a worst-case consistency rate over distributions in $\mathscr{P}$. I focus on the case of very weak overlap, because (ref) shows that under somewhat weak overlap, clipped and traditional AIPW achieve a traditional $\sqrt{n}$ consistency rate.

corollary[Worst-case consistency rate] Suppose $\gamma_0 < 2$ and let $b_n$ be a fixed sequence of $b_n$ satisfying $1 \gg b_n \gg n^{-1/2}$. There exists a $C' > 0$ such that for $\mathscr{P}$ satisfying (ref), $C' n^{-1} b_n^{\gamma_0 - 2} \geq \sup_{P \in \mathscr{P}} \sigma_n^2$ for all $n$ large enough. Further, there exists a family $\mathscr{P}$ satisfying (ref) and a $C'' \in (0, C')$ such that $\sup_{P \in \mathscr{P}} \sigma_n^2 \geq C'' n^{-1} b_n^{\gamma_0 - 2}$ for all $n$ large enough.

The rate $n^{-1} b_n^{\gamma_0 - 2}$ is a worst-case consistency rate in $b_n$: every distribution in $\mathscr{P}$ achieves a consistency at least as fast as $n^{-1} b_n^{\gamma_0 - 2}$, and it is possible to find a distribution for which the consistency rate is no faster. The combination of (ref) and (ref) yields a trade-off under very weak overlap: smaller values of $b_n$ yield laxer requirements on regression estimation near singularities, but lead to larger variance and slower consistency.

Degradation of Black-Box Nuisance Requirements

Under very weak overlap, the black-box rates of (ref) are more stringent than the usual $r_{\mu,n} r_{e,n} \ll n^{-1/2}$ condition. The main requirement is that that $r_{\mu,n} r_{e,n}^{\min\{ \gamma_0 / 2, 1\}}$ goes to zero faster than $n^{-1/2}$. As a result, outcome regression rates are more valuable than nominally equivalent propensity rates under very weak overlap.

The usual product-of-errors condition under strict overlap often takes a form like $r_{\mu,n} r_{e,n} \ll \sigma_n$, where $\sigma_n$ is the standard deviation of the Oracle estimator. For example, HeilerKazak argue that this condition is sufficient for estimated unthresholded AIPW to be first-order equivalent to an oracle estimator. An alternative characterization of the usual product-of-errors condition that $r_{\mu,n} r_{e,n} \ll n^{-1/2}$; this requirement is equivalent under somewhat weak overlap, but is more stringent under very weak overlap. Even this stronger product-of-errors requirement is insufficient for Wald confidence intervals constructed with thresholded AIPW to be valid.

corollary[Under very weak overlap, faster rates are necessary] For any overlap bound $\gamma_0 \in (1, 2)$, there exists a $\mathscr{P}$ and cross-fit nuisance estimators satisfying: \begin{enumerate} • The model is regular. $\mathscr{P}$ satisfies (ref) for this $\gamma_0$. • Classic product of errors. $\sup_{P} E_{P}\left[ \| \hat{\mu} - \mu \|_{\infty} \right] E_{P}\left[ \| \hat{e} - e \|_{\infty} \right] \ll n^{-1/2}$ and $b_n \gg E_{P}\left[ \| \hat{e} - e \|_{\infty} \right]$. • Wald inference fails. For any fixed target coverage level $\alpha \in (0, 1)$, $\sup_{P \in \mathscr{P}} P( \psi(P) \in \hat{\mathcal{C}}_n) \to 0$. \end{enumerate}

A heuristic sufficient condition for Wald confidence interval validity is $r_{\mu,n} r_{e,n}^{\min\{ \gamma_0, 2 \} / 2} \ll n^{-1/2}$. When $\gamma_0 < 2$, there is a range of nuisance estimates such that $r_{\mu,n} r_{e,n} \ll n^{-1/2} \ll r_{\mu,n} r_{e,n}^{\min\{ \gamma_0, 2 \} / 2} $ and estimation bias can be of a higher order than the oracle variance.

I calculate the black-box rate requirements under (ref)(ref) in a few special cases. I omit an analysis of the stronger rate requirement in (ref)(ref) that would be needed to handle degenerate distributions.

assumptionAssumptions (ref), (ref), and (ref)(ref) hold, and $r_{e,n}, r_{\mu,n} \to 0$.

I characterize the following special cases.

example[Somewhat weak overlap] Suppose (ref) holds, $\gamma_0 > 2$, and $r_{\mu,n} r_{e,n} \ll n^{-1/2}$. Then there exists a $b_n \to 0$ such that clipped AIPW t-statistics are asymptotically well-calibrated.
example[Second moments barely fail to exist] Suppose (ref) holds for $\gamma_0 = 2$ and there is some $\eta > 0$ such that $r_{\mu,n} r_{e,n} \log(1/r_{e,n}) \ll n^{-1/2}$. Then there exists a $b_n \to 0$ such that clipped AIPW t-statistics are asymptotically well-calibrated.
example[Shared rates, very weak overlap] Suppose (ref) holds for some $\gamma_0 > 1$ and $r_{\mu,n}, r_{e,n} \ll n^{-1/3}$. Then there exists a $b_n \to 0$ such that clipped AIPW t-statistics are asymptotically well-calibrated.
example[Parametric rates] Suppose (ref) holds for some $\gamma_0 > 1$ and either (i) $r_{\mu,n} = O(n^{-1/2})$ and $r_{e,n} = o(1)$ or (ii) $r_{e,n} = O(n^{-1/2})$ and $r_{\mu,n} = o(n^{(\gamma_0-2) / 4})$. Then there exists a $b_n \to 0$ such that clipped AIPW t-statistics are asymptotically well-calibrated.

I now unpack the black box and quantify the degree to which weak overlap makes a given outcome regression rate more difficult to achieve.

Necessary Smoothness Conditions

Weak overlap makes outcome regression more difficult: there may be few treated observations in precisely the regions in which thresholded AIPW depends most acutely on outcome regression. In this section, I show that for pointwise rates, even somewhat weak overlap can be viewed as degrading effective outcome smoothness for optimal local polynomial estimators. I also derive a blessing of weak overlap: the optimal global rate is equal to the optimal pointwise rate, without the usual polylogarithmic penalty.

I characterize optimal nonparametric regression rates under Hölder continuity. For convenience, I fix the covariates to be uniform over a specific hypercube in $\mathbb{R}^d$ with constant variance.

assumption[Hölder smoothness and fixed domain] $\mathscr{P}$ satisfies Assumptions (ref) and (ref)(ref), and for all $P \in \mathscr{P}$, $X \sim Unif([-1, 1]^d)$ and $Y \mid X, D \sim \mathcal{N}( D \mu_{P}(X) + (1-D) \mu_{P}'(X), \sigma_{\min}^2 )$, with $\mu, \mu'$ in the Hölder smoothness class $\Sigma(\beta_{\mu}, L)$ for some fixed $\beta_{\mu}, L > 0$.

Most of these assumptions are standard assumptions for studying local polynomial regression under strict overlap Stone1982. I assume normal outcomes in order to simplify the characterization of the optimal rate. It is known that in this case but under strict overlap, the optimal pointwise rate is $n^{\frac{-\beta_{\mu}}{2 \beta_{\mu} + d}}$, and the optimal global rate $\left( n / \log(n) \right)^{\frac{-\beta_{\mu}}{2 \beta_{\mu} + d}}$ has a polylogarithmic penalty, in the sense that the optimal global rate is worse by some polynomial factor of $\log(n)$.

I require that treated observations cannot concentrate in small regions.

assumption[Non-trivial concentration] There are parameters $\rho, \gamma > 0$ such that for all $h > 0$ small enough and all $P \in \mathscr{P}$ and $x_0 \in [-1, 1]^d$ $P( e(X) \geq \rho \sup_{\| x - x_0 \| \leq h} e(x) \mid D = 1, \| X - x_0 \| \leq h ) > \gamma$.

I show in Appendix (ref) that this condition holds if the propensity function is sufficiently smooth. When the propensity function is nonsmooth, it is possible for nature to concentrate treated observations within a given bandwidth in a region that is too small, introducing local polynomial degeneracy issues that are outside the scope of this work.

In the worst case, weak overlap of order $\gamma_0 > 1$ plays a role equivalent to scaling the effective outcome smoothness downward by $\left( 1 - 1 / \gamma_0 \right)$, but also removes the polylogarithmic factor in the optimal global rate.

theorem[Weak overlap reduces effective outcome smoothness] Define $\psi_n = n^{\frac{-\beta^*}{2 \beta^* + d} }$, where $\beta^* = \beta_{\mu} \left(1 - 1 / \gamma_0 \right)$. Then \begin{enumerate}[label=(\roman*)] • $\psi_n$ is a pointwise (and global) rate upper bound. There exists a $c > 0$ and a $\mathscr{P}$ satisfying Assumptions (ref) and (ref) such that $$\liminf_{n \to \infty} \inf_{\hat{\mu}} \sup_{P \in \mathscr{P}} P\left( |\hat{\mu}(0) - \mu(0)| > c \psi_n \right) > 0.$$$\psi_n$ is an achievable global (and pointwise) rate. Suppose $\mathscr{P}$ satisfies Assumptions (ref) and (ref). Then there exists an estimator $\hat{\mu}(x)$ such that for all $\epsilon > 0$, there is a finite $c(\epsilon)$ such that $$\limsup_{n \to \infty} \sup_{P \in \mathscr{P}} P\left( \| \hat{\mu} - \mu \|_{\infty} > c(\epsilon) \psi_n \right) \leq \epsilon.$$ Further, the estimator can be computed without knowledge of the overlap bound $\gamma_0$. \end{enumerate}

Recall that under strict overlap, the optimal pointwise convergence rate for a Hölder-smooth regression function is $n^{\frac{-\beta_{\mu}}{2 \beta_{\mu} + d}}$. (ref) shows that weak overlap has the effect of degrading the effective smoothness rate from $\beta_{\mu}$ to $\beta_{\mu} (1 - 1 / \gamma_0 )$. As overlap is allowed to become increasingly weak and other parameters are held constant, there can be regions of the covariate space with increasingly few treated observations so that the optimal pointwise regression rate is slower. This penalty occurs even under somewhat weak overlap. In the limit in which $\gamma_0$ tends to one, the rate in (ref) can become arbitrarily poor. For example, when $\gamma_0 = 2$, the difficulty of estimating a twice continuously differentiable function is comparable to the difficulty of estimating a Lipschitz-continuous function under strict overlap.

(ref) also presents a blessing of weak overlap: the optimal global rate is equal to the pointwise rate, with no polylogarithmic penalty. Usually, optimal global rates are worse than optimal pointwise rates by a polylogarithmic factor due to the need to have accuracy at a number of gridpoints that grows with $n$. Under weak overlap, the global rate still must be consistent at all gridpoints, but most gridpoints must satisfy strict overlap and have a negligible contribution to the global consistency rate. The challenge is to partition the remaining observations in a way which does not impose a polylogarithmic penalty. Naively partitioning the remaining observations into regions of increasing overlap with a constant penalty yields $\log(\log(n))$ partitions and a penalty potentially as large as $\log(\log(\log(n)))$. The proof of (ref) instead uses a more subtle construction, partitioning $[-1, 1]^d$ into observations with increasing overlap in such a way as each partition's worst-case global penalty is half as large as the previous worst-case penalty, even after accounting for the increasing number of gridpoints. Appendix (ref) shows that if smallest partition's penalty is sufficiently large, then this construction eventually includes all relevant observations. Then, summing over the penalties yields a global penalty that is large but bounded, so that the optimal pointwise rate is an achievable (and therefore optimal) global rate.

(ref) yields minimal smoothness assumptions for Wald confidence interval validity.

corollary[Minimal smoothness conditions] Suppose (ref) holds and there is a $\beta_{e} > 0$ such that $e(X) \in \Sigma(\beta_{e}, L)$ and \begin{align} \frac{\beta_{\mu}}{2 \beta_{\mu} + d \gamma_0 / (\gamma_0 - 1)} + \frac{\min\{ \gamma_0 / 2, 1\} \beta_e}{2 \beta_e + d} > 1 / 2. \end{align} Then there is a sequence of nuisance estimators and thresholds that are independent of $\gamma_0$ such that for all $\gamma_0 > 1$, the associated Wald confidence interval $\hat{\mathcal{C}}_n(\alpha)$ constructed using $\hat{\psi}_{clip}^{AIPW}(b_n)$ satisfies $$\limsup_{n \to \infty} \sup_{P \in \mathscr{P}} \left| P( \psi(P) \in \hat{\mathcal{C}}_n(\alpha)) - (1-\alpha) \right| = 0.$$

In one dimension, thresholded AIPW with Lipschitz-continuous conditional outcome mean and propensity function can handle weak overlap of order $\gamma_0 > \frac{4 + 1 / \beta_e}{3}$. In multiple dimensions, the econometrician must assume stronger smoothness restrictions than Lipschitz continuity in order to achieve the necessary nuisance rate guarantees under even strict overlap. When the propensity function is infinitely-differentiable, thresholded AIPW with a Lipschitz-continuous conditional outcome mean can handle weak overlap of order $\gamma_0 > \frac{2 (d + 1)}{d + 2}$. More generally, under very weak overlap, if $\beta_{\mu}, \beta_e > \frac{d \left( \sqrt{\gamma_0^2 + 4 \gamma_0 - 4} + 2 - \gamma_0 \right)}{4 (\gamma_0-1)}$, then it is feasible to achieve standard inference with thresholded AIPW.

This concludes the substantive theoretical analysis. In the next section, I use these results to infer lessons for empirical practice.

Lessons for Empirical Practice

This section leverages the theoretical analysis to consider misspecified parametric estimators and some rules of thumb for empirical use.

Parametric Estimators and Misspecification

When both nuisance functions are estimated nonparametrically, then consistency is achievable and AIPW is generally preferable under strict overlap. When both nuisance functions are estimated parametrically, then it is possible for one or both nuisance function to be inconsistent and the choice of estimator may be ambiguous. I now provide some intuition on the two estimators when nuisance functions are estimated parametrically and through cross-fitting. I will consider IPW and AIPW with the same sequence of thresholds $b_n$ satisfying $1 \gg b_n \gg n^{-1/2}$. I write that a nuisance estimate $\hat{\eta}$ is consistent if it tends to the correct limit $\eta$, and I write that $\hat{\eta}$ is inconsistent otherwise.

In this subsection, I will assume that parametric nuisance estimators $\hat{\eta}$ achieve an $L_\infty$ error relative to a limiting nuisance function $\bar{\eta}$ that is the order of $n^{-1/2}$. For example, consider logit estimation of a propensity model of the form $\bar{e}(X) = \frac{exp(X' \beta)}{1 + exp(X' \beta)}$ for a pseudo-true parameter $\beta$. If the support of $X$ is bounded, then $n^{-1/2}$-consistent estimate of $\beta$ is sufficient to achieve $n^{-1/2}$-consistent estimation of $\bar{e}(X)$ everywhere. However, weak overlap may emerge from unbounded tails, in which case the $L_\infty$ rate may not go to zero. Unbounded covariates are an important case in general. For example, XinweiMaRobustIPW motivate weak overlap tails through the distribution of covariates under a logistic propensity model. Nevertheless, a careful treatment of parametric estimation of nuisances with unbounded covariates is outside the scope of this work.

The analysis above is easiest to extend when either both or neither nuisance function is consistent. If both nuisance estimates are consistent, then the AIPW and IPW estimators will be consistent and will have variance on the same order, but the IPW estimator may have higher-order bias than the AIPW estimator. This higher-order bias follows because IPW can be viewed as a particular case of AIPW with an inconsistent outcome regression estimator. If both the propensity and outcome regression estimates are inconsistent, then both the IPW and AIPW estimators fail to be consistent, and as in the case of inconsistent nuisance functions with strict overlap, there is no general reason to prefer one or the other.

When the outcome regression estimate is inconsistent, there is no general reason to prefer IPW or AIPW, but both estimators may have bias that is of a higher-order than the estimator's standard deviation. When $\hat{\mu}$ is inconsistent, both IPW and AIPW can be viewed as instances of AIPW with an inconsistent outcome regression estimate. Suppose $P$ is a distribution from the second half of (ref), which has $P( e(X) \leq \pi ) \sim \pi^{\gamma_0 - 1}$ for all $\pi$ small enough. The bias in the thresholded region with an inconsistent outcome regression estimate is generally on the order of $P(e(X) \leq b_n) \sim b_n^{\gamma_0 - 1}$. However, by (ref), the oracle AIPW (and oracle IPW) standard deviation is on the order of $n^{-1/2} b_n^{\gamma_0 / 2 - 1} \ll b_n^{\gamma_0 - 1}$. This heuristic analysis suggests that in many cases, IPW or AIPW-with-inconsistent-outcome-regression will have bias that is of a higher order than the estimator's standard error. That intuition is similar to ma2023doubly's analysis of trimmed AIPW with a tailored debiasing procedure.

The case of a consistent outcome regression estimate with inconsistent propensity estimates is more interesting. In this case, AIPW should have lower-order bias than IPW, because IPW will be inconsistent. The Berry-Esseen argument for AIPW asymptotic normality with known nuisance functions only requires cross-fitting and $b_n$ to go to zero slower than $n^{-1/2}$, so that thresholded AIPW should also be asymptotically normal under appropriate error product conditions. However, it is unclear how the bias compares to sampling error. In any event, this robustness intuition is useful, because I apply parametric nuisance estimators in the application to right heart catheterization. Careful treatment of the parametric case is left for future work.

Before proceeding to apply clipped AIPW, I derive some rules of thumb for choosing a threshold.

Choice of Threshold

The theoretical analysis above provides conditions under which there is some sequence of thresholds for which AIPW is asymptotically normal and centered around the true causal estimand. I now provide guidance for how to choose the threshold.

algorithm[algorithm omitted — 1,803 chars of source]

I propose different rules of thumb based on whether the econometrician is willing to provide an upper bound on the rate of convergence for the propensity estimate, the outcome regression estimate, or both. The combined proposal is presented in Algorithm (ref).

In practice, it is often relatively easy to identify an upper bound on the propensity rate of convergence. For instance, if the propensity score is estimated with local polynomial regression of order $\ell_e$, then the econometrician is implicitly asserting a Hölder smoothness of some order $\beta_e > \ell_e$, and a feasible consistency rate of $(n / \log(n))^{-\beta_e / (2 \beta_e + d)}$ Stone1982. In this case, the econometrician can safely conjecture that if their estimator is well-founded, then it will achieve a global consistency rate $r_{e,n} \ll n^{-\ell_e / (2 \ell_e + d)}$, and a threshold of $b_n = n^{-\ell_e / (2 \ell_e + d)} = ``\bar{r}_{e,n}"$. This rule of thumb is practical and the safest rule with respect to outcome regression that does not impose stronger requirements on the propensity estimate, but also often corresponds to a relatively slow consistency rate with relatively little power.

Three alternative rules of thumb target faster consistency rates and laxer propensity requirements. While the main text focuses on black-box requirements on consistency rates in terms $\gamma_0$, the proof goes through showing $b_n$ tends to zero slower than $n^{-1/2}$, $r_{e,n} \ll b_n$, and

align*[align* omitted — 259 chars of source]

In these alternative cases, I propose replacing $b_n$ in the right-hand side with the putative threshold $b$ and $r_{\mu,n}$ and $r_{e,n}$ in the left-hand side with predicted upper bounds $\bar{r}_{\mu,n}$ and $\bar{r}_{e,n}$ where feasible and with the putative threshold $b$ otherwise. This yields an empirical function error_bound($b$). I then solve for the putative threshold $b$ where error_bound($b$) crosses $n^{-1/2}$. The result is a threshold $b_n$ aimed to target maximal efficiency and minimal propensity consistency requirements, calculated without use of the outcome data or direct knowledge of $\gamma_0$, and with the property that if the nuisance errors go to zero faster than the specified upper bound (or estimated threshold), then the resulting Wald confidence intervals will be well-calibrated. An interesting avenue for future work is whether there is a convenient choice of context-dependent constant multiples in the error_bound function.

This rule of thumb is also always feasible, for example in the case of nonspecified upper bounds.

lemma[Well-defined rule of thumb] Suppose $\hat{e} \in (0, 1]$ and $\sum D / \hat{e} > 0$. Let $f_n(b)$ be the error_bound function with no upper bound nuisance rates given. Then there is exactly one $b_n$ such that $\limsup_{b \to b_n^-} f_n(b) \leq 0 \leq \liminf_{b \to b_n^+} f_n(b_n)$.

A rule of thumb with nonspecified rates seems particularly attractive, since it is an entirely data-driven way to choose the AIPW threshold. However, in practice, a given outcome regression rate is more difficult to achieve than a given propensity rate under weak overlap, so that rules of thumb with specified nuisance rates is more appropriate for nonparametric outcome regression estimates.

Applications

In this section, I present simulated results for the clipped AIPW estimator as well as empirical results from an application to right heart catheterization. I find that clipped AIPW performs well asymptotically, producing near-perfect calibration of p-values with 100,000 observations, but exhibits some undercoverage in small samples. When studying the right heart catheterization data, I find that the rule of thumb approach increases the estimated harm of the procedure by 0.17 standard errors relative to the usual 10% trimming rule, while increasing the estimated standard error by 5.1%.

Simulation Evidence

I now study the performance of the clipped AIPW estimator in simulations.

My simulation design is based on the design in XinweiMaRobustIPW. As in their work, I simulate data with $P(e(X) \leq \pi) = \pi^{\gamma - 1}$ and $D Y = \kappa D (1-e(X)) + D (\varepsilon - 4) / \sqrt{8}$, where $\varepsilon \mid X, D \sim \xi_4^2$ is scaled to achieve zero mean and unit variance. However, I increase $\gamma$ from $1.5$ to $1.8$ to ensure feasible outcome regression rates, set $\kappa = 2$ rather than $\kappa = 1$ to avoid coincidental offsetting bias of IPW lower and upper tails in small samples, and reduce $D Y$ by $\kappa E[ D (1-e(X)) ]$ so that the true average potential outcome is zero. I achieve this propensity distribution by taking $X \sim Unif([0, 1])$ i.i.d. and setting $e(X) = X^{1/(\gamma_0-1)}$. I present results for 5,000 simulations of increasingly large samples.

I estimate both the propensity and outcome regressions with five-fold cross-fitting. I use shrinkage cubic splines and REML estimation, as implemented by the mgcv package in R. In this setting, (ref) establishes that kernel regression can achieve a pointwise rate of $n^{-1 / (3 + 1 / (\gamma_0-1))}$. I conjecture that $r_{\mu,n} \ll n^{-1/5}$, which is feasible if $\gamma_0 > 1.5$, and choose the clipping threshold $b_n$ based on Algorithm (ref).

figure[figure omitted — 478 chars of source]

I begin by summarizing point estimates in (ref). The unthresholded estimators are approximately median-unbiased, but possess sufficiently heavy inverse propensity tails that the mean performance degrades with increasing sample size. The clipped estimators perform much better, but the clipped IPW estimator exhibits its known first-order bias. The clipped AIPW estimator exhibits less bias than the clipped IPW estimator, and has slightly better performance in terms of mean squared error.

figure[figure omitted — 446 chars of source]

I find in (ref) that the clipped AIPW estimator's t-statistics are reasonably well-calibrated. The plot presents t-statistics on the true average potential outcome. The t-statistics of unthresholded IPW and AIPW estimators are visibly non-Gaussian, and often exhibit a multimodal distribution. This poor performance is unsurprising: unthresholded estimators are known to fail to be asymptotically normal in this setting. Both thresholded estimators are known to be asymptotically normal in this setting when the propensity score is known, and both the asymptotic normality and the clipped IPW estimator's first-order bias are visible to the naked eye, although the clipped IPW estimator also exhibits visible skew in small samples. I test for t-statistic normality using a Shapiro-Wilk test. The test rejects normality for both clipped estimates. Still, the clipped AIPW estimator's violations are less severe by this criterion.

figure[figure omitted — 408 chars of source]

I find in (ref) that the clipped AIPW estimator's p-values are well-calibrated in large samples. I use Wald confidence intervals to calculate two-sided p-values on the null of the true average potential outcome. If Wald confidence intervals are well-calibrated, then the simulated p-values on the true average potential outcome will be exactly uniformly distributed. The unthresholded IPW and AIPW estimators exhibit known poor performance. The clipped IPW estimator exhibits overrejection even with large samples, as even oracle clipped IPW would provide well-calibrated inference for a biased estimand. The clipped AIPW estimator also overrejects in small samples, but the bias is less severe: with 1,000 observations, clipped IPW rejects the true null in 12.0% of simulations, while clipped AIPW rejects in 8.8% of simulations. As the sample size increases, the asymptotic calibration of (ref) becomes apparent. With 100,000 observations, clipped IPW rejects the true null hypothesis in 12.8% of simulations, while clipped AIPW rejects in 5.3% of simulations. The Kolmogorov-Smirnov p-value on exact calibration of the two-sided test statistics for clipped AIPW with 100,000 observations is 0.692. This is a remarkable result: despite the known extreme difficulty of statistical inference in this setting, 5,000 simulated draws are insufficient to detect a meaningful failure of Wald confidence intervals based on the clipped AIPW estimator.

figure[figure omitted — 544 chars of source]

In moderate samples, clipped AIPW can undercover due to the difficulty of outcome regression in this setting. (ref) presents an example with 1,000 observations. It is rare to have treated observations with small values of $e(X)$. As a result, when such observations are treated, a small number of observations can receive substantial leverage in outcome regression, and the predictions of $E[Y \mid X = 0, D = 1]$ can be driven by a small number of observations. In (ref) (Figures (ref) through (ref)), I conduct the same experiments, but with the estimated outcome regression function replaced by the true outcome regression function. The root-mean-squared error and failures of normality are comparable, suggesting these non-inferential patterns are driven by propensity estimation and clipping. However, the two-sided p-values exhibit better performance in small samples, and if anything slightly underreject with 100,000 observations.

In (ref) (Figures (ref) through (ref)), I show that these conclusions would largely carry through if clipping were replaced by trimming. The notable differences are that trimmed AIPW exhibits slightly better estimation performance in small samples, while if anything trimmed IPW is slightly worse; trimmed t-statistics exhibit less severe violations of normality; and p-values based on trimmed propensities exhibit more severe undercoverage for both IPW and AIPW.

Application to Right Heart Catheterization

I apply the clipped AIPW estimator to study the effect of right-heart cathterization (RHC) on survival. This dataset was first analyzed by ConnorsEtAl1996, and is a common benchmark in the weak overlap literature CrumpEtAlOptimizePrecision, ArmstrongKolesarBandwidthSnooping.

I analyze a version of the dataset from ArmstrongKolesarBandwidthSnooping. The dataset is comprised of 5{,}735 adult patients, and the treatment $D$ corresponds to receiving RHC within 24 hours of admission. The target causal effect is the average treatment effect of RHC on 30-day survival. The data includes 52 covariates $X$ (72 covariates if counting factor levels separately). I estimate the nuisance functions $e(X)$ and $\mu(X)$ using five-fold cross-fitting. I estimate nuisance functions with logistic regression to align with CrumpEtAlOptimizePrecision's empirical application. I estimate standard errors by bootstrapping the procedure. I keep fold assignment fixed in bootstraps to minimize the risk of over-fitting.

CrumpEtAlOptimizePrecision propose a weak overlap rule of thumb that estimates the treatment effect for the subpopulation with propensity scores between 10% and 90%. This rule-of-thumb trimming rule is chosen to approximately minimize asymptotic variance. This strategy ensures asymptotic normality, but changes the target estimand even asymptotically. By comparison, the clipped and trimmed AIPW estimators I analyze have thresholds $b_n$ that tend to zero asymptotically. As a result, the estimators proposed here are able to target full population average treatment effect, potentially at the cost of increased variance. I compare these procedures to the 10% rule and other fixed trimming rules using the same nuisance estimates.

figure[figure omitted — 431 chars of source]

I present the distribution of estimated propensity scores for treated and control units in (ref). The figure is an analog of CrumpEtAlOptimizePrecision's Figure 1. There is a meaningful density of units with estimated propensities near zero, suggesting weak overlap. This pattern is similar to the findings of CrumpEtAlOptimizePrecision, although there are slight differences, presumably due to my use of cross-fitting.

figure[figure omitted — 668 chars of source]

I compare AIPW estimators for various trimmed subsamples to the clipped AIPW estimator. I choose the clipping threshold $b_n$ through the no-specified-upper-bound version of Algorithm (ref) because I estimate both nuisance functions parametrically. I plot the functions used in choosing $b_n$ in (ref). The estimated lower clipping threshold is 0.068 and affects 10.5% of observations. The CrumpEtAlOptimizePrecision 10% rule of thumb would exclude 16.3% of observations below. The estimated upper clipping threshold is 0.09 below one: there are few observations with large estimated propensities, so the rule of thumb concludes there is no need to trim observations with large estimated propensities. This upper threshold affects 1.4% of observations, comparable to the 1.8% of observations excluded above by the 10% rule of thumb.

figure[figure omitted — 690 chars of source]

I present estimated effects and confidence intervals for various potential fixed trimming rules in (ref). The 10% trimming rule yields an estimated reduction in survival rates of 5.79 percentage points among the trimmed sample, with an estimated 95% Wald confidence interval of [-9.14, -2.43]. Other trimming rules would yield larger confidence intervals, as expected because the 10% rule is chosen to roughly minimize asymptotic variance over target populations.

I compare the fixed-trimmed-sample AIPW estimates to a clipped AIPW estimator that targets the full population treatment effect. The estimated harm increases to -6.07 percentage points, a change of 0.168 standard errors under the 10% rule of thumb estimator. The clipped AIPW confidence interval of [-9.6, -2.55], has a 5.14% larger width than the 10% trimmed sample interval. The clipped AIPW point estimates are similar to the point estimates under a 1% or 5% trimming rule, but the associated confidence interval is narrower under the full-population estimator. Part of the added width is driven by inverse propensities among clipped observations: if I used a trimmed, rather than clipped, AIPW estimator, the estimated effect would move by 0.256 standard errors, and the standard error would only increase by 0.54%. However, the simulation results of (ref) suggest that trimmed AIPW may slightly undercover.

Taken together, these results illustrate that under weak overlap, targeting the causal effect within the full population need not come at a large precision cost. In this application, clipped AIPW with a rule-of-thumb clipping rate yields similar estimates to estimators that target a fixed trimmed sample, while targeting a population that is often more relevant and adding only a small precision cost.

Conclusion

This work shows that standard Wald confidence intervals for clipped AIPW can achieve target coverage for standard causal effects under plausible conditions. I provide sufficient conditions on nuisance regression rates for clipped (or trimmed) AIPW to be uniformly valid over distributions with even very weak overlap. I use these theoretical results to derive new rules of thumb for choosing a threshold. I find that Wald confidence intervals perform well in simulations, especially in large samples, and can achieve comparable precision to a fixed 10% trimming rule in practice.

These results can be extended in many interesting directions. This work exploits Neyman orthogonality to achieve standard statistical inference in the presence of a small region of irregular identification. SasakiUra2022 and ma2023doubly propose estimators for ratio estimands beyond IPW; the arguments here are likely to extend to their more general framework. Issues of weak overlap hold for inverse propensity and other importance sampling estimators in settings like difference-in-difference estimation CallawaySantAnna or statistical inference for parameters that are identified at infinity AndrewsSchafgans, KhanNekipelov; the results and rules of thumb here can likely be adapted to those settings. semenova2024aggregatedintersectionboundsaggregated applies thresholding strategies to intersection bounds, where at a high level a margin condition plays the role of the minimal overlap bound here. Perhaps similar ideas could apply to other forms of irregular identification. The regression analysis in (ref) may also extend to estimating the effects of continuous treatments.

The results here suggest that thresholded AIPW is a viable alternative to fixed-trimming rules. I provide rules of thumb that enable practitioners to easily report results that target the population average effect. When, as in my empirical application, the fixed-trimming and sequence-of-thresholds approaches yield similar causal conclusions, then there is strong evidence that causal conclusions are driven by causal effects, and not how the researcher treats observations with extreme propensity scores.