EconBase
← Back to paper

Doubly Robust Estimators with Weak Overlap

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,499 characters · 0 sections · 45 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Doubly Robust Estimators with Weak Overlap

bibunit\begin{abstract} Doubly robust (DR) estimators guard against model misspecification but remain sensitive to weak covariate overlap. We show that trimming propensity scores reduces variance but eliminates double robustness. We introduce DR estimators that retain double robustness after trimming through bias correction, preserving the original causal targets across unconfoundedness, instrumental variables, and difference-in-differences designs. In four applications, the proposed estimator yields more precise estimates: ruling out large mortality effects of Medicaid expansion, detecting workforce growth from mental health reform, recovering the Black--White test score gap without strong functional form restrictions, and recovering a positive 401(k) savings effect consistent with the prior literature. JEL Codes: C10, C14, C21, C23, C26. \\ Keywords: Difference-in-Differences; LATE; Unconfoundedness; Instrumental Variables; Propensity Score; Trimming; Bias Correction. \end{abstract} \allowdisplaybreaks \section{Introduction} Accounting for confounders is essential for reliable observational causal inference, but doing so poses practical challenges due to model misspecification and weak covariate overlap. Doubly robust (DR) estimators address the first concern: they remain consistent for a prespecified causal estimand even when one component of the working model is incorrect Robins1994.\footnote{See Sloczynski2018 for general identification results with doubly robustness in the causal inference.} Yet DR estimators remain vulnerable to the second problem: when propensity scores are close to zero or one, they can become unreliable. A handful of observations with extreme inverse probability weights can dominate the entire estimate, inflating standard errors by factors of two to six, or worse, flipping the sign of point estimates. Weak overlap does not merely reduce precision; it can render observational studies uninformative or misleading. As we highlight in different applications, this concern is indeed warranted. For instance, we find that the standard DR 95% simultaneous confidence bands in a staggered difference-in-differences (DiD) analysis of Medicaid expansion on mortality span on average 68 deaths per 100,000 people over the post-treatment horizons---too wide to inform policy. In a 401(k) savings application, the standard DR local average treatment effect (LATE) estimate implies that participation reduces net financial assets by \$13,042---a conclusion at odds with the prior literature---because two ineligible observations with extreme instrument propensity scores carry effective weights of roughly 31 and 98 each, dominating the remaining observations. Removing the direct influence of these observations recovers a sensible point estimate; the bias correction we develop below allows this trimmed estimator to target the original LATE with valid inference, rather than a redefined trimmed estimand. A natural first response is to trim observations with extreme propensity scores. Although intuitive, we highlight that trimming a DR estimator eliminates its double robustness, the very property that justified using DR methods in the first place. We demonstrate this via simulations: when the outcome regression is misspecified but the propensity score is correct, the standard trimmed DR estimator should remain consistent, yet its 95% coverage collapses to 34% at a trimming threshold of $h = 0.10$ (Section (ref)). The double robustness guarantee is gone. The alternatives to trimming also come with important caveats. Changing the target parameter of interest CrumpEtAl2006_WP_MovingGoalposts, li2018balancing, Sloczynski2022_Restat,Sloczynski2026_Restud,Blandhol_etal_2026_Restud addresses the variance problem by redefining the causal question being answered---paraphrasing the song When I'm Not Near the Girl I Love, “When I'm not near the parameter I love, I love the parameter I'm near.”\footnote{Our understanding is that Art Goldberger is the source of this paraphrase. We thank Matt Masten for bringing this to our attention.} A prominent example is imposing a linear specification and interpreting the regression treatment coefficient as a weighted average of heterogeneous treatment effects; the resulting estimand can have unattractive properties such as placing negative weight on some subpopulations and being hard to motivate via policy justification, and must be derived on a case-by-case (and design-by-design) basis Sloczynski2022_Restat, Sloczynski2026_Restud, Mogstad_Torgovitsky_2024, Caetano_Callaway_2024. This paper shows that robustness to weak covariate overlap need not come at the cost of redefining the causal estimand. We make three contributions. First, we establish a negative result: trimming without bias correction eliminates the double robustness property of DR estimators. The trimmed estimator is consistent only if the outcome regression is correctly specified, losing robustness to the very misspecification it was designed to guard against (Section (ref)). Second, we propose a class of bias-corrected DR estimators that is also robust against weak overlap. We show that a key condition for bias correction---that the conditional mean of the quantity being averaged vanishes as the propensity score approaches its boundary---holds under either route to double robustness. When the propensity score is correctly specified, no observations of the relevant type exist at the boundary; when the outcome regression is correctly specified, the conditional expectation of the DR residual is zero everywhere. Either way, the conditional mean of the ratio has a well-defined limit at the boundary, and a low-dimensional polynomial can reconstruct the contribution of the trimmed observations, so the bias correction preserves double robustness. Third, we highlight that our unified framework covers the ATE, LATE, and staggered DiD designs and derive a closed-form influence function that enables standard inference. Implementation adds a single polynomial regression step to any existing DR procedure. Rich conditioning sets are often needed for identification---interviewer fixed effects for the test score gap, county characteristics for parallel trends, income controls for instrument validity---but the same covariates that make identification credible can make estimation infeasible through weak overlap damour2021overlap. Our estimator addresses this tension: researchers can condition on whatever covariates identification requires while substantially reducing the variance penalty. Four empirical applications---two DiD, one under unconfoundedness, and one using instrumental variables---demonstrate that the proposed estimator produces tighter confidence intervals and recovers informative estimates: in a DiD analysis of Medicaid expansion, the 95% simultaneous confidence bands narrow from roughly 68 to 17 deaths per 100,000 on average over the post-treatment horizons; in an IV analysis of 401(k) participation, the estimator recovers a positive and significant savings effect of \$8,864, consistent with prior findings; and under unconfoundedness, it detects the Black--White test score gap at age two more precisely than linear regressions while directly targeting the ATE without imposing linearity. Across the four applications, standard errors from existing DR methods are 1.8 to 6.6 times larger than those from our proposed doubly robust bias-corrected (DR-BC) method. Since reducing the standard error by a factor of $c$ translates into a $c^2$-fold reduction in the sample size needed to detect a given effect at conventional power, these gains expand what applied researchers can learn from existing datasets. Our approach builds on and extends several strands of work on causal inference under weak overlap crump2009dealing,khan/tamer:2010.\footnote{See sasaki2018estimation for more discussions on this literature.} We are mostly related to yang2018asymptotic and heiler2019valid, who study DR-type estimators in unconfoundedness settings; however, yang2018asymptotic relies on trimming strategies that redefine the estimand, while heiler2019valid does not consider trimming at all. None of these papers considers more general research designs as we do. sasaki2018estimation propose bias-corrected inverse probability weighting estimators for a single ratio moment under weak overlap, relying on the boundary condition when the propensity score is correctly specified. Our framework differs in that it accommodates doubly robust estimands, establishes boundary conditions under both DR routes, and handles the nonlinear multi-moment structure required for the ATE, LATE, and DiD designs. In particular, the H\'ajek-type normalization that is standard in causal inference creates a denominator dependence across moments---requiring us to jointly bias-correct the normalization moments---and the residual-versus-normalization moment asymmetry we exploit under the outcome-correct route has no analog in the unnormalized, single-moment setting of sasaki2018estimation. See, e.g., sloczynski2024abadie for a recent discussion of why H\'ajek-type estimators are predominant in causal inference problems. \textbf{Notation.} We write $\mathop{}\!\mathbb{E}[\cdot]$ for the expectation operator and $\mathbb{E}_{n}\left[\cdot\right]=n^{-1}\sum_{i=1}^n(\cdot)_i$ for the sample average. The indicator function is $\mathbbm{1}\{ \cdot \}$. For a nuisance parameter $\gamma$, we let $\gamma_0$ denote the true value, $\hat\gamma$ an estimator, and $\gamma^*$ the probability limit of $\hat\gamma$. \section{Overview and practical relevance} This section introduces the setup, demonstrates that trimming compromises double robustness, presents four empirical applications, and describes the proposed estimator. \subsection{Setup and DR estimands} Our framework covers different research designs commonly used in empirical work, including unconfoundedness (selection on observables), LATE (instrumental variables), and DiD. Although the notation and identification assumptions differ across designs, they share a common structure: each involves a propensity score that can take values near its boundary, leading to weak overlap, and each admits a DR estimand that combines an outcome regression with inverse probability weighting. We present the DR estimands for these three designs using normalized weights that sum to one within each treatment group. \subsubsection{Unconfoundedness} A common goal in observational studies is to estimate the average effect of a binary treatment on an outcome, while adjusting for pre-treatment covariates to make the identification assumptions more plausible. Let $Y(1)$ and $Y(0)$ respectively denote the potential outcomes with and without treatment, $D \in \{0,1\}$ the treatment indicator, and $X$ a vector of pre-treatment confounding variables. The propensity score is $p_D(X) = \mathop{}\!\mathbb{E}[D \mid X]$, and the outcome regression is $m_Y(d, X) = \mathop{}\!\mathbb{E}[Y \mid D = d, X]$. Under unconfoundedness ($\{Y(0), Y(1)\} \mathbin{ \mathpalette{\@indep}{} } D \mid X$) and overlap ($0 < p_D(X) < 1$), the $\text{ATE} = \mathop{}\!\mathbb{E}[Y(1) - Y(0)]$ is identified and admits the following DR representation Sloczynski2018 \begin{equation} \resizebox{0.92\linewidth}{!}{$\displaystyle \text{ATE}_{\text{DR}} = \mathop{}\!\mathbb{E}\left[m_Y(1,X)-m_Y(0,X)+w_{D=1}^{\text{ATE}}(D,X)\left(Y - m_Y(1,X)\right) - w_{D=0}^{\text{ATE}}(D,X)\left(Y - m_Y(0,X)\right)\right], $} \end{equation} where \begin{align} w_{D=1}^{\text{ATE}}(D,X) = \frac{D/p_D(X)}{\mathop\!\mathbb{E}\left[D/p_D(X)\right]}, \quad w_{D=0}^{\text{ATE}}(D,X) = \frac{(1-D)/(1-p_D(X))}{\mathop\!\mathbb{E}\left[(1-D)/(1-p_D(X))\right]}. \end{align} When $p_D(X)$ is close to zero or one for some values of $X$, the weights in (ref) can become large. This is the weak overlap problem. Standard two-step estimators for (ref) do not achieve the $\sqrt{n}$-consistency in such weak-overlap cases khan/tamer:2010. \subsubsection{Instrumental variables and LATE} When treatment is endogenous, and unconfoundedness is not empirically plausible, researchers often use an instrument to identify the average treatment effect among units whose treatment status is shifted by the instrument, i.e., the compliers imbens/angrist:1994. Consider a setting with a binary instrument $Z$, binary treatment $D$, and observed covariates $X$, with potential treatments $D(1)$ and $D(0)$ and potential outcomes $Y(1), Y(0)$. Define the instrument propensity score as $p_Z(X) = \mathop{}\!\mathbb{E}[Z \mid X]$, the reduced-form and first-stage regressions as $m_{Y}^{\text{LATE}}(z, X) = \mathop{}\!\mathbb{E}[Y \mid Z = z, X]$ and $m_{D}^{\text{LATE}}(z, X) = \mathop{}\!\mathbb{E}[D \mid Z = z, X]$, respectively. Under the standard IV-LATE assumptions, the local average treatment effect $\text{LATE} = \mathop{}\!\mathbb{E}[Y(1) - Y(0) \mid D(1) > D(0)]$ is identified and the DR estimand for the LATE Belloni2017, Sloczynski2022, sloczynski2024abadie can be written as \begin{equation} \resizebox{0.92\linewidth}{!}{$\displaystyle \text{LATE}_{\text{DR}} = \frac{\mathop{}\!\mathbb{E}\left[m_Y^{\text{LATE}}(1,X)-m_Y^{\text{LATE}}(0,X)+w_{Z=1}^{\text{LATE}}(Z,X)\left(Y - m_Y^{\text{LATE}}(1,X)\right) - w_{Z=0}^{\text{LATE}}(Z,X)\left(Y - m_Y^{\text{LATE}}(0,X)\right)\right]} {\mathop{}\!\mathbb{E}\left[m_D^{\text{LATE}}(1,X)-m_D^{\text{LATE}}(0,X)+w_{Z=1}^{\text{LATE}}(Z,X)\left(D - m_D^{\text{LATE}}(1,X)\right) - w_{Z=0}^{\text{LATE}}(Z,X)\left(D - m_D^{\text{LATE}}(0,X)\right)\right]} $} \end{equation} where \begin{align*} w_{Z=1}^{\text{LATE}}(Z,X) = \frac{Z/p_Z(X)}{\mathop\!\mathbb{E}\left[Z/p_Z(X)\right]}, \quad w_{Z=0}^{\text{LATE}}(Z,X) = \frac{(1-Z)/(1-p_Z(X))}{\mathop\!\mathbb{E}\left[(1-Z)/(1-p_Z(X))\right]}. \end{align*} The numerator identifies the reduced-form effect and the denominator the compliance share, so the ratio recovers the LATE. Weak overlap arises when $p_Z(X)$ is close to zero or one, potentially implying a high variance of two-step estimators based on (ref). \subsubsection{Difference-in-Differences} Another widely used research design that accommodates selection on unobservables is the DiD design. Consider a staggered DiD panel data setup with $T$ time periods Callaway2021. Let $G$ denote the period of first treatment ($G = \infty$ if never treated), $X$ be a vector of pre-treatment covariates, and let $Y_t(g)$ be the potential outcome in period $t$ if a unit was first treated in period $g$. Let $\delta \geq 0$ denote the number of anticipation periods. Under limited anticipation and a conditional parallel trends assumption Callaway2021, the group-time average treatment effect $\text{ATT}(g,t) = \mathop{}\!\mathbb{E}[Y_t(g) - Y_t(\infty) \mid G = g]$ is identified for $t \geq g - \delta$. The more aggregated event-study estimand \begin{align} \text{ES}(e) = \sum_{g} w_{g,e}^{\text{es}} \cdot \text{ATT}(g, g+e), \end{align} where $w_{g,e}^{\text{es}} = \mathop{}\!\mathbb{P}(G = g \mid G + e \leq T, G \leq T)$ is also identified. Estimation of each $\text{ATT}(g,t)$ relies on a comparison group $\mathcal{C}_{g,t}$ (either never-treated or not-yet-treated units) with indicator $C_{g,t}$, generalized propensity score $p_{g,t}(X) = \mathop{}\!\mathbb{P}(G = g \mid X, \mathbbm{1}\{ G=g \} + C_{g,t} = 1)$, and outcome regression $m_{g,t}(X) = \mathop{}\!\mathbb{E}[Y_t - Y_{g-\delta-1} \mid X, C_{g,t} = 1]$. The DR estimand for the $\text{ATT}(g,t)$ parameters is given by \begin{align} \text{ATT}_{\text{DR}}(g,t) = \mathop\!\mathbb{E}\left[\left(w_{G=g}^{\text{DiD}}(G) - w_{g,t}^{\text{DiD}}(G,X)\right)\left(Y_t-Y_{g-\delta-1} -m_{g,t}(X)\right)\right], \end{align} with weights given by \begin{align*} w_{G=g}^{\text{DiD}}(G) = \frac{\mathbbm{1}\{ G=g \}}{\mathop\!\mathbb{E}\left[\mathbbm{1}\{ G=g \}\right]}, \quad w_{g,t}^{\text{DiD}}(G,X) = \frac{C_{g,t}\,p_{g,t}(X)/(1-p_{g,t}(X))}{\mathop\!\mathbb{E}\left[C_{g,t}\,p_{g,t}(X)/(1-p_{g,t}(X))\right]}. \end{align*} DR estimands for $\text{ES}(e)$ are defined by replacing $\text{ATT}(g, g+e)$ in (ref) with $\text{ATT}_{\text{DR}}(g,g+e)$ as defined in (ref). Weak overlap arises when $p_{g,t}(X)$ is close to one, potentially implying a high variance of DiD and ES estimators based on (ref). \subsubsection{The common challenge} Across all three designs, the DR estimand involves weights that divide by either the propensity score or one minus the propensity score (for the ATE, both treated and control weights can be extreme; for the LATE, both instrument groups; for DiD, only the comparison-group weights). When any such denominator is near zero, the resulting weights become very large. A handful of such observations can dominate the entire estimate. Trimming them reduces variance but introduces bias for the original estimand. As we show next, trimming a standard DR estimator can \emph{eliminate its double robustness property}. \subsection{Trimming compromises double robustness} A key appeal of the DR estimands presented above is their robustness to model misspecification: estimators based on them remain consistent for the causal parameter of interest even if one of the two working models (the propensity score or the outcome regression) is misspecified. A natural question is: Does this property survive when one trims observations with extreme propensity scores? The answer is no, as we illustrate via a simulation exercise. We illustrate this with a DiD setup similar to Kang2007 and sant2020doubly, with $n = 10{,}000$, a true ATT of zero, and weak overlap by design. To assess double robustness, we consider two complementary configurations: the propensity score is misspecified while the outcome regression is correct, and vice versa. We repeat each exercise $10{,}000$ times; full details of the data-generating process are in Appendix (ref), which also reports results for additional configurations (including correct specification of both models)---all yield qualitatively similar conclusions. We use the standard DR DiD estimator based on (ref), with logit for the propensity score and OLS for the outcome regression. We vary the trimming threshold over $h \in \{0, 0.01, 0.025, 0.05, 0.10\}$, where $h=0$ is the no-trimming case. Figure (ref) summarizes the results, showing ridge density plots alongside zipper confidence interval plots (black segments indicate non-coverage). \begin{figure}[htp!] \resizebox{0.93\textwidth}{!}{ \begin{minipage}{\textwidth} \begin{subfigure}[t]{.48\textwidth} \caption{PS misspecified -- Standard trimmed DR} \end{subfigure} \begin{subfigure}[t]{.48\textwidth} \caption{PS misspecified -- DR-BC (our proposal)} \end{subfigure}\\[0.3cm] \begin{subfigure}[t]{.48\textwidth} \centering \includegraphics[width = \textwidth]{Figures/Simulations/figure1_panelA_dgp3.pdf} \caption{Outcome misspecified -- Standard trimmed DR} \label{fig:sim_trim_dgp3} \end{subfigure} \hfill \begin{subfigure}[t]{.48\textwidth} \centering \includegraphics[width = \textwidth]{Figures/Simulations/figure1_panelB_dgp3.pdf} \caption{Outcome misspecified -- DR-BC (our proposal)} \label{fig:sim_bc_dgp3} \end{subfigure} \end{minipage} } \vspace{-0.2cm} \caption{Effect of trimming on DR DiD estimators under model misspecification} \label{fig:Sims_DiD} \justifying \vspace{0.1cm}\noindent\scriptsize{\textit{Notes.} Simulation design as discussed in Appendix \ref{appendix:simulation_dgp}, with $n=10{,}000$ and $10{,}000$ Monte Carlo repetitions. The true ATT is zero. Top row (Panels~(a)--(b)): propensity score misspecified, outcome regression correct. Bottom row (Panels~(c)--(d)): propensity score correct, outcome regression misspecified. In both cases, double robustness should ensure consistent estimation. Left column (Panels~(a),~(c)): standard trimmed DR estimator. Right column (Panels~(b),~(d)): our bias-corrected estimator (DR-BC). Ridge plots show the density of point estimates; zipper plots show 95\% confidence intervals across $10{,}000$ Monte Carlo draws, with black segments indicating non-coverage.} \end{figure} The left column of Figure~\ref{fig:Sims_DiD} presents the results for the standard trimmed DR DiD estimator. When the propensity score is misspecified but the outcome regression is correct (Panel~(a)), the trimmed estimator performs well: bias is negligible and coverage remains near 95\% across all positive thresholds, as one would expect from a doubly robust estimator. However, the reverse case (Panel~(c)) reveals the problem: when the outcome regression is misspecified and the propensity score is correct, trimming introduces growing bias. At $h = 0.05$, the 95\% empirical coverage drops to 71\%, and at $h = 0.10$, it collapses to 34\%. The double robustness property is lost, as the reliability of this procedure depends on the outcome model being correctly specified. We expand on this rationale in Section~\ref{sec:dr_bc}. The right column of Figure~\ref{fig:Sims_DiD} (Panels~(b) and~(d)) shows that our proposed weak-overlap-robust DR estimator resolves this trade-off. It trims comparison-group observations with extreme propensity scores, just as the standard trimmed estimator does, but adds a bias-correction term that compensates for the trimmed observations. In both configurations, whether the propensity score or the outcome regression is the correctly specified component, the estimator remains well-centered at the truth, and coverage stays near the nominal 95\% level (94--95\%). Confidence intervals at $h = 0.05$ are narrower than at $h = 0$, so the bias-corrected estimator retains double robustness while achieving the variance reduction that trimming provides. Thus, although the standard trimming procedure compromises the DR property, our proposed bias-correction procedure restores it and yields estimators with attractive statistical guarantees. \subsection{Empirical applications}\label{sec:applications} We compare DR-BC ($h = 0.05$) with the standard untrimmed DR estimator ($h = 0$) in four applications spanning all three research designs: unconfoundedness, IV and DiD. Because both estimators target the same causal parameter, differences in standard errors reflect genuine precision gains, not changes in what is being estimated or the causal question being addressed.\footnote{In all applications, propensity scores are estimated via logit regression and outcome regressions via least squares. Standard errors are computed analytically from the influence function. For the DiD event studies, standard errors are clustered at the county level (Medicaid) and municipality level (Dias--Fontes), and simultaneous confidence bands are based on multiplier bootstrap.} Because we focus on applications where weak overlap is a concern, the precision gains below are large; in settings with adequate overlap, the gains would be more modest. \subsubsection{Difference-in-Differences}\label{sec:app_did} We apply our method to two staggered DiD settings using the DR framework of \citet{Callaway2021}. In both cases, a rich set of covariates is used to make conditional parallel trends more plausible, but conditioning on them leads to weak overlap. \paragraph{Medicaid expansion and mortality.} Does Medicaid expansion save lives? \citet{baker2025difference} provides a DiD practical guide and uses this question as their running example, using county-level all-cause mortality rates for adults ages 20--64 and the staggered adoption of Medicaid expansion across states from 2009 to 2019. We re-analyze their data using the DR DiD framework of \citet{Callaway2021} with never-treated counties as the comparison group, conditioning on six county-level covariates (poverty rate, unemployment rate, median household income, percent female, percent White, and percent Hispanic) and using population weights. Because these covariates differ between expanding and non-expanding states, the generalized propensity score is close to 1 for some comparison counties, resulting in weak overlap; at $h = 0.05$, up to 3.3\% of comparison-group observations are trimmed in the most affected group-time cell. Panel~(a) of Figure~\ref{fig:DiD_applications} shows the event-study estimates. Post-treatment standard errors are on average 3.5 times larger for the standard DR estimator. While neither estimator rejects a zero effect, the standard DR 95\% simultaneous confidence bands span on average 68 deaths per 100,000, which is too wide to be policy relevant. DR-BC narrows these bands to roughly 17 deaths per 100,000 on average, ruling out large effects in either direction and providing a much more informative bound on the plausible magnitude of Medicaid expansion's mortality impact. \begin{figure}[h!t] \centering \resizebox{0.7\textwidth}{!}{ \begin{minipage}{\textwidth} \begin{subfigure}[t]{\textwidth} \centering \includegraphics[width = \textwidth]{Figures/Baker_et_al/ES_DiD_Medicaid_nevertreated.pdf} \vspace{-0.3cm}\caption{Medicaid expansion on mortality} \label{fig:DiD_Medicaid} \end{subfigure} \vspace{0.1cm} \begin{subfigure}[t]{\textwidth} \centering \includegraphics[width = \textwidth]{Figures/Dias_Fontes/ES_DiasFontes_pf_total_notyettreated_ant1.pdf} \vspace{-0.3cm}\caption{CAPS adoption on mental health practitioners} \label{fig:DiD_Dias_Fontes} \end{subfigure} \end{minipage} } \caption{DR DiD event-study estimators across applications} \label{fig:DiD_applications} \justifying \vspace{0.1cm}\noindent\scriptsize{\textit{Notes.} Panel~(a) shows population-weighted event-study estimates with staggered treatment timing using the doubly robust estimation method of \citet{Callaway2021}. The comparison group consists of counties that remained untreated by 2019. The outcome is the crude mortality rate for adults ages 20--64. Data are from \citet{baker2025difference}. Panel~(b) shows event-study estimates with not-yet-treated municipalities as the comparison group. The outcome is mental health practitioners per 10,000 population. Covariates include state fixed effects and 29 baseline characteristics. Three small northern states with fewer than 10 treated municipalities are excluded. We allow for a one-year anticipation. Data are from \citet{DiasFontes2024_AEJPolicy}. In both panels, DR-BC is our proposed estimator with trimming threshold $h = 0.05$; Standard DR is the untrimmed estimator of \citet{Callaway2021}. Points denote point estimates, vertical bars denote 95\% simultaneous confidence bands, and error bars with caps denote 95\% pointwise confidence intervals. } \end{figure} \paragraph{Mental health reform in Brazil.} Did Brazil's 2002 psychiatric reform expand community mental health capacity? \citet{DiasFontes2024_AEJPolicy} study this question using the staggered rollout of community-based Psychosocial Care Centers (CAPS) across municipalities, which replaced centralized hospital-based care. Because the timing of CAPS adoption was correlated with municipal characteristics, covariate adjustment is essential for the parallel trends argument. We re-analyze their data using the DR DiD framework of \citet{Callaway2021}, focusing on mental health practitioners per 10,000 population; results for additional outcomes (service utilization, hospital spending, violence) are in Appendix \ref{appendix:did_outcomes}. We restrict the sample to states with at least 10 treated municipalities and adjust for state fixed effects and 29 baseline covariates that are motivated by the original study. Because this rich covariate set leaves limited overlap between treated and not-yet-treated municipalities, a small number of comparison observations receive estimated propensity scores near one and therefore extremely large weights; at $h = 0.05$, up to 0.2\% of comparison-group observations are trimmed in the most affected group-time cell. Panel~(b) of Figure~\ref{fig:DiD_applications} presents the event-study estimates. Post-treatment standard errors are on average 2.8 times larger for the standard DR estimator, and at longer horizons the standard DR confidence intervals become too wide to be informative. DR-BC yields estimates precise enough to trace the workforce expansion over time, with a sustained increase of roughly 1 mental health practitioner per 10,000 population following CAPS adoption, consistent with the findings of \citet{DiasFontes2024_AEJPolicy}. See Appendix \ref{appendix:did_outcomes} and Appendix \ref{appendix:h_sensitivity} for additional outcomes and sensitivity analysis. \subsubsection{Unconfoundedness}\label{sec:app_ate} How large is the Black--White test score gap at age two, and can we estimate it without imposing functional form restrictions? Using data from the Early Childhood Longitudinal Study, Birth Cohort (ECLS-B), \citet{FryerLevitt2013_AER} document that the gap is negligible at 9 months but grows to roughly 0.2--0.4 standard deviations by age two, conditioning on a rich set of family background variables via OLS. Because their OLS specification imposes linearity and does not interact race dummies with covariates, it may not recover the ATE under treatment effect heterogeneity \citep{Sloczynski2022_Restat}. We re-analyze their data using DR estimation, treating race as the ``treatment'' in an operational sense: following \citet{holland1986}, the ``ATE'' here is interpreted as a covariate-adjusted descriptive contrast between Black and White children. Their conditioning set includes interviewer fixed effects, fifth-degree polynomials in parental age, and numerous interactions. The interviewer fixed effects create many small cells --- interviewers who observed children of mostly one race generate near-zero or near-one propensity scores --- pushing the propensity score to its boundaries for a non-negligible share of the sample. At $h = 0.05$, a small number of observations with extreme arm-specific weights drive the instability in the untrimmed DR estimator (see Appendix~\ref{appendix:ps_density} for the exact decomposition and Appendix~\ref{appendix:h_sensitivity} for sensitivity analysis). \begin{figure}[htbp] \centering \resizebox{0.7\textwidth}{!}{ \begin{minipage}{\textwidth} \centering \includegraphics[width = \textwidth]{Figures/Fryer_Levitt/forest_fryer_levitt_ecls.pdf} \end{minipage} } \caption{DR estimates of the Black--White test score gap} \label{fig:Fryer_Levitt} \justifying \vspace{0.1cm}\noindent\scriptsize{\textit{Notes.} Forest plot of ATE estimates for the Black--White test score gap at ages 9 months and 24 months from the ECLS-B. DR-BC is our proposed estimator with $h = 0.05$; Standard DR is the untrimmed estimator. Covariates and sample restrictions follow \citet{FryerLevitt2013_AER}. Error bars denote 95\% confidence intervals based on analytical standard errors. } \end{figure} Figure~\ref{fig:Fryer_Levitt} presents the results. At 24 months, the standard error drops from 0.072 to 0.025 (a 2.9-fold reduction), while at 9 months the reduction is 1.8-fold (from 0.027 to 0.015). DR-BC yields $-0.175$ (SE $= 0.025$) under a much more flexible specification that does not impose linearity, while the standard DR estimate of $-0.306$ (SE $= 0.072$) illustrates the precision cost of flexibility without overlap correction. DR-BC thus delivers OLS-level precision while directly targeting the ATE, without strong functional-form restrictions; the $h$-sensitivity analysis in Appendix \ref{appendix:h_sensitivity} (Figure~\ref{fig:h_sensitivity_ATE}) confirms that point estimates remain stable across $h$. \subsubsection{Instrumental variables, LATE, and ITT}\label{sec:app_late} Does 401(k) participation increase household savings? A central challenge is that participation is endogenous. A widely used identification strategy exploits employer-determined \emph{eligibility} as an instrument for participation \citep{abadie2003semiparametric, Chernozhukov2004, Benjamin2003a}. Using the 1991 Survey of Income and Program Participation (SIPP) data and the same covariates as \citet{abadie2003semiparametric}, we estimate both the ITT effect of eligibility and the LATE of participation for two outcomes: net financial assets and total wealth. We consider two samples: a full sample requiring only positive income \citep{Chernozhukov2004, Benjamin2003a}, and a restricted sample with income between \$10,000 and \$200,000 \citep{abadie2003semiparametric}. The income restriction was motivated precisely by weak overlap: as \citet[p.~249]{abadie2003semiparametric} notes, ``outside this interval, 401(k) eligibility in the sample is rare.'' Indeed, there are a few variance-inflating observations in this dataset: in the full sample, the 2 ineligible observations in the $\hat{p}_Z(X)>0.95$ tail have effective weights $1/(1-\hat{p}_Z(X))$ of roughly 31 and 98. It is this small set of high-weight ineligible observations that drives the variance inflation, yet removing their direct contribution via trimming is enough to reverse the sign of the untrimmed estimate. Figure~\ref{fig:401K} presents the results. The full-sample LATE for net financial assets illustrates how weak overlap can render standard DR estimates unreliable: the standard DR point estimate is $-\$13{,}042$ (SE $= \$16{,}223$), suggesting that 401(k) participation \emph{reduces} savings---a conclusion that is at odds with standard economic theory and the existing literature. The untrimmed DR estimator is uninformative in this sample, with a confidence interval spanning from $-\$45{,}000$ to $+\$19{,}000$. DR-BC yields $\$8{,}864$ (SE $= \$2{,}471$), recovering a positive, statistically significant effect consistent with prior findings \citep{abadie2003semiparametric}. The sign reversal is not a change in estimand---both estimators target the same LATE---but reflects removing the outsized influence of the two ineligible observations whose effective weights dominate the untrimmed estimator. Decomposing the Wald ratio reveals that the instability is entirely in the reduced form: the first-stage estimate is stable at $0.68$ regardless of trimming, but the intention-to-treat (ITT) swings from $-\$8{,}879$ (SE $= \$11{,}044$) to $\$6{,}035$ (SE $= \$1{,}686$). The precision gains here come primarily from variance reduction through trimming; the bias correction ensures theoretical integrity by targeting the original LATE and providing the correct influence function. DR-BC point estimates are stable across $h \in [0.03, 0.10]$ (Appendix~\ref{appendix:h_sensitivity}). For total wealth, the standard error reduction is 3.2-fold. In the restricted sample, gains are more modest (1.35- to 1.70-fold), but point estimates are stable across both samples. \begin{figure}[htbp] \centering \resizebox{0.8\textwidth}{!}{ \begin{minipage}{\textwidth} \begin{subfigure}[t]{.48\textwidth} \centering \includegraphics[width = \textwidth]{Figures/401K/forest_itt_nfa.pdf} \caption{DR ITT estimates of having access to 401(k) on net financial assets} \label{fig:401K_itt_nfa} \end{subfigure} \hfill \begin{subfigure}[t]{.48\textwidth} \centering \includegraphics[width = \textwidth]{Figures/401K/forest_itt_tw.pdf} \caption{DR ITT estimates of having access to 401(k) on total wealth} \label{fig:401K_itt_tw} \end{subfigure}\\[0.3cm] \begin{subfigure}[t]{.48\textwidth} \centering \includegraphics[width = \textwidth]{Figures/401K/forest_late_nfa.pdf} \caption{DR LATE estimates of 401(k) participation on net financial assets} \label{fig:401K_late_nfa} \end{subfigure} \hfill \begin{subfigure}[t]{.48\textwidth} \centering \includegraphics[width = \textwidth]{Figures/401K/forest_late_tw.pdf} \caption{DR LATE estimates of 401(k) participation on total wealth} \label{fig:401K_late_tw} \end{subfigure} \end{minipage} } \caption{Effect of 401(k) retirement plans on asset accumulation} \label{fig:401K} \justifying \vspace{0.1cm}\noindent\scriptsize{\textit{Notes.} Full sample restricts to households with positive income \citep{Benjamin2003a, Chernozhukov2004}. Restricted sample further restricts to income between \$10,000 and \$200,000 \citep{abadie2003semiparametric}. DR-BC trims observations with estimated instrument propensity scores $\widehat{p}_Z(X)$ outside $[0.05, 0.95]$; Standard DR is the untrimmed DR estimator based on \eqref{eq:dr_late}. Covariates: income, age, age squared, marital status, family size. Points denote estimates; horizontal lines denote 95\% confidence intervals. } \end{figure} \subsection{Trimming with bias correction to retain the DR property}\label{sec:dr_bc} Each of the DR estimands in Section~\ref{sec:setup}, whether for the ATE, LATE, or DiD ATT, can be written as a known function of one or more ratio moments $\alpha_\ell \equiv \mathop{}\!\mathbb{E}\left[B_\ell / A_\ell\right]$, where $\ell\in\{1,\cdots,L\}$ indexes the ratio moments entering the estimand. Here, $A_\ell$ involves the propensity score and can be close to zero, and $B_\ell$ contains a group indicator---multiplied by an outcome residual for the effect-estimating moments, or standing alone for the companion normalization moments. The specific $(A_\ell, B_\ell)$ decomposition depends on the design, but the weak overlap problem is always the same: when $A_\ell$ is near zero, $B_\ell / A_\ell$ is unstable with high variance. Standard trimming discards observations with $|A_\ell| < h$ for some threshold $h > 0$. This stabilizes the estimator but changes what is being estimated: \[ \mathop{}\!\mathbb{E}\left[\frac{B_\ell}{A_\ell} \cdot \mathbbm{1}\{ |A_\ell| \geq h \}\right] \neq \mathop{}\!\mathbb{E}\left[\frac{B_\ell}{A_\ell}\right]. \] The trimming bias equals $\mathop{}\!\mathbb{E}\left[(B_\ell / A_\ell) \cdot \mathbbm{1}\{ |A_\ell| < h \}\right]$, which does not vanish for fixed $h$. Our bias-correction term uses the behavior of the conditional expectation $\xi_\ell(a) \equiv \mathop{}\!\mathbb{E}\left[B_\ell \mid A_\ell = a\right]$ near the propensity score boundary to compensate for the discarded observations. The key observation is that, under either condition for double robustness, the relevant conditional mean at the propensity score boundary equals zero. There are two routes. \emph{First}, when the propensity score is correctly specified, $A_\ell = 0$ places the propensity score at its boundary, so one group has probability zero at that covariate value. The group indicator in $B_\ell$ then forces $B_\ell = 0$ almost surely conditional on $A_\ell = 0$, giving $\xi_\ell(0) = 0$. \emph{Second}, when the outcome regression is correctly specified, the residual-containing moments have a conditional mean of zero given $X$. Since $A_\ell$ is a function of $X$, the law of iterated expectations gives $\xi_\ell(a) = 0$ for those moments, and in particular $\xi_\ell(0) = 0$. The companion normalization moments need not satisfy this stronger statement, but they are irrelevant along this route because the corresponding residual moments are already zero. One can verify this for each estimand in Section (ref); the formal general statement is Assumption (ref) in Section (ref). Note the asymmetry: when the outcome regression is correct, $\xi_\ell(a) = 0$ for \emph{all} $a$ for the residual-containing moments, so trimming introduces no bias and the correction is inoperative. It is needed only when the propensity score is correct and the outcome model is not. In either case, $\xi_\ell(0) = 0$ implies that $\xi_\ell(a)/a$ has a removable singularity at zero and can be approximated by a polynomial in $a$. This explains the results in Figure (ref). Our proposed estimator replaces each problematic ratio moment $\mathop{}\!\mathbb{E}\left[B_\ell / A_\ell\right]$ with its bias-corrected counterpart: \begin{align*} \widehat{\alpha}_\ell(h) = \underbrace{\mathbb{E}_{n}\left[\frac{B_\ell}{A_\ell} \cdot \mathbbm{1}\{ |A_\ell| \geq h \}\right]}_{\text{trimmed mean}} + \underbrace{\sum_{\kappa=1}^{k} \frac{\mathbb{E}_{n}\left[A_\ell^{\kappa-1} \cdot \mathbbm{1}\{ |A_\ell| < h \}\right]}{\kappa!} \cdot \widehat{\xi}_\ell^{(\kappa)}(0)}_{\text{bias correction}}, \end{align*} where $\widehat{\xi}_\ell^{(\kappa)}(0)$ is the $\kappa$-th derivative of a sieve estimator for $\xi_\ell(\cdot)$, evaluated at zero, using a shifted Legendre polynomial basis of maximum degree $K$, and $k$ is the order of the bias-correction terms. The first part is the standard trimmed estimator. The second part uses the observations excluded by trimming ($|A_\ell| < h$), together with the estimated derivatives, to reconstruct their contribution. The final estimator applies a known function $\Lambda$ to the corrected moments $(\widehat{\alpha}_1(h), \ldots, \widehat{\alpha}_L(h))$; with the normalized weights in Section (ref), $\Lambda$ involves ratios for all three designs. This structure preserves double robustness. Under either DR condition, $\xi_\ell(0)=0$ holds for the residual-containing moments under correct outcome regression, and all moments under correct propensity score, so the polynomial approximation of $\xi_\ell(a)/a$ is valid for those moments. Under standard regularity conditions on the propensity score density and the smoothness of $\xi_\ell$, the resulting estimator is consistent and asymptotically normal for the original estimand, with a closed-form influence function that enables standard inference; the standard errors account for all sources of estimation uncertainty, including the sieve-based bias correction (Section (ref)). In practice, we recommend $h = 0.05$, $k = 1$ (first-order bias correction), and $K = 3$. Figure (ref) provides evidence: coverage remains near 95% under both misspecification configurations, while confidence intervals narrow relative to $h = 0$. Higher-order corrections ($k \geq 2$) degrade finite-sample performance. Sensitivity to $h$ is examined in Appendix (ref). Section (ref) presents the formal theory. \subsection{Practical diagnostics and when to use DR-BC} DR-BC and the standard (untrimmed) DR estimator target the same causal parameter, and DR-BC reduces to standard DR when no observations are trimmed ($h = 0$ or no propensity scores near the boundary). The cost of using DR-BC when overlap is adequate is therefore minimal: it adds a polynomial regression step that has a negligible impact on the estimates. The benefit when overlap is weak can be substantial, as the applications above demonstrate. When the number of units trimmed is fairly small, the bias correction has little to reconstruct, and the precision gains come primarily from variance reduction. We also recommend plotting estimates as a function of $h$ (Appendix (ref)) to assess stability; instability may indicate that the conditions underlying our results are not met. \section{Theory} This section formalizes the general estimator introduced in Section (ref), states the conditions for asymptotic normality, and verifies them for each research design. \subsection{General framework} The DR estimands in Section (ref) share a common structure. Each can be written as a known function of $L$ ratio moments: \begin{equation*} \theta_0=\Lambda\left(\mathop\!\mathbb{E}\left[\frac{B_1(\gamma_0)}{A_1(\gamma_0)}\right],\ldots,\mathop\!\mathbb{E}\left[\frac{B_L(\gamma_0)}{A_L(\gamma_0)}\right]\right), \end{equation*} where $(A_\ell(\gamma),B_\ell(\gamma))_{\ell=1}^L$ are known functions of observed data, $\gamma_0=(m,p)$ collects the true outcome regression $m$ and propensity score $p$, and $\Lambda$ is a known function.\footnote{$\Lambda$ may also depend on $\gamma$ through regression adjustments $\mathop{}\!\mathbb{E}\left[m(d,X)\right]$ that are smooth, root-$n$ estimable, and not subject to weak overlap; see Section (ref).} Weak overlap means some $A_\ell(\gamma_0)$ can be close to zero, making $B_\ell/A_\ell$ potentially unstable with high variance. Given i.i.d.\ observations $W_1,\ldots,W_n$, a preliminary estimator $\hat\gamma=(\hat m,\hat p)$ with probability limit $\gamma^*=(m^*,p^*)$ (so $m^*=m$ when the outcome model is correctly specified, or $p^*=p$ when the propensity score model is correctly specified), a trimming threshold $h > 0$, and positive integers $k \leq K$, our estimator is $\hat\theta=\Lambda(\hat\alpha_1(h,\hat\gamma),\ldots,\hat\alpha_L(h,\hat\gamma))$, where \begin{equation} \hat\alpha_\ell(h,\gamma)=\mathbb{E}_{n}\left[\frac{B_\ell(\gamma)}{A_\ell(\gamma)}\mathbbm{1}\{ |A_\ell(\gamma)|\geq h \}\right]+\sum_{\kappa=1}^{k}\frac{\mathbb{E}_{n}\left[A_\ell(\gamma)^{\kappa-1}\mathbbm{1}\{ |A_\ell(\gamma)|< h \}\right]}{\kappa !} \cdot \hat{\xi}_\ell^{(\kappa)}(0;\gamma). \end{equation} Here $\hat{\xi}_\ell^{(\kappa)}(0;\gamma)$ is the $\kappa$-th derivative at zero of a sieve estimator for ${\xi}_\ell(a;\gamma)\equiv\mathop{}\!\mathbb{E}\left[B_\ell(\gamma)\mid A_\ell(\gamma)=a\right]$, the conditional mean of $B_\ell$ given $A_\ell$. The dependence on $\gamma$ was suppressed in Section (ref). The sieve uses a shifted Legendre polynomial basis $q_K(\cdot)$ of degree $K$, natural since propensity scores lie in $[0,1]$; with $K$ fixed, the sieve basis is well-conditioned. The asymptotic theory allows $K$ to grow with $n$, and we recommend $K=3$ as a practical default for typical sample sizes in applications. Details are in Appendix (ref). \subsection{Asymptotic theory} We decompose $\hat\theta-\theta_0=(\hat\theta-\theta_h)+(\theta_h-\theta_0)$, where $\theta_h\equiv\theta_h(\gamma_0)$ with $\theta_h(\gamma)=\Lambda(\alpha_1(h,\gamma),\ldots,\alpha_L(h,\gamma))$ is the population analog of $\hat\theta$. The first term is the estimation error; the second is the trimming bias. We require the following assumptions. \begin{assumption}[Approximate double robustness] If either $m^*=m$ or $p^*=p$ (i.e., the outcome regression or propensity score is correctly specified), then $\theta_h(\gamma^*)=\theta_0+o(h^k)$. \end{assumption} \begin{assumption}[Conditional mean at zero] For each $\ell=1,\ldots,L$ with $0\in\mathrm{support}(A_\ell(\gamma^*))$: \textnormal{(i)} ${\xi}_\ell(0;\gamma^*)=0$ if $p^*=p$, and ${\xi}_\ell(a;\gamma^*)=0$ for all $a$ for the residual moments if $m^*=m$; \textnormal{(ii)} ${\xi}_\ell(\cdot;\gamma^*)$ is $(k+1)$-times continuously differentiable near $0$. \end{assumption} \begin{assumption}[Smoothness of $\Lambda$] $\Lambda(\cdot)$ is twice continuously differentiable near $(\alpha_1(0,\gamma^*),\ldots,$ $\alpha_L(0,\gamma^*))$. \end{assumption} Assumptions (ref)--(ref) are the key substantive conditions. Assumption (ref) extends double robustness from $\theta_0$ to its trimmed counterpart $\theta_h$; we verify it for each design in Section (ref). Assumption (ref)(i) is the key condition from Section (ref). When the propensity score is specified correctly, this gives $\xi_\ell(0;\gamma^*)=0$. When the outcome regression model is specified correctly, the residual-containing moments satisfy $\xi_\ell(a;\gamma^*)=0$ for all $a$, so in particular $\xi_\ell(0;\gamma^*)=0$; the companion normalization moments need not satisfy this statement. Part (ii) controls the polynomial approximation error. Assumption (ref) holds whenever the population moments $\alpha_\ell(0,\gamma^*)$ that appear as denominators in $\Lambda$ are bounded away from zero; this is a condition on the target estimand (e.g., compliance share bounded away from zero for the LATE, group share bounded away from zero for DiD), not on individual observations, and therefore remains compatible with the weak-overlap regime. \begin{assumption}[Regularity conditions] The following conditions hold for each $\ell=1,\ldots,L$: \begin{enumerate} • \textit{First-stage influence function.} $\alpha_\ell(h,\hat\gamma)-\alpha_\ell(h,\gamma^*)=(\mathbb{E}_{n}\left[\cdot\right]-\mathop{}\!\mathbb{E}[\cdot])[\phi_\ell]+o_p(n^{-1/2})$, where $\phi_\ell$ is the influence function of $\alpha_\ell(h,\gamma)$ at $\gamma^*$. • \textit{Sieve estimation.} For each $\kappa=1,\ldots,k$, \[ \hat{\xi}_\ell^{(\kappa)}(0;\gamma^*)-{\xi}_\ell^{(\kappa)}(0;\gamma^*)-(\mathbb{E}_{n}\left[\cdot\right]-\mathop{}\!\mathbb{E}[\cdot])[\psi_{\ell,\kappa}(\gamma^*)]=o_p(n^{-1/2}h^{1-\kappa}), \] where $\psi_{\ell,\kappa}(\gamma)=q_{K}^{(\kappa)}(0)'\mathop{}\!\mathbb{E}\left[q_{K}(A_\ell(\gamma))q_{K}(A_\ell(\gamma))'\right]^{-1}q_{K}(A_\ell(\gamma))(B_\ell(\gamma)-{\xi}_\ell(A_\ell(\gamma);\gamma))$. • \textit{Moment bound.} $\mathop{}\!\mathbb{E}\left[\omega_\ell(h,\gamma^*)^2\right]=o(n^{1/2})$, where $\omega_\ell$ is defined in (ref) below. • \textit{Stochastic equicontinuity.} $\hat\alpha_\ell(h,\hat\gamma)-\alpha_\ell(h,\hat\gamma)-\hat\alpha_\ell(h,\gamma^*)+\alpha_\ell(h,\gamma^*)=o_p(n^{-1/2})$. • \textit{Rate condition.} $n h^{2k}=O(1)$ as $n \rightarrow \infty$. \end{enumerate} \end{assumption} Parts (a)--(d) require the first-stage estimator, sieve regression, influence function moments, and the sample criterion to behave well; lower-level sufficient conditions (e.g., for parametric first stages) are Appendix (ref). Part (e) imposes an upper bound on $h$ that controls the trimming bias; a complementary lower bound of the form $n h^{4} \to \infty$, needed to control the variance of the sieve bias-correction term, is stated in Appendix (ref). The influence function $\omega_\ell$ of each moment $\hat\alpha_\ell$ has four components: \begin{align} \omega_\ell(h,\gamma) &=\underbrace{\frac{B_\ell(\gamma)}{A_\ell(\gamma)}\mathbbm{1}\{ |A_\ell(\gamma)|\geq h \}}_{\text{trimmed ratio}} \;+\;\underbrace{\sum_{\kappa=1}^{k}\frac{A_\ell(\gamma)^{\kappa-1}\mathbbm{1}\{ |A_\ell(\gamma)|<h \}}{\kappa !} \cdot {\xi}_\ell^{(\kappa)}(0;\gamma)}_{\text{bias correction}} \nonumber\\[6pt] &\quad+\;\underbrace{\sum_{\kappa=1}^{k}\frac{\mathop\!\mathbb{E}\left[A_\ell(\gamma)^{\kappa-1}\mathbbm{1}\{ |A_\ell(\gamma)|<h \}\right]}{\kappa !} \cdot \psi_{\ell,\kappa}(\gamma)}_{\text{sieve estimation}} \;+\;\underbrace{\phi_\ell}_{\text{first stage}}. \end{align} Let $\Lambda_\ell$ denote the partial derivative of $\Lambda$ with respect to its $\ell$-th argument, and define the overall influence function \begin{equation} \varphi=\sum_{\ell=1}^L\Lambda_\ell(\alpha_1(h,\gamma^*),\ldots,\alpha_L(h,\gamma^*))\,\omega_\ell(h,\gamma^*). \end{equation} Note that $\Lambda_\ell$ is evaluated at the population moments $\alpha_\ell(h,\gamma^*)$---the probability limits of the bias-corrected moment estimators---rather than at $\alpha_\ell(0,\gamma^*)$; these are the correct expansion points for the delta method because, under the outcome-correct route, $\alpha_\ell(h,\gamma^*)$ need not equal $\alpha_\ell(0,\gamma^*)$ for the normalization moments. Under Assumption (ref), when $h$ goes to zero at the rate in Assumption (ref)(e), $\alpha_\ell(h,\gamma^*)$ converges to $\alpha_\ell(0,\gamma^*)$ and the influence function targets $\theta_0$. \begin{theorem} Under Assumptions (ref)--(ref): \begin{enumerate} • $\hat\theta-\theta_0= (\mathbb{E}_{n}\left[\cdot\right]-\mathop{}\!\mathbb{E}[\cdot])[\varphi]+o_p(n^{-1/2})$. • If in addition $\mathop{}\!\mathbb{E}\left[\varphi^2\right]$ is bounded away from zero and ${\mathop{}\!\mathbb{E}\left[|\varphi-\mathop{}\!\mathbb{E}\left[\varphi\right]|^{2+\eta}\right]ta}}}/({n^{\eta/2}\mathop{}\!\mathbb{E}\left[(\varphi-\mathop{}\!\mathbb{E}\left[\varphi\right])^2\right])^2}^{(2+\eta)/2}})=o(1)$ for some $\eta>0$, then $\displaystyle(\hat\theta-\theta_0)\Big/\sqrt{\mathop{}\!\mathbb{E}\left[(\varphi-\mathop{}\!\mathbb{E}\left[\varphi\right])^2\right])^2}/n}\stackrel{d}{\rightarrow}\mathcal{N}(0,1).$ \end{enumerate} \end{theorem} A proof is in Appendix (ref). A key difficulty relative to sasaki2018estimation is that $\theta_0$ is a nonlinear function of multiple ratio moments with heterogeneous convergence rates, so the standard delta method does not directly apply. The variance is estimated by $\hat\sigma^2 = \mathbb{E}_{n}\left[(\hat\varphi - \mathbb{E}_{n}\left[\hat\varphi\right])^2\right]hat\varphi})^2}$, where $\hat\varphi$ plugs sample analogs into (ref)--(ref). \begin{remark}[Diverging variance] Assumption (ref)(c) allows $\mathop{}\!\mathbb{E}\left[\omega_\ell^2\right]$ to diverge, accommodating heavy tails from $B_\ell/A_\ell$ when $h$ is small. The Lyapunov condition ensures the CLT holds despite this. \end{remark} \begin{remark}[Verification across designs] The design-specific conditions underlying Assumptions (ref)--(ref) are satisfied for the ATE, LATE, and staggered DiD estimands introduced in Section (ref); see Appendix Table (ref) for the corresponding $(A_\ell,B_\ell)$ decompositions and Appendix (ref) for the formal verification. \end{remark} {\singlespacing {1pt plus 0.3ex} \putbib }
bibunit