Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
178,471 characters · 26 sections · 88 citation commands
Robust Inference for Weighted Estimands
In empirical research, many estimands can be expressed as weighted averages of group-level parameters.\footnote{Under various assumptions, ordinary least squares (OLS) estimands can be expressed as weighted averages of conditional average treatment effects (CATEs), conditional mean derivatives, or other regression parameters that may vary across covariate cells and treatment intensities yitzhaki1996using, angrist1998estimating, aronow2016does, sloczynski2022interpreting, goldsmith2024contamination; two-stage least squares (TSLS) estimands can be expressed as weighted averages of local average treatment effects (LATEs), which may vary across compliance types and covariate cells imbens1994identification, heckman2005structural, angrist2010extrapolate, kolesar2013ivheterogeneity, sloczynski2020should, huntington2020instruments, coussens2021improving, blandhol2022tsls, abadie2024instrumental; and two-way fixed effects (TWFE) estimands can be expressed as weighted averages of average treatment effects on the treated (ATTs), which may vary across treatment cohorts and time periods de2020two, callaway2021difference, goodman2021difference, sun2021estimating, athey2022design, gardner2022two, borusyak2024revisiting, wing2024stacked.} These weighted estimands give concise summaries of the parameters of interest, making them frequent targets for statistical inference. However, conventional inference procedures can be sensitive to the choice of weights: conclusions reached under a researcher’s baseline weights may not hold under alternative weights entertained by readers.
For example, in event studies with multiple treatment cohorts, many common estimands recover weighted averages of cohort-level causal effects, including conventional two-way fixed effects (TWFE) regression estimands and recently proposed alternatives.\footnote{For a review, see roth2023s.} TWFE has a convenient and flexible structure, which can yield lower-variance estimators in practice armstrong2025adapting. However, TWFE can place negative weight on some of the cohort-level effects, complicating the interpretation of the resulting estimands under heterogeneous effects de2020two, goodman2021difference, sun2021estimating, borusyak2024revisiting. Recent alternatives allow a researcher to target weighted estimands that rule out negative weights by design, but this introduces an additional degree of freedom: the researcher must decide which positive weights to use.
This choice matters because different readers may prefer different weights, depending on the empirical or policy question that they have in mind callaway2021difference. One reader may wish to upweight early-treated cohorts, another later-treated cohorts; one may care about short-run effects, another long-run effects; one may prefer weights that represent the characteristics of a policy-relevant target population, another may favor weights that trade off representativeness against estimation precision. These alternatives may all be reasonable, but they need not lead to the same conclusion under heterogeneous effects. Given the potential for disagreement, researchers may wish to assess the robustness of conclusions to the choice of weights. Current practice seems to reflect this desire by reporting results for different weighting choices, such as default implementations of the recently proposed alternatives. But such exercises consider only a finite set of weighting choices---which may not cover the range of plausible weighting schemes---and they generally do not provide formal inferential guarantees for claims that conclusions are robust to the choice of weights.
In this paper, I develop robust inference procedures that directly account for ambiguity and disagreement over the choice of weights. I consider a general framework in which a researcher observes asymptotically normal estimates for a vector of parameters. The researcher reports a conventional point estimator and confidence interval (CI) for a baseline estimand defined by a vector of baseline weights. At the same time, the researcher faces a broad class of alternative weights, giving rise to alternative estimands for which the baseline estimator may be biased and the baseline CI may undercover. To address these concerns, I develop robust estimators that mitigate worst-case bias and robust CIs that ensure valid coverage.
To proceed, I first establish a sharp bound on the difference between the baseline weighted estimand and any given alternative weighted estimand. This bound is the product of (i) the heterogeneity in parameters, defined as the square root of the residual sum of squares from a generalized least squares (GLS) regression of the parameters on a constant and (ii) the distance between weights, defined as the standard deviation of the difference in estimators under the baseline and alternative weights.
To develop my estimator, I study the problem of choosing weights to minimize the maximum distance between weights over a class of alternative weights. For a broad range of such classes, I show that this problem admits a unique solution. Moreover, I show that the corresponding estimator is optimal for minimizing the maximum bias over the class of alternative estimands under any bound on the heterogeneity in parameters. I therefore refer to this estimator as the robust estimator, and to the corresponding weights as the minimax-bias weights. I propose that researchers report the robust estimator alongside the baseline estimator. Intuitively, the robust estimator accounts for the bias concerns of readers who may disagree with the baseline weights. Moreover, since the minimax-bias weights depend only on the class of alternative weights, they provide a natural default for researchers facing ambiguity over their initial choice of baseline weights.
To develop my CI, I study the problem of constructing an upper confidence bound (UCB) for the heterogeneity in parameters. I show that a monotonic transformation of the heterogeneity in estimates produces an optimal UCB. This heterogeneity UCB—together with the maximum distance between weights—yields a simple adjustment to the critical values of a baseline CI. The resulting robust CI provides uniformly valid coverage over the class of alternative estimands. In particular, for each alternative estimand, the coverage probability is at least one minus the sum of (i) the significance level of the baseline CI and (ii) the significance level of the heterogeneity UCB. For example, suppose a 95% baseline CI excludes zero, leading the researcher to reject the null of no average effect at the 5% level. If zero remains excluded from a robust CI constructed using a 95% heterogeneity UCB, then readers with alternative weights can robustly reject the null of no average effect at the 10% level. To facilitate such inference procedures, I propose that researchers report the robust CI alongside the baseline CI.
The robust estimator and CI require the researcher to specify a class of alternative weights. My framework accommodates any compact and convex class of alternatives. I highlight three such classes that cover a broad range of empirical contexts. Below I give a high-level overview of these classes, deferring precise definitions and details to Section (ref) of the paper.
The first class is the bounded variance class, which restricts attention to weights that yield estimators with variance no larger than a given bound, ensuring that the alternative estimands can be estimated with a reasonable degree of precision. For example, when assessing robustness of results, a researcher may wish to consider alternative estimands that can be estimated at least as precisely as the baseline. Under the bounded variance class, my proposed procedures admit a convenient reduction to the GLS regression of the group-level estimates on a constant, where (i) the GLS estimator coincides with the robust estimator, (ii) the GLS variance determines the maximum distance between weights for any given choice of variance bound, and (iii) the GLS residual sum of squares determines the heterogeneity UCB. This reduction has two key implications. First, because the GLS estimator minimizes variance, the robust estimator is optimal for both bias and variance under the bounded variance class. Second, due to the structure in (ii) and (iii), readers can construct robust CIs for their own choices of variance bounds, provided the researcher reports the GLS variance and heterogeneity UCB. These convenient properties make the bounded variance class a natural benchmark.
The simplex is the set of nonnegative weights, which yields weighted estimands that do not extrapolate beyond the range of the parameters. However, the unrestricted simplex allows for some groups to receive zero weight, which can be undesirable in practice. To ensure that groups are represented, one can impose a floor on the simplex. The truncated simplex class is precisely the set of simplex weights that are bounded below by the given floor. Under positive baseline weights, the floor can be parameterized so that the truncated simplex yields the set of weights that are a given fraction between the baseline weights and the unrestricted simplex weights. In this form, the truncated simplex class allows one to investigate perturbations of a chosen magnitude from the baseline weights.
The simplex weights can be viewed as probability distributions over the groups, and thus as different target populations that readers may be interested in. For instance, when the groups index the sites of an experiment, baseline results may not generalize to future sites of interest to policymakers allcott2015site. If the researcher observes site-level covariates, however, the differences in covariate means under the baseline and alternative weights can inform the extent of external validity. Based on this idea, the covariate balance class considers simplex weights whose covariate means differ from the baseline mean by at most a given constant. Under this class, readers can assess whether results are robust to alternative populations whose covariate means are not too far from the baseline population.
I illustrate my framework in two empirical applications. The first application is an event study in lakdawala2023dynamic, which uses the staggered rollout of internet access in Peruvian public schools to estimate the causal effects of school-based internet access on second grade test scores. The authors find a delayed achievement response: TWFE estimates are initially small but grow over time, indicating that schools require time to adapt to new internet access. The authors find similar patterns and magnitudes when using the sun2021estimating estimator to account for negative weighting concerns, concluding that results are not sensitive to the use of TWFE. My robust estimator and CI give a stronger conclusion: results are robust to classes of nonnegative weights that yield comparable estimation precision to the sun2021estimating weights. In particular, robustness only fails when attempting to cover weighted estimands that cannot be estimated precisely in the first place.
The second application is Tennessee's Project STAR (Student/Teacher Achievement Ratio) experiment, which randomized students in seventy-nine Tennessee public elementary schools to classrooms of different sizes to estimate the causal effects of class size on test scores achilles2008tennessee, krueger1999experimental. Project STAR has been used in many policy discussions, but the experiment's selection of schools raises concerns about whether results are representative of Tennessee or U.S. schools more broadly schanzenbach2006have. To investigate these issues, I use an equal weights baseline to model the distribution of STAR schools and the truncated simplex class to model departures from the STAR empirical distribution. I show that, while the baseline CI implies medium-sized positive effects, the robust CI includes small-to-zero effect sizes even under small departures from equal weighting. The robust CI shrinks when I intersect the truncated simplex with the covariate balance class based on school-level covariates, but results remain sensitive to small departures from the baseline weights.
While I focus on event studies and multisite experiments as the motivating examples, my framework applies broadly to settings where group-level parameters are averaged using weights that sum to one. My inference procedures are not designed for continuously indexed groups, but the bound on differences between weighted estimands has a natural analogue in that case.\footnote{For a discussion, see footnote (ref) under Proposition (ref).} Finally, the inferential guarantees of the robust estimator and CI do not depend on how the underlying parameters are defined. For example, in event studies one can define the parameters as cohort-level difference-in-differences (DiD) estimands rather than cohort-level average causal effects sun2021estimating. This distinction affects the interpretation of the weighted estimands, but not the validity of the inference procedures.
A large literature highlights issues that arise when interpreting weighted estimands. In the context of my framework, these issues represent different sources of ambiguity and disagreement over the choice of weights. For example, regression-based estimands often recover weighted averages of treatment effects, but allow for negative weights de2020two, sloczynski2020should, goodman2021difference, mogstad2021causal, sun2021estimating, blandhol2022tsls, bhuller20242sls, goldsmith2024contamination. Negative weights can flip the signs of treatment effects, complicating the interpretation of an estimand. In such contexts, there may be (i) ambiguity over how much negative weighting matters and (ii) disagreement over what alternative weighting schemes to consider.\footnote{abadie2025harvesting and chiu2026causal give different takeaways for issue (i) in the event study context.} Regarding the latter, a separate issue is that a positively weighted estimand can still lack empirical or policy relevance yitzhaki1996using, heckman2005structural, crump2006moving, angrist2010extrapolate, aronow2016does, li2018balancing, sloczynski2022interpreting, mogstad2024instrumental, poirier2024quantifying. In this case, disagreement may arise from policymakers interested in the effects of treatment on their own target populations, while ambiguity may arise from researchers who must navigate such disagreement when summarizing results.
Common approaches to these issues are to report the extent of negative weighting in one's estimator or to examine the stability of results under alternative weighting schemes.\footnote{See roth2023s for a review of such approaches in the event study context.} Intuitively, such approaches convey information about how far one's weights deviate from a benchmark, the degree of underlying heterogeneity, or some combination of both. My procedures build on this same intuition, but in a GLS-based geometry that yields a notion of heterogeneity amenable to inference with confidence bounds, drawing on statistical results on optimal quantile-unbiased estimation pfanzagl1994parametric. This allows one to directly account for weighting issues by constructing robust CIs and estimators with explicit inferential guarantees, thereby facilitating robust inference for weighted estimands.
The remainder of this paper proceeds as follows. Section (ref) develops the model setting and notation. Section (ref) defines the heterogeneity and distance measures and establishes the bound on differences between weighted estimands. Section (ref) develops the robust estimator and CI and establishes their properties. These results are developed under the assumption that group-level estimates are normally distributed with a known covariance matrix. Section (ref) shows that, with asymptotically normal estimates and a consistent covariance matrix estimator, my proposed procedures are uniformly asymptotically valid over a broad class of data generating processes. Section (ref) discusses the practical implementation of these procedures. Section (ref) presents the empirical applications. Section (ref) concludes. The supplemental appendices contain proofs and additional results.
Consider parameters $\theta_{k} \in \mathbb{R}$ indexed by groups $k \in \{1, \ldots, K\}$, where $K \geq 2$. Let $\theta = (\theta_{1}, \ldots, \theta_{K})'$ denote the vector of parameters. Researchers often conduct inference on estimands that can be expressed as weighted averages of the group-level parameters:
where $w \in \mathcal{W}$ is a vector of weights, $\mathcal{W}$ is the set of all weights, and $\bm{1}$ is the vector of ones. I refer to $\tau_{w}(\theta)$ as a weighted estimand. I allow the weights $w$ to be negative unless stated otherwise. For the case of nonnegative (i.e., convex) weights, I define the simplex $\mathcal{W}_{+} = \{w \in \mathcal{W}: w \geq 0\}$.
The researcher observes a vector of estimates $\hat{\theta}$ for the parameters $\theta$. I model the relationship between $\hat{\theta}$ and $\theta$ as follows.
I denote probabilities, expectations, and variances under $\hat{\theta} \sim N(\theta, \Sigma)$ as $\mathbb{P}_{\theta}\left\{\cdot\right\}$, $\mathbb{E}_{\theta}\left[\cdot\right]$, and $\textnormal{Var}_{\theta}\left(\cdot\right)$, respectively. Moreover, $\Phi(z)$ denotes the cumulative distribution function (CDF) of the standard normal distribution $N(0,1)$ and $z_{\alpha}$ denotes its $\alpha$-quantile for $\alpha \in (0,1)$.
Assumption (ref) is motivated by standard large-sample asymptotic results. For example, the central limit theorem yields asymptotically normal estimates, which motivates the normality condition. Likewise, the law of large numbers yields consistent covariance matrix estimation, which motivates the condition that $\Sigma$ is known. I therefore develop my procedures under Assumption (ref). In Section (ref), I relax Assumption (ref) and establish the asymptotic validity of my procedures.
For each vector of weights $w \in \mathcal{W}$, Assumption (ref) implies $w'\hat{\theta} \sim N(w'\theta, \sigma_{w}^{2})$, where $\sigma_{w}^{2} = w'\Sigma w$. For a given target estimand $\tau_{w}(\theta) = w'\theta$, the conventional point estimator is defined as $\hat{\tau}_{w} = w'\hat{\theta}$ and the conventional CI at significance level $\alpha$ is defined as
The conventional estimator $\hat{\tau}_{w}$ is unbiased for $\tau_{w}(\theta)$:
The conventional $CI_{w}$ has exact coverage for $\tau_{w}(\theta)$:
In this sense, the conventional statistics $(\hat{\tau}_{w}, CI_{w})$ provide valid inference for $\tau_{w}(\theta)$.
The researcher reports $(\hat{\tau}_{w}, \sigma_{w})$ for some baseline choice of weights $w \in \mathcal{W}$. This report yields conventional statistics $(\hat{\tau}_{w}, CI_{w})$ that provide valid inference for the baseline estimand $\tau_{w}(\theta)$. However, the researcher may also consider a class $\Lambda \subseteq \mathcal{W}$ of alternative weights $\lambda \in \Lambda$, yielding alternative estimands $\tau_{\lambda}(\theta)$ for which $(\hat{\tau}_{w}, CI_{w})$ may be uninformative.
Whether the issue at hand is researcher ambiguity, reader disagreement, or some combination of both, my framework assumes that it can be represented by a class of alternative weights $\Lambda$. To make this modeling assumption a practical one, I restrict attention to classes $\Lambda$ that have the following structure.
This structure ensures the existence and uniqueness of solutions to upcoming optimization problems. I now provide examples of classes $\Lambda$ that satisfy Assumption (ref) and map to various empirical contexts.
I conclude this section by formalizing the inference distortions that arise when one uses the baseline statistics $(\hat{\tau}_{w}, CI_{w})$ to learn about alternative estimands $\tau_{\lambda}(\theta)$. In particular, I consider (i) the absolute bias of $\hat{\tau}_{w}$ for $\tau_{\lambda}(\theta)$, given by
and (ii) the noncoverage probability of $CI_{w}$ for $\tau_{\lambda}(\theta)$, given (in the two-sided case) by
The inference distortions are governed by the (absolute) difference in estimands $|\tau_{\lambda}(\theta)- \tau_{w}(\theta)|$: the larger this difference, the higher the bias and noncoverage. In particular, when $\lambda \neq w$, there exist parameter values $\theta$ where $|\tau_{\lambda}(\theta)- \tau_{w}(\theta)|$ is arbitrarily large so that bias is arbitrarily large and noncoverage is arbitrarily close to one. In such cases, $(\hat{\tau}_{w}, CI_{w})$ is completely uninformative for $\tau_{\lambda}(\theta)$. Thus, to obtain useful inferences for $\tau_{\lambda}(\theta)$, one must account for potential differences $|\tau_{\lambda}(\theta)- \tau_{w}(\theta)|$ in the weighted estimands across $\lambda \in \Lambda$ and $\theta \in \mathbb{R}^{K}$.
In this section, I bound the difference in estimands in terms of the heterogeneity in parameters and the distance between weights. I then show how to infer these quantities, which will be the basis for constructing the robust estimator and CI in Section (ref).
\paragraph{Heterogeneity in Parameters.} Consider the generalized least squares (GLS) regression of the parameters $\theta$ on a constant $\bm{1}$ under weighting matrix $\Sigma^{-1}$. I define the heterogeneity in parameters as the square root of the GLS residual sum of squares:
By construction, the heterogeneity in $\theta$ is zero if and only if $\theta_{k}$ is constant across $k$. The above regression is uniquely minimized at the GLS estimand
which yields the formula
where $A$ is the annihilator matrix for $\Sigma^{-1/2}\bm{1}$ and $I$ is the identity matrix. Thus, the squared heterogeneity can be represented as a quadratic form $\theta'Q\theta$ of the parameters. This particular quadratic form will facilitate quantile-unbiased inference in Section (ref).
\paragraph{Distance Between Weights.} The standard deviation $\left\|v\right\|_{\Sigma} = \displaystyle\sqrt{v'\Sigma v}$ of a given linear estimator $v'\hat{\theta}$ defines a norm on $v \in \mathbb{R}^{K}$. Given weights $\lambda$ and $w$, I define the distance between weights as the corresponding norm of their difference:
Intuitively, the distance $\left\|\lambda - w\right\|_{\Sigma}$ measures the disagreement between $\lambda$ and $w$ by taking the standard deviation of the corresponding difference in estimators.
Proposition (ref) shows that the difference in estimands is bounded by the product of (i) the heterogeneity $H(\theta)$, which is unknown due to the parameters $\theta$ being unobserved, and (ii) the distance $\left\|\lambda - w\right\|_{\Sigma}$, which is unknown due to ambiguity or disagreement over the alternative weights $\lambda \in \Lambda$.\footnote{Proposition (ref) is based on the Cauchy-Schwarz inequality, similar to Scheffé-style arguments for bounding data-dependent linear combinations to obtain uniformly valid inference scheffe1953method, lehmann2024testing. However, Proposition (ref) considers nonrandom linear combinations and obtains bounds on the inference distortions themselves. Analogous bounds hold if one replaces $\Sigma$ with another positive definite matrix when defining the heterogeneity and distance measures. It also has a natural analogue for continuously indexed groups, where one may define the weighted estimands as $\tau_{w}(\theta) = \int w(k)\theta(k)d\nu(k)$ and apply the same Cauchy-Schwarz argument in an $L^{2}(\nu)$ inner product. } In Section (ref), I account for the latter by taking the maximum distance across alternative weights. In Section (ref), I account for the former by constructing an upper confidence bound (UCB) on the heterogeneity.
Given baseline $w$ and class of alternatives $\Lambda$, the maximum distance between weights is
Intuitively, the maximum distance measures the worst-case disagreement between the baseline and alternative weights across $\Lambda$. Under the structure on $\Lambda$ from Assumption (ref), the maximum is attained at some $\lambda^{*} \in \Lambda$, which can be viewed as the weights of a reader who disagrees with $w$ the most. Below I analyze the structure of the maximum distance under the example classes from Section (ref) and illustrate how this structure can facilitate communication between researchers and readers---I defer practical recommendations and implementation choices to Section (ref).
I now show how to infer the unknown heterogeneity in parameters $H(\theta)$. In particular, I derive an upper confidence bound (UCB) for $H(\theta)$: given a significance level $\beta \in (0,1)$, I construct an estimator $\hat{\eta}_{1-\beta}$ such that
In words, $\hat{\eta}_{1-\beta}$ upper bounds the unknown heterogeneity $H(\theta)$ with probability at least $1-\beta$, for any value of $\theta$. I refer to $\hat{\eta}_{1-\beta}$ as a heterogeneity UCB at confidence level $1-\beta$.
\paragraph{Derivation.} To derive $\hat{\eta}_{1-\beta}$, I use the representation of $H(\theta) = \displaystyle\sqrt{\theta'Q\theta}$ as a quadratic form in $\theta$. In particular, since $\hat{\theta} \sim N(\theta, \Sigma)$, the definition of $Q$ from (ref) implies that the heterogeneity in estimates $H(\hat{\theta}) = \sqrt{\hat{\theta}'Q\hat{\theta}}$ satisfies
where $\chi_{K-1}^{2}(\eta)$ denotes the noncentral chi-squared distribution with $K-1$ degrees of freedom and noncentrality parameter $\eta^{2}$. Let $F_{\chi^{2}}(x; \eta)$ denote its CDF and consider the corresponding pivot function $\eta \mapsto F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; \eta)$. The probability integral transform yields
In particular, the probability of observing $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; H(\theta)) \geq \beta$ is equal to $1-\beta$. For each $x > 0$, the function $\eta \mapsto F_{\chi^{2}}(x;\eta)$ is strictly decreasing sun2010monotonicity. Therefore, when $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; 0) > \beta$, I define $\hat{\eta}_{1-\beta}$ as the unique solution to $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; \hat{\eta}_{1-\beta}) = \beta$. Otherwise, when $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; 0) \leq \beta$, I define $\hat{\eta}_{1-\beta} = 0$.\footnote{This construction is consistent with the general approach in pfanzagl1994parametric, which constructs confidence bounds for one-parameter distributions satisfying appropriate monotonicity conditions.} In summary,
This construction depends only on $\hat{\theta}'Q\hat{\theta}$, which is conveniently obtained in (ref) as the residual sum of squares from the GLS regression of the estimates $\hat{\theta}$ on a constant $\bm{1}$ under weighting matrix $\Sigma^{-1}$. The inversion for $\hat{\eta}_{1-\beta}$ is computationally efficient, since it amounts to finding the root of a strictly monotone function.
The next result shows that the above $\hat{\eta}_{1-\beta}$ is a valid UCB for $H(\theta)$. Moreover, it is quantile-unbiased under heterogeneity in the sense that its (weak) overestimation probability for $H(\theta)$ is equal to $1-\beta$ when $H(\theta) > 0$. Note that when $H(\theta) = 0$, the overestimation probability is equal to one, as would be the case for any nonnegative estimator of $H(\theta)$.
The quantile-unbiasedness in Proposition (ref) facilitates a variety of inference procedures for $H(\theta)$. First, $[0, \hat{\eta}_{1-\beta}]$ is a one-sided upper confidence interval for $H(\theta)$, with exact coverage rate $1-\beta$ under heterogeneous $\theta$. Next, the one-sided lower interval $[\hat{\eta}_{\beta}, \infty)$ satisfies
Likewise, one can construct a two-sided interval $[\hat{\eta}_{\beta/2}, \hat{\eta}_{1-\beta/2}]$ that satisfies
Finally, one can construct a point estimator $\hat{\eta}_{1/2}$ that is median-unbiased in the sense that
That is, the median of $\hat{\eta}_{1/2}$ is equal to $H(\theta)$ under heterogeneity.\footnote{For point estimation of heterogeneity, one can also consider mean-unbiased estimation of $\theta'Q\theta$, following a similar construction to, e.g., kline2020leave. In particular, one can show that $\hat{\theta}'Q\hat{\theta} - (K-1)$ is a mean-unbiased estimator of $\theta'Q\theta$. However, the quantile-unbiased approach adopted here is more amenable to the construction of confidence bounds for my objects of interest.} Given my focus on upper bounding differences in weighted estimands, I focus on properties (ref) and (ref). Nevertheless, properties (ref)-(ref) illustrate the versatility of $\hat{\eta}_{1-\beta}$ for inference on parameter heterogeneity. I conclude this section with comparative statics and an optimality result for $\hat{\eta}_{1-\beta}$.
\paragraph{Comparative Statics.} For a given $\eta$, the CDF $F_{\chi^{2}}(x; \eta)$ is increasing in $x$ and decreasing in its degrees of freedom sun2010monotonicity. Thus, $\hat{\eta}_{1-\beta}$ is increasing in $\hat{\theta}'Q\hat{\theta}$ given a fixed $K$, and decreasing in $K$ given a fixed $\hat{\theta}'Q\hat{\theta}$. In this sense, $\hat{\eta}_{1-\beta}$ is increasing in heterogeneity of the estimates relative to the number of groups: $H(\hat{\theta})/\sqrt{K}$. Notice that, under homoskedasticity $\Sigma = \sigma^{2}I$, the expression $\hat{\theta}'Q\hat{\theta}/K$ reduces to the empirical variance of the estimates relative to the common sampling variance $\sigma^{2}$ of the estimates:
Thus, the proposed heterogeneity measure $\hat{\eta}_{1-\beta}$ has intuitive comparative statics in terms of $\hat{\theta}'Q\hat{\theta}/K$. Finally, holding $\hat{\theta}'Q\hat{\theta}$ and $K$ fixed, a larger confidence level $1-\beta$ implies a larger $\hat{\eta}_{1-\beta}$. This highlights a tradeoff when constructing $\hat{\eta}_{1-\beta}$: to upper bound $H(\theta)$ with higher confidence, one must tolerate a larger UCB, corresponding to a lower significance level $\beta$.
\paragraph{Optimality.} Let $\Bar{\Theta}^{\eta} = \{\theta: H(\theta) = \eta\}$ denote the set of parameters for which the heterogeneity is equal to $\eta$, and consider the class of potentially randomized estimators $\Tilde{\eta}_{1-\beta}$ that are quantile-unbiased in the sense of (ref). The following result shows that $\hat{\eta}_{1-\beta}$ is minimax most accurate in the sense that it minimizes the worst-case probability of overestimating $H(\theta)$, for any degree of underlying heterogeneity $\eta$.
The above notion of optimality is a minimax analogue of the notion of a uniformly most accurate confidence bound lehmann2024testing. In Appendix (ref), I show that $\hat{\eta}_{1-\beta}$ is uniformly most accurate in the class of quantile-unbiased estimators that depend on $\hat{\theta}$ through $\hat{\theta}'Q\hat{\theta}$. In fact, $\hat{\eta}_{1-\beta}$ minimizes expected loss in this class under any quasiconvex loss function that attains its minimum at $H(\theta)$, for any heterogeneous $\theta$. The latter optimality statement is based on results from pfanzagl1994parametric on optimal quantile-unbiased estimation in models with monotone likelihood ratios.
Given the maximum distance between weights $\max_{\lambda \in \Lambda}\left\|\lambda - w\right\|_{\Sigma}$ and the heterogeneity UCB $\hat{\eta}_{1-\beta}$, I define the bias UCB
The following result shows that $\widehat{B}_{w}^{\beta}(\Lambda)$ is indeed a valid bias UCB.
In words, $\widehat{B}_{w}^{\beta}(\Lambda)$ upper bounds the maximum bias with probability at least $1-\beta$, for any value of $\theta$. This provides an avenue for addressing the bias and undercoverage of the baseline estimator $\hat{\tau}_{w}$ and $CI_{w}$ for inference on alternative estimands $\tau_{\lambda}(\theta)$ across $\lambda \in \Lambda$. To this end, I develop the robust estimator in Section (ref) and the robust CI in Section (ref).
Since $\widehat{B}_{w}^{\beta}(\Lambda)$ is a valid bias UCB under any $w \in \mathcal{W}$, it provides a criterion for optimizing $w$. In the optimization, I allow for consideration sets $\mathcal{W}^{*} \subseteq \mathcal{W}$ that are nonempty, closed, and convex. I distinguish the weights in the optimization problem from the researcher's baseline weights, denoting the former by $\Bar{w} \in \mathcal{W}^{*}$ and the latter by $w \in \mathcal{W}$. I allow for classes $\Lambda$ that depend on the baseline $w$, but not on the optimizer variable $\Bar{w}$.
Given a consideration set $\mathcal{W}^{*}$, I define the robust estimator $\hat{\tau}^{*}$ as the weighted estimator induced by the minimax-bias weights $w^{*} \in \mathcal{W}^{*}$, given by
Let $\widehat{B}_{\min}^{\beta}(\Lambda) = \hat{\eta}_{1-\beta}\max_{\lambda \in \Lambda}\left\|\lambda - w^{*}\right\|_{\Sigma}$ denote the corresponding minimax-bias UCB.
Proposition (ref) establishes optimality of the minimax-bias weights $w^{*}$ from two perspectives: first, at the observed data, $w^{*}$ minimizes the bias $\widehat{B}_{\Bar{w}}^{\beta}(\Lambda)$ that is inferred ex post; and second, across hypothetical data realizations, $w^{*}$ minimizes the bias that is possible ex ante. Conveniently, the maximum bias from either perspective is proportional to the maximum distance: $w^{*}$ does not depend on the estimated heterogeneity $\hat{\eta}_{1-\beta}$ or on the heterogeneity bound $\eta$. In the above setup, the optimality of $w^{*}$ among weights $\Bar{w} \in \mathcal{W}^{*}$ is equivalent to optimality of the robust estimator $\hat{\tau}^{*}$ among weighted estimators $\hat{\tau}_{\Bar{w}}$.
By default, I take the consideration set to be the class of alternative weights: $\mathcal{W}^{*} = \Lambda$. In this case, $w^{*}$ can be interpreted as the weights that minimize worst-case disagreement across the class of alternative weights $\Bar{w} \in \Lambda$. Moreover, because the minimax-bias weights depend only on the class of alternative weights, they provide a natural default for researchers facing ambiguity over their initial choice of baseline weights. I now analyze the structure and interpretation of the minimax-bias weights $w^{*}$ for the example classes of $\Lambda$ from Section (ref).
I define the robust CI centered at $w \in \mathcal{W}$ as
where the critical value function $\text{cv}_{1-\alpha}\left(b\right)$ gives the $(1-\alpha)$-quantile of the folded normal distribution $|N(b,1)|$. Thus, the robust $CI_{w}^{*}$ parallels the baseline $CI_{w}$, but uses critical values that adjust for the inferred bias $\widehat{B}_{w}^{\beta}(\Lambda)$. This adjustment widens the baseline endpoints: $CI_{w} \subseteq CI_{w}^{*}$, with equality if and only if $\widehat{B}_{w}^{\beta}(\Lambda) = 0$. Provided that $\Lambda \neq \{w\}$, this equality holds if and only if the heterogeneity UCB is insignificant at level $\beta$ in the sense that $\hat{\eta}_{1-\beta} = 0$.
Proposition (ref) shows that $CI_{w}^{*}$ covers $\tau_{\lambda}(\theta)$ at confidence level $1-(\alpha+\beta)$, uniformly over $\lambda \in \Lambda$.\footnote{$CI_{w} \subseteq CI_{w}^{*}$ implies $\mathbb{P}_{\theta}\left\{\tau_{w}(\theta) \in CI_{w}^{*}\right\} \geq \mathbb{P}_{\theta}\left\{\tau_{w}(\theta) \in CI_{w}\right\} = 1-\alpha$ for all $\theta$. Thus, the robust $CI_{w}^{*}$ still covers the baseline estimand $\tau_{w}(\theta)$ at conventional confidence levels. The additional error level $\beta$ when considering alternative estimands $\tau_{\lambda}(\theta)$ is a Bonferroni correction to account for the first-step inference on heterogeneity; such Bonferroni corrections have been used for two-step construction of test statistics and CIs in other contexts, such as romano2014practical and mccloskey2017bonferroni.} This uniform coverage differs from the conventional coverage in (ref). The latter provides inference guarantees for one given $w$, while the former extends such guarantees to every $\lambda \in \Lambda$. Relative to conventional coverage, the price of uniform coverage is (i) a higher error level from using $\hat{\eta}_{1-\beta}$ to infer heterogeneity and (ii) longer intervals to obtain coverage that is robust to alternative weights $\lambda \in \Lambda$.
\paragraph{Interpretation.} Intuitively, if $\lambda \in \Lambda$ indexes the preferred weights of different readers, then the robust $CI_{w}^{*}$ provides valid coverage for each reader's estimand $\tau_{\lambda}(\theta)$. For example, suppose the baseline $CI_{w}$ excludes zero, leading the researcher to reject the $w$-null of no average effect $H_{0,w}: \tau_{w}(\theta) = 0$ at significance level $\alpha$. If zero remains excluded from the robust $CI_{w}^{*}$, then a reader with alternative weights $\lambda \in \Lambda$ can reject the $\lambda$-null of no average effect $H_{0,\lambda}: \tau_{\lambda}(\theta) = 0$ at significance level $\alpha + \beta$. In this sense, $CI_{w}^{*}$ facilitates robust inference.
\paragraph{Additional Coverage Properties.} In Appendix (ref), I show that $CI_{w}^{*}$ also provides a form of simultaneous coverage---but at a lower confidence level in the two-sided case. I compare this to the uniform coverage in (ref). In comparing the two coverage notions, I draw connections to the literature on inference for partially identified parameters molinari2020microeconometrics. In Appendix (ref), I derive an upper bound on the coverage rate of $CI_{w}^{*}$. The bound is strictly less than one for any $\lambda \in \Lambda$ under any degree of heterogeneity. This implies that $CI_{w}^{*}$ has nontrivial coverage for $\tau_{\lambda}(\theta)$ across all values of $\lambda \in \Lambda$ and $\theta$.
\paragraph{Choice of Centering Weights.} The robust $CI_{w}^{*}$ can be centered at any $w \in \mathcal{W}$ while maintaining uniform coverage over $\lambda \in \Lambda$. By default, I center the robust CIs at the minimax-bias weights $w^{*}$ and denote $CI^{*} = CI_{w^{*}}^{*}$. In particular, given the minimax-bias UCB $\widehat{B}_{\min}^{\beta}(\Lambda)$, robust estimator $\hat{\tau}^{*}$, and standard deviation $\sigma^{*} = \sqrt{(w^{*})'\Sigma w^{*}}$, the robust CIs become
This choice of center yields the smallest bias UCB $\widehat{B}_{\min}^{\beta}(\Lambda)$ across the consideration set weights. In this sense, $w^{*}$ prioritizes the component of robust CI length that comes from the ambiguity or disagreement over $\Lambda$. However, the resulting standard deviation $\sigma^{*}$ may be larger than the baseline $\sigma_{w}$, so the overall effect on length is ambiguous ex ante.\footnote{Nevertheless, when the maximum bias is fixed and positive, the bias UCB is asymptotically non-negligible while the variance shrinks to zero with the sample size, which supports choosing $w^{*}$ as the default center.} Note that when $\hat{\eta}_{1-\beta}$ is small enough, it is even possible for $CI^{*}$ to be shorter than $CI_{w}$, and hence shorter than $CI_{w}^{*}$. For example, since the bounded variance class yields $w^{*}=w_{\mathrm{GLS}}$, and GLS minimizes variance, the corresponding $CI^{*}$ must be weakly shorter than $CI_{w}$ in realizations where $\hat{\eta}_{1-\beta} = 0$.
\paragraph{Relationship to Bias-Aware CIs.} My robust CIs have a similar structure to the bias-aware CIs advanced in armstrong2018optimal,armstrong2020simple,armstrong2021finite,armstrong2021sensitivity, but there are important differences in model setup and inferential objective. The bias-aware CIs are designed to provide valid inference for a fixed target estimand $\lambda'\theta$ subject to bounds on the parameter space for $\theta$, such as bounds on heterogeneity---see kwon2025estimating for results of this flavor. By contrast, my robust CIs are designed to provide valid inference for a class of target estimands $\{\lambda'\theta: \lambda \in \Lambda\}$ and impose no bound on heterogeneity, opting instead to infer the heterogeneity using $\hat{\eta}_{1-\beta}$. A downside of bounding heterogeneity is the potential for undercoverage when the bound is incorrect, while a downside of inferring heterogeneity is the additional error level $\beta$ in the uniform coverage statement (ref).
The foregoing results are developed under normally distributed estimates $\hat{\theta} \sim N(\theta, \Sigma)$, a known covariance matrix $\Sigma$, known baseline weights $w$, and a known class of alternative weights $\Lambda$. In this section, I establish the uniform asymptotic validity of my procedures under asymptotically normal estimates, consistent covariance matrix estimators, consistent estimators for the baseline weights, and consistent estimators for the class of alternative weights. Sections (ref) and (ref) formalize the asymptotic setup and assumptions. Sections (ref), (ref), and (ref) establish asymptotic results for the class of alternative weights, maximum distance, and robust estimator. Sections (ref) and (ref) establish asymptotic results for the heterogeneity UCB and robust CI. For self-contained practical implementations of my inference procedures, see Section (ref).
There is a sample of size $n$ drawn from some unknown distribution $P_{n} \in \mathcal{P}_{n}$, where $\mathcal{P}_{n}$ is a class of distributions for samples of size $n$. Based on this data, the researcher constructs a vector of estimates $\hat{\theta}_{n} \in \mathbb{R}^{K}$ and a positive definite covariance matrix estimator $\Tilde{\Sigma}_{n} \in \mathbb{R}^{K \times K}$. Let $\widehat{\Sigma}_{n} = n\Tilde{\Sigma}_{n}$ denote the normalized covariance matrix estimator. Under distribution $P_{n}$, the sample objects $(\hat{\theta}_{n}, \widehat{\Sigma}_{n})$ have population analogues $(\theta(P_{n}), \Sigma(P_{n}))$. The number of groups $K \geq 2$ is fixed while the sample size $n \overset{}{\rightarrow}\infty$ grows. I leave the case of growing $K$ to future work.
I assume that the normalized vector of estimates $\sqrt{n}(\hat{\theta}_{n} - \theta(P_{n}))$ is uniformly asymptotically normal in the sense that it converges in bounded Lipschitz (BL) metric to the normal distribution $N(0, \Sigma(P_{n}))$, uniformly over $P_{n} \in \mathcal{P}_{n}$. Moreover, I assume that the vector of parameters $\theta(P_{n})$ is uniformly bounded (over $P_{n} \in \mathcal{P}_{n}$ and $n \geq 1$).
Uniform convergence in BL metric is a standard way to formalize uniform convergence in distribution. For example, when the components of $\hat{\theta}_{n}$ are regression coefficients or sample averages computed over unit-level observations, Assumption (ref) follows from bounds on the moments of the observations and bounds on the dependence across observations.
Next, I assume that the normalized sample covariance matrix $\widehat{\Sigma}_{n} = n\Tilde{\Sigma}_{n}$ is uniformly $\sqrt{n}$-consistent for the population covariance matrix $\Sigma(P_{n})$, and that the eigenvalues of $\Sigma(P_{n})$ are uniformly bounded above and away from zero.
Assumption (ref) is a rate condition on the accuracy of covariance matrix estimation. In iid settings, it can be obtained from higher-moment bounds, while in dependent settings it requires moment and dependence conditions strong enough to yield a uniform $\sqrt{n}$ rate. The uniform eigenvalue bounds ensure that $N(0, \Sigma(P_{n}))$ is uniformly tight and nondegenerate. Note that Assumption (ref) implies, for each $\varepsilon > 0$,
That is, $\widehat{\Sigma}_{n}$ is uniformly consistent for $\Sigma(P_{n})$ under Assumption (ref). The stronger $\sqrt{n}$-consistency is imposed to obtain uniformly valid inference on heterogeneity in Section (ref).
The researcher constructs a vector of baseline weights $\hat{w}_{n} \in \mathcal{W}$ and a nonempty, compact, and convex class of alternative weights $\hat{\Lambda}_{n} \subseteq \mathcal{W}$. The population analogues are $(w(P_{n}), \Lambda(P_{n}))$. The baseline and alternative estimators are $\hat{\tau}_{\hat{w}_{n},n}= \hat{w}_{n}'\hat{\theta}_{n}$ and $\hat{\tau}_{\lambda,n} = \lambda'\hat{\theta}_{n}$, where $\lambda \in \hat{\Lambda}_{n}$. The baseline and alternative estimands are $\tau_{w_{n}}(P_{n}) = w(P_{n})'\theta(P_{n})$ and $\tau_{\lambda}(P_{n}) = \lambda'\theta(P_{n})$, where $\lambda \in \Lambda(P_{n})$. If the baseline and alternative weights are known, then I define $\hat{w}_{n} = w(P_{n})$ and $\hat{\Lambda}_{n} = \Lambda(P_{n})$.
I assume that the sample baseline weights $\hat{w}_{n}$ are uniformly consistent for the population baseline weights $w(P_{n})$. Moreover, I assume that $w(P_{n})$ is uniformly bounded---i.e., the baseline weighting scheme cannot place arbitrarily large negative weight on any group.
For example, under Assumption (ref), it follows that Assumption (ref) holds for GLS weights
Assumption (ref) also holds for weights considered in event studies, such as those in callaway2021difference under uniform versions of their asymptotic linearity assumptions, and those in sun2021estimating under uniform versions of their moment conditions.
I suppose that the class of alternative weights $\hat{\Lambda}_{n}$ depends on the data through statistics $\hat{S}_{n} \in \mathcal{S}$ taking values in a metric space $(\mathcal{S},d_{\mathcal{S}})$. In particular, $\hat{\Lambda}_{n} = \Lambda(\hat{S}_{n})$ and $\Lambda(P_{n}) = \Lambda(S(P_{n}))$, where $S(P_{n})$ is the population analogue of $\hat{S}_{n}$. I assume the sample statistics $\hat{S}_{n}$ are uniformly consistent for the population statistics $S(P_{n})$. Moreover, $S(P_{n})$ is contained in a compact set $\mathbb{S} \subseteq \mathcal{S}$ and $\hat{S}_{n}$ is contained in $\mathbb{S}$ with probability uniformly approaching one.
For a bounded variance class with $\hat{S}_{n} = (\widehat{\Sigma}_{n}, \hat{w}_{n})$, Assumption (ref) follows from Assumptions (ref) and (ref), where $\mathbb{S}$ can be defined with constants $2\Bar{e}$ and $2\Bar{C}_{w}$; see the example below for details. For a truncated simplex class with $\hat{S}_{n} = \hat{w}_{n} \in \mathcal{W}_{+}$, Assumption (ref) follows from Assumption (ref) and $\mathbb{S} = \mathcal{W}_{+}$. For a covariate balance class with $\hat{S}_{n} = (\hat{w}_{n}, \hat{\bm{X}}_{n}) \in \mathcal{W}_{+} \times \mathbb{R}^{K \times M}$, where $\hat{\bm{X}}_{n}$ is an estimated covariate mean matrix, Assumption (ref) follows from Assumption (ref) and uniform consistency of $\hat{\bm{X}}_{n}$ to a population matrix $\bm{X}(P_{n})$ with $\left\|\bm{X}(P_{n})\right\| \leq \Bar{C}_{X}$ uniformly for some constant $\Bar{C}_{X} > 0$ and $\min_{m}\text{sd}(\bm{X}_{m}(P_{n})) \geq \varsigma$ uniformly for some constant $\varsigma > 0$, where the corresponding component of $\mathbb{S}$ can be defined with constants $2\Bar{C}_{X}$ and $\varsigma/2$; see the example below for details.
Let $\operatorname*{diam}(A) = \sup_{a,b \in A}\left\|a - b\right\|$ denote the diameter of a given set $A \subseteq \mathbb{R}^{K}$. I assume that the class of alternative weights has the following structure.
Condition (ref)(i) requires that, on the set $\mathbb{S}$ containing the population statistics $S(P_{n})$, the classes of interest can be represented as inequality constraints of continuous and convex functions $\lambda \mapsto g(\lambda,S)$ on a compact $\mathbb{W}$, which implies that the population classes $\Lambda(P_{n}) \subseteq \mathbb{W}$ are compact and convex. Condition (ref)(ii) requires that the functions $S \mapsto g(\lambda,S)$ in the representation are uniformly Lipschitz on $\mathbb{S}$, which helps in establishing continuity of $S \mapsto \Lambda(S)$ in the Hausdorff metric defined below in Section (ref). Condition (ref)(iii) requires that the inequality constraints are uniformly strictly feasible, which helps in establishing continuity of $S \mapsto \Lambda(S)$ and implies that the population classes $\Lambda(P_{n})$ are nonempty. Moreover, it implies classes are nondegenerate in the sense that the maximum distance between weights is uniformly bounded away from zero. In particular, the maintained assumptions combined with Lemma (ref) yield
Thus, condition (ref)(iii) rules out asymptotic regimes in which the class of alternatives collapses to a singleton, which helps in formulating general arguments for the asymptotic validity of the robust CIs. I now map the conditions of Assumption (ref) to the example classes of interest.
\paragraph{Notation for Population Objects.} When convenient, I will use the shorthands
and likewise for other objects that depend on $P_{n}$.
To formulate uniform consistency of the estimated class $\hat{\Lambda}_{n}$ for the population class $\Lambda_{n}$, I use a standard notion of distance between sets. Given two nonempty sets $A \subseteq \mathbb{R}^{K}$ and $B \subseteq \mathbb{R}^{K}$, let $\operatorname{dist}(a,B) = \inf_{b \in B}\left\|a-b\right\|$ and $\operatorname{dist}(b,A) = \inf_{a \in A}\left\|b-a\right\|$ denote the distance from $a \in A$ to $B$ and from $b \in B$ to $A$, respectively. The Hausdorff distance between $A$ and $B$ is defined as
Under Assumption (ref), Lemma (ref) shows that $S \mapsto \Lambda(S)$ is Lipschitz continuous in the Hausdorff metric $d_{H}$:
Based on this continuity, the following result shows that $\hat{\Lambda}_{n} = \Lambda(\hat{S}_{n})$ is a uniformly consistent estimator of $\Lambda_{n} = \Lambda(S_{n})$ in the Hausdorff metric.
The estimated and population maximum distances between weights are
Intuitively, $\left\|\lambda - w_{n}\right\|_{\Sigma_{n}}/\sqrt{n}$ is the standard deviation for $(\lambda - w_{n})'\hat{\theta}_{n}$ under the normal approximation $\hat{\theta}_{n} \overset{a}{\sim} N(\theta_{n}, \Sigma_{n}/n)$, while $\left\|\lambda - \hat{w}_{n}\right\|_{\widehat{\Sigma}_{n}}/\sqrt{n} = \left\|\lambda - \hat{w}_{n}\right\|_{\Tilde{\Sigma}_{n}}$ is the plug-in standard error. The following result shows that the estimated maximum distance is uniformly consistent for the population maximum distance.
The estimated and population minimax-bias weights are
The following result shows that the minimax-bias weights satisfy Assumption (ref).
To establish asymptotic optimality of the robust estimator
I first characterize the asymptotic maximum bias of weighted estimators $\hat{\tau}_{\hat{w}_{n},n}= \hat{w}_{n}'\hat{\theta}_{n}$ under the following asymptotic uniform integrability (UI) condition:
UI condition (ref) controls the tails of the squared estimation errors $\|\hat{\theta}_{n} - \theta_{n}\|^{2}$ and $\|\hat{w}_{n} - w_{n}\|^{2}$, ensuring that the corresponding mean squared errors uniformly converge to zero under Assumptions (ref) and (ref). Similar to the discussions for Assumptions (ref) and (ref), UI condition (ref) follows from moment and dependence bounds on the underlying observations used to construct $(\hat{\theta}_{n}, \hat{w}_{n})$.\footnote{By Markov's inequality, UI condition (ref) follows from bounds strong enough to yield uniformly bounded $L^{2+\delta}(P_{n})$ norms on the estimation errors for some $\delta > 0$; e.g., see van2000asymptotic.}
For the population heterogeneity in parameters $H_{n}(\theta_{n}) = \min_{\gamma \in \mathbb{R}}\left\|\theta_{n} - \bm{1}\gamma\right\|_{\Sigma_{n}^{-1}}$, denote the largest heterogeneity over $\mathcal{P}_{n}$ as $\Bar{\eta}(\mathcal{P}_{n}) = \supPnH_{n}(\theta_{n})$, which is uniformly bounded over $n$ under the eigenvalue bounds on $\Sigma_{n}$ and the norm bound on $\theta_{n}$. For each $w \in \mathcal{W}$, $P_{n} \in \mathcal{P}_{n}$, and $n$, let $P_{n}^{\dagger}(w)$ denote a distribution where $\Sigma(P_{n}^{\dagger}(w)) = \Sigma_{n}$, $S(P_{n}^{\dagger}(w)) = S_{n}$, $w(P_{n}^{\dagger}(w)) = w$, and
which satisfies $H_{n}(\theta(P_{n}^{\dagger}(w))) = \Bar{\eta}(\mathcal{P}_{n})$ whenever $\Lambda_{n} \neq \{w\}$. For example, if $P_{n}^{\dagger}(w_{n}) \in \mathcal{P}_{n}$ for all $P_{n} \in \mathcal{P}_{n}$ and $n$, then the class of distributions will be rich enough for the bias bound below to be sharp, in analogy to the sharpness of the bias bound from Proposition (ref). In this sense, the distributions $P_{n}^{\dagger}(w_{n})$ are least-favorable. In what follows, I denote $\tau_{\lambda,n} = \lambda'\theta_{n}$.
The following result establishes asymptotic optimality of the robust estimator, serving as the asymptotic analogue of Proposition (ref) from the normal model.
This optimality result presumes that $\hat{w}^{*}_{n}\,$ satisfies UI condition (ref). This holds, for example, when the class of alternatives $\Lambda(S)$ is uniformly bounded over $S \in \mathcal{S}$---not just $S \in \mathbb{S}$---as with any subset of the simplex. In other cases, it may be possible to argue the UI condition directly. For example, $\hat{w}^{*}_{n}\,$ reduces to the GLS weights in formula (ref) under the bounded variance class, from which it suffices to impose appropriate UI conditions on $\widehat{\Sigma}_{n}$. Note that even without UI conditions, $\hat{w}^{*}_{n}\,$ is still uniformly consistent for the weights $w^{*}_{n}$ that minimize the population maximum distance, as established in Proposition (ref).
The sample heterogeneity in parameters is the square root of the residual sum of squares from the GLS regression of the parameters $\theta_{n}$ on a constant $\bm{1}$ weighted by $\widehat{\Sigma}_{n}^{-1}$:
where $\hat{A}_{n}$ is the annihilator matrix for $\widehat{\Sigma}_{n}^{-1/2}\bm{1}$. The function $\hat{H}_{n}$ used to measure heterogeneity is random due to $\widehat{\Sigma}_{n}$. A different way to measure heterogeneity is the population heterogeneity in parameters, which replaces $\widehat{\Sigma}_{n}$ with $\Sigma_{n}$:
The heterogeneity in estimates is $\hat{H}_{n}(\hat{\theta}_{n})$. The heterogeneity UCB at significance level $\beta$ is
The following result establishes uniform asymptotic validity of $\hat{\eta}_{1-\beta,n}$ for inference on the sample heterogeneity in parameters $\hat{H}_{n}(\theta_{n})$, which will turn out to be sufficient for establishing uniform asymptotic validity of the robust CI---see Remark (ref) below for why I consider inference on the sample measure $\hat{H}_{n}(\theta_{n})$ instead of the population measure $H_{n}(\theta_{n})$.
Note that when $H_{n}(\theta_{n}) = 0$, the UCB $\hat{\eta}_{1-\beta,n} \geq 0$ always covers $\hat{H}_{n}(\theta_{n})$, so multiplication by the indicator $\mathds{1}\left\{H_{n}(\theta_{n}) > 0\right\}$ restricts attention to sequences where heterogeneity is nontrivial, in analogy to quantile-unbiasedness criterion (ref) from the normal model. Thus, (ref) implies
In this sense, $\hat{\eta}_{1-\beta,n}$ is a uniformly asymptotically valid UCB, in analogy to criterion (ref).
I now establish the asymptotic validity of robust CIs centered at weight vectors $\hat{w}_{n}$ that satisfy Assumption (ref), which include the minimax-bias weights $\hat{w}^{*}_{n}\,$ in view of Proposition (ref).
For centering weights $\hat{w}_{n}$, the bias UCB is
For standard error $\Tilde{\sigma}_{\hat{w}_{n},n} = \sqrt{\hat{w}_{n}'\Tilde{\Sigma}_{n} \hat{w}_{n}} = \sqrt{\hat{w}_{n}'\widehat{\Sigma}_{n} \hat{w}_{n}}/\sqrt{n} = \hat{\sigma}_{\hat{w}_{n},n}/\sqrt{n}$, the robust CI is
The validity of $CI^{*}_{\hat{w}_{n},n}$ depends on its coverage for weights in the population class $\Lambda_{n} \subseteq \mathbb{W}$, where $\mathbb{W} \subseteq \mathcal{W}$ is the compact convex set with positive diameter from Assumption (ref). The uniform consistency in Proposition (ref) shows that the estimated class $\hat{\Lambda}_{n}$ is close to the population class $\Lambda_{n}$ with high probability as $n \overset{}{\rightarrow}\infty$. However, this is not the same as $\hat{\Lambda}_{n}$ containing the alternatives from $\Lambda_{n}$ with high probability. To see this, recall that Assumption (ref) yields the inequality constraint representations
where the former holds with probability uniformly approaching one under Assumption (ref). The issue is at the boundary of the inequality constraint, where small estimation errors can lead to exclusion, $g(\lambda, \hat{S}_{n}) > 0$, of population alternatives where $g(\lambda, S_{n}) = 0$. Since membership is discontinuous at the boundary, uniform consistency of $\hat{S}_{n}$ to $S_{n}$ generally does not ensure that $\hat{\Lambda}_{n}$ contains the alternatives in $\Lambda_{n}$. However, as I show in examples further below, there will typically exist a subclass $\Lambda_{0}(S) \subseteq \Lambda(S)$ such that, for $\Lambda_{0,n} = \Lambda_{0}(S_{n})$,
For subclasses $\Lambda_{0,n}$ satisfying containment condition (ref), the robust CI constructed with $\hat{\Lambda}_{n}$ provides uniformly asymptotically valid coverage for the target alternatives $\lambda \in \Lambda_{0,n}$.
The asymptotic coverage in (ref) is analogous to the coverage in (ref) from the normal model, but stated for the subclass $\Lambda_{0,n}$ rather than the population class $\Lambda_{n}$. This suggests viewing the class $\Lambda_{0,n}$ in Proposition (ref) as a target class, $\Lambda_{n} \supseteq \Lambda_{0,n}$ as an enlargement of the target class, and $\hat{\Lambda}_{n}$ as an estimator of the enlargement $\Lambda_{n}$. In other words, to ensure asymptotic coverage for target alternatives $\Lambda_{0,n}$, it suffices to use an enlarged estimator $\hat{\Lambda}_{n} \supseteq \hat{\Lambda}_{0,n}$ to construct the robust CI rather than using the sample analogue $\hat{\Lambda}_{0,n}$ of the target. For example, if the class $\hat{\Lambda}_{0,n} = \Lambda(\hat{S}_{n}; \kappa_{0})$ increases in a scalar parameter $\kappa$, this suggests using $\hat{\Lambda}_{n} = \Lambda(\hat{S}_{n};\kappa)$ with a larger $\kappa > \kappa_{0}$. Importantly, unlike the enlarged $\Lambda_{n}$, I do not require the target $\Lambda_{0,n}$ to satisfy the conditions of Assumption (ref), so long as it satisfies containment condition (ref) relative to $\Lambda_{n}$. For example, a target bounded variance class with $r_{0}=1$ may not satisfy the Slater condition, but nevertheless satisfy containment condition (ref) relative to an enlargement with $r > 1$ that does satisfy the Slater condition; see the examples below.
To establish containment condition (ref), a practical approach is to show that the inequality constraint from $\Lambda_{n}$ is uniformly slack when evaluated over $\Lambda_{0,n}$. In particular, it suffices to show the existence of a constant $\nu > 0$ such that for all $P_{n} \in \mathcal{P}_{n}$ and $n$,
By Lemma (ref), slack condition (ref) implies containment condition (ref) under Assumptions (ref) and (ref). I now show how to derive $\Lambda_{0,n}$ and $\nu$ for the benchmark classes of interest.
Motivated by the asymptotic results in Section (ref), but suppressing the sample size index $n$ for conciseness, I consider asymptotically normal estimates $\hat{\theta} \overset{a}{\sim} N(\theta, \Sigma/n)$, consistent covariance matrix estimator $\Tilde{\Sigma} = \widehat{\Sigma}/n \overset{a}{\approx} \Sigma/n$, consistent baseline weights estimator $\hat{w} \overset{a}{\approx} w$, and classes of alternative weights $\hat{\Lambda} = \Lambda(\hat{S})$ characterized by consistently estimated statistics $\hat{S} \overset{a}{\approx} S$: e.g., $\hat{S} = (\hat{w}, \widehat{\Sigma})$ for bounded variance class, $\hat{S} = \hat{w}$ for truncated simplex class, and $\hat{S} = (\hat{w}, \hat{\bm{X}})$ for covariate balance class with estimated covariate mean matrix $\hat{\bm{X}}$.
Section (ref) summarizes my robust inference procedures in terms of the above objects and discusses specification choices for $\hat{\Lambda}$. Sections (ref) and (ref) discuss practical implementations of these procedures in the context of event studies and multisite experiments, respectively.
The GLS regression of $\hat{\theta}$ on a constant $\bm{1}$ weighted by $\Tilde{\Sigma}^{-1}$ yields
where $\hat{\tau}_{\mathrm{GLS}}$ is the GLS estimator, $\Tilde{\sigma}_{\mathrm{GLS}}$ is the GLS standard error, $\Tilde{H}^{2}(\hat{\theta})$ is the GLS residual sum of squares, which yields heterogeneity in estimates $\Tilde{H}(\hat{\theta})$ with explicit formula
Following the steps of Section (ref), the heterogeneity UCB is
Following the steps of Section (ref), the minimax-bias weights and bias UCB are
For standard error $\Tilde{\sigma}^{*} = \sqrt{(\hat{w}^{*})'\Tilde{\Sigma} \hat{w}^{*}}$, the robust estimator and corresponding robust CI are
where $\text{cv}_{1-\alpha}\left(b\right)$ can be computed as the square root of the $(1-\alpha)$-quantile of the noncentral chi-squared distribution with one degree of freedom and noncentrality parameter $b^{2}$.
Recipe (ref) summarizes the implementation of my inference procedures for inputs given by (i) the estimates and covariance matrix estimator $(\hat{\theta}, \Tilde{\Sigma})$, for which I describe typical constructions in Section (ref) for event studies and Section (ref) for multisite experiments; (ii) the significance levels $\alpha$ and $\beta$, which by default can be set to $\alpha = \beta = 0.05$; and (iii) the class of alternatives $\hat{\Lambda}$.
In some cases there may be ambiguity over which class of alternatives to use. In particular, the class $\hat{\Lambda} = \hat{\Lambda}(\kappa)$ may be indexed by a scalar parameter $\kappa \in \mathbb{R}$ over which there is ambiguity. To address this, one can execute Recipe (ref) for a range of plausible $\kappa$. Alternatively, assuming that $\hat{\Lambda}(\kappa)$ is increasing in $\kappa$, one can compute the breakdown value $\kappa^{*}$, which is the smallest value of $\kappa$ at which the robust $CI^{*}(\kappa)$ includes a given threshold value $\tau_{0}$, such as $\tau_{0} = 0$.
\paragraph{Context.} Following the exposition in Section (ref), an event study has units $i$ treated at different time periods $t$, allowing one to estimate the dynamic effects of treatment on outcomes $Y_{it}$. In this setting, two key sources of heterogeneity are (i) across treatment cohorts $G_{i}$, which denote the periods of initial treatment for units $i$; and (ii) across event times $\ell \geq 0$, which denote the number of periods since initial treatment. Given the $\operatorname{ATT}_{g,t}$ object in (ref), the average treatment effect $\ell$ periods after initial treatment for cohort $G_{i} = g$ is
Under event-study assumptions, sun2021estimating show that event-time $\ell$ coefficients from conventional TWFE regression specifications recover estimands of the form
which allow treatment effects from different event times $\ell' \neq \ell$ and negative weights in $\Tilde{w}_{\ell}$. As surveyed in roth2023s, the recent event studies literature proposes heterogeneity-robust (HR) approaches that, for weights $w_{\ell} > 0$, recover event-time estimands of the form
preventing negative weighting and excluding treatment effects from different event times.
For each target event time $\ell$, a researcher may report conventional estimators and CIs for both a TWFE specification and one of the proposed HR approaches and examine the stability of results. There are two ways to map this setup to the notation of Recipe (ref).
While it is possible to do the former (see sun2021estimating), the latter has the advantage of automatically shutting off the $\ell' \neq \ell$ channels and reducing the number of objects to be estimated. In fact, the ingredients required for implementing Recipe (ref) are available in many applications of HR approaches. For example, the sun2021estimating procedure already requires estimation of $(w_{\ell,g}, \operatorname{ATT}_{g,g+\ell})$, and so Recipe (ref) can be readily executed in applications that were already using the sun2021estimating procedure, or any other procedure that first estimates ATT objects and their corresponding standard errors, such as the doubly robust procedure of callaway2021difference, or stacked DiD procedures wing2024stacked.
\paragraph{Implementation.} I now use $k$ to index the groups $g$. For each target event time $\ell$, denote the number of available cohorts by $K_{\ell}$. Take an existing HR approach for which one can estimate baseline weights $\hat{w}_{\ell} \overset{a}{\approx} w_{\ell}$ and ATT objects
where $\Tilde{\Sigma}_{\ell}$ is a covariance matrix estimator. Let $\Tilde{\sigma}_{\hat{w},\ell} = \displaystyle\sqrt{\hat{w}_{\ell}'\Tilde{\Sigma}_{\ell}\hat{w}_{\ell}}$ denote the plug-in standard error for $\hat{\tau}_{\hat{w},\ell} = \hat{w}_{\ell}'\hat{\theta}_{\ell}$. Following equation (ref), consider the bounded variance simplex class
Executing Recipe (ref) with $\hat{\Lambda}_{\sigma, \ell}^{+}(r)$ produces $(\hat{\tau}_{\ell}^{*}(r), CI_{\ell}^{*}(r))$. For the choice of standard error ratio bound $r$, I propose the following implementation.
I apply this implementation to lakdawala2023dynamic in Section (ref).
Following the exposition in Section (ref), a multisite experiment has sites $k$ containing units $i$ over which ATEs are estimated:
where $\Tilde{\Sigma}$ is a covariance matrix estimator. Given independent sites, it suffices to obtain site-level objects $(\hat{\theta}_{k}, \Tilde{\sigma}_{k}, \widehat{E}_{k}[X_{i}])$, where estimated variances $\Tilde{\sigma}_{k}^{2}$ can be collected into a diagonal matrix $\Tilde{\Sigma}$ and estimated covariate means $\widehat{E}_{k}[X_{i}]$ can be collected into a matrix $\hat{\bm{X}} = (\widehat{E}_{1}[X_{i}], \ldots, \widehat{E}_{K}[X_{i}])'$. Given microdata on units $i$ from the sites, one can directly estimate $\widehat{E}_{k}[X_{i}]$ with sample means and $(\hat{\theta}_{k}, \Tilde{\sigma}_{k})$ via regression. For example, letting $\mathds{1}\left\{i \in k\right\}$ be an indicator for membership in site $k$, one can estimate the (no-intercept) saturated model
with OLS to obtain ATE estimates $\hat{\theta}_{k}$ and heteroskedasticity-robust standard errors $\Tilde{\sigma}_{k}$.
The baseline weights $\hat{w} \in \mathcal{W}_{+}$ represent a particular population/distribution over the set of sites. To model departures from the baseline population, one can use the truncated simplex class $\hat{\Lambda}_{+}(\epsilon)$ defined in (ref) or the covariate balance class $\hat{\Lambda}_{X}(\Bar{c})$ defined in (ref).
To obtain benefits from both approaches, I propose the following implementation based on the truncated covariate balance class $\hat{\Lambda}_{X}^{\epsilon}(\Bar{c})$ defined in (ref).
I apply this implementation to Project STAR in Section (ref).
I now apply my procedures to the event study in lakdawala2023dynamic and the multisite experiment in Project STAR, following the implementations described in Sections (ref) and (ref). Below I use significance levels $\alpha = \beta = 0.05$ and hence confidence levels $1-\alpha = 95\%$ for conventional CIs, $1-\beta = 95\%$ for heterogeneity UCBs, and $1-(\alpha+\beta)=90\%$ for robust CIs.
lakdawala2023dynamic---henceforth LNK---study the rollout of school-based internet access in Peruvian public primary schools. The estimating sample contains schools that either gained internet access between 2007 and 2020 or remained unconnected by 2020, with second grade math and reading test scores observed from 2007--2016. LNK use a staggered-adoption event-study design: after absorbing school fixed effects, calendar-time controls, and time-varying school and student controls, the event-time coefficients compare changes in scores around a school's installation date to contemporaneous changes among schools that are not yet connected or remain unconnected. LNK's event-study figures report school-level dynamic effects for successive cohorts of second graders, with scores standardized within calendar year and effects interpreted relative to the year before internet installation.
LNK's Figure 2 shows a delayed achievement response. The TWFE event-study estimates, replicated in Figure (ref), are small immediately after internet installation and grow over time. In LNK's page 234 description of the medium-run effect, “By year 5, scores are 0.110 standard deviations higher for math” and “0.063 standard deviations higher for reading” relative to the year before installation. These dynamics are central to the interpretation of the intervention because LNK argue that schools require time to adapt to new internet access. The empirical target is therefore a dynamic path of treatment effects.
\paragraph{Weighting Issues.} Following Section (ref), let cohorts $k$ index the first year in which a school gains internet access and event times $\ell$ denote the time relative to first internet access. LNK's pages 241-42 discuss weighting issues: “Recent work \ldots raises some important issues with using two-way fixed effects estimation when there is staggered timing of treatment. We show that these issues do not explain our main estimates by illustrating robustness to the estimator proposed in Sun and Abraham (2021) \ldots The patterns and magnitudes of these estimates are very similar to our main estimates.” The sun2021estimating---henceforth SA---estimator corresponds to weights $\hat{w}_{\mathrm{SA},\ell}$. LNK's Figure A.12, replicated in Figure (ref), is therefore a robustness check that replaces the TWFE event-study aggregation with the SA aggregation in order to assess whether results are robust to alternative weights.\footnote{LNK use a bootstrap procedure to construct the SA CIs. Their replication files estimate cohort-specific event-study coefficients using the same controls and fixed effects as the main specification, drop cohorts not identified relative to the omitted period, aggregate cohort-event coefficients using the SA event-time weights, and bootstrap the aggregated coefficients using 500 replications stratified by first access year and clustered by school. The CIs reported here instead use the plug-in standard error computed from the estimated covariance matrix $\Tilde{\Sigma}_{\ell}$ for the cohort-event coefficient vector $\hat{\theta}_{\ell}$, treating the SA weights as fixed. Thus, while I replicate the same SA estimates, my reported SA CIs are not identical to LNK's bootstrap SA CIs.} This suggests a sharper robustness question: are the substantive conclusions robust only to the SA weights, or to a broader class of plausible event-time weights? To answer this question, I compute robust estimators and CIs for classes of alternative weights that nest the SA weights.
\paragraph{Measuring Heterogeneity.} Table (ref) reports the estimated event-time heterogeneity UCBs $\Tilde{\eta}_{\ell}$ used to construct the robust CIs. For event times with $\Tilde{\eta}_{\ell}=0$, the heterogeneity adjustment in the respective robust CI drops out. The last row reports $K_{\ell}$, the total number of cohorts $k$ at each event time $\ell$. Across all but one event time, the heterogeneity UCB is lower for the math cohort-level estimates than for the reading cohort-level estimates, previewing the greater degree of robustness for the math results.
\paragraph{Bounded Variance Simplex Class.} Consider the bounded variance simplex
with $r=1$, which corresponds to the class of simplex-weighted estimands that can be estimated at least as precisely as the SA estimand. For each event time $\ell$, I compute the robust $CI^{*}_{\ell}(1)$ centered at the robust estimator $\hat{\tau}_{\ell}^{*}(1)$. Figure (ref) presents three series: the TWFE CI centered at the TWFE estimate, the SA CI centered at the SA estimate, and the robust CI centered at the robust estimate. For math, the robust CIs are positive for event times 1 through 6 and only include zero at event time 0. At event time 5, the robust estimate is $0.099$ with CI $[0.055,0.143]$, close to both the TWFE and SA estimates. Thus, LNK's medium-run math conclusion is robust to alternative simplex weights that yield estimators at least as precise as the SA estimator. For reading, the robust estimates remain close to the SA path, but the conclusion that results are significantly positive is somewhat more sensitive. The robust CIs are positive for event times 1 through 4 and 6, while event time 5 barely includes zero: the robust estimate is $0.042$ with CI $[-0.0002,0.084]$.
Table (ref) presents breakdown values $r^{*}_{\ell}$, which for each $\ell$ gives the smallest value of $r$ at which the robust $CI^{*}_{\ell}(r)$ includes zero for the class $\hat{\Lambda}_{\sigma,\ell}^{+}(r)$. A value of $1.60$, for example, means that the robust CI begins to include zero once the allowable $\Tilde{\sigma}$ bound is expanded to about $1.60$ times the SA $\Tilde{\sigma}$. A dash means that zero is not reached even when considering the simplex without the variance constraint.\footnote{If $\Tilde{\eta}_{\ell}=0$, I still compute breakdown values using the same definition: report the smallest feasible value of the robustness parameter at which the robust CI contains zero, and report “--” if zero is never reached over the full class. This convention implies, for example, that math at $\ell=0$ has $r^{*}_{0}<1$ because zero is already included at the smallest feasible bounded-variance simplex class, while math at $\ell=6$ has no finite breakdown value because zero is not reached even when the variance constraint is no longer binding over the simplex.} The breakdown values reinforce the visual evidence. For math, event times 2 through 5 tolerate substantial increases in $r$ before the robust CI includes zero, with $r^*_5=1.60$ at the headline year-5 event time. Event time 6 is even stronger in this sense: zero is not reached over the full simplex. For reading, by contrast, the breakdown values are close to one for event times 1 through 5, and event time 5 already includes zero at $r=1$. Overall, these results support the baseline math conclusions strongly, but not as strongly for the reading conclusions.
The Tennessee Project STAR (Student/Teacher Achievement Ratio) experiment randomized students in seventy-nine Tennessee public elementary schools to classes with different numbers of students to estimate the causal effects of class size on test scores achilles2008tennessee, krueger1999experimental. The experimental arms were (i) regular-sized classes (20-25 students), small classes (13-17 students), and regular-sized classes with a teacher aide. To define the sample for my analysis, I drop observations in the teacher-aide arm and focus on the effect of being assigned to a small class versus a regular-sized class. Following goldsmith2024contamination, I focus on the effects of treatment in kindergarten, which mitigates attrition and other complications. This leaves $K=78$ schools and a total of $n=3783$ students. Following schanzenbach2006have, I define the outcome as the average of a student's math and reading scores on the SAT, where the scores are demeaned and standardized relative to control-group scores. This puts the outcome into effect size units that have benchmarks in causal studies of preK-12 education interventions on student achievement: <0.05 = small, 0.05-0.20 = medium, >0.20 = large kraft2020interpreting.
In my notation, $\theta_{k}$ is the ATE of being assigned to a small class on the test score outcome at school $k$, $\hat{\theta}_{k}$ is the estimated ATE at school $k$ from regressing the outcome on school-specific treatment indicators and fixed effects, $\Tilde{\Sigma}$ is the heteroskedasticity-robust covariance matrix from this regression, and the baseline weights are equal weights (EW) $w = w_{\mathrm{EW}} = \bm{1}/K$ so that $\tau_{w}(\theta) = \tau_{\mathrm{EW}}(\theta)$ is the ATE in the empirical distribution of Project STAR schools. The baseline results establish medium-sized positive effects of small class size on test scores. In particular, the EW estimate is $\hat{\tau}_{\mathrm{EW}}=0.188$ with standard error $\Tilde{\sigma}_{\mathrm{EW}}=0.028$, yielding conventional EW CI $[0.133, 0.242]$. My baseline estimate is consistent with the one in schanzenbach2006have. However, my standard error is about 0.01 smaller. In particular, schanzenbach2006have clusters her standard errors at the classroom level. By contrast, I follow goldsmith2024contamination and use heteroskedasticity-robust standard errors since the randomization of students to classrooms was at the individual level.\footnote{This is consistent with the when-to-cluster framework of abadie2023should.}
\paragraph{Weighting Issues.} As reviewed by schanzenbach2006have, Project STAR has been used in many policy discussions. However, schanzenbach2006have notes external validity concerns:
Thus, the Project STAR results may not be representative of Tennessee (or U.S.) schools more broadly. Formally, the EW estimand averages over the empirical distribution of STAR schools, while alternative estimands of interest may place unequal weights across the STAR schools. This suggests a robustness question: are the baseline STAR results robust to alternative choices of weights/distributions? To answer this question, I compute robust estimators and CIs for classes of alternative weights that nest the EW weights, following implementations from Section (ref).
\paragraph{Measuring Heterogeneity.} The estimated heterogeneity UCB is $\Tilde{\eta} = 16.214$, which is over twice as large as the baseline $t$-statistic $\hat{\tau}_{\mathrm{EW}}/\Tilde{\sigma}_{\mathrm{EW}} = 6.714$. This suggests a meaningful degree of heterogeneity in the school-level ATE estimates, which previews the sensitivity of baseline results to the choice of weights.
\paragraph{Truncated Simplex Class.} I first consider the truncated simplex class
which collects the set of $\epsilon$-deviations in weight distribution from the STAR empirical distribution. For each $\tau_{0} \in \{0, 0.05\}$, Table (ref) presents the breakdown value $\epsilon^{*}$, which is the smallest value of $\epsilon$ at which $\tau_{0}$ is included in the robust $CI^{*}(\epsilon)$ centered at robust estimate $\hat{\tau}^{*}(\epsilon)$ with standard error $\Tilde{\sigma}(\epsilon^{*})$. The breakdown value for $\tau_{0}=0.05$ is $\epsilon^{*}=0.0169$, while the breakdown value for $\tau_{0}=0$ is $\epsilon^{*}=0.0261$. These are both small departures from EW: even though the corresponding robust estimates are larger than the baseline EW estimate $\hat{\tau}_{\mathrm{EW}}=0.188$, the robust CIs indicate that baseline inferences are robust to only small departures from EW.
\paragraph{Truncated Covariate Balance Class.} I intersect the truncated simplex with the covariate balance class, yielding
Table (ref) reports the Project STAR school-level covariates $\hat{\bm{X}}_{m}$ used in my analysis.
I fix $\epsilon = \epsilon^{*}(0) = 0.0261$ at the truncated simplex breakdown value for $\tau_{0} = 0$ from the previous analysis and consider balance gap bounds $(\Bar{c}_{d})_{d=1}^{9}$ equal to the deciles of the school-level balance gaps $c_{k}(\hat{\bm{X}})$. Intuitively, $c_{k}(\hat{\bm{X}})$ is the balance gap from a population that has covariate means $\hat{\bm{X}}_{m,k}$ equal to those of school $k$. The deciles $\Bar{c}_{d}$ presented in Table (ref) therefore represent typical values for the balance gaps.
Figure (ref) presents robust intervals $CI^{*}(\Bar{c}_{d})$ centered at robust estimates $\hat{\tau}^{*}(\Bar{c}_{d})$ under $\epsilon^{*}(0)$. Across every balance gap decile $\Bar{c}_{d}$, the robust CIs include $\tau_{0} = 0.05$. Thus, while the covariate balance constraints allow the robust CIs to exclude zero relative to imposing only the truncated simplex constraint, this is not enough to allow one to robustly conclude that effects are medium-sized under $\epsilon^{*}(0)$-departures from EW. Overall, these results provide further support for the conclusion that baseline STAR inferences are robust to only small departures from EW. There is too much heterogeneity for the conclusions of a baseline CI to be (uniformly) informative for alternative weights that depart meaningfully from the baseline weights.
Weighted estimands are ubiquitous in empirical research. The choice of weights is not merely a statistical nuisance: it shapes the empirical and policy relevance of a researcher's weighted estimand. Conventional inference procedures can therefore be limited when readers entertain alternative weights, leading to ambiguity and disagreement over the choice of weights.
This paper develops inference procedures that account for these issues. The key geometric observation is that the difference between two weighted estimands can be controlled by two objects: the heterogeneity of the underlying parameters and the distance between the corresponding weights. This decomposition leads to two practical tools. First, I derive minimax-bias weights that minimize the maximum distance to a class of alternative weights. The resulting robust estimator provides a natural default when researchers face ambiguity over the choice of weights. Second, I construct robust confidence intervals by combining a confidence bound for the heterogeneity in parameters with the maximum distance between the centering weights and the alternative weights. These intervals provide coverage guarantees for classes of weighted estimands rather than for a single baseline estimand.
My framework accommodates many empirically relevant classes of alternatives. The bounded variance class captures concerns about comparing the baseline estimand to alternatives that can be estimated with reasonable precision. The truncated simplex class captures departures from a baseline target population while ensuring that all represented groups continue to receive positive weight. The covariate balance class captures external validity concerns by restricting attention to alternative populations whose covariate means remain close to the baseline. These classes can also be intersected, allowing researchers to combine precision, positivity, and covariate-balance restrictions in a single robustness analysis.
The empirical applications illustrate that robustness to the choice of weights can be context-specific. For the event study in lakdawala2023dynamic, conclusions are broadly robust to classes of weights that yield reasonable estimation precision. For the experiment in Project STAR, however, conclusions are sensitive even to small departures from the baseline weighting scheme. These findings highlight the value of reporting robust estimators and intervals alongside conventional ones: doing so makes explicit whether the conclusions reached under a researcher's baseline weights are informative for alternative weighted estimands that may be of interest to readers.