EconBase
← Back to paper

Robust Inference for Weighted Estimands

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

178,471 characters · 26 sections · 88 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Inference for Weighted Estimands

abstractResearchers often conduct inference on weighted estimands, defined as weighted averages of group-level effects. Example settings include event studies with cohort-level effects and experiments with site-level effects. Under heterogeneous effects, different weighting schemes yield estimands with distinct empirical and policy interpretations, leading to ambiguity and disagreement over the choice of weights. I establish bounds on differences between weighted estimands and confidence bounds on effect heterogeneity, which I use to construct estimators that minimize worst-case bias and confidence intervals that are uniformly valid over classes of weighted estimands. I apply these methods to an event study in Lakdawala, Nakasone, and Kho (2023), which studies the effects of school-based internet access on test scores. I find that results are robust to broad classes of weights. I then apply the methods to Tennessee's Project STAR experiment and find that results are sensitive to small departures from baseline weights.

Introduction

In empirical research, many estimands can be expressed as weighted averages of group-level parameters.\footnote{Under various assumptions, ordinary least squares (OLS) estimands can be expressed as weighted averages of conditional average treatment effects (CATEs), conditional mean derivatives, or other regression parameters that may vary across covariate cells and treatment intensities yitzhaki1996using, angrist1998estimating, aronow2016does, sloczynski2022interpreting, goldsmith2024contamination; two-stage least squares (TSLS) estimands can be expressed as weighted averages of local average treatment effects (LATEs), which may vary across compliance types and covariate cells imbens1994identification, heckman2005structural, angrist2010extrapolate, kolesar2013ivheterogeneity, sloczynski2020should, huntington2020instruments, coussens2021improving, blandhol2022tsls, abadie2024instrumental; and two-way fixed effects (TWFE) estimands can be expressed as weighted averages of average treatment effects on the treated (ATTs), which may vary across treatment cohorts and time periods de2020two, callaway2021difference, goodman2021difference, sun2021estimating, athey2022design, gardner2022two, borusyak2024revisiting, wing2024stacked.} These weighted estimands give concise summaries of the parameters of interest, making them frequent targets for statistical inference. However, conventional inference procedures can be sensitive to the choice of weights: conclusions reached under a researcher’s baseline weights may not hold under alternative weights entertained by readers.

For example, in event studies with multiple treatment cohorts, many common estimands recover weighted averages of cohort-level causal effects, including conventional two-way fixed effects (TWFE) regression estimands and recently proposed alternatives.\footnote{For a review, see roth2023s.} TWFE has a convenient and flexible structure, which can yield lower-variance estimators in practice armstrong2025adapting. However, TWFE can place negative weight on some of the cohort-level effects, complicating the interpretation of the resulting estimands under heterogeneous effects de2020two, goodman2021difference, sun2021estimating, borusyak2024revisiting. Recent alternatives allow a researcher to target weighted estimands that rule out negative weights by design, but this introduces an additional degree of freedom: the researcher must decide which positive weights to use.

This choice matters because different readers may prefer different weights, depending on the empirical or policy question that they have in mind callaway2021difference. One reader may wish to upweight early-treated cohorts, another later-treated cohorts; one may care about short-run effects, another long-run effects; one may prefer weights that represent the characteristics of a policy-relevant target population, another may favor weights that trade off representativeness against estimation precision. These alternatives may all be reasonable, but they need not lead to the same conclusion under heterogeneous effects. Given the potential for disagreement, researchers may wish to assess the robustness of conclusions to the choice of weights. Current practice seems to reflect this desire by reporting results for different weighting choices, such as default implementations of the recently proposed alternatives. But such exercises consider only a finite set of weighting choices---which may not cover the range of plausible weighting schemes---and they generally do not provide formal inferential guarantees for claims that conclusions are robust to the choice of weights.

In this paper, I develop robust inference procedures that directly account for ambiguity and disagreement over the choice of weights. I consider a general framework in which a researcher observes asymptotically normal estimates for a vector of parameters. The researcher reports a conventional point estimator and confidence interval (CI) for a baseline estimand defined by a vector of baseline weights. At the same time, the researcher faces a broad class of alternative weights, giving rise to alternative estimands for which the baseline estimator may be biased and the baseline CI may undercover. To address these concerns, I develop robust estimators that mitigate worst-case bias and robust CIs that ensure valid coverage.

To proceed, I first establish a sharp bound on the difference between the baseline weighted estimand and any given alternative weighted estimand. This bound is the product of (i) the heterogeneity in parameters, defined as the square root of the residual sum of squares from a generalized least squares (GLS) regression of the parameters on a constant and (ii) the distance between weights, defined as the standard deviation of the difference in estimators under the baseline and alternative weights.

To develop my estimator, I study the problem of choosing weights to minimize the maximum distance between weights over a class of alternative weights. For a broad range of such classes, I show that this problem admits a unique solution. Moreover, I show that the corresponding estimator is optimal for minimizing the maximum bias over the class of alternative estimands under any bound on the heterogeneity in parameters. I therefore refer to this estimator as the robust estimator, and to the corresponding weights as the minimax-bias weights. I propose that researchers report the robust estimator alongside the baseline estimator. Intuitively, the robust estimator accounts for the bias concerns of readers who may disagree with the baseline weights. Moreover, since the minimax-bias weights depend only on the class of alternative weights, they provide a natural default for researchers facing ambiguity over their initial choice of baseline weights.

To develop my CI, I study the problem of constructing an upper confidence bound (UCB) for the heterogeneity in parameters. I show that a monotonic transformation of the heterogeneity in estimates produces an optimal UCB. This heterogeneity UCB—together with the maximum distance between weights—yields a simple adjustment to the critical values of a baseline CI. The resulting robust CI provides uniformly valid coverage over the class of alternative estimands. In particular, for each alternative estimand, the coverage probability is at least one minus the sum of (i) the significance level of the baseline CI and (ii) the significance level of the heterogeneity UCB. For example, suppose a 95% baseline CI excludes zero, leading the researcher to reject the null of no average effect at the 5% level. If zero remains excluded from a robust CI constructed using a 95% heterogeneity UCB, then readers with alternative weights can robustly reject the null of no average effect at the 10% level. To facilitate such inference procedures, I propose that researchers report the robust CI alongside the baseline CI.

The robust estimator and CI require the researcher to specify a class of alternative weights. My framework accommodates any compact and convex class of alternatives. I highlight three such classes that cover a broad range of empirical contexts. Below I give a high-level overview of these classes, deferring precise definitions and details to Section (ref) of the paper.

The first class is the bounded variance class, which restricts attention to weights that yield estimators with variance no larger than a given bound, ensuring that the alternative estimands can be estimated with a reasonable degree of precision. For example, when assessing robustness of results, a researcher may wish to consider alternative estimands that can be estimated at least as precisely as the baseline. Under the bounded variance class, my proposed procedures admit a convenient reduction to the GLS regression of the group-level estimates on a constant, where (i) the GLS estimator coincides with the robust estimator, (ii) the GLS variance determines the maximum distance between weights for any given choice of variance bound, and (iii) the GLS residual sum of squares determines the heterogeneity UCB. This reduction has two key implications. First, because the GLS estimator minimizes variance, the robust estimator is optimal for both bias and variance under the bounded variance class. Second, due to the structure in (ii) and (iii), readers can construct robust CIs for their own choices of variance bounds, provided the researcher reports the GLS variance and heterogeneity UCB. These convenient properties make the bounded variance class a natural benchmark.

The simplex is the set of nonnegative weights, which yields weighted estimands that do not extrapolate beyond the range of the parameters. However, the unrestricted simplex allows for some groups to receive zero weight, which can be undesirable in practice. To ensure that groups are represented, one can impose a floor on the simplex. The truncated simplex class is precisely the set of simplex weights that are bounded below by the given floor. Under positive baseline weights, the floor can be parameterized so that the truncated simplex yields the set of weights that are a given fraction between the baseline weights and the unrestricted simplex weights. In this form, the truncated simplex class allows one to investigate perturbations of a chosen magnitude from the baseline weights.

The simplex weights can be viewed as probability distributions over the groups, and thus as different target populations that readers may be interested in. For instance, when the groups index the sites of an experiment, baseline results may not generalize to future sites of interest to policymakers allcott2015site. If the researcher observes site-level covariates, however, the differences in covariate means under the baseline and alternative weights can inform the extent of external validity. Based on this idea, the covariate balance class considers simplex weights whose covariate means differ from the baseline mean by at most a given constant. Under this class, readers can assess whether results are robust to alternative populations whose covariate means are not too far from the baseline population.

I illustrate my framework in two empirical applications. The first application is an event study in lakdawala2023dynamic, which uses the staggered rollout of internet access in Peruvian public schools to estimate the causal effects of school-based internet access on second grade test scores. The authors find a delayed achievement response: TWFE estimates are initially small but grow over time, indicating that schools require time to adapt to new internet access. The authors find similar patterns and magnitudes when using the sun2021estimating estimator to account for negative weighting concerns, concluding that results are not sensitive to the use of TWFE. My robust estimator and CI give a stronger conclusion: results are robust to classes of nonnegative weights that yield comparable estimation precision to the sun2021estimating weights. In particular, robustness only fails when attempting to cover weighted estimands that cannot be estimated precisely in the first place.

The second application is Tennessee's Project STAR (Student/Teacher Achievement Ratio) experiment, which randomized students in seventy-nine Tennessee public elementary schools to classrooms of different sizes to estimate the causal effects of class size on test scores achilles2008tennessee, krueger1999experimental. Project STAR has been used in many policy discussions, but the experiment's selection of schools raises concerns about whether results are representative of Tennessee or U.S. schools more broadly schanzenbach2006have. To investigate these issues, I use an equal weights baseline to model the distribution of STAR schools and the truncated simplex class to model departures from the STAR empirical distribution. I show that, while the baseline CI implies medium-sized positive effects, the robust CI includes small-to-zero effect sizes even under small departures from equal weighting. The robust CI shrinks when I intersect the truncated simplex with the covariate balance class based on school-level covariates, but results remain sensitive to small departures from the baseline weights.

While I focus on event studies and multisite experiments as the motivating examples, my framework applies broadly to settings where group-level parameters are averaged using weights that sum to one. My inference procedures are not designed for continuously indexed groups, but the bound on differences between weighted estimands has a natural analogue in that case.\footnote{For a discussion, see footnote (ref) under Proposition (ref).} Finally, the inferential guarantees of the robust estimator and CI do not depend on how the underlying parameters are defined. For example, in event studies one can define the parameters as cohort-level difference-in-differences (DiD) estimands rather than cohort-level average causal effects sun2021estimating. This distinction affects the interpretation of the weighted estimands, but not the validity of the inference procedures.

A large literature highlights issues that arise when interpreting weighted estimands. In the context of my framework, these issues represent different sources of ambiguity and disagreement over the choice of weights. For example, regression-based estimands often recover weighted averages of treatment effects, but allow for negative weights de2020two, sloczynski2020should, goodman2021difference, mogstad2021causal, sun2021estimating, blandhol2022tsls, bhuller20242sls, goldsmith2024contamination. Negative weights can flip the signs of treatment effects, complicating the interpretation of an estimand. In such contexts, there may be (i) ambiguity over how much negative weighting matters and (ii) disagreement over what alternative weighting schemes to consider.\footnote{abadie2025harvesting and chiu2026causal give different takeaways for issue (i) in the event study context.} Regarding the latter, a separate issue is that a positively weighted estimand can still lack empirical or policy relevance yitzhaki1996using, heckman2005structural, crump2006moving, angrist2010extrapolate, aronow2016does, li2018balancing, sloczynski2022interpreting, mogstad2024instrumental, poirier2024quantifying. In this case, disagreement may arise from policymakers interested in the effects of treatment on their own target populations, while ambiguity may arise from researchers who must navigate such disagreement when summarizing results.

Common approaches to these issues are to report the extent of negative weighting in one's estimator or to examine the stability of results under alternative weighting schemes.\footnote{See roth2023s for a review of such approaches in the event study context.} Intuitively, such approaches convey information about how far one's weights deviate from a benchmark, the degree of underlying heterogeneity, or some combination of both. My procedures build on this same intuition, but in a GLS-based geometry that yields a notion of heterogeneity amenable to inference with confidence bounds, drawing on statistical results on optimal quantile-unbiased estimation pfanzagl1994parametric. This allows one to directly account for weighting issues by constructing robust CIs and estimators with explicit inferential guarantees, thereby facilitating robust inference for weighted estimands.

The remainder of this paper proceeds as follows. Section (ref) develops the model setting and notation. Section (ref) defines the heterogeneity and distance measures and establishes the bound on differences between weighted estimands. Section (ref) develops the robust estimator and CI and establishes their properties. These results are developed under the assumption that group-level estimates are normally distributed with a known covariance matrix. Section (ref) shows that, with asymptotically normal estimates and a consistent covariance matrix estimator, my proposed procedures are uniformly asymptotically valid over a broad class of data generating processes. Section (ref) discusses the practical implementation of these procedures. Section (ref) presents the empirical applications. Section (ref) concludes. The supplemental appendices contain proofs and additional results.

Model Setting

Consider parameters $\theta_{k} \in \mathbb{R}$ indexed by groups $k \in \{1, \ldots, K\}$, where $K \geq 2$. Let $\theta = (\theta_{1}, \ldots, \theta_{K})'$ denote the vector of parameters. Researchers often conduct inference on estimands that can be expressed as weighted averages of the group-level parameters:

align*[align* omitted — 175 chars of source]

where $w \in \mathcal{W}$ is a vector of weights, $\mathcal{W}$ is the set of all weights, and $\bm{1}$ is the vector of ones. I refer to $\tau_{w}(\theta)$ as a weighted estimand. I allow the weights $w$ to be negative unless stated otherwise. For the case of nonnegative (i.e., convex) weights, I define the simplex $\mathcal{W}_{+} = \{w \in \mathcal{W}: w \geq 0\}$.

example*[Event Studies] Given units $i$ and time periods $t \in \{0, 1, \ldots, T\}$, where $T \geq 2$, the researcher observes outcomes $Y_{it} \in \mathbb{R}$ and treatment indicators $D_{it} \in \{0,1\}$. A unit $i$'s cohort group $G_{i} = \min\left\{t: D_{it}=1\right\}$ is the time period when first treated---if never treated, then $G_{i} = \infty$. There are no treated units in the base period (i.e., $G_{i} > 0$). Moreover, treatment is “absorbing” in the sense that a treated unit remains treated (i.e., $t \geq G_{i}$ implies $D_{it} = 1$).\footnote{For any treatment that is not absorbing, one can define an indicator for ever having received the treatment, which will be an absorbing treatment by construction; see sun2021estimating for an example.} Let $Y_{it} = Y_{it}(G_{i})$, where $Y_{it}(g)$ denotes the potential outcome for unit $i$ when assigned to cohort $g$. Consider the group-time average treatment effect on the treated (ATT): \begin{align} \operatorname{ATT}_{g,t} = E[Y_{it}(g) - Y_{it}(\infty)|G_{i} = g], \quad t \geq g. \end{align} This parameter gives the period $t$ causal effect of being treated in cohort $g$ versus being never-treated, averaged among the units $i$ in cohort $G_{i} = g$. Under assumptions of parallel trends and no anticipation, $\operatorname{ATT}_{g,t}$ is identified for different group-time pairs callaway2021difference, roth2023s. Let $k$ index such pairs $(g_{k},t_{k})$. For parameters $\theta_{k} = \operatorname{ATT}_{k} = \operatorname{ATT}_{g_{k},t_{k}}$ and weights $w \in \mathcal{W}$, the corresponding weighted estimand is \begin{align*} \tau_{w}(\theta) = \sum_{k=1}^{K} w_{k}\operatorname{ATT}_{k}. \end{align*} Many common event-study estimands can be expressed as weighted estimands, including conventional TWFE regression coefficients and recently proposed alternatives de2020two, callaway2021difference, goodman2021difference, sun2021estimating, gardner2022two, roth2023efficient, borusyak2024revisiting, wing2024stacked. In this setting, the different weights $w \in \mathcal{W}$ represent different ways of summarizing treatment effect heterogeneity across cohort groups and time periods; see callaway2021difference for examples.
example*[Multisite Experiments] For units $i$, there are outcomes $Y_{i} \in \mathbb{R}$, treatment indicators $D_{i} \in \{0,1\}$, and covariates $X_{i} \in \mathbb{R}^{M}$. Let $(Y_{i}(1), Y_{i}(0))$ denote potential outcomes under treatment and control, so that $Y_{i} = D_{i}Y_{i}(1) + (1-D_{i})Y_{i}(0)$. The units come from different site populations $k$, represented by distributions $P_{k}$ over the unit-level random variables. Letting $E_{k} = E_{P_{k}}$ denote expectations under $P_{k}$, the site-level average treatment effects (ATEs) are \begin{align*} \operatorname{ATE}_{k} = E_{k}[Y_{i}(1) - Y_{i}(0)], \quad k = 1, \ldots, K. \end{align*} Under assumptions of random treatment assignment within each site, $\operatorname{ATE}_{k}$ is identified for each site $k$ hotz2005predicting. For parameters $\theta_{k} = \operatorname{ATE}_{k}$ and weights $w \in \mathcal{W}$, the corresponding weighted estimand is \begin{align*} \tau_{w}(\theta) = \sum_{k=1}^{K} w_{k}\operatorname{ATE}_{k}. \end{align*} For instance, the equal weights (EW) vector $w_{\mathrm{EW}} = \bm{1}/K$ represents an empirical distribution over the set of sites---this has been considered in, for example, allcott2015site. More generally, each simplex vector $w \in \mathcal{W}_{+}$ represents a distribution over the set of sites.

Conventional Inference for Weighted Estimands

The researcher observes a vector of estimates $\hat{\theta}$ for the parameters $\theta$. I model the relationship between $\hat{\theta}$ and $\theta$ as follows.

assumption$\hat{\theta} \sim N(\theta, \Sigma)$, where $\Sigma$ is a known positive definite covariance matrix.

I denote probabilities, expectations, and variances under $\hat{\theta} \sim N(\theta, \Sigma)$ as $\mathbb{P}_{\theta}\left\{\cdot\right\}$, $\mathbb{E}_{\theta}\left[\cdot\right]$, and $\textnormal{Var}_{\theta}\left(\cdot\right)$, respectively. Moreover, $\Phi(z)$ denotes the cumulative distribution function (CDF) of the standard normal distribution $N(0,1)$ and $z_{\alpha}$ denotes its $\alpha$-quantile for $\alpha \in (0,1)$.

Assumption (ref) is motivated by standard large-sample asymptotic results. For example, the central limit theorem yields asymptotically normal estimates, which motivates the normality condition. Likewise, the law of large numbers yields consistent covariance matrix estimation, which motivates the condition that $\Sigma$ is known. I therefore develop my procedures under Assumption (ref). In Section (ref), I relax Assumption (ref) and establish the asymptotic validity of my procedures.

For each vector of weights $w \in \mathcal{W}$, Assumption (ref) implies $w'\hat{\theta} \sim N(w'\theta, \sigma_{w}^{2})$, where $\sigma_{w}^{2} = w'\Sigma w$. For a given target estimand $\tau_{w}(\theta) = w'\theta$, the conventional point estimator is defined as $\hat{\tau}_{w} = w'\hat{\theta}$ and the conventional CI at significance level $\alpha$ is defined as

align[align omitted — 398 chars of source]

The conventional estimator $\hat{\tau}_{w}$ is unbiased for $\tau_{w}(\theta)$:

align*[align* omitted — 106 chars of source]

The conventional $CI_{w}$ has exact coverage for $\tau_{w}(\theta)$:

align*[align* omitted — 113 chars of source]

In this sense, the conventional statistics $(\hat{\tau}_{w}, CI_{w})$ provide valid inference for $\tau_{w}(\theta)$.

example*[Event Studies, continued] Let $\hat{\theta}_{k} = \widehat{\operatorname{ATT}}_{k}$ denote estimators of $\operatorname{ATT}_{k}$, such as DiD-based estimators. Under standard conditions, such estimators are asymptotically normal with consistently estimable covariance matrices as the number of units goes to infinity callaway2021difference. Often one must also estimate the weights $w$. For example, the sun2021estimating weights depend on the population distribution of treatment $D_{it}$, so that in practice one must use the sample distribution of $D_{it}$ to estimate $w$. For ease of exposition, I develop my procedures assuming that $w$ is known and later establish asymptotic validity under estimated weights in Section (ref).
example*[Multisite Experiments, continued] Let $\hat{\theta}_{k} = \widehat{\operatorname{ATE}}_{k}$ denote estimators of $\operatorname{ATE}_{k}$, such as those based on sample differences in means between treatment and control groups. Under site-level CLTs, such estimators are asymptotically normal with consistently estimable covariance matrices as the number of units goes to infinity. Note that even in cases where the unit-level variables $(Y_{i}, D_{i}, X_{i})$ are unobserved or confidential, one may still have access to site-level ATE estimates $\hat{\theta} = (\widehat{\operatorname{ATE}}_{1}, \ldots, \widehat{\operatorname{ATE}}_{K})'$ and covariate means $\bm{X} = (E_{1}[X_{i}], \ldots, E_{K}[X_{i}])'$, as in the case of the metadata in allcott2015site. The application of my framework to multisite experiments requires only the site-level variables. I accommodate the case of estimated covariate means $\widehat{E}_{k}[X_{i}]$ in my asymptotic results.

Classes of Alternative Weights

The researcher reports $(\hat{\tau}_{w}, \sigma_{w})$ for some baseline choice of weights $w \in \mathcal{W}$. This report yields conventional statistics $(\hat{\tau}_{w}, CI_{w})$ that provide valid inference for the baseline estimand $\tau_{w}(\theta)$. However, the researcher may also consider a class $\Lambda \subseteq \mathcal{W}$ of alternative weights $\lambda \in \Lambda$, yielding alternative estimands $\tau_{\lambda}(\theta)$ for which $(\hat{\tau}_{w}, CI_{w})$ may be uninformative.

example*[Event Studies, continued] callaway2021difference advocate choosing $w$ to address well-posed empirical or policy questions.\footnote{This perspective is also prevalent in other empirical settings, such as instrumental variables estimation under heterogeneous treatment effects heckman2007econometric, mogstad2024instrumental.} While this principle is ideal, two practical issues can arise, leading to consideration of a class of alternatives $\Lambda$. \paragraph{Researcher Ambiguity.} First, even with a well-posed question, the researcher may find it difficult to articulate corresponding weights $w$. This costly introspection presumably underlies the various default weighting schemes considered in the event studies literature.\footnote{See callaway2021difference and roth2023s for examples.} A researcher may choose a default specification for $w$ and then assess robustness of results to other defaults. But to the extent that these defaults fail to capture the scope of plausible weights, such robustness exercises may not adequately account for researcher ambiguity over the choice of $w$. To do better, it seems useful to consider a broader class of alternatives for $w$, such as the example $\Lambda$ developed below. \paragraph{Reader Disagreement.} Second, even absent introspection costs, the researcher may have to communicate results to readers who are interested in different questions and hence disagree over the choice of weights. For example, colleagues with different priors about the empirical setting may wish to highlight different types of treatment effect heterogeneity: e.g., across cohorts versus across time callaway2021difference. In such cases, $\Lambda$ represents the range of questions that readers may be interested in, which can differ from the question addressed by $w$.
example*[Multisite Experiments, continued] Given the site-level populations $P_{1}, \ldots, P_{K}$, a baseline vector of simplex weights $w \in \mathcal{W}_{+}$ represents the population $P_{w} = \sum_{k}w_{k}P_{k}$. The set of readers may include policymakers interested in the effect of treatment on their own target populations, leading to disagreement over the choice of weights. For concreteness, consider a policymaker with target population $P_{0}$ and a corresponding target parameter \begin{align*} \operatorname{ATE}_{0} = E_{0}[Y_{i}(1) - Y_{i}(0)]. \end{align*} This policymaker may find the baseline weights $w$ to be palatable when $P_{0}$ resembles $P_{w}$, such as when the mean $\mu_{0} = E_{0}[X_{i}]$ under $P_{0}$ is close to the mean $\mu_{w}(\bm{X}) = w'\bm{X}$ under $P_{w}$.\footnote{This mirrors the logic advanced in aronow2016does for interpreting the representativeness of weighted estimands in regression contexts. In their language, $\mu_{w}(\bm{X})$ is the covariate mean for the effective sample of units spanned by the $w$-weighted sites.} In view of this, $\Lambda \subseteq \mathcal{W}_{+}$ can be specified as a class of populations that a range of policymakers may find to be palatable. Note that if there were a single known policymaker, one could employ canonical approaches for extrapolating from the site-level populations $(P_{1}, \ldots, P_{K})$ to the target population $\operatorname{ATE}_{0}$, such as density reweighting under covariate shift assumptions.\footnote{There are applications of this approach in economics hotz2005predicting, stuart2011use, allcott2015site, dehejia2021local, health cole2010generalizing, hartman2015sample, and machine learning farahani2021brief, zhou2022domain.} However, such avenues are less tractable in the face of multiple and potentially unknown policymakers.

Whether the issue at hand is researcher ambiguity, reader disagreement, or some combination of both, my framework assumes that it can be represented by a class of alternative weights $\Lambda$. To make this modeling assumption a practical one, I restrict attention to classes $\Lambda$ that have the following structure.

assumption$\Lambda \subseteq \mathcal{W}$ is nonempty, compact, and convex.

This structure ensures the existence and uniqueness of solutions to upcoming optimization problems. I now provide examples of classes $\Lambda$ that satisfy Assumption (ref) and map to various empirical contexts.

example*[Bounded Variance] To analyze the robustness of baseline inferences to alternative weights, one may wish to account for differences in how well the corresponding estimands can be estimated. Formally, one can consider weights $\lambda$ for which the standard deviation of $\hat{\tau}_{\lambda}$ is at most $r$ times that of the baseline $\hat{\tau}_{w}$, leading to the bounded variance class \begin{align} \Lambda_{\sigma}(r) = \left\{\lambda \in \mathcal{W}: \sigma_{\lambda} \leq r\sigma_{w}\right\}, \quad r \geq \frac{\sigma_{\min}}{\sigma_{w}}, \quad \sigma_{\min}^{2} = \min_{w \in \mathcal{W}} \sigma_{w}^{2} = \frac{1}{\bm{1}'\Sigma^{-1}\bm{1}}. \end{align} This class reflects a preference for estimands that can be estimated with some reasonable degree of precision.\footnote{For example, mogstad2024instrumental note that “How interesting a target parameter is also cannot be divorced from the difficulty involved in estimating it.”} For instance, $r=1$ represents the class of estimands that can be estimated at least as precisely as the baseline estimand. When convenient, I leave the dependence of $\Lambda_{\sigma}(r)$ on the standard deviation ratio bound $r$ implicit, denoting $\Lambda_{\sigma} = \Lambda_{\sigma}(r)$.
example*[Truncated Simplex] The simplex $\mathcal{W}_{+} = \{w \in \mathcal{W}: w \geq 0\}$ is the set of nonnegative weights, which yields estimands that do not extrapolate beyond the range of the parameters: $\min_{k}\theta_{k} \leq \tau_{w}(\theta) \leq \max_{k}\theta_{k}$. However, the simplex allows for some groups to receive zero weight, which can be undesirable in practice. To ensure that groups are represented, one can impose a floor on the simplex. In particular, given baseline simplex weights $w \in \mathcal{W}_{+}$, the floor $(1-\epsilon)w$ yields the truncated simplex class \begin{align} \Lambda_{+}(\epsilon) = \left\{\lambda \in \mathcal{W}_{+}: \lambda \geq (1-\epsilon)w\right\} = \left\{(1-\epsilon)w + \epsilon \lambda: \lambda \in \mathcal{W}_{+}\right\} \quad \epsilon \in [0,1]. \end{align} This class allows one to flexibly model departures from the baseline simplex weights $w$, where the truncation parameter $\epsilon$ is the fraction of a given alternative $\lambda \in \Lambda_{+}$ that is allowed to deviate from $w$. To ensure that a group $k$ is represented in $\lambda \in \Lambda_{+}(\epsilon)$, one must use weights with $w_{k} > 0$. In the case of equal weighting, $w = w_{\mathrm{EW}} = \bm{1}/K$, one can interpret $\epsilon$ as the maximum possible discrepancy in weights $|\lambda_{k}-\lambda_{k'}|$ between two groups $k$ and $k'$. I give additional interpretations of $\epsilon$ when discussing practical implementations in Section (ref).
example*[Covariate Balance] For $m \in \{1, \ldots, M\}$, consider group-level covariates $\bm{X}_{m} \in \mathbb{R}^{K}$. A baseline simplex weight vector $w \in \mathcal{W}_{+}$ is a probability distribution over the set of groups, and can thus be viewed as a population with covariate means $w'\bm{X}_{m}$. Likewise, alternative simplex weights $\lambda \in \mathcal{W}_{+}$ yield populations with covariate means $\lambda'\bm{X}_{m}$. Letting $|(\lambda - w)'\bm{X}_{m}|$ denote the covariate $m$ balance gap and $\text{sd}(\bm{X}_{m})$ the standard deviation of $\bm{X}_{m,k}$ across groups $k$, the set of populations with balance gaps no larger than $\Bar{c}$ standard deviations for all $m$ is given by the covariate balance class \begin{align} \Lambda_{X}(\Bar{c}) = \left\{\lambda \in \mathcal{W}_{+}: \max_{m \in \{1,\ldots,M\}}\frac{\left|(\lambda - w)'\bm{X}_{m}\right|}{sd(\bm{X}_{m})} \leq \Bar{c}\right\}, \quad \Bar{c} \geq 0, \quad \min_{m \in \{1,\ldots,M\}}sd(\bm{X}_{m}) > 0. \end{align} This class allows one to interrogate a range of populations with different covariate profiles from the baseline. For ease of reference, I abbreviate the balance gap measure in $\Lambda_{X}$ as \begin{align*} c_{\lambda}(\bm{X}) = \max_{m \in \{1, \ldots, M\}}\frac{\left|(\lambda - w)'\bm{X}_{m}\right|}{sd(\bm{X}_{m})}, \quad \bm{X} = (\bm{X}_{1}, \ldots, \bm{X}_{M}). \end{align*} Thus, I can write $\Lambda_{X}(\Bar{c}) = \left\{\lambda \in \mathcal{W}_{+}: c_{\lambda}(\bm{X}) \leq \Bar{c}\right\}$ for balance gap bound $\Bar{c}$.
remark[Intersections] One can also take intersections of the above classes. For example, intersecting the bounded variance class $\Lambda_{\sigma}(r)$ with the unrestricted simplex $\Lambda_{+}(1) = \mathcal{W}_{+}$ yields the bounded variance simplex class \begin{align} \Lambda_{\sigma}^{+}(r) = \Lambda_{\sigma}(r) \cap \Lambda_{+}(1) = \left\{\lambda \in \mathcal{W}_{+}: \sigma_{\lambda} \leq r\sigma_{w}\right\}, \quad r \geq \frac{\min_{w \in \mathcal{W}_{+}} \sigma_{w}}{\sigma_{w}}, \end{align} which considers convex weights that yield estimators with reasonable estimation precision. As a second example, one can intersect the covariate balance class $\Lambda_{X}(\Bar{c})$ with the truncated simplex $\Lambda_{+}(\epsilon)$ to obtain the truncated covariate balance class \begin{align} \Lambda_{X}^{\epsilon}(\Bar{c}) = \Lambda_{X}(\Bar{c}) \cap \Lambda_{+}(\epsilon) = \left\{(1-\epsilon)w + \epsilon \lambda: \lambda \in \Lambda_{X}(\Bar{c}/\epsilon)\right\}, \quad \epsilon > 0, \end{align} which considers $\epsilon$-departures from the baseline $w$ towards populations within a covariate mean radius $\Bar{c}/\epsilon$ from the baseline. I will consider these intersections in the empirical applications, but I focus on the benchmark classes $(\Lambda_{\sigma}, \Lambda_{+}, \Lambda_{X})$ when interpreting my theoretical results.

I conclude this section by formalizing the inference distortions that arise when one uses the baseline statistics $(\hat{\tau}_{w}, CI_{w})$ to learn about alternative estimands $\tau_{\lambda}(\theta)$. In particular, I consider (i) the absolute bias of $\hat{\tau}_{w}$ for $\tau_{\lambda}(\theta)$, given by

align*[align* omitted — 166 chars of source]

and (ii) the noncoverage probability of $CI_{w}$ for $\tau_{\lambda}(\theta)$, given (in the two-sided case) by

align*[align* omitted — 287 chars of source]

The inference distortions are governed by the (absolute) difference in estimands $|\tau_{\lambda}(\theta)- \tau_{w}(\theta)|$: the larger this difference, the higher the bias and noncoverage. In particular, when $\lambda \neq w$, there exist parameter values $\theta$ where $|\tau_{\lambda}(\theta)- \tau_{w}(\theta)|$ is arbitrarily large so that bias is arbitrarily large and noncoverage is arbitrarily close to one. In such cases, $(\hat{\tau}_{w}, CI_{w})$ is completely uninformative for $\tau_{\lambda}(\theta)$. Thus, to obtain useful inferences for $\tau_{\lambda}(\theta)$, one must account for potential differences $|\tau_{\lambda}(\theta)- \tau_{w}(\theta)|$ in the weighted estimands across $\lambda \in \Lambda$ and $\theta \in \mathbb{R}^{K}$.

Bounding the Difference in Estimands

In this section, I bound the difference in estimands in terms of the heterogeneity in parameters and the distance between weights. I then show how to infer these quantities, which will be the basis for constructing the robust estimator and CI in Section (ref).

\paragraph{Heterogeneity in Parameters.} Consider the generalized least squares (GLS) regression of the parameters $\theta$ on a constant $\bm{1}$ under weighting matrix $\Sigma^{-1}$. I define the heterogeneity in parameters as the square root of the GLS residual sum of squares:

align*[align* omitted — 126 chars of source]

By construction, the heterogeneity in $\theta$ is zero if and only if $\theta_{k}$ is constant across $k$. The above regression is uniquely minimized at the GLS estimand

align[align omitted — 273 chars of source]

which yields the formula

align[align omitted — 219 chars of source]

where $A$ is the annihilator matrix for $\Sigma^{-1/2}\bm{1}$ and $I$ is the identity matrix. Thus, the squared heterogeneity can be represented as a quadratic form $\theta'Q\theta$ of the parameters. This particular quadratic form will facilitate quantile-unbiased inference in Section (ref).

\paragraph{Distance Between Weights.} The standard deviation $\left\|v\right\|_{\Sigma} = \displaystyle\sqrt{v'\Sigma v}$ of a given linear estimator $v'\hat{\theta}$ defines a norm on $v \in \mathbb{R}^{K}$. Given weights $\lambda$ and $w$, I define the distance between weights as the corresponding norm of their difference:

align*[align* omitted — 178 chars of source]

Intuitively, the distance $\left\|\lambda - w\right\|_{\Sigma}$ measures the disagreement between $\lambda$ and $w$ by taking the standard deviation of the corresponding difference in estimators.

propositionFor any $w \in \mathcal{W}$ and $\lambda \in \mathcal{W}$, the difference in estimands is bounded as \begin{align*} |\tau_{\lambda}(\theta) - \tau_{w}(\theta)| \leq H(\theta)\left\|\lambda - w\right\|_{\Sigma}, \quad \forall \theta. \end{align*} This bound is sharp in the sense that, given $\lambda \neq w$ and any $\eta \geq 0$, there exists $\theta$ with $H(\theta) = \eta$ for which the bound holds with equality.
proofSee Appendix (ref).

Proposition (ref) shows that the difference in estimands is bounded by the product of (i) the heterogeneity $H(\theta)$, which is unknown due to the parameters $\theta$ being unobserved, and (ii) the distance $\left\|\lambda - w\right\|_{\Sigma}$, which is unknown due to ambiguity or disagreement over the alternative weights $\lambda \in \Lambda$.\footnote{Proposition (ref) is based on the Cauchy-Schwarz inequality, similar to Scheffé-style arguments for bounding data-dependent linear combinations to obtain uniformly valid inference scheffe1953method, lehmann2024testing. However, Proposition (ref) considers nonrandom linear combinations and obtains bounds on the inference distortions themselves. Analogous bounds hold if one replaces $\Sigma$ with another positive definite matrix when defining the heterogeneity and distance measures. It also has a natural analogue for continuously indexed groups, where one may define the weighted estimands as $\tau_{w}(\theta) = \int w(k)\theta(k)d\nu(k)$ and apply the same Cauchy-Schwarz argument in an $L^{2}(\nu)$ inner product. } In Section (ref), I account for the latter by taking the maximum distance across alternative weights. In Section (ref), I account for the former by constructing an upper confidence bound (UCB) on the heterogeneity.

Maximum Distance

Given baseline $w$ and class of alternatives $\Lambda$, the maximum distance between weights is

align*[align* omitted — 146 chars of source]

Intuitively, the maximum distance measures the worst-case disagreement between the baseline and alternative weights across $\Lambda$. Under the structure on $\Lambda$ from Assumption (ref), the maximum is attained at some $\lambda^{*} \in \Lambda$, which can be viewed as the weights of a reader who disagrees with $w$ the most. Below I analyze the structure of the maximum distance under the example classes from Section (ref) and illustrate how this structure can facilitate communication between researchers and readers---I defer practical recommendations and implementation choices to Section (ref).

example*[Bounded Variance, continued] One can show that \begin{align} \max_{\lambda \in \Lambda_{\sigma}}\left\|\lambda - w\right\|_{\Sigma} = \sqrt{r^{2}\sigma_{w}^{2} - \sigma_{\min}^{2}} + \sqrt{\sigma_{w}^{2} - \sigma_{\min}^{2}}, \quad \sigma_{\min}^{2} = \min_{w \in \mathcal{W}} \sigma_{w}^{2} = \frac{1}{\bm{1}'\Sigma^{-1}\bm{1}}, \end{align} which is a function of $(r, \sigma_{w}, \sigma_{\min})$. If the researcher reports the minimum variance $\sigma_{\min}^{2}$ alongside the baseline variance $\sigma_{w}^{2}$, then a reader can compute the maximum distance for their own choice of $r$. Conveniently, $\sigma_{\min}^{2}$ coincides with the variance of the GLS estimator \begin{align} \hat{\tau}_{\mathrm{GLS}} = w_{\mathrm{GLS}}'\hat{\theta} = \operatorname*{argmin}_{\gamma \in \mathbb{R}}(\hat{\theta} - \bm{1} \gamma)'\Sigma^{-1}(\hat{\theta} - \bm{1} \gamma), \quad \sigma_{\min}^{2} = \sigma_{\mathrm{GLS}}^{2} = w_{\mathrm{GLS}}'\Sigma w_{\mathrm{GLS}}, \end{align} where $w_{\mathrm{GLS}}$ are the weights of the GLS estimand defined in (ref). The above GLS regression is an input to my inference procedures for $H(\theta)$ in Section (ref), so the GLS variance can be obtained at essentially no further cost.
example*[Truncated Simplex, continued] Letting $v_{1}, \ldots, v_{K}$ denote the standard unit vectors, one can show that \begin{align} \max_{\lambda \in \Lambda_{+}}\left\|\lambda - w\right\|_{\Sigma} = \epsilon \max_{\lambda \in \mathcal{W}_{+}}\left\|\lambda - w\right\|_{\Sigma} = \epsilon \max_{j \in \{1,\ldots,K\}}\left\|v_{j} - w\right\|_{\Sigma}, \end{align} which is a function of $(\epsilon, \max_{j}\left\|v_{j} - w\right\|_{\Sigma})$. If the researcher reports $\max_{j}\left\|v_{j} - w\right\|_{\Sigma}$, then a reader can compute the maximum distance for their own choice of $\epsilon$. In the case of an equal weights baseline $w = w_{\mathrm{EW}} = \bm{1}/K$ and a reader with weights $\lambda_{0} \in \mathcal{W}_{+}$ in the simplex, $\epsilon = 1-K\min_{k}\lambda_{0,k}$ is the smallest truncation parameter at which $\Lambda_{+}(\epsilon)$ contains the reader's weights $\lambda_{0}$.
example*[Covariate Balance, continued] Compared to the above examples, there does not appear to be a transparent expression for $\max_{\lambda \in \Lambda_{X}}\left\|\lambda - w\right\|_{\Sigma}$. However, the researcher can plot it for a range of $\Bar{c}$, which allows a reader to examine the maximum distance at their preferred value of $\Bar{c}$ from the reported range. To support this exercise, the researcher can additionally report covariate statistics $(w'\bm{X}_{m}, \text{sd}(\bm{X}_{m}))_{m=1}^{M}$. This allows a given policymaker with covariate means $\mu_{0} \in \mathbb{R}^{M}$ for a target population $P_{0}$ to compute the balance gap measure \begin{align*} c_{0}(\bm{X}) = \max_{m \in \{1, \ldots, M\}}\frac{|\mu_{0,m} - w'\bm{X}_{m}|}{sd(\bm{X}_{m})}, \end{align*} and check the maximum distance at $\Bar{c}_{0} = \min\{c: c_{0}(\bm{X}) \leq c\}$. Intuitively, $\Lambda_{X}(\Bar{c}_{0})$ is the smallest class for which there can exist weights $\lambda \in \Lambda_{X}(\Bar{c}_{0})$ that induce balance gaps consistent with the policymaker's target population. In this sense, $\Lambda_{X}(\Bar{c}_{0})$ is a minimal set of plausible candidates for $P_{0}$ and the maximum distance at $\Bar{c}_{0}$ gives a corresponding measure of ambiguity. Note that the above discussion implicitly assumes that the number of covariates $M$ is small enough for the policymaker to easily compute $c_{0}(\bm{X})$.
remark[Reporting Constraints] $\Lambda$ represents the broad range of beliefs and objectives that researchers and readers may have, making it impractical for researchers to report $(\hat{\tau}_{\lambda}, \sigma_{\lambda})$ for all $\lambda \in \Lambda$. Moreover, if there are confidentiality restrictions or communication costs, reporting the entire data $(\hat{\theta}, \Sigma)$ may not be a broadly applicable solution.\footnote{allcott2015site is one example where confidentiality restrictions preclude the reporting of $(\hat{\theta}, \Sigma)$, but not of $(\hat{\tau}_{w}, \sigma_{w})$. Even absent such restrictions, researchers still often focus on reporting statistics of the form $(\hat{\tau}_{w}, \sigma_{w})$, as discussed in athey2023thirdnumber. This reporting convention may be due to communication costs that induce researchers to report low-dimensional statistics for their readers.} With these issues in mind, my framework develops methods that facilitate inferences across $\lambda \in \Lambda$ based simply on (i) the baseline report $(\hat{\tau}_{w}, \sigma_{w})$ and (ii) a supplementary report of low-dimensional statistics, where the latter loosely refers to statistics whose dimension (e.g., the number of columns and rows occupied in a table) does not grow with the number of groups $K$, thus precluding $(\hat{\theta}, \Sigma)$.

Heterogeneity UCB

I now show how to infer the unknown heterogeneity in parameters $H(\theta)$. In particular, I derive an upper confidence bound (UCB) for $H(\theta)$: given a significance level $\beta \in (0,1)$, I construct an estimator $\hat{\eta}_{1-\beta}$ such that

align[align omitted — 165 chars of source]

In words, $\hat{\eta}_{1-\beta}$ upper bounds the unknown heterogeneity $H(\theta)$ with probability at least $1-\beta$, for any value of $\theta$. I refer to $\hat{\eta}_{1-\beta}$ as a heterogeneity UCB at confidence level $1-\beta$.

\paragraph{Derivation.} To derive $\hat{\eta}_{1-\beta}$, I use the representation of $H(\theta) = \displaystyle\sqrt{\theta'Q\theta}$ as a quadratic form in $\theta$. In particular, since $\hat{\theta} \sim N(\theta, \Sigma)$, the definition of $Q$ from (ref) implies that the heterogeneity in estimates $H(\hat{\theta}) = \sqrt{\hat{\theta}'Q\hat{\theta}}$ satisfies

align*[align* omitted — 157 chars of source]

where $\chi_{K-1}^{2}(\eta)$ denotes the noncentral chi-squared distribution with $K-1$ degrees of freedom and noncentrality parameter $\eta^{2}$. Let $F_{\chi^{2}}(x; \eta)$ denote its CDF and consider the corresponding pivot function $\eta \mapsto F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; \eta)$. The probability integral transform yields

align*[align* omitted — 104 chars of source]

In particular, the probability of observing $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; H(\theta)) \geq \beta$ is equal to $1-\beta$. For each $x > 0$, the function $\eta \mapsto F_{\chi^{2}}(x;\eta)$ is strictly decreasing sun2010monotonicity. Therefore, when $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; 0) > \beta$, I define $\hat{\eta}_{1-\beta}$ as the unique solution to $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; \hat{\eta}_{1-\beta}) = \beta$. Otherwise, when $F_{\chi^{2}}(\hat{\theta}'Q\hat{\theta}; 0) \leq \beta$, I define $\hat{\eta}_{1-\beta} = 0$.\footnote{This construction is consistent with the general approach in pfanzagl1994parametric, which constructs confidence bounds for one-parameter distributions satisfying appropriate monotonicity conditions.} In summary,

align*[align* omitted — 267 chars of source]

This construction depends only on $\hat{\theta}'Q\hat{\theta}$, which is conveniently obtained in (ref) as the residual sum of squares from the GLS regression of the estimates $\hat{\theta}$ on a constant $\bm{1}$ under weighting matrix $\Sigma^{-1}$. The inversion for $\hat{\eta}_{1-\beta}$ is computationally efficient, since it amounts to finding the root of a strictly monotone function.

The next result shows that the above $\hat{\eta}_{1-\beta}$ is a valid UCB for $H(\theta)$. Moreover, it is quantile-unbiased under heterogeneity in the sense that its (weak) overestimation probability for $H(\theta)$ is equal to $1-\beta$ when $H(\theta) > 0$. Note that when $H(\theta) = 0$, the overestimation probability is equal to one, as would be the case for any nonnegative estimator of $H(\theta)$.

propositionThe above $\hat{\eta}_{1-\beta}$ satisfies (ref). Moreover, under heterogeneous $\theta$, the confidence level $1-\beta$ is attained: \begin{align} \mathbb{P}_{\theta}\left\{H(\theta) \leq \hat{\eta}_{1-\beta}\right\} = 1 - \beta, \quad \forall \theta: H(\theta) > 0. \end{align} Under homogeneous $\theta$, the coverage probability is equal to one.
proofSee Appendix (ref).

The quantile-unbiasedness in Proposition (ref) facilitates a variety of inference procedures for $H(\theta)$. First, $[0, \hat{\eta}_{1-\beta}]$ is a one-sided upper confidence interval for $H(\theta)$, with exact coverage rate $1-\beta$ under heterogeneous $\theta$. Next, the one-sided lower interval $[\hat{\eta}_{\beta}, \infty)$ satisfies

align[align omitted — 179 chars of source]

Likewise, one can construct a two-sided interval $[\hat{\eta}_{\beta/2}, \hat{\eta}_{1-\beta/2}]$ that satisfies

align[align omitted — 212 chars of source]

Finally, one can construct a point estimator $\hat{\eta}_{1/2}$ that is median-unbiased in the sense that

align[align omitted — 246 chars of source]

That is, the median of $\hat{\eta}_{1/2}$ is equal to $H(\theta)$ under heterogeneity.\footnote{For point estimation of heterogeneity, one can also consider mean-unbiased estimation of $\theta'Q\theta$, following a similar construction to, e.g., kline2020leave. In particular, one can show that $\hat{\theta}'Q\hat{\theta} - (K-1)$ is a mean-unbiased estimator of $\theta'Q\theta$. However, the quantile-unbiased approach adopted here is more amenable to the construction of confidence bounds for my objects of interest.} Given my focus on upper bounding differences in weighted estimands, I focus on properties (ref) and (ref). Nevertheless, properties (ref)-(ref) illustrate the versatility of $\hat{\eta}_{1-\beta}$ for inference on parameter heterogeneity. I conclude this section with comparative statics and an optimality result for $\hat{\eta}_{1-\beta}$.

\paragraph{Comparative Statics.} For a given $\eta$, the CDF $F_{\chi^{2}}(x; \eta)$ is increasing in $x$ and decreasing in its degrees of freedom sun2010monotonicity. Thus, $\hat{\eta}_{1-\beta}$ is increasing in $\hat{\theta}'Q\hat{\theta}$ given a fixed $K$, and decreasing in $K$ given a fixed $\hat{\theta}'Q\hat{\theta}$. In this sense, $\hat{\eta}_{1-\beta}$ is increasing in heterogeneity of the estimates relative to the number of groups: $H(\hat{\theta})/\sqrt{K}$. Notice that, under homoskedasticity $\Sigma = \sigma^{2}I$, the expression $\hat{\theta}'Q\hat{\theta}/K$ reduces to the empirical variance of the estimates relative to the common sampling variance $\sigma^{2}$ of the estimates:

align*[align* omitted — 227 chars of source]

Thus, the proposed heterogeneity measure $\hat{\eta}_{1-\beta}$ has intuitive comparative statics in terms of $\hat{\theta}'Q\hat{\theta}/K$. Finally, holding $\hat{\theta}'Q\hat{\theta}$ and $K$ fixed, a larger confidence level $1-\beta$ implies a larger $\hat{\eta}_{1-\beta}$. This highlights a tradeoff when constructing $\hat{\eta}_{1-\beta}$: to upper bound $H(\theta)$ with higher confidence, one must tolerate a larger UCB, corresponding to a lower significance level $\beta$.

\paragraph{Optimality.} Let $\Bar{\Theta}^{\eta} = \{\theta: H(\theta) = \eta\}$ denote the set of parameters for which the heterogeneity is equal to $\eta$, and consider the class of potentially randomized estimators $\Tilde{\eta}_{1-\beta}$ that are quantile-unbiased in the sense of (ref). The following result shows that $\hat{\eta}_{1-\beta}$ is minimax most accurate in the sense that it minimizes the worst-case probability of overestimating $H(\theta)$, for any degree of underlying heterogeneity $\eta$.

propositionFor any quantile-unbiased estimator $\Tilde{\eta}_{1-\beta}$ and any degree of heterogeneity $\eta$, the heterogeneity measure $\hat{\eta}_{1-\beta}$ yields a lower worst-case probability of overestimating $H(\theta)$: \begin{align*} \sup_{\theta \in \Bar{\Theta}^{\eta}}\mathbb{P}_{\theta}\left\{\hat{\eta}_{1-\beta} \geq H(\theta) + \varepsilon\right\} \leq \sup_{\theta \in \Bar{\Theta}^{\eta}}\mathbb{P}_{\theta}\left\{\Tilde{\eta}_{1-\beta} \geq H(\theta) + \varepsilon\right\}, \quad \forall (\eta,\varepsilon) > 0. \end{align*}
proofSee Appendix (ref).

The above notion of optimality is a minimax analogue of the notion of a uniformly most accurate confidence bound lehmann2024testing. In Appendix (ref), I show that $\hat{\eta}_{1-\beta}$ is uniformly most accurate in the class of quantile-unbiased estimators that depend on $\hat{\theta}$ through $\hat{\theta}'Q\hat{\theta}$. In fact, $\hat{\eta}_{1-\beta}$ minimizes expected loss in this class under any quasiconvex loss function that attains its minimum at $H(\theta)$, for any heterogeneous $\theta$. The latter optimality statement is based on results from pfanzagl1994parametric on optimal quantile-unbiased estimation in models with monotone likelihood ratios.

Robust Inference Procedures

Given the maximum distance between weights $\max_{\lambda \in \Lambda}\left\|\lambda - w\right\|_{\Sigma}$ and the heterogeneity UCB $\hat{\eta}_{1-\beta}$, I define the bias UCB

align*[align* omitted — 135 chars of source]

The following result shows that $\widehat{B}_{w}^{\beta}(\Lambda)$ is indeed a valid bias UCB.

propositionFor any $w \in \mathcal{W}$, the above $\widehat{B}_{w}^{\beta}(\Lambda)$ satisfies \begin{align*} \mathbb{P}_{\theta}\left\{\max_{\lambda \in \Lambda}|\tau_{\lambda}(\theta) - \tau_{w}(\theta)| \leq \widehat{B}_{w}^{\beta}(\Lambda)\right\} \geq 1-\beta, \quad \forall \theta. \end{align*} Moreover, when $\Lambda \neq \{w\}$, there exists $\theta$ for which the coverage probability is equal to $1-\beta$.
proofSee Appendix (ref).

In words, $\widehat{B}_{w}^{\beta}(\Lambda)$ upper bounds the maximum bias with probability at least $1-\beta$, for any value of $\theta$. This provides an avenue for addressing the bias and undercoverage of the baseline estimator $\hat{\tau}_{w}$ and $CI_{w}$ for inference on alternative estimands $\tau_{\lambda}(\theta)$ across $\lambda \in \Lambda$. To this end, I develop the robust estimator in Section (ref) and the robust CI in Section (ref).

Robust Estimator

Since $\widehat{B}_{w}^{\beta}(\Lambda)$ is a valid bias UCB under any $w \in \mathcal{W}$, it provides a criterion for optimizing $w$. In the optimization, I allow for consideration sets $\mathcal{W}^{*} \subseteq \mathcal{W}$ that are nonempty, closed, and convex. I distinguish the weights in the optimization problem from the researcher's baseline weights, denoting the former by $\Bar{w} \in \mathcal{W}^{*}$ and the latter by $w \in \mathcal{W}$. I allow for classes $\Lambda$ that depend on the baseline $w$, but not on the optimizer variable $\Bar{w}$.

Given a consideration set $\mathcal{W}^{*}$, I define the robust estimator $\hat{\tau}^{*}$ as the weighted estimator induced by the minimax-bias weights $w^{*} \in \mathcal{W}^{*}$, given by

align*[align* omitted — 198 chars of source]

Let $\widehat{B}_{\min}^{\beta}(\Lambda) = \hat{\eta}_{1-\beta}\max_{\lambda \in \Lambda}\left\|\lambda - w^{*}\right\|_{\Sigma}$ denote the corresponding minimax-bias UCB.

propositionThe minimax-bias weights $w^{*}$ exist uniquely and satisfy \begin{align*} \widehat{B}_{\min}^{\beta}(\Lambda) \leq \widehat{B}_{\Bar{w}}^{\beta}(\Lambda), \quad \forall \Bar{w} \in \mathcal{W}^{*}. \end{align*} Moreover, for any parameter space $\Theta^{\eta} = \{\theta: H(\theta) \leq \eta\}$ with a bound $\eta > 0$ on heterogeneity, the minimax-bias weights $w^{*}$ solve \begin{align*} w^{*} = \arg \adjustlimits \min_{\Bar{w} \in \mathcal{W}^{*}} \max_{\lambda \in \Lambda} \left.\max_{\theta \in \Theta^{\eta}}\right. \left|\mathbb{E}_{\theta}\left[\hat{\tau}_{\Bar{w}}\right] - \tau_{\lambda}(\theta)\right|, \quad \forall \eta > 0. \end{align*} In this sense, the robust estimator $\hat{\tau}^{*} = \hat{\tau}_{w^{*}}$ is optimal for minimizing worst-case bias under any bound on heterogeneity.
proofSee Appendix (ref).

Proposition (ref) establishes optimality of the minimax-bias weights $w^{*}$ from two perspectives: first, at the observed data, $w^{*}$ minimizes the bias $\widehat{B}_{\Bar{w}}^{\beta}(\Lambda)$ that is inferred ex post; and second, across hypothetical data realizations, $w^{*}$ minimizes the bias that is possible ex ante. Conveniently, the maximum bias from either perspective is proportional to the maximum distance: $w^{*}$ does not depend on the estimated heterogeneity $\hat{\eta}_{1-\beta}$ or on the heterogeneity bound $\eta$. In the above setup, the optimality of $w^{*}$ among weights $\Bar{w} \in \mathcal{W}^{*}$ is equivalent to optimality of the robust estimator $\hat{\tau}^{*}$ among weighted estimators $\hat{\tau}_{\Bar{w}}$.

By default, I take the consideration set to be the class of alternative weights: $\mathcal{W}^{*} = \Lambda$. In this case, $w^{*}$ can be interpreted as the weights that minimize worst-case disagreement across the class of alternative weights $\Bar{w} \in \Lambda$. Moreover, because the minimax-bias weights depend only on the class of alternative weights, they provide a natural default for researchers facing ambiguity over their initial choice of baseline weights. I now analyze the structure and interpretation of the minimax-bias weights $w^{*}$ for the example classes of $\Lambda$ from Section (ref).

example*[Bounded Variance, continued] Let $w^{*}(r)$ denote the minimax-bias weights under the bounded variance class $\Lambda_{\sigma}(r)$. Following equation (ref), one can show that \begin{align} \hat{\tau}^{*}(r) = \hat{\tau}_{\mathrm{GLS}}, \quad w^{*}(r) = w_{\mathrm{GLS}} = \frac{\Sigma^{-1}\bm{1}}{\bm{1}'\Sigma^{-1}\bm{1}}, \quad \min_{\Bar{w} \in \Lambda_{\sigma}}\max_{\lambda \in \Lambda_{\sigma}}\left\|\lambda - \Bar{w}\right\|_{\Sigma} = \sqrt{r^{2}\sigma_{w}^{2} - \sigma_{\mathrm{GLS}}^{2}}. \end{align} The corresponding standard deviation is $\sigma^{*}(r) = \sigma_{\mathrm{GLS}}$. In summary, the minimax-bias weights coincide with the GLS weights so that the robust estimator and corresponding variance reduce to the GLS estimator and variance, for any choice of standard deviation ratio bound $r$.\footnote{GLS weights have appeared in discussions about how to aggregate CATEs---li2018balancing call them overlap weights while goldsmith2024contamination call them easiest-to-estimate weights. Established properties include variance-efficiency crump2006moving, li2018balancing, goldsmith2024contamination and interpretation as policy effect weights from a marginal increase in the log odds of treatment kennedy2019nonparametric, zhou2022marginal. I establish a complementary property: GLS weights minimize worst-case bias when the only consensus over weighting schemes is that they should yield estimators with bounded variance.} Since the GLS weights minimize variance, $\sigma_{\mathrm{GLS}}^{2} = \sigma_{\min}^{2}$, it follows that the robust estimator is optimal for both bias and variance under the bounded variance class.\footnote{This double-optimality of $\hat{\tau}_{\mathrm{GLS}}$ parallels recent results in adusumilli2026you and sarfati2026integrating, which show that variance-efficient estimators are also bias-optimal under unrestricted---but bounded---forms of model misspecification. In the weighted estimand context, bounding heterogeneity in the GLS metric and the alternative weights in the standard deviation norm can be viewed as unrestricted bounds on the misspecification of a target estimand $\tau_{\lambda}(\theta)$. If one imposes further restrictions, such as intersecting $\Lambda_{\sigma}$ with the simplex, then the robust estimator generally differs from GLS.} From Proposition (ref), it follows that $\hat{\tau}^{*} = \hat{\tau}_{\mathrm{GLS}}$ is minimax-optimal for estimating $\tau_{\lambda}(\theta)$ over $(\lambda, \theta) \in \Lambda \times \Theta^{\eta}$ under squared error in the class of weighted estimators.\footnote{Recent work studies how to optimally aggregate treatment effects from a decision-theoretic perspective armstrong2021finite, de2021trading, kwon2025estimating, lau2026aggregating. In contrast to this work, my framework allows for ambiguity and disagreement over the target estimand.} When $\Sigma$ is diagonal, the GLS weights are guaranteed to be strictly positive, weighting each group $k$ by the precision $1/\sigma_{k}^{2}$ of its estimate relative to the other groups: \begin{align*} w_{\mathrm{GLS}, k} = \frac{1/\sigma_{k}^{2}}{\sum_{k=1}^{K}1/\sigma_{k}^{2}} > 0, \quad \Sigma = \begin{pmatrix} \sigma_{1}^{2} & \cdots & 0 \\ \vdots & \ddots & \vdots \\ 0 & \cdots & \sigma_{K}^{2} \end{pmatrix}. \end{align*} This occurs, for example, when $\hat{\theta}$ is a vector of estimates from statistically independent sites $k$, or when $\hat{\theta}$ is a vector of CATE estimates computed on independent observations across covariate cells $k$. But when $\Sigma$ is non-diagonal, $w_{\mathrm{GLS}}$ may place negative weight on some groups in order to minimize variance. To prevent this negative weighting, one can restrict the bounded variance class to the simplex as in (ref). In this case, generally $w^{*} \neq w_{\mathrm{GLS}}$.
example*[Truncated Simplex, continued] Let $w^{*}(1) = \arg\min_{\Bar{w} \in \mathcal{W}_{+}}\max_{\lambda \in \mathcal{W}_{+}}\left\|\lambda - \Bar{w}\right\|_{\Sigma}$ denote the minimax-bias weights under $\Lambda_{+}(1) = \mathcal{W}_{+}$ and $w^{*}(0) = w$ the minimax-bias weights under $\Lambda_{+}(0) = \{w\}$. Following equation (ref), one can show that \begin{align} w^{*}(\epsilon) = (1-\epsilon)w^{*}(0) + \epsilon w^{*}(1), \quad \min_{\Bar{w} \in \Lambda_{+}}\max_{\lambda \in \Lambda_{+}}\left\|\lambda - \Bar{w}\right\|_{\Sigma} = \epsilon\min_{\Bar{w} \in \mathcal{W}_{+}}\max_{\lambda \in \mathcal{W}_{+}}\left\|\lambda - \Bar{w}\right\|_{\Sigma}. \end{align} Thus, the minimax-bias weights $w^{*}(\epsilon)$ are a fraction $1-\epsilon$ consistent with the baseline weights and a fraction $\epsilon$ consistent with the weights that minimize disagreement over the unrestricted simplex. The robust estimator is therefore $\hat{\tau}^{*}(\epsilon) = (1-\epsilon)\hat{\tau}^{*}(0) + \epsilon\hat{\tau}^{*}(1)$, which is between the baseline estimator $\hat{\tau}^{*}(0) = \hat{\tau}_{w}$ and the simplex-robust estimator $\hat{\tau}^{*}(1)$.
example*[Covariate Balance, continued] There does not appear to be a transparent formula for $w^{*}(\Bar{c})$ under $\Lambda_{X}(\Bar{c})$. However, in the case of $\Bar{c} = 0$, one can interpret $\Lambda_{X}(0)$ in (ref) as the set of solutions to a synthetic control problem abadie2003economic, abadie2010synthetic. In particular, if the baseline $w$ represents synthetic control weights derived under the max-norm and standardized covariates, the corresponding predictor means $w'\bm{X}$ produce the same synthetic control fit as each alternative $\lambda \in \Lambda_{X}(0)$, meaning that $\Lambda_{X}(0)$ is a set of synthetic control weights.\footnote{See liu2025synthetic for a recent identification framework predicated on such ambiguity sets.} The minimax-bias weights $w^{*}(0)$ under $\Lambda_{X}(0)$ can then be viewed as a point of centrality among the set of synthetic control weights.

Robust CI

I define the robust CI centered at $w \in \mathcal{W}$ as

align*[align* omitted — 515 chars of source]

where the critical value function $\text{cv}_{1-\alpha}\left(b\right)$ gives the $(1-\alpha)$-quantile of the folded normal distribution $|N(b,1)|$. Thus, the robust $CI_{w}^{*}$ parallels the baseline $CI_{w}$, but uses critical values that adjust for the inferred bias $\widehat{B}_{w}^{\beta}(\Lambda)$. This adjustment widens the baseline endpoints: $CI_{w} \subseteq CI_{w}^{*}$, with equality if and only if $\widehat{B}_{w}^{\beta}(\Lambda) = 0$. Provided that $\Lambda \neq \{w\}$, this equality holds if and only if the heterogeneity UCB is insignificant at level $\beta$ in the sense that $\hat{\eta}_{1-\beta} = 0$.

proposition$CI_{w}^{*}$ provides uniformly valid coverage at confidence level $1-(\alpha+\beta)$: \begin{align} \mathbb{P}_{\theta}\left\{\tau_{\lambda}(\theta) \in CI_{w}^{*}\right\} \geq 1-(\alpha+\beta), \quad \forall \lambda \in \Lambda, \quad \forall \theta. \end{align}
proofSee Appendix (ref).

Proposition (ref) shows that $CI_{w}^{*}$ covers $\tau_{\lambda}(\theta)$ at confidence level $1-(\alpha+\beta)$, uniformly over $\lambda \in \Lambda$.\footnote{$CI_{w} \subseteq CI_{w}^{*}$ implies $\mathbb{P}_{\theta}\left\{\tau_{w}(\theta) \in CI_{w}^{*}\right\} \geq \mathbb{P}_{\theta}\left\{\tau_{w}(\theta) \in CI_{w}\right\} = 1-\alpha$ for all $\theta$. Thus, the robust $CI_{w}^{*}$ still covers the baseline estimand $\tau_{w}(\theta)$ at conventional confidence levels. The additional error level $\beta$ when considering alternative estimands $\tau_{\lambda}(\theta)$ is a Bonferroni correction to account for the first-step inference on heterogeneity; such Bonferroni corrections have been used for two-step construction of test statistics and CIs in other contexts, such as romano2014practical and mccloskey2017bonferroni.} This uniform coverage differs from the conventional coverage in (ref). The latter provides inference guarantees for one given $w$, while the former extends such guarantees to every $\lambda \in \Lambda$. Relative to conventional coverage, the price of uniform coverage is (i) a higher error level from using $\hat{\eta}_{1-\beta}$ to infer heterogeneity and (ii) longer intervals to obtain coverage that is robust to alternative weights $\lambda \in \Lambda$.

\paragraph{Interpretation.} Intuitively, if $\lambda \in \Lambda$ indexes the preferred weights of different readers, then the robust $CI_{w}^{*}$ provides valid coverage for each reader's estimand $\tau_{\lambda}(\theta)$. For example, suppose the baseline $CI_{w}$ excludes zero, leading the researcher to reject the $w$-null of no average effect $H_{0,w}: \tau_{w}(\theta) = 0$ at significance level $\alpha$. If zero remains excluded from the robust $CI_{w}^{*}$, then a reader with alternative weights $\lambda \in \Lambda$ can reject the $\lambda$-null of no average effect $H_{0,\lambda}: \tau_{\lambda}(\theta) = 0$ at significance level $\alpha + \beta$. In this sense, $CI_{w}^{*}$ facilitates robust inference.

\paragraph{Additional Coverage Properties.} In Appendix (ref), I show that $CI_{w}^{*}$ also provides a form of simultaneous coverage---but at a lower confidence level in the two-sided case. I compare this to the uniform coverage in (ref). In comparing the two coverage notions, I draw connections to the literature on inference for partially identified parameters molinari2020microeconometrics. In Appendix (ref), I derive an upper bound on the coverage rate of $CI_{w}^{*}$. The bound is strictly less than one for any $\lambda \in \Lambda$ under any degree of heterogeneity. This implies that $CI_{w}^{*}$ has nontrivial coverage for $\tau_{\lambda}(\theta)$ across all values of $\lambda \in \Lambda$ and $\theta$.

\paragraph{Choice of Centering Weights.} The robust $CI_{w}^{*}$ can be centered at any $w \in \mathcal{W}$ while maintaining uniform coverage over $\lambda \in \Lambda$. By default, I center the robust CIs at the minimax-bias weights $w^{*}$ and denote $CI^{*} = CI_{w^{*}}^{*}$. In particular, given the minimax-bias UCB $\widehat{B}_{\min}^{\beta}(\Lambda)$, robust estimator $\hat{\tau}^{*}$, and standard deviation $\sigma^{*} = \sqrt{(w^{*})'\Sigma w^{*}}$, the robust CIs become

align*[align* omitted — 520 chars of source]

This choice of center yields the smallest bias UCB $\widehat{B}_{\min}^{\beta}(\Lambda)$ across the consideration set weights. In this sense, $w^{*}$ prioritizes the component of robust CI length that comes from the ambiguity or disagreement over $\Lambda$. However, the resulting standard deviation $\sigma^{*}$ may be larger than the baseline $\sigma_{w}$, so the overall effect on length is ambiguous ex ante.\footnote{Nevertheless, when the maximum bias is fixed and positive, the bias UCB is asymptotically non-negligible while the variance shrinks to zero with the sample size, which supports choosing $w^{*}$ as the default center.} Note that when $\hat{\eta}_{1-\beta}$ is small enough, it is even possible for $CI^{*}$ to be shorter than $CI_{w}$, and hence shorter than $CI_{w}^{*}$. For example, since the bounded variance class yields $w^{*}=w_{\mathrm{GLS}}$, and GLS minimizes variance, the corresponding $CI^{*}$ must be weakly shorter than $CI_{w}$ in realizations where $\hat{\eta}_{1-\beta} = 0$.

\paragraph{Relationship to Bias-Aware CIs.} My robust CIs have a similar structure to the bias-aware CIs advanced in armstrong2018optimal,armstrong2020simple,armstrong2021finite,armstrong2021sensitivity, but there are important differences in model setup and inferential objective. The bias-aware CIs are designed to provide valid inference for a fixed target estimand $\lambda'\theta$ subject to bounds on the parameter space for $\theta$, such as bounds on heterogeneity---see kwon2025estimating for results of this flavor. By contrast, my robust CIs are designed to provide valid inference for a class of target estimands $\{\lambda'\theta: \lambda \in \Lambda\}$ and impose no bound on heterogeneity, opting instead to infer the heterogeneity using $\hat{\eta}_{1-\beta}$. A downside of bounding heterogeneity is the potential for undercoverage when the bound is incorrect, while a downside of inferring heterogeneity is the additional error level $\beta$ in the uniform coverage statement (ref).

Uniform Asymptotic Validity

The foregoing results are developed under normally distributed estimates $\hat{\theta} \sim N(\theta, \Sigma)$, a known covariance matrix $\Sigma$, known baseline weights $w$, and a known class of alternative weights $\Lambda$. In this section, I establish the uniform asymptotic validity of my procedures under asymptotically normal estimates, consistent covariance matrix estimators, consistent estimators for the baseline weights, and consistent estimators for the class of alternative weights. Sections (ref) and (ref) formalize the asymptotic setup and assumptions. Sections (ref), (ref), and (ref) establish asymptotic results for the class of alternative weights, maximum distance, and robust estimator. Sections (ref) and (ref) establish asymptotic results for the heterogeneity UCB and robust CI. For self-contained practical implementations of my inference procedures, see Section (ref).

Environment

There is a sample of size $n$ drawn from some unknown distribution $P_{n} \in \mathcal{P}_{n}$, where $\mathcal{P}_{n}$ is a class of distributions for samples of size $n$. Based on this data, the researcher constructs a vector of estimates $\hat{\theta}_{n} \in \mathbb{R}^{K}$ and a positive definite covariance matrix estimator $\Tilde{\Sigma}_{n} \in \mathbb{R}^{K \times K}$. Let $\widehat{\Sigma}_{n} = n\Tilde{\Sigma}_{n}$ denote the normalized covariance matrix estimator. Under distribution $P_{n}$, the sample objects $(\hat{\theta}_{n}, \widehat{\Sigma}_{n})$ have population analogues $(\theta(P_{n}), \Sigma(P_{n}))$. The number of groups $K \geq 2$ is fixed while the sample size $n \overset{}{\rightarrow}\infty$ grows. I leave the case of growing $K$ to future work.

I assume that the normalized vector of estimates $\sqrt{n}(\hat{\theta}_{n} - \theta(P_{n}))$ is uniformly asymptotically normal in the sense that it converges in bounded Lipschitz (BL) metric to the normal distribution $N(0, \Sigma(P_{n}))$, uniformly over $P_{n} \in \mathcal{P}_{n}$. Moreover, I assume that the vector of parameters $\theta(P_{n})$ is uniformly bounded (over $P_{n} \in \mathcal{P}_{n}$ and $n \geq 1$).

assumptionUFor the set $BL_{1}(\mathbb{R}^{K})$ of functions $f: \mathbb{R}^{K} \overset{}{\rightarrow}\mathbb{R}$ that are bounded by one, i.e., $|f(x)| \leq 1$ for all $x \in \mathbb{R}^{K}$, and have Lipschitz constant bounded by one, i.e., $|f(x) - f(z)| \leq \left\|x-z\right\|$ for all $x,z \in \mathbb{R}^{K}$, \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \sup_{f \in BL_{1}(\mathbb{R}^{K})} \left|\mathbb{E}_{P_{n}}\left[f(\sqrt{n}(\hat{\theta}_{n} - \theta(P_{n})))\right] - \mathbb{E}_\left[f(Z_{n})\right]\right| = 0, \quad Z_{n} \sim N(0, \Sigma(P_{n})). \end{align*} Moreover, there exists constant $\Bar{C}_{\theta} > 0$ such that $\sup_{P_{n} \in \mathcal{P}_{n}} \left\|\theta(P_{n})\right\| \leq \Bar{C}_{\theta}$ for all $n$.

Uniform convergence in BL metric is a standard way to formalize uniform convergence in distribution. For example, when the components of $\hat{\theta}_{n}$ are regression coefficients or sample averages computed over unit-level observations, Assumption (ref) follows from bounds on the moments of the observations and bounds on the dependence across observations.

Next, I assume that the normalized sample covariance matrix $\widehat{\Sigma}_{n} = n\Tilde{\Sigma}_{n}$ is uniformly $\sqrt{n}$-consistent for the population covariance matrix $\Sigma(P_{n})$, and that the eigenvalues of $\Sigma(P_{n})$ are uniformly bounded above and away from zero.

assumptionUFor each $\varepsilon > 0$, there exists constant $C_{\varepsilon} > 0$ such that \begin{align*} \limsup_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{\sqrt{n}\left\|\widehat{\Sigma}_{n} - \Sigma(P_{n})\right\| > C_{\varepsilon}\right\} \leq \varepsilon, \end{align*} where $\left\|\Sigma\right\|$ denotes the matrix operator norm. Moreover, there exists constant $\Bar{e} > 0$ such that \begin{align*} 1/\Bar{e} \leq \inf_{n}\inf_{P_{n} \in \mathcal{P}_{n}} e_{\min}(\Sigma(P_{n})) \leq \sup_{n}\sup_{P_{n} \in \mathcal{P}_{n}} e_{\max}(\Sigma(P_{n})) \leq \Bar{e}, \end{align*} where $e_{\min}(\Sigma)$ and $e_{\max}(\Sigma)$ denote the minimum and maximum eigenvalues of a matrix $\Sigma$.

Assumption (ref) is a rate condition on the accuracy of covariance matrix estimation. In iid settings, it can be obtained from higher-moment bounds, while in dependent settings it requires moment and dependence conditions strong enough to yield a uniform $\sqrt{n}$ rate. The uniform eigenvalue bounds ensure that $N(0, \Sigma(P_{n}))$ is uniformly tight and nondegenerate. Note that Assumption (ref) implies, for each $\varepsilon > 0$,

align*[align* omitted — 193 chars of source]

That is, $\widehat{\Sigma}_{n}$ is uniformly consistent for $\Sigma(P_{n})$ under Assumption (ref). The stronger $\sqrt{n}$-consistency is imposed to obtain uniformly valid inference on heterogeneity in Section (ref).

remark[Analogy to the Normal Model] The normal model maintains exact normality $\hat{\theta} \sim N(\theta, \Sigma)$ and known $\Sigma$. The asymptotic environment instead has normal approximation $\hat{\theta}_{n} \overset{a}{\sim} N(\theta(P_{n}), \Sigma(P_{n})/n)$ and estimated $\Tilde{\Sigma}_{n} = \widehat{\Sigma}_{n}/n \overset{a}{\approx} \Sigma(P_{n})/n$. Thus, the objects $(\hat{\theta}, \Sigma, \theta)$ in the normal model are analogous to the objects $(\hat{\theta}_{n}, \Tilde{\Sigma}_{n}, \theta(P_{n}))$ in the asymptotic environment.

Baseline and Alternative Weights

The researcher constructs a vector of baseline weights $\hat{w}_{n} \in \mathcal{W}$ and a nonempty, compact, and convex class of alternative weights $\hat{\Lambda}_{n} \subseteq \mathcal{W}$. The population analogues are $(w(P_{n}), \Lambda(P_{n}))$. The baseline and alternative estimators are $\hat{\tau}_{\hat{w}_{n},n}= \hat{w}_{n}'\hat{\theta}_{n}$ and $\hat{\tau}_{\lambda,n} = \lambda'\hat{\theta}_{n}$, where $\lambda \in \hat{\Lambda}_{n}$. The baseline and alternative estimands are $\tau_{w_{n}}(P_{n}) = w(P_{n})'\theta(P_{n})$ and $\tau_{\lambda}(P_{n}) = \lambda'\theta(P_{n})$, where $\lambda \in \Lambda(P_{n})$. If the baseline and alternative weights are known, then I define $\hat{w}_{n} = w(P_{n})$ and $\hat{\Lambda}_{n} = \Lambda(P_{n})$.

I assume that the sample baseline weights $\hat{w}_{n}$ are uniformly consistent for the population baseline weights $w(P_{n})$. Moreover, I assume that $w(P_{n})$ is uniformly bounded---i.e., the baseline weighting scheme cannot place arbitrarily large negative weight on any group.

assumptionUFor each $\varepsilon > 0$, \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{\left\|\hat{w}_{n} - w(P_{n})\right\| > \varepsilon\right\} = 0. \end{align*} Moreover, there exists constant $\Bar{C}_{w} > 0$ such that $\sup_{P_{n} \in \mathcal{P}_{n}} \left\|w(P_{n})\right\| \leq \Bar{C}_{w}$ for all $n$.

For example, under Assumption (ref), it follows that Assumption (ref) holds for GLS weights

align[align omitted — 207 chars of source]

Assumption (ref) also holds for weights considered in event studies, such as those in callaway2021difference under uniform versions of their asymptotic linearity assumptions, and those in sun2021estimating under uniform versions of their moment conditions.

I suppose that the class of alternative weights $\hat{\Lambda}_{n}$ depends on the data through statistics $\hat{S}_{n} \in \mathcal{S}$ taking values in a metric space $(\mathcal{S},d_{\mathcal{S}})$. In particular, $\hat{\Lambda}_{n} = \Lambda(\hat{S}_{n})$ and $\Lambda(P_{n}) = \Lambda(S(P_{n}))$, where $S(P_{n})$ is the population analogue of $\hat{S}_{n}$. I assume the sample statistics $\hat{S}_{n}$ are uniformly consistent for the population statistics $S(P_{n})$. Moreover, $S(P_{n})$ is contained in a compact set $\mathbb{S} \subseteq \mathcal{S}$ and $\hat{S}_{n}$ is contained in $\mathbb{S}$ with probability uniformly approaching one.

assumptionUFor each $\varepsilon > 0$, \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{d_{\mathcal{S}}(\hat{S}_{n}, S(P_{n})) > \varepsilon\right\} = 0. \end{align*} Moreover, $S(P_{n}) \in \mathbb{S}$ for all $P_{n} \in \mathcal{P}_{n}$ and $n$, and \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{\hat{S}_{n} \notin \mathbb{S}\right\} = 0. \end{align*}

For a bounded variance class with $\hat{S}_{n} = (\widehat{\Sigma}_{n}, \hat{w}_{n})$, Assumption (ref) follows from Assumptions (ref) and (ref), where $\mathbb{S}$ can be defined with constants $2\Bar{e}$ and $2\Bar{C}_{w}$; see the example below for details. For a truncated simplex class with $\hat{S}_{n} = \hat{w}_{n} \in \mathcal{W}_{+}$, Assumption (ref) follows from Assumption (ref) and $\mathbb{S} = \mathcal{W}_{+}$. For a covariate balance class with $\hat{S}_{n} = (\hat{w}_{n}, \hat{\bm{X}}_{n}) \in \mathcal{W}_{+} \times \mathbb{R}^{K \times M}$, where $\hat{\bm{X}}_{n}$ is an estimated covariate mean matrix, Assumption (ref) follows from Assumption (ref) and uniform consistency of $\hat{\bm{X}}_{n}$ to a population matrix $\bm{X}(P_{n})$ with $\left\|\bm{X}(P_{n})\right\| \leq \Bar{C}_{X}$ uniformly for some constant $\Bar{C}_{X} > 0$ and $\min_{m}\text{sd}(\bm{X}_{m}(P_{n})) \geq \varsigma$ uniformly for some constant $\varsigma > 0$, where the corresponding component of $\mathbb{S}$ can be defined with constants $2\Bar{C}_{X}$ and $\varsigma/2$; see the example below for details.

Let $\operatorname*{diam}(A) = \sup_{a,b \in A}\left\|a - b\right\|$ denote the diameter of a given set $A \subseteq \mathbb{R}^{K}$. I assume that the class of alternative weights has the following structure.

assumptionUFor a compact and convex set $\mathbb{W} \subseteq \mathcal{W}$ with positive diameter $\operatorname*{diam}(\mathbb{W}) = \max_{\lambda,w \in \mathbb{W}}\left\|\lambda - w\right\| > 0$, there exists a function $g: \mathbb{W} \times \mathbb{S} \overset{}{\rightarrow}\mathbb{R}$ with constants $L_{g}, \delta_{g} > 0$ such that \begin{enumerate}[label=(\roman*)] • for each $S \in \mathbb{S}$, the class of alternative weights can be represented as \begin{align*} \Lambda(S) = \left\{\lambda \in \mathbb{W}: g(\lambda, S) \leq 0\right\}, \end{align*} where the functions $\{\lambda \mapsto g(\lambda,S): S \in \mathbb{S}\}$ are continuous and convex on $\mathbb{W}$; • the functions $\{S \mapsto g(\lambda,S): \lambda \in \mathbb{W}\}$ are uniformly $L_{g}$-Lipschitz on $\mathbb{S}$ in the sense that \begin{align*} \max_{\lambda \in \mathbb{W}}\left|g(\lambda, S_{1}) - g(\lambda, S_{2})\right| \leq L_{g}d_{\mathcal{S}}(S_{1},S_{2}), \quad \forall S_{1},S_{2} \in \mathbb{S}; \end{align*} • for each $S \in \mathbb{S}$, there exists $\lambda^{\circ}(S) \in \mathbb{W}$ satisfying the Slater condition \begin{align*} g(\lambda^{\circ}(S), S) \leq -\delta_{g}. \end{align*} \end{enumerate}

Condition (ref)(i) requires that, on the set $\mathbb{S}$ containing the population statistics $S(P_{n})$, the classes of interest can be represented as inequality constraints of continuous and convex functions $\lambda \mapsto g(\lambda,S)$ on a compact $\mathbb{W}$, which implies that the population classes $\Lambda(P_{n}) \subseteq \mathbb{W}$ are compact and convex. Condition (ref)(ii) requires that the functions $S \mapsto g(\lambda,S)$ in the representation are uniformly Lipschitz on $\mathbb{S}$, which helps in establishing continuity of $S \mapsto \Lambda(S)$ in the Hausdorff metric defined below in Section (ref). Condition (ref)(iii) requires that the inequality constraints are uniformly strictly feasible, which helps in establishing continuity of $S \mapsto \Lambda(S)$ and implies that the population classes $\Lambda(P_{n})$ are nonempty. Moreover, it implies classes are nondegenerate in the sense that the maximum distance between weights is uniformly bounded away from zero. In particular, the maintained assumptions combined with Lemma (ref) yield

align*[align* omitted — 206 chars of source]

Thus, condition (ref)(iii) rules out asymptotic regimes in which the class of alternatives collapses to a singleton, which helps in formulating general arguments for the asymptotic validity of the robust CIs. I now map the conditions of Assumption (ref) to the example classes of interest.

example*[Bounded Variance, continued] Let $S = (\Sigma,w)$, where $\Sigma$ is a positive definite matrix and $w \in \mathcal{W}$. Given $r \geq 1$, define the bounded variance class \begin{align*} \Lambda_{\sigma}(\Sigma,w) = \left\{\lambda \in \mathcal{W}: g(\lambda,\Sigma,w) \leq 0\right\}, \quad g(\lambda,\Sigma,w) = \sqrt{\lambda'\Sigma\lambda} -r\sqrt{w'\Sigma w}. \end{align*} Take the eigenvalue and norm bounds in Assumptions (ref) and (ref) as given and define \begin{align*} \mathbb{S} = \left\{(\Sigma,w): 1/(2\Bar{e}) \leq e_{\min}(\Sigma) \leq e_{\max}(\Sigma) \leq 2\Bar{e}, w \in \mathcal{W}, \left\|w\right\| \leq 2\Bar{C}_{w}\right\}. \end{align*} Choose a finite constant $\Bar{C}_{\Lambda} > 4r\Bar{e}\Bar{C}_{w}$ and define \begin{align*} \mathbb{W} = \left\{\lambda \in \mathcal{W}: \left\|\lambda\right\| \leq \Bar{C}_{\Lambda}\right\}. \end{align*} If $\lambda \in \Lambda_{\sigma}(\Sigma,w)$ and $(\Sigma,w) \in \mathbb{S}$, then \begin{align*} \left\|\lambda\right\|^{2}/(2\Bar{e}) \leq \lambda'\Sigma\lambda \leq r^{2}w'\Sigma w \leq 2r^{2}\Bar{e}\Bar{C}_{w}^{2}, \end{align*} which implies $\left\|\lambda\right\| \leq 2r\Bar{e}\Bar{C}_{w} \leq \Bar{C}_{\Lambda}$. This means $\Lambda_{\sigma}(\Sigma,w) \subseteq \mathbb{W}$ for each $(\Sigma,w) \in \mathbb{S}$. One can thus consider $g(\lambda,\Sigma,w)$ over $(\lambda, \Sigma, w) \in \mathbb{W} \times \mathbb{S}$, which is continuous and convex in $\lambda$, and Lipschitz in $(\Sigma, w)$. For the Slater condition, consider \begin{align*} w_{\mathrm{GLS}}(\Sigma) = \frac{\Sigma^{-1}\bm{1}}{\bm{1}'\Sigma^{-1}\bm{1}}, \quad \sigma_{\min}(\Sigma) = \min_{w \in \mathcal{W}}\sqrt{w'\Sigma w} = \sqrt{w_{\mathrm{GLS}}(\Sigma)'\Sigma w_{\mathrm{GLS}}(\Sigma)} = \frac{1}{\sqrt{\bm{1}'\Sigma^{-1}\bm{1}}}. \end{align*} If $r > 1$, then $\lambda^{\circ}(\Sigma,w) = w_{\mathrm{GLS}}(\Sigma)$ is a Slater point: \begin{align*} g(\lambda^{\circ}(\Sigma,w),\Sigma,w) \leq (1-r)\sigma_{\min}(\Sigma) \leq -\delta_{g}, \quad \delta_{g} = (r-1)/\sqrt{2\Bar{e}K} > 0. \end{align*} For $r=1$, it suffices to assume that $w$ is bounded away from $\lambda^{\circ}(\Sigma,w) = w_{\mathrm{GLS}}(\Sigma)$. But since this may be difficult to verify in practice, a simple alternative is to use $r = 1+\delta$ for some small fixed $\delta > 0$.\footnote{If one wishes to allow for data-dependent $\delta$ and $r$, it suffices to embed these parameters into $\mathbb{S}$ and then investigate the conditions of Assumption (ref) (and Assumption (ref)) relative to that new class.} The same logic applies to the boundary cases in the examples below.
example*[Truncated Simplex, continued] Let $S = w \in \mathcal{W}_{+}$ and $\mathbb{S} = \mathbb{W} = \mathcal{W}_{+}$. Given $\epsilon \in [0,1]$, define the truncated simplex class \begin{align*} \Lambda_{+}(w) = \left\{\lambda \in \mathbb{W}: g(\lambda,w) \leq 0\right\}, \quad g(\lambda,w) = \max_{k \in \{1,\ldots,K\}}\left((1-\epsilon)w_{k}-\lambda_{k}\right), \quad w \in \mathbb{S}. \end{align*} The function $g(\lambda,w)$ is continuous and convex in $\lambda$ and Lipschitz in $w$. For the Slater condition, consider the mixture $\lambda^{\circ}(w) = (1-\epsilon)w + \epsilon \bm{1}/K$. If $\epsilon > 0$, then \begin{align*} g(\lambda^{\circ}(w),w) = -\delta_{g}, \quad \delta_{g} = \epsilon/K > 0. \end{align*}
example*[Covariate Balance, continued] Let $S = (w,\bm{X})$, where $\bm{X} = (\bm{X}_{1},\ldots,\bm{X}_{M}) \in \mathbb{R}^{K \times M}$ and $w \in \mathcal{W}_{+}$. Given $\Bar{c} \geq 0$, define the balance gap measure \begin{align*} c_{\lambda}(w,\bm{X}) = \max_{m \in \{1,\ldots,M\}} \frac{\left|(\lambda-w)'\bm{X}_{m}\right|}{sd(\bm{X}_{m})}. \end{align*} Take $\mathbb{W} = \mathcal{W}_{+}$ and define the covariate balance class \begin{align*} \Lambda_{X}(w,\bm{X}) = \left\{\lambda \in \mathbb{W}: g(\lambda,w,\bm{X}) \leq 0\right\}, \quad g(\lambda,w,\bm{X}) = c_{\lambda}(w,\bm{X}) - \Bar{c}. \end{align*} Define the set \begin{align*} \mathbb{S} = \left\{(w,\bm{X}): w \in \mathcal{W}_{+}, \left\|\bm{X}\right\| \leq 2\Bar{C}_{X}, \min_{m \in \{1,\ldots,M\}}sd(\bm{X}_{m}) \geq \varsigma/2\right\}. \end{align*} Over $(w,\bm{X}) \in \mathbb{S}$, the standard deviation functions $\bm{X}_{m} \mapsto \text{sd}(\bm{X}_{m})$ are Lipschitz and bounded away from zero and the functions $\{(w,\bm{X}_{m}) \mapsto \left|(\lambda-w)'\bm{X}_{m}\right|: \lambda \in \mathbb{W}\}$ are (uniformly) Lipschitz. Thus, the function $g(\lambda, w, \bm{X})$ is Lipschitz in $(w, \bm{X})$. It is also continuous and convex in $\lambda$. For the Slater condition, consider $\lambda^{\circ}(w,\bm{X}) = w$. If $\Bar{c} > 0$, then \begin{align*} g(\lambda^{\circ}(w, \bm{X}),w,\bm{X}) = -\delta_{g}, \quad \delta_{g} = \Bar{c} > 0. \end{align*}
remark[Mapping to Intersections] Finite intersections of the above classes can be handled by considering the maximum of their corresponding $g$-functions, yielding \begin{align*} \Lambda(S) = \bigcap_{j=1}^{J}\left\{\lambda \in \mathbb{W}: g_{j}(\lambda,S) \leq 0\right\} = \left\{\lambda \in \mathbb{W}: g(\lambda,S) \leq 0\right\}, \quad g(\lambda,S)=\max_{j \in \{1,\ldots,J\}} g_{j}(\lambda,S). \end{align*} If each function $g_{j}(\lambda,S)$ is continuous and convex in $\lambda$ and Lipschitz in $S$, then so is $g(\lambda,S)$. The Slater condition holds when there is a common point $\lambda^{\circ}(S) \in \mathbb{W}$ satisfying \begin{align*} g_{j}(\lambda^{\circ}(S),S) \leq -\delta_{g}, \quad \forall j \in \{1,\ldots,J\}. \end{align*} The above observations can be used to map the conditions of Assumption (ref) to the bounded variance simplex class and the truncated covariate balance class. For example, the bounded variance simplex $\Lambda_{\sigma}^{+}(\Sigma,w) = \{\lambda \in \mathcal{W}_{+}: \sqrt{\lambda'\Sigma \lambda} - r\sqrt{w'\Sigma w} \leq 0\}$ can be mapped to $\mathbb{W} = \mathcal{W}_{+}$, $\mathbb{S} = \{(\Sigma, w): 1/(2\Bar{e}) \leq e_{\min}(\Sigma) \leq e_{\max}(\Sigma) \leq 2\Bar{e}, w \in \mathcal{W}_{+}\}$, $g(\lambda, \Sigma, w) = \sqrt{\lambda'\Sigma \lambda} - r\sqrt{w'\Sigma w}$, and Slater point $\lambda^{\circ}(\Sigma,w) = \operatorname*{argmin}_{\lambda \in \mathcal{W}_{+}}\sqrt{\lambda'\Sigma \lambda}$ for $r > 1$ given by the simplex-constrained variance minimizing weights. For $r=1$, the aforementioned caveats apply.

\paragraph{Notation for Population Objects.} When convenient, I will use the shorthands

align*[align* omitted — 139 chars of source]

and likewise for other objects that depend on $P_{n}$.

Consistency of the Class of Alternatives

To formulate uniform consistency of the estimated class $\hat{\Lambda}_{n}$ for the population class $\Lambda_{n}$, I use a standard notion of distance between sets. Given two nonempty sets $A \subseteq \mathbb{R}^{K}$ and $B \subseteq \mathbb{R}^{K}$, let $\operatorname{dist}(a,B) = \inf_{b \in B}\left\|a-b\right\|$ and $\operatorname{dist}(b,A) = \inf_{a \in A}\left\|b-a\right\|$ denote the distance from $a \in A$ to $B$ and from $b \in B$ to $A$, respectively. The Hausdorff distance between $A$ and $B$ is defined as

align*[align* omitted — 155 chars of source]

Under Assumption (ref), Lemma (ref) shows that $S \mapsto \Lambda(S)$ is Lipschitz continuous in the Hausdorff metric $d_{H}$:

align*[align* omitted — 187 chars of source]

Based on this continuity, the following result shows that $\hat{\Lambda}_{n} = \Lambda(\hat{S}_{n})$ is a uniformly consistent estimator of $\Lambda_{n} = \Lambda(S_{n})$ in the Hausdorff metric.

propositionUUnder Assumptions (ref) and (ref), for each $\varepsilon > 0$, \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{d_{H}(\hat{\Lambda}_{n}, \Lambda_{n}) > \varepsilon\right\} = 0. \end{align*}
proofSee Appendix (ref).

Consistency of the Maximum Distance

The estimated and population maximum distances between weights are

align*[align* omitted — 393 chars of source]

Intuitively, $\left\|\lambda - w_{n}\right\|_{\Sigma_{n}}/\sqrt{n}$ is the standard deviation for $(\lambda - w_{n})'\hat{\theta}_{n}$ under the normal approximation $\hat{\theta}_{n} \overset{a}{\sim} N(\theta_{n}, \Sigma_{n}/n)$, while $\left\|\lambda - \hat{w}_{n}\right\|_{\widehat{\Sigma}_{n}}/\sqrt{n} = \left\|\lambda - \hat{w}_{n}\right\|_{\Tilde{\Sigma}_{n}}$ is the plug-in standard error. The following result shows that the estimated maximum distance is uniformly consistent for the population maximum distance.

propositionUUnder Assumptions (ref)-(ref), for each $\varepsilon > 0$, \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{\left|\max_{\lambda \in \hat{\Lambda}_{n}}\left\|\lambda - \hat{w}_{n}\right\|_{\widehat{\Sigma}_{n}} - \max_{\lambda \in \Lambda_{n}}\left\|\lambda - w_{n}\right\|_{\Sigma_{n}}\right| > \varepsilon\right\} = 0. \end{align*}
proofSee Appendix (ref).

Optimality of the Robust Estimator

The estimated and population minimax-bias weights are

align*[align* omitted — 333 chars of source]

The following result shows that the minimax-bias weights satisfy Assumption (ref).

propositionUUnder Assumptions (ref), (ref), and (ref), for each $\varepsilon > 0$, \begin{align*} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \mathbb{P}_{P_{n}}\left\{\left\|\hat{w}^{*}_{n}\, - w^{*}_{n}\right\| > \varepsilon\right\} = 0. \end{align*} Moreover, there exists constant $\Bar{C}_{w} > 0$ such that $\sup_{P_{n} \in \mathcal{P}_{n}} \left\|w^{*}_{n}\right\| \leq \Bar{C}_{w}$ for all $n$.
proofSee Appendix (ref).

To establish asymptotic optimality of the robust estimator

align*[align* omitted — 109 chars of source]

I first characterize the asymptotic maximum bias of weighted estimators $\hat{\tau}_{\hat{w}_{n},n}= \hat{w}_{n}'\hat{\theta}_{n}$ under the following asymptotic uniform integrability (UI) condition:

align[align omitted — 470 chars of source]

UI condition (ref) controls the tails of the squared estimation errors $\|\hat{\theta}_{n} - \theta_{n}\|^{2}$ and $\|\hat{w}_{n} - w_{n}\|^{2}$, ensuring that the corresponding mean squared errors uniformly converge to zero under Assumptions (ref) and (ref). Similar to the discussions for Assumptions (ref) and (ref), UI condition (ref) follows from moment and dependence bounds on the underlying observations used to construct $(\hat{\theta}_{n}, \hat{w}_{n})$.\footnote{By Markov's inequality, UI condition (ref) follows from bounds strong enough to yield uniformly bounded $L^{2+\delta}(P_{n})$ norms on the estimation errors for some $\delta > 0$; e.g., see van2000asymptotic.}

For the population heterogeneity in parameters $H_{n}(\theta_{n}) = \min_{\gamma \in \mathbb{R}}\left\|\theta_{n} - \bm{1}\gamma\right\|_{\Sigma_{n}^{-1}}$, denote the largest heterogeneity over $\mathcal{P}_{n}$ as $\Bar{\eta}(\mathcal{P}_{n}) = \supPnH_{n}(\theta_{n})$, which is uniformly bounded over $n$ under the eigenvalue bounds on $\Sigma_{n}$ and the norm bound on $\theta_{n}$. For each $w \in \mathcal{W}$, $P_{n} \in \mathcal{P}_{n}$, and $n$, let $P_{n}^{\dagger}(w)$ denote a distribution where $\Sigma(P_{n}^{\dagger}(w)) = \Sigma_{n}$, $S(P_{n}^{\dagger}(w)) = S_{n}$, $w(P_{n}^{\dagger}(w)) = w$, and

align*[align* omitted — 311 chars of source]

which satisfies $H_{n}(\theta(P_{n}^{\dagger}(w))) = \Bar{\eta}(\mathcal{P}_{n})$ whenever $\Lambda_{n} \neq \{w\}$. For example, if $P_{n}^{\dagger}(w_{n}) \in \mathcal{P}_{n}$ for all $P_{n} \in \mathcal{P}_{n}$ and $n$, then the class of distributions will be rich enough for the bias bound below to be sharp, in analogy to the sharpness of the bias bound from Proposition (ref). In this sense, the distributions $P_{n}^{\dagger}(w_{n})$ are least-favorable. In what follows, I denote $\tau_{\lambda,n} = \lambda'\theta_{n}$.

propositionUUnder Assumptions (ref)-(ref) and UI condition (ref), \begin{align*} \limsup_{n \overset{\rightarrow}\infty} \left(\sup_{P_{n} \in \mathcal{P}_{n}} \max_{\lambda \in \Lambda_{n}} \left|\mathbb{E}_{P_{n}}\left[\hat{\tau}_{\hat{w}_{n},n}\right] - \tau_{\lambda,n}\right|\right) \leq \limsup_{n \overset{\rightarrow}\infty} \left(\Bar{\eta}(\mathcal{P}_{n}) \sup_{P_{n} \in \mathcal{P}_{n}} \max_{\lambda \in \Lambda_{n}} \left\|\lambda - w_{n}\right\|_{\Sigma_{n}}\right), \end{align*} where the right-hand side is finite. Moreover, equality holds if $P_{n}^{\dagger}(w_{n}) \in \mathcal{P}_{n}$ for all $P_{n} \in \mathcal{P}_{n}$ and $n$.
proofSee Appendix (ref).

The following result establishes asymptotic optimality of the robust estimator, serving as the asymptotic analogue of Proposition (ref) from the normal model.

propositionULet Assumptions (ref), (ref), (ref), and (ref) be satisfied. Moreover, suppose that $\hat{w}^{*}_{n}\,$ satisfies UI condition (ref) and $P_{n}^{\dagger}(\Bar{w}) \in \mathcal{P}_{n}$ for all $\Bar{w} \in \Lambda_{n}$, $P_{n} \in \mathcal{P}_{n}$, and $n$. Then for any estimator $\hat{\tau}_{\hat{w}_{n},n}$ with weights $\hat{w}_{n}$ satisfying Assumption (ref), UI condition (ref), and $w_{n} \in \Lambda_{n}$ for all $P_{n} \in \mathcal{P}_{n}$ and $n$, the robust estimator $\hat{\tau}_{n}^{*} = \hat{w}^{*}_{n}\,'\hat{\theta}_{n}$ has a lower asymptotic maximum bias in the sense that \begin{align*} \limsup_{n \overset{\rightarrow}\infty} \left(\sup_{P_{n} \in \mathcal{P}_{n}} \max_{\lambda \in \Lambda_{n}} \left|\mathbb{E}_{P_{n}}\left[\hat{\tau}_{n}^{*}\right] - \tau_{\lambda,n}\right|\right) \leq \limsup_{n \overset{\rightarrow}\infty} \left(\sup_{P_{n} \in \mathcal{P}_{n}} \max_{\lambda \in \Lambda_{n}} \left|\mathbb{E}_{P_{n}}\left[\hat{\tau}_{\hat{w}_{n},n}\right] - \tau_{\lambda,n}\right|\right), \end{align*} where the right-hand side is finite.
proofSee Appendix (ref).

This optimality result presumes that $\hat{w}^{*}_{n}\,$ satisfies UI condition (ref). This holds, for example, when the class of alternatives $\Lambda(S)$ is uniformly bounded over $S \in \mathcal{S}$---not just $S \in \mathbb{S}$---as with any subset of the simplex. In other cases, it may be possible to argue the UI condition directly. For example, $\hat{w}^{*}_{n}\,$ reduces to the GLS weights in formula (ref) under the bounded variance class, from which it suffices to impose appropriate UI conditions on $\widehat{\Sigma}_{n}$. Note that even without UI conditions, $\hat{w}^{*}_{n}\,$ is still uniformly consistent for the weights $w^{*}_{n}$ that minimize the population maximum distance, as established in Proposition (ref).

Validity of the Heterogeneity UCB

The sample heterogeneity in parameters is the square root of the residual sum of squares from the GLS regression of the parameters $\theta_{n}$ on a constant $\bm{1}$ weighted by $\widehat{\Sigma}_{n}^{-1}$:

align*[align* omitted — 412 chars of source]

where $\hat{A}_{n}$ is the annihilator matrix for $\widehat{\Sigma}_{n}^{-1/2}\bm{1}$. The function $\hat{H}_{n}$ used to measure heterogeneity is random due to $\widehat{\Sigma}_{n}$. A different way to measure heterogeneity is the population heterogeneity in parameters, which replaces $\widehat{\Sigma}_{n}$ with $\Sigma_{n}$:

align*[align* omitted — 321 chars of source]

The heterogeneity in estimates is $\hat{H}_{n}(\hat{\theta}_{n})$. The heterogeneity UCB at significance level $\beta$ is

align*[align* omitted — 369 chars of source]

The following result establishes uniform asymptotic validity of $\hat{\eta}_{1-\beta,n}$ for inference on the sample heterogeneity in parameters $\hat{H}_{n}(\theta_{n})$, which will turn out to be sufficient for establishing uniform asymptotic validity of the robust CI---see Remark (ref) below for why I consider inference on the sample measure $\hat{H}_{n}(\theta_{n})$ instead of the population measure $H_{n}(\theta_{n})$.

propositionUUnder Assumptions (ref) and (ref), \begin{align} \lim_{n \overset{\rightarrow}\infty} \sup_{P_{n} \in \mathcal{P}_{n}} \left|\mathbb{P}_{P_{n}}\left\{\hat{H}_{n}(\theta_{n}) \leq \hat{\eta}_{1-\beta,n}\right\} - (1-\beta)\right|\mathds{1}\left\{H_{n}(\theta_{n}) > 0\right\} = 0. \end{align}
proofSee Appendix (ref).

Note that when $H_{n}(\theta_{n}) = 0$, the UCB $\hat{\eta}_{1-\beta,n} \geq 0$ always covers $\hat{H}_{n}(\theta_{n})$, so multiplication by the indicator $\mathds{1}\left\{H_{n}(\theta_{n}) > 0\right\}$ restricts attention to sequences where heterogeneity is nontrivial, in analogy to quantile-unbiasedness criterion (ref) from the normal model. Thus, (ref) implies

align*[align* omitted — 191 chars of source]

In this sense, $\hat{\eta}_{1-\beta,n}$ is a uniformly asymptotically valid UCB, in analogy to criterion (ref).

remark[Sample versus Population Heterogeneity] Consider a sequence where the normalized population heterogeneity diverges: $\sqrt{n}H_{n}(\theta_{n}) \overset{}{\rightarrow}\infty$. Given the norm bound on $\theta_{n}$ and the $\sqrt{n}$-consistency of $\widehat{\Sigma}_{n}$, this implies $\sqrt{n}\hat{H}_{n}(\theta_{n}) \overset{p}{\rightarrow} \infty$. To see what this means for validity of $\hat{\eta}_{1-\beta,n}$, consider noncentral chi-squared asymptotics for the heterogeneity in estimates $\hat{H}_{n}(\hat{\theta}_{n})$ relative to the sample heterogeneity in parameters $\hat{H}_{n}(\theta_{n})$: \begin{align} \frac{\sqrt{n}(\hat{H}_{n}^{2}(\hat{\theta}_{n}) - \hat{H}_{n}^{2}(\theta_{n}))}{2\hat{H}_{n}(\theta_{n})} = \frac{\hat{Z}_{n}' \hat{Q}_{n} \hat{Z}_{n}}{2\sqrt{n}\hat{H}_{n}(\theta_{n})} + \frac{(\hat{A}_{n} \widehat{\Sigma}_{n}^{-1/2}\theta_{n})'}{\left\|\hat{A}_{n} \widehat{\Sigma}_{n}^{-1/2}\theta_{n}\right\|} \widehat{\Sigma}_{n}^{-1/2}\hat{Z}_{n}, \quad \hat{Z}_{n} = \sqrt{n}(\hat{\theta}_{n} - \theta_{n}). \end{align} In this asymptotic regime, the first term is $o_{p}(1)$ while the second term converges in distribution to a standard normal. The asymptotic validity of $\hat{\eta}_{1-\beta,n}$ follows from expression (ref) combined with bounds from seri2015tight that imply the uniform convergence of a noncentral chi-squared CDF with diverging noncentrality parameter to a normal CDF; see Appendix (ref) for details. By comparison, the population heterogeneity analogue of (ref) yields \begin{align*} \frac{\sqrt{n}(\hat{H}_{n}^{2}(\hat{\theta}_{n}) - H_{n}^{2}(\theta_{n}))}{2H_{n}(\theta_{n})} = \frac{\hat{Z}_{n}' \hat{Q}_{n} \hat{Z}_{n}}{2\sqrt{n}H_{n}(\theta_{n})} + \frac{(\hat{A}_{n} \widehat{\Sigma}_{n}^{-1/2}\theta_{n})'}{\left\|A_{n} \Sigma_{n}^{-1/2}\theta_{n}\right\|} \widehat{\Sigma}_{n}^{-1/2}\hat{Z}_{n} + \frac{\sqrt{n}(\hat{H}_{n}^{2}(\theta_{n}) - H_{n}^{2}(\theta_{n}))}{2H_{n}(\theta_{n})}, \end{align*} where the first two terms behave like before, but now there is a third term where \begin{align*} \left|\frac{\sqrt{n}(\hat{H}_{n}^{2}(\theta_{n}) - H_{n}^{2}(\theta_{n}))}{2H_{n}(\theta_{n})}\right| \lesssim_{p} \frac{\sqrt{n}\left\|\widehat{\Sigma}_{n} - \Sigma_{n}\right\|}{H_{n}(\theta_{n})} \lesssim_{p} \frac{1}{H_{n}(\theta_{n})}, \end{align*} which need not be $o_{p}(1)$. Proposition (ref) avoids this term altogether.
remark[Analogy to Heterogeneity in the Normal Model] As noted in Remark (ref), objects $(\hat{\theta}, \Sigma, \theta)$ in the normal model are analogous to objects $(\hat{\theta}_{n}, \Tilde{\Sigma}_{n}, \theta_{n})$ in the asymptotic environment, where $\Tilde{\Sigma}_{n} = \widehat{\Sigma}_{n}/n$. The corresponding GLS weighting matrices are $\Sigma^{-1}$ and $\Tilde{\Sigma}_{n}^{-1} = n\widehat{\Sigma}_{n}^{-1}$. Thus, using $\hat{\eta}_{1-\beta}$ for inference on $H(\theta)$ in the normal model is analogous to using $\Tilde{\eta}_{1-\beta,n} = \sqrt{n}\hat{\eta}_{1-\beta,n}$ for inference on $\Tilde{H}_{n}(\theta_{n}) = \sqrt{n}\hat{H}_{n}(\theta_{n})$ in the asymptotic environment.

Validity of the Robust CI

I now establish the asymptotic validity of robust CIs centered at weight vectors $\hat{w}_{n}$ that satisfy Assumption (ref), which include the minimax-bias weights $\hat{w}^{*}_{n}\,$ in view of Proposition (ref).

For centering weights $\hat{w}_{n}$, the bias UCB is

align*[align* omitted — 366 chars of source]

For standard error $\Tilde{\sigma}_{\hat{w}_{n},n} = \sqrt{\hat{w}_{n}'\Tilde{\Sigma}_{n} \hat{w}_{n}} = \sqrt{\hat{w}_{n}'\widehat{\Sigma}_{n} \hat{w}_{n}}/\sqrt{n} = \hat{\sigma}_{\hat{w}_{n},n}/\sqrt{n}$, the robust CI is

align*[align* omitted — 706 chars of source]

The validity of $CI^{*}_{\hat{w}_{n},n}$ depends on its coverage for weights in the population class $\Lambda_{n} \subseteq \mathbb{W}$, where $\mathbb{W} \subseteq \mathcal{W}$ is the compact convex set with positive diameter from Assumption (ref). The uniform consistency in Proposition (ref) shows that the estimated class $\hat{\Lambda}_{n}$ is close to the population class $\Lambda_{n}$ with high probability as $n \overset{}{\rightarrow}\infty$. However, this is not the same as $\hat{\Lambda}_{n}$ containing the alternatives from $\Lambda_{n}$ with high probability. To see this, recall that Assumption (ref) yields the inequality constraint representations

align*[align* omitted — 233 chars of source]

where the former holds with probability uniformly approaching one under Assumption (ref). The issue is at the boundary of the inequality constraint, where small estimation errors can lead to exclusion, $g(\lambda, \hat{S}_{n}) > 0$, of population alternatives where $g(\lambda, S_{n}) = 0$. Since membership is discontinuous at the boundary, uniform consistency of $\hat{S}_{n}$ to $S_{n}$ generally does not ensure that $\hat{\Lambda}_{n}$ contains the alternatives in $\Lambda_{n}$. However, as I show in examples further below, there will typically exist a subclass $\Lambda_{0}(S) \subseteq \Lambda(S)$ such that, for $\Lambda_{0,n} = \Lambda_{0}(S_{n})$,

align[align omitted — 225 chars of source]

For subclasses $\Lambda_{0,n}$ satisfying containment condition (ref), the robust CI constructed with $\hat{\Lambda}_{n}$ provides uniformly asymptotically valid coverage for the target alternatives $\lambda \in \Lambda_{0,n}$.

propositionULet Assumptions (ref)-(ref) be satisfied and suppose that $\Lambda_{0,n} \subseteq \Lambda_{n}$ satisfies containment condition (ref). Then \begin{align} \liminf_{n \overset{\rightarrow}\infty} \inf_{P_{n} \in \mathcal{P}_{n}} \inf_{\lambda \in \Lambda_{0,n}}\mathbb{P}_{P_{n}}\left\{\tau_{\lambda,n} \in CI^{*}_{\hat{w}_{n},n}\right\} \geq 1-(\alpha+\beta). \end{align}
proofSee Appendix (ref).

The asymptotic coverage in (ref) is analogous to the coverage in (ref) from the normal model, but stated for the subclass $\Lambda_{0,n}$ rather than the population class $\Lambda_{n}$. This suggests viewing the class $\Lambda_{0,n}$ in Proposition (ref) as a target class, $\Lambda_{n} \supseteq \Lambda_{0,n}$ as an enlargement of the target class, and $\hat{\Lambda}_{n}$ as an estimator of the enlargement $\Lambda_{n}$. In other words, to ensure asymptotic coverage for target alternatives $\Lambda_{0,n}$, it suffices to use an enlarged estimator $\hat{\Lambda}_{n} \supseteq \hat{\Lambda}_{0,n}$ to construct the robust CI rather than using the sample analogue $\hat{\Lambda}_{0,n}$ of the target. For example, if the class $\hat{\Lambda}_{0,n} = \Lambda(\hat{S}_{n}; \kappa_{0})$ increases in a scalar parameter $\kappa$, this suggests using $\hat{\Lambda}_{n} = \Lambda(\hat{S}_{n};\kappa)$ with a larger $\kappa > \kappa_{0}$. Importantly, unlike the enlarged $\Lambda_{n}$, I do not require the target $\Lambda_{0,n}$ to satisfy the conditions of Assumption (ref), so long as it satisfies containment condition (ref) relative to $\Lambda_{n}$. For example, a target bounded variance class with $r_{0}=1$ may not satisfy the Slater condition, but nevertheless satisfy containment condition (ref) relative to an enlargement with $r > 1$ that does satisfy the Slater condition; see the examples below.

remark[Caveats Regarding Class Estimation Error] If there is no class estimation error (i.e., $\hat{\Lambda}_{n} = \Lambda_{n}$), then containment condition (ref) holds for $\Lambda_{0,n} = \Lambda_{n}$ and Proposition (ref) applies with $\Lambda_{n} = \Lambda_{0,n}$. However, Proposition (ref) still assumes that $\Lambda_{n}$ satisfies Assumption (ref). Thus, even when there is no class estimation error, it can still be useful to apply Proposition (ref) to $\Lambda_{n} \supseteq \Lambda_{0,n}$ when the target class $\Lambda_{0,n}$ may violate the conditions of Assumption (ref), as with the aforementioned bounded variance class with $r=1$ and the Slater condition.

To establish containment condition (ref), a practical approach is to show that the inequality constraint from $\Lambda_{n}$ is uniformly slack when evaluated over $\Lambda_{0,n}$. In particular, it suffices to show the existence of a constant $\nu > 0$ such that for all $P_{n} \in \mathcal{P}_{n}$ and $n$,

align[align omitted — 111 chars of source]

By Lemma (ref), slack condition (ref) implies containment condition (ref) under Assumptions (ref) and (ref). I now show how to derive $\Lambda_{0,n}$ and $\nu$ for the benchmark classes of interest.

example*[Bounded Variance, continued] For $S_{n} = (\Sigma_{n}, w_{n})$ and $r_{0} \geq 1$, the target class is \begin{align*} \Lambda_{0,n} = \left\{\lambda \in \mathcal{W}: g(\lambda, \Sigma_{n}, w_{n}; r_{0}) \leq 0\right\} = \left\{\lambda \in \mathcal{W}: \sqrt{\lambda'\Sigma_{n} \lambda} \leq r_{0}\sqrt{w_{n}'\Sigma_{n} w_{n}}\right\}. \end{align*} For $r > r_{0}$, the enlarged class is \begin{align*} \Lambda_{n} = \left\{\lambda \in \mathcal{W}: g(\lambda, \Sigma_{n}, w_{n}; r) \leq 0\right\} = \left\{\lambda \in \mathcal{W}: \sqrt{\lambda'\Sigma_{n} \lambda} \leq r\sqrt{w_{n}'\Sigma_{n} w_{n}}\right\}. \end{align*} Observe that \begin{align*} \sup_{\lambda \in \Lambda_{0,n}}g(\lambda, \Sigma_{n}, w_{n}; r) = \sup_{\lambda \in \Lambda_{0,n}}\sqrt{\lambda'\Sigma_{n} \lambda} - r\sqrt{w_{n}'\Sigma_{n} w_{n}} \leq -(r - r_{0})\sqrt{w_{n}'\Sigma_{n} w_{n}} \leq 0. \end{align*} Since $e_{\min}(\Sigma_{n}) \geq 1/\Bar{e}$ and $\left\|w_{n}\right\| \geq 1/\sqrt{K}$, slack condition (ref) is satisfied: \begin{align*} \sup_{\lambda \in \Lambda_{0,n}}g(\lambda, \Sigma_{n}, w_{n}; r) \leq -\nu, \quad \nu = (r - r_{0})/\sqrt{\Bar{e}K} > 0. \end{align*} By Proposition (ref), a robust CI constructed with estimated class \begin{align*} \hat{\Lambda}_{n} = \left\{\lambda \in \mathcal{W}: g(\lambda, \widehat{\Sigma}_{n}, \hat{w}_{n}; r) \leq 0\right\} = \left\{\lambda \in \mathcal{W}: \sqrt{\lambda'\widehat{\Sigma}_{n} \lambda} \leq r\sqrt{\hat{w}_{n}'\widehat{\Sigma}_{n} \hat{w}_{n}}\right\} \end{align*} provides uniformly asymptotically valid coverage for $\lambda \in \Lambda_{0,n}$. These derivations also apply to the bounded variance simplex by replacing $\mathcal{W}$ with $\mathcal{W}_{+}$ and restricting to $\hat{w}_{n}, w_{n} \in \mathcal{W}_{+}$.
example*[Truncated Simplex, continued] For $S_{n} = w_{n} \in \mathcal{W}_{+}$ and $\epsilon_{0} \in [0,1)$, the target class is \begin{align*} \Lambda_{0,n} = \left\{\lambda \in \mathcal{W}_{+}: g(\lambda, w_{n}; \epsilon_{0}) \leq 0\right\} = \left\{\lambda \in \mathcal{W}_{+}: \lambda \geq (1-\epsilon_{0})w_{n}\right\}. \end{align*} For $\epsilon \in (\epsilon_{0}, 1]$, the enlarged class is \begin{align*} \Lambda_{n} = \left\{\lambda \in \mathcal{W}_{+}: g(\lambda, w_{n}; \epsilon) \leq 0\right\} = \left\{\lambda \in \mathcal{W}_{+}: \lambda \geq (1-\epsilon)w_{n}\right\}. \end{align*} Observe that \begin{align*} \sup_{\lambda \in \Lambda_{0,n}}g(\lambda, w_{n}; \epsilon) = \max_{k \in \{1,\ldots,K\}}\left((1-\epsilon)w_{n,k}-\inf_{\lambda \in \Lambda_{0,n}}\lambda_{k}\right) \leq -(\epsilon - \epsilon_{0})\min_{k \in \{1,\ldots,K\}}w_{n,k}. \end{align*} Suppose that $w_{n,k}$ is uniformly bounded away from zero: \begin{align*} \inf_{n}\inf_{P_{n} \in \mathcal{P}_{n}} w_{n,k} \geq w_{\min} > 0, \quad \forall k \in \{1,\ldots,K\}. \end{align*} This holds, for instance, for the equal weights vector $w_{n} = \bm{1}/K$. Given this uniform lower bound on $w_{n,k}$, slack condition (ref) is satisfied: \begin{align*} \sup_{\lambda \in \Lambda_{0,n}}g(\lambda, w_{n}; \epsilon) \leq -\nu, \quad \nu = (\epsilon - \epsilon_{0})w_{\min} > 0. \end{align*} By Proposition (ref), a robust CI constructed with estimated class \begin{align*} \hat{\Lambda}_{n} = \left\{\lambda \in \mathcal{W}_{+}: g(\lambda, \hat{w}_{n}; \epsilon) \leq 0\right\} = \left\{\lambda \in \mathcal{W}_{+}: \lambda \geq (1-\epsilon)\hat{w}_{n}\right\} \end{align*} provides uniformly asymptotically valid coverage for $\lambda \in \Lambda_{0,n}$.
example*[Covariate Balance, continued] For $S_{n} = (w_{n}, \bm{X}_{n}) \in \mathcal{W}_{+} \times \mathbb{R}^{K \times M}$ and $\Bar{c}_{0} \geq 0$, the target class is \begin{align*} \Lambda_{0,n} = \left\{\lambda \in \mathcal{W}_{+}: g(\lambda, w_{n}, \bm{X}_{n}; \Bar{c}_{0}) \leq 0\right\} = \left\{\lambda \in \mathcal{W}_{+}: c_{\lambda}(w_{n},\bm{X}_{n}) \leq \Bar{c}_{0}\right\}. \end{align*} For $\Bar{c} > \Bar{c}_{0}$, the enlarged class is \begin{align*} \Lambda_{n} = \left\{\lambda \in \mathcal{W}_{+}: g(\lambda, w_{n}, \bm{X}_{n}; \Bar{c}) \leq 0\right\} = \left\{\lambda \in \mathcal{W}_{+}: c_{\lambda}(w_{n},\bm{X}_{n}) \leq \Bar{c}\right\}. \end{align*} Slack condition (ref) is satisfied: \begin{align*} \sup_{\lambda \in \Lambda_{0,n}}g(\lambda, w_{n}, \bm{X}_{n}; \Bar{c}) = -\nu, \quad \nu = \Bar{c} - \Bar{c}_{0} > 0. \end{align*} By Proposition (ref), a robust CI constructed with estimated class \begin{align*} \hat{\Lambda}_{n} = \left\{\lambda \in \mathcal{W}_{+}: g(\lambda, \hat{w}_{n}, \hat{\bm{X}}_{n}; \Bar{c}) \leq 0\right\} = \left\{\lambda \in \mathcal{W}_{+}: c_{\lambda}(\hat{w}_{n},\hat{\bm{X}}_{n}) \leq \Bar{c}\right\} \end{align*} provides uniformly asymptotically valid coverage for $\lambda \in \Lambda_{0,n}$.
remark[Reporting Convention for Boundary Alternatives] The takeaway from the above examples and discussions is that asymptotic coverage for boundary alternatives requires some care. For instance, a target bounded variance simplex class \begin{align*} \Lambda_{0,n} = \left\{\lambda \in \mathcal{W}_{+}: \sqrt{\lambda'\Sigma_{n} \lambda} \leq r_{0}\sqrt{w_{n}'\Sigma_{n} w_{n}}\right\}, \quad r_{0} \geq 1, \end{align*} may contain alternative weights that lie on the boundary of the variance constraint: $\displaystyle\sqrt{\lambda'\Sigma_{n} \lambda} = r_{0}\sqrt{w_{n}'\Sigma_{n} w_{n}}$. Proposition (ref) says that a sufficient avenue for asymptotic coverage of the entire $\Lambda_{0,n}$ is to use a robust CI constructed with an enlarged estimated class \begin{align*} \hat{\Lambda}_{n} = \left\{\lambda \in \mathcal{W}_{+} : \sqrt{\lambda'\widehat{\Sigma}_{n} \lambda} \leq r\sqrt{\hat{w}_{n}'\widehat{\Sigma}_{n} \hat{w}_{n}}\right\}, \quad r = r_{0}+\delta, \quad \delta > 0. \end{align*} The asymptotic coverage in Proposition (ref) applies to the $r_{0}$ target class, with coverage obtained through a small enlargement, $r = r_{0}+\delta$, of the estimated class. However, in the implementation and empirical applications, I will suppress this caveat and talk about coverage as if $\delta = 0$. The reason is that, in practice, replacing $r=r_{0}$ by $r=r_{0} +\delta$ for $\delta = 0.0001$ leaves the reported CIs unchanged at the displayed precision.

Practical Implementation

Motivated by the asymptotic results in Section (ref), but suppressing the sample size index $n$ for conciseness, I consider asymptotically normal estimates $\hat{\theta} \overset{a}{\sim} N(\theta, \Sigma/n)$, consistent covariance matrix estimator $\Tilde{\Sigma} = \widehat{\Sigma}/n \overset{a}{\approx} \Sigma/n$, consistent baseline weights estimator $\hat{w} \overset{a}{\approx} w$, and classes of alternative weights $\hat{\Lambda} = \Lambda(\hat{S})$ characterized by consistently estimated statistics $\hat{S} \overset{a}{\approx} S$: e.g., $\hat{S} = (\hat{w}, \widehat{\Sigma})$ for bounded variance class, $\hat{S} = \hat{w}$ for truncated simplex class, and $\hat{S} = (\hat{w}, \hat{\bm{X}})$ for covariate balance class with estimated covariate mean matrix $\hat{\bm{X}}$.

Section (ref) summarizes my robust inference procedures in terms of the above objects and discusses specification choices for $\hat{\Lambda}$. Sections (ref) and (ref) discuss practical implementations of these procedures in the context of event studies and multisite experiments, respectively.

Feasible Inference Procedures

The GLS regression of $\hat{\theta}$ on a constant $\bm{1}$ weighted by $\Tilde{\Sigma}^{-1}$ yields

align[align omitted — 418 chars of source]

where $\hat{\tau}_{\mathrm{GLS}}$ is the GLS estimator, $\Tilde{\sigma}_{\mathrm{GLS}}$ is the GLS standard error, $\Tilde{H}^{2}(\hat{\theta})$ is the GLS residual sum of squares, which yields heterogeneity in estimates $\Tilde{H}(\hat{\theta})$ with explicit formula

align[align omitted — 326 chars of source]

Following the steps of Section (ref), the heterogeneity UCB is

align[align omitted — 315 chars of source]

Following the steps of Section (ref), the minimax-bias weights and bias UCB are

align[align omitted — 361 chars of source]

For standard error $\Tilde{\sigma}^{*} = \sqrt{(\hat{w}^{*})'\Tilde{\Sigma} \hat{w}^{*}}$, the robust estimator and corresponding robust CI are

align[align omitted — 676 chars of source]

where $\text{cv}_{1-\alpha}\left(b\right)$ can be computed as the square root of the $(1-\alpha)$-quantile of the noncentral chi-squared distribution with one degree of freedom and noncentrality parameter $b^{2}$.

recipeImplementation of Inference Procedures \begin{recipenumerate} • Construct $\Tilde{\eta}_{1-\beta}$ as in (ref), based on (ref) or (ref). • Construct $\hat{w}^{*}$ and $\widehat{B}_{\min}^{\beta}(\hat{\Lambda})$ as in (ref). • Construct $\hat{\tau}^{*}$ and $CI^{*}$ as in (ref). • Report $(\hat{\tau}^{*}, CI^{*})$. \end{recipenumerate}

Recipe (ref) summarizes the implementation of my inference procedures for inputs given by (i) the estimates and covariance matrix estimator $(\hat{\theta}, \Tilde{\Sigma})$, for which I describe typical constructions in Section (ref) for event studies and Section (ref) for multisite experiments; (ii) the significance levels $\alpha$ and $\beta$, which by default can be set to $\alpha = \beta = 0.05$; and (iii) the class of alternatives $\hat{\Lambda}$.

In some cases there may be ambiguity over which class of alternatives to use. In particular, the class $\hat{\Lambda} = \hat{\Lambda}(\kappa)$ may be indexed by a scalar parameter $\kappa \in \mathbb{R}$ over which there is ambiguity. To address this, one can execute Recipe (ref) for a range of plausible $\kappa$. Alternatively, assuming that $\hat{\Lambda}(\kappa)$ is increasing in $\kappa$, one can compute the breakdown value $\kappa^{*}$, which is the smallest value of $\kappa$ at which the robust $CI^{*}(\kappa)$ includes a given threshold value $\tau_{0}$, such as $\tau_{0} = 0$.

example*[Bounded Variance, continued] From the derivations in (ref), it follows that Recipe (ref) is entirely determined by the GLS output (ref). In particular, for a lower one-sided $CI^{*}$ and baseline weights $\hat{w}$, Recipe (ref) under $\hat{\Lambda}_{\sigma}(r) = \{\lambda \in \mathcal{W}: \Tilde{\sigma}_{\lambda} \leq r\Tilde{\sigma}_{\hat{w}}\}$ for standard error ratio bound $r$ and standard error $\Tilde{\sigma}_{\hat{w}} = \displaystyle\sqrt{\hat{w}'\Tilde{\Sigma}\hat{w}}$ yields \begin{align*} \hat{\tau}^{*} = \hat{\tau}_{\mathrm{GLS}}, \quad \quad CI^{*}(r) = \left[\hat{\tau}_{\mathrm{GLS}} - z_{1-\alpha}\Tilde{\sigma}_{\mathrm{GLS}} - \Tilde{\eta}_{1-\beta}\sqrt{r^{2}\Tilde{\sigma}_{\hat{w}}^{2} - \Tilde{\sigma}_{\mathrm{GLS}}^{2}}, \infty\right). \end{align*} The breakdown value for a given threshold $\tau_{0}$ is \begin{align*} r^{*} = \frac{\Tilde{\sigma}_{\mathrm{GLS}}}{\Tilde{\sigma}_{\hat{w}}}\sqrt{\left(\frac{\max\left\{T_{\mathrm{GLS}}(\tau_{0}) - z_{1-\alpha}, 0\right\}}{\Tilde{\eta}_{1-\beta}}\right)^{2} + 1}, \quad T_{\mathrm{GLS}}(\tau_{0}) = \frac{\hat{\tau}_{\mathrm{GLS}} - \tau_{0}}{\Tilde{\sigma}_{\mathrm{GLS}}}. \end{align*} $CI^{*}(r)$ includes $\tau_{0}$ when $r \geq r^{*}$. If $\tau_{0} = 0$, this happens when the baseline $t$-statistic $T_{\mathrm{GLS}}(0)$ is sufficiently large relative to the degree of inferred heterogeneity $\Tilde{\eta}_{1-\beta}$.
example*[Truncated Simplex, continued] Given baseline weights $\hat{w}$, one can execute Recipe (ref) with $\hat{\Lambda}_{+}(\epsilon) = \{\lambda \in \mathcal{W}_{+}: \lambda \geq (1-\epsilon)\hat{w}\} = \{(1-\epsilon)\hat{w} + \epsilon \lambda: \lambda \in \mathcal{W}_{+}\}$ for some range of truncation parameters $\epsilon$, or compute breakdown values $\epsilon^{*}$. To interpret magnitudes of $\epsilon$, consider the equal weights baseline $\hat{w} = w_{\mathrm{EW}} = \bm{1}/K$. In this case, $\epsilon$ is the maximum possible discrepancy in weights $|\lambda_{k}-\lambda_{k'}|$ between two groups $k$ and $k'$. Moreover, note that the maximum amount of weight that can be reallocated across groups when moving from $w_{\mathrm{EW}}$ to $\lambda \in \Lambda_{+}(\epsilon)$ is \begin{align*} \left((1-\epsilon)\frac{1}{K} + \epsilon\right) - \frac{1}{K} = \frac{K-1}{K}\epsilon. \end{align*} On the other hand, the amount of weight that is reallocated when moving all the weight from $J \in \{1,\ldots,K-1\}$ of the groups to the other $K-J$ groups (i.e., a leave-$J$-out procedure) is equal to $J/K$. Equating these two reallocation quantities yields the $\epsilon$ calibration \begin{align} \frac{K-1}{K}\epsilon = \frac{J}{K} \iff \epsilon = \frac{J}{K-1}. \end{align} Thus, in the EW baseline, choosing $\epsilon$ according to (ref) allows for the same maximal reallocation of weight as simply dropping $J$ of the groups. In this sense, $\epsilon$ is a rescaled leave-$J$-out magnitude: $(K-1)\epsilon$ is the corresponding number of groups dropped, and $(K-1)\epsilon/K$ is the corresponding fraction of groups dropped. This calibration allows one to maintain the mathematical structure of $\Lambda_{+}(\epsilon)$ while casting the logic of $\epsilon$-departures in terms of the potentially more familiar logic of leave-$J$-out procedures.
example*[Covariate Balance, continued] Given covariates $\hat{\bm{X}}$, one can execute Recipe (ref) with $\hat{\Lambda}_{X}(\Bar{c}) = \left\{\lambda \in \mathcal{W}_{+}: c_{\lambda}(\hat{\bm{X}}) \leq \Bar{c}\right\}$ for some range of balance gap bounds $\Bar{c}$, or compute breakdown values $\Bar{c}^{*}$. To interpret magnitudes of $\Bar{c}$, consider the group-level balance gaps \begin{align} c_{k}(\hat{\bm{X}}) = \max_{m \in \{1,\ldots,M\}}\frac{\left|\hat{\bm{X}}_{m,k} - \hat{w}'\hat{\bm{X}}_{m}\right|}{sd(\hat{\bm{X}}_{m})}. \end{align} Intuitively, $c_{k}(\hat{\bm{X}})$ is the balance gap that arises from a population with covariate means equal to those of group $k$. If one assumes that readers are interested in balance gaps within the range of $\{c_{k}(\hat{\bm{X}})\}_{k=1}^{K}$, then the deciles $(\Bar{c}_{d})_{d=1}^{9}$ of the group-level balance gaps provide one simple way to determine relevant magnitudes of $\Bar{c}$. For example, $\Bar{c} = \Bar{c}_{5}$ allows one to assess robustness to populations that have covariate profiles consistent with the median group-level balance gap.

Implementation in Event Studies

\paragraph{Context.} Following the exposition in Section (ref), an event study has units $i$ treated at different time periods $t$, allowing one to estimate the dynamic effects of treatment on outcomes $Y_{it}$. In this setting, two key sources of heterogeneity are (i) across treatment cohorts $G_{i}$, which denote the periods of initial treatment for units $i$; and (ii) across event times $\ell \geq 0$, which denote the number of periods since initial treatment. Given the $\operatorname{ATT}_{g,t}$ object in (ref), the average treatment effect $\ell$ periods after initial treatment for cohort $G_{i} = g$ is

align*[align* omitted — 102 chars of source]

Under event-study assumptions, sun2021estimating show that event-time $\ell$ coefficients from conventional TWFE regression specifications recover estimands of the form

align*[align* omitted — 137 chars of source]

which allow treatment effects from different event times $\ell' \neq \ell$ and negative weights in $\Tilde{w}_{\ell}$. As surveyed in roth2023s, the recent event studies literature proposes heterogeneity-robust (HR) approaches that, for weights $w_{\ell} > 0$, recover event-time estimands of the form

align*[align* omitted — 104 chars of source]

preventing negative weighting and excluding treatment effects from different event times.

For each target event time $\ell$, a researcher may report conventional estimators and CIs for both a TWFE specification and one of the proposed HR approaches and examine the stability of results. There are two ways to map this setup to the notation of Recipe (ref).

itemize• Consider TWFE as the baseline and estimate $(\Tilde{w}_{\ell, (g,\ell')}, \operatorname{ATT}_{g,g+\ell'})$ across $(g,\ell')$. • Consider an HR approach as the baseline and estimate $(w_{\ell,g}, \operatorname{ATT}_{g,g+\ell})$ across $g$.

While it is possible to do the former (see sun2021estimating), the latter has the advantage of automatically shutting off the $\ell' \neq \ell$ channels and reducing the number of objects to be estimated. In fact, the ingredients required for implementing Recipe (ref) are available in many applications of HR approaches. For example, the sun2021estimating procedure already requires estimation of $(w_{\ell,g}, \operatorname{ATT}_{g,g+\ell})$, and so Recipe (ref) can be readily executed in applications that were already using the sun2021estimating procedure, or any other procedure that first estimates ATT objects and their corresponding standard errors, such as the doubly robust procedure of callaway2021difference, or stacked DiD procedures wing2024stacked.

\paragraph{Implementation.} I now use $k$ to index the groups $g$. For each target event time $\ell$, denote the number of available cohorts by $K_{\ell}$. Take an existing HR approach for which one can estimate baseline weights $\hat{w}_{\ell} \overset{a}{\approx} w_{\ell}$ and ATT objects

align*[align* omitted — 179 chars of source]

where $\Tilde{\Sigma}_{\ell}$ is a covariance matrix estimator. Let $\Tilde{\sigma}_{\hat{w},\ell} = \displaystyle\sqrt{\hat{w}_{\ell}'\Tilde{\Sigma}_{\ell}\hat{w}_{\ell}}$ denote the plug-in standard error for $\hat{\tau}_{\hat{w},\ell} = \hat{w}_{\ell}'\hat{\theta}_{\ell}$. Following equation (ref), consider the bounded variance simplex class

align*[align* omitted — 268 chars of source]

Executing Recipe (ref) with $\hat{\Lambda}_{\sigma, \ell}^{+}(r)$ produces $(\hat{\tau}_{\ell}^{*}(r), CI_{\ell}^{*}(r))$. For the choice of standard error ratio bound $r$, I propose the following implementation.

enumerate• Choose $r=1$, allowing one to assess robustness to the class of nonnegative weights that yield standard errors no larger than the HR baseline $\Tilde{\sigma}_{\hat{w},\ell}$. This is the minimal $r$ for which the baseline weights are included in $\hat{\Lambda}_{\sigma, \ell}^{+}(r)$. • Report breakdown value $r_{\ell}^{*}$, which gives the smallest value of $r$ at which the robust $CI_{\ell}^{*}(r)$ includes a threshold value $\tau_{0}$ (e.g., $\tau_{0} = 0$). If $r_{\ell}^{*}$ is large, then the robust CI only includes $\tau_{0}$ when it attempts to cover estimands that cannot be estimated precisely.

I apply this implementation to lakdawala2023dynamic in Section (ref).

Implementation in Multisite Experiments

Following the exposition in Section (ref), a multisite experiment has sites $k$ containing units $i$ over which ATEs are estimated:

align*[align* omitted — 144 chars of source]

where $\Tilde{\Sigma}$ is a covariance matrix estimator. Given independent sites, it suffices to obtain site-level objects $(\hat{\theta}_{k}, \Tilde{\sigma}_{k}, \widehat{E}_{k}[X_{i}])$, where estimated variances $\Tilde{\sigma}_{k}^{2}$ can be collected into a diagonal matrix $\Tilde{\Sigma}$ and estimated covariate means $\widehat{E}_{k}[X_{i}]$ can be collected into a matrix $\hat{\bm{X}} = (\widehat{E}_{1}[X_{i}], \ldots, \widehat{E}_{K}[X_{i}])'$. Given microdata on units $i$ from the sites, one can directly estimate $\widehat{E}_{k}[X_{i}]$ with sample means and $(\hat{\theta}_{k}, \Tilde{\sigma}_{k})$ via regression. For example, letting $\mathds{1}\left\{i \in k\right\}$ be an indicator for membership in site $k$, one can estimate the (no-intercept) saturated model

align*[align* omitted — 155 chars of source]

with OLS to obtain ATE estimates $\hat{\theta}_{k}$ and heteroskedasticity-robust standard errors $\Tilde{\sigma}_{k}$.

The baseline weights $\hat{w} \in \mathcal{W}_{+}$ represent a particular population/distribution over the set of sites. To model departures from the baseline population, one can use the truncated simplex class $\hat{\Lambda}_{+}(\epsilon)$ defined in (ref) or the covariate balance class $\hat{\Lambda}_{X}(\Bar{c})$ defined in (ref).

itemize• Departures in $\hat{\Lambda}_{+}(\epsilon)$ ensure that groups $k$ represented in the baseline population remain represented in the alternative population: $\hat{w}_{k} > 0 \implies \lambda_{k} > 0$ for each $\lambda \in \hat{\Lambda}_{+}(\epsilon)$. • Departures in $\hat{\Lambda}_{X}(\Bar{c})$ incorporate information about group-level covariates $\hat{\bm{X}}$.

To obtain benefits from both approaches, I propose the following implementation based on the truncated covariate balance class $\hat{\Lambda}_{X}^{\epsilon}(\Bar{c})$ defined in (ref).

enumerate• Choose the equal weights baseline $\hat{w} = w_{\mathrm{EW}} = \bm{1}/K$, which represents an empirical distribution over the sites, under which all groups $k$ will be represented in the alternative populations within $\hat{\Lambda}_{X}^{\epsilon}(\Bar{c}) = \Lambda_{+}(\epsilon) \cap \hat{\Lambda}_{X}(\Bar{c})$. • Formulate a choice of truncation parameter $\epsilon$. Following (ref), one can interpret $\epsilon$ in terms of the leave-$J$-out calibration. • Execute Recipe (ref) with $\hat{\Lambda}_{X}^{\epsilon}(\Bar{c})$ for a range of balance gap bounds $\Bar{c}$. Following (ref), one can use the site-level balance gap deciles $(\Bar{c}_{d})_{d=1}^{9}$.

I apply this implementation to Project STAR in Section (ref).

Empirical Applications

I now apply my procedures to the event study in lakdawala2023dynamic and the multisite experiment in Project STAR, following the implementations described in Sections (ref) and (ref). Below I use significance levels $\alpha = \beta = 0.05$ and hence confidence levels $1-\alpha = 95\%$ for conventional CIs, $1-\beta = 95\%$ for heterogeneity UCBs, and $1-(\alpha+\beta)=90\%$ for robust CIs.

Application to Lakdawala, Nakasone and Kho (2023)

lakdawala2023dynamic---henceforth LNK---study the rollout of school-based internet access in Peruvian public primary schools. The estimating sample contains schools that either gained internet access between 2007 and 2020 or remained unconnected by 2020, with second grade math and reading test scores observed from 2007--2016. LNK use a staggered-adoption event-study design: after absorbing school fixed effects, calendar-time controls, and time-varying school and student controls, the event-time coefficients compare changes in scores around a school's installation date to contemporaneous changes among schools that are not yet connected or remain unconnected. LNK's event-study figures report school-level dynamic effects for successive cohorts of second graders, with scores standardized within calendar year and effects interpreted relative to the year before internet installation.

LNK's Figure 2 shows a delayed achievement response. The TWFE event-study estimates, replicated in Figure (ref), are small immediately after internet installation and grow over time. In LNK's page 234 description of the medium-run effect, “By year 5, scores are 0.110 standard deviations higher for math” and “0.063 standard deviations higher for reading” relative to the year before installation. These dynamics are central to the interpretation of the intervention because LNK argue that schools require time to adapt to new internet access. The empirical target is therefore a dynamic path of treatment effects.

\paragraph{Weighting Issues.} Following Section (ref), let cohorts $k$ index the first year in which a school gains internet access and event times $\ell$ denote the time relative to first internet access. LNK's pages 241-42 discuss weighting issues: “Recent work \ldots raises some important issues with using two-way fixed effects estimation when there is staggered timing of treatment. We show that these issues do not explain our main estimates by illustrating robustness to the estimator proposed in Sun and Abraham (2021) \ldots The patterns and magnitudes of these estimates are very similar to our main estimates.” The sun2021estimating---henceforth SA---estimator corresponds to weights $\hat{w}_{\mathrm{SA},\ell}$. LNK's Figure A.12, replicated in Figure (ref), is therefore a robustness check that replaces the TWFE event-study aggregation with the SA aggregation in order to assess whether results are robust to alternative weights.\footnote{LNK use a bootstrap procedure to construct the SA CIs. Their replication files estimate cohort-specific event-study coefficients using the same controls and fixed effects as the main specification, drop cohorts not identified relative to the omitted period, aggregate cohort-event coefficients using the SA event-time weights, and bootstrap the aggregated coefficients using 500 replications stratified by first access year and clustered by school. The CIs reported here instead use the plug-in standard error computed from the estimated covariance matrix $\Tilde{\Sigma}_{\ell}$ for the cohort-event coefficient vector $\hat{\theta}_{\ell}$, treating the SA weights as fixed. Thus, while I replicate the same SA estimates, my reported SA CIs are not identical to LNK's bootstrap SA CIs.} This suggests a sharper robustness question: are the substantive conclusions robust only to the SA weights, or to a broader class of plausible event-time weights? To answer this question, I compute robust estimators and CIs for classes of alternative weights that nest the SA weights.

\paragraph{Measuring Heterogeneity.} Table (ref) reports the estimated event-time heterogeneity UCBs $\Tilde{\eta}_{\ell}$ used to construct the robust CIs. For event times with $\Tilde{\eta}_{\ell}=0$, the heterogeneity adjustment in the respective robust CI drops out. The last row reports $K_{\ell}$, the total number of cohorts $k$ at each event time $\ell$. Across all but one event time, the heterogeneity UCB is lower for the math cohort-level estimates than for the reading cohort-level estimates, previewing the greater degree of robustness for the math results.

table[table omitted — 620 chars of source]

\paragraph{Bounded Variance Simplex Class.} Consider the bounded variance simplex

align*[align* omitted — 297 chars of source]

with $r=1$, which corresponds to the class of simplex-weighted estimands that can be estimated at least as precisely as the SA estimand. For each event time $\ell$, I compute the robust $CI^{*}_{\ell}(1)$ centered at the robust estimator $\hat{\tau}_{\ell}^{*}(1)$. Figure (ref) presents three series: the TWFE CI centered at the TWFE estimate, the SA CI centered at the SA estimate, and the robust CI centered at the robust estimate. For math, the robust CIs are positive for event times 1 through 6 and only include zero at event time 0. At event time 5, the robust estimate is $0.099$ with CI $[0.055,0.143]$, close to both the TWFE and SA estimates. Thus, LNK's medium-run math conclusion is robust to alternative simplex weights that yield estimators at least as precise as the SA estimator. For reading, the robust estimates remain close to the SA path, but the conclusion that results are significantly positive is somewhat more sensitive. The robust CIs are positive for event times 1 through 4 and 6, while event time 5 barely includes zero: the robust estimate is $0.042$ with CI $[-0.0002,0.084]$.

figure[figure omitted — 608 chars of source]

Table (ref) presents breakdown values $r^{*}_{\ell}$, which for each $\ell$ gives the smallest value of $r$ at which the robust $CI^{*}_{\ell}(r)$ includes zero for the class $\hat{\Lambda}_{\sigma,\ell}^{+}(r)$. A value of $1.60$, for example, means that the robust CI begins to include zero once the allowable $\Tilde{\sigma}$ bound is expanded to about $1.60$ times the SA $\Tilde{\sigma}$. A dash means that zero is not reached even when considering the simplex without the variance constraint.\footnote{If $\Tilde{\eta}_{\ell}=0$, I still compute breakdown values using the same definition: report the smallest feasible value of the robustness parameter at which the robust CI contains zero, and report “--” if zero is never reached over the full class. This convention implies, for example, that math at $\ell=0$ has $r^{*}_{0}<1$ because zero is already included at the smallest feasible bounded-variance simplex class, while math at $\ell=6$ has no finite breakdown value because zero is not reached even when the variance constraint is no longer binding over the simplex.} The breakdown values reinforce the visual evidence. For math, event times 2 through 5 tolerate substantial increases in $r$ before the robust CI includes zero, with $r^*_5=1.60$ at the headline year-5 event time. Event time 6 is even stronger in this sense: zero is not reached over the full simplex. For reading, by contrast, the breakdown values are close to one for event times 1 through 5, and event time 5 already includes zero at $r=1$. Overall, these results support the baseline math conclusions strongly, but not as strongly for the reading conclusions.

table[table omitted — 563 chars of source]

Application to Project STAR

The Tennessee Project STAR (Student/Teacher Achievement Ratio) experiment randomized students in seventy-nine Tennessee public elementary schools to classes with different numbers of students to estimate the causal effects of class size on test scores achilles2008tennessee, krueger1999experimental. The experimental arms were (i) regular-sized classes (20-25 students), small classes (13-17 students), and regular-sized classes with a teacher aide. To define the sample for my analysis, I drop observations in the teacher-aide arm and focus on the effect of being assigned to a small class versus a regular-sized class. Following goldsmith2024contamination, I focus on the effects of treatment in kindergarten, which mitigates attrition and other complications. This leaves $K=78$ schools and a total of $n=3783$ students. Following schanzenbach2006have, I define the outcome as the average of a student's math and reading scores on the SAT, where the scores are demeaned and standardized relative to control-group scores. This puts the outcome into effect size units that have benchmarks in causal studies of preK-12 education interventions on student achievement: <0.05 = small, 0.05-0.20 = medium, >0.20 = large kraft2020interpreting.

In my notation, $\theta_{k}$ is the ATE of being assigned to a small class on the test score outcome at school $k$, $\hat{\theta}_{k}$ is the estimated ATE at school $k$ from regressing the outcome on school-specific treatment indicators and fixed effects, $\Tilde{\Sigma}$ is the heteroskedasticity-robust covariance matrix from this regression, and the baseline weights are equal weights (EW) $w = w_{\mathrm{EW}} = \bm{1}/K$ so that $\tau_{w}(\theta) = \tau_{\mathrm{EW}}(\theta)$ is the ATE in the empirical distribution of Project STAR schools. The baseline results establish medium-sized positive effects of small class size on test scores. In particular, the EW estimate is $\hat{\tau}_{\mathrm{EW}}=0.188$ with standard error $\Tilde{\sigma}_{\mathrm{EW}}=0.028$, yielding conventional EW CI $[0.133, 0.242]$. My baseline estimate is consistent with the one in schanzenbach2006have. However, my standard error is about 0.01 smaller. In particular, schanzenbach2006have clusters her standard errors at the classroom level. By contrast, I follow goldsmith2024contamination and use heteroskedasticity-robust standard errors since the randomization of students to classrooms was at the individual level.\footnote{This is consistent with the when-to-cluster framework of abadie2023should.}

\paragraph{Weighting Issues.} As reviewed by schanzenbach2006have, Project STAR has been used in many policy discussions. However, schanzenbach2006have notes external validity concerns:

itemize• “A few aspects of the sample may limit the validity of generalizing the study to other settings. In order to be eligible to participate in the program, schools were required to have a minimum-size cohort of 57 students \ldots As a result, the schools that participated were about 25% larger, on average, than other Tennessee schools”; • “Because of requirements imposed by the legislature for geographic diversity, schools in inner cities were overrepresented, and the students included were more economically disadvantaged and more likely to be African American”; • “Finally, average school spending in Tennessee was about three-fourths of the nationwide average, and teachers were less likely to have a master’s degree.”

Thus, the Project STAR results may not be representative of Tennessee (or U.S.) schools more broadly. Formally, the EW estimand averages over the empirical distribution of STAR schools, while alternative estimands of interest may place unequal weights across the STAR schools. This suggests a robustness question: are the baseline STAR results robust to alternative choices of weights/distributions? To answer this question, I compute robust estimators and CIs for classes of alternative weights that nest the EW weights, following implementations from Section (ref).

\paragraph{Measuring Heterogeneity.} The estimated heterogeneity UCB is $\Tilde{\eta} = 16.214$, which is over twice as large as the baseline $t$-statistic $\hat{\tau}_{\mathrm{EW}}/\Tilde{\sigma}_{\mathrm{EW}} = 6.714$. This suggests a meaningful degree of heterogeneity in the school-level ATE estimates, which previews the sensitivity of baseline results to the choice of weights.

\paragraph{Truncated Simplex Class.} I first consider the truncated simplex class

align*[align* omitted — 207 chars of source]

which collects the set of $\epsilon$-deviations in weight distribution from the STAR empirical distribution. For each $\tau_{0} \in \{0, 0.05\}$, Table (ref) presents the breakdown value $\epsilon^{*}$, which is the smallest value of $\epsilon$ at which $\tau_{0}$ is included in the robust $CI^{*}(\epsilon)$ centered at robust estimate $\hat{\tau}^{*}(\epsilon)$ with standard error $\Tilde{\sigma}(\epsilon^{*})$. The breakdown value for $\tau_{0}=0.05$ is $\epsilon^{*}=0.0169$, while the breakdown value for $\tau_{0}=0$ is $\epsilon^{*}=0.0261$. These are both small departures from EW: even though the corresponding robust estimates are larger than the baseline EW estimate $\hat{\tau}_{\mathrm{EW}}=0.188$, the robust CIs indicate that baseline inferences are robust to only small departures from EW.

table[table omitted — 542 chars of source]

\paragraph{Truncated Covariate Balance Class.} I intersect the truncated simplex with the covariate balance class, yielding

align*[align* omitted — 371 chars of source]

Table (ref) reports the Project STAR school-level covariates $\hat{\bm{X}}_{m}$ used in my analysis.

table[table omitted — 670 chars of source]

I fix $\epsilon = \epsilon^{*}(0) = 0.0261$ at the truncated simplex breakdown value for $\tau_{0} = 0$ from the previous analysis and consider balance gap bounds $(\Bar{c}_{d})_{d=1}^{9}$ equal to the deciles of the school-level balance gaps $c_{k}(\hat{\bm{X}})$. Intuitively, $c_{k}(\hat{\bm{X}})$ is the balance gap from a population that has covariate means $\hat{\bm{X}}_{m,k}$ equal to those of school $k$. The deciles $\Bar{c}_{d}$ presented in Table (ref) therefore represent typical values for the balance gaps.

table[table omitted — 426 chars of source]

Figure (ref) presents robust intervals $CI^{*}(\Bar{c}_{d})$ centered at robust estimates $\hat{\tau}^{*}(\Bar{c}_{d})$ under $\epsilon^{*}(0)$. Across every balance gap decile $\Bar{c}_{d}$, the robust CIs include $\tau_{0} = 0.05$. Thus, while the covariate balance constraints allow the robust CIs to exclude zero relative to imposing only the truncated simplex constraint, this is not enough to allow one to robustly conclude that effects are medium-sized under $\epsilon^{*}(0)$-departures from EW. Overall, these results provide further support for the conclusion that baseline STAR inferences are robust to only small departures from EW. There is too much heterogeneity for the conclusions of a baseline CI to be (uniformly) informative for alternative weights that depart meaningfully from the baseline weights.

figure[figure omitted — 279 chars of source]

Conclusion

Weighted estimands are ubiquitous in empirical research. The choice of weights is not merely a statistical nuisance: it shapes the empirical and policy relevance of a researcher's weighted estimand. Conventional inference procedures can therefore be limited when readers entertain alternative weights, leading to ambiguity and disagreement over the choice of weights.

This paper develops inference procedures that account for these issues. The key geometric observation is that the difference between two weighted estimands can be controlled by two objects: the heterogeneity of the underlying parameters and the distance between the corresponding weights. This decomposition leads to two practical tools. First, I derive minimax-bias weights that minimize the maximum distance to a class of alternative weights. The resulting robust estimator provides a natural default when researchers face ambiguity over the choice of weights. Second, I construct robust confidence intervals by combining a confidence bound for the heterogeneity in parameters with the maximum distance between the centering weights and the alternative weights. These intervals provide coverage guarantees for classes of weighted estimands rather than for a single baseline estimand.

My framework accommodates many empirically relevant classes of alternatives. The bounded variance class captures concerns about comparing the baseline estimand to alternatives that can be estimated with reasonable precision. The truncated simplex class captures departures from a baseline target population while ensuring that all represented groups continue to receive positive weight. The covariate balance class captures external validity concerns by restricting attention to alternative populations whose covariate means remain close to the baseline. These classes can also be intersected, allowing researchers to combine precision, positivity, and covariate-balance restrictions in a single robustness analysis.

The empirical applications illustrate that robustness to the choice of weights can be context-specific. For the event study in lakdawala2023dynamic, conclusions are broadly robust to classes of weights that yield reasonable estimation precision. For the experiment in Project STAR, however, conclusions are sensitive even to small departures from the baseline weighting scheme. These findings highlight the value of reporting robust estimators and intervals alongside conventional ones: doing so makes explicit whether the conclusions reached under a researcher's baseline weights are informative for alternative weighted estimands that may be of interest to readers.