Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,999 characters · 7 sections · 37 citation commands
Inference for the Marginal Value of Public Funds
\doparttoc \faketableofcontents \thispagestyle{empty}
\setcounter{page}{1}
Policymakers increasingly rely on evidence from randomized evaluations to guide decisions about which programs to fund and scale. These evaluations often report multiple estimated causal effects---for example, a tax credit program's impacts on earnings, after-tax income, and labor force participation---but policymakers typically care about a summary measure that aggregates these effects into a measure of overall cost-effectiveness of the policy, such as the Marginal Value of Public Funds (MVPF). Determining whether a program delivers “bang for the buck” thus requires inference on a scalar function of several estimated causal effects, rather than on any individual effect alone.
Conducting valid inference on such summary measures of policy cost-effectiveness is often difficult in practice. In many cases, researchers observe only the estimated causal effects and their standard errors but lack information about the correlations between them. These correlations are crucial for quantifying uncertainty: the variance of any scalar function that aggregates multiple effects depends not only on the precision of each individual estimate but also on how the estimates co-vary. If the underlying microdata were available, the covariances between causal effect estimates could be computed directly, allowing the variance of the function to be estimated zellner1962efficient. However, in many settings of practical interest, the microdata are unavailable for ex-post analysis, leaving researchers to rely only on published estimates and their standard errors. In this paper, we study the problem of inference on functions of multiple causal effects when the correlation structure across these causal effects is unknown.
To illustrate the challenge, we focus on the problem of conducting inference for the Marginal Value of Public Funds (MVPF) of a policy hendren2020unified. The MVPF is a widely used metric for evaluating the welfare consequences of government expenditure. It is defined as a non-linear function of multiple causal effects: the benefits a policy provides to its recipients are divided by the policy's net cost to the government. hendren2020unified construct MVPFs for more than one hundred policies using causal effects reported in existing studies. In most cases, only the point estimates and their standard errors are available to them, while the microdata underlying these estimates are inaccessible for such ex-post analysis. The challenge for inference is that the variance of the MVPF depends on the correlations across causal effects, which are not reported and not estimable.
We propose a simple inference procedure that delivers valid confidence intervals for functions of causal effects, even when the correlation structure across effects is unknown. The idea is straightforward: we ask what is the largest possible variance of the function given the available information and we identify the correlation structure under which this upper bound is attained. Using this worst-case variance, we construct conservative confidence intervals that guarantee valid coverage. We show that this conservative approach enables meaningful inference when computing the confidence intervals for MVPF estimates. Importantly, we formulate the problem of finding the variance upper bound as an optimization problem, which also makes it straightforward to incorporate other setting-specific information---for example, known independence between causal effects---to further tighten the confidence intervals and improve statistical precision.
Second, we show how confidence intervals can be tightened further still when the causal effects correspond to the impacts of a randomized treatment on multiple outcomes. In this setting, we characterize the off-diagonal entries of the covariance matrix and show that they take a particularly interpretable form, the sign of which may be known from prior studies, economic theory, or other data sources. Incorporating such information allows us to meaningfully increase statistical power to reject null hypotheses of interest.
Finally, we introduce a complementary approach to inference and ask a different question from worst-case inference. Instead of focusing on the largest possible variance given the available information, we ask how sensitive a policy-relevant conclusion is to uncertainty about the correlation structure. For example, a policymaker may wish to test whether a dollar spent on a policy provides beneficiaries with at least one dollar of benefits, i.e., $H_0: \text{MVPF} < 1$ against $H_1: \text{MVPF} \geq 1$. We introduce a “breakdown statistic” that quantifies how robust this conclusion is to different correlation structures: it measures the proportion of admissible correlation structures under which the null hypothesis would not be rejected. The statistic takes values between 0 and 1, where a value of 0 implies that we can conclude that the MVPF is greater than 1 under all plausible correlation structures, while a value of 1 implies that we cannot reject the null under any correlation structure. Unlike inference based on the worst-case variance, which guarantees valid coverage but might be conservative, the breakdown statistic facilitates comparisons of the robustness with which we can arrive at a policy conclusion---for example, whether the MVPF of a policy is greater than 1---across settings.
We illustrate our inference procedure by conducting inference on the MVPF for eight different policies. First, we show that meaningful inference is possible even in the absence of any microdata, using the upper bound of the variance alone. Second, hendren2020unified note that because the MVPF reflects the shadow price of redistribution, a welfare-maximizing government should have a positive willingness-to-pay to reduce the statistical uncertainty around the cost of redistribution. We demonstrate how this uncertainty can be reduced by leveraging setting-specific information about the sign of correlations across outcomes. In fact, our novel characterization of the covariance structure in randomized trials allows us to tighten MVPF confidence intervals beyond the worst-case by up to 30% in the policies we consider. Finally, we compute the breakdown statistic for the MVPF across multiple policies and illustrate how this metric can guide policymakers choosing among alternative policies.
Our work contributes to the literature on welfare analyses of government expenditure chetty2009sufficient,heckman2010rate,hendren2020unified. While existing tools provide a unified framework for evaluating the welfare consequences of government policies, statistical methods for conducting inference on welfare metrics under frequently encountered data limitations have been less developed. Our inference procedures strengthen the MVPF framework by providing a formal approach to quantifying statistical uncertainty in welfare metrics. hendren2020unified show that increasing spending on Policy A is welfare-improving by reducing spending on Policy B if and only if the MVPF of Policy A exceeds that of Policy B; the methods developed in this paper provide a valid test for the policy-relevant null hypothesis, $H_0: \text{MVPF}_A < \text{MVPF}_B$.
hendren2020unified propose a parametric bootstrap approach that constructs confidence intervals for the MVPF under a user-specified correlation structure. While such an approach can yield valid inference when the specified structure is indeed the worst case, misspecification may lead to confidence intervals that fail to achieve nominal coverage. Our method avoids this risk by formally identifying---rather than assuming---the correlation structure that maximizes the variance, solving an optimization problem that guarantees valid inference regardless of the true correlation structure. Moreover, our setup allows us to incorporate additional setting-specific information---for example, known independence across estimates or theory-driven sign restrictions---thereby increasing statistical power when such information is available.
Our methods might be applicable beyond the MVPF as well, in other settings when the correlation structure across causal effects might be difficult to obtain. First, researchers are frequently interested in functions of causal effects reported in existing publications, but the underlying microdata may be inaccessible. This can occur when the effects are estimated using privately held administrative data or when replication files are not publicly released. Replication data are missing for nearly half of all empirical papers published in the American Economic Review christensen2018transparency, underscoring how common this problem is. Second, even when the underlying data are technically available, computing the correlations can be prohibitively costly when the effects come from distinct datasets with common units but difficult-to-merge identifiers. For example, unique identifiers may be missing in historical decennial Census data ruggles2018historical, or the relevant data sources may be stored across separate federal agencies, as is the case when linking administrative tax data and administrative crime records for the full population rose2018effects.
Finally, cocci2023standard develop a closely related procedure in a different context: bounding the asymptotic variance of an estimate for a structural parameter when calibrating models to match empirical moments. Their work shows how to compute worst-case standard errors when the off-diagonal elements of the variance-covariance matrix are unknown, using only the variances of the empirical moments. We extend their framework in three important ways. First, we exploit the structure of randomized treatments to characterize the covariance matrix, which enables us to impose theory-motivated sign restrictions and potentially yields substantial power improvements. Second, our paper applies this variance-bounding approach to a new domain---inference on welfare metrics, such as the MVPF---where analysts frequently lack access to the underlying microdata. Third, we introduce a “breakdown” approach that quantifies the robustness of policy conclusions to uncertainty about the correlation structure.
Our starting point is a vector of estimated causal effects, denoted by $\bm{\hat{\beta}} \in \mathbb{R}^d$, that we seek to aggregate into a measure of the cost-effectiveness of a policy. We assume that $\bm{\hat{\beta}} $ asymptotically follows a joint Normal distribution with variance-covariance matrix $\bfV$. Since it is standard practice to report standard errors for individual estimates, we assume we have access to consistent estimates of the diagonal entries of $\bfV$. In contrast, the covariances between estimated causal effects---the off-diagonal entries of $\bfV$---are rarely reported. We focus on a setting where the underlying microdata are unavailable, so these off-diagonal entries cannot be directly estimated. We summarize the available information in Assumption (ref).
In our setting, $f(\bm{\hat{\beta}})$ represents an estimate of a policy's cost-effectiveness. We assume that the function aggregating the causal effects, $f: \mathbb{R}^d \to \mathbb{R}$, is continuously differentiable at $\bm\beta$, with $f'(\bm\beta)\neq0$. We place no additional restrictions on $f(\cdot)$ and allow it to be non-linear, since in practice its form depends on the economic mapping between the estimated causal effects and the cost-effectiveness measure for the policy being analyzed. This is summarized in Assumption (ref).
In summary, Assumption (ref) states that we observe consistent estimates of the causal effects and their standard errors, but lack reliable information about the covariances between them. Assumption (ref) requires that the function mapping estimated causal effects into the cost-effectiveness measure is smooth and well-behaved. We maintain Assumptions (ref) and (ref) throughout the paper.
Under these assumptions, we apply the delta method to obtain the asymptotic distribution of $f(\bm{\hat{\beta}})$: \[ \sqrt{n}\Big(f(\bm{\hat{\beta}}) - f(\bm{\beta})\Big) \xrightarrow{d} \mathcal{N}(0,\tau^2) \] where \[ \tau^2 = \sum_{i=1}^{d}\sum_{j=1}^{d} \sigma_{ij} \frac{\partial f(\bm{\beta})}{\partial \beta_i} \frac{\partial f(\bm{\beta})}{\partial \beta_j} = \left[ \sum_{i=1}^{d}\left(\sigma_i \frac{\partial f(\bm{\beta})}{\partial \beta_i}\right)^2 \right] + \Bigg[\mathop{\sum\limits_{i=1}^{d}\sum\limits_{j=1}^{d}}_{\{i,j : i \neq j\}} \sigma_{ij} \frac{\partial f(\bm{\beta})}{\partial \beta_i} \frac{\partial f(\bm{\beta})}{\partial \beta_j}\Bigg] \tag{2.1}\label{eqn:target-var} \] and $\sigma_{ij}$ denotes the covariance between $\beta_i$ and $\beta_j$.\footnote{ The delta method relies on a first-order (linear) approximation of the function $f(\cdot)$ around $\bm{\beta}$. Under the maintained assumptions, this approximation is valid asymptotically. However, in finite samples, if the variance of $\bm{\hat{\beta}}$ is large, $\bm{\hat{\beta}}$ may deviate from $\bm{\beta}$ with non-negligible probability, making the linear approximation less accurate.} We define \[ \rho_{ij} \equiv \frac{\sigma_{ij}}{\sigma_i \sigma_j}, \] as the correlation coefficient between $\beta_i$ and $\beta_j$.
The objective of this paper is to learn about $\tau^2$, the asymptotic variance of $f(\bm{\hat{\beta}})$. The central challenge is that $\tau^2$ depends on the covariances $\sigma_{ij}$, which are not estimable in our setting because the underlying microdata are unavailable. This raises the key question: what can be learned about $\tau^2$ when $\sigma_{ij}$ for $i \neq j$ cannot be estimated?
Since the asymptotic variance of $f(\bm{\hat{\beta}})$ depends on the correlation structure across the estimated causal effects---and this correlation structure cannot be estimated in the absence of microdata---we consider an alternative approach to inference on $f(\bm\beta)$. We ask: given the observed information, how large could the asymptotic variance of $f(\bm{\hat{\beta}})$ be? We then use an estimate of this variance upper bound to conduct valid hypothesis tests on $f(\bm\beta)$.
To motivate this approach and provide a rationale for focusing on the variance upper-bound, consider testing:
When the variance $\tau^2$ can be consistently estimated, standard $t$-tests control size and are uniformly most powerful. The difficulty arises when the correlations across effects ($\rho_{ij}$) are unknown and $\tau^2$ cannot be estimated. In this case, the problem can be framed as hypothesis testing with nuisance parameters, $\rho_{ij}$ for $i\neq j$. Finding the UMP test in this setting corresponds to identifying the least favorable distribution of the nuisance parameters romano2005testing.\footnote{A least favorable distribution is the distribution on the nuisance parameters under which the test performs the worst, or in other words, the distribution under which the probability of correctly rejecting a false null is smallest. If a test controls size and has good power even under this “worst-case” scenario, then it will perform at least as well under all other admissible distributions.} While least favorable distributions are often challenging to characterize elliott2015nearly, our setting is simplified by the fact that the nuisance parameters enter the distribution of $f(\hat{\bm\beta})$ only through its variance. Since power is minimized when variance is maximized, finding the least favorable distribution---and hence the uniformly most powerful test---corresponds to finding the correlation structure that maximizes $\tau^2$. This provides the statistical rationale for focusing on the variance upper bound.
In Section (ref), we consider the general setting where no additional structure is imposed on the correlation structure, so the variance upper bound is determined solely by the mathematical constraints of the correlation matrix. In Section (ref), we specialize to cases where treatment assignment is either completely randomized or randomized conditional on observed covariates. In this setting, we provide a novel characterization of the covariance structure that allows us to incorporate information from prior studies, economic theory, or other data sources to impose sign constraints on elements of the covariance matrix.
We can re-express Equation (ref) in terms of the correlations $\rho_{ij}$ as follows:
Since $\beta_i$ and $\sigma_i$ can be consistently estimated from the observed data, obtaining an upper bound for $\tau^2$ amounts to maximizing Equation (ref) with respect to $\rho_{ij}$ for $i \neq j$, subject to certain constraints. We formulate it as the following convex optimization problem:
Constraint (ref) requires the variance-covariance matrix to be positive semidefinite; Constraint (ref) enforces symmetry of the variance-covariance matrix; and Constraint (ref) ensures that all pairwise correlations lie within $[-1,1]$. Problem (ref) is therefore a well-defined semidefinite program (SDP) that can be solved using existing optimization tools grant2008graph,grant2014cvx. A key advantage of this formulation is its flexibility: we can incorporate available information about the correlations as additional constraints. For instance, in some cases, it may be known that two estimates are uncorrelated, such as when they are constructed using independent, non-overlapping samples. This information can be incorporated into the optimization problem by fixing the corresponding correlation to be zero. This flexibility is particularly important for the analysis in Section (ref), where we leverage our characterization of the off-diagonal entries of the covariance matrix to impose theory-motivated sign restrictions.
We denote the maximum variance obtained by solving (ref) as $\tau^2_{\max}$ and note the following. First, confidence intervals constructed using $\tau_{\max}$ will, by construction, have weakly higher coverage probability than those based on $\tau$. While this reduces power, it guarantees size control; coverage is exact only when $\tau_{\max} = \tau$. Second, in settings where estimating the covariance across estimates is feasible but costly, we recommend that researchers first test their hypotheses using $\tau_{\max}$. Rejecting a null hypothesis under the worst-case variance implies that the null would also be rejected using the true variance. This allows researchers to conduct valid inference while avoiding the costs of computing the full covariance matrix. Finally, in Appendix Section (ref), we compare our approach to that of cocci2023standard who propose a method for worst-case inference when matching structural parameters to empirical moments in overidentified settings. We show how tighter bounds can be found than what is implied by Lemma 1 in cocci2023standard, and illustrate through an example why maximizing the variance is challenging even in the simple case where there are no additional constraints beyond those in (ref).
In contrast to the generic case considered in Section (ref), this section leverages information about treatment assignment to derive more powerful tests on cost-effectiveness parameters. When treatment is randomized---either completely or conditional on observables---we obtain a novel, interpretable characterization of the covariance matrix. This characterization allows researchers to impose sign or independence restrictions on the correlation matrix, grounded in theory, prior evidence, or auxiliary data. Incorporating these restrictions can substantially sharpen variance bounds and deliver more precise inference on policy-relevant parameters.
Let $Y_{ij}(1)$ denote the treated potential outcome $j$ for unit $i$ and $Y_{ij}(0)$ denote the control potential outcome $j$ for unit $i$, where $j \in \{1,...,d\}$. Let $Z_i \in \{0,1\}$ indicate treatment assignment, where $Z_i = 1$ if unit $i$ is treated and $Z_i =0$ otherwise. We assume random assignment, meaning that treatment is independent of the full vector of potential outcomes:
The observed outcome is $Y_{ij} = Z_i\cdot Y_{ij}(1)+(1-Z_i)\cdot Y_{ij}(0)$. For each outcome $j$, the average treatment effect (ATE) is given by $\beta_j = \mathbb{E}[Y_{ij}(1) - Y_{ij}(0)]$, and we estimate it using the difference in sample means between the treated and control groups:
where $n_1$ and $n_0$ are the number of treated and control units, respectively. Let $\hat{\bm\beta} \in \mathbb{R}^d$ denote the vector of estimated treatment effects. In this setting, the asymptotic variance-covariance matrix $\bfV$ has a structure that allows for a simple and interpretable characterization, summarized in the following proposition.
The proof of Proposition (ref) is in Appendix Section (ref). The proposition shows that the asymptotic covariance between estimated treatment effects $\widehat{\beta}_{p}$ and $\widehat{\beta}_{q}$ depends only on the covariances of outcomes $Y_{ip}$ and $Y_{iq}$ within the treatment and control groups. In particular, if the outcomes are positively correlated within both groups, the treatment effects on those outcomes must also move in the same direction. For example, in the case of a randomized tax credit expansion called Paycheck Plus, the MVPF depends on the effects of the program on after-tax income, earnings, and labor force participation.\footnote{The Paycheck Plus program is studied in miller2017expanding and the MVPF for this program is computed in hendren2020unified.} Since individuals with higher earnings also tend to have weakly higher after-tax income and are weakly more likely to participate in the labor force, Proposition (ref) implies that the off-diagonal entries of $\bfV$ are non-negative. This information can be incorporated as additional constraints in the optimization problem (ref), thereby producing a (weakly) tighter upper bound. If it is known that all covariances between outcome pairs are non-negative as in the Paycheck Plus MVPF, we can solve the following optimization problem to obtain the variance upper bound:
In Section (ref), we show that adding Constraint (ref) to the optimization problem reduces the width of the Paycheck Plus MVPF confidence intervals by nearly 30%.
We also extend Proposition (ref) to settings where treatment assignment is random only conditional on covariates, such that $$ \left(Y_{i j}(1), Y_{i j}(0)\right) \perp Z_i \mid \mathbf{X}_i \quad \text{for all} \quad j = 1,\ldots,d $$ where $\mathbf{X}_i \in \mathbb{R}^k$ is a vector of observed covariates. A similar characterization of the asymptotic covariance under unconfoundedness is provided in Appendix Section (ref).
In Section (ref), we proposed a method that constructs worst-case confidence intervals for cost-effectiveness parameters, guaranteeing appropriate coverage regardless of the true correlation structure across causal effects. The strength of this method is its robustness: it delivers valid inference without making parametric assumptions about which correlation structures are more likely than others. The drawback is that tests based on the worst-case variance may leave policymakers underpowered to reject relevant null hypotheses. In this section, we develop a complementary approach to inference: we specify a set of plausible correlation structures, place a probability distribution over them, and ask: {how likely is it that a given hypothesis would be rejected if the true correlation structure were drawn from this set?}
We begin by specifying three elements. First, we fix the null hypothesis of interest. For example, in the case of the MVPF, a policymaker may wish to test whether one dollar of government spending generates more than one dollar of benefits for recipients, i.e., $H_0:\,\text{MVPF} < 1$ vs. $H_1:\,\text{MVPF} \geq 1$. Second, we specify the set of admissible correlation structures. For instance, Proposition (ref) may imply that correlations across the estimated causal effects are non-negative, so the admissible set is all correlation matrices with non-negative off-diagonal elements. Finally, we specify a probability distribution over this admissible set. For example, a policymaker might assume that all correlation structures in the admissible set are equally plausible. Alternatively, they might want to assume that correlation structures closer to independence are more plausible in their setting. We operationalize this by placing an LKJ prior lewandowski2009generating over the admissible set. The LKJ distribution has density \( \pi(\rho) \propto \det(\rho)^{\eta - 1}, \) where $\eta$ is the parameter governing which correlation structures are more likely than others. When $\eta = 1$, the prior is uniform over all correlation matrices. Larger values of $\eta$ place more mass near the identity matrix, favoring weaker correlations.
Next, we repeatedly draw from the specified distribution of correlation structures and test whether the null hypothesis is rejected under each draw. We define the breakdown statistic as the share of correlation structures under which we are unable to reject the null hypothesis. For policymakers, the breakdown statistic provides a transparent measure of how fragile a conclusion is to uncertainty about correlations. A breakdown statistic close to zero implies that the conclusion is robust to most correlation structures, whereas a value close to one indicates that the null hypothesis is unlikely to be rejected under any plausible correlation structure. We refer to this approach of assessing how easily a conclusion “breaks down” under alternative correlation structures as breakdown analysis.\footnote{See manski2018right,masten2020inference,diegert2022assessing,rambachan2023more,spini2021robustness for similar approaches.} Finally, note that if a null hypothesis can be rejected even under the worst-case correlation structure derived in Section (ref), the breakdown statistic must equal zero: by definition, all other admissible correlation structures imply a (weakly) smaller asymptotic variance than the worst case and therefore also lead to rejection of the null.
The exact algorithm for estimating the breakdown statistic is described in Appendix Section (ref); we provide a sketch of the algorithm here. We aim to assess the robustness of inference on $f({\bm\beta})$ to uncertainty about the asymptotic correlation structure of the estimated causal effects $\hat{\bm\beta} \in \mathbb{R}^d$. We define the robust region $\text{RR}_f$ as the set of admissible correlation matrices under which the null hypothesis $H_0: f(\bm\beta) < k$ is rejected at level $\alpha$: \[ \text{RR}_f = \left\{ \rho \in \mathcal{R} : f(\hat{\bm\beta}) - z_\alpha \cdot \tau(\rho) \geq k \right\}, \] where $\tau^2(\rho)$ denotes the asymptotic variance of $f(\hat{\bm\beta})$ under the correlation matrix $\rho$, $\mathcal{R}$ is the set of all admissible correlation matrices, and $z_\alpha$ is the $1-\alpha$ quantile of the standard normal distribution. We then define the breakdown statistic as the probability that the null is not rejected under an LKJ prior $\pi$ on $\rho$: \[ {\text{BR}}_f = 1 - \Pr_{\rho \sim \pi} \left[ \rho \in \text{RR}_f \right]. \] To estimate the Breakdown Statistic, we sample $\rho^{(1)}, \dots, \rho^{(N)}$ from the specified LKJ prior distribution $\pi$. For each draw, we compute the implied standard error $\tau^{(m)}$ and determine whether the null hypothesis is rejected. The estimated Breakdown Statistic is the proportion of draws under which we are unable to reject the null hypothesis of interest: \[ \widehat{\text{BR}}_f = 1 - \frac{1}{N} \sum_{m=1}^N R^{(m)}. \] In Section (ref), we describe how the Breakdown Statistic can help a policy-maker choose from a menu of policies to fund.
We illustrate our method by conducting inference on the Marginal Value of Public Funds (MVPF), a widely used metric for evaluating the welfare consequences of government policies. We first outline the MVPF framework and explain why our approach is particularly well suited for inference in this setting, before applying the tools developed in Sections (ref) and (ref) to construct valid confidence intervals for MVPFs across eight policies.
hendren2020unified popularized the Marginal Value of Public Funds (MVPF) as a unified metric for evaluating the “bang-for-the-buck” of public spending. An MVPF of 1 means that a policy delivers one dollar of benefits to recipients for each dollar of net government cost. Formally, the MVPF is defined as the benefits provided to recipients of a policy divided by the net cost borne by the government: \[ MVPF = \frac{Benefits}{Net\;Government\;Costs} = \frac{\Delta W}{\Delta E - \Delta C}, \] where $\Delta W$ denotes the estimated benefits to individuals, $\Delta E$ is the government's initial expenditure on the policy, and $\Delta C$ is the estimated reduction in government costs induced by the policy's causal effects.
Four features of the MVPF framework make our proposed method particularly well suited for valid inference. First, the MVPF is a non-linear function of multiple causal effects. To illustrate, consider the MVPF of the expanded Earned Income Tax Credit (EITC) program, Paycheck Plus. miller2017expanding estimate the causal effects of the program on several outcomes, including earnings, employment, and after-tax income. These estimates, reported in Table (ref), form the input vector:
From these causal effects, the MVPF for Paycheck Plus is constructed as
Second, in most applications, the only available information are the reported causal effect estimates and their standard errors. For example, the effects of Paycheck Plus are estimated using confidential administrative tax data, and the original study does not report the correlation structure across outcomes. Thus, the information available for inference is limited to the estimates and standard errors in Table (ref). To conduct inference in this setting, hendren2020unified assume a correlation structure across estimates. As we illustrate in Appendix Section (ref), relying on an assumed correlation structure can imply confidence intervals that are not guaranteed to have the correct coverage. Our method ensures valid inference without requiring the variance-covariance matrix to be assumed or consistently estimated.
Third, hendren2020unified show that reallocating spending from Policy B to Policy A is welfare-improving if and only if $MVPF_A > MVPF_B$. Testing the hypothesis $H_0: MVPF_A < MVPF_B$, then, is central to the policy choice problem.\footnote{Here, we assume that the beneficiaries of both policies receive equal welfare weights.} Our method provides a test for this hypothesis that controls size under any correlation structure.
Fourth, because the MVPF reflects the shadow price of redistribution, a welfare-maximizing government should, in principle, be willing to pay to reduce statistical uncertainty in its estimated cost of redistribution hendren2020unified. In Section (ref), we show how mild, setting-specific assumptions can be used to sharpen inference on the MVPF, offering a systematic way to reduce statistical uncertainty.
We apply our inference method to the MVPF of eight government policies spanning different domains of public expenditure: three job-training programs (Job Start, Work Advance, Year Up), two cash transfers (Paycheck Plus, Alaska Universal Basic Income), a health insurance expansion (Medicare Part D), childcare spending (foster care provision), and an unemployment insurance (UI) expansion.\footnote{The MVPFs for Job Start, Work Advance, Year Up, Paycheck Plus, and Alaska Universal Basic Income are computed in hendren2020unified. The MVPF for Medicare Part D is computed in wettstein2020retirement. The MVPF for foster care provision is computed in baron2022there. The MVPF for the UI expansion is computed in huang2021welfare.} The estimated MVPFs and 95% confidence intervals constructed by solving (ref) are shown in Figure (ref). Details of each policy and its MVPF calculation are provided in Appendix Section (ref).
Several lessons emerge from Figure (ref). First, even without assumptions on the off-diagonal entries of the variance-covariance matrix, we can reject the null that the MVPF of Job Start or Year Up exceeds one under any correlation structure, implying that a dollar spent on these job-training programs delivers less than a dollar in benefits. Second, using the variance upper bound, we test $H_0: MVPF_{\text{Alaska UBI}} < MVPF_{\text{Job Start}}$.\footnote{Since these MVPFs are based on independent samples, we assume they are uncorrelated.} We reject this null hypothesis, suggesting that reallocating spending from job-training programs to universal basic income programs could be welfare-enhancing. Finally, our estimates highlight meaningful statistical uncertainty in the relative ranking of some policies. For example, while the point estimates suggest that reallocating funds from job training to UI extensions is welfare-improving, our inference exercise shows that the uncertainty in these estimates precludes such a conclusion.
The only policy for which we have access to the underlying microdata is Medicare Part D. Using this data, we can recover the full variance-covariance matrix and compute exact confidence intervals, something that is infeasible for the other policies we study. Table (ref) compares three sets of confidence intervals for the estimated MVPF of Medicare Part D: exact intervals using the estimated correlation structure, intervals assuming all causal effects are uncorrelated, and worst-case intervals from Problem (ref). The exact confidence intervals rule out MVPF values below 0.80 and above 1.95, whereas the worst-case confidence intervals rule out values below 0.17 and above 2.57. These results show that our worst-case intervals remain informative even without microdata, but also highlight a key takeaway for practitioners: reporting the estimated covariance matrix across causal effects, when feasible, can substantially improve the precision of ex-post inference.
A policymaker choosing among policies may care about whether we can robustly conclude that a policy “pays for itself,” rather than focusing only on the statistical uncertainty surrounding the estimated returns to each policy. To answer this question, we use the Breakdown approach from Section (ref). The Breakdown Statistic summarizes robustness by reporting the share of admissible correlation structures under which the null hypothesis $H_0:\text{MVPF} < 1$ is not rejected. We compute this statistic in Table (ref) using a uniform prior over the space of correlation matrices.\footnote{Because the policy-relevant threshold is $MVPF > 1$, our Breakdown Analysis focuses on policies with point estimates above 1; if the point estimate of the MVPF is less than 1, the null can never be rejected and the Breakdown Statistic is mechanically equal to 1.} Comparing Medicare Part D, Foster Care Provision, and UI Extension, we find that the conclusion that the $MVPF \geq 1$ is most robust for Foster Care Provision, which has a Breakdown Statistic of 0.67. Put differently, a Breakdown Statistic of 0.67 means that the conclusion that foster-care provision pays for itself fails to hold under roughly two-thirds of admissible correlation structures. This illustrates the value of the Breakdown Statistic: it makes clear not only whether a policy appears cost-effective, but also how fragile that conclusion is to uncertainty about the correlation structure.
Finally, we turn to policies evaluated using randomized trials, the setting of interest in Section (ref). In these cases, Proposition (ref) provides an interpretable characterization of the covariance structure that allows us to impose sign restrictions on correlations across outcomes. For example, in the case of Paycheck Plus, it is plausible to assume that individuals with higher after-tax income also have higher earnings and are more likely to participate in the labor force. Incorporating such restrictions, we compute the MVPF confidence intervals by solving Problem (ref). Figure (ref) reports the resulting intervals for Job Start, Paycheck Plus, Work Advance, and Year Up. A key takeaway is that sign restrictions can meaningfully sharpen inference: for Paycheck Plus, the data allow us to rule out MVPF values below -0.38 and above 2.37, reducing the width of the confidence interval by nearly 30% relative to the worst-case bound. This demonstrates how even mild, theory-motivated assumptions can deliver substantially more informative inference for policy analysis.
\pagenumbering{arabic}