EconBase
← Back to paper

What's the Magic Formula Instrument?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

61,063 characters · 14 sections · 110 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

What's the Magic Formula Instrument?

\thispagestyle{empty}

\abstract{ Two recent papers by borusyakhull2023e, borusyakhull2026e propose using known formulas to adjust linear instrumental variable estimators for confounding covariates. Implementing this “formula instrument” approach requires making a parametric assumption on the distribution of the unobserved shocks that generated the instrument. We develop a method for systematically evaluating the sensitivity of formula instrument estimates to this parametric assumption. The method is straightforward to implement using our companion R package formulaiv. We use our method to reanalyze the applications in both borusyakhull2023e and borusyakhull2026e. In both applications, we find that a variety of estimates of different signs and magnitudes can be recovered by slightly changing the shock distribution. }

\onehalfspacing

Introduction

In two recent papers, borusyakhull2023e, borusyakhull2026e consider causal inference strategies based on “formula instruments,” where an instrumental variable (IV) is created by applying a known formula to other variables. Similar approaches have long been used for policy eligibility curriegruber1996tqjoe, tax liability grubersaez2002jope, and the ubiquitous bartik1991 (or shift-share) instruments blanchardkatz1992bpoea, goldsmith-pinkhamsorkinswift2020aera. borusyakhull2023e take this idea a step further by proposing that researchers specify the entire data generating process for the formula instrument, including the ex-ante distribution of counterfactual shocks that produces its exogenous variation.

In this paper, we develop a method for conducting sensitivity analysis to this assumed distribution of shocks. As borusyakhull2021a note,

quotation\em The key challenge of applying our framework, absent true randomization, is in specifying plausible shock counterfactuals.

This challenge raises a basic question: how sensitive are causal conclusions to the researcher's specification of the shock assignment process? The method we develop enables researchers to evaluate the sensitivity of their estimates to small or large deviations away from an assumed baseline distribution of shocks. The method can be reliably implemented at scale with linear programming techniques through our companion R package formulaiv.

We use our method to reanalyze the empirical applications in both borusyakhull2023e and borusyakhull2026e.

borusyakhull2023e analyze the effect of market access on employment using the roll-out of the high-speed rail system in China. The authors find that a naive OLS estimate yields large positive effects, while their formula instrument approach produces a small positive effect that is indistinguishable from zero. Our sensitivity analysis shows that small changes in the distribution of shocks used in their formula instrument lead to instrumental variable estimates that are anywhere from large negative effects to large positive effects. We show that the specification test proposed by borusyakhull2023e is unable to reject the null hypothesis that any of these alternative distributions are correctly specified.

borusyakhull2026e analyze the effect of the Medicaid expansions on the take-up of private insurance. The authors show that using a formula instrument allows one to tighten the precision on a simulated instrument approach freangrubersommers2017johe by focusing attention on the population that is potentially affected by the reform. However, doing so requires taking a stance on the distribution of Medicaid expansion shocks across states. The authors do so by assuming that the ex-ante probability of expansion only depends on the party of the governor, so that, for example, Republican-led states like Michigan (did expand) and Alabama (did not) had equal probabilities of expanding, while Democrat-led states like Delaware (did expand) and Missouri (did not) also had equal ex-ante probabilities of expanding. Our sensitivity analysis shows that changing the Republican-led probabilities to be non-homogeneous allows for formula instrument estimates that are consistent with a broad range of possible effects of both Medicaid eligibility and take-up. The implication is that the reduction in variance obtained by borusyakhull2026e comes with the risk of substantial bias from misspecification of the distribution of counterfactual shocks.

Our paper is relevant for a growing empirical literature that applies the borusyakhull2023e formula instrument method. Examples include dellolken2020res, bosshartweigand2025, buhlerdickens2025, and moroninicolettisalvanestominey2025. Our results suggest that formula instrument approaches can be sensitive to the parametric assumption about the shock distribution. Our method provides researchers an easy way to assess this sensitivity in their applications.

Our paper is also related to an old but growing theoretical literature on sensitivity analysis in statistics and econometrics. More recent examples include conleyhansenrossi2010roeas, nevorosen2012roeas, and klinesantos2013qe; see mastenpoirier2025wp for a survey with an emphasis on linear models. More related to our contribution is a smaller literature focused on sensitivity to parametric distributional assumptions in nonlinear models, for example chentamertorgovitsky2011cfdp1, bonhommeweidner2022q, christensenconnault2023e, and gurussell2024wp, although all of these authors consider settings much different than formula instruments.

The structure of the paper is as follows. In Section (ref), we explain the formula instrument approach and the recentered IV estimator that comes out of it. In Section (ref), we develop our method for sensitivity analysis. In Section (ref), we use our method to reanalyze the application to market access in borusyakhull2023e. In Section (ref), we use our method to reanalyze the application to Medicaid expansion in borusyakhull2026e. Section (ref) provides some brief concluding remarks.

Formula instruments and the recentered IV estimator

We briefly review the methodology developed by borusyakhull2023e.

The authors consider the linear model

align[align omitted — 83 chars of source]

where $i$ indexes the unit for $i = 1,\ldots,N$, $y_{i}$ is an outcome, $x_{i}$ is an endogenous treatment variable, and $\varepsilon_{i}$ is a latent residual. Both $y_{i}$ and $x_{i}$ are scalar and assumed to be sample mean zero for simplicity. The authors assume access to an instrumental variable $z_{i}$. Their methodology is also applicable to the case when $z_{i} = x_{i}$, which is the case they analyze in the application we revisit in Section (ref).

The authors assume that each $z_{i}$ is determined as a known function (or formula) of two types of observable variables: a vector of exogenous shocks, $g \equiv (g_{1},\ldots,g_{K})$, and a vector of predetermined covariates, $w_{i}$. To allow for spillovers, each $z_{i}$ can in general be determined by the collection of covariates $w \equiv (w_{1},\ldots,w_{N})$ from other units. The formula is a function $f_{i}$---possibly depending on $i$---that maps $g$ and $w$ to $z_{i}$:

align[align omitted — 67 chars of source]

The function $f_{i}$ is assumed to be known for all $i$. The shocks are assumed to be exogenous in the sense of being conditionally independent of the entire vector of latent residuals $\varepsilon \equiv (\varepsilon_{1},\ldots,\varepsilon_{N})$ borusyakhull2023e.

assumption(Shock Exogeneity) $g \independent \varepsilon \vert w$, where $\independent$ denotes independence.

Assumption (ref) implies that $z_{i}$ is independent of $\varepsilon_{i}$, conditional on $w_{i}$, but not unconditionally. As a consequence, Assumption (ref) is not sufficient to ensure that the linear IV estimator that uses $z_{i}$ as an instrument for $x_{i}$ will be consistent for $\beta$. To see this, write (ref) as

align[align omitted — 317 chars of source]

where the second equality invokes Assumption (ref) and defines $\eta_{i} \equiv \varepsilon_{i} - \Exp[\varepsilon_{i} \vert g, w]$. The new residual, $\eta_{i}$, satisfies $\Exp[\eta_{i} \vert z] = 0$ because of Assumption (ref) and the formula relationship (ref):

align[align omitted — 242 chars of source]

However, if $\Exp[\varepsilon_{i} \vert w]$ is a non-constant function of $w$, then $z_{i}$, which is also a function of $w$ via the formula (ref), will generally be correlated with the original residual, $\varepsilon_{i}$.

The traditional solution to this problem is to control for $w$. In the context of (ref), this means specifying a functional form for $\Exp[\varepsilon_{i} \vert w]$. For example, if $\Exp[\varepsilon_{i} \vert w] = w_{i}'\alpha$, then both $\beta$ and $\alpha$ can be consistently estimated by the linear IV estimator that uses $z_{i}$ as an instrument for $x_{i}$ while controlling for $w_{i}$, assuming sufficient independent variation in $z_{i}$. The motivation for using the borusyakhull2023e approach of leveraging the formula (ref) is that specifying the correct functional form for $\Exp[\varepsilon_{i} \vert w]$ may be difficult, especially when $w$ is a complex set of controls. borusyakhull2023e argue in the context of three empirical examples that it would be challenging to choose the correct functional form to control for $w$.

The alternative proposed by borusyakhull2023e is to instead model the conditional expectation of the instrument, $\mu_{i} \equiv \Exp[z_{i} \vert w]$. If this conditional expectation were known, then the instrument could be recentered as $\tilde{z}_{i} \equiv z_{i} - \mu_{i}$. While Assumption (ref) is not sufficient to ensure that the original instrument, $z_{i}$, is uncorrelated with $\varepsilon_{i}$, it is sufficient to ensure that the recentered instrument, $\tilde{z}_{i}$, is uncorrelated with $\varepsilon_{i}$:

align[align omitted — 283 chars of source]

This suggests using the linear IV estimator that instruments for $x_{i}$ with $\tilde{z}_{i}$ instead of $z_{i}$, which borusyakhull2023e describe as the “recentered IV” (RIV):

align[align omitted — 144 chars of source]

Under the usual statistical conditions, $\hat{\beta}_{\textsc{riv}}$ will be a consistent estimator of $\beta$. Earlier examples of this argument can be found in the literature on partially linear models, notably robinson1988e, ideas from which feature prominently in the modern literature on using machine learning to control for covariates in IV regressions chernozhukovchetverikovdemirerduflohansenetal2018ej, and have also been used in the literature on marginal treatment effects carneiroheckmanvytlacil2011aer, andresen2018tsj.

The benefit of using $\hat{\beta}_{\textsc{riv}}$ is that there is no need to specify the functional form of $\Exp[\varepsilon_{i} \vert w]$. The appeal of recentering the instrument turns on the relative difficulty of modeling $\mu_{i} \equiv \Exp[z_{i} \vert w]$ and $\Exp[\varepsilon_{i} \vert w]$. Both are potentially complicated functions when $w$ is a complex vector of covariates. The novel proposal of borusyakhull2023e is that one can model $\mu_{i}$ by combining the formula (ref) with the assumption that the conditional distribution of the shocks $g$, denoted $G(\cdot \vert w)$, is known by the researcher. This requires maintaining the following assumption borusyakhull2023e, which the authors describe as a “Known Assignment Process”.

assumption(Known Assignment Process) The distribution of $g$ conditional on $w$ is known and given by $G(g \vert w)$ for all supported $w$.

Assumption (ref) and the formula (ref) enable direct computation of $\mu_{i}$ through simulation. For example, borusyakhull2023e suggest choosing $G(\cdot \vert w) = G(\cdot)$ to be the uniform distribution over the set of all permutations of the observed realization of $g \equiv (g_{1},\ldots,g_{K})$, independently of $w$. There are $K!$ permutations of the $K$ elements of $g$, so this suggestion implies the assumption that $G(\cdot \vert w)$ places equal mass $1/K!$ on each permutation formed from the components of the realized $g$. When $K!$ is a large number, the authors propose approximating $\mu_{i}$ with a subset of $S$ permutations. With $\mu_{i}$ (or a sufficient approximation) in hand, the recentered IV estimator $\hat{\beta}_{\textsc{riv}}$ in (ref) can then be constructed by using $\tilde{z}_{i} \equiv z_{i} - \mu_{i}$ as an instrument for $x_{i}$, without controlling for covariates.

Sensitivity to the known assignment process

In this section, we develop a systematic sensitivity analysis that relaxes Assumption (ref).

Our object of interest is the joint distribution $G(\cdot \vert w)$ of the shock vector $g \equiv (g_{1},\ldots,g_{K})$. We assume for simplicity that $G(\cdot \vert w) = G(\cdot)$ does not depend on $w$, since this is the case in both of the applications we consider; however, this is not essential to what follows. We represent $G$ through a finite support $\{(g_{1s},\ldots,g_{Ks})\}_{s=1}^{S}$ of shock realizations, together with a vector of probabilities $p \equiv (p_{1},\ldots,p_{S})$ assigned to them, where $p_{s} \equiv \Prob_{G}[g = (g_{1s},\ldots,g_{Ks})]$.\footnote{ Our analysis can be extended to cases where $G$ has a continuous distribution; see Appendix (ref). } Then

align[align omitted — 182 chars of source]

where $f_{is} \equiv f_{i}((g_{1s},\ldots,g_{Ks}), w)$. The vector $p$ must live in the $S$-dimensional simplex of non-negative numbers that sum to one, which we denote by $\Delta^{S}$.

We consider sensitivity of the recentered IV estimate to the choice of $p$ as it deviates from some baseline $\bar{p}$ across some pre-determined set $\mathcal{P} \subseteq \Delta^{S}$. For example, $\bar{p}$ might be the uniform distribution used by borusyakhull2023e, which has $\bar{p}_{s} = 1/S$ for all $s$. We consider two ways of specifying the sensitivity set $\mathcal{P}$, intended to capture different ways of measuring deviations between $p$ and $\bar{p}$.

The first way is to require each component of $p$ to be within $\kappa \geq 1$ multiples of its corresponding component of $\bar{p}$ by restricting $p$ to the set

align[align omitted — 286 chars of source]

Setting $\kappa = 1$ makes $\mathcal{P}_{\textsc{j}}(1 \vert \bar{p}) = \{\bar{p}\}$ a singleton, while as $\kappa \rightarrow \infty$, the set $\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p})$ becomes closer to the simplex, $\Delta^{S}$. We call $\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p})$ the joint sensitivity set because it measures deviations from $\bar{p}$ in terms of the entire joint distribution $p$. This measure is used in the robust Bayes literature where $\bar{p}$ is viewed as a baseline prior lavine1991jotasa, wassermankadane1992jotasa.

The second way is to constrain the marginal distributions of each $g_{k}$ rather than the entire joint distribution. For a joint distribution $p \in \Delta^{S}$, the implied marginal probability that $g_{k} = h$ is

align[align omitted — 119 chars of source]

Let $\bar{q} = (\bar{q}_{1},\ldots,\bar{q}_{K}) \equiv (q_{1}(\cdot \vert \bar{p}),\ldots,q_{K}(\cdot \vert \bar{p}))$ denote the baseline collection of marginals produced from the baseline joint distribution $\bar{p}$ via (ref). We define the marginal sensitivity set to be the set of joint distributions whose implied marginals are within $\delta \geq 1$ multiples of $\bar{q}$,

align*[align* omitted — 304 chars of source]

For example, in the application in Section (ref), each shock $g_{k}$ is a binary event, so setting $\mathcal{H}_{k} = \{1\}$ for each $k$ makes $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q})$ the set of $p$ whose event probabilities for each shock $k$ are within $\delta$ multiples of the baseline event probabilities $\bar{q}_{k}(1)$. Note that unlike the joint sensitivity set, the marginal sensitivity set does not collapse to a singleton at $\delta = 1$, because many joint distributions can share the same marginals.

Each choice of $p \in \mathcal{P}$ produces a different recentered IV estimator by changing $\mu_{i}$ in (ref). We denote this dependence by writing $\mu_{i}(p)$. The recentered IV using $p$ is then $\tilde{z}_{i}(p) \equiv z_{i} - \mu_{i}(p)$ and the recentered IV estimator is

align[align omitted — 198 chars of source]

noting again that $y_{i}$ and $x_{i}$ are assumed to have sample mean zero for simplicity. The recentered IV estimator varies as $p$ ranges across a sensitivity set $\mathcal{P}$, such as $\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p})$ or $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q})$. To handle the possibility that $\hat{\beta}_{\textsc{riv}}(p)$ is undefined because the denominator $D(p) \equiv\sum_{i=1}^{N}x_{i}\tilde{z}_{i}(p)$ of $\hat{\beta}_{\textsc{riv}}(p)$ is zero, we define the set

align*[align* omitted — 117 chars of source]

Then the smallest and largest values that $\hat{\beta}_{\textsc{riv}}(p)$ can take are

align[align omitted — 322 chars of source]

The following proposition shows that these extremal values can be found by solving linear programs as long as $\mathcal{P}$ is a polyhedron, such as $\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p})$ or $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q})$.

propositionSuppose that $\mathcal{P} \subseteq \Delta^{S}$ is a polyhedron, written as $\mathcal{P} = \{p \in \Delta^{S} : Ap \leq c\}$ for some known matrix $A$ and vector $c$. If $D(p) \geq 0$ for all $p \in \mathcal{P}$ then \begin{align} \beta_{riv}(\mathcal{P}) = \min_{ \phi \in \re^{S}, \tau \in \re } \quad &\sum_{i=1}^{N} y_{i}z_{i}\tau - \sum_{i=1}^{N} \sum_{s=1}^{S} y_{i}f_{is}\phi_{s} \notag \\ s.t. \quad &\tau \geq 0, \phi_{s} \geq 0 for all $s = 1,\ldots,S$ \notag \\ &\sum_{s=1}^{S} \phi_{s} = \tau \notag \\ &A\phi \leq c\tau \notag \\ \quad &\sum_{i=1}^{N}x_{i}z_{i}\tau - \sum_{i=1}^{N} \sum_{s=1}^{S} x_{i}f_{is}\phi_{s} = 1, \end{align} and $\overline{\beta}_{\textsc{riv}}(\mathcal{P})$ is given by the corresponding maximization problem.\footnote{ We use the usual convention here of setting $\underline{\beta}_{\textsc{riv}}(\mathcal{P}) = -\infty$ if the minimization problem is unbounded and $\overline{\beta}_{\textsc{riv}}(\mathcal{P}) = +\infty$ if the maximization problem is unbounded. } Moreover, for any real number $b \in [\underline{\beta}_{\textsc{riv}}(\mathcal{P}), \overline{\beta}_{\textsc{riv}}(\mathcal{P})]$, there exists a $p \in \mathcal{P}$ such that $\hat{\beta}_{\textsc{riv}}(p) = b$. If instead $D(p) \leq 0$ for all $p \in \mathcal{P}$ then the same statement is true after two changes: (i) change the last constraint in (ref) from $1$ to $-1$ and (ii) take $-\overline{\beta}_{\textsc{riv}}(\mathcal{P})$ to be the optimal value of the minimization problem and take $-\underline{\beta}_{\textsc{riv}}(\mathcal{P})$ to be the optimal value of the corresponding maximization problem. If $D(p)$ takes both positive and negative values as $p$ ranges over $\mathcal{P}$, and if it is not the case that $\hat{\beta}_{\textsc{riv}}(p)$ is constant for all $p \in \mathcal{P}_{D \neq 0}$, then $\underline{\beta}_{\textsc{riv}}(\mathcal{P}) = -\infty$ and $\overline{\beta}_{\textsc{riv}}(\mathcal{P}) = +\infty$.

Proposition (ref) provides a computationally tractable way to compute all of the possible values that the recentered IV estimator $\hat{\beta}_{\textsc{riv}}(p)$ can take as $p$ varies across the sensitivity set $\mathcal{P}$. The linear programs have $S + 1$ variables and a similar number of constraints if $\mathcal{P}$ is taken to be the joint or marginal sensitivity set. This makes the programs straightforward to solve even if $S$ is quite large. The justification of Proposition (ref) recognizes that $\underline{\beta}_{\textsc{riv}}(\mathcal{P})$ and $\overline{\beta}_{\textsc{riv}}(\mathcal{P})$ are the optimal values of linear fractional programs because $\hat{\beta}_{\textsc{riv}}(p)$ is the ratio of two affine functions of $p$. Applying the charnescooper1962nrlq transformation to (ref) yields the linear program (ref); see, for example, boydvandenberghe2004.

Reevaluating the effects of market access in China

In this section, we use Proposition (ref) to reanalyze the application in borusyakhull2023e.

Replication

borusyakhull2023e use a two-period panel of 275 subprovince-level administrative divisions (“prefectures”) in mainland China. Regional employment in prefecture $i$ is defined as urban employment as taken from the Chinese City Statistical Yearbooks. Market access in prefecture $i$, year $t$, is defined as

align[align omitted — 173 chars of source]

where $\textsc{pop}_{j, 2000}$ is the population of prefecture $j$ in $2000$, and $\tau_{ijt}$ is the predicted travel time between prefectures $i$ and $j$ in year $t$.

The travel time $\tau_{ijt}$ is determined in part by the presence of high-speed rail (HSR) connections between the prefectures. borusyakhull2023e compute $\tau_{ijt}$ using comprehensive data on the evolution of the Chinese HSR network. The network includes 150 potential total lines: 83 lines that opened between 2007 and 2016, 66 additional lines that were planned or under construction by April 2019, but had not yet opened by the end of 2016, as well as one line between Qinhuangdao and Shenyang that opened in 2003.

This roll-out of HSR lines produces variation in MA over time. The authors define their endogenous variable $x_{i}$ as this change over the course of their two-period panel: $x_{i} \equiv \log \textsc{ma}_{i,2016} - \log \textsc{ma}_{i,2007}$. They take the outcome $y_{i}$ to be the corresponding change in urban employment between 2007 and 2016. The empirical challenge is to determine the causal effect of $x_{i}$ on $y_{i}$. As borusyakhull2023e discuss, this is difficult because $x_{i}$ is correlated with geography $w_{i}$, which may be correlated with unobserved determinants of employment growth, such as local productivity shocks.

borusyakhull2023e apply the recentered IV approach to this problem. The shock sequence $g \equiv (g_{1},\ldots,g_{K})$ is a vector of $K = 150$ binary shocks for each HSR line $k$, with $g_{k} = 1$ denoting that a line opened by 2016 and $g_{k} = 0$ denoting that it did not open. Assumption (ref) requires these line openings to be independent of unobserved determinants of employment growth, perhaps conditional on geographic controls $w_{i}$. To operationalize Assumption (ref), borusyakhull2023e assume that $G$ is a uniform distribution over a fixed support of $S = 1999$ draws of $g$.\footnote{ The authors need to do this because the formula $f$ implied by the market access function (ref) is non-separable across $g_{k}$ through their interdependence in $\tau_{ijt}$. In Section (ref), we consider an example where the support of $g$ is not constrained in this way. } For each draw, $(g_{1s},\ldots,g_{Ks})$, they recompute the travel time variable $\tau_{ijt}$, then construct $\mu_{i} \equiv \Exp[x_{i} \vert w_{i}] = S^{-1}\sum_{s=1}^{S}f_{i}((g_{1s},\ldots,g_{Ks}), w)$ using the formula for $x_{i} \equiv \log \textsc{ma}_{i,2016} - \log \textsc{ma}_{i,2007}$ implied by (ref). Note that this application has $z_{i} = x_{i}$, which is a special case of the formula IV framework that might be more appropriately called “formula OLS.”

We are able to replicate the results in borusyakhull2023e exactly by using the same sample of $S$ permuted $g$ vectors, which the authors included in their replication package. We briefly review these results, which are the same as in Table I of borusyakhull2023e. An unadjusted OLS estimate of $y_{i}$ on $x_{i}$ yields a statistically significant estimate of $.232$, which would be interpreted as an elasticity of employment with respect to the market access measure. Controlling for geographic measures lowers this to $.133$, which is still statistically significant (standard error $.064$). By contrast, the authors' recentered IV estimate with no covariates produces a statistically insignificant point estimate of $.084$ with a similar standard error of $.097$. Controlling for covariates lowers the recentered IV estimate to $.056$, with a standard error of $.089$.

Sensitivity to the assumed assignment process

In the notation of Section (ref), the shock distribution $\bar{p}_{\textsc{bh}}$ used by borusyakhull2023e amounts to setting the probability of each of the $s=1,\ldots,1999 \equiv S$ drawn simulations to be $\bar{p}_{\textsc{bh},s} \equiv 1/1999 \approx .0005$. Their reported estimate is $\hat{\beta}_{\textsc{riv}}(\bar{p}_{\textsc{bh}})$. Figure (ref) shows how sensitive $\hat{\beta}_{\textsc{riv}}(\bar{p}_{\textsc{bh}})$ is to this choice of $\bar{p}_{\textsc{bh}}$, with sensitivity measured in terms of the joint sensitivity set $\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p}_{\textsc{bh}})$ and its parameter $\kappa$. For example, a value of $\kappa = 5$ on the x-axis allows for a distribution of counterfactual network configurations with $p_{s}$ between $[1/(5 \times 1999), 5/1999] \approx [.0001, .0025]$ for each $s$, while still requiring $\sum_{s=1}^{S}p_{s} = 1$. The y-axis of Figure (ref) shows the set $[\underline{\beta}_{\textsc{riv}}(\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p}_{\textsc{bh}})), \overline{\beta}_{\textsc{riv}}(\mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p}_{\textsc{bh}}))]$, which contains all values of $\hat{\beta}_{\textsc{riv}}(p)$ that one could obtain for a $p \in \mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p}_{\textsc{bh}})$.

\fig {figs/china/est_joint_no_controls.pdf} {fig:est_joint_no_controls} {Sensitivity to assumed joint distribution in borusyakhull2023e} { Bounds from solving (ref) with $\mathcal{P} = \mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p}_{\textsc{bh}})$ for different values of $\kappa$. The horizontal lines show the OLS and recentered IV estimates from borusyakhull2023e. Spatially-clustered conley1999joe standard errors are shown in parentheses, following the same specification as in borusyakhull2023e. For the bounds, these show the standard errors at the optimizer for the program. }

At $\kappa = 1$, the bounds collapse to the baseline estimate reported by borusyakhull2023e. As $\kappa$ increases, the bounds widen, reflecting ambiguity in the specification of Assumption (ref). For example, with $\kappa = 5$, the set of recentered IV estimates one can obtain includes everything from substantial negative employment effects of about $-.100$ to substantial positive employment effects that are about equal to the unadjusted OLS estimate of $.232$. The implication is that changes in the assumed shock distribution (Assumption (ref)) can lead the recentered IV estimate to be as potentially misleading about positive employment effects as the uncontrolled OLS estimate, while also leaving open the possibility of negative employment effects. Figure (ref) shows that controlling for geographic covariates leads to similar conclusions.

Is $\kappa = 5$ large or small? The baseline choice of $\bar{p}_{\textsc{bh},s} = 1/1999 \approx .0005$ made by borusyakhull2023e requires each of the $1999$ shocks to have an equal probability that is small, with no single shock realization occurring in more than $.05\%$ of potential draws of the underlying data generating process. Setting $\kappa = 5$ means that none of the $1999$ shock configurations can occur in more than $.25\%$ or less than $.01\%$ of these draws. It is not clear how one could reason about the exact magnitude of so many small probabilities simultaneously, suggesting that $\kappa = 5$ is rather small compared to the baseline of $\kappa = 1$. As Figure (ref) shows, increasing $\kappa$ to $10$ leads to even greater ambiguity, while still imposing the mild restriction that no possible shock realization occurs in more than $.5\%$ of draws.

\fig {figs/china/est_marginal_no_controls.pdf} {fig:est_marginal_no_controls} {Sensitivity to assumed marginal distribution in borusyakhull2023e} { Bounds from solving (ref) with $\mathcal{P} = \mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{\textsc{bh}})$ for different values of $\delta$. See notes for Figure (ref). }

Figure (ref) shows sensitivity measured across the marginal set $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{\textsc{bh}})$ with $\mathcal{H}_{k} = \{1\}$ for each $k$. The probability of each shock $g_{k}$ being equal to one represents the probability that HSR line $k$ opened by 2016. The joint distribution $\bar{p}_{\textsc{bh}}$ used by borusyakhull2023e implies marginal probabilities $\bar{q}_{\textsc{bh}}$ that have most line opening probabilities between roughly $.4$ and $.6$, with a few lines pegged to an opening probability of one. For a line with a probability of $.5$, setting $\delta = 1.25$ means that $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{\textsc{bh}})$ contains joint distributions $p$ that admit marginal (line opening) probabilities between $.4$ and $.625$. As Figure (ref) shows, even this mild relaxation is consistent with a broad range of recentered IV estimates that produce anything from large negative to large positive estimates.

Specification tests

borusyakhull2023e suggest that Assumption (ref) can be tested using randomization inference with test statistic equal to the sample covariance between the recentered instrument and the implied residual. They conduct this test for their market access application and report a p-value of $.711$, failing to reject the null that the known assignment process is correctly specified. They interpret this result as “validating” their specification of the HSR assignment process borusyakhull2023e.

\fig {figs/china/false_no_controls.pdf} {fig:false_no_controls} {P-values from the borusyakhull2023e specification test} { P-values from the randomization inference test proposed by borusyakhull2023e. The left-hand facet shows results with $\mathcal{P} = \mathcal{P}_{\textsc{j}}(\kappa \vert \bar{p}_{\textsc{bh}})$ and the right-hand facet shows results with $\mathcal{P} = \mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{\textsc{bh}})$. The dotted line is the same p-value reported by borusyakhull2023e for their baseline specification. }

Figure (ref) shows p-values from the same test conducted for assignment processes that yield the lower and upper bounds for each $\kappa$ and $\delta$ considered in Figures (ref) and (ref). The p-values do not cross even a conservative conventional threshold such as $.10$ for any value of $\kappa$ or $\delta$: the specification test never rejects. This is despite the fact that we know that the assignment processes at the lower and upper bounds and at different values of $\kappa/\delta$ are inconsistent with one another. The implication is that the test proposed by borusyakhull2023e has low power for detecting violations of Assumption (ref).

Alternative baseline distributions

The results in Figures (ref)--(ref) show that recentered IV estimates are sensitive to deviations from the baseline shock distribution $\bar{p}_{\textsc{bh}}$ chosen by borusyakhull2023e in a way that is not detectable through their specification test. In this section, we examine whether $\bar{p}_{\textsc{bh}}$ is a sensible starting point.

\fig {figs/china/histogram_ml.pdf} {fig:histogram_ml} {Marginal probability of a line opening} { Each facet shows the out-of-sample histogram of marginal probabilities of opening across the 150 HSR lines for a machine learning model compared to the marginal probabilities $\bar{q}_{\textsc{bh}}$ generated by the uniform shock distribution $\bar{p}_{\textsc{bh}}$ used by borusyakhull2023e. The specification and training of the models is discussed in Appendix (ref). }

While the choice of $\bar{p}_{\textsc{bh}}$ specifies only one joint probability over the $S = 1999$ possible shock realizations, it implies $K = 150$ marginal probabilities for each of the HSR lines in the data. This suggests a data-driven exercise: for each HSR line, we train machine learning algorithms that use the predetermined characteristics of the line in 2007 to predict whether the line would be opened by 2016. We fit four learners: a random forest, penalized logistic regression, gradient-boosted trees, and $k$-nearest neighbors; Appendix (ref) contains details on how we specified and trained them.

Figure (ref) shows that---unsurprisingly---each of these learners provides better out-of-sample predictions than the implicit prediction $\bar{q}_{\textsc{bh}}$ generated by the uniform shock distribution $\bar{p}_{\textsc{bh}}$ used by borusyakhull2023e. Figure (ref) compares the histograms of line openings for the four models to $\bar{q}_{\textsc{bh}}$. Whereas $\bar{q}_{\textsc{bh}}$ has many line opening probabilities concentrated around $.4$ and $.6$, the learners recognize that some lines were considerably more or less likely to open for reasons that could be predicted from their predetermined characteristics. This provides additional evidence against the suggestion that the shocks should be viewed as “exchangeable,” a condition which borusyakhull2023e appeal to as a sufficient condition to support their choice of the uniform distribution $\bar{p}_{\textsc{bh}}$.

\fig {figs/china/est_marginal_ml_no_controls.pdf} {fig:est_marginal_ml_no_controls} {Sensitivity to machine learning marginal distributions} { Bounds from solving (ref) with $\mathcal{P} = \mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{m})$ for different values of $\delta$ and the $\bar{q}_{1},\ldots,\bar{q}_{4}$ generated by the four learners shown in Figure (ref). }

Figure (ref) reports a sensitivity analysis comparable to Figure (ref) when $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{m})$ is specified relative to the marginal distributions produced by the four machine learning models, $\bar{q}_{1},\ldots,\bar{q}_{4}$. We start the x-axis for each model at the first value of $\delta$ for which it's possible to find any valid probability distribution $p \in \mathcal{P}_{\textsc{m}}(\delta \vert \bar{q}_{m})$ that rationalizes $\bar{q}_{m}$. For the best-performing model, this requires taking $\delta$ past two, implying a relaxation of $[.25, 1.00]$ for a line-opening probability of $.5$. This suggests that the support of $S = 1999$ counterfactual shocks used by borusyakhull2023e is itself hard to rationalize with the data. Even so, deviations around each of the baselines provided by the machine learning models show the same type of sensitivity as for the baseline used by borusyakhull2023e. This suggests that the sensitivity found in Figures (ref) and (ref) is a consequence of the formula IV idea itself, rather than the specific choice of baseline reference distribution.

Reevaluating the impacts of Medicaid expansion

borusyakhull2026e apply the formula instrument idea to evaluate the impact of Medicaid eligibility on private insurance take-up using the partial state-level expansion of Medicaid that occurred under the Affordable Care Act (ACA) in 2014. The authors use a repeated cross-section of individuals $i$ from the American Community Survey (ACS). The endogenous variable $x_{i} \in \{0,1\}$ is individual $i$'s eligibility for Medicaid and the outcome $y_{i}$ is a measure of private insurance take-up. The authors propose a linear model of the form

align[align omitted — 123 chars of source]

where $s(i)$ and $t(i)$ denote the state and year of individual $i$, $\alpha_{s(i)}$ are state fixed effects, $r_{k}$ is an indicator for whether the government of state $k$ in 2013 is a Republican, $\tau_{r, t}$ are party-by-year fixed effects, and $\varepsilon_{i}$ is an unobservable. The concern is that $x_{i}$ and $\varepsilon_{i}$ may be correlated through the individual characteristics $c_{i}$.

The authors use the binary expansion decisions $g_{k} \in \{0,1\}$ for each state $k$ as the “shocks.” One way to do this is to instrument for $x_{i}$ using $z_{i} \equiv g_{s(i)}\IndicSmall{t(i) = \text{2014}}$, which produces an instrumented difference-in-differences estimate of $\beta$. The authors describe this as a simulated instrument along the lines of curriegruber1996tqjoe or freangrubersommers2017johe.\footnote{ The simulated instrument terminology may be misleading here because $z_{i}$ is binary, so lacks any variation intensity across states. We are following borusyakhull2023e in our usage of the phrase. } A drawback of this approach is that many individuals have no variation in $x_{i}$ regardless of the value of $z_{i}$, for example if they are ineligible for Medicaid either with or without the expansion. This dilutes the relevance of $z_{i}$ for $x_{i}$, making estimates of $\beta$ relatively imprecise.

borusyakhull2026e propose a formula instrument alternative based on knowledge of how Medicaid eligibility is determined:

align[align omitted — 129 chars of source]

where $h^{t(i)}$ is a known, year-specific function that determines Medicaid eligibility, $c_{i}$ are individual characteristics such as income, work status, or parental status, $e_{k}^{\text{2013}}$ is the Medicaid eligibility policy of state $k$ in 2013, and $e_{k}^{\Delta}$ includes other changes in 2014 to Medicaid coverage in state $k$. They propose recentering the instrument $z_{i} \equiv h^{t(i)}(c_{i}, e_{s(i)}^{\text{2013}}, g_{s(i)}, \emptyset)$ that ignores the non-ACA eligibility changes $e_{k}^{\Delta}$. Recentering this instrument via Assumption (ref) is necessary for it to be exogenous because $z_{i}$ depends on individual characteristics $c_{i}$ that are likely also reflected in $\varepsilon_{i}$.

The model that the authors propose for Assumption (ref) is based on the assumption that

align[align omitted — 187 chars of source]

where $w_{i}$ collects $c_{i}, s(i), t(i), e_{s(i)}^{\text{2013}}$, and $r_{s(i)}$. That is, the probability that state $k$ expands is a constant function $\pi(r_{k})$ of whether its governor in 2013 was a Republican, $r_{k}$. This implies that, for example, two Republican-led states like Michigan and Alabama---one of which expanded and one of which did not---had ex-ante equal probabilities of adopting the ACA expansion. Given (ref), the conditional expectation of $z_{i}$ given $w_{i}$ is

align[align omitted — 286 chars of source]

where $a_{i}$ is a binary indicator for whether individual $i$'s eligibility would have been affected by an expansion in 2014:

align*[align* omitted — 165 chars of source]

From (ref) we get that the recentered instrument $\tilde{z}_{i} \equiv z_{i} - \mu_{i}$ is

align*[align* omitted — 90 chars of source]

Unlike the simulated instrument, this recentered instrument $\tilde{z}_{i}$ is mechanically zero for individuals with $a_{i} = 0$, whose eligibility would have been unaffected by an expansion in their state. Using $\tilde{z}_{i}$ as an instrument therefore numerically drops these individuals, raising the hope that the resulting recentered IV estimator may be more precise than the simulated IV estimator. As a practical matter, it also means that the recentered IV estimator using $\tilde{z}_{i}$ is numerically equivalent to an IV estimator that uses $g_{s(i)} - \pi(r_{s(i)})$ as an instrument for $x_{i}$ among the subsample of affected individuals $a_{i} = 1$. Because the authors already include party-by-year fixed effects $\tau_{r_{s(i)}, t(i)}$ in (ref), this in turn is equivalent to just using $g_{s(i)}\IndicSmall{t(i) = \text{2014}}$ as an instrument for $x_{i}$, the same as in the simulated instrument but now only among the subsample with $a_{i} = 1$.

The key to this equivalence is (ref), which requires all Republican-led states to have had the same ex-ante probability of expanding. This assumption may be concerning to observers of U.S.\@ politics. Without it, one would be unable to recenter the instrument without specifying the distribution over the expansion indicators $g \equiv (g_{1},\ldots,g_{43})$, as in the market access application in Section (ref). If there is within-party heterogeneity in the expansion probability, then the expansion probabilities will no longer be absorbed by the party-by-year fixed effects used in (ref).

To evaluate sensitivity to (ref), we apply Proposition (ref) to allow states with the same party to have different expansion probabilities. We take $\{g_{1s},\ldots,g_{Ks}\}_{s=1}^{S}$ to be the full set of $S = 2^{43}$ possible binary realizations.\footnote{ Because $\mu_{i}$ only depends on the marginal distributions of each $g_{k}$ separately, the program in Proposition (ref) is equivalent to one that only has $K = 43$ variables. } We take the sensitivity set to be $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q})$ where $\bar{q}$ puts marginal probability $8/30$ on Republican-led states and probability $11/13$ on the Democrat-led states. These are the empirical ex-post probabilities of states expanding by party, which is what borusyakhull2026e use to specify their shock distribution in their Monte Carlo simulations. We abuse notation slightly by not applying the $\delta$ expansion in $\mathcal{P}_{\textsc{m}}(\delta \vert \bar{q})$ to the Democrat-led states.\footnote{ This would be like having $\delta_{k}$ depend on $k$ in the definition of $\mathcal{P}_{\textsc{m}}$, with $\delta_{k}$ fixed at one for states $k$ that are Democrat-led. } This is intended to keep the exercise simple by considering sensitivity to the Republican-led states only.

\fig[\textwidth][p] {figs/medicaid/medicaid_sensitivity_facets.pdf} {fig:medicaid} {Sensitivity to heterogeneous Republican-led expansion probabilities} { Bounds from solving (ref) with $\mathcal{P} = \mathcal{P}_{\textsc{m}}(\delta \vert \bar{q})$ for different values of $\delta$. Different outcomes $y_{i}$ are in the columns and different endogenous variables $x_{i}$ are in the rows. Following borusyakhull2026e, we set the baseline $\bar{q}$ to be 8/30 for Republican-led states and 11/13 for Democrat-led states. However, we only consider sensitivity to allowing deviations from $\bar{q}$ for the Republican-led states. At $\delta = 1$, all Republican-led states have an equal ex-ante probability of expansion, which reproduces the borusyakhull2026e recentered IV estimate. Standard errors clustered by state are shown in parentheses, following the same approach to inference as in borusyakhull2026e. For the bounds, these show the standard errors at the optimizer for the program. }

Figure (ref) shows the results for different values of $\delta$, together with the two instrumented difference-in-differences estimates reported by borusyakhull2026e. borusyakhull2026e point out that the precision gains in their recentered IV estimate lead to standard errors for the impact of Medicaid eligibility on private insurance take-up that are 70% smaller than for the simulated IV estimate. The first row of Figure (ref) shows that this conclusion comes at the risk of bias from incorrectly specifying the expansion probabilities. If Republican-led states are allowed to have ex-ante probabilities between $.11$ and $.67$ ($\delta = 2.5$) rather than a uniform $.27$ ($\delta = 1$), then a wide range of conclusions are available: eligibility could have a negative effect on private insurance take-up similar to that found by the simulated instrument or it could have a null effect. The impacts on employer-sponsored health insurance are even more stark and show that the sign-flip found by borusyakhull2026e is highly fragile to their assumed expansion probabilities. The second row of Figure (ref) changes the endogenous variable from Medicaid eligibility to Medicaid enrollment, as in borusyakhull2026e Table 2, Panel B. Even greater sensitivity is found here; in particular the negative effect on employer-sponsored insurance can be statistically insignificant if ex-ante Republican-led expansion probabilities can vary between $.18$ and $.40$ ($\delta = 1.5$) and positive if these probabilities are allowed to vary between $.13$ and $.53$ ($\delta = 2.0$).

Conclusion

We developed a computationally tractable method for systematically assessing the sensitivity of estimators based on formula instruments to the assumed distribution of counterfactual shocks. The estimator can be implemented in our companion package formulaiv. We applied our estimator to both of the empirical applications in borusyakhull2023e and borusyakhull2026e and found both to exhibit substantial sensitivity to the assumptions on the distribution of counterfactual shocks.

Our analysis suggests that researchers using formula instruments should be cautious about the specification of counterfactual shocks. When these shocks represent events such as the opening of a high-speed rail line or a state policy change, it seems like a challenging exercise to divine the “correct” shock distribution. Other examples suggested in borusyakhull2023e, such as the probability of earthquakes, likely face similar challenges, which can be assessed quantitatively using our methods. These uses of formula instruments have begun to be adopted by empirical researchers: see, for example, dellolken2020res, bosshartweigand2025, buhlerdickens2025, moroninicolettisalvanestominey2025, and doellingsenlim2025wp.

However, there are other uses of formula instruments that rely on institutional knowledge of how the instrument was assigned. Examples include chaureynayyarsharmaverhoogen2025, hollenbeckhristakevauetake2025wp, cailinszeidl2026wp, baguesmakanyvattuonezinovyeva2026wp, jensenkumarpoensgen2026wp, and gao2026wp. Sensitivity to these formulas is likely a smaller concern, because the distribution of counterfactual shocks is determined by the randomization protocol. For these applications, our method can be used to provide a robustness check to deviations from the stated protocol.

\inputbibliography{formulaiv}

\startappendix

Proof of Proposition (ref)

We first establish the case where $D(p) \geq 0$ for all $p \in \mathcal{P}$. The case where $D(p) \leq 0$ for all $p \in \mathcal{P}$ follows symmetrically after the noted changes.

The linear program (ref) is the charnescooper1962nrlq transform of the linear-fractional program (ref). The two programs yield the same optimal values when $D(p) > 0$ for all $p \in \mathcal{P}$; see boydvandenberghe2004 for a textbook treatment. If $D(p) = 0$ for some $p \in \mathcal{P}$, then the optimal value may be unbounded. After the transformation, the simplex membership $p \in \Delta^{S}$ becomes $\phi \geq 0$ and $\sum_{s=1}^{S}\phi_{s} = \tau$, while the remaining constraints $Ap \leq c$ that define $\mathcal{P}$ become $A\phi \leq c\tau$.

Now suppose that $b \in [\underline{\beta}_{\textsc{riv}}(\mathcal{P}), \overline{\beta}_{\textsc{riv}}(\mathcal{P})]$ is a real number. The objective function of (ref) is continuous in $(\tau, \phi)$, and the constraint set of (ref) is convex, hence connected. So the image of the objective function over the constraint set is an interval with infimum $\underline{\beta}_{\textsc{riv}}(\mathcal{P})$ and supremum $\overline{\beta}_{\textsc{riv}}(\mathcal{P})$ rudin1976. Because a feasible linear program attains any finite optimal value, this interval contains its finite endpoints and therefore contains every real $b \in [\underline{\beta}_{\textsc{riv}}(\mathcal{P}), \overline{\beta}_{\textsc{riv}}(\mathcal{P})]$. It follows that there exists a $\phi(b), \tau(b)$ pair that is feasible in (ref) that produces objective value $b$. Suppose momentarily that $\tau(b) > 0$. Let $p(b) = \phi(b)/\tau(b)$. Then $p(b) \in \mathcal{P}$ and

align*[align* omitted — 609 chars of source]

This shows that there exists a $p(b) \in \mathcal{P}$ that produces $\hat{\beta}_{\textsc{riv}}(p(b)) = b$.

We conclude the proof by showing that a feasible pair $\tau(b), \phi(b)$ cannot have $\tau(b) = 0$. If $\tau(b) = 0$, then the constraints $\phi_{s}(b) \geq 0$ for all $s$ and $\sum_{s=1}^{S}\phi_{s}(b) = \tau(b) = 0$ force $\phi_{s}(b) = 0$ for all $s$. But then the normalization constraint

align*[align* omitted — 104 chars of source]

reduces to $0 = 1$, contradicting the feasibility of $\tau(b), \phi(b)$.

Finally, suppose that $D(p)$ takes both positive and negative values on $\mathcal{P}$ and that $\hat{\beta}_{\textsc{riv}}(p)$ is not constant on $\mathcal{P}_{D \neq 0}$. Because $\mathcal{P}$ is connected and $D$ is continuous, there exists some $p_{0}$ such that $D(p_{0}) = 0$ while the numerator of $\hat{\beta}_{\textsc{riv}}(p)$ is non-zero; suppose it is positive for concreteness. Then taking a feasible sequence of $p$ that approaches $p_{0}$ from within the set $\{p \in \mathcal{P} : D(p) > 0\}$ produces arbitrarily large values of $\hat{\beta}_{\textsc{riv}}(p)$, while taking a feasible sequence from within the set $\{p \in \mathcal{P} : D(p) < 0\}$ produces arbitrarily small values of $\hat{\beta}_{\textsc{riv}}(p)$. We conclude that $\underline{\beta}_{\textsc{riv}}(\mathcal{P}) = -\infty$ and $\overline{\beta}_{\textsc{riv}}(\mathcal{P}) = +\infty$.

Extension to general assignment processes

In Section (ref), we assumed that $G(\cdot \vert w)$ is independent of $w$ with discrete support. In this appendix, we relax this assumption by assuming instead that $G(\cdot \vert w)$ has a density $\gamma(\cdot \vert w)$ with respect to some known dominating measure $\lambda$. Then

align[align omitted — 149 chars of source]

Suppose that $\gamma$ can be written using a finite basis expansion as

align[align omitted — 107 chars of source]

where $\gamma_{s}$ are known basis functions. This nests the case considered in the main text by taking $\lambda$ to be counting measure on the finite set $\{g_{1},\ldots,g_{S}\}$ and $\gamma_{s}(g \vert w) = \IndicSmall{g = g_{s}}$ for $s = 1,\ldots,S$. Substituting (ref) into (ref) produces

align[align omitted — 153 chars of source]

with $f_{is}$ redefined (generalized) from the main text. This has the same form as equation (ref), except that instead of being known directly, the quantities $f_{is}$ need to be computed, for example by drawing from $\lambda(g)$. Having done that, Proposition (ref) proceeds unchanged as long as $p$ is constrained to a polyhedral sensitivity set $\mathcal{P}$.

Predicting HSR line openings

In this section, we describe how we use machine learning algorithms to predict the opening probability of HSR lines for the application in Section (ref).

The variable being predicted is a binary indicator for whether the HSR line was open by 2016. The predictors are variables predetermined as of 2007: the 2007 opening status, anticipated railway speed, line length, line type, number of links, and plan type. A few plan type categories appear for only one or two lines, which we pool into a separate “other” category. borusyakhull2023e set some line opening probabilities to one across all of their scenarios; we continue to do this in our prediction exercise, while focusing our attention on the other lines.

\fig {figs/china/model_eval.pdf} {fig:model_eval} {Out-of-fold forecast performance} { See the text of Appendix (ref) for details. }

We train four different learners: a random forest, penalized logistic regression, gradient-boosted trees, and $k$-nearest neighbors. For each one, we use nested cross-validation and evaluate performance using the area under the ROC (true positive rate/false positive rate) curve, often abbreviated as the AUC. In the outer loop, we split the data into five folds. In the inner loop, we use ten-fold cross-validation within the four training folds (repeated five times) to select tuning parameters by cross-validated AUC. We then use the optimal tuning parameters to construct a prediction for lines in the left-out fifth fold. Repeating this process for each of the five folds produces out-of-fold predictions for each line.

Figure (ref) reports the out-of-fold AUC for each learner together with the AUC for the naive data-agnostic predictions implied by the borusyakhull2023e baseline $\bar{q}_{\textsc{bh}}$. Unsurprisingly, the learners that use data perform substantially better with out-of-fold AUC between roughly $.79$ and $.83$, compared to $.65$ for the $\bar{q}_{\textsc{bh}}$ baseline. As we saw in Figure (ref) in the main text, the four learners produce substantially different marginal distributions $\bar{q}_{m}$ even while performing comparably on the AUC measure.

Additional figures

\fig[\textwidth][h!] {figs/china/est_joint_with_controls.pdf} {fig:est_joint_with_controls} {Adding geographic controls to Figure (ref)} { The figure is the same as Figure (ref) but with controls for distance to Beijing, latitude, and longitude, as in Panel B of Table I in borusyakhull2023e. See notes for Figure (ref). }