Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
128,266 characters · 18 sections · 68 citation commands
Policy Evaluation during a Pandemic
\abstract{National and local governments have implemented a large number of policies in response to the Covid-19 pandemic. Evaluating the effects of these policies, both on the number of Covid-19 cases as well as on other economic outcomes is a key ingredient for policymakers to be able to determine which policies are most effective as well as the relative costs and benefits of particular policies. In this paper, we consider the relative merits of common identification strategies that exploit variation in the timing of policies across different locations by checking whether the identification strategies are compatible with leading epidemic models in the epidemiology literature. We argue that unconfoundedness type approaches, that condition on the pre-treatment “state” of the pandemic, are likely to be more useful for evaluating policies than difference-in-differences type approaches due to the highly nonlinear spread of cases during a pandemic. For difference-in-differences, we further show that a version of this problem continues to exist even when one is interested in understanding the effect of a policy on other economic outcomes when those outcomes also depend on the number of Covid-19 cases. We propose alternative approaches that are able to circumvent these issues. We apply our proposed approach to study the effect of state level shelter-in-place orders early in the pandemic.}
JEL Codes: C21, C23, I1
Keywords: Policy Evaluation, Difference-in-Differences, Unconfoundedness, Covid-19, Pandemic, Mediators
\onehalfspacing
There have been a large number of policies implemented in order to decrease the spread of Covid-19. During the early part of the pandemic, the most important of these policies were non-pharmaceutical interventions such as requirements to wear masks, making Covid-19 tests widely available, contact tracing, school closures, lockdowns, and others. These sorts of policies are likely to come with a number of tradeoffs in terms of effectiveness in reducing the spread of Covid-19 as well as their effects on individuals' economic and psychological well-being. Thus, understanding the effects of different policies along a number of dimensions (both effects on number of cases as well as effects on other outcomes) is a key ingredient for researchers, policymakers, and governments to consider when evaluating Covid-19 related policies.
The main way that these policies have been studied by researchers is to compare outcomes in locations that implemented some policy to outcomes in another location that did not implement the policy. Researchers typically exploit having access to panel data --- data on cases, testing, and economic variables is generally widely available for particular locations over multiple time periods --- to try to understand these effects. This sort of setup is very familiar to many researchers in economics, and, with this sort of data availability, an almost default strategy of empirical researchers is to use difference-in-differences. And, indeed, difference-in-differences has been widely used to study the effects of policies in response to Covid-19.
In the current paper, we argue that difference-in-differences has properties that make it relatively less attractive for conducting policy evaluation during a pandemic than it typically would be for most applications in economics. The intuition for our results is that difference-in-differences methods are typically motivated by a two-way fixed effects model for untreated potential outcomes (see, for example, blundell-dias-2009). The key feature of these models is that they include additively separable unit-level unobserved heterogeneity (i.e., a fixed effect). This sort of heterogeneity is very common in applications in economics; a textbook example would be an application on the effect of some policy on individuals' earnings where the researcher is worried that “ability” is unobserved, affects earnings, and is distributed differently between the group of individuals that are affected by the policy and the group of individuals not affected by the policy. In this setup, the unobserved heterogeneity can be differenced out and paths of outcomes for the group of treated units and the group of untreated units can be compared to each other to deliver the effect of the policy.
However, this sort of motivation does not apply in the case of Covid-19. In particular, the main epidemiological models for Covid-19 transmission are highly nonlinear and depend on (i) the number of currently infected individuals in a particular location, (ii) the number of susceptible individuals in a location, and (iii) the transmission properties of Covid-19. In other words, the key challenge for identifying effects of Covid-19 related policies on the number of Covid-19 cases is not that different locations are different in terms of unobserved heterogeneity, but rather but that pre-policy differences in the “state of the pandemic” between treated and untreated locations (e.g., differences in the number of Covid-19 cases before the policy is implemented) can lead to substantial differences between locations in how the pandemic would have evolved absent the policy being implemented. These differences can make it difficult to evaluate the effects of Covid-19 related policies. Moreover, these differences are a major concern for evaluating early pandemic policies where different locations' policy choices often depended on the state of the pandemic in those locations.
Another central issue in the economics literature is to understand the effect of Covid-19 related policies on various economic outcomes.\footnote{Throughout the text, we use the term “economic outcomes” but our results apply to any outcome of interest that is outside the epidemic model that we consider in the paper.} Understanding the effects of Covid-19 related policies on economic outcomes is essential in order to understand the costs and benefits of various policies that are aimed at reducing the number of Covid-19 cases. For economics outcomes, we focus on the simple leading case where untreated potential outcomes are generated by a two way fixed effects model that also depends on the current number of Covid-19 cases in a particular location. This setup allows both for the active number of Covid-19 cases to have an effect on the outcome of interest and for the active number of cases to themselves be affected by the policy. In this case, we consider (i) a “standard” version of difference-in-differences that directly compares paths of economic outcomes among treated and untreated locations, and (ii) difference-in-differences when the current number of Covid-19 cases is included as a regressor. We show that, generally, neither approach can deliver the average effect of the treatment on economic outcomes across treated locations. The first strategy breaks down because of differences in the current number of cases across treated and untreated locations. The second strategy breaks down when the policy affects the number of Covid-19 cases (which is the goal of the policy).
We propose alternative approaches that address both sets of issues mentioned above. In particular, for evaluating the effect of policies on Covid-19 cases, we first show that unconfoundedness-type identification strategies (i.e., strategies that compare locations that have the same pre-treatment values of key Covid-19 related variables) do not suffer from the same drawbacks as difference-in-differences approaches. For this case, we propose an estimation strategy that involves estimating (i) the propensity score (i.e., the probability of experiencing the policy conditional on pre-treatment values of Covid-19 related variables) and (ii) an outcome regression for Covid-19 cases in the absence of the policy that is related to the epidemic model. Our approach is doubly robust in the sense that it delivers consistent estimates of policy effects if either the propensity score or outcome regression is correctly specified. This is important as it implies that we can circumvent having to estimate a full structural epidemic model while still delivering estimates of policy effects that are compatible with the epidemic model.
Second, for evaluating the effect of Covid-19 related policies on economic outcomes, we propose a two-step approach where the parameters of a two way fixed effects model that additionally allows for current cases to affect the outcome are identified using untreated locations in the first step. Then, in the second step, we recover the path of active cases that treated locations would have experienced on average if they had not been treated (this follows under similar unconfoundedness type arguments used for cumulative cases above). Using these two pieces of information, we are able to construct the average economic outcome that treated locations would have experienced if they had not participated in the policy --- and, therefore, the average effect of the policy is identified for treated locations. Unlike other common approaches, this approach allows both for the policy to affect the current number of cases and for the current number of cases to affect untreated potential outcomes.
We conclude the paper by studying the effect of state-level shelter-in-place orders (SIPOs) on the number of Covid-19 cases and on travel early in the pandemic. These are challenging policies to evaluate because states that implemented these policies also tended to have a large number of Covid-19 cases earlier than states that did not implement this type of policy (or implemented it later). This correlation mechanically leads to larger increases in Covid-19 cases among early treated states relative to untreated and later-treated states. This additionally implies that parallel trends is violated and can lead to difference-in-differences estimates that these policies increased the number of Covid-19 cases which is clearly unreasonable. We additionally show that difference-in-differences estimates are very sensitive to minor changes in how they are specified. And, interestingly, despite very different post-policy estimates across different types of DID specifications, none of them are rejected in pre-treatment periods. This suggests that it is not feasible to choose between alternative difference-in-differences specifications based on their performance in pre-treatment periods. Using difference-in-differences, we also sometimes spuriously estimate meaningfully large effects of placebo SIPOs in states that did not actually implement a SIPO. In contrast, we document notably better performance of the unconfoundedness strategy along several dimensions: this approach does not lead to any estimates that SIPOs increased Covid-19 cases, and it does not lead to large or statistically significant effects of placebo SIPOs among states that did not actually implement a SIPO.
There are a number of recent papers at the intersection of economics, Covid-19, and policy evaluation, and here we only briefly summarize some of the most related ones.
The most related papers to ours are several methodological papers on evaluating Covid-19 related policies. allcott-boxell-conway-ferguson-gentzkow-goldman-2020 propose an event-study regression estimator that is motivated by SIRD models (the same type of model that we consider below) though their approach ends up being substantially different from ours. goodman-marcus-2020,gauthier-2021 discuss two-way fixed effects regressions in the context of evaluating Covid-19 policies. chernozhukov-kasahara-schrimpf-2021 propose an alternative approach to evaluate Covid-19 related policies that is motivated by a SIRD model. Their approach is more structural than ours which comes with tradeoffs. For example, their approach can be useful for evaluating counterfactual policies while ours is geared towards evaluating the effects of policies that were actually implemented. On the other hand, our approach generally requires fewer assumptions to evaluate policies that were actually enacted and is set up to be robust to general forms of treatment effect heterogeneity (which is likely to be important in contexts like many early pandemic policies where policies were often implemented at different times in different locations; see goodman-marcus-2020 for some related discussion). aleman-busch-ludwig-santaeulalia-2020 provide a way to transform different pandemic “states” across locations and evaluate pandemic-related policies exploiting different locations being in different “stages”.
Difference-in-differences has been widely used in empirical work to study the effects of Covid-19 policies. Some examples of papers that consider Covid-19 related policies using difference-in-differences types of identification strategies include bartik-bertrand-lin-rothstein-unrath-2020, berry-fowler-glazer-handel-macmillen-2021, chetty-friedman-hendren-stepner-2020, courtemanche-garuccio-le-pinkston-yelowitz-2020, dave-friedson-matsuzawa-mcnichols-sabia-2020, dave-friedson-matsuzawa-sabia-safford-2020, gapen-millar-blerina-sriram-2020, glaeser-jin-leyden-luca-2021, goolsbee-syverson-2021, gupta-montenovo-nguyen-rojas-schmutte-simon-weinburg-2020, haynes-kulkarnia-li-siddique-2022, hsiang-et-al-2020, juranek-zoutman-2021, kong-prinz-2020, villas-sears-villas-villas-2020, wright-sonin-driscoll-wilson-2020, and ziedan-simon-wing-2020. haber-et-al-2022 provides a recent survey of empirical strategies that have been used to evaluate Covid-19 related policies.
Finally, on the econometrics side, our paper is related to a large literature on unconfoundedness and difference-in-differences (see imbens-wooldridge-2009 for a survey of this literature). More notably, our results on checking the compatibility of structural epidemic models with reduced form identification strategies is broadly similar to a number of papers in econometrics; for example, just in the context of panel data, heckman-robb-1985,heckman-ichimura-todd-1997,athey-imbens-2006,blundell-dias-2009, chabe-2015,ghanem-santanna-wuthrich-2022,marx-tamer-tang-2022 all provide connections between structural models and conditions under which various reduced form approaches, such as difference-in-differences, can be compatible with these models. Our contributions on allowing for infections to both be affected by the policy and to have a direct effect on economic outcomes appear to be conceptually new but related to work on mediation analysis (see huber-2020 for a summary of this literature) which is also broadly related to work on simultaneous equation models (e.g., griliches-1977,imbens-newey-2009). Viewing the “untreated potential” number of Covid-19 infections as a covariate in the model for untreated potential outcomes is related to the idea of difference-in-differences with covariates that can be affected by the treatment; see, bonhomme-sauder-2011,lechner-2011,caetano-callaway-payne-rodrigues-2022 for related discussion along these lines.
To start with, in this section we provide examples of the main types of issues that can confound policy analysis during a pandemic using simulated data from the leading type of epidemic model that has been widely used in the context of Covid-19. (ref) shows the paths of the key variables during a simulated pandemic coming from a stochastic SIRD model. SIRD models categorize individuals in a population into being S-Susceptible, I-Infected, R-Recovered, or D-Dead. We discuss this model in substantially more detail in the next section. The shapes of the path of each variable is typical of a SIRD model. In particular, at some point in time, some small number of cases shows up in a particular location. Then, the number of infections rise in early periods when there are a large number of susceptible individuals in that location combined with an increasing number of currently infected (which also implies contagious). As the number of susceptible decreases (i.e., as infected individuals recover or die), eventually the number of infected individuals decreases. Simultaneously, the cumulative number of cases, number of recovered individuals, and number of deaths all initially grow before eventually leveling off.
In this example, we consider the case where a new policy is implemented in some locations in period 150. For simplicity, we consider the case where the policy has no effect on Covid-19 cases. Locations that participate in the treatment and locations that do not participate in the treatment are alike in all ways except that treated locations tend to experience their first cases earlier than untreated locations. Panel (a) of (ref) shows plots of the average paths of cumulative cases for treated locations and untreated locations in this setup.
Panel (b) of (ref) shows event study-type estimates of the effect of the treatment on the number of cases. To be precise, these are difference-in-differences type estimates where the estimated effect comes from the average change in cases experienced by the treated group of locations relative to the change in cases experienced by the untreated group of locations over the same time periods. Taken at face value, the estimated effects in Panel (b) suggest that the policy decreased the number of Covid-19 cases in treated locations relative to what they would have been in the absence of the policy. However, recall that, in our simulation setup, the policy has no effect on Covid-19 cases. Thus, this example demonstrates that difference-in-differences can perform poorly in the context of trying to evaluate the effect of a policy on the number of Covid-19 cases. The key driver of this poor performance is (i) the nonlinearity of the model for Covid-19 transmission and (ii) differences in the timing of the first cases between locations that participate in the treatment and those that do not. The first of these is an inherent feature of trying to evaluate the effects of policies on Covid-19 cases. For the latter, generally, the bias of difference-in-differences approaches for policy evaluation becomes more severe as the timing of first cases becomes more different between treated and untreated locations.\footnote{Interestingly, the best case for difference-in-differences is when the timing of first cases is the same across treated and untreated locations. However, this is also a case where there is no need to take a time difference at all and one could just make level comparisons of Covid-19 cases across locations.}
Next, we consider the effect of the policy on some economic outcome of interest. (ref) continues with the same simulated policy as above. As above, we consider the case where the policy has no effect on cases or on economic outcomes. However, in this simulation we allow for the economic outcome to depend on the number of active Covid-19 cases in a particular location (here, more active cases tend to decrease the economic outcome), but, otherwise, the economic outcome would follow parallel trends. Panel (a) shows average paths of outcomes for treated and untreated locations in this setup. Panel (b) shows event study type estimates under the assumption of parallel trends. As before, and even in this very simple example, differences in the timing of first cases lead to violations of parallel trends that lead to poor estimates of the effect of the policy on the economic outcome of interest.
Finally, we contrast the poor performance of difference-in-differences in both of these contexts with using an unconfoundedness type strategy to deal with the pandemic-related variables. In particular, when Covid-19 cases is the outcome, we effectively compare locations that were in a similar pandemic “state” in the period right before the policy was implemented. For the economic outcome, we continue to use a version of difference-in-differences, but one that, in the absence of the policy, accounts for economic outcomes depending on the number of Covid-19 infections that would have occurred if the policy had not been implemented using an unconfoundedness strategy (see (ref) below for more details). Estimates using these approaches are provided (ref). In both cases, these strategies perform notably better at evaluating the effects of the policy.
In this section, we briefly discuss a stochastic SIRD model which is the workhorse model of epidemic spread in epidemiology and has been used extensively to forecast the spread of Covid-19 cases. SIRD models have a long history in epidemiology --- a deterministic version of this kind of model was proposed by kermack-mckendrick-1927. Stochastic SIRD models are discussed in allen-2008,allen-2017 and have been considered by economists in oka-wei-zhu-2021, fernandez-jones-2022,ellison-2020,acemoglu-chernozhukov-werning-whinston-2021, among others.
\paragraph{Notation:} Let $N_l$ denote the number of individuals in location $l$. Let $\mathcal{T}$ denote the total number of time periods. The number of susceptible individuals in location $l$ in a particular time period $t$ is denoted by $S_{lt}$, the number of currently infected individuals in location $l$ at time period $t$ is denoted by $I_{lt}$, the cumulative number of recovered individuals is denoted by $R_{lt}$, and the number of cumulative deaths is denoted by $\delta_{lt}$. All individuals in the population are in exactly one of these states at a particular point in time so that
in all time periods. Later, we will be interested in the effect of the policy on the cumulative number of cases by time period $t$, and we denote this variable by $C_{lt}$ and note that $C_{lt} = N_l - S_{lt}$.
In a SIRD model, the paths of all of these variables are governed by some transition equations. The transition equations have the Markov property; i.e., the path of each outcome over time only depends on the “state” of location $l$ in the immediately preceding period. And, in particular, these transition equations are given by
where for some time period $t$, we define $\mathcal{F}_{lt} = (S_{lt},I_{lt},R_{lt},\delta_{lt})$, and often refer to this as the “state” of the pandemic in location $l$ in time period $t$. It is worth considering each of these equations in some more detail. To start with, consider the term $\beta \frac{I_{lt-1}}{N_l}S_{lt-1}$ which shows up in (ref). This is the expected number of new cases in time period $t$ conditional on the state of the pandemic in location $l$ in time period $t-1$. The expected number of new cases from one period to the next depends on three things. First, it depends on $I_{lt-1}/N_l$ which is the fraction of individuals that are infected in period $t-1$. Holding other things constant, when more individuals are infected, it implies an expected larger increase in the number of cases. Second, the expected number of new cases depends on the number of susceptible individuals in the population. Intuitively, when there are more susceptible individuals, the number of cases grows more rapidly (other things constant). The spread of a pandemic stops when the number of susceptible becomes small enough which can happen either through “herd immunity” or by decreasing the number of susceptible (for example, through the introduction of a vaccine). Finally, the change in the number of cases depends on the parameter $\beta$ which is called the infection rate. Most non-pharmaceutical interventions are aimed at changing the infection rate --- here, there are two potential benefits: (i) decreasing the infection rate through non-pharmaceutical interventions decreases the total number of cases that need to occur before reaching herd immunity,\footnote{This is also one explanation for repeated “waves” of Covid-19 cases. That is, the infection rate may be temporarily reduced by policy intervention or individual choices but then increases again once these interventions are relaxed.} and (ii) if there is a vaccine on the horizon, it also would decrease the total number of cases that occur before herd immunity is reached through the vaccine.
Next, consider (ref). This transition equation says that, on average, the number of total recoveries in location $l$ in time period $t$ (conditional on the state of the pandemic in period $t-1$) is equal to number of individuals in location $l$ that have already recovered by time period $t-1$ plus some fraction of infected individuals in period $t-1$. This fraction is determined by the parameter $\lambda$ which is the recovery rate from Covid-19. (ref) is the transition equation for deaths. The key parameter is $\gamma$ which parameterizes the death rate from being infected with Covid-19.
Next, consider (ref). This is the transition equation for active Covid-19 cases. The expected number of infections in period $t$ thus depends on (i) the remaining cases after accounting for recoveries and deaths (this is the first term in (ref)), and (ii) the expected number of new cases (this is the second term in (ref)). Finally, in (ref), the expected number of susceptible individuals is equal to the number of susceptible individuals in time period $t-1$ minus the expected number of new cases; likewise, the expected number of cumulative cases by time period $t$ is equal to the number of cumulative cases in time period $t-1$ plus the expected number of new cases.
The previous section presented a basic stochastic SIRD model. This section connects that sort of model with the treatment effects literature and considers the relative merits of difference-in-differences and unconfoundedness strategies for evaluating the effect of a policy on the number of Covid-19 cases.
The strategy of this section is to impose the stochastic SIRD model for untreated potential outcomes and to check if difference-in-differences and/or unconfoundedness are compatible with the stochastic SIRD model. This setup does not place restrictions on how treated potential outcomes (i.e., Covid-19 cases under the policy) are generated. In particular, this is consistent with Covid-19 cases under the policy continuing to follow a stochastic SIRD model but where the values of the parameters potentially change in response to the policy; but it is also more general than that in the sense that there are no substantive restrictions on treated potential outcomes. Perhaps more importantly, this setup also allows for heterogeneous effects of policies across different locations.
\paragraph{Additional Treatment Effects Notation} To make the connection with the treatment effects literature, we start by introducing some additional notation. First, we define $D_l$ as a binary variable indicating whether or not location $l$ participated in the treatment. We also define treated and untreated versions of all of the variables in the stochastic SIRD model. In particular, for generic time period $t$, $S_{lt}(0)$, $I_{lt}(0)$, $R_{lt}(0)$, and $\delta_{lt}(0)$ are the number of susceptible, infected, recovered, and dead individuals in location $l$ in time period $t$ if the policy had not been enacted. We also define $C_{lt}(0)$ as the cumulative number of cases in location $l$ by time period $t$ if the policy had not been enacted. Similarly, we define $S_{lt}(1)$, $I_{lt}(1)$, $R_{lt}(1)$, $\delta_{lt}(1)$, and $C_{lt}(1)$ to be the corresponding treated potential variables; i.e., the values of each of these if the policy had been enacted. Following a large literature on policy evaluation which exploits having access to panel data, we consider the case where the researcher has access to some pre-treatment periods. We suppose that the policy is implemented for treated locations in time period $t^*$ where $1 < t^* \leq \mathcal{T}$.\footnote{In practice, the timing of implementing a particular policy may vary across different locations. Extending our arguments to this case is relatively straightforward, and, therefore, this section considers the case where the policy is implemented at the same time across all treated locations. See (ref) below for additional discussion on this point.} For random variables indexed by time periods, we define $\Delta X_t := (X_t - X_{t-1})$. Because we are also interested in how policy effects vary across time, some of our arguments involve “long differences” where, for $t_2 > t_1$, we define $\Delta^{(t_1,t_2)} X_t := X_{t_2} - X_{t_1}$. In (ref), we write the SIRD model given in the previous section in terms of untreated potential outcomes, and we refer to this model as the (ref) throughout the remainder of the paper.
Our main interest for this part of the paper is the effect of the policy on the cumulative number of Covid-19 cases. Typically, the main parameter of interest in DID applications (and the parameter that we focus on in the current paper) is the Average Treatment Effect on the Treated (ATT). It is given by
where we index the $ATT$ by $C$ to indicate that we are considering the effect of the policy on the cumulative number of cases in time period $t$. $ATT^C_t$ is the difference between cumulative cases under the policy relative to cumulative cases in the absence of the policy on average among locations that participated in the treatment. That this parameter is disaggregated by time period makes it straightforward to report across time periods (as in an event study), but it is also straightforward to, for example, average it across post-treatment time periods in order to report an overall average effect of participating in the treatment.
The main underlying motivation for considering a DID approach is when a researcher thinks that untreated potential outcomes are generated from a two-way fixed effects model (see, for example, blundell-dias-2009). These sorts of models are attractive in many applications in economics where there are thought to be important unobserved differences between individuals (or firms, etc.) that are not observed by the researcher. In labor economics, these are often thought of as being unobserved skill; in industrial organization, these may be unobserved differences in productivity across firms; and, in health economics, these may be thought of as proneness to particular health conditions. However, there is an important difference of Covid-19 relative to all of these cases. In general, particular locations do not have time invariant unobservables that make them more or less likely to have a large number of cases; instead, the key differences between locations are (i) the timing of their initial case(s), and (ii) the pandemic response (both in terms of policies and in terms of actions taken by the populations in different locations).\footnote{One caveat to this is that different locations may have characteristics that are related to the parameters of the SIRD model discussed above. See (ref) below for more discussion along these lines.}
The main result in this section is that there are likely to be major drawbacks to using DID to evaluate the effects of Covid-19 related policies on the number of Covid-19 cases. The two primary reasons for this are (i) the highly nonlinear spread of Covid-19 cases during a pandemic and (ii) that the key difference between locations is the current number of Covid-19 cases rather than some fixed unobserved difference between locations in terms of “proneness” to having a large number of cases.
In this section, we consider whether difference-in-differences approaches are compatible with the stochastic SIRD model presented above. We begin by providing some background on using difference-in-differences to identify the effect of some policy. The key identifying assumption in a DID application is the following parallel trends assumption.
The parallel trends assumption says that the path of Covid-19 cases that locations in the treated group would have experienced if they had not participated in the treatment is the same as the path of Covid-19 cases that locations in the untreated group did experience. Invoking this assumption leads to the following estimand for the $ATT^C_t$ for $t \geq t^*$
where we use the notation $DID^C_t$ to highlight that it may not be equal to $ATT^C_t$. $DID^C_t$ is equal to the path of Covid-19 cases that treated locations experienced adjusted by the path of Covid-19 cases that untreated location experienced; if the parallel trends assumption holds, then the latter is the path of Covid-19 cases that treated locations would have experienced on average if they had not experienced the policy, and $DID^C_t$ would be equal to $ATT^C_t$. And, regardless of whether or not the parallel trends assumption holds, $DID^C_t$ is the population quantity for what is estimated in DID applications on Covid-19. Before providing our main result on using difference-in-differences to identify/estimate the effect of a policy on Covid-19 cases, it is also worth mentioning that the primary motivating model for difference-in-differences identification strategies is one where
where $\theta_t$ is a time fixed effect, $\eta_l$ is location-specific unobserved heterogeneity that can be distributed differently between the treated group and untreated group and $v_{lt}$ is an idiosyncratic time varying unobservable. Comparing (ref) to the equation for cumulative Covid-19 cases in (ref), it is immediately clear that these are notably different. In the stochastic SIRD model, the important difference between treated and untreated locations is not unobserved heterogeneity, but rather differences in the current number of Covid-19 cases and the number of susceptible individuals across locations. This immediately provides a suggestive piece of evidence that the parallel trends assumption is unlikely to hold when $C_{lt}(0)$ is generated from a stochastic SIRD model.
The next result makes explicit that the parallel trends assumption is generally violated in stochastic SIRD models and provides an expression for the bias resulting from incorrectly imposing the parallel trends assumption in cases where untreated potential outcomes are generated by a stochastic SIRD model.
The proof of (ref) is provided in (ref). (ref) shows that difference-in-differences generally delivers (potentially severely) biased estimates of the effect of a policy on cumulative Covid-19 cases. It is worth making a few additional comments before proceeding. First, the key reason why the difference-in-differences strategy breaks down is that, in general, the distribution of pandemic related variables immediately before the policy (contained in $\mathcal{F}_{t^*-1}$) is not the same across treated and untreated locations. Due to the nonlinearity of the SIRD model, this leads to violations of the parallel trends assumption. Second, the sign of the bias cannot generally be determined from these expressions. For example, in (ref) above, difference-in-differences resulted in downward biased estimates of the effect of the policy, but the direction of the bias is sensitive to both (i) timing of first cases in treated and untreated locations, and (ii) the timing of the policy itself (this can be clearly seen in Panel (a) of (ref) where setting the policy at an alternative time period could result in parallel trends being violated in the opposite direction).
Some of the expressions in (ref) seem complicated. One special case of this result that is worth pointing out is when $t=t^*$ (so that we are considering the effect of the policy on Covid-19 cases “on impact”). In that case, the bias from using DID is given by
This bias is the difference between the expected number of new cases that treated locations would have experienced in the absence of the policy relative to the expected number of new cases for untreated locations. And, here, it is straightforward to see key reasons why difference-in-differences can perform poorly: if the joint distribution of currently infected and number of susceptible individuals is different among treated and untreated locations, then they would have experienced a different number of new Covid-19 cases even if the policy had not been implemented. In the context of Covid-19, there are some cases where these biases could be substantial. Perhaps the leading example is when the timing of initial Covid-19 cases varied across locations and Covid-19 related policies were implemented earlier in locations that tended to have cases earlier.
A main alternative to difference-in-differences for evaluating the effects of policies is to assume some version of unconfoundedness. Unconfoundedness means that, after conditioning on some covariates, treatment assignment is as good as randomly assigned. In other words, in order to identify the effect of some policy on Covid-19 cases, one can compare Covid-19 cases in locations that experienced the treatment to Covid-19 cases in locations that did not participate in the treatment and had the same characteristics related to the pandemic as treated locations. In this section, we consider a particular version of unconfoundedness that does not suffer from the same limitations as difference-in-differences for evaluating the effects of policies on the number of Covid-19 cases.
Intuitively, the reason why an unconfoundedness strategy works better for studying policy effects of Covid-19 is that the key differences between locations are the current amount of cases and the current number of susceptible individuals rather than differences in location-specific unobserved heterogeneity. Therefore, conditioning on current cases and the current number of susceptible individuals is sufficient for comparisons of treated and untreated locations to deliver causal effects of policies on Covid-19 cases; while the differencing strategy of difference-in-differences is not able to do the same.
The next result is a main result on the validity of identifying policy effects under the assumption of unconfoundedness. Before stating this result, define the propensity score as
which is the probability of being treated conditional on pre-treatment characteristics $\mathcal{F}_{t^*-1}$ and make the following assumption
(ref) is a standard assumption in the treatment effects literature. In the context of Covid-19 related policies, the first part says that there are some locations that participate in the treatment, and the second part says that, for all values of $\mathcal{F}_{t^*-1}$, one can find untreated locations that have those characteristics. This implies that, for all treated locations, there exists matching untreated locations with the same pre-treatment characteristics. In practice, if this condition is violated, one can identify treatment effects that are local to the region of common support (see, for example, crump-hotz-imbens-mitnik-2009).
The proof of (ref) is provided in (ref). This is an important result and implies that, on average, the unobserved number of cumulative cases that locations that participated in the treatment would have experienced if they had not participated in the treatment is the same as the cumulative number of cases that untreated locations actually did experience among locations that had the same pre-treatment characteristics.
Finally, in this section, we provide an identification result for $ATT^C_t$ which is valid under the SIRD model for Covid-19 cases.
(ref) says that, under a stochastic SIRD model, we can evaluate the effect of a policy using an unconfoundedness strategy that compares the number of cases in locations that participated in the treatment to the number of cases in locations that did not participate in the treatment and which had the same Covid-19 related characteristics in the period before the policy was implemented.
It is worth making several additional comments related to the result in (ref). First, estimating $ATT^C_t$ from the expression in (ref) involves estimating the propensity score, $p(\mathcal{F}_{t^*-1})$ and the outcome regression $m^C_{0,t}(\mathcal{F}_{t^*-1})$. It is also possible to derive alternative expressions for $ATT^C_t$ that only require either estimating the propensity score (these would be similar to propensity score re-weighting estimators as in hirano-imbens-ridder-2003) or estimating the outcome regression (these would be similar to regression adjustment estimators). However, the expression for $ATT^C_t$ in (ref) possesses the double robustness property.\footnote{For completeness, we provide a proof in the Supplementary Appendix, but the arguments follow along the same lines as arguments for existing doubly robust estimators under unconfoundedness.} A main advantage of a doubly robust estimator is that it provides consistent estimates of $ATT^C_t$ if either the propensity score model or the outcome regression model are correctly specified (see, for example, bang-robins-2005,sloczynski-wooldridge-2018). Double robustness is particularly appealing in this context as it enables us to side-step the problem of estimating the full SIRD model and instead involves estimating a model of the treatment assignment process which is both familiar to economists and may be substantially more feasible to do with a simple parametric model. In unreported simulations, we found that imposing flexible parametric models for both the propensity score and the outcome regression performed notably better than either the pure outcome regression approach or the propensity score re-weighting approach.
Second, it is worth briefly mentioning that the weights in (ref) are normalized to have mean one in finite samples. This type of normalized weights is said to be of the H{\'a}jek-type (hajek-1971) and typically results in estimators with improved finite sample properties relative to its unnormalized counterpart (busso-dinardo-mccrary-2014). Finally, we provide the asymptotic properties of our estimator in the Supplementary Appendix. In order to conduct inference, we use a multiplier bootstrap procedure that involves perturbing the influence function of the estimator of $ATT^C_t$; we also discuss how to conduct uniform inference across different time periods to account for multiple testing. These results primarily follow from recent results on doubly robust estimators with H{\'a}jek-type weights in santanna-zhao-2020.
Another interest of economists is studying the effect of Covid-19 related policies on other (particularly economic) outcomes. This is likely to be useful for thinking about a cost-benefit analysis of particular policies. Relative to textbook versions of difference-in-differences, what is different in this section is that we allow for economic outcomes to depend on the current number of Covid-19 cases in a particular location.\footnote{As above, because the target parameter is an ATT-type parameter, the setup in this section does not require assumptions on how treated potential outcomes are generated and, therefore, the discussion about the effect of current cases and SIRD models in this section need only apply for untreated potential outcomes.} We denote the economic outcome of interest by $Y_{lt}$ which is the observed economic outcome for location $l$ in time period $t$. We also define treated potential outcomes, $Y_{lt}(1)$, and untreated potential outcomes, $Y_{lt}(0)$, and note that $Y_{lt} = D_l Y_{lt}(1) + (1-D_l)Y_{lt}(0)$. The target parameter in this section is given by
which is the difference between treated potential outcomes and untreated potential outcomes on average, in time period $t$, and among treated locations. As discussed above, difference-in-differences is closely related to two-way fixed effects models for untreated potential outcomes, and, in this section, we consider the following model for untreated potential outcomes
where $\tau_t$ is a common macro shock. For economic outcomes, there is clear evidence of common macroeconomic shocks which can be motivated by, for example, common information about the health risks of Covid-19 across locations. $\xi_l$ is a location-specific fixed effect allowing for time-invariant location-specific differences in economic outcomes, and $v_{lt}$ are idiosyncratic, time varying unobservables.
What is different about this model from standard DID is the term involving $I_{lt}(0)$ where $I_{lt}(0)$ is the number of Covid-19 cases in location $l$ in time period $t$ if the policy were not implemented. It is likely to be very important to include this sort of term during the pandemic as it allows for economic outcomes to depend on the local spread of cases. In particular, this allows for current cases to directly affect outcomes as well as individuals and/or firms taking more Covid-19 related precautions when the number of local cases is high.
In this section, we propose an approach that is able to deliver consistent estimates of $ATT^Y_t$ in the case when policies can have an effect on current Covid-19 cases and current Covid-19 cases can, in turn, have an effect on the outcome of interest. Throughout this section, we contrast our suggested approach with two very common DID-type approaches. First, we consider the case where a researcher compares the path of outcomes of treated locations to the path of outcomes among untreated locations without accounting for the current number of Covid-19 cases. Throughout this section, we refer to this case as “standard DID”. Second, we consider a version of DID that includes the number of cases as a regressor. Throughout this section, we refer to this case as “regression DID”. We show that both of these approaches generally deliver biased estimates of $ATT^Y_t$ under the model in (ref). In the standard DID case, biased estimates arise because the researcher does not account for current cases in a particular location having a direct effect on outcomes. In the regression DID case, biased estimates arise because the approach does not accommodate the possibility that the policy has an effect on the current number of cases (which in turn has an effect on outcomes).
In light of this discussion, we propose an alternative approach that simultaneously addresses both of these issues. We call our approach “adjusted regression DID”. Our idea is to include an adjustment term that accounts for the possibility that the policy affects the current number of cases. This adjustment term is closely related to the arguments in the previous section; in particular, we can recover an estimate of the number of active cases that a treated location would experience in a particular time period by recovering the number of active cases in untreated locations with similar pre-treatment pandemic-related characteristics.
Before stating the main result in this section, it is helpful to notice that, in the model in (ref),
where we define $\tilde{\tau}_t := (\tau_t - \tau_{t^*-1})$. Also note that $\tilde{\tau}_t$ and $\alpha$ are both identified using the untreated group (in that case, untreated potential outcomes and untreated potential active cases are observed in all time periods which implies that the parameters are identified as this amounts to a simple linear regression of $\Delta^{(t^*-1,t)} Y_t$ on $\Delta^{(t^*-1,t)} I_t$ using untreated locations).
Next, to fix ideas, under standard DID, the estimator of $ATT^Y_t$ is the sample analogue of
Likewise, under regression DID, the estimator of $ATT^Y_t$ is the sample analogue of
Including covariates in this sort of way is a common strategy\footnote{Notice that this estimand is similar in spirit, though not exactly the same, as two way fixed effects regressions that include a treatment dummy variable along with other time varying covariates. Besides the issues pointed out in this section (related to the covariates), those sorts of regressions do not generally deliver an interpretable treatment effect parameter in the case with multiple time periods and variation in treatment timing (see, for example, goodman-2021). The estimand mentioned above avoids the issues related to multiple periods and variation in treatment timing but, as we point in this section, still suffers from issues stemming from actual cases in treated locations not being equal to what cases would have been if the policy had not been implemented.} and would amount to comparing paths of outcomes for treated and untreated locations that experienced the same change in cases over time.
The next result provides an alternative identification result for $ATT^Y_t$ as well as results for the bias of standard DID and regression DID.
It is worth sketching the arguments underlying the result in (ref). To start with, notice that
The bias of standard DID arises from (incorrectly) setting $\mathbb{E}[\Delta^{(t^*-1,t)}Y_t(0) | D=1] = \mathbb{E}[\Delta^{(t^*-1,t)}Y_t | D=0]$. In general, this sort of substitution is not appropriate because the path of outcomes that treated locations would have experienced in the absence of participating in the treatment depends on the path of active cases (which is not accounted for here).
Next, based on the model in (ref), it follows from (ref) that
The bias of regression DID (that directly includes current cases as a covariate) comes from (incorrectly) setting $\mathbb{E}[\Delta^{(t^*-1,t)} I_t(0) | D=1] = \mathbb{E}[\Delta^{(t^*-1,t)} I_t | D=1]$. This strategy is also not generally appropriate because the policy can change (and is likely targeted at changing) the path of active cases.
By contrast, our approach uses the expression in (ref) for $\mathbb{E}[\Delta^{(t^*-1,t)} I_t(0)|D=1]$. This expression takes the observed path of active cases and subtracts from it the effect of the policy on active cases (which is the term $\mathbb{E}[\omega(D,\mathcal{F}_{t^*-1}) (I_t - m^I_{0,t}(\mathcal{F}_{t^*-1}))]$ and holds under the (ref)). Notice that this term is analogous to the expression for the effect of the policy on cumulative cases in (ref) in the previous section. The difference between the observed path of active cases among treated locations and the effect of the policy on active cases recovers the path of active cases that would have occurred if the policy had not been implemented. Given this expression, it can be plugged into (ref) to recover the path of untreated potential outcomes and, hence, to recover $ATT^Y_t$.
\paragraph{Estimation:} The above discussion suggests the following estimation strategy:
In the Supplementary Appendix, we provide the asymptotic distribution of our estimator of $ATT^Y_t$. The estimation procedure involves several steps, but each step is parametric and the limiting distribution of the estimate of $ATT^Y_t$ can be obtained following well-known arguments about multiple step estimation procedures that account for estimation effects of each step. In particular, the term $\mathbb{E}[\omega(D,\mathcal{F}_{t^*-1}) (I_t - m_{0,t}^I(\mathcal{F}_{t^*-1}))]$ can be handled using exactly the same arguments as in the previous section. The other steps in the estimation procedure only involve either running simple parametric regressions or directly calculating averages and are therefore straightforward to account for. As earlier, in practice, we use the multiplier bootstrap to conduct inference and discuss how to conduct uniform inference across different time periods.
In this section, we provide some Monte Carlo simulations to demonstrate the performance of the main estimation strategies considered in the paper. To begin with, we consider estimating the effect of a policy on cumulative Covid-19 cases. In order to generate the data, we consider the case where untreated potential outcomes are generated by the (ref). The values for the main parameters in the SIRD model are provided in (ref) in (ref). We also suppose that the policy has no effect on the pandemic so that all treatment effects are equal to 0.
Throughout this section, we consider the case where there are 250 locations (we vary this number in a few cases), where the probability of a location being treated is equal to 0.5, and where there are 1000 individuals in each location. We report bias, root mean squared error, and rejection probabilities for $H_0: ATT=0$ for the average effect of the policy across the first 50 post-treatment time periods (i.e., we compute event-study type estimates for 50 periods following the treatment, average them across event time to get an overall average treatment effect parameter, and compute the properties of this estimator). To implement our doubly robust estimator, we include a third order polynomial (also including all interactions) in the pre-treatment number of infected individuals and pre-treatment number of susceptible individuals both for the outcome regression and for the propensity score. Across simulations, we primarily focus on varying the timing of initial Covid-19 cases among treated and untreated locations, and on varying the treatment timing across treated and untreated locations.
{ {10pt}
}
The results for our first set of simulations are provided in (ref). The high level takeaway from this table is that the unconfoundedness approach uniformly appears to perform better than difference-in-differences. Difference-in-differences is severely biased when the timing of initial cases is different between treated and untreated locations (this is in line with our earlier discussion). The magnitude of the bias of difference-in-differences is also sensitive to the timing of the policy (this holds because the direction/magnitude of violations of parallel trends depends on the shape of pandemic related variables which are, in turn, dependent on how long ago the pandemic started). Across simulations, difference-in-differences also tends to over-reject.
On the other hand, the doubly robust unconfoundedness approach performs much better with good performance across each specification. Interestingly, even in the case where the first cases show up, on average, at the same time across treated and untreated locations (in this case, as expected, DID appears to be unbiased), the unconfoundedness approach suggested in the paper has notably smaller root mean squared error.
{ {10pt}
}
Next, we provide analogous results but for the effect of the policy on economic outcomes. For this part, we generate untreated potential outcomes according to (ref). We set $\alpha=-0.1$, $\tau_t = (50 + 20 \times t/ \mathcal{T})$, and we set $\eta|D=d \sim N(\mu_d, 1)$ where $\mu_d = 20 - 10d$ for $d \in \{0,1\}$. We also set the parameters for the pandemic related variables as in the baseline specification discussed above.
These results are provide in (ref) where we vary the timing of initial cases as well as the number of locations across simulations. As in the previous case, the approach suggested in the paper (adjusted regression DID) performs well uniformly across DGPs. By contrast, standard DID performs less well particularly in cases where the timing of initial cases is systematically different across treated and untreated locations.
To conclude the paper, we apply our approach to study the effect of shelter-in-place orders (SIPOs) on Covid-19 cases and travel. We start by using state-level data to evaluate the effect of SIPOs on Covid-19 cases. This approach is broadly similar to a number of other papers including berry-fowler-glazer-handel-macmillen-2021, bendavid-oh-bhattacharya-ioannidis-2021, courtemanche-garuccio-le-pinkston-yelowitz-2020, dave-friedson-matsuzawa-sabia-2021, dave-friedson-matsuzawa-sabia-safford-2020, hsiang-et-al-2020, and villas-sears-villas-villas-2020, among others. We consider a number of variations of difference-in-differences estimation strategies in this context as well as implementing the unconfoundedness approach discussed above. Importantly, we document substantial differences in important pre-treatment characteristics such as the number of Covid-19 cases between states that were early adopters of SIPOs, late adopters of SIPOs, or never implemented a SIPO. As emphasized above, this suggests that the parallel trends assumption underlying the DID approach is likely to be violated. We find that DID approaches can, in some cases, lead to unreasonable estimates that SIPOs increased Covid-19 cases. Another of our main findings is that the DID approach is quite sensitive to seemingly minor modifications to the specification such as using the logarithm or level of the outcome or whether or not one includes location-specific linear trends.
For any approach using state-level data, there are some important challenges. One of these is that a number of additional Covid-19 related policies such as emergency declarations, school closures, and business closures were often implemented around the same time. Moreover, there is variation across states both in terms of exactly which policies were implemented as well as the timing of these policies; we discuss some other challenges below as well. Therefore, as a second step, we use county-level data and provide difference-in-differences and unconfoundedness estimates of policy effects in states where a SIPO was implemented relative to bordering states that never implemented a SIPO among states that are similar both in terms of their pre-treatment pandemic-related characteristics (such as having experienced a similar number of Covid-19 cases) and in terms of the mix and timing of other policies that were implemented.
In general, using variations of DID, we tend to find a hard-to-interpret mix of policy effects that include estimates indicating large reductions in Covid-19 cases due to SIPOs, no effect of SIPOs, or even relatively large increases in Covid-19 cases due to SIPOs. This contrasts with our results using unconfoundedness which are more consistent across different state-level policies and where we tend to find relatively smaller reductions in Covid-19 cases due to SIPOs. Finally, we also use the county-level data to study the effects of SIPOs on travel. In line with the literature (e.g., goolsbee-syverson-2021), we tend to find relatively small reductions in travel due to SIPOs.
Our first set of results come from analyzing state-level data. We consider a period early in the pandemic --- March 10, 2020 to May 1, 2020 --- when a large number of states implemented shelter-in-place orders.
We follow dave-friedson-matsuzawa-sabia-2021 in terms of definitions of shelter-in-place orders and the timing of implementation across states. In order to facilitate estimating conditional treatment assignment probabilities, we assign states into “groups” on the basis of the timing when they adopted a SIPO. And, in particular, we assign states that adopted a SIPO within a five day window, starting on March 19, into the same group. For example, California was the first state to implement a shelter-in-place order on March 19; Illinois and New Jersey followed on March 21; New York on March 22; and Connecticut, Louisiana, Oregon, and Washington on March 23. These form a group of states that we refer to as the March 19 group. Sixteen other states adopted shelter-in-place orders between March 24 and March 28 and form the group that we refer to as the March 24 group. We include four such groups total as well as an untreated group of ten states that did not adopt a shelter-in-place order over the time period that we consider.
Next, we obtained data on state-level Covid-19 cases and testing from the Centers for Disease Control COVID Data Tracker (\url{https://covid.cdc.gov/covid-data-tracker/}). We also use 2019 state-level populations from the Census Bureau. In order to deal with heterogeneity in terms of state populations, we use versions of pandemic related variables in terms of their number per thousand individuals in a particular state (e.g., cumulative cases per thousand individuals). In terms of the SIRD model, this amounts to dividing all variables by $N_l$ and multiplying by one thousand; this transformation is compatible with the SIRD model. We construct the current number of active Covid-19 cases (and therefore contagious individuals) as the total number of newly reported cases over the past seven days; as for the other pandemic related variables, we use the number of current cases per thousand individuals in a state. Finally, we use travel data from Google's Covid-19 Community Mobility Report (\url{https://www.google.com/covid19/mobility}). We focus on state-level retail and recreation travel (these are aggregated together) which is reported as a percentage change relative to pre-Covid travel baselines.
{ {16pt}
}
Summary statistics for the data that we use are provided in (ref). There are some things that are immediately notable from the summary statistics. First, early in the pandemic, the number of cases were substantially different for states that adopted shelter-in-place orders earlier relative to states that adopted them later or that did not adopt them at all. This immediately suggests that it will be challenging for DID to perform well at evaluating the effect of shelter-in-place orders on the number of cases. In addition, notice that early treated states (particularly, the March 19 group) experienced very large increases in their number of cases relative to later- and never-adopters of the policy. Finally, the second panel of the table shows changes in retail and recreation trips. The most notable feature of this part of the table is that there were large decreases in travel across all states regardless of their shelter-in-place policies.
Our first set of results come from using state-level data and several different types of difference-in-differences estimation strategies. In (ref), we use two-way fixed effects (TWFE) event study regressions of the form
where $D_{lt}^e$ is a binary variable that is equal to 1 for location $l$ in time period $t$ if that location has been treated for exactly $e$ periods in period $t$ and is otherwise equal to 0. These sorts of regressions have been widely used in work evaluating the effects of Covid-19 related policies. The panels in (ref) differ along two dimensions: first, whether the outcome is the logarithm or the level of the number of Covid-19 cases; and second, whether or not the specification additionally includes a location-specific linear time trend. All of these specifications are common in the literature, and, in particular, the specification where the outcomes is in logarithms and includes a location-specific linear trend is a main specification in dave-friedson-matsuzawa-sabia-2021.
(ref) highlights that the qualitative conclusion as to whether or not SIPOs affected the number of Covid-19 cases is highly sensitive to functional form assumptions made by the researcher. First, when the outcome is the logarithm of the number of Covid-19 cases and the specification includes a location-specific linear time trend, then the estimates indicate a large reduction in Covid-19 cases due to shelter-in-place policies (see panel (d) of (ref)). The results in panel (c), which include also include a linear trend but where the outcome is the level of the number of cases, also indicate that SIPOs may have reduced the number of Covid-19 cases (these results are closer to zero and only marginally statistically significant). On the other hand, the results in panels (a) and (b) of (ref), neither of which include a location-specific linear trend, are much different and suggest that SIPOs led to an increase in the number of Covid-19 cases. It seems very hard to rationalize these sorts of results; in particular, it would seem that shelter-in-place orders could either have no effect or decrease Covid-19 cases, but it is difficult to see how they could increase cases. However, even from the summary statistics, one can see that DID estimates are likely to be positive as early policy adopters were tending to experience larger increases in cases. In our view, a better explanation for these results is that Covid-19 was more prevalent earlier in locations that were early policy adopters and that the strong, early exponential growth of Covid-19 cases overwhelms any reduction in the infection rate due to the policy. It is exactly in this case where DID would be susceptible to attributing faster growth in Covid-19 cases to the policy rather than to simply a larger number of early cases in treated locations.
Another noteworthy feature of (ref) is that none of the four specifications used in the figure result in either large or statistically significant violations of the underlying parallel trends assumption in pre-treatment periods (i.e., estimates different from 0 for $e<0$).\footnote{To be more specific, there are 23 pre-treatment estimates reported in each panel. No pre-treatment estimates are statistically significant at the 5% level in panels (a), (b), or (c). One estimate is statistically significant at the 5% level in panel (d); this occurs for $e=-3$ and the p-value is 0.037.} This suggests that it is not possible to use a purely data-driven/reduced form model selection procedure to infer that one specification is likely to perform better than the others, despite the choice of the model largely driving the results.\footnote{In one sense, it is somewhat surprising that that parallel trends cannot be rejected in pre-treatment periods for any of the estimation strategies considered here. The (ref) implies that, for all of the models considered here, the parallel trends assumption is also violated in pre-treatment periods. Instead, the results here indicate that we are not able to detect violations of parallel trends in pre-treatment periods (even though parallel trends is probably actually violated). That we are not able to detect violations of parallel trends is not altogether surprising as pre-tests are often under-powered (roth-2022), and this is likely to be an acute issue early in the pandemic when the number of Covid-19 cases is very small in most pre-treatment periods.} Finally, we should emphasize that none of the TWFE event study specifications discussed here are compatible with the SIRD model that we have discussed in the paper.
One possible concern with the previous results is related to limitations of two-way fixed effects regressions when there is variation in treatment timing and treatment effect heterogeneity (see, for example, chaisemartin-dhaultfoeuille-2020,goodman-2021,sun-abraham-2021,borusyak-jaravel-spiess-2021). There is certainly variation in treatment timing for SIPOs and heterogeneous effects (which would include things like variation in policy effects according the time period when the policy is adopted or effects that vary with length of exposure to the policy, among others) are very likely as well. These issues are considered in dave-friedson-matsuzawa-sabia-2021 and discussed extensively in goodman-marcus-2020. In order to address these possible concerns, (ref) provides estimates using “heterogeneity robust” DID estimation strategies from callaway-santanna-2021,gardner-2021. The setup for these results mirrors the setup from (ref) as the panels differ based on whether the outcome is in levels or logarithms and by whether or not the results include location-specific linear trends. The results in panels (a) and (b) that do not include a linear time trend are broadly similar to the TWFE event study regressions from above but still suggest that SIPOs increased Covid-19 cases. On the other hand, the results that include location-specific linear trends are notably different from the results in (ref). Using heterogeneity robust versions of DID estimation strategies changes the sign of the estimates and results in estimates that SIPOs increased Covid-19 cases. This is disconcerting as these estimation strategies provide a number of advantages relative to the TWFE event studies; however, as discussed above, it seems hard to understand how SIPOs could lead to an increase in Covid-19 cases. As before, we do not find large or any statistically significant estimates of Covid-19 policies in pre-treatment periods.\footnote{Depending on the specification, there are either 23 or 24 pre-treatment estimates in each panel of (ref). Across all four panels, none of these estimates are statistically different from 0 at the 5% level.}
Next, we turn to estimates using the approach based on unconfoundedness. As discussed above, these estimates require estimating a model for treatment participation and an outcome regression model. For both of these models, we include a cubic polynomial in the current number of cases per 1000 people in a state (we define the number of current cases as the change in cumulative cases over the previous seven days); we do not include the number of susceptible individuals as this is very close to the full population in all states during the period early in the pandemic that we consider; we additionally include a dummy variable for region of the country so that states are compared to other states in the same region as well as the logarithm of the state's population; finally, we include the 7-day lag of the cumulative cases, the 7-day lag of the number of tests run per 1000 people in the state and the change in tests from the pre-treatment period to the current period in order to control for the possibility that some states were detecting Covid-19 cases better than others. When we estimate the propensity score, we find evidence that the overlap condition is violated indicating that there are a substantial number of states that do not have reasonable comparisons among never-treated and late-treated states. From this step, we drop 14 states from our analysis.\footnote{The states that we drop are Alabama, Arizona, California, Florida, Georgia, Mississippi, Missouri, New York, North Carolina, Pennsylvania, South Carolina, Texas, Virginia, and Washington.} The omitted states include states such as New York and California which had large numbers of early Covid-19 cases and were early adopters of SIPOs. It is, therefore, hard to find reasonable comparisons for these states. Another large number of states from the South are dropped due to the timing of their policies being very similar which makes it challenging to find reasonable comparison states when region is included as a conditioning variable.
These results are provided in (ref). Using the unconfoundedness approach, we estimate relatively small and statistically insignificant effects of SIPOs.
There are several complications with any state-level analysis that are worth noting. First, there are several other Covid-19 policies that were being implemented in a number of states around the same time. Several recent papers have pointed out challenges with “controlling for” other policies in linear models like the one in (ref) (e.g., chaisemartin-dhaultfoeuille-2021b,goldsmith-hull-kolesar-2021) especially in a framework (like the current one) with treatment effect heterogeneity. Moreover, using state-level data, the number of observations is very small, and, relatedly, it is challenging to find treated and untreated states that are similar enough to each other to reasonably interpret differences in outcomes as being due to the policy.\footnote{To give a simple specific example, population density is likely to be an important determinant of Covid-19 transmission rates. One of the largest states that did not implement a SIPO was Arkansas, and a natural strategy is to try to compare Covid-19 cases in Arkansas to, say, Missouri which implemented a SIPO on April 6. At the state-level, though, the population density of Arkansas and Missouri is much different. Missouri's population is double Arkansas's population, and it has two major cities, Kansas City and St.\ Louis, while Arkansas does not have a major city. This suggests that, at the state-level, Arkansas may not be very useful for delivering counterfactual outcomes for Missouri had it not implemented the policy.} On the other hand, at the county-level, there are often a large number of counties that are located in neighboring states and have similar characteristics, including both pre-treatment pandemic-related characteristics as well as population, median income, or demographic characteristics.\footnote{There are some potential drawbacks to using county-level data. One issue is that, by restricting comparisons to counties with similar characteristics, it potentially changes the target parameter. There are other issues related to inference such as spatial correlations and clustering standard errors at the county-level rather than the state-level as well. We discuss these issues in more detail in the Supplementary Appendix.}
Our strategy for the remainder of this section is to use county-level data and compare counties in states that implemented a SIPO to counties in states that did not implement a SIPO while (i) having a similar mix of other policies (both in terms of which types of policies were implemented and their timing) and (ii) making tight comparisons between counties with similar characteristics. We focus on Arkansas and Iowa which were two states that did not implement a SIPO; we compare them, in turn, to their surrounding states that implemented SIPOs.\footnote{haynes-kulkarnia-li-siddique-2022 similarly make comparisons between counties in states that implemented a SIPO to counties in states that did not. Compared to our approach, they use TWFE event study regressions and find mixed results with estimates sometimes indicating reductions in Covid-19 cases, sometimes indicating no effect, and sometimes indicating increases in Covid-19 cases.} Among their surrounding states, Nebraska, Oklahoma, and South Dakota also did not implement a SIPO. This provides an opportunity to estimate placebo policy effects; that is, to use different estimation strategies in a case where we know that the “policy effects” should be equal to 0, which, therefore, provides a way to assess the performance of different estimation strategies.
As for the state-level data, we obtain the number of county-level Covid-19 cases from the CDC COVID Data Tracker and 2019 county-level population from the Census Bureau. Some of our results below use the number of Covid-19 tests in a particular county which also comes from the CDC COVID Data Tracker.
We initially include all states that border either Arkansas or Iowa with two exceptions. We exclude Texas which shares a small border with Arkansas, and we include Kansas which “almost” borders both Arkansas and Iowa. Thus, we start with county-level data for thirteen states; these are listed in (ref).
A major concern for our application is the timing of other Covid-19 related policies across states. (ref) provides the date when each state declared an emergency, closed schools, implemented a SIPO, closed non-essential businesses, or imposed gathering restrictions. Overall, Arkansas tended to implement other policies with very similar timing as all of its surrounding states especially with respect to the timing of the emergency declaration and school closures. There is more variation with respect to non-essential business closures and gathering restrictions though we note that these policies are less clearly defined than the other main Covid-related policies. Our interpretation is that it is reasonable to consider the mix and timing of other policies as being quite similar to all of its neighboring states. The timing of Iowa's policies are somewhat more different from its surrounding states. The primary difference is that Iowa closed its schools somewhat later than surrounding states with the exception of Nebraska. This arguably suggests that the results below that use Iowa as the comparison state are somewhat less credible (with a viable alternative explanation that differences in school closure policy could be at least partially driving the results).
A second major concern is being able to find counties with similar pre-SIPO pandemic related characteristics. For each pair of states, we use the same model for the outcome regression and propensity score as for the state-level data that includes a cubic polynomial in the current number of cases, the 7-day lag of cumulative cases, the 7-day lag and change in the number of Covid-19 tests in the county, and the logarithm of county population. In order to enforce the overlap condition between treated and untreated states, we drop treated counties with an estimated propensity score greater than 0.95.\footnote{In practice, for our main results, we do not drop many counties from this procedure. When we compare Missouri to Arkansas, we do not drop any counties using this criteria. When we compare Minnesota to Iowa, we drop Ramsey County which is the county where St.\ Paul is located. As discussed above, dropping counties this way does change the target parameter from being the $ATT$ to being an $ATT$-type parameter that is local to the region of common support. And, for example, our approach tends to drop urban counties relative to suburban or rural counties; if SIPOs tended to reduce Covid-19 cases more in urban counties than in other counties, then our estimation procedure would tend to result in smaller in magnitude estimates of policy effects.} With the remaining set of counties, including both treated and untreated counties, we compute the propensity score weights from (ref) using the same set of covariates as mentioned above. Given these weights, we compute the effective sample size for untreated states, which is defined as $\sum_{l \in \mathcal{U}} w_l^2 / \Big(\sum_{l \in \mathcal{U}} w_l^2\Big)$ where $\mathcal{U}$ denotes the set of untreated county (see, for example, kish-1965,chattopadhyay-zubizarreta-2022,shook-hudgens-2022). Then, we exclude combinations of states where the effective sample size is less than or equal to 25 counties for either the treated or untreated state. Following this procedure, we end up with four combinations of states that satisfy this criteria. Two of these, (i) Oklahoma relative to Arkansas and (ii) South Dakota relative to Iowa, arise for states that did not implement a SIPO. We estimate placebo policy effects for these states. The other two combinations of states are (iii) Missouri relative to Arkansas and (iv) Minnesota relative to Iowa. These states provide a chance to estimate effects of policies that were actually implemented.
We start by providing placebo policy effect estimates for Oklahoma and South Dakota --- neither of which implemented a SIPO. The reason for starting here is that this serves as a good way to evaluate estimation strategies; policy effects should, in principle, be equal to 0 in all time periods. In order to facilitate comparisons to other results, for states that did not implement a SIPO, we use a “placebo policy” date of April 1. These results are provided in (ref). Panels (a) and (b) provide estimates using the unconfoundedness approach. These estimates are generally small and not statistically significantly different from zero. Estimates using a difference-in-differences approach are provided in panels (c) and (d). For Oklahoma, none of the post-policy estimates are statistically different from zero though the standard errors are notably larger using DID than for the unconfoundedness approach. More notably, using DID, we spuriously estimate that the placebo policy in South Dakota reduced the number of Covid-19 cases in South Dakota. These results suggest that, at least for these two policies, the unconfoundedness approach performs better than DID. It is also worth mentioning that, due to the relatively similar pre-pandemic characteristics of Oklahoma to Arkansas and South Dakota to Iowa, this is still a relatively favorable setting for DID estimation strategies.
Next, we move to estimating effects of policies that were actually implemented in Missouri (using counties from Arkansas as the comparison) and in Minnesota (using counties from Iowa as the comparison). The results under unconfoundedness are provided in (ref). For Missouri, we do not find evidence of policy effects on Covid-19 cases. On the other hand, for Minnesota, we estimate that its SIPO decreased Covid-19 cases at least in some periods. For example, on April 25, we estimate that Minnesota's SIPO reduced Covid-19 cases by about 0.46 cases per 1000 people in the state. In results provided in the Supplementary Appendix, using DID, we estimate a large reduction in Covid-19 cases in Missouri due to the policy though the standard errors are much larger than for the results based on unconfoundedness and not statistically different from 0. For Minnesota, the DID estimates are quite similar to the unconfoundedness results presented here.
In the Supplementary Appendix, we provide additional estimates for all combinations of states discussed in this section. Of the remaining 8 combinations of states that (i) implemented a policy and (ii) exhibit similar pre-treatment paths of Covid-19 cases, using the same DID estimation strategy used in this section, two of the estimates are positive and statistically significant, two of the estimates are negative and statistically significant, and four of the estimates are not statistically different from zero. Thus, like the case with state-level data above, using DID can lead to hard-to-explain positive estimates of SIPOs on Covid-19 cases and, more generally, estimates that seem inconsistent with each other. Arguably, these differences could be explained by heterogeneous effects of policies implemented in different states, though, in our view, the large magnitude of the differences in estimated policy effects cuts against heterogeneous policy effects being a full explanation of the differences. Instead, a better explanation seems to be that these differences are driven by pre-treatment differences in the state of the pandemic between counties in states that implemented SIPOs relative to counties in neighboring states that did not.
Finally, we consider the effect of SIPOs on travel. We focus on the percentage change in retail and recreation travel from a pre-Covid baseline. In the main text, we focus on the effect of the policy on travel in Missouri using Arkansas as the comparison state. These results are available in (ref). The results in Panel (a) come from a standard DID approach that implicitly imposes that Covid-19 cases do not directly affect travel; the results in Panel (b) come from the regression DID approach that includes current cases as a covariate but not accounting for the possibility that SIPOs could have affected the number of cases directly; and the results in Panel (c) use the adjusted regression DID approach proposed in the current paper that allows for the policy to have had an effect on Covid-19 cases.
In this case, the estimates are more broadly similar than they were for cumulative Covid-19 cases. Standard DID estimates indicate a relatively small but persistent negative effect of SIPOs on retail and recreation travel. Using standard DID, the overall estimate of the effect of SIPOs across post-treatment time periods is -1.30 (p-value: 0.16); i.e., across the first several weeks of the SIPO, the policy reduced travel by about 1.3 percent relative to what travel would have been if the policy had not been implemented. The point estimates from regression DID (these estimates just include observed current cases as a covariate) are somewhat larger in magnitude; in this case, the estimated overall effect of SIPOs on travel is -1.63 (p-value: 0.08). Finally, the adjusted regression DID overall estimated effect of SIPOs is -1.25 (p-value: 0.19) when we allow for the policy to have had an effect on cases. In the Supplementary Appendix, we provide analogous results using the state-level data from earlier in this section. Those results are broadly similar to the ones presented here though we estimate somewhat larger-in-magnitude effects (around a 3 percent reduction in travel due to SIPOs) that do not vary much by estimation strategy and are somewhat more precisely estimated than the results presented here.
Our results, especially those about the effect of SIPOs on the number of Covid-19 cases, are substantially different from existing estimates, and it is worth making a few additional comments. First, our results on the effect of SIPOs on travel are more similar to existing estimates (e.g., goolsbee-syverson-2021). Since the primary channel through which SIPOs would likely reduce Covid-19 cases is through reducing travel/contact with other individuals, it seems reasonable to simultaneously estimate relatively small effects of SIPOs on travel coinciding with small effects of SIPOs on the number of Covid-19 cases; but harder to rationalize SIPOs strongly decreasing Covid-19 cases while having only a small effect on travel.
That being said, we hesitate to interpret our results as providing strong evidence that SIPOs did not reduce Covid-19 cases. For one thing, the interpretation of our treatment effect parameters is somewhat subtle. Untreated potential outcomes here do not correspond to particular locations not reacting at all to Covid-19 but rather to outcomes that would have occurred if a state had not implemented the policy (but other things about the state remained the same). This can be seen to be clearly relevant from the summary statistics in (ref) where all states (not just those who implemented the policy) were experiencing massive decreases in travel over the period that we consider. Importantly, this indicates that our results are not at all saying that staying at home did not have an effect on Covid-19 cases. Second, the statistical power of all of the estimation procedures considered above is relatively low (this is true both for the unconfoundedness approach and the DID approaches) and leads to generally wide confidence intervals on estimated policy effects. And even though a number of our estimates are not statistically significant, many of them are large enough to be compatible with a wide variety of possible effects of SIPOs on Covid-19 cases.\footnote{To give an example, in our results for Missouri, which is a state for which we do not estimate a statistically significant policy effects on Covid-19 cases in post-treatment periods (see (ref)), the lower end of a 90% confidence interval for our estimated effect of the policy on Covid-19 cases 21 days after the policy was implemented is that it reduced Covid-19 cases by about 0.13 per 1000 people. If we divide this by the estimated number of Covid-19 cases that treated locations would have experienced in the same period if they had not implemented the policy (i.e., $\mathbb{E}[C_t(0)|D=1]$ in (ref) which is identified and can be recovered in our setup), we would estimate that SIPOs decreased Covid-19 cases by 30%. In other words, our estimates do not necessarily rule out the possibility that SIPOs may have had quite large effects on Covid-19 cases.} Instead, we interpret the results from the application as indicating that these sorts of policies are likely to be very challenging to precisely evaluate due to both data limitations as well as trying to deal with a highly nonlinear outcome and a policy that was adopted at different times by states whose exposure to the pandemic also varied widely and was correlated with the timing of the policy being adopted.
\FloatBarrier
In this paper, we have considered several different policy evaluation strategies and how compatible they are with a leading epidemic model. For identifying the direct effects of policies on the number of Covid-19 cases, our results suggest that strategies based on unconfoundedness type assumptions are likely to perform better than difference-in-differences type strategies due to the highly nonlinear nature of the spread of Covid-19.
Our second main set of results were about evaluating the effects of policies on other economic outcomes when (i) the policy can affect the number of Covid-19 cases and (ii) the number of Covid-19 cases can have a direct effect on the economic outcome of interest. For this case, we also showed that two of the most common ways to evaluate these policies (difference-in-differences directly or including the number of cases as a covariate in a DID setup) do not generally deliver an average effect of the policy. We proposed an alternative estimator that is valid in this case.
We applied our approach to study the effects of shelter-in-place orders on Covid-19 cases and travel early in the pandemic. We showed that our theoretical arguments were indeed relevant in this context and led to notably different estimates (particularly for the number of Covid-19 cases) relative to the most common approaches used in applications.
There remain a number of interesting possible extensions to this work, and we conclude by mentioning two of them. First, evaluating the effects of various policies (especially early in the pandemic) is complicated by limited and nonrandom Covid-19 testing during the early part of the pandemic (see, for example, callaway-li-2021,manski-molinari-2020), and it would be interesting to extend our results along these dimensions. Second, another common policy evaluation approach in the context of Covid-19 related policies is the synthetic control method (examples include cho-2020,dave-friedson-matsuzawa-mcnichols-sabia-2020,friedson-mchnichols-sabia-dave-2021,mitze-kosfeld-rode-walde-2020). It appears that, most often, a researcher's decision between difference-in-differences or synthetic controls is driven by whether the number of treated locations is large or small. However, it is less clear under what conditions synthetic control approaches are compatible with the sorts of epidemiological models that we considered in the current paper.
\singlespacing
\printbibliography