Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,578 characters · 13 sections · 37 citation commands
Lockdown effects in US states: an artificial counterfactual approach
The evolution of the Covid-19 has been posing several challenges to policymakers. Decisions have to be made in a timely fashion, without much undisputed evidence to support them. Being a new disease, and despite the enormous research effort to understand it, estimates of the transmission, recovery and death rates remain uncertain. Nevertheless, these are key pieces of information to assess potential pressures on the health system capacity, as well as the need of a lockdown policy and its intensity if implemented.
Not surprisingly, similar regions have implemented different strategies regarding lockdowns. The leading example in the media is the looser social distancing policy in Sweden versus strict policies in its Scandinavian peers. By informally comparing the evolution of the pandemics in Sweden and Denmark (or Norway), many commentators argue that several Covid-19 cases and deaths in Sweden would be avoided in the short-run were a strict lockdown in place.\footnote{juranek2020 explore this case study to construct proper counterfactuals. Hospitalizations and ICU patients would be much higher in Denmark and Norway were Sweden's more lenient measures adopted. andersen2020 argue that despite the divergence in deaths in Sweden relative to Denmark, at least in the very short-run, there was not a large difference in the aggregate spending drop due to the stricter lockdown strategy in Denmark.}
Aiming to provide a quantitative assessment on the short-run effects of lockdowns, this paper takes this exercise seriously in the context of US states. Given that the timing US states adopted lockdown policies differs among them, we adopt techniques based on synthetic control (SC) approach of aAjG2003 and aAaDjH2010 to assess the impact of lockdowns on the short-run evolution of the number of cases (and deaths) in the treated US states.\footnote{Throughout the main text in this paper, we focus on the number of cases. Results concerning the number of deaths are relegated to the Appendix. The timing of most lockdowns was soon enough such that there is not enough in-sample observations of deaths to apply the synthetic control method, so we use an alternative methodology we explain below.} More specifically, we consider an extension of the original SC method called Artificial Counterfactual (ArCo) which was put forward by CarvalhoMasiniMedeiros2018ArCo. Due to the nonstationary nature of the data, the correction of rMmcM2019 is necessary.
It is hard to downplay the importance of finding out the effects of lockdown policies, especially now that several countries are experiencing even harsher second waves of Covid-19. Our results point to a substantial short-run taming of the cumulative number cases due to the adoption of lockdown policies. On average, for treated states, the counterfactual accumulated number of cases, according to the method adopted here, would be two times larger were lockdown policies not implemented.
The decision to implement a lockdown policy is not taken out of the blue. It might complement or substitute other types of containment policies implemented in control or treated states, such as, for example, mask mandates. Hence, in principle, our estimates are better interpreted as capturing the effect of a “combo" of policies that include lockdowns, relative to another “combo" of policies that do not include them. Nonetheless, to address some confounding effects, we use a simple causal model similar to vChKpS2021 to claim that a sizable part of the estimated effects is arguably attributable to lockdown policies.
A key feature of our approach is that it is purely data-driven. In the beginning of the crisis, the majority of papers written by economists to evaluate the effectiveness of lockdowns relied on epidemiological models for analysis, including the most recent ones that incorporate behavioral responses.\footnote{Descriptions of epidemiological models and simulations concerning the evolution of the Covid-19 pandemic can be found in atkeson2020will and berger2020seir. alvarez2020simple, bethune2020covid, callum2020coronavirus, eichenbaum2020macroeconomics, among many others, incorporate behavioral responses and evaluate several containment policies.} These models are hard to discipline quantitatively. Many calibrated parameters remain uncertain,\footnote{See, for example, atkeson2020death on the uncertainty regarding estimates of the fatality rate.} and models that incorporate behavioral responses need time to mature and agree on a reliable set of ingredients and moments to be matched.
Model-free approaches like ours or medeiros2020forecast should complement policy discussions or forecasting exercises based on those models, especially from a quantitative point of view. There are related papers using state or county level US data.\footnote{There are also related papers for other countries. For example, fang2020china for China.} At least one of them, california, uses a synthetic control approach but it is restricted solely to California. Other papers, such as brzezinski2020, dave and sears2020stayinghome, use variations in the timing of statewide adoption of containment policies, and difference-in-differences models to document substantial reductions in mobility and improvements of health outcomes. The key identification assumption in these papers is that variations in the timing are random after controlling for covariates. brzezinski2020 also consider an instrumental-variable approach. fowler2020stayinghome and barrot2020 follow similar empirical strategies but at county level, and also find substantial reductions in cases and fatalities in counties that adopted stay-at-home orders and state-mandated business closures, respectively. Our analysis, that rests on alternative identification assumption and method, should be seen as complementary.
The paper is organized as follows. Section (ref) describes the data, while Section (ref) presents the empirical strategy. The results are discussed in Section (ref). In Section (ref), we discuss our estimates and address some confounding effects. Finally, Section (ref) concludes the paper. Additional results are included in the Appendix.
Data on Covid-19 (confirmed) cases are obtained from the repository at the Johns Hopkins University Center for Systems Science and Engineering (JHU CSSE). We consider the cumulative cases for a subset of the 50 US states and the District of Columbia. Instead of using the chronological time across the states, we consider the epidemiological time, which means that the day one in a given state is the day that the first Covid-19 case was confirmed there.
The econometric approach adopted here relies on the fact that some states adopted a lockdown strategy (the treatment), whereas others did not adopt social distancing measures (control group) and are used to construct the counterfactual.\footnote{The timing of those policies at each state were obtained, and double checked, in several press articles, e.g., https://www.businessinsider.com/us-map-stay-at-home-orders-lockdowns-2020-3 and https://www.nytimes.com/interactive/2020/us/coronavirus-stay-at-home-order.html.} Lockdown strategies include a mix of state-wide non-pharmaceutical measures aiming to limit social interactions, such as restrictions on non-essential activities and requirements that residents stay at home.
In this section, we describe how we assign states to control and treatment groups, and then, describe the method used to construct the counterfactuals.
Aiming to balance control and treatment states, and at the same time obtain enough observations to estimate properly the model before the lockdown policy was implemented, we divide US states into three groups.
For a state to be included in the analysis, a state-wide lockdown policy must be established at least twenty days after the first case. We assume that whenever an individual becomes infected, it takes an average of ten days to show up as a confirmed case in the statistics.\footnote{This assumption is motivated by the incubation period of the virus. According to the World Health Organization, the “[...] the incubation period for COVID-19, which is the time between exposure to the virus (becoming infected) and symptom onset, is on average 5-6 days, however can be up to 14 days." See https://www.who.int/docs/default-source/coronaviruse/situation-reports/20200402-sitrep-73-covid-19.pdf.} Hence, the in-sample period used to estimate the synthetic control (“before" the lockdown policy) for each treated state (to be defined below) is the number of days between the tenth day after the first confirmed case and the tenth day after the lockdown strategy was implemented. We choose to start the in-sample from the tenth day as a way to smooth the initial volatility of the data.
We adopt a criteria that a state must have at least twenty observations in the in-sample period to be included in the analysis. This criteria excludes states that adopted a state-wide lockdown strategy too early, such as Connecticut, New Jersey, Ohio, among others. These are the unmarked states in Table (ref), which reports the dates of the first case and lockdown policy, as well the difference in days between them, and also helps visualize the three groups of states.
The remaining states must be divided into treated and control groups. The idea is to find a synthetic control for each of the treated states. The group of potential controls should consist of states that adopted a lockdown policy too late (or never adopted), such that counterfactuals are not contaminated by lockdown policies implemented in those states. At the same time, and for a similar reasoning, the lockdown strategies adopted in treated states must be in place during the period of analysis.\footnote{In the Appendix (ref), Table (ref) shows the reopen dates for the treated states.}
Fortunately, there are horizons that can balance both goals: enough states to build the synthetic controls and a relative extensive period to construct the counterfactuals. In particular, we restrict the analysis up to the 58th epidemiological day. This figure accommodates at least ten control states to build the synthetic controls,\footnote{That is, to be in the control group, whenever a lockdown policy was implemented in a given control state, it was implemented at least 48th days after the first epidemiological day. Given the aforementioned assumption, its effects on Covid-19 confirmed cases only show up in the statistics ten days later, on average.} at the same time it maximizes the out-of-sample days to run the counterfactuals. In this sense, our analysis concerns the very short-run impact of lockdowns, up to nearly three weeks.
The treated states are marked in blue in Table (ref), and include twenty states: Alabama, Colorado, Florida, Georgia, Kansas, Kentucky, Maine, Maryland, Mississippi, Missouri, Nevada, New Hampshire, New York, North Carolina, Oregon, Pennsylvania, Rhode Island, South Carolina, Tennessee, and Texas. The potential control states are marked in red, and include ten states: Arizona, Arkansas, California, Illinois, Iowa, Massachusetts, Nebraska, North Dakota, South Dakota, and Washington. Nonetheless, due to the lack of variation within the in-sample period, we exclude four states from this control pool as we explain below.
Importantly, Oklahoma, Utah and Wyoming only implemented partial lockdowns (not reported in the table). Therefore, they are hard to classify as either treated or control states. We opt to exclude them from the analysis.
Figure (ref) illustrates the empirical strategy, which is formalized in the next subsection. It plots the evolution of (log) cumulative cases along the epidemiological time. The first vertical dashed line represents the tenth day after the first confirmed case. The in-sample period is represented in between the first and second vertical dashed lines, which mark the tenth day and the following twenty days, respectively. Similarly, the out-of-sample period is in between the second and third vertical dashed lines, which mark the 31th and 58th epidemiological day, respectively.
Blue lines represent the treated states, whereas the red ones the potential control states. The turning points from blue full- to dashed-lines represent the days lockdowns were implemented (plus ten days) in treated states. Note that New York is clearly an outlier among the treated states, exhibiting a huge amount of cases (more on that below). We use the red lines to build synthetic controls for each full blue-line up to the turning point, and then construct counterfactuals by simulating the synthetic controls forward up to the 58th day. The idea is to compare counterfactuals with the blue dashed-lines that capture actual cases, and obtain the effect of lockdowns.\footnote{A lockdown policy might complement or substitute other types of containment policies implemented in control or treated states. Hence, our empirical strategy is arguably capturing the effect of a “combo" of policies that include lockdowns, relative to another “combo" of policies that do not include them. We further discuss below how to disentangle the role of lockdown policies from alternative policies.}
As Figure (ref) highlights, some states display lack of variation within the in-sample period. Just to give an example, Washington had had only one confirmed case for the first 36 days since its first confirmed Covid-19 infection. Therefore, we exclude it from the control group. For similar reasons, we also exclude Arizona, Illinois, and Massachusetts from the control pool. The analysis ended up relying on six control states.
We propose a two-step approach using the artificial counterfactual (ArCo) method introduced by CarvalhoMasiniMedeiros2018ArCo with the correction of rMmcM2019 to estimate the number of cases for each US state.
Let $t=10,11,\ldots,58$ represents the number of days after the first confirmed case of Covid-19 in a given state. Define $y_{t}$ as the natural logarithm of the number of confirmed cases $t$ days after the outbreak of Covid-19 in this specific treated state, and $\boldsymbol{x}_t$ contains the logarithm of the number of cases for $p$ control states $t$ days after the first case, as well as a logarithmic trend, $\log(t)$. The inclusion of the trend is important to capture the shape of the curve.
The model is estimated as follows. We use the weighted least absolute and shrinkage operator (WLASSO) as described in rMmcM2019 to select the control states that will be used to estimate counterfactuals. The goal of the WLASSO is to balance the trade-off between bias and variance and is an useful tool to select the relevant peers in an environment with very few data points. The estimator is given as:
where $\kappa_j=|x_{j,L}|$, $j=1,\ldots,p-1$, and $\kappa_p=1$. $L$ is, for each state, the number of days from the first reported case until the lockdown plus ten extra days, and $\lambda>0$ is the penalty parameter which is selected by the Bayesian Information Criterion (BIC), in accordance with MedeirosMendes2016L1. The weight correction in the WLASSO is necessary in order to control for the nonstationarity of the data; see rMmcM2019 for a detailed discussion.
The counterfactual for $t=L+1,\ldots,$ is computed as $\widehat{y}_t=\boldsymbol{x}_t'\widehat{\boldsymbol{\omega}}$. We also report 95% confidence intervals based on the resampling procedure proposed in rMmcM2019.
We are interested in examining the effects of lockdown policies not only on the number of cases, but also on the number of deaths. However, we cannot implement the strategy described above because there is not enough variation in deaths for the in-sample period. Some states, for instance, implemented a state-wide lockdown policies before the first confirmed death.
Thus, we propose an alternative method. We consider a counterfactual state for the number of deaths based on the counterfactual estimated for the number of cases. This is not straightforward as in the traditional synthetic control method because the ArCo methodology described above includes an intercept in the estimation, which is measured in the log of the number of cases, and the counterfactual is not only a convex combination of other states. Intuitively, the methodology described above chooses a combination of states that is at a fixed distance from the treated unit at the in-sample period and not a convex combination of states that matches exactly the actual number of cases. The intercept controls for all time-invariant characteristics that define the counterfactual.
Then, we proceed as follows. Let $y_{st}$ be the number of accumulated deaths in state $s$ at the day $t$. Also, let $\boldsymbol{\beta}^s$ be the vector of estimated coefficients for the state $s$ as in expression (ref) above and used to construct the counterfactual for cases. In addition, let $\boldsymbol{y}_t$ be a vector of the number of deaths for all states in the control pool at time $t$. We define the counterfactual number of deaths in that state as
where $\bar{t}$ is the day that state $s$ implemented the lockdown policies. That is, we maintain the weights estimated above and adjust the intercept so that the counterfactual series for deaths matches the number of actual observed deaths in the beginning of the quarantine. For the sake of exposition, we relegate the results on cumulative deaths to Appendix (ref).
To illustrate how the method works, Figure (ref) presents the ArCo counterfactuals for the states of Alabama, Colorado, and Maine. The timing of the policy intervention ($T_0 + 10$) corresponds to the lockdown date plus ten days. The gray area represents 95% confidence intervals.
The counterfactual analysis makes it clear the importance of lockdown policies in mitigating the acceleration of the number of Covid-19 confirmed cases in the treated states. As shown in Figure (ref), for example, our results point to a substantial increase in the number of cases in Alabama if it had not adopted an early lockdown. Similarly, Figures (ref) and (ref) reveal the same behavior for the cumulative curves in the other selected states. Counterfactuals are constructed with the estimated weights and cumulative cases of the six states that compose the control group. These weights are reported in Table (ref) in Appendix (ref). In Appendix (ref), we present similar counterfactual plots for the remaining treated states. Similar results apply for most of the treated states.
In order to assure that the proposed methodology is producing proper counterfactual analysis, we generate placebo results by producing a “synthetic control" for each control state using the remaining control states as donor pool. Results are displayed in Figure (ref), which shows the ratio of the estimated counterfactual cumulative cases to the actual ones for treated states except New York (black lines), and non-treated states (red lines). We assume that the epidemiological day of the placebo intervention is $T_0 = 36$, marked by the vertical dashed line, which is the median (and the mean) timing of the policy interventions in the treated states.
It is reassuring that for half of the placebo counterfactuals, these ratios fluctuate around one, whereas for the majority of treated states ratios grew above one at some point (likely around the actual timing of policy intervention). The latter result means that lockdown policies were effective to tame the spread of the virus, whereas the former suggests that results are not driven by chance.
Regarding South Dakota, the only placebo counterfactual that reached a ratio well above one, by using Google Mobility Data (described in Appendix (ref)), we show that mobility in residential areas increased whereas mobility in outdoor areas decreased substantially once compared to the period before the pandemic (see Figures (ref) and (ref) in Appendix (ref)). This is suggestive that South Dakota's population endogenously decided to stay more at home, and avoided environments prone to the risk of contamination. At the time, a proper lockdown policy was not necessary, and South Dakota's non-conformity to the placebo test does not seem to invalidate our approach.
In contrast, for Nebraska and California, the counterfactuals are pointing to a smaller number of cases than the actual ones, which goes against finding that lockdowns were effective to reduce cases of Covid-19. The case of California is quite emblematic, as the number of cases during the estimation window remained very small and with very low variation. However, the number of cases started to grow at a fast rate much after the cut-off date. The state of Nebraska displays a similar pattern.
To gauge the quantitative impact of lockdown policies, for each state, whether treated or control used as placebo, we compute the ratio of the counterfactual estimated cumulative cases (“without" a lockdown strategy in place) to actual ones on the 58th epidemiological day, which is the last day used to compute the counterfactual. Table (ref) reports the mean and median of the ratios across states, whereas Table (ref) in Appendix (ref) reports these ratios for each state. The first row corresponds the case in which controls are used as placebos, whereas the second considers the treated states only. As we discuss below, New York is clearly an outlier, whose ratio reached an implausible value of 16.5 as reported in Table (ref). Hence, our preferred specification is displayed in the third row which excludes New York from the pool of treated states. We also compute other two versions of these ratios using the lower bound (lb) and upper bound (up) of the 95% confidence interval in the numerator.
The ratios are clearly above one for the treated units, whether New York is excluded or not. According to our preferred specification, counterfactual estimates suggest that the number of cases would be nearly two times larger were lockdown policies absent. Again, it is reassuring that among the controls used as placebo, these average ratios remain around one.
Of course, a lockdown policy might complement or substitute other types of containment policies implemented in control or treated states (e.g., mask mandates). Hence, our estimates are arguably better interpreted as capturing the effect of a “combo" of policies that include lockdowns, relative to another “combo" of policies that do not include them. Nonetheless, in the next section, we use a simple casual model to argue that a sizable part of the counterfactual is attributable to lockdown policies.
Regarding the effects of lockdowns on cumulative deaths, we present the results for all treated states in Appendix (ref). For some states, the counterfactual cumulative deaths exhibit similar patterns to those regarding cumulative cases. But, for many other states, they are not statistically significant at least for the first days after the policy implementation. One possible explanation is that there is a delay between cases and deaths, as the latter is a consequence of the former. Hence, deaths only show up in the official statistics days after cases. Perhaps, if we could estimate counterfactuals for longer periods, the synthetic accumulated deaths would further decouple from the actual ones. In addition, since weights on the controls are estimated considering the (log) cumulative number of cases, the counterfactuals for cumulative deaths are arguably noisier.
As discussed above and presented in Table (ref) in Appendix (ref), we obtain an implausible ratio (of counterfactuals to actual cumulative cases) of 16.5 to New York. This section zooms on this state. In particular, Figure (ref) displays the estimated cumulative number of cases for New York “without" lockdown, as well as extrapolations of the cumulative number of cases based on the mean and median growth rate of the last ten days of the in-sample period.
As reported in Table (ref), among the treated states, New York was the fastest one to react to the pandemic, and established a state-wide lockdown policy only 20 days after the first case. Figure (ref) extrapolates the last in-sample observations by using both the observed mean and median growth rates for the last ten days, which yields a similar pattern to the result obtained by applying the ArCo approach. Due to the progression of the virus, particularly in New York City, the in-sample observed rates are quite high once compared to other states as illustrated in Figure (ref), which can be explained not only by the dynamics of the city but also by its high population density. Hence, New York is clearly an outlier and might not be amenable to our synthetic control approach, which justifies reporting results excluding New York.
Several factors may act as possible confounders to the estimaded effects of lockdowns.
First, individuals may react to the pandemic and change their behavior endogenously independent of the adoption of stricter lockdown rules imposed by the authorities. To the extent that this endogenous change of behavior would be similar across control and treated states, this is less of a concern. After all, we would like to report the impact of lockdown policies above and beyond individual responses to the pandemic that would occur in the absence of lockdowns.
Second, lockdown is not the only policy in the menu. Control states may not implement a lockdown, but may enact other alternative policies, such as, for example, mandatory mask-wearing or massive campaigning for people to stay at home, that may contain the pandemic evolution. This would introduce a negative bias in our estimates, suggesting even more sizable effects of lockdowns. Alternatively, lockdown policies in treated states may be designed altogether with other containment measures. In this case, our results should be better interpreted as the average effect of a “combo" of policies that include a lockdown strategy.
In order to guide the interpretation of the empirical estimates, and try to isolate the role of lockdowns, we describe a simple causal model similar to vChKpS2021. The model helps to organize ideas on how lockdown policies affect the variables of interest and how they interact with confounders. We argue below that, through the lens of this simple model and some auxiliary evidence, lockdowns (rather than other confounders or alternative policies) explain a sizable part of our estimated effects.
The model incorporates the interactions between the following variables: (i) $Y_{s,t+l}$, which are the cases of or deaths by Covid-19 in state $s$ and period $t+l$; (ii) $L_{st}$, an indicator variable that a state adopted a lockdown policy in state $s$ and period $t$; (iii) $P_{st}$, another indicator variable of alternative policies implemented; (iv) $I_{st}$, which is available information to individuals that maybe be useful to affect behavior and/or contain the pandemic;\footnote{In the vChKpS2021 model, this variable generates inter-temporal dependence between periods. In our model, this variable is modeled differently as an alternative mechanism through which policies can contain the pandemic.} (v) $B_{st}$ summarizes the relevant behavior of individuals such as adherence to social distancing, use of masks, etc; and (vi) $U_{st}$ is the set of confounders that might affect the determination of policies, individuals' behavior, and the pandemic evolution. We represent these interactions through the Direct Acyclic Graph (DAG) below.\footnote{See jP1995 and jP2009 for a discussion of DAGs and causal models.}
The sequence of events is the following. First, every period, potential confounders are determined. Second, public policies ($L_{st}$ and $P_{st}$) are set, and note that $P_{st}$ already encodes the transmission mechanisms (e.g., behavioral responses) through which policies other than lockdowns affect cases or deaths. Third, conditional on confounders and policies, individuals' information is updated. Fourth, individuals behavior are determined by confounders, policies and information. Finally, the variable of interest ($Y_{s,t+l}$) is determined.
Importantly, we assume that a lockdown policy, $L_{st}$, does not directly affect the number cases or deaths, $Y_{s,t+l}$. In fact, a lockdown policy only has effects on $Y_{s,t+l}$ to the extent that it affects some mediating variables, such as individuals' behavior (above and beyond endogenous responses to the pandemic in the absence of lockdowns) or available information.
Ideally, we would like to identify the causal effect of a lockdown policy ($L_{st}$) on the pandemic evolution ($Y_{st+l}$). By using the back-door criteria suggested by jP1993, we can envision two threats to the identification of the causal effect of interest.
First, non-observed variables ($U_{st}$) affect the probability of lockdown adoption and the evolution of the pandemic simultaneously. This is the traditional omitted variable bias in public policy evaluation. We discuss the problem of omitted variable bias below in the end of this section.
Second, as mentioned above, the potential additional problem related to the simultaneous implementation of alternative policies. Local authorities could adopt other containment policies $P_{st}$ as a substitute to the lockdown policy $L_{st}$ in control states, or as a complement in the treated ones. Hence, the adoption of alternative policies may affect the probability of implementing a lockdown and, simultaneously, affect individuals' behavior and the number of Covid-19 cases and deaths. Thus, we need to control for this possibility to identify the causal effect of interest.
The model also incorporates the possibility of endogenous responses of individuals to the evolution of the pandemic (or the accumulation of information in terms of the model). However, this is a mechanism through which our treatment acts. Therefore, according to the front-door criteria in jP1993, we should not control for these variables, $I_{st}$ and $B_{st}$ (more on that below).
In what follows we discuss how we can control for some of the threats to the identification scheme described above. We also discuss the nature of the treatment.
Consider the implementation of simultaneous policies. Beyond policies such as the closure of schools, firms, and services included in $L_{st}$, mandatory mask-wearing is arguably the most important alternative implemented policy. We focus, here, on this alternative policy.
If the state is in lockdown, the mandatory mask-wearing is less of a concern as social contact and mobility are substantially curtailed (as we document in the next subsection). Regarding control states that did not adopt a lockdown strategy, in order to deal with policy simultaneity, we leverage on the temporal mismatch between policies.
In the beginning of the pandemic, lockdown policies were widely suggested by international institutions. In contrast, mandatory mask-wearing was not encouraged by the World Health Organization (WHO), which only changed its recommendation in the beginning of July. Note, however, that we restrict the sample to the first 56 days after the first Covid-19 case in each state. Thus, the last calendar day in the sample is May 11 (Arkansas).
In Table (ref), we report the dates a mandatory mask-wearing policy was adopted in control states.\footnote{We manually collected data for these dates from state-level executive orders. For an example of these executive orders, see: https://governor.arkansas.gov/images/uploads/executiveOrders/EO_20-43.pdf.} Note that these dates are not within our sample. Indeed, the earliest adoption of such policy was in June 26 in both Illinois and Washington.
We conclude that the most important alternative policy was not implemented until the end of the period considered in this paper, and that lockdowns implemented during this period were exogenous to this policy. Hence, mandatory mask wearing among control states does not seem to attenuate the effect of interest. Below, we change the casual DAG above to accommodate this insight.
Before discussing the other threat to identification of causal effects, i.e. the omitted variable bias, we explore the nature of the treatment we are considering. What a lockdown policy does? Why would it affect the variables of interest?
The causal diagram in Figure (ref) suggests two mechanisms. First, a lockdown affects individuals' behavior by reducing mobility. Second, the policy might affect the available information. For instance, its implementation can increase the awareness of individuals about the pandemic and provide incentives to further changes in behavior.
Despite considering the theoretical possibility of an informational transmission channel, some auxiliary empirical evidence suggests that its relevance is quite limited. Figure (ref) plots the number of (changes in) pandemic-related Google searches in the days immediately before and after the policy is implemented, obtained through an event-study design for the treated states.\footnote{We collected data on total Google searches for terms related to the pandemic for all states in the sample. The search terms include “pandemic" and “Covid-19". Google analyses a sample of total searches and makes available the relative amount of searches at each point in time. The results are normalized to a fraction of the highest number of searches in time for each state. In order to make the data comparable across states, we use the SD2014 methodology.} We also plot the 95% confidence interval for each estimate. Note there is a small increase in searches a few days after the lockdown announcement, but its magnitude is very limited and it vanishes almost immediately.
We also examine if lockdown measures impact Google searchers related to traditional and alternative methods to fight the pandemic. The search terms for traditional methods include “social distancing", “mask use" and “washing hands". The search terms for alternative treatment include “zinc", “hydroxychloroquine", and “Covid alternative treatments".\footnote{Again, we use SD2014 methodology to standardize the data.} Results are reported in Figure (ref). The top (bottom) panel plots searchers for traditional (alternative) methods in both treated and control sates.
We find little evidence that the search patterns are systematically different between treatment and control groups, or that the lockdown implementation affected these search patterns. Hence, information seem to evolve in an aggregate way and not to be affected by policies. We further rewrite the causal diagram to incorporate this insight.
By eliminating the casual link between $L_{st}$ and $I_t$, this simplification allows a straightforward interpretation of the treatment. The lockdown policy affects pandemic evolution mainly through its effects on behavior. To confirm this, we evaluate the impact of lockdown policies in mobility, captured by Google Mobility Data (described in Appendix (ref)), above and beyond the effects that would have happened endogenously.
Similar to the way we compute the counterfactual to cumulative deaths, we compute the counterfactual to mobility in residential and outdoor areas. Figure (ref) plots the average measures of mobility that in fact realized in treated states (blue lines), and the average counterfactual measures were lockdowns not implemented there (red lines). We consider the same control states above and we use the same weights as in the main estimates. The top panel considers outside activities, whereas the bottom panel considers residential activities.
As the figure makes it clear, lockdown policies affected the pandemic evolution mainly through its effects on behavior, captured by this substantial decrease (increase) in outside (residential) mobility in treated states relative to the counterfactuals (that is, above and beyond endogenous behavioral responses in the absence of lockdowns).
In this subsection we discuss the problem of omitted variable bias. To do so, we write a linear model based on the DAG in Figure (ref). The variable of interest depends on available information ($I_{st}$), individuals' behavior ($B_{st}$) and confounders ($U_{st}$). That is, \[ Y_{s,t+l}=\pi B_{st}+\mu I_{st}+\delta U_{st}+\epsilon^Y_{st}, \] where $\pi$, $\mu$, and $\delta$ are unknown parameters and $\epsilon^Y_{st}$ is a zero-mean random term.
The behavior of individuals also depends on the lockdown policies implemented ($L_{st}$) and can be expressed as \[ B_{st} = \alpha L_{st}+\eta I_{st}+\epsilon^B_{st}, \] where $\alpha$ and $\eta$ are unknown parameters and $\epsilon^B_{st}$ is a zero-mean random term.
As a consequence, we can write the reduced-form of the model above as \[ Y_{s,t+l}=\beta L_{st}+\gamma I_{st}+\delta U_{st}+\epsilon_{st}, \] where $\beta=\alpha\pi$, $\gamma=\pi\eta+\mu$, and $\epsilon_{st}=\pi\epsilon^B_{st} + \epsilon^Y_{st}$.
Since we do not observe $U_{st}$ and these confounders are correlated with $L_{st}$, a simple ordinary least-squares (OLS) estimation of the reduced-form equation does not identify $\beta$. The ArCo methodology used in this paper helps to mitigate the effects of potential confounders.
Indeed, for each state, we find a (non-necessarily convex) combination $\boldsymbol{w}$ of states in the control pool that most closely matches the number of cases in the treated state. Thus, we propose the following estimator, \[ \widehat{\boldsymbol{\beta}}_{s,t+l}=Y_{s,{t+l}}-\omega_0-\boldsymbol{\omega}'\boldsymbol{Y}^C_{t+l}, \] where $\boldsymbol{Y}^C_{t}$ is the vector of cases for the control states. Also, note that \[ \widehat{\beta}_{s,t+l}=\beta+\delta(U_{st}-\omega_0-\boldsymbol{\omega}'\boldsymbol{U}^C_{t}), \] where $\boldsymbol{U}^C_t$ is the vector of non-observed confounders for the control group. It is clear that if \[ U_{st}=\omega_0+\boldsymbol{\omega}'\boldsymbol{U}^C_{t}, \] then the proposed estimator recovers the effect of the lockdown policy on the number of registered cases. That is, the additional identification hypothesis is that the combination of estimated weights does not only reproduce the trends in the in-sample period but also the non-observed relevant variables. Finally, note that $\beta$ recovers precisely the effect of lockdowns through the behavioral channel.
In this paper, as opposed to most of the early and incipient literature on the lockdown effects during the Covid-19 crisis, we consider a purely data-driven approach to assess the impact of lockdowns on the short-run evolution of the number of cases and deaths in some US states. Also, as opposed to some recent papers that use a difference-in-difference approach, we adopt a variant of the synthetic control approach. On average, according to the synthetic controls, the counterfactual accumulated number of cases would be two times larger were lockdown policies not implemented in treated states.