EconBase
← Back to paper

Structural Econometric Estimation of the Basic Reproduction Number for Covid-19 Across U.S. States and Selected Countries

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

43,245 characters · 9 sections · 45 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Structural Econometric Estimation of the Basic Reproduction Number for Covid-19 Across U.S. States and Selected Countries

abstractThis paper proposes a structural econometric approach to estimating the basic reproduction number ($\mathcal{R}_{0}$) of Covid-19. This approach identifies $\mathcal{R}_{0}$ in a panel regression model by filtering out the effects of mitigating factors on disease diffusion and is easy to implement. We apply the method to data from 48 contiguous U.S. states and a diverse set of countries. Our results reveal a notable concentration of $\mathcal{R}_{0}$ estimates with an average value of 4.5. Through a counterfactual analysis, we highlight a significant underestimation of the $\mathcal{R}_{0}$ when mitigating factors are not appropriately accounted for. Keywords: basic reproduction number, Covid-19, panel threshold regression model JEL Classifications: C13, C33, I12, I18, J18

\pagenumbering{gobble}

\pagenumbering{arabic} \setcounter{page}{1}

Introduction

Since the Covid-19 outbreak, a rapidly growing body of literature has been devoted to estimating its reproduction numbers, which are standard epidemiological metrics used to quantify the rate at which the disease spreads at the initial and later stages of the epidemic. Obtaining accurate estimates of the reproduction numbers is critical in understanding the transmissibility of infectious agents and formulating public health responses. Different formal definitions of reproduction numbers have been proposed in the literature. One of the most widely adopted metrics is the basic reproduction number, denoted by $\mathcal{R}_{0}$, which is defined as "the average number of secondary cases produced by one infected individual during the infected individual's entire infectious period assuming a fully susceptible population" DelValle2013. An essential assumption of this definition is that the population must be entirely susceptible without any immunity or interventions. In other words, the value of $\mathcal{R}_{0}$ represents the maximum rate at which an epidemic can spread in the absence of any mitigating actions, whether voluntary or mandatory, such as improved personal hygiene, mask-wearing, social distancing, isolation, or vaccination. In practice, the day-to-day evolution of the epidemic is governed by public health directives and individual behavior, as well as the internal dynamics of the epidemic, which on its own eventually slows down its rate of spread as the number of susceptible individuals declines (due to immunity, death, or mitigation practices). A measure that quantifies the time-varying transmissibility during the life of the epidemic is the effective reproduction number, $\mathcal{R}_{e,t}$.\footnote{The effective reproduction number is the expected number of secondary cases produced by one infected individual in a population that includes both susceptible and non-susceptible individuals at time $t$ (assuming that the conditions remained the same after time $t$). The effective reproduction number differs from another time-varying metric, the case reproduction number Fraser2007. The latter is defined as the average number of secondary cases that a primary case actually infects at time $t$, and thus its estimation can only be undertaken retrospectively.} The focus of the present study is on the structural estimation of $\mathcal{R}_{0}$ for Covid-19 across U.S. states and countries.

Obtaining a reliable estimate of $\mathcal{R}_{0}$, although seemingly straightforward, can be quite challenging in practice. This is largely due to the difficulty of accurately counting early infections during an epidemic, especially for newly emerging pathogens like SARS-CoV-2. Even if public health officials may strive to establish surveillance systems promptly, the data quality during the initial stages of the outbreak is inevitably poor. Additionally, many existing time-series estimation methods encounter difficulties in identifying an appropriate sample period to estimate $\mathcal{R}_{0}$. Unlike $\mathcal{R}_{e,t}$, which varies over time and is affected by mitigation policies, to estimate $\mathcal{R}_{0}$ it is important to focus on the initial stages of the epidemic when the infections are spreading exponentially, without impediments. If the selected sample is too long, it will likely include periods of reduced social interactions due to mandatory and/or voluntary behavioral changes, which slow down the spread of the infection and result in a downward bias in the $ \mathcal{R}_{0}$ estimates. On the other hand, if the sample selected at the initial stages of the epidemic is too short, the under-recording of infected cases can be particularly serious, resulting in an underestimation of the disease's transmission rate and again leading to the underestimation of $ \mathcal{R}_{0}$.

The availability of Covid-19 data from diverse populations and locations presents an unprecedented opportunity for researchers to compare the $ \mathcal{R}_{0}$ estimates using various models and estimation methods. Researchers around the world have reported a wide range of $\mathcal{R}_{0}$ values of Covid-19, as summarized in Table (ref). Some studies have highlighted the heterogeneity of $\mathcal{R}_{0}$ across different countries and regions Korolev2021, FernandezVillaverde2022, while others find evidence of similarity Katul2020, Chudik2022. Many of the existing estimates were obtained based on complex models that require large amounts of data and involve numerous assumptions made by the modeler. As it is rare to have high-quality data for all components of the model, researchers often have to calibrate multiple model parameters Delamater2019. As our knowledge of the disease and the quality of data have improved over time, an increasing number of researchers have argued that earlier estimates of $\mathcal{R}_{0}$ were too conservative Katul2020, Ke2021.

In the current paper, we estimate $\mathcal{R}_{0}$ of Covid-19 using a structural econometric approach, which allows us to identify $\mathcal{R}_{0} $ as the intercept in panel regressions of outcomes on a number of key factors across the contiguous states in the U.S. and a diverse range of countries. This structural method is grounded in the individual-based stochastic network model for epidemic diffusion recently developed by PY2022, who established a moment condition for the evolution of the infections. Chudik2022 use the same moment condition but allow for time variations in the transmission rate, modeled as a function of various factors, including government responses (mandatory mitigation policies and economic support), precautionary behavioral changes such as voluntary social distancing, vaccinations, and virus mutations. These authors focus on nine European countries, which had similar starts at the outset of the epidemic in March 2020 but ended up with differing outcomes. They find robust statistically significant effects of reduced mobility, increased government economic support to comply with containment policies, vaccination, and virus mutations on the transmission rate of Covid-19.

We employ the structural approach proposed by Chudik2022, CPR hereafter, and estimate $\mathcal{R}_{0}$ as the intercept in panel data regressions where suitably transformed case numbers are explained in terms of the mitigating factors identified in the CPR study. Such an estimate can be viewed as a counterfactual outcome that could have resulted if none of the public health policies had been enacted. This method has several advantages. First, by controlling for a number of mitigation measures, we are able to estimate $\mathcal{R}_{0}$ using relatively long samples, thus avoiding the difficulties of selecting an early sample. Second, the model easily accommodates the differences in the initial start dates of the Covid-19 outbreaks across countries (or regions) by using an unbalanced panel data model. Third, country-specific $\mathcal{R}_{0}$ can be estimated by least squares under very weak assumptions and only requires the regressors in the panel data model to be weakly exogenous. We also use lagged values of the regressors to avoid simultaneity bias, and we take account of error serial correlation when computing the standard errors to deal with the implications of using overlapping seven-day moving averages, a practice routinely employed in the literature to deal with uneven recording of infected cases over the different days of the week. Fourth, our estimation of $\mathcal{R}_{0}$ requires only Covid-19 case data and an assumed value for the recovery rate, $\gamma $. This approach serves as a useful complement to other methods that rely on mortality data. Instead of estimating the recovery rate, we set it a priori using clinical evidence with a high degree of reliability. Fifth, we allow for the under-reporting of infected cases, which is known to vary across different phases of the epidemic.

We find remarkably similar estimates of $\mathcal{R}_{0}$ across 48 contiguous U.S. states, averaging at 4.7 with a range of 4.2 to 4.9 (4.5 to 5.3) under the assumption of a lower (higher) degree of under-reporting. In addition, for a selection of diverse nations, $\mathcal{R}_{0}$ remains within a narrow range of 3.4 to 5 across all scenarios considered, with an average value of 4.3. Overall, these estimates align well with the recent findings in the literature. Furthermore, our counterfactual analysis underscores a substantial underestimation of $\mathcal{R}_{0}$ that results from neglecting the mitigating factors governing the time profile of the Covid-19 transmission rate.

The rest of the paper is organized as follows. Section (ref) reviews the related literature. Section (ref) outlines the empirical methodology. Section (ref) presents the main findings, and Section (ref) concludes. Additional tables and figures are provided in an online supplement.

Related Literature

The estimation of the reproduction number for infectious diseases, including Covid-19, has been a topic of extensive research in the literature. In the interest of space, we will highlight a selection of recent contributions that are closely related to our study.

There are many different approaches to the estimation of $\mathcal{R}_{0}$ that can be broadly categorized into mathematical and statistical approaches White2020. The mathematical approaches rely on constructing a theoretical model, which makes explicit hypotheses about the biological mechanisms driving the transmission rate and its dynamics. Such hypotheses range from simple representations of the time needed to complete some part of the disease process (such as the incubation period) to complex agent-based models with explicit modeling of social interactions Lessler2016. In contrast, statistical methods derive estimators of $ \mathcal{R}_{0}$ directly from data using probabilistic models. They do not require specifying a disease-approximating mathematical model and can be implemented using standard statistical software.

The most widely used mathematical models for disease transmission are the compartmental models, most notably the classical SIR model that divides a population into three compartments (susceptible, infected, and recovered) and its variants, such as the SEIR and SIRD models (with an additional exposed or deceased compartment, respectively). Tang2020a provide a recent review of the compartmental infectious disease models. Many researchers have made efforts to infer $\mathcal{R}_{0}$ of Covid-19 from different mathematical models and adopted various frequentist methods to estimate or calibrate the model parameters Buckman2020, Katul2020, Ke2021, Korolev2021, FernandezVillaverde2022. Another strand of literature estimates similar compartmental models using Bayesian techniques Wu2020, Arias2023, Atkeson2023.\footnote{Some authors categorize this strand of methods separately as stochastic approaches or stochastic Markov Chain Monte Carlo (MCMC) methods.}

The statistical methods typically require an estimate of the generation time, which is the time lag between infections in primary and secondary cases Obadia2012. A frequently used estimator of $\mathcal{R}_{0}$ was proposed by Wallinga2007, who show how to estimate $\mathcal{R} _{0}$ from the estimated exponential growth rate using data on infected cases during the early phase of an outbreak and the moment generating function of the generation time distribution. Since the time lag between all infectee/infector pairs is not directly observable, the generation time distribution in practice is often substituted with the serial interval distribution that measures the time between symptoms onset. Besides the method of statistical exponential growth, other statistical approaches to estimating $\mathcal{R}_{0}$ include sequential Bayesian Bettencourt2008 and maximum likelihood estimations White2008. See White2020 for a recent review of the statistical estimation of the reproduction number. Many researchers have adopted the statistical exponential growth model to estimate the $\mathcal{R}_{0}$ of Covid-19 using early case data for China Li2020, Liu2020a, Sanche2020, Zhao2020.

table[table omitted — 1,829 chars of source]

To facilitate comparisons, Table (ref) provides a summary of reported $\mathcal{R}_{0}$ estimates and respective methods from selected studies. This table complements earlier summaries that focus on the $ \mathcal{R}_{0}$ estimates for China Alimohamadi2020, Liu2020 by showcasing the more recent estimates for various locations, with special emphasis on the U.S. estimates.\footnote{Alimohamadi2020 found that the estimates of $\mathcal{R}_{0}$ based on 23 studies using Chinese data ranged from 1.90 to 6.49, with a mean value of 3.38. Liu2020 surveyed 12 studies reporting a mean $\mathcal{R}_{0}$ estimate for China of 3.28, with a range of 1.4 to 6.49.}

Empirical Methodology

Our econometric approach follows from a simplified moment condition of an individual-based stochastic network SIR model developed by PY2022. Assuming a homogeneous population,\footnote{PY2022 also considered a heterogeneous population segmented into multiple groups that can be based on demographic and socioeconomic factors as well as contact locations/schedules. Using simulations, they find that when $n$ is sufficiently large and the rate of infection is reasonably rapid, there are little differences in the time profiles of infected cases at the aggregate level whether one uses a single-group or multi-group network model, which partly justify our use of single-group moment conditions for estimation of $\mathcal{R}_{0}$.} they derive the following moment condition in $c_{t}$ (per capita number of cumulative cases on day $t$) and $i_{t}$ (per capita number of active infected cases on day $t$):

equation[equation omitted — 127 chars of source]

where $\gamma $ is the recovery rate (per day) and is assumed to be time-invariant; $\beta _{t}$ denotes the transmission rate, which can vary over time; and $n$ stands for the population size. Under this setup, $i_{t}$ can be computed from current and past values of $c_{t}$ for a given choice of $\gamma$:

equation*[equation* omitted — 99 chars of source]

or recursively, using $i_{t}=(1-\gamma )i_{t-1}+\Delta c_{t},$ where $\Delta c_{t}=c_{t}-c_{t-1}$ denotes per-capita daily new cases, with $c_{t}=i_{t}=0$ , for $t\leq 0$ (before the start of the epidemic).

Since $n$ is large and $\beta _{t}i_{t}$ is small, Eq. ((ref)) implies that $\beta _{t}$ can be approximated by $-i_{t}^{-1}\ln \left( \frac{1-c_{j,t+1}}{1-c_{jt}} \right)$. To account for additional factors that may influence disease transmission, CPR model the time variations of $\beta _{t}$ in terms of a set of variables measuring mandatory or voluntary social distancing, economic support, vaccine uptake, and virus mutations. We adopt CPR's strategy and estimate the following threshold panel data model:

equation[equation omitted — 238 chars of source]

in $n$ states or countries indexed by $j=1,2,\ldots ,n,$ over the days $ t=1,2,\ldots ,T$. The dependent variable, $\beta _{jt}/\gamma $, measures the transmission rate in country $j$ scaled by the recovery rate, $\gamma $. One can also interpret $1/{\gamma }$ as the mean infectious period. Based on the clinical evidence for Covid-19 and in line with quarantine policies, we set $\gamma =1/14$, which is also a commonly adopted value in previous studies.\footnote{If data on the recovered cases are available, one can also estimate the recovery rate using, for example, the recovery moment condition outlined in PY2022. However, we have opted to calibrate the recovery rate value based on clinical information due to its greater reliability, and this is the only parameter we need to calibrate.}

Turning to the right-hand side variables of Eq. ((ref)), $ \boldsymbol{x}_{j,t-p}$ is a vector of explanatory variables lagged $p$ days, and $u_{j,t+1}$ are the idiosyncratic errors. The precautionary behavior is captured by the term $I(\Delta c_{j,t-p}>\tau )$, where $I(\cdot )$ is the indicator function equal to one if $\Delta c_{j,t-p}>\tau $ and zero otherwise. $\Delta c_{j,t-p}$ is the (seven-day moving average of) reported daily new cases per 100,000 people also lagged $p$ periods. We chose this as the threshold variable since it has been the most closely monitored metric on the spread of Covid-19 in media reports worldwide. $\tau $ is a threshold parameter and, for simplicity, is assumed to be the same across countries. Intuitively, higher values of $\Delta c_{j,t-p}$ imply greater perceived risk; individuals are likely to practice voluntary social distancing more vigorously if $\Delta c_{j,t-p}$ exceeds the threshold value $\tau $.

The parameters to be estimated include $\alpha _{j}$, $\boldsymbol{\psi }$, $ \kappa $, and $\tau $. In the present study, the key structural parameter of interest is the state- or country-specific intercept, $\alpha _{j}$, which we use as the estimator of the mean basic production number, $\mathcal{R}_{0} $, for state (or country) $j$.\footnote{In contrast to CPR who also considered a common intercept term for the nine European countries in their study, we focus on the fixed-effects (FE) estimation because our study covers a larger set of heterogeneous countries and states in the U.S.} This is valid because in the absence of any mitigating interventions, whether mandatory and/or voluntary, both $\mathbf{x }_{j,t-p}$ and $I(\Delta c_{j,t-p}>\tau )$ would equal to zero, and then $ \beta _{jt}/\gamma =\beta _{j0}/\gamma $ for all $t$. Hence, the intercept serves as an estimate of $\beta _{j0}/\gamma $, which is indeed the same as $ \mathcal{R}_{0}$ in the standard SIR models. To identify $\alpha _{j}$, however, it is important that we take into account the mitigating effects of behavioral changes in response to the spread of the epidemic, uptake of vaccination, and possible mutations of the virus. It is also worth noting that $c_{jt}$ is very small at the onset of the outbreak, and hence $-\ln \left( \frac{1-c_{j,t+1}}{1-c_{jt}}\right) \approx \Delta c_{j,t+1}$. This approximation, coupled with the fact that $\mathbf{x}_{j,t-p}$ and $I(\Delta c_{j,t-p}>\tau )$ are both equal to zero at the initial stages of the epidemic, makes Eq. ((ref)) consistent with the basic idea of estimating $\mathcal{R}_{0}$ by the average growth of infections per active cases over its infectious period.

We estimate model ((ref)) over two sample periods. The first sample covers the period from March 6, 2020, to January 31, 2021, prior to the roll-out of Covid-19 vaccinations. The second sample extends the first one to November 30, 2021. We will refer to these two samples as the pre-vaccination sample and the full sample, respectively. Both samples are unbalanced since the dates of the outbreaks of Covid-19 and the available case data differ across U.S. states and countries.

Apart from the threshold indicator variable, we also include policy stringency and economic support indices as mitigating factors for both sample periods. For the full sample, we include two additional factors, namely the share of the population that is fully vaccinated and the share of the Delta variant of Covid-19. The rationale behind including the latter factor stems from its greater transmissibility as compared to the earlier variants.

For the U.S. state-level regressions, we also include the interaction between the economic support index and a dummy variable for a Republican governor. This allows us to account for the varied effects of economic policies on the spread of Covid-19 across states with different political leanings, since the economic support index does not include support to firms or take into account the total fiscal value of economic support provided Hale2020.

We utilized data from multiple sources to conduct our analysis. The policy stringency and economic support indices were retrieved from the Oxford Covid-19 Government Response Tracker (OxCGRT) project.\footnote{Available at \url{https://github.com/OxCGRT/covid-policy-dataset}.} The data on Delta variants were obtained from CoVariants.org Hodcroft2021. Our World in Data was the source of the data on Covid-19 vaccinations as well as the country-level Covid case numbers.\footnote{Available at \url{https://github.com/owid/covid-19-data/tree/master/public/data}.} The Centers for Disease Control and Prevention's Covid data tracker was used to obtain state-by-state Covid cases in the U.S.\footnote{Downloaded from \url{https://data.cdc.gov/api/views/9mfq-cb36/rows.csv?accessType=DOWNLOAD} (last accessed: May 2023).}

To reduce the impact of within-week seasonality in the reported daily cases due to delayed reporting or reduced testing on weekends, we follow the standard practice and take seven-day moving averages of the reported data before estimation. Another important data issue that needs to be addressed is the under-reporting of confirmed cases. This occurs mainly due to asymptomatic infections and lack of testing, especially during the early stages of an outbreak. In the epidemiological literature, the magnitude of under-reporting is often measured by the multiplication factor (MF), which is defined as the ratio of true to reported cases Gibbons2014. For Covid-19, with increased testing we would expect MF to fall over time, and simulations carried out by PY2022 support this. To account for this declining under-reporting, we consider two different specifications of the MF values: one where the MF linearly declines from 5 to 2, and another where it declines from 8 to 2.5.

We compute three types of standard errors in all estimations. The first type is the “usual" standard error assuming no cross-sectional or serial correlation in the errors. The second type, labeled as “robust1" in the tables, is the Newey-West type heteroskedasticity and autocorrelation consistent standard error NeweyWest1987. The third type, labeled “robust2", is the DriscollKraay1998 standard error, which accounts for both cross-sectional and serial correlation. For the latter two robust standard errors, we choose the truncation lag as the integer part of $T^{1/3}$, where $T$ denotes the longest time span of the unbalanced panel.

The Main Findings

This section presents the estimates of $\mathcal{R}_{0}$ for the U.S. states and 19 countries. In assessing the estimates' statistical significance, we will focus on the most conservative standard errors that are robust to both error cross-sectional and serial correlation. Estimates for the mitigating covariates and their standard errors are provided in the online supplement.

Estimates of $\mathcal{R}_{0}$ for U.S. states

table[table omitted — 4,123 chars of source]

We use the estimates of $\alpha _{j}$ from the panel threshold regressions, ((ref)), after filtering out the effects of the mitigating factors, to obtain an estimate of $\mathcal{R}_{0}$ for the $j^{th}$ state. These estimates are summarized in Table (ref) for all 48 contiguous states in the U.S. We report two sets of estimates for each of the two sample periods that end on January 31, 2021, and November 30, 2021, respectively. We scale up all reported cases by $MF_{t}$ that declines linearly from 5 to 2 in one scenario and from 8 to 2.5 in another scenario.\footnote{The parameter estimates for the mitigating factors are provided in Table (ref) of the online supplement.}

Overall, the results exhibit remarkable similarities across all states and sample periods, with all estimates being statistically highly significant, even when the most robust standard errors are used. The results are also reasonably robust to the choice of MF and the sample period. When using lower MF values, the estimates of $\mathcal{R}_{0}$ vary between 4.21 (Oklahoma) to 4.88 (Rhode Island) in the pre-vaccination sample, and between 4.44 (Mississippi) and 4.95 (Rhode Island) in the full sample. With higher MF values, the estimates rise slightly and vary between 4.47 (New Hampshire) to 5.29 (Rhode Island) in the case of the pre-vaccination sample, and 4.47 (Vermont) to 5.26 (Rhode Island) when we use the full sample.

A closer inspection of results in Table (ref) reveals that the estimates of $\mathcal{R}_{0}$ tend to be slightly higher in states such as Rhode Island, California, Connecticut, and New Jersey, and lower in states such as Missouri, Oklahoma, and Alabama.\footnote{For a visual representation of the spatial distribution of the $\mathcal{R}_{0}$ estimates, refer to the hexagon maps presented in Figures (ref) and (ref) of the online supplement.} Despite these differences, the estimates are very tightly clustered. This can be observed from the histograms in Figure (ref), which give the distribution of the estimates for the two sample periods and MF specifications. With lower MF values, the estimates are primarily clustered around 4.2 to 4.8. When the MF is higher, about 40 out of the 48 states have estimates concentrated between 4.6 and 5.1. To summarize, an average estimate (across states, periods, and MF values) of 4.7 seems to provide a reasonable summary number for the U.S.

figure[figure omitted — 986 chars of source]

Country-specific Estimates

table[table omitted — 2,312 chars of source]

Table (ref) reports the estimates of $\mathcal{R}_{0}$ for 19 countries.\footnote{The associated estimation results for the mitigating factors of the country regressions are provided in Table (ref) in the online supplement, where we also present the $\mathcal{R}_{0}$ estimates in descending bar charts in Figures (ref) and (ref).} Similar to our findings for the U.S. states, the estimates fall within a narrow range. Moreover, the results are fairly robust to the choice of MF, with higher MF values only marginally increasing the $\mathcal{R}_{0}$ estimates. Argentina, Chile, Peru, and Spain are found to have the highest average $\mathcal{R}_{0}$ values, while Nigeria, South Korea, and Indonesia have the lowest average $\mathcal{R}_{0}$ values. Specifically, for the pre-vaccination period, the estimates range from 3.85 (South Korea) to 5.01 (Argentina) under the lower MF values, and 3.84 (South Korea) to 5.04 (Argentina) under the higher MF values. The full-sample estimates lie between 3.42 for Nigeria and 4.46 for Chile (3.44 for Nigeria and 4.56 for Spain) in the low (high) MF scenarios. The slight differences in estimated $\mathcal{R}_{0}$ values across countries might be associated with varying degrees of under-reporting. To further examine the distribution of the estimates, Figure (ref) displays the histograms for each sample period and MF specification. We see that the estimates from the pre-vaccination sample are relatively evenly distributed within the range, whereas about 15 out of the 19 countries have full sample estimates falling in the range of 3.8 to 4.5. In sum, the average estimate of $ \mathcal{R}_{0}$ for 19 countries across both sample periods and MF values is around 4.3, which is quite close to the average U.S. estimate of 4.7. Overall, both our U.S. states and international estimates align with the recent literature and suggest that earlier studies have underestimated $ \mathcal{R}_{0}$, as documented in Table (ref).

figure[figure omitted — 1,003 chars of source]

Estimates of $\mathcal{R}_{0}$ without Mitigating Factors

To demonstrate the importance of filtering out the effects of the mitigating factors in the estimation of $\mathcal{R}_{0}$, we estimated Eq. ((ref)) without any of the mitigating factors (or by setting both $ \mathbf{x}_{j,t-p}$ and $I(\Delta c_{j,t-p}>\tau )$ to zero). The state- and country-specific estimates are presented in Tables (ref) and (ref) in the online supplement. Evidently, the estimated $\mathcal{R}_{0}$'s are significantly biased downward when the mitigating factors are not accounted for. The average estimate across sample periods and MF specifications is only 1.5 (1.3) for U.S. states (19 countries). These results underscore the problem of omitted variable bias resulting from neglecting mitigating factors.

Conclusions

In this paper, we propose a novel approach to the estimation of the basic reproduction number, $\mathcal{R}_{0}$, of Covid-19, and provide estimates for U.S. states and a selected number of countries. Our approach falls under the category of counterfactual causal analysis, where the focus is to filter out the effects of mitigating factors on the diffusion of the virus, and thus identify $\mathcal{R}_{0}$ when the values of the mitigating factors are set to zero in the counterfactual exercise. Our estimates of $\mathcal{R} _{0}$ turn out to be centered around $4.5$, clustered closely across U.S. states and a selected number of countries. Not allowing for the mitigating factors results in estimates of $\mathcal{R}_{0}$ that suffer from substantial downward bias.

While our estimation approach is relatively simple to implement and yields satisfactory estimates, it is subject to an important limitation. In order to identify and provide accurate estimates of $\mathcal{R}_{0}$, it requires the availability of reliable data on mitigating factors, ideally taking into account all such factors in the analysis. In practice, this means that the method might not be applicable at the very early stages of an epidemic. Nevertheless, we believe our approach offers a useful alternative to the existing methods for $\mathcal{R}_{{0}}$ estimation, which also face the challenge of obtaining reliable early samples that are not subject to any mitigating interventions.

\onehalfspacing

Statements and Declarations

{ The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. The authors have no relevant financial or non-financial interests to disclose.}