EconBase
← Back to paper

What can we learn about SARS-CoV-2 prevalence from testing and hospital data?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,261 characters · 21 sections · 53 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

What can we learn about SARS-CoV-2 prevalence from testing and hospital data?

spacing{0.9} \begin{titlepage} \thispagestyle{empty} \begin{abstract} Measuring the prevalence of active SARS-CoV-2 infections in the general population is difficult because tests are conducted on a small and non-random segment of the population. However, people admitted to the hospital for non-COVID reasons are tested at very high rates, even though they do not appear to be at elevated risk of infection. This sub-population may provide valuable evidence on prevalence in the general population. We estimate upper and lower bounds on the prevalence of the virus in the general population and the population of non-COVID hospital patients under weak assumptions on who gets tested, using Indiana data on hospital inpatient records linked to SARS-CoV-2 virological tests. The non-COVID hospital population is tested fifty times as often as the general population, yielding much tighter bounds on prevalence. We provide and test conditions under which this non-COVID hospitalization bound is valid for the general population. The combination of clinical testing data and hospital records may contain much more information about the state of the epidemic than has been previously appreciated. The bounds we calculate for Indiana could be constructed at relatively low cost in many other states. \end{abstract} {{ Key words: COVID-19, SARS-CoV-2, prevalence, partial identification }} \end{titlepage}

Introduction

Constructing credible estimates of the current prevalence of SARS-CoV-2 in the United States is challenging. Despite growing since the start of the epidemic, testing rates remain low in most of the country. Moreover, tests are often allocated to people exhibiting COVID-19 symptoms or who are thought to have come into contact with the virus wsjTest. For example, New York and Texas both use a self-diagnostic tool to screen people for testing. \footnote{See \url{https://covid19screening.health.ny.gov/covid-19-screening/} and \url{https://txctt.force.com/ct/s/assessment?language=en_US}.} The low rate of testing in the general population means that the number of confirmed SARS-CoV-2 cases almost certainly understates the true number of infections in the population. At the same time, statistics like the fraction of tests that are positive likely overstate population prevalence because the tested population is more likely to be infected than the population as a whole.

In this paper, we propose a new approach to measuring the point-in-time prevalence of active SARS-CoV-2 infections in the overall population using data on patients who are hospitalized for non-COVID reasons. There is less uncertainty about SARS-CoV-2 prevalence among non-COVID hospital patients because people in the hospital are tested at much higher rates than the general population, even if they are hospitalized for reasons unrelated to COVID-19 SuttonEtAl2020. We show how a family of weak instrumental variable assumptions can be used to reduce the inferential uncertainty about the prevalence of SARS-CoV-2. We combine these assumptions with linked testing-hospital data from Indiana to estimate relatively tight upper and lower bounds on the prevalence of active SARS-CoV-2 infections in the overall population in Indiana in each week from mid-March to mid-December. The detailed Indiana data allows us to conduct robustness checks that partially validate some of our assumptions. Our basic method (without the validation checks) could be implemented using data that many states are already collecting and partially reporting. Thus, our approach could help states extract timely prevalence information using existing surveillance data. Importantly, our paper is focused on estimates of the fraction of the population that would test positive in each week. These estimates are distinct from recent efforts to estimate the share of the population ever infected with SARS-CoV-2 ManskiMolinari2020. The distinction between active prevalence and cumulative prevalence is important because the prevalence of active infections is a key determinant of the spread of the epidemic, given that the level of immunity in the population is thought to be quite low.

Estimates and forecasts of the prevalence of active SARS-CoV-2 infections are crucial for public and private responses to the disease. They have shaped decisions about disruptive non-pharmaceutical interventions such as school closures, non-essential business closures, gathering restrictions, stay-at-home mandates, and temporary increases in the generosity of the unemployment insurance system gupta2020. Reported estimates of prevalence likely also motivate individual precautionary behaviors ranging from wearing a mask to reducing demand for goods and services that require physical interaction gupta_brookings_2020,AllcottEtAl2020, philipson2000economic, philipson1996private, kremer1996integrating. Finally, estimates of prevalence are a necessary input into efforts to measure other quantities of interest, like the infection fatality rate and the infection hospitalization rate.

Given its importance, researchers have developed several approaches to measuring SARS-CoV-2 prevalence given non-representative testing. One of the most credible is to conduct a biometric survey in which tests are offered to a representative sample of the population NirWave1, NirWave2,GudbjartssonEtAl2020. Other studies have tested a census of smaller populations such as cruise ships DiamondPrincessFeb2020. However, it is difficult and costly to regularly implement a survey with accurate coverage and high response rates, especially a new survey that has not been in the field for long. A second approach involves backcalculation methods, which use data on observed hospitalizations or deaths to infer disease prevalence at earlier dates using assumptions about the unobserved parameters that determine the progression of the disease, hospitalization rates, and case fatality rates brookmeyer1988method, egan2015review, FlaxmanEtAl2020, SaljeEtAl2020. Backcalculation may work well if hospitalizations or deaths are well measured, and if previous research has already reached consensus about key parameters related to the disease. However, backcalculation may be less credible for a novel virus because the scientific knowledge base is smaller. And even when backcalculation is based on credible assumptions, it may be of limited value for public health decision making because both hospitalizations and deaths lag current infections VerityEtAl2020.

An alternative to biometric surveys and backcalculation is to combine non-random clinical testing data with weak distributional assumptions to construct bounds on population prevalence (e.g. manski1999identification, wing2010three). In the COVID-19 epidemic, stock2020identification provide bounds on population prevalence using a testing encouragement design, which is not always available. ManskiMolinari2020 bound the share of the population that has ever been infected under a “test monotonicity assumption" that the infection rate is weakly higher among the tested than among the untested. This assumption is appealingly credible, and the bounds can be calculated from widely available test data. However, even when the focus is cumulative prevalence under test monotonicity, the prevalence bounds are often wide because testing is so rare.

The strategy we pursue in this paper is based on the insight that test rates are especially high among hospitalized patients, even patients hospitalized for reasons that are apparently unrelated to COVID-19, such as labor and delivery or vehicle accidents. The upper and lower bounds on prevalence in these populations will tend to be much tighter than similar bounds in the general population. In addition, it is plausible that some types of hospitalizations occur for reasons that are independent of infection risk. For example, people who are hospitalized for injuries sustained in a traffic accident might be expected to have about the same risk of SARS-CoV-2 infection as the general population. We build on this insight to estimate more informative upper and lower bounds on weekly SARS-CoV-2 prevalence in the population.

Our paper makes methodological contributions that may be relevant to the development of more informative public health surveillance systems. We describe the conditions under which upper and lower bounds on active prevalence among non-COVID hospitalizations are valid estimates of the upper and lower bounds on prevalence in the general population. We maintain the test monotonicity assumption throughout, and we derive upper and lower bounds on prevalence in the population under two alternative assumptions about the representativeness of non-COVID hospitalizations for the broader population. The first assumption is a relatively weak “hospital monotonicity" assumption that prevalence is at least as high among non-COVID hospital patients as it is in the general population. This assumption would be satisfied even if hospitalized patients are at greater risk of COVID because, for example, people who get into car accidents have more social interactions. The second assumption is a stronger "hospital independence" assumption that prevalence is the same among non-COVID hospital patients as it is in the general population. The resulting bounds are informative for prevalence at a point in time, not just cumulative prevalence. This is important because the information required to estimate the bounds is closely related to the simple statistics that many states already report. It appears possible for many states to use our method to report upper and lower bounds on prevalence in near real time using data that they already collect and report. In particular, states already report the COVID-hospitalization rate, as well as overall test rates and test positivity rates. To report a version of the upper and lower bounds we describe in this paper, states would simply have to compute testing and positivity rates among non-COVID hospitalizations. \footnote{Our methodological contribution is contemporaneous work by gelman2021. Like us, they study how non-COVID hospitalizations can improve COVID-19 surveillance. They study a hospital system which tests all of its patients and is therefore free of selection into testing. Whereas we develop bounds to account for selective testing, they focus on adjusting for demographic differences between hospitalized patients and the general community. In principle both approaches can be combined.}

Our first empirical contribution uses this framework to estimate a collection of upper and lower bounds on weekly prevalence of SARS-CoV-2 in Indiana. To operationalize the basic idea, we work with two definitions of non-COVID hospitalizations. The first, which we call non-Influenza- or COVID-like-illness (non-ICLI) hospitalizations, simply excludes all hospitalizations with diagnosis for ICLI cdcCovidCodes,afhsc2015. This definition would be easy to implement in many different hospital data sets and yields a large population of hospital patients. However, we also work with a definition using a narrower set of patients who are hospitalized for six groups of clear non-COVID causes: (i) cancer; (ii) appendicitis and vehicle accidents; (iii) labor and delivery; (iv) AMI and stroke; (v) fractures, crushes, and open wounds; and (vi) other accidents. The clear-cause analysis is more intuitive and transparent, but it might be harder to implement as part of a public health surveillance system. In addition, the clear cause analysis often involves the analysis of smaller samples of data. In practice, this may create a trade-off between the inferential uncertainty that stems from the width of the bounds under strong vs weak assumptions, and the statistical uncertainty that arises when estimates are based on fewer observations.

We show that, in Indiana, test rates are much higher among non-COVID hospital patients than in the general population. For example, in June, about 0.4 percent of the general population was tested in a given week, compared with 24 percent of non-COVID hospital patients. Test positivity rates are lower among non-COVID hospital patients than in the general population. In the general population, 4.5 percent of tests were positive. In contrast, among non-COVID hospital patients who were tested, only 2.2 percent of tests were positive. These testing and positivity rates can be combined to estimate upper and lower bounds on prevalence. Under the test monotonicity assumption, between 0.05 and 4.5 percent of the general population was infected with SARS-CoV-2 on June 15. Under the same test monotonicity assumption, the bounds are half as wide for non-COVID hospital patients: prevalence was between 0.7 and 2.2 percent.

By the end of November, with the pandemic worsening, the test rate had increased to about 3 percent in the general population, and to more than 40 percent in the non-COVID hospitalization population. Although testing rates had increased in the population, the test positivity rate had also risen substantially. As a result, the test monotonicity bounds on prevalence in the general population widened to 0.7 to 22.5 percent in November. In contrast, among non-COVID hospitalization patients, the bounds at the end of November were 3.1 to 6.2 percent - unambiguously higher than in June. Low testing rates in the population mean that selection bias in testing creates so much ambiguity that it is impossible to confirm that that prevalence was rising using the state's testing data. In contrast, the bounds on prevalence based on non-COVID hospital patients are tight enough that they show an unambiguous rise in SARS-CoV-2 prevalence from summer to late fall.

Under stronger assumptions, these tight hospital bounds can be interpreted as bounds on population SARS-CoV-2 prevalence. Specifically, under a hospital monotonicity assumption, the upper bound on prevalence in the non-COVID hospital population is a valid upper bound on population prevalence. And under the stronger hospital independence assumption, the upper and lower bounds on prevalence in the non-COVID hospital population are valid upper and lower bounds on population prevalence.

In the results section of the paper we present bounds under alternative hospital representativeness assumptions (none, monotonicity, independence) to allow readers to make their own judgements about which assumptions are credible, and to better understand how much information about prevalence is derived from the data versus assumptions. To characterize the statistical uncertainty in our work, we estimated 95% confidence interval around each set of upper and lower bounds using a version of the methods developed in the partial identification literature horowitz2000nonparametric,ImbensManski2004.

We also report bounds based on the non-ICLI definition as well as multiple clear cause definitions and find that the two approaches give similar results. In addition, we estimated the upper and lower bounds on prevalence among ICLI hospitalizations. We find, of course, much higher prevalence among this group, but our bounds still rule out very high prevalence. Specifically, we find that SARS-CoV-2 prevalence among ICLI hospitalizations is no higher than 50 percent at its peak, and no higher than about 30 percent by mid-June. This shows that there is value to testing even highly symptomatic patients, as their SARS-CoV-2 rate is far from 100 percent, and testing outcomes would be informative for treatment and quarantine decisions.

Our second empirical contribution is to assess the credibility of key assumptions that would be difficult to study in other data sets. ManskiMolinari2020 point out that the accuracy of SARS-CoV-2 virological tests is not well understood. Incorporating information about testing errors alters the bounds on prevalence. We use data on people who were tested twice in a two day period to shed some light on the fraction of people who test negative but are actually infected. We tentatively conclude that test errors have a negligible effect on the upper and lower bounds on prevalence reported in our paper.

We also assess the credibility of the hospital representativeness assumptions at the core of the paper. The most restrictive hospital independence condition assumes that SARS-CoV-2 prevalence is the same in the non-COVID and general populations, and the weaker hospital monotonicity condition assumes that prevalence is at least as high among the hospitalized. Although we are not able to directly validate these assumptions, we probe their credibility in three ways. First, we compute age-standardized estimates of the upper and lower bounds by estimating age group specific bounds and then averaging over the population age distribution. This helps address concerns that the hospital population may have a different age distribution than the general population. Second, we compare the hospitalization bounds to estimates of population prevalence obtained from tests of a random sample of Indiana residents in April and June NirWave1, NirWave2. Third, we examine the pre-hospitalization test rates of non-COVID hospitalized patients. These validity checks are roughly consistent with the the hospital independence assumption, and highly consistent with hospital monotonicity. This suggests that these assumptions might be considered reasonable in other states where it would be easy to estimate the upper and lower bounds but harder to perform elaborate validation exercises.

Overall, our results indicate that combining testing data and information on non-COVID hospitalizations may be a feasible and informative way of measuring SARS-CoV-2 prevalence. In the most recent weeks of our data (and under test monotonicity but without hospitalization data), we can only conclude that at most 20 percent of the overall Indiana population was actively infected. With hospitalization data and the hospital monotonicity assumption, we can conclude that at most 256 percent of the Indiana population was infected. The bounds are tight enough that we can conclude that prevalence rose unambiguously from mid-summer to late fall. Similar bounds could be constructed in other states using aggregate data on non-COVID hospitalizations and their testing and test positivity rates, potentially improving SARS-CoV-2 surveillance systems across the country.

Inferring COVID Prevalence from Incomplete Testing

The empirical goal of our study is to establish upper and lower bounds on the prevalence of active SARS-CoV-2 infections in the Indiana population in each week. Measuring prevalence is challenging because only a small fraction of people are tested in any given week. One way to overcome this challenge is to combine the data with assumptions about testing and infection risk that are strong enough to point identify prevalence. For example, under the assumption that the people who are tested are a representative sample from the overall population, prevalence among the tested equals the population prevalence. However this assumption is not very credible because of non-random selection into testing, which may occur because people who exhibit symptoms are tested at higher rates. In this paper, we explore the identifying power of weak instrumental variable assumptions related to hospital based testing. These assumptions are reasonably credible, but they only partially identify prevalence. This means that the assumptions establish upper and lower bounds on prevalence but do not restrict it to a single value.

The instrumental variable methods we use are standard in the literature on partial identification econometrics manski1999identification, manski2009identification. And they are closely related to methods from research on sample selection bias (e.g. heckman1979sample). But the approach may seem unfamiliar to readers who mainly work in the context of causal inference (e.g. AngristPischke2008). In the causal inference literature, an instrumental variable is a factor which affects treatment exposure but is mean-independent of potential outcomes and satisfies an exclusion restriction. In the sample selection literature, an instrumental variable affects sample selection, but is mean-independent of the outcome of interest and satisfies an exclusion restriction. The classic application involves the distribution of wages among women. Wages are observed for women who work, and unobserved for women who do not work. An instrumental variable is something that affects a woman's employment, but is not associated with her wage offer. Our application is conceptually similar: SARS-CoV-2 status is observed for people who are tested, and unobserved for people who are not tested. An instrument would be a variable that affects testing but is not otherwise associated with SARS-CoV-2 infection risk. The classical sample selection literature combines instrumental variable assumptions with a parametric model of the selection decision and the outcome to point identify parameters of the outcome distribution. The partial identification literature is non-parametric and relaxes the functional form assumptions. It also examines weaker forms of instrumental variable assumptions. For example, in this paper we consider replacing mean independence assumptions weak sign restrictions, generating “monotone" rather than mean independent instrumental variables manski2000monotone, Manski2020. All of these different approaches are “instrumental variable methods" because they rely on assumptions to restrict the relationship between an instrument and an outcome.

An appealing feature of the partial identification approach is that inferences can be tightly linked to assumptions, making it transparent how stronger assumptions can generate tighter inferences. Our approach introduces two monotone instrumental variable assumptions: test monotonicity and hospitalization monotonicity. We show how to use these relatively weak assumptions to narrow the worst-case (no-assumption) bounds on prevalence. To tighten the bounds further, we introduce a stronger assumption of hospital independence. \hyperref[fig:flowchart]{Figure \ref*{fig:flowchart}} gives a schematic representation of the data, assumptions, and results we present in the study. The figure shows how our assumptions and data are combined to generate inferences on SARS-CoV-2 prevalence, with increasingly strong assumptions yielding increasingly tight bounds. After introducing, discussing, and justifying our key assumptions, we return to the figure to summarize how different assumptions yield different bounds.

Notation and Worst Case Bounds

We use $i=1...N$ to index members of the population of Indiana. Let $C_{it} = 1$ indicate that person $i$ is currently infected with SARS-CoV-2 on date $t$. Leaving conditioning on the date implicit to reduce clutter, the population prevalence of active SARS-CoV-2 infections in Indiana at date $t$ is $Pr(C_{it} = 1) = \frac{1}{N} \sum_{i=1}^{N}C_{it} $. We are also interested in prevalence among hospital inpatients with various COVID- and non-COVID-related diagnoses. Let $H_{it}$ be a binary indicator set to 1 if the person was hospitalized with a specified non-COVID-related diagnosis. Then $Pr(C_{it} = 1 | H_{it} = 1) $ is the prevalence of active SARS-CoV-2 infections in the sub-population of people who were admitted with a non-COVID-related diagnosis on date $t$.

The central problem in estimating prevalence is that values of $C_{it} $ are unknown for most people on most days, because testing is rare. Let $D_{it} = 1$ if person $i$ was tested on $t$ and $D_{it} = 0$ if the person was not tested. $Pr(D_{it} = 1) $ is the proportion of the population tested on date $t$, where conditioning on $t$ is implicit. Continuing with the notation laid out above, $Pr(C_{it} | D_{it} = 1)$ and $Pr(C_{it} | D_{it} = 0)$ represent the prevalence of SARS-CoV-2 among people who are tested and not tested, respectively. The value of $C_{it} $ is observed for people with $D_{it} = 1 $, but unknown for people with $D_{it} = 0$. This means that $Pr(C_{it} | D_{it} = 0)$ is not identified by the data on testing and test outcomes. Uncertainty about the prevalence of SARS-CoV-2 among people who have not been tested is the main reason why it is difficult to estimate population prevalence using data from clinical tests.

In the absence of any distributional assumptions, the observed clinical tests partially identify prevalence overall, and in any sub-populations that can be defined by observable covariates. To see the point, use the law of total probability to decompose population prevalence:

equation[equation omitted — 129 chars of source]

The only unknown quantity on the right-hand side of the expression is $ Pr(C_{it}=1|D_{it} = 0)$, which is prevalence among people who were not tested. All that is known is that this value lies between $0$ and $1$. Substituting 0 and 1 for the unknown prevalence yields worst-case lower and upper bounds $L_w$ and $U_w$ on population prevalence:

alignat*{2} L_{w} & = \underbrace{Pr(C_{it}=1|D_{it} = 1) Pr(D_{it} = 1)}_Confirmed Positive Rate \\ U_{w} &= \underbrace{Pr(C_{it}=1|D_{it} = 1) Pr(D_{it} = 1)}_Confirmed Positive Rate + \underbrace{Pr(D_{it} = 0)}_Untested Rate .

These bounds defines the set of values for prevalence that are compatible with the observed data. From the definition of joint and conditional probability, we can rewrite the lower bound as the joint probability that a person is both tested and infected: $L_{w} = Pr(C_{it} = 1, D_{it})$. In other words, the worst case lower bound is the fraction of the population that has actually tested positive for SARS-CoV-2 on a given date. Since this group of people is definitely infected, the lower bound is the confirmed positive rate and it makes intuitive sense that overall prevalence cannot be any lower than confirmed positive rate. The lower bound is the population prevalence in a scenario where none of the untested people are infected. The upper bound gives the opposite extreme in which all of the untested people are infected. Put differently, the worst case upper bound is the confirmed positive rate plus the untested rate. The components of the bounds can be estimated using the appropriate proportions in test and hospital data. To account for statistical uncertainty in the estimates of the upper and lower bounds, we use methods developed in horowitz2000nonparametric and ImbensManski2004 to construct confidence intervals on the bounds. See \hyperref[app:ci]{Appendix (ref)} for details.

Test monotonicity

The worst case bounds are very wide when testing rates are low. To narrow the bounds, ManskiMolinari2020 propose the “test monotonicity" condition. In our context, test monotonicity requires

assumption(Test monotonicity) $Pr(C_{it}=1|D_{it}=1) \geq Pr(C_{it}=1|D_{it} = 0)$

Assumption (ref) requires that the prevalence of SARS-CoV-2 is at least as high in the tested population as it is in the untested population. This assumption is untestable because prevalence is unknown in the tested sub-population. However, the assumption has credibility because virological tests are typically allocated to symptomatic individuals, who have a higher than average likelihood of infection. To apply the assumption, recall that the identification problem is that prevalence in the untested sub-population -- $Pr(C_{it}=1 | D_{it}=0)$ -- is unknown. Assumption (ref) implies that $0 \leq Pr(C_{it} =1| D_{it}=0) \leq Pr(C_{it=1} | D_{it}=1)$. Substituting the values $0$ and $Pr(C_{it} =1| D_{it}=1)$ into \hyperref[eq:tp]{Equation \ref*{eq:tp}} yields a new set of bounds on overall prevalence:

align*[align* omitted — 186 chars of source]

The new upper bound is the prevalence in the tested population, which is often called the test positivity rate. This new test monotonicity upper bound will be lower than the worst-case upper bound as long as prevalence in the tested sub-population is less than 1. In our data, test rates are often less than 1 percent and positivity rates in the population are often 10 percent or less, so this assumption brings the upper bound down from 99 percent to 10 percent or less. A similar assumption could be made for the non-COVID hospitalization population, yielding bounds $L_{m}^H$ and $U_{m}^H$ on prevalence among the non-COVID hospitalized population.

Test monotonicity appears to be a credible assumption given the reliance on symptomatic testing for SARS-CoV-2 in the United States. Nevertheless, it is worth considering scenarios under which the assumption would fail. In the abstract, test monotonicity could fail if the demand for voluntary testing is higher among people with lower risk of infection. The “worried well” phenomenon that occurred the HIV/AIDS epidemic is one example cochran1989women. Theoretically, a worried well effect could arise because of preferences that increase a person's demand for testing are correlated with preferences that reduce the person's SARS-CoV-2 infection risks. A worried well could also arise if testing is cheaper or more convenient for sub-populations that are actually at low risk of infection. This could happen -- for example -- because of socio-economic, racial, or geographic disparities in the severity of the epidemic and available testing capacity. Although these scenarios are plausible, they seem unlikely to dominate the effects of symptomatic testing procedures that support test monotonicity.

Inferring Population Prevalence From Non-Covid Hospital Patients

Test monotonicity can be used to narrow the the bounds on prevalence in the population and also among non-COVID hospitalized patients. Because testing rates are much higher in hospitals than in the general population, the bounds on prevalence in hospitalized sub-populations are much narrower. Thus, assumptions that link hospital and population prevalence may be a powerful way to reduce uncertainty about population prevalence. We pursue two types of assumptions that enable extrapolation from non-COVID hospital populations to the general population: (i) hospitalization monotonicity and (ii) hospitalization independence. These are both forms of hospital instrumental variable assumptions, and we refer to them collectively as hospital IV assumptions.

Hospitalization Monotonicity

In some contexts, it may be plausible to assume that people who are hospitalized for a non-COVID related health condition will not have a lower risk of SARS-CoV-2 infection than than the general population. Stated formally, the hospitalization monotonicity assumption is:

assumption(Hospitalization Monotonicity) $Pr(C_{it}=1 | H_{it}=1) \geq Pr(C_{it}=1)$

Assumption (ref) requires that the prevalence of active SARS-CoV-2 infections among non-COVID hospital patients is at least as high as the prevalence of active infections in the general (non-hospitalized) population. When prevalence is bounded in the hospitalized and general populations, the hospital monotonicity assumption may further reduce the width of both sets of bounds by ruling out values that would violate the restriction. To see the idea, suppose that we apply Assumption (ref) in both the hospitalized population and the general population. Let $U_{m}^{H}$ and $L_{m}^{H}$ represent the upper and lower bounds on SARS-CoV-2 prevalence in the hospitalized sub-population under the test monotonicity assumption. Following the notation above, $U_m$ and $L_m$ represent the test monotonicity bounds in the general population. Layering the additional hospitalization monotonicity assumption, (Assumption (ref)) creates a cross-population restriction that implies that the upper bound on population prevalence $(U_m)$ cannot be larger than the upper bound on hospital prevalence $(U_{m}^{H})$. Maintaining both Assumption (ref) (test monotonicity) and Assumption (ref) (hospitalization monotonicity) yields the following upper bound on SARS-CoV-2 prevalence in the population:

align*[align* omitted — 265 chars of source]

In our data, the hospital upper bound is typically lower than the population upper bound, so the hospital monotonicity condition implies that the positivity rate among non-COVID hospitalizations is an upper bound population prevalence. In principle, the monotone hospitalization assumption could be used to derive a potentially higher lower bound on prevalence in the hospital population. But this lower bound condition is essentially non-binding in practice, so we ignore it here.

Hospitalization Independence

In some situations it may be credible to assume a stronger condition than monotone hospitalizations, that hospitalization for a non-COVID health condition is mean-independent of SARS-CoV-2 infection risk. Formally, the hospitalization independence assumption can be written:

assumption(Hospitalization Independence) $Pr(C_{it} =1| H_{it} = 1) = Pr(C_{it}=1)$

The independence assumption is stronger than the weak sign restriction embodied by the hospitalization monotonicity assumption. Assumption (ref) implies that people who are hospitalized for a specified non-COVID health condition have the same probability of being infected with the virus as the general population. An equivalent statement of Assumption (ref) is that people who are infected with SARS-CoV-2 have the same probability of being hospitalized for a non-COVID condition as people who are not infected with SARS-CoV-2, which implies $Pr(H_{it} = 1 | C_{it}=1) = Pr(H_{it} = 1 |C_{it}=0) $.

Under the hospital independence assumption, the bounds on population prevalence are defined by the intersection of the hospital and population bounds. As before, $U_m$ and $L_m$ are the upper and lower bounds on prevalence in the general population under the test monotonicity assumption alone. Likewise, $U_{m}^{H}$ and $L_{m}^{H}$ are the upper and lower bounds on prevalence in the hospital population under test monotonicity. Assumption (ref) implies that prevalence is the same in the hospitalized and general populations. That implies that the lower bound on prevalence in the general population must be no lower than the larger of the two lower bounds. By the same logic, the upper bound on prevalence in the general population must be no larger than the smaller of the two upper bounds. Formally, layering the hospital independence assumption on top of the test monotonicity assumption yields a new set of lower and upper bounds:

align*[align* omitted — 348 chars of source]

As it turns out, $U_{m,ind} = U_{m,h}$, so that the upper bound on prevalence is the same under (ref) and (ref) as it is under Assumption (ref) and Assumption (ref). What the hospital independence assumption buys us is a potentially tighter lower bound, which is now the greater of the lower bounds on population and hospital prevalence under test monotonicity. In practice we find that the lower bound is always higher in the non-COVID hospitalization sub-population than in the general population, so in practice this assumption implies that the lower bound on population prevalence is the confirmed positive rate among non-COVID hospitalizations.

Hospital Instrumental Variables In Perspective

Assumptions (ref) and (ref) are strong instrumental variable restrictions that may hold for some non-COVID hospitalizations and not others. The independence assumption, Assumption (ref), is the strongest condition. It would be satisfied if non-COVID hospitalizations occur at random in the population, from the perspective of SARS-CoV-2 infection risk. For example, people who are hospitalized for cancers that are mainly due to genetic rather than behavioral factors are unlikely to be selected on differential SARS-CoV-2 risk. And people who are hospitalized because of car accidents are likely a fairly representative draw from the population of regular drivers. Hospitalization for stroke and heart attack also seems plausibly unrelated to a person's SARS-CoV-2 risk exposure. People who are in the hospital to deliver a baby or who are having “routine" inpatient procedures like joint replacement are selected on the basis of decisions and behaviors that are mainly determined before the epidemic. They too might be expected to have the same distribution of SARS-CoV-2 infection risks as the general population.

Although hospital independence is plausible for some health conditions, it is also easy to think of reasons why it might fail. The basic demographics of people with particular health conditions may differ from the general population. This is easiest to notice for specific health conditions. Young people will be underrepresented among stroke and heart attack patients. Pregnant women are selected on gender. And older women will be underrepresented among pregnant women. In the empirical section of the paper, we report estimates of age-standardized upper and lower bounds that correct for differences in the age distribution across different hospital samples, which addresses some of these issues. But a more general concern with the independence assumption is that people who are hospitalized are simply sicker and more exposed to a variety of health risks than the general population. Furthermore, some types of hospital events may occur in part because people are engaged in activities that are higher risk during the epidemic. For example, people who drive regularly might have higher SARS-CoV-2 risk than the general population because they have more social interactions.

These concerns motivate our effort to move away from full independence and consider the weaker hospitalization monotonicity condition, Assumption (ref) (monotone hospitalization). Of course, it is possible that even this weaker condition does not hold for some types of non-COVID hospitalizations. For instance, it would fail for non-COVID hospitalizations that are generated by a health condition that led a person to be more cautious about avoiding SARS-CoV-2 infection than the general population. In that case, we would expect the hospitalized group to have lower infection rates than the general population. While our discussion has focused on how individual behaviors might cause violations of our identifying assumptions, we emphasize that Assumptions (ref), (ref), and (ref) are restrictions on population and sub-population averages. The existence of a single person whose behavior or characteristics are counter to the assumption does not necessarily imply a violation of the population level condition. Our approach is to present upper and lower bounds under a range of weak and strong assumptions and for multiple non-COVID hospital sub-populations. Readers can inspect the evidence and judge for themselves which assumptions are most plausible and what that implies about the prevalence of COVID-19 during the epidemic.

Measurement Error in Testing

Virological tests for the presence of SARS-CoV-2 may not be perfectly accurate, and so far there are no detailed studies of the performance of the PCR tests that Indiana is using to test people for SARS-CoV-2. To clarify how error-ridden tests complicate our prevalence estimates, we augment the notation to distinguish between test results and virological status. We continue to use $C_{it}$ and $D_it$ to represent a person's true infection and testing status at date $t$. But now we introduce $R_{it}$, which is a binary measure set to 1 if the person tests positive and 0 if the person tests negative. Using this notation, $Pr(C_{it} = 1 | D_{it} = 1, R_{it} = 1)$ is called the Positive Predictive Value (PPV) of the test among people who are tested and who test positive. $Pr(C_{it} = 0 | D_{it} = 1, R_{it} = 0)$ is called the Negative Predictive Value (NPV) among people who are tested and who test negative. $1-NPV = Pr(C_{it} = 1 | D_{it} = 1, R_{it} = 0)$ is the fraction of people who test negative who are actually infected with SARS-CoV-2.

Our initial worst case bounds assumed no test errors. Relaxing that assumption yields a different set of upper and lower bounds on prevalence. Following ManskiMolinari2020, we assume that (i) $PPV = 1$ so that none of the positive tests are false, but (ii) $Pr(C_{it} = 1 | D_{it} = 1, R_{it} = 0) \in [\lambda_{l}, \lambda_{u}]$. The second condition imposes a bound on $1-NPV$, which is the fraction of people who test negative who are actually infected. Under these two restrictions, the new worst case bounds work out to:

align*[align* omitted — 180 chars of source]

Allowing for test errors increases the worst case lower bound by the best-case fraction of missing positives, and increases the worst case upper bound by the worst-case fraction of missing positives. Similar expressions hold for prevalence bounds under test monotonicity and other independence assumptions.

The upshot is that knowledge of test accuracy is important for efforts to learn about prevalence. In their study of the cumulative prevalence of SARS-CoV-2 infections, ManskiMolinari2020 computed upper and lower bounds on prevalence under the assumption that $\lambda_{l} = .1$ and $\lambda_{u}=.4$, citing peci2014performance. ManskiMolinari2020 view this choice of $.1 \leq 1-NPV \leq .4$ as an expression of scientific uncertainty about test errors, and they refer to the resulting prevalence bounds as “illustrative.” However, the structure of the test error bounds makes it clear that assumptions about the numerical magnitude of test errors have inferential consequences. For example, setting $ \lambda_{u}=.4$ implies that, regardless of the outcome of the test, at least 40 percent of the people who are tested for SARS-CoV-2 are infected.

Although there is little published evidence on the properties of the SARS-CoV-2 PCR test, previous research suggests that PCR test errors are uncommon in other settings. For example, peci2014performance study the performance of rapid influenza tests using PCR-based tests as a gold standard. PCR tests are used as a gold standard because they are expected to have very high PPV and NPV.

To shed more light on test errors, we constructed a sample of people who are tested and retested in a short interval, specifically people who were (i) tested on day $t$, (ii) not tested on day $t-1$, and (iii) were tested again on day $t+1$. We show in \hyperref[app:npv]{Appendix \ref*{app:npv}} how these data can be used to estimate error rates, under assumptions of random retesting and no false positives. Our data include 835,000 test-retest events. Using $R1_i$ and $R2_i$ to represent the results of a person's first and second test, we found that $Pr(R1_i = 1, R2_i = 1) = .11 $ and $Pr(R1_i = 0, R2_i = 0) = .88$ among the people in the twice-tested sample. The two tests were discordant for less than 1 percent of the twice-tested sample. These results imply a negative predictive value of 99.8 percent.

This estimate of NPV depends on our assumptions of random retesting and no false positives. While the no false positive assumption appears plausible, random retesting is not necessarily satisfied. In particular, a patient with a suspected COVID case who initially tests negative may be retested; this selective retesting would bias us towards finding false negatives. Another reason for retesting is delays in processing results. If a patient was tested prior to a planned hospitalization, and the result is not available at the time of the hospitalization, the attending physician may order an in-hospital test, which would be available within hours. This type of retesting is less likely to lead to bias. As we explain in \hyperref[app:npv]{Appendix \ref*{app:npv}}, we can test for selection into retesting by looking for symmetry in test results. Under random retesting (and no false positives), the sequences “positive-then-negative" and “negative-then-positive" should be equally likely. In practice we find that “negative-then-positive" is slightly more common, meaning that our test-retest sample likely disproportionately selects people with initial false negatives.

Overall, we think that a plausible value for $\lambda_{l}$ is nearly zero, and a plausible value for $\lambda_{u}$ is 0.005. Accounting for test errors in this range would have almost no effect on the upper and lower bounds reported in the paper. Test-retest data are potentially informative about test errors, but a limitation of is that retested people are not necessarily representative of the population.

Summary and data requirements

The analysis above shows how to combine test and hospital data with increasingly strong assumptions to obtain bounds on population prevalence. The overall approach requires data on the tested population that can be linked to hospital inpatient records that ideally include diagnosis codes to identify different types of hospitalizations. The diagnosis information is important because the hospitalization instrumental variable assumptions are more plausible for some types of hospitalizations. At a minimum, it is important to exclude non-COVID hospitalizations.

\hyperref[fig:flowchart]{Figure \ref*{fig:flowchart}} summarizes our methodological results and serves as a guide for interpreting our empirical findings. We work with three main assumptions: two weak monotonicity assumptions, and one conditional independence assumption. The flow chart shows which assumptions yield which bounds on population prevalence. Using only data and no assumptions, we have worst-case bounds for prevalence in the general population and for hospitalized sub-populations. Under Assumption (ref), test monotonicity, we have tighter bounds on prevalence for both populations.

Assumptions (ref) and (ref) let us extrapolate from the hospitalized sub-population to the general population. Under Assumption (ref), hospitalization monotonicity, the upper bound on population prevalence tightens to the least upper bound under test monotonicity among the hospitalized subpopulation and the general population. Under Assumption (ref), hospitalization independence, the lower bound on population prevalence tightens to the greatest lower bound among the hospitalized subpopulation and the general population.

The bounds turn out to be fairly simple objects. Under test monotonicity the lower bound is the confirmed positive rate, the share of the population that tests positive. The upper bound under monotonicity is the test positivity rate, the share of tests that are positive. Under hospital monotonicity, the upper bound in the general population becomes the test positivity rate among non-COVID hospitalizations. And under hospital representativeness, the lower bound in the general population becomes the confirmed positive rate among the non-COVID hospitalized subpopulation. When presenting our results, we show confirmed positive rate and test positivity rates for the overall population and for the non-COVID hospitalized subpopulation. Thus readers can easily calculate bounds under any assumption, such as test monotonicity alone or combined with hospitalization monotonicity.

An appealing feature of these bounds is that they can be calculated with little additional data beyond what public health organizations already report. Every state already reports the number of tests and the number of positive tests, and many states report the number of COVID-related hospitalizations. \footnote{See, e.g., covidTracking.} States would only have to report test and positivity rates for non-COVID-related hospitalizations. This appears possible because many states already report "suspected" or "under investigation" COVID hospitalizations, defined as hospitalized patients exhibiting COVID-like illness. Some states actually report both the number of hospitalizations of patients with COVID- or influenza-like illness and, separately, the number of hospitalizations of patients with a positive SARS-CoV-2 test (e.g. Arizona and Illinois azDashboard,ilHosp, ilSyndromic). \footnote{States reporting both confirmed SARS-CoV-2 hospitalizations and hospitalizations of suspected cases or cases under investigation include California caDashboard, Colorado coDashboard, Mississippi msDashboard, Tennessee tnDashboard, and Vermont vtDashboard. } Thus states have the capacity to identify ICLI-related hospitalizations and link hospitalization and testing data.

Indiana Hospital and Testing Data

Data sets

Our analysis is based on two main data sources managed by the Regenstrief Institute. First, we obtained data on the near universe of polymerase chain reaction (PCR) tests for SARS-CoV-2 conducted in Indiana between January 1, 2020 and December 18, 2020. Second, we obtained data on all inpatient hospital admissions from hospitals that belong to the Indiana Network for Patient Care (INPC), which is a health information exchange that centralizes and stores data from health providers across the state of Indiana, including all hospitals with emergency departments. \footnote{See grannisINPC2005 for more details.} The hospital data are derived from the same database that the state uses for reporting hospitalizations on its dashboard inCovidDashboard. We linked the individual testing and hospital inpatient records using an encrypted common identifier. Of course, only a subset of hospital patients are tested and only a subset of tested people appear in the inpatient hospital data. For both data sets, we are also able to link the individual testing and inpatient records with basic demographic data collected by the INPC; this information is available only for a subset of patients.

The test data contain individual records for nearly all of the SARS-CoV-2 tests conducted in Indiana during 2020. A small number of tests are excluded from our data because some institutions that conduct tests provide data to INPC but do not allow the data to be used for research purposes, and some tests are not included in our sample because the data are reported with a delay. The consequence of these exclusions is that we are missing some tests, which will result in a reduced lower bound in our framework. Despite these exclusions, our data set tracks the state's official case counts quite closely until November, 2020, when the state count exceeds ours; see \hyperref[fig:compare_to_state]{Appendix Figure \ref*{fig:compare_to_state}}. As an aggregate summary, our data contain 383,976 people with a positive test as of December 15, 2020, whereas the state reports 439,916 mphData. \footnote{We compare positives rather than tests because some testing institutions appear not to report negative tests to the state. (Such institutions are not included in our data.)} Each test record in our testing data includes information on the date the test specimen was obtained, the outcome of the test (positive, negative, or inconclusive), and a patient identifier that we use to link the test data to demographic files and inpatient hospital files.

The hospital inpatient data contain separate observations for each admission. We always observe admission time and a patient identifier that we use to link to the test data and inpatient files. Discharge time and diagnosis information are only observed for a subset of admissions. (Not all fields are available for all admissions because different institutions contribute different information to the INPC.) Because the INPC data come from health care providers and payers, the same hospitalization can appear in the data set multiple times. To de-duplicate these records, we keep one observation per admission time (defined second-by-second), keeping the observation with the most diagnosis codes. This procedure could in principle result in us dropping relevant diagnostic information, but in practice this happens only very rarely. There are 707,734 admissions that are unique in terms of patient identifier and time, and 613,036 non-unique admission records. In 788 of these non-unique cases are there multiple records with the same number of diagnostic codes; we keep one of these multiples at random. In 122 of the 613,064 cases there are multiple records with diagnosis codes that provide conflicting first diagnoses.

Measuring tests and cases

In-hospital testing, positivity rate, and confirmed positives The fraction of people who are tested in the hospital is an important quantity of interest in our analysis because hospital patients are tested at a higher rate than the general population, and because hospital testing may be less correlated with COVID symptoms. Our data do not distinguish whether a person was tested in the hospital or whether the test was initiated independently of the hospital visit, but we can match tests to hospitalizations based on the test date and hospitalization date. We say that a hospitalized patient is tested in-hospital if she had at least one SARS-CoV-2 test dated between 5 days prior to admission to 1 day after admission, and we say she had a positive if she had at least one positive test in that window. Several considerations justify this definition. First, we want to include tests that are part of the admission. For patients with planned procedures, these tests may happen a few days prior to admission. For patients admitted from the emergency department, these tests may happen on the day of admission or the day after. Consistent with this view, \hyperref[fig:hospital_test_times]{Appendix Figure \ref*{fig:hospital_test_times}}, shows that among non-ICLI hospitalizations, test rates begin to rise several days before admission and fall rapidly 3-4 days after admission. Second, we want keep the post-admission window relatively short to ensure that we do not pick up hospital-acquired SARS-CoV-2 infections, although this may not be a major concern since rhee2020incidence indicate that hospital-acquired SARS-CoV-2 infections are quite rare. Finally, we limit the window to seven days, so that we can compare hospital testing rates to population testing rates.

Population Testing and Positivity Rates In some analyses we compare hospital testing and positivity to population testing and positivity. Hospital testing and positivity are defined over a week-long span for a given hospitalization. To make the comparison with the general population clean, we examine test rates and positivity in a given week-long period. We say that a person was tested if she was tested at least once in a given week, and we say she was positive if she was positive at least once in that period.

Sample construction

Throughout, a patient is in the “test sample" if they are tested at least once. We say a patient is in the “inpatient sample" if they are hospitalized at least once. We limit our analysis to admissions with non-missing diagnostic information, and we say a patient is in the “inpatient diagnoses sample" if they meet this restriction. This limitation is important because diagnostic information is necessary for distinguishing COVID-related admissions from non-COVID-related admissions.

ICLI and non-ICLI Hospitalizations We construct three analytic samples from the inpatient data. We start by defining hospitalizations for influenza- and COVID-like illness (ICLI) using ICD-10 codes. We identify admissions with any of a standard set of ICD-10 codes for ICLI following afhsc2015. Then we identify admissions with any of the additional ICD-10 codes that the CDC recommends using for coding COVID hospitalizations cdcCovidCodes. Both the influenza-like and COVID-like diagnoses include general symptoms such as cough or fever, as well as more specific diagnoses like acute pneumonia, viral influenza, or COVID-19. We classify hospitalizations as ICLI-related if they have any influenza- or COVID-like illness (ICLI) diagnoses, and we classify hospitalizations as non-ICLI if they are not ICLI-related. \hyperref[app:cause]{Appendix \ref*{app:cause}} lists the ICD-10 codes used to define the analytic samples. ICLI hospitalizations may contain diagnostic codes for other (non-ICLI) illnesses. A given hospitalization cannot be both ICLI and non-ICLI, but a given patient with multiple hospitalizations can have ICLI and non-ICLI hospitalizations. \footnote{We view it as an advantage of this definition that a given patient can have a non-ICLI and an ICLI hospitalization, because we want our hospital-based bounds to count initially mild COVID-19 cases that nonetheless develop into serious cases. For example, a patient hospitalized for labor and delivery with asymptomatic COVID that later develops into a serious case, requiring an ICLI hospitalization, should be counted as an initial negative but a subsequent positive. Our definition counts this patient properly.}

We view the non-ICLI sample as a useful starting point for our analysis for two reasons. First, our hospital IV assumptions are most plausible for hospitalizations that are not obviously COVID-related, and this sample meets that criteria. Second, as we have noted, many states already classify hospitalizations as ICLI-related; thus non-ICLI hospitalizations are identifiable and measurable in near-real time, so this sample can be studied more broadly.

We acknowledge, however, that the non-ICLI sample may not satisfy the hospital IV assumptions for at least two reasons. First, it may condition on COVID itself, since a patient with a reported COVID diagnosis would be excluded from it. (In practice we observe many patients with positive COVID tests but no COVID diagnosis.) Second, COVID is a new disease with heterogeneous symptoms, so even if a patient is hospitalized because of COVID, she may not have one of our flagged diagnoses, and we may incorrectly call her hospitalization non-ICLI YangEtAl2020.

Clear-cause sample To avoid these problems, we study a third sample, which we call the “clear cause" sample. These are hospitalizations with a clear cause that is not obviously COVID-related. We define clear-cause hospitalizations as hospitalizations with a diagnosis code for labor and delivery, AMI, stroke, fractures, crushes, open wounds, appendicitis, vehicle accidents, other accidents, or cancer. For all of these conditions except cancer, we flag hospitalizations with a diagnosis at any priority. For cancer, we flag hospitalizations with a cancer diagnosis code as the admitting diagnosis, the primary final diagnosis, or any chemotherapy diagnosis. Note that we do not include among the clear causes respiratory disorders or other diagnoses contributing to our ICLI measure. However admissions can have many diagnosis codes, so it is possible for a clear cause hospitalization to be an ICLI hospitalization.

We view the clear-cause sample as important for two reasons. First, we believe the hospital IV assumptions are most plausible for this sample, so we believe the bounds on prevalence are most likely to be valid. Second, we view the clear-cause sample as offering a test of the validity of the non-ICLI sample. To the extent that the two samples generate similar bounds, we can be more confident that the non-ICLI sample is informative of broader population COVID prevalence, despite the problems with the non-ICLI classification. This would be valuable because classifying hospitalizations as ICLI-related or not requires less information than ascertaining a clear cause of the hospitalization.

Summary statistics We show summary statistics for all of our samples in \hyperref[tab:ss]{Table \ref*{tab:ss}}, as well as for the state as a whole (from Census Fact Finder and inCQF2019). The average tested and hospitalized patient is substantially older than the population as a whole, and also more likely to be female. Because the tested and hospitalized samples are not age representative of the general population, in what follows we reweight all samples to match the population age distribution. \footnote{Specifically, we calculate test rates and positivity rates in week-by-age-group cells, for age groups 0-17, 18-30, 30-50, 50-64, 65-74, and 75 and older. Then we average these age-specific rates across the age groups, weighting each group by its population share.} The tested and hospitalized samples are fairly similar to the general population in terms of racial composition. Limiting the inpatient sample to admissions with diagnoses reduces our sample size substantially, but it does not appear to change its demographic profile. About one-in-three Hoosiers has ever had a COVID test, whereas about half of hospitalized Hoosiers have had a test.

Although hospitalized patients are about 44 percent (i.e. 49%/34%) more likely to have ever been tested than the general public, during the period of their actual hospitalization they are vastly more likely to be tested, as we show in \hyperref[fig:test_rates]{Figure \ref*{fig:test_rates}}. The figure plots age-adjusted weekly testing rates for the whole population and each of our hospitalized samples. \footnote{We report the exact values of each of the test rates and the weekly number of admissions in \hyperref[tab:test_rates]{Appendix Table \ref*{tab:test_rates}}. We report age-unweighted test rates in \hyperref[tab:test_rates_unwtd]{Appendix Table \ref*{tab:test_rates_unwtd}}.}. The testing rate in the general population grew from 0.2 percent in April to between 0.4 and 0.6 percent in May and June, and it peaked at about 3 percent in mid-November. So despite increasing more than 10 fold, the weekly test in the Indiana population rate remained below 5 percent for the entire period covered by our study. In contrast, people hospitalized for ICLI were tested at a very high rate, between 60 and 75 percent in most weeks. Testing rates among non-ICLI hospital patients and among the clear-cause non-COVID hospital patients were lower than the ICLI sample but much higher than the population overall, 25-40 percent in May and later months. Testing rates among non-COVID hospital inpatients are 10-25 times higher than testing rates in the general population, but they are typically less than half as high as testing rates in the ICLI population. Despite their very high test rates, hospitalisations are sufficiently rare that hospitalized patients account for a low share of overall tested population, about 4.5 percent in the typical week.

\hyperref[fig:test_rates]{Figure \ref*{fig:test_rates}} shows that although hospitalized patients are tested more often than the general population, they are not always tested. ICLI patients are tested only about two-thirds of the times, and non-ICLI patients only about a third of the time. Several factors contribute to observed tests rates. Highly symptomatic patients may not be tested because a test would not necessarily influence care, and could generate a false negative, and testing capacity was sometimes limited. They also might not receive a SARS-CoV-2 test if they had a positive influenza test, as that provides an alternative explanation for the symptoms. For symptomatic patients, hospital policy seems to be to encourage testing, but not always require it. The Chief Medical Officer of one large hospital system in the state indicted that asymptomatic patients would typically be tested at admission, but this might vary across hospitals depending on their capacity to isolate patients in private or semi-private rooms weaver2020. Another Chief Medical Officer of a large hospital system other reported that testing was at times based on capacity, but patients coming into particular divisions were more likely to be tested, as were patients coming in for operations crabb2020. Our personal experience was that hospitals encouraged testing but did not strictly require it. \footnote{One of us had a child born in August at a hospital in our sample. The mother was encouraged to obtain a SARS-CoV-2 test prior to admission, but the father (who attended the birth) was not. The mother's test result was not available until after admission (we are happy to report it was negative) and the hospital did not require a rapid test.} Overall we view this anecdotal evidence as indicating that non-ICLI hospitalizations provides a strong encouragement but not a mandate for testing among asymptomatic or mildly symptomatic patients. Greater testing of mildly symptomatic patients would be consistent with our test monotonicity assumption applied to non-ICLI hospitalizations. \footnote{Interestingly, test monotonicity might not hold for ICLI patients, as testing might be redundant for the most symptomatic patients. Clinicians indicated that this is more likely in an emergency room context.}

Bounds on COVID-19 prevalence

Bounds by broad samples

The high test rates among the hospitalized populations shown in \hyperref[fig:test_rates]{Figure \ref*{fig:test_rates}} imply tight bounds on population prevalence under our monotonicity assumptions. We plot these bounds in \hyperref[fig:bounds]{Figure \ref*{fig:bounds}} and report them in \hyperref[tab:bounds]{Appendix Tables \ref*{tab:bounds} and \ref*{tab:bounds2}}. These bounds are age-adjusted using the same age-group weighting scheme we used for test rates. We report unweighted bounds in \hyperref[tab:bounds_unwtd]{Appendix Tables \ref*{tab:bounds_unwtd} and \ref*{tab:bounds_unwtd2}}. Each panel in \hyperref[fig:bounds]{Figure \ref*{fig:bounds}} plots the test monotonicity bounds for one of our three hospitalizations samples, along with the test monotonicity bounds in the overall population based on the test data. The dashed lines are 95% confidence intervals around the estimated upper and lower bounds. (The estimates in the population test data are precise enough that the confidence intervals are indistinguishable from the bounds.)

Several patterns are clear in the figure. First, the ICLI hospitalized population has higher upper and lower bounds on prevalence than the other groups. For the ICLI patients, the prevalence bounds begin at 6-18 percent in the first week of our sample, increase to 33-39 percent in the last week of March, decline steadily to roughly 12-18 percent in the summer and 8-15 percent in early fall, and increase dramatically in November. Although high, these bounds rule out the possibility that even a majority symptomatic patients are infected with SARS-CoV-2 in any week. In all weeks weeks the ICLI bounds lie outside the other groups' bounds, implying unambiguously higher COVID prevalence among patients hospitalized with influenza- or COVID-like illness. This unsurprising separation shows that the data are sensible and that the bounds are at least informative enough to tell apart these highly distinct populations.

The second clear pattern in the figure is that the test monotonicity prevalence bounds are much tighter for the non-ICLI and clear-cause hospitalization samples than for the all-test sample. In fact the bounds for both of these hospitalizations samples are always contained within the population-based bounds. At their tightest, the bounds for the all-test sample are as wide as 0.05 to 4.5 percent, for the week of June, 12. In that week the bounds for non-ICLI hospitalizations are [0.7%, 2.2%] and for clear-cause hospitalizations they are [0.5%, 1.8%]. These tight bounds imply that the hospitalization data could substantially reduce uncertainty about population prevalence under stronger assumptions like hospitalization monotonicity or hospitalization independence.

Third, the bounds for the non-ICLI hospitalization sample and for the clear-cause hospitalization sample are nearly indistinguishable. The only noticeable difference is that the upper bound for non-ICLI hospitalization is perhaps slightly higher. This fact is important because non-ICLI hospitalizations are potentially easier to measure, but they may be negatively selected in the sense that by construction they may exclude COVID-likely cases. The similarity of the non-ICLI bounds with the bounds for the clear-cause sample (which is not selected based on COVID-likelihood) provides some evidence in support of using non-ICLI hospitalizations to measure general prevalence.

The upper bound for all samples shows a U-shaped pattern, with lower and upper bounds high in the spring, falling in the summer, and rising rapidly in the fall. This pattern does not necessarily indicate that prevalence follows a U-shaped trend, because the lower bound for the population as a whole remains fixed at essentially zero. However the non-ICLI hospitalization bounds are sufficiently tight to confirm that prevalence is lower in mid-summer than late fall. For example, the upper bound the week of September 18 is 1.6 percent; this is lower than the lower bound in any week after October 30. Thus under our monotonicity assumptions, our hospital based bounds are tight enough to show that prevalence unambiguously rose from summer to fall. This rise in prevalence is also evident in mortality data inCovidDashboard, but the mortality data show it only with a lag, so the results here show that combining hospital and testing data can be useful for high-frequency monitoring of a pandemic.

Finally, the figures show that although the hospital-based bounds are estimated with less precision than the population-based bounds, the estimates are nonetheless precise enough that in many weeks the confidence intervals remain inside the population bounds. This is especially true for the non-ICLI hospitalizations, which are more common than the “clear cause" hospitalization. Thus while there is a precision trade-off arising from using a smaller sample with a higher test rate, this trade-off favors the non-ICLI hospitalizations, at least for our data.

Bounds by cause of admission

Our overall clear-cause hospitalization sample pools many distinct causes, including among others labor and delivery, vehicle accidents, and other accidents, including falls. In principle these hospitalizations may differ in their SARS-CoV-2 infection risk. One might worry, for example, that pregnant women are especially cautious and careful not to become infected. In contrast, people who get into vehicle accidents during the epidemic might be a less cautious group either because they are not careful drivers, or because they are out of the house at all.

Since the credibility of key assumptions may vary across different clear causes, we estimated test rates and bounds separately for each of our clear causes of hospitalizations. Because each individual cause has relatively few hospitalizations, we aggregate across all time periods to form these estimates. We focus on nine sets of causes: AMI (i.e. heart attack), appendicitis, cancer, fractures, labor/delivery, non-vehicles accidents, stroke, vehicle accidents, and wounds. These six groups have between 2,000 and 14,000 hospitalizations each. The age profile varies considerably across groups, as we show in \hyperref[tab:ss_group]{Appendix Table \ref*{tab:ss_group}} for age profiles of admitted patients by cause of admission. All ages are represented in the cancer sample. Appendicitis and vehicle accidents both afflict more young people. AMI, stroke, and other accidents---primarily falls---afflict older people; and labor and delivery is limited, of course, to women of childbearing age. \footnote{These groups are not necessarily mutually exclusive, and in particular there is overlap between injury and accidents.} Because not all age groups are represented in every category, we do not age-weight these results.

We report bounds by clear cause of admission in \hyperref[tab:bounds_cause]{Table \ref*{tab:bounds_cause}}. Labor and delivery is tested at the lowest rate, about 20 percent; the other groups are tested 25-40 percent of the time. The bounds are similar across all groups; the (time-pooled) lower bound ranges from .6 percent for vehicle accidents to 3.2 percent for AMI, and the upper bound ranges from 2.1 percent for cancer to 8.7 percent for other accidents. These bounds therefore separate slightly; cancer and vehicle accident stand out as low-prevalence admission types, whereas AMI and other accidents are high prevalence types. However given the small sample sizes, some amount of separation might be expected, and indeed the 95 percent confidence intervals for these bounds overlap. This evidence shows that patients admitted to the hospital for different reasons and with different demographic profiles are all nonetheless tested at a high rate and with similar bounds on prevalence. Indeed, we might have hypothesized ex ante that cancer patients would be particularly cautious and have the lowest prevalence, but vehicle accident patients (who may be younger and/or more likely to work in essential occupations, given the fact that they were involved in vehicle accidents) would have the highest. Instead we see both groups are on the low end. Thus overall there is no clear evidence that prevalence bounds differ substantially across types of admissions. This is perhaps reassuring for the view that pooling many distinct causes of admissions can nonetheless generate meaningful bounds on prevalence.

Assessing the hospital representativeness assumptions

The results so far show that the test monotonicity bounds on prevalence are much tighter for the non-ICLI hospitalized population than for the population as a whole. These tighter bounds are informative for general population prevalence only under additional assumptions about hospital representativeness, either a monotonicity assumption or an equal prevalence assumption. How valid are these assumptions? Assessing them directly is of course impossible because we lack data on prevalence in the population as a whole or in the hospital sample.

We have already provided one type of indirect evidence in support of our hospital representativeness assumptions. The non-ICLI and clear-cause samples generate similar bounds, and, within the clear-cause sample, there are not large differences in bounds across different causes of admission. This suggests that prevalence does not vary with the exact set of hospitalizations studied, although of course this does not prove hospitalization monotonicity or hospitalization independence are credible assumptions.

In this section, we provide two additional pieces of evidence on the hospital IV assumptions. First we show that the hospital bounds are consistent with the estimates of population prevalence from the Indiana COVID-19 Random Sample Study NirWave1,NirWave2. \footnote{Our data do not contain the test results from the Random Sample Study, so we compare our bounds to the published results.} Second, we compare the hospital sample to the general population in terms of their likelihood of prior testing (prior to the hospital data) and the test rate of their home counties. We take these to be proxies for their concern about COVID, although other interpretations are possible.

Comparison to random sample testing

A valuable benchmark for the hospital-based prevalence bounds comes from a large-scale study of SARS-CoV-2 prevalence in Indiana. The study invited a representative sample of Indiana residents (aged 12 and older) to obtain a SARS-CoV-2 test. The first wave of the study took place April 25-29, and the second wave took place June 3-7. The preliminary results are reported in NirWave1 and NirWave2. The response rate was roughly 25 percent, and no attempt was made to correct for non-random response. Nonetheless this survey appears to be the best benchmark available. We report the point estimates for prevalence (assuming random nonresponse) and their confidence intervals in the top panel of \hyperref[tab:pop_prev]{Table \ref*{tab:pop_prev}}. The first wave estimates 1.7 percent prevalence and the second 0.5 percent. \footnote{The estimates in \hyperref[tab:pop_prev]{Table \ref*{tab:pop_prev}} are slightly different from those reported by NirWave2. We report updated calculations, based on correspondence with the authors.}

We compare our prevalence bound during the same time periods in the bottom panel of the table. We limit our sample to tests of people aged 12 and older, for comparison with the population study. Using population testing we obtain very wide bounds that contain the random sample study estimates. This fact provides some support for the test monotonicity assumption. For both the non-ICLI hospitalization and clear-cause hospitalization samples, the bounds are much tighter and the point estimate from the random sample survey lies near the bottom of the identification region in both samples. The confidence interval around upper and lower bounds includes the point estimate from the random sample survey in both April and June for both the non-ICLI and clear-cause samples. Thus for both dates the prevalence point estimates are consistent with the bounds obtained from the non-COVID hospitalizations. As a comparison we also report the bounds from the ICLI-related hospitalizations, which always exclude the random sample estimates.

Comparison of prior testing and community testing

A standard way of measuring representativeness is to compare the distribution of covariates in a study population to their distribution in the target population. In our case, this approach is most convincing if we have well-measured covariates that proxy for SARS-CoV-2 infection risk. Two candidate covariates are the community SARS-CoV-2 testing rate and the prior testing rate. The idea behind these proxies is that people who come from areas with high test rates, or who have been tested in the past, may themselves have a higher current likelihood of being infected with the virus.

To operationalize these measures, we define the community testing rate for person $i$ as the fraction of people in $i$'s county who have ever been tested, as of the end of our sample period. We define the prior test rate of person $i$ as of date $t$ as the probability that $i$ was tested at least once during the week-long period $[t-15, t-9]$. We focus on this window because it is the second week prior to our hospital testing window (which runs from $t-2$ to $t+4$ for a patient admitted at $t$). We allow for a week of time to elapse between the hospitalization and the “prior" testing because it is possible that some pre-hospital testing would occur in the window $[t-8, t-3]$. When studying prior tests, we limit the sample to each person's first hospitalization after March 1, 2020, to avoid picking up the higher testing that mechanically results from the fact that people hospitalized once are more likely than the general population to have been previously hospitalized. As with our bounds, we weight the data to match the population age distribution.

\hyperref[tab:county_rates]{Table \ref*{tab:county_rates}} shows the community testing rate. The average county in Indiana has a testing rate of 25%, with an interquartile range of 22% to 28%. The average person lives in a county with a test rate of 26.7%. The average non-ICLI hospitalized patient comes from a county with a test rate of 26.6%, and the average clear-cause hospitalization patient comes from a county with a test rate of 26.5%. Among ICLI hospitalizations it is 26.7%. Our sample size is large enough that these differences are all statistically significant. Practically, however, the differences are very small. Hospitalized patients appear to come from counties that are roughly representative in terms of their testing rates. These rates are all significantly different from the population average.

\hyperref[fig:prior_rates]{Figure \ref*{fig:prior_rates}} shows the prior testing rate as a function of admission date for the non-ICLI hospitalization sample, the clear-cause hospitalization sample, and the general population (for which the prior test rate on day $t$ is defined as the fraction tested between $t-15$ and $t-9$). The rates in the hospitalization samples are initially close to the population rate (when testing is low in general), but the lines diverge. By the last week of the sample, the prior testing rate is 1-2 percentage points lower in the hospitalization samples, than in the population. Although the differences in weekly testing rates are not statistically significant, the lower prior testing rate in the hospital sample could indicate that the hospital sample is negatively selected on SARS-CoV-2 infection risk.

Conclusion

We have calculated weekly bounds on the prevalence of SARS-CoV-2 for the Indiana population as a whole and for three hospitalized populations: people hospitalized for influenza- and COVID-like illness, people hospitalized for other reasons not related to influenza or COVID, and people with clear (and clearly not COVID) causes of hospitalization. The bounds we report are valid under weak test monotonicity assumptions. The bounds for the general population are wide but narrow over time. The bounds for the hospitalized population are much tighter, because the hospitalized populations are tested at a much higher rate than the general population.

The hospitalized populations are informative for the general population only under additional, stronger assumptions. In particular, if the hospitalized population is representative of the general population in terms of SARS-CoV-2 prevalence, then both the upper and lower bounds are valid. If the hospitalized population has a higher prevalence than the general population, then only the upper bound is valid. We assess these assumptions in multiple ways. We find that the confidence intervals around our non-COVID hospitalized population bounds contained the point estimates of prevalence from random sample testing surveys conducted in late April and early June. Since the summer months, hospitalized patients also appear to have lower prior testing rates than the general population, even outside the hospital. The gap in prior testing rates in recent months is not statistically significant, but the point estimates suggest that the hospitalization monotonicity and hospitalization independence assumption could be violated, although the magnitude may not be too severe. Overall, the two hospital representativeness assumptions we consider appear to be plausible, and they have substantial identifying power. Even under the weak hospitalization monotonicity condition, the hospital data bringing down the upper bound on population prevalence by a third or more.

The basic strategy we pursue in this paper is built on the observation that testing rates differ substantially in the general population and the hospitalized population. The difference in testing rates means that assumptions that permit limited forms of extrapolation from the high testing rate hospitalized population to the general population can substantially reduce uncertainty about the overall prevalence of SARS-CoV-2, which is an important measure of population health. Although we focus on the hospitalized population, our basic strategy might also be applicable to other sub-populations that are tested at higher than usual rates, such as health care workers or university students.

We believe that the main promise of the non-COVID hospitalization population is that it can provide near-real time information about population prevalence. Only three numbers are necessary to calculate bounds from the non-COVID hospitalizations: the count of non-COVID admissions, the number of tests among this group, and the number of positive results. Although these numbers are not currently reported, many states already report related numbers, including both the number of COVID tests and the count of ICLI-related hospitalizations. Thus the infrastructure largely exists already to calculate these bounds. The results here help validate this approach for real-time surveillance.

figure[figure omitted — 846 chars of source]
figure[figure omitted — 1,182 chars of source]
figure[figure omitted — 1,291 chars of source]
figure[figure omitted — 646 chars of source]
table[table omitted — 2,331 chars of source]
table[table omitted — 1,617 chars of source]
table[table omitted — 1,516 chars of source]
table[table omitted — 1,098 chars of source]