Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,033 characters · 15 sections · 74 citation commands
Binary Classification Tests, Imperfect Standards, and Ambiguous Information
\thispagestyle{empty}
\cleardoublepage \setcounter{page}{1} \setstretch{1.3}
An important aspect of evaluating a new diagnostic test is to assess its accuracy. Intuitively, a sensible binary test should have test results highly correlated with the underlying health condition. In other words, a positive test result should be likely if and only the tested person is indeed infected or sick.\footnote{In the following, I will not differentiate between being infected and being sick.} However, establishing whether a person is truly infected is often costly or even impossible. Therefore, a new test is analyzed relative to an established test. An established test is perfect when a positive test result occurs if and only if the person is truly infected. The medical literature calls these perfect tests a “gold standard” watson-etal-2020. In these situations the joint distribution of the new test's outcomes and the underlying true health condition is the same as the joint distribution of test results from both tests. Thus, this observed joint distribution can be used to evaluate the new test's accuracy.
In practice, however, a perfect reference test does not exist. In such a case, the researcher would need the joint distribution of the health conditions and the outcomes of both tests.\footnote{Depending on the question, it might suffice to consider the joint distribution of the new test's outcomes and the underlying true health condition. For example, the analysis (ref) requires only this bivariate joint distribution. The trivariate distribution is needed to evaluate the informativeness of performing both tests as discussed in (ref).} This overall joint distribution is not observable (or maybe only if the researcher incurs a high costs for obtaining the data). This missing data problem leads to two distinct problems: ($i$) the marginal distribution of the underlying health condition is missing and ($ii$) the correlation between new test's outcome and health status is missing too.\footnote{In the introduction, I use the word correlation loosely and informal.} The latter of these problems will introduce ambiguity in the information provided by the new test.
The first of these problems, missing data about the underlying health condition, is well-known. Recently, manski-molinari-2021 use methods known from the literature on partial identification to provide bounds on prevalence---the fraction of infected people in the population.\footnote{Similar approaches were used by stoye-2020 and sacks-etal-2020.} Measuring prevalence is different from the usual inference problem because the tested population might not be representative of the overall population. Here the data are observed selectively which corresponds to a selection problem as introduced by manski-1989. Furthermore, manski-2020 illustrates how this problem carries over to evaluating accuracy of new tests in the context of COVID-19 Antibody tests maintaining the assumption of a perfect reference test.
The second problem of missing data about the correlation is different in nature and avoided when a perfect reference is available. Even if one would assume knowledge of prevalence, potentially multiple 'correlation structures' are consistent with the observed data. The reason for this multiplicity is well-known from copulas as studied in probability theory. Knowledge of prevalence provides the marginal distribution of the health condition, whereas the observed testing data provides (a bivariate) marginal distribution. In general, there are multiple (trivariate) joint distributions with these marginal distributions. Due to this multiplicity a simple, unambiguous interpretation of the new test is not possible. Without knowledge of prevalence, the problem identified before carries over and therefore exacerbates the overall multiplicity. However, as discussed in more detail later, the ambiguous information stems only from the missing data on correlation and therefore occurs whether or not the researcher has knowledge about prevalence.
In this paper, I provide a theoretic framework that combines insights from manski-molinari-2021 and stoye-2020 about selective testing with the missing correlation data due to an imperfect reference test. Within this framework, it is possible to address informativeness of both tests. First, (ref) shows that the established test's negative predictive value\footnote{A test's negative predictive values is the probability of being healthy conditional on obtaining a negative test result. Another important informativeness measure is the positive predictive value, which is the probability of being infected conditional on a positive test result. I will assume throughout that the established has a perfect predictive value in line with the application to SARS-CoV-2 testing.} is usually not given by a unique number, but it always informative nevertheless. This multiplicity arises because of problem ($i$) only. Then, I analyze the new test's informativeness for the test population only. The focus on the tested population simplifies the algebra and furthermore shuts down the ambiguity about prevalence (cf.\ problem ($i$)) and therefore allows me to study the essence of ambiguous information for the new test in separation (problem ($ii$) only). Finally, I study the implications on informativeness if both effects are present.
Studying the informativeness of tests has a long tradition in probability theory, statistics, economics, and philosophy. blackwell-1951,blackwell-1953 introduces a notion of “(more) informative” for (what is now called Blackwell) experiments.\footnote{deoliveira-2018 provides a more recent treatment.} An experiment is a mapping from states of the world to a distribution over signals. In the current setting, an experiment is a function that associates a distribution over test results to each of the possible health conditions, i.e.\ for being infected and for being healthy. In such a setting, the value of information is defined as the amount a Bayesian decision maker is willing to pay for the experiment. Since every experiment is more informative than an uninformative experiment,\footnote{An experiment is uninformative if the mapping mentioned above is a constant function?} blackwell-1951's theorem shows that the value of information is (weakly) positive for every Bayesian decision maker.\footnote{More generally, blackwell-1951 characterizes his notion of “more informative” with the requirement that every Bayesian decision maker has a higher value of information for the more informative experiment.} Ideally a diagnostic test should satisfy blackwell-1951's definition of an experiment in order to ensure that it is always informative. However, this is typically only true for the established test in my framework.
The new test fails to be a Blackwell experiment because it does not map each state to a unique distribution over test results. Rather, due to the multiplicity of joint distributions, there is a set of distributions over test results for a given health condition.\footnote{Formally, the new test can be seen a correspondence or set-valued function.} Therefore, blackwell-1951's informativeness notion does not apply to the new test. Furthermore, the value of information needs to be adjusted because a Bayesian analysis does not readily apply with sets of probabilities. Such a situation is usually referred to as a situation of “ambiguity” and the literature has identified several extensions of Bayesian decision making to the realm of ambiguity.\footnote{machina-siniscalchi-2014 provide a recent overview about this topic.}
Instead of defining the value of information for a specific decision criterion in such a situation, I adopt a very weak notion of informativeness: the diagnostic test is informative if and only if it is not a dilation. seidenfeld-wasserman-1993 introduce the notation of dilation for situations with sets of probabilities. In the current context, a dilation occurs if, no matter what test result is obtained, the set of probabilities conditional on this information contains the original set of probabilities. (ref) illustrates an example of a dilation. Here, the set of probabilities indicating the infection likelihood before the test (black set) lies within both sets after the test result (blue for a positive result and red indicating the set after a negative result). Thus, in a sense, the decision maker is worse-off after taking the test than before taking the test no matter what the test result is. For this reason, seidenfeld-wasserman-1993 call a dilation a “counterintuitive phenomenon” and gul-pesendorfer-2018 refer to it as “all news is bad news”.
My framework allows to fully characterize when a new diagnostic test is a dilation (cf.\ (ref)). Since the definition of informativeness for tests is rather weak, any reasonable test should satisfy this criterion. The characterization provides a method to verify whether the new test is informative.
The who-2020 recommends a minimum standard of accuracy for rapid Antigen tests.\footnote{For the informed reader, the who-2020 recommends a sensitivity of at least $80\%$ and a specificity of at least $97\%$. These measures will be formally introduced and defined later.} Usually a PCR test is the established test used to evaluate these Antigen tests esbin-etal-2020. (ref) illustrates hypothetical test data, which fulfill the minimal requirements of the who-2020. However, as the analysis will reveal, this test is actually a dilation and therefore not informative.\footnote{This statement depends on the accuracy of the established PCR test. A dilation occurs only if the PCR has sensitivity at the lower range identified by the literature.}
For minimum required accuracy standards, the dilation characterization provides an easy to verify sufficient condition to avoid dilation. The new test is informative (in the population of tested people) if
Besides theoretic applications of dilation, there is not much empirical evidence in the literature yet. Recently, economists started to investigate dilation and ambiguous information experimentally. The only experiment focusing on how decision makers react to dilation and relate the behavior to the value of information is conduced by shishkin-ortoleva-2020. To the best of my knowledge, the possible occurrence of dilation with diagnostic tests (or SARS-CoV-2 tests more specifically) is the first observation of this phenomenon `in the field.'\footnote{manski-2018 mentions that a dilation might occur in a different medical context, but does not address this issue further.}
Of course, researchers studying diagnostic tests are well aware of the general issues addressed here. The problem of selection leading to unobserved prevalence is known as Verification Bias, whereas the problem arising from unobserved correlation due to an imperfect reference test is descriptively named Imperfect Gold Standard Bias. zhou-etal-2014 This paper is not the first to document that either of of these problem leads to non-identified models; rather the novelty of this paper comes in the approach. Diagnostic test research seeks to avoid non-identified models by introducing additional assumptions and then address resulting biases relative to a baseline assumption. Proposed methods include simply imputing missing data or considering more sophisticated correction methods. Moreover, the two problems are often addressed separately. By contrast, my framework requires minimal assumptions and addresses both problems simultaneously.\footnote{reitsma-etal-2009 provide a flowchart as guidance for applied researchers to address several problems arising when establishing accuracy of diagnostic tests. The two problems addressed here are in two distinct branches of the flowchart.}
I consider the following situation. Let $x=1$ denote that a person is infected and $x=0$ if the person is healthy. Initially, there is binary test available where $y=1$ indicates a positive test result and $y=0$ a negative result. Finally, a new test is introduced which again can be either positive ($z=1$) or negative ($z=0$).
Let $P(x,y,z)$ denote the population distribution under consideration with $p:=P(x=1)$ denoting prevalence. However, the population distribution is not directly observable. This is, of course, almost always the case, because a researcher usually only observes a sample from the population distribution. This leads to the usual inference problem. Throughout, I will abstract away from inference altogether. Instead, the data is given for people who were tested to obtain data on the new test. For this denote tested people with $t=1$ and $t=0$ otherwise. Then, the data are given by $P(y,z|t=1)$ and I assume that $P(t=1)>0$.\footnote{Furthermore, the following logical implications of (not) being tested hold: ($i$) $t=0 \implies z=0$ and ($ii$) $ z=1 \implies t=1$. Note that $y=1$ is possible even if not tested, because the participation pool concerns only the new test.}\textsuperscript{,}\footnote{Equivalently, the data is given by sensitivity and specificity of the new test relative to the established test with the additional information about how many established or new tests had a positive result.}
Furthermore, since the established test is well-known, precise information about the sensitivity and specificity of this test is available as well. The following assumption ensures that both of these measures are well defined.
With this assumption, sensitivity and specificity for the initial test are respectively defined as:
As discussed in manski-2020, for decision making sensitivity and specificity are not the relevant measures. The relevant measures are positive predictive value (PPV) and negative predictive value (NPV). For the established test these measures can be obtained from specificity and sensitivity via Bayes' rule if prevalence $p$ and $P(y=1)$ are known:
Since the tested people are usually not representative of the overall population,\footnote{For example, supposedly infected people may be oversampled in order to get meaningful results.} even for the established test these two measures are not point-identified. manski-molinari-2021,manski-2020,stoye-2020
To simplify the analysis and in-line with the application to SARS-CoV-2 testing, I also consider the following three baseline assumptions.
(ref) implies that the established test achieves a maximum specificity and $\operatorname{PPV}_y$ of $1$.\footnote{This holds because $ P(x=0,y=0) = P(x=0,y=0) + P(x=0, y=1) = P(x=0) = 1-p$ and $P(x=1, y=1) = P(x=1, y=1) + P(x=0, y=1) = P(y=1)$.}
Additionally, I will assume test-monotonicity as in manski-molinari-2021, meaning conditional on being tested the probability of being infected is greater than if not being tested.\footnote{This might not be true, if there is voluntary enrollment into the testing pool. However, for establishing the accuracy of new tests this assumptions seems to be applicable often. See (ref).}
Lastly, I assume that the established test's sensitivity does depend on the underlying health status $x$, but not on whether the person is in the testing pool $t=1$.\footnote{Recall that the testing pool is obtained for the new test. This assumption might be violated, if, for example, the medical staff performing the established test for the testing pool is extra careful. In this case, the established test might be more sensitive for the testing pool.}
To reduce cumbersome lengthy notation in the following, I will use this simplified notation henceforth:
To avoid trivial cases, assume that $\gamma, \zeta, \tau >0$. Note that $\tau$ has a slightly different interpretation as in manski-molinari-2021 or stoye-2020. Here, $\tau=1$ means the data $P(y,z|t=1)$ is perfectly representative of the overall population. In particular, such a parameter value implies no oversampling of infected participants.\footnote{See (ref) for why such an assumption might be problematic.} In particular, even if the participation pool is small (as is often the case), this does not mean that $\tau$ should be close to zero.\footnote{A small participation pool might worsen the statistical inference problem: suppose the participation pool is perfectly representative but small. In this case $\tau=1$, but inference usually relies on some sort of central limit theorem which would not be appropriate in this scenario. However, recall that I abstract away from inference problems as mentioned above.}
With this notation, we have $P(z=1)=\tau \zeta$ since the new test is positive only if the person was tested. Furthermore, (ref) combined with (ref) gives $P(x=1|t=1) = \nicefrac{\gamma}{\sigma}$. Then, the Law of Total Probability together with (ref) provides sharp bounds\footnote{A bound for a given set is called sharp if the bound itself is a member of this set.} on prevalence $p \in \left[ \tau \nicefrac{\gamma}{\sigma}, \nicefrac{\gamma}{\sigma} \right]=:\left[\underline{\chi}, \overline{\chi}\right]$ because
In turn, bounds on the established test's overall positivity rate are implied by sensitivity $\sigma$ and (ref): $P(y=1) = p\sigma \in \left[\tau \gamma, \gamma\right]$.
Since we consider the non-trivial case of $p\in(0,1)$, consistency of the data with the maintained assumptions requires the established test's sensitivity to be sufficiency high , i.e.\ $\gamma < \sigma \leq 1$. In turn, the assumptions imply $P(y=1) \in (0,1)$.
(ref) implies a perfect positive predictive value for the established test ($\operatorname{PPV}_y=1$). However, the negative predictive value is only partially identified and (ref) provides sharp bounds.
With this in hand, the established test's informativeness can be analyzed. (ref) summarizes the prevalence before and after observing a test result from the established test, which are the relevant measures for defining informativeness (cf.\ (ref)). Formally, the established test is a dilation if every possible prevalence $p \in [\underline \chi, \overline \chi]$ is a possible value of both $P(x=1|y=1)=\operatorname{PPV}_y$ and $P(x=1|y=0)=1-\operatorname{NPV}_y$. Obviously, a positive test result gives perfect knowledge due to the maintained assumptions. Thus, the established test cannot be a dilation. On the other hand, a negative result lowers the lower and upper bound of prevalence conditional on a negative result because $\tau \gamma \leq \gamma < \sigma$.\footnote{Of course, this has to hold since the Law of Total Probability holds pointwise.} Furthermore, the interval width for any test result shrinks the set of possible values for prevalence conditional on either test result.\footnote{It is obvious for a positive test result. For a negative result, note that the width strictly increases if and only if $1-\sigma > (1- \gamma)(1-\tau\gamma)$, which is equivalent to $\tau(1+\gamma) > \frac{\sigma}{\gamma}+1 \geq 2$ leading to a contradiction.} In this sense, the established test is not just informative (i.e. not a dilation), but also strictly shrinks the size of the set of possible prevalence values after a negative test result.
It is well known that knowledge of prevalence is needed in order to apply Bayes' rule to obtain NPV. Since in most applications prevalence is not known, a common practice is to assume a given prevalence level. For example the United States Food and Drug Administration fda-2020 assumes a prevalence of $5\%$ to calculate PPV and NPV. If such an assumption ($p=\chi$) is added to the maintained assumptions, then $P(y=1)=\chi \sigma$ and furthermore $P(y=1|t=0) = \frac{ \chi \sigma - \gamma \tau}{1-\tau}$.\footnote{Alternatively, one could drop the assumption that $P(t=1)=\tau$ is known exactly. In this case (and allowing $P(y=1|t=0) \in [0, \gamma]$ as in the general case) the assumed prevalence bounds $\tau$. Calculations show that $\tau \in \left[0, \frac{\chi \sigma}{\gamma}\right]$. Since the lower bound is always $\tau_{\min}=0$, we do not find this case very interesting.} This additional assumption allows to exactly pin down the established test's NPV as $\frac{1-\chi}{1-\chi\sigma}$ and therefore $P(x=1|y=0) = \chi \frac{1-\sigma}{1-\chi \sigma}$. Thus, this additional assumption not only assumes away the ambiguity about prevalence, but also illustrates that the established test does not provide ambiguous information itself.\footnote{Technically, the established test is an experiment \'{a} la blackwell-1951, where sensitivity and specificity can be seen as functions mapping (health) states to distributions over signals (i.e. test results). As mentioned in the introduction, this implies that the established test's value of information is (weakly) positive under these assumptions.} The apparent ambiguity reflected in the non-trivial interval for values of prevalence after a negative test result (cf.\ (ref)) or NPV (cf.\ (ref)) is only a manifestation of the ambiguity about prevalence, but it is not due to the test itself.
Next, the new test's informativeness is analyzed. First, I will discuss informativeness only based on the tested population. For this subpopulation the prevalence is given by $\overline \chi=\nicefrac{\gamma}{\sigma}$ and therefore the ambiguity about prevalence is muted. (ref) extends the analysis then to the informativeness of the new test for the overall population. For the test population, the relevant measures are again positive-predictive value (PPV) and negative-predicative value (NPV), but now they are also conditional on being tested:
To obtain these measures, the distribution $P(x,z| t=1)$ is needed. For a fixed $\tau$, I use a result from joe-1997 that provides the set of all possible joint distributions $P(x,y,z)$ compatible with the data $P(x,y|t=1)$ (cf.\ (ref)). Setting $\tau=1$ in this construction gives the possible distributions $P(x,y,z|t=1)$. Finally, $P(x,z| t=1)$ is obtained by marginalization.
To simplify the algebraic expressions it will be useful to differentiate between four cases defined in (ref). Fixing the established test's sensitivity $\sigma$, the test data $P(x,y|t=1)$ immediately reveals the case the test belongs to. (ref) illustrates this for three real tests considered later (StQ, BiN, CT) and three hypothetical tests (including the dilation test from (ref)). When $\sigma \rightarrow 1$, then all but the informative case (I) cease to be relevant. For SARS-CoV-2 detecting Antigen test the WHO recommends a minimum specificity close to one. Tests close to the (top-right) frontier in (ref) satisfy this criterion.\footnote{CT, Uni, and Anti do not satisfy the WHO minimum requirement of a $97\%$ minimum specificity.} Thus, for most applications, either the confirmatory (if $\sigma < 1$) or the informative case (if $\sigma \approx 1$) will be the relevant ones.
In contrast to the established test, the new test's PPV could be less than one and is, in general, only set-identified. The reason for set-identification is that $P(x,z| t=1)$ is not directly observed. As explained above, there are multiple distributions $P(x,z| t=1)$ consistent with the data and each distribution leads to a potentially different PPV. (ref) establishes the sharp identified set for values of PPV.
To avoid partially identified predictive values, these measures for the new tests are often reported as if the reference test is perfect. In this case, the data $P(y,z)$ alone delivers a unique predictive value:
We saw before that the established test always achieves a maximal PPV of one and therefore provides a lot of information in case it delivers a positive result. How informative is a positive result of the new test? To answer this question, note that since we condition on being tested, there is no prior uncertainty as the prevalence in the testing pool is given by $\overline{\chi}$. Even without this prior ambiguity there remains ambiguity in the test result. For example, in the confirmatory case (C) the interval's width of possible values for $\operatorname{PPV}_z$ is $1-P(y=1|z=1, t=1)$, which is usually small---but non-zero---in applications. Thus, the information obtained from a new test is ambiguous at least after a positive test result.\footnote{This observation alone implies the test is not an experiment \`{a} la blackwell-1951. Similarly to deriving PPV, it is possible to derive the new test's sensitivity. This will be a set in general too. Thus, there is a correspondence from (health) states to distributions over signals (test results).}
In contrast to the established test as discussed in (ref), the ambiguity arising from the new test allows for the occurrence of dilation. In the current setting a dilation occurs if $\overline{\chi} = P(x=1|t=1)$ is contained in the intersection of the two sets with possible values for $P(x=1|y=i, t=1)$ for each test result $i \in \{0,1\}$. Is it possible that after a positive test result the set of possible values for $P(x=1|y=1, t=1)$ contain $\overline{\chi}=P(x=1|t=1)$? (ref) provides a full characterization. The corresponding case after a negative test result will be discussed afterwards.
The inequality of (ref) becomes non-trivial in case the testing data does not correspond to an independent distribution, which will be the case for most applications. In these cases, a dilation cannot occur when the non-trivial inequality of (ref) is violated.
It remains to analyze the information contained in a negative test result. (ref) establishes sharp bounds for the negative predictive value of the new test. This is the relevant measure for analyzing how informative a negative test result is.
The uninformative (U) and contradictory (X) case seem problematic in light of (ref). In both cases, the lower bound is zero and also the width of the interval is rather large. This is another indication that any reasonable test should not fall in either of these two cases. However, even for the other cases---and like for PPV---the NPV is generally only set-identified. Therefore a negative test result also produces ambiguous information.
Avoiding this ambiguity can be achieved with a perfect reference test. (ref) verifies that if the reference test is perfect, (ref) reduces to the expression used in many applications and can be calculated directly from the data $P(y,z)$.
If there is no perfect reference test available, the negative new test's result leads to ambiguity. Similar to the case of a positive test result, this ambiguity allows for the occurrence of dilation. Using (ref), (ref) provides a characterization of when the set of possible values of $P(x=1|z=0,t=1)=1-\operatorname{NPV}_z$ contains the prior information $P(x=1|t=1) = \overline{\chi}$.
(ref) combined with (ref) provides an exact characterization for when the new test is a dilation. In fact, as the conditions are the same a dilation occurs if and only if
When evaluating a new test's accuracy it is important to make sure the data violates (ref). Otherwise, the test is uninformative in an extreme sense. In typical applications, the data often satisfies $P(y=1, z=0|t=1)\leq\gamma (1-\zeta)$.\footnote{Even data that regards a test as inadequate as in cassaniti-etal-2020 satisfies this inequality. I thank Filip Obradovic for making me aware of this report.} test In these cases, a dilation can only occur if $\sigma \leq \frac{\gamma \zeta}{P(y=1, z=1|t=1)}$.
The who-2020 recommends minimum quality requirements using only information directly provided by the data $P(y,z|t=1)$. In light of this analysis, an evaluation should also take $\sigma$, the established test's sensitivity, into account and with this also make sure that the test is not a dilation. $\sigma \leq \frac{\gamma \zeta}{P(y=1, z=1|t=1)}$ combined with a given minimum standard provides an easy-to-verify sufficient condition to avoid dilation.
For this, let $\underline{\Sigma}$ be a minimum (apparent) sensitivity threshold below which a test is deemed not reliable and denote the new test's apparent sensitivity with $\Sigma = P(z=1|y=1, t=1)$, so that a test is reliable if $\Sigma > \underline{\Sigma}$.\footnote{Usually, the minimum requirements include also a threshold for specificity, but this does not matter here.} Then, the application relevant case from (ref) to avoid a dilation can be expressed as $\sigma > \zeta /\Sigma$ or equivalently as $\Sigma > \zeta/\sigma$. If $\underline{\Sigma} \geq \zeta/\sigma$, then any test meeting the minimum requirement cannot be a dilation. Thus, it suffices to make sure the new test's yield is not too high:\footnote{It is worth recalling that this is a sufficient condition when, additionally, $\frac{\gamma \zeta}{P(y=1, z=1|t=1)} \leq \frac{\gamma (1-\zeta)}{P(y=1, z=0|t=1)}$ and in many applications this inequality becomes irrelevant for (ref) because the right-hand side is greater than one.}
If the established test is highly specific, i.e. $\sigma \approx 1$, then (ref) is satisfied unless the new test's yield is extremely high.
For SARS-CoV-2 Antigen tests the WHO recommendation is $\underline{\Sigma} = 0.8$ and if the PCR test is not highly specific then (ref) might be violated. For example, the dilation test of (ref) has $\zeta=0.49$ and if the PCR has sensitivity of $\sigma=0.6$ then not only (ref) is violated but the test is a dilation. More specifically, (ref) can be used to find the exact threshold sensitivity $\sigma^*$ below which a given test turns into a dilation. For the dilation test this value is $\sigma^* = 60.79\%$.
Whereas the new test produces ambiguous information, the established test is always informative. Therefore, practitioners might want to perform an additional established test depending on whether a person obtains a negative or positive result from the new test. For example, if an Antigen test is used to detect SARS-CoV-2 and the result is positive, a common practice is verifying the result by means of a PCR test. Since PCR tests are the reference test for evaluating the accuracy of Antigen tests, the current framework can be used to shed light on how informative this additional test is.
(ref) once more reveals that tests in the category (U) and (X) should be avoided. Even if the two test results match and are both negative, the possibility of zero (negative) predictive value cannot be ruled out. (ref) also makes clear the naming convention of the cases defined in (ref). A confirmatory test (C) provides accurate information when both test produce a negative result, but is completely uninformative if and only if the new test has a positive result. An informative test (I), however, always provides some information in the sense of producing not completely trivial bounds. A contradictory test (X) provides information, but leans against the result of the established test. The uninformative test (U) provides no information at all even when both tests agree on a negative result. Of course, performing the additional test is always informative in the sense of not being a dilation. The established test does not produce false-negatives and therefore a positive result from the established test is always a perfect predictor of being infected regardless of the new test.
(ref) bounds the established test's NPV for the overall population, not only for the tested population. The new test, on the other hand, was analyzed for the testing pool only so far. The full characterization in (ref) allows to extend the analysis of the new test to make an evaluation for the overall population. Since this involves more cumbersome notation, I only illustrate the resulting bounds for the $\operatorname{NPV} = P(x=0|z=0)$. The analysis of PPV would proceed in a similar matter.
If instead of predictive values the interest lies in the new test's sensitivity or specificity in the whole population another complication arises. Conditional on the testing pool, both of these measures can be derived as in (ref). For example, for sensitivity one could use the proof of (ref) and divide by $P(x=1|t=1)=\overline{\chi}$ instead of $P(z=1|t=1)=\zeta$. The bounds for sensitivity are again determined by considering the extremes of (ref) and (ref). For the unconditional sensitivity, however, the the numerator and the denominator are both set identified because $P(x=1) \in [\underline{\chi}, \overline{\chi}]$. Therefore, the lower bound might not be attained at either of the extreme distributions. This makes solving for a closed-form expression for sensitivity intractable. Nonetheless, the bounds can easily be obtained computationally by considering a fixed $\Gamma := P(y=1) \in [\tau \gamma, \gamma]$ with corresponding $p =\Gamma/\sigma$. For this $\Gamma$, sharp bounds of sensitivity, say $[L_\Gamma, H_\Gamma]$, can be obtained by using (ref) and (ref). To find the overall bounds for sensitivity, two (non-linear) optimization problems across all values of $\Gamma$ need to be performed to give $[\min_\Gamma L_\Gamma, \max_\Gamma H_\Gamma]$.
In this section, the theoretic framework will be illustrated with several applications. First, I analyze the (hypothetical) dilation test presented in the introduction. Then, I examine two real SARS-CoV-2 detecting tests. Finally, I show that CT-scanning procedures to detect COVID-19 are prone to being dilations.
As argued before the hypothetical test data in (ref) corresponds to a dilation. Suppose the test data is derived for an Antigen test to detect SARS-CoV-2 and the reference test is a PCR test.\footnote{Recall that for SARS-CoV-2 detection a PCR test is the established test used to evaluate other tests. esbin-etal-2020}. The test satisfies the who-2020's (who-2020) minimum requirements with apparent sensitivity ($\Sigma=80.6\%$) and specificity ($97.1\%$) above the specified thresholds of $80\%$ and $97\%$, respectively.\footnote{These numbers are calculated as if the reference test is perfect. This is similar to (ref) and (ref).} For such a setting the current framework is applicable. Especially, (ref) seems to be warranted because a PCR test is highly specific. However, it is known that a PCR test might lack high sensitivity. alcoba-florez-etal-2020 report sensitivity for several PCR tests with point estimates ranging from $\sigma=60.2\%$ to $\sigma=97.9\%$.\footnote{alcoba-florez-etal-2020 differentiate values based on the targeted gene. The range reported here is across all genes and tests.} All of the $95\%$ confidence intervals exclude perfect sensitivity, $\sigma=1$.
Using the results from (ref), (ref) summarizes some key statistics for the dilation test. When the PCR sensitivity is close to one, the new (hypothetical) test produces relative accurate measurements with PPV close to one and NPV above $75\%$. However, if the PCR test lacks high sensitivity then we cannot be sure of the dilation test's quality. In the worst-case for PCR sensitivity ($\sigma=0.6$), the new test is indeed a dilation: Before a test result was obtained the prevalence (in the testing pool) is $98.8\%$, after obtaining either dilation test's result the possible probability of being infected is at least the interval $[97.7\%, 100\%]$. In fact, potentially even more puzzling is that the lowest value after a negative test is strictly higher than after a positive result. Using (ref), $\sigma^*=60.8\%$ represents the cutoff PCR sensitivity below which a dilation occurs.
Next, consider the Standard Q (StQ) COVID-19 Rapid Antigen Test of SD Biosensor/Roche for detection of SARS-CoV-2 as analyzed by kaiser-etal-2020. They use results of PCR tests as comparison (see (ref)). The testing data is summarized in (ref).
When $\sigma=1$, then StQ's PPV and NPV are obtained with (ref) and (ref) which yields $99.42\%$ and $94.13\%$, respectively. These are the reported values of kaiser-etal-2020. However, as explained above PCR are not perfectly sensitive.\footnote{ kaiser-etal-2020 use PCR tests targeting $E$ genes, which tend to have higher sensitivity in the analysis of alcoba-florez-etal-2020. The lowest reported sensitivity for a PCR test targeting $E$ genes is $65.33\%$.} Thus, to evaluate the StQ test the current framework is applicable.
Focusing first on the testing pool only, (ref) summarizes PPV and NPV for different values of PCR sensitivity ($\sigma$) using (ref) and (ref). Even if the PCR test lacks high sensitivity, StQ has a close to perfect positive predicative value ($\operatorname{PPV}_z \approx 1$). However, the values for NPV drop significantly as $\sigma$ decreases. In the worst case, a negative StQ result becomes close to a fair coin flip. However, the test is very informative overall as can be seen by the low dilation threshold $\sigma^*=36.3\%$.
kaiser-etal-2020 state “study participants were representative of the usual population seeking testing in our center (main testing center in Geneva). The majority were presenting with symptoms compatible with a SARS-CoV2 infection and a minority were asymptomatic but with a known positive contact or were asymptomatic healthcare workers.” The current framework allows to use the obtained testing data to evaluate StQ's quality for the overall population (of Geneva) as analyzed in (ref). Furthermore, this explanation supports (ref).
(ref) shows bounds on prevalence using the baseline analysis in (ref). For low values of $\tau$, i.e. the testing pool was highly non-representative of the overall population, the width of the intervals is rather wide. However, even the lowest number is close to $2\%$ indicating a thorough spread of the virus in Geneva at the time of testing.\footnote{Note that this is a one time analysis. It does not answer the question of how many people were cumulatively infected by SARS-CoV-2 up to the time of testing.} When testing becomes representative ($\tau \rightarrow 1$) the prevalence converges to the prevalence in the testing pool.
At this time, if a Genevese obtains a negative PCR result, what is the probability of her being infected? If testing is not competently representative, a unique number cannot be given. However, (ref) provides sharp bounds for this case and the results are shown in (ref).
How do these PCR results compare to results from StQ? (ref) provides the numbers using (ref). The lower bounds are significantly lower than for the PCR test. This makes the width of the interval also significantly wider. The widening is a reflection of the combination of the two missing data problems inherit in the testing procedure without a perfect reference test: ($i$) unknown overall prevalence (which also affects PCR's NPV) and ($ii$) missing correlation data (which does not affect the PCR's NPV).
The BiaxNOW (BiN) Covid-19 Ag Home Test of Abbott is one of the first rapid Antigen tests for use at home which is able to detect the SARS-CoV-2 virus that was emergency approved by the fda-2020a. BiN's clinical performance for approval by the FDA was conducted with a PCR test as a reference. The data are shown in (ref) and (ref) shows the implied accuracy measures.
Relative to StQ, BiN has significantly lower PPV and also rules out perfect PPV for lower values of PCR specificity. On the other hand, NPV is uniformly greater for BiN compared to StQ. Even for the worst-case PCR sensitivity, BiN's possible NPV values are reasonably high. Furthermore, the dilation threshold is extremely low at $\sigma^*=26.7\%$.
The fda-2020a also provides additional data about BiN results by including the cycle threshold obtained by the PCR test.\footnote{publichealthengland-2020 explains: “Cycle threshold (Ct) is a semi-quantitative value that can broadly categorise the concentration of viral genetic material in a patient sample following testing by RT PCR as low, medium or high –- that is, it tells us approximately how much viral genetic material is in the sample. A low Ct indicates a high concentration of viral genetic material, which is typically associated with high risk of infectivity. A high Ct indicates a low concentration of viral genetic material which is typically associated with a lower risk of infectivity.”} (ref) shows this data.
This is additional data a PCR test produces, which can be used to refine bounds on predictive values of the Antigen test. However, an extension of the current setting is needed, because such additional information is not accounted for in the current setting with binary tests. (ref) discusses a possible extension of the current setting to allow for this additional data.
ai-etal-2020 and gietema-etal-2020 propose using chest CT scans for early identifying COVID-19 in patients. In gietema-etal-2020's study, all COVID-19 symptomatic patients of a single Dutch emergency department have a chest CT scan and a PCR test for detecting SARS-CoV-2. Their study design exactly fits the framework of the current paper: ($i$) non-representative sampling of the testing pool and ($ii$) missing correlation data due to use of an imperfect reference test (with perfect specificity). The testing data are reproduced in (ref).
Compared to the previously studies Antigen tests, the data for CT scans seems less aligned with the PCR test results. This is an indication that such a CT test is less informative: the dilation threshold of $\sigma^*=63.35\%$ is higher than for the Antigen tests. Thus, this testing procedure is completely uninformative if $\sigma=60\%$---the lowest sensitivity of a PCR test reported by alcoba-florez-etal-2020. In this case, (ref) shows (sharp bounds on) population prevalence, PCR NPVs, and CT scan NPVs.
The PCR's (assumed) low sensitivity means that its NPV might be quite low, but it as at least close to $50\%$, irrespective of $\tau$. On the other hand, the CT scan has both a sizable interval and a low lower bound of possible NPVs. Since $\sigma=0.6$ is below the dilation threshold, the CT scan is completely uninformative for the tested population (equivalently for $\tau=1$). Furthermore, this remains true for the overall population if $\tau \geq 1/2$ as shown in (ref). For example, for a non-COVID-indicative CT scan the set of possible infection probabilities $P(x=1|z=0)$ increases relative to the prior information $p$.\footnote{Recall that $P(x=1|z=0)=1-\operatorname{NPV}_z$.}
Even more striking is the data of ai-etal-2020 shown in (ref).\footnote{I thank Filip Obradovic for providing this reference.} ai-etal-2020 also use Chest CT scans to test for COVID-19. Their data is obtained in Wuhan, China and like the study of gietema-etal-2020 a PCR test is used as a reference. The data reveals a very low apparent specificity but a high apparent sensitivity of $\Sigma = 96.51\%$. (ref) indicates that a high yield of the new test, $\zeta$, might be problematic. Here, this yield is very high with $\zeta=87.57\%$. The problem becomes even more apparent by looking at the dilation threshold, which is high with a value of $\sigma^* = 90.74$. This implies that even if the PCR test is quite sensitive, the CT scan is completely uninformative for the tested people in Wuhan.
(ref) might be too strong for some applications. Although, this assumption simplifies the algebraic expression sometimes significantly, it is not a crucial assumption conceptually. The crucial characterization of joint distributions in (ref) can easily be extended to allow for false-positives of the established test.
Evaluations of a new test sometimes have more data available than just $P(y,z|t=1)$. For example, blood samples from before the existence of a virus can serve as true-negative samples. On the other hand, specific blood samples could be analyzed with more sophisticated (and usually much more expensive) methods than just using an established test as reference. These methods would lead to samples with true positives (or at least with very high probability).\footnote{For example, olbrich-etal-2020 combine these methods to evaluate SARS-CoV-2 antibody tests.} Either of these methods would be provide additional data and therefore would also reduce the missing data problem. In general, this supplementary knowledge leads to narrower bounds, but unless these extra methods are performed for the whole tested population, the missing correlation issues remains. Of course, these methods cannot be applied for the untested population. Therefore the the missing data on prevalence cannot be avoided with these extraneous data.
The current framework only allows for binary outcomes for both tests and also for the underlying health state. This seems to be the most common situation studied in the literature on diagnostic testing. zhou-etal-2014 Often tests provide ternary results (with the additional result of `inconclusive' or `invalid'), or allow for even more detailed information, like the Cycle Threshold Count of a PCR test as mentioned in (ref). In such situations, the theoretic analysis does not provide the appropriate machinery. However, the crucial application to characterize the set of all joint distribution is a result in copula theory joe-1997, which does not rely on any dimension being binary. Indeed, the result even works for continuous outcomes on each dimension.
Similarly, one could use other results in joe-1997 to characterize the set of possible joint distributions if multiple tests are conducted simultaneously as studied in zhou-etal-2014. In this case, and like in the characterization of (ref), the testing data are higher-dimensional marginal distribution of an overall joint distribution with an additional dimension (the health state). When considering such an extension, a caution has to be taken because sometimes sharp bounds on the set of possible higher-dimensional distributions may not be known.
Since testing has a potentially big impact on the economy, an accurate description of the available testing technology is crucial. From the microeconomic perspective, the testing technology affects how test should be optimally allocated (see ely-etal-2020, lipnowski-ravid-2020) and also how much people engage in social distancing as studied by acemoglu-etal-2020. But also the macroeconomy is highly affected by testing strategies and an optimal choice might reduce the economic costs of pandemics considerably. alvarez-etal-2020,eichenbaum-etal-2020 Although, these studies establish the importance of testing and also address varying testing technologies, all of them assume that a test corresponds to an experiment \`{a} la blackwell-1951 and therefore is always informative (sometimes the assumption is even that the test itself provides perfect information).
This paper demonstrates that the assumption of unambiguous information in test results is only applicable if a perfect reference is available when evaluating new tests. In particular, new Antigen test for detection of SARS-CoV-2 are evaluated relative to an imperfect PCR test and therefore---as shown in this paper---these Antigen test produce ambiguous information. An optimal testing procedure should take this ambiguity into account. Similarly, practitioner guides (like galeotti-etal-2020,watson-etal-2020) might want to consider addressing uncertainty in test results in more detail.