Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,474 characters · 19 sections · 0 citation commands
Inference for ROC Curves Based on Estimated Predictive Indices
Binary prediction or classification is a fundamental problem in statistics and machine learning with applications in many scientific disciplines including economics. The predictive ability of statistical models used for binary classification is frequently evaluated through the receiver operating characteristic (ROC) curve, designed to summarize the tradeoffs between the probability of a true positive prediction (vertical axis) and a false positive prediction (horizontal axis) as one combines a predictive index with a varying classification threshold. Though its origins are in the signal detection and medical diagnostics literature, in recent years ROC analysis has become increasingly common in financial and economic applications as well (e.g., Anjali and Bossaerts 2014; Bazzi et al.\ 2021; Bonfim et al.\ 2021; Berge and Jorda 2011; Kleinberg et al.\ 2018; Lahiri and Wang 2013; Lahiri and Yang 2018; McCracken et al.\ 2021; Schularik and Taylor 2012 and many others).
While there is a large literature on the statistical properties of empirical ROC curves, the standard distributional theory assumes that the signal or predicitive index used for classification is either directly observed---it is “raw data”---or that it is a fixed function of raw data. However, if the signal itself is generated from an underlying regression model with estimated coefficients, conducting in-sample inference about ROC curves based on the traditional theory can be highly misleading. For instance, Demler et al.\ (2012) point out that the standard DeLong et al.\ (1988) test for comparing AUCs for different (but potentially correlated) signals can lead to flawed inference if the signals come from nested models with estimated coefficients.\footnote{AUC stands for “area under [the ROC] curve”. It is an overall performance measure for binary prediction models. AUC=1/2 corresponds to no predictive power.} Similarly, Lieli and Hsu (2019) demonstrate that the asymptotic normality results in Bamber (1975) are inappropriate for testing AUC=1/2 for models with estimated parameters.
The central contribution of this paper is the development of a general functional limit theory for the empirical ROC curve that takes the pre-estimation effect into account. Regarding the ROC curve as a random function defined over the [0,1] interval, we provide a uniform influence function representation theorem, and show that the difference between the empirical and population ROC curves converges weakly to a mean zero Gaussian process with a given covariance structure at the parametric rate. These results constitute a non-trivial extension of the functional limit result in Hsieh and Turnbull (1996), who work under the assumption that the observations available on the predictive index are i.i.d.\ conditional on the outcome. (If the predictive index depends on coefficients estimated in-sample, this assumption is no longer valid, as the parameter estimates depend on all data points.) Our results not only allow for the construction of a uniform confidence band for the ROC curve but also facilitate model selection through handling virtually any comparison between two correlated ROC curves (e.g., testing dominance or partial dominance; testing the difference between AUCs or partial AUCs, etc.) In terms implementation, we propose two methods to simulate the limiting distribution of the empirical ROC curve, one of which is the weighted bootstrap by Ma and Kosorok (2005). Although some type of bootstrap procedure would be a natural way to approach the pre-estimation problem in practice even without a theory, our results provide rigorous justification and guidance for doing this.
A second contribution of the theory developed in this paper is that it provides insight into what determines the impact of the first stage estimation error on the asymptotic distribution of the ROC curve. The derivatives of the true and false positive prediction rates with respect to the first stage model parameters play a central role --- if these gradients vanish at the pseudo-true parameter values, then so does the estimation effect. Nevertheless, the gradients also depend on the classification threshold and will not generally be negligible along the entire ROC curve. For associated functionals, it is the gradient of the functional that drives the estimation effect. Some functionals, e.g., the area under the curve, have the property that they are maximal when the predictive index is given by $p(X)=P(Y=1|X)$, the conditional probability of a positive outcome $Y$ given the covariates $X$. For such functionals the estimation effect is negligible when the first stage model is correctly specified for $p(X)$ and the first stage estimator converges at the root-n rate. Nevertheless, the first stage estimation error will generally affect the asymptotic distribution under misspecification; thus, our results facilitate robust inference.
There are additional technical contributions that are more subtle. In employing standard empirical process techniques to derive our results, we make most of the fact that the population and sample ROC curves are invariant to monotone increasing transformations of the predictive index. This observation allows us to leverage powerful assumptions that may seem restrictive at first glance. In particular, we use the assumption that the density $f_0$ of the predictive index conditional on $Y=0$ is bounded away from zero to derive various uniform approximations. The problem is that even if the individual predictors have densities bounded away from zero (already a big if), the predictive index may not share this property, as it often involves a linear combination of the predictors.\footnote{Think of the sum of two independent uniform[0,1] random variables or the central limit theorem for that matter.} Nevertheless, one can always find a strictly increasing transformation of the predictive index, say, the probability integral transform, so that the post-transform $f_0$ will be greater than some $\epsilon>0$ across the whole support. What matters for the asymptotic theory is the properties of the likelihood ratio $f_1/f_0$, which is invariant to monotone increasing transformations (here $f_1$ is the conditional density of the predictive index given $Y=1$). In particular, uniform inference is possible only for parts of the ROC curve that are generated by thresholds falling into some interval $[c_L,c_U]$ over which $f_1/f_0$ is bounded and bounded away from zero.\footnote{That such an interval exists is a weak assumption; that it coincides with the support of $f_0$ is a much stronger one.} But given the properties of the likelihood ratio, one is free to assume the theoretically most convenient scenario about the individual density $f_0$ that is achievable through monotone transformations, even if this transformation is not implemented or even identified.
We must also point out some technical limitations of the paper. First, we do not allow for serial dependence in the data, precluding time series applications such as the evaluation of recession forecasting models. Nevertheless, our proofs rely mostly on high level conditions; specifically, the asymptotically linear representation of the first stage estimator, the stochastic equicontinuity of the empirical process defined by the (pseudo-true) predictive index, and the uniform continuity of some derivatives. Given the availability of these conditions for stationary, weakly dependent time series, we conjecture that our representation results should generalize to this setting with relatively straightforward modifications. However, simulating the asymptotic distribution of the limiting process would require more complex procedures and we do not pursue this extension here. Second, the predictive model evaluated at the pseudo-true parameter values must have strictly positive variance. This is not an innocent assumption in that it rules out a completely uninformative predictive model. For example, our results are not suitable for testing the hypothesis that AUC=1/2; see Lieli and Hsu (2019) for some specialized results in this very non-standard case. More generally, in using our results to compare two ROC curves, the difference of the two influence functions evaluated at the pseudo-true parameter values must have strictly positive variance as well. This condition can be violated when the first stage models are nested, and we are currently working on some test procedures that are applicable in this scenario as well. Finally, we only consider parametric estimators of $p(X)$ as the first stage model; nonparametric estimators that converge slower than the root-n rate are ruled out.
One might discount the practical relevance of our theoretical results discussed above based on the fact that the first stage estimation problem can be avoided by conducting out-of-sample evaluation. If the ROC curve is constructed over a test sample that is independent of the training sample used to estimate the first stage model, then the asymptotic validity of the standard inference procedures is restored. We acknowledge this point but offer two responses. First, out-of-sample evaluation is costly: it leads to power loss in model comparisons and the potential dependence of the results on the particular split(s) used. In fact, one could argue that out-of-sample evaluation is a necessity forced on practitioners by the fact that it is often very difficult to characterize analytically or in a practically useful way how goodness-of-fit measures behave over the training sample so that one can compensate for overfitting. In this case we do provide such a result. Second, apart from dealing with the pre-estimation problem, our results provide a unified framework for conducting uniform inference, and comparing ROC curves estimated over the same sample in virtually any way.
This work has ties to several strands of the statistics and econometrics literature. We have already cited a number of classic works on the statistical properties of the empirical ROC curve that maintain the assumption of a directly observed signal (Bamber 1975, DeLong et al.\ 1988, Hsieh and Turnbull 1996). It is the last of these papers that is closest to ours; however, the pre-estimation effect is obviously missing from their framework and they actually do not exploit their functional limit result for inference apart from re-deriving the asymptotic normality result of Bamber (1975) for the empirical AUC. Instead, they focus on estimating the ROC curve under an additional “binormal” assumption, i.e., when the signal has a normal distribution conditional on both outcomes.
Pre-estimation problems have a long history in the literature; for example, Pagan (1984) studied the distributional consequences of including “generated regressors” into a regression model. As mentioned above, Demler et al.\ (2012) pointed out the relevance of the pre-estimation effect in the context of ROC analysis. More generally, our work is related to papers dealing with two-step estimators where the first step involves estimating some nuisance parameter whose sampling variation potentially affects the otherwise well-understood second stage. Abadie and Imbens (2016) is a relatively recent example in the context of matching estimators.
The application of our results in testing for dominance relations and AUC differences across ROC curves has similarities to stochastic dominance tests; see, e.g., Barrett and Donald (2003), Linton, Maasoumi and Whang (2005), Linton, Song and Whang (2010) and Donald and Hsu (2016). Some papers, such as Linton, Maasoumi and Whang (2005) and Linton, Song and Whang (2010), even allow for generated variables in this context. Finally, the paper speaks indirectly to the forecasting literature on the relative merits of in-sample vs.\ out-of-sample model evaluation (e.g., Inoue and Kilian 2004, Clark and McCracken 2012). The connection lies in the fact that we extend the scope of in-sample evaluation methods in binary prediction.
The rest of the paper is organized as follows. Section (ref) sets up the prediction framework and introduces the ROC curve along with some of its basic properties. Section (ref) discusses the estimation effect in detail and presents pointwise (fixed-cutoff) asymptotic results. The functional limit theory is contained in Section (ref). In Section (ref) we show how to use the abstract results for conducting inference about the ROC curve; we discuss dominance testing, AUC comparisons, etc. Section (ref) presents Monte Carlo results highlighting the impact of first stage estimation on the distribution of the ROC curve. Section (ref) concludes.
Let $Y\in\{0,1\}$ be a Bernoulli random variable representing some outcome of interest and $X$ be a $k\times 1$ vector of covariates (predictors). We consider point forecasts (classifications) of $Y$ that are constructed by combining a scalar predictive index $G(X)$ with a suitable cutoff (threshold) $c$. More specifically, the prediction rule for $Y$ is given by
where $1[.]$ is the indicator function. The role of the function $G: \mathbb{R}^k\rightarrow\mathbb{R}$ is to aggregate the information that the predictors contain about $Y$ while the choice of $c$ governs the use of this information. With $G$ and $c$ unrestricted, ((ref)) represents a very general class of prediction rules.
In many binary prediction problems there is also a loss function $\ell(\hat y, y)$ that describes the cost of predicting $Y=\hat y$ when the realized outcome is $Y=y$. If the decision maker's objective is to minimize expected loss conditional on $X$, the optimal choice of $G$ and $c$ is determined as follows. Given the observed value of $X$, the optimal point forecast of $Y$ solves
where $p(X)=\mathbb{P}(Y=1|X)$ is the conditional probability of $Y=1$ given $X$. Adopting the normalization $\ell(0,0)=\ell(1,1)=0$ and assuming $\ell(0,1)>0$ and $\ell(1,0)>0$, it is straightforward to verify that the optimal prediction rule is given by
where $c_\ell=\ell(1,0)/[\ell(0,1)+\ell(1,0)]$.\footnote{In case $p(X)=c_\ell$, which is often a zero probability event, the decision maker is indifferent between predicting 0 or 1. The formula stated above arbitrarily specifies $\hat Y=0$ in this case.}
Equation ((ref)) reveals that the optimal predictive index is $p(X)$ regardless of $\ell$ while the optimal choice of $c$ is fully determined by $\ell$ (specifically, by the relative cost of a false alarm versus a miss). This simple observation motivates a two-step empirical strategy in binary prediction.\footnote{See Elliott and Lieli (2013) for an alternative approach where the decision rule ((ref)) is estimated in a single step based on a specific loss function.} First, one models and estimates $p(X)$ using data on $(Y,X)$; a common approach is to specify a parametric model $G(X,\beta)$ for $p(X)$ and to estimate it by maximum likelihood (logistic regression is a leading example). In the second step a point forecast is obtained by combining the estimated conditional probability $G(X,\hat\beta)$ with a suitable cutoff, which depends on the forecaster's or forecaster user's preferences. Thus, there is a separation between the construction of the predictive index, representing the objective information available to the forecaster, and the use of that information, governed by the loss function.
We will now define the population receiver operating characteristic (ROC) curve. Let $G(X,\beta)$ be a predictive index with a fixed value of the parameter $\beta$.\footnote{The following definitions do not depend on the parametric structure and generalize immediately to any predictive index $G(X)$. We work with a parametric specification in anticipation of studying the pre-estimation step.} Combined with a cutoff $c$, the resulting prediction rule produces true positive predictions and false positive predictions (false alarms) with the following probabilities:
where TP and FP stand for the rate of “true positive” and “false positive” predictions, respectively. As the cutoff $c$ varies, both quantities change in the same direction; in general, TP can be increased only at the cost of increasing FP as well. The ROC curve traces out all attainable (FP, TP) pairs in the $[0,1]\times[0,1]$ unit square, i.e., it is the locus \[ \Big\{\big(FP(c,\beta), TP(c,\beta)\big): c\in\mathbb{R}\Big\}. \]
Intuitively, the ROC curve is a way of summarizing the information content of $G(X,\beta)$ about the outcome $Y$ without committing to any particular cutoff, i.e., loss function. The use of such a forecast evaluation tool is particularly appropriate in situations in which there many potential forecast users with diverse loss functions; see Lieli and Nieto-Barthaburu (2010). It is also clear from the definition that the ROC curve is invariant to strictly monotone transformations of $G(X,\beta)$.
The ROC curve based on the true conditional probability function $p(X)$ possesses some optimality properties. To state these in a parametric modeling framework, we introduce the following correct specification assumption.
We state the following result.
\paragraph{Remarks:}
Throughout the paper we maintain the assumption that the available data consists of a random sample. More formally:
Given the sample and a fixed value of $\beta$, the empirical ROC curve is defined as the locus $\big\{(\widehat{FP}(c,\beta), \widehat{TP}(c,\beta)): c\in\mathbb{R}\big\}\subset [0,1]\times[0,1]$, where
and $n_1=\sum_{i=1}^n Y_i$, $n_0=n-n_1$.
The simplest type of inference about an ROC curve involves constructing (joint) confidence intervals for $TP(c,\beta)$ and $FP(c,\beta)$ for one threshold $c$ at a time and for fixed values of the coefficient vector $\beta$. In this case one can use the CLT to arrive at the normal approximations:
where TP is a shorthand for $TP(c,\beta)$ (and similarly for FP). These results immediately provide asymptotic confidence intervals for TP and FP, and a joint confidence rectangle is also easy to construct due to the independence of $\widehat{TP}(c,\beta)$ and $\widehat{FP}(c,\beta)$; see Pepe (2003), Section 2.2.2 for details.\footnote{The two statistics are independent because they are computed from two disjoint sets of observations; namely, the $Y=1$ and $Y=0$ subsamples.}
Furthermore, for fixed $\beta$ one can apply the asymptotic normality results in Bamber (1975) to conduct inference about the AUC, and the DeLong et.\ al.\ (1988) test for comparing the areas under ROC curves based on different (but non-random) values of $\beta$. The nonparametric functional limit result in Hsieh and Turnbull (1996) also applies.
We will now develop a comprehensive theory of in-sample inference about individual points on the ROC curve taking the pre-estimation effect into account. We present both analytical results and results based on the weighted bootstrap. We start by describing the setup and stating some technical conditions.
The sample $\{(Y_i,X_i)\}_{i=1}^n$ now plays a dual role. First, it is used to construct an estimated parameter vector $\hat\beta=\hat\beta_n$. Typically, $\hat\beta$ consists of an intercept and slope coefficients from some type of regression of $Y$ on $X$ (e.g., linear, logit or probit). Second, the same sample is used to compute the predictive index values $G(X_i,\hat\beta)$, $i=1,\ldots, n$, and to construct the empirical ROC curve as described in Section (ref). We impose the following high level condition on $\hat\beta$.
Assumption (ref) states that $\hat\beta$ is an $M$-estimator with an asymptotically linear representation, implying that $\hat\beta$ is asymptotically normally distributed. The stated conditions do not require the first stage model to be correctly specified for $p(X)$; $\beta^*$ simply stands for the probability limit of $\hat\beta$, i.e., the pseudo-true value of the parameter vector. We nevertheless assume that $\hat\beta$ is consistent for $\beta^\circ$ under correct specification (Assumption (ref)). The reason for allowing for misspecification is twofold. First, the first stage predictive model is often simply an approximation of $p(X)$, e.g., a linear regression. Second, as we will see, the estimation effect can depend on whether or not the model is correctly specified.
The stochastic equicontinuity requirement in Assumption (ref) limits the complexity of the model $G(X,\beta)$ and plays an important role in handling the estimation effect. It holds, for example, if $G(X,\beta)=X'\beta$ or $G(X,\beta)=G(X'\beta)$ with $G$ bounded (see the definition of a type I class in Andrews 1994). Apart from a small degree of added generality, we state stochastic equicontinuity as a high level condition to make it more transparent what is required for our results.
The final assumption states the differentiability of $TP$ and $FP$ with respect to the components of $\beta$. Let $\nabla_\beta$ denote the corresponding gradient operator and $B^*(r)$ the open ball with radius $r>0$ centered on $\beta^*$.
As we will shortly see, the first stage estimation of $\beta$ affects the asymptotic distribution of the ROC curve through the derivatives presented in Assumption (ref).
Let $c$ be a given value of the cutoff; we want to conduct inference about the corresponding point $(FP(c,\beta^*), TP(c,\beta^*))$ on the limiting ROC curve. To isolate the effect of the first stage estimator $\hat\beta$ on the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$, we can write
where the second equality is due to the fact that the process $\sqrt{n}(\widehat{TP}-TP)$ is stochastically equicontinuous (Assumption (ref)), implying \[ \sqrt{n}[\widehat{TP}(c,\hat\beta)-TP(c,\hat\beta)]-\sqrt{n}[\widehat{TP}(c,\beta^*)-TP(c,\beta^*)]=o_p(1), \] given that $\hat\beta\rightarrow_p\beta^*$. As $\beta^*$ is fixed and $Var[G(X,\beta^*)]>0$, the first term in equation ((ref)) has the asymptotic distribution given by ((ref)) and the second term represents the effect of estimating $\beta^*$. Does this term have a non-negligible effect on the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$, and if yes, how do we characterize it?
To address these questions, we can use Assumption (ref) to expand the second term in ((ref)) around $\beta^*$ to obtain
Equation ((ref)) shows that the estimation effect is negligible whenever $\nabla_\beta TP(c,\beta^*)=0$. However, as Proposition (ref)(ii) shows, this condition does not generally hold even if $G(X,\beta)$ is correctly specified (i.e., $\beta^*=\beta^\circ$), because $TP(c,\beta^\circ)$ solves a constrained (rather than unconstrained) optimization problem.\footnote{The first order conditions are $\nabla_\beta TP(c,\beta^\circ)=\lambda\nabla_\beta FP(c,\beta^\circ)$ for some scalar $\lambda$ and $FP(c,\beta)=F^\circ_c$. The Lagrange multiplier $\lambda$ is generally non-zero, at least when $TP(c,\beta^\circ)<1$.} Under misspecification Proposition (ref)(ii) does not apply, but of course there is still no general reason for $\nabla_\beta TP(c,\beta^*)$ to vanish. Therefore, in either case the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$ and $\widehat{FP}(c,\hat\beta)$ will generally differ from that stated under ((ref)) and ((ref)) because $\sqrt{n}(\hat\beta-\beta^*)=O_p(1)$.\footnote{Proposition (ref)(i) implies that under correct specification one can conduct inference about the linear combination $(1-c)\piTP-c(1-\pi)FP$ without the need to consider the pre-estimation effect. This is because the true value of $\beta$ maximizes this linear combination and hence the corresponding gradient driving the estimation effect vanishes. More generally, the estimation effect is negligible for a functional of the ROC curve if (i) the ROC curve based on $p(X)$ maximizes that functional and (ii) Assumption (ref) holds.}
To describe the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$ in more detail, we can further expand the decomposition in ((ref)) by substituting in the asymptotically linear (influence function) representation of the two terms. Using the definition of $\widehat{TP}(c,\beta^*)$ and Assumption (ref), it is straightforward to verify that
where the definition of $\psi_{TP}$ is enclosed by the braces on the previous line.
Of course, $\widehat{FP}(c,\hat\beta)$ has a corresponding asymptotically linear representation with influence function \[ \psi_{FP}(Y_i,X_i,c,\beta^*)=\frac{1-Y_i}{1-\pi}\big[1(G(X_i,\beta^*)>c)-FP(c,\beta^*)\big]+\nabla_\betaFP(c,\beta^*)\psi_\beta(Y_i,X_i,\beta^*). \] Stacking the influence functions as $$\psi(Y_i,X_i,c,\beta^*)=[\psi_{TP}(Y_i,X_i,c,\beta^*),\psi_{FP}(Y_i,X_i,c,\beta^*)]'$$ and applying the multivariate CLT gives the asymptotic joint distribution of an individual point $(\widehat{TP}(c,\hat\beta),\widehat{FP}(c,\hat\beta))$ on the sample ROC curve.
\paragraph{Remarks}
We supplement Proposition (ref) by some results that reveal the structure of $\nabla_\beta TP$ and $\nabla_\beta FP$ and facilitate their estimation. Let $\partial_{j}$ denote the partial derivative operator with respect to the $j$th component of $\beta$.
Finally, we specialize Propositions (ref) and (ref) by imposing a logit first stage.
\paragraph{Remarks:}
Here we provide an alternative method for making pointwise inference about the ROC curve by utilizing the weighted bootstrap for M-estimators proposed by Ma and Kosorok (2005). The main advantage of this approach is that it sidesteps the estimation of the gradient vectors $\nabla_\beta TP(c,\beta^*)$ and $\nabla_\beta FP(c,\beta^*)$. Furthermore, the method is similar to the simulation-based procedure that we propose for functional inference in Section (ref).
The weighted bootstrap employs a sequence of (pseudo) random variables as multipliers to simulate the sampling variation of an estimator.
We first define the weighted bootstrap version of the first stage estimator of $\beta$:
Given $\hat{\beta}^w$, the weighted bootstrap estimators of $TP(c,\beta)$ and $FP(c,\beta)$ are defined as
Assumption (ref) ensures that the weighted bootstrap is valid for the first stage estimator, i.e., conditional on the data, $\sqrt{n}(\hat{\beta}^w-\hat{\beta})$ has the same limiting distribution as $\sqrt{n}(\hat\beta-\beta^*)$ unconditionally.
Furthermore, by Theorem 2 of Ma and Kosorok (2005), the validity of the weighted bootstrap for $\widehat{TP}^w(c,\hat\beta^w)$ follows from showing that (i) $\widehat{TP}$, $\widehat{FP}$, $\widehat{TP}^w$ and $\widehat{FP}^w$ can be represented as M-estimators and (ii) that these estimators are $\sqrt{n}$-consistent and asymptotically linear.
Item (i) is verified by noting that
and similarly for $\widehat{FP}$ and $\widehat{FP}^w$. As for item (ii), Proposition (ref) establishes the asymptotically linear representation of $(\widehat{TP}(c,\hat\beta), \widehat{FP}(c,\hat\beta))$; essentially the same argument also yields
Thus, we obtain the following result.
\paragraph{Remarks}
To derive uniform results, we first express the ROC curve explicitly as a function over the interval [0,1]. Let the inverse of the decreasing function $c\mapsto FP(c,\beta)$ be defined as \[ FP^{-1}_{\beta}(t)=\inf\{c: FP(c,\beta)\le t\},\; t\in [0,1]. \] The more compact notation on the l.h.s.\ emphasizes that the inverse is taken with respect to the cutoff $c$ for a fixed value of $\beta$. Thus, is $FP^{-1}_{\beta}(t)$ as the “first” (smallest) cutoff value at which the false positive rate is equal to $t$ or falls below $t$.\footnote{Of course, if $FP(c,\beta)$ is strictly decreasing and continuous in $c$, then $FP^{-1}_{\beta}(t)$ is the unique solution to the equation $FP(c,\beta)=t$.} Because $1-FP(c,\beta)$ is the c.d.f.\ of the conditional distribution of $G(X,\beta)$ given $Y=0$, an equivalent interpretation of $FP^{-1}_{\beta}(t)$ is that it is the $(1-t)$-quantile of this distribution.
We can now represent the ROC curve as a function that returns the true positive rate associated with given false positive rate $t$:
For a given parameter value $\beta$, the sample ROC curve is defined by replacing $TP(\cdot,\beta)$ and $FP^{-1}_\beta(\cdot)$ by sample analogs: $\widehat R(t,\beta)=\widehat{TP}\big(\widehat{FP}^{-1}_{\beta}(t),\beta\big),\;t\in[0,1]$.
Our goal is to characterize the statistical behavior of the random function $t\mapsto \hat R(t,\hat\beta)$ over the interval $[0,1]$. This requires some additional assumptions.
Assumption (ref) merits careful discussion. An immediate practical implication of part (i) is that the limiting model $G(X,\beta^*)$ must depend on at least one continuous predictor in a nontrivial way. For instance, if the model is based on a linear index, this rules out $X$ being completely independent of $Y$; see Remark 1 after Proposition (ref). Part (iii) implies that $supp(f_0^*)$ and $supp(f_1^*)$ overlap, ensuring that the classification problem is nontrivial. Nevertheless, the overlap does not need to be complete; we allow for applications in which extreme values of the index are associated exclusively with one of the two outcomes.
From a technical standpoint, the main purpose of Assumption (ref) is to facilitate uniform inference by controlling the behavior of the likelihood ratio $f_1^*/f_0^*$. In particular, $f_1^*/f_0^*$ is required to be bounded and bounded away from zero on an interval $[c_{0,L},c_{0,U}]$. Our uniform influence function representation result for $\hat R(t,\hat\beta)$ holds only for quantiles $t$ satisfying $FP^{-1}_{\beta^*}(t)\in [c_{0,L},c_{0,U}]$ or, equivalently, for $t\in [FP(c_{0,U},\beta^*), FP(c_{0,L},\beta^*)]$. While this representation depends on $f_1^*$ and $f_0^*$ only through $f_1^*/f_0^*$, the derivation of the result relies on the additional condition that $f_0^*$ is bounded away from zero (Assumption (ref)(i)). This may seem overly restrictive at first glance---for example, if $G(X,\beta^*)=X'\beta^*$, then predictors with unbounded support are ruled out. Furthermore, it is easy to see that even if all components of $X$ have densities bounded away from zero, their linear combinations will generally not share this property.\footnote{For example, consider the sum of two independent uniform [-0.5,0.5] random variables. The resulting density is $(1-|x|)1_{[-1,1]}(x)$, which tends to zero as $x$ approaches $-1$ or $1$. } However, one can always find a monotone increasing transformation $\Phi(\cdot)$ such that the density of $\Phi[G(X,\beta^*)]$ conditional on $Y=0$ is bounded away from zero, e.g., one can use the probability integral transform to arrive at a uniform[0,1] density. At the same time, such a transformation leaves the ROC curve as well as the range of the likelihood ratio $f_1^*/f_0^*$ unchanged. Thus, the last part of Assumption (ref)(i) is simply a theoretical normalization that does not need to be imposed on the data in practice (see Figure 1 for an illustration).
Of course, Assumption (ref) allows for scenarios in which $[FP(c_{0,U},\beta^*), FP(c_{0,L},\beta^*)]=[0,1]$, i.e., uniform inference is possible along the entire ROC curve. This is the case, for example, if the “propensity score” function $P(Y=1|X=x)$ takes values from an interval $[\delta, 1-\delta]$ for some $0<\delta<1/2$, which implies $supp(f_0^*)=supp(f_1^*)$ and that (iii) holds with $[c_{0,L},c_{0,U}]=[a_0,b_0]$.\footnote{To see this, let $f_x(x)$ denote the density function of $X$. Note that
It follows that
which is bounded below by $ \delta^2/(1-\delta)^2$ and bounded above by $ (1-\delta)^2/\delta^2$. } More generally, Assumption (ref)(iii) allows $f^*_1(c)/f^*_0(c)$ to reach zero or explode for cutoffs $c$ outside the range $[c_{0,L}, c_{1,L}]$. For example, the likelihood ratio vanishes as $c$ approaches $a_0$ from above whenever $a_0<a_1$. In this case the lowest index values imply $Y=0$ and the ROC curve reaches the top of the unit square for some $FP$ rate below unity. Similarly, $f^*_1(c)/f^*_0(c)$ may become unbounded as $c$ approaches $b_0$ from below. This can easily happen when $b_0<b_1$, i.e., the largest index values are associated exclusively with the $Y=1$ outcome. In this case the ROC curve has a positive vertical intercept at $FP=0$. Again, see Figure 1 for an example.
The next assumption is a strengthening of Assumption (ref). These stricter conditions on the gradient vectors $\nabla_\beta TP(c,\beta)$ and $\nabla_\beta FP(c,\beta)$ also play a key role in establishing a uniform influence function representation for the sample ROC curve. Recall that $B^*(r)$ denote the open ball with radius $r>0$ centered on $\beta^*$.
Letting $c^*_t={FP}_{\beta^*}^{-1}(t)\in[a_0,b_0]$ and $\hat c_t= \widehat{FP}_{\hat\beta}^{-1}(t)\in[a_0,b_0]$, we start from a decomposition of $\sqrt{n}[\widehat{TP}(\hat c_t,\hat\beta)-TP(c^*_t,\beta^*)]$ similar to ((ref)). There are two added layers of difficulty. First, functional results require uniform approximations to these terms as $t$ varies over the $[0,1]$ interval. Second, instead of being fixed, the cutoff is now estimated for any given value of $t$. The sampling variation in $\hat c_t$ contributes another non-trivial term to the asymptotic distribution.
We express the centered and scaled empirical ROC curve as the sum of three terms:
The first term in equation ((ref)) can be expanded similarly to the second equality in ((ref)):
where $\sup_{t\in[0,1]}|R_{1n}(t)|=o_p(1)$. The uniform convergence of the remainder term is a consequence of the stochastic equicontinuity of the process $(c,\beta)\mapsto \sqrt{n}(\widehat{TP}(c,\beta)-TP(c,\beta))$, stated directly in Assumption (ref), coupled with the fact that $\hat\beta\rightarrow_p\beta^*$ (Assumption (ref)) and $\sup_{t\in[0,1]}|\hat c_t-c^*_t|\rightarrow_p 0$ (Lemma (ref) in Appendix A). This last result makes use of Assumption (ref)(i), which requires that the density $f_0^*$ be bounded away from zero on its compact support.
The second term in equation ((ref)) is due to the estimation of $\beta$ and is again handled by a standard mean value expansion:
where $\sup_{t\in[0,1]}|R_{2n}(t)|=o_p(1)$. The uniformity of the approximation is ensured by Assumption (ref), which implies that $\nabla_\betaTP(c,\beta^*)$ is uniformly continuous.
Finally, the third term in ((ref)) arises because of the need to estimate the cutoff value associated with a given false positive rate $t$; it therefore does not arise in the fixed-cutoff setting. Starting with a mean value expansion of $TP(\hat c_t,\beta^*)$ around $c^*_t$, one can write
The remainder term $R_{3n}(t)$ converges in probability to zero uniformly over the interval \[ \big\{t: c^*_t\in[c_{0,L}, c_{0,U}]\big\}=\big[FP(c_{0,U}, \beta^*), FP(c_{0,L},\beta^*)\big], \] where $c_{0,L}$ and $c_{0,U}$ are specified in Assumption (ref)(iii). The asymptotic distribution of the process $t\mapsto \sqrt{n}[\widehat{FP}_{\hat\beta}^{-1}(t)- FP^{-1}_{\beta^*}(t)]$ can be analyzed in two steps: First, we establish an asymptotically linear representation for the “base process” $c\mapsto \sqrt{n}[\widehat{FP}(c,\hat\beta)-FP(c,\beta^*)]$ that holds uniformly in $c$ (and implies a mean zero Gaussian limit process). Second, we apply the functional delta method under the inverse functional $\phi(F)=F^{-1}$ to characterize the contribution of the term ((ref)) to the asymptotic distribution of the empirical ROC curve.
Lemma (ref) summarizes and completes the development of the approximations presented in equations ((ref)), ((ref)) and ((ref)).
\paragraph{Remarks}
Combining equations ((ref)) through ((ref)) with the influence function representations of $\sqrt{n}[\widehat{TP}(c,\beta^*)-TP(c,\beta^*)]$ and $\sqrt{n}(\hat\beta-\beta^*)$ yields the following proposition, which is the central result of the paper.
\paragraph{Remarks}
In order to employ Proposition (ref) for statistical inference, we need a method to approximate $\Psi_{h_{R}}(t)$, the distributional limit of the process $\sqrt{n}(\widehat{R}_n(t,\beta^*)-{R}(t,\beta^*))$. To this end, offer two methods: the weighted bootstrap as in Ma and Korosok (2005) and the multiplier bootstrap as in Hsu (2016).
We first present the discussion of the weighted bootstrap. Define $\widehat{TP}^w(c,\hat{\beta}^w)$ and $\widehat{FP}^w(c,\hat{\beta}^w)$ precisely as in Section (ref) and let $\hat c^w_t= \big(\widehat{FP}^w_{\hat{\beta}^w}\big)^{-1}(t)$. We can then construct the weighted ROC curve and its estimated limit process as $\widehat{R}^w_n(t,\hat\beta^w)=\widehat{TP}^w(\hat{c}^w_t,\hat{\beta}^w)$ and $\widehat{\Psi}^w_{R,n}(t)=\sqrt{n}(\widehat{R}^w_n(t,\hat\beta^w)- \widehat{R}_n(t,\hat\beta))$.
Under the conditions of Proposition (ref), one can apply the arguments in Theorem 2 of Ma and Kosorok (2005) to show that $\widehat{\Psi}^w_{R,n}(t)$ also approximates the distribution of $\Psi_{h_{R,2}}(t)$ in the sense of Proposition (ref). That is, conditional on the sample path of the data, $\widehat{\Psi}^w_{R,n}(t)\Rightarrow \Psi_{h_{R,2}}(t)$ with probability approaching one.
We now turn to the discussion of the multiplier bootstrap method that is based on the conditional multiplier central limit theorem (see, e.g., van der Vaart and Wellner 1996, Section 2.9). The method requires consistent estimation of the components of the influence function $\psi_R$, uniformly in $t$. However, this estimation needs to be performed only once, over the original data set, given that the method does not rely on successive resampling and reestimation.
Let $\widehat{\psi}_\beta(y,x, \hat{\beta})$ denote the estimated influence function of $\hat\beta$, where we replace any unknown parameters or functions within $\psi_\beta$ with consistent estimators (note that this function does not depend on $t$). We make the general assumption that the asymptotic variance-covariance matrix of $\hat\beta$ is consistently estimable using $\widehat{\psi}_\beta$ and the sample analog principle:
Let $\mathcal{C}=[c_{0,L},c_{0,U}]$. We further assume that there exist uniformly consistent estimators for $\nabla_\beta FP(c,\beta^*)$, $\nabla_\beta TP(c,\beta^*)$ and $f_{1}^*(c)/f_{0}^*(c)$ on $\mathcal{C}$. Here we state the existence of these estimators as a high level assumption and provide concrete implementations and additional assumptions in Appendix C.
We now present the multiplier bootstrap. Let $U_1,\ldots, U_n$ be i.i.d.\ random variables independent of the data with moments $E[U]=0$, $E[U^2]=1$, and $E|U|^{2+\delta_u}<\infty$ for some $\delta_u>0$. For $t\in[0,1]$, we define the simulated stochastic process $\widehat{\Psi}^u_{R,n}(t)$ as
where
The next result shows that the distribution of the simulated process $\widehat{\Psi}^u_{R,n}(t)$ approximates that of the true limiting process $\Psi_{h_{R,2}}(t)$ in large samples.
The weighted bootstrap has the advantage that it does not require explicit estimation of this function, but it is computationally somewhat more costly. On the other hand, the proof of weighted bootstrap is less involved because the multiplier method relies heavily on Assumption (ref), i.e., the availability of uniformly consistent estimators for the components of $\psi_R$. To obtain estimators satisfy Assumption (ref) additional assumptions are needed and in Appendix C, we provide estimators and additional assumptions so that Assumption (ref) can be satisfied and we can apply multiplier method.
In this section, we provide some examples that we can apply the results in Section (ref).
Let $\widehat{\sigma}^2_t$ denote a uniform consistent estimator for ${\sigma}^2_t$, the asymptotic variance of $\sqrt{n}(\widehat{R}_n(t,\hat\beta)-{R}(t,\beta^*))$ for $t\in T$. Later we will provide two estimators based on weighted bootstrap method and analytic results. Let $\widehat{\sigma}_{t,\epsilon}=\max\{\widehat{\sigma}_t,\epsilon\}$ in which $\epsilon>0$ is a fixed and small number. We are interested in a standardized version of confidence bands and by truncating $\widehat{\sigma}_t$ by $\epsilon$, we can make sure that we will not divide something close to zero when $t$ is close to 0 or 1.
For a nominal significance level $\alpha$ and for $\tau_\ell,\tau_u\in T$ with $\tau_\ell\leq \tau_u$, let $\widehat{C}^\text{1-sided}_\alpha$ and $\widehat{C}^\text{2-sided}_\alpha$ respectively denote the one- and two-sided critical values that satisfy
Here, $\widehat{C}^\text{1-sided}_\alpha$ and $\widehat{C}^\text{2-sided}_\alpha$ are, respectively, the $(1-\alpha)$th quantile of $\sup_{t\in[\tau_\ell,\tau_u]}\widehat{\Psi}^w_{R,n}(t) \big/\widehat{\sigma}_{t,\epsilon}$ and $(1-\alpha)$th quantile of $\sup_{t\in[\tau_\ell,\tau_u]}\big|\widehat{\Psi}^w_{R,n}(t) \big/\widehat{\sigma}_{t,\epsilon}\big|$. Note that one can replace $\widehat{\Psi}^w_{R,n}(t) $ with $\widehat{\Psi}^u_{R,n}(t) $ to construct $\widehat{C}^\text{1-sided}_\alpha$ and $\widehat{C}^\text{2-sided}_\alpha$ as well.
Once the critical values are constructed, we can also obtain one- and two-sided uniform confidence bands for ${R}(t,\beta^*)$ over $[\tau_\ell,\tau_u]$. Specifically, the one-sided $(1-\alpha)$ uniform confidence band is given by
and the two-sided $(1-\alpha)$ uniform confidence band is
{\bf Implementation of Uniform Confidence Bands}
We now provide a step-by-step implementation for constructing uniform confidence bands.
{\bf Uniformly consistent estimators for ${\sigma}^2_t$}\\ We consider two estimators here. First estimator is based on weighted bootstrap that is similar to Remark 2 after Proposition (ref). Let $\widehat{\Psi}^w_{R,n}(t)$ denote the ROC estimate from the $w$th bootstrap cycle, $w=1,\ldots, \mathcal{W}$. Then one can estimate ${\sigma}^2_t$ by
We have that conditional on sample path with probability approaching one,
where $\stackrel{p}{\rightarrow}_w$ denotes probability limit under the law of the $W_i$'s. It follows that uniformly over $t\in[t_\ell,t_u]$,
The second estimator is based on analytic results. Recall that $\widehat{\psi}_R(Y_i,X_i,t,\hat\beta)$ is the estimated influence function for $\widehat{R}_n(t,\hat\beta)$ used in the multiplier bootstrap method. A uniformly consistent estimator for $ {\sigma}^2_t$ is given by
and this is shown in the proof of (ref).
For two predictive index models $G_1(X,\beta_1)$ and $G_2(X,\beta_2)$, we may want to test whether $G_1$ has strictly better predictive power than $G_2$ in the sense that the ROC curve associated with $G_1$ dominates the ROC curve associated with $G_2$. What domination means is that for any given false positive rate $G_1$ delivers a higher true positive rate, i.e., the ROC curve for $G_1$ always lies above the ROC curve for $G_2$. Any decision maker, regardless of their loss function and their optimal cutoff, would then prefer model $G_1$ over $G_2$.
Let $R_1(t,\beta_1^*)$ and $R_2(t,\beta_2^*)$ denote the ROC curves associated with $G_1$ and $G_2$, respectively. The hypotheses that $R_1(t,\beta_1^*)$ dominates $R_2(t,\beta_2^*)$ can be formally stated as
Our test for ROC dominance is similar to the test for first order stochastic dominance in Barrett and Donald (2003) and Donald and Hsu (2016) except that we need to consider the estimation effect of $\hat{\beta}$ as in Linton, Massoumi and Whang (2005) and Linton, Song and Whang (2010).
Let $\widehat{R}_{j,n}(t,\hat{\beta}_j)$ be the estimators for $R_j(t,\beta_j^*)$ for $j=1,2$. Define $\psi_{j,R}(Y_i,X_i,t,\beta_1^*)$ for $j=1$ and 2 as above. Let $\widehat{\sigma}^2_{RD}(t)$ denote a uniform consistent estimator for ${\sigma}^2_{RD}(t)$, the asymptotic variance of $\sqrt{n}(\widehat{R}_{2,n}(t,\hat{\beta}_2)-\widehat{R}_{1,n}(t,\hat{\beta}_1)-{R}_{2}(t,{\beta}^*_2)-{R}_{1}(t,{\beta}^*_1))$. Let $\widehat{\sigma}_{RD,\epsilon}(t)=\max\{\widehat{\sigma}_{RD}(t),\epsilon\}$ in which $\epsilon>0$ is a fixed and small number. Uniform consistent estimator $\widehat{\sigma}^2_{RD}(t)$ can be obtained similar to $\widehat{\sigma}^2_t$ in Section (ref), so we omit the details.
We define the test statistic as $\widehat{S}_n=\sqrt{n}\sup_{t\in[0,1]}(\widehat{R}_{2,n}(t,\hat{\beta}_2)-\widehat{R}_{1,n}(t,\hat{\beta}_1))/\widehat{\sigma}_{RD,\epsilon}(t)$. Define the weighted bootstrap process $\Psi^w_{RD,n}(t)$ as $\widehat{\Psi}^w_{R,n}(t)=\sqrt{n}\big(\widehat{R}^w_{2,n}(t,\hat\beta^w_2)-\widehat{R}^w_{1,n}(t,\hat\beta^w_1) -(\widehat{R}_{2,n}(t,\hat{\beta}_2)-\widehat{R}_{1,n}(t,\hat{\beta}_1))\big) $ and define the multiplier bootstrap process $\Psi^u_{RD,n}(t)$ as
Under the least favorable configuration, we define the weighted bootstrap critical value as
with significance level $\alpha$. The decision rule is
Then one can use $\widehat{\Psi}^u_{RD,n}(t)$ to construct critical value $\hat{c}_n$ as well.
Similar to the stochastic dominance test literature, we can show that under the null hypothesis the asymptotic size of a test with decision rule defined in ((ref)) is less than or equal to $\alpha$. That is, we can control the asymptotic size of our ROC dominance test well. Also, under the fixed alternative, we have the test statistic converging to positive infinity and the critical value converging to a finite number, so the test is consistent. Our test is based on least favorable configuration, so it is conservative in that the asymptotic size is strictly smaller than $\alpha$ unless $R_2(t,\beta_2^*)= R_1(t,\beta_1^*)$ for all $t\in[0,1]$. One can improve the power of our test by using the recentering method in Hansen (2005), Donald and Hsu (201) which is similar to the generalized moment selection method in Andrews and Soreas (2010), and Andrews and Shi (2013), and the contact set approach in Linton, Song and Whang (2010). In this paper, we do not adopt this approach but the extension is straightforward.
Recall that AUC is defined as the integral of ROC curve from 0 to one. Following Section (ref), let $R_1(t,\beta_1^*)$ and $R_2(t,\beta_2^*)$ denote the ROC curves for two predictive index models $G_1(X,\beta_1)$ and $G_2(X,\beta_2)$. Let $AUC_j=\int_{0}^1 R_j(t,\beta_j^*) dt$ and its estimator be $\widehat{AUC}_j=\int_{0}^1 \widehat{R}_{j,n}(t,\hat{\beta}_j)dt $ for $j=1$ and 2. Then it is true that
and
To make inference, one can use a weighted bootstrap method to approximate the limiting distribution of $N[0, \mathcal{V}_{a}]$ or one can estimate $ \mathcal{V}_{a}$ analytically by
For brevity, we omit the details here.
{\bf Remark}\\ Suppose that $G(X,\beta)$ is a correct specification for the propensity score function, $P(Y=1|X)$, in that there exists $\beta^*$ such that $G(X,\beta^*)=P(Y=1|X)$ a.s. Then the estimation effect of $\beta^*$ on the distribution of AUC is negligible. It is then true that when two predictive predictive index models, $G_1(X,\beta_1)$ and $G_2(X,\beta_2)$, are both correctly specified for $P(Y=1|X)$, we have $\mathcal{V}_{a}=0$, i.e., the limiting distribution of $\sqrt{n}(\widehat{AUC}_2-\widehat{AUC}_1)$ is degenerate.
We now present a small Monte Carlo simulation to illustrate the theoretical discussion of the estimation effect and the pointwise asymptotic results in Section (ref). The data generating process (DGP) is specified as follows. Let $X=(X_1, X_2, X_3)'$ be a $3\times 1$ vector of predictors and $\tilde X=(1,X')'$. The components of $X$ are either independent $N(0,1)$ or unif$[-.5,1.5]$ variables. The outcomes $Y$ are generated according to the conditional probability function \[ p(X)=G(X,\beta^\circ)=G(\tilde X'\beta^\circ)\text{ with } \beta^\circ=(0, 0.5, 0.25, 1)'\text{ and }G\in\{\text{logit},\text{cauchit}\}. \] In the majority of the exercises we use a logistic link in the DGP (so that the logit first stage is correctly specified), but we also conduct some simulations with a cauchit link (so that the logit first stage is mildly misspecified).
Given a sample of observations and a cutoff $c$, we construct nominally 90% confidence intervals for TP$(c,\beta^\circ)$ and TP$(c,\beta^\circ)-FP(c,\beta^\circ)$ in three different ways: (i) using the true predictive index $p(X)=G(\tilde X'\beta^\circ)$ with the conventional limit distributions ((ref)) and ((ref)); (ii) using the estimated predictive index $\Lambda(\tilde X'\hat\beta)$ with the conventional limit distributions (so that the estimation effect is ignored); (iii) using the estimated predictive index $\Lambda(\tilde X'\hat\beta)$ with the corrected limit distribution ((ref)). We simulate 10,000 samples and compute the actual coverage probability of these intervals.
Table (ref) reports the results from this exercise for $c\in\{0.2, 0.33, 0.5, 0.67, 0.8\}$ and $n\in\{200, 500, 5000\}$. The first message is that failing to account for the pre-estimation effect can cause substantial distortions in the coverage probability of the conventional CIs. In panels (A) through (D) the estimation effect can be seen by comparing the columns titled “Conventional $G(\tilde X'\beta^\circ)$” and “Conventional $\Lambda(\tilde X'\hat\beta)$.” In the former case there is no estimation effect and any deviation from the nominal confidence level of 90% is a small sample phenomenon.\footnote{For example, for $c=0.2$ the value of TP$(c,\beta^\circ)$ is close to the upper bound 1, and the coverage probability of the fixed-$\beta$ CI is only 80% for $n=200$.} Over the various cases, the estimation effect ranges from essentially zero to as large as a 30 to 40 percentage point difference in coverage probability. In panel (E), the comparison between the same two columns includes the estimation effect as well as some “bias” due to the fact that the first stage logit regression is misspecified.
The theory presented in Section (ref) gives insight into why the estimation effect is negligible in some cases. In particular, consider the parameter $TP-FP$ on panels (A) through (C) with $c=0.5$. As the predictors are independent standard normal variables and there is no constant in the DGP, the symmetry of the logistic cdf gives $\pi=\mathbb{P}(Y=1)=0.5$. Therefore, when $c=0.5$, $TP-FP$ is a scalar multiple of $(1-c)\piTP-c(1-\pi)FP$. As explained in footnote (ref), inference about this particular linear combination is not impacted by the pre-estimation effect. This is clearly reflected in the estimation results. By contrast, in panel (D) the predictor distribution is not symmetric around zero, so $\pi\neq 0.5$, and the estimation effect is indeed present for $TP-FP$ even when $c=0.5$.
The second main message is that the proposed analytical correction works well in virtually all the cases considered here. This includes panel (A), where the sample size is small, and panel (E), where the first stage logit model is misspecified. Not surprisingly, under misspecification the corrected CI can also fall somewhat short of the 90% confidence level, but it still represents a large improvement over conventional inference.
We provided both pointwise and uniform asymptotic results that describe the distribution of an empirical ROC curve based on a pre-estimated index. The core theory is complete. Ongoing work consists of: (i) developing appropriate test procedures when the first stage models are nested and the ROC influence functions are the same under the null; (ii) additional simulations that illustrate the small sample performance of the uniform asymptotic results, the practical use of the tests, and the power gains afforded by in-sample inference.