EconBase
← Back to paper

Inference for ROC Curves Based on Estimated Predictive Indices

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,474 characters · 19 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for ROC Curves Based on Estimated Predictive Indices

abstractWe provide a comprehensive theory of conducting in-sample statistical inference about receiver operating characteristic (ROC) curves that are based on predicted values from a first stage model with estimated parameters (such as a logit regression). The term “in-sample” refers to the practice of using the same data for model estimation (training) and subsequent evaluation, i.e., the construction of the ROC curve. We show that in this case the first stage estimation error has a generally non-negligible impact on the asymptotic distribution of the ROC curve and develop the appropriate pointwise and functional limit theory. We propose methods for simulating the distribution of the limit process and show how to use the results in practice in comparing ROC curves. Keywords: classification, ROC curve, uniform asymptotics, estimation effect JEL codes: C25, C38, C46, C52

Introduction

Binary prediction or classification is a fundamental problem in statistics and machine learning with applications in many scientific disciplines including economics. The predictive ability of statistical models used for binary classification is frequently evaluated through the receiver operating characteristic (ROC) curve, designed to summarize the tradeoffs between the probability of a true positive prediction (vertical axis) and a false positive prediction (horizontal axis) as one combines a predictive index with a varying classification threshold. Though its origins are in the signal detection and medical diagnostics literature, in recent years ROC analysis has become increasingly common in financial and economic applications as well (e.g., Anjali and Bossaerts 2014; Bazzi et al.\ 2021; Bonfim et al.\ 2021; Berge and Jorda 2011; Kleinberg et al.\ 2018; Lahiri and Wang 2013; Lahiri and Yang 2018; McCracken et al.\ 2021; Schularik and Taylor 2012 and many others).

While there is a large literature on the statistical properties of empirical ROC curves, the standard distributional theory assumes that the signal or predicitive index used for classification is either directly observed---it is “raw data”---or that it is a fixed function of raw data. However, if the signal itself is generated from an underlying regression model with estimated coefficients, conducting in-sample inference about ROC curves based on the traditional theory can be highly misleading. For instance, Demler et al.\ (2012) point out that the standard DeLong et al.\ (1988) test for comparing AUCs for different (but potentially correlated) signals can lead to flawed inference if the signals come from nested models with estimated coefficients.\footnote{AUC stands for “area under [the ROC] curve”. It is an overall performance measure for binary prediction models. AUC=1/2 corresponds to no predictive power.} Similarly, Lieli and Hsu (2019) demonstrate that the asymptotic normality results in Bamber (1975) are inappropriate for testing AUC=1/2 for models with estimated parameters.

The central contribution of this paper is the development of a general functional limit theory for the empirical ROC curve that takes the pre-estimation effect into account. Regarding the ROC curve as a random function defined over the [0,1] interval, we provide a uniform influence function representation theorem, and show that the difference between the empirical and population ROC curves converges weakly to a mean zero Gaussian process with a given covariance structure at the parametric rate. These results constitute a non-trivial extension of the functional limit result in Hsieh and Turnbull (1996), who work under the assumption that the observations available on the predictive index are i.i.d.\ conditional on the outcome. (If the predictive index depends on coefficients estimated in-sample, this assumption is no longer valid, as the parameter estimates depend on all data points.) Our results not only allow for the construction of a uniform confidence band for the ROC curve but also facilitate model selection through handling virtually any comparison between two correlated ROC curves (e.g., testing dominance or partial dominance; testing the difference between AUCs or partial AUCs, etc.) In terms implementation, we propose two methods to simulate the limiting distribution of the empirical ROC curve, one of which is the weighted bootstrap by Ma and Kosorok (2005). Although some type of bootstrap procedure would be a natural way to approach the pre-estimation problem in practice even without a theory, our results provide rigorous justification and guidance for doing this.

A second contribution of the theory developed in this paper is that it provides insight into what determines the impact of the first stage estimation error on the asymptotic distribution of the ROC curve. The derivatives of the true and false positive prediction rates with respect to the first stage model parameters play a central role --- if these gradients vanish at the pseudo-true parameter values, then so does the estimation effect. Nevertheless, the gradients also depend on the classification threshold and will not generally be negligible along the entire ROC curve. For associated functionals, it is the gradient of the functional that drives the estimation effect. Some functionals, e.g., the area under the curve, have the property that they are maximal when the predictive index is given by $p(X)=P(Y=1|X)$, the conditional probability of a positive outcome $Y$ given the covariates $X$. For such functionals the estimation effect is negligible when the first stage model is correctly specified for $p(X)$ and the first stage estimator converges at the root-n rate. Nevertheless, the first stage estimation error will generally affect the asymptotic distribution under misspecification; thus, our results facilitate robust inference.

There are additional technical contributions that are more subtle. In employing standard empirical process techniques to derive our results, we make most of the fact that the population and sample ROC curves are invariant to monotone increasing transformations of the predictive index. This observation allows us to leverage powerful assumptions that may seem restrictive at first glance. In particular, we use the assumption that the density $f_0$ of the predictive index conditional on $Y=0$ is bounded away from zero to derive various uniform approximations. The problem is that even if the individual predictors have densities bounded away from zero (already a big if), the predictive index may not share this property, as it often involves a linear combination of the predictors.\footnote{Think of the sum of two independent uniform[0,1] random variables or the central limit theorem for that matter.} Nevertheless, one can always find a strictly increasing transformation of the predictive index, say, the probability integral transform, so that the post-transform $f_0$ will be greater than some $\epsilon>0$ across the whole support. What matters for the asymptotic theory is the properties of the likelihood ratio $f_1/f_0$, which is invariant to monotone increasing transformations (here $f_1$ is the conditional density of the predictive index given $Y=1$). In particular, uniform inference is possible only for parts of the ROC curve that are generated by thresholds falling into some interval $[c_L,c_U]$ over which $f_1/f_0$ is bounded and bounded away from zero.\footnote{That such an interval exists is a weak assumption; that it coincides with the support of $f_0$ is a much stronger one.} But given the properties of the likelihood ratio, one is free to assume the theoretically most convenient scenario about the individual density $f_0$ that is achievable through monotone transformations, even if this transformation is not implemented or even identified.

We must also point out some technical limitations of the paper. First, we do not allow for serial dependence in the data, precluding time series applications such as the evaluation of recession forecasting models. Nevertheless, our proofs rely mostly on high level conditions; specifically, the asymptotically linear representation of the first stage estimator, the stochastic equicontinuity of the empirical process defined by the (pseudo-true) predictive index, and the uniform continuity of some derivatives. Given the availability of these conditions for stationary, weakly dependent time series, we conjecture that our representation results should generalize to this setting with relatively straightforward modifications. However, simulating the asymptotic distribution of the limiting process would require more complex procedures and we do not pursue this extension here. Second, the predictive model evaluated at the pseudo-true parameter values must have strictly positive variance. This is not an innocent assumption in that it rules out a completely uninformative predictive model. For example, our results are not suitable for testing the hypothesis that AUC=1/2; see Lieli and Hsu (2019) for some specialized results in this very non-standard case. More generally, in using our results to compare two ROC curves, the difference of the two influence functions evaluated at the pseudo-true parameter values must have strictly positive variance as well. This condition can be violated when the first stage models are nested, and we are currently working on some test procedures that are applicable in this scenario as well. Finally, we only consider parametric estimators of $p(X)$ as the first stage model; nonparametric estimators that converge slower than the root-n rate are ruled out.

One might discount the practical relevance of our theoretical results discussed above based on the fact that the first stage estimation problem can be avoided by conducting out-of-sample evaluation. If the ROC curve is constructed over a test sample that is independent of the training sample used to estimate the first stage model, then the asymptotic validity of the standard inference procedures is restored. We acknowledge this point but offer two responses. First, out-of-sample evaluation is costly: it leads to power loss in model comparisons and the potential dependence of the results on the particular split(s) used. In fact, one could argue that out-of-sample evaluation is a necessity forced on practitioners by the fact that it is often very difficult to characterize analytically or in a practically useful way how goodness-of-fit measures behave over the training sample so that one can compensate for overfitting. In this case we do provide such a result. Second, apart from dealing with the pre-estimation problem, our results provide a unified framework for conducting uniform inference, and comparing ROC curves estimated over the same sample in virtually any way.

This work has ties to several strands of the statistics and econometrics literature. We have already cited a number of classic works on the statistical properties of the empirical ROC curve that maintain the assumption of a directly observed signal (Bamber 1975, DeLong et al.\ 1988, Hsieh and Turnbull 1996). It is the last of these papers that is closest to ours; however, the pre-estimation effect is obviously missing from their framework and they actually do not exploit their functional limit result for inference apart from re-deriving the asymptotic normality result of Bamber (1975) for the empirical AUC. Instead, they focus on estimating the ROC curve under an additional “binormal” assumption, i.e., when the signal has a normal distribution conditional on both outcomes.

Pre-estimation problems have a long history in the literature; for example, Pagan (1984) studied the distributional consequences of including “generated regressors” into a regression model. As mentioned above, Demler et al.\ (2012) pointed out the relevance of the pre-estimation effect in the context of ROC analysis. More generally, our work is related to papers dealing with two-step estimators where the first step involves estimating some nuisance parameter whose sampling variation potentially affects the otherwise well-understood second stage. Abadie and Imbens (2016) is a relatively recent example in the context of matching estimators.

The application of our results in testing for dominance relations and AUC differences across ROC curves has similarities to stochastic dominance tests; see, e.g., Barrett and Donald (2003), Linton, Maasoumi and Whang (2005), Linton, Song and Whang (2010) and Donald and Hsu (2016). Some papers, such as Linton, Maasoumi and Whang (2005) and Linton, Song and Whang (2010), even allow for generated variables in this context. Finally, the paper speaks indirectly to the forecasting literature on the relative merits of in-sample vs.\ out-of-sample model evaluation (e.g., Inoue and Kilian 2004, Clark and McCracken 2012). The connection lies in the fact that we extend the scope of in-sample evaluation methods in binary prediction.

The rest of the paper is organized as follows. Section (ref) sets up the prediction framework and introduces the ROC curve along with some of its basic properties. Section (ref) discusses the estimation effect in detail and presents pointwise (fixed-cutoff) asymptotic results. The functional limit theory is contained in Section (ref). In Section (ref) we show how to use the abstract results for conducting inference about the ROC curve; we discuss dominance testing, AUC comparisons, etc. Section (ref) presents Monte Carlo results highlighting the impact of first stage estimation on the distribution of the ROC curve. Section (ref) concludes.

Making and evaluating binary predictions

Cutoff rules and the ROC curve

Let $Y\in\{0,1\}$ be a Bernoulli random variable representing some outcome of interest and $X$ be a $k\times 1$ vector of covariates (predictors). We consider point forecasts (classifications) of $Y$ that are constructed by combining a scalar predictive index $G(X)$ with a suitable cutoff (threshold) $c$. More specifically, the prediction rule for $Y$ is given by

equation[equation omitted — 63 chars of source]

where $1[.]$ is the indicator function. The role of the function $G: \mathbb{R}^k\rightarrow\mathbb{R}$ is to aggregate the information that the predictors contain about $Y$ while the choice of $c$ governs the use of this information. With $G$ and $c$ unrestricted, ((ref)) represents a very general class of prediction rules.

In many binary prediction problems there is also a loss function $\ell(\hat y, y)$ that describes the cost of predicting $Y=\hat y$ when the realized outcome is $Y=y$. If the decision maker's objective is to minimize expected loss conditional on $X$, the optimal choice of $G$ and $c$ is determined as follows. Given the observed value of $X$, the optimal point forecast of $Y$ solves

equation[equation omitted — 164 chars of source]

where $p(X)=\mathbb{P}(Y=1|X)$ is the conditional probability of $Y=1$ given $X$. Adopting the normalization $\ell(0,0)=\ell(1,1)=0$ and assuming $\ell(0,1)>0$ and $\ell(1,0)>0$, it is straightforward to verify that the optimal prediction rule is given by

equation[equation omitted — 76 chars of source]

where $c_\ell=\ell(1,0)/[\ell(0,1)+\ell(1,0)]$.\footnote{In case $p(X)=c_\ell$, which is often a zero probability event, the decision maker is indifferent between predicting 0 or 1. The formula stated above arbitrarily specifies $\hat Y=0$ in this case.}

Equation ((ref)) reveals that the optimal predictive index is $p(X)$ regardless of $\ell$ while the optimal choice of $c$ is fully determined by $\ell$ (specifically, by the relative cost of a false alarm versus a miss). This simple observation motivates a two-step empirical strategy in binary prediction.\footnote{See Elliott and Lieli (2013) for an alternative approach where the decision rule ((ref)) is estimated in a single step based on a specific loss function.} First, one models and estimates $p(X)$ using data on $(Y,X)$; a common approach is to specify a parametric model $G(X,\beta)$ for $p(X)$ and to estimate it by maximum likelihood (logistic regression is a leading example). In the second step a point forecast is obtained by combining the estimated conditional probability $G(X,\hat\beta)$ with a suitable cutoff, which depends on the forecaster's or forecaster user's preferences. Thus, there is a separation between the construction of the predictive index, representing the objective information available to the forecaster, and the use of that information, governed by the loss function.

We will now define the population receiver operating characteristic (ROC) curve. Let $G(X,\beta)$ be a predictive index with a fixed value of the parameter $\beta$.\footnote{The following definitions do not depend on the parametric structure and generalize immediately to any predictive index $G(X)$. We work with a parametric specification in anticipation of studying the pre-estimation step.} Combined with a cutoff $c$, the resulting prediction rule produces true positive predictions and false positive predictions (false alarms) with the following probabilities:

eqnarray*[eqnarray* omitted — 173 chars of source]

where TP and FP stand for the rate of “true positive” and “false positive” predictions, respectively. As the cutoff $c$ varies, both quantities change in the same direction; in general, TP can be increased only at the cost of increasing FP as well. The ROC curve traces out all attainable (FP, TP) pairs in the $[0,1]\times[0,1]$ unit square, i.e., it is the locus \[ \Big\{\big(FP(c,\beta), TP(c,\beta)\big): c\in\mathbb{R}\Big\}. \]

Intuitively, the ROC curve is a way of summarizing the information content of $G(X,\beta)$ about the outcome $Y$ without committing to any particular cutoff, i.e., loss function. The use of such a forecast evaluation tool is particularly appropriate in situations in which there many potential forecast users with diverse loss functions; see Lieli and Nieto-Barthaburu (2010). It is also clear from the definition that the ROC curve is invariant to strictly monotone transformations of $G(X,\beta)$.

The ROC curve based on the true conditional probability function $p(X)$ possesses some optimality properties. To state these in a parametric modeling framework, we introduce the following correct specification assumption.

assumptionThere exists some point $\beta^\circ$ in the parameter space $\mathcal{B}\subseteq\mathbb{R}^p$ such that $G(X,\beta^\circ)=p(X)$ almost surely.

We state the following result.

proposition(i) Given Assumption (ref), $\beta^\circ$ solves the following maximization problem for any value of $c$: \begin{equation*} \max_\beta \big[ (1-c)\pi TP(c,\beta)-c(1-\pi)FP(c,\beta)\big], \end{equation*} where $\pi=\mathbb{P}(Y=1)$. (ii) Given Assumption (ref), define $F^\circ_c= FP(c,\beta^\circ)$. Then $\beta^\circ$ also solves the following constrained maximization problem for any value of $c$: \begin{equation*} \max_\beta TP(c,\beta) s.t. FP(c,\beta)=F^\circ_c \end{equation*}

\paragraph{Remarks:}

enumerate• Part (i) is a consequence of the predictor ((ref)) solving ((ref)) for any given value of $X$. To see this, note that by the law of iterated expectations and the monotonicity of the expectation operator, $\hat Y^*(c_\ell)$ also solves the unconditional expected loss minimization problem $\min_{\hat Y}E_{XY}[\ell(\hat Y,Y)]$, where the minimization is over all random variables $\hat Y$ that are (measurable) transformations of $X$. It is easy to verify that $E[\ell(\hat Y,Y)]$ can be written as \[ [\ell(0,1)+\ell(1,0)]\cdot\big[ c_\ell(1-\pi)\mathbb{P}(\hat Y=1|\, Y=0)-(1-c_\ell)\pi \mathbb{P}(\hat Y=1|\, Y=1)\big]+\pi\ell(0,1), \] which immediately implies the result. • Part (ii) is a consequence of part (i) and it means that for any given FP rate, it is the ROC curve based on $p(X)$ that achieves the largest possible TP rate. In other words, the ROC curve associated with $p(X)$ weakly dominates any other ROC curve that is constructed based on some index $G(X)$.\footnote{To see this, fix a false positive rate $F_0\in[0,1]$, and find the cutoff $c_0$ that produces $FP(\beta^\circ, c_0)=F_0$ (for simplicity, assume that exact equality can be achieved). Let $\beta'$ be any other parameter value satisfying $FP(\beta', c_0)=F_0$. Proposition (ref)(i) implies $$(1-c_0)\pi TP(c_0,\beta^\circ)-c_0(1-\pi)FP(c_0,\beta^\circ)\ge (1-c_0)\pi TP(c_0,\beta')-c_0(1-\pi)FP(c_0,\beta')$$ so that $TP(c_0,\beta^\circ)\ge TP(c_0,\beta')$.} • These results are not new; they have appeared in the ROC literature in alternative formulations. See, e.g., Egan (1975) and Pepe (2003, Section 4).

The sample ROC curve and conventional inference

Throughout the paper we maintain the assumption that the available data consists of a random sample. More formally:

assumptionThe sample $\{(Y_i,X_i)\}_{i=1}^n$ consists of independent and identically distributed observations on the random vector $(Y,X)\in\{0,1\}\times\mathbb{R}^k$.

Given the sample and a fixed value of $\beta$, the empirical ROC curve is defined as the locus $\big\{(\widehat{FP}(c,\beta), \widehat{TP}(c,\beta)): c\in\mathbb{R}\big\}\subset [0,1]\times[0,1]$, where

eqnarray*[eqnarray* omitted — 169 chars of source]

and $n_1=\sum_{i=1}^n Y_i$, $n_0=n-n_1$.

The simplest type of inference about an ROC curve involves constructing (joint) confidence intervals for $TP(c,\beta)$ and $FP(c,\beta)$ for one threshold $c$ at a time and for fixed values of the coefficient vector $\beta$. In this case one can use the CLT to arrive at the normal approximations:

eqnarray[eqnarray omitted — 242 chars of source]

where TP is a shorthand for $TP(c,\beta)$ (and similarly for FP). These results immediately provide asymptotic confidence intervals for TP and FP, and a joint confidence rectangle is also easy to construct due to the independence of $\widehat{TP}(c,\beta)$ and $\widehat{FP}(c,\beta)$; see Pepe (2003), Section 2.2.2 for details.\footnote{The two statistics are independent because they are computed from two disjoint sets of observations; namely, the $Y=1$ and $Y=0$ subsamples.}

Furthermore, for fixed $\beta$ one can apply the asymptotic normality results in Bamber (1975) to conduct inference about the AUC, and the DeLong et.\ al.\ (1988) test for comparing the areas under ROC curves based on different (but non-random) values of $\beta$. The nonparametric functional limit result in Hsieh and Turnbull (1996) also applies.

In-sample inference: pointwise asymptotics

We will now develop a comprehensive theory of in-sample inference about individual points on the ROC curve taking the pre-estimation effect into account. We present both analytical results and results based on the weighted bootstrap. We start by describing the setup and stating some technical conditions.

First stage estimation and technical assumptions

The sample $\{(Y_i,X_i)\}_{i=1}^n$ now plays a dual role. First, it is used to construct an estimated parameter vector $\hat\beta=\hat\beta_n$. Typically, $\hat\beta$ consists of an intercept and slope coefficients from some type of regression of $Y$ on $X$ (e.g., linear, logit or probit). Second, the same sample is used to compute the predictive index values $G(X_i,\hat\beta)$, $i=1,\ldots, n$, and to construct the empirical ROC curve as described in Section (ref). We impose the following high level condition on $\hat\beta$.

assumption(i) Let $\hat\beta$ be an M-estimator of $\beta$ so that \begin{align*} \hat{\beta}\equiv \arg\max_{\beta\in \mathcal{B}} \frac{1}{n} \sum_{i=1}^n q(Y_i,X_i,\beta). \end{align*} (ii) There is a point $\beta^*$ in the interior of the compact parameter space $\mathcal{B}\subset\mathbb{R}^p$ so that \begin{align} \sqrt{n}(\hat{\beta}_n-\beta^*)=\frac{1}{\sqrt{n}}\sum_{i=1}^n\psi_\beta(Y_i,X_i,\beta^*)+o_p(1), \end{align} where $\psi_\beta: \mathbb{R}^{1+k+p}\rightarrow \mathbb{R}^p$ is a given function with $E[\psi_\beta(Y,X, \beta^*)]=0$ and $E\|\psi_\beta(Y,X, \beta^*)\|^{2+\epsilon}<\infty$ for some $\epsilon>0$. Furthermore, $\beta^*=\beta^\circ$ under Assumption (ref).

Assumption (ref) states that $\hat\beta$ is an $M$-estimator with an asymptotically linear representation, implying that $\hat\beta$ is asymptotically normally distributed. The stated conditions do not require the first stage model to be correctly specified for $p(X)$; $\beta^*$ simply stands for the probability limit of $\hat\beta$, i.e., the pseudo-true value of the parameter vector. We nevertheless assume that $\hat\beta$ is consistent for $\beta^\circ$ under correct specification (Assumption (ref)). The reason for allowing for misspecification is twofold. First, the first stage predictive model is often simply an approximation of $p(X)$, e.g., a linear regression. Second, as we will see, the estimation effect can depend on whether or not the model is correctly specified.

assumptionThe empirical processes \[ (c,\beta)\mapsto\sqrt{n}(\hat \mathbb{P}_0-\mathbb{P}_0)1[G(X,\beta)\le c]\text{ and }(c,\beta)\mapsto\sqrt{n}(\hat \mathbb{P}_1-\mathbb{P}_1)1[G(X,\beta)\le c] \] are stochastically equicontinuous over $\mathbb{R}\times \mathcal{B}$, where $\mathbb{P}_j$ denotes probability conditional on $Y=j$ and $\hat\mathbb{P}_j$ is the corresponding empirical measure in the $Y=j$ subsample.

The stochastic equicontinuity requirement in Assumption (ref) limits the complexity of the model $G(X,\beta)$ and plays an important role in handling the estimation effect. It holds, for example, if $G(X,\beta)=X'\beta$ or $G(X,\beta)=G(X'\beta)$ with $G$ bounded (see the definition of a type I class in Andrews 1994). Apart from a small degree of added generality, we state stochastic equicontinuity as a high level condition to make it more transparent what is required for our results.

The final assumption states the differentiability of $TP$ and $FP$ with respect to the components of $\beta$. Let $\nabla_\beta$ denote the corresponding gradient operator and $B^*(r)$ the open ball with radius $r>0$ centered on $\beta^*$.

assumptionFor any given cutoff $c$, the gradient vectors $\nabla_\betaTP(c,\beta)$ and $\nabla_\betaFP(c,\beta)$ exist and are continuous over $B^*(r)$ for some $r>0$.

As we will shortly see, the first stage estimation of $\beta$ affects the asymptotic distribution of the ROC curve through the derivatives presented in Assumption (ref).

Theoretical illustration of the estimation effect

Let $c$ be a given value of the cutoff; we want to conduct inference about the corresponding point $(FP(c,\beta^*), TP(c,\beta^*))$ on the limiting ROC curve. To isolate the effect of the first stage estimator $\hat\beta$ on the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$, we can write

eqnarray[eqnarray omitted — 319 chars of source]

where the second equality is due to the fact that the process $\sqrt{n}(\widehat{TP}-TP)$ is stochastically equicontinuous (Assumption (ref)), implying \[ \sqrt{n}[\widehat{TP}(c,\hat\beta)-TP(c,\hat\beta)]-\sqrt{n}[\widehat{TP}(c,\beta^*)-TP(c,\beta^*)]=o_p(1), \] given that $\hat\beta\rightarrow_p\beta^*$. As $\beta^*$ is fixed and $Var[G(X,\beta^*)]>0$, the first term in equation ((ref)) has the asymptotic distribution given by ((ref)) and the second term represents the effect of estimating $\beta^*$. Does this term have a non-negligible effect on the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$, and if yes, how do we characterize it?

To address these questions, we can use Assumption (ref) to expand the second term in ((ref)) around $\beta^*$ to obtain

eqnarray[eqnarray omitted — 214 chars of source]

Equation ((ref)) shows that the estimation effect is negligible whenever $\nabla_\beta TP(c,\beta^*)=0$. However, as Proposition (ref)(ii) shows, this condition does not generally hold even if $G(X,\beta)$ is correctly specified (i.e., $\beta^*=\beta^\circ$), because $TP(c,\beta^\circ)$ solves a constrained (rather than unconstrained) optimization problem.\footnote{The first order conditions are $\nabla_\beta TP(c,\beta^\circ)=\lambda\nabla_\beta FP(c,\beta^\circ)$ for some scalar $\lambda$ and $FP(c,\beta)=F^\circ_c$. The Lagrange multiplier $\lambda$ is generally non-zero, at least when $TP(c,\beta^\circ)<1$.} Under misspecification Proposition (ref)(ii) does not apply, but of course there is still no general reason for $\nabla_\beta TP(c,\beta^*)$ to vanish. Therefore, in either case the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$ and $\widehat{FP}(c,\hat\beta)$ will generally differ from that stated under ((ref)) and ((ref)) because $\sqrt{n}(\hat\beta-\beta^*)=O_p(1)$.\footnote{Proposition (ref)(i) implies that under correct specification one can conduct inference about the linear combination $(1-c)\piTP-c(1-\pi)FP$ without the need to consider the pre-estimation effect. This is because the true value of $\beta$ maximizes this linear combination and hence the corresponding gradient driving the estimation effect vanishes. More generally, the estimation effect is negligible for a functional of the ROC curve if (i) the ROC curve based on $p(X)$ maximizes that functional and (ii) Assumption (ref) holds.}

Pointwise inference based on analytical results

To describe the asymptotic distribution of $\widehat{TP}(c,\hat\beta)$ in more detail, we can further expand the decomposition in ((ref)) by substituting in the asymptotically linear (influence function) representation of the two terms. Using the definition of $\widehat{TP}(c,\beta^*)$ and Assumption (ref), it is straightforward to verify that

align[align omitted — 349 chars of source]

where the definition of $\psi_{TP}$ is enclosed by the braces on the previous line.

Of course, $\widehat{FP}(c,\hat\beta)$ has a corresponding asymptotically linear representation with influence function \[ \psi_{FP}(Y_i,X_i,c,\beta^*)=\frac{1-Y_i}{1-\pi}\big[1(G(X_i,\beta^*)>c)-FP(c,\beta^*)\big]+\nabla_\betaFP(c,\beta^*)\psi_\beta(Y_i,X_i,\beta^*). \] Stacking the influence functions as $$\psi(Y_i,X_i,c,\beta^*)=[\psi_{TP}(Y_i,X_i,c,\beta^*),\psi_{FP}(Y_i,X_i,c,\beta^*)]'$$ and applying the multivariate CLT gives the asymptotic joint distribution of an individual point $(\widehat{TP}(c,\hat\beta),\widehat{FP}(c,\hat\beta))$ on the sample ROC curve.

propositionSuppose that Assumptions (ref) to (ref) are satisfied. Then \begin{equation} \sqrt{n} \begin{pmatrix} \widehat{TP}(c,\hat\beta)-TP(c,\beta^*)\\ \widehat{FP}(c,\hat\beta)-FP(c,\beta^*)\\ \end{pmatrix} =\frac{1}{\sqrt{n}}\sum_{i=1}^n \psi(Y_i,X_i,c,\beta^*)+o_p(1) \rightarrow_d N[0,E(\psi\psi')] \end{equation} for cutoffs $c$ for which $E[\psi^2_{TP}(Y_i,X_i,c,\beta^*)]>0$ and $E[\psi^2_{FP}(Y_i,X_i,c,\beta^*)]>0$.

\paragraph{Remarks}

enumerate• Using Proposition (ref), it is easy to obtain the asymptotic distribution of any linear combination $a\widehat{TP}(c,\hat\beta)+b\widehat{FP}(c,\hat\beta)$. • Proposition (ref) is a “pointwise” result in the sense that the cutoff $c$ is assumed to be fixed. It is straightforward to generalize the setup so that one can make joint inference about points that are associated with a finite number of different cutoffs. One can simply stack the values of the influence function $\psi$ evaluated at these cutoffs and a result analogous to ((ref)) will continue to hold. • The variance condition $E[\psi^2_{TP}]>0$ will generally hold for interior points $TP(c,\beta^*)\in (0,1)$ but fail for $TP(c,\beta^*)\in \{0,1\}$. The same is true for $FP$.

We supplement Proposition (ref) by some results that reveal the structure of $\nabla_\beta TP$ and $\nabla_\beta FP$ and facilitate their estimation. Let $\partial_{j}$ denote the partial derivative operator with respect to the $j$th component of $\beta$.

assumption(i) $G(X,\beta)$ is twice continuously differentiable (a.s.) w.r.t.\ $\beta$ on $B^*(r)$ for some $r>0$ with $\sup_{\beta\in B^*(r)}|\partial_{jj}G(X,\beta)|\le M$ (a.s.) for some $M>0$. (ii) The conditional density of $G(X,\beta^*)$ given $Y=0,1$ exists. The conditional density of $G(X,\beta^*)$ given $\partial_jG(X,\beta^*)$ and $Y=y$ also exists and is bounded uniformly by some $M>0$ for almost all values of $\partial_jG(X,\beta^*)$, $y=0,1$, and all $j$. (iii) $E\big[|\partial_j G(X,\beta^*)|\,\big|\,Y=1\big]<\infty$ for all $j$.
propositionSuppose that Assumptions (ref) and (ref) hold. Then: \begin{equation} \nabla_\beta TP(c,\beta^*)=E\Big[\nabla G_\beta(X,\beta^*)\Big|\,G(X,\beta^*)=c, Y=1 \Big]f^*_1(c), \end{equation} where $f^*_1(c)$ is the conditional density of $G(X,\beta^*)$ given $Y=1$. If, in addition, Assumption (ref) is satisfied ($\beta^*=\beta^\circ$), then the expectation in equation ((ref)) does not need to be conditioned on $Y=1$. The formula for $\nabla_\beta FP(c,\beta^*)$ is analogous; it conditions on $Y=0$ throughout.

Finally, we specialize Propositions (ref) and (ref) by imposing a logit first stage.

assumptionSuppose that the first stage estimation consists of a logit regression of $Y$ on $X$ and a constant so that $G(X,\hat\beta)=\Lambda(\tilde X'\hat\beta)$, where $\Lambda(\cdot)$ is the logistic c.d.f., $\tilde X=(1, X')'$ and $\hat\beta$ is the maximum likelihood estimator.
propositionSuppose that Assumption (ref) is satisfied. Then: \begin{itemize} • $\psi_\beta(Y_i,X_i,\beta)=A_\beta^{-1}X_i[Y_i-\Lambda(\tilde X_i'\beta)]$, where $A_\beta=E\{\Lambda(\tilde X'\beta)[1-\Lambda(\tilde X'\beta)]\tilde X\tilde X'\}$. • The components of $\nabla_\beta TP(c,\beta^*)$ are given by: \begin{equation} c(1-c)E\big[X_j\big|\,\Lambda(\tilde X'\beta^*)=c, Y=1 \big]f_1^*(c),\; j=0,1,\ldots,d, \end{equation} where $X_0\equiv 1$, $X_j$, $j=1,\ldots,d$ is the $j$th component of $X$, and $f_1^*(c)$ is the conditional density of $\Lambda(\tilde X'\beta^*)$ given $Y=1$. \end{itemize}

\paragraph{Remarks:}

enumerate• The proofs of Propositions (ref) and (ref) are presented in Appendix B. • The existence of $f_1^*(c)$ requires that $X$ has a continuous component and the corresponding coefficient in $\beta^*$ is nonzero. This rules out $X$ and $Y$ being independent. • The expression for $\psi_\beta$ follows from formulas (12.16), (15.18) and (15.19) in Wooldridge (2002) when specialized to the logit case. • One can estimate the unknown quantities in ((ref)) nonparametrically to obtain a semiparametric estimator for $\nabla_\beta TP(c,\beta^*)$. More precisely, expression ((ref)) may actually be estimated in a single step as \begin{equation} c(1-c)\frac{1}{n_1h}\sum_{i: Y_i=1}X_{ji}K\left(\frac{\Lambda(\tilde X_i'\hat\beta)-c}{h}\right), \end{equation} where $K(\cdot)$ is a kernel function and $h$ is a bandwidth that may be chosen according to Silverman's rule of thumb. • Alternatively, if correct specification is assumed in the first stage ($\beta^*=\beta^\circ$), then one can estimate $E[X_j|\,\Lambda=c, Y=1]=E[X_j|\,\Lambda=c]$ by a kernel regression on the full sample and $f^*_1(c)$ by a kernel estimator on the $Y=1$ subsample.

Pontwise inference based on the weighted bootstrap

Here we provide an alternative method for making pointwise inference about the ROC curve by utilizing the weighted bootstrap for M-estimators proposed by Ma and Kosorok (2005). The main advantage of this approach is that it sidesteps the estimation of the gradient vectors $\nabla_\beta TP(c,\beta^*)$ and $\nabla_\beta FP(c,\beta^*)$. Furthermore, the method is similar to the simulation-based procedure that we propose for functional inference in Section (ref).

The weighted bootstrap employs a sequence of (pseudo) random variables as multipliers to simulate the sampling variation of an estimator.

assumptionLet $\{W_i\}_{i=1}^n$ be a sequence of i.i.d.\ (pseudo) random variables, independent of the sample path $\{(Y_i,X_i)\}_{i=1}^n$, with $E(W_i)=1$ and $Var(W_i)=1$.

We first define the weighted bootstrap version of the first stage estimator of $\beta$:

align*[align* omitted — 113 chars of source]

Given $\hat{\beta}^w$, the weighted bootstrap estimators of $TP(c,\beta)$ and $FP(c,\beta)$ are defined as

eqnarray*[eqnarray* omitted — 243 chars of source]
assumptionAssume that \begin{align} \sqrt{n}(\hat{\beta}^w-\beta^*)=\frac{1}{\sqrt{n}}\sum_{i=1}^nW_i \cdot \psi_\beta(Y_i,X_i,\beta^*)+o_p(1), \end{align} where $\beta^*$ and $\psi_\beta(Y_i,X_i,\beta^*)$ are given in Assumption (ref).

Assumption (ref) ensures that the weighted bootstrap is valid for the first stage estimator, i.e., conditional on the data, $\sqrt{n}(\hat{\beta}^w-\hat{\beta})$ has the same limiting distribution as $\sqrt{n}(\hat\beta-\beta^*)$ unconditionally.

Furthermore, by Theorem 2 of Ma and Kosorok (2005), the validity of the weighted bootstrap for $\widehat{TP}^w(c,\hat\beta^w)$ follows from showing that (i) $\widehat{TP}$, $\widehat{FP}$, $\widehat{TP}^w$ and $\widehat{FP}^w$ can be represented as M-estimators and (ii) that these estimators are $\sqrt{n}$-consistent and asymptotically linear.

Item (i) is verified by noting that

eqnarray*[eqnarray* omitted — 291 chars of source]

and similarly for $\widehat{FP}$ and $\widehat{FP}^w$. As for item (ii), Proposition (ref) establishes the asymptotically linear representation of $(\widehat{TP}(c,\hat\beta), \widehat{FP}(c,\hat\beta))$; essentially the same argument also yields

equation*[equation* omitted — 270 chars of source]

Thus, we obtain the following result.

propositionSuppose that Assumptions (ref)-(ref), (ref) and (ref) are satisfied. Then, conditional on the sample path of the data, \begin{equation*} \sqrt{n} \begin{pmatrix} \widehat{TP}^w(c,\hat\beta^w)-\widehat{TP}(c,\hat{\beta})\\ \widehat{FP}^w(c,\hat\beta^w)-\widehat{FP}(c,\hat{\beta})\\ \end{pmatrix} =\frac{1}{\sqrt{n}}\sum_{i=1}^n (W_i-1)\cdot \psi(Y_i,X_i,c,\beta^*) \rightarrow_d N[0,E(\psi\psi')] \end{equation*} with probability approaching one for cutoffs $c$ such that $E[\psi^2_{TP}]>0$ and $E[\psi^2_{FP}]>0$.

\paragraph{Remarks}

enumerate• In applications we suggest letting the weights $W_i$ take the values 0 and 2 with equal probability. The main reason is that with non-negative weights the weighted objective function remains concave if the $q(Y_i,X_i,\beta)$ is concave in $\beta$. This makes it computationally easier to obtain $\hat\beta^w$. • The weighted bootstrap estimator of the asymptotic variance-covariance matrix $\Psi(c)\equiv E(\psi\psi')$ can be constructed as follows. With a minor abuse of notation, let $\widehat R^w(c)=(\widehat{TP}^w(c,\hat\beta^w), \widehat{FP}^w(c,\hat\beta^w))'$ denote the ROC estimate from the $w$th bootstrap cycle, $w=1,\ldots, \mathcal{W}$. Then one can estimate $\Psi(c)$ by \begin{align*} &\widehat\Psi_\mathcal{W}(c)=\frac{n}{\mathcal{W}}\sum_{w=1}^\mathcal{W} \big(\widehat{R}^w(c)-\overline{\widehat{R}}^w(c)\big)\big(\widehat{R}^w(c)-\overline{\widehat{R}}^w(c)\big)', where\\ &\overline{\widehat{R}}^w(c)=\frac{1}{\mathcal{W}}\sum_{w=1}^\mathcal{W}\widehat{R}^w(c). \end{align*} We have that conditional on sample path with probability approaching one, \begin{align*} &\widehat{\Psi}_{\mathcal{W}}(c)\stackrel{p}{\rightarrow}_w \frac{1}{n}\sum_{i=1}^n \psi(Y_i,X_i,c,\beta^*)\psi(Y_i,X_i,c,\beta^*)'+o_p(1), \end{align*} where $\stackrel{p}{\rightarrow}_w$ denotes probability limit under the law of the $W_i$'s. It follows that \begin{align*} \lim_{\mathcal{W}\rightarrow \infty} \widehat{\Psi}_\mathcal{W}(c)\stackrel{p}{\rightarrow} {\Psi}(c). \end{align*}

In-sample inference: uniform asymptotics

To derive uniform results, we first express the ROC curve explicitly as a function over the interval [0,1]. Let the inverse of the decreasing function $c\mapsto FP(c,\beta)$ be defined as \[ FP^{-1}_{\beta}(t)=\inf\{c: FP(c,\beta)\le t\},\; t\in [0,1]. \] The more compact notation on the l.h.s.\ emphasizes that the inverse is taken with respect to the cutoff $c$ for a fixed value of $\beta$. Thus, is $FP^{-1}_{\beta}(t)$ as the “first” (smallest) cutoff value at which the false positive rate is equal to $t$ or falls below $t$.\footnote{Of course, if $FP(c,\beta)$ is strictly decreasing and continuous in $c$, then $FP^{-1}_{\beta}(t)$ is the unique solution to the equation $FP(c,\beta)=t$.} Because $1-FP(c,\beta)$ is the c.d.f.\ of the conditional distribution of $G(X,\beta)$ given $Y=0$, an equivalent interpretation of $FP^{-1}_{\beta}(t)$ is that it is the $(1-t)$-quantile of this distribution.

We can now represent the ROC curve as a function that returns the true positive rate associated with given false positive rate $t$:

equation[equation omitted — 109 chars of source]

For a given parameter value $\beta$, the sample ROC curve is defined by replacing $TP(\cdot,\beta)$ and $FP^{-1}_\beta(\cdot)$ by sample analogs: $\widehat R(t,\beta)=\widehat{TP}\big(\widehat{FP}^{-1}_{\beta}(t),\beta\big),\;t\in[0,1]$.

Additional technical assumptions for uniform inference

Our goal is to characterize the statistical behavior of the random function $t\mapsto \hat R(t,\hat\beta)$ over the interval $[0,1]$. This requires some additional assumptions.

assumption(i) The conditional distribution of $G(X,\beta^*)$ given $Y=0$ has compact support $[a_0,b_0]$ and probability density function $f_0^*(c)$ that is continuous (and hence bounded) over $[a_0,b_0]$ and satisfies $\inf\{f_0^*(c): c\in[a_0,b_0]\}\ge\delta$ for some $\delta>0$. (ii) The conditional distribution of $G(X,\beta^*)$ given $Y=1$ has compact support $[a_1,b_1]$ and a probability density function $f_1^*(c)$. (iii) There exits a subinterval $[c_{0,L},c_{0,U}]\subseteq [a_0,b_0]$ such that $f^*_1(c)/f^*_0(c)$ is continuous (and hence bounded) over $[c_{0,L},c_{0,U}]$ and satisfies $\inf\{f_1^*(c)/f_0^*(c): c\in[c_{0,L},c_{0,U}]\}\ge\delta$ for some $\delta>0$.

Assumption (ref) merits careful discussion. An immediate practical implication of part (i) is that the limiting model $G(X,\beta^*)$ must depend on at least one continuous predictor in a nontrivial way. For instance, if the model is based on a linear index, this rules out $X$ being completely independent of $Y$; see Remark 1 after Proposition (ref). Part (iii) implies that $supp(f_0^*)$ and $supp(f_1^*)$ overlap, ensuring that the classification problem is nontrivial. Nevertheless, the overlap does not need to be complete; we allow for applications in which extreme values of the index are associated exclusively with one of the two outcomes.

From a technical standpoint, the main purpose of Assumption (ref) is to facilitate uniform inference by controlling the behavior of the likelihood ratio $f_1^*/f_0^*$. In particular, $f_1^*/f_0^*$ is required to be bounded and bounded away from zero on an interval $[c_{0,L},c_{0,U}]$. Our uniform influence function representation result for $\hat R(t,\hat\beta)$ holds only for quantiles $t$ satisfying $FP^{-1}_{\beta^*}(t)\in [c_{0,L},c_{0,U}]$ or, equivalently, for $t\in [FP(c_{0,U},\beta^*), FP(c_{0,L},\beta^*)]$. While this representation depends on $f_1^*$ and $f_0^*$ only through $f_1^*/f_0^*$, the derivation of the result relies on the additional condition that $f_0^*$ is bounded away from zero (Assumption (ref)(i)). This may seem overly restrictive at first glance---for example, if $G(X,\beta^*)=X'\beta^*$, then predictors with unbounded support are ruled out. Furthermore, it is easy to see that even if all components of $X$ have densities bounded away from zero, their linear combinations will generally not share this property.\footnote{For example, consider the sum of two independent uniform [-0.5,0.5] random variables. The resulting density is $(1-|x|)1_{[-1,1]}(x)$, which tends to zero as $x$ approaches $-1$ or $1$. } However, one can always find a monotone increasing transformation $\Phi(\cdot)$ such that the density of $\Phi[G(X,\beta^*)]$ conditional on $Y=0$ is bounded away from zero, e.g., one can use the probability integral transform to arrive at a uniform[0,1] density. At the same time, such a transformation leaves the ROC curve as well as the range of the likelihood ratio $f_1^*/f_0^*$ unchanged. Thus, the last part of Assumption (ref)(i) is simply a theoretical normalization that does not need to be imposed on the data in practice (see Figure 1 for an illustration).

figure[figure omitted — 845 chars of source]

Of course, Assumption (ref) allows for scenarios in which $[FP(c_{0,U},\beta^*), FP(c_{0,L},\beta^*)]=[0,1]$, i.e., uniform inference is possible along the entire ROC curve. This is the case, for example, if the “propensity score” function $P(Y=1|X=x)$ takes values from an interval $[\delta, 1-\delta]$ for some $0<\delta<1/2$, which implies $supp(f_0^*)=supp(f_1^*)$ and that (iii) holds with $[c_{0,L},c_{0,U}]=[a_0,b_0]$.\footnote{To see this, let $f_x(x)$ denote the density function of $X$. Note that

align*[align* omitted — 186 chars of source]

It follows that

align*[align* omitted — 170 chars of source]

which is bounded below by $ \delta^2/(1-\delta)^2$ and bounded above by $ (1-\delta)^2/\delta^2$. } More generally, Assumption (ref)(iii) allows $f^*_1(c)/f^*_0(c)$ to reach zero or explode for cutoffs $c$ outside the range $[c_{0,L}, c_{1,L}]$. For example, the likelihood ratio vanishes as $c$ approaches $a_0$ from above whenever $a_0<a_1$. In this case the lowest index values imply $Y=0$ and the ROC curve reaches the top of the unit square for some $FP$ rate below unity. Similarly, $f^*_1(c)/f^*_0(c)$ may become unbounded as $c$ approaches $b_0$ from below. This can easily happen when $b_0<b_1$, i.e., the largest index values are associated exclusively with the $Y=1$ outcome. In this case the ROC curve has a positive vertical intercept at $FP=0$. Again, see Figure 1 for an example.

The next assumption is a strengthening of Assumption (ref). These stricter conditions on the gradient vectors $\nabla_\beta TP(c,\beta)$ and $\nabla_\beta FP(c,\beta)$ also play a key role in establishing a uniform influence function representation for the sample ROC curve. Recall that $B^*(r)$ denote the open ball with radius $r>0$ centered on $\beta^*$.

assumptionLet $\mathcal{C}=[a_1,b_1]$. $\nabla_\betaTP(c,\beta)$ exits and is continuous over $\mathcal{C}\times B^*(r)$ for some $r>0$ with $\sup_{(c,\beta)\in\mathcal{C}\times B^*(r)}\|\nabla_\betaTP(c,\beta)\|\le M$ for some $M>0$. The same applies to $\nabla_\betaFP(c,\beta)$ with $\mathcal{C}=[a_0,b_0]$.

Functional limit results

Letting $c^*_t={FP}_{\beta^*}^{-1}(t)\in[a_0,b_0]$ and $\hat c_t= \widehat{FP}_{\hat\beta}^{-1}(t)\in[a_0,b_0]$, we start from a decomposition of $\sqrt{n}[\widehat{TP}(\hat c_t,\hat\beta)-TP(c^*_t,\beta^*)]$ similar to ((ref)). There are two added layers of difficulty. First, functional results require uniform approximations to these terms as $t$ varies over the $[0,1]$ interval. Second, instead of being fixed, the cutoff is now estimated for any given value of $t$. The sampling variation in $\hat c_t$ contributes another non-trivial term to the asymptotic distribution.

We express the centered and scaled empirical ROC curve as the sum of three terms:

multline[multline omitted — 322 chars of source]

The first term in equation ((ref)) can be expanded similarly to the second equality in ((ref)):

equation[equation omitted — 162 chars of source]

where $\sup_{t\in[0,1]}|R_{1n}(t)|=o_p(1)$. The uniform convergence of the remainder term is a consequence of the stochastic equicontinuity of the process $(c,\beta)\mapsto \sqrt{n}(\widehat{TP}(c,\beta)-TP(c,\beta))$, stated directly in Assumption (ref), coupled with the fact that $\hat\beta\rightarrow_p\beta^*$ (Assumption (ref)) and $\sup_{t\in[0,1]}|\hat c_t-c^*_t|\rightarrow_p 0$ (Lemma (ref) in Appendix A). This last result makes use of Assumption (ref)(i), which requires that the density $f_0^*$ be bounded away from zero on its compact support.

The second term in equation ((ref)) is due to the estimation of $\beta$ and is again handled by a standard mean value expansion:

equation[equation omitted — 151 chars of source]

where $\sup_{t\in[0,1]}|R_{2n}(t)|=o_p(1)$. The uniformity of the approximation is ensured by Assumption (ref), which implies that $\nabla_\betaTP(c,\beta^*)$ is uniformly continuous.

Finally, the third term in ((ref)) arises because of the need to estimate the cutoff value associated with a given false positive rate $t$; it therefore does not arise in the fixed-cutoff setting. Starting with a mean value expansion of $TP(\hat c_t,\beta^*)$ around $c^*_t$, one can write

eqnarray[eqnarray omitted — 238 chars of source]

The remainder term $R_{3n}(t)$ converges in probability to zero uniformly over the interval \[ \big\{t: c^*_t\in[c_{0,L}, c_{0,U}]\big\}=\big[FP(c_{0,U}, \beta^*), FP(c_{0,L},\beta^*)\big], \] where $c_{0,L}$ and $c_{0,U}$ are specified in Assumption (ref)(iii). The asymptotic distribution of the process $t\mapsto \sqrt{n}[\widehat{FP}_{\hat\beta}^{-1}(t)- FP^{-1}_{\beta^*}(t)]$ can be analyzed in two steps: First, we establish an asymptotically linear representation for the “base process” $c\mapsto \sqrt{n}[\widehat{FP}(c,\hat\beta)-FP(c,\beta^*)]$ that holds uniformly in $c$ (and implies a mean zero Gaussian limit process). Second, we apply the functional delta method under the inverse functional $\phi(F)=F^{-1}$ to characterize the contribution of the term ((ref)) to the asymptotic distribution of the empirical ROC curve.

Lemma (ref) summarizes and completes the development of the approximations presented in equations ((ref)), ((ref)) and ((ref)).

lemmaSuppose that Assumptions (ref), (ref), (ref), (ref) and (ref) are satisfied. Then: (i) $\sup_{t\in [0,1]}R_{1n}(t)=o_p(1)$; (ii) $\sup_{t\in [0,1]}R_{2n}(t)=o_p(1)$; (iii) $\sup_{t\in T} R_{3n}(t)=o_p(1)$, where $T=\big[FP(c_{0,U}, \beta^*), FP(c_{0,L},\beta^*)\big]$; (iv) $\widehat{FP}(c,\hat\beta)$ admits asymptotically linear representation that holds uniformly in $c$: \begin{equation} \sqrt{n}\big(\widehat{FP}(c,\hat\beta)-FP(c,\beta^*)\big)=\frac{1}{\sqrt{n}}\sum_{i=1}^n \psi_{FP}(Y_i,X_i,c,\beta^*)+R_{4n}(t), \end{equation} where $\sup_{c_\in[a_0,b_0]}R_{4n}(c)=o_p(1)$; (v) and, by the functional delta method, \begin{equation} \sqrt{n}[\widehat{FP}_{\hat\beta}^{-1}(t)- FP^{-1}_{\beta^*}(t)]= -\frac{1}{f_0^*(c^*_t)}\sqrt{n}[\widehat{FP}(c^*_t, \hat\beta)- FP(c^*_t,\beta^*)]+R_{5n}(t), \end{equation} where $\sup_{t\in(0,1)}|R_{5n}(t)|=o_p(1)$.

\paragraph{Remarks}

enumerate• The proof of Lemma (ref) is provided in Appendix B; it simply adds some technical details to the arguments outlined in the main text. • The fact that $f^*_0(c)$ is bounded away from zero (Assumption (ref)(i)) plays a critical role in ensuring that the remainder term $R_{5n}(t)$ associated with the delta method converges to zero uniformly over the entire $[0,1]$ interval.

Combining equations ((ref)) through ((ref)) with the influence function representations of $\sqrt{n}[\widehat{TP}(c,\beta^*)-TP(c,\beta^*)]$ and $\sqrt{n}(\hat\beta-\beta^*)$ yields the following proposition, which is the central result of the paper.

propositionSuppose that Assumptions (ref), (ref), (ref), (ref) and (ref) are satisfied. Define \begin{align} &\psi_{R}(y,x,t,\beta^*)=\psi_{TP}(y,x,c^*_t,\beta^*)- \frac{f_{1}^*(c^*_t)}{f_{0}^*(c^*_t)}\psi_{FP}(y,x,c^*_t,\beta^*),\nonumber \end{align} where $c^*_t=FP^{-1}_{\beta^*}(t)$. Then: (i) The empirical ROC curve admits an asymptotically linear representation that holds uniformly over $T=\big[FP(c_{0,U}, \beta^*), FP(c_{0,L},\beta^*)\big]$: \begin{align} \sup_{t\in T}\Big|\sqrt{n}&\big(\widehat{R}(t,\hat\beta)-R(t,\beta^*)\big)-\frac{1}{\sqrt{n}}\sum_{i=1}^n \psi_{R}(Y_i,X_i,t,\beta^*)\Big|=o_p(1), \end{align} where $c_{0,L}$ and $c_{0,U}$ are chosen in accordance with Assumption (ref)(iii), i.e., $f_1^*/f_0^*$ is continuous and bounded away from zero on $[c_{0,L},c_{0,U}]$ . (ii) The process $t\mapsto \frac{1}{\sqrt{n}}\sum_{i=1}^n \psi_{R}(Y_i,X_i,t,\beta^*)$ is stochastically equicontinuous over $T$. (iii) Therefore, \begin{align*} \sqrt{n}(\widehat{R}_n(t,\hat\beta)-{R}(t,\beta^*))\Rightarrow \Psi_{h_{R}}(t) in the space $L^\infty(T)$, \end{align*} where “$\Rightarrow$” denotes weak convergence, $L^\infty(T)$ is the space of bounded functions over $T$, and $\Psi_{h_R}(\cdot)$ is a zero mean Gaussian process defined on $T$ with covariance kernel $h_{R}(t_1,t_2)=E[\psi_{R}(Y,X,t_1,\beta^*)\psi_{R}(Y,X,t_2,\beta^*)]$.

\paragraph{Remarks}

enumerate• The precise notion of weak convergence employed in part (iii) is given by Definition 1.3.3 of van der Vaart and Wellner (1996). • Given the arguments leading up to Proposition (ref), the proof of part (i) is practically complete (technically, it still requires showing that the influence function representation of $\sqrt{n}[\widehat{TP}(c,\beta^*)-TP(c,\beta^*)]$ holds uniformly in $c$, but this is essentially covered by Lemma (ref)(iv)). The proof of Part (ii) relies on Assumptions (ref), (ref) and (ref). Part (iii) follows immediately from parts (i) and (ii). Details are presented in Appendix B.

Simulating the asymptotic distribution of the ROC curve

In order to employ Proposition (ref) for statistical inference, we need a method to approximate $\Psi_{h_{R}}(t)$, the distributional limit of the process $\sqrt{n}(\widehat{R}_n(t,\beta^*)-{R}(t,\beta^*))$. To this end, offer two methods: the weighted bootstrap as in Ma and Korosok (2005) and the multiplier bootstrap as in Hsu (2016).

We first present the discussion of the weighted bootstrap. Define $\widehat{TP}^w(c,\hat{\beta}^w)$ and $\widehat{FP}^w(c,\hat{\beta}^w)$ precisely as in Section (ref) and let $\hat c^w_t= \big(\widehat{FP}^w_{\hat{\beta}^w}\big)^{-1}(t)$. We can then construct the weighted ROC curve and its estimated limit process as $\widehat{R}^w_n(t,\hat\beta^w)=\widehat{TP}^w(\hat{c}^w_t,\hat{\beta}^w)$ and $\widehat{\Psi}^w_{R,n}(t)=\sqrt{n}(\widehat{R}^w_n(t,\hat\beta^w)- \widehat{R}_n(t,\hat\beta))$.

propositionSuppose that (ref)-(ref), and (ref)-(ref) are satisfied. Then, conditional on the sample path of the data, $\widehat{\Psi}^w_{R,n}(\cdot)\Rightarrow \Psi_{h_{R,2}}(\cdot)$ in the space $L^\infty(T)$ with probability approaching one.

Under the conditions of Proposition (ref), one can apply the arguments in Theorem 2 of Ma and Kosorok (2005) to show that $\widehat{\Psi}^w_{R,n}(t)$ also approximates the distribution of $\Psi_{h_{R,2}}(t)$ in the sense of Proposition (ref). That is, conditional on the sample path of the data, $\widehat{\Psi}^w_{R,n}(t)\Rightarrow \Psi_{h_{R,2}}(t)$ with probability approaching one.

We now turn to the discussion of the multiplier bootstrap method that is based on the conditional multiplier central limit theorem (see, e.g., van der Vaart and Wellner 1996, Section 2.9). The method requires consistent estimation of the components of the influence function $\psi_R$, uniformly in $t$. However, this estimation needs to be performed only once, over the original data set, given that the method does not rely on successive resampling and reestimation.

Let $\widehat{\psi}_\beta(y,x, \hat{\beta})$ denote the estimated influence function of $\hat\beta$, where we replace any unknown parameters or functions within $\psi_\beta$ with consistent estimators (note that this function does not depend on $t$). We make the general assumption that the asymptotic variance-covariance matrix of $\hat\beta$ is consistently estimable using $\widehat{\psi}_\beta$ and the sample analog principle:

assumptionLet $V=E[{\psi}_\beta(Y,X, \beta^*){\psi}_\beta(Y,X, \beta^*)']$. Then: \begin{align*} \widehat{V}_n=n^{-1}\sum_{i=1}^n \widehat{\psi}_\beta(Y_i,X_i, \hat{\beta}_n)\widehat{\psi}_\beta(Y_i,X_i, \hat{\beta}_n)'\stackrel{p}{\rightarrow} V. \end{align*}

Let $\mathcal{C}=[c_{0,L},c_{0,U}]$. We further assume that there exist uniformly consistent estimators for $\nabla_\beta FP(c,\beta^*)$, $\nabla_\beta TP(c,\beta^*)$ and $f_{1}^*(c)/f_{0}^*(c)$ on $\mathcal{C}$. Here we state the existence of these estimators as a high level assumption and provide concrete implementations and additional assumptions in Appendix C.

assumptionThe estimators $\nabla_\beta \widehat{FP}(c,\hat{\beta}_n)$, $\nabla_\beta \widehat{TP}(c,\hat{\beta}_n)$, and $\hat{f}_{1}(c)/\hat{f}_{0}(c)$ are Lipschitz continuous in $c$ on $\mathcal{C}$ which is compact and satisfy \begin{align*} &\sup_{c\in\mathcal{C}}\|\nabla_\beta\widehat{FP}(c,\hat{\beta}_n)-\nabla_\beta FP(c,\beta^*)\|=o_p(1),\\ &\sup_{c\in\mathcal{C}}\|\nabla_\beta\widehat{TP}(c,\hat{\beta}_n)-\nabla_\beta TP(c,\beta^*)\|=o_p(1),\\ &\sup_{c\in\mathcal{C}}\left|\frac{\hat{f}_{1}(c)}{\hat{f}_{0}(c)}-\frac{f_{1}(c)}{f_{0}(c)}\right|=o_p(1). \end{align*} In addition, the estimator $\hat{c}_t$ is uniformly consistent for $c_t$ for $t\in T$.

We now present the multiplier bootstrap. Let $U_1,\ldots, U_n$ be i.i.d.\ random variables independent of the data with moments $E[U]=0$, $E[U^2]=1$, and $E|U|^{2+\delta_u}<\infty$ for some $\delta_u>0$. For $t\in[0,1]$, we define the simulated stochastic process $\widehat{\Psi}^u_{R,n}(t)$ as

align[align omitted — 136 chars of source]

where

align*[align* omitted — 556 chars of source]

The next result shows that the distribution of the simulated process $\widehat{\Psi}^u_{R,n}(t)$ approximates that of the true limiting process $\Psi_{h_{R,2}}(t)$ in large samples.

propositionSuppose that Assumptions (ref)-(ref) and (ref)-(ref) are satisfied. Then, conditional on the sample path of the data, $\widehat{\Psi}^u_{R,n}(\cdot)\Rightarrow \Psi_{h_{R,2}}(\cdot)$ in the space $L^\infty(T)$ with probability approaching one.

The weighted bootstrap has the advantage that it does not require explicit estimation of this function, but it is computationally somewhat more costly. On the other hand, the proof of weighted bootstrap is less involved because the multiplier method relies heavily on Assumption (ref), i.e., the availability of uniformly consistent estimators for the components of $\psi_R$. To obtain estimators satisfy Assumption (ref) additional assumptions are needed and in Appendix C, we provide estimators and additional assumptions so that Assumption (ref) can be satisfied and we can apply multiplier method.

Applications to various inference problems

In this section, we provide some examples that we can apply the results in Section (ref).

Uniform confidence bands

Let $\widehat{\sigma}^2_t$ denote a uniform consistent estimator for ${\sigma}^2_t$, the asymptotic variance of $\sqrt{n}(\widehat{R}_n(t,\hat\beta)-{R}(t,\beta^*))$ for $t\in T$. Later we will provide two estimators based on weighted bootstrap method and analytic results. Let $\widehat{\sigma}_{t,\epsilon}=\max\{\widehat{\sigma}_t,\epsilon\}$ in which $\epsilon>0$ is a fixed and small number. We are interested in a standardized version of confidence bands and by truncating $\widehat{\sigma}_t$ by $\epsilon$, we can make sure that we will not divide something close to zero when $t$ is close to 0 or 1.

For a nominal significance level $\alpha$ and for $\tau_\ell,\tau_u\in T$ with $\tau_\ell\leq \tau_u$, let $\widehat{C}^\text{1-sided}_\alpha$ and $\widehat{C}^\text{2-sided}_\alpha$ respectively denote the one- and two-sided critical values that satisfy

align[align omitted — 446 chars of source]

Here, $\widehat{C}^\text{1-sided}_\alpha$ and $\widehat{C}^\text{2-sided}_\alpha$ are, respectively, the $(1-\alpha)$th quantile of $\sup_{t\in[\tau_\ell,\tau_u]}\widehat{\Psi}^w_{R,n}(t) \big/\widehat{\sigma}_{t,\epsilon}$ and $(1-\alpha)$th quantile of $\sup_{t\in[\tau_\ell,\tau_u]}\big|\widehat{\Psi}^w_{R,n}(t) \big/\widehat{\sigma}_{t,\epsilon}\big|$. Note that one can replace $\widehat{\Psi}^w_{R,n}(t) $ with $\widehat{\Psi}^u_{R,n}(t) $ to construct $\widehat{C}^\text{1-sided}_\alpha$ and $\widehat{C}^\text{2-sided}_\alpha$ as well.

Once the critical values are constructed, we can also obtain one- and two-sided uniform confidence bands for ${R}(t,\beta^*)$ over $[\tau_\ell,\tau_u]$. Specifically, the one-sided $(1-\alpha)$ uniform confidence band is given by

equation[equation omitted — 196 chars of source]

and the two-sided $(1-\alpha)$ uniform confidence band is

equation[equation omitted — 296 chars of source]

{\bf Implementation of Uniform Confidence Bands}

We now provide a step-by-step implementation for constructing uniform confidence bands.

enumerate• Obtain $\widehat{R}_n(t,\hat\beta)$ from Section (ref) and $\widehat{\sigma}_{t,\epsilon}$ from Section (ref) with $t\in\{\tau_\ell,\tau_\ell+0.01,\dotsc,\tau_u\}$. • Draw i.i.d.\ pseudo random variables $\{W_1,\dotsc,W_n\}$ where $W_i$'s are normal distributions with mean and variance equal to one $B$ times for, say, $B=1000$. For each repetition $b=1,\dotsc,B$, calculate the simulated process $\widehat{\Psi}^w_{R,n}(t)$ according to ((ref)). • For the one-sided case, store the maximum value of ${\widehat{\Psi}^w_{R,n}(t) }\big/{\widehat{\sigma}_{t,\epsilon}}$ over the grid of $t$ values set up in Step 1; that is, let $M_b=\max_{t\in\{\tau_\ell,\tau_\ell+0.01,\dotsc,\tau_u\}}{\widehat{\Psi}^w_{R,n}(t) }\big/{\widehat{\sigma}_{t,\epsilon}}$ for $b=1,\dotsc,B$. • Rank the $M_b$ values in an ascending order so that $M_{(1)}\leq\dotsc\leq M_{(B)}$. Next, define $M_{(\lfloor(1-\alpha)B\rfloor)}$ as the critical value $\widehat{C}^\text{1-sided}_\alpha$, where $\lfloor a \rfloor$ is the floor function returning the largest integer not greater than $a$. The one-sided $(1-\alpha)$ uniform confidence bands for $\{{R}(t,\beta^*):t\in[\tau_\ell,\tau_u]\}$ are given by (ref). • For the two-sided case, simply replace ${\widehat{\Psi}^w_{R,n}(t) }\big/{\widehat{\sigma}_{t,\epsilon}}$ in Step 3 with $\big|{\widehat{\Psi}^w_{R,n}(t) }\big|\big/{\widehat{\sigma}_{t,\epsilon}}$ and repeat Step 4 for the critical value $\widehat{C}^\text{2-sided}_\alpha$. The two-sided $(1-\alpha)$ uniform confidence band for $\{{R}(t,\beta^*):t\in[\tau_\ell,\tau_u]\}$ is given by (ref).

{\bf Uniformly consistent estimators for ${\sigma}^2_t$}\\ We consider two estimators here. First estimator is based on weighted bootstrap that is similar to Remark 2 after Proposition (ref). Let $\widehat{\Psi}^w_{R,n}(t)$ denote the ROC estimate from the $w$th bootstrap cycle, $w=1,\ldots, \mathcal{W}$. Then one can estimate ${\sigma}^2_t$ by

align*[align* omitted — 383 chars of source]

We have that conditional on sample path with probability approaching one,

align*[align* omitted — 125 chars of source]

where $\stackrel{p}{\rightarrow}_w$ denotes probability limit under the law of the $W_i$'s. It follows that uniformly over $t\in[t_\ell,t_u]$,

align*[align* omitted — 111 chars of source]

The second estimator is based on analytic results. Recall that $\widehat{\psi}_R(Y_i,X_i,t,\hat\beta)$ is the estimated influence function for $\widehat{R}_n(t,\hat\beta)$ used in the multiplier bootstrap method. A uniformly consistent estimator for $ {\sigma}^2_t$ is given by

align*[align* omitted — 98 chars of source]

and this is shown in the proof of (ref).

ROC dominance test

For two predictive index models $G_1(X,\beta_1)$ and $G_2(X,\beta_2)$, we may want to test whether $G_1$ has strictly better predictive power than $G_2$ in the sense that the ROC curve associated with $G_1$ dominates the ROC curve associated with $G_2$. What domination means is that for any given false positive rate $G_1$ delivers a higher true positive rate, i.e., the ROC curve for $G_1$ always lies above the ROC curve for $G_2$. Any decision maker, regardless of their loss function and their optimal cutoff, would then prefer model $G_1$ over $G_2$.

Let $R_1(t,\beta_1^*)$ and $R_2(t,\beta_2^*)$ denote the ROC curves associated with $G_1$ and $G_2$, respectively. The hypotheses that $R_1(t,\beta_1^*)$ dominates $R_2(t,\beta_2^*)$ can be formally stated as

align[align omitted — 205 chars of source]

Our test for ROC dominance is similar to the test for first order stochastic dominance in Barrett and Donald (2003) and Donald and Hsu (2016) except that we need to consider the estimation effect of $\hat{\beta}$ as in Linton, Massoumi and Whang (2005) and Linton, Song and Whang (2010).

Let $\widehat{R}_{j,n}(t,\hat{\beta}_j)$ be the estimators for $R_j(t,\beta_j^*)$ for $j=1,2$. Define $\psi_{j,R}(Y_i,X_i,t,\beta_1^*)$ for $j=1$ and 2 as above. Let $\widehat{\sigma}^2_{RD}(t)$ denote a uniform consistent estimator for ${\sigma}^2_{RD}(t)$, the asymptotic variance of $\sqrt{n}(\widehat{R}_{2,n}(t,\hat{\beta}_2)-\widehat{R}_{1,n}(t,\hat{\beta}_1)-{R}_{2}(t,{\beta}^*_2)-{R}_{1}(t,{\beta}^*_1))$. Let $\widehat{\sigma}_{RD,\epsilon}(t)=\max\{\widehat{\sigma}_{RD}(t),\epsilon\}$ in which $\epsilon>0$ is a fixed and small number. Uniform consistent estimator $\widehat{\sigma}^2_{RD}(t)$ can be obtained similar to $\widehat{\sigma}^2_t$ in Section (ref), so we omit the details.

We define the test statistic as $\widehat{S}_n=\sqrt{n}\sup_{t\in[0,1]}(\widehat{R}_{2,n}(t,\hat{\beta}_2)-\widehat{R}_{1,n}(t,\hat{\beta}_1))/\widehat{\sigma}_{RD,\epsilon}(t)$. Define the weighted bootstrap process $\Psi^w_{RD,n}(t)$ as $\widehat{\Psi}^w_{R,n}(t)=\sqrt{n}\big(\widehat{R}^w_{2,n}(t,\hat\beta^w_2)-\widehat{R}^w_{1,n}(t,\hat\beta^w_1) -(\widehat{R}_{2,n}(t,\hat{\beta}_2)-\widehat{R}_{1,n}(t,\hat{\beta}_1))\big) $ and define the multiplier bootstrap process $\Psi^u_{RD,n}(t)$ as

align*[align* omitted — 173 chars of source]

Under the least favorable configuration, we define the weighted bootstrap critical value as

align[align omitted — 175 chars of source]

with significance level $\alpha$. The decision rule is

align[align omitted — 87 chars of source]

Then one can use $\widehat{\Psi}^u_{RD,n}(t)$ to construct critical value $\hat{c}_n$ as well.

Similar to the stochastic dominance test literature, we can show that under the null hypothesis the asymptotic size of a test with decision rule defined in ((ref)) is less than or equal to $\alpha$. That is, we can control the asymptotic size of our ROC dominance test well. Also, under the fixed alternative, we have the test statistic converging to positive infinity and the critical value converging to a finite number, so the test is consistent. Our test is based on least favorable configuration, so it is conservative in that the asymptotic size is strictly smaller than $\alpha$ unless $R_2(t,\beta_2^*)= R_1(t,\beta_1^*)$ for all $t\in[0,1]$. One can improve the power of our test by using the recentering method in Hansen (2005), Donald and Hsu (201) which is similar to the generalized moment selection method in Andrews and Soreas (2010), and Andrews and Shi (2013), and the contact set approach in Linton, Song and Whang (2010). In this paper, we do not adopt this approach but the extension is straightforward.

Comparing AUCs

Recall that AUC is defined as the integral of ROC curve from 0 to one. Following Section (ref), let $R_1(t,\beta_1^*)$ and $R_2(t,\beta_2^*)$ denote the ROC curves for two predictive index models $G_1(X,\beta_1)$ and $G_2(X,\beta_2)$. Let $AUC_j=\int_{0}^1 R_j(t,\beta_j^*) dt$ and its estimator be $\widehat{AUC}_j=\int_{0}^1 \widehat{R}_{j,n}(t,\hat{\beta}_j)dt $ for $j=1$ and 2. Then it is true that

align*[align* omitted — 179 chars of source]

and

align*[align* omitted — 207 chars of source]

To make inference, one can use a weighted bootstrap method to approximate the limiting distribution of $N[0, \mathcal{V}_{a}]$ or one can estimate $ \mathcal{V}_{a}$ analytically by

align*[align* omitted — 183 chars of source]

For brevity, we omit the details here.

{\bf Remark}\\ Suppose that $G(X,\beta)$ is a correct specification for the propensity score function, $P(Y=1|X)$, in that there exists $\beta^*$ such that $G(X,\beta^*)=P(Y=1|X)$ a.s. Then the estimation effect of $\beta^*$ on the distribution of AUC is negligible. It is then true that when two predictive predictive index models, $G_1(X,\beta_1)$ and $G_2(X,\beta_2)$, are both correctly specified for $P(Y=1|X)$, we have $\mathcal{V}_{a}=0$, i.e., the limiting distribution of $\sqrt{n}(\widehat{AUC}_2-\widehat{AUC}_1)$ is degenerate.

Simulations: the relevance of the estimation effect

We now present a small Monte Carlo simulation to illustrate the theoretical discussion of the estimation effect and the pointwise asymptotic results in Section (ref). The data generating process (DGP) is specified as follows. Let $X=(X_1, X_2, X_3)'$ be a $3\times 1$ vector of predictors and $\tilde X=(1,X')'$. The components of $X$ are either independent $N(0,1)$ or unif$[-.5,1.5]$ variables. The outcomes $Y$ are generated according to the conditional probability function \[ p(X)=G(X,\beta^\circ)=G(\tilde X'\beta^\circ)\text{ with } \beta^\circ=(0, 0.5, 0.25, 1)'\text{ and }G\in\{\text{logit},\text{cauchit}\}. \] In the majority of the exercises we use a logistic link in the DGP (so that the logit first stage is correctly specified), but we also conduct some simulations with a cauchit link (so that the logit first stage is mildly misspecified).

Given a sample of observations and a cutoff $c$, we construct nominally 90% confidence intervals for TP$(c,\beta^\circ)$ and TP$(c,\beta^\circ)-FP(c,\beta^\circ)$ in three different ways: (i) using the true predictive index $p(X)=G(\tilde X'\beta^\circ)$ with the conventional limit distributions ((ref)) and ((ref)); (ii) using the estimated predictive index $\Lambda(\tilde X'\hat\beta)$ with the conventional limit distributions (so that the estimation effect is ignored); (iii) using the estimated predictive index $\Lambda(\tilde X'\hat\beta)$ with the corrected limit distribution ((ref)). We simulate 10,000 samples and compute the actual coverage probability of these intervals.

table[table omitted — 3,706 chars of source]

Table (ref) reports the results from this exercise for $c\in\{0.2, 0.33, 0.5, 0.67, 0.8\}$ and $n\in\{200, 500, 5000\}$. The first message is that failing to account for the pre-estimation effect can cause substantial distortions in the coverage probability of the conventional CIs. In panels (A) through (D) the estimation effect can be seen by comparing the columns titled “Conventional $G(\tilde X'\beta^\circ)$” and “Conventional $\Lambda(\tilde X'\hat\beta)$.” In the former case there is no estimation effect and any deviation from the nominal confidence level of 90% is a small sample phenomenon.\footnote{For example, for $c=0.2$ the value of TP$(c,\beta^\circ)$ is close to the upper bound 1, and the coverage probability of the fixed-$\beta$ CI is only 80% for $n=200$.} Over the various cases, the estimation effect ranges from essentially zero to as large as a 30 to 40 percentage point difference in coverage probability. In panel (E), the comparison between the same two columns includes the estimation effect as well as some “bias” due to the fact that the first stage logit regression is misspecified.

The theory presented in Section (ref) gives insight into why the estimation effect is negligible in some cases. In particular, consider the parameter $TP-FP$ on panels (A) through (C) with $c=0.5$. As the predictors are independent standard normal variables and there is no constant in the DGP, the symmetry of the logistic cdf gives $\pi=\mathbb{P}(Y=1)=0.5$. Therefore, when $c=0.5$, $TP-FP$ is a scalar multiple of $(1-c)\piTP-c(1-\pi)FP$. As explained in footnote (ref), inference about this particular linear combination is not impacted by the pre-estimation effect. This is clearly reflected in the estimation results. By contrast, in panel (D) the predictor distribution is not symmetric around zero, so $\pi\neq 0.5$, and the estimation effect is indeed present for $TP-FP$ even when $c=0.5$.

The second main message is that the proposed analytical correction works well in virtually all the cases considered here. This includes panel (A), where the sample size is small, and panel (E), where the first stage logit model is misspecified. Not surprisingly, under misspecification the corrected CI can also fall somewhat short of the 90% confidence level, but it still represents a large improvement over conventional inference.

Conclusions

We provided both pointwise and uniform asymptotic results that describe the distribution of an empirical ROC curve based on a pre-estimated index. The core theory is complete. Ongoing work consists of: (i) developing appropriate test procedures when the first stage models are nested and the ROC influence functions are the same under the null; (ii) additional simulations that illustrate the small sample performance of the uniform asymptotic results, the practical use of the tests, and the power gains afforded by in-sample inference.

apabibAbadie, A.\ and G.W.\ Imbens (2016): “Matching on the Estimated Propensity Score". Econometrica, 84: 781-807. Andrews, D.W.K.\ (1994): “Empirical Process Methods in Econometrics,” in Handbook of Econometrics, vol.\ IV, eds.\ R.F.\ Engle and D.L.\ McFadden, Elsevier. Andrews, D.\ W.\ K.\ and G.\ Soares (2010): “Inference for Parameters Defined by Moment Inequalities Using Generalized Moment Selection". Econometrica, 78: 119-157. Andrews, D.\ W.\ K.\ and X.\ Shi (2013): “Inference Based on Conditional Moment Inequalities". Econometrica, 81: 609-666. Anjali D.N.\ and P. Bossaerts (2014): “Risk and Reward Preferences under Time Pressure”. Review of Finance, 18: 999-1022. Bamber, D.\ (1975): “The Area above the Ordinal Dominance Graph and the Area below the Receiver Operating Characteristic Graph”. Journal of Mathematical Psychology 12: 387-415. Barrett, G.F.\ and S.G.\ Donald (2003): “Consistent Tests for Stochastic Dominance”. Econometrica, 71: 71-104. Bazzi, S., R.A.\ Blair, C.\ Blattman, O.\ Dube, M.\ Gudgeon and R.\ Peck (2021): “The Promise and Pitfalls of Conflict Prediction: Evidence from Colombia and Indonesia”. \textit{The Review of Economics and Statistics}, forthcoming. Berge, T.J. and O.\ Jorda (2011): “Evaluating the Classification of Economic Activity into Recessions and Expansions”. \textit{American Economic Journal: Macroeconomics} 3: 246-247. Bonfim, D., G. Nogueira and S.\ Ongena (2021): “ `Sorry, We're Closed' Bank Branch Closures, Loan Pricing, and Information Asymmetries”. \textit{Review of Finance}, 25: 1211-1259. Clark, T.E.\ and M.W.\ McCracken (2012): “In-sample Tests of Predictive Ability: A New Approach”. \textit{Journal of Econometrics} 170: 1-14. DeLong, E.R., D.M. DeLong and D.L. Clarke-Pearson (1988): “Comparing Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach”. \textit{Biometrics} 44: 837-845. Demler, O.V., M.J.\ Pencina and R.B.\ D'Agostino, Sr. (2012): “Misuse of DeLong Test to Compare AUCs for Nested Models”. \textit{Statistics in Medicine} 31: 2577-2587. Egan, J.P.\ (1975): \emph{Signal Detection Theory and ROC Analysis}. Academic Press: New York. Donald, S.G.\ and Y.-C.\ Hsu (2014): “Estimation and Inference for Distribution Functions and Quantile Functions in Treatment Effect Models". \textit{Journal of Econometrics}, 178: 383-397. Donald, S.G.\ and Y.-C.\ Hsu (2016): “Improving the Power of Tests of Stochastic Dominance". \textit{Econometric Reviews}, 35: 553-585. Donald, S.G.\ , Y.-C.\ Hsu and G.F.\ Barrett (2012): “Incorporating Covariates in the Measurement of Welfare and Inequality: Methods and Applications". \textit{Econometrics Journal}, 15: C1-C30. Elliott, G.\ and R.P.\ Lieli (2013): “Predicting Binary Outcomes”. \textit{Journal of Econometrics}, 174: 15-26. Hansen, P.\ R.\ (2005): “A Test for Superior Predictive Ability". \textit{Journal of Business and Economic Statistics}, 23: 365--380. Hsieh, F.\ and Turnbull, B.W.\ (1996): “Nonparametric and Semiparametric Estimation of the Receiver Operating Characteristic Curve”. \textit{The Annals of Statistics}, 24: 25-40. Inoue, A.\ and Kilian, L.\ (2004): “In-sample or out-of-sample tests of predictability? Which one should we use?”. \textit{Econometric Reviews}, 23: 371-402. Kleinberg, J., H.\ Lakkaraju, J.\ Leskovec, J.\ Ludwig and S.\ Mullainathan (2018): “Human Decisions and Machine Predictions”. \textit{The Quarterly Journal of Economics}, 133: 237–293. Lahiri, K.\ and L.\ Yang (2018): “Confidence Bands for ROC Curves With Serially Dependent Data”. \textit{Journal of Business and Economic Statistics}, 36: 115-130. Lahiri, K. and J.G. Wang (2013): “Evaluating Probability Forecasts for GDP Declines Using Alternative Methodologies”. \textit{International Journal of Forecasting}, 29: 175-190. Lieli, R.P.\ and Y-C.\ Hsu (2019): “Using the Estimated AUC to Test the Adequacy of Binary Predictors”. \textit{Journal of Nonparametric Statistics}, 31: 100-130. Lieli, R.P.\ and A.\ Nieto-Barthaburu (2010): “Optimal Binary Prediction for Group Decision Making”. \textit{Journal of Business and Economic Statistics}, 28: 308-319. Linton, O., E.\ Maasoumi and Y.-J.\ Whang (2005): “Consistent Testing for Stochastic Dominance under General Sampling Schemes". \textit{The Review of Economic Studies}, 72: 735-765. Linton, O., K.\ Song and Y.-J.\ Whang (2010): “An Improved Bootstrap Test of Stochastic Dominance". \textit{Journal of Econometrics}, 154: 186-202. Ma, S.\ and M.R.\ Kosorok (2005): “Robust Semiparametric M-estimation and the Weighted Bootstrap”. \textit{Journal of Multivariate Analysis}, 96: 190-270. McCracken, M.W., J.T.\ McGillicuddy and M.T.\ Owyang (2021): “Binary Conditional Forecasts,” \textit{Journal of Business and Economic Statistics}, forthcoming. Pagan, A.\ (1984): “Econometric Issues in the Analysis of Regressions with Generated Regressors”. \textit{International Economic Review}, 25: 221–247. Pepe, M.S.\ (2003): \emph{The Statistical Evaluation of Medical Tests for Classification and Prediction.} Oxford University Press: Oxford. Pollard, D.\ (1990): \emph{Empirical Processes: Theory and Application.} CBMS Conference Series in Probability and Statistics, Vol.\ 2. Hayward, CA: Institute of Mathematical Statistics. Schularik, M. and A.M. Taylor (2012): “Credit Booms Gone Bust: Monetary Policy, Leverage Cycles, and Financial Crises, 1870-2008”. \textit{American Economic Review} 102: 1029-1061. Van der Vaart, A.\ W.\ and J.\ A.\ Wellner (1996): \emph{Weak Convergence and Empirical Processes: With Application to Statistics.} New York: Springer-Verlag. Wooldridge, J.M.\ (2002): \emph{Econometric Analysis of Cross Section and Panel Data.} The MIT Press: Cambridge.