Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,979 characters · 24 sections · 74 citation commands
Asymptotic Inference for Rank Correlations
Keywords: Kendall's tau; Spearman's rho; Goodman-Kruskal's gamma; Confidence Intervals; Statistical Tests \newline
Correlation coefficients are fundamental tools for data analysis, quantifying the mutual dependence between two random variables. The classical and most popular correlation coefficients are Pearson correlation pearson1896, Kendall's $\tau$ kendall1938 and Spearman's $\rho$ spearman1904, which indicate the direction of dependence by their sign, and the strength of dependence by their closeness to 1 in absolute value. Despite being the most widely used of the three coefficients, Pearson correlation, which we denote by $r$, has well-known shortcomings. It lacks fundamental theoretical properties such as invariance to monotonic transformations, making it dependent on data transformations, and attainability, in the sense that it can be far away from 1 in absolute value under very strong or perfect dependence, thereby compromising its interpretability embrechts2002, fissler2023. Furthermore, it suffers from statistical problems in that its estimator is not robust to outliers and inefficient for heavy-tailed distributions. Finally, its applicability is limited: It is only defined for random variables on a metric scale, excluding ordinal variables, and only for variables with finite second moments, excluding heavy-tailed distributions. In the case of continuous random variables, the classical rank correlations Kendall's $\tau$ and Spearman's $\rho$ cure all of these shortcomings, but for discrete random variables, they also suffer from non-attainability genest2007, neslehova2007. Popular generalizations of $\rho$ and $\tau$ mitigating this problem are Kendall's $\tau_b$ kendall1948 and grade correlation, which we denote by $\rho_b$. For $\tau$, there is even a simple and appealing generalization in Goodman-Kruskal's $\gamma$ goodmanKruskal1954, which fully cures this problem pohle2025.
Given the widespread use of these rank correlations and their favorable theoretical properties, the availability of confidence intervals to quantify estimation uncertainty around point estimators, and tests, e.g., of the null hypotheses of uncorrelatedness or independence, ought to be taken for granted. Surprisingly, we have been unable to find a comprehensive treatment of asymptotic inference for rank correlations in the extant literature. In particular, often only the cases of continuous and independent and identically distributed (iid) observations are considered, excluding time series and discrete (and mixed) variables. Contrary to the case of Pearson correlation, the formulas for the asymptotic variances of rank correlations in the discrete case do not just carry over from the continuous case due to the role that ties play for rank correlations.
This lack of key results is even more surprising given that hoeffding1948 in his seminal paper on U-statistics already in the iid case established asymptotic normality and derived the corresponding limiting variances of empirical Kendall's $\tau$, $\widehat{\tau}$ (here even including the discrete case), and Spearman's $\rho$, $\widehat{\rho}$ (limiting himself to the continuous case), making use of the facts that $\widehat{\tau}$ is a U-statistic and $\widehat{\rho}$ is asymptotically equivalent to a U-statistic. Also in the iid case, genest2013estimation provide the limiting distribution of multivariate generalizations of $\widehat{\rho}$, paying particular attention to the discrete case.
In the time series case, dehling2017 derive an invariance principle for U-statistics of order 2 and in particular $\widehat{\tau}$ under weak dependence. The limiting distributions for the other rank correlations mentioned above, in particular $\widehat{\rho}$, do not even seem to be available in the continuous case from the extant literature.
Moreover, estimators for asymptotic variances of rank correlations have been neglected in the literature with the consequence that classical tools for statistical inference -- asymptotic confidence intervals and tests -- are not available. The only exceptions are for Kendall's $\tau$, where noether1967 discusses an estimator for the asymptotic variance (see also samara1988) and again dehling2017, who propose a variance estimator in the continuous time series case. Instead of consistent variance estimators, approximations to the asymptotic variances of $\widehat{\tau}$ and $\widehat{\rho}$ in the iid case seem to be in widespread use, see the proposal by fieller1957 being derived under bivariate normality, or modifications thereof bonett2000. Already borkowf2002 demonstrated them to be inaccurate.
The only settings, where classical asymptotic inference for $\tau$ and $\rho$ seems to be in widespread use, is in testing independence in the iid case. There, the limiting distributions under independence assuming continuous variables are used, where the limiting variances of $\widehat{\tau}$ and $\widehat{\rho}$ are $4/9$ and $1$, respectively.\footnote{The basic R function cor.test from the stats package RCoreTeam only provides a confidence interval for Pearson correlation. For $\tau$ and $\rho$, only those tests of independence are available.} The former result is due to kendall1938, and according to hoeffding1948, the latter is due to Student and was published by pearson1907. Thus, here, the variances are constant and do not have to be estimated. In the discrete case, however, the limiting variances under independence have not been available from the literature. hoeffding1948 conjectured that they are different from their values in the continuous case and depend on single-tie and double-tie probabilities of the two variables of interest.
We provide a comprehensive treatment of asymptotic inference for classical rank correlations, namely for Kendall's $\tau$ and Spearman's $\rho$ as well as their generalizations to the discrete case, Goodman-Kruskal's $\gamma$, $\tau_b$ and $\rho_b$. First, we establish asymptotic normality of their empirical versions, $\widehat{\tau}$, $\widehat{\rho}$, $\widehat{\gamma}$, $\widehat{\tau}_b$ and $\widehat{\rho}_b$, respectively, and derive the corresponding asymptotic variances under classical time series assumptions , which includes the important special case of iid processes. Here, we make use of asymptotic theory for U-statistics for weakly dependent stochastic processes (see dehling2006 for an overview). Secondly, we introduce consistent estimators for the asymptotic variances. Those two elements -- the asymptotic distribution and consistent variance estimators -- allow for the construction of confidence intervals and tests. Thirdly, we derive formulas for the asymptotic variances under independence between the two variables or processes of interest, which is important to execute widely-used tests of independence. In the iid case, our results generalize the aforementioned classical results from the continuous case. The variance formulas indeed depend on the double-tie probabilities and are easily implementable, providing simple and valid tests of independence also in the discrete case. In time series settings, our results generalize the results of lun2023, who establish the asymptotic distribution of $\widehat{\tau}$ and $\widehat{\rho}$ in the continuous case under independence between the two stochastic processes of interest. We also provide a multivariate central limit theorem for a vector of U-statistics, which enables the asymptotic analysis of $\widehat{\gamma}$, $\widehat{\tau}_b$ and $\widehat{\rho}_b$ as they are functions of several U-statistics. This result also demonstrates that vectors of empirical rank correlations are asymptotically multivariate normal and allows for the calculation of the corresponding covariance matrices.
In an extensive simulation study, we analyze the finite-sample performance of our proposed confidence intervals and independence tests for a variety of continuous and discrete iid as well as time series data generating processes. We illustrate the use of our inferential procedures in two real-world data applications.
The remainder of the paper is structured as follows. Section (ref) introduces fundamental concepts and the classical rank correlations $\rho$ and $\tau$ as well as their generalizations to the discrete case, while Section (ref) considers their empirical counterparts. In Section (ref), we provide the corresponding asymptotic distributions and propose consistent estimators for the asymptotic variances in the iid case. Furthermore, we discuss the limiting distributions and variance estimation under independence. Section (ref) presents generalizations of the results in Section (ref) to the time series case. Section (ref) summarizes simulation results and Section (ref) discusses empirical applications. Finally, we conclude in Section (ref), where we outline directions for future research.
Appendix (ref) contains a discussion of asymptotic theory for U-statistics under weak dependence and, in particular, the above-mentioned multivariate central limit theorem. Appendix (ref) presents all proofs, and Appendix (ref) an illustrative example involving a bivariate geometric distribution. An extensive description of the simulated DGPs from Section (ref) and a detailed interpretation of the simulation results as well as the results itself in a tabulated form are shown in Appendix (ref). We also provide code implementing our confidence intervals and tests in an R package RCoreTeam at https://github.com/jan-lukas-wermuth/RCor and replication material at https://github.com/jan-lukas-wermuth/replication_RCor.
Throughout the paper, let $(X,Y) \sim F_{X,Y}$ be a bivariate random variable with cumulative distribution function (CDF) $F_{X,Y}$. Similarly, $X\sim F_X$ denotes a univariate random variable with CDF $F_X$. We say that a random variable $X$ is continuous if its CDF $F_X$ is a continuous function. We use $(X^\prime,Y^\prime)$ and $(X^{\prime\prime},Y^{\prime\prime})$ to denote independent copies of $(X,Y)$. For a function $f$, we denote the left limit at $x$ as $f(x^{-}) = \lim_{a \rightarrow x, a < x} f(a)$. We write $\stackrel{p} {\rightarrow}$ for convergence in probability and $\stackrel{d} {\rightarrow}$ for convergence in distribution.
We define the mid-distribution function (MDF) of $X$ as $G_X(x):=(F_X(x) + F_X(x^{-}))/2$ for $x \in \mathbb{R}$ (and analogously for $Y$), see e.g.\ parzen97 for a discussion. Note that the MDF can be rewritten as
denotes the Heaviside step function using the half-maximum convention. Essentially, the MDF counts the probability mass for values smaller than $x$ fully, and the probability mass for values equal to $x$ half. It is thus closely related to the concept of a midrank, see (ref) in the next section. Furthermore, we define the bivariate MDF of $(X, Y)$ as hoeffding1948
Plugging-in a univariate or bivariate random variable itself into the respective MDF, we obtain the univariate grade hoeffding1948 or mid-distribution transform parzen2004, $G_X(X)$, and the bivariate grade $G_{X,Y}(X,Y)$, which are of fundamental importance for this paper. The univariate grade is a variant of the probability integral transform (PIT) $F_X(X)$ and better-suited for arbitrary real-valued random variables. It reduces to the PIT if $F_X$ is continuous. While the univariate grade does not follow a standard uniform distribution in the discrete case like the PIT (and itself) in the continuous case ($F_X(X) \sim U(0,1)$ if $X$ is continuous), contrary to the PIT its expectation is always $1/2$. Also, the variance is smaller than $1/12$ by a factor only depending on the double-tie probability of $X$, i.e.\ on $\mathrm{P}(X=X^\prime=X^{\prime\prime})$, also see the subsequent Lemma (ref). For the moments of the bivariate grade, there are no general results available, and its distribution depends on the distribution of $(X,Y)$. We often need its expectation in the paper, which we consequently need to estimate.
A further crucial concept for this paper is the sign function
more precisely the sign function of a difference. We collect important results about the grade, the sign difference $\mathrm{sgn}(X-X')$ and their relation, as often needed throughout the rest of the paper, in the following Lemma (ref). The first result just amounts to (ref) rewritten in terms of $\mathrm{sgn}(x-X)$, noting that $H(x) = 1/2 ( 1 + \mathrm{sgn}(x))$ holds.
The single-tie probability $\zeta(X):=\mathrm{P}(X=X^\prime)$ and the double-tie probability $\zeta_2(X):=\mathrm{P}(X=X^\prime=X^{\prime\prime})$ showing up in the variance formulas are crucial for this paper as well. These probabilities equal zero if $X$ is continuous and can be computed from the probability mass function (PMF) of $X$ in the discrete (or discrete-continuous) case:
We denote the probability of a tie in $X$ or $Y$ or both by
Throughout the rest of the paper, we assume non-degenerate random variables such that $\zeta_X$, $\zeta_2(X)$ and $\nu(X,Y)$ are strictly smaller than 1.
Let us conclude this section by referring to a running example in Appendix (ref), where we illustrate the terms, concepts, and results on the special case of a bivariate geometric distribution.
The idea behind all classical correlation coefficients is to measure if two random variables $X$ and $Y$ tend to move in the same or in opposite directions, and how strong this tendency is. Making this idea specific requires a reference point. In the case of covariance and Pearson correlation,
the reference point is the expectation $(\mathrm{E}[X],\mathrm{E}[Y])$ of $(X,Y)$.
Kendall's $\tau$ uses $(X^\prime,Y^\prime)$, an independent copy of $(X,Y)$, as the reference point and calculates the probability that $X$ and $Y$ move in the same direction relative to $X^\prime$ and $Y^\prime$ minus the probability that they move in the opposite direction, that is, it is defined as the probability of concordance minus the probability of discordance. It can be rewritten in terms of the sign function from (ref).
By part (iii) of Lemma (ref), it holds that $\tau(X,Y) = \mathrm{Cov}\big(\mathrm{sgn}(X-X^\prime),\mathrm{sgn}(Y-Y^\prime)\big)$, and in the continuous case, we even have $\tau(X,Y) = \mathrm{Cor}\big(\mathrm{sgn}(X-X^\prime),\mathrm{sgn}(Y-Y^\prime)\big)$. Kendall's $\tau$ can also be compactly written in terms of the grade by using parts (ii) and (iv) of Lemma (ref), namely $\tau(X,Y) = 4 \mathrm{E}[G_{X,Y} (X,Y)] - 1$, which generalizes its relation to the joint CDF and copula from the continuous case schweizer1981. For the special case of the bivariate geometric distribution, see Appendix (ref).
Note that the definition of $\tau(X,Y)$ uses a difference between random variables for notational simplicity. However, we do not assume that $X$ and $Y$ are necessarily quantitative random variables, but we also allow for ordinal random variables with categories, say, $s_0<\cdots<s_d$. The above difference is valid in the ordinal case if we use the numerical coding $s_k := k$ for $k=0,\ldots,d$, which simply counts how many categories $s_k$ is apart from the lowest category $s_0$. Equivalently, one could avoid differences by using (expectations of) the indicator function instead, where $X-X^\prime>0$ iff $\mathds{1}(X>X^\prime)=1$ etc., also see (ref) above on how to avoid differences. But to keep notations compact, we shall continue writing differences as in Definition (ref).
A suitable population version of Spearman's $\rho$ involves the covariance of the grades, $G_X(X)$ and $G_Y(Y)$, see e.g.\ hoeffding1948.
Note that if $X$ and $Y$ are continuous, then $\mathrm{Var}(G_X(X))=\mathrm{Var}(G_Y(Y))= 1/12$ and $\rho(X,Y)=\mathrm{Cor}(G_X(X), G_Y(Y))$. Spearman's $\rho$ carries the idea behind covariance or Pearson correlation to the rank scale, leading to a measure with much more favourable properties. For the special case of the bivariate geometric distribution, see again Appendix (ref).
Using the relation between the sign function of a difference and the grade from Lemma (ref), $\rho$ can be rewritten as
Thus, it can also be seen as a modification of $\tau$, which does not use an independent copy $(X^\prime,Y^\prime)$ of the pair $(X,Y)$ as a reference point, but $(X^\prime,Y^{\prime\prime})$, a pair of independent copies of $X$ and $Y$, respectively, which are also independent of each other. Rewriting (ref) again such that it becomes symmetric with respect to $(X,Y)$, $(X^\prime,Y^\prime)$ and $(X^{\prime\prime},Y^{\prime\prime})$, see (ref) in the appendix and the discussion around it, we get
where
The kernel $k^{(\rho)}$ can be written more compactly in the form of (ref) in the appendix, but we prefer not to use the more cumbersome notation from the appendix here, at the cost of writing out the sum in (ref). The empirical analogue of representation (ref) is a U-statistic, which is the starting point for our asymptotic analysis of $\widehat{\rho}$, see (ref) below.
Kendall's $\tau$ and Spearman's $\rho$ possess all major desirable properties for dependence measures in the continuous case embrechts2002. However, for discrete random variables, they lack the property of attainability neslehova2007, genest2007, that is, they do not take the values $-1$ and $1$ under perfect negative and positive dependence (counter- and comonotonicity), respectively. In fact, they may be very far away from those values under perfect or strong dependence. This leads to problems with their interpretability as strong dependence should be indicated by values close to 1 in absolute value. pohle2025 analyze attainable generalizations of $\tau$ and $\rho$, which thus fulfill all desirable properties in general, and refer to them as “proper”. For Kendall's $\tau$, there exists a particularly nice and simple proper counterpart, namely Goodman-Kruskal's $\gamma$. This measure just computes Kendall's $\tau$ conditional on the event that there are no ties in both $X$ and $Y$, recall (ref).
A popular generalization of Kendall's $\tau$, which is often used to address its attainability problem, is Kendall's $\tau_b$, which uses the univariate tie probabilities (ref) instead of $\nu$.
However, while $\tau_b$ mitigates the problem in that it is weakly larger in absolute value than $\tau$, pohle2025 show that it is still non-attainable and thus inferior to $\gamma$. Nevertheless, we consider $\tau_b$ due to its popularity and the widespread use of its empirical analogue in statistical software.
Unfortunately, for Spearman's $\rho$, there is no simple proper version. We consequently recommend using the grade correlation coefficient, which constitutes, in a way, the analogue to $\tau_b$ as a generalized version of $\rho$ in the discrete case. However, it only mitigates the attainability problem (like $\tau_b$ does), but does not cure it.
Again, we refer to Appendix (ref) for the special case of the bivariate geometric distribution.
We now define the empirical versions of Kendall's $\tau$ and Spearman's $\rho$ and discuss their (asymptotic in the case of the latter) representations as U-statistics, which are the basis for our approach to establishing asymptotic theory. From now on, let $\{X_i,Y_i\}_{i=1}^n$ be a bivariate time series of length $n$, that is a sample of size $n$ of a stochastic process $\{X_i,Y_i\}_{i\in \mathbb{Z}}$ with $\mathbb{Z} = \{\ldots, -1,0,1, \ldots\}$. This includes the case of an iid sample of size $n$ as a special case. As we will assume strict stationarity, the time index does not play a role when we deal with marginal distributions involving $X_i$ or $Y_i$, so we omit it in these cases, e.g.\ writing $F_X$ instead of $F_{X_i}$. We also omit the time index when we deal with joint distributions of $(X_i,Y_i)$ only at a single point in time $i$.
We note that $\widehat{\tau}$ is a U-statistic with kernel
We present a short introduction to U-statistics in Appendix (ref). In particular, see Definition (ref).
Denote the empirical distribution function computed from a univariate time series $\{X_i\}_{i=1}^n$ by $\widehat{F}_X(x)$, and the empirical MDF by $\widehat{G}_X(x):=\big(\widehat{F}_X(x) + \widehat{F}_X(x^{-})\big)/2$ (and similarly for $\{Y_i\}_{i=1}^n$). Furthermore, define the mid-rank of the observation $X_j$ in $\{X_i\}_{i=1}^n$ as
It holds that $\widehat{G}_X(X_j) = \frac 1 n \big( \widehat{R}_X(X_j) - 1/2\big)$.
Contrary to $\widehat{\tau}$, $\widehat{\rho}$ itself is not a U-statistic, but asymptotically equivalent to the empirical analogue of (ref):
where $k^{(\rho)}$ is defined in (ref). More precisely, the relation between $\widehat{\rho}$ and $\widetilde{\rho}$ is the following, see hoeffding1948, which arises essentially by the relation between midranks and the difference sign function from (ref):\footnote{The slightly different factors involving $n$ in hoeffding1948 arise because Hoeffding defines empirical Spearman's $\rho$ as $n^2/(n^2-1)\, \widehat{\rho}$.}
where the second summand is $\mathcal{O}_{\mathrm{P}}(n^{-1})$ and thus plays no role asymptotically. So $\widetilde{\rho}$ and $\widehat{\rho}$ are asymptotically equivalent in probability, $\widehat{\rho}/\widetilde{\rho} \stackrel{p}{\to} 1$.
We now consider the empirical counterparts of Goodman-Kruskal's $\gamma$, Kendall's $\tau_b$, and grade correlation from Section (ref).
We note that $\widehat{\nu}$ is a U-statistic with kernel $$k^{(\nu)} \left( (x^\prime, y^\prime),(x,y) \right) = \mathds{1}\big((x-x^\prime)(y-y^\prime)=0\big).$$
Empirical grade correlation is just the empirical Pearson correlation of the empirical grades. $\widehat{\gamma}$, $\widehat{\tau}_b$, and $\widehat{\rho}_b$ are all differentiable functions of U-statistics (at least asymptotically in the case of the latter).
To show consistency and to derive the asymptotic distributions of the five empirical rank correlations from the previous section, we exploit that $\widehat{\tau}$ is a U-statistic, $\widehat{\rho}$ is asymptotically equivalent to the U-statistic $\widetilde{\rho}$, $\widehat{\gamma}$ and $\widehat{\tau}_b$ are functions of U-statistics, and $\widehat{\rho}_b$ is asymptotically equivalent to a function of U-statistics. Thus, we can employ limit theorems for U-statistics. Note that for notational convenience, we will usually suppress the arguments $X$ and $Y$ when writing down our dependence measures, e.g.\ $\tau:=\tau(X,Y)$. In this section, we work under the iid assumption.
Note that we do not need moment conditions as the considered rank-based measures do not operate with the observations themselves, but with their ranks, grades, or CDFs, respectively. All limit theorems from this section are special cases of the more general results in the next section, where we work under classical time series assumptions, so that the formulas for the asymptotic variances get more involved. We nevertheless dedicate a separate section to the results for iid data because they are of high relevance in their own right and new to a large part. Furthermore, it is beneficial to appreciate the less involved results in this section, in particular regarding variance estimation, before tackling the cumbersome time series results. In Appendix (ref), we review a univariate central limit theorem (Proposition (ref)) for U-statistics under the iid assumption hoeffding1948 and under weak dependence denker1983. Furthermore, we present a multivariate central limit theorem for U-statistics (Proposition (ref)) under both assumptions, which in combination with the delta method beutner24 allows for the treatment of $\widehat{\gamma}$, $\widehat{\tau}_b$, and $\widehat{\rho}_b$.
The following consistency result follows from the classical law of large numbers for U-statistics hoeffding1948 and the continuous mapping theorem.
The following proposition on the asymptotic distributions of $\widehat{\tau}$ and $\widehat{\rho}$ follows from a central limit theorem for U-statistics, namely Proposition (ref) in the appendix.
The variance formulas arise from the general formula for the asymptotic variance of a U-statistic (see again Proposition (ref) in the appendix), where $k_1$ is the leading term of the Hoeffding decomposition of U-statistics (see (ref) and (ref) in the appendix), and where the factors 4 and 9, respectively, arise from the squared orders of the U-statistics, here $2^2$ and $3^2$. For example, this leading term is computed for $\tau$ as $$ k_1^{(\tau)}(x,y) \ =\ \mathrm{E}\big[k^{(\tau)}\big((x,y), (X_i,Y_i)\big)\big]-\tau \ =\ \mathrm{E}\big[\mathrm{sgn}(x-X_i)\, \mathrm{sgn}(y-Y_i)\big]-\tau, $$ and the latter is computed by part (ii) of Lemma (ref). The full derivations of the other leading terms appearing in Proposition (ref) and (ref) are provided by Lemma (ref) in the appendix. The asymptotic distribution of $\widehat{\tau}$ has already been established by hoeffding1948 and, in the continuous case, also the asymptotic distribution of $\widehat{\rho}$.
The asymptotic distributions of the generalized rank correlations in the iid case follow from our multivariate central limit theorem for U-statistics in Proposition (ref) in the appendix together with the delta method.
Proposition (ref) in the appendix is not only the key ingredient to obtain this proposition, but also shows that vectors of the empirical rank correlations considered here are asymptotically multivariate normal and enables the calculation of the corresponding variance matrix.
In the iid case, our proposed variance estimator requires, in a first step, the estimation of $k_1^{(\tau)}, k_1^{(\tau_X)}, k_1^{(\tau_Y)}, k_1^{(\rho)}, k_1^{(\rho_X)}, k_1^{(\rho_Y)}$ and $k_1^{(\nu)}$ (depending on which of those functionals are relevant for the chosen coefficient). In a second step, the expectation in the asymptotic variances is replaced by a sample mean. In the time series case, a third step comes in additionally, namely a heteroscedasticity and autocorrelation consistent (HAC) estimator being well known from variance estimation of the sample mean under temporal dependence.
To go into more detail, let us have a closer look at the asymptotic (co)variances in the Subsection (ref), which are spelled out in general form in the formula for $\sigma_{lm,\textup{iid}}$ in Proposition (ref). On the one hand, they contain the factor $r^{(l)}$, which is the (known) order of the $i$-th U-statistic. On the other hand, they contain the functionals $k_1^{(l)}$ depending on the joint CDF $F_{X,Y}$ of $(X,Y)$, and these need to be estimated. To estimate them, we replace the theoretical CDFs $F_{X,Y}$ by empirical CDFs $\widehat{F}_{X,Y}$, and accordingly the theoretical univariate or multivariate MDFs or PMFs by the empirical ones. For example, for $\tau$, we replace the MDFs $G_X$ and $G_Y$ in (ref) by their empirical counterparts $\widehat{G}_X$ and $\widehat{G}_Y$, which arise by replacing the theoretical CDFs $F_X$ and $F_Y$ in the definitions of the MDFs in and above $\eqref{eq:bivariate_mid-distribution_function}$ with empirical CDFs $\widehat{F}_X$ and $\widehat{F}_Y$:
In the cases of $\widehat{\rho}$ and $\widehat{\rho}_b$, an additional difficulty arises due to functions of the form $g_X(x):=\mathrm{E}[G_{X,Y} (x,Y) ]$ arising in $k_1^{(\rho)}$. However, they can be estimated by $\widehat{g}_X (x) = \frac 1 n \sum_{i=1}^n \widehat{G}_{X,Y} (x,Y_i)$, and thus we estimate $k_1^{(\rho)}$ by
We finally estimate the expectation $\sigma_{lm,\textup{iid}}$ via
We establish the consistency of this estimator. Because the estimators contain averages over random functions, e.g.\ $\widehat{k}^{(\tau)}_1 (x,y)$, and the asymptotic (co)variances featuring $\rho$ even contain averages of averages of random functions, the consistency is not trivial. Yet, it follows by standard arguments on random functions from empirical process theory using the Glivenko-Cantelli theorem. For details, see the proof in Appendix (ref) and vandervaart2000.
Building on the derived asymptotic distributions and proposed variance estimators, classical asymptotic confidence intervals and tests can be constructed. For any of our dependence measures, say $\delta$, the corresponding consistent estimator $\widehat{\delta}$ and variance estimator $\widehat{\sigma}^2_{\delta,\textup{iid}}$, the confidence interval at level $1-\alpha$ is
where $z_{1-{\alpha}/2}$ denotes the $1-{\alpha}/2$-quantile of the standard normal distribution. The test statistic for the test for the null hypothesis $H_0: \delta=\delta_0$ is
and the corresponding p-value against $H_1: \delta \neq \delta_0$ is $2(1-\Phi (|T_{\delta_0}|))$, where $\Phi$ denotes the CDF of the standard normal distribution.
The Fisher transformation introduced next leads to improved finite-sample performance of confidence intervals and tests if the true parameter is close to the bounds $\pm 1$.
It is well known that for the empirical Pearson correlation, the normal limiting distribution does not yield a good approximation if the true Pearson correlation is close to $-1$ or $1$, because then the true sampling distribution is heavily skewed. A solution to this problem is to transform the empirical correlation to an unbounded scale (the real line) by the Fisher transformation fisher1915, and to use the limiting distribution of the Fisher transformation instead. This yields a much better approximation close to the boundaries and leaves the situation essentially unchanged away from the boundaries. This reasoning carries over to other correlation coefficients as well, see pohle2024measuring for a detailed discussion of the Fisher transformation for generic dependence measures. The Fisher transformation is defined as
for a generic dependence measure $\delta$ (or an empirical dependence measure $\widehat{\delta}$). By the delta method, if the asymptotic distribution of $\widehat{\delta}$ is normal with variance $\sigma^2_{\delta}$, the asymptotic distribution of $z(\widehat{\delta})$ is normal as well, but with variance $\sigma^2_{\delta}/(1-\delta^2)^2 $, which can be estimated by plugging in estimators $\widehat{\delta}$ and $\widehat{\sigma}^2_{\delta}$. Tests can then be executed directly based on the limiting distribution of the Fisher transformation. For confidence intervals, the bounds of the confidence interval from this limiting distribution need to be transformed back to the original scale by applying the inverse Fisher transformation $z^{-1}(x) = \mathrm{tanh}(x) = (e^{2x}-1)/(e^{2x}+1)$.
Rank correlations equal 0 under independence of $X$ and $Y$. An important and classical application of rank correlations is testing the null hypothesis of independence via the implication of uncorrelatedness. We thus derive the formulas for the asymptotic variances under independence. As the basis for this, note that $k_1^{(\tau)}$ and $k_1^{(\rho)}$ from Proposition (ref) simplify considerably and become equal to each other, $k_1^{(\tau)}=k_1^{(\rho)}$, under independence, see Lemma (ref) in the appendix.
Corollary (ref) nicely extends the classical results for the continuous case, where the asymptotic variances are $4/9$ and $1$, respectively, to the general case, where the additional factors involving the double-tie probabilities of $X$ and $Y$ arise. Even though the asymptotic variances are quite fundamental and well known for the continuous case, see kendall1938 and hoeffding1948, we are not aware of any reference for their neat formulas in the general case. Interestingly, already hoeffding1948 conjectured that they depend on the tie probabilities $\mathrm{P}(X=X^\prime)$ and $\mathrm{P}(X=X^\prime=X^{\prime\prime})$.
Note that by the equality of $k_1^{(\tau)}$ and $k_1^{(\rho)}$, and our multivariate CLT for U-statistics (Proposition (ref) in the appendix), the asymptotic correlation between $\widehat{\tau}$ and $\widehat{\rho}$ under independence is 1. Considering the ratio of the standard deviations, this implies that they asymptotically have the deterministic relationship $\widehat{\rho} = 3/2\, \widehat{\tau}$, which was already proven by daniels1944. Thus, the higher asymptotic variance of $\widehat{\rho}$ is not caused by it being less efficient, but by it taking higher (absolute) values under independence.
As an immediate consequence from Proposition (ref), using that $\tau=\rho=\gamma=0$ under independence of $X$ and $Y$, we obtain the limiting distribution of the generalized rank correlations under independence. For $\widehat{\gamma}$, we also use that $1-\nu$ from (ref) simplifies to
under independence.
The higher asymptotic variance of $\widehat{\tau}_b$ compared to $\widehat{\tau}$ (as well as that of $\widehat{\rho}_b$ compared to $\widehat{\rho}$) reflects the fact that it is weakly larger, mitigating the attainability problem. The even higher variance of $\widehat{\gamma}$ is again due to the fact that it takes weakly larger values than $\widehat{\tau}_b$ and solves the attainability problem. Nicely, the asymptotic variance of the grade correlation $\widehat{\rho}_b$ is 1, like the variance of $\widehat{\rho}$ itself in the continuous case. Thus, $\widehat{\rho}_b$ carries the nice property of an asymptotic variance not containing any unknown quantities that need to be estimated, which $\widehat{\rho}$ and $\widehat{\tau}$ possess in the continuous case, to the discrete world. Note that for the sake of testing independence, one could also define a variant of $\tau$ where the empirical counterpart has a constant asymptotic variance of $4/9$ under independence, namely by using the denominator of $\rho_b$:
To apply Corollaries (ref) or (ref) in practice, in order to test for cross-dependence, the probabilities $\zeta(X) = \mathrm{P}(X=X^\prime)$ and $\zeta_2(X) = \mathrm{P}(X=X^\prime=X^{\prime\prime})$ (analogously for $Y$) in the asymptotic variances need to be specified (except for the case of grade correlation $\rho_b$ as well as $\tau_{b,\textup{mod}}$). In the continuous case, $\zeta(X)=\zeta_2(X)=0$. In the discrete case, these probabilities have to be estimated from the given data. Recall that by (ref), it suffices to estimate the marginal distribution of $\{X_i\}$. In a non-parametric setup, this is done by computing relative frequencies instead of the probability masses, whereas simplifications are possible in parametric setups. For example, if $X$ is geometrically distributed according to $\mathrm{Geom}(\pi)$, then $\zeta(X) = \pi/(2-\pi)$ and $\zeta_2(X) = \pi^2/(3-(3-\pi)\pi)$, recall Appendix (ref). Then, only the estimation of $\pi$ is necessary. A plot of the corresponding theoretical tie probabilities for $X\sim \mathrm{Geom}\left(1/(1+\mu) \right)$ is provided by Figure (ref) (left). The tie probabilities quickly converge towards 0 for increasing mean $\mu$. Figure (ref) (right) shows an analogous plot for the asymptotic variances from Corollaries (ref) or (ref), where the variables $X,Y$ are assumed to be independent and identically distributed according to $\mathrm{Geom}\left(1/(1+\mu) \right)$.
Estimating the tie probabilities as described above and plugging those estimators into the variance formulas from the two preceding corollaries, we arrive at estimators for $\widehat{\sigma}^2_{\delta,\textup{iid},\textup{ind}}$. For testing the null hypothesis of independence via the implication of uncorrelatedness, we simply use the test statistic from (ref), plugging in $\delta_0=0$ and $\widehat{\sigma}_{\delta,\textup{iid},\textup{ind}}$ as the estimator for the standard deviation.
In this section, we work under classical time series assumptions.
Absolute regularity or $\beta$-mixing is a form of asymptotic independence. It is stronger than strong mixing and weaker than uniform mixing, see bradley2005 for details.
Assumption (ref) is a special case of Assumption (ref), so we have to generalize the results from the previous section. The proofs for the limiting theorems are analogous, we just have to replace the law of large numbers and the central limit theorem under the iid assumption by more general limit theorems under Assumption (ref).
Strong consistency of all five empirical rank correlations follows directly by an appropriate law of large numbers for U-statistics under weak dependence dehling2006, together with the continuous mapping theorem in the cases of $\widehat{\gamma}$, $\widehat{\tau}_b$, and $\widehat{\rho}_b$.
To generalize the results from Propositions (ref) and (ref), we need a univariate central limit theorem under Assumption (ref) (see denker1983 and Proposition (ref) in the appendix), a multivariate version of it (Proposition (ref) in the appendix), and again the delta method. Furthermore, for the derivation of the asymptotic variances, we need Lemma (ref) in the appendix.
Thus, only the expressions for the asymptotic variances and covariances of the U-statistics occurring in the definitions of the rank correlations change compared to the iid case. In the continuous case, dehling2017 obtained the limiting distribution for Kendall's $\tau$ under similar assumptions.
We denote the cross- and autocovariances between two zero-mean processes $\{V_i\}_{i \in \mathbb{Z}}$ and $\{W_i\}_{i \in \mathbb{Z}}$ by $\alpha_{V,W} (h) := \mathrm{Cov}(V_i,W_{i+h})$.\footnote{Note that we denote the autocovariance function by $\alpha$ instead of using the classical notation $\gamma$ as we already use this letter for Goodman-Kruskal's $\gamma$.} An infinite sum over the cross-autocovariances between the processes $\{ k^{(l)}_1(X_i,Y_i) \}_{i \in \mathbb{Z}}$ and $\{ k^{(m)}_1(X_i,Y_i) \}_{i \in \mathbb{Z}}$ emerges in $\sigma_{lm}$, that is, it can be rewritten as
The following lemma shows that the asymptotic covariance $\sigma_{lm}$ is always finite under our assumptions.
Such an absolute summability condition on the autocovariances is very common as an assumption in time series analysis, ruling out long memory (e.g.\ hassler2018). Here, it is already implied by Assumption (ref) and the boundedness of our kernels.
We note that the structure of the asymptotic covariance $\sigma_{lm}$ is similar to the asymptotic variance of a sample mean $\overline{f}=\frac 1 n \sum_{i=1}^n f(X_i,Y_i)$, where $f$ is a known function and $\{X_i,Y_i\}_{i\in \mathbb{Z}}$ a bivariate stochastic process satisfying Assumption (ref). More precisely, it holds that $\sigma^2_{\mu_f} = n\, \mathrm{Var}( \overline{f}) = \sum_{h=-\infty}^{\infty} \alpha_{f(X,Y),f(X,Y)} (h)$ and similarly for the asymptotic covariance.
Compared to the estimator $\widehat{\sigma}_{lm,\textup{iid}}$ from Subsection (ref), the additional difficulty arises from the infinite sum in (ref). Besides estimating the functions $k_1^{(l)}$ and $k_1^{(m)}$ as laid out in Subsection (ref), we now employ classical techniques for long-run variance estimation, which are well known from the case of the sample mean newey1987. That is, we estimate $\sigma_{lm}$ via
where the weighting function $w(h/(b_n+1))$ is also called a kernel or a window, $b_n$ is the bandwidth, and we define the empirical cross-autocovariance function between the zero-mean processes $\{V_i\}_{i \in \mathbb{Z}}$ and $\{W_i\}_{i \in \mathbb{Z}}$ by $$ \widehat{\alpha}_{V,W} (h) := \frac 1 n \sum_{i=1}^{n-h}V_iW_{i+h}. $$ We refer to lazarus2018 for an overview of the vast literature on long-run variance estimation, or heteroscedastiticy and autocorrelation consistent (HAC) estimation, in the case of sample means or regression coefficients. dehling2017 propose the use of a HAC estimator of this type for U-statistics of order 2 and especially for Kendall's $\tau$, but under a different set of assumptions. We establish consistency of our variance estimators under our assumptions on the time series and standard assumptions on the kernel and the bandwidth.
We use the Bartlett kernel, or triangular kernel, which has been popularized by newey1987: \[ w_B\biggl(\frac{h}{b_n+1}\biggr)=\left\{
\right. . \] As the bandwidth, we choose $b_n = \lfloor 2 n^{1/3} \rfloor$ as in dehling2017. With these choices, Assumption (ref) is satisfied. We conjecture that the bounded support assumption for the kernel can be replaced by an assumption on its rate of decay.
It is well known that variance estimation under temporal dependence is a notoriously difficult problem, usually plagued by oversized tests and confidence intervals with undercoverage, with the problem getting worse with the strength of temporal dependence and only disappearing for large sample sizes, see again lazarus2018 on this issue. Alternatives to the HAC estimator such as the moving block bootstrap also exhibit this behaviour fitzenberger1998. As expected, these problems also arise in our simulations, see Section (ref).
Except for replacing the variance estimators for the iid case with the ones proposed in this section, confidence intervals and tests are constructed as described around (ref) and (ref).
The case where $\{X_i\}_{i \in \mathbb{Z}}$ and $\{Y_i\}_{i \in \mathbb{Z}}$ are mutually independent, but themselves serially dependent processes, is of interest for testing the null hypothesis of independence between the two processes. We thus generalize Corollaries (ref) and (ref) from the iid case to serial dependence. The resulting expressions are considerably simpler than in Proposition (ref), which is essentially due to the simpler form of $k_1^{(\tau)}$ and $k_1^{(\rho)}$ stated in Lemma (ref). The latter leads to a neat simplification of the autocovariances in (ref), also see Lemma (ref) in the appendix.
Note that in order to achieve compact formulas, we have expressed the asymptotic variances for $\tau, \rho, \gamma$ and $\tau_b$ via the autocorrelation in terms of $\rho$, and the asymptotic variance for the grade correlation via the autocorrelation in terms of $\rho_b$. We could also express all results in terms of one or the other, see Definition (ref) and Lemma (ref) in the appendix for their relation. In the continuous case, this result for $\widehat{\tau}$ and $\widehat{\rho}$ has recently been derived by lun2023.
In order to estimate the asymptotic variances from Corollary (ref), we estimate the Spearman autocorrelations of $X$ on the sample $\{X_i,X_{i+h}\}_{i=1}^{n-h}$ via $\widehat{\rho}_X (h) := \widehat{\rho} (X_i,X_{i+h})$, see Definition (ref) (and analogously for $Y$), plug them in for their theoretical counterparts and downweight the summands in the same way as for our variance estimator discussed in and below (ref). Consistency of this estimator follows by standard arguments for the consistency of a HAC estimator as employed in the proof of Proposition (ref). For $\rho_b$, we use the values of the Spearman autocorrelations but estimate the denominator (i.e.\ the tie probabilities) only once on the full sample.
In order to evaluate the finite-sample performance of the derived asymptotic distributions and proposed variance estimators for the different dependence measures, we report on a comprehensive simulation study. For many different types of data-generating processes (DGPs) -- both continuous and discrete, and both iid and serially dependent, we simulated $MC=1,000$ bivariate iid samples or time series, respectively, which are used for either testing for independence, or for computing a confidence interval for the respective dependence measure. We summarize the most important results of the simulation study here. For details on the DGPs, tabulated simulation results, and a more in-depth analysis and interpretation, we refer to Appendix (ref).
We start with the independence tests for iid DGPs (see Corollaries (ref) and (ref) and the corresponding variance estimators), where we also compare the size and power values of the tests based on rank correlations to tests based on Pearson correlation (see equations (ref) and (ref) for its asymptotic distribution under independence in the iid and the time series case). Pearson correlation shows mild size distortions outside the bivariate normal case and strong size distortions for heavy-tailed DGPs. This pattern does not change between continuous and discrete DGPs. However, the size distortions further deteriorate with growing sample size when second moments are infinite and improve otherwise. The rank correlations on the other hand are robust to distributional characteristics such as heavy tails and hold the nominal size well across all DGPs. Only Kendall's $\tau$ shows mild oversizing in small samples ($n=50$). In terms of power, Pearson correlation again shows a worse performance than the rank correlations for the heavy-tailed DGPs. Additionally, asymmetries in the distributions seem to worsen the performance of Pearson correlation in terms of both size and power. Besides that, all coefficients show rather similar rejection rates with slight advantages for Spearman's $\rho$ in the continuous and for Goodman-Kruskal's $\gamma$ in the discrete case. Expectedly, all independence tests based on rank correlations exhibit increasing power with increasing sample size.
Turning to the coverage rates of the confidence intervals for iid DGPs (see Propositions (ref), (ref) and (ref)), all coefficients show an empirical coverage close to the theoretical confidence level. Only for strong cross-dependencies, the former sometimes falls short of the latter, a misbehavior that diminishes when applying the Fisher transformation (see Subsection (ref)).
For the time series DGPs, independence tests amount to tests of whether the two considered time series are independent of each other (and we thus employ Corollary (ref) and the corresponding variance estimators). Due to the rather strong serial dependence we use in our time series DGPs (our DGPs involve AR(1) processes with a coefficient value of 0.8) and the well-known difficulties with variance estimation under temporal dependence discussed in Subsection (ref), we observe strong oversizing for all DGPs and all coefficients even for $n=800$. Again, Pearson correlation performs worst for the heavy-tailed and asymmetric DGPs. For the independence tests based on the asymptotic theory in this paper, we again observe improving power with increasing length of the time series.
For the confidence intervals in the time series case, the empirical coverages, in turn, suffer from the biased variance estimates (see Propositions (ref) and (ref)) in the same way as the sizes of the independence tests. We observe strong undercoverage, especially for short time series, that improves with growing length of the time series. In addition to that, the strength of cross-dependence does also seem to play an important role for the empirical coverage, with some (continuous) DGPs showing decreasing coverage for increasing strength of dependence and some other DGPs exhibiting the opposite relationship.
To illustrate the practical application of the different dependence measures and the proposed confidence intervals and tests, we consider two data examples. The first one involves an ecological data set, where the data can be assumed to arise from an iid DGP, see Section (ref) for details. In Section (ref), we consider a time series of daily numbers of road accidents. In both cases, the data are discrete counts such that our novel asymptotics are crucial for executing dependence tests and computing confidence intervals.
We consider the bivariate count data $(x_1,y_1),\ldots,(x_{100},y_{100})$ in Table 1 of holgate66, which were further analyzed by kokonendji18. The data arose in a study of a secondary rain forest in Trinidad, and they express the numbers of plants of the species Lacistema aggregatum ($x_i$) and Protium guianense ($y_i$) in each of 100 contiguous quadrants. Figure (ref) visualizes the data using a bubble plot.
Due to the experimental design, the data can be understood as the realization of an iid DGP, so we make use of the asymptotics in Section (ref). Analyzing the bivariate data for possible cross-dependence is of practical relevance in order to find out if the species tend to occur together, to be unrelated or to be mutually exclusive. In Table (ref) we report the values of empirical Pearson correlation $\widehat{r}$, the classical empirical rank correlation $\widehat{\tau}$ and $\widehat{\rho}$ and their generalizations $\widehat{\tau}_b$, $\widehat{\gamma}$ and $\widehat{\rho}_b$, which are better suited for the discrete case. This table also reports $90\%$ confidence intervals for the theoretical counterparts of these coefficients as given in (ref) and p-values for tests of independence and uncorrelatedness as in and around (ref), where for the independence tests the simplified variance formulas from Corollaries (ref) and (ref) (and for Pearson correlation from (ref)) were used.
When first being interested in the question if there is any dependence at all, that is, considering the tests of independence, all rank-correlation-based tests speak the same language, yielding p-values below $5\%$. The test based on Pearson correlation, on the contrary, yields a p-value way above $10\%$. This could be due to a monotonic, but nonlinear form of dependence, which the rank correlations can capture, implying that the respective tests have power in this direction, whereas Pearson correlation is tailored to measuring linear dependence. The tests for uncorrelatedness, i.e., for the respective correlations being 0, have very similar p-values as the tests for independence. Thus, in this case using the null hypothesis of independence to simplify the variance formula does not lead to efficiency gains.
The values of the correlation coefficients inform us about direction and strength of dependence. All coefficients are positive, but Pearson correlation with a value of 0.114 is smaller than all the rank correlations (and insignificant as discussed above), in particular than the generalized rank correlations, which range from 0.199 to 0.302 and which should be used here for measuring strength of dependence due to the discreteness of the data. Thus, we can conclude on a (mild) tendency of co-occurrence for both species, where the dependence does not appear to be a linear one.
Our second example is about bivariate time series data, namely the bivariate counts $(x_i,y_i)$, $i=1,\ldots,365$, which refer to the daily numbers of daytime ($x_i$) and nighttime ($y_i$) road accidents in the Schiphol area (Netherlands) during the year 2001. The data have been read from Figure 1 in pedeli11. These authors showed that the data exhibit significant serial dependence and propose to model the data by a bivariate INAR$(1)$ model. In Figure (ref) we again depict the empirical joint distribution in a bubble plot. Figure (ref) presents a time plot of both series.
Here, our interest is in the cross-dependence between the daytime and nighttime accidents, where hypothesis tests and confidence intervals have to account for the serial dependence (without imposing a specific time series model) as described in Section (ref). All rank correlations and Pearson correlation indicate mild positive positive dependence (see Table (ref)) with values around 0.1. Interestingly and contrary to the first data example, Pearson correlation is slightly larger than all of the rank correlations here. Usually, Pearson correlation being larger than Kendall's $\tau$ or Spearman's $\rho$ is associated with linear dependence, that is, a form of dependence that can be captured well by Pearson correlation. A more in-depth explanation is as follows: The shape of the bubble plot suggests that the joint distribution can be well approximated by an elliptical distribution. For elliptical distributions, firstly, the dependence is fully characterized by Pearson correlation $r$, and secondly, there are well-known functional relationships between $r$ and the classical rank correlations $\tau$ and $\rho$, which show that these rank correlations are always smaller than $r$ embrechts2002. The respective generalized rank correlations are greater than $\tau$ and $\rho$, but still smaller than $r$ here.
When it comes to inference, all correlations are significantly different from 0. Interestingly, again contrary to the first data example, the p-values nevertheless strongly depend on the actual kind of estimating the variance. If we make use of the component-wise independence as postulated by the null hypothesis (see Section (ref)), we get p-values $<2\%$ throughout, i.e., we conclude on significant dependence even on the 5%-level. If estimating the variance without imposing component-wise independence as in Section (ref) (as when computing confidence intervals), the p-values of the tests become larger with values around 6%. Thus, there are efficiency gains from imposing the null hypothesis of independence on the variance. Note that since the temporal dependence is mild in this application (all first order-autocorrelations and cross-autocorrelations for the two series are between 0.1 and 0.2), our HAC-type variance estimators should not suffer from substantial size distortions.
Altogether, we conclude on mild positive dependence between the daytime and nighttime accidents on a day $i$, which appears reasonable as for example general weather conditions or holidays agree between $x_i$ and $y_i$ for a given day $i$.
We provide a comprehensive treatment of asymptotic inference for the classical rank correlations Kendall's $\tau$ and Spearman's $\rho$ as well as their generalizations for discrete random variables $\tau_b$, Goodman-Kruskal's $\gamma$, and grade correlation. Under iid or time series assumptions, and not restricting ourselves to continuous variables, we derive limiting distributions by using asymptotic results for U-statistics and propose variance estimators that enable the construction of confidence intervals and tests. Furthermore, we derive simplified variance formulas in the case of independence, which form the basis for independence tests.
Although asymptotic confidence intervals and tests are the most classical approach to statistical inference, and Kendall's $\tau$ and Spearman's $\rho$ have favorable properties compared to Pearson correlation and are commonly employed in descriptive statistics, it is quite surprising that asymptotic inference for these rank correlations has been lacking fundamental results and, thus, has rarely been used in inductive statistics. Consequently, sound inferential procedures (with the exception of independence tests) are not implemented in most statistical software. The results of this paper, which of course build on earlier foundational work such as hoeffding1948 and dehling2017, allow for a change in this respect.
In the iid setting, the proposed asymptotic inference performs very well and we posit that it ought to be the method of choice as in related settings. Certainly, resampling-based inference for rank correlations constitutes a valid alternative. It thus seems to be a fruitful endeavor for future research to compare the performance of the two approaches.
For time series data, we discover some problems with variance estimation under strong temporal dependence and in small to medium-sized samples. This, however, is not surprising as variance estimation under temporal dependence in similar settings (e.g.\ for the sample mean or linear regression coefficients) is a notoriously difficult problem, plagued by oversized tests and confidence intervals with undercoverage. Despite these issues, we consider our asymptotic confidence intervals and tests to be a substantial improvement on the status quo in time series settings, where inferential procedures are hardly available so far. Future research in this direction is undoubtedly crucial. For example, self-normalization might be a way to tackle these problems shao2015, or at least the improvement of the HAC estimator by systematically comparing different choices of the bandwidth or kernel.
While Kendall's $\tau$, Spearman's $\rho$, and their variants for the discrete case, are the classical and most widely-used rank correlations, inference for other representatives of this class is also relevant and can be approached by using the results of this paper. For example, there are asymmetric versions of $\tau_b$ and grade correlation: Somers' $D$ defined as $\tau(X,Y)/\tau(X,X)$ somers1962, newson2002, and the recently introduced asymmetric grade correlation defined as $\rho(X,Y)/\rho(X,X)$ walz2025, measure directed dependence of $Y$ on $X$, e.g.\ to assess the predictive potential. In the continuous case, they equal $\tau$ and $\rho$, respectively, and for a dichotomous $Y$, they are both related by a one-to-one mapping to the area under the receiver operating characteristic (ROC) curve, short area under the curve (AUC), which is a popular measure to assess the performance of binary classifiers walz2025. For all these measures, inference tools can be constructed by using the techniques introduced in this paper. walz2025 do this for the case of asymmetric grade correlation and develop tests for the equality of two grade correlations as well. Spearman's $\rho$ and asymmetric grade correlation are also closely related to the slope coefficients of rank-rank regressions, for which inference has recently been treated by chetverikov2023.