EconBase
← Back to paper

Generalised Covariances and Correlations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,748 characters · 21 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Generalised Covariances and Correlations

abstractAbstract. The covariance of two random variables measures the average joint deviations from their respective means. We generalise this well-known measure by replacing the means with other statistical functionals such as quantiles, expectiles, or thresholds. Deviations from these functionals are defined via generalised errors, often induced by identification or moment functions. As a normalised measure of dependence, a generalised correlation is constructed. Replacing the common Cauchy--Schwarz normalisation by a novel Fr\'echet--Hoeffding normalisation, we obtain attainability of the entire interval $[-1, 1]$ for any given marginals. We uncover favourable properties of these new dependence measures and establish consistent estimators. The families of quantile and threshold correlations give rise to function-valued distributional correlations, exhibiting the entire dependence structure. They lead to tail correlations, which should arguably supersede the coefficients of tail dependence. Finally, we construct summary covariances (correlations), which arise as (normalised) weighted averages of distributional covariances. We retrieve Pearson covariance and Spearman correlation as special cases. The applicability and usefulness of our new dependence measures is illustrated on demographic data from the Panel Study of Income Dynamics.

Keywords: dependence measure; statistical functional; identification function; quantile correlation; copula; tail dependence

Introduction

Measuring the dependence of two random variables $X$ and $Y$ has been a long-standing task in statistics with relevance for almost any empirical field of science. The two key approaches to this task are regression analysis, considering $Y$ conditional on $X$, and mutual dependence measures. The most popular measures of dependence are covariance, Pearson correlation and the rank correlations coefficients Spearman's $\rho$ and Kendall's $\tau$. Overviews of the vast literature on dependence measures are given in Mari2001, Balakrishnan2009, and Tjostheim2022. Directed dependence measures are usually normalised to range between $-1$ and $1$, and they indicate the direction of dependence by their sign and the strength of dependence by the proximity of their absolute value to 1. Crucial for the usefulness and interpretability of dependence measures are certain properties, called R\'enyi's axioms Renyi1959 and often modified subsequently Schweizer1981, Embrechts2002, Balakrishnan2009. In particular, a dependence measure should indicate the extreme forms of independence, perfect positive and negative dependence by attaining the values 0, 1 and $-1$, respectively. Pearson correlation suffers from some well-known shortcomings Embrechts2002, most importantly, attainability issues: For given marginal distributions of $X$ and $Y$ there are in general no joint distributions with these marginals achieving a Pearson correlation of $1$ and $-1$, respectively. This seriously impacts the interpretability of Pearson correlation.

Let us now consider the definitions of covariance $\mathrm{Cov}(X,Y)$ and Pearson correlation $r(X,Y)$ more closely:

equation[equation omitted — 178 chars of source]

where $\mathrm{Var}(X) = \mathrm{Cov}(X,X)$ denotes the variance of $X$. Covariance measures the average co-movements of $X$ and $Y$ around their respective means, $\mu(X)$ and $\mu(Y)$. Pearson correlation is a normalised version of covariance, relying on the Cauchy--Schwarz inequality. We aim at measuring the dependence of $X$ and $Y$ around general statistical functionals $T_1(X)$ and $T_2(Y)$ such as quantiles, expectiles or thresholds. Thus, we strive at providing a more complete picture of the dependence structure. This widening of the perspective is akin to the methodological advancement in other key branches of statistics. In univariate statistics, statistical functionals constitute valuable summary measures complementing the mean. In regression analysis, the shift to modelling functionals of the conditional distribution other than the mean was initiated by the advent of quantile and expectile regression koenker1978, newey1987.

In Section (ref), we introduce generalised covariance. Replacing the mean by another functional necessitates a different, in general non-linear, measurement of deviation from the functional of interest to preserve the property that independence implies nullity. We replace the classical error or deviation from the mean, $X-\mu(X)$, by generalised errors, typically constructed via identification functions. Then, our new generalised covariances are defined as the expectation of the product of these generalised errors.

Next, we normalise the generalised covariance accordingly to arrive at our generalised correlation (Section (ref)). We show that the classical Cauchy--Schwarz normalisation employed in Pearson correlation should not be used here as it leads to serious attainability issues. Instead, we propose an alternative and natural normalisation via what we call the Fr\'echet--Hoeffding bounds, which are sharp by construction. This new normalisation can also straightforwardly be used for classical covariance, which leads to what we call mean correlation, an attainable version of Pearson correlation. The generalised correlations have favourable properties, allow for measuring new forms of dependence and thus for gaining additional insights about the dependence structure between $X$ and $Y$. In particular, quantile correlation and the closely related threshold correlation arise, which allow for measuring dependence locally around a pair of quantiles of $X$ and $Y$ or around any point in the codomain of $(X,Y)$. Quantile correlation can be viewed as the correlation analogue to quantile regression.

Considering all the information on local dependence jointly, that is, using the whole families of quantile or of threshold covariances and correlations, respectively, leads to what we call distributional covariances and correlations (Section (ref)). They are function-valued objects revealing the entire dependence structure, with one-to-one connections to the copula (for the former) and the joint distribution of $X$ and $Y$ (for the latter). Since they are normalised, attaining values in $[-1,1]$, their interpretation is a lot easier than interpreting copulas or joint distribution functions. Interestingly, a constant zero (one\,/\,minus one) of these distributional correlations implies independence (perfect positive\,/\,perfect negative dependence) of $X$ and $Y$. The distributional correlations can be interpreted as generalised correlations of the identity functional, constituting the correlation counterpart of distributional regression chernozhukov2013, kneib2021.

In Section (ref) we introduce tail correlations, which arise as limits of quantile correlations, and elaborate on their connection to the widely-used coefficients of tail dependence Coles1999, Joe2014. Since they are derived from correlations, they attain values between $-1$ and $1$. Our new measures coincide with the coefficients of tail dependence if they are positive, but nicely distinguish between different strengths of negative tail dependence and tail independence, whereas the latter are 0 in these cases. Thus, constituting one of the rare occasions for a Pareto improvement in statistical methodology, they should arguably supersede the traditional coefficients of tail dependence.

Section (ref) goes back to the classical task of dependence measures, summarising overall dependence in a single number, and elaborates on the idea of integrating over distributional correlations to construct such summary covariances and correlations. Strikingly, this recovers classical and Spearman covariance as well as mean and Spearman correlation as canonical special cases.

We elaborate on the finite sample counterparts of generalised covariances and correlations and and show that they constitute consistent estimators (Section (ref)). In Section (ref), we illustrate the usage of our newly introduced dependence measures on demographic data stemming from the Panel Study of Income Dynamics. The Appendix contains proofs and additional details on the relation of generalised errors and identification functions and on the empirical applications. We provide an R package accompanying this paper at \url{https://github.com/MarcPohle/GCor}.

Generalised covariances

Generalised errors

Let $(\Omega, \mathfrak{F}, \mathbb P)$ be a non-atomic probability space. Denote by $L^0(\mathbb{R}^d)$, $d=1,2$, the space of all $\mathbb{R}^d$-valued random variables. For $p\in[1,\infty)$, let $L^p(\mathbb{R}) = \{X\in L^0(\mathbb{R}) \mid \mathbb{E}[|X|^p]<\infty\}$ and $L^p(\mathbb{R}^2) = \{(X,Y) \in L^0(\mathbb{R}^2) \mid X, Y \in L^p(\mathbb{R})\}$. We consider statistical functionals $T$ as law-determined maps from some collection of random variables $\mathcal L\subseteq L^0(\mathbb{R})$ to a set $\mathsf{A}\subseteq \mathbb{R}$, meaning that for any $X,X'\in\mathcal{L}$ it holds that $T(X) = T(X')$ whenever $F_X=F_{X'}$, where $F_X,F_{X'}$ are the distribution functions of $X,X'$, respectively.

Generalised covariances and generalised correlations are law-determined maps from a class $\mathcal{D} \subseteq L^0(\mathbb{R}^2)$ of bivariate random vectors $(X,Y)$ to $\mathbb{R}$ and to $[-1,1]$, respectively. Recall the classical covariance and Pearson correlation on $\mathcal{D} = L^2(\mathbb{R}^2)$ from (ref). The rationale behind the definition of covariance is to measure average co-movements of $X$ and $Y$ around their respective means, $\mu(X)$ and $\mu(Y)$. The very idea behind generalised covariances is to measure average co-movements around functionals $T_1$ and $T_2$ other than the mean, e.g., around certain quantiles of $X$ and $Y$. A naive ansatz is to merely replace $\mu$ by $T$ in the definition of the covariance. However, this is not a suitable way to measure co-movements around arbitrary functionals. For example the fundamental property that independence of $X$ and $Y$ implies nullity of the generalised covariance would be violated. Covariance is constructed via deviations from the means of $X$ and $Y$, or errors, $X-\mu(X)$ and $Y-\mu(Y)$. We need to find a suitable way to measure deviations of a random variable $X$ from an arbitrary functional $T(X)$, which leads to the notion of generalised errors for $T$, capturing the most important properties of the prototypical error $X-\mu(X)$: having mean zero, being positive (negative) if $X$ realises above (below) $T(X)$, and being (weakly) larger the further $X$ realises away from $T(X)$.

definition[Generalised error] For a given functional $T\colon\mathcal{L}\to\mathsf{A}\subseteq\mathbb{R}$, we call a map $e_T\colon\mathcal{L}\to L^1(\mathbb{R})$ a generalised error for $T$ if the following properties hold for all $X\in\mathcal{L}$. \begin{enumerate}[(i)] • Centred: $\mathbb{E}[e_{T}(X)]=0$. • Increasing: for $\mathbb{P}\otimes \mathbb{P}$-almost all $(\omega, \omega')\in\Omega^2$ \[ X(\omega)\ge X(\omega') \implies e_{T}(X)(\omega)\ge e_{T}(X)(\omega'). \] • Sign change at $T$: for all $X\in\mathcal L$ and for $\mathbb{P}$-almost all $\omega\in\Omega$ \begin{equation} \big(X(\omega) - T(X)\big) e_{T}(X)(\omega)\ge0. \end{equation} \end{enumerate}

Further examples of generalised errors besides the prototypical $e_\mu(X) = X - \mu(X)$ are discussed in Subsection (ref). A natural way to construct a generalised error map is via so-called identification functions.

definition[Identification function] A map $v\colon \mathsf{A}\times\mathbb{R}\to\mathbb{R}$ is called $\mathcal{L}$-integrable if for all $t\in\mathsf{A}$ and $X\in\mathcal{L}$ it holds that $\mathbb{E}|v(t,X)|<\infty$. Moreover, $v$ is called increasing\,/\,non-constant if for any $t\in\mathsf{A}$ the map $x\mapsto v(t,x)$ is increasing\,/\,non-constant. An $\mathcal{L}$-integrable map $v\colon \mathsf{A}\times\mathbb{R}\to\mathbb{R}$ is an $\mathcal{L}$-identification function for a functional $T\colon\mathcal{L}\to\mathsf{A}\subseteq \mathbb{R}$ if $\mathbb{E}\big[v(T(X),X)\big]=0$ for all $X\in\mathcal{L}$. It is a strict $\mathcal{L}$-identification function if additionally \[ \mathbb{E}\big[v(t,X)\big]=0 \implies t=T(X) \] for all $t\in\mathsf{A}$ and for all $X\in\mathcal{L}$. $T$ is identifiable on $\mathcal{L}$ if there exists a strict $\mathcal{L}$-identification function for it.

In the field of forecast evaluation, identification functions are a central tool to assess forecast calibration NoldeZiegel2017,DimiPattonSchmidt2019. In econometrics, they are often known as moment functions, and they give rise to Z-estimation or the (generalised) method of moments estimation Huber1967, Hansen1982, NeweyMcFadden1994. An example for a strict $L^1(\mathbb{R})$-identification function for the mean, which induces the error for the mean $e_\mu(X)$ from above, is $v_\mu(t,x) = x-t$ (see again Subsection (ref) for further examples). The following proposition provides a recipe how to build generalised errors from identification functions, which will be the construction principle for all generalised errors in this paper but the ones for thresholds (Example (ref)) and quantiles in the non-continuous case (Example (ref)), where there is still a very close connection to identification functions. Assumption (ref) is spelled out in the Appendix.

propositionLet $v_T:\mathsf{A}\times \mathbb{R}\to\mathbb{R}$ be an increasing, non-constant $\mathcal{L}$-identification function for the functional $T:\mathcal{L}\to\mathsf{A}\subseteq \mathbb{R}$ satisfying Assumption (ref). Then, for a random variable $X\in\mathcal{L}$, the quantity \begin{equation} \omega\mapsto e_{v_T}(X)(\omega):= v_T\big(T(X),X(\omega)\big), \end{equation} is a generalised error of $X$ for $T$.

Usually, an estimator or forecast for the functional $T$ is plugged in as the first argument of the identification function. Just plugging in the true functional itself yields a generalised error.

Definition and properties

definition[Generalised covariance] Let $T_1\colon\mathcal{L}_1\to\mathsf{A}_1\subseteq\mathbb{R}$, $T_2\colon\mathcal{L}_2\to\mathsf{A}_2\subseteq\mathbb{R}$ be two functionals and $e_{T_1}$ and $e_{T_2}$ generalised errors for these functionals. Let $\mathcal{D} = \{(X,Y) \in L^0(\mathbb{R}^2)\,|\,X \in\mathcal{L}_1,\ Y \in \mathcal{L}_2,\ e_{T_1}(X) e_{T_2}(Y) \in L^1(\mathbb{R})\}$. Then, the generalised covariance at $T_1$ and $T_2$ induced by $e_{T_1}$ and $e_{T_2}$, or the $T_1-T_2$-covariance induced by $e_{T_1}$ and $e_{T_2}$, is defined on $\mathcal{D}$ via \begin{equation} \mathrm{Cov}_{T_1, T_2}(X,Y) := \mathbb{E}\big[e_{T_1}(X) e_{T_2}(Y)\big]. \end{equation}

A classical sufficient condition for the integrability of the product $e_{T_1}(X) e_{T_2}(Y)$ is that the factors are square integrable, exploiting the Cauchy--Schwarz inequality. An alternative and weaker condition is provided in Proposition (ref).

Since the generalised errors are centred by definition, $\mathbb{E}\big[e_{T_1}(X)\big] = \mathbb{E}\big[e_{T_2}(Y)\big]=0$, independence implies nullity.

propositionFor the generalised covariance from (ref) it holds that $\mathrm{Cov}_{T_1, T_2}(X,Y) = 0$ if $X$ and $Y$ are independent.
remarkThe fact that the error terms are centred also implies that the generalised covariance can equivalently be written as the covariance of the generalised errors. That is, \begin{equation} \mathrm{Cov}_{T_1, T_2}(X,Y) = \mathrm{Cov}\big(e_{T_1}(X), e_{T_2}(Y)\big). \end{equation}

This nicely illustrates the rationale of generalised covariances at $T_1$ and $T_2$, measuring average co-movements of $X$ and $Y$ around their respective reference functionals. An increasing likelihood of joint positive deviations of $X$ and $Y$ from $T_1(X)$ and $T_2(Y)$ leads to an increase in $\mathrm{Cov}_{T_1, T_2}(X,Y)$. On the other hand, an increasing likelihood of countermovements decreases this covariance. Moreover, thanks to the errors being increasing, the value of the covariance is also sensitive to the magnitude of the deviations of $X$ and $Y$ from their reference functionals. Finally, if there is no systematic mutual influence between $X$ and $Y$, i.e., they are independent, the covariance vanishes.

Obviously, $\mathrm{Cov}_{T_1, T_2}(X,Y)$ depends on the choice of the generalised errors, $e_{T_1}$, $e_{T_2}$. For the leading situation when the generalised error is induced by an identification function (see Proposition (ref)), we characterise this dependence in Section (ref) of the Appendix (Remark (ref)) and remark that generalised correlations are actually independent of the choice of the identification function, subject to regularity conditions (Proposition (ref)). For the examples discussed in the following subsection we utilise the canonical identification functions as suggested by GneitingResin2021.

Examples

Some of the examples of generalised covariance we discuss here have appeared in the literature. We discuss relations to the literature in Subsection (ref) when introducing the respective generalised correlations.

example[Mean and expectile covariance] The mean has an increasing, non-constant strict $L^1(\mathbb{R})$-identification function \( v_\mu(t,x) := x-t, \ x,t\in\mathbb{R}. \) The induced error (ref) leads to the classical covariance (ref). Likewise, its asymmetric version, the $\tau$-expectile, admits an increasing, non-constant strict $L^1(\mathbb{R})$-identification function \begin{equation} v_{\mu_\tau}(t,x) := 2|\mathds{1}\{x\le t\} - \tau|(x-t), \qquad x,t\in\mathbb{R}, \end{equation} where $\tau\in(0,1)$. Clearly, for $\tau = 1/2$, this recovers the case of the mean. The induced expectile covariance at levels $\tau,\eta\in(0,1)$ is \begin{align} \mathrm{ECov}_{\tau,\eta}(X,Y) &:= \mathrm{Cov}_{\mu_{\tau}, \mu_{\eta}}(X,Y)\\ \nonumber &= 4 \mathbb{E}\big[|\mathds{1}\{X\le \mu_{\tau}(X)\} - \tau|(X-\mu_{\tau}(X)) |\mathds{1}\{Y\le \mu_{\eta}(Y)\} - \eta|(Y-\mu_{\eta}(Y)) \big], \end{align} where $X,Y, XY \in L^1(\mathbb{R})$. Clearly, for $\tau = \eta = 1/2$, one recovers the usual covariance.

Just as the mean, the $\tau$-expectile is translation equivariant in the sense that $\mu_\tau(X + c) = \mu_\tau(X)+c$ for all $X\in L^1(\mathbb{R})$ and $c\in\mathbb{R}$, and it is positively homogeneous, i.e., $\mu_\tau(\lambda X) = \lambda \mu_\tau(X)$ for all $X\in L^1(\mathbb{R})$ and $\lambda>0$. The canonical expectile identification function (ref) shares similar properties: It is positively homogeneous and translation invariant. Hence, the expectile covariance is also translation invariant and positively homogeneous in both arguments.

propositionFor all $\tau,\eta\in(0,1)$, for all $X,Y\in L^1(\mathbb{R})$ such that $XY\in L^1(\mathbb{R})$, for all $c\in\mathbb{R}$ and $\lambda>0$ it holds that \begin{align*} \mathrm{ECov}_{\tau,\eta}(X + c,Y) &= \mathrm{ECov}_{\tau,\eta}(X,Y + c) = \mathrm{ECov}_{\tau,\eta}(X,Y),\\ \mathrm{ECov}_{\tau,\eta}(\lambda X,Y) &=\mathrm{ECov}_{\tau,\eta}( X,\lambda Y) = \lambda \mathrm{ECov}_{\tau,\eta}( X,Y) . \end{align*}
example[Threshold covariance] The arguably simplest situation is to consider dependence around a point $(a,b) \in \mathbb{R}^2$, that is, to measure the average joint deviation of $X$ from an absolute threshold $a\in\mathbb{R}$ and of $Y$ from $b\in\mathbb{R}$. This requires a special treatment since, formally, the functionals are constant. As such, they are identifiable, but the identification function for the constant $a$, $v_a(t,x) = t-a$, is constant in its second argument $x$. Thus, it would only yield a trivial generalised error, which would be constant 0. To circumvent this problem, consider \begin{equation} e_a(X) = F_X(a) - \mathds{1}\{X \le a\}, \end{equation} which is indeed for all $X\in L^0(\mathbb{R})$ a generalised error for the constant functional $a$. Hence, we can define the threshold covariance at points $a,b\in\mathbb{R}$ \begin{align} \mathrm{TCov}_{a,b}(X,Y) := \mathbb{E}\big[\left( F_X(a) - \mathds{1} \{ X \leq a\} \right) \left( F_Y(b) - \mathds{1} \{ Y \leq b\} \right)\big] = F_{X,Y}(a,b) - F_X(a)F_Y(b). \end{align} Interestingly, the generalised error in (ref) can also arise via (ref) for the evaluation functional $T_a(F):= F(a)$ with the increasing, non-constant strict $L^0(\mathbb{R})$-identification function \begin{equation*} v_{T_a} (t,x) := t - \mathds{1}\{x \le a\}, \qquad x,t\in\mathbb{R}. \end{equation*}
example[Quantile covariance, continuous case] For the (lower) $\alpha$-quantile $q_\alpha(X) = \inf\{x\in\mathbb{R}\,|\, F_X(x)\ge\alpha\}$, $X\in L^0(\mathbb{R})$, $\alpha\in(0,1)$, the function \begin{equation} v_{q_\alpha}(t,x) := \alpha - \mathds{1}\{x\le t\}, \qquad x,t\in\mathbb{R}, \end{equation} is an increasing, non-constant $L_\alpha$-identification function, where $L_\alpha = \{X\in L^0(\mathbb{R}) \,|\,F_X(q_\alpha(X))=\alpha\}$.\footnote{$L_\alpha$ is a superclass of random variables with continuous distributions. On the subclass $L_{(\alpha)} = \{X\in L_\alpha\,|\, F_X\big(q_\alpha(X)+\varepsilon\big)>\alpha\ \forall \varepsilon>0\}$, $v_{q_\alpha}$ in (ref) is also a strict identification function for $q_\alpha$.} The induced quantile covariance at levels $\alpha,\beta\in(0,1)$ is \begin{align} \mathrm{QCov}_{\alpha, \beta}(X,Y) &:= \mathrm{Cov}_{q_{\alpha}, q_{\beta}}(X,Y) = \mathbb{E}\big[ (\alpha - \mathds{1} \{X\leq q_{\alpha}(X)\}) (\beta - \mathds{1} \{Y\leq q_{\beta}(Y)\}) \big]. \end{align} If $X\in L_\alpha$ and $Y\in L_\beta$, we get that \begin{equation} \mathrm{QCov}_{\alpha, \beta}(X,Y) = C_{X,Y}(\alpha,\beta) - \alpha\beta, \end{equation} where $C_{X,Y}$ is a copula of $(X,Y)$. Since all copulas for $(X,Y)$ coincide on $\text{range}(F_X) \times \text{range}(F_Y)$ and $F_X(q_\alpha(X))=\alpha$, $F_Y(q_\beta(Y))=\beta$ by assumption, the expression (ref) is well defined, i.e., independent of the choice of the copula. Again, it is convenient that the quantile identification function (ref) is bounded in $x$ such that we can dispense with integrability assumptions on $X$ and $Y$.
example[Quantile covariance, general case] On the entire $L^0(\mathbb{R})$, $v_{q_\alpha}$ in (ref) fails to identify the $\alpha$-quantile, due to possible discontinuities in the cumulative distribution function (CDF). However, a natural modification of the generalised error from the continuous case leads to a suitable error for the general case: \begin{equation} e_{q_\alpha}(X) := F_X( q_\alpha (X) ) - \mathds{1}\{X \le q_\alpha(X)\}. \end{equation} This can be seen as a correction of the generalised error from the continuous case, replacing the quantile level $\alpha$ with the corrected quantile level $\alpha^*:=F_X(q_\alpha(X))$ accounting for a jump in the CDF and ensuring that $e_{q_\alpha}(X)$ is centred. On the other hand, it is just the natural analogue to the threshold error (ref). This leads to the general definition of quantile covariance at levels $\alpha,\beta\in(0,1)$ as \begin{align} \mathrm{QCov}_{\alpha, \beta}(X,Y) = \mathbb{E}\big[ \big(F_X(q_\alpha(X)) - \mathds{1} \{X\leq q_{\alpha}(X)\}\big) \big(F_Y(q_\beta(Y)) - \mathds{1} \{Y\leq q_{\beta}(Y)\}\big) \big]. \end{align} Similar to (ref), (ref) can also be expressed in terms of a copula of $C_{X,Y}$ of $(X,Y)$: \begin{equation} \mathrm{QCov}_{\alpha, \beta}(X,Y) = C_{X,Y}\big(F_X(q_\alpha(X)), F_Y(q_\beta(Y))\big) - F_X(q_\alpha(X)) F_Y(q_\beta(Y)). \end{equation} Invoking the same arguments as above, this expression does not depend on the choice of the copula.

The quantile enjoys even more invariance properties than the expectile. It is equivariant under all strictly increasing transformations: $q_\alpha\big(g(X)\big) = g\big(q_\alpha(X)\big)$ for any $X\in L^0(\mathbb{R})$ and for any strictly increasing $g\colon\mathbb{R}\to\mathbb{R}$. The quantile error (ref) inherits this invariance, $e_{q_{\alpha}} (g(X)) = e_{q_{\alpha}} (X)$. Hence, the induced quantile covariance is also invariant with respect to strictly increasing transformations in both arguments.

propositionFor all $\alpha,\beta\in(0,1)$, for all $X,Y\in L^0(\mathbb{R})$, and for all strictly increasing transformations $g\colon\mathbb{R}\to\mathbb{R}$ it holds that \begin{align*} \mathrm{QCov}_{\alpha,\beta}\big(g(X),Y\big) &= \mathrm{QCov}_{\alpha,\beta}\big(X,g(Y)\big) = \mathrm{QCov}_{\alpha,\beta}(X,Y) . \end{align*}
remark[Local covariances] Threshold and quantile covariance are closely connected and complementary. Indeed, if $(a,b) = (q_{\alpha}(X),q_{\beta}(Y))$, the two measures coincide, $\mathrm{TCov}_{a,b}(X,Y)=\mathrm{QCov}_{\alpha, \beta}(X,Y)$, as the generalised errors coincide, $e_{a}(X)=e_{q_{\alpha}}(X)$ and $e_{b}(Y)=e_{q_{\beta}}(Y)$. What distinguishes them is that threshold covariance is a measure on the observation scale, i.e., one chooses a point $(a,b) \in \mathbb{R}^2$, while quantile covariance measures dependence on the quantile scale, i.e., one chooses two quantile levels $(\alpha,\beta) \in (0,1)^2$. If one choose a point $(a,b)$, computes the respective quantile levels $F_X(a)$ and $F_Y(b)$ to obtain the quantile covariance at this point, this just leads to the threshold covariance at this point, $\mathrm{QCov}_{F_X(a),F_Y(b)} (X,Y) = \mathrm{TCov}_{a,b}(X,Y)$ and vice versa, $\mathrm{TCov}_{q_{\alpha}(X),q_{\beta}(Y)}(X,Y)=\mathrm{QCov}_{\alpha,\beta} (X,Y)$. Both measures allow to measure dependence locally around the specified point: They only depend on the joint exceedances of this point, or in other words, the joint CDF $F_{X,Y}$ or copula $C_{X,Y}$ and the marginal distributions $F_X$, $F_Y$ there. Thus, the corresponding generalised errors are naturally binary random variables, amounting to centred exceedance indicators.
example[Quantile-mean covariance] We can also pair functionals $T_1$ and $T_2$ which belong to different families. E.g., we can define the quantile-mean covariance as \begin{equation} \mathrm{Cov}_{q_\alpha,\mu}(X,Y) := \mathbb{E}\big[(\alpha - \mathds{1}\{X \le q_\alpha(X)\})(Y - \mu(Y))\big] = - \mathbb{E}\big[F_{X\,\mid Y}(q_\alpha(X))(Y - \mu(Y))\big], \end{equation} where $\alpha\in(0,1)$, $X\in L^0(\mathbb{R})$ and $Y\in L^1(\mathbb{R})$. The second identity is due to the tower property of the conditional expectation, where one first conditions on $Y$. Conveniently, we may use this definition even if $F_X$ has a jump at its $\alpha$-quantile such that the generalised error for the quantile fails to be centred. The reason is that the generalised error for the mean is always centred.
remarkNaturally, generalised covariances involving the quantile error $e_{q_\alpha}(X)$ (ref) are only interesting for $\alpha$ such that $q_\alpha(X) < \mathrm{ess\,sup}(X) = q_1(X)$. For larger levels $\alpha$, $X$ cannot vary around $q_\alpha(X)$, but $X\le q_\alpha(X)$, implying that $e_{q_\alpha}(X)$ is constant and the quantile covariance is 0 for such situations.

Generalised correlations

Normalisation: Fr\'echet--Hoeffding vs. Cauchy--Schwarz

To turn a covariance into a correlation, it is essential to ensure normalisation of the correlation meaning that it only attains values in $[-1,1]$. Normalisation enhances the interpretability of a correlation. It can be achieved by bounding the covariance in absolute values with a constant $K_{X,Y}$, which may depend on the marginal distributions of $X$ and $Y$

equation[equation omitted — 75 chars of source]

Then, trivially $\mathrm{Cov}_{T_1,T_2}(X,Y)/ K_{X,Y}\in[-1,1]$. Pearson correlation $r(X,Y)$ (ref) builds on this idea and exploits the Cauchy--Schwarz inequality to bound the covariance. That is, for $X,Y\in L^2(\mathbb{R})$,

equation[equation omitted — 104 chars of source]

For the generalised covariance, we could hence simply exploit the identity (ref) and obtain

equation[equation omitted — 146 chars of source]

provided that $e_{T_1}(X), e_{T_2}(Y) \in L^2(\mathbb{R})$. The implied quantity is just the Pearson correlation of the generalised errors

equation[equation omitted — 209 chars of source]

However, it is well-known that the Cauchy--Schwarz inequality (ref) is not sharp in general. Consequently, the Pearson correlation of the generalised errors (ref) does not always attain all values in $[-1,1]$ for given marginal distributions of $X$ and $Y$. In particular, its minimum and maximum may be far away from $-1$ and 1, respectively, and their absolute values may differ strongly, depending on the specific marginal distributions of $X$ and $Y$, see Embrechts2002. This compromises its interpretability in that it is not really able to indicate the strength of dependence (by the closeness of its absolute value to 1). In fact, a value close to 0 may arise even if the dependence is quite strong for certain marginal distributions.

We revisit the conditions for attainability of Pearson correlation more closely now. For $X,Y\in L^2(\mathbb{R})$, $r(X,Y)$ is $1$ ($-1$) if and only if $X$ and $Y$ have perfect positive (negative) linear dependence. That is, if and only if for some $a, a'\in\mathbb{R}$, $b, b'>0$ it holds almost surely that $Y = a + bX$ ($Y = a' - b'X$). Therefore, for given non-degenerate marginals $F_X$ and $F_Y$, there exists a joint distribution $F_{X,Y}$ with Pearson correlation $1$ ($-1$) if and only if $Y$ and $X$ ($-X$) are of the same type, meaning that $Y \stackrel{\mathrm{d}}{=} a + bX$ ($Y \stackrel{\mathrm{d}}{=} a' - b'X$) for some $a, a'\in\mathbb{R}$, $b, b'>0$. This observation leads to the following insight.

lemma[Attainability of Pearson correlation] Let $X,Y\in L^2(\mathbb{R})$ be non-constant. Pearson correlation $\mathrm{Cor}(X,Y)$ is attainable, that is, there exist joint distributions $F$, $\tilde F$ with marginals $F_X$, $F_Y$ and Pearson correlations $1$ and $-1$, respectively, if and only if $X$ and $Y$ are of the same type and the distributions are symmetric, that is, there exist $c,d\in\mathbb{R}$ such that $X-c \stackrel{\mathrm{d}}{=} - (X-c)$ and $Y-d \stackrel{\mathrm{d}}{=} - (Y-d)$.

The restriction to symmetric distributions of the same type is substantial beyond the confinements of a Gaussian world. In the context of generalised covariances, which can be written as covariances of generalised errors (see (ref)), it is definitely too restrictive. Here, the distributions of the generalised errors $e_{T_1}(X)$ and $e_{T_2}(Y)$ are generally neither symmetric nor of the same type.

exampleAn example illustrates how severe the attainability problem can become. For the Pearson correlation (ref) of the quantile errors in the continuous case for $\alpha, \beta \in (0,1)$ from Example (ref) the upper and lower bound do not even depend on the marginal distributions of $X$ and $Y$, but only on the quantile levels $\alpha$ and $\beta$ as will become clear in Example (ref). Figure (ref) presents the bounds for all $(\alpha,\beta) \in (0,1)$. This quantity can only attain 1 if $\alpha=\beta$, and $-1$ if $\alpha=1-\beta$, implying that it is only attainable at the medians, $\alpha=\beta=0.5$. What is more, the upper and lower bounds can be very far away from 1 and $-1$, and their absolute values are usually very different from each other.
figure[figure omitted — 230 chars of source]

The following proposition provides a sharp version of the inequality (ref), exploiting previous results by hoeffding1940, Frechet1957, and Embrechts2002. Recall that $(X,Y)$ are called comonotonic if $(X,Y)\stackrel{\mathrm{d}}{=} \big(\nu_1(Z), \nu_2(Z)\big)$ for some random variable $Z$ and two increasing functions $\nu_1$, $\nu_2$. Similarly, $(X,Y)$ are countermonotonic if $(X,Y)\stackrel{\mathrm{d}}{=} \big(\nu_1(Z), \nu_2(Z)\big)$ for some random variable $Z$ with $\nu_1$ increasing and $\nu_2$ decreasing. In other words, comonotonicity (countermonotonicity) corresponds to the situation of perfect positive (negative) dependence between $X$ and $Y$.

propositionFor any pair of random variables $(X,Y)$, let $(X,Y')$ and $(X,Y'')$ be pairs with the same marginal distributions such that $(X,Y')$ is countermonotonic and $(X,Y'')$ is comonotonic. Let $T_1$ and $T_2$ be functionals with generalised errors $e_{T_1}$ and $e_{T_2}$ such that $\mathrm{Cov}_{T_1,T_2}(X,Y')$ and $\mathrm{Cov}_{T_1,T_2}(X,Y'')$ exist and are finite. Then the following holds. \begin{enumerate}[(i)] • $\mathrm{Cov}_{T_1,T_2}(X,Y)$ is finite and \begin{equation} \mathrm{Cov}_{T_1,T_2}(X,Y') \le \mathrm{Cov}_{T_1,T_2}(X,Y) \le \mathrm{Cov}_{T_1,T_2}(X,Y”). \end{equation} • If $e_{T_1}(X)$ and $e_{T_2}(Y)$ are non-constant, then \begin{equation} \mathrm{Cov}_{T_1,T_2}(X,Y') <0 < \mathrm{Cov}_{T_1,T_2}(X,Y”). \end{equation} • If $e_{T_1}(X)$ is a strictly increasing function of $X$ and $e_{T_2}(Y)$ is a strictly increasing function of $Y$, then an equality in (ref) is attained only if $(X,Y)$ is co- or countermonotonic. \end{enumerate}
remarkPrime examples where the assumption that the generalised errors are strictly increasing functions of $X$ and $Y$, respectively, from part (iii) of Proposition (ref) is violated are the local covariances discussed in and before Remark (ref), where the generalised errors are binary. Due to their local nature, they can and should not determine global properties such as co- and countermonotonicity. In fact, for them equality in (ref) holds under perfect local dependence, that is, if the corresponding exceedance indicators (or equivalently the generalised errors) are perfectly dependent.
example[Perfect local dependence] Suppose $X$ follows a uniform distribution on $[0,1]$ and let $Y = 1/2-X$ for $X\in[0,1/2]$ and $Y = 3/2 - X$ for $X\in(1/2,1]$. Then $Y$ is also uniformly distributed on $[0,1]$. The pair $(X,Y)$ is neither co- nor countermonotonic. However, if $T_1$ and $T_2$ are the median and if we use the generalised errors induced by (ref), then $\big(e_{q_{1/2}}(X) , e_{q_{1/2}}(Y)\big)$ only attains the values $(-1/2,-1/2)$ and $(1/2, 1/2)$. Hence, the generalised errors are comonotonic, inducing perfect positive dependence locally around the medians such that $\mathrm{Cov}_{q_{1/2},q_{1/2}}(X,Y) = \mathrm{Cov}_{q_{1/2},q_{1/2}}(X,Y'')$.

Definition and properties

Combining generalised covariance from Definition (ref) with the Fr\'echet--Hoeffding normalisation from Proposition (ref) leads to generalised correlation.

definition[Generalised correlation] Let $T_1\colon\mathcal{L}_1\to\mathsf{A}_1\subseteq\mathbb{R}$, $T_2\colon\mathcal{L}_2\to\mathsf{A}_2\subseteq\mathbb{R}$ be two functionals and $e_{{T_1}}, e_{{T_2}}$ generalised errors for $T_1$ and $T_2$. Let $\mathcal{D}$ be the set of random variables $(X,Y) \in L^0(\mathbb{R}^2)$ such that $e_{{T_1}}(X) e_{{T_2}}(Y')\in L^1(\mathbb{R})$ and $e_{{T_1}}(X) e_{{T_2}}(Y'')\in L^1(\mathbb{R})$, where $(X,Y')$ and $(X,Y'')$ are pairs with the same marginal distributions such that $(X,Y')$ is countermonotonic and $(X,Y'')$ is comonotonic. If neither $e_{{T_1}}(X)$ nor $e_{{T_2}}(Y)$ are constant almost surely, the generalised correlation at $T_1$ and $T_2$ induced by $e_{{T_1}}$ and $e_{{T_2}}$, or the $T_1-T_2$-correlation induced by $e_{{T_1}}$ and $e_{{T_2}}$, is defined on $\mathcal{D}$ via \begin{equation} \mathrm{Cor}_{T_1, T_2}(X,Y) := \begin{cases} \frac{\mathrm{Cov}_{T_1, T_2}(X,Y)}{|\mathrm{Cov}_{T_1, T_2}(X,Y'\,)|}, &if \mathrm{Cov}_{T_1, T_2}(X,Y)<0, \\[0.5em] \frac{\mathrm{Cov}_{T_1, T_2}(X,Y)}{|\mathrm{Cov}_{T_1, T_2}(X,Y''\,)|}, &\text{if } \mathrm{Cov}_{T_1, T_2}(X,Y)\ge0. \end{cases} \end{equation} If one of the generalised errors is constant almost surely, so in particular if $X$ or $Y$ is constant, then \[ \mathrm{Cor}_{T_1, T_2}(X,Y) :=0. \]

We summarise the most important properties of generalised correlation.

theorem[Properties of generalised correlation] The generalised correlation at $T_1$ and $T_2$ induced by $e_{{T_1}}$ and $e_{{T_2}}$ satisfies the following properties. \begin{enumerate}[(i)] • Normalisation: $\mathrm{Cor}_{T_1, T_2}(X,Y) \in [-1,1]$. • Independence implies nullity: $\mathrm{Cor}_{T_1, T_2}(X,Y) =0$ if $X$ and $Y$ are independent. • Perfect dependence: Suppose $e_{T_1}(X)$ and $e_{T_2}(Y)$ are not constant almost surely. \begin{enumerate} • $\mathrm{Cor}_{T_1, T_2}(X,Y) =1(-1)$ if $X$ and $Y$ are comonotonic (countermonotonic). • If $e_{T_1}(X)$ is a strictly increasing function of $X$ and $e_{T_2}(Y)$ is a strictly increasing function of $Y$, then $\mathrm{Cor}_{T_1, T_2}(X,Y) =1(-1)$ implies that $X$ and $Y$ are comonotonic (countermonotonic). \end{enumerate} • Symmetry: It holds that $\mathrm{Cor}_{T_1, T_2}(X,Y) = \mathrm{Cor}_{T_2, T_1}(Y,X)$. In particular, if $T_1 = T_2$, then $\mathrm{Cor}_{T_1, T_2}(X,Y) = \mathrm{Cor}_{T_1, T_2}(Y,X)$. \end{enumerate}

Properties (i), (ii) and part (a) of (iii) are fundamental properties that every correlation-type dependence measure should fulfil. They ensure interpretability in that the measure takes the right values in the extreme cases of independence and perfect positive and negative dependence. Part (b) of (iii) is not always desirable, e.g., when it comes to local dependence measures, see Remark (ref).

We would like to highlight the novelty of normalising covariances with Fr\'echet--Hoeffding bounds, which are sharp by construction. We are aware of similar constructions only in the context of a dependence measure for binary random variables cole1949 and to combat the non-attainability of Spearman's $\rho$ and Kendall's $\tau$ in the discrete case VandenhendeLambert2003, genest2007.

Examples

We are going to review the examples of generalised covariances from Subsection (ref) and see how they translate into generalised correlations. To that end, we will mainly focus on the normalisation terms, that is, the denominators in (ref). Throughout this section, let again $(X,Y)$, $(X,Y')$, $(X,Y'')$ be pairs with the same marginal distribution, where $(X,Y')$ is countermonotonic and $(X,Y'')$ is comonotonic.

example[Mean correlation] We write $\mathrm{MCor}(X,Y):= \mathrm{Cor}_{\mu,\mu} (X,Y)$ for mean correlation. Due to Hoeffding's formula mcneil2015, it holds that \begin{align*} \mathrm{Cov}(X,Y')&= \iint \max \big(F_X(z_1) + F_Y(z_2) - 1,0\big) - F_X(z_1)F_Y(z_2)\,\mathrm{d} z_1 \mathrm{d} z_2, \\ \mathrm{Cov}(X,Y”)&= \iint \min \big(F_X(z_1), F_Y(z_2)\big) - F_X(z_1)F_Y(z_2)\,\mathrm{d} z_1 \mathrm{d} z_2. \end{align*} In particular, for the special case that the full range $[-1,1]$ is attainable by Pearson correlation, that is, if $X$ and $Y$ are of the same type and symmetric (Lemma (ref)), e.g., under bivariate normality, mean correlation coincides with Pearson correlation, and $|\mathrm{Cov}(X,Y')| = |\mathrm{Cov}(X,Y'')| = \sqrt{\mathrm{Var}(X)\mathrm{Var}(Y)}$, provided that $X,Y\in L^2(\mathbb{R})$. Thus, indeed Pearson correlation arises as a special case of generalised correlation. Mean correlation, in turn, can be viewed as an improved version of Pearson correlation, which solves the attainability problem.
example[Expectile correlation] $\mathrm{ECov}_{\tau,\eta}(X,Y')$ and $\mathrm{ECov}_{\tau,\eta}(X,Y'')$ can be calculated via (ref), i.e., the representation of generalised covariance as covariance of generalised errors, and again Hoeffding's formula.

It follows from Proposition (ref) that expectile correlation and hence in particular mean correlation is invariant under linear transformations.

propositionFor all $\tau,\eta\in(0,1)$, for all $X,Y\in L^1(\mathbb{R})$ such that $XY', XY'\in L^1(\mathbb{R})$, for all $c\in\mathbb{R}$ and $\lambda>0$ it holds that \begin{align*} \mathrm{ECor}_{\tau,\eta}(\lambda X + c,Y) &= \mathrm{ECor}_{\tau,\eta}(X,\lambda Y + c) = \mathrm{ECor}_{\tau,\eta}(X,Y). \end{align*}
example[Threshold correlation] For the threshold correlation, the classical Fr\'echet--Hoeffding bounds for joint CDFs arise as normalisations, that is, for $a,b\in\mathbb{R}$ \begin{align*} \mathrm{TCov}_{a,b}(X,Y') &= \max\big(F_X(a) + F_Y(b) - 1,0\big) - F_X(a)F_Y(b),\\ \mathrm{TCov}_{a,b}(X,Y”) &= \min\big(F_X(a), F_Y(b)\big) - F_X(a)F_Y(b). \end{align*}
example[Quantile correlation] For $\alpha,\beta\in(0,1)$ and $X \in L_\alpha$, $Y\in L_\beta$ (e.g., when $F_X$ and $F_Y$ are continuous), the Fr\'echet--Hoeffding bounds for copulas arise as normalising terms, that is, for $\alpha,\beta\in(0,1)$ \begin{align*} \mathrm{QCov}_{\alpha,\beta}(X,Y') = \max (\alpha + \beta - 1,0 ) - \alpha\beta, \qquad \mathrm{QCov}_{\alpha,\beta}(X,Y”) = \min ( \alpha, \beta ) - \alpha\beta. \end{align*} If $F_X$ and $F_Y$ possibly have jumps at their respective $\alpha$- and $\beta$-quantiles, applying the Fr\'echet--Hoeffding bounds to (ref) yields \begin{align*} \mathrm{QCov}_{\alpha,\beta}(X,Y') &= \max \big(F_X(q_\alpha(X)) + F_Y(q_\beta(Y)) - 1,0 \big) - F_X(q_\alpha(X))F_Y(q_\beta(Y)), \\ \mathrm{QCov}_{\alpha,\beta}(X,Y”) &= \min \big( F_X(q_\alpha(X)),F_Y(q_\beta(Y)) \big) - F_X(q_\alpha(X))F_Y(q_\beta(Y)). \end{align*}

Quantile correlation is invariant with respect to strictly increasing transformations, which follows from Proposition (ref). Thus, quantile correlation belongs to the family of rank correlations.

propositionFor all $\alpha,\beta\in(0,1)$, for all $X,Y\in L^0(\mathbb{R})$, and for all strictly increasing transformations $g\colon\mathbb{R}\to\mathbb{R}$ it holds that \begin{align*} \mathrm{QCor}_{\alpha,\beta}\big(g(X),Y\big) &= \mathrm{QCor}_{\alpha,\beta}\big(X,g(Y)\big) = \mathrm{QCor}_{\alpha,\beta}(X,Y) . \end{align*}
example[Median correlation] Median correlation, $\mathrm{QCor}_{0.5,0.5}$, arises as a special case of quantile correlation. What is particular about it is that, for continuous marginals $F_X$ and $F_Y$, the generalised errors $e_{q_{0.5}}(X)$ and $e_{q_{0.5}}(Y)$ are symmetric and of the same type. Hence, they fulfil the conditions of Lemma (ref), implying that the Cauchy--Schwarz and the Fr\'echet--Hoeffding normalisation coincide, $\mathrm{QCov}_{\alpha,\beta}(X,Y')=-\mathrm{QCov}_{\alpha,\beta}(X,Y'')=\sqrt{\mathrm{Var}(e_{q_{0.5}}(X))\mathrm{Var}(e_{q_{0.5}}(Y))}=1/4$. Indeed, then median correlation is equal to Blomqvist1950's (Blomqvist1950) $\beta$, defined as $\beta(X,Y) := 4 \mathbb{P} (X \leq q_{0.5}(X), Y \leq q_{0.5}(Y)) - 1$, which is also sometimes referred to as median correlation in the literature.

Quantile correlation does not only generalise median correlation. When moving from the center to the tails, that is, considering the limit of $\mathrm{QCor}_{\alpha,\alpha}$ for $\alpha \to 0$ or $\alpha \to 1$, a quantity closely related to the well-known coefficient of tail dependence shows up which is discussed in Section (ref).

example[Quantile-mean correlation] Again, $\mathrm{Cov}_{q_\alpha,\mu}(X,Y')$ and $\mathrm{Cov}_{q_\alpha,\mu}(X,Y')$ can be calculated via (ref) and Hoeffding's formula. If $X$ and $Y$ have continuous and strictly increasing marginal distributions, they take a particularly convenient form: \begin{align*} \mathrm{Cov}_{q_\alpha,\mu}(X,Y') = \mathbb{E}\big[(\alpha - \mathds{1}\{X \le q_\alpha(X)\})(Y' - \mu(Y'))\big] = \alpha\big(\mu(Y) - \mathrm{ES}_{1-\alpha}^+(Y)\big)<0, \end{align*} where $\mathrm{ES}_{1-\alpha}^+(Y)$ is the upper expected shortfall $\mathrm{ES}_{1-\alpha}^+(Y) := \mathbb{E}[Y|Y> q_{1-\alpha}(Y)] = \frac{1}{\alpha}\mathbb{E}[Y \mathds{1}\{Y > q_{1-\alpha}(Y)\}]$, and \begin{align*} \mathrm{Cov}_{q_\alpha,\mu}(X,Y”) = \mathbb{E}\big[(\alpha - \mathds{1}\{X \le q_\alpha(X)\})(Y” - \mu(Y”))\big] = \alpha\big(\mu(Y) - \mathrm{ES}_\alpha^-(Y)\big) >0, \end{align*} where $\mathrm{ES}_\alpha^-(Y)$ is the lower expected shortfall $\mathrm{ES}_\alpha^-(Y) := \mathbb{E}[Y|Y\le q_\alpha(Y)] = \frac{1}{\alpha}\mathbb{E}[Y \mathds{1}\{Y \le q_\alpha(Y)\}]$.

There are some measures in the literature related to quantile, quantile-mean and threshold covariances and correlations. Linton2007 employ a closely related quantity in their quantilogram, namely our quantile covariance (see Example (ref)), but normalised with the Cauchy--Schwarz normalisation, see also Han2016. Similarly, what Li2015 call quantile correlation amounts to our quantile-mean covariance (see Example (ref)), but again normalised with the Cauchy--Schwarz normalisation. Further, so-called indicator covariances from spatial statistics Dubrule2017 are closely related to our threshold covariances (see Example (ref)).

As discussed in Remark (ref), threshold and quantile covariance can be seen as local measures of dependence around a certain point. Accordingly, we call threshold and quantile correlation local correlations. Surprisingly, such local dependence measures have hardly been touched upon in the literature. Noteworthy exceptions are the local dependence function of Holland1987, see also Jones1996, and the local Gaussian correlation of Tjostheim2013.

Distributional covariances and correlations

Definition and properties

The generalised covariances and correlations presented so far measure the dependence of two random variables $X$, $Y$ around the two functionals $T_1(X)$, $T_2(Y)$. As such, they focus on a certain aspect of the dependence structure of $X$ and $Y$. This section proposes measures that uncover the entire dependence structure of $X$ and $Y$. The idea behind those measures is to consider all the local information contained in the whole families of local covariances (correlations) jointly (see Remark (ref)), leading to dependence measures that are functions in two arguments -- the thresholds $(a,b) \in \mathbb{R}^2$ or quantile levels $(\alpha,\beta) \in (0,1)^2$.

An alternative way to arrive at the same measures is to approach them from the angle of generalised covariances and correlations by not considering point-valued functionals $T_1(X)$ and $T_2(Y)$ as in Sections (ref) and (ref), but the entire CDF or the quantile function themselves. An $L^0(\mathbb{R})$-identification function for the identity or CDF-functional is the function-valued map $v_{\mathrm{CDF}}(F,x) = \big(F(a) - \mathds{1}\{x\le a\}\big)_{a\in\mathbb{R}}$, leading to the function-valued generalised error $e_{\mathrm{CDF}}(X) = \big(F(a) - \mathds{1}\{X\le a\}\big)_{a\in\mathbb{R}}$. Alternatively, one may consider the quantile function (QF) functional, mapping a CDF, $F$, to its generalised inverse, $F^{-1}$. On the class of random variables with a continuous CDF, denoted by $L_{\text{con}}$, we have the $L_{\text{con}}$-identification function $v_{\mathrm{QF}}(F^{-1},x) = \big(\alpha - \mathds{1} \{x \le F^{-1}(\alpha)\}\big)_{\alpha\in(0,1)}$. The modification discussed in Example (ref) for the general case leads to the generalised error $e_{\mathrm{QF}}(X) := \big( F_X( q_\alpha (X) ) - \mathds{1}\{X \le q_\alpha(X)\} \big)_{\alpha\in(0,1)}$. We can construct generalised covariances (and correlations) from $e_{\mathrm{CDF}}$ and $e_{\mathrm{QF}}$ via an outer product ansatz, generalising (ref). This reasoning explains why we call the resulting function-valued dependence measures CDF and quantile function covariance (correlation), and subsume both under the name distributional covariances (correlations).

definition[Distributional covariances and correlations] For any $X,Y\in L^0(\mathbb{R})$, the CDF covariance and the CDF correlation are \begin{align*} &\mathrm{CDFCov}(X,Y)\colon \mathbb{R}^2 \to \mathbb{R}, &&(a,b) \mapsto \mathrm{TCov}_{a,b}(X,Y),\\ &\mathrm{CDFCor}(X,Y)\colon \mathbb{R}^2 \to [-1,1], &&(a,b) \mapsto \mathrm{TCor}_{a,b}(X,Y). \end{align*} Likewise, the quantile function covariance and the quantile function correlation are \begin{align*} &\mathrm{QFCov}(X,Y)\colon (0,1)^2 \to \mathbb{R}, &&(\alpha, \beta) \mapsto \mathrm{QCov}_{\alpha, \beta}(X,Y),\\ &\mathrm{QFCor}(X,Y)\colon (0,1)^2 \to [-1,1], &&(\alpha, \beta) \mapsto \mathrm{QCor}_{\alpha, \beta}(X,Y). \end{align*}

Just as in univariate statistics when characterising the distribution of a single random variable, we may either take the perspective of the quantile function or of the CDF when characterising dependence between two random variables, as there are two corresponding families of distributional dependence measures. Indeed, the distributional covariances and correlations characterise the dependence structure between $X$ and $Y$ fully: Given the marginals $F_X$ and $F_Y$, it is possible to recover the joint CDF $F_{X,Y}$ or the copula $C_{X,Y}$, which contain the information on the full dependence structure, from $\mathrm{CDFCor}$ or $\mathrm{QFCor}$ (as well as $\mathrm{CDFCov}$ or $\mathrm{QFCov}$). This is clear from the representations of $\mathrm{TCov}$ in terms of $F_{X,Y}$ and the marginals in (ref) and $\mathrm{QCov}$ in terms of the $C_{X,Y}$ and the marginals in (ref) as well as the normalisations, which only depend on the marginals, see Examples (ref) and (ref). If $F_X$ and $F_Y$ are continuous, $\mathrm{QFCor}$ and $\mathrm{QFCov}$ are independent of the marginals, and by (ref) and again the normalisations from Example (ref), there is even a one-to-one mapping between $\mathrm{QFCov}$ and $C_{X,Y}$ (as well as $\mathrm{QFCor}$ and $C_{X,Y}$). While of course the joint CDF and copula contain the information on the dependence structure as well, we emphasise that particularly the normalised quantities, that is, the distributional correlations, have the great advantage that they are really dependence measures. Thus, they uncover the full dependence structure, and regions of stronger and weaker dependence can be identified relatively easily, see Subsection (ref) and Section (ref) for theoretical and empirical examples, respectively.

Distributional correlations are able to characterise the limiting cases of independence and perfect positive and negative dependence properly. In particular, nullity of any of the distributional covariances or correlations implies independence (Proposition (ref)). Further, unity (negative unity) of distributional correlations implies perfect positive (negative) dependence (Corollary (ref)). Additionally, $\mathrm{QFCor}$ is invariant to strictly increasing transformations.

propositionFor any $X,Y\in L^0(\mathbb{R})$ and for any of the four distributional dependence measures $\mathrm{D} \in \{\mathrm{CDFCov}, \mathrm{CDFCor}, \mathrm{QFCov}, \mathrm{QFCor}\}$ it holds that $X$ and $Y$ are independent if and only if $\mathrm D(X,Y)$ vanishes identically.
remarkThe functions $\mathrm{CDFCov}, \mathrm{CDFCor}, \mathrm{QFCov}$ and $\mathrm{QFCor}$ are naturally constant between jumps of the joint CDF $F_{X,Y}$, see (ref) and (ref). Furthermore, along the lines of Remark (ref), the generalised errors are constant for values $a\ge \mathrm{ess\,sup}(X)$ and $\alpha$ such that $q_\alpha(X)=q_1(X) = \mathrm{ess\,sup}(X)$. Thus, $\mathrm{CDFCov}$ and $\mathrm{CDFCor}$ are naturally only interesting on the restricted domain $[\mathrm{ess\,inf}(X),\mathrm{ess\,sup}(X)) \times [\mathrm{ess\,inf}(Y),\mathrm{ess\,sup}(Y))$ and $\mathrm{QFCov}$ and $\mathrm{QFCor}$ on $[F_X(\mathrm{ess\,inf}(X)), F_X(\mathrm{ess\,sup}(X))) \times [F_Y(\mathrm{ess\,inf}(Y)), F_Y(\mathrm{ess\,sup}(Y)))$.
corollaryFor any $X,Y \in L^0(\mathbb{R})$ and any of the two distributional correlations $\mathrm{DCor} \in \{\mathrm{QFCor},\mathrm{CDFCor}\}$ it holds that \begin{enumerate}[(i)] • Normalisation: $\mathrm{DCor}(X,Y) \in [-1,1]$. • Perfect dependence: $\mathrm{CDFCor}(X,Y) =1(-1)$ on $[\mathrm{ess\,inf}(X),\mathrm{ess\,sup}(X)) \times [\mathrm{ess\,inf}(Y),\mathrm{ess\,sup}(Y))$ if and only if $X$ and $Y$ are comonotonic (countermonotonic). \\ $\mathrm{QFCor}(X,Y) =1(-1)$ on $[F_X(\mathrm{ess\,inf}(X)), F_X(\mathrm{ess\,sup}(X))) \times [F_Y(\mathrm{ess\,inf}(Y)), F_Y(\mathrm{ess\,sup}(Y)))$ if and only if $X$ and $Y$ are comonotonic (countermonotonic). \end{enumerate}

Examples

figure[figure omitted — 262 chars of source]

Figures (ref) and (ref) present plots of quantile function correlations for classical examples of bivariate copulas, respectively. Figure (ref) depicts the quantile function correlation of a bivariate Cauchy copula with a Spearman correlation of 0 alongside a scatter plot of 1000 random draws from this copula. While joint CDFs or copulas and their respective densities are often difficult to interpret graphically, such a scatter plot has usually been considered the best, yet informal, tool to depict and understand the dependence structure. Quantile function correlation provides a formal tool that uncovers the full dependence structure: Even though $X$ and $Y$ have a Spearman correlation of 0, there is quite a strong dependence in the tails. In the first and third quadrants the dependence is positive, meaning that larger values (exceedances of certain quantiles) of $X$ are associated with larger values (exceedances of certain quantiles) of $Y$, while in the second and fourth quadrants it is negative, meaning that larger values (exceedances of certain quantiles) of $X$ are associated with smaller values (falling short of certain quantiles) of $Y$. Actually, Spearman's $\rho$ is equal to a properly normalised Lebesgue integral over $\mathrm{QFCov}$ (see Example (ref)) and thus the positive and negative dependence here cancels out when summarising the full dependence structure represented by $\mathrm{QFCor}$ in a single number.

figure[figure omitted — 304 chars of source]

Figure (ref) contains quantile function correlations stemming from four different copulas, all having a Spearman correlation of 0.5: a Gaussian, a Cauchy, a Clayton, and a Gumbel copula. Despite their Spearman's $\rho$ being the same, they exhibit very different dependence structures. For example, the Gaussian copula has a weak dependence in the lower left and upper right corner, while the Cauchy copula shows a strong dependence in both corners, and the Clayton and Gumbel copula exhibit strong dependence in one corner, but not the other. In the upper left and lower right corners, for all but the Cauchy copula there is strong positive dependence, which just means, e.g., for the lower right corner that exceedances of a high quantile of $X$ are strongly positively associated with exceedances of a small quantile of $Y$. This means that very large values of $X$ and very small values of $Y$ virtually never occur jointly for those copulas. In contrast, for the Cauchy copula, despite the overall positive dependence between $X$ and $Y$ as indicated by Spearman's $\rho$, the dependence in the lower right and upper left corners even becomes negative, reflecting the tail behaviour of the Cauchy distribution also seen in Figure (ref). Bivariate t-distributions with more degrees of freedom exhibit a similar behaviour as the Cauchy examples in both figures. Plots of the closely related CDF correlations (see Remark (ref)) usually look like distorted versions of quantile function correlations, where the distortion originates from the influence of the marginals. We discuss examples of CDF correlations in Section (ref).

Global positive and negative dependence

There is a vast literature on the question what it means or what it should mean that two random variables $X$ and $Y$ are globally positively (negatively) dependent. We refer to Mari2001 and Balakrishnan2009 for overviews of this strand of literature. The distributional covariances and correlations suggest a natural definition.

definitionAny $X,Y\in L^0(\mathbb{R})$ are globally positively dependent if $\mathrm{CDFCov}(X,Y)\ge0$. They are globally negatively dependent if $\mathrm{CDFCov}(X,Y)\le0$.
remarkDefinition (ref) of global positive dependence coincides with lehmann1966's (lehmann1966) definition of positive quadrant dependence, which holds for two random variables $X,Y$ if \begin{equation*} \mathbb{P} (X \leq a, Y \leq b) \geq \mathbb{P} (X \leq a) \mathbb{P} (Y \leq b) \quad for all a,b \in \mathbb{R}. \end{equation*}

The following proposition shows that we can define global positive and negative dependence also in terms of quantile function correlation.

propositionAny $X,Y\in L^0(\mathbb{R})$ are globally positively (negatively) dependent if and only if $\mathrm{QCov}(X,Y)\ge0$ ($\mathrm{QCov}(X,Y)\le0$).

For example, the first three copulas from Figure (ref) are globally positively dependent, while the Cauchy copulas in Figures (ref) and (ref) represent cases of mixed dependence, with regions of local positive as well as local negative dependence.

Each generalised covariance and correlation directly gives rise to a specific concept of positive (negative) dependence between two random variables $(X,Y)$ as well. The following proposition shows that such a dependence is implied by global dependence.

propositionAssume that $X,Y\in L^0(\mathbb{R})$ and $T_1,T_2$ are such that the generalised covariance $\mathrm{Cov}_{T_1,T_2}(X,Y)$ exists. If $X$ and $Y$ are globally positively dependent, it holds that $\mathrm{Cov}_{T_1,T_2}(X,Y) \geq 0.$ If $X$ and $Y$ are globally negatively dependent, $\mathrm{Cov}_{T_1,T_2}(X,Y) \leq 0$.

Proposition (ref) can also be viewed as a complement to Theorem (ref) in that it contains a further desirable property of generalised correlations: They not only take the correct values of $-1$\,/\,0\,/\,1 in the extreme cases, but also the correct sign under positive and negative dependence.

Tail correlations and tail dependence

It is often of interest to analyse co-movement in the tails, i.e., if for example very large values in $X$ and $Y$ tend to occur together, if there is no dependence in the tails, or if large values of $X$ render large values of $Y$ rather more unlikely and vice versa. Such questions are subsumed under the term of tail, extremal or asymptotic dependence in the literature, see, e.g., Joe2014 or Coles1999.

Natural measures for lower and upper tail dependence are the respective limits of quantile correlation or the limits of quantile function correlation when moving to the upper right and lower left corner, respectively.

definition[Tail correlations] For any $X,Y \in L^0(\mathbb{R})$ the lower and upper tail correlations are defined as \begin{equation*} \mathrm{LTCor}(X,Y) := \lim_{\alpha \to 0} \mathrm{QCor}_{\alpha,\alpha}(X,Y), \qquad \mathrm{UTCor}(X,Y) := \lim_{\alpha \to 1} \mathrm{QCor}_{\alpha,\alpha}(X,Y), \end{equation*} respectively, provided that the limits exist.

On top of the upper and lower tail correlation, one might also consider $\lim_{\alpha \to 0} \mathrm{QCor}_{\alpha,1-\alpha}(X,Y)$ and $\lim_{\alpha \to 0} \mathrm{QCor}_{1-\alpha,\alpha}(X,Y)$, provided that these limits exist. The discussion is similar to what follows and is therefore omitted.

definition[Tail dependence] Any $X,Y \in L^0(\mathbb{R})$ are positively lower tail dependent if $\mathrm{LTCor}(X,Y) > 0$, negatively lower tail dependent if $\mathrm{LTCor}(X,Y) < 0$, and lower tail independent if $\mathrm{LTCor}(X,Y) = 0$. They are lower tail comonotonic if $\mathrm{LTCor}(X,Y) = 1$ and lower tail countermonotonic if $\mathrm{LTCor}(X,Y) = -1$. For the upper tail notions, replace $\mathrm{LTCor}(X,Y)$ by $\mathrm{UTCor}(X,Y)$.

The by far most prominent measure of tail dependence is the coefficient of tail dependence, see Joe1993, Coles1999.\footnote{Fiebig2017 contains a literature review on the use and naming of the coefficient in different fields.} To facilitate the following discussion of the relation between coefficients of tail dependence and tail correlations we assume continuity of the marginals $F_X$ and $F_Y$ throughout, stated by $X, Y\in L_{\mathrm{con}}$. The coefficients of lower and upper tail dependence are defined as \[ \lambda_l(X,Y) := \lim_{\alpha \to0} \mathbb{P} ( Y \leq q_{\alpha} (Y) \mid X \leq q_{\alpha} (X)) = \lim_{\alpha \to0} \frac{C_{X,Y}(\alpha,\alpha)}{\alpha}, \quad X, Y\in L_{\text{con}}, \] and \[ \lambda_u(X,Y) := \lim_{\alpha \to1} \mathbb{P} ( Y > q_{\alpha} (Y) \mid X > q_{\alpha} (X)) = \lim_{\alpha \to1} \frac{\overline{C}_{X,Y}(\alpha,\alpha)}{ 1 - \alpha}, \quad X, Y\in L_{\text{con}}, \] where $\overline{C}_{X,Y}$ denotes the survival function of the copula $C_{X,Y}$ and where we assume that the limits exist. The following lemma clarifies the relation between the coefficients of tail dependence and the tail correlations.

lemmaFor $X, Y\in L_{\mathrm{con}}$, the following assertions hold. \begin{enumerate}[(a)] • If the coefficient of lower (upper) tail dependence or the lower (upper) tail correlation exist and are positive, the other quantity exists as well and the two quantities coincide. That is, \begin{align*}\nonumber \lambda_l(X,Y) &= \lim_{\alpha \downarrow0} \frac{C_{X,Y}(\alpha,\alpha)}{\alpha} = \lim_{\alpha \downarrow0} \frac{C_{X,Y}(\alpha,\alpha) - \alpha^2}{\alpha - \alpha^2} =\mathrm{LTCor}(X,Y)\,,\\ \lambda_u(X,Y) &= \lim_{\alpha \uparrow 1} \frac{\overline{C}_{X,Y}(\alpha,\alpha)}{1-\alpha} = \lim_{\alpha \uparrow 1} \frac{\overline{C}_{X,Y}(\alpha,\alpha) - (1-\alpha)^2}{\alpha(1-\alpha)} = \lim_{\alpha \uparrow 1} \frac{C_{X,Y}(\alpha,\alpha) - \alpha^2}{\alpha - \alpha^2} = \mathrm{UTCor}(X,Y)\,. \end{align*} • If $X, Y$ are lower (upper) tail independent, the coefficient of lower (upper) tail dependence is 0 as well. • If $X,Y$ are negatively lower (upper) tail dependent, it holds that \begin{align*} \mathrm{LTCor}(X,Y) &= \lim_{\alpha \downarrow 0} \frac{C_{X,Y}(\alpha,\alpha) - \alpha^2}{\alpha^2} = \lim_{\alpha \downarrow 0} \frac{C_{X,Y}(\alpha,\alpha)}{\alpha^2}-1 and \lambda_l(X,Y) = 0 \,,\\ \mathrm{UTCor}(X,Y) &= \lim_{\alpha \uparrow 1} \frac{C_{X,Y}(\alpha,\alpha) - \alpha^2}{(1-\alpha)^2} = \lim_{\alpha \uparrow 1} \frac{\overline{C}_{X,Y}(\alpha,\alpha)}{(1-\alpha)^2} - 1 and \lambda_u(X,Y) = 0 \,. \end{align*} \end{enumerate}

In the literature, the cases $\lambda_l=0$ and $\lambda_u=0$ are usually called tail or asymptotic independence and $\lambda_l>0$ and $\lambda_u>0$ tail dependence mcneil2015. The lower and upper tail correlations provide a more nuanced picture of the tail behaviour. While under positive tail dependence as introduced in Definition (ref), the coefficients of tail dependence and the tail correlations coincide, the latter measures are able to classify the situation of $\lambda_l=0$ and $\lambda_u=0$ into actual tail independence (the tail correlations are 0) on the one hand and negative tail correlation on the other hand -- also indicating the strength of negative dependence. In fact, a countermonotonic pair $(X,Y)$ yields a lower (upper) tail correlation of $-1$, while the coefficients of tail dependence are still 0, deeming the pair asymptotically independent.

Hence, we make the case for replacing the coefficients of tail dependence with the tail correlations: Since no information is lost, but strictly more information is gained, this is one of the rare cases in statistical methodology where a Pareto improvement is possible and should therefore be implemented.

Summary covariances and correlations

Summary covariances

The distributional covariances and correlations from Section (ref) reveal the full dependence structure between $X$ and $Y$. Nevertheless, it is often required or useful to summarise the dependence structure in a single number. In fact, this is what most classical dependence measures aim for. A natural way to construct such summary measures from distributional covariances is to compute weighted averages.

definition[Summary covariances] For any $X,Y\in L^0(\mathbb{R})$ the summary covariance induced by quantile function covariance with respect to a measure $\kappa$ on $[0,1]^2$ is \begin{equation} \mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y) := \int_{[0,1]^2} \mathrm{QCov}_{\alpha,\beta} (X,Y) \,\mathrm{d}\kappa (\alpha,\beta). \end{equation} Likewise, the summary covariance induced by CDF covariance with respect to a measure $\nu$ on $\mathbb{R}^2$ is \begin{equation} \mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y) := \int_{\mathbb{R}^2} \mathrm{TCov}_{a,b} (X,Y) \,\mathrm{d}\nu (a,b). \end{equation}

We tacitly assume that the integrals in (ref) and (ref) exist and are finite. Since both $\mathrm{QCov}$ and $\mathrm{TCov}$ are bounded, a sufficient condition is that $\nu$ and $\kappa$ are finite. The summary covariances inherit the properties of the distributional covariances. They are 0 under independence and nonnegative (nonpositive) under global positive (negative) dependence. Further, the quantile function summary covariance is invariant under strictly increasing transformations.

Interestingly, two of the most popular dependence measures arise as canonical special cases of summary covariances.

example[Covariance] If we plug in the Lebesgue measure $\lambda$ for $\nu$ in (ref), Hoeffding's formula mcneil2015 implies that \begin{equation*} \mathrm{SCov}_{\mathrm{CDF},\lambda} (X,Y) = \int_{\mathbb{R}^2} \mathrm{TCov}_{a,b} (X,Y) \,\mathrm{d} (a,b) = \mathrm{Cov}(X,Y), \end{equation*} which is the classical covariance.
example[Spearman covariance] Recall that Spearman's rank correlation coefficient $\rho$ can be defined as the Pearson correlation of the probability integral transforms, \begin{equation} \rho(X,Y) = \mathrm{Cor} \big(F_X(X),F_Y(Y)\big). \end{equation} From Example (ref) and the relation between quantile and threshold correlation discussed in Remark (ref), it follows for $\kappa$ being the Lebesgue measure $\lambda$ that (ref) becomes \begin{equation*} \mathrm{SCov}_{\mathrm{QF},\lambda} (X,Y) = \int_{[0,1]^2} \mathrm{QCov}_{\alpha,\beta} (X,Y) \,\mathrm{d}(\alpha, \beta) = \mathrm{Cov}\big(F_X(X),F_Y(Y)\big), \end{equation*} which is the Spearman covariance.

We discuss further examples, which focus on specific regions of interest, when dealing with the respective correlations below.

Summary correlations

Summary correlations arise when summary covariances are appropriately normalised, again utilising co- and countermonotonic random couplings with identical marginals as $X$ and $Y$.

definition[Summary correlations] Consider measures $\kappa$ on $[0,1]^2$ and $\nu$ on $\mathbb{R}^2$. For any non-constant $X,Y\in L^0(\mathbb{R})$ consider pairs with the same marginal distributions $(X,Y')$ and $(X,Y'')$ such that $(X,Y')$ is countermonotonic and $(X,Y'')$ is comonotonic. Then the summary correlation induced by quantile function covariance with respect to $\kappa$ is \begin{equation} \mathrm{SCor}_{\mathrm{QF},\kappa} (X,Y) := \begin{cases} \frac{\mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y)}{|\mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y'\,)|} &if \mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y) <0, \\[0.5em] \frac{\mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y)}{|\mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y”\,)|} &if \mathrm{SCov}_{\mathrm{QF},\kappa} (X,Y)\ge0. \end{cases} \end{equation} Likewise, the summary correlation induced by CDF covariance with respect to $\nu$ is \begin{equation} \mathrm{SCor}_{\mathrm{CDF},\nu} (X,Y) := \begin{cases} \frac{\mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y)}{|\mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y'\,)|} &if \mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y) <0, \\[0.5em] \frac{\mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y)}{|\mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y”\,)|} &if \mathrm{SCov}_{\mathrm{CDF},\nu} (X,Y)\ge0. \end{cases} \end{equation} Provided that the involved quantities exist and are finite. If $X$ or $Y$ is constant, then the two measures are set to be 0.

The normalisation terms in (ref) and (ref) can be computed from (ref) and (ref) and the Fr\'echet--Hoeffding bounds presented in Examples (ref) and (ref). The summary correlations inherit the appealing properties of the distributional correlations.

corollaryFor any $X,Y \in L^0(\mathbb{R})$ and for any of the two summary correlations $\mathrm{SCor} \in \{\mathrm{SCor}_{\mathrm{QF},\kappa},\mathrm{SCor}_{\mathrm{CDF},\nu}\}$ such that $\mathrm{SCor}(X,Y)$ exists, the following properties hold. \begin{enumerate}[(i)] • Normalisation: $\mathrm{SCor}(X,Y) \in [-1,1]$. • Independence implies nullity: $\mathrm{SCor}(X,Y) =0$ if $X$ and $Y$ are independent. • Perfect dependence: $\mathrm{SCor}(X,Y) =1(-1)$ if $X$ and $Y$ are comonotonic (countermonotonic). If $\kappa$ and $\nu$ are strictly positive,\footnote{That means they assign a strictly positive mass to any open non-empty set.} then $\mathrm{SCor}(X,Y) =1(-1)$ implies that $X$ and $Y$ are comonotonic (countermonotonic). • Symmetry: If $\kappa$ and $\nu$ are invariant in their arguments in the sense that $\,\mathrm{d}\kappa (\alpha,\beta) = \,\mathrm{d}\kappa (\beta,\alpha)$ and $\,\mathrm{d}\nu (a,b) = \,\mathrm{d}\nu (b,a)$ for all $\alpha, \beta\in [0,1]$ and for all $a,b\in\mathbb{R}$, then $\mathrm{SCor}(X,Y) = \mathrm{SCor}(Y,X)$. \end{enumerate}

By Proposition (ref), $\mathrm{SCor}_{\mathrm{QF},\kappa}$ is invariant under strictly increasing transformations of $X$ and $Y$ as well.

example[Mean and Pearson correlation] From Example (ref) it follows that \begin{equation*} \mathrm{SCor}_{\mathrm{CDF},\lambda} (X,Y) = \mathrm{MCor}(X,Y), \end{equation*} where $\lambda$ is the Lebesgue measure on $\mathbb{R}^2$. Under the conditions of Lemma (ref) it holds that $\mathrm{SCor}_{\mathrm{CDF},\lambda} (X,Y) = \mathrm{Cor}(X,Y)$.
example[Spearman correlation] By Example (ref) it holds for the Lebesgue measure $\lambda$ on $\mathbb{R}^2$ that \begin{equation*} \mathrm{SCor}_{\mathrm{QF},\lambda} (X,Y) = \mathrm{MCor}\big(F_X(X),F_Y(Y)\big). \end{equation*} If $F_X$ and $F_Y$ are continuous, Spearman's $\rho$ arises, $\mathrm{SCor}_{\mathrm{QF},\lambda} (X,Y)=\rho(X,Y)$, since the probability integral transforms are standard uniform, $F_X(X) \sim U(0,1)$, $F_Y(Y) \sim U(0,1)$, and thus fulfil the conditions of Lemma (ref). In this case the normalisation does not depend on the sign of $\mathrm{Cov}(F_X(X),F_Y(Y))$ and always equals $\frac{1}{12}$, which leads to the well-known formula $\rho(X,Y) = 12 \mathrm{Cov}(F_X(X),F_Y(Y))$. In the discrete case, $\mathrm{MCor}(F_X(X),F_Y(Y))$ recovers a proposal by genest2007 to combat the non-attainability of Spearman's $\rho$ for discrete random variables.

In practice, Pearson and Spearman correlation are most often interpreted as summaries of the full dependence structure, expressed in a single number. However, by definition covariance and Pearson correlation measure dependence around the means (which was the starting point of this paper), while Spearman's $\rho$ does the same on the rank scale, see (ref) and (ref). Our formal approach to summary correlations, where they arise as canonical special cases, provides a powerful justification for this practical use. Of course, the dependence structure cannot be fully described by a single number, for example for all the bivariate copulas in Figure (ref), Spearman's $\rho$ has the same value, $\rho=0.5$, but they have very different dependence structures as the distributional correlations uncover. The two closely related families (threshold and quantile family) of local, distributional and canonical summary correlations provide dependence measures for different purposes and should be chosen according to the statistical problem at hand: measuring dependence locally, characterising the full dependence structure or condensing it in a single number.

example[Regional measures of dependence] Let us now integrate with the Lebesgue measure only over some parts of $[0,1]^2$ or $\mathbb{R}^2$. For example in the quantile function case and for $A \subset [0,1]^2$, the respective summary covariance reads: \begin{equation*} \mathrm{SCov}_{\mathrm{QF},\lambda_A} (X,Y) := \int_{A} \mathrm{QCov}_{\alpha,\beta} (X,Y) \,\mathrm{d} (\alpha, \beta). \end{equation*} Considering the respective correlation $\mathrm{SCor}_{\mathrm{QF},\lambda_A}$ and letting $A$ be located in one of the tails, e.g., $A = [0,c]^2$ for small $c$, this leads to an alternative measure of tail dependence that has a similar relation to $\mathrm{QCor}_{c,c}$ as the expected shortfall has to the value at risk. Letting $A$ be a region in the centre of the distribution, e.g., again a rectangle $[0.5-d,0.5+d]^2$, this leads to a measure of dependence in the centre, which is in a similar relation to median covariance $\mathrm{QCor}_{0.5,0.5}$. When defining the corresponding quantity for the summary correlation induced by CDF correlation with a region $B \subset \mathbb{R}^2$, we get \begin{equation*} \mathrm{SCov}_{\mathrm{CDF},\lambda_B} (X,Y) := \int_{B} \mathrm{TCov}_{a,b} (X,Y) \,\mathrm{d} (a, b). \end{equation*} Here, one could focus on a whole quadrant, e.g., $B=[-\infty,0]^2$. This resembles the idea behind so-called semi-correlations Joe2014.

Estimation

The empirical analogues of generalised covariances and correlations are fairly straightforward, exploiting plug-in estimators and the method of moments. They constitute natural and consistent estimators for the respective quantities on the population level. Suppose we have a random sample $\{(X_i,Y_i),\ i=1,\ldots, n\}$ from $F_{X,Y}$ and we are interested in estimating the generalised covariance, $\mathrm{Cov}_{T_1,T_2}(X,Y)$, or generalised correlation, $\mathrm{Cor}_{T_1,T_2}(X,Y)$, at $T_1$, $T_2$. Further, suppose that the corresponding generalised errors are induced by increasing and non-constant $\mathcal{L}_1$- and $\mathcal{L}_2$-identification functions $v_{1}$ and $v_{2}$, invoking Proposition (ref). To start with, suppose that $\mathcal{L}_1$ and $\mathcal{L}_2$ contain all random variables with any empirical distribution. Then, we can simply apply the definition of the generalised covariance and correlation to the empirical distribution and use this as an estimator for the population quantity. To obtain the sample analogue $\hat t_1^n := \widehat T_1^n(X)$ of an identifiable functional $T_1(X)$, we use the solution in $t$ of

equation[equation omitted — 88 chars of source]

Similarly, we write $\hat t_2^n := \widehat T_2^n(Y)$. Then, the generalised covariance on the sample level is

equation[equation omitted — 149 chars of source]

To obtain the normalisation for the generalised correlation, we take the empirical marginal distributions from the random sample $\{(X_i,Y_i),\ i=1,\ldots, n\}$, but couple the observations with the corresponding co- and countermonotonicity copulas, utilising the increasing order statistics $X_{(1)}\le \cdots \le X_{(n)}$ and $Y_{(1)}\le \cdots \le Y_{(n)}$. Hence, we obtain for the comonotonic coupling $(X,Y'')$ and the countermonotonic coupling $(X,Y')$

align[align omitted — 301 chars of source]

Finally, we set

equation[equation omitted — 414 chars of source]

For threshold and quantile correlation the estimation simplifies as the normalisation does not involve the co- and countermonotonic coupling. For threshold correlation, $\mathrm{TCor}_{a,b}(X,Y)$, we only need to estimate $F_{X,Y} (a,b)$, $F_X(a)$ and $F_Y(b)$ via

equation[equation omitted — 306 chars of source]

and replace the theoretical quantities by their empirical counterparts in $\mathrm{TCov}_{a,b}(X,Y)$ from (ref) and the normalisation terms in Example (ref).

To estimate quantile correlation, $\mathrm{QCor}_{\alpha,\beta}(X,Y)$, we first obtain the sample $\alpha$- and $\beta$-quantiles, or more formally, for the canonical identification functions $v_{q_\alpha}$ and $v_{q_\beta}$ in (ref) we set \[ \hat q^n_\alpha := \inf \{t \mid \frac{1}{n}\sum_{i=1}^n v_{q_\alpha}(t,X_i)\le0\}, \qquad \hat q^n_\beta := \inf \{t \mid \frac{1}{n}\sum_{i=1}^n v_{q_\beta}(t,Y_i)\le0\}. \] Then, we obtain estimates for $C_{X,Y}\big(F_X(q_\alpha(X)), F_Y(q_\beta(Y))\big)$, $F_X(q_\alpha(X))$ and $F_Y(q_\beta(Y))$ by replacing the thresholds $a$, $b$ in (ref) with $q^n_\alpha$ and $q^n_\beta$ and replacing the theoretical quantities in (ref) and the normalisation terms in Example (ref) with those estimators.

To construct estimators for distributional correlations, we use the respective estimators for the local correlations just described on a grid (of size 10000 in the following applications). We close this section by establishing the consistency of the described estimators, subject to typical regularity conditions.

propositionSuppose that for $i=1,2$ the identification functions $v_i$ are increasing, strict $\mathcal{L}_i$-identification functions for $T_i$, such that the families $v_1(\cdot,x)_{x\in\mathbb{R}}$ and $v_2(\cdot, y)_{y\in \mathbb{R}}$ are pointwise equicontinuous and let $X\in\mathcal{L}_1, Y \in \mathcal{L}_2$. Further, suppose that $\{F_X: X\in\mathcal{L}_1\}$ and $\{F_Y:Y\in\mathcal{L}_2\}$ contain all empirical distribution functions. \\ Then, the estimators $\widehat \mathrm{Cov}_{T_1,T_2}^n(X,Y)$ (ref) and $\widehat \mathrm{Cor}_{T_1, T_2}^n(X,Y)$ (ref) based on a random sample are strongly consistent. That is, they converge almost surely to $\mathrm{Cov}_{T_1,T_2}(X,Y)$ and $\mathrm{Cor}_{T_1,T_2}(X,Y)$, respectively.

Threshold covariance and threshold correlation do not satisfy the conditions of Proposition (ref), but the strong consistency follows directly from the strong law of large numbers and the continuous mapping theorem. For quantile covariance and quantile correlation, we establish strong consistency if the marginal distributions are continuous at the respective quantiles, see Proposition (ref).

Data examples

To illustrate the use of local, distributional and summary correlations in practice, we extract data on mixed-sex couples from the 2019 wave of the Panel Study of Income Dynamics.\footnote{Citation: Panel Study of Income Dynamics, public use dataset. Produced and distributed by the Survey Research Center, Institute for Social Research, University of Michigan, Ann Arbor, MI (2023).} After data cleaning we have a sample of 4417 couples living in the same households, of whom 85 % are married. As all of the variables we analyse are discrete (either due to being inherently discrete or, e.g., heights being recorded in full inches), we use bubble plots for our scatter plots to avoid overplotting. Further, note that CDF and quantile function correlation only change their value at jumps of the CDF (see Remark (ref)). In all graphical representations we focus on the ranges between the 2.5%- and the 97.5%-quantiles.

figure[figure omitted — 310 chars of source]

We first consider the heights of the couples in cm. Figure (ref) depicts a scatter plot and the corresponding CDF correlation. Heights of men and women are globally positively dependent as CDF correlation is positive everywhere, indicating a preference for assortative mating stulp2013. Mean correlation equals $\mathrm{MCor}=0.215$ (Spearman's $\rho=0.189$), indicating a weak positive relation on average. The local dependence structure, however, varies strongly: In the lower right corner, the dependence is quite strong, while elsewhere (with the exception of the upper left corner) it is weak. This reflects the male-taller norm in Western societies stulp2013: For women and men fairly close to the average height difference (indicated by the dotted line), CDF correlation is close to 0, indicating that mating behaviour in this region is hardly influenced by the partner's (in relation to the own) height. As the height of the man approaches the height of the woman (the diagonal is represented by the solid line), CDF correlation rises abruptly, suggesting that the mating behaviour is strongly influenced by height in this region: heights of partners are strongly positively associated there.\footnote{Actually, the norm is rather that the woman should be at least a few centimetres smaller than the man.} When computing the regional summary correlation $\mathrm{SCor}_{\mathrm{CDF},\lambda_B}$ from Example (ref) below and above the diagonal we consequently get 0.601 and 0.201, respectively. In the far upper left corner the dependence gets quite strong as well, reflecting the male-not-too-tall norm stulp2013. Quantile function correlation yields qualitatively the same picture, see Figure (ref) in the Appendix and the discussion below.

figure[figure omitted — 243 chars of source]

Figure (ref) depicts the bubble plot, the CDF and quantile function correlation for the body mass indices (BMI, mass in kg divided by the square of height in cm) of the couples. Again the two variables are (almost) globally positively dependent and the canonical summary correlations from examples (ref) and (ref) take the values $\mathrm{MCor}=0.304$ and $\rho=0.287$. In the lower left corner below a BMI of 25, which corresponds to the threshold between having a normal weight and being overweight according to the World Health Organization (WHO),\footnote{\url{https://www.who.int/europe/news-room/fact-sheets/item/a-healthy-lifestyle---who-recommendations}, accessed: 7th May 2023} $\mathrm{CDFCor}$ and $\mathrm{QCor}$ are close to 0, whereas elsewhere they are larger, in particular if one of the partners exceeds the threshold of 30, from which on a person is classified as obese according to the WHO. Thus, if both partners are in a normal weight range, there seems to be no association between BMIs, but once one of the partners is obsese, the assocation gets strongly positive.\footnote{Here, the dependence is of course not only generated by mating behaviour, but also by mutual influence in lifestyle.} This example nicely illustrates how CDF and quantile correlation may differ and how they complement each other: While $\mathrm{CDFCor}$ is often nicely interpretable as it operates on the observation scale itself, sometimes a lot of space in the plot is occupied by regions where few observations lie (which is why we chose to present only the central regions of the axes with 95% of the probability mass in the first place). For example here, due to the marginal distributions of the BMIs being right-skewed, regions with higher values occupy comparatively more space in the plot. $\mathrm{QFCor}$, on the other hand, always has a solid interpretation in terms of quantile levels and naturally assigns space in the plot according to probability mass.

figure[figure omitted — 263 chars of source]

Next, we consider an example exhibiting regions of positive and of negative dependence. Figure (ref) contains the scatter plot as well as the CDF correlation of the number of strength trainings per week and the BMI of the men in our sample. On average, frequency of strength training is negatively associated with BMI, $\mathrm{MCor}=-0.112$ and $\rho=-0.088$. Locally, the negative dependence starts to arise with about the overweight threshold of 25 and gets stronger the higher BMI gets, being particularly strong for people above the obesity threshold of 30. For men in the upper normal range of BMI, $\mathrm{CDFCor}$ is close to 0, whereas for men in the lower normal range, the association is positive. Thus, for people with a low BMI strength training is associated with a gain in body weight -- probably due to increased muscle mass, while for overweight people it is associated with a loss in body weight -- probably due to fat loss outweighing increased muscle mass.

In Part (ref) of the Appendix we provide the above-mentioned additional two figures and compare mean correlation from Example (ref) and Pearson correlation from (ref) for the three examples above and a fourth one.

Conclusion

We present new concepts and measures of dependence. On the one hand, we consider dependence from the perspective of statistical functionals and put forward generalised covariances and correlations as corresponding dependence measures. On the other hand, with our local and distributional correlations we introduce local dependence (including tail dependence) measures as well as function-valued measures uncovering the full dependence structure. Summary correlations average over distributional correlations and close the loop to classical measures of dependence like covariance, Pearson correlation and Spearman's $\rho$. We analyse the properties of the new measures and present first applications.

These measures open many opportunities for future research. First of all, they will be useful in a wide variety of applications, possibly providing deeper insights about dependence structures than classical measures. Our measures are concerned with dependence between two random variables. Thus, they naturally extend to settings where pairwise dependence plays a role, for example correlation matrices, temporal or spatial dependence. Extending our measures to examine dependence for a vector of variables jointly is naturally more difficult, as negative dependence turns into a subtle concept for more than two variables Mari2001. Further aspects of statistical inference, beyond those related to estimation and discussed in the paper, are relegated to future research. Our measures are fundamentally different from recent popular measures of functional dependence szekely2007, szekely2009, chatterjee2021, reshef2011, which are not measures of directional dependence, but only of strength of dependence, thus mapping to $[0,1]$. Analysing possible connections of those measures to distributional or summary correlations is certainly interesting. Finally, exploring the links between generalised correlations and generalised regression approaches might be fruitful.

Acknowledgements

We are grateful to Patrick Cheridito, Timo Dimitriadis, Tilmann Gneiting, Bettina Gr\"un, Alexander Jordan, Johanna Ne\v{s}lehov\'a, Melanie Schienle and Jan-Lukas Wermuth for valuable discussions about the topic. We further thank seminar participants at Heidelberg University, Heidelberg Institute for Theoretical Studies, Vienna University of Economics and Business, University of Sussex and ETH Z\"urich and conference participants at DAGStat 2022 and the Bernoulli Young Researcher Event 2022 for helpful comments. Marc-Oliver Pohle is grateful for support by the Klaus Tschira Foundation, Germany. The collection of data used in this study was partly supported by the National Institutes of Health under grant number R01 HD069609 and R01 AG040213, and the National Science Foundation under award numbers SES 1157698 and 1623684.