EconBase
← Back to paper

Various issues around the L1-norm distance

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

171,399 characters · 42 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Various issues around the $L_1$-norm distance

titlepage\begin{abstract} Beyond the new results mentioned hereafter, this article aims at familiarizing researchers working in applied fields -- such as physics or economics -- with notions or formulas that they use daily without always identifying all their theoretical features or potentialities. Various situations where the $L_1$-norm distance $E|X-Y|$ between real-valued random variables intervene are closely examined. The axiomatic surrounding this distance is also explored. We constantly try to build bridges between the concrete uses of $E|X-Y|$ and the underlying probabilistic model. An alternative interpretation of this distance is also examined, as well as its relation to the Gini index (economics) and the Lukaszyk-Karmovsky distance (physics). The main contributions are the following: (a) We show that under independence, triangle inequality holds for the normalized form $E|X-Y| /(E|X| + E|Y|)$. (b) In order to present a concrete advance, we determine the analytic form of $E|X-Y|$ and of its normalized expression when $X$ and $Y$ are independent with Gaussian or uniform distribution. The resulting formulas generalize relevant tools already in use in areas such as physics and economics. (c) We propose with all the required rigor a brief one-dimensional introduction to the optimal transport problem, essentially for a $L_1$ cost function. The chosen illustrations and examples should be of great help for newcomers to the field. New proofs and new results are proposed. \\ \\ \noindentKeywords: $L_p$-distance, Normalized $L_1$-distance, independence, Lukaszyk-Karmowski metric, Gini index, coupling, optimal transport.\\ \\ \noindentMSC2020 Codes: 60A05, 62P20, 91B80, 91B82\\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

Introduction

The notion of distance is fundamental in human experience; human beings constantly need to represent some degree of closeness between objects, whether the latter are physical or symbolic, concrete or abstract. Quantifying the closeness between random objects has become a task of vital interest to virtually all researchers working in applied sciences. This text is designated for a broad readership and we tried to make it as self-contained as we reasonably can, with as few “it can be shown that” as possible. As a result, the presentation contains a relatively larger amount of background material than it is usually found in articles dealing with comparable subjects. Our hope is that readers mainly interested in applications will find a theory that is accessible to them.

Conceptual metric spaces are usually meant to have properties similar to those of the “natural” metric $|x-y|$ of the real line. We can ask ourselves the following question: which distance should we use when the real numbers $x$ and $y$ are replaced by real-valued random variables $X \hbox{ and } Y$? The answer is not unique, but it immediately comes to mind to look for it in the family of distances resulting from the $L_p$ norm. Let us first recall the context in which this norm is used. In what follows a probability space will be denoted by $(\Omega,{\cal A},P)$, where the sample space $\Omega$ is endowed with a $\sigma$-field ${\cal A}$ and a probability function $P$, and ${\cal B}_1$ will designate the $\sigma$-field of Borel subsets of $I\!\!R$. If $p \in [1, \infty )$, which is the range of interest of most applications and the range we will consider in this article, the space $L_p (\Omega, {\cal A}, P)$ -- also denoted by $L_p (I\!\!R)$ or even more simply by $L_p$ -- consists of all $p$-(Lebesgue) integrable random variables $(\Omega, {\cal A},P) \rightarrow (I\!\!R, {\cal B}_1)$, i.e. random variables that satisfy $E|X|^p < \infty$. Then, if $X \in L_p$, we define the $L_p$ norm of $X$ by $||X||_p = (E|X|^p )^{1/p}$. Note that we are facing the following technical point: $||X||_p = 0$ does not imply that $X=0$, but only that $X=0$ almost surely. To be quite precise, the set of $p$-integrable real-valued random variables together with the function $||\cdot ||_p$ is a seminormed vector space denoted by ${\cal L}_p (\Omega, {\cal A}, P)$, ${\cal L}_p(I\!\!R)$ or simply ${\cal L}_p$ . Then the quotient space $L_p (\Omega, {\cal A}, P)$ is defined as the normed vector space of the equivalence classes for the equivalence relation: $X \sim Y $ if and only if $X = Y$ almost surely, where $X,Y \in {\cal L}_p(\Omega, {\cal A}, P)$. In other words the random variables which agree almost surely are identified. The passage of ${\cal L}_p$ to $L_p$ is convenient, but also a bit confusing. However, many authors use the notation $L_p$ to refer to either space. In practice, we often forget that we are in the presence of equivalence classes rather than random variables. In short, ${\cal L}_p$ refers to a set of random variables while $L_p$ is a set of classes for the a.s. equality relation. We will make a clear distinction when other equivalence relations are involved, e.g. in Subsection (ref).

The spaces $L_p$ are complete normed vector spaces, that is Banach spaces. “Complete” means that the limit of any Cauchy sequence is within the space itself. Among all $p \in [1, \infty )$, the case $p=2\ $ is special: the norm $||X||_2$ follows from the inner product of $L_2$, and $(L_2,||X||_2)$ is the only Hilbert space\footnote{A Hilbert space is a Banach space whose norm is determined by an inner product. } of the family. Finally $1 \le p < q < \infty$ implies that $||X||_p = (E|X|^p )^{1/p} \le (E|X|^q )^{1/q} = ||X||_q$ and hence $L_q \subset L_p$ (in particular $L_2 \subset L_1$).

We still haven't answered the question of replacing the distance $|x-y|$ between two real numbers $x$ and $y$ when instead of them we have to deal with real-valued random variables $X \hbox{ and } Y$. Let $\delta_x$ (resp. $\delta_y$) denote the Dirac delta measure supported on the singleton $\{ x \}$ (resp. $\{ y \}$). If $X \sim \delta_x$ and $Y \sim \delta_y$, then the distance $||X-Y||_p$ equalizes to $|x-y|$. Indeed, $E|X-Y|^p = \int_{I\!\!R} \ \{ \int_{I\!\!R} |s-t|^p \delta(t-y) dt\}\ \delta(s-x) ds = |x-y|^p$, and thus $||X-Y||_p = (E|X-Y|^p )^{1/p} = |x-y|$. So all distances $||X-Y||_p$ generalize $|x-y|$ in the sense described above. Which $p$ should we choose? A serious candidate is $p = 2$, by virtue of the Hilbertian character of the space $(L_2,||\cdot ||_2)$. But if in a given study we are not interested in the mathematical facilities that the existence of an inner product allows (ability to treat random variables as vectors, projection on a closed convex set, orthogonality, etc.), then the choice of $(L_1,||\cdot ||_1)$ seems more suitable: let us look in this respect at the $L_2$-norm distance $||X-Y||_2 = (E|X-Y|^2 )^{1/2}$. Its use gives rise to a distortion in that it squares the differences: this distance tends to underestimate small differences (smaller than 1) and overestimate large differences (larger than 1), although the problem is mitigated by taking the square root at the end of the calculation. As the $L_1$ distance is not subject to this type of deformation, it comes first in terms of simplicity and interpretation. And all the more so since it is less than or equal to all other $L_p$ distances: $||X-Y||_1 \le ||X-Y||_p$ $\forall p > 1$. For all these reasons we will mainly focus on the distance $||X-Y||_1 = E|X-Y|$ between real-valued integrable random variables $X \hbox{ and } Y$ defined on the same probability space $(\Omega,{\cal A},P)$, namely

equation[equation omitted — 153 chars of source]

where $P_{(X,Y)}$ is the pushforward probability of $P$ induced on $I\!\!R^2$ by the random pair $(X,Y)$. Equation ((ref)) can be given the following interpretation: a pair of random values $(x,y)$ taken by the jointly distributed $X \hbox{ and } Y$ is observed. The absolute difference between $x$ and $y$ is recorded. The sampling procedure is repeated independently an infinite number of times and the observed absolute differences $|x-y|$ are averaged. This endless process will yield $E|X-Y|$.

The fields concerned with the expected absolute difference are numerous and various. They include in particular data analysis, clustering, optimal transport, physics, biology, economics, finance, engineering, image analysis. Interestingly, an identical distance measurement is subject to sporadic reappearances in areas that have a priori nothing in common. For example, the Gini mean difference (GMD) used in inequality economics -- and sometimes also considered as an $L_1$ alternative to the standard deviation --, or the so-called Lukaszyk-Karmowski metric used in mechanical physics or in quantum physics, have been proposed independently, according to the specific needs of their domain. Both are expressions of the statistical distance $E|X-Y|$ between two random variables $X$ and $Y$. In the case of GMD, $X \hbox{ and } Y$ are independent and identically distributed (i.i.d), whilst in the case of the Lukaszyk-Karmowski metric they are usually assumed to be independent. Not surprisingly, $E|X-Y|$ appears under different names in the literature, including expected (mean, average) absolute difference (deviation) between variables, ${\cal L}_1$- distance, $L_1$-distance, $L_1$-norm distance, $L_1$-metric, 1-average compound metric. In this document, $E|X-Y|$ will almost always be referred to as the expected absolute difference between $X \hbox{ and } Y$.

As we will frequently encounter the notion of “distance”, “semimetric”, “metric” and “metametric”, it is certainly not useless to specify the mathematical properties that these words cover. Let $E$ be a set and consider a function $d:E\times E\longrightarrowI\!\!R_+ := [0,\infty )$. This non-negative function may have various properties that must hold for all $x$, $y$, $z\in E$:

enumerate$d(x,x)=0$ \hskip .3cm(reflexivity) • $d(x,y)=0 \Rightarrow x=y$ \hskip .3cm(reverse reflexivity) • $d(x,y)=d(y,x)$ \hskip .3cm(symmetry) • $d(x,y) \leq d(x,z) + d(z,y)$ \hskip .3cm(triangle inequality)

If $d$ satisfies reflexivity and symmetry, it is called a distance and the ordered pair $(E,d)$ is called a distance space. If $d$ satisfies reflexivity, symmetry and triangle inequality, it is a semimetric and $(E,d)$ is a semimetric space. If $d$ satisfies reflexivity, reverse reflexivity, symmetry and triangle inequality, it is a metric and $(E,d)$ is a metric space. Reflexivity and reverse reflexivity together constitute the so-called \emph{identity of indiscernibles}. We will exercise some latitude when using the word “distance” even if we are actually talking about metrics or semimetrics.

This work organizes the reflexion around $E|X-Y|$ in the following way: Section (ref) shows how the $L_1$-distance between distribution functions\footnote{which is also the 1-Wasserstein distance between the corresponding distributions.} and the expected absolute difference are related and how two separate experiments can be consistently unified. Section (ref) deals with the axiomatic of probability metrics, $E|X-Y|$ being what is called a compound metric. On this occasion, a well-hidden logical inconsistency tainting published work is brought to light. Section (ref) focuses on the behavior of $E|X-Y|$ in relation to independence, almost sure equality, or equality of distribution of the variables $X \hbox{ and } Y$. The normalized expected absolute difference is discussed in Section (ref), where the prominent role of independence is emphasized. A primary metric defined on the distributions of pairs of random variables is specified in \textbf{Section (ref)}, in order to provide a new interpretation of the expected absolute difference. This leads to a very general expression for the \emph{Gini mean difference} and the \emph{Gini index}. In \textbf{Section (ref)}, we give in analytic form the expected absolute difference between two independent normally distributed random variables. We end up with a result generalizing formulas used in applied physics an in economics. In the process, we also give the analytic form of the average distance between coordinates of points falling at random into a proper rectangle of $I\!\!R^2$. We envision that \textbf{Section (ref)} can provide the basic background material to understand the main concepts of the optimal transport theory. The latter is consciously presented in a restricted framework, as a first step in the access to a complex field in full expansion. It is precisely these restrictions that allow to bring to light very telling results, sometimes even spectacular, in any case of a indeniable mathematical beauty. In this context, the presence of closed-form solutions to the optimal transport problem allows - or at least greatly facilitates - a good understanding of the subject through important special cases.

Here are some of the notations used throughout this article: we write $I\!\!R_+$ for the set of non-negative real numbers, and ${\cal B}_d$ refers to the $\sigma$-field of Borel subsets of $I\!\!R^d$, $d\ge 1$. Moreover, ${\cal P}(I\!\!R^d)$ denotes the set of probability measures on $(I\!\!R^d,{\cal B}_d)$ and ${\cal P}_p(I\!\!R^d)$ the set of probability measures on $(I\!\!R^d,{\cal B}_d)$ with finite $p$-th moment. We will be mainly interested in random variables $X,Y,Z,\ldots$ defined on $(\Omega,{\cal A},P)$ taking their values in $(I\!\!R,{\cal B}_1)$. Let $V:(\Omega,{\cal A},P) \rightarrow (I\!\!R^d,{\cal B}_d)$ be a random variable or a random vector defined on a given probability space. We denote by $P_V$ the pushforward probability of $P$ induced by $V$ on $(I\!\!R^d,{\cal B}_d)$. For example, $P_X$, $P_{(X,Y)}$, $P_{(X,Y,Z)}$ refer to pushforward probability measures of $P$ on $(I\!\!R,{\cal B}_1)$ (resp. $(I\!\!R^2,{\cal B}_2)$, $(I\!\!R^3,{\cal B}_3)$) induced by $X$ (resp. $(X,Y)$, $(X,Y,Z)$). The notation ${\cal L}_1(I\!\!R^d)$, $d\ge 1$, will refer to the space of integrable random variables or vectors defined on a probability space $(\Omega,{\cal A},P)$ which take their values in $(I\!\!R^d,{\cal B}_d)$. Almost sure equality (resp. equality of distribution) of random variables $X \hbox{ and } Y$ are denoted by $X\stackrel{a.s.}{=} Y$ (resp. $X\stackrel{d}{=} Y$). The respective abbreviations cdf and pdf stand for cumulative distribution function and probability density function.

On distances based on absolute difference

Subsection (ref) focuses on the relationship between the Gini-Kantorovich distance (a $L_1$-distance between cdf's) and the expected absolute difference (a $L_1$-norm distance between random variables). We discuss properties of these distances and examine the historical premises of the optimization problem at the origin of the link that unites them. Subsection (ref) discusses how two separate experiments can be consistently unified.

How $L_1$-distance between cumulative distribution functions and ${\cal L}_1$-distance between random variables are related

A statement such as “random variables $X$ and $Y$ are defined on the same probability space” implies that $X$ and $Y$ have a joint distribution, in which case we say that they are coupled. First, consider two integrable real-valued random variables $X$ and $Y$ that may not be defined on the same probability space. If one knows their (individual) distributions only, one can define a distance between them by using their respective cumulative distribution functions $F_X$ and $F_Y$. A rather intuitive way of measuring this distance is to calculate the Gini-Kantorovich distance

equation[equation omitted — 77 chars of source]

which can be easily visualized as a surface between two curves. One can show that

equation[equation omitted — 120 chars of source]

where $F_X^-$ and $F_Y^-$ are the quantile functions or generalized inverses of $F_X$ and $F_Y$, respectively. A proof of this remarkable coincidence is given in Thorpe (2018)\nocite{Thorpe2018}, see also Rachev and Rueschendorf (1998)\nocite{Rachev1998}. The generalized inverse is defined in Subsection (ref) (Definition (ref)). For more details about quantile functions, see Karr (1993)\nocite{Karr1993} p. 63, or Embrechts and Hofer (2014)\nocite{Embrechts2014}.

$GK$, also known as the Gini index of dissimilarity\footnote{$GK$ is also called the $L_1$-metric between distribution functions or Monge-Kantorovich metric or Hutchinson metric, 1-Wasserstein metric, Fortet-Mourier metric, see Deza and Deza (2014)\nocite{Deza2014}.}, is a special case of a more general Gini-Kantorovich metric $GK_p$ when $p=1$ (Ortobelli et al. (2006)\nocite{Ortobelli2006}). The Gini index of dissimilarity should not be confused with the Gini mean difference (GMD) or the Gini index discussed later in this article (Subsections (ref), (ref) and (ref)). Rachev et al. (2013)\nocite{Rachev2013}, note that ((ref)) is the explicit solution of a minimization problem studied by Gini (1914)\nocite{Gini1914a} and solved by Salvemini (1943)\nocite{Salvemini1943} for discrete cdf's and Dall'Aglio (1956)\nocite{Dallaglio1956} in the general case. More precisely, let ${\cal F}(F_1,F_2)$ denote the set of all bivariate cdf's $F$ with marginal cdf's $F_1$ and $F_2$. Then the analytic solution of the minimization problem

equation[equation omitted — 111 chars of source]

is “simply”

equation[equation omitted — 70 chars of source]

The optimization problem ((ref)) and its solution are often expressed as

equation[equation omitted — 140 chars of source]

where $X \hbox{ and } Y$ are given random variables and where “$\stackrel{d}{=}$” refers to equality of distribution (Rachev et al. (2007), Eq. 3.23\nocite{Rachev2007})\footnote{Rachev et al. (2013)\nocite{Rachev2013} formulate concisely the more general problem of mass transportation studied by Kantorovich, of which the classic transportation problem in linear programming and the minimization problem ((ref)) are special cases.}. Actually, as the map $c(x,y) := |x - y|$ is continuous\footnote{or even lower semi-continuous, a weaker condition. In optimal transport, $c(x,y)$ denotes the transportation cost function.}, a minimimizer does exist (Gangbo (2004), Th. 2.4)\nocite{Gangbo2004} and we can replace “inf” by “min” in ((ref)) and ((ref)). Typically, ((ref)) shows that the infimum runs over all couplings\footnote{Couplings are defined in Subsection (ref). } $(\tilde{X},\tilde{Y})$ of $X \hbox{ and } Y$, where $X \hbox{ and } Y$ may or may not be defined on the same probability space. Now suppose that $X \hbox{ and } Y$ are both defined on a probability space $(\Omega,{\cal A},P)$, i.e. are jointly distributed. Then a look at ((ref)) -- where $GK$ depends only on the individual cdf's of $X \hbox{ and } Y$ -- confirms that $GK$ ignores any structure of dependence or independence inside the pair $(X,Y)$. In the process of minimization described in ((ref)), of which $GK(X,Y)$ is the solution, any dependence or independence structure is swept away: we are left only with a probability metric measuring the $L_1$-distance between the cdf's $F_X$ and $F_Y$ of $X \hbox{ and } Y$, respectively. Such a situation is unsatisfactory because it ignores valuable information that can be available in practice, for example when one can assume that $X \hbox{ and } Y$ are independent. This can be illustrated with a simple example: Table (ref) shows two joint distributions of binary $\{0,1\}$-valued random variables.

table[table omitted — 997 chars of source]

Distribution (a) in Table (ref) reflects a dependence between $X$ and $Y$, while distribution (b) corresponds to independence. In both cases, $GK(X,Y) = 0.5$, ignoring the dependence structure between the variables. Note that $E|X-Y| = 0.7$ (case (a)) and 0.62 (case (b)). A probability metric such as $GK(X,Y)$ makes sense if $X$ and $Y$ are uncoupled, i.e. if we only know their one-dimensional cdf's. When $X \hbox{ and } Y$ are coupled, there are more informative ways of determining how far apart they are from each other. As $GK$ cannot take full account of the information of the model, it may be replaced by the expected absolute difference. As a matter of fact, $E|X-Y|$ uses all the information contained in the probability space $(\Omega,{\cal A},P)$ governing the distribution of the pair $(X,Y)$ to determine how far $X$ is from $Y$. An emblematic case occurs when $X \hbox{ and } Y$ can be assumed to be independent. It is well-known that under the independence assumption $X: (\Omega_1,{\cal A}_1,P_1) \rightarrow (I\!\!R, {\cal B}_1)$ and $Y: (\Omega_2,{\cal A}_2,P_2) \rightarrow (I\!\!R, {\cal B}_1)$ can be defined trivially on the same probability space, namely the product space $(\Omega,{\cal A},P) = (\Omega_1 \times \Omega_2,{\cal A}_1 \otimes {\cal A}_2,P_1 \otimes P_2)$. Consequently, $X \hbox{ and } Y$ are jointly distributed and the use of $GK$ (or any other similar metric) would be inappropriate.

Consistent unification of two separate experiments

(The reader familiar with measure or probability theory may skim this subsection). Fundamental probabilistic concepts, although often trivialized in applied papers, are not always sufficiently understood. We have seen in the previous section how the fact that random variables are jointly distributed or not can affect the choice of an adequate distance function. In connection with the content of Subsection (ref), we recall the rules that must be respected so that two separate experiments can be adequately combined into a joint experiment.

Consider two random experiments ${\cal E}_1$ and ${\cal E}_2$. Suppose that the information on the experiments is captured by real numbers; that is, ${\cal E}_1$ and ${\cal E}_2$ are completely described by the respective probability spaces $(I\!\!R, {\cal B}_1,P_1)$ and $(I\!\!R, {\cal B}_1,P_2)$\footnote{For simplicity, we suppose that $I\!\!R$ is the common sample space of ${\cal E}_1$ and ${\cal E}_2$ and that $I\!\!R$ is endowed with ${\cal B}_1$, to obtain the common measurable space $(I\!\!R,{\cal B}_1)$. }. If the unification operation is conducted in a coherent way, then ${\cal E}_1$ and ${\cal E}_2$ can be seen as “marginal” experiments of ${\cal E}$. A probability model $(\Omega,{\cal A},P)$ has to be defined to describe the “joint experiment” ${\cal E}$.

Random variables $X$, $Y$ -- and the resulting pair $(X,Y)$ taking values in $(I\!\!R^2, {\cal B}_2)$ -- can be defined on a common probability space $(\Omega,{\cal A},P)$ so that $P_1$ and $P_2$ are the marginal probability measures of $P$. Indeed, we can define $\Omega = I\!\!R^2$, ${\cal A} = {\cal B}_1 \otimes {\cal B}_1 = {\cal B}_2 $ and $(X,Y)= Id_2 = (q_1,q_2)$, where $Id_2$ is the identity map on $I\!\!R^2$ and the $q_i$'s are the corresponding projection functions (i.e. for $(x_1,x_2)\in I\!\!R^2$, $q_i(x_1,x_2) = x_i$, $i=1,2$). So $X \hbox{ and } Y$ can be interpreted indifferently as random variables or as projections. In order for $P_1$ and $P_2$ to be the marginals of some probability measure $P$, and noting that $P(q_1^{-1} B_1) = P(B_1 \times I\!\!R)$ and $P(q_2^{-1} B_2) = P(I\!\!R \times B_2)$, the following consistency conditions are imposed:

equation[equation omitted — 141 chars of source]

for all $B_1, B_2 \in {\cal B}_1$. Of course the conditions in ((ref)) are not sufficient to fully determine $P$, a feature that was predictable since the link between the two experiments was not specified. In the particular case where the two experiments (and hence the two random variables $X$ and $Y$) are assumed to be independent, $P$ must satisfy the additional condition

equation[equation omitted — 75 chars of source]

for all $B_1,B_2 \in {\cal B}_1$. In other words, $P$ is the product probability $P_1\otimes P_2$, which is uniquely defined on ${\cal B}_2$. Moreover $P_1 = P_X$, $P_2 = P_Y$ and $P = P_{(X,Y)}$ in the above construction. The process just described consists of two steps that are worth distinguishing. \vskip .2cm First step: Two random experiments ${\cal E}_1$ and ${\cal E}_2$ are united to obtain a measurable space for the resulting joint experiment ${\cal E}$. \vskip .2cm Second step: A probability measure (and the resulting dependence structure between $X \hbox{ and } Y$) is enforced on the model set up in the first step.

Figuratively, one could say\footnote{without reference to the Marxist phraseology...} that the first step corresponds to an “infrastructure" on which a probabilistic “superstructure" is built in the second step. \vskip .2cm

One can illustrate, in terms of $\sigma$-fields, the qualitative leap following the coupling of two previously separate (stand-alone) random experiments ${\cal E}_1$ and ${\cal E}_2$. Under the above assumptions and the ensuing construction, the information that ${\cal E}_1$ and ${\cal E}_2$ provide is carried by the random variables $X:(I\!\!R^2,{\cal B}_2) \longrightarrow (I\!\!R,{\cal B}_1 )$ and $Y:(I\!\!R^2,{\cal B}_2) \longrightarrow (I\!\!R,{\cal B}_1 )$. The minimal $\sigma$-fields generated by $X$ and $Y$ are denoted by $\sigma (X)$ and $\sigma (Y)$, respectively. In turn, $\sigma (X)$ and $\sigma (Y)$ generate the $\sigma$-field

equation[equation omitted — 101 chars of source]

As by construction $X = q_1$ and $Y = q_2$, we have: $\sigma (X) = \{B_1 \times I\!\!R : B_1 \in {\cal B}_1 \}$ and $\sigma (Y) = \{I\!\!R \times B_2: B_2 \in {\cal B}_1 \}$. It can be shown without much difficulty that $\sigma (X) \vee \sigma (Y) = \sigma ((X,Y))$, the $\sigma$-field on $I\!\!R^2$ induced by the pair $(X,Y)$, noting that $\sigma ((X,Y)) = {\cal B}_2 = \sigma({\cal C})$, where ${\cal C} = \{B_1 \times B_2 : B_1, B_2 \in {\cal B}_1 \}$\footnote{That is, ${\cal B}_2$ is generated by the $(B_1 \times B_2)$'s, but ${\cal B}_2 = \sigma (X) \vee \sigma (Y)$ implies that ${\cal B}_2$ is also generated by the union of the $(B_1 \times I\!\!R)$'s and the $(I\!\!R \times B_2)$'s.}. Considering $X$ and $Y$ together as a pair of random variables instead of two stand-alone random variables allows to prepare a much wider portion of ${\cal B}_2$ (infrastructure) on which a probability measure (superstructure) can be defined. The representation $\sigma (X) \vee \sigma (Y)$, probably more telling than $\sigma ((X,Y))$, is symbolized in Figure (ref).

figure[figure omitted — 376 chars of source]

Incidentaly, note that the intersection $\sigma (X) \cap \sigma (Y)$ is not empty, since it contains $\emptyset$ and $I\!\!R^2$.

Axiomatic approach of probability metrics

Errors of interpretation about the notion of distance between random variables occur because important concepts are not formally stated or not sufficiently explained. The purpose of this section is to eliminate ambiguities encountered here and there in applied science works published in journals or on the Internet. The role of a metric is to define a distance between elements of a same set. We mentioned in the introduction that to verify precisely whether a functional is adequate to cover what is commonly understood by the notion of distance, it is unavoidable to specify certain elementary, natural and intuitive rules -- axioms -- that this functional should satisfy. Again, we limit the discussion to real-valued random variables. Rachev (1991)\nocite{Rachev1991} provides a more general treatment.

The idea of distance between two random variables $X \hbox{ and } Y$ is linked in a decisive way to what is meant by $X \hbox{ and } Y$ "are the same", "are coincident", "are indistinguishable". And the concept of sameness, coincidence or indistinguishably -- treated here as synonyms -- can be mathematically captured by an equivalence relation: $X \hbox{ and } Y$ are declared to be the same, coincident or indistinguishable if they are in the same equivalence class. It is therefore natural to define a metric not on an initial set of random variables, but on the quotient space resulting from the adequate equivalence relation.

The theory of probability metrics considers three categories of metrics defined according to the type of equivalence deemed useful in a given context.

Primary, simple and compound metrics

We assume throughout this section that real-valued random variables are defined on one and the same probability space $(\Omega,{\cal A},P)$. Denote by ${\cal V}$ the set of all random variables on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R, {\cal B}_1)$.\footnote{${\cal V} = {\cal L}_1(I\!\!R)$ when the random variables are assumed to be integrable.} So ${\cal V}^2 = {\cal V}\times {\cal V}$ is the set of all random pairs defined on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R^2, {\cal B}_2)$.

Let $\psi : {\cal V} \rightarrow I\!\!R$ be a functional candidate to be a distance. We would like $({\cal V},\psi )$ to be a metric space from the outset, but this wish is thwarted by the fact that we are working with random variables -- that is relatively complex mathematical objects. It is nevertheless possible to use $\psi$ to build a metric space defined on equivalence classes. For this purpose, let us define an equivalence relation on ${\cal V}$ by

equation[equation omitted — 89 chars of source]

What properties of $\psi$ will ensure that ((ref)) is a reflexive, symmetric and transitive relation, i.e. is an equivalence relation? It is easily verified that the following axioms meet this requirement: for all $X,Y,Z \in {\cal V}$, \newline (a) $\psi (X,X) = 0$ \hskip .3cm(reflexivity) \newline (b) $\psi (X,Y) = \psi (Y,X)$ \hskip .3cm(symmetry) \newline (c) $\psi (X,Z) \leq \psi (X,Y) + \psi (Y,Z)$ \hskip .3cm(triangle inequality), \newline noting that these axioms imply the non-negativity of $\psi$. So (a), (b) and (c), beyond their natural and intuitive content, are sufficient to ensure that $\stackrel{\psi}{\sim}$ is an equivalence relation. We call a functionnal satisfying (a), (b) and (c) a probability metric, although it is rigorously a semimetric\footnote{The terminology is still fluctuating: in topology, what we define here as a semimetric is called a pseudometric. }. Now, $X \stackrel{\psi}{\sim} Y$ means different things depending on the choice of $\psi$, because $\psi(X,Y) = 0$ may be true if and only if $X \hbox{ and } Y$ \newline (i) share a given set $c_1, c_2,\ldots$ of characteristics such that $c_i(X) = c_i(Y)$, $i= 1,2,\ldots$ (e.g. $X \hbox{ and } Y$ have the same mean and the same variance, as in ((ref)) below), or\newline (ii) have the same distribution, or \newline (iii) are almost surely equal.

So $\stackrel{\psi}{\sim}$ in ((ref)) is a generic notation for specific equivalence relations: \newline $X \stackrel{c}{=} Y$ when $X \hbox{ and } Y$ share a given set of characteristics, or \newline $X \stackrel{d}{=} Y$ when $X \hbox{ and } Y$ have the same distribution, or \newline $X \stackrel{a.s.}{=} Y$ when $X \hbox{ and } Y$ are almost surely equal.

defi(Rachev et al. (2011) \nocite{Rachev2011})\newline If $\stackrel{\psi}{\sim}$ means $\stackrel{c}{=}$, then $\psi$ is called a primary (probability) metric.\newline If $\stackrel{\psi}{\sim}$ means $\stackrel{d}{=}$, then $\psi$ is called a simple metric.\newline If $\stackrel{\psi}{\sim}$ means $\stackrel{a.s.}{=}$, then $\psi$ is called a compound metric\footnote{The intervention of three random variables (instead of two) in the triangle inequality axiom raises a theoretical issue in the compound metrics case. The pairs $(X,Y)$, $(X,Z)$ and $(Y,Z)$ can be chosen in such a way that there exists a random vector $(X,Y,Z)$ ensuring that the three pairs are its two-dimensional projections. For more information, see the so-called "gluing lemma" (Thorpe (2018, Lemma 5.5)) which allows to "glue" two (or more) bivariate (multivariate) distributions so as to respect the different marginals. We refer to our discussion in paragraph (ref), where the consistency rule is stated. The triangle inequality does not hold for all random variables $X,Y,Z$, but only for those satisfying this rule. }.

Primary metrics correspond to the weakest form of equivalence. Random variables $X \hbox{ and } Y$ can be considered equivalent if they have the same mean and the same standard deviation. A plain example of primary metric is

equation[equation omitted — 101 chars of source]

where $\sigma$ refers to the standard deviation\footnote{Note that ((ref)) makes sense because the standard deviation and the mean are defined in the same unity. }.

The simple metrics imply a stronger form of sameness: $X \hbox{ and } Y$ are considered equivalent if their cdf's are identical (remembering that a random variable is completely described by its cdf). An example of simple metric is the Gini-Kantorovich metric GK given in ((ref)). GK measures the distance between $X \hbox{ and } Y$ -- which are assumed to have finite first moment -- by a distance between their respective cdf's.\footnote{If $X \hbox{ and } Y$have respective probability distributions $\mu$ and $\nu$, it is remarkable that $GK(X,Y)$ coincides with the 1-Wasserstein metric $W_1(\mu,\nu)$. It turns out that the latter is none other than the minimal cost in the Monge-Kantorowich transport problem (Kantorovich (1942)\nocite{Kantorovich1942}, see e.g. Villani (2008)\nocite{Villani2008} or Thorpe (2018)). Rachev (2007)\nocite{Rachev2007} uses the notation $GK(X,Y) = \min E|X - Y|$. This is very telling since (i) due to the fact that $|x - y|$ is a continuous cost function in the Monge-Kantorovich optimal transport problem, the infimum is realized (Gangbo (2004)\nocite{Gangbo2004}) and (ii) it remains us that $GK(X,Y)$ may be seen as the solution of this celebrated minimization problem. }

The compound metrics represent the strongest form of sameness. The simplest example of compound metric is probably the expected absolute difference $\psi (X,Y) = E|X - Y|$. Importantly, here $X \hbox{ and } Y$ are necessarily defined on the same probability space (i.e. are jointly distributed) and have finite first moment.

We will now show that there is a (true) metric $\tilde{\psi}$ derived from the semimetric $\psi$, where $\tilde{\psi}$ is defined on equivalence classes. The classes, denoted by $[\cdot]$, stem from the equivalence relation $\stackrel{\psi}{\sim}$, which can mean, as we have seen above, $\stackrel{c}{=}$, $\stackrel{d}{=}$ or $\stackrel{a.s.}{=}$. Define in a canonical way

equation[equation omitted — 69 chars of source]

We must show that $\tilde{\psi}$ is well-defined, i.e. does not depend on the representatives chosen to designate the classes. To show that ((ref)) makes sense, we need the following lemma.

lemme(quadrilateral inequality, proof in the appendix) \newline Let $\psi : {\cal V} \rightarrow I\!\!R$ be a functional satisfying non-negativity, symmetry and triangle inequality. Then for any $X,Y,X_1, Y_1 \in {\cal V}$, \begin{equation} |\psi (X,Y) - \psi (X_1,Y_1) | \leq \psi (X,X_1) + \psi (Y,Y_1). \end{equation}

Assume that $\psi$ satisfies reflexivity, symmetry and triangle inequality (and therefore nonnegativity, so that the conditions of Lemma (ref) are fulfilled) and let the equivalence relation on ${\cal V}$ be given by ((ref)). Suppose that $X_1\stackrel{\psi}{\sim}X$ and $Y_1\stackrel{\psi}{\sim}Y$. Then, using ((ref)) and the symmetry of $\psi$, we get $\psi (X_1,Y_1) = \psi (X,Y)$, that is $\tilde{\psi} ([X_1],[Y_1]) = \tilde{\psi} ([X],[Y])$, which proves that $\tilde{\psi}$ in ((ref)) is well-defined.

Let us write $\tilde{{\cal V}}$ for the quotient space ${\cal V}/\stackrel{\psi}{\sim} \ = \{[X]: X \in {\cal V} \}$. We are now able to state that $(\tilde{{\cal V}},\tilde{\psi})$ is a metric space, i.e. that $\tilde{\psi}$ satisfies the following axioms:

enumerate$[X] = [Y] \Rightarrow \tilde{\psi} ([X],[Y])=0 $ \hskip .3cm(reflexivity) • $\tilde{\psi} ([X],[Y])=0 \Rightarrow [X] = [Y]$ \hskip .3cm(reverse reflexivity) • $\tilde{\psi}([X],[Y])=\tilde{\psi} ([Y],[X])$ \hskip .3cm(symmetry) • $\tilde{\psi} ([X],[Z]) \leq \tilde{\psi} ([X],[Y]) + \tilde{\psi} ([Y],[Z])$ \hskip .3cm(triangle inequality).

The first two axioms -- known as the identity of indiscernibles when taken together -- are the consequence of the definitions of $\stackrel{\psi}{\sim}$ and $\tilde{\psi}$, while axioms 3 and 4 stem directly from the symmetry and triangle inequality property of $\psi$.

Note that if ${\cal V} = {\cal L}_1(I\!\!R)$ and $\psi (X,Y) = E|X-Y|$, then $\stackrel{\psi}{\sim}$ means $\stackrel{a.s.}{=}$. In this case $\tilde{{\cal V}} = L_1(I\!\!R)$ and $\tilde{\psi}$ is the metric induced by the $L_1$-norm.

Uncovering a logical inconsistency

We limit our discussion to $E|\cdot -\cdot |$, which implies that ${\cal V} = {\cal L}_1(I\!\!R)$ (${\cal V}$ defined in Section (ref)), but our conclusions can be generalized to other compound metrics.

Focusing on the two equivalence relations $\stackrel{a.s.}{=}$ and $\stackrel{d}{=}$ which group variables belonging to ${\cal L}_1(I\!\!R)$, we denote by $[\cdot]_{a.s.}$ and $[\cdot]_{d}$, the corresponding equivalence classes. Why are we so eager to identify certain elements of ${\cal L}_1(I\!\!R)$? It is because it allows $E|X-Y|$ to switch from a semimetric to a metric. Indeed, if we define $\tilde{\psi} ([X]_{a.s.},[Y]_{a.s.}) = \psi (X,Y) = E|X-Y|$, then $E|\cdot - \cdot|$ represents both $\tilde{\psi}$ (a metric defined on classes) and $\psi$ (a semimetric defined on random variables). We saw in Section (ref) that $\tilde{\psi}$ realizes the identity of indiscernibles (reflexivity and reverse reflexivity), whereas $\psi$ satisfies reflexivity ($\psi (X,X) = 0$), but not reverse reflexivity ($\psi (X,Y) = 0$ does not imply $X = Y$). Moreover, $\psi$ and $\tilde{\psi}$ both satisfy symmetry and triangle inequality.

That said, some authors using $E|X-Y|$ have fallen into the trap of identifying within ${\cal L}_1(I\!\!R)$ the identically distributed random variables rather than the almost surely equal random variables. Unfortunately, this leads to a logical impasse. Seeking a contradiction, suppose that we set

equation[equation omitted — 70 chars of source]

where $X,Y \in {\cal L}_1(I\!\!R)$ are identically distributed without being almost surely equal. We have $[X]_d = [Y]_d$ and $E|X-Y| > 0$ (since $E|X-Y| = 0 \Leftrightarrow X\stackrel{a.s.}{=} Y$), and we end up with the following contradiction: \newline $\tilde{\psi} ([X]_d,[X]_d) \stackrel{(\ref{interdit})}{=} E|X-X|=0 < E|X-Y| \stackrel{(\ref{interdit})}{=} \tilde{\psi} ([X]_d,[Y]_d) = \tilde{\psi} ([X]_d,[X]_d)$, meaning that ((ref)) is ill-defined.

To convince ourselves that the above discussion is not in vain, take the case of the so-called Lukaszyk-Karmowski metric (Lukaszyk 2004)\nocite{Lukaszyk2004} which is actually the functional $\tilde{\psi} ([\cdot]_d,[\cdot]_d) = E|\cdot - \cdot|$ set in ((ref)). It is only when this author asserts that the identity of indiscernibles property is not realized by the metric $E|\cdot - \cdot|$ he uses that we end up understanding he implicitly identifies the identically distributed random variables of ${\cal L}_1(I\!\!R)$. In other words, he reasons as if ((ref)) were well-defined\footnote{ In later publications of Lukaszyk (and on Wikipedia, etc.), where $D(X,Y)$ refers to $E|X-Y|$, we find the notation $D(X,X) > 0$, which proves that the identically distributed variables of ${\cal L}_1(I\!\!R)$ are implicitly identified. This leads to the logical contradiction we have put forward. }. On this erroneous basis, Lukaszyk claims to have used a new operator which, as such, would deserve a special denomination. This does not make sense and the so-called Lukaszyk-Karmowski metric is none other than the good old $L_1$-norm distance. Note that this clarification does not greatly affect the merit of this author's 2004 article, where otherwise conclusive results in applied physics are presented. Lukaszyk correctly computes the distance $D(X,Y) = E|X-Y|$ between independent elements of ${\cal L}_1(I\!\!R)$ -- notably when $X \hbox{ and } Y$ are Gaussian -- but he should not pretend that the reflexivity condition does not hold.

Expected absolute difference of independent random variables

General considerations

Why is the notion of independence so important? This section contains a few remainders and general thoughts about independence. Two random variables are independent if they come from phenomena such that the result observed for one of them has no influence on the other. We should admit that the assumption of independence is often a question of intuition or common sense -- although in this respect caution is needed. When the hypothesis of independence can be made reasonably, great mathematical simplifications follow (resulting in particular from the use of the Fubini-Tonelli theorem).

Contrary to a common belief -- at least among researchers having somewhat forgotten the fundamentals of probability theory -- the notion of independence between two random variables $X$ and $Y$ does not imply that they “have nothing to do with each other”. Indeed, independence only makes sense if these random variables are coupled, i.e. defined on the same probability space. They are coupled to each other, albeit in a particular way, by the fact that they are independent.

In applied sciences, two observations, phenomena or experiments can be perceived as independent. Independence may be imposed from the outside or may be organized in full awareness by the experimenter. In both cases, he or she will be interested in forming a probability model for a joint experiment such that the original two experiments are carried out independently.

Calculating a distance between independent random variables turns out to be often very useful. For example two measurement devices $X$ and $Y$ independently measure unknown quantities with some random error, or multiple researchers independently measure the same object and compare their results. To take an example related to economic inequalities, suppose that $X$ (resp. $ Y $) is the income of a household drawn at random from a statistical population ${\cal P} _1$ (resp. ${\cal P} _2$). Let $\mu \hbox{ and } \nu$ denote the income distribution of $X \hbox{ and } Y$, respectively. The random variables $X \hbox{ and } Y$ are assumed to be independent, and we use this information to compute $E|X-Y|$, which is interpreted as a measure of income disparity between the two populations. In other words, $X=x$ and $Y=y$ are independently observed and $E|X-Y|$ is the weighted average of the $|x-y|$'s. We have already mentioned that a metric such as $GK(X,Y)$, which depends only on the stand-alone distributions of $X$ and $Y$, cannot take account of the independence information.\footnote{We will see in Section (ref) that in case of a single population, i.e. if ${\cal P}_1 = {\cal P} _2$ and $\mu = \nu$, and if the independent random variables $X$ and $Y$ are non-negative, then $E|X-Y|$ is none other than the Gini mean difference, a measure of income inequality within a population. The normalized Gini mean difference is the celebrated Gini index.}

Expected absolute difference in the context of almost sure equality, equality of distribution and independence

Consider $(X,Y) \in {\cal L}_1(I\!\!R^2)$, the space of all pairs of integrable real-valued random variables defined on a probability space $(\Omega,{\cal A},P)$. It is well-known that $X\stackrel{a.s.}{=} Y$ ($X = Y$ almost surely) if and only if $E|X-Y| =0$. A pair $(X,Y) \in {\cal L}_1(I\!\!R^2)$ with $E|X-Y| =0$ is such that the probability mass $P_{(X,Y)}$ is concentrated on the diagonal $\Delta:=\{(x,x):x\inI\!\!R\}$ of $I\!\!R^2$, (that $P_{(X,Y)}(\Delta ) = 1$ is formalized in Subsection (ref), Proposition (ref)).

Here are some remarks about the values $E|X-Y|$ can take: if the distribution of the random variables $X$ and $Y$ differ, then $X$ and $Y$ cannot be almost surely equal, and consequently $E|X-Y| > 0$. Moreover, the fact that $X$ and $Y$ have the same distribution by no means implies that $E|X-Y| =0$: suppose that $X$ and $Y$ are two continuous independent and identically distributed (i.i.d.) random variables. In this case $P({X\neq Y}) = 1$, i.e. $X \neq Y$ a.s., which implies that $E|X-Y| > 0$. For example, take two i.i.d. random variables $X \hbox{ and } Y$ having standard normal distribution. Then ((ref)) -- a consequence of Theorem (ref) -- implies that $E|X-Y| = 2/\sqrt{\pi}$.

It is interesting to take a closer look at the behavior of $E|X-Y|$ when $X$ and $Y$ are independent.

prop(Proof in the appendix) Let $X, Y \in {\cal L}_1(I\!\!R)$ be independent. The following statements are equivalent \begin{description} • $X\stackrel{a.s.}{=} Y$. • $X$ and $Y$ are a.s. equal to a same constant, i.e. there exists $c\in I\!\!R$ such that $X\stackrel{a.s.}{=} Y\stackrel{a.s.}{=}c$. \end{description}

A random variable is said to be degenerate if it is almost surely constant. So, if $X \hbox{ and } Y$ are independent, then $E|X-Y| = 0$ occurs if and only if they are (identically distributed and) degenerate, i.e. if and only if their distribution in concentrated on the same constant.

The above considerations can be summarized as follows: \newline \[ E|X-Y| = 0 \Rightarrow \left\{

array[array omitted — 212 chars of source]

\right. \] a situation illustrated in Figure (ref).

figure[figure omitted — 409 chars of source]

Partitioning ${\cal L}_1(I\!\!R^2)$ into six categories

We are interested in random pairs $(X,Y) \in {\cal L}_1(I\!\!R^2 )$. In order to bring together and clarify the concepts encountered in Section (ref), define \newline ${\cal E}_{a.s} = \{ (X,Y) \in {\cal L}_1(I\!\!R^2 ) : X \stackrel{a.s.}{=} Y \}$ and ${\cal E}_{d} = \{ (X,Y) \in {\cal L}_1(I\!\!R^2 ) : X\stackrel{d}{=} Y \}$. That is, the subsets ${\cal E}_{a.s}$ and ${\cal E}_{d} $ of ${\cal L}_1(I\!\!R^2 )$ contain the random pairs whose components are equivalent in the almost sure sense and in the equality of distribution sense, respectively. Taking into account a possible independence between $X \hbox{ and } Y$, the pairs $(X,Y) \in {\cal L}_1(I\!\!R^2 )$ fall into six mutually exclusive categories described in Table (ref).

table[table omitted — 791 chars of source]

Figure (ref) depicts the six categories of Table (ref) as a partition of ${\cal L}_1(I\!\!R^2 )$.

figure[figure omitted — 291 chars of source]
table[table omitted — 3,038 chars of source]
example\vskip .5cm Table (ref) displays the distributions of six pairs $(X,Y)$ of (binary) $\{0,1\}$-valued random variables. These bivariate distributions correspond to the categories defined in Table (ref) and represented in Figure (ref).

Independence and entropy

It would be incomplete to talk about independence without giving a brief overview of the concept of entropy, whose definition comes from the pioneers of information theory. The entropy of a distribution was established by Shannon for discrete laws and then extended to continuous laws characterized by their density. The notion of entropy has been generalized in many ways and has undergone vast and important developments. The real-valued random variables $X$ and $Y$ defined on $(\Omega,{\cal A},P)$, where $\Omega$ is a finite set, are independent when the degree of uncertainty is maximal in their joint distribution. This means that the entropy of the bivariate distribution must be maximal.

Entropy is a measure of uncertainty or randomness. The following equation holds: \newline $H(X,Y) = H(X) + H(Y) - I(X,Y) $, where $H(X,Y)$ is the (joint) entropy of the pair $(X,Y)$, $H(X)$ and $H(Y)$ being the entropy of $X$ and $Y$, respectively. The non-negative $I(X,Y)$ is the so-called mutual information, and one has $I(X,Y) = 0$ if and only if $X$ and $Y$ are independent, in which case $H(X,Y)$ is maximal. In the finite case that we are dealing with $H(X) = \sum_i p_i \ln(1/p_i)$, $H(Y) = \sum_j p_j \ln(1/p_j)$ and $H(X,Y) = \sum_i \sum_j p_{ij} \ln(1/p_{ij})$, where $p_i = P(X=x_i)$, $p_j = P(Y=y_j)$ and $p_{ij} = P(X=x_i, Y=y_j)$.

Example (ref) below illustrates that the distance $E|X-Y|$ between dependent $X$ and $Y$ can be both smaller or larger than it is when $X$ and $Y$ are independent.

example\footnote{Inspired from Prof. Dr. Svetlozar Rachev's online lecture on probability metrics (summer semester 2008), Institute for Statistics and Mathematical Economics, University of Karlsruhe.} Table (ref) shows the joint distributions of three pairs $(X,Y)$ of binary $\{0,1\}$-valued random variables. In the three cases, the marginal distributions are the same. \newline Case (a) refers to a distribution of dependent variables yielding minimal $E|X-Y|$. \newline Case (b) refers to independent random variables. \newline Case (c) displays a distribution of dependent variables yielding maximal $E|X-Y|$. \vskip .1cm The expected absolute difference, the joint entropy and the Gini-Kantorovich distance for the three distributions are summarized in Table (ref). We realize that, unlike entropy, $E|X-Y|$ does not culminate with independence.
table[table omitted — 1,217 chars of source]
table[table omitted — 730 chars of source]

Normalized expected absolute difference

This section focuses on the normalized form of the expected absolute difference and its characteristics. In particular, we address the issue of triangle inequality in relation to this functional. In an interesting paper, Yianilos (2002)\nocite{Yianilos2002} shows that symmetric set difference and Euclidean distance on $I\!\!R^d$ have normalized forms that remain metrics. We examine in this section the conditions under which $E|X-Y|$ can be $[0,1]$-normalized. The additional difficulty here is that we are not in a deterministic context. With our usual notation, ${\cal L}_1(I\!\!R)$ is the vector space of all integrable random variables on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R, {\cal B}_1)$, and the random variables used below belong to this set.

It is natural to consider forming relative distance measures. Converting $D(X,Y) := E|X-Y|$ to a normalized (or standardized) form may be very useful in the solution of certain problems, especially when relative errors are at stake, as is often the case in numerical analysis. Define a normalized counterpart $D_{norm}(\cdot,\cdot )$ of $D(\cdot,\cdot )$ by

equation[equation omitted — 195 chars of source]

so that $0\le D_{norm}(X,Y) \le 1$. The upper bound is reached when, say, $Y\stackrel{a.s.}{=}0$ with $E(|X|) > 0$, while the lower bound is reached when $X$ and $Y$ are almost surely equal and $E(|X|)$ $(= E(|Y|))$ is strictly positive. It is not our intention to comment here on the pros and cons of using a normalized distance measure. We will simply check whether the generally accepted axioms for a distance are verified. What we can say though is that $D_{norm}$ is likely to share the strengths and weaknesses of relative deterministic measures such as the Canberra metric (Lance and Williams (1967))\nocite{Lance1967}. Clearly, $D_{norm}$ is a distance since it satisfies nonnegativity, reflexivity ($D_{norm}(X,X) = 0$) and symmetry ($D_{norm}(X,Y) = D_{norm}(Y,X)$). Let $(X_i)_{i \in I} \subset {\cal L}_1(I\!\!R)$ denote a finite family of independent random variables. We show in Subsection (ref) that the triangle inequality holds in case of independence, but does not hold in general. More precisely, while $({\cal L}_1(I\!\!R),D_{norm})$ is a distance space, we show that $((X_i)_{i \in I},D_{norm})$ is a semimetric space (i.e. $D_{norm}$ is a distance satisfying the triangle inequality).

The triangle inequality issue

The independent case

A corollary of Theorem (ref) below is that $D_{norm}(\cdot ,\cdot )$ defined in ((ref)) satisfies the triangle inequality in the independence case.

theor(Proof in the appendix) Let $X,Y,Z \in {\cal L}_1(I\!\!R)$ be three (mutually) independent random variables and assume that at most one of these variables is almost surely zero. Then $E|X| + E|Z| > 0 $, $E|X| + E|Y| > 0 $ and $E|Y| + E|Z| > 0 $, and the following property is realized \begin{equation} \frac{E|X-Z|}{E|X| + E|Z|} \le \frac{E|X-Y|}{E|X| + E|Y|} + \frac{E|Y-Z|}{E|Y| + E|Z|} . \end{equation}

Proving this inequality was a particularly thorny exercise (see the proof in the appendix). Using ((ref)) and the definition of $D_{norm}$ in ((ref)), one can easily check that

equation[equation omitted — 82 chars of source]

for mutually independent $X,Y,Z$. Note that independence of these variables implies that the three pairs $(X,Z)$, $(X,Y)$ and $(Y,Z)$ intervening in the triangle inequality are the two-dimensional projections of the three-dimensional random vector $(X,Y,Z)$ having the product distribution $P_X \otimes P_Y \otimes P_Z$. In other words, the three pairs can be consistently embedded in a three-dimensional vector so that the triangle inequality makes sense (more details on this are given below).

The general case

A question naturally arises: would the triangle inequality ((ref)) hold in all cases if the assumption of independence were lifted in Theorem (ref) ? The answer is negative: with a little of craftmanship, one can find counterexamples such as the one resulting from the three bivariate distributions (A), (B), (C) of pairs of dependent random variables shown in Table (ref).

table[table omitted — 3,332 chars of source]

It turns out that $E|X| = E|Y| = E|Z| = 1$ and

equation[equation omitted — 97 chars of source]

which contradicts ((ref)). The statement that the triangle inequality should hold for any $X,Y,Z$ is actually pretty vague. As $E|X-Y|$ is a compound probability metric (see Definition (ref)), the choice of the three pairs $(X,Z)$, $(X,Y)$ and $(Y,Z)$ cannot be totally free. Indeed, suppose we fix the distributions (A) and (B) in Table (ref). Then the choice of distribution (C) cannot be arbitrary, because the (internal) dependence structure of $(Y,Z)$ depends on the dependence structures of $(X,Y)$ and $(X,Z)$. Rachev et al. (2007, p. 93) \nocite{Rachev2007} give the following consistency rule: ”The three pairs of random variables $(X,Z)$, $(X,Y)$ and $(Y,Z)$ should be chosen in such a way that there exists a consistent three-dimensional random vector $(X,Y,Z)$ and the three pairs are its two-dimensional projections.” In other words, if this rule is respected, then the three pairs can be safely embedded in a three-dimensional random vector. To validate our counterexample, we must make sure that the distributions (A), (B), \textbf{(C)} abide by the consistency rule. This is indeed the case, for the matrix

equation[equation omitted — 159 chars of source]

is positive definite (with eigenvalues $\lambda_1 = 0.072$, $\lambda_2 = 0.44$ and $\lambda_3 = 2.128$), which means that $V$ is a valid covariance matrix.

An illegitimate counterexample

For completeness, we give what in appearance only is a counterexample. The distributions (D), (E), (F) in Table (ref) are such that the triangle inequality does not hold because \newline $(E|X-Y| + E|Y-Z| - E|X-Z|)/2 = -0.2$. However, (D), (E), (F) do not constitute a valid counterexample, because the matrix $V$ in ((ref)) resulting from these distributions is indefinite (with eigenvalues $\lambda_1 = -0.076$, $\lambda_2 = 1.093$ and $\lambda_3 = 1.623$), and thus simply cannot be a covariance matrix of a three-dimensional vector.

We end this subsection with an example illustrating the importance of the independence assumption in Theorem (ref). In Table (ref), $(X,Z)$, $(X,Y)$ and $(Y,Z)$ are now three pairs of independent variables with the same marginal distributions as those shown in Table (ref).

table[table omitted — 1,671 chars of source]

From the distributions (G), (H), (I), we find that $(E|X-Y| + E|Y-Z| - E|X-Z|)/2 = 0.5$ (instead of -0.02 in ((ref))), which means that the triangle inequality holds. Note also that $V$ is the three-dimensional diagonal matrix diag(0.96,0.84,0.84), which is of course a valid covariance matrix.

Link with the Gini index

The computation of a distance between identically distributed random variables of ${\cal L}_1(I\!\!R)$ is helpful in various domains. Examples are the Gini mean difference (GMD)\footnote{The GMD is twice the $L$-scale (the second $L$-moment). It is sometimes considered as a competitor of the standard deviation. } and the Gini index, used in particular in inequality economics to measure the amount of inequality included in a distribution of income (alternatively consumption or wealth, etc.). Let $\mu$ be such a distribution. The GMD and the Gini index are defined as GMD$(\mu) = E|X-Y|$ ($= D(X,Y))$ and $Gini(\mu) = E|X-Y|/[2E(X)]$ ($= D_{norm}(X,Y)$), where $X \sim \mu$ and $Y \sim \mu$ are assumed to be independent and (usually) non-negative (see e.g. Yitzhaki (1998)\nocite{Yitzhaki1998}, Yitzhaki and Schechtman (2013)\nocite{YitzhakiSchechtman2013}, or Xu (2003)\nocite{Xu2003})\footnote{When the range of $X$ encroaches on $(-\infty ,0]$, we know from ((ref)) that there is no mathematical reason why we should not define $Gini(\mu) = E|X-Y|/[2E|X|]$.}. Looking at ((ref)), we can say that the Gini index is the distance (semimetric) $D_{norm}$ from $X$ to an i.i.d. “copy” of itself. In that sense, it is sometimes called an “autodistance” in the literature. However, a copy must be clearly distinguished from the original and there is some confusion on this point. “Copy” is to be understood here in the equality of distribution sense (in the almost sure equality sense, the distance is trivially zero).

Independence of $X \hbox{ and } Y$ allows in many cases to use Fubini-Tonelli to represent the GMD and the Gini index in closed form. Independence also implies that, except in the degenerate case where $X\stackrel{a.s.}{=} Y\stackrel{a.s.}{=}c$ for some $c \in I\!\!R_+$, $GMD(\mu) > 0$. Although seemingly simple, the Gini index is actually a quite proteiform measure of inequality. It can be expressed in an astonishing number of ways, some of which can be found in Yitzhaky (1998)\nocite{Yitzhaki1998}.

We end this subsection by probabilistic considerations on the values the Gini index can take. Let $X\sim \mu$ and $Y\sim \mu$ be two i.i.d. non-negative random variables where $\mu$ is (say) the income distribution of a population. Then $Gini(\mu) = E|X-Y|/[2E(X)]$ is not defined if and only if $X\stackrel{a.s.}{=}0$. Leaving this uninteresting case aside, $Gini(\mu)= 0 \Leftrightarrow E|X-Y| = 0 \stackrel{Prop.\ref{equivalence}}{\Longleftrightarrow} \exists c > 0$ such that $X\stackrel{a.s.}{=} Y\stackrel{a.s.}{=}c$. Moreover, the triangle inequality $E|X-Y| \leq |X| + |Y|$ implies that $Gini(\mu) \in [0,1]$.

In economic applications, non-negative real numbers $x_1,x_2,\ldots ,x_n$ are typically incomes earned respectively by individuals $i_1,i_2,\ldots , i_n$ belonging to a population or a statistical sample. It is known that $0\le g \leq (n-1)/n$, where $g$ is the Gini index. A low value of $g$ corresponds to a more equal income distribution, with 0 indicating perfect equality (all individuals have the same income). The most unequal society is the one where a single individual receives all the income and the remaining individuals receive nothing. In that case, $g = (n-1)/n$. Generalizing a bit, we can say that a tiny proportion $\epsilon > 0$ of the population receives an income $b>0$, while a proportion $1 - \epsilon$ gets nothing. Such an extreme case can be described thanks to i.i.d. random variables $X \hbox{ and } Y$ having distribution $\mu = (1 - \epsilon) \delta_0 + \epsilon \delta_b$ where $b$ is a strictly positive income level and where $\epsilon > 0$, $\epsilon$ small, is the probability that $X$ (or $Y$) takes the value $b$. The distribution of the pair $(X,Y)$ is summarized in Table (ref).

table[table omitted — 817 chars of source]

The Gini index then becomes $Gini(\mu) = E|X-Y|/[2E(X)] = 1 - \epsilon$. The upper bound 1 corresponding to $\epsilon = 0$ cannot be reached, because in such a case $X$ would take the value 0 with probability 1 , and the Gini index would not be defined ($E(X)$ would be zero).

Alternative interpretation of the ${\cal L}_p$-distance

For $p \ge 1$, let ${\cal L}_p (I\!\!R)$ denote the class of all real-valued random variables on a probability space $(\Omega,{\cal A},P)$ that have finite $p$-th moment, and consider $X,Y\in {\cal L}_p (I\!\!R)$. Suppose that $X \hbox{ and } Y$ are identically distributed. In this section, we show in particular that the ${\cal L}_p$-distance $(E|X -Y|^p)^{1/p}$ certainly represents a distance between $X \hbox{ and } Y$, but also -- in a sense to be specified -- tells us how far from almost sure equality these variables are. In a symbolic way, we will show that $(E|X -Y|^p)^{1/p}$ can be conceived as a distance $Dist(X \stackrel{d}{=} Y,X \stackrel{a.s.}{=} Y)$ between two possible states of the pair $(X,Y)$. In (ref), we introduce the so-called diagonal coupling of a probability measure with itself, and we clarify its connection to almost sure equality. In (ref), we define a distance\footnote{in fact a semimetric, i.e. a distance satisfying the triangle inequality.} $\eta_p(\pi_1, \pi_2)$ between bidimensional probability measures $\pi_1$ and $\pi_2$ having finite $p$-th moment. In (ref), we recall the notion of coupling and observe what becomes of $\eta_p$ when $\pi_1$ and $\pi_2$ are couplings of unidimensional probability measures $\mu$ and $\nu$ of finite $p$-th moment. Equating $\mu$ and $\nu$ and taking $\pi_2$ as the diagonal coupling of $\mu$ with itself allows then to interpret $\eta_p$ as a distance indicating how far $X\sim \mu$ and $Y\sim \mu$ are from almost sure equality. The results are illustrated with the bivariate normal distribution. Finally, thanks to $\eta_p$, we propose a probabilistic representation of the Gini mean difference and of the Gini index.

Diagonal coupling and almost sure equality

Consider a probability space $(I\!\!R, {\cal B}_1,\mu)$. We denote by $\mu \triangle\mu$ a probability measure on the product space $(I\!\!R^2, {\cal B}_2)$ defined by

equation[equation omitted — 129 chars of source]

Then $\mu \triangle\mu$ can be extended to the whole ${\cal B}_2$ by using Carath\'eodory's theorem (Kalikow and McCutchean (2010) \nocite{Kalikow2010}).

defi$\mu \triangle\mu$, as defined above, is called the diagonal coupling of $\mu$ with itself.

Actually, $(I\!\!R, {\cal B}_1,\mu )$ and $(I\!\!R^2, {\cal B}_2,\mu \triangle \mu )$ are closely related: the map $s:I\!\!R \longrightarrow I\!\!R^2$ defined by $s(x) = (x,x)$ is a measurable isomorphism from $(I\!\!R, {\cal B}_1,\mu )$ to $(I\!\!R^2, {\cal B}_2,\mu \triangle\mu )$. An example of diagonal coupling is given in Table (ref) (A), where $\mu$ is defined by $\mu (0) = 0.3$ and $\mu (1) = 0.7$). Incidentally, Proposition (ref) tells us that if $X\stackrel{a.s.}{=} Y$, with $X \hbox{ and } Y$ independent, then there exists $c \in I\!\!R$ such that $X\stackrel{a.s.}{=} c$ and $Y\stackrel{a.s.}{=} c$, which implies that $\mu \otimes \mu = \mu \triangle \mu = \delta_{(c,c)}$ (the Dirac delta measure concentrated on $(c,c)$ ), where $\mu \otimes \mu$ is the product measure of $\mu$ with itself. So the probability measures $\mu \otimes \mu$ and $\mu \triangle \mu$ on $(I\!\!R^2,{\cal B}_2)$ coincide in this trivial case.

figure[figure omitted — 299 chars of source]

Let $\Delta$ denote the diagonal $\{(x,x):x\inI\!\!R\}$ of $I\!\!R^2$. One can show quite easily that $B_1\cap B_2 = s^{-1}((B_1 \times B_2)\cap \Delta)$ for any subsets $B_1,B_2$ of $I\!\!R$. Therefore, the definition of the diagonal coupling in ((ref)) may be replaced by $(\mu \triangle\mu)(B_1\times B_2) = \mu(s^{-1}((B_1 \times B_2)\cap \Delta))$, which is visually more telling (see Figure (ref)). Let $X,Y:(\Omega,{\cal A},P) \rightarrow (I\!\!R, {\cal B}_1)$ (e.g. $X,Y \in {\cal L}_p (I\!\!R)$ if $X \hbox{ and } Y$ have finite $p$-th moment) be two identically distributed random variables. Then the diagonal coupling of $P_X$ with itself is the distribution of the pair $(X,Y)$ if and only if $X\stackrel{a.s.}{=} Y$:\footnote{By definition $X\stackrel{a.s.}{=} Y$ makes sense only if $X \hbox{ and } Y$ are defined on the same probability space. } the entire probability mass of $P_{(X,Y)}$ is concentrated on the diagonal $\Delta$. This result is in line with our intuition. It is formally stated in Proposition (ref).

prop(Proof in the appendix) Let $X \hbox{ and } Y$ be two identically distributed random variables defined on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R, {\cal B}_1)$. Then $X\stackrel{a.s.}{=} Y$ if and only if $P_{(X,Y)} = P_X \triangle P_X$.

Lemma (ref) below -- added for completeness -- is not directly related to what we need in this article. However, we would like to answer the following question: how can we construct a probability space on $\Delta$ consistent with $(I\!\!R^2, {\cal B}_2, \mu \triangle \mu)$? Two ways come to mind: (i) take the trace space of $(I\!\!R^2, {\cal B}_2, \mu \triangle \mu)$ with respect to $\Delta$, or (ii) take the pushforward space of $(I\!\!R, {\cal B}_1, \mu )$ induced by $s$ ($s$ defined above). It turns out that the two methods produce the same space, as evidenced by Lemma (ref).

lemme(Proof in the appendix) Consider the trace probability space $(\Delta, {\cal B}_2 \cap \Delta, (\mu \triangle\mu)\left\lfloor_\Delta \right.)$ of $(I\!\!R^2, {\cal B}_2, \mu \triangle \mu)$, where $(\mu \triangle\mu)\left\lfloor_\Delta \right.$ denotes the probability measure $\mu \triangle\mu$ restricted to $\Delta$, and the pushforward probability space $(\Delta, s({\cal B}_1), \mu_s)$ of $(I\!\!R, {\cal B}_1, \mu )$ induced by $s:I\!\!R \longrightarrow I\!\!R^2$ defined by $s(x) = (x,x)$. Then the trace space and the pushforward space coincide, i.e. \newline (a) ${\cal B}_2 \cap \Delta = s({\cal B}_1)$ (equality of $\sigma$-fields) \newline (b) $(\mu \triangle\mu)\left\lfloor_\Delta \right. = \mu_s$ (equality of measures)\footnote{Incidentally, $\mu_s$ is a so-called deterministic transport plan in the Monge-Kantorovich transport problem. Denote by $Id$ the identity map on $I\!\!R$. Let $\pi_T := \mu_{(Id,T)}$ be the pushforward probability measure of $\mu$ induced by the function $(Id,T):I\!\!R \rightarrow I\!\!R^2$, where $T$ is a transport map. $\pi_T$ is called a deterministic transport plan. In reference to the optimal transport problem, we have here $\mu_s = \mu_{(Id,T)}$ where $T = Id$, i.e. $\mu_s = \mu_{(Id,Id)}$. }.

A primary metric defined on distributions of pairs of random variables

Let $(X_1,Y_1)$ (resp. $(X_2,Y_2)$) be two pairs of random variables having joint distribution $\pi_1$ (resp. $\pi_2$) on $(I\!\!R^2, {\cal B}_2)$. What is meant by $\pi_1$ and $\pi_2$ being “close to each other”? A possible answer -- serving what we wish to show in this subsection -- is to measure their proximity by using a primary metric, i.e. to consider that $\pi_1$ and $\pi_2$ coincide when they share a given set of relevant characteristics. Accordingly, we consider here that the distance between $\pi_1$ and $\pi_2$ is zero if (i) the centers (mathematical expectations, means) of $(X_1,Y_1)$ and $(X_2,Y_2)$ are the same and (ii) the deviation between $X_1$ and $Y_1$ is the same as the deviation between $X_2$ and $Y_2$.

For $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2)\sim \pi_2$, assume that the marginals of $\pi_1$ and $\pi_2$ have finite $p$-th moment, $p\in [1,\infty )$. For $i=1,2$, we will use the following notations: \newline $E|X_i -Y_i|^p = \int_{I\!\!R^2} |x_i -y_i|^p d\pi_i(x_i,y_i)$, $EX_i = \int_{I\!\!R^2} x_i d\pi_i(x_i,y_i)$, $EY_i = \int_{I\!\!R^2} y_i d\pi_i(x_i,y_i)$ and $C(\pi_i) = E([X_i, Y_i]) = [EX_i, EY_i]$. We can now define

equation[equation omitted — 147 chars of source]

where $||\cdot||$ is the Euclidean norm on $I\!\!R^2$. So $\eta_p$, which takes finite values, integrates two sources of deviation: between pairs and within pairs of random variables. Obviously, the characteristics entering the definition of $\eta_p$ in ((ref)) do not fully describe what differentiates $\pi_1$ from $\pi_2$.

propLet ${\cal P}(I\!\!R^2)$ denote the set of probability measures on $(I\!\!R^2,{\cal B}_2)$. For $p\in [1,\infty )$, let ${\cal P}_p(I\!\!R^2)$ denote the set of probability measures on the Borel subsets of $I\!\!R^2$ whose marginals have finite moment of order $p$, i.e. \newline ${\cal P}_p(I\!\!R^2) = \{\pi \in {\cal P}(I\!\!R^2): \int_{I\!\!R^2} |x|^p d\pi (x,y)< \infty \hbox{ and } \int_{I\!\!R^2} |y|^p d\pi (x,y)< \infty \}$. Then $\eta_p : {\cal P}_p(I\!\!R^2) \times {\cal P}_p(I\!\!R^2) \rightarrow I\!\!R_+$ defined in {\rm ((ref))} is a semimetric, i.e. $\eta_p$ satisfies non-negativity, reflexivity ($\eta_p(\pi, \pi) = 0$), symmetry ($\eta_p(\pi_1, \pi_2) = \eta_p(\pi_2, \pi_1)$) and triangle inequality \newline ($\eta_p(\pi_1, \pi_3) \le \eta_p(\pi_1, \pi_2) + \eta_p(\pi_2, \pi_3)$ for all $\pi_1,\pi_2,\pi_3 \in {\cal P}_p(I\!\!R^2)$.

The (easy) proof of Proposition (ref) is omitted. As $\eta_p$ is in particular non-negative, reflexive and symmetric, it can be called a distance between elements of ${\cal P}_p(I\!\!R^2)$.\footnote{Formally, as already mentioned, the difference between a semimetric and a distance is the relaxation of the triangle inequality.} Moreover, $\eta_p$ is a semimetric without being a metric\footnote{As a semimetric, $\eta_p$ can be transformed into a metric between equivalence classes. Define an equivalence relation between the elements of ${\cal P}_p(I\!\!R^2)$ by $\pi_1 \sim \pi_2 :\Leftrightarrow \eta_p(\pi_1, \pi_2) = 0$. Then $\tilde{\eta_p}([\pi_1], [\pi_2]) := \eta_p(\pi_1, \pi_2)$ is a metric on the set $\left\{ [\pi] : \pi \in {\cal P}_p(I\!\!R^2) \right\}$ of classes. One can show that equivalent elements are equidistant from any other element of ${\cal P}_p(I\!\!R^2)$. }. Indeed, $\eta_p$ does not satisfy the reverse reflexivity condition $\eta_p(\pi_1, \pi_2) = 0 \Rightarrow \pi_1 = \pi_2$. Example (ref) shows us that there exists probability measures $\pi_1,\pi_2 \in {\cal P}_p(I\!\!R^2)$ such that $\pi_1 \neq \pi_2$ while $\eta_p(\pi_1, \pi_2) =0$.

exampleLet $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2)\sim \pi_2$, where $\pi_1, \pi_2$ are described in Table (ref). Clearly, $\pi_1, \pi_2 \in {\cal P}_p(I\!\!R^2)$ for any $p\ge 1$. Note that $X_1$, $Y_1$, $X_2$ and $Y_2$ all have the same distribution $\mu$ given by $\mu (0) = .20$, $\mu (1) = .35$ and $\mu (2) = .45$. Since $\eta_p(\pi_1, \pi_2) = 0$, $\pi_1$ and $\pi_2$ are in the same equivalence class when classes are defined with respect to the equivalence relation $\pi_1 \sim \pi_2 : \Leftrightarrow \eta_p(\pi_1, \pi_2) = 0$.
center[center omitted — 1,103 chars of source]

The special case of couplings

Couplings between distributions and between random variables

We now need the general definition of coupling, of which the diagonal coupling (Definition (ref)) is a special case.

defi(coupling of probability measures on the real line) A coupling of two given probability measures $\mu \hbox{ and } \nu$ on $(I\!\!R, {\cal B}_1)$ is any probability measure $\pi$ on $(I\!\!R^2, {\cal B}_2)$ whose marginals are $\mu \hbox{ and } \nu$, that is, $\mu =\pi \circ q_1^{-1}$ and $\nu =\pi \circ q_2^{-1}$, where the $q_i$'s are the projection functions defined by $q_i(x_1,x_2) = x_i$ for all $(x_1,x_2)\in I\!\!R^2$, $i=1,2$.

By definition, couplings are multiple. The class of all couplings between $\mu \hbox{ and } \nu$ is denoted by $\Pi (\mu ,\nu)$. For example, the distributions $\pi_1$ and $\pi_2$ shown in Table (ref) belong to the set $\Pi (\mu,\mu)$, i.e. $\nu = \mu$ in this case.

defi(coupling of real-valued random variables) A coupling of two given random variables $X \hbox{ and } Y$ taking values in $(I\!\!R, {\cal B}_1)$ is any pair of random variables $(\tilde{X},\tilde{Y})$ taking values in $(I\!\!R^2, {\cal B}_2)$ such that $\tilde{X}$ and $\tilde{Y}$ are defined on the same probability space $(\tilde{\Omega},\tilde{\cal A},\tilde{P})$, with $\tilde{X}\stackrel{d}{=} X$ and $\tilde{Y}\stackrel{d}{=} Y$.

We observe that the law $\tilde{\pi}$ of $(\tilde{X},\tilde{Y})$ is a coupling of the laws $\mu$ of $X$ and $\nu$ of $Y$. An important point of Definition (ref) is that the coupled random variables are defined on the same probability space, while $X \hbox{ and } Y$ may not be defined on a common probability space. If $X \hbox{ and } Y$ are defined on the same probability space $(\Omega,{\cal A},P)$, then $(X,Y)$ is also defined on $(\Omega,{\cal A},P)$ and $P_{(X,Y)}$ is a coupling of $P_X$ and $P_Y$. Two trivial couplings are (i) the diagonal coupling $\mu \triangle\mu$ of $\mu$ with itself defined in ((ref)), and (ii) the product coupling $\mu \otimes \nu$. If $X \sim \mu$ and $Y \sim \nu$ are independent, then the law of $(X,Y)$ is the product probability measure $\mu \otimes \nu$. Table (ref) (D) and (E) are examples of product couplings.

Distance between couplings, $\mu \hbox{ and } \nu$ arbitrary

Let $\Pi_p (\mu,\nu)$ denote the set of couplings of two probability measures $\mu \hbox{ and } \nu$ of finite $p$-th moment, both defined on $(I\!\!R, {\cal B}_1)$. For $\pi_1,\pi_2 \in \Pi_p (\mu,\nu)$, suppose that $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2)\sim \pi_2$. As couplings of $\mu$ and $\nu$ have identical centers, $||C(\pi_1) - C(\pi_2)|| $ disappears in ((ref)) and we have

equation[equation omitted — 115 chars of source]
example(Discrete case) Equation ((ref)) has a particularly simple form when $p=1$ and when $\mu$ and $\nu$ are discrete. Suppose that $X_1$ and $X_2$ take values in $\{ s_i:i=1,2,\ldots ,n_X \}$, and that $Y_1$ and $Y_2$ take values in $\{ t_j:j=1,2,\ldots ,n_Y \}$. Then ((ref)) becomes \begin{equation} \eta_1(\pi_1, \pi_2) = \left| \sum_{i=1}^{n_X} \sum_{j=1}^{n_Y} |s_i - t_j| (p_{ij} - q_{ij}) \right| , \end{equation} where $p_{ij} = \pi_1(\{ (s_i,t_j) \})$ and $q_{ij} = \pi_2(\{ (s_i,t_j) \})$. Equation ((ref)) indicates that $\pi_1 = \pi_2$ implies $\eta_p(\pi_1, \pi_2) = 0$, but we know that the converse is not true (see the counterexample in Table (ref)).

Distance between couplings when $\nu = \mu$ and $\pi_2 = \mu \triangle \mu$

Let us now consider the case where $\nu = \mu$, when $\mu \hbox{ and } \nu$ have finite $p$-th moment. Assume that $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2) \sim \mu \triangle \mu$. Note that $\mu \triangle \mu \in \Pi_p (\mu,\mu)$, and assume that $\pi_1 \in \Pi_p (\mu,\mu)$. For ease of reading, write $(X,Y) \sim \pi$ instead of $(X_1,Y_1)\sim \pi_1$. Clearly, $X,Y,X_2$ and $Y_2$ have all the same distribution $\mu$. We wish to measure the distance $\eta_p$ between $\pi$ and $\mu \triangle \mu$. As $X_2\stackrel{a.s.}{=} Y_2$, we have $E|X_2-Y_2|^p=0$ and ((ref)) becomes

equation[equation omitted — 105 chars of source]

That is, the ${\cal L}_p$-distance $(E|X -Y|^p)^{1/p}$ between identically distributed $X \hbox{ and } Y$ also represents a distance between the law $\pi$ of $(X,Y)$ and the diagonal coupling built from the marginals of $\pi$. Note that both the ${\cal L}_p$-distance and $\eta_p$ are semimetrics. \vskip .2cm Distance between equality of distribution and almost sure equality \vskip .2cm Suppose that $X$, $Y$ (and therefore $(X,Y)$) are defined on some probability space $(\Omega,{\cal A},P)$. Moreover, suppose that $X \hbox{ and } Y$ are identically distributed. The following question comes to mind: how far from almost sure equality are $X \hbox{ and } Y$? Writing $\pi = P_{(X,Y)}$ and $\mu = P_X$ in ((ref)), we obtain

equation[equation omitted — 96 chars of source]

That is, when jointly distributed random variables $X$ and $Y$ are identically distributed, $||X-Y||_p$ is actually a distance between the distribution of the pair $(X,Y)$ and the diagonal coupling of $P_X$ with itself. It is in this sense that $||X-Y||_p$ may be symbolically interpreted as a distance between $X\stackrel{d}{=} Y$ and $X\stackrel{a.s.}{=} Y$. \vskip .2cm Illustration of the above developments: the case of bivariate normal distributions \vskip .2cm Consider the bivariate normal density function

equation[equation omitted — 100 chars of source]
equation[equation omitted — 238 chars of source]

with parameters $m_X$, $m_Y$ (marginal means), $\sigma_X > 0$, $\sigma_Y > 0$ (marginal standard deviations) and $|\rho| < 1$ (correlation coefficient). Let $(X,Y) \sim \pi$, where $\pi$ is characterized by the density function ((ref)). The marginals $\mu \hbox{ and } \nu$ of $\pi$ are $\mu = N(m_X,\sigma^2_X)$ and $\nu = N(m_Y,\sigma^2_Y)$. Importantly, they do not depend on $\rho$, which means that each bivariate normal distribution $\pi$ is a coupling of $\mu \hbox{ and } \nu$. An infinite number of couplings can be created by just changing the value of $\rho$.

When $m_X = m_Y$ and $\sigma_X = \sigma_Y$, i.e. when $X \hbox{ and } Y$ are identically distributed, then $\mu \triangle \mu$ is the diagonal coupling of $N(m_X,\sigma^2_X)$ with itself, which is supported on the diagonal $\Delta$ of $I\!\!R^2$. Setting for example $p=1$ in ((ref)), we obtain

equation[equation omitted — 110 chars of source]

from which we draw the conclusion that the smaller the value of $E|X-Y|$, the more the graph of $f_{XY}(x,y)$ is concentrated along the diagonal $\Delta$.

figure[figure omitted — 799 chars of source]

Figure (ref) graphs the densities of three bivariate normal distributions, all three having parameters $m_X = m_Y =0$ and $\sigma_X = \sigma_Y =1$. Consequently, the two marginal laws of each of these distributions are the standard normal $N(0,1)$. The only difference between the three bivariate distributions of Figure (ref) is the value of $\rho$. The distribution on the left, where $\rho = 0$, corresponds to i.i.d. standard normal random variables $X \hbox{ and } Y$. Looking at the distribution on the right, where $\rho = 0.99$, we can guess the bell-shaped form the distribution will have when $\rho = 1$. In this case $\mu \triangle \mu = N(0,1) \triangle N(0,1)$, defined on $I\!\!R^2$, concentrates its probability mass on $\Delta$.

It is interesting to visualize towards which distribution the densities of Figure (ref) converge when $\rho\nearrow 1$ and $\Delta$ is seen as a standalone line of numbers. In this case the manifold $\Delta$ differs from $I\!\!R$ only by its coordinate system: $I\!\!R$ is “stretched” to obtain $\Delta$ and $\Delta$ can be seen as $I\!\!R$ with a new coordinate system given by $y=y(x)=\sqrt{2}x$. Let $f$ (resp. $g$) be the pdf representing $\mu = N(0,1)$ in the old (resp. new) coordinate system. Then, for any $B\in {\cal B}_1$, $\int_B f(x)dx = \int_B g(y)dy$, where, using the Jacobian rule for a change of coordinates, $g(y) = |\frac{dx}{dy}| f(x) = \frac{1}{\sqrt{2}} f (\frac{y}{\sqrt{2}} )$, which is the density of the distribution $N(0,2)$. That is, we have shown that the bivariate normal density represented on the right of Figure (ref) is close to the density $N(0,2)$ defined on $\Delta$ when $\Delta$ is considered as a simple line of numbers.

Effect of independence

If, in addition to being identically distributed, two jointly distributed random variables $X \hbox{ and } Y$ are assumed to be independent (i.e. if $X \hbox{ and } Y$ are i.i.d.), then $P_{(X,Y)} = P_X \otimes P_Y = P_X \otimes P_X$. Moreover, Proposition (ref) implies that, for $c:=E(X)$, $P_X \triangle P_X = \delta_{(c,c)}$, the Dirac delta measure concentrated on $(c,c)\in \Delta$. In that case, ((ref)) becomes

equation[equation omitted — 108 chars of source]

i.e. $(E|X -Y|^p)^{1/p}$, whose value depends only on $P_X$, is a measure of the distance between the product coupling of $P_X$ with itself and the measure concentrated on a point of the diagonal of $I\!\!R^2$. Note that $\delta_{(E(X),E(X))}$ is not in $\Pi (P_X,P_X)$. However, since $\pi_1 := P_X \otimes P_X$ and $\pi_2 := \delta_{(E(X),E(X))}$ have the same center, $||C(\pi_1) - C(\pi_2)|| = 0$ in ((ref)).

The Gini mean difference as a distance between measures

Consider two random variables $X,Y \in {\cal L}_1 (I\!\!R)$. As we have seen in Subsection (ref), the Gini mean difference (GMD) of an income distribution $\mu$ is defined as $E|X-Y|$, where $X\sim \mu$ and $Y\sim \mu$ are non-negative and independent. We are now able to give a probabilistic definition of the GMD, perhaps the most general definition that can be given to this index. Looking at ((ref)), and setting $p=1$, we obtain

equation[equation omitted — 98 chars of source]

In other words, the GMD of an income (or a wealth, etc.) distribution $\mu$ represents a distance between the product coupling $\mu \otimes \mu$ of $\mu$ with itself and the Dirac delta measure supported on the single point $(E(X),E(X))$. Consequently, GMD$(\mu) = E|X-Y|$ measures how far the i.i.d. $X\sim \mu$ and $Y\sim \mu$ are from almost sure equality. Note that $\mu \otimes \mu$ distributes its mass symmetrically with respect to $\Delta$. The GMD thus measures a distance between the distribution $\mu \otimes \mu$ on $I\!\!R^2$ (or on a set of $A\times A$, $A\subset I\!\!R$) and the measure concentrated on the center of this same distribution.

Expected absolute difference for independent random variables: applications to physics and to economics

Many examples can be given of the usefulness of $E|X-Y|$ for independent variables, and they range from physics to economics. In physics, for instance, Lukaszyk (2004)\nocite{Lukaszyk2004} presents a modified Shephard-Liszka\nocite{Liszka1984,Shephard1967} approximation, where $E|X-Y|$ proves to be more reliable than a plain Euclidean metric, suggesting that an analogous improvement can be achieved in various numerical methods, and in particular in approximation algorithms. In the same paper, Lukaszyk suggests further applications in fringe pattern analysis, or in quantum mechanics, to estimate the distance of two quantum particles described by their wave functions. As pointed out in paragraph (ref), $E|X-Y|$ has also important applications in inequality economics. Other applications include clustering systems, pattern recognition or finance.

In Subsection (ref), we indicate the forms in which $E|X-Y|$ is usually represented in scientific publications, especially of physics, engineering, finance and economics. In (ref), we express $E|X-Y|$ in analytic form when $X \hbox{ and } Y$ are two independent normally distributed random variables. This new result generalizes specific formulas already present in the literature, in particular that of physics. In (ref), we give closed-form formulas for the average distance and the normalized average distance between coordinates of points falling at random into a rectangle of the plane. In (ref), we take advantage of results of (ref) to express in closed form the mean value of the set $\{|a-b|: a \in A, b \in B \}$, where $A$ and $B$ are bounded intervals of real numbers.

Formulations in use in applied fields

For independent $X,Y \in {\cal L}_1(I\!\!R )$, let $P_{(X,Y)} = \mu \otimes \nu$ be the product distribution of two probability distributions $\mu \hbox{ and } \nu$ defined on $(I\!\!R,{\cal B}_1)$. The Fubini-Tonelli theorem applies, and we can interchange the order of integration or summation so that

eqnarray[eqnarray omitted — 289 chars of source]

Thereafter, $F$ (resp. $G$) will refer to the cumulative distribution function (cdf) of $X$ (resp. $Y$), and $f$ (resp. $g$) will refer to the probability density function (pdf) of $X$ (resp. $Y$). Particularly interesting are the cases where independent $X$ and $Y$ are (i) both discrete, (ii) both absolutely continuous, or (iii) one of them is discrete and the other absolutely continuous.

(i) The (independent) discrete random variables $X$ (resp. $Y$), take values in the countable sets $\Omega_X =\{ x_1,x_2,\ldots,x_i,\ldots \} \subset I\!\!R$ (resp. $\Omega_Y =\{ y_1,y_2,\ldots,y_j,\ldots \} \subset I\!\!R$). In that case, $f$ (resp. $g$) are probability mass functions\footnote{i.e. have densities with respect to the counting measure on $\Omega_X$ (resp. $\Omega_Y$). }. In this context, ((ref)) can be rewritten as $E|X-Y|=\sum_i\sum_j\left|x_{i}-y_{j}\right|p_{i}q_{j}$, where $p_i = P(X = x_i)$ and $q_j = P(Y = y_j)$\footnote{This double sum is just an integral with respect to the counting measure on $\Omega_X \times \Omega_Y$.}. In the finite equiprobable case, when independent random variables $X$ and $Y$ both take the non-negative values $x_1,x_2,\ldots , x_N$ with respective probabilities $p_1=p_2=\cdots = p_N=1/N$, then $$ E|X-Y|= \frac{1}{N^2}\sum_{i=1}^N\sum_{j=1}^N\left|x_{i}-x_{j}\right| $$ is the Gini mean difference in its discrete form, see for example Gini (1912)\nocite{Gini1912}, Kendall and Stuart (1958)\nocite{Kendall1958}, or Xu (2003)\nocite{Xu2003}.

(ii) The independent $X \sim \mu$ and $Y \sim \nu$ are both absolutely continuous\footnote{i.e. absolutely continuous with respect to the Lebesgue measure.}. Equation ((ref)) usually appears in the following form in the literature (notably in physics and economics) $$ E|X-Y|= \int^{\infty}_{x=-\infty} \int^{\infty}_{y=-\infty}|x-y|f(x)g(y)dxdy. $$ When $X,Y \in {\cal L}_1(I\!\!R )$ are i.i.d. non-negative absolutely continuous random variables ($f=g$), $$ E|X-Y|= \int^{\infty}_{x=-\infty} \int^{\infty}_{y=-\infty}|x-y|f(x)f(y)dxdy. $$ is known in the economic literature as the continuous Gini mean difference (see, for example, Yitzhaki (1998)\nocite{Yitzhaki1998}, or Yitzhaki and Schechtman (2013)\nocite{YitzhakiSchechtman2013}).

(iii) $X$ is discrete, whereas $Y$, independent of $X$, is absolutely continuous. Then $$ E|X-Y|= \sum_i p_i \int^{\infty}_{y=-\infty}|x_i-y|g(y)dy. $$ A special case of this formula was used by Lukaszyk (2004)\nocite{Lukaszyk2004}, when he proposed a modified Liszka method to handle an experimental mechanics issue.

Analytic form of the expected absolute difference between two independent normally distributed random variables

The normal case plays a crucial part in a great many of the techniques used in applied statistics. The Central-limit theorem alone ensures that this will be the case, but there are other important reasons extensively discussed in the literature.

We begin this section by writing $E|X-Y|$ in a form facilitating the calculation of its analytic expression when independent $X \hbox{ and } Y$ are both absolutely continuous.

prop(Proof in the appendix) Let $X,Y \in {\cal L}_1(I\!\!R )$ be two independent absolutely continuous random variables with means $\mu_X = E(X)$ (resp. $\mu_Y = E(Y)$) and cdf's $F$ (resp. $G$). Then \begin{equation} E|X-Y| = 2 \left\{E[XG(X)] + E[YF(Y)] \right\} - \mu_X - \mu_Y . \end{equation}

Consider the special case where $X \hbox{ and } Y$ in Proposition (ref) are i.i.d. Then $F=G$, $\mu_X = \mu_Y$, and ((ref)) becomes

equation[equation omitted — 57 chars of source]

Moreover, since $X$ absolutely continuous $\Rightarrow$ $F$ continuous $\Rightarrow$ $F(X) \sim U(0,1)$ $\Rightarrow$ $E[F(X)] = 1/2$, we have: $\hbox{\rm cov}[X,F(X)] = E[XF(X)] - \mu_X/2 \stackrel{(\ref{XFX})}{=} E|X-Y|/4$, i.e.

equation[equation omitted — 66 chars of source]

This result -- of which ((ref)) is a generalization -- can be found in Lerman and Yitzhaki (1984)\nocite{Lerman1984}.

In the following (new) theorem (Theorem (ref)), we give the analytic form of $E|X-Y|$ for normally distributed independent random variables.

theor(Two alternative proofs can be found in the appendix) Assume that $X\sim N(\mu_X,\sigma^2_X)$ and $Y\sim N(\mu_Y,\sigma^2_Y)$ are independent normally distributed random variables. Let $\phi$ (resp. $\Phi$) be the pdf (resp. the cdf) of the standard normal distribution. Then the expected absolute difference between $X$ and $Y$ is given by \begin{eqnarray} E|X-Y| &=& \frac{2\sigma_X^2}{\sqrt{\sigma_X^2+\sigma_Y^2}} \phi \left( \frac{|\mu_X-\mu_Y|}{\sigma_Y} \right) \exp \left( \frac{\sigma_X^2 (\mu_X-\mu_Y)^2 }{2\sigma_Y^2 (\sigma_X^2+\sigma_Y^2)} \right) \nonumber \\ &+& \frac{2\sigma_Y^2}{\sqrt{\sigma_X^2+\sigma_Y^2}} \phi \left( \frac{|\mu_X-\mu_Y|}{\sigma_X} \right) \exp \left( \frac{\sigma_Y^2 (\mu_X-\mu_Y)^2 }{2\sigma_X^2 (\sigma_X^2+\sigma_Y^2)} \right) \nonumber \\ &+& 2|\mu_X-\mu_Y| \Phi \left( \frac{|\mu_X-\mu_Y|}{\sqrt{\sigma_X^2+\sigma_Y^2}} \right) - |\mu_X-\mu_Y| . \\ \nonumber \end{eqnarray}

Equation ((ref)) can also be written

equation[equation omitted — 256 chars of source]

Note that ((ref)) is the formula we end up with if we use Proposition (ref). To obtain ((ref)), we used the fact that the convolution of two Gaussian distributions is a Gaussian distribution. The proof of ((ref)) is shorter than that of ((ref)) resulting from Proposition (ref). However, Proposition (ref) applies to absolutely continuous variables and we wanted to test it in the particular case where the variables are Gaussian. It is left to the reader to show that ((ref)) and ((ref)) are equivalent.

The rest of Subsection (ref) is devoted to corollaries of Theorem (ref). Equation ((ref)) generalizes formulas that have already proven their usefulness in the applied sciences, especially in physics, as illustrated by the following examples. First, consider the case of a degenerate normal random variable $Y$ with mean $\mu_Y$ and variance $\sigma^2_Y \searrow 0$. Symbolically, we may write $Y\sim N(\mu_Y,0) = \delta_{\mu_Y}$, where $\delta_{\mu_Y}$ is the Dirac measure supported on the singleton $\{\mu_Y\}$. Let us write $\sigma$ instead of $\sigma_X$ and $\mu_{XY}$ instead of $|\mu_X-\mu_Y|$. Taking the limit $\sigma^2_Y \searrow 0$ in ((ref)) and ((ref)), we obtain the respective formulations ((ref)) and ((ref)) below

eqnarray[eqnarray omitted — 372 chars of source]

Using the equality ${\displaystyle \Phi(z)=1-\hbox{\rm erfc}\left(\frac{z}{\sqrt{2}} \right)/2 }$, ((ref)) becomes

equation[equation omitted — 234 chars of source]

Lukaszyk (2004)\nocite{Lukaszyk2004} used ((ref)) to successfully implement a modified Liszka approximation method to experimental mechanics. Note that ((ref)), unlike ((ref)), leads directly to Lukaszyk's formula ((ref)).

If $X\sim N(\mu,\sigma^2)$, a direct consequence of ((ref)) is that $E|X - \mu| = \sqrt{2}\sigma/\sqrt{\pi} \approx 0.7979 \sigma$. Indeed, setting $Y\sim \delta_{\mu}$ implies that $\mu_X = \mu_Y = \mu$ and $\mu_{XY} = 0$.

In the same paper, Lukaszyk studied the case of two normal distributions having the same variance\footnote{The assumption of homoscedasticity greatly facilitated the integral calculations undertaken by Lukaszyk, as can be seen in his PH.D. thesis (2001)\nocite{Lukaszyk2001} .}, i.e. $X\sim N(\mu_X,\sigma^2)$ and $Y\sim N(\mu_Y,\sigma^2)$. Setting $\sigma_X=\sigma_Y=:\sigma$ in ((ref)) and ((ref)), we obtain successively

eqnarray[eqnarray omitted — 409 chars of source]

In Lukaszyk's paper $E|X-Y|$ appears in the form

equation[equation omitted — 209 chars of source]

which follows directly from ((ref)).

So Lukaszyk (2001, 2004) found (and used) the formulas when $\sigma_X = \sigma_Y =: \sigma$. When $\sigma_X$ and $\sigma_Y$ are arbitrary, but $\mu_X = \mu_Y$, ((ref)) or ((ref)) directly imply

equation[equation omitted — 103 chars of source]

In other words, ${\displaystyle ||X-Y||_1 = \sqrt{\frac{2}{\pi}} ||X-Y||_2 \approx 0.7979 ||X-Y||_2 }$: the ${\cal L}_1$-distance between $X \hbox{ and } Y$ is approximately one fifth smaller than the ${\cal L}_2$-distance in this case. We used the fact that $\hbox{\rm var} (X-Y) = \sigma_X^2+\sigma_Y^2$, since $X \hbox{ and } Y$ are independent and that $E(X) = E(Y)$ implies that $\hbox{\rm var} (X-Y) = E(X-Y)^2$.

Next, consider the case where $\mu_X = \mu_Y =: \mu$ and $\sigma_X = \sigma_Y =: \sigma$, i.e. $X \hbox{ and } Y$ are i.i.d. and both follow $N(\mu,\sigma^2)$. Setting $\sigma_X = \sigma_Y = \sigma$ in ((ref)), we obtain

equation[equation omitted — 85 chars of source]

which is the Gini mean difference when the underlying income distribution is Gaussian, see e.g. Yitzhaki and Schechtman (2013)\nocite{YitzhakiSchechtman2013}. Note that $E|X-Y|$ in ((ref)), which does not depend on $\mu$, is a measure of dispersion of the same nature as the standard deviation $\sigma$.

Next, if $\sigma^2_X \searrow 0$ and $\sigma^2_Y \searrow 0$, i.e. if $X\sim N(\mu_X,0) = \delta_{\mu_X}$ and $Y\sim N(\mu_Y,0) = \delta_{\mu_Y}$, equations ((ref)) or ((ref)) become

equation[equation omitted — 74 chars of source]

Of course, instead of taking the limits $\sigma^2_X \searrow 0$ and $\sigma^2_Y \searrow 0$ in ((ref)) or ((ref)), one can calculate directly $E|X-Y| = \int_{I\!\!R} \ \{ \int_{I\!\!R} |x-y| \delta(y-\mu_Y) dy\}\ \delta(x-\mu_X) dx = |\mu_X-\mu_Y |$. When $X\stackrel{a.s.}{=} a$ and $Y\stackrel{a.s.}{=} b$, $a,b \in I\!\!R$, $E|X-Y|$ simply transforms into $|a-b|$, i.e. $E|\cdot - \cdot|$ becomes the metric on $I\!\!R$ induced by the norm $|\cdot|$.

We end this subsection by observing that equation ((ref)) provides an analytic formula for $E|X|$, which enables to express $D_{norm}(X,Y)$ in ((ref)) in analytic form when $X$ and $Y$ are independent and normally distributed. To get $E|X|$, assume that $Y\stackrel{a.s.}{=} 0$, i.e. that $Y \sim \delta_0$. Then $\mu_{XY} = |\mu|$ and ((ref)) becomes

equation[equation omitted — 153 chars of source]

Average distance between coordinates of points falling at random into a proper rectangle of $I\!\!R^2$

In this section, uppercase letters $A,B$ refer to bounded proper intervals of real numbers and $L_A,L_B$ refer to their respective length. A proper interval is an interval that is neither empty (an example of empty interval is $[a,a[ = \emptyset$, for some $a \in I\!\!R$)\footnote{To avoid any confusion, we use the notation $]a,b[$ instead of $(a,b)$ in Subsections (ref) and (ref). } nor degenerate (i.e. of the form $[a,a] = \{a\}$). A proper bounded rectangle in $I\!\!R^2$ is the cartesian product $A \times B$ of two proper bounded intervals. We are interested in univariate or bivariate continuous uniform distributions such as $U(A)$ or $U(A \times B)$. Moreover, for $a \in I\!\!R$, we identify $U(\{a\})$ to the Dirac delta measure $\delta_{a}$ supported on $\{a\}$.

For a given set $C$, let $I_C$ denote the indicator function (defined by $I_C(z) = 1$ if $z\in C$, $I_C(z) = 0$ if $z\notin C$). Consider the random pair $(X,Y) \sim U(A \times B)$. Noting that $L_A^{-1}L_B^{-1} I_{A\times B}(x,y) = L_A^{-1}I_A(x) \cdot L_B^{-1}I_B(y)$ for all $(x,y) \in I\!\!R^2$, one can easily show the following (intuitive) equivalence: [$(X,Y) \sim U(A \times B)$] if and only if [$X \sim U(A)$, $Y \sim U(B)$, $X \hbox{ and } Y$ independent]. We are now ready to ask the question: a point $(x,y)$, which is a realization of $(X,Y)$, falls at random into $A\times B$. What is the average distance between its coordinates? Answer: $E|X-Y|$. Put slightly differently, $E|X-Y|$ reflects the expected absolute difference of two independent variables $X \hbox{ and } Y$ following continuous uniform distributions $X\sim U(A)$ and $Y\sim U(B)$. Theorem (ref) below expresses $E|X-Y|$ in closed form. Without loss of generality, the intervals $A$ and $B$ are assumed to be open in the theorem.

theor(Proof in the appendix) Let $A=]a_1,a_2[$ and $B=]b_1,b_2[$ be two open bounded intervals of real numbers, let $L_A = a_2 - a_1$, $L_B = b_2 - b_1$ be their lengths, and $m_A=(a_1+a_2)/2$, $m_B=(b_1+b_2)/2$ be their midpoints. Assume that $X\sim U(A)$ and $Y\sim U(B)$ are independent (or, equivalently, that $(X,Y) \sim U(A \times B)$) and consider the following three possible cases \[ \begin{array}{lll} \textbf{Case 1} & a_1 \leq b_1 <a_2 < b_2 & \hbox{ (overlap without inclusion)}\\ \textbf{Case 2} & a_1 \leq b_1 < b_2 \leq a_2 & B \subset A \hbox{ (inclusion, with $B=A$ possible)} \\ \textbf{Case 3} & a_1 < a_2 < b_1 < b_2 & A \cap B = \emptyset \hbox{ (separation).} \end{array} \] Then the expected absolute difference of $X \hbox{ and } Y$ is given in closed form by \newline $E|X-Y| =$ \[ \begin{array}{lr} L_A^{-1}L_B^{-1}[(b_2-b_1)(b_1-a_1)(b_2-a_1)/2 + (b_2-a_2)(a_2-b_1)(b_2-b_1)/2 + \frac{(a_2-b_1)^3}{3}] & (\textbf{Case 1} )\\ L_A^{-1}L_B^{-1}[(b_2-b_1)(b_1-a_1)(b_2-a_1)/2 - (b_2-a_2)(a_2-b_1)(b_2-b_1)/2 + \frac{(b_2-b_1)^3}{3}] & (\textbf{Case 2} )\\ |m_A-m_B| . & (\textbf{Case 3} ) \end{array} \] Moreover, $$ \begin{array}{lc} E|X| = \left\{ \begin{array}{ll} L_A^{-1} (a_1^2 + a_2^2)/2 & if \hskip .5cm 0\in A \\ |m_A|& if \hskip .5cm 0\notin A \end{array} \right. & \hbox{ and } \ \ E|Y| = \left\{ \begin{array}{ll} L_B^{-1} (b_1^2 + b_2^2)/2 & if \hskip .5cm 0\in B \\ |m_B|& if \hskip .5cm 0\notin B. \end{array} \right. \end{array} $$

Theorem (ref) can also be used to give closed formulas for $E|X-Y|$ when one of the two intervals $A$ or $B$ is degenerate. Suppose for example that $X\sim U(A)$ and $Y\sim \delta_b$. One way of expressing $E|X-Y|$ is then to take the cases 2 and 3 of Theorem (ref) and to calculate the limit $b_2\searrow b_1$. We obtain

equation[equation omitted — 340 chars of source]

A more direct way to proceed is to use the Lebesgue integral with the coupling $\pi = \mu \otimes \nu$ as the measure used for integration, where $\mu = U(A)$ and $\nu = \delta_b$. Indeed, let $h:I\!\!R^2 \rightarrow I\!\!R$ be given by $h(x,y) = |x-y|$. Then $E|X-Y| = \int h d\pi = \int_A \{ \int_{I\!\!R} |x-y| \delta(y-b)dy\} L_A^{-1}I_A(x) dx = L_A^{-1} \int_A |x-b| dx$. For $A=]a_1,a_2[$, distinguishing the respective situations where (i) $b\in A$ and (ii) $b\notin A$ with $b \le a_1$ and $b \ge a_2$, we find the formulas of ((ref)). If both intervals are degenerate, i.e. of type $\{ a \}$ and $\{ b \}$, then $\mu = \delta_a$, $\nu = \delta_b$, $\pi = \delta_a \otimes \delta_b = \delta_{(a,b)}$, and $\int h d\pi = |a-b|$.

Theorem (ref) allows to easily calculate the normalized average distance $D_{norm} (X,Y) = E|X-Y|/(E|X| + E|Y|) \in [0,1]$ defined in ((ref)). As evidenced by Theorem (ref), the fact that $X \hbox{ and } Y$ (and an additional $Z$) are independent implies that $D_{norm}(\cdot , \cdot)$ satisfies the triangle inequality.

The average of the distances $|a-b|$, $ a \in A, b \in B$, is not a distance

Let $\mathcal I$ denote the set of bounded proper intervals of real numbers. As $A\in \mathcal I$ and the probability measure $U(A)$ are in bijection, the closed formulas for $E|X-Y|$ in Theorem (ref) can also be taken in a purely deterministic sense to compute the “mean value” $D(A,B)$ of the set $\{|a-b|: a \in A, b \in B \}$, where $A,B \in \mathcal I$. Formally, we just have to adopt new notations: replace, say, $E|X-Y|$ by $D(A,B)$, $E|X|$ by $S_A$, $E|Y|$ by $S_B$ and $E|X-Y|/(E|X| + E|Y|)$ by $D_{norm}(A,B) = D(A,B)/(S_A + S_B)$ in Theorem (ref). More precisely, we have the weighted means $D(A,B) = \int_A \int_B |x-y| w(x,y) dxdy / \int_A \int_B w(x,y) dxdy$, $S_A = \int_A |x|u(x)dx / \int_A u(x)dx$ and $S_B = \int_B |y|v(y)dy / \int_B v(y)dy$, where $w$, $u$ and $v$ are the respective weight functions $w=I_{A \times B}$, $u = I_A$ and $v = I_B$. As $D(A,B)$ stems from $E|X-Y|$ and $D_{norm}(A,B)$ stems from $E|X-Y|/(E|X| + E|Y|)$, $D(A,B)$ is invariant under location transformations, while $D_{norm}(A,B)$ is invariant under scale transformations.

It turns out that the functionals $D$ and $D_{norm}$, both ${\mathcal I} \times {\mathcal I} \longrightarrow I\!\!R_+$, satisfy the axioms of symmetry and triangle inequality (the fact that $D_{norm}(\cdot,\cdot)$ satisfies the triangle inequality, far from obvious, is a consequence of Theorem (ref)). However, reflexivity does not hold: indeed $D(A,A) > 0$ and $D_{norm}(A,A) > 0$ for any $A \in {\mathcal I}$, which means that $D$ and $D_{norm}$ are neither distances, nor -- a fortiori -- metrics on $\mathcal I$.\footnote{Actually, $D$ and $D_{norm}$ are metametrics. The term “metametric” is specified in Deza and Deza (2014)\nocite{Deza2014}. Metametrics appear in the study of Gromov hyperbolic metric spaces. They were first defined by V\"ais\"al\"a (2005)\nocite{Vaisala2005}.}

Finally, let $b$ be any fixed real number. To obtain the closed-form formula for the “mean value” $D(A,\{ b\})$ of the set $\{ |a-b|$, $ a \in A \}$, replace $E|X-Y|$ by $D(A,\{ b\})$ in ((ref)).

Optimal transport problem for probability measures on the real line

The following brief introduction to the problem of optimal transport is intended for the many non-specialists in the field. We will stay at a rather heuristic level, focusing on the founding ideas of transport theory. For a detailed account of the theory, the reader is referred to Villani (2003, 2008)\nocite{Villani2008} for example. The problem of optimal transport can be presented in two related ways. The formulation of Monge is ancient and dates back to the 18th century. Kantorovich's work is much more recent and was published during World War II. It can be interpreted as a generalization or a relaxation of Monge's approach. In practice, the latter seems to be more direct and easier to interpret, but its resolution is mathematically more complicated.

This text is designed for a broad readership and focuses on the main ideas of the optimal transport problem. In this perspective of relative simplicity, our discussion here is limited to probability measures $\mu \hbox{ and } \nu$ defined on the real line rather than on more general spaces, so as not to lose sight of the main issues. The originality of this short presentation consists in exploiting known results (if possible with a slightly shifted look) while using notations familiar to practitioners of applied statistics, physics or econometrics. This does not prevent some new results from emerging.

The optimal mass transport problem tries to find the most efficient way to transport a source measure $\mu$ over a target measure $\nu$ taking account of a given cost function. A transport cost determines in some way the difference or distance between these measures. For the sake of simplicity, the cost function $c:I\!\!R^2 \rightarrow I\!\!R_+$ that will be used in this article is mainly of type $c(x,y) = |x-y|$, with sometimes a slight generalization: $c(x,y) = h(x-y)$, where $h$ is convex and continuous. We will see later how the choice of a cost function of type $c(x,y) = |x-y|$ leads to the expected absolute difference $E|X-Y|$ between $X \hbox{ and } Y$.

It should be borne in mind that the optimal transport problem is much more general than the particular cases treated here. This concerns notably -- as we have pointed out -- the type of spaces on which $\mu \hbox{ and } \nu$ are defined, but no less significantly the characteristics of the cost function. The restriction to the one-dimensional case and the use of a simple cost function make it possible to define the problem without technicalities -- sometimes severe -- related to more general cases. When $\mu \hbox{ and } \nu$ are defined on $I\!\!R$, or on a subset of $I\!\!R$, the problems of Monge and Kantorovich have easily interpretable closed form solutions in some important cases. This is a significant property as it alleviates the need for optimization.

When, as in this article, the source measure $\mu$ and the target measure $\nu$ are defined on the same space, the transport of measures has applications in many fields. For instance, if objects are initially distributed according to $\mu$, then they are arranged after transport according to $\nu$. In inequality economics, $\mu$ represents a distribution of income and the problem is to find a planning carrying $\mu$ over a less unequal target distribution $\nu$. In finance, $\mu$ can be the return distribution of a portfolio of stocks and $\nu$ the return of another portfolio or a benchmark.

In the next two sections, we present in detail how the Monge and Kantorovich approaches of the optimal transport unfold when $\mu \hbox{ and } \nu$ are probability measures on the real line (or have a support on the real line).

The Monge formulation

In his M\'emoire sur la Th\'eorie des D\'eblais et Remblais, Monge (1781),\nocite{Monge1781} was interested in minimizing the cost of transporting sand from a dune to fill a ditch, or transporting stones from an excavation to build a fortification. Monge's historical modeling was in $I\!\!R^3$ and the cost function was the Euclidean distance. In the generalizations that followed, $I\!\!R^3$ became for example a Polish metric space and the cost function took various forms that were quite different from the original Euclidean distance. Above all, the piles of sand or pebbles and the cavities to be filled became over time distributions of probability, of income, of wealth, configurations of physical particles, return distributions of financial assets, and so on.

In order to state Monges's problem in the one-dimensional real case, we need the following definition.

defi(transport map) Consider $\mu, \nu \in {\cal P}(I\!\!R)$, where ${\cal P}(I\!\!R)$ is the set of probability measures on $(I\!\!R ,{\cal B}_1)$. We say that a measurable map $T:(I\!\!R ,{\cal B}_1) \rightarrow (I\!\!R ,{\cal B}_1)$ transports $\mu$ to $\nu$, and we call $T$ a transport map, if $\nu(B) = \mu(T^{-1}(B))$ for all Borel subsets $B$ of $I\!\!R$.

When $T$ transports $\mu$ to $\nu$, we use the notation $\mu_T = \nu$ (rather than $T_{\#}\mu = \nu$). We denote by ${\cal T}(\mu,\nu) = \{ T: (I\!\!R, {\cal B}_1,\mu) \rightarrow (I\!\!R, {\cal B}_1)| T \hbox{ measurable with } \mu_T = \nu\}$ the set of transport maps pushing forward $\mu$ to $\nu$.

defi(Monge's formulation of the transport problem for probability measures on the real line -- or with support on the real line) Let ${\cal P}_1(I\!\!R)$ be the set of probability measures on $(I\!\!R, {\cal B}_1)$ that have finite first moment. Given $\mu, \nu \in {\cal P}_1(I\!\!R)$ and the cost function $c(x,y) = |x-y|$, find a transport map that realizes the infimum \begin{equation} \inf\{\int_{I\!\!R}|x-T(x)| \ d\mu(x): T \in {\cal T}(\mu,\nu) \}. \end{equation}

Unfortunately, we may not find any measurable map such that $\mu_T = \nu$. In other words, ${\cal T}(\mu,\nu)$ may be empty. We do not have to look very far: let us take $\mu = \delta_{x_1}$ (the Dirac delta measure supported on $\{x_1\}$) and $\nu = \frac{1}{2}\delta_{y_1} + \frac{1}{2}\delta_{y_2}$, with $y_1\neq y_2$. In this case, no map $T$ with $\mu_T = \nu$ can be found. Important cases where ${\cal T}(\mu,\nu) \ne \emptyset$ are (Thorpe 2018)\nocite{Thorpe2018}: (i) the discrete case when $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$, i.e. when $\mu \hbox{ and } \nu$ are supported on the same number of points with equal mass. And (ii) the absolutely continuous case, when $d\mu (x) = f(x)dx$ and $d\nu (y) = g(y)dy$. Moreover, even if ${\cal T}(\mu,\nu) \ne \emptyset$, the constraint in Monge's problem is usually very non-linear and difficult to handle with the classical tools of the calculus of variations. Kantorovich's approach alleviates these problems by seeking an optimal transport plan rather than an optimal transport map.

The Kantorovich formulation (often called Monge-Kantorovich formulation)

defi(Transport plan) Consider $\mu, \nu \in {\cal P}(I\!\!R)$, where ${\cal P}(I\!\!R)$ is the set of probability measures on $(I\!\!R ,{\cal B}_1)$. Let ${\cal P}(I\!\!R^2)$ denote the set of probability measures on $(I\!\!R^2, {\cal B}_2)$. We say that a probability measure $\pi \in {\cal P}(I\!\!R^2)$ whose marginals are $\mu \hbox{ and } \nu$, transports $\mu$ to $\nu$. The measure $\pi$ is called a transport plan. We say that $\pi$ has first marginal $\mu$ and second marginal $\nu$ if $\pi (A\times I\!\!R) = \mu(A)$ and $\pi (I\!\!R\times B) = \nu(B)$ for all $A,B \in {\cal B}_1$. Equivalently, if $q_1(x,y)=x$ and $q_2(x,y)=y$ are the first and second projection functions, respectively, then $\mu = \pi_{q_1}$, $\nu = \pi_{q_2}$, and $\mu(A) = \pi_{q_1}(A) = \pi(q_1^{-1}A)$, $\nu(B) = \pi_{q_2}(B) = \pi(q_2^{-1}B)$ for all $A,B\in {\cal B}_1$. The class of transport plans is denoted by $\Pi (\mu ,\nu)$; it is also called the class of all couplings between $\mu \hbox{ and } \nu$.

Note that the set $\Pi (\mu ,\nu)$ of transport plans is never empty since it contains the trivial plan $\mu \otimes \nu$. For any $A,B\in {\cal B}_1$, the quantity $\pi (A\times B)$ tells us how much mass in set $A$ is being moved to set $B$. The total amount of mass removed from $A$ has to be equal to $\mu(A)$ and the total amount of mass moved to $B$ must be $\nu (B)$. Hence the constraints: $\pi (A\times I\!\!R) = \mu(A)$ and $\pi (I\!\!R\times B) = \nu(B)$ for all $A,B \in {\cal B}_1$.

Kantorovich (1942)\nocite{Kantorovich1942} proposed a general formulation of the problem by considering optimal transport plans which allow mass to be split. This is a very important difference between the two approaches; Monge's problem, unlike Kantorovich's, requires that each mass in $x$ is sent to a single position $y$: there is no possible separation of a unit of mass of $\mu$ into several pieces during the transport. Still restricting ourselves to probabilities $\mu \hbox{ and } \nu$ defined on the real line, or having a support on the real line, we can formulate the following definition.

defi(Kantorovich's form of the transport problem) Given $\mu, \nu \in {\cal P}_1(I\!\!R)$ and the cost function $c(x,y) = |x-y|$, find a transport plan that realizes the infimum \begin{equation} \inf\{\int_{I\!\!R^2}|x-y| \ d\pi(x,y): \pi \in \Pi (\mu,\nu) \}. \end{equation}

The term $\int_{I\!\!R^2}|x-y| \ d\pi(x,y)$ represents the transport cost from $\mu$ to $\nu$ under $\pi$. We can think of $d\pi (x,y)$ as the amount of mass transferred from $x$ to $y$. Actually, as $c(x,y)=|x - y|$ is continuous\footnote{or even lower semi-continuous, a weaker condition implying the existence of a minimizer.}, a minimizer $\pi^*$ always exists (Gangbo (2004), th. 2.4)\nocite{Gangbo2004} and we can replace “inf” by “min” in ((ref)).

Probabilistic point of view: some helpful clarifications

Before resuming the substantive discussion on the problem of transport, we turn for a moment to considerations of a purely formal or didactic nature. The optimal transport domain is mainly conceived in terms of (probability) measures. Our experience is that many newcomers to this field feel more comfortable with concepts such as random variables or mathematical expectation, while the language of measure theory seems less telling to them, at least initially. We think that a brief development will clarify some aspects that specialists may consider as futile or obvious.

(i) Classically, if we have a workspace $(I\!\!R^2,{\cal B}_2, \pi)$, but wish to reason in terms of ramdom variables, we formally introduce a general probability space $(\Omega,{\cal A},P)$ and a pair of random variables $(X,Y)$ so as to obtain the scheme $(\Omega,{\cal A},P) \stackrel{(X,Y)}{\longrightarrow} (I\!\!R^2,{\cal B}_2, \pi)$, where $\Omega :=I\!\!R^2$, ${\cal A} := {\cal B}_2$ and $(X,Y) := Id_2$, the identity map on $I\!\!R^2$. As a consequence, $\pi = P_{(X,Y)} = P_{(Id_2)} = P$, noting the obvious fact that $Id_2 = (q_1,q_2)$, (as in Definition (ref), we use the notations $q_1$ and $q_2$ for the projection functions). By doing so, the representation $(\Omega,{\cal A},P) \stackrel{(X,Y)}{\longrightarrow} (I\!\!R^2,{\cal B}_2, P_{(X,Y)})$ becomes $(I\!\!R^2,{\cal B}_2, \pi) \stackrel{Id_2}\longrightarrow (I\!\!R^2,{\cal B}_2, \pi)$, i.e., in particular, the variables $X \hbox{ and } Y$ transform themselves into canonical projections\footnote{This reminds us that the term “random variable” is rather unfortunate since a random variable is neither random nor a variable. }. So $X \hbox{ and } Y$ can be interpreted indifferently as random variables or as projections. For instance, a notation such as $E_\pi |X-Y|$ instead of $\int_{I\!\!R^2}|x-y| d\pi(x,y))$ makes sense as a mean absolute deviation between two random variables. This convention will be applied below, notably in Figure (ref), which will help us articulate the Kantorovich relaxation of the Monge problem.

(ii) A transport plan $\pi \in \Pi (\mu,\nu)$ is a coupling between $\mu \hbox{ and } \nu$, i.e. a joint distribution $(X,Y) \sim \pi$ such as marginally $X \sim \mu$ and $Y \sim \nu$. By abuse of language the expression “coupling between $X \hbox{ and } Y$” is used to mean “coupling between the distribution of $X$ and the distribution of $Y$”. Considered in probabilistic form, the problem stated in ((ref)) amounts to finding the distribution of a pair of random variables $(X^*,Y^*)$ minimizing the mean absolute deviation $E|X-Y|$ among all jointly distributed pairs $(X,Y)$ such that $P_X = \mu$ and $P_Y = \nu$.

Using Kantorovich's relaxation to solve Monge's problem

In some special cases, Kantorovich's approach can be used to solve the (difficult) Monge optimisation problem. In this context, deterministic plans play an essential role.

defi(Deterministic transport plan). Let $\mu, \nu \in {\cal P}(I\!\!R)$, and denote by $Id$ the identity map on $I\!\!R$. Let $\pi_T := \mu_{(Id,T)}$ be the pushforward probability measure of $\mu$ induced by the function $(Id,T):I\!\!R \rightarrow I\!\!R^2$, where $T$ is a transport map that pushes forward $\mu$ to $\nu$. Then $\pi_T \in \Pi (\mu,\nu)$, and $\pi_T$ is called a deterministic transport plan.

Note that $\pi_T = \mu_{(Id,T)}$ is supported on the graph of $T$. Let us check that the marginals of $\mu_{(Id,T)}$ are $\mu$ and $\nu$, respectively. In this regard, consider $A,B \in {\cal B}_1$. Then, using in particular Lemma (ref) (ii) in Subsection (ref): \newline $\pi_T (A\times I\!\!R) = \mu_{(Id,T)} (A\times I\!\!R) = \mu [(Id,T)^{-1}(A\times I\!\!R)] = \mu [Id^{-1}(A) \cap T^{-1}(I\!\!R)) = \mu (A)$.\newline $\pi_T (I\!\!R \times B) = \mu_{(Id,T)} (I\!\!R \times B) = \mu [(Id,T)^{-1}(I\!\!R \times B)] = \mu (T^{-1}(B)) = \mu_T (B) = \nu (B)$, since $T$ is a transport map.

Figure (ref) provides a convenient overview of the situation. In particular, and in relation to Subsection (ref), it clearly shows that we can express the probability measures in terms of law (${\cal L}$) of random variables: there exists random variables $X \hbox{ and } Y$ such that $\pi = {\cal L}((X,Y))$, $\mu ={\cal L}(X)$, $\nu = {\cal L}(Y)$, $\mu_T = {\cal L}(T(X))$ and $\pi_T = {\cal L}((X,T(X)))$.

figure[figure omitted — 316 chars of source]

This may seem like a detail, but the way to represent the transport cost associated with a transport plan $\pi$ or a transport map $T$ varies according to the research field. Three common types of notations are shown in Table (ref) for the cost function $c(x,y) = |x-y|)$, where $c_{\pi}$ and $c_T$ denote the cost resulting from the transformation of $\mu$ into $\nu$ by means of $\pi$ and $ T$, respectively. \vskip .4cm

table[table omitted — 1,249 chars of source]

It turns out that the formulas in (b), (c) and (d) of Table (ref) are different expressions of the same quantity $c_T = \int c\circ (Id,T)\circ X d\pi$. This results immediately from a repeated use of the change of variable formula in the context of Figure (ref), recalling that $X = q_1$ under the convention adopted in Subsection (ref). For example, let us show that (c) = (b). Remembering that $c(x,y) = |x-y|$, we have $E_\pi |X-T(X)|= \int c \circ (Id,T) \circ X d\pi$ (see Figure (ref)). On the other hand, $E_{\pi_T} |X-Y| = \int c d\pi_T = \int c d\mu_{(Id,T)} = \int c\circ (Id,T) d\mu = \int c\circ (Id,T) d\pi_{q_1} = \int c \circ (Id,T) d\pi_X = \int c\circ (Id,T)\circ X d\pi$.

A heuristic interpretation of (a) is as follows: a pair of random variables $(X,Y)$ is defined on a probability space $(\Omega,{\cal A},P) = (I\!\!R^2,{\cal B}_2, \pi)$. We observe independently an infinity of realizations $(x,y)$ of $(X,Y)$ falling on $I\!\!R^2$ according to $\pi$, we calculate the distance $|x-y|$ between the coordinates and average these distances to obtain $E_\pi |X-Y|$. The three equivalent expressions (b), (c) and (d) are interpreted in an almost identical way. Take $E_{\pi_T} |X-Y|$: we observe independently an infinity of realizations $(x,y)$ of $(X,Y) \sim \pi$. But this time, we forget $y$ to replace it by $T(x)$. In other words, an infinite number of pairs $(x,y)$ fall (almost surely) on the graph of $T$. We calculate the distances between the coordinates of these pairs and average them to find $E_{\pi_T} |X-Y|$.

The most important equality in Table (ref) is

equation[equation omitted — 65 chars of source]

which shows that any transport map $T$ induces a transport plan of the same cost, i.e. can be canonically embedded into the set of transport plans. Adopting the convention that $\inf \emptyset = \infty$ if ${\cal T}(\mu,\nu) = \emptyset$, this means that

equation[equation omitted — 125 chars of source]

Equality in ((ref)) holds under fairly general assumptions when any plan can be approximated by transport maps (see e.g. Ambrosio and Pratelli (2003)\nocite{Ambrosio2003})\footnote{The presence of atoms can seriously impede the existence of transport maps. Under fairly general assumptions, it can be shown that if $\mu$ is atomless (in our setting, this means that $\mu (\{ x\}) = 0 \ \forall x\in I\!\!R$), then the set $\{ \pi_T : \mu_T = \nu \}$ is weak-*dense in $\Pi (\mu,\nu)$, which implies equality in ((ref)) (see Carlier (2010)\nocite{Carlier2010} or Ambrosio et al. (2004)\nocite{Ambrosio2004}). Ambrosio (2002)\nocite{Ambrosio2002} notes that the infimum of the Kantorovich problem “is attained on an extremal element of $\Pi (\mu,\nu)$”. However, all extremal points are not induced by transport maps “otherwise one would get existence of transport maps directly from the Kantorovich formulation”. It can be shown that deterministic transport plans are extremal in $\Pi (\mu,\nu)$. “Unfortunately, the extremal points of $\Pi (\mu,\nu)$ are not all transport plans, except in very particular cases. It turns out that the existence of optimal transport maps depends not only on the geometry of $\Pi (\mu,\nu)$, but also (in a quite sensible way) of the choice of the cost function $c$”.}.

We are now ready to show the following result: (i) if the Kantorovich problem admits an optimal plan (minimizer) $\pi$, and (ii) this plan turns out to be deterministic i.e. of the form $\pi_T$, then the transport map $T$ is Monge-optimal. To see that, let us assume that $\pi_T$ is Kantorovich-optimal. Then \newline $\inf_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| \leq E_\mu |X-T(X)| = E_{\pi_T} |X-Y| =^{(\hbox{hyp.})} \min_{\pi \in \Pi (\mu,\nu)} E_{\pi} |X-Y|$ \newline $ \leq^{(\ref{infinf})} \inf_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| $ and therefore \[ E_\mu |X-T(X)| = \inf_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| = \min_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| . \] This relaxation of Monge by Kantorovich occurs in a few important cases. We will see two examples below (in (ref) and (ref)).

Closed-form solution of the optimal transport problem in dimension one

The following theorem and its proof can be found in Thorpe (2018), \nocite{Thorpe2018} see also Villani (2003) \nocite{Villani2003} and Santambrogio (2015). \nocite{Santambrogio2015} It is a powerful tool for the treatment of the two cases mentioned above.

theorLet $\mu , \nu \in {\cal P}(I\!\!R)$, with cumulative distribution functions $F$ and $G$, respectively. Assume that $c(x,y) = h(x - y)$ where $h$ is convex and continuous. Let $\bar{\pi}$ be the probability measure on $I\!\!R^2$ with cdf $H(x,y) = \min \{F(x), G(y) \}$. Then $\bar{\pi} \in \Pi (\mu,\nu)$ and, furthermore, $\bar{\pi}$ is optimal for Kantorovich's optimal transport problem with cost function $c$.

In this subsection (Subsection (ref)), we are mainly interested in a cost function of type $c(x,y) = |x-y|$. Nevertheless, all the results obtained remain valid for a cost function such as the one specified in Theorem (ref).

Still limited to one-dimensional probability measures $\mu , \nu \in {\cal P}(I\!\!R)$, we give two examples where the Kantorovich relaxation approach leads to a solution of the Monge problem. We base ourselves on the criterion stated in Theorem (ref).\newline (i) When the continuous cost function satisfies a certain convexity condition, an optimal plan is supported on a “curve” in $I\!\!R^2$ depending on the quantile functions associated with respective cdf's $F$ of $\mu$ and $G$ of $\nu$ (quantile functions are defined in Definition (ref)) below. Moreover, if $\mu$ is atomless, that is if $F$ is continuous, then this optimal plan is deterministic, i.e. associated with a Monge-optimal transport map. \newline (ii) Theorem (ref), applied this time to the special discrete case where $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$, also leads to a deterministic optimal plan. The optimization process will provide an interesting by-product in that case (Proposition (ref)).

We now need the following definition to give an analytic representation of the optimal plan $\bar{\pi}$ referred to in Theorem (ref).

defi(Quantile function, Karr (1993)\nocite{Karr1993} p. 63). Consider a measure $\mu \in {\cal P}(I\!\!R)$ with cumulative distribution function $F$, i.e. $F(x) = \mu((-\infty , x])$. The generalized inverse $F^-$ of $F$, or quantile function associated with $F$, is defined by \begin{equation} F^-(t) = \inf \{ x \in I\!\!R:F(x) \geq t \} \ \ \ t \in [0,1]. \end{equation} Another generalized inverse can be defined: $$F^+(t) = \sup \{ x \in I\!\!R:F(x) \leq t \}.$$

The function $F^-$ always exists, even when $F$ is not continuous or not strictly increasing. As both $F$ and $F^-$ are monotonically increasing, they are also measurable, an important property that we will use later. The notation $F^{-1}$ instead of $F^-$ is used by many authors. However, it can sometimes be confusing (it may be mistaken for the preimage operator of a set, see below). If $F$ is continuous and strictly increasing, then the two generalized inverses are equal to the ordinary inverse (on the range of $F$). One can often work with $F^-$ as if it were an ordinary inverse. Note that ((ref)) implies $F^-(0) = -\infty$, and we adopt the convention that $\inf \emptyset = \infty$. Figure (ref) illustrates the inequalities $F^- \circ F(x_0) \leq x_0 \leq F^+ \circ F(x_0)$ when $x_0$ corresponds to a flat part of $F$, but these two inequalities are in fact valid for any $x_0 \in I\!\!R$.

figure[figure omitted — 317 chars of source]

An important property of the quantile function $F^-$ is the following: for each $t$ and $x$

equation[equation omitted — 77 chars of source]

(noting that to prove the implication “$\Rightarrow$”, one uses the right continuity of $F$).

We are now ready to represent in a more useful way a plan which -- like the one of Theorem (ref) -- has a cdf of type $H(x,y) = \min \{F(x), G(y) \}$, where $F$ and $G$ are the cdf's characterizing the probability measures $\mu , \nu \in {\cal P}(I\!\!R)$, respectively. To that aim, consider the “curve” $K: [0,1]\rightarrow I\!\!R^2$ given by $K(t) = (F^-(t),G^-(t))$.\footnote{We use the term “curve” by abuse of language, even if $K$ is not continuous. $K$ is a parametric curve in the usual sense when $F^-$ and $G^-$ are continuous, which is true if and only if $F$, resp. $G$, are strictly increasing. } Note that $K$ is measurable, since $F^-$ and $G^-$ are measurable. Examples of such “curves” are given in Figure (ref). We need the following lemma:

figure[figure omitted — 231 chars of source]
lemme(i) For all $b \in [0,1]$, one has $b=\lambda([0,b])$.\newline (ii) If $f = (f_1,f_2)$ is a function $E \rightarrow E_1 \times E_2$, then $f^{-1} (A \times B) = (f^{-1}_1 A)\cap (f^{-1}_2 B)$ for all $A \subset E_1$, $B \subset E_2$. \newline (iii) Define $A_x = (-\infty,x]$ and let $F^-$ be the quantile function associated with the cdf $F$. Then $(F^-)^{-1} (A_x) = [0,F(x)]$.

(i) is trivial, (ii) is well-known. Property (iii) is a consequence of ((ref)). Indeed, \newline $(F^-)^{-1} (A_x) = \{ t \in [0,1]: F^-(t) \leq x \} = \{ t \in [0,1]: t \leq F(x) \} = [0,F(x)]$.

Let us designate by $\lambda$ the Lebesgue measure restricted to $[0,1]$. Then $\lambda_K$ denotes the pushforward probability measure of the Lebesgue measure on $[0,1]$ induced by $K$ (on $(I\!\!R^2,{\cal B}_2)$).\footnote{Another way of looking at $\lambda_K$: let $X \sim F$ (resp. $Y \sim G$) be the cdf characterizing $\mu$ (resp. $\nu$), and let $U\sim U(0,1)$ be a random variable uniformly distributed on $[0,1]$. Then $\hat{X} := F^-(U) \sim \mu$, $\hat{Y} := G^-(U) \sim \nu$ and $\lambda_K$ is the law of the pair $(\hat{X},\hat{Y})$. That is, the transport plan $\lambda_K$ is a coupling of $\mu \hbox{ and } \nu$ or, in other words, $(\hat{X},\hat{Y})$ is a coupling of $X \hbox{ and } Y$. } Adopting, as in Theorem (ref), the notation $\bar{\pi}$ for a plan with cdf $H(x,y) = \min \{F(x), G(y) \}$, we show that: (A) $\lambda_K \in \Pi (\mu,\nu)$, and (B) $\lambda_K = \bar{\pi}$.

To prove (A), i.e. to prove that $\mu \hbox{ and } \nu$ are the marginals of $\lambda_K$, we use the rule governing composition of maps: the first marginal of $\lambda_{(F^-,G^-)}$ is $(\lambda_{(F^-,G^-)})_{q_1} = \lambda_{q_1 \circ (F^-,G^-)}) = \lambda_{F^-} = \mu$. To see that $\lambda_{F^-} = \mu$, we can use the points (i) and (iii) of Lemma (ref): define $A_x = (-\infty,x]$. We only have to show that $\lambda_{F^-}(A_x) = \mu (A_x)$ for any $x\in I\!\!R$. But $\lambda_{F^-}(A_x) = \lambda ((F^-)^{-1} (A_x)) \stackrel{(iii)}{=} \lambda([0,F(x)]) \stackrel{(i)}{=} F(x) = \mu (A_x)$. Using the same rule and $q_2$ instead of $q_1$, we see immediately that the second marginal of $\lambda_{(F^-,G^-)}$ is $\lambda_{G^-} = \nu$.

Next, using (i), (ii) and (iii) of Lemma (ref), we are now ready to prove (B), i.e. to prove that $\bar{\pi} = \lambda_K$. For $A_x = (-\infty,x]$ and $B_y = (-\infty,y]$, all we need to show is that $\bar{\pi}(A_x \times B_y) = \lambda_K (A_x \times B_y)$. Then $\bar{\pi} (A_x \times B_y) = H(x,y) = \min \{F(x), G(y) \} =^{(i)} \lambda([0,\min \{F(x), G(y) \}]) \newline = \lambda([0,F(x)] \cap [0,G(y)]) =^{(iii)} \lambda([(F^-)^{-1} A_x] \cap [(G^-)^{-1} B_y]) =^{(ii)} \lambda( (F^-,G^-)^{-1}(A_x \times B_y)) \newline = \lambda( K^{-1} (A_x \times B_y)) = \lambda_K( A_x \times B_y )$.

Taking into account the results we have just stated, the great contribution of Theorem (ref) is to establish that $\lambda_K$ is Kantorovich-optimal as long as the cost function $c(x,y)$ satisfies the required convexity condition.

exampleIn this example, $F$ is the cdf of a random variable $X \sim N(0,1) = \mu$ and $G$ is the cdf of $Y \sim \exp (X)$, i.e. $Y$ has lognormal distribution $logN(0,1) = \nu$. Figure (ref) shows the curve $C = K((0,1]) = (F^-,G^-)((0,1])$. The interval $(0,1]$ has been divided into ten parts of equal length: $I_1 = (0,0.1]$, $I_2 = (0.1,0.2], \ldots$, $I_{10} = (0.9,1]$. Consequently, $C$ is divided into ten portions of curve $C_1 = K(I_1)$, $C_2 = K(I_2), \ldots$, $C_{10} = K(I_{10})$. We observe that the mass is stronger on the central portions of $C$ rather than on its extremities. Next, let $A\subset I\!\!R$ and $B\subset I\!\!R_+$ be the intervals shown in Figure (ref). Intuitively, $\lambda_K(A \times B)$ can be interpreted as the amount of mass contained in $A$ that is moved to $B$ by $\lambda_K$. Since $C$ is the support of $\lambda_K$, $0.2=\lambda_K(A\times B) = \lambda_K(A\times I\!\!R_+) = \lambda_K(I\!\!R\times B) = \mu (A) = \nu (B)$. As expected, $\mu (A)$ -- the total amount of mass removed from $A$ -- and $\nu (B)$ -- the total amount of mass transferred to $B$ -- are equal.
figure[figure omitted — 445 chars of source]

Now, suppose that the cost function $c(x,y) = h(x-y)$ satisfies the convexity condition stated in Theorem (ref). Since $\bar{\pi} = \lambda_K$, the optimal cost is given by

equation[equation omitted — 74 chars of source]

We simply used the change of variables formula: $\bar{c} = \int c \ d\lambda_{(F^-,G^-)} = \int c \circ (F^-,G^-) d\lambda = \int_0^1 h(F^-(t) - G^-(t))dt$. Taking $c(x,y) = |x-y|$ as a special case, ((ref)) becomes

equation[equation omitted — 71 chars of source]

Note that $\bar{c}$ in ((ref)) is not only the $L_1$-distance between quantile functions, it is also the $L_1$-distance between the corresponding cumulative distribution functions, i.e.

equation[equation omitted — 67 chars of source]

A proof of this remarkable coincidence is given in Thorpe (2018)\nocite{Thorpe2018}, see also Rachev and Rueschendorf (1998)\nocite{Rachev1998}. It should also be noted that Theorem (ref) does not make any particular assumption on $F$ or $G$ and that the optimal plan $\bar{\pi} = \lambda_K$ has not necessarily the deterministic form $\pi = \pi_T =\mu_{(Id,T)}$ for a certain transport map $T$; therefore does not, as it stands, help to solve Monge's problem. The question now is to give additional assumptions about $\mu \hbox{ and } \nu$ (or $F$ and $G$) ensuring that $K((0,1])$ is the graph of a transport map $T$: the very fact that this curve is the graph of a transport map $T$ means that $T$ is Monge-optimal. We give below two special instances where this happens: $K((0,1])$ is the graph of a transport map $T$ (i) when $F$ is continuous (Subsection (ref)) and (ii) in the discrete case, when $\mu \hbox{ and } \nu$ are supported on the same number of points of identical mass (Subsection (ref)). Indeed, these two subsections yield two situations where, under the assumptions of Theorem (ref), $\lambda_{(F^-,G^-)} = \mu_{(Id,T)}$.

Deterministic optimal plan when $F$ is continuous

Examples of continuous cdf's appear in Figure (ref). We now show that if $F$ is continuous, then the deterministic plan $\pi_{G^-\circ F} = \mu_{(Id,G^-\circ F)}$ has distribution function $H(x,y) = \min \{F(x), G(y)\}$. Theorem (ref) then implies that $\mu_{(Id,G^-\circ F)}$ is Kantorovich-optimal for any cost function having the convexity property specified in this theorem. This means that $T = G^-\circ F$ is Monge-optimal for the same cost function. We need the following lemma:

lemme(Proof in the appendix) For $\mu \in {\cal P} (I\!\!R )$ and $x \in I\!\!R$, define $A_x = (-\infty ,x]$ and assume that $F(x) = \mu (A_x)$ is continuous. Then \newline {\rm (a)} $\mu_F = \lambda$, where $\lambda$ denotes the Lebesgue measure restricted to $[0,1]$.\newline {\rm (b)} For any Borel subset $C$ of $I\!\!R$ and any $x_0 \in I\!\!R$, $\mu (A_{x_0} \cap C) = \mu (\ F^{-1} ([0, F(x_0)]) \cap C \ )$.
figure[figure omitted — 651 chars of source]

We will use the properties (i), (ii) and (iii) of Lemma (ref), as well as the properties (a) and (b) of Lemma (ref). Let $(x,y) \in I\!\!R^2$. According to Theorem (ref), all we have to do is show that $\pi_{(G^-\circ F)} (A_x \times B_y) = \min \{ F(x),G(y)\}$, where $A_x=(-\infty ,x]$ and $B_y=(-\infty ,y]$.

We have \newline $\pi_{(G^-\circ F)} (A_x \times B_y) = \mu_{(Id,G^-\circ F)}(A_x \times B_y) =^{(ii)} \mu [\ A_x \cap F^{-1} ((G^-)^{-1}(B_y))\ ]$ \newline $=^{(iii)} \mu (\ A_x \cap F^{-1} ([0,G(y)]\ ) =^{(b)} \mu (\ F^{-1} ([0, F(x)]) \cap F^{-1} ([0,G(y)] \ ) $ \newline $= \mu (\ F^{-1} ([0, F(x)] \cap [0,G(y)])\ ) = \mu_F (\ [0, F(x)] \cap [0,G(y)]\ ) =^{(a)} \lambda (\ [0,\min \{ F(x),G(y)\}]\ )$ \newline $=^{(i)}\min \{ F(x),G(y)\} = H(x,y)$.

In Example (ref), $F$ (standard normal) and $G$ (lognormal) are both continuous and correspond to case (a) in Figure (ref). If $c(x,y) = h(x-y)$ is a cost function with $h$ convex and continuous, the curve $C$ in Figure (ref) is the graph of a (strictly increasing) optimal transport map $T(x) = G^- \circ F (x)$.

Deterministic optimal plan in a special discrete case

We now propose a new case where the solution of Monge's problem involves Kantorovich' relaxation. Let the cost function $c(x,y) = h(x-y)$ be as in Theorem (ref) (i.e. $h$ is convex and continuous) and assume that $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$. As we are dealing here with dimension one, the $x_i$'s and $y_j$'s are real numbers, and we assume that they are ordered: $x_1 \leq x_2 \leq \cdots \leq x_n$ and $y_1 \leq y_2 \leq \cdots \leq y_n$. Define $t_i = F(x_i)$, and $s_j = G(y_j)$, $i,j \in \{ 1, \ldots , n \}$. As all points $\{x_i\}$ and $\{y_j\}$ have the same mass ($1/n$), $t_i = s_i$, $i = 1, \ldots , n$. To prove that $\lambda_{(F^-,G^-)} = \mu_{(Id,G^-\circ F)}$, all we have to show is that $\lambda_{(F^-,G^-)}(\{(x_i,y_j)\}) = \mu_{(Id,G^-\circ F)} (\{(x_i,y_j)\})$ for all $i,j \in \{ 1, \ldots , n \}$. Consider the partition $\{(0,t_1],(t_1,t_2], \ldots , (t_{n-1},t_n =1]\}$ of the interval $(0,1]$, each element of the partition being of length $1/n$. Since $(F^-,G^-)((0,t_1]) = (x_1,y_1), (F^-,G^-)((t_1,t_2]) = (x_2,y_2), \ldots , (F^-,G^-)((t_{n-1},t_n]) = (x_n,y_n)$, we have \newline $\lambda_{(F^-,G^-)}(\{(x_1,y_1)\}) = \lambda_{(F^-,G^-)}(\{(x_2,y_2)\}) = \cdots = \lambda_{(F^-,G^-)}(\{(x_n,y_n)\}) =1/n$. And since \newline $\sum_{i=1}^n \lambda_{(F^-,G^-)}(\{(x_i,y_i)\}) = 1$, we have $\lambda_{(F^-,G^-)}(\{(x_i,y_j)\}) = \frac{1}{n}\delta_{ij}$ where $\delta_{ij}$ is the Kronecker delta ($\delta_{ij}$ equals one if $i=j$, zero otherwise). On the other hand, note that $x_i \stackrel{F}{\longmapsto} t_i = s_i \stackrel{G^-}{\longmapsto} y_i$ (because $G^-$ is left-continuous), and therefore

equation[equation omitted — 80 chars of source]

Then $\mu_{(Id,G^-\circ F)} (\ \{(x_i,y_j)\}\ ) = \mu [\ (Id,G^-\circ F)^{-1} \{(x_i,y_j)\}\ ] = \mu (\ \{x_i\}\cap (G^-\circ F)^{-1} \{y_j\}\ ) =^{(\ref{optiT1})} \mu (\{ x_i\} \cap \{ x_j\}) = \frac{1}{n}\delta_{ij}$, and we have thus proved that $\lambda_{(F^-,G^-)} = \mu_{(Id,G^-\circ F)}$. Denoting by $T^* = G^-\circ F$ the Monge-optimal transport map, one gets from ((ref))

equation[equation omitted — 71 chars of source]

Proposition (ref) is a consequence of the above.

prop(Proof in the appendix) Consider two ordered sets of real numbers $x_1 \leq x_2 \leq \cdots \leq x_n$ and $y_1 \leq y_2 \leq \cdots \leq y_n$. Let $S$ be the set of permutations $\sigma: \{ 1,\ldots ,n \} \rightarrow \{1,\ldots ,n \}$ and let $h$ be a convex continuous function. Then \begin{equation} \min_{\sigma \in S} \sum_{i=1}^n h(x_i - y_{\sigma(i)}) = \sum_{i=1}^n h(x_i - y_i) . \hbox{ In particular } \end{equation} \begin{equation} \min_{\sigma \in S} \sum_{i=1}^n |x_i - y_{\sigma(i)}| = \sum_{i=1}^n |x_i - y_i| = ||x - y||_1 , \end{equation} the Manhattan distance between $x=(x_1, \ldots ,x_n)$ and $y=(y_1, \ldots ,y_n)$.

Note that the results shown in Subsection (ref) have a multidimensional extension when the cost function $c$ is general and the optimization problem restricts to discrete measures \newline $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$, where $x_i , y_j \in I\!\!R^d$, $d \ge 2$. However, the $x_i$'s and $y_j$'s cannot be totally ordered in that case and the Monge-optimal transport map $T_{\sigma}$ defined by $T_\sigma (x_i) = y_{\sigma(i)}$, $i = 1,\ldots ,n$, is not necessarily the one corresponding to the trivial permutation $\bar{\sigma}(i) := i$. Using the Minkowski-Carath\'eodory Theorem and the Birkhoff Theorem, Thorpe (2018, Th. 2.5 and 2.6)\nocite{Thorpe2018} shows that any solution $\pi^*$ to Kantorovich's optimal transport problem is a permutation matrix, i.e. there exists a permutation $\sigma^* \in S$ such that $\pi_{ij}^* = \frac{1}{n}\delta_{j=\sigma^*(i)}$, which implies that $T^*: I\!\!R^d \rightarrow I\!\!R^d$ defined by $T^*(x_i) = y_{\sigma^*(i)}$ is a Monge-optimal transport map, see also Villani (2003) \nocite{Villani2003}.

The results of Section (ref) can be summarized as follows: let probability measures $\mu ,\nu$ on $(I\!\!R, {\cal B}_1)$ have the respective cdf's $F$ and $G$. For a cost function $c(x,y) = |x-y|$, or more generally $c(x,y) = h(x-y)$ with $h$ convex and continuous, the transport plan $\lambda_{(F^-,G^-)}$ supported on the “curve” $C = (F^-,G^-)(\ (0,1]\ )$ is Kantorovich-optimal. In this context, we gave two examples (special cases) where the Kantorovich-optimal transport plan $\lambda_{(F^-,G^-)}$ becomes $\mu_{(Id,G^-\circ F)}$ and the curve $C$ is the graph of a Monge-optimal transport map $T = G^- \circ F$. In the first example, $F$ is assumed to be continuous, while in the second example, $\mu \hbox{ and } \nu$ are assumed to be discrete and supported on the same number of ordered real numbers having the same mass. Proposition (ref) is a consequence of the latter case.

Optimal transport cost as a metric

We add this section for completeness, noting that the optimal transport problem can be put forward to define a distance between $\mu \hbox{ and } \nu$, namely the so-called Wasserstein metric (or Wasserstein distance). It is also known, in computer sciences, as the earth mover's distance. Like the optimal transport problem discussed in this paper, the definition of the Wasserstein distance is set in a much broader context than the one presented below. This concerns the shape or the properties of the cost function, the characteristics of the spaces on which $\mu$ and $\nu$ are defined (usually $I\!\!R^d$ or subsets of $I\!\!R^d$), as well as the value of $p$ in the definitions that follow.

The set of probability measures on $(I\!\!R ,{\cal B}_1)$ (or on $(E,E \cap {\cal B}_1)$, $E \subset I\!\!R$) with finite $p$-th moment is defined as \[ {\cal P}_p (I\!\!R ) = \left\{ \mu \in {\cal P} (I\!\!R ) :\int_{I\!\!R}|x|^p d\mu(x) < \infty \right\}, \] noting that if $E \subset I\!\!R$ is bounded, then ${\cal P}_p (E ) = {\cal P} (E )$.

defiFor $\mu, \nu \in {\cal P}_p (I\!\!R )$ and $p\in [1,\infty)$, the Wasserstein distance between $\mu \hbox{ and } \nu$ is defined as \begin{equation} W_p(\mu ,\nu) = \left( \inf_{\pi \in \Pi (\mu,\nu)} \int_{I\!\!R^2}|x-y|^p d\pi(x,y) \right)^{1/p}, \end{equation} that is, the Wasserstein distance is the $p^{th}$ root of the minimum of the Kantorovich optimal transport problem for cost function $c(x,y) = |x-y|^p$.

It can be shown that the distance $ W_p: {\cal P}_p(I\!\!R) \times {\cal P}_p(I\!\!R) \rightarrow [0,\infty)$ is a metric on ${\cal P}_p(I\!\!R)$. Under the probabilistic notation adopted above, where the random variables are projections (i.e. $X = q_1$ and $Y = q_2$), ((ref)) becomes \[ W_p(\mu ,\nu) = \left( \inf_{\pi \in \Pi (\mu,\nu)} E_\pi |X-Y|^p \right)^{1/p}. \] In particular, for $p = $1

equation[equation omitted — 95 chars of source]

that is, since $c(x,y) = |x-y|$ is convex and continuous

equation[equation omitted — 94 chars of source]

as we have seen in ((ref)) and ((ref)).

Concluding remarks

In this article, we focused on the $L_p$ distances and in priority on $L_1$. Its content is intended to be a compromise between theoretical considerations and applications. We did not dwell on the already well-known properties of these distances, but we went back to basics of probability theory to clarify various aspects that applied science papers tend to neglect. We have examined the relationship between the $L_1$ distance and the Gini-Kantorovich distance (an $L_1$ distance between cfd's), as well as certain uses of the former, such as the Gini mean difference or the Lukaszyk-Karmowski metric. Unlike simple metrics such as the Gini-Kantorovich distance, $E|X-Y|$ integrates the dependency structure between $X \hbox{ and } Y$ and makes it possible in particular to take account of the assumption of independence. We then studied the axiomatic in which $E|X-Y|$ is inscribed; this allowed us to uncover a interpretive error that crept into the literature on the subject. The properties of $E|X-Y|$ have been clarified in the cases of independence, equality of distribution and almost sure equality. The problem of the normalization of $E|X-Y|$ has been solved, especially the question of triangle inequality. The Gini index is a special case of this [0,1]-normalized form. We have also shown that, for identically distributed variables $X \hbox{ and } Y$, $E|X-Y|$ can be interpreted as a distance to almost sure equality. In a section reserved more specifically for applications, $E|X-Y|$ is expressed in analytic form when $X\sim N(\mu_X,\sigma^2_X)$ and $Y\sim N(\mu_Y,\sigma^2_Y)$ are independent. The resulting formula generalizes tools used in applied physics: it allows in particular to lift the assumption of homoscedasticity ($\sigma^2_X = \sigma^2_Y$), and thus allows more flexibility in the use of the Lukaszyk-Karmowski metric. Closed-form formulas are also determined when independently distributed $X \hbox{ and } Y$ have continuous uniform distributions. These formulas can be used to compute the average distance between intervals. Moreover, leads are opened for obtaining analytical forms when $X \hbox{ and } Y$ follow non-normal distributions, noting that such an attempt can be very complicated or even hopeless in some cases. Finally, for two probability measures $\mu$ and $\nu$ defined on the real line, $E|X-Y|$ is a key ingredient in the optimal transport problem when the issue is how to transport $\mu$ to $\nu$, whilst minimizing a cost function of the form $c(x,y) = |x-y|$. The question of optimal transport is developed in a relatively complete way within this restricted framework. The way the problem is presented has been thought of as a first approach to a particularly demanding domain.