Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
171,399 characters · 42 sections · 0 citation commands
Various issues around the $L_1$-norm distance
The notion of distance is fundamental in human experience; human beings constantly need to represent some degree of closeness between objects, whether the latter are physical or symbolic, concrete or abstract. Quantifying the closeness between random objects has become a task of vital interest to virtually all researchers working in applied sciences. This text is designated for a broad readership and we tried to make it as self-contained as we reasonably can, with as few “it can be shown that” as possible. As a result, the presentation contains a relatively larger amount of background material than it is usually found in articles dealing with comparable subjects. Our hope is that readers mainly interested in applications will find a theory that is accessible to them.
Conceptual metric spaces are usually meant to have properties similar to those of the “natural” metric $|x-y|$ of the real line. We can ask ourselves the following question: which distance should we use when the real numbers $x$ and $y$ are replaced by real-valued random variables $X \hbox{ and } Y$? The answer is not unique, but it immediately comes to mind to look for it in the family of distances resulting from the $L_p$ norm. Let us first recall the context in which this norm is used. In what follows a probability space will be denoted by $(\Omega,{\cal A},P)$, where the sample space $\Omega$ is endowed with a $\sigma$-field ${\cal A}$ and a probability function $P$, and ${\cal B}_1$ will designate the $\sigma$-field of Borel subsets of $I\!\!R$. If $p \in [1, \infty )$, which is the range of interest of most applications and the range we will consider in this article, the space $L_p (\Omega, {\cal A}, P)$ -- also denoted by $L_p (I\!\!R)$ or even more simply by $L_p$ -- consists of all $p$-(Lebesgue) integrable random variables $(\Omega, {\cal A},P) \rightarrow (I\!\!R, {\cal B}_1)$, i.e. random variables that satisfy $E|X|^p < \infty$. Then, if $X \in L_p$, we define the $L_p$ norm of $X$ by $||X||_p = (E|X|^p )^{1/p}$. Note that we are facing the following technical point: $||X||_p = 0$ does not imply that $X=0$, but only that $X=0$ almost surely. To be quite precise, the set of $p$-integrable real-valued random variables together with the function $||\cdot ||_p$ is a seminormed vector space denoted by ${\cal L}_p (\Omega, {\cal A}, P)$, ${\cal L}_p(I\!\!R)$ or simply ${\cal L}_p$ . Then the quotient space $L_p (\Omega, {\cal A}, P)$ is defined as the normed vector space of the equivalence classes for the equivalence relation: $X \sim Y $ if and only if $X = Y$ almost surely, where $X,Y \in {\cal L}_p(\Omega, {\cal A}, P)$. In other words the random variables which agree almost surely are identified. The passage of ${\cal L}_p$ to $L_p$ is convenient, but also a bit confusing. However, many authors use the notation $L_p$ to refer to either space. In practice, we often forget that we are in the presence of equivalence classes rather than random variables. In short, ${\cal L}_p$ refers to a set of random variables while $L_p$ is a set of classes for the a.s. equality relation. We will make a clear distinction when other equivalence relations are involved, e.g. in Subsection (ref).
The spaces $L_p$ are complete normed vector spaces, that is Banach spaces. “Complete” means that the limit of any Cauchy sequence is within the space itself. Among all $p \in [1, \infty )$, the case $p=2\ $ is special: the norm $||X||_2$ follows from the inner product of $L_2$, and $(L_2,||X||_2)$ is the only Hilbert space\footnote{A Hilbert space is a Banach space whose norm is determined by an inner product. } of the family. Finally $1 \le p < q < \infty$ implies that $||X||_p = (E|X|^p )^{1/p} \le (E|X|^q )^{1/q} = ||X||_q$ and hence $L_q \subset L_p$ (in particular $L_2 \subset L_1$).
We still haven't answered the question of replacing the distance $|x-y|$ between two real numbers $x$ and $y$ when instead of them we have to deal with real-valued random variables $X \hbox{ and } Y$. Let $\delta_x$ (resp. $\delta_y$) denote the Dirac delta measure supported on the singleton $\{ x \}$ (resp. $\{ y \}$). If $X \sim \delta_x$ and $Y \sim \delta_y$, then the distance $||X-Y||_p$ equalizes to $|x-y|$. Indeed, $E|X-Y|^p = \int_{I\!\!R} \ \{ \int_{I\!\!R} |s-t|^p \delta(t-y) dt\}\ \delta(s-x) ds = |x-y|^p$, and thus $||X-Y||_p = (E|X-Y|^p )^{1/p} = |x-y|$. So all distances $||X-Y||_p$ generalize $|x-y|$ in the sense described above. Which $p$ should we choose? A serious candidate is $p = 2$, by virtue of the Hilbertian character of the space $(L_2,||\cdot ||_2)$. But if in a given study we are not interested in the mathematical facilities that the existence of an inner product allows (ability to treat random variables as vectors, projection on a closed convex set, orthogonality, etc.), then the choice of $(L_1,||\cdot ||_1)$ seems more suitable: let us look in this respect at the $L_2$-norm distance $||X-Y||_2 = (E|X-Y|^2 )^{1/2}$. Its use gives rise to a distortion in that it squares the differences: this distance tends to underestimate small differences (smaller than 1) and overestimate large differences (larger than 1), although the problem is mitigated by taking the square root at the end of the calculation. As the $L_1$ distance is not subject to this type of deformation, it comes first in terms of simplicity and interpretation. And all the more so since it is less than or equal to all other $L_p$ distances: $||X-Y||_1 \le ||X-Y||_p$ $\forall p > 1$. For all these reasons we will mainly focus on the distance $||X-Y||_1 = E|X-Y|$ between real-valued integrable random variables $X \hbox{ and } Y$ defined on the same probability space $(\Omega,{\cal A},P)$, namely
where $P_{(X,Y)}$ is the pushforward probability of $P$ induced on $I\!\!R^2$ by the random pair $(X,Y)$. Equation ((ref)) can be given the following interpretation: a pair of random values $(x,y)$ taken by the jointly distributed $X \hbox{ and } Y$ is observed. The absolute difference between $x$ and $y$ is recorded. The sampling procedure is repeated independently an infinite number of times and the observed absolute differences $|x-y|$ are averaged. This endless process will yield $E|X-Y|$.
The fields concerned with the expected absolute difference are numerous and various. They include in particular data analysis, clustering, optimal transport, physics, biology, economics, finance, engineering, image analysis. Interestingly, an identical distance measurement is subject to sporadic reappearances in areas that have a priori nothing in common. For example, the Gini mean difference (GMD) used in inequality economics -- and sometimes also considered as an $L_1$ alternative to the standard deviation --, or the so-called Lukaszyk-Karmowski metric used in mechanical physics or in quantum physics, have been proposed independently, according to the specific needs of their domain. Both are expressions of the statistical distance $E|X-Y|$ between two random variables $X$ and $Y$. In the case of GMD, $X \hbox{ and } Y$ are independent and identically distributed (i.i.d), whilst in the case of the Lukaszyk-Karmowski metric they are usually assumed to be independent. Not surprisingly, $E|X-Y|$ appears under different names in the literature, including expected (mean, average) absolute difference (deviation) between variables, ${\cal L}_1$- distance, $L_1$-distance, $L_1$-norm distance, $L_1$-metric, 1-average compound metric. In this document, $E|X-Y|$ will almost always be referred to as the expected absolute difference between $X \hbox{ and } Y$.
As we will frequently encounter the notion of “distance”, “semimetric”, “metric” and “metametric”, it is certainly not useless to specify the mathematical properties that these words cover. Let $E$ be a set and consider a function $d:E\times E\longrightarrowI\!\!R_+ := [0,\infty )$. This non-negative function may have various properties that must hold for all $x$, $y$, $z\in E$:
If $d$ satisfies reflexivity and symmetry, it is called a distance and the ordered pair $(E,d)$ is called a distance space. If $d$ satisfies reflexivity, symmetry and triangle inequality, it is a semimetric and $(E,d)$ is a semimetric space. If $d$ satisfies reflexivity, reverse reflexivity, symmetry and triangle inequality, it is a metric and $(E,d)$ is a metric space. Reflexivity and reverse reflexivity together constitute the so-called \emph{identity of indiscernibles}. We will exercise some latitude when using the word “distance” even if we are actually talking about metrics or semimetrics.
This work organizes the reflexion around $E|X-Y|$ in the following way: Section (ref) shows how the $L_1$-distance between distribution functions\footnote{which is also the 1-Wasserstein distance between the corresponding distributions.} and the expected absolute difference are related and how two separate experiments can be consistently unified. Section (ref) deals with the axiomatic of probability metrics, $E|X-Y|$ being what is called a compound metric. On this occasion, a well-hidden logical inconsistency tainting published work is brought to light. Section (ref) focuses on the behavior of $E|X-Y|$ in relation to independence, almost sure equality, or equality of distribution of the variables $X \hbox{ and } Y$. The normalized expected absolute difference is discussed in Section (ref), where the prominent role of independence is emphasized. A primary metric defined on the distributions of pairs of random variables is specified in \textbf{Section (ref)}, in order to provide a new interpretation of the expected absolute difference. This leads to a very general expression for the \emph{Gini mean difference} and the \emph{Gini index}. In \textbf{Section (ref)}, we give in analytic form the expected absolute difference between two independent normally distributed random variables. We end up with a result generalizing formulas used in applied physics an in economics. In the process, we also give the analytic form of the average distance between coordinates of points falling at random into a proper rectangle of $I\!\!R^2$. We envision that \textbf{Section (ref)} can provide the basic background material to understand the main concepts of the optimal transport theory. The latter is consciously presented in a restricted framework, as a first step in the access to a complex field in full expansion. It is precisely these restrictions that allow to bring to light very telling results, sometimes even spectacular, in any case of a indeniable mathematical beauty. In this context, the presence of closed-form solutions to the optimal transport problem allows - or at least greatly facilitates - a good understanding of the subject through important special cases.
Here are some of the notations used throughout this article: we write $I\!\!R_+$ for the set of non-negative real numbers, and ${\cal B}_d$ refers to the $\sigma$-field of Borel subsets of $I\!\!R^d$, $d\ge 1$. Moreover, ${\cal P}(I\!\!R^d)$ denotes the set of probability measures on $(I\!\!R^d,{\cal B}_d)$ and ${\cal P}_p(I\!\!R^d)$ the set of probability measures on $(I\!\!R^d,{\cal B}_d)$ with finite $p$-th moment. We will be mainly interested in random variables $X,Y,Z,\ldots$ defined on $(\Omega,{\cal A},P)$ taking their values in $(I\!\!R,{\cal B}_1)$. Let $V:(\Omega,{\cal A},P) \rightarrow (I\!\!R^d,{\cal B}_d)$ be a random variable or a random vector defined on a given probability space. We denote by $P_V$ the pushforward probability of $P$ induced by $V$ on $(I\!\!R^d,{\cal B}_d)$. For example, $P_X$, $P_{(X,Y)}$, $P_{(X,Y,Z)}$ refer to pushforward probability measures of $P$ on $(I\!\!R,{\cal B}_1)$ (resp. $(I\!\!R^2,{\cal B}_2)$, $(I\!\!R^3,{\cal B}_3)$) induced by $X$ (resp. $(X,Y)$, $(X,Y,Z)$). The notation ${\cal L}_1(I\!\!R^d)$, $d\ge 1$, will refer to the space of integrable random variables or vectors defined on a probability space $(\Omega,{\cal A},P)$ which take their values in $(I\!\!R^d,{\cal B}_d)$. Almost sure equality (resp. equality of distribution) of random variables $X \hbox{ and } Y$ are denoted by $X\stackrel{a.s.}{=} Y$ (resp. $X\stackrel{d}{=} Y$). The respective abbreviations cdf and pdf stand for cumulative distribution function and probability density function.
Subsection (ref) focuses on the relationship between the Gini-Kantorovich distance (a $L_1$-distance between cdf's) and the expected absolute difference (a $L_1$-norm distance between random variables). We discuss properties of these distances and examine the historical premises of the optimization problem at the origin of the link that unites them. Subsection (ref) discusses how two separate experiments can be consistently unified.
A statement such as “random variables $X$ and $Y$ are defined on the same probability space” implies that $X$ and $Y$ have a joint distribution, in which case we say that they are coupled. First, consider two integrable real-valued random variables $X$ and $Y$ that may not be defined on the same probability space. If one knows their (individual) distributions only, one can define a distance between them by using their respective cumulative distribution functions $F_X$ and $F_Y$. A rather intuitive way of measuring this distance is to calculate the Gini-Kantorovich distance
which can be easily visualized as a surface between two curves. One can show that
where $F_X^-$ and $F_Y^-$ are the quantile functions or generalized inverses of $F_X$ and $F_Y$, respectively. A proof of this remarkable coincidence is given in Thorpe (2018)\nocite{Thorpe2018}, see also Rachev and Rueschendorf (1998)\nocite{Rachev1998}. The generalized inverse is defined in Subsection (ref) (Definition (ref)). For more details about quantile functions, see Karr (1993)\nocite{Karr1993} p. 63, or Embrechts and Hofer (2014)\nocite{Embrechts2014}.
$GK$, also known as the Gini index of dissimilarity\footnote{$GK$ is also called the $L_1$-metric between distribution functions or Monge-Kantorovich metric or Hutchinson metric, 1-Wasserstein metric, Fortet-Mourier metric, see Deza and Deza (2014)\nocite{Deza2014}.}, is a special case of a more general Gini-Kantorovich metric $GK_p$ when $p=1$ (Ortobelli et al. (2006)\nocite{Ortobelli2006}). The Gini index of dissimilarity should not be confused with the Gini mean difference (GMD) or the Gini index discussed later in this article (Subsections (ref), (ref) and (ref)). Rachev et al. (2013)\nocite{Rachev2013}, note that ((ref)) is the explicit solution of a minimization problem studied by Gini (1914)\nocite{Gini1914a} and solved by Salvemini (1943)\nocite{Salvemini1943} for discrete cdf's and Dall'Aglio (1956)\nocite{Dallaglio1956} in the general case. More precisely, let ${\cal F}(F_1,F_2)$ denote the set of all bivariate cdf's $F$ with marginal cdf's $F_1$ and $F_2$. Then the analytic solution of the minimization problem
is “simply”
The optimization problem ((ref)) and its solution are often expressed as
where $X \hbox{ and } Y$ are given random variables and where “$\stackrel{d}{=}$” refers to equality of distribution (Rachev et al. (2007), Eq. 3.23\nocite{Rachev2007})\footnote{Rachev et al. (2013)\nocite{Rachev2013} formulate concisely the more general problem of mass transportation studied by Kantorovich, of which the classic transportation problem in linear programming and the minimization problem ((ref)) are special cases.}. Actually, as the map $c(x,y) := |x - y|$ is continuous\footnote{or even lower semi-continuous, a weaker condition. In optimal transport, $c(x,y)$ denotes the transportation cost function.}, a minimimizer does exist (Gangbo (2004), Th. 2.4)\nocite{Gangbo2004} and we can replace “inf” by “min” in ((ref)) and ((ref)). Typically, ((ref)) shows that the infimum runs over all couplings\footnote{Couplings are defined in Subsection (ref). } $(\tilde{X},\tilde{Y})$ of $X \hbox{ and } Y$, where $X \hbox{ and } Y$ may or may not be defined on the same probability space. Now suppose that $X \hbox{ and } Y$ are both defined on a probability space $(\Omega,{\cal A},P)$, i.e. are jointly distributed. Then a look at ((ref)) -- where $GK$ depends only on the individual cdf's of $X \hbox{ and } Y$ -- confirms that $GK$ ignores any structure of dependence or independence inside the pair $(X,Y)$. In the process of minimization described in ((ref)), of which $GK(X,Y)$ is the solution, any dependence or independence structure is swept away: we are left only with a probability metric measuring the $L_1$-distance between the cdf's $F_X$ and $F_Y$ of $X \hbox{ and } Y$, respectively. Such a situation is unsatisfactory because it ignores valuable information that can be available in practice, for example when one can assume that $X \hbox{ and } Y$ are independent. This can be illustrated with a simple example: Table (ref) shows two joint distributions of binary $\{0,1\}$-valued random variables.
Distribution (a) in Table (ref) reflects a dependence between $X$ and $Y$, while distribution (b) corresponds to independence. In both cases, $GK(X,Y) = 0.5$, ignoring the dependence structure between the variables. Note that $E|X-Y| = 0.7$ (case (a)) and 0.62 (case (b)). A probability metric such as $GK(X,Y)$ makes sense if $X$ and $Y$ are uncoupled, i.e. if we only know their one-dimensional cdf's. When $X \hbox{ and } Y$ are coupled, there are more informative ways of determining how far apart they are from each other. As $GK$ cannot take full account of the information of the model, it may be replaced by the expected absolute difference. As a matter of fact, $E|X-Y|$ uses all the information contained in the probability space $(\Omega,{\cal A},P)$ governing the distribution of the pair $(X,Y)$ to determine how far $X$ is from $Y$. An emblematic case occurs when $X \hbox{ and } Y$ can be assumed to be independent. It is well-known that under the independence assumption $X: (\Omega_1,{\cal A}_1,P_1) \rightarrow (I\!\!R, {\cal B}_1)$ and $Y: (\Omega_2,{\cal A}_2,P_2) \rightarrow (I\!\!R, {\cal B}_1)$ can be defined trivially on the same probability space, namely the product space $(\Omega,{\cal A},P) = (\Omega_1 \times \Omega_2,{\cal A}_1 \otimes {\cal A}_2,P_1 \otimes P_2)$. Consequently, $X \hbox{ and } Y$ are jointly distributed and the use of $GK$ (or any other similar metric) would be inappropriate.
(The reader familiar with measure or probability theory may skim this subsection). Fundamental probabilistic concepts, although often trivialized in applied papers, are not always sufficiently understood. We have seen in the previous section how the fact that random variables are jointly distributed or not can affect the choice of an adequate distance function. In connection with the content of Subsection (ref), we recall the rules that must be respected so that two separate experiments can be adequately combined into a joint experiment.
Consider two random experiments ${\cal E}_1$ and ${\cal E}_2$. Suppose that the information on the experiments is captured by real numbers; that is, ${\cal E}_1$ and ${\cal E}_2$ are completely described by the respective probability spaces $(I\!\!R, {\cal B}_1,P_1)$ and $(I\!\!R, {\cal B}_1,P_2)$\footnote{For simplicity, we suppose that $I\!\!R$ is the common sample space of ${\cal E}_1$ and ${\cal E}_2$ and that $I\!\!R$ is endowed with ${\cal B}_1$, to obtain the common measurable space $(I\!\!R,{\cal B}_1)$. }. If the unification operation is conducted in a coherent way, then ${\cal E}_1$ and ${\cal E}_2$ can be seen as “marginal” experiments of ${\cal E}$. A probability model $(\Omega,{\cal A},P)$ has to be defined to describe the “joint experiment” ${\cal E}$.
Random variables $X$, $Y$ -- and the resulting pair $(X,Y)$ taking values in $(I\!\!R^2, {\cal B}_2)$ -- can be defined on a common probability space $(\Omega,{\cal A},P)$ so that $P_1$ and $P_2$ are the marginal probability measures of $P$. Indeed, we can define $\Omega = I\!\!R^2$, ${\cal A} = {\cal B}_1 \otimes {\cal B}_1 = {\cal B}_2 $ and $(X,Y)= Id_2 = (q_1,q_2)$, where $Id_2$ is the identity map on $I\!\!R^2$ and the $q_i$'s are the corresponding projection functions (i.e. for $(x_1,x_2)\in I\!\!R^2$, $q_i(x_1,x_2) = x_i$, $i=1,2$). So $X \hbox{ and } Y$ can be interpreted indifferently as random variables or as projections. In order for $P_1$ and $P_2$ to be the marginals of some probability measure $P$, and noting that $P(q_1^{-1} B_1) = P(B_1 \times I\!\!R)$ and $P(q_2^{-1} B_2) = P(I\!\!R \times B_2)$, the following consistency conditions are imposed:
for all $B_1, B_2 \in {\cal B}_1$. Of course the conditions in ((ref)) are not sufficient to fully determine $P$, a feature that was predictable since the link between the two experiments was not specified. In the particular case where the two experiments (and hence the two random variables $X$ and $Y$) are assumed to be independent, $P$ must satisfy the additional condition
for all $B_1,B_2 \in {\cal B}_1$. In other words, $P$ is the product probability $P_1\otimes P_2$, which is uniquely defined on ${\cal B}_2$. Moreover $P_1 = P_X$, $P_2 = P_Y$ and $P = P_{(X,Y)}$ in the above construction. The process just described consists of two steps that are worth distinguishing. \vskip .2cm First step: Two random experiments ${\cal E}_1$ and ${\cal E}_2$ are united to obtain a measurable space for the resulting joint experiment ${\cal E}$. \vskip .2cm Second step: A probability measure (and the resulting dependence structure between $X \hbox{ and } Y$) is enforced on the model set up in the first step.
Figuratively, one could say\footnote{without reference to the Marxist phraseology...} that the first step corresponds to an “infrastructure" on which a probabilistic “superstructure" is built in the second step. \vskip .2cm
One can illustrate, in terms of $\sigma$-fields, the qualitative leap following the coupling of two previously separate (stand-alone) random experiments ${\cal E}_1$ and ${\cal E}_2$. Under the above assumptions and the ensuing construction, the information that ${\cal E}_1$ and ${\cal E}_2$ provide is carried by the random variables $X:(I\!\!R^2,{\cal B}_2) \longrightarrow (I\!\!R,{\cal B}_1 )$ and $Y:(I\!\!R^2,{\cal B}_2) \longrightarrow (I\!\!R,{\cal B}_1 )$. The minimal $\sigma$-fields generated by $X$ and $Y$ are denoted by $\sigma (X)$ and $\sigma (Y)$, respectively. In turn, $\sigma (X)$ and $\sigma (Y)$ generate the $\sigma$-field
As by construction $X = q_1$ and $Y = q_2$, we have: $\sigma (X) = \{B_1 \times I\!\!R : B_1 \in {\cal B}_1 \}$ and $\sigma (Y) = \{I\!\!R \times B_2: B_2 \in {\cal B}_1 \}$. It can be shown without much difficulty that $\sigma (X) \vee \sigma (Y) = \sigma ((X,Y))$, the $\sigma$-field on $I\!\!R^2$ induced by the pair $(X,Y)$, noting that $\sigma ((X,Y)) = {\cal B}_2 = \sigma({\cal C})$, where ${\cal C} = \{B_1 \times B_2 : B_1, B_2 \in {\cal B}_1 \}$\footnote{That is, ${\cal B}_2$ is generated by the $(B_1 \times B_2)$'s, but ${\cal B}_2 = \sigma (X) \vee \sigma (Y)$ implies that ${\cal B}_2$ is also generated by the union of the $(B_1 \times I\!\!R)$'s and the $(I\!\!R \times B_2)$'s.}. Considering $X$ and $Y$ together as a pair of random variables instead of two stand-alone random variables allows to prepare a much wider portion of ${\cal B}_2$ (infrastructure) on which a probability measure (superstructure) can be defined. The representation $\sigma (X) \vee \sigma (Y)$, probably more telling than $\sigma ((X,Y))$, is symbolized in Figure (ref).
Incidentaly, note that the intersection $\sigma (X) \cap \sigma (Y)$ is not empty, since it contains $\emptyset$ and $I\!\!R^2$.
Errors of interpretation about the notion of distance between random variables occur because important concepts are not formally stated or not sufficiently explained. The purpose of this section is to eliminate ambiguities encountered here and there in applied science works published in journals or on the Internet. The role of a metric is to define a distance between elements of a same set. We mentioned in the introduction that to verify precisely whether a functional is adequate to cover what is commonly understood by the notion of distance, it is unavoidable to specify certain elementary, natural and intuitive rules -- axioms -- that this functional should satisfy. Again, we limit the discussion to real-valued random variables. Rachev (1991)\nocite{Rachev1991} provides a more general treatment.
The idea of distance between two random variables $X \hbox{ and } Y$ is linked in a decisive way to what is meant by $X \hbox{ and } Y$ "are the same", "are coincident", "are indistinguishable". And the concept of sameness, coincidence or indistinguishably -- treated here as synonyms -- can be mathematically captured by an equivalence relation: $X \hbox{ and } Y$ are declared to be the same, coincident or indistinguishable if they are in the same equivalence class. It is therefore natural to define a metric not on an initial set of random variables, but on the quotient space resulting from the adequate equivalence relation.
The theory of probability metrics considers three categories of metrics defined according to the type of equivalence deemed useful in a given context.
We assume throughout this section that real-valued random variables are defined on one and the same probability space $(\Omega,{\cal A},P)$. Denote by ${\cal V}$ the set of all random variables on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R, {\cal B}_1)$.\footnote{${\cal V} = {\cal L}_1(I\!\!R)$ when the random variables are assumed to be integrable.} So ${\cal V}^2 = {\cal V}\times {\cal V}$ is the set of all random pairs defined on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R^2, {\cal B}_2)$.
Let $\psi : {\cal V} \rightarrow I\!\!R$ be a functional candidate to be a distance. We would like $({\cal V},\psi )$ to be a metric space from the outset, but this wish is thwarted by the fact that we are working with random variables -- that is relatively complex mathematical objects. It is nevertheless possible to use $\psi$ to build a metric space defined on equivalence classes. For this purpose, let us define an equivalence relation on ${\cal V}$ by
What properties of $\psi$ will ensure that ((ref)) is a reflexive, symmetric and transitive relation, i.e. is an equivalence relation? It is easily verified that the following axioms meet this requirement: for all $X,Y,Z \in {\cal V}$, \newline (a) $\psi (X,X) = 0$ \hskip .3cm(reflexivity) \newline (b) $\psi (X,Y) = \psi (Y,X)$ \hskip .3cm(symmetry) \newline (c) $\psi (X,Z) \leq \psi (X,Y) + \psi (Y,Z)$ \hskip .3cm(triangle inequality), \newline noting that these axioms imply the non-negativity of $\psi$. So (a), (b) and (c), beyond their natural and intuitive content, are sufficient to ensure that $\stackrel{\psi}{\sim}$ is an equivalence relation. We call a functionnal satisfying (a), (b) and (c) a probability metric, although it is rigorously a semimetric\footnote{The terminology is still fluctuating: in topology, what we define here as a semimetric is called a pseudometric. }. Now, $X \stackrel{\psi}{\sim} Y$ means different things depending on the choice of $\psi$, because $\psi(X,Y) = 0$ may be true if and only if $X \hbox{ and } Y$ \newline (i) share a given set $c_1, c_2,\ldots$ of characteristics such that $c_i(X) = c_i(Y)$, $i= 1,2,\ldots$ (e.g. $X \hbox{ and } Y$ have the same mean and the same variance, as in ((ref)) below), or\newline (ii) have the same distribution, or \newline (iii) are almost surely equal.
So $\stackrel{\psi}{\sim}$ in ((ref)) is a generic notation for specific equivalence relations: \newline $X \stackrel{c}{=} Y$ when $X \hbox{ and } Y$ share a given set of characteristics, or \newline $X \stackrel{d}{=} Y$ when $X \hbox{ and } Y$ have the same distribution, or \newline $X \stackrel{a.s.}{=} Y$ when $X \hbox{ and } Y$ are almost surely equal.
Primary metrics correspond to the weakest form of equivalence. Random variables $X \hbox{ and } Y$ can be considered equivalent if they have the same mean and the same standard deviation. A plain example of primary metric is
where $\sigma$ refers to the standard deviation\footnote{Note that ((ref)) makes sense because the standard deviation and the mean are defined in the same unity. }.
The simple metrics imply a stronger form of sameness: $X \hbox{ and } Y$ are considered equivalent if their cdf's are identical (remembering that a random variable is completely described by its cdf). An example of simple metric is the Gini-Kantorovich metric GK given in ((ref)). GK measures the distance between $X \hbox{ and } Y$ -- which are assumed to have finite first moment -- by a distance between their respective cdf's.\footnote{If $X \hbox{ and } Y$have respective probability distributions $\mu$ and $\nu$, it is remarkable that $GK(X,Y)$ coincides with the 1-Wasserstein metric $W_1(\mu,\nu)$. It turns out that the latter is none other than the minimal cost in the Monge-Kantorowich transport problem (Kantorovich (1942)\nocite{Kantorovich1942}, see e.g. Villani (2008)\nocite{Villani2008} or Thorpe (2018)). Rachev (2007)\nocite{Rachev2007} uses the notation $GK(X,Y) = \min E|X - Y|$. This is very telling since (i) due to the fact that $|x - y|$ is a continuous cost function in the Monge-Kantorovich optimal transport problem, the infimum is realized (Gangbo (2004)\nocite{Gangbo2004}) and (ii) it remains us that $GK(X,Y)$ may be seen as the solution of this celebrated minimization problem. }
The compound metrics represent the strongest form of sameness. The simplest example of compound metric is probably the expected absolute difference $\psi (X,Y) = E|X - Y|$. Importantly, here $X \hbox{ and } Y$ are necessarily defined on the same probability space (i.e. are jointly distributed) and have finite first moment.
We will now show that there is a (true) metric $\tilde{\psi}$ derived from the semimetric $\psi$, where $\tilde{\psi}$ is defined on equivalence classes. The classes, denoted by $[\cdot]$, stem from the equivalence relation $\stackrel{\psi}{\sim}$, which can mean, as we have seen above, $\stackrel{c}{=}$, $\stackrel{d}{=}$ or $\stackrel{a.s.}{=}$. Define in a canonical way
We must show that $\tilde{\psi}$ is well-defined, i.e. does not depend on the representatives chosen to designate the classes. To show that ((ref)) makes sense, we need the following lemma.
Assume that $\psi$ satisfies reflexivity, symmetry and triangle inequality (and therefore nonnegativity, so that the conditions of Lemma (ref) are fulfilled) and let the equivalence relation on ${\cal V}$ be given by ((ref)). Suppose that $X_1\stackrel{\psi}{\sim}X$ and $Y_1\stackrel{\psi}{\sim}Y$. Then, using ((ref)) and the symmetry of $\psi$, we get $\psi (X_1,Y_1) = \psi (X,Y)$, that is $\tilde{\psi} ([X_1],[Y_1]) = \tilde{\psi} ([X],[Y])$, which proves that $\tilde{\psi}$ in ((ref)) is well-defined.
Let us write $\tilde{{\cal V}}$ for the quotient space ${\cal V}/\stackrel{\psi}{\sim} \ = \{[X]: X \in {\cal V} \}$. We are now able to state that $(\tilde{{\cal V}},\tilde{\psi})$ is a metric space, i.e. that $\tilde{\psi}$ satisfies the following axioms:
The first two axioms -- known as the identity of indiscernibles when taken together -- are the consequence of the definitions of $\stackrel{\psi}{\sim}$ and $\tilde{\psi}$, while axioms 3 and 4 stem directly from the symmetry and triangle inequality property of $\psi$.
Note that if ${\cal V} = {\cal L}_1(I\!\!R)$ and $\psi (X,Y) = E|X-Y|$, then $\stackrel{\psi}{\sim}$ means $\stackrel{a.s.}{=}$. In this case $\tilde{{\cal V}} = L_1(I\!\!R)$ and $\tilde{\psi}$ is the metric induced by the $L_1$-norm.
We limit our discussion to $E|\cdot -\cdot |$, which implies that ${\cal V} = {\cal L}_1(I\!\!R)$ (${\cal V}$ defined in Section (ref)), but our conclusions can be generalized to other compound metrics.
Focusing on the two equivalence relations $\stackrel{a.s.}{=}$ and $\stackrel{d}{=}$ which group variables belonging to ${\cal L}_1(I\!\!R)$, we denote by $[\cdot]_{a.s.}$ and $[\cdot]_{d}$, the corresponding equivalence classes. Why are we so eager to identify certain elements of ${\cal L}_1(I\!\!R)$? It is because it allows $E|X-Y|$ to switch from a semimetric to a metric. Indeed, if we define $\tilde{\psi} ([X]_{a.s.},[Y]_{a.s.}) = \psi (X,Y) = E|X-Y|$, then $E|\cdot - \cdot|$ represents both $\tilde{\psi}$ (a metric defined on classes) and $\psi$ (a semimetric defined on random variables). We saw in Section (ref) that $\tilde{\psi}$ realizes the identity of indiscernibles (reflexivity and reverse reflexivity), whereas $\psi$ satisfies reflexivity ($\psi (X,X) = 0$), but not reverse reflexivity ($\psi (X,Y) = 0$ does not imply $X = Y$). Moreover, $\psi$ and $\tilde{\psi}$ both satisfy symmetry and triangle inequality.
That said, some authors using $E|X-Y|$ have fallen into the trap of identifying within ${\cal L}_1(I\!\!R)$ the identically distributed random variables rather than the almost surely equal random variables. Unfortunately, this leads to a logical impasse. Seeking a contradiction, suppose that we set
where $X,Y \in {\cal L}_1(I\!\!R)$ are identically distributed without being almost surely equal. We have $[X]_d = [Y]_d$ and $E|X-Y| > 0$ (since $E|X-Y| = 0 \Leftrightarrow X\stackrel{a.s.}{=} Y$), and we end up with the following contradiction: \newline $\tilde{\psi} ([X]_d,[X]_d) \stackrel{(\ref{interdit})}{=} E|X-X|=0 < E|X-Y| \stackrel{(\ref{interdit})}{=} \tilde{\psi} ([X]_d,[Y]_d) = \tilde{\psi} ([X]_d,[X]_d)$, meaning that ((ref)) is ill-defined.
To convince ourselves that the above discussion is not in vain, take the case of the so-called Lukaszyk-Karmowski metric (Lukaszyk 2004)\nocite{Lukaszyk2004} which is actually the functional $\tilde{\psi} ([\cdot]_d,[\cdot]_d) = E|\cdot - \cdot|$ set in ((ref)). It is only when this author asserts that the identity of indiscernibles property is not realized by the metric $E|\cdot - \cdot|$ he uses that we end up understanding he implicitly identifies the identically distributed random variables of ${\cal L}_1(I\!\!R)$. In other words, he reasons as if ((ref)) were well-defined\footnote{ In later publications of Lukaszyk (and on Wikipedia, etc.), where $D(X,Y)$ refers to $E|X-Y|$, we find the notation $D(X,X) > 0$, which proves that the identically distributed variables of ${\cal L}_1(I\!\!R)$ are implicitly identified. This leads to the logical contradiction we have put forward. }. On this erroneous basis, Lukaszyk claims to have used a new operator which, as such, would deserve a special denomination. This does not make sense and the so-called Lukaszyk-Karmowski metric is none other than the good old $L_1$-norm distance. Note that this clarification does not greatly affect the merit of this author's 2004 article, where otherwise conclusive results in applied physics are presented. Lukaszyk correctly computes the distance $D(X,Y) = E|X-Y|$ between independent elements of ${\cal L}_1(I\!\!R)$ -- notably when $X \hbox{ and } Y$ are Gaussian -- but he should not pretend that the reflexivity condition does not hold.
Why is the notion of independence so important? This section contains a few remainders and general thoughts about independence. Two random variables are independent if they come from phenomena such that the result observed for one of them has no influence on the other. We should admit that the assumption of independence is often a question of intuition or common sense -- although in this respect caution is needed. When the hypothesis of independence can be made reasonably, great mathematical simplifications follow (resulting in particular from the use of the Fubini-Tonelli theorem).
Contrary to a common belief -- at least among researchers having somewhat forgotten the fundamentals of probability theory -- the notion of independence between two random variables $X$ and $Y$ does not imply that they “have nothing to do with each other”. Indeed, independence only makes sense if these random variables are coupled, i.e. defined on the same probability space. They are coupled to each other, albeit in a particular way, by the fact that they are independent.
In applied sciences, two observations, phenomena or experiments can be perceived as independent. Independence may be imposed from the outside or may be organized in full awareness by the experimenter. In both cases, he or she will be interested in forming a probability model for a joint experiment such that the original two experiments are carried out independently.
Calculating a distance between independent random variables turns out to be often very useful. For example two measurement devices $X$ and $Y$ independently measure unknown quantities with some random error, or multiple researchers independently measure the same object and compare their results. To take an example related to economic inequalities, suppose that $X$ (resp. $ Y $) is the income of a household drawn at random from a statistical population ${\cal P} _1$ (resp. ${\cal P} _2$). Let $\mu \hbox{ and } \nu$ denote the income distribution of $X \hbox{ and } Y$, respectively. The random variables $X \hbox{ and } Y$ are assumed to be independent, and we use this information to compute $E|X-Y|$, which is interpreted as a measure of income disparity between the two populations. In other words, $X=x$ and $Y=y$ are independently observed and $E|X-Y|$ is the weighted average of the $|x-y|$'s. We have already mentioned that a metric such as $GK(X,Y)$, which depends only on the stand-alone distributions of $X$ and $Y$, cannot take account of the independence information.\footnote{We will see in Section (ref) that in case of a single population, i.e. if ${\cal P}_1 = {\cal P} _2$ and $\mu = \nu$, and if the independent random variables $X$ and $Y$ are non-negative, then $E|X-Y|$ is none other than the Gini mean difference, a measure of income inequality within a population. The normalized Gini mean difference is the celebrated Gini index.}
Consider $(X,Y) \in {\cal L}_1(I\!\!R^2)$, the space of all pairs of integrable real-valued random variables defined on a probability space $(\Omega,{\cal A},P)$. It is well-known that $X\stackrel{a.s.}{=} Y$ ($X = Y$ almost surely) if and only if $E|X-Y| =0$. A pair $(X,Y) \in {\cal L}_1(I\!\!R^2)$ with $E|X-Y| =0$ is such that the probability mass $P_{(X,Y)}$ is concentrated on the diagonal $\Delta:=\{(x,x):x\inI\!\!R\}$ of $I\!\!R^2$, (that $P_{(X,Y)}(\Delta ) = 1$ is formalized in Subsection (ref), Proposition (ref)).
Here are some remarks about the values $E|X-Y|$ can take: if the distribution of the random variables $X$ and $Y$ differ, then $X$ and $Y$ cannot be almost surely equal, and consequently $E|X-Y| > 0$. Moreover, the fact that $X$ and $Y$ have the same distribution by no means implies that $E|X-Y| =0$: suppose that $X$ and $Y$ are two continuous independent and identically distributed (i.i.d.) random variables. In this case $P({X\neq Y}) = 1$, i.e. $X \neq Y$ a.s., which implies that $E|X-Y| > 0$. For example, take two i.i.d. random variables $X \hbox{ and } Y$ having standard normal distribution. Then ((ref)) -- a consequence of Theorem (ref) -- implies that $E|X-Y| = 2/\sqrt{\pi}$.
It is interesting to take a closer look at the behavior of $E|X-Y|$ when $X$ and $Y$ are independent.
A random variable is said to be degenerate if it is almost surely constant. So, if $X \hbox{ and } Y$ are independent, then $E|X-Y| = 0$ occurs if and only if they are (identically distributed and) degenerate, i.e. if and only if their distribution in concentrated on the same constant.
The above considerations can be summarized as follows: \newline \[ E|X-Y| = 0 \Rightarrow \left\{
\right. \] a situation illustrated in Figure (ref).
We are interested in random pairs $(X,Y) \in {\cal L}_1(I\!\!R^2 )$. In order to bring together and clarify the concepts encountered in Section (ref), define \newline ${\cal E}_{a.s} = \{ (X,Y) \in {\cal L}_1(I\!\!R^2 ) : X \stackrel{a.s.}{=} Y \}$ and ${\cal E}_{d} = \{ (X,Y) \in {\cal L}_1(I\!\!R^2 ) : X\stackrel{d}{=} Y \}$. That is, the subsets ${\cal E}_{a.s}$ and ${\cal E}_{d} $ of ${\cal L}_1(I\!\!R^2 )$ contain the random pairs whose components are equivalent in the almost sure sense and in the equality of distribution sense, respectively. Taking into account a possible independence between $X \hbox{ and } Y$, the pairs $(X,Y) \in {\cal L}_1(I\!\!R^2 )$ fall into six mutually exclusive categories described in Table (ref).
Figure (ref) depicts the six categories of Table (ref) as a partition of ${\cal L}_1(I\!\!R^2 )$.
It would be incomplete to talk about independence without giving a brief overview of the concept of entropy, whose definition comes from the pioneers of information theory. The entropy of a distribution was established by Shannon for discrete laws and then extended to continuous laws characterized by their density. The notion of entropy has been generalized in many ways and has undergone vast and important developments. The real-valued random variables $X$ and $Y$ defined on $(\Omega,{\cal A},P)$, where $\Omega$ is a finite set, are independent when the degree of uncertainty is maximal in their joint distribution. This means that the entropy of the bivariate distribution must be maximal.
Entropy is a measure of uncertainty or randomness. The following equation holds: \newline $H(X,Y) = H(X) + H(Y) - I(X,Y) $, where $H(X,Y)$ is the (joint) entropy of the pair $(X,Y)$, $H(X)$ and $H(Y)$ being the entropy of $X$ and $Y$, respectively. The non-negative $I(X,Y)$ is the so-called mutual information, and one has $I(X,Y) = 0$ if and only if $X$ and $Y$ are independent, in which case $H(X,Y)$ is maximal. In the finite case that we are dealing with $H(X) = \sum_i p_i \ln(1/p_i)$, $H(Y) = \sum_j p_j \ln(1/p_j)$ and $H(X,Y) = \sum_i \sum_j p_{ij} \ln(1/p_{ij})$, where $p_i = P(X=x_i)$, $p_j = P(Y=y_j)$ and $p_{ij} = P(X=x_i, Y=y_j)$.
Example (ref) below illustrates that the distance $E|X-Y|$ between dependent $X$ and $Y$ can be both smaller or larger than it is when $X$ and $Y$ are independent.
This section focuses on the normalized form of the expected absolute difference and its characteristics. In particular, we address the issue of triangle inequality in relation to this functional. In an interesting paper, Yianilos (2002)\nocite{Yianilos2002} shows that symmetric set difference and Euclidean distance on $I\!\!R^d$ have normalized forms that remain metrics. We examine in this section the conditions under which $E|X-Y|$ can be $[0,1]$-normalized. The additional difficulty here is that we are not in a deterministic context. With our usual notation, ${\cal L}_1(I\!\!R)$ is the vector space of all integrable random variables on $(\Omega,{\cal A},P)$ taking values in $(I\!\!R, {\cal B}_1)$, and the random variables used below belong to this set.
It is natural to consider forming relative distance measures. Converting $D(X,Y) := E|X-Y|$ to a normalized (or standardized) form may be very useful in the solution of certain problems, especially when relative errors are at stake, as is often the case in numerical analysis. Define a normalized counterpart $D_{norm}(\cdot,\cdot )$ of $D(\cdot,\cdot )$ by
so that $0\le D_{norm}(X,Y) \le 1$. The upper bound is reached when, say, $Y\stackrel{a.s.}{=}0$ with $E(|X|) > 0$, while the lower bound is reached when $X$ and $Y$ are almost surely equal and $E(|X|)$ $(= E(|Y|))$ is strictly positive. It is not our intention to comment here on the pros and cons of using a normalized distance measure. We will simply check whether the generally accepted axioms for a distance are verified. What we can say though is that $D_{norm}$ is likely to share the strengths and weaknesses of relative deterministic measures such as the Canberra metric (Lance and Williams (1967))\nocite{Lance1967}. Clearly, $D_{norm}$ is a distance since it satisfies nonnegativity, reflexivity ($D_{norm}(X,X) = 0$) and symmetry ($D_{norm}(X,Y) = D_{norm}(Y,X)$). Let $(X_i)_{i \in I} \subset {\cal L}_1(I\!\!R)$ denote a finite family of independent random variables. We show in Subsection (ref) that the triangle inequality holds in case of independence, but does not hold in general. More precisely, while $({\cal L}_1(I\!\!R),D_{norm})$ is a distance space, we show that $((X_i)_{i \in I},D_{norm})$ is a semimetric space (i.e. $D_{norm}$ is a distance satisfying the triangle inequality).
A corollary of Theorem (ref) below is that $D_{norm}(\cdot ,\cdot )$ defined in ((ref)) satisfies the triangle inequality in the independence case.
Proving this inequality was a particularly thorny exercise (see the proof in the appendix). Using ((ref)) and the definition of $D_{norm}$ in ((ref)), one can easily check that
for mutually independent $X,Y,Z$. Note that independence of these variables implies that the three pairs $(X,Z)$, $(X,Y)$ and $(Y,Z)$ intervening in the triangle inequality are the two-dimensional projections of the three-dimensional random vector $(X,Y,Z)$ having the product distribution $P_X \otimes P_Y \otimes P_Z$. In other words, the three pairs can be consistently embedded in a three-dimensional vector so that the triangle inequality makes sense (more details on this are given below).
A question naturally arises: would the triangle inequality ((ref)) hold in all cases if the assumption of independence were lifted in Theorem (ref) ? The answer is negative: with a little of craftmanship, one can find counterexamples such as the one resulting from the three bivariate distributions (A), (B), (C) of pairs of dependent random variables shown in Table (ref).
It turns out that $E|X| = E|Y| = E|Z| = 1$ and
which contradicts ((ref)). The statement that the triangle inequality should hold for any $X,Y,Z$ is actually pretty vague. As $E|X-Y|$ is a compound probability metric (see Definition (ref)), the choice of the three pairs $(X,Z)$, $(X,Y)$ and $(Y,Z)$ cannot be totally free. Indeed, suppose we fix the distributions (A) and (B) in Table (ref). Then the choice of distribution (C) cannot be arbitrary, because the (internal) dependence structure of $(Y,Z)$ depends on the dependence structures of $(X,Y)$ and $(X,Z)$. Rachev et al. (2007, p. 93) \nocite{Rachev2007} give the following consistency rule: ”The three pairs of random variables $(X,Z)$, $(X,Y)$ and $(Y,Z)$ should be chosen in such a way that there exists a consistent three-dimensional random vector $(X,Y,Z)$ and the three pairs are its two-dimensional projections.” In other words, if this rule is respected, then the three pairs can be safely embedded in a three-dimensional random vector. To validate our counterexample, we must make sure that the distributions (A), (B), \textbf{(C)} abide by the consistency rule. This is indeed the case, for the matrix
is positive definite (with eigenvalues $\lambda_1 = 0.072$, $\lambda_2 = 0.44$ and $\lambda_3 = 2.128$), which means that $V$ is a valid covariance matrix.
For completeness, we give what in appearance only is a counterexample. The distributions (D), (E), (F) in Table (ref) are such that the triangle inequality does not hold because \newline $(E|X-Y| + E|Y-Z| - E|X-Z|)/2 = -0.2$. However, (D), (E), (F) do not constitute a valid counterexample, because the matrix $V$ in ((ref)) resulting from these distributions is indefinite (with eigenvalues $\lambda_1 = -0.076$, $\lambda_2 = 1.093$ and $\lambda_3 = 1.623$), and thus simply cannot be a covariance matrix of a three-dimensional vector.
We end this subsection with an example illustrating the importance of the independence assumption in Theorem (ref). In Table (ref), $(X,Z)$, $(X,Y)$ and $(Y,Z)$ are now three pairs of independent variables with the same marginal distributions as those shown in Table (ref).
From the distributions (G), (H), (I), we find that $(E|X-Y| + E|Y-Z| - E|X-Z|)/2 = 0.5$ (instead of -0.02 in ((ref))), which means that the triangle inequality holds. Note also that $V$ is the three-dimensional diagonal matrix diag(0.96,0.84,0.84), which is of course a valid covariance matrix.
The computation of a distance between identically distributed random variables of ${\cal L}_1(I\!\!R)$ is helpful in various domains. Examples are the Gini mean difference (GMD)\footnote{The GMD is twice the $L$-scale (the second $L$-moment). It is sometimes considered as a competitor of the standard deviation. } and the Gini index, used in particular in inequality economics to measure the amount of inequality included in a distribution of income (alternatively consumption or wealth, etc.). Let $\mu$ be such a distribution. The GMD and the Gini index are defined as GMD$(\mu) = E|X-Y|$ ($= D(X,Y))$ and $Gini(\mu) = E|X-Y|/[2E(X)]$ ($= D_{norm}(X,Y)$), where $X \sim \mu$ and $Y \sim \mu$ are assumed to be independent and (usually) non-negative (see e.g. Yitzhaki (1998)\nocite{Yitzhaki1998}, Yitzhaki and Schechtman (2013)\nocite{YitzhakiSchechtman2013}, or Xu (2003)\nocite{Xu2003})\footnote{When the range of $X$ encroaches on $(-\infty ,0]$, we know from ((ref)) that there is no mathematical reason why we should not define $Gini(\mu) = E|X-Y|/[2E|X|]$.}. Looking at ((ref)), we can say that the Gini index is the distance (semimetric) $D_{norm}$ from $X$ to an i.i.d. “copy” of itself. In that sense, it is sometimes called an “autodistance” in the literature. However, a copy must be clearly distinguished from the original and there is some confusion on this point. “Copy” is to be understood here in the equality of distribution sense (in the almost sure equality sense, the distance is trivially zero).
Independence of $X \hbox{ and } Y$ allows in many cases to use Fubini-Tonelli to represent the GMD and the Gini index in closed form. Independence also implies that, except in the degenerate case where $X\stackrel{a.s.}{=} Y\stackrel{a.s.}{=}c$ for some $c \in I\!\!R_+$, $GMD(\mu) > 0$. Although seemingly simple, the Gini index is actually a quite proteiform measure of inequality. It can be expressed in an astonishing number of ways, some of which can be found in Yitzhaky (1998)\nocite{Yitzhaki1998}.
We end this subsection by probabilistic considerations on the values the Gini index can take. Let $X\sim \mu$ and $Y\sim \mu$ be two i.i.d. non-negative random variables where $\mu$ is (say) the income distribution of a population. Then $Gini(\mu) = E|X-Y|/[2E(X)]$ is not defined if and only if $X\stackrel{a.s.}{=}0$. Leaving this uninteresting case aside, $Gini(\mu)= 0 \Leftrightarrow E|X-Y| = 0 \stackrel{Prop.\ref{equivalence}}{\Longleftrightarrow} \exists c > 0$ such that $X\stackrel{a.s.}{=} Y\stackrel{a.s.}{=}c$. Moreover, the triangle inequality $E|X-Y| \leq |X| + |Y|$ implies that $Gini(\mu) \in [0,1]$.
In economic applications, non-negative real numbers $x_1,x_2,\ldots ,x_n$ are typically incomes earned respectively by individuals $i_1,i_2,\ldots , i_n$ belonging to a population or a statistical sample. It is known that $0\le g \leq (n-1)/n$, where $g$ is the Gini index. A low value of $g$ corresponds to a more equal income distribution, with 0 indicating perfect equality (all individuals have the same income). The most unequal society is the one where a single individual receives all the income and the remaining individuals receive nothing. In that case, $g = (n-1)/n$. Generalizing a bit, we can say that a tiny proportion $\epsilon > 0$ of the population receives an income $b>0$, while a proportion $1 - \epsilon$ gets nothing. Such an extreme case can be described thanks to i.i.d. random variables $X \hbox{ and } Y$ having distribution $\mu = (1 - \epsilon) \delta_0 + \epsilon \delta_b$ where $b$ is a strictly positive income level and where $\epsilon > 0$, $\epsilon$ small, is the probability that $X$ (or $Y$) takes the value $b$. The distribution of the pair $(X,Y)$ is summarized in Table (ref).
The Gini index then becomes $Gini(\mu) = E|X-Y|/[2E(X)] = 1 - \epsilon$. The upper bound 1 corresponding to $\epsilon = 0$ cannot be reached, because in such a case $X$ would take the value 0 with probability 1 , and the Gini index would not be defined ($E(X)$ would be zero).
For $p \ge 1$, let ${\cal L}_p (I\!\!R)$ denote the class of all real-valued random variables on a probability space $(\Omega,{\cal A},P)$ that have finite $p$-th moment, and consider $X,Y\in {\cal L}_p (I\!\!R)$. Suppose that $X \hbox{ and } Y$ are identically distributed. In this section, we show in particular that the ${\cal L}_p$-distance $(E|X -Y|^p)^{1/p}$ certainly represents a distance between $X \hbox{ and } Y$, but also -- in a sense to be specified -- tells us how far from almost sure equality these variables are. In a symbolic way, we will show that $(E|X -Y|^p)^{1/p}$ can be conceived as a distance $Dist(X \stackrel{d}{=} Y,X \stackrel{a.s.}{=} Y)$ between two possible states of the pair $(X,Y)$. In (ref), we introduce the so-called diagonal coupling of a probability measure with itself, and we clarify its connection to almost sure equality. In (ref), we define a distance\footnote{in fact a semimetric, i.e. a distance satisfying the triangle inequality.} $\eta_p(\pi_1, \pi_2)$ between bidimensional probability measures $\pi_1$ and $\pi_2$ having finite $p$-th moment. In (ref), we recall the notion of coupling and observe what becomes of $\eta_p$ when $\pi_1$ and $\pi_2$ are couplings of unidimensional probability measures $\mu$ and $\nu$ of finite $p$-th moment. Equating $\mu$ and $\nu$ and taking $\pi_2$ as the diagonal coupling of $\mu$ with itself allows then to interpret $\eta_p$ as a distance indicating how far $X\sim \mu$ and $Y\sim \mu$ are from almost sure equality. The results are illustrated with the bivariate normal distribution. Finally, thanks to $\eta_p$, we propose a probabilistic representation of the Gini mean difference and of the Gini index.
Consider a probability space $(I\!\!R, {\cal B}_1,\mu)$. We denote by $\mu \triangle\mu$ a probability measure on the product space $(I\!\!R^2, {\cal B}_2)$ defined by
Then $\mu \triangle\mu$ can be extended to the whole ${\cal B}_2$ by using Carath\'eodory's theorem (Kalikow and McCutchean (2010) \nocite{Kalikow2010}).
Actually, $(I\!\!R, {\cal B}_1,\mu )$ and $(I\!\!R^2, {\cal B}_2,\mu \triangle \mu )$ are closely related: the map $s:I\!\!R \longrightarrow I\!\!R^2$ defined by $s(x) = (x,x)$ is a measurable isomorphism from $(I\!\!R, {\cal B}_1,\mu )$ to $(I\!\!R^2, {\cal B}_2,\mu \triangle\mu )$. An example of diagonal coupling is given in Table (ref) (A), where $\mu$ is defined by $\mu (0) = 0.3$ and $\mu (1) = 0.7$). Incidentally, Proposition (ref) tells us that if $X\stackrel{a.s.}{=} Y$, with $X \hbox{ and } Y$ independent, then there exists $c \in I\!\!R$ such that $X\stackrel{a.s.}{=} c$ and $Y\stackrel{a.s.}{=} c$, which implies that $\mu \otimes \mu = \mu \triangle \mu = \delta_{(c,c)}$ (the Dirac delta measure concentrated on $(c,c)$ ), where $\mu \otimes \mu$ is the product measure of $\mu$ with itself. So the probability measures $\mu \otimes \mu$ and $\mu \triangle \mu$ on $(I\!\!R^2,{\cal B}_2)$ coincide in this trivial case.
Let $\Delta$ denote the diagonal $\{(x,x):x\inI\!\!R\}$ of $I\!\!R^2$. One can show quite easily that $B_1\cap B_2 = s^{-1}((B_1 \times B_2)\cap \Delta)$ for any subsets $B_1,B_2$ of $I\!\!R$. Therefore, the definition of the diagonal coupling in ((ref)) may be replaced by $(\mu \triangle\mu)(B_1\times B_2) = \mu(s^{-1}((B_1 \times B_2)\cap \Delta))$, which is visually more telling (see Figure (ref)). Let $X,Y:(\Omega,{\cal A},P) \rightarrow (I\!\!R, {\cal B}_1)$ (e.g. $X,Y \in {\cal L}_p (I\!\!R)$ if $X \hbox{ and } Y$ have finite $p$-th moment) be two identically distributed random variables. Then the diagonal coupling of $P_X$ with itself is the distribution of the pair $(X,Y)$ if and only if $X\stackrel{a.s.}{=} Y$:\footnote{By definition $X\stackrel{a.s.}{=} Y$ makes sense only if $X \hbox{ and } Y$ are defined on the same probability space. } the entire probability mass of $P_{(X,Y)}$ is concentrated on the diagonal $\Delta$. This result is in line with our intuition. It is formally stated in Proposition (ref).
Lemma (ref) below -- added for completeness -- is not directly related to what we need in this article. However, we would like to answer the following question: how can we construct a probability space on $\Delta$ consistent with $(I\!\!R^2, {\cal B}_2, \mu \triangle \mu)$? Two ways come to mind: (i) take the trace space of $(I\!\!R^2, {\cal B}_2, \mu \triangle \mu)$ with respect to $\Delta$, or (ii) take the pushforward space of $(I\!\!R, {\cal B}_1, \mu )$ induced by $s$ ($s$ defined above). It turns out that the two methods produce the same space, as evidenced by Lemma (ref).
Let $(X_1,Y_1)$ (resp. $(X_2,Y_2)$) be two pairs of random variables having joint distribution $\pi_1$ (resp. $\pi_2$) on $(I\!\!R^2, {\cal B}_2)$. What is meant by $\pi_1$ and $\pi_2$ being “close to each other”? A possible answer -- serving what we wish to show in this subsection -- is to measure their proximity by using a primary metric, i.e. to consider that $\pi_1$ and $\pi_2$ coincide when they share a given set of relevant characteristics. Accordingly, we consider here that the distance between $\pi_1$ and $\pi_2$ is zero if (i) the centers (mathematical expectations, means) of $(X_1,Y_1)$ and $(X_2,Y_2)$ are the same and (ii) the deviation between $X_1$ and $Y_1$ is the same as the deviation between $X_2$ and $Y_2$.
For $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2)\sim \pi_2$, assume that the marginals of $\pi_1$ and $\pi_2$ have finite $p$-th moment, $p\in [1,\infty )$. For $i=1,2$, we will use the following notations: \newline $E|X_i -Y_i|^p = \int_{I\!\!R^2} |x_i -y_i|^p d\pi_i(x_i,y_i)$, $EX_i = \int_{I\!\!R^2} x_i d\pi_i(x_i,y_i)$, $EY_i = \int_{I\!\!R^2} y_i d\pi_i(x_i,y_i)$ and $C(\pi_i) = E([X_i, Y_i]) = [EX_i, EY_i]$. We can now define
where $||\cdot||$ is the Euclidean norm on $I\!\!R^2$. So $\eta_p$, which takes finite values, integrates two sources of deviation: between pairs and within pairs of random variables. Obviously, the characteristics entering the definition of $\eta_p$ in ((ref)) do not fully describe what differentiates $\pi_1$ from $\pi_2$.
The (easy) proof of Proposition (ref) is omitted. As $\eta_p$ is in particular non-negative, reflexive and symmetric, it can be called a distance between elements of ${\cal P}_p(I\!\!R^2)$.\footnote{Formally, as already mentioned, the difference between a semimetric and a distance is the relaxation of the triangle inequality.} Moreover, $\eta_p$ is a semimetric without being a metric\footnote{As a semimetric, $\eta_p$ can be transformed into a metric between equivalence classes. Define an equivalence relation between the elements of ${\cal P}_p(I\!\!R^2)$ by $\pi_1 \sim \pi_2 :\Leftrightarrow \eta_p(\pi_1, \pi_2) = 0$. Then $\tilde{\eta_p}([\pi_1], [\pi_2]) := \eta_p(\pi_1, \pi_2)$ is a metric on the set $\left\{ [\pi] : \pi \in {\cal P}_p(I\!\!R^2) \right\}$ of classes. One can show that equivalent elements are equidistant from any other element of ${\cal P}_p(I\!\!R^2)$. }. Indeed, $\eta_p$ does not satisfy the reverse reflexivity condition $\eta_p(\pi_1, \pi_2) = 0 \Rightarrow \pi_1 = \pi_2$. Example (ref) shows us that there exists probability measures $\pi_1,\pi_2 \in {\cal P}_p(I\!\!R^2)$ such that $\pi_1 \neq \pi_2$ while $\eta_p(\pi_1, \pi_2) =0$.
We now need the general definition of coupling, of which the diagonal coupling (Definition (ref)) is a special case.
By definition, couplings are multiple. The class of all couplings between $\mu \hbox{ and } \nu$ is denoted by $\Pi (\mu ,\nu)$. For example, the distributions $\pi_1$ and $\pi_2$ shown in Table (ref) belong to the set $\Pi (\mu,\mu)$, i.e. $\nu = \mu$ in this case.
We observe that the law $\tilde{\pi}$ of $(\tilde{X},\tilde{Y})$ is a coupling of the laws $\mu$ of $X$ and $\nu$ of $Y$. An important point of Definition (ref) is that the coupled random variables are defined on the same probability space, while $X \hbox{ and } Y$ may not be defined on a common probability space. If $X \hbox{ and } Y$ are defined on the same probability space $(\Omega,{\cal A},P)$, then $(X,Y)$ is also defined on $(\Omega,{\cal A},P)$ and $P_{(X,Y)}$ is a coupling of $P_X$ and $P_Y$. Two trivial couplings are (i) the diagonal coupling $\mu \triangle\mu$ of $\mu$ with itself defined in ((ref)), and (ii) the product coupling $\mu \otimes \nu$. If $X \sim \mu$ and $Y \sim \nu$ are independent, then the law of $(X,Y)$ is the product probability measure $\mu \otimes \nu$. Table (ref) (D) and (E) are examples of product couplings.
Let $\Pi_p (\mu,\nu)$ denote the set of couplings of two probability measures $\mu \hbox{ and } \nu$ of finite $p$-th moment, both defined on $(I\!\!R, {\cal B}_1)$. For $\pi_1,\pi_2 \in \Pi_p (\mu,\nu)$, suppose that $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2)\sim \pi_2$. As couplings of $\mu$ and $\nu$ have identical centers, $||C(\pi_1) - C(\pi_2)|| $ disappears in ((ref)) and we have
Let us now consider the case where $\nu = \mu$, when $\mu \hbox{ and } \nu$ have finite $p$-th moment. Assume that $(X_1,Y_1)\sim \pi_1$ and $(X_2,Y_2) \sim \mu \triangle \mu$. Note that $\mu \triangle \mu \in \Pi_p (\mu,\mu)$, and assume that $\pi_1 \in \Pi_p (\mu,\mu)$. For ease of reading, write $(X,Y) \sim \pi$ instead of $(X_1,Y_1)\sim \pi_1$. Clearly, $X,Y,X_2$ and $Y_2$ have all the same distribution $\mu$. We wish to measure the distance $\eta_p$ between $\pi$ and $\mu \triangle \mu$. As $X_2\stackrel{a.s.}{=} Y_2$, we have $E|X_2-Y_2|^p=0$ and ((ref)) becomes
That is, the ${\cal L}_p$-distance $(E|X -Y|^p)^{1/p}$ between identically distributed $X \hbox{ and } Y$ also represents a distance between the law $\pi$ of $(X,Y)$ and the diagonal coupling built from the marginals of $\pi$. Note that both the ${\cal L}_p$-distance and $\eta_p$ are semimetrics. \vskip .2cm Distance between equality of distribution and almost sure equality \vskip .2cm Suppose that $X$, $Y$ (and therefore $(X,Y)$) are defined on some probability space $(\Omega,{\cal A},P)$. Moreover, suppose that $X \hbox{ and } Y$ are identically distributed. The following question comes to mind: how far from almost sure equality are $X \hbox{ and } Y$? Writing $\pi = P_{(X,Y)}$ and $\mu = P_X$ in ((ref)), we obtain
That is, when jointly distributed random variables $X$ and $Y$ are identically distributed, $||X-Y||_p$ is actually a distance between the distribution of the pair $(X,Y)$ and the diagonal coupling of $P_X$ with itself. It is in this sense that $||X-Y||_p$ may be symbolically interpreted as a distance between $X\stackrel{d}{=} Y$ and $X\stackrel{a.s.}{=} Y$. \vskip .2cm Illustration of the above developments: the case of bivariate normal distributions \vskip .2cm Consider the bivariate normal density function
with parameters $m_X$, $m_Y$ (marginal means), $\sigma_X > 0$, $\sigma_Y > 0$ (marginal standard deviations) and $|\rho| < 1$ (correlation coefficient). Let $(X,Y) \sim \pi$, where $\pi$ is characterized by the density function ((ref)). The marginals $\mu \hbox{ and } \nu$ of $\pi$ are $\mu = N(m_X,\sigma^2_X)$ and $\nu = N(m_Y,\sigma^2_Y)$. Importantly, they do not depend on $\rho$, which means that each bivariate normal distribution $\pi$ is a coupling of $\mu \hbox{ and } \nu$. An infinite number of couplings can be created by just changing the value of $\rho$.
When $m_X = m_Y$ and $\sigma_X = \sigma_Y$, i.e. when $X \hbox{ and } Y$ are identically distributed, then $\mu \triangle \mu$ is the diagonal coupling of $N(m_X,\sigma^2_X)$ with itself, which is supported on the diagonal $\Delta$ of $I\!\!R^2$. Setting for example $p=1$ in ((ref)), we obtain
from which we draw the conclusion that the smaller the value of $E|X-Y|$, the more the graph of $f_{XY}(x,y)$ is concentrated along the diagonal $\Delta$.
Figure (ref) graphs the densities of three bivariate normal distributions, all three having parameters $m_X = m_Y =0$ and $\sigma_X = \sigma_Y =1$. Consequently, the two marginal laws of each of these distributions are the standard normal $N(0,1)$. The only difference between the three bivariate distributions of Figure (ref) is the value of $\rho$. The distribution on the left, where $\rho = 0$, corresponds to i.i.d. standard normal random variables $X \hbox{ and } Y$. Looking at the distribution on the right, where $\rho = 0.99$, we can guess the bell-shaped form the distribution will have when $\rho = 1$. In this case $\mu \triangle \mu = N(0,1) \triangle N(0,1)$, defined on $I\!\!R^2$, concentrates its probability mass on $\Delta$.
It is interesting to visualize towards which distribution the densities of Figure (ref) converge when $\rho\nearrow 1$ and $\Delta$ is seen as a standalone line of numbers. In this case the manifold $\Delta$ differs from $I\!\!R$ only by its coordinate system: $I\!\!R$ is “stretched” to obtain $\Delta$ and $\Delta$ can be seen as $I\!\!R$ with a new coordinate system given by $y=y(x)=\sqrt{2}x$. Let $f$ (resp. $g$) be the pdf representing $\mu = N(0,1)$ in the old (resp. new) coordinate system. Then, for any $B\in {\cal B}_1$, $\int_B f(x)dx = \int_B g(y)dy$, where, using the Jacobian rule for a change of coordinates, $g(y) = |\frac{dx}{dy}| f(x) = \frac{1}{\sqrt{2}} f (\frac{y}{\sqrt{2}} )$, which is the density of the distribution $N(0,2)$. That is, we have shown that the bivariate normal density represented on the right of Figure (ref) is close to the density $N(0,2)$ defined on $\Delta$ when $\Delta$ is considered as a simple line of numbers.
If, in addition to being identically distributed, two jointly distributed random variables $X \hbox{ and } Y$ are assumed to be independent (i.e. if $X \hbox{ and } Y$ are i.i.d.), then $P_{(X,Y)} = P_X \otimes P_Y = P_X \otimes P_X$. Moreover, Proposition (ref) implies that, for $c:=E(X)$, $P_X \triangle P_X = \delta_{(c,c)}$, the Dirac delta measure concentrated on $(c,c)\in \Delta$. In that case, ((ref)) becomes
i.e. $(E|X -Y|^p)^{1/p}$, whose value depends only on $P_X$, is a measure of the distance between the product coupling of $P_X$ with itself and the measure concentrated on a point of the diagonal of $I\!\!R^2$. Note that $\delta_{(E(X),E(X))}$ is not in $\Pi (P_X,P_X)$. However, since $\pi_1 := P_X \otimes P_X$ and $\pi_2 := \delta_{(E(X),E(X))}$ have the same center, $||C(\pi_1) - C(\pi_2)|| = 0$ in ((ref)).
Consider two random variables $X,Y \in {\cal L}_1 (I\!\!R)$. As we have seen in Subsection (ref), the Gini mean difference (GMD) of an income distribution $\mu$ is defined as $E|X-Y|$, where $X\sim \mu$ and $Y\sim \mu$ are non-negative and independent. We are now able to give a probabilistic definition of the GMD, perhaps the most general definition that can be given to this index. Looking at ((ref)), and setting $p=1$, we obtain
In other words, the GMD of an income (or a wealth, etc.) distribution $\mu$ represents a distance between the product coupling $\mu \otimes \mu$ of $\mu$ with itself and the Dirac delta measure supported on the single point $(E(X),E(X))$. Consequently, GMD$(\mu) = E|X-Y|$ measures how far the i.i.d. $X\sim \mu$ and $Y\sim \mu$ are from almost sure equality. Note that $\mu \otimes \mu$ distributes its mass symmetrically with respect to $\Delta$. The GMD thus measures a distance between the distribution $\mu \otimes \mu$ on $I\!\!R^2$ (or on a set of $A\times A$, $A\subset I\!\!R$) and the measure concentrated on the center of this same distribution.
Many examples can be given of the usefulness of $E|X-Y|$ for independent variables, and they range from physics to economics. In physics, for instance, Lukaszyk (2004)\nocite{Lukaszyk2004} presents a modified Shephard-Liszka\nocite{Liszka1984,Shephard1967} approximation, where $E|X-Y|$ proves to be more reliable than a plain Euclidean metric, suggesting that an analogous improvement can be achieved in various numerical methods, and in particular in approximation algorithms. In the same paper, Lukaszyk suggests further applications in fringe pattern analysis, or in quantum mechanics, to estimate the distance of two quantum particles described by their wave functions. As pointed out in paragraph (ref), $E|X-Y|$ has also important applications in inequality economics. Other applications include clustering systems, pattern recognition or finance.
In Subsection (ref), we indicate the forms in which $E|X-Y|$ is usually represented in scientific publications, especially of physics, engineering, finance and economics. In (ref), we express $E|X-Y|$ in analytic form when $X \hbox{ and } Y$ are two independent normally distributed random variables. This new result generalizes specific formulas already present in the literature, in particular that of physics. In (ref), we give closed-form formulas for the average distance and the normalized average distance between coordinates of points falling at random into a rectangle of the plane. In (ref), we take advantage of results of (ref) to express in closed form the mean value of the set $\{|a-b|: a \in A, b \in B \}$, where $A$ and $B$ are bounded intervals of real numbers.
For independent $X,Y \in {\cal L}_1(I\!\!R )$, let $P_{(X,Y)} = \mu \otimes \nu$ be the product distribution of two probability distributions $\mu \hbox{ and } \nu$ defined on $(I\!\!R,{\cal B}_1)$. The Fubini-Tonelli theorem applies, and we can interchange the order of integration or summation so that
Thereafter, $F$ (resp. $G$) will refer to the cumulative distribution function (cdf) of $X$ (resp. $Y$), and $f$ (resp. $g$) will refer to the probability density function (pdf) of $X$ (resp. $Y$). Particularly interesting are the cases where independent $X$ and $Y$ are (i) both discrete, (ii) both absolutely continuous, or (iii) one of them is discrete and the other absolutely continuous.
(i) The (independent) discrete random variables $X$ (resp. $Y$), take values in the countable sets $\Omega_X =\{ x_1,x_2,\ldots,x_i,\ldots \} \subset I\!\!R$ (resp. $\Omega_Y =\{ y_1,y_2,\ldots,y_j,\ldots \} \subset I\!\!R$). In that case, $f$ (resp. $g$) are probability mass functions\footnote{i.e. have densities with respect to the counting measure on $\Omega_X$ (resp. $\Omega_Y$). }. In this context, ((ref)) can be rewritten as $E|X-Y|=\sum_i\sum_j\left|x_{i}-y_{j}\right|p_{i}q_{j}$, where $p_i = P(X = x_i)$ and $q_j = P(Y = y_j)$\footnote{This double sum is just an integral with respect to the counting measure on $\Omega_X \times \Omega_Y$.}. In the finite equiprobable case, when independent random variables $X$ and $Y$ both take the non-negative values $x_1,x_2,\ldots , x_N$ with respective probabilities $p_1=p_2=\cdots = p_N=1/N$, then $$ E|X-Y|= \frac{1}{N^2}\sum_{i=1}^N\sum_{j=1}^N\left|x_{i}-x_{j}\right| $$ is the Gini mean difference in its discrete form, see for example Gini (1912)\nocite{Gini1912}, Kendall and Stuart (1958)\nocite{Kendall1958}, or Xu (2003)\nocite{Xu2003}.
(ii) The independent $X \sim \mu$ and $Y \sim \nu$ are both absolutely continuous\footnote{i.e. absolutely continuous with respect to the Lebesgue measure.}. Equation ((ref)) usually appears in the following form in the literature (notably in physics and economics) $$ E|X-Y|= \int^{\infty}_{x=-\infty} \int^{\infty}_{y=-\infty}|x-y|f(x)g(y)dxdy. $$ When $X,Y \in {\cal L}_1(I\!\!R )$ are i.i.d. non-negative absolutely continuous random variables ($f=g$), $$ E|X-Y|= \int^{\infty}_{x=-\infty} \int^{\infty}_{y=-\infty}|x-y|f(x)f(y)dxdy. $$ is known in the economic literature as the continuous Gini mean difference (see, for example, Yitzhaki (1998)\nocite{Yitzhaki1998}, or Yitzhaki and Schechtman (2013)\nocite{YitzhakiSchechtman2013}).
(iii) $X$ is discrete, whereas $Y$, independent of $X$, is absolutely continuous. Then $$ E|X-Y|= \sum_i p_i \int^{\infty}_{y=-\infty}|x_i-y|g(y)dy. $$ A special case of this formula was used by Lukaszyk (2004)\nocite{Lukaszyk2004}, when he proposed a modified Liszka method to handle an experimental mechanics issue.
The normal case plays a crucial part in a great many of the techniques used in applied statistics. The Central-limit theorem alone ensures that this will be the case, but there are other important reasons extensively discussed in the literature.
We begin this section by writing $E|X-Y|$ in a form facilitating the calculation of its analytic expression when independent $X \hbox{ and } Y$ are both absolutely continuous.
Consider the special case where $X \hbox{ and } Y$ in Proposition (ref) are i.i.d. Then $F=G$, $\mu_X = \mu_Y$, and ((ref)) becomes
Moreover, since $X$ absolutely continuous $\Rightarrow$ $F$ continuous $\Rightarrow$ $F(X) \sim U(0,1)$ $\Rightarrow$ $E[F(X)] = 1/2$, we have: $\hbox{\rm cov}[X,F(X)] = E[XF(X)] - \mu_X/2 \stackrel{(\ref{XFX})}{=} E|X-Y|/4$, i.e.
This result -- of which ((ref)) is a generalization -- can be found in Lerman and Yitzhaki (1984)\nocite{Lerman1984}.
In the following (new) theorem (Theorem (ref)), we give the analytic form of $E|X-Y|$ for normally distributed independent random variables.
Equation ((ref)) can also be written
Note that ((ref)) is the formula we end up with if we use Proposition (ref). To obtain ((ref)), we used the fact that the convolution of two Gaussian distributions is a Gaussian distribution. The proof of ((ref)) is shorter than that of ((ref)) resulting from Proposition (ref). However, Proposition (ref) applies to absolutely continuous variables and we wanted to test it in the particular case where the variables are Gaussian. It is left to the reader to show that ((ref)) and ((ref)) are equivalent.
The rest of Subsection (ref) is devoted to corollaries of Theorem (ref). Equation ((ref)) generalizes formulas that have already proven their usefulness in the applied sciences, especially in physics, as illustrated by the following examples. First, consider the case of a degenerate normal random variable $Y$ with mean $\mu_Y$ and variance $\sigma^2_Y \searrow 0$. Symbolically, we may write $Y\sim N(\mu_Y,0) = \delta_{\mu_Y}$, where $\delta_{\mu_Y}$ is the Dirac measure supported on the singleton $\{\mu_Y\}$. Let us write $\sigma$ instead of $\sigma_X$ and $\mu_{XY}$ instead of $|\mu_X-\mu_Y|$. Taking the limit $\sigma^2_Y \searrow 0$ in ((ref)) and ((ref)), we obtain the respective formulations ((ref)) and ((ref)) below
Using the equality ${\displaystyle \Phi(z)=1-\hbox{\rm erfc}\left(\frac{z}{\sqrt{2}} \right)/2 }$, ((ref)) becomes
Lukaszyk (2004)\nocite{Lukaszyk2004} used ((ref)) to successfully implement a modified Liszka approximation method to experimental mechanics. Note that ((ref)), unlike ((ref)), leads directly to Lukaszyk's formula ((ref)).
If $X\sim N(\mu,\sigma^2)$, a direct consequence of ((ref)) is that $E|X - \mu| = \sqrt{2}\sigma/\sqrt{\pi} \approx 0.7979 \sigma$. Indeed, setting $Y\sim \delta_{\mu}$ implies that $\mu_X = \mu_Y = \mu$ and $\mu_{XY} = 0$.
In the same paper, Lukaszyk studied the case of two normal distributions having the same variance\footnote{The assumption of homoscedasticity greatly facilitated the integral calculations undertaken by Lukaszyk, as can be seen in his PH.D. thesis (2001)\nocite{Lukaszyk2001} .}, i.e. $X\sim N(\mu_X,\sigma^2)$ and $Y\sim N(\mu_Y,\sigma^2)$. Setting $\sigma_X=\sigma_Y=:\sigma$ in ((ref)) and ((ref)), we obtain successively
In Lukaszyk's paper $E|X-Y|$ appears in the form
which follows directly from ((ref)).
So Lukaszyk (2001, 2004) found (and used) the formulas when $\sigma_X = \sigma_Y =: \sigma$. When $\sigma_X$ and $\sigma_Y$ are arbitrary, but $\mu_X = \mu_Y$, ((ref)) or ((ref)) directly imply
In other words, ${\displaystyle ||X-Y||_1 = \sqrt{\frac{2}{\pi}} ||X-Y||_2 \approx 0.7979 ||X-Y||_2 }$: the ${\cal L}_1$-distance between $X \hbox{ and } Y$ is approximately one fifth smaller than the ${\cal L}_2$-distance in this case. We used the fact that $\hbox{\rm var} (X-Y) = \sigma_X^2+\sigma_Y^2$, since $X \hbox{ and } Y$ are independent and that $E(X) = E(Y)$ implies that $\hbox{\rm var} (X-Y) = E(X-Y)^2$.
Next, consider the case where $\mu_X = \mu_Y =: \mu$ and $\sigma_X = \sigma_Y =: \sigma$, i.e. $X \hbox{ and } Y$ are i.i.d. and both follow $N(\mu,\sigma^2)$. Setting $\sigma_X = \sigma_Y = \sigma$ in ((ref)), we obtain
which is the Gini mean difference when the underlying income distribution is Gaussian, see e.g. Yitzhaki and Schechtman (2013)\nocite{YitzhakiSchechtman2013}. Note that $E|X-Y|$ in ((ref)), which does not depend on $\mu$, is a measure of dispersion of the same nature as the standard deviation $\sigma$.
Next, if $\sigma^2_X \searrow 0$ and $\sigma^2_Y \searrow 0$, i.e. if $X\sim N(\mu_X,0) = \delta_{\mu_X}$ and $Y\sim N(\mu_Y,0) = \delta_{\mu_Y}$, equations ((ref)) or ((ref)) become
Of course, instead of taking the limits $\sigma^2_X \searrow 0$ and $\sigma^2_Y \searrow 0$ in ((ref)) or ((ref)), one can calculate directly $E|X-Y| = \int_{I\!\!R} \ \{ \int_{I\!\!R} |x-y| \delta(y-\mu_Y) dy\}\ \delta(x-\mu_X) dx = |\mu_X-\mu_Y |$. When $X\stackrel{a.s.}{=} a$ and $Y\stackrel{a.s.}{=} b$, $a,b \in I\!\!R$, $E|X-Y|$ simply transforms into $|a-b|$, i.e. $E|\cdot - \cdot|$ becomes the metric on $I\!\!R$ induced by the norm $|\cdot|$.
We end this subsection by observing that equation ((ref)) provides an analytic formula for $E|X|$, which enables to express $D_{norm}(X,Y)$ in ((ref)) in analytic form when $X$ and $Y$ are independent and normally distributed. To get $E|X|$, assume that $Y\stackrel{a.s.}{=} 0$, i.e. that $Y \sim \delta_0$. Then $\mu_{XY} = |\mu|$ and ((ref)) becomes
In this section, uppercase letters $A,B$ refer to bounded proper intervals of real numbers and $L_A,L_B$ refer to their respective length. A proper interval is an interval that is neither empty (an example of empty interval is $[a,a[ = \emptyset$, for some $a \in I\!\!R$)\footnote{To avoid any confusion, we use the notation $]a,b[$ instead of $(a,b)$ in Subsections (ref) and (ref). } nor degenerate (i.e. of the form $[a,a] = \{a\}$). A proper bounded rectangle in $I\!\!R^2$ is the cartesian product $A \times B$ of two proper bounded intervals. We are interested in univariate or bivariate continuous uniform distributions such as $U(A)$ or $U(A \times B)$. Moreover, for $a \in I\!\!R$, we identify $U(\{a\})$ to the Dirac delta measure $\delta_{a}$ supported on $\{a\}$.
For a given set $C$, let $I_C$ denote the indicator function (defined by $I_C(z) = 1$ if $z\in C$, $I_C(z) = 0$ if $z\notin C$). Consider the random pair $(X,Y) \sim U(A \times B)$. Noting that $L_A^{-1}L_B^{-1} I_{A\times B}(x,y) = L_A^{-1}I_A(x) \cdot L_B^{-1}I_B(y)$ for all $(x,y) \in I\!\!R^2$, one can easily show the following (intuitive) equivalence: [$(X,Y) \sim U(A \times B)$] if and only if [$X \sim U(A)$, $Y \sim U(B)$, $X \hbox{ and } Y$ independent]. We are now ready to ask the question: a point $(x,y)$, which is a realization of $(X,Y)$, falls at random into $A\times B$. What is the average distance between its coordinates? Answer: $E|X-Y|$. Put slightly differently, $E|X-Y|$ reflects the expected absolute difference of two independent variables $X \hbox{ and } Y$ following continuous uniform distributions $X\sim U(A)$ and $Y\sim U(B)$. Theorem (ref) below expresses $E|X-Y|$ in closed form. Without loss of generality, the intervals $A$ and $B$ are assumed to be open in the theorem.
Theorem (ref) can also be used to give closed formulas for $E|X-Y|$ when one of the two intervals $A$ or $B$ is degenerate. Suppose for example that $X\sim U(A)$ and $Y\sim \delta_b$. One way of expressing $E|X-Y|$ is then to take the cases 2 and 3 of Theorem (ref) and to calculate the limit $b_2\searrow b_1$. We obtain
A more direct way to proceed is to use the Lebesgue integral with the coupling $\pi = \mu \otimes \nu$ as the measure used for integration, where $\mu = U(A)$ and $\nu = \delta_b$. Indeed, let $h:I\!\!R^2 \rightarrow I\!\!R$ be given by $h(x,y) = |x-y|$. Then $E|X-Y| = \int h d\pi = \int_A \{ \int_{I\!\!R} |x-y| \delta(y-b)dy\} L_A^{-1}I_A(x) dx = L_A^{-1} \int_A |x-b| dx$. For $A=]a_1,a_2[$, distinguishing the respective situations where (i) $b\in A$ and (ii) $b\notin A$ with $b \le a_1$ and $b \ge a_2$, we find the formulas of ((ref)). If both intervals are degenerate, i.e. of type $\{ a \}$ and $\{ b \}$, then $\mu = \delta_a$, $\nu = \delta_b$, $\pi = \delta_a \otimes \delta_b = \delta_{(a,b)}$, and $\int h d\pi = |a-b|$.
Theorem (ref) allows to easily calculate the normalized average distance $D_{norm} (X,Y) = E|X-Y|/(E|X| + E|Y|) \in [0,1]$ defined in ((ref)). As evidenced by Theorem (ref), the fact that $X \hbox{ and } Y$ (and an additional $Z$) are independent implies that $D_{norm}(\cdot , \cdot)$ satisfies the triangle inequality.
Let $\mathcal I$ denote the set of bounded proper intervals of real numbers. As $A\in \mathcal I$ and the probability measure $U(A)$ are in bijection, the closed formulas for $E|X-Y|$ in Theorem (ref) can also be taken in a purely deterministic sense to compute the “mean value” $D(A,B)$ of the set $\{|a-b|: a \in A, b \in B \}$, where $A,B \in \mathcal I$. Formally, we just have to adopt new notations: replace, say, $E|X-Y|$ by $D(A,B)$, $E|X|$ by $S_A$, $E|Y|$ by $S_B$ and $E|X-Y|/(E|X| + E|Y|)$ by $D_{norm}(A,B) = D(A,B)/(S_A + S_B)$ in Theorem (ref). More precisely, we have the weighted means $D(A,B) = \int_A \int_B |x-y| w(x,y) dxdy / \int_A \int_B w(x,y) dxdy$, $S_A = \int_A |x|u(x)dx / \int_A u(x)dx$ and $S_B = \int_B |y|v(y)dy / \int_B v(y)dy$, where $w$, $u$ and $v$ are the respective weight functions $w=I_{A \times B}$, $u = I_A$ and $v = I_B$. As $D(A,B)$ stems from $E|X-Y|$ and $D_{norm}(A,B)$ stems from $E|X-Y|/(E|X| + E|Y|)$, $D(A,B)$ is invariant under location transformations, while $D_{norm}(A,B)$ is invariant under scale transformations.
It turns out that the functionals $D$ and $D_{norm}$, both ${\mathcal I} \times {\mathcal I} \longrightarrow I\!\!R_+$, satisfy the axioms of symmetry and triangle inequality (the fact that $D_{norm}(\cdot,\cdot)$ satisfies the triangle inequality, far from obvious, is a consequence of Theorem (ref)). However, reflexivity does not hold: indeed $D(A,A) > 0$ and $D_{norm}(A,A) > 0$ for any $A \in {\mathcal I}$, which means that $D$ and $D_{norm}$ are neither distances, nor -- a fortiori -- metrics on $\mathcal I$.\footnote{Actually, $D$ and $D_{norm}$ are metametrics. The term “metametric” is specified in Deza and Deza (2014)\nocite{Deza2014}. Metametrics appear in the study of Gromov hyperbolic metric spaces. They were first defined by V\"ais\"al\"a (2005)\nocite{Vaisala2005}.}
Finally, let $b$ be any fixed real number. To obtain the closed-form formula for the “mean value” $D(A,\{ b\})$ of the set $\{ |a-b|$, $ a \in A \}$, replace $E|X-Y|$ by $D(A,\{ b\})$ in ((ref)).
The following brief introduction to the problem of optimal transport is intended for the many non-specialists in the field. We will stay at a rather heuristic level, focusing on the founding ideas of transport theory. For a detailed account of the theory, the reader is referred to Villani (2003, 2008)\nocite{Villani2008} for example. The problem of optimal transport can be presented in two related ways. The formulation of Monge is ancient and dates back to the 18th century. Kantorovich's work is much more recent and was published during World War II. It can be interpreted as a generalization or a relaxation of Monge's approach. In practice, the latter seems to be more direct and easier to interpret, but its resolution is mathematically more complicated.
This text is designed for a broad readership and focuses on the main ideas of the optimal transport problem. In this perspective of relative simplicity, our discussion here is limited to probability measures $\mu \hbox{ and } \nu$ defined on the real line rather than on more general spaces, so as not to lose sight of the main issues. The originality of this short presentation consists in exploiting known results (if possible with a slightly shifted look) while using notations familiar to practitioners of applied statistics, physics or econometrics. This does not prevent some new results from emerging.
The optimal mass transport problem tries to find the most efficient way to transport a source measure $\mu$ over a target measure $\nu$ taking account of a given cost function. A transport cost determines in some way the difference or distance between these measures. For the sake of simplicity, the cost function $c:I\!\!R^2 \rightarrow I\!\!R_+$ that will be used in this article is mainly of type $c(x,y) = |x-y|$, with sometimes a slight generalization: $c(x,y) = h(x-y)$, where $h$ is convex and continuous. We will see later how the choice of a cost function of type $c(x,y) = |x-y|$ leads to the expected absolute difference $E|X-Y|$ between $X \hbox{ and } Y$.
It should be borne in mind that the optimal transport problem is much more general than the particular cases treated here. This concerns notably -- as we have pointed out -- the type of spaces on which $\mu \hbox{ and } \nu$ are defined, but no less significantly the characteristics of the cost function. The restriction to the one-dimensional case and the use of a simple cost function make it possible to define the problem without technicalities -- sometimes severe -- related to more general cases. When $\mu \hbox{ and } \nu$ are defined on $I\!\!R$, or on a subset of $I\!\!R$, the problems of Monge and Kantorovich have easily interpretable closed form solutions in some important cases. This is a significant property as it alleviates the need for optimization.
When, as in this article, the source measure $\mu$ and the target measure $\nu$ are defined on the same space, the transport of measures has applications in many fields. For instance, if objects are initially distributed according to $\mu$, then they are arranged after transport according to $\nu$. In inequality economics, $\mu$ represents a distribution of income and the problem is to find a planning carrying $\mu$ over a less unequal target distribution $\nu$. In finance, $\mu$ can be the return distribution of a portfolio of stocks and $\nu$ the return of another portfolio or a benchmark.
In the next two sections, we present in detail how the Monge and Kantorovich approaches of the optimal transport unfold when $\mu \hbox{ and } \nu$ are probability measures on the real line (or have a support on the real line).
In his M\'emoire sur la Th\'eorie des D\'eblais et Remblais, Monge (1781),\nocite{Monge1781} was interested in minimizing the cost of transporting sand from a dune to fill a ditch, or transporting stones from an excavation to build a fortification. Monge's historical modeling was in $I\!\!R^3$ and the cost function was the Euclidean distance. In the generalizations that followed, $I\!\!R^3$ became for example a Polish metric space and the cost function took various forms that were quite different from the original Euclidean distance. Above all, the piles of sand or pebbles and the cavities to be filled became over time distributions of probability, of income, of wealth, configurations of physical particles, return distributions of financial assets, and so on.
In order to state Monges's problem in the one-dimensional real case, we need the following definition.
When $T$ transports $\mu$ to $\nu$, we use the notation $\mu_T = \nu$ (rather than $T_{\#}\mu = \nu$). We denote by ${\cal T}(\mu,\nu) = \{ T: (I\!\!R, {\cal B}_1,\mu) \rightarrow (I\!\!R, {\cal B}_1)| T \hbox{ measurable with } \mu_T = \nu\}$ the set of transport maps pushing forward $\mu$ to $\nu$.
Unfortunately, we may not find any measurable map such that $\mu_T = \nu$. In other words, ${\cal T}(\mu,\nu)$ may be empty. We do not have to look very far: let us take $\mu = \delta_{x_1}$ (the Dirac delta measure supported on $\{x_1\}$) and $\nu = \frac{1}{2}\delta_{y_1} + \frac{1}{2}\delta_{y_2}$, with $y_1\neq y_2$. In this case, no map $T$ with $\mu_T = \nu$ can be found. Important cases where ${\cal T}(\mu,\nu) \ne \emptyset$ are (Thorpe 2018)\nocite{Thorpe2018}: (i) the discrete case when $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$, i.e. when $\mu \hbox{ and } \nu$ are supported on the same number of points with equal mass. And (ii) the absolutely continuous case, when $d\mu (x) = f(x)dx$ and $d\nu (y) = g(y)dy$. Moreover, even if ${\cal T}(\mu,\nu) \ne \emptyset$, the constraint in Monge's problem is usually very non-linear and difficult to handle with the classical tools of the calculus of variations. Kantorovich's approach alleviates these problems by seeking an optimal transport plan rather than an optimal transport map.
Note that the set $\Pi (\mu ,\nu)$ of transport plans is never empty since it contains the trivial plan $\mu \otimes \nu$. For any $A,B\in {\cal B}_1$, the quantity $\pi (A\times B)$ tells us how much mass in set $A$ is being moved to set $B$. The total amount of mass removed from $A$ has to be equal to $\mu(A)$ and the total amount of mass moved to $B$ must be $\nu (B)$. Hence the constraints: $\pi (A\times I\!\!R) = \mu(A)$ and $\pi (I\!\!R\times B) = \nu(B)$ for all $A,B \in {\cal B}_1$.
Kantorovich (1942)\nocite{Kantorovich1942} proposed a general formulation of the problem by considering optimal transport plans which allow mass to be split. This is a very important difference between the two approaches; Monge's problem, unlike Kantorovich's, requires that each mass in $x$ is sent to a single position $y$: there is no possible separation of a unit of mass of $\mu$ into several pieces during the transport. Still restricting ourselves to probabilities $\mu \hbox{ and } \nu$ defined on the real line, or having a support on the real line, we can formulate the following definition.
The term $\int_{I\!\!R^2}|x-y| \ d\pi(x,y)$ represents the transport cost from $\mu$ to $\nu$ under $\pi$. We can think of $d\pi (x,y)$ as the amount of mass transferred from $x$ to $y$. Actually, as $c(x,y)=|x - y|$ is continuous\footnote{or even lower semi-continuous, a weaker condition implying the existence of a minimizer.}, a minimizer $\pi^*$ always exists (Gangbo (2004), th. 2.4)\nocite{Gangbo2004} and we can replace “inf” by “min” in ((ref)).
Before resuming the substantive discussion on the problem of transport, we turn for a moment to considerations of a purely formal or didactic nature. The optimal transport domain is mainly conceived in terms of (probability) measures. Our experience is that many newcomers to this field feel more comfortable with concepts such as random variables or mathematical expectation, while the language of measure theory seems less telling to them, at least initially. We think that a brief development will clarify some aspects that specialists may consider as futile or obvious.
(i) Classically, if we have a workspace $(I\!\!R^2,{\cal B}_2, \pi)$, but wish to reason in terms of ramdom variables, we formally introduce a general probability space $(\Omega,{\cal A},P)$ and a pair of random variables $(X,Y)$ so as to obtain the scheme $(\Omega,{\cal A},P) \stackrel{(X,Y)}{\longrightarrow} (I\!\!R^2,{\cal B}_2, \pi)$, where $\Omega :=I\!\!R^2$, ${\cal A} := {\cal B}_2$ and $(X,Y) := Id_2$, the identity map on $I\!\!R^2$. As a consequence, $\pi = P_{(X,Y)} = P_{(Id_2)} = P$, noting the obvious fact that $Id_2 = (q_1,q_2)$, (as in Definition (ref), we use the notations $q_1$ and $q_2$ for the projection functions). By doing so, the representation $(\Omega,{\cal A},P) \stackrel{(X,Y)}{\longrightarrow} (I\!\!R^2,{\cal B}_2, P_{(X,Y)})$ becomes $(I\!\!R^2,{\cal B}_2, \pi) \stackrel{Id_2}\longrightarrow (I\!\!R^2,{\cal B}_2, \pi)$, i.e., in particular, the variables $X \hbox{ and } Y$ transform themselves into canonical projections\footnote{This reminds us that the term “random variable” is rather unfortunate since a random variable is neither random nor a variable. }. So $X \hbox{ and } Y$ can be interpreted indifferently as random variables or as projections. For instance, a notation such as $E_\pi |X-Y|$ instead of $\int_{I\!\!R^2}|x-y| d\pi(x,y))$ makes sense as a mean absolute deviation between two random variables. This convention will be applied below, notably in Figure (ref), which will help us articulate the Kantorovich relaxation of the Monge problem.
(ii) A transport plan $\pi \in \Pi (\mu,\nu)$ is a coupling between $\mu \hbox{ and } \nu$, i.e. a joint distribution $(X,Y) \sim \pi$ such as marginally $X \sim \mu$ and $Y \sim \nu$. By abuse of language the expression “coupling between $X \hbox{ and } Y$” is used to mean “coupling between the distribution of $X$ and the distribution of $Y$”. Considered in probabilistic form, the problem stated in ((ref)) amounts to finding the distribution of a pair of random variables $(X^*,Y^*)$ minimizing the mean absolute deviation $E|X-Y|$ among all jointly distributed pairs $(X,Y)$ such that $P_X = \mu$ and $P_Y = \nu$.
In some special cases, Kantorovich's approach can be used to solve the (difficult) Monge optimisation problem. In this context, deterministic plans play an essential role.
Note that $\pi_T = \mu_{(Id,T)}$ is supported on the graph of $T$. Let us check that the marginals of $\mu_{(Id,T)}$ are $\mu$ and $\nu$, respectively. In this regard, consider $A,B \in {\cal B}_1$. Then, using in particular Lemma (ref) (ii) in Subsection (ref): \newline $\pi_T (A\times I\!\!R) = \mu_{(Id,T)} (A\times I\!\!R) = \mu [(Id,T)^{-1}(A\times I\!\!R)] = \mu [Id^{-1}(A) \cap T^{-1}(I\!\!R)) = \mu (A)$.\newline $\pi_T (I\!\!R \times B) = \mu_{(Id,T)} (I\!\!R \times B) = \mu [(Id,T)^{-1}(I\!\!R \times B)] = \mu (T^{-1}(B)) = \mu_T (B) = \nu (B)$, since $T$ is a transport map.
Figure (ref) provides a convenient overview of the situation. In particular, and in relation to Subsection (ref), it clearly shows that we can express the probability measures in terms of law (${\cal L}$) of random variables: there exists random variables $X \hbox{ and } Y$ such that $\pi = {\cal L}((X,Y))$, $\mu ={\cal L}(X)$, $\nu = {\cal L}(Y)$, $\mu_T = {\cal L}(T(X))$ and $\pi_T = {\cal L}((X,T(X)))$.
This may seem like a detail, but the way to represent the transport cost associated with a transport plan $\pi$ or a transport map $T$ varies according to the research field. Three common types of notations are shown in Table (ref) for the cost function $c(x,y) = |x-y|)$, where $c_{\pi}$ and $c_T$ denote the cost resulting from the transformation of $\mu$ into $\nu$ by means of $\pi$ and $ T$, respectively. \vskip .4cm
It turns out that the formulas in (b), (c) and (d) of Table (ref) are different expressions of the same quantity $c_T = \int c\circ (Id,T)\circ X d\pi$. This results immediately from a repeated use of the change of variable formula in the context of Figure (ref), recalling that $X = q_1$ under the convention adopted in Subsection (ref). For example, let us show that (c) = (b). Remembering that $c(x,y) = |x-y|$, we have $E_\pi |X-T(X)|= \int c \circ (Id,T) \circ X d\pi$ (see Figure (ref)). On the other hand, $E_{\pi_T} |X-Y| = \int c d\pi_T = \int c d\mu_{(Id,T)} = \int c\circ (Id,T) d\mu = \int c\circ (Id,T) d\pi_{q_1} = \int c \circ (Id,T) d\pi_X = \int c\circ (Id,T)\circ X d\pi$.
A heuristic interpretation of (a) is as follows: a pair of random variables $(X,Y)$ is defined on a probability space $(\Omega,{\cal A},P) = (I\!\!R^2,{\cal B}_2, \pi)$. We observe independently an infinity of realizations $(x,y)$ of $(X,Y)$ falling on $I\!\!R^2$ according to $\pi$, we calculate the distance $|x-y|$ between the coordinates and average these distances to obtain $E_\pi |X-Y|$. The three equivalent expressions (b), (c) and (d) are interpreted in an almost identical way. Take $E_{\pi_T} |X-Y|$: we observe independently an infinity of realizations $(x,y)$ of $(X,Y) \sim \pi$. But this time, we forget $y$ to replace it by $T(x)$. In other words, an infinite number of pairs $(x,y)$ fall (almost surely) on the graph of $T$. We calculate the distances between the coordinates of these pairs and average them to find $E_{\pi_T} |X-Y|$.
The most important equality in Table (ref) is
which shows that any transport map $T$ induces a transport plan of the same cost, i.e. can be canonically embedded into the set of transport plans. Adopting the convention that $\inf \emptyset = \infty$ if ${\cal T}(\mu,\nu) = \emptyset$, this means that
Equality in ((ref)) holds under fairly general assumptions when any plan can be approximated by transport maps (see e.g. Ambrosio and Pratelli (2003)\nocite{Ambrosio2003})\footnote{The presence of atoms can seriously impede the existence of transport maps. Under fairly general assumptions, it can be shown that if $\mu$ is atomless (in our setting, this means that $\mu (\{ x\}) = 0 \ \forall x\in I\!\!R$), then the set $\{ \pi_T : \mu_T = \nu \}$ is weak-*dense in $\Pi (\mu,\nu)$, which implies equality in ((ref)) (see Carlier (2010)\nocite{Carlier2010} or Ambrosio et al. (2004)\nocite{Ambrosio2004}). Ambrosio (2002)\nocite{Ambrosio2002} notes that the infimum of the Kantorovich problem “is attained on an extremal element of $\Pi (\mu,\nu)$”. However, all extremal points are not induced by transport maps “otherwise one would get existence of transport maps directly from the Kantorovich formulation”. It can be shown that deterministic transport plans are extremal in $\Pi (\mu,\nu)$. “Unfortunately, the extremal points of $\Pi (\mu,\nu)$ are not all transport plans, except in very particular cases. It turns out that the existence of optimal transport maps depends not only on the geometry of $\Pi (\mu,\nu)$, but also (in a quite sensible way) of the choice of the cost function $c$”.}.
We are now ready to show the following result: (i) if the Kantorovich problem admits an optimal plan (minimizer) $\pi$, and (ii) this plan turns out to be deterministic i.e. of the form $\pi_T$, then the transport map $T$ is Monge-optimal. To see that, let us assume that $\pi_T$ is Kantorovich-optimal. Then \newline $\inf_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| \leq E_\mu |X-T(X)| = E_{\pi_T} |X-Y| =^{(\hbox{hyp.})} \min_{\pi \in \Pi (\mu,\nu)} E_{\pi} |X-Y|$ \newline $ \leq^{(\ref{infinf})} \inf_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| $ and therefore \[ E_\mu |X-T(X)| = \inf_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| = \min_{S \in {\cal T} (\mu,\nu)} E_\mu |X-S(X)| . \] This relaxation of Monge by Kantorovich occurs in a few important cases. We will see two examples below (in (ref) and (ref)).
The following theorem and its proof can be found in Thorpe (2018), \nocite{Thorpe2018} see also Villani (2003) \nocite{Villani2003} and Santambrogio (2015). \nocite{Santambrogio2015} It is a powerful tool for the treatment of the two cases mentioned above.
In this subsection (Subsection (ref)), we are mainly interested in a cost function of type $c(x,y) = |x-y|$. Nevertheless, all the results obtained remain valid for a cost function such as the one specified in Theorem (ref).
Still limited to one-dimensional probability measures $\mu , \nu \in {\cal P}(I\!\!R)$, we give two examples where the Kantorovich relaxation approach leads to a solution of the Monge problem. We base ourselves on the criterion stated in Theorem (ref).\newline (i) When the continuous cost function satisfies a certain convexity condition, an optimal plan is supported on a “curve” in $I\!\!R^2$ depending on the quantile functions associated with respective cdf's $F$ of $\mu$ and $G$ of $\nu$ (quantile functions are defined in Definition (ref)) below. Moreover, if $\mu$ is atomless, that is if $F$ is continuous, then this optimal plan is deterministic, i.e. associated with a Monge-optimal transport map. \newline (ii) Theorem (ref), applied this time to the special discrete case where $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$, also leads to a deterministic optimal plan. The optimization process will provide an interesting by-product in that case (Proposition (ref)).
We now need the following definition to give an analytic representation of the optimal plan $\bar{\pi}$ referred to in Theorem (ref).
The function $F^-$ always exists, even when $F$ is not continuous or not strictly increasing. As both $F$ and $F^-$ are monotonically increasing, they are also measurable, an important property that we will use later. The notation $F^{-1}$ instead of $F^-$ is used by many authors. However, it can sometimes be confusing (it may be mistaken for the preimage operator of a set, see below). If $F$ is continuous and strictly increasing, then the two generalized inverses are equal to the ordinary inverse (on the range of $F$). One can often work with $F^-$ as if it were an ordinary inverse. Note that ((ref)) implies $F^-(0) = -\infty$, and we adopt the convention that $\inf \emptyset = \infty$. Figure (ref) illustrates the inequalities $F^- \circ F(x_0) \leq x_0 \leq F^+ \circ F(x_0)$ when $x_0$ corresponds to a flat part of $F$, but these two inequalities are in fact valid for any $x_0 \in I\!\!R$.
An important property of the quantile function $F^-$ is the following: for each $t$ and $x$
(noting that to prove the implication “$\Rightarrow$”, one uses the right continuity of $F$).
We are now ready to represent in a more useful way a plan which -- like the one of Theorem (ref) -- has a cdf of type $H(x,y) = \min \{F(x), G(y) \}$, where $F$ and $G$ are the cdf's characterizing the probability measures $\mu , \nu \in {\cal P}(I\!\!R)$, respectively. To that aim, consider the “curve” $K: [0,1]\rightarrow I\!\!R^2$ given by $K(t) = (F^-(t),G^-(t))$.\footnote{We use the term “curve” by abuse of language, even if $K$ is not continuous. $K$ is a parametric curve in the usual sense when $F^-$ and $G^-$ are continuous, which is true if and only if $F$, resp. $G$, are strictly increasing. } Note that $K$ is measurable, since $F^-$ and $G^-$ are measurable. Examples of such “curves” are given in Figure (ref). We need the following lemma:
(i) is trivial, (ii) is well-known. Property (iii) is a consequence of ((ref)). Indeed, \newline $(F^-)^{-1} (A_x) = \{ t \in [0,1]: F^-(t) \leq x \} = \{ t \in [0,1]: t \leq F(x) \} = [0,F(x)]$.
Let us designate by $\lambda$ the Lebesgue measure restricted to $[0,1]$. Then $\lambda_K$ denotes the pushforward probability measure of the Lebesgue measure on $[0,1]$ induced by $K$ (on $(I\!\!R^2,{\cal B}_2)$).\footnote{Another way of looking at $\lambda_K$: let $X \sim F$ (resp. $Y \sim G$) be the cdf characterizing $\mu$ (resp. $\nu$), and let $U\sim U(0,1)$ be a random variable uniformly distributed on $[0,1]$. Then $\hat{X} := F^-(U) \sim \mu$, $\hat{Y} := G^-(U) \sim \nu$ and $\lambda_K$ is the law of the pair $(\hat{X},\hat{Y})$. That is, the transport plan $\lambda_K$ is a coupling of $\mu \hbox{ and } \nu$ or, in other words, $(\hat{X},\hat{Y})$ is a coupling of $X \hbox{ and } Y$. } Adopting, as in Theorem (ref), the notation $\bar{\pi}$ for a plan with cdf $H(x,y) = \min \{F(x), G(y) \}$, we show that: (A) $\lambda_K \in \Pi (\mu,\nu)$, and (B) $\lambda_K = \bar{\pi}$.
To prove (A), i.e. to prove that $\mu \hbox{ and } \nu$ are the marginals of $\lambda_K$, we use the rule governing composition of maps: the first marginal of $\lambda_{(F^-,G^-)}$ is $(\lambda_{(F^-,G^-)})_{q_1} = \lambda_{q_1 \circ (F^-,G^-)}) = \lambda_{F^-} = \mu$. To see that $\lambda_{F^-} = \mu$, we can use the points (i) and (iii) of Lemma (ref): define $A_x = (-\infty,x]$. We only have to show that $\lambda_{F^-}(A_x) = \mu (A_x)$ for any $x\in I\!\!R$. But $\lambda_{F^-}(A_x) = \lambda ((F^-)^{-1} (A_x)) \stackrel{(iii)}{=} \lambda([0,F(x)]) \stackrel{(i)}{=} F(x) = \mu (A_x)$. Using the same rule and $q_2$ instead of $q_1$, we see immediately that the second marginal of $\lambda_{(F^-,G^-)}$ is $\lambda_{G^-} = \nu$.
Next, using (i), (ii) and (iii) of Lemma (ref), we are now ready to prove (B), i.e. to prove that $\bar{\pi} = \lambda_K$. For $A_x = (-\infty,x]$ and $B_y = (-\infty,y]$, all we need to show is that $\bar{\pi}(A_x \times B_y) = \lambda_K (A_x \times B_y)$. Then $\bar{\pi} (A_x \times B_y) = H(x,y) = \min \{F(x), G(y) \} =^{(i)} \lambda([0,\min \{F(x), G(y) \}]) \newline = \lambda([0,F(x)] \cap [0,G(y)]) =^{(iii)} \lambda([(F^-)^{-1} A_x] \cap [(G^-)^{-1} B_y]) =^{(ii)} \lambda( (F^-,G^-)^{-1}(A_x \times B_y)) \newline = \lambda( K^{-1} (A_x \times B_y)) = \lambda_K( A_x \times B_y )$.
Taking into account the results we have just stated, the great contribution of Theorem (ref) is to establish that $\lambda_K$ is Kantorovich-optimal as long as the cost function $c(x,y)$ satisfies the required convexity condition.
Now, suppose that the cost function $c(x,y) = h(x-y)$ satisfies the convexity condition stated in Theorem (ref). Since $\bar{\pi} = \lambda_K$, the optimal cost is given by
We simply used the change of variables formula: $\bar{c} = \int c \ d\lambda_{(F^-,G^-)} = \int c \circ (F^-,G^-) d\lambda = \int_0^1 h(F^-(t) - G^-(t))dt$. Taking $c(x,y) = |x-y|$ as a special case, ((ref)) becomes
Note that $\bar{c}$ in ((ref)) is not only the $L_1$-distance between quantile functions, it is also the $L_1$-distance between the corresponding cumulative distribution functions, i.e.
A proof of this remarkable coincidence is given in Thorpe (2018)\nocite{Thorpe2018}, see also Rachev and Rueschendorf (1998)\nocite{Rachev1998}. It should also be noted that Theorem (ref) does not make any particular assumption on $F$ or $G$ and that the optimal plan $\bar{\pi} = \lambda_K$ has not necessarily the deterministic form $\pi = \pi_T =\mu_{(Id,T)}$ for a certain transport map $T$; therefore does not, as it stands, help to solve Monge's problem. The question now is to give additional assumptions about $\mu \hbox{ and } \nu$ (or $F$ and $G$) ensuring that $K((0,1])$ is the graph of a transport map $T$: the very fact that this curve is the graph of a transport map $T$ means that $T$ is Monge-optimal. We give below two special instances where this happens: $K((0,1])$ is the graph of a transport map $T$ (i) when $F$ is continuous (Subsection (ref)) and (ii) in the discrete case, when $\mu \hbox{ and } \nu$ are supported on the same number of points of identical mass (Subsection (ref)). Indeed, these two subsections yield two situations where, under the assumptions of Theorem (ref), $\lambda_{(F^-,G^-)} = \mu_{(Id,T)}$.
Examples of continuous cdf's appear in Figure (ref). We now show that if $F$ is continuous, then the deterministic plan $\pi_{G^-\circ F} = \mu_{(Id,G^-\circ F)}$ has distribution function $H(x,y) = \min \{F(x), G(y)\}$. Theorem (ref) then implies that $\mu_{(Id,G^-\circ F)}$ is Kantorovich-optimal for any cost function having the convexity property specified in this theorem. This means that $T = G^-\circ F$ is Monge-optimal for the same cost function. We need the following lemma:
We will use the properties (i), (ii) and (iii) of Lemma (ref), as well as the properties (a) and (b) of Lemma (ref). Let $(x,y) \in I\!\!R^2$. According to Theorem (ref), all we have to do is show that $\pi_{(G^-\circ F)} (A_x \times B_y) = \min \{ F(x),G(y)\}$, where $A_x=(-\infty ,x]$ and $B_y=(-\infty ,y]$.
We have \newline $\pi_{(G^-\circ F)} (A_x \times B_y) = \mu_{(Id,G^-\circ F)}(A_x \times B_y) =^{(ii)} \mu [\ A_x \cap F^{-1} ((G^-)^{-1}(B_y))\ ]$ \newline $=^{(iii)} \mu (\ A_x \cap F^{-1} ([0,G(y)]\ ) =^{(b)} \mu (\ F^{-1} ([0, F(x)]) \cap F^{-1} ([0,G(y)] \ ) $ \newline $= \mu (\ F^{-1} ([0, F(x)] \cap [0,G(y)])\ ) = \mu_F (\ [0, F(x)] \cap [0,G(y)]\ ) =^{(a)} \lambda (\ [0,\min \{ F(x),G(y)\}]\ )$ \newline $=^{(i)}\min \{ F(x),G(y)\} = H(x,y)$.
In Example (ref), $F$ (standard normal) and $G$ (lognormal) are both continuous and correspond to case (a) in Figure (ref). If $c(x,y) = h(x-y)$ is a cost function with $h$ convex and continuous, the curve $C$ in Figure (ref) is the graph of a (strictly increasing) optimal transport map $T(x) = G^- \circ F (x)$.
We now propose a new case where the solution of Monge's problem involves Kantorovich' relaxation. Let the cost function $c(x,y) = h(x-y)$ be as in Theorem (ref) (i.e. $h$ is convex and continuous) and assume that $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$. As we are dealing here with dimension one, the $x_i$'s and $y_j$'s are real numbers, and we assume that they are ordered: $x_1 \leq x_2 \leq \cdots \leq x_n$ and $y_1 \leq y_2 \leq \cdots \leq y_n$. Define $t_i = F(x_i)$, and $s_j = G(y_j)$, $i,j \in \{ 1, \ldots , n \}$. As all points $\{x_i\}$ and $\{y_j\}$ have the same mass ($1/n$), $t_i = s_i$, $i = 1, \ldots , n$. To prove that $\lambda_{(F^-,G^-)} = \mu_{(Id,G^-\circ F)}$, all we have to show is that $\lambda_{(F^-,G^-)}(\{(x_i,y_j)\}) = \mu_{(Id,G^-\circ F)} (\{(x_i,y_j)\})$ for all $i,j \in \{ 1, \ldots , n \}$. Consider the partition $\{(0,t_1],(t_1,t_2], \ldots , (t_{n-1},t_n =1]\}$ of the interval $(0,1]$, each element of the partition being of length $1/n$. Since $(F^-,G^-)((0,t_1]) = (x_1,y_1), (F^-,G^-)((t_1,t_2]) = (x_2,y_2), \ldots , (F^-,G^-)((t_{n-1},t_n]) = (x_n,y_n)$, we have \newline $\lambda_{(F^-,G^-)}(\{(x_1,y_1)\}) = \lambda_{(F^-,G^-)}(\{(x_2,y_2)\}) = \cdots = \lambda_{(F^-,G^-)}(\{(x_n,y_n)\}) =1/n$. And since \newline $\sum_{i=1}^n \lambda_{(F^-,G^-)}(\{(x_i,y_i)\}) = 1$, we have $\lambda_{(F^-,G^-)}(\{(x_i,y_j)\}) = \frac{1}{n}\delta_{ij}$ where $\delta_{ij}$ is the Kronecker delta ($\delta_{ij}$ equals one if $i=j$, zero otherwise). On the other hand, note that $x_i \stackrel{F}{\longmapsto} t_i = s_i \stackrel{G^-}{\longmapsto} y_i$ (because $G^-$ is left-continuous), and therefore
Then $\mu_{(Id,G^-\circ F)} (\ \{(x_i,y_j)\}\ ) = \mu [\ (Id,G^-\circ F)^{-1} \{(x_i,y_j)\}\ ] = \mu (\ \{x_i\}\cap (G^-\circ F)^{-1} \{y_j\}\ ) =^{(\ref{optiT1})} \mu (\{ x_i\} \cap \{ x_j\}) = \frac{1}{n}\delta_{ij}$, and we have thus proved that $\lambda_{(F^-,G^-)} = \mu_{(Id,G^-\circ F)}$. Denoting by $T^* = G^-\circ F$ the Monge-optimal transport map, one gets from ((ref))
Proposition (ref) is a consequence of the above.
Note that the results shown in Subsection (ref) have a multidimensional extension when the cost function $c$ is general and the optimization problem restricts to discrete measures \newline $\mu = \frac{1}{n}\sum_{i=1}^n \delta_{x_i}$ and $\nu = \frac{1}{n}\sum_{j=1}^n \delta_{y_j}$, where $x_i , y_j \in I\!\!R^d$, $d \ge 2$. However, the $x_i$'s and $y_j$'s cannot be totally ordered in that case and the Monge-optimal transport map $T_{\sigma}$ defined by $T_\sigma (x_i) = y_{\sigma(i)}$, $i = 1,\ldots ,n$, is not necessarily the one corresponding to the trivial permutation $\bar{\sigma}(i) := i$. Using the Minkowski-Carath\'eodory Theorem and the Birkhoff Theorem, Thorpe (2018, Th. 2.5 and 2.6)\nocite{Thorpe2018} shows that any solution $\pi^*$ to Kantorovich's optimal transport problem is a permutation matrix, i.e. there exists a permutation $\sigma^* \in S$ such that $\pi_{ij}^* = \frac{1}{n}\delta_{j=\sigma^*(i)}$, which implies that $T^*: I\!\!R^d \rightarrow I\!\!R^d$ defined by $T^*(x_i) = y_{\sigma^*(i)}$ is a Monge-optimal transport map, see also Villani (2003) \nocite{Villani2003}.
The results of Section (ref) can be summarized as follows: let probability measures $\mu ,\nu$ on $(I\!\!R, {\cal B}_1)$ have the respective cdf's $F$ and $G$. For a cost function $c(x,y) = |x-y|$, or more generally $c(x,y) = h(x-y)$ with $h$ convex and continuous, the transport plan $\lambda_{(F^-,G^-)}$ supported on the “curve” $C = (F^-,G^-)(\ (0,1]\ )$ is Kantorovich-optimal. In this context, we gave two examples (special cases) where the Kantorovich-optimal transport plan $\lambda_{(F^-,G^-)}$ becomes $\mu_{(Id,G^-\circ F)}$ and the curve $C$ is the graph of a Monge-optimal transport map $T = G^- \circ F$. In the first example, $F$ is assumed to be continuous, while in the second example, $\mu \hbox{ and } \nu$ are assumed to be discrete and supported on the same number of ordered real numbers having the same mass. Proposition (ref) is a consequence of the latter case.
We add this section for completeness, noting that the optimal transport problem can be put forward to define a distance between $\mu \hbox{ and } \nu$, namely the so-called Wasserstein metric (or Wasserstein distance). It is also known, in computer sciences, as the earth mover's distance. Like the optimal transport problem discussed in this paper, the definition of the Wasserstein distance is set in a much broader context than the one presented below. This concerns the shape or the properties of the cost function, the characteristics of the spaces on which $\mu$ and $\nu$ are defined (usually $I\!\!R^d$ or subsets of $I\!\!R^d$), as well as the value of $p$ in the definitions that follow.
The set of probability measures on $(I\!\!R ,{\cal B}_1)$ (or on $(E,E \cap {\cal B}_1)$, $E \subset I\!\!R$) with finite $p$-th moment is defined as \[ {\cal P}_p (I\!\!R ) = \left\{ \mu \in {\cal P} (I\!\!R ) :\int_{I\!\!R}|x|^p d\mu(x) < \infty \right\}, \] noting that if $E \subset I\!\!R$ is bounded, then ${\cal P}_p (E ) = {\cal P} (E )$.
It can be shown that the distance $ W_p: {\cal P}_p(I\!\!R) \times {\cal P}_p(I\!\!R) \rightarrow [0,\infty)$ is a metric on ${\cal P}_p(I\!\!R)$. Under the probabilistic notation adopted above, where the random variables are projections (i.e. $X = q_1$ and $Y = q_2$), ((ref)) becomes \[ W_p(\mu ,\nu) = \left( \inf_{\pi \in \Pi (\mu,\nu)} E_\pi |X-Y|^p \right)^{1/p}. \] In particular, for $p = $1
that is, since $c(x,y) = |x-y|$ is convex and continuous
as we have seen in ((ref)) and ((ref)).
In this article, we focused on the $L_p$ distances and in priority on $L_1$. Its content is intended to be a compromise between theoretical considerations and applications. We did not dwell on the already well-known properties of these distances, but we went back to basics of probability theory to clarify various aspects that applied science papers tend to neglect. We have examined the relationship between the $L_1$ distance and the Gini-Kantorovich distance (an $L_1$ distance between cfd's), as well as certain uses of the former, such as the Gini mean difference or the Lukaszyk-Karmowski metric. Unlike simple metrics such as the Gini-Kantorovich distance, $E|X-Y|$ integrates the dependency structure between $X \hbox{ and } Y$ and makes it possible in particular to take account of the assumption of independence. We then studied the axiomatic in which $E|X-Y|$ is inscribed; this allowed us to uncover a interpretive error that crept into the literature on the subject. The properties of $E|X-Y|$ have been clarified in the cases of independence, equality of distribution and almost sure equality. The problem of the normalization of $E|X-Y|$ has been solved, especially the question of triangle inequality. The Gini index is a special case of this [0,1]-normalized form. We have also shown that, for identically distributed variables $X \hbox{ and } Y$, $E|X-Y|$ can be interpreted as a distance to almost sure equality. In a section reserved more specifically for applications, $E|X-Y|$ is expressed in analytic form when $X\sim N(\mu_X,\sigma^2_X)$ and $Y\sim N(\mu_Y,\sigma^2_Y)$ are independent. The resulting formula generalizes tools used in applied physics: it allows in particular to lift the assumption of homoscedasticity ($\sigma^2_X = \sigma^2_Y$), and thus allows more flexibility in the use of the Lukaszyk-Karmowski metric. Closed-form formulas are also determined when independently distributed $X \hbox{ and } Y$ have continuous uniform distributions. These formulas can be used to compute the average distance between intervals. Moreover, leads are opened for obtaining analytical forms when $X \hbox{ and } Y$ follow non-normal distributions, noting that such an attempt can be very complicated or even hopeless in some cases. Finally, for two probability measures $\mu$ and $\nu$ defined on the real line, $E|X-Y|$ is a key ingredient in the optimal transport problem when the issue is how to transport $\mu$ to $\nu$, whilst minimizing a cost function of the form $c(x,y) = |x-y|$. The question of optimal transport is developed in a relatively complete way within this restricted framework. The way the problem is presented has been thought of as a first approach to a particularly demanding domain.