EconBase
← Back to paper

Principal Component Analysis: A Generalized Gini Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,335 characters · 18 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Principal Component Analysis: A Generalized Gini Approach

abstractA principal component analysis based on the generalized Gini correlation index is proposed (Gini PCA). The Gini PCA generalizes the standard PCA based on the variance. It is shown, in the Gaussian case, that the standard PCA is equivalent to the Gini PCA. It is also proven that the dimensionality reduction based on the generalized Gini correlation matrix, that relies on city-block distances, is robust to outliers. Monte Carlo simulations and an application on cars data (with outliers) show the robustness of the Gini PCA and provide different interpretations of the results compared with the variance PCA.

keywords : Generalized Gini ; PCA ; Robustness.

\break

Introduction

This late decade, a line of research has been developed and focused on the Gini methodology, see Yitzhaki13 for a general review of different Gini approaches applied in Statistics and in Econometrics.\footnote{See giorgi for an overview of the "Gini methodology".} Among the Gini tools, the Gini regression has received a large audience since the Gini regression initiated by ref9. Gini regressions have been generalized by Yitzhaki13 in different areas and particularly in time series analysis. ref13 and ref2 investigated ARMA processes with an identification and an estimation procedure based on Gini autocovariance functions. This robust Gini approach has been shown to be relevant to heavy tailed distributions such as Pareto processes. Also, shelef16 proposed a unit root test based on Gini regressions to deal with outlying observations in the data.

In parallel to the above literature, a second line of research on multidimensional Gini indices arose. This literature paved the way on the valuation of inequality about multiple commodities or dimensions such as education, health, income, etc., that is, to find a real-valued function that quantifies the inequality between the households of a population over each dimension, see among others, List, Gajdos, Decancq. More recently, Banerjee shows that it is possible to construct multidimensional Gini indices by exploring the projection of the data in reduced subspaces based on the Euclidean norm. Accordingly, some notions of linear algebra have been increasingly included in the axiomatization of multidimensional Gini indices.

In this paper, in the same vein as in the second line of research mentioned above, we start from the recognition that linear algebra may be closely related to the maximum level of inequality that arises in a given dimension. In data analysis, the variance maximization is mainly used to further analyze projected data in reduced subspaces. The variance criterion implies many problems since it captures a very precise notion of dispersion, which does not always match some basic properties satisfied by variability measures such as the Gini index. Such a property may be, for example, an invariance condition postulating that a dispersion measure remains constant when the data are transformed by monotonic maps.\footnote{See furman for the link between variability (risk) measures and the Gini correlation index.} Another property typically related to the Gini index is its robustness to outlying observations, see e.g. Olkin in the case of linear regressions. Accordingly, it seems natural to analyze multidimensional dispersion with the Gini index, instead of the variance, in order to provide a Principal Components Analysis (PCA) in a Gini sense (Gini PCA).

In the field of PCA, Baccini and Korhonen are among the first authors dealing with a $\ell_1$-norm PCA framework. Their idea was to robustify the standard PCA by means of the Gini Mean Difference metric introduced by Gini12, which is a city-block distance measure of variability. The authors employ the Gini Mean Difference as an estimator of the standard deviation of each variable before running the singular value decomposition leading to a robust PCA. In the same vein, Ding2006 make use of a rotational $\ell_1$ norm PCA to robustify the variance-covariance matrix in such a way that the PCA is rotational invariant. Recent PCAs derive latent variables thanks to regressions based on elastic net (a $\ell_1$ regularization) that improves the quality of the regression curve estimation, see Zou.

In this paper, it is shown that the variance may be seen as an inappropriate criterion for dimensionality reduction in the case of data contamination or outlying observations. A generalized Gini PCA is investigated by means of Gini correlations matrices. These matrices contain generalized Gini correlation coefficients (see Yitzhaki03) based on the Gini covariance operator introduced by Schechtman87 and Schechtman03. The generalized Gini correlation coefficients are: (i) bounded, (ii) invariant to monotonic transformations, (iii) and symmetric whenever the variables are exchangeable. It is shown that the standard PCA is equivalent to the Gini PCA when the variables are Gaussians. Also, it is shown that the generalized Gini PCA may be realized either in the space of the variables or in the space of the observations. In each case, some statistics are proposed to perform some interpretations of the variables and of the observations (absolute and relative contributions). To be precise, an $U$-statistics test is introduced to test for the significance of the correlations between the axes of the new subspace and the variables in order to assess their significance. Monte Carlo simulations are performed in order to show the superiority of the Gini PCA compared with the usual PCA when outlying observations contaminate the data. Finally, with the aid of the well-known cars data, which contain outliers, it is shown that the generalized Gini PCA leads to different results compared with the usual PCA.

The outline of the paper is as follows. Section (ref) sets the notations and presents some $\ell_2$ norm approaches of PCA. Section (ref) reviews the Gini-covariance operator. Section (ref) is devoted to the generalized Gini PCA. Section (ref) focuses on the interpretation of the Gini PCA. Sections (ref) and (ref) present some Monte Carlo simulations and applications, respectively.

Motivations for the use of Gini PCA

In this Section, the notations are set. Then, some assumptions are imposed and some $\ell_2$-norm PCA techniques are reviewed in order to motivate the employ of the Gini PCA.

Notations and definitions

Let $\mathbb{N}^{*}$ be the set of integers and $\mathbb{R}$ $[\mathbb{R}_{++}]$ the set of [positive] real numbers. Let $\mathcal{M}$ be the set of all $N\times K$ matrix $\boldsymbol{X}=[x_{ik}]$ that describes $N$ observations on $K$ dimensions such that $N \gg K$, with elements $x_{ik}\in \mathbb{R}$, and $\mathbb{I}_n$ the $n\times n$ identity matrix. The $N\times 1$ vectors representing each variable are expressed as $\mathbf{x}_{\cdot k}$, for all $k \in \{1,\ldots,K\}$ and we assume that $\mathbf{x}_{\cdot k} \neq c\mathbf{1}_{N}$, with $c$ a real constant and $\mathbf{1}_{N}$ a $N$-dimensional column vector of ones. The $K\times 1$ vectors representing each observation $i$ (the transposed $i$th line of $\boldsymbol{X}$) are expressed as $\mathbf{x}_{i\cdot}$, for all $i \in \{1,\ldots,N\}$. It is assumed that $\mathbf{x}_{\cdot k}$ is the realization of the random variable $X_k$, with cumulative distribution function $F_k$. The arithmetic mean of each column (line) of the matrix $\boldsymbol{X}$ is given by $\bar{\mathbf{x}}_{\cdot k}$ ($\bar{\mathbf{x}}_{i\cdot}$). The cardinal of set $A$ is denoted $\# \{A\}$. The $\ell_1$ norm, for any given real vector $\mathbf{x}$, is $\| \mathbf x \|_1 = \sum_{k=1}^K|x_{\cdot k}|$, whereas the $\ell_2$ norm is $\| \mathbf{x} \|_2 =(\sum_{k=1}^Kx_{\cdot k}^2)^{1/2}$.

assumption\ The random variables $X_k$ are such that $\mathbb E[|X_k|] < \infty$ for all $k \in \{1,\ldots,K\}$, but no assumption is made on the second moments (that may not exist).

This assumption imposes less structure compared with the classical PCA in which the existence of the second moments are necessary, as can be seen in the next subsection.

Variants of PCA based on the $\ell_2$ norm

The classical formulation of the PCA, to obtain the first component, can be obtained by solving

equation[equation omitted — 170 chars of source]

or equivalently

equation[equation omitted — 176 chars of source]

where $\omega\in\mathbb{R}^K$, and $\boldsymbol{\Sigma}$ is the (symmetric positive semi-definite) $K\times K$ sample covariance matrix. Mardia suggest to write$$ \omega_1^*\in\text{argmax}\left\lbrace\sum_{j=1}^K\text{Var}[\mathbf{x}_{\cdot, j}]\cdot\text{Cor}[\mathbf{x}_{\cdot,j},\boldsymbol{X}\omega]\right\rbrace \text{ subject to }\|\omega\|_2^2=\omega^{\top}\omega=1.$$ With scaled variables\footnote{In most cases, PCA is performed on scaled (and centered) variables, otherwise variables with large scales might alter interpretations. Thus, it will make sense, later on, to assume that components of $\boldsymbol{X}$ have identical distributions. At least the first two moments will be equal.} (i.e. $\text{Var}[\mathbf{x}_{\cdot, j}]=1$, $\forall j$)

equation[equation omitted — 209 chars of source]

Then a Principal Component Pursuit can start: we consider the `residuals', $\boldsymbol{X}_{(1)}=\boldsymbol{X}-\boldsymbol{X}\omega_1^*\omega_1^{*\top}$, its covariance matrix $\boldsymbol{\Sigma}_{(1)}$, and we solve $$ \omega_2^*\in\text{argmax}\left\lbrace\omega^{\top}\boldsymbol{\Sigma}_{(1)}\omega\right\rbrace \text{ subject to }\|\omega\|_2^2=\omega^{\top}\omega=1.$$ The part $\boldsymbol{X}\omega_1^*\omega_1^{*\top}$ is actually a constraint that we add to ensure the orthogonality of the two first components. This problem is equivalent to finding the maxima of $\text{Var}[\boldsymbol{X}\omega]$ subject to $\|\omega\|_2^2=1$ and $\omega\perp\omega_1^*$. This idea is also called Hotelling (or Wielandt) deflation technique. On the $k$-th iteration, we extract the leading eigenvector $$ \omega_k^*\in\text{argmax}\left\lbrace\omega^{\top}\boldsymbol{\Sigma}_{(k-1)}\omega\right\rbrace \text{ subject to }\|\omega\|_2^2=\omega^{\top}\omega=1,$$ where $\boldsymbol{\Sigma}_{(k-1)}=\boldsymbol{\Sigma}_{(k-2)}-\omega_{k-1}^{*}\omega_{k-1}^{*\top}\boldsymbol{\Sigma}_{(k-1)}\omega_{k-1}^{*}\omega_{k-1}^{*\top}$ (see e.g. Saad). Note that, following Hotelling and Eckart, that it is also possible to write this problem as $$ \min\left\lbrace\|\boldsymbol{X}-\tilde{\boldsymbol{X}}\|_\star\right\rbrace \text{ subject to }\text{rank}[\tilde{\boldsymbol{X}}]\leq k $$ where $\|\cdot\|_\star$ denotes the nuclear norm of a matrix (i.e. the sum of its singular values)\footnote{but other norms have also been considered in statistical literature, such as the Froebenius norm in the Eckart-Young theorem, or the maximum of singular values -- also called $2$-(induced)-norm.}.

One extension, introduced in dAspremont, was to add a constraint based on the cardinality of $\omega$ (also called $\ell_0$ norm) corresponding to the number of non-zero coefficients of $\omega$. The penalized objective function is then $$ \max\left\lbrace\omega^{\top}\boldsymbol{\Sigma}\omega-\lambda\text{card}[\omega]\right\rbrace \text{ subject to }\|\omega\|_2^2=\omega^{\top}\omega=1, $$ for some $\lambda>0$. This is called {\em sparse PCA}, and can be related to sparse regression, introduced in Tibshirani. But as pointed out in Mackey, interpretation is not easy and the components obtained are not orthogonal. Gorban considered an extension to nonlinear Principal Manifolds to take into account nonlinearities.

Another direction for extensions was to consider Robust Principal Component Analysis. Candes suggested an approach based on the fact that principal component pursuit can be obtained by solving $$ \min\left\lbrace\|\boldsymbol{X}-\tilde{\boldsymbol{X}}\|_\star+\lambda\|\tilde{\boldsymbol{X}}\|_1\right\rbrace.$$ But other methods were also considered to obtain Robust PCA. A natural `scale-free' version is obtained by considering a rank matrix instead of $\boldsymbol{X}$. This is also called `ordinal' PCA in the literature, see Korhonen. The first `ordinal' component is

equation[equation omitted — 189 chars of source]

where $\mathcal{R}$ denotes some rank based correlation, e.g. Spearman's rank correlation, as an extention of Equation ((ref)). So, quite naturally, one possible extension of Equation ((ref)) would be $$ \omega_1^*\in\text{argmax}\left\lbrace\omega^{\top}\mathcal{R}[\boldsymbol{X}]\omega\right\rbrace \text{ subject to }\|\omega\|_2^2=\omega^{\top}\omega=1$$ where $\mathcal{R}[\boldsymbol{X}]$ denotes Spearman's rank correlation. In this section, instead of using Pearson's correlation (as in Equation ((ref)) when the variables are scaled) or Spearman's (as in this ordinal PCA), we will consider the multidimensional Gini correlation based on the $h$-covariance operator.

Geometry of Gini PCA: Gini-Covariance Operators

The first PCA was introduced by Pearson, projecting $\boldsymbol{X}$ onto the eigenvectors of its covariance matrix, and observing that the variances of those projections are the corresponding eigenvalues. One of the key property is that $\boldsymbol{X}^\top\boldsymbol{X}$ is a positive matrix. Most statistical properties of PCAs (see FluryRiedwyl or Anderson) are obtained under Gaussian assumptions. Furthermore, geometric properties can be obtained using the fact that the covariance defines an inner product on the subspace of random variables with finite second moment (up to a translation, i.e. we identify any two that differ by a constant).

We will discuss in this section the properties of the Gini Covariance operator with the special case of Gaussian random variables, and the property of the Gini correlation matrix that will be used in the next Section for the Gini PCA.

The Gini-covariance operator

In this section, $\boldsymbol{X}=(X_1,\cdots,X_K)$ denotes a random vector. The covariance matrix between $\boldsymbol{X}$ and $\boldsymbol{Y}$, two random vectors, is defined as the inner product between centered versions of the vectors,

equation[equation omitted — 243 chars of source]

Hence, it is the matrix where elements are regular covariances between components of the vectors, $\text{Cov}(\boldsymbol{X},\boldsymbol{Y})=[\text{Cov}(X_i,Y_j)]$. It is the upper-right block of the covariance matrix of $(\boldsymbol{X},\boldsymbol{Y})$. Note that $\text{Cov}(\boldsymbol{X},\boldsymbol{X})$ is the standard variance-covariance matrix of vector $\boldsymbol{X}$.

definitionLet $\boldsymbol{X}=(X_1,\cdots,X_{K})$ be collections of $K$ identically distributed random variables. Let $h:\mathbb{R}\rightarrow\mathbb{R}$ denote a non-decreasing function. Let $h(\boldsymbol{X})$ denote the random vector $(h(X_1),\cdots,h(X_K))$, and assume that each component has a finite variance. Then, operator $\Gamma C_{h}(\boldsymbol{X})=\emph{Cov}(\boldsymbol{X},h(\boldsymbol{X}))$ is called $h$-Gini covariance matrix.

Since $h$ is a non-decreasing mapping, then $\boldsymbol{X}$ and $h(\boldsymbol{X})$ are componentwise comonotonic random vectors. Assuming that components of $\boldsymbol{X}$ are identically distributed is a reasonable assumption in the context of scaled (and centered) PCA, as discussed in footnote (ref). Nevertheless, a stronger technical assumption will be necessary: pairwise-exchangeability.

definition$\boldsymbol{X}$ is said to be pairwise-exchangeable if for all pair $(i,j)\in\{1,\cdots,K\}^2$, $(X_i,X_j)$ is exchangeable, in the sense that $(X_i,X_j)\overset{\mathcal{L}}{=}(X_j,X_i)$.

Pairwise-exchangeability is a stronger concept than having only one vector with identically distributed components, and a weaker concept than (full) exchangeability. In the Gaussian case where $h(X_k) = \Phi(X_k)$ with $\Phi(X_k)$ being the normal cdf of $X_k$ for all $k=1,\ldots,K$, pairwise-exchangeability is equivalent to components identically distributed.

propositionIf $\boldsymbol{X}$ is a Gaussian vector with identically distributed components, then $\boldsymbol{X}$ is pairwise-exchangeable.
proofFor simplicity, assume that components of $\boldsymbol{X}$ are $\mathcal{N}(0,1)$ random variables, then $\boldsymbol{X}\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{\rho})$ where $\boldsymbol{\rho}$ is a correlation matrix. In that case $$ \left[ \begin{matrix} X_i \\ X_j \end{matrix} \right]\sim\mathcal{N}\left( \left[ \begin{matrix} 0 \\ 0 \end{matrix} \right], \left[ \begin{matrix} 1 & \rho_{ij} \\ \rho_{ji} & 1 \\ \end{matrix} \right] \right),$$ with Pearson correlation $\rho_{ij}=\rho_{ji}$, thus $(X_i,X_j)$ is exchangeable.

Let us now introduce the Gini-covariance. Gini12 introduced the Gini mean difference operator $\Delta$, defined as:

equation[equation omitted — 158 chars of source]

for some random variable $X$ (or more specifically for some distribution $F$ with $X\sim F$, because this operator is law invariant). One can rewrite: $$ \Delta(X)=4\text{Cov}[X,F(X)]=\frac{1}{3}\frac{\text{Cov}[X,F(X)]}{\text{Cov}[F(X),F(X)]} $$ where the term on the right is interpreted as the slope of the regression curve of the observed variable $X$ and its `ranks' (up to a scaling coefficient). Thus, the Gini-covariance is obtained when the function $h$ is equal to the cumulative distribution function of the second term, see Schechtman87.

definitionLet $\boldsymbol{X}=(X_1,\cdots,X_{K})$ be a collection of $K$ identically distributed random variables, with cumulative distribution function $F$. Then, the Gini covariance is $\Gamma C_F(\boldsymbol{X})=\emph{\text{Cov}}(\boldsymbol{X},F(\boldsymbol{X}))$.

On this basis, it is possible to show that the Gini covariance matrix is a positive semi-definite matrix.

theoremLet $Z \sim \mathcal{N}(0,1)$. If $\boldsymbol{X}$ represents identically distributed Gaussian random variables, with distribution $\mathcal{N}(\mu,\sigma^2)$, then the two following assertions hold: (i) $\Gamma C_F(\boldsymbol{X})=\sigma^{-1} \emph{Cov}(Z,\Phi(Z)) \emph{Var}(\boldsymbol{X}).$ (ii) $\Gamma C_F(\boldsymbol{X})$ is a positive-semi definite matrix.
proof(i) In the Gaussian case, if $h$ is the cumulative distribution function of the $X_k$'s, then $\text{Cov}(X_k,h(X_\ell))=r\sigma\cdot\text{Cov}(Z,\Phi(Z))$, where $\Phi$ is the normal cdf, see Yitzhaki13, Chapter 3. Observe that $\text{Cov}(X_k,h(X_k))=\sigma\cdot \text{Cov}(Z,\Phi(Z))$, if $h$ is the cdf of $X_k$. Thus, $\lambda := \text{Cov}(Z,\Phi(Z))$ yields: $$\Gamma C_F((X_k,X_\ell))=\lambda \left[ \begin{matrix} \sigma & \rho \sigma \\ \rho\sigma & \sigma \\ \end{matrix} \right]= \lambda \sigma \left[ \begin{matrix} 1 & \rho \\ \rho & 1\\ \end{matrix} \right]=\frac{\lambda}{\sigma} \text{Var}((X_k,X_\ell)).$$ (ii) We have $\text{Cov}(Z,\Phi(Z)) \geq 0$, then it follows that $C_F((X_k,X_\ell)) \geq 0$: $$ \mathbf{x}^{\top}C_F(\boldsymbol{X}) \mathbf{x} = \mathbf{x}^{\top}\frac{\text{Cov}(Z,\Phi(Z))}{\sigma}\text{Var}(\boldsymbol{X}) \mathbf{x} \geq 0, $$ which ends the proof.

Note that $\Gamma C_F(\boldsymbol{X})=\text{Cov}(\boldsymbol{X},-\overline{F}(\boldsymbol{X}))=\Gamma C_{-\overline{F}}(\boldsymbol{X})$, where $\overline{F}$ denotes the survival distribution function.

definitionLet $\boldsymbol{X}=(X_1,\cdots,X_{K})$ be a collection of $K$ identically distributed random variables, with survival distribution function $\overline{F}$. Then, the generalized Gini covariance is $G\Gamma C_{\nu}(\boldsymbol{X})=\Gamma C_{-\overline{F}^{\nu-1}}(\boldsymbol{X})=\text{\em Cov}(\boldsymbol{X},-\overline{F}^{\nu-1}(\boldsymbol{X}))$, for $\nu>1$.

This operator is related to the one introduced in Schechtman03, called generalized Gini mean difference $GMD_\nu$ operator. More precisely, an estimator of the generalized Gini mean difference is given by: \[ GMD_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k}) := -\frac{2}{N-1}\nu \text{Cov}(\mathbf{x}_{\cdot\ell},\textsc{\textbf{r}}_{\mathbf{x}_{\cdot k}}^{\nu-1}) , \ \nu > 1, \] where $\textsc{\textbf{r}}_{\mathbf{x}_{\cdot k}} = (R(x_{1k}),\ldots,R(x_{nk}))$ is the decumulative rank vector of $\mathbf{x}_{\cdot k}$, that is, the vector that assigns the smallest value (1) to the greatest observation $x_{ik}$, and so on. The rank of observation $i$ with respect to variable $k$ is: \[ R(x_{ik}) := \left\{

array[array omitted — 147 chars of source]

\right. \] Hence $GMD_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k})$ is the empirical version of $$ 2\nu \Gamma C_\nu(X_{\ell},X_{ k}) := -2\nu\text{Cov}\big(X_\ell,\overline{F}_k(X_k)^{\nu-1}\big). $$ The index $GMD_\nu$ is a generalized version of the $GMD_2$ proposed earlier by Schechtman87, and can also be written as: $$ GMD_2(X_{k},X_{k}) = 4\text{Cov}\big(X_k,F_k(X_k)\big) = \Delta(X_k). $$ When $k=\ell$, $GMD_\nu$ represents the variability of the variable $\mathbf{x}_{\cdot k}$ itself. Focus is put on the lower tail of the distribution $\mathbf{x}_{\cdot k}$ whenever $\nu \to \infty$, the approach is said to be max-min in the sense that $GMD_\nu$ inflates the minimum value of the distribution. On the contrary, whenever $\nu \to 0$, the approach is said to be max-max, in this case focus is put on the upper tail of the distribution $\mathbf{x}_{\cdot k}$. As mentioned in Yitzhaki13, the case $\nu < 1$ does not entail simple interpretations, thereby the parameter $\nu$ is used to be set as $\nu > 1$ in empirical applications.\footnote{In risk analysis $\nu \in (0,1)$ denotes risk lover decision makers (max-max approach), whereas $\nu > 1$ stands for risk averse decision makers, and $\nu \to \infty$ extreme risk aversion (max-min approach).}

Note that even if $X_k$ and $X_\ell$ have the same distribution, we might have $GMD_\nu(X_k,X_\ell)\neq GMD_\nu(X_\ell,X_k)$, as shown on the example of Figure (ref). In that case $\mathbb{E}[X_kh(X_\ell)]\neq \mathbb{E}[X_\ell h(X_k)]$ if $h(2)\neq 2h(1)$ (this property is nevertheless valid if $h$ is linear). We would have $GMD_\nu(X_k,X_\ell)= GMD_\nu(X_\ell,X_k)$ when $X_k$ and $X_\ell$ are exchangeable. But since generally $GMD_\nu$ is not symmetric, we have for $\mathbf{x}_{\cdot k}$ being not a monotonic transformation of $\mathbf{x}_{\cdot\ell}$ and $\nu > 1$, $GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot\ell})\neq GMD_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k})$.

figure[figure omitted — 345 chars of source]

Generalized Gini correlation

In this section, $\boldsymbol{X}$ is a matrix in $\mathcal{M}$. The Gini correlation coefficient ($G$-correlation frown now on), is a normalized $GMD_\nu$ index such that for all $\nu > 1$, see Schechtman03, \[ GC_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k}) := \frac{GMD_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k})}{GMD_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot\ell})} \ ; \ GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot\ell}) := \frac{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot\ell})}{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})}, \] with $GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k}) = 1$ and $GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k}) \neq 0$, for all $k,\ell=1,\ldots,K$. Following Schechtman03, the $G$-correlation is well-suited for the measurement of correlations between non-normal distributions or in the presence of outlying observations in the sample.

property-- Schechtman and Yitzhaki (2013): (i) $GC_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k}) \leq 1$. (ii) If the variables $\mathbf{x}_{\cdot\ell}$ and $\mathbf{x}_{\cdot k}$ are independent, for all $k\neq \ell$, then $GC_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k}) = GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot\ell}) =0$. (iii) For any given monotonic increasing transformation $\varphi$, $GC_\nu(\mathbf{x}_{\cdot\ell},\varphi(\mathbf{x}_{\cdot k})) = GC_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k})$. (iv) If $(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k})$ have a bivariate normal distribution with Pearson correlation $\rho$, then $GC_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k}) = GC_\nu(\mathbf{x}_{\cdot k}, \mathbf{x}_{\cdot\ell}) = \rho$. \emph{(v)} If $\mathbf{x}_{\cdot k}$ and $\mathbf{x}_{\cdot\ell}$ are exchangeable up to a linear transformation, then $GC_\nu(\mathbf{x}_{\cdot\ell},\mathbf{x}_{\cdot k}) = GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot\ell})$.

Whenever $\nu \to 1$, the variability of the variables is attenuated so that $GMD_\nu$ tends to zero (even if the variables exhibit a strong variance). The choice of $\nu$ is interesting to perform generalized Gini PCA with various values of $\nu$ in order to robustify the results of the PCA, since the standard PCA (based on the variance) is potentially of bad quality if outlying observations drastically affect the sample.

A $G$-correlation matrix is proposed to analyze the data into a new vector space. Following Property (ref) (iv), it is possible to rescale the variables $\mathbf{x}_{\cdot \ell}$ thanks to a linear transformation, then the matrix of standardized observation is,

equation[equation omitted — 187 chars of source]

The variable $z_{i\ell}$ is a real number without dimension. The variables $\mathbf{x}_{\cdot k}$ are rescaled such that their Gini variability is equal to unity. Now, we define the $N\times K$ matrix of decumulative centered rank vectors of $\boldsymbol{Z}$, which are the same compared with those of $\boldsymbol{X}$: \[ \boldsymbol{R}_\mathbf{z}^c \equiv [R^c(z_{i\ell})] := [R(z_{i\ell})^{\nu-1} - \bar{\textsc{\textbf{r}}}^{\nu-1}_{\mathbf{z}_{\cdot\ell}}] = [R(x_{i\ell})^{\nu-1} - \bar{\textsc{\textbf{r}}}_{\mathbf{x}_{\cdot\ell}}^{\nu-1}]. \] Note that the last equality holds since the standardization ((ref)) is a strictly increasing affine transformation.\footnote{By definition $GMD_\nu(\mathbf{x}_{\cdot \ell},\mathbf{x}_{\cdot \ell}) \geq 0$ for all $\ell=1,\ldots,K$. As we impose that $\mathbf{x}_{\cdot \ell} \neq c\mathbf{1}_N$, the condition becomes $GMD_\nu(\mathbf{x}_{\cdot \ell},\mathbf{x}_{\cdot \ell}) >0$.} The $K\times K$ matrix containing all $G$-correlation indices between all couples of variables $\mathbf{z}_{\cdot k}$ and $\mathbf{z}_{\cdot \ell}$, for all $k,\ell=1,\ldots,K$ is expressed as:

equation[equation omitted — 117 chars of source]

Indeed, if $GMD_\nu(\boldsymbol{Z}) \equiv [GMD_\nu(\mathbf{z}_{\cdot k},\mathbf{z}_{\cdot \ell})]$, then we get the following.

propositionFor each standardized matrix $\boldsymbol{Z}$ defined in ((ref)), the following relations hold: \begin{align} & GMD_\nu(\boldsymbol{Z}) = GC_\nu(\boldsymbol{X}) = GC_\nu(\boldsymbol{Z}).\\ & GMD_\nu(\mathbf{z}_{\cdot k},\mathbf{z}_{\cdot k}) = 1, \ \forall k=1,\ldots,K. \end{align}
proofWe have $GMD_\nu(\boldsymbol{Z}) \equiv [GMD_\nu(\mathbf{z}_{\cdot k},\mathbf{z}_{\cdot \ell})]$ being a $K\times K$ matrix. The extra diagonal terms may be rewritten as, \begin{align} GMD_\nu(\mathbf{z}_{\cdot k},\mathbf{z}_{\cdot \ell}) &= -\frac{2}{N-1}\nu Cov(\mathbf{z}_{\cdot k},r_{\mathbf{z}_{\cdot \ell}}^{\nu-1}) \notag \\ &= -\frac{2}{N-1}\nu Cov \left( \frac{\mathbf{x}_{\cdot k} - \bar{\mathbf{x}}_{\cdot k}\mathbf{1}_N }{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})},r_{\mathbf{z}_{\cdot \ell}}^{\nu-1} \right) \notag \\ & = -\frac{2}{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})}\left[ \frac{\nu \text{Cov}( \mathbf{x}_{\cdot k}, \textsc{\textbf{r}}_{\mathbf{z}_{\cdot \ell}}^{\nu-1})}{N-1} - \frac{\nu \text{Cov}( \bar{\mathbf{x}}_{\cdot k}\mathbf{1}_N,\textsc{\textbf{r}}_{\mathbf{z}_{\cdot \ell}}^{\nu-1})}{N-1}\right] \notag \\ &= \frac{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot \ell})}{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})} = GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot \ell}).\notag \end{align} Finally, using the same approach as before, we get: \begin{align} GMD_\nu(\mathbf{z}_{\cdot k},\mathbf{z}_{\cdot k}) &= -\frac{2}{N-1} \nu\text{Cov}(\mathbf{z}_{\cdot k},\textsc{\textbf{r}}_{\mathbf{z}_{\cdot k}}^{\nu-1}) \notag \\ &= \frac{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})}{GMD_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})} \notag \\ &= GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot k})=1.\notag \end{align} By Property (ref) (iv), since $\textsc{\textbf{r}}_{\mathbf{x}_{\cdot k}} = \textsc{\textbf{r}}_{\mathbf{z}_{\cdot k}}$, then $GC_\nu(\mathbf{x}_{\cdot k},\mathbf{x}_{\cdot \ell}) = GC_\nu(\mathbf{z}_{\cdot k},\mathbf{z}_{\cdot \ell})$. Thus, \[ GMD_\nu(\boldsymbol{Z}) = -\frac{2\nu}{N(N-1)} \boldsymbol{Z}^{\top}\boldsymbol{R}_\mathbf{z}^c = GC_\nu(\boldsymbol{X}) = GC_\nu(\boldsymbol{Z}), \] which ends the proof.

Finally, under a normality assumption, the generalized Gini covariance matrix $GC_\nu(\boldsymbol{X}) \equiv [GMD_\nu(X_{ k},X_{\ell})]$ is shown to be a positive semi-definite matrix.

theoremLet $Z \sim \mathcal{N}(0,1)$. If $\boldsymbol{X}$ represents identically distributed Gaussian random variables, with distribution $\mathcal{N}(\mu,\sigma^2)$, then the two following assertions holds: (i) $GC_\nu(\boldsymbol{X})= \sigma^{-1}\emph{Cov}(Z,\Phi(Z)) \emph{Var}(\boldsymbol{X}).$ (ii) $GC_\nu(\boldsymbol{X})$ is a positive semi-definite matrix.
proofThe first part (i) follows from Yitzhaki13, Chapter 6. The second part follows directly from (i).

Theorem (ref) shows that under the normality assumption, the variance is a special case of the Gini methodology. As a consequence, for multivariate normal distributions, it is shown in Section (ref) that Gini PCAs and classical PCA (based on the $\ell_2$ norm and the covariance matrix) are equivalent.

Generalized Gini PCA

In this section, the multidimensional Gini variability of the observations $i=1,\ldots,N$, embodied by the matrix $GC_\nu(\boldsymbol{Z})$, is maximized in the $\mathbb{R}^{K}$-Euclidean space, i.e., in the set of variables $\{\mathbf{z}_{\cdot 1},\ldots,\mathbf{z}_{\cdot K}\}$. This allows the observations to be projected onto the new vector space spanned by the eigenvectors of $GC_\nu(\boldsymbol{Z})$. Then, the projection of the variables is investigated in the $\mathbb{R}^{N}$-Euclidean space induced by $GC_\nu(\boldsymbol{Z})$. Both observations and variables are analyzed through the prism of absolute and relative contributions to propose relevant interpretations of the data in each subspace.

The $\mathbb{R}^{K}$-Euclidean space

It is possible to investigate the projection of the data $\boldsymbol{Z}$ onto the new vector space induced by $GMD_\nu(\boldsymbol{Z})$ or alternatively by $GC_\nu(\boldsymbol{Z})$ since $GMD_\nu(\boldsymbol{Z}) = GC_\nu(\boldsymbol{Z})$. Let $\mathbf{f}_{\cdot k}$ be the $k$th principal component, i.e. the $k$th axis of the new subspace, such that the $N\times K$ matrix $\boldsymbol{F}$ is defined by $\boldsymbol{F} \equiv [\mathbf{f}_{\cdot 1},\ldots,\mathbf{f}_{\cdot K}]$ with $\boldsymbol{R}_\mathbf{f}^c \equiv [\textsc{\textbf{r}}_{c,\mathbf{f}_1}^{\nu-1},\ldots,\textsc{\textbf{r}}_{c,\mathbf{f}_{\cdot K}}^c]$ its corresponding decumulative centered rank matrix (where each decumulative rank vector is raised to an exponent of $\nu-1$). The $K\times K$ matrix $\boldsymbol{B} \equiv [\mathbf{b}_{\cdot 1},\ldots,\mathbf{b}_{\cdot K}]$ is the projector of the observations, with the normalization condition $\mathbf{b}_{\cdot k}^{\top}\mathbf{b}_{\cdot k} = 1$, such that $\boldsymbol{F}=\boldsymbol{Z}\boldsymbol{B}$. We denote by $\lambda_{\cdot k}$ (or $2\mu_{\cdot k}$) the eigenvalues of the matrix $[GC_\nu(\boldsymbol{Z})+GC_\nu(\boldsymbol{Z})^{\top}]$. Let the basis $\mathscr{B}:=\{\mathbf{b}_{\cdot1},\ldots,\mathbf{b}_{\cdot h}\}$ with $h\leq K$ issued from the maximization of the overall Gini variability: \[ \max \mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k} \ \Longrightarrow \ [GC_\nu(\boldsymbol{Z})+GC_\nu(\boldsymbol{Z})^{\top}] \mathbf{b}_{\cdot k} = 2\mu_{\cdot k} \mathbf{b}_{\cdot k}, \ \forall k=1,\ldots,K. \] Indeed, from the Lagrangian, \[ L = \mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k} - \mu_{\cdot k} [1- \mathbf{b}_{\cdot k}^{\top}\mathbf{b}_{\cdot k}], \] because of the non-symmetry of $GC_\nu(\boldsymbol{Z})$, the eigenvalue equation is, \[ [GC_\nu(\boldsymbol{Z}) + GC_\nu(\boldsymbol{Z})^{\top}]\mathbf{b}_{\cdot k} = 2\mu_{\cdot k}\mathbf{b}_{\cdot k}, \] that is,

equation[equation omitted — 153 chars of source]

The new subspace $\{\mathbf{f}_{\cdot 1},\ldots,\mathbf{f}_{\cdot h}\}$ such that $h\leq K$ is issued from the maximization of the Gini variability between the observations on each axis $\mathbf{f}_{\cdot k}$. Although the result of the generalized Gini PCA seems to be close to the classical PCA, some differences exist.

propositionLet $\mathscr{B} = \{\mathbf{b}_{\cdot 1},\ldots,\mathbf{b}_{\cdot h}\}$ with $h\leq K$ be the basis issued from the maximization of $\mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k}$ for all $k=1,\ldots,K$, then the following assertions hold: (i) $\max GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = \mu_{\cdot k}$ for all $k=1,\ldots,K$, if and only if $\textsc{\textbf{r}}_{c,\mathbf{f}_{\cdot k}}^{\nu-1} = \boldsymbol{R}_\mathbf{z}^c \mathbf{b}_{\cdot k}$. (ii) $\mathbf{b}_{\cdot k} \mathbf{b}_{\cdot h}^{\top} = 0$, for all $k\neq h$. (iii) $\mathbf{b}_{\cdot k} \mathbf{b}_{\cdot k}^{\top} = 1$, for all $k=1,\ldots,K$.
proof(i) Note that $GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = -\frac{2\nu}{N(N-1)}(\mathbf{f}_{\cdot k} - \bar{\mathbf{f}})^{\top}\textsc{\textbf{r}}_{c,\mathbf{f}_{\cdot k}}^{\nu-1}$, where $\textsc{\textbf{r}}_{c,\mathbf{f}_{\cdot k}}^{\nu-1}$ is the $k$th column of the centered (decumulative) rank matrix $\boldsymbol{R}_\mathbf{f}^c$. Since $\mathbf{f}_{\cdot k} = \boldsymbol{Z} \mathbf{b}_{\cdot k}$ and $\bar{\mathbf{f}} = (\bar{\mathbf{f}}_{\cdot 1},\ldots,\bar{\mathbf{f}}_{\cdot K}) = \mathbf{0}$ then:\footnote{We have: \[ \bar{\mathbf{f}}_{\cdot k} = 1/N\sum_{i=1}^N f_{ik} = 1/N \left[ \sum_{i=1}^N \mathbf{z}^{\top}_{i\cdot} \mathbf{b}_{\cdot k} \right] = 1/N\left[ \sum_{i=1}^N \mathbf{z}^{\top}_{i\cdot} \right]\mathbf{b}_{\cdot k} = \mathbf{0}.\]} \begin{align} \mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k} &= -\frac{2\nu}{N(N-1)}\mathbf{b}_{\cdot k}^{\top} \boldsymbol{Z}^{\top} \boldsymbol{R}_{\mathbf{z}}^c \mathbf{b}_{\cdot k} \\ &= -\frac{2\nu}{N(N-1)}\mathbf{b}_{\cdot k}^{\top} \boldsymbol{Z}^{\top} r_{c,\mathbf{f}_{\cdot k}}^{\nu-1} \ \ \ \ (by \ r_{c,\mathbf{f}_{\cdot k}}^{\nu-1} = \boldsymbol{R}_\mathbf{z}^c \mathbf{b}_{\cdot k}) \notag \\ &= -\frac{2\nu}{N(N-1)}\mathbf{f}^{\top}_{\cdot k} \textbf{r}_{c,\mathbf{f}_{\cdot k}}^{\nu-1} \notag \\ &= GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}).\notag \end{align} Then, maximizing the multidimensional variability $\mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k}$ yields from ((ref)): \begin{align} \mathbf{b}_{\cdot k}^{\top} [GC_\nu(\boldsymbol{Z})+GC_\nu(\boldsymbol{Z})^{\top}]\mathbf{b}_{\cdot k} &= \mathbf{b}_{\cdot k}^{\top} \lambda_{\cdot k}\mathbf{b}_{\cdot k}\notag \\ \Longleftrightarrow \ \mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z})\mathbf{b}_{\cdot k} + \mathbf{b}_{\cdot k}^{\top}GC_\nu(\boldsymbol{Z})^{\top} \mathbf{b}_{\cdot k} &= \lambda_{\cdot k}.\notag \end{align} Since $\mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k} = \mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z})^{\top} \mathbf{b}_{\cdot k}$, then \[ \mathbf{b}_{\cdot k}^{\top}GC_\nu(\boldsymbol{Z})^{\top} \mathbf{b}_{\cdot k} = \lambda_{\cdot k}/2 = \mu_{\cdot k}, \] and so $GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = \lambda_{\cdot k}$ for all $k=1,\ldots,K$. The results (ii) and (iii) are straightforward.

Discussion

Condition (i) shows that the maximization of the multidimensional variability (in the Gini sense) $\mathbf{b}_{\cdot k}^{\top} GC_\nu(\boldsymbol{Z}) \mathbf{b}_{\cdot k}$ does not necessarily coincide with the maximization of the variability of the observations projected onto the new axis $\mathbf{f}_{\cdot k}$ embodied by $GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})$. Since in general, the rank of the observations on axis $\mathbf{f}_{\cdot k}$ does not coincide with the projected ranks, that is, \[ \textsc{\textbf{r}}_{c,\mathbf{f}_{\cdot k}}^{\nu-1} \neq \boldsymbol{R}_\mathbf{z}^c \mathbf{b}_{\cdot k}, \] then, \[ \max \mathbf{b}_{\cdot k}^{\top} GC(\boldsymbol{Z}) \mathbf{b}_{\cdot k} \neq GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}). \] In other words, maximizing the quadratic form $\mathbf{b}_{\cdot k}^{\top} GC(\boldsymbol{Z})_\nu \mathbf{b}_{\cdot k}$ does not systematically maximize the overall Gini variability $GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})$. However, it maximizes the following generalized Gini index:

align[align omitted — 328 chars of source]

In the literature on inequality indices, this kind of index is rather known as a generalized Gini index, because of the product between a variable $\mathbf{f}_{\cdot k}$ and a function $\Psi$ of its ranks, $\Psi (\textsc{\textbf{r}}_{\mathbf{f}_{k}}) := \mathbf{b}_{\cdot k}^{\top} (\boldsymbol{R}_{\mathbf{z}}^{c})^{\top}$, such that: \[ GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = -\frac{2\nu}{N(N-1)}\mathbf{f}_{\cdot k} \ \Psi (\textsc{\textbf{r}}_{\mathbf{f}_{k}}). \] Yaari1987 and subsequently Yaari1988 proposes generalized Gini indices with a rank distortion function $\Psi$ that describes the behavior of the decision maker (being either max-min or max-max).\footnote{Strictly speaking Yaari1987 and Yaari1988 suggests probability distortion functions $\Psi: [0,1] \rightarrow [0,1]$, which does not necessarily coincide to our case.}

It is noteworthy that this generalized Gini index of variability is very different from Banerjee's multidimensional Gini index. The author proposes to extract the first eigenvector $\mathbf{e}_{\cdot 1}$ of $\boldsymbol{X}^{\top}\boldsymbol{X}$ and to project the data $\boldsymbol{X}$ such that $\mathbf{s}:=\boldsymbol{X}\mathbf{e}_{\cdot 1}$ so that the multidimensional Gini index is $G(\mathbf{s}) = \mathbf{s}^{\top} \tilde{\Psi}(\textsc{\textbf{r}}_s)$, with $\textsc{\textbf{r}}_s$ the rank vector of $\mathbf{s}$ and with $\tilde{\Psi}$ a function that distorts the ranks. Banerjee's index is derived from the matrix $\boldsymbol{X}^{\top}\boldsymbol{X}$. To be precise, the maximization of the variance-covariance matrix $\boldsymbol{X}^{\top}\boldsymbol{X}$ (based on the $\ell_2$ metric) yields the projection of the data on the first component $\mathbf{f}_{\cdot 1}$, which is then employed in the multidimensional Gini index (based on the $\ell_1$ metric). This approach is legitimated by the fact that $G(\mathbf{s})$ has some desirable properties linked with the Gini index. However, this Gini index deals with an information issued from the variance, because the vector $\mathbf{s}$ relies on the maximization of the variance of component $\mathbf{f}_{\cdot 1}$. Alternatively, it is possible to make use of the Gini variability, in a first stage, in order to project the data onto a new subspace, and in a second stage, to use the generalized Gini index of the projected data for the interpretations. In such as case, the Gini metric enables outliers to be attenuated. The employ of $G(\mathbf{s})$ as a result of the variance-covariance maximization may transform the data so that outlying observations would capture an important part of the information (variance) on the first component. This case occurs in the classical PCA. This fact will be proven in the next sections with Monte Carlo simulations. Let us before investigate the employ of the generalized Gini index ${GGMD}_{\nu}$.

Properties of ${GGMD}_{\nu}$

Since the Gini PCA relies on the generalized Gini index ${GGMD}_{\nu}$, let us explore its properties.

propositionLet the eigenvalues of $GC_{\nu}(\boldsymbol{Z})+GC_{\nu}(\boldsymbol{Z})^{\top}$ be such that $\lambda_1 = \mu_1/2 \geq \cdots \geq \lambda_{\cdot K} = \mu_{\cdot K}/2$. Then, (i) $GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot \ell}) = GGMD_\nu(\mathbf{f}_{\cdot \ell},\mathbf{f}_{\cdot k}) = 0, \ \text{for all } \ell=1,\ldots,K$, if and only if, $\lambda_{\cdot k} =0$. (ii) $\max_{k=1,\ldots,K} {GGMD}_{\nu}(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = \mu_1$. (iii) $\min_{k=1,\ldots,K} {GGMD}_{\nu}(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) = \mu_{\cdot K}$.
proof(i) The result comes from the rank-nullity theorem. From the eigenvalue Equation ((ref)), we have: \[ \mathbf{b}_{\cdot k}^{\top}GC_{\nu}(\boldsymbol{Z}) \mathbf{b}_{\cdot k} = \lambda_{\cdot k}/2 = \mu_{\cdot k}. \] Let $f$ be the linear application issued from the matrix $GC_{\nu}(\boldsymbol{Z})$. Whenever $\lambda_{\cdot k} = 0$, two columns (or rows) of $GC_{\nu}(\boldsymbol{Z})$ are collinear, then the dimension of the image set of $f$ is $\dim(f) = K-1$. Hence, $\mathbf{f}_{\cdot k}=\mathbf{0}$. Since $\mathbf{b}_{\cdot k}^{\top}GC_{\nu}(\boldsymbol{Z})^{\top} \mathbf{b}_{\cdot k} = {GGMD}_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})$ for all $k=1,\ldots,K$, then for $\lambda_{\cdot k}$ we get: \[ \mathbf{b}_{\cdot k}^{\top}GC_{\nu}(\boldsymbol{Z})^{\top} \mathbf{b}_{\cdot k} = {GGMD}_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k}) =\lambda_{\cdot k}/2 = \mu_{\cdot k} = 0. \] On the other hand, since $\mathbf{f}_{\cdot k} = \mathbf{0}$, it follows that $GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot \ell}) = 0$ for all $\ell=1,\ldots,K$. Also, if $\mathbf{f}_{\cdot k} = \mathbf{0}$ then the centered rank vector $\textsc{\textbf{r}}^c_{\mathbf{f}_{\cdot k}} = \mathbf{0}$, and so $GGMD_\nu(\mathbf{f}_{\cdot \ell},\mathbf{f}_{\cdot k}) = 0$ for all $\ell=1,\ldots,K$. (ii) The proof comes from the Rayleigh-Ritz identity: \[ \lambda_{\max} := \max \frac{\mathbf{b}_{\cdot 1}^{\top}[GC_\nu(\boldsymbol{Z})+GC_\nu(\boldsymbol{Z})^{\top}]\mathbf{b}_1}{\mathbf{b}_1^{\top}\mathbf{b}_1} = \lambda_{1}. \] Since $\mathbf{b}_1^{\top}GC_\nu(\boldsymbol{Z})\mathbf{b}_{\cdot 1} = \lambda_1/2$ and because $\mathbf{b}_1^{\top}GC_\nu(\boldsymbol{Z})\mathbf{b}_1 = GGMD_{\nu}(\mathbf{f}_1,\mathbf{f}_1)$, the result follows. (iii) Again, the Rayleigh-Ritz identity yields: \[ \lambda_{\min} := \min \frac{\mathbf{b}_{\cdot K}^{\top}[GC_\nu(\boldsymbol{Z})+GC_\nu(\boldsymbol{Z})^{\top}]\mathbf{b}_{\cdot K}}{\mathbf{b}_{\cdot K}^{\top}\mathbf{b}_{\cdot K}} = \lambda_{K}. \] Then, $\mathbf{b}_{\cdot K}^{\top}GC_\nu(\boldsymbol{Z})\mathbf{b}_{\cdot K} = GGMD_{\nu}(\mathbf{f}_{\cdot K},\mathbf{f}_{\cdot K}) = \lambda_{\cdot K}/2$.

The index $GGMD_{\nu}(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})$ represents the variability of the observations projected onto component $\mathbf{f}_{\cdot k}$. When this variability is null, then the eigenvalue is null (i). In the same time, there is neither co-variability in the Gini sense between $\mathbf{f}_{\cdot k}$ and another axis $\mathbf{f}_{\cdot \ell}$, that is $GGMD_{\nu}(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot \ell})=0$.

In the Gaussian case, because the Gini correlation matrix is positive semi-definite, the eigenvalues are non-negative, then $GGMD$ is null whenever it reaches its minimum.

propositionLet $Z \sim \mathcal N (0,1)$ and let $\boldsymbol{X}$ represent identically distributed Gaussian random variables, with distribution $\mathcal{N}(\mathbf{0},\boldsymbol{\rho})$ such that $\emph{Var}(X_k)=1$ for all $k=1,\ldots,K$ and let $\gamma_1,\ldots,\gamma_K$ be the eigenvalues of $\emph{Var}(\boldsymbol{X})$. Then the following assertions holds: (i) $\emph{Tr}[GC_\nu(\boldsymbol{X})] = \emph{Cov}(Z,\Phi(Z))\emph{Tr}[\emph{Var}(\boldsymbol{X})]$. (ii) $\mu_k = \emph{Cov}(Z,\Phi(Z)) \gamma_k$ for all $k=1,\ldots,K$. (iii) $|GC_\nu(\boldsymbol{X})| = \emph{Cov}^K(Z,\Phi(Z))|\emph{Var}(\boldsymbol{X})|$. (iv) For all $\nu > 1$: $$ \frac{\mu_k}{\emph{Tr}[GC_\nu(\boldsymbol{X})]} = \frac{\gamma_k}{\emph{Tr}[\emph{Var}(\boldsymbol{X})]}, \ \forall k=1,\ldots,K. $$
proofFrom Theorem (ref): $$GC_\nu(\boldsymbol{X})=\sigma^{-1} \text{Cov}(Z,\Phi(Z)) \emph{Var}(\boldsymbol{X}).$$ From Abramowitz (Chapter 26), when $Z \sim \mathcal{N}(0,1)$, $$ \text{Cov}(Z,\Phi(Z))=\frac{1}{2\sqrt{\pi}}\approx 0.2821. $$ Then the results follow directly.

Point (iv) shows that the eigenvalues of the standard PCA are proportional to those issued from the generalized Gini PCA. Because each eigenvalue (in proportion of the trace) represents the variability (or the quantity of information) inherent to each axis, then both PCA techniques are equivalent when $\boldsymbol{X}$ is Gaussian: $$ \frac{\mu_k}{\text{Tr}[GC_\nu(\boldsymbol{X})]} = \frac{\gamma_k}{\text{Tr}[\text{Var}(\boldsymbol{X})]}, \ \forall k=1,\ldots,K \ ; \ \forall \nu > 1. $$

The $\mathbb{R}^{N}$-Euclidean space

In classical PCA, the duality between $\mathbb{R}^{N}$ and $\mathbb{R}^{K}$ enables the eigenvectors and eigenvalues of $\mathbb{R}^{N}$ to be deduced from those of $\mathbb{R}^{K}$ and conversely. This duality is not so obvious in the Gini PCA case. Indeed, in $\mathbb{R}^{N}$ the Gini variability between the observations would be measured by $GC_\nu(\widetilde{\boldsymbol{Z}}):=\frac{-2\nu}{N(N-1)}(\boldsymbol{R}_\mathbf{z}^{c})^{\top}\boldsymbol{Z}$, and subsequently the idea would be to derive the eigenvalue equation related to $\mathbb{R}^N$, \[ [GC_\nu(\widetilde{\boldsymbol{Z}})+GC_\nu(\widetilde{\boldsymbol{Z}})^{\top}]\ \widetilde{\mathbf{b}}_{\cdot k} = \widetilde{\lambda}_{\cdot k}\widetilde{\mathbf{b}}_{\cdot k}. \] The other option is to define a basis of $\mathbb{R}^N$ from a basis already available in $\mathbb{R}^K$. In particular, the set of principal components $\{\mathbf{f}_1,\ldots,\mathbf{f}_{\cdot k}\}$ provides by construction a set of normalized and orthogonal vectors. Let us rescale the vectors $\mathbf{f}_{\cdot k}$ such that: \[ \widetilde{\mathbf{f}}_{\cdot k} = \frac{\mathbf{f}_{\cdot k}}{GMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})}. \] Then, $\{\widetilde{\mathbf{f}}_1,\ldots,\widetilde{\mathbf{f}}_{\cdot k}\}$ constitutes an orthonormal basis of $\mathbb{R}^K$ in the Gini sense since $GMD_\nu(\widetilde{\mathbf{f}}_{\cdot k},\widetilde{\mathbf{f}}_{\cdot k})=1$. This basis may be used as a projector of the variables $\mathbf{z}_{\cdot k}$ onto $\mathbb{R}^N$. Let $\widetilde{\boldsymbol{F}}$ be the $N\times K$ matrix with $\widetilde{\mathbf{f}}_{\cdot k}$ in columns. The projection of the variables $\mathbf{z}_{\cdot k}$ in $\mathbb{R}^N$ is given by the following Gini correlation matrix: \[ \boldsymbol{V} := \frac{-2\nu}{N(N-1)}\ \widetilde{\boldsymbol{F}}^{\top}\boldsymbol{R}_\mathbf{z}^c, \] whereas it is given by $\frac{1}{N}\widetilde{\boldsymbol{F}}^{\top}\boldsymbol{Z}$ in the standard PCA, that is, the matrix of Pearson correlation coefficients between all $\widetilde{\mathbf{f}}_{\cdot k}$ and $\mathbf{z}_{\cdot \ell}$. The same interpretation is available in the Gini case. The matrix $\boldsymbol{V}$ is normalized in such a way that $\boldsymbol{V} \equiv [v_{k\ell}]$ are the $G$-correlations indices between $\widetilde{\mathbf{f}}_{\cdot k}$ and $\mathbf{z}_{\cdot \ell}$. This yields the ability to make easier the interpretation of the variables projected onto the new subspace.

Interpretations of the Gini PCA

The analysis of the projections of the observations and of the variables are necessary to provide accurate interpretations. Some criteria have to be designed in order to bring out, in the new subspace, the most significant observations and variables.

Observations

The absolute contribution of an observation $i$ to the variability of a principal component $\mathbf{f}_{\cdot k}$ is: \[ ACT_{ik} = \frac{f_{ik}\ \Psi(R(f_{ik}))}{GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})}. \] The absolute contribution of each observation $i$ to the generalized Gini Mean Difference of $\mathbf{f}_{\cdot k}$ ($ACT_{ik}$) is interpreted as a percentage of variability of $GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})$, such that $\sum_{i=1}^N ACT_{ik}=1$. This provides the most important observations $i$ related to component $\mathbf{f}_{\cdot k}$ with respect to the information $GGMD_\nu(\mathbf{f}_{\cdot k},\mathbf{f}_{\cdot k})$. On the other hand, instead of employing the Euclidean distance between one observation $i$ and the component $\mathbf{f}_{\cdot k}$, the Manhattan distance is used. The relative contribution of an observation $i$ to component $\mathbf{f}_{\cdot k}$ is then: \[ RCT_{ik} = \frac{\left|f_{ik}\right |}{ \| \mathbf{f}_{i\cdot}\|_1}. \] Remark that the gravity center of $\{\mathbf{f}_1,\ldots,\mathbf{f}_{\cdot K}\}$ is $\mathbf{g}:=(\bar{\mathbf{f}}_1,\ldots,\bar{\mathbf{f}}_{\cdot K}) =\mathbf{0}$. The Manhattan distance between observation $i$ and $\mathbf{g}$ is then $\sum_{k=1}^K\left|f_{ik} - 0\right |$, and so \[ RCT_{ik} = \frac{\left|f_{ik}\right |}{\| \mathbf{f}_{i\cdot} - \mathbf g\|_1}. \] The relative contribution $RCT_{ik}$ may be interpreted rather as the contribution of dimension $k$ to the overall distance between observation $i$ and $\mathbf{g}$.

Variables

The most significant variables must be retained for the analysis and the interpretation of the data in the new subspace. It would be possible, in the same manner as in the observations case, to compute absolute and relative contributions from the Gini correlation matrix $\boldsymbol{V} \equiv [v_{k\ell}]$. Instead, it is possible to test directly for the significance of the elements $v_{k\ell}$ of $\boldsymbol{V}$ in order to capture the variables that significantly contribute to the Gini variability of components $\mathbf{f}_{\cdot k}$. Let us denote $\widetilde{U}_{\ell k}:=\text{Cov}(\mathbf{f}_{\cdot \ell},\boldsymbol{R}_{\mathbf{z}_{\cdot k}}^c)$ with $\boldsymbol{R}_{\mathbf{z}_{\cdot k}}^c$ the (decumulative) centered rank vector of $\mathbf{z}_{\cdot k}$ raised to an exponent of $\nu-1$ and $U_{\cdot \ell} :=\text{Cov}(\mathbf{f}_{\cdot \ell},\boldsymbol{R}^c_{\mathbf{f}_{\cdot \ell}})$. Those two Gini covariances yield the following $U$-statistics:

equation[equation omitted — 94 chars of source]

Let $U_{\ell k}^0$ be the expectation of $U_{\ell k}$, that is $U_{\ell k}^0:=\mathbb{E}[U_{\ell k}]$. From Yitzhaki13, $U_{\ell k}$ is an unbiased and consistent estimator of $U_{\ell k}^0$. From Theorem 10.4 in Yitzhaki13, Chapter 10, we asymptotically get that $\sqrt{N}(U_{\ell k} - U_{\ell k}^0) \stackrel{a}{\sim} \mathcal{N}$. Then, it is possible to test for:

equation[equation omitted — 120 chars of source]

Let $\hat{\sigma}^2_{\ell k}$ the Jackknife variance of $U_{\ell k}$, then it is possible to test for the null under the assumption $N \to \infty$ as follows:\footnote{As indicated by Yitzhaki1990, the efficient Jackknife method may be used to find the variance of any $U$-statistics.}

equation[equation omitted — 100 chars of source]

The usual PCA enables the variables to be analyzed in the circle of correlation, which outlines the correlations between the variables $\mathbf{z}_{\cdot k}$ and the components $\mathbf{f}_{\cdot \ell}$. In order to make a comparison with the usual PCA, let us rescale the $U$-statistics $U_{\ell k}$. Let $\mathbf{U}$ be the $K\times K$ matrix such that $\mathbf{U} \equiv [U_{\ell k}]$, and $\mathbf{u}_{\cdot k}$ the $k$-th column of $\mathbf{U}$. Then, the absolute contribution of the variable $\mathbf{z}_{\cdot k}$ to the component $\mathbf{f}_{\cdot \ell}$ is: \[ \widetilde{ACT}_{k \ell} = \frac{U_{\ell k}}{\| \mathbf{u}_{\cdot k}\|_2}. \] The measure $\widetilde{ACT}_{k \ell}$ yields a graphical tool aiming at comparing the standard PCA with the Gini PCA. In the standard PCA, $\cos^2\theta$ (see Figure (ref) below) provides the Pearson correlation coefficient between $\mathbf{f}_1$ and $\mathbf{z}_{\cdot k}$. In the Gini PCA, $\cos^2\theta$ is the normalized Gini correlation coefficient $\widetilde{ACT}_{k 1}$ thanks to the $\ell_2$ norm.

figure[figure omitted — 112 chars of source]

It is worth mentioning that the circle of correlation does not provide the significance of the variables. This significance relies on the statistical test based on the $U$-statistics exposed before. Because $\widetilde{ACT}$ depends on the $\ell_2$ metric, it is sensitive to outliers, and as such, the choice of the variables must rely on the test of $U_{\ell k}^0$ only.

Monte Carlo Simulations

In this Section, it is shown with the aid of Monte Carlo simulations that the usual PCA yields irrelevant results when outlying observations contaminate the data. To be precise, the absolute contributions computed in the standard PCA based on the variance may lead to select outlying observations on the first component in which there is the most important variability (a direct implication of the maximization of the variance). In consequence, the interpretation of the PCA may inflate the role of the first principal components. The Gini PCA dilutes the importance of the outliers to make the interpretations more robust and therefore more relevant.

algorithm[algorithm omitted — 677 chars of source]

The mean squared errors of the eigenvalues are computed as follows:

equation[equation omitted — 103 chars of source]

where $\lambda_k^{oi}$ is the eigenvalue computed with outlying observations in the sample. The MSE of $ACT$ et $RCT$ are computed in the same manner.

We first investigate the case where the variables are highly correlated in order to gauge the robustness of each technique (Gini for $\nu=2,4,6$ and variance). The correlation matrix between the variables is given by:

equation[equation omitted — 184 chars of source]

As can be seen in the matrix above, we can expect that all the information be gathered on the first axis because each pair of variables records an important linear correlation. The repartition of the information on each component, that is, each eigenvalue in percentage of the sum of the eigenvalues is the following.

table[table omitted — 1,303 chars of source]

The first axis captures around 82% of the variability of the overall sample (before contamination). Although each PCA method yields the same repartition of the information over the different components before the contamination of the data, it is possible to show that the classical PCA is not robust. For this purpose, let us analyze Figures (ref)-(ref) below that depict the MSE of each observation with respect to the contamination process described in Algorithm 1 above.

figure[figure omitted — 590 chars of source]

On the first axis of Figure (ref), the absolute contribution of each observation (among 500 observations) is not stable because of the contamination of the data, however the Gini PCA performs better. The MSE of the ACTs measured during to the contamination process provides lower values for the Gini index compared with the variance. On the other hand, if we compute the standard deviation of all these MSEs over the two first axis, again the Gini methodology provides lower variations (see Table (ref)). \break

table[table omitted — 341 chars of source]

Let us take now an example with less correlations between the variables in order to get a more equal repartition of the information on the first two axes.

equation[equation omitted — 184 chars of source]

The repartition of the information over the new axes (percentage of each eigenvalue) is given in Table (ref). When the information is less concentrated on the first axis (55% on axis 1 and around 35% on axis 2), the MSE of the eigenvalues after contamination are much more important for the standard PCA compared with the Gini approach (2 to 3 times more important). Although the fourth axis reports an important MSE for the Gini method ($\nu=6$), the eigenvalue percentage is not significant (1.56%).

table[table omitted — 1,310 chars of source]

Let us now have a look on the MSE of the absolute contributions of each observation ($N=500$) for each PCA technique ((ref)-(ref)). We obtain the same kind of results, with less variability on the second axis. In Figures (ref)-(ref), it is apparent that the classical PCA based on the $\ell_2$ norm exhibits much more ACT variability (black points). This means that the contamination of the data can lead to the interpretation of some observations as significant (important contribution to the variance of the axis) while they are not (and vice versa). On the other hand, the MSE of the RCTs after contamination of the data, Figures (ref)-(ref), are less spread out for the Gini technique for $\nu=4$ and $\nu=6$, however for $\nu=2$ there is more variability of the MSE compared with the variance. This means that the distance from one observation to an axis may not be reliable (although the interpretation of the data rather depends on the ACTs).

figure[figure omitted — 588 chars of source]

The results of Figures (ref)-(ref) can be synthesized by measuring the standard deviation of the MSE over the ACTs of the 500 observations along the two first axes.

table[table omitted — 346 chars of source]

As in the previous example of simulation, Table (ref) indicates that the PCA based on the variance is less stable about the values of the ACTs that provide the most important observations of the sample. This may lead to irrelevant interpretations.

Application on cars data

We propose a simple application with the celebrated cars data (see the Appendix).\footnote{An R markdown for Gini PCA is available: \scriptsize{\textcolor{blue}{https://github.com/freakonometrics/GiniACP/}} \\ Data from Michel Tenenhaus's website (see also the Appendix):\\ \scriptsize{\textcolor{blue}{https://studies2.hec.fr/jahia/webdav/site/hec/shared/sites/tenenhaus/acces_anonyme/home/fichier_excel/auto_2004.xls}} } The dataset is particularly interesting since there are highly correlated variables as can be seen in the Pearson correlation matrix given in Table (ref).

table[table omitted — 729 chars of source]

Also, the dataset is composed of some outlying observations (Figure (ref)): Ferrari enzo ($x_1$, $x_2$, $x_5$), Bentley continental ($x_2$), Aston Martin ($x_2$), Land Rover discovery ($x_5$), Mercedes class S ($x_5$), Smart ($x_5, x_6$).

figure[figure omitted — 430 chars of source]

The overall information (variability) is partitioned over six components (Table (ref)).

table[table omitted — 836 chars of source]

Two axes may be chosen to analyze the data. As shown in the previous Section about the simulations, when the data are highly correlated such that two axes are sufficient to project the data, the Gini PCA and the standard PCA yield the same share of information on each axis. However, we can expect some differences for absolute contributions $ACT$ and relative contributions $RCT$.

The projection of the data is depicted in Figure (ref), for each method.

figure[figure omitted — 572 chars of source]

As depicted in Figure (ref), the projection is very similar for each technique. The cars with extraordinary (or very low) abilities are in the same relative position in the four projections: Land Rover Discovery Td5 at the top, Ferrari Enzo at the bottom right, Smart Fortwo coup\'e at the bottom left. However, when we improve the coefficient of variability $\nu$ to look for what happens at the tails of the distributions (of the two axes), we see that more cars are distinguishable: Land Rover Defender, Audi TT, BMW Z4, Renault Clio 3.0 V6, Bentley Contiental GT. Consequently, contrary to the case $\nu=2$ or the variance, the projections with $\nu=4,6$ allow one to find other important observations, which are not outlying observations that contribute to the overall amount of variability. For this purpose, let us first analyze the correlations between the variables and the new axes in order to interpret the results, see Tables (ref) to (ref).

Some slight differences appear between the Gini PCA and the classical one based on the variance. The theoretical Section (ref) indicates that the Gini methodology for $\nu=2$ is equivalent to the variance when the variables are Gaussian. On cars data, we observe this similarity. In each PCA, all variables are correlated with Axis 1 and weight with Axis 2. However, when $\nu$ increases, the Gini methodology allows outlying observations to be diluted so that some variables may appear to be significant, whereas they are not in the variance case.

table[table omitted — 1,025 chars of source]
table[table omitted — 1,022 chars of source]
table[table omitted — 1,049 chars of source]
table[table omitted — 1,032 chars of source]

Tables (ref) and (ref) ($\nu=4,6$) show that Axis 2 is correlated to speed (not weight as in the variance PCA). In this respect the absolute contributions must describe the cars associated with speed on Axis 2. Indeed, the Land Rover discovery, a heavy weight car, is no more available on Axis 2 for the Gini PCA for $\nu=2,4,6$ (Figures (ref), (ref), (ref)). Note that the red line in the Figures represents the mean share of the information on each axis, i.e. 100%/24 cars = 4.16% of information per car.

figure[figure omitted — 308 chars of source]
figure[figure omitted — 312 chars of source]
figure[figure omitted — 312 chars of source]
figure[figure omitted — 313 chars of source]

Finally, some cars are not correlated with axis 2 in the standard PCA, see Figures (ref)--(ref), while this is the case in the Gini PCA. Indeed some cars are now associated with speed: Aston Martin, Bentley Continental GT, Renault Clio 3.0 V6 and Mini 1.6 170. This example of application shows that the use of the Gini metric robust to outliers may involve some serious changes in the interpretation of the results.

Conclusion

In this paper, it has been shown that the geometry of the Gini covariance operator allows one to perform Gini PCA, that is, a robust principal component analysis based on the $\ell_1$ norm.

To be precise, the variance may be replaced by the Gini Mean Difference, which captures the variability of couples of variables based on the rank of the observations in order to attenuate the influence of the outliers. The Gini Mean Difference may be rather interpreted with the aid of the generalized Gini index $GGMD_\nu$ in the new subspace for a better understanding of the variability of the components, that is, $GGMD_\nu$ is both a rank-dependent measure of variability in Yaari1987 sense and also an eigenvalue of the Gini correlation matrix.

Contrary to many approaches in multidimensional statistics in which the standard variance-covariance matrix is used to project the data onto a new subspace before deriving multidimensional Gini indices (see e.g. Banerjee), we propose to employ the Gini correlation indices (see Yitzhaki13). This provides the ability to interpret the results with the $\ell_1$ norm and the use of $U$-statistics to measure the significance of the correlation between the new axes and the variables.

This research may open the way on data analysis based on Gini metrics in order to study multivariate correlations with categorical variables or discriminant analyses when outlying observations drastically affect the sample.

thebibliography{99} \bibitem[Abramowitz & Stegun(1964)]{Abramowitz} Abramowitz, M. & I. Stegun. (1964). Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables. National Bureau of Standards Applied Mathematics Series No. 55. \bibitem[Anderson(1963)]{Anderson} Anderson, T.W. (1963) Asymptotic Theory for Principal Component Analysis. {\em The Annals of Mathematical Statistics}, {\bf 34}, 122--148. \bibitem[Baccini, Besse & de Falguerolles(1996)]{Baccini} Baccini, A., P. Besse & A. de Falguerolles (1996), A $L_1$ norm PCA and a heuristic approach, in Ordinal and Symbolic Data Analysis, E Didday, Y. Lechevalier and O. Opitz (eds), Springer, 359--368. \bibitem[Banerjee(2010)]{Banerjee} Banerjee, A.K. (2010), A multidimensional Gini index, Mathematical Social Sciences, {\bf 60}: 87--93. \bibitem[Candes {\em et al.}(2009)]{Candes} Candes, E.J., Xiaodong Li, Yi Ma & John Wright. (2009) Robust Principal Component Analysis?. arXiv:0912.3599. \bibitem[Carcea and Serfling(2015)]{ref2} Carcea, M. & R. Serfling (2015), A Gini autocovariance function for time series modeling. Journal of Time Series Analysis 36: 817--38. \bibitem[Dalton(1920)]{Dalton1920} Dalton, H. 1920. The Measurement of the Inequality of Incomes. {\em The Economic Journal}, {\bf 30}:119, 348--361. \bibitem[d'Aspremont {\em et al.}(1920)]{dAspremont} d'Aspremont, A., L. El Ghaoui, M.I. Jordan, & G. R. G. Lanckriet (2007). A direct formulation for sparse PCA using semidefinite programming. {\em SIAM Review}, {\bf 49}:3, 434--448. \bibitem[Decancq & Lugo(2013)]{Decancq} Decancq, K. & M.-A. Lugo (2013), Weights in Multidimensional Indices of Well-Being: An Overview, Econometric Reviews, {\bf 32}:1, 7--34. \bibitem[Ding {\em et al.}(2006)]{Ding2006} Ding, C., Zhou, D., He, X. & Zha, H. (2006). $R_1$-PCA: rotational invariant L1-norm principal component analysis for robust subspace factorization. {\em ICML '06 Proceedings of the 23rd international conference on Machine learning}, 281--288 \bibitem[Eckart & Young(1936)]{Eckart} Eckart, C. & G. Young (1936) The approximation of one matrix by another of lower rank. {\em Psychometrika},{\bf 1}, 211--218. \bibitem[Flury & Riedwyl(1988)]{FluryRiedwyl} Flury & Riedwyl (1988). Multivariate Statistics: A Practical Approach. Chapman & Hall \bibitem[Furman & Zitikis(2017)]{furman} Furman, E. & R. Zitikis (2017), Beyond the Pearson Correlation: Heavy-Tailed Risks, Weighted Gini Correlations, and A Gini-Type Weighted Insurance Pricing Model, \emph{ASTIN Bulletin: The Journal of the International Actuarial Association}, \textbf{47(03)}: 919-942. \bibitem[Gajdos & Weymark(2005)]{Gajdos} Gajdos, T. and J. Weymark (2005), Multidimensional generalized Gini indices, \emph{Economic Theory}, {\bf 26}:3, 471-496. \bibitem[Gini(1912)]{Gini12} Gini, C. (1912), Variabilit\`{a} e mutabilit\`{a}, Memori di Metodologia Statistica, Vol. 1, Variabilit\`{a} e Concentrazione. Libreria Eredi Virgilio Veschi, Rome, 211--382. \bibitem[Giorgi(2013)]{giorgi} Giorgi, G.M. (2013), Back to the future: some considerations on Shlomo Yitzhaki and Edna Schechtman's book "`The Gini Methodology: A Primer on a Statistical Methodology"', \emph{Metron}, \textbf{71(2)}: 189-195. \bibitem[Gorban {\em et al.}(2007)]{Gorban} Gorban, A.N. , B. Kegl, D.C. Wunsch, & A. Zinovyev (Eds.) (2007) Principal Manifolds for Data Visualisation and Dimension Reduction. LNCSE 58, Springer Verlag. \bibitem[Hotelling(1933)]{Hotelling} Hotelling, H. (1933) Analysis of a complex of statistical variables into principal components. {\em Journal of Educational Psychology}, {\bf 24}, 417--441, 1933. \bibitem[Korhonen & Siljam\"aki(1998)]{Korhonen} Korhonen, P. & Siljam\"aki, A. (1998). Ordinal principal component analysis theory and an application. {\em Computational Statistics & Data Analysis}, {\bf 26}:4, 411--424. \bibitem[List(1999)]{List} List, C. (1999), Multidimensional Inequality Measurement: A Proposal, \emph{Working paper}, Nuffield College. \bibitem[Mackey(2009)]{Mackey} Mackey, L. (2009) Deflation methods for sparse PCA. {\em Advances in Neural Information Processing Systems}, {\bf 21}: 1017--1024. \bibitem[Mardia, Kent & Bibby(1979)]{Mardia} Mardia, K, Kent, J. & Bibby, J. (1979). Multivariate Analysis. Academic Press, London. \bibitem[Olkin and Yitzhaki(1992)]{ref9} Olkin, Ingram, and Shlomo Yitzhaki. 1992. Gini regression analysis. \textit{International Statistical Review} \textbf{60}: 185--96. \bibitem[Pearson(1901)]{Pearson} Pearson, K. (1901), On Lines and Planes of Closest Fit to System of Points in Space, {\em Philosophical Magazine}, {\bf 2}: 559--572. \bibitem[Saad(1998)]{Saad} Saad, Y. (1998). Projection and deflation methods for partial pole assignment in linear state feedback .{\em IEEE Trans. Automat. Contr.}, {\bf 33}: 290--297. \bibitem[Schechtman & Yitzhaki(1987)]{Schechtman87} Schechtman, E. & S. Yitzhaki (1987), A Measure of Association Based on Gini's Mean Difference, \emph{Communications in Statistics: A}, {\bf 16}: {207--231}. \bibitem[Yitzhaki & Schechtman(2003)]{Schechtman03} Schechtman, E. & S. Yitzhaki (2003), A family of correlation coefficients based on the extended Gini index, \emph{Journal of Economic Inequality}, {\bf 1}:2, {129--146}. \bibitem[Shelef(2016)]{shelef16}Shelef, A. (2016), A Gini-based unit root test. \emph{Computational Statistics & Data Analysis}, \textbf{100}: 763--772. \bibitem[Shelef & Schechtman(2011)]{ref13} Shelef, A., and E. Schechtman (2011), A Gini-based methodology for identifying and analyzing time series with non-normal innovations. \emph{SSNR Electronic Journal} July: 1--26. \bibitem[Tibshirani(1996)]{Tibshirani} Tibshirani, R. (1996) Regression shrinkage and selection via the LASSO. {\em Journal of the Royal statistical society, series B}, {\bf 58}:1, 267--288. \bibitem[Yaari (1987)]{Yaari1987}\ Yaari, M.E. (1987), The Dual Theory of Choice Under Risk, \textit{Econometrica}, {\bf55}: 99--115. \bibitem[Yaari (1988)]{Yaari1988}\ Yaari, M.E. (1988), A Controversial Proposal Concerning Inequality Measurement, \textit{Journal of Economic Theory}, {\bf 44}: 381--397. \bibitem[Yitzhaki(1991)]{Yitzhaki1990} Yitzhaki, S. (1991), Calculating Jackknife Variance Estimators for Parameters of the Gini Method. {\em Journal of Business and Economic Statistics}, {\bf 9}: 235--239. \bibitem[Yitzhaki(2003)]{Yitzhaki03} Yitzhaki, S. (2003), Gini's Mean difference: a superior measure of variability for non-normal distributions, \emph{Metron}, \textbf{LXI(2)}: 285-316. \bibitem[Yitzhaki & Olkin(1991)]{Olkin} Yitzhaki, S. & Olkin, I. (1991). Concentration indices and concentration curves. {\em Institute of Mathematical Statistics Lecture Notes}, {\bf 19}: 380--392. \bibitem[Yitzhaki & Schechtman(2013)]{Yitzhaki13} Yitzhaki, S. & E. Schechtman (2013), \emph{The Gini Methodology. A Primer on a Statistical Methodology}, Springer. \bibitem[Zou, Hastie & Tibshirani(2006)]{Zou} Zou, H., Hastie, T. & R. Tibshirani (2006), Sparse Principal Component Analysis, \emph{Journal of Computational and Graphical Statistics}, {\bf15}:2, 265-286.

\break