EconBase
← Back to paper

Pairwise Difference Representations of Moments: Gini and Generalized Lagrange identities

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,772 characters · 8 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Article title for specific publication (Journal)

Slides\begin{comment} \begin{Slides} \end{comment} \end{Slides} \begin{Tentative} \begin{comment} \pagenumbering{Roman}\setcounter{page}{1} \begin{center} PLAN \quad \end{center} \begin{enumerate} • To do \begin{enumerate} • Abandon the terminology \textquotedblleft decoupling\textquotedblright : use $D$-representation. Done. • Introduce the notion of $D$-statistics • The first differences have symmetric distributions centered at zero. • Relate to symmetrization inequalities. • Relate to the coupling method. • Moment decompositions based on Gini representations. Relate to Samuelson inequality. • Sign-skewness statistics. • Link with Spearman and Kendall measures of dependence. • Improve the results on higher-order moments. • Relate to: \begin{enumerate} • $h$-statistics; • $U$-statistics. \end{enumerate} • Unbiased estimation of higher-order moments \begin{enumerate} • Document the bias of natural sample moments • Relate to Heffernan (1997, JRSS B). \end{enumerate} • Improved procedures based on unbiased estimators of higher-order moments. • Relate to cumulative distributions. • Applications \begin{enumerate} • Conceptual: pairwise differences vs. deviations from the mean • Categorical data • First differences are not contaminated by the location estimator • Representations of empirical moments in terms of symmetric random variables • Alternative method for deriving unbiased estimators of moments • Decompositions of empirical moments in terms of observations • Tests based on pairwise differences • Filtered and robust moments • Multivariate extensions • Lorenz curves in finance • Conditioning • Ranks as $U$ -statistics \end{enumerate} \end{enumerate} \end{enumerate} \quad Potential journal for submission \begin{enumerate} • The American Statistician • Statistics and Probability Letters • Communications in Statistics • Econometrics and Statistics • Journal of Statistics and Data Science Education • Teaching Statistics \end{enumerate} \begin{Tentative} \end{comment} \end{Tentative} \pagestyle{plain}\pagenumbering{roman}\setcounter{page}{1} \begin{center} ABSTRACT \quad \end{center} We provide pairwise-difference (Gini-type) representations of higher-order central moments for both general random variables and empirical moments. Such representations do not require a measure of location. For third and fourth moments, this yields pairwise-difference representations of skewness and kurtosis coefficients. We show that all central moments possess such representations, so no reference to the mean is needed for moments of any order. This is done by considering i.i.d. replications of the random variables considered, by observing that central moments can be interpreted as covariances between a random variable and powers of the same variable, and by giving recursions which link the pairwise-difference representation of any moment to lower order ones. Numerical summation identities are deduced. Through a similar approach, we give analogues of the Lagrange and Binet-Cauchy identities for general random variables, along with a simple derivation of the classic Cauchy-Schwarz inequality for covariances. Finally, an application to unbiased estimation of centered moments is discussed. \quad Key words: Gini; Covariance; Skewness; Kurtosis; Pairwise differences; Lagrange identity; Moments; Unbiased estimation. \begin{comment} \end{comment} \quad \begin{Tentative} \begin{comment} \listoftables \listoffigures \begin{Tentative} \end{comment} \end{Tentative} \begin{comment}

Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes \@mkboth{\MakeUppercase Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes} {\MakeUppercase Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes}

\@starttoc{lth}

\addcontentsline{toc}{section}{ Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes} \end{comment}

\pagenumbering{arabic} \setcounter{section}{0} \setcounter{page}{1}

Introduction

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

The sample variance is certainly the most widely used measure of dispersion among observations in statistics. It is typically defined as the average of the squared deviations of the observations $x_{1},\ldots ,\,x_{n}$ from their sample mean:

equation[equation omitted — 126 chars of source]

where $\bar{x}=\overset{n}{\underset{i=1}{\sum }}x_{i}/n$. At first sight, the variance depends crucially on \textquotedblleft centering\textquotedblright\ the observations with respect to their mean. A century ago, however, \nocite*{Gini(1912)} noticed that $s_{X}^{2}(n)$ can be rewritten as the average of pairwise differences between all the observations:

equation[equation omitted — 274 chars of source]

see \nocite*{Heffernan(1988)}. The Gini representation underscores that $ s_{X}^{2}(n)$ does not depend on using the sample mean as location parameter, and thus may be viewed as a measure of \textquotedblleft intrinsic variability\textquotedblright\ between observations. In particular, it provides a way of measuring dispersion when there is no natural location parameter, as with categorical data; see \citename{Light-Margolin(1971)} (\citeyear*{Light-Margolin(1971)}, \citeyear*{Light-Margolin(1974)}) . Further, by focusing on pairwise differences $x_{i}-x_{j}$, it may be easier to determine which observations have the greatest influence on $ s_{X}^{2}(n)$, because deviations from the mean $x_{i}-\bar{x}$ depend on all the observations (through $\bar{x})$. A comprehensive review of statistical methods based on the Gini approach is presented by \nocite* {Yitzhaki-Schechtman(2013)}.

If we have bivariate observations $z_{i}:=(x_{i},$\thinspace $y_{i})^{\prime }$, $i=1,\ldots ,\,n,$ the sample covariance

equation[equation omitted — 136 chars of source]

is in turn the most widely used \textquotedblleft measure of association\textquotedblright\ between $x$ and $y.$ For the covariance, a pairwise representation is provided by the formula

equation[equation omitted — 291 chars of source]

see \nocite*{Hayes(2011)}. Clearly, ((ref)) is a special case of ((ref)) obtained by taking $ x_{i}=y_{i},\;i=1,\ldots ,\,n.$ Formula ((ref) ) shows that the sample covariance measures the tendency of $x$ and $y$ to move in the same direction, without reference to sample means. Similar expressions for the variances and covariances of general random variables have also been used in the literature on $U$-statistics; see \nocite*[Chapter 1] {LeeAJ(1990)}. Since centering is not needed, formula ((ref)) provides a basis for developing alternative measures of dependence, such as Kendall's and Spearman's measures of dependence [see \nocite*{Lehmann(1966)}]. Again, pairwise differences $x_{i}-x_{j}$ and $ y_{i}-y_{j}$ may be easier to interpret than $x_{i}-\bar{x}$ and $y_{i}-\bar{ y}$, because $\bar{x}$ and $\bar{y}$ depend on all the observations.

Another similar result we consider here is the Lagrange identity:

gather[gather omitted — 505 chars of source]

see \nocite*{Wright(1992)} and \nocite*[Chapter 3]{Steele(2004)}. An interesting feature of this identity is that it involves pairwise comparisons between cross-products of the components of $x:=(x_{1},\ldots ,\,x_{n})^{\prime }$ and\ $y:=(y_{1},\ldots ,\,y_{n})^{\prime }$. Since the right-hand side is non-negative, the Lagrange identity provides a simple way of proving the discrete Cauchy-Schwarz inequality:

equation[equation omitted — 190 chars of source]

Interestingly, ((ref)) directly shows when the bound is an equality $(x_{i}y_{j}=x_{j}y_{i}\;$for all $i$ and $j)$ and what determines its tightness. The Lagrange identity is itself a special case of the Binet-Cauchy identity

equation[equation omitted — 353 chars of source]

where $a_{i},\,b_{i},\,c_{i},\,d_{i},$ $i=1,\ldots ,\,n$, are arbitrary real constants [see \nocite*[p. 49]{Steele(2004)}]. Formula ((ref)) relates several inner products to pairwise comparisons between the components of the vectors considered.

Since the representations in ((ref)) and ((ref)) only depend on differences between observations, we will call such representations (and similar ones) $D$ -representations (or Gini-type representations). In this paper, we provide $D$-representations of higher-order central moments for both general random variables and empirical moments. For third and fourth moments, this yields new representations of skewness and kurtosis coefficients, where reference to a location parameter is eliminated. We show that all central moments possess $D$-representations, so again no reference to the mean is needed for moments of any order. This is done by considering i.i.d. replications (or realizations) of the random variables considered, by observing that central moments can be interpreted as covariances between a random variable and powers of the same variable (a \emph{nonlinear} relation), and by giving \emph{recursions} which link the $ D $-representation of any moment to lower order ones. Numerical summation identities similar to ((ref))\thinspace -\thinspace ((ref)) are also deduced. Finally, through a similar approach, we give the analogues of the Lagrange and Binet-Cauchy identities for general random variables, along with a simple derivation of the classic Cauchy-Schwarz inequality for covariances.

The paper is organized as follows. In Section (ref), we discuss $D$-representations for covariances and regression coefficients, and we show how corresponding numerical identities can be derived from these. In Section (ref), we provide $D$-representations for third and fourth moments, along with corresponding expressions for skewness and kurtosis coefficients. In Section (ref), we show that all moments possess such representations, and we give recursive formulae for computing a $D$ -representation for any moment. Generalized Lagrange and Binet-Cauchy identities are given in Section (ref). In Section (ref), we discuss some applications of the proposed moment representations. We conclude in Section (ref) . Proofs are given in an online Appendix.

$D$-representation of Covariances

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

In order to generalize the expressions ((ref) )\thinspace -\thinspace ((ref)) to higher-order moments, it will be useful to discuss the pairwise difference representation of the covariance between general random variables. Let $X$ and $Y$ be random variables with finite second moments, and means ${\Greekmath 0116} _{X}= \mathbb{E}\{X\}$, ${\Greekmath 0116} _{Y}=\mathbb{E}\{Y\}.$ We denote by $\mathsf{C} [X,\,Y]:=\mathbb{E}\{(X-{\Greekmath 0116} _{X})(Y-{\Greekmath 0116} _{Y})\}$ the covariance between $X$ and $Y$. Consider two independent and identically distributed replications of the random vector $(X,\,Y)^{\prime }$: $\;(X_{1},\,Y_{1})^{\prime }$ and $ (X_{2},\,Y_{2})^{\prime }$. On observing

eqnarray[eqnarray omitted — 424 chars of source]

it follows that

equation[equation omitted — 104 chars of source]

It is also easy to see that

eqnarray[eqnarray omitted — 213 chars of source]

In other words, only one of the variables in $(X,$\thinspace $Y)^{\prime }$ need be replicated to avoid centering with respect to the means (${\Greekmath 0116} _{X}$ and ${\Greekmath 0116} _{Y}$). ((ref))\thinspace -\thinspace ( (ref)) also hold under the weaker assumption (without identical distributions) that $(X_{1},\,Y_{1})^{\prime }$ and $ (X_{2},\,Y_{2})^{\prime }$ have the same first and second moments as $ (X,\,Y) $ with

equation[equation omitted — 100 chars of source]

For $X_{1}=Y_{1}$ and $X_{2}=Y_{2},$ we obtain the following formulae for the variance of $X:$

eqnarray[eqnarray omitted — 203 chars of source]

If $X_{3}$ is a third replication of $X$ [so that $X_{1}$, $X_{2}$, $X_{3}$ are i.i.d.], it is also easy to check that

equation[equation omitted — 108 chars of source]

The uncentered second moment of $X$ can be written as

eqnarray[eqnarray omitted — 244 chars of source]

and the linear regression coefficient of $Y$ on $X$ is

eqnarray[eqnarray omitted — 355 chars of source]

In view of the above discussion on the covariance and the variance, we will define more formally what we mean by the $D$-representation of a parameter.

definitionA parameter ${\Greekmath 010D} $ has a $D$-representation for the variables $ X_{1},\ldots ,\,X_{n}$ if there is a function $g(x_{1},\ldots ,\,x_{n})=h(x_{2}-x_{1},\ldots ,\,x_{n}-x_{n-1})$ such that \begin{equation} {\Greekmath 010D} =\mathbb{E}\{h(X_{2}-X_{1},\ldots ,\,X_{n}-X_{n-1})\}\,. \end{equation}

In other words, ${\Greekmath 010D} $ has a $D$-representation for the variables $ X_{1},\ldots ,\,X_{n}$ if it can be written as the expected value of a function $g(X_{1},\ldots ,\,X_{n})$ which depends on $X_{1},\ldots ,\,X_{n}$ only through the differences between the variables $X_{1},\ldots ,\,X_{n}$. Equivalently, this means that

equation[equation omitted — 73 chars of source]

where $g(x_{1},\ldots ,\,x_{n})$ is invariant to uniform location changes, i.e.

equation[equation omitted — 152 chars of source]

A potentially useful feature of the representation ((ref)) follows from the fact that the variables $(X_{1}-X_{2})$ and $(Y_{1}-Y_{2})$ have marginal distributions symmetric around zero. Consequently, the odd moments of $(X_{1}-X_{2})$ and $(Y_{1}-Y_{2})$ are all zero (when they exist). More generally, the distribution of the vector $(X_{1}-X_{2},\,Y_{1}-Y_{2})^{ \prime }$ is jointly symmetric around zero in the sense that

equation[equation omitted — 97 chars of source]

This type of symmetry can be useful for proving unbiasedness results in various setups; see \citename{Dufour(1984)} (\citeyear*{Dufour(1984)}, \citeyear*{Dufour(1985)}). This property is shared by all $D$-representations when the paired differences are based on i.i.d. observations.

It is of interest to note that the numerical identity ((ref)) can be derived from the general moment formula ((ref)). Let $(I,$\thinspace $J)$ be a pair of random integers such that

equation[equation omitted — 81 chars of source]

and consider the random vector

equation[equation omitted — 58 chars of source]

The values $x_{1},\ldots ,\,x_{n}$ and $y_{1},\ldots ,\,y_{n}$ need not be distinct. It is then easy to see that:

equation[equation omitted — 190 chars of source]
equation[equation omitted — 127 chars of source]

and by the definition of the covariance,

equation[equation omitted — 209 chars of source]

Consider now two i.i.d. replications of $(X^{\ast },\,Y^{\ast })$:\ \ $ (X_{1}^{\ast },\,Y_{1}^{\ast })$ and $(X_{2}^{\ast },\,Y_{2}^{\ast })$. By applying ((ref)) and ((ref)) to $(X^{\ast },$\thinspace $ Y^{\ast })$, the covariance between $X^{\ast }$ and $Y^{\ast }$ is given by:

equation[equation omitted — 224 chars of source]

Since the above sum contains a zero whenever $i=j$, it is natural to divide by $(n^{2}-n)\;$rather than $n^{2}$. This leads to:

equation[equation omitted — 191 chars of source]

In the special case where $X^{\ast }=Y^{\ast },$ we have

equation[equation omitted — 182 chars of source]

hence, on setting $s_{X^{\ast }}^{2}:=(\frac{n}{n-1}){\Greekmath 011B} _{X^{\ast }}^{2}$ , the Gini formula ((ref)) for the variance is:

equation[equation omitted — 123 chars of source]

Skewness and Kurtosis

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

In this section, we express skewness and kurtosis coefficients in terms of pairwise differences. Such representations can be written as functions of differences of the form $X_{i}-X_{j}$ or differences between powers of variables [$X_{i}^{2}-X_{j}^{2}$]. We first underscore that the third and fourth central moments of a random variable $X$ can be interpreted as covariances, which in turn can be represented in terms of pairwise differences between replicated random variables. In each case, we give several expressions which may have independent interest.

propositionCovariance representation of third and fourth central moments. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : Covariance representation of third and fourth central moments}} Let $X$ be a random variable with finite mean ${\Greekmath 0116} _{X}$ and variance ${\Greekmath 011B} _{X}^{2}$. Let $X_{1}$, $X_{2}$ two i.i.d. replications of $X$. If\ $X$ has finite third moment, then \begin{eqnarray} \mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\} &=&\mathsf{C}[X,\,(X-{\Greekmath 0116} _{X})^{2}]=\mathsf{C} [X,\,X^{2}]-2{\Greekmath 0116} _{X}{\Greekmath 011B} _{X}^{2} \notag \\ &=&\mathbb{E}\{(X_{1}-X_{2})(X_{1}-{\Greekmath 0116} _{X})^{2}\}=\mathbb{E} \{(X_{1}-X_{2})X_{1}^{2}\}-2{\Greekmath 0116} _{X}{\Greekmath 011B} _{X}^{2} \notag \\ &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})[(X_{1}-{\Greekmath 0116} _{X})^{2}-(X_{2}-{\Greekmath 0116} _{X})^{2}]\} \notag \\ &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})(X_{1}^{2}-X_{2}^{2})\}-{\Greekmath 0116} _{X} \mathbb{E}\{(X_{1}-X_{2})^{2}\}\,. \end{eqnarray} If $X$ has finite fourth moment, then \begin{eqnarray} \mathbb{E}\{(X-{\Greekmath 0116} _{X})^{4}\} &=&\mathsf{C}[X,\,(X-{\Greekmath 0116} _{X})^{3}]=\mathsf{C} [X,\,X^{3}]-3{\Greekmath 0116} _{X}\mathsf{C}[X,\,X^{2}]+3{\Greekmath 0116} _{X}^{2}{\Greekmath 011B} _{X}^{2} \notag \\ &=&\mathbb{E}\{(X_{1}-X_{2})(X_{1}-{\Greekmath 0116} _{X})^{3}\}=\mathbb{E} \{(X_{1}-X_{2})(X_{1}^{2}-2{\Greekmath 0116} _{X}X_{1})\} \notag \\ &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})[(X_{1}-{\Greekmath 0116} _{X})^{3}-(X_{2}-{\Greekmath 0116} _{X})^{3}]\} \notag \\ &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})(X_{1}^{3}-X_{2}^{3})\}-3{\Greekmath 0116} _{X} \mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}-3{\Greekmath 0116} _{X}^{2}{\Greekmath 011B} _{X}^{2}\,. \end{eqnarray}

To interpret Proposition (ref), consider first the case where the mean is zero ($ {\Greekmath 0116} _{X}=0$), so the third moment does not depend on the mean or the variance. When the distribution is skewed to the right [$\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}>0$\}, equation ((ref)) shows that an increase in the value of $X$ ($X_{1}-X_{2}>0$) is associated with an increase in the absolute value of $X$ [$X_{1}^{2}-X_{2}^{2}>0$], while a shift to the left would go with a decrease in the absolute value of $X$. More generally, $\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}$ can be written as

equation[equation omitted — 222 chars of source]

so that $\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}$ measures the association between a change in the deviation from the mean and the corresponding change in the absolute values of the deviations. A positive (negative) association entails positive (negative) skewness, which is clearly exhibited by pairwise differences.

The formulae given by Proposition (ref) depend on unknown parameters (${\Greekmath 0116} _{X}$ and ${\Greekmath 011B} _{X}^{2}$). It is possible to avoid this by using a larger number of replications. For the third moment, only one additional replication of $X$ is needed. If $X_{1}$, $X_{2}$, $X_{3}$ are three i.i.d. replications of $X$, we see [from ((ref))] that

eqnarray[eqnarray omitted — 311 chars of source]

For the fourth moment, this can be done with four replications. If $X_{1}$, $ X_{2}$, $X_{3}$, $X_{4}$ are four i.i.d. replications of $X$, we have:

align*[align* omitted — 289 chars of source]
equation[equation omitted — 200 chars of source]

((ref)) - ((ref)) still do not provide $D$-representations, because $X_{3}$ and $X_{4}$ are in levels. This is done in the following propositions by using additional replicates of $X$. A number of alternative expressions are provided in the following proposition.

proposition$D$-representation of third central moment. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : $D$-representation of third central moment}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, variance ${\Greekmath 011B} _{X}^{2}$ , and finite third moment. If $X_{1}$, $X_{2}$, $X_{3}$ are three i.i.d. replications of $X$, then \begin{eqnarray} \mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\} &=&\mathbb{E}\{(X_{1}-X_{3})(X_{1}-X_{2})^{2} \}=\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})[(X_{1}-X_{3})^{2}-(X_{2}-X_{3})^{2}] \} \notag \\ &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})^{2}[(X_{1}-X_{3})+(X_{2}-X_{3})]\}= \frac{1}{6}\mathbb{E}\{D_{1}^{(3)}D_{2}^{(3)}D_{3}^{(3)}\} \notag \\ &=&\frac{9}{2}\mathbb{E}\{(X_{1}-\bar{X}^{(3)})(X_{2}-\bar{X}^{(3)})(X_{3}- \bar{X}^{(3)})\}\quad \quad \end{eqnarray} where\ \ $\bar{X}^{(3)}:=\,\underset{j=1}{\overset{3}{\sum }}X_{j}/3\;$and \begin{equation} D_{i}^{(3)}:=\,\underset{j\neq i}{\sum }(X_{i}-X_{j})=3(X_{i}-\bar{X} ^{(3)})\,,\quad i=1,\,2,\,3. \end{equation}

As the variables $D_{i}^{(3)}$ are sums of pairwise differences, ((ref)) provides simple $D$-representations of the third central moment of $X$. Note also that $\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}$ can be rewritten as a function of deviations with respect to a \textquotedblleft reference\textquotedblright\ value $X_{3}$:

eqnarray[eqnarray omitted — 262 chars of source]

We now consider the fourth central moment.

proposition$D$-representation of fourth central moment. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : $D$-representation of fourth central moment}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, variance ${\Greekmath 011B} _{X}^{2}$, and finite fourth moment. If $X_{1}$, $X_{2}$, $X_{3}$, $X_{4}$ four i.i.d. replications of $X$, then \begin{eqnarray} \mathbb{E}\{(X-{\Greekmath 0116} _{X})^{4}\} &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})^{4} \}-3{\Greekmath 011B} _{X}^{4}=\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})^{4}\}-\frac{3}{4}( \mathbb{E}\{(X_{1}-X_{2})^{2}\})^{2} \notag \\ &=&\mathbb{E}\big\{\frac{1}{2}(X_{1}-X_{2})^{4}-\frac{3}{4} (X_{1}-X_{2})^{2}(X_{3}-X_{4})^{2}\big\}\,. \end{eqnarray}

The last two identities in ((ref)) provide simple $D$-representations of the fourth central moment of $X$. We now apply the above results to skewness, kurtosis and excess kurtosis coefficients:

equation[equation omitted — 294 chars of source]

$\mathrm{Kur}(X)$ is also called the \textquotedblleft Pearson kurtosis\textquotedblright\ coefficient, while the excess kurtosis $\mathrm{ EKur}(X)$ is the \textquotedblleft Fisher kurtosis\textquotedblright .

proposition$D$-representation of skewness and kurtosis. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : $D$-representation of skewness and kurtosis}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$ and variance ${\Greekmath 011B} _{X}^{2}$, and $X_{1}$, $X_{2}$, $X_{3}$, $X_{4}$\ four i.i.d. replications of $X$. If $X$ has finite third moment, then \begin{eqnarray} \mathrm{Sk}(X) &=&\frac{\mathbb{E}\{(X_{1}-X_{2})(X_{1}^{2}-X_{2}^{2})\}}{ 2{\Greekmath 011B} _{X}^{3}}-2\frac{{\Greekmath 0116} _{X}}{{\Greekmath 011B} _{X}}=\frac{\mathbb{E} \{(X_{1}-X_{3})(X_{1}-X_{2})^{2}\}}{{\Greekmath 011B} _{X}^{3}} \notag \\ &=&\frac{\sqrt{8}\,\mathbb{E}\{(X_{1}-X_{3})(X_{1}-X_{2})^{2}\}}{(\mathbb{E} \{(X_{1}-X_{2})^{2}\})^{^{3/2}}}=\frac{\sqrt{2}}{3}\frac{\mathbb{E} \{D_{1}^{(3)}D_{2}^{(3)}D_{3}^{(3)}\}}{(\mathbb{E}\{(X_{1}-X_{2})^{2} \})^{^{3/2}}} \end{eqnarray} where $D_{i}^{(3)}$ is defined by $(\ref{eq: Sum of differences})$. If $X$ has finite fourth moment, then \begin{equation} \mathrm{Kur}(X)=\frac{2\mathbb{E}\{(X_{1}-X_{2})^{4}\}}{(\mathbb{E} \{(X_{1}-X_{2})^{2}\})^{2}}-3\,. \end{equation}

If $X_{1}$ and $X_{2}$ are i.i.d. normal, it is easy to see that

equation[equation omitted — 161 chars of source]

so that $\mathrm{EKur}(X)=0$.

Tentative\begin{comment} ******************* Discrete analogues. ******************* \begin{Tentative} \end{comment}

Higher-order Moments

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

In this section, we study how higher-order central moments can be represented in terms of pairwise differences between i.i.d. realizations of a random variable. Consider a random variable $X$ with mean ${\Greekmath 0116} _{X}$ and finite moments up to order $n\geq 1$. We denote the $k$-th central moment of $X$ by

equation[equation omitted — 109 chars of source]

Suppose that $X_{0},$ $X_{1},\ldots ,\,X_{n}$\ are i.i.d. replications of $X$ , and define

equation[equation omitted — 168 chars of source]

Since $\mathbb{E}\{\tilde{X}_{i}\}={\Greekmath 0116} _{1}=0$, for $i=0,\ldots ,\,n$, the independence of $\tilde{X}_{0},\ldots ,\,\tilde{X}_{n}$ entails:

equation[equation omitted — 225 chars of source]

In other words, $P_{k}$ can be viewed as an unbiased estimator of ${\Greekmath 0116} _{k}$ , for $k=1,\ldots ,\,n$ . This immediately shows that all the central moments of $X$ have a $D$-representation for moments up to order $n$, provided $X$ has moments up to order $n$.

The function $P_{n}$ requires $n+1$ replications of $X$ to represent the $n$ -th central moment of $X$. However, for $n=2$, $3,$ $4$, we have given representations which only require $n$ replications (see Section (ref)). We will now show this is indeed feasible for $n\geq 5$. We first state some useful recursions.

propositionRecursions for higher-order moments. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : Recursions for higher-order moments}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, and finite central moments ${\Greekmath 0116} _{j}$ for $j=1,\ldots ,\,n+1$, where $n\geq 1.$ Then \begin{equation} {\Greekmath 0116} _{n+1}=\mathsf{C}[X,\,(X-{\Greekmath 0116} _{X})^{n}]=\mathbb{E}\{X(X-{\Greekmath 0116} _{X})^{n}\}-{\Greekmath 0116} _{X}{\Greekmath 0116} _{n}\,. \end{equation} If $X_{1}$, $X_{2}$, $X_{3}\;$are i.i.d. replications of $X$, then \begin{eqnarray} {\Greekmath 0116} _{n+1} &=&\mathbb{E}\{(X_{1}-X_{2})(X_{1}-{\Greekmath 0116} _{X})^{n}\} \notag \\ &=&\mathbb{E}\{(X_{1}-X_{3})(X_{1}-X_{2})^{n}\}- \sum_{j=2}^{n-1}(-1)^{j}C_{n}^{j}\,{\Greekmath 0116} _{j}{\Greekmath 0116} _{n+1-j} \end{eqnarray} where $C_{n}^{j}=n!/[j!(n-j)!]$ and, for $n+1$ even, \begin{equation} {\Greekmath 0116} _{n+1}=\frac{1}{2}\big[\mathbb{E}\{(X_{1}-X_{2})^{n+1}\}- \sum_{j=2}^{n-1}(-1)^{j}C_{n+1}^{j}\,{\Greekmath 0116} _{j}{\Greekmath 0116} _{n+1-j}\big]\,. \end{equation}

When the summations in ((ref))\thinspace -\thinspace ((ref)) are empty (for $n=1$ and $n=2$), the corresponding formulae yield [as expected from ((ref)) and ( (ref))]:

equation[equation omitted — 203 chars of source]

Now, let $X_{1},\ldots ,\,X_{n}$ be i.i.d. replications of $X$ with $n\geq 5$ , and consider the following sequence of functions:

equation[equation omitted — 151 chars of source]
equation[equation omitted — 121 chars of source]
equation[equation omitted — 173 chars of source]

and for $k\geq 4$,

equation[equation omitted — 271 chars of source]

If $k+1$ is even (with $k\geq 1$), we can also consider the functions

eqnarray[eqnarray omitted — 313 chars of source]

From the above definitions, it is clear that the functions $\bar{{\Greekmath 0116}} _{k}(x_{1},\ldots ,\,x_{k})$ and $\tilde{{\Greekmath 0116}}_{k}(x_{1},\ldots ,\,x_{k})$ depend on their arguments only through differences $x_{i}-x_{j}$ ($ i,j=1,\ldots ,\,k$). Note also that $\bar{{\Greekmath 0116}}_{j}$ can be replaced by $ \tilde{{\Greekmath 0116}}_{j}$ in the summations of ((ref))\thinspace -\thinspace ((ref)) whenever $j$ is even.

propositionRecursive $D$-representations for higher-order moments. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : Recursive $D$-representations for higher-order moments}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, and finite moments up to order $n+1$, where $n\geq 1$, and let $\bar{{\Greekmath 0116}} _{k}(x_{1},\ldots ,\,x_{k})$ and $\tilde{{\Greekmath 0116}}_{k}(x_{1},\ldots ,\,x_{k})$ be defined by $(\ref{eq: muhat(2)})\,$-\thinspace $\,(\ref{eq: mutilde(k)}),$ for $k\geq 1.$ If $X_{1},\ldots ,\,X_{n+1}$ are i.i.d. replications of $X$, then \begin{equation} \mathbb{E}\{\bar{{\Greekmath 0116}}_{k+1}(X_{1},\ldots ,\,X_{k+1})\}={\Greekmath 0116} _{k+1}\,,\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for }k=1,\ldots ,\,n\,, \end{equation} and, if $k+1$ is even, \begin{equation} \mathbb{E}\{\tilde{{\Greekmath 0116}}_{k+1}(X_{1},\ldots ,\,X_{k+1})\}={\Greekmath 0116} _{k+1}\,,\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for\ }1\leq k\leq n\,. \end{equation}

The above proposition shows that the $k$-th central moment of $X$ has a $D$ -representation which only requires $k$ replications of $X$ (provided ${\Greekmath 0116} _{k}$ is finite).

Tentative\begin{comment} ******************* Discrete analogues. ******************* \begin{Tentative} \end{comment}

Generalized Lagrange Identities

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

In this section, we give some generalizations of the Lagrange identity. We first give a moment-based Lagrange-type identity using independent random vectors.

propositionGeneralized Lagrange identity. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : Generalized Lagrange identity}} Let $(X_{1},$ \thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $Y_{2})^{\prime }$ be two independent random vectors with finite second moments. Then \begin{equation} \frac{1}{2}\big(\mathbb{E}\{X_{1}^{2}\}\mathbb{E}\{Y_{2}^{2}\}+\mathbb{E} \{X_{2}^{2}\}\mathbb{E}\{Y_{1}^{2}\}\big)-\mathbb{E}\{X_{1}Y_{1}\}\mathbb{E} \{X_{2}Y_{2}\}=\frac{1}{2}\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}\,. \end{equation}

In Proposition (ref), $(X_{1},$\thinspace $ Y_{1})^{\prime }$ and $(X_{2},$\thinspace $Y_{2})^{\prime }$ need not be identically distributed. Since the right-hand side of ((ref)) is nonnegative, it follows that

equation[equation omitted — 205 chars of source]

and, on replacing each variable $X_{1}$ , $X_{2}$ , $Y_{1}$ , $Y_{2}$ by its deviation from the mean ($X_{1}-\mathbb{E}\{X_{1}\},$ etc.),

equation[equation omitted — 239 chars of source]

The latter inequality holds even if the random variables $X_{1}$ , $X_{2}$ , $Y_{1}$ , $Y_{2}$ have different means and variances. When $(X_{1},$ \thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $Y_{2})^{\prime }$ have the same covariance matrix (though possibly different means), ((ref)) yields the usual Cauchy-Schwarz inequality for covariances:

equation[equation omitted — 139 chars of source]

Interestingly, ((ref)) provides a remarkably simple way of proving the Cauchy-Schwarz inequality for covariances.

The distance $\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}$ can be viewed as a measure of non-proportionality between $Y$ and $X$ : if $Y_{1}=b_{1}X_{1}$ and $Y_{2}=b_{2}X_{2}$, then

equation[equation omitted — 205 chars of source]

with $\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}=0$ when $b_{1}=b_{2}$. When $ \mathbb{E}\{(X_{1}X_{2})^{2}\}\neq 0$, the condition $b_{1}=b_{2}$ is necessary and sufficient for $\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}=0$. The identity ((ref)) holds even if $X_{1}$ and $X_{2}$ do not have the same distribution.

When additional homogeneity restrictions are imposed, we get simpler identities which are spelled out in the following corollary.

corollaryCovariance generalized Lagrange identity. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : Covariance generalized Lagrange identity}} Let $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $ Y_{2})^{\prime }$ be two independent random vectors with finite second moments, $\tilde{X}_{i}:=X_{i}-\mathbb{E}\{X_{i}\}$ and $\tilde{Y} _{i}:=Y_{i}-\mathbb{E}\{Y_{i}\}$, $i=1,2$. Then \begin{equation} \frac{1}{2}({\Greekmath 011B} _{X_{1}}^{2}{\Greekmath 011B} _{Y_{2}}^{2}+{\Greekmath 011B} _{X_{2}}^{2}{\Greekmath 011B} _{Y_{1}}^{2})-\mathsf{C}[X_{1},\,Y_{1}]\mathsf{C}[X_{2},\,Y_{2}]=\frac{1}{2} \mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}_{2}\tilde{Y}_{1})^{2}\}\,. \end{equation} If ${\Greekmath 011B} _{X_{1}}^{2}={\Greekmath 011B} _{X_{2}}^{2}$ and ${\Greekmath 011B} _{Y_{1}}^{2}={\Greekmath 011B} _{Y_{2}}^{2}$, then \begin{equation} {\Greekmath 011B} _{X_{1}}^{2}{\Greekmath 011B} _{Y_{1}}^{2}-\mathsf{C}[X_{1},\,Y_{1}]\mathsf{C} [X_{2},\,Y_{2}]=\frac{1}{2}\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X} _{2}\tilde{Y}_{1})^{2}\}\,. \end{equation} If the two random vectors $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $ (X_{2}, $\thinspace $Y_{2})^{\prime }$ have the same covariance matrix, then \begin{equation} {\Greekmath 011B} _{X_{1}}^{2}{\Greekmath 011B} _{Y_{1}}^{2}-\mathsf{C}[X_{1},\,Y_{1}]^{2}=\frac{1 }{2}\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}_{2}\tilde{Y} _{1})^{2}\}\,. \end{equation}

The identity ((ref)) holds a fortiori when $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$ \thinspace $Y_{2})^{\prime }$ are i.i.d. (with finite second moments). When $ (X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $ Y_{2})^{\prime }$ have the same covariance matrix, the correlation ${\Greekmath 011A} (X_{1},\,Y_{1})$ between $X_{1}$ and $Y_{1}$ satisfies:

equation[equation omitted — 241 chars of source]

Thus, if $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $ Y_{2})^{\prime }$ are i.i.d. replications of the random vector $(X,$ \thinspace $Y)^{\prime },$ ${\Greekmath 011A} (X_{1},\,Y_{1})^{2}$ measures how close two independent replications of $(X,$\thinspace $Y)^{\prime }$ are according to the distance $\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}_{2}\tilde{Y} _{1})^{2}\}$.

The following proposition extends in a similar way the Binet-Cauchy identity ((ref)).

propositionGeneralized Binet-Cauchy identity. \addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}} { : Generalized Binet-Cauchy identity}} Let $Z_{i}:=(A_{i},$\thinspace $B_{i},\,C_{i},\,D_{i})^{\prime }\;$and\ $ \tilde{Z}_{i}:=Z_{i}-\mathbb{E}\{Z_{i}\}=(\tilde{A}_{i},$\thinspace $\tilde{B }_{i},\,\tilde{C}_{i},\,\tilde{D}_{i})^{\prime }$, $i=1,2$, be independent random vectors with finite fourth moments. Then \begin{eqnarray} \mathbb{E}\{(A_{1}B_{2}-A_{2}B_{1})(C_{1}D_{2}-C_{2}D_{1})\} &=&\mathbb{E} \{A_{1}C_{1}\}\mathbb{E}\{B_{2}D_{2}\}+\mathbb{E}\{A_{2}C_{2}\}\mathbb{E} \{B_{1}D_{1}\} \notag \\ &-&[\mathbb{E}\{A_{1}D_{1}\}\mathbb{E}\{B_{2}C_{2}\}+\mathbb{E}\{A_{2}D_{2}\} \mathbb{E}\{B_{1}C_{1}\}]\,.\quad \quad \end{eqnarray} If $(A_{1},$\thinspace $B_{1},\,C_{1},\,D_{1})^{\prime }$ and $(A_{2},$ \thinspace $B_{2},\,C_{2},\,D_{2})^{\prime }$ have the same covariance matrix, then \begin{equation} \mathsf{C}[A_{1},\,C_{1}]\mathsf{C}[B_{1},\,D_{1}]-\mathsf{C}[A_{1},\,D_{1}] \mathsf{C}[B_{1},\,C_{1}]=\frac{1}{2}\mathbb{E}\{(\tilde{A}_{1}\tilde{B}_{2}- \tilde{A}_{2}\tilde{B}_{1})(\tilde{C}_{1}\tilde{D}_{2}-\tilde{C}_{2}\tilde{D} _{1})\}\,. \end{equation}

The above proposition yields restrictions on the covariances of a four-dimensional random vector. From ((ref) ), we see that $A_{1}B_{2}-A_{2}B_{1}=0$ (a.s.) or $C_{1}D_{2}-C_{2}D_{1}=0$ (a.s.) entails

equation[equation omitted — 121 chars of source]

On taking\ $A_{i}=C_{i}$\ and\ $B_{i}=D_{i}\;$for $i=1,$\thinspace $2$, we get the Lagrange identity:

equation[equation omitted — 198 chars of source]

Applications

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

The $D$-representations given above show that the central moments of a real random variable $X$ can all be represented as expected values of polynomial functions in differences between a small number of replications of $X$. For $ k\geq 2,$ the $k$-th central moment of $X$ only requires $k$ replications $ X_{1},\ldots ,\,X_{k}$. More precisely, we can write:

equation[equation omitted — 215 chars of source]

where

equation[equation omitted — 66 chars of source]
equation[equation omitted — 106 chars of source]
equation[equation omitted — 129 chars of source]
align[align omitted — 211 chars of source]

In other words, $h_{k}(X_{1},\ldots ,\,X_{k})$ is an unbiased estimator of $ {\Greekmath 0116} _{k}$ based on $k$ observations.

If $n$ observations $X_{1},\ldots ,\,X_{n}$ are available, any (fixed or randomly selected) subset of $k$ observations from $\{X_{1},\ldots ,\,X_{n}\} $ yields an unbiased estimator of ${\Greekmath 0116} _{k},$ and similarly for the arithmetic (or a weighted) average of such subsets. In particular, the arithmetic average of $h_{k}(X_{i_{1}},\ldots ,\,X_{i_{k}})$ over $N$ randomly selected $\{i_{1},\ldots ,\,i_{k}\}$ of $\{1,\ldots ,\,n\}$ is an unbiased estimator of ${\Greekmath 0116} _{k}$. A fortiori, this will also hold if the average is taken over all subsets of $k$ observations from $ \{X_{1},\ldots ,\,X_{n}\}$, but the large number of such subsets can easily make this approach computationally prohibitive. We will call such estimators based on observation differences $D$-estimators.

It is of interest to compare $D$-estimators of ${\Greekmath 0116} _{k}$ with the \textquotedblleft natural\textquotedblright\ estimators based on deviations from the sample mean:

equation[equation omitted — 135 chars of source]

For $k=2,$ $\hat{m}_{k}(n)$ is unbiased, but it is not generally unbiased for $k\geq 3$. To see how big the bias difference can be, we take $n=k$ and compare by simulation $\hat{{\Greekmath 0116}}_{k}(k)$ with the corresponding $D$ -estimator with $h_{k}(X_{1},\ldots ,\,X_{k});$ see Table (ref). We use the Exponential distribution as an example because its central moments are nonzero at all orders, which helps us assess estimator bias without interference from sign-canceling or offsetting effects that occur when the true moment is zero. We see from these results that $\hat{m}_{k}(k)$ is heavily biased, while the corresponding $D$-estimators exhibit no bias as predicted by the results presented above.

table[table omitted — 1,352 chars of source]

Of course, if $n>k$, more efficient unbiased estimators can be obtained by averaging several minimal estimators. This can be easily implemented through computer programming by selecting the index tuples $(i_{1},\ldots ,i_{k+1})$ independently and identically distributed (i.i.d.) from a uniform distribution over the index set $I$. Due to its stochastic nature, we refer to this estimator as a Monte Carlo approximation of the $D$-estimator, $\hat{ {\Greekmath 0116}}_k(n)^{MC}$. Table (ref) presents simulation results comparing the bias of the \textquotedblleft natural\textquotedblright\ estimators $\hat{m}_k(n)$ with the proposed $D$-estimator Monte Carlo approximation for central moments of orders $k=2,\ldots ,8$ under an $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Exponential}(2)$ distribution. For each replication, we generate 30,000 Monte Carlo samples.

table[table omitted — 1,641 chars of source]

The simulation results demonstrate that the proposed $D$-estimator Monte Carlo approximation substantially reduces bias across all orders of the central moment compared to the natural estimator. This improvement is particularly notable for higher-order moments where the naive estimator exhibits larger negative bias. Moreover, increasing the sample size from 50 to 100 decreases the bias for both estimators, but the $D$-estimator approach consistently outperforms the natural estimator at each sample size. These findings highlight the effectiveness of the Monte Carlo approximation in providing more accurate and unbiased estimates of central moments, especially for higher-order moments that are typically more challenging to estimate.

Tentative\begin{comment} It is of interest to compare $D$-estimators of ${\Greekmath 0116} _{k}$ with the \textquotedblleft natural\textquotedblright\ estimators based on deviations from the sample mean: \begin{equation} \hat{m}_{k}(n)=\frac{1}{n-1}\sum_{i=i}^{n}(X_{i}-\bar{X}_{n})^{k}\,,\quad \bar{X}_{n}=\frac{1}{n}\sum_{i=i}^{n}X_{i}\,. \end{equation} For $k=2,$ $\hat{m}_{k}(n)$ is unbiased, but it is not generally unbiased for $k\geq 3$. To see how big the bias difference can be, we take $n=k$ and compare by simulation $\hat{{\Greekmath 0116}}_{k}(k)$ with the corresponding $D$ -estimator with $h_{k}(X_{1},\ldots ,\,X_{k});$ see Table (ref). We see from these results that $\hat{m}_{k}(k)$ is heavily biased, especially for non-zero central moments, while the corresponding $D$-estimators exhibit no bias as predicted by the results presented above. \begin{table}[tbph] \caption{Add caption} \resizebox{\columnwidth}{!}{ \begin{tabular}{rlrrrrrrr} \toprule \multicolumn{9}{c}{ Table 1. Simulation Comparison of Bias for Two Distributions when n=k (\# of Simulation = 20000000) } \\ \midrule & Central Moment & \multicolumn{1}{l}{Ord. 2} & \multicolumn{1}{l}{Ord. 3} & \multicolumn{1}{l}{Ord. 4} & \multicolumn{1}{l}{Ord. 5} & \multicolumn{1}{l}{\textbf{Ord. 6}} & \multicolumn{1}{l}{\textbf{Ord. 7}} & \multicolumn{1}{l}{\textbf{Ord. 8}} \\ & \textbf{Number of Observations} & \textbf{2} & \textbf{3} & \textbf{4} & \textbf{5} & \textbf{6} & \textbf{7} & \textbf{8} \\ \midrule \multicolumn{1}{l}{\textbf{Exponential(2)}} & \textit{True Value} & \textit{0.25} & \textit{0.25} & \textit{0.56} & \textit{1.38} & \textit{4.00} & \textit{14.48} & \textit{57.94} \\ \cmidrule{2-9} & \textbf{Natural Estimator $\hat{m}_k(k)$} & \textbf{0.000} & -0.167 & -0.258 & -0.768 & -2.171 & -8.321 & -33.739 \\ &\textbf{D-estimators $\hat{{\Greekmath 0116}}_k(k)$} & \textbf{0.000} & \textbf{0.000} & \textbf{0.000} & \textbf{0.003} & \textbf{0.094} & \textbf{0.131} & \textbf{0.053} \\ \midrule \multicolumn{1}{l}{\textbf{Normal(0,1)}} & \textit{True Value} & \textit{1} & \textit{0} & \textit{3} & \textit{0} & \textit{15} & \textit{0} & \textit{105} \\ \cmidrule{2-9} &\textbf{Natural Estimator $\hat{m}_k(k)$} & \textbf{0.000} & \textbf{0.000} & -0.748 & \textbf{0.001} & -4.582 & \textbf{0.009} & -34.650 \\ & \textbf{D-estimators $\hat{{\Greekmath 0116}}_k(k)$} & \textbf{0.000} & -0.001 & \textbf{-0.003} & -0.061 & \textbf{0.246} & -0.292 & \textbf{0.493} \\ \bottomrule \bottomrule \end{tabular} } \end{table} Of course, if $n>k$, more efficient unbiased estimators can be obtained by averaging several minimal estimators. This can be easily implemented through computer programming by selecting the index tuples $(i_{1},\ldots ,i_{k+1})$ independently and identically distributed (i.i.d.) from a uniform distribution over the index set $I$. Due to its stochastic nature, we refer to this estimator as a Monte Carlo approximation of the $D$-estimator, $\hat{ {\Greekmath 0116}}_k(n)^{MC}$. Table (ref) presents simulation results comparing the bias of the \textquotedblleft natural\textquotedblright\ estimators $\hat{m}_k(n)$ with the proposed $D$-estimator Monte Carlo approximation for central moments of orders $k=2,\ldots ,8$ under an $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Exponential}(2)$ distribution and a $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Normal}(0,1)$ distribution. For each replication, we generate 30,000 Monte Carlo samples. \begin{table}[htbp] \caption{Add caption} \resizebox{\columnwidth}{!}{ \begin{tabular}{rlrrrrrrr} \toprule \multicolumn{9}{c}{\textbf{Table 2. Simulation Comparison of Bias for Two Distributions when n>k (\# of Simulation = 1000)}} \\ \midrule & \textbf{Central Moment} & \multicolumn{1}{l}{\textbf{Ord.2}} & \multicolumn{1}{l}{\textbf{Ord.3}} & \multicolumn{1}{l}{\textbf{Ord.4}} & \multicolumn{1}{l}{\textbf{Ord.5}} & \multicolumn{1}{l}{\textbf{Ord.6}} & \multicolumn{1}{l}{\textbf{Ord.7}} & \multicolumn{1}{l}{\textbf{Ord.8}} \\ \midrule \multicolumn{1}{l}{\textbf{Exponential (2)}} & \textit{True Value} & \textit{0.25} & \textit{0.25} & \textit{0.56} & \textit{1.38} & \textit{4.00} & \textit{14.48} & \textit{57.94} \\ \cmidrule{1-9} \multicolumn{1}{l}{\textbf{n=50}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & -0.004 & -0.018 & -0.059 & -0.240 & -0.921 & -5.203 & -27.483 \\ & \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & \textbf{-0.004} & \textbf{-0.009} & \textbf{-0.035} & \textbf{-0.162} & \textbf{-0.592} & \textbf{-4.048} & \textbf{-22.484} \\ \cmidrule{2-9} \multicolumn{1}{l}{\textbf{n=100}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & \textbf{-0.001} & -0.008 & -0.020 & -0.084 & -0.295 & -2.621 & -16.574 \\ & \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & -0.001 & \textbf{-0.003} & \textbf{-0.005} & \textbf{-0.035} & \textbf{-0.133} & \textbf{-1.991} & \textbf{-13.020} \\ \midrule & & & & & & & & \\ \midrule \multicolumn{1}{l}{\textbf{N(0,1)}} & \textit{True Value} & \textit{1} & \textit{0} & \textit{3} & \textit{0} & \textit{15} & \textit{0} & \textit{105} \\ \cmidrule{1-9} \multicolumn{1}{l}{\textbf{n=50}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & 0.003 & -0.003 & -0.055 & \textbf{0.014} & -0.767 & \textbf{0.403} & -11.345 \\ & \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & \textbf{0.002} & \textbf{-0.003} & \textbf{0.000} & 0.041 & \textbf{-0.287} & 0.516 & \textbf{5.071} \\ \cmidrule{2-9} \multicolumn{1}{l}{\textbf{n=100}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & -0.002 & \textbf{0.002} & -0.036 & \textbf{0.081} & -0.381 & \textbf{1.355} & -5.460 \\ & \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & \textbf{-0.002} & 0.003 & \textbf{-0.007} & 0.118 & \textbf{-0.228} & 2.115 & \textbf{-3.893} \\ \bottomrule \bottomrule \end{tabular} } \end{table} The simulation results demonstrate that the proposed $D$-estimator Monte Carlo approximation substantially reduces bias across all orders of the central moment compared to the natural estimator in Exponential distributions. This improvement is particularly notable for higher-order moments where the natural estimator exhibits larger negative bias. Moreover, increasing the sample size from 50 to 100 decreases the bias for both estimators, but the $D$-estimator approach consistently outperforms the natural estimator at each sample size. For Normal(0,1) distributions where the odd order central moments are zero, the performance of estimator may be partially masked by sign calcellation in the simulation data. In cases of non-zero central moments, the $D$-estimator approach outperforms the natural estimator. These findings highlight the effectiveness of the Monte Carlo approximation in providing more accurate and unbiased estimates of central moments, especially for higher-order moments that are typically more challenging to estimate. \begin{Tentative} \end{comment} \end{Tentative} From ((ref)), one can see that several choices of minimal estimator $h_{3}$ are possible (and similarly for higher-order moments), which may have different efficiency properties. The above results can also be used to build unbiased estimators of cumulants as well as to analyze which observations deviate from the rest of the sample. The latter application can be based on analyzing the population of subgroups of observations $(X_{i_{1}},\ldots ,\,X_{i_{k}})$ along with the corresponding unbiased estimates $h_{k}(X_{i_{1}},\ldots ,\,X_{i_{k}})$. Discussing in detail variants of $D$-estimators, as well as other applications, goes beyond the scope of this paper.

Conclusion

{\Alph{section}} {\Alph{section}} \setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}

In this paper, following an approach (apparently) initiated by \nocite* {Gini(1912)} for the variance, we have studied the general problem of representing the central moments of a random variable in terms of pairwise differences between i.i.d. replications, without reference to a location parameter. We first gave pairwise-difference representations (defined formally as $D$-representations) for the third and fourth central moments of a general random variable. These yield intuitive interpretations of the familiar skewness and kurtosis coefficients. For general moments, we have observed that central moments can be interpreted as covariances between a random variable and powers of the same variable (a nonlinear relation). These then provide recursions which link the $D$ -representation of any moment to lower order ones, hence $D$-representations for all central moments (as long as they are finite). Through a similar approach, analogues of Lagrange and Binet-Cauchy identities were established for general random variables. These provide a simple derivation of the classic Cauchy-Schwarz inequality for covariances. Finally, it is of interest to note that the formulas given in this paper can be interpreted as applications of the \textquotedblleft coupling\textquotedblright\ method [see \nocite*{Thorisson(2000)}], which opens the way to additional applications.

Acknowledgements

Cette recherche a b\'{e}n\'{e}fici\'{e} du support financier de la Chaire William Dow (McGill University), de la Banque du Canada (Bourse de recherche), d'une Bourse Guggenheim, de la Bourse Konrad-Adenauer (Fondation Alexander von Humboldt, Allemagne), du Conseil de recherche en sciences humaines du Canada, du Conseil de recherche en sciences naturelles et en g\'{e}nie du Canada, du R\'{e}seau canadien de centres d'excellence (projet MITACS), et du Fonds de recherche sur la soci\'{e}t\'{e} et la culture (Qu\'{e}bec).