Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Article title for specific publication (Journal)
Slides\begin{comment}
\begin{Slides}
\end{comment}
\end{Slides}
\begin{Tentative}
\begin{comment}
\pagenumbering{Roman}\setcounter{page}{1}
\begin{center}
PLAN
\quad
\end{center}
\begin{enumerate}
• To do
\begin{enumerate}
• Abandon the terminology \textquotedblleft decoupling\textquotedblright
: use $D$-representation. Done.
• Introduce the notion of $D$-statistics
• The first differences have symmetric distributions centered at zero.
• Relate to symmetrization inequalities.
• Relate to the coupling method.
• Moment decompositions based on Gini representations. Relate to
Samuelson inequality.
• Sign-skewness statistics.
• Link with Spearman and Kendall measures of dependence.
• Improve the results on higher-order moments.
• Relate to:
\begin{enumerate}
• $h$-statistics;
• $U$-statistics.
\end{enumerate}
• Unbiased estimation of higher-order moments
\begin{enumerate}
• Document the bias of natural sample moments
• Relate to Heffernan (1997, JRSS B).
\end{enumerate}
• Improved procedures based on unbiased estimators of higher-order
moments.
• Relate to cumulative distributions.
• Applications
\begin{enumerate}
• Conceptual: pairwise differences vs. deviations from the mean
• Categorical data
• First differences are not contaminated by the location estimator
• Representations of empirical moments in terms of symmetric random
variables
• Alternative method for deriving unbiased estimators of moments
• Decompositions of empirical moments in terms of observations
• Tests based on pairwise differences
• Filtered and robust moments
• Multivariate extensions
• Lorenz curves in finance
• Conditioning
• Ranks as $U$ -statistics
\end{enumerate}
\end{enumerate}
\end{enumerate}
\quad
Potential journal for submission
\begin{enumerate}
• The American Statistician
• Statistics and Probability Letters
• Communications in Statistics
• Econometrics and Statistics
• Journal of Statistics and Data Science Education
• Teaching Statistics
\end{enumerate}
\begin{Tentative}
\end{comment}
\end{Tentative}
\pagestyle{plain}\pagenumbering{roman}\setcounter{page}{1}
\begin{center}
ABSTRACT
\quad
\end{center}
We provide pairwise-difference (Gini-type) representations of
higher-order central moments for both general random variables and empirical
moments. Such representations do not require a measure of location. For
third and fourth moments, this yields pairwise-difference representations of
skewness and kurtosis coefficients. We show that all central moments possess
such representations, so no reference to the mean is needed for moments of
any order. This is done by considering i.i.d. replications of the
random variables considered, by observing that central moments can be
interpreted as covariances between a random variable and powers of the same
variable, and by giving recursions which link the pairwise-difference
representation of any moment to lower order ones. Numerical summation
identities are deduced. Through a similar approach, we give analogues of the
Lagrange and Binet-Cauchy identities for general random variables, along
with a simple derivation of the classic Cauchy-Schwarz inequality for
covariances. Finally, an application to unbiased estimation of centered
moments is discussed.
\quad
Key words: Gini; Covariance; Skewness; Kurtosis; Pairwise
differences; Lagrange identity; Moments; Unbiased estimation.
\begin{comment}
\end{comment}
\quad
\begin{Tentative}
\begin{comment}
\listoftables
\listoffigures
\begin{Tentative}
\end{comment}
\end{Tentative}
\begin{comment}
Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes
\@mkboth{\MakeUppercase Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes}
{\MakeUppercase Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes}
\@starttoc{lth}
\addcontentsline{toc}{section}{ Liste de d\'efinitions, hypoth\`eses, propositions et th\'eor\`emes}
\end{comment}
\pagenumbering{arabic} \setcounter{section}{0} \setcounter{page}{1}
Introduction
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
The sample variance is certainly the most widely used measure of dispersion
among observations in statistics. It is typically defined as the average of
the squared deviations of the observations $x_{1},\ldots ,\,x_{n}$ from
their sample mean:
equation[equation omitted — 126 chars of source]
where $\bar{x}=\overset{n}{\underset{i=1}{\sum }}x_{i}/n$. At first sight,
the variance depends crucially on \textquotedblleft
centering\textquotedblright\ the observations with respect to their mean. A
century ago, however, \nocite*{Gini(1912)} noticed that $s_{X}^{2}(n)$ can be
rewritten as the average of pairwise differences between all the
observations:
equation[equation omitted — 274 chars of source]
see \nocite*{Heffernan(1988)}. The Gini representation underscores that $
s_{X}^{2}(n)$ does not depend on using the sample mean as location
parameter, and thus may be viewed as a measure of \textquotedblleft
intrinsic variability\textquotedblright\ between observations. In
particular, it provides a way of measuring dispersion when there is no
natural location parameter, as with categorical data; see
\citename{Light-Margolin(1971)} (\citeyear*{Light-Margolin(1971)}, \citeyear*{Light-Margolin(1974)})
. Further, by focusing on pairwise differences $x_{i}-x_{j}$, it may be
easier to determine which observations have the greatest influence on $
s_{X}^{2}(n)$, because deviations from the mean $x_{i}-\bar{x}$ depend on
all the observations (through $\bar{x})$. A comprehensive review of
statistical methods based on the Gini approach is presented by \nocite*
{Yitzhaki-Schechtman(2013)}.
If we have bivariate observations $z_{i}:=(x_{i},$\thinspace $y_{i})^{\prime
}$, $i=1,\ldots ,\,n,$ the sample covariance
equation[equation omitted — 136 chars of source]
is in turn the most widely used \textquotedblleft measure of
association\textquotedblright\ between $x$ and $y.$ For the covariance, a
pairwise representation is provided by the formula
equation[equation omitted — 291 chars of source]
see \nocite*{Hayes(2011)}. Clearly, ((ref)) is a
special case of ((ref)) obtained by taking $
x_{i}=y_{i},\;i=1,\ldots ,\,n.$ Formula ((ref)
) shows that the sample covariance measures the tendency of $x$ and $y$ to
move in the same direction, without reference to sample means. Similar
expressions for the variances and covariances of general random variables
have also been used in the literature on $U$-statistics; see \nocite*[Chapter 1]
{LeeAJ(1990)}. Since centering is not needed, formula ((ref)) provides a basis for developing alternative measures of
dependence, such as Kendall's and Spearman's measures of dependence [see
\nocite*{Lehmann(1966)}]. Again, pairwise differences $x_{i}-x_{j}$ and $
y_{i}-y_{j}$ may be easier to interpret than $x_{i}-\bar{x}$ and $y_{i}-\bar{
y}$, because $\bar{x}$ and $\bar{y}$ depend on all the observations.
Another similar result we consider here is the Lagrange identity:
gather[gather omitted — 505 chars of source]
see \nocite*{Wright(1992)} and \nocite*[Chapter 3]{Steele(2004)}. An interesting
feature of this identity is that it involves pairwise comparisons between
cross-products of the components of $x:=(x_{1},\ldots ,\,x_{n})^{\prime }$
and\ $y:=(y_{1},\ldots ,\,y_{n})^{\prime }$. Since the right-hand side is
non-negative, the Lagrange identity provides a simple way of proving the
discrete Cauchy-Schwarz inequality:
equation[equation omitted — 190 chars of source]
Interestingly, ((ref)) directly shows when the bound
is an equality $(x_{i}y_{j}=x_{j}y_{i}\;$for all $i$ and $j)$ and what
determines its tightness. The Lagrange identity is itself a special case of
the Binet-Cauchy identity
equation[equation omitted — 353 chars of source]
where $a_{i},\,b_{i},\,c_{i},\,d_{i},$ $i=1,\ldots ,\,n$, are arbitrary real
constants [see \nocite*[p. 49]{Steele(2004)}]. Formula ((ref)) relates several inner products to pairwise comparisons between
the components of the vectors considered.
Since the representations in ((ref)) and ((ref)) only depend on differences between
observations, we will call such representations (and similar ones) $D$
-representations (or Gini-type representations). In this paper, we
provide $D$-representations of higher-order central moments for both
general random variables and empirical moments. For third and fourth
moments, this yields new representations of skewness and kurtosis
coefficients, where reference to a location parameter is eliminated. We show
that all central moments possess $D$-representations, so again no
reference to the mean is needed for moments of any order. This is done by
considering i.i.d. replications (or realizations) of the random
variables considered, by observing that central moments can be interpreted
as covariances between a random variable and powers of the same variable (a
\emph{nonlinear} relation), and by giving \emph{recursions} which link the $
D $-representation of any moment to lower order ones. Numerical summation
identities similar to ((ref))\thinspace
-\thinspace ((ref)) are also deduced.
Finally, through a similar approach, we give the analogues of the Lagrange
and Binet-Cauchy identities for general random variables, along with a
simple derivation of the classic Cauchy-Schwarz inequality for covariances.
The paper is organized as follows. In Section (ref), we discuss $D$-representations for covariances and
regression coefficients, and we show how corresponding numerical identities
can be derived from these. In Section (ref), we
provide $D$-representations for third and fourth moments, along with
corresponding expressions for skewness and kurtosis coefficients. In Section
(ref), we show that all moments possess such
representations, and we give recursive formulae for computing a $D$
-representation for any moment. Generalized Lagrange and Binet-Cauchy
identities are given in Section (ref).
In Section (ref), we discuss some applications of the
proposed moment representations. We conclude in Section (ref)
. Proofs are given in an online Appendix.
$D$-representation of Covariances
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
In order to generalize the expressions ((ref)
)\thinspace -\thinspace ((ref)) to
higher-order moments, it will be useful to discuss the pairwise difference
representation of the covariance between general random variables. Let $X$
and $Y$ be random variables with finite second moments, and means ${\Greekmath 0116} _{X}=
\mathbb{E}\{X\}$, ${\Greekmath 0116} _{Y}=\mathbb{E}\{Y\}.$ We denote by $\mathsf{C}
[X,\,Y]:=\mathbb{E}\{(X-{\Greekmath 0116} _{X})(Y-{\Greekmath 0116} _{Y})\}$ the covariance between $X$
and $Y$. Consider two independent and identically distributed replications
of the random vector $(X,\,Y)^{\prime }$: $\;(X_{1},\,Y_{1})^{\prime }$ and $
(X_{2},\,Y_{2})^{\prime }$. On observing
eqnarray[eqnarray omitted — 424 chars of source]
it follows that
equation[equation omitted — 104 chars of source]
It is also easy to see that
eqnarray[eqnarray omitted — 213 chars of source]
In other words, only one of the variables in $(X,$\thinspace $Y)^{\prime }$
need be replicated to avoid centering with respect to the means (${\Greekmath 0116} _{X}$
and ${\Greekmath 0116} _{Y}$). ((ref))\thinspace -\thinspace (
(ref)) also hold under the weaker assumption (without
identical distributions) that $(X_{1},\,Y_{1})^{\prime }$ and $
(X_{2},\,Y_{2})^{\prime }$ have the same first and second moments as $
(X,\,Y) $ with
equation[equation omitted — 100 chars of source]
For $X_{1}=Y_{1}$ and $X_{2}=Y_{2},$ we obtain the following formulae for
the variance of $X:$
eqnarray[eqnarray omitted — 203 chars of source]
If $X_{3}$ is a third replication of $X$ [so that $X_{1}$, $X_{2}$, $X_{3}$
are i.i.d.], it is also easy to check that
equation[equation omitted — 108 chars of source]
The uncentered second moment of $X$ can be written as
eqnarray[eqnarray omitted — 244 chars of source]
and the linear regression coefficient of $Y$ on $X$ is
eqnarray[eqnarray omitted — 355 chars of source]
In view of the above discussion on the covariance and the variance, we will
define more formally what we mean by the $D$-representation of a parameter.
definitionA parameter ${\Greekmath 010D} $ has a $D$-representation for the variables $
X_{1},\ldots ,\,X_{n}$ if there is a function $g(x_{1},\ldots
,\,x_{n})=h(x_{2}-x_{1},\ldots ,\,x_{n}-x_{n-1})$ such that
\begin{equation}
{\Greekmath 010D} =\mathbb{E}\{h(X_{2}-X_{1},\ldots ,\,X_{n}-X_{n-1})\}\,.
\end{equation}
In other words, ${\Greekmath 010D} $ has a $D$-representation for the variables $
X_{1},\ldots ,\,X_{n}$ if it can be written as the expected value of a
function $g(X_{1},\ldots ,\,X_{n})$ which depends on $X_{1},\ldots ,\,X_{n}$
only through the differences between the variables $X_{1},\ldots ,\,X_{n}$.
Equivalently, this means that
equation[equation omitted — 73 chars of source]
where $g(x_{1},\ldots ,\,x_{n})$ is invariant to uniform location changes,
i.e.
equation[equation omitted — 152 chars of source]
A potentially useful feature of the representation ((ref)) follows
from the fact that the variables $(X_{1}-X_{2})$ and $(Y_{1}-Y_{2})$ have
marginal distributions symmetric around zero. Consequently, the odd moments
of $(X_{1}-X_{2})$ and $(Y_{1}-Y_{2})$ are all zero (when they exist). More
generally, the distribution of the vector $(X_{1}-X_{2},\,Y_{1}-Y_{2})^{
\prime }$ is jointly symmetric around zero in the sense that
equation[equation omitted — 97 chars of source]
This type of symmetry can be useful for proving unbiasedness results in
various setups; see \citename{Dufour(1984)} (\citeyear*{Dufour(1984)}, \citeyear*{Dufour(1985)}). This property is
shared by all $D$-representations when the paired differences are based on
i.i.d. observations.
It is of interest to note that the numerical identity ((ref)) can be derived from the general moment formula ((ref)). Let $(I,$\thinspace $J)$ be a pair of random integers such that
equation[equation omitted — 81 chars of source]
and consider the random vector
equation[equation omitted — 58 chars of source]
The values $x_{1},\ldots ,\,x_{n}$ and $y_{1},\ldots ,\,y_{n}$ need not be
distinct. It is then easy to see that:
equation[equation omitted — 190 chars of source]
equation[equation omitted — 127 chars of source]
and by the definition of the covariance,
equation[equation omitted — 209 chars of source]
Consider now two i.i.d. replications of $(X^{\ast },\,Y^{\ast })$:\ \ $
(X_{1}^{\ast },\,Y_{1}^{\ast })$ and $(X_{2}^{\ast },\,Y_{2}^{\ast })$. By
applying ((ref)) and ((ref)) to $(X^{\ast },$\thinspace $
Y^{\ast })$, the covariance between $X^{\ast }$ and $Y^{\ast }$ is given by:
equation[equation omitted — 224 chars of source]
Since the above sum contains a zero whenever $i=j$, it is natural to divide
by $(n^{2}-n)\;$rather than $n^{2}$. This leads to:
equation[equation omitted — 191 chars of source]
In the special case where $X^{\ast }=Y^{\ast },$ we have
equation[equation omitted — 182 chars of source]
hence, on setting $s_{X^{\ast }}^{2}:=(\frac{n}{n-1}){\Greekmath 011B} _{X^{\ast }}^{2}$
, the Gini formula ((ref)) for the variance is:
equation[equation omitted — 123 chars of source]
Skewness and Kurtosis
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
In this section, we express skewness and kurtosis coefficients in terms of
pairwise differences. Such representations can be written as functions of
differences of the form $X_{i}-X_{j}$ or differences between powers of
variables [$X_{i}^{2}-X_{j}^{2}$]. We first underscore that the third and
fourth central moments of a random variable $X$ can be interpreted as
covariances, which in turn can be represented in terms of pairwise
differences between replicated random variables. In each case, we give
several expressions which may have independent interest.
propositionCovariance representation of third and
fourth central moments.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : Covariance representation of third and
fourth central moments}} Let $X$ be a random variable with finite mean ${\Greekmath 0116}
_{X}$ and variance ${\Greekmath 011B} _{X}^{2}$. Let $X_{1}$, $X_{2}$ two i.i.d.
replications of $X$. If\ $X$ has finite third moment, then
\begin{eqnarray}
\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\} &=&\mathsf{C}[X,\,(X-{\Greekmath 0116} _{X})^{2}]=\mathsf{C}
[X,\,X^{2}]-2{\Greekmath 0116} _{X}{\Greekmath 011B} _{X}^{2} \notag \\
&=&\mathbb{E}\{(X_{1}-X_{2})(X_{1}-{\Greekmath 0116} _{X})^{2}\}=\mathbb{E}
\{(X_{1}-X_{2})X_{1}^{2}\}-2{\Greekmath 0116} _{X}{\Greekmath 011B} _{X}^{2} \notag \\
&=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})[(X_{1}-{\Greekmath 0116} _{X})^{2}-(X_{2}-{\Greekmath 0116}
_{X})^{2}]\} \notag \\
&=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})(X_{1}^{2}-X_{2}^{2})\}-{\Greekmath 0116} _{X}
\mathbb{E}\{(X_{1}-X_{2})^{2}\}\,.
\end{eqnarray}
If $X$ has finite fourth moment, then
\begin{eqnarray}
\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{4}\} &=&\mathsf{C}[X,\,(X-{\Greekmath 0116} _{X})^{3}]=\mathsf{C}
[X,\,X^{3}]-3{\Greekmath 0116} _{X}\mathsf{C}[X,\,X^{2}]+3{\Greekmath 0116} _{X}^{2}{\Greekmath 011B} _{X}^{2}
\notag \\
&=&\mathbb{E}\{(X_{1}-X_{2})(X_{1}-{\Greekmath 0116} _{X})^{3}\}=\mathbb{E}
\{(X_{1}-X_{2})(X_{1}^{2}-2{\Greekmath 0116} _{X}X_{1})\} \notag \\
&=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})[(X_{1}-{\Greekmath 0116} _{X})^{3}-(X_{2}-{\Greekmath 0116}
_{X})^{3}]\} \notag \\
&=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})(X_{1}^{3}-X_{2}^{3})\}-3{\Greekmath 0116} _{X}
\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}-3{\Greekmath 0116} _{X}^{2}{\Greekmath 011B} _{X}^{2}\,.
\end{eqnarray}
To interpret Proposition (ref), consider first the case where the mean is zero ($
{\Greekmath 0116} _{X}=0$), so the third moment does not depend on the mean or the
variance. When the distribution is skewed to the right [$\mathbb{E}\{(X-{\Greekmath 0116}
_{X})^{3}\}>0$\}, equation ((ref)) shows that an
increase in the value of $X$ ($X_{1}-X_{2}>0$) is associated with an
increase in the absolute value of $X$ [$X_{1}^{2}-X_{2}^{2}>0$], while a
shift to the left would go with a decrease in the absolute value of $X$.
More generally, $\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}$ can be written as
equation[equation omitted — 222 chars of source]
so that $\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}$ measures the association between a
change in the deviation from the mean and the corresponding change in the
absolute values of the deviations. A positive (negative) association entails
positive (negative) skewness, which is clearly exhibited by pairwise
differences.
The formulae given by Proposition (ref) depend on unknown parameters (${\Greekmath 0116} _{X}$
and ${\Greekmath 011B} _{X}^{2}$). It is possible to avoid this by using a larger
number of replications. For the third moment, only one additional
replication of $X$ is needed. If $X_{1}$, $X_{2}$, $X_{3}$ are three i.i.d.
replications of $X$, we see [from ((ref))] that
eqnarray[eqnarray omitted — 311 chars of source]
For the fourth moment, this can be done with four replications. If $X_{1}$, $
X_{2}$, $X_{3}$, $X_{4}$ are four i.i.d. replications of $X$, we have:
align*[align* omitted — 289 chars of source]
equation[equation omitted — 200 chars of source]
((ref)) - ((ref))
still do not provide $D$-representations, because $X_{3}$ and $X_{4}$ are in
levels. This is done in the following propositions by using additional
replicates of $X$. A number of alternative expressions are provided in the
following proposition.
proposition$D$-representation of third central moment.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : $D$-representation of third central moment}}
Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, variance ${\Greekmath 011B} _{X}^{2}$
, and finite third moment. If $X_{1}$, $X_{2}$, $X_{3}$ are three i.i.d.
replications of $X$, then
\begin{eqnarray}
\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\} &=&\mathbb{E}\{(X_{1}-X_{3})(X_{1}-X_{2})^{2}
\}=\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})[(X_{1}-X_{3})^{2}-(X_{2}-X_{3})^{2}]
\} \notag \\
&=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})^{2}[(X_{1}-X_{3})+(X_{2}-X_{3})]\}=
\frac{1}{6}\mathbb{E}\{D_{1}^{(3)}D_{2}^{(3)}D_{3}^{(3)}\} \notag \\
&=&\frac{9}{2}\mathbb{E}\{(X_{1}-\bar{X}^{(3)})(X_{2}-\bar{X}^{(3)})(X_{3}-
\bar{X}^{(3)})\}\quad \quad
\end{eqnarray}
where\ \ $\bar{X}^{(3)}:=\,\underset{j=1}{\overset{3}{\sum }}X_{j}/3\;$and
\begin{equation}
D_{i}^{(3)}:=\,\underset{j\neq i}{\sum }(X_{i}-X_{j})=3(X_{i}-\bar{X}
^{(3)})\,,\quad i=1,\,2,\,3.
\end{equation}
As the variables $D_{i}^{(3)}$ are sums of pairwise differences, ((ref)) provides simple $D$-representations of the
third central moment of $X$. Note also that $\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{3}\}$
can be rewritten as a function of deviations with respect to a
\textquotedblleft reference\textquotedblright\ value $X_{3}$:
eqnarray[eqnarray omitted — 262 chars of source]
We now consider the fourth central moment.
proposition$D$-representation of fourth central
moment.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : $D$-representation of fourth central
moment}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, variance ${\Greekmath 011B}
_{X}^{2}$, and finite fourth moment. If $X_{1}$, $X_{2}$, $X_{3}$, $X_{4}$
four i.i.d. replications of $X$, then
\begin{eqnarray}
\mathbb{E}\{(X-{\Greekmath 0116} _{X})^{4}\} &=&\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})^{4}
\}-3{\Greekmath 011B} _{X}^{4}=\frac{1}{2}\mathbb{E}\{(X_{1}-X_{2})^{4}\}-\frac{3}{4}(
\mathbb{E}\{(X_{1}-X_{2})^{2}\})^{2} \notag \\
&=&\mathbb{E}\big\{\frac{1}{2}(X_{1}-X_{2})^{4}-\frac{3}{4}
(X_{1}-X_{2})^{2}(X_{3}-X_{4})^{2}\big\}\,.
\end{eqnarray}
The last two identities in ((ref))
provide simple $D$-representations of the fourth central moment of $X$. We
now apply the above results to skewness, kurtosis and excess kurtosis
coefficients:
equation[equation omitted — 294 chars of source]
$\mathrm{Kur}(X)$ is also called the \textquotedblleft Pearson
kurtosis\textquotedblright\ coefficient, while the excess kurtosis $\mathrm{
EKur}(X)$ is the \textquotedblleft Fisher kurtosis\textquotedblright .
proposition$D$-representation of skewness and kurtosis.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : $D$-representation of skewness and kurtosis}}
Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$ and variance ${\Greekmath 011B}
_{X}^{2}$, and $X_{1}$, $X_{2}$, $X_{3}$, $X_{4}$\ four i.i.d. replications
of $X$. If $X$ has finite third moment, then
\begin{eqnarray}
\mathrm{Sk}(X) &=&\frac{\mathbb{E}\{(X_{1}-X_{2})(X_{1}^{2}-X_{2}^{2})\}}{
2{\Greekmath 011B} _{X}^{3}}-2\frac{{\Greekmath 0116} _{X}}{{\Greekmath 011B} _{X}}=\frac{\mathbb{E}
\{(X_{1}-X_{3})(X_{1}-X_{2})^{2}\}}{{\Greekmath 011B} _{X}^{3}} \notag \\
&=&\frac{\sqrt{8}\,\mathbb{E}\{(X_{1}-X_{3})(X_{1}-X_{2})^{2}\}}{(\mathbb{E}
\{(X_{1}-X_{2})^{2}\})^{^{3/2}}}=\frac{\sqrt{2}}{3}\frac{\mathbb{E}
\{D_{1}^{(3)}D_{2}^{(3)}D_{3}^{(3)}\}}{(\mathbb{E}\{(X_{1}-X_{2})^{2}
\})^{^{3/2}}}
\end{eqnarray}
where $D_{i}^{(3)}$ is defined by $(\ref{eq: Sum of differences})$. If $X$
has finite fourth moment, then
\begin{equation}
\mathrm{Kur}(X)=\frac{2\mathbb{E}\{(X_{1}-X_{2})^{4}\}}{(\mathbb{E}
\{(X_{1}-X_{2})^{2}\})^{2}}-3\,.
\end{equation}
If $X_{1}$ and $X_{2}$ are i.i.d. normal, it is easy to see that
equation[equation omitted — 161 chars of source]
so that $\mathrm{EKur}(X)=0$.
Tentative\begin{comment}
*******************
Discrete analogues.
*******************
\begin{Tentative}
\end{comment}
Higher-order Moments
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
In this section, we study how higher-order central moments can be
represented in terms of pairwise differences between i.i.d. realizations of
a random variable. Consider a random variable $X$ with mean ${\Greekmath 0116} _{X}$ and
finite moments up to order $n\geq 1$. We denote the $k$-th central moment of
$X$ by
equation[equation omitted — 109 chars of source]
Suppose that $X_{0},$ $X_{1},\ldots ,\,X_{n}$\ are i.i.d. replications of $X$
, and define
equation[equation omitted — 168 chars of source]
Since $\mathbb{E}\{\tilde{X}_{i}\}={\Greekmath 0116} _{1}=0$, for $i=0,\ldots ,\,n$, the
independence of $\tilde{X}_{0},\ldots ,\,\tilde{X}_{n}$ entails:
equation[equation omitted — 225 chars of source]
In other words, $P_{k}$ can be viewed as an unbiased estimator of ${\Greekmath 0116} _{k}$
, for $k=1,\ldots ,\,n$ . This immediately shows that all the central
moments of $X$ have a $D$-representation for moments up to order $n$,
provided $X$ has moments up to order $n$.
The function $P_{n}$ requires $n+1$ replications of $X$ to represent the $n$
-th central moment of $X$. However, for $n=2$, $3,$ $4$, we have given
representations which only require $n$ replications (see Section (ref)). We will now show this is indeed feasible for $n\geq
5$. We first state some useful recursions.
propositionRecursions for higher-order moments.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : Recursions for higher-order moments}} Let $X$
be a random variable with mean ${\Greekmath 0116} _{X}$, and finite central moments ${\Greekmath 0116}
_{j}$ for $j=1,\ldots ,\,n+1$, where $n\geq 1.$ Then
\begin{equation}
{\Greekmath 0116} _{n+1}=\mathsf{C}[X,\,(X-{\Greekmath 0116} _{X})^{n}]=\mathbb{E}\{X(X-{\Greekmath 0116}
_{X})^{n}\}-{\Greekmath 0116} _{X}{\Greekmath 0116} _{n}\,.
\end{equation}
If $X_{1}$, $X_{2}$, $X_{3}\;$are i.i.d. replications of $X$, then
\begin{eqnarray}
{\Greekmath 0116} _{n+1} &=&\mathbb{E}\{(X_{1}-X_{2})(X_{1}-{\Greekmath 0116} _{X})^{n}\} \notag \\
&=&\mathbb{E}\{(X_{1}-X_{3})(X_{1}-X_{2})^{n}\}-
\sum_{j=2}^{n-1}(-1)^{j}C_{n}^{j}\,{\Greekmath 0116} _{j}{\Greekmath 0116} _{n+1-j}
\end{eqnarray}
where $C_{n}^{j}=n!/[j!(n-j)!]$ and, for $n+1$ even,
\begin{equation}
{\Greekmath 0116} _{n+1}=\frac{1}{2}\big[\mathbb{E}\{(X_{1}-X_{2})^{n+1}\}-
\sum_{j=2}^{n-1}(-1)^{j}C_{n+1}^{j}\,{\Greekmath 0116} _{j}{\Greekmath 0116} _{n+1-j}\big]\,.
\end{equation}
When the summations in ((ref))\thinspace -\thinspace ((ref)) are empty (for $n=1$ and $n=2$), the
corresponding formulae yield [as expected from ((ref)) and (
(ref))]:
equation[equation omitted — 203 chars of source]
Now, let $X_{1},\ldots ,\,X_{n}$ be i.i.d. replications of $X$ with $n\geq 5$
, and consider the following sequence of functions:
equation[equation omitted — 151 chars of source]
equation[equation omitted — 121 chars of source]
equation[equation omitted — 173 chars of source]
and for $k\geq 4$,
equation[equation omitted — 271 chars of source]
If $k+1$ is even (with $k\geq 1$), we can also consider the functions
eqnarray[eqnarray omitted — 313 chars of source]
From the above definitions, it is clear that the functions $\bar{{\Greekmath 0116}}
_{k}(x_{1},\ldots ,\,x_{k})$ and $\tilde{{\Greekmath 0116}}_{k}(x_{1},\ldots ,\,x_{k})$
depend on their arguments only through differences $x_{i}-x_{j}$ ($
i,j=1,\ldots ,\,k$). Note also that $\bar{{\Greekmath 0116}}_{j}$ can be replaced by $
\tilde{{\Greekmath 0116}}_{j}$ in the summations of ((ref))\thinspace
-\thinspace ((ref)) whenever $j$ is even.
propositionRecursive $D$-representations for
higher-order moments.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : Recursive $D$-representations for
higher-order moments}} Let $X$ be a random variable with mean ${\Greekmath 0116} _{X}$, and
finite moments up to order $n+1$, where $n\geq 1$, and let $\bar{{\Greekmath 0116}}
_{k}(x_{1},\ldots ,\,x_{k})$ and $\tilde{{\Greekmath 0116}}_{k}(x_{1},\ldots ,\,x_{k})$ be
defined by $(\ref{eq: muhat(2)})\,$-\thinspace $\,(\ref{eq: mutilde(k)}),$
for $k\geq 1.$ If $X_{1},\ldots ,\,X_{n+1}$ are i.i.d. replications of $X$,
then
\begin{equation}
\mathbb{E}\{\bar{{\Greekmath 0116}}_{k+1}(X_{1},\ldots ,\,X_{k+1})\}={\Greekmath 0116} _{k+1}\,,\quad
\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for }k=1,\ldots ,\,n\,,
\end{equation}
and, if $k+1$ is even,
\begin{equation}
\mathbb{E}\{\tilde{{\Greekmath 0116}}_{k+1}(X_{1},\ldots ,\,X_{k+1})\}={\Greekmath 0116} _{k+1}\,,\quad
\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for\ }1\leq k\leq n\,.
\end{equation}
The above proposition shows that the $k$-th central moment of $X$ has a $D$
-representation which only requires $k$ replications of $X$ (provided ${\Greekmath 0116}
_{k}$ is finite).
Tentative\begin{comment}
*******************
Discrete analogues.
*******************
\begin{Tentative}
\end{comment}
Generalized Lagrange Identities
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
In this section, we give some generalizations of the Lagrange identity. We
first give a moment-based Lagrange-type identity using independent random
vectors.
propositionGeneralized Lagrange identity.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : Generalized Lagrange identity}} Let $(X_{1},$
\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $Y_{2})^{\prime }$ be
two independent random vectors with finite second moments. Then
\begin{equation}
\frac{1}{2}\big(\mathbb{E}\{X_{1}^{2}\}\mathbb{E}\{Y_{2}^{2}\}+\mathbb{E}
\{X_{2}^{2}\}\mathbb{E}\{Y_{1}^{2}\}\big)-\mathbb{E}\{X_{1}Y_{1}\}\mathbb{E}
\{X_{2}Y_{2}\}=\frac{1}{2}\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}\,.
\end{equation}
In Proposition (ref), $(X_{1},$\thinspace $
Y_{1})^{\prime }$ and $(X_{2},$\thinspace $Y_{2})^{\prime }$ need not be
identically distributed. Since the right-hand side of ((ref)) is nonnegative, it follows that
equation[equation omitted — 205 chars of source]
and, on replacing each variable $X_{1}$ , $X_{2}$ , $Y_{1}$ , $Y_{2}$ by its
deviation from the mean ($X_{1}-\mathbb{E}\{X_{1}\},$ etc.),
equation[equation omitted — 239 chars of source]
The latter inequality holds even if the random variables $X_{1}$ , $X_{2}$ ,
$Y_{1}$ , $Y_{2}$ have different means and variances. When $(X_{1},$
\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $Y_{2})^{\prime }$
have the same covariance matrix (though possibly different means), ((ref)) yields the usual Cauchy-Schwarz inequality for covariances:
equation[equation omitted — 139 chars of source]
Interestingly, ((ref)) provides a remarkably
simple way of proving the Cauchy-Schwarz inequality for covariances.
The distance $\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}$ can be viewed as a
measure of non-proportionality between $Y$ and $X$ : if $Y_{1}=b_{1}X_{1}$
and $Y_{2}=b_{2}X_{2}$, then
equation[equation omitted — 205 chars of source]
with $\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}=0$ when $b_{1}=b_{2}$. When $
\mathbb{E}\{(X_{1}X_{2})^{2}\}\neq 0$, the condition $b_{1}=b_{2}$ is
necessary and sufficient for $\mathbb{E}\{(X_{1}Y_{2}-X_{2}Y_{1})^{2}\}=0$.
The identity ((ref)) holds even if $X_{1}$
and $X_{2}$ do not have the same distribution.
When additional homogeneity restrictions are imposed, we get simpler
identities which are spelled out in the following corollary.
corollaryCovariance generalized Lagrange identity.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : Covariance generalized Lagrange identity}}
Let $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $
Y_{2})^{\prime }$ be two independent random vectors with finite second
moments, $\tilde{X}_{i}:=X_{i}-\mathbb{E}\{X_{i}\}$ and $\tilde{Y}
_{i}:=Y_{i}-\mathbb{E}\{Y_{i}\}$, $i=1,2$. Then
\begin{equation}
\frac{1}{2}({\Greekmath 011B} _{X_{1}}^{2}{\Greekmath 011B} _{Y_{2}}^{2}+{\Greekmath 011B} _{X_{2}}^{2}{\Greekmath 011B}
_{Y_{1}}^{2})-\mathsf{C}[X_{1},\,Y_{1}]\mathsf{C}[X_{2},\,Y_{2}]=\frac{1}{2}
\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}_{2}\tilde{Y}_{1})^{2}\}\,.
\end{equation}
If ${\Greekmath 011B} _{X_{1}}^{2}={\Greekmath 011B} _{X_{2}}^{2}$ and ${\Greekmath 011B} _{Y_{1}}^{2}={\Greekmath 011B}
_{Y_{2}}^{2}$, then
\begin{equation}
{\Greekmath 011B} _{X_{1}}^{2}{\Greekmath 011B} _{Y_{1}}^{2}-\mathsf{C}[X_{1},\,Y_{1}]\mathsf{C}
[X_{2},\,Y_{2}]=\frac{1}{2}\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}
_{2}\tilde{Y}_{1})^{2}\}\,.
\end{equation}
If the two random vectors $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $
(X_{2}, $\thinspace $Y_{2})^{\prime }$ have the same covariance matrix, then
\begin{equation}
{\Greekmath 011B} _{X_{1}}^{2}{\Greekmath 011B} _{Y_{1}}^{2}-\mathsf{C}[X_{1},\,Y_{1}]^{2}=\frac{1
}{2}\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}_{2}\tilde{Y}
_{1})^{2}\}\,.
\end{equation}
The identity ((ref)) holds a
fortiori when $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$
\thinspace $Y_{2})^{\prime }$ are i.i.d. (with finite second moments). When $
(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $
Y_{2})^{\prime }$ have the same covariance matrix, the correlation ${\Greekmath 011A}
(X_{1},\,Y_{1})$ between $X_{1}$ and $Y_{1}$ satisfies:
equation[equation omitted — 241 chars of source]
Thus, if $(X_{1},$\thinspace $Y_{1})^{\prime }$ and $(X_{2},$\thinspace $
Y_{2})^{\prime }$ are i.i.d. replications of the random vector $(X,$
\thinspace $Y)^{\prime },$ ${\Greekmath 011A} (X_{1},\,Y_{1})^{2}$ measures how close two
independent replications of $(X,$\thinspace $Y)^{\prime }$ are according to
the distance $\mathbb{E}\{(\tilde{X}_{1}\tilde{Y}_{2}-\tilde{X}_{2}\tilde{Y}
_{1})^{2}\}$.
The following proposition extends in a similar way the Binet-Cauchy identity
((ref)).
propositionGeneralized Binet-Cauchy identity.
\addcontentsline{lth}{theorem}{{\bf Proposition \numberline{\theproposition}}
{ : Generalized Binet-Cauchy identity}}
Let $Z_{i}:=(A_{i},$\thinspace $B_{i},\,C_{i},\,D_{i})^{\prime }\;$and\ $
\tilde{Z}_{i}:=Z_{i}-\mathbb{E}\{Z_{i}\}=(\tilde{A}_{i},$\thinspace $\tilde{B
}_{i},\,\tilde{C}_{i},\,\tilde{D}_{i})^{\prime }$, $i=1,2$, be independent
random vectors with finite fourth moments. Then
\begin{eqnarray}
\mathbb{E}\{(A_{1}B_{2}-A_{2}B_{1})(C_{1}D_{2}-C_{2}D_{1})\} &=&\mathbb{E}
\{A_{1}C_{1}\}\mathbb{E}\{B_{2}D_{2}\}+\mathbb{E}\{A_{2}C_{2}\}\mathbb{E}
\{B_{1}D_{1}\} \notag \\
&-&[\mathbb{E}\{A_{1}D_{1}\}\mathbb{E}\{B_{2}C_{2}\}+\mathbb{E}\{A_{2}D_{2}\}
\mathbb{E}\{B_{1}C_{1}\}]\,.\quad \quad
\end{eqnarray}
If $(A_{1},$\thinspace $B_{1},\,C_{1},\,D_{1})^{\prime }$ and $(A_{2},$
\thinspace $B_{2},\,C_{2},\,D_{2})^{\prime }$ have the same covariance
matrix, then
\begin{equation}
\mathsf{C}[A_{1},\,C_{1}]\mathsf{C}[B_{1},\,D_{1}]-\mathsf{C}[A_{1},\,D_{1}]
\mathsf{C}[B_{1},\,C_{1}]=\frac{1}{2}\mathbb{E}\{(\tilde{A}_{1}\tilde{B}_{2}-
\tilde{A}_{2}\tilde{B}_{1})(\tilde{C}_{1}\tilde{D}_{2}-\tilde{C}_{2}\tilde{D}
_{1})\}\,.
\end{equation}
The above proposition yields restrictions on the covariances of a
four-dimensional random vector. From ((ref)
), we see that $A_{1}B_{2}-A_{2}B_{1}=0$ (a.s.) or $C_{1}D_{2}-C_{2}D_{1}=0$
(a.s.) entails
equation[equation omitted — 121 chars of source]
On taking\ $A_{i}=C_{i}$\ and\ $B_{i}=D_{i}\;$for $i=1,$\thinspace $2$, we
get the Lagrange identity:
equation[equation omitted — 198 chars of source]
Applications
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
The $D$-representations given above show that the central moments of a real
random variable $X$ can all be represented as expected values of polynomial
functions in differences between a small number of replications of $X$. For $
k\geq 2,$ the $k$-th central moment of $X$ only requires $k$ replications $
X_{1},\ldots ,\,X_{k}$. More precisely, we can write:
equation[equation omitted — 215 chars of source]
where
equation[equation omitted — 66 chars of source]
equation[equation omitted — 106 chars of source]
equation[equation omitted — 129 chars of source]
align[align omitted — 211 chars of source]
In other words, $h_{k}(X_{1},\ldots ,\,X_{k})$ is an unbiased estimator of $
{\Greekmath 0116} _{k}$ based on $k$ observations.
If $n$ observations $X_{1},\ldots ,\,X_{n}$ are available, any (fixed or
randomly selected) subset of $k$ observations from $\{X_{1},\ldots
,\,X_{n}\} $ yields an unbiased estimator of ${\Greekmath 0116} _{k},$ and similarly for
the arithmetic (or a weighted) average of such subsets. In particular, the
arithmetic average of $h_{k}(X_{i_{1}},\ldots ,\,X_{i_{k}})$ over $N$
randomly selected $\{i_{1},\ldots ,\,i_{k}\}$ of $\{1,\ldots ,\,n\}$ is an
unbiased estimator of ${\Greekmath 0116} _{k}$. A fortiori, this will also hold if
the average is taken over all subsets of $k$ observations from $
\{X_{1},\ldots ,\,X_{n}\}$, but the large number of such subsets can easily
make this approach computationally prohibitive. We will call such estimators
based on observation differences $D$-estimators.
It is of interest to compare $D$-estimators of ${\Greekmath 0116} _{k}$ with the
\textquotedblleft natural\textquotedblright\ estimators based on deviations
from the sample mean:
equation[equation omitted — 135 chars of source]
For $k=2,$ $\hat{m}_{k}(n)$ is unbiased, but it is not generally unbiased
for $k\geq 3$. To see how big the bias difference can be, we take $n=k$ and
compare by simulation $\hat{{\Greekmath 0116}}_{k}(k)$ with the corresponding $D$
-estimator with $h_{k}(X_{1},\ldots ,\,X_{k});$ see Table (ref). We use the
Exponential distribution as an example because its central moments are
nonzero at all orders, which helps us assess estimator bias without
interference from sign-canceling or offsetting effects that occur when the
true moment is zero. We see from these results that $\hat{m}_{k}(k)$ is
heavily biased, while the corresponding $D$-estimators exhibit no bias as
predicted by the results presented above.
table[table omitted — 1,352 chars of source]
Of course, if $n>k$, more efficient unbiased estimators can be obtained by
averaging several minimal estimators. This can be easily implemented through
computer programming by selecting the index tuples $(i_{1},\ldots ,i_{k+1})$
independently and identically distributed (i.i.d.) from a uniform
distribution over the index set $I$. Due to its stochastic nature, we refer
to this estimator as a Monte Carlo approximation of the $D$-estimator, $\hat{
{\Greekmath 0116}}_k(n)^{MC}$. Table (ref) presents simulation results comparing the bias of
the \textquotedblleft natural\textquotedblright\ estimators $\hat{m}_k(n)$
with the proposed $D$-estimator Monte Carlo approximation for central
moments of orders $k=2,\ldots ,8$ under an $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Exponential}(2)$
distribution. For each replication, we generate 30,000 Monte Carlo samples.
table[table omitted — 1,641 chars of source]
The simulation results demonstrate that the proposed $D$-estimator Monte
Carlo approximation substantially reduces bias across all orders of the
central moment compared to the natural estimator. This improvement is
particularly notable for higher-order moments where the naive estimator
exhibits larger negative bias. Moreover, increasing the sample size from 50
to 100 decreases the bias for both estimators, but the $D$-estimator
approach consistently outperforms the natural estimator at each sample size.
These findings highlight the effectiveness of the Monte Carlo approximation
in providing more accurate and unbiased estimates of central moments,
especially for higher-order moments that are typically more challenging to
estimate.
Tentative\begin{comment}
It is of interest to compare $D$-estimators of ${\Greekmath 0116} _{k}$ with the
\textquotedblleft natural\textquotedblright\ estimators based on deviations
from the sample mean:
\begin{equation}
\hat{m}_{k}(n)=\frac{1}{n-1}\sum_{i=i}^{n}(X_{i}-\bar{X}_{n})^{k}\,,\quad
\bar{X}_{n}=\frac{1}{n}\sum_{i=i}^{n}X_{i}\,.
\end{equation}
For $k=2,$ $\hat{m}_{k}(n)$ is unbiased, but it is not generally unbiased
for $k\geq 3$. To see how big the bias difference can be, we take $n=k$ and
compare by simulation $\hat{{\Greekmath 0116}}_{k}(k)$ with the corresponding $D$
-estimator with $h_{k}(X_{1},\ldots ,\,X_{k});$ see Table (ref). We see from
these results that $\hat{m}_{k}(k)$ is heavily biased, especially for
non-zero central moments, while the corresponding $D$-estimators exhibit no
bias as predicted by the results presented above.
\begin{table}[tbph]
\caption{Add caption}
\resizebox{\columnwidth}{!}{ \begin{tabular}{rlrrrrrrr}
\toprule
\multicolumn{9}{c}{ Table 1. Simulation Comparison of Bias for Two Distributions when n=k (\# of Simulation = 20000000) } \\
\midrule
& Central Moment & \multicolumn{1}{l}{Ord. 2} & \multicolumn{1}{l}{Ord. 3} & \multicolumn{1}{l}{Ord. 4} & \multicolumn{1}{l}{Ord. 5} & \multicolumn{1}{l}{\textbf{Ord. 6}} & \multicolumn{1}{l}{\textbf{Ord. 7}} & \multicolumn{1}{l}{\textbf{Ord. 8}} \\
& \textbf{Number of Observations} & \textbf{2} & \textbf{3} & \textbf{4} & \textbf{5} & \textbf{6} & \textbf{7} & \textbf{8} \\
\midrule
\multicolumn{1}{l}{\textbf{Exponential(2)}} & \textit{True Value} & \textit{0.25} & \textit{0.25} & \textit{0.56} & \textit{1.38} & \textit{4.00} & \textit{14.48} & \textit{57.94} \\
\cmidrule{2-9} & \textbf{Natural Estimator $\hat{m}_k(k)$} & \textbf{0.000} & -0.167 & -0.258 & -0.768 & -2.171 & -8.321 & -33.739 \\
&\textbf{D-estimators $\hat{{\Greekmath 0116}}_k(k)$} & \textbf{0.000} & \textbf{0.000} & \textbf{0.000} & \textbf{0.003} & \textbf{0.094} & \textbf{0.131} & \textbf{0.053} \\
\midrule
\multicolumn{1}{l}{\textbf{Normal(0,1)}} & \textit{True Value} & \textit{1} & \textit{0} & \textit{3} & \textit{0} & \textit{15} & \textit{0} & \textit{105} \\
\cmidrule{2-9} &\textbf{Natural Estimator $\hat{m}_k(k)$} & \textbf{0.000} & \textbf{0.000} & -0.748 & \textbf{0.001} & -4.582 & \textbf{0.009} & -34.650 \\
& \textbf{D-estimators $\hat{{\Greekmath 0116}}_k(k)$} & \textbf{0.000} & -0.001 & \textbf{-0.003} & -0.061 & \textbf{0.246} & -0.292 & \textbf{0.493} \\
\bottomrule
\bottomrule
\end{tabular} }
\end{table}
Of course, if $n>k$, more efficient unbiased estimators can be obtained by
averaging several minimal estimators. This can be easily implemented through
computer programming by selecting the index tuples $(i_{1},\ldots ,i_{k+1})$
independently and identically distributed (i.i.d.) from a uniform
distribution over the index set $I$. Due to its stochastic nature, we refer
to this estimator as a Monte Carlo approximation of the $D$-estimator, $\hat{
{\Greekmath 0116}}_k(n)^{MC}$. Table (ref) presents simulation results comparing the bias of
the \textquotedblleft natural\textquotedblright\ estimators $\hat{m}_k(n)$
with the proposed $D$-estimator Monte Carlo approximation for central
moments of orders $k=2,\ldots ,8$ under an $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Exponential}(2)$
distribution and a $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Normal}(0,1)$ distribution. For each replication,
we generate 30,000 Monte Carlo samples.
\begin{table}[htbp]
\caption{Add caption}
\resizebox{\columnwidth}{!}{ \begin{tabular}{rlrrrrrrr}
\toprule
\multicolumn{9}{c}{\textbf{Table 2. Simulation Comparison of Bias for Two Distributions when n>k (\# of Simulation = 1000)}} \\
\midrule
& \textbf{Central Moment} & \multicolumn{1}{l}{\textbf{Ord.2}} & \multicolumn{1}{l}{\textbf{Ord.3}} & \multicolumn{1}{l}{\textbf{Ord.4}} & \multicolumn{1}{l}{\textbf{Ord.5}} & \multicolumn{1}{l}{\textbf{Ord.6}} & \multicolumn{1}{l}{\textbf{Ord.7}} & \multicolumn{1}{l}{\textbf{Ord.8}} \\
\midrule
\multicolumn{1}{l}{\textbf{Exponential (2)}} & \textit{True Value} & \textit{0.25} & \textit{0.25} & \textit{0.56} & \textit{1.38} & \textit{4.00} & \textit{14.48} & \textit{57.94} \\
\cmidrule{1-9} \multicolumn{1}{l}{\textbf{n=50}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & -0.004 & -0.018 & -0.059 & -0.240 & -0.921 & -5.203 & -27.483 \\
& \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & \textbf{-0.004} & \textbf{-0.009} & \textbf{-0.035} & \textbf{-0.162} & \textbf{-0.592} & \textbf{-4.048} & \textbf{-22.484} \\
\cmidrule{2-9} \multicolumn{1}{l}{\textbf{n=100}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & \textbf{-0.001} & -0.008 & -0.020 & -0.084 & -0.295 & -2.621 & -16.574 \\
& \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & -0.001 & \textbf{-0.003} & \textbf{-0.005} & \textbf{-0.035} & \textbf{-0.133} & \textbf{-1.991} & \textbf{-13.020} \\
\midrule
& & & & & & & & \\
\midrule
\multicolumn{1}{l}{\textbf{N(0,1)}} & \textit{True Value} & \textit{1} & \textit{0} & \textit{3} & \textit{0} & \textit{15} & \textit{0} & \textit{105} \\
\cmidrule{1-9} \multicolumn{1}{l}{\textbf{n=50}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & 0.003 & -0.003 & -0.055 & \textbf{0.014} & -0.767 & \textbf{0.403} & -11.345 \\
& \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & \textbf{0.002} & \textbf{-0.003} & \textbf{0.000} & 0.041 & \textbf{-0.287} & 0.516 & \textbf{5.071} \\
\cmidrule{2-9} \multicolumn{1}{l}{\textbf{n=100}} & \textbf{Natural Estimator $\hat{m}_k(n)$} & -0.002 & \textbf{0.002} & -0.036 & \textbf{0.081} & -0.381 & \textbf{1.355} & -5.460 \\
& \textbf{D-Estimator MC $\hat{{\Greekmath 0116}}_k(n)^{MC}$} & \textbf{-0.002} & 0.003 & \textbf{-0.007} & 0.118 & \textbf{-0.228} & 2.115 & \textbf{-3.893} \\
\bottomrule
\bottomrule
\end{tabular} }
\end{table}
The simulation results demonstrate that the proposed $D$-estimator Monte
Carlo approximation substantially reduces bias across all orders of the
central moment compared to the natural estimator in Exponential
distributions. This improvement is particularly notable for higher-order
moments where the natural estimator exhibits larger negative bias. Moreover,
increasing the sample size from 50 to 100 decreases the bias for both
estimators, but the $D$-estimator approach consistently outperforms the
natural estimator at each sample size. For Normal(0,1) distributions where
the odd order central moments are zero, the performance of estimator may be
partially masked by sign calcellation in the simulation data. In cases of
non-zero central moments, the $D$-estimator approach outperforms the natural
estimator. These findings highlight the effectiveness of the Monte Carlo
approximation in providing more accurate and unbiased estimates of central
moments, especially for higher-order moments that are typically more
challenging to estimate.
\begin{Tentative}
\end{comment}
\end{Tentative}
From ((ref)), one can see that several
choices of minimal estimator $h_{3}$ are possible (and similarly for
higher-order moments), which may have different efficiency properties. The
above results can also be used to build unbiased estimators of cumulants as
well as to analyze which observations deviate from the rest of the sample.
The latter application can be based on analyzing the population of subgroups
of observations $(X_{i_{1}},\ldots ,\,X_{i_{k}})$ along with the
corresponding unbiased estimates $h_{k}(X_{i_{1}},\ldots ,\,X_{i_{k}})$.
Discussing in detail variants of $D$-estimators, as well as other
applications, goes beyond the scope of this paper.
Conclusion
{\Alph{section}}
{\Alph{section}}
\setcounter{theorem}{0} \setcounter{definition}{0} \setcounter{equation}{0}
In this paper, following an approach (apparently) initiated by \nocite*
{Gini(1912)} for the variance, we have studied the general problem of
representing the central moments of a random variable in terms of pairwise
differences between i.i.d. replications, without reference to a location
parameter. We first gave pairwise-difference representations (defined
formally as $D$-representations) for the third and fourth central
moments of a general random variable. These yield intuitive interpretations
of the familiar skewness and kurtosis coefficients. For general moments, we
have observed that central moments can be interpreted as covariances between
a random variable and powers of the same variable (a nonlinear
relation). These then provide recursions which link the $D$
-representation of any moment to lower order ones, hence $D$-representations
for all central moments (as long as they are finite). Through a similar
approach, analogues of Lagrange and Binet-Cauchy identities were established
for general random variables. These provide a simple derivation of the
classic Cauchy-Schwarz inequality for covariances. Finally, it is of
interest to note that the formulas given in this paper can be interpreted as
applications of the \textquotedblleft coupling\textquotedblright\ method
[see \nocite*{Thorisson(2000)}], which opens the way to additional
applications.
Acknowledgements
Cette recherche a b\'{e}n\'{e}fici\'{e}
du support financier de la Chaire William Dow (McGill University),
de la Banque du Canada (Bourse de recherche),
d'une Bourse Guggenheim, de la Bourse Konrad-Adenauer
(Fondation Alexander von Humboldt, Allemagne),
du Conseil de recherche en sciences humaines du Canada, du
Conseil de recherche en sciences naturelles et en g\'{e}nie du
Canada, du R\'{e}seau canadien de centres d'excellence (projet MITACS), et du
Fonds de recherche sur la soci\'{e}t\'{e} et la culture
(Qu\'{e}bec).