The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
82,251 characters
A Robust Similarity Estimator
\title{A Robust Similarity Estimator \thanks{The author is grateful for helpful comments from Peter Reinhard Hansen,
Jun Yu, Han Chen, Yijie Fei, and Marine Carrasco. The author also
thanks participants of the NUS Quantitative Finance Conference in
Singapore (July, 2025), the Econometrics Seminar at the University
of Macau (October, 2025) and the Virtual Time Series Seminar (December,
2025) for valuable discussions and feedback. }{\normalsize\emph{\medskip{}
}}}
\author{\textbf{Ilya Archakov}\bigskip{}
\\
{\normalsize\emph{York University; Department of Economics\medskip{}
}}}
\date{{\normalsize\emph{\date{}}}}
\maketitle
\begin{abstract}
We construct and analyze an estimator of association between random
variables based on their similarity in both direction and magnitude.
Under special conditions, the proposed measure becomes a robust and
consistent estimator of the linear correlation, for which an exact
sampling distribution is available. This distribution is intrinsically
insensitive to heavy tails and outliers, thereby facilitating robust
inference for correlations. The measure can be naturally extended
to higher dimensions, where it admits an interpretation as an indicator
of joint similarity among multiple random variables. We investigate
the empirical performance of the proposed measure with financial return
data at both high and low frequencies. Specifically, we apply the
new estimator to construct confidence intervals for correlations based
on intraday returns and to develop a new specification for multivariate
GARCH models.
\bigskip{}
\end{abstract}
{\small\textit{Keywords:}}{\small{} Correlation, Robust Estimation,
Robust Inference, Fisher Transformation, Matrix Logarithm, High Frequency
Data, Multivariate GARCH}{\small\par}
\noindent{\small\textit{JEL Classification:}}{\small{} C13, C30, C38,
C58 \newpage}{\small\par}
\section{Introduction}
Measuring statistical association and dependence between random variables
is a broad topic in statistics with a long history. The most popular
and widely used measure of association is the linear correlation coefficient,
or the Pearson correlation coefficient (denoted by $\rho$), which
represents the covariance between two random variables scaled by the
product of their standard deviations. Pearson's correlation naturally
appears in many statistical and econometric frameworks such as linear
regression, multivariate GARCH models, or network analysis. In financial
econometrics literature, correlations play a critical role in risk
management, hedging, optimal portfolio allocation, analysis of systemic
risk, etc.
Both estimation and inference for correlations are practically challenging,
especially in the settings where the sample size is limited. For example,
under sufficiently mild assumptions, the well-known sample correlation
estimator ($\hat{\rho}$) is a consistent estimator of $\rho$. Although
it is an efficient estimator when observations are independent and
normally distributed, its variance is inflated in the presence of
heavy-tailed data, and the estimator remains notoriously sensitive
to outliers. In addition, the finite sample distribution of $\hat{\rho}$
converges very slowly to the asymptotic limit with both the shape
and spread often strongly depend on the properties of underlying data.
As a result, potential distortions in both the central tendency and
the sampling distribution of $\hat{\rho}$ undermine robust estimation
of the correlation coefficient and complicate statistical inference.
In the series of seminal papers, Ronald A. Fisher proposed a continuous
transformation (now known as the Fisher transformation) for $\hat{\rho}$
that offers several advantages including variance stabilization and
a symmetric, nearly Gaussian sampling distribution for the transformed
sample correlation, even in relatively small samples (see \citet{Fisher_1921},
\citet{Hotelling_1953}). The rapid convergence to the asymptotic
distribution has made the Fisher transformation popular for conducting
statistical inference about $\rho$, even with limited data. The estimation
and inference, however, remain fragile, as the transformation does
not provide robustness to outliers, and the sampling variance remains
inflated under heavy-tailed data distributions.
The lack of robustness of the sample correlation is commonly addressed
through the use of alternative correlation measures that are, by construction,
insensitive to extreme observations. Notable examples include the
Quadrant estimator (\citet{Greiner_1909}, \citet{Kendall_1949},
\citet{Blomqvist_1950}) and the Kendall rank correlation coefficient
(\citet{esscher1924method}, \citet{Kendall_1938}). These measures
are based on the relative number of concordant and discordant pairs
of observations and are therefore intrinsically robust to outliers.
Under mild assumptions, they can be transformed into consistent (though
not efficient) estimators of $\rho$ (see, for example, \citet{Croux2010}).
Despite their robustness and consistency, estimators based on rank
correlation measures typically have sampling distributions that remain
sensitive to the underlying data generating process. This sensitivity
poses a substantial challenge for inference on correlations in applied
empirical analysis, where the true distribution of the data is typically
unknown.
In this paper, we develop a new estimator of statistical association
between two random variables that is consistent for a specific functional
of the covariance matrix of the underlying variables. We refer to
it as the similarity estimator. The new estimator is inspired by works
of \citet{Thorndike_1905} and \citet{Fisher_1919}, and is based
on a measure of similarity between two variables that accounts for
both sign and magnitude. While the estimator retains a meaningful
interpretation on its own and can be used as an alternative measure
to gauge statistical relationship between random variables, under
the special conditions of elliptically distributed data with homogeneous
variances it becomes a consistent estimator of the Pearson correlation
$\rho$, on the Fisher scale, and exhibits excellent finite sample
properties. In particular, it admits a \textit{robust} sampling distribution
that is invariant over the class of elliptically distributed data
with arbitrary kurtosis parameters. The sampling distribution is available
via the known characteristic function. This property enables not only
robust estimation but also reliable inference for correlations, even
in small samples. For example, it allows to conduct robust interval
estimation for a given coverage probability. Naturally, the efficiency
of the similarity estimator is lower than that of, for example, the
sample correlation estimator, reflecting the price paid for its intrinsic
robustness. This results in more conservative inference and, on average,
wider confidence intervals. However, these intervals remain robust
to both outliers and extremely heavy-tailed data distributions.
We propose a natural generalization of the similarity estimator to
the multivariate setting involving an arbitrary number of variables.
The resulting generalized similarity measure captures the relative
variation of the observed data along a similarity direction, defined
by the vector of ones, and can be interpreted as a measure of joint
similarity across multiple variables. This estimator naturally inherits
robustness against outliers and nests the bivariate similarity estimator
as a special case. For elliptically distributed data with homogeneous
variances and identical pairwise correlations, the multivariate similarity
estimator becomes a consistent estimator of the (transformed) equicorrelation
parameter, thus permitting a correlation-based interpretation analogous
to the bivariate case.
Estimating financial correlations from high-frequency data represents
a particularly promising area for applications. While intraday data
provide a rich set of observations, offering the potential for efficient
estimation and accurate inference, the analysis of such data is accompanied
by multiple econometric challenges. For example, large instantaneous
price movements, or jumps, which are prevalent in financial markets,
generate extreme (outlying) return observations that often induce
a downward bias in estimated correlations. In addition, at sufficiently
high frequencies, the estimation of realized covariances and correlations
may be affected by various forms of market microstructure noise (see
\citet{Hansen_Lunde_2006}, \citet{BandiRussellRES08}), data asynchronicity
and the Epps effect (see \citet{Ren`o2003}), as well as other adverse
artifacts.
The literature on covariance and correlation estimation using high-frequency
data proposes a wide range of estimators with varying degrees of robustness
to the aforementioned challenges. In practice, realized correlations
are most commonly obtained from multivariate volatility estimators
by rescaling estimated covariances to the corresponding Pearson correlations
(see \citet{BNS:2004}, \citet{AitSahaliaFanXiu2010}, \citet{BNHLS:2011},
\citet{HansenHorelLundeArchakov}, among others). While many of these
methods are explicitly designed to handle various microstructural
effects, they nevertheless remain vulnerable to price jumps and other
outliers. A common remedy is the application of truncation or thresholding
techniques to filter out extreme returns (see, for example, \citet{Mancini2001},
\citet{AndersenDobrevSchaumburg2012}). An alternative approach is
to employ robust correlation measures, such as the Kendall and Quadrant
correlation estimators, along with their modifications and adaptations
for high-frequency data (see \citet{VanderElst_Veredas_2015}, \citet{Hansen_Luo_2023},
among others).
The intrinsic robustness of the similarity estimator to extreme observations
and heavy-tailed data makes it a promising tool for correlation estimation
in high-frequency data settings. A key feature that distinguishes
the similarity estimator from existing alternatives is its ability
to deliver robust interval estimates for correlations, enabled by
the availability of a robust sampling distribution. We provide an
empirical illustration by constructing robust confidence intervals
for daily correlations between several selected stocks from the U.S.
stock market during the COVID-19 outbreak. In our empirical application,
we estimate correlations at moderately low frequencies (ranging from
1 to 10 minutes), which helps mitigate microstructure effects while
still providing a sufficient number of intraday observations for accurate
estimation. Our results are notable in several respects. First, despite
the intervals are inherently conservative, pronounced correlation
dynamics are clearly identified over the sample period. Second, the
estimated intervals exhibit visible clustering over sequences of several
consecutive trading days, indicating estimation stability and suggesting
persistent dynamics in the underlying correlation process. Finally,
the constructed intervals show strong agreement with robust correlation
estimates based on the Kendall coefficient, supporting the reliable
performance of the similarity estimator in high-frequency data applications.
We note that, in our empirical illustration, the similarity estimator
is applied while ignoring several features of intraday data that are
inconsistent with the assumptions underlying its theoretical properties.
Accordingly, we leave a careful adaptation of the proposed estimator
to high-frequency settings for future research.
Another area for applications is the modeling of correlation dynamics
for vectors of asset returns. The traditional approach builds on extensive
literature on multivariate GARCH or stochastic volatility models,
in which the conditional correlation process is assumed, either directly
or indirectly, to evolve over time (see the respective surveys in
\citet{Bauwens_Laurent_Rombouts_2006} and \citet{Asai_McAleer_Yu_2006}).
The models are typically applied to relatively low-frequency data,
such as daily or weekly asset returns. Key challenges of this class
of models include preserving positive definiteness of the conditional
correlation structures, as well as the rapidly increasing computational
complexity as the number of assets grows.
We propose a new class of multivariate GARCH models that employ the
similarity measure to model correlation dynamics. Our approach follows
the Dynamic Conditional Correlation (DCC) framework introduced in
\citet{Engle_2002}, in which the conditional correlation process
is modeled in isolation from the conditional volatilities. In the
bivariate version of our model, correlation dynamics is specified
directly in the Fisher scale, ensuring positive definiteness of the
correlation matrix by construction. The similarity measure, applied
to past (standardized) returns, serves as a robust observation-driven
update for the Fisher-transformed conditional correlation process.
We further develop a parsimonious model specification in which an
arbitrary number of assets can be accommodated without substantially
increasing computational complexity. To this end, we impose the equicorrelation
assumption, similar to \citet{Engle_Kelly_2012}, and specify dynamics
for the conditional equicorrelation parameter under the matrix logarithmic
transformation. This formulation eliminates the need for additional
constraints to guarantee positive definiteness, in contrast to the
classical DCC framework. The multivariate similarity measure, computed
on standardized returns, is then used as an observation-driven update
for the dynamic equicorrelation parameter. An important feature of
the proposed multivariate GARCH model is the intrinsic robustness
of the estimated correlation process to extreme observations and fat-tailed
distributions, which are common in financial returns.
We apply both proposed specifications to daily returns from the U.S.
stock market over a 16-year sample period. The resulting trajectory
of the estimated correlation process appears to be broadly in line
with both the realized sample correlations and the conditional correlations
obtained from standard DCC--GARCH models, however, it exhibits noticeable
local deviations. We attribute these differences to the improved robustness
of the proposed approach. Although the estimation results indicate
that the new model delivers a stable and tractable estimation procedure
with sensible outcomes, a more thorough assessment of its statistical
and economic performance is left for future research.
\section{\label{sec:Measure-of-Similarity}A Measure of Similarity for Random
Variables}
Let $x_{1}$ and $x_{2}$ are two real-valued random variables. Assume
additionally that $x_{1}$ and $x_{2}$ have zero means, $E(x_{1})=E(x_{2})=0$,
and the norm of random vector $x=(x_{1},x_{2})^{\prime}$ is positive
with probability one. We consider the following variable,
\begin{equation}
r=\frac{2x_{1}x_{2}}{x_{1}^{2}+x_{2}^{2}},\label{eq:res}
\end{equation}
which can be interpreted as a measure of statistical similarity between
$x_{1}$ and $x_{2}$. For empirical analysis, this quantity was firstly
introduced in \citet{Thorndike_1905} to measure statistical association
between a pair of twins with respect to a set of considered characteristics,
and was originally named \textit{resemblance}. In that study, this
term was used as a synonym for the \textit{coefficient of correlation}.
Indeed, multiple aspects allow to consider the Thorndike's resemblance
variable, $r$, as a measure of correlation. The quantity is confined
within the fixed interval, $r\in[-1,1]$, and the value of $r$ is
higher when $x_{1}$ and $x_{2}$ are more similar in magnitude, while
sharing the same sign. The magnitude of $r$ is larger when magnitudes
of $x_{1}$ and $x_{2}$ are more similar to each other, while the
sign of $r$ is positive (negative) if $x_{1}$ and $x_{2}$ have
the same (opposite) directions.
Variable $r$ has a particularly simple form when the random vector
is represented in polar coordinates ($x_{1}=s\cos\theta$ and $x_{2}=s\sin\theta$).
In this case, $r=\sin2\theta$, so the resemblance measure does not
depend on the vector length, $s$. Intuitively, it implies that $r$
depends only on the angle between the observed vector and coordinate
axes and ignores the magnitude of observations. This points on intrinsic
insensitivity of the resemblance measure, $r$, to the presence of
outliers in the data.
If we additionally assume that vector $x=(x_{1},x_{2})^{\prime}$
has finite second moments, we can define the linear correlation coefficient,
or the Pearson correlation, between $x_{1}$ and $x_{2}$ as
\begin{equation}
\rho=\frac{\sigma_{12}}{\sqrt{\sigma_{1}^{2}\sigma_{2}^{2}}},\label{eq: Pearson_corr}
\end{equation}
where $\sigma_{1}^{2}=V(x_{1})=E(x_{1}^{2})$, $\sigma_{2}^{2}=V(x_{2})=E(x_{2})^{2}$
and $\sigma_{12}=E(x_{1}x_{2})$ are the second moments of $x$. The
Pearson correlation is the most popular measure of statistical association
that arises naturally in multiple econometric frameworks, from a simple
linear regression to complex statistical learning algorithms. The
Pearson correlation is often considered as the default benchmark for
measuring dependence in empirical analysis. Although $\rho$ is able
to capture the precise association between random variables only when
the true relationship is linear, it can still provide a reasonable
approximation when the underlying dependence is monotonic and not
heavily non-linear.
While the correlation coefficient $\rho$ is bounded between $-1$
and $1$ by construction, it is sometimes convenient to work with
an unconstrained correlation measure and this can be achieved by means
of a suitable transformation. The most prominent example is the Fisher
transformation defined as,
\[
\phi_{\rho}=\frac{1}{2}\log\Bigl(\frac{1+\rho}{1-\rho}\Bigl),
\]
for $\rho\ensuremath{\in(-1,1)}$. The Fisher transformation represents
a strictly monotone transformation of $\rho$ onto the set of real
numbers, such that $\phi_{\rho}\in\mathbb{R}$. The transformation
was proposed by Ronald A. Fisher in a series of seminal papers (see
\citet{Fisher_1915}, \citet{Fisher_1921}), where he also demonstrated
that it improves the distributional properties of the sample correlation
coefficient.
\begin{figure}[h]
\begin{centering}
\includegraphics[scale=0.6]{plots/contour_gamma_ib}
\par\end{centering}
\caption{\footnotesize\label{fig:contour_gamma} A heatmap plot for $\phi_{r}$
as a function of $x_{1}$ and $x_{2}$. Areas with the red color indicate
higher values of $\gamma$, while areas with the blue color indicate
lower (negative) values of $\phi_{r}$. The white lines corresponding
to $x_{1}=x_{2}$ ($r=1$) and $x_{1}=-x_{2}$ ($r=-1$) represent
the loci where the function is undefined.}
\end{figure}
In what follows, we will refer to the Thorndike's resemblance measure,
$r$, under the Fisher transformation,
\begin{equation}
\phi_{r}=\frac{1}{2}\log\Bigl(\frac{1+r}{1-r}\Bigl)=\frac{1}{2}\log\frac{(x_{1}+x_{2})^{2}}{(x_{1}-x_{2})^{2}},\label{eq:phi_r}
\end{equation}
as to the \textit{measure of similarity}, or the similarity variable.
In contrast to $r$, the similarity measure $\phi_{r}$ has an unrestricted
range, $\phi_{r}\in\mathbb{R}$, and its magnitude increases as the
values of $x_{1}$ and $x_{2}$ become more similar. The sign of $\phi_{r}$
is positive (negative) when the two observations have the same (opposite)
signs. Figure \ref{fig:contour_gamma} illustrates the magnitudes
of $\phi_{r}$ as a function of $x_{1}$ and $x_{2}$, where warmer
(red) colors indicate higher values of $\phi_{r}$ and cooler (blue)
colors indicate lower values. We also note that, similar to $r$,
$\phi_{r}$ depends only on the angular coordinate of the vector $x$,
which underlies its intrinsic robustness to observations with extreme
magnitudes.
An important special case arises when the random variables $x_{1}$
and $x_{2}$ follow a bivariate elliptical distribution, for which
the dependence structure is inherently linear. In this situation,
the Pearson correlation completely characterizes the statistical association
between the variables. Assume additionally that variances of $x_{1}$
and $x_{2}$ are homogeneous, such that $\sigma_{1}^{2}=\sigma_{2}^{2}=\sigma^{2}$.
Then the covariance matrix $\Sigma=V(x)$ reads
\begin{equation}
\Sigma=\sigma^{2}\left(\begin{array}{cc}
1 & \rho\\
\rho & 1
\end{array}\right).\label{eq:Sigma_homo}
\end{equation}
Under this assumption, $\phi_{r}$ and $\phi_{\rho}$ are elegantly
connected. This result was originally formulated in \citet{Fisher_1919}
for the Gaussian case, and below we provide an extension of the original
result to the entire class of elliptical distributions.
\begin{prop}
\label{prop:P1} Assume that $x=(x_{1},x_{2})^{\prime}$ is a bivariate
random vector which follows some elliptical distribution with zero
mean and positive-definite covariance matrix $\Sigma$ with homogeneous
variances, and let $\rho$ denotes the Pearson correlation between
$x_{1}$ and $x_{2}$. Denote the measure of resemblance by $r=\frac{2x_{1}x_{2}}{x_{1}^{2}+x_{2}^{2}}$,
and the corresponding Fisher transformation by $\phi_{r}=\frac{1}{2}\log\Bigl(\frac{1+r}{1-r}\Bigl)$.
Then, $\phi_{r}$ is a random variable with the probability density
function
\[
f(\phi_{r})=\frac{1}{\pi}\text{sech}(\phi_{r}-\phi_{\rho}),
\]
where $\phi_{\rho}=\frac{1}{2}\log\Bigl(\frac{1+\rho}{1-\rho}\Bigl)$
is the Fisher transformation of $\rho$.
\end{prop}
Proposition \ref{prop:P1} implies that $\phi_{r}$ is symmetrically
distributed, according to the hyperbolic secant distribution, around
the Fisher transformation of $\rho$. The distribution does not depend
on the variance parameter, $\sigma^{2}$, as well as on the underlying
correlation coefficient, $\rho$, and remains invariant once $x=(x_{1},x_{2})^{\prime}$
belongs to the elliptical family. Therefore, under the assumptions
of the proposition, the similarity variable $\phi_{r}$ provides an
unbiased and robust signal of the latent correlation level (on the
Fisher scale) with stable sampling properties, and thus emerges as
an attractive statistical tool for correlation estimation.
\section{\label{sec:Realized-Similarity-Estimator}The Similarity Estimator }
Let a random sample is given, $\{x_{t}\}_{t=1}^{T}$, where $x_{t}=(x_{1,t},x_{2,t})^{\prime}$
are independent observations from some bivariate distribution with
zero mean and finite second moments. Probably the most popular estimator
of the Pearson correlation is the sample correlation estimator which
is given by
\[
\hat{\rho}=\frac{\sum_{t=1}^{T}x_{1,t}x_{2,t}}{\sqrt{\Bigl(\sum_{t=1}^{T}x_{1,t}^{2}\Bigl)\Bigl(\sum_{t=1}^{T}x_{2,t}^{2}\Bigl)}},
\]
where we internalize that $\mathbb{E}x=0$. The sample correlation,
$\hat{\rho}$, is a consistent estimator of the population Pearson
correlation, and, under mild assumptions, it is asymptotically normal
with $\sqrt{T}(\hat{\rho}-\rho)\overset{d}{\rightarrow}N\Bigl(0,V_{\rho}\Bigl)$.
For relatively small $T$, however, the sampling properties of $\hat{\rho}$
are often poorly approximated by the asymptotic results, especially
in the presence of sufficiently heavy-tailed data. The inference is
additionally complicated by the fact that the asymptotic variance
$V_{\rho}$ generally depends on the unknown value of $\rho$. For
example, for the Gaussian case, $V_{\rho}=(1-\rho^{2})^{2}$ (see
\citet{Fisher_1915}).
The Fisher transformation is particularly useful in improving sampling
properties of $\hat{\rho}$. We denote the Fisher transformation of
$\hat{\rho}$ by $\phi_{\hat{\rho}}$. Thus, for normally distributed
observations, the asymptotic distribution of $\phi_{\hat{\rho}}$
is $\sqrt{T}(\phi_{\hat{\rho}}-\phi_{\rho})\overset{d}{\rightarrow}N\Bigl(0,1\Bigl)$,
and the asymptotic variance is independent of $\rho$. More importantly,
the Fisher transformation offers multiple advantages in finite samples.
In particular, when the data is close to be normally distributed,
it provides the variance stabilization for the sampling distribution
of $\phi_{\hat{\rho}}$ with making it symmetric and nearly Gaussian
even for very small samples (see \citet{Fisher_1921}, \citet{Hotelling_1953},
etc.). These properties often motivate to analyze the sample correlation
coefficient in the Fisher scale when conducting inference for correlations.
The critical drawback of the sample correlation estimator, that may
compromise both estimation and inference, is its notorious sensitivity
to outliers and, more generally, to extreme observations. A popular
example of a robust correlation measure is the Kendall rank correlation
coefficient, or Kendall\textquoteright s tau coefficient (\citet{esscher1924method},
\citet{Kendall_1938}), given by
\[
\hat{\tau}=\frac{2}{T(T-1)}\sum_{i<j}\text{sign}(x_{1,i}-x_{2,j})\text{sign}(x_{2,i}-x_{2,j}).
\]
The measure captures the strength of monotonic dependence between
two variables by quantifying the relative frequency of sign-concordant/discordant
pairs of observations. For elliptical distributions, it can be transformed
to match the Pearson correlation via the Greiner's equality which
links the correlation coefficient with the quadrant probabilities,
$\rho=\sin\Bigl(\frac{\pi}{2}E\hat{\tau}\Bigl)$, see \citet{Greiner_1909}.
A similar alternative is the class of quadrant estimators of correlation
which are based on the sample proportion of sign-concordant observations
(see \citet{Sheppard_1899}, \citet{Kendall_1949}, \citet{Blomqvist_1950},
\citet{Hansen_Luo_2023}, etc.)
The results in Proposition \ref{prop:P1} motivate an alternative
estimator of statistical association that is based on the Fisher transformed
similarity, $\phi_{r}$. For each observation $x_{t}$, $t=1,...,T$,
we can construct the corresponding (local) empirical measure of similarity
$\phi_{r,t}$, as given in \eqref{eq:phi_r}. In case $x_{t}$ are
drawn from an elliptical distribution, and if the marginal variances
of $x_{t}$ are homogeneous, it follows from Proposition \ref{prop:P1}
that $\mathbb{E}\phi_{r,t}=\phi_{\rho}$. Therefore, $\phi_{r,t}$
is an unbiased and robust signal of the correlation coefficient, on
the Fisher scale, and this naturally motivates suggesting the following
moment-based estimator,
\begin{equation}
\hat{\gamma}=\frac{1}{T}\sum_{t=1}^{T}\phi_{r,t}=\frac{1}{2T}\sum_{t=1}^{T}\log\frac{(x_{1,t}+x_{2,t})^{2}}{(x_{1,t}-x_{2,t})^{2}}.\label{eq:realized_sim}
\end{equation}
In what follows, we will refer to $\hat{\gamma}$ as to the similarity
estimator. In the elliptical case with homogeneous variances, the
similarity estimator is i) a consistent and unbiased estimator of
the Pearson correlation coefficient (in the Fisher scale), ii) has
robust mean and robust sampling distribution that does not depend
on $\rho$, iii) the sampling distribution is available via the characteristic
function, for any $T$, allowing for exact inference on correlations.
In a more general scenario, when variances of $x_{1}$ and $x_{2}$
may differ, the estimator $\hat{\gamma}$ can be treated as a standalone,
robust measure of statistical association between the two random variables,
and it retains an additional interpretation as a lower bound on the
correlation coefficient $\rho$.
\subsection{\label{subsec:Variance-Homogeneity:-Robust}Robust Estimation and
Inference for Correlations under Variance Homogeneity}
We assume that a random sample of vectors $x_{t}=(x_{1,t},x_{2,t})^{\prime}$,
for $t=1,...,T$, is independently drawn from some elliptical distribution
with finite second moments, identical marginal variances, $\sigma_{1}=\sigma_{2}$,
and the Pearson correlation coefficient, $\rho\in(-1,1)$. Under these
assumptions, it directly follows from Proposition \ref{prop:P1} that
\begin{equation}
\sqrt{T}(\hat{\gamma}-\phi_{\rho})\overset{d}{\rightarrow}N\Bigl(0,\frac{\pi^{2}}{4}\Bigl),\label{eq:gamma_clt}
\end{equation}
because the distribution of $f(\phi_{r,t})$ implies that the variance
of $\phi_{r,t}$ is given by $V(\phi_{r,t})=\frac{\pi^{2}}{4}$.
The asymptotic distribution of $\hat{\gamma}$ does not depend on
the actual correlation coefficient $\rho$. In contrast to $\phi_{\hat{\rho}}$,
which represents the transformation of the sample correlation estimator
$\hat{\rho}$, the similarity estimator, $\hat{\gamma}$, directly
targets $\phi_{\rho}$ by estimating the correlation coefficient under
the Fisher scale. The asymptotic variance of $\hat{\gamma}$ is $\frac{\pi^{2}}{4}$,
and this value is, in general, larger than the asymptotic variance
of $\phi_{\hat{\rho}}$ in the unconstrained scenario. The efficiency
reduction is not surprising since $\hat{\gamma}$ incorporates only
information about the relative magnitudes of $(x_{1,t},x_{2,t})$
and their signs, but ignores information about the total magnitude
of $x_{t}$. A lower efficiency of $\hat{\gamma}$ is nonetheless
compensated by its robustness to outliers and stability of the sampling
distribution.
\begin{figure}[h]
\begin{centering}
\includegraphics[scale=0.8]{plots/simi_density_pdf}
\par\end{centering}
\caption{\footnotesize\label{fig:pdf1} The probability density functions
of $z_{\hat{\gamma}}=\frac{\sqrt{T}(\hat{\gamma}-\phi_{\rho})}{\pi/2}$
for $T=1,3,10$ (colored lines) and the standard normal probability
density (dashed line) which is the distribution limit of $z_{\hat{\gamma}}$
for $T\rightarrow\infty$.}
\end{figure}
A remarkable feature of the similarity estimator in the considered
scenario is that not only the asymptotic distribution, but also the
finite sample distribution of $\hat{\gamma}-\phi_{\rho}$ is invariant
over the entire class of elliptical distributions of $x_{t}$. Furthermore,
the characteristic function of $\phi_{r,t}$ is available and allows
to recover the exact sampling distribution of the similarity estimator
for any finite $T$. Specifically, the distribution $f(\phi_{r,t})$
provided in Proposition \ref{prop:P1} implies that the characteristic
function for $\phi_{r}-\phi_{\rho}$ is given by $\varphi_{\phi_{r}}(u)=\text{sech}\bigl(\frac{\pi}{2}u\bigl)$
for $|u|<1$. Denote the standardized similarity estimator by $z_{\hat{\gamma}}=\frac{2\sqrt{T}}{\pi}(\hat{\gamma}-\phi_{\rho})$.
Then, the characteristic function of $z_{\hat{\gamma}}$ can be written
as
\begin{equation}
\varphi_{z}(u)=\Pi_{t=1}^{T}\varphi_{\phi_{r,t}}\Bigl(\frac{2u}{\pi\sqrt{T}}\Bigl)=\Biggl[\text{sech}\Bigl(\frac{u}{\sqrt{T}}\Bigl)\Biggl]^{T},\label{eq:cf}
\end{equation}
for $|u|<1$. Therefore, the sampling distribution of $\hat{\gamma}$
becomes available in semi-explicit form (via the characteristic function)
for any finite $T$, and this allows to conduct exact inference for
the estimated correlation parameter. The sampling distribution of
$z_{\hat{\gamma}}$ is symmetric with a positive excess kurtosis for
any sample size. It converges to the standard normal distribution
very quickly as $T$ increases. This is illustrated in Figure \ref{fig:pdf1}
and in Table \ref{tab:quantiles_table}, where the quantiles of the
finite sample distribution are reported for a range of $T$.
It is worth mentioning that the properties of $\hat{\gamma}$ are
generally retained for zero-centered elliptical distributions with
infinite moments. For such distributions, the correlation coefficient
is not possible to define using \eqref{eq: Pearson_corr}, however,
it can be defined alternatively via the quadrant probabilities (see
\citet{Sheppard_1899}, \citet{Greiner_1909}, etc.) as $\rho=\sin\Bigl(\pi P-\frac{\pi}{2}\Bigl)=-\cos(\pi P)$,
where $P$ denotes the probability of both $x_{1,t}$ and $x_{2,t}$
are of the same sign. Note that, for an elliptical random vector with
finite second moments, the correlation defined in this way is equivalent
to the Pearson correlation coefficient given by \eqref{eq: Pearson_corr}.
Thus, the similarity estimator remains a robust and consistent estimator
of $\phi_{\rho}$ even for extremely heavy-tailed elliptical distributions,
such as multivariate Cauchy distribution.
\begin{figure}[h]
\begin{centering}
\includegraphics[scale=0.8]{plots/sum_study_density_plots_v2pres2b}
\par\end{centering}
\caption{\footnotesize\label{fig:finite_sample_densities} Finite sample distributions
of $\phi_{\hat{\rho}}$ (red dashed lines) and $\hat{\gamma}$ (blue
solid lines) obtained on 10,000 simulated samples of sizes $T=8$
(top plots) and $T=40$ (bottom plots). Data vectors $x_{t}$ were
simulated out of the three selected distributions -- normal, $t$-distribution
with 5 degrees of freedom, and Cauchy -- with the true correlation
parameter $\rho=0.5$.}
\end{figure}
Figure \ref{fig:finite_sample_densities} provides small sample distributions
of $\hat{\gamma}$ and $\phi_{\hat{\rho}}$ resulted from the simulation
analysis. In this illustration we consider two selected sample sizes
($T=8$ and $T=40$) and three selected bivariate distributions of
vector $x_{t}$: normal, $t$-distribution with 5 degrees of freedom,
and Cauchy, for which the correlation coefficient $\rho$ is defined
via quadrant probabilities. The figure shows an apparent instability
of the sampling density of $\phi_{\hat{\rho}}$ across the considered
elliptical distributions with different kurtosis parameters, thus,
highlighting the sensitivity of the sample correlation estimator $\hat{\rho}$
to the presence of extreme observations. In contrast, the sampling
density of $\hat{\gamma}$ is in line with the theoretically predicted
results and is stable across all distribution specifications for both
considered sample sizes.
The similarity estimator $\hat{\gamma}$ can be used as a robust and
consistent estimator of the Pearson correlation coefficient once the
inverse Fisher transformation is applied, $\phi^{-1}(\hat{\gamma})$,
where
\[
\phi^{-1}(\gamma)=\frac{e^{2\gamma}-1}{e^{2\gamma}+1}=\tanh(\gamma),
\]
which is a hyperbolic tangent function. We note that despite $\hat{\gamma}$
is an unbiased estimator of $\phi_{\rho}$, the inverse transformation,
$\phi^{-1}(\hat{\gamma})$, is not an unbiased estimator of $\rho$
because $\mathbb{E}\phi^{-1}(\hat{\gamma})\neq\rho$, in general,
for $T>1$. However, with an exact finite sample distribution of $\hat{\gamma}$
and due to monotonicity of the Fisher transformation, $\hat{\gamma}$
can be used for \textit{interval estimation} of $\rho$ by providing
exact confidence intervals for any sample size $T$.
Under the variance homogeneity assumption, another interesting feature
of the similarity estimator is related to the matrix logarithm transformation.
Assuming $x_{t}$ has a positive definite covariance matrix $\Sigma$
as in \eqref{eq:Sigma_homo}, $\lambda_{+}=\sigma^{2}(1+\rho)$ and
$\lambda_{-}=\sigma^{2}(1-\rho)$ are the two eigenvalues of $\Sigma$
with the corresponding eigenvectors $q_{+}=\frac{1}{\sqrt{2}}\iota_{2}$
and $q_{-}=\frac{1}{\sqrt{2}}\iota_{2}^{\perp}$, where $\iota_{2}=(1,1)^{\prime}$
and $\iota_{2}^{\perp}=(1,-1)^{\prime}$. For $n=2$, the matrix logarithm
transformation of $\Sigma$ has an explicit analytic expression and
reads
\[
\log\Sigma=\left(\begin{array}{cc}
\frac{1}{2}\log(\lambda_{+}\cdot\lambda_{-}) & \mathinner{\color{purple}\frac{1}{2}}{\color{purple}\log\Bigl(\frac{\lambda_{+}}{\lambda_{-}}}\mathopen{\color{purple}\Bigl)}\\
\mathinner{\color{purple}\frac{1}{2}}{\color{purple}\log\Bigl(\frac{\lambda_{+}}{\lambda_{-}}}\mathopen{\color{purple}\Bigl)} & \frac{1}{2}\log(\lambda_{+}\cdot\lambda_{-})
\end{array}\right),
\]
where the off-diagonal entry, representing the (half) log-condition
number of $\Sigma$, coincides with the Fisher transformation of the
correlation coefficient, $\phi_{\rho}=\frac{1}{2}\log\Bigl(\frac{\lambda_{+}}{\lambda_{-}}\Bigl)$.
Therefore, under the considered assumptions, $\hat{\gamma}=\frac{1}{2T}\sum_{t=1}^{T}\log\frac{(q_{+}^{\prime}x_{t})^{2}}{(q_{-}^{\prime}x_{t})^{2}}$
consistently estimates the off-diagonal element of $\log\Sigma$.
This suggests a promising direction for using $\hat{\gamma}$ to estimate
correlation matrices directly under the matrix logarithm transformation
which offers many convenient properties for correlation analysis (see
\citet{Archakov_Hansen_2021}). We further explore this idea in Section
\ref{subsec:Equicorrelation-Scenario}.
\subsection{\label{subsec:Realized-Similarity-Estimator}Similarity Estimator
under Variance Heterogeneity}
In a more general scenario, where the variables are not necessarily
assumed to have identical variances, the interpretation of the similarity
estimator is different. Let the covariance matrix of a random vector
$x=(x_{1},x_{2})^{\prime}$ is positive definite and is given by
\begin{equation}
\Sigma=\left(\begin{array}{cc}
\sigma_{1}^{2} & \sigma_{12}\\
\sigma_{12} & \sigma_{2}^{2}
\end{array}\right),\label{eq:Sigma_2x2_hetero}
\end{equation}
where $\sigma_{1}^{2}$ and $\sigma_{2}^{2}$ are not necessarily
identical. Consider quantity $\xi=\frac{2\sigma_{12}}{\sigma_{1}^{2}+\sigma_{2}^{2}}$,
which we refer to as the \textit{coefficient of resemblance}. This
coefficient is closely related to the Pearson correlation coefficient,
$\rho$, and can be interpreted as an alternative measure of statistical
association between random variables. More particularly, $\rho$ and
$\xi$ are proportionally related,
\begin{equation}
\xi=\frac{2\sigma_{1}\sigma_{2}}{\sigma_{1}^{2}+\sigma_{2}^{2}}\rho,\label{eq:cor}
\end{equation}
and the coefficient of proportionality depends only on the relative
size of variances, $\sigma_{1}^{2}/\sigma_{2}^{2}$.
Note that the signs of $\xi$ and $\rho$ are always identical, while
the magnitude of $\xi$ never exceeds the magnitude of $\rho$, i.e.
$|\xi|\leq|\rho|$, because $\frac{2\sigma_{1}\sigma_{2}}{\sigma_{1}^{2}+\sigma_{2}^{2}}\leq1$
due to the Cauchy-Schwartz inequality. It then follows that $\xi$
is also constrained between $-1$ and $1$, and attains the limits
in case of perfect correlation ($\rho=\pm1$) and variance homogeneity
($\sigma_{1}^{2}=\sigma_{2}^{2}$). In other words, the coefficient
of resemblance, $\xi$, can be interpreted as a measure of statistical
similarity between the random variables, where both the correlation
and scale are taken into consideration. Namely, $\xi$ is higher when
the variables tend to exhibit more similar correlation components,
along with more similar variance components (magnitudes).
If $x$ is an elliptical random vector, an important feature of $\xi$
is that, under the Fisher transformation, it becomes the mean of the
transformed resemblance measure, $\phi_{r}=\frac{1}{2}\log\Bigl(\frac{1+r}{1-r}\Bigl)$,
where $r$ introduced in \eqref{eq:res}. This property is reflected
in the following proposition.
\begin{prop}
\label{prop:P2} Assume that $x=(x_{1},x_{2})^{\prime}$ is a bivariate
random vector which follows some elliptical distribution with zero
mean and positive-definite covariance matrix $\Sigma$. Denote the
measure of resemblance by $r=\frac{2x_{1}x_{2}}{x_{1}^{2}+x_{2}^{2}}$,
and the corresponding Fisher transformation by $\phi_{r}=\frac{1}{2}\log\Bigl(\frac{1+r}{1-r}\Bigl)$.
Then, $E(\phi_{r})=\phi_{\xi}=\frac{1}{2}\log\Bigl(\frac{1+\xi}{1-\xi}\Bigl)$,
where $\xi=\frac{2\sigma_{12}}{\sigma_{1}^{2}+\sigma_{2}^{2}}$ is
the coefficient of resemblance, and the variance of $\phi_{r}$ is
given by
\[
V_{\phi_{r}}=\frac{\pi^{2}}{6}-\sum_{k=1}^{\infty}\frac{\cos2k\vartheta}{k^{2}},
\]
where $\vartheta$ is a function of elements in $\Sigma$ (the exact
expression is provided in the proof).
\end{prop}
It is important to mention that $\phi_{r}$, as well as its distribution
and moments, retain robustness to outliers due to intrinsic insensitivity
of $r$ to the total magnitude of $x$. In contrast to the homoskedastic
case ($\sigma_{1}^{2}=\sigma_{2}^{2}$), the variance of $\phi_{r}$
does depend on the underlying covariance matrix of $x$. However,
for any positive definite $\Sigma$, we have that $V_{\phi_{r}}\leq\frac{\pi^{2}}{4}$,
with the equality holds only for $\sigma_{1}^{2}=\sigma_{2}^{2}$.
Therefore, $V_{\phi_{r}}$ reaches its maximum value under the variance
homogeneity, and has a lower value otherwise.
The results in Proposition \ref{prop:P2} allow to generalize the
asymptotic properties of the similarity estimator $\hat{\gamma}$
for the case of potentially heterogeneous variances. We assume a sample
of random vectors $x_{t}=(x_{1,t},x_{2,t})^{\prime}$, for $t=1,...,T$,
are independent and elliptically distributed with a non-singular covariance
matrix $\Sigma$. Then $\hat{\gamma}$ has the limit distribution
\begin{equation}
\sqrt{T}(\hat{\gamma}-\phi_{\xi})\overset{d}{\rightarrow}N\Bigl(0,V_{\phi_{r}}\Bigl),\label{eq:gamma_clt_2}
\end{equation}
where $\phi_{\xi}$ is the Fisher transformation of $\xi$ and $V_{\phi_{r}}$
is provided in Proposition \ref{prop:P2}. As a result, $\hat{\gamma}$
is a consistent and robust estimator of the resemblance coefficient
(on the Fisher scale), and its sampling distribution remains stable
across the wide class of elliptical densities for the sample observations.
Since the resemblance coefficient $\xi$ represents a downward scaled
version of the Pearson correlation $\rho$, with the scaling coefficient
is provided in \eqref{eq:cor}, $\hat{\gamma}$, in general, does
not directly estimates $\phi_{\rho}$. When the aim is to estimate
the correlation coefficient, the variables have to be standardized
such that the variances of $x_{1,t}$ and $x_{2,t}$ become homogeneous,
and thus $\phi_{\xi}=\phi_{\rho}$, as shown in Section \ref{subsec:Variance-Homogeneity:-Robust}.
For example, this can be done via a two-step approach. In the first
step, the individual variances $\hat{\sigma}_{1}^{2}$ and $\hat{\sigma}_{2}^{2}$
of $x_{1}$ and $x_{2}$, respectively, are estimated, and the standardized
observations are constructed, $z_{t}=\Bigl(\frac{x_{1,t}}{\hat{\sigma}_{1}},\frac{x_{2,t}}{\hat{\sigma}_{2}}\Bigl)^{\prime}$.
In the second step, the similarity estimator is applied to the standardized
variables $z_{t}$. If $\hat{\sigma}_{1}^{2}$ and $\hat{\sigma}_{2}^{2}$
consistently (and robustly) estimate the corresponding variances,
$\hat{\gamma}$ becomes a consistent estimator of $\phi_{\rho}$ with
the asymptotic distribution given in \eqref{eq:gamma_clt}, however,
the finite sample distribution of $\hat{\gamma}$, in general, will
differ from the one characterized in Section \ref{subsec:Variance-Homogeneity:-Robust}.
Alternatively, the information about marginal volatilities, $\sigma_{1}^{2}$
and $\sigma_{2}^{2}$, can be inferred from dynamic filters, such
as GARCH models.
In case variance homogeneity is not ensured, the similarity estimator
$\hat{\gamma}$ can be interpreted as a consistent lower bound estimator
for the magnitude of $\phi_{\rho}$ due to $|\phi_{\xi}|\leq|\phi_{\rho}|$.
Moreover, the asymptotic distribution in \eqref{eq:gamma_clt} can
be used to construct conservative confidence intervals for $\phi_{\xi}$
(and $\xi$) because $V_{\phi_{r}}\leq\frac{\pi^{2}}{4}$. In this
situation, estimator $\hat{\gamma}$ can be used for conservative,
but theoretically justified, estimation and inference for the latent
correlation coefficient. This aspect can be useful, for example, for
testing whether the correlation coefficient is equal to zero.
\section{\label{sec:Measuring-Similarity-for}Measuring Similarity for Multiple
Variables}
The quantity $\phi_{r,t}$ formulated in \eqref{eq:phi_r} can be
represented as a log-ratio of the \textit{local} variations of vector
$x_{t}$ -- variation of the sum and variation of the difference
of the vector components, where $x_{t}=(x_{1,t},x_{2,t})^{\prime}$
is a zero mean random vector with finite second moments. For fixed
$\sigma_{1}^{2}$ and $\sigma_{2}^{2}$, an increase in $\rho$ makes
the first variation higher and the second variation lower, and conversely.
Naturally, for $\rho\approx\pm1$ the contrast between the variations
is the highest, while for $\rho=0$ the two variations are identical.
Therefore, $\phi_{r,t}$ can be viewed as a local measure of the divergence
between $V(x_{1,t}+x_{2,t})$ and $V(x_{1,t}-x_{2,t})$ which nests
information about the correlation coefficient $\rho$, and this information
can be explicitly recovered once $\sigma_{1}^{2}=\sigma_{2}^{2}$
(see Section \ref{subsec:Variance-Homogeneity:-Robust}).
We note that $V(x_{1,t}+x_{2,t})$ can alternatively be represented
as the variance of the projection of $x_{t}$ onto the direction of
the vector of ones, i.e. the direction of perfect\textit{ similarity}.
Analogously, $V(x_{1,t}-x_{2,t})$ is the variance of the projection
of $x_{t}$ onto the orthogonal direction, i.e. the direction of \textit{dissimilarity}.
Therefore, $\phi_{r,t}$, and consequently $\hat{\gamma}$, indicate
the degree of similarity in the joint variation of $x_{1,t}$ and
$x_{2,t}$ measured by the relative magnitude of variation along the
vector of ones. This idea can be directly generalized to random vectors
of arbitrarily large dimension.
Assume that $x_{t}=(x_{1,t},x_{2,t},...,x_{n,t})^{\prime}$ is a zero
mean $n$-dimensional random vector with finite second moments and
positive definite covariance matrix $\Sigma$. We denote the $n$-dimensional
vector of ones by $\iota_{n}=(1,1,...,1)^{\prime}$. The matrix $P_{n}=\frac{1}{n}\iota_{n}\iota_{n}^{\prime}$
is the projection matrix such that the projection of $x_{t}$ onto
the direction defined by the vector $\iota_{n}$ is given by $P_{n}x_{t}$.
Then, $P_{n}^{\perp}=I_{n}-P_{n}$, where $I_{n}$ is the identity
matrix of dimension $n$, is also the projection matrix such that
$P_{n}^{\perp}x_{t}$ represents the projection of $x_{t}$ onto the
subspace orthogonal to $\iota_{n}$. Then, the local similarity measure
$\phi_{r,t}$ can be naturally extended to the general multivariate
case via
\begin{equation}
\phi_{r,t}=\frac{1}{n}\log\frac{x_{t}^{\prime}P_{n}x_{t}}{x_{t}^{\prime}P_{n}^{\perp}x_{t}},\label{eq:local_sim_mult}
\end{equation}
which captures the variation of $x_{t}$ along the vector of \textit{perfect
similarity}, $\iota_{n}$, relative to the variation of $x_{t}$ along
its orthogonal complement. Note that, for the special case $n=2$,
equation \eqref{eq:local_sim_mult} is reduced to the same expression
for $\phi_{r,t}$ as in \eqref{eq:phi_r}.
As in the bivariate case, the similarity estimator $\hat{\gamma}$
is constructed as $\hat{\gamma}=\frac{1}{T}\sum_{t=1}^{T}\phi_{r,t}$.
Intuitively, $\gamma$ measures the average similarity in how all
$n$ variables move, in terms of both magnitude and direction.
\subsection{\label{subsec:Equicorrelation-Scenario} Equicorrelation Scenario}
In the multivariate scenario with $n>2$, the similarity estimator
$\hat{\gamma}$ has, in general, less tractable interpretation as
compared to the bivariate setting. In some special cases, however,
$\hat{\gamma}$ retains explicit finite sample and asymptotic limits,
and is immediately related to the Pearson correlation coefficient.
Let the entries of covariance matrix $\Sigma$ are denoted by
\[
\Sigma=\left(\begin{array}{cccc}
\sigma_{1}^{2} & \cdot & \cdot & \cdot\\
\sigma_{12} & \sigma_{2}^{2} & \cdot & \cdot\\
\vdots & \vdots & \ddots & \cdot\\
\sigma_{1n} & \sigma_{2n} & \cdots & \sigma_{n}^{2}
\end{array}\right),
\]
and here we assume that all variances are homogeneous, i.e. $\sigma_{1}^{2}=\sigma_{2}^{2}=...=\sigma_{n}^{2}=\sigma^{2}$,
and all covariance (off-diagonal) elements are identical too. This
implies that all pairwise correlations are also identical, and, so,
we can write $\sigma_{ij}=\sigma^{2}\rho$ for all $i$ and $j$,
such that $i\neq j$, where $\rho$ is the common correlation parameter.
Note that, once $\Sigma$ is positive definite by assumption, it implies
that $\rho\in\bigl(-\frac{1}{n-1},1\bigl)$.
The matrix logarithm transformation suggests a convenient method for
reparametrization of correlation matrices in such way that the transformed
correlation elements can be represented as an unconstrained real vector
which can always be mapped to a unique positive definite correlation
matrix. In this sense, the parametrization based on the matrix logarithm
can be viewed as a multi-dimensional generalization of the Fisher
transformation, see \citet{Archakov_Hansen_2021} for more details.
When the matrix logarithm transformation is applied to $\Sigma$ with
homogeneous variances and equal correlations, the general structure
of the matrix is preserved after the transformation, i.e. all diagonal
and all off-diagonal elements remain identical,
\setlength{\arraycolsep}{5pt}
\[
\Sigma=\sigma^{2}\left(\begin{array}{ccccc}
1 & \cdotp & \cdotp & & \cdotp\\
{\color{purple}{\footnotesize \rho}} & 1 & \cdotp & & \cdotp\\
{\color{purple}{\footnotesize \rho}} & {\color{purple}{\footnotesize \rho}} & 1 & \cdots & \cdotp\\
& & \vdots & \ddots & \vdots\\
{\color{purple}{\footnotesize \rho}} & {\color{purple}{\footnotesize \rho}} & {\color{purple}{\footnotesize \rho}} & \cdots & 1
\end{array}\right),\qquad\log\Sigma=\left(\begin{array}{ccccc}
\delta & \cdotp & \cdotp & & \cdotp\\
{\color{blue}\phi_{\rho}} & \delta & \cdotp & & \cdotp\\
{\color{blue}\phi_{\rho}} & {\color{blue}\phi_{\rho}} & \delta & \cdots & \cdotp\\
& & \vdots & \ddots & \vdots\\
{\color{blue}\phi_{\rho}} & {\color{blue}\phi_{\rho}} & {\color{blue}\phi_{\rho}} & \cdots & \delta
\end{array}\right),
\]
as it is shown in \citet{Archakov_Hansen_2024}. The eigenvalues of
matrix $\Sigma$ are $\lambda_{+}=\sigma^{2}\bigl(1+(n-1)\rho\bigl)$
and $\lambda_{-}=\sigma^{2}(1-\rho)$ with multiplicities $1$ and
$n-1$, respectively, see \citet{Olkin_Pratt_1958}. The entries of
$\log\Sigma$ are analytically available and given by $\delta=\frac{1}{n}\log\lambda_{+}+\frac{n-1}{n}\log\lambda_{-}$
and
\begin{equation}
\phi_{\rho}=\frac{1}{n}\log\frac{\lambda_{+}}{\lambda_{-}}=\frac{1}{n}\log\Bigl(\frac{1+(n-1)\rho}{1-\rho}\Bigl).\label{eq:gamma_equi}
\end{equation}
We note that the off-diagonal element $\phi_{\rho}$ depends only
on $\rho$ and $n$, and does not depend on variance $\sigma^{2}$.
Naturally, this result is a generalization of the matrix logarithm
result for $n=2$ in Section \ref{subsec:Variance-Homogeneity:-Robust},
so we preserve notation $\phi_{\rho}$ introduced earlier.
If we additionally assume that random vector $x_{t}$ is elliptical,
the distribution of $\phi_{r,t}$ in \eqref{eq:local_sim_mult} can
be derived explicitly which is reflected in the following result.
\begin{prop}
\label{prop:P3} Assume that $x=(x_{1},x_{2},...,x_{n})^{\prime}$
is a $n$-variate random vector which follows some elliptical distribution
with zero mean and positive-definite covariance matrix $\Sigma$ with
homogeneous variances and identical correlations (equi-correlation
structure), and let denote the common correlation coefficient by $\rho$.
Denote the measure of joint resemblance by $\phi_{r}=\frac{1}{n}\log\frac{x^{\prime}P_{n}x}{x^{\prime}P_{n}^{\perp}x}$,
where $P_{n}=\frac{1}{n}\iota_{n}\iota_{n}^{\prime}$ and $P_{n}^{\perp}=I_{n}-P_{n}$
are orthogonal projection matrices, $i_{n}$ is the $n$-dimensional
vector of ones, and $I_{n}$ is the n-dimensional identity matrix.
Then, $\phi_{r}$ is a random variable from the Logistic-Beta family
with the probability density function
\[
f(\phi_{r})=\frac{1}{\mathcal{B}\Bigl(\frac{1}{2},\frac{n-1}{2}\Bigl)}\cdot\frac{ne^{\frac{1}{2}n(\phi_{r}-\phi_{\rho})}}{\Bigl(1+e^{n(\phi_{r}-\phi_{\rho})}\Bigl)^{\frac{n}{2}}},
\]
where $\phi_{\rho}=\frac{1}{n}\log\Bigl(\frac{1+(n-1)\rho}{1-\rho}\Bigl)$
is the off-diagonal element of $\log\Sigma$.
\end{prop}
Note that $\phi_{r}$ is not an unbiased measure of $\phi_{\rho}$
because $E(\phi_{r})=\phi_{\rho}-\omega_{n}$, where $\omega_{n}\ge0$
is a non-monotone function of $n$ such that $\omega_{2}=0$ and $\omega_{n}\rightarrow0$
as $n\rightarrow\infty$. The exact expression for $\omega_{n}$ is
provided in Appendix. Variance of $\phi_{\rho}$ is inversely proportional
to $n^{2}$ and is given by
\[
V(\phi_{r})=\frac{1}{n^{2}}\biggl(\psi^{\prime}\Bigl(\frac{n-1}{2}\Bigl)+\frac{\pi^{2}}{2}\biggl),
\]
where $\psi^{\prime}(t)$ is the trigamma function. Naturally, when
vector dimension $n$ increases, $\phi_{r}$ absorbs more information
about the common correlation coefficient from the cross-sectional
dimension and, so, provides an increasingly more efficient signal
about the latent correlation level. In Figure \ref{fig:pdf2}, we
illustrate the Logistic-Beta probability density function of $\phi_{r}-\phi_{\rho}$
for several selected values of $n$.
In the special case of homogeneous variances and correlations, Proposition
\ref{prop:P3} helps to characterize the asymptotic properties of
the similarity estimator $\hat{\gamma}$. With a bias correction,
$\hat{\gamma}+\omega_{n}$ becomes a consistent and robust estimator
for the off-diagonal element $\phi_{\rho}$ of the log-transformed
covariance matrix $\Sigma$, where $\phi_{\rho}$ represents a monotone
transformation of the equicorrelation coefficient $\rho$, similarly
to the Fisher transformation function in the bivariate case. The asymptotic
variance of $\hat{\gamma}$ is given by $V(\phi_{r})$, and the estimator
becomes more efficient with dimension $n$ due to the growth of relevant
cross-sectional information.
\begin{figure}[h]
\begin{centering}
\includegraphics[scale=0.8]{plots/simi_density_equi_pdf}
\par\end{centering}
\caption{\footnotesize\label{fig:pdf2} Probability density functions of $\phi_{r}-\phi_{\rho}$
for different dimensions $n$ of vector $x$.}
\end{figure}
Despite the structure assumed for $\Sigma$ is restrictive and does
not occur in practice frequently, the results reveal several qualitative
aspects which are characteristic of the similarity estimator, $\hat{\gamma}$,
in more general and realistic scenarios. At first, $\hat{\gamma}$
is a robust estimator of an aggregate measure of association between
considered variables. Although this does not directly correspond to
the average correlation level for an arbitrary covariance matrix $\Sigma$,
it still can be interpreted as an indicator of \textit{joint} statistical
similarity between variables, i.e. how similar they are in terms of
both magnitude and direction. At second, if elements in $\Sigma$
are sufficiently homogeneous, $\hat{\gamma}$ is expected to be more
efficient once $n$ gets larger. This is because each additional variable
provides some extra information about the average level of similarity.
\section{Empirical Applications}
Although the similarity estimator can be used as a standalone measure
of statistical association, in this section, we consider empirical
applications where the latter is used for estimation and modeling
the Pearson correlation coefficient. The two considered applications
utilize financial market data for measuring correlations between stock
returns.
\subsection{Robust Confidence Intervals for Stock Return Correlations }
We apply the similarity estimator for robust interval estimation of
financial correlations. For this analysis, we use intraday transaction
data from the TAQ database cleaned according to the recommendations
provided in \citet{Barndorff-Nielsen2009}. For exposition purposes
we consider daily correlations between Apple (AAPL) and Exxon Mobil
(XOM) returns during the period between February and April 2020 (62
trading days) which corresponds to the COVID-19 outbreak.
We note that AAPL and XOM belong to different sectors of the economy
(Tech vs Energy) which implies that the correlation between these
stocks is not supposed to be particularly strong in ordinary periods
of time. During the COVID-19 crisis, however, the overall correlation
level has significantly elevated across the entire market due to i)
an initial decline of the market that affected almost all sectors
and industries (largely began on February 19), and ii) the subsequent
common recovery (began after March 23). Therefore, we may expect notable
changes in the latent correlation trajectory for the considered assets
during the analyzed time period.
For each considered trading day, we construct samples of intraday
(logarithmic) returns, $r_{t}=(r_{1,t},r_{2,t})^{\prime}$, $t=1,...,T$,
for a range of selected frequencies by using the previous tick interpolation
scheme which was introduced in \citet{Wasserfallen_Zimmermann_1985}
and is used routinely in high frequency econometrics (see, for example,
\citet{Hansen_Lunde_2006}). Due to the duration of a typical trading
day is $6.5$ hours, we obtain a sample with $T=78$ observations
for the frequency $\Delta=5$ min, while we have $T=390$ observations
for $\Delta=1$ min.
\begin{figure}[!ph]
\begin{centering}
\subfloat[Intraday returns are observed at $\Delta=10$ min frequency ($T=39$
observations).]{\centering{}\includegraphics[scale=0.52]{plots/corr_conf_xom_aapl_covid2_10m_v2}}
\par\end{centering}
\begin{centering}
\subfloat[Intraday returns are observed at $\Delta=5$ min frequency ($T=78$
observations).]{\centering{}\includegraphics[scale=0.52]{plots/corr_conf_xom_aapl_covid2_5m_v2}}
\par\end{centering}
\centering{}\subfloat[Intraday returns are observed at $\Delta=1$ min frequency ($T=390$
observations).]{\centering{}\includegraphics[scale=0.52]{plots/corr_conf_xom_aapl_covid2_1m_v2}}\caption{\footnotesize\label{fig:spring_2020} Confidence intervals for daily
correlations between Apple (AAPL) and Exxon Mobil (XOM) estimated
using intraday returns. The analyzed period is between February and
April 2020 (62 trading days). The intervals are obtained for coverage
probabilities 90\% (boxes) and 95\% (whiskers). Red circles correspond
to daily Realized (sample) Correlations, while blue squares correspond
to daily correlations estimated by the Kendall tau coefficient. }
\end{figure}
To construct an interval with a specified coverage probability, we
estimate the correlation using the similarity estimator $\hat{\gamma}$
on standardized intraday returns, $z_{t}=(z_{1,t},z_{2,t})^{\prime}$,
where $z_{i,t}=r_{i,t}/\hat{\sigma}_{i}$ with $\hat{\sigma}_{i}$
is a sample standard deviation of $r_{i,t}$. Next, we obtain the
required interval around $\hat{\gamma}$ using the exact quantiles\footnote{The results only marginally differ in case the intervals are constructed
using the asymptotic distribution of $\hat{\gamma}$ given in \eqref{eq:gamma_clt}.
We nonetheless report intervals based on exact quantiles as they provide
more conservative assessment of the sampling error.} of the sampling distribution characterized in \eqref{eq:cf}, and
map the interval from the Fisher correlation scale to the Pearson
correlation scale by using the inverse Fisher transformation. The
resulting intervals for daily correlations are obtained for 90\% and
95\% coverage probabilities for the three selected intraday frequencies
($\Delta=10$, $5$, and $1$ min) and are shown in Figure \ref{fig:spring_2020}.
It also provides point estimates of correlation calculated with two
benchmark estimators -- the sample (realized) correlation estimator
and the Kendall rank coefficient (see Section \ref{sec:Realized-Similarity-Estimator}
for the corresponding expressions). The results in Figure \ref{fig:spring_2020}
reveal several interesting aspects.
When constructed with returns sampled at $\Delta=10$ min frequency,
the intervals are very wide and thus not informative about the underlying
correlation level. For the majority of trading days, the intervals
do not even allow to reject the null hypothesis of zero correlation.
Such conservative interval widths reflect the trade-off required to
achieve robustness and invariance of the constructed intervals under
arbitrary fat-tailed data. At higher frequencies, where more observations
are available, intervals naturally shrink, and variability of the
correlation level over time becomes more apparent. Thus, when constructed
at $\Delta=1$ min frequency, the intervals are sufficiently narrow
allowing to clearly observe a sharp surge of the correlation level
in the second half of February and gradual non-monotone subsequent
decay. Interestingly, the estimated intervals appear to cluster within
weeks, which may indicate highly persistent correlation dynamics.
\begin{figure}[ph]
\begin{centering}
\includegraphics[scale=0.75]{plots/scatter_xom_aapl_1m}
\par\end{centering}
\caption{\footnotesize\label{fig:spring_2020_scatter} Scatter plot for 1-min
intraday log-returns ($\times$100) for XOM and AAPL on February 27,
2020 (left plot, estimators disagree), and February 28, 2020 (right
plot, estimators agree).}
\end{figure}
It is important to mention that since we standardize raw intraday
returns by using sample standard deviations, which is not a robust
estimator of dispersion, the resulting standardized returns do not
necessarily have homogeneous variances. Recall that under non-perfectly
homogeneous variances, the similarity estimator underestimates the
true correlation level, as it is discussed in Section \ref{subsec:Realized-Similarity-Estimator}.
Despite this, the constructed robust intervals based on the similarity
measure show good agreement with both benchmark estimators for almost
all trading days even at the highest considered frequency, where the
intervals are sufficiently narrow. For a number of trading days, however,
we observe that the sample correlations fall below the constructed
robust intervals as well as below the estimates obtained with the
Kendall estimator. This can be explained by the common presence of
price jumps in the market data inducing outliers in high-frequency
returns. While the realized correlation estimator exhibits a downward
bias in such cases, the Kendall and similarity estimators remain robust.
Figure \ref{fig:spring_2020_scatter} illustrates intraday returns
on AAPL and XOM sampled at 1-min frequency on February 27 and 28.
Despite $\hat{\gamma}$ indicates similar correlation levels for both
days, the realized correlation and the Kendall coefficient show sufficiently
lower correlation on February 27, and align with $\hat{\gamma}$ on
February 28. As we may see from the scatter plot, on February 27 we
observe a sufficient number of potentially outlying observations distributed
in all directions around a more compact and regularly shaped core
of the sample. This may drive the benchmark measures downwards and
cause disagreement between the estimators. In contrast, on February
28 outliers appear to be more directionally aligned with the core
part of the sample. This results in a more regular elliptical shape
for the sample data, with all estimators largely agreeing on the estimated
correlation value.
We emphasize that the presented analysis serves rather for experimental
and illustration purposes. The assumption of independent and identically
distributed observations can hardly be justified due to stochastic
volatility, market microstructural effects, and other stylized artifacts
which are typically attributed to high frequency returns. Moreover,
we may expect that the correlation level change over a trading day,
so the obtained estimates should rather be interpreted as indicators
of average daily correlation. A more careful adaptation of the similarity
estimator for high frequency financial data is a challenging and interesting
question which is a subject of ongoing work.
\subsection{A Robust Multivariate GARCH Model}
We suggest a new specification for the multivariate GARCH model that
incorporates the similarity measures introduced in the paper. Let
$r_{t}=(r_{1,t},r_{2,t},...,r_{n,t})^{\prime}$ be $n$-dimensional
vector of asset returns, for $n\geq2$, observed at discrete time
moments $t=1,...,T$. The key object of interest is the conditional
covariance matrix of $r_{t}$, which we denote by $H_{t}=V(r_{t}|\mathcal{F}_{t-1})$,
where $\{\mathcal{F}_{t}\}$ is the natural filtration for $r_{t}$.
We follow the logic of the Dynamic Conditional Correlation (DCC) approach
introduced in \citet{Engle_2002}, and decompose the $H_{t}$ into
variance and correlation components,
\[
H_{t}=\Lambda_{h_{t}}^{1/2}C_{t}\Lambda_{h_{t}}^{1/2},
\]
where $\Lambda_{h_{t}}=\text{diag}(h_{1,t},h_{2,t},...,h_{n,t})^{\prime}$
with $h_{i,t}$ is the conditional variance of an individual asset
return $i$, such that $h_{i,t}=[H_{t}]_{ii}$, for $i=1,...,n$,
and $C_{t}=\text{corr}(r_{t}|\mathcal{F}_{t-1})$ is the positive-definite
conditional correlation matrix of $r_{t}$. This structure allows
to effectively split the modeling of $H_{t}$ into separate modeling
of conditional variances and correlations.
We assume that all conditional means are constant and denote them
by $\mu_{i}=E(r_{i,t}|\mathcal{F}_{t-1})$, for $i=1,...,n$, which
is a standard assumption in the GARCH literature. Then, we can formulate
return equations as follows,
\begin{equation}
r_{i,t}=\mu_{i}+h_{i,t}^{1/2}z_{i,t},\qquad i=1,...,n\quad\text{and}\quad t=1,...,T,\label{eq:ret_eq}
\end{equation}
where $z_{i,t}$ are standardized returns, such that $E(z_{i,t}|\mathcal{F}_{t-1})=0$
and $V(z_{i,t}|\mathcal{F}_{t-1})=1$. Conditional variances $h_{i,t}$,
for $i=1,...,T$, can be modeled with any appropriate univariate dynamic
GARCH equation. For example, EGARCH(1,1) specification by \citet{Nelson_1991}
can be used,
\begin{equation}
\log h_{i,t}=\alpha_{i}+\beta_{i}\cdot\log h_{i-1,t}+\kappa_{i}\cdot\underbrace{\Bigl(z_{i,t-1}+\eta_{i}\cdot\bigl(|z_{i,t-1}|-E|z_{i,t-1}|\bigl)\Bigl)}_{=g_{i}(z_{i,t-1})},\label{eq:var_eq}
\end{equation}
where $g_{i}(z_{i,t-1})\in\mathcal{F}_{t-1}$ and, so, it provides
an observation-driven update for $\log h_{i,t}$.\footnote{The classification of time-varying parameter models into the classes
of observation-driven and parameter-driven models goes back to \citet{Cox_1981}.
See also an instructive discussion about observation-driven modeling
in \citet{Koopman_Lucas_Scharth_2016}.} Denote the vector of standardized returns at period $t$ by $z_{t}=(z_{1,t},z_{2,t},...,z_{n,t})^{\prime}$,
and note that $\text{corr}(z_{t}|\mathcal{F}_{t-1})=C_{t}$. Therefore,
$z_{t}$ carries the information about the conditional correlation
structure of raw asset returns.
The methodological novelty of the suggested model is related to the
way how the dynamics for conditional correlation matrix $C_{t}$ is
modeled. Following the approach in \citet{Archakov_Hansen_Lunde_2025},
we set up the dynamics for the off-diagonal elements of the log-transformed
conditional correlation matrix, $\log C_{t}$. The novel idea is to
use the similarity measure $\phi_{r,t}$, characterized in Sections
\ref{sec:Measure-of-Similarity} and \ref{sec:Measuring-Similarity-for},
as a natural signal about the current level of $\log C_{t}$. Since
the conditional variances of standardized returns in $z_{t}$ are
homogeneous, and assuming the return vector follows an elliptical
distribution, we have that $\phi_{r,t}$, constructed with $z_{t}$,
represents a robust empirical measure of the transformed correlation
coefficients, as it was shown in Propositions \ref{prop:P1} and \ref{prop:P3}
for the bivariate and the equicorrelation structures, respectively.
\subsubsection{Conditional Correlations for Bivariate Model}
We begin with the bivariate case ($n=2$), where $\log C_{t}$ can
be fully characterized by only a single correlation parameter, $\phi_{\rho,t}$,
which is the Fisher transformation of the single correlation parameter
$\rho_{t}$ in $C_{t}$. We consider the following dynamic specification
for $\phi_{\rho,t}$,
\begin{equation}
\phi_{\rho,t}=\alpha+\beta\cdot\phi_{\rho,t-1}+\kappa\cdot\underbrace{\frac{1}{2}\log\frac{(z_{1,t-1}+z_{2,t-1})^{2}}{(z_{1,t-1}-z_{2,t-1})^{2}}}_{=\phi_{r,t-1}},\label{eq:bivar_garch}
\end{equation}
where $\phi_{r,t}$ is a local similarity measure defined in \eqref{eq:phi_r}.
The structure of this recursive equation is analogous to the classical
GARCH equation for volatility. The autoregressive term $\beta\cdot\phi_{\rho,t-1}$
accounts for persistence in $\phi_{\rho,t}$, while term $\kappa\cdot\phi_{r,t-1}$
represents an observation-driven innovation for the correlation dynamics.
Note that the vector of standardized returns $z_{t}=(z_{1,t},z_{2,t})^{\prime}$
is elliptical, by assumption, and have homogeneous unit variances.
Therefore, by Proposition \ref{prop:P1}, $\phi_{r,t}$ becomes an
unbiased signal about the latent correlation, $\phi_{\rho,t}$, on
the Fisher scale. At the same time, $\phi_{r,t}$ is robust to the
presence of outliers and fat-tailed distributions of the observed
returns. This is a particularly useful feature in modeling financial
returns, where the fat tails and jumps are commonly found among stylized
empirical regularities.
An apparent advantage of the model suggested in \eqref{eq:bivar_garch}
is an unconstrained support of $\phi_{\rho,t}$ which allows to avoid
any additional restrictions on the dynamic process for ensuring positive
definiteness of $C_{t}$. Recall that, in the classical DCC-GARCH
models, the conditional correlation dynamics is specified through
the matrix process, such that the resulting matrix needs an extra
adjustment for positive definiteness and for the estimated conditional
correlations remain consistent (see also \citet{Aielli_2013}, \citet{Brownlees_Llorens_2023}).
The specification in \eqref{eq:bivar_garch} is closely related to
the angular DCC-GARCH model suggested in \citet{Jarjour_Chan_2020},
where the conditional correlation dynamics was specified for non-transformed
$\rho_{t}$, and $r_{t}=\frac{2z_{1,t}z_{2,t}}{z_{1,t}^{2}+z_{2,t}^{2}}$
was used as a dynamic innovation term. Although $r_{t}$ is a robust
signal of correlation, it is not an unbiased measure of $\rho_{t}$
because, in general, $E(r_{t}|\mathcal{F}_{t-1})\neq\rho_{t}$. Furthermore,
extra restrictions for the parameter coefficients must be imposed
to ensure positive definiteness. In light of this, the specification
in \eqref{eq:bivar_garch} appears to be a more natural and convenient
approach to modeling dynamic correlations.
Since daily returns are typically fat-tailed, it is popular to model
the conditional distribution of $r_{t}$ by some elliptical distribution
that allows for excess kurtosis, such as the Student's $t$-distribution,
the symmetric multivariate stable distribution, etc.\footnote{Note that the assumption of elliptical distribution for returns requires,
in general, to introduce an additional parameter for the stochastic
radial component which controls tail behavior.} Once the distribution for $r_{t}$ is specified, the model given
by \eqref{eq:ret_eq}-\eqref{eq:bivar_garch} can be estimated using
the Maximum Likelihood method. We note that many extensions, which
are common for the DCC-GARCH models, are also readily available for
our model. For example, these extensions may include correlation targeting,
an increased number of lags in \eqref{eq:bivar_garch}, or incorporating
realized measures of correlation in spirit of \citet{Archakov_Hansen_Lunde_2025}.
\begin{figure}[ph]
\begin{centering}
\includegraphics[scale=0.65]{plots/simi_garch_CVX_MRO}
\par\end{centering}
\caption{\footnotesize\label{fig:rmg_bivar} Robust bivariate GARCH estimation
with daily returns on CVX and MRO. The estimated conditional correlation
(blue line) and 5-min realized correlations calculated on a daily
basis (gray line) are shown for the period between January 2005 and
December 2020 (4744 trading days).}
\end{figure}
In Figure \ref{fig:rmg_bivar}, we provide an example of the estimated
robust conditional correlation trajectory. For this example, we analyze
the correlation between Chevron Corp. (CVX) and Marathon Oil Corporation
(MRO) for the 16-year sample period between 2005 and 2020. We use
close-to-close daily returns, adjusted for stock splits and dividends,
from the CRSP US Stock Database. While the estimated trajectory (solid
blue line) has an apparent time-varying dynamics, we observe persistently
high correlations over the entire sample period, which is not a surprising
empirical evidence due to both companies belong to the energy sector.
It is interesting that the estimated correlations are systematically
higher than the daily realized (sample) correlations calculated with
5-min intraday returns (gray line). A possible explanation is that
the (non-robust) realized correlations are often biased towards zero
due to presence of price jumps and extreme returns.
\subsubsection{Dynamic Equicorrelation Model}
Based on results in Section \ref{subsec:Equicorrelation-Scenario},
we can specify an elegant and parsimonious multivariate GARCH specification
for arbitrary large dimension $n$. For this, we assume an equicorrelation
structure for $C_{t}$, such that all conditional correlations are
identical and parametrized by a single dynamic coefficient, $\rho_{t}$.
The DECO-GARCH model was originally introduced in \citet{Engle_Kelly_2012},
and was extended to accommodate realized measures of variances and
correlations in \citet{Archakov_Hansen_Lunde_2025}. We formulate
the dynamic correlation process for the off-diagonal elements of the
log-matrix transformation. In case $C_{t}$ has the equicorrelation
structure, the off-diagonal entries of $\log C_{t}$ are all identical
and equal to $\phi_{\rho,t}$, and the analytical relationship between
$\phi_{\rho,t}$ and $\rho_{t}$ is provided in \eqref{eq:gamma_equi}.
We formulate dynamics for $\phi_{\rho,t}$ as follows,
\begin{equation}
\phi_{\rho,t}=\alpha+\beta\cdot\phi_{\rho,t-1}+\kappa\cdot\underbrace{\frac{1}{n}\log\frac{z_{t-1}^{\prime}P_{n}z_{t-1}}{z_{t-1}^{\prime}P_{n}^{\perp}z_{t-1}}}_{=\phi_{r,t-1}},\label{eq:deco_garch}
\end{equation}
where $\phi_{r,t}$ is the local similarity measure introduced in
Section \ref{sec:Measuring-Similarity-for}, $P_{n}=\frac{1}{n}\iota_{n}\iota_{n}^{\prime}$
and $P_{n}^{\perp}=I_{n}-P_{n}$ are the orthogonal projection matrices,
$i_{n}$ is the $n$-dimensional vector of ones, and $I_{n}$ is the
n-dimensional identity matrix. This specification effectively retains
all the benefits of specification \eqref{eq:bivar_garch} formulated
for the bivariate case. Namely, $\phi_{r,t}$ is intrinsically robust
to extreme observation and is a relevant signal of $\phi_{\rho,t}$
with a known constant bias term, $E(\phi_{r,t}|\mathcal{F}_{t-1})=\phi_{\rho,t}+\omega_{n}$.
Such fixed bias is harmless as it is absorbed by the constant coefficient
$\alpha$ in \eqref{eq:deco_garch}. Unconstrained range of $\phi_{\rho,t}$
implies that no extra restrictions are needed in order to ensure that
estimated matrices $C_{t}$ are positive definite.
The equicorrelation assumption imposes a tight constraint on $C_{t}$,
which is rarely empirically plausible. This structure, however, provides
a substantial dimension reduction since, instead of modeling $\frac{n(n-1)}{2}$
dynamic correlations, we effectively model only a single common correlation
coefficient. The estimated dynamics can be interpreted as a time-varying
average correlation level, or an index of market co-movement, and
can serve as a useful state variable in various contexts. For example,
when a sufficiently representative sample of assets is available,
$\phi_{\rho,t}$ can be used as a barometer of diversification potential
for portfolio management purposes, or as an aggregate correlation
index which can be informative about overall market uncertainty and
risk.
As in the previous case, the model can be estimated using the Maximum
Likelihood. An important advantage of the DCC structure for multivariate
GARCH is that the model can be effectively estimated in two steps.
In the first step, the univariate GARCH models \eqref{eq:ret_eq}-\eqref{eq:var_eq}
are estimated individually for $i=1,...,n$, and the standardized
returns, $\hat{z}_{t}$, are obtained. In the second step, the conditional
correlation dynamics of $C_{t}$, given in \eqref{eq:deco_garch},
is estimated by using $\hat{z}_{t}$ obtained in the first step. Such
two-stage estimation can facilitate computation complexity dramatically,
especially if high dimensional return vectors are modeled.
\begin{figure}[ph]
\begin{centering}
\includegraphics[scale=0.65]{plots/simi_garch_equi_annotated_deco}
\par\end{centering}
\caption{\footnotesize\label{fig:rmg_index} Robust multivariate DECO-GARCH
estimation with daily returns on 9 selected assets. The conditional
equicorrelation index estimated with the new robust specification
(blue line), the equicorrelation index estimated with the standard
DECO-GARCH model (red line), and the 5-min average realized correlations
calculated on a daily basis (gray line) are shown for the period between
January 2005 and December 2020 (4744 trading days).}
\end{figure}
For illustration purposes, we estimate the new model for the sample
of nine stocks using the daily close-to-close returns, adjusted for
stock splits and dividends, from the CRSP US Stock Database. In particular,
we consider three stocks from the energy sector (CVX, MRO, and OXY),
three stocks from the health care sector (JNJ, LLY, and MRK), and
three stock from the information technology sector (AAPL, MU, and
ORCL), and estimate the model for the 16-year sample period between
2005 and 2020. The estimation results are illustrated in Figure \ref{fig:rmg_index}.
The conditional equicorrelation index estimated with the new robust
model is shown with a blue line, while the index estimated with a
standard DECO-GARCH model by \citet{Engle_Kelly_2012} is shown in
red. We observe that both indices exhibit a significant amount of
variation and often display visible reactions around major historical
events. Although the correlations estimated by the two models follow
similar trajectories, numerous local divergences are evident over
the sample period, which can be attributed to the intrinsic robustness
of the newly proposed model. For example, we observe that the new
correlation index exhibits a more modest response during the European
Debt Crisis in 2011 and the outbreak of COVID-19, periods in which
many co-directional extreme returns were observed. This evidence points
to distinctive informational content in the new robust correlation
index, and its evaluation suggests a promising direction for further
empirical analysis.
\section{Conclusion}
In this paper, we suggest a new measure of statistical association
which captures the similarity between multiple random variables and
is grounded in the ideas of \citet{Thorndike_1905} and \citet{Fisher_1919}.
The similarity is defined as the relative extent of joint variation
along the direction of the vector of ones, which can be viewed as
the direction of perfect similarity for random outcomes, measured
in both sign and magnitude. The suggested similarity estimator is
intrinsically insensitive to extreme observations, caused by fat-tailed
data and outliers, and this ensures robustness for the entire sampling
distribution of the estimator.
We analyze statistical properties of the similarity estimator for
the entire class of elliptical random vectors. In the bivariate case
and under assumption of variance homogeneity, we demonstrate that
the similarity estimator is a consistent and robust estimator of the
Pearson correlation coefficient (on the Fisher scale). The robustness
of the estimator also emerges in exact and analytically available
finite sample distribution. For example, the similarity estimator
can be used for construction of exact confidence intervals for correlations
in the presence of noisy and/or heavy-tailed data as well as for many
other robust inference applications. In case variances are not identical,
the similarity estimator preserves its connection to the Pearson correlation
coefficient by becoming the estimator of its lower bound, and thus
can be useful for conservative estimation and inference. When the
estimator is applied to multiple (more than two) variables, it can
be interpreted as a robust estimator of an aggregate correlation level
among the considered variables, while in the special case, when variances
and correlations are homogeneous, it is explicitly related to the
correlation coefficient.
An intrinsic robustness of the similarity estimator comes at the expense
of lower efficiency. This motivates to consider efficiency improving
modifications of the estimator based, for example, on sub-sampling
methods. This research direction has particularly high potential in
the context of intraday financial data, where a substantial amount
of observations can be additionally exploited at high frequencies.
Another potentially useful methodological contribution lies in the
development of new methods for composite estimation of large-scale
correlation matrices. For instance, the correlation elements can be
estimated separately using the bivariate similarity estimator, and
then the resulting matrix is projected to the sub-space of a positive
definite correlation matrices according to some appropriate criterion
of optimality. We reserve these topics for future work.
We illustrate the empirical performance of the similarity estimator
by applying it to intraday stock returns. The robust confidence intervals
constructed by using the new method show strong agreement with widely
used robust alternatives to the sample correlation estimator. This
evidence suggests that the similarity estimator can be a reliable
tool for estimation and inference of financial correlations with high-frequency
data. It migth be especially useful when dealing with particularly
noisy asset classes, such as crypto-currencies.
As an econometric application, we develop a novel robust multivariate
GARCH model in which the conditional correlation process is modeled
using the matrix logarithm transformation, and its dynamics are driven
by the similarity measure. The suggested specification naturally retains
positive definiteness for the filtered correlation structure and ensures
robustness in the presence of fat tails and outliers in the data.
A straightforward extension would be to accommodate robust realized
measures of correlation calculated using the similarity estimator
with intraday data in spirit of \citet{Archakov_Hansen_Lunde_2025}.
Furthermore, the similarity estimator can be also appropriate for
modeling dynamic correlations with the score-driven approach suggested
in \citet{CKL13}, or with the parameter driven models such as state-space
or stochastic volatility models. We leave these avenues for future
research.
\bibliographystyle{apalike}
\bibliography{literature}
\noindent{\small\newpage}{\small\par}