Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,251 characters · 13 sections · 64 citation commands
A Robust Similarity Estimator
{\smallKeywords:}{ Correlation, Robust Estimation, Robust Inference, Fisher Transformation, Matrix Logarithm, High Frequency Data, Multivariate GARCH}
{\smallJEL Classification:}{ C13, C30, C38, C58 }
Measuring statistical association and dependence between random variables is a broad topic in statistics with a long history. The most popular and widely used measure of association is the linear correlation coefficient, or the Pearson correlation coefficient (denoted by $\rho$), which represents the covariance between two random variables scaled by the product of their standard deviations. Pearson's correlation naturally appears in many statistical and econometric frameworks such as linear regression, multivariate GARCH models, or network analysis. In financial econometrics literature, correlations play a critical role in risk management, hedging, optimal portfolio allocation, analysis of systemic risk, etc.
Both estimation and inference for correlations are practically challenging, especially in the settings where the sample size is limited. For example, under sufficiently mild assumptions, the well-known sample correlation estimator ($\hat{\rho}$) is a consistent estimator of $\rho$. Although it is an efficient estimator when observations are independent and normally distributed, its variance is inflated in the presence of heavy-tailed data, and the estimator remains notoriously sensitive to outliers. In addition, the finite sample distribution of $\hat{\rho}$ converges very slowly to the asymptotic limit with both the shape and spread often strongly depend on the properties of underlying data. As a result, potential distortions in both the central tendency and the sampling distribution of $\hat{\rho}$ undermine robust estimation of the correlation coefficient and complicate statistical inference.
In the series of seminal papers, Ronald A. Fisher proposed a continuous transformation (now known as the Fisher transformation) for $\hat{\rho}$ that offers several advantages including variance stabilization and a symmetric, nearly Gaussian sampling distribution for the transformed sample correlation, even in relatively small samples (see Fisher_1921, Hotelling_1953). The rapid convergence to the asymptotic distribution has made the Fisher transformation popular for conducting statistical inference about $\rho$, even with limited data. The estimation and inference, however, remain fragile, as the transformation does not provide robustness to outliers, and the sampling variance remains inflated under heavy-tailed data distributions.
The lack of robustness of the sample correlation is commonly addressed through the use of alternative correlation measures that are, by construction, insensitive to extreme observations. Notable examples include the Quadrant estimator (Greiner_1909, Kendall_1949, Blomqvist_1950) and the Kendall rank correlation coefficient (esscher1924method, Kendall_1938). These measures are based on the relative number of concordant and discordant pairs of observations and are therefore intrinsically robust to outliers. Under mild assumptions, they can be transformed into consistent (though not efficient) estimators of $\rho$ (see, for example, Croux2010). Despite their robustness and consistency, estimators based on rank correlation measures typically have sampling distributions that remain sensitive to the underlying data generating process. This sensitivity poses a substantial challenge for inference on correlations in applied empirical analysis, where the true distribution of the data is typically unknown.
In this paper, we develop a new estimator of statistical association between two random variables that is consistent for a specific functional of the covariance matrix of the underlying variables. We refer to it as the similarity estimator. The new estimator is inspired by works of Thorndike_1905 and Fisher_1919, and is based on a measure of similarity between two variables that accounts for both sign and magnitude. While the estimator retains a meaningful interpretation on its own and can be used as an alternative measure to gauge statistical relationship between random variables, under the special conditions of elliptically distributed data with homogeneous variances it becomes a consistent estimator of the Pearson correlation $\rho$, on the Fisher scale, and exhibits excellent finite sample properties. In particular, it admits a robust sampling distribution that is invariant over the class of elliptically distributed data with arbitrary kurtosis parameters. The sampling distribution is available via the known characteristic function. This property enables not only robust estimation but also reliable inference for correlations, even in small samples. For example, it allows to conduct robust interval estimation for a given coverage probability. Naturally, the efficiency of the similarity estimator is lower than that of, for example, the sample correlation estimator, reflecting the price paid for its intrinsic robustness. This results in more conservative inference and, on average, wider confidence intervals. However, these intervals remain robust to both outliers and extremely heavy-tailed data distributions.
We propose a natural generalization of the similarity estimator to the multivariate setting involving an arbitrary number of variables. The resulting generalized similarity measure captures the relative variation of the observed data along a similarity direction, defined by the vector of ones, and can be interpreted as a measure of joint similarity across multiple variables. This estimator naturally inherits robustness against outliers and nests the bivariate similarity estimator as a special case. For elliptically distributed data with homogeneous variances and identical pairwise correlations, the multivariate similarity estimator becomes a consistent estimator of the (transformed) equicorrelation parameter, thus permitting a correlation-based interpretation analogous to the bivariate case.
Estimating financial correlations from high-frequency data represents a particularly promising area for applications. While intraday data provide a rich set of observations, offering the potential for efficient estimation and accurate inference, the analysis of such data is accompanied by multiple econometric challenges. For example, large instantaneous price movements, or jumps, which are prevalent in financial markets, generate extreme (outlying) return observations that often induce a downward bias in estimated correlations. In addition, at sufficiently high frequencies, the estimation of realized covariances and correlations may be affected by various forms of market microstructure noise (see Hansen_Lunde_2006, BandiRussellRES08), data asynchronicity and the Epps effect (see Ren`o2003), as well as other adverse artifacts.
The literature on covariance and correlation estimation using high-frequency data proposes a wide range of estimators with varying degrees of robustness to the aforementioned challenges. In practice, realized correlations are most commonly obtained from multivariate volatility estimators by rescaling estimated covariances to the corresponding Pearson correlations (see BNS:2004, AitSahaliaFanXiu2010, BNHLS:2011, HansenHorelLundeArchakov, among others). While many of these methods are explicitly designed to handle various microstructural effects, they nevertheless remain vulnerable to price jumps and other outliers. A common remedy is the application of truncation or thresholding techniques to filter out extreme returns (see, for example, Mancini2001, AndersenDobrevSchaumburg2012). An alternative approach is to employ robust correlation measures, such as the Kendall and Quadrant correlation estimators, along with their modifications and adaptations for high-frequency data (see VanderElst_Veredas_2015, Hansen_Luo_2023, among others).
The intrinsic robustness of the similarity estimator to extreme observations and heavy-tailed data makes it a promising tool for correlation estimation in high-frequency data settings. A key feature that distinguishes the similarity estimator from existing alternatives is its ability to deliver robust interval estimates for correlations, enabled by the availability of a robust sampling distribution. We provide an empirical illustration by constructing robust confidence intervals for daily correlations between several selected stocks from the U.S. stock market during the COVID-19 outbreak. In our empirical application, we estimate correlations at moderately low frequencies (ranging from 1 to 10 minutes), which helps mitigate microstructure effects while still providing a sufficient number of intraday observations for accurate estimation. Our results are notable in several respects. First, despite the intervals are inherently conservative, pronounced correlation dynamics are clearly identified over the sample period. Second, the estimated intervals exhibit visible clustering over sequences of several consecutive trading days, indicating estimation stability and suggesting persistent dynamics in the underlying correlation process. Finally, the constructed intervals show strong agreement with robust correlation estimates based on the Kendall coefficient, supporting the reliable performance of the similarity estimator in high-frequency data applications.
We note that, in our empirical illustration, the similarity estimator is applied while ignoring several features of intraday data that are inconsistent with the assumptions underlying its theoretical properties. Accordingly, we leave a careful adaptation of the proposed estimator to high-frequency settings for future research.
Another area for applications is the modeling of correlation dynamics for vectors of asset returns. The traditional approach builds on extensive literature on multivariate GARCH or stochastic volatility models, in which the conditional correlation process is assumed, either directly or indirectly, to evolve over time (see the respective surveys in Bauwens_Laurent_Rombouts_2006 and Asai_McAleer_Yu_2006). The models are typically applied to relatively low-frequency data, such as daily or weekly asset returns. Key challenges of this class of models include preserving positive definiteness of the conditional correlation structures, as well as the rapidly increasing computational complexity as the number of assets grows.
We propose a new class of multivariate GARCH models that employ the similarity measure to model correlation dynamics. Our approach follows the Dynamic Conditional Correlation (DCC) framework introduced in Engle_2002, in which the conditional correlation process is modeled in isolation from the conditional volatilities. In the bivariate version of our model, correlation dynamics is specified directly in the Fisher scale, ensuring positive definiteness of the correlation matrix by construction. The similarity measure, applied to past (standardized) returns, serves as a robust observation-driven update for the Fisher-transformed conditional correlation process.
We further develop a parsimonious model specification in which an arbitrary number of assets can be accommodated without substantially increasing computational complexity. To this end, we impose the equicorrelation assumption, similar to Engle_Kelly_2012, and specify dynamics for the conditional equicorrelation parameter under the matrix logarithmic transformation. This formulation eliminates the need for additional constraints to guarantee positive definiteness, in contrast to the classical DCC framework. The multivariate similarity measure, computed on standardized returns, is then used as an observation-driven update for the dynamic equicorrelation parameter. An important feature of the proposed multivariate GARCH model is the intrinsic robustness of the estimated correlation process to extreme observations and fat-tailed distributions, which are common in financial returns.
We apply both proposed specifications to daily returns from the U.S. stock market over a 16-year sample period. The resulting trajectory of the estimated correlation process appears to be broadly in line with both the realized sample correlations and the conditional correlations obtained from standard DCC--GARCH models, however, it exhibits noticeable local deviations. We attribute these differences to the improved robustness of the proposed approach. Although the estimation results indicate that the new model delivers a stable and tractable estimation procedure with sensible outcomes, a more thorough assessment of its statistical and economic performance is left for future research.
Let $x_{1}$ and $x_{2}$ are two real-valued random variables. Assume additionally that $x_{1}$ and $x_{2}$ have zero means, $E(x_{1})=E(x_{2})=0$, and the norm of random vector $x=(x_{1},x_{2})^{\prime}$ is positive with probability one. We consider the following variable,
which can be interpreted as a measure of statistical similarity between $x_{1}$ and $x_{2}$. For empirical analysis, this quantity was firstly introduced in Thorndike_1905 to measure statistical association between a pair of twins with respect to a set of considered characteristics, and was originally named resemblance. In that study, this term was used as a synonym for the coefficient of correlation.
Indeed, multiple aspects allow to consider the Thorndike's resemblance variable, $r$, as a measure of correlation. The quantity is confined within the fixed interval, $r\in[-1,1]$, and the value of $r$ is higher when $x_{1}$ and $x_{2}$ are more similar in magnitude, while sharing the same sign. The magnitude of $r$ is larger when magnitudes of $x_{1}$ and $x_{2}$ are more similar to each other, while the sign of $r$ is positive (negative) if $x_{1}$ and $x_{2}$ have the same (opposite) directions.
Variable $r$ has a particularly simple form when the random vector is represented in polar coordinates ($x_{1}=s\cos\theta$ and $x_{2}=s\sin\theta$). In this case, $r=\sin2\theta$, so the resemblance measure does not depend on the vector length, $s$. Intuitively, it implies that $r$ depends only on the angle between the observed vector and coordinate axes and ignores the magnitude of observations. This points on intrinsic insensitivity of the resemblance measure, $r$, to the presence of outliers in the data.
If we additionally assume that vector $x=(x_{1},x_{2})^{\prime}$ has finite second moments, we can define the linear correlation coefficient, or the Pearson correlation, between $x_{1}$ and $x_{2}$ as
where $\sigma_{1}^{2}=V(x_{1})=E(x_{1}^{2})$, $\sigma_{2}^{2}=V(x_{2})=E(x_{2})^{2}$ and $\sigma_{12}=E(x_{1}x_{2})$ are the second moments of $x$. The Pearson correlation is the most popular measure of statistical association that arises naturally in multiple econometric frameworks, from a simple linear regression to complex statistical learning algorithms. The Pearson correlation is often considered as the default benchmark for measuring dependence in empirical analysis. Although $\rho$ is able to capture the precise association between random variables only when the true relationship is linear, it can still provide a reasonable approximation when the underlying dependence is monotonic and not heavily non-linear.
While the correlation coefficient $\rho$ is bounded between $-1$ and $1$ by construction, it is sometimes convenient to work with an unconstrained correlation measure and this can be achieved by means of a suitable transformation. The most prominent example is the Fisher transformation defined as, \[ \phi_{\rho}=\frac{1}{2}\log\Bigl(\frac{1+\rho}{1-\rho}\Bigl), \] for $\rho\ensuremath{\in(-1,1)}$. The Fisher transformation represents a strictly monotone transformation of $\rho$ onto the set of real numbers, such that $\phi_{\rho}\in\mathbb{R}$. The transformation was proposed by Ronald A. Fisher in a series of seminal papers (see Fisher_1915, Fisher_1921), where he also demonstrated that it improves the distributional properties of the sample correlation coefficient.
In what follows, we will refer to the Thorndike's resemblance measure, $r$, under the Fisher transformation,
as to the measure of similarity, or the similarity variable. In contrast to $r$, the similarity measure $\phi_{r}$ has an unrestricted range, $\phi_{r}\in\mathbb{R}$, and its magnitude increases as the values of $x_{1}$ and $x_{2}$ become more similar. The sign of $\phi_{r}$ is positive (negative) when the two observations have the same (opposite) signs. Figure (ref) illustrates the magnitudes of $\phi_{r}$ as a function of $x_{1}$ and $x_{2}$, where warmer (red) colors indicate higher values of $\phi_{r}$ and cooler (blue) colors indicate lower values. We also note that, similar to $r$, $\phi_{r}$ depends only on the angular coordinate of the vector $x$, which underlies its intrinsic robustness to observations with extreme magnitudes.
An important special case arises when the random variables $x_{1}$ and $x_{2}$ follow a bivariate elliptical distribution, for which the dependence structure is inherently linear. In this situation, the Pearson correlation completely characterizes the statistical association between the variables. Assume additionally that variances of $x_{1}$ and $x_{2}$ are homogeneous, such that $\sigma_{1}^{2}=\sigma_{2}^{2}=\sigma^{2}$. Then the covariance matrix $\Sigma=V(x)$ reads
Under this assumption, $\phi_{r}$ and $\phi_{\rho}$ are elegantly connected. This result was originally formulated in Fisher_1919 for the Gaussian case, and below we provide an extension of the original result to the entire class of elliptical distributions.
Proposition (ref) implies that $\phi_{r}$ is symmetrically distributed, according to the hyperbolic secant distribution, around the Fisher transformation of $\rho$. The distribution does not depend on the variance parameter, $\sigma^{2}$, as well as on the underlying correlation coefficient, $\rho$, and remains invariant once $x=(x_{1},x_{2})^{\prime}$ belongs to the elliptical family. Therefore, under the assumptions of the proposition, the similarity variable $\phi_{r}$ provides an unbiased and robust signal of the latent correlation level (on the Fisher scale) with stable sampling properties, and thus emerges as an attractive statistical tool for correlation estimation.
Let a random sample is given, $\{x_{t}\}_{t=1}^{T}$, where $x_{t}=(x_{1,t},x_{2,t})^{\prime}$ are independent observations from some bivariate distribution with zero mean and finite second moments. Probably the most popular estimator of the Pearson correlation is the sample correlation estimator which is given by \[ \hat{\rho}=\frac{\sum_{t=1}^{T}x_{1,t}x_{2,t}}{\sqrt{\Bigl(\sum_{t=1}^{T}x_{1,t}^{2}\Bigl)\Bigl(\sum_{t=1}^{T}x_{2,t}^{2}\Bigl)}}, \] where we internalize that $\mathbb{E}x=0$. The sample correlation, $\hat{\rho}$, is a consistent estimator of the population Pearson correlation, and, under mild assumptions, it is asymptotically normal with $\sqrt{T}(\hat{\rho}-\rho)\overset{d}{\rightarrow}N\Bigl(0,V_{\rho}\Bigl)$. For relatively small $T$, however, the sampling properties of $\hat{\rho}$ are often poorly approximated by the asymptotic results, especially in the presence of sufficiently heavy-tailed data. The inference is additionally complicated by the fact that the asymptotic variance $V_{\rho}$ generally depends on the unknown value of $\rho$. For example, for the Gaussian case, $V_{\rho}=(1-\rho^{2})^{2}$ (see Fisher_1915).
The Fisher transformation is particularly useful in improving sampling properties of $\hat{\rho}$. We denote the Fisher transformation of $\hat{\rho}$ by $\phi_{\hat{\rho}}$. Thus, for normally distributed observations, the asymptotic distribution of $\phi_{\hat{\rho}}$ is $\sqrt{T}(\phi_{\hat{\rho}}-\phi_{\rho})\overset{d}{\rightarrow}N\Bigl(0,1\Bigl)$, and the asymptotic variance is independent of $\rho$. More importantly, the Fisher transformation offers multiple advantages in finite samples. In particular, when the data is close to be normally distributed, it provides the variance stabilization for the sampling distribution of $\phi_{\hat{\rho}}$ with making it symmetric and nearly Gaussian even for very small samples (see Fisher_1921, Hotelling_1953, etc.). These properties often motivate to analyze the sample correlation coefficient in the Fisher scale when conducting inference for correlations.
The critical drawback of the sample correlation estimator, that may compromise both estimation and inference, is its notorious sensitivity to outliers and, more generally, to extreme observations. A popular example of a robust correlation measure is the Kendall rank correlation coefficient, or Kendall\textquoteright s tau coefficient (esscher1924method, Kendall_1938), given by \[ \hat{\tau}=\frac{2}{T(T-1)}\sum_{i<j}\text{sign}(x_{1,i}-x_{2,j})\text{sign}(x_{2,i}-x_{2,j}). \] The measure captures the strength of monotonic dependence between two variables by quantifying the relative frequency of sign-concordant/discordant pairs of observations. For elliptical distributions, it can be transformed to match the Pearson correlation via the Greiner's equality which links the correlation coefficient with the quadrant probabilities, $\rho=\sin\Bigl(\frac{\pi}{2}E\hat{\tau}\Bigl)$, see Greiner_1909. A similar alternative is the class of quadrant estimators of correlation which are based on the sample proportion of sign-concordant observations (see Sheppard_1899, Kendall_1949, Blomqvist_1950, Hansen_Luo_2023, etc.)
The results in Proposition (ref) motivate an alternative estimator of statistical association that is based on the Fisher transformed similarity, $\phi_{r}$. For each observation $x_{t}$, $t=1,...,T$, we can construct the corresponding (local) empirical measure of similarity $\phi_{r,t}$, as given in (ref). In case $x_{t}$ are drawn from an elliptical distribution, and if the marginal variances of $x_{t}$ are homogeneous, it follows from Proposition (ref) that $\mathbb{E}\phi_{r,t}=\phi_{\rho}$. Therefore, $\phi_{r,t}$ is an unbiased and robust signal of the correlation coefficient, on the Fisher scale, and this naturally motivates suggesting the following moment-based estimator,
In what follows, we will refer to $\hat{\gamma}$ as to the similarity estimator. In the elliptical case with homogeneous variances, the similarity estimator is i) a consistent and unbiased estimator of the Pearson correlation coefficient (in the Fisher scale), ii) has robust mean and robust sampling distribution that does not depend on $\rho$, iii) the sampling distribution is available via the characteristic function, for any $T$, allowing for exact inference on correlations. In a more general scenario, when variances of $x_{1}$ and $x_{2}$ may differ, the estimator $\hat{\gamma}$ can be treated as a standalone, robust measure of statistical association between the two random variables, and it retains an additional interpretation as a lower bound on the correlation coefficient $\rho$.
We assume that a random sample of vectors $x_{t}=(x_{1,t},x_{2,t})^{\prime}$, for $t=1,...,T$, is independently drawn from some elliptical distribution with finite second moments, identical marginal variances, $\sigma_{1}=\sigma_{2}$, and the Pearson correlation coefficient, $\rho\in(-1,1)$. Under these assumptions, it directly follows from Proposition (ref) that
because the distribution of $f(\phi_{r,t})$ implies that the variance of $\phi_{r,t}$ is given by $V(\phi_{r,t})=\frac{\pi^{2}}{4}$.
The asymptotic distribution of $\hat{\gamma}$ does not depend on the actual correlation coefficient $\rho$. In contrast to $\phi_{\hat{\rho}}$, which represents the transformation of the sample correlation estimator $\hat{\rho}$, the similarity estimator, $\hat{\gamma}$, directly targets $\phi_{\rho}$ by estimating the correlation coefficient under the Fisher scale. The asymptotic variance of $\hat{\gamma}$ is $\frac{\pi^{2}}{4}$, and this value is, in general, larger than the asymptotic variance of $\phi_{\hat{\rho}}$ in the unconstrained scenario. The efficiency reduction is not surprising since $\hat{\gamma}$ incorporates only information about the relative magnitudes of $(x_{1,t},x_{2,t})$ and their signs, but ignores information about the total magnitude of $x_{t}$. A lower efficiency of $\hat{\gamma}$ is nonetheless compensated by its robustness to outliers and stability of the sampling distribution.
A remarkable feature of the similarity estimator in the considered scenario is that not only the asymptotic distribution, but also the finite sample distribution of $\hat{\gamma}-\phi_{\rho}$ is invariant over the entire class of elliptical distributions of $x_{t}$. Furthermore, the characteristic function of $\phi_{r,t}$ is available and allows to recover the exact sampling distribution of the similarity estimator for any finite $T$. Specifically, the distribution $f(\phi_{r,t})$ provided in Proposition (ref) implies that the characteristic function for $\phi_{r}-\phi_{\rho}$ is given by $\varphi_{\phi_{r}}(u)=\text{sech}\bigl(\frac{\pi}{2}u\bigl)$ for $|u|<1$. Denote the standardized similarity estimator by $z_{\hat{\gamma}}=\frac{2\sqrt{T}}{\pi}(\hat{\gamma}-\phi_{\rho})$. Then, the characteristic function of $z_{\hat{\gamma}}$ can be written as
for $|u|<1$. Therefore, the sampling distribution of $\hat{\gamma}$ becomes available in semi-explicit form (via the characteristic function) for any finite $T$, and this allows to conduct exact inference for the estimated correlation parameter. The sampling distribution of $z_{\hat{\gamma}}$ is symmetric with a positive excess kurtosis for any sample size. It converges to the standard normal distribution very quickly as $T$ increases. This is illustrated in Figure (ref) and in Table (ref), where the quantiles of the finite sample distribution are reported for a range of $T$.
It is worth mentioning that the properties of $\hat{\gamma}$ are generally retained for zero-centered elliptical distributions with infinite moments. For such distributions, the correlation coefficient is not possible to define using (ref), however, it can be defined alternatively via the quadrant probabilities (see Sheppard_1899, Greiner_1909, etc.) as $\rho=\sin\Bigl(\pi P-\frac{\pi}{2}\Bigl)=-\cos(\pi P)$, where $P$ denotes the probability of both $x_{1,t}$ and $x_{2,t}$ are of the same sign. Note that, for an elliptical random vector with finite second moments, the correlation defined in this way is equivalent to the Pearson correlation coefficient given by (ref). Thus, the similarity estimator remains a robust and consistent estimator of $\phi_{\rho}$ even for extremely heavy-tailed elliptical distributions, such as multivariate Cauchy distribution.
Figure (ref) provides small sample distributions of $\hat{\gamma}$ and $\phi_{\hat{\rho}}$ resulted from the simulation analysis. In this illustration we consider two selected sample sizes ($T=8$ and $T=40$) and three selected bivariate distributions of vector $x_{t}$: normal, $t$-distribution with 5 degrees of freedom, and Cauchy, for which the correlation coefficient $\rho$ is defined via quadrant probabilities. The figure shows an apparent instability of the sampling density of $\phi_{\hat{\rho}}$ across the considered elliptical distributions with different kurtosis parameters, thus, highlighting the sensitivity of the sample correlation estimator $\hat{\rho}$ to the presence of extreme observations. In contrast, the sampling density of $\hat{\gamma}$ is in line with the theoretically predicted results and is stable across all distribution specifications for both considered sample sizes.
The similarity estimator $\hat{\gamma}$ can be used as a robust and consistent estimator of the Pearson correlation coefficient once the inverse Fisher transformation is applied, $\phi^{-1}(\hat{\gamma})$, where \[ \phi^{-1}(\gamma)=\frac{e^{2\gamma}-1}{e^{2\gamma}+1}=\tanh(\gamma), \] which is a hyperbolic tangent function. We note that despite $\hat{\gamma}$ is an unbiased estimator of $\phi_{\rho}$, the inverse transformation, $\phi^{-1}(\hat{\gamma})$, is not an unbiased estimator of $\rho$ because $\mathbb{E}\phi^{-1}(\hat{\gamma})\neq\rho$, in general, for $T>1$. However, with an exact finite sample distribution of $\hat{\gamma}$ and due to monotonicity of the Fisher transformation, $\hat{\gamma}$ can be used for interval estimation of $\rho$ by providing exact confidence intervals for any sample size $T$.
Under the variance homogeneity assumption, another interesting feature of the similarity estimator is related to the matrix logarithm transformation. Assuming $x_{t}$ has a positive definite covariance matrix $\Sigma$ as in (ref), $\lambda_{+}=\sigma^{2}(1+\rho)$ and $\lambda_{-}=\sigma^{2}(1-\rho)$ are the two eigenvalues of $\Sigma$ with the corresponding eigenvectors $q_{+}=\frac{1}{\sqrt{2}}\iota_{2}$ and $q_{-}=\frac{1}{\sqrt{2}}\iota_{2}^{\perp}$, where $\iota_{2}=(1,1)^{\prime}$ and $\iota_{2}^{\perp}=(1,-1)^{\prime}$. For $n=2$, the matrix logarithm transformation of $\Sigma$ has an explicit analytic expression and reads \[ \log\Sigma=\left(
\right), \] where the off-diagonal entry, representing the (half) log-condition number of $\Sigma$, coincides with the Fisher transformation of the correlation coefficient, $\phi_{\rho}=\frac{1}{2}\log\Bigl(\frac{\lambda_{+}}{\lambda_{-}}\Bigl)$. Therefore, under the considered assumptions, $\hat{\gamma}=\frac{1}{2T}\sum_{t=1}^{T}\log\frac{(q_{+}^{\prime}x_{t})^{2}}{(q_{-}^{\prime}x_{t})^{2}}$ consistently estimates the off-diagonal element of $\log\Sigma$. This suggests a promising direction for using $\hat{\gamma}$ to estimate correlation matrices directly under the matrix logarithm transformation which offers many convenient properties for correlation analysis (see Archakov_Hansen_2021). We further explore this idea in Section (ref).
In a more general scenario, where the variables are not necessarily assumed to have identical variances, the interpretation of the similarity estimator is different. Let the covariance matrix of a random vector $x=(x_{1},x_{2})^{\prime}$ is positive definite and is given by
where $\sigma_{1}^{2}$ and $\sigma_{2}^{2}$ are not necessarily identical. Consider quantity $\xi=\frac{2\sigma_{12}}{\sigma_{1}^{2}+\sigma_{2}^{2}}$, which we refer to as the coefficient of resemblance. This coefficient is closely related to the Pearson correlation coefficient, $\rho$, and can be interpreted as an alternative measure of statistical association between random variables. More particularly, $\rho$ and $\xi$ are proportionally related,
and the coefficient of proportionality depends only on the relative size of variances, $\sigma_{1}^{2}/\sigma_{2}^{2}$.
Note that the signs of $\xi$ and $\rho$ are always identical, while the magnitude of $\xi$ never exceeds the magnitude of $\rho$, i.e. $|\xi|\leq|\rho|$, because $\frac{2\sigma_{1}\sigma_{2}}{\sigma_{1}^{2}+\sigma_{2}^{2}}\leq1$ due to the Cauchy-Schwartz inequality. It then follows that $\xi$ is also constrained between $-1$ and $1$, and attains the limits in case of perfect correlation ($\rho=\pm1$) and variance homogeneity ($\sigma_{1}^{2}=\sigma_{2}^{2}$). In other words, the coefficient of resemblance, $\xi$, can be interpreted as a measure of statistical similarity between the random variables, where both the correlation and scale are taken into consideration. Namely, $\xi$ is higher when the variables tend to exhibit more similar correlation components, along with more similar variance components (magnitudes).
If $x$ is an elliptical random vector, an important feature of $\xi$ is that, under the Fisher transformation, it becomes the mean of the transformed resemblance measure, $\phi_{r}=\frac{1}{2}\log\Bigl(\frac{1+r}{1-r}\Bigl)$, where $r$ introduced in (ref). This property is reflected in the following proposition.
It is important to mention that $\phi_{r}$, as well as its distribution and moments, retain robustness to outliers due to intrinsic insensitivity of $r$ to the total magnitude of $x$. In contrast to the homoskedastic case ($\sigma_{1}^{2}=\sigma_{2}^{2}$), the variance of $\phi_{r}$ does depend on the underlying covariance matrix of $x$. However, for any positive definite $\Sigma$, we have that $V_{\phi_{r}}\leq\frac{\pi^{2}}{4}$, with the equality holds only for $\sigma_{1}^{2}=\sigma_{2}^{2}$. Therefore, $V_{\phi_{r}}$ reaches its maximum value under the variance homogeneity, and has a lower value otherwise.
The results in Proposition (ref) allow to generalize the asymptotic properties of the similarity estimator $\hat{\gamma}$ for the case of potentially heterogeneous variances. We assume a sample of random vectors $x_{t}=(x_{1,t},x_{2,t})^{\prime}$, for $t=1,...,T$, are independent and elliptically distributed with a non-singular covariance matrix $\Sigma$. Then $\hat{\gamma}$ has the limit distribution
where $\phi_{\xi}$ is the Fisher transformation of $\xi$ and $V_{\phi_{r}}$ is provided in Proposition (ref). As a result, $\hat{\gamma}$ is a consistent and robust estimator of the resemblance coefficient (on the Fisher scale), and its sampling distribution remains stable across the wide class of elliptical densities for the sample observations.
Since the resemblance coefficient $\xi$ represents a downward scaled version of the Pearson correlation $\rho$, with the scaling coefficient is provided in (ref), $\hat{\gamma}$, in general, does not directly estimates $\phi_{\rho}$. When the aim is to estimate the correlation coefficient, the variables have to be standardized such that the variances of $x_{1,t}$ and $x_{2,t}$ become homogeneous, and thus $\phi_{\xi}=\phi_{\rho}$, as shown in Section (ref). For example, this can be done via a two-step approach. In the first step, the individual variances $\hat{\sigma}_{1}^{2}$ and $\hat{\sigma}_{2}^{2}$ of $x_{1}$ and $x_{2}$, respectively, are estimated, and the standardized observations are constructed, $z_{t}=\Bigl(\frac{x_{1,t}}{\hat{\sigma}_{1}},\frac{x_{2,t}}{\hat{\sigma}_{2}}\Bigl)^{\prime}$. In the second step, the similarity estimator is applied to the standardized variables $z_{t}$. If $\hat{\sigma}_{1}^{2}$ and $\hat{\sigma}_{2}^{2}$ consistently (and robustly) estimate the corresponding variances, $\hat{\gamma}$ becomes a consistent estimator of $\phi_{\rho}$ with the asymptotic distribution given in (ref), however, the finite sample distribution of $\hat{\gamma}$, in general, will differ from the one characterized in Section (ref). Alternatively, the information about marginal volatilities, $\sigma_{1}^{2}$ and $\sigma_{2}^{2}$, can be inferred from dynamic filters, such as GARCH models.
In case variance homogeneity is not ensured, the similarity estimator $\hat{\gamma}$ can be interpreted as a consistent lower bound estimator for the magnitude of $\phi_{\rho}$ due to $|\phi_{\xi}|\leq|\phi_{\rho}|$. Moreover, the asymptotic distribution in (ref) can be used to construct conservative confidence intervals for $\phi_{\xi}$ (and $\xi$) because $V_{\phi_{r}}\leq\frac{\pi^{2}}{4}$. In this situation, estimator $\hat{\gamma}$ can be used for conservative, but theoretically justified, estimation and inference for the latent correlation coefficient. This aspect can be useful, for example, for testing whether the correlation coefficient is equal to zero.
The quantity $\phi_{r,t}$ formulated in (ref) can be represented as a log-ratio of the local variations of vector $x_{t}$ -- variation of the sum and variation of the difference of the vector components, where $x_{t}=(x_{1,t},x_{2,t})^{\prime}$ is a zero mean random vector with finite second moments. For fixed $\sigma_{1}^{2}$ and $\sigma_{2}^{2}$, an increase in $\rho$ makes the first variation higher and the second variation lower, and conversely. Naturally, for $\rho\approx\pm1$ the contrast between the variations is the highest, while for $\rho=0$ the two variations are identical. Therefore, $\phi_{r,t}$ can be viewed as a local measure of the divergence between $V(x_{1,t}+x_{2,t})$ and $V(x_{1,t}-x_{2,t})$ which nests information about the correlation coefficient $\rho$, and this information can be explicitly recovered once $\sigma_{1}^{2}=\sigma_{2}^{2}$ (see Section (ref)).
We note that $V(x_{1,t}+x_{2,t})$ can alternatively be represented as the variance of the projection of $x_{t}$ onto the direction of the vector of ones, i.e. the direction of perfect similarity. Analogously, $V(x_{1,t}-x_{2,t})$ is the variance of the projection of $x_{t}$ onto the orthogonal direction, i.e. the direction of dissimilarity. Therefore, $\phi_{r,t}$, and consequently $\hat{\gamma}$, indicate the degree of similarity in the joint variation of $x_{1,t}$ and $x_{2,t}$ measured by the relative magnitude of variation along the vector of ones. This idea can be directly generalized to random vectors of arbitrarily large dimension.
Assume that $x_{t}=(x_{1,t},x_{2,t},...,x_{n,t})^{\prime}$ is a zero mean $n$-dimensional random vector with finite second moments and positive definite covariance matrix $\Sigma$. We denote the $n$-dimensional vector of ones by $\iota_{n}=(1,1,...,1)^{\prime}$. The matrix $P_{n}=\frac{1}{n}\iota_{n}\iota_{n}^{\prime}$ is the projection matrix such that the projection of $x_{t}$ onto the direction defined by the vector $\iota_{n}$ is given by $P_{n}x_{t}$. Then, $P_{n}^{\perp}=I_{n}-P_{n}$, where $I_{n}$ is the identity matrix of dimension $n$, is also the projection matrix such that $P_{n}^{\perp}x_{t}$ represents the projection of $x_{t}$ onto the subspace orthogonal to $\iota_{n}$. Then, the local similarity measure $\phi_{r,t}$ can be naturally extended to the general multivariate case via
which captures the variation of $x_{t}$ along the vector of perfect similarity, $\iota_{n}$, relative to the variation of $x_{t}$ along its orthogonal complement. Note that, for the special case $n=2$, equation (ref) is reduced to the same expression for $\phi_{r,t}$ as in (ref).
As in the bivariate case, the similarity estimator $\hat{\gamma}$ is constructed as $\hat{\gamma}=\frac{1}{T}\sum_{t=1}^{T}\phi_{r,t}$. Intuitively, $\gamma$ measures the average similarity in how all $n$ variables move, in terms of both magnitude and direction.
In the multivariate scenario with $n>2$, the similarity estimator $\hat{\gamma}$ has, in general, less tractable interpretation as compared to the bivariate setting. In some special cases, however, $\hat{\gamma}$ retains explicit finite sample and asymptotic limits, and is immediately related to the Pearson correlation coefficient.
Let the entries of covariance matrix $\Sigma$ are denoted by \[ \Sigma=\left(
\right), \] and here we assume that all variances are homogeneous, i.e. $\sigma_{1}^{2}=\sigma_{2}^{2}=...=\sigma_{n}^{2}=\sigma^{2}$, and all covariance (off-diagonal) elements are identical too. This implies that all pairwise correlations are also identical, and, so, we can write $\sigma_{ij}=\sigma^{2}\rho$ for all $i$ and $j$, such that $i\neq j$, where $\rho$ is the common correlation parameter. Note that, once $\Sigma$ is positive definite by assumption, it implies that $\rho\in\bigl(-\frac{1}{n-1},1\bigl)$.
The matrix logarithm transformation suggests a convenient method for reparametrization of correlation matrices in such way that the transformed correlation elements can be represented as an unconstrained real vector which can always be mapped to a unique positive definite correlation matrix. In this sense, the parametrization based on the matrix logarithm can be viewed as a multi-dimensional generalization of the Fisher transformation, see Archakov_Hansen_2021 for more details. When the matrix logarithm transformation is applied to $\Sigma$ with homogeneous variances and equal correlations, the general structure of the matrix is preserved after the transformation, i.e. all diagonal and all off-diagonal elements remain identical,
{5pt} \[ \Sigma=\sigma^{2}\left(
\right),\qquad\log\Sigma=\left(
\right), \] as it is shown in Archakov_Hansen_2024. The eigenvalues of matrix $\Sigma$ are $\lambda_{+}=\sigma^{2}\bigl(1+(n-1)\rho\bigl)$ and $\lambda_{-}=\sigma^{2}(1-\rho)$ with multiplicities $1$ and $n-1$, respectively, see Olkin_Pratt_1958. The entries of $\log\Sigma$ are analytically available and given by $\delta=\frac{1}{n}\log\lambda_{+}+\frac{n-1}{n}\log\lambda_{-}$ and
We note that the off-diagonal element $\phi_{\rho}$ depends only on $\rho$ and $n$, and does not depend on variance $\sigma^{2}$. Naturally, this result is a generalization of the matrix logarithm result for $n=2$ in Section (ref), so we preserve notation $\phi_{\rho}$ introduced earlier.
If we additionally assume that random vector $x_{t}$ is elliptical, the distribution of $\phi_{r,t}$ in (ref) can be derived explicitly which is reflected in the following result.
Note that $\phi_{r}$ is not an unbiased measure of $\phi_{\rho}$ because $E(\phi_{r})=\phi_{\rho}-\omega_{n}$, where $\omega_{n}\ge0$ is a non-monotone function of $n$ such that $\omega_{2}=0$ and $\omega_{n}\rightarrow0$ as $n\rightarrow\infty$. The exact expression for $\omega_{n}$ is provided in Appendix. Variance of $\phi_{\rho}$ is inversely proportional to $n^{2}$ and is given by \[ V(\phi_{r})=\frac{1}{n^{2}}\biggl(\psi^{\prime}\Bigl(\frac{n-1}{2}\Bigl)+\frac{\pi^{2}}{2}\biggl), \] where $\psi^{\prime}(t)$ is the trigamma function. Naturally, when vector dimension $n$ increases, $\phi_{r}$ absorbs more information about the common correlation coefficient from the cross-sectional dimension and, so, provides an increasingly more efficient signal about the latent correlation level. In Figure (ref), we illustrate the Logistic-Beta probability density function of $\phi_{r}-\phi_{\rho}$ for several selected values of $n$.
In the special case of homogeneous variances and correlations, Proposition (ref) helps to characterize the asymptotic properties of the similarity estimator $\hat{\gamma}$. With a bias correction, $\hat{\gamma}+\omega_{n}$ becomes a consistent and robust estimator for the off-diagonal element $\phi_{\rho}$ of the log-transformed covariance matrix $\Sigma$, where $\phi_{\rho}$ represents a monotone transformation of the equicorrelation coefficient $\rho$, similarly to the Fisher transformation function in the bivariate case. The asymptotic variance of $\hat{\gamma}$ is given by $V(\phi_{r})$, and the estimator becomes more efficient with dimension $n$ due to the growth of relevant cross-sectional information.
Despite the structure assumed for $\Sigma$ is restrictive and does not occur in practice frequently, the results reveal several qualitative aspects which are characteristic of the similarity estimator, $\hat{\gamma}$, in more general and realistic scenarios. At first, $\hat{\gamma}$ is a robust estimator of an aggregate measure of association between considered variables. Although this does not directly correspond to the average correlation level for an arbitrary covariance matrix $\Sigma$, it still can be interpreted as an indicator of joint statistical similarity between variables, i.e. how similar they are in terms of both magnitude and direction. At second, if elements in $\Sigma$ are sufficiently homogeneous, $\hat{\gamma}$ is expected to be more efficient once $n$ gets larger. This is because each additional variable provides some extra information about the average level of similarity.
Although the similarity estimator can be used as a standalone measure of statistical association, in this section, we consider empirical applications where the latter is used for estimation and modeling the Pearson correlation coefficient. The two considered applications utilize financial market data for measuring correlations between stock returns.
We apply the similarity estimator for robust interval estimation of financial correlations. For this analysis, we use intraday transaction data from the TAQ database cleaned according to the recommendations provided in Barndorff-Nielsen2009. For exposition purposes we consider daily correlations between Apple (AAPL) and Exxon Mobil (XOM) returns during the period between February and April 2020 (62 trading days) which corresponds to the COVID-19 outbreak.
We note that AAPL and XOM belong to different sectors of the economy (Tech vs Energy) which implies that the correlation between these stocks is not supposed to be particularly strong in ordinary periods of time. During the COVID-19 crisis, however, the overall correlation level has significantly elevated across the entire market due to i) an initial decline of the market that affected almost all sectors and industries (largely began on February 19), and ii) the subsequent common recovery (began after March 23). Therefore, we may expect notable changes in the latent correlation trajectory for the considered assets during the analyzed time period.
For each considered trading day, we construct samples of intraday (logarithmic) returns, $r_{t}=(r_{1,t},r_{2,t})^{\prime}$, $t=1,...,T$, for a range of selected frequencies by using the previous tick interpolation scheme which was introduced in Wasserfallen_Zimmermann_1985 and is used routinely in high frequency econometrics (see, for example, Hansen_Lunde_2006). Due to the duration of a typical trading day is $6.5$ hours, we obtain a sample with $T=78$ observations for the frequency $\Delta=5$ min, while we have $T=390$ observations for $\Delta=1$ min.
To construct an interval with a specified coverage probability, we estimate the correlation using the similarity estimator $\hat{\gamma}$ on standardized intraday returns, $z_{t}=(z_{1,t},z_{2,t})^{\prime}$, where $z_{i,t}=r_{i,t}/\hat{\sigma}_{i}$ with $\hat{\sigma}_{i}$ is a sample standard deviation of $r_{i,t}$. Next, we obtain the required interval around $\hat{\gamma}$ using the exact quantiles\footnote{The results only marginally differ in case the intervals are constructed using the asymptotic distribution of $\hat{\gamma}$ given in (ref). We nonetheless report intervals based on exact quantiles as they provide more conservative assessment of the sampling error.} of the sampling distribution characterized in (ref), and map the interval from the Fisher correlation scale to the Pearson correlation scale by using the inverse Fisher transformation. The resulting intervals for daily correlations are obtained for 90% and 95% coverage probabilities for the three selected intraday frequencies ($\Delta=10$, $5$, and $1$ min) and are shown in Figure (ref). It also provides point estimates of correlation calculated with two benchmark estimators -- the sample (realized) correlation estimator and the Kendall rank coefficient (see Section (ref) for the corresponding expressions). The results in Figure (ref) reveal several interesting aspects.
When constructed with returns sampled at $\Delta=10$ min frequency, the intervals are very wide and thus not informative about the underlying correlation level. For the majority of trading days, the intervals do not even allow to reject the null hypothesis of zero correlation. Such conservative interval widths reflect the trade-off required to achieve robustness and invariance of the constructed intervals under arbitrary fat-tailed data. At higher frequencies, where more observations are available, intervals naturally shrink, and variability of the correlation level over time becomes more apparent. Thus, when constructed at $\Delta=1$ min frequency, the intervals are sufficiently narrow allowing to clearly observe a sharp surge of the correlation level in the second half of February and gradual non-monotone subsequent decay. Interestingly, the estimated intervals appear to cluster within weeks, which may indicate highly persistent correlation dynamics.
It is important to mention that since we standardize raw intraday returns by using sample standard deviations, which is not a robust estimator of dispersion, the resulting standardized returns do not necessarily have homogeneous variances. Recall that under non-perfectly homogeneous variances, the similarity estimator underestimates the true correlation level, as it is discussed in Section (ref). Despite this, the constructed robust intervals based on the similarity measure show good agreement with both benchmark estimators for almost all trading days even at the highest considered frequency, where the intervals are sufficiently narrow. For a number of trading days, however, we observe that the sample correlations fall below the constructed robust intervals as well as below the estimates obtained with the Kendall estimator. This can be explained by the common presence of price jumps in the market data inducing outliers in high-frequency returns. While the realized correlation estimator exhibits a downward bias in such cases, the Kendall and similarity estimators remain robust.
Figure (ref) illustrates intraday returns on AAPL and XOM sampled at 1-min frequency on February 27 and 28. Despite $\hat{\gamma}$ indicates similar correlation levels for both days, the realized correlation and the Kendall coefficient show sufficiently lower correlation on February 27, and align with $\hat{\gamma}$ on February 28. As we may see from the scatter plot, on February 27 we observe a sufficient number of potentially outlying observations distributed in all directions around a more compact and regularly shaped core of the sample. This may drive the benchmark measures downwards and cause disagreement between the estimators. In contrast, on February 28 outliers appear to be more directionally aligned with the core part of the sample. This results in a more regular elliptical shape for the sample data, with all estimators largely agreeing on the estimated correlation value.
We emphasize that the presented analysis serves rather for experimental and illustration purposes. The assumption of independent and identically distributed observations can hardly be justified due to stochastic volatility, market microstructural effects, and other stylized artifacts which are typically attributed to high frequency returns. Moreover, we may expect that the correlation level change over a trading day, so the obtained estimates should rather be interpreted as indicators of average daily correlation. A more careful adaptation of the similarity estimator for high frequency financial data is a challenging and interesting question which is a subject of ongoing work.
We suggest a new specification for the multivariate GARCH model that incorporates the similarity measures introduced in the paper. Let $r_{t}=(r_{1,t},r_{2,t},...,r_{n,t})^{\prime}$ be $n$-dimensional vector of asset returns, for $n\geq2$, observed at discrete time moments $t=1,...,T$. The key object of interest is the conditional covariance matrix of $r_{t}$, which we denote by $H_{t}=V(r_{t}|\mathcal{F}_{t-1})$, where $\{\mathcal{F}_{t}\}$ is the natural filtration for $r_{t}$. We follow the logic of the Dynamic Conditional Correlation (DCC) approach introduced in Engle_2002, and decompose the $H_{t}$ into variance and correlation components, \[ H_{t}=\Lambda_{h_{t}}^{1/2}C_{t}\Lambda_{h_{t}}^{1/2}, \] where $\Lambda_{h_{t}}=\text{diag}(h_{1,t},h_{2,t},...,h_{n,t})^{\prime}$ with $h_{i,t}$ is the conditional variance of an individual asset return $i$, such that $h_{i,t}=[H_{t}]_{ii}$, for $i=1,...,n$, and $C_{t}=\text{corr}(r_{t}|\mathcal{F}_{t-1})$ is the positive-definite conditional correlation matrix of $r_{t}$. This structure allows to effectively split the modeling of $H_{t}$ into separate modeling of conditional variances and correlations.
We assume that all conditional means are constant and denote them by $\mu_{i}=E(r_{i,t}|\mathcal{F}_{t-1})$, for $i=1,...,n$, which is a standard assumption in the GARCH literature. Then, we can formulate return equations as follows,
where $z_{i,t}$ are standardized returns, such that $E(z_{i,t}|\mathcal{F}_{t-1})=0$ and $V(z_{i,t}|\mathcal{F}_{t-1})=1$. Conditional variances $h_{i,t}$, for $i=1,...,T$, can be modeled with any appropriate univariate dynamic GARCH equation. For example, EGARCH(1,1) specification by Nelson_1991 can be used,
where $g_{i}(z_{i,t-1})\in\mathcal{F}_{t-1}$ and, so, it provides an observation-driven update for $\log h_{i,t}$.\footnote{The classification of time-varying parameter models into the classes of observation-driven and parameter-driven models goes back to Cox_1981. See also an instructive discussion about observation-driven modeling in Koopman_Lucas_Scharth_2016.} Denote the vector of standardized returns at period $t$ by $z_{t}=(z_{1,t},z_{2,t},...,z_{n,t})^{\prime}$, and note that $\text{corr}(z_{t}|\mathcal{F}_{t-1})=C_{t}$. Therefore, $z_{t}$ carries the information about the conditional correlation structure of raw asset returns.
The methodological novelty of the suggested model is related to the way how the dynamics for conditional correlation matrix $C_{t}$ is modeled. Following the approach in Archakov_Hansen_Lunde_2025, we set up the dynamics for the off-diagonal elements of the log-transformed conditional correlation matrix, $\log C_{t}$. The novel idea is to use the similarity measure $\phi_{r,t}$, characterized in Sections (ref) and (ref), as a natural signal about the current level of $\log C_{t}$. Since the conditional variances of standardized returns in $z_{t}$ are homogeneous, and assuming the return vector follows an elliptical distribution, we have that $\phi_{r,t}$, constructed with $z_{t}$, represents a robust empirical measure of the transformed correlation coefficients, as it was shown in Propositions (ref) and (ref) for the bivariate and the equicorrelation structures, respectively.
We begin with the bivariate case ($n=2$), where $\log C_{t}$ can be fully characterized by only a single correlation parameter, $\phi_{\rho,t}$, which is the Fisher transformation of the single correlation parameter $\rho_{t}$ in $C_{t}$. We consider the following dynamic specification for $\phi_{\rho,t}$,
where $\phi_{r,t}$ is a local similarity measure defined in (ref). The structure of this recursive equation is analogous to the classical GARCH equation for volatility. The autoregressive term $\beta\cdot\phi_{\rho,t-1}$ accounts for persistence in $\phi_{\rho,t}$, while term $\kappa\cdot\phi_{r,t-1}$ represents an observation-driven innovation for the correlation dynamics. Note that the vector of standardized returns $z_{t}=(z_{1,t},z_{2,t})^{\prime}$ is elliptical, by assumption, and have homogeneous unit variances. Therefore, by Proposition (ref), $\phi_{r,t}$ becomes an unbiased signal about the latent correlation, $\phi_{\rho,t}$, on the Fisher scale. At the same time, $\phi_{r,t}$ is robust to the presence of outliers and fat-tailed distributions of the observed returns. This is a particularly useful feature in modeling financial returns, where the fat tails and jumps are commonly found among stylized empirical regularities.
An apparent advantage of the model suggested in (ref) is an unconstrained support of $\phi_{\rho,t}$ which allows to avoid any additional restrictions on the dynamic process for ensuring positive definiteness of $C_{t}$. Recall that, in the classical DCC-GARCH models, the conditional correlation dynamics is specified through the matrix process, such that the resulting matrix needs an extra adjustment for positive definiteness and for the estimated conditional correlations remain consistent (see also Aielli_2013, Brownlees_Llorens_2023).
The specification in (ref) is closely related to the angular DCC-GARCH model suggested in Jarjour_Chan_2020, where the conditional correlation dynamics was specified for non-transformed $\rho_{t}$, and $r_{t}=\frac{2z_{1,t}z_{2,t}}{z_{1,t}^{2}+z_{2,t}^{2}}$ was used as a dynamic innovation term. Although $r_{t}$ is a robust signal of correlation, it is not an unbiased measure of $\rho_{t}$ because, in general, $E(r_{t}|\mathcal{F}_{t-1})\neq\rho_{t}$. Furthermore, extra restrictions for the parameter coefficients must be imposed to ensure positive definiteness. In light of this, the specification in (ref) appears to be a more natural and convenient approach to modeling dynamic correlations.
Since daily returns are typically fat-tailed, it is popular to model the conditional distribution of $r_{t}$ by some elliptical distribution that allows for excess kurtosis, such as the Student's $t$-distribution, the symmetric multivariate stable distribution, etc.\footnote{Note that the assumption of elliptical distribution for returns requires, in general, to introduce an additional parameter for the stochastic radial component which controls tail behavior.} Once the distribution for $r_{t}$ is specified, the model given by (ref)-(ref) can be estimated using the Maximum Likelihood method. We note that many extensions, which are common for the DCC-GARCH models, are also readily available for our model. For example, these extensions may include correlation targeting, an increased number of lags in (ref), or incorporating realized measures of correlation in spirit of Archakov_Hansen_Lunde_2025.
In Figure (ref), we provide an example of the estimated robust conditional correlation trajectory. For this example, we analyze the correlation between Chevron Corp. (CVX) and Marathon Oil Corporation (MRO) for the 16-year sample period between 2005 and 2020. We use close-to-close daily returns, adjusted for stock splits and dividends, from the CRSP US Stock Database. While the estimated trajectory (solid blue line) has an apparent time-varying dynamics, we observe persistently high correlations over the entire sample period, which is not a surprising empirical evidence due to both companies belong to the energy sector. It is interesting that the estimated correlations are systematically higher than the daily realized (sample) correlations calculated with 5-min intraday returns (gray line). A possible explanation is that the (non-robust) realized correlations are often biased towards zero due to presence of price jumps and extreme returns.
Based on results in Section (ref), we can specify an elegant and parsimonious multivariate GARCH specification for arbitrary large dimension $n$. For this, we assume an equicorrelation structure for $C_{t}$, such that all conditional correlations are identical and parametrized by a single dynamic coefficient, $\rho_{t}$. The DECO-GARCH model was originally introduced in Engle_Kelly_2012, and was extended to accommodate realized measures of variances and correlations in Archakov_Hansen_Lunde_2025. We formulate the dynamic correlation process for the off-diagonal elements of the log-matrix transformation. In case $C_{t}$ has the equicorrelation structure, the off-diagonal entries of $\log C_{t}$ are all identical and equal to $\phi_{\rho,t}$, and the analytical relationship between $\phi_{\rho,t}$ and $\rho_{t}$ is provided in (ref). We formulate dynamics for $\phi_{\rho,t}$ as follows,
where $\phi_{r,t}$ is the local similarity measure introduced in Section (ref), $P_{n}=\frac{1}{n}\iota_{n}\iota_{n}^{\prime}$ and $P_{n}^{\perp}=I_{n}-P_{n}$ are the orthogonal projection matrices, $i_{n}$ is the $n$-dimensional vector of ones, and $I_{n}$ is the n-dimensional identity matrix. This specification effectively retains all the benefits of specification (ref) formulated for the bivariate case. Namely, $\phi_{r,t}$ is intrinsically robust to extreme observation and is a relevant signal of $\phi_{\rho,t}$ with a known constant bias term, $E(\phi_{r,t}|\mathcal{F}_{t-1})=\phi_{\rho,t}+\omega_{n}$. Such fixed bias is harmless as it is absorbed by the constant coefficient $\alpha$ in (ref). Unconstrained range of $\phi_{\rho,t}$ implies that no extra restrictions are needed in order to ensure that estimated matrices $C_{t}$ are positive definite.
The equicorrelation assumption imposes a tight constraint on $C_{t}$, which is rarely empirically plausible. This structure, however, provides a substantial dimension reduction since, instead of modeling $\frac{n(n-1)}{2}$ dynamic correlations, we effectively model only a single common correlation coefficient. The estimated dynamics can be interpreted as a time-varying average correlation level, or an index of market co-movement, and can serve as a useful state variable in various contexts. For example, when a sufficiently representative sample of assets is available, $\phi_{\rho,t}$ can be used as a barometer of diversification potential for portfolio management purposes, or as an aggregate correlation index which can be informative about overall market uncertainty and risk.
As in the previous case, the model can be estimated using the Maximum Likelihood. An important advantage of the DCC structure for multivariate GARCH is that the model can be effectively estimated in two steps. In the first step, the univariate GARCH models (ref)-(ref) are estimated individually for $i=1,...,n$, and the standardized returns, $\hat{z}_{t}$, are obtained. In the second step, the conditional correlation dynamics of $C_{t}$, given in (ref), is estimated by using $\hat{z}_{t}$ obtained in the first step. Such two-stage estimation can facilitate computation complexity dramatically, especially if high dimensional return vectors are modeled.
For illustration purposes, we estimate the new model for the sample of nine stocks using the daily close-to-close returns, adjusted for stock splits and dividends, from the CRSP US Stock Database. In particular, we consider three stocks from the energy sector (CVX, MRO, and OXY), three stocks from the health care sector (JNJ, LLY, and MRK), and three stock from the information technology sector (AAPL, MU, and ORCL), and estimate the model for the 16-year sample period between 2005 and 2020. The estimation results are illustrated in Figure (ref). The conditional equicorrelation index estimated with the new robust model is shown with a blue line, while the index estimated with a standard DECO-GARCH model by Engle_Kelly_2012 is shown in red. We observe that both indices exhibit a significant amount of variation and often display visible reactions around major historical events. Although the correlations estimated by the two models follow similar trajectories, numerous local divergences are evident over the sample period, which can be attributed to the intrinsic robustness of the newly proposed model. For example, we observe that the new correlation index exhibits a more modest response during the European Debt Crisis in 2011 and the outbreak of COVID-19, periods in which many co-directional extreme returns were observed. This evidence points to distinctive informational content in the new robust correlation index, and its evaluation suggests a promising direction for further empirical analysis.
In this paper, we suggest a new measure of statistical association which captures the similarity between multiple random variables and is grounded in the ideas of Thorndike_1905 and Fisher_1919. The similarity is defined as the relative extent of joint variation along the direction of the vector of ones, which can be viewed as the direction of perfect similarity for random outcomes, measured in both sign and magnitude. The suggested similarity estimator is intrinsically insensitive to extreme observations, caused by fat-tailed data and outliers, and this ensures robustness for the entire sampling distribution of the estimator.
We analyze statistical properties of the similarity estimator for the entire class of elliptical random vectors. In the bivariate case and under assumption of variance homogeneity, we demonstrate that the similarity estimator is a consistent and robust estimator of the Pearson correlation coefficient (on the Fisher scale). The robustness of the estimator also emerges in exact and analytically available finite sample distribution. For example, the similarity estimator can be used for construction of exact confidence intervals for correlations in the presence of noisy and/or heavy-tailed data as well as for many other robust inference applications. In case variances are not identical, the similarity estimator preserves its connection to the Pearson correlation coefficient by becoming the estimator of its lower bound, and thus can be useful for conservative estimation and inference. When the estimator is applied to multiple (more than two) variables, it can be interpreted as a robust estimator of an aggregate correlation level among the considered variables, while in the special case, when variances and correlations are homogeneous, it is explicitly related to the correlation coefficient.
An intrinsic robustness of the similarity estimator comes at the expense of lower efficiency. This motivates to consider efficiency improving modifications of the estimator based, for example, on sub-sampling methods. This research direction has particularly high potential in the context of intraday financial data, where a substantial amount of observations can be additionally exploited at high frequencies. Another potentially useful methodological contribution lies in the development of new methods for composite estimation of large-scale correlation matrices. For instance, the correlation elements can be estimated separately using the bivariate similarity estimator, and then the resulting matrix is projected to the sub-space of a positive definite correlation matrices according to some appropriate criterion of optimality. We reserve these topics for future work.
We illustrate the empirical performance of the similarity estimator by applying it to intraday stock returns. The robust confidence intervals constructed by using the new method show strong agreement with widely used robust alternatives to the sample correlation estimator. This evidence suggests that the similarity estimator can be a reliable tool for estimation and inference of financial correlations with high-frequency data. It migth be especially useful when dealing with particularly noisy asset classes, such as crypto-currencies.
As an econometric application, we develop a novel robust multivariate GARCH model in which the conditional correlation process is modeled using the matrix logarithm transformation, and its dynamics are driven by the similarity measure. The suggested specification naturally retains positive definiteness for the filtered correlation structure and ensures robustness in the presence of fat tails and outliers in the data. A straightforward extension would be to accommodate robust realized measures of correlation calculated using the similarity estimator with intraday data in spirit of Archakov_Hansen_Lunde_2025. Furthermore, the similarity estimator can be also appropriate for modeling dynamic correlations with the score-driven approach suggested in CKL13, or with the parameter driven models such as state-space or stochastic volatility models. We leave these avenues for future research.