EconBase
← Back to paper

Correlation Matrices in High Dimensions: The Elliptope as a Sample-Correlation Ensemble

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

64,387 characters

Correlation Matrices in High Dimensions: The Elliptope as a Sample-Correlation Ensemble


\maketitle

\begin{abstract}
\noindent The set of $n\times n$ correlation matrices, the elliptope $\mathcal{E}_n$, has volume decaying at the super-exponential rate $\exp\{-\tfrac14 n^2\log n\}$. We characterize where this vanishing volume concentrates.
A uniform draw is entrywise close to the identity yet globally far from it and nearly singular: its maximum absolute correlation is of order $\sqrt{\log n/n}$, its Frobenius distance is asymptotic to $\sqrt n$, its empirical spectral distribution converges to the Marchenko--Pastur law with ratio one, and its smallest eigenvalue has the exact $\operatorname{Beta}(1,d)$ distribution, where $d=n(n-1)/2$, and is therefore of order $n^{-2}$.
More generally, distinct off-diagonal entries are exactly pairwise independent under every $\operatorname{LKJ}(\eta)$ law. For the uniform law, this yields a Chen--Stein proof of the extreme-correlation point-process limit and an $O(n^{-1})$ total-variation bound for finite-dimensional exceedance counts relative to Poisson laws with their exact finite-$n$ means. We also identify two distinct scales: $\eta_n\asymp n$ alters the limiting spectrum, whereas $\eta_n\asymp n^2$ is needed to keep the Frobenius distance bounded. Finally, for a bounded, centered i.i.d.\ off-diagonal specification, projection to the nearest correlation matrix incurs a squared repair cost asymptotically at least one-half of the squared Frobenius norm of its off-diagonal part.

\medskip
\noindent\textbf{Keywords:} correlation matrices; correlation networks; elliptope; extreme-value theory; LKJ; Marchenko--Pastur; Poisson point process.\\
\textbf{MSC2020 subject classifications:} Primary 60B20, 62H10; secondary 15B48, 60G70.
\end{abstract}

\section{Introduction}

Correlation matrices play a central role in multivariate statistics, econometrics, and applied probability.
Despite their importance, the geometry of the set of all $n\times n$ correlation matrices,
$$
\mathcal{E}_{n}=\{ C\in\mathbb{R}^{n\times n} : C=C^{\prime}, C\succeq 0, C_{ii}=1 \},
$$
is relatively unexplored in high dimensions. The constraints defining $\mathcal{E}_{n}$ are highly nonlinear: the positive semidefinite requirement constrains all off-diagonal elements, while the unit diagonal fixes an affine subspace of dimension $d=n(n-1)/2$. An intersection of the positive semidefinite (PSD) cone $\mathbb{S}_{+}^{n}$ with an affine subspace is a spectrahedron, and the particular spectrahedron that arises from unit diagonal elements is called the elliptope.
The exact closed-form formula for $\operatorname{Vol}(\mathcal{E}_n)$ is a classical result, first derived by \citet{JohnsonNaevdal:1998}, and has been applied mainly to bound the failure rates of rejection-sampling algorithms; its history is reviewed in Section~\ref{sec:volume}. The main contribution of this paper is not that formula itself, but the high-dimensional probabilistic consequences of the restricted-Wishart/LKJ equivalence (Proposition~\ref{prop:WishartGram}): the known identification of the uniform distribution on $\mathcal{E}_n$ with the law of the uncentered sample correlation matrix of $n$ independent Gaussian variables (equivalently, the Pearson correlation matrix at $T=n+2$), together with the refined volume asymptotics, organizes much of the analysis that follows.

Proposition~\ref{prop:VolumeCn} restates the volume formula together with a refined asymptotic expansion, equivalent to that of \citet{BohmHornik:2014}, in the natural dimension $d=n(n-1)/2$.

The vanishing volume of $\mathcal{E}_n$ is not spread evenly over the elliptope. We prove that it concentrates around the identity matrix: as $n$ increases, the algebraic restrictions imposed by positive semidefiniteness make large pairwise correlations exceptional, severely restricting the entrywise max-norm, while simultaneously the mass expands globally in the Frobenius distance. This two-scale geometry is an intrinsic structural feature of the elliptope itself, not an artifact of any particular estimator or prior. Via Proposition~\ref{prop:WishartGram}, a uniform draw from $\mathcal{E}_n$ is exactly distributed as the uncentered sample correlation matrix of $n$ independent Gaussian variables observed $T=n+1$ times, so the empirical eigenvalue distribution of a uniformly drawn correlation matrix converges to the Marchenko--Pastur law with ratio one. The smallest eigenvalue has the exact law $\lambda_{\min}(C)\sim\operatorname{Beta}(1,d)$ (Theorem~\ref{thm:specfloor}), obtained from an elementary homothety; to our knowledge this finite-dimensional law is new.

There is an extensive literature on extreme values of sample correlations. For sample correlation matrices of independent Gaussian variables, the Gumbel limit of the largest off-diagonal entry was established by \citet{Jiang:2004maxentry}; pairwise independence of distinct sample correlations under Gaussian and spherical null models is classical, and follows from the representation of the sample correlations as inner products of independent vectors uniform on the sphere, which \citet{HeroRajaratnam:2011} use to derive the mean number and Poisson behavior of threshold exceedances in their correlation-screening framework; and Poisson point-process convergence for the off-diagonal entries of sample covariance and correlation matrices is developed in generality by \citet[theorem~3.7]{HeinyMikoschYslas:2021}. Via the sample-correlation identification, these results apply at $T=n+2$ and therefore already cover the uniform law on $\mathcal{E}_n$. We do not claim them as new; our contribution on the extremes is threefold. First, Lemma~\ref{lem:pairwise} extends exact pairwise independence from the classical Gaussian/spherical sample-correlation setting to \emph{every} $\operatorname{LKJ}(\eta)$ distribution, including noninteger effective Wishart degrees, where no sample-correlation representation is available; the proof is through the restricted-Wishart/Bartlett representation. Second, this exactness yields a short, self-contained proof of the Poisson point-process limit (Theorem~\ref{thm:PPP}), together with an $O(n^{-1})$ total-variation bound for the finite-dimensional exceedance counts relative to Poisson laws with their exact finite-$n$ means: the covariance contribution from edges sharing a vertex vanishes identically rather than asymptotically, and the entire argument fits in two pages. Third, the point process is put to work as a structurally coherent null model for thresholded correlation networks (Section~\ref{sec:extremes}).

Our other two main results are Theorem~\ref{thm:LKJscaling} and Theorem~\ref{thm:nearest}. Theorem~\ref{thm:LKJscaling} shows that the LKJ$(\eta_n)$ prior has a two-scale phase structure: the Marchenko--Pastur ratio changes only when $\eta_n\asymp n$, whereas the Frobenius distance from the identity is bounded only at the much larger rate $\eta_n\asymp n^2$; these two critical scalings are not visible from the LKJ density alone, but emerge from the sample-correlation identification. Theorem~\ref{thm:nearest} establishes that replacing an entrywise-specified matrix by its nearest valid correlation matrix is not a minor correction: the squared Frobenius repair cost is asymptotically at least one-half of the squared Frobenius norm of the off-diagonal part, a lower bound that follows from the PSD projection inequality and the Wigner semicircle law for bounded, centered i.i.d.\ entrywise specifications with positive variance.

The volume and structure of $\mathcal{E}_{n}$ are relevant for several reasons.
First, they help explain why empirical correlation matrices are noisy, ill-conditioned, and strongly shaped by random-matrix effects whenever $n$ is large relative to the sample size. When $T\asymp n$, the accumulated entrywise sampling noise produces a Frobenius displacement from $I_n$ of order $\sqrt{n}$, which is exactly the global scale of the elliptope itself (Proposition~\ref{prop:C-I}); and at the critical sample size $T=n+1$ (uncentered), equivalently $T=n+2$ (centered), the sample correlation matrix is not merely at that scale but is exactly uniform on $\mathcal{E}_n$ (Proposition~\ref{prop:WishartGram}). The associated spectral law follows from the Wishart/LKJ identity and is the Marchenko--Pastur law of Theorem~\ref{thm:MP}.
Second, the volume of $\mathcal{E}_{n}$ quantifies how rapidly semidefinite feasible sets shrink in high dimensions. The elliptope is arguably the simplest nontrivial spectrahedron, and its vanishing volume illustrates how rapidly the feasible set shrinks as positivity constraints interact with linear constraints. This sheds light on the structure of feasible sets in high-dimensional semidefinite programs, where feasible regions often become extremely thin even when described by relatively few constraints. Related geometric constraints appear in machine learning, including PSD-constrained network layers \citep{HuangVanGool:2017} and correlation-parameterized kernels in Gaussian process models.
Third, as a consequence of this volume collapse, naive parameterizations of correlation matrices, such as choosing the off-diagonal entries independently in $[-1,1]$, almost never yield a positive semidefinite matrix when $n$ is large. In practice, one therefore relies on parameterizations that enforce positive semidefiniteness by construction, such as the Generalized Fisher Transformation (GFT) by \citet{ArchakovHansen:Correlation}, partial correlations \citep{Joe:2006}, vines \citep{LewandowskiKurowickaJoe:2009}, or hyperspherical angles \citep{PourahmadiWang:2015}; see also \citet{ArchakovHansenLuo-RandomCorr:2024}.
Fourth, in Bayesian modeling, prior distributions over correlation matrices have the elliptope as their support. The separation strategy of \citet{BarnardMcCullochMeng:2000}, which models a covariance matrix as $\Sigma=DCD$ with independent priors on the standard deviations $D$ and the correlation matrix $C$, makes the geometry of $\mathcal{E}_n$ directly relevant to prior elicitation. This geometry causes volume-based priors, including fixed-$\eta$ LKJ laws, to exhibit dimension-induced shrinkage of individual correlations. However, because the space simultaneously expands globally in the Frobenius norm, the probability mass is often pushed outward toward the singular boundary. This tension shapes the behavior of priors, such as the LKJ$(\eta)$ distribution, in high dimensions, and it is a consequence of the underlying space rather than of the prior specification itself.
Fifth, these constraints help explain the well-documented instability of minimum-variance portfolios: the portfolio solution requires the precision matrix $\Sigma^{-1}$, and in the decomposition $\Sigma=DCD$ the boundary-seeking behavior of the noisy empirical correlation matrix drives large and erratic changes in portfolio weights even when the volatilities in $D$ are perfectly estimated; Section~\ref{sec:precision} quantifies this.



\section{Theoretical Results on Volume}\label{sec:volume}
The set of valid $n\times n$ correlation matrices $\mathcal{E}_n$ is a subset of the hypercube $[-1,1]^{d}$,
where $d=n(n-1)/2$.
We will use both the $n$- and $d$-parameterizations as convenient, where the inverse relation is $n=\frac{1+\sqrt{1+8d}}{2}$. All limits and asymptotic statements are as $n\rightarrow\infty$, equivalently $d\rightarrow\infty$, unless stated otherwise.

The exact formula (\ref{eq:VolCn}) is a classical result, first derived by \citet{JohnsonNaevdal:1998} via the Schur parametrization of positive semidefinite matrices; equivalent expressions were subsequently obtained by \citet{Joe:2006} and \citet{PourahmadiWang:2015}, and rediscovered via the hypersphere decomposition by \citet{EastmanHollisNumpacharoenSchlieper:2016}, extending \citet{RousseeuwMolenberghs:1994}; see \citet{ForresterZhang:2020} for a survey. The following proposition states this classical formula together with the asymptotic expansion (\ref{eq:VolCnAsym}), which is a reparameterization of the expansion of \citet{BohmHornik:2014}. Throughout, $\zeta(s)$ denotes the Riemann zeta function; $\zeta(-1)=-1/12$ and $\zeta^\prime(-1)\approx -0.1654$.

\begin{proposition}\label{prop:VolumeCn}The Lebesgue volume of $\mathcal{E}_n$ satisfies the recursion
$$
\operatorname{Vol}(\mathcal{E}_{n+1})
=\operatorname{Vol}(\mathcal{E}_{n})\times
\left[B(\tfrac{n+1}{2},\tfrac{1}{2})\right]^{n},\qquad n\geq1,
$$
with $\operatorname{Vol}({\mathcal{E}_{1}})\equiv 1$; consequently,
\begin{equation}\label{eq:VolCn}
\operatorname{Vol}(\mathcal{E}_n)
=\pi^{\frac{n(n-1)}{4}}
\prod_{j=1}^{n}\frac{\Gamma\big(\tfrac{j+1}{2}\big)}{\Gamma\big(\tfrac{n+1}{2}\big)},
\qquad\text{for }n\geq 1.
\end{equation}
For large $n$, $\log\operatorname{Vol}(\mathcal{E}_n)=-\tfrac{n^2}{4}\log \frac{n}{2\pi\sqrt{e}}+O(n\log n)$ and in terms of $d=\tfrac{n(n-1)}{2}$ we have
\begin{equation}\label{eq:VolCnAsym}
\log\operatorname{Vol}(\mathcal{E}_n)
=-\tfrac{d}{4}\log(\tfrac{d}{2e\pi^2})-\tfrac{1}{2\sqrt{2}}\sqrt{d}-\tfrac{1}{48}\log d + \kappa
+O\big(\tfrac{1}{\sqrt{d}}\big),
\end{equation}
where $\kappa=\tfrac{1}{4}\left[(1+\zeta(-1))\log 2-\zeta(-1)+2\zeta^\prime(-1)\right]\approx 0.097$.
\end{proposition}
The expansion of \citet{BohmHornik:2014}, obtained in their analysis of rejection sampling, is stated in the $n$-parameterization for the hypercube probability $\operatorname{Vol}(\mathcal{E}_n)/2^d$; its constant involves the Glaisher--Kinkelin constant $A$, which is equivalent to $\kappa$ via $\log A=\tfrac{1}{12}-\zeta^\prime(-1)$. The form (\ref{eq:VolCnAsym}) is compact in the intrinsic dimension and accurate even for moderate $n$: its absolute error is below $0.02$ already at $n=5$ and below $10^{-3}$ by $n=100$, in line with the $O(d^{-1/2})$ remainder, whereas the leading-order term alone has absolute error that diverges with $n$; see Figure~\ref{fig:VolCorr} in the Supplement. Exact closed-form expressions and numerical values for selected dimensions up to $n=30$ are given in Table~\ref{tab:CnVolume} in the Supplement. For perspective: at $n=10$ the feasible set occupies roughly $2\times10^{-14}$ of the hypercube $[-1,1]^{45}$; at $n=20$ this drops to roughly $4\times10^{-86}$; and by $n=30$ to below $10^{-233}$.


\section{Concentration, Spectral Structure, and Extreme Correlations}\label{sec:concentration}

While the volume of the elliptope shrinks rapidly, we now show where this vanishing volume concentrates. The answer depends on the norm. Proposition~\ref{prop:entrywise} shows that almost all of $\mathcal{E}_n$ lies in a thin entrywise neighborhood of $I_n$: the largest correlation in a typical matrix is of order $\sqrt{\log n/n}$ and vanishes as $n\rightarrow\infty$. Proposition~\ref{prop:C-I} shows that the same typical matrix is nevertheless far from the identity in the Frobenius norm, at a distance of order $\sqrt{n}$, and Theorem~\ref{thm:specfloor} gives the exact law of the smallest eigenvalue, $\lambda_{\min}(C)\sim\operatorname{Beta}(1,d)$. There is no contradiction: individually the correlations are tiny, but there are $d=n(n-1)/2$ of them, and collectively they add up. An exact identity connects these results to sample correlation matrices: the uniform distribution on $\mathcal{E}_n$ is the distribution of an uncentered sample correlation matrix based on $T=n+1$ Gaussian observations (Proposition~\ref{prop:WishartGram}), so the empirical eigenvalue distribution of a typical correlation matrix converges to the Marchenko--Pastur law (Theorem~\ref{thm:MP}). The section ends with the Poisson point-process limit of the extreme correlations (Theorem~\ref{thm:PPP}).

\subsection{Entrywise and Frobenius Concentration}\label{sec:twoscale}

For $r\in[0,1]$ we define
$$\mathcal{E}_n(r)=\{C\in \mathcal{E}_n:\max_{i<j}|C_{ij}|\leq r\},$$
which is the subset of $\mathcal{E}_n=\mathcal{E}_n(1)$ about the identity matrix, $\mathcal{E}_n(0)=\{I_n\}$.

\begin{proposition} \label{prop:entrywise}
    For any fixed $r\in(0,1)$,
    $$\log\left(1-\frac{\operatorname{Vol}(\mathcal{E}_n(r))}{\operatorname{Vol}(\mathcal{E}_n)}\right)=-\frac{n}{2}\left|\log(1-r^2)\right|+O(\log n),\quad\text{as }n\rightarrow\infty.$$
\end{proposition}

The statement is at the exponential scale: the fraction of $\operatorname{Vol}(\mathcal{E}_n)$ that lies outside $\mathcal{E}_n(r)$ vanishes at the rate $\exp\{-\tfrac{n}{2}|\log(1-r^2)|\}$, up to a polynomial factor that is absorbed by the $O(\log n)$ term. The result follows from the Beta tail of a single off-diagonal entry, $p_n(r)=\Pr(|C_{12}|>r)$, combined with Boole's inequality. The proposition shows that almost all high-dimensional correlation matrices are in the vicinity of the identity matrix. The convergence is slow, however, as the following typical-value calculation for the largest correlation coefficient shows.

Let $N_r=\sum_{i<j}1_{\{|C_{ij}|>r\}}$. By linearity of expectation and the symmetry of the uniform distribution, $\mathbb{E}(N_r)=d\Pr(|C_{ij}|>r)$ exactly. Note that $\Pr(\max_{1\leq i<j\leq n}|C_{ij}|>r)=\Pr(N_r\geq 1)$. We can select $r$ such that $\mathbb{E}(N_r)\approx 1$, by solving $\Pr(|C_{ij}|>r)=1/d$. Under the uniform distribution on $\mathcal{E}_n$ we have $(C_{ij}+1)/2\sim \operatorname{Beta}(\tfrac{n}{2},\tfrac{n}{2})$, so $\operatorname{var}(C_{ij})=\tfrac{1}{n+1}$ and $\sqrt{n+1}C_{ij}\overset{d}{\rightarrow}N(0,1)$ as $n\rightarrow\infty$; we use this Gaussian approximation to evaluate $\Pr(|C_{ij}|>r)$.

Next, we use the standard extreme-value expansion: the solution to $\Phi(-q)=\tfrac{1}{m}$ is
$$q=\sqrt{2\log m} -\frac{\log\log m +\log4\pi}{2\sqrt{2\log m}} + o\left(\tfrac{1}{\sqrt{\log m}}\right).$$
With $m=2d\approx n^2$ and $\sqrt{n+1}r =  q$
we arrive at
\begin{equation}
M_n\equiv\max_{1\leq i<j\leq n}|C_{ij}|\approx\left(\frac{1}{n+1}\log\frac{n^4}{8\pi\log n}\right)^{1/2}.
\label{eq:maxrhoapprox}
\end{equation}

The approximation (\ref{eq:maxrhoapprox}) is shown in Figure~\ref{fig:CorrNearIdentity}. We have simulated random correlation matrices of varying dimensions to verify this asymptotic behavior, showing close agreement with the theoretical approximation.
\begin{figure}[!htb]
\centering{}\includegraphics[width=0.7\textwidth]{Figures/MaximumCorrelation.pdf}
\caption{{\small{}The largest absolute correlation coefficient in a uniformly distributed $n\times n$ correlation matrix. The theoretical approximation (\ref{eq:maxrhoapprox}) is the solid curve; circles show simulated averages of $M_n$ over independent draws from the uniform distribution on $\mathcal{E}_n$, with the number of replications decreasing from $10{,}000$ at $n=50$ to $50$ draws for $n=10{,}000$. Fewer replications are needed in high dimensions because $M_n$ concentrates, cf.\ Corollary~\ref{cor:gumbel}.\label{fig:CorrNearIdentity}}}
\end{figure}

The approximation (\ref{eq:maxrhoapprox}) is a heuristic typical-value calculation; the rigorous fluctuation limit is the Gumbel law of Corollary~\ref{cor:gumbel} in Section~\ref{sec:extremes}.

Entrywise concentration does not, however, imply global concentration. The next result measures the typical distance from $I_n$ in the Frobenius norm.

\begin{proposition} \label{prop:C-I}
Let $C$ be drawn uniformly from $\mathcal{E}_n$. Then
$$ \mathbb{E}\|C - I_n\|_F^2 = \frac{n(n-1)}{n+1}=n(1+o(1)), $$
and $\|C-I_n\|_F^2/\mathbb{E}\|C-I_n\|_F^2\rightarrow1$ in probability as $n\rightarrow\infty$.
\end{proposition}
Proposition~\ref{prop:entrywise} and Proposition~\ref{prop:C-I} together characterize the typical correlation matrix at two different scales. Entrywise, it is nearly indistinguishable from the identity: by (\ref{eq:maxrhoapprox}), even its single largest correlation is only of order $\sqrt{\log n/n}$. Globally, it is far from the identity: each of the $d=n(n-1)/2$ off-diagonal pairs contributes $\mathbb{E}[C_{ij}^{2}]=\tfrac{1}{n+1}$ to the squared Frobenius distance, and while every individual term is negligible, their sum, $2d/(n+1)\sim n$, is not. In summary,
$$
\max_{i<j}|C_{ij}| \approx 2\sqrt{\tfrac{\log n}{n}}\rightarrow 0,
\qquad\text{while}\qquad
\|C-I_n\|_F \approx \sqrt{n}\rightarrow\infty.
$$
A typical high-dimensional correlation matrix is therefore locally near, but globally far from, the identity matrix. This two-scale structure resolves the apparent tension between entrywise shrinkage and determinant behavior, and it is the key to understanding why empirical eigenvalue distributions behave the way they do in high dimensions, as we show next.

\subsection{Spectral Structure and the Sample-Correlation Representation}\label{sec:spectral}

A third measurement looks at the spectrum directly. The following identity is exact at every $n$; no asymptotics are involved.

\begin{theorem}[Spectral floor]\label{thm:specfloor}
For $n\geq2$ and $x\in[0,1]$,
$$
\operatorname{Vol}\{C\in\mathcal{E}_n:\lambda_{\min}(C)\geq x\}=(1-x)^{d}\operatorname{Vol}(\mathcal{E}_n).
$$
\end{theorem}

Under the uniform law, $\Pr(\lambda_{\min}(C)\geq x)=(1-x)^d$; equivalently, $\lambda_{\min}(C)\sim\operatorname{Beta}(1,d)$. Despite the simple proof, given in the appendix, this exact finite-dimensional law is, to our knowledge, new for the elliptope. The underlying homothety is elementary, and has a close analogue for complex density matrices under the Hilbert--Schmidt measure \citep{MajumdarBohigasLakshminarayan:2008}. Two consequences follow. First, the two constraints live on different exponential scales: the entrywise ceiling $\max_{i<j}|C_{ij}|\leq r$ excludes only a fraction $\exp\{-\tfrac{n}{2}|\log(1-r^2)|+O(\log n)\}$ of the volume (Proposition~\ref{prop:entrywise}), whereas the spectral floor $\lambda_{\min}(C)\geq x$ is highly restrictive: it is satisfied by only the fraction $(1-x)^d$, which decays at rate $n^2$ in the exponent. Second, $\Pr(n^2\lambda_{\min}(C)>t)=(1-t/n^2)^d\rightarrow e^{-t/2}$, so $n^2\lambda_{\min}(C)$ converges in distribution to an exponential with mean two, which recovers the hard-edge scale discussed in Remark~\ref{rem:lambdamin} by elementary means. The probability identity is special to the uniform law: for $\eta\neq1$ the determinant tilt breaks the affine invariance, while the set homothety and the volume identity are geometric and hold regardless.

The connection between the elliptope and the Marchenko--Pastur phenomenon is more than an analogy. The density factorization in Proposition~\ref{prop:WishartGram}(i) is the restricted-Wishart/LKJ equivalence established by \citet{WangWuChu:2018}, building on the LKJ construction of \citet{LewandowskiKurowickaJoe:2009} and the separation-strategy lineage of \citet{BarnardMcCullochMeng:2000}. The integer-degree Gram representation in part~(ii) is also standard. Our contribution is to exploit this bridge systematically: at $\eta=1$, the determinant exponent vanishes, so the uniform law on $\mathcal{E}_n$ is exactly a sample-correlation ensemble at the critical edge $T/n\to1$. This observation transfers random-matrix and extreme-value structure to the elliptope and is the source of the Marchenko--Pastur law, the Frobenius expansion, the LKJ scaling theorem, and the Poisson point-process result below.

\begin{proposition}[Wishart and Gram representations]\label{prop:WishartGram}
(i) Let $\nu>n-1$ be real, let $W$ have the Wishart distribution $W_n(\nu,I_n)$, and let $C=D^{-1}WD^{-1}$ where $D=\operatorname{diag}(\sigma_1,\ldots,\sigma_n)$ with $\sigma_i^2=W_{ii}$. Then $C$ has density proportional to $\det(C)^{(\nu-n-1)/2}$ on $\mathcal{E}_n$; equivalently, $C\sim\operatorname{LKJ}(\eta)$ with $\nu=n+2\eta-1$, for every $\eta>0$.\newline
(ii) For integer $\nu=k\geq n$, $C$ is distributed as the Gram matrix $C=V^\prime V$, $V=(v_1,\ldots,v_n)$, of $n$ independent vectors $v_i$ that are uniformly distributed on the unit sphere in $\mathbb{R}^k$. In particular, $C$ is uniformly distributed on $\mathcal{E}_n$ if and only if $k=n+1$.
\end{proposition}

Part~(i) is the restricted-Wishart/LKJ equivalence of \citet{WangWuChu:2018}, written in the notation used here. Part~(ii) is the standard integer-degree Gram representation; \citet{Joe:2006} gives this construction explicitly, and \citet{LewandowskiKurowickaJoe:2009} use the same spherical representation.

The equivalence also gives a direct Bartlett sampler for $\operatorname{LKJ}(\eta)$ correlation matrices: draw $W\sim W_n(n+2\eta-1,I_n)$ by the Bartlett decomposition and normalize to a correlation matrix. This is the restricted-Wishart sampler proposed by \citet{WangWuChu:2018}. In the implementation used here, one may row-normalize the Bartlett factor itself, producing a Cholesky factor of the LKJ draw without first forming the dense Wishart matrix; forming $C$ then requires the additional multiplication $\tilde{L}\tilde{L}^\prime$.

Proposition~\ref{prop:WishartGram}(i) also identifies $\operatorname{Vol}(\mathcal{E}_n)$ as the normalizing constant of the sample-correlation ensemble at $T=n+1$.  The LKJ$(\eta)$ density is proportional to $\det(C)^{\eta-1}$ on $\mathcal{E}_n$, with normalizing constant
$$
Z_n(\eta) = \int_{\mathcal{E}_n}\det(C)^{\eta-1}dC;
$$
setting $\eta=1$ gives $Z_n(1)=\operatorname{Vol}(\mathcal{E}_n)$.  The product formula~(\ref{eq:VolCn}) is the explicit evaluation of $Z_n(1)$ via the Bartlett decomposition underlying Proposition~\ref{prop:WishartGram}(i), and the asymptotic expansion~(\ref{eq:VolCnAsym}) is its large-$n$ consequence.  The volume formula and the sample-correlation identification therefore rest on the same Wishart normalizing constant.

Writing $v_i=z_i/\|z_i\|$ with $z_1,\ldots,z_n$ independent $N(0,I_k)$, the entries $C_{ij}=z_i^\prime z_j/(\|z_i\|\|z_j\|)$ are the (uncentered) sample correlations of $n$ independent Gaussian variables computed from $T=k$ observations. So, Proposition~\ref{prop:WishartGram} shows that a uniform draw from $\mathcal{E}_n$ is distributed exactly as the empirical correlation matrix of $n$ independent Gaussian variables observed $T=n+1$ times, one more observation than the dimension. Only the spherical symmetry of the Gaussian distribution is used here, so the identity extends to any spherically distributed observation vectors. Equivalently, centering $T=n+2$ independent Gaussian observations projects each column onto the $(T-1)=n+1$ dimensional subspace orthogonal to the vector of ones, and after normalization the centered columns are independent and uniformly distributed on the unit sphere in that subspace; a uniform draw is therefore also distributed exactly as the centered (Pearson) sample correlation matrix based on $T=n+2$ Gaussian observations. In either form, the uniform distribution corresponds to a sample correlation matrix with $n/T\rightarrow1$, and the limiting theory of sample correlation matrices applies directly.

\begin{theorem}[Marchenko--Pastur law]\label{thm:MP}
Let $C_n$ be uniformly distributed on $\mathcal{E}_n$, with eigenvalues $\lambda_1,\ldots,\lambda_n$, and let $F_n(x)=\frac{1}{n}\#\{i:\lambda_i\leq x\}$ denote the empirical spectral distribution. Then, as $n\rightarrow\infty$, $F_n$ converges weakly, in probability, to the Marchenko--Pastur law with ratio one, whose density is
$$
f_{\mathrm{MP}}(x)=\frac{1}{2\pi x}\sqrt{x(4-x)},\qquad x\in(0,4].
$$
\end{theorem}

\begin{remark}[Coupling]\label{rem:coupling}
Almost-sure convergence holds when the $C_n$ are coupled on a common probability space as follows. Let $Z=(Z_{ti})_{t\geq1,i\geq1}$ be an infinite array of independent $N(0,1)$ entries. For each $n$, let $z_i^{(n)}=(Z_{1i},\ldots,Z_{n+1,i})^\prime\in\mathbb{R}^{n+1}$ and set $v_i^{(n)}=z_i^{(n)}/\|z_i^{(n)}\|$; then $C_n=(v_i^{(n)\prime}v_j^{(n)})_{i,j=1}^n$ is uniformly distributed on $\mathcal{E}_n$ by Proposition~\ref{prop:WishartGram}(ii). Since $C_n$ is the uncentered sample correlation matrix of the first $n$ variables observed $T=n+1$ times (equivalently, the Pearson sample correlation matrix based on $T=n+2$ observations; see the paragraph following Proposition~\ref{prop:WishartGram}), almost-sure weak convergence of $F_n$ follows from \citet{HeinyMikosch:2018}.
\end{remark}

Figure~\ref{fig:MPlaw} illustrates the convergence by comparing empirical spectral distributions of uniform draws from $\mathcal{E}_n$ with the Marchenko--Pastur density.

The limiting spectrum reproduces the two scales established above. The second moment of the Marchenko--Pastur law about one is $\int(x-1)^2f_{\mathrm{MP}}(x)dx=1$, so $\|C-I_n\|_F^2=\sum_i(\lambda_i-1)^2\approx n$, which is shown in Proposition~\ref{prop:C-I}. Likewise, the approximation (\ref{eq:maxrhoapprox}) for the largest correlation is consistent with the Gumbel fluctuation limit established in Section~\ref{sec:extremes}. Entrywise, a uniform draw is nearly the identity matrix; spectrally, its eigenvalues spread over the entire interval $[0,4]$, with substantial mass near zero. A uniform draw therefore becomes asymptotically ill-conditioned: the typical matrix sits close to the singular boundary of $\mathcal{E}_n$, even though every individual correlation is tiny. Section~\ref{sec:precision} quantifies this ill-conditioning.

\begin{figure}[htbp!]
\centering{}\includegraphics[width=0.8\textwidth]{Figures/EigenvalueDistributionMP.pdf}
\caption{{\small{}Empirical spectral distributions of uniformly distributed $n\times n$ correlation matrices for $n\in\{50,200,500\}$, based on pooled eigenvalues from independent draws from $\mathcal{E}_n$ via the LKJ$(1)$ representation. Each histogram pools $1{,}000$ eigenvalues: $20$ draws at $n=50$, $5$ draws at $n=200$, and $2$ draws at $n=500$. The Marchenko--Pastur density with ratio one, $f_{\mathrm{MP}}(x)=(2\pi x)^{-1}\sqrt{x(4-x)}$, is shown in red. Convergence to the limiting law is visually apparent already at moderate $n$.\label{fig:MPlaw}}}
\end{figure}

\subsection{Extreme Correlations and Poisson Limits}\label{sec:extremes}

The preceding results describe the bulk spectrum and the largest entry separately. We now characterize the full collection of extreme off-diagonal entries. The key input is exact pairwise independence, which holds throughout the LKJ family.

\begin{lemma}[Pairwise independence]\label{lem:pairwise}
Let $C\sim\operatorname{LKJ}_n(\eta)$ with $\eta>0$. Any two distinct off-diagonal entries $C_{ij}$ and $C_{kl}$, $\{i,j\}\neq\{k,l\}$, are independent. In particular this holds for the uniform distribution on $\mathcal{E}_n$ ($\eta=1$).
\end{lemma}

\begin{lemma}[Beta-tail intensity]\label{lem:betatail}
Let $\mu_n=\log\tfrac{n^4}{\log n}$, and let $\Lambda$ be the measure on $\mathbb{R}$ with density
$$
\lambda(t)=\tfrac{1}{2\sqrt{8\pi}}e^{-t/2},\qquad\text{so that }\Lambda(t,\infty)=\tfrac{1}{\sqrt{8\pi}}e^{-t/2}.
$$
Under the uniform law on $\mathcal{E}_n$,
$$
d\cdot\Pr\{(n+1)C_{12}^2-\mu_n>t\}\rightarrow\Lambda(t,\infty),\qquad n\to\infty,
$$
for every $t\in\mathbb{R}$. More generally, $d\cdot\Pr\{(n+1)C_{12}^2-\mu_n\in B\}\to\Lambda(B)$ for every Borel set $B\subseteq\mathbb{R}$ that is bounded away from $-\infty$, with $\Lambda(B)<\infty$ and $\Lambda(\partial B)=0$.
\end{lemma}

Thresholding $C$ at level $\tau\in(0,1)$ produces a correlation network, with edge count
$$
N_n(\tau)=\#\{i<j:|C_{ij}|>\tau\},
$$
and we write $p_n(\tau)=\Pr(|C_{12}|>\tau)$ for the common exceedance probability of a single entry.
The exceedance point process associated with $C$ is the random counting measure defined, for a Borel set $B\subseteq\mathbb{R}$, by
$$
\Xi_n(B)=\#\{i<j:(n+1)C_{ij}^2-\mu_n\in B\};
$$
in particular, $\Xi_n(t,\infty)$ counts how many of the $d$ normalized squared correlations exceed $t$. Lemma~\ref{lem:betatail} gives $\mathbb{E}[\Xi_n(t,\infty)]=d\cdot\Pr\{(n+1)C_{12}^2-\mu_n>t\}\to\Lambda(t,\infty)$, so the expected count converges to the intensity of a Poisson process with measure $\Lambda(dt)=\frac{1}{2\sqrt{8\pi}}e^{-t/2}dt$. Theorem~\ref{thm:PPP} below shows that $\Xi_n$ converges to that Poisson process in distribution; formally, convergence is in the vague topology on the space of Radon measures on $(-\infty,\infty]$.

\begin{theorem}[Poisson point process]\label{thm:PPP}
Let $C$ be uniformly distributed on $\mathcal{E}_n$. The exceedance process $\Xi_n$ converges to a Poisson point process with intensity $\Lambda(dt)=\frac{1}{2\sqrt{8\pi}}e^{-t/2}dt$. Equivalently, for any finite collection of disjoint Borel sets $B_1,\ldots,B_m$ bounded away from $-\infty$, with $\Lambda(B_r)<\infty$ and $\Lambda(\partial B_r)=0$,
$$
\bigl(\Xi_n(B_1),\ldots,\Xi_n(B_m)\bigr)\overset{d}{\rightarrow}(Z_1,\ldots,Z_m),
$$
where $Z_1,\ldots,Z_m$ are independent with $Z_r\sim\operatorname{Poisson}(\Lambda(B_r))$.
\end{theorem}

This theorem is not new as a statement: via Proposition~\ref{prop:WishartGram}(ii), the uniform law on $\mathcal{E}_n$ is the sample correlation matrix of $n$ independent Gaussian series with $T=n+1$ observations, and point-process convergence of the off-diagonal entries in that setting is covered by the general theory of \citet[theorem~3.7]{HeinyMikoschYslas:2021}. The contribution here is the proof: Lemma~\ref{lem:pairwise} gives pairwise independence of off-diagonal entries exactly (not just asymptotically) and for every LKJ parameter, so the joint-expectation terms in the Chen--Stein bound reduce to products of marginal probabilities, and both error terms in the bound of \citet{ArratiaGoldsteinGordon:1989} are of order $n^{-1}$. The argument is elementary and self-contained, and it delivers an explicit $O(n^{-1})$ total-variation bound for the finite-dimensional exceedance counts relative to Poisson laws with the exact means $dp_n(B_r)$; the passage to the limiting process uses only the convergence of these means and carries no rate.

Two consequences follow directly by projecting the point process onto specific functionals.

\begin{corollary}[Edge counts and triangles]\label{cor:ppp_consequences}
Under the uniform law on $\mathcal{E}_n$: (a) if the thresholds $\tau_n\in(0,1)$ satisfy $dp_n(\tau_n)\rightarrow\lambda\in(0,\infty)$, which holds for $\tau_n^2=(\mu_n-\log(8\pi\lambda^2)+o(1))/(n+1)$, then $N_n(\tau_n)\overset{d}{\rightarrow}\operatorname{Poisson}(\lambda)$; and (b) at these thresholds, the probability that the thresholded graph contains a triangle tends to zero.
\end{corollary}

\begin{corollary}[Gumbel limit]\label{cor:gumbel}
Let $C$ be uniformly distributed on $\mathcal{E}_n$, and let $M_n=\max_{1\leq i<j\leq n}|C_{ij}|$. Then, as $n\rightarrow\infty$,
$$
(n+1)M_n^2-\mu_n\overset{d}{\rightarrow}\mathcal{G},\qquad \mathcal{G}(t)=\exp\{-\tfrac{1}{\sqrt{8\pi}}e^{-t/2}\},
$$
which is a Gumbel distribution. In particular, $\sqrt{(n+1)/\log n}M_n\overset{p}{\rightarrow}2$.
\end{corollary}

The rescaled maximum $(n+1)M_n^2-\mu_n$ is the largest atom of $\Xi_n$, so the Gumbel limit follows from Theorem~\ref{thm:PPP} and $\Pr\{\Xi_n(t,\infty)=0\}\rightarrow\exp\{-\Lambda(t,\infty)\}$. Via the sample-correlation identification, Corollary~\ref{cor:gumbel} is the specialization to $T=n+2$ of the classical largest-entry result of \citet{Jiang:2004maxentry}; it is the rigorous counterpart to the typical-value approximation (\ref{eq:maxrhoapprox}), whose square agrees with the implied typical value $M_n^2\approx\log\bigl(\tfrac{n^4}{2\pi\log n^4}\bigr)/(n+1)$ to second order.

Under the uniform law, the edge indicators of the thresholded network have common probability $p_n(\tau)$ and are pairwise independent by Lemma~\ref{lem:pairwise}. Consequently, every edge count has exactly the same mean and variance as under the Erd\H{o}s--R\'enyi model $G(n,p_n)$, a canonical benchmark in which edges are independent, although higher-order dependence remains because all edges arise from a single positive semidefinite matrix. At extreme thresholds with $dp_n\rightarrow\lambda$, this dependence is negligible to first order: edge counts converge to a Poisson distribution and triangles disappear with probability tending to one (Corollary~\ref{cor:ppp_consequences}). The uniform elliptope law therefore provides a structurally coherent finite-dimensional null whose sparse extreme limit agrees with the Erd\H{o}s--R\'enyi benchmark. The setting is close in spirit to the correlation-screening framework of \citet{HeroRajaratnam:2011}, extended to compound-Poisson characterizations in ultra-high dimension by \citet{WeiRajaratnamHero:2023}. Figure~\ref{fig:PPP} in the Supplement illustrates the convergence of $\Xi_n$ to the limit.

\section{Implications for High-Dimensional Modeling}\label{sec:implications}

Whether considering the $n$ assets in a minimum-variance portfolio or the $n$ covariates in a regression model, the matrix $C\in\mathcal{E}_n$ governs the dependency structure of the inputs. This section develops the consequences of the results above for sampling noise, prior specification, precision matrices, and entrywise specification.

\subsection{Sampling Noise in Empirical Correlation Matrices}\label{sec:MP}

The framework of Proposition~\ref{prop:WishartGram} and Theorem~\ref{thm:LKJscaling} covers sample correlation matrices at general aspect ratios; Theorem~\ref{thm:MP} is the critical case in which the sample correlation matrix is exactly uniform on $\mathcal{E}_n$. Three statements should be kept separate.

First, let $\hat{C}$ be the sample correlation matrix computed from $T$ independent observations of $n$ independent standardized Gaussian variables, the setting in which the Wishart identity is exact. Then $\operatorname{var}(\hat{C}_{ij})=1/(T-1)$ exactly for the Pearson correlation matrix, and hence
$$
\mathbb{E}\|\hat{C}-I_n\|_F^2=\frac{n(n-1)}{T-1},
$$
so when $T\asymp n$ the accumulated entrywise noise produces a Frobenius displacement of order $\sqrt{n}$: the same global scale as a uniform draw from the elliptope (Proposition~\ref{prop:C-I}). (For non-Gaussian data the $1/T$ variance scale and the spectral limit below continue to hold under standard moment conditions, e.g.\ independent standardized entries with finite fourth moments \citep{Jiang:2004,BaiSilverstein:2010}, but the exact distributional identities used in this paper are Gaussian, or more generally spherical, facts.)

Second, the spectral consequence is not read off from that scale but from the Wishart/LKJ identity: for $n/T\rightarrow c\in(0,1]$ the limiting spectral distribution is the Marchenko--Pastur law with ratio $c$, whose left edge reaches zero exactly at $c=1$, so the eigenvalues spread out and accumulate near the singular boundary.

Third, at the critical sample size $T=n+1$ uncentered, equivalently $T=n+2$ centered, the two coincide in the strongest possible sense: the law of $\hat{C}$ is then \emph{exactly} the uniform distribution on $\mathcal{E}_n$ (Proposition~\ref{prop:WishartGram}), and the uniform measure is the extreme case $c=1$. Away from that critical value the sample-correlation law is determinant-tilted, $\pi(C)\propto\det(C)^{\eta-1}$ with $\eta=(T-n+1)/2$ in the uncentered parameterization, so matching the Frobenius scale of the elliptope does not by itself mean that $\hat{C}$ is spread uniformly over it.

\subsection{Concentration of Correlation Priors in High Dimensions}\label{sec:LKJ}

By Proposition~\ref{prop:WishartGram}(i), the LKJ$(\eta)$ prior of \citet{LewandowskiKurowickaJoe:2009}, with density $\pi_\eta(C)\propto\det(C)^{\eta-1}$ on $\mathcal{E}_n$, is the correlation matrix of a Wishart with $\nu=n+2\eta-1$ degrees of freedom. At $\eta=1$ it reduces to the uniform distribution studied in Section~\ref{sec:concentration}, for which Proposition~\ref{prop:VolumeCn} supplies the exact normalizing constant, a quantity that is difficult to compute numerically in high dimensions. The Wishart identification allows the limit theory of sample correlation matrices to extend to the full LKJ family with aspect ratio $c_n=n/\nu_n$. The following result characterizes the prior when the hyperparameter is allowed to grow with the dimension.

\FloatBarrier
\begin{theorem}[LKJ scaling]\label{thm:LKJscaling}
Let $C_n\sim\operatorname{LKJ}_n(\eta_n)$ with $\eta_n>0$, and let $\nu_n=n+2\eta_n-1$.\newline
(i) If $\eta_n/n\rightarrow\theta\in[0,\infty)$, the empirical spectral distribution of $C_n$ converges weakly, in probability, to the Marchenko--Pastur law with ratio $c_\theta=1/(1+2\theta)$.\newline
(ii) $\mathbb{E}\|C_n-I_n\|_F^2=n(n-1)/\nu_n$, and $\|C_n-I_n\|_F^2/\mathbb{E}\|C_n-I_n\|_F^2\rightarrow1$ in probability. In particular, $\tfrac{1}{n}\|C_n-I_n\|_F^2\rightarrow1/(1+2\theta)$ in probability when $\eta_n/n\rightarrow\theta$, and when $\eta_n/n\rightarrow\infty$ the Frobenius distance diverges if $\eta_n=o(n^2)$, is bounded in probability if $\eta_n\asymp n^2$, and vanishes if $n^2=o(\eta_n)$.
\end{theorem}

Theorem~\ref{thm:MP} is the special case $\eta_n=1$ of part~(i): the uniform law is LKJ$(1)$, so $\eta_n/n\to0$, giving $\theta=0$ and $c_0=1$. The spectral result~(i) is the more interesting of the two. For any fixed $\eta$, however large, $\eta_n/n\to 0$ and the limiting spectrum is the Marchenko--Pastur law with ratio one, indistinguishable from the uniform distribution. Only when $\eta_n$ grows proportionally to $n$ does the aspect ratio shift away from one. The Frobenius result~(ii) is more mechanical: $\mathbb{E}\|C_n-I_n\|_F^2 = n(n-1)/(n+2\eta_n-1)\asymp n^2/(n+\eta_n)$, and when $\eta_n/n\rightarrow\infty$ this is $\sim n^2/(2\eta_n)$, so boundedness requires $\eta_n\asymp n^2$ to offset the $d\sim n^2/2$ off-diagonal terms. When $\nu_n$ is an integer, the prior is moreover a sample correlation matrix in the sense of Proposition~\ref{prop:WishartGram}(ii). Figure~\ref{fig:LKJphase} illustrates the two-scale structure.

\begin{figure}[!htbp]
\centering
\includegraphics[width=0.8\textwidth]{Figures/MPandFrobenius.pdf}
\caption{Dimension-dependent LKJ scaling. The left panel shows the Marchenko--Pastur limits from Theorem~\ref{thm:LKJscaling} when $\eta_n/n\rightarrow\theta$, so that the limiting aspect ratio is $c_\theta=1/(1+2\theta)$. The right panel shows the \emph{normalized} Frobenius scale under linear scaling $\eta_n=\theta n$, namely $\tfrac{1}{n}\mathbb{E}\|C_n-I_n\|_F^2=(n-1)/(n+2\eta_n-1)$, evaluated at $n=20$ and $n=100$, together with its limit $1/(1+2\theta)$ from Theorem~\ref{thm:LKJscaling}(ii). Linear scaling thus changes the limiting spectrum and the normalized Frobenius scale, but leaves $\|C_n-I_n\|_F^2$ of order $n$; quadratic scaling $\eta_n\asymp n^2$ is needed to keep the unnormalized Frobenius distance from the identity bounded.}
\label{fig:LKJphase}
\end{figure}

The central fact about fixed hyperparameters is that they are all alike to first order. Under $\operatorname{LKJ}(\eta)$ the marginal variance of an off-diagonal entry is exactly
\begin{equation}\label{eq:LKJvar}
\operatorname{var}(C_{ij})=\frac{1}{n+2\eta-1}=\frac{1}{\nu_n}\sim\frac{1}{n}\qquad\text{for every fixed }\eta>0,
\end{equation}
and by Theorem~\ref{thm:LKJscaling}(i) with $\theta=0$ the limiting spectrum is the Marchenko--Pastur law with ratio one, again for every fixed $\eta>0$. Entrywise concentration near $I_n$ is therefore not a feature of the tilt $\det(C)^{\eta-1}$ with $\eta>1$; it is a feature of the elliptope, and it occurs for $0<\eta<1$ as well. What a fixed $\eta$ does change is the finite-$n$ constant in (\ref{eq:LKJvar}) and, more visibly, the determinant and conditioning behavior, as we now describe.

For any correlation matrix $C$, $0\leq\det(C)\leq1$ with equality if and only if $C=I_n$. \citet{Joe:2006} shows that $\det(C)=\prod_{i<j}(1-\varrho_{ij}^2)$, so nonzero partial correlations reduce the determinant and move the matrix toward the boundary in the log-determinant sense. The scale of this effect is quantified by Proposition~\ref{prop:WishartGram}(i): the Bartlett decomposition of $W\sim W_n(n+1,I_n)$ gives the exact expression for $\mathbb{E}\log\det(C)$ stated in Remark~\ref{rem:logdet}.

This determinant penalty interacts with the volume collapse of the elliptope. As shown in Section~\ref{sec:volume}, the total volume $\operatorname{Vol}(\mathcal{E}_n)$ decays super-exponentially, meaning the set of all valid correlation matrices is already vanishingly small and entrywise confined near $I_n$. When $\eta>1$, the factor $\det(C)^{\eta-1}$ in the LKJ$(\eta)$ density compounds this thinness by down-weighting matrices away from the center; when $0<\eta<1$, the same factor favors small determinants and tilts mass toward the singular boundary instead.

For $C$ uniformly distributed on $\mathcal{E}_{n}$, the mean of $\log\det(C)$ is available in closed form. Writing $C=D^{-1}WD^{-1}$ as in Proposition~\ref{prop:WishartGram}, where $W$ is Wishart with $k=n+1$ degrees of freedom and $D=\operatorname{diag}(\sigma_1,\ldots,\sigma_n)$ with $\sigma_i^2=W_{ii}$, the Bartlett decomposition gives $\det(W)\overset{d}{=}\prod_{i=1}^{n}\chi^2_{k-i+1}$ while $W_{ii}\sim\chi^2_{k}$, and $\mathbb{E}[\log\chi^2_m]=\psi(\tfrac{m}{2})+\log2$ then yields
$$
\mathbb{E}[\log\det(C)]=\sum_{j=2}^{n+1}\psi(\tfrac{j}{2})-n\psi(\tfrac{n+1}{2})=-n+\tfrac{1}{2}\log(2n)+\tfrac{1+\gamma}{2}+o(1),
$$
where $\psi$ is the digamma function and $\gamma$ is the Euler--Mascheroni constant. On the log scale, the determinant of a typical high-dimensional correlation matrix is therefore of order $-n$, even under a uniform prior, in agreement with Theorem~\ref{thm:MP} and $\int\log xf_{\mathrm{MP}}(x)dx=-1$; the distributional asymptotics of $\det(C)$ for the uniform law are studied in \citet{HaneaNane:2018}.

Because the LKJ$(\eta)$ density is proportional to $\det(C)^{\eta-1}$, this exponentially small determinant contributes a factor of order $\exp\{-(\eta-1)n\}$ to the density. For $\eta>1$ the prior therefore places exponentially more mass on matrices with relatively large determinants, and for $0<\eta<1$ the negative exponent runs the same argument in reverse, tilting mass toward smaller determinants and more ill-conditioned matrices. Neither tilt is strong enough to matter at first order. By (\ref{eq:LKJvar}) the marginal variance is $1/\nu_n\sim1/n$ in both cases, by Theorem~\ref{thm:LKJscaling}(i) the limiting spectrum is the Marchenko--Pastur law with ratio one in both cases, and $\mathbb{E}\|C-I_n\|_F^2=n(n-1)/\nu_n\approx n$ diverges in both cases. In summary, a fixed $\eta>0$ alters finite-dimensional and lower-order features of the prior, in particular its determinant and conditioning behavior, but changes neither the first-order marginal variance nor the limiting spectral distribution; linear growth $\eta_n\asymp n$ is required to change the spectral limit, and quadratic growth $\eta_n\asymp n^2$ is required to keep the Frobenius distance from the identity bounded. This is what gives Theorem~\ref{thm:LKJscaling} its force: the hyperparameter must be tied to the dimension before it does anything to the two scales of Section~\ref{sec:concentration}.

A natural question within the LKJ family is whether dimension-dependent tuning of $\eta$ can prevent the entrywise shrinkage altogether. By (\ref{eq:LKJvar}), keeping $\operatorname{var}(C_{ij})$ bounded away from zero as $n\to\infty$ would require $\eta_n\sim -n/2$. This is a constraint specific to the LKJ parameterization, not a universal impossibility for Bayesian priors on correlation matrices: the density $\pi_\eta(C)\propto\det(C)^{\eta-1}$ is integrable only for $\eta>0$, so no valid member of the LKJ family can maintain dimension-stable marginal variances as $n$ grows. Priors built through an unconstrained reparameterization, such as the Generalized Fisher Transformation \citep{ArchakovHansen:Correlation} or partial-correlation sequences \citep{Joe:2006,LewandowskiKurowickaJoe:2009}, can place independent, dimension-stable distributions on the transformed parameters and thereby escape this constraint; see \citet{ArchakovHansenLuo-RandomCorr:2024} for an implementation.

Thus every admissible LKJ sequence satisfies $\operatorname{var}(C_{ij})\leq1/(n-1)\rightarrow0$. The two scales in Theorem~\ref{thm:LKJscaling} determine when its spectral and Frobenius behavior can nevertheless change.

\subsection{Precision Matrices and the Edge of Invertibility}\label{sec:precision}

Many of the applications described above require not the correlation matrix itself but its inverse: minimum-variance portfolio weights are proportional to $\Sigma^{-1}\iota$, generalized least squares is built on $C^{-1}$, and partial correlations are read off the precision matrix. The spectral law of Theorem~\ref{thm:MP} implies that the inverse of a typical correlation matrix is poorly behaved. The following proposition states two known consequences of the Marchenko--Pastur theory \citep{BaiSilverstein:2010,Jiang:2004} applied to the uniform law and to sample correlation matrices.

\begin{proposition}\label{prop:precision}
(i) Let $C_n$ be uniformly distributed on $\mathcal{E}_n$. Then
$$
\tfrac{1}{n}\operatorname{tr}(C_n^{-1})\rightarrow\infty,\qquad\text{in probability}.
$$
(ii) Let $\hat{C}_n$ be the sample correlation matrix of $n$ independent Gaussian variables based on $T$ observations, where $n/T\rightarrow c\in(0,1)$. Then
$$
\tfrac{1}{n}\operatorname{tr}(\hat{C}_n^{-1})\rightarrow\frac{1}{1-c},\qquad\text{almost surely}.
$$
\end{proposition}

Both parts follow from $\tfrac{1}{n}\operatorname{tr}(C^{-1})=\tfrac{1}{n}\sum_i\lambda_i^{-1}\to\int x^{-1}f_{\mathrm{MP},c}(x)dx=\tfrac{1}{1-c}$ for $c<1$, with part~(i) the limiting case $c\to1$. Part (i) is the spectral counterpart of boundary concentration. The Marchenko--Pastur density with ratio one behaves like $\pi^{-1}x^{-1/2}$ near zero, placing enough mass on near-singular matrices that the average inverse eigenvalue diverges: a typical correlation matrix has, per coordinate, infinite average precision. Part (ii) shows how this divergence builds up as the sample size shrinks toward the dimension. The limit $1/(1-c)$ is the familiar risk-inflation factor of plug-in estimators, and it explodes precisely as $c\rightarrow1$. The uniform measure on $\mathcal{E}_n$, which by Proposition~\ref{prop:WishartGram} corresponds to $T=n+1$, sits exactly at the edge where average precision ceases to exist. The same threshold appears in the exact Wishart identity $\mathbb{E}[\hat{\Sigma}^{-1}]=\tfrac{T}{T-n-1}\Sigma^{-1}$ for Gaussian covariance matrices, which is finite only for $T>n+1$.

This makes the instability of inverse problems quantitative. Any procedure that inverts an unregularized correlation matrix estimated with $T$ close to $n$ operates in a regime where the relevant population benchmark, the average precision of a typical feasible matrix, is infinite (Proposition~\ref{prop:precision}(i)). Shrinkage, factor structure, or explicit regularization of the spectrum, such as the condition-number regularization of \citet{WonLimKimRajaratnam:2013}, is therefore not merely advisable but necessary for the inverse to be statistically meaningful.

The Marchenko--Pastur law describes the bulk of the spectrum; the following two remarks quantify the smallest eigenvalue and the log-determinant.

\begin{remark}[Hard edge]\label{rem:lambdamin}
The Marchenko--Pastur law places the left edge of its support at zero when the ratio is one, so $\lambda_{\min}(C_n)\to0$ in probability (almost surely under the coupling of Remark~\ref{rem:coupling}). The rate and the limit law are exact consequences of Theorem~\ref{thm:specfloor}: since $\lambda_{\min}(C_n)\sim\operatorname{Beta}(1,d)$,
$$
\Pr\bigl(n^2\lambda_{\min}(C_n)>t\bigr)=\bigl(1-t/n^2\bigr)^{d}\rightarrow e^{-t/2},
\qquad\text{i.e.}\qquad
n^2\lambda_{\min}(C_n)\overset{d}{\rightarrow}\operatorname{Exp}(\text{mean }2),
$$
so $\lambda_{\min}(C_n)=O_p(n^{-2})$, far smaller than the $O(1)$ bulk eigenvalue scale. This is consistent with the classical hard-edge theory: the uniform law is the sample-correlation ensemble with $T=n+1$, so the ensemble sits at the hard edge of the real Laguerre family, where \citet[proposition~7.2.1]{Forrester:2010} gives $4n\lambda_{\min}(W)\overset{d}{\rightarrow}\xi_{\min}$ for $W\sim W_n(n+1,I_n)$, with $\xi_{\min}$ the smallest point of the $\beta=1$ Bessel point process. Passing to $C_n$ by the diagonal-normalization argument of Step~3 in the proof of Theorem~\ref{thm:LKJscaling} yields $n^2\lambda_{\min}(C_n)\overset{d}{\rightarrow}\tfrac14\xi_{\min}$, and matching the two limits shows that $\tfrac14\xi_{\min}$ is exponential with mean two. The exponential law itself is classical: at $\nu=n+1$ the smallest Wishart eigenvalue satisfies $\Pr(\lambda_{\min}(W)>x)=e^{-nx/2}$ exactly for every $n$, a case of the finite-$n$ results of \citet{Edelman:1991}; Theorem~\ref{thm:specfloor} gives an independent derivation on the correlation side.
\end{remark}

\begin{remark}[CLT for $\log\det C_n$]\label{rem:logdet}
For $C_n$ uniformly distributed on $\mathcal{E}_n$, \citet[theorem~3]{HaneaNane:2018} establish asymptotic normality of $\log\det C_n$ using the Beta factorization $\log\det C_n=\sum_{j=1}^{n-1}\log B_j$, $B_j\sim\operatorname{Beta}(\tfrac{j+1}{2},\tfrac{n-j}{2})$, which underlies Proposition~\ref{prop:WishartGram}(i), via a Lyapunov CLT for triangular arrays. Equivalently, using the expansion of the exact mean, the result reads
$$
\frac{\log\det C_n + n - \tfrac{1}{2}\log(2n) - \tfrac{1+\gamma}{2}}{\sqrt{2\log(n/2)}} \overset{d}{\rightarrow} N(0,1),
$$
because $\mathbb{E}[\log\det C_n]=-n+\tfrac{1}{2}\log(2n)+\tfrac{1+\gamma}{2}+o(1)$ by the digamma expression in Section~\ref{sec:LKJ}. The centering displayed in \citet[theorem~3]{HaneaNane:2018} differs from this in the sign of the logarithmic term; the discrepancy appears to be a misprint; Section~\ref{sec:centering} of the Supplement gives the numerical evidence and the reconciliation with \citet{ParolyaHeinyKurowicka:2024}. The $-n$ centering comes from the bulk: $\int_0^4\log(x)f_{\mathrm{MP}}(x)dx=-1$, so $\sum_i\log\lambda_i\approx -n$.
\end{remark}

\subsection{Entrywise Specification and the Nearest Correlation Matrix}\label{sec:nearest}

In practice, correlation matrices are often specified entrywise rather than estimated jointly: experts elicit pairwise correlations one at a time, regulators prescribe correlation shocks in stress scenarios, and pairwise estimates are assembled from different samples, time windows, or data sources. The result is a symmetric matrix $\tilde{C}$ with unit diagonal and entries in $[-1,1]$ that need not be positive semidefinite. If the $d$ off-diagonal entries are specified independently from any distribution whose density is bounded by $K$, the probability of validity is at most $K^d\operatorname{Vol}(\mathcal{E}_n)$, which vanishes super-exponentially by Proposition~\ref{prop:VolumeCn} for every fixed $K$. The failure concerns nondegenerate, dimension-invariant independent entry distributions: under the bounded-density condition above, the probability of validity vanishes super-exponentially. For the centered Wigner-type specifications considered below, scaling the entries by $o(n^{-1/2})$ restores positive definiteness with probability tending to one, as quantified at the end of this subsection. Validity at a fixed entrywise scale can only be achieved by coordinating the entries, as structured specifications such as equicorrelation do.

The standard remedy is to replace $\tilde{C}$ by the nearest correlation matrix,
$$
C^{*}=\arg\min_{C\in\mathcal{E}_n}\|\tilde{C}-C\|_F,
$$
see \citet{Higham:2002}. The following result gives a lower bound on the repair cost: since $\mathcal{E}_n$ is contained in the positive semidefinite cone, the distance from $\tilde{C}$ to $\mathcal{E}_n$ is at least as large as its distance to the cone, which is easier to compute via the Wigner semicircle law.

\begin{theorem}\label{thm:nearest}
Let $E_{ij}$, $i<j\in\mathbb{N}$, be independent and identically distributed random variables with values in $[-1,1]$, mean zero, and variance $\sigma^2>0$. For each $n$, let $\tilde{C}=I_n+E$, where $E$ is the symmetric $n\times n$ matrix with zero diagonal and off-diagonal entries $E_{ij}$, and let $C^{*}$ be the nearest correlation matrix to $\tilde{C}$ in the Frobenius norm. Then
$$
\liminf_{n\rightarrow\infty}\frac{\|\tilde{C}-C^{*}\|_F^2}{\|\tilde{C}-I_n\|_F^2}\geq\frac{1}{2}\qquad\text{almost surely}.
$$
\end{theorem}

The theorem is a statement about repair cost, not about the deletion of individual entries: the squared Frobenius distance from $\tilde{C}$ to the nearest correlation matrix is asymptotically at least one-half of the squared Frobenius norm of the off-diagonal part, $\|\tilde{C}-I_n\|_F^2$, of the original specification. The proof rests on two observations: the distance from $\tilde{C}$ to $\mathcal{E}_n$ is bounded below by the distance to the positive semidefinite cone, and the semicircle law spreads the eigenvalues of $E$ over $(-2\sigma\sqrt{n},2\sigma\sqrt{n})$, so that close to half of the eigenvalues of $\tilde{C}$ are negative and of order $\sigma\sqrt{n}$. The constant $\tfrac{1}{2}$ is an asymptotic lower bound; the unit-diagonal constraint can only increase the distance.

The PSD lower bound used in the proof is itself asymptotically exact as a bound: by symmetry of the Wigner semicircle, $\int_{-2}^{0}s^{2}\rho_{\mathrm{sc}}(s)ds=\tfrac{1}{2}$, so the PSD-projection loss converges to $\tfrac{1}{2}$ rather than merely exceeding it. Whether the unit-diagonal constraint makes the actual limiting repair cost strictly larger than $\tfrac{1}{2}$ is an open question. The finite-$n$ simulations reported in Table~\ref{tab:ncm} of the Supplement show the realized repair ratio above the finite-$n$ PSD bound at every dimension considered, with a gap that widens over the range $n\leq400$, which is suggestive but not conclusive: both quantities are still far from their limits at these dimensions. Determining the limiting ratio appears to require an analysis of the nearest-correlation projection in the random-matrix limit, likely within the framework of free probability theory, and we leave this as an open problem.

The bound also identifies $n^{-1/2}$ as the natural transition scale for the specified entries. Consider the scaled specification $\tilde{C}_n=I_n+a_nE_n$, where $E_n$ is the Wigner matrix of Theorem~\ref{thm:nearest} and $a_n>0$ is deterministic. The eigenvalues of $a_nE_n$ spread over $\pm2\sigma a_n\sqrt{n}$. If $a_n\sqrt{n}\rightarrow\infty$, the unit diagonal is asymptotically negligible relative to the perturbation, the proof of Theorem~\ref{thm:nearest} applies verbatim, and the repair cost is again at least half of $\|a_nE_n\|_F^2$. If $a_n\sqrt{n}\rightarrow0$, then $\lambda_{\min}(\tilde{C}_n)\geq1-2\sigma a_n\sqrt{n}(1+o(1))>0$ eventually, so $\tilde{C}_n$ is a valid correlation matrix with probability tending to one and the repair cost vanishes. At the critical scale $a_n=c/\sqrt{n}$ the outcome depends on the constant: for $2\sigma c<1$ the matrix is asymptotically valid, whereas for $2\sigma c>1$ a positive fraction of the spectrum is negative and the repair cost remains a positive fraction of $\|a_nE_n\|_F^2$, a fraction that grows with $c$. There is thus no single unconditional transition point, but $n^{-1/2}$ is the boundary between the two regimes. In high dimensions, entrywise elicitation and positive semidefinite repair are therefore conflicting rather than complementary steps, and coherent specification requires parameterizations that enforce validity from the outset, such as the unconstrained transformations discussed in the Introduction \citep{ArchakovHansen:Correlation}. A related positive-definiteness issue arises when a correlation matrix is hard-thresholded to obtain a sparse estimate; see \citet{GuillotRajaratnam:2012}.

\section{Conclusion}

The requirement that a correlation matrix must be positive semidefinite imposes severe, highly nonlinear constraints on its individual elements. In this paper, we have quantified exactly how restrictive these constraints become in high dimensions, using the classical volume formula for the elliptope, $\mathcal{E}_n$, together with a refined asymptotic expansion. The volume of $\mathcal{E}_n$ has logarithmic leading term $-\tfrac{1}{4}n^2\log n$, meaning the space of valid correlation matrices becomes vanishingly small compared to its enclosing hypercube; our contribution lies in the probabilistic consequences that follow from it. Furthermore, we established a tension between local and global behavior: while the elliptope concentrates sharply around the identity matrix in the entrywise max norm, it simultaneously expands globally in the Frobenius norm, placing the bulk of the volume near the singular boundary. This is matched by an exact sample-correlation representation, which implies the Marchenko--Pastur spectral limit: the uniform distribution on $\mathcal{E}_n$ coincides with the distribution of an uncentered sample correlation matrix computed from $T=n+1$ Gaussian observations, so the empirical eigenvalue distribution of a typical correlation matrix converges to the Marchenko--Pastur law with ratio one. A uniform draw is therefore asymptotically ill-conditioned, despite having uniformly small correlations. There is an interesting but imperfect analogy between $\mathcal{E}_n$ and high-dimensional unit balls; we discuss this in Section~\ref{sec:unitball} of the Supplement.

This structure is relevant for high-dimensional models in statistics, finance, and machine learning. In portfolio optimization, it explains the instability of classical minimum-variance portfolios when the dimension is comparable to the sample size. In Bayesian modeling, it clarifies how the LKJ family behaves as the dimension grows: every fixed hyperparameter already produces entrywise marginal variance of order $1/n$, growth $\eta_n\asymp n$ is needed to change the limiting spectrum, and growth $\eta_n\asymp n^2$ is needed to control the global Frobenius distance from the identity; no choice of $\eta$ avoids the dimension-induced entrywise shrinkage, which is a property of the elliptope itself. The ill-conditioning we quantified in Proposition~\ref{prop:precision} makes the remedy precise: any inference that requires inverting an unregularized high-dimensional correlation matrix operates in a regime where the average precision diverges, motivating shrinkage estimators, factor-model decompositions, or graphical-lasso regularization as structural necessities rather than optional refinements.

Beyond explaining existing phenomena, our results provide a starting point for structurally coherent null models in correlation-network analysis, where the independent-edge benchmark ignores the dependence induced by thresholding a single positive semidefinite matrix. Theorem~\ref{thm:PPP} shows that the exceedance set of a uniformly drawn correlation matrix converges to a Poisson point process, from which the Poisson edge count, triangle absence, and Gumbel maximum-correlation fluctuations all follow as corollaries; the statement specializes known point-process theory for sample correlations \citep{HeinyMikoschYslas:2021}, but the exact pairwise-independence proof extends to all LKJ parameters and yields an explicit $O(n^{-1})$ total-variation bound for finite-dimensional exceedance counts, and a more complete treatment of correlation-network inference is left to future work.

Ultimately, the super-exponential volume collapse of $\mathcal{E}_n$ means that unstructured parameterizations of correlation matrices almost never yield a valid matrix in high dimensions. Theorem~\ref{thm:nearest} makes the repair cost explicit: for bounded, centered i.i.d.\ off-diagonal specifications, replacing the specified matrix by its nearest valid correlation matrix has a squared repair cost that is asymptotically at least one-half of the squared Frobenius norm of the off-diagonal part, so post-hoc projection is a substantial reconstruction rather than a minor correction. Robust modeling therefore requires parameterizations that enforce the positive semidefinite constraint from the outset; the Generalized Fisher Transformation \citep{ArchakovHansen:Correlation} provides an elegant solution by mapping the interior of $\mathcal{E}_n$ bijectively onto an unconstrained space, enabling dimension-stable priors and unrestricted optimization without ever leaving the feasible set.