EconBase
← Back to paper

Local Optimality and Rigidity of Frobenius Tests for Dense High-Dimensional Covariance Alternatives

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

162,583 characters

Local Optimality and Rigidity of Frobenius Tests for Dense High-Dimensional Covariance Alternatives



\begin{frontmatter}
\title{Local Optimality and Rigidity of Frobenius Tests for Dense High-Dimensional Covariance Alternatives}
\runtitle{Local Optimality and Rigidity of Frobenius Tests}

\begin{aug}
\author[A]{\fnms{Peter Reinhard}~\snm{Hansen}\ead[label=e1]{[email removed]}}
\author[B]{\fnms{Werner}~\snm{Ploberger}\ead[label=e2]{[email removed]}}
\author[C]{\fnms{Chen}~\snm{Tong}\ead[label=e3]{[email removed]}}

\address[A]{Department of Economics,
University of North Carolina at Chapel Hill\printead[presep={,\ }]{e1}}

\address[B]{Department of Economics,
Washington University in St.\ Louis\printead[presep={,\ }]{e2}}

\address[C]{School of Economics,
Xiamen University\printead[presep={,\ }]{e3}}
\end{aug}

\begin{abstract}
We study identity testing for high-dimensional covariance matrices
against dense alternatives of unknown direction, with $p/n\to\gamma$.
Along a globally positive quadratic precision path, mixing Gaussian
alternatives over a Gaussian Orthogonal Ensemble direction yields a
contiguous experiment whose log likelihood reduces to the corrected
Frobenius statistic; its upper-tail test attains the limiting
weighted-power envelope at every fixed strength.  Fixing the prior's
Frobenius radius perturbs the mixture by only $O(p^{-1/2})$ in total
variation, and exact whitening carries the experiment, the statistic, and
its null law to any known null covariance.  Separately, under a
product-coordinate null, feasibility needs only $4+\eta$ moments, plus
identical distributions over time when means are estimated;
studentization and an exact degrees-of-freedom correction preserve the
local power.  A stability inequality turns near-envelope attainment into
null agreement with the Frobenius rule, so uniform noninferiority on the
typical dense bulk precludes gains at any contiguous alternative.  For
trace-matched rank-one alternatives, the corrected statistic is the first
likelihood direction when $\vartheta_n\to0$ and $n\vartheta_n\to\infty$;
at fixed strength, the log likelihood ratio in the
Onatski--Moreira--Hallin fixed-spike benchmark is governed by a richer
linear spectral statistic below the Baik--Ben Arous--P\'ech\'e threshold,
while eigenvalue separation permits cost-free largest-eigenvalue
enhancement above it.  Simulations illustrate the theory.
\end{abstract}

\begin{keyword}[class=MSC2020]
\kwdgroup[type=primary]{\kwd{62H15}}
\kwdgroup[type=secondary]{\kwd{62C15}\kwd{62E20}\kwd{60B20}}
\end{keyword}

\begin{keyword}
\kwd{High-dimensional covariance matrices}
\kwd{covariance identity testing}
\kwd{local power}
\kwd{Bayes optimality}
\kwd{Gaussian Orthogonal Ensemble}
\kwd{dense alternatives}
\end{keyword}

\end{frontmatter}

\section{Introduction}
\label{sec:introduction}

Testing restrictions on covariance and correlation matrices is a central problem in multivariate statistics. For fixed dimension $p$ and increasing sample size $n$, classical Gaussian likelihood-ratio statistics have chi-square null limits. The situation is different when $p/n$ tends to a positive constant: sample eigenvalues remain dispersed even under the identity null, and the usual likelihood-ratio chi-square calibration fails. Standard log-determinant tests also become unusable when the sample covariance is singular. Quadratic alternatives include the Frobenius criteria of \citet{John:1971} and \citet{Nagao:1973}, whose proportional-regime behavior was studied by
 \citet{LedoitWolf:2002}, with related high-dimensional tests developed by \citet{Srivastava:2005} and \citet{Schott:2005}.  For fixed $p$, John's criterion is locally most powerful invariant; Theorem~\ref{thm:goe-optimality} and Corollary~\ref{cor:composite} give the proportional-regime counterpart of that optimality.  Other approaches include
random-matrix corrections to likelihood-ratio and sphericity tests
\citep{BaiJiangYaoZheng:2009,WangYao:2013,JiangYang:2013}, and quadratic
U-statistic tests \citep{ChenZhangZhong:2010,CaiMa:2013,LiChen:2012}.
More recent work has broadened null validity, covariance structures, and
dense/sparse adaptivity: \citet{ZhengChengGuoZhu:2019} treat
high-dimensional correlation matrices directly, \citet{HanWu:2020} allow
general covariance structures with nonclassical calibration,
\citet{YuLiXue:2024} combine quadratic and maximum-type procedures to adapt
between dense and sparse signals, and \citet{Xiong:2025} develops
U-statistic identity tests.  Our contribution is complementary: we identify
a likelihood experiment, its full local-power envelope, and the rigidity of
attaining that envelope.

This literature provides reliable null calibration, minimax separation
rates, and power calculations for important alternatives.  A different
question remains: \emph{which local experiment makes a high-dimensional
Frobenius test optimal, and how unique is the result?}  A rate
result alone cannot answer this question.  Nor can pointwise power at a
selected sequence of alternatives: a quadratic test and a largest-eigenvalue
test respond to different distributions of the departure across
eigendirections, even when the two alternatives have the same Frobenius
norm.  We address the question by specifying an
isotropic dense experiment, deriving its mixture likelihood, and comparing
that likelihood directly with the feasible corrected Frobenius statistic.

The experiment uses a Gaussian Orthogonal Ensemble (GOE) matrix $H$ as the
direction of departure.  The GOE law is orthogonally invariant and its
direction is exactly uniform on the Frobenius sphere of symmetric matrices.
For fixed $\omega\ge0$, set $A=\omega H/n$ and define
\begin{equation}
\Sigma^{-1}_\omega(H)=I_p+A+A^2.
\label{eq:alternative}
\end{equation}
The path is positive definite for every realization of $H$, because
$1+x+x^2>0$ for all $x\in\mathbb{R}$.  It is spectrally dense: a nonvanishing
fraction of the eigenvalues move on the $n^{-1/2}$ scale, while the aggregate
Frobenius departure has a nondegenerate limit.  Mixing over $H$ assigns no preferred orientation to the alternatives.
This provides a dense counterpart to spiked covariance models with uniformly
distributed spike directions, which are used to study the local power of invariant eigenvalue tests,
\citep{Onatski:2009,OnatskiMoreiraHallin:2013}.
In spiked covariance models, the leading sample eigenvalue undergoes the
Baik--Ben Arous--P{\'e}ch{\'e} (BBP) phase transition: a separated sample
eigenvalue emerges only after the population spike crosses a critical
strength \citep{BaikBenArousPeche:2005,BaikSilverstein:2006}.

The first contribution is a likelihood characterization.  The integrated
likelihood follows the weighted-average-power tradition of
\citet{AndrewsPloberger:1994} and \citet{CarrascoHuPloberger:2014}.  The GOE mixture
likelihood is analytically intractable in its original form, but its
data-dependent part is only quadratic in $H$.  We use an exponential
embedding as a proof device and compare the resulting factorized likelihood
with the exact likelihood under (\ref{eq:alternative}).  Their ratio is a
function of $H$ alone, not of the sample, and the order-one fourth-order term
cancels after GOE averaging.  The resulting expansion is
\[
\log g_n(\omega)
=\omega^2 U_n-\frac{\omega^4\gamma^2}{8}+o_{P_n}(1),
\]
where $U_n$ is the corrected Frobenius martingale statistic.  Thus the
Frobenius threshold test is asymptotically equivalent to the
Neyman--Pearson test of the identity null against the GOE mixture.  Moreover,
the likelihood ratios at every fixed finite collection of strengths converge
jointly to a one-sided Gaussian location experiment.  The same test therefore
attains the mixture-power envelope at every fixed strength.

This conclusion is stronger and more specific than minimax rate
optimality.  Our statistic is closely related to the U-statistic of
\citet{ChenZhangZhong:2010}, whose Frobenius separation rate
\citet{CaiMa:2013} showed to be minimax under Gaussianity; the
contribution is an experiment in which the
corrected form is selected by the likelihood itself, together with its
asymptotic local-power envelope.  Two ingredients do the selecting:
the isotropic direction picks out the Frobenius family, and the
scale-neutral quadratic path, the unique trace-orthogonal member of
$I_p+A+cA^2$ (Proposition~\ref{prop:path-selection}), picks out the
data-driven correction.  The radial part of
the GOE prior plays no asymptotic role: replacing it by the uniform
prior on a Frobenius sphere of matched radius changes the mixture
experiment by only $O(p^{-1/2})$ in total variation, uniformly over
bounded strengths (Proposition~\ref{prop:fixed-radius}).

An exact whitening isomorphism (Lemma~\ref{lem:whitening}) carries the
Gaussian experiment, the statistic, and its pivotal null law to any
known null covariance $\Sigma_0$: the optimal statistic is the same statistic in
whitened coordinates, its null law does not depend on $\Sigma_0$, and
it uses $\Sigma_0$ only through the inner products
$x^\prime\Sigma_0^{-1}y$.  A second experiment replaces the dense
alternative by a single random factor of strength $\vartheta_n$, with
the null scale matched in trace so that the trace statistic carries no
linear mean signal.  Its character depends on the strength.
For every vanishing strength the hypotheses remain contiguous, and when
in addition $n\vartheta_n\to\infty$ the corrected Frobenius statistic is
the first nonzero direction of the likelihood, and its first-order power
gain, relative to the test's null rejection probability, agrees with the
small-shift expansion of the dense-mixture envelope
(Theorem~\ref{thm:weak-factor}).  At a fixed
strength below the Baik--Ben Arous--P\'ech\'e (BBP) threshold
$\vartheta=\sqrt\gamma$, the fixed-spike benchmark of
\citet{OnatskiMoreiraHallin:2013}, an invariant fixed-radius experiment
that remains contiguous, has a log likelihood ratio governed by a richer
linear spectral statistic, of which the corrected statistic is only the
leading quadratic term.  Above the threshold the largest eigenvalue
separates \citep{BaikBenArousPeche:2005,BaikSilverstein:2006}.  The
threshold thus bounds cost-free largest-eigenvalue enhancement within
this fixed-strength rank-one family.  The fixed-strength comparison
imports random-matrix limits and concerns that benchmark; the transfer
to the random-radius mixture is left unproved.

The second contribution makes that optimality feasible.  Under the null we
allow the coordinate distributions to be heterogeneous, asymmetric, and
non-Gaussian.  Independence, unit variances, and a uniformly bounded
$4+\eta$ moment suffice for a martingale central limit theorem.  The
observable corrections delete all coincident-time products, leaving a
degenerate distinct-time martingale whose null variance is free of trace
and fourth-moment fluctuations; the diagonal correction additionally
absorbs marginal kurtosis.  An exact scaling identity then shows that replacing sample
covariances by sample correlations is asymptotically innocuous.  When means
are estimated, an exact $n-1$ degrees-of-freedom identity gives the correct
centering, and demeaning has no first-order effect.  By contiguity, the fully
feasible statistic inherits the GOE local power
$1-\Phi(z_{1-\alpha}-\omega^2\gamma/2)$.

The third contribution is a rigidity result.  Classical weighted-likelihood
admissibility arguments are developed by \citet{AndrewsPloberger:1995}; in
the proportional regime, Bayes optimality rules out a
test that is nowhere worse almost everywhere under the mixing distribution
and strictly better on a set of positive prior probability.  It does not rule
out gains on one region that are offset by losses elsewhere.  Theorem~\ref{thm:rigidity}
does: a competing same-level test that is nowhere worse than the
invariant Frobenius benchmark uniformly over the typical dense bulk must merge
with the Frobenius test under the null and its power difference from
the Frobenius test vanishes along every contiguous sequence, so its
gains are confined to
direction sets whose GOE probability vanishes.  The premise is
transparent because the benchmark's directionwise power is
asymptotically constant over that bulk (Lemma~\ref{lem:flatness}).
We also provide a finite-$n$ stability
inequality: mixture-power regret of order $\delta$ implies null disagreement
of order $\sqrt\delta$ and, under $L^2$-bounded alternatives, power
disagreement of order $\delta^{1/4}$, in each case up to explicit
approximation terms that vanish without a claimed rate.  The
attainment argument is related to the near-optimality bounds of
\citet{ElliottMuellerWatson:2015}.

Rigidity is deliberately not called unrestricted admissibility.  In growing
dimension, power enhancement can add a detector that is asymptotically silent
under the null but powerful against selected noncontiguous signals
\citep{FanLiaoYao:2015,KockPreinerstorfer:2019,YuLiXue:2024}.  For example, a
largest-eigenvalue screen improves the known-scale Frobenius procedure at a
super-critical spike without affecting asymptotic size.  The positive and
negative statements identify the boundary: no contiguous power can be gained
without sacrificing power somewhere on the dense bulk, while improvements
remain possible on prior-negligible or noncontiguous alternatives.  This is
related to the negligibility phenomena studied by
\citet{KockPreinerstorfer:2024} in a different model; here the power
difference vanishes along all contiguous sequences.

Null calibration and alternative optimality are
separate questions.  Corrected likelihood-ratio and linear spectral
statistics recover valid null limits when $p/n$ does not vanish
\citep{BaiJiangYaoZheng:2009,JiangYang:2013,WangYao:2013}, but that does
not identify the alternatives against which a statistic is most
powerful; our likelihood calculation addresses optimality, the
product-coordinate limit theory makes the statistic feasible, and
Gaussianity is essential only for the former.

Separation-rate and local-power optimality likewise answer different
questions.  Quadratic U-statistics achieve the optimal Frobenius
separation rate \citep{ChenZhangZhong:2010,CaiMa:2013}, but at the
boundary rate alternatives with the same Frobenius norm induce different
limiting experiments: a rank-one spike concentrates its signal in a
single population eigenvalue, which yields a separated sample
eigenvalue only above the BBP threshold, while an isotropic dense
perturbation spreads it over many eigendirections.  The GOE mixture resolves this ambiguity by
specifying the angular distribution and the radial scale, and yields the
full local power curve rather than only the detection boundary.

Optimality here is deliberately weighted and local: for each
fixed strength the Neyman--Pearson lemma maximizes average power against
a simple mixture, the logic used when a nuisance parameter is present
only under the alternative
\citep{AndrewsPloberger:1994,CarrascoHuPloberger:2014}.  It implies
neither a uniformly most-powerful test nor the impossibility of
reallocating power across positive-prior regions; the rigidity theorem
obtains a directionwise conclusion only under uniform noninferiority on
a typical dense bulk, and the super-critical spike shows why.

The form of the alternative matters as well.  Against a
trace-matched weak random one-factor alternative, the corrected Frobenius
statistic is the leading likelihood direction when the factor strength
satisfies $\vartheta_n\to0$ and $n\vartheta_n\to\infty$;
at a fixed subcritical strength below the BBP threshold $\vartheta=\sqrt\gamma$
the fixed-spike benchmark stays contiguous, such that consistent
separation is impossible there, and its likelihood-optimal test is a
richer linear spectral statistic, while
above the threshold a largest-eigenvalue screen is consistent.  The rigidity
result of Section~\ref{sec:rigidity} binds on the contiguous side of this
boundary (Section~\ref{subsec:weak-factor}).

The main paper proceeds as follows: Section~\ref{sec:setup}
defines the experiment and statistic; Section~\ref{sec:likelihood} gives
the mixture likelihood, Gaussian-shift limit, local optimality, and
fixed-radius equivalence; Section~\ref{sec:implementation} establishes the
feasible null theory and local power; Section~\ref{sec:rigidity} gives
quantitative and qualitative rigidity; Section~\ref{sec:scope} delineates
scope, and Section~\ref{sec:simulation} reports the principal simulations.
The supplementary material, which follows the main text, contains the
auxiliary calculations, the admissibility, power-enhancement, and
conditional-flatness results, additional simulations, and complete
proofs.

\section{Dense alternatives and the corrected statistic}
\label{sec:setup}

Let $X_t=(X_{1t},\ldots,X_{pt})^\prime\in\mathbb R^p$, $t=1,\ldots,n$, denote
independent observations.  In the Gaussian optimality experiment,
$X_t\stackrel{\mathrm{i.i.d.}}{\sim}N_p(0,\Sigma)$.  The non-Gaussian null
theory in Section~\ref{sec:implementation} permits the marginal laws to vary
with both $i$ and $t$; temporal identical distribution is imposed only when
means are estimated.

The Gaussian experiment uses known coordinate scales and tests
\begin{equation}
H_0:\Sigma=I_p
\qquad\text{against}\qquad
H_1:\Sigma\ne I_p.
\label{eq:identity-hypothesis}
\end{equation}
Under alternatives, the marginal variances need not remain one.  When the
scales are unknown, the feasible procedure replaces the sample covariance
matrix by the sample correlation matrix and tests the scale-invariant
hypothesis that the population correlation matrix is $I_p$.  The distinction
is important: our formal likelihood-optimality statement concerns the
known-scale identity experiment, while the studentized procedure inherits its
first-order local power by an equivalence argument, and
Corollary~\ref{cor:composite} shows that the same envelope holds for
the composite unknown-scale null.

\begin{assumption}[Proportional asymptotics]
\label{ass:gamma}
$p=p(n)\to\infty$ and $p/n\to\gamma\in(0,\infty)$.
\end{assumption}

Write $\gamma_n=p/n$, such that $\gamma_n\rightarrow\gamma$.  We use
$\gamma_n$ in finite-$n$ expressions and reserve $\gamma$ for limiting
quantities.

This is the proportional regime in which the empirical spectral distribution
of $\hat R_n$ converges under the Gaussian null to the
Mar\v{c}enko--Pastur law rather than collapsing at one
\citep{MarchenkoPastur:1967}.  It is also the regime in which a classical
chi-square approximation to a fixed-dimensional likelihood ratio ceases to
describe the statistic.  The parameter $\gamma$ appears both in the null
variance and in the normalization of local alternatives, so power
comparisons across aspect ratios must specify what measure of signal strength
is held fixed.

Write
\[
\hat R_n=\frac1n\sum_{t=1}^nX_tX_t^\prime,
\qquad \hat r_{ij}=[\hat R_n]_{ij}.
\]
Under the Gaussian identity null, let $P_n$ be the law of the sample.  The
alternative direction $H$ has the GOE law
$Q_n$: the upper-triangular elements are independent, with
$H_{ij}\sim N(0,1)$ for $i<j$ and $H_{ii}\sim N(0,2)$.  Orthogonal
conjugation leaves this law unchanged.  More precisely, if
$R_p^{\mathrm{GOE}}=\|H\|_F$ and $U_p=H/\|H\|_F$, then $U_p$ is uniform on
the Frobenius sphere of the $p(p+1)/2$-dimensional symmetric-matrix space,
$U_p$ is independent of $R_p^{\mathrm{GOE}}$, and
$(R_p^{\mathrm{GOE}})^2/2$ is chi-square with $p(p+1)/2$ degrees of
freedom.

The polar representation gives a precise meaning to ``isotropic.''  The
angular component is exactly uniform over symmetric Frobenius directions and
the radius satisfies $R_p^{\mathrm{GOE}}/p=1+O_{Q_n}(p^{-1})$.  Thus the
prior is isotropic, and its spectrum is dense:
\[
p^{-1/2}\|H\|_{\mathrm{op}}\to_{Q_n}2,
\qquad
p^{-2}\operatorname{tr}(H^2)\to_{Q_n}1,
\qquad
p^{-3}\operatorname{tr}(H^4)\to_{Q_n}2.
\]
The empirical eigenvalue distribution of $p^{-1/2}H$ converges to the
semicircle law.  These facts distinguish the GOE mixture from a Haar-rotated
finite-rank spike.  They also explain why a nonvanishing fraction of the
precision eigenvalues moves at order $n^{-1/2}$, even though the aggregate
Frobenius displacement remains of constant order.

The angular component of the GOE is exactly uniform on the Frobenius
sphere.  Proposition~\ref{prop:fixed-radius} below shows that fixing its
radius at $p$ changes the induced mixture law by $O(p^{-1/2})$ in total
variation, with a bound uniform over bounded strengths.  Thus the
fixed-radius sphere prior is the canonical isotropic prior, while the
Gaussian radius is used to evaluate the mixture likelihood.  The simulations
include fixed dense spectra as an additional robustness diagnostic.

\subsection{The quadratic GOE path}
\label{sec:goe}

For fixed $\omega\ge0$, set $A=\omega H/n$ and use the path
(\ref{eq:alternative}).  Its leading Frobenius radius satisfies
\begin{equation}
\|\Sigma_{\omega}^{-1}(H)-I_p\|_F^2
=\frac{\omega^2}{n^2}\operatorname{tr}(H^2)+o_{Q_n}(1)
\rightarrow \omega^2\gamma^2.
\label{eq:frobenius-radius}
\end{equation}
Thus the limiting Frobenius radius is
$\delta=\omega\gamma$.  Holding $\omega$ fixed while increasing $\gamma$
increases this radius; holding $\delta$ fixed instead requires
$\omega=\delta/\gamma$.  This distinction later changes the local mean shift
from $\omega^2\gamma/2$ to $\delta^2/(2\gamma)$.

The experiment has two defining ingredients, and both are needed for
the selection of the corrected statistic: an isotropic direction, which
selects the Frobenius family, and a scale-neutral local path, of which
$I_p+A+A^2$ is the unique member of the family $I_p+A+cA^2$ whose
likelihood carries no trace channel (Proposition~\ref{prop:path-selection}).
The path is the globally positive representative of the additive
covariance model $\Sigma=I_p-A$: since
$(I_p+A+A^2)^{-1}-(I_p-A)=A^3(I_p+A+A^2)^{-1}$, the GOE mixture of
$N_p(0,I_p-A)$, with the prior conditioned on $\|A\|_{\mathrm{op}}\le1/2$,
differs from $G_n(\omega)$ by $O(n^{-1/2})$ in total variation,
uniformly over bounded strengths (Supplement Section~S.3).
The use of the inverse covariance is convenient because the Gaussian log
likelihood is affine in the precision matrix, and the positivity of the
path means that no truncation of the GOE is needed to define the
alternative.  The
diagonal perturbations generated by the path are smaller in aggregate than
the dense off-diagonal component, but they are retained in the exact
identity experiment.  Studentization removes their first-order relevance for
the correlation implementation.
Let $P_{n,\Sigma_{\omega}(H)}$ denote the Gaussian sample law conditional on
$H$, which is drawn first, independently of the Gaussian innovations, and
define the mixture
\begin{equation}
G_n(\omega)(B)=\int P_{n,\Sigma_{\omega}(H)}(B)\mathrm dQ_n(H),
\qquad
g_n(\omega)=\frac{\mathrm dG_n(\omega)}{\mathrm dP_n}.
\label{eq:mixture}
\end{equation}

The corrected statistic forced by this mixture is
\begin{equation}
U_n=\frac1{n^2}\sum_{i<j}\sum_{t=2}^nX_{it}X_{jt}\sum_{s<t}X_{is}X_{js}
+\frac1{2n^2}\sum_i\sum_{t=2}^n(X_{it}^2-1)\sum_{s<t}(X_{is}^2-1).
\label{eq:Tn-martingale}
\end{equation}
The first line is the off-diagonal component; the second is asymptotically
negligible.  Equivalently, $4U_n$ is the squared Frobenius distance
$\operatorname{tr}\{(\hat R_n-I_p)^2\}$ after removing the contributions from
coincident time indices.  This data-driven correction is essential outside
the Gaussian model.

For later use, the exact decomposition is worth displaying.  Write
\begin{align}
L&=2\sum_{i<j}\hat r_{ij}^{2}=C_L+S_L,
&C_L&=\frac{2}{n^2}\sum_{i<j}\sum_tX_{it}^2X_{jt}^2,
\label{eq:offdiag-decomposition}\\
D&=\sum_i(\hat r_{ii}-1)^2=C_D+S_D,
&C_D&=\frac{1}{n^2}\sum_i\sum_t(X_{it}^2-1)^2.
\label{eq:diag-decomposition}
\end{align}
Here $S_L$ and $S_D$ collect the distinct-time products, and
\begin{equation}
\operatorname{tr}\{(\hat R_n-I_p)^2\}-(C_L+C_D)=S_L+S_D=4U_n.
\label{eq:exact-correction}
\end{equation}
This identity has no asymptotic remainder.  Under Gaussianity,
$C_L+C_D$ matches the random centering produced by the mixture likelihood.
The correction matters for the variance: under coordinate independence
a deterministic off-diagonal centering is already mean-correct but
leaves trace and fourth-moment fluctuations in the null variance.
Deleting the coincident-time products removes them, and the diagonal
part of the correction absorbs marginal kurtosis.

\subsection{A feasible corrected Frobenius statistic}

The off-diagonal correction has a simple U-statistic form.  For $i<j$,
\begin{equation}
D_{ij}=\hat r_{ij}^{2}
-\frac1{n^2}\sum_tX_{it}^2X_{jt}^2
=\frac{2}{n^2}\sum_{s<t}X_{is}X_{js}X_{it}X_{jt}.
\label{eq:pairwise-u-form}
\end{equation}
Thus the null centering is achieved by deleting equal-time products rather
than estimating a marginal kurtosis.  Under temporally i.i.d.\ observations with covariance entries $\sigma_{ij}$,
\begin{equation}
\mathbb{E} D_{ij}=\frac{n-1}{n}\sigma_{ij}^2,
\qquad
\mathbb{E} Z_n=\sqrt{\tfrac{n(n-1)}{p(p-1)}}
\sum_{i<j}\sigma_{ij}^2.
\label{eq:signal-interpretation}
\end{equation}
Here $Z_n$ denotes the normalized known-scale statistic defined formally
in (\ref{eq:known-scale-statistic}).

Equation (\ref{eq:signal-interpretation}) shows what the statistic
accumulates: squared off-diagonal covariances, with
(\ref{eq:pairwise-u-form}) removing their sampling bias under the identity
null, as in high-dimensional quadratic U-statistics
\citep{ChenZhangZhong:2010,CaiMa:2013}.  The key point here is that the
same corrected statistic emerges from the GOE mixture likelihood.

The diagonal martingale has only $p$ rowwise components, compared with
$p(p-1)/2$ off-diagonal pairs, and is $o_p(1)$ after the local
normalization.  Thus, omitting it after studentization loses no
first-order information in the dense experiment.  A low-rank spike can have
the same aggregate Frobenius departure but distribute it very differently across
sample eigenvalues; this is why (\ref{eq:signal-interpretation}) does not
imply spectral optimality.

\subsection{Extension to a known null covariance}
\label{subsec:known-sigma}

The identity null extends exactly to any known positive definite null
covariance $\Sigma_0$.  Let
$Y_t=\Sigma_0^{-1/2}X_t$ be the whitened observations, and define the
known-$\Sigma_0$ corrected statistic directly in the original coordinates,
\[
U_n^{\Sigma_0}
=\frac1{2n^2}\sum_{1\le s<t\le n}h_{\Sigma_0}(X_s,X_t),
\qquad
h_{\Sigma_0}(x,y)=(x^\prime\Sigma_0^{-1}y)^2-x^\prime\Sigma_0^{-1}x-y^\prime\Sigma_0^{-1}y+p .
\]
The alternative departs from $\Sigma_0$ in the conjugated direction, through the
precision path $\Sigma_\omega^{-1}(H;\Sigma_0)
=\Sigma_0^{-1/2}(I_p+A+A^2)\Sigma_0^{-1/2}$ with $A=\omega H/n$ and $H\sim Q_n$,
such that the identity experiment of (\ref{eq:alternative}) is the case
$\Sigma_0=I_p$.

\begin{lemma}[Exact whitening isomorphism]
\label{lem:whitening}
In the Gaussian experiment, fix $n$, $p$, and $\omega$.  Almost surely and for every $n$ and $p$:
(i) under $\Sigma=\Sigma_0$ the whitened sample $(Y_t)$ has the identity-null law
$P_n$, and under the conditional alternative it has the identity-experiment law
$P_{n,\Sigma_{\omega}(H)}$;
(ii) the mixture likelihood ratios coincide,
$g_n^{\Sigma_0}(\omega)(X_1,\ldots,X_n)=g_n(\omega)(Y_1,\ldots,Y_n)$;
(iii) $U_n^{\Sigma_0}(X_1,\ldots,X_n)=U_n(Y_1,\ldots,Y_n)$, with $U_n$ the
identity-null statistic (\ref{eq:Tn-martingale}).
\end{lemma}

Because Lemma~\ref{lem:whitening} holds pathwise, every conclusion of
Theorem~\ref{thm:goe-optimality} transfers verbatim to the
known-$\Sigma_0$ null, including the local-power envelope
$1-\Phi(z_{1-\alpha}-\omega^2\gamma/2)$.  The null
law of $U_n^{\Sigma_0}$ under $\Sigma=\Sigma_0$ coincides, for every $n$ and $p$,
with that of $U_n$ under the identity null: the test is \emph{exactly pivotal} in
$\Sigma_0$, so a single set of critical values serves every known null.  The
rule depends on $\Sigma_0$ only through the inner products $x^\prime\Sigma_0^{-1}y$;
no square root is formed, and any factorization $\Sigma_0=CC^\prime$ gives the identical
value.  The whitening carries the Gaussian likelihood and pivotality theory, but
it does not automatically transfer the non-Gaussian product-coordinate central
limit theorem of Section~\ref{sec:implementation}: applying $\Sigma_0^{-1/2}$ to a
non-Gaussian vector can destroy the coordinate independence that theorem assumes.

The correction also makes the statistic exactly uncorrelated with the
null scale.
\begin{lemma}
\label{lem:trace-frobenius}
Let $W_n^{\Sigma_0}=\operatorname{tr}(\Sigma_0^{-1}\hat R_n)-p$ be the trace statistic, with
$\hat R_n=n^{-1}\sum_tX_tX_t^\prime$ (so $W_n^{I}=\operatorname{tr}(\hat R_n)-p$).  Then, under
$\Sigma=\Sigma_0$ and for every $n$ and $p$,
$\operatorname{cov}(U_n^{\Sigma_0},W_n^{\Sigma_0})=0$.
\end{lemma}

The two statistics thus provide orthogonal first-order channels: the trace
statistic $W_n^{\Sigma_0}$ responds linearly to scale perturbations, whereas the
corrected Frobenius statistic $U_n^{\Sigma_0}$ begins only at quadratic
order: under $\Sigma=c\Sigma_0$ its mean is
$\binom n2(2n^2)^{-1}p(c-1)^2$, quadratic in the scale departure $c-1$.
Deleting the coincident-time products in (\ref{eq:exact-correction}) is what
makes $U_n$ orthogonal to the trace at every $n$;
Section~\ref{subsec:weak-factor} uses this exact orthogonality in the
trace-matched weak-factor experiment.

\section{Mixture likelihood and local optimality}
\label{sec:likelihood}

For fixed $H$, the conditional Gaussian log likelihood is
\begin{equation}
\log\ell_n(H)
=\frac n2\log\det(I_p+A+A^2)
-\frac\omega2\operatorname{tr}(\hat R_nH)
-\frac{\omega^2}{2n}\operatorname{tr}(\hat R_nH^2).
\label{eq:loglr}
\end{equation}
The data-dependent part is exactly quadratic in $H$.  To exploit the
Gaussian factorization, expand
\begin{equation}
\log(I_p+A+A^2)
=A+\frac12A^2-\frac23A^3+\frac14A^4
+O(\|A\|_{\mathrm{op}}^5).
\label{eq:log-quadratic-path}
\end{equation}
The exponential matrix $\exp\{A+A^2/2-2A^3/3\}$ is used only as an
analytical bridge; it is not the alternative.  The associated factorized
approximation $\psi_n(H)$ satisfies
\[
\log\{\psi_n(H)/\ell_n(H)\}=f_n(H),
\]
where $f_n$ is independent of the data and vanishes in $Q_n$ probability on
the likelihood-relevant GOE bulk.  This data-independent comparison is the
key technical simplification.

Diagonalize $\hat R_n=O_n'D_nO_n$, write the sample eigenvalues as
$\lambda_1,\ldots,\lambda_p$, and set $W=O_nHO_n^\prime$.  Conditional on the
data, $W$ remains a GOE.  The approximation rotates to
\begin{equation}
\log\psi_n
=\frac{\omega^4\gamma_n^3}{4}
+\frac\omega2\sum_i(1-\lambda_i)W_{ii}
+\frac{\omega^2}{4n}\sum_{i,j}(1-2\lambda_i)W_{ij}^2.
\label{eq:psi-rotated}
\end{equation}
It therefore factorizes through
\begin{equation}
\mathbb{E}\exp\{aZ-bZ^2/2\}
=(1+b\tau^2)^{-1/2}
\exp\left\{\tfrac{a^2\tau^2}{2(1+b\tau^2)}\right\},
\qquad Z\sim N(0,\tau^2),\ 1+b\tau^2>0.
\label{eq:gaussian-integral}
\end{equation}
If $h_n=\int\psi_n(H)\mathrm dQ_n(H)$, direct integration and expansion
give
\begin{equation}
\log h_n=\frac{\omega^2}{4}S_n+\kappa_n+r_n,
\qquad \kappa_n\to-\frac{\omega^4\gamma^2}{8},
\qquad r_n=o_{P_n}(1),
\label{eq:hn}
\end{equation}
where
\begin{equation}
S_n
=\operatorname{tr}\{(\hat R_n-I_p)^2\}
-\tfrac{p(p+1)}{n}
-2\gamma_n\{\operatorname{tr}(\hat R_n)-p\}+o_{P_n}(1)
=4U_n+o_{P_n}(1).
\label{eq:Sn}
\end{equation}
The final equality explains both the data-driven centering and the statistic
selected by the mixture.

\begin{theorem}[Mixture likelihood and local optimality]
\label{thm:goe-optimality}
Suppose Assumption~\ref{ass:gamma} holds and $\omega>0$ is fixed.  Under the
Gaussian null $P_n$:
\begin{enumerate}
\item[(i)]
\[
\log g_n(\omega)
=\omega^2U_n-\tfrac{\omega^4\gamma^2}{8}+o_{P_n}(1),
\qquad
U_n\Rightarrow N\left(0,\tfrac{\gamma^2}{4}\right).
\]
Thus $G_n(\omega)$ is contiguous to $P_n$.
\item[(ii)] Fix $\alpha\in(0,1)$, and let $\Psi_n^{\mathrm{NP}}$ be an
exact level-$\alpha$ Neyman--Pearson test of $P_n$ against $G_n(\omega)$,
with likelihood-ratio critical value $k_{n,\alpha}$ and possible
randomization at equality.  If $s_\omega=\omega^2\gamma/2$, then
\[
k_{n,\alpha}\rightarrow
k_\alpha(\omega)
=\exp\{s_\omega z_{1-\alpha}-s_\omega^2/2\},
\qquad
\mathbb{E}_{P_n}\left|
\Psi_n^{\mathrm{NP}}-
1\{2U_n/\gamma_n>z_{1-\alpha}\}\right|\rightarrow0.
\]
By contiguity, the same $L^1$ equivalence holds under $G_n(\omega)$.
Thus the one-sided test based on $2U_n/\gamma_n$ is locally
asymptotically Bayes-optimal with GOE weights.
\item[(iii)] Let $\tau^2=\gamma^2/4$.  For fixed $m<\infty$ and fixed
$\theta_1,\ldots,\theta_m\ge0$,
\[
\left(U_n,\log g_n(\sqrt{\theta_1}),\ldots,
\log g_n(\sqrt{\theta_m})\right)
\Rightarrow
\left(U,\theta_1U-\tfrac12\theta_1^2\tau^2,\ldots,
\theta_mU-\tfrac12\theta_m^2\tau^2\right),
\]
where $U\sim N(0,\tau^2)$.  Thus every fixed finite subexperiment converges
in likelihood-ratio law to $Y\sim N(\theta\tau^2,\tau^2)$, $\theta\ge0$.
In particular, under $G_n(\omega)$,
\[
\tfrac{2U_n}{\gamma_n}\Rightarrow
N\left(\tfrac{\omega^2\gamma}{2},1\right),
\]
and the level-$\alpha$ upper-tail test has limiting power
$1-\Phi(z_{1-\alpha}-\omega^2\gamma/2)$.
\end{enumerate}
\end{theorem}

Part (iii) is a finite-dimensional statement; the expansion is not asserted
uniformly over a continuum of $\omega$.  Every statistic $U_n^\prime$ satisfying
$U_n^\prime-U_n=o_{P_n}(1)$ yields the same conclusions after the corresponding
positive rescaling and critical-value adjustment; this transfer is used
repeatedly below without further comment.

The Gaussian radial law of the GOE is a computational convenience.
Let $Q_n^{\mathrm{sph}}$ be the uniform law on the
Frobenius sphere $\{S=S^\prime:\operatorname{tr}(S^2)=p^2\}$, let
$G_n^{\mathrm{sph}}(\omega)$ be the corresponding mixture law, and let
$g_n^{\mathrm{sph}}(\omega)$ be its likelihood ratio with respect to $P_n$.

\begin{proposition}
\label{prop:fixed-radius}
Suppose Assumption~\ref{ass:gamma} holds.  For every finite $\Omega$,
\[
\sup_{0\le\omega\le\Omega}
\big\|G_n^{\mathrm{sph}}(\omega)-G_n(\omega)\big\|_{\mathrm{TV}}
=O\big(p^{-1/2}\big).
\]
\end{proposition}

\begin{corollary}
\label{cor:fixed-radius}
For every fixed $\omega>0$,
$\log g_n^{\mathrm{sph}}(\omega)=\omega^2U_n-\omega^4\gamma^2/8+o_{P_n}(1)$,
and every conclusion of Theorem~\ref{thm:goe-optimality} holds after
replacing $(G_n,g_n)$ by $(G_n^{\mathrm{sph}},g_n^{\mathrm{sph}})$ and ``GOE
weights'' by ``uniform-sphere weights''.
\end{corollary}

The proof couples the two priors
through the exact polar decomposition $H=\rho S$ and bounds the total
variation between the conditional experiments at strengths $\omega$ and
$\omega\rho$ by a Fisher-information estimate along the quadratic path;
positivity, $1+x+x^2\ge3/4$, makes the information bound uniform over the
sphere.  Thus, for every test $0\le\phi_n\le1$,
\[
\sup_{\omega\le\Omega}
\big|\mathbb{E}_{G_n^{\mathrm{sph}}(\omega)}\phi_n
-\mathbb{E}_{G_n(\omega)}\phi_n\big|
=O\big(p^{-1/2}\big).
\]

\subsection{The effective local parameter}

Although $\omega$ appears linearly in the conditional precision matrix, the
mixing distribution is symmetric under $H\mapsto-H$.  The effective
parameter of the averaged experiment is therefore
$\theta=\omega^2\ge0$.  With $\tau^2=\gamma^2/4$, the limiting log
likelihood is
\begin{equation}
\log\frac{\mathrm dP_\theta}{\mathrm dP_0}(U)
=\theta U-\frac12\theta^2\tau^2,
\qquad U\sim N(0,\tau^2)\ \text{under }P_0.
\label{eq:limit-likelihood}
\end{equation}
This is a regular Gaussian shift in $\theta$ with a one-sided parameter
space.  Under $P_\theta$, $U$ has mean $\theta\tau^2$ and unchanged
variance.  The limiting Kullback--Leibler divergence in either direction has
magnitude $\theta^2\tau^2/2$, and the likelihood-ratio threshold is
equivalent to an upper threshold in $U$ for every $\theta>0$.

Although each precision eigenvalue moves on the $n^{-1/2}$ scale, because
$\|H\|_{\mathrm{op}}=O_{Q_n}(\sqrt p)$ and $p\asymp n$, a typical
fixed-direction alternative is not contiguous to the null when its
direction is known: a matched directional test could separate it, and its
$n$-sample Kullback--Leibler divergence is of order $n$.  Contiguity is
created by averaging over the high-dimensional unknown direction, which
converts the many signed first-order movements into an order-one quadratic
score.  Thus ``local'' refers to the integrated experiment, whose
likelihood ratio has an order-one limit, rather than to each conditional
Gaussian sequence; the difficulty of the problem comes from directional
uncertainty.  In standardized units the shift is
$\theta\tau^2/\tau=\omega^2\gamma/2$; indexed by the limiting Frobenius
radius $\delta=\omega\gamma$, which facilitates comparisons at equal
aggregate signal strength, it is $\delta^2/(2\gamma)$.

\subsection{Likelihood reduction to the Frobenius statistic}

We give the essential steps because they explain why the statistic is
selected by the experiment.

\emph{Step 1: compare before integrating.}
Let $\lambda_1(H),\ldots,\lambda_p(H)$ denote the eigenvalues of $H$ and
put $\phi(u)=u+u^2/2-\log(1+u+u^2)$.
The quadratic, data-dependent terms of $\log\psi_n(H)$ and
$\log\ell_n(H)$ cancel exactly.  Hence
\begin{equation}
f_n(H)=\log\frac{\psi_n(H)}{\ell_n(H)}
=\frac n2\sum_{j=1}^p
\phi\left(\frac{\omega\lambda_j(H)}n\right)
+\frac{\omega^4\gamma_n^3}{4},
\label{eq:data-independent-comparison}
\end{equation}
which contains no sample quantity.  Since
$\phi(u)=2u^3/3-u^4/4+O(u^5)$, the GOE trace bounds give, uniformly on a
high-probability likelihood bulk,
\begin{equation}
f_n(H)=
\frac{\omega^3}{3}\frac{\operatorname{tr}(H^3)}{n^2}
-\frac{\omega^4}{8}\frac{\operatorname{tr}(H^4)}{n^3}
+\frac{\omega^4\gamma_n^3}{4}+o(1)
=o_{Q_n}(1).
\label{eq:comparison-expansion}
\end{equation}
The last equality uses $\operatorname{tr}(H^3)=O_{Q_n}(p^{3/2})$ and
$p^{-3}\operatorname{tr}(H^4)\to2$.  In particular, the fourth-order trace cancels the
explicit $\omega^4\gamma_n^3/4$ normalization.  Separately, the cubic
coefficient $-2/3$ of the exponential bridge is the value that matches
the quadratic path through third order.

The comparison also identifies which path coefficients matter.

\begin{proposition}
\label{prop:path-selection}
Fix $c>\tfrac14$ and let $g_n^{(c)}(\omega)$ be the likelihood ratio of
the GOE mixture along $\Sigma^{-1}=I_p+A+cA^2$, $A=\omega H/n$, which is
positive definite for every $H$.  Then, under $P_n$,
\[
\log g_n^{(c)}(\omega)
=\omega^2U_n
+(1-c)\,\frac{\omega^2\gamma}{2}\,\{\operatorname{tr}(\hat R_n)-p\}
+\kappa_c(\omega)+o_{P_n}(1),
\]
with $\kappa_c(\omega)=-\omega^4\gamma^2/8-(1-c)^2\omega^4\gamma^3/4$.
The two channels are exactly uncorrelated under $P_n$ and
asymptotically independent, so $c=1$ is the unique member free of the
trace channel; equivalently, by
$\Sigma_c-I_p=-A+(1-c)A^2+O(A^3)$, the unique member whose covariance
perturbation is linear in $H$ to second order.  For $c\le\tfrac14$,
including $c=0$, the same expansion holds after truncating the prior to
a $Q_n$-set of probability tending to one.
\end{proposition}

The data terms are exactly quadratic for every $c$, so only the
Gaussian-integral bookkeeping changes.  GOE isotropy selects the Frobenius family; $c=1$, singled out
by asymptotic trace matching within the quadratic family, selects the
equal-time-deleted centering, and $c=0$ the deterministic one.  Under
the positivity, localization, and remainder conditions of Supplement
Section~S.3, cubic and higher-order modifications of the path alter
the mixture by $O(n^{-1/2})$ in total variation and preserve the
limiting experiment, by the same comparison of conditional sample laws
that identifies the path with the additive covariance model.

\emph{Step 2: integrate the factorized approximation.}
Conditional on the data, orthogonal invariance makes
$W=O_nHO_n^\prime$ another GOE.  Its independent diagonal and off-diagonal
entries turn (\ref{eq:psi-rotated}) into the product
\begin{align*}
h_n={}&e^{\omega^4\gamma_n^3/4}
\prod_i \mathbb{E}\exp\left\{
\tfrac\omega2(1-\lambda_i)W_{ii}
+\tfrac{\omega^2}{4n}(1-2\lambda_i)W_{ii}^2\right\}\\
&\quad\times
\prod_{i<j}\mathbb{E}\exp\left\{
\tfrac{\omega^2}{2n}(1-\lambda_i-\lambda_j)W_{ij}^2\right\}.
\end{align*}
The one-dimensional Gaussian identity (\ref{eq:gaussian-integral}) evaluates
every factor.  Expanding the resulting logarithms and using the first three
Mar\v{c}enko--Pastur moments yields two different kinds of terms.  The
data-dependent terms combine into $\omega^2S_n/4$.  The deterministic
terms satisfy
\[
\tfrac{\omega^4\gamma_n^3}{4}
-\tfrac{\omega^4}{4}(\gamma^2+2\gamma^3)
+\tfrac{\omega^4}{4}\gamma^2(\gamma+1/2)
=-\tfrac{\omega^4\gamma^2}{8}+o(1),
\]
since $\gamma_n^3\rightarrow\gamma^3$.  All order-$\gamma^3$ pieces cancel
in the limit.  The remaining constant equals minus
one half of the limiting variance of $\omega^2U_n$, as a mean-one
lognormal likelihood ratio requires.  This calculation gives
(\ref{eq:hn}); it does not make that display a finite-sample identity.

\emph{Step 3: identify the observable centering.}
Under the Gaussian null,
\begin{equation}
C_L+C_D=
\frac1{n^2}\sum_t\Bigl[
\Bigl(\sum_iX_{it}^2\Bigr)^2+p-2\sum_iX_{it}^2\Bigr]
=\frac{p(p+1)}n
+2\gamma_n\{\operatorname{tr}(\hat R_n)-p\}+o_{P_n}(1).
\label{eq:centering-match}
\end{equation}
Combining (\ref{eq:centering-match}) with the exact decomposition
(\ref{eq:exact-correction}) proves $S_n=4U_n+o_{P_n}(1)$.  This step is
also the bridge to the non-Gaussian implementation.  The deterministic
equivalent on the right of (\ref{eq:centering-match}) uses Gaussian fourth
moments, whereas $C_L+C_D$ itself remains observable; its diagonal part
adjusts to the marginal fourth moments, and the deleted coincident-time
products no longer contribute fluctuations to the null variance.

\emph{Step 4: control the tails of the prior.}
Convergence of $f_n(H)$ in $Q_n$ probability is not enough to replace an
integral of $\ell_n$ by an integral of $\psi_n$: likelihood mass could, in
principle, concentrate on a rare set.  The Supplement constructs a bulk that
controls $\|H\|_{\mathrm{op}}$ and the trace quantities needed for the
likelihood and tail bounds, through order six, proves that
$f_n$ is uniformly bounded there, and shows that the complementary integral
is negligible.  This transfers the factorized expansion from $h_n$ to the
exact mixture $g_n$.

Finally, the martingale central limit theorem gives
$U_n\Rightarrow N(0,\gamma^2/4)$.  The limit of $g_n(\omega)$ is therefore
lognormal with expectation one, which gives contiguity by Le Cam's first
lemma.  Le Cam's third lemma supplies the mean shift under the mixture.  As
all limiting likelihood ratios are functions of the same scalar Gaussian
observation, an upper threshold in $U_n$ is the Neyman--Pearson rule for
every fixed finite subexperiment.  This completes the statistical, as
distinct from purely algebraic, part of the argument.

\section{Feasible null theory and local power}
\label{sec:implementation}

The Gaussian experiment identifies the optimal statistic.  We now separate
that experiment from the assumptions needed to implement the statistic.

\subsection{Data-driven correction and the known-scale limit}
\label{subsec:decomposition}

\begin{assumption}[Product-coordinate null]
\label{ass:moments}
The array $\{X_{it}:1\le i\le p,1\le t\le n\}$ consists of independent
random variables with
\[
\mathbb{E}X_{it}=0,\qquad \mathbb{E}X_{it}^2=1,
\qquad
\sup_{i,t}\mathbb{E}|X_{it}|^{4+\eta}<\infty
\]
for some $\eta>0$.
\end{assumption}

The marginal distributions may vary with both indices and need not be
symmetric.  The assumption retains the coordinate independence implied by
the Gaussian identity null, but permits substantially more general tails and
heterogeneity.  It does not cover dependent-but-uncorrelated elliptical or
common-scale observations.

Removing the negligible diagonal term from (\ref{eq:Tn-martingale}), define
the known-scale statistic
\begin{equation}
Z_n=
\frac{n^2}{\sqrt{p(p-1)n(n-1)}}
\sum_{i<j}\left[
\hat r_{ij}^{2}
-\tfrac1{n^2}\sum_{t=1}^nX_{it}^2X_{jt}^2
\right].
\label{eq:known-scale-statistic}
\end{equation}
The subtraction deletes all coincident-time products and leaves a
degenerate distinct-time U-statistic.  Under coordinate independence,
deterministic off-diagonal centering is already mean-correct, since
$\mathbb{E}(X_{it}^2X_{jt}^2)=1$ for $i\ne j$, but it retains trace and
fourth-moment fluctuations that alter the null variance.  The deletion
yields the universal martingale normalization used below; for the full
Frobenius criterion, the analogous diagonal correction also absorbs
marginal kurtosis.

The normalization in (\ref{eq:known-scale-statistic}) is exact.  Indeed,
with $\mathcal F_{nt}=\sigma(X_1,\ldots,X_t)$, set
\[
S_{ij,t-1}=\sum_{s<t}X_{is}X_{js},
\qquad
\Delta_{nt}=\frac1{n^2}\sum_{i<j}S_{ij,t-1}X_{it}X_{jt}.
\]
Then $U_n^{od}=\sum_{t=2}^n\Delta_{nt}$ and
$\mathbb{E}(\Delta_{nt}|\mathcal F_{n,t-1})=0$.  Independence and unit variances
make all cross-pair terms disappear from the conditional variance:
\[
\mathbb{E}(\Delta_{nt}^2|\mathcal F_{n,t-1})
=\frac1{n^4}\sum_{i<j}S_{ij,t-1}^2.
\]
Therefore
\begin{equation}
\mathbb{E}\sum_{t=2}^n\mathbb{E}(\Delta_{nt}^2|\mathcal F_{n,t-1})
=\frac{p(p-1)n(n-1)}{4n^4}
\rightarrow\frac{\gamma^2}{4}.
\label{eq:conditional-variance-mean}
\end{equation}
Only pairs sharing an index can contribute to the variance of the
conditional-variance sum $\sum_t\mathbb{E}(\Delta_{nt}^2|\mathcal F_{n,t-1})$.
Counting identical and overlapping pairs gives an $O(n^{-2})+O(n^{-1})$
bound, so the conditional variance converges in probability to the same
limit.

The moment condition is matched to the quadratic structure.  Choose
$0<\delta\le\min(1,\eta/2)$.  A conditional moment inequality for
homogeneous quadratic forms and a Marcinkiewicz--Zygmund bound for
$S_{ij,t-1}$ give
\begin{equation}
\sum_{t=2}^n\mathbb{E}|\Delta_{nt}|^{2+\delta}
\le C n^{-\delta/2}\rightarrow0.
\label{eq:martingale-lyapunov}
\end{equation}
Thus the martingale Lindeberg condition follows under the maintained
$4+\eta$ moments.  The diagonal component of $U_n$ has variance
$O(pn^2/n^4)=O(n^{-1})$ and is negligible.  Equations (\ref{eq:conditional-variance-mean}) and
(\ref{eq:martingale-lyapunov}) are the core of the null central limit
theorem.

\subsection{Studentization and demeaning}
\label{subsec:optimal-scaling}
\label{subsec:drop-diagonal}

Passing from covariances to correlations gives invariance to the units
in which each coordinate is measured, and it has a minimum-distance
interpretation.  Multiply row
$i$ by $d_i\ge0$ and write $V_i=d_i^2$.  Over diagonal rescalings
$D=\operatorname{diag}(d_1,\ldots,d_p)$, the uncorrected discrepancy
$\|I_p-D\hat R_nD\|_F^2$ is a convex quadratic in $V$, and the corrected
Frobenius criterion
\begin{equation}
Q(V)=p+V'A_nV-2b_n'V,
\qquad
[A_n]_{ij}=\hat r_{ij}^{2}
-\frac1{n^2}\sum_tX_{it}^2X_{jt}^2,
\quad
[b_n]_i=(1-n^{-1})\hat r_{ii}.
\label{eq:scaling-criterion}
\end{equation}
Moreover,
\[
A_n=\frac1{n^2}\sum_{s\ne t}
(X_s\circ X_t)(X_s\circ X_t)^\prime
\]
is positive semidefinite, so $Q$ is convex.  The uncorrected criterion has
the literal minimum-distance interpretation, while $Q$ is its
bias-corrected analogue.  At a diagonal population covariance with positive
diagonal entries, ordinary studentization is the exact population minimizer
of the uncorrected criterion.  The feasible test neither optimizes $Q$ nor evaluates it: it uses only
the off-diagonal part of the criterion at $V_i=1/\hat r_{ii}$; the
diagonal contribution retained by $Q$ there is of order one, so the
minimum-distance interpretation attaches to the full criterion, not to
$\tilde Z_n$ itself.
Theorem~\ref{thm:martingale-clt} shows that this studentization
is first-order innocuous under the product-coordinate null.

Let $\tilde X_{it}=X_{it}/\hat r_{ii}^{1/2}$ and
$\tilde r_{ij}=n^{-1}\sum_t\tilde X_{it}\tilde X_{jt}$.  The
known-mean, studentized statistic is
\begin{equation}
\tilde Z_n=
\frac{n^2}{\sqrt{p(p-1)n(n-1)}}
\sum_{i<j}\left[
\tilde r_{ij}^{2}
-\frac1{n^2}\sum_{t=1}^n\tilde X_{it}^2\tilde X_{jt}^2
\right].
\label{eq:implemented-statistic}
\end{equation}
On the event that a sample variance is zero, the corresponding
standardized row, and every ratio with that denominator, is defined to be
zero; this convention is used throughout.  Assumption~\ref{ass:moments}
implies $\inf_{i,t}\mathbb{P}(X_{it}\ne0)>0$, so independence and $p=O(n)$ make the
probability of any such row vanish exponentially.
The diagonal is omitted because sample correlations have unit diagonal.
The exact identity
\[
\tilde r_{ij}^{2}
-\frac1{n^2}\sum_t\tilde X_{it}^2\tilde X_{jt}^2
=\frac{1}{\hat r_{ii}\hat r_{jj}}
\left(\hat r_{ij}^{2}
-\frac1{n^2}\sum_tX_{it}^2X_{jt}^2\right)
\]
reduces the studentization error to a weighted sum of conditionally
mean-zero pair terms.

A maximum bound $\max_i|\hat r_{ii}-1|=o_p(1)$ is too crude after
summing $p(p-1)/2$ terms.  The decisive structure is conditional
orthogonality.  If
\[
D_{ij}=\hat r_{ij}^{2}
-\frac1{n^2}\sum_tX_{it}^2X_{jt}^2,
\]
then $\mathbb{E}(D_{ij}|X_{i1},\ldots,X_{in})=0$, and symmetrically after
conditioning on row $j$.  More strongly,
\begin{equation}
\mathbb{E}\{D_{ij}(\hat r_{jj}-1)|X_{i1},\ldots,X_{in}\}=0,
\label{eq:studentization-orthogonality}
\end{equation}
with the analogous identity after interchanging the rows.  Expanding
$(\hat r_{ii}\hat r_{jj})^{-1}$, equation
(\ref{eq:studentization-orthogonality}) removes the accumulated linear term.
The remaining nonlinear term is localized to rows whose sample variances are
close to one and controlled by truncation.  This argument is what permits the
same $4+\eta$ moment condition as the martingale central limit theorem.

When the means are unknown, put
$X_{it}^c=X_{it}-\bar X_i$, standardize by
$\hat r_{ii}^c=n^{-1}\sum_t(X_{it}^c)^2$ (with the zero convention on
$\{\hat r_{ii}^c=0\}$), and write the resulting
observations and correlations as $\tilde X_{it}^c$ and
$\tilde r_{ij}^c$.  The fully feasible statistic is
\begin{equation}
\tilde Z_n^c=
\frac{n^2}{\sqrt{p(p-1)n(n-1)}}
\sum_{i<j}\left[
(\tilde r_{ij}^c)^2
-\tfrac1{n(n-1)}\sum_{t=1}^n
(\tilde X_{it}^c)^2(\tilde X_{jt}^c)^2
\right].
\label{eq:demeaned-statistic}
\end{equation}
Under temporally i.i.d.\ sampling (Assumption~\ref{ass:time-iid} below),
the divisor $n(n-1)$ is exact.  Let
$\pi_{j,n}=\mathbb{P}(\hat r_{jj}^c>0)$, which is exponentially close to one
(quantified below).  Conditional on $\{\hat r_{jj}^c>0\}$, a standardized
centered row lies in the $(n-1)$-dimensional subspace orthogonal to the
vector of ones and has squared length $n$, and exchangeability in time
gives
\[
\mathbb{E}(\tilde X_{jt}^c\tilde X_{js}^c\mid\hat r_{jj}^c>0)
=\begin{cases}+1,&t=s,\\-1/(n-1),&t\ne s.
\end{cases}
\]
Conditioning on the independent row $i$ and using $\sum_t\tilde X_{it}^c=0$
yields the exact identities
\[
\mathbb{E}\{(\tilde r_{ij}^c)^2\mid\tilde X_i^c\}
=\mathbb{E}\Big\{\tfrac1{n(n-1)}\textstyle\sum_t
(\tilde X_{it}^c)^2(\tilde X_{jt}^c)^2\Big|\tilde X_i^c\Big\}
=\frac{\pi_{j,n}}{n-1}1\{\hat r_{ii}^c>0\},
\]
so each pairwise summand in (\ref{eq:demeaned-statistic}) is exactly
conditionally centered, and both of its terms have conditional
expectation $1/(n-1)$ on the positive-variance events (Supplement
Section~S.5).  Replacing $n(n-1)$ by the known-mean divisor $n^2$ therefore
creates a systematic finite-sample location error; the simulation section
shows that it is visible at conventional sample sizes.

\begin{assumption}[Estimated means]
\label{ass:time-iid}
For each $n$ and $i$, $X_{i1},\ldots,X_{in}$ are identically distributed.
Their laws may vary with $i$ and $n$.
\end{assumption}

No positivity or anti-concentration condition on the centered sample
variances is needed: the moment bound of Assumption~\ref{ass:moments}
already makes the zero-variance event exponentially negligible.  Let
$r=4+\eta$ and $M=\sup_{i,t}\mathbb{E}|X_{it}|^{r}$.  If an atom $a$ of a
coordinate law has mass $q>1/2$, then $|a|^{r}\le2M$, and H\"older applied
to $1\le\mathbb{E}(X-a)^2\le\{\mathbb{E}|X-a|^{r}\}^{2/r}(1-q)^{1-2/r}$
shows $q\le1-c$ for a constant $c=c(M,r)>0$.  For an i.i.d.\ row,
$\mathbb{P}(X_{i1}=\cdots=X_{in})=\sum_a\mathbb{P}(X_{i1}=a)^n
\le(1-c)^{n-1}$, so
\[
\mathbb{P}\Big(\min_{i\le p}\hat r_{ii}^c=0\Big)
\le p(1-c)^{n-1}=o(1),
\]
and the zero convention for the standardized rows on this event affects
no asymptotic statement.

\begin{theorem}
\label{thm:martingale-clt}
Suppose Assumptions~\ref{ass:gamma} and \ref{ass:moments} hold.
\begin{enumerate}
\item[(i)] Under the product-coordinate null,
\[
Z_n\Rightarrow N(0,1),
\qquad
\tilde Z_n-Z_n=o_{P_n^0}(1),
\qquad
\tilde Z_n\Rightarrow N(0,1),
\]
where $P_n^0$ denotes any null sequence satisfying the assumptions.  The
diagonal component of $U_n$ is $o_{P_n^0}(1)$.
\item[(ii)] If Assumption~\ref{ass:time-iid} also holds, then
\[
\tilde Z_n^c-\tilde Z_n=o_{P_n^0}(1),
\qquad
\tilde Z_n^c\Rightarrow N(0,1).
\]
\item[(iii)] The $o_{P_n^0}(1)$ equivalences in (i)--(ii) transfer to
every sequence contiguous to $P_n^0$.  In particular, under the Gaussian mixture
$G_n(\omega)$,
\[
\tilde Z_n^c\Rightarrow
N\left(\tfrac{\omega^2\gamma}{2},1\right).
\]
Thus rejecting when $\tilde Z_n^c>z_{1-\alpha}$ has asymptotic size
$\alpha$ and power
$1-\Phi(z_{1-\alpha}-\omega^2\gamma/2)$.
\end{enumerate}
\end{theorem}

The martingale argument was
summarized in (\ref{eq:conditional-variance-mean})--
(\ref{eq:martingale-lyapunov}).  Studentization uses exact pairwise
equivariance and conditional orthogonality;
its localization step uses a uniform sample-variance rate.  With $q=2+\eta/2$,
choose $0<a<(q/2-1)/q=\eta/(8+2\eta)$ and $b_n=n^{-a}$.
For the sample variances $\bar r_{ii}$ computed from the array truncated
at $n^{1/2}$, Rosenthal's inequality gives, uniformly in $i$,
\[
\mathbb{P}\{|\bar r_{ii}-1|>b_n\}\le Cn^{-1-\kappa},
\qquad \kappa=q/2-1-aq>0.
\]
Since $p\le Cn$, the promised union bound is
\begin{equation}
\mathbb{P}\left\{\max_{1\le i\le p}|\bar r_{ii}-1|>b_n\right\}
\le Cp n^{-1-\kappa}\le C'n^{-\kappa}\rightarrow0.
\label{eq:variance-union-bound}
\end{equation}
The truncation event, on which $\bar r_{ii}=\hat r_{ii}$ for every
$i$, has probability tending to one under the same $4+\eta$ moment
condition, so the bound transfers to the observed variances.  On the complement of
the event in (\ref{eq:variance-union-bound}), a row-wise Hoeffding decomposition and
(\ref{eq:studentization-orthogonality}) control the linear and quadratic
Taylor terms.

Demeaning uses the exact $n-1$ identity and a centered-row perturbation
bound.  Because mean subtraction is a rank-one projection in the time
dimension, its aggregate contribution is of lower order once the exact
centering has removed the deterministic degrees-of-freedom effect.  Finally,
contiguity transfers every $o_{P_n}(1)$ implementation error from the
Gaussian null to $G_n(\omega)$; Le Cam's third lemma supplies the shift of
$Z_n$.

\subsection{Computation of the feasible statistic}

The pairwise formula need not be implemented with nested loops:
(\ref{eq:demeaned-statistic}) can be evaluated from one sample
correlation matrix and two columnwise sums, at a dominant cost of
$O(np^2)$ operations for forming $\tilde R_n^c=n^{-1}YY^\prime$ from
the standardized data matrix $Y=(\tilde X_{it}^c)_{i,t}$.  The two
identities are recorded in Supplement Section~S.4, where the matrix and
pairwise implementations are also checked against each other.

The recommended workflow is therefore: demean and studentize every row,
form $\tilde R_n^c$, evaluate the two compact identities with the
$n(n-1)$ divisor, and reject for $\tilde Z_n^c>z_{1-\alpha}$.  The
standard-normal critical value is justified under
Assumptions~\ref{ass:gamma}--\ref{ass:time-iid}; other calibration is
needed when coordinate independence is not credible.

\section{Quantitative and qualitative rigidity}
\label{sec:rigidity}

Fix $\alpha\in(0,1)$ and write
\[
\varphi_n^*=1\{\tilde Z_n^c>z_{1-\alpha}\},
\qquad
\varphi_n^{U}=1\{F_n>z_{1-\alpha}\},
\qquad
\beta^*(\omega)=1-\Phi\left(z_{1-\alpha}-\frac{\omega^2\gamma}{2}\right),
\]
with $F_n=2U_n/\gamma_n$.
The benchmark $\varphi_n^{U}$ is the upper-tail test of the full
invariant statistic selected by the likelihood in
Theorem~\ref{thm:goe-optimality}, standardized to an asymptotically
unit null variance.
Under the Gaussian null both rules satisfy
$\mathbb{E}_{P_n}\varphi_n\to\alpha$ and
$\mathbb{E}_{G_n(\omega)}\varphi_n\to\beta^*(\omega)$, and
$\mathbb{E}_{P_n}|\varphi_n^*-\varphi_n^{U}|\to0$
(Theorems~\ref{thm:goe-optimality} and~\ref{thm:martingale-clt}); the
off-diagonal $Z_n$ of (\ref{eq:known-scale-statistic}) is kept for the
implementation.  Define the typical GOE bulk
\[
\mathcal K_n^{\mathrm{rig}}=
\left\{H:\ \|H\|_{\mathrm{op}}\le3\sqrt p,\quad
\left|p^{-2}\operatorname{tr}(H^2)-1\right|\le n^{-1/4}\right\},
\]
an orthogonally invariant set with $Q_n(\mathcal K_n^{\mathrm{rig}})\to1$,
and let $\nu_n(H,\omega)=P_{n,\Sigma_\omega(H)}$.

Ordinary mixture optimality is an average statement:
\[
\mathbb{E}_{G_n(\omega)}\varphi
=\int \mathbb{E}_{\nu_n(H,\omega)}\varphi\mathrm dQ_n(H).
\]
It rules out a same-size test that is nowhere worse $Q_n$-almost everywhere
and strictly better on a positive-probability set, but not gains on one
region offset by losses on another.  Condition (\ref{eq:noninferiority})
below excludes such compensation on the typical dense set; integration then
converts directionwise noninferiority into mixture-power attainment, which
forces the two rejection rules to merge under the null.

The finite-sample stability statement quantifies the last step.  It
measures the loss in the Neyman--Pearson objective rather than in raw
power, so tests of slightly different sizes are handled correctly; for
exact-size tests $\mathcal R_n(\varphi)$ is the mixture-power gap, and
$e_n$ and $\rho_n$ account for the feasible-versus-exact
Neyman--Pearson equivalence and the finite-sample likelihood
approximation.

For the quantitative statement, let $\Psi_n^{\mathrm{NP}}$ be an exact
level-$\alpha$ Neyman--Pearson test of $P_n$ against $G_n(\omega)$, with
threshold $k_{n,\alpha}$, and set
\[
\mathcal R_n(\varphi)=
\mathbb{E}_{G_n(\omega)}(\Psi_n^{\mathrm{NP}}-\varphi)
-k_{n,\alpha}\mathbb{E}_{P_n}(\Psi_n^{\mathrm{NP}}-\varphi).
\]
If $F_\omega$ and $f_\omega$ are the cdf and density of the limiting
lognormal likelihood ratio in Theorem~\ref{thm:goe-optimality}, define
\[
\rho_n(\omega)=\sup_x|P_n\{g_n(\omega)\le x\}-F_\omega(x)|,
\qquad
e_n(\omega)=\mathbb{E}_{P_n}|\Psi_n^{\mathrm{NP}}-\varphi_n^*|.
\]

Throughout, a sequence of tests $\varphi_n$ has \emph{asymptotic level}
$\alpha$ if $\limsup_n\mathbb{E}_{P_n}\varphi_n\le\alpha$.  The first
statement is a general Neyman--Pearson regret inequality; it does not use
the covariance structure.

\begin{proposition}[Neyman--Pearson regret inequality]
\label{prop:np-regret}
Fix $\omega>0$.  For any test $0\le\varphi\le1$ and every $\varepsilon>0$,
\begin{equation}
\mathbb{E}_{P_n}|\varphi-\varphi_n^*|
\le e_n(\omega)+2\|f_\omega\|_\infty\varepsilon+2\rho_n(\omega)
+\frac{\mathcal R_n(\varphi)}{\varepsilon},
\label{eq:quantitative-stability}
\end{equation}
where $e_n(\omega)\to0$, $\rho_n(\omega)\to0$, and
$\mathcal R_n(\varphi)\ge0$.  Optimizing gives
\begin{equation}
\mathbb{E}_{P_n}|\varphi-\varphi_n^*|
\le e_n(\omega)+2\rho_n(\omega)
+2\sqrt{2\|f_\omega\|_\infty\mathcal R_n(\varphi)}.
\label{eq:quantitative-stability-optimized}
\end{equation}
If $\nu_n\ll P_n$ and
$\sup_n\mathbb{E}_{P_n}(\mathrm d\nu_n/\mathrm dP_n)^2\le K$, then
\begin{equation}
|\mathbb{E}_{\nu_n}\varphi-\mathbb{E}_{\nu_n}\varphi_n^*|
\le K^{1/2}\{\mathbb{E}_{P_n}|\varphi-\varphi_n^*|\}^{1/2}.
\label{eq:quantitative-transfer}
\end{equation}
For an exact-size test, $\mathcal R_n$ is its mixture-power deficiency.
\end{proposition}

\begin{theorem}[Rigidity of dense-mixture optimality]
\label{thm:rigidity}
Fix $\omega>0$.  Let $\varphi_n$ have asymptotic level $\alpha$ and suppose it is
asymptotically nowhere worse than the invariant benchmark
$\varphi_n^{U}$ uniformly on the GOE bulk:
\begin{equation}
\liminf_n\inf_{H\in\mathcal K_n^{\mathrm{rig}}}
\{\mathbb{E}_{\nu_n(H,\omega)}\varphi_n
-\mathbb{E}_{\nu_n(H,\omega)}\varphi_n^{U}\}\ge0.
\label{eq:noninferiority}
\end{equation}
Then
\[
\mathbb{E}_{P_n}\varphi_n\to\alpha,
\qquad
\mathbb{E}_{P_n}|\varphi_n-\varphi_n^*|\to0.
\]
For every sequence $\nu_n$ contiguous to $P_n$,
\[
\mathbb{E}_{\nu_n}\varphi_n-\mathbb{E}_{\nu_n}\varphi_n^*\to0.
\]
Moreover, for each $\epsilon>0$,
\[
Q_n\left\{H\in\mathcal K_n^{\mathrm{rig}}:
\mathbb{E}_{\nu_n(H,\omega)}\varphi_n
-\mathbb{E}_{\nu_n(H,\omega)}\varphi_n^{U}>\epsilon\right\}\to0.
\]
\end{theorem}

The premise (\ref{eq:noninferiority}) is interpretable because the
benchmark's directionwise power is asymptotically constant across the
bulk.

\begin{lemma}[Conditional flatness of the invariant benchmark]
\label{lem:flatness}
Fix $\omega>0$ and $\alpha\in(0,1)$.  Then
\[
\sup_{H\in\mathcal K_n^{\mathrm{rig}}}
\big|\mathbb{E}_{\nu_n(H,\omega)}\varphi_n^{U}-\beta^*(\omega)\big|
\rightarrow0 .
\]
\end{lemma}

By Lemma~\ref{lem:flatness},
condition (\ref{eq:noninferiority}) says precisely that the
competitor's directionwise power on the typical dense bulk is
asymptotically nowhere below the constant level $\beta^*(\omega)$.
The proof couples $X_t=\Sigma^{1/2}Y_t$ to a null sample and uses the
exact decomposition of the invariant kernel; the full statistic is
essential, since the off-diagonal $Z_n$ is blind to diagonal-heavy
directions such as $H=\sqrt p\,\operatorname{diag}(\pm1,\ldots,\pm1)$, on which its
power stays at $\alpha$ (Supplement Section~S.6).  Null
merging in the theorem is stated relative to the feasible rule
$\varphi_n^*$; the directionwise comparison uses
$\varphi_n^{U}$.  The uniform form of (\ref{eq:noninferiority}) is
chosen for interpretability, not out of necessity: the proof uses only
its integral against $Q_n$, so the premise may be weakened to the
corresponding integrated-shortfall condition with the same
conclusions.  We retain the uniform statement because
Lemma~\ref{lem:flatness} gives it a directionwise reading that the
averaged condition lacks.

\begin{corollary}
\label{cor:envelope}
For every asymptotic-level-$\alpha$ test $\varphi_n$ and every
fixed $\omega>0$,
\[
\limsup_n\mathbb{E}_{G_n(\omega)}\varphi_n\le\beta^*(\omega),
\]
and the same test $\varphi_n^*$ attains the bound for each fixed
$\omega>0$.  This assertion is pointwise in $\omega$; it does not claim
uniform convergence over a continuum of strengths.
\end{corollary}

\begin{corollary}
\label{cor:composite}
Let the composite Gaussian null be
$\{N_p(0,D)^{\otimes n}:D$ diagonal, positive definite$\}$, and let
$G_n^{D}(\omega)$ be the law of $\{D^{1/2}X_t\}_{t\le n}$ with
$\{X_t\}\sim G_n(\omega)$.
(i) $\varphi_n^*$ is invariant under $X_{it}\mapsto d_iX_{it}$,
$d_i>0$; hence for every $n$ its rejection probability is constant in
$D$ over the null and over each orbit, so
$\sup_D\mathbb{E}_{N_p(0,D)^{\otimes n}}\varphi_n^*\to\alpha$ and
$\mathbb{E}_{G_n^{D}(\omega)}\varphi_n^*\to\beta^*(\omega)$ for
every $D$.
(ii) Any test sequence with
$\limsup_n\sup_D\mathbb{E}_{N_p(0,D)^{\otimes n}}\varphi_n\le\alpha$
satisfies
$\limsup_n\mathbb{E}_{G_n^{D_n}(\omega)}\varphi_n\le\beta^*(\omega)$
for every sequence of diagonal $D_n$ and every $\omega>0$.
The unknown-scale problem thus carries the same envelope, and the
feasible test attains it: adaptation to unknown scales is
asymptotically cost-free.
\end{corollary}

The proof is a change of variables carrying
the composite level to the simple null and the scaled mixture to
$G_n(\omega)$, followed by Corollary~\ref{cor:envelope}; attainment is
exact scale invariance of $\tilde Z_n^c$.

An exact decision-theoretic benchmark stands behind the asymptotic
argument: truncating the mixing law to
$\{H:\|H\|_{\mathrm{op}}\le3\sqrt p\}$ yields an atomless
compact-mixture likelihood ratio whose level-$\alpha$ Neyman--Pearson
rule is almost surely unique and admissible for the corresponding
finite-dimensional family (Supplement Section~S.6).  The benchmark
concerns the compact-mixture rule, which
Theorem~\ref{thm:goe-optimality} makes asymptotically equivalent to the
Frobenius threshold; admissibility rules out a competitor that
dominates everywhere but allows reallocation across directions; the
final conclusion of Theorem~\ref{thm:rigidity} strengthens this only
under the explicit uniform-noninferiority premise.

\subsection{Why regret controls disagreement}

Write $g_n=g_n(\omega)$ and $k_n=k_{n,\alpha}$.  Apart from immaterial
randomization on the threshold, the Neyman--Pearson rule satisfies
\begin{equation}
\mathcal R_n(\varphi)
=\mathbb{E}_{P_n}\{(g_n-k_n)(\Psi_n^{\mathrm{NP}}-\varphi)\}
=\mathbb{E}_{P_n}\{|g_n-k_n||\Psi_n^{\mathrm{NP}}-\varphi|\}.
\label{eq:regret-identity}
\end{equation}
The second equality holds because the signs of the two factors agree on
both sides of the rejection threshold.  Splitting the disagreement event
at $|g_n-k_n|=\varepsilon$, the complement contributes at most
$\mathcal R_n(\varphi)/\varepsilon$ by (\ref{eq:regret-identity}),
while in the band the likelihood-ratio cdf convergence and the bounded
lognormal density give
\[
P_n\{|g_n-k_n|\le\varepsilon\}
\le2\|f_\omega\|_\infty\varepsilon+2\rho_n(\omega).
\]
Adding the approximation error between $\Psi_n^{\mathrm{NP}}$ and
$\varphi_n^*$ proves (\ref{eq:quantitative-stability}); optimizing in
$\varepsilon$ gives the square-root rate, and a Cauchy--Schwarz change
of measure gives (\ref{eq:quantitative-transfer}).  A regret of order
$\delta$ thus entails null disagreement of order $\delta^{1/2}$ up to
the additive term $a_n=e_n(\omega)+2\rho_n(\omega)$ and, under an $L^2$
likelihood-ratio bound, power disagreement of order $\delta^{1/4}$ up
to $(Ka_n)^{1/2}$: the bounds are finite-$n$ moduli in $\delta$ whose
approximation terms vanish without a claimed rate, not single-sequence
rates in $\delta$ alone.  With $\alpha_n=\mathbb{E}_{P_n}\varphi$ and
$\Delta_n=\mathbb{E}_{G_n(\omega)}(\Psi_n^{\mathrm{NP}}-\varphi)$ the
regret is $\mathcal R_n(\varphi)=\Delta_n+k_{n,\alpha}(\alpha_n-\alpha)$,
so it equals the mixture-power deficiency for an exact-size test and is
bounded by it when $\alpha_n\le\alpha$.

For the qualitative part, integrating (\ref{eq:noninferiority}) over
$\mathcal K_n^{\mathrm{rig}}$, whose GOE probability tends to one,
shows that the mixture power of $\varphi_n$ is asymptotically no
smaller than that of $\varphi_n^{U}$, which attains
$\beta^*(\omega)$ by Theorem~\ref{thm:goe-optimality}; the envelope
forces equality.  The attainment argument, the zero-regret limit of
(\ref{eq:regret-identity}), then gives
$\mathbb{E}_{P_n}|\varphi_n-\varphi_n^*|\to0$; contiguity and
boundedness extend this to equality of limiting power under every
contiguous sequence; and a fixed positive gain on a GOE set of
nonvanishing probability would contradict mixture-power equality, which
proves the negligible-direction-set conclusion.

\subsection{Why the contiguity boundary matters}

The restriction to contiguous alternatives is substantive.  Consider in the
known-scale Gaussian experiment the equicorrelation matrix
\[
R_p=(1-\rho_p)I_p+\rho_p\mathbf1\mathbf1^\prime,
\qquad \rho_p=\tfrac{\vartheta}{p-1},
\qquad \vartheta>\sqrt\gamma.
\]
The leading eigenvalue of $R_p$ is $1+\vartheta$, whereas the remaining
$p-1$ eigenvalues equal $1-\vartheta/(p-1)$ and converge to one.  Thus
$\vartheta$ is the asymptotic excess of the population spike above the bulk,
and $\vartheta=\sqrt\gamma$ is the BBP sample-eigenvalue separation
threshold \citep{BaikBenArousPeche:2005}.  The corresponding almost-sure
outlier limits for general real and complex spiked covariance models are
given by \citet{BaikSilverstein:2006}.  Because
$\vartheta>\sqrt\gamma$, this is a super-critical rank-one spike.
The corrected Frobenius statistic has the nondegenerate limit
$Z_n\Rightarrow N(\vartheta^2/(2\gamma),1)$,
so its limiting power is strictly below one.  In contrast, the largest
sample eigenvalue separates from the null edge and is consistently detectable
\citep{Johnstone:2001,BaiSilverstein:2010}.  If $(1+\sqrt\gamma)^2<m<(1+\vartheta)(1+\gamma/\vartheta)$,
then adjoining the screen
$1\{\lambda_{\max}(\hat R_n)>m\}$ to the known-scale Frobenius rejection
rule leaves asymptotic null size unchanged and raises power at this spike to
one.  Fixed dense directions can likewise be detected by matched filters.

These alternatives are noncontiguous and negligible under the GOE mixing
law.  The enhancement therefore does not contradict Theorem~\ref{thm:rigidity}:
a procedure that preserves power uniformly over the dense bulk cannot gain
against a contiguous sequence, but unrestricted asymptotic admissibility is
false.  The formal finite-sample compact-mixture admissibility statement,
attainment lemma, and enhancement proof are in Supplement Section~S.6 and
Appendix~G.  The spike claim is deliberately stated for the known-scale
statistic; its fully studentized noncontiguous analogue would require a
separate argument.

\subsection{The weak-factor experiment and the BBP boundary}
\label{subsec:weak-factor}

The super-critical spike above singles out one direction in which the Frobenius
rule can be improved.  A second experiment locates that boundary:
exactly in the fixed-radius benchmark below, one-sidedly for the
random-radius experiment studied here.  Couple
a strength $\vartheta=\vartheta_n>0$ to both hypotheses and match the null scale
to the alternative in trace:
\[
H_0:\ \Sigma=\Big(1+\frac{\vartheta}{p}\Big)\Sigma_0
\qquad\text{against}\qquad
H_1:\ \Sigma=\Sigma_0^{1/2}(I_p+ff^\prime)\Sigma_0^{1/2},
\quad f\sim N_p\Big(0,\frac{\vartheta}{p}I_p\Big),
\]
where $f$ is drawn first, independently of the Gaussian innovations, so
that conditional on $f$ the observations are i.i.d.\ $N_p(0,\Sigma)$, and
$p/n\to\gamma$.  Here $\vartheta$ is the
spike-excess parameter of the equicorrelation example above: the alternative is a
single random factor with Frobenius radius
$\|ff^\prime\|_F=\|f\|^2=\vartheta(1+o_P(1))$, while the null spends the \emph{same}
expected trace budget
$\operatorname{tr}\{(1+\vartheta/p)I_p\}=p+\vartheta=\mathbb{E}\operatorname{tr}(I_p+ff^\prime)$
isotropically.  The matching equates the expected trace statistic
under the null and the mixture alternative, removing its linear mean
signal exactly, and the trace and corrected Frobenius statistics are
exactly uncorrelated under the null (Lemma~\ref{lem:trace-frobenius}).
These identities concern means and covariances; the full trace laws
differ at second order.  Within the decay regime of
Theorem~\ref{thm:weak-factor}, the corrected Frobenius statistic
carries the first nonzero likelihood direction.  Write
$V_n=2U_n^{(1+\vartheta/p)\Sigma_0}/\gamma_n$ for the standardized corrected
Frobenius statistic of Section~\ref{subsec:known-sigma} evaluated at the
trace-matched null $(1+\vartheta/p)\Sigma_0$; by Lemma~\ref{lem:whitening} it is
exactly pivotal and $V_n\Rightarrow N(0,1)$ under the null law $P_0$, the sample
law under $H_0$.  Let $G_n$ denote the marginal law of the sample under $H_1$,
mixed over the random factor $f$.

\begin{theorem}[Weak-factor decay regime]
\label{thm:weak-factor}
Let $p/n\to\gamma\in(0,\infty)$ in the trace-matched weak-factor experiment, and
let $P_0$ be the sample law under $H_0$.  If $\vartheta_n\to0$, then $G_n$ is
contiguous to $P_0$.  If in addition $n\vartheta_n\to\infty$, then, with
$\sigma_n=\vartheta_n^2/(2\gamma_n)$,
\[
\frac{\mathrm dG_n}{\mathrm dP_0}=1+\sigma_nV_n+o_{L^1(P_0)}(\sigma_n),
\qquad
\|G_n-P_0\|_{\mathrm{TV}}=\frac{\sigma_n}{\sqrt{2\pi}}+o(\sigma_n),
\]
with $V_n\Rightarrow N(0,1)$ under $P_0$.  Consequently, for every
$\alpha\in(0,1)$ and every sequence of tests $0\le\psi_n\le1$ with
$\mathbb{E}_{P_0}\psi_n\to\alpha$,
\[
\limsup_{n\to\infty}\sigma_n^{-1}\{\mathbb{E}_{G_n}\psi_n-\mathbb{E}_{P_0}\psi_n\}
\le\Phi^\prime(z_{1-\alpha}),
\]
with equality for $\psi_n=1\{V_n>z_{1-\alpha}\}$, where $\Phi^\prime$ is the
standard normal density: the upper-tail Frobenius test maximizes the
first-order power gain, which is of exact order $\sigma_n$.
\end{theorem}

The decay-regime statement is exact and self-contained (Supplement
Section~S.8).  The expansion is stated in $L^1(P_0)$, which is the
strongest norm available for the true likelihood ratio: its untruncated
second moment is infinite for every $n$, so no $L^2$ expansion can hold
for $\mathrm dG_n/\mathrm dP_0$ itself.  Supplement Section~S.8
establishes the $L^2$ expansion for a radially truncated surrogate and
bounds the gap to the true ratio by $2e^{-p/8}=o(\sigma_n)$ in
$L^1(P_0)$; the $L^1$ form suffices for the total-variation formula and
for the power statements, which concern bounded tests only.  Its shift $\sigma_n=\vartheta_n^2/(2\gamma_n)$ is the
radius-parametrized shift $\delta^2/(2\gamma)$ of Section~\ref{sec:likelihood} at
the factor's Frobenius radius $\delta=\vartheta_n$, the trace matching having
removed the isotropic contribution.  The first-order power gain,
measured relative to the test's null rejection probability, therefore
agrees with the small-shift expansion of the dense-mixture envelope,
and the experiment selects the same $V_n$ as the GOE mixture even
though the alternative is spiked: $V_n$ is the first nonzero direction
of the likelihood in the stated regime $\vartheta_n\to0$ with
$n\vartheta_n\to\infty$.  The first-order gain transfers to the
feasible test.

\begin{corollary}
\label{cor:feasible-weak-factor}
Under the conditions of the expansion in Theorem~\ref{thm:weak-factor},
let $\psi_n^F=1\{\tilde Z_n^c>z_{1-\alpha}\}$ be the fully feasible
test computed from $\Sigma_0^{-1/2}X_t$.  Then
\[
\mathbb{E}_{G_n}\psi_n^F-\mathbb{E}_{P_0}\psi_n^F
=\sigma_n\Phi^\prime(z_{1-\alpha})+o(\sigma_n).
\]
\end{corollary}

The feasible test thus attains the first-order gain, measured relative
to its own null rejection probability.  The proof uses only that the
gain is linear in $\mathrm dG_n/\mathrm dP_0-1$: with
$L_n=1+\sigma_nV_n+r_n$, the gains of two tests differ by at most
$\sigma_n(\mathbb{E}_{P_0}V_n^2)^{1/2}
(\mathbb{E}_{P_0}|\psi_n^F-\varphi_n^V|)^{1/2}+\mathbb{E}_{P_0}|r_n|$,
where $\varphi_n^V=1\{V_n>z_{1-\alpha}\}$, and
Theorem~\ref{thm:martingale-clt} makes the middle factor vanish.  No
rate for the feasible approximation is needed, and no claim is made
that the nominal-size error is $o(\sigma_n)$.

At a fixed strength $\vartheta_n\to\vartheta\in(0,\infty)$ we compare the
experiment with the invariant fixed-radius spiked experiment of
\citet{OnatskiMoreiraHallin:2013}, the fixed-spike benchmark, whose
conclusions we import rather than reprove: conditional on the radius $\|f\|^2\to_p\vartheta$, the factor
direction is Haar-uniform, so the conditional alternative is theirs at
strength $\|f\|^2$, tested against the trace-matched null; the transfer
of their conclusions to the random-radius mixture is not proved here
(Supplement Section~S.8.5).  In their subcritical experiment,
$\vartheta<\sqrt\gamma$, the hypotheses remain contiguous and the
asymptotically optimal test is not $V_n$ but a likelihood-based
\emph{linear spectral statistic}: with $Y_t=\Sigma_0^{-1/2}X_t$ and
$\hat S_Y=n^{-1}\sum_tY_tY_t^\prime$, a centered sum
$\sum_j\log\{z_0(\vartheta)-\bar\lambda_j\}$ of the eigenvalues
$\bar\lambda_j$ of the trace-matched matrix
$(1+\vartheta_n/p)^{-1}\hat S_Y$, with the trace correction
$\operatorname{tr}\{(1+\vartheta_n/p)^{-1}\hat S_Y\}-p$, organized around the saddlepoint
$z_0(\vartheta)=(1+\vartheta)(\gamma+\vartheta)/\vartheta$.  The corrected
Frobenius statistic $V_n$ is only the leading quadratic term of this statistic as
$\vartheta\downarrow0$, so it is first-order optimal in the vanishing-strength
limit and is generally suboptimal at fixed subcritical strength.  For a
supercritical strength $\vartheta>\sqrt\gamma$ contiguity fails in the
trace-matched experiment itself: by the
Baik--Ben Arous--P\'ech\'e and Baik--Silverstein spike phase transition
\citep{BaikBenArousPeche:2005,BaikSilverstein:2006}, the largest whitened
eigenvalue separates,
$\lambda_{\max}\to(1+\vartheta)(1+\gamma/\vartheta)>(1+\sqrt\gamma)^2$, while the
null edge stays at $(1+\sqrt\gamma)^2$, so a largest-eigenvalue test has size
$\to0$ and power $\to1$.  We do not analyze the critical case
$\vartheta=\sqrt\gamma$.

This makes the escape clause of Section~\ref{sec:rigidity} precise:
cost-free enhancement by an asymptotically null largest-eigenvalue screen
is possible above $\vartheta=\sqrt\gamma$, the regime of the
equicorrelation example and Supplement Section~S.6.  The supercritical
statement is proved for the random-radius experiment itself; that the
boundary is exact, with no cost-free enhancement below $\sqrt\gamma$,
holds in the imported fixed-radius benchmark.  There a richer
linear spectral statistic may improve power at a fixed spiked alternative,
but Theorem~\ref{thm:rigidity} implies that such a gain cannot coexist with
uniform noninferiority over the dense bulk.

\section{Scope of the theory}
\label{sec:scope}

The three main results concern different probability models.
Theorem~\ref{thm:goe-optimality} is a likelihood statement for known-scale
Gaussian covariance identity testing.  Theorem~\ref{thm:martingale-clt} is a
null and implementation result for correlation identity testing under
independent, possibly non-Gaussian coordinates.  Contiguity links the two by
transferring the feasible-statistic equivalence to the GOE mixture.  It does
not convert the non-Gaussian null class into a non-Gaussian likelihood
optimality experiment.

The angular law and the radial scale are both essential.  GOE normalization
gives $p^{-2}\operatorname{tr}(H^2)\to1$ and spreads the perturbation over order $p$
eigendirections.  Thus the mixture retains the aggregate Frobenius departure but
averages away signed first-order movements, leaving the quadratic parameter
$\theta=\omega^2$.  A different directional prior, even at the same limiting
Frobenius radius, need not have the same likelihood reduction or power
envelope.  The fixed-spectrum simulations in Section~\ref{sec:simulation}
are therefore robustness checks.

The dense and spiked experiments mark another boundary: at equal Frobenius
norm a rank-one spike and a GOE perturbation can reverse the ordering of
Frobenius and largest-eigenvalue tests.  The super-critical spike in
Section~\ref{sec:rigidity} is noncontiguous and has asymptotically
negligible GOE weight, so power enhancement there
\citep{FanLiaoYao:2015,YuLiXue:2024} is consistent with rigidity on the
typical dense bulk.

The whitening isomorphism of
Section~\ref{subsec:known-sigma} is exact at every $n$ and $p$: it
carries every optimality and pivotality statement to an arbitrary known null
covariance $\Sigma_0$, requiring only the inner products
$x^\prime\Sigma_0^{-1}y$.  The trace-matched weak-factor experiment of
Section~\ref{subsec:weak-factor} is asymptotic: contiguity holds for
every vanishing strength, the first-order power gain matches the
small-shift expansion of the dense-mixture envelope under
the additional rate $n\vartheta_n\to\infty$
(Theorem~\ref{thm:weak-factor}), and the BBP threshold
$\vartheta=\sqrt\gamma$ is the boundary for cost-free
largest-eigenvalue enhancement in the fixed-strength rank-one family,
exactly so in the fixed-radius benchmark and from above only in the
random-radius experiment;
the subcritical fixed-strength picture rests on imported random-matrix
limits and on the benchmark comparison of
Section~\ref{subsec:weak-factor}, and the critical case
$\vartheta=\sqrt\gamma$ is not analyzed.

The product-coordinate null is substantive: the observable centering
deletes the coincident-time products, so the null
variance is free of heterogeneous trace and fourth-moment fluctuations,
and the $4+\eta$ condition controls the martingale and studentization
remainders, but coordinate independence supplies the conditional
orthogonality used throughout.  Dependent-but-uncorrelated scale mixtures can
have identity covariance while violating the variance formula; the final
simulation makes this boundary visible.  Extending the feasible theory to
such models requires a new long-run or cross-sectional variance analysis,
not merely a different critical value.

\section{Simulation evidence}
\label{sec:simulation}

The experiments assess five observable implications of the theory: null
calibration under the product-coordinate model; convergence of the
studentization and demeaning errors; the exact $n-1$ centering effect; the
GOE local-power formula and its dependence on $\gamma$; and the contrast
between dense and spiked alternatives.  Two further descriptive stress
tests, one varying the dense spectrum and the other violating coordinate
independence through a common random scale, are reported in Supplement
Section~S.9.

The primary procedure is the fully feasible statistic
(\ref{eq:demeaned-statistic}).  Every row is demeaned and studentized, the
diagonal is omitted, and the test rejects when
$\tilde Z_n^c>z_{0.95}$.  For targeted comparisons we also compute the
known-scale statistic $Z_n$, its known-mean studentized version
$\tilde Z_n$, and a deliberately naive demeaned statistic that retains
the known-mean divisor $n^2$ instead of $n(n-1)$.  The naive version
isolates the consequence of ignoring the residual degree of freedom.

The deterministically centered statistic
\[
Z_n^{\mathrm{det}}=
\frac{n^2}{\sqrt{p(p-1)n(n-1)}}
\left\{\sum_{i<j}(\tilde r_{ij}^c)^2-
\tfrac{1}{n-1}\binom{p}{2}\right\}
\]
isolates the role of the data-driven correction.  A largest-eigenvalue
test targets low-rank signals; its statistic is the largest eigenvalue
of the demeaned, studentized sample correlation matrix, the feasible
analogue of the covariance spectrum in the enhancement discussion of
Section~\ref{sec:rigidity}; because raw Tracy--Widom critical values can
be conservative at the displayed dimensions, its critical value is the
corresponding quantile from an independent Gaussian null simulation of
the same statistic at the same $(n,p)$.  The two tests are thus
calibrated under the Gaussian null and approximately size matched;
the spectral power standard errors are conditional on the realized
calibration threshold.

Unless stated otherwise, results use $10{,}000$ evaluation replications at
nominal level $0.05$ and an independent $10{,}000$-replication sample for
spectral calibration.  The exact-GOE implementation comparison also uses
$10{,}000$ replications.  Common random numbers are used across signal
strengths within a design.  Monte Carlo standard errors are retained in the
replication outputs; with $10{,}000$ replications, the largest possible
standard error is $0.0050$, and the standard error at rejection
probability $0.05$ is about $0.0022$.

\subsection{Null approximation}

Table~\ref{tab:adm-size-focused} uses $n=200$ and
$p/n\in\{0.5,1,2\}$.  The Gaussian, standardized $t_{10}$, $t_8$, and $t_5$,
and centered-and-standardized skewed chi-square marginals all satisfy
Assumption~\ref{ass:moments}; the $t_5$ design lies close to its moment
boundary.  Among these covered designs, the corrected-test rejection
frequencies range from $0.045$ to $0.053$, with no systematic deterioration
under heavy tails, skewness, or the larger aspect ratio.  The standardized
$t_3$ row is an out-of-assumption stress test: it has unit variance but an
infinite fourth moment.  In that row, the corrected test is mildly
conservative, with rejection frequencies from $0.036$ to $0.039$, whereas
the deterministic-centering and Gaussian-calibrated largest-eigenvalue tests
overreject.  These $t_3$ results are finite-sample diagnostics and do not
extend the theorem beyond its moment condition.  For the designs covered by
Assumption~\ref{ass:moments}, deterministic centering also performs reasonably
because $\mathbb{E}(X_{it}^2X_{jt}^2)=1$ for $i\ne j$; the common-scale experiment
below shows why that observation does not extend to dependent squares.  The
independently calibrated largest-eigenvalue test is included to put the later
power comparison on an approximately size-matched footing.

\begin{table}[t]
\centering
\caption{Monte Carlo size under product-coordinate nulls}
\label{tab:adm-size-focused}
\footnotesize
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}lccccccccc}
\toprule
&\multicolumn{3}{c}{$\gamma=0.5$}
&\multicolumn{3}{c}{$\gamma=1$}
&\multicolumn{3}{c}{$\gamma=2$}\\
\cmidrule(lr){2-4}\cmidrule(lr){5-7}\cmidrule(lr){8-10}
Design
& $\tilde Z_n^c$ & $Z_n^{\mathrm{det}}$
& $\lambda_{\max}^{\mathrm{cal}}$
& $\tilde Z_n^c$ & $Z_n^{\mathrm{det}}$
& $\lambda_{\max}^{\mathrm{cal}}$
& $\tilde Z_n^c$ & $Z_n^{\mathrm{det}}$
& $\lambda_{\max}^{\mathrm{cal}}$\\
\midrule
Gaussian & 0.051 & 0.052 & 0.046 & 0.045 & 0.048 & 0.047 & 0.050 & 0.052 & 0.050 \\
$t_{10}$ & 0.053 & 0.056 & 0.051 & 0.051 & 0.052 & 0.048 & 0.049 & 0.049 & 0.050 \\
$t_8$ & 0.053 & 0.055 & 0.049 & 0.049 & 0.051 & 0.046 & 0.046 & 0.048 & 0.046 \\
$t_5$ & 0.048 & 0.052 & 0.050 & 0.047 & 0.055 & 0.050 & 0.047 & 0.051 & 0.049 \\
$t_3^{\star}$ & 0.039 & 0.074 & 0.070 & 0.038 & 0.079 & 0.070 & 0.036 & 0.083 & 0.077 \\
Skewed $\chi^2_4$ & 0.049 & 0.053 & 0.047 & 0.051 & 0.055 & 0.051 & 0.047 & 0.052 & 0.048 \\
\bottomrule
\end{tabular*}
\par
\vspace{0.05cm}
\begin{minipage}{\textwidth}
\footnotesize\emph{Notes:}
Rejection frequencies at nominal level $0.05$, based on $10{,}000$
evaluation replications with $n=200$.
$\tilde Z_n^c$ is the demeaned, studentized, diagonal-free corrected
statistic; $Z_n^{\mathrm{det}}$ replaces the data-driven correction by its
deterministic product-null counterpart; and
$\lambda_{\max}^{\mathrm{cal}}$ denotes the largest-eigenvalue test using
an independently simulated Gaussian critical value.  The Student designs use
$X_{it}=\sqrt{(\nu-2)/\nu}T_{it}$ with
$T_{it}{\sim}\ {\mathrm{i.i.d.}}\ t_\nu$.  The skewed chi-square design uses $X_{it}=(Y_{it}-4)/\sqrt{8}$ with
$Y_{it}{\sim}\ {\mathrm{i.i.d.}}\ \chi_4^2$. ${}^{\star}$ The standardized $t_3$ design has an infinite fourth moment and lies
outside Assumption~\ref{ass:moments}; it is included only as a
finite-sample stress test.
\end{minipage}
\end{table}

\subsection{Studentization, demeaning, and exact centering}

Theorem~\ref{thm:martingale-clt} predicts the successive equivalences
\[
\tilde Z_n-Z_n=o_p(1),
\qquad
\tilde Z_n^c-\tilde Z_n=o_p(1).
\]
Table~\ref{tab:adm-implementation-main} examines both steps on common
samples under Gaussian, standardized $t_5$, and skewed chi-square
marginals.  Both mean absolute errors roughly halve between $n=p=100$ and
$400$ in every design; the slower studentization convergence under $t_5$
is consistent with the heavier tails allowed by the $4+\eta$ theorem,
while the chi-square design shows that the same equivalences extend to
asymmetric marginals.

Across the nine designs and dimensions, the three correctly centered
rejection frequencies range from $0.044$ to $0.057$. The naive demeaned version
behaves very differently: its rejection probability ranges from $0.115$
to $0.127$. When $p=n$, retaining the divisor $n^2$ after demeaning
creates the leading location shift $n/\{2(n-1)\}$, approximately one half
over this range, explaining the persistent overrejection.

\begin{table}[t]
\centering
\caption{Implementation equivalence under product-coordinate nulls}
\label{tab:adm-implementation-main}
\footnotesize
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}lccccccc}
\toprule
&& \multicolumn{2}{c}{Mean absolute difference}
& \multicolumn{4}{c}{Rejection frequency}\\
\cmidrule(lr){3-4}\cmidrule(lr){5-8}
Design & $n=p$
& $\mathbb{E}|\tilde Z_n-Z_n|$
& $\mathbb{E}|\tilde Z_n^c-\tilde Z_n|$
& $Z_n$ & $\tilde Z_n$ & $\tilde Z_n^c$ & Naive\\
\midrule
Gaussian & 100 & 0.157 & 0.113 & 0.051 & 0.049 & 0.051 & 0.126 \\
Gaussian & 200 & 0.111 & 0.080 & 0.048 & 0.048 & 0.045 & 0.123 \\
Gaussian & 400 & 0.080 & 0.056 & 0.050 & 0.049 & 0.050 & 0.121 \\
$t_5$ & 100 & 0.271 & 0.109 & 0.050 & 0.044 & 0.046 & 0.115 \\
$t_5$ & 200 & 0.199 & 0.078 & 0.051 & 0.047 & 0.047 & 0.124 \\
$t_5$ & 400 & 0.146 & 0.055 & 0.054 & 0.051 & 0.051 & 0.122 \\
Skewed $\chi_4^2$ & 100 & 0.240 & 0.121 & 0.057 & 0.047 & 0.048 & 0.120 \\
Skewed $\chi_4^2$ & 200 & 0.173 & 0.083 & 0.054 & 0.051 & 0.051 & 0.127 \\
Skewed $\chi_4^2$ & 400 & 0.126 & 0.057 & 0.049 & 0.046 & 0.047 & 0.121 \\
\bottomrule
\end{tabular*}
\par\smallskip
\begin{minipage}{\textwidth}
\footnotesize\emph{Notes:} Results are based on $10{,}000$ common-sample
replications per row at nominal level $0.05$. The first two numerical
columns report sample mean absolute differences; the remaining columns
report rejection frequencies. The skewed chi-square design uses
$X_{it}=(Y_{it}-4)/\sqrt{8}$, where
$Y_{it}{\sim}\ {\mathrm{i.i.d.}}\ \chi_4^2$. ``Naive'' denotes the
demeaned statistic retaining the divisor $n^2$ instead of the exact
$n(n-1)$ divisor.
\end{minipage}
\end{table}





Under the exact GOE mixture the same pattern holds.  At $\omega=2$, the
rejection probabilities of $(Z_n,\tilde Z_n,\tilde Z_n^c)$ are
$(0.508,0.485,0.487)$ at $n=p=100$ and
$(0.564,0.558,0.557)$ at $n=p=200$: the feasible implementation
preserves the local power, with gaps shrinking in $n$.


\subsection{The local-power formula}

For each replication we draw a new GOE matrix and simulate from the exact
quadratic alternative (\ref{eq:alternative}).  The top row of
Figure~\ref{fig:adm-sim} compares rejection frequencies with
$\beta^*(\omega,\gamma)=1-\Phi(z_{0.95}-\omega^2\gamma/2)$.
At $\gamma=1$, the finite-sample curves approach the envelope as $n=p$
increases.  At $\omega=2$, power rises from $0.487$ to $0.557$ and $0.594$
as the common dimension increases from $100$ to $200$ and $400$, compared
with the limit $0.639$.  At $\omega=2.5$, the corresponding values are
$0.740$, $0.842$, and $0.891$, compared with $0.931$.

At $n=200$, the ordering across $\gamma=0.5,1,2$ agrees with the theory
throughout the grid; at $\omega=1$, for example, the simulated rejection
frequencies are $0.080$, $0.121$, and $0.243$, against theoretical values
$0.082$, $0.126$, and $0.260$.  This comparison holds $\omega$ fixed; at a
fixed limiting Frobenius radius $\delta=\omega\gamma$ the shift is
$\delta^2/(2\gamma)$, so the comparative statics reverse.

\begin{figure}[t]
\centering
\includegraphics[width=\textwidth]{figures/adm_local_power_validation.pdf}\\[2pt]
\includegraphics[width=\textwidth]{figures/adm_signal_geometry.pdf}
\caption{Top row: finite-sample validation of the GOE local-power formula;
the left panel shows convergence at $\gamma=1$, the right panel varies the
aspect ratio at $n=200$; curves are the Gaussian-shift power functions,
markers are simulated rejection frequencies.  Bottom row: power under
dense and spiked alternatives at $n=p=200$, with the largest-eigenvalue
test calibrated under the Gaussian null; the left panel uses dense
GOE alternatives, the right panel rank-one spiked alternatives, with the
vertical line at $\vartheta=\sqrt\gamma=1$ marking the BBP threshold.}
\label{fig:adm-sim}
\end{figure}

\subsection{Dense and spiked alternatives}

The bottom row of Figure~\ref{fig:adm-sim} compares the corrected Frobenius
test with a Gaussian-calibrated largest-eigenvalue test at $n=p=200$.  Under the GOE
alternatives, Frobenius power is $0.557$ at $\omega=2$ against $0.158$
for the spectral test; at $\omega=3$ the values are $0.976$ and $0.359$.

Under a shrinking equicorrelation spike, the ordering reverses above the
BBP sample-eigenvalue separation threshold
$\vartheta=\sqrt\gamma=1$ when $\gamma=1$.  At $\vartheta=0.5$, Frobenius and
spectral powers are $0.065$ and $0.055$; at $\vartheta=1.5$ they are $0.304$
and $0.675$; at $\vartheta=2$ they are $0.606$ and $0.967$; and at
$\vartheta=2.5$ they are $0.858$ and $0.998$.

Dense-spectrum, path-robustness, and common-scale stress tests are
reported in Supplement Section~S.9.

\section{Discussion}
\label{sec:conclusion}

Along the quadratic precision path, the upper-tail test of the
corrected Frobenius statistic is asymptotically equivalent to the
Neyman--Pearson test against the GOE mixture.  It therefore attains
the limiting mixture-power envelope at every fixed strength, a full
local power curve rather than a detection rate.

The fully feasible correlation statistic has the same null limit and,
by contiguity, the same GOE local power as the known-scale statistic,
under $4+\eta$ moments for studentization and temporally i.i.d.\
sampling for demeaning; its null variance formula uses coordinate
independence.

Rigidity is the converse of the envelope: a same-level test that is
asymptotically nowhere worse than the benchmark uniformly over the
typical dense bulk merges with the Frobenius test under the Gaussian
null, and its power difference from the Frobenius test vanishes along
every sequence contiguous to that null, so gains are confined to
noncontiguous or prior-negligible signals.  The projection lemma of Supplement
Section~S.7, by which the closest positive-semidefinite rank-$k$
Frobenius approximation of a symmetric discrepancy retains its $k$
largest positive eigencomponents, suggests fitting and deflating one
factor at a time as a residual diagnostic; its calibration is open.

\begin{acks}[Acknowledgments]
We thank Michael Wolf for detailed comments, Martin Wagner, and
participants at the 6th Vienna Workshop on High-Dimensional Time
Series in Macroeconomics and Finance (May 2024) and at the W\&R Talk
of the Quantitative Economics Division at the University of Klagenfurt
(June 2025) for helpful comments and discussions.
\end{acks}

\begin{funding}
Werner Ploberger acknowledges financial support from the Weidenbaum
Center on the Economy, Government, and Public Policy at Washington
University in St.\ Louis.  Chen Tong acknowledges financial support
from the National Natural Science
Foundation of China (72301227) and the Fujian Provincial Natural Science
Foundation of China (2025J08008).
\end{funding}



\clearpage
\setcounter{table}{0}
\begin{center}
{\large\bfseries Supplementary Material}\\[6pt]
{\bfseries Supplement to ``Local Optimality and Rigidity of Frobenius Tests for
Dense High-Dimensional Covariance Alternatives''}
\end{center}
\medskip
\noindent This supplementary material contains the auxiliary calculations, the
finite-sample admissibility, power-enhancement, and conditional-flatness
results, the trace-matched weak-factor decay analysis, additional
simulations, and complete proofs for the main text.  Sections are
numbered S.1, S.2, \ldots\ and appendices A--G; plain-numbered references
(Theorem~1, Assumption~2, equation~(16), Section~3) refer to the main
text above.
\medskip

\setcounter{section}{0}
\setcounter{equation}{0}
\setcounter{theorem}{0}
\numberwithin{theorem}{section}
\numberwithin{proposition}{section}
\numberwithin{lemma}{section}
\numberwithin{corollary}{section}
\numberwithin{definition}{section}
\numberwithin{assumption}{section}
\numberwithin{remark}{section}

\numberwithin{equation}{section}

\section{GOE geometry and trace calculations}

Let $\mathbb S^p$ be the real symmetric $p\times p$ matrices with the
Frobenius inner product, and let $d_p=p(p+1)/2$.  A GOE matrix $H$ has
independent upper-triangular entries with $H_{ij}\sim N(0,1)$ for $i<j$ and
$H_{ii}\sim N(0,2)$.

\begin{proposition}[Exact GOE polar decomposition]
\label{prop:goe-polar}
Let $R_p^{\mathrm{GOE}}=\|H\|_F$ and $U_p=H/\|H\|_F$, with an arbitrary
definition on the null event $\{H=0\}$.  Then $U_p$ is uniform on the
Frobenius unit sphere of $\mathbb S^p$, $U_p$ and
$R_p^{\mathrm{GOE}}$ are independent, and
\[
\frac{(R_p^{\mathrm{GOE}})^2}{2}\sim\chi^2_{d_p}.
\]
In particular, $R_p^{\mathrm{GOE}}/p=1+O_p(p^{-1})$.
\end{proposition}

\begin{proposition}[GOE spectrum]
\label{prop:goe-spectrum}
As $p\to\infty$,
\[
p^{-1/2}\|H\|_{\mathrm{op}}\to_p2,
\]
and the empirical distribution of the eigenvalues of $p^{-1/2}H$ converges
weakly, in probability, to the semicircle law with density
$(2\pi)^{-1}\sqrt{4-x^2}$ on $[-2,2]$.
\end{proposition}

Both assertions are classical.  In the normalization of
\citet{AndersonGuionnetZeitouni:2010}, a Wigner matrix has off-diagonal
entries of variance $1/p$ and diagonal entries with all moments finite,
which is exactly $p^{-1/2}H$ here; the diagonal variance $2$ affects
neither limit.  Wigner's theorem is their Theorem~2.1.1, and
$\lambda_{\max}(p^{-1/2}H)\to_p2$ is their Theorem~2.1.22, whose
moment-growth condition is satisfied by Gaussian entries; since $-H$ is
again a GOE matrix, $\lambda_{\min}(p^{-1/2}H)\to_p-2$ as well, which
gives the operator-norm statement.  Almost-sure versions are in
\citet[Chapters~2 and~5]{BaiSilverstein:2010}; only convergence in
probability is used below.

\begin{lemma}[Trace orders]
\label{lem:traces}
For a $p\times p$ GOE matrix,
\begin{align*}
p^{-1/2}\operatorname{tr}(H)&\sim N(0,2),
&p^{-2}\operatorname{tr}(H^2)&\to_p1,\\
p^{-3/2}\operatorname{tr}(H^3)&\Rightarrow N(0,24),
&p^{-3}\operatorname{tr}(H^4)&\to_p2,
\end{align*}
and $\operatorname{tr}(H^5)=o_p(p^{7/2})$.  More generally,
$p^{-(k+1)}\operatorname{tr}(H^{2k})\to_p C_k$, where
$C_k=(k+1)^{-1}\binom{2k}{k}$.
\end{lemma}

\section{Auxiliary likelihood results}

Let $\ell_n(H)$ be the exact conditional likelihood ratio under the
quadratic path in equation~(1) of the main paper, and let
$\psi_n(H)$ be the factorized approximation leading to
(13).

\begin{lemma}[Likelihood comparison]
\label{lem:approximation}
If $\lambda_1(H),\ldots,\lambda_p(H)$ are the eigenvalues of $H$ and
$\phi(u)=u+u^2/2-\log(1+u+u^2)$, then
\[
f_n(H):=\log\frac{\psi_n(H)}{\ell_n(H)}
=\frac n2\sum_{j=1}^p
\phi\left(\frac{\omega\lambda_j(H)}n\right)
+\frac{\omega^4\gamma_n^3}{4}.
\]
The ratio is independent of the data.  On the likelihood bulk
$\mathcal K_n^{\mathrm{LR}}$ defined in Lemma~\ref{lem:psi-tail}, $f_n$ is
uniformly bounded, and $f_n(H)\to0$ in $Q_n$ probability.
\end{lemma}

\begin{lemma}[Null centering]
\label{lem:centering}
Under the Gaussian null and under each fixed contiguous GOE mixture,
\[
S_n=
\operatorname{tr}\{(\hat R_n-I_p)^2\}
-\frac1{n^2}\sum_{t=1}^n
\left[\left(\sum_iX_{it}^2\right)^2+p-2\sum_iX_{it}^2\right]
+o_p(1)
=4U_n+o_p(1),
\]
where $S_n$ and $U_n$ are defined in (16) and
(5) of the main paper.
\end{lemma}

\subsection{Whitening isomorphism and channel orthogonality}
\label{app:whitening}

The next two proofs support the known-$\Sigma_0$ reduction of Section~2.3 of the
main paper.  Throughout, $Y_t=\Sigma_0^{-1/2}X_t$ are the whitened observations,
$h_{\Sigma_0}(x,y)=(x^\prime\Sigma_0^{-1}y)^2-x^\prime\Sigma_0^{-1}x-y^\prime\Sigma_0^{-1}y+p$, and
$U_n^{\Sigma_0}=(2n^2)^{-1}\sum_{s<t}h_{\Sigma_0}(X_s,X_t)$.

\begin{proof}[Proof of Lemma~1]
(i) $\operatorname{cov}(Y_t)=\Sigma_0^{-1/2}\Sigma\Sigma_0^{-1/2}$ is $I_p$
under the null and $(I_p+A+A^2)^{-1}$ under the conditional alternative, with
$A=\omega H/n$; Gaussianity and independence across $t$ are preserved by the
fixed linear map.  (ii) The Jacobian $|\det\Sigma_0^{-1/2}|^{-n}$ is common to the
numerator and denominator of the likelihood ratio of two laws of the same sample
and cancels; integrating over $H$ preserves the equality.  (iii) is the
coordinatewise identity~(5): at the whitened sample the off-diagonal part of a
pair $(s,t)$ equals $(Y_s'Y_t)^2-\sum_iY_{is}^2Y_{it}^2$ and the diagonal part
equals $\sum_iY_{is}^2Y_{it}^2-\|Y_s\|^2-\|Y_t\|^2+p$, so the coincident products
$\sum_iY_{is}^2Y_{it}^2$ cancel and the pair contributes
\[
(Y_s'Y_t)^2-\|Y_s\|^2-\|Y_t\|^2+p=h_{\Sigma_0}(X_s,X_t),
\]
using $Y_s'Y_t=X_s^\prime\Sigma_0^{-1}X_t$ and $\|Y_t\|^2=X_t^\prime\Sigma_0^{-1}X_t$.
Summing over pairs and dividing by $2n^2$ gives~(iii).
\end{proof}

\begin{proof}[Proof of Lemma~2 (trace--Frobenius orthogonality)]
Let $g(x)=x^\prime\Sigma_0^{-1}x-p$, such that
$W_n^{\Sigma_0}=n^{-1}\sum_{r=1}^ng(X_r)$, and recall that the kernel is degenerate
under the null: $\mathbb{E}[h_{\Sigma_0}(x,X)]=0$ for every fixed $x$ when
$X\sim N_p(0,\Sigma_0)$, because
$\mathbb{E}(x^\prime\Sigma_0^{-1}X)^2=x^\prime\Sigma_0^{-1}x$ and
$\mathbb{E}X^\prime\Sigma_0^{-1}X=p$.  For each pair $s<t$ and index $r$: if
$r\notin\{s,t\}$ the three vectors are independent and the summand has mean
zero; if $r\in\{s,t\}$, say $r=s$,
\[
\mathbb{E}[h_{\Sigma_0}(X_s,X_t)g(X_s)]
=\mathbb{E}\big[g(X_s)\mathbb{E}\{h_{\Sigma_0}(X_s,X_t)|X_s\}\big]=0 .
\]
Since $\mathbb{E}h_{\Sigma_0}=0$, every covariance term vanishes, so
$\operatorname{cov}(U_n^{\Sigma_0},W_n^{\Sigma_0})=0$.
\end{proof}

\section{Fixed-radius equivalence: proof of
Proposition~1}
\label{app:fixed-radius}

Write $P_{n,S,u}$ for the $n$-sample Gaussian law with precision
$B_{S,u}=I_p+uS/n+u^2S^2/n^2$, such that
$G_n^{\mathrm{sph}}(\omega)=\mathbb{E}_S[P_{n,S,\omega}]$ with
$S\sim Q_n^{\mathrm{sph}}$ and, by the polar decomposition
(Proposition~\ref{prop:goe-polar}) with $\rho=\|H\|_F/p$, $S=pH/\|H\|_F$,
$G_n(\omega)=\mathbb{E}_{(\rho,S)}[P_{n,S,\omega\rho}]$: the likelihood
(11) depends on $(M,\omega)$ only through $A=\omega M/n$, whence
$\ell_n(H;\omega)=\ell_n(S;\omega\rho)$ exactly.

\emph{Step 1: radial moments.}  $\rho^2=2V/p^2$ with $V\sim\chi^2_{d_p}$,
$d_p=p(p+1)/2$, so $\mathbb{E}\rho^2=1+p^{-1}$ and
$\operatorname{var}(\rho^2)=8d_p/p^4=4(p+1)/p^3\le8p^{-2}$; hence
$\mathbb{E}|\rho^2-1|\le\{\operatorname{var}(\rho^2)\}^{1/2}+|\mathbb{E}\rho^2-1|
\le4p^{-1}$ and, since
$\rho>0$, $\mathbb{E}|\rho-1|\le\mathbb{E}|\rho^2-1|\le4p^{-1}$.

\emph{Step 2: Fisher information along the path.}  For fixed $S$ on the
sphere, the $n$-sample Fisher information of $u\mapsto P_{n,S,u}$ is
$I_{n,S}(u)=\tfrac n2\operatorname{tr}\{(B_{S,u}^{-1}B_{S,u}^\prime)^2\}$ with
$B_{S,u}^\prime=S/n+2uS^2/n^2$.  Since $B_{S,u}$ and $B_{S,u}^\prime$ are polynomials in
$S$, they commute and the trace is a sum over eigenvalues.  Every eigenvalue
of $B_{S,u}$ has the form $1+t+t^2\ge3/4$, so
$\|B_{S,u}^{-1}\|_{\mathrm{op}}\le4/3$; and
$\|B_{S,u}^\prime\|_F^2\le2\operatorname{tr}(S^2)/n^2+8u^2\operatorname{tr}(S^4)/n^4
\le2\gamma_n^2+8u^2\gamma_n^4$, using $\operatorname{tr}(S^4)\le\{\operatorname{tr}(S^2)\}^2=p^4$.
Hence, uniformly over the sphere and for every $u\ge0$,
\[
\sqrt{I_{n,S}(u)}\le C\sqrt n\big(\gamma_n+u\gamma_n^2\big),
\]
with an absolute constant $C$.

\emph{Step 3: path total-variation bound.}  The Gaussian densities $p_t$
of $P_{n,S,t}$ have analytic, uniformly positive-definite precision
matrices, so $t\mapsto p_t(x)$ is absolutely continuous with integrable
derivative on compact intervals, and Cauchy--Schwarz gives
$\int|\partial_tp_t|=\int|\partial_t\log p_t|p_t\le\{I_{n,S}(t)\}^{1/2}$.
Hence, for $0\le u\le v$,
\begin{align*}
\|P_{n,S,u}-P_{n,S,v}\|_{\mathrm{TV}}
&=\frac12\int|p_u-p_v|
\le\frac12\int_u^v\{I_{n,S}(t)\}^{1/2}dt\\
&\le C\sqrt n\Big\{\gamma_n(v-u)+\gamma_n^2\frac{v^2-u^2}{2}\Big\}.
\end{align*}

\emph{Step 4: coupling.}  By joint convexity of total variation and the
coupling in the opening display, with $\{u,v\}=\{\omega,\omega\rho\}$ and no
truncation of $\rho$,
\[
\big\|G_n^{\mathrm{sph}}(\omega)-G_n(\omega)\big\|_{\mathrm{TV}}
\le\mathbb{E}_{(\rho,S)}
\big\|P_{n,S,\omega}-P_{n,S,\omega\rho}\big\|_{\mathrm{TV}}
\le C_\Omega\sqrt n\gamma_n
\mathbb{E}\big[|\rho-1|+\gamma_n|\rho^2-1|\big],
\]
uniformly in $\omega\le\Omega$.  Step 1 bounds the expectation by
$4(1+\gamma_n)p^{-1}$, and $\sqrt n\gamma_n/p=n^{-1/2}=O(p^{-1/2})$
exactly, proving the proposition.

\emph{Proof of Corollary~1.}  Both mixture laws are
dominated by $P_n$ (each conditional likelihood is strictly positive by
$1+x+x^2>0$), so
$\mathbb{E}_{P_n}|g_n^{\mathrm{sph}}(\omega)-g_n(\omega)|
=2\|G_n^{\mathrm{sph}}(\omega)-G_n(\omega)\|_{\mathrm{TV}}\rightarrow0$,
whence $g_n^{\mathrm{sph}}(\omega)-g_n(\omega)\rightarrow_{P_n}0$.  By
Theorem~1(i), $g_n(\omega)$ converges to a strictly
positive lognormal limit, so it is bounded away from zero in
$P_n$-probability; dividing gives
$g_n^{\mathrm{sph}}(\omega)/g_n(\omega)\rightarrow_{P_n}1$ and therefore
$\log g_n^{\mathrm{sph}}(\omega)=\log g_n(\omega)+o_{P_n}(1)$.  For the
optimality transfers, note that the $L^1(P_n)$ closeness holds jointly for
every fixed finite collection of strengths $\omega_1,\ldots,\omega_m$, so
the joint likelihood-ratio limits of Theorem~1(iii) are unchanged, and
continuity of the limiting lognormal distribution at its quantiles carries
the Neyman--Pearson critical values over to the sphere mixture; every
limiting log likelihood ratio remains the same increasing affine function of
$U_n$.\qed

\subsection{Equivalence to additive covariance alternatives}
\label{supp:additive}

The quadratic precision path is the globally positive representative of
the additive covariance model $\Sigma=I_p-A$; the two mixtures are
asymptotically indistinguishable, and so are all modifications of the
precision path at cubic or higher order.

\begin{lemma}
\label{lem:additive-covariance}
Suppose $p/n\to\gamma\in(0,\infty)$ and fix $0<\Omega<\infty$.  For
$0\le\omega\le\Omega$ put $A=\omega H/n$ and let
$E_n=\{\|H\|_{\mathrm{op}}\le n/(2\Omega)\}$, so that
$\|A\|_{\mathrm{op}}\le1/2$ on $E_n$.  Let $G_n^{\mathrm{lin}}(\omega)$
be the mixture of $N_p(0,I_p-A)^{\otimes n}$ over the GOE law
conditioned on $E_n$.  Then
\[
\sup_{0\le\omega\le\Omega}
\|G_n^{\mathrm{lin}}(\omega)-G_n(\omega)\|_{\mathrm{TV}}=O(n^{-1/2}).
\]
The same bound holds when $I_p-A$ is replaced by
$(I_p+A+A^2+\sum_{k\ge3}a_kA^k)^{-1}$ for any fixed real coefficients
with $\sum_{k\ge3}|a_k|2^{-k}<3/4$, in particular for any single cubic
coefficient $|a_3|<6$.
\end{lemma}

\begin{proof}
Write $\Sigma_q=(I_p+A+A^2)^{-1}$ and $\Sigma_l=I_p-A$.  On $E_n$ the
eigenvalues of both matrices lie in $[1/2,3/2]$, uniformly in
$\omega\le\Omega$.  By the exact identity
$(1+u+u^2)^{-1}-(1-u)=u^3(1+u+u^2)^{-1}$,
\[
\Sigma_q-\Sigma_l=A^3\Sigma_q,
\qquad
\|\Sigma_q-\Sigma_l\|_F^2\le\|\Sigma_q\|_{\mathrm{op}}^2\operatorname{tr}(A^6)
\le2\operatorname{tr}(A^6).
\]
For Gaussian laws with eigenvalues in a fixed compact subset of
$(0,\infty)$, the Kullback--Leibler divergence of $n$ independent
observations satisfies
$D_{\mathrm{KL}}(N_p(0,\Sigma_q)^{\otimes n}\,\|\,N_p(0,\Sigma_l)^{\otimes n})
\le Cn\|\Sigma_q-\Sigma_l\|_F^2\le Cn\operatorname{tr}(A^6)$.  The Wick expansion of
GOE moments gives $\mathbb{E}\operatorname{tr}(H^6)\le Cp^4$, so
\[
\mathbb{E}\operatorname{tr}(A^6)\le C\Omega^6p^4/n^6=O(n^{-2}),
\qquad
Q_n(E_n^c)\le(2\Omega/n)^6\,\mathbb{E}\operatorname{tr}(H^6)=O(n^{-2}),
\]
the latter by Markov's inequality.  Let
$G_n^{q,E}(\omega)$ be the quadratic-path mixture over the prior
conditioned on $E_n$.  Convexity of total variation in the mixing law,
Pinsker's inequality, and Jensen's inequality give
\[
\|G_n^{q,E}(\omega)-G_n^{\mathrm{lin}}(\omega)\|_{\mathrm{TV}}
\le\mathbb{E}\big[\{D_{\mathrm{KL}}/2\}^{1/2}\,\big|\,E_n\big]
\le\Big\{\frac{Cn\,\mathbb{E}\operatorname{tr}(A^6)}{2Q_n(E_n)}\Big\}^{1/2}
=O(n^{-1/2}),
\]
while $\|G_n(\omega)-G_n^{q,E}(\omega)\|_{\mathrm{TV}}\le Q_n(E_n^c)
=O(n^{-2})$.  The triangle inequality proves the first claim.  For the
second, the modified precision matrix has eigenvalues
$1+u+u^2+\sum_{k\ge3}a_ku^k\ge3/4-\sum_{k\ge3}|a_k|2^{-k}>0$ on $E_n$,
and its inverse differs from $\Sigma_q$ by $O(\|A\|_{\mathrm{op}}^3)$ in
operator norm with Frobenius norm squared bounded by $C\operatorname{tr}(A^6)$, so the
same argument applies.
\end{proof}

The lemma is a statement about the sample laws conditional on $H$,
mixed afterwards; it does not require the likelihood expansion, and it
shows that, under the positivity, localization, and remainder
conditions stated in the lemma, the limit experiment of Theorem~1 is
that of the additive covariance model and of every path agreeing with
it through second order.  Without localization a modified precision
need not be positive definite on the full GOE support: $1+u+u^2+u^3$
is negative for $u<-1$, a region of positive prior probability at
every $n$.

\section{Scaling and studentization identities}

For the corrected criterion evaluated after diagonal scaling, write
$Q(V)=p+V'A_nV-2b_n'V$, where
\[
[A_n]_{ij}=\hat r_{ij}^{2}
-\frac1{n^2}\sum_tX_{it}^2X_{jt}^2,
\qquad
[b_n]_i=(1-n^{-1})\hat r_{ii}.
\]
Then
\[
A_n=\frac1{n^2}\sum_{s\ne t}(X_s\circ X_t)(X_s\circ X_t)^\prime
\]
is positive semidefinite.  If $A_n$ is nonsingular and $A_n^{-1}b_n$ is
componentwise positive, it is the minimizer over nonnegative diagonal
scalings.  Under uniformly bounded eighth moments, for each fixed
coordinate the first-order residual at $v_i^0=1/\hat r_{ii}$ is
$O_p(\sqrt p/n+1/n)$ (Appendix~D).  Under the additional conditioning
and uniform-row assumptions stated there, the coordinatewise difference
between the optimizer and ordinary studentization is $O_p(n^{-1/2})$,
which is agreement at order one.  None of the formal testing results
relies on this variational condition.

\begin{lemma}[Exact scaling and orthogonality]
\label{lem:scaling-orthogonality}
Let
\[
D_{ij}=\hat r_{ij}^{2}
-\frac1{n^2}\sum_tX_{it}^2X_{jt}^2,
\qquad i<j,
\]
and let $\tilde D_{ij}$ be the same corrected quantity after
studentization.  Then (i)
\[
\tilde D_{ij}=\frac{D_{ij}}{\hat r_{ii}\hat r_{jj}};
\]
(ii) $D_{ij}$ is conditionally mean zero given either row; and (iii)
\[
\mathbb{E}\{D_{ij}(\hat r_{jj}-1)|X_{i1},\ldots,X_{in}\}=0,
\]
with the symmetric identity after interchanging $i$ and $j$.
\end{lemma}

\subsection{Computational identities}
\label{supp:computation}

Let $Y=(\tilde X_{it}^c)_{i,t}$ and $\tilde R_n^c=n^{-1}YY^\prime$.
With $q_i^c=n^{-1}\sum_t(X_{it}-\bar X_i)^2$ the centered sample
variance, the diagonal element $\tilde r_{ii}^c$ equals $1\{q_i^c>0\}$
under the zero-row convention, so
\[
\sum_{i<j}(\tilde r_{ij}^c)^2
=\frac12\Big\{\|\tilde R_n^c\|_F^2-\sum\nolimits_i1\{q_i^c>0\}\Big\},
\]
which reduces to $\frac12\{\|\tilde R_n^c\|_F^2-p\}$ when every
centered sample variance is positive, and
the elementary identity
$\sum_{i<j}a_ia_j=\{(\sum_i a_i)^2-\sum_i a_i^2\}/2$ gives
\[
\sum_{t=1}^n\sum_{i<j}
(\tilde X_{it}^c)^2(\tilde X_{jt}^c)^2
=\frac12\sum_{t=1}^n
\left\{\Big(\sum_iY_{it}^2\Big)^2-\sum_iY_{it}^4\right\}.
\]
Thus equation~(27) of the main paper can be evaluated from one sample
correlation matrix and two columnwise sums; the dominant cost is forming
$YY^\prime$, namely $O(np^2)$ operations, and blocked matrix
multiplication avoids storing additional pairwise arrays.  The matrix
and pairwise implementations agree to machine precision in the supplied
code.

\section{Exact centering after demeaning}

\begin{lemma}[Exact degrees-of-freedom centering]
\label{lem:demeaning-centering}
Let $q_j^c=n^{-1}\sum_t(X_{jt}-\bar X_j)^2$ be the centered sample
variance of row $j$ ($\hat r_{jj}^c$ in the main paper), set the
standardized centered row
$\tilde X_j^c=(\tilde X_{j1}^c,\ldots,\tilde X_{jn}^c)$ to zero on the
event $\{q_j^c=0\}$, and let $\pi_{j,n}=\mathbb{P}(q_j^c>0)$.  Under
Assumptions~2 and 3 of the main paper, for every $i\ne j$,
\[
\mathbb{E}\{(\tilde r_{ij}^c)^2\mid\tilde X_i^c\}
=\mathbb{E}\left\{\left.
\frac1{n(n-1)}\sum_t
(\tilde X_{it}^c)^2(\tilde X_{jt}^c)^2
\right|
\tilde X_i^c\right\}
=\frac{\pi_{j,n}}{n-1}\,1\{q_i^c>0\}.
\]
The corrected pairwise summand is therefore exactly conditionally
degenerate.  Each side differs from $(n-1)^{-1}1\{q_i^c>0\}$ by at most
$(1-\pi_{j,n})/(n-1)$ and from $1/(n-1)$ by at most
$\{1-\pi_{j,n}+1\{q_i^c=0\}\}/(n-1)$; both discrepancies are
exponentially negligible, uniformly in $i,j\le p$, by the bound
$\mathbb{P}(\min_{i\le p}q_i^c=0)\le p(1-c)^{n-1}$ of the main paper.
Both statements also hold after interchanging the two rows.
\end{lemma}

\section{Admissibility benchmarks and power enhancement}

Let
\[
\mathcal K_n=\{H:\|H\|_{\mathrm{op}}\le3\sqrt p\},
\qquad
Q_n^{\mathcal K}=Q_n(\cdot|\mathcal K_n),
\]
and write $G_n^{\mathcal K}(\omega)$ for the corresponding compact mixture.

\begin{proposition}[Finite-sample compact-mixture admissibility]
\label{prop:finite-sample}
Fix $n,p,\omega>0$ and $\alpha\in(0,1)$ with
$Q_n(\mathcal K_n)>0$.  The likelihood ratio of
$G_n^{\mathcal K}(\omega)$ relative to $P_n$ is atomless under $P_n$.
Hence its level-$\alpha$ Neyman--Pearson test is the $P_n$-almost-surely
unique most powerful test against the compact mixture and is an admissible
level-$\alpha$ test against
$\{\nu_n(H,\omega):H\in\mathcal K_n\}$.
\end{proposition}

\begin{lemma}[Attainment]
\label{lem:np-rigidity}
Fix $\omega>0$.  If $\mathbb{E}_{P_n}\varphi_n\to\alpha$ and
$\mathbb{E}_{G_n(\omega)}\varphi_n\to\beta^*(\omega)$, then
$\mathbb{E}_{P_n}|\varphi_n-\varphi_n^*|\to0$.
\end{lemma}

For the enhancement result, fix $\vartheta>\sqrt\gamma$ and let
$\nu_n^{\mathrm{sp}}$ be the Gaussian law with equicorrelation matrix
\[
R_p=(1-\rho_p)I_p+\rho_p\mathbf1\mathbf1^\prime,
\qquad \rho_p=\frac{\vartheta}{p-1}.
\]
Let $\varphi_n^\circ=1\{Z_n>z_{1-\alpha}\}$ be the known-scale Frobenius
test.

In our parameterization, $\vartheta=\sqrt\gamma$ is the
Baik--Ben Arous--P{\'e}ch{\'e} (BBP) sample-eigenvalue separation
threshold \citep{BaikBenArousPeche:2005}.  The null-edge limit used below
follows from \citet{Johnstone:2001}, while the super-critical outlier limit
for the real spiked covariance model follows from
\citet{BaikSilverstein:2006}.

\begin{lemma}[Frobenius statistic under a super-critical spike]
\label{lem:tightness}
Under $\nu_n^{\mathrm{sp}}$,
\[
Z_n\Rightarrow N\left(\frac{\vartheta^2}{2\gamma},1\right).
\]
Thus the limiting power of $\varphi_n^\circ$ is strictly below one.
\end{lemma}

\begin{proposition}[Power enhancement]
\label{prop:hodges}
Let
$m\in((1+\sqrt\gamma)^2,(1+\vartheta)(1+\gamma/\vartheta))$ and define
\[
\varphi_n=\max\{\varphi_n^\circ,
1\{\lambda_{\max}(\hat R_n)>m\}\}.
\]
Then $\varphi_n\ge\varphi_n^\circ$ pointwise,
$\mathbb{E}_{P_n}\varphi_n\to\alpha$, and
$\mathbb{E}_{\nu_n^{\mathrm{sp}}}\varphi_n\to1$, whereas the limiting power of
$\varphi_n^\circ$ is below one.  Thus the Gaussian known-scale Frobenius
procedure is not asymptotically admissible over unrestricted alternatives.
\end{proposition}

This proposition concerns the known-scale statistic.  Extending it to the
fully feasible statistic at the noncontiguous equicorrelation alternative
requires a separate studentization argument and is not claimed.

\subsection{Conditional flatness of the invariant benchmark}
\label{supp:flatness}

This subsection proves Lemma~3 of the main paper.  Throughout,
$\nu_n(H,\omega)$ is the Gaussian sample law with precision
$\Sigma^{-1}(H)=I_p+A+A^2$, $A=\omega H/n$,
$F_n=2U_n/\gamma_n$ with $U_n$ the invariant statistic
(equation~(5) of the main paper),
$\varphi_n^{U}=1\{F_n>z_{1-\alpha}\}$,
$\gamma_n=p/n\to\gamma\in(0,\infty)$, and
\[
\mathcal K_n^{\mathrm{rig}}
=\Big\{H:\ \|H\|_{\mathrm{op}}\le3\sqrt p,\
\big|p^{-2}\operatorname{tr}(H^2)-1\big|\le n^{-1/4}\Big\}.
\]
All matrices below are functions of $H$ and commute with one another,
and $\|\Sigma\|_{\mathrm{op}}\le2$ for $n$ large, uniformly on
$\mathcal K_n^{\mathrm{rig}}$.

\medskip
\noindent\emph{Step A: geometry of the bulk.}\quad
The exact identity
\begin{equation}
\Delta:=\Sigma-I_p=-A+A^3\Sigma,
\qquad\text{from }(1+u+u^2)^{-1}=1-u+\frac{u^3}{1+u+u^2},
\label{eq:supp-exact-D}
\end{equation}
gives, on $\mathcal K_n^{\mathrm{rig}}$: (i)
$\|\Delta\|_{\mathrm{op}}\le\|A\|_{\mathrm{op}}
+\|A\|_{\mathrm{op}}^3\|\Sigma\|_{\mathrm{op}}=O(n^{-1/2})$; (ii)
\begin{align*}
\operatorname{tr}(\Delta^2)
&=\operatorname{tr}(A^2)-2\operatorname{tr}(A^4\Sigma)+\operatorname{tr}(A^3\Sigma A^3\Sigma)\\
&=\operatorname{tr}(A^2)\{1+O(n^{-1})\}
=\omega^2\gamma_n^2\{1+O(n^{-1/4})\},
\end{align*}
using $|\operatorname{tr}(A^4\Sigma)|\le\|A\|_{\mathrm{op}}^2\|\Sigma\|_{\mathrm{op}}
\operatorname{tr}(A^2)$ and
$\operatorname{tr}(A^3\Sigma A^3\Sigma)\le\|A\|_{\mathrm{op}}^4
\|\Sigma\|_{\mathrm{op}}^2\operatorname{tr}(A^2)$.  In particular
$\|\Delta\|_F=O(1)$.  No condition on the diagonal of $H$ is used.

\medskip
\noindent\emph{Step B: exact decomposition under a null coupling.}\quad
Write $h(x,y)=(x^\prime y)^2-\|x\|^2-\|y\|^2+p
=\operatorname{tr}\{(xx^\prime-I_p)(yy^\prime-I_p)\}$, so that
$U_n=(2n^2)^{-1}\sum_{s<t}h(X_s,X_t)$ and
$F_n=(\gamma_nn^2)^{-1}\sum_{s<t}h(X_s,X_t)$.  Couple
$X_t=\Sigma^{1/2}Y_t$ with $Y_1,\ldots,Y_n$ independent $N_p(0,I_p)$,
and put $B_t=Y_tY_t^\prime-I_p$.  Then
$X_tX_t^\prime-I_p=\Sigma^{1/2}B_t\Sigma^{1/2}+\Delta$, and because
$\Sigma^{1/2}$ and $\Delta$ commute,
\[
h(X_s,X_t)
=h(Y_s,Y_t)+\operatorname{tr}\{B_s(\Sigma B_t\Sigma-B_t)\}
+\operatorname{tr}(B_s\Sigma\Delta)+\operatorname{tr}(B_t\Sigma\Delta)+\operatorname{tr}(\Delta^2).
\]
Summing over $s<t$,
\begin{equation}
F_n(X)=F_n(Y)+\frac{n-1}{2p}\operatorname{tr}(\Delta^2)+R_n^{d}+R_n^{l},
\label{eq:supp-flat-decomp}
\end{equation}
where
\[
R_n^{d}=\frac1{\gamma_nn^2}\sum_{s<t}\operatorname{tr}\{B_s\mathcal M(B_t)\},
\qquad
R_n^{l}=\frac{n-1}{\gamma_nn^2}\sum_{t}\operatorname{tr}(B_t\Sigma\Delta),
\]
and
where $\mathcal M(C)=\Sigma C\Sigma-C$ acts on the space of symmetric
matrices with the Frobenius inner product and is self-adjoint there.
The decomposition is exact for every $n$, $p$, and $H$.

\medskip
\noindent\emph{Step C: second moments of the remainders.}\quad
For symmetric deterministic $C$ and $D$ and $Y\sim N_p(0,I_p)$,
$\operatorname{tr}\{(YY^\prime-I_p)C\}=Y^\prime CY-\operatorname{tr} C$, so
\begin{equation}
\mathbb{E}\,\operatorname{tr}(B_tC)\operatorname{tr}(B_tD)=2\operatorname{tr}(CD).
\label{eq:supp-flat-isotropy}
\end{equation}
For $R_n^{l}$, the summands are independent across $t$, and
(\ref{eq:supp-flat-isotropy}) gives
\[
\mathbb{E}(R_n^{l})^2
=\frac{2(n-1)^2}{\gamma_n^2n^3}\operatorname{tr}\{(\Sigma\Delta)^2\}
\le\frac{2(n-1)^2\|\Sigma\|_{\mathrm{op}}^2}{\gamma_n^2n^3}\operatorname{tr}(\Delta^2)
=O(n^{-1}).
\]
For $R_n^{d}$, each summand $\xi_{st}=\operatorname{tr}\{B_s\mathcal M(B_t)\}$ has
conditional mean zero given $Y_s$ and given $Y_t$, so summands with
distinct unordered pairs are uncorrelated, and $\xi_{st}$ is
uncorrelated with every summand of $R_n^{l}$.  Conditioning on $Y_t$
and applying (\ref{eq:supp-flat-isotropy}) twice, once in $Y_s$ and
once in $Y_t$ over an orthonormal basis of symmetric matrices,
\[
\mathbb{E}\xi_{st}^2
=2\,\mathbb{E}\|\mathcal M(B_t)\|_F^2
=4\|\mathcal M\|_{\mathrm{HS}}^2 .
\]
Since $\mathcal M$ is the restriction to symmetric matrices of
$\Sigma\otimes\Sigma-I_p\otimes I_p
=\Delta\otimes I_p+I_p\otimes\Delta+\Delta\otimes\Delta$,
\[
\|\mathcal M\|_{\mathrm{HS}}
\le2\sqrt p\,\|\Delta\|_F+\|\Delta\|_F^2=O(\sqrt p),
\]
hence
\[
\mathbb{E}(R_n^{d})^2
=\frac{2(n-1)}{\gamma_n^2n^3}\|\mathcal M\|_{\mathrm{HS}}^2
=O\Big(\frac p{n^2}\Big)=O(n^{-1}).
\]
All bounds are uniform over $\mathcal K_n^{\mathrm{rig}}$, so
$\sup_{H\in\mathcal K_n^{\mathrm{rig}}}
\mathbb{E}(R_n^{d}+R_n^{l})^2=O(n^{-1})$.

\medskip
\noindent\emph{Step D: conclusion.}\quad
By Step A, $(n-1)\operatorname{tr}(\Delta^2)/(2p)
=\tfrac{n-1}{2n}\,\omega^2\gamma_n\{1+O(n^{-1/4})\}
\to\omega^2\gamma/2$ uniformly on $\mathcal K_n^{\mathrm{rig}}$.
The statistic $F_n(Y)$ in (\ref{eq:supp-flat-decomp}) is the null
statistic, whose law does not depend on $H$, and
$F_n(Y)\Rightarrow N(0,1)$ by Theorem~1.  Hence for every sequence
$H_n\in\mathcal K_n^{\mathrm{rig}}$,
$F_n(X)\Rightarrow N(\omega^2\gamma/2,\,1)$ under $\nu_n(H_n,\omega)$,
so $\mathbb{E}_{\nu_n(H_n,\omega)}\varphi_n^{U}\to\beta^*(\omega)$ by
continuity of $\Phi$ at $z_{1-\alpha}$; the subsequence criterion
turns this into the uniform statement of Lemma~3.
\qed

\medskip
\noindent\emph{Remark (the invariant kernel is essential).}\quad
The argument uses the full kernel $h$.  The off-diagonal known-scale
statistic $Z_n=a_n\sum\nolimits_{i<j}D_{ij}$ of the main paper is
not flat on $\mathcal K_n^{\mathrm{rig}}$; here
$a_n=n^2/\sqrt{p(p-1)n(n-1)}$.  Take
$H^{\mathrm d}=\sqrt p\,\operatorname{diag}(\varepsilon_1,\ldots,\varepsilon_p)$
with signs $\varepsilon_i=\pm1$; then
$\|H^{\mathrm d}\|_{\mathrm{op}}=\sqrt p$ and
$\operatorname{tr}\{(H^{\mathrm d})^2\}=p^2$ exactly, so
$H^{\mathrm d}\in\mathcal K_n^{\mathrm{rig}}$ for every $n$, while
$\Sigma(H^{\mathrm d})$ is diagonal.  Write $X_{it}=s_iY_{it}$ with
$s_i^2=\Sigma_{ii}(H^{\mathrm d})=1+O(n^{-1/2})$ uniformly in $i$ and
$Y_{it}$ i.i.d.\ $N(0,1)$.  The exact scaling identity
$D_{ij}(X)=s_i^2s_j^2D_{ij}(Y)$ gives
$Z_n(X)-Z_n(Y)=a_n\sum_{i<j}(s_i^2s_j^2-1)D_{ij}(Y)$, and, because
distinct pairs are uncorrelated,
\[
\mathbb{E}\{Z_n(X)-Z_n(Y)\}^2
\le\max_{i<j}|s_i^2s_j^2-1|^2\,a_n^2\sum_{i<j}\mathbb{E}D_{ij}(Y)^2
=O(n^{-1}).
\]
Hence $Z_n(X)=Z_n(Y)+o_p(1)\Rightarrow N(0,1)$ and the
directionwise power of $1\{Z_n>z_{1-\alpha}\}$ at $H^{\mathrm d}$
converges to $\alpha$ for every $\omega>0$.  Such directions are
$Q_n$-negligible, and they are not contiguous to the null even though
their GOE mixture is, so they do not affect any integrated statement;
they show that a correlation-based benchmark would require a
diagonal-mass restriction such as $\sum_iH_{ii}^2\le p^{3/2}$, which
the invariant benchmark avoids.

\section{Minimum-distance projections}
\label{app:projections}

Two deterministic facts underlie the scaling interpretation in Section~4.2
and the deflation heuristic in the Discussion.

\emph{Diagonal scaling.}  For a population covariance $R$ with
$r=\operatorname{diag}(R)$, and for $D=\operatorname{diag}(d_1,\ldots,d_p)$ with $v_i=d_i^2\ge0$,
\[
\|I_p-DRD\|_F^2
=p-2r'v+v^\prime(R\circ R)v,
\]
a convex quadratic in $v$ because the Hadamard square $R\circ R$ is
positive semidefinite.  The constrained minimizer over $v\ge0$ satisfies
the Karush--Kuhn--Tucker conditions
\[
v\ge0,\qquad (R\circ R)v-r\ge0,\qquad v_i\{(R\circ R)v-r\}_i=0
\quad\text{for every }i,
\]
and when it is interior and $R\circ R\succ0$ it is unique and equals
$v=(R\circ R)^{-1}r$.  At a diagonal $R$ with $R_{ii}>0$ this gives
$v_i=R_{ii}^{-1}$: ordinary studentization is the exact population
minimizer.

\begin{lemma}[Best positive-semidefinite rank-$k$ approximation]
\label{lem:psd-rank-k}
Let $1\le k\le p$.  Let $S$ be symmetric with eigenvalues
$\lambda_1\ge\cdots\ge\lambda_p$ and orthonormal eigenvectors
$u_1,\ldots,u_p$, write $\lambda^+=\max\{\lambda,0\}$, and set
$\lambda_{p+1}^+=0$.  Then
\[
\min_{F\in\mathbb R^{p\times k}}\|S-FF^\prime\|_F^2
=\operatorname{tr}(S^2)-\sum_{j=1}^k(\lambda_j^+)^2,
\]
and one minimizing matrix is $FF^\prime=\sum_{j=1}^k\lambda_j^+u_ju_j^\prime$.  The
minimizing matrix is unique unless $k<p$ and $\lambda_k=\lambda_{k+1}>0$,
in which case the minimizing eigenspace is not unique.
\end{lemma}

\begin{proof}
Write $S=S^+-S^-$ with $S^+=\sum_j\lambda_j^+u_ju_j^\prime$ and
$S^-=\sum_j(-\lambda_j)^+u_ju_j^\prime$, so $\langle S^+,S^-\rangle=0$.  For any
positive-semidefinite $M=FF^\prime$ of rank at most $k$,
\[
\|S-M\|_F^2
=\|S^+-M\|_F^2+2\langle M,S^-\rangle+\|S^-\|_F^2
\ge\|S^+-M\|_F^2+\|S^-\|_F^2,
\]
since $\langle M,S^-\rangle\ge0$ for positive-semidefinite $M$ and $S^-$.
By the Eckart--Young theorem applied to the positive-semidefinite matrix
$S^+$, $\|S^+-M\|_F^2\ge\sum_{j>k}(\lambda_j^+)^2$, with equality at the
top-$k$ eigencomponents of $S^+$, which are themselves positive
semidefinite of rank at most $k$.  Adding
$\|S^-\|_F^2=\sum_j\{(-\lambda_j)^+\}^2$ and using
$\operatorname{tr}(S^2)=\sum_j(\lambda_j^+)^2+\sum_j\{(-\lambda_j)^+\}^2$ gives the
displayed minimum.  Uniqueness of the Eckart--Young projection holds
exactly when $\lambda_k^+>\lambda_{k+1}^+$ or $\lambda_k^+=0$; the stated
qualification follows.
\end{proof}

Applied to the covariance discrepancy $S=\hat R_n-I_p$, removing the
largest positive eigencomponents sets the corresponding eigenvalues of $S$
to zero, equivalently resets the corresponding eigenvalues of $\hat R_n$ to
one; it does not set eigenvalues of $\hat R_n$ to zero, which would produce
a singular residual covariance farther from the identity.  These are
deterministic projection facts: they do not by themselves provide the null
distribution of any statistic computed after data-dependent deflation,
where factor estimation, rescaling, and sequential stopping all matter.

\section{The weak-factor experiment and the BBP boundary}
\label{app:weak-factor}

This section proves Theorem~4 of the main paper.  Throughout,
$\gamma_n=p/n\to\gamma\in(0,\infty)$, and $\vartheta=\vartheta_n>0$ is the
factor strength, and $\varepsilon=\vartheta_n/p$.  By the whitening
isomorphism (Lemma~1) it suffices to argue at the identity null
covariance: setting $Y_t=\Sigma_0^{-1/2}X_t$ carries the experiment onto
its whitened version, in which
\[
H_0:\ Y_t\stackrel{\mathrm{i.i.d.}}{\sim}N_p\big(0,(1+\varepsilon)I_p\big)
\quad\text{against}\quad
H_1:\ Y_t\mid f\stackrel{\mathrm{i.i.d.}}{\sim}N_p(0,I_p+ff^\prime),
\ \ f\sim N_p\big(0,\tfrac{\vartheta_n}{p}I_p\big);
\]
under $H_1$ the factor $f$ is drawn first, independently of the Gaussian
innovations that generate the conditional sample, so the sample law
depends on $f$ and the observed sample is a mixture over $f$.  The
Jacobian cancels from every likelihood ratio, and each statistic below is
exactly pivotal (translate to a general known $\Sigma_0$ by
$x'y\mapsto x^\prime\Sigma_0^{-1}y$).  Write
\[
\kappa=\frac{1+2\varepsilon}{1+\varepsilon},\qquad
\mu=\frac{\varepsilon}{1+2\varepsilon},\qquad
\sigma_n=\frac{\vartheta_n^2}{2\gamma_n},
\]
$\hat S_Y=n^{-1}\sum_tY_tY_t^\prime$, $T_n=\sum_t\|Y_t\|^2$, and the standardized trace
statistic
\[
\zeta_n=\frac{T_n-np(1+\varepsilon)}{\sqrt{2np}(1+\varepsilon)} .
\]
Let $P_0$ be the law of $Y_1,\ldots,Y_n$ under $H_0$, $P_f$ the conditional law
under $H_1$ given $f$, $\pi=N_p(0,(\vartheta_n/p)I_p)$,
$\ell(f)=\mathrm dP_f/\mathrm dP_0$, $\mathrm{LR}_n=\int\ell(f)\pi(\mathrm df)$,
and $G_n$ the mixture law of the sample under $H_1$.  The corrected Frobenius
statistic evaluated at the null covariance $(1+\varepsilon)I_p$ is
$V_n=2\tilde U_n/\gamma_n$, where
$\tilde U_n=(2n^2)^{-1}\sum_{s<t}h_\varepsilon(Y_s,Y_t)$ with
\[
h_\varepsilon(y,\tilde y)
=\frac{(y^\prime\tilde y)^2}{(1+\varepsilon)^2}
-\frac{\|y\|^2+\|\tilde y\|^2}{1+\varepsilon}+p ;
\]
by Lemma~1 (with $\Sigma_0=(1+\varepsilon)I_p$) and Theorem~1(i),
$V_n\Rightarrow N(0,1)$ under $P_0$.  We take $0<\vartheta_n\le\vartheta_0$ for a
small absolute constant $\vartheta_0\le\tfrac18$, $p\ge2$, $n\ge2$; constants $C$
depend only on $\gamma$-bounds and $\vartheta_0$ and may change from line to line.
Everything in Sections~S.8.1--S.8.4 is exact and self-contained; the
fixed-strength picture at the BBP boundary is set out at the end
(Section~S.8.5) as a comparison with imported random-matrix limits, following
\citet{OnatskiMoreiraHallin:2013}.

\subsection{Exact structure of the likelihood ratio}

Conditional on $f$, both laws are centered Gaussian, so with
$(I_p+ff^\prime)^{-1}=I_p-ff^\prime/(1+\|f\|^2)$ and $\det(I_p+ff^\prime)=1+\|f\|^2$,
\begin{equation}
\log\ell(f)
=\underbrace{\frac{np}2\log(1+\varepsilon)-\frac{\varepsilon}{2(1+\varepsilon)}T_n}_{A_n\ \text{(scale term)}}
-\frac n2\log(1+\|f\|^2)+\frac n2\frac{f^\prime\hat S_Yf}{1+\|f\|^2},
\label{eq:s8-condLR}
\end{equation}
using $\sum_tY_t'ff'Y_t=nf^\prime\hat S_Yf$.

\begin{lemma}[Exact orthogonality of the trace channel]
\label{lem:s8-orth}
For every $n,p,\vartheta_n$: $\mathbb{E}_{P_0}\ell(f)=1$ for each fixed $f$;
$\mathbb{E}_{G_n}\zeta_n=0$; and
$\operatorname{cov}_{P_0}(\mathrm{LR}_n,\zeta_n)=0$.  These are exact identities.
\end{lemma}

\begin{proof}
$\ell(f)$ is a ratio of Gaussian densities of the same sample, so
$\mathbb{E}_{P_0}\ell(f)=\int\mathrm dP_f=1$.  Under $P_f$,
$\mathbb{E}\|Y_t\|^2=\operatorname{tr}(I_p+ff^\prime)=p+\|f\|^2$, hence
$\mathbb{E}_{P_f}\zeta_n=n(\|f\|^2-\vartheta_n)/\{\sqrt{2np}(1+\varepsilon)\}$;
since $\mathbb{E}_\pi\|f\|^2=p\cdot(\vartheta_n/p)=\vartheta_n$ exactly, Fubini
gives $\mathbb{E}_{G_n}\zeta_n=\mathbb{E}_\pi\mathbb{E}_{P_f}\zeta_n=0$.  Finally
$\mathbb{E}_{P_0}[\mathrm{LR}_n\zeta_n]
=\mathbb{E}_\pi\mathbb{E}_{P_0}[\ell(f)\zeta_n]
=\mathbb{E}_\pi\mathbb{E}_{P_f}\zeta_n=0$ and $\mathbb{E}_{P_0}\zeta_n=0$.
\end{proof}

\begin{lemma}[Exact expansion of the scale term]
\label{lem:s8-scale}
For every realization,
$A_n=-\vartheta_n(2\gamma_n)^{-1/2}\zeta_n-\vartheta_n^2/(4\gamma_n)+r_n$ with
$r_n$ deterministic and $|r_n|\le\vartheta_n^3/(6p\gamma_n)$.
\end{lemma}

\begin{proof}
Substituting $T_n=np(1+\varepsilon)+\sqrt{2np}(1+\varepsilon)\zeta_n$ into
(\ref{eq:s8-condLR}) gives
$A_n=\tfrac{np}2[\log(1+\varepsilon)-\varepsilon]-\tfrac\varepsilon2\sqrt{2np}\zeta_n$.
Now $\log(1+\varepsilon)-\varepsilon=-\varepsilon^2/2+r$, $0\le r\le\varepsilon^3/3$;
$\tfrac{np}2\cdot\tfrac{\varepsilon^2}2=n\vartheta_n^2/(4p)=\vartheta_n^2/(4\gamma_n)$,
$\tfrac\varepsilon2\sqrt{2np}=\vartheta_n\sqrt{n/(2p)}=\vartheta_n/\sqrt{2\gamma_n}$,
and $|npr/2|\le np\varepsilon^3/6=\vartheta_n^3/(6p\gamma_n)$.
\end{proof}

Lemmas~\ref{lem:s8-orth}--\ref{lem:s8-scale} identify the mechanism.
Lemma~\ref{lem:s8-orth} is an exact statement about the likelihood ratio
$\mathrm{LR}_n$ itself: it is uncorrelated with $\zeta_n$ under $P_0$,
which is the finite-sample form of the channel orthogonality in Lemma~2
of the main paper.  The additive decomposition of $\log\ell(f)$ is a
different object, and $\operatorname{cov}(\mathrm{LR}_n,\zeta_n)=0$ does
not by itself give $\operatorname{cov}(\log\mathrm{LR}_n,\zeta_n)=0$.
What the lemmas show is that the scale term loads on $\zeta_n$ with
coefficient $-\vartheta_n/\sqrt{2\gamma_n}$ and that the $\zeta_n$
loading of the mixed factor term is $+\vartheta_n/\sqrt{2\gamma_n}$ to
first order in the small-strength expansion, so the linear $\zeta_n$
contribution to the log likelihood cancels at first order; the exact
integrations in the second-moment proof below do not rely on this
first-order reading.  In the unmatched variant,
with null $(1+\vartheta_n)\Sigma_0$, the same computation gives
$\mathbb{E}_{G_n}\zeta_n\asymp-\vartheta_n\sqrt{np}$, so the trace channel would
dominate and force the rate $\vartheta_n\asymp(np)^{-1/2}$; matching removes this
channel and, with it, the rate constraint.

\subsection{Second-moment machinery (exact and self-contained)}

\begin{lemma}[Gaussian ratio integral]
\label{lem:s8-ratio}
Let $A,B,C\succ0$ and $\varphi_\Sigma$ the $N_p(0,\Sigma)$ density.  If
$M:=A^{-1}+B^{-1}-C^{-1}\succ0$ then
\[
\int\frac{\varphi_A(x)\varphi_B(x)}{\varphi_C(x)}\mathrm dx
=\frac{|C|^{1/2}}{|A|^{1/2}|B|^{1/2}|M|^{1/2}},
\]
and the integral is $+\infty$ if $M$ has a nonpositive eigenvalue.
\end{lemma}

\begin{proof}
The integrand equals
$(2\pi)^{-p/2}|C|^{1/2}|A|^{-1/2}|B|^{-1/2}\exp\{-\tfrac12x'Mx\}$; integrate.
\end{proof}

\begin{lemma}[Exact pair identity]
\label{lem:s8-pair}
Let $a=\|f\|^2$, $b=\|\tilde f\|^2$, $c=f^\prime\tilde f$, and set
$\mathfrak a=\kappa(1+\mu a)$, $\mathfrak b=\kappa(1+\mu b)$.  If
$a,b\le\tfrac14$ then, per observation,
\[
J(f,\tilde f)
:=\int\frac{\varphi_{I+ff^\prime}\varphi_{I+\tilde f\tilde f^\prime}}{\varphi_{(1+\varepsilon)I}}\mathrm dx
=(1+\varepsilon)^{p/2}\kappa^{-(p-2)/2}(\mathfrak a\mathfrak b-c^2)^{-1/2},
\]
and consequently $\mathbb{E}_{P_0}[\ell(f)\ell(\tilde f)]=J^n$.
\end{lemma}

\begin{proof}
Apply Lemma~\ref{lem:s8-ratio} with $A=I+ff^\prime$, $B=I+\tilde f\tilde f^\prime$,
$C=(1+\varepsilon)I$.  Since $2-1/(1+\varepsilon)=\kappa$,
$M=\kappa I-\alpha ff^\prime-\beta\tilde f\tilde f^\prime$ with $\alpha=1/(1+a)$,
$\beta=1/(1+b)$.  The nonzero spectrum of $\alpha ff^\prime+\beta\tilde f\tilde f^\prime$
equals that of $\bigl(\begin{smallmatrix}\alpha a&\sqrt{\alpha\beta}c\\
\sqrt{\alpha\beta}c&\beta b\end{smallmatrix}\bigr)$, so
\[
|M|=\kappa^{p-2}\{(\kappa-\alpha a)(\kappa-\beta b)-\alpha\beta c^2\}
=\kappa^{p-2}\frac{[(1+a)\kappa-a][(1+b)\kappa-b]-c^2}{(1+a)(1+b)} .
\]
Now $(1+a)\kappa-a=\kappa+a(\kappa-1)=\kappa(1+\mu a)=\mathfrak a$ (using
$\kappa-1=\kappa\mu$), and likewise $(1+b)\kappa-b=\mathfrak b$, so
$|M|=\kappa^{p-2}(\mathfrak a\mathfrak b-c^2)/\{(1+a)(1+b)\}$.  On $a,b\le\tfrac14$,
$\alpha a+\beta b\le\tfrac25<1\le\kappa$, so $M\succ0$.  Assembling
Lemma~\ref{lem:s8-ratio} with $|A|=1+a$, $|B|=1+b$, $|C|=(1+\varepsilon)^p$, the
factors $(1+a)(1+b)$ cancel and the displayed formula follows.  The $n$
observations are i.i.d., so $\mathbb{E}_{P_0}[\ell(f)\ell(\tilde f)]=J^n$.
\end{proof}

\begin{lemma}[Radial truncation]
\label{lem:s8-trunc}
Let $\mathcal F=\{\|f\|^2\le\tfrac14\}$, $\theta_p=\pi(\mathcal F^c)$, and
$\widetilde{\mathrm{LR}}_n=\int_{\mathcal F}\ell(f)\pi(\mathrm df)+\theta_p$,
which keeps $\mathbb{E}_{P_0}\widetilde{\mathrm{LR}}_n=1$.  Then for
$\vartheta_n\le\tfrac18$,
$\mathbb{E}_{P_0}|\mathrm{LR}_n-\widetilde{\mathrm{LR}}_n|\le2\theta_p$ and
$\theta_p\le\exp\{-\tfrac p8(\tfrac1{4\vartheta_n}-1)\}\le e^{-p/8}$.
\end{lemma}

\begin{proof}
$\mathbb{E}_{P_0}|\mathrm{LR}_n-\widetilde{\mathrm{LR}}_n|
\le\mathbb{E}_\pi[\mathbb{E}_{P_0}\ell(f);\mathcal F^c]+\theta_p=2\theta_p$ by
Lemma~\ref{lem:s8-orth}.  With $\|f\|^2=(\vartheta_n/p)\chi^2_p$,
$\theta_p=\mathbb{P}(\chi^2_p\ge p/(4\vartheta_n))$; the Chernoff bound
$\mathbb{P}(\chi^2_p\ge py)\le(ye^{1-y})^{p/2}$ for $y\ge1$, at
$y=1/(4\vartheta_n)\ge2$, gives
$\theta_p\le\exp\{-\tfrac p8(\tfrac1{4\vartheta_n}-1)\}$, using
$y-1-\log y\ge(y-1)/4$ for $y\ge2$.
\end{proof}

The untruncated second moment $\mathbb{E}_{P_0}\mathrm{LR}_n^2$ is $+\infty$ for
every $n$ (on the $\pi\otimes\pi$-positive event $\{\mathfrak a\mathfrak b\le c^2\}$
the integral of Lemma~\ref{lem:s8-ratio} diverges); the radial truncation
of Lemma~\ref{lem:s8-ratio} removes this tail at a cost of only
$e^{-p/8}$ in $L^1(P_0)$.

\subsection{Moments of the corrected Frobenius statistic}

\begin{proposition}[Exact moments]
\label{prop:s8-moments}
For every $n,p,\varepsilon$:
(i) $\mathbb{E}_{P_0}V_n=0$ and
$\operatorname{var}_{P_0}(V_n)=\dfrac{(n-1)(p+1)}{np}\to1$;
(ii) for every fixed $f$, with $\Omega=I_p+ff^\prime$,
\[
\mathbb{E}_{P_f}V_n
=\frac{n-1}{2n\gamma_n}\operatorname{tr}\Big[\Big(\frac{\Omega}{1+\varepsilon}-I_p\Big)^2\Big]
=\frac{n-1}{2n\gamma_n(1+\varepsilon)^2}\big[\|f\|^4-2\varepsilon\|f\|^2+p\varepsilon^2\big];
\]
(iii) $\mathbb{E}_{G_n}V_n=\sigma_n(1+O(n^{-1}+p^{-1}+\varepsilon))$, with
$\sigma_n=\vartheta_n^2/(2\gamma_n)$.
\end{proposition}

\begin{proof}
(i) By pivotality it suffices to take $\varepsilon=0$ and $Z_t$ i.i.d.\ $N_p(0,I)$
with kernel $h(z,\tilde z)=(z^\prime\tilde z)^2-\|z\|^2-\|\tilde z\|^2+p$.  Degeneracy
$\mathbb{E}[h(z,Z)]=0$ (Section~S.8.1 form with $\varepsilon=0$) kills all
covariances between pairs sharing at most one index, so
$\operatorname{var}(\tilde U_n)=\tfrac1{4n^4}\binom n2\mathbb{E}h^2$.  With
$Q=(Z_1'Z_2)^2$ and $R=\|Z_1\|^2+\|Z_2\|^2-p$, conditioning on $Z_2$ gives
$\mathbb{E}Q=p$, $\mathbb{E}Q^2=3p(p+2)$, $\mathbb{E}[Q\|Z_2\|^2]=\mathbb{E}[Q\|Z_1\|^2]=p(p+2)$,
$\mathbb{E}R=p$, $\mathbb{E}R^2=p^2+4p$, whence
$\mathbb{E}h^2=3p(p+2)-2\{2p(p+2)-p^2\}+p^2+4p=2p(p+1)$ and
$\operatorname{var}(\tilde U_n)=(n-1)p(p+1)/(4n^3)$, so
$\operatorname{var}(V_n)=4\gamma_n^{-2}\operatorname{var}(\tilde U_n)=(n-1)(p+1)/(np)$.
(ii) For $s\ne t$ with $Y\sim N_p(0,\Omega)$,
$\mathbb{E}(Y_s'Y_t)^2=\operatorname{tr}(\Omega^2)$ and $\mathbb{E}\|Y\|^2=\operatorname{tr}(\Omega)$, so
$\mathbb{E}h_\varepsilon=\operatorname{tr}(\Omega^2)/(1+\varepsilon)^2-2\operatorname{tr}(\Omega)/(1+\varepsilon)+p
=\operatorname{tr}[(\Omega/(1+\varepsilon)-I)^2]$; multiply by $\binom n2/(2n^2)\cdot2/\gamma_n$
and expand $\operatorname{tr}[(ff^\prime-\varepsilon I)^2]=\|f\|^4-2\varepsilon\|f\|^2+p\varepsilon^2$.
(iii) uses $\mathbb{E}_\pi\|f\|^4=\vartheta_n^2(1+2/p)$,
$\mathbb{E}_\pi\|f\|^2=\vartheta_n$, $p\varepsilon^2=\vartheta_n^2/p$; the bracket
integrates to $\vartheta_n^2(1+O(p^{-1}))$, and
$(n-1)/(2n\gamma_n(1+\varepsilon)^2)=\gamma_n^{-1}(1+O(n^{-1}+\varepsilon))/2$.
\end{proof}

\subsection{The decay regime}

Throughout this subsection $\vartheta_n\to0$; for the sharp expansion we assume in
addition $n\vartheta_n\to\infty$ (equivalently $\varepsilon=o(\sigma_n)$).

\begin{proposition}[Second moment]
\label{prop:s8-second}
As $\vartheta_n\to0$, $\mathbb{E}_{P_0}[\widetilde{\mathrm{LR}}_n^2]\to1$; if
moreover $n\vartheta_n\to\infty$, then
$\mathbb{E}_{P_0}[\widetilde{\mathrm{LR}}_n^2]=1+\sigma_n^2(1+o(1))$.
\end{proposition}

\begin{proof}
Since $\widetilde{\mathrm{LR}}_n=\int_{\mathcal F}\ell\mathrm d\pi+\theta_p$
and $\mathbb{E}_{P_0}\int_{\mathcal F}\ell\mathrm d\pi=\pi(\mathcal F)$,
Lemma~\ref{lem:s8-pair} and Fubini's theorem give
\[
\mathbb{E}_{P_0}[\widetilde{\mathrm{LR}}_n^2]
=\mathbb{E}_{\pi\otimes\pi}[J^n;\mathcal F\times\mathcal F]
+2\theta_p\pi(\mathcal F)+\theta_p^2
=\mathbb{E}_{\pi\otimes\pi}[J^n;\mathcal F\times\mathcal F]+O(e^{-p/8}).
\]
When $n\vartheta_n\to\infty$ we have $\vartheta_n\ge n^{-1}$ eventually,
hence $\sigma_n\ge(2np)^{-1}$ and $\log\sigma_n^{-1}=O(\log n)$, such that
$e^{-cp}=o(\sigma_n^2)$ for every fixed $c>0$; this is used repeatedly.
For the plain contiguity bound only $\theta_p\to0$ is needed.

\emph{Step 1: coordinates and the exactly integrable part.}
Fix $\tilde f\ne0$, put $u=\tilde f/\|\tilde f\|$, $s=u^\prime f$,
$f_\perp=f-su$, and $r=\|f_\perp\|^2$.  Because $f\sim N_p(0,\varepsilon I_p)$
is independent of $\tilde f$ and rotation invariant, conditionally on
$\tilde f$ the variables $s\sim N(0,\varepsilon)$ and
$r\sim\varepsilon\chi^2_{p-1}$ are independent, and
\[
a=s^2+r,\qquad c=s\sqrt b,\qquad c^2=s^2b .
\]
Let $A_n=\tfrac{np}2[\log(1+\varepsilon)-\log\kappa]
=\tfrac{np}2\{2\log(1+\varepsilon)-\log(1+2\varepsilon)\}$; since
$np\varepsilon^2=2\sigma_n$,
\begin{equation}
A_n=\tfrac{np}2\{\varepsilon^2-2\varepsilon^3+O(\varepsilon^4)\}
=\sigma_n+O(\sigma_n\varepsilon).
\label{eq:s8-An}
\end{equation}
With $J$ from Lemma~\ref{lem:s8-pair}, using
$\log\mathfrak a=\log\kappa+\log(1+\mu a)$ and recombining $p\log\kappa$,
\[
n\log J
=A_n-\tfrac n2\log(1+\mu a)-\tfrac n2\log(1+\mu b)
-\tfrac n2\log(1-w),
\qquad w=\frac{c^2}{\mathfrak a\mathfrak b},
\]
which we split as $n\log J=E_0+\rho$ with
\[
E_0
=A_n-\frac{n\varepsilon}2b-\frac{n\varepsilon}2r+\frac n2(b-\varepsilon)s^2
=A_n-\frac{n\varepsilon}2(a+b)+\frac n2c^2
\]
and
\[
\rho=-\frac n2\{\log(1+\mu a)-\varepsilon a\}
-\frac n2\{\log(1+\mu b)-\varepsilon b\}
-\frac n2\{\log(1-w)+c^2\}.
\]
The $s^2$-part of $a$ contributes $-n\varepsilon s^2/2$ to $E_0$; it is
integrated jointly with $nc^2/2=nbs^2/2$ and is not dropped: its mean
$n\varepsilon^2/2=\sigma_n/p$ is not $o(\sigma_n^2)$ unless
$n\vartheta_n^2\to\infty$, and it is the cancellation below that removes
it.  Put $x=n\varepsilon(b-\varepsilon)$ and $y=n\varepsilon^2=2\sigma_n/p$.
If $x<1$, the Gaussian and chi-square moment generating functions give,
conditionally on $\tilde f$,
\begin{equation}
\mathbb{E}_{s,r}\,e^{E_0}
=\exp\Big\{A_n-\frac{n\varepsilon b}2-\frac12\log(1-x)
-\frac{p-1}2\log(1+y)\Big\},
\label{eq:s8-exact-E0}
\end{equation}
and under the tilted law with density $e^{E_0}/\mathbb{E}_{s,r}e^{E_0}$ the
variables $s$ and $r$ remain independent, with
$s^2\sim\{\varepsilon/(1-x)\}\chi^2_1$ and
$r\sim\{\varepsilon/(1+y)\}\chi^2_{p-1}$.  Writing $\mathbb{E}^\star$ for
expectation under this tilt, if $x\le\tfrac12$ then, for every fixed $k$,
\begin{equation}
\mathbb{E}^\star s^{2k}\le C_k\varepsilon^k,\qquad
\mathbb{E}^\star r^k\le C_k(p\varepsilon)^k,\qquad
\mathbb{E}^\star a^k\le C_k(p\varepsilon)^k .
\label{eq:s8-tilted-moments}
\end{equation}

\emph{Step 2: the remainder on $\mathcal F\times\mathcal F$.}
On $a,b\le\tfrac14$ we have $|\log(1+\mu a)-\mu a|\le\mu^2a^2/2
\le\varepsilon^2a^2/2$ and $|\mu-\varepsilon|a\le2\varepsilon^2a$;
moreover $\mathfrak a\mathfrak b=\kappa^2(1+\mu a)(1+\mu b)\in[1,1+3\varepsilon]$,
so $0\le w\le c^2\le ab\le\tfrac1{16}$, $|w-c^2|\le3\varepsilon c^2$, and
$w\le-\log(1-w)\le w+w^2$.  Hence
\begin{equation}
|\rho|\le\bar\rho
:=\frac{n\varepsilon^2}4(a^2+b^2)+n\varepsilon^2(a+b)
+\frac{3n\varepsilon}2c^2+\frac n2c^4 .
\label{eq:s8-rho-bound}
\end{equation}

\emph{Step 3: a good event and the complementary set.}
Let $\tau_n\to\infty$ satisfy $\varepsilon\tau_n\to0$ and
$\tau_n^2\vartheta_n^2/p\to0$; for the sharp expansion take
$\tau_n=12\log(p/\sigma_n)=O(\log n)$.  Define
\[
\mathcal E=\{b\le\tfrac32\vartheta_n\}\cap\{r\le\tfrac32\vartheta_n\}
\cap\{s^2\le\varepsilon\tau_n\},
\]
which is contained in $\mathcal F\times\mathcal F$ for all large $n$.  On
$\mathcal E$, $x\le\tfrac32n\varepsilon\vartheta_n=3\sigma_n$ and, by
$n\varepsilon^2=2\sigma_n/p$ and $n\varepsilon\vartheta_n=2\sigma_n$,
$\bar\rho\le C\sigma_n(\vartheta_n^2/p+\varepsilon+\varepsilon\tau_n
+\tau_n^2\vartheta_n^2/p)\to0$, such that $|e^{\rho}-1|\le3\bar\rho$ on
$\mathcal E$ for all large $n$.

On $\mathcal F\times\mathcal F$, dropping the two nonpositive logarithms,
$J^n\le e^{A_n}\exp\{-\tfrac n2\log(1-w)\}\le e^{A_n}\exp(\tfrac{2n}{15}s^2)$,
because $-\log(1-w)\le\tfrac{16}{15}w$ for $w\le\tfrac1{16}$ and
$w\le s^2b\le s^2/4$.  Since $n\varepsilon=\vartheta_n/\gamma_n\to0$,
$\mathbb{E}_s\exp(\tfrac{2n}{15}s^2)=(1-\tfrac{4n\varepsilon}{15})^{-1/2}\le2$ and
$\mathbb{E}_s[\exp(\tfrac{2n}{15}s^2);s^2>\varepsilon\tau_n]
\le2\mathbb{P}(\chi^2_1>\tau_n/2)\le4e^{-\tau_n/4}$ for all large $n$.
Because $(\mathcal F\times\mathcal F)\setminus\mathcal E$ is covered by
$\{b>\tfrac32\vartheta_n\}\cup\{r>\tfrac32\vartheta_n\}\cup\{s^2>\varepsilon\tau_n\}$,
the independence of $s$, $r$, and $\tilde f$ and the Chernoff bound
$\mathbb{P}(\chi^2_m\ge\tfrac32m)\le e^{-cm}$ give
\begin{equation}
\mathbb{E}[J^n;(\mathcal F\times\mathcal F)\setminus\mathcal E]
\le Ce^{A_n}\{e^{-cp}+e^{-\tau_n/4}\}.
\label{eq:s8-bad-set}
\end{equation}

\emph{Step 4: the leading integral.}
On $\{b\le\tfrac32\vartheta_n\}$, $-y\le x\le3\sigma_n$, so
$-\tfrac12\log(1-x)=\tfrac x2+\tfrac{x^2}4+O(|x|^3)$ and
$-\tfrac{p-1}2\log(1+y)=-\tfrac{(p-1)y}2+O(py^2)$, and
(\ref{eq:s8-exact-E0}) becomes
\begin{align*}
\log\mathbb{E}_{s,r}e^{E_0}
&=A_n-\frac{n\varepsilon b}2+\frac{n\varepsilon(b-\varepsilon)}2
-\frac{(p-1)n\varepsilon^2}2+\frac{x^2}4+O(\sigma_n^3+\sigma_n^2/p)\\
&=\frac{x^2}4+O(\sigma_n\varepsilon+\sigma_n^3+\sigma_n^2/p),
\end{align*}
uniformly on $\{b\le\tfrac32\vartheta_n\}$: the terms linear in $b$ cancel
exactly, and the deterministic linear terms sum to
$A_n-np\varepsilon^2/2=A_n-\sigma_n$, which is $O(\sigma_n\varepsilon)$ by
(\ref{eq:s8-An}).  This is Lemma~\ref{lem:s8-orth} reappearing at the
second-moment level: the two apparent contributions of order $\sigma_n/p$,
from $-n\varepsilon s^2/2$ and from the $p-1$ orthogonal coordinates of
$f$, cancel against each other.  Since $x^2/4\le\tfrac94\sigma_n^2\to0$,
$\mathbb{E}_{s,r}e^{E_0}=1+x^2/4+O(\sigma_n\varepsilon+\sigma_n^3+\sigma_n^2/p)$
uniformly on $\{b\le\tfrac32\vartheta_n\}$.  As $b\sim\varepsilon\chi^2_p$,
$\mathbb{E}x^2=n^2\varepsilon^2\mathbb{E}(b-\varepsilon)^2
=n^2\varepsilon^4(p^2+1)=4\sigma_n^2(1+p^{-2})$, and the Cauchy--Schwarz
inequality bounds the contribution of $\{b>\tfrac32\vartheta_n\}$ to
$\mathbb{E}x^2$ by $O(\sigma_n^2e^{-cp/2})$.  Therefore
\begin{equation}
\mathbb{E}[e^{E_0};b\le\tfrac32\vartheta_n]
=1+\sigma_n^2+O(\sigma_n\varepsilon+\sigma_n^3+\sigma_n^2/p+e^{-cp}).
\label{eq:s8-leading}
\end{equation}
The same exact integrals control the part of $\{b\le\tfrac32\vartheta_n\}$
outside $\mathcal E$: $\mathbb{E}_r[e^{-n\varepsilon r/2};r>\tfrac32\vartheta_n]
\le\mathbb{P}(r>\tfrac32\vartheta_n)\le e^{-cp}$ and
$\mathbb{E}_s[e^{n(b-\varepsilon)s^2/2};s^2>\varepsilon\tau_n]
=(1-x)^{-1/2}\mathbb{P}\{\chi^2_1>(1-x)\tau_n\}\le4e^{-\tau_n/4}$, while the
remaining factors in (\ref{eq:s8-exact-E0}) are bounded by $Ce^{A_n}$.
Hence
\begin{equation}
\mathbb{E}[e^{E_0};\mathcal E]
=\mathbb{E}[e^{E_0};b\le\tfrac32\vartheta_n]+O(e^{-cp}+e^{-\tau_n/4}).
\label{eq:s8-E0-good}
\end{equation}

\emph{Step 5: the remainder under the tilt.}
By (\ref{eq:s8-exact-E0}), $\mathbb{E}_{s,r}e^{E_0}\le e^{A_n}(1-x)^{-1/2}\le C$
on $\{b\le\tfrac32\vartheta_n\}$.  Since $e^{E_0}\bar\rho\ge0$, we may
enlarge $\mathcal E$ to $\{b\le\tfrac32\vartheta_n\}$, integrate $s$ and
$r$ first, and use (\ref{eq:s8-tilted-moments}),
(\ref{eq:s8-rho-bound}), and $b\le\tfrac32p\varepsilon$:
\begin{align*}
\mathbb{E}[e^{E_0}\bar\rho;\mathcal E]
&\le C\mathbb{E}_b\big[\mathbb{E}^\star\bar\rho;b\le\tfrac32\vartheta_n\big]
\le C\{n\varepsilon^2(p\varepsilon)^2+n\varepsilon^2p\varepsilon
+n\varepsilon\cdot p\varepsilon\cdot\varepsilon+n(p\varepsilon)^2\varepsilon^2\}\\
&=O(\sigma_n\vartheta_n^2/p+\sigma_n\varepsilon)
=O(\sigma_n^2/n+\sigma_n\varepsilon),
\end{align*}
using $np^2\varepsilon^4=2\sigma_n\vartheta_n^2/p$,
$np\varepsilon^3=2\sigma_n\varepsilon$, and
$\vartheta_n^2/p=2\sigma_n/n$.

\emph{Step 6: assembly.}
On $\mathcal E$, $J^n=e^{E_0}e^{\rho}$ with $|e^{\rho}-1|\le3\bar\rho$, so
(\ref{eq:s8-bad-set}), (\ref{eq:s8-leading}), (\ref{eq:s8-E0-good}), and
Step~5 give
\[
\mathbb{E}[J^n;\mathcal F\times\mathcal F]
=1+\sigma_n^2
+O\big(\sigma_n\varepsilon+\sigma_n^3+\sigma_n^2/p+\sigma_n^2/n
+e^{-cp}+e^{-\tau_n/4}\big).
\]
If $n\vartheta_n\to\infty$, then $\varepsilon/\sigma_n=2/(n\vartheta_n)\to0$,
$e^{-cp}=o(\sigma_n^2)$, and $e^{-\tau_n/4}=(\sigma_n/p)^3=o(\sigma_n^2)$,
which proves the sharp expansion.  If only $\vartheta_n\to0$, take
$\tau_n=\log p$: every error term and $\sigma_n^2$ itself tend to zero, so
$\mathbb{E}[J^n;\mathcal F\times\mathcal F]\to1$.
\end{proof}

\begin{theorem}[Decay regime; Theorem~4]
\label{thm:s8-decay}
Let $\vartheta_n\to0$.  Then $G_n$ is contiguous to $P_0$ at every decay rate.
If in addition $n\vartheta_n\to\infty$, then
\[
\mathbb{E}_{P_0}[(\widetilde{\mathrm{LR}}_n-1-\sigma_nV_n)^2]=o(\sigma_n^2),
\qquad
\mathrm{LR}_n=1+\sigma_nV_n+o_{L^1(P_0)}(\sigma_n),
\]
and $\|G_n-P_0\|_{\mathrm{TV}}=\sigma_n/\sqrt{2\pi}+o(\sigma_n)$,
with $V_n\Rightarrow N(0,1)$ under $P_0$; no $L^2$ expansion holds for
$\mathrm{LR}_n$ itself, whose second moment is infinite.  Moreover, for
every $\alpha\in(0,1)$ and every sequence of tests $0\le\psi_n\le1$ with
$\mathbb{E}_{P_0}\psi_n\to\alpha$,
\begin{equation}
\limsup_{n\to\infty}\sigma_n^{-1}\big\{\mathbb{E}_{G_n}\psi_n-\mathbb{E}_{P_0}\psi_n\big\}
\le\Phi^\prime(z_{1-\alpha}),
\label{eq:s8-first-order}
\end{equation}
with equality for $\psi_n=1\{V_n>z_{1-\alpha}\}$, where $\Phi^\prime$ is
the standard normal density.  If the size constraint is weakened to
$\limsup_n\mathbb{E}_{P_0}\psi_n\le\alpha$, the right side of
(\ref{eq:s8-first-order}) is replaced by
$\sup_{0\le\beta\le\alpha}\Phi^\prime(z_{1-\beta})$, which equals
$\Phi^\prime(z_{1-\alpha})$ exactly when $\alpha\le\tfrac12$ and equals
$\Phi^\prime(0)$ otherwise.
\end{theorem}

\begin{proof}
Expanding the square and using $\mathbb{E}_{P_0}\widetilde{\mathrm{LR}}_n=1$,
$\mathbb{E}_{P_0}V_n=0$, $\operatorname{var}(V_n)=(n-1)(p+1)/(np)$
(Proposition~\ref{prop:s8-moments}(i)),
\[
\mathbb{E}(\widetilde{\mathrm{LR}}_n-1-\sigma_nV_n)^2
=\mathbb{E}\widetilde{\mathrm{LR}}_n^2-1-2\sigma_n\mathbb{E}[\widetilde{\mathrm{LR}}_nV_n]+\sigma_n^2\operatorname{var}(V_n).
\]
By Fubini's theorem and Proposition~\ref{prop:s8-moments}(ii)--(iii),
$\mathbb{E}[\widetilde{\mathrm{LR}}_nV_n]
=\mathbb{E}_\pi[\mathbb{E}_{P_f}V_n;\mathcal F]
=\mathbb{E}_{G_n}V_n-\mathbb{E}_\pi[\mathbb{E}_{P_f}V_n;\mathcal F^c]
=\sigma_n(1+o(1))$, because $|\mathbb{E}_{P_f}V_n|\le C(\|f\|^4+1)$ and
$\mathbb{E}_\pi[\|f\|^4+1;\mathcal F^c]\le C\vartheta_n^2\theta_p^{1/2}+\theta_p
=O(e^{-p/16})=o(\sigma_n)$ by the Cauchy--Schwarz inequality and
Lemma~\ref{lem:s8-trunc}.  With Proposition~\ref{prop:s8-second},
\[
\mathbb{E}(\widetilde{\mathrm{LR}}_n-1-\sigma_nV_n)^2
=[1+\sigma_n^2(1+o(1))]-1-2\sigma_n^2(1+o(1))+\sigma_n^2(1+o(1))=o(\sigma_n^2).
\]
Combining this $L^2$ expansion of the truncated surrogate with
$\mathbb{E}_{P_0}|\mathrm{LR}_n-\widetilde{\mathrm{LR}}_n|\le2e^{-p/8}=o(\sigma_n)$
from Lemma~\ref{lem:s8-trunc} gives
$\mathbb{E}_{P_0}|\mathrm{LR}_n-1-\sigma_nV_n|=o(\sigma_n)$, the $L^1$
expansion of the true likelihood ratio; in particular
$\mathrm{LR}_n=1+\sigma_nV_n+o_P(\sigma_n)$, and $\log(1+z)=z+O(z^2)$ with
$\sigma_nV_n=O_P(\sigma_n)\to0$ gives the equivalent logarithmic form
$\log\mathrm{LR}_n=\sigma_nV_n+o_P(\sigma_n)$ (the second-order term
$-\tfrac12\sigma_n^2$ is of order $o(\sigma_n)$ and is absorbed into the
remainder).
Contiguity for any $\vartheta_n\to0$: Proposition~\ref{prop:s8-second} gives
$\sup_n\mathbb{E}_{P_0}\widetilde{\mathrm{LR}}_n^2<\infty$, so for $A_n$ with
$P_0(A_n)\to0$,
$G_n(A_n)\le(\mathbb{E}\widetilde{\mathrm{LR}}_n^2)^{1/2}P_0(A_n)^{1/2}+2\theta_p\to0$.

\emph{Total variation.}  Since $V_n\Rightarrow N(0,1)$ with
$\sup_n\mathbb{E}_{P_0}V_n^2<\infty$, the family $\{V_n\}$ is uniformly
integrable and $\mathbb{E}_{P_0}|V_n|\to\sqrt{2/\pi}$; hence
$\|G_n-P_0\|_{\mathrm{TV}}=\tfrac12\mathbb{E}_{P_0}|\mathrm{LR}_n-1|
=\tfrac12\sigma_n\mathbb{E}_{P_0}|V_n|+o(\sigma_n)
=\sigma_n/\sqrt{2\pi}+o(\sigma_n)$.

\emph{First-order optimality.}  For any test $0\le\psi_n\le1$, the $L^1$
expansion gives
\[
\mathbb{E}_{G_n}\psi_n-\mathbb{E}_{P_0}\psi_n
=\mathbb{E}_{P_0}[(\mathrm{LR}_n-1)\psi_n]
=\sigma_n\mathbb{E}_{P_0}[V_n\psi_n]+o(\sigma_n),
\]
with a remainder bounded by $\mathbb{E}_{P_0}|\mathrm{LR}_n-1-\sigma_nV_n|$,
uniformly over tests.  Let $\alpha_n=\mathbb{E}_{P_0}\psi_n\to\alpha\in(0,1)$,
let $q_n$ be an upper $\alpha_n$-quantile of $V_n$ under $P_0$, and let
$\psi_n^\circ=1\{V_n>q_n\}+\eta_n1\{V_n=q_n\}$ with $\eta_n\in[0,1]$
chosen such that $\mathbb{E}_{P_0}\psi_n^\circ=\alpha_n$.  Since
$(V_n-q_n)(\psi_n^\circ-\psi_n)\ge0$ pointwise,
\[
\mathbb{E}_{P_0}[V_n\psi_n^\circ]-\mathbb{E}_{P_0}[V_n\psi_n]
=\mathbb{E}_{P_0}[(V_n-q_n)(\psi_n^\circ-\psi_n)]\ge0,
\]
which is the Neyman--Pearson argument with $V_n$ in the role of the
likelihood ratio.  Because the limit law of $V_n$ has a continuous,
strictly increasing distribution function, $q_n\to z_{1-\alpha}$ and
$P_0(V_n=q_n)\to0$; uniform integrability then gives
$\mathbb{E}_{P_0}[V_n\psi_n^\circ]\to\mathbb{E}[Z1\{Z>z_{1-\alpha}\}]
=\Phi^\prime(z_{1-\alpha})$ for $Z\sim N(0,1)$.  This proves
(\ref{eq:s8-first-order}).  The same uniform integrability gives
$P_0(V_n>z_{1-\alpha})\to\alpha$ and
$\mathbb{E}_{P_0}[V_n1\{V_n>z_{1-\alpha}\}]\to\Phi^\prime(z_{1-\alpha})$, which
is the equality case.  Under the weaker constraint
$\limsup_n\mathbb{E}_{P_0}\psi_n\le\alpha$, pass to a subsequence along which
$\alpha_n\to\beta\in[0,\alpha]$; the same argument bounds the limit of the
normalized gain by $\Phi^\prime(z_{1-\beta})$, with the value $0$ when
$\beta=0$ because then $q_n\to\infty$ and
$\mathbb{E}_{P_0}[V_n\psi_n^\circ]\to0$ by uniform integrability.  Since
$\beta\mapsto\Phi^\prime(z_{1-\beta})$ increases on $(0,\tfrac12]$ and
decreases on $[\tfrac12,1)$, the supremum over $\beta\le\alpha$ is
$\Phi^\prime(z_{1-\alpha})$ for $\alpha\le\tfrac12$ and $\Phi^\prime(0)$,
attained at asymptotic size $\tfrac12$, otherwise.
\end{proof}

\begin{proof}[Proof of Corollary~4 of the main paper]
Put $\varphi_n^V=1\{V_n>z_{1-\alpha}\}$ and
$L_n=\mathrm dG_n/\mathrm dP_0=1+\sigma_nV_n+r_n$ with
$\mathbb{E}_{P_0}|r_n|=o(\sigma_n)$, which is the $L^1$ expansion of
Theorem~4.  For any test $\psi$,
$\mathbb{E}_{G_n}\psi-\mathbb{E}_{P_0}\psi
=\mathbb{E}_{P_0}\{(L_n-1)\psi\}
=\sigma_n\mathbb{E}_{P_0}(V_n\psi)+\mathbb{E}_{P_0}(r_n\psi)$, so, by
Cauchy--Schwarz and $|\psi_n^F-\varphi_n^V|^2=|\psi_n^F-\varphi_n^V|$,
\[
\big|(\mathbb{E}_{G_n}-\mathbb{E}_{P_0})\psi_n^F
-(\mathbb{E}_{G_n}-\mathbb{E}_{P_0})\varphi_n^V\big|
\le\sigma_n\{\operatorname{var}_{P_0}(V_n)\}^{1/2}d_n^{1/2}
+\mathbb{E}_{P_0}|r_n|,
\]
where $d_n=\mathbb{E}_{P_0}|\psi_n^F-\varphi_n^V|$.
Under $P_0$ the whitened observations $\Sigma_0^{-1/2}X_t$ are i.i.d.\
$N_p(0,(1+\vartheta_n/p)I_p)$.  The scalar factor does not affect the
demeaned and studentized statistic $\tilde Z_n^c$, and $V_n$ is the
statistic $2U_n/\gamma_n$ of the corresponding standard Gaussian
sample, so Theorem~2(i)--(ii) of the main paper give
$\tilde Z_n^c-V_n=o_{P_0}(1)$.  Since $V_n\Rightarrow N(0,1)$ with a
continuous limit, $d_n\to0$.  With
$\operatorname{var}_{P_0}(V_n)=(n-1)(p+1)/(np)$ the right side is
$o(\sigma_n)$, and the equality case of Theorem~4 for $\varphi_n^V$
gives the claim.  No rate for $d_n$ is used, and the argument
concerns the gain relative to $\mathbb{E}_{P_0}\psi_n^F$, not
relative to the nominal level.
\end{proof}

\subsection{Fixed strength: comparison with Onatski--Moreira--Hallin (2013)}

At a fixed strength $\vartheta_n\to\vartheta\in(0,\infty)$ the trace-matched
experiment is no longer covered by the exact second-moment machinery above.
We record the picture as a comparison with the invariant fixed-radius
spiked experiment whose asymptotic power is characterized by
\citet{OnatskiMoreiraHallin:2013}, not as a theorem about the
random-radius mixture.  The connection is exact at the level of the
conditional experiment: given $r_n=\|f\|^2$, the direction $f/\|f\|$ is
Haar-uniform and independent of $r_n$, so the conditional alternative is
their fixed-radius alternative at strength $r_n$, tested here against
the trace-matched null, and $r_n\to_p\vartheta$ with
$r_n-\vartheta_n=O_p(\vartheta_n p^{-1/2})$.  Transferring their conclusions to
the mixture over $r_n$ would require, in addition, that their
likelihood-ratio process $L_n(\cdot)$ be asymptotically equicontinuous in
the strength near $\vartheta$, such that
$\int L_n(r)\pi_r(\mathrm dr)=L_n(\vartheta)\{1+o_{P_0}(1)\}$ for the law
$\pi_r$ of $r_n$ (the contribution of $\{|r_n-\vartheta|>\delta_n\}$ being
$o_{L^1(P_0)}(1)$ for suitable $\delta_n\to0$ because
$\mathbb{E}_{P_0}L_n(r)=1$), together with the joint limit of the trace term
generated by the matched null scale.  We do not carry this out, and no
theorem of the paper relies on it.  The imported statements are applied
to the trace-matched, rescaled matrix
$\bar S_Y=(1+\vartheta_n/p)^{-1}\hat S_Y$, defined with the finite-$n$
matching scale $\vartheta_n\to\vartheta$, whose null covariance is
exactly $I_p$; all spectral quantities below are those of $\bar S_Y$.
The statements below are theirs, imported and not reproved, and they
describe their fixed-radius experiment.

For a subcritical strength $\vartheta<\sqrt\gamma$ their experiment remains
contiguous and its asymptotically optimal test is a likelihood-based
\emph{linear spectral statistic} of $\bar S_Y$: a
centered sum $\sum_j\log\{z_0(\vartheta)-\bar\lambda_j\}$ over the
eigenvalues $\bar\lambda_j$ of $\bar S_Y$, with the trace correction
$\operatorname{tr}(\bar S_Y)-p$, organized around the
saddlepoint $z_0(\vartheta)=(1+\vartheta)(\gamma+\vartheta)/\vartheta$
\citep{OnatskiMoreiraHallin:2013,BaiSilverstein:2004}.  Explicitly, in
the form of equations (4.1)--(4.2) of
\citet{OnatskiMoreiraHallin:2013}, the log likelihood ratio of the
subcritical experiment is, up to deterministic centering,
\[
-\frac12\,\Delta_{p,n}\{z_0(\vartheta)\}
-\frac{\vartheta}{2\gamma}\,\big\{\operatorname{tr}(\bar S_Y)-p\big\}
+o_P(1),
\qquad
\Delta_{p,n}(z)=\sum_j\log\{z-\bar\lambda_j\}-c_{p,n}(z),
\]
with $c_{p,n}(z)$ the deterministic Mar\v{c}enko--Pastur centering.
The trace term arises from a density-ratio identity.  With
$a=\vartheta_n/p$, the likelihood ratio of the spiked alternative
against the matched null $N_p(0,(1+a)I_p)$ equals its likelihood ratio
against $N_p(0,I_p)$ times the ratio of the two null densities, so
\[
L_{\mathrm{match}}(\vartheta;\bar\lambda)
=\exp\Big\{\frac{np}{2}\log(1+a)-\frac{na}{2}\sum_j\bar\lambda_j\Big\}
\,L_{\mathrm{OMH}}\big(\vartheta;(1+a)\bar\lambda\big),
\]
where $L_{\mathrm{OMH}}$ is their known-scale likelihood ratio
evaluated at the eigenvalues $(1+a)\bar\lambda_j$ of $\hat S_Y$.  Since
$\frac{np}{2}\log(1+a)-\frac{na}{2}\sum_j\bar\lambda_j
=-\frac{\vartheta_n}{2\gamma_n}\{\operatorname{tr}(\bar S_Y)-p\}
-\frac{\vartheta_n^2}{4\gamma_n}+O\{\vartheta_n^3/(p\gamma_n)\}$,
the prefactor supplies the trace term of the display up to
deterministic centering.  Evaluating the centered spectral sum at
$(1+a)\bar\lambda_j$ rather than at $\bar\lambda_j$ changes it by
$-(\vartheta_n/p)\sum_j\bar\lambda_j/\{z_0-\bar\lambda_j\}+o_P(1)$,
a deterministic constant plus $o_P(1)$ that is absorbed in
$c_{p,n}$.  Their equation (4.2) concerns instead the trace-normalized
eigenvalues of the unknown-scale invariant experiment; the matched-null
experiment uses deterministic scale matching, and the two
normalizations are distinct.  As $\vartheta\downarrow0$
this statistic collapses onto its leading quadratic term: the linear ($k=1$)
spectral moment is removed by the exact cancellation of Section~S.8.1, and the
quadratic ($k=2$) term is $\sigma_nV_n+o_P(\sigma_n)$, recovering the corrected
Frobenius statistic $V_n$ of the decay regime.  Thus $V_n$ is first-order optimal
only as the strength vanishes; at fixed subcritical $\vartheta$ the higher
spectral moments carry information about the rank-one alternative beyond
its Frobenius norm, and $V_n$ is generally suboptimal.  The removal of
the linear spectral moment is the first-order counterpart of the exact
mean and covariance identities of Lemma~\ref{lem:s8-orth}: the linear
trace signal is removed at first order, and the quadratic-and-higher
spectral content remains.

For a supercritical strength $\vartheta>\sqrt\gamma$ contiguity fails in
the trace-matched experiment itself, without any transfer argument.
Conditional on $f$ the alternative is a spiked model with spike
$1+\|f\|^2$, and for fixed Gaussian innovations $\lambda_{\max}(\bar S_Y)$
is nondecreasing in the spike size, so $\|f\|^2\to_p\vartheta$ and the
Baik--Ben Arous--P\'ech\'e and Baik--Silverstein spike phase transition
\citep{BaikBenArousPeche:2005,BaikSilverstein:2006}, applied at the
spikes $1+\vartheta\pm\delta$, give
$\lambda_{\max}(\bar S_Y)\to_p(1+\vartheta)(1+\gamma/\vartheta)>(1+\sqrt\gamma)^2$
under $H_1$, while under $H_0$ the null edge is $(1+\sqrt\gamma)^2$; a
threshold test on $\lambda_{\max}$ then has size $\to0$ and power $\to1$,
so $\|G_n-P_0\|_{\mathrm{TV}}\to1$.  We do not analyze the critical case
$\vartheta=\sqrt\gamma$.

\section{Additional simulations: robustness and stress tests}
\label{supp:extra-sims}

This section reports the two descriptive stress tests referenced in
Section~7 of the main paper: a dense-spectrum and path-robustness
comparison, and a common-scale design outside the product-coordinate
null.  The simulation setup (sample sizes, replication counts, and
procedures) is that of the main paper.

\subsection{Dense-spectrum and path robustness}
\label{supp:robustness}

The GOE supplies both Haar eigenvectors and a random semicircle-type
spectrum.  Panel A of Table~\ref{tab:adm-robustness-main} retains Haar
eigenvectors but replaces the eigenvalues by two deterministic dense spectra
normalized such that $\operatorname{tr}(H^2)=p^2$.  Power is similar across the three
designs.  At $\omega=2$, the rejection frequencies are $0.557$, $0.562$,
and $0.582$, compared with the asymptotic GOE benchmark $0.639$.  This is a
descriptive robustness check, not a proof of fixed-radius or spectral
universality.

Panel B compares the canonical quadratic path with the exponential bridge
using the same GOE matrix and Gaussian innovations in every replication.
The two paths remain numerically very close throughout the reported grid;
their largest absolute power difference is $0.0004$.  This numerical
proximity supports the bridge as a finite-sample approximation, while the theorem and reported
power envelope continue to concern only the quadratic path.

\begin{table}[t]
\centering
\caption{Dense spectra and the exponential proof bridge ($n=p=200$)}
\label{tab:adm-robustness-main}
\footnotesize
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}lcccc}
\toprule
Design & $\omega=1$ & $\omega=1.5$ & $\omega=2$ & $\omega=2.5$\\
\midrule
\multicolumn{5}{l}{\emph{Panel A: direction spectrum, quadratic path}}\\
GOE spectrum &0.121&0.273&0.557&0.842\\
Fixed smooth spectrum &0.127&0.275&0.562&0.855\\
Balanced two-point spectrum &0.125&0.283&0.582&0.874\\
Asymptotic theory &0.126&0.302&0.639&0.931\\
\addlinespace
\multicolumn{5}{l}{\emph{Panel B: canonical path and proof bridge}}\\
Canonical quadratic path &0.121&0.273&0.557&0.842\\
Exponential proof bridge &0.121&0.273&0.557&0.841\\
\bottomrule
\end{tabular*}
\par
\vspace{0.05cm}
\begin{minipage}{\textwidth}
\footnotesize\emph{Notes:} Rejection frequencies of $\tilde Z_n^c$
from $10{,}000$ replications.  Deterministic spectra satisfy
$\operatorname{tr}(H^2)=p^2$.  Panel B uses common random numbers across paths.
\end{minipage}
\end{table}

\subsection{A stress test outside the null model}
\label{supp:stress}

The last experiment makes the scope of Assumption~2 of the main paper
visible.  Generate $X_{it}=\sigma_te_{it}$ with independent standard-normal
$e_{it}$ and $\sigma_t^2\sim\chi^2_\nu/\nu$.  The covariance and correlation
matrices are $I_p$, but coordinates are dependent through their squares.
The design therefore lies outside the maintained product-coordinate null.

Table~\ref{tab:adm-scale-main} separates centering from variance effects.
The data-driven correction keeps the mean of $\tilde Z_n^c$ relatively
close to zero, but the variance is inflated.  With $\nu=5$, for example, its
mean is $0.182$, its standard deviation is $1.351$, and its rejection
frequency is $0.137$.  Distortion declines as the common scale becomes less
variable.  Deterministic centering fails completely because
$\mathbb{E}(X_{it}^2X_{jt}^2)=\mathbb{E}(\sigma_t^4)>1$.  The experiment shows that correcting
the location does not by itself deliver the product-coordinate variance
formula under dependent-but-uncorrelated observations.

\begin{table}[t]
\centering
\caption{Common-scale stress test outside the product-coordinate null
($n=p=200$)}
\label{tab:adm-scale-main}
\footnotesize
\begin{tabular*}{\textwidth}{@{\extracolsep{\fill}}lccccc}
\toprule
$\nu$ & Rej. $\tilde Z_n^c$ & Mean & S.d.
& Rej. $Z_n^{\rm det}$ & Rej. $\lambda_{\max}$\\
\midrule
$5$ &0.137&0.182&1.351&1.000&1.000\\
$10$ &0.095&0.102&1.184&1.000&1.000\\
$20$ &0.077&0.073&1.093&1.000&0.973\\
\bottomrule
\end{tabular*}
\par
\vspace{0.05cm}
\begin{minipage}{\textwidth}
\footnotesize\emph{Notes:} $10{,}000$ replications at nominal level $0.05$.
Although $\operatorname{var}(X_t)=I_p$, coordinate independence fails.
\end{minipage}
\end{table}