Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,604 characters · 12 sections · 27 citation commands
Variance or Standard Deviation? Shell Geometry and Global-Scale Priors in High-Dimensional Shrinkage
A recurring default choice in hierarchical Bayes and high-dimensional shrinkage is how to place a prior on a common global Gaussian scale. Should the benchmark default be flat on the variance component or flat on the standard deviation? The two formulations are often treated as nearby modeling choices, but in high dimensions, they induce different scale geometries near the origin. The question is whether these different ways of weighting scales have systematically different frequentist risk consequences in high-dimensional settings, and how this should inform proper prior choices in practice.
We study these questions in the Gaussian normal means problem through a radial-power benchmark that nests the two canonical defaults: one corresponding to Stein's harmonic prior matches a flat prior on the variance, while the other matches a flat prior on the standard deviation. We analyze the generalized Bayes posterior mean across three asymptotic regimes for signal energy, meaning the squared Euclidean size of the unknown mean vector, as the dimension $d$ increases to infinity: weak-signal sequences with bounded energy, a critical regime in which signal energy is on the natural $\sqrt d$ fluctuation scale of Gaussian noise, and strong-signal sequences with energy of order $d$.
High dimension is not only of increasing relevance in applied work, but also provides a theoretical framework that makes the prior-scale question geometrically sharp. In high-dimensional Euclidean space, volume is concentrated in spherical shells: in polar coordinates the volume element contains the factor $r^{d-1}$. Here and below, a shell means a spherical shell or thin annulus: a set of points whose Euclidean norm lies in a small interval. Likewise, a standard Gaussian noise vector in $d$ dimensions typically lies in a thin annulus around radius $\sqrt d$; equivalently, its squared norm is $d$ plus random fluctuations of order $\sqrt d$. Such concentration of measure is one of the central lessons of high-dimensional geometry: mass can concentrate on sets that look negligible from low-dimensional intuition, and small changes in radial parametrization can become first-order under the relevant near-zero scaling Ball1997,Ledoux2001,Vershynin2018. Related geometric effects also underlie phase-transition phenomena in high-dimensional data analysis and signal processing DonohoTanner2009.
In the Gaussian normal-means experiment considered here, rotational invariance makes this geometry explicit. The benchmark posterior depends on the data through their squared Euclidean norm, and the asymptotic regimes are organized by the squared Euclidean size of the unknown mean. The two defaults (variance-flat versus SD-flat) differ exactly in how they weight radial shells near the zero-scale boundary. The high-dimensional limit therefore converts a seemingly local reparameterization choice (over variance or SD) into a quantitative risk comparison.
Our main results are as follows. In the weak regime, the SD-flat default improves on the variance-flat benchmark by one asymptotic risk unit. In the critical regime, we obtain an explicit limit for the centered risk and show that the SD-scale advantage holds at lower critical-regime signal strengths but can reverse as signal strength within that regime grows. In the strong regime, the risk expansion is universal across fixed radial exponents up to second order, and thus the two defaults yield asymptotically equivalent risk.
We then prove that these benchmark results transfer to a much wider range of settings, formalizing the heuristic that the key tradeoff in the choice of variance vs SD is driven by the high-dimensional scale geometry rather than by the particular radial-power benchmark. First, proper single global-scale hyperpriors inherit the same low-signal behavior through the near-zero exponent of their global SD density. Second, the same conclusion holds for bounded coordinate-multiplier normal mixtures. This class contains bounded global--local normal scale mixtures when the multipliers are positive continuous local scales, and Bernoulli spike-and-slab priors with fixed inclusion probability when the multipliers are binary indicators. The theorem separates the contribution of the common scale from additional mechanisms (heavy local-scale tails, dimension-dependent Dirichlet weight vectors, and model-size priors) that govern risk for common sparse priors outside the bounded-multiplier class.
The benchmark comparison belongs to the classical normal-means shrinkage literature because the benchmark rules are generalized Bayes estimators in the same radial family studied by Stein1956,Brown1971,Strawderman1971,MaruyamaTakemura2008,BrownZhao2012. Our emphasis is different, however. We do not study admissibility, minimaxity, or other fixed-dimensional decision-theoretic properties. The object here is frequentist risk under high dimension ($d\to\infty$) asymptotics. On the hierarchical-Bayes side, Gelman2006 emphasizes that default priors for variance components are not innocuous under reparameterization and argues for working on the standard-deviation scale, leading to the half-$t$ family as a default for hierarchical variance parameters. PolsonScott2012 single out the half-Cauchy as a useful prior for a common global scale, highlighting its heavy tails, substantial mass near zero, and convenient scale-mixture representation. PiironenVehtari2017Hyperprior,PiironenVehtari2017Sparsity make a complementary point for the horseshoe: the numerical scale of the global parameter can be calibrated from prior sparsity information, for example through the induced effective number of nonzero coefficients. The present analysis addresses a different component of the prior specification: it does not choose the numerical scale, but identifies the near-zero SD-scale exponent that controls high-dimensional low-signal risk. Thus the prior shape near zero and the numerical scale calibration are distinct components of the global-scale specification.
The comparison also connects to the modern shrinkage literature motivated by sparsity. Global--local priors such as the horseshoe and its variants were developed for nearly-black signals and related sparse normal means problems \citep*{CarvalhoPolsonScott2010,vanDerPasKleijnvanDerVaart2015,BhadraDattaPolsonWillard2016}. Spike-and-slab priors give a second major route to sparse shrinkage and variable selection: the standard specification introduces Bernoulli inclusion indicators, a model-size prior or inclusion-probability prior, and a slab distribution for included coefficients. Fully Bayes and empirical Bayes treatments of the inclusion probability provide multiplicity adjustment in high-dimensional model selection ScottBerger2010,CastilloSzabo2020. Rockova2018Continuous studies continuous spike-and-slab priors in sparse normal means, while RockovaGeorge2018SSL develop the spike-and-slab LASSO for high-dimensional regression with adaptive nonseparable penalties and multiplicity adjustment. In the regression setting with unknown noise level, MoranRockovaGeorge2019Variance show that variance-prior form can materially affect high-dimensional Bayesian variable selection. These papers are essential background, but they answer a different question from the one studied here. We deliberately average over directions and focus on the common global-scale component, because many applications feature many weak effects rather than a clean split between a few large coordinates and exact zeros. This viewpoint is compatible with sequence-space work outside exact sparsity JohnstoneSilverman2004, with hybrid dense-plus-sparse formulations ChernozhukovHansenLiao2017, and with recent econometric arguments that exact sparsity can be empirically fragile GiannoneLenzaPrimiceri2021,KolesarMullerRoelsgaard2025.
The practical message is therefore more transferable than a mere ranking of named priors. We identify the near-zero exponent of the common shrinkage scale as the quantity that governs low-signal isotropic risk when the remaining prior components preserve the same local near-zero asymptotic problem. For heavy-tailed global--local priors, Dirichlet--Laplace and R2-D2 shrinkage priors, and sparse model-selection priors, this exponent remains a useful specification of the common scale, while the full frequentist risk also depends on the additional component that defines the prior architecture.
The rest of the paper is organized as follows. Section (ref) introduces the benchmark model, the scale-mixture representation, the three asymptotic risk regimes, and our main results. Section (ref) transfers the benchmark results to proper priors, to common single global-scale hyperpriors, and to bounded coordinate-multiplier priors. Section (ref) gives numerical illustrations, and Section (ref) concludes. The appendices contain proofs and auxiliary calculations.
Consider the following normal means experiment \[ X_d\sim N_d(\theta_d,I_d), \] under a high-dimensional setting where $d\to\infty$. We estimate the normal mean vector $\theta_d\in\mathbb{R}^d$ by shrinkage rules. In the formulation used below, the shrinkage rules are generalized Bayes posterior means induced by appropriate priors on $\theta_d$, so the choice of prior determines the corresponding shrinkage rule.\footnote{We use “generalized Bayes” in the standard decision-theoretic sense: the formal posterior mean is computed from an improper prior when the marginal integral is finite, so the resulting estimator need not arise from a proper probability prior. Generalized Bayes shrinkage rules based on improper radial priors are classical in the normal-means problem; see, for example, Brown1971,Strawderman1971 and BergerStrawdermanTang2005.}
Our benchmark family of priors is the (improper) radial prior \[ \pi_{d,c}(\theta)\propto \left\lVert \theta \right\rVert^{-(d-c)}, \quad c>0, \] which nests the two commonly used global-scale geometries of interest. The family is a natural benchmark because rotational invariance removes directional modeling choices, and the polar-coordinate calculation below shows that the single exponent $c$ controls the radial-coordinate weighting. It therefore isolates the global-scale coordinate question without mixing it with sparsity, tail, or coordinate-selection mechanisms.
In polar coordinates $\theta=r\omega$, the prior $\pi_{d,c}(\theta)\propto\left\lVert \theta \right\rVert^{-(d-c)}$ induces marginal measure \[ r^{c-1}\,dr \] for the radial coordinate $r=\left\lVert \theta \right\rVert$. Hence $c=1$ is flat in $r$, corresponding to the SD-flat benchmark. Equivalently, it gives equal mass to spherical shells of equal radial thickness. By contrast, $c=2$ is flat in $r^2=\left\lVert \theta \right\rVert^2$, corresponding to the variance-flat benchmark.
Alternatively, we may also view the radial prior from the following perspective.
Proposition (ref) implies that $\pi_{d,c}$ is the marginal prior induced by the conjugate hierarchy $\theta\mid g\sim N_d(0,gI_d)$ with hyperprior $p(g)\propto g^{c/2-1}$ on the global variance component $g$.
Hence, the two defaults of interest are:
Both defaults have been well studied in previous literature. The variance-flat choice is algebraically conjugate and, through Proposition (ref), coincides with the harmonic radial prior that anchors much of the classical Stein-shrinkage literature Stein1956,Brown1971,Strawderman1971,MaruyamaTakemura2008,BrownZhao2012. The SD-flat choice is the default advocated in hierarchical variance-component modeling when the scale itself is the interpretable quantity; it is also the near-zero behavior of half-normal, half-$t$, and half-Cauchy scale priors with positive density at zero Gelman2006,PolsonScott2012.
Motivated by this radial scale geometry, we compare the two defaults by the frequentist squared-error risk of their generalized Bayes posterior means. For an estimator $\delta$, define \[ \mathcal{R}(\theta,\delta)=\mathbb{E}_\theta \left\lVert \delta(X)-\theta \right\rVert^2. \] For later use, write $T_d=\left\lVert X_d \right\rVert^2$ for the quadratic statistic, or equivalently the observed squared norm, and write $Y=g/(1+g)\in(0,1)$ for the variance-mixture shrinkage factor. We refer to $\left\lVert \theta_d \right\rVert$ as the signal norm and to $\left\lVert \theta_d \right\rVert^2$ as the signal energy. The generalized Bayes posterior mean under $\pi_{d,c}$ is the isotropic shrinkage rule \[ \delta_{d,c}(x)=s_{d,c}(\left\lVert x \right\rVert^2)x, \quad s_{d,c}(t)=\mathbb{E}[Y\mid T_d=t]. \]
The relevant asymptotic scale is the scale of the quadratic statistic. The same Gaussian sequence model underlies much of the high-dimensional shrinkage and sparse-normal-means literature, where regimes are often defined by sparsity or by signal energy rather than by fixed dimension alone.\footnote{For sparsity-based regimes in the normal-means problem, one typically writes $s_d=|\{j:\theta_{j,d}\ne0\}|$ and distinguishes nearly-black or sparse classes, such as $s_d=o(d)$, from dense classes with $s_d\asymp d$; see JohnstoneSilverman2004,CarvalhoPolsonScott2010,vanDerPasKleijnvanDerVaart2015. Detection formulations instead parameterize alternatives by an $\ell_2$ radius or total signal energy relative to Gaussian noise; see IngsterSuslina2000.}
We study three asymptotic regimes for the signal energy: as $d\to\infty$, \[
\] The weak regime has bounded total energy, the critical regime has signal energy on the null fluctuation scale of $T_d$, and the strong regime has energy proportional to dimension.
The three regimes are tied to fluctuations of the quadratic statistic rather than chosen ad hoc. When $\theta_d=0$, $T_d\sim\chi^2_d$, so $\mathbb{E} T_d=d$ and $\mathrm{Var}(T_d)=2d$, and $T_d-d$ is on the scale of $O(\sqrt d)$ with high probability LaurentMassart2000, consistent with the Gaussian-annulus viewpoint in high-dimensional probability Ledoux2001,Vershynin2018. A bounded signal energy (in the weak regime) is therefore below the intrinsic null fluctuation scale, $\left\lVert \theta_d \right\rVert^2\asymp \sqrt d$ is the critical scaling at which the signal shifts $T_d$ by the same order as noise, and $\left\lVert \theta_d \right\rVert^2\asymp d$ changes the normalized value $T_d/d$ at leading order.
The following two lemmas are useful for the frequentist-risk calculations. Lemma (ref) reduces the posterior shrinkage factor to a one-dimensional conditional distribution given the quadratic statistic, and Lemma (ref) rewrites squared-error risk in terms of that shrinkage factor.
We start with the weak regime, where the total signal energy remains bounded. In this regime the posterior mean is evaluated close to the origin, where the radial exponent $c$ directly controls the risk constant.
Thus the SD-flat benchmark has a one-unit asymptotic risk advantage over the variance-flat benchmark in the weak regime.
Theorem (ref) gives the first main message of the paper: in the weak regime, the SD-flat prior is preferable to the variance-flat one as measured by the asymptotic risk $\mathcal{R}$. This complements the argument of Gelman2006, who recommends working on the standard-deviation scale because the standard deviation is directly interpretable in the original model, the improper uniform density on that scale can be understood as a limit of proper half-$t$ priors, and common alternatives such as inverse-gamma priors or a flat prior on the variance can be sensitive near zero or miscalibrated toward larger variance components. It also clarifies how the half-Cauchy recommendation of PolsonScott2012 enters the present comparison: their case for the half-Cauchy is that it is a proper top-level scale prior with substantial mass near zero and heavy tails, giving good frequentist risk near the origin without severe compromises elsewhere. In our benchmark, its positive finite density at $\tau=0$ places its common-scale behavior in the same near-zero $c=1$ class as the SD-flat default.
Theorem (ref) is the bottom endpoint of the critical-regime analysis: a bounded signal energy lies below the $\sqrt d$ null-fluctuation scale of $T_d$, so it does not alter the limiting geometry of the statistic, and the result follows from the proof of Theorem (ref). The key proof heuristic is therefore deferred until after Theorem (ref). Here we only briefly explain why the limit under the weak regime is so clean and simple. In this regime, given the weakness of the signal, the only relevant feature of the prior is its near-zero exponent $c$: the centered risk converges to a constant that a Gaussian integration-by-parts identity (established for the critical regime) collapses to exactly $c$, with every other term vanishing. The total energy $\left\lVert \theta_d \right\rVert^2\to\nu$ then simply adds back, because the estimator barely shrinks this close to the origin, giving the simple additive form $\nu+c$ and the immediate one-unit gap between the two benchmarks.
We confirm this asymptotic risk gap numerically. Figure (ref) plots the excess risk $\mathcal{R}(\theta_d,\delta_{d,c})-\lambda_d$, with $\lambda_d=\left\lVert \theta_d \right\rVert^2$, for the weak-signal calibration at $d=2000$. The horizontal reference lines are the limiting constants $c=1$ and $c=2$ from Theorem (ref). The finite-$d$ curves are already close to these targets, and the vertical distance between the two curves is close to the one-unit asymptotic gap throughout the displayed bounded-energy range. Section (ref) gives the simulation design and Monte Carlo calculation.
The next regime asks how far this one-unit weak-signal advantage persists once the signal energy changes the quadratic statistic by the same order as its null fluctuation.
The intrinsic fluctuation scale of $T_d=\left\lVert X_d \right\rVert^2$ is $\sqrt d$, because \[ T_d=\left\lVert \varepsilon_d \right\rVert^2+2\theta_d^\top\varepsilon_d+\left\lVert \theta_d \right\rVert^2, \quad \varepsilon_d\sim N_d(0,I_d), \] and $\left\lVert \varepsilon_d \right\rVert^2-d$ is of order $\sqrt d$. The critical-regime scaling $\left\lVert \theta_d \right\rVert^2\asymp \sqrt d$ is therefore the first scaling at which signal energy changes $T_d$ by the same order as its stochastic fluctuations.
Through the exact risk identity of Lemma (ref), the proof reduces to the analysis of the limit of the single term $\mathbb{E}[T_d\,s_{d,c}(T_d)^2]$, which is delivered by the following two ingredients. First, at the critical-rigeme scaling the signal displaces the quadratic statistic by the same order as its stochastic fluctuation, so the centered statistic $Z_d=(T_d-d)/\sqrt d$ converges to $N(\beta,2)$: the variance is the $\chi^2$ noise and the mean is the signal energy, while the cross term between signal and noise washes out. Second, because the statistic sits at a distance of order $\sqrt d$ from the origin, the latent variance $g$ is pinned near the boundary $g=0$ on the scale $1/\sqrt d$, and a local boundary-Laplace limit shows the rescaled shrinkage factor converges to the function $h_c$ in which the prior exponent $c$ survives. Combining the two, $T_d\,s_{d,c}(T_d)^2$ behaves like $h_c(Z_\beta)^2$, and centering by $\left\lVert \theta_d \right\rVert^2$ leaves $L_c(\beta)$.
The critical regime is the transition regime in which the signal contribution to the quadratic statistic is comparable to the statistic's null fluctuation. The limiting risk comparison is therefore indexed by the single signal-strength parameter $\beta$: the SD-flat default is favored when the signal energy remains close to the origin ($\beta$ small), but the variance-flat harmonic benchmark can have smaller centered risk once $\beta$ is large enough. Thus the benchmark comparison is not a uniform dominance claim; it describes where each near-zero scale behavior is preferable in frequentist risk.
We numerically plot the finite-$d$ critical-regime risk gap $\mathcal{R}(\theta_d,\delta_{d,1})-\mathcal{R}(\theta_d,\delta_{d,2})$ together with the limiting curve $\Delta(\beta)$ in Figure (ref). Negative values favor the SD-scale benchmark and positive values favor the variance-flat benchmark. In the numerical calculation, the plotted asymptotic curve has a numerically identified zero at $\beta_{*}\approx 2.080$, where the centered risks of the two benchmarks are nearly tied. Section (ref) gives the detailed design.
The shape of $\Delta(\beta)$ is also informative. It starts at $\Delta(0)=-1$, so the weak-signal ordering survives for small signal energy in the critical regime. The upper-tail expansion $\Delta(\beta)=5\beta^{-2}+O(\beta^{-4})$ then forces the sign to become positive for sufficiently large signals in the critical regime.
In the strong regime, $\left\lVert \theta_d \right\rVert^2$ is of order $d$. The quadratic statistic $T_d=\left\lVert X_d \right\rVert^2$ concentrates around $d+\left\lVert \theta_d \right\rVert^2$, and the posterior for the global variance component moves away from the near-zero scaling that drives the critical-regime analysis.
In the following, Theorem (ref) gives the first-order limit $\mathcal{R}(\theta_d,\delta_{d,c})/d\to \rho/(1+\rho)$ whenever $\left\lVert \theta_d \right\rVert^2/d\to\rho>0$. Theorem (ref) refines this statement by identifying the $O(1)$ correction, and Corollary (ref) shows that any two fixed $c$ values are asymptotically indistinguishable at the $O(1)$ scale.
The main mechanism underlying Theorem (ref) is that strong signals wash out the prior. When $\left\lVert \theta_d \right\rVert^2\asymp d$, the normalized statistic $\tau_d=T_d/d$ obeys a law of large numbers and concentrates at the deterministic value $\tau_0=1+\rho$, so the stochastic fluctuations that drive the critical regime disappear. The exponent $c$ is therefore asymptotically irrelevant at this scale, which is the source of the universality. The random Bayes shrinkage factor accordingly collapses to the deterministic ridge coefficient $s_{d,c}(T_d)\to a(\tau_0)=1-1/\tau_0=\rho/(1+\rho)$, so $\delta_{d,c}$ behaves like constant shrinkage $a(\tau_0)X$.
The refinement in Theorem (ref) tracks two corrections that the first-order argument discards, and universality survives because they exactly cancel. We then have the following immediate corollary.
Figure (ref) illustrates this result numerically, with the details of the numerical exercise available in Section (ref).
Taken together, the three regimes give a phase diagram for the default comparison. At bounded energy, the SD-flat prior has a fixed one-unit advantage. Under the critical-regime scaling, that advantage holds only for sufficiently small signal energy and crosses over at a finite threshold on the signal-energy growth rate. At energy proportional to $d$, the posterior is no longer governed by the zero-scale boundary, and the two fixed-exponent benchmarks share the same asymptotic risk through the second-order risk expansion.
Overall, this three-regime transition gives a geometric risk map for the SD-flat versus variance-flat choice. The SD-flat prior is preferable in the weak regime and for sufficiently small critical-regime signal strength, while the variance-flat prior can have smaller centered risk for larger critical-regime signals. Once signal energy is proportional to $d$, however, the zero-scale geometry no longer affects risk through the second-order expansion.
Section (ref) compares the two benchmark geometries inside the radial power family. We now transfer those conclusions to proper priors, to common single global-scale hyperpriors, and to bounded coordinate-multiplier priors. We then use the same near-zero scale calculus to interpret common priors beyond a single global scale that fall outside the bounded-multiplier theorem. The key quantity is the near-zero exponent of the relevant standard-deviation scale.
Fix $c>0$. Let $\ell:[0,\infty)\to[0,\infty)$ be a nonincreasing function that is continuous at $0$ with $\ell(0)\in (0,\infty)$ and \[ \int_0^\infty g^{c/2-1}\ell(g)\,dg<\infty. \] Based on $\ell$, we may define the following proper hierarchical prior
and let $\delta_{d,c,\ell}$ be the corresponding posterior mean.\footnote{Equivalently, on the SD scale, if $\theta\mid\tau\sim N_d(0,\tau^2I_d)$ and $p_m(\tau)\propto \tau^{c-1}m(\tau)$, where $m$ is bounded, nonincreasing, continuous and positive at zero, and $\int_0^\infty \tau^{c-1}m(\tau)\,d\tau<\infty$, then the same transfer conclusions apply.}
The properizing factor $\ell$ does not affect the limiting low-signal risks because those limits are determined at the variance-scale endpoint $g=0$. In the weak and critical regimes, $T_d=d+O_{\mathbb{P}}(\sqrt d)$, so the posterior contribution to the shrinkage factor comes from $g=O(d^{-1/2})$. On that scale, $\ell(g)=\ell(0)+o(1)$, and the constant $\ell(0)$ cancels from the numerator and denominator of the posterior mean. Thus the local power $g^{c/2-1}$ fixes the local near-zero asymptotic problem, while monotonicity and integrability of $\ell$ only keep the properized prior from contributing mass at larger, nonlocal scales.
Proposition (ref) covers several proper defaults used in practice, including the half-$t$ family advocated by Gelman2006 and the half-Cauchy special case emphasized by PolsonScott2012: \[ p(\tau)\propto \mathbf 1_{[0,A]}(\tau) \quad\Longleftrightarrow\quad p(g)\propto g^{-1/2}\mathbf 1_{[0,A^2]}(g) \] (truncated flat SD), \[ p(g)\propto g^{-1/2}\Bigl(1+\frac{g}{\nu_0 s^2}\Bigr)^{-(\nu_0+1)/2} \] (half-$t$ on SD, including half-Cauchy when $\nu_0=1$), and \[ p(g)\propto \mathbf 1_{[0,A]}(g) \] (truncated flat variance). Thus the one-unit weak-signal advantage of flat on SD over flat on variance is not an artifact of the improper power family; it persists for the corresponding proper defaults as well.
The SD-scale formulation in Proposition (ref) turns the benchmark radial index $c$ into a classification rule on the standard-deviation scale. For a single global-scale hyperprior, the weak-signal and critical-regime limits depend only on the near-zero exponent of the global SD density.
This subsection classifies commonly used single global-scale hyperpriors according to whether their near-zero standard-deviation density has the SD-flat geometry ($c=1$), the variance-flat geometry ($c=2$), or a more general power $c$. The SD-scale formulation in Proposition (ref) gives immediate consequences for several priors that are widely used in high-dimensional shrinkage work. Consider the single global-scale Gaussian hierarchy \[ \theta\mid \tau\sim N_d(0,\tau^2 I_d), \] and let $\delta_{d,p}$ denote the posterior mean under hyperprior $p(\tau)$.
\paragraph{Flat-SD, half-\texorpdfstring{$t$}{t}, half-Cauchy, half-normal, exponential, and truncated-flat priors.} If $p(\tau)$ is half-$t_{\nu_0}(s)$, half-Cauchy$(s)$, half-normal$(s)$, exponential$(b)$, or truncated flat on $[0,A]$, then the density on the SD scale is positive and finite at zero, so Theorems (ref) and (ref) continue to apply to $\delta_{d,p}$ with $c=1$.
\paragraph{\texorpdfstring{$R^2$}{R2}-based priors.} Suppose $W=w_0\tau^2$ for some fixed $w_0>0$ and \[ R^2=\frac{W}{1+W}\sim \mathrm{Beta}(a,b), \quad a,b>0. \] Then \[ p(\tau)=\frac{2w_0^a}{\mathrm{B}(a,b)}\tau^{2a-1}(1+w_0\tau^2)^{-(a+b)}, \quad \tau>0, \] and therefore the near-zero SD-scale exponent is $c=2a$, so Theorems (ref) and (ref) continue to apply to $\delta_{d,p}$ with $c=2a$. In particular, $a=\tfrac12$ gives the SD-scale class, whereas $a=1$ gives the variance-flat class; for any $a_1<a_2$, Theorem (ref) gives the weak-signal gap $2(a_1-a_2)$.
\paragraph{Gamma priors on variance.} If $g=\tau^2\sim \mathrm{Gamma}(a,b)$, then \[ p(\tau)=\frac{2b^a}{\Gamma(a)}\tau^{2a-1}e^{-b\tau^2}, \quad \tau>0, \] the near-zero SD-scale exponent is again $c=2a$, so gamma variance priors fall under Theorems (ref) and (ref) with $c=2a$. In particular, exponential priors on $g$ ($a=1$) belong to the same variance-flat class as the improper flat prior on $g$.
Therefore, by Proposition (ref), the Section (ref) results translate directly: SD-scale priors inherit the $c=1$ weak-signal and critical-regime limits, variance-scale priors inherit the $c=2$ limits, and the more general beta or gamma powers use their corresponding exponent $c$.
\paragraph{Scale calibration.} This exponent classification concerns the local shape of $p(\tau)$ at zero; it does not choose the numerical scale. For the horseshoe in regression, PiironenVehtari2017Hyperprior,PiironenVehtari2017Sparsity propose calibrating that scale through prior information about sparsity. In their Gaussian-regression normalization, if $p_0$ is the prior guess for the number of relevant predictors among $D$ predictors, then the natural global-scale level is \[ \tau_0=\frac{p_0}{D-p_0}\frac{\sigma}{\sqrt n}. \] This separates two components of the prior specification: the weak-regime comparison favors the $c=1$ SD-scale shape over the $c=2$ variance-flat benchmark, while the width of that prior can be chosen from substantive sparsity information rather than from a universal unit-scale convention.
Many high-dimensional priors add coefficient-specific latent structure on top of a common scale. This subsection studies a precise transfer question: after such a component of the hierarchy is added, which part of the Section (ref) comparison still comes from the common scale, and which part can be changed by the additional component?
We separate the problem into two cases. For bounded coordinate-multiplier priors, an exact risk-transfer theorem shows that the same near-zero SD-scale exponent $c$ determines the weak-signal and critical-regime limits, just as in the single-global benchmark. For common sparse or heavy-tailed priors outside this bounded class, the exponent still classifies the common global-scale component, but it does not by itself yield a complete risk theorem: local-scale tails, model-size priors, or dimension-dependent allocation priors can also enter the leading risk calculation.
We use the standard names for the resulting architectures. A global--local prior is a continuous shrinkage prior with a common scale and positive local scales, typically \[ \theta_j\mid \tau,\lambda_j\sim N(0,\tau^2\lambda_j^2), \quad \lambda_j>0. \] A spike-and-slab prior is a model-selection prior with inclusion indicators and, in the point-mass case, exact prior mass on coordinate subspaces. Dirichlet--Laplace and R2-D2 priors BhattacharyaPatiPillaiDunson2015DL,ZhangNaughtonBondellReich2022R2D2 add a different type of structure: a common total scale or model-fit parameter is distributed across coordinates through a Dirichlet weight vector. These architectures all contain a common scale, but the additional component has a different mathematical role in each case.
\paragraph{Exact transfer for bounded coordinate multipliers.}
The formal transfer theorem uses the following coordinate-multiplier envelope. It keeps a single common variance scale $g=\tau^2$, but allows each coordinate to receive a bounded multiplier $A_j$.
Consider the coordinate-multiplier hierarchy
with common variance prior \[ \pi(g)\propto g^{c/2-1}\ell(g), \quad g=\tau^2. \] When $A_j=\lambda_j^2$ with positive continuous local scales, (ref) is a bounded global--local normal scale mixture. When $A_j=Z_j\in\{0,1\}$ is Bernoulli with fixed success probability $q$, it gives the fixed-$q$ point-mass spike-and-slab prior \[ Z_j\overset{\mathrm{iid}}{\sim}\operatorname{Bernoulli}(q), \quad \theta_j\mid g,Z_j\sim (1-Z_j)\delta_0+Z_jN(0,g). \] Thus the coordinate-multiplier model is a common mathematical representation; the subclasses retain their usual statistical interpretations.
The posterior mean under (ref) is generally non-radial, because different coordinates can be shrunk by different posterior multiplier information. Proposition (ref) shows that this non-radiality does not change the low-signal local near-zero limit as long as the multipliers are bounded and have positive mean. The analytic posterior-mean identity and the coordinate marginal used in the proof are given in Appendix (ref).
The proof also explains why the multiplier distribution does not affect the limit. The multiplier distribution enters the local near-zero likelihood through its mean $\mu=\mathbb{E} A$. After the change of variables $u=\mu\sqrt d\,g$, this factor disappears from the limiting likelihood, leaving the same local near-zero asymptotic problem as in Section (ref). Hence, within the bounded coordinate-multiplier class, an SD-scale common prior retains the one-unit weak-signal advantage over its variance-scale analogue.
\paragraph{What does not transfer as an exact theorem.}
Proposition (ref) is not a universal theorem for every prior that contains a global scale. Outside the bounded-multiplier envelope, the common-scale exponent still says whether the common scale is SD-like or variance-like near zero, but leading risk can also be affected by other prior components: local-scale tails, model-size priors, or dimension-dependent allocation priors. Hence the constants $L_c(\beta)$ should not be read as applying automatically to horseshoe-type, sparse spike-and-slab, Dirichlet--Laplace, or R2-D2 priors. Appendix (ref) gives the corresponding calculations. The simulations in Section (ref) provide qualitative support outside the theorem: in the many-weak designs considered there, the $c=1$ common-scale law has lower weak-signal excess risk than its $c=2$ analogue within both the global--local and spike-and-slab architectures.
This section gives finite-dimensional illustrations of the risk regimes in Section (ref), the transfer to representative proper single global-scale hyperpriors, and the weak-signal comparison within two representative architectures beyond a single global scale. For the benchmark family and for single global-scale hyperpriors, the posterior mean is radial, so the displayed risks are exact finite-$d$ evaluations up to numerical integration error. The final numerical comparison is different: the posterior mean is non-radial, and the reported risks combine Monte Carlo evaluation with deterministic quadrature over the common global or slab scale.
To represent many weak effects, the benchmark and single-global-scale calculations use the equal-effects vector \[ \theta_d^{\mathrm{eq}}(\lambda_d) = \sqrt{\lambda_d/d}\,\mathbf 1_d, \quad \lambda_d=\left\lVert \theta_d \right\rVert^2. \] For the radial rules studied in the first three numerical exercises, the risk depends on $\theta_d$ only through $\lambda_d$, so this representation is interpretive rather than restrictive. In the final comparison beyond a single global scale, the displayed many-weak vector is part of the simulation design.
\paragraph{Default priors.}
Figures (ref) and (ref) give the benchmark checks next to the corresponding theoretical statements. Figure (ref) verifies Theorem (ref): already at $d=2000$, the excess risk $\mathcal{R}(\theta_d,\delta_{d,c})-\lambda_d$ is close to its limiting constant $c$. In Table (ref), for example, the $c=1$ and $c=2$ risks are $(1.981,2.971)$ at $\lambda=1$ and $(4.922,5.882)$ at $\lambda=4$, so the SD-scale benchmark retains an approximately one-unit advantage throughout the displayed weak-signal range. Figure (ref) visualizes Theorem (ref) and Corollary (ref): the finite-$d$ critical-regime gap $\mathcal{R}(\theta_d,\delta_{d,1})-\mathcal{R}(\theta_d,\delta_{d,2})$ tracks the limit $\Delta(\beta)$ closely, is zero at $\beta=\beta_{*}\approx 2.080$, and reverses sign by $\beta=3$ ($-6.952$ versus $-7.229$).
Figure (ref) verifies Theorem (ref). At both $d=200$ and $d=1000$, the scaled risks $\mathcal{R}(\theta_d,\delta_{d,c})/d$ for $c=1$ and $c=2$ are nearly indistinguishable, while the second-order approximation materially sharpens the first-order limit when $d$ is moderate. Table (ref) records the same collapse at $d=200$: the two scaled risks agree to the displayed precision at $\rho=1$ and $\rho=4$, and differ by only $0.002$ at $\rho=0.25$. The calculations illustrate Corollary (ref): once the signal energy is of order $d$, the distinction between flat on variance and flat on SD disappears at the $O(1)$ scale.
Table (ref) gives selected finite-$d$ values from the weak-signal, critical-regime, and strong-signal calibrations. At the null, the risks are exactly $1.000$ and $2.000$ for $c=1$ and $c=2$. At the displayed critical-regime calibration $\beta_{*}\approx2.080$ the two centered risks are nearly tied, while at $\beta=3$ the ordering has reversed. In the strong-signal rows, the scaled risks for $c=1$ and $c=2$ coincide to the displayed precision except for the smallest $\rho$ value.
\paragraph{Proper hyperpriors.}
Figure (ref) and Table (ref) verify the transfer results from Section (ref), namely Proposition (ref) and the single-global-scale dictionary in Section (ref). The $c=1$ class---half-Cauchy, half-normal, truncated flat on $\tau$, and Beta-$R^2$ with $a=\tfrac12$---clusters around the SD-flat benchmark: the displayed null risks range from $0.965$ to $1.000$, and the critical-regime values at $\beta=1$ range from $-0.302$ to $-0.272$. The $c=2$ class---Gamma on $g$ with $a=1$, truncated flat on $g$, and Beta-$R^2$ with $a=1$---clusters around the variance-flat benchmark, with null risks from $1.929$ to $2.000$ and $\beta=1$ values from $0.184$ to $0.250$. Finite-$d$ differences within each class remain, as expected for proper priors, but the classwise ordering is stable and becomes especially sharp away from the critical-regime crossover.
\paragraph{Beyond a single global scale.}
Section (ref) separates exact transfer results from architecture-specific risk behavior. The numerical exercise here focuses on a narrower comparison: within a fixed architecture, only the common-scale law is changed from the $c=1$ SD-scale class to the $c=2$ variance-scale class.
We consider the many-weak design \[ m_d=\lfloor d^{3/4}\rfloor, \quad \theta_{d,j}^{\mathrm{mw}}(\lambda)=(-1)^{j-1}\sqrt{\lambda/m_d}\,\mathbf 1\{j\le m_d\}, \quad j=1,\ldots,d, \] so that $\left\lVert \theta_d^{\mathrm{mw}}(\lambda) \right\rVert^2=\lambda$, the number of active coordinates diverges, and each active coordinate remains weak. The simulations use $d=500$ and $\lambda\in\{0,1,2,4\}$.
The first architecture is a horseshoe-type global--local normal scale mixture with half-Cauchy local scales, \[ \theta_j\mid \lambda_j,\tau\sim N(0,\tau^2\lambda_j^2), \quad \lambda_j\sim \mathrm{half\mbox{-}Cauchy}. \] Its $c=1$ version assigns a half-Cauchy prior to the common SD scale $\tau$, whereas its $c=2$ version assigns an exponential prior to the common variance $g=\tau^2$. The second architecture is the standard Bernoulli spike-and-slab model, \[ \theta_j\mid z_j,\gamma\sim (1-z_j)\delta_0+z_jN(0,\gamma^2), \quad z_j\overset{\mathrm{iid}}{\sim}\operatorname{Bernoulli}(q_d). \] For the simulation we set $q_d=m_d/d$, so that the prior expected model size $dq_d$ equals the number of active coordinates in the design. This oracle calibration is used only to isolate the effect of the common slab-scale law; in data analysis $q$ is typically assigned a hyperprior or estimated by marginal likelihood. The $c=1$ version places a half-Cauchy prior on the common slab SD $\gamma$, and the $c=2$ version replaces that component by an exponential prior on the slab variance $\gamma^2$.
Figure (ref) and Table (ref) show that the within-architecture comparison follows the same direction in both examples. Replacing the $c=1$ common-scale law by the corresponding $c=2$ law increases weak-signal excess risk at every displayed value of $\lambda$. For the global--local architecture, the increase ranges from $0.86$ to $0.94$; for the spike-and-slab architecture, it ranges from $0.64$ to $0.72$. At $\lambda=4$, for instance, the excess risks are $0.97$ and $1.83$ in the global--local architecture, and $0.62$ and $1.26$ in the spike-and-slab architecture. The risk curves remain architecture-specific. Nevertheless, after the architecture is fixed, the common-scale exponent continues to organize the weak-signal comparison in the same direction as in the single-global benchmark.
Overall, the numerical evidence reinforces the main conclusions. The weak-signal one-unit advantage is already visible at moderate dimension, the critical regime exhibits the predicted crossover, strong-signal sequences erase the benchmark distinction up to second order, representative proper single global-scale hyperpriors inherit the same low-signal behavior through their near-zero SD-scale exponent, and the two representative architectures beyond a single global scale display the same within-architecture $c=1$ versus $c=2$ ordering in many-weak signals.
This paper isolates the near-zero geometry of a global Gaussian shrinkage scale as the relevant object in high-dimensional low-signal risk. For the two canonical defaults, flat on the standard deviation improves on flat on the variance by one asymptotic risk unit in the weak regime, the critical regime exhibits a signal-strength crossover, and strong-signal sequences erase the distinction up to second order.
The transfer results show how this benchmark comparison propagates to the priors used in practice. For a single global-scale hyperprior, the near-zero exponent of the global SD density is the only statistic that matters for weak-signal and critical-regime isotropic risk. For bounded coordinate-multiplier normal mixtures, the same limit survives even though the posterior mean is non-radial. This bounded class contains bounded global--local normal scale mixtures and Bernoulli spike-and-slab priors with fixed inclusion probability as distinct subclasses. For commonly used priors outside this class, the classification identifies the common global-scale component and the additional component that enters the asymptotic problem: finite-mean local mixtures have the same first-order small-scale geometry, horseshoe-type local-scale distributions create a heavy-tail near-zero behavior in which the effective small-scale parameter is $\tau$, and sparse spike-and-slab priors are governed partly by the model-size prior.
Several questions remain open. A sharper characterization of the critical-regime crossover set, including uniqueness of the zero of $\Delta(\beta)$, would refine the benchmark phase diagram. On the prior side, the next theoretical steps are to extend the bounded coordinate-multiplier theorem to unbounded finite-moment multipliers, to develop the heavy-tail near-zero theory needed for horseshoe-type local-scale distributions, and to analyze triangular model-size priors, such as sparse spike-and-slab, and dimension-dependent Dirichlet-weight priors, such as Dirichlet--Laplace and R2-D2. These extensions would separate the contribution of the common global scale from the additional local, Dirichlet-weight, or model-size prior that determines risk outside the bounded class.
Disclosure of Generative AI Usage: We used generative AI tools (ChatGPT 5.5, Claude Opus 4.8, Gemini 3.5, and Refine.ink) for assistance with coding, drafting/editing, and critical reviews. We assume full responsibility for the final manuscript.