EconBase
← Back to paper

Gaussian Boundary Inference in a Hypergeometric Heavy-Tailed Family

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

117,423 characters

Gaussian Boundary Inference in a Hypergeometric Heavy-Tailed Family



\title{Gaussian Boundary Inference in a Hypergeometric Heavy-Tailed Family}

\author{
    Steve Lawford\thanks{ENAC (University of Toulouse), 7 avenue Edouard Belin, BP 54005, 31055, Toulouse, Cedex 4, France (email: [email removed]).
    }
}

\date{}

\maketitle

\begin{abstract}
\noindent This paper develops inference for a Gaussian-nested hypergeometric family of distribution functions. The family
\[
    G_c(z) = \frac12 + z\,\frac{\Gamma(c-1/2)}{2\sqrt2\,\Gamma(c)}\,{}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right),\quad c\ge\frac32,
\]
contains the standard normal distribution at the boundary
$c=3/2$. Away from the boundary, the density has algebraic
tail behaviour $g_c(z)\sim(c-3/2)|z|^{-3}$, so the parameter $c$ indexes a directed heavy-tailed deformation of the Gaussian law. The distribution also admits an equivalent beta-precision normal scale-mixture representation, in which the Gaussian boundary corresponds to degenerate unit precision. I establish the admissibility of the hypergeometric family, derive minimum-distance estimators, and obtain both regular interior asymptotics and nonstandard boundary asymptotics under the normal null. The standardised fit-improvement statistic converges to the mixture distribution $\tfrac12\delta_0+\tfrac12\chi_1^2$. I extend the theory to plug-in location-scale procedures, including a robust
median/IQR version. Simulations document accurate null size
and directed power against heavy-tailed alternatives. An
application to daily S\&P~500 returns, both unconditionally
and after GARCH(1,1) filtering, illustrates the empirical
implications for tail fitting and risk quantiles, and documents that GARCH filtering substantially reduces but does not eliminate the symmetric heavy-tailed departure detected by the test.
\end{abstract}

\noindent\textbf{Keywords:}
hypergeometric functions; normality testing; heavy tails; boundary inference; minimum-distance estimation; empirical process; financial returns.

\medskip

\noindent\textbf{JEL classification:} C12; C14; C46; C58.

\newpage

\section{Introduction}\label{sec:introduction}

Normality is one of the most useful reference assumptions in econometrics, statistics, and financial modelling. It provides tractable likelihoods, familiar asymptotic approximations, and a benchmark against which departures in skewness, kurtosis, and tail behaviour are measured. Yet in many empirical settings, and especially in economics and finance, the Gaussian model is best understood as a local approximation rather than a literal description of the data. Asset returns, forecast errors, regression residuals, and macro-financial shocks frequently display more mass in the shoulders and tails than the Gaussian law permits \citep{mandelbrot63, fama65, cont01, gabaix09}. Detecting such departures is important, but detection alone is not enough:  rejection of normality by an omnibus test \citep{anderson_darling52, jarque_bera87, bai_ng05, doornik_hansen08} says that something is wrong with the Gaussian approximation; it does not say what distributional direction has been found, nor what model should replace the Gaussian benchmark.

This paper develops a new directed test of normality based on a Gaussian-nested hypergeometric family of distribution functions. The family is
\[
    G_c(z) = \frac12 + z\,\frac{\Gamma(c-1/2)}{2\sqrt2\,\Gamma(c)}\,
    {}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right),
    \quad c\ge\frac32,
\]
where ${}_1F_1$ is Kummer's confluent hypergeometric function \citep{abramowitz_stegun72, olver_etal10}. The standard normal distribution is obtained exactly at the finite boundary value
\[
    G_{3/2} = \Phi.
\]
For every $c > 3/2$, the density $g_c = G_c'$ is symmetric and has algebraic tails,
\[
    g_c(z) \sim \left(c-\frac32\right)|z|^{-3}, \quad |z|\to\infty.
\]
Thus the Gaussian law appears as a critical boundary case: at $c = 3/2$, the algebraic tail component vanishes and the density reduces to the standard normal; for $c > 3/2$, the distribution has finite mean but infinite variance. The parameter $c$ therefore indexes a specific heavy-tailed deformation of the Gaussian law.

The construction is not an arbitrary use of a special function. It proceeds from the classical identity
\[
    \Phi(z) = \frac12 + \frac{z}{\sqrt{2\pi}}\,
    {}_1F_1\!\left(\frac12;\frac32;-\frac{z^2}{2}\right).
\]
A natural hypergeometric ansatz would allow both the numerator and denominator parameters of ${}_1F_1$ to vary. However, the endpoint restrictions required of a distribution function force the numerator parameter to be $1/2$. The remaining denominator parameter $c$ is then free, and monotonicity holds on the admissible region $c \geq 3/2$. The result is a genuine one-parameter family of distribution functions, not merely a formal perturbation of the Gaussian density.

This feature distinguishes the present approach from many standard flexible approximations to normality. Gram-Charlier, Edgeworth, and seminonparametric expansions \citep{gallant_nychka87, jondeau_rockinger01} can be useful for representing deviations from a benchmark density, but they often require truncation choices, sign restrictions, or additional corrections to ensure positivity and integrability. By contrast, $G_c$ is an admissible distribution function by construction. It also differs from standard heavy-tailed parametric families. In the Student $t$ family, normality is approached as the degrees of freedom tend to infinity, and the tail index varies with the shape parameter. In the hypergeometric family, normality occurs at the finite boundary $c=3/2$, and all non-Gaussian members share the same  tail index. The parameter $c$ controls the scale of the algebraic tail rather than its index. The family is therefore best understood as a Gaussian-boundary tail-deformation model: a one-parameter perturbation of the normal law in a single, well-defined distributional direction. A companion paper in progress frees the tail index, yielding a two-parameter family that nests the present case.

The hypergeometric family admits a useful probabilistic representation. For $c>3/2$,
\[
    Z \sim G_c \quad \Longleftrightarrow \quad
    Z = \frac{\varepsilon}{\sqrt{\Lambda}},
    \quad \varepsilon\sim \operatorname{N}(0,1), \quad
    \Lambda\sim \operatorname{Beta}\left(1,c-\frac32\right),
\]
with $\varepsilon$ and $\Lambda$ independent. Thus $G_c$ is a beta-precision normal scale mixture. The Gaussian boundary $c=3/2$ corresponds to the degenerate case $\Lambda\equiv1$, while $c>3/2$ introduces beta-distributed precision heterogeneity. Small realisations of $\Lambda$ generate high-variance conditional Gaussian observations and hence the algebraic tails of $G_c$. This representation connects the hypergeometric family to the normal variance-mixture literature \citep{andrews_mallows74, barndorff-nielsen_etal82}, but the inference problem here is not general estimation of a normal mixture: it is the boundary test of Gaussian degeneracy within a specific hypergeometric path.

The paper is also motivated by a broader line of work on hypergeometric functions as econometric modelling tools. Hypergeometric functions have long appeared in econometrics and statistics, especially in exact finite-sample distribution theory, unit-root asymptotics, ratios of quadratic forms, and related problems where nonstandard distributions can be written in terms of special functions \citep{phillips83, abadir93, abadir95, hillier01, forchini02}. In that literature, hypergeometric functions usually arise as analytic representations of distributions or moments arising from a given model or inferential problem. A separate and more constructive line of work, initiated by \citet{abadir99}, proposed estimating hypergeometric parameters directly as a flexible but structured way to model unknown nonlinear functions, with applications to omitted nonlinearity, constructive nonlinear modelling, and density estimation in finance. That programme was developed further in \citet{lawford01} and \citet{abadir_rockinger03}.

The appeal of that programme is clear. Generalised hypergeometric functions contain many familiar functional forms as special or limiting cases; they possess analytic identities, transformations, integral representations, and asymptotic expansions; and their parameters can alter functional shape parsimoniously. They therefore offer a middle ground between rigid parametric specifications and fully nonparametric methods: flexible enough for empirical realism, yet structured enough for interpretation and computation.

However, the generality of that programme also made it difficult. Estimating unrestricted ${}_pF_q$ functions raises questions about admissibility, identification, numerical evaluation, analytic continuation, and asymptotic theory, and the resulting procedures can be computationally delicate and theoretically hard to characterise. The present paper takes a narrower and more disciplined route. Rather than using a general ${}_pF_q$ as a flexible approximator, it isolates one specific confluent hypergeometric object that is analytically tractable, is an admissible distribution function, nests the Gaussian law exactly, and yields a well-defined boundary testing problem. In this sense, the paper contributes a full theoretical treatment of one canonical hypergeometric distributional path, as a foundation for broader hypergeometric modelling.

The econometric problem is to estimate $c$ and test the Gaussian boundary. Let $Z_1,\ldots,Z_n$ be i.i.d. with empirical distribution function $\widehat F_n$. For a fixed finite measure $\nu$, define the minimum-distance criterion
\[
    Q_n(c) = \int_\mathbb{R} \{\widehat F_n(z)-G_c(z)\}^2\,d\nu(z),
    \quad c\in\mathcal C=[3/2,c_U],
\]
and let
\[
    \widehat c_n\in\arg\min_{c\in\mathcal C}Q_n(c).
\]
The Gaussian null corresponds to the boundary value $c_0=3/2$. The test statistic is the improvement in fit from estimating $c$ rather than imposing $c=c_0$:
\[
    T_n = n\,\{Q_n(c_0)-Q_n(\widehat c_n)\}.
\]
After standardisation by null constants determined by $\nu$, $\Phi$, and the hypergeometric score direction $\dot G_{c_0}$, the statistic has the nonstandard boundary limit
\[
    \widetilde T_n \Rightarrow \frac12\delta_0 + \frac12\chi_1^2,
\]
where $\delta_0$ denotes a point mass at zero. The limiting distribution reflects the fact that the Gaussian null lies on the boundary of the admissible parameter space: negative fluctuations of the unconstrained estimator would imply $c<3/2$, which is inadmissible, and are therefore projected to the boundary. The limit is consistent with the general boundary theory of \citet{chernoff54}, \citet{self_liang87} and \citet{andrews01}; the contribution is to derive it rigorously for a hypergeometric minimum-distance criterion in which the Gaussian null arises as an exact finite boundary rather than an asymptotic limit.

The paper develops the corresponding minimum-distance theory. First, I prove consistency of $\widehat{c}_n$, under correct specification and under misspecification, where the limit is the pseudo-true hypergeometric projection of the underlying distribution. Second, I derive asymptotic normality when the pseudo-true value is interior. Third, I derive the boundary distribution under the Gaussian null and the associated fit-improvement statistic. Fourth, I establish directed consistency: the test has power tending to one against fixed alternatives for which the best-fitting hypergeometric distribution improves strictly on the Gaussian boundary. Finally, I derive local asymptotic power under $\sqrt{n}$-local alternatives. The local power index is
\[
    \delta_\nu(\eta) =
    \frac{2\int_\mathbb{R} \eta(z)\,\dot G_{c_0}(z)\,d\nu(z)}{\sqrt{\Omega_0(\nu)}},
\]
where $\eta$ is the local departure from normality and $\Omega_0(\nu)$ is the null asymptotic variance of the score functional. This expression makes explicit that the hypergeometric test is directed rather than omnibus: it is locally sensitive to departures that project positively onto the hypergeometric score direction $\dot G_{c_0}$.

For empirical use, I develop plug-in location-scale versions of the test. The first uses the sample mean and standard deviation: the null theory follows by replacing the Brownian bridge covariance kernel with the projected empirical-process kernel induced by estimating location and scale. The second, and operationally more useful, version uses the sample median and normal-calibrated IQR scale: this robust plug-in statistic is motivated by the fact that the hypergeometric alternatives have infinite variance for $c>3/2$, so standard-deviation standardisation can be unstable under the alternatives of greatest interest. Median/IQR standardisation preserves tail-shape information more effectively and has the same boundary-mixture null limit, with the plug-in variance constant replaced by the corresponding robust projected-process variance.

The choice of $\nu$ is part of the definition of the minimum-distance criterion. The paper focuses on a pre-specified tail-sensitive measure
\[
    d\nu_{1/2}(z) \propto [\Phi(z)\{1-\Phi(z)\}]^{-1/2}\,d\Phi(z).
\]
This choice increases the relative importance of the tails while keeping $\nu$ finite and symmetric. The local power calculation shows that one could, in principle, choose $\nu$ to maximise power against a specified direction, but the optimal weighting would depend on the unknown alternative and could place excessive weight in regions where the empirical distribution is noisy. Instead, the paper fixes a transparent tail-weighted criterion and studies its finite-sample behaviour.

The contribution of this paper is not a more powerful normality test: omnibus procedures such as Jarque--Bera, Anderson--Darling, and Shapiro--Wilk often have higher power against classical non-Gaussian alternatives. The contribution is an inferential architecture: admissibility, estimation, boundary asymptotics, plug-in implementation, and finite-sample calibration, for a specific hypergeometric distributional model in which the Gaussian law arises as an exact finite boundary. The normality testing application is the natural vehicle for demonstrating that architecture; the fitted parameter $\widehat c$ and its tail-risk implications are what the framework delivers. The simulations support the theory. Under the Gaussian null, the fixed-scale statistic has accurate finite-sample size from $n=25$ onwards, validating the boundary mixture approximation in small samples. The plug-in statistics are well calibrated at moderate $n$; the robust median/IQR plug-in is mildly liberal in small samples, with the distortion decreasing with $n$. The mean/SD plug-in has correct null size but essentially zero power against hypergeometric alternatives with $c>3/2$, because sample-standard-deviation standardisation is not regular under infinite-variance alternatives and destroys the tail signal; this validates the recommendation of the robust version. Joint estimation of the shape parameter together with unknown location and scale is severely size-distorted: the exact-profiled statistic has size that increases markedly with $n$ over the simulated range, reaching $31\%$ at nominal $10\%$ when $n=1600$, confirming the necessity of the plug-in approach.

Power simulations confirm that the fixed-scale statistic is highly powerful against hypergeometric alternatives in the designed direction, and that the robust plug-in has good power against hypergeometric and symmetric heavy-tailed alternatives, including Student $t$ and Laplace distributions, though with a cost in power relative to the fixed-scale test against hypergeometric alternatives, due to the projection loss from median/IQR standardisation. The fixed-scale test has essentially zero power against non-hypergeometric alternatives, including the Student $t_3$, reflecting the directed nature of the test: despite being heavy-tailed, the $t_3$ departure has negligible positive projection onto the hypergeometric score direction; its survival-tail index is 3, rather than the index 2 shared by the hypergeometric alternatives. As expected from the local power theory, both hypergeometric procedures have negligible power against skewed alternatives; this is a consequence of the symmetry of the hypergeometric family and not a defect. Omnibus tests such as Jarque--Bera, Anderson--Darling, and Shapiro--Wilk often have higher power against classical non-Gaussian alternatives, but they provide neither a fitted heavy-tail direction nor an interpretable hypergeometric parameter estimate with direct tail-risk implications.

The empirical illustration applies the robust plug-in hypergeometric statistic to daily S\&P~500 returns, both unconditionally and after GARCH(1,1) filtering. In both exercises the Gaussian boundary is rejected with $\widehat c>3/2$; GARCH filtering substantially reduces the estimated
tail departure but does not eliminate it. For raw returns the Student $t$ and Laplace fit better overall, while for GARCH-filtered residuals the hypergeometric ranks second by log score. In both cases the fitted $\widehat c$ implies materially more extreme far-tail quantiles than the
Student $t$, despite the latter's better overall fit. The hypergeometric model is not proposed as a universal replacement for conventional heavy-tailed distributions but as a constructive diagnostic: it asks whether the departure from normality is consistent with a particular Gaussian-boundary algebraic-tail deformation, and what tail-risk implications that deformation carries.

The paper contributes to several literatures. It contributes to normality and goodness-of-fit testing \citep{shapiro_wilk65, jarque_bera87, bai_ng05, doornik_hansen08}: unlike omnibus tests, rejection of the hypergeometric statistic has a direction and an associated fitted distribution. It connects to empirical-process and minimum-distance inference \citep{shorack_wellner86, bai03}, where the weak convergence arguments used here are standard tools applied in a nonstandard boundary setting. It relates to boundary inference and one-sided testing, where limiting distributions involve projections and mixtures of chi-squared laws \citep{chernoff54, self_liang87, andrews01}. It builds on flexible density modelling and seminonparametric methods \citep{gallant_nychka87, gallant_tauchen89, jondeau_rockinger01}, from which it differs by working with an admissible distribution function rather than a truncated expansion. It contributes to the use of hypergeometric functions in econometrics, including exact distribution theory, unit-root asymptotics, and direct estimation of hypergeometric parameters \citep{phillips83, abadir93, abadir95, abadir99, lawford01, abadir_rockinger03}, and to hypergeometric distribution modelling \citep{gordy98}. Finally, it connects to heavy-tailed financial return modelling \citep{mandelbrot63, fama65, bollerslev87, nelson91, mcneil_etal15}, contributing a directed test and fitted distribution rather than a model selected by information criteria.

The paper provides a full theoretical treatment of a hypergeometric distribution family with an exact Gaussian boundary, minimum-distance estimation, nonstandard boundary inference, and a robust plug-in implementation with documented finite-sample behaviour. The resulting test complements existing normality tests rather than replacing them. It is not designed to detect all departures from normality, but to detect, estimate, and interpret a specific heavy-tailed deformation of the Gaussian law; and to do so with a fitted parameter that carries direct tail-risk meaning. More broadly, the paper suggests that hypergeometric econometrics need not begin with a fully general ${}_{p}F_q$ specification. One productive route is to identify special-function structures that satisfy the shape restrictions required by the econometric object of interest, and to develop inference for the resulting low-dimensional family. The present family is one such object. Future work can extend the foundation to varying tail-index and asymmetric hypergeometric families, generated residuals from volatility models, multivariate settings, likelihood-based estimation, and broader classes of hypergeometric distributional deformations.

The rest of the paper is organised as follows. Section~\ref{sec:hgm_family} introduces the hypergeometric family and its main distributional properties. Section~\ref{sec:inference} defines the minimum-distance criterion, states the regularity conditions, and derives consistency, interior asymptotics, boundary inference, and local asymptotic power. Section~\ref{sec:plugin} develops plug-in location-scale versions and their asymptotic theory, including the robust median/IQR statistic. Section~\ref{sec:simulations} reports size and power simulations. Section~\ref{sec:application} presents the S\&P~500 empirical illustration. Section~\ref{sec:conclusion} concludes and discusses extensions. Proofs are in Appendix~\ref{app:proofs}. Simulation implementation details, including the full quadrature specification, null constants, and power design, are in Appendix~\ref{app:simulations}. Additional empirical tables and figures are in Appendix~\ref{app:empirical}. Appendix~\ref{app:crossover} derives the tail dominance and Value-at-Risk (VaR) implications of a departure from the Gaussian boundary. Appendix~\ref{app:joint} provides a short self-contained theoretical treatment of joint location-scale estimation as a reference against which the plug-in procedures of Section~\ref{sec:plugin} are compared. The logical dependencies among the main results are given in Appendix~\ref{app:dependencies}.

\section{The hypergeometric family}\label{sec:hgm_family}

This section introduces the hypergeometric family $\{G_c:c \geq 3/2\}$, establishes that it consists of admissible distribution functions, and derives its main distributional properties. The starting point is a classical identity that expresses the standard normal distribution function in terms of Kummer's confluent hypergeometric function ${}_1F_1$. Generalising the denominator parameter of that representation from $3/2$ to a free parameter $c \geq 3/2$ generates a one-parameter family of symmetric distribution functions that contains the Gaussian as an exact boundary case. The properties of the family (admissibility, tail behaviour, moment existence, the phase transition at $c=3/2$, and the beta-precision scale-mixture representation) are derived from the endpoint constraints of Lemma~\ref{lem:endpoint} and standard properties of the ${}_1F_1$ function, and are stated in the lemmas, proposition, and corollaries below; unimodality and the characteristic function are established as consequences of the mixture representation; the location-scale extension follows from the basic distributional properties of the family. Readers unfamiliar with hypergeometric functions will find the necessary background in \citet{abramowitz_stegun72}, \citet{abadir99}, and \citet{olver_etal10}; each property of the ${}_1F_1$ function used in the proofs is stated explicitly when first needed.

\begin{lemma}[Confluent hypergeometric representation of $\Phi$]\label{lem:normalcdf}
    For all \(z\in\mathbb R\),
    \begin{equation}\label{eq:normalcdf}
        \Phi(z) = \frac12 + \frac{z}{\sqrt{2\pi}}\,
        {}_1F_1\!\left(\frac12;\frac32;-\frac{z^2}{2}\right).
    \end{equation}
\end{lemma}

\noindent Equation~\eqref{eq:normalcdf}
is the constructive starting point of the paper. Replacing the denominator
parameter $3/2$ by a free parameter $c\geq 3/2$ generates a candidate
one-parameter family. The next result shows that the endpoint restrictions
required of a distribution function force the numerator parameter to equal
$1/2$ and determine the normalising constant uniquely.

\begin{lemma}[Endpoint constraints on the hypergeometric ansatz]
\label{lem:endpoint}
Let $a,c\in\mathbb{R}$ with $c\notin\mathbb{Z}_{0,-}$ and
$c-a\notin\mathbb{Z}_{0,-}$, and let $\kappa(a,c)$ be a real constant. Define
\[
    G_{a,c}(z) = \frac12 + z\,\kappa(a,c)\,{}_1F_1\!\left(a;c;-\frac{z^2}{2}\right), \quad z\in\mathbb{R}.
\]
If $G_{a,c}$ satisfies the endpoint conditions
\[
    \lim_{z\to-\infty}G_{a,c}(z)=0,
    \quad
    \lim_{z\to+\infty}G_{a,c}(z)=1,
\]
then $a=1/2$, and the normalising constant is determined uniquely by
\begin{equation}\label{eq:Khalf}
    \kappa\!\left(\tfrac12,c\right) = \frac{\Gamma(c-1/2)}{2\sqrt{2}\,\Gamma(c)}.
\end{equation}
\end{lemma}

\begin{remark}
At $c=3/2$, using $\Gamma(3/2)=\sqrt{\pi}/2$,
\[
    \kappa\!\left(\tfrac12,\tfrac32\right) = \frac{\Gamma(1)}{2\sqrt{2}\,\Gamma(3/2)} = \frac{1}{\sqrt{2\pi}},
\]
which matches the constant in Lemma~\ref{lem:normalcdf}. Hence
$G_{1/2,3/2}=\Phi$, confirming that the normal distribution is
recovered exactly at the boundary.
\end{remark}

\noindent Having fixed $a=1/2$, we henceforth write
\begin{equation}\label{eq:Gc_def}
    G_c := G_{1/2,c}, \quad \kappa(c) := \kappa\!\left(\tfrac12,c\right) = \frac{\Gamma(c-1/2)}{2\sqrt{2}\,\Gamma(c)},
    \quad c\geq\frac32.
\end{equation}

Lemma~\ref{lem:endpoint} fixes $a=1/2$ and determines the normalising
constant $\kappa(c)$. It remains to verify that $G_c$ is non-decreasing,
so that $g_c = G_c' \geq 0$ and $G_c$ is a genuine distribution
function. This is the content of the next lemma.

\begin{lemma}[Monotonicity of $G_c$]\label{lem:monotonicity}
For every $c\geq 3/2$, the function $G_c$ defined in \eqref{eq:Gc_def} satisfies $G_c'(z)\geq 0$ for all $z\in\mathbb{R}$.
\end{lemma}

\noindent The proof in fact establishes strict positivity:
$G_c'(z)>0$ for all $z\in\mathbb{R}$ and all $c\geq\tfrac32$. Combining Lemmas~\ref{lem:endpoint} and~\ref{lem:monotonicity} yields two immediate corollaries.

\begin{corollary}[$G_c$ is a distribution function]\label{cor:Gc_distn}
For every $c\geq 3/2$, $G_c$ defined in \eqref{eq:Gc_def} is a
distribution function on $\mathbb{R}$.
\end{corollary}

\begin{corollary}[Density of $G_c$]\label{cor:density}
For every $c\geq 3/2$, $G_c$ is absolutely continuous with density
\begin{equation}\label{eq:gc}
    g_c(z) = \kappa(c)\,{}_1F_1\!\left(\frac32;c;-\frac{z^2}{2}\right).
\end{equation}
The density $g_c$ is even, and $g_{3/2}=\phi$.
\end{corollary}

\noindent The family $\{G_c: c\geq 3/2\}$ is therefore a well-defined one-parameter family of distribution functions indexed by $c$. At the boundary $c=3/2$, the density reduces to the standard normal: $g_{3/2}=\phi$. The family varies continuously and smoothly in its parameter, as the next two lemmas show.

\begin{lemma}[Continuity of $G_c$ in $c$]\label{lem:continuity}
For each fixed $z\in\mathbb{R}$, the map $c\mapsto G_c(z)$ is
continuous on $[3/2,\infty)$. Moreover, $G_c\to\Phi$ uniformly
in $z$ as $c\downarrow 3/2$.
\end{lemma}

\begin{lemma}[Differentiability of $G_c$ in $c$]\label{lem:differentiability}
For each fixed $z\in\mathbb{R}$, the map $c\mapsto G_c(z)$ is infinitely differentiable on $(3/2,\infty)$. In particular, the first derivative is
\begin{equation}\label{eq:Gdot}
    \dot G_c(z) := \frac{\partial G_c(z)}{\partial c}
    = z\left[\dot \kappa(c)\,{}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right)
    + \kappa(c)\,\frac{\partial}{\partial c}\,{}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right)
    \right],
\end{equation}
where
\begin{equation}\label{eq:Kdot}
    \dot \kappa(c) = \frac{\partial \kappa(c)}{\partial c}
    = \kappa(c)\left[\psi(c-\tfrac12)-\psi(c)\right],
\end{equation}
$\psi=\Gamma'/\Gamma$ is the digamma function, and
\begin{equation}\label{eq:1F1dot}
    \frac{\partial}{\partial c}\,{}_1F_1\!\left(\frac12;c;-\frac{z^2}{2}\right)
    = -\sum_{k=1}^{\infty}
    \frac{(1/2)_k}{(c)_k\,k!}\left(-\frac{z^2}{2}\right)^k\,
    \left[\psi(c+k)-\psi(c)\right].
\end{equation}
\end{lemma}

\noindent The admissibility of $G_c$ shown above does not yet reveal how the family departs from the Gaussian. The next lemma makes this precise by characterising the tail behaviour of $g_c$.

\begin{lemma}[Tail behaviour of $g_c$]
\label{lem:tail}
For $c\geq 3/2$, as $|z|\to\infty$:
\begin{enumerate}[label=\textup{(\roman*)}]
    \item if $c>3/2$, the density has an algebraic tail,
    \begin{equation}\label{eq:tail_algebraic}
        g_c(z) = \left(c-\frac32\right)|z|^{-3} + o(|z|^{-3});
    \end{equation}
    \item if $c=3/2$, then $g_{3/2}=\phi$
exactly, with super-exponential tail decay
$g_{3/2}(z)=\frac{1}{\sqrt{2\pi}}\,e^{-z^2/2}
\to0$ faster than any power of $|z|^{-1}$
as $|z|\to\infty$.
\end{enumerate}
\end{lemma}

\noindent The boundary $c=3/2$ therefore separates Gaussian from
algebraic tail behaviour: for $c>3/2$ the density decays as
$|z|^{-3}$, with the scale of the algebraic component governed
by $c-3/2$. This is the sense in which $c$ indexes a directed
heavy-tailed deformation of the Gaussian law. Integrating the tail of Lemma~\ref{lem:tail} gives the corresponding survival function asymptotic.

\begin{corollary}[Survival function tail]\label{cor:survival}
For $c>3/2$, as $z\to\infty$,
\begin{equation}\label{eq:survival}
    1-G_c(z) \sim \frac{c-3/2}{2}\,z^{-2}.
\end{equation}
In particular, the survival function has tail index $2$ for every $c>3/2$: the index is fixed across the family and $c$ governs only the scale of the tail.
\end{corollary}

\noindent Lemma~\ref{lem:tail} and Corollary~\ref{cor:survival} together characterise the tail structure of the family completely. Two remarks place this structure in context.

\begin{remark}[Phase transition at $c=3/2$]\label{rem:phasetransition}
For $c>3/2$ the algebraic tail coefficient $c-3/2>0$ decreases continuously to zero as $c\downarrow3/2$. At the boundary it vanishes
exactly and $g_{3/2}=\phi$ by Corollary~\ref{cor:density}: the tail switches from algebraic to super-exponential.
\end{remark}

\begin{remark}[Comparison with standard heavy-tailed families]
\label{rem:comparison}
The non-Gaussian members $\{G_c : c>3/2\}$ of the family are distinct from standard symmetric heavy-tailed families. By Corollary~\ref{cor:survival}, the survival function satisfies $1-G_c(z)\sim\frac{c-3/2}{2}z^{-2}$, so the tail index is fixed at $2$ for every $c>3/2$: the parameter $c$ governs the scale of the algebraic tail, not its index. Symmetric $\alpha$-stable laws have tail index $\alpha\in(0,2)$ in the heavy-tailed case, so no non-Gaussian stable law has tail index $2$. The Student $t_k$ family (equivalently, the Pearson Type~VII family with shape parameter $m=(k+1)/2$) has tail index $k$, so index $2$ forces $k=2$; but $G_c$ is not a rescaling of $t_2$, since the peak $g_c(0)=\kappa(c)$ and the tail coefficient $c-3/2$ vary with $c$ at different rates, whereas a scale family ties them through a single parameter: rescaling $z\mapsto z/\sigma$ multiplies the peak by $\sigma^{-1}$ and the tail coefficient by $\sigma^2$, so their product
$\kappa(c)\,(c-3/2)^{1/2}$ would be constant across a scale family but varies with $c$ in the hypergeometric family. The hypergeometric family shares with $t_2$ only the moment structure (finite mean, infinite variance), not its functional form. More broadly, the Student $t_k$ family for general $k$ is symmetric and, like the hypergeometric family, unable to capture skewness; it places the Gaussian at the asymptotic limit $k\to\infty$ rather than at a finite boundary point. The structural advantage of $\{G_c\}$ over the Student $t$ family is precisely this finite boundary: the exact location of the Gaussian at $c=3/2$ enables the $\frac12\delta_0+\frac12\chi_1^2$ boundary limit
of Theorem~\ref{thm:fit_improvement}.
\end{remark}

\noindent The fixed tail index has immediate implications for the existence of moments.

\begin{corollary}[Moment existence]\label{cor:moments}
Let $c>3/2$ and $Z\sim G_c$. For $r>0$,
\[
    \operatorname{E}[|Z|^r]<\infty
    \quad\Longleftrightarrow\quad
    r<2.
\]
In particular, $\operatorname{E}[Z]=0$ by symmetry and $\operatorname{E}[Z^2]=\infty$: every
non-Gaussian member of the family has zero mean and infinite variance.
\end{corollary}

\begin{remark}[Tail departure from normality]\label{rem:heavytail}
The moment structure is markedly non-Gaussian for every $c>3/2$,
regardless of how close $c$ is to the boundary: $G_c$ is
symmetric with finite mean but infinite variance, since the
algebraic tail $|z|^{-3}$ is too heavy for a second moment to
exist. The parameter $c$ indexes departure from normality in the tails.
\end{remark}

\begin{figure}[t]
    \centering
    \includegraphics[width=\textwidth]{hgm_family.pdf}
    \caption{The hypergeometric family $G_c$ and its density $g_c$ for $c\in\{1.5,1.6,2.0,3.0\}$, with $c=3/2$ the standard normal. \emph{Left:} densities $g_c$ on linear axes, showing that the family is a smooth perturbation of the normal near the centre, with peak height decreasing in $c$. \emph{Right:} survival functions $1-G_c$ on log-log axes; for $c>3/2$ the curves are asymptotically linear with common slope $-2$, reflecting the fixed tail index in
    $1-G_c(z)\sim\tfrac{c-3/2}{2}\,z^{-2}$ for every $c>3/2$; the normal ($c=3/2$) falls away with super-exponential decay.}
    \label{fig:hgm_family}
\end{figure}

Figure~\ref{fig:hgm_family} illustrates the two features of the family established above: proximity to the normal at the centre and qualitatively distinct tails. The left panel plots the densities $g_c$: near the origin the curves are close to the standard normal, to which they reduce exactly at $c=3/2$; as $c$ increases the peak falls and mass shifts toward the tails, interpolating continuously away from normality. The right panel plots the survival functions $1-G_c$ on log-log axes, on which an algebraic tail appears as a straight line with slope equal to
the negative tail index. For every $c>3/2$ the slope is $-2$, consistent with Corollary~\ref{cor:survival}; $c$ shifts the intercept but not the slope, confirming that the tail index is fixed and $c$ governs only the tail scale. The normal case $c=3/2$ is qualitatively different: its survival function curves steeply downward, reflecting super-exponential decay, and falls below every algebraic tail beyond a moderate threshold. The figure thus displays the phase transition of
Remark~\ref{rem:phasetransition} directly.

The distributional properties established above (algebraic tails, fixed tail index, and finite mean with infinite variance) all follow from the hypergeometric density \eqref{eq:gc}, but they also admit a unified probabilistic explanation that connects the construction to a classical mixture structure and provides intuition for the role of $c$.

\begin{proposition}[Beta-precision normal scale mixture]\label{prop:mixture}
For every $c>3/2$, let $\varepsilon\sim\operatorname{N}(0,1)$ and
$\Lambda\sim\operatorname{Beta}(1,c-3/2)$ be independent. Then
$Z=\varepsilon/\sqrt{\Lambda}$ has distribution $G_c$. Equivalently,
\begin{equation}\label{eq:mixture_density}
    g_c(z) = \int_0^1 \frac{\sqrt{\lambda}}{\sqrt{2\pi}}\,
    e^{-\lambda z^2/2}\,\left(c-\tfrac32\right)(1-\lambda)^{c-5/2}\,d\lambda,
\end{equation}
the marginal density obtained by averaging $\operatorname{N}(0,\lambda^{-1})$
densities over $\Lambda\sim\operatorname{Beta}(1,c-3/2)$.
\end{proposition}

\begin{remark}[Interpretation and boundary]\label{rem:mixture}
The representation clarifies the role of $c$ at two levels.
Probabilistically, $c-3/2$ is the second shape parameter of the Beta distribution governing the latent precision $\Lambda\in(0,1)$: small values of $\Lambda$ generate high-variance conditional Gaussian observations and hence the algebraic tails of $G_c$.

As $c \downarrow 3/2$, the $\mathrm{Beta}(1,c-3/2)$ distribution converges weakly to a point mass at one, $\Lambda \equiv 1$. We define the boundary case by this degenerate limiting distribution, so $Z\mid(\Lambda=1)\sim\mathrm{N}(0,1)$ and $G_{3/2}=\Phi$ exactly. Nonnegativity of $g_c$ is immediate from \eqref{eq:mixture_density}, since the integrand is a product of a $\operatorname{N}(0,\lambda^{-1})$ density and a $\operatorname{Beta}(1,c-3/2)$ density, both non-negative. The representation connects the hypergeometric family to the normal variance-mixture literature
\citep{andrews_mallows74, barndorff-nielsen_etal82}, but the inference problem studied here is not the general estimation of a normal mixture: it is the boundary test of Gaussian degeneracy within this particular beta-precision hypergeometric path.
\end{remark}

\begin{corollary}[Unimodality]\label{cor:unimodal}
For every $c\geq\tfrac32$, the density $g_c$ is strictly unimodal with mode at $z=0$: $g_c(0)>g_c(z)$ for all $z\neq0$, and $g_c$ is strictly decreasing on $(0,\infty)$.
\end{corollary}

\begin{corollary}[Characteristic function]\label{cor:charfun}
Let $Z\sim G_c$ with $c>3/2$. Then
\begin{equation}\label{eq:charfun}
    \varphi_c(t) = \operatorname{E}[e^{itZ}] = \Gamma(c-1/2)\,\exp\!\left(-\frac{t^2}{2}\right)\,
    \Psi\!\left(c-\frac32;\,0;\,\frac{t^2}{2}\right), \quad t\in\mathbb{R},
\end{equation}
where $\Psi$ is Tricomi's confluent hypergeometric function. Equivalently,
\begin{equation}\label{eq:charfun_integral}
    \varphi_c(t) = (c-3/2)\int_0^1 \exp\!\left(-\tfrac{t^2}{2\lambda}\right)\,
    (1-\lambda)^{c-5/2}\,d\lambda, \quad t\in\mathbb{R}.
\end{equation}
As $c\downarrow 3/2$, $\varphi_c(t)\to\exp(-t^2/2)$,
the characteristic function of the standard normal.
\end{corollary}

\begin{remark}[Tricomi versus Kummer]\label{rem:charfun}
The characteristic function involves Tricomi's $\Psi$ rather than the Kummer ${}_1F_1$ that appears in $g_c$. This is not accidental: the density averages the conditional Gaussian density over the beta-precision distribution via the Euler integral, giving ${}_1F_1$; the characteristic function averages $\exp(-t^2/(2\lambda))$ over the same distribution, but the reciprocal $\lambda^{-1}$ in the exponent produces the Tricomi representation instead. Because the second parameter is zero, there is no simple ${}_1F_1$ expression for $\Psi(c-\tfrac32;0;\tfrac{t^2}{2})$; the
nearest identity \citep[eq.~13.1.7]{abramowitz_stegun72} is
$\Psi(c-\tfrac32;0;\tfrac{t^2}{2})=\tfrac{t^2}{2}\,\Psi(c-\tfrac12;2;\tfrac{t^2}{2})$. Since $\operatorname{E}[Z^2]=\infty$ for all $c>3/2$
by Corollary~\ref{cor:moments}, and since twice differentiability of $\varphi_c$ at the origin is equivalent to finite second moment, $\varphi_c$ is not twice differentiable at $t=0$. For every $c>3/2$, methods that rely on second moments, moment-generating functions, or cumulant expansions are unavailable. Although these methods are well defined at the Gaussian boundary $c=3/2$, they are not available on any right neighbourhood of that boundary and therefore do not provide a regular basis for inference on c. This provides a further motivation for the minimum-distance distribution-function approach of Section~\ref{sec:inference}.
\end{remark}

\begin{corollary}[Location-scale extension]\label{cor:location_scale_family}
For $\mu\in\mathbb{R}$, $\sigma>0$, and $c\geq\tfrac32$, define
$G_{\mu,\sigma,c}(x)=G_c\bigl((x-\mu)/\sigma \bigr)$. Then $G_{\mu,\sigma,c}$ is a distribution function on $\mathbb{R}$ with
density $g_{\mu,\sigma,c}(x)=\sigma^{-1} g_c\bigl((x-\mu)/\sigma\bigr)$, and
$G_{\mu,\sigma,3/2}(x)=\Phi\bigl((x-\mu)/ \sigma\bigr)$. For $c>\tfrac32$, the density is symmetric about $\mu$ with
\[
    g_{\mu,\sigma,c}(x) \sim \bigl(c-\tfrac32\bigr)\sigma^2|x-\mu|^{-3},
    \quad|x|\to\infty.
\]
If $X\sim G_{\mu,\sigma,c}$ with $c>\tfrac32$, then $\operatorname{E}[|X-\mu|^r]<\infty
\Leftrightarrow r<2$; in particular $\operatorname{E}[X]=\mu$, while the variance does
not exist.
\end{corollary}

\section{Minimum-distance inference}\label{sec:inference}

This section develops the inferential theory for the hypergeometric family. The minimum-distance criterion measures the squared $\nu$-distance between the empirical distribution function and the parametric family $\{G_c:c\in\mathcal{C}\}$, and the estimator $\widehat c_n$ minimises it over the admissible parameter space. The Gaussian null $H_0:c=3/2$ lies on the boundary of $\mathcal{C}$, which is the source of the nonstandard limiting distribution of the test statistic. Consistency, interior asymptotics, boundary inference, directed consistency, and local asymptotic power are derived in turn.

\subsection{Criterion, estimator, and test statistic}\label{subsec:criterion}

Let $Z_1,\ldots,Z_n$ be i.i.d.\ with distribution function $F_0$ and empirical distribution function $\widehat F_n$. For a fixed finite measure $\nu$ on $\mathbb{R}$, define the minimum-distance criterion
\[
    Q_n(c) = \int_{\mathbb{R}} \{\widehat F_n(z)-G_c(z)\}^2\,d\nu(z),\quad c\in\mathcal{C}=\left[\frac32,c_U\right],
\]
and its population counterpart
\[
    Q(c) = \int_{\mathbb{R}}\{F_0(z)-G_c(z)\}^2\,d\nu(z).
\]
The parameter space $\mathcal{C}=[3/2,c_U]$ is compact, with the lower bound $3/2$ dictated by the admissibility region of the family and $c_U\in(3/2,\infty)$ a fixed finite upper bound chosen large enough not to bind in applications. The choice of $\nu$ is part of the definition of the criterion. Since the parameter $c$ controls a Gaussian-to-heavy-tail deformation, a natural class of tail-sensitive measures is
\begin{equation}\label{eq:nu_lambda}
    d\nu_\lambda(z) \propto
    [\Phi(z)\{1-\Phi(z)\}]^{-\lambda}\,d\Phi(z),
    \quad 0\leq\lambda<1,
\end{equation}
which increases the relative weight on the tails as $\lambda$ increases while keeping $\nu_\lambda$ finite. The substitution $u=\Phi(z)$ transforms $\nu_\lambda$ to a $\operatorname{Beta}(1-\lambda,1-\lambda)$ measure on $(0,1)$ with total mass given by the beta function $\operatorname{B}(1-\lambda,1-\lambda)=(\Gamma(1-\lambda))^2/\Gamma(2-2\lambda)$, which is finite if and only if $\lambda<1$. The simulations and empirical illustration use  $\lambda=1/2$, giving
\[
    d\nu_{1/2}(z) \propto [\Phi(z)\{1-\Phi(z)\}]^{-1/2}\,d\Phi(z).
\]
The rationale for the $d\Phi(z)$ base measure, the tail-weighting structure, and the choice $\lambda=1/2$ are discussed in Remark~\ref{rem:nu_choice}. The minimum-distance estimator is any
\[
    \widehat c_n\in\arg\min_{c\in\mathcal{C}}Q_n(c).
\]
The Gaussian null $H_0:F_0=\Phi$ corresponds to the boundary value $c_0=3/2$. The test statistic is the improvement in fit from estimating $c$ rather than imposing $c=c_0$:
\[
    T_n = n\{Q_n(c_0)-Q_n(\widehat c_n)\}.
\]
Since $\widehat c_n\in\mathcal{C}$ and $c_0=3/2$ is the left
endpoint of $\mathcal{C}$, $T_n\geq 0$ by construction. The
asymptotic properties of $\widehat c_n$ and $T_n$ are derived in
the subsections below.

\begin{remark}[Structure of $\nu_\lambda$ and the choice $\lambda=1/2$]
\label{rem:nu_choice} Two structural features of the family $\nu_\lambda$ deserve comment: (i) \textit{Role of the $d\Phi(z)$ base.}
Using $d\Phi(z)=\phi(z)\,dz$ as the base measure rather than Lebesgue measure $dz$ is not just a normalisation convenience. On the quantile scale the weight $[u(1-u)]^{-\lambda}$ diverges at $u=0$ and $u=1$, and against Lebesgue measure this divergence is not integrable for any $\lambda>0$. The factor $\phi(z)$ in $d\Phi(z)$ provides super-exponential damping in both tails, converting the divergent weight into the integrable unnormalised $\operatorname{Beta}(1-\lambda,1-\lambda)$ density on $(0,1)$. It follows that $\lambda=0$ gives $d\nu_0\propto d\Phi(z)$, a measure concentrated near the centre, and Lebesgue measure $dz$ (Cram\'er--von Mises) is not a member of the $\nu_\lambda$ family for any $\lambda\in[0,1)$. (ii) \textit{Choice of $\lambda$.}
Fixing $\lambda=1/2$ delivers a finite, symmetric, tail-sensitive criterion with analytically computable null constants. In principle $\lambda$ could be selected to maximise power against a target alternative class, but the optimal choice depends on the direction of departure from the Gaussian null, which is unknown in practice; a data-adaptive procedure would require controlling the joint distribution of the selection step and the test statistic, a non-standard problem requiring bootstrap calibration. The cost of fixing $\lambda=1/2$ relative to an oracle-optimal choice is expected to be small against symmetric heavy-tailed alternatives, precisely the class for which the hypergeometric score direction $\dot G_{c_0}$ and the $\nu_{1/2}$ weighting are well-aligned; the local power analysis of
Section~\ref{subsec:localpower} makes this precise.
\end{remark}

\begin{remark}[Comparison with likelihood-based estimation]\label{rem:mle}
Maximum likelihood estimation (MLE) is feasible in principle using the explicit density in Corollary~\ref{cor:density}, or through the beta-precision mixture representation of Proposition~\ref{prop:mixture}. Its boundary theory is, however, non-regular: the model is not differentiable in quadratic mean at the Gaussian boundary, and the usual $\sqrt{n}$ likelihood expansion and regular Fisher-information theory do not apply. The minimum-distance approach avoids this problem and extends naturally to hypergeometric models where explicit densities may not be available.
\end{remark}

\subsection{Assumptions}\label{subsec:assumptions}

\begin{enumerate}
    \item[(A0)] \textit{(Family regularity.)}
    For each $c\in\mathcal{C}$, $G_c$ is a distribution
    function; for each $z\in\mathbb{R}$, the map
    $c\mapsto G_c(z)$ is continuous.

    \item[(A1)] \textit{(Compact parameter space.)}
    $\mathcal{C}=[3/2,c_U]$ with $c_U\in(3/2,\infty)$ fixed.

    \item[(A2)] \textit{(Finite measure.)}
    $\nu(\mathbb{R})<\infty$.

    \item[(A3)] \textit{(Identification.)}
    $Q$ has a unique minimiser $c^\star\in\mathcal{C}$.

    \item[(A4)] \textit{(Tail sensitivity.)}
    $\nu((M,\infty))>0$ and $\nu((-\infty,-M))>0$ for all
    $M>0$.
\end{enumerate}

\noindent Assumption (A0) is not a maintained assumption but a verified property of the hypergeometric family: $G_c$ is a distribution function for every $c\geq 3/2$ by Corollary~\ref{cor:Gc_distn}, and $c\mapsto G_c(z)$ is continuous for each fixed $z$ by Lemma~\ref{lem:continuity}. It is listed explicitly because it is the condition on the parametric family that drives both uniform convergence of $Q_n$ and continuity of $Q$, and would need to be re-verified for any extension of the family. Assumption (A1) ensures minimisers exist and the argmin theorem applies. The upper bound $c_U$ is a technical device that does not affect the asymptotic distribution under the null and is chosen large enough not to
bind in practice. Assumption (A2) is part of the definition of $\nu$ and makes the criterion well defined. It also converts the pointwise Glivenko--Cantelli bound into a uniform bound on $Q_n-Q$. Assumption (A3) is the substantive condition. Under correct specification it holds by identifiability: distinct values of $c$ yield distinct tail coefficients in Lemma~\ref{lem:tail}, hence distinct distributions. Under misspecification $c^\star$ is the pseudo-true value and (A3) is a genuine condition on the geometry of $F_0$ relative to $\{G_c\}$. In practice this is a weak regularity condition: multiple minimisers would require a non-generic alignment between $F_0$ and the hypergeometric family, analogous to the standard identification conditions in minimum-distance and GMM estimation. Assumption (A4) requires $\nu$ to assign positive mass to every tail region. It ensures the criterion is sensitive to the tail differences that distinguish members of the family; it is needed only for identification under correct specification, and is
satisfied by $\nu_{1/2}$ by construction.

\begin{remark}[Identification, misspecification, and the boundary]
\label{rem:identification}
Three features of the setup merit comment. (i)~The identifying signal for $c$ weakens as $c\downarrow 3/2$: by Lemma~\ref{lem:continuity}, $G_c\to\Phi$ uniformly, so all features that distinguish $G_c$ from the Gaussian, including the algebraic tail and the peak height $g_c(0)=\kappa(c)$, vanish continuously at the boundary. This is intrinsic to any test of normality against alternatives that approach the
null in distribution, rather than a limitation of the present
procedure; it is precisely what the boundary distribution theory
below addresses. (ii)~Under misspecification, $\widehat c_n$ converges to the pseudo-true value $c^\star$, the minimum-distance projection of $F_0$ onto $\{G_c:c\in\mathcal{C}\}$: an interpretable index of tail heaviness whether or not $F_0\in\{G_c\}$. (iii) Under the Gaussian null, $c^\star=c_0=3/2$ lies on the boundary of $\mathcal{C}$. Consistency is unaffected, but the boundary position means that standard $\sqrt{n}$-asymptotic normality fails: the limit is of one-sided type, as derived below.

Taken together, rejection of $H_0:c=3/2$ carries a direction
(departure from normality toward heavier algebraic tails, quantified by $\widehat c_n-3/2$) that omnibus tests do not provide. The cost is reduced power near the null in small samples; the procedure is most informative at larger sample sizes and on data where the departure is of the symmetric heavy-tailed kind it is designed to detect. The simulations
below examine this trade-off directly.
\end{remark}

\subsection{Consistency}\label{subsec:consistency}

The first main result establishes almost sure convergence of
$\widehat c_n$ to the pseudo-true value $c^\star$, both under
correct specification and under misspecification. Under correct specification $c^\star$ is the true parameter and $Q(c^\star)=0$; under
misspecification $c^\star$ is the closest member of the family
to $F_0$ in the minimum-distance sense and $Q(c^\star)>0$.

\begin{theorem}[Consistency]\label{thm:consistency}
Under \textup{(A0)--(A3)}, any $\widehat c_n\in\arg\min_{c\in\mathcal{C}}Q_n(c)$ satisfies
$\widehat c_n\to c^\star$ almost surely.
\end{theorem}

\begin{remark}[Pseudo-true value and misspecification]\label{rem:pseudotrue}
Under misspecification, $\widehat c_n$ consistently estimates the pseudo-true value $c^\star$: the parameter indexing the hypergeometric distribution closest to $F_0$ in the minimum-distance sense. A rejection of $H_0:c=3/2$ therefore retains meaning under misspecification: it indicates that the best hypergeometric approximation to $F_0$ is not the Gaussian boundary; and $\widehat c_n-3/2$ quantifies the degree of tail departure from normality in the minimum-distance sense.
\end{remark}

\noindent Under correct specification, \textup{(A3)} is verified by the following lemma, which uses \textup{(A4)} to establish that distinct members of the family are distinguished by the criterion $Q$.

\begin{lemma}[Identification under correct specification]
\label{lem:identifiability}
Under \textup{(A4)}, the map $c\mapsto G_c$ is identified on $\mathcal{C}$: $G_c=G_{c_0}$ $\nu$-almost everywhere if and only if $c=c_0$. Consequently, if $F_0=G_{c_0}$ for some $c_0\in\mathcal{C}$, then $c^\star=c_0$ is the unique minimiser of $Q$, so \textup{(A3)} holds.
\end{lemma}

\begin{corollary}[Consistency under correct specification]
\label{cor:consistency_correct}
Under \textup{(A0)--(A2)} and \textup{(A4)}, if $F_0=G_{c_0}$ for some $c_0\in\mathcal{C}$, then $\widehat c_n\to c_0$ almost surely. In particular, under the Gaussian null $F_0=\Phi=G_{3/2}$, $\widehat c_n\to 3/2$ almost surely.
\end{corollary}

\subsection{Interior asymptotics}\label{subsec:interior}

\noindent Theorem~\ref{thm:consistency} establishes that $\widehat c_n$ converges to $c^\star$ almost surely. The next step is to determine the rate of convergence and the limiting distribution. When $c^\star$ lies in the interior of $\mathcal{C}$, differentiability of $c\mapsto G_c(z)$
established in Lemma~\ref{lem:differentiability} justifies differentiation of $Q_n$ under the integral sign under the regularity conditions of Theorem~\ref{thm:interior} below, so $Q_n$ is locally smooth in $c$ and the standard $\sqrt{n}$-theory applies. The asymptotic distribution is normal with a sandwich variance determined by three objects: the derivative $\dot G_{c^\star}(z):=\partial G_c(z)/\partial c|_{c=c^\star}$, a function of $z\in\mathbb{R}$ given explicitly by Lemma~\ref{lem:differentiability} (I write $\dot G_{c^\star}$ suppressing $z$ hereafter); the scalar curvature $H_{c^\star}:=\partial^2 Q(c)/\partial c^2|_{c=c^\star}$; and the scalar limiting variance $\Omega_{c^\star}$ of $\sqrt{n}\,\partial Q_n(c)/\partial c|_{c=c^\star}$. These objects also drive the boundary theory and plug-in extensions below, giving a modular structure that carries through the paper. The following additional assumptions govern the local behaviour
of $Q_n$ near $c^\star$.

\begin{enumerate}
    \item[(B1)] \textit{(Smoothness.)} For each $z\in\mathbb{R}$, the map $c\mapsto G_c(z)$ is twice continuously differentiable on a neighbourhood $\mathcal{N}$ of $c^\star$.

    \item[(B2)] \textit{(Uniform domination.)} There exist
    $\nu$-integrable functions $M_1,M_2$ with
    $|\dot G_c(z)|\leq M_1(z)$ and
    $|\ddot G_c(z)|\leq M_2(z)$ for all $c\in\mathcal{N}$,
    with $\int M_1^2\,d\nu<\infty$ and
    $\int M_2\,d\nu<\infty$.

    \item[(B3)] \textit{(Local identification.)}
    $H_{c^\star}:=Q''(c^\star)>0$.
\end{enumerate}

\noindent Assumption~(B1) is a verified property of the hypergeometric family. It is listed explicitly because it is the condition required for differentiation under the integral sign in $Q_n$, and would need re-verification for extensions beyond $G_c$. By Lemma~\ref{lem:differentiability}, $c\mapsto G_c(z)$ is already known to be smooth on $(3/2,\infty)$ for each fixed $z$; this is strengthened to real-analyticity by Lemma~\ref{lem:B1}, using the analyticity of $K(c)$ and the meromorphic structure of $c\mapsto{}_1F_1(\frac12;c;-z^2/2)$. Real-analyticity is stronger than required for \textup{(B1)}, which needs only twice continuous differentiability; it is established here for completeness and for use in extensions of the framework.

\begin{lemma}[Verification of \textup{(B1)}]\label{lem:B1}
For each fixed $z\in\mathbb{R}$, the map $c\mapsto G_c(z)$ is real-analytic on $(1/2,\infty)$, hence smooth on any neighbourhood of $c^\star$ in $\mathcal{C}=[3/2,c_U]$. In particular, \textup{(B1)} holds.
\end{lemma}

\noindent Assumption (B2) is likewise verified for the present family: $\dot G_c$ and $\ddot G_c$ are uniformly bounded in $z$ and $c\in\mathcal{N}$, so constant dominators suffice. Uniform domination plays two distinct roles in the proof: square-integrability of $M_1$ controls the limiting variance $\Omega$, while integrability of $M_2$ controls the Taylor remainder and the uniform convergence of $Q_n''$ to $Q''$.

\begin{lemma}[Verification of \textup{(B2)}]\label{lem:B2}
For any compact neighbourhood $\mathcal{N}$ of $c^\star$ in $(1/2,\infty)$, there exist finite constants $C_1,C_2$ such that $|\dot G_c(z)|\leq C_1$ and $|\ddot G_c(z)|\leq C_2$ for all $z\in\mathbb{R}$ and all $c\in\mathcal{N}$. In particular, for $c^\star\geq 3/2$, setting $M_1(z)=C_1$ and $M_2(z)=C_2$, both integrability conditions of \textup{(B2)} hold under \textup{(A2)}.
\end{lemma}

\noindent Assumption~(B3) requires the population criterion to have strictly positive curvature at $c^\star$, ruling out a locally flat criterion. Under correct specification, (B3) is verified by the following lemma. Under misspecification, $H_{c^\star}$ includes the additional term
$-2\int\{F_0-G_{c^\star}\}\ddot G_{c^\star}\,d\nu$, and \textup{(B3)} is a genuine maintained assumption.

\begin{lemma}[Positivity of the score-direction integral]\label{lem:B3}
Suppose $\nu$ assigns positive mass to some open interval. Then, for any $c^\star\geq3/2$,
\[
    2\int_{\mathbb{R}}\dot G_{c^\star}(z)^2\,d\nu(z) > 0.
\]
In particular, if $F_0=G_{c^\star}$, this quantity equals $H_{c^\star}=Q''(c^\star)$, verifying \textup{(B3)} under correct specification.
\end{lemma}

\noindent The score variance $\Omega_{c^\star}$ appearing in the asymptotic normality result below is, by the same argument, non-degenerate under the same hypothesis on $\nu$.

\begin{lemma}[Positivity of $\Omega$]\label{lem:Omega_pos}
Suppose $\nu$ assigns positive mass to some open interval. Then, for any $c^\star\geq3/2$, $\Omega_{c^\star}>0$.
\end{lemma}

\noindent With assumptions \textup{(B1)--(B3)} in place, together with the non-degeneracy of $\Omega_{c^\star}$ just established, $\widehat c_n$ is asymptotically normal at the standard $\sqrt{n}$ rate, with a non-degenerate limiting variance.

\begin{theorem}[Interior asymptotic normality]\label{thm:interior}
Suppose the conditions of Theorem~\ref{thm:consistency} and \textup{(B1)--(B3)} hold, and that $c^\star\in\operatorname{int}(\mathcal{C})$. Then
\[
    \sqrt{n}\,(\widehat c_n-c^\star)\;\Rightarrow\;
    \operatorname{N}\!\left(0,\,\frac{\Omega_{c^\star}}{H_{c^\star}^2}
    \right),
\]
where
\begin{equation}\label{eq:Omega}
    \Omega_{c^\star} = 4\iint_{\mathbb{R}^2}
    \bigl[F_0(z\wedge y)-F_0(z)F_0(y)\bigr]
    \dot G_{c^\star}(z)\,\dot G_{c^\star}(y)
    \,d\nu(z)\,d\nu(y),
\end{equation}
and $z\wedge y:=\min(z,y)$.
\end{theorem}

\begin{remark}[Structure of the sandwich variance]\label{rem:sandwich}
The asymptotic variance $\Omega_{c^\star}/H_{c^\star}^2$ in Theorem~\ref{thm:interior} has the sandwich form standard in minimum-distance asymptotics. The numerator $\Omega_{c^\star}$, given in \eqref{eq:Omega}, is the variance of the limiting score $\sqrt{n}\,Q_n'(c^\star)$: a weighted integral of the Brownian bridge covariance $F_0(z\wedge y)-F_0(z)F_0(y)$ against the score direction $\dot G_{c^\star}$. The denominator $H_{c^\star}^2$ is the squared curvature of the population criterion at $c^\star$. Under correct specification, $F_0=G_{c^\star}$, so $\sqrt{n}(\widehat F_n-G_{c^\star})\Rightarrow\mathbb{B}_{G_{c^\star}}$ and the covariance kernel in \eqref{eq:Omega} reduces to $G_{c^\star}(z\wedge y)-G_{c^\star}(z)G_{c^\star}(y)$, so that $\Omega_{c^\star}$ becomes the variance of the $G_{c^\star}$-Brownian bridge projected onto $\dot G_{c^\star}$. For the plug-in tests of Section~\ref{sec:plugin}, the same sandwich structure holds with the Brownian bridge covariance replaced by the projected covariance kernel that accounts for estimation of location and scale; $H_{c^\star}$ and $\dot G_{c^\star}$ are unchanged.
\end{remark}

\begin{remark}[Assumption \textup{(B3)} under misspecification]\label{rem:B_misspec}
Under misspecification, (B1) and (B2) remain verified properties of $G_c$. Only (B3) changes character: the curvature $H$ involves the additional term
$-2\int_\mathbb{R}\{F_0(z)-G_{c^\star}(z)\}\ddot G_{c^\star}(z)\,d\nu(z)$, and positive curvature is no longer automatic. This is the standard second-order local identification condition in minimum-distance asymptotics under misspecification \citep{newey_mcfadden94}, maintained as an assumption.
\end{remark}

\subsection{Boundary inference}\label{subsec:boundary}

This subsection establishes the limiting distribution of $\widehat c_n$ at the Gaussian null $c_0=3/2$, which lies on the boundary of $\mathcal{C}$, where $Q_n'(\widehat c_n)=0$ need not hold, and the interior argument of Theorem~\ref{thm:interior} does not apply. Beyond the failure of the first-order condition, the one-sided constraint $c\geq c_0$ reshapes the limit itself: sample paths along which the unconstrained Gaussian fluctuation would be negative are instead redirected to $\widehat c_n=c_0$, so the limit places mass $1/2$ at zero and is half-normal on $(0,\infty)$, rather than normal. The proof uses a localisation argument. Writing $c=c_0+h/\sqrt{n}$ with $h\geq 0$, the scaled criterion $L_n(h)=n\{Q_n(c_0+h/\sqrt{n})-Q_n(c_0)\}$ converges to a quadratic limit $L(h)$ whose minimiser over $h\geq 0$ is the censored Gaussian $\max\{0,Z\}$. The key ingredients are the score limit from Donsker's theorem, as in the proof of Theorem~\ref{thm:interior}; the two properties of $Q_n''$ established in Lemma~\ref{lem:curvature}: strict positivity of $H_{c_0}$ and uniform convergence of $Q_n''$ to $Q''$ on a neighbourhood of $c_0$; and the tightness of $\sqrt{n}(\widehat c_n-c_0)$ established in Lemma~\ref{lem:tightness}.

\begin{lemma}[Curvature at the Gaussian null]
\label{lem:curvature}
Let $\mathcal N\subset(1/2,\infty)$ be any compact neighbourhood of $c_0$, as in \textup{(B1)--(B2)}. Under \textup{(B1)--(B2)}, $F_0=\Phi=G_{3/2}$, and $\nu$ assigning positive mass to some open interval,
\begin{equation}\label{eq:Hc0}
    H_{c_0} = Q''(c_0) = 2\int_{\mathbb{R}}
    \dot G_{c_0}(z)^2\,d\nu(z) > 0.
\end{equation}
Moreover,
\begin{equation}\label{eq:Qn_uniform}
    \sup_{c\in\mathcal{N}}|Q_n''(c)-Q''(c)|\;\to_p\; 0.
\end{equation}
\end{lemma}

\noindent Although $c_0=3/2$ is the left endpoint of the parameter space $\mathcal C=[3/2,c_U]$, it is interior to $(1/2,\infty)$, so $\mathcal N$ is necessarily two-sided and extends strictly below $c_0$ as well as above. This costs nothing: (B1)--(B2) are verified on all of $(1/2,\infty)$ by Lemmas~\ref{lem:B1} and~\ref{lem:B2}, and in the proofs of Theorems~\ref{thm:interior} and~\ref{thm:boundary} the supremum in \eqref{eq:Qn_uniform} is applied only at points $c\geq c_0$ within $\mathcal C$. The following lemma establishes that $\sqrt{n}(\widehat c_n-c_0)$ remains bounded in probability: the tightness condition required by the argmax continuous mapping theorem.

\begin{lemma}[Tightness of the boundary estimator]\label{lem:tightness}
Suppose the conditions of
Corollary~\ref{cor:consistency_correct} and
\textup{(B1)--(B2)} hold, that $F_0=\Phi=G_{3/2}$, and that
$\nu$ assigns positive mass to some open interval. Then
\[
    \sqrt{n}\,(\widehat c_n-c_0) = O_p(1).
\]
\end{lemma}

\noindent With both lemmas in place, the boundary limit follows from a localisation argument and the argmax continuous mapping theorem of \citet{vandervaart_wellner23}.

\begin{theorem}[Boundary limit under the Gaussian null]\label{thm:boundary}
Suppose the conditions of Corollary~\ref{cor:consistency_correct} and
\textup{(B1)--(B2)} hold, that $F_0=\Phi=G_{3/2}$, and that $\nu$ assigns positive mass to some open interval. Then
\[
    \sqrt{n}\,(\widehat c_n-c_0)\;\Rightarrow\;\max\{0,\,Z\},
    \quad Z\sim\operatorname{N}\!\left(0,\,\frac{\Omega_{c_0}}{H_{c_0}^2}\right),
\]
where $H_{c_0}=2\int_{\mathbb{R}}\dot G_{c_0}(z)^2\,d\nu(z)$ and
\begin{equation}\label{eq:Omega_c0}
    \Omega_{c_0} = 4\iint_{\mathbb{R}^2}
    \bigl[\Phi(z\wedge y)-\Phi(z)\Phi(y)\bigr]
    \dot G_{c_0}(z)\,\dot G_{c_0}(y)
    \,d\nu(z)\,d\nu(y).
\end{equation}
The limit places mass $1/2$ at zero and is half-normal on $(0,\infty)$.
\end{theorem}

\begin{remark}[Interior versus boundary]
Theorems~\ref{thm:interior} and~\ref{thm:boundary} describe qualitatively different regimes. When $c^\star>3/2$, the estimator is regular: $\sqrt{n}$-consistent and asymptotically normal with sandwich variance $\Omega_{c^\star}/H_{c^\star}^2$. When $c^\star=c_0=3/2$, the one-sided constraint instead produces the point-mass-plus-half-normal limit of
Theorem~\ref{thm:boundary}.
\end{remark}

\noindent This nonstandard limit has an immediate consequence for testing: normal critical values from Theorem~\ref{thm:interior} do not apply at $c_0$. I construct a test from the fit-improvement statistic
$T_n=n\{Q_n(c_0)-Q_n(\widehat c_n)\}\geq0$, the minimum-distance analogue of a likelihood-ratio statistic; Remark~\ref{rem:test_principle} discusses this choice relative to the Wald and score alternatives.

\begin{theorem}[Boundary limit of the fit-improvement
statistic] \label{thm:fit_improvement}
Under the conditions of Theorem~\ref{thm:boundary}, the fit-improvement statistic $T_n=n\{Q_n(c_0)-Q_n(\widehat c_n)\}$ satisfies
\begin{equation}\label{eq:Tn_limit}
    T_n \;\Rightarrow\;\frac12\delta_0
    + \frac12\!\left(\frac{\Omega_{c_0}}{2H_{c_0}}\chi_1^2\right),
\end{equation}
where $\delta_0$ is a point mass at zero, and the standardised statistic $\widetilde T_n=(2H_{c_0}/\Omega_{c_0})\,T_n$ satisfies
\begin{equation}\label{eq:Tntilde_limit}
    \widetilde T_n \;\Rightarrow\;
    \frac12\delta_0+\frac12\chi_1^2.
\end{equation}
\end{theorem}

\begin{remark}[Choice of test principle]\label{rem:test_principle}
Three statistics suggest themselves for testing $c=c_0$: a Wald statistic based on $\widehat c_n-c_0$, studentised by $\Omega_{c_0}^{1/2}/H_{c_0}$; a score (Lagrange multiplier) statistic based on $Q_n'(c_0)$, studentised by $\Omega_{c_0}^{1/2}$; and the standardised criterion-difference statistic $\widetilde T_n$ above. Since $H_{c_0}$ and $\Omega_{c_0}$ are known exactly under the null, all three statistics are immediately operational, with no nuisance parameters to estimate. At an interior $c^\star$, these three coincide to first order, exactly as in the classical Wald--LR--LM trinity. This equivalence fails at the boundary. The score is evaluated at the fixed point $c_0$, before any constrained optimisation occurs, so $\sqrt{n}\,Q_n'(c_0)\Rightarrow\operatorname{N}(0,\Omega_{c_0})$ regardless of the constraint $c\geq c_0$: an LM test retains a standard null distribution. The Wald and criterion-difference statistics, by contrast, are both built from $\widehat c_n$ and inherit its censoring, giving the nonstandard limits $\max\{0,Z\}$ and $\tfrac12\delta_0+\tfrac12\chi_1^2$ respectively: the familiar finding of the literature on testing at the boundary of the parameter space \citep{chernoff54, self_liang87, andrews01}. I use standardised $\widetilde T_n$ rather than the regular LM alternative for two reasons. First, regularity of the null distribution says nothing about power at the boundary, where the local-power equivalence of Wald, LR, and LM at interior points need not carry over. Second, $\widetilde T_n$ is built from the same minimum-distance criterion $Q_n$ used throughout the paper, whereas an LM test requires a separate, score-based apparatus. I prefer $\widetilde T_n$ to the Wald alternative because $\widetilde T_n$ is the exact difference in criterion values at $c_0$ and $\widehat c_n$, whereas the Wald statistic relies on a single local quadratic approximation to $Q_n$ anchored at $c_0$; the criterion-difference statistic therefore tracks the finite-sample shape of $Q_n$ more closely, the standard rationale for preferring likelihood-ratio-type statistics over Wald statistics more generally.
\end{remark}

\begin{remark}[Null constants and critical values]\label{rem:critical_values}
The standardised statistic $\widetilde T_n$ is operational once $\nu$ has been chosen. The null constants $H_{c_0}$ and $\Omega_{c_0}$, given in \eqref{eq:Hc0} and \eqref{eq:Omega_c0}, are determined entirely by $\Phi$, $\dot G_{c_0}$, and $\nu$; they are not unknown population parameters requiring estimation. Both, moreover, reduce to closed-form expressions in the same special functions used to define the model itself: $\dot G_{c_0}$ is given explicitly by \eqref{eq:Gdot}--\eqref{eq:1F1dot} of Lemma~\ref{lem:differentiability} in terms of the Gamma and digamma functions and the $c$-derivative of ${}_1F_1$, while $\Phi$ itself is available exactly via the confluent hypergeometric representation \eqref{eq:normalcdf} of Lemma~\ref{lem:normalcdf}, with no numerical integration of the Gaussian density required. Evaluating $H_{c_0}$ and $\Omega_{c_0}$ at any point therefore requires only ${}_1F_1$, $\Gamma$, and $\psi$: the same three special functions on which the entire hypergeometric family is built; though the integrals defining $H_{c_0}$ and $\Omega_{c_0}$ over $\nu$ are one- and two-dimensional respectively, and generally still require numerical quadrature; what is avoided is any further numerical approximation of the integrands.

For the Gaussian-weighted choice $d\nu(z)=d\Phi(z)$: the $\lambda=0$ member of the tail-weighted family $\nu_\lambda$ in \eqref{eq:nu_lambda}, rather than the $\lambda=1/2$ choice used in the simulations and empirical illustration below (see Remark~\ref{rem:nu_choice}); writing $a(u)=\dot G_{c_0} \{\Phi^{-1}(u)\}$ for $u\in(0,1)$, the constants take the form
\[
    H_{c_0} = 2\int_0^1 a(u)^2\,du, \quad \Omega_{c_0}
    = 4\int_0^1\!\int_0^1 \{\min(u,v)-uv\}\,a(u)\,a(v)\,du\,dv,
\]
both of which are evaluable by numerical quadrature. For general $\nu_\lambda$, including the tail-weighted $\nu_{1/2}$ used in this paper, the analogous integrals can be computed over $\mathbb{R}$, or, after the same substitution, over the corresponding $\operatorname{Beta}(1-\lambda,1-\lambda)$ measure on $(0,1)$. Since $\nu_\lambda$ in \eqref{eq:nu_lambda} is defined only up to a positive constant, $H_{c_0}$ and $\Omega_{c_0}$ individually depend on that constant for $\lambda>0$; the estimator $\widehat c_n$, the asymptotic variance $\Omega_{c_0}/H_{c_0}^2$, and the standardised statistic $\widetilde T_n$ do not. The limiting null law $\frac12\delta_0+\frac12\chi_1^2$ yields a simple rejection rule. For an asymptotic level-$\alpha$ upper-tail test,
\[
    \Pr\!\left(\tfrac12\delta_0+\tfrac12\chi_1^2>k_\alpha\right)
    = \tfrac12\Pr(\chi_1^2>k_\alpha) = \alpha,
\]
so $k_\alpha=\chi_{1,1-2\alpha}^2$, the $(1-2\alpha)$-quantile of the $\chi_1^2$ distribution. The rejection rule is
\[
    \text{reject }H_0:F_0=\Phi
    \quad\Longleftrightarrow\quad
    \widetilde T_n > \chi_{1,1-2\alpha}^2.
\]
For $\alpha=0.05$, the critical value is $\chi_{1,0.90}^2\approx 2.706$; for $\alpha=0.01$, it is $\chi_{1,0.98}^2\approx 5.412$. These are strictly smaller than the standard $\chi_1^2$ critical values ($3.841$ and $6.635$).
\end{remark}

\noindent The moments of the limiting distribution are also available in closed form.
\begin{corollary}[Moments of $\widetilde T_\infty$]
\label{cor:moments_Ttilde}
Under the conditions of Theorem~\ref{thm:fit_improvement},
$\widetilde T_n\Rightarrow\widetilde T_\infty\sim\frac12
\delta_0+\frac12\chi_1^2$, with
\[
    \operatorname{E}[\widetilde T_\infty] = \frac12, \quad
    \operatorname{Var}(\widetilde T_\infty) = \frac54.
\]
\end{corollary}

\noindent The simulated mean and variance of $\widetilde T_n$ are reported alongside the limiting values $1/2$ and $5/4$ from Corollary~\ref{cor:moments_Ttilde} as an informal diagnostic of how closely the finite-sample distribution tracks the boundary asymptotics. Theorems~\ref{thm:boundary} and~\ref{thm:fit_improvement}, together with
Corollary~\ref{cor:moments_Ttilde}, complete the boundary inference theory: $\widetilde T_n$ has a known null limit, an explicit rejection rule, and null constants computable from $\nu$ and $\dot G_{c_0}$ alone. The next two results establish directed consistency against fixed alternatives and non-trivial local asymptotic power against $\sqrt{n}$-local alternatives.

\subsection{Directed consistency}\label{subsec:directed}

Theorems~\ref{thm:boundary} and~\ref{thm:fit_improvement} establish the null distribution of $\widetilde T_n$. The following theorem shows that the test is also consistent: under any fixed alternative for which the best hypergeometric approximation strictly improves on the Gaussian boundary, the statistic diverges to infinity and the test rejects with
probability tending to one.

\begin{theorem}[Directed consistency]\label{thm:directed}
Suppose the conditions of Theorem~\ref{thm:consistency} hold,
and that $\nu$ assigns positive mass to some open interval. Let
\begin{equation}\label{eq:Delta}
    \Delta = Q(c_0)-Q(c^\star)
\end{equation}
be the population fit improvement. If $\Delta>0$, then
$Q_n(c_0)-Q_n(\widehat c_n)\to_p\Delta$, and consequently
\[
    \widetilde T_n \;=\; \left(\frac{2H_{c_0}}{\Omega_{c_0}}\right)
    n\{Q_n(c_0)-Q_n(\widehat c_n)\} \;\to_p\; \infty.
\]
Hence $\Pr_{F_0}(\widetilde T_n>k_\alpha)\to 1$ for any fixed
critical value $k_\alpha<\infty$.
\end{theorem}

\begin{remark}[Directed, not omnibus]\label{rem:directed}
Consistency is directed at alternatives for which the best hypergeometric approximation strictly improves on the Gaussian boundary, that is, $\Delta>0$: the heavy-tailed deformation encoded by $\{G_c:c>3/2\}$. It need not detect every departure from normality: under a skewed alternative the symmetric family cannot capture, the minimum-distance projection $c^\star$ can remain at $c_0=3/2$, giving $\Delta=0$ and no guaranteed power. This is precisely the sense in which the test complements rather than replaces omnibus procedures such as Jarque--Bera.
\end{remark}

\noindent Theorem~\ref{thm:directed} establishes power against fixed alternatives; the next result characterises power against local alternatives approaching the null at rate $n^{-1/2}$.

\subsection{Local asymptotic power}\label{subsec:localpower}

Theorem~\ref{thm:directed} establishes power tending to one against fixed alternatives with $\Delta>0$. The complementary local result characterises power against $\sqrt{n}$-local alternatives. The key is the projection of the local departure from normality onto the score direction $\dot G_{c_0}$. Consider a triangular array with distribution functions
\begin{equation}\label{eq:local_alt}
    F_n(z) = \Phi(z) + \frac{\eta(z)}{\sqrt{n}} + r_n(z),
\end{equation}
where $\eta$ is the local departure from normality. The score functional and its variance under the local alternative are determined by the inner product
\[
    A_\nu(\eta) = \int_{\mathbb{R}}\eta(z)\,\dot G_{c_0}(z)\,d\nu(z),
\]
and the Brownian bridge covariance kernel $K(z,y)=\Phi(z\wedge y)-\Phi(z)\Phi(y)$.

\begin{theorem}[Local asymptotic power]\label{thm:local_power}
Suppose \textup{(B1)--(B2)} hold and $\nu$ assigns positive mass to some open interval, so that Theorem~\ref{thm:fit_improvement} applies at $F_0=\Phi$. Let $F_n$ satisfy \eqref{eq:local_alt} with
\[
    \int_{\mathbb{R}}|\eta(z)\,\dot G_{c_0}(z)|\,d\nu(z)<\infty,
\]
and $\sqrt{n}(\widehat F_n(z)-\Phi(z))\Rightarrow
\mathbb{B}_\Phi(z)+\eta(z)$ in $\ell^\infty(\mathbb{R})$,
where $\mathbb{B}_\Phi$ is a $\Phi$-Brownian bridge. Then
\begin{equation}\label{eq:Tn_local}
    \widetilde T_n\;\Rightarrow\;
    (Z-\delta_\nu(\eta))^2\,\mathbf{1}\{Z-\delta_\nu(\eta)<0\},
    \quad Z\sim\operatorname{N}(0,1),
\end{equation}
where the local power index is
\begin{equation}\label{eq:delta}
    \delta_\nu(\eta) = \frac{2A_\nu(\eta)}{\sqrt{\Omega_{c_0(\nu)}}}.
\end{equation}
The asymptotic rejection probability of the level-$\alpha$ test is
\begin{equation}\label{eq:power}
    \pi_\nu(\eta) =
    \Phi\bigl\{\delta_\nu(\eta)-\Phi^{-1}(1-\alpha)\bigr\}.
\end{equation}
\end{theorem}

\begin{remark}[Directed local sensitivity]
\label{rem:local_power}
The functional limit $\sqrt{n}(\widehat F_n-\Phi)\Rightarrow \mathbb{B}_\Phi+\eta$ is maintained as the standard local empirical process limit for a triangular array drifting toward $\Phi$ at rate $n^{-1/2}$ \citep[cf.][Thm.~3.10.12] {vandervaart_wellner23}, rather than re-derived here; setting $\eta\equiv0$ recovers Donsker's theorem under the fixed null (Theorem~\ref{thm:fit_improvement}). The local power index
$\delta_\nu(\eta)=2A_\nu(\eta)/\sqrt{\Omega_{c_0}(\nu)}$ measures the projection of $\eta$ onto the score direction $\dot G_{c_0}$, standardised by the null standard deviation. Alternatives with positive projection have power above size; alternatives orthogonal to $\dot G_{c_0}$, including skewed departures since $\dot G_{c_0}$ is odd, have $\delta_\nu(\eta)=0$ and power equal to size asymptotically. This is the local counterpart of Remark~\ref{rem:directed}: the test is locally sensitive only to departures aligned with the hypergeometric score direction.
\end{remark}

\begin{corollary}[Local hypergeometric alternatives]\label{cor:local_hgm}
Under the local hypergeometric alternatives
$F_n(z)=G_{c_0+\tau/\sqrt{n}}(z)$ with $\tau\geq 0$,
\[
    \delta_\nu(\tau) =
    \frac{\tau H_{c_0}(\nu)}{\sqrt{\Omega_{c_0}(\nu)}},
\]
and the asymptotic rejection probability is
\[
    \pi_\nu(\tau) = \Phi\!\left\{
    \frac{\tau H_{c_0}(\nu)}{\sqrt{\Omega_{c_0}(\nu)}}
    -\Phi^{-1}(1-\alpha)\right\}.
\]
\end{corollary}

\noindent Theorems~\ref{thm:boundary}, \ref{thm:fit_improvement}, \ref{thm:directed} and \ref{thm:local_power} complete the asymptotic theory for the fixed-scale test. The next section develops plug-in location-scale versions.

\section{Plug-in location-scale procedures}\label{sec:plugin}

The fixed-scale test assumes known location and scale. In practice one tests the composite Gaussian null $H_0:F_0\in \{\operatorname{N}(\mu,\sigma^2):\mu\in\mathbb{R},\sigma>0\}$ by standardising with estimated location and scale before applying the fixed-scale procedure. Two standardisations are considered. The first uses the sample mean and standard deviation, which are natural under the Gaussian null. The second uses the sample median and normal-calibrated IQR, which is motivated by the infinite variance of $G_c$ for $c>3/2$, under which the sample standard deviation is an unstable scale estimator and median/IQR standardisation preserves the tail signal more effectively. In both cases, the fixed-scale theory extends immediately: consistency, interior asymptotics, boundary inference, and directed consistency all carry over with $\widehat F_n$ replaced by the standardised empirical distribution. The only structural change is the projected covariance kernel, which replaces the Brownian bridge kernel $K(z,y)=\Phi(z\wedge y)-\Phi(z)\Phi(y)$ to account for estimation of location and scale. The curvature $H_{c_0}$, score direction $\dot G_{c_0}$, null limit $\frac12\delta_0+\frac12\chi_1^2$, and rejection rule $\widetilde T_n^\diamond>\chi_{1,1-2\alpha}^2$ are all unchanged.

\subsection{Location-scale standardisation and plug-in
statistics}\label{subsec:plugin_stats}

Let $X_1,\ldots,X_n$ be the observed sample. I consider two location-scale standardisations, denoted by superscript $\diamond\in\{P,R\}$: mean/standard-deviation ($P$) and median/IQR ($R$).\\
\\
\textit{Mean/standard-deviation ($P$).} Define
\[
    \widehat\mu_n^P = \bar X_n, \quad (\widehat\sigma_n^P)^2 =
    \frac{1}{n}\sum_{i=1}^n(X_i-\bar X_n)^2,
\]
and let $Y_i^P=(X_i-\widehat\mu_n^P)/\widehat\sigma_n^P$, with empirical distribution $\widehat F_n^P$.\\
\\
\textit{Median/IQR ($R$).} Let $\widehat q_n(p)$ denote the $p$-th sample quantile and $a=\Phi^{-1}(3/4)$. Let
\begin{equation}\label{eq:median_iqr}
    \widehat\mu_n^R = \widehat q_n(1/2), \quad
    \widehat\sigma_n^R =
    \frac{\widehat q_n(3/4)-\widehat q_n(1/4)}{2a},
\end{equation}
where $2a$ is the interquartile range of $\Phi$, used to calibrate $\widehat\sigma_n^R$. Let $Y_i^R=(X_i-\widehat\mu_n^R)/\widehat\sigma_n^R$, with empirical distribution $\widehat F_n^R$. For each standardisation $\diamond\in\{P,R\}$, define the plug-in criterion, estimator, and fit-improvement statistic
\[
    Q_n^\diamond(c) = \int_{\mathbb{R}}
    \{\widehat F_n^\diamond(z)-G_c(z)\}^2\,d\nu(z),
    \quad c\in\mathcal{C},
\]
\[
    \widehat c_n^\diamond \in
    \arg\min_{c\in\mathcal{C}}Q_n^\diamond(c), \quad
    T_n^\diamond = n\{Q_n^\diamond(c_0)
    -Q_n^\diamond(\widehat c_n^\diamond)\}.
\]
\noindent Under the Gaussian null, $\widehat\mu_n^P\to\mu_0$ and $\widehat\sigma_n^P\to\sigma_0$ almost surely by the strong law of large numbers, since $\mathrm{E}[X_i^2] <\infty$; similarly $\widehat\mu_n^R\to\mu_0$ and $\widehat\sigma_n^R\to\sigma_0$ almost surely by the consistency of sample quantiles at the population quantiles
corresponding to probabilities $1/4$, $1/2$, and $3/4$, where the Gaussian
density is continuous and positive. The two conditions are not equivalent: the mean/SD case requires a finite second moment, while the median/IQR case requires only local density regularity, with no moment condition.

\subsection{Asymptotic theory}\label{subsec:plugin_asymptotics}

Under the Gaussian null, both standardisations yield a mean-zero
Gaussian empirical-process limit. The only change relative to the
fixed-scale theory is that the Brownian-bridge covariance kernel is
replaced by the projected kernel induced by the chosen location-scale
standardisation. This projection accounts for the first-order effect of estimating location and scale and leads to the same boundary limit after using the corresponding plug-in score variance. The next two lemmas give the projected empirical processes for the mean/standard-deviation and median/IQR standardisations, respectively.

\begin{lemma}[Projected empirical process: mean/SD]\label{lem:plugin_ep}
Suppose $X_i\stackrel{\text{i.i.d.}}{\sim}\operatorname{N}(\mu_0,\sigma_0^2)$.
Then, in $\ell^\infty(\mathbb{R})$,
\begin{equation}\label{eq:ZP}
    \sqrt{n}\{\widehat F_n^P(z)-\Phi(z)\}
    \;\Rightarrow\;
    \mathbb{Z}_P(z),
\end{equation}
where $\mathbb{Z}_P$ is the mean-zero Gaussian process admitting the
representation
\[
    \mathbb{Z}_P(z) = \mathbb{B}_\Phi(z) + \phi(z)Z_1 + z\phi(z)Z_2.
\]
Here $\mathbb{B}_\Phi$ is a $\Phi$-Brownian bridge and
$(\mathbb{B}_\Phi,Z_1,Z_2)$ is jointly Gaussian with
\[
    Z_1\sim\operatorname{N}(0,1),\quad Z_2\sim\operatorname{N}(0,1/2),\quad \operatorname{Cov}(Z_1,Z_2)=0,
\]
and cross-covariances
\[
    \operatorname{Cov}(\mathbb{B}_\Phi(z),Z_1)=-\phi(z),\quad
    \operatorname{Cov}(\mathbb{B}_\Phi(z),Z_2)=-\frac12 z\phi(z).
\]
Consequently, the covariance kernel of $\mathbb{Z}_P$ is
\begin{equation}\label{eq:KP}
    K_P(z,y) = \Phi(z\wedge y)-\Phi(z)\Phi(y) -\phi(z)\phi(y)
    -\frac12 zy\,\phi(z)\phi(y).
\end{equation}
\end{lemma}

\noindent The two correction terms in $K_P$ reflect the components of the empirical process absorbed by estimating $\mu$ and $\sigma$ respectively.

\begin{lemma}[Projected empirical process: median/IQR]\label{lem:robust_ep}
Suppose $X_i\stackrel{\text{i.i.d.}}{\sim}\operatorname{N}(\mu_0,\sigma_0^2)$.
Let $a=\Phi^{-1}(3/4)$ and define the influence functions
\[
    \operatorname{IF}_m(w) = \frac{1/2-\mathbf{1}\{w\leq 0\}}{\phi(0)}, \quad \operatorname{IF}_s(w)
    =
    \frac{1}{2a}\left\{\frac{3/4-\mathbf{1}\{w\leq a\}}{\phi(a)} -
    \frac{1/4-\mathbf{1}\{w\leq -a\}}{\phi(a)}\right\}.
\]
Then, in $\ell^\infty(\mathbb{R})$,
\[
    \sqrt{n}\{\widehat F_n^R(z)-\Phi(z)\}\;\Rightarrow\;\mathbb{Z}_R(z),
\]
where $\mathbb{Z}_R$ is the mean-zero Gaussian process admitting the
representation
\[
    \mathbb{Z}_R(z) = \mathbb{B}_\Phi(z) + \phi(z)Z_m + z\phi(z)Z_s.
\]
Here $\mathbb{B}_\Phi$ is a $\Phi$-Brownian bridge and $(\mathbb{B}_\Phi,Z_m,Z_s)$ is jointly Gaussian, where
\[
    Z_m\sim \operatorname{N}(0,\operatorname{E}[\operatorname{IF}_m(W)^2]), \quad
    Z_s\sim \operatorname{N}(0,\operatorname{E}[\operatorname{IF}_s(W)^2]), \quad
    \operatorname{Cov}(Z_m,Z_s)=\operatorname{E}[\operatorname{IF}_m(W)\operatorname{IF}_s(W)],
\]
with $W\sim\operatorname{N}(0,1)$, and
\[
    \operatorname{Cov}(\mathbb{B}_\Phi(z),Z_m) =
    \operatorname{E}[(\mathbf{1}\{W\leq z\}-\Phi(z))\operatorname{IF}_m(W)],
\]
\[
    \operatorname{Cov}(\mathbb{B}_\Phi(z),Z_s) =
    \operatorname{E}[(\mathbf{1}\{W\leq z\}-\Phi(z))\operatorname{IF}_s(W)].
\]
Consequently, the covariance kernel of $\mathbb{Z}_R$ is
\begin{equation}\label{eq:KR}
    K_R(z,y) = \operatorname{E}[\xi_z(W)\xi_y(W)],
\end{equation}
where
\[
    \xi_z(w) = \mathbf{1}\{w\leq z\}-\Phi(z) + \phi(z)\operatorname{IF}_m(w) +
    z\phi(z)\operatorname{IF}_s(w).
\]
\end{lemma}

\begin{remark}[Modularity of the plug-in correction]
\label{rem:plugin_modularity}
The plug-in score variance for standardisation
$\diamond\in\{P,R\}$ is
\begin{equation}\label{eq:plugin_Omega}
    \Omega_{c_0}^\diamond = 4\iint_{\mathbb{R}^2}
    K_\diamond(z,y)\,\dot G_{c_0}(z)\,\dot G_{c_0}(y)\,d\nu(z)\,d\nu(y).
\end{equation}
Relative to the fixed-scale case, the minimum-distance criterion, the boundary argument, and the limiting one-sided quadratic form are unchanged. What changes is the Gaussian empirical-process kernel: the Brownian-bridge kernel is replaced by the projected kernel $K_\diamond$, and consequently the score variance $\Omega_{c_0}$ is replaced by $\Omega_{c_0}^\diamond$. Intuitively, estimating location and scale absorbs from the empirical process the components that look, to first order, like small location and scale perturbations of the null distribution. For mean/SD standardisation this gives the usual normal location-scale projection. For median/IQR standardisation the same mechanism applies, but the projection is driven by the quantile influence functions instead. This replacement has no effect on the form of the boundary limit after standardisation: the normalising constant in the statistic is simply computed with $K_\diamond$ through \eqref{eq:plugin_Omega}. In numerical implementation, the same quadrature scheme used in the fixed-scale case can be used after replacing the kernel. This is the familiar phenomenon in empirical-process goodness-of-fit theory that nuisance-parameter estimation changes the covariance structure of the limiting process, but not the basic form of the resulting projected test statistic.
\end{remark}

\begin{theorem}[Plug-in fit-improvement limit]
\label{thm:plugin_fit_improvement}
Suppose the Gaussian location-scale null holds and the conditions of Theorem~\ref{thm:fit_improvement} hold with $\widehat F_n$ replaced by $\widehat F_n^\diamond$. Then
\[
    \widetilde T_n^\diamond =
    \frac{2H_{c_0}}{\Omega_{c_0}^\diamond}
    \,n\{Q_n^\diamond(c_0)-Q_n^\diamond(\widehat c_n^\diamond)\}
    \;\Rightarrow\; \tfrac12\delta_0+\tfrac12\chi_1^2.
\]
An asymptotic level-$\alpha$ test rejects the Gaussian location-scale null when $\widetilde T_n^\diamond>\chi_{1,1-2\alpha}^2$.
\end{theorem}

\noindent The same substitution principle applies to the remaining minimum-distance results. Consistency of $\widehat c_n^\diamond$, interior asymptotic normality with the corresponding projected score variance, and directed consistency against alternatives satisfying
$\Delta^\diamond=Q^\diamond(c_0)-Q^\diamond(c_\diamond^\star)>0$ follow from the arguments of Theorems~\ref{thm:consistency}, \ref{thm:interior}, and \ref{thm:directed}, with the fixed-scale empirical-process variance replaced by the appropriate projected variance $\Omega^\diamond$.

\begin{corollary}[Moments]\label{cor:plugin_moments}
Under the conditions of Theorem~\ref{thm:plugin_fit_improvement}, $\operatorname{E}[\widetilde T_\infty^\diamond]=1/2$ and $\operatorname{Var}(\widetilde T_\infty^\diamond)=5/4$ for each $\diamond\in\{P,R\}$, by the same calculation as Corollary~\ref{cor:moments_Ttilde}.
\end{corollary}

\begin{corollary}[Plug-in local power]\label{cor:plugin_local_power}
Under local alternatives $F_n^\diamond(z)=\Phi(z)+\eta_\diamond(z)/
\sqrt{n} + o(n^{-1/2})$ satisfying the conditions of
Theorem~\ref{thm:local_power} with $K$ replaced by $K_\diamond$,
\[
    \pi_\nu^\diamond(\eta_\diamond) = \Phi\!\left\{
    \delta_\nu^\diamond(\eta_\diamond)-\Phi^{-1}(1-\alpha)
    \right\}, \quad \delta_\nu^\diamond(\eta_\diamond) =
    \frac{2\int_\mathbb{R} \eta_\diamond(z)\,\dot G_{c_0}(z)\,d\nu(z)}
    {\sqrt{\Omega_{c_0}^\diamond(\nu)}}.
\]
\end{corollary}

\begin{corollary}[Plug-in consistency]\label{cor:plugin_consistency} For $\diamond\in\{P,R\}$, let $F_0^\diamond$ denote the distribution of $(X-\mu_0^\diamond)/\sigma_0^\diamond$, where $(\mu_0^\diamond,\sigma_0^\diamond)$ are the population limits of the corresponding standardising estimators, with $\sigma_0^\diamond>0$. Define \[ Q^\diamond(c) = \int_{\mathbb R}\{F_0^\diamond(z)-G_c(z)\}^2\,d\nu(z), \] and suppose that $Q^\diamond$ has a unique minimiser $c_\diamond^\star\in\mathcal C$. Then $\widehat c_n^\diamond\to c_\diamond^\star$ almost surely. In particular, under the Gaussian location-scale null, $F_0^\diamond=G_{c_0}=\Phi$ for both $\diamond\in\{P,R\}$, and hence $\widehat c_n^\diamond\to c_0$ almost surely.
\end{corollary}

\begin{corollary}[Plug-in directed consistency]\label{cor:plugin_directed} For $\diamond\in\{P,R\}$, let $\Delta^\diamond = Q^\diamond(c_0)-Q^\diamond(c_\diamond^\star)$, where $c_\diamond^\star$ is the unique minimiser of $Q^\diamond$ over $\mathcal C$. If $\Delta^\diamond>0$, then $ n\{Q_n^\diamond(c_0)-Q_n^\diamond(\widehat c_n^\diamond)\} \to_p \infty$. Consequently, $\Pr_{F_0}(\widetilde T_n^\diamond>k_\alpha)\to1$ for any fixed $k_\alpha<\infty$.
\end{corollary}

\begin{remark}[Plug-in versus joint estimation]\label{rem:plugin_vs_joint}
An alternative to the plug-in approach is to estimate $\theta=(\mu,\sigma,c)$ jointly by minimising
\[
    (\widehat\mu_n,\widehat\sigma_n,\widehat c_n)
    \in
    \arg\min_{\mu,\sigma,c}
    \int_{\mathbb{R}}\left\{
    \widehat F_n(z)
    -G_c\!\left(\frac{z-\mu}{\sigma}\right)
    \right\}^2 d\nu(z).
\]
The joint theory is developed in Appendix~\ref{app:joint}. I prefer the plug-in approach for three reasons. First, the criterion has a nearly degenerate score-variance Schur complement in the $c$-direction after optimising over $(\mu,\sigma)$: under the Gaussian null, the $c$-score direction $\dot G_{c_0}$ nearly lies in the span of $\{\phi(z),z\phi(z)\}$ in the Brownian bridge covariance metric, making the residual score variance
$\Omega_{cc\cdot\lambda}\approx0.000013$ very small
(Table~\ref{tab:null_size_tail_scaling_robust}). The curvature Schur complement $H_{cc\cdot\lambda}\approx0.0023$ is not similarly reduced, giving a normalisation constant $2H_{cc\cdot\lambda}/\Omega_{cc\cdot\lambda}
\approx355$ that amplifies any small spurious criterion improvement into a large test statistic; since this is structural, the size distortion worsens with $n$. Simulation evidence confirms the severity: the $10\%$ rejection rate rises from $0.089$ at $n=200$ to $0.309$ at $n=1600$, with the boundary rate falling away from the theoretical $1/2$
(Table~\ref{tab:null_size_tail_exact_profile_appendix}). Second, the plug-in approach preserves the modular structure of the theory: the limit distribution, curvature, and rejection rule are invariant across all three versions of the test; only the projected covariance kernel must be recomputed. Third, the two-step structure allows the location-scale estimator to be adapted to the operating conditions: median/IQR standardisation remains valid when the variance is infinite ($c>3/2$), while mean/SD standardisation does not.
\end{remark}

\section{Finite-sample performance}\label{sec:simulations}

The simulations examine the finite-sample size and power of three hypergeometric (HGM) statistics: the fixed-scale statistic $\widetilde T_n$, the mean/SD plug-in statistic $\widetilde T_n^P$, and the median/IQR plug-in statistic $\widetilde T_n^R$. All use the tail-weighted criterion $d\nu_{1/2}(z)\propto [\Phi(z)\{1-\Phi(z)\}]^{-1/2}\,d\Phi(z)$, and asymptotic critical values from $\frac12\delta_0+\frac12\chi_1^2$. The criterion is evaluated on the Gaussian quantile grid \eqref{eq:grid} with $J=301$ nodes and normalised weights \eqref{eq:weights}. The null constants $H_{c_0}$, $\Omega_{c_0}$, $\Omega_{c_0}^P$, and $\Omega_{c_0}^R$ are computed once and reported in Table~\ref{tab:null_size_tail_scaling_robust}. The parameter space is $\mathcal C=[3/2,8]$, searched over a fine grid concentrated near the boundary. Unless stated otherwise, size results are based on $R=10^5$ replications and power results on $R=10^4$ replications.

\subsection{Null size}\label{subsec:sim_size}

Table~\ref{tab:null_size_tail_main_robust} reports empirical rejection
frequencies under the Gaussian null $F_0=\Phi$ for
$n\in\{25,50,100,200,400,800,1600\}$ and nominal levels
$10\%$, $5\%$, and $1\%$. The boundary frequency
$\Pr(\widehat c=c_0)$ and the mean of the standardised statistic provide
diagnostics for the boundary limit, whose corresponding values are
$1/2$ and $1/2$. The fixed-scale statistic is accurately sized throughout the design. At the $5\%$ level, rejection frequencies range from $0.050$ to $0.052$, and both diagnostic columns are close to their limiting values even for $n=25$. Thus the fixed-scale boundary approximation is reliable in quite small samples.

The mean/SD plug-in statistic converges more slowly but is well calibrated by moderate sample sizes. It mildly over-rejects at the $10\%$ level for small $n$ and is conservative at the $1\%$ level for small $n$, but these distortions diminish as $n$ increases. The slower convergence of the boundary frequency reflects the additional first-order effect of estimating location and scale, which is accounted for asymptotically by the projected kernel $K_P$. The median/IQR plug-in statistic is the most demanding finite-sample case. It is liberal for small samples, with $5\%$ size equal to $0.095$ at $n=25$, and remains mildly oversized at the largest sample sizes reported. This is consistent with the greater sampling variability of quantile-based standardisation. In practice, the robust plug-in test should be used with caution at $n<200$, although its size distortion declines substantially as $n$ grows. Appendix~\ref{app:simulations} reports an additional check: under $\operatorname{N}(3,4)$ the robust plug-in rejection frequencies match those under $\operatorname{N}(0,1)$ up to simulation error, confirming location-scale invariance (Table~\ref{tab:null_size_tail_robust_plugin_invariance_N3_2sq}). Overall, the simulations support use of the boundary-mixture critical values for the fixed-scale and mean/SD plug-in statistics, and show that the robust median/IQR version requires larger samples for accurate null calibration.

\begin{table}[!p]
\centering
\caption{Null size of the HGM tests with tail-weighted criterion}
\label{tab:null_size_tail_main_robust}
\begin{tabular}{rcccccc}
\toprule
$n$ & $10\%$ & $5\%$ & $1\%$ & $\Pr(\widehat c_n=c_0)$ & mean $\widehat c_n$ & mean $\widetilde T_n$ \\
\midrule
\multicolumn{7}{l}{Panel A: fixed-scale $G_c$ test of $\Phi$} \\
25 & 0.100 & 0.052 & 0.011 & 0.511 & 1.5898 & 0.503 \\
50 & 0.100 & 0.051 & 0.010 & 0.505 & 1.5603 & 0.500 \\
100 & 0.100 & 0.051 & 0.011 & 0.506 & 1.5413 & 0.506 \\
200 & 0.101 & 0.051 & 0.011 & 0.507 & 1.5284 & 0.503 \\
400 & 0.101 & 0.050 & 0.010 & 0.502 & 1.5198 & 0.502 \\
800 & 0.100 & 0.050 & 0.010 & 0.502 & 1.5138 & 0.502 \\
1600 & 0.100 & 0.050 & 0.010 & 0.504 & 1.5096 & 0.499 \\
\addlinespace
\multicolumn{7}{l}{Panel B: mean/standard-deviation plug-in location-scale test} \\
25 & 0.116 & 0.043 & 0.002 & 0.373 & 1.5443 & 0.565 \\
50 & 0.114 & 0.047 & 0.004 & 0.412 & 1.5290 & 0.546 \\
100 & 0.112 & 0.050 & 0.006 & 0.442 & 1.5194 & 0.537 \\
200 & 0.109 & 0.050 & 0.008 & 0.457 & 1.5133 & 0.531 \\
400 & 0.106 & 0.050 & 0.008 & 0.473 & 1.5090 & 0.517 \\
800 & 0.104 & 0.050 & 0.009 & 0.483 & 1.5062 & 0.514 \\
1600 & 0.103 & 0.050 & 0.009 & 0.491 & 1.5043 & 0.509 \\
\addlinespace
\multicolumn{7}{l}{Panel C: robust median/IQR plug-in location-scale test} \\
25 & 0.161 & 0.095 & 0.029 & 0.412 & 1.6145 & 0.815 \\
50 & 0.131 & 0.073 & 0.020 & 0.438 & 1.5659 & 0.668 \\
100 & 0.126 & 0.068 & 0.017 & 0.455 & 1.5429 & 0.632 \\
200 & 0.119 & 0.065 & 0.016 & 0.472 & 1.5282 & 0.600 \\
400 & 0.114 & 0.060 & 0.014 & 0.477 & 1.5190 & 0.569 \\
800 & 0.111 & 0.058 & 0.014 & 0.484 & 1.5130 & 0.558 \\
1600 & 0.107 & 0.057 & 0.013 & 0.492 & 1.5090 & 0.543 \\
\bottomrule
\end{tabular}
\begin{minipage}{0.95\linewidth}
\footnotesize
Notes: The table reports empirical rejection frequencies under the Gaussian null $F_0=\Phi$. All tests use the tail-weighted measure $d\nu_{1/2}(z)\propto[\Phi(z)\{1-\Phi(z)\}]^{-1/2}d\Phi(z)$. Critical values are from the boundary limit $\frac12\delta_0+\frac12\chi_1^2$. The mean--standard-deviation plug-in statistic uses the projected-process variance $\Omega_{c_0}^P(\nu_{1/2})$; the robust plug-in statistic uses the median/IQR projected-process variance $\Omega_{c_0}^R(\nu_{1/2})$. The boundary column reports $\Pr(\widehat c_n\le c_0+10^{-4})$ with $c_0=3/2$.
\end{minipage}
\end{table}

\subsection{Power}\label{subsec:sim_power}

The test is directed: by Theorem~\ref{thm:local_power}, local power is determined by the projection of the departure onto the hypergeometric score direction $\dot G_{c_0}$, giving power against symmetric heavy-tailed alternatives and limited power against skewed alternatives by construction. Three tests are compared: the fixed-scale $G_c$ test, the robust plug-in $\widetilde T_n^R$, and the omnibus benchmarks Jarque--Bera (JB), Anderson--Darling (AD), and Shapiro--Wilk (SW). Mean/SD plug-in and size-adjusted results are in Appendix~\ref{app:simulations}, where the mean/SD plug-in is shown to fail completely under HGM alternatives with infinite variance.

\subsubsection{HGM alternatives}

Table~\ref{tab:power_hgm_alternatives_revised} and Figure~\ref{fig:power_curves_robust} (upper panel) report rejection frequencies against HGM alternatives $G_c$ with $c\in\{1.55,1.60,1.75,2.00,3.00\}$ at nominal level $5\%$. The fixed-scale test $\widetilde T_n$ has strong power against HGM alternatives, as the theory predicts. At $c=2.00$ power reaches $0.590$ at $n=25$ and $0.967$ by $n=100$; at $c=3.00$ it reaches $0.964$ at $n=25$. This confirms the directed consistency of Theorem~\ref{thm:directed} in finite samples. For the closest alternative, $c=1.55$, the local-power approximation in Corollary~\ref{cor:local_hgm}, $\pi_n(c)
\approx \Phi\left\{ 1.052 \sqrt{n}\left(c-\frac32\right)-1.645
\right\}$, is also quantitatively accurate: it predicts rejection frequencies of $0.132$, $0.184$, $0.277$, and $0.438$ for $n=100,200,400$, and $800$, respectively, compared with the simulated frequencies $0.135$, $0.186$, $0.275$, and $0.413$. The robust plug-in test $\widetilde T_n^R$ has substantially lower power. At $c=1.60$ it reaches only $0.167$ at $n=800$; at $c=3.00$ it reaches $0.986$ at $n=800$. The gap relative to the fixed-scale test reflects the loss of signal from the median/IQR standardisation: the projected kernel $K_R$ absorbs part of the HGM tail signal together with the location-scale nuisance, as Corollary~\ref{cor:plugin_local_power} predicts. The omnibus benchmarks substantially outperform the robust plug-in at close alternatives: JB reaches $0.931$ at $c=1.60$, $n=400$, against $0.129$ for $\widetilde T_n^R$, reflecting the high finite-sample kurtosis of HGM alternatives with $c>3/2$, to which JB's kurtosis component is directly sensitive; $\widehat c$ by contrast provides a fitted distribution with explicit tail-risk implications rather than a rejection decision.

\subsubsection{Standard non-HGM alternatives}

Table~\ref{tab:power_standard_alternatives_revised} and Figure~\ref{fig:power_curves_robust} (lower panel) report rejection frequencies against six alternatives: contaminated normal, Laplace, chi-squared $\chi^2_3$, $t_3$, $t_5$, and $t_{10}$. The fixed-scale test has essentially zero power against every non-HGM alternative. Against $t_3$ it
never exceeds zero. The standardised $t_3$ distribution has density tail index $4$ (survival $\propto|z|^{-3}$), which is lighter than the HGM tail-index-2 direction (density $\propto|z|^{-3}$, survival $\propto|z|^{-2}$) targeted by the test: the best HGM approximation to $t_3$ data is therefore at the Gaussian boundary $c_0$, giving $\Delta=0$ and no directed power. Against the chi-squared, power is essentially zero because the departure is in the skewness direction, which is orthogonal to $\dot G_{c_0}$ by the symmetry of the
HGM family.

The robust plug-in has useful power against symmetric heavy-tailed alternatives. Against Laplace it reaches $0.272$ at $n=25$ and $0.999$ at $n=800$; against $t_3$ it reaches $0.195$ at $n=25$ and $0.977$ at $n=800$. Power against the contaminated normal and $t_{10}$ is moderate, reflecting their relative proximity to the Gaussian. Against the chi-squared, power falls to zero at large $n$, consistent with Theorem~\ref{thm:local_power}: skewed departures are orthogonal to the even function $\dot G_{c_0}$. The omnibus benchmarks dominate the robust plug-in against non-HGM alternatives: JB, AD, and SW have substantially higher power against the contaminated normal, chi-squared, and $t_5$. At $n=25$ the robust plug-in reaches $0.122$ against $t_{10}$, slightly above JB ($0.100$) and AD ($0.080$). This reflects the local sensitivity of directed tests near their targeted direction; after size adjustment the advantage disappears (Table~\ref{tab:appendix_sizeadj_robust_standard_power}). The HGM test complements rather than competes with omnibus procedures: it isolates the symmetric heavy-tailed component of non-Gaussianity and provides a fitted parameter $\widehat c$ with direct tail-risk interpretation.

\begin{table}[!p]
\centering
\caption{Power against HGM alternatives}
\label{tab:power_hgm_alternatives_revised}
\begin{tabular}{llcccccc}
\toprule
Alternative & Test & $n=25$ & $n=50$ & $n=100$ & $n=200$ & $n=400$ & $n=800$ \\
\midrule
$c=1.55$ & Fixed scale & 0.088 & 0.108 & 0.135 & 0.186 & 0.275 & 0.413 \\
 & Robust plug-in & 0.104 & 0.078 & 0.083 & 0.078 & 0.087 & 0.099 \\
 & JB & 0.117 & 0.207 & 0.335 & 0.535 & 0.745 & 0.922 \\
 & AD & 0.103 & 0.143 & 0.200 & 0.308 & 0.465 & 0.684 \\
 & SW & 0.129 & 0.199 & 0.303 & 0.484 & 0.700 & 0.899 \\
\addlinespace
$c=1.60$ & Fixed scale & 0.146 & 0.187 & 0.262 & 0.398 & 0.618 & 0.855 \\
 & Robust plug-in & 0.107 & 0.091 & 0.100 & 0.107 & 0.129 & 0.167 \\
 & JB & 0.184 & 0.329 & 0.518 & 0.745 & 0.931 & 0.995 \\
 & AD & 0.153 & 0.236 & 0.344 & 0.527 & 0.759 & 0.943 \\
 & SW & 0.188 & 0.309 & 0.474 & 0.702 & 0.906 & 0.992 \\
\addlinespace
$c=1.75$ & Fixed scale & 0.319 & 0.467 & 0.694 & 0.910 & 0.993 & 1.000 \\
 & Robust plug-in & 0.125 & 0.124 & 0.144 & 0.178 & 0.258 & 0.390 \\
 & JB & 0.308 & 0.533 & 0.778 & 0.944 & 0.997 & 1.000 \\
 & AD & 0.259 & 0.408 & 0.622 & 0.841 & 0.977 & 1.000 \\
 & SW & 0.307 & 0.501 & 0.736 & 0.920 & 0.995 & 1.000 \\
\addlinespace
$c=2.00$ & Fixed scale & 0.590 & 0.814 & 0.967 & 1.000 & 1.000 & 1.000 \\
 & Robust plug-in & 0.157 & 0.159 & 0.215 & 0.321 & 0.486 & 0.730 \\
 & JB & 0.419 & 0.682 & 0.902 & 0.990 & 1.000 & 1.000 \\
 & AD & 0.364 & 0.572 & 0.810 & 0.967 & 0.999 & 1.000 \\
 & SW & 0.417 & 0.648 & 0.879 & 0.985 & 1.000 & 1.000 \\
\addlinespace
$c=3.00$ & Fixed scale & 0.964 & 0.999 & 1.000 & 1.000 & 1.000 & 1.000 \\
 & Robust plug-in & 0.199 & 0.266 & 0.395 & 0.603 & 0.859 & 0.986 \\
 & JB & 0.524 & 0.807 & 0.967 & 0.999 & 1.000 & 1.000 \\
 & AD & 0.495 & 0.749 & 0.942 & 0.997 & 1.000 & 1.000 \\
 & SW & 0.535 & 0.790 & 0.959 & 0.999 & 1.000 & 1.000 \\
\addlinespace
\bottomrule
\end{tabular}
\begin{minipage}{0.95\linewidth}
\footnotesize
Notes: Entries are rejection frequencies at nominal level $5\%$ based on $R=10^4$ replications, against HGM alternatives $G_c: c \in \{1.55, 1.60, 1.75, 2.00, 3.00\}$. The reported HGM procedures are the fixed-scale test $\widetilde{T}_n$ and the robust median/IQR plug-in test $\widetilde{T}_n^R$; both tests use the tail-weighted criterion $d\nu_{1/2}(z)\propto[\Phi(z)\{1-\Phi(z)\}]^{-1/2}d\Phi(z)$ and asymptotic boundary critical values. Benchmark tests are Jarque–Bera (JB), Anderson–Darling (AD), and Shapiro–Wilk (SW).
\end{minipage}
\end{table}

\begin{table}[!p]
\centering
\caption{Power against standard non-HGM alternatives}
\label{tab:power_standard_alternatives_revised}
\begin{tabular}{llcccccc}
\toprule
Alternative & Test & $n=25$ & $n=50$ & $n=100$ & $n=200$ & $n=400$ & $n=800$ \\
\midrule
$\chi^2_3$ & Fixed scale & 0.008 & 0.004 & 0.001 & 0.000 & 0.000 & 0.000 \\
 & Robust plug-in & 0.041 & 0.018 & 0.008 & 0.004 & 0.000 & 0.000 \\
 & JB & 0.466 & 0.860 & 1.000 & 1.000 & 1.000 & 1.000 \\
 & AD & 0.675 & 0.958 & 1.000 & 1.000 & 1.000 & 1.000 \\
 & SW & 0.782 & 0.987 & 1.000 & 1.000 & 1.000 & 1.000 \\
\addlinespace
$t_{10}$ & Fixed scale & 0.031 & 0.024 & 0.014 & 0.008 & 0.004 & 0.000 \\
 & Robust plug-in & 0.122 & 0.109 & 0.123 & 0.155 & 0.206 & 0.303 \\
 & JB & 0.100 & 0.172 & 0.293 & 0.443 & 0.667 & 0.894 \\
 & AD & 0.080 & 0.098 & 0.148 & 0.216 & 0.373 & 0.660 \\
 & SW & 0.109 & 0.152 & 0.233 & 0.357 & 0.569 & 0.834 \\
\addlinespace
$t_{3}$ & Fixed scale & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\
 & Robust plug-in & 0.195 & 0.248 & 0.380 & 0.569 & 0.818 & 0.977 \\
 & JB & 0.394 & 0.661 & 0.896 & 0.992 & 1.000 & 1.000 \\
 & AD & 0.364 & 0.592 & 0.837 & 0.982 & 1.000 & 1.000 \\
 & SW & 0.403 & 0.639 & 0.874 & 0.989 & 1.000 & 1.000 \\
\addlinespace
$t_{5}$ & Fixed scale & 0.011 & 0.006 & 0.002 & 0.000 & 0.000 & 0.000 \\
 & Robust plug-in & 0.144 & 0.155 & 0.207 & 0.302 & 0.469 & 0.711 \\
 & JB & 0.211 & 0.393 & 0.626 & 0.863 & 0.983 & 0.999 \\
 & AD & 0.170 & 0.277 & 0.451 & 0.716 & 0.936 & 0.998 \\
 & SW & 0.217 & 0.354 & 0.561 & 0.816 & 0.971 & 0.999 \\
\addlinespace
Cont. normal & Fixed scale & 0.010 & 0.006 & 0.001 & 0.000 & 0.000 & 0.000 \\
 & Robust plug-in & 0.117 & 0.094 & 0.098 & 0.116 & 0.141 & 0.187 \\
 & JB & 0.232 & 0.411 & 0.642 & 0.855 & 0.977 & 0.999 \\
 & AD & 0.174 & 0.269 & 0.405 & 0.599 & 0.831 & 0.975 \\
 & SW & 0.230 & 0.381 & 0.587 & 0.815 & 0.964 & 0.999 \\
\addlinespace
Laplace & Fixed scale & 0.010 & 0.004 & 0.001 & 0.000 & 0.000 & 0.000 \\
 & Robust plug-in & 0.272 & 0.369 & 0.566 & 0.804 & 0.968 & 0.999 \\
 & JB & 0.290 & 0.507 & 0.783 & 0.968 & 0.999 & 1.000 \\
 & AD & 0.302 & 0.517 & 0.807 & 0.982 & 1.000 & 1.000 \\
 & SW & 0.321 & 0.520 & 0.798 & 0.977 & 1.000 & 1.000 \\
\addlinespace
\bottomrule
\end{tabular}
\begin{minipage}{0.95\linewidth}
\footnotesize
Notes: Entries are rejection frequencies at nominal level $5\%$ based on $R=10^4$ replications. Continuous alternatives are standardised to mean zero and unit variance before testing; this is relevant for the fixed-scale test, which assumes known location and scale, but immaterial for the plug-in and omnibus tests, which are location-scale invariant. The contaminated normal is the unit-variance version of $0.95\operatorname{N}(0,1)+0.05\operatorname{N}(0,9)$.
\end{minipage}
\end{table}

\begin{figure}[!p]
\centering
\includegraphics[width=0.78\textwidth]{power_curve_hgm_alternatives_robust.pdf}
\vspace{0.4cm}
\includegraphics[width=0.78\textwidth]{power_curve_standard_alternatives_robust.pdf}
\caption{Power curves for the robust plug-in statistic $\widetilde T_n^R$ at nominal level $5\%$, $R=10^4$ replications. Upper panel: HGM alternatives $G_c$, $c\in\{1.55,1.60,1.75,2.00,3.00\}$. Lower panel: standard alternatives standardised to mean zero and unit variance (contaminated normal, Laplace, chi-squared $\chi^2_3$, $t_3$, $t_5$, $t_{10}$).}
\label{fig:power_curves_robust}
\end{figure}

\section{Empirical illustration}\label{sec:application}

This section provides two descriptive illustrations using daily S\&P~500 index data from the FRED database \citep{fred_sp500}. Daily log returns $r_t=100\ln(P_t/P_{t-1})$ give $n=1255$ observations from 2~June~2021 to 1~June~2026. The first exercise applies the HGM test to raw returns; the second applies it to standardised residuals from a GARCH(1,1) model
\citep{bollerslev86},
\[
    r_t = \mu+\varepsilon_t, \quad
    h_t = \omega+\alpha\varepsilon_{t-1}^2+\beta h_{t-1},
    \quad z_t = \varepsilon_t/\sqrt{h_t},
\]
where $h_t$ is the conditional variance and $z_t$ is the standardised residual. The model is estimated by Gaussian MLE, with $\widehat\alpha=0.105$, $\widehat\beta=0.863$, $\widehat\alpha+\widehat\beta=0.968$. Both exercises are descriptive: the formal theory of Sections~\ref{sec:inference}--\ref{sec:plugin} assumes i.i.d. observations. The GARCH residual exercise additionally involves a generated-regressors problem; it serves as a robustness check on whether tail non-normality persists in the shape of the standardised residuals after filtering out volatility clustering. Figure~\ref{fig:garch_volatility} shows the estimated conditional volatility $\widehat h_t^{1/2}$: the 2022 equity decline and the tariff-driven market stress of early 2025 both produce sustained episodes of higher conditional volatility. Full descriptive statistics, goodness-of-fit tables, and GARCH diagnostics are in Tables \ref{tab:sp500_returns_desc}--\ref{tab:sp500_garch_residuals_tail} of Appendix~\ref{app:empirical}.

\subsection{Unconditional distribution of raw returns}\label{subsec:emp_raw}

Returns are standardised by $(\widehat\mu_n^R, \widehat\sigma_n^R)$ of~\eqref{eq:median_iqr} and $c$ is estimated by minimising $Q_n^R(c)$ with $d\nu_{1/2}$. Table~\ref{tab:sp500_returns_tests} reports normality test results: Jarque--Bera, Anderson--Darling, and Shapiro--Wilk all strongly reject the Gaussian null. The HGM statistic is $\widetilde T_n^R=26.65$ ($p\approx1.2\times10^{-7}$), with tail coefficient $\widehat c-3/2=0.135$, locating the departure in the symmetric heavy-tailed direction. Figure~\ref{fig:sp500_density_overlay} overlays the fitted densities; Figure~\ref{fig:sp500_tail_survival} plots two-sided survival functions on a logarithmic scale. In the observable range ($z\in[2,5]$) HGM lies below Student $t$ because the tail coefficient is small and the HGM asymptotic dominance sets in only beyond the sample range. By average log score, Student $t$ provides the best fit and HGM ranks third behind Laplace (Table~\ref{tab:sp500_returns_gof}). Lower-tail quantiles are reported in Table~\ref{tab:sp500_returns_var}. The models diverge materially at the $0.1\%$ level: Gaussian $-3.24$, Laplace $-4.60$, empirical $-4.82$, Student $t$ $-5.51$, HGM $-6.57$. The HGM quantile falls below the sample minimum ($-6.16$), illustrating that $\widehat c$ carries tail-risk implications that differ substantially from standard alternatives even when $\widehat c-3/2$ is small.

\subsection{Conditional innovation distribution: GARCH(1,1) residuals}\label{subsec:emp_garch}

The standardised residuals $\widehat z_t$ are treated as the sample, with $(\widehat\mu_n^R,\widehat\sigma_n^R)$ applied before fitting.  Excess kurtosis falls from $6.54$ (raw returns) to $1.56$, confirming that the GARCH filter accounts for most tail heaviness through time-varying volatility. Table~\ref{tab:sp500_garch_residuals_tests} reports test results. The HGM statistic falls to $\widetilde T_n^R=6.32$ ($p=0.006$), with $\widehat c =1.563$, much closer to the Gaussian boundary. Omnibus tests still reject normality, driven partly by residual left skewness ($-0.49$) that lies outside the scope of the symmetric HGM family. By log score, HGM now ranks second behind Student $t$ only, ahead of Laplace and Gaussian, reflecting closer alignment of the filtered residual tail structure with the hypergeometric family (Table~\ref{tab:sp500_garch_residuals_gof}). Figures~\ref{fig:sp500_garch_density} and~\ref{fig:sp500_garch_tail} display fitted densities and survival functions. Lower-tail quantiles at the $0.1\%$ level are (Table~\ref{tab:sp500_garch_residuals_var}): Gaussian $-3.12$, Student $t$ $-4.05$, Laplace $-4.68$, empirical $-4.70$, HGM $-5.02$. Student $t$ falls short of the empirical quantile; HGM gives the most extreme extrapolation, as for raw returns.

\subsection{Tail dominance and Value-at-Risk}\label{subsec:crossover}

For every $c>3/2$ and $z>0$, $1-G_c(z)>1-\Phi(z)$: the HGM distribution assigns strictly more probability to tail events than the Gaussian. The correction factor $\rho(z,c):=(1-G_c(z))/(1-\Phi(z))>1$ is strictly increasing in both $z$ and $c$ (Lemma~\ref{lem:R_monotone}), with exact values in Table~\ref{tab:correction}. HGM VaR quantiles are computed exactly by inverting $1-G_c(q_p)=1-p$; Table~\ref{tab:var_ratio} reports exact values alongside Gaussian quantiles. The ratio $q_p^{\mathrm{HGM}}/q_p^{\Phi}$ diverges as $p\to1^-$ for any $c>3/2$ (Proposition~\ref{prop:var_divergence}): any statistically significant rejection of $H_0:c=c_0$, however small $\widehat c-3/2$, implies understatement of extreme tail risk by the Gaussian model at sufficiently high confidence levels. Full details are in Appendix~\ref{app:crossover}.

\subsection{Limitations}\label{subsec:emp_limitations}

These exercises are not intended to establish that the HGM distribution fits better than all competitors: in both exercises Student $t$ provides a better overall log score, and omnibus tests have broader power. The purpose is to show what the fitted parameter $\widehat c$ implies for tail risk, and how that implication changes when volatility clustering is removed by GARCH filtering. Three limitations apply. First, the formal null theory assumes i.i.d.\ observations; raw returns exhibit serial dependence and volatility clustering \citep{engle82,bollerslev86}, and GARCH residuals are generated regressors, so size control is not formally guaranteed in either case. Second, the symmetric HGM family cannot capture the residual left skewness ($-0.49$) in the GARCH residuals, motivating future extensions to asymmetric families. Third, minimum-distance estimation rather than MLE means comparisons of log scores are descriptive rather than formal likelihood-ratio tests.

\begin{figure}[p]
\centering
\includegraphics[width=0.82\textwidth]{sp500_returns_density_overlay_standardised.pdf}
\caption{Fitted densities of robust-standardised raw returns: HGM and Student $t$ are similar near the centre; both are more peaked than the Gaussian.}
\label{fig:sp500_density_overlay}
\end{figure}

\begin{figure}[p]
\centering
\includegraphics[width=0.82\textwidth]{sp500_returns_tail_survival_standardised.pdf}
\caption{Two-sided survival functions $\Pr(|Y|>z)$ for robust-standardised raw returns, logarithmic scale: HGM lies below Student $t$ in the 2--5 standard deviation range despite the heavier asymptotic tail index, reflecting the small tail coefficient $\widehat c-3/2=0.135$.}
\label{fig:sp500_tail_survival}
\end{figure}

\begin{figure}[p]
\centering
\includegraphics[width=0.82\textwidth]
    {sp500_garch_residuals_density_overlay_standardised.pdf}
\caption{Fitted densities of robust-standardised GARCH(1,1) residuals: HGM and Student $t$ are the closest competitors; excess kurtosis is reduced relative to raw returns.}
\label{fig:sp500_garch_density}
\end{figure}

\begin{figure}[p]
\centering
\includegraphics[width=0.82\textwidth]
    {sp500_garch_residuals_tail_survival_standardised.pdf}
\caption{Two-sided survival functions $\Pr(|Y|>z)$ for robust-standardised GARCH(1,1) residuals, logarithmic scale: HGM and Student $t$ track the empirical tail more closely than for raw returns; Laplace overstates tail mass at intermediate thresholds.}
\label{fig:sp500_garch_tail}
\end{figure}

\clearpage

\section{Conclusion}\label{sec:conclusion}

This paper develops a complete inferential framework for a one-parameter family of distribution functions built from the confluent hypergeometric function ${}_1F_1$. The family $\{G_c:c\geq3/2\}$ places the standard normal at an exact finite boundary, with algebraic tails of fixed index $2$ for $c>3/2$ and a beta-precision Gaussian scale-mixture representation. These properties are mathematical consequences of the endpoint constraints that determine the family from the ${}_1F_1$ representation of $\Phi$, not modelling assumptions. More broadly, the construction reverses the usual role of hypergeometric functions in econometrics. Rather than arising as analytic solutions to a model or inferential problem specified in advance, the hypergeometric representation is used here as the starting point for model construction: selected parameters are freed, the restrictions required of the econometric object are imposed, and the resulting admissible parameters are estimated from the data.

The central result is that the standardised fit-improvement statistic $\widetilde T_n$ converges to $\frac{1}{2}\delta_0+\frac{1}{2}\chi_1^2$ under the Gaussian null, with the explicit rejection rule $\widetilde T_n>\chi^2_{1,1-2\alpha}$. This limit holds for all three versions of the test: fixed-scale, mean/SD plug-in, and robust median/IQR plug-in, with only the projected covariance kernel changing across versions. The modular structure is deliberate: extensions to new standardisation procedures require only a new kernel, leaving the limit distribution, curvature, and rejection rule unchanged. The fitted parameter $\widehat c$ provides a constructive complement to omnibus normality tests: it locates the departure from normality in the hypergeometric tail direction, quantifies the departure as a single interpretable index, and implies tail-risk extrapolations that differ materially from Gaussian benchmarks even when $\widehat c-3/2$ is small.

The framework opens a natural research programme. The most immediate extension frees the tail index, generating a two-parameter family that nests the present symmetric case and extends the boundary inference to a two-dimensional constrained problem. A closely related development introduces asymmetry, allowing the family to capture signed skewness alongside heavy tails and generating a richer mixture limit at the boundary. A further line of research develops the formal two-step asymptotic theory for the test applied to residuals from parametric volatility models, where first-step GARCH estimation uncertainty propagates through an additional projected-process correction to the covariance kernel. A fourth extension pursues likelihood-based inference using the explicit density $g_c$ of Corollary~\ref{cor:density}, yielding an efficiency comparison and a likelihood-ratio counterpart to the fit-improvement statistic. A longer-horizon direction pursues multivariate normality testing via matrix-argument hypergeometric functions, replacing the scalar boundary $c_0=3/2$ with a matrix boundary at the identity. The present paper establishes the foundations on which this programme builds.

\newpage