EconBase
← Back to paper

Inference under First-Order Degeneracy

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,900 characters · 15 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference under First-Order Degeneracy

\thispagestyle{empty}

abstractWe study inference in models where a transformation of parameters exhibits first-order degeneracy --- that is, its gradient is zero or close to zero, making the standard delta method invalid. A leading example is causal mediation analysis, where the indirect effect is a product of coefficients and the gradient degenerates near the origin. In these local regions of degeneracy the limiting behaviors of plug-in estimators depend on nuisance parameters that are not consistently estimable. We show that this failure is intrinsic --- around points of degeneracy, both regular and quantile-unbiased estimation are impossible. Despite these restrictions, we develop minimum-distance methods that deliver uniformly valid confidence intervals. We establish sufficient conditions under which standard chi-square critical values remain valid, and propose a simple bootstrap procedure when they are not. We demonstrate favorable power in simulations and in an empirical application linking teacher gender attitudes to student outcomes. Keywords: Delta Method, Higher Order Asymptotics, Impossibility, Minimum Distance Inference JEL Codes: C12, C13, C18

Introduction

The delta method is a fundamental tool in econometric analysis for deriving the asymptotic distribution of smooth functions of estimators. In its standard form, the delta method relies on a non-zero gradient of the smooth function with respect to the primitive parameter. When this condition holds, first-order linearization provides an accurate approximation to the sampling distribution. However, in many empirically relevant scenarios, the gradient may be zero at certain points in which case higher-order terms become essential. As a result, the limiting distribution near points of degeneracy typically differs substantially from that obtained under standard regularity. A leading example arises in causal mediation analysis, where the indirect effect is the product of two primitive parameters: the effect of the treatment on the mediator and the effect of the mediator on the outcome. When these effects are zero, the gradient of the indirect effect degenerates, leading to a nonstandard limiting distribution of the Wald statistic sobel1982asymptotic.

In practice, the researcher does not know whether the true parameter lies near such points of degeneracy, and thus cannot know which asymptotic approximation is appropriate for inference. This challenge has motivated a number of recent papers on hypothesis testing when the gradient may be degenerate, see, for example, van_garderen_nearly_2024 for an analysis of the mediation model mentioned above and Dufour2025WaldTW in a more general treatment of Wald-type statistics. These papers acknowledge discontinuities in limiting distributions at the point of degeneracy, and analyze the resulting distortions in Wald-type statistics. Yet, they do not propose a unified asymptotic framework for studying local regions of degeneracy. Moreover, most existing papers are interested in specific point hypotheses, leaving open the broader question of how to construct uniformly valid confidence intervals.

This paper makes two main contributions to the literature. Our first contribution is a formal asymptotic framework for studying the behavior of statistics in local regions of first-order degeneracy. In our framework, the primitive parameter is modeled as local to the point of degeneracy, in the spirit of weak identification asymptotics, where identifying information is local to zero. Under this setup, the behavior of simple plug-in estimators becomes nonstandard and depends on local parameters that cannot be consistently estimated, formally capturing an observation in the existing literature that the standard delta method does not properly approximate the behavior of plug-in estimators near degenerate points miller2024testing.

Leveraging Le Cam's limit of experiments framework lecam1970assumptions,le1972limits, we show that the problem of estimating a smooth function under degeneracy is asymptotically equivalent to estimating a quadratic form of the shift parameter in a Gaussian shift model. Within this model, we show that equivariant-in-law and quantile-unbiased estimators cannot be obtained. Translated back to the original estimation problem, these results imply that regular estimation is impossible --- the limiting behavior of any properly scaled and centered estimator in local regions of degeneracy depends on nuisance parameters that cannot be consistently estimated. Moreover, there do not exist asymptotically similar confidence intervals for the transformation of interest in local regions of degeneracy.

Our second contribution is to construct confidence intervals that are both uniformly valid and exhibit favorable power in local regions of degeneracy compared to the few existing alternatives. The impossibility results mentioned above imply that standard delta-method based confidence intervals may not be uniformly valid in such regions. Indeed, in the context of mediation analysis our simulation study shows that standard Wald-statistic based tests lead to confidence intervals that undercover when the true primitive parameter is near the origin. We thus take a different approach and propose confidence intervals based on test inversion with a minimum-distance test statistic. We first show that when the parameter dimension is two, the standard chi-square critical value is uniformly valid under either of two conditions: (i) the curvature of the null curve is not too large, or (ii) the two branches of the null curve are sufficiently close. These conditions hold in the leading mediation example. In more general settings, we propose a bootstrap critical value based on a quadratic approximation of the test statistic, and show that when the true parameter is well separated from the point of degeneracy, the bootstrap critical value is nearly identical to the efficient one.

We demonstrate the empirical relevance of our results in both simulation study and real world. In simulation study our proposed methods are shown to control size uniformly over the parameter space while standard Wald-based inference can overreject when the true primitive parameter is close to points of degeneracy, a finding also noted by Dufour2025WaldTW. We additionally demonstrate favorable power properties of our proposed methods when compared to a method proposed by AM-2016-Geometric, which turns out to also be applicable in this setting.\footnote{In their original analysis, AM-2016-Geometric were interested in testing non-linear restrictions in the context of weak identification.} These improvements in power are also seen in an application to the data of AEM-2018, who consider the effect of teachers' gender attitudes on student outcomes. In revising their mediation analysis, we find that our proposed methods deliver tighter confidence intervals than existing methods for all parameters. Our results support a conclusion by van_garderen_nearly_2024 that the mediation effect of a one-year exposure to a teacher with traditional negative views is negative, while existing methods cannot rule out an effect of size zero.

The rest of this paper proceeds as follows. This section concludes with a review of the related literature. (ref) gives examples of when first-order degeneracy may be a concern. (ref) formally establishes the impossibility results mentioned above and discusses implications for hypothesis testing. (ref) introduces the minimum distance based inference procedures and discusses how uniformly valid critical values may be constructed. (ref) contain, respectively, the simulation study and empirical application to the data of AEM-2018. (ref) concludes. Proofs are deferred to (ref).

Literature Review

Our paper is related to previous literature on econometrics and statistics studying inference under degeneracy, statistical impossibility results, and testing non-linear restrictions.

There is a growing literature examining hypothesis tests in which the null includes points of singularity. gaffke1999asymptotic show that the distribution of the Wald statistic at points of degeneracy is nonstandard, and gaffke2002asymptotic derive its asymptotic distribution under a variety of singular null hypotheses. drton2016wald demonstrate the conservativeness of the Wald test at degeneracy points for quadratic forms and for bivariate monomials of arbitrary degree. dufour2016rank propose rank-robust regularized Wald-type tests allowing for singular covariance matrices. Dufour2025WaldTW analyze Wald tests for polynomial restrictions with possibly multiple constraints and show that such tests can under-reject, over-reject, or even diverge under the null; see also Dufour2025WaldTW for additional references in this area. Our paper differs from this literature in two key ways. First, prior work focuses on testing problems where the null itself contains the singularity, while our interest lies in constructing uniformly valid confidence intervals when the null may be near a singularity. In simulations, we show that the upper bound on the Wald statistic derived in Dufour2025WaldTW does not, in general, yield valid confidence intervals. Second, instead of using a Wald statistic, we employ a minimum distance–based test statistic, which is bounded in probability by construction. This approach avoids the divergence issues documented in Dufour2025WaldTW.

Our paper is also related to the hypothesis testing problem with a curved null. AM-2016-Geometric study this problem and show that the distribution of minimum-distance statistics is dominated by a tractable distribution that depends only on the maximal curvature of the null manifold relative to the known variance matrix. Inverting their test leads to uniformly valid confidence intervals. However, when the curvature of the null hypothesis is large, for example, in testing the significance of an indirect effect, their procedure yields critical values that are close to those from projection-based methods. By contrast, our procedure exploits the possibility that the null hypothesis may include multiple manifolds that are close to one another, which in turn reduces the critical value.

Finally, our paper contributes to the econometric literature on statistical impossibility results. In particular, it is related to work by hirano2012impossibility who show that regular estimation of directionally, but not fully, differentiable functions is unattainable. Our paper takes a similar approach to that of hirano2012impossibility in that we rule out properties of estimators by analyzing a limiting experiment lecam1970assumptions,le1972limits. However, the target functional in our limit experiment is distinct from that of hirano2012impossibility. This approach of ruling out behaviors by analyzing limit experiments has also been utilized successfully by kaji2021theory and andrewsgmm in the study of weak identification. The present work is also related to work by chen2019inference who show that all standard bootstrap procedures necessarily fail at points of degeneracy, similarly to how fang2019inference establish that the bootstrap necessarily fails as an inference procedure for the functionals considered in hirano2012impossibility.

Overview and Examples

Consider a parameter \(\theta \in \Theta \subseteq \SR^d\) and a twice continuously differentiable function \(g: \Theta \to \SR\). We are interested in inference on \(g(\theta)\) in local neighborhoods of a point \(\theta_\star\) for which \(\nabla g(\theta_\star) = 0\). Below, we give some empirically relevant examples of when such a phenomenon may occur.

example[Mediation Analysis] Consider a causal mediation analysis with parameter \(\theta = (\theta_1,\theta_2)'\), where \(\theta_1\) represents the effect of a treatment variable on a mediator and \(\theta_2\) represents the effect of the mediator on the outcome. The indirect effect of the treatment on the outcome is then given by \(g(\theta) = \theta_1\theta_2\). At \(\theta_\star = (0,0)'\), we have \(\nabla g(\theta_\star) = 0\), which complicates inference on \(g(\theta)\) in local regions of \(\theta_\star\). As a result, recent works have proposed tests for the specific null-alternate pair, \(H_0: g(\theta) = 0\) against \(H_1: g(\theta) \neq 0\), see van_garderen_nearly_2024 or hillier_improved_2024. However, these works do not consider the more general problem of constructing confidence intervals in local regions of the origin. \qed
example[Impulse Response Function] Consider an autoregressive \(\text{AR}(1)\) model of the form \(y_t = \theta y_{t-1} + u_t\) where \(y_t,y_{t-1} \in \SR\), \(\theta \in \SR\), and \(u_t \in \SR\) is a white noise process. The “impulse response function” is defined as \(g(\theta) = \theta^h\) and measures the impact at time period \(h\) of an initial shock. Due to the importance of this in macroeconomic analysis, inference on \(g(\theta)\) has received attention from the econometric literature IK-2002, Gospodinov-2004, Mikusheva-2012, mostly related to inference when \(\theta\) is close to one --- the so-called “unit-root” problem. However, due to degeneracy, inference on the impulse response function can also be complicated when \(\theta\) is close to \(\theta_\star = 0\) as \(\frac{\partial }{\partial \theta} g(\theta_\star) = h\theta_\star^{h-1} = 0\) benkwitz2000problems,lutkepohl2013introduction. \qed
example[Breakdown Point Analysis] Consider a missing data setup in which the researcher observes \(\{Y_iD_i, D_i, X_i\}_{i=1}^n\), where \(D_i \in \{0,1\}\) represents whether or not an observation's outcome $Y_i$ is observed, and \(X_i\) is a set of discrete covariates, i.e., \(X_i \in \calX \coloneqq \{x_1,\dots,x_{K}\}\). To achieve identification of parameters of interest, assumptions are typically made about the selection mechanism, such as “missing conditionally at random”, i.e, \(Y_i \perp D_i \mid X_i\). ober2026robustness proposes assessing the robustness of results to these assumptions through a breakdown point analysis. This assessment involves generating a confidence interval for the squared Hellinger distance between \(P_0\), the distribution of \(X_i \mid D_i = 0\), and \(P_1\), the distribution of \(X_i \mid D_i = 1\). Since \(X_i\) is discrete, this distance can be written as \[ g(\theta) = H^2(P_0,P_1) = \frac{1}{2}\sum_{k=1}^K (\sqrt{\theta_{0,k}} - \sqrt{\theta_{1,k}})^2 \] where \(\theta_{d,k} = \Pr(X = x_k \mid D = d)\). Simple sample analog estimators of \(\theta_{0,k}\) and \(\theta_{1,k}\) are \(\sqrt{n}\)-consistent and asymptotically normal under mild assumptions. However, the derivative of \(g(\theta)\) with respect to the quantities \(\theta_{0,k}\) and \(\theta_{1,k}\) are given by \begin{align*} \frac{\partial}{\partial\, \theta_{0,k}}g(\theta) = \frac{1}{2}\frac{\sqrt{\theta_{0,k}} - \sqrt{\theta_{1,k}}}{\sqrt{\theta_{0,k}}}, \;\;\;\;\;\;\;\;\;\; \frac{\partial}{\partial \theta_{1,k}} g(\theta) = -\frac{1}{2}\frac{\sqrt{\theta_{0,k}} - \sqrt{\theta_{1,k}}}{\sqrt{\theta_{1,k}}}. \end{align*} Let \(\theta_\star\) be a point such that \(\theta_{0,k} = \theta_{1,k}\) for all \(k\). This occurs if the data is missing completely at random, that is \(D \perp (Y,X)\). At \(\theta_\star\), the derivatives above are uniformly equal to zero and the squared Hellinger distance \(g(\theta)\) between \(P_0\) and \(P_1\) is zero. Thus, when \(\theta\) is close to \(\theta_\star\), that is when \(P_0\) is close to \(P_1\), standard approaches to inference on \(g(\theta)\) will fail. \qed
example[Weak IV Bias and Size Distortion] Consider a standard homoskedastic linear IV model, \begin{equation*} \begin{split} y_i &= x_i\beta + \eps_i \\ x_i &= z_i'\theta + v_i \end{split} \end{equation*} where \(y_i, x_i \in \SR\), \(\mathbb{E}[(\eps_i, v_i)'] = 0\), and \(Z = (z_i',\dots,z_n')' \in \SR^{n \times d_z}\) is treated as fixed. When \(\theta\) is close to zero, identification is referred to as “weak” and it is well known that standard inference procedures for \(\beta\) fail to control size staiger1994instrumental. StockYogo-2005 provide bounds on the size distortion of Wald tests for \(\beta\) in terms of the concentration parameter, \[ g(\theta) = \theta'(Z'Z)\theta/\sigma_v^2. \] GIR-2021 extend this analysis and develop confidence intervals for the bias and size distortion. In both papers, the researcher makes inferences about the concentration parameter by examining the distribution of the \(F\)-statistic, a scaled version of \(g(\hat\theta)\), where \(\hat\theta\) is the OLS estimate of \(\theta\). The analyses of both StockYogo-2005 and GIR-2021 are complicated by the fact that the limiting distribution of \(g(\hat\theta)\) is non-standard when \(\theta\) is close to zero, that is, when identification is weak. This can also be seen as inference in local regions of degeneracy --- at \(\theta_\star = 0\) we have that \(\nabla g(\theta_\star) = 2(Z'Z)\theta/\sigma_v^2 = 0\). \qed
example[Explained Variance in Linear Regression] Consider a linear regression model, \begin{equation*} Y = X'\theta + \eps, \;\;\; \mathbb{E}[\eps X] = 0 \end{equation*} and define \(\sigma_Y^2 = \text{Var}(Y)\) and \(\Sigma_X = \mathbb{E}[XX']\). A parameter of interest is the proportion of variance in \(Y\) explained by the linear model with \(X\), i.e \begin{equation*} g(\theta) = \theta'\Sigma_X \theta/\sigma_Y^2, \end{equation*} that is, the population $R^2$. Although empirical work typically reports only a point estimate of $R^2$, reporting a confidence set for $R^2$ is informative for comparing the explanatory power or predictive performance of competing models hawinkel2024out. When $R^2$ is bounded away from zero and one, standard errors and confidence intervals can be obtained using conventional asymptotic approximations cohen2013applied. However, at \(\theta_\star = 0\) we have that \(\nabla g(\theta) = 2\Sigma_X\theta_\star/\sigma_Y^2 = 0\). Consequently, inference on the explained variance is non-standard when \(\theta\) is close to zero, or equivalently, when $R^2$ is close to zero. This type of parameter is also of interest in labor economics when explaining variation in wage regressions. If a model under consideration can only explain a weak amount of variation in wage dispersion, we may expect \(\theta\) to be close to zero. CHK-2013 compare the baseline \citet*{AKM-1999} (AKM) model with various extensions in terms of each models ability to explain increases in wage inequality in West Germany. They find that these extensions provide little explanatory power on top of the baseline AKM model. The additional variance explained by these extensions corresponds to the linear regression model with \(Y\) equal to the residual from the AKM model and \(X\) equal to the new fixed-effect terms introduced by the extended models. They find that these new fixed-effect terms are close to zero suggesting that degeneracy may be a concern when conducting inference on \(g(\theta)\).\qed

Impossibility Results

In this section we establish that standard approaches to inference on $g(\theta)$ necessarily fail in local regions of first-order degeneracy. We begin in (ref)\textcolor{pastelred}{.1} by introducing a parametric framework to study this problem and defining what it means for an estimator to be “regular” in this setting. (ref)\textcolor{pastelred}{.2} then uses a representation theorem to show that the problem reduces to estimation of quadratic forms in a Gaussian shift experiment, where we prove that well-behaved estimators cannot be constructed. (ref)\textcolor{pastelred}{.3} extends the analysis in two directions: first, to hypothesis testing problems where the null hypothesis that \(g(\theta) = g(\theta_\star)\) may hold on a nontrivial subset of the parameter space, and second, to infinite-dimensional models, where we show that the impossibility results remain valid so long as the model contains a suitable parametric submodel. Together, these results demonstrate that standard approaches to inference necessarily break down in local regions of degeneracy.

Preliminaries

We begin by assuming that the researcher observes data \(X^{(n)} = (X_1,\dots,X_n)\) drawn from a parametric model \(P_{n,\theta}\),

equation[equation omitted — 108 chars of source]

where \(\theta \in \Theta \subset \SR^d\) and \(\Theta\) is a compact set with a nonempty interior, \(\Theta^\circ \neq \emptyset\). Let \(\calX_i\) denote the support of \(X_i\), which could be a general space, and denote \(\calX^n = \bigtimes_{i=1}^n \calX_i\). We assume that the sequence of statistical models \(\left(P_{n,\theta}: \theta \in \Theta^\circ\right)\), indexed by the sample size \(n\), is locally asymptotically normal in the sense of lecam1960lan.

assumption[Local Asymptotic Normality] There exists a sequence \(r_n \to \infty\) such that for every \(\theta \in \Theta^\circ\) and every sequence \(h_n \to h \in \SR^d\) \begin{equation} \begin{split} \log \left(\frac{dP_{n,\theta + h_n/r_n}}{dP_{n,\theta}}(X^{(n)})\right) = h'\Delta_n - \frac{1}{2}h'\Gamma_\theta h + Z_n(h) \end{split} \end{equation} where \(\Delta_n\) converges in distribution to \(N(0,\Gamma_\theta)\) under the sequence of measures \(P_{n,\theta}\), \(\Delta_n \overset{\theta}{\rightsquigarrow} N(0,\Gamma_\theta)\), and \(Z_n(h)\) converges in probability to zero under \(P_{n,\theta}\) for every \(h \in \SR^d\), \(Z_n(h) \xrightarrow p 0\).
example[Smooth Parametric Models] A leading example of a model that satisfies (ref) is when the researcher observes i.i.d data, \(X_i \overset{iid}{\sim} P_\theta\) where \(\theta \in \Theta\). Assume that there exists a dominating measure \(\mu\) such that \(P_\theta \ll \mu\) for all \(\theta \in \Theta^\circ\) and the Radon-Nikodym densities \(p_\theta = dP_\theta/d\mu\) are differentiable in quadratic mean, that is there is a function \(\dot\ell_\theta\) such that for any \(\theta \in \Theta^\circ\), \begin{equation} \int \left[\sqrt{p_{\theta + h}} - \sqrt{p_\theta} - \frac{1}{2}h'\dot\ell_\theta \sqrt{p_\theta}\right]^2d\mu = o(\|h\|^2),\;\;\; h \to 0 \end{equation} and such that the Fisher information, \(\Gamma_\theta = P_\theta\dot\ell_\theta\dot\ell_\theta'\), is nonsingular. Let \(P_{n,\theta} = \otimes_{i=1}^n P_\theta\), then, for any \(\theta \in \Theta^\circ\) and any \(h_n \to h \in \SR^d\), the sequence of log likelihood ratios satisfies (vanDerVaart1998, Theorem 7.2): \begin{equation*} \begin{split} \log\left(\frac{dP_{n,\theta + h/\sqrt{n}}}{dP_\theta}(X^{(n)})\right) &= \log \prod_{i=1}^n \frac{p_{\theta + h/\sqrt{n}}}{p_\theta}(X_i) \\ &= \frac{1}{\sqrt{n}}\sum_{i=1}^n h'\dot\ell_\theta(X_i) - \frac{1}{2} h'\Gamma_\theta h + R_n(h), \end{split} \end{equation*} where \(R_n(h) = o_{P_{n,\theta}}(1)\) for all \(h \in \SR^d\). By the central limit theorem, \(\frac{1}{\sqrt{n}}\sum_{i=1}^n \dot\ell_\theta(X_i) \overset{P_{n,\theta}}{\rightsquigarrow} N(0, \Gamma_{\theta})\). Thus, by letting \(\Delta_n = \frac{1}{\sqrt{n}}\sum_{i=1}^n \dot\ell_\theta(X_i)\) we see that (ref) is satisfied with \(r_n = \sqrt{n}\). \qed

We study a twice continuously differentiable scalar functional \(g:\Theta \to \SR\), and the behavior of estimators of \(g(\theta)\) in local regions of a point \(\theta_\star \in \Theta^\circ\) which is such that the first-order derivatives of \(g(\cdot)\) at \(\theta_\star\) are zero. We will refer to \(\theta_\star\) as the “point of degeneracy” and local neighborhoods of \(\theta_\star\) as “local regions of degeneracy.”

assumption[Differentiability] The function \(g:\Theta \to \SR\) is twice continuously differentiable on \(\Theta\), a compact subset of \(\SR^d\), with \(\nabla g(\theta_\star) = 0\) and \(\nabla^2g(\theta_\star) \neq 0\) for some \(\theta_\star \in \Theta^\circ\).

Given the maintained assumption of a locally asymptotically normal model, it is useful to examine regions close to \(\theta_\star\) by adopting a local parameterization around \(\theta_\star\), defining \[ \theta_{n, h} = \theta_\star + h/r_n \] and letting \(P_{n,h} = P_{n, \theta_\star + h/r_n}\). In our framework an estimator is an arbitrary measurable function of the data, \(\Psi_n: \calX^n \to \SR\). We consider sequences of estimators, \(\Psi_n\), that converge in distribution under every sequence of alternative distributions \(P_{n,h}\) to some limiting law, \(\calL_h\). This is denoted

equation[equation omitted — 121 chars of source]

where we note that the convergence rate is \(r_n^2\) instead of \(r_n\) due to the fact that \(g(\theta)\) is “flat” around \(\theta_\star\). It is straightforward to show that tests for \(g(\theta)\) based on estimators whose convergence rates are slower than \(r_n^2\) when \(\theta\) is close to \(\theta_\star\) have trivial power against local alternatives of the form \(g(\theta_\star + h/r_n)\).

example[Plug-In Estimators] Suppose the researcher has access to an estimator \(\hat\theta\) of \(\theta\) that satisfies \[ r_n(\hat\theta - \theta_{n,h}) \overset{h}{\rightsquigarrow} \calW \] for every \(h \in \SR^d\). In the smooth parametric models described in (ref), such an estimator could be the maximum likelihood estimator or a Bayes estimator such as the posterior mean. Since \(g(\cdot)\) is assumed to be twice continuously differentiable, the limiting behavior of the plug-in estimator \(g(\hat\theta)\) can be found via the second order delta method \[ r_n^2(g(\hat\theta) - g(\theta_{n,h})) \overset{h}{\rightsquigarrow} h'\nabla^2g(\theta_\star)\calW + \frac{1}{2}\calW'\nabla^2g(\theta_\star)\calW \] The behavior of the plug in estimator, \(g(\hat\theta)\), depends on the local parameter \(h\). \qed

We focus on ruling out regular and locally asymptotically \(\alpha\)-quantile unbiased estimation of \(g(\theta_{n,h})\) in local regions of \(\theta_\star\). Formally, these notions are defined as follows.

definition[Regularity] Let $\Psi_n$ be an estimator satisfying (ref), and let $\alpha \in (0,1)$. \begin{enumerate}[(i)] • $\Psi_n$ is regular if its limiting distribution does not depend on $h$, i.e.\ there exists a distribution $\calL$ on $\mathbb{R}$ such that $\calL_h = \calL$ for all $h \in \mathbb{R}^d$. • $\Psi_n$ is locally asymptotically $\alpha$-quantile unbiased if its limiting $\alpha$-quantile is zero for every $h$, i.e.\ $\calL_h\{(-\infty,0]\} = \alpha$ for all $h \in \mathbb{R}^d$. \end{enumerate}

The existence of regular estimators is closely tied to the validity of Wald-type inference procedures --- without regular estimators, standard Wald-type inference procedures that compare test statistics to fixed critical values will not have correct asymptotic size. Similarly, the existence of locally asymptotically \(\alpha\)-quantile unbiased estimators is closely tied to the existence of asymptotically similar confidence intervals for \(g(\theta)\) in local regions of \(\theta_\star\). Since any asymptotically similar confidence interval of the form \((-\infty, \hat c]\) can be converted into a locally asymptotically \(\alpha\)-quantile unbiased estimator by taking \(\Psi_n = \hat c\), by ruling out such estimators we also rule out the possibility of similar one-sided confidence intervals.\footnote{The focus on one-sided confidence intervals is largely for simplicity of exposition. If \((-\infty, \hat c_1]\) and \([\hat c_2, \infty)\) are two asymptotically similar confidence intervals for \(g(\theta)\) each with coverage rate \(1 - \alpha/2\) and \(\hat c_1 \leq \hat c_2\) with probability approaching one, then \([\hat c_2, \hat c_1]\) is asymptotically similar with coverage rate \(1 - \alpha\).}

Analysis in the Limiting Experiment

To examine the possible behavior of such estimators, we make use of a representation result, given below in (ref), which is a slight adaptation of Theorem 8.3 in vanDerVaart1998. This earlier result is, in turn, a version of Le Cam's limit of experiments analysis for locally asymptotically normal models lecam1970assumptions,le1972limits.

prop[Limit Experiment] Suppose (ref) holds, and let $\Psi_n$ be a sequence of estimators satisfying (ref). Then there exists a randomized statistic $\Psi(Z,U)$, where $Z$ is drawn from the Gaussian shift experiment \[ Z \sim N(h, \Gamma^{-1}_{\theta_\star}), \quad h \in \mathbb{R}^d, \] and $U \sim \mathrm{Unif}(0,1)$ independent of $Z$, such that \[ \Psi(Z,U) - \tfrac{1}{2} h^\top \nabla^2 g(\theta_\star) h \;\sim\; \calL_h \quad \text{for all } h \in \mathbb{R}^d. \]

(ref) establishes an equivalence between estimating \(g(\theta)\) in local regions of first-order degeneracy and estimating of a quadratic form of the mean parameter in a Gaussian shift model in which one observes a single draw \(Z \sim N(h, \Gamma_{\theta_\star}^{-1})\), where \(\Gamma_{\theta_\star}^{-1}\) is known but \(h\) is not. In particular, in a spirit similar to the approach in hirano2012impossibility, we can rule out sufficiently regular behavior of estimators of \(g(\theta)\) in local regions of \(\theta_\star\) if the corresponding behavior is not permissible in the Gaussian shift model. Intuitively, sufficiently regular estimation of quadratic forms in the Gaussian shift experiment is not possible since the parameter of interest changes non-linearly as the mean parameter \(h\) varies over \(\SR^d\) while the distribution of \(Z\) changes in a linear fashion.

To illustrate, suppose that there was an estimator, \(\Psi(Z,U)\), and law, \(\calL\) with \(\int x^2\,d\calL(x) < \infty\), such that \(\Psi(Z,U) - \frac{1}{2}h'\nabla^2 g(\theta_\star)h\) is distributed according to \(\calL\) for all \(h \in \SR^d\). Since any such estimator can be turned into an unbiased estimator by subtracting off the mean of \(\calL\), it is without loss of generality to assume that \(\calL\) is mean zero and thus that \(\Psi(Z,U)\) is unbiased. On the other hand, the Cramér-Rao lower bound for the variance of any unbiased estimator of \(\frac{1}{2}h'\nabla^2g(\theta_\star)h\) in the Gaussian shift model yields

align[align omitted — 168 chars of source]

By letting \(h\) vary over \(\SR^d\), the right hand side of (ref) can be made arbitrarily large while the left hand side is bounded by the second moment of \(\calL\). Thus, no such estimator can exist. Our full argument relies on analyzing characteristic functions, but the intuition is similar.

remark[] It is instructive to compare the argument sketched above to the argument of hirano2012impossibility, who rule out regular estimation of \(g(\theta)\) when \(g\) is directionally, but not fully differentiable at a point \(\theta_\star\). The hirano2012impossibility argument relies on analyzing the behavior of a potential regular estimator as the local parameter \(h\) approaches zero. Our arguments, on the other hand, rule out regular estimation by analyzing the “global” behavior of a potential regular estimator, that is, the behavior as \(h\) varies over \(\SR^d\). The approach taken by hirano2012impossibility does not apply in the present setting as the parameter of interest in the limit experiment is non-linear but continuously differentiable at zero. In contrast, in the limit experiment of hirano2012impossibility the parameter of interest is a function \(\kappa(h)\) which is exactly linear around values of \(h \neq 0\), but is not continuously differentiable at zero. The argument of hirano2012impossibility is able to additionally rule out locally unbiased estimation whereas in our setting locally unbiased estimation is possible. \qed
prop[] Let \(Z \sim N(h, \Gamma_{\theta_\star}^{-1})\) and \(U \sim \mathrm{Unif}(0,1)\) independently of \(Z\). Let \(J\) be a \(d \times d\) non-zero, symmetric matrix. \begin{enumerate} • There is no randomized statistic \(\Psi(Z,U)\) and law \(\calL\) on \(\SR\) with \(\Psi(Z,U) - h'Jh \overset{h}{\sim} \calL\) for all \(h \in \SR^d\). • Let \(\{\calL_h\}_{h \in \SR^d}\) be a system of probability measures on \(\SR\) such that (i) \(\calL_h\{(-\infty,0]\} =\alpha\) for some \(\alpha \in (0,1)\) and (ii) the CDFs associated with \(\calL_h\), \(F_h(\cdot)\), are differentiable at zero with derivative bounded below by some \(\eps > 0\). Then, there does not exist a randomized statistic \(\Psi(Z,U)\) such that \(\Psi(Z,U) - h'Jh \sim \calL_h\) for all \(h \in \SR^d\). \end{enumerate}

Together, (ref) can be combined for the main result of this section, which rules out sufficiently regular estimation in local areas of first-order degeneracy.\footnote{In the application of (ref) to our setting, take \(J = \frac{1}{2}\nabla^2 g(\theta_\star)\). The result in (ref) rules out well-behaved estimation of any quadratic form of the shift parameter, not just those associated with the Hessian of \(g(\cdot)\).}

thm[Impossibility of Regular Estimation] Suppose (ref) hold. \begin{enumerate} • There is no estimator sequence $\Psi_n$ and law \(\calL\) on \(\SR\) such that \[ r_n^2\big(\Psi_n - g(\theta_{n,h})\big) \;\overset{h}{\rightsquigarrow}\; \calL \quad \text{for all } h \in \mathbb{R}^d. \] • Let $\{\calL_h\}_{h \in \mathbb{R}^d}$ be a family of distributions such that (i) $\calL_h\{(-\infty,0]\} = \alpha$ for some fixed $\alpha \in (0,1)$ and all $h$, and (ii) the CDFs, $F_h(\cdot)$, of $\mathcal{L}_h$ are differentiable at zero with derivatives bounded below by $\epsilon>0$. Then there is no estimator sequence $\Psi_n$ such that \[ r_n^2\big(\Psi_n - g(\theta_{n,h})\big) \;\overset{h}{\rightsquigarrow}\; \calL_h \quad \text{for all } h \in \mathbb{R}^d. \] \end{enumerate}

(ref) rules out sufficiently well-behaved estimation of \(g(\theta)\) when the true parameter is “close” to \(\theta_\star\). In particular, (ref)(a) rules out the possibility of regular estimation --- the properly scaled and centered behavior of any estimator \(\Psi_n\) of \(g(\theta)\) must depend, in local regions of \(\theta_\star\), on the local parameter \(h\), which cannot be consistently estimated. Similarly, (ref)(b) rules out the possibility of \(\alpha\)-quantile unbiased estimation. As mentioned below (ref), this result has profound implications for inference on \(g(\theta)\) in local regions of \(\theta_\star\). In particular, both asymptotically exact Wald-type inference procedures and asymptotically similar confidence intervals for \(g(\theta)\) are unavailable in local regions of degeneracy \(\theta_\star\).

(ref)(a) also has implications for the construction of efficient estimators of \(g(\theta)\) in local regions of \(\theta_\star\). Standard notions of efficiency are tied to comparing the asymptotic risk of regular estimators. (ref)(a) shows that, after properly scaling, such regular estimators are not available in local regions of degeneracy. Consequently, alternative notions of efficiency must be considered in these settings and standard estimators may not be optimal in these regions. As an example, one can show that estimators of \(g(\theta)\) that are efficient under a standard asymptotic regime can be dominated by alternative estimators in local asymptotic mean squared error around points of degeneracy.

remarkAs with the impossibility of locally asymptotically \(\alpha\)-quantile unbiased estimation for directionally but not fully differentiable parameters established in hirano2012impossibility, the result in (ref) requires some regularity conditions on the system of limiting laws \(\{\calL_h\}_{h \in \SR^d}\).\footnote{Let \(F_0(\cdot)\) be the CDF associated with \(\calL_0\). hirano2012impossibility show that, for any \(\alpha\)-quantile unbiased estimate, it must be the case that either \(F_0(\cdot)\) is not differentiable at zero or must satisfy \(F_h'(0) = 0\).} The regularity condition in (ref)(b) implies that, if a locally asymptotically \(\alpha\)-quantile unbiased estimator were to exist, its associated limiting laws \(\calL_h\) must be able to be made arbitrarily flat. In particular, if each limiting law \(\calL_h\) has a density with respect to Lebesgue measure, these densities evaluated at zero, which is by definition the \(\alpha\)-quantile of each \(\calL_h\), must be able to be made arbitrarily small. As an example, suppose that \(\{\calL_h\}_{h \in \SR^d}\) is a family of Gaussian distributions on \(\SR\) associated with a locally asymptotically \(\alpha\)-quantile unbiased estimator. Then, the variance of these Gaussian distributions must be able to become arbitrarily large as \(h\) ranges over \(\SR^d\). \qed

Hypothesis Testing and Infinite Dimensional Models

The above analysis rules out standard approaches to inference on \(g(\theta)\) in local regions of \(\theta_\star\). These results are informative when one is interested in constructing confidence intervals for \(g(\theta)\) around points of first-order degeneracy when the data is drawn from a parametric model satisfying (ref). In this subsection, we consider two extensions of our results. In the first, we consider the somewhat simpler problem of testing the null hypothesis \(H_0: g(\theta) = g(\theta_\star)\). We show that if the null hypothesis contains a sufficiently rich set of values and a similar test exists, this similar test must have low power in local regions of \(\theta_\star\). In the second extension we generalize (ref) to infinite dimensional, i.e, semiparametric or nonparametric, models.

Hypothesis Testing

In this subsection we consider the problem of testing the null hypothesis \(H_0: g(\theta) = g(\theta_\star)\) where the alternative can be one sided, i.e, \(H_1: g(\theta) > g(\theta_\star)\) or two-sided, \(H_1: g(\theta) \neq g(\theta_\star)\). To setup, define \(\calH_\star\) to be the set of local parameters, \(h \in \SR^d\), such that \(g(\theta_{n,h})\) is asymptotically indistinguishable from \(g(\theta_\star)\), i.e, \(r_n^2(g(\theta_{n,h}) - g(\theta_\star)) \to 0\); \[ \calH_\star = \{h \in \SR^d: h'\nabla^2g(\theta_\star)h = 0\}. \] We study the behavior of asymptotically similar tests in local regions of degeneracy. In this setup, an asymptotically similar test is a statistic \(\Xi_n: \calX^n \to \{0,1\}\) such that \[ \limsup_{n\to\infty} P_{\theta_{n,h}}(\Xi_n = 1) = \alpha \quad \text{for all } h \in \calH_\star. \] Equivalently, the test is similar if \(\limsup_{n\to\infty }P_{\theta_{n,h}}(1 - \Xi_n \leq 0) = \alpha\). Letting \(\Psi_n = 1 - \Xi_n\) it is apparent that this is a nearly identical requirement to that of local \(\alpha\)-quantile unbiasedness in (ref), with the key difference being that the requirement only needs to hold for local parameters \(h \in \calH_\star\) rather than for all \(h \in \SR^d\).

However, unlike quantile unbiased estimation, which is ruled out in (ref), asymptotically similar tests can exist --- one can imagine constructing a similar test by flipping a weighted coin. Such a test, though, may not be powerful against local alternatives close to \(\theta_\star\). Our main result in this subsection establishes this formally: if such test exists then its local asymptotic power curve must be flat at \(\theta_\star\) in the sense that the derivative of the local asymptotic power curve with respect to the local parameter \(h\) exists and is equal to zero.

Define the local asymptotic power curve as

equation[equation omitted — 112 chars of source]
prop[] Let \(\Xi_n\) be an asymptotically similar test such that \(\calP(h)\) is differentiable at \(h = 0\). Then, the directional derivative of the local asymptotic power curve, \(\calP(h)\), in directions \(h \in \calH_\star\) is equal to zero: \[ D_h\calP(0) = 0\;\;\text{ for all }\; h \in \calH_\star. \] In particular, if \(\nabla^2 g(\theta_\star)\) is indefinite then \(\calH_\star\) spans \(\SR^d\) and \(\nabla\calP(0)\) is equal to zero.
remark[Differentiability of \(\calP\)] A common strategy in hypothesis testing is to compare a test statistic \(\Psi_n^\circ\) to a possibly data-dependent critical value \(\hat c_n\), rejecting when the former exceeds the latter. Let \(\Psi_n=\Psi_n^\circ-\hat c_n\), so the rejection rule can be written as \(\Xi_n=\bm{1}\{\Psi_n\geq 0\}\). If \(\Psi_n \overset{h}{\rightsquigarrow} \calL_h\) for each \(h \in \SR^d\), and if the CDF of \(\calL_0\) is continuous at zero, then the resulting local asymptotic power curve \(\calP\) is differentiable at \(0\); see (ref). This is a milder version of the regularity condition imposed on quantile unbiased estimators in hirano2012impossibility. \qed

The first statement in (ref) follows immediately from the definition of similarity along with the fact that \(\calH_\star\) is a cone: because \(\mathcal{P}(h)\) is constant on \(\calH_\star\), its directional derivatives in directions \(h \in \calH_\star\) must be zero. The force of the result, however, lies in the structure of \(\mathcal{H}^\star\) near points of degeneracy. In standard inference problems, i.e, when the true parameter is well separated from points of degeneracy, the parameter of interest in the limit experiment is a linear function of the shift parameter \(h\), so \(\mathcal{H}^\star = \{0\}\) and the zero-derivative condition carries no information about the shape of the power curve. By contrast, when \(g\) exhibits first-order degeneracy \(\mathcal{H}_\star\), can be a non-trivial cone — that is, it may contain directions other than zero — and the constraint \(D_h \mathcal{P}(0) = 0\) for \(h \in \calH_\star\) becomes substantive. When \(\nabla^2 g(\theta_\star)\) is indefinite, \(\calH_\star\) spans \(\SR^d\) and the entire gradient of the local asymptotic power curve vanishes at the origin, implying that power cannot increase at a linear rate in any direction away from \(\theta_\star\). The following examples illustrate the strength of this restriction.

Example (ref), cont. Consider again the mediation model, where the original model is given \((P_\theta: \theta \in \Theta \subseteq \SR^2)\). Suppose the researcher is interested in testing the null hypothesis \(H_0:\theta_1\theta_2 = 0\), that is \(g(\theta) = \theta_1\theta_2\) and \(\theta_\star = 0\). We can show that \(\calH_\star = \{h \in \SR^2: h_1h_2 = 0\}\). Since \(\calH_\star\) is the union of the two coordinate axes, \(\mathrm{span}(\calH_\star) = \SR^2\). Thus, we have that \(\nabla\calP(0) = 0\) for any asymptotically similar test.

We can equivalently show that \(\calH_\star\) must span \(\SR^2\) by noting that the second derivative matrix of \(g(\cdot)\) at \(\theta_\star\) is given by \[

pmatrix[pmatrix omitted — 27 chars of source]

, \] which is indefinite. \qed

example[Squared Mean] On the other hand consider the case where \(\theta = (\theta_1,\theta_2) \in \SR^2\) and the researcher is interested in testing the null hypothesis \(H_0: \theta_1^2 + \theta_2^2 = 0\). In this case \(g(\theta) = \theta_1^2 + \theta_2^2\), \(\theta_\star = 0\), and \(\calH_\star = \{(0,0)\}\). Since \(\nabla^2 g(\theta_\star)\) is positive definite, \(\mathrm{span}(\calH_\star) = \{0\}\). Thus, the results of (ref) do not apply and powerful similar tests for the null hypothesis \(H_0:\theta_1^2 + \theta_2^2 = 0\) can be constructed, see e.g chen2019improved. \qed
example[Standard Inference] Suppose that the primitive parameter is univariate \(\theta_{n,h} = \theta_0 + h/r_n \in \SR\), and local to a point \(\theta_0 \) such that \(g'(\theta_0) > 0\), that is we are well separated from points of degeneracy. In this setting, the researcher typically has access to an estimator \(\Psi_n\) that satisfies \(r_n(\Psi_n -g(\theta_{n,h}))\overset{h}{\rightsquigarrow}N(0,\sigma^2)\). This estimator is regular and thus \(\alpha\)-quantile-unbiased for all \(\alpha \in (0,1)\). Based on this estimator, an asymptotically similar one sided test for the null hypothesis, $H_0: g(\theta) = g(\theta_0)$, can be constructed with local asymptotic power curve \(\calP(h) = 1 - \Phi(c_{1-\alpha} - g'(\theta_0)h/\sigma)\), where \(\Phi(\cdot)\) is the standard normal CDF and \(c_{1-\alpha}\) is its \(1-\alpha\) quantile. Here, \(\frac{\partial }{\partial h}\calP(h)\big|_{h=0} = g'(\theta_0)\phi(c_{1-\alpha})/\sigma > 0\).
remark*[] Recent papers by van_garderen_nearly_2024 and Dufour2025WaldTW also study tests of the null hypothesis \(H_0:g(\theta) = g(\theta_\star)\) in various contexts. van_garderen_nearly_2024 consider the case of the mediation model, that is where \(\theta = (\theta_1, \theta_2)'\) and \(g(\theta) = \theta_1\theta_2\). They assume that the researcher has access to an asymptotically normal estimate of \(\theta\), \(\hat\theta = (\hat\theta_1, \hat\theta_2)'\) and show that there is no reasonable similar test of the form: reject if \(\max\{|\hat\theta_1|,|\hat\theta_2|\} > g(\min\{|\hat\theta_1|,|\hat\theta_2|\})\), where \(g(\cdot)\) may be an arbitrary function. Similarly, Dufour2025WaldTW consider the behavior of Wald type tests based on the test statistic \(W_n = \frac{g(\hat\theta) - g(\theta_\star)}{\nabla g(\hat\theta)'\Sigma\nabla g(\hat\theta)}\), where \(\Sigma\) represents the asymptotic variance of \(\hat\theta\). The authors show that, when \(\theta\) is close to \(\theta_\star\), the behavior of the Wald statistic can be irregular and propose alternate critical values for testing the null hypothesis \(H_0: g(\theta) = g(\theta_\star)\) using \(W_n\). We view our results as complementary to these existing results. The results in Proposition (ref) are narrower in their conclusion --- we establish only that similar tests must have flat power at $\theta^\star$ --- but broader in their scope, since they apply to any testing procedure that depends on the data, rather than only to procedures based on a specific initial estimator $\hat{\theta}$. \qed

Infinite Dimensional Models

In many settings the researcher may not be willing to assume that the data comes from a finite dimensional parametric model as described in the previous section. Following a tradition on studying semiparametric efficiency bickel1993efficient, we show that this does not affect our impossibility results so long as the larger model contains a parametric submodel satisfying (ref).

Formally, let the model \(\calP\) be a collection of sequences of probability measures on the sample space \(\calX^n\) from the previous section. That is, each element of \(\calP\) is a sequence of probability measures \(\{P_n\}\), where each probability measure \(P_n\) is defined on the sample space \(\calX^n\). A finite dimensional submodel, \(\calP_f\), is some smaller collection of sequences of probability measures that can be parameterized as \(\calP_{f} = (\{P_{n,\theta}\}_{n\in\SN}: \theta \in \Theta)\) for an open set \(\Theta \in \SR^{d_f}\). Fix a “centering” sequence of probability measures \(\{P_{0,n}\}_{n \in \SN} \in \calP\). We say that the submodel passes through \(\{P_{0,n}\}\) if \(\{P_{0,n}\} \in \calP_f\), that is \(\{P_{0,n}\} = \{P_{n,\theta}\}\) for some \(\theta \in \Theta\). We will call such a parametric model “regular” if (ref) holds and the model passes through \(\{P_{0,n}\}\).

We suppose that the object of interest is a quantity that depends on the sequence of underlying probability measures, that is we can think of the estimand \(g[\{P_n\}]\) as a functional defined on \(\calP\). For any regular parametric model, \(\calP_f\), this implicitly defines a function on \(\theta\) via the relation \(g_f(\theta) = g[\{P_{n,\theta}\}]\). With this notation defined, we show that the results of (ref) can be extended in a straightforward fashion to infinite dimensional models.

remark[Semiparametric Models with i.i.d Data] In the literature on semiparametric estimation with i.i.d data, where the researcher observes repeated observations drawn independently from a probability distribution \(P\) on \(\calX\) belonging to a model \(\mathcal{P}\) bickel1993efficient, one can associate the entire sequence of probability measures \(\{\otimes_{i=1}^n P: n \in \SN\}\) with the underlying common distribution \(P\). With this association, one can consider the model \(\calP\) described above as a collection of probability measures on \(\calX\) rather than a collection of sequences of probability measures \(\{P_n\}\) where each \(P_n\) is defined on \(\calX^n\). The parameter, in turn, can be defined as a function of the underlying distribution \(P\) rather than as a function of the entire sequence by \(g[P] \equiv g[\{\otimes_{i=1}^n P: n \in \SN\}]\). However, when dealing with time-series or network data, it may not be possible to define the parameter as a function of some representative underlying distribution and is instead a property of the sequence \(\{P_n\}\). \qed
corollary[Impossibility in Infinite-Dimensional Models] Suppose the data are generated from a sequence of distributions $\{P_{0,n}\} \in \mathcal{P}$. Let $\mathcal{P}_f \subset \mathcal{P}$ be a regular parametric submodel passing through $\{P_{0,n}\}$, and suppose that \(g_f\) satisfies (ref) with \(\{P_{n,\theta_\star}\} = \{P_{0,n}\}\). \begin{enumerate} • There is no estimator sequence $\Psi_n$ and law \(\calL\) on \(\SR\) such that, along every regular parametric submodel $\mathcal{P}_f$, \[ r_n^2\big( \Psi_n - g_f(\theta_\star + h/r_n) \big) \;\overset{h}{\rightsquigarrow}\; \calL \quad \text{for all } h \in \mathbb{R}^{d_f}. \] • Let $\{\calL_h\}_{h \in \mathbb{R}^{d_f}}$ be a family of distributions such that (i) $\calL_h\{(-\infty,0]\} = \alpha$ for some $\alpha \in (0,1)$ and all $h$, and (ii) each $\calL_h$ has a CDF, $F_h$, differentiable at zero with derivative bounded below by $\epsilon_f > 0$. Then there is no estimator sequence $\Psi_n$ such that, along every regular parametric submodel $\mathcal{P}_f$, \[ r_n^2\big( \Psi_n - g_f(\theta_\star + h/r_n) \big) \;\overset{h}{\rightsquigarrow}\; \calL_h \quad \text{for all } h \in \mathbb{R}^{d_f}. \] \end{enumerate}

Minimum Distance Based Inference

In this section, we construct a uniformly valid confidence interval for $g(\theta):\Theta\rightarrow\mathbb{R}$ using estimator $\hat{\theta}$. The confidence interval is obtained by inverting the hypothesis $H_{0}:g(\theta)=\tau$, and we use a minimum distance (MD) test statistic \[ \hat{T}_{n}(\tau)=\inf_{\theta\in\Theta:g(\theta)=\tau}r_{n}^{2}(\hat{\theta}-\theta)^{\prime}\Sigma^{-1}(\hat{\theta}-\theta). \] We focus on the settings where the standard first order approximation of $g(\theta)$ fails at $\theta_{\star}$, but the second order derivative is nondegenerate; that is, $\frac{\partial^{2}g}{\partial\theta\partial\theta^{\prime}}(\theta_{\star})=H$ with $H$ full rank.

In Section (ref), we discuss a simple case where $\Theta \subseteq \mathbb{R}^{2}$ and $H$ is indefinite, and we provide sufficient conditions under which the standard critical value $Q(\chi_{1}^{2},1-\alpha)$ is uniformly valid. In Section (ref), we propose a computationally simple method for a general $g$, which can be generalized to cases with higher order singularity.

Two-Dimensional $\theta$ and Indefinite $H$

To simplify notation, consider the null hypothesis

equation[equation omitted — 112 chars of source]

with $|\rho|<1$ and $\tau\geq0$. The restriction $|\rho|<1$ guarantees that $H$ is indefinite, while $\tau\geq0$ is a normalization. For simplicity, let $n=r_{n}=1$, and assume $\hat{\theta}-\theta\sim N(0,I_{2}).$ The quadratic form $g$ and the normality of $\hat{\theta}$ can be viewed as second order approximations, with general asymptotic results provided in Theorem (ref).

Let $X_{2}(\theta_{1})\in \mathbb{R}_+$ be the positive solution for $\theta_{2}$ such that ((ref)) holds. Let $\mathcal{S}_{0}(\tau)$ be the null parameter space, which contains two separate curves, $\mathcal{S}_{0}^{+}(\tau)$ and $\mathcal{S}_{0}^{-}(\tau)$,

align*[align* omitted — 93 chars of source]

where \[ \mathcal{S}_{0}^{+}(\tau)=\left\{ \left(x_{1},X_{2}(x_{1})\right):x_{1}\in\mathbb{R}\right\} ,\quad\mathcal{S}_{0}^{-}(\tau)=\left\{ \left(x_{1},-X_{2}(x_{1})\right):x_{1}\in\mathbb{R}\right\} . \] Let $\mathcal{S}(\tau,c)$ be the acceptance region with critical value $c^{2}$, i.e., the $c$-enlargement of $\mathcal{S}_{0}(\tau)$, \[ \mathcal{S}(\tau,c)=\left\{ (x_{1},x_{2}):(x_{1}-\theta_{1})^{2}+(x_{2}-\theta_{2})^{2}\leq c^{2},(\theta_{1},\theta_{2})\in\mathcal{S}_{0}(\tau)\right\} . \]

propLet $c=\sqrt{Q(\chi_{1}^{2},1-\alpha)}$ and $\hat{\theta}-\theta\sim N(0,I_{2})$. Suppose either $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}\leq\frac{1}{c}$ or $\rho\geq0$. For all $\theta\in\mathcal{S}_{0}(\tau)$, it holds that \[ P\left(\hat{\theta}\in\mathcal{S}(\tau,c)\right)\geq1-\alpha. \]

Proposition (ref) shows that the standard MD test remains valid under a curved null hypothesis when the maximum curvature of $\mathcal{S}_{0}^+(\tau)$, given by $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}$, is sufficiently small, or when the two branches $\mathcal{S}_{0}^{+}(\tau)$ and $\mathcal{S}_{0}^{-}(\tau)$ are sufficiently close. The argument proceeds by comparing the coverage of $\mathcal{S}(\tau,c)$ with that of the auxiliary acceptance set, \[\mathcal{S}_{\text{aux}}=\left\{ (x_{1},x_{2}):(x_{2}-\theta_{2})^{2}\leq c^{2}\right\} .\] whose coverage is exactly $1-\alpha$. Expressed in polar coordinates, the coverage depends on the fraction of each circle of radius $r$ centered at $\theta\in \mathcal{S}^{+}_0(\tau)$, denoted $\partial B(\theta,r)$, that is contained in the acceptance region. Consequently, it suffices to show that, for each $r$, the arc length of $\partial B(\theta,r)$ contained in $\mathcal{S}(\tau,c)$ is no smaller than that of $\mathcal{S}_{aux}$.

When $r\leq c$, the entire circle is covered by both $\mathcal{S}_{aux}$ and $\mathcal{S}(\tau,c)$ by construction. For $r>c$, let $\mathcal{C}_{u}(\tau)$ and $\mathcal{C}_{\ell}(\tau)$ denote the upper and lower boundaries of $\mathcal{S}^{+}(\tau,c)$; see Figure (ref) left panel. If $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}\leq\frac{1}{c}$, the circle $\partial B\left(\theta,r\right)$ intersects $\mathcal{C}_{u}(\tau)$ at points $A$ and $B$, and $\mathcal{C}_{\ell}(\tau)$ at points $C$ and $D$. We can show that the lengths of chords $\overline{AC}$ and $\overline{BD}$ are no smaller than $2c$. Otherwise, $B(G,c)\not\subseteq\mathcal{S}^{+}(\tau,c)$, where $G$ denotes the intersection of $AC$ with $\mathcal{S}_{0}^{+}(\tau)$, contradicting the definition of $\mathcal{S}^{+}(\tau,c)$. Note that the arcs of $\partial B(\theta,c)$ covered by $\mathcal{S}_{aux}$ correspond to chords of length $2c$. It therefore follows that the portion of $\partial B(\theta,c)$ covered by $\mathcal{S}(\tau,c)$ is larger than that covered by $\mathcal{S}_{aux}$.

figure[figure omitted — 648 chars of source]

If $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}>\frac{1}{c}$, $\mathcal{C}_{u}(\tau)$ has a kink due to the high curvature. Thus, there exist $\theta\in\mathcal{S}_{0}^{+}(\tau)$ and $r>c$ such that $\partial B\left(\theta,r\right)$ does not intersect $\mathcal{C}_{u}(\tau)$; see Figure (ref) right panel. For such $(\theta,r)$, the argument based on chord lengths no longer applies. However, when $\rho\geq0$, $\partial B(\theta,r)$ is sufficiently close to $\mathcal{S}_{0}^{-}(\tau)$, and we can show that $\partial B\left((\theta_{1},\theta_{2}),r\right)\subseteq\mathcal{S}(\tau,c)$.

Next, we present the asymptotic results for general data generating processes.

assumptionSuppose that \begin{enumerate} • \(\nabla^2 g(\theta_\star)\) has full rank. • Let $BL_{1}$ denote the set of Lipschitz functions which are bounded by $1$ in absolute value and have Lipschitz constant bounded by $1$. Assume there exists $r_{n}\rightarrow\infty$ such that \[ \lim_{n\rightarrow\infty}\sup_{P\in\mathcal{P}}\sup_{f\in BL_{1}}\left|E_{P}\left[f\left(\sqrt{r_{n}}\left(\hat{\theta}-\theta_{P}\right)\right)\right]-E_{P}\left[f\left(\xi_{P}\right)\right]\right|=0, \] where $\xi_{P}\sim N(0,\Sigma_{P})$. • Let $\mathcal{S}$ denote the set of matrices with eigenvalues bounded below by $\underline{e}>0$ and above by $\bar{e}\geq\underline{e}$. For all $P\in\mathcal{P}$, $\Sigma_{P}\in\mathcal{S}$. • For all $\varepsilon>0$, \[ \lim_{n\rightarrow\infty}\sup_{P\in\mathcal{P}}P\left(\left\Vert \hat{\Sigma}-\Sigma_{P}\right\Vert >\varepsilon\right)=0. \] \end{enumerate}

(ref).(ref) assumes that the second order derivative of \(g\) at \(\theta_\star\) is of full rank, so that there is no higher order degeneracy. (ref).(ref) requires that the researcher has access to an estimator \(\hat\theta\) that is uniformly asymptotically normal over the class of DGPs considered while (ref).(ref) and (ref).(ref) require that the asymptotic variance of this estimator is well-behaved and consistently estimable. Since degeneracy is a property of the transformation of interest, $g$, rather than of the primitive parameter $\theta$, these are mild conditions that can be verified for most common estimators $\hat\theta$.

thmSuppose $d=2$, and let $c=\sqrt{Q(\chi_{1}^{2},1-\alpha)}$. Let $(\lambda_{P,1},\lambda_{P,2})$ be the eigenvalues of $\text{sign}\left(g(\theta_{P})-g(\theta_{\star})\right)\Sigma_{P}^{1/2}H\Sigma_{P}^{1/2}$, and define $\rho_{P}=\frac{\lambda_{P,1}+\lambda_{P,2}}{\left|\lambda_{P,1}-\lambda_{P,2}\right|}.$ Assume that Assumptions (ref) and (ref) hold. If for some $\eta>0$, it holds that either \begin{equation} \mathcal{P}_{n}\subseteq\left\{ P\in\mathcal{P}:\frac{(1-\rho_{P})\sqrt{\left|\lambda_{P,1}-\lambda_{P,2}\right|}}{2r_{n}\sqrt{\left|g(\theta_{P})-g(\theta_{\star})\right|(1+\rho_{P})}}\leq\frac{1}{c},\;\rho_P\in[\eta-1,1-\eta]\right\} \end{equation} or \begin{equation} \mathcal{P}_{n}\subseteq\left\{ P\in\mathcal{P}:\rho_{P}\in[0,1-\eta]\right\}, \end{equation} then \[ \liminf_{n}\inf_{P\in\mathcal{P}_{n}}P\left(\hat{T}_{n}\left(g(\theta_{P})\right)\leq c^{2}\right)\geq1-\alpha. \]

Theorem (ref) follows from Proposition (ref). To see this, let $\tilde{g}(h) = g\!\left(\theta_{\star} + r_{n}^{-1}\Sigma^{1/2} h\right),$ where \( r_{n} \) governs the rate at which \( \theta \) is estimated and \( \Sigma \) adjusts for the covariance. For \( h = O(1) \), \( \tilde{g}(h) \) can be approximated by a hyperbola, as in (ref). Condition (ref) ensures that the curvature of \( \tilde{g}(h) \) is not too large, while (ref) implies that the two branches of the hyperbola are sufficiently close. The result remains uniformly valid even as \( \|h\| \to \infty \).

remarkTo illustrate Theorem (ref), consider $g(\theta)=\theta_{1}\theta_{2}$, as motivated by Example (ref). In this case, $H=\left[\begin{array}{cc} 0 & 1\\ 1 & 0 \end{array}\right]$. With $\Sigma_{P}=\left[\begin{array}{cc} \sigma_{1}^{2} & r\sigma_{1}\sigma_{2}\\ r\sigma_{1}\sigma_{2} & \sigma_{2}^{2} \end{array}\right]$, we have $\lambda_{P,1}=\text{sign}\left(g(\theta_{P})\right)\left(r-1\right)\sigma_{1}\sigma_{2}$, $\lambda_{P,2}=\text{sign}\left(g(\theta_{P})\right)\left(r+1\right)\sigma_{1}\sigma_{2}$, and $\rho_{P}=\text{sign}\left(g(\theta_{P})\right)r$. Therefore, ((ref)) holds when $r=0$, and the MD test with the simple critical value yields a uniformly valid confidence interval for the mediation effect. It is worth noting that, the rejection region for $\theta_{1}\theta_{2}=0$ is given by $\min\left\{ |\hat{\theta}_{1}|,|\hat{\theta}_{2}|\right\} >Q(\chi_{1}^{2},1-\alpha)$, which coincides with the rejection region of the likelihood ratio test. The latter is the uniformly most powerful invariant test among information- and size-coherent tests (van2022optimality). If $\text{sign}\left(g(\theta_{P})\right)r<0$, then ((ref)) is satisfied when $\frac{\sqrt{\sigma_{1}\sigma_{2}}(1+|r|)}{ r_n\sqrt{2(1-|r|)\left|g(\theta_{P})\right|}}\leq\frac{1}{c}.$ Since $\sigma_1$, $\sigma_2$ and $r$ can be consistently estimated, and $g(\theta_P)$ is known under $H_0$, conditions ((ref)) and ((ref)) are straightforward to verify in practice.

General Cases

In this section, we present the inference procedure for a general function $g$. The procedure is based on a local approximation of the test statistic. First, consider the case where the true parameter value $\theta_{n}$ satisfies $\theta_{n}=\theta_{\star}+h_{n}/r_{n}$. Under $H_0$, the test statistic is given by\\ \[ \hat{T}_{n}(g(\theta_n))=\inf_{\vartheta:g(\vartheta)=g(\theta_{n})}r_{n}^{2}\left(\hat{\theta}-\vartheta\right)^{\prime}\hat{\Sigma}^{-1}\left(\hat{\theta}-\vartheta\right)=r_{n}^{2}\left\Vert \hat{\Sigma}^{-1/2}(\hat{\theta}-\tilde{\theta}_{n})\right\Vert ^{2} \] where $\tilde{\theta}_{n}$ denotes the minimizer. Since $\hat{T}_n(g(\theta_n))\leq r_{n}^{2}\left\Vert \hat{\Sigma}^{-1/2}(\hat{\theta}-\theta_{n})\right\Vert ^{2}=O_p (1)$, we have $\tilde{\theta}_{n}=\theta_{n}+O_{p}(\frac{1}{r_{n}}).$ If $h_n = O(1)$, a second order Taylor expansion of $g(\tilde{\theta}_{n})=g(\theta_{n})$ gives

align*[align* omitted — 137 chars of source]

In addition, let $\mathbb{Z}_{n}=r_{n}\hat{\Sigma}^{-1/2}(\hat{\theta}-\theta_{n})$, we can write

align*[align* omitted — 302 chars of source]

In sum, let $t=r_n(\tilde{\theta}_n-\theta_\star)$, given $h_{n}$, we can approximate $\hat{T}_{n}$ by

equation[equation omitted — 160 chars of source]

where $\mathbb{Z}|\hat{T}_{n}(g(\theta_n))\sim N(0,I_{d})$. We can show that $\hat{T}_{n}(g(\theta_n))$ and $\hat{T}_{n}^{*}(h_{n})$ have the same asymptotic distribution, regardless of whether $h_{n}$ converges to $h\in\mathbb{R}$ or diverges to infinity (Lemmas (ref) and (ref)). Intuitively, if $h_{n}\rightarrow\infty$, the restriction for the optimizer $\tilde{t}$ in ((ref)) is approximately linear. In this case, both $\hat{T}_{n}(g(\theta_n))$ and $\hat{T}_{n}^{*}(h_{n})$ are approximated $\chi_{1}^{2}$.

Given $h_{n}$, we can easily get the quantile of $\hat{T}_{n}^{*}(h_n)$ by simulation. However, $h_{n}$ is a nuisance parameter that cannot be consistently estimated. Next, we propose a two step feasible critical value. Suppose set $\mathcal{H}_{z}$ satisfies $P\left(N(0,I_{d})\in\mathcal{H}_{z}\right)=1-\eta$.\footnote{For instance, $\mathcal{H}_z=\{z\in \mathbb{R}^d:z^\prime z\leq Q(\chi^2_d,1-\eta)\}.$} In the first step, we construct a $(1-\eta)$ confidence set for $h_{n}$,

equation[equation omitted — 116 chars of source]

In the second step, we construct the critical value based on the $\frac{1-\alpha}{1-\eta}$ quantile of $\hat{T}_{n}^{*}(h)$ conditional on the first step. That is, let

align[align omitted — 179 chars of source]

and reject $H_{0}:g(\theta)=\tau$ if $\hat{T}_{n}(\tau)>\hat{c}$. In ((ref)), the construction of $\hat{c}$ takes into account the first step selection, thus it is less conservative than simple Bonferroni correction.

thmUnder Assumptions (ref) and (ref), it holds that \[ \liminf_{n}\inf_{P\in\mathcal{P}}P\left(\hat{T}_{n}(g(\theta_{P_n}))\leq\hat{c}\right)\geq1-\alpha. \] In addition, if $\left\Vert r_{n}(\theta_{P_{n}}-\theta_{\star})\right\Vert \rightarrow\infty$, \begin{equation} \lim_{n}P_{n}\left(\hat{T}_{n}(g(\theta_{P_n}))\leq\hat{c}\right)\in\left[1-\alpha,1-\alpha+\eta\right). \end{equation}

The slight conservativeness arises from the two-step procedure. Alternatively, we can introduce a pretest to check whether $h_{n}$ is far away from zero, e.g. $||h_{n}||>\ln r_{n}$. If so, we can use the standard critical value $Q(\chi_{1}^{2},1-\alpha)$. The cost is that we need to introduce an extra tuning parameter.

remarkIn general, if $H$ is singular and $g$ is higher order identified, we can construct the critical value using a similar two step procedure. In the first step, we construct a $1-\eta$ confidence set $\hat{\Theta}$ for $\theta$. In the second step, we define the critical value as \[ \hat{c}=\sup_{\theta\in\hat{\Theta}}Q\left(\inf_{\vartheta:g(\vartheta)=g(\theta)}\left\Vert \mathbb{Z}+r_{n}\hat{\Sigma}^{-1/2}(\theta-\vartheta)\right\Vert ^{2};1-\alpha+\eta\right). \] $H_{0}:g(\theta)=\tau$ is rejected if $\hat{T}_{n}(\tau)>\hat{c}$. \qed
remarkDufour2025WaldTW show that when $g$ is a vector-valued function and the degree of singularity differs across elements of $g$, the Wald-type test statistic may diverge, complicating inference. In contrast, the MD test considered in this paper yields a test statistic that is first-order stochastically dominated by $\chi_{d}^{2}$, regardless of the level of singularity in $g$. Moreover, Dufour2025WaldTW focus solely on hypothesis tests at a fixed point, i.e., testing $g(\theta)=g(\theta_{\star})$, whereas this paper aims to construct uniformly valid confidence intervals. \qed
remarkAM-2016-Geometric construct a uniformly valid MD test based on a geometric approach that incorporates the curvature of the null restriction $g(\theta)=\tau$. When the curvature is large, their procedure may yield overly conservative critical values. For example, consider the mediation analysis problem in (ref) where one is interested in testing the null hypothesis \(H_0:\theta_1\theta_2 = \tau\). As \(\tau\) approaches zero, the curvature of the null manifold can be made arbitrarily large and the critical value of AM-2016-Geometric approaches $Q(\chi_{2}^{2},1-\alpha)$. However, as shown in Section (ref) of this paper, a uniformly valid critical value in this setting is $Q(\chi_{1}^{2},1-\alpha)$. Indeed, even when \(\tau\) is far from zero, the AM-2016-Geometric critical value is always larger than \(Q(\chi_1^2, 1 - \alpha)\). \qed

Simulation

In this section, we examine the size and power properties of the proposed procedures and compare them with several alternatives. We focus on the context of Example (ref), namely the construction of confidence intervals for the mediation effect. In addition to the two MD-based methods proposed in Section (ref), one using the $Q(\chi^2_1,1-\alpha)$ critical value (BN1; Section (ref)) and one using a bootstrapped critical value (BN2; Section (ref)), we consider two uniformly valid MD-based alternatives: (i) the procedure of AM-2016-Geometric (AM),\footnote{We report results using their Section 4.1 implementation, which computes curvature over a restricted set with tuning parameter $\eta=\alpha/10$. We also implemented their worst-case curvature procedure from Section 2. The two procedures have nearly identical power, with the latter performing slightly worse.} and (ii) the MD-based method with projection critical value $Q(\chi_{2}^{2},1-\alpha)$. For comparison, we also include the naive delta method, i.e. a Wald-type test with critical value $Q(\chi_{1}^{2},1-\alpha)$, and a naive bootstrap method. The nominal rejection rate is $\alpha=0.05$, and the tuning parameter for BN2 is $\eta=\alpha/10$.

We study confidence intervals for $g(\theta)=\theta_{1}\theta_{2}$, where the estimators are simulated from \[ \hat{\theta}-\theta\sim\mathcal{N}\left(0,\left[

array[array omitted — 30 chars of source]

\right]\right). \] Without loss of generality, we normalize the variance of $\hat{\theta}$ to one. We consider $r=0$ and $r=0.5$, with $\theta_{2}\in\{2,6\}$ and $\theta_{1}=[-1:0.2:1]\times\theta_{2}$.

In Figure (ref), we plot the probability that the confidence intervals exclude the true value $g(\theta)$. The naive method does not overreject when $\theta_{1}\theta_{2}=0$, but its rejection probability is very low near the origin, consistent with earlier findings (e.g., Dufour2025WaldTW). Away from the origin (see, e.g., $\theta_1=\theta_{2}=2$), the naive Wald test overrejects. According to Remark (ref), BN1 is valid for $r=0$. For $r=0.5$, Theorem (ref) does not guarantee validity when $\theta_{1}<0$, $\theta_{2}=2$, or when $\theta_1\in[-1.4,0)$, $\theta_2 = 6$. Nevertheless, BN1 maintains correct size across all designs, even when these conditions fail, suggesting that the condition is sufficient but not necessary. As expected, all other MD-based methods control size.

Figure (ref) shows the probability that the confidence intervals exclude zero, i.e., the probability of obtaining a significant result. When $\theta$ is close to the origin ($\theta_2=2$), our methods have substantially higher power than AM, whose performance is close to that of the simple projection method. When $\theta$ is further from the origin ($\theta_{2}=6$), power curves across methods are nearly identical.

Finally, Figure (ref) reports the median length of the confidence intervals, computed across $S$ replications. BN1 consistently yields the shortest intervals, with BN2 close behind. The projection method is the most conservative, producing intervals 19--30% longer than BN1. AM lies between BN1 and the projection method, with median lengths 5--18% longer than those of BN1. The differences are most pronounced when $\theta$ is near the origin.

center[center omitted — 1,565 chars of source]

Empirical Application

We illustrate the empirical relevance of our results using the setting analyzed by AEM-2018. Their study takes advantage of a distinctive feature of the Turkish education system, in which elementary school teachers are randomly allocated across schools. This institutional detail generates plausibly exogenous variation in teacher characteristics that can be used to study how teachers' gender role attitudes influence student outcomes. The data include roughly 4,000 third- and fourth-grade students taught by 145 teachers, and students can be grouped according to the length of their exposure to a given teacher --- at most one year, two to three years, or up to four years. The treatment variable is whether a teacher is identified as holding traditional rather than progressive gender beliefs, while the mediator of interest is the student's own gender role beliefs. Following a similar analysis of this data in van_garderen_nearly_2024, we focus on verbal test scores as the outcome. AEM-2018 argue that, after controlling for an extensive set of student, family, teacher, and school characteristics, the identifying assumptions for causal mediation analysis are satisfied in this context.

table[table omitted — 634 chars of source]

(ref) reports estimates from the AEM-2018 analysis linking teachers’ gender role attitudes to students’ verbal test performance. The first coefficient, $\hat{\theta}_1$, comes from a regression of students’ gender role beliefs on the gender role attitudes of their teachers, with the standard set of student, family, teacher, and school controls included. This coefficient summarizes the extent to which traditional teachers transmit their views to students. The second coefficient, $\hat{\theta}_2$, is estimated from a regression of test scores on both student gender beliefs and teacher attitudes, again with the full set of controls. It reflects how student beliefs are associated with verbal performance once teacher attitudes are held constant. Multiplying these two coefficients gives the mediated, or indirect, effect: the part of the teacher’s influence on scores that operates through the channel of student beliefs. The estimates show that this indirect pathway is negative and relatively small, although it varies across exposure groups, being largest in the one-year sample and smallest for students exposed for four years.

Because the true mediation, or indirect, effect appears to be close to zero, the results of (ref) suggest that standard approaches to inference will fail. In particular, we cannot construct valid confidence intervals via the typical approach of inverting a \(t\)-test. Instead, we construct confidence intervals using the newly proposed methods of (ref). van_garderen_nearly_2024 show that the correlation coefficient between $\hat{\theta}_1$ and $\hat{\theta}_2$ is zero; consequently, both of our methods are uniformly valid.\footnote{The validity of BN1 follows from Remark (ref).}

table[table omitted — 2,634 chars of source]

We compare our confidence intervals to two other inference procedures that might be applied in this setting, both of which are based on the minimum distance statistic. The first alternate procedure is that of AM-2016-Geometric.\footnote{In implementing the test, we follow the empirical application in the working paper version of AM-2016-Geometric and only calculate the maximum curvature over a set “close” to the point estimate, adjusting the critical value accordingly.} This testing procedure technically does not cover the case where we are testing the null that the mediation effect is equal to zero since the null manifold is not smooth in this case. However, the AM-2016-Geometric critical value approaches \(Q(\chi^2_2, 1-\alpha)\) from below as the null hypothesis value approaches zero and \(Q(\chi^2_2, 1 - \alpha)\) is a valid critical value for testing the null that the mediation effect is equal to zero so we simply modify the procedure slightly to directly use a \(Q(\chi^2_2, 1 - \alpha)\) when the null value is equal to zero. The second method, “Projection”, simply uses the \(Q(\chi^2_2, 1 -\alpha)\) at all points, which is justified since, under the null hypothesis, the distance to the null manifold is always less than the distance to the point \((\theta_1,\theta_2)'\).

Consistent with the discussion in (ref), the confidence intervals based on either the $\chi^2_{1}$ critical value (BN1) or the two-step procedure (BN2) are uniformly tighter than those obtained from the AM-2016-Geometric simulated critical value (AM); in all specifications, our intervals are strict subsets of theirs. The difference is not only theoretical but also empirically relevant. Using the $\chi^2_{1}$ critical value, for instance, the BN1 confidence interval supports the conclusion of van_garderen_nearly_2024 that the mediation effect of a one-year exposure to a teacher with traditional views is negative, whereas the alternative methods cannot reject a null of zero at the five-percent level. As expected, the AM intervals lie strictly inside those generated by the projection method, which uses a \(\chi^2_2\) critical value at all points. However, because the true mediation effect appears small in this setting, their simulated critical value converges toward $Q(\chi^2_2, 1 -\alpha)$, which accounts for the close similarity between the two sets of intervals.

Conclusion

We examine inference in local regions of first-order degeneracy, meaning that the gradient of the transformation is zero or nearly zero so that first-order approximations alone do not provide reliable information and second-order terms must also be considered. In such regions of local degeneracy, we show that neither regular estimation nor quantile-unbiased procedures are feasible, paralleling impossibility results for nondifferentiable functionals and ruling out standard approaches to inference. We then develop alternate inference procedures based on minimum-distance statistics that deliver uniformly valid confidence intervals. Simulation studies indicate that these procedures control size while maintaining favorable power, and the empirical application to teacher gender attitudes shows that they yield tighter confidence intervals than existing approaches.