Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
86,900 characters · 15 sections · 81 citation commands
Inference under First-Order Degeneracy
\thispagestyle{empty}
The delta method is a fundamental tool in econometric analysis for deriving the asymptotic distribution of smooth functions of estimators. In its standard form, the delta method relies on a non-zero gradient of the smooth function with respect to the primitive parameter. When this condition holds, first-order linearization provides an accurate approximation to the sampling distribution. However, in many empirically relevant scenarios, the gradient may be zero at certain points in which case higher-order terms become essential. As a result, the limiting distribution near points of degeneracy typically differs substantially from that obtained under standard regularity. A leading example arises in causal mediation analysis, where the indirect effect is the product of two primitive parameters: the effect of the treatment on the mediator and the effect of the mediator on the outcome. When these effects are zero, the gradient of the indirect effect degenerates, leading to a nonstandard limiting distribution of the Wald statistic sobel1982asymptotic.
In practice, the researcher does not know whether the true parameter lies near such points of degeneracy, and thus cannot know which asymptotic approximation is appropriate for inference. This challenge has motivated a number of recent papers on hypothesis testing when the gradient may be degenerate, see, for example, van_garderen_nearly_2024 for an analysis of the mediation model mentioned above and Dufour2025WaldTW in a more general treatment of Wald-type statistics. These papers acknowledge discontinuities in limiting distributions at the point of degeneracy, and analyze the resulting distortions in Wald-type statistics. Yet, they do not propose a unified asymptotic framework for studying local regions of degeneracy. Moreover, most existing papers are interested in specific point hypotheses, leaving open the broader question of how to construct uniformly valid confidence intervals.
This paper makes two main contributions to the literature. Our first contribution is a formal asymptotic framework for studying the behavior of statistics in local regions of first-order degeneracy. In our framework, the primitive parameter is modeled as local to the point of degeneracy, in the spirit of weak identification asymptotics, where identifying information is local to zero. Under this setup, the behavior of simple plug-in estimators becomes nonstandard and depends on local parameters that cannot be consistently estimated, formally capturing an observation in the existing literature that the standard delta method does not properly approximate the behavior of plug-in estimators near degenerate points miller2024testing.
Leveraging Le Cam's limit of experiments framework lecam1970assumptions,le1972limits, we show that the problem of estimating a smooth function under degeneracy is asymptotically equivalent to estimating a quadratic form of the shift parameter in a Gaussian shift model. Within this model, we show that equivariant-in-law and quantile-unbiased estimators cannot be obtained. Translated back to the original estimation problem, these results imply that regular estimation is impossible --- the limiting behavior of any properly scaled and centered estimator in local regions of degeneracy depends on nuisance parameters that cannot be consistently estimated. Moreover, there do not exist asymptotically similar confidence intervals for the transformation of interest in local regions of degeneracy.
Our second contribution is to construct confidence intervals that are both uniformly valid and exhibit favorable power in local regions of degeneracy compared to the few existing alternatives. The impossibility results mentioned above imply that standard delta-method based confidence intervals may not be uniformly valid in such regions. Indeed, in the context of mediation analysis our simulation study shows that standard Wald-statistic based tests lead to confidence intervals that undercover when the true primitive parameter is near the origin. We thus take a different approach and propose confidence intervals based on test inversion with a minimum-distance test statistic. We first show that when the parameter dimension is two, the standard chi-square critical value is uniformly valid under either of two conditions: (i) the curvature of the null curve is not too large, or (ii) the two branches of the null curve are sufficiently close. These conditions hold in the leading mediation example. In more general settings, we propose a bootstrap critical value based on a quadratic approximation of the test statistic, and show that when the true parameter is well separated from the point of degeneracy, the bootstrap critical value is nearly identical to the efficient one.
We demonstrate the empirical relevance of our results in both simulation study and real world. In simulation study our proposed methods are shown to control size uniformly over the parameter space while standard Wald-based inference can overreject when the true primitive parameter is close to points of degeneracy, a finding also noted by Dufour2025WaldTW. We additionally demonstrate favorable power properties of our proposed methods when compared to a method proposed by AM-2016-Geometric, which turns out to also be applicable in this setting.\footnote{In their original analysis, AM-2016-Geometric were interested in testing non-linear restrictions in the context of weak identification.} These improvements in power are also seen in an application to the data of AEM-2018, who consider the effect of teachers' gender attitudes on student outcomes. In revising their mediation analysis, we find that our proposed methods deliver tighter confidence intervals than existing methods for all parameters. Our results support a conclusion by van_garderen_nearly_2024 that the mediation effect of a one-year exposure to a teacher with traditional negative views is negative, while existing methods cannot rule out an effect of size zero.
The rest of this paper proceeds as follows. This section concludes with a review of the related literature. (ref) gives examples of when first-order degeneracy may be a concern. (ref) formally establishes the impossibility results mentioned above and discusses implications for hypothesis testing. (ref) introduces the minimum distance based inference procedures and discusses how uniformly valid critical values may be constructed. (ref) contain, respectively, the simulation study and empirical application to the data of AEM-2018. (ref) concludes. Proofs are deferred to (ref).
Our paper is related to previous literature on econometrics and statistics studying inference under degeneracy, statistical impossibility results, and testing non-linear restrictions.
There is a growing literature examining hypothesis tests in which the null includes points of singularity. gaffke1999asymptotic show that the distribution of the Wald statistic at points of degeneracy is nonstandard, and gaffke2002asymptotic derive its asymptotic distribution under a variety of singular null hypotheses. drton2016wald demonstrate the conservativeness of the Wald test at degeneracy points for quadratic forms and for bivariate monomials of arbitrary degree. dufour2016rank propose rank-robust regularized Wald-type tests allowing for singular covariance matrices. Dufour2025WaldTW analyze Wald tests for polynomial restrictions with possibly multiple constraints and show that such tests can under-reject, over-reject, or even diverge under the null; see also Dufour2025WaldTW for additional references in this area. Our paper differs from this literature in two key ways. First, prior work focuses on testing problems where the null itself contains the singularity, while our interest lies in constructing uniformly valid confidence intervals when the null may be near a singularity. In simulations, we show that the upper bound on the Wald statistic derived in Dufour2025WaldTW does not, in general, yield valid confidence intervals. Second, instead of using a Wald statistic, we employ a minimum distance–based test statistic, which is bounded in probability by construction. This approach avoids the divergence issues documented in Dufour2025WaldTW.
Our paper is also related to the hypothesis testing problem with a curved null. AM-2016-Geometric study this problem and show that the distribution of minimum-distance statistics is dominated by a tractable distribution that depends only on the maximal curvature of the null manifold relative to the known variance matrix. Inverting their test leads to uniformly valid confidence intervals. However, when the curvature of the null hypothesis is large, for example, in testing the significance of an indirect effect, their procedure yields critical values that are close to those from projection-based methods. By contrast, our procedure exploits the possibility that the null hypothesis may include multiple manifolds that are close to one another, which in turn reduces the critical value.
Finally, our paper contributes to the econometric literature on statistical impossibility results. In particular, it is related to work by hirano2012impossibility who show that regular estimation of directionally, but not fully, differentiable functions is unattainable. Our paper takes a similar approach to that of hirano2012impossibility in that we rule out properties of estimators by analyzing a limiting experiment lecam1970assumptions,le1972limits. However, the target functional in our limit experiment is distinct from that of hirano2012impossibility. This approach of ruling out behaviors by analyzing limit experiments has also been utilized successfully by kaji2021theory and andrewsgmm in the study of weak identification. The present work is also related to work by chen2019inference who show that all standard bootstrap procedures necessarily fail at points of degeneracy, similarly to how fang2019inference establish that the bootstrap necessarily fails as an inference procedure for the functionals considered in hirano2012impossibility.
Consider a parameter \(\theta \in \Theta \subseteq \SR^d\) and a twice continuously differentiable function \(g: \Theta \to \SR\). We are interested in inference on \(g(\theta)\) in local neighborhoods of a point \(\theta_\star\) for which \(\nabla g(\theta_\star) = 0\). Below, we give some empirically relevant examples of when such a phenomenon may occur.
In this section we establish that standard approaches to inference on $g(\theta)$ necessarily fail in local regions of first-order degeneracy. We begin in (ref)\textcolor{pastelred}{.1} by introducing a parametric framework to study this problem and defining what it means for an estimator to be “regular” in this setting. (ref)\textcolor{pastelred}{.2} then uses a representation theorem to show that the problem reduces to estimation of quadratic forms in a Gaussian shift experiment, where we prove that well-behaved estimators cannot be constructed. (ref)\textcolor{pastelred}{.3} extends the analysis in two directions: first, to hypothesis testing problems where the null hypothesis that \(g(\theta) = g(\theta_\star)\) may hold on a nontrivial subset of the parameter space, and second, to infinite-dimensional models, where we show that the impossibility results remain valid so long as the model contains a suitable parametric submodel. Together, these results demonstrate that standard approaches to inference necessarily break down in local regions of degeneracy.
We begin by assuming that the researcher observes data \(X^{(n)} = (X_1,\dots,X_n)\) drawn from a parametric model \(P_{n,\theta}\),
where \(\theta \in \Theta \subset \SR^d\) and \(\Theta\) is a compact set with a nonempty interior, \(\Theta^\circ \neq \emptyset\). Let \(\calX_i\) denote the support of \(X_i\), which could be a general space, and denote \(\calX^n = \bigtimes_{i=1}^n \calX_i\). We assume that the sequence of statistical models \(\left(P_{n,\theta}: \theta \in \Theta^\circ\right)\), indexed by the sample size \(n\), is locally asymptotically normal in the sense of lecam1960lan.
We study a twice continuously differentiable scalar functional \(g:\Theta \to \SR\), and the behavior of estimators of \(g(\theta)\) in local regions of a point \(\theta_\star \in \Theta^\circ\) which is such that the first-order derivatives of \(g(\cdot)\) at \(\theta_\star\) are zero. We will refer to \(\theta_\star\) as the “point of degeneracy” and local neighborhoods of \(\theta_\star\) as “local regions of degeneracy.”
Given the maintained assumption of a locally asymptotically normal model, it is useful to examine regions close to \(\theta_\star\) by adopting a local parameterization around \(\theta_\star\), defining \[ \theta_{n, h} = \theta_\star + h/r_n \] and letting \(P_{n,h} = P_{n, \theta_\star + h/r_n}\). In our framework an estimator is an arbitrary measurable function of the data, \(\Psi_n: \calX^n \to \SR\). We consider sequences of estimators, \(\Psi_n\), that converge in distribution under every sequence of alternative distributions \(P_{n,h}\) to some limiting law, \(\calL_h\). This is denoted
where we note that the convergence rate is \(r_n^2\) instead of \(r_n\) due to the fact that \(g(\theta)\) is “flat” around \(\theta_\star\). It is straightforward to show that tests for \(g(\theta)\) based on estimators whose convergence rates are slower than \(r_n^2\) when \(\theta\) is close to \(\theta_\star\) have trivial power against local alternatives of the form \(g(\theta_\star + h/r_n)\).
We focus on ruling out regular and locally asymptotically \(\alpha\)-quantile unbiased estimation of \(g(\theta_{n,h})\) in local regions of \(\theta_\star\). Formally, these notions are defined as follows.
The existence of regular estimators is closely tied to the validity of Wald-type inference procedures --- without regular estimators, standard Wald-type inference procedures that compare test statistics to fixed critical values will not have correct asymptotic size. Similarly, the existence of locally asymptotically \(\alpha\)-quantile unbiased estimators is closely tied to the existence of asymptotically similar confidence intervals for \(g(\theta)\) in local regions of \(\theta_\star\). Since any asymptotically similar confidence interval of the form \((-\infty, \hat c]\) can be converted into a locally asymptotically \(\alpha\)-quantile unbiased estimator by taking \(\Psi_n = \hat c\), by ruling out such estimators we also rule out the possibility of similar one-sided confidence intervals.\footnote{The focus on one-sided confidence intervals is largely for simplicity of exposition. If \((-\infty, \hat c_1]\) and \([\hat c_2, \infty)\) are two asymptotically similar confidence intervals for \(g(\theta)\) each with coverage rate \(1 - \alpha/2\) and \(\hat c_1 \leq \hat c_2\) with probability approaching one, then \([\hat c_2, \hat c_1]\) is asymptotically similar with coverage rate \(1 - \alpha\).}
To examine the possible behavior of such estimators, we make use of a representation result, given below in (ref), which is a slight adaptation of Theorem 8.3 in vanDerVaart1998. This earlier result is, in turn, a version of Le Cam's limit of experiments analysis for locally asymptotically normal models lecam1970assumptions,le1972limits.
(ref) establishes an equivalence between estimating \(g(\theta)\) in local regions of first-order degeneracy and estimating of a quadratic form of the mean parameter in a Gaussian shift model in which one observes a single draw \(Z \sim N(h, \Gamma_{\theta_\star}^{-1})\), where \(\Gamma_{\theta_\star}^{-1}\) is known but \(h\) is not. In particular, in a spirit similar to the approach in hirano2012impossibility, we can rule out sufficiently regular behavior of estimators of \(g(\theta)\) in local regions of \(\theta_\star\) if the corresponding behavior is not permissible in the Gaussian shift model. Intuitively, sufficiently regular estimation of quadratic forms in the Gaussian shift experiment is not possible since the parameter of interest changes non-linearly as the mean parameter \(h\) varies over \(\SR^d\) while the distribution of \(Z\) changes in a linear fashion.
To illustrate, suppose that there was an estimator, \(\Psi(Z,U)\), and law, \(\calL\) with \(\int x^2\,d\calL(x) < \infty\), such that \(\Psi(Z,U) - \frac{1}{2}h'\nabla^2 g(\theta_\star)h\) is distributed according to \(\calL\) for all \(h \in \SR^d\). Since any such estimator can be turned into an unbiased estimator by subtracting off the mean of \(\calL\), it is without loss of generality to assume that \(\calL\) is mean zero and thus that \(\Psi(Z,U)\) is unbiased. On the other hand, the Cramér-Rao lower bound for the variance of any unbiased estimator of \(\frac{1}{2}h'\nabla^2g(\theta_\star)h\) in the Gaussian shift model yields
By letting \(h\) vary over \(\SR^d\), the right hand side of (ref) can be made arbitrarily large while the left hand side is bounded by the second moment of \(\calL\). Thus, no such estimator can exist. Our full argument relies on analyzing characteristic functions, but the intuition is similar.
Together, (ref) can be combined for the main result of this section, which rules out sufficiently regular estimation in local areas of first-order degeneracy.\footnote{In the application of (ref) to our setting, take \(J = \frac{1}{2}\nabla^2 g(\theta_\star)\). The result in (ref) rules out well-behaved estimation of any quadratic form of the shift parameter, not just those associated with the Hessian of \(g(\cdot)\).}
(ref) rules out sufficiently well-behaved estimation of \(g(\theta)\) when the true parameter is “close” to \(\theta_\star\). In particular, (ref)(a) rules out the possibility of regular estimation --- the properly scaled and centered behavior of any estimator \(\Psi_n\) of \(g(\theta)\) must depend, in local regions of \(\theta_\star\), on the local parameter \(h\), which cannot be consistently estimated. Similarly, (ref)(b) rules out the possibility of \(\alpha\)-quantile unbiased estimation. As mentioned below (ref), this result has profound implications for inference on \(g(\theta)\) in local regions of \(\theta_\star\). In particular, both asymptotically exact Wald-type inference procedures and asymptotically similar confidence intervals for \(g(\theta)\) are unavailable in local regions of degeneracy \(\theta_\star\).
(ref)(a) also has implications for the construction of efficient estimators of \(g(\theta)\) in local regions of \(\theta_\star\). Standard notions of efficiency are tied to comparing the asymptotic risk of regular estimators. (ref)(a) shows that, after properly scaling, such regular estimators are not available in local regions of degeneracy. Consequently, alternative notions of efficiency must be considered in these settings and standard estimators may not be optimal in these regions. As an example, one can show that estimators of \(g(\theta)\) that are efficient under a standard asymptotic regime can be dominated by alternative estimators in local asymptotic mean squared error around points of degeneracy.
The above analysis rules out standard approaches to inference on \(g(\theta)\) in local regions of \(\theta_\star\). These results are informative when one is interested in constructing confidence intervals for \(g(\theta)\) around points of first-order degeneracy when the data is drawn from a parametric model satisfying (ref). In this subsection, we consider two extensions of our results. In the first, we consider the somewhat simpler problem of testing the null hypothesis \(H_0: g(\theta) = g(\theta_\star)\). We show that if the null hypothesis contains a sufficiently rich set of values and a similar test exists, this similar test must have low power in local regions of \(\theta_\star\). In the second extension we generalize (ref) to infinite dimensional, i.e, semiparametric or nonparametric, models.
In this subsection we consider the problem of testing the null hypothesis \(H_0: g(\theta) = g(\theta_\star)\) where the alternative can be one sided, i.e, \(H_1: g(\theta) > g(\theta_\star)\) or two-sided, \(H_1: g(\theta) \neq g(\theta_\star)\). To setup, define \(\calH_\star\) to be the set of local parameters, \(h \in \SR^d\), such that \(g(\theta_{n,h})\) is asymptotically indistinguishable from \(g(\theta_\star)\), i.e, \(r_n^2(g(\theta_{n,h}) - g(\theta_\star)) \to 0\); \[ \calH_\star = \{h \in \SR^d: h'\nabla^2g(\theta_\star)h = 0\}. \] We study the behavior of asymptotically similar tests in local regions of degeneracy. In this setup, an asymptotically similar test is a statistic \(\Xi_n: \calX^n \to \{0,1\}\) such that \[ \limsup_{n\to\infty} P_{\theta_{n,h}}(\Xi_n = 1) = \alpha \quad \text{for all } h \in \calH_\star. \] Equivalently, the test is similar if \(\limsup_{n\to\infty }P_{\theta_{n,h}}(1 - \Xi_n \leq 0) = \alpha\). Letting \(\Psi_n = 1 - \Xi_n\) it is apparent that this is a nearly identical requirement to that of local \(\alpha\)-quantile unbiasedness in (ref), with the key difference being that the requirement only needs to hold for local parameters \(h \in \calH_\star\) rather than for all \(h \in \SR^d\).
However, unlike quantile unbiased estimation, which is ruled out in (ref), asymptotically similar tests can exist --- one can imagine constructing a similar test by flipping a weighted coin. Such a test, though, may not be powerful against local alternatives close to \(\theta_\star\). Our main result in this subsection establishes this formally: if such test exists then its local asymptotic power curve must be flat at \(\theta_\star\) in the sense that the derivative of the local asymptotic power curve with respect to the local parameter \(h\) exists and is equal to zero.
Define the local asymptotic power curve as
The first statement in (ref) follows immediately from the definition of similarity along with the fact that \(\calH_\star\) is a cone: because \(\mathcal{P}(h)\) is constant on \(\calH_\star\), its directional derivatives in directions \(h \in \calH_\star\) must be zero. The force of the result, however, lies in the structure of \(\mathcal{H}^\star\) near points of degeneracy. In standard inference problems, i.e, when the true parameter is well separated from points of degeneracy, the parameter of interest in the limit experiment is a linear function of the shift parameter \(h\), so \(\mathcal{H}^\star = \{0\}\) and the zero-derivative condition carries no information about the shape of the power curve. By contrast, when \(g\) exhibits first-order degeneracy \(\mathcal{H}_\star\), can be a non-trivial cone — that is, it may contain directions other than zero — and the constraint \(D_h \mathcal{P}(0) = 0\) for \(h \in \calH_\star\) becomes substantive. When \(\nabla^2 g(\theta_\star)\) is indefinite, \(\calH_\star\) spans \(\SR^d\) and the entire gradient of the local asymptotic power curve vanishes at the origin, implying that power cannot increase at a linear rate in any direction away from \(\theta_\star\). The following examples illustrate the strength of this restriction.
Example (ref), cont. Consider again the mediation model, where the original model is given \((P_\theta: \theta \in \Theta \subseteq \SR^2)\). Suppose the researcher is interested in testing the null hypothesis \(H_0:\theta_1\theta_2 = 0\), that is \(g(\theta) = \theta_1\theta_2\) and \(\theta_\star = 0\). We can show that \(\calH_\star = \{h \in \SR^2: h_1h_2 = 0\}\). Since \(\calH_\star\) is the union of the two coordinate axes, \(\mathrm{span}(\calH_\star) = \SR^2\). Thus, we have that \(\nabla\calP(0) = 0\) for any asymptotically similar test.
We can equivalently show that \(\calH_\star\) must span \(\SR^2\) by noting that the second derivative matrix of \(g(\cdot)\) at \(\theta_\star\) is given by \[
, \] which is indefinite. \qed
In many settings the researcher may not be willing to assume that the data comes from a finite dimensional parametric model as described in the previous section. Following a tradition on studying semiparametric efficiency bickel1993efficient, we show that this does not affect our impossibility results so long as the larger model contains a parametric submodel satisfying (ref).
Formally, let the model \(\calP\) be a collection of sequences of probability measures on the sample space \(\calX^n\) from the previous section. That is, each element of \(\calP\) is a sequence of probability measures \(\{P_n\}\), where each probability measure \(P_n\) is defined on the sample space \(\calX^n\). A finite dimensional submodel, \(\calP_f\), is some smaller collection of sequences of probability measures that can be parameterized as \(\calP_{f} = (\{P_{n,\theta}\}_{n\in\SN}: \theta \in \Theta)\) for an open set \(\Theta \in \SR^{d_f}\). Fix a “centering” sequence of probability measures \(\{P_{0,n}\}_{n \in \SN} \in \calP\). We say that the submodel passes through \(\{P_{0,n}\}\) if \(\{P_{0,n}\} \in \calP_f\), that is \(\{P_{0,n}\} = \{P_{n,\theta}\}\) for some \(\theta \in \Theta\). We will call such a parametric model “regular” if (ref) holds and the model passes through \(\{P_{0,n}\}\).
We suppose that the object of interest is a quantity that depends on the sequence of underlying probability measures, that is we can think of the estimand \(g[\{P_n\}]\) as a functional defined on \(\calP\). For any regular parametric model, \(\calP_f\), this implicitly defines a function on \(\theta\) via the relation \(g_f(\theta) = g[\{P_{n,\theta}\}]\). With this notation defined, we show that the results of (ref) can be extended in a straightforward fashion to infinite dimensional models.
In this section, we construct a uniformly valid confidence interval for $g(\theta):\Theta\rightarrow\mathbb{R}$ using estimator $\hat{\theta}$. The confidence interval is obtained by inverting the hypothesis $H_{0}:g(\theta)=\tau$, and we use a minimum distance (MD) test statistic \[ \hat{T}_{n}(\tau)=\inf_{\theta\in\Theta:g(\theta)=\tau}r_{n}^{2}(\hat{\theta}-\theta)^{\prime}\Sigma^{-1}(\hat{\theta}-\theta). \] We focus on the settings where the standard first order approximation of $g(\theta)$ fails at $\theta_{\star}$, but the second order derivative is nondegenerate; that is, $\frac{\partial^{2}g}{\partial\theta\partial\theta^{\prime}}(\theta_{\star})=H$ with $H$ full rank.
In Section (ref), we discuss a simple case where $\Theta \subseteq \mathbb{R}^{2}$ and $H$ is indefinite, and we provide sufficient conditions under which the standard critical value $Q(\chi_{1}^{2},1-\alpha)$ is uniformly valid. In Section (ref), we propose a computationally simple method for a general $g$, which can be generalized to cases with higher order singularity.
To simplify notation, consider the null hypothesis
with $|\rho|<1$ and $\tau\geq0$. The restriction $|\rho|<1$ guarantees that $H$ is indefinite, while $\tau\geq0$ is a normalization. For simplicity, let $n=r_{n}=1$, and assume $\hat{\theta}-\theta\sim N(0,I_{2}).$ The quadratic form $g$ and the normality of $\hat{\theta}$ can be viewed as second order approximations, with general asymptotic results provided in Theorem (ref).
Let $X_{2}(\theta_{1})\in \mathbb{R}_+$ be the positive solution for $\theta_{2}$ such that ((ref)) holds. Let $\mathcal{S}_{0}(\tau)$ be the null parameter space, which contains two separate curves, $\mathcal{S}_{0}^{+}(\tau)$ and $\mathcal{S}_{0}^{-}(\tau)$,
where \[ \mathcal{S}_{0}^{+}(\tau)=\left\{ \left(x_{1},X_{2}(x_{1})\right):x_{1}\in\mathbb{R}\right\} ,\quad\mathcal{S}_{0}^{-}(\tau)=\left\{ \left(x_{1},-X_{2}(x_{1})\right):x_{1}\in\mathbb{R}\right\} . \] Let $\mathcal{S}(\tau,c)$ be the acceptance region with critical value $c^{2}$, i.e., the $c$-enlargement of $\mathcal{S}_{0}(\tau)$, \[ \mathcal{S}(\tau,c)=\left\{ (x_{1},x_{2}):(x_{1}-\theta_{1})^{2}+(x_{2}-\theta_{2})^{2}\leq c^{2},(\theta_{1},\theta_{2})\in\mathcal{S}_{0}(\tau)\right\} . \]
Proposition (ref) shows that the standard MD test remains valid under a curved null hypothesis when the maximum curvature of $\mathcal{S}_{0}^+(\tau)$, given by $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}$, is sufficiently small, or when the two branches $\mathcal{S}_{0}^{+}(\tau)$ and $\mathcal{S}_{0}^{-}(\tau)$ are sufficiently close. The argument proceeds by comparing the coverage of $\mathcal{S}(\tau,c)$ with that of the auxiliary acceptance set, \[\mathcal{S}_{\text{aux}}=\left\{ (x_{1},x_{2}):(x_{2}-\theta_{2})^{2}\leq c^{2}\right\} .\] whose coverage is exactly $1-\alpha$. Expressed in polar coordinates, the coverage depends on the fraction of each circle of radius $r$ centered at $\theta\in \mathcal{S}^{+}_0(\tau)$, denoted $\partial B(\theta,r)$, that is contained in the acceptance region. Consequently, it suffices to show that, for each $r$, the arc length of $\partial B(\theta,r)$ contained in $\mathcal{S}(\tau,c)$ is no smaller than that of $\mathcal{S}_{aux}$.
When $r\leq c$, the entire circle is covered by both $\mathcal{S}_{aux}$ and $\mathcal{S}(\tau,c)$ by construction. For $r>c$, let $\mathcal{C}_{u}(\tau)$ and $\mathcal{C}_{\ell}(\tau)$ denote the upper and lower boundaries of $\mathcal{S}^{+}(\tau,c)$; see Figure (ref) left panel. If $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}\leq\frac{1}{c}$, the circle $\partial B\left(\theta,r\right)$ intersects $\mathcal{C}_{u}(\tau)$ at points $A$ and $B$, and $\mathcal{C}_{\ell}(\tau)$ at points $C$ and $D$. We can show that the lengths of chords $\overline{AC}$ and $\overline{BD}$ are no smaller than $2c$. Otherwise, $B(G,c)\not\subseteq\mathcal{S}^{+}(\tau,c)$, where $G$ denotes the intersection of $AC$ with $\mathcal{S}_{0}^{+}(\tau)$, contradicting the definition of $\mathcal{S}^{+}(\tau,c)$. Note that the arcs of $\partial B(\theta,c)$ covered by $\mathcal{S}_{aux}$ correspond to chords of length $2c$. It therefore follows that the portion of $\partial B(\theta,c)$ covered by $\mathcal{S}(\tau,c)$ is larger than that covered by $\mathcal{S}_{aux}$.
If $\frac{1-\rho}{\sqrt{\tau(1+\rho)}}>\frac{1}{c}$, $\mathcal{C}_{u}(\tau)$ has a kink due to the high curvature. Thus, there exist $\theta\in\mathcal{S}_{0}^{+}(\tau)$ and $r>c$ such that $\partial B\left(\theta,r\right)$ does not intersect $\mathcal{C}_{u}(\tau)$; see Figure (ref) right panel. For such $(\theta,r)$, the argument based on chord lengths no longer applies. However, when $\rho\geq0$, $\partial B(\theta,r)$ is sufficiently close to $\mathcal{S}_{0}^{-}(\tau)$, and we can show that $\partial B\left((\theta_{1},\theta_{2}),r\right)\subseteq\mathcal{S}(\tau,c)$.
Next, we present the asymptotic results for general data generating processes.
(ref).(ref) assumes that the second order derivative of \(g\) at \(\theta_\star\) is of full rank, so that there is no higher order degeneracy. (ref).(ref) requires that the researcher has access to an estimator \(\hat\theta\) that is uniformly asymptotically normal over the class of DGPs considered while (ref).(ref) and (ref).(ref) require that the asymptotic variance of this estimator is well-behaved and consistently estimable. Since degeneracy is a property of the transformation of interest, $g$, rather than of the primitive parameter $\theta$, these are mild conditions that can be verified for most common estimators $\hat\theta$.
Theorem (ref) follows from Proposition (ref). To see this, let $\tilde{g}(h) = g\!\left(\theta_{\star} + r_{n}^{-1}\Sigma^{1/2} h\right),$ where \( r_{n} \) governs the rate at which \( \theta \) is estimated and \( \Sigma \) adjusts for the covariance. For \( h = O(1) \), \( \tilde{g}(h) \) can be approximated by a hyperbola, as in (ref). Condition (ref) ensures that the curvature of \( \tilde{g}(h) \) is not too large, while (ref) implies that the two branches of the hyperbola are sufficiently close. The result remains uniformly valid even as \( \|h\| \to \infty \).
In this section, we present the inference procedure for a general function $g$. The procedure is based on a local approximation of the test statistic. First, consider the case where the true parameter value $\theta_{n}$ satisfies $\theta_{n}=\theta_{\star}+h_{n}/r_{n}$. Under $H_0$, the test statistic is given by\\ \[ \hat{T}_{n}(g(\theta_n))=\inf_{\vartheta:g(\vartheta)=g(\theta_{n})}r_{n}^{2}\left(\hat{\theta}-\vartheta\right)^{\prime}\hat{\Sigma}^{-1}\left(\hat{\theta}-\vartheta\right)=r_{n}^{2}\left\Vert \hat{\Sigma}^{-1/2}(\hat{\theta}-\tilde{\theta}_{n})\right\Vert ^{2} \] where $\tilde{\theta}_{n}$ denotes the minimizer. Since $\hat{T}_n(g(\theta_n))\leq r_{n}^{2}\left\Vert \hat{\Sigma}^{-1/2}(\hat{\theta}-\theta_{n})\right\Vert ^{2}=O_p (1)$, we have $\tilde{\theta}_{n}=\theta_{n}+O_{p}(\frac{1}{r_{n}}).$ If $h_n = O(1)$, a second order Taylor expansion of $g(\tilde{\theta}_{n})=g(\theta_{n})$ gives
In addition, let $\mathbb{Z}_{n}=r_{n}\hat{\Sigma}^{-1/2}(\hat{\theta}-\theta_{n})$, we can write
In sum, let $t=r_n(\tilde{\theta}_n-\theta_\star)$, given $h_{n}$, we can approximate $\hat{T}_{n}$ by
where $\mathbb{Z}|\hat{T}_{n}(g(\theta_n))\sim N(0,I_{d})$. We can show that $\hat{T}_{n}(g(\theta_n))$ and $\hat{T}_{n}^{*}(h_{n})$ have the same asymptotic distribution, regardless of whether $h_{n}$ converges to $h\in\mathbb{R}$ or diverges to infinity (Lemmas (ref) and (ref)). Intuitively, if $h_{n}\rightarrow\infty$, the restriction for the optimizer $\tilde{t}$ in ((ref)) is approximately linear. In this case, both $\hat{T}_{n}(g(\theta_n))$ and $\hat{T}_{n}^{*}(h_{n})$ are approximated $\chi_{1}^{2}$.
Given $h_{n}$, we can easily get the quantile of $\hat{T}_{n}^{*}(h_n)$ by simulation. However, $h_{n}$ is a nuisance parameter that cannot be consistently estimated. Next, we propose a two step feasible critical value. Suppose set $\mathcal{H}_{z}$ satisfies $P\left(N(0,I_{d})\in\mathcal{H}_{z}\right)=1-\eta$.\footnote{For instance, $\mathcal{H}_z=\{z\in \mathbb{R}^d:z^\prime z\leq Q(\chi^2_d,1-\eta)\}.$} In the first step, we construct a $(1-\eta)$ confidence set for $h_{n}$,
In the second step, we construct the critical value based on the $\frac{1-\alpha}{1-\eta}$ quantile of $\hat{T}_{n}^{*}(h)$ conditional on the first step. That is, let
and reject $H_{0}:g(\theta)=\tau$ if $\hat{T}_{n}(\tau)>\hat{c}$. In ((ref)), the construction of $\hat{c}$ takes into account the first step selection, thus it is less conservative than simple Bonferroni correction.
The slight conservativeness arises from the two-step procedure. Alternatively, we can introduce a pretest to check whether $h_{n}$ is far away from zero, e.g. $||h_{n}||>\ln r_{n}$. If so, we can use the standard critical value $Q(\chi_{1}^{2},1-\alpha)$. The cost is that we need to introduce an extra tuning parameter.
In this section, we examine the size and power properties of the proposed procedures and compare them with several alternatives. We focus on the context of Example (ref), namely the construction of confidence intervals for the mediation effect. In addition to the two MD-based methods proposed in Section (ref), one using the $Q(\chi^2_1,1-\alpha)$ critical value (BN1; Section (ref)) and one using a bootstrapped critical value (BN2; Section (ref)), we consider two uniformly valid MD-based alternatives: (i) the procedure of AM-2016-Geometric (AM),\footnote{We report results using their Section 4.1 implementation, which computes curvature over a restricted set with tuning parameter $\eta=\alpha/10$. We also implemented their worst-case curvature procedure from Section 2. The two procedures have nearly identical power, with the latter performing slightly worse.} and (ii) the MD-based method with projection critical value $Q(\chi_{2}^{2},1-\alpha)$. For comparison, we also include the naive delta method, i.e. a Wald-type test with critical value $Q(\chi_{1}^{2},1-\alpha)$, and a naive bootstrap method. The nominal rejection rate is $\alpha=0.05$, and the tuning parameter for BN2 is $\eta=\alpha/10$.
We study confidence intervals for $g(\theta)=\theta_{1}\theta_{2}$, where the estimators are simulated from \[ \hat{\theta}-\theta\sim\mathcal{N}\left(0,\left[
\right]\right). \] Without loss of generality, we normalize the variance of $\hat{\theta}$ to one. We consider $r=0$ and $r=0.5$, with $\theta_{2}\in\{2,6\}$ and $\theta_{1}=[-1:0.2:1]\times\theta_{2}$.
In Figure (ref), we plot the probability that the confidence intervals exclude the true value $g(\theta)$. The naive method does not overreject when $\theta_{1}\theta_{2}=0$, but its rejection probability is very low near the origin, consistent with earlier findings (e.g., Dufour2025WaldTW). Away from the origin (see, e.g., $\theta_1=\theta_{2}=2$), the naive Wald test overrejects. According to Remark (ref), BN1 is valid for $r=0$. For $r=0.5$, Theorem (ref) does not guarantee validity when $\theta_{1}<0$, $\theta_{2}=2$, or when $\theta_1\in[-1.4,0)$, $\theta_2 = 6$. Nevertheless, BN1 maintains correct size across all designs, even when these conditions fail, suggesting that the condition is sufficient but not necessary. As expected, all other MD-based methods control size.
Figure (ref) shows the probability that the confidence intervals exclude zero, i.e., the probability of obtaining a significant result. When $\theta$ is close to the origin ($\theta_2=2$), our methods have substantially higher power than AM, whose performance is close to that of the simple projection method. When $\theta$ is further from the origin ($\theta_{2}=6$), power curves across methods are nearly identical.
Finally, Figure (ref) reports the median length of the confidence intervals, computed across $S$ replications. BN1 consistently yields the shortest intervals, with BN2 close behind. The projection method is the most conservative, producing intervals 19--30% longer than BN1. AM lies between BN1 and the projection method, with median lengths 5--18% longer than those of BN1. The differences are most pronounced when $\theta$ is near the origin.
We illustrate the empirical relevance of our results using the setting analyzed by AEM-2018. Their study takes advantage of a distinctive feature of the Turkish education system, in which elementary school teachers are randomly allocated across schools. This institutional detail generates plausibly exogenous variation in teacher characteristics that can be used to study how teachers' gender role attitudes influence student outcomes. The data include roughly 4,000 third- and fourth-grade students taught by 145 teachers, and students can be grouped according to the length of their exposure to a given teacher --- at most one year, two to three years, or up to four years. The treatment variable is whether a teacher is identified as holding traditional rather than progressive gender beliefs, while the mediator of interest is the student's own gender role beliefs. Following a similar analysis of this data in van_garderen_nearly_2024, we focus on verbal test scores as the outcome. AEM-2018 argue that, after controlling for an extensive set of student, family, teacher, and school characteristics, the identifying assumptions for causal mediation analysis are satisfied in this context.
(ref) reports estimates from the AEM-2018 analysis linking teachers’ gender role attitudes to students’ verbal test performance. The first coefficient, $\hat{\theta}_1$, comes from a regression of students’ gender role beliefs on the gender role attitudes of their teachers, with the standard set of student, family, teacher, and school controls included. This coefficient summarizes the extent to which traditional teachers transmit their views to students. The second coefficient, $\hat{\theta}_2$, is estimated from a regression of test scores on both student gender beliefs and teacher attitudes, again with the full set of controls. It reflects how student beliefs are associated with verbal performance once teacher attitudes are held constant. Multiplying these two coefficients gives the mediated, or indirect, effect: the part of the teacher’s influence on scores that operates through the channel of student beliefs. The estimates show that this indirect pathway is negative and relatively small, although it varies across exposure groups, being largest in the one-year sample and smallest for students exposed for four years.
Because the true mediation, or indirect, effect appears to be close to zero, the results of (ref) suggest that standard approaches to inference will fail. In particular, we cannot construct valid confidence intervals via the typical approach of inverting a \(t\)-test. Instead, we construct confidence intervals using the newly proposed methods of (ref). van_garderen_nearly_2024 show that the correlation coefficient between $\hat{\theta}_1$ and $\hat{\theta}_2$ is zero; consequently, both of our methods are uniformly valid.\footnote{The validity of BN1 follows from Remark (ref).}
We compare our confidence intervals to two other inference procedures that might be applied in this setting, both of which are based on the minimum distance statistic. The first alternate procedure is that of AM-2016-Geometric.\footnote{In implementing the test, we follow the empirical application in the working paper version of AM-2016-Geometric and only calculate the maximum curvature over a set “close” to the point estimate, adjusting the critical value accordingly.} This testing procedure technically does not cover the case where we are testing the null that the mediation effect is equal to zero since the null manifold is not smooth in this case. However, the AM-2016-Geometric critical value approaches \(Q(\chi^2_2, 1-\alpha)\) from below as the null hypothesis value approaches zero and \(Q(\chi^2_2, 1 - \alpha)\) is a valid critical value for testing the null that the mediation effect is equal to zero so we simply modify the procedure slightly to directly use a \(Q(\chi^2_2, 1 - \alpha)\) when the null value is equal to zero. The second method, “Projection”, simply uses the \(Q(\chi^2_2, 1 -\alpha)\) at all points, which is justified since, under the null hypothesis, the distance to the null manifold is always less than the distance to the point \((\theta_1,\theta_2)'\).
Consistent with the discussion in (ref), the confidence intervals based on either the $\chi^2_{1}$ critical value (BN1) or the two-step procedure (BN2) are uniformly tighter than those obtained from the AM-2016-Geometric simulated critical value (AM); in all specifications, our intervals are strict subsets of theirs. The difference is not only theoretical but also empirically relevant. Using the $\chi^2_{1}$ critical value, for instance, the BN1 confidence interval supports the conclusion of van_garderen_nearly_2024 that the mediation effect of a one-year exposure to a teacher with traditional views is negative, whereas the alternative methods cannot reject a null of zero at the five-percent level. As expected, the AM intervals lie strictly inside those generated by the projection method, which uses a \(\chi^2_2\) critical value at all points. However, because the true mediation effect appears small in this setting, their simulated critical value converges toward $Q(\chi^2_2, 1 -\alpha)$, which accounts for the close similarity between the two sets of intervals.
We examine inference in local regions of first-order degeneracy, meaning that the gradient of the transformation is zero or nearly zero so that first-order approximations alone do not provide reliable information and second-order terms must also be considered. In such regions of local degeneracy, we show that neither regular estimation nor quantile-unbiased procedures are feasible, paralleling impossibility results for nondifferentiable functionals and ruling out standard approaches to inference. We then develop alternate inference procedures based on minimum-distance statistics that deliver uniformly valid confidence intervals. Simulation studies indicate that these procedures control size while maintaining favorable power, and the empirical application to teacher gender attitudes shows that they yield tighter confidence intervals than existing approaches.