EconBase
← Back to paper

Inference for Linear Systems with Unknown Coefficients

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

76,569 characters · 11 sections · 64 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for Linear Systems with Unknown Coefficients

spacing{1.1} \begin{abstract} This paper considers the problem of testing whether there exists a solution satisfying certain non-negativity constraints to a linear system of equations. Importantly and in contrast to some prior work, we allow all parameters in the system of equations, including the slope coefficients, to be unknown. For this reason, we describe the linear system as having unknown (as opposed to known) coefficients. This hypothesis testing problem arises naturally when constructing confidence sets for possibly partially identified parameters in the analysis of nonparametric instrumental variables models, treatment effect models, and random coefficient models, among other settings. To rule out certain instances in which the testing problem is impossible, in the sense that the power of any test will be bounded by its size, we begin our analysis by characterizing the closure of the null hypothesis with respect to the total variation distance. We then use this characterization to develop novel testing procedures based on sample-splitting. We establish the validity of our testing procedures under weak and interpretable conditions on the linear system. An important feature of these conditions is that they permit the dimensionality of the problem to grow rapidly with the sample size. A further attractive property of our tests is that they do not require simulation to compute suitable critical values. We illustrate the practical relevance of our theoretical results in a simulation study. \end{abstract}

KEYWORDS: Linear programming, linear (in)equalities, partial identification, uniform inference, treatment effects, nonparametric instrumental variables

JEL classification codes: C31, C35, C36

\thispagestyle{empty} \setcounter{page}{1}

Introduction

Given an independent and identically distributed (i.i.d.) sample $\{Z_i\}_{i=1}^n$ with $Z_i$ distributed according to $P \in \mathbf P$, this paper studies the hypothesis testing problem

equation[equation omitted — 126 chars of source]

where $\mathbf P$ is a “large” set of distributions satisfying conditions described below and

align[align omitted — 187 chars of source]

Here, “$x_1 \geq 0$” signifies that all coordinates of $x_1 \in \mathbb R^{d_1}$ are non-negative, $A_0(P)$ is a $p \times d_0$ matrix with $d_0 \geq 0$, $A_1(P)$ is a $p \times d_1$ matrix, and $\beta(P)$ is a $p \times 1$ vector.

As discussed further in Section (ref), the testing problem described above arises naturally in many settings of empirical interest, including (i) inference for linear functionals of structural functions in nonparametric instrumental variables (NPIV) models with shape restrictions, as in freyberger2015identification; (ii) inference for average marginal effects in nonlinear models with random coefficients, as in fox2011simple; (iii) inference for treatment effect parameters that are partially identified through the marginal treatment response (MTR) framework of mogstad2018using; (iv) inference under “synthetic parallel trends” with convex weights as considerd in liu2025synthetic; and (v) inference for counterfactual choice probabilities in the distribution-free binary choice model in gurussell2023joe. We also note that the null hypothesis (ref) subsumes as a special case the setting where $A_0(P)$ and $A_1(P)$ are known, for which there exist many other examples fang2023inference.

We demonstrate in Section (ref) that testing (ref) in certain special cases is impossible (in the sense that the power of any test is bounded by its size), by observing that the null set $\mathbf{P}_0$ is dense in $\mathbf{P}$ with respect to the total variation metric. Accordingly, the first contribution of the paper is to obtain a characterization of the closure of $\mathbf P_0$ which guides the construction of our test. Exploiting Farkas' lemma, we show that (a subset of) the closure of $\mathbf P_0$ with respect to the total variation metric can be described in terms of a set of linear inequality restrictions involving the projection of $(A_1(P), \beta(P))$ onto the orthogonal complement of the column span of $A_0(P)$. Specifically, the characterization amounts to verifying if, for every unit vector, the minimum of a collection of linear inequalities is non-positive.

Building on this observation, in Section (ref) we develop an inference procedure based on sample splitting. Using the first subsample, we construct a unit vector for which we expect that our characterization of the null is most strongly violated. We then use the second subsample to formally test whether or not our inequalities are violated. By virtue of this sample-splitting construction, the resulting test is straightforward to implement and involves no simulation: in particular, the construction of the “violating” unit vector requires solving two linear programs, and the test statistic is provided in closed form with rejection based on a comparison with the appropriate quantile of a standard normal distribution. Moreover, the uniform asymptotic validity of our test is established under weak and interpretable regularity conditions while allowing the dimensions of $p$, $d_0$, $d_1$ to grow with the sample size $n$.

We propose two variants of our test. The first, which we term the “direct” method, tests if the union of all the inequalities in our characterization hold. The second, which we term the “screening” method, tests whether or not a single inequality in our characterization is non-positive, under the restriction that the other inequalities are positive with high probability. In our simulation study, the screening method typically yields shorter confidence intervals through test inversion, at the cost of introducing one tuning parameter when selecting the unit vector in the first subsample.

Inference procedures for testing (ref) and closely related null hypotheses have gained increasing attention in the literature. bai2022testing, andrews2023inference, cox2023simple, and fang2023inference all propose inference procedures which could be applied to (ref) whenever $A_0(P)$ and $A_1(P)$ do not depend on $P$ (i.e. they are known quantities). The method proposed in our paper immediately applies to this setting as a special case. There is also a literature for the closely related problem of inference for the value of a linear program: see in particular freyberger2015identification, cho2024simple, gafarov2025simple, voronin2025linear, goff2025inference. As explained in cox2025testing, these methods can be used to address the same empirical problems as those that we consider in Section (ref), but there is a subtle technical difference between the two settings, and in general these methods are neither tuning parameter nor simulation-free.

Two recent papers that propose inference procedures which could be used to test (ref) are cox2025testing and goff2025inference. The theoretical results in cox2025testing and goff2025inference require rank conditions which, as argued in liu2025synthetic, could be difficult to verify or may even be violated in certain settings of empirical interest. In contrast, we are able to establish the asymptotic validity of our procedures under weak and interpretable regularity conditions on the ranks of $A_0(P)$ and $A_1(P)$. Moreover, neither paper establishes the validity of their tests in a high-dimensional regime where $p$, $d_1$ and $d_0$ are allowed to grow with sample size, as we do in this paper.\footnote{We note that goff2025inference explain that they derive their results in a semi high-dimensional regime, where the number of unknown coefficients in the linear system must remain bounded.} We illustrate the finite-sample performance of our procedures using simulation designs based on the ones considered previously in these papers, as well as some high-dimensional counterparts that these papers do not consider.

The remainder of the paper is organized as follows. In Section (ref), we describe several examples of empirical problems of interest that can be accommodated in our framework. Our main results are contained in Sections (ref) and (ref): Section (ref) presents our characterization of the closure of the null hypothesis, whereas Section (ref) describes the test and establishes its uniform asymptotic validity. In Section (ref), we study the finite-sample behavior of our proposed tests in a simulation study. Proofs of all results are collected in the Appendix.

Examples

In this section, we present a collection of motivating examples, two of which we revisit in the simulation study in Section (ref). goff2025inference and cox2025testing discuss several additional examples that provide further motivation.

example[Nonparametric instrumental variables with shape restrictions] Consider the NPIV model studied in freyberger2015identification, in which \begin{gather} Y = g(X) + U ,\notag \\ E_P[U|W] = 0 , \end{gather} with $X$ a possibly endogenous explanatory variable supported on $\mathcal{X} = \{x_1, \dots, x_H\}$, and $W$ an instrumental variable supported on $\mathcal{W} = \{w_1, \dots, w_K\}$ with $K < H$. Define \begin{align*} \pi_{hk}(P) & = P \{X = x_h, W = w_k\} \\ m_k(P) & = E_P[Y 1\{W = w_k\}] \\ \Pi(P) & = (\pi_{hk}(P))_{1 \leq h \leq H, 1 \leq k \leq K} \\ m(P) & = (m_1(P), \dots, m_K(P))' . \end{align*} Then, $g = (g(x_1), \dots, g(x_H)) \in \mathbb R^H$ satisfies $ \Pi(P)' g = m(P)$. Assume additionally that the vector $g$ satisfies shape constraints encoded as $ Sg \leq 0$ for some matrix $S \in \mathbb R^{M \times H}$. Suppose we wish to test the null hypothesis $H_0: L(g) = L_0 \in \mathbb R$, where $L(g)$ represents a generic linear functional $L(g) = c'g$. Putting everything together, we obtain the following system: \[ A_0(P) = \begin{pmatrix} \Pi(P)'\\ S\\ c' \end{pmatrix} \hspace{3em} A_1(P) = \begin{pmatrix} \mathbf{0}_{K \times M} \\ \mathbf{I}_{M} \\ \mathbf{0}_{1\times M} \end{pmatrix} \hspace{3em} \beta(P) = \begin{pmatrix} m(P) \\ \mathbf{0}_{M\times 1} \\ L_0 \end{pmatrix}~. \] In this example, $A_1(P)$ is known and $A_0(P)$ and $\beta(P)$ are unknown.
example[Average marginal effects in nonlinear models with random coefficients] fox2011simple consider a class of nonlinear mixture models with discrete unobserved heterogeneity. A simple example is a static, binary choice logit model with random coefficients: \begin{align*} Y = 1\{C'W - U \geq 0\} , \end{align*} where $Y \in \{0,1\}$ is an observed choice, $W$ is a vector of observed explanatory variables, $C$ is a vector of latent random coefficients, and $U$ is a latent random variable that follows a standard logistic distribution, independently of $(C,W)$. Under these assumptions, a consumer of type $c$ with observables $w$ chooses $Y = 1$ with probability \begin{align} P \{Y = 1 \vert W = w, C = c\} = \frac{1}{1 + \exp(-c'w)} := \ell(c'w) , \end{align} where $\ell(\cdot)$ is the standard logistic distribution function. bajarifoxryan2007aer and fox2011simple assume $C$ is independent of $W$ and approximate the distribution of $C$ using a discrete distribution with known support points $(c_{1},\ldots,c_{d})$ and unknown respective probabilities $\pi := (\pi_{1},\ldots,\pi_{d})$. Then (ref) implies observed probabilities \begin{align} P \{Y = 1 \vert W = w\} = \sum_{j=1}^{d} \pi_{j}\ell(c_{j}'w) . \end{align} A target parameter in this model is the average marginal effect (AME) of the $k$th explanatory variable wooldridge2010: \begin{align} \alpha(P) := E_{P}\left[ \frac{\partial}{\partial w_{k}} \ell(C'W) \right] = E_{P}[ C_{k}\ell(C'W)(1-\ell(C'W)) ] = \sum_{j=1}^{d} \pi_{j}\underbrace{ c_{jk}E_{P}[\ell(c_{j}'W)(1-\ell(c_{j}'W))] }_{\alpha_{j}(P)} , \end{align} where $c_{jk}$ is the $k$th component of the $j$th support point, $c_{j}$. Suppose we wish to test the null hypothesis $H_{0}: \alpha(P) = \alpha_{0}$ for a vector of probabilities $\pi$ that satisfies (ref) at $p-2$ support points $w_{1},\ldots,w_{p-2}$. This problem fits into form (ref) with $x_{0}$ null, $x_{1} = \pi$, \begin{align*} A_{1}(P) = \begin{pmatrix} \ell(c_{1}'w_{1}) & \cdots & \ell(c_{d}'w_{1}) \\ \vdots & \vdots & \vdots \\ \ell(c_{1}'w_{p-2}) & \cdots & \ell(c_{d}'w_{p-2}) \\ 1 & \cdots & 1 \\ \alpha_{1}(P) & \cdots & \alpha_{d}(P) \end{pmatrix} \beta(P) = \begin{pmatrix} P\{Y=1|W = w_1\} \\ \vdots \\ P\{Y=1|W=w_{p-2}\} \\ 1 \\ \alpha_{0} \end{pmatrix} . \end{align*} Notice that the dependence of $A_{1}(P)$ on $P$ comes through the row corresponding to the AME. The same structure will generally appear in a linear program with known coefficients fang2023inference when the target parameter is an object that averages over observed heterogeneity.
example[Instrumental variables with heterogeneous treatment effects] mogstad2018using develop an approach to marginal treatment effect analysis heckman2005structural that allows for shape constraints and partial identification. A binary treatment $D$ produces two real-valued potential outcomes $Y(0)$, $Y(1)$, and an observed outcome $Y = (1-D)Y(0) + DY(1)$. The researcher additionally observes an instrument $Z$. Treatment assignment satisfies the imbens1994identification monotonicity condition, which can be equivalently written with the threshold-crossing model $D = 1\{p(Z) \geq U\}$, where $U$ is a uniformly distributed unobservable that is independent of $Z$ and $p(z) := P\{D = 1 \vert Z = z\}$ is the propensity score vytlacil2002independence. The marginal treatment response functions are assumed to take a linear-in-parameters form \begin{align} E[Y(d) \vert U = u] = \theta(d)'b(d \vert u), \end{align} where $\theta(d)$ are unknown parameters and $b(d \vert u)$ are known basis functions. The linear parameterization (ref) implies that common target parameters can also be written as linear functions of the parameters, $\theta := (\theta(0), \theta(1))$, taking the general form \begin{align} \tau^{\star} = \theta(0)' E\left[ \int_0^1 b(0 \vert u) \omega^\star(0 \vert u,Z)du \right] + \theta(1)' E\left[ \int_0^1 b(1 \vert u)\omega^\star(1 \vert u,Z)du \right] = \sum_{d \in \{0,1\}} \theta(d)'E[t(d \vert Z)], \end{align} where $\omega^{\star}(d \vert u,z)$ are scalar weights that are known or identified and $t(d \vert z) = \int_0^1 b(d \vert u)\omega^\star(d \vert u,z)du$. mogstadtorgovitsky2024hole observe that if $Y(d)$ is mean independent of $Z$, conditional on $U$, then the linear-in-parameters form (ref) also has implications for the observed outcome: \begin{align} E_P[Y \vert D, Z] = \phi(D,Z)'\theta, \end{align} where $\theta := (\theta(0), \theta(1))$ and $\phi$ is a known function of $(D,Z)$ that depends on the basis functions, $b$, and the propensity score, $p(Z)$. Bounds on $\tau^{\star}$ can be found by considering all values of (ref) that can be produced by $\theta$ that satisfy (ref) either for all $(D,Z)$ or for some implied moments. mogstad2018using propose using moments of the form $E_P[Ys(D,Z)]$, which can be shown from (ref) to also be linear in $\theta$. sheatorgovitsky2023os alternatively propose using the normal equations implied by (ref): \begin{align} E_P[\phi(D,Z)\phi(D,Z)']\theta = E_P[\phi(D,Z)Y], \end{align} which has the advantage of always being the same dimension as $\theta$ and not requiring one to choose the $s$ functions. In either case, shape constraints can be imposed on the marginal treatment response functions by constraining $\theta$. Consider testing whether $\tau^* = \tau_0$ for $\tau_0 \in \mathbb R$. Without shape constraints, the normal equation version of the problem can be phrased as (ref) by taking $x_{0} = \theta$, then setting \begin{align} A_{0}(P) = \begin{pmatrix} E_{P}[\phi(D,Z)\phi(D,Z)'] \\ E_{P}[t(Z)]' \end{pmatrix} \quad and \quad \beta(P) = \begin{pmatrix} E_{P}[\phi(D,Z)Y] \\ \tau_0 \end{pmatrix} , \end{align} where $t(Z) = (t(0 \vert Z)', t(1 \vert Z)')'$. Shape constraints can be imposed by including appropriate slack variables. For example, if $Y \in \{0,1\}$ is binary and $b(d \vert u)$ are Bernstein polynomials, then the implied MTR function can be constrained to lie in $[0,1]$ by restricting all elements of $\theta$ to lie within $[0,1]$. To incorporate these shape constraints into (ref), we would now let $x_{0}$ be null and set $x_{1} = [\theta', s']'$, where $s$ are slack variables. Instead of (ref), we would take $A_{0}$ to be empty and set \begin{align*} A_{1}(P) = \begin{pmatrix} E_{P}[\phi(D,Z)\phi(D,Z)'] & 0_{d_{\theta} \times d_{\theta}} \\ E_{P}[t(Z)]' & 0_{1 \times d_{\theta}} \\ \mathbf I_{d_{\theta}} & \mathbf I_{d_{\theta}} \end{pmatrix} \quad and \quad \beta(P) = \begin{pmatrix} E_{P}[\phi(D,Z)Y] \\ \tau_0 \\ 1_{d_{\theta} \times 1} \end{pmatrix}, \end{align*} where $0_{d_{\theta} \times d_{\theta}}$ is a $d_{\theta}$-dimensional square matrix of zeros, $\theta_{1 \times d_{\theta}}$ is a $d_{\theta}$-dimensional row vector of zeros, $1_{d_{\theta} \times 1}$ is a $d_{\theta}$-dimensional column vector of ones, and $\mathbf I_{d_{\theta}}$ is a $d_{\theta}$-dimensional identity matrix. The new rows relative to (ref) correspond to the constraint $\theta + s \leq 1$, which requires $\theta \leq 1$ because the slack variable $s$ is non-negative.
example[Synthetic parallel trends with convex weights] Consider the causal panel data setting presented in liu2025synthetic. There are $K$ aggregate units indexed by $k\in\{1,\dots,K\}$, observed over periods $t\in\{1,\dots,T_0,T\}$, where $\{1, \ldots, T_0\}$ denote pre-treatment periods and $T$ denotes a treatment period at which unit $k=1$ is treated (units $k\ge 2$ are never treated). Let $\tilde \mu_t^k(1)$ and $\tilde \mu_t^k(0)$ denote potential aggregate outcomes for unit $k$ at time $t$ with and without treatment, respectively, and let the observed aggregate outcome be given by \[ \mu_t^k(P) = \tilde \mu_t^k(0) + \big(\tilde \mu_t^k(1)-\tilde \mu_t^k(0)\big)1\{k=1,t=T\}~. \] The target parameter is the effect on the treated unit at time $T$: \[ \tau = \tilde \mu_T^1(1)- \tilde \mu_T^1(0)~. \] liu2025synthetic maintains the assumption of (convex) synthetic parallel trends (SPT); that is, there exists a set of weights $(\omega_k: 2 \le k \le K) \in \mathbb R^{K-1}$ with $\sum_{2 \le k \le K}\omega_k = 1$, $\omega_k \ge 0$ for all $k$ such that for every $t\in\{2,\dots,T\}$, \[ \sum_{k=2}^K \omega_k\,\Delta\tilde \mu_t^k(0)=\Delta\tilde \mu_t^1(0)~, \] where $\Delta\tilde\mu^k_t(0) = \tilde\mu^k_t(0) - \tilde\mu^k_{t-1}(0)$. Let $\Delta \mu_t^k(P) = \Delta \tilde\mu_t^k(0)$ for all $k \geq 2$ and $t \geq 2$. Suppose we wish to test the null hypothesis $H_0: \tau = \tau_0$. Then, under the convex SPT assumption, we obtain the following system: \[ A_1(P) = \begin{pmatrix} \Delta\mu_{2}^{2}(P) & \cdots & \Delta\mu_{2}^{K}(P)\\ \vdots & \ddots & \vdots\\ \Delta\mu_{T_0}^{2}(P) & \cdots & \Delta\mu_{T_0}^{K}(P) \\ \Delta\mu_{T}^{2}(P) & \cdots & \Delta\mu_{T}^{K}(P) \\ 1 & \cdots & 1 \end{pmatrix} \hspace{3em} \beta(P) = \begin{pmatrix} \Delta\mu_{2}^{1}(P)\\ \vdots\\ \Delta\mu_{T_0}^{1}(P) \\ \mu^1_T(P) - \tau_0 - \mu^1_{T_0}(P) \\ 1 \end{pmatrix}~, \] and $A_0(P)$ does not exist. liu2025synthetic also analyzes the setting where we drop the assumption of convexity, so that $\omega_k$ are not restricted to be non-negative. In this case, $A_0(P)$ is given by the above matrix of aggregate-outcome differences and $A_1(P)$ does not exist.
example[Distribution-free binary choice] gurussell2023joe consider binary choice models of the form \begin{align} Y = 1\left\{\varphi(D, Z, U) \geq 0\right\} , \end{align} where $Y$ is a binary outcome, $\varphi$ is an unknown function, $D$ is an endogenous regressor, $Z$ is vector of exogenous regressors, and $U$ is a vector of unobservables. The distribution of $U$ is not restricted to lie in a parametric family, raising the possibility of partial identification. The authors observe that if $D$ and $Z$ are discrete, then the conditional distribution of $Y$ implied by the model is determined by the mass placed on a finite partition of the support of $U$ into sets $\mathcal{U}_{j}$, $j = 1,\ldots,d_{U}$. In particular: \begin{align} P\{Y = 1 \vert D = d, Z = z\} = \sum_{j=1}^{d_{U}} 1\left\{j \in \mathcal{J}(d, z)\right\} \theta_{j}(d,z) , \end{align} where $\theta_{j}(d,z)$ is the mass that the distribution of $U$ places on $\mathcal{U}_{j}$, conditional on $D = d, Z = z$, and the set $\mathcal{J}(d,z)$ collects the appropriate indices for sets that lead to $Y = 1$ when $D = d$ and $Z = z$. If the instrument $Z$ is independent with $U$, then also \begin{align} \sum_{d} \theta_{j}(d,z)P\{D = d \vert Z = z\} = \sum_{d} \theta_{j}(d,z')P\{D = d \vert Z = z'\} \quad for all $z, z'$, and $j$. \end{align} A natural target parameter in this model is the counterfactual choice probability $\pi(P) := P\{Y(d^{\star}) = 1\} = P\{\varphi(d^{\star}, Z, U) \geq 0\}$ at some fixed $d^{\star}$, which can also expressed as a linear function of the $\theta_{j}(d,z)$ if the sets $\mathcal{U}_{j}$ have been constructed to be sufficiently fine: \begin{align} \pi(P) = \sum_{j=1}^{d_{U}} E_{P} \left[ 1\left\{ j \in \mathcal{J}_{\pi}(d^{\star},Z) \right\} \theta_{j}(D,Z) \right] = \sum_{j=1}^{d_{U}} \sum_{d, z} 1\left\{ j \in \mathcal{J}_{\pi}(d^{\star},z) \right\} \theta_{j}(d,z)P\{D = d, Z = z\} . \end{align} Similar observations have been used for multinomial choice models by manski2007ier, tebalditorgovitskyyang2023e, and gurussellstringham2024. Suppose we wish to test the null hypothesis $H_{0}: \pi(P) = \pi_{0}$ for a vector of probabilities $\{\theta_{j}(d,z)\}_{j,d,z}$ that satisfies (ref) and (ref) when $D$ and $Z$ are both discrete. This problem fits into form (ref) with $x_{0}$ null, $x_{1}$ taken to be the $\theta_{j}(d,z)$ arranged in a vector across $(j,d,z)$, and $A_{1}(P)$ and $\beta(P)$ constructed from the linear functions (ref)--(ref), with (ref) set equal to $\pi_{0}$ and additional sum-to-one constraints for each $(d,z)$. The rows of $A_{1}(P)$ corresponding to (ref)--(ref) both depend on $P$.

A Useful Characterization of $\mathbf P_0$

figure[figure omitted — 1,266 chars of source]

In this section, we provide a characterization of $\mathbf P_0$ that informs the construction of the test we present in Section (ref). Before discussing the characterization formally, we motivate the need for such a characterization by demonstrating the impossibility of testing the null hypothesis that $P\in \mathbf P_0$ in a seemingly simple example. In particular, consider the case in which, for some scalars $a(P)$ and $b(P)$, the set $\mathbf{P}_0$ is given by \[\mathbf{P}_0 = \{P \in \mathbf{P}: a(P)x = b(P) \text{ for some } x \in \mathbb{R}\}~.\] This is a special case of (ref) where $d_0 = 1$ and $A_1(P)$ does not exist. In Figure (ref)(i) we plot the set \[\mathbf{C}_0 = \{(a, b) \in \mathbb{R}^2: ax = b \text{ for some } x \in \mathbb{R}\}~,\] which represents the set of points $(a(P),b(P)) \in \mathbb{R}^2$ for which there exists a distribution $P \in \mathbf{P}_0$. From Figure (ref)(i) we notice immediately that the closure of $\mathbf{C}_0$ (as a subset of $\mathbb{R}^2$) is the entire space $\mathbb{R}^2$. This observation suggests that the closure of $\mathbf{P}_0$ (with respect to the total variation metric) coincides with the entire set of distributions $\mathbf{P}$, and indeed we discuss conditions under which this is the case below. In contrast, Figure (ref)(ii) depicts the analogous set $\mathbf{C}_0$ for testing the null hypothesis given by \[\mathbf{P}_0 = \{P \in \mathbf{P}: a(P)x = b(P) \text{ for some } x \in \mathbb{R}, x \ge 0\}~,\] which is a special case of (ref) where $d_1$ = 1 and $A_0(P)$ does not exist. Here, we see that the closure of $\mathbf{C}_0$ (as a subset of $\mathbb{R}^2$) is a strict subset of $\mathbb{R}^2$, which suggests that the closure of $\mathbf{P}_0$ in this example is a strict subset of $\mathbf{P}$.

Because it is impossible to test a null hypothesis for which $\mathbf{P}_0$ is dense in $\mathbf{P}$ with respect to the total variation metric, and more generally that it is impossible for any test to have non-trivial power against alternatives which lie on the boundary of $\mathbf P_0$ romano2004non, our test is based on a characterization of the closure of $\mathbf P_0$ relative to the total variation metric, which we denote by $\mathrm{cl}(\mathbf P_0)$. Towards that end, define

align[align omitted — 250 chars of source]

so that $\mathbf P_0 = \{P \in \mathbf P: (A_0(P),A_1(P),\beta(P)) \in \mathbf{C}_0\}$. Accordingly, we begin by deriving a characterization of the closure of $\mathbf {C}_0$ with respect to the Euclidean topology, and then relate this characterization back to $\mathrm{cl}(\mathbf P_0)$. First, consider the following alternative representation of $\mathbf{C}_0$ based on pre-multiplying the equation in (ref) by the annihilator of $A_0$, which we denote by $M_0$.

lemmaLet $M_0$ denote the projection operator onto the orthogonal complement of the column space of $A_0$. Then \[\mathbf C_0 = \{(A_0, A_1, b): M_0 A_1 x_1 = M_0b \text{ for some } x_1 \in \mathbb R^{d_1}, x_1 \geq 0\}~.\]

Let $a_1 , \dots, a_{d_1}$ denote the columns of $A_1$. The columns of $M_0A_1$ are then given by $M_0a_1, \ldots, M_0a_{d_1}$. Given this alternative representation of $\mathbf C_0$, we can apply Farkas' lemma to conclude that $(A_0,A_1,b) \in \mathbf C_0$ if and only if for all $y \in \mathbb R^p$, either there exists some $1 \le j \le d_1$ such that $a_j'M_0y < 0$ or $b' M_0 y \geq 0$. Consider the set of triples $(A_0, A_1, b)$ obtained by weakening the strict inequalities $a_j'M_0y < 0$ to weak inequalities:

equation[equation omitted — 270 chars of source]

Theorem (ref) formalizes the sense in which replacing these strict inequalities with weak inequalities relates to the closure of the set $\mathbf{C}_0$.

theoremLet \begin{equation} \mathbf{C}^{\rm RD} := \{(A_0, A_1, b) \in \mathbb R^{p \times d_0}\times \mathbb R^{p \times d_1} \times \mathbb R^p: \mathrm{rank}(A_0) < d_0\} . \end{equation} Then, \[\mathrm{cl}(\mathbf{C}_0) = \bar{\mathbf{C}}_0 \cup \mathbf{C}^{\rm RD} ~,\] where $\mathrm{cl}(\mathbf{C}_0)$ denotes the closure of $\mathbf{C}_0$ in the Euclidean topology.

Theorem (ref) shows that the closure of $\mathbf{C}_0$ can be characterized by combining the set $\bar{\mathbf{C}}_0$, which describes the set of triples obtained by weakening the inequalities in the conclusion of Farkas' lemma, along with the set of triples for which $A_0$ is rank deficient. Note the theorem implies that, if $p < d_0$, then the closure of $\mathbf{C}_0$ becomes $\mathbb R^{p \times d_0}\times \mathbb R^{p \times d_1} \times \mathbb R^p$. As a result, we implicitly assume that $p \ge d_0$ for the rest of the paper. This characterization of the closure forms the basis for the test which we present in Section (ref).

Finally, we relate $\mathrm{cl}(\mathbf{C}_0)$ to the closure $\mathrm{cl}(\mathbf{P}_0)$ in the total variation distance. Let $\widetilde{\mathbf P}_0 := \{P \in \mathbf P: (A_0(P), A_1(P), \beta(P)) \in \mathrm{cl}(\mathbf C_0)\}$ denote the pre-image of $\mathrm{cl}(\mathbf C_0)$ in $\mathbf{P}$. We claim that in general $\mathbf P_0 \subseteq \widetilde{\mathbf P}_0 \subseteq \mathrm{cl}(\mathbf P_0)$. Indeed, it follows by construction that $\mathbf{P}_0 \subseteq \widetilde{\mathbf P}_0$. As a result, any test that controls size on $\widetilde{\mathbf P}_0$ will necessarily control size on $\mathbf{P}_0$. Meanwhile, we generally expect $\widetilde{\mathbf P}_0 \subseteq \mathrm{cl}(\mathbf P_0)$, in which case we will not have power against any distribution in $\widetilde{\mathbf P}_0$. By definition, this will be the case whenever $\mathbf P$ is “rich” enough in the sense that for every $P \in \mathbf P$ such that $(A_0(P), A_1(P), \beta(P)) \in \mathrm{cl}(\mathbf C_0)$, there exists a sequence $P_n \in \mathbf P_0$ such that $P_n$ converges to $P$ in the total variation metric.

Low-level sufficient conditions for this “richness" property could be obtained, for example, if we view $P$ as the distribution of a random vector in $\mathbb R^{p(d_0 + d_1 + 1)}$ and define $\mathrm{vec}(A_0(P), A_1(P), \beta(P))$ to be the corresponding vector of means, where $\mathrm{vec}$ is the vec-operator magnus2019matrix. In this case, $\mathbf P$ is rich enough if it contains a sufficiently large collection of normal location families. To illustrate, let $\mu = \operatorname*{vec}(A_0(P), A_1(P), \beta(P))$ where $(A_0(P), A_1(P), \beta(P)) \in \mathrm{cl}(\mathbf{C}_0)$. Then by definition there exists a sequence of vectors $\mu_n$ corresponding to triplets in $\mathbf{C}_0$ such that $\mu_n \rightarrow \mu$. Let $H = \mathrm{span}\{\mu_n - \mu: n \ge 1\}$ and let $\Sigma$ be a positive semi-definite matrix whose range is exactly $H$, so that $\mu_n - \mu \in \mathrm{range}(\Sigma)$ for all $n \ge 1$. Then, the sequence of distributions $N(\mu_n, \Sigma)$ converges to $N(\mu, \Sigma)$ in the total variation metric by Lemma (ref) in the appendix.

remarkFollowing fang2021inference, it may seem natural to first transform (ref) into standard form \begin{equation} \mathbf P_0 = \{ P \in \mathbf P : (A(P), \beta(P)) \in \mathbf{C}^{\rm alt}_0 \} , \end{equation} where $\mathbf{C}^{\rm alt}_0 = \{(A,b) \in \mathbb R^{p \times (2d_0 + d_1)}\times\mathbb R^p: Ax = b \text{ for some } x \in \mathbb R^{2d_0 + d_1}, x \geq 0\}$. Indeed, given a triple $(A_0, A_1, b) \in \mathbf{C}_0$, we can obtain a pair $(A,b) \in \mathbf{C}^{\rm alt}_0$ by defining $A = (A_0 \quad {-A_0} \quad A_1)$. However, this transformation would not help provide a useful characterization of the closure as presented in this section. Let $\tilde{a}_j$ for $1 \le j \le 2d_0 + d_1$ denote the columns of $A$. By applying the reasoning we used to obtain $\bar{\mathbf{C}}_0$ in (ref), we obtain the set \[\bar{\mathbf{C}}^{\rm alt}_0 = \left\{(A,b) \in \mathbb R^{p \times (2d_0 + d_1)}\times\mathbb R^p: \sup_{y \in \mathbb R^p} \min \left \{ \min_{1 \leq j \leq 2d_0 + d_1} \tilde a_j' y, -b'y \right \} \leq 0\right\}~.\] Note that the $d_0 + 1$ to $2d_0$-th columns of $A$ are simply the negatives of the first $d_0$ columns, so that there always exists a $j$ for which $\tilde{a}_j'y \leq 0$. As a result, $\bar{\mathbf{C}}^{\rm alt}_0$ recovers the entire space $\mathbb R^{p \times (2d_0 + d_1)}\times\mathbb R^p$ and thus does not help to provide a useful characterization of the closure.

The Test

Description of the Test

Let $\{Z_i\}_{i=1}^n$ be i.i.d.\ with $Z_i$ distributed according to $P \in \mathbf P$. Recall from Section (ref) that the closure of the null space can be characterized as the set of distributions $P$ for which either $A_0(P)$ is rank-deficient, or

equation[equation omitted — 151 chars of source]

Recognizing that whether condition (ref) holds is not affected by norm constraints on $y$, and defining $b_j(P) := a_j(P)$ for $1 \le j \le d_1$, $b_{d_1+1}(P) := -\beta(P)$, and $J = \{1, 2, \ldots, d_1 + 1\}$, we may rewrite (ref) as

equation[equation omitted — 158 chars of source]

where $\|y\|_1 = \sum_{i=1}^p |y_i|$ for any $(y_1,\ldots, y_p)\in \mathbb R^p$. In what follows, we consider the following equivalent formulation of (ref): For any $J^* \subseteq J$ and $(J^*)^c := J \setminus J^*$, condition (ref) is equivalent to the statement

equation[equation omitted — 208 chars of source]

Our test uses a sample-splitting procedure in which one sample split is used to select a $y\in \mathbb R^p$ and a second split is used to test whether the inequalities in (ref) hold at the selected $y$. We consider two proposals for $(J^*)^c$. Our first proposal takes $(J^*)^c$ to be those $j\in J$ for which $b_j(P)'M_0(P)$ is known deterministically -- i.e.\ for which $b_j(P)^\prime M_0(P)$ does not depend on $P$. In other words, we test the condition $\min_{j \in J^*} b_j(P)'M_0(P)y \le 0$ using a unit vector $y$ that is known to satisfy $\min_{j \in (J^*)^c}b_j(P)'M_0(P)y > 0$. We call this method the “direct” method in what follows. Our second proposal sets $J^*=\{j^*\}$ for some non-random $j^*$. In this case, we test the condition $b_{j^*}(P)'M_0(P)y \le 0$ using a unit vector $y$ such that $\min_{j \in (J^*)^c}b_j(P)'M_0(P)y > 0$ holds with high probability. We call this method the “screening” method. In all of our examples, we set $j^* = d_1 + 1$, which is natural in settings where we perform test inversion to construct a confidence set for a scalar parameter whose null value only enters the vector $b_{d_1 + 1}(P)$; see, e.g., the examples in Section (ref). We show via simulation in Section (ref) that the screening method often generates shorter confidence intervals than the direct method, at the cost of introducing an additional tuning parameter which determines the amount of “screening” that is performed in the first sample split.

We next present a high-level description of the test and defer the details of the construction of its specific components to Sections (ref) and (ref). To construct the test, we first randomly split the data into two samples $\{Z_i\}_{i \in I_{1, n}}$, $\{Z_i\}_{i \in I_{2, n}}$ of sizes $n_1$ and $n_2$, where $I_{1, n} \cup I_{2, n} = \{1, \dots, n\}$ and $I_{1, n} \cap I_{2, n} = \emptyset$. In what follows, we always assume that $n_2 \to \infty$ as $n \to \infty$ and allow $n_1$ to be fixed for the direct method but require $n_1 \to \infty$ for the screening method. Throughout, we use the superscript $(k)$ to denote when a given quantity is a function of only the $k$th split. Using the first sample split $\{Z_i\}_{i\in I_{1,n}}$, we construct a vector $\hat y_n^{(1)}$ which represents a direction in which the weak inequality in (ref) appears to be “most violated” --- we discuss how to construct such a vector in Section (ref).

Next, given suitable estimators $\hat{b}_{j, n}^{(2)}$ and $\hat{M}_{0, n}^{(2)}$ for $b_j(P)$ and $M_0(P)$ computed in the second sample split $\{Z_i\}_{i\in I_{2,n}}$, we define the test statistic

equation[equation omitted — 179 chars of source]

where $\hat{\sigma}_{j, n}^{(2)}(\hat y_n^{(1)})$ is an estimator for the asymptotic standard deviation of $\sqrt{n_2}(\hat b_{j, n}^{(2)})'\hat{M}_{0, n}^{(2)} \hat y_n^{(1)}$. Finally, we set

equation[equation omitted — 70 chars of source]

where, for $\alpha \in (0, 1)$, $z_{1 - \alpha}$ denotes the $1 - \alpha$ quantile of a standard normal distribution.

remarkIt may occur that for some $j\in J^*$ the asymptotic variance of $\sqrt{n_2}(\hat b_{j, n}^{(2)})'\hat M_{0,n}^{(2)}\hat y_n^{(1)}$ is zero. To avoid degeneracy of the corresponding standard error, we define $\hat\sigma_{j, n}^{(2)}(\hat y_n^{(1)})$ using a small truncation, as described in Section (ref). This modification is technically motivated and typically has no effect on the test in practice. Indeed, if the numerator of (ref) is negative for any $j \in J^*$, then we fail to reject for any choice of truncation. If the numerator of (ref) is positive for all $j \in J^*$ then the choice of truncation can only induce a failure to reject if the estimated variance falls below the truncation threshold for at least one index that attains the minimum in the truncated version of $T_n$.
remarkIn practice, researchers may want to reduce the uncertainty introduced by sample splitting by aggregating the test results obtained from multiple different splits of the data. This can be accomplished by appropriately aggregating the (upper bounds on) $p$-values produced by the test described in (ref). To that end, we have found the exchangeable improvement to the “twice the average” $p$-value, as described in gasparin2025combining, works well in simulations.

Properties of the Test

In this section, we present results establishing the asymptotic validity of our test and a detailed construction of the standard deviation in (ref). When stating our assumptions, we will suppress the superscript $(k)$, with the understanding that all assumptions stated on the entire sample will also hold when applied on the sample splits $\{Z_i\}_{I_{1, n}}$ and $\{Z_i\}_{i\in I_{2, n}}$, under suitable scaling.

Our first assumption imposes conditions on the estimators $\hat A_{0,n}$ for $A_0(P)$ and $\hat b_{j,n}$ for $b_j(P)$ with $1\leq j \leq d_1+1$. In its statement, $\|\cdot\|_2$ denotes the Euclidean norm, $\|\cdot\|_{2, 2}$ denotes the operator norm of a matrix when the domain and range are endowed with $\|\cdot\|_2$, and $a \vee b = \max(a, b)$ for any $a, b \in \mathbb R$.

assumptionLet $\{Z_i\}_{i=1}^n$ be i.i.d.\ with marginal distribution $P \in \mathbf{P}$. Then, \begin{packed_enum} • There are $\Psi(Z_i, P) \in \mathbb R^{p \times d_0}$ with $E_P[\Psi(Z_i, P)] = 0$, $\varphi_j(Z_i, P) \in \mathbb R^p$ with $E_P[\varphi_j(Z_i, P)] = 0$ for $1 \leq j \leq d_1 + 1$, and $a_n$ for which $a_n / \sqrt n \to 0$ such that uniformly in $P\in \mathbf P$, \begin{align} \Big \| \sqrt n(\hat A_{0, n} - A_0(P)) - \frac{1}{\sqrt n} \sum_{1 \leq i \leq n} \Psi(Z_i, P) \Big \|_{2, 2} & = O_P(a_n / \sqrt n) \\ \max_{1 \leq j \leq d_1 + 1} \Big \| \sqrt n(\hat b_{j, n} - b_j(P)) - \frac{1}{\sqrt n} \sum_{1 \leq i \leq n} \varphi_j(Z_i, P) \Big \|_2 & = O_P(a_n / \sqrt n) . \end{align} • For each $p \geq 1$, there are $1\leq K_{0, p}, K_{1, p} < \infty$ such that $\|\Psi(Z_i, P)\|_{2,2} \leq K_{0, p}$ with probability one and \begin{align*} & \sup_{P \in \mathbf P} \Big ( \| E_P[\Psi(Z_i, P)\Psi(Z_i, P)']\|_{2,2} \vee \|E_P[\Psi(Z_i, P)' \Psi(Z_i, P)]\|_{2,2} \\ & \vee \max_{1 \le j \le d_1 + 1}E_P[\varphi_j(Z_i, P)'\varphi_j(Z_i, P)] \Big ) \leq K_{1, p} . \end{align*} \end{packed_enum}

Assumption (ref)(a) requires our estimators for $A_0(P)$ and $b_j(P)$ for $1 \leq j \leq d_1 + 1$ to be asymptotically linear with influence functions whose moments are disciplined by Assumption (ref)(b). Assumption (ref)(a) is automatically satisfied with $a_n = 0$ whenever the entries of $A_0(P)$ and $b_j(P)$ are expectations and the entries of $\hat A_{0,n}$ and $\hat b_{j,n}$ are the corresponding sample means.

Our second assumption imposes boundedness and non-degeneracy conditions on $A_0(P)$ and $b_j(P)$. In its statement, $\bar s(A_0(P))$ and $\underline{s}(A_0(P))$ denote the maximum and minimum singular value of $A_0(P)$.

assumption$A_0(P)$ and $b_j(P)$ are such that \begin{packed_enum} • $\sup_{P \in \mathbf P} \bar s(A_0(P)) \leq \bar s_p$ for some $\bar s_p$ satisfying $1\leq \bar s_p < \infty$. • $\inf_{P \in \mathbf P} \underline s(A_0(P)) \geq \underline s > 0$ for some $\underline s$ not depending on $p$. • $\sup_{P \in \mathbf P} \max_{1 \leq j \leq d_1 + 1} \|b_j(P)\|_2 \leq K_{2, p}$ for some $K_{2,p}$ satisfying $1\leq K_{2,p} < \infty$. \end{packed_enum}

Assumption (ref)(a) requires the maximum singular value of $A_0(P)$ is bounded above uniformly in $\mathbf{P}$, with the bound possibly depending on $p$. Assumption (ref)(b) ensures that $A_0(P)'A_0(P)$ is bounded away from degeneracy. This rules out the set of rank deficient matrices $\mathbf{C}^{\rm RD}$ in our characterization of the closure of the null hypothesis, as defined in Theorem (ref). Assumption (ref)(c) requires that the Euclidean norm of $b_j(P)$ is bounded uniformly in $P \in \mathbf P$, with the bound possibly depending on $p$.

We note that Assumption (ref)(b) can be dropped if we let $n_1 \to \infty$ and apply the test $\phi_n$ in (ref) only when $\underline s(\hat{A}^{(1)}_{0,n}) > \tau$ for some pre-specified small value $\tau > 0$ (and do not reject the null hypothesis when $s(\hat{A}^{(1)}_{0,n}) \leq \tau$). We emphasize, however, that even if we maintain Assumption (ref)(b), our assumptions impose no requirements on the rank of $A_1(P)$. In contrast, the assumptions underlying cox2025testing and goff2025inference implicitly restrict the ranks of both $A_0(P)$ and $A_1(P)$.

Assumptions (ref) and (ref) are instrumental in obtaining an asymptotic expansion for our test statistic. In particular, letting $A_0^\dagger(P)$ denote the Moore-Penrose pseudoinverse of $A_0(P)$, we will show that for every $1\leq j \leq d_1+1$ the vector $\sqrt n(\hat b_{j, n}' \hat M_{0, n} - b_j(P)' M_0(P))'$ is asymptotically linear with influence function

equation[equation omitted — 170 chars of source]

Our third assumption imposes moment restrictions on the influence function $\xi_j(Z,P)$.

assumptionThere is a constant $K_\xi < \infty$ not depending on $p$ such that the following holds: $$\sup_{P \in \mathbf P} \sup_{\|y\|_1\leq 1}\max_{1\leq j \leq d_1+1} E_P[|\xi_j(Z_i, P)'y|^{3}]^{1/3} \leq K_{\xi}.$$

Assumption (ref) ensures that we are able to couple an influence function for our test statistic to a Gaussian random variable uniformly in $P\in \mathbf P$. The requirement of Assumption (ref) is satisfied, for example, if the third moments of the entries of $\xi_j(Z,P)$ are uniformly bounded across $1\leq j\leq d_1+1$ and $P\in \mathbf P$.

Our fourth assumption ensures that the inequalities not examined by our test statistic (i.e., those in $(J^*)^c$) are indeed positive when evaluated at a $\hat y^{(1)}_n$ not equal to zero, as required by the characterization of the null hypothesis in (ref). We describe methods to construct such a $\hat y_n^{(1)}$ in Section (ref) below.

assumptionFor $\mathcal{Y}(P;J^*) := \{y \in \mathbb R^p : \|y\|_1 \leq 1 \text{ and } b_j^\prime(P)M_0(P)y > 0 \text{ for all } j\in (J^*)^c\}$, we have \[ \lim_{n \to \infty} \inf_{P \in \mathbf P_0} P \big \{ \{\hat y_n^{(1)} \in \mathcal{Y}(P;J^*)\} \cup \{\hat y_n^{(1)} = 0 \} \big \} = 1~. \]

Assumptions (ref) automatically holds, for instance, for the direct method which either sets $J^*$ to equal $J$ (so $(J^*)^c$ is empty) or $(J^*)^c$ to only contains coordinates $j$ for which $b_j^\prime M_0(P)$ is known. In this case, $\hat y_n^{(1)}$ can be chosen to belong to $\mathcal{Y}(P;J^*)$ with probability one for any $n_1$, and our asymptotics only require that $n_2 \to \infty$. In contrast, the screening method intuitively conducts a pre-test in the first fold to ensure that $\hat y_n^{(1)}\in \mathcal Y(P;J^*)$ with high probability. In this case, our asymptotics therefore require $n_1\to \infty$ in order for Assumption (ref) to be satisfied. We also note that Assumption (ref) allows $\hat y_n^{(1)}$ to equal zero when it does not belong to $\mathcal Y(P;J^*)$. This flexibility is important because our test never rejects when $\hat y_n^{(1)}$ is zero (since then $T_n =0$; see (ref)). Therefore, setting $\hat y_n^{(1)} =0$ allows us to decide not to reject after examining the first sample split --- e.g., if we fail to reject in the screening method pre-test.

We now turn to the construction of the standard error $\hat \sigma_{j,n}(y)$ employed in the construction of our test statistic. To this end, we let $\operatorname*{vec}$ denote the vec-operator and $\otimes$ denote the Kronecker product magnus2019matrix. We further define the asymptotic covariance matrix \[ V_j(P) := \operatorname{Var}_P \bigg [

pmatrix[pmatrix omitted — 66 chars of source]

\bigg ] \] for our estimator $\hat A_{0,n}$ and $\hat b_{j,n}$ and let $\hat V_{j, n}$ denote an estimator for $V_j(P)$. For each fixed $y \in \mathbb R^p$ and $1\leq j \leq d_1+1$, we apply the Delta method to the function $(A_0(P), b_j(P)) \mapsto b_j(P)' M_0(P) y$ to obtain the asymptotic variance of the estimator $\hat b_{j,n}^\prime \hat M_{0,n}y$. Accordingly, defining the gradient \[ D_j(P; y) :=

pmatrix[pmatrix omitted — 126 chars of source]

, \] it is possible to show that the asymptotic variance of $\hat b_{j,n}^\prime \hat M_{0,n}y$ equals $\sigma_j^2(P; y) := D_j(P; y)' V_j(P) D_j(P; y)$. By analogy, for $\hat A_{0,n}^\dagger$ the Moore-Penrose pseudoinverse of $\hat A_{0,n}$, we estimate $D_j(P; y)$ by setting \[ \hat D_{j, n}(y) =

pmatrix[pmatrix omitted — 175 chars of source]

. \] As an estimator for the asymptotic variance we then set $\hat{\sigma}^2_{j, n}(y) = (\hat D_{j, n}(y)' \hat{V}_{j, n} \hat D_{j, n}(y)) \vee \underline{\sigma}^2$ for $1 \leq j \leq d_1 + 1$, where, as previously discussed in Remark (ref), $\underline{\sigma} > 0$ is a small positive constant which we include to address potential (near) degeneracies in $\sigma^2_j(P;y)$.

Our fifth assumption imposes that our estimator $\hat V_{j,n}$ is suitably uniformly consistent.

assumption$\|\hat V_{j, n} - V_j(P)\|_{2, 2} = o_P(K_{2,p}^{-2})$ uniformly in $P \in \mathbf P$ and $1 \leq j \leq d_1 + 1$.

Assumption (ref) enables us to show that the standard errors $\hat \sigma_{j,n}(y)$ are suitably uniformly consistent in both $P\in \mathbf P$ and $1\leq j \leq d_1+1$. It requires that $\hat V_{j,n}$ converge to $V_j(P)$ at a rate faster than $K_{2,p}^2$, where $K_{2,p}$ depends on $p$ and is specified in Assumption (ref)(c). We state Assumption (ref) as a high level condition because the structure of $V_j(P)$ is dictated by the influence functions of our estimators (as introduced in Assumption (ref)). In applications for which our estimators are sample means, and hence $\hat V_{j,n}$ are sample covariance matrices, sufficient conditions for Assumption (ref) are readily available from the literature; see, e.g., Chapter 6 in wainwright2019high.

Our final assumption imposes conditions on the rates of convergence of our moment and error bounds.

assumptionThe following rate restrictions hold as $n \to \infty$: \begin{packed_enum} • $\displaystyle K_{2,p}(K_{0,p}\vee K_{1,p})\log(1+p) = o(\sqrt {n_2})$. • $\bar s_p^2(K_{0,p}\vee K_{1,p})\log(1+p) = o(n_2)$. • $\displaystyle K_{2,p} a_{n_2} = o(\sqrt {n_2})$. \end{packed_enum}

In Remark (ref) below, we discuss high-level sufficient conditions which guarantee Assumption (ref) holds, as well as how these conditions differ when $A_0(P)$ needs to be estimated versus when it is known (or does not exist). At this point we emphasize, however, that Assumptions (ref) are automatically satisfied in asymptotic regimes in which $p$ and $d_1$ are fixed. Moreover, we note that Assumption (ref) only restricts the dimension $d_1$ through Assumptions (ref) and (ref)(c), which impose that the linearization error, the second moment of the influence function $\varphi_j(Z,P)$, and the norm of $\|b_j(P)\|_2$ be bounded uniformly in $1\leq j \leq d_1$. If these bounds do not grow with $d_1$, then Assumption (ref) in fact leaves the dimension $d_1$ unrestricted. As we discuss in the next section, however, certain approaches for selecting $\hat y_n^{(1)}$ may impose restrictions on the dimension $d_1$ relative to the sample size in the first split $\{Z_i\}_{i\in I_{n,1}}$.

Our next, main, result establishes that our proposed test is uniformly consistent in level over $\mathbf P_0$.

theoremSuppose Assumptions (ref)--(ref) hold. Then, for any $\alpha < 0.5$ it follows that \[ \limsup_{n \to \infty} \sup_{P \in \mathbf P_0} E_P[\phi_n] \leq \alpha~. \]
remarkWhile the constants specified by Assumptions (ref) and (ref) are application specific, in many instances we can expect the bounds $K_{0, p} \vee K_{1, p} \lesssim p$, $K_{2, p} \vee \overline{s}_p \lesssim \sqrt{p}$, and $a_n \lesssim p$ (or $a_n =0$) to hold. Under these conditions, Assumption (ref) holds provided that $p^3 = o(n_2)$ (up to logs), while Assumption (ref) can also be shown to hold when $p^3 = o(n_2)$ (up to logs) under suitable moment restrictions. These rates simplify when the matrix $A_0(P)$ does not exist or does not depend on $P$. In this case, an inspection of the proof of Theorem (ref) reveals that Assumptions (ref)(a)(b) are no longer needed, Assumption (ref)(c) can be weakened to $a_n = o(\sqrt{n_2})$, and Assumption (ref) can be weakened by replacing $K_{2,p}$ with one. Thus, when $A_0(P)$ does not exist or is known, Theorem (ref) requires $p =o(n_2)$ and $a_n = o(\sqrt{n_2})$, with the latter requirement automatically holding when there is no linearization error or being implied by $p^2 = o(n_2)$ when $a_n \lesssim p$.

Procedure for selecting $\hat y_n^{(1)}$

In this section we describe a procedure for selecting $\hat y_n^{(1)}$ that satisfies Assumption (ref). Given suitable estimators $\hat{b}_{j, n}^{(1)}$ and $\hat{M}_{0, n}^{(1)}$, we select $\hat y_n^{(1)}$ as the solution to the following optimization problem:

equation[equation omitted — 360 chars of source]

where $\hat{\omega}_{j, n} > 0$ are positive scaling parameters that we specify below. If the optimization problem in (ref) is found to be infeasible, then we set $\hat y_n^{(1)} = 0$. In Appendix (ref), we explain how this optimization problem can be re-formulated as a linear program. Note that the specific choice of $\hat\omega_{j,n}$ for $j \in J^*$ will not affect the validity of the procedure, although these should be carefully selected to ensure the test has good power; see the discussion following Lemma (ref) for details. In contrast, as we demonstrate in Lemma (ref), the choice of $\hat \omega_{j,n}$ for $j \in (J^*)^c$ is relevant for the screening method. In the case of the direct method, that is, when $(J^*)^c$ contains only those $j$ for which $b_j(P)'M_0(P)$ is known, the specific choice of $\hat{\omega}_{j, n}$ for $j \in (J^*)^c$ is immaterial in practice: we can simply set the constraints in (ref) to ensure that $b_j(P)'M_0(P)y$ is strictly positive for every $j \in (J^*)^c$.

lemmaLet $\hat{y}_n^{(1)}$ be defined as in (ref) and $\hat{\omega}_{j, n} > 0$ for all $j \in J$. Then, the following statements hold: \begin{packed_enum} • Suppose that $b_j(P)'M_0(P)$ is known for all $j \in (J^*)^c$. Then, it follows that Assumption (ref) holds. • Suppose Assumptions (ref), (ref), (ref) and (ref) hold, $K_{1,p}(K_{0, p} \vee K_{1, p})\log (1 + p) d_1^{1/3} = o(np^{2/3})$, and $\max_{j \in (J^*)^c} \{(pd_1)^{1/3}/\hat \omega_{j,n}\}= o_P(1)$ uniformly in $P\in \mathbf P$. Then, it follows that Assumption (ref) holds. • Suppose Assumptions (ref), (ref), (ref) and (ref) hold, and that the following moment restrictions hold \begin{equation} \sup_{P\in \mathbf P} E_P\left[\max_{1\leq j \leq d_1+1} \|\xi_j(Z,P)\|_\infty^2\right] \leq C_{\xi,p}^2 \sup_{P\in \mathbf P} E_P\left[\max_{1\leq j \leq d_1+1}\|\varphi_j(Z,P)\|_2^2\right] \leq C_{\varphi,p}^2 . \end{equation} If $pC_{\varphi,p}^2\log^2(p+d_1)(K_{0,p}\vee K_{1,p}) = o(n)$ and $\max_{j \in (J^*)^c} \{C_{\xi,p}\sqrt{\log(p+d_1)}/\hat \omega_{j,n}\}= o_P(1)$ uniformly in $P\in \mathbf P$, then it follows that Assumption (ref) holds. \end{packed_enum}

Part (a) of Lemma (ref) formally states that for the direct method, the sole constraint on the weights $\hat {\omega}_{j,n}$ is that they be positive. Parts (b) and (c) specify the requirements on the weights $\hat \omega_{j,n}$ when the coordinates $j\in (J^*)^c$ may be such that $b_j(P)^\prime M_0(P)$ depends on $P$, as in the screening method. In particular part (b) shows that Assumption (ref) is satisfied under rate restrictions on the growth of $d_1$ and that $\hat \omega_{j,n}$ diverge to infinity faster than $(pd_1)^{1/3}$. As we show in part (c), however, these restrictions can be considerably weakened provided we strengthen our moment conditions by assuming that (ref) holds. Under such a requirement, part (c) weakens the rate restrictions on $d_1$ and the rate at which $\hat {\omega}_{j,n}$ must diverge to be logarithmic.

Although the conditions of Lemma (ref) are satisfied for a wide range of choices of $\hat \omega_{j,n}$, careful choices of these weights should be used to ensure that our test has good power properties. First, the weights $\hat{\omega}_{j, n}$ for $j \in J^*$ should be selected in a way that “appropriately weights” the degree of violation of each component. For this we set $\hat \omega_{j,n} = \hat \sigma_{j,n}^{(1)}(\hat y_{0,n})$ for $\hat \sigma_{j,n}^{(1)}(y)$ our (truncated) estimate of the asymptotic standard deviation of $\sqrt{n_1}(\hat b_{j,n}^{(1)})^\prime \hat M_{0,n} y$ evaluated at $\hat y_{0,n}$. To maintain the linear structure of the optimization problem, we employ a preliminary direction $\hat y_{0,n}$ computed by solving the linear program

equation[equation omitted — 315 chars of source]

The weights $\hat \omega_{j,n}$ for $j \in (J^*)^c$ should further be selected so that the assumptions of Lemma (ref) are satisfied and so that, under the alternative, (ref) is feasible with probability tending to one: for this we specify $\hat \omega_{j,n}$ to satisfy $\hat \omega_{j,n} = o_P(\sqrt{n_1})$. To this end, we've found that in simulations it works well to set $\hat \omega_{j,n} = c_n \cdot \hat \sigma_{j,n}^{(1)}(\hat y_{0,n})$ for some sequence $c_n$ diverging to infinity. We have found that the following choices work well in practice: in a low dimensional regime where $p$ and $d_1$ could be considered fixed, we set $c_n = \sqrt{\log(\log(n_1))}$; in a high dimensional regime where $p$ or $d_1$ diverges with sample size, we set $c_n = \sqrt{\log(\log(\log(n_1)))\times \log(p + d_1)}$.

remarkIn the case of the direct method, in which $b_j(P)'M_0(P)$ is known for all $j \in (J^*)^c$, it is straightforward to verify that Assumption (ref) is in fact satisfied for any $\hat y_n^{(1)}$ such that $$\sqrt{n_1}(\hat{b}_{j, n}^{(1)})'\hat{M}_{0,n}^{(1)}y \ge \eta \text{ for all } j \in (J^*)^c$$ for some $\eta > 0$, not just $\hat y_n^{(1)}$ defined in (ref) and discussed in Lemma (ref).

Simulations

In this section, we illustrate the finite-sample performance of our procedure via simulation. We consider three distinct settings, each inspired by earlier related papers. In all designs, we examine the performance of the direct and the screening method with $n_1 = n_2 = n/2$. In all cases we set $\hat \omega_{j,n}$ for $j \in J^*$ to equal $\hat \sigma^{(1)}_{j,n}(\hat y_{0n})$ for $\hat y_{0,n}$ as defined in (ref) and $\hat \omega_{j,n}$ for $j \in (J^*)^c$ to equal $c_n \cdot \hat \sigma^{(1)}_{j,n}(\hat y_{0n})$ with $c_n = \sqrt{\log(\log(\log(n_1)))\times \log(p + d_1)}$.

cox2025testing

We first consider the simple one-sided model described in cox2025testing. Let $C := (C_1, \ldots C_H)'$ and $X := (X_1, \ldots, X_H)'$ be random vectors and let $H \ge 0$ index the number of inequalities to be considered. In our notation, the design can be described as $A_0(P) = \nu(P) := (E_P[C_1], \dots, E_P[C_H])'$, $A_1(P) = \mathbf{I}_H$, $\beta(P) = -\mu(P) - \mathbf v \theta$, where $\mu(P) := (E_P[X_1], \ldots, E_P[X_H])'$ and $\mathbf v := (1, 1, 0, \dots, 0)' \in \mathbb R^H$. It can be shown that the identified set of $\theta$ is $(-\infty, 0]$. Under $P$, $(C,X)$ are distributed according to

align*[align* omitted — 67 chars of source]

with $\mu(P) = (-1, 1, 1, \ldots, 1)'$, $\nu(P) = (1, -1, -1, \ldots, -1)'$. Accordingly, given a random sample from $(X,C)$, $\hat{A}_{0,n}$ and $\hat{b}_{j,n}$ are computed by taking sample averages.

Figures (ref) and (ref) present the rejection probabilities from $1,000$ draws for both of our proposed tests for the null hypothesis that $P \in \mathbf P_0$ at a $5\%$ significance level, over a grid of values of $\theta$ and across different choices of dimension $H$ and sample size $n$ (the direct method is represented by the solid line and the screening method by the dashed line). We see that both the direct and screening methods control size in all cases within the identified set for $\theta$. Outside of the identified set, the rejection probabilities for the screening method are slightly larger than for the direct method, although the differences become negligible for large $H$ and/or large $n$.

figure[figure omitted — 1,375 chars of source]
figure[figure omitted — 1,073 chars of source]

goff2025inference

We revisit a simple instance of Example (ref) proposed by goff2025inference, which is loosely based on the application in mogstadsantostorgovitsky2017nwp to the bed net data used by dupas2014e. The outcome $Y$ is a binary indicator of using a new type of anti-malarial bed net, the treatment $D$ is an indicator for purchasing the bed net, and the instrument $Z$ is a randomly-assigned price for the bed net. The context implies that $Y(0) = 0$. Suppose that the instrument is binary and that there are no covariates. The researcher assumes that $E[Y(1) \vert U = u] = \theta_{0} + \theta_{1}u + \theta_{2}u^{2}$ is a weakly decreasing, quadratic function of $u$, which must be contained within $[0,1]$ because $Y(1)$ is binary. These shape restrictions are imposed through four linear constraints:

align[align omitted — 251 chars of source]

which imply $E[Y(1) \vert U = u] \leq 1$ for all $u$, because $E[Y(1) \vert U = 0] = \theta_{0} \leq 1$ and the derivative of $E[Y(1) \vert U = u]$ is $\theta_{1} + 2\theta_{2} \leq 0$.

It will be convenient to use an alternative parameterization. Set $\tilde{\theta}_{1} = -\theta_{1}$, $\delta = \theta_{0} + \theta_{1} + \theta_{2}$, so that $\theta_{0}, \tilde{\theta}_{1}, \delta \geq 0$. Then the constraints (ref) can be simplified to

align[align omitted — 187 chars of source]

where $s_1, s_2 \geq 0$ are slack variables. The researcher matches the moments $E_{P}[YD \vert Z = z]$ for $z = 0,1$, creating the two restrictions

align[align omitted — 407 chars of source]

The null hypothesis of interest is $H_{0}: E_{P}[Y(1)] = \tau_{0}$, where

align[align omitted — 282 chars of source]

We set $x_{1} = (\theta_{0}, \tilde{\theta}_{1}, \delta, s_{1}, s_{2})'$, so that the problem defined by (ref)--(ref) fits in the form (ref) with

align*[align* omitted — 518 chars of source]

and with both $x_{0}$ and $A_{0}(P)$ null.

The simulation design is specified as $P\{Z = 1\} = 0.5 = P\{Z = 0\}$, $p(0) = 1/3$, $p(1) = 2/3$, $Y(0) = 0$, $Y(1) = \mathbf{1}\{V \le \theta_{0} + \theta_{1} U + \theta_{2} U^2\}$ where $V|Z,U \sim U[0,1]$ and $(\theta_{0},\theta_{1},\theta_{2}) = (1, -1, 0.5)$. Given this design, the identified set for the ATE is $[0.58,0.67]$. Using a random sample from $(Y,D,Z)$, $\hat{b}_{j,n}$ is once again computed using sample analogs (recall that $A_0(P)$ does not exist given this parametrization).

Figure (ref) presents the rejection probabilities from $1,000$ draws for tests of the null hypothesis $H_0: E[Y(1)] = \tau_{0}$ at a $5\%$ significance level, for a grid of values of $\tau_0$. Both tests control size within the identified set for all sample sizes, however, the rejection probability remains below $5\%$ well outside of the identified set in small samples. At all sample sizes, we find that the direct method is more conservative outside of the identified set than the screening method in this example.

figure[figure omitted — 913 chars of source]

freyberger2015identification

Finally, we consider the model presented in freyberger2015identification, as described in Example (ref). The simulation design is specified as follows: $\mathcal{X} = \{2,3,4,5,6,7\}$ and $\mathcal{W} = \{0,1\}$, with $\pi_{h,k} = \Pr(X = x_h, W = w_k)$ given by \[ \Pi =

pmatrix[pmatrix omitted — 96 chars of source]

. \] The vector $g = (g(2), \ldots, g(7))'$ is given by $(23, 17, 13, 11, 9, 8)'$. The instrument is distributed as $Z_i \sim N(0,1)$, independently of $(X_i, W_i)$. Finally, the outcome equation is \[Y_i = g(X_i) + U_i~,\] where the error is given by \[U_i = X_iZ_i^2 - E_P[X_i|W_i]~.\] The shape constraint we maintain is that the structural function $g$ is decreasing, i.e., $Sg \leq 0$ for \[S =

pmatrix[pmatrix omitted — 162 chars of source]

, \] and our functional of interest is $L(g) = g(2)$. Given this design the identified set for $L(g)$ is $[20.21, 24.61]$. Using a random sample from $(Y, X, W, Z)$, $\hat{A}_{0,n}$ and $\hat{b}_{j,n}$ are computed using sample analogs, using the known value of $S$ and $c$.

Figure (ref) presents the rejection probabilities from $1,000$ draws for tests of the null hypothesis $H_0: L(g) = L_0$ at a $5\%$ significance level, for a grid of values of $L_0$. This design seems particularly challenging, with the rejection probabilities outside of the identified set being highly asymmetric. One possible source for this asymmetry is that the shape constraints are violated to the left of the identified set, but are not violated to the right. Beyond this observation, our findings are similar to the previous two examples; both tests control size within the identified set for all sample sizes, and we find that the direct method is more conservative outside of the identified set than the screening method.

figure[figure omitted — 932 chars of source]