EconBase
← Back to paper

Testing Inequalities Linear in Nuisance Parameters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

121,682 characters · 14 sections · 64 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing Inequalities Linear in Nuisance Parameters

abstractThis paper proposes a new test for inequalities that are linear in possibly partially identified nuisance parameters. This type of hypothesis arises in a broad set of problems, including subvector inference for linear unconditional moment (in)equality models, specification testing of such models, and inference for parameters bounded by linear programs. The new test uses a two-step test statistic and a chi-squared critical value with data-dependent degrees of freedom that can be calculated by an elementary formula. Its simple structure and tuning-parameter-free implementation make it attractive for practical use. We establish uniform asymptotic validity of the test, demonstrate its finite-sample size and power in simulations, and illustrate its use in an empirical application that analyzes women's labor supply in response to a welfare policy reform.

{\bf Keywords:} Linear Program, Moment Inequalities, Quadratic Programming, Subvector Inference, Uniform Inference

Introduction

We propose a simple new test for hypotheses of the form $H_0$: there exists a $\delta$ such that $C\delta\leq b$, where elements of the Jacobian matrix $C$ and the intercept vector $b$ are reduced-form parameters that can be consistently estimated, and elements of $\delta$ are unknown parameters whose values are partially identified by the inequalities under $H_0$. Since the inequalities, rather than $\delta$, are of central interest, $\delta$ is a nuisance parameter vector. Hypotheses of this form arise in specification testing and subvector inference for linear unconditional moment (in)equality models and in inference for parameters bounded by linear programs, including discrete instrumental variable (IV) models with shape restrictions and policy relevant treatment effect models. These models have wide applications in empirical work. We explain the applications and give examples in Section (ref).

Testing this hypothesis is non-standard both because the nuisance parameter $\delta$ may not be point-identified and because the hypothesis involves inequalities. As a result, commonly used test statistics have non-standard asymptotic distributions involving parameters that cannot be consistently estimated, in particular, the local slackness of the inequalities evaluated at the true value of $\delta$. This complicates the design of critical values. A common approach is to simulate the asymptotic distribution with a conservative estimator of the local slackness plugged in. However, the conservative estimators typically involve user-chosen tuning parameters that introduce arbitrariness to the procedure. Moreover, the simulated critical values can be computationally burdensome.

In the special case that the Jacobian is {\em known}, CoxShi2023, hereafter CS23, propose the subvector conditional chi-squared (sCC) test that does not require simulation or user-chosen tuning parameters and yet still has uniform asymptotic size control and good power. The simplicity of the test is achieved by considering the conditional distribution of the quasi-likelihood ratio (QLR) statistic given the identity of the active inequalities.\footnote{An inequality is active if it holds with equality at our null-imposed estimator of $\delta$.} The conditional distribution is shown to be bounded by a chi-squared distribution with degrees of freedom (DoF) dependent on the conditioning event. CS23 recommends computing the DoF by solving a sequence of linear programming problems.

Our first contribution is to derive an elementary formula for the DoF that replaces the recommendation in CS23. The formula makes the computation of the critical value elementary. Implementing the sCC test is now no harder than calculating the test statistic, which is a convex quadratic programming problem (CQPP). The formula also reveals an intuitive interpretation of the sCC test: The sCC test turns out to be the same as the classic Sargan-Hansen's J test for a moment equality model, where the equalities are determined by the active inequalities.

While the sCC test is attractive, the known Jacobian assumption significantly restricts its applicability. In both linear unconditional moment inequality models and models with a parameter of interest bounded by linear programs, a known Jacobian only applies in special cases. The known Jacobian assumption even rules out conditional moment inequality models where $\delta$ includes coefficients on endogenous covariates. The second contribution of this paper is to propose a new test, called the generalized conditional chi-squared (GCC) test, that accounts for the estimation error of an {\em unknown} Jacobian while maintaining the simplicity, size, and power properties of the sCC test.

Estimating the Jacobian presents two challenges. The first challenge is finding the limit of the constraint set in the definition of the QLR statistic. In general, convergence of the matrix of coefficients in a system of inequalities is insufficient for the set defined by those inequalities to converge in a setwise sense. We show that a simple stable rank condition on the Jacobian is sufficient for convergence of the constraint set. The stable rank condition requires the rank of certain submatrices of the Jacobian to not change in the limit. The submatrices that need to satisfy the condition are minimal in some sense. They are associated with collections of inequalities that (implicitly or explicitly) define equality restrictions in the limit. While the stable rank condition is not innocuous, it relaxes the commonly used strong identification assumption in moment equality models, which requires the Jacobian to be full rank.

The second challenge is finding a consistent variance estimator. The effect of the estimated Jacobian on the variance depends on $\delta$, but $\delta$ is only partially identified. This makes the additional variance term difficult to account for. Our solution is to use a first-stage estimator of $\delta$ that converges to a point in the identified set for $\delta$. The procedure resembles optimal weighting in two-step generalized method of moments (GMM). The GCC test compares this two-step statistic to a chi-squared critical value with DoF determined by the formula from our first contribution. We also define a refinement of the GCC test in the supplemental appendix that has slightly more power.

We next review three strands of related literature. The first consists of papers that propose tests of moment inequalities that are possibly nonlinear in nuisance parameters. These include BugniCanayShi2015, BugniCanayShi2017, CCT2018, BelloniBugniChernozhukov2018, KMS2019, and Bei2023. These methods use critical values that are nontrivial to compute and require user-chosen tuning parameters to be adaptive to the unknown slackness of the inequalities.\footnote{An exception is Procedure 3 in CCT2018 in that it does not require any tuning parameter or simulation. We include this procedure in the simulations.} In contrast, the GCC test uses an algebraic critical value and is adaptive to the unknown slackness of the inequalities without any tuning parameter, at the price of requiring linearity.

The second strand of related literature is composed of CS23 and AndrewsRothPakes2023, which consider conditional moment inequalities that are linear in the nuisance parameters. These papers rely on an idea that if the Jacobian depends only on random variables on which the inequalities hold conditionally, then the Jacobian can be treated as known. However, this idea does not apply if the moment inequalities are unconditional or, more generally, if the Jacobian depends on a random variable for which the moment inequalities do not hold conditionally. Our method applies regardless of whether the Jacobian can be treated as known.

When our test is applied to inference for a scalar parameter (say $\theta$) bounded by linear programs, it is related to the third strand of literature. This literature addresses the problem of inference for the value of a linear programming problem (LPP). The papers include FreybergerHorowitz2015, BaiSantosShaikh2022, FSST2023, ChoRussell2024, Gafarov2025, Voronin2025, and GoffMbakop2025, among others. There is a subtle technical difference between our setting and this literature: while our test is inverted to yield a confidence interval for $\theta$, the LPP literature aims at constructing confidence intervals for the upper or lower bound for $\theta$ defined by the value of a LPP. While the two problems are distinct, they are closely related. Our confidence interval by design can cover either bound with correct nominal coverage probability (asymptotically), and in the LPP literature, one-sided confidence intervals of the appropriate direction for either bound are also valid confidence intervals for $\theta$.\footnote{In the LPP literature, two-sided confidence intervals for $\theta$ can be obtained by combining two one-sided confidence intervals via a Bonferroni adjustment.} Thus, the methods can be used for the same empirical problems. Notably, our method is the only one that is tuning parameter and simulation free.

We include a simple one-sided specification in the simulations in order to compare the GCC test to representative papers in all three strands of the literature. In addition to the simple simulation, we also evaluate the GCC test in two realistic simulation examples: one of an interval outcome instrumental variables (IV) model as in GandhiLuShi2023, hereafter GLS23, and the other of bounds on policy relevant treatment effects as in MogstadSantosTorgovitsky2018, hereafter MST18. The simulations show that the GCC test is computationally very fast with good size and power.

We also implement the GCC test in an empirical analysis of female labor supply in response to a welfare policy reform. KlineTartari2016, hereafter KT16, estimate bounds on the treatment responses by manually eliminating the nuisance parameters from revealed preference inequalities. The GCC test provides uniformly valid inference for the treatment responses based directly on the revealed preference inequalities. Overall, we find statistically significant heterogeneous responses to the policy change, which agree with the results in KT16.

The rest of the paper is organized as follows. Section (ref) describes the setup and applications. Section (ref) defines the GCC test. Section (ref) describes the theoretical properties of the GCC test. Section (ref) presents the simulations. Section (ref) presents the empirical illustration. Section (ref) concludes. An appendix contains proofs of the theorems, while a supplemental appendix includes additional results, proofs, simulations, and discussion.

Setup and Examples

We are interested in testing the hypothesis

equation[equation omitted — 110 chars of source]

where $C$ is a $d_C\times d_\delta$ matrix of reduced-form parameters, $b$ is a $d_C$-dimensional vector of reduced-form parameters, $d_\delta$ is the dimension of $\delta$, and $d_C$ is the number of inequalities. To make an invertibility assumption imposed later as unrestrictive as possible, we add some structure to $C$ and $b$. We assume that

equation[equation omitted — 68 chars of source]

where $B$ is a known $d_C\times d_\mu$ matrix, $D$ is a known $d_C\times d_\delta$ matrix, $d$ is a known $d_C$-dimensional vector, $\Pi$ is an unknown $d_\mu\times d_\delta$ matrix of reduced-form parameters, and $\mu$ is an unknown $d_\mu$-dimensional vector of reduced-form parameters. This structure separates the unknown and estimated components from the known components in the Jacobian $C$ and the intercept $b$. It is satisfied in all the examples considered below. Typically, $B$ has more rows than columns, and it absorbs the linear dependence across rows for the estimation noise of the inequalities. This allows us to accommodate inequalities with linearly dependent estimation errors, which arise when we write an equality as a pair of opposing inequalities, when the model contains a deterministic constraint such as a shape or sign restriction, or when the law of total probability dictates that a weighted sum of the inequalities involves no unknown quantities under $H_0$.

With the structure in ((ref)), $H_0$ can be equivalently written as:

align[align omitted — 116 chars of source]

Let $\overline{\mu}_n$ and $\overline{\Pi}_n$ be estimators of $\mu$ and $\Pi$. Let $\overline{C}_n=D+B\overline{\Pi}_n$ and $\bar b_n=d-B\overline{\mu}_n$. In the next two subsections, we describe two classes of models that are covered by this framework.

Moment (In)equality Models

Moment (in)equality models are used to address data or modeling incompleteness issues, including missing data, multiple equilibria, and large intractable games.\footnote{For empirical applications and current statistical methods for such models, see the survey papers by CanayShaikh2017, HoRosen2017, and Molinari2020.} Here, we show that both specification testing and subvector inference for moment inequality models fit the hypothesis in equation ((ref)) when the moments are linear in the parameters.

Consider the moment (in)equality model:

equation[equation omitted — 208 chars of source]

where $\overline{m}_n(\beta) = (\overline{m}^{eq}_n(\beta)',\overline{m}_n^{ineq}(\beta)')'$ is a $\mathbb{R}^{d_m}$-valued sample moment function, $\beta$ is an unknown vector of parameters, and ${\cal B}$ is its parameter space. Specifically, let $\overline{m}_n(\beta) = n^{-1}\sum_{i=1}^n m(W_i,\beta)$, where $\{W_i\}_{i=1}^n$ is a sample of observable variables and $m(\cdot, \beta)$ is a function known up to the unknown parameter $\beta$. Let the number of equalities be denoted $d_{eq}$ and the number of inequalities be denoted $d_{ineq}$, so that $d_m = d_{eq}+d_{ineq}$.

Suppose $\overline{m}_n(\beta)$ is linear in $\beta$. That is, $\overline{m}_n(\beta) = \overline{\Gamma}_n\beta + \overline{\eta}_n$ for $\overline{\Gamma}_n = \partial\overline{m}_n(\beta)/\partial\beta'$ and $\overline{\eta}_n = \overline{m}_n(\mathbf{0})$. Suppose ${\cal B} = \mathbb{R}^{d_\beta}$.\footnote{More generally, if ${\cal B}$ is a polyhedral set, then the deterministic inequalities that define ${\cal B}$ should be included when writing the hypothesis in the form of ((ref)). We show how deterministic constraints can be incorporated in the next subsection.} Consider the following types of problems:

enumerate• Specification Testing. When specification testing, one evaluates whether there exists a $\beta\in\mathbb{R}^{d_\beta}$ such that the moment (in)equalities hold. If not, then the model is misspecified. The hypothesis is \[ H_0: \mathbb{E}[\overline{m}_n^{eq}(\beta)] = \mathbf{0}\text{ and }\mathbb{E}[\overline{m}_n^{ineq}(\beta)]\geq \mathbf{0} \text{ for some }\beta\in \mathbb{R}^{d_\beta}. \] This is the type of hypothesis considered in BugniCanayShi2015. It can be written in the form of ((ref)) with \[ B = \left(\begin{smallmatrix}-I_{d_{eq}}&\mathbb{O}_{d_{eq}\times d_{ineq}}\\I_{d_{eq}}&\mathbb{O}_{d_{eq}\times d_{ineq}}\\\mathbb{O}_{d_{ineq}\times d_{eq}}&-I_{d_{ineq}}\end{smallmatrix}\right),~\mu = \mathbb{E}[\overline{\eta}_n], ~\Pi = \mathbb{E}[\overline{\Gamma}_n], ~\delta = \beta, ~D=\mathbb{O}_{(d_{eq}+d_m)\times d_\beta}, ~d = \mathbf{0}, \] where $I_a$ is an identity matrix of size $a$ and $\mathbb{O}_{a\times b}$ is a $a\times b$ zero matrix. Note that we write the equalities as pairs of opposing inequalities via the first $2d_{eq}$ rows of $B$. • Subvector Inference. In subvector inference, one constructs a confidence set for a subvector of the parameters. Suppose the subvector of interest is composed of the first $\ell$ elements of $\beta$ and denote it by $\theta$. Let $\delta$ denote the rest of the elements in $\beta$. Then a confidence set for $\theta$ can be constructed by testing the following hypothesis at each value of $\theta$ and collecting the values of $\theta$ at which the hypothesis is not rejected: \[ H_0: \mathbb{E}[\overline{m}^{eq}(\theta,\delta)] = \mathbf{0}\text{ and }\mathbb{E}[\overline{m}^{ineq}(\theta,\delta)]\geq \mathbf{0} \text{ for some }\delta \in \mathbb{R}^{d_\delta}. \] This is the type of hypotheses considered in BugniCanayShi2017, and it can be written in the form of ((ref)) with \[ B = \left(\begin{smallmatrix}-I_{d_{eq}}&\mathbb{O}_{d_{eq}\times d_{ineq}}\\I_{d_{eq}}&\mathbb{O}_{d_{eq}\times d_{ineq}}\\\mathbb{O}_{d_{ineq}\times d_{eq}}&-I_{d_{ineq}}\end{smallmatrix}\right),~\mu = \mathbb{E}[\overline{m}_n(\theta,\mathbf{0})], ~\Pi = \mathbb{E}[\overline{\Gamma}_n^{\delta}], ~D=\mathbb{O}_{(d_{eq}+d_m)\times (d_\beta-\ell)}, ~d = \mathbf{0}, \] where $\overline{\Gamma}_n^{\delta} = \partial\overline{m}_n(\theta,\delta)/\partial\delta'$, which is the last $d_\beta -\ell$ columns of $\overline{\Gamma}_n$.

Two remarks are in order regarding subvector inference:

{\bf Remarks:} (1) {\it Linearity in $\theta$ is not needed for subvector inference. The discussion remains unchanged if $\overline{m}_n(\theta,\delta)$ is linear in $\delta$ and nonlinear in $\theta$. }

(2) {\it Sometimes, the parameter of interest is not a subvector of $\beta$, but instead a linear function of $\beta$. That is, $\theta = \Lambda\beta$ for a known $d_\theta\times d_\beta$ full-rank matrix $\Lambda$ with $d_\theta<d_\beta$. One approach is to reparameterize $\beta$ so that $\theta$ becomes a subvector of the new parameter. Let $\Lambda^c$ be a $(d_\beta-d_\theta) \times d_\beta$ row-augmenting matrix so that $\left(

smallmatrix\Lambda\\\Lambda^c

\right)$ is nonsingular.\footnote{In Matlab, one can find such a matrix by applying the function \texttt{null}(~) on $\Lambda$.} The reparameterization is given by $\gamma = \left(

smallmatrix\Lambda\\\Lambda^c

\right)\beta$. Then, plugging $\beta = \left(

smallmatrix\Lambda\\\Lambda^c

\right)^{-1}\gamma$ into \textup{(\ref{mi})} reparameterizes the model so that $\theta$ is a subvector of $\gamma$. Equivalently, one can add $\theta= \Lambda\beta$ to the model as deterministic constraints and treat $\beta$ as the nuisance parameter. }

We end this subsection with two examples of linear moment inequality models.

example[Interval Outcome IV Regression] Consider a linear model $Y^\ast = X'\beta+\varepsilon$ with $\mathbb{E}[\varepsilon Z]=\mathbf{0}$, where $Z$ is a vector of instruments. The dependent variable $Y^\ast$ is not observed. Instead, we observe $Y^L$ and $Y^U$ that satisfy: $\mathbb{E}[Y^LZ]\leq \mathbb{E}[Y^\ast Z]\leq \mathbb{E}[Y^U Z]$. Then, we have the following unconditional moment inequalities: \begin{equation} \mathbb{E}[Y^L Z -ZX'\beta]\leq \mathbf{0} and \mathbb{E}[ZX'\beta-Y^UZ]\leq \mathbf{0}. \end{equation} This is an example of ((ref)). The interval outcome IV regression model was proposed in ManskiTamer2002. A generalization of such a model to a non-standard aggregate demand estimation problem is studied in GLS23. In their generalization, $Y^L$ and $Y^U$ can be nonlinear functions of the parameter of interest. {CS23 and AndrewsRothPakes2023 cover a related model where the inequalities in \textup{((ref))} hold conditionally on $Z$. Their tests apply to hypotheses that fix the coefficients on all the endogenous regressors. Then, the Jacobian of the inequalities with respect to the nuisance parameter is known after conditioning on $Z$ since it is not a function of the endogenous regressors. Their tests do not apply when a nuisance parameter is a coefficient on an endogenous regressor.}
example[Panel Data Multinomial Choice Model] Consider a panel data multinomial choice model where individual $i$ at time $t$ obtains utility $u_{ijt}$ from choosing option $j$. Let $y_{ijt}=1$ if $i$ chooses $j$ at time $t$, and $y_{ijt}=0$ otherwise. The random utility model stipulates that $y_{ijt}=1$ if and only if $u_{ijt}\geq u_{ij't}\text{ for all }j'\in \{0,1,2,\dots,J\}$. Consider the linear index model of the random utility: $u_{ijt} = X_{ijt}'\gamma+\lambda_{ij}+\varepsilon_{ijt}$, where $X_{ijt}$ is a vector of observed covariates, $\lambda_{ij}$ is an unobserved fixed effect, and $\varepsilon_{ijt}$ is an idiosyncratic taste shock. Normalize $X_{i0t}=\mathbf{0}$. For illustration, let there be only two time periods ($t=1,2$) and let the individuals be independent and identically distributed. Under a conditional time homogeneity assumption on $\varepsilon_{ijt}$, ShiShumSong2018 show that the following moment inequality holds: \begin{equation} \mathbb{E}[\Delta y_i' \Delta X_{i}\gamma | X_{i}]\geq 0, \end{equation} where $\Delta y_i$ is a $J$-dimensional vector with its $j$th element being $y_{ij2}-y_{ij1}$, $\Delta X_{i}$ is a $J\times d_x$ dimensional matrix with its $j$th row being $(X_{ij2}-X_{ij1})'$, and $X_{i}$ collects $X_{ijt}$ for all $j\in \{1,2,\dots,J\}$ and $t\in\{1,2\}$. Note that none of the elements of $\Delta y_i' \Delta X_{i}$ can be considered exogenous because they depend on $y_{ijt}$. Thus, the inequalities do not fit into the conditional moment inequality setup in CS23 or AndrewsRothPakes2023. Let ${\cal I}(X_i)$ be a non-negative vector-valued instrumental function. Then, \begin{equation} \mathbb{E}[{\cal I}(X_i)\Delta y_i' \Delta X_{i}\gamma]\geq \mathbf{0} \end{equation} rewrites the inequalities in ((ref)) into the form of ((ref)).\footnote{In this model, a normalization is usually imposed on $\gamma$, such as the first element being one, that can be accommodated by simply setting that element to 1.}

Parameters Bounded by Linear Programming

Recently, an important class of models have arisen in the structural estimation literature where a scalar parameter of interest is not point-identified but bounded by the values of linear programming problems (LPPs).\footnote{Some notable examples appear in KT16, MST18, KKLS2021, and SyrgkanisTamerZiani2021.} The constraint sets for the LPPs are defined by linear (in)equalities, where the constants and coefficients in the (in)equalities are unknown parameters to be estimated. To fix ideas, suppose the parameter of interest is

equation[equation omitted — 54 chars of source]

where $\gamma\in\mathbb{R}^{d_\delta}$ is a known vector that defines a linear combination of the nuisance parameters $\delta$. The constraint set is defined by a collection of (in)equalities:

equation[equation omitted — 77 chars of source]

where the elements of $m\in \mathbb{R}^{d_{\Gamma}}$ and $\Gamma\in\mathbb{R}^{d_\Gamma\times d_\delta}$ are known or can be consistently estimated, $A$ is a known $d_A\times d_\delta$ matrix, and $b\in\mathbb{R}^{d_A}$ is a known vector.\footnote{We focus on the case $A$ and $b$ are known because it is common in applications. Unknown and estimated $A$ and/or $b$ can also be covered.} The known inequalities in ((ref)) define a parameter space for $\delta$. For example, if $\delta$ is a vector of weights, then $\delta$ takes values in the simplex, which can be represented by an appropriate choice of $A$ and $b$.

Inference on $\theta$ can be based on the GCC test for the existence of a value of $\delta$ that simultaneously satisfies ((ref)) and ((ref)) at a hypothesized value of $\theta$. A confidence interval for $\theta$ can be calculated by inverting a family of tests. The restrictions in ((ref)) and ((ref)) can be written in the form of ((ref)) with

align[align omitted — 570 chars of source]

Note that the first two rows and the last $d_A$ rows of $B$ are zeros to accommodate the deterministic constraints $\theta = \gamma\delta$ and $A\delta\leq b$.

Another approach to inference on $\theta$ is to use LPPs. From the point of view of identification, $\theta$ is bounded sharply by $\theta_\min = \min_{\delta:~\text{(\ref{deltaID}) holds}}\gamma'\delta$ and $\theta_\max = \max_{\delta:~\text{(\ref{deltaID}) holds}}\gamma'\delta$. Then one can construct one-sided confidence intervals that are bounded from below (above) for $\theta_{\min}$ ($\theta_{\max}$) and use them as confidence intervals for $\theta$. Inference for the value of a LPP is generally based on plugging the estimators of $m$ and $\Gamma$ into the LPPs and simulating or bootstrapping the asymptotic distributions of these estimators. However, the asymptotic distributions depend on which corner or face of the constraint set solves the LPP, which is not smooth as a function of the estimated reduced-form parameters. Thus, the na\"{i}ve strategy of bootstrapping the value of a LPP is generally invalid. In order to obtain valid inference, the LPP literature recommends various modifications that involve tuning parameters and/or simulating/bootstrapping nonstandard distributions. Using the GCC test to obtain a confidence interval for $\theta$ avoids these complications.

{\bf Remarks:} (1){\it The GCC test is valid for any hypothesized value of $\theta$ in $[\theta_\min, \theta_\max]$, including the endpoints. Thus, the GCC test is a valid way to do inference for the value of a LPP. However, the GCC test may direct power in a one-sided way. In particular, when $\theta_\max-\theta_\min$ is not small, the GCC test is effectively one-sided for the value of the LPP (though still two-sided for $\theta$). Therefore, when the parameter of interest is $\theta_\min$ or $\theta_\max$ instead of $\theta$ and the researcher desires two-sided inference, then some of the inference recommendations from the LPP literature are preferred. }

(2) {\it In some applications, $\gamma$ is unknown and estimated. There is a convenient way to write ((ref)) and ((ref)) in the form of ((ref)). The idea is to add an element to $\mu$ that is always zero, while adding $\gamma'$ as a row of $\Pi$. Specifically, ((ref)) can be satisfied when $\gamma$ is unknown by letting

equation[equation omitted — 904 chars of source]

More generally, if one of the equations in ((ref)) has a known intercept but the corresponding row of the Jacobian is unknown, then the intercept should be included in $d$ while the row of the Jacobian should be included in $\Pi$, possibly including a zero in $\mu$ and augmenting the columns of $B$.\footnote{This idea can be applied more generally. Consider a generalization of (2), where $C$ can be written as $C=B_1\Pi_1+B_2\Pi_2+D$ and $b=d-B_1\mu_1$. That is, $B_2\Pi_2$ is a component of $C$ that needs to be estimated but cannot be written as a linear combination of the columns of $B_1$. Then, we can satisfy the structure in (2) by taking $B=[B_1,B_2]$ and $\mu=(\mu'_1,\mathbf{0}')'$. This works because we do not require the estimator of $\mu$ to have a nonsingular variance matrix, but we only require the estimator of $\mu+\Pi\delta$ to have a nonsingular variance matrix for a value of $\delta$ in the identified set; see Assumption (ref)(iii), below.} This demonstrates the flexibility of the specification in ((ref)). }

We demonstrate the relevance of the setup in ((ref)) and ((ref)) with examples.

example[Discrete IV Regression with Shape Restrictions] IV regressions with discrete regressors and instruments are common in practice. Prominent examples include PermuttHebel1989, AngristKrueger1991, and Angrist1998. FreybergerHorowitz2015 consider an IV model with discrete $X_i$ and $Z_i$: \begin{equation} Y_i = \delta(X_i)+\varepsilon_i, \mathbb{E}[\varepsilon_i|Z_i]=0, \end{equation} where $Y_i$ is the dependent variable, $X_i$ a discrete endogenous regressor, $Z_i$ is a discrete instrument, $\delta(\cdot)$ an unknown function that represents the structural relationship between $X_i$ and $Y_i$, and $\varepsilon_i$ is the error term. While linearity of $\delta(\cdot)$ is often assumed, FreybergerHorowitz2015 emphasize that no such functional form restrictions are needed. Let the support of $X_i$ and $Z_i$ be $\{x_1,\dots,x_{d_x}\}$ and $\{z_1,\dots,z_{d_z}\}$, respectively. Let $\delta_k = \delta(x_k)$ for $k\in\{1,\dots,d_x\}$, and $\delta= (\delta_1,\dots,\delta_{d_x})'$. The model in ((ref)) implies that $\delta$ satisfies the equalities in ((ref)) with \begin{equation} m = \left(\begin{smallmatrix}\mathbb{E}[Y_i1\{Z_i=z_1\}]\\ \mathbb{E}[Y_i1\{Z_i=z_2\}]\\ \vdots\\ \mathbb{E}[Y_i1\{Z_i=z_{c_z}\}]\end{smallmatrix}\right) and \Gamma =\left(\begin{smallmatrix}\mathbb{P}(X_i =x_1,Z_i=z_1)&\mathbb{P}(X_i =x_2,Z_i=z_1)&\cdots&\mathbb{P}(X_i =x_{d_x},Z_i=z_1)\\ \mathbb{P}(X_i =x_1,Z_i=z_2)&\mathbb{P}(X_i =x_2,Z_i=z_2)&\cdots&\mathbb{P}(X_i =x_{d_x},Z_i=z_2)\\ \vdots&\vdots&\ddots&\vdots\\ \mathbb{P}(X_i =x_1,Z_i=z_{d_z})&\mathbb{P}(X_i =x_2,Z_i=z_{d_z})&\cdots&\mathbb{P}(X_i =x_{d_x},Z_i=z_{d_z}) \end{smallmatrix}\right). \end{equation} Note that $m$ and $\Gamma$ are reduced-form parameters that can be estimated by sample averages. When $d_z<d_x$, $\delta$ is not point identified by the equalities in ((ref)). To sharpen identification, FreybergerHorowitz2015 add shape restrictions of the form $A\delta\le b$ for some known matrix $A$ and known vector $b$. This covers several types of shape restrictions including monotonicity and/or convexity of $\delta(\cdot)$. The parameter of interest is typically a linear function of $\delta$. For example, $\theta = [-1,1,0,...,0]\delta=\delta_2-\delta_1$ is the effect of changing $X$ from $x_1$ to $x_2$. Thus, inference for the structural function in ((ref)) falls into the framework of ((ref))-((ref)).
example[Policy Relevant Treatment Effects (PRTE)] Treatment effects that are relevant for policy are often not equal to the local average treatment effects (LATEs) associated with any available instrument. In a standard program evaluation model, they are weighted averages of underlying marginal treatment responses (MTRs), where the weights are identified or known, but the MTRs are not. MST18 show that the MTRs can be partially identified from the LATEs, or more generally from IV-like estimands, because these estimands are weighted averages of the MTRs. Then, bounds on a PRTE can be deduced from the identified set of the MTRs. Under a parameterization of the MTRs, MST18 show that the bounds are values of LPPs.\footnote{In some cases, MST18 show that even the nonparametric bounds can be written as finite-dimensional LPPs; see their Proposition 4.} To be specific, consider an outcome $Y$, a binary treatment indicator $D$, and covariates $Z=(X,Z_0)$, where $X$ is a vector of control variables and $Z_0$ is the vector of excluded instruments. Let $Y_0$ and $Y_1$ denote the potential outcomes corresponding to the two treatment arms. Suppose treatment is determined by a weakly separable selection equation: $D=\mathds{1}\{p(Z)\ge U\}$ for some unobserved uniformly distributed variable $U$, where $p(Z)$ is the propensity score. The MTR functions are defined to be \begin{equation} \kappa_0(u,x)=\mathbb{E}[Y_0|U=u, X=x] and \kappa_1(u,x)=\mathbb{E}[Y_1|U=u, X=x]. \end{equation} A wide range of PRTEs can be written as weighted averages of the MTRs. The MTRs can be parameterized as a linear combination of functions belonging to some basis. Let $\delta$ be a vector of coefficients on the basis functions. Then, a PRTE, say $\theta$, can be written as $\theta=\gamma'\delta$, where $\gamma$ is a vector of weighted averages applied to each basis function. This writes the parameter of interest in the form of ((ref)). An IV-like estimand is a parameter of the form $m_s=\mathbb{E}[s(D,Z)Y]$ for some identified or known function $s(D,Z)$. MST18 show that every IV-like estimand can be written as a weighted average of the MTRs with a simple formula for the weights. This means that each $m_s$ can be written as a linear combination of $\delta$. If $m$ denotes a vector of finitely many IV-like estimands, then $m$ satisfies the equalities in ((ref)) with each row of $\Gamma$ being the weighted average of the basis functions with weights corresponding to the IV-like estimand. In addition, MST18 allow the researcher to specify additional shape restrictions on the MTEs, in a similar manner as Example (ref). Depending on the choice of basis functions, these can sometimes be written as deterministic linear inequalities on $\delta$. Overall, this shows that $\theta$ is bounded by LPPs and satisfies the structure in ((ref)) and ((ref)). We demonstrate the GCC test in a simulation of this example in Appendix (ref).
example[State Transition Probabilities] Consider a model where individuals choose between finitely many states, $s\in\mathcal{S}$. A policy change may induce individuals to choose a different state. The state transition probabilities are defined by \begin{equation} \delta_{s, s'}=\Pr(S_a=s'|S_b=s) for s, s'\in\mathcal{S}, \end{equation} where $S_a$ denotes an individual's choice after a policy change and $S_b$ denotes an individual's choice before a policy change. These transition probabilities represent the fraction of individuals who start in state $s$ and change to state $s'$ in response to the policy change. The transition probabilities are not identified from the data if all one has is a repeated cross-section of individuals before and after the treatment, or a cross-section of individuals randomly assigned to different policy regimes. On the other hand, the marginal probabilities of $S_a$ and $S_b$, denoted $p_a$ and $p_b$, respectively, are identified. Then, the transition probabilities are partially identified through their relationship with $p_a$ and $p_b$: \begin{equation} p_a=\Delta p_b, \end{equation} where $\Delta$ is the matrix of $\delta_{s,s'}$ values for $s,s'\in\mathcal{S}$. These equations fit the structure of the equalities in ((ref)), with elements of $\Delta$ forming the nuisance parameter vector $\delta$.\footnote{One of the equations in ((ref)) is redundant because the sum of the elements in $p_a$ is one. This is not a problem.} In addition, the transition probabilities satisfy: \begin{equation} \delta_{s,s'}\ge 0 for all s, s'\in\mathcal{S} and \sum_{s'\in\mathcal{S}}\delta_{s,s'}=1 for all s\in\mathcal{S}. \end{equation} These restrictions can be written as the inequalities in ((ref)) with known $A$ and $b$. Then, inference for a particular transition probability fits \textup{((ref))} and \textup{((ref))}. KT16 use such a model to study women's labor supply in response to a welfare policy change. In their model, the transition probabilities represent the fraction of women who gain/lose employment and/or register for welfare. KT16 analyze the data from an experiment that randomized exposure to a new welfare policy. The policy change could have heterogeneous effects on labor supply through the extensive margin (encouraging some women to gain employment) and the intensive margin (encouraging some women to decrease their hours or wages in order to qualify for welfare). The experiment identifies the distribution of employment, welfare participation, and income for women under two welfare policies. KT16 point out that the details of the policy change, combined with weak assumptions on the utility functions of the individuals, restrict many of the transition probabilities to zero. They then manually solve for bounds on each of the remaining transition probabilities from \textup{((ref))} and \textup{((ref))} and find significant labor market effects along both the extensive and intensive margins. In Section (ref), we employ the GCC test to construct confidence intervals for this example, avoiding the need to solve for the bounds manually.

The Generalized Conditional Chi-Squared Test

In this section, we define the generalized conditional chi-squared (GCC) test. The test depends on $\overline{\mu}_n$ and $\overline{\Pi}_n$, consistent and asymptotically normal estimators of $\mu$ and $\Pi$. The test also depends on $\overline{\Omega}_n$, a consistent estimator of the asymptotic variance of $(\overline{\mu}'_n,\textup{vec}(\overline{\Pi}_n)')'$, where $\textup{vec}(A)$ denotes the vectorization of a matrix, $A$.

We start with preliminary $H_0$-restricted estimators for $\mu$ and $\delta$:

equation[equation omitted — 208 chars of source]

where $\widehat\Upsilon_n$ is a preliminary weight matrix that is converging in probability to a deterministic positive definite limit.\footnote{This is similar to the weight matrix used in the first step of two-step GMM; the limit theory is invariant to the choice of weight matrix used. We take $\widehat{\Upsilon}_n$ to be the identity matrix.} While $\widetilde\mu_n$ is always unique, $\widetilde\delta_n$ need not be. In that case, we can take $\widetilde\delta_n$ to be the minimizer that has the smallest norm:

equation[equation omitted — 147 chars of source]

Equation ((ref)) is a tie-breaking procedure that is only used if the value of $\widetilde\delta_n$ that minimizes ((ref)) is not unique.\footnote{Equation ((ref)) uses the Euclidean norm, although it could be replaced with any other norm.} Later, we add assumptions so that $\widetilde\delta_n$ is consistent for its population analogue, $\delta^\ast_F$.\footnote{We formally define $\delta^\ast_F$ in ((ref)), below. For now, it is enough to think of $\delta^\ast_F$ as the probability limit of $\widetilde\delta_n$.}

We next estimate the asymptotic variance of $\overline{\mu}_n+\overline{\Pi}_n\delta^\ast_F$ using

equation[equation omitted — 237 chars of source]

where $\otimes$ denotes the Kronecker product. Define the QLR test statistic as

equation[equation omitted — 155 chars of source]

Let $(\widehat\mu_n,\widehat\delta_n)$ solve the minimization problem in ((ref)). Similar to the initial estimators, $\widehat\mu_n$ is always unique and $\widehat\delta_n$ may not be unique. In that case, we can take $\widehat\delta_n$ to be any minimizer.\footnote{Note the subtlety in the definitions of $\widetilde\delta_n$ and $\widehat\delta_n$: $\widetilde\delta_n$ is required to minimize ((ref)) because it has to be consistent for $\delta^\ast_F$, while $\widehat\delta_n$ can be arbitrary because its consistency is not essential.}

For any $K\subseteq\{1,...,d_C\}$, let $|K|$ denote the cardinality of $K$, and let $I_K$ denote the submatrix of the $d_C\times d_C$ identity matrix formed by taking the rows corresponding to the indices in $K$. In this way, conformable premultiplication of $I_K$ to a matrix $B$ selects the rows of $B$ corresponding to the indices in $K$ to form the submatrix of $B$ with $|K|$ rows.

We are ready to define the DoF and the critical value. Let $\widehat b_n=d-B\widehat\mu_n$ and $\widehat{K}= \{j\in\{1,\dots,d_C\}: e'_j\overline{C}_n\widehat\delta_n= e'_j\widehat b_n\}$, where $e_j$ is the $j$th standard normal basis vector. Note that $\widehat{K}$ denotes the set of indices at which the inequality constraint holds as equality for the minimizers in ((ref)). Let

equation[equation omitted — 149 chars of source]

For a significance level $\alpha$, let $cv(s,\alpha)$ denote the $1-\alpha$ quantile of the $\chi^2$ distribution with DoF equal to $s$. The GCC test rejects if $T_n$ is greater than $cv(\widehat s_n,\alpha)$.

We end this section with some remarks on the definition of the GCC test.

\noindentRemarks: {\it (1) Equation ((ref)) gives an algebraic formula for the DoF. In CS23, the DoF, $\widehat{r}_n$, is defined as the dimension of the span of a polyhedral cone. Theorem (ref), below, shows that $\widehat{r}_n$ is equal to $\widehat s_n$ with probability one in the limit. For computational reasons, we recommend $\widehat s_n$ over $\widehat r_n$.

(2) The GCC test is in some sense a “na\"ive” test. It is equivalent to first selecting the inequalities that are active according to the finite sample CQPP in} ((ref)){\it, pretending that these inequalities are equalities, and then forming a Sargan-Hansen's $J$-test for overidentification of a model defined by these equalities. To see this, note that the typical case is when $\textup{rk}(I_{\widehat K}[B,D])=|\widehat K|$ and $\textup{rk}(I_{\widehat K}C)=d_\delta$. Then, the DoF used by the GCC test is $|\widehat K|-d_\delta$, which represents the number of active inequalities minus the number of nuisance parameters.

(3) The weight matrix, $\widetilde{\Sigma}_n$, used in the definition of the GCC test is an estimator of the asymptotic variance of $\sqrt{n}(\overline{\mu}_n+ \overline\Pi_{n}\delta^\ast_F)$ instead of that of $\sqrt{n}\overline{\mu}_n $. The latter is used in CS23 for the sCC test. The new weight matrix accounts for the estimation error in $\overline{\Pi}_n$. To gain some intuition for this weight matrix, note that $T_n$ can be rewritten as

align[align omitted — 136 chars of source]

where $\gamma = \sqrt{n}(\delta-\delta^\ast_F)$, $\eta = \sqrt{n}(\mu-\mu_F+(\overline{\Pi}_n-\Pi_F)\delta_F^\ast)$, $X_n = \sqrt{n}(\overline{\mu}_n - \mu_F +(\overline{\Pi}_n-\Pi_F)\delta_F^\ast)$, and $h_n = \sqrt{n}(d-C_F\delta_F^\ast-B\mu_F)$, where $\mu_F$, $\Pi_F$, and $C_F$ stand for the true values of $\mu$, $\Pi$, and $C$. With this change of variables, one can see that $\widetilde{\Sigma}_n$ estimates the asymptotic variance of $X_n$.

(4) The GCC test is very easy to compute since it only requires solving two CQPPs. Efficient interior-point algorithms for CQPPs are available in most commonly used software, and they are known to have a worst-case computational complexity of $O((d_C+d_\delta)^4)$, where $d_C$ is the number of inequalities and $d_\delta$ is the dimension of the nuisance parameter. This is only slightly slower than the computational complexity of LPPs, which is $O((d_C+d_\delta)^{3.5})$; see Karmarkar1984 and YeTse1989. Most importantly, no simulation or bootstrap is needed to perform the test.}

Theoretical Properties

In this section, we present the three main theoretical results of the paper: (1) a theorem that justifies using $\widehat s_n$ and hence simplifies the rank calculation, (2) a theorem that shows the consistency of $\widetilde\delta_n$ for $\delta^\ast_F$, and (3) a theorem that shows the uniform asymptotic validity of the GCC test.

Rank Calculation Theorem

We now present the result that justifies using $\widehat s_n$. One can also define the DoF in the GCC test using the Karush-Kuhn-Tucker (KKT) multipliers. Let $\widehat{L}=\{j\in\{1,...,d_C\}: e'_j\widehat\psi_n>0\}$, where $\widehat\psi_n$ is a vector of nonnegative multipliers that satisfy the KKT conditions for ((ref)). Then

equation[equation omitted — 150 chars of source]

is another way to define the DoF. Also note that

equation[equation omitted — 136 chars of source]

is the definition of the DoF for the sCC test in CS23, where $\textup{dim}(\cdot)$ denotes the dimension of a set, or the maximum number of linearly independent elements. The following theorem shows that $\widehat s_n$, $\widehat t_n$, and $\widehat r_n$ are equal with probability one in the limit. This theorem plays a vital role in the proof of the uniform asymptotic validity of the GCC test, below.

theoremSuppose $\widetilde{\Sigma}_n$ is positive definite. Then, (a) $\widehat t_n\le\widehat r_n\le \widehat s_n$. (b) For fixed $\overline{C}_n$ and $\widetilde\Sigma_n$, there is a Lebesgue measure zero subset of $\mathbb{R}^{d_\mu}$, ${\cal M}_0$, such that \begin{equation} \widehat r_n=\widehat s_n=\widehat t_n, unless \overline{\mu}_n\in{\cal M}_0. \end{equation}

\noindentRemarks: (1) {\it Theorem (ref) justifies the use of $\widehat s_n$ as the DoF. This overrides CS23, which recommends calculating $\widehat r_n$ using an algorithm that includes a series of LPPs. The new recommendation applies to both the sCC test in CS23 and to the GCC test.}

(2) {\it Theorem (ref) is not random---it does not rely on the distribution of $\overline{\mu}_n$, $\overline{C}_n$, or $\widetilde\Sigma_n$. The result is a general feature of CQPPs. Part (a) shows that, regardless of the distribution, $\widehat s_n$ is (weakly) more conservative than $\widehat r_n$. Part (b) shows that equality holds with probability one if the conditional distribution of $\overline{\mu}_n$ given $\overline{C}_n$ and $\widetilde\Sigma_n$ is absolutely continuous. A key case where this holds is in the limit, where $\overline{C}_n$ and $\widetilde\Sigma_n$ are deterministic and $\overline{\mu}_n$ is Gaussian. Thus, a simple corollary of Theorem (ref) is that $\widehat r_n=\widehat s_n$ with probability one in the limit. }

(3) {\it The expression for $\mathcal{M}_0$ can be found in the Supplemental Appendix, equation ((ref)). The value of $\mathcal{M}_0$ may depend on the value of $\overline{C}_n$ or $\widetilde\Sigma_n$ that is fixed. Note that $\widehat{\delta}_n$ (and $\widehat{K}$) may not be unique for the definition of $\widehat s_n$, and $\widehat\psi_n$ (and $\widehat{L}$) may not be unique for the definition of $\widehat t_n$. When they are not unique, $\mathcal{M}_0$ does not depend on the choice of $\widehat{\delta}_n$ or $\widehat \psi_n$. }

(4) {\it In general, $\mathcal{M}_0$ is not the empty set. To clarify the necessity of $\mathcal{M}_0$ in part (b), we give a simple example to show that $\widehat r_n<\widehat s_n$ is possible, albeit on a set of measure zero. Suppose $d_\mu=d_\delta=1$ and $d_C=2$. Let $B=(0,1)'$, $\overline{C}_n=D=(1,1)'$, and $d=(0,0)'$. If $\overline{\mu}_n=0$, then $\widehat\mu_n=\widehat\delta_n=0$ solves ((ref)) (for any positive scalar $\widetilde\Sigma_n$). From ((ref)), $\widehat s_n=1$ because $\widehat K=\{1,2\}$, but from ((ref)), $\widehat r_n=0$. This example requires degenerate features that make $\widehat r_n<\widehat s_n$ unlikely to occur in practice.}

Uniform Asymptotic Validity of the GCC Test

Before describing the assumptions and theorems, we first clarify the true values of the parameters and the underlying distribution of the estimators. Let $F$ denote the joint distribution of $\overline{\mu}_n$, $\overline{C}_n$, and $\overline{\Omega}_n$, and let $P_F(\cdot)$ denote probabilities taken with respect to $F$. Let $\mathcal{F}_n$ be a parameter space for $F$.\footnote{We subscript $\mathcal{F}_n$ with $n$ because $\overline{\mu}_n$, $\overline{C}_n$, and $\overline{\Omega}_n$ are typically functions of a sample, $\{W_i\}_{i=1}^n$, with sample size $n$. Thus, their distribution naturally depends on $n$. We allow $\mathcal{F}_n$ to depend arbitrarily on $n$.} Let $\mu_F$ and $\Pi_{F}$ denote the true values of $\mu$ and $\Pi$. Also let $C_F=B\Pi_F+D$ and $b_F=d_n-B\mu_F$. The $F$ in the subscript makes explicit that these quantities depend on $F$. We allow the values of these parameters, together with the value of $d$, to depend on $n$ to incorporate the situation where we test a sequence of null hypotheses.\footnote{Testing a sequence of null hypotheses is required to evaluate the uniform coverage probability of a confidence set for a parameter of interest, $\theta$. Then, the inequalities that define the null hypothesis may depend on the hypothesized value of $\theta$.} We make explicit the dependence of $d$ on $n$, denoting it by $d_n$. For notational simplicity, we keep the dependence of $\mu_F$, $\Pi_F$, $C_F$, and $b_F$ on $n$ implicit.

Let $\mathcal{F}_{n0}$ be the subset of $\mathcal{F}_n$ that satisfies the null hypothesis: $\mathcal{F}_{n0}=\{F\in\mathcal{F}_n: C_F\delta\le b_F \text{ for some }\delta\in\mathbb{R}^{d_\delta}\}$. For $F\in\mathcal{F}_{n0}$, let $\delta^\ast_{F}$ be the value of $\delta$ that satisfies $C_F\delta\le b_F$. If there is more than one such value of $\delta$, we take $\delta^\ast_{F}$ to be the one that has minimum norm:

equation[equation omitted — 110 chars of source]

This mimics the definition of $\widetilde\delta_n$.

The following assumption ensures asymptotic normality of the estimators of the reduced-form parameters and consistent estimation of the asymptotic variance, at least along a subsequence of true data generating processes. It is used to show consistency of $\widetilde\delta_n$ and asymptotic uniform validity of the GCC test.

assumptionFor every sequence $\{F_n\}_{n=1}^{\infty}$ with $F_n\in {\cal F}_{n0}$ and for every subsequence, $\{n_m\}$, there exists a further subsequence, $\{n_q\}$, a vector $\mu_\infty$, a vector $d_\infty$, a matrix $\Pi_\infty$, a positive semi-definite matrix, $\Omega_\infty$, and a positive definite matrix, $\Upsilon_\infty$, such that: (i) $\mu_{F_{n_q}}\rightarrow\mu_\infty$, $\Pi_{F_{n_q}}\rightarrow \Pi_\infty$, and $d_{n_q}\to d_\infty$ (ii) $\sqrt{n_q}\left(\begin{array}{c}\overline{\mu}_{n_q}-\mu_{F_{n_q}}\\\textup{vec}(\overline\Pi_{n_q}-\Pi_{F_{n_q}})\end{array}\right)\rightarrow_d N(\mathbf{0},\Omega_\infty)$ (iii) $\Sigma_\infty:=\left(\begin{smallmatrix}I\\\delta^\ast_\infty\otimes I\end{smallmatrix}\right)'\Omega_\infty\left(\begin{smallmatrix}I\\\delta^\ast_\infty\otimes I\end{smallmatrix}\right)$ is positive definite, where $\delta^\ast_\infty:=\underset{\delta: B\mu_\infty+(D+B\Pi_\infty)\delta\le d_\infty}{\textup{argmin}}\|\delta\|$, (iv) $\overline{\Omega}_{n_q}\rightarrow_p \Omega_\infty$, (v) $\widehat\Upsilon_{n_q}\rightarrow_p \Upsilon_\infty$, and (vi) $\{\delta\in \mathbb{R}^{d_{\delta}}:B\mu_\infty+(D+B\Pi_{\infty})\delta\leq d_\infty\} \neq \emptyset$.

{Remarks:} (1) {\it Assumption (ref) is stated using subsequences in order to ensure uniformity over $\mathcal{F}_{n0}$. Part (i) assumes that $d_n$ and the sequence of true parameter values for the reduced-form parameters converge to some limits along a subsequence. This is equivalent to assuming that the parameter space for these parameters is compact.

(2) Part (ii) assumes asymptotic normality of the estimators for the reduced-form parameters along a subsequence. Part (iii) requires $\Sigma_\infty$ to be positive definite. While this may seem restrictive, it is mitigated by the way the inequalities are specified in equation ((ref)). Specifically, the $B$ matrix allows us to write the inequalities as a linear function of a core collection of reduced-form parameters that admit an estimator with a positive definite asymptotic variance matrix. The matrix $B$ absorbs any linear dependence among the estimation errors of the inequalities. Also note that $\delta^\ast_\infty$ is well-defined by Assumption (ref)(vi). Part (iv) assumes consistency of the estimator of the asymptotic variance. Part \textup{(v)} assumes consistency of the first-step weight matrix in equation \textup{((ref))} for a positive definite limit. Parts \textup{(ii)}, \textup{(iv)}, and \textup{(v)} can be verified using standard consistency and asymptotic normality arguments. For example, when the data are i.i.d.\ and the model is a moment (in)equality model, they can be verified by the Lindeberg-Feller central limit theorem and a law of large numbers for triangular arrays.

(3) Part (vi) assumes the constraint set for the limit is nonempty. Part (vi) is guaranteed under part (i) if, for example, there is a fixed compact set $\Delta$ such that $\{\delta\in \mathbb{R}^{d_\delta}: C_{F}\delta\leq b_{F}\}\subseteq\Delta$ for all $F\in{\cal F}_{n0}$.\footnote{This type of assumption is common in the literature. For example, it is assumed by Voronin2025 and GoffMbakop2025.} In that case, for any sequence $F_{n_q}\in{\cal F}_{n_q0}$, there exists a $\delta_{F_{n_q}}$ such that $B\mu_{F_{n_q}}+(B\Pi_{F_{n_q}}+D)\delta_{F_{n_q}}\leq d_{n_q}$. This sequence has a subsequence that converges to some limit $\delta_\infty$ that satisfies $B\mu_{\infty}+(B\Pi_\infty+D)\delta_\infty\leq d_\infty$, showing part (vi). Appendix (ref) shows another way to verify part (vi) under a strengthened version of Assumption \textup{(ref)}, below.}

Next, we state the stable rank condition mentioned in the introduction. We first introduce some new notation. Let $K^=\subseteq \{1,...,d_C\}$ be a set that contains the indices for the inequalities that were originally equalities. This set is special because it is always included in the set of active inequalities: $K^=\subseteq\widehat K$. For any $d_C\times d_\delta$ matrix $C$ and for any $d_C$-dimensional vector $b$, let $\mathcal{A}(C, b)=\{K\subseteq\{1,...,d_C\}: Cx\le b \text{ and } I_{K}(C x-b) = \mathbf{0} \text{ for some }x\in\mathbb{R}^{d_\delta}\}$ be the collection of all subsets of inequalities that could be simultaneously active for the system of inequalities defined by $C$ and $b$.\footnote{A combination of inequalities cannot be simultaneously active if, for example, it involves an upper and a lower bound that are parallel and separated. Such combinations are excluded from $\mathcal{A}(C, b)$.}

assumption[Stable Rank] For every sequence $\{F_n\}_{n=1}^\infty$ with $F_n\in\mathcal{F}_{n0}$ and for any subsequence, $\{n_m\}$, satisfying Assumption (ref)(i) with $C_\infty=B\Pi_\infty+D$ and $b_\infty=d_\infty-B\mu_\infty$, there is a further subsequence, $\{n_q\}$, along which \begin{equation} P_{F_{n_q}}(rk(I_K\overline{C}_{n_q}) =rk(I_K{C}_{F_{n_q}})=rk(I_KC_\infty))\to 1, as q\to \infty, \end{equation} for any $K\in{\cal A}(C_\infty, b_\infty)$ that satisfies $K^{=}\subseteq K$.

\noindentRemarks: (1) {\it We refer to Assumption (ref) as a “stable rank” condition because perturbations of the $C_\infty$ matrix in the directions of the estimation error do not change the rank. }

(2) {\it Assumption (ref) is not a necessary condition for the validity of the GCC test. {It is possible to relax {\it Assumption (ref) by reducing the number of collections of indices, $K$, for which the rank equality in ((ref)) needs to be assumed. In Appendix (ref), we show that the rank equality only needs to hold for index sets, $K$ corresponding to collections of inequalities that define a linear subspace in the limit. In cases where there are no equalities and the identified set for the nuisance parameters has a positive volume in the parameter space, every inequality could be slack, and no stable rank condition is needed.} Due to the nuances of this discussion, it is relegated to the Supplemental Appendix.}}

(3) {\it While Assumption (ref) is not necessary, it is used in an essential way in the proofs of consistency of $\widetilde{\delta}_n$ and asymptotic validity of the GCC test. Moreover, a more than superficial connection of Assumption (ref) with the weak IV problem in linear IV regression models suggests that relaxing Assumption (ref) completely may require insights from that literature. We discuss this connection in Section (ref), below.}

(4) {\it Assumptions playing a similar role as Assumption (ref) are common in the literature on subvector inference in moment inequality models and in models defined by linear systems. One type of such assumptions is a known and fixed $C$, as in GuggenbergerHahnKim2008, KaidoSantos2014, and FSST2023. In that case $\overline{C}_n=C_F= C$, and Assumption (ref) holds trivially. Other types of assumptions appear in PPHI2015, BugniCanayShi2017, ChoRussell2024, and GoffMbakop2025.\footnote{Assumption (ref) differs from constraint qualification, as considered in KMS2022. Constraint qualification restricts a fixed collection of constraints to ensure KKT conditions are necessary or sufficient or the KKT multipliers are unique. In contrast, Assumption (ref) concerns a sequence of linear inequality constraints and restricts the way they converge to a limiting set of constraints.} We discuss the connection between these assumptions in a simple example in Section (ref)}.

The following theorem states an important preliminary result: consistency of $\widetilde{\delta}_n$.

theoremSuppose Assumption (ref) holds. Let $\{F_n\}_{n=1}^\infty$ be a sequence with $F_n\in\mathcal{F}_{0n}$ and let $\{n_q\}$ be a subsequence satisfying Assumptions (ref)(i), (ii), (v), and (vi). We have that \[ \widetilde{\delta}_{n_q}\to_p\delta_\infty^\ast~\text{ and }~\delta^\ast_{F_{n_q}}\rightarrow \delta^\ast_\infty~\text{ as }~q\to\infty, \] where $\delta^\ast_\infty$ is defined in Assumption (ref)(iii).

\noindentRemark: Consistency of estimators defined by tie-breaking procedures, such as the norm minimization in the definition of $\widetilde{\delta}_n$, is especially challenging. It is surprising that Assumption (ref) is sufficient in this case. The proof uses a novel argument that establishes setwise convergence of the constraint set.

The following theorem is the main theoretical result of the paper.

theoremIf Assumptions (ref) and (ref) hold, then \begin{equation*} \underset{n\to\infty}{limsup}\sup_{F\in{\cal F}_{n0}}P_F(T_n>cv(\widehat s_n,\alpha)) \leq\alpha. \end{equation*}

{Remarks:} (1) {\it Theorem (ref) establishes the uniform asymptotic validity of the GCC test. This extends the result for the sCC test from CS23 to allow $C$ to be estimated, as long as the estimator is consistent and asymptotically normal and a stable rank condition is satisfied. The generalization is essential for handling the applications discussed in Section} (ref).

(2) {\it The asymptotic validity of the GCC test is surprising because, intuitively, the active inequalities are not necessarily binding in population and even when all inequalities are binding, the limit distribution of $T_n$ is not $\chi^2$. Indeed, the set of active inequalities does not converge to the set of binding-in-population inequalities but remains random in the limit. The key to validity of the GCC test is that the limit {\em conditional} distribution of $T_n$ given the set of active inequalities is bounded by the $\chi^2$ distribution with the associated DoF. }

(3) {\it When $P_F(\widehat{s}_n = 0)>0$, the GCC test can be slightly conservative: Its null rejection probability is between $\alpha (1-P_F(\widehat{s}_n = 0))$ and $\alpha$ asymptotically. The refinement in CS23 can be used to remove the conservativeness. We define the refined GCC (RGCC) test in Appendix} (ref){\it. The RGCC test differs from the GCC test only when $\widehat{s}_n=1$ and is also tuning parameter and simulation free. However, the refinement requires calculating $A$ and $g$ such that $\{\mu\in\mathbb{R}^{d_\mu}:B\mu+\overline{C}_n \delta\leq d \text{ for some } \delta\in\mathbb{R}^{d_\delta}\} = \{\mu\in\mathbb{R}^{d_\mu}:A \mu\leq g\}$. The computation is relatively easy when $d_C$ and $d_\delta$ are small but gets exponentially harder when $d_C$ and $d_\delta$ increase. In particular, it can have a high memory requirement. We investigate the performance of the RGCC test along with the GCC test in the simulations and the empirical illustration.}

Assumption (ref) and Instrumental Variable Regressions

To better understand Assumption (ref), we now relate it to the linear IV regression model.

example[IV Regression] Let $Y$ be a scalar dependent variable and $X$ be a $d_x$-vector of potentially endogenous regressors. Consider the IV regression model: $Y = X'\beta +\varepsilon$, where $\varepsilon$ is the error term. Let $Z$ be a $d_z$-vector of instruments that satisfy $\mathbb{E}[Z\varepsilon]=\mathbf{0}$. Suppose we are interested in the first element of $\beta$, denoted by $\theta$. Then the rest of the elements of $\beta$ are nuisance parameters, denoted by $\delta$. Let $X_1$ denote the first element of $X$ and $X_{-1}$ denote the rest of the elements. Then, the model can be represented by the following moment conditions: $\mathbb{E}[ZY] - \theta \mathbb{E}[ZX_1] - \mathbb{E}[ZX'_{-1}]\delta = \mathbf{0}$. If we conduct inference for $\theta$ by test inversion, the hypothesis to be tested for each $\theta$ value is \begin{equation} H_0: \mathbb{E}[ZY]-\theta \mathbb{E}[ZX_1] - \mathbb{E}[ZX_{-1}']\delta = \mathbf{0} for some \delta\in\mathbb{R}^{d_\delta}. \end{equation} This hypothesis is a special case of that in equation ((ref)) with $B= \left(\begin{smallmatrix}I_{d_z}\\-I_{d_z}\end{smallmatrix}\right), \mu = \mathbb{E}[ZY]-\theta \mathbb{E}[ZX_1]$, $\Pi = -\mathbb{E}[ZX_{-1}']$, $D = \mathbb{O}$, and $d = \mathbf{0}.$

In this model, Assumption (ref) allows $\mathbb{E}[ZX_{-1}']$ to change with $n$ as long as the rank does not change in the limit. Equivalently, the smallest nonzero eigenvalue of $\mathbb{E}[ZX_{-1}']\mathbb{E}[X_{-1}Z']$ does not converge to zero. Notably, zero eigenvalues are allowed. This happens, for example, when the number of instruments is smaller than the number of nuisance parameters. Assumption (ref) is weaker than the usual rank condition for strong identification of $\delta$ (under a hypothesis that fixes a value of $\theta$), which is that the smallest eigenvalue of $\mathbb{E}[ZX_{-1}']\mathbb{E}[X_{-1}Z']$ is bounded away from zero. This means that Assumption (ref) can be thought of as a “no weak identification” condition, where linear combinations of $\delta$ can be strongly identified or non-identified as long as they are not weakly identified.

Even in the weak instruments/weak identification literature, strong identification of the nuisance parameters is a useful assumption. For example, StockWright2000, Kleibergen2005, and AndrewsMikusheva2016 propose identification-robust hypothesis tests for subvectors only when the nuisance parameters are strongly identified under the null. Papers that cover inference with weakly identified nuisance parameters, including ChaudhuriZivot2011, Andrews2018, and GKM2024, recommend some version of two-step inference requiring a tuning parameter, among other complications.\footnote{An exception is GKMC2012. They focus on a homoskedastic linear IV model and show that the plug-in Anderson-Rubin test remains valid with weakly identified nuisance parameters.} Cox2022 states separate limit theory depending on whether the nuisance parameters are strongly identified under the null. In this literature, strong identification of the nuisance parameters under the null is used to guarantee that the null-imposed estimator of the nuisance parameters is consistent and asymptotically normal. In contrast, we use Assumption (ref) to ensure convergence of the constraint set in the QLR statistic to its limit in a setwise sense. These different purposes reflect the compounding complications that arise when trying to relax Assumption (ref).

Simulations

This section evaluates the finite-sample performance of the GCC test for testing inequalities that are linear in nuisance parameters. We also evaluate the RGCC test, defined in Appendix (ref). Section (ref) considers a simple one-sided hypothesis testing problem with one nuisance parameter. Section (ref) considers a more realistic model with more nuisance parameters: an interval outcome IV regression model. Also, a simulation of Example (ref) can be found in Section (ref). The bottom line of all the simulations is that the GCC and RGCC tests are easy to compute and have good size and power.

A Simple One-Sided Model

Consider a simple one-sided hypothesis testing problem with one nuisance parameter and normally distributed randomness. The model is designed to abstract from computational and asymptotic complications, so as to focus on size and power. The validity of any test for linear inequalities in this simple specification should be a necessary condition for implementing the test in practice. Also note that, because the bound on the parameter of interest is one-sided, we can compare to methods from the literature on one-sided inference for the value of a LPP.

The simple model has one nuisance parameter, no equalities, and $J$ inequalities:

align[align omitted — 245 chars of source]

This model is a special case of ((ref)) with $B=I_J$, $\mu = (\mu_1,\mu_2,...,\mu_J)'$, $\Pi = C=(c_1,c_2,...,c_J)'$, $D = \mathbf{0}_J$, and $d = -(1, 1, 0,...,0)'\theta$. The first two inequalities give an upper bound on $\theta$, while the remaining $J-2$ inequalities only bound $\delta$.

Suppose $\mu$ is estimated by $\overline{\mu}_n$ and $C$ by $\overline{C}_n$. Suppose $\overline{\mu}_n$ and $\overline{C}_n$ are sample means of independent random samples from $N(\mu,I_J)$ and $N(C,2I_J)$, respectively. The covariance matrix of $\overline{\mu}_n$ and $\overline{C}_n$ are estimated by the sample variances and covariances. Below, we consider $\mu = (-1,1,1-qn^{-1/2},...,1-qn^{-1/2})'$ and $C = (1,-1,-1,...,-1)'$ with $n=500$, $J\in\{3,10,50\}$, and $q\in\{0,4\}$. Inequalities 3 through $J$ may be binding or slack depending on the value of $q$. The identified set for $\theta$ is $(-\infty,0]$. For a fixed a value of $\theta$ in $(-\infty,0]$, the identified set for $\delta$ is $[\max(1+\theta,1-qn^{-1/2}), 1-\theta]$.

table[table omitted — 1,726 chars of source]

We first consider testing the null hypothesis $H_0: \theta=0$. Table (ref) reports the simulated rejection probabilities for various tests when $q=0$. This means that all $J$ inequalities are binding. We implement four groups of tests. The first group consists of the GCC and RGCC tests. The second group consists of the sCC and sRCC tests from CS23 and the hybrid test (ARP) from AndrewsRothPakes2023. These tests are implemented using $\overline{C}_n$ as if it were the true value. The third group consists of the MR test (BCS) from BugniCanayShi2017, procedure 3 (CCT) from CCT2018, and the recommended test (Bei) from Bei2023, which are designed for nonlinear inequalities. The fourth group consists of the recommended test (FSST) from FSST2023, the recommended test (Gaf) from Gafarov2025, and three tests (CR1-CR3) from ChoRussell2024 implemented with three choices of the tuning parameter.\footnote{CR1, CR2, and CR3 are implemented with $\underline{\epsilon}=0.1$, $\underline{\epsilon}=0.01$, and $\underline{\epsilon}=0.001$, respectively.} These tests are from the literature on inference for the value of a LPP. They are implemented for one-sided inference on the upper bound of $\theta$.\footnote{Gaf and CR1-CR3 report confidence intervals instead of hypothesis tests. For these methods, we say that a value of $\theta$ is rejected if $\theta$ does not belong to the confidence interval.}

As we can see in Table (ref):

(1) {The GCC and RGCC tests are valid for every $J$. The GCC test is somewhat conservative when $J=3$, which is expected. Unexpected is that the GCC and RGCC tests are somewhat conservative when $J=10$. This could be due to simulation noise}.

(2) {The sCC, sRCC, and ARP tests are invalid. This demonstrates the need to account for the estimation error in $\overline{C}_n$.\footnote{Another way to implement these tests is to consider the case that the inequalities in ((ref)) hold conditionally on $\overline{C}_n$. Then, the sCC, sRCC, and ARP tests can be implemented with the conditional variance of $\overline{\mu}_n$ given $\overline{C}_n$. Since $\overline{\mu}_n$ is independent of $\overline{C}_n$, the conditional variance is the same as the unconditional variance and the implementation is the same. Thus, Table (ref) shows that neither way to implement these tests is valid. The problem is that the inequalities in ((ref)) do not hold conditionally on $\overline{C}_n$.} }

(3) {Concerning the tests in the third group, CCT is invalid for $J>3$. It is only valid under a high-level condition that is not satisfied in this model; see Assumption 4.7 in CCT2018. BCS is valid and becomes quite conservative when $J=50$. Bei is a more computationally feasible version of BCS. Table (ref) suggests Bei has mild to moderate over-rejection. It is unclear why the null rejection probabilities for BCS and Bei are so different. }

(4) {Concerning the tests in the fourth group, FSST is understandably invalid because it requires known Jacobian. Surprisingly, Gaf is also invalid for $J>3$. This could be because a rank condition is not satisfied; see Assumption 2 in Gafarov2025. CR1-CR3 appear to be very sensitive to the choice of the tuning parameter. CR1 is very conservative while CR2 and CR3 have moderate over-rejection when $J=50$. }

(5) {Computationally, the GCC and RGCC tests are by far the fastest among the tests considered. Comparing their times to the sCC and sRCC tests demonstrates the computational improvement of the new algebraic formula for the DoF. For CR1-CR3 and Gaf, the computational time is for the calculation of a confidence interval and thus is not comparable to the computation time for a single hypothesis. }

figure[figure omitted — 12,303 chars of source]

\addtocounter{figure}{-1}

We also consider the power functions of some of the tests as a function of the hypothesized value of $\theta$. In addition to the GCC and RGCC tests, we include the BCS and Bei tests as benchmarks and the CR1 and CR2 tests to see whether the sensitivity of their NRPs to the tuning parameter carries over to the power functions. These tests are included because they are the ones that are valid or exhibit only moderate over-rejection in Table (ref).

We test the inequalities in ((ref)) for a grid of values of $\theta$ between $-1/\sqrt{n}$ and $6/\sqrt{n}$. Dividing the grid by $\sqrt{n}$ ensures the power function approaches the asymptotic local power function as $n\rightarrow\infty$. For $\theta\le 0$ (the shaded region), the rejection probabilities are under the null and thus should be less than or equal to $\alpha=5\%$. For $\theta>0$, the rejection probabilities represent the power of the tests. Figure (ref) reports the power curves for $q\in\{0,4\}$ and $J\in\{3,10,50\}$. When $q=0$, the 3rd to the $J$th inequalities are binding at $\delta=1$, and when $q=4$, these inequalities are slack at $\delta=1$.

As we can see in Figure (ref):

(1) {The GCC and RGCC tests are very fast with good power for all $J\in\{3,10,50\}$ and $q\in\{0,4\}$. The power is especially impressive when $q=4$, so $J-2$ of the inequalities are slack. Also note that, when the number of inequalities increases, the difference between the GCC and RGCC power curves decreases, especially when all the inequalities bind.}

(2) {The BCS test has good power for $J=3$, but becomes more conservative, and therefore less powerful, for larger $J$. The Bei test is more powerful than BCS, GCC, and RGCC, with especially high power when $q=0$ and $J\in\{10,50\}$. These specifications correspond to the cases of moderate over-rejection of the Bei test at $\theta=0$.}

(3) {The sensitivity that the CR tests show in Table (ref) is reflected in power. CR1 has low power, and the power of CR2 is closely related to the over-rejection of the test at $\theta=0$.}

An Interval Outcome IV Regression Model

This subsection considers the interval outcome IV regression model from Example (ref). We simulate the power curves of the GCC and RGCC tests. We also include the BCS and Bei tests as benchmarks. Since this is a two-sided problem, we do not include recommendations from the literature on inference for the value of a LPP.

The model is based on the aggregate demand model considered in GLS23, where the market shares are noisy measures of conditional choice probabilities and may contain zero values. The model boils down to the interval outcome IV regression in Example (ref). Write out $X=(X_1, X_2, W')'$, where $X_1$ is a scalar endogenous regressor, $X_2$ is a scalar exogenous regressor, and $W$ is a $d_W$-vector of additional exogenous controls. Similarly, write out $\beta=(\theta_1, \theta_2, \gamma')'$ with $\gamma\in\mathbb{R}^{d_W}$. Also let $Z_e$ be an excluded exogenous instrument. We take $Z=\mathcal{I}(X_2, W, Z_e)$ to be a non-negative vector-valued instrumental function.

In this model, there is a latent market share that satisfies a logit specification:

equation[equation omitted — 166 chars of source]

where $\theta_{10}$ and $\theta_{20}$ denote the true values of $\theta_1$ and $\theta_2$ and $\epsilon$ is an error term. (The true value of $\gamma$ is zero.) In the framework of GLS23, $s^\ast$ is the (unobserved) conditional probability that a representative consumer buys a product, and $Y^\ast=\log(s^\ast)-\log(1-s^\ast)$ is the mean utility of the product. We observe a market share $s_N\sim Binomial(N,s^\ast)/N$, where $N$ is the number of participants in the market. GLS23 argue that $Y^U = \log(s_N+2/N)-\log(1-s_N+0.00125)$ and $Y^L = \log(s_N+0.00125)-\log(1-s_N+2/N)$ satisfy $\mathbb{E}[Y^L|Z]\leq \mathbb{E}[Y^\ast|Z]\leq \mathbb{E}[Y^U|Z]$, which justifies the use of the model in Example (ref).

We take $X_2$, $Z_e$, and the components of $W$ to be independent $Bernoulli(0.5)$ random variables, except for the first element of $W$, which is taken to be the constant one. We take $Z=\mathcal{I}(X_2,W,Z_e)$ to be the vector of indicators for each point in the support of $(X_2,W,Z_e)$. When $d_W=1,2,3$, the dimension of $Z$ is $4,8,16$, respectively. Also, the fact that there are two inequalities in ((ref)) means that there are $8$, $16$, and $32$ inequalities, respectively. Independently of $Z$, let $\epsilon\sim \min\left(\max\left(N(0,1),-4\right),4\right)$. Then, let $X_1=\mathds{1}\{Z_e+\epsilon/2>0\}$ be the endogenous regressor. We calculate $s^\ast$ according to ((ref)) with $\theta_{10}=\theta_{20}=-1$. We simulate $s_N$ with $N=100$ independently from $Z$ and $\epsilon$. This specifies the data generating process for all the observed variables: $s_N$, $X_1$, $X_2$, $W$, and $Z_e$. We simulate a sample of size $n$ from this model for $n\in\{500, 1000\}$ and calculate confidence intervals for $\theta_2$ treating $\delta=(\theta_1,\gamma')'$ as nuisance parameters.\footnote{The same model is considered in Section 5.2 of CS23. However, CS23 construct confidence intervals for $\theta_1$, treating $(\theta_2,\gamma')'$ as nuisance parameters. Since $X_2$ and $W$ are both exogenous, it is valid to conduct inference conditional on $(X_2,W,Z_e)$ using the sCC test in CS23 because the Jacobian of the sample moments with respect to $(\theta_2,\gamma')'$ is known given the sample for $(X_2,W,Z_e)$. In contrast, this paper takes $(\theta_1,\gamma')'$ to be the nuisance parameters. Then, the Jacobian with respect to $\theta_1$ is not known even after conditioning on $(X_2,W,Z_e)$. Thus, the sCC test in CS23 is invalid.}

Figure (ref) plots simulated power curves for the GCC, RGCC, BCS, and Bei tests. In each graph, the horizontal axis represents the value of $\theta_2$, while the vertical axis represents the rejection probability.\footnote{The rejection probabilities reported are frequencies that each $\theta_2$ value lies outside the confidence interval for $\theta_2$. The confidence intervals are computed using a bisection algorithm for the endpoints.} The shaded region indicates the identified set for $\theta_2$. Since the control variables have zero coefficients and are independent of the other random variables, they do not affect the identified set of $\theta_2$. In the legend, each number in the square brackets is the median computational time (in seconds) to compute one confidence interval. We make the following remark on Figure (ref).

figure[figure omitted — 9,036 chars of source]

\addtocounter{figure}{-1}

\noindentRemark: The power curves of the GCC, RGCC, BCS, and Bei tests are remarkably similar. All four have rejection probabilities below the nominal level 5% in the shaded region. All four have increasing power as $\theta_2$ deviates from its identified set and as the sample size increases from 500 to 1000. Overall, the GCC and RGCC tests are able to match the size and power performance of the BCS and Bei tests while being much faster computationally.

Empirical Illustration: Female Labor Supply

This section demonstrates the GCC test in the model of female labor supply considered in KT16. KT16 use a model of state transition probabilities with states that indicate (1) the earnings of the individual (zero or not employed, positive and below the federal poverty line, or above the federal poverty line), (2) whether the individual participates in welfare, and (3) whether the individual under-reports her earnings in order to qualify for welfare. KT16 use data from the Manpower Development Research Corporation (MDRC) Jobs First study, which is a random experiment that assigned women with young children to one of two welfare programs: the Aid to Families with Dependent Children (AFDC) welfare program or the Jobs First Temporary Family Assistance (JF) program.\footnote{Instructions for accessing the datasets and the replication codes provided by KT16 can be found on the AEA webpage: \url{https://www.aeaweb.org/articles?id=10.1257/aer.20130824}.}

Relative to AFDC, the JF program primarily changes how eligibility for and the amount of government transfers respond to earnings.\footnote{The JF welfare reform introduces other changes too, such as stricter work requirements and changed administration of the Food Stamps program. We focus on the changes that are relevant for bounding the state transition probabilities. For more information, see the description in KT16.} Under JF, for individuals earning under the federal poverty line (FPL), the government transfer does not decrease as earnings increase: everyone under the FPL receives the same government transfer. This is more generous than the policy under AFDC, which had government transfers decrease after earnings reached a threshold. This more generous policy may induce unemployed women (or women who would be unemployed under AFDC) to gain employment that pays below the FPL. This represents a labor supply response along the extensive margin. Another noteworthy feature of the JF program is that the government transfers abruptly drop to zero when earnings cross the FPL. This may incentivize some women who would be employed with earnings above the FPL (under AFDC) to decrease their earnings (or under-report their earnings) in order to be eligible for welfare. This represents a labor supply response along the intensive margin. The goal in KT16 is to distinguish these two responses without imposing strong assumptions on the utility functions of the individuals.

KT16 distinguish 7 labor supply/welfare participation states:

enumerate• zero earnings, welfare nonparticipation, • positive earnings below FPL, welfare nonparticipation, • earnings above FPL, welfare nonparticipation, • zero earnings, welfare participation, truthful reporting of earnings, • positive earnings below FPL, welfare participation, truthful reporting of earnings, • positive earnings below FPL, welfare participation, underreporting of earnings, and • earnings above FPL, welfare participation, underreporting of earnings.

In the label for each state, the number indicates the level of earnings: “0” indicates zero earnings or not employed, “1” indicates positive earnings below the FPL, and “2” indicates earnings above the FPL. The letter indicates welfare participation and under-reporting of earnings: “n” indicates nonparticipation in welfare, “r” indicates welfare participation with truthful reporting of earnings, and “u” indicates welfare participation with under-reporting of earnings.

Each individual is associated with two states, the one she would choose under AFDC and the one she would choose under JF. Because AFDC is the status quo, we label the state transition probabilities with an individual's choice under AFDC first. For example, $\pi_{\text{0n},\text{1r}}$ is the conditional probability that an individual who chooses 0n under AFDC would choose 1r under JF. Because of the features of the JF reform, KT16 argue that only nine state transition probabilities need to be considered. The first group are flows out of 0r: $(\pi_{\text{0r},\text{0n}}, \pi_{\text{0r},\text{1n}}, \pi_{\text{0r},\text{2n}}, \pi_{\text{0r},\text{1r}}, \pi_{\text{0r},\text{2u}})$. Individuals with zero earnings and participating in welfare under AFDC may transition to any other state except 1u.\footnote{Under JF, no one chooses 1u because everyone earning under the FPL receives the same government transfer, so there is no reason to under-report earnings. Also note that $\pi_{\text{0r},\text{0r}}$ is not needed because it can be calculated as one minus the others. This is true in general. We do not need to include the transition probabilities from one state into itself because they are determined by the other transition probabilities.} The second group are flows into 1r: $(\pi_{\text{0n},\text{1r}}, \pi_{\text{1n},\text{1r}}, \pi_{\text{2n},\text{1r}}, \pi_{\text{0r},\text{1r}}, \pi_{\text{2u},\text{1r}})$. Individuals who choose 1r under JF may have chosen any state under AFDC.\footnote{Individuals who choose 1u under AFDC are guaranteed to choose 1r under JF, so $\pi_{\text{1u},\text{1r}}=1$ and there is no need to include it as a free parameter.} KT16 argue that all other transition probabilities can be set to zero either because the budget sets for the individuals is unchanged between the two policies or because the combination of choices would violate weak assumptions on individuals' utility functions. Let \[ \delta = (\pi_{\text{0n},\text{1r}}, \pi_{\text{1n},\text{1r}}, \pi_{\text{2n},\text{1r}}, \pi_{\text{0r},\text{0n}}, \pi_{\text{0r},\text{1n}}, \pi_{\text{0r},\text{2n}}, \pi_{\text{0r},\text{1r}}, \pi_{\text{0r},\text{2u}}, \pi_{\text{2u},\text{1r}})' \] collect the transition probabilities into a vector of nuisance parameters. (Note that one of the transition probabilities is common to both groups.)

There is an additional problem: it is unobserved whether an individual under-reports her income. Thus, the marginal probabilities of only six observable states---combining 1r and 1u---is identified. KT16 show that after accounting for this problem, the resulting equalities are still linear in $\delta$ and they can be rewritten into five non-redundant equalities of the form $\Gamma\delta = m$, where

equation[equation omitted — 1,066 chars of source]

where $p_{\text{s}}^{\text{A}}$ denotes the marginal probability that the individual chooses state s under the AFDC policy for $s\in\{\text{0n}, \text{1n}, \text{2n}, \text{0r}, \text{2u}\}$, and similarly for $p_{\text{s}}^{\text{J}}$ for the JF policy. In addition, $\delta \in [0,1]^9$ and $\pi_{\text{0r},\text{0n}}+\pi_{\text{0r},\text{1n}}+\pi_{\text{0r},\text{2n}}+\pi_{\text{0r},\text{1r}}+\pi_{\text{0r},\text{2u}}\leq 1$, which can be imposed by an appropriate choice of $A$ and $b$ in ((ref)).\footnote{In Section (ref), we give a simple sufficient condition for Assumption (ref) in this model.}

We follow KT16 and report confidence intervals for each transition probability. In addition to the GCC test, we also implement the RGCC test, defined in Section (ref). The GCC and RGCC tests are implemented by rewriting the restrictions in terms of $B$, $\mu$, $\Pi$, $D$, and $d$ using ((ref)). We can similarly define estimators of $\mu$ and $\Pi$ from the sample averages that estimate $p_s^t$ for $s\in\{\text{0n}, \text{1n}, \text{2n}, \text{0r}, \text{2u}\}$ and $t\in\{\text{A},\text{J}\}$.\footnote{We copy KT16 and use weighted sample averages with propensity score weights to adjust for baseline differences.} We estimate the asymptotic variance by bootstrapping the sample averages with a cluster bootstrap, clustered at the case level and with 1000 bootstrap draws.\footnote{This is the same implementation of the bootstrap that KT16 use, except that we bootstrap the estimators of $p_s^t$ while they bootstrap the formulas for the bounds that they calculate after eliminating the nuisance parameters manually. This means that our variance estimators, while asymptotically equivalent, are numerically different.} The endpoints of the GCC and RGCC confidence intervals are calculated using a bisection algorithm.

KT16 manually eliminate the nuisance parameters and work out explicit formulas for the bounds of each element of $\delta$. To avoid the dependence on tuning parameters that is prevalent in the literature on testing inequalities, they report two confidence intervals. One confidence interval, called Na\"ive, is constructed by ignoring the uncertainty in which bounds bind. This interval is formed by the single lowest upper (highest lower) estimated bound plus (minus) its standard error. The asymptotic coverage probability of this interval is unknown. The other interval, called Conservative, assumes all population bounds bind simultaneously, leading to asymptotically valid but often overly conservative inference.

table[table omitted — 2,026 chars of source]

Table (ref) reports the confidence intervals for each transition probability.\footnote{The values for the Na\"ive and Conservative confidence intervals are slightly different from the ones in the published version of KT16. We calculated these values using the KT16 replication code without changes. The differences are likely due to differences in random number generation across versions of Stata when implementing the bootstrap.} As Table (ref) shows:

(1) { All the confidence intervals are qualitatively similar. They provide evidence for the same heterogeneous labor supply responses: statistically significant outflows from state 0r, corresponding to an increase in labor supply along the extensive margin, and statistically significant inflows into state 1r, especially from state 2n, corresponding to a decrease in labor supply along the intensive margin as women decrease their earnings to qualify for welfare. }

(2) { One would expect the endpoints of the GCC and RGCC confidence intervals to lie between the endpoints of the Na\"{i}ve and Conservative confidence intervals. That is mostly true, but there are a few noteworthy exceptions. } (a) { For $\pi_{\text{0r},\text{1n}}$, the GCC and RGCC confidence intervals are narrower than the Na\"{i}ve confidence interval. In this case, the two smallest upper bounds are very close to each other. Together, they provide stronger statistical evidence than just the one that is active. The GCC and RGCC tests respond to this statistical evidence in a way that the Na\"{i}ve confidence interval does not. } (b) {For $\pi_{\text{1n},\text{1r}}$ and $\pi_{\text{2n},\text{1r}}$, the GCC and RGCC confidence intervals are wider than the Conservative confidence intervals. This is not surprising for the GCC confidence interval because the GCC test is conservative when there is one binding inequality.\footnote{For all the transition probabilities in Table (ref), there is only one nontrivial lower bound. This explains both why the Na\"{i}ve lower bounds are equal to the Conservative lower bounds and why the GCC confidence interval appears conservative for the lower bounds.} This is surprising for the RGCC confidence interval because with one binding inequality, the RGCC test is asymptotically equivalent to the optimal one-sided test. This is likely due to the numerical difference between the variance matrix estimators used in the RGCC and the Conservative confidence intervals.}

(3) {To compute all nine confidence intervals, the GCC and RGCC methods took about 4 seconds and 220 seconds, respectively. Both approaches are quite feasible, especially considering the fact that manual elimination of the nuisance parameters is not needed to implement the GCC and RGCC tests. }

Conclusion

This paper proposes a simple, tuning-parameter-free test that is designed for inequality testing problems that are linear in nuisance parameters, including specification testing and subvector inference in moment (in)equality models and inference for parameters bounded by linear programs. We prove asymptotic uniform validity of the test under a stable rank condition and demonstrate its size, power, and computational performance in simulations and an empirical illustration.