Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
121,682 characters · 14 sections · 64 citation commands
Testing Inequalities Linear in Nuisance Parameters
{\bf Keywords:} Linear Program, Moment Inequalities, Quadratic Programming, Subvector Inference, Uniform Inference
We propose a simple new test for hypotheses of the form $H_0$: there exists a $\delta$ such that $C\delta\leq b$, where elements of the Jacobian matrix $C$ and the intercept vector $b$ are reduced-form parameters that can be consistently estimated, and elements of $\delta$ are unknown parameters whose values are partially identified by the inequalities under $H_0$. Since the inequalities, rather than $\delta$, are of central interest, $\delta$ is a nuisance parameter vector. Hypotheses of this form arise in specification testing and subvector inference for linear unconditional moment (in)equality models and in inference for parameters bounded by linear programs, including discrete instrumental variable (IV) models with shape restrictions and policy relevant treatment effect models. These models have wide applications in empirical work. We explain the applications and give examples in Section (ref).
Testing this hypothesis is non-standard both because the nuisance parameter $\delta$ may not be point-identified and because the hypothesis involves inequalities. As a result, commonly used test statistics have non-standard asymptotic distributions involving parameters that cannot be consistently estimated, in particular, the local slackness of the inequalities evaluated at the true value of $\delta$. This complicates the design of critical values. A common approach is to simulate the asymptotic distribution with a conservative estimator of the local slackness plugged in. However, the conservative estimators typically involve user-chosen tuning parameters that introduce arbitrariness to the procedure. Moreover, the simulated critical values can be computationally burdensome.
In the special case that the Jacobian is {\em known}, CoxShi2023, hereafter CS23, propose the subvector conditional chi-squared (sCC) test that does not require simulation or user-chosen tuning parameters and yet still has uniform asymptotic size control and good power. The simplicity of the test is achieved by considering the conditional distribution of the quasi-likelihood ratio (QLR) statistic given the identity of the active inequalities.\footnote{An inequality is active if it holds with equality at our null-imposed estimator of $\delta$.} The conditional distribution is shown to be bounded by a chi-squared distribution with degrees of freedom (DoF) dependent on the conditioning event. CS23 recommends computing the DoF by solving a sequence of linear programming problems.
Our first contribution is to derive an elementary formula for the DoF that replaces the recommendation in CS23. The formula makes the computation of the critical value elementary. Implementing the sCC test is now no harder than calculating the test statistic, which is a convex quadratic programming problem (CQPP). The formula also reveals an intuitive interpretation of the sCC test: The sCC test turns out to be the same as the classic Sargan-Hansen's J test for a moment equality model, where the equalities are determined by the active inequalities.
While the sCC test is attractive, the known Jacobian assumption significantly restricts its applicability. In both linear unconditional moment inequality models and models with a parameter of interest bounded by linear programs, a known Jacobian only applies in special cases. The known Jacobian assumption even rules out conditional moment inequality models where $\delta$ includes coefficients on endogenous covariates. The second contribution of this paper is to propose a new test, called the generalized conditional chi-squared (GCC) test, that accounts for the estimation error of an {\em unknown} Jacobian while maintaining the simplicity, size, and power properties of the sCC test.
Estimating the Jacobian presents two challenges. The first challenge is finding the limit of the constraint set in the definition of the QLR statistic. In general, convergence of the matrix of coefficients in a system of inequalities is insufficient for the set defined by those inequalities to converge in a setwise sense. We show that a simple stable rank condition on the Jacobian is sufficient for convergence of the constraint set. The stable rank condition requires the rank of certain submatrices of the Jacobian to not change in the limit. The submatrices that need to satisfy the condition are minimal in some sense. They are associated with collections of inequalities that (implicitly or explicitly) define equality restrictions in the limit. While the stable rank condition is not innocuous, it relaxes the commonly used strong identification assumption in moment equality models, which requires the Jacobian to be full rank.
The second challenge is finding a consistent variance estimator. The effect of the estimated Jacobian on the variance depends on $\delta$, but $\delta$ is only partially identified. This makes the additional variance term difficult to account for. Our solution is to use a first-stage estimator of $\delta$ that converges to a point in the identified set for $\delta$. The procedure resembles optimal weighting in two-step generalized method of moments (GMM). The GCC test compares this two-step statistic to a chi-squared critical value with DoF determined by the formula from our first contribution. We also define a refinement of the GCC test in the supplemental appendix that has slightly more power.
We next review three strands of related literature. The first consists of papers that propose tests of moment inequalities that are possibly nonlinear in nuisance parameters. These include BugniCanayShi2015, BugniCanayShi2017, CCT2018, BelloniBugniChernozhukov2018, KMS2019, and Bei2023. These methods use critical values that are nontrivial to compute and require user-chosen tuning parameters to be adaptive to the unknown slackness of the inequalities.\footnote{An exception is Procedure 3 in CCT2018 in that it does not require any tuning parameter or simulation. We include this procedure in the simulations.} In contrast, the GCC test uses an algebraic critical value and is adaptive to the unknown slackness of the inequalities without any tuning parameter, at the price of requiring linearity.
The second strand of related literature is composed of CS23 and AndrewsRothPakes2023, which consider conditional moment inequalities that are linear in the nuisance parameters. These papers rely on an idea that if the Jacobian depends only on random variables on which the inequalities hold conditionally, then the Jacobian can be treated as known. However, this idea does not apply if the moment inequalities are unconditional or, more generally, if the Jacobian depends on a random variable for which the moment inequalities do not hold conditionally. Our method applies regardless of whether the Jacobian can be treated as known.
When our test is applied to inference for a scalar parameter (say $\theta$) bounded by linear programs, it is related to the third strand of literature. This literature addresses the problem of inference for the value of a linear programming problem (LPP). The papers include FreybergerHorowitz2015, BaiSantosShaikh2022, FSST2023, ChoRussell2024, Gafarov2025, Voronin2025, and GoffMbakop2025, among others. There is a subtle technical difference between our setting and this literature: while our test is inverted to yield a confidence interval for $\theta$, the LPP literature aims at constructing confidence intervals for the upper or lower bound for $\theta$ defined by the value of a LPP. While the two problems are distinct, they are closely related. Our confidence interval by design can cover either bound with correct nominal coverage probability (asymptotically), and in the LPP literature, one-sided confidence intervals of the appropriate direction for either bound are also valid confidence intervals for $\theta$.\footnote{In the LPP literature, two-sided confidence intervals for $\theta$ can be obtained by combining two one-sided confidence intervals via a Bonferroni adjustment.} Thus, the methods can be used for the same empirical problems. Notably, our method is the only one that is tuning parameter and simulation free.
We include a simple one-sided specification in the simulations in order to compare the GCC test to representative papers in all three strands of the literature. In addition to the simple simulation, we also evaluate the GCC test in two realistic simulation examples: one of an interval outcome instrumental variables (IV) model as in GandhiLuShi2023, hereafter GLS23, and the other of bounds on policy relevant treatment effects as in MogstadSantosTorgovitsky2018, hereafter MST18. The simulations show that the GCC test is computationally very fast with good size and power.
We also implement the GCC test in an empirical analysis of female labor supply in response to a welfare policy reform. KlineTartari2016, hereafter KT16, estimate bounds on the treatment responses by manually eliminating the nuisance parameters from revealed preference inequalities. The GCC test provides uniformly valid inference for the treatment responses based directly on the revealed preference inequalities. Overall, we find statistically significant heterogeneous responses to the policy change, which agree with the results in KT16.
The rest of the paper is organized as follows. Section (ref) describes the setup and applications. Section (ref) defines the GCC test. Section (ref) describes the theoretical properties of the GCC test. Section (ref) presents the simulations. Section (ref) presents the empirical illustration. Section (ref) concludes. An appendix contains proofs of the theorems, while a supplemental appendix includes additional results, proofs, simulations, and discussion.
We are interested in testing the hypothesis
where $C$ is a $d_C\times d_\delta$ matrix of reduced-form parameters, $b$ is a $d_C$-dimensional vector of reduced-form parameters, $d_\delta$ is the dimension of $\delta$, and $d_C$ is the number of inequalities. To make an invertibility assumption imposed later as unrestrictive as possible, we add some structure to $C$ and $b$. We assume that
where $B$ is a known $d_C\times d_\mu$ matrix, $D$ is a known $d_C\times d_\delta$ matrix, $d$ is a known $d_C$-dimensional vector, $\Pi$ is an unknown $d_\mu\times d_\delta$ matrix of reduced-form parameters, and $\mu$ is an unknown $d_\mu$-dimensional vector of reduced-form parameters. This structure separates the unknown and estimated components from the known components in the Jacobian $C$ and the intercept $b$. It is satisfied in all the examples considered below. Typically, $B$ has more rows than columns, and it absorbs the linear dependence across rows for the estimation noise of the inequalities. This allows us to accommodate inequalities with linearly dependent estimation errors, which arise when we write an equality as a pair of opposing inequalities, when the model contains a deterministic constraint such as a shape or sign restriction, or when the law of total probability dictates that a weighted sum of the inequalities involves no unknown quantities under $H_0$.
With the structure in ((ref)), $H_0$ can be equivalently written as:
Let $\overline{\mu}_n$ and $\overline{\Pi}_n$ be estimators of $\mu$ and $\Pi$. Let $\overline{C}_n=D+B\overline{\Pi}_n$ and $\bar b_n=d-B\overline{\mu}_n$. In the next two subsections, we describe two classes of models that are covered by this framework.
Moment (in)equality models are used to address data or modeling incompleteness issues, including missing data, multiple equilibria, and large intractable games.\footnote{For empirical applications and current statistical methods for such models, see the survey papers by CanayShaikh2017, HoRosen2017, and Molinari2020.} Here, we show that both specification testing and subvector inference for moment inequality models fit the hypothesis in equation ((ref)) when the moments are linear in the parameters.
Consider the moment (in)equality model:
where $\overline{m}_n(\beta) = (\overline{m}^{eq}_n(\beta)',\overline{m}_n^{ineq}(\beta)')'$ is a $\mathbb{R}^{d_m}$-valued sample moment function, $\beta$ is an unknown vector of parameters, and ${\cal B}$ is its parameter space. Specifically, let $\overline{m}_n(\beta) = n^{-1}\sum_{i=1}^n m(W_i,\beta)$, where $\{W_i\}_{i=1}^n$ is a sample of observable variables and $m(\cdot, \beta)$ is a function known up to the unknown parameter $\beta$. Let the number of equalities be denoted $d_{eq}$ and the number of inequalities be denoted $d_{ineq}$, so that $d_m = d_{eq}+d_{ineq}$.
Suppose $\overline{m}_n(\beta)$ is linear in $\beta$. That is, $\overline{m}_n(\beta) = \overline{\Gamma}_n\beta + \overline{\eta}_n$ for $\overline{\Gamma}_n = \partial\overline{m}_n(\beta)/\partial\beta'$ and $\overline{\eta}_n = \overline{m}_n(\mathbf{0})$. Suppose ${\cal B} = \mathbb{R}^{d_\beta}$.\footnote{More generally, if ${\cal B}$ is a polyhedral set, then the deterministic inequalities that define ${\cal B}$ should be included when writing the hypothesis in the form of ((ref)). We show how deterministic constraints can be incorporated in the next subsection.} Consider the following types of problems:
Two remarks are in order regarding subvector inference:
{\bf Remarks:} (1) {\it Linearity in $\theta$ is not needed for subvector inference. The discussion remains unchanged if $\overline{m}_n(\theta,\delta)$ is linear in $\delta$ and nonlinear in $\theta$. }
(2) {\it Sometimes, the parameter of interest is not a subvector of $\beta$, but instead a linear function of $\beta$. That is, $\theta = \Lambda\beta$ for a known $d_\theta\times d_\beta$ full-rank matrix $\Lambda$ with $d_\theta<d_\beta$. One approach is to reparameterize $\beta$ so that $\theta$ becomes a subvector of the new parameter. Let $\Lambda^c$ be a $(d_\beta-d_\theta) \times d_\beta$ row-augmenting matrix so that $\left(
\right)$ is nonsingular.\footnote{In Matlab, one can find such a matrix by applying the function \texttt{null}(~) on $\Lambda$.} The reparameterization is given by $\gamma = \left(
\right)\beta$. Then, plugging $\beta = \left(
\right)^{-1}\gamma$ into \textup{(\ref{mi})} reparameterizes the model so that $\theta$ is a subvector of $\gamma$. Equivalently, one can add $\theta= \Lambda\beta$ to the model as deterministic constraints and treat $\beta$ as the nuisance parameter. }
We end this subsection with two examples of linear moment inequality models.
Recently, an important class of models have arisen in the structural estimation literature where a scalar parameter of interest is not point-identified but bounded by the values of linear programming problems (LPPs).\footnote{Some notable examples appear in KT16, MST18, KKLS2021, and SyrgkanisTamerZiani2021.} The constraint sets for the LPPs are defined by linear (in)equalities, where the constants and coefficients in the (in)equalities are unknown parameters to be estimated. To fix ideas, suppose the parameter of interest is
where $\gamma\in\mathbb{R}^{d_\delta}$ is a known vector that defines a linear combination of the nuisance parameters $\delta$. The constraint set is defined by a collection of (in)equalities:
where the elements of $m\in \mathbb{R}^{d_{\Gamma}}$ and $\Gamma\in\mathbb{R}^{d_\Gamma\times d_\delta}$ are known or can be consistently estimated, $A$ is a known $d_A\times d_\delta$ matrix, and $b\in\mathbb{R}^{d_A}$ is a known vector.\footnote{We focus on the case $A$ and $b$ are known because it is common in applications. Unknown and estimated $A$ and/or $b$ can also be covered.} The known inequalities in ((ref)) define a parameter space for $\delta$. For example, if $\delta$ is a vector of weights, then $\delta$ takes values in the simplex, which can be represented by an appropriate choice of $A$ and $b$.
Inference on $\theta$ can be based on the GCC test for the existence of a value of $\delta$ that simultaneously satisfies ((ref)) and ((ref)) at a hypothesized value of $\theta$. A confidence interval for $\theta$ can be calculated by inverting a family of tests. The restrictions in ((ref)) and ((ref)) can be written in the form of ((ref)) with
Note that the first two rows and the last $d_A$ rows of $B$ are zeros to accommodate the deterministic constraints $\theta = \gamma\delta$ and $A\delta\leq b$.
Another approach to inference on $\theta$ is to use LPPs. From the point of view of identification, $\theta$ is bounded sharply by $\theta_\min = \min_{\delta:~\text{(\ref{deltaID}) holds}}\gamma'\delta$ and $\theta_\max = \max_{\delta:~\text{(\ref{deltaID}) holds}}\gamma'\delta$. Then one can construct one-sided confidence intervals that are bounded from below (above) for $\theta_{\min}$ ($\theta_{\max}$) and use them as confidence intervals for $\theta$. Inference for the value of a LPP is generally based on plugging the estimators of $m$ and $\Gamma$ into the LPPs and simulating or bootstrapping the asymptotic distributions of these estimators. However, the asymptotic distributions depend on which corner or face of the constraint set solves the LPP, which is not smooth as a function of the estimated reduced-form parameters. Thus, the na\"{i}ve strategy of bootstrapping the value of a LPP is generally invalid. In order to obtain valid inference, the LPP literature recommends various modifications that involve tuning parameters and/or simulating/bootstrapping nonstandard distributions. Using the GCC test to obtain a confidence interval for $\theta$ avoids these complications.
{\bf Remarks:} (1){\it The GCC test is valid for any hypothesized value of $\theta$ in $[\theta_\min, \theta_\max]$, including the endpoints. Thus, the GCC test is a valid way to do inference for the value of a LPP. However, the GCC test may direct power in a one-sided way. In particular, when $\theta_\max-\theta_\min$ is not small, the GCC test is effectively one-sided for the value of the LPP (though still two-sided for $\theta$). Therefore, when the parameter of interest is $\theta_\min$ or $\theta_\max$ instead of $\theta$ and the researcher desires two-sided inference, then some of the inference recommendations from the LPP literature are preferred. }
(2) {\it In some applications, $\gamma$ is unknown and estimated. There is a convenient way to write ((ref)) and ((ref)) in the form of ((ref)). The idea is to add an element to $\mu$ that is always zero, while adding $\gamma'$ as a row of $\Pi$. Specifically, ((ref)) can be satisfied when $\gamma$ is unknown by letting
More generally, if one of the equations in ((ref)) has a known intercept but the corresponding row of the Jacobian is unknown, then the intercept should be included in $d$ while the row of the Jacobian should be included in $\Pi$, possibly including a zero in $\mu$ and augmenting the columns of $B$.\footnote{This idea can be applied more generally. Consider a generalization of (2), where $C$ can be written as $C=B_1\Pi_1+B_2\Pi_2+D$ and $b=d-B_1\mu_1$. That is, $B_2\Pi_2$ is a component of $C$ that needs to be estimated but cannot be written as a linear combination of the columns of $B_1$. Then, we can satisfy the structure in (2) by taking $B=[B_1,B_2]$ and $\mu=(\mu'_1,\mathbf{0}')'$. This works because we do not require the estimator of $\mu$ to have a nonsingular variance matrix, but we only require the estimator of $\mu+\Pi\delta$ to have a nonsingular variance matrix for a value of $\delta$ in the identified set; see Assumption (ref)(iii), below.} This demonstrates the flexibility of the specification in ((ref)). }
We demonstrate the relevance of the setup in ((ref)) and ((ref)) with examples.
In this section, we define the generalized conditional chi-squared (GCC) test. The test depends on $\overline{\mu}_n$ and $\overline{\Pi}_n$, consistent and asymptotically normal estimators of $\mu$ and $\Pi$. The test also depends on $\overline{\Omega}_n$, a consistent estimator of the asymptotic variance of $(\overline{\mu}'_n,\textup{vec}(\overline{\Pi}_n)')'$, where $\textup{vec}(A)$ denotes the vectorization of a matrix, $A$.
We start with preliminary $H_0$-restricted estimators for $\mu$ and $\delta$:
where $\widehat\Upsilon_n$ is a preliminary weight matrix that is converging in probability to a deterministic positive definite limit.\footnote{This is similar to the weight matrix used in the first step of two-step GMM; the limit theory is invariant to the choice of weight matrix used. We take $\widehat{\Upsilon}_n$ to be the identity matrix.} While $\widetilde\mu_n$ is always unique, $\widetilde\delta_n$ need not be. In that case, we can take $\widetilde\delta_n$ to be the minimizer that has the smallest norm:
Equation ((ref)) is a tie-breaking procedure that is only used if the value of $\widetilde\delta_n$ that minimizes ((ref)) is not unique.\footnote{Equation ((ref)) uses the Euclidean norm, although it could be replaced with any other norm.} Later, we add assumptions so that $\widetilde\delta_n$ is consistent for its population analogue, $\delta^\ast_F$.\footnote{We formally define $\delta^\ast_F$ in ((ref)), below. For now, it is enough to think of $\delta^\ast_F$ as the probability limit of $\widetilde\delta_n$.}
We next estimate the asymptotic variance of $\overline{\mu}_n+\overline{\Pi}_n\delta^\ast_F$ using
where $\otimes$ denotes the Kronecker product. Define the QLR test statistic as
Let $(\widehat\mu_n,\widehat\delta_n)$ solve the minimization problem in ((ref)). Similar to the initial estimators, $\widehat\mu_n$ is always unique and $\widehat\delta_n$ may not be unique. In that case, we can take $\widehat\delta_n$ to be any minimizer.\footnote{Note the subtlety in the definitions of $\widetilde\delta_n$ and $\widehat\delta_n$: $\widetilde\delta_n$ is required to minimize ((ref)) because it has to be consistent for $\delta^\ast_F$, while $\widehat\delta_n$ can be arbitrary because its consistency is not essential.}
For any $K\subseteq\{1,...,d_C\}$, let $|K|$ denote the cardinality of $K$, and let $I_K$ denote the submatrix of the $d_C\times d_C$ identity matrix formed by taking the rows corresponding to the indices in $K$. In this way, conformable premultiplication of $I_K$ to a matrix $B$ selects the rows of $B$ corresponding to the indices in $K$ to form the submatrix of $B$ with $|K|$ rows.
We are ready to define the DoF and the critical value. Let $\widehat b_n=d-B\widehat\mu_n$ and $\widehat{K}= \{j\in\{1,\dots,d_C\}: e'_j\overline{C}_n\widehat\delta_n= e'_j\widehat b_n\}$, where $e_j$ is the $j$th standard normal basis vector. Note that $\widehat{K}$ denotes the set of indices at which the inequality constraint holds as equality for the minimizers in ((ref)). Let
For a significance level $\alpha$, let $cv(s,\alpha)$ denote the $1-\alpha$ quantile of the $\chi^2$ distribution with DoF equal to $s$. The GCC test rejects if $T_n$ is greater than $cv(\widehat s_n,\alpha)$.
We end this section with some remarks on the definition of the GCC test.
\noindentRemarks: {\it (1) Equation ((ref)) gives an algebraic formula for the DoF. In CS23, the DoF, $\widehat{r}_n$, is defined as the dimension of the span of a polyhedral cone. Theorem (ref), below, shows that $\widehat{r}_n$ is equal to $\widehat s_n$ with probability one in the limit. For computational reasons, we recommend $\widehat s_n$ over $\widehat r_n$.
(2) The GCC test is in some sense a “na\"ive” test. It is equivalent to first selecting the inequalities that are active according to the finite sample CQPP in} ((ref)){\it, pretending that these inequalities are equalities, and then forming a Sargan-Hansen's $J$-test for overidentification of a model defined by these equalities. To see this, note that the typical case is when $\textup{rk}(I_{\widehat K}[B,D])=|\widehat K|$ and $\textup{rk}(I_{\widehat K}C)=d_\delta$. Then, the DoF used by the GCC test is $|\widehat K|-d_\delta$, which represents the number of active inequalities minus the number of nuisance parameters.
(3) The weight matrix, $\widetilde{\Sigma}_n$, used in the definition of the GCC test is an estimator of the asymptotic variance of $\sqrt{n}(\overline{\mu}_n+ \overline\Pi_{n}\delta^\ast_F)$ instead of that of $\sqrt{n}\overline{\mu}_n $. The latter is used in CS23 for the sCC test. The new weight matrix accounts for the estimation error in $\overline{\Pi}_n$. To gain some intuition for this weight matrix, note that $T_n$ can be rewritten as
where $\gamma = \sqrt{n}(\delta-\delta^\ast_F)$, $\eta = \sqrt{n}(\mu-\mu_F+(\overline{\Pi}_n-\Pi_F)\delta_F^\ast)$, $X_n = \sqrt{n}(\overline{\mu}_n - \mu_F +(\overline{\Pi}_n-\Pi_F)\delta_F^\ast)$, and $h_n = \sqrt{n}(d-C_F\delta_F^\ast-B\mu_F)$, where $\mu_F$, $\Pi_F$, and $C_F$ stand for the true values of $\mu$, $\Pi$, and $C$. With this change of variables, one can see that $\widetilde{\Sigma}_n$ estimates the asymptotic variance of $X_n$.
(4) The GCC test is very easy to compute since it only requires solving two CQPPs. Efficient interior-point algorithms for CQPPs are available in most commonly used software, and they are known to have a worst-case computational complexity of $O((d_C+d_\delta)^4)$, where $d_C$ is the number of inequalities and $d_\delta$ is the dimension of the nuisance parameter. This is only slightly slower than the computational complexity of LPPs, which is $O((d_C+d_\delta)^{3.5})$; see Karmarkar1984 and YeTse1989. Most importantly, no simulation or bootstrap is needed to perform the test.}
In this section, we present the three main theoretical results of the paper: (1) a theorem that justifies using $\widehat s_n$ and hence simplifies the rank calculation, (2) a theorem that shows the consistency of $\widetilde\delta_n$ for $\delta^\ast_F$, and (3) a theorem that shows the uniform asymptotic validity of the GCC test.
We now present the result that justifies using $\widehat s_n$. One can also define the DoF in the GCC test using the Karush-Kuhn-Tucker (KKT) multipliers. Let $\widehat{L}=\{j\in\{1,...,d_C\}: e'_j\widehat\psi_n>0\}$, where $\widehat\psi_n$ is a vector of nonnegative multipliers that satisfy the KKT conditions for ((ref)). Then
is another way to define the DoF. Also note that
is the definition of the DoF for the sCC test in CS23, where $\textup{dim}(\cdot)$ denotes the dimension of a set, or the maximum number of linearly independent elements. The following theorem shows that $\widehat s_n$, $\widehat t_n$, and $\widehat r_n$ are equal with probability one in the limit. This theorem plays a vital role in the proof of the uniform asymptotic validity of the GCC test, below.
\noindentRemarks: (1) {\it Theorem (ref) justifies the use of $\widehat s_n$ as the DoF. This overrides CS23, which recommends calculating $\widehat r_n$ using an algorithm that includes a series of LPPs. The new recommendation applies to both the sCC test in CS23 and to the GCC test.}
(2) {\it Theorem (ref) is not random---it does not rely on the distribution of $\overline{\mu}_n$, $\overline{C}_n$, or $\widetilde\Sigma_n$. The result is a general feature of CQPPs. Part (a) shows that, regardless of the distribution, $\widehat s_n$ is (weakly) more conservative than $\widehat r_n$. Part (b) shows that equality holds with probability one if the conditional distribution of $\overline{\mu}_n$ given $\overline{C}_n$ and $\widetilde\Sigma_n$ is absolutely continuous. A key case where this holds is in the limit, where $\overline{C}_n$ and $\widetilde\Sigma_n$ are deterministic and $\overline{\mu}_n$ is Gaussian. Thus, a simple corollary of Theorem (ref) is that $\widehat r_n=\widehat s_n$ with probability one in the limit. }
(3) {\it The expression for $\mathcal{M}_0$ can be found in the Supplemental Appendix, equation ((ref)). The value of $\mathcal{M}_0$ may depend on the value of $\overline{C}_n$ or $\widetilde\Sigma_n$ that is fixed. Note that $\widehat{\delta}_n$ (and $\widehat{K}$) may not be unique for the definition of $\widehat s_n$, and $\widehat\psi_n$ (and $\widehat{L}$) may not be unique for the definition of $\widehat t_n$. When they are not unique, $\mathcal{M}_0$ does not depend on the choice of $\widehat{\delta}_n$ or $\widehat \psi_n$. }
(4) {\it In general, $\mathcal{M}_0$ is not the empty set. To clarify the necessity of $\mathcal{M}_0$ in part (b), we give a simple example to show that $\widehat r_n<\widehat s_n$ is possible, albeit on a set of measure zero. Suppose $d_\mu=d_\delta=1$ and $d_C=2$. Let $B=(0,1)'$, $\overline{C}_n=D=(1,1)'$, and $d=(0,0)'$. If $\overline{\mu}_n=0$, then $\widehat\mu_n=\widehat\delta_n=0$ solves ((ref)) (for any positive scalar $\widetilde\Sigma_n$). From ((ref)), $\widehat s_n=1$ because $\widehat K=\{1,2\}$, but from ((ref)), $\widehat r_n=0$. This example requires degenerate features that make $\widehat r_n<\widehat s_n$ unlikely to occur in practice.}
Before describing the assumptions and theorems, we first clarify the true values of the parameters and the underlying distribution of the estimators. Let $F$ denote the joint distribution of $\overline{\mu}_n$, $\overline{C}_n$, and $\overline{\Omega}_n$, and let $P_F(\cdot)$ denote probabilities taken with respect to $F$. Let $\mathcal{F}_n$ be a parameter space for $F$.\footnote{We subscript $\mathcal{F}_n$ with $n$ because $\overline{\mu}_n$, $\overline{C}_n$, and $\overline{\Omega}_n$ are typically functions of a sample, $\{W_i\}_{i=1}^n$, with sample size $n$. Thus, their distribution naturally depends on $n$. We allow $\mathcal{F}_n$ to depend arbitrarily on $n$.} Let $\mu_F$ and $\Pi_{F}$ denote the true values of $\mu$ and $\Pi$. Also let $C_F=B\Pi_F+D$ and $b_F=d_n-B\mu_F$. The $F$ in the subscript makes explicit that these quantities depend on $F$. We allow the values of these parameters, together with the value of $d$, to depend on $n$ to incorporate the situation where we test a sequence of null hypotheses.\footnote{Testing a sequence of null hypotheses is required to evaluate the uniform coverage probability of a confidence set for a parameter of interest, $\theta$. Then, the inequalities that define the null hypothesis may depend on the hypothesized value of $\theta$.} We make explicit the dependence of $d$ on $n$, denoting it by $d_n$. For notational simplicity, we keep the dependence of $\mu_F$, $\Pi_F$, $C_F$, and $b_F$ on $n$ implicit.
Let $\mathcal{F}_{n0}$ be the subset of $\mathcal{F}_n$ that satisfies the null hypothesis: $\mathcal{F}_{n0}=\{F\in\mathcal{F}_n: C_F\delta\le b_F \text{ for some }\delta\in\mathbb{R}^{d_\delta}\}$. For $F\in\mathcal{F}_{n0}$, let $\delta^\ast_{F}$ be the value of $\delta$ that satisfies $C_F\delta\le b_F$. If there is more than one such value of $\delta$, we take $\delta^\ast_{F}$ to be the one that has minimum norm:
This mimics the definition of $\widetilde\delta_n$.
The following assumption ensures asymptotic normality of the estimators of the reduced-form parameters and consistent estimation of the asymptotic variance, at least along a subsequence of true data generating processes. It is used to show consistency of $\widetilde\delta_n$ and asymptotic uniform validity of the GCC test.
{Remarks:} (1) {\it Assumption (ref) is stated using subsequences in order to ensure uniformity over $\mathcal{F}_{n0}$. Part (i) assumes that $d_n$ and the sequence of true parameter values for the reduced-form parameters converge to some limits along a subsequence. This is equivalent to assuming that the parameter space for these parameters is compact.
(2) Part (ii) assumes asymptotic normality of the estimators for the reduced-form parameters along a subsequence. Part (iii) requires $\Sigma_\infty$ to be positive definite. While this may seem restrictive, it is mitigated by the way the inequalities are specified in equation ((ref)). Specifically, the $B$ matrix allows us to write the inequalities as a linear function of a core collection of reduced-form parameters that admit an estimator with a positive definite asymptotic variance matrix. The matrix $B$ absorbs any linear dependence among the estimation errors of the inequalities. Also note that $\delta^\ast_\infty$ is well-defined by Assumption (ref)(vi). Part (iv) assumes consistency of the estimator of the asymptotic variance. Part \textup{(v)} assumes consistency of the first-step weight matrix in equation \textup{((ref))} for a positive definite limit. Parts \textup{(ii)}, \textup{(iv)}, and \textup{(v)} can be verified using standard consistency and asymptotic normality arguments. For example, when the data are i.i.d.\ and the model is a moment (in)equality model, they can be verified by the Lindeberg-Feller central limit theorem and a law of large numbers for triangular arrays.
(3) Part (vi) assumes the constraint set for the limit is nonempty. Part (vi) is guaranteed under part (i) if, for example, there is a fixed compact set $\Delta$ such that $\{\delta\in \mathbb{R}^{d_\delta}: C_{F}\delta\leq b_{F}\}\subseteq\Delta$ for all $F\in{\cal F}_{n0}$.\footnote{This type of assumption is common in the literature. For example, it is assumed by Voronin2025 and GoffMbakop2025.} In that case, for any sequence $F_{n_q}\in{\cal F}_{n_q0}$, there exists a $\delta_{F_{n_q}}$ such that $B\mu_{F_{n_q}}+(B\Pi_{F_{n_q}}+D)\delta_{F_{n_q}}\leq d_{n_q}$. This sequence has a subsequence that converges to some limit $\delta_\infty$ that satisfies $B\mu_{\infty}+(B\Pi_\infty+D)\delta_\infty\leq d_\infty$, showing part (vi). Appendix (ref) shows another way to verify part (vi) under a strengthened version of Assumption \textup{(ref)}, below.}
Next, we state the stable rank condition mentioned in the introduction. We first introduce some new notation. Let $K^=\subseteq \{1,...,d_C\}$ be a set that contains the indices for the inequalities that were originally equalities. This set is special because it is always included in the set of active inequalities: $K^=\subseteq\widehat K$. For any $d_C\times d_\delta$ matrix $C$ and for any $d_C$-dimensional vector $b$, let $\mathcal{A}(C, b)=\{K\subseteq\{1,...,d_C\}: Cx\le b \text{ and } I_{K}(C x-b) = \mathbf{0} \text{ for some }x\in\mathbb{R}^{d_\delta}\}$ be the collection of all subsets of inequalities that could be simultaneously active for the system of inequalities defined by $C$ and $b$.\footnote{A combination of inequalities cannot be simultaneously active if, for example, it involves an upper and a lower bound that are parallel and separated. Such combinations are excluded from $\mathcal{A}(C, b)$.}
\noindentRemarks: (1) {\it We refer to Assumption (ref) as a “stable rank” condition because perturbations of the $C_\infty$ matrix in the directions of the estimation error do not change the rank. }
(2) {\it Assumption (ref) is not a necessary condition for the validity of the GCC test. {It is possible to relax {\it Assumption (ref) by reducing the number of collections of indices, $K$, for which the rank equality in ((ref)) needs to be assumed. In Appendix (ref), we show that the rank equality only needs to hold for index sets, $K$ corresponding to collections of inequalities that define a linear subspace in the limit. In cases where there are no equalities and the identified set for the nuisance parameters has a positive volume in the parameter space, every inequality could be slack, and no stable rank condition is needed.} Due to the nuances of this discussion, it is relegated to the Supplemental Appendix.}}
(3) {\it While Assumption (ref) is not necessary, it is used in an essential way in the proofs of consistency of $\widetilde{\delta}_n$ and asymptotic validity of the GCC test. Moreover, a more than superficial connection of Assumption (ref) with the weak IV problem in linear IV regression models suggests that relaxing Assumption (ref) completely may require insights from that literature. We discuss this connection in Section (ref), below.}
(4) {\it Assumptions playing a similar role as Assumption (ref) are common in the literature on subvector inference in moment inequality models and in models defined by linear systems. One type of such assumptions is a known and fixed $C$, as in GuggenbergerHahnKim2008, KaidoSantos2014, and FSST2023. In that case $\overline{C}_n=C_F= C$, and Assumption (ref) holds trivially. Other types of assumptions appear in PPHI2015, BugniCanayShi2017, ChoRussell2024, and GoffMbakop2025.\footnote{Assumption (ref) differs from constraint qualification, as considered in KMS2022. Constraint qualification restricts a fixed collection of constraints to ensure KKT conditions are necessary or sufficient or the KKT multipliers are unique. In contrast, Assumption (ref) concerns a sequence of linear inequality constraints and restricts the way they converge to a limiting set of constraints.} We discuss the connection between these assumptions in a simple example in Section (ref)}.
The following theorem states an important preliminary result: consistency of $\widetilde{\delta}_n$.
\noindentRemark: Consistency of estimators defined by tie-breaking procedures, such as the norm minimization in the definition of $\widetilde{\delta}_n$, is especially challenging. It is surprising that Assumption (ref) is sufficient in this case. The proof uses a novel argument that establishes setwise convergence of the constraint set.
The following theorem is the main theoretical result of the paper.
{Remarks:} (1) {\it Theorem (ref) establishes the uniform asymptotic validity of the GCC test. This extends the result for the sCC test from CS23 to allow $C$ to be estimated, as long as the estimator is consistent and asymptotically normal and a stable rank condition is satisfied. The generalization is essential for handling the applications discussed in Section} (ref).
(2) {\it The asymptotic validity of the GCC test is surprising because, intuitively, the active inequalities are not necessarily binding in population and even when all inequalities are binding, the limit distribution of $T_n$ is not $\chi^2$. Indeed, the set of active inequalities does not converge to the set of binding-in-population inequalities but remains random in the limit. The key to validity of the GCC test is that the limit {\em conditional} distribution of $T_n$ given the set of active inequalities is bounded by the $\chi^2$ distribution with the associated DoF. }
(3) {\it When $P_F(\widehat{s}_n = 0)>0$, the GCC test can be slightly conservative: Its null rejection probability is between $\alpha (1-P_F(\widehat{s}_n = 0))$ and $\alpha$ asymptotically. The refinement in CS23 can be used to remove the conservativeness. We define the refined GCC (RGCC) test in Appendix} (ref){\it. The RGCC test differs from the GCC test only when $\widehat{s}_n=1$ and is also tuning parameter and simulation free. However, the refinement requires calculating $A$ and $g$ such that $\{\mu\in\mathbb{R}^{d_\mu}:B\mu+\overline{C}_n \delta\leq d \text{ for some } \delta\in\mathbb{R}^{d_\delta}\} = \{\mu\in\mathbb{R}^{d_\mu}:A \mu\leq g\}$. The computation is relatively easy when $d_C$ and $d_\delta$ are small but gets exponentially harder when $d_C$ and $d_\delta$ increase. In particular, it can have a high memory requirement. We investigate the performance of the RGCC test along with the GCC test in the simulations and the empirical illustration.}
To better understand Assumption (ref), we now relate it to the linear IV regression model.
In this model, Assumption (ref) allows $\mathbb{E}[ZX_{-1}']$ to change with $n$ as long as the rank does not change in the limit. Equivalently, the smallest nonzero eigenvalue of $\mathbb{E}[ZX_{-1}']\mathbb{E}[X_{-1}Z']$ does not converge to zero. Notably, zero eigenvalues are allowed. This happens, for example, when the number of instruments is smaller than the number of nuisance parameters. Assumption (ref) is weaker than the usual rank condition for strong identification of $\delta$ (under a hypothesis that fixes a value of $\theta$), which is that the smallest eigenvalue of $\mathbb{E}[ZX_{-1}']\mathbb{E}[X_{-1}Z']$ is bounded away from zero. This means that Assumption (ref) can be thought of as a “no weak identification” condition, where linear combinations of $\delta$ can be strongly identified or non-identified as long as they are not weakly identified.
Even in the weak instruments/weak identification literature, strong identification of the nuisance parameters is a useful assumption. For example, StockWright2000, Kleibergen2005, and AndrewsMikusheva2016 propose identification-robust hypothesis tests for subvectors only when the nuisance parameters are strongly identified under the null. Papers that cover inference with weakly identified nuisance parameters, including ChaudhuriZivot2011, Andrews2018, and GKM2024, recommend some version of two-step inference requiring a tuning parameter, among other complications.\footnote{An exception is GKMC2012. They focus on a homoskedastic linear IV model and show that the plug-in Anderson-Rubin test remains valid with weakly identified nuisance parameters.} Cox2022 states separate limit theory depending on whether the nuisance parameters are strongly identified under the null. In this literature, strong identification of the nuisance parameters under the null is used to guarantee that the null-imposed estimator of the nuisance parameters is consistent and asymptotically normal. In contrast, we use Assumption (ref) to ensure convergence of the constraint set in the QLR statistic to its limit in a setwise sense. These different purposes reflect the compounding complications that arise when trying to relax Assumption (ref).
This section evaluates the finite-sample performance of the GCC test for testing inequalities that are linear in nuisance parameters. We also evaluate the RGCC test, defined in Appendix (ref). Section (ref) considers a simple one-sided hypothesis testing problem with one nuisance parameter. Section (ref) considers a more realistic model with more nuisance parameters: an interval outcome IV regression model. Also, a simulation of Example (ref) can be found in Section (ref). The bottom line of all the simulations is that the GCC and RGCC tests are easy to compute and have good size and power.
Consider a simple one-sided hypothesis testing problem with one nuisance parameter and normally distributed randomness. The model is designed to abstract from computational and asymptotic complications, so as to focus on size and power. The validity of any test for linear inequalities in this simple specification should be a necessary condition for implementing the test in practice. Also note that, because the bound on the parameter of interest is one-sided, we can compare to methods from the literature on one-sided inference for the value of a LPP.
The simple model has one nuisance parameter, no equalities, and $J$ inequalities:
This model is a special case of ((ref)) with $B=I_J$, $\mu = (\mu_1,\mu_2,...,\mu_J)'$, $\Pi = C=(c_1,c_2,...,c_J)'$, $D = \mathbf{0}_J$, and $d = -(1, 1, 0,...,0)'\theta$. The first two inequalities give an upper bound on $\theta$, while the remaining $J-2$ inequalities only bound $\delta$.
Suppose $\mu$ is estimated by $\overline{\mu}_n$ and $C$ by $\overline{C}_n$. Suppose $\overline{\mu}_n$ and $\overline{C}_n$ are sample means of independent random samples from $N(\mu,I_J)$ and $N(C,2I_J)$, respectively. The covariance matrix of $\overline{\mu}_n$ and $\overline{C}_n$ are estimated by the sample variances and covariances. Below, we consider $\mu = (-1,1,1-qn^{-1/2},...,1-qn^{-1/2})'$ and $C = (1,-1,-1,...,-1)'$ with $n=500$, $J\in\{3,10,50\}$, and $q\in\{0,4\}$. Inequalities 3 through $J$ may be binding or slack depending on the value of $q$. The identified set for $\theta$ is $(-\infty,0]$. For a fixed a value of $\theta$ in $(-\infty,0]$, the identified set for $\delta$ is $[\max(1+\theta,1-qn^{-1/2}), 1-\theta]$.
We first consider testing the null hypothesis $H_0: \theta=0$. Table (ref) reports the simulated rejection probabilities for various tests when $q=0$. This means that all $J$ inequalities are binding. We implement four groups of tests. The first group consists of the GCC and RGCC tests. The second group consists of the sCC and sRCC tests from CS23 and the hybrid test (ARP) from AndrewsRothPakes2023. These tests are implemented using $\overline{C}_n$ as if it were the true value. The third group consists of the MR test (BCS) from BugniCanayShi2017, procedure 3 (CCT) from CCT2018, and the recommended test (Bei) from Bei2023, which are designed for nonlinear inequalities. The fourth group consists of the recommended test (FSST) from FSST2023, the recommended test (Gaf) from Gafarov2025, and three tests (CR1-CR3) from ChoRussell2024 implemented with three choices of the tuning parameter.\footnote{CR1, CR2, and CR3 are implemented with $\underline{\epsilon}=0.1$, $\underline{\epsilon}=0.01$, and $\underline{\epsilon}=0.001$, respectively.} These tests are from the literature on inference for the value of a LPP. They are implemented for one-sided inference on the upper bound of $\theta$.\footnote{Gaf and CR1-CR3 report confidence intervals instead of hypothesis tests. For these methods, we say that a value of $\theta$ is rejected if $\theta$ does not belong to the confidence interval.}
As we can see in Table (ref):
(1) {The GCC and RGCC tests are valid for every $J$. The GCC test is somewhat conservative when $J=3$, which is expected. Unexpected is that the GCC and RGCC tests are somewhat conservative when $J=10$. This could be due to simulation noise}.
(2) {The sCC, sRCC, and ARP tests are invalid. This demonstrates the need to account for the estimation error in $\overline{C}_n$.\footnote{Another way to implement these tests is to consider the case that the inequalities in ((ref)) hold conditionally on $\overline{C}_n$. Then, the sCC, sRCC, and ARP tests can be implemented with the conditional variance of $\overline{\mu}_n$ given $\overline{C}_n$. Since $\overline{\mu}_n$ is independent of $\overline{C}_n$, the conditional variance is the same as the unconditional variance and the implementation is the same. Thus, Table (ref) shows that neither way to implement these tests is valid. The problem is that the inequalities in ((ref)) do not hold conditionally on $\overline{C}_n$.} }
(3) {Concerning the tests in the third group, CCT is invalid for $J>3$. It is only valid under a high-level condition that is not satisfied in this model; see Assumption 4.7 in CCT2018. BCS is valid and becomes quite conservative when $J=50$. Bei is a more computationally feasible version of BCS. Table (ref) suggests Bei has mild to moderate over-rejection. It is unclear why the null rejection probabilities for BCS and Bei are so different. }
(4) {Concerning the tests in the fourth group, FSST is understandably invalid because it requires known Jacobian. Surprisingly, Gaf is also invalid for $J>3$. This could be because a rank condition is not satisfied; see Assumption 2 in Gafarov2025. CR1-CR3 appear to be very sensitive to the choice of the tuning parameter. CR1 is very conservative while CR2 and CR3 have moderate over-rejection when $J=50$. }
(5) {Computationally, the GCC and RGCC tests are by far the fastest among the tests considered. Comparing their times to the sCC and sRCC tests demonstrates the computational improvement of the new algebraic formula for the DoF. For CR1-CR3 and Gaf, the computational time is for the calculation of a confidence interval and thus is not comparable to the computation time for a single hypothesis. }
\addtocounter{figure}{-1}
We also consider the power functions of some of the tests as a function of the hypothesized value of $\theta$. In addition to the GCC and RGCC tests, we include the BCS and Bei tests as benchmarks and the CR1 and CR2 tests to see whether the sensitivity of their NRPs to the tuning parameter carries over to the power functions. These tests are included because they are the ones that are valid or exhibit only moderate over-rejection in Table (ref).
We test the inequalities in ((ref)) for a grid of values of $\theta$ between $-1/\sqrt{n}$ and $6/\sqrt{n}$. Dividing the grid by $\sqrt{n}$ ensures the power function approaches the asymptotic local power function as $n\rightarrow\infty$. For $\theta\le 0$ (the shaded region), the rejection probabilities are under the null and thus should be less than or equal to $\alpha=5\%$. For $\theta>0$, the rejection probabilities represent the power of the tests. Figure (ref) reports the power curves for $q\in\{0,4\}$ and $J\in\{3,10,50\}$. When $q=0$, the 3rd to the $J$th inequalities are binding at $\delta=1$, and when $q=4$, these inequalities are slack at $\delta=1$.
As we can see in Figure (ref):
(1) {The GCC and RGCC tests are very fast with good power for all $J\in\{3,10,50\}$ and $q\in\{0,4\}$. The power is especially impressive when $q=4$, so $J-2$ of the inequalities are slack. Also note that, when the number of inequalities increases, the difference between the GCC and RGCC power curves decreases, especially when all the inequalities bind.}
(2) {The BCS test has good power for $J=3$, but becomes more conservative, and therefore less powerful, for larger $J$. The Bei test is more powerful than BCS, GCC, and RGCC, with especially high power when $q=0$ and $J\in\{10,50\}$. These specifications correspond to the cases of moderate over-rejection of the Bei test at $\theta=0$.}
(3) {The sensitivity that the CR tests show in Table (ref) is reflected in power. CR1 has low power, and the power of CR2 is closely related to the over-rejection of the test at $\theta=0$.}
This subsection considers the interval outcome IV regression model from Example (ref). We simulate the power curves of the GCC and RGCC tests. We also include the BCS and Bei tests as benchmarks. Since this is a two-sided problem, we do not include recommendations from the literature on inference for the value of a LPP.
The model is based on the aggregate demand model considered in GLS23, where the market shares are noisy measures of conditional choice probabilities and may contain zero values. The model boils down to the interval outcome IV regression in Example (ref). Write out $X=(X_1, X_2, W')'$, where $X_1$ is a scalar endogenous regressor, $X_2$ is a scalar exogenous regressor, and $W$ is a $d_W$-vector of additional exogenous controls. Similarly, write out $\beta=(\theta_1, \theta_2, \gamma')'$ with $\gamma\in\mathbb{R}^{d_W}$. Also let $Z_e$ be an excluded exogenous instrument. We take $Z=\mathcal{I}(X_2, W, Z_e)$ to be a non-negative vector-valued instrumental function.
In this model, there is a latent market share that satisfies a logit specification:
where $\theta_{10}$ and $\theta_{20}$ denote the true values of $\theta_1$ and $\theta_2$ and $\epsilon$ is an error term. (The true value of $\gamma$ is zero.) In the framework of GLS23, $s^\ast$ is the (unobserved) conditional probability that a representative consumer buys a product, and $Y^\ast=\log(s^\ast)-\log(1-s^\ast)$ is the mean utility of the product. We observe a market share $s_N\sim Binomial(N,s^\ast)/N$, where $N$ is the number of participants in the market. GLS23 argue that $Y^U = \log(s_N+2/N)-\log(1-s_N+0.00125)$ and $Y^L = \log(s_N+0.00125)-\log(1-s_N+2/N)$ satisfy $\mathbb{E}[Y^L|Z]\leq \mathbb{E}[Y^\ast|Z]\leq \mathbb{E}[Y^U|Z]$, which justifies the use of the model in Example (ref).
We take $X_2$, $Z_e$, and the components of $W$ to be independent $Bernoulli(0.5)$ random variables, except for the first element of $W$, which is taken to be the constant one. We take $Z=\mathcal{I}(X_2,W,Z_e)$ to be the vector of indicators for each point in the support of $(X_2,W,Z_e)$. When $d_W=1,2,3$, the dimension of $Z$ is $4,8,16$, respectively. Also, the fact that there are two inequalities in ((ref)) means that there are $8$, $16$, and $32$ inequalities, respectively. Independently of $Z$, let $\epsilon\sim \min\left(\max\left(N(0,1),-4\right),4\right)$. Then, let $X_1=\mathds{1}\{Z_e+\epsilon/2>0\}$ be the endogenous regressor. We calculate $s^\ast$ according to ((ref)) with $\theta_{10}=\theta_{20}=-1$. We simulate $s_N$ with $N=100$ independently from $Z$ and $\epsilon$. This specifies the data generating process for all the observed variables: $s_N$, $X_1$, $X_2$, $W$, and $Z_e$. We simulate a sample of size $n$ from this model for $n\in\{500, 1000\}$ and calculate confidence intervals for $\theta_2$ treating $\delta=(\theta_1,\gamma')'$ as nuisance parameters.\footnote{The same model is considered in Section 5.2 of CS23. However, CS23 construct confidence intervals for $\theta_1$, treating $(\theta_2,\gamma')'$ as nuisance parameters. Since $X_2$ and $W$ are both exogenous, it is valid to conduct inference conditional on $(X_2,W,Z_e)$ using the sCC test in CS23 because the Jacobian of the sample moments with respect to $(\theta_2,\gamma')'$ is known given the sample for $(X_2,W,Z_e)$. In contrast, this paper takes $(\theta_1,\gamma')'$ to be the nuisance parameters. Then, the Jacobian with respect to $\theta_1$ is not known even after conditioning on $(X_2,W,Z_e)$. Thus, the sCC test in CS23 is invalid.}
Figure (ref) plots simulated power curves for the GCC, RGCC, BCS, and Bei tests. In each graph, the horizontal axis represents the value of $\theta_2$, while the vertical axis represents the rejection probability.\footnote{The rejection probabilities reported are frequencies that each $\theta_2$ value lies outside the confidence interval for $\theta_2$. The confidence intervals are computed using a bisection algorithm for the endpoints.} The shaded region indicates the identified set for $\theta_2$. Since the control variables have zero coefficients and are independent of the other random variables, they do not affect the identified set of $\theta_2$. In the legend, each number in the square brackets is the median computational time (in seconds) to compute one confidence interval. We make the following remark on Figure (ref).
\addtocounter{figure}{-1}
\noindentRemark: The power curves of the GCC, RGCC, BCS, and Bei tests are remarkably similar. All four have rejection probabilities below the nominal level 5% in the shaded region. All four have increasing power as $\theta_2$ deviates from its identified set and as the sample size increases from 500 to 1000. Overall, the GCC and RGCC tests are able to match the size and power performance of the BCS and Bei tests while being much faster computationally.
This section demonstrates the GCC test in the model of female labor supply considered in KT16. KT16 use a model of state transition probabilities with states that indicate (1) the earnings of the individual (zero or not employed, positive and below the federal poverty line, or above the federal poverty line), (2) whether the individual participates in welfare, and (3) whether the individual under-reports her earnings in order to qualify for welfare. KT16 use data from the Manpower Development Research Corporation (MDRC) Jobs First study, which is a random experiment that assigned women with young children to one of two welfare programs: the Aid to Families with Dependent Children (AFDC) welfare program or the Jobs First Temporary Family Assistance (JF) program.\footnote{Instructions for accessing the datasets and the replication codes provided by KT16 can be found on the AEA webpage: \url{https://www.aeaweb.org/articles?id=10.1257/aer.20130824}.}
Relative to AFDC, the JF program primarily changes how eligibility for and the amount of government transfers respond to earnings.\footnote{The JF welfare reform introduces other changes too, such as stricter work requirements and changed administration of the Food Stamps program. We focus on the changes that are relevant for bounding the state transition probabilities. For more information, see the description in KT16.} Under JF, for individuals earning under the federal poverty line (FPL), the government transfer does not decrease as earnings increase: everyone under the FPL receives the same government transfer. This is more generous than the policy under AFDC, which had government transfers decrease after earnings reached a threshold. This more generous policy may induce unemployed women (or women who would be unemployed under AFDC) to gain employment that pays below the FPL. This represents a labor supply response along the extensive margin. Another noteworthy feature of the JF program is that the government transfers abruptly drop to zero when earnings cross the FPL. This may incentivize some women who would be employed with earnings above the FPL (under AFDC) to decrease their earnings (or under-report their earnings) in order to be eligible for welfare. This represents a labor supply response along the intensive margin. The goal in KT16 is to distinguish these two responses without imposing strong assumptions on the utility functions of the individuals.
KT16 distinguish 7 labor supply/welfare participation states:
In the label for each state, the number indicates the level of earnings: “0” indicates zero earnings or not employed, “1” indicates positive earnings below the FPL, and “2” indicates earnings above the FPL. The letter indicates welfare participation and under-reporting of earnings: “n” indicates nonparticipation in welfare, “r” indicates welfare participation with truthful reporting of earnings, and “u” indicates welfare participation with under-reporting of earnings.
Each individual is associated with two states, the one she would choose under AFDC and the one she would choose under JF. Because AFDC is the status quo, we label the state transition probabilities with an individual's choice under AFDC first. For example, $\pi_{\text{0n},\text{1r}}$ is the conditional probability that an individual who chooses 0n under AFDC would choose 1r under JF. Because of the features of the JF reform, KT16 argue that only nine state transition probabilities need to be considered. The first group are flows out of 0r: $(\pi_{\text{0r},\text{0n}}, \pi_{\text{0r},\text{1n}}, \pi_{\text{0r},\text{2n}}, \pi_{\text{0r},\text{1r}}, \pi_{\text{0r},\text{2u}})$. Individuals with zero earnings and participating in welfare under AFDC may transition to any other state except 1u.\footnote{Under JF, no one chooses 1u because everyone earning under the FPL receives the same government transfer, so there is no reason to under-report earnings. Also note that $\pi_{\text{0r},\text{0r}}$ is not needed because it can be calculated as one minus the others. This is true in general. We do not need to include the transition probabilities from one state into itself because they are determined by the other transition probabilities.} The second group are flows into 1r: $(\pi_{\text{0n},\text{1r}}, \pi_{\text{1n},\text{1r}}, \pi_{\text{2n},\text{1r}}, \pi_{\text{0r},\text{1r}}, \pi_{\text{2u},\text{1r}})$. Individuals who choose 1r under JF may have chosen any state under AFDC.\footnote{Individuals who choose 1u under AFDC are guaranteed to choose 1r under JF, so $\pi_{\text{1u},\text{1r}}=1$ and there is no need to include it as a free parameter.} KT16 argue that all other transition probabilities can be set to zero either because the budget sets for the individuals is unchanged between the two policies or because the combination of choices would violate weak assumptions on individuals' utility functions. Let \[ \delta = (\pi_{\text{0n},\text{1r}}, \pi_{\text{1n},\text{1r}}, \pi_{\text{2n},\text{1r}}, \pi_{\text{0r},\text{0n}}, \pi_{\text{0r},\text{1n}}, \pi_{\text{0r},\text{2n}}, \pi_{\text{0r},\text{1r}}, \pi_{\text{0r},\text{2u}}, \pi_{\text{2u},\text{1r}})' \] collect the transition probabilities into a vector of nuisance parameters. (Note that one of the transition probabilities is common to both groups.)
There is an additional problem: it is unobserved whether an individual under-reports her income. Thus, the marginal probabilities of only six observable states---combining 1r and 1u---is identified. KT16 show that after accounting for this problem, the resulting equalities are still linear in $\delta$ and they can be rewritten into five non-redundant equalities of the form $\Gamma\delta = m$, where
where $p_{\text{s}}^{\text{A}}$ denotes the marginal probability that the individual chooses state s under the AFDC policy for $s\in\{\text{0n}, \text{1n}, \text{2n}, \text{0r}, \text{2u}\}$, and similarly for $p_{\text{s}}^{\text{J}}$ for the JF policy. In addition, $\delta \in [0,1]^9$ and $\pi_{\text{0r},\text{0n}}+\pi_{\text{0r},\text{1n}}+\pi_{\text{0r},\text{2n}}+\pi_{\text{0r},\text{1r}}+\pi_{\text{0r},\text{2u}}\leq 1$, which can be imposed by an appropriate choice of $A$ and $b$ in ((ref)).\footnote{In Section (ref), we give a simple sufficient condition for Assumption (ref) in this model.}
We follow KT16 and report confidence intervals for each transition probability. In addition to the GCC test, we also implement the RGCC test, defined in Section (ref). The GCC and RGCC tests are implemented by rewriting the restrictions in terms of $B$, $\mu$, $\Pi$, $D$, and $d$ using ((ref)). We can similarly define estimators of $\mu$ and $\Pi$ from the sample averages that estimate $p_s^t$ for $s\in\{\text{0n}, \text{1n}, \text{2n}, \text{0r}, \text{2u}\}$ and $t\in\{\text{A},\text{J}\}$.\footnote{We copy KT16 and use weighted sample averages with propensity score weights to adjust for baseline differences.} We estimate the asymptotic variance by bootstrapping the sample averages with a cluster bootstrap, clustered at the case level and with 1000 bootstrap draws.\footnote{This is the same implementation of the bootstrap that KT16 use, except that we bootstrap the estimators of $p_s^t$ while they bootstrap the formulas for the bounds that they calculate after eliminating the nuisance parameters manually. This means that our variance estimators, while asymptotically equivalent, are numerically different.} The endpoints of the GCC and RGCC confidence intervals are calculated using a bisection algorithm.
KT16 manually eliminate the nuisance parameters and work out explicit formulas for the bounds of each element of $\delta$. To avoid the dependence on tuning parameters that is prevalent in the literature on testing inequalities, they report two confidence intervals. One confidence interval, called Na\"ive, is constructed by ignoring the uncertainty in which bounds bind. This interval is formed by the single lowest upper (highest lower) estimated bound plus (minus) its standard error. The asymptotic coverage probability of this interval is unknown. The other interval, called Conservative, assumes all population bounds bind simultaneously, leading to asymptotically valid but often overly conservative inference.
Table (ref) reports the confidence intervals for each transition probability.\footnote{The values for the Na\"ive and Conservative confidence intervals are slightly different from the ones in the published version of KT16. We calculated these values using the KT16 replication code without changes. The differences are likely due to differences in random number generation across versions of Stata when implementing the bootstrap.} As Table (ref) shows:
(1) { All the confidence intervals are qualitatively similar. They provide evidence for the same heterogeneous labor supply responses: statistically significant outflows from state 0r, corresponding to an increase in labor supply along the extensive margin, and statistically significant inflows into state 1r, especially from state 2n, corresponding to a decrease in labor supply along the intensive margin as women decrease their earnings to qualify for welfare. }
(2) { One would expect the endpoints of the GCC and RGCC confidence intervals to lie between the endpoints of the Na\"{i}ve and Conservative confidence intervals. That is mostly true, but there are a few noteworthy exceptions. } (a) { For $\pi_{\text{0r},\text{1n}}$, the GCC and RGCC confidence intervals are narrower than the Na\"{i}ve confidence interval. In this case, the two smallest upper bounds are very close to each other. Together, they provide stronger statistical evidence than just the one that is active. The GCC and RGCC tests respond to this statistical evidence in a way that the Na\"{i}ve confidence interval does not. } (b) {For $\pi_{\text{1n},\text{1r}}$ and $\pi_{\text{2n},\text{1r}}$, the GCC and RGCC confidence intervals are wider than the Conservative confidence intervals. This is not surprising for the GCC confidence interval because the GCC test is conservative when there is one binding inequality.\footnote{For all the transition probabilities in Table (ref), there is only one nontrivial lower bound. This explains both why the Na\"{i}ve lower bounds are equal to the Conservative lower bounds and why the GCC confidence interval appears conservative for the lower bounds.} This is surprising for the RGCC confidence interval because with one binding inequality, the RGCC test is asymptotically equivalent to the optimal one-sided test. This is likely due to the numerical difference between the variance matrix estimators used in the RGCC and the Conservative confidence intervals.}
(3) {To compute all nine confidence intervals, the GCC and RGCC methods took about 4 seconds and 220 seconds, respectively. Both approaches are quite feasible, especially considering the fact that manual elimination of the nuisance parameters is not needed to implement the GCC and RGCC tests. }
This paper proposes a simple, tuning-parameter-free test that is designed for inequality testing problems that are linear in nuisance parameters, including specification testing and subvector inference in moment (in)equality models and inference for parameters bounded by linear programs. We prove asymptotic uniform validity of the test under a stable rank condition and demonstrate its size, power, and computational performance in simulations and an empirical illustration.