EconBase
← Back to paper

Inference for Moment Inequalities: A Constrained Moment Selection Procedure

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

86,061 characters

Inference for Moment Inequalities: A Constrained Moment Selection Procedure


\maketitle
\begin{abstract}
Inference in models where the parameter is defined by moment inequalities is of interest in many areas of economics. This paper develops a new method for improving the performance of generalized moment selection (GMS) testing procedures in finite-samples. The method modifies GMS tests by tilting the empirical distribution in its moment selection step by an amount that maximizes the empirical likelihood subject to the restrictions of the null hypothesis. We characterize sets of population distributions on which a modified GMS test is (i) asymptotically equivalent to its non-modified version to first-order, and (ii) superior to its non-modified version according to local power when the sample size is large enough. An important feature of the proposed modification is that it remains computationally feasible even when the number of moment inequalities is large. We report simulation results that show the modified tests control size well, and have markedly improved local power over their non-modified counterparts.
\end{abstract}
\small
\noindent \textbf{\emph{Keywords}}: empirical likelihood, moment inequality model, statistical information. \\
\noindent \textbf{\emph{JEL Classification}}: C12, C14, C21
\newpage
\section{Introduction}
Statistical inference in models defined by moment inequalities is a frequently encountered topic in econometrics. Examples of applications include games of entry with multiple equilibria (e.g.,~\citealp{ciliberto2009market}), single/multiple agent optimization problems (e.g.,~\citealp{pakes2015moment}), censored and missing data (e.g.,~ \citealp{manski2002inference,imbens2004confidence}), model selection tests (e.g.,~\citealp{Shi-KL}, and~\citealp{Shi-Hsu}), event-study designs (e.g.,~\citealp{rambachanhonest}), stochastic dominance comparisons (e.g. \citealp{whang_2019}) and New-Keynesian DSGE models (e.g., \citealp{moon2009estimation}). This paper considers inference for a finite-dimensional parameter defined by a finite number of unconditional moment inequalities.

\par We suppose that there exists a true value of the parameter $\theta_0\in\Theta\subseteq\mathbb{R}^{d}$ that satisfies the moment inequality restrictions
\begin{align}\label{eq - moment inequality}
E_{F_{0}}\big(g_{j}(W_i,\theta_{0})\big) \geq 0\quad\text{for}\quad j=1,...,J,
\end{align}
where $\{g_{j}(\cdot,\theta): j=1,...,J\}$ are known real-valued functions, $\{W_i:i\leq n\}$ are independent and identically distributed (i.i.d.) with unknown distribution $F_0,$ and $W_i\in\mathbb{R}^{\text{dim}(W_i)}.$  Under these moment conditions,  the set $\Theta_{I}(F_0)\equiv\left\{\theta\in\Theta: E_{F_{0}}\big(g_{j}(W_i,\theta)\big) \geq 0\;\forall j=1,...,J\right\}$
denotes the so-called \emph{identified set} while any $\theta\in\Theta_{I}(F_0)$ is termed an \emph{identifiable parameter}. Thus, the true value of the parameter might not be uniquely identified by $F_0$ and the economic model.

\par We are interested in confidence sets for $\theta_0$ constructed by test inversion. The test is based on a statistic $T_n,$ for testing individual hypotheses for each $\theta$ that have the form
\begin{align}\label{eq - Test Problem}
H_0:\,\theta\in\Theta_{I}(F_0)\quad\text{versus}\quad H_1:\,\theta\notin\Theta_{I}(F_0).
\end{align}
Inference in this model is challenging because the pointwise limiting null distribution of conventional test statistics are discontinuous in the parameter -- the dependence on the parameter is through the index set of moment inequalities~(\ref{eq - moment inequality}) that are binding. In particular, a moment inequality enters the pointwise asymptotic null distribution of the test statistic $T_n$ whenever it holds as an equality. Tests of~(\ref{eq - Test Problem}) that have good properties incorporate information about which moments $E_{F_{0}}\big(g_{j}(W_i,\theta)\big)$ are ``positive'', in order to exclude them from the computation of a critical value. Tests of this sort are known as two-step procedures in the literature, examples of which include~\cite{andrews2010inference},~\cite{canay2010inference},~\cite{andrews2012inference}, and~\cite{romano2014practical}. The first step of those testing procedures use the data to determine whether the moment inequalities~(\ref{eq - moment inequality}) are close to or far from being equalities. The second step uses the outcome of the first step to yield information about which moment inequalities are ``positive'' when constructing tests of~(\ref{eq - Test Problem}).

\par The literature on two-step tests of~(\ref{eq - Test Problem}) is vast, and almost all of these tests use the sample-analogue estimator of the moments $E_{F_{0}}\big(g_{j}(W_i,\theta)\big)$ in the first step to determine the slackness of the moment inequalities. This feature ignores the information present in the restrictions~(\ref{eq - moment inequality}), because the sample-analogue estimator does not exploit the fact that the moments satisfy these restrictions under the null hypothesis in~(\ref{eq - Test Problem}). Thus, we conjecture that implementing this information in such tests can improve their accuracy in finite-samples under the null and alternative hypotheses. This paper provides such a modification for the broad class of \emph{generalized moment selection} (GMS) testing procedures put forward by \cite{andrews2010inference}, and finds that our conjecture is in the right direction.

\par We propose a modification of GMS testing procedures that implements the information present in~(\ref{eq - moment inequality}) using the method of empirical likelihood (\citealp{owen2001empirical}). The modification is to replace the sample-analogue estimator of the moments in the first step of the GMS procedure with its constrained empirical likelihood counterpart, where the constraints are the moment inequalities~(\ref{eq - moment inequality}). We label this modification \emph{constrained moment selection} (CMS). For a given test statistic and moment selection function, the CMS and GMS tests only differ in terms of which moments they select for the computation of the critical value in tests of~(\ref{eq - Test Problem}). The motivation for our proposal is that the detection of the ``positive'' moment inequalities in the first step would be more accurate because we are using additional information that is available to us, which the sample-analogue estimator of the moments ignores. Consequently, the CMS procedure alters the GMS critical value for testing~(\ref{eq - Test Problem}) in a data-dependent way that incorporates the information contained in~(\ref{eq - moment inequality}) through a reduction of the parameter space for $F_0.$ For this reason, we expect CMS tests of~(\ref{eq - Test Problem}) to be more accurate than their GMS counterparts in finite-samples.

\par This paper characterises the parameter space for $(\theta_0,F_0)$ over which the CMS and GMS testing procedures are asymptotically equivalent, to first-order, under the null, local alternatives, and distant alternatives. We focus, though, on the GMS class of testing procedures in which the moment selection function is given by the moment selection $t$-test. This focus is without loss of generality, as the results extend naturally, with appropriate modifications, to the more general setup in~\cite{andrews2010inference} using their assumptions. This means that for a given test statistic, CMS tests inherit all of the asymptotic properties of GMS tests. Specifically, under the null, CMS confidence sets are asymptotically valid with uniformity over the parameter space, not asymptotically conservative, and not asymptotically similar. Furthermore, CMS tests of~(\ref{eq - Test Problem}) have greater asymptotic local power than tests based on subsampling or fixed asymptotic critical values, and are consistent against distant alternatives. The parameter space imposes only three conditions in addition to the conditions that define the parameter space~\cite{andrews2010inference} introduce. These conditions are part of Assumption~GEL in~\cite{andrews2009validity}: (i) a uniform bound on the variances of the moment functions, (ii) a lower bound on the determinant of their correlation matrix, and (iii) a regularity condition on an estimator of the degree of slackness of the moments arising from the dual formulation of the constrained empirical likelihood problem. Collectively, the conditions that define our parameter space enables the use of results from~\cite{andrews2009validity} on constrained empirical likelihood estimation in our proofs of the aforementioned asymptotic results.

\par While GMS and CMS are asymptotically equivalent procedures, we characterise local alternatives under which the power of CMS tests dominate their GMS counterparts for sufficiently large, but finite, samples. These are directions in the alternative that have some non-violated moment inequalities (SNVIs) and a non-negative correlational structure. That is, configurations where some of the moments $E_{F_{0}}\big(g_{j}(W_i,\theta_{0})\big)$ under the alternative hypothesis are ``positive", and the covariance matrix of $\{g_{j}(W_i,\theta): j=1,...,J\}$ has non-negative entries only. The non-negative correlational structure arises in empirical applications; see, for example,~\cite{Lok-Tabri-inpress} who point to that structure for moment inequalities characterising stochastic dominance comparisons. It is quite difficult to determine the extent of this difference in local powers analytically. However, using a Monte Carlo simulation experimental design based on~\cite{andrews2012inference}, who focus on finite-sample comparisons of the maximum null rejection probability (MNRP), we show using the modified method of moments (MMM) statistic that along such local alternatives the differences in MNRP-corrected powers of CMS and GMS tests can be approximately 36 percentage points when $J=4$ and $n=250,$ which is strikingly large. See Section~\ref{Setion - MC Sims} for more details.

\par The two-step tests in this literature that exploit the information~(\ref{eq - moment inequality}) are the procedures put forward by~\cite{andrews2009validity} and~\cite{canay2010inference}. They implement this information using (generalised) empirical likelihood.~\cite{andrews2009validity} and~\cite{canay2010inference} develop subsampling and bootstrap tests of~(\ref{eq - Test Problem}), respectively, using empirical-likelihood-type test statistics. Both tests have correct asymptotic size in a uniform sense and are shown not to be asymptotically conservative. However,~\cite{canay2010inference}'s test has higher asymptotic power because it is a GMS procedure. More generally,~\cite{andrews2010inference} show the asymptotic power of GMS tests dominate that of subsampling and plug-in asymptotic tests. A disadvantage of~\cite{canay2010inference}'s procedure is that it may be more computationally burdensome than other GMS tests. Thus, our modification of GMS tests can improve finite-sample performance without incurring a high computational cost.

\par \cite{andrews2012inference} proposed a refinement of GMS termed \emph{refined moment selection} (RMS) and discussed the reasons why such an approach is preferable. However, the RMS procedure is quite computationally expensive when $J>10.$ By contrast, the CMS procedure remains computationally feasible when $J$ is large. The reason is that the constrained empirical likelihood optimization problem it is based upon has a strictly concave objective function, convex feasible set, and the choice variables enter linearly into the constraints. As a consequence, there is a unique global solution to this optimization problem and its implementation involves an of-the-shelf programming routine. More recently,~\cite{romano2014practical} proposed a two-step testing procedure for moment inequalities that is similar in spirit to the RMS procedure and remains computationally feasible when $J$ is large. An important distinction between the CMS testing procedure and these tests is that, like GMS tests, neither of them exploits the information present in the moment inequality constraints~(\ref{eq - moment inequality}), because they employ the sample-analogue estimator of the moments in their first step.

\par We examine the finite-sample performance of CMS tests using the MMM and adjusted quasi-likelihood-ratio (AQLR) test statistics in Monte Carlo simulations based on the experimental design in~\cite{andrews2012inference}. The experiment compares the performance of CMS to its GMS, RMS, RSW counterparts in terms of MNRP and MNRP-corrected local power. The inclusion of the RMS and RSW procedures in the simulation experiment is to benchmark the performance of CMS. Overall, the simulation results showcase the value of implementing the information~(\ref{eq - moment inequality}) in the CMS procedure in terms of finite-sample size and power properties, and corroborate its theoretical superior performance over GMS. The simulation results also show the performance of CMS and RMS tests based on the AQLR statistic are comparable. This finding is encouraging as the RMS test has desirable asymptotic properties but can be computationally expensive when $J$ is large, while the CMS procedure isn't costly to compute at all.

\par The idea of exploiting information on parameters defined by constraints for improving performance in statistical problems, through constrained estimation, is one of the most natural ideas in statistics. The literature on constrained estimation via tilting the empirical distribution overlaps with this paper, where the problem is that the constraints/information are not adequately reflected by the empirical distribution (e.g.,~\citealp{Hall-Presnell}). Tilting the empirical distribution allows one to incorporate information selectively into a statistical procedure without changing the procedure itself.~\cite{Lok-Tabri-inpress}
apply this idea to modifying two-step bootstrap tests for restricted stochastic dominance orderings using empirical likelihood and semi-infinite programming. The parameter of interest in their setup is infinite-dimensional and there is a continuum of moment inequality restrictions, which are defined by moment functions that have a particular form. The form of the moment functions in their setup yields a correlational structure that facilitates the analysis of such moment inequalities. Contrastingly, in the setup of this paper, $J$ is finite and the form of the moment functions $\{g_{j}(\cdot,\theta): j=1,...,J\}$ is arbitrary. The implementation of empirical likelihood in their setup has a data-driven number of inequalities that increases with the sample size, which can be as large as 500 in moderate sample sizes. The ability of empirical likelihood to straightforwardly execute with a large number of moment inequality restrictions transfers to the CMS procedure for models with large $J.$ This computational feasibility of CMS is an important feature of our approach. Similar to~\cite{Lok-Tabri-inpress}, this paper is also part of the econometrics literature on shape restrictions (e.g.,~\citealp{Chetverikov-Santos-Shaikh}, and the references therein), as the inequalities~(\ref{eq - moment inequality}) can be thought of as finite-dimensional analogues of shape restrictions on nonparametric functions.


\par We organize the paper as follows. Section~\ref{Section - Steup and Background} introduces the statistical framework, as well as the GMS and CMS procedures. Section~\ref{Section - CMS} introduces the main results of the paper. Section~\ref{Setion - MC Sims} reports the results of Monte Carlo simulations, and Section~\ref{Section - Conclusion} concludes.

\par For notational simplicity, throughout the paper we write partitioned column vectors as $h=(h_1,h_2).$ rather than $h=(h_{1}',h_{2}')'$. Let $\mathbb{R}_{+}=\{x\in\mathbb{R}:x\geq0\}$, $\mathbb{R}_{+,\infty}=\mathbb{R}_{+} \cup\{+\infty\}$,
$\mathbb{R}_{[+\infty]}=\mathbb{R} \cup\{+\infty\}$,$\mathbb{R}_{[\pm\infty]}=\mathbb{R} \cup\{\pm\infty\}$, ``$:=$" denote the definitional identity, and $\overline{A}$ denote the closure of a set $A$.

\section{Setup}\label{Section - Steup and Background}

\subsection{Moment Inequality Model and Test statistic}
The object of interest is a parameter $\theta_{0} \in \Theta \subseteq \mathbb{R}^{d}$, $d<+\infty$, defined by a finite number of known moment functions $g_{j}: \mathcal{W} \times \Theta \rightarrow \mathbb{R}$ that satisfy the following unconditional moment inequality restrictions:
\begin{align}\label{Moment inequality restrictions S3}
E_{F_{0}}\big(g_{j}(W,\theta_{0})\big) \geq 0 \ \forall \ j\in\mathcal{J},
\end{align}
where $F_{0}$ denotes the true distribution of the observed data $W$ and $\mathcal{J}:=\{1,...,J\}$ with $J<\infty$. In general, the identified set, $\Theta_{I}(F_0)=\{\theta \in \Theta: \ E_{F_0}(g_{j}(W,\theta)) \geq 0 \ \forall \ j \in \mathcal{J}\}$, is not a singleton meaning that the parameter is partially identified.

\par The moment inequality model is given by the following definition.
\begin{definition}\label{def1}[\textbf{Moment Inequality Model}] Let $\mathcal{F}$ be the set of parameters $(\theta, F)$ that satisfy:
\begin{enumerate}
\item $\theta \in \Theta\subseteq \mathbb{R}^{d}$. \label{A1}
\item $\{W_{i}: i \geq 1\}$ are i.i.d. under $F$. \label{A2}
\item $E_{F}\big(g_{j}(W_{i},\theta)\big) \geq 0 \ \text{for} \ j\in\mathcal{J}$. \label{A3}
\item $\sigma^{2}_{F,j}(\theta) := Var_{F}\big(g_{j}(W_{i},\theta)\big) \in [\varepsilon_{*}, M_{*}]$ for some $M_{*}>\varepsilon_{*}>0$. \label{A4}
\item $\Omega(\theta,F) \in \varPsi_{2}$, where $\Omega(\theta,F)$ is the $J\times J$ correlation matrix of $\{g_{j}(W_{i},\theta),j=1,\ldots,J\},$ and $\varPsi_{2}$ is the space of correlation matrices whose determinant is greater than $\varepsilon>0$. \label{A5}
\item $\exists \delta$ and $M >0:$ $E_{F}|g_{j}(W_{i},\theta)/\sigma_{F,j}(\theta)|^{2+\delta} \leq M \ \forall \ j\in\mathcal{J}$. \label{A6}
\end{enumerate}
\end{definition}
\noindent All of the conditions in this definition, except for Conditions~\ref{A4} and~\ref{A5}, are those presented in (2.2) of~\citet{andrews2010inference}. Condition~\ref{A4} is a strengthening of Condition (v) in~\citet{andrews2010inference} so that the variances of the moment functions are uniformly bounded. Condition~\ref{A5} specifies the nonsingularity of the matrix $\Omega(\theta,F).$  These conditions are relatively unrestrictive and are part of Assumption GEL in~\citet{andrews2009validity}. Furthermore, they arise frequently in papers that consider empirical likelihood inference for moment inequalities (e.g.,~\citealp{canay2010inference}, and~\citealp{Lok-Tabri-inpress}).

\par For a given value of the parameter, $\theta=\theta_0,$ we invert tests of the hypothesis $H_0:\,\theta_0\in\Theta_{I}(F_0)$ to construct confidence sets of the form $CS_{n}=\{\theta \in \Theta: T_{n}(\theta) \leq c_{1-\alpha}(\theta)\},$ where $T_{n}(\theta)$ denotes a test statistic and $c_{1-\alpha}(\theta)$ is a critical value for tests with nominal level $\alpha \in (0, 1/2)$. We say $CS_{n}$ is a uniformly valid confidence set for $\theta$ if
\begin{align}\label{Uniform Validity}
\liminf_{n\rightarrow+\infty}\inf_{(\theta,F)\in\mathcal{F}}P_{F}(T_{n}(\theta) \leq c_{1-\alpha}(\theta))\geq 1-\alpha,
\end{align}
where $P_{F}(\cdot)$ is the probability measure induced by repeated sampling from $F$. Uniformity is essential in order for asymptotic size to be a good approximation to the finite-sample size of confidence sets, because the test statistic exhibits a discontinuity in its asymptotic distribution (as a function of the distribution generating the data), but not in its finite-sample distribution. Discontinuities of this type can create asymptotic size problems that are analogous to those that arise with parameters that are near a boundary (e.g.,~\citealp{andrews2009validity}).

\par A test statistic is a function $S: \mathbb{R}_{[+\infty]}^{J}\times \mathcal{V}_{J\times J} \rightarrow \mathbb{R}$ given by $T_{n}(\theta_0):=S\left(n^{\frac{1}{2}}\hat{g}_{n}(\theta_0), \hat{\Sigma}_{n}(\theta_0)\right),$ where $\mathcal{V}_{J\times J}$ is the set of invertible $J\times J$ variance matrices,
\begin{align*}
\hat{g}_{n}(\theta_0) &  = \left[\frac{1}{n}\sum_{i=1}^{n}g_{1}(W_{i},\theta_0),...,\frac{1}{n}\sum_{i=1}^{n}g_{J}(W_{i},\theta_0)\right]^{\top},\;g(W_{i},\theta_0)=\big[g_{1}(W_{i},\theta_0),...,g_{J}(W_{i},\theta_0)\big]^{\top},\\
\text{and} & \quad \hat{\Sigma}_{n}(\theta_0)= n^{-1}\sum_{i=1}^{n}(g(W_{i},\theta_0)-\hat{g}_{n}(\theta))(g(W_{i},\theta_0)-\hat{g}_{n}(\theta_0))^{\top}.
\end{align*}
 Two examples are the modified method of moments (MMM) and adjusted quasi-likelihood-ratio (AQLR) statistics. In the context of the moment inequality model $\mathcal{F}$ given by Definition~\ref{def1}, these test statistics are defined as
\begin{align}
S_{1}\big(n^{\frac{1}{2}}\hat{g}_{n}(\theta_0), \hat{\Sigma}_{n}(\theta_0)\big) &  = n\sum_{j=1}^{J}\Big(\min\big\{0, \hat{g}_{n,j}(\theta_0)/\hat{\sigma}_{n,j}(\theta_0)\big\}\Big)^{2}\;\text{and} \label{MMM} \\
S_{2A}\big(n^{\frac{1}{2}}\hat{g}_{n}(\theta_0), \hat{\Sigma}_{n}(\theta_0)\big) & = n\inf_{t \in \mathbb{R}_{+,\infty}^{J}}(\hat{g}_{n}(\theta_0)-t)^{\intercal}\big(\tilde{\Sigma}_{n}(\theta_0)\big)^{-1}(\hat{g}_{n}(\theta_0)-t),\label{QLR}
\end{align}
respectively, where $\tilde{\Sigma}_{n}(\theta_0) =  \hat{\Sigma}_{n}(\theta_0)+\max\{0, 0.012-|\hat{\Omega}_{n}(\theta_0)|\}\hat{D}_{n}(\theta_0),\; \hat{\Omega}_{n}(\theta_{0})= \hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})\hat{\Sigma}_{n}(\theta_{0})\hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})$ and $\hat{D}_{n}(\theta_{0})=\operatorname*{diag} \hat{\Sigma}_{n}(\theta_{0})$, where $\operatorname*{diag} \hat{\Sigma}_{n}(\theta_{0})$ is a diagonal matrix with dimensions equal to those of $\hat{\Sigma}_{n}(\theta_{0})$ whose diagonal elements equal those of $\hat{\Sigma}_{n}(\theta_{0})$.


\subsection{GMS and CMS Procedures}
\par The point of departure for establishing that~(\ref{Uniform Validity}) holds for the GMS procedure is to consider the asymptotic distribution of $T_n(\theta_0)$ under a suitable sequence of null distributions. For any sequence $\{F_n: n\geq1\}$ in the model of the null hypothesis, the test statistic satisfies
\begin{align}\label{eq - Asy GMS}
T_n(\theta_0)\stackrel{d}{\longrightarrow}S\left(\Omega_0^{1/2}Z^{*}+h_1,\Omega_0\right)\quad\text{where}\quad Z^{*}\sim N(0_J,I_J),
\end{align}
where $h_1\in\mathbb{R}^{J}_{+,\infty},$ and $\Omega_0$ is a $J\times J$ correlation matrix.\footnote{Specifically, this large-sample result~(\ref{eq - Asy GMS}) follows from the form of the test statistic, the Central Limit Theorem, and the convergence in probability of the sample correlation matrix.} The vector $h_1=(h_{1,1},...,h_{1,J})'$ has elements given by $\lim_{n\rightarrow+\infty}n^{1/2}(E_{F_n}(g_{j}(W_{i},\theta_0))/ \sigma_{F,j}(\theta_0)$ and measures the degree of slackness of the moment inequalities. The crux of this asymptotic construction is that the limiting distribution in~(\ref{eq - Asy GMS}) now depends continuously on the degree of slackness of the moment inequalities via the parameter $h_1,$ which reflects the finite-sample situation.

\par The asymptotic implementation of the GMS critical value is the $1-\alpha$ quantile of a data-dependent version of the asymptotic null distribution in~(\ref{eq - Asy GMS}). It replaces $\Omega_0$ by a consistent estimator and replaces $h_1$  with a function $\varphi : \mathbb{R}^{J} \times \varPsi_{2} \rightarrow \mathbb{R}^{J}_{+,\infty}$, which measures the slackness of moment inequalities through $\hat{\xi}_{n}(\theta_{0})=\kappa_{n}^{-1}n^{\frac{1}{2}}\hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})\hat{g}_{n}(\theta_{0}),$ where $\{\kappa_{n}: n \geq 1\}$ is a divergent sequence of scalars \citep{andrews2010inference}. The GMS critical value, $\hat{c}_{n}(\theta_{0},1-\alpha),$ is the $1-\alpha$ quantile of
\begin{align}\label{eq - L_n GMS}
L_n(\theta_0,Z^{*})=S\left(\hat{\Omega}_{n}^{\frac{1}{2}}(\theta_{0})Z^{*}+\varphi(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0})), \hat{\Omega}_{n}(\theta_{0})\right),
\end{align}
where $Z^{*}\sim N(0_{J}, I_{J})$ and is independent of $\{W_i: i\geq 1\}.$ That is,
\begin{align}\label{GMScv}
\hat{c}_{n}(\theta_{0},1-\alpha) := \inf\bigg\{x \in \mathbb{R}: P\Big(L_n(\theta,Z^{*})\leq x\Big) \geq 1-\alpha \bigg\}.
\end{align}
where $P\Big(L_n(\theta_0,Z^{*})\leq x\Big)$ denotes the conditional CDF at $x$ of $L_n(\theta_0,Z^{*}),$ conditional upon $(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0})).$ In practice,
the calculation of $\hat{c}_{n}(\theta_{0},1-\alpha)$ is by simulating $L_n(\theta_0,Z^{*})$ using $R$ i.i.d. draws from $Z^{*}\sim N(0_{J}, I_{J})$ and computing the $1-\alpha$ quantile of the empirical CDF from $\{L_n(\theta_0,Z_r^{*}): r =1,...,R\}.$

\par Alternatively, one may compute the GMS critical value using the bootstrap. We briefly describe this approach. Let $\{W_{i}^{*}: i \leq n\}$ be a bootstrap sample drawn from the empirical distribution of the data $\{W_{i}: i \leq n\}$, and define $\hat{g}_{n}^{*}(\theta_{0}) = n^{-1}\sum_{i=1}^{n}g(W_{i}^{*},\theta_{0})$, $\hat{\Sigma}_{n}^{*}(\theta_{0})=n^{-1}\sum_{i=1}^{n}(g(W_{i}^{*},\theta_{0})-\hat{g}_{n}^{*}(\theta_{0}))(g(W_{i}^{*},\theta_{0})-\hat{g}_{n}^{*}(\theta_{0}))^{\intercal},$ $\hat{D}_{n}^{*}(\theta_{0}) = \text{diag} \hat{\Sigma}_{n}^{*}(\theta_{0})$, and $\hat{\Omega}_{n}^{*}(\theta_{0}) = (\hat{D}_{n}^{*}(\theta_{0}))^{-\frac{1}{2}}\hat{\Sigma}_{n}^{*}(\theta_{0})(\hat{D}_{n}^{*}(\theta_{0}))^{-\frac{1}{2}}$. The bootstrap implementation of the GMS procedure replaces $L_{n}(\theta_{0},Z^{*})$ in~(\ref{eq - L_n GMS}) with $$L_{n}(\theta_{0},\{W_{i}^{*}: i \leq n\})=S\left(G_{n}^{*}(\theta_{0})+\varphi(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0})), \hat{\Omega}_{n}^{*}(\theta_{0})\right),$$ where $G_{n}^{*}(\theta_{0}) = n^{\frac{1}{2}}(\hat{D}_{n}^{*}(\theta_{0}))^{-\frac{1}{2}}(\hat{g}_{n}^{*}(\theta_{0})-\hat{g}_{n}(\theta_{0})),$ and defines a critical value analogous to (\ref{GMScv}). In practice, this critical value is the empirical $1-\alpha$ quantile of the bootstrap statistics $\{L_{n}(\theta_{0},\{W_{i,r}^{*}: i \leq n\}):r=1,...,R\}$, where $\{\{W_{i,r}^{*}:i\leq n\}:r=1,...,R\}$ are bootstrap samples drawn from the empirical distribution of the data $\{W_{i}: i \leq n\}$. The asymptotic results of this paper hold for the bootstrap provided that $G_{n}^{*}(\theta_{n,h}) \overset{d}{\rightarrow} \Omega_{0}^{\frac{1}{2}}Z^{*}$, where the convergence is conditional on $\{W_i:i\leq n\}$ for almost every sample path, for all sequences $\{(\theta_{n,h},F_{n,h}): n \geq 1\}$ in $\mathcal{F}$.

\par There are numerous choices for $\varphi$ and $\{\kappa_{n}: n \geq 1\}$. \cite{chernozhukov2007estimation} and~\cite{andrews2010inference} recommend using $\kappa_{n} = (\ln n)^{\frac{1}{2}}$. Another option is to set $\kappa_{n} = (2\ln \ln n)^{\frac{1}{2}}$, which is used in~\cite{canay2010inference}.
Our main results set $\varphi=\varphi^{(1)}$, where
\begin{align}\label{t-test MS function}
\varphi^{(1)}_{j}(\xi, \Omega)=
\begin{cases}
0 &\text{if $\xi_{j} \leq 1$} \\
+\infty &\text{if $\xi_{j} > 1$}
\end{cases}
\end{align}
for each $j \in \mathcal{J}$, and is referred to as the `moment selection $t$-test' because it resembles a $t$-test with deterministic critical value $\kappa_{n}$. The decision reflects the recommendations of~\cite{andrews2012inference}, and is essentially without loss of generality because our results extend to any choice of $\varphi$ that satisfies the assumptions of~\cite{andrews2010inference}. Appendix~\ref{OtherGMS} discusses how to generalize our results to other suitable choices of $\varphi$.

\par The advantage of the GMS procedure is that it asymptotically detects the ``positive'' moments  $E_{F_0}(g_{j}(W_{i},\theta_{0}))$ and excludes them from the computation of the critical value, so as to mimic the discontinuity in the asymptotic null distribution of $T_n(\theta).$ This ability of GMS tests to detect such moments is the source of its improvements over the subsampling and plug-in procedures under the null and alternative hypotheses.

\par Although GMS tests are computationally simple and have desirable asymptotic properties, their performance in finite-samples depends crucially on how well they detect the ``positive'' moments, so as to omit them from the computation of the critical value. Their use of the sample-analogue estimator of the moments for detecting the positive moments does not implement the information embedded in~(\ref{Moment inequality restrictions S3}) and implementing this information appropriately can improve the detection accuracy of ``positive" moments in finite-samples.


\par For a given moment selection function $\varphi,$ the CMS procedure implements the information present in~(\ref{Moment inequality restrictions S3}) through a surgical modification of the GMS procedure. The modification is to replace $\hat{g}_{n}(\theta_0)$ with its constrained empirical likelihood counterpart, where the constraints impose the inequality restrictions~(\ref{Moment inequality restrictions S3}). Specifically, CMS replaces $\hat{\xi}_{n}(\theta_{0})$ with $\acute{\xi}_{n}(\theta_{0}) =\kappa_{n}^{-1}n^{\frac{1}{2}}\hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})\acute{g}_{n}(\theta_{0})$, where $\acute{g}_{n}(\theta_{0}) = \sum_{i=1}^{n}\acute{p}_{i}g(W_{i},\theta_{0})$ and the probabilities $\acute{p}_{1},...,\acute{p}_{n}$ solve
\begin{align}\label{eq - EL problem}
\max_{p_{1},...,p_{n}}\Bigg\{\sum_{i=1}^{n}\ln(p_{i}): \sum_{i=1}^{n}p_{i}g_{j}(W_{i},\theta_{0}) \geq 0 \ \forall \ j\in\mathcal{J}, \ \sum_{i=1}^{n}p_{i}=1, \ p_{i} \geq 0 \ \forall \ i\Bigg\},
\end{align}
and then computes a critical value as described in~(\ref{GMScv}), but replaces $\varphi(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$ with $\varphi(\acute{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$ in~(\ref{eq - L_n GMS}). The CMS modification of GMS can easily be applied to all choices of $\varphi$ and $\{\kappa_{n}: n \geq 1\}$ presented in \citet{andrews2010inference} because it only replaces $\hat{g}_{n}(\theta_{0})$ with $\acute{g}_{n}(\theta_{0}).$ The estimator $\acute{g}_{n,j}(\theta_{0})$ of $E_{F_{0}}\big(g_{j}(W,\theta_{0})\big)>0$ is more accurate than $\hat{g}_{n,j}(\theta_{0})$ because the optimization problem~(\ref{eq - EL problem}), which gives rise to $\acute{g}_{n}(\theta_{0}),$ imposes a correct constraint $E_{F_{0}}\big(g_{j}(W,\theta_{0})\big)\geq0,$ while $\hat{g}_{n}(\theta_{0})$ ignores such information. Thus, when $E_{F_{0}}\big(g_{j}(W,\theta_{0})\big)>0$ (under the null or alternative), the moment selection function based on $\acute{g}_{n,j}$ detects this configuration more reliably than $\varphi_{j}(\hat{\xi}_{n}, \hat{\Omega}_{n}(\theta_{0})),$ and therefore, takes it into account by delivering a critical value that is suitable for the case where this moment inequality is omitted. This feature of CMS leads to it having better finite-sample properties than GMS under the null and alternative hypotheses.

\par The CMS procedure is not computationally expensive because the empirical likelihood optimization problem~(\ref{eq - EL problem}) has a strictly concave objective function and a convex feasible set that is characterised by affine functions of the choice variables \citep{owen2001empirical}. This means that the optimization problem~(\ref{eq - EL problem}) has a unique global solution, and it can be computed numerically using standard optimization routines in software such as Matlab, R, or GAUSS. This computational simplicity of the optimization problem~(\ref{eq - EL problem}) is an important feature of the CMS procedure.

\begin{remark}
\normalfont
One can `fully constrain' the CMS procedure by using restricted estimators of the correlation matrix. In this case, we evaluate $\varphi(\acute{\xi}_{n}^{FC}(\theta_{0}), \acute{\Omega}_{n}(\theta_{0}))$, where
\begin{align*}
\acute{\xi}_{n}^{FC}(\theta_{0}) & =\kappa_{n}^{-1}n^{\frac{1}{2}}\acute{D}_{n}^{-\frac{1}{2}}(\theta_{0})\acute{g}_{n}(\theta_{0}),\quad\acute{\Omega}_{n}(\theta_{0}) = \acute{D}_{n}^{-\frac{1}{2}}(\theta_{0})\acute{\Sigma}_{n}(\theta_{0})\acute{D}_{n}^{-\frac{1}{2}}(\theta_{0}), \\
\acute{\Sigma}_{n}(\theta_{0}) & = \sum_{i=1}^{n}\acute{p}_{i}(g(W_{i},\theta_{0})-\acute{g}_{n}(\theta_{0}))(g(W_{i},\theta_{0})-\acute{g}_{n}(\theta_{0}))^{\top},\;\text{and}\;
\acute{D}_{n}^{-\frac{1}{2}}(\theta_{0})= \operatorname*{diag} \acute{\Sigma}_{n}(\theta_{0}).
\end{align*}
In our simulations not presented in this paper, we find limited practical difference between the $\varphi(\acute{\xi}_{n}^{FC}(\theta_{0}), \acute{\Omega}_{n}(\theta_{0}))$ and $\varphi(\acute{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$. Consequently, the rest of the paper focuses on $\varphi(\acute{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$ because it is simpler to show that there are power advantages over GMS.
\end{remark}

\section{Main Results}\label{Section - CMS}

\par We start by introducing the assumptions that beget the main results of this paper. They are conditions on the test statistic $S$, the moment selection function $\varphi$, and the parameter space $\mathcal{F}$. The assumptions on $S$ we consider are from~\cite{andrews2010inference}, and are stated as Assumptions 1-7 in Appendix~\ref{Section - Test Stat Assump} for ease of exposition. Recall that we set $\varphi=\varphi^{(1)}$ in~(\ref{t-test MS function}), and the main results we present are based on this choice of moment selection function. It should be noted that this choice of $\varphi$ is without loss of generality as one can employ assumptions identical to those in~\cite{andrews2010inference} on $\varphi$ to deduce the same conclusions, because the CMS procedure does not alter the moment selection function in the GMS procedure. See Appendix~\ref{OtherGMS} for the details on other choices of $\varphi.$


\par The first assumption concerns the sequence $\{\kappa_n: n\geq1\}.$
\begin{kappaassumption}
\begin{enumerate*}
\item $\kappa_{n} \rightarrow +\infty$ as $n\rightarrow +\infty$.\label{Assump - Kappa 1}
\item $\kappa_{n}^{-1}n^{\frac{1}{2}} \rightarrow +\infty$ as $n\rightarrow +\infty$. \label{Assump - Kappa 2}
\end{enumerate*}
\end{kappaassumption}
\noindent The conditions in this assumption are not restrictive -- the aforementioned examples of $\{\kappa_{n}: n \geq 1\}$ satisfy them. The `optimal' choice of $\{\kappa_{n}: n \geq 1\}$ is an important question, but the goal of our paper is more modest: to demonstrate how incorporating statistical information can improve finite-sample inference for moment inequalities in a computationally simple way and, for this purpose, our analysis conditions on an arbitrary choice of $\{\kappa_{n}: n \geq 1\}$. For our Monte Carlo experiment (Section \ref{Setion - MC Sims}), we set $\kappa_{n}=(\ln n)^{\frac{1}{2}}$ which is the recommended choice in \cite{chernozhukov2007estimation} and \cite{andrews2010inference}.

\par The next assumption we present is the first part in Part (d) of Assumption GEL in~\cite{andrews2009validity}. It is helpful in establishing that $\acute{g}_{n}(\theta_0)$ is a uniformly consistent estimator of the moments under the null hypothesis $H_0:\theta_0\in\Theta_{I}(F_0).$ To introduce this assumption, for each $t\in \mathbb{R}^{J},$ define $g_{i}(t,\theta) = g(W_{i},\theta) - t.$ The vector $t$ is a nuisance parameter that captures the slackness of the moment inequalities. Using the dual formulation of the empirical likelihood problem~(\ref{eq - EL problem}), the amount of slackness is captured by $\acute{t}_n=\operatorname*{arg\,min}_{t \in \mathbb{R}_{+}^{J}}\sup_{\lambda \in \acute{\Lambda}_n(t,\theta)}n^{-1}\sum_{i=1}^{n}\ln\Big(1-\lambda^{\top}g_{i}(t,\theta)\Big),$ where $\acute{\Lambda}_{n}(t,\theta)=\{\lambda\in\mathbb{R}^{J}:\lambda^{\top}g_{i}(t,\theta)\in Q\,\forall i=1,\ldots,n\},$ $Q$ is an open interval of $\mathbb{R}$ containing $0.$ This reformulation of the empirical likelihood problem~(\ref{eq - EL problem}) is feasible because the linear constraint qualification applies to it. The part of Assumption GEL we include in our setup is a regularity condition concerning the uniform asymptotic behavior of $\acute{t},$ and is stated in terms of the following reparametrization of $\mathcal{F}.$
\begin{definition}\label{def Gamma} Let $\Gamma$ be defined as the set of all $\gamma=(\gamma_1,\gamma_2,\gamma_3)$ such that for some $(\theta, F)\in\mathcal{F}$ where
\begin{enumerate}
\item $\mathcal{F}$ is defined in Definition~\ref{def1}.
\item $\gamma_1=(E_{F}(g_{1}(W_{i},\theta))/ \sigma_{F,1}(\theta),\ldots,E_{F}(g_{J}(W_{i},\theta))/\sigma_{F,J}(\theta)).$
\item $\gamma_2= \big(\theta, \text{vech}_{*}(\Omega(\theta,F))\big),$ where $\text{vech}_{*}(\Omega(\theta,F))$ is the vector of lower off-diagonal elements of $\Omega(\theta,F).$
\item $\gamma_{3} = F.$
\end{enumerate}
\end{definition}
\noindent~\cite{andrews2010inference} indicate that there is a one-to-one mapping from $\gamma$ to $(\theta,F);$ see Appendix~A of their paper for the details.  Denote by $\{\gamma_{n,h}: n \geq 1\}\subset \Gamma$ a sequence of parameters in $\Gamma$ such that $n^{1/2}\gamma_{n,h,1} \rightarrow h_{1}\in \mathbb{R}^{J}_{+,\infty}$ and $\gamma_{n,h,2} \rightarrow h_{2} \in \mathbb{R}_{[\pm\infty]}^{q}$ as $n\rightarrow\infty,$ where $q =\dim(\Theta)+ \dim(\text{vech}_{*}(\Omega(\theta,F))).$ The part of Assumption GEL that we include in our setup is given by the following assumption.
\begin{assumptiont}\label{AT}
For all subsequences $\{w_n\}$ of $\{n\}$  and all sequence $\{\gamma_{w_{n},h}: n \geq 1\}\subset \Gamma$ and corresponding $\{(\theta_{w_{n},h}, F_{w_{n},h}): n \geq 1\}\subset\mathcal{F},$
$$\acute{t}_{w_n}=\operatorname*{arg\,min}_{t \in \mathbb{R}_{+}^{J}}\sup_{\lambda \in \acute{\Lambda}_{w_n}(t,\theta_{w_{n},h})}\frac{1}{n}\sum_{i=1}^{n}\ln\Big(1-\lambda^{\top}g_{i}(t,\theta_{w_{n},h})\Big)$$
exists and satisfies $\sup_{n\geq 1}||\acute{t}_{w_n}||_{\ell^{2}_{J}} \leq K$ with probability approaching 1 as $n\rightarrow+\infty$ for some constant $K<+\infty,$ where $||\cdot||_{\ell^{2}_{J}}$ is the usual Euclidean norm on $\mathbb{R}^{J}$.
\end{assumptiont}

\par We also include an assumption from~\cite{andrews2010inference} for the case in which $\text{Int}(\Theta_{I}(F_0)) \neq \emptyset$ for some data-generating process in the model. It is required to show that when there are no binding moment inequalities, the maximum asymptotic coverage probability is equal to $1.$
\begin{assumptionm}
There exists $(\theta, F) \in \mathcal{F}$ that satisfies $E_{F}\big(g_{j}(W_{i},\theta)\big)>0$ for all $j\in \mathcal{J}$.
\end{assumptionm}

\subsection{Asymptotic Size Results}\label{Subsection Asymptotic Size results}
We now present the first main result of the paper. It mirrors Theorem~1 of~\cite{andrews2010inference} which concerns the asymptotic size of GMS confidence sets. Denote by $\acute{c}_{n}(\theta_0,1-\alpha)$ the CMS critical value under the nominal level $1-\alpha$ for testing the null hypothesis $H_0:\theta_0\in\Theta_{I}(F_0).$
\begin{theorem}\label{asymptoticsize}
Suppose $S$ satisfies Assumptions~\ref{Test1} -~\ref{T2}, $\varphi=\varphi^{(1)}$ in~(\ref{t-test MS function}), the sequence $\{\kappa_n:n\geq1\}$ satisfies Part 1 of Assumption~K, and $\alpha \in (0,1/2)$. Furthermore, let $\mathcal{F}_{+}=\{(\theta, F) \in \mathcal{F}: \text{Assumption T holds}\},$ and $\mathcal{F}_{++}=\{(\theta, F) \in \mathcal{F}: \text{Assumptions T and M hold}\}.$ Then, the nominal level $(1-\alpha)$ CMS confidence set based on the statistic $T_n(\theta)$ satisfies the following statements:
\begin{enumerate}
\item $\liminf_{n\rightarrow+\infty}\inf_{(\theta, F) \in \mathcal{F}_+} P_{F}\Big(T_{n}(\theta) \leq \acute{c}_{n}(\theta,1-\alpha)\Big)\geq1-\alpha.$ \label{result1}
\item $\liminf_{n\rightarrow+\infty}\inf_{(\theta, F) \in \mathcal{F}_+} P_{F}\Big(T_{n}(\theta) \leq \acute{c}_{n}(\theta,1-\alpha)\Big)=1-\alpha$, if in addition $S$ and $\{\kappa_n:n\geq1\}$ satisfy Assumption~\ref{Assump test stat 7} and Part 2 of Assumption~K, respectively. \label{result2}
\item $\limsup_{n\rightarrow+\infty}\sup_{(\theta, F) \in \mathcal{F}_{++}}P_{F}\Big(T_{n}(\theta) \leq \acute{c}_{n}(\theta, 1-\alpha)\Big)=1.$  \label{result3}
\end{enumerate}
\begin{proof}
See Appendix~\ref{Appendix - proof Thm 1}.
\end{proof}
\end{theorem}
The first result of Theorem~\ref{asymptoticsize} establishes the uniform validity of CMS confidence sets over the parameter space $\mathcal{F}_{+},$ and the second result of this theorem shows that they are not asymptotically conservative. The third result shows the maximum coverage probability of CMS confidence sets is equal to 1 over the parameter space $\mathcal{F}_{++}.$ The parameter space $\mathcal{F}_+$ is a subset of the one used in Theorem~1 of~\cite{andrews2010inference}, because it imposes Assumption~T and Condition 4 in Definition~\ref{def1} in addition to the conditions they set for their parameter space.

\par The proof of Theorem~\ref{asymptoticsize} establishes that CMS and GMS procedures are asymptotically equivalent with uniformity over the parameter space $\mathcal{F}_+.$ The essence of this result is that for every sequence $\{\gamma_{n,h}: n \geq 1\}\subset \Gamma$ such that Assumption~T holds, along the corresponding sequence $\{(\theta_{w_{n},h}, F_{w_{n}}): n \geq 1\}\subset\mathcal{F}_+$ we have $\acute{\xi}_{w_{n}}(\theta_{w_{n},h})=\hat{\xi}_{w_{n}}(\theta_{w_{n},h})+o_{p}(1).$ This asymptotic equivalence is a consequence of
$\acute{g}_{w_{n}}(\theta_{w_{n},h})-\hat{g}_{w_{n}}(\theta_{w_{n},h})=O_{p}(w_{n}^{-\frac{1}{2}})$ (see Lemma~\ref{L2CS} in Appendix~\ref{Section - Technical Lemmas for Thm 1}) and Assumption K on $\kappa_{w_{n}}.$
In particular, these arguments are used after re-writing the expression of $\acute{\xi}_{w_{n}}(\theta_{w_{n},h})$
in terms of $\hat{\xi}_{w_{n}}(\theta_{w_{n},h}),$ as such
\begin{equation}
\begin{aligned}
\acute{\xi}_{w_{n}}(\theta_{w_{n},h}) & =\kappa_{w_{n}}^{-1}w_{n}^{\frac{1}{2}}\hat{D}_{w_{n}}^{-\frac{1}{2}}(\theta_{w_{n},h})\acute{g}_{w_{n}}(\theta_{w_{n},h})\\
& =\kappa_{w_{n}}^{-1}w_{n}^{\frac{1}{2}}\hat{D}_{w_{n}}^{-\frac{1}{2}}(\theta_{w_{n},h})\big(\acute{g}_{w_{n}}(\theta_{w_{n},h})-\hat{g}_{w_{n}}(\theta_{w_{n},h})\big) + \hat{\xi}_{w_{n}}(\theta_{w_{n},h}),
\end{aligned}
\label{eq - Asy equiv null seq}
\end{equation}
to obtain the asymptotic equivalence.
\subsection{Limiting Local Power Function of CMS Tests}\label{Subsection CMS Local Power Funcion}
This section employs the setup in Section 8 of~\cite{andrews2010inference} to show the limiting local power function of the CMS tests coincide with their GMS counterparts when the null parameter space is $\mathcal{F}_{+}=\{(\theta, F) \in \mathcal{F}: \text{Assumption T holds}\}.$ For sequences of parameters $\{(\theta_{n,*}, F_{n}): n \geq 1\}$, consider the testing problem
\begin{align}\label{LAhyp}
H_{0}: E_{F_{n}}\big(g_{j}(W_{i},\theta_{n,*})\big) \geq 0 \ \forall \ j \in \mathcal{J} \ \text{vs.} \ H_{1}: \text{$H_{0}$ is false}
\end{align}
where $\theta_{n,*}=\theta_{n}+\eta n^{-\frac{1}{2}}(1+o(1))$ for all $n \geq 1$, $(\theta_{n},F_{n}) \in \mathcal{F}_+$ for all $n \geq 1$, and $\eta \in \mathbb{R}^{d}$ where $d=\dim(\theta_{n})<+\infty$. The idea is to study the behavior of the testing procedure along sequences of parameters $\{(\theta_{n,*}, F_{n}): n \geq 1)\}$ that differ locally from a point in the true parameter space $\mathcal{F}_+$ by $O(n^{-\frac{1}{2}})$. The local power function is defined as $P_{F_{n}}\big(T_{n}(\theta_{n,*}) > \acute{c}_{n}(\theta_{n,*},1-\alpha)\big)$, where $P_{F_{n}}(\cdot)$ is the probability measure induced by random sampling from $F_{n}$ for all $n \geq 1$. The objective is to derive an expression for the limiting local power function, $\lim_{n\rightarrow+\infty}P_{F_{n}}\big(T_{n}(\theta_{n,*})> \acute{c}_{n}(\theta_{n,*},1-\alpha)\big),$ and compare it to its GMS counterpart.

\par To this end, we introduce technical assumptions for deriving the limiting local power function for CMS tests. These assumptions are from Section~8 of~\cite{andrews2010inference}.
\begin{lassumption}\label{LAssumption1}
The true parameters $\{(\theta_{n}, F_{n}): n \geq 1\}$ satisfy:
\begin{enumerate}
\item $\theta_{n}=\theta_{n,*}-\eta n^{-\frac{1}{2}}(1+o(1))$ for some $\eta \in \mathbb{R}^{d}$, $\theta_{n,*} \rightarrow \theta_{0}$ and $F_{n}\rightarrow F_{0}$ as $n\rightarrow+\infty$, where $(\theta_{0},F_{0})\in\mathcal{F}_+$.
\item For each $j\in \mathcal{J}$, there exists $h_{1,j} \in \mathbb{R}_{+,\infty}$ such that $n^{\frac{1}{2}}E_{F_{n}}\big(g_{j}(W_{i},\theta_{n})\big)/\sigma_{F_{n},j}(\theta_{n}) \rightarrow h_{1,j}$ as $n\rightarrow+\infty$.
\item $\sup\big\{E_{F_{n}}|g_{j}(W_{i},\theta_{n,*})/\sigma_{F_{n},j}(\theta_{n,*})|^{2+\delta}: n \geq 1\big\}<+\infty$ for all $j \in \mathcal{J}$ for some $\delta>0.$
\end{enumerate}
\end{lassumption}
\noindent The first two parts of Assumptions~LA1 show that the sequence of true parameters, $\{\theta_{n}: n \geq 1\}$, is $n^{-\frac{1}{2}}$-local to the sequence of parameters under the null hypothesis, $\{\theta_{n,*}: n\geq 1\}$ and provides the limit of the sequence of normalised moment functions when evaluated at the sequence of true parameters $\{\theta_{n}: n\geq 1\}$. The third part of this assumption is a uniform integrability condition that permits the use of stochastic limit theorems for triangular arrays of row-wise IID random variables. The second assumption is as follows.
\begin{lassumption}\label{LAssumption2}
$\Pi(\theta, F) := (\partial/\partial \theta^{\top})[D^{-\frac{1}{2}}(\theta, F)E_{F}(g(W_{i}, \theta))] \in \mathbb{R}^{J\times \dim(\Theta)}$ exists and is a continuous function in a neighbourhood of $(\theta_{0}, F_{0})$.
\end{lassumption}
\noindent Both Assumptions LA\ref{LAssumption1} and LA\ref{LAssumption2} are important for proving the large sample properties of CMS tests under $n^{-\frac{1}{2}}$-local alternatives. Namely, they allow one to mean value expand the normalised moment functions under $H_{0}$ around $\theta=\theta_{n}$ and show that
\begin{align*}
\lim_{n\rightarrow \infty}n^{\frac{1}{2}}D^{-\frac{1}{2}}(\theta_{n,*}, F)E_{F}(g(W_{i}, \theta_{n,*}))=h_{1}+\Pi(\theta_{0},F_{0})\eta
\end{align*}
which can then be used to show that $T_{n}(\theta_{n,*}) \overset{d}{\rightarrow} J_{h_{1}, \eta},$ where $J_{h_{1}, \eta}$ is the distribution function of $S\big(\Omega_{0}^{\frac{1}{2}}Z^{*}+h_{1}+\Pi(\theta_{0},F_{0})\eta, \Omega_{0}\big)$ and $Z^{*}\sim N(0_{J}, I_{J})$ \citep{andrews2010inference}.
\begin{lassumption}\label{LAssumption3}
$\lim_{n\rightarrow+\infty}\kappa_{n}^{-1}n^{\frac{1}{2}}D^{-\frac{1}{2}}(\theta_{n},F_{n})E_{F_{n}}(g(W_{i},\theta_{n}))=\pi_{1} \in \mathbb{R}^{J}_{+,\infty}$.
\end{lassumption}
\noindent The last assumption involves the set
$C(\varphi)=\{\tilde{\pi}_{1}\in \mathbb{R}^{J}_{[+\infty]}: \forall j \in \mathcal{J},\;\text{either}\;\tilde{\pi}_{1,j}=+\infty\;\text{or}\;\varphi_{j}(\xi, \Omega) \rightarrow \varphi_{j}(\tilde{\pi}_{1}, \Omega_{0})\;\text{as}\; (\xi, \Omega) \rightarrow (\tilde{\pi}_{1}, \Omega_{0})\}.$ Loosely, $C(\varphi)$ is the set of all vectors in $\mathbb{R}^{J}_{[+\infty]}$ for which $\varphi$ is continuous at $(\tilde{\pi}_{1}, \Omega_{0})$. With $\varphi=\varphi^{(1)},$ this set is $C(\varphi^{(1)})=\{\tilde{\pi}_{1}\in \mathbb{R}^{J}_{[+\infty]}:\tilde{\pi}_{1,j}\neq 1,\forall j\in\mathcal{J}\}.$
\begin{lassumption}\label{LAssumption4}
\begin{enumerate*}
\item $\pi_{1} \in C(\varphi^{(1)})$ and \item $P_{F}\Big(S\big(\Omega_{0}^{\frac{1}{2}}Z^{*}+\varphi^{(1)}(\pi_{1}, \Omega_{0}), \Omega_{0}\big) \leq x\Big)$ is continuous and strictly increasing at $x=c_{\pi_{1}}(\varphi^{(1)} , 1-\alpha)$, the $1-\alpha$ quantile of the distribution function of $S\big(\Omega_{0}^{\frac{1}{2}}Z^{*}+\varphi^{(1)}(\pi_{1}, \Omega_{0}), \Omega_{0}\big)$.
\end{enumerate*}
\end{lassumption}
\noindent Assumptions~LA3 and LA4 are imposed so that we can use Theorem 2(a) of~\cite{andrews2010inference} to obtain the form of the GMS limiting local power function.

\par Next, we present the second main result of this paper. This result states that GMS and CMS tests are asymptotically equivalent, to first-order, under $n^{-1/2}$-local alternatives.
\begin{theorem}\label{LAthm}
Suppose $S$ satisfies Assumptions 1-5, $\varphi=\varphi^{(1)}$ in~(\ref{t-test MS function}), the sequence $\{\kappa_n:n\geq1\}$ satisfies Assumption~K, and that Assumptions~LA1 - LA4, hold. Then $\lim_{n\rightarrow+\infty}P_{F_{n}}\big(T_{n}(\theta_{n,*})>\acute{c}_{n}(\theta_{n,*}, 1-\alpha)\big)=1-J_{h_{1},\eta}\big(c_{\pi_{1}}(\varphi, 1-\alpha)\big).$
\end{theorem}
\begin{proof}
See Appendix~\ref{Appendix - proof Thm 2}.
\end{proof}
\noindent The intuition behind Theorem~\ref{LAthm} is essentially the same as Theorem~\ref{asymptoticsize}. For a given sequence $\{(\theta_{n,*}, F_{n}): \ n \geq 1\}$ of $n^{-1/2}$-local alternatives, we show
$\acute{\xi}_{n}(\theta_{n,*})=\hat{\xi}_{n}(\theta_{n,*})+o_{p}(1)$. This asymptotic equivalence is a consequence of applying
$\acute{g}_{n}(\theta_{n,*})-\hat{g}_{n}(\theta_{n,*})=O_{p}(n^{-\frac{1}{2}})$ (see Lemma~\ref{LP4} in Appendix~\ref{Section Local Pwr Thms 2 3}) and Assumption~K to a decomposition of $\acute{\xi}_{n}(\theta_{n,*})$ identical to~(\ref{eq - Asy equiv null seq}). Therefore, the pairs $(\acute{\xi}_{n}(\theta_{n,*}), \hat{\Omega}_{n}(\theta_{n,*}))$ and $(\hat{\xi}_{n}(\theta_{n,*}), \hat{\Omega}_{n}(\theta_{n,*}))$ are asymptotically equivalent along sequences of $n^{-1/2}$-local alternatives. As this is the only point of difference between CMS and GMS, Theorem~\ref{LAthm} follows from Theorem~2(a) of~\cite{andrews2010inference}. An important corollary to Theorem~\ref{LAthm} is that CMS inherits the first-order improvements that GMS exhibits over subsampling and plug-in asymptotic critical values (see \citealp{andrews2010inference}).

\subsection{Local Power Comparison Between CMS and GMS Tests}\label{Subsection CMS and GMS Local Power comparison}
While Theorem~\ref{LAthm} establishes the equality of the limiting local power functions of CMS and GMS tests under $n^{-1/2}$-local alternatives, this section presents results that characterize sequences of local alternatives under which the power of CMS tests dominate their GMS counterparts for sufficiently large, but finite, samples. First, we must establish when it is meaningful to compare tests along sequences of $n^{-1/2}$-local alternatives. Under the conditions of part 2 of Theorem~\ref{asymptoticsize},
for every $r>0,$ there exists $N_{r,0},N_{r,1}\in\mathbb{Z}_{+}$ (depending on $r$) such that
\begin{align}
\left|\sup_{m\geq n}\sup_{(\theta, F) \in \mathcal{F}_{+}}P_{F}(T_{m}(\theta) > \acute{c}_{m}(\theta, 1-\alpha))-\alpha\right| & \leq \frac{r}{2}\quad\forall n\geq N_{r,0}\quad\text{and}\;\label{eq - LP-size CMS}\\
\left|\sup_{m\geq n}\sup_{(\theta, F) \in \mathcal{F}_{+}}P_{F}(T_{m}(\theta) > \hat{c}_{m}(\theta, 1-\alpha))-\alpha\right| & \leq \frac{r}{2}\quad\forall n\geq N_{r,1},\label{eq - LP-size GMS}
\end{align}
by the definition of limit superior (with respect to $n$). Then by the triangular inequality,
\begin{align*}
\left|\sup_{m\geq n}\sup_{(\theta, F) \in \mathcal{F}_{+}}P_{F}(T_{m}(\theta) > \acute{c}_{m}(\theta, 1-\alpha))-\sup_{m\geq n}\sup_{(\theta, F) \in \mathcal{F}_{+}}P_{F}(T_{m}(\theta) > \hat{c}_{m}(\theta, 1-\alpha))\right|\leq r
\end{align*}
for all $n\geq N_r=\max\{N_{r,0},N_{r,1}\},$ holds. In words, given an error tolerance $r,$ the tails of the sequences of exact sizes of CMS and GMS tests are within $r$ of $\alpha$ and of each other, when $n\geq N_r.$ Thus, given $r$ (e.g., 0.0001), it is meaningful to compare the rejection probabilities along sequences of local alternatives when $n\geq N_r.$

\par Let $\mathcal{H}$ denote the set of all sequences $\{(\theta_{n,*},F_{n}): n \geq1\}$ that satisfy Assumption~LA1 and LA2. The family we consider for the comparisons is defined as
\begin{align}\label{eq - local alternatives family}
\mathcal{M}=\left\{\{(\theta_{n,*},F_{n}): n \geq1\}\in\mathcal{H}:\text{$\Omega(\theta_{n,*},F_{n})$ has nonnegative off-diagonal elements, $\forall n$}\right\}.
\end{align}
 For $\{(\theta_{n,*},F_{n}): n \geq 1\} \in \mathcal{H}$, let $\hat{\Upsilon}_{n}(\theta_{n,*}) := \{j\in\mathcal{J}:\varphi_j^{(1)}(\hat{\xi}_{n}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*}))=0\}$ and $\acute{\Upsilon}_{n}(\theta_{n,*}) := \{j\in\mathcal{J}:\varphi_j^{(1)}(\acute{\xi}_{n}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*}))=0\}$ for each $n \geq 1$. We have the following result.
\begin{theorem}\label{LPorder}
Let $\mathcal{M}$ be as in~(\ref{eq - local alternatives family}). Suppose that $S$ satisfies Part 1 of Assumption~1, $\varphi=\varphi^{(1)}$ in~(\ref{t-test MS function}), and the sequence $\{\kappa_n:n\geq1\}$ satisfies Assumption~K. For every $\{(\theta_{n,*}, F_{n}): n \geq 1\} \in \mathcal{M},$ there exists $N(\theta_{n,*}, F_{n})  \in \mathbb{Z}_{+}$ such that
\begin{align}\label{LPorder - 1}
P_{F_{n}}\Big(T_{n}(\theta_{n})>\hat{c}_{n}(\theta_{n},1-\alpha)\Big) \leq P_{F_{n}}\Big(T_{n}(\theta_{n}) > \acute{c}_{n}(\theta_{n},1-\alpha)\Big)\quad \forall n \geq N(\theta_{n,*}, F_{n}).
\end{align}
If in addition $S$ satisfies part 1 of Assumption 2 and Part 2 of Assumption 5, and the event $$\bigg\{\acute{\Upsilon}_{n}(\theta_{n,*})\subsetneq \hat{\Upsilon}_{n}(\theta_{n,*})\bigg\} \bigcap \bigg\{\hat{c}_{n}(\theta_{n,*},1-\alpha)>0\bigg\} \bigcap \bigg\{\acute{c}_{n}(\theta_{n,*},1-\alpha)<T_{n}(\theta_{n,*}) \leq \hat{c}_{n}(\theta_{n,*},1-\alpha)\bigg\}$$
has positive probability for each $n \geq N(\theta_{n,*},F_{n})$, then the weak inequalities in~(\ref{LPorder - 1}) are strict.
\end{theorem}
\begin{proof}
See Appendix~\ref{Appendix - proof Thm 3 and Prop 1}.
\end{proof}
\noindent Theorem~\ref{LPorder} states the rejection probabilities of CMS tests are no less than their GMS counterparts in large enough, but finite, sample sizes, under local alternatives in $\mathcal{M}.$ It also provides a sufficient condition for the ordering to hold strictly. Thus, for each sequence of local alternatives in $\mathcal{M}$ and small $r>0,$ the local power of a CMS test is larger than its GMS counterpart when $n \geq \max\{N(\theta_{n,*}, F_{n}),N_r\},$ where $N_r=\max\{N_{r,0},N_{r,1}\}$ and $N_{r,0}$ and $N_{r,1}$ defined in~(\ref{eq - LP-size CMS}) and~(\ref{eq - LP-size GMS}), respectively.

\par The key message from Theorem~\ref{LPorder} is that a comparison of GMS and CMS tests based on first-order asymptotics can be misleading, as it does not reflect the finite-sample situation for certain local alternatives. The result of Theorem~\ref{LPorder} is similar to Corollary~6.1 \cite{Lok-Tabri-inpress}; however, it is important to note that their result is specific to moment inequalities arising from restricted stochastic dominance orderings. Consequently, Theorem~\ref{LPorder} provides a nontrivial extension of their result to the moment inequality model with finitely many inequalities and arbitrary moment functions, when the off-diagonal elements of $\Omega(\theta_{n,*},F_{n})$ are non-negative for each $n.$

\par At the heart of this result is the marriage of the non-negative correlational structure on $\Omega(\theta_{n,*},F_{n})$ and constrained empirical likelihood estimation. This marriage begets $\hat{g}_{n,j}(\theta_{n,*}) \leq \acute{g}_{n,j}(\theta_{n,*})$ with probability approaching $1,$ for all sequences in $\mathcal{M}$ (see Lemma~\ref{LPC2}). This ordering of the estimators implies that $\varphi^{(1)}_{j}(\acute{\xi}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*})) \geq \varphi^{(1)}_{j}(\hat{\xi}_{n}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*}))$ holds with probability approaching 1, for all sequences in $\mathcal{M}$ (see Lemma~\ref{GMSorder}). It is this ordering of the moment selection functions under such sequences that gives rise to the result of Theorem~\ref{LPorder}.

\par While Theorem~\ref{LPorder} indicates that the local powers of the GMS and CMS tests can be ordered under a class of local alternatives $\mathcal{M}$ for large enough $n,$ it does not specify the extent of the discrepancy in the local powers. It is quite difficult to determine the extent of this discrepancy analytically. However, Section~\ref{Setion - MC Sims} presents Monte Carlo evidence that the discrepancy that Theorem \ref{result3} implies can be very large for local alternative sequences which have some non-violated inequalities (SNVIs). That is, sequences $\{(\theta_{n,*},F_{n}): n \geq1\}$ in $\mathcal{M}$ where there exists $j\in\mathcal{J}$ such that $E_{F_{n}}\big(g_{j}(W_i,\theta_{n,*})\big)>0\;\forall n$ and $\lim_{n\rightarrow+\infty}E_{F_{n}}\big(g_{j}(W_i,\theta_{n,*})\big)=0$.

\subsection{Power Against Distant Alternatives}\label{Subsection - Distant Alternatives}
\par This section shows CMS tests are consistent against distant alternatives. Distant alternatives include fixed alternatives and alternatives that differ from the null by greater than $O(n^{-\frac{1}{2}})$. The next assumption is useful for deducing this result, and it is the same one introduced by~\cite{andrews2010inference} in Section~9 of their paper.
\begin{dassumption}
Let $g_{n,j}^{*}=E_{F_{n}}(g_{j}(W_{i},\theta_{n,*}))/\sigma_{F_{n},j}(\theta_{n,*})$ for each $j \in \mathcal{J},$ and $\upsilon_{n}=\max_{j \in \mathcal{J}}\{-g_{n,j}^{*}\}$.
\begin{enumerate*}
\item $n^{\frac{1}{2}}\upsilon_{n}\rightarrow+\infty$ as $n\rightarrow+\infty$
\item $\Omega(\theta_{n,*},F_{n})\rightarrow \Omega_{1}$, $\Omega_{1} \in \varPsi_{2}$.
\end{enumerate*}
\end{dassumption}
\noindent The key part of this assumption is the first part, which indicates that there exists $j \in \mathcal{J}$ such that $g_{n,j}^{*}<0$ and that the violation of the non-negativity constraint is not $O(n^{-\frac{1}{2}})$. This condition differs from the setup with $n^{-\frac{1}{2}}$-local alternatives, where the sequences of alternatives $\{(\theta_{n,*},F_{n}): n \geq 1\}$ are within a $n^{-\frac{1}{2}}$-neighbourhood of $\mathcal{F}_+$.
\par We have the following result.
\begin{theorem}\label{distant}
Suppose $S$ satisfies Assumptions 1,3,4 and 6, $\varphi=\varphi^{(1)}$ in~(\ref{t-test MS function}), and the sequence $\{\kappa_n:n\geq1\}$ satisfies Assumption~K. Then
$\lim_{n\rightarrow+\infty}P_{F_{n}}\Big(T_{n}(\theta_{n,*})>\acute{c}_{n}(\theta_{n,*},1-\alpha)\Big)=1.$
\end{theorem}
\begin{proof}
See Appendix~\ref{Appendix - proof Thm 4}.
\end{proof}

\section{Simulation Results}\label{Setion - MC Sims}
\par This section studies the finite-sample performance of the CMS procedure and compares it to the GMS procedure using a simulation experiment based on the designs in~\cite{andrews2012inference}. The study uses the test statistics $S_1$ in~(\ref{MMM}) and $S_{2A}$ in~(\ref{QLR}), the recommended moment selection function $\varphi=\varphi^{(1)}$ in~(\ref{t-test MS function}), and the recommended localisation parameter $\kappa_n=(\ln n)^{1/2}.$ The nominal level is set to $\alpha=0.05,$ and we considered sample sizes $n=50,100$, and $250$. We also report simulation results for (i) the RSW procedure that use $S_1$ and $S_{2A}$, and (ii) the recommended RMS testing procedure, which combines $S=S_{2A}$, $\varphi=\varphi^{(1)}$, and $\kappa$-auto (a data-driven choice of $\kappa_n$), as additional benchmarks in studying the finite-sample performance of CMS; see Appendices~\ref{RSWdetails} and~\ref{RMSdetails}, respectively, for further details on these testing procedures. Only bootstrap versions of the tests were implemented, with 10000 bootstrap samples per Monte Carlo replication.  The computations were implemented using \textsf{R}.


\par For a given $\theta,$ the null hypothesis is $H_0:\theta\in\Theta(F_0).$ The experimental design in~\cite{andrews2012inference,andrews2012supplement} is a general formulation of that testing problem that does not require the specification of a particular form for the moment functions $\{g_{j}(\cdot,\theta): j=1,...,J\}.$ They note that the finite-sample properties of tests of $H_0$ depend on the moment functions only through (i) the vector $\mu=[E_{F_{0}}\big(g_{1}(W_i,\theta)\big),\ldots,E_{F_{0}}\big(g_{J}(W_i,\theta)\big)],$ (ii) the correlation matrix $\Omega=\text{Corr}\left(g_{1}(W_i,\theta),\ldots,g_{J}(W_i,\theta)\right),$ and (iii) the distribution of the mean zero, variance $I_J$ random vector $Z^{\dagger}=[Z_1^{\dagger},\ldots,Z_J^{\dagger}]$ where
$$
Z_j^{\dagger}=\text{Var}_{F_0}^{-1/2}\left(g_{j}(W_i,\theta)\right)\left(g_{j}(W_i,\theta)-E_{F_{0}}\big(g_{j}(W_i,\theta)\big)\right),\quad j=1,\ldots, J.
$$
We consider the case $Z^{\dagger}\sim N(0_J,I_J)$ and three correlation matrices, $\Omega_{\text{Neg}},$ $\Omega_{\text{Zero}},$ and $\Omega_{\text{Pos}},$ which exhibit negative, zero, and positive correlations.

\par The assertion of the null hypothesis in this general formulation is $H_0:\mu_j\geq0\;\forall j=1,\ldots,J$. For comparisons under the null hypothesis, we follow~\cite{andrews2012inference} by comparing the tests' maximum null rejection probabilities (MNRPs). The MNRPs are computed over the mean vectors $\mu$ in the null parameter space given the correlation matrix $\Omega\in\{\Omega_{\text{Neg}},\Omega_{\text{Zero}},\Omega_{\text{Pos}}\}$ and under the assumption of normally distributed moment inequalities. Based on simulation evidence, they conjecture that the MNRPs occur for mean vectors $\mu$ whose elements are $0$'s and $+\infty$'s. Thus, given a nominal level $\alpha,$ they compute MNRP results over the set of mean vectors $\mu$ which have that form. The results we report are for $J=2,4$ and $10.$ The matrix $\Omega_{\text{Zero}}$ equals the $J$-dimensional identity matrix. The matrices $\Omega_{\text{Neg}}$ and $\Omega_{\text{Pos}}$ are Toeplitz matrices with correlations given by the following: for $J=2:$ $\rho=-.9$ for $\Omega_{\text{Neg}}$ and $\rho=.5$ for $\Omega_{\text{Pos}};$ for $J=4:$ $\rho=(-.9,.7,-.5)$ for $\Omega_{\text{Neg}}$ and $\rho=(.9,.7,.5)$ for $\Omega_{\text{Pos}};$ for $J=10:$ $\rho=(-.9,.8,-.7,.6,-.5,.4,-.3,.2,-.1)$ for $\Omega_{\text{Neg}}$ and $\rho=(.9,.8,.7,.6,.5,\ldots,.5)$ for $\Omega_{\text{Pos}}$. As in~\cite{andrews2012inference}, the simulation study treats the correlation matrices as unknown in the implementation of all of the tests.

\par For power comparisons, we also follow~\cite{andrews2012inference,andrews2012supplement}. They compare the power of different tests by comparing their empirical power for a chosen set of alternative parameter vectors
$\mu\in\mathbb{R}^{J}$ for a given correlation matrix $\Omega\in\{\Omega_{\text{Neg}},\Omega_{\text{Zero}},\Omega_{\text{Pos}}\}$. The sets of $\mu$ vectors in the alternative are similar to the ones described in~\cite{andrews2012inference,andrews2012supplement}. We adjust those sets so as to compare the local power properties of the testing procedures. The adjustment is as follows. For each $J\in\{2,4,10\}$ the set of $\mu$ vectors are given by $\mathcal{M}_{J,n}(\Omega) = \left\{\mu/\sqrt{n}: \mu\in\mathcal{M}_J(\Omega)\right\},$ where the set $\mathcal{M}_J(\Omega)$ of $\mu$ vectors is described in Section~7.1 of~\cite{andrews2012supplement}. The $\mu$ vectors in $\mathcal{M}_{J,n}(\Omega)$ are scaled versions of those in $\mathcal{M}_J(\Omega),$ where the scaling is by $n^{-1/2}$ to create the $n^{-1/2}$-local alternatives. There are $7,24,$ and $40$ elements in $\mathcal{M}_{J,n}(\Omega)$ for $J=2,4$ and $10,$ respectively. We omit their description for brevity.

\par As the MNRPs of the tests can differ in finite-samples, the simulation results on power comparisons are based on a MNRP correction that is similar to the one employed by~\cite{andrews2012inference}. For each test statistic $S$, the MNRP correction of the CMS, GMS and RMS procedures is to add a constant based on the true matrix $\Omega$ to their corresponding critical values, so that their resulting MNRPs match that of the RSW testing procedure with nominal level $\alpha=0.05;$ see Section~\ref{scdetails} in the appendix for the details. The simulation studies in~\cite{andrews2012inference} and~\cite{romano2014practical} compare tests under the alternative using average MNRP-corrected power, where the average is computed over alternative $\mu$ vectors in $\mathcal{M}_J(\Omega).$ We report simulation results graphically using boxplots of the MNRP-corrected \emph{local} powers over sets of $\mu$ vectors $\mathcal{M}_{J,n}(\Omega)$ for the 54 different combinations of $(J,\Omega,S,n)$ for each of the CMS, GMS and RSW procedures, and 9 different combination of $(J,\Omega,S_{2A},n)$ for the recommended RMS test. Additionally, we report average MNRP-corrected local powers of the different tests across the aforementioned configurations using the symbol $\oplus$ in these plots.

\par While the average MNRP-corrected power is a useful criterion for comparing tests across $\mu$ vectors in a given set $\mathcal{M}_{J,n}(\Omega)$, it does not convey the whole picture of the tests' performance over elements in $\mathcal{M}_{J,n}(\Omega).$ Reporting boxplots, as we do, reveals the variation in powers of the tests across elements in $\mathcal{M}_{J,n}(\Omega);$ thus, presenting a broader and more extensive approach to comparing the tests under the alternative. These plots are especially useful for detecting differences in the performances of tests when the averages of their MNRP-corrected powers are close, but exhibit different distributional variations in MNRP-corrected power across $\mu$ vectors in $\mathcal{M}_{J,n}(\Omega).$

\subsection{Maximum Null Rejection Probabilities}
\par As in~\cite{andrews2010inference},~\cite{andrews2012inference}, and~\cite{romano2014practical}, empirical MNRPs are simulated as the maximum rejection probability over all $\mu$ vectors whose components are $0$ and $+\infty,$ with at least one component equal to zero. Table~\ref{Table - MNRPs} reports the MNRPs for tests. Each experiment used 10000 Monte Carlo replications when $J \in \{2,4\}$ and 2500 when $J=10$.
\begin{table}[pt]
  \caption{Empirical Maximum Null Rejection Probabilities}\label{Table - MNRPs}

\centering
\resizebox{16.5cm}{!}{
  \begin{tabular}{lccccccccccccc}
  \toprule
  & & \multicolumn{3}{r}{$J=2$} & & \multicolumn{3}{r}{$J=4$} & & \multicolumn{3}{r}{$J=10$} \smallskip\smallskip \\ \cline{4-6} \cline{8-10} \cline{12-14}
$n$ & Procedure & Statistic & $\Omega_{\text{Neg}}$ & $\Omega_{\text{Zero}}$ & $\Omega_{\text{Pos}}$ & & $\Omega_{\text{Neg}}$ & $\Omega_{\text{Zero}}$ & $\Omega_{\text{Pos}}$ & & $\Omega_{\text{Neg}}$ & $\Omega_{\text{Zero}}$ & $\Omega_{\text{Pos}}$ \smallskip\smallskip  \\
 \midrule

50   &  GMS  & $S_{1}$ &         0.06         &    0.052             &     0.052      & &   0.065               &       0.054             &         0.053      & & 0.078  & 0.054 & 0.053   \smallskip\smallskip \\
   &    & $S_{2A}$ &    0.07               &   0.052                 &  0.052         & &     0.077                 &      0.054              &      0.058         & & 0.083 & 0.054 & 0.067  \smallskip\smallskip \\
     &  CMS   &  $S_{1}$ &     0.053              &      0.052            &      0.052  & &        0.052             &       0.054          & 0.053   & & 0.062 & 0.054 & 0.053   \smallskip\smallskip \\
     &    &  $S_{2A}$ &        0.053            &       0.052          &   0.053     & &      0.054               &            0.054          &  0.058  & & 0.061 &  0.054 & 0.067 \smallskip\smallskip \\
     &  RSW   & $S_{1}$  &    0.047    & 0.047 & 0.047 & &0.047 &0.047 & 0.047 & &0.054 &0.047 & 0.047
      \smallskip\smallskip \\
           &    & $S_{2A}$  &  0.047	& 0.046 &	0.047 & &	0.047 &	0.047 &	0.047 & &	0.054	& 0.047 & 0.047
      \smallskip\smallskip \\
           &  RMS   & $S_{2A}$ & 0.053 &	0.052 &	0.056 & &	0.049 &	0.050 &	0.048 & & 0.048 &0.050 & 0.047
      \smallskip\smallskip \\
\midrule

100   &  GMS  & $S_{1}$ &  0.056	& 0.052 &	0.051 & & 	0.059 &	0.052 &	0.051 & & 	0.070	& 0.052 &	0.058  \smallskip\smallskip \\
   &    & $S_{2A}$ &   0.063	&0.052 &	0.052& &	0.067 &	0.050	& 0.055 & &	0.072	& 0.052	&  0.062 \smallskip\smallskip \\
     &  CMS   &  $S_{1}$ &  0.051&	0.052 &	0.051 && 	0.052& 	0.052 &	0.051 & &	0.061	 &0.052	 &0.058   \smallskip\smallskip \\
     &    &  $S_{2A}$ & 0.051&	0.053	&0.052	& &0.053&	0.050&	0.055 & &	0.062	&0.052	& 0.062 \smallskip\smallskip \\
     &  RSW   & $S_{1}$  & 0.047	&0.048&	0.046	&&0.046&	0.045	&0.046& &	0.055&	0.048&	0.050
      \smallskip\smallskip \\
           &    & $S_{2A}$  &  0.047	&0.048&	0.046 & &	0.046	&0.045&	0.046 & &0.056	&	0.048	& 0.051
      \smallskip\smallskip \\
           &  RMS   & $S_{2A}$ &   0.051	& 0.052	 & 0.053 &	& 0.048	 & 0.046	&0.046	& & 0.052 &	0.042	& 0.043
      \smallskip\smallskip \\
\midrule
250   &  GMS  & $S_{1}$ &  0.049 &	0.049	 & 0.051& &	0.053	&0.052 &	0.052 & &	0.053	& 0.058	 & 0.057  \smallskip\smallskip \\
   &    & $S_{2A}$ &       	0.052 &	0.049	&0.051 & &	0.057&	0.052&	0.054& &	0.057 &0.055	& 0.059 \smallskip\smallskip \\
     &  CMS   &  $S_{1}$ &    0.049 &	0.049	& 0.051 & &	0.051& 	0.052	& 0.052 & &	0.051 &	0.058 &	0.057  \smallskip\smallskip \\
     &    &  $S_{2A}$ &  0.049	 & 0.049	 &0.051	 & &0.052&	0.052&	0.054 & &	0.052 & 0.055	& 0.059 \smallskip\smallskip \\
     &  RSW   & $S_{1}$  &     0.044 &	0.043 &	0.046 & &	0.046 &	0.047 &	0.048 & &	0.046	& 0.049	& 0.052     \smallskip\smallskip \\
           &    & $S_{2A}$  &    	0.044	 & 0.043 &	0.046 & &	0.046 &	0.047 &	0.047	 & & 0.046 & 	0.049	 & 0.052
      \smallskip\smallskip \\
           &  RMS   & $S_{2A}$ &  0.049 &	0.049 & 	0.051	 &  & 0.046 &	0.048 & 	0.048	& & 0.046 &	0.046 &	0.045
      \smallskip\smallskip \\
\bottomrule

    \end{tabular}
    }
\end{table}

\par Overall, the procedures achieve a satisfactory performance for all cases considered. The RMS and RSW tests perform the best, as their MNRPs are closest to the 5\% nominal level across all of the cases considered. For the RSW procedure: the MNRPs fall within the ranges [.043,.056] and [.043,.055] when using $S_{2A}$ and $S_1$ test statistics, respectively. For the RMS test: the MNRPs fall into the range [.042,.053]. The CMS tests over-reject the null slightly: the MNRPs fall within the ranges [.049,.067] and [.049,.061] when using $S_{2A}$ and $S_1$ test statistics, respectively. The tables also show CMS tests have better MNRPs than their GMS versions as the latter tend to over-reject more: the GMS MNRPs fall within the ranges [.049,.083] and [.049,.078] for $S_{2A}$ and $S_1$, respectively. The largest MNRPs arise in the configurations where $\Omega=\Omega_{\text{Neg}},$ and these MNRPs increase with larger $J$, for the CMS, GMS and RSW tests. However, the MNRPs of all of these tests do get closer to the 5\% nominal level with larger sample sizes, across all configurations, and for CMS tests, this numerical result is a consequence of Theorem~\ref{asymptoticsize}.

\par While we don't have a theoretical result on improved size control of CMS tests over their GMS versions, Table~\ref{Table - MNRPs} provides simulation-based evidence of such an improvement. Hence, these results point to the potential benefit of implementing the information~(\ref{eq - moment inequality}), as we do, in two step testing procedures, under the null. The next section presents simulation results on MNRP-corrected power of these tests, under local alternatives, and illustrates the result of Theorem~\ref{LPorder}.

\subsection{Local Power}
\par Figures~\ref{Figure - Power QLR} and~\ref{Figure - Power MMM} below report boxplots of the MNRP-corrected powers of the tests under $S_{2A}$ and $S_1$, respectively. The results can be summarised as follows. For each test statistic, the MNRP-corrected power values of the tests are generally distributed in a similar way in configurations where $\Omega=\Omega_{\text{Neg}},$ and all of the tests have comparable average powers in those configurations. By contrast, in configurations where $\Omega=\Omega_{\text{Zero}}$, for each test statistic, the boxplots show the RSW tests' MNRP-corrected power values tend to be (i) more dispersed (as shown by the lengths of their boxes), (ii) have a wider overall range, and (iii) have lower average power in comparison to the remaining tests, which all behave similarly as can be seen by their boxplots. For example, the average power of the RSW test when $S=S_{2A}$, $J=10$, and $n=250$ is approximately equal to 0.57, while the averages of the remaining procedures in that scenario are all approximately equal to 0.66, which is a large difference.

\par More noticeable differences in the tests' performance arise in configurations where $\Omega=\Omega_{\text{Pos}}.$ For each test statistic, there is evidence for the following ranking in terms of average MNRP-corrected power, uniformly in $J$ and $n$: CMS in first place, GMS in second place, and RSW in third place, with the RMS test tied in first place with the CMS-$S_{2A}$ test. The boxplots also show:
\begin{itemize}
\item The MNRP-corrected power values for the RMS and CMS-$S_{2A}$ tests are generally distributed in a similar way for each $n$ and $J$, except when $J=2$ the CMS-$S_{2A}$ power values are slightly more dispersed (as shown by the length of the boxes) than their RMS counterparts for each $n$.
\item For each $S$, the MNRP-corrected power values of CMS tests are markedly less dispersed and have smaller overall ranges than their GMS and RSW counterparts.

\item The difference among the CMS, GMS and RSW tests in these configurations with $S=S_1$ can be strikingly large in terms of average power; for example, with $J=4$ and $n=250,$ the average powers of CMS, GMS and RSW tests are approximately equal to 0.75, 0.65, and 0.60, respectively. By contrast, with $S=S_{2A}$ the difference among these tests is less pronounced, which is on account of using a more effective test statistic. For example, in the aforementioned configuration, the average powers of CMS, GMS and RSW tests are approximately equal to 0.76, 0.74, and 0.73, respectively. However, this less pronounced difference in average powers does not mean that these procedures behave similarly, as evidenced by the radically different boxplots of the tests' power values.
\end{itemize}


\begin{landscape}
 \begin{figure}
\centering
\includegraphics[height=16cm, width=18cm]{figure2}
\caption{Boxplots of MNRP-corrected powers of $S_{2A}$-based tests. For each configuration, the symbol $\oplus$ marks the location of the average MNRP-corrected power of a test. }\label{Figure - Power QLR}
\end{figure}
\end{landscape}

\begin{landscape}
 \begin{figure}
\centering
\includegraphics[height=16cm, width=18cm]{figure3}
\caption{Boxplots of MNRP-corrected powers of $S_{1}$-based tests. For each configuration, the symbol $\oplus$ marks the location of the average MNRP-corrected power of a test. }\label{Figure - Power MMM}
\end{figure}
\end{landscape}



\par The result of Theorem~\ref{LAthm} implies that the average power of CMS and GMS tests should get closer together with larger sample sizes. The simulations reflect this implication across all configurations, but indicate that it happens slowly when $\Omega=\Omega_{\text{Pos}}.$ Consequently, there is simulation-based evidence that shows the implementation of the information~(\ref{Moment inequality restrictions S3}), as we do with CMS, may not improve the local power of GMS tests for configurations in which $\Omega\neq\Omega_{\text{Pos}}.$ The reason is that the boxplots for MNRP-corrected power values of CMS and GMS tests are generally quite similar in those configurations. By contrast, the result of Theorem~\ref{LPorder} points to such an improvement in local power for configurations in which $\Omega=\Omega_{\text{Pos}}$, and this result is reflected in the simulations as described above.

\par Table~\ref{Table - LP Pos Omega} reports the average MNRP-corrected powers of the tests when $\Omega=\Omega_{\text{Pos}}$ and we use these results to further contextualize the local power improvement associated with CMS tests over GMS and RSW tests. We benchmark our analysis to RMS because simulation evidence in~\cite{andrews2012inference} suggests that it is superior in terms of asymptotic average power and is therefore the recommended test. The CMS-$S_{2A}$ and RMS tests are neck and neck as their average powers are essentially identical and achieve the highest average powers in all of those scenarios, with the CMS-$S_1$ test having slightly lower average powers than those tests. For a given $S$, the RSW tests are the worst performing, as they achieve the lowest average powers in each corresponding scenario, and the difference between them and the RMS test can be quite large. For example, when $J=10$ and $n=250$, the difference between RSW-$S_1$ and RMS is 0.252, and with RSW-$S_{2A}$ it is 0.055 which is a much smaller on account of using a more effective test statistic. The CMS-$S_1$ test dominates the GMS-$S_{1}$ test in each of those scenarios, where the difference can be as large as 10 percentage points -- see the scenarios with $J=10$. Consequently, the importance of incorporating the statistical information from the constraints, as we do with CMS, picks up the difference in average powers between the RMS and GMS test when $S=S_{2A}$ and most of the difference when $S=S_1,$ in each of the those scenarios.

\begin{table}[pt]
  \caption{MNRP-Corrected Average Powers: $\Omega = \Omega_{pos}$}\label{Table - LP Pos Omega}

\centering
\resizebox{13cm}{!}{
  \begin{tabular}{lccccccccc}
  \toprule
$J$ & $n$ & & GMS-$S_{1}$ & GMS-$S_{2A}$ & CMS-$S_{1}$ & CMS-$S_{2A}$ & RSW-$S_{1}$ & RSW-$S_{2A}$ & RMS \smallskip\smallskip  \\
 \midrule
& $50$ & & 0.658 &	0.676	&0.677	&0.691	&0.611	&0.652	&0.692  \smallskip\smallskip \\
2 & $100$ & & 0.654 &	0.681&	0.682	&0.697	&0.620&	0.661&	0.700 \smallskip\smallskip \\
 & $250$ & &0.662&	0.687	&0.689	&0.701&	0.630	&0.672&	0.709\smallskip\smallskip \\
 \midrule
 & $50$ & & 0.675	&0.743	&0.752&	0.759	&0.575	&0.715&	0.750\smallskip\smallskip \\
4 & $100$ & & 0.667	&0.747	&0.728&0.763&	0.587&	0.729&	0.760 \smallskip\smallskip \\
 & $250$ & & 0.656	&0.746&	0.738&	0.761&	0.593	&0.733&	0.761\smallskip\smallskip \\
 \midrule
 & $50$ & & 0.692	&0.782	&0.787&	0.803&	0.557	&0.746	&0.797 \smallskip\smallskip \\
10 & $100$ & & 0.696&	0.796&	0.799	&0.819&	0.573	&0.765&	0.825\smallskip\smallskip \\
 & $250$ & & 0.695	&0.810	&0.802	&0.830	&0.585	&0.782&	0.837\smallskip\smallskip \\
 \bottomrule

    \end{tabular}
    }
\end{table}

\par While the focus above has been on average power, for individual $\mu$ vectors the power differences can be massive with $\Omega=\Omega_{\text{Pos}}$. Consider, for example, the element $\mu/\sqrt{n}\in\mathcal{M}_{4,n}(\Omega_{\text{Pos}})$ with $\mu=(-2.4705,1,1,1)^{\intercal}$. This mean vector is an example of an SNVI local alternative. Table \ref{Table - Illustration Output} reports the MNRP-corrected power estimates for the tests under this local alternative for $n=50,100,250$. The estimates indicate that:
\begin{table}[pt]
  \caption{MNRP-Corrected Power: $\mu/\sqrt{n}\in\mathcal{M}_{4,n}(\Omega_{\text{Pos}})$ with $\mu=(-2.4705,1,1,1)^{\intercal}$.} \label{Table - Illustration Output}
  \smallskip
\centering
\resizebox{13cm}{!}{
  \begin{tabular}{lccccccccc}
  \toprule
$J$ & $n$ & & GMS-$S_{1}$ & GMS-$S_{2A}$ & CMS-$S_{1}$ & CMS-$S_{2A}$ & RSW-$S_{1}$ & RSW-$S_{2A}$ & RMS \smallskip\smallskip  \\
 \midrule
  & $50$ &   & 0.3        &	0.668	     & 0.684	   & 0.733	      & 0.237	    & 0.633	       & 0.734  \smallskip\smallskip \\
4 & $100$ &  & 0.275      &	0.663        & 0.625	   & 0.726	      & 0.236       & 0.645        & 0.74 \smallskip\smallskip \\
  & $250$ &  & 0.256      &	0.665	     & 0.616	   & 0.718        &	0.234	    & 0.654        & 0.744\smallskip\smallskip \\
\bottomrule

    \end{tabular}
    }
\end{table}
 \begin{figure}[H]
\centering
\includegraphics[height=6.5cm, width=15cm]{figure1}
\caption{ECDFs of MNRP-corrected CMS (solid line), GMS (dashed line), RMS (dash-dot line), and RSW (dotted line) critical values using the $S_1$ (MMM) and $S_{2A}$ (AQLR) test statistics.}\label{Figure - ECDFs_new}
\end{figure}

\begin{itemize}
\item There can be \textit{extremely} large power improvements associated with CMS relative to RSW and GMS when $S=S_{1}$. Indeed, the improvement in power of CMS over GMS is approximately 36 percentage points and 40 percentages points over RSW.
\item The improvements persist with $S=S_{2A}$, but are not as large. The AQLR statistic results in CMS experiencing a six percentage point improvement over GMS and eight percentages point improvement over RSW. In absolute terms, all procedures experience higher local power with $S=S_{2A}$.
\item The MNRP-corrected powers of CMS-$S_{2A}$ are comparable to their RMS counterparts.
\end{itemize}



 To gain a deeper insight into the behavior of the tests under this local alternative, Figure \ref{Figure - ECDFs_new} reports the empirical distribution functions (ECDFs) of the MNRP-corrected critical values for $n=250.$ The focus on this sample size is without loss of generality as similar graphs of the critical values' ECDFs arise in all of the other values of $n$ we considered. For either test statistic, the ECDFs in Figure~\ref{Figure - ECDFs_new} show strong evidence of a first-order stochastic dominance ranking among the critical values of the CMS, GMS, and RSW, tests. Specifically, for both types of test statistics, there is evidence for the ordering $\acute{c}_{n} \leq \hat{c}_{n} \leq \check{c}_{n}$, where $\check{c}_{n}$ denotes the RSW critical value. By contrast, the ECDF of the recommended RMS test crosses that of CMS with $S=S_{2A},$ which means that there isn't evidence of a clear ordering of their critical values. Overall, the differences between the ECDFs is quite striking and indicates that there is a big difference in the behavior of the tests even in moderately large sample sizes. The stochastic ordering of the CMS and GMS critical values is a reflection of Theorem \ref{result3} and provides evidence for local power improvements under SNVI local alternatives which have positively correlated moment functions. Finally, we discuss the behavior of the RSW procedure. The RSW procedure rejects on the event $\{M_{n}(\beta) \nsubseteq \mathbb{R}_{+}^{J}\}\bigcap \{S>\check{c}_{n}\}$, where $M_{n}(\beta)$ is a lower confidence rectangle that is used to detect ``positive'' moments in the first step of their two-step procedure (see Appendix \ref{RSWdetails}). Across the two test statistics, our simulations indicate that (i) the event $\{M_{n}(\beta) \nsubseteq \mathbb{R}_{+}^{J}\}$ occurs with empirical probability close to $1$, and (ii) in $9465$ times out of $10000$ Monte Carlo replications, their critical value $\check{c}_{n}$ corresponds to the case where none of the moment inequalities have been omitted from its calculation. These findings show the RSW procedure fails to reliably detect the ``positive'' moments in $\mu/\sqrt{n}$ in most of the $10000$ Monte Carlo replications, resulting in it having low empirical power.

\section{Conclusion}\label{Section - Conclusion}
\par This paper has proposed a surgical modification of the generalized moment selection (GMS) procedure put forward by~\cite{andrews2010inference} that improves its performance, called constrained moment selection (CMS). The basic idea of the CMS procedure is to use empirical likelihood to incorporate the information embedded in the moment inequality constraints into the moment selection step of the GMS procedure. Our analyses highlights the importance of using this information to more reliably detect the binding moments, which is the source of the improvement of CMS over GMS tests.

\par There are a number of directions for future research. Although we focus on modifying GMS tests, the intuition of incorporating the information embedded in the identified set transcends this choice and we conjecture that similar finite-sample benefits would arise in a similar modification to the two-step procedure of~\cite{romano2014practical}. There is also an emerging literature that focuses on testing with `many' moments, where the number of inequalities grow exponentially with the sample size (e.g.,~\citealp{chernozhukov2019inference}, and~\citealp{bai2019practical}). Extending the empirical likelihood modification to such testing procedures may improve their performance, but different theoretical tools must be employed to account for the increasing number of constraints. Finally, our paper is related to the semi-infinite programming empirical likelihood procedure proposed by~\cite{Lok-Tabri-inpress} for two-step bootstrap tests of stochastic dominance, where the continuum of unconditional moment inequalities is akin to inference for conditional moment inequalities. Their results are limited to restricted stochastic dominance tests and it would be interesting to extend the insights from this paper to the general conditional moment inequality models of~\cite{andrews2013inference, andrews2017inference}.

\section{Acknowledgements}
We are grateful to Jonathan Roth for providing valuable comments. We are also appreciative of feedback from participants at the Graduate Student Workshop in Econometrics, Harvard University. The computations in this paper were run on the FASRC Cannon cluster supported by the FAS Division of Science Research Computing Group at Harvard University. All errors are our own.
\bibliography{CMS}
\newpage