EconBase
← Back to paper

Inference for Moment Inequalities: A Constrained Moment Selection Procedure

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,061 characters · 14 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for Moment Inequalities: A Constrained Moment Selection Procedure

abstractInference in models where the parameter is defined by moment inequalities is of interest in many areas of economics. This paper develops a new method for improving the performance of generalized moment selection (GMS) testing procedures in finite-samples. The method modifies GMS tests by tilting the empirical distribution in its moment selection step by an amount that maximizes the empirical likelihood subject to the restrictions of the null hypothesis. We characterize sets of population distributions on which a modified GMS test is (i) asymptotically equivalent to its non-modified version to first-order, and (ii) superior to its non-modified version according to local power when the sample size is large enough. An important feature of the proposed modification is that it remains computationally feasible even when the number of moment inequalities is large. We report simulation results that show the modified tests control size well, and have markedly improved local power over their non-modified counterparts.

Keywords: empirical likelihood, moment inequality model, statistical information. \\ JEL Classification: C12, C14, C21

Introduction

Statistical inference in models defined by moment inequalities is a frequently encountered topic in econometrics. Examples of applications include games of entry with multiple equilibria (e.g., ciliberto2009market), single/multiple agent optimization problems (e.g., pakes2015moment), censored and missing data (e.g., manski2002inference,imbens2004confidence), model selection tests (e.g., Shi-KL, and Shi-Hsu), event-study designs (e.g., rambachanhonest), stochastic dominance comparisons (e.g. whang_2019) and New-Keynesian DSGE models (e.g., moon2009estimation). This paper considers inference for a finite-dimensional parameter defined by a finite number of unconditional moment inequalities.

We suppose that there exists a true value of the parameter $\theta_0\in\Theta\subseteq\mathbb{R}^{d}$ that satisfies the moment inequality restrictions

align[align omitted — 121 chars of source]

where $\{g_{j}(\cdot,\theta): j=1,...,J\}$ are known real-valued functions, $\{W_i:i\leq n\}$ are independent and identically distributed (i.i.d.) with unknown distribution $F_0,$ and $W_i\in\mathbb{R}^{\text{dim}(W_i)}.$ Under these moment conditions, the set $\Theta_{I}(F_0)\equiv\left\{\theta\in\Theta: E_{F_{0}}\big(g_{j}(W_i,\theta)\big) \geq 0\;\forall j=1,...,J\right\}$ denotes the so-called identified set while any $\theta\in\Theta_{I}(F_0)$ is termed an identifiable parameter. Thus, the true value of the parameter might not be uniquely identified by $F_0$ and the economic model.

We are interested in confidence sets for $\theta_0$ constructed by test inversion. The test is based on a statistic $T_n,$ for testing individual hypotheses for each $\theta$ that have the form

align[align omitted — 126 chars of source]

Inference in this model is challenging because the pointwise limiting null distribution of conventional test statistics are discontinuous in the parameter -- the dependence on the parameter is through the index set of moment inequalities ((ref)) that are binding. In particular, a moment inequality enters the pointwise asymptotic null distribution of the test statistic $T_n$ whenever it holds as an equality. Tests of ((ref)) that have good properties incorporate information about which moments $E_{F_{0}}\big(g_{j}(W_i,\theta)\big)$ are “positive”, in order to exclude them from the computation of a critical value. Tests of this sort are known as two-step procedures in the literature, examples of which include andrews2010inference, canay2010inference, andrews2012inference, and romano2014practical. The first step of those testing procedures use the data to determine whether the moment inequalities ((ref)) are close to or far from being equalities. The second step uses the outcome of the first step to yield information about which moment inequalities are “positive” when constructing tests of ((ref)).

The literature on two-step tests of ((ref)) is vast, and almost all of these tests use the sample-analogue estimator of the moments $E_{F_{0}}\big(g_{j}(W_i,\theta)\big)$ in the first step to determine the slackness of the moment inequalities. This feature ignores the information present in the restrictions ((ref)), because the sample-analogue estimator does not exploit the fact that the moments satisfy these restrictions under the null hypothesis in ((ref)). Thus, we conjecture that implementing this information in such tests can improve their accuracy in finite-samples under the null and alternative hypotheses. This paper provides such a modification for the broad class of generalized moment selection (GMS) testing procedures put forward by andrews2010inference, and finds that our conjecture is in the right direction.

We propose a modification of GMS testing procedures that implements the information present in ((ref)) using the method of empirical likelihood (owen2001empirical). The modification is to replace the sample-analogue estimator of the moments in the first step of the GMS procedure with its constrained empirical likelihood counterpart, where the constraints are the moment inequalities ((ref)). We label this modification constrained moment selection (CMS). For a given test statistic and moment selection function, the CMS and GMS tests only differ in terms of which moments they select for the computation of the critical value in tests of ((ref)). The motivation for our proposal is that the detection of the “positive” moment inequalities in the first step would be more accurate because we are using additional information that is available to us, which the sample-analogue estimator of the moments ignores. Consequently, the CMS procedure alters the GMS critical value for testing ((ref)) in a data-dependent way that incorporates the information contained in ((ref)) through a reduction of the parameter space for $F_0.$ For this reason, we expect CMS tests of ((ref)) to be more accurate than their GMS counterparts in finite-samples.

This paper characterises the parameter space for $(\theta_0,F_0)$ over which the CMS and GMS testing procedures are asymptotically equivalent, to first-order, under the null, local alternatives, and distant alternatives. We focus, though, on the GMS class of testing procedures in which the moment selection function is given by the moment selection $t$-test. This focus is without loss of generality, as the results extend naturally, with appropriate modifications, to the more general setup in andrews2010inference using their assumptions. This means that for a given test statistic, CMS tests inherit all of the asymptotic properties of GMS tests. Specifically, under the null, CMS confidence sets are asymptotically valid with uniformity over the parameter space, not asymptotically conservative, and not asymptotically similar. Furthermore, CMS tests of ((ref)) have greater asymptotic local power than tests based on subsampling or fixed asymptotic critical values, and are consistent against distant alternatives. The parameter space imposes only three conditions in addition to the conditions that define the parameter space andrews2010inference introduce. These conditions are part of Assumption GEL in andrews2009validity: (i) a uniform bound on the variances of the moment functions, (ii) a lower bound on the determinant of their correlation matrix, and (iii) a regularity condition on an estimator of the degree of slackness of the moments arising from the dual formulation of the constrained empirical likelihood problem. Collectively, the conditions that define our parameter space enables the use of results from andrews2009validity on constrained empirical likelihood estimation in our proofs of the aforementioned asymptotic results.

While GMS and CMS are asymptotically equivalent procedures, we characterise local alternatives under which the power of CMS tests dominate their GMS counterparts for sufficiently large, but finite, samples. These are directions in the alternative that have some non-violated moment inequalities (SNVIs) and a non-negative correlational structure. That is, configurations where some of the moments $E_{F_{0}}\big(g_{j}(W_i,\theta_{0})\big)$ under the alternative hypothesis are “positive", and the covariance matrix of $\{g_{j}(W_i,\theta): j=1,...,J\}$ has non-negative entries only. The non-negative correlational structure arises in empirical applications; see, for example, Lok-Tabri-inpress who point to that structure for moment inequalities characterising stochastic dominance comparisons. It is quite difficult to determine the extent of this difference in local powers analytically. However, using a Monte Carlo simulation experimental design based on andrews2012inference, who focus on finite-sample comparisons of the maximum null rejection probability (MNRP), we show using the modified method of moments (MMM) statistic that along such local alternatives the differences in MNRP-corrected powers of CMS and GMS tests can be approximately 36 percentage points when $J=4$ and $n=250,$ which is strikingly large. See Section (ref) for more details.

The two-step tests in this literature that exploit the information ((ref)) are the procedures put forward by andrews2009validity and canay2010inference. They implement this information using (generalised) empirical likelihood. andrews2009validity and canay2010inference develop subsampling and bootstrap tests of ((ref)), respectively, using empirical-likelihood-type test statistics. Both tests have correct asymptotic size in a uniform sense and are shown not to be asymptotically conservative. However, canay2010inference's test has higher asymptotic power because it is a GMS procedure. More generally, andrews2010inference show the asymptotic power of GMS tests dominate that of subsampling and plug-in asymptotic tests. A disadvantage of canay2010inference's procedure is that it may be more computationally burdensome than other GMS tests. Thus, our modification of GMS tests can improve finite-sample performance without incurring a high computational cost.

andrews2012inference proposed a refinement of GMS termed refined moment selection (RMS) and discussed the reasons why such an approach is preferable. However, the RMS procedure is quite computationally expensive when $J>10.$ By contrast, the CMS procedure remains computationally feasible when $J$ is large. The reason is that the constrained empirical likelihood optimization problem it is based upon has a strictly concave objective function, convex feasible set, and the choice variables enter linearly into the constraints. As a consequence, there is a unique global solution to this optimization problem and its implementation involves an of-the-shelf programming routine. More recently, romano2014practical proposed a two-step testing procedure for moment inequalities that is similar in spirit to the RMS procedure and remains computationally feasible when $J$ is large. An important distinction between the CMS testing procedure and these tests is that, like GMS tests, neither of them exploits the information present in the moment inequality constraints ((ref)), because they employ the sample-analogue estimator of the moments in their first step.

We examine the finite-sample performance of CMS tests using the MMM and adjusted quasi-likelihood-ratio (AQLR) test statistics in Monte Carlo simulations based on the experimental design in andrews2012inference. The experiment compares the performance of CMS to its GMS, RMS, RSW counterparts in terms of MNRP and MNRP-corrected local power. The inclusion of the RMS and RSW procedures in the simulation experiment is to benchmark the performance of CMS. Overall, the simulation results showcase the value of implementing the information ((ref)) in the CMS procedure in terms of finite-sample size and power properties, and corroborate its theoretical superior performance over GMS. The simulation results also show the performance of CMS and RMS tests based on the AQLR statistic are comparable. This finding is encouraging as the RMS test has desirable asymptotic properties but can be computationally expensive when $J$ is large, while the CMS procedure isn't costly to compute at all.

The idea of exploiting information on parameters defined by constraints for improving performance in statistical problems, through constrained estimation, is one of the most natural ideas in statistics. The literature on constrained estimation via tilting the empirical distribution overlaps with this paper, where the problem is that the constraints/information are not adequately reflected by the empirical distribution (e.g., Hall-Presnell). Tilting the empirical distribution allows one to incorporate information selectively into a statistical procedure without changing the procedure itself. Lok-Tabri-inpress apply this idea to modifying two-step bootstrap tests for restricted stochastic dominance orderings using empirical likelihood and semi-infinite programming. The parameter of interest in their setup is infinite-dimensional and there is a continuum of moment inequality restrictions, which are defined by moment functions that have a particular form. The form of the moment functions in their setup yields a correlational structure that facilitates the analysis of such moment inequalities. Contrastingly, in the setup of this paper, $J$ is finite and the form of the moment functions $\{g_{j}(\cdot,\theta): j=1,...,J\}$ is arbitrary. The implementation of empirical likelihood in their setup has a data-driven number of inequalities that increases with the sample size, which can be as large as 500 in moderate sample sizes. The ability of empirical likelihood to straightforwardly execute with a large number of moment inequality restrictions transfers to the CMS procedure for models with large $J.$ This computational feasibility of CMS is an important feature of our approach. Similar to Lok-Tabri-inpress, this paper is also part of the econometrics literature on shape restrictions (e.g., Chetverikov-Santos-Shaikh, and the references therein), as the inequalities ((ref)) can be thought of as finite-dimensional analogues of shape restrictions on nonparametric functions.

We organize the paper as follows. Section (ref) introduces the statistical framework, as well as the GMS and CMS procedures. Section (ref) introduces the main results of the paper. Section (ref) reports the results of Monte Carlo simulations, and Section (ref) concludes.

For notational simplicity, throughout the paper we write partitioned column vectors as $h=(h_1,h_2).$ rather than $h=(h_{1}',h_{2}')'$. Let $\mathbb{R}_{+}=\{x\in\mathbb{R}:x\geq0\}$, $\mathbb{R}_{+,\infty}=\mathbb{R}_{+} \cup\{+\infty\}$, $\mathbb{R}_{[+\infty]}=\mathbb{R} \cup\{+\infty\}$,$\mathbb{R}_{[\pm\infty]}=\mathbb{R} \cup\{\pm\infty\}$, “$:=$" denote the definitional identity, and $\overline{A}$ denote the closure of a set $A$.

Setup

Moment Inequality Model and Test statistic

The object of interest is a parameter $\theta_{0} \in \Theta \subseteq \mathbb{R}^{d}$, $d<+\infty$, defined by a finite number of known moment functions $g_{j}: \mathcal{W} \times \Theta \rightarrow \mathbb{R}$ that satisfy the following unconditional moment inequality restrictions:

align[align omitted — 128 chars of source]

where $F_{0}$ denotes the true distribution of the observed data $W$ and $\mathcal{J}:=\{1,...,J\}$ with $J<\infty$. In general, the identified set, $\Theta_{I}(F_0)=\{\theta \in \Theta: \ E_{F_0}(g_{j}(W,\theta)) \geq 0 \ \forall \ j \in \mathcal{J}\}$, is not a singleton meaning that the parameter is partially identified.

The moment inequality model is given by the following definition.

definition[Moment Inequality Model] Let $\mathcal{F}$ be the set of parameters $(\theta, F)$ that satisfy: \begin{enumerate} • $\theta \in \Theta\subseteq \mathbb{R}^{d}$. • $\{W_{i}: i \geq 1\}$ are i.i.d. under $F$. • $E_{F}\big(g_{j}(W_{i},\theta)\big) \geq 0 \ \text{for} \ j\in\mathcal{J}$. • $\sigma^{2}_{F,j}(\theta) := Var_{F}\big(g_{j}(W_{i},\theta)\big) \in [\varepsilon_{*}, M_{*}]$ for some $M_{*}>\varepsilon_{*}>0$. • $\Omega(\theta,F) \in \varPsi_{2}$, where $\Omega(\theta,F)$ is the $J\times J$ correlation matrix of $\{g_{j}(W_{i},\theta),j=1,\ldots,J\},$ and $\varPsi_{2}$ is the space of correlation matrices whose determinant is greater than $\varepsilon>0$. • $\exists \delta$ and $M >0:$ $E_{F}|g_{j}(W_{i},\theta)/\sigma_{F,j}(\theta)|^{2+\delta} \leq M \ \forall \ j\in\mathcal{J}$. \end{enumerate}

All of the conditions in this definition, except for Conditions (ref) and (ref), are those presented in (2.2) of andrews2010inference. Condition (ref) is a strengthening of Condition (v) in andrews2010inference so that the variances of the moment functions are uniformly bounded. Condition (ref) specifies the nonsingularity of the matrix $\Omega(\theta,F).$ These conditions are relatively unrestrictive and are part of Assumption GEL in andrews2009validity. Furthermore, they arise frequently in papers that consider empirical likelihood inference for moment inequalities (e.g., canay2010inference, and Lok-Tabri-inpress).

For a given value of the parameter, $\theta=\theta_0,$ we invert tests of the hypothesis $H_0:\,\theta_0\in\Theta_{I}(F_0)$ to construct confidence sets of the form $CS_{n}=\{\theta \in \Theta: T_{n}(\theta) \leq c_{1-\alpha}(\theta)\},$ where $T_{n}(\theta)$ denotes a test statistic and $c_{1-\alpha}(\theta)$ is a critical value for tests with nominal level $\alpha \in (0, 1/2)$. We say $CS_{n}$ is a uniformly valid confidence set for $\theta$ if

align[align omitted — 157 chars of source]

where $P_{F}(\cdot)$ is the probability measure induced by repeated sampling from $F$. Uniformity is essential in order for asymptotic size to be a good approximation to the finite-sample size of confidence sets, because the test statistic exhibits a discontinuity in its asymptotic distribution (as a function of the distribution generating the data), but not in its finite-sample distribution. Discontinuities of this type can create asymptotic size problems that are analogous to those that arise with parameters that are near a boundary (e.g., andrews2009validity).

A test statistic is a function $S: \mathbb{R}_{[+\infty]}^{J}\times \mathcal{V}_{J\times J} \rightarrow \mathbb{R}$ given by $T_{n}(\theta_0):=S\left(n^{\frac{1}{2}}\hat{g}_{n}(\theta_0), \hat{\Sigma}_{n}(\theta_0)\right),$ where $\mathcal{V}_{J\times J}$ is the set of invertible $J\times J$ variance matrices,

align*[align* omitted — 402 chars of source]

Two examples are the modified method of moments (MMM) and adjusted quasi-likelihood-ratio (AQLR) statistics. In the context of the moment inequality model $\mathcal{F}$ given by Definition (ref), these test statistics are defined as

align[align omitted — 467 chars of source]

respectively, where $\tilde{\Sigma}_{n}(\theta_0) = \hat{\Sigma}_{n}(\theta_0)+\max\{0, 0.012-|\hat{\Omega}_{n}(\theta_0)|\}\hat{D}_{n}(\theta_0),\; \hat{\Omega}_{n}(\theta_{0})= \hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})\hat{\Sigma}_{n}(\theta_{0})\hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})$ and $\hat{D}_{n}(\theta_{0})=\operatorname*{diag} \hat{\Sigma}_{n}(\theta_{0})$, where $\operatorname*{diag} \hat{\Sigma}_{n}(\theta_{0})$ is a diagonal matrix with dimensions equal to those of $\hat{\Sigma}_{n}(\theta_{0})$ whose diagonal elements equal those of $\hat{\Sigma}_{n}(\theta_{0})$.

GMS and CMS Procedures

The point of departure for establishing that ((ref)) holds for the GMS procedure is to consider the asymptotic distribution of $T_n(\theta_0)$ under a suitable sequence of null distributions. For any sequence $\{F_n: n\geq1\}$ in the model of the null hypothesis, the test statistic satisfies

align[align omitted — 165 chars of source]

where $h_1\in\mathbb{R}^{J}_{+,\infty},$ and $\Omega_0$ is a $J\times J$ correlation matrix.\footnote{Specifically, this large-sample result ((ref)) follows from the form of the test statistic, the Central Limit Theorem, and the convergence in probability of the sample correlation matrix.} The vector $h_1=(h_{1,1},...,h_{1,J})'$ has elements given by $\lim_{n\rightarrow+\infty}n^{1/2}(E_{F_n}(g_{j}(W_{i},\theta_0))/ \sigma_{F,j}(\theta_0)$ and measures the degree of slackness of the moment inequalities. The crux of this asymptotic construction is that the limiting distribution in ((ref)) now depends continuously on the degree of slackness of the moment inequalities via the parameter $h_1,$ which reflects the finite-sample situation.

The asymptotic implementation of the GMS critical value is the $1-\alpha$ quantile of a data-dependent version of the asymptotic null distribution in ((ref)). It replaces $\Omega_0$ by a consistent estimator and replaces $h_1$ with a function $\varphi : \mathbb{R}^{J} \times \varPsi_{2} \rightarrow \mathbb{R}^{J}_{+,\infty}$, which measures the slackness of moment inequalities through $\hat{\xi}_{n}(\theta_{0})=\kappa_{n}^{-1}n^{\frac{1}{2}}\hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})\hat{g}_{n}(\theta_{0}),$ where $\{\kappa_{n}: n \geq 1\}$ is a divergent sequence of scalars andrews2010inference. The GMS critical value, $\hat{c}_{n}(\theta_{0},1-\alpha),$ is the $1-\alpha$ quantile of

align[align omitted — 210 chars of source]

where $Z^{*}\sim N(0_{J}, I_{J})$ and is independent of $\{W_i: i\geq 1\}.$ That is,

align[align omitted — 148 chars of source]

where $P\Big(L_n(\theta_0,Z^{*})\leq x\Big)$ denotes the conditional CDF at $x$ of $L_n(\theta_0,Z^{*}),$ conditional upon $(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0})).$ In practice, the calculation of $\hat{c}_{n}(\theta_{0},1-\alpha)$ is by simulating $L_n(\theta_0,Z^{*})$ using $R$ i.i.d. draws from $Z^{*}\sim N(0_{J}, I_{J})$ and computing the $1-\alpha$ quantile of the empirical CDF from $\{L_n(\theta_0,Z_r^{*}): r =1,...,R\}.$

Alternatively, one may compute the GMS critical value using the bootstrap. We briefly describe this approach. Let $\{W_{i}^{*}: i \leq n\}$ be a bootstrap sample drawn from the empirical distribution of the data $\{W_{i}: i \leq n\}$, and define $\hat{g}_{n}^{*}(\theta_{0}) = n^{-1}\sum_{i=1}^{n}g(W_{i}^{*},\theta_{0})$, $\hat{\Sigma}_{n}^{*}(\theta_{0})=n^{-1}\sum_{i=1}^{n}(g(W_{i}^{*},\theta_{0})-\hat{g}_{n}^{*}(\theta_{0}))(g(W_{i}^{*},\theta_{0})-\hat{g}_{n}^{*}(\theta_{0}))^{\intercal},$ $\hat{D}_{n}^{*}(\theta_{0}) = \text{diag} \hat{\Sigma}_{n}^{*}(\theta_{0})$, and $\hat{\Omega}_{n}^{*}(\theta_{0}) = (\hat{D}_{n}^{*}(\theta_{0}))^{-\frac{1}{2}}\hat{\Sigma}_{n}^{*}(\theta_{0})(\hat{D}_{n}^{*}(\theta_{0}))^{-\frac{1}{2}}$. The bootstrap implementation of the GMS procedure replaces $L_{n}(\theta_{0},Z^{*})$ in ((ref)) with $$L_{n}(\theta_{0},\{W_{i}^{*}: i \leq n\})=S\left(G_{n}^{*}(\theta_{0})+\varphi(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0})), \hat{\Omega}_{n}^{*}(\theta_{0})\right),$$ where $G_{n}^{*}(\theta_{0}) = n^{\frac{1}{2}}(\hat{D}_{n}^{*}(\theta_{0}))^{-\frac{1}{2}}(\hat{g}_{n}^{*}(\theta_{0})-\hat{g}_{n}(\theta_{0})),$ and defines a critical value analogous to ((ref)). In practice, this critical value is the empirical $1-\alpha$ quantile of the bootstrap statistics $\{L_{n}(\theta_{0},\{W_{i,r}^{*}: i \leq n\}):r=1,...,R\}$, where $\{\{W_{i,r}^{*}:i\leq n\}:r=1,...,R\}$ are bootstrap samples drawn from the empirical distribution of the data $\{W_{i}: i \leq n\}$. The asymptotic results of this paper hold for the bootstrap provided that $G_{n}^{*}(\theta_{n,h}) \overset{d}{\rightarrow} \Omega_{0}^{\frac{1}{2}}Z^{*}$, where the convergence is conditional on $\{W_i:i\leq n\}$ for almost every sample path, for all sequences $\{(\theta_{n,h},F_{n,h}): n \geq 1\}$ in $\mathcal{F}$.

There are numerous choices for $\varphi$ and $\{\kappa_{n}: n \geq 1\}$. chernozhukov2007estimation and andrews2010inference recommend using $\kappa_{n} = (\ln n)^{\frac{1}{2}}$. Another option is to set $\kappa_{n} = (2\ln \ln n)^{\frac{1}{2}}$, which is used in canay2010inference. Our main results set $\varphi=\varphi^{(1)}$, where

align[align omitted — 162 chars of source]

for each $j \in \mathcal{J}$, and is referred to as the `moment selection $t$-test' because it resembles a $t$-test with deterministic critical value $\kappa_{n}$. The decision reflects the recommendations of andrews2012inference, and is essentially without loss of generality because our results extend to any choice of $\varphi$ that satisfies the assumptions of andrews2010inference. Appendix (ref) discusses how to generalize our results to other suitable choices of $\varphi$.

The advantage of the GMS procedure is that it asymptotically detects the “positive” moments $E_{F_0}(g_{j}(W_{i},\theta_{0}))$ and excludes them from the computation of the critical value, so as to mimic the discontinuity in the asymptotic null distribution of $T_n(\theta).$ This ability of GMS tests to detect such moments is the source of its improvements over the subsampling and plug-in procedures under the null and alternative hypotheses.

Although GMS tests are computationally simple and have desirable asymptotic properties, their performance in finite-samples depends crucially on how well they detect the “positive” moments, so as to omit them from the computation of the critical value. Their use of the sample-analogue estimator of the moments for detecting the positive moments does not implement the information embedded in ((ref)) and implementing this information appropriately can improve the detection accuracy of “positive" moments in finite-samples.

For a given moment selection function $\varphi,$ the CMS procedure implements the information present in ((ref)) through a surgical modification of the GMS procedure. The modification is to replace $\hat{g}_{n}(\theta_0)$ with its constrained empirical likelihood counterpart, where the constraints impose the inequality restrictions ((ref)). Specifically, CMS replaces $\hat{\xi}_{n}(\theta_{0})$ with $\acute{\xi}_{n}(\theta_{0}) =\kappa_{n}^{-1}n^{\frac{1}{2}}\hat{D}_{n}^{-\frac{1}{2}}(\theta_{0})\acute{g}_{n}(\theta_{0})$, where $\acute{g}_{n}(\theta_{0}) = \sum_{i=1}^{n}\acute{p}_{i}g(W_{i},\theta_{0})$ and the probabilities $\acute{p}_{1},...,\acute{p}_{n}$ solve

align[align omitted — 231 chars of source]

and then computes a critical value as described in ((ref)), but replaces $\varphi(\hat{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$ with $\varphi(\acute{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$ in ((ref)). The CMS modification of GMS can easily be applied to all choices of $\varphi$ and $\{\kappa_{n}: n \geq 1\}$ presented in andrews2010inference because it only replaces $\hat{g}_{n}(\theta_{0})$ with $\acute{g}_{n}(\theta_{0}).$ The estimator $\acute{g}_{n,j}(\theta_{0})$ of $E_{F_{0}}\big(g_{j}(W,\theta_{0})\big)>0$ is more accurate than $\hat{g}_{n,j}(\theta_{0})$ because the optimization problem ((ref)), which gives rise to $\acute{g}_{n}(\theta_{0}),$ imposes a correct constraint $E_{F_{0}}\big(g_{j}(W,\theta_{0})\big)\geq0,$ while $\hat{g}_{n}(\theta_{0})$ ignores such information. Thus, when $E_{F_{0}}\big(g_{j}(W,\theta_{0})\big)>0$ (under the null or alternative), the moment selection function based on $\acute{g}_{n,j}$ detects this configuration more reliably than $\varphi_{j}(\hat{\xi}_{n}, \hat{\Omega}_{n}(\theta_{0})),$ and therefore, takes it into account by delivering a critical value that is suitable for the case where this moment inequality is omitted. This feature of CMS leads to it having better finite-sample properties than GMS under the null and alternative hypotheses.

The CMS procedure is not computationally expensive because the empirical likelihood optimization problem ((ref)) has a strictly concave objective function and a convex feasible set that is characterised by affine functions of the choice variables owen2001empirical. This means that the optimization problem ((ref)) has a unique global solution, and it can be computed numerically using standard optimization routines in software such as Matlab, R, or GAUSS. This computational simplicity of the optimization problem ((ref)) is an important feature of the CMS procedure.

remark\normalfont One can `fully constrain' the CMS procedure by using restricted estimators of the correlation matrix. In this case, we evaluate $\varphi(\acute{\xi}_{n}^{FC}(\theta_{0}), \acute{\Omega}_{n}(\theta_{0}))$, where \begin{align*} \acute{\xi}_{n}^{FC}(\theta_{0}) & =\kappa_{n}^{-1}n^{\frac{1}{2}}\acute{D}_{n}^{-\frac{1}{2}}(\theta_{0})\acute{g}_{n}(\theta_{0}),\quad\acute{\Omega}_{n}(\theta_{0}) = \acute{D}_{n}^{-\frac{1}{2}}(\theta_{0})\acute{\Sigma}_{n}(\theta_{0})\acute{D}_{n}^{-\frac{1}{2}}(\theta_{0}), \\ \acute{\Sigma}_{n}(\theta_{0}) & = \sum_{i=1}^{n}\acute{p}_{i}(g(W_{i},\theta_{0})-\acute{g}_{n}(\theta_{0}))(g(W_{i},\theta_{0})-\acute{g}_{n}(\theta_{0}))^{\top},\;and\; \acute{D}_{n}^{-\frac{1}{2}}(\theta_{0})= \operatorname*{diag} \acute{\Sigma}_{n}(\theta_{0}). \end{align*} In our simulations not presented in this paper, we find limited practical difference between the $\varphi(\acute{\xi}_{n}^{FC}(\theta_{0}), \acute{\Omega}_{n}(\theta_{0}))$ and $\varphi(\acute{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$. Consequently, the rest of the paper focuses on $\varphi(\acute{\xi}_{n}(\theta_{0}), \hat{\Omega}_{n}(\theta_{0}))$ because it is simpler to show that there are power advantages over GMS.

Main Results

We start by introducing the assumptions that beget the main results of this paper. They are conditions on the test statistic $S$, the moment selection function $\varphi$, and the parameter space $\mathcal{F}$. The assumptions on $S$ we consider are from andrews2010inference, and are stated as Assumptions 1-7 in Appendix (ref) for ease of exposition. Recall that we set $\varphi=\varphi^{(1)}$ in ((ref)), and the main results we present are based on this choice of moment selection function. It should be noted that this choice of $\varphi$ is without loss of generality as one can employ assumptions identical to those in andrews2010inference on $\varphi$ to deduce the same conclusions, because the CMS procedure does not alter the moment selection function in the GMS procedure. See Appendix (ref) for the details on other choices of $\varphi.$

The first assumption concerns the sequence $\{\kappa_n: n\geq1\}.$

kappaassumption\begin{enumerate*} • $\kappa_{n} \rightarrow +\infty$ as $n\rightarrow +\infty$. • $\kappa_{n}^{-1}n^{\frac{1}{2}} \rightarrow +\infty$ as $n\rightarrow +\infty$. \end{enumerate*}

The conditions in this assumption are not restrictive -- the aforementioned examples of $\{\kappa_{n}: n \geq 1\}$ satisfy them. The `optimal' choice of $\{\kappa_{n}: n \geq 1\}$ is an important question, but the goal of our paper is more modest: to demonstrate how incorporating statistical information can improve finite-sample inference for moment inequalities in a computationally simple way and, for this purpose, our analysis conditions on an arbitrary choice of $\{\kappa_{n}: n \geq 1\}$. For our Monte Carlo experiment (Section (ref)), we set $\kappa_{n}=(\ln n)^{\frac{1}{2}}$ which is the recommended choice in chernozhukov2007estimation and andrews2010inference.

The next assumption we present is the first part in Part (d) of Assumption GEL in andrews2009validity. It is helpful in establishing that $\acute{g}_{n}(\theta_0)$ is a uniformly consistent estimator of the moments under the null hypothesis $H_0:\theta_0\in\Theta_{I}(F_0).$ To introduce this assumption, for each $t\in \mathbb{R}^{J},$ define $g_{i}(t,\theta) = g(W_{i},\theta) - t.$ The vector $t$ is a nuisance parameter that captures the slackness of the moment inequalities. Using the dual formulation of the empirical likelihood problem ((ref)), the amount of slackness is captured by $\acute{t}_n=\operatorname*{arg\,min}_{t \in \mathbb{R}_{+}^{J}}\sup_{\lambda \in \acute{\Lambda}_n(t,\theta)}n^{-1}\sum_{i=1}^{n}\ln\Big(1-\lambda^{\top}g_{i}(t,\theta)\Big),$ where $\acute{\Lambda}_{n}(t,\theta)=\{\lambda\in\mathbb{R}^{J}:\lambda^{\top}g_{i}(t,\theta)\in Q\,\forall i=1,\ldots,n\},$ $Q$ is an open interval of $\mathbb{R}$ containing $0.$ This reformulation of the empirical likelihood problem ((ref)) is feasible because the linear constraint qualification applies to it. The part of Assumption GEL we include in our setup is a regularity condition concerning the uniform asymptotic behavior of $\acute{t},$ and is stated in terms of the following reparametrization of $\mathcal{F}.$

definitionLet $\Gamma$ be defined as the set of all $\gamma=(\gamma_1,\gamma_2,\gamma_3)$ such that for some $(\theta, F)\in\mathcal{F}$ where \begin{enumerate} • $\mathcal{F}$ is defined in Definition (ref). • $\gamma_1=(E_{F}(g_{1}(W_{i},\theta))/ \sigma_{F,1}(\theta),\ldots,E_{F}(g_{J}(W_{i},\theta))/\sigma_{F,J}(\theta)).$$\gamma_2= \big(\theta, \text{vech}_{*}(\Omega(\theta,F))\big),$ where $\text{vech}_{*}(\Omega(\theta,F))$ is the vector of lower off-diagonal elements of $\Omega(\theta,F).$$\gamma_{3} = F.$ \end{enumerate}

andrews2010inference indicate that there is a one-to-one mapping from $\gamma$ to $(\theta,F);$ see Appendix A of their paper for the details. Denote by $\{\gamma_{n,h}: n \geq 1\}\subset \Gamma$ a sequence of parameters in $\Gamma$ such that $n^{1/2}\gamma_{n,h,1} \rightarrow h_{1}\in \mathbb{R}^{J}_{+,\infty}$ and $\gamma_{n,h,2} \rightarrow h_{2} \in \mathbb{R}_{[\pm\infty]}^{q}$ as $n\rightarrow\infty,$ where $q =\dim(\Theta)+ \dim(\text{vech}_{*}(\Omega(\theta,F))).$ The part of Assumption GEL that we include in our setup is given by the following assumption.

assumptiontFor all subsequences $\{w_n\}$ of $\{n\}$ and all sequence $\{\gamma_{w_{n},h}: n \geq 1\}\subset \Gamma$ and corresponding $\{(\theta_{w_{n},h}, F_{w_{n},h}): n \geq 1\}\subset\mathcal{F},$ $$\acute{t}_{w_n}=\operatorname*{arg\,min}_{t \in \mathbb{R}_{+}^{J}}\sup_{\lambda \in \acute{\Lambda}_{w_n}(t,\theta_{w_{n},h})}\frac{1}{n}\sum_{i=1}^{n}\ln\Big(1-\lambda^{\top}g_{i}(t,\theta_{w_{n},h})\Big)$$ exists and satisfies $\sup_{n\geq 1}||\acute{t}_{w_n}||_{\ell^{2}_{J}} \leq K$ with probability approaching 1 as $n\rightarrow+\infty$ for some constant $K<+\infty,$ where $||\cdot||_{\ell^{2}_{J}}$ is the usual Euclidean norm on $\mathbb{R}^{J}$.

We also include an assumption from andrews2010inference for the case in which $\text{Int}(\Theta_{I}(F_0)) \neq \emptyset$ for some data-generating process in the model. It is required to show that when there are no binding moment inequalities, the maximum asymptotic coverage probability is equal to $1.$

assumptionmThere exists $(\theta, F) \in \mathcal{F}$ that satisfies $E_{F}\big(g_{j}(W_{i},\theta)\big)>0$ for all $j\in \mathcal{J}$.

Asymptotic Size Results

We now present the first main result of the paper. It mirrors Theorem 1 of andrews2010inference which concerns the asymptotic size of GMS confidence sets. Denote by $\acute{c}_{n}(\theta_0,1-\alpha)$ the CMS critical value under the nominal level $1-\alpha$ for testing the null hypothesis $H_0:\theta_0\in\Theta_{I}(F_0).$

theoremSuppose $S$ satisfies Assumptions (ref) - (ref), $\varphi=\varphi^{(1)}$ in ((ref)), the sequence $\{\kappa_n:n\geq1\}$ satisfies Part 1 of Assumption K, and $\alpha \in (0,1/2)$. Furthermore, let $\mathcal{F}_{+}=\{(\theta, F) \in \mathcal{F}: \text{Assumption T holds}\},$ and $\mathcal{F}_{++}=\{(\theta, F) \in \mathcal{F}: \text{Assumptions T and M hold}\}.$ Then, the nominal level $(1-\alpha)$ CMS confidence set based on the statistic $T_n(\theta)$ satisfies the following statements: \begin{enumerate} • $\liminf_{n\rightarrow+\infty}\inf_{(\theta, F) \in \mathcal{F}_+} P_{F}\Big(T_{n}(\theta) \leq \acute{c}_{n}(\theta,1-\alpha)\Big)\geq1-\alpha.$$\liminf_{n\rightarrow+\infty}\inf_{(\theta, F) \in \mathcal{F}_+} P_{F}\Big(T_{n}(\theta) \leq \acute{c}_{n}(\theta,1-\alpha)\Big)=1-\alpha$, if in addition $S$ and $\{\kappa_n:n\geq1\}$ satisfy Assumption (ref) and Part 2 of Assumption K, respectively. • $\limsup_{n\rightarrow+\infty}\sup_{(\theta, F) \in \mathcal{F}_{++}}P_{F}\Big(T_{n}(\theta) \leq \acute{c}_{n}(\theta, 1-\alpha)\Big)=1.$ \end{enumerate} \begin{proof} See Appendix (ref). \end{proof}

The first result of Theorem (ref) establishes the uniform validity of CMS confidence sets over the parameter space $\mathcal{F}_{+},$ and the second result of this theorem shows that they are not asymptotically conservative. The third result shows the maximum coverage probability of CMS confidence sets is equal to 1 over the parameter space $\mathcal{F}_{++}.$ The parameter space $\mathcal{F}_+$ is a subset of the one used in Theorem 1 of andrews2010inference, because it imposes Assumption T and Condition 4 in Definition (ref) in addition to the conditions they set for their parameter space.

The proof of Theorem (ref) establishes that CMS and GMS procedures are asymptotically equivalent with uniformity over the parameter space $\mathcal{F}_+.$ The essence of this result is that for every sequence $\{\gamma_{n,h}: n \geq 1\}\subset \Gamma$ such that Assumption T holds, along the corresponding sequence $\{(\theta_{w_{n},h}, F_{w_{n}}): n \geq 1\}\subset\mathcal{F}_+$ we have $\acute{\xi}_{w_{n}}(\theta_{w_{n},h})=\hat{\xi}_{w_{n}}(\theta_{w_{n},h})+o_{p}(1).$ This asymptotic equivalence is a consequence of $\acute{g}_{w_{n}}(\theta_{w_{n},h})-\hat{g}_{w_{n}}(\theta_{w_{n},h})=O_{p}(w_{n}^{-\frac{1}{2}})$ (see Lemma (ref) in Appendix (ref)) and Assumption K on $\kappa_{w_{n}}.$ In particular, these arguments are used after re-writing the expression of $\acute{\xi}_{w_{n}}(\theta_{w_{n},h})$ in terms of $\hat{\xi}_{w_{n}}(\theta_{w_{n},h}),$ as such

equation[equation omitted — 450 chars of source]

to obtain the asymptotic equivalence.

Limiting Local Power Function of CMS Tests

This section employs the setup in Section 8 of andrews2010inference to show the limiting local power function of the CMS tests coincide with their GMS counterparts when the null parameter space is $\mathcal{F}_{+}=\{(\theta, F) \in \mathcal{F}: \text{Assumption T holds}\}.$ For sequences of parameters $\{(\theta_{n,*}, F_{n}): n \geq 1\}$, consider the testing problem

align[align omitted — 160 chars of source]

where $\theta_{n,*}=\theta_{n}+\eta n^{-\frac{1}{2}}(1+o(1))$ for all $n \geq 1$, $(\theta_{n},F_{n}) \in \mathcal{F}_+$ for all $n \geq 1$, and $\eta \in \mathbb{R}^{d}$ where $d=\dim(\theta_{n})<+\infty$. The idea is to study the behavior of the testing procedure along sequences of parameters $\{(\theta_{n,*}, F_{n}): n \geq 1)\}$ that differ locally from a point in the true parameter space $\mathcal{F}_+$ by $O(n^{-\frac{1}{2}})$. The local power function is defined as $P_{F_{n}}\big(T_{n}(\theta_{n,*}) > \acute{c}_{n}(\theta_{n,*},1-\alpha)\big)$, where $P_{F_{n}}(\cdot)$ is the probability measure induced by random sampling from $F_{n}$ for all $n \geq 1$. The objective is to derive an expression for the limiting local power function, $\lim_{n\rightarrow+\infty}P_{F_{n}}\big(T_{n}(\theta_{n,*})> \acute{c}_{n}(\theta_{n,*},1-\alpha)\big),$ and compare it to its GMS counterpart.

To this end, we introduce technical assumptions for deriving the limiting local power function for CMS tests. These assumptions are from Section 8 of andrews2010inference.

lassumptionThe true parameters $\{(\theta_{n}, F_{n}): n \geq 1\}$ satisfy: \begin{enumerate} • $\theta_{n}=\theta_{n,*}-\eta n^{-\frac{1}{2}}(1+o(1))$ for some $\eta \in \mathbb{R}^{d}$, $\theta_{n,*} \rightarrow \theta_{0}$ and $F_{n}\rightarrow F_{0}$ as $n\rightarrow+\infty$, where $(\theta_{0},F_{0})\in\mathcal{F}_+$. • For each $j\in \mathcal{J}$, there exists $h_{1,j} \in \mathbb{R}_{+,\infty}$ such that $n^{\frac{1}{2}}E_{F_{n}}\big(g_{j}(W_{i},\theta_{n})\big)/\sigma_{F_{n},j}(\theta_{n}) \rightarrow h_{1,j}$ as $n\rightarrow+\infty$. • $\sup\big\{E_{F_{n}}|g_{j}(W_{i},\theta_{n,*})/\sigma_{F_{n},j}(\theta_{n,*})|^{2+\delta}: n \geq 1\big\}<+\infty$ for all $j \in \mathcal{J}$ for some $\delta>0.$ \end{enumerate}

The first two parts of Assumptions LA1 show that the sequence of true parameters, $\{\theta_{n}: n \geq 1\}$, is $n^{-\frac{1}{2}}$-local to the sequence of parameters under the null hypothesis, $\{\theta_{n,*}: n\geq 1\}$ and provides the limit of the sequence of normalised moment functions when evaluated at the sequence of true parameters $\{\theta_{n}: n\geq 1\}$. The third part of this assumption is a uniform integrability condition that permits the use of stochastic limit theorems for triangular arrays of row-wise IID random variables. The second assumption is as follows.

lassumption$\Pi(\theta, F) := (\partial/\partial \theta^{\top})[D^{-\frac{1}{2}}(\theta, F)E_{F}(g(W_{i}, \theta))] \in \mathbb{R}^{J\times \dim(\Theta)}$ exists and is a continuous function in a neighbourhood of $(\theta_{0}, F_{0})$.

Both Assumptions LA(ref) and LA(ref) are important for proving the large sample properties of CMS tests under $n^{-\frac{1}{2}}$-local alternatives. Namely, they allow one to mean value expand the normalised moment functions under $H_{0}$ around $\theta=\theta_{n}$ and show that

align*[align* omitted — 149 chars of source]

which can then be used to show that $T_{n}(\theta_{n,*}) \overset{d}{\rightarrow} J_{h_{1}, \eta},$ where $J_{h_{1}, \eta}$ is the distribution function of $S\big(\Omega_{0}^{\frac{1}{2}}Z^{*}+h_{1}+\Pi(\theta_{0},F_{0})\eta, \Omega_{0}\big)$ and $Z^{*}\sim N(0_{J}, I_{J})$ andrews2010inference.

lassumption$\lim_{n\rightarrow+\infty}\kappa_{n}^{-1}n^{\frac{1}{2}}D^{-\frac{1}{2}}(\theta_{n},F_{n})E_{F_{n}}(g(W_{i},\theta_{n}))=\pi_{1} \in \mathbb{R}^{J}_{+,\infty}$.

The last assumption involves the set $C(\varphi)=\{\tilde{\pi}_{1}\in \mathbb{R}^{J}_{[+\infty]}: \forall j \in \mathcal{J},\;\text{either}\;\tilde{\pi}_{1,j}=+\infty\;\text{or}\;\varphi_{j}(\xi, \Omega) \rightarrow \varphi_{j}(\tilde{\pi}_{1}, \Omega_{0})\;\text{as}\; (\xi, \Omega) \rightarrow (\tilde{\pi}_{1}, \Omega_{0})\}.$ Loosely, $C(\varphi)$ is the set of all vectors in $\mathbb{R}^{J}_{[+\infty]}$ for which $\varphi$ is continuous at $(\tilde{\pi}_{1}, \Omega_{0})$. With $\varphi=\varphi^{(1)},$ this set is $C(\varphi^{(1)})=\{\tilde{\pi}_{1}\in \mathbb{R}^{J}_{[+\infty]}:\tilde{\pi}_{1,j}\neq 1,\forall j\in\mathcal{J}\}.$

lassumption\begin{enumerate*} • $\pi_{1} \in C(\varphi^{(1)})$ and • $P_{F}\Big(S\big(\Omega_{0}^{\frac{1}{2}}Z^{*}+\varphi^{(1)}(\pi_{1}, \Omega_{0}), \Omega_{0}\big) \leq x\Big)$ is continuous and strictly increasing at $x=c_{\pi_{1}}(\varphi^{(1)} , 1-\alpha)$, the $1-\alpha$ quantile of the distribution function of $S\big(\Omega_{0}^{\frac{1}{2}}Z^{*}+\varphi^{(1)}(\pi_{1}, \Omega_{0}), \Omega_{0}\big)$. \end{enumerate*}

Assumptions LA3 and LA4 are imposed so that we can use Theorem 2(a) of andrews2010inference to obtain the form of the GMS limiting local power function.

Next, we present the second main result of this paper. This result states that GMS and CMS tests are asymptotically equivalent, to first-order, under $n^{-1/2}$-local alternatives.

theoremSuppose $S$ satisfies Assumptions 1-5, $\varphi=\varphi^{(1)}$ in ((ref)), the sequence $\{\kappa_n:n\geq1\}$ satisfies Assumption K, and that Assumptions LA1 - LA4, hold. Then $\lim_{n\rightarrow+\infty}P_{F_{n}}\big(T_{n}(\theta_{n,*})>\acute{c}_{n}(\theta_{n,*}, 1-\alpha)\big)=1-J_{h_{1},\eta}\big(c_{\pi_{1}}(\varphi, 1-\alpha)\big).$
proofSee Appendix (ref).

The intuition behind Theorem (ref) is essentially the same as Theorem (ref). For a given sequence $\{(\theta_{n,*}, F_{n}): \ n \geq 1\}$ of $n^{-1/2}$-local alternatives, we show $\acute{\xi}_{n}(\theta_{n,*})=\hat{\xi}_{n}(\theta_{n,*})+o_{p}(1)$. This asymptotic equivalence is a consequence of applying $\acute{g}_{n}(\theta_{n,*})-\hat{g}_{n}(\theta_{n,*})=O_{p}(n^{-\frac{1}{2}})$ (see Lemma (ref) in Appendix (ref)) and Assumption K to a decomposition of $\acute{\xi}_{n}(\theta_{n,*})$ identical to ((ref)). Therefore, the pairs $(\acute{\xi}_{n}(\theta_{n,*}), \hat{\Omega}_{n}(\theta_{n,*}))$ and $(\hat{\xi}_{n}(\theta_{n,*}), \hat{\Omega}_{n}(\theta_{n,*}))$ are asymptotically equivalent along sequences of $n^{-1/2}$-local alternatives. As this is the only point of difference between CMS and GMS, Theorem (ref) follows from Theorem 2(a) of andrews2010inference. An important corollary to Theorem (ref) is that CMS inherits the first-order improvements that GMS exhibits over subsampling and plug-in asymptotic critical values (see andrews2010inference).

Local Power Comparison Between CMS and GMS Tests

While Theorem (ref) establishes the equality of the limiting local power functions of CMS and GMS tests under $n^{-1/2}$-local alternatives, this section presents results that characterize sequences of local alternatives under which the power of CMS tests dominate their GMS counterparts for sufficiently large, but finite, samples. First, we must establish when it is meaningful to compare tests along sequences of $n^{-1/2}$-local alternatives. Under the conditions of part 2 of Theorem (ref), for every $r>0,$ there exists $N_{r,0},N_{r,1}\in\mathbb{Z}_{+}$ (depending on $r$) such that

align[align omitted — 422 chars of source]

by the definition of limit superior (with respect to $n$). Then by the triangular inequality,

align*[align* omitted — 244 chars of source]

for all $n\geq N_r=\max\{N_{r,0},N_{r,1}\},$ holds. In words, given an error tolerance $r,$ the tails of the sequences of exact sizes of CMS and GMS tests are within $r$ of $\alpha$ and of each other, when $n\geq N_r.$ Thus, given $r$ (e.g., 0.0001), it is meaningful to compare the rejection probabilities along sequences of local alternatives when $n\geq N_r.$

Let $\mathcal{H}$ denote the set of all sequences $\{(\theta_{n,*},F_{n}): n \geq1\}$ that satisfy Assumption LA1 and LA2. The family we consider for the comparisons is defined as

align[align omitted — 213 chars of source]

For $\{(\theta_{n,*},F_{n}): n \geq 1\} \in \mathcal{H}$, let $\hat{\Upsilon}_{n}(\theta_{n,*}) := \{j\in\mathcal{J}:\varphi_j^{(1)}(\hat{\xi}_{n}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*}))=0\}$ and $\acute{\Upsilon}_{n}(\theta_{n,*}) := \{j\in\mathcal{J}:\varphi_j^{(1)}(\acute{\xi}_{n}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*}))=0\}$ for each $n \geq 1$. We have the following result.

theoremLet $\mathcal{M}$ be as in ((ref)). Suppose that $S$ satisfies Part 1 of Assumption 1, $\varphi=\varphi^{(1)}$ in ((ref)), and the sequence $\{\kappa_n:n\geq1\}$ satisfies Assumption K. For every $\{(\theta_{n,*}, F_{n}): n \geq 1\} \in \mathcal{M},$ there exists $N(\theta_{n,*}, F_{n}) \in \mathbb{Z}_{+}$ such that \begin{align} P_{F_{n}}\Big(T_{n}(\theta_{n})>\hat{c}_{n}(\theta_{n},1-\alpha)\Big) \leq P_{F_{n}}\Big(T_{n}(\theta_{n}) > \acute{c}_{n}(\theta_{n},1-\alpha)\Big)\quad \forall n \geq N(\theta_{n,*}, F_{n}). \end{align} If in addition $S$ satisfies part 1 of Assumption 2 and Part 2 of Assumption 5, and the event $$\bigg\{\acute{\Upsilon}_{n}(\theta_{n,*})\subsetneq \hat{\Upsilon}_{n}(\theta_{n,*})\bigg\} \bigcap \bigg\{\hat{c}_{n}(\theta_{n,*},1-\alpha)>0\bigg\} \bigcap \bigg\{\acute{c}_{n}(\theta_{n,*},1-\alpha)<T_{n}(\theta_{n,*}) \leq \hat{c}_{n}(\theta_{n,*},1-\alpha)\bigg\}$$ has positive probability for each $n \geq N(\theta_{n,*},F_{n})$, then the weak inequalities in ((ref)) are strict.
proofSee Appendix (ref).

Theorem (ref) states the rejection probabilities of CMS tests are no less than their GMS counterparts in large enough, but finite, sample sizes, under local alternatives in $\mathcal{M}.$ It also provides a sufficient condition for the ordering to hold strictly. Thus, for each sequence of local alternatives in $\mathcal{M}$ and small $r>0,$ the local power of a CMS test is larger than its GMS counterpart when $n \geq \max\{N(\theta_{n,*}, F_{n}),N_r\},$ where $N_r=\max\{N_{r,0},N_{r,1}\}$ and $N_{r,0}$ and $N_{r,1}$ defined in ((ref)) and ((ref)), respectively.

The key message from Theorem (ref) is that a comparison of GMS and CMS tests based on first-order asymptotics can be misleading, as it does not reflect the finite-sample situation for certain local alternatives. The result of Theorem (ref) is similar to Corollary 6.1 Lok-Tabri-inpress; however, it is important to note that their result is specific to moment inequalities arising from restricted stochastic dominance orderings. Consequently, Theorem (ref) provides a nontrivial extension of their result to the moment inequality model with finitely many inequalities and arbitrary moment functions, when the off-diagonal elements of $\Omega(\theta_{n,*},F_{n})$ are non-negative for each $n.$

At the heart of this result is the marriage of the non-negative correlational structure on $\Omega(\theta_{n,*},F_{n})$ and constrained empirical likelihood estimation. This marriage begets $\hat{g}_{n,j}(\theta_{n,*}) \leq \acute{g}_{n,j}(\theta_{n,*})$ with probability approaching $1,$ for all sequences in $\mathcal{M}$ (see Lemma (ref)). This ordering of the estimators implies that $\varphi^{(1)}_{j}(\acute{\xi}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*})) \geq \varphi^{(1)}_{j}(\hat{\xi}_{n}(\theta_{n,*}),\hat{\Omega}_{n}(\theta_{n,*}))$ holds with probability approaching 1, for all sequences in $\mathcal{M}$ (see Lemma (ref)). It is this ordering of the moment selection functions under such sequences that gives rise to the result of Theorem (ref).

While Theorem (ref) indicates that the local powers of the GMS and CMS tests can be ordered under a class of local alternatives $\mathcal{M}$ for large enough $n,$ it does not specify the extent of the discrepancy in the local powers. It is quite difficult to determine the extent of this discrepancy analytically. However, Section (ref) presents Monte Carlo evidence that the discrepancy that Theorem (ref) implies can be very large for local alternative sequences which have some non-violated inequalities (SNVIs). That is, sequences $\{(\theta_{n,*},F_{n}): n \geq1\}$ in $\mathcal{M}$ where there exists $j\in\mathcal{J}$ such that $E_{F_{n}}\big(g_{j}(W_i,\theta_{n,*})\big)>0\;\forall n$ and $\lim_{n\rightarrow+\infty}E_{F_{n}}\big(g_{j}(W_i,\theta_{n,*})\big)=0$.

Power Against Distant Alternatives

This section shows CMS tests are consistent against distant alternatives. Distant alternatives include fixed alternatives and alternatives that differ from the null by greater than $O(n^{-\frac{1}{2}})$. The next assumption is useful for deducing this result, and it is the same one introduced by andrews2010inference in Section 9 of their paper.

dassumptionLet $g_{n,j}^{*}=E_{F_{n}}(g_{j}(W_{i},\theta_{n,*}))/\sigma_{F_{n},j}(\theta_{n,*})$ for each $j \in \mathcal{J},$ and $\upsilon_{n}=\max_{j \in \mathcal{J}}\{-g_{n,j}^{*}\}$. \begin{enumerate*} • $n^{\frac{1}{2}}\upsilon_{n}\rightarrow+\infty$ as $n\rightarrow+\infty$$\Omega(\theta_{n,*},F_{n})\rightarrow \Omega_{1}$, $\Omega_{1} \in \varPsi_{2}$. \end{enumerate*}

The key part of this assumption is the first part, which indicates that there exists $j \in \mathcal{J}$ such that $g_{n,j}^{*}<0$ and that the violation of the non-negativity constraint is not $O(n^{-\frac{1}{2}})$. This condition differs from the setup with $n^{-\frac{1}{2}}$-local alternatives, where the sequences of alternatives $\{(\theta_{n,*},F_{n}): n \geq 1\}$ are within a $n^{-\frac{1}{2}}$-neighbourhood of $\mathcal{F}_+$. We have the following result.

theoremSuppose $S$ satisfies Assumptions 1,3,4 and 6, $\varphi=\varphi^{(1)}$ in ((ref)), and the sequence $\{\kappa_n:n\geq1\}$ satisfies Assumption K. Then $\lim_{n\rightarrow+\infty}P_{F_{n}}\Big(T_{n}(\theta_{n,*})>\acute{c}_{n}(\theta_{n,*},1-\alpha)\Big)=1.$
proofSee Appendix (ref).

Simulation Results

This section studies the finite-sample performance of the CMS procedure and compares it to the GMS procedure using a simulation experiment based on the designs in andrews2012inference. The study uses the test statistics $S_1$ in ((ref)) and $S_{2A}$ in ((ref)), the recommended moment selection function $\varphi=\varphi^{(1)}$ in ((ref)), and the recommended localisation parameter $\kappa_n=(\ln n)^{1/2}.$ The nominal level is set to $\alpha=0.05,$ and we considered sample sizes $n=50,100$, and $250$. We also report simulation results for (i) the RSW procedure that use $S_1$ and $S_{2A}$, and (ii) the recommended RMS testing procedure, which combines $S=S_{2A}$, $\varphi=\varphi^{(1)}$, and $\kappa$-auto (a data-driven choice of $\kappa_n$), as additional benchmarks in studying the finite-sample performance of CMS; see Appendices (ref) and (ref), respectively, for further details on these testing procedures. Only bootstrap versions of the tests were implemented, with 10000 bootstrap samples per Monte Carlo replication. The computations were implemented using R.

For a given $\theta,$ the null hypothesis is $H_0:\theta\in\Theta(F_0).$ The experimental design in andrews2012inference,andrews2012supplement is a general formulation of that testing problem that does not require the specification of a particular form for the moment functions $\{g_{j}(\cdot,\theta): j=1,...,J\}.$ They note that the finite-sample properties of tests of $H_0$ depend on the moment functions only through (i) the vector $\mu=[E_{F_{0}}\big(g_{1}(W_i,\theta)\big),\ldots,E_{F_{0}}\big(g_{J}(W_i,\theta)\big)],$ (ii) the correlation matrix $\Omega=\text{Corr}\left(g_{1}(W_i,\theta),\ldots,g_{J}(W_i,\theta)\right),$ and (iii) the distribution of the mean zero, variance $I_J$ random vector $Z^{\dagger}=[Z_1^{\dagger},\ldots,Z_J^{\dagger}]$ where $$ Z_j^{\dagger}=\text{Var}_{F_0}^{-1/2}\left(g_{j}(W_i,\theta)\right)\left(g_{j}(W_i,\theta)-E_{F_{0}}\big(g_{j}(W_i,\theta)\big)\right),\quad j=1,\ldots, J. $$ We consider the case $Z^{\dagger}\sim N(0_J,I_J)$ and three correlation matrices, $\Omega_{\text{Neg}},$ $\Omega_{\text{Zero}},$ and $\Omega_{\text{Pos}},$ which exhibit negative, zero, and positive correlations.

The assertion of the null hypothesis in this general formulation is $H_0:\mu_j\geq0\;\forall j=1,\ldots,J$. For comparisons under the null hypothesis, we follow andrews2012inference by comparing the tests' maximum null rejection probabilities (MNRPs). The MNRPs are computed over the mean vectors $\mu$ in the null parameter space given the correlation matrix $\Omega\in\{\Omega_{\text{Neg}},\Omega_{\text{Zero}},\Omega_{\text{Pos}}\}$ and under the assumption of normally distributed moment inequalities. Based on simulation evidence, they conjecture that the MNRPs occur for mean vectors $\mu$ whose elements are $0$'s and $+\infty$'s. Thus, given a nominal level $\alpha,$ they compute MNRP results over the set of mean vectors $\mu$ which have that form. The results we report are for $J=2,4$ and $10.$ The matrix $\Omega_{\text{Zero}}$ equals the $J$-dimensional identity matrix. The matrices $\Omega_{\text{Neg}}$ and $\Omega_{\text{Pos}}$ are Toeplitz matrices with correlations given by the following: for $J=2:$ $\rho=-.9$ for $\Omega_{\text{Neg}}$ and $\rho=.5$ for $\Omega_{\text{Pos}};$ for $J=4:$ $\rho=(-.9,.7,-.5)$ for $\Omega_{\text{Neg}}$ and $\rho=(.9,.7,.5)$ for $\Omega_{\text{Pos}};$ for $J=10:$ $\rho=(-.9,.8,-.7,.6,-.5,.4,-.3,.2,-.1)$ for $\Omega_{\text{Neg}}$ and $\rho=(.9,.8,.7,.6,.5,\ldots,.5)$ for $\Omega_{\text{Pos}}$. As in andrews2012inference, the simulation study treats the correlation matrices as unknown in the implementation of all of the tests.

For power comparisons, we also follow andrews2012inference,andrews2012supplement. They compare the power of different tests by comparing their empirical power for a chosen set of alternative parameter vectors $\mu\in\mathbb{R}^{J}$ for a given correlation matrix $\Omega\in\{\Omega_{\text{Neg}},\Omega_{\text{Zero}},\Omega_{\text{Pos}}\}$. The sets of $\mu$ vectors in the alternative are similar to the ones described in andrews2012inference,andrews2012supplement. We adjust those sets so as to compare the local power properties of the testing procedures. The adjustment is as follows. For each $J\in\{2,4,10\}$ the set of $\mu$ vectors are given by $\mathcal{M}_{J,n}(\Omega) = \left\{\mu/\sqrt{n}: \mu\in\mathcal{M}_J(\Omega)\right\},$ where the set $\mathcal{M}_J(\Omega)$ of $\mu$ vectors is described in Section 7.1 of andrews2012supplement. The $\mu$ vectors in $\mathcal{M}_{J,n}(\Omega)$ are scaled versions of those in $\mathcal{M}_J(\Omega),$ where the scaling is by $n^{-1/2}$ to create the $n^{-1/2}$-local alternatives. There are $7,24,$ and $40$ elements in $\mathcal{M}_{J,n}(\Omega)$ for $J=2,4$ and $10,$ respectively. We omit their description for brevity.

As the MNRPs of the tests can differ in finite-samples, the simulation results on power comparisons are based on a MNRP correction that is similar to the one employed by andrews2012inference. For each test statistic $S$, the MNRP correction of the CMS, GMS and RMS procedures is to add a constant based on the true matrix $\Omega$ to their corresponding critical values, so that their resulting MNRPs match that of the RSW testing procedure with nominal level $\alpha=0.05;$ see Section (ref) in the appendix for the details. The simulation studies in andrews2012inference and romano2014practical compare tests under the alternative using average MNRP-corrected power, where the average is computed over alternative $\mu$ vectors in $\mathcal{M}_J(\Omega).$ We report simulation results graphically using boxplots of the MNRP-corrected local powers over sets of $\mu$ vectors $\mathcal{M}_{J,n}(\Omega)$ for the 54 different combinations of $(J,\Omega,S,n)$ for each of the CMS, GMS and RSW procedures, and 9 different combination of $(J,\Omega,S_{2A},n)$ for the recommended RMS test. Additionally, we report average MNRP-corrected local powers of the different tests across the aforementioned configurations using the symbol $\oplus$ in these plots.

While the average MNRP-corrected power is a useful criterion for comparing tests across $\mu$ vectors in a given set $\mathcal{M}_{J,n}(\Omega)$, it does not convey the whole picture of the tests' performance over elements in $\mathcal{M}_{J,n}(\Omega).$ Reporting boxplots, as we do, reveals the variation in powers of the tests across elements in $\mathcal{M}_{J,n}(\Omega);$ thus, presenting a broader and more extensive approach to comparing the tests under the alternative. These plots are especially useful for detecting differences in the performances of tests when the averages of their MNRP-corrected powers are close, but exhibit different distributional variations in MNRP-corrected power across $\mu$ vectors in $\mathcal{M}_{J,n}(\Omega).$

Maximum Null Rejection Probabilities

As in andrews2010inference, andrews2012inference, and romano2014practical, empirical MNRPs are simulated as the maximum rejection probability over all $\mu$ vectors whose components are $0$ and $+\infty,$ with at least one component equal to zero. Table (ref) reports the MNRPs for tests. Each experiment used 10000 Monte Carlo replications when $J \in \{2,4\}$ and 2500 when $J=10$.

table[table omitted — 3,708 chars of source]

Overall, the procedures achieve a satisfactory performance for all cases considered. The RMS and RSW tests perform the best, as their MNRPs are closest to the 5% nominal level across all of the cases considered. For the RSW procedure: the MNRPs fall within the ranges [.043,.056] and [.043,.055] when using $S_{2A}$ and $S_1$ test statistics, respectively. For the RMS test: the MNRPs fall into the range [.042,.053]. The CMS tests over-reject the null slightly: the MNRPs fall within the ranges [.049,.067] and [.049,.061] when using $S_{2A}$ and $S_1$ test statistics, respectively. The tables also show CMS tests have better MNRPs than their GMS versions as the latter tend to over-reject more: the GMS MNRPs fall within the ranges [.049,.083] and [.049,.078] for $S_{2A}$ and $S_1$, respectively. The largest MNRPs arise in the configurations where $\Omega=\Omega_{\text{Neg}},$ and these MNRPs increase with larger $J$, for the CMS, GMS and RSW tests. However, the MNRPs of all of these tests do get closer to the 5% nominal level with larger sample sizes, across all configurations, and for CMS tests, this numerical result is a consequence of Theorem (ref).

While we don't have a theoretical result on improved size control of CMS tests over their GMS versions, Table (ref) provides simulation-based evidence of such an improvement. Hence, these results point to the potential benefit of implementing the information ((ref)), as we do, in two step testing procedures, under the null. The next section presents simulation results on MNRP-corrected power of these tests, under local alternatives, and illustrates the result of Theorem (ref).

Local Power

Figures (ref) and (ref) below report boxplots of the MNRP-corrected powers of the tests under $S_{2A}$ and $S_1$, respectively. The results can be summarised as follows. For each test statistic, the MNRP-corrected power values of the tests are generally distributed in a similar way in configurations where $\Omega=\Omega_{\text{Neg}},$ and all of the tests have comparable average powers in those configurations. By contrast, in configurations where $\Omega=\Omega_{\text{Zero}}$, for each test statistic, the boxplots show the RSW tests' MNRP-corrected power values tend to be (i) more dispersed (as shown by the lengths of their boxes), (ii) have a wider overall range, and (iii) have lower average power in comparison to the remaining tests, which all behave similarly as can be seen by their boxplots. For example, the average power of the RSW test when $S=S_{2A}$, $J=10$, and $n=250$ is approximately equal to 0.57, while the averages of the remaining procedures in that scenario are all approximately equal to 0.66, which is a large difference.

More noticeable differences in the tests' performance arise in configurations where $\Omega=\Omega_{\text{Pos}}.$ For each test statistic, there is evidence for the following ranking in terms of average MNRP-corrected power, uniformly in $J$ and $n$: CMS in first place, GMS in second place, and RSW in third place, with the RMS test tied in first place with the CMS-$S_{2A}$ test. The boxplots also show:

itemize• The MNRP-corrected power values for the RMS and CMS-$S_{2A}$ tests are generally distributed in a similar way for each $n$ and $J$, except when $J=2$ the CMS-$S_{2A}$ power values are slightly more dispersed (as shown by the length of the boxes) than their RMS counterparts for each $n$. • For each $S$, the MNRP-corrected power values of CMS tests are markedly less dispersed and have smaller overall ranges than their GMS and RSW counterparts. • The difference among the CMS, GMS and RSW tests in these configurations with $S=S_1$ can be strikingly large in terms of average power; for example, with $J=4$ and $n=250,$ the average powers of CMS, GMS and RSW tests are approximately equal to 0.75, 0.65, and 0.60, respectively. By contrast, with $S=S_{2A}$ the difference among these tests is less pronounced, which is on account of using a more effective test statistic. For example, in the aforementioned configuration, the average powers of CMS, GMS and RSW tests are approximately equal to 0.76, 0.74, and 0.73, respectively. However, this less pronounced difference in average powers does not mean that these procedures behave similarly, as evidenced by the radically different boxplots of the tests' power values.
landscape\begin{figure} \caption{Boxplots of MNRP-corrected powers of $S_{2A}$-based tests. For each configuration, the symbol $\oplus$ marks the location of the average MNRP-corrected power of a test. } \end{figure}
landscape\begin{figure} \caption{Boxplots of MNRP-corrected powers of $S_{1}$-based tests. For each configuration, the symbol $\oplus$ marks the location of the average MNRP-corrected power of a test. } \end{figure}

The result of Theorem (ref) implies that the average power of CMS and GMS tests should get closer together with larger sample sizes. The simulations reflect this implication across all configurations, but indicate that it happens slowly when $\Omega=\Omega_{\text{Pos}}.$ Consequently, there is simulation-based evidence that shows the implementation of the information ((ref)), as we do with CMS, may not improve the local power of GMS tests for configurations in which $\Omega\neq\Omega_{\text{Pos}}.$ The reason is that the boxplots for MNRP-corrected power values of CMS and GMS tests are generally quite similar in those configurations. By contrast, the result of Theorem (ref) points to such an improvement in local power for configurations in which $\Omega=\Omega_{\text{Pos}}$, and this result is reflected in the simulations as described above.

Table (ref) reports the average MNRP-corrected powers of the tests when $\Omega=\Omega_{\text{Pos}}$ and we use these results to further contextualize the local power improvement associated with CMS tests over GMS and RSW tests. We benchmark our analysis to RMS because simulation evidence in andrews2012inference suggests that it is superior in terms of asymptotic average power and is therefore the recommended test. The CMS-$S_{2A}$ and RMS tests are neck and neck as their average powers are essentially identical and achieve the highest average powers in all of those scenarios, with the CMS-$S_1$ test having slightly lower average powers than those tests. For a given $S$, the RSW tests are the worst performing, as they achieve the lowest average powers in each corresponding scenario, and the difference between them and the RMS test can be quite large. For example, when $J=10$ and $n=250$, the difference between RSW-$S_1$ and RMS is 0.252, and with RSW-$S_{2A}$ it is 0.055 which is a much smaller on account of using a more effective test statistic. The CMS-$S_1$ test dominates the GMS-$S_{1}$ test in each of those scenarios, where the difference can be as large as 10 percentage points -- see the scenarios with $J=10$. Consequently, the importance of incorporating the statistical information from the constraints, as we do with CMS, picks up the difference in average powers between the RMS and GMS test when $S=S_{2A}$ and most of the difference when $S=S_1,$ in each of the those scenarios.

table[table omitted — 1,144 chars of source]

While the focus above has been on average power, for individual $\mu$ vectors the power differences can be massive with $\Omega=\Omega_{\text{Pos}}$. Consider, for example, the element $\mu/\sqrt{n}\in\mathcal{M}_{4,n}(\Omega_{\text{Pos}})$ with $\mu=(-2.4705,1,1,1)^{\intercal}$. This mean vector is an example of an SNVI local alternative. Table (ref) reports the MNRP-corrected power estimates for the tests under this local alternative for $n=50,100,250$. The estimates indicate that:

table[table omitted — 818 chars of source]
figure[figure omitted — 294 chars of source]
itemize• There can be extremely large power improvements associated with CMS relative to RSW and GMS when $S=S_{1}$. Indeed, the improvement in power of CMS over GMS is approximately 36 percentage points and 40 percentages points over RSW. • The improvements persist with $S=S_{2A}$, but are not as large. The AQLR statistic results in CMS experiencing a six percentage point improvement over GMS and eight percentages point improvement over RSW. In absolute terms, all procedures experience higher local power with $S=S_{2A}$. • The MNRP-corrected powers of CMS-$S_{2A}$ are comparable to their RMS counterparts.

To gain a deeper insight into the behavior of the tests under this local alternative, Figure (ref) reports the empirical distribution functions (ECDFs) of the MNRP-corrected critical values for $n=250.$ The focus on this sample size is without loss of generality as similar graphs of the critical values' ECDFs arise in all of the other values of $n$ we considered. For either test statistic, the ECDFs in Figure (ref) show strong evidence of a first-order stochastic dominance ranking among the critical values of the CMS, GMS, and RSW, tests. Specifically, for both types of test statistics, there is evidence for the ordering $\acute{c}_{n} \leq \hat{c}_{n} \leq \check{c}_{n}$, where $\check{c}_{n}$ denotes the RSW critical value. By contrast, the ECDF of the recommended RMS test crosses that of CMS with $S=S_{2A},$ which means that there isn't evidence of a clear ordering of their critical values. Overall, the differences between the ECDFs is quite striking and indicates that there is a big difference in the behavior of the tests even in moderately large sample sizes. The stochastic ordering of the CMS and GMS critical values is a reflection of Theorem (ref) and provides evidence for local power improvements under SNVI local alternatives which have positively correlated moment functions. Finally, we discuss the behavior of the RSW procedure. The RSW procedure rejects on the event $\{M_{n}(\beta) \nsubseteq \mathbb{R}_{+}^{J}\}\bigcap \{S>\check{c}_{n}\}$, where $M_{n}(\beta)$ is a lower confidence rectangle that is used to detect “positive” moments in the first step of their two-step procedure (see Appendix (ref)). Across the two test statistics, our simulations indicate that (i) the event $\{M_{n}(\beta) \nsubseteq \mathbb{R}_{+}^{J}\}$ occurs with empirical probability close to $1$, and (ii) in $9465$ times out of $10000$ Monte Carlo replications, their critical value $\check{c}_{n}$ corresponds to the case where none of the moment inequalities have been omitted from its calculation. These findings show the RSW procedure fails to reliably detect the “positive” moments in $\mu/\sqrt{n}$ in most of the $10000$ Monte Carlo replications, resulting in it having low empirical power.

Conclusion

This paper has proposed a surgical modification of the generalized moment selection (GMS) procedure put forward by andrews2010inference that improves its performance, called constrained moment selection (CMS). The basic idea of the CMS procedure is to use empirical likelihood to incorporate the information embedded in the moment inequality constraints into the moment selection step of the GMS procedure. Our analyses highlights the importance of using this information to more reliably detect the binding moments, which is the source of the improvement of CMS over GMS tests.

There are a number of directions for future research. Although we focus on modifying GMS tests, the intuition of incorporating the information embedded in the identified set transcends this choice and we conjecture that similar finite-sample benefits would arise in a similar modification to the two-step procedure of romano2014practical. There is also an emerging literature that focuses on testing with `many' moments, where the number of inequalities grow exponentially with the sample size (e.g., chernozhukov2019inference, and bai2019practical). Extending the empirical likelihood modification to such testing procedures may improve their performance, but different theoretical tools must be employed to account for the increasing number of constraints. Finally, our paper is related to the semi-infinite programming empirical likelihood procedure proposed by Lok-Tabri-inpress for two-step bootstrap tests of stochastic dominance, where the continuum of unconditional moment inequalities is akin to inference for conditional moment inequalities. Their results are limited to restricted stochastic dominance tests and it would be interesting to extend the insights from this paper to the general conditional moment inequality models of andrews2013inference, andrews2017inference.

Acknowledgements

We are grateful to Jonathan Roth for providing valuable comments. We are also appreciative of feedback from participants at the Graduate Student Workshop in Econometrics, Harvard University. The computations in this paper were run on the FASRC Cannon cluster supported by the FAS Division of Science Research Computing Group at Harvard University. All errors are our own.